cs.CL · 2026-06-24 · No. 33

Computation and Language, 2026-06-24.

23 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

23 entries
  1. 01

    L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models

    Hariom Ingle, Ronit Ghode, Ishwari Gondkar, Jidnyasa Harad, Raviraj Joshi

    cs.CL · cs.LG

    Part-of-Speech (POS) tagging is a foundational NLP task underpinning machine translation, information extraction, and syntactic parsing. Despite Marathi being spoken by over 83 million people and ranking among the top twenty most spoken languages worldwide, it remains severely under-resourced in annotated corpora and standardised evaluation benchmarks. Marathi presents unique challenges for computational modelling owing to its rich...

    arxiv.org/abs/2606.24825 · PDF

  2. 02

    Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce

    Filippos Ventirozos, Matthew Shardlow

    cs.CL · cs.AI

    Commercial NLP treats the shopping chatbot as a recommender or a conversion tool: its job is to match a user to a catalogue entry and close a sale. We argue that the arrival of agent-native micro-payment rails (e.g., x402, AP2) changes what is scarce. When the buyer is an autonomous agent that can investigate exhaustively, the bottleneck is no longer matching products but acquiring trustworthy, decision-relevant information about them. We...

    arxiv.org/abs/2606.24783 · PDF

  3. 03

    Task Decomposition for Efficient Annotation

    Nupoor Gandhi, Emma Strubell

    cs.CL · cs.AI · cs.HC

    High-quality annotations of structured representations are expensive to collect over large corpora. Manual annotation of structure is laborious, and model-based annotation, although cheaper to generate, requires expensive validation and potentially significant supervision to ensure that the annotation quality is strong enough to be useful downstream. In traditional annotation workflows, annotation of each complete example is performed...

    arxiv.org/abs/2606.24734 · PDF

  4. 04

    AI-PAVE-Br: Leveraging Large Language Models for Enhanced Product Attribute Value Extraction through a Golden Set Approach

    Murilo Gazzola, Hugo Gobato Souto, Samuel Silva, Júlia Schubert Peixoto, Felipe Siqueira, André Luis Pedroso de...

    cs.CL · cs.AI · cs.LG · cs.PF

    The explosive growth and complexity of product data within the dynamic Brazilian e-commerce landscape demand robust and specialized methods for structured information extraction. Traditional approaches to Product Attribute Value Extraction (PAVE) often struggle with the linguistic nuances and sheer diversity of product descriptions in Portuguese. To address this critical gap, this paper introduces two major contributions. First, we present...

    arxiv.org/abs/2606.24655 · PDF

  5. 05

    Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity

    Yuanhe Zhao, Tianyu Zhang, Huafei Xing, Derek F. Wong, Jianbin Li, Tao Fang

    cs.CL · cs.AI

    Retrieval-Augmented Generation enhances large language models by incorporating external knowledge, but deploying it in sensitive scenarios risks privacy leakage via malicious prompts. To address this, we propose a multi-agent framework that sanitizes retrieved content through semantic rewriting. By employing three specialized agents for privacy extraction, semantic analysis, and reconstruction, our approach collaboratively removes sensitive...

    arxiv.org/abs/2606.24623 · PDF

  6. 06

    Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

    Arda Eren, Micheal Cheung, Youqian Zhang, Grace Ngai, Eugene Yujun Fu

    cs.CL · cs.AI

    Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource languages. In low-resource settings such as Turkish, detection is especially difficult, as annotated data is scarce and technological defenses remain limited. This research investigates how large language models (LLMs) can support scam detection in Turkish by introducing the first public...

    arxiv.org/abs/2606.24523 · PDF

  7. 07

    The African Language Tax: Quantifying the Cost, Latency, and Context Penalty of Tokenizing African Languages in Frontier LLMs

    Olaoye Anthony Somide

    cs.CL · cs.AI

    Commercial large language models bill, scale latency, and budget context per token. Yet tokenizers assign more subword tokens to the same meaning in some languages than in others, so speakers of languages with high token-fertility pay a structural penalty before a model is ever invoked. This penalty is documented for multilingual settings in general, but it has not been measured systematically for African languages at the level of enterprise...

    arxiv.org/abs/2606.24460 · PDF

  8. 08

    On the Stability of Prompt Ranking in Large Language Model Evaluation

    Shaoshuai Du, Penghao Liang, Yixian Shen, Chuanqi Shi, Hang Zhang, Lun Wang

    cs.CL · cs.AI

    Prompt-based interaction has become a dominant paradigm for using large language models (LLMs), where multiple candidate prompts are evaluated and the top-ranked one is selected for downstream use. This workflow implicitly assumes that prompt rankings are stable under minor variations in evaluation conditions. In this paper, we systematically study prompt ranking stability under common sources of variability, including random seeds and...

    arxiv.org/abs/2606.24381 · PDF

  9. 09

    CALIBER: Calibrating Confidence Before and After Reasoning in Language Models

    Conor Finlay, Joshua Kurien, Saurabh Dash, Marzieh Fadaee, Beyza Ermis

    cs.CL · cs.AI

    Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success. Existing methods typically elicit confidence only once: either before thinking or after answering. We argue that confidence in reasoning models is state-dependent: before thinking, confidence should estimate the chance of the model correctly solving the prompt, while after thinking it should predict...

    arxiv.org/abs/2606.24281 · PDF

  10. 10

    Pigeonholing: Bad prompts hurt models to collapse and make mistakes

    Hyunji Nam, Keertana Chidambaram, Dorottya Demszky, Natasha Jaques

    cs.CL · cs.AI

    While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing." **Unintentionally bad** contexts can happen without malicious jailbreaking intents: For example, a user asks the model to justify an incorrect math theorem or fails to correct the model's buggy code. Specifically, we investigate ``pigeonholing" in...

    arxiv.org/abs/2606.24267 · PDF

  11. 11

    SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization

    Noor Islam S. Mohammad, Ulug Bayazit

    cs.CL · cs.AI

    Fine-tuned encoders deployed across heterogeneous NLP tasks face three compounding problems: mismatched inductive biases, class-imbalance corruption of feature statistics, and no mechanism to condition attention on external lexical knowledge. We introduce \textbf{\surgellm}, a unified transformer framework that addresses each with a dedicated lightweight module: a \emph{surgical feature gate} (learned per-dimension sigmoid over curated...

    arxiv.org/abs/2606.24259 · PDF

  12. 12

    MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval

    Junhyeok Lee, Han Jang, Hyeonjin Goh, Kyu Sung Choi

    cs.CL · cs.AI · cs.IR

    Retrieval-augmented generation (RAG) in clinical settings increasingly requires multilingual retrieval against predominantly English evidence corpora. Multilingual medical retrieval demands three capabilities: cross-lingual alignment, concept discrimination, and evidence retrieval. However, existing benchmarks evaluate these only in isolation, leaving the interaction between biomedical expertise and multilingual coverage unmeasured. We...

    arxiv.org/abs/2606.24200 · PDF

  13. 13

    A Pāninian Foundation for Indic Language Processing

    Ritwik Banerjee, Lav R. Varshney

    cs.CL · cs.AI · cs.LG

    More than a billion people communicate in Indic languages, yet the natural language processing infrastructure serving them remains fragmented and underdeveloped. The cause is structural: the field organizes its tools and benchmarks around individual languages or small subsets of genealogical language families, building separate analyzers, parsers, and datasets for each language and starting over for the next. This overlooks a deep regularity....

    arxiv.org/abs/2606.24172 · PDF

  14. 14

    BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks

    Jin Huang, Yutong Xie, Wanli Song, Xingjian Zhang, Walter Yuan, Matthew O. Jackson, Qiaozhu Mei

    cs.CL · cs.LG

    Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show promise in individual tasks such as survey response prediction and human-subject experiment simulation, there remains no systematic understanding of how well they perform across diverse behavioral science tasks, contexts, and populations. We introduce BehaviorBench, a comprehensive benchmark that...

    arxiv.org/abs/2606.24162 · PDF

  15. 15

    Metis: Bridging Text and Code Memory for Self-Evolving Agents

    Zijie Dai, Siuhin He, Hui Li, Qihui Zhou, Jiajun Li, Mingcong Song, Guoping Long, Hongjie Si, Xin Yao, Lin Zhang,...

    cs.CL · cs.AI

    Self-evolving agents improve over time by distilling experience from past executions and reusing it in future tasks. Existing systems represent such experience either as natural-language text injected into the agent context or as code exposed as callable tools. However, the choice between these representations is typically made at design time rather than derived from the characteristics of the experience itself, leaving the trade-offs between...

    arxiv.org/abs/2606.24151 · PDF

  16. 16

    PORTER: Language-Grounded Event Representations for Portable Structured EHR Foundation Models

    Lin Lawrence Guo, Adam Paul Yan, Emily Vettese, Lillian Sung

    cs.CL · cs.LG

    Most electronic health record (EHR) foundation models encode clinical events as discrete event tokens from a fixed vocabulary and therefore cannot directly represent events containing unseen concepts or new combinations of concepts and attributes such as numeric values. This limits transfer across institutions and even across deployment pipelines within the same institution. We introduce PORTER, a language-grounded structured EHR foundation...

    arxiv.org/abs/2606.24102 · PDF

  17. 17

    Predicting Poets' Origins from Verse: A Computational Analysis of Regional Linguistic Fingerprints in the Complete Tang Poems

    Chi-Sheng Chen, Hung-Yun Liu

    cs.CL · cs.AI

    We ask whether the geographic origin of Tang-dynasty poets leaves a detectable linguistic trace in their work. Aggregating every poem attributed to each author in the Complete Tang Poems (Quan Tang Shi) and linking poets to their administrative circuit of origin via the China Biographical Database (CBDB), we build a poet-level corpus of 357 poets across the ten Tang circuits and frame origin prediction as multi-class classification. Using...

    arxiv.org/abs/2606.24093 · PDF

  18. 18

    CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

    Morayo Danielle Adeyemi, Ryan A. Rossi, Franck Dernoncourt

    cs.CL · cs.AI · cs.LG

    "Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves anything depends on which channel (the user's prompt or the model's response) is being compressed. We present Cavewoman, a two-channel evaluation protocol that scores every generation on task accuracy, realized per-item cost, and reference-text agreement against the model's unconstrained reference. We...

    arxiv.org/abs/2606.24083 · PDF

  19. 19

    Selective Capability Unlearning in End-to-End Spoken Language Understanding

    Akanksha Singh, Vinod Kumar Kurmi

    cs.CL · cs.AI

    Modern spoken language understanding (SLU) systems are increasingly deployed in real-world settings, where specific functionalities may need to be removed due to policy or safety constraints. In SLU, a functionality corresponds to an intent and its associated slot-generation behavior. However, in autoregressive models, suppressing a target intent does not eliminate the conditional mapping that generates slots conditioned on that intent. When...

    arxiv.org/abs/2606.24063 · PDF

  20. 20

    Towards Version-aware Operations and Transaction Memories for Multi-layer MeMo

    Peiran Li

    cs.CL · cs.AI · cs.SC

    MeMo proposes language models with explicit multi-layer correlation matrix memories (CMMs), where memorization, retrieval, and forgetting are architectural operations. This paper asks how such memories can reduce the need for retraining when knowledge changes. For changes expressible as MeMo memory associations, the model's accessible knowledge can be updated by editing explicit memories rather than retraining the whole model. We propose a...

    arxiv.org/abs/2606.24040 · PDF

  21. 21

    Towards Spec Learning: Inference-Time Alignment from Preference Pairs

    Dhriti Krishnan, Tejas Goyal, Jaromir Savelka

    cs.CL · cs.AI

    Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on a careful inspection of the model's responses. This is an involved, brittle, and error-prone process. Preference-based fine-tuning is a more rigorous but often prohibitively expensive solution. We propose spec learning, a framework that relies on a brief user instruction and a small set of preference...

    arxiv.org/abs/2606.24004 · PDF

  22. 22

    RASC+: Retrieval-Constrained LLM Adjudication for Clinical Value Set Authoring

    Sumit Mukherjee

    cs.CL · cs.AI · cs.LG

    Clinical value sets define the standardized terminology codes used in quality measurement, phenotyping, cohort construction, and clinical decision support. The recently introduced Retrieval-Augmented Set Completion (RASC) benchmark showed that direct zero-shot large language model (LLM) generation is poorly suited to this task: clinical code systems are large, version-controlled, and not reliably memorized by language models. We study a...

    arxiv.org/abs/2606.23992 · PDF

  23. 23

    Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization

    Shuo Guan

    cs.CL · cs.AI

    End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions they offer are typically coarse (whole documents or passages) and generated post hoc, leaving each summary statement hard to verify. We revisit the modular Extract--Select--Rewrite paradigm and recast its intermediate representation as the unit of attribution. We present CAMS, a Claim-Anchored Multi-document...

    arxiv.org/abs/2606.23989 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.