cs.CL · 2026-09-02 · No. 103

Computation and Language, 2026-09-02.

23 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

23 entries
  1. 01

    Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

    Himil Vasava, Ming Jiang

    cs.CL · cs.LG

    LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. We investigate this procedure mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt...

    arxiv.org/abs/2609.01604 · PDF

  2. 02

    CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?

    Damien Sileo, Dimitri Kachler

    cs.CL · cs.AI

    Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a local plugin change can propagate through dependencies and cleanup. We introduce CordisBench, a 1,200-question benchmark of this lifecycle reasoning. It combines a controlled formal setting with programs executed against Cordis, a runtime that manages component dependencies and cleanup, and asks...

    arxiv.org/abs/2609.01600 · PDF

  3. 03

    The Rise of Verbal Reinforcement Learning

    Kshitij Tayal, Arun Sharma, Genta Indra Winata, Anirban Das, Sambit Sahu

    cs.CL · cs.AI

    Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Reinforcement Learning (VRL) and offer the first unified account of it. We organize the field around a single axis, \textit{when} verbal feedback takes effect in an agent's lifecycle and...

    arxiv.org/abs/2609.01597 · PDF

  4. 04

    Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

    Jingtan Wang, Arun Verma, Xiaoqiang Lin, Zhengyuan Liu, Nancy F. Chen, Daniela Rus, Bryan Kian Hsiang Low

    cs.CL · cs.AI · cs.LG

    How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during LLM post-training remains an open problem. Existing work characterizes only broad trends (e.g., SFT dominates in low-data regimes), lacks a principled allocation framework, and does not examine whether the optimal ratio transfers across model sizes. We frame this problem in terms of near-optimality: rather than seeking a single...

    arxiv.org/abs/2609.01573 · PDF

  5. 05

    From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification

    Manish Gupta, Chaitanya Giri, Jayasimha Talur

    cs.CL · cs.AI

    Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the distinctions are domain-specific and not captured by pre-training. To handle large label spaces, a common approach retrieves top-$K$ candidate labels by embedding similarity and prompt the LLM to choose among them. However, top-$K$ retrieval reduces the number of candidates but does not help the model tell similar ones apart....

    arxiv.org/abs/2609.01564 · PDF

  6. 06

    GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions

    Elias Stengel-Eskin, Newton Sander, Carlos Bonetti, Sasha Boguraev, James Bowler, Hale Sirin, Simon Kirby

    cs.CL · cs.AI · cs.MA

    The growing rate at which LLM agents interact with one another raises key questions about language evolution in multi-LLM-agent settings, with implications for safety and monitorability as well as for linguistic accounts of LLMs. To address these questions, we introduce GlossoGen, a novel platform for studying multi-agent language evolution in complex scenarios. Within GlossoGen, we build the SaveVeyru scenario, which requires agents with...

    arxiv.org/abs/2609.01491 · PDF

  7. 07

    Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA

    Nishant Mishra, Ameen Abu-Hanna, Iacer Calixto

    cs.CL · cs.LG

    Linear classifiers trained on hidden states of a large language model (LLM), linear probes, can flag factual errors from a single forward pass. Geometrically, that implies that true and false statements separate along a stable direction in hidden state space, i.e., the truth direction. Prior work disagrees on whether this generalises across input shifts, but the disagreement is hard to interpret because cross-dataset probe transfer...

    arxiv.org/abs/2609.01361 · PDF

  8. 08

    Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR

    Esther Xin

    cs.CL · cs.LG

    Reinforcement learning with verifiable rewards (RLVR) and standard benchmark evaluation both rely on an automatic verifier that turns a free text answer into a binary reward. Prior work reports that one evaluation harness accepts only about 94% of its own ground truth answers, blaming LaTeX parsing. That is an aggregate: it does not say which answer forms consume the error budget. We supply the decomposition. We apply metamorphic testing to...

    arxiv.org/abs/2609.01354 · PDF

  9. 09

    CHARM: Character Hallucination for Multicultural Role Play Benchmark

    Sunkyung Han, Nahyeon Park, Gaeun Seo, Seunghyun Yoon, JinYeong Bak

    cs.CL · cs.AI

    Role-playing large language models (LLMs) are expected to adopt a character's style while also respecting that character's knowledge boundaries. Prior evaluations detect character hallucination but rarely distinguish whether errors arise from failure to recognize a boundary or from failure to comply despite recognition. We introduce CHARM, a multicultural benchmark of 40 real and fictional characters drawn from five cultural-linguistic...

    arxiv.org/abs/2609.01352 · PDF

  10. 10

    Probing Factual Knowledge Transfer with Training Data Interventions

    Romina Oji, Marc Braun, Marcel Bollmann, Marco Kuhlmann, Jenny Kunz

    cs.CL · cs.AI

    Do multilingual language models transfer factual knowledge across languages during continued pretraining, or do they mostly recall facts learned directly from the target-language data? To answer this question more reliably, we propose an intervention-based framework: starting from an English-pretrained model, we continue pretraining on Persian data from which specific facts have been systematically removed at varying levels of granularity. We...

    arxiv.org/abs/2609.01341 · PDF

  11. 11

    Exploring Sparse Autoencoders in Text-Based Causal Confounding Adjustment

    Mian Zhong, Katherine A. Keith, Anjalie Field

    cs.CL · cs.LG

    In many settings, studying causal questions based on text data requires adjusting for confounding information within texts. Yet there is a tradeoff in constructing text representations for adjustment: they must be sufficiently large and/or dense to preserve the confounding variables necessary for unbiased effect estimation, but sufficiently small and/or sparse to satisfy finite-sample overlap and yield low-variance estimates. To address this...

    arxiv.org/abs/2609.01322 · PDF

  12. 12

    Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models

    Tian Fang, Gaël Guibon, Davide Buscaldi

    cs.CL · cs.AI

    Emotion is expressed in text along a wide spectrum, from surface lexical cues to inferences entangled with content. Most layer-wise analyses of emotion in LLMs use a single corpus, leaving open whether the depth at which emotion becomes accessible is a property of the model or also of the text source. We investigate this across three datasets spanning different degrees of explicitness and contextualization in emotion expression (Twitter...

    arxiv.org/abs/2609.01279 · PDF

  13. 13

    Towards AI-Assisted Clinical Trial Matching: Practical Considerations, Multicenter Evaluation, and Real-World Deployment

    Yin Fang, Qiao Jin, Shubo Tian, Lauren He, Maya Geer, Noor Naffakh, Ryan Huu-Tuan Nguyen, Zifeng Wang, Jimeng Sun,...

    cs.CL · cs.AI · cs.CY

    Clinical trials are essential for advancing cancer care and drug development, but many fail because of insufficient patient enrollment. While there is growing interest in using AI to support patient recruitment, existing systems largely perform eligibility assessment alone and have rarely been evaluated in real-world oncology workflows. Here we present TrialGPT 2.0, an AI-assisted clinical trial recommendation system designed for real-world...

    arxiv.org/abs/2609.01202 · PDF

  14. 14

    EDRAC: Benchmarking Arabic Dialect Reading Comprehension

    Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn, Nizar Habash, Reham Marzouk, Malik H. Altakrori, Younes Samih,...

    cs.CL · cs.AI

    Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly for machine reading comprehension (MRC) and question answering (QA). Existing Arabic QA benchmarks primarily focus on formal written MSA or multiple-choice QA, with limited coverage of naturally spoken dialects. Here, we aim to bridge this gap. We introduce EDRAC, the first large-scale benchmark for dialectal Arabic machine reading...

    arxiv.org/abs/2609.01113 · PDF

  15. 15

    StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions

    Chao Gao, Haijiang Liu, Qiyuan Li, Caicai Guo, Frank van Harmelen, Jinguang Gu

    cs.CL · cs.AI

    Large language models often answer the same multiple-choice question inconsistently when it is posed under support-oriented and elimination-oriented framings. We investigate whether these discrepancies arise from different internal representations induced by the two framings. We introduce a dual-framing protocol with minimally varied prompts that use either support- or elimination-oriented framing while keeping the evaluation target fixed. To...

    arxiv.org/abs/2609.01081 · PDF

  16. 16

    Lagged Coupling: Internal Representations Become Readable Before They Become Causal

    Xining Xun

    cs.CL · cs.AI

    Across the full Pythia suite (160M-12B, eight checkpoints, four task families), a linear probe can read a target variable from the residual stream as early as step 1,000 at every scale -- yet steering along that same reading direction remains null-equivalent in 43 of 48 model-checkpoint cells. Internal readability systematically outruns causal efficacy, and the lag does not shrink with scale. We call this structure lagged coupling and...

    arxiv.org/abs/2609.01048 · PDF

  17. 17

    Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close

    Rania Elbadry, Ahmed Heakl, Saeed Almheiri, Fan Zhang, Muhra AlMahri, Xueqing Peng, Mohsinul Kabir, Shuyao Wang, Yi...

    cs.CL · cs.AI · cs.LG

    When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer correctly within it. We call this setting normative pluralism and study it in Islamic finance using a four-choice taxonomy that separates framework selection from within-framework correctness. This separation reveals the stereotype trap: a cultural cue steers a model toward one framework, but...

    arxiv.org/abs/2609.00999 · PDF

  18. 18

    Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages

    Michele Ciletti

    cs.CL · cs.AI

    Word Sense Disambiguation has advanced rapidly for English and a handful of well-resourced modern languages, but it continues to assume the existence of a sense inventory and a word-to-sense mapping in the source language (Navigli, 2026). These assumptions break down for most historical and low-resource languages, whose dedicated WordNets are either incomplete or still under construction. We present Inspicio, an open-vocabulary retrieval...

    arxiv.org/abs/2609.00998 · PDF

  19. 19

    Disclosure-Gated User Simulation for Companion-Agent Evaluation

    Yao Liu, Yu He

    cs.CL · cs.AI · cs.HC

    Using a large language model to play the user is now standard in scalable evaluation. It has a repeatedly diagnosed failure: the simulated user is excessively cooperative, so a system under test can score by the sheer number of questions it asks rather than by making the user willing to speak. We answer with a disclosure gate conditioning information release on the companion agent's behaviour: its state is a ladder of five ordered gates,...

    arxiv.org/abs/2609.00982 · PDF

  20. 20

    Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

    Kangjia Zhao, Jiajun Li, Haozhan Shen, Wei Chow, Linfeng Li, Hang Song, Lingdong Kong, Chen Zhi, Tiancheng Zhao,...

    cs.CL · cs.AI

    Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass closed-source frontier models in aggregate accuracy. However, this metric averages over many different multi-turn situations and obscures whether progress is balanced across them. We propose an action-class-oriented diagnostic framework that decomposes multi-turn...

    arxiv.org/abs/2609.00949 · PDF

  21. 21

    DualStake: Dual-Path Confidence Calibration in Deep Research Agents

    Yinuo Xu, Yuwei Liang, Jianjie Cheng, Meng Wang, Yongcan Yu, Shuo Lu, Jian Liang

    cs.CL · cs.AI · cs.LG

    Deep Research agents tackle knowledge-intensive tasks through multi-round retrieval and decision-oriented generation. However, these agents suffer from severe overconfidence, making their expressed confidence unreliable for user trust and downstream abstention. To address this, we augment the Deep Research pipeline with step confidence elicitation after each retrieval, building on the commonly used post-answer verbalized confidence....

    arxiv.org/abs/2609.00935 · PDF

  22. 22

    Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO

    Prakhar Gupta, Vaibhav Gupta

    cs.CL · cs.AI · cs.LG

    Language models can ignore prompt evidence when it conflicts with memorized knowledge. Post-training can make models follow such evidence more reliably, but it is unclear whether these gains require new machinery or strengthen machinery already present. We compare nine post-training arms spanning GRPO, SFT, and DPO from one starting checkpoint, with key comparisons extended across scales and families. We estimate a grounding direction from...

    arxiv.org/abs/2609.00925 · PDF

  23. 23

    Dense Process Supervision for Search Agents via Fact Utility Estimation

    Rongzhi Zhu, Xiangyu Liu, Yi Liu, Shuo Zhang, Ruirui Zhang, Rui Wu, Tao Jiang, Zequn Sun, Wenhao Xu, Wei Hu

    cs.CL · cs.LG

    Reinforcement learning (RL) for search agents typically relies on outcome rewards. However, it often fails to achieve effective credit assignment, due to the unclear value of intermediate steps. It is hard to separate their contributions from the final result. In this paper, we propose a dense process supervision method based on fact utility estimation, which models the reasoning process as the accumulation of discrete evidence facts. We...

    arxiv.org/abs/2609.00833 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.