cs.CL · 2026-10-01 · No. 130

Computation and Language, 2026-10-01.

17 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

17 entries
  1. 01

    How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

    Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

    cs.CL · cs.LG

    Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild...

    arxiv.org/abs/2609.40295 · PDF

  2. 02

    Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning

    Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav

    cs.CL · cs.AI

    Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neither scalable nor desirable as it amplifies damage to unrelated model capabilities. We introduce the task of language budgeted multilingual...

    arxiv.org/abs/2609.40286 · PDF

  3. 03

    Comparison of techniques for fine-tuning open-weight models for entity extraction from radiology reports

    Aawez Mansuri, Kush Mehta, Mohammadreza Chavoshi, Jahanzaib Malik, Theodorus Dapamede, Frank Li, Rohan Isaac,...

    cs.CL · cs.LG

    Converting free-text radiology reports into structured labels supports cohort building, quality assurance, and monitoring of clinical imaging models, but the strongest label extractors are hosted proprietary models whose use raises privacy, cost, and reproducibility concerns. We asked whether a fine-tuned open-weight model (Gemma-3-12B) can match GPT-4o at multi-label intracranial hemorrhage (ICH) acuity extraction from non-contrast head-CT...

    arxiv.org/abs/2609.40236 · PDF

  4. 04

    SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models

    Kanpat Vesessook, Saksorn Ruangtanusak

    cs.CL · cs.AI · cs.SD

    Speech-to-speech systems must solve tasks whose requirements emerge across conversational turns. We introduce SpeechConversationBench (SCB), a focused evaluation of spoken mathematical reasoning using 103 sharded GSM8K problems. The framework compares the original problem delivered in one turn (full), its concatenated information shards delivered together (concat), and incremental spoken disclosure across turns (sharded). We report...

    arxiv.org/abs/2609.40198 · PDF

  5. 05

    On the (In)effectiveness of AMR Augmentation for Large Language Models

    Hoa Quynh Nhung Nguyen, Jacopo Staiano, Michael Sullivan

    cs.CL · cs.AI

    While Abstract Meaning Representation (AMR) has historically improved performance on a range of NLP tasks, the benefit---or lack thereof---of AMR augmentation for modern LLMs is thus far unclear. In this paper, we attempt to reproduce recent work that reported substantial downstream gains from AMR augmentation, finding that these are likely due to specific choices in the experimental settings used: using a consistent and unified protocol for...

    arxiv.org/abs/2609.40121 · PDF

  6. 06

    OPTS-TTPO: Enhancing Finite-Sample Policy-Gradient Learning with Tree Search

    Junyu Lu, Shichao Weng, Zhiqiang Wang, Haojie Luo, Jingfan Zhang, Yuhua Zhou, Cheng Du, Yuzhuo Zhang, Xi Li, Jinwei...

    cs.CL · cs.LG

    The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We study whether tree search improves their coverage within a fixed budget while controlling gradient bias. We introduce On-Policy Parallel Tree Search (OPTS) and Tree Trajectory Policy Optimization (TTPO) using on-policy tree trajectories, which sample new suffixes from the current policy at...

    arxiv.org/abs/2609.40035 · PDF

  7. 07

    Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

    Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan, Yonggan Fu, Jindong Jiang, Mingjie Liu, Ehsan...

    cs.CL · cs.AI · cs.MA

    Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory...

    arxiv.org/abs/2609.39982 · PDF

  8. 08

    Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering

    Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, Martin Krallinger, Miguel Rodríguez-Ortega, Eduard...

    cs.CL · cs.AI · cs.IR

    This paper presents an overview of the fourteenth edition of the BioASQ challenge, organized in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2026. BioASQ is an international challenge series that supports progress in biomedical language processing tasks ranging from semantic indexing and information extraction to question answering and summarization. In 2026, BioASQ included six shared tasks: a) Task 14b on biomedical...

    arxiv.org/abs/2609.39975 · PDF

  9. 09

    LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

    Juyi Lin, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi, Xinru Jiang, Yanzhi Wang,...

    cs.CL · cs.AI · cs.CV

    Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight...

    arxiv.org/abs/2609.39938 · PDF

  10. 10

    OPSRD: On-Policy Self-Role Distillation

    Weijie Ren, Yanwen Zhang, Hao Li, Zhuolin Qi, Hengyi Zhang, Naibo Wang

    cs.CL · cs.AI · cs.LG

    Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD,...

    arxiv.org/abs/2609.39884 · PDF

  11. 11

    LLM Persona Unlearning

    Kemou Li, Zhuan Shi, Qizhou Wang, Fengpeng Li, Negar Rostamzadeh, Golnoosh Farnadi, Jiantao Zhou

    cs.CL · cs.LG

    Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-training teaches conditional enactment and makes a helpful Assistant the default, but it does not erase alternative modes from the weights; explicit prompts can therefore elicit personas that repeatedly shape judgment, language, and action. In open-weight settings, runtime controls can be...

    arxiv.org/abs/2609.39882 · PDF

  12. 12

    When a Kindergartener Solves Calculus: Measuring Capability Leakage in Role-Prompted Reasoning Models

    Pakhapoom Sarapat, Saksorn Ruangtanusak, Kunat Pipatanakul, Pittawat Taveekitworachai

    cs.CL · cs.AI · cs.HC

    We investigate the problem of role-capability leakage (RCL), in which a role-prompted reasoning model generates convincing in-role text while continuing to exhibit capabilities on benchmarks that exceed those implied by the assigned role. For example, when a model is prompted to assume the role of a kindergarten student, one might expect its performance on a mathematics benchmark to reflect kindergarten-level ability rather than expert-level...

    arxiv.org/abs/2609.39846 · PDF

  13. 13

    Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior

    Atsuki Yamaguchi, Tatsuro Inaba, Joel Niklaus, Michal Štefánik, Aline Villavicencio, Nikolaos Aletras

    cs.CL · cs.AI · cs.LG

    Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective...

    arxiv.org/abs/2609.39827 · PDF

  14. 14

    Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations

    Maximilian von Klinski, Sebastian Lapuschkin, Wojciech Samek, Lennart Bürger

    cs.CL · cs.AI · cs.LG

    Lie detection probes aim to predict from a language model's internal states whether its output is truthful or dishonest. However, role-play complicates what "truth" means for an LLM: language models can adopt a wide range of personas that take very different claims to be true, including personas whose beliefs clearly contradict reality, such as a conspiracy theorist. In this work, we investigate whether lie detection probes reliably flag...

    arxiv.org/abs/2609.39807 · PDF

  15. 15

    Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation

    Xincheng Wei, Yifan Ding, Yoshua Li, Yuquan Lu, Ziheng Li, Yi Lu, Dongsheng Ma, Rongxiang Weng, Xunliang Cai

    cs.CL · cs.LG

    On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these...

    arxiv.org/abs/2609.39687 · PDF

  16. 16

    SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration

    Weijie Ren, Yanwen Zhang, Hao Li, Zhuolin Qi, Hengyi Zhang, Naibo Wang

    cs.CL · cs.AI · cs.LG

    Multi-agent collaboration lets large language models (LLMs) improve question answering through deliberation and feedback. Yet shared discussion couples correction with exposure to the same mistakes, which can erode the diversity needed for voting. Self-consistency offers sampling diversity without feedback, while single-pair Actor-Critic collaboration refines only one candidate. We introduce SEPAL, which assigns three private Actor-Critic...

    arxiv.org/abs/2609.39645 · PDF

  17. 17

    Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies

    Dalton Raphael Harmsen, Swier Garst, Thomas van Osch, Zarè Palanciyan, Joaquin Vanschoren

    cs.CL · cs.AI

    Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological...

    arxiv.org/abs/2609.39640 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.