cs.CL · 2026-09-30 · No. 129

Computation and Language, 2026-09-30.

19 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

19 entries
  1. 01

    STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

    Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu, Ziyu Guo, Renrui Zhang, Zhichao Lu, Zhenan Sun, Ying Wei

    cs.CL · cs.AI · cs.LG

    Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in...

    arxiv.org/abs/2609.38169 · PDF

  2. 02

    How Local Mixing Encodes Relative Position in Global NoPE Attention

    Cutter Dawes, Nick Alonso, Tom Figliolia, Beren Millidge

    cs.CL · cs.AI · cs.LG

    The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encodings have long been assumed to be required, recent methods that interleave local mixing layers, such as sliding window attention (SWA) and gated...

    arxiv.org/abs/2609.38109 · PDF

  3. 03

    Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces

    Ratish Puduppully, Pranabendu Misra, Paarth Iyer, Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati

    cs.CL · cs.AI

    Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechanically verifiable. We revisit it in iGSM, a synthetic grade-school mathematics benchmark designed to study thinking traces and used to support claims of learned reasoning and planning. Crucially,...

    arxiv.org/abs/2609.38107 · PDF

  4. 04

    Gender bias across LLMs is common and highly heterogenous

    Edoardo Bolzoni, Valerio Capraro

    cs.CL · cs.AI · cs.CY · cs.HC

    Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms:...

    arxiv.org/abs/2609.38036 · PDF

  5. 05

    Auditable Long-Term Memory: A Deterministic Retrieval Chain Measured at 479/475 of 500 on LongMemEval-S

    Christopher J. Chanhnourack

    cs.CL · cs.AI · cs.IR

    We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, coverage-first packet compilation, and deterministic reasoning scaffolds; an LLM is used only as a replaceable final reader. The chain places all gold sessions in the candidate pool for 468/470 answerable questions and produces gold-complete packets for 462/470. With a Claude Opus reader called...

    arxiv.org/abs/2609.38021 · PDF

  6. 06

    BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals

    Julien Knafou, Luc Mottin, Alexandre Flament, Paul van Rijen, Esteban Gaillac, Patrick Ruch

    cs.CL · cs.AI · cs.IR

    The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admitted only once an entailment cascade has checked it against the passage it cites, and an answer is released only once enough checked evidence stands behind it. Each question is run three...

    arxiv.org/abs/2609.37993 · PDF

  7. 07

    Learning What to Remember: Long-horizon Counterfactual Memory Optimization

    Jiaming Tang, Mingyan Liu, Armin Sarabi

    cs.CL · cs.LG

    Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before the rewrite. We introduce Memory Gain Policy Optimization (MGPO), which isolates the incremental value of each memory rewrite...

    arxiv.org/abs/2609.37930 · PDF

  8. 08

    The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment

    Gonçalo Paulo, Louis Jaburi, Nora Belrose, Lucia Quirke, Stella Biderman

    cs.CL · cs.AI

    Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or 'evil' persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute...

    arxiv.org/abs/2609.37914 · PDF

  9. 09

    It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

    Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov, Carlos Rosas-Hinostroza, Neil Si Smail,...

    cs.CL · cs.LG

    Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, eg, with reasoning traces to address cold-start problems. While demonstratively...

    arxiv.org/abs/2609.37891 · PDF

  10. 10

    How Many Labels Does a Language Need? Annotation Budgets and Cross-Lingual Pooling for African-Language Text Classification

    Bhanu Prakash Vangala, Sowmya Guda, Navya Vangala

    cs.CL · cs.LG

    Every text classifier for an African language begins with a budgeting question: how many labelled examples are needed, and can labels from other African languages stand in for them? We answer both questions empirically for 28 language-task pairs, news topic classification in 16 languages (MasakhaNEWS) and tweet sentiment in 12 languages (AfriSenti), using a character n-gram linear model that trains in seconds on two CPU cores with no...

    arxiv.org/abs/2609.37882 · PDF

  11. 11

    Retrieval Capacity of Self-Attention Under Competition

    Timur Mudarisov, Mikhail Burtsev, Radu State

    cs.CL · cs.LG

    How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention. Without retraining, we retain only the tokens with the highest attention weights at each head, layer, and query, keeping their original weights unchanged. By varying the selected set size and measuring the increase in negative log-likelihood (NLL), we estimate the effective attention set size...

    arxiv.org/abs/2609.37879 · PDF

  12. 12

    It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them

    Nagham Omar, Mahmoud Jabarin, Kinan Ibraheem, Lotem Peled-Cohen

    cs.CL · cs.AI · cs.CV

    Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no...

    arxiv.org/abs/2609.37863 · PDF

  13. 13

    One Threshold Does Not Fit All Languages: Language-Conditional Deferral for Reliable and Efficient Low-Resource Text Classification

    Bhanu Prakash Vangala, Vangala Navya

    cs.CL · cs.LG

    In the Global South, the lower-income countries of Africa, Asia, and Latin America where most of the world's languages are spoken, a deployed text classifier usually runs on ordinary CPUs, serves many languages with a single model, has few labeled examples in any of them, and relies on people to catch its mistakes. Such a system is only useful if it can promise how often it will be wrong: at most a fixed fraction of the labels it assigns on...

    arxiv.org/abs/2609.37861 · PDF

  14. 14

    The Geometry of Inference in Transformer Residual Streams

    Timur Mudarisov, Mikhail Burtsev, Radu State

    cs.CL · cs.LG

    Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states from other contexts. Across six pretrained language models, the own endpoint becomes preferable to the average alternative early, while many...

    arxiv.org/abs/2609.37824 · PDF

  15. 15

    Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue

    Shengbo Cai, Yuxiang Wang, Jingran Xie, Zhisheng Zhang, Shun Lei, Di Cao, Teddy Sun, Zhiyong Wu

    cs.CL · cs.AI

    Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the perception-reasoning gap. In addition, CoT may not fully capture acoustic cues in words, and generating it adds inference latency. To address...

    arxiv.org/abs/2609.37818 · PDF

  16. 16

    CompOrca: Corpus-Scale Compliance Labelling of Instruction-Tuning Data

    Philipp E. Glass, Alina Miron

    cs.CL · cs.LG

    Studying how fine-tuning shapes refusal and noncompliance behaviour requires identifying training examples that refuse, evade or otherwise fail to fulfil the requested task. But existing annotation covers evaluation sets of a few thousand prompts at most. We present CompOrca, a compliance labelling over the entirety of the 4,233,923-example OpenOrca corpus. Every example was classified as compliant or noncompliant by five independent passes...

    arxiv.org/abs/2609.37807 · PDF

  17. 17

    A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses

    Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo

    cs.CL · cs.AI

    Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR.BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the...

    arxiv.org/abs/2609.37788 · PDF

  18. 18

    Evaluating and Benchmarking the System One Model Jev

    Tobias Deußer, Lorenz Sparrenberg, Rafet Sifa

    cs.CL · cs.AI

    Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor describes as calibrated. Such models target small decisions in information access pipelines, such as routing queries, checking grounding, moderating content, or rating against a rubric. We...

    arxiv.org/abs/2609.37647 · PDF

  19. 19

    Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision

    Jacob Epifano

    cs.CL · cs.AI · cs.LG

    Fine-tuning a language model on a narrow set of harmful demonstrations, such as bad medical advice, can make it broadly misaligned on unrelated questions, a phenomenon known as emergent misalignment (EM). The usual defense is to find the offending rows and delete them, but a row locator failed our held-out test and deleting rows helps less than expected. We ask a different question: given a fixed set of poisoned rows, is it better to correct...

    arxiv.org/abs/2609.37624 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.