cs.CL · 2026-09-09 · No. 110

Computation and Language, 2026-09-09.

11 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

11 entries
  1. 01

    Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

    Leyuan Tang, Kangda Wei, Tianyu Jiang, Ruihong Huang

    cs.CL · cs.AI

    Large language models (LLMs) may abandon correct positions when users push back, exhibiting a failure mode known as sycophancy. Existing evaluations typically use short, pre-specified conversations and may therefore miss failures that emerge under sustained, adaptive disagreement. We introduce SPINE, a benchmark in which an LLM proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns. We evaluate...

    arxiv.org/abs/2609.09090 · PDF

  2. 02

    Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics

    Andy Nkansah, Hanna Plotnitskaya, Stanislau Salavei, Anna Kozlova, Piotr Gibas, Julian Milek, Viktar Harbachou,...

    cs.CL · cs.AI · cs.HC

    Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We compared Doctorina, eight physicians and four standalone frontier language models in 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 concordance versus 57.0% for physicians (difference, 25.0 percentage points; 95% confidence interval, 17.7-32.7) and 97.3% versus 85.0%...

    arxiv.org/abs/2609.09070 · PDF

  3. 03

    The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits

    Siddharth Vohra, Manikandan Ravikiran

    cs.CL · cs.AI · cs.CY

    Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lending, and medical triage: 40,726 requests to five models, applications differing only in the applicant's name, and a primary...

    arxiv.org/abs/2609.09048 · PDF

  4. 04

    Evaluation of Contextual Understanding in Large Language Models

    Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan,...

    cs.CL · cs.LG

    Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingual Evaluation Understudy (BLEU), or surface-level accuracy fail to reveal how well LLMs extract, integrate, and reason over contextual information--a gap particularly critical in question answering, where models must...

    arxiv.org/abs/2609.09004 · PDF

  5. 05

    Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence Ablation

    Xingyu Deng, Mingzi Cao, Nikolaos Aletras, Xi Wang, Mark Stevenson

    cs.CL · cs.AI · cs.IR

    Automatic fact-checking systems assess the veracity of claims given evidence from relevant documents. Large Language Models (LLMs) have demonstrated strong performance in fact-checking due to their general reasoning capabilities. However, it remains unclear whether they faithfully make use of the evidence provided to reach veracity judgments or rely on parametric knowledge. To investigate this, we introduce Fact-Ablated Evaluation (FAE), a...

    arxiv.org/abs/2609.08943 · PDF

  6. 06

    Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations

    Antonin Poché, Fanny Jourdan, Nils Feldhus, Qianli Wang, Jing Yang, Simon Ostermann, Nicholas Asher, Philippe...

    cs.CL · cs.LG

    Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is costly, automated simulatability replaces human explainees with LLM simulators, as proposed in ConSim (Poché et al., 2025) for large-scale experiments. We qualitatively replicate and extend ConSim's ranking of explanation methods across the tested datasets,...

    arxiv.org/abs/2609.08585 · PDF

  7. 07

    Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?

    Sara Rizwan, Samaanah Abdus Salam

    cs.CL · cs.AI

    Long context language models now advertise windows of one million tokens, but two habits limit how much of that window is used. Attention heads with nothing useful to read still spend their budget on the first token, which is called the attention sink, and where a fact sits in the context changes whether the model finds it. Gated attention cut first token attention from 46.7 percent to 4.8 percent at NeurIPS 2025, and Kimi K3 pairs that idea...

    arxiv.org/abs/2609.08574 · PDF

  8. 08

    Same Values, Different Languages? From Multilingual Probing to Steering LLMs Toward Chinese Social Values

    Yuemei Xu, Kexin Xu, Jian Zhou, Haoyu Lu, Yequan Wang, Aishan Liu

    cs.CL · cs.AI

    As Large Language Models (LLMs) are increasingly integrated into human society, aligning them with pluralistic social values has become a critical priority. However, whether LLMs exhibit consistent value preferences across languages remains underexplored, particularly for culturally grounded values, which are more abstract and difficult to evaluate and align than safety-centric principles. We investigate this issue through Chinese Social...

    arxiv.org/abs/2609.08515 · PDF

  9. 09

    Do Reviewers Still Reward Lexical Complexity? A Frozen-Rater Study of Preference Drift in 124K ICLR Reviews

    Jiabin Zheng

    cs.CL · cs.DL · cs.LG

    Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about the evaluator, not about the text. When the association between a writing cue and review scores moves across years, the reviewers may have changed, the submissions may have changed, or both, and a regression of scores on text cannot say which. We separate the two with a frozen rater: 81,850...

    arxiv.org/abs/2609.08475 · PDF

  10. 10

    Tracing Stereotypes from Representation to Output in Multilingual LLMs

    Ariun-Erdene Tumurchuluun, Yusser Al Ghussin, Pinzhen Chen, Josef van Genabith, Koel Dutta Chowdhury

    cs.CL · cs.AI

    Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is represented or how it affects the output. To investigate these internal mechanisms, we compare linear probing, attribution patching, sparse autoencoders (SAEs) and feature ablation in Llama-3.1-8B, Qwen3-8B, and Gemma-2-9B. Probe performance peaks substantially earlier than attribution in all...

    arxiv.org/abs/2609.08322 · PDF

  11. 11

    What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory

    Chen Shen

    cs.CL · cs.AI · cs.DB

    Agent memory systems must discard stored information when their history exceeds a fixed token budget. Existing budget-accuracy frontiers quantify the resulting loss in accuracy, but do not distinguish irreversible losses caused by eviction from recoverable retrieval failures. We introduce the restore counterfactual, a per-question paired intervention that reinstates the question's gold evidence in the read-time context and reruns the same...

    arxiv.org/abs/2609.08279 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.