cs.CL · 2026-09-24 · No. 123

Computation and Language, 2026-09-24.

21 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

21 entries
  1. 01

    Contrastive Learning for Authorship Verification

    Peter Kirby

    cs.CL · cs.LG

    Our results show that contrastive learning outperforms a classification-based approach to authorship verification under the tested settings. We identify loss function, batch size, training duration, pre-trained model, input context length, and random text span data augmentation as important factors of model performance. Based on these considerations, we develop a ModernBERT Bi-Encoder model that achieves 98.4% accuracy on the PAN21 authorship...

    arxiv.org/abs/2609.28471 · PDF

  2. 02

    Agent-Editing World Model: Rethinking World Modeling for LLM Agents

    Shuang Sun, Guoxin Chen, Fanzhe Meng, Jia Deng, Huatong Song, Jinhao Jiang, Wayne Xin Zhao, Hongteng Xu, Ji-Rong Wen

    cs.CL · cs.AI · cs.LG

    Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limited value when real feedback is available. Meanwhile, agents suffer from \emph{task-state contamination}, where unsupported...

    arxiv.org/abs/2609.28416 · PDF

  3. 03

    Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following

    Niklas Scholz, David Thulke, Abdallah Nasir, Will Allred, Evgeny Matusov, Hermann Ney

    cs.CL · cs.LG

    Fine-tuning large language models on parallel data improves translation quality but can cause catastrophic forgetting. Mitigation methods are generally evaluated by retention on general benchmarks. We ask whether these findings transfer to machine translation (MT) fine-tuning and to MT-specific instruction following (MT-IF): instructions that modify a translation, such as formality, grammatical gender, and length control. We compare methods...

    arxiv.org/abs/2609.28395 · PDF

  4. 04

    Predicting Quantization Price for Selecting PTQ Configurations Before Deployment

    Junbin Qiu, Jian Mu, Weitong Zhang, Yao Shu

    cs.CL · cs.LG

    Weight-space post-training quantization (PTQ) must choose finite formats, granularities, quantizer families, transformations, and bits before the completed quantized model reveals its output-distribution drift. Existing PTQ methods predict important pieces of this degradation, including reconstruction error, Hessian sensitivity, transformation effects, and downstream loss, but these pieces are usually scored after fixing the quantization...

    arxiv.org/abs/2609.28270 · PDF

  5. 05

    Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?

    AbdulRahman A. Morsy, Aya Zirikly

    cs.CL · cs.AI

    Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored. Prior work has focused largely on modern language varieties and poetry, while classical prose traditions such as maqama remain largely unstudied. The maqama is a classical literary genre characterized by rhymed prose (saj), dense rhetorical...

    arxiv.org/abs/2609.28245 · PDF

  6. 06

    Scaling Attention Head Analysis via Gradient-Based Attribution in Context-Aware Machine Translation

    Paweł Mąka, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis

    cs.CL · cs.AI

    In this paper, we introduce a gradient-based head attribution strategy where the Token-level Max-Margin loss is backpropagated to the attention maps. This framework enables a large-scale causal analysis of attention heads, making it suitable for LLMs. We evaluate our method on the task of disambiguation in Context-aware Machine Translation, where we analyze 50 phenomena across 4 models and 4 language directions. We empirically show the...

    arxiv.org/abs/2609.28117 · PDF

  7. 07

    Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark

    Makar Ulesov, Vladislav Smirnov, Omar Ibrahim, Arsenii Bobovnikov

    cs.CL · cs.AI · cs.CE · cs.SE

    Backtest auditing is a calibration problem: high flaw recall is not useful when the model falsely flags matched clean strategies. We build a 96-item paired benchmark in which every flawed backtest has a clean control that holds strategy, dates, code style, labels, and reporting scaffold fixed while changing one methodology detail. A deterministic scorer separates flaw recall, clean-control false positives, evidence localization, and fix...

    arxiv.org/abs/2609.28090 · PDF

  8. 08

    TEMPS: Temporal Sentence Embeddings for Temporal Information Retrieval

    Mourad Hassani, Julien Romero, Amel Bouzeghoub, Christian Jacquelinet

    cs.CL · cs.AI

    Modern information retrieval (IR) systems rarely represent time, yet many information needs depend on it: in clinical, journalistic, and legal search, when an event occurred can decide whether a document is relevant. Dense retrievers and Retrieval-Augmented Generation (RAG) pipelines match queries to documents well on topic but poorly on time, so they surface content that is on-topic yet temporally wrong. We introduce Temporal Textual...

    arxiv.org/abs/2609.28048 · PDF

  9. 09

    Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing

    Norah Almousa, Shayan Peyghambari Oskoui, Raquel Coelho, Gayle Rogers, Xiang Lorraine Li, Diane Litman

    cs.CL · cs.AI

    We investigate whether state-of-the-art large language models (LLMs) generate feedback that reflects the pedagogical practices of expert teachers in terms of feedback focus and adaptivity. Previous evaluation efforts have examined feedback characteristics, its impact on learning, and its target, yet the focus of feedback and its adaptivity remains largely overlooked. To bridge this gap, we adopt and refine Narciss's taxonomy into seven...

    arxiv.org/abs/2609.28026 · PDF

  10. 10

    Evaluating Open-Weight LLMs for Turkish Domain Documents Under Retrieval and Hardware Constraints

    Imtiaz Ul Hassan, Öykü Akbulut, Onur Kaya, Ardhendu Behera, Swagat Kumar, Peter Matthew, Yonghuai Liu

    cs.CL · cs.IR · cs.LG

    Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain documents. This paper evaluates five open-weight 7B-8B models for Turkish document question answering under a resource-constrained local deployment setting. The primary benchmark contains 100 systematically validated questions derived from a 109-page industrial R&D report, and the evaluation protocol...

    arxiv.org/abs/2609.28007 · PDF

  11. 11

    Controlled Attribute-Specific Summarization of Interrogative Dialogues

    A Aditya Bhardwaj, Arjit Singh Arora, Md Shad Akhtar

    cs.CL · cs.AI

    Effective summarization of interrogative dialogues is a critical task in forensic and investigative settings, requiring high factual accuracy, coherence, and attribute-specific relevance. In this work, we introduce CASPER, a novel Chain-of-Thought Attribute-Specific Prompting for Evaluative Summarization framework that leverages structured prompting and iterative refinement to generate high-quality summaries of interrogator-witness...

    arxiv.org/abs/2609.28004 · PDF

  12. 12

    Risk-Controlled KV-Cache Eviction: From Memory Budgets to Risk Targets

    Beomgu Kang, SoJin Yun, Hojoon Kim, Hyunseok Seo

    cs.CL · cs.LG

    KV-cache eviction is typically evaluated through average quality-memory trade-offs, yet a small average loss can hide requests whose utility degrades materially. We reformulate eviction as a deployment risk-control problem: a material degradation occurs when eviction lowers task utility by more than a deployment-specified tolerance relative to full-KV inference on the same request, and deployment risk is the population frequency of such...

    arxiv.org/abs/2609.27981 · PDF

  13. 13

    Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

    Rasmus Aagaard, Nicki Skafte Detlefsen

    cs.CL · cs.LG

    Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the {\tt whisper-large-v3-turbo} variant reduced the decoder from 32 to 4 layers, while Distill-Whisper similarly reduced the decoder to only 2 layers. Although some attention has been put towards reducing the size of the encoder, no approach has...

    arxiv.org/abs/2609.27980 · PDF

  14. 14

    The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA

    Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia, Nurassyl Askar, Stefano Rini

    cs.CL · cs.AI

    Small language models (SLMs) are increasingly paired with knowledge graphs (KGs), yet end-to-end KG question answering conflates graph access, search, navigation, reasoning, and answer generation. This coupling makes it difficult both to determine whether an SLM can faithfully execute the reasoning path implied by a question and to attribute failures to navigation rather than to other stages of the pipeline. We isolate this capability by...

    arxiv.org/abs/2609.27669 · PDF

  15. 15

    Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality

    Jiaju Huang, Hao Yang, Xinyu Ma, Xinglong Liang, Kunyan Cai, Junqiang Ma, Shaobin Chen, Yue Sun, Tao Tan

    cs.CL · cs.AI

    An AI-generated radiology report can resemble a physician's report while omitting an abnormality, adding an unsupported finding, or reversing its presence. Measuring these factual differences is essential for evaluating report generators. We study Jev, a System One decision model, as a simple, low-cost judge of agreement with physician-written reference reports. Our evaluator checks whether each statement is supported by the other report and...

    arxiv.org/abs/2609.27607 · PDF

  16. 16

    When Context Misleads: In-context Learning with Jurisdiction in Large Language Models

    Pei-lin Li, Qingle Liu, Junyang Feng, Siyu Li, Sunqi Fan, Xin-Sheng Chen, Shuojin Yang

    cs.CL · cs.AI

    In-Context Learning (ICL) has become a cornerstone of modern LLM deployment. However, existing ICL post-training methods have a critical blind spot: they excel at extracting patterns from demonstrations while often neglecting context authority, the ability to determine whether contextual information should govern the final answer. To benchmark this capability, we introduce FakeContextBench, which contains pseudoscientific claims across seven...

    arxiv.org/abs/2609.27603 · PDF

  17. 17

    Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models

    Kaifeng Tan, Yudong Li, Linlin Shen

    cs.CL · cs.AI

    Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results. Reliable evaluation is particularly challenging for base models, whose limited instruction-following ability complicates task-based assessment. We introduce Uncheatable Eval, a dynamic benchmark that regularly collects newly published text to...

    arxiv.org/abs/2609.27510 · PDF

  18. 18

    Planned Test-Time Scaling with Coordinated Reasoning Paths

    Xueqing Wu, Langxing Bai, Hritik Bansal, Po-Nien Kung, Shuo Li, Hao Liu, Nanyun Peng, Kai-Wei Chang

    cs.CL · cs.AI

    Test-time scaling with parallel branches is widely adopted to improve performance on challenging reasoning tasks. The predominant approach, repeated sampling, draws branches independently from a single policy, which can produce redundant attempts and thereby limit the gains from additional inference compute. To address this limitation, we propose Planned Test-Time Scaling (PTTS), which replaces independent sampling with a coordinated joint...

    arxiv.org/abs/2609.27374 · PDF

  19. 19

    Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models

    Kian Shamsaie, Iman Modarressi

    cs.CL · cs.AI · cs.HC · cs.SD

    Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from...

    arxiv.org/abs/2609.27372 · PDF

  20. 20

    Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

    Ke Wan, Chen Chen

    cs.CL · cs.LG

    Recurrent language models repeatedly apply shared network blocks to refine latent representations, but standard inference recomputes global attention at every recurrent step. We study attention dynamics across recurrent depth and find that attention support and distributions stabilize substantially earlier than hidden states and attention outputs. This suggests a two-stage structure: early steps discover a sparse working set of relevant...

    arxiv.org/abs/2609.27373 · PDF

  21. 21

    Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition

    Hao Shi, Yun Liu, Xuehao Yang, Jun Liu, Chuanbo Hua, Xuanjun Chen, Lianbo Liu, Shiao Zhu, Zixiong Su

    cs.CL · cs.AI

    Conventional Japanese automatic speech recognition (ASR) is supervised by an orthographic transcript, although the same written form can correspond to different lexical readings realized in speech. Such utterances receive an identical target, so their reading distinction is absent from the supervision interface and cannot be recovered reliably by post-hoc text-only grapheme-to-phoneme conversion. We present Ruby-ASR, which refines the...

    arxiv.org/abs/2609.27289 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.