cs.CL · 2026-09-15 · No. 114

Computation and Language, 2026-09-15.

18 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

18 entries
  1. 01

    Disentangling Representation Evolution in Transformers through Directional Decomposition

    Shwai He, Haichao Zhang, Shen Yan

    cs.CL · cs.LG

    Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study this evolution as a functional geometry, decomposing learned updates into parallel and perpendicular components. Across pretrained models, we find substantial parallel components beyond the residual identity path. We then apply the decomposition in two spaces: to attention and MLP updates relative...

    arxiv.org/abs/2609.15975 · PDF

  2. 02

    Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States

    Zixuan Wang, Yufan Zhou, Jinzhou Tang, Xinle Yu, Chengjun Wu, Lyumanshan Ye, Zhaoxiang Feng, Letian Peng, Adyasha...

    cs.CL · cs.LG

    As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental supervision gap because current datasets for LLM assistant training contain few if any well-informed responses explicitly grounded in users' unspoken beliefs and goals. Scaling such supervision is inherently...

    arxiv.org/abs/2609.15972 · PDF

  3. 03

    K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

    Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel...

    cs.CL · cs.AI · cs.LG

    % !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configurations representing 33 base models from 14 providers across a fixed cohort of 200 multi-turn vignettes involving suicide, self-harm, domestic violence, substance...

    arxiv.org/abs/2609.15855 · PDF

  4. 04

    Before You Poll with LLMs: A Deliberative Diagnostic Framework

    Ahmed Wali, Hassaan Tayyab

    cs.CL · cs.AI · cs.CY

    Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM personas simulate public opinion at scale. Current evaluations test only whether personas hold the right opinions -- a static snapshot. But opinion research increasingly depends on dynamic fidelity: whether personas update beliefs in response to new arguments, as humans do during deliberation. No...

    arxiv.org/abs/2609.15849 · PDF

  5. 05

    CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering

    Sumit Barua, Guan Hong, Halil Dursunoglu, Charles Rodgers, Alvis Fong

    cs.CL · cs.AI · cs.IR

    Retrieval-augmented generation (RAG) can improve access to complex information; however, retrieving evidence alone does not ensure that answers are grounded, citation-valid, or appropriately refused. This paper introduces CiteGuard-RAG, a validation-centered AI system for evidence-grounded question answering. The system integrates hybrid semantic-lexical retrieval, citation-constrained generation, sentence-level grounding validation, and...

    arxiv.org/abs/2609.15830 · PDF

  6. 06

    Look Before You Leap: Factual Decoding with Internal Attribution Signals

    Hayeong Ryu, JungMin Yun, Byeonggeuk Lim, Sunhee Jo, YoungBin Kim

    cs.CL · cs.AI

    Hallucination remains a critical challenge in large language models (LLMs), where early factual errors compound through autoregressive generation in a snowballing effect that neither post-hoc correction nor weight-level intervention can effectively preempt. We propose DescaPE (DEcoding Signal Control Against Path Error-snowballing), a decoding framework that leverages internal model signals to suppress hallucination-prone trajectories at...

    arxiv.org/abs/2609.15745 · PDF

  7. 07

    Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection

    Ana-Maria Luisa Mocanu, Sebastian Mocanu, Ciprian-Octavian Truică, Elena-Simona Apostol

    cs.CL · cs.AI · cs.CV · cs.LG

    Detecting sexism on the internet is a fundamentally subjective task; our team, VANGUARD, addresses this challenge in the EXIST 2026 Task 2 by proposing a human-centered multimodal framework that analyses and incorporates the psychological and demographic characteristics of human annotators into the detection pipeline. We fuse five input modalities through a cross-attention architecture with Feature-wise Linear Modulation conditioning. Meme...

    arxiv.org/abs/2609.15608 · PDF

  8. 08

    Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction

    Hayeong Ryu, Sunhee Jo, Seunguk Yu, YoungBin Kim

    cs.CL · cs.AI

    Grammatical error correction (GEC) evaluation has traditionally relied on reference or edit overlap, which can penalize valid rewrites that differ from gold corrections. Reference-free metrics reduce this dependence, but evaluating whether a fluent output is a valid correction of the source remains challenging. We propose SURE, a source-conditioned reward evaluator trained on within-source preferences spanning minimal-edit and...

    arxiv.org/abs/2609.15559 · PDF

  9. 09

    Authorship attribution and aesthetic evaluation of AI poetry: a case study with Haiku

    Livia Oddi, Simone Scardapane, Toru Sugimoto, Donatella Genovese

    cs.CL · cs.AI

    This paper investigates the generation and human evaluation of Japanese haiku by contemporary Large Language Models (LLMs), focusing on authorship perception and aesthetic judgment within a constrained poetic form. Using a few-shot prompting strategy, Japanese haiku were generated across a heterogeneous set of large language models, including open- and closed-source systems, medium-scale and large-scale architectures, models with native or...

    arxiv.org/abs/2609.15511 · PDF

  10. 10

    How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus

    Ilya Koziev, Leonid Sinev, Ivan Oseledets

    cs.CL · cs.AI

    Orthrus is a hybrid autoregressive-diffusion architecture that accelerates autoregressive language-model inference by generating multiple tokens in parallel while using a frozen autoregressive backbone. Its central claim is that an intra-model consensus mechanism enables lossless speculative decoding, producing the same output sequence as the autoregressive model. We independently reproduce Orthrus and examine this claim under different...

    arxiv.org/abs/2609.15504 · PDF

  11. 11

    Temperature Fragility and the Conditional Benefits of Truncation Sampling

    Francesco La Rosa

    cs.CL · cs.LG

    Large language models generate text by sampling each token from a predicted distribution, and a temperature parameter sets how far the draw strays from the most probable tokens. Truncation samplers such as top-p and min-p discard the least probable tokens before the draw, so that sampling at high temperature stays coherent. Their reported accuracy gains come from temperatures of 1.5 to 3, while the defaults of deployed systems cluster between...

    arxiv.org/abs/2609.15476 · PDF

  12. 12

    Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice Evaluation

    M. Ali Bayram

    cs.CL · cs.AI · cs.PF

    Adding answer options can lower multiple-choice scores without improving assessment validity. Turkish MMLU Pro examines this distinction using 12,000 Turkish-source questions across 58 sections. Each question retains its stem, five original options and source key, and receives five options copied from other questions in the same section. Sentence-embedding retrieval proposes candidates; a language model selects existing identifiers....

    arxiv.org/abs/2609.15467 · PDF

  13. 13

    Dynamic Semantic Compression for Efficient Latent-Space Inference in Large Language Models

    Peipei Li, Dongsen Zhang, Yuchen Liu, Wenjun Xu

    cs.CL · cs.AI

    Large Language Models (LLMs) primarily perform inference at the token level, resulting in substantial memory overhead and compromised computational efficiency. In this paper, we propose a Dynamic Semantic Extraction and Inference (DSEI) framework, which achieves segment-level inference within the latent space through a two-stage training strategy. First, we construct a Dynamic Semantic Autoencoder (DSAE) via self-supervised learning. DSAE...

    arxiv.org/abs/2609.15338 · PDF

  14. 14

    Artificial entrepreneurial cognition: Locating and causally steering an opportunity recognition dial inside large language models (LLMs)

    Christian Fisch, Angela Altmeier, Martin Obschonka, Michal Kosinski, Pin Ni

    cs.CL · cs.LG

    Entrepreneurial cognition is a foundation of entrepreneurship research. Yet the growing involvement of large language models (LLMs) in entrepreneurial work extends the cognition question beyond human actors to systems whose internal representations remain largely unexplored. We introduce artificial entrepreneurial cognition, the functional organisation of entrepreneurship-relevant representations and computations inside artificial...

    arxiv.org/abs/2609.15277 · PDF

  15. 15

    EMR: Self-Evolving Medical Multi-Agent System via Experience Mining and Reuse

    Dongsheng Shi, Yue Li, Xin Yi, Linlin Wang

    cs.CL · cs.AI

    Large language model (LLM) driven multi-agent systems have shown promise in complex clinical reasoning, yet existing approaches rely on static strategies and lack persistent clinical memory, preventing self-evolving from prior diagnostic successes and failures. We present EMR, a self-evolving medical multi-agent system via Experience Mining and Reuse. EMR introduces a hierarchical clinical experience library that organizes accumulated...

    arxiv.org/abs/2609.15161 · PDF

  16. 16

    MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup

    Muchen Li, Leonid Sigal, Renjie Liao

    cs.CL · cs.AI

    Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more recently, conditional memory: token-indexed embedding tables that augment the backbone with cheap parametric lookups. Existing memory-embedding methods retrieve via a deterministic function of the surface form, which collapses different contextual senses of the same token (e.g., python the language vs. the animal) into a...

    arxiv.org/abs/2609.15126 · PDF

  17. 17

    Translating the Translator: Decomposing the Cost of English-Forced Inter-Agent Communication

    Kushagra Agrawal, Yuming Feng, Man-Fai Leung

    cs.CL · cs.AI · cs.MA

    Multi-agent LLM architectures, such as LangChain and AutoGen, largely assume English as the lingua franca for internal inter-agent communication, even when the end-user task is non-English. We fill this gap by evaluating a two-agent extraction-answer core, with an additional back-translation agent in the English-forced condition, across four typologically diverse languages (Hindi, Chinese, Spanish, Arabic; n = 300 per language) using the...

    arxiv.org/abs/2609.15079 · PDF

  18. 18

    Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

    Zixiang Chen, Sufeng Niu, Yingchi Liu, Wenting Zhao, Akshara Prabhakar, Shubham Mehrotra, Bin Bi, Zhujun Lan,...

    cs.CL · cs.AI · cs.LG

    We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Salesforce Koa is trained on public and synthetically generated data, with no customer data, to improve tool use and agentic capabilities while preserving strong general-purpose performance. Its distinctive component is a...

    arxiv.org/abs/2609.15066 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.