cs.CL · 2026-06-05 · No. 14

Computation and Language, 2026-06-05.

19 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

19 entries
  1. 01

    Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection

    Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tianjun Yao, Xinyi Shang, Yi Tang, Jiacheng Cui, Ahmed Elhagry,...

    cs.CL · cs.AI · cs.LG

    As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-written or AI-generated, but instead result from progressive human-AI co-editing. However, existing AI-text detection benchmarks largely focus on final outputs and provide limited understanding of how AI authorship signals emerge, accumulate, or disappear throughout the revision process. We...

    arxiv.org/abs/2606.06481 · PDF

  2. 02

    Self-Augmenting Retrieval for Diffusion Language Models

    Paul Jünger, Justin Lovelace, Linxi Zhao, Dongyoung Go, Kilian Q. Weinberger

    cs.CL · cs.AI · cs.LG

    Discrete diffusion language models generate text by iteratively denoising an entire response in parallel. At each step, they predict tentative tokens for every masked position, committing the confident predictions to the output and discarding the unconfident ones. We show that the discarded tokens are in fact a useful lookahead signal for retrieval-augmented generation: even low-confidence tokens often surface salient entities early in the...

    arxiv.org/abs/2606.06474 · PDF

  3. 03

    You Only Index Once: Cross-Layer Sparse Attention with Shared Routing

    Yutao Sun, Yanqi Zhang, Li Dong, Jianyong Wang, Furu Wei

    cs.CL · cs.AI · cs.LG

    Long-context inference in modern LLMs is increasingly constrained by decoding efficiency, especially in reasoning-heavy settings where models generate long intermediate chains of thought. Existing sparse attention methods often face a practical efficiency-quality trade-off. Structured block sparse methods typically provide stronger acceleration but incur noticeable quality loss, while token sparse methods are usually more accurate yet deliver...

    arxiv.org/abs/2606.06467 · PDF

  4. 04

    Latent Reasoning with Normalizing Flows

    Guancheng Tu, Xiangjun Fu, Suhao Yu, Yao Tang, Haoqiang Kang, Lianhui Qin, Yizhe Zhang, Jiatao Gu

    cs.CL · cs.LG

    Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. However, textual CoT forces this computation through a discrete, serial, and communication-oriented token stream: each reasoning step must be verbalized before the model can proceed, even when the underlying update is semantic, uncertain, or only partially formed. Latent reasoning offers a...

    arxiv.org/abs/2606.06447 · PDF

  5. 05

    Emergent Language as an Approach to Conscious AI

    Zengqing Wu, Chuan Xiao

    cs.CL · cs.AI · cs.MA · cs.NE

    The question of whether artificial systems can be conscious remains open, in part because existing approaches either evaluate systems against theory-derived checklists (discriminative) or engineer consciousness-inspired modules directly (architectural); both leave open whether observed structures are artifacts of human language priors. We propose a generative methodology: emergent language (EL) in multi-agent reinforcement learning, where...

    arxiv.org/abs/2606.06380 · PDF

  6. 06

    LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs

    Gianluca Barmina, Peter Schneider-Kamp, Lukas Galke Poech

    cs.CL · cs.AI

    Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather than whether they do so under ordinary use. We introduce PropMe, a propensity-aware framework for memorization evaluation that contrasts prefix-based capability attacks with non-adversarial evaluations. We propose a metric transformation that, applied to existing functions, allows to create...

    arxiv.org/abs/2606.06286 · PDF

  7. 07

    Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

    AJ Carl P. Dy, Aivin V. Solatorio

    cs.CL · cs.AI · cs.CV · cs.IR

    Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables. Current approaches for extracting visual content from documents are largely built around generic document layout analysis, where figures and tables are treated as uniformly relevant document objects rather than semantically meaningful analytical artifacts. In this work, we introduce a benchmark dataset and...

    arxiv.org/abs/2606.06242 · PDF

  8. 08

    Dense Contexts Are Hard Contexts: Lexical Density Limits Effective Context in LLMs

    Giovanni Dettori, Matteo Boffa, Danilo Giordano, Idilio Drago, Marco Mellia

    cs.CL · cs.AI

    Input length and the position of relevant information are widely cited as the primary causes of degraded LLM long-context performance. Here, we study lexical density -- the rate at which a context introduces distinct information -- as a third, largely overlooked factor that systematically reduces the effective context window of LLMs. We quantify the impact of lexical density on open-weight LLMs (9B-685B) using three "find-the-needle" style...

    arxiv.org/abs/2606.06203 · PDF

  9. 09

    Improving Answer Extraction in Context-based Question Answering Systems Using LLMs

    Hafez Abdelghaffar, Ahmed Alansary, Ali Hamdi

    cs.CL · cs.AI

    Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs). However, they still face challenges in accurately extracting and generating precise answers from given contexts, particularly when dealing with complex or ambiguous queries. Existing approaches often struggle with contextual understanding, answer consistency, and generalization across diverse domains. In this work, we propose a...

    arxiv.org/abs/2606.06197 · PDF

  10. 10

    Harnessing Structural Context for Entity Alignment Foundation Models

    Xingyu Chen, Yuanning Cui, Zequn Sun, Wei Hu

    cs.CL · cs.AI

    Entity alignment (EA) aims to identify equivalent entities across heterogeneous knowledge graphs (KGs) and is a key component of knowledge fusion and cross-KG reasoning. The recent EA foundation model demonstrates that alignment knowledge, once pretrained, can be directly applied to diverse previously unseen KG pairs. However, it still underuses structural context in two places: cross-KG interaction is weak during encoding, and final...

    arxiv.org/abs/2606.06109 · PDF

  11. 11

    IR3DE: A Linear Router for Large Language Models

    Eros Fanì, Oğuzhan Ersoy

    cs.CL · cs.LG

    Foundational Large Language Models (LLMs) demonstrate proficiency on a wide range of general tasks, and achieve remarkable results on various specialized tasks via domain-expert LLMs. With the ever-growing list of available LLMs, inference routers are being proposed to select the most appropriate LLM for each prompt. However, existing routing methods either optimize cost across weak-to-strong generalist LLMs or require substantial training to...

    arxiv.org/abs/2606.06098 · PDF

  12. 12

    LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents

    Aofan Yu, Chenyu Zhou, Tianyi Xu, Zihan Guo, Rong Shan, Zhihui Fu, Jun Wang, Weiwen Liu, Yong Yu, Weinan Zhang, Jianghao Lin

    cs.CL · cs.AI

    Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substantial context overhead and exposes skill content as plaintext. We present LatentSkill, a framework that converts textual skills into plug-and-play LoRA adapters through a pretrained hypernetwork. LatentSkill stores skill knowledge in weight space rather than context space, removing per-step...

    arxiv.org/abs/2606.06087 · PDF

  13. 13

    EGTR-Review: Efficient Evidence-Grounded Scientific Peer Review Generation via Multi-Agent Teacher Distillation

    Xinpeng Qiu, Wang Yihu, Zhifeng Liu, Xiaochen Wang, Jimin Wang

    cs.CL · cs.AI

    Scientific peer review generation has attracted increasing attention for reducing reviewing burdens and providing timely feedback. However, existing Large Language Model (LLM)-based methods often produce generic comments with insufficient evidence support and weak source traceability, while complex multi-agent systems incur high inference costs. To address these challenges, we propose EGTR-Review, an Evidence-Grounded and Traceable Review...

    arxiv.org/abs/2606.06025 · PDF

  14. 14

    Measuring the sensitivity of LLM-based structured extraction to prompt, model, and schema choices in clinical discharge summaries

    Martin Murin

    cs.CL · cs.AI · cs.LG

    Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream configuration choices is less understood than their accuracy on fixed benchmarks. This work measures that sensitivity without human-annotated ground truth, by holding the extraction task fixed and varying one choice at a time. The fixed schema comprises 17 clinical documentation flags on a...

    arxiv.org/abs/2606.05970 · PDF

  15. 15

    To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection

    Erfan Loweimi, Mengjie Qian, Kate Knill, Guanfeng Wu, Chi-Ho Chan, Abbas Haider, Muhammad Awan, Josef Kittler, Hui...

    cs.CL · cs.AI · cs.CV · cs.IR · cs.LG · cs.MM · eess.AS

    When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curated benchmarks, a target may be heard but unseen, seen but unheard, or both. Fusing scores from an absent modality injects noise, degrading precision below the best unimodal system. We propose a query-adaptive framework that detects active modalities via cross-modal score consistency: when both...

    arxiv.org/abs/2606.05931 · PDF

  16. 16

    Better Literary Translation: A Multi-Aspect Data Generation and LLM Training Approach

    Zhihao Lin, Ziqi Zhu, Hao Huang, Guanghui Wang, Peiyang He

    cs.CL · cs.AI

    Literary translation poses unique challenges due to the scarcity of high-quality annotated data and the need to balance expression fluency with literary effect. We present a multi-aspect iterative refinement framework that generates high-quality translation references and preference data through specialized LLM translators, each targeting a distinct quality dimension. We leverage the generated data for supervised fine-tuning and reinforcement...

    arxiv.org/abs/2606.05924 · PDF

  17. 17

    Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)

    Christopher J. Wedge, Joshua Stutter, Danny Dixon, Jacek Cała

    cs.CL · cs.AI

    Large language models (LLMs) have fundamentally transformed the landscape of Natural Language Processing. Despite these advances, LLMs and LLM-based systems remain prone to a variety of failure modes. Retrieval-augmented generation (RAG) systems have emerged as a common deployment scenario seeking to both avoid the well known risk of the LLM "hallucinating" information, and to enable reasoning and question answering over proprietary...

    arxiv.org/abs/2606.05901 · PDF

  18. 18

    Representing Research Attention as Contextually Structured Flows

    Jessica Rodrigues, Angelo Salatino, Gard Jenset, Scott Hale

    cs.CL · cs.LG

    Research attention is widely used as an indicator of visibility, influence, and societal uptake, yet it is typically represented as aggregated counts that do not preserve how attention develops across contexts over time. This creates a mismatch between how attention is interpreted and how it is represented. We propose attention flows as contextually structured representations that encode the organisation of attention and its evolution over...

    arxiv.org/abs/2606.05895 · PDF

  19. 19

    Staying with the Uncertainty: Uncertainty-Scaffolding Strategies for Artificial Moral Advisors in LLM-to-LLM Simulated Conversations

    Salvatore Greco, Hainiu Xu, Jacopo Domenicucci, Yulan He, Sylvie Delacroix

    cs.CL · cs.AI

    LLMs are increasingly deployed as Artificial Moral Advisors (AMA) in a variety of contexts: what kind of conversational patterns should they display? In this paper, we study how AMA can help their interlocutors "stay with the uncertainty". We propose three modes of uncertainty (Perspective-Multiplying, Tension-Preserving, Process-Reflecting) and compare them against three control conditions (Baseline, Persuasive, Sycophantic). A user-agent...

    arxiv.org/abs/2606.05890 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.