cs.CL · 2026-08-24 · No. 94

Computation and Language, 2026-08-24.

18 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

18 entries
  1. 01

    EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering

    Xuanyu Meng, Jiashuo Sun, Jash Rajesh Parekh, Jiawei Han

    cs.CL · cs.AI · cs.DB · cs.IR

    Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their relationships. Existing retrieval-augmented generation (RAG) methods typically index documents as raw chunks and retrieve them through embedding similarity. Their performance degrades when chunk boundaries separate entities from supporting evidence or when a question requires multi-hop reasoning across the...

    arxiv.org/abs/2608.21252 · PDF

  2. 02

    No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation

    Dimitri Staufer, David Hartmann, Ibrahim Baroud

    cs.CL · cs.AI · cs.LG

    Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's evidential status is uncontrolled, measurements may conflate memorisation, retrieval, name priors and wrong-person attribution. We operationalise an unknown name as one with plausible First-Last form, no indexed full-name evidence, and no ambiguity signals under a documented validation run, and introduce...

    arxiv.org/abs/2608.21206 · PDF

  3. 03

    PromptResponse: Optimizing Prompts for LLM Coding Tasks

    Erik Thureck, Robert Kühnen, Tim Jacobowitz

    cs.CL · cs.AI · cs.HC · cs.SE

    Large language models (LLMs) are increasingly used in research workflows and software development pipelines, yet their output remains sensitive to input prompt variations. This paper presents $\unicode{x00AB}$PromptResponse$\unicode{x00BB}$, a controlled study examining how formatting and LLM-based tuning of coding task prompts affect the resulting code's performance, efficiency, and stability. Using five semantically identical yet...

    arxiv.org/abs/2608.21074 · PDF

  4. 04

    Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge

    Rishiraj Sengupta, Sotiris Chatzimiltis, Mohammad Shojafar, Xiatian Zhu

    cs.CL · cs.AI · cs.NI

    Real-world fault analysis in 5G and emerging 6G networks demands domain expertise to analyze free-text diagnostics, including root-cause explanations and recommended actions. LLMs have emerged as a promising approach to automating this, yet whether lightweight, edge-deployable models are capable of performing in-depth free-text diagnostics remains an open question. While existing benchmarks rely on restrictive MCQs with fixed answer keys,...

    arxiv.org/abs/2608.21021 · PDF

  5. 05

    Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models

    Zhen Yang, Sizai Hou, Kaiwen Zheng, Yaofang Liu, Liang He, Yixuan Chen, Kangning Cui

    cs.CL · cs.AI

    Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is rarely treated as a primary objective. We frame calibration-data selection for quantization as a target-dependent uncertainty-preservation problem. Different deployments emphasize different regions of the input distribution, yet prior work mainly optimizes accuracy-oriented compression metrics or...

    arxiv.org/abs/2608.21019 · PDF

  6. 06

    Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric

    Sami Shames El Deen, Mariette Awad

    cs.CL · cs.AI

    In this research, we introduce SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers for extractive summarization tasks. To ensure that the summaries generated by SAraBERT achieve a high coverage of the document's main ideas, we propose Semantic Siamese Similarity, a novel evaluation metric that measures the level of similarity between two text inputs. We validated using BLEU, ROUGE, and Semantic Siamese...

    arxiv.org/abs/2608.20964 · PDF

  7. 07

    Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

    Bakbergen Ryskulov, Iker García-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene,...

    cs.CL · cs.AI · cs.LG · cs.PF

    Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. The default recipe, quantization-aware training (QAT), re-fits the compressed, quantized model to hard labels; in our...

    arxiv.org/abs/2608.20953 · PDF

  8. 08

    MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation

    Ziwu Liu, Guozhong Li, Chen Qiu, Weiyang Kong, Panos Kalnis

    cs.CL · cs.AI

    Cross-model latent guidance lets a frozen large mentor encode an input once and a frozen small student generate from the resulting signal. Existing methods keep this signal fixed, assuming it stays useful as the output grows; we show this fails in long-form generation. On multi-turn instruction following, static guidance pushes a 4B student's constraint satisfaction 2.5 points below its no-guidance baseline; a training-free refresh every 16...

    arxiv.org/abs/2608.20927 · PDF

  9. 09

    Source-Free MT Evaluation Is Not MT Evaluation

    Baban Gain, Ramakrishna Appicharla, Asif Ekbal

    cs.CL · cs.AI

    Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments. As a result, source-free, reference-based evaluation has become the practical norm, even though it is unfaithful to the definition of translation adequacy and unfair to systems whose outputs preserve the source meaning while differing from the reference. This paper...

    arxiv.org/abs/2608.20925 · PDF

  10. 10

    KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs

    Xubin Chen, Yipeng Zhou, Wen Sun, Chengkai Huang, Xiaoming Fu, Quan Z. Sheng

    cs.CL · cs.AI

    Automatic Medical Coding (AMC), which assigns standardized International Classification of Diseases (ICD) codes to clinical notes, is essential for medical reimbursement, quality reporting, and clinical research. Existing pre-trained language model (PLM)-based methods typically formulate AMC as an extreme multi-label classification problem over a predefined code set, while recent large language model (LLM)-based approaches instead frame it as...

    arxiv.org/abs/2608.20887 · PDF

  11. 11

    SAC-Copula: Quality-Preserving Watermarking for Diffusion Language Models via Smooth Correlated Gumbel Fields

    Baixin Li, Haiyun He

    cs.CL · cs.CR · cs.LG

    Watermarking diffusion language models (DLMs) requires mechanisms compatible with iterative parallel unmasking rather than autoregressive decoding. Existing sampling-based watermarking methods typically inject position-wise i.i.d. perturbations, which can be poorly aligned with DLM decoding dynamics and degrade generation quality. We propose SAC-Copula, a quality-preserving watermarking method for DLMs based on smooth, locally correlated...

    arxiv.org/abs/2608.20839 · PDF

  12. 12

    STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction

    Tong Sun, Mingyang Ma, Jiayang Yu

    cs.CL · cs.AI

    Aspect-based sentiment analysis (ABSA) quadruple extraction requires jointly predicting target, aspect, opinion, and sentiment over reviews that often contain multiple fine-grained sentiment tuples. While large chain-of-thought (CoT) models perform well on this task, distilling them into smaller deployable models remains difficult. We identify a task-specific failure mode in distilled ABSA extraction: student errors at the target-aspect...

    arxiv.org/abs/2608.20831 · PDF

  13. 13

    Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation

    Yanglei Gan, Peng He, Run Lin, Peiyuan Jiang, Yifan Wang, Qiao Liu

    cs.CL · cs.AI

    Temporal Knowledge Graph (TKG) extrapolation seeks to infer future facts from time-varying relational histories. Recent diffusion-based approaches improve uncertainty modeling through generative denoising, but their aggregated conditioning on subject histories may insufficiently distinguish query-specific evidence from non-salient historical facts, thereby diluting target-discriminative signals. To bridge this gap, we propose FreqDiff, a...

    arxiv.org/abs/2608.20804 · PDF

  14. 14

    PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering

    Srikar Kashyap Pulipaka

    cs.CL · cs.AI · cs.LG

    We describe the PSK submission to the WMT 2026 Multilingual Instruction Shared Task. Our system uses the 3.35B-parameter Tiny Aya Global model with three QLoRA adapters, one for each task. The adapters are trained on multilingual document-summary pairs, passage-based question answering, and filtered standalone question answering. The summarization data also includes scientific papers with their author-written abstracts. On our held-out split,...

    arxiv.org/abs/2608.20757 · PDF

  15. 15

    MIL-BERT: Classification of Arbitrarily Large Text with Performance and Explanatory Guarantees

    John Cadigan, Dayne Freitag, Eric Yeh

    cs.CL · cs.LG

    Many text classification decisions are viable based on constituent excerpts alone. Taking inspiration from the field of multiple instance learning, we present an algorithm for training a neural network to classify text by selecting such excerpts. We show that our approach is also scalable with demonstrated learning against samples with nearly 1M tokens. We evaluate our methods on 7 datasets with emphasis on long-textual collections that far...

    arxiv.org/abs/2608.20636 · PDF

  16. 16

    AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale

    Minbyul Jeong, Chanwoong Yoon

    cs.CL · cs.AI

    Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthesized around predefined tasks and benchmarks. This task-centric paradigm makes it difficult to scale environments that reflect realistic and evolving workflows where diverse tasks can naturally emerge from the underlying world. We introduce AgentMercury, a scalable framework for synthesizing executable...

    arxiv.org/abs/2608.20634 · PDF

  17. 17

    When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation

    Lauren Pothuru

    cs.CL · cs.AI

    Agentic retrieval-augmented generation (RAG) interleaves retrieval, reasoning, and answer generation across multiple hops. A retrieval error at hop 1 can surface only as a wrong answer at hop 3, while later retrieval can also repair the trajectory. This paper introduces AgenticRAG-FP, an interventional benchmark for causal failure attribution in agentic RAG. The benchmark injects a certified fault at a specified hop, re-executes the...

    arxiv.org/abs/2608.20627 · PDF

  18. 18

    JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

    Tianxin Zhou, Ruixi Lin

    cs.CL · cs.AI · cs.LG

    Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared false-negative blind spots rather than independent evidence. We introduce JuryProbe, an empirical consensus-risk diagnostic for reference-free factuality judge panels, paired with a calibration-based routing policy....

    arxiv.org/abs/2608.20607 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.