cs.CL · 2026-08-31 · No. 101

Computation and Language, 2026-08-31.

16 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

16 entries
  1. 01

    NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry

    Samuel Xiao, Judy Song, Rory Hu, Ziliang Zong

    cs.CL · cs.AI

    Recent advances in large language models (LLMs) have demonstrated strong capabilities in natural language understanding and mathematical reasoning. However, their ability to translate informal mathematical problems into formal representations remains underexplored. This limitation is particularly important for neuro-symbolic geometry systems such as AlphaGeometry, whose theorem-proving engine requires inputs in a specialized domain-specific...

    arxiv.org/abs/2608.28481 · PDF

  2. 02

    Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents

    Nan Li

    cs.CL · cs.LG

    Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state across turns, interpret feedback, and choose valid actions under changing constraints. We study this setting in the LM Playschool Challenge with a 2B open-weight model, and find that many failures are not only broad knowledge failures but also local decision failures: repeated guesses, malformed actions, and violations of...

    arxiv.org/abs/2608.28458 · PDF

  3. 03

    Sliding-window beats linear attention

    Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, Emy Gervais

    cs.CL · cs.LG

    Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy. Every new token costs more than the previous one. For each additional token, the keys and values must be stored in memory indefinitely, which is unsustainable. Several alternatives have been proposed to fix the quadratic scaling problem, one of which is retrofitting LLMs to use Linear Attention. This idea has attracted a lot of...

    arxiv.org/abs/2608.28444 · PDF

  4. 04

    Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

    Qing Ye, Meng-Hsuan Lin

    cs.CL · cs.AI

    One model passed our fidelity check without ever opening the datasheet. We found it while qualifying models for an internal extraction service: a structured-output constraint had silently disabled tool use, and the model answered anyway, with fabricated source text. Only the per-tool trace exposed it. Fidelity -- whether an extracted value matches the source -- is the standard measure for agentic document extraction, and it scores that run a...

    arxiv.org/abs/2608.28439 · PDF

  5. 05

    Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL

    Jiayan Lin, Yujia Liu, Zijin Hong, Zheng Yuan, Yilin Xiao, Hao Chen, Qinggang Zhang, Xiao Huang, Feiran Huang

    cs.CL · cs.AI · cs.DB

    Recent advances in in-context learning (ICL) text-to-SQL have substantially improved execution accuracy on public benchmarks by assembling increasingly elaborate pipelines around the base generator, yet existing studies typically report aggregate end-to-end accuracy, without quantifying the marginal accuracy-cost contribution of individual design choices. Consequently, providing a unified, paradigm-level cost-accuracy quantification remains a...

    arxiv.org/abs/2608.28432 · PDF

  6. 06

    When Linguistic and Internal Confidence Diverge in Large Language Models

    Hefan Zhang, Bingquan Zhang, Ming Cheng, Saeed Hassanpour, Weicheng Ma, Soroush Vosoughi

    cs.CL · cs.AI

    Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study this question across 8 classification tasks, 2 generation tasks and 30 models from three families. For classification, we compare linguistic confidence with logits-based confidence along three axes: association, magnitude agreement and calibration. For generation,...

    arxiv.org/abs/2608.28382 · PDF

  7. 07

    BanglaMed-QA: A Question Answering System for Healthcare Support in Bangla

    Rowzatul Zannat, Abdullah Al Shafi, K. M. Azharul Hasan, Atia Shahnaz Ipa

    cs.CL · cs.AI · cs.LG

    Medical question answering (QA) systems have become crucial tools for providing reliable health information. But they remain very unexplored for low-resource languages like Bangla due to limited datasets and systems tailored to these languages. To address this, we introduce BanglaMed-QA, a robust QA system specifically designed for the Bangla medical domain. The process begins with building a structured medical knowledge base that includes...

    arxiv.org/abs/2608.28329 · PDF

  8. 08

    A Probabilistic Interpretation of KV Cache Eviction

    Renato Geh, Alex Chen, Daniel Israel, Aditya Grover, Guy Van den Broeck

    cs.CL · cs.AI

    The premise and promise of KV (cache) eviction is simple: higher throughput can be achieved by evicting some entries from the KV cache, at a negligible cost to quality. This holds empirically for many existing methods, though most rely on creative heuristics for selecting which entries to drop. Despite recent advances, the problem of KV eviction has remained informal in the literature. This paper aims to properly formalize this problem...

    arxiv.org/abs/2608.28293 · PDF

  9. 09

    Embedding Models for Stance-Aware Argument Retrieval

    Angelo Sparacino, Francesca Toni, Adam Dejl

    cs.CL · cs.AI

    In computational argumentation, obtaining arguments that explicitly support or attack given claims is a critical precursor to downstream reasoning tasks. When these supporting and attacking arguments are to be retrieved using semantic search methods, they need to be assessed for topic-relevance to the claims of interest as well as for correctness of their (positive or negative) stance towards the claims. In this paper we explore how dense...

    arxiv.org/abs/2608.28283 · PDF

  10. 10

    Text Restoration of Ancient Documents with Language Models

    Shibingfeng Zhang, Edoardo Caraffa, Annafelicia Zuffrano, Maddalena Modesti, Giovanni Colavizza

    cs.CL · cs.AI

    Purpose - This study investigates the feasibility of restoring missing text caused by physical lacunae in damaged ancient manuscripts using language models. Methodology - The study proposes different scenarios to replicate real-world conditions. Language models of different architectures are applied according to their suitability to each scenario. We also propose several decoding strategies that further enhance performance and address the...

    arxiv.org/abs/2608.28170 · PDF

  11. 11

    Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result

    Christos Koutsiaris

    cs.CL · cs.AI · cs.IR

    A byte-level BPE tokenizer is an ordered list of merge rules, so applying only a prefix yields a vocabulary whose token identifiers are the first rows of the full vocabulary. This prefix nesting allows one language model to operate at several vocabulary sizes, use a control token to indicate the active size, and be deployed at any trained size by slicing its embedding and output head. We pre-registered five claims, including margins, seeds,...

    arxiv.org/abs/2608.28151 · PDF

  12. 12

    SimpCue: Cue-Based Prompting for Multilingual Text Simplification

    Mehrzad Tareh, Horacio Saggion, Stefan Bott

    cs.CL · cs.AI

    Text simplification aims to make complex texts easier to understand while preserving their original meaning. Recent large language models can perform simplification through prompting, but it remains unclear whether adding explicit linguistic information about sentence complexity to the prompt improves their outputs. We investigate this question for multilingual sentence-level Easy-to-Read simplification in Catalan, Spanish, and Italian. Using...

    arxiv.org/abs/2608.28042 · PDF

  13. 13

    Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning

    Vy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He, Feng Xia, Renqiang Luo, Erik Cambria, Xiuzhen Zhang

    cs.CL · cs.AI · cs.LG

    Knowledge-intensive reasoning requires Large Language Models (LLMs) to ground answers in provided evidence. When evidence is insufficient, it is desirable that models abstain rather than confidently generating unsupported answers. Existing abstention methods rely on uncertainty estimation or evidence sufficiency checks, but neither tests whether the reasoning process for generation, driven by the interaction of provided evidence and the...

    arxiv.org/abs/2608.28018 · PDF

  14. 14

    LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages

    Injun Baek, HyeongSeok Lee, Yearim Kim, Junhoo Lee, Nojun Kwak

    cs.CL · cs.AI

    Landing pages are goal-oriented web interfaces that must communicate a target-specific value proposition while organizing information flow, visual hierarchy, and calls to action (CTA). Although large language models can generate plausible webpage code from natural-language prompts, direct generation often yields generic templates and unsupported persuasive claims. We study target-grounded, reference-guided landing-page generation, where a...

    arxiv.org/abs/2608.27902 · PDF

  15. 15

    OpenStamp: A Watermark for Open-Source Language Models

    Miroojin Bakshi, Saksham Rastogi, Danish Pruthi

    cs.CL · cs.AI · cs.LG

    With the growing prevalence of large language model (LLM) generated content, watermarking is considered a promising approach for attributing text to LLMs and distinguishing it from human-written content. A prominent class of techniques embeds subtle but detectable signals in generated text by modifying token sampling probabilities. However, such methods are unsuitable for open-source models, where users have white-box access and can easily...

    arxiv.org/abs/2608.27899 · PDF

  16. 16

    Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict

    Adarsh Sudheer, David Li, Omar Elbanna, Ishaan Kodarapu, Arjun Bahuguna, Vasu Sharma

    cs.CL · cs.AI

    We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2-7B-AV, three alignment configurations remain nearchance on the scored exact-string Yes/No subset of AVHBench, even though their output priors shift substantially. Similarly, off-the-shelf InternVideo2 experienced a...

    arxiv.org/abs/2608.27785 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.