cs.CL · 2026-08-05 · No. 75

Computation and Language, 2026-08-05.

18 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

18 entries
  1. 01

    TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

    Changle Qu, Sunhao Dai, Hengyi Cai, Yuqi Zhou, Xinran Chen, Simon, Jun Xu

    cs.CL · cs.AI

    Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit assignment in long-horizon TIR scenarios. On-policy self-distillation offers denser signals through teacher branches with privileged context, but existing approaches typically derive such context from ground-truth...

    arxiv.org/abs/2608.04007 · PDF

  2. 02

    Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

    Jo-Ku Cheng, Nikolaos Aletras, Marco Valentino

    cs.CL · cs.AI · cs.LG

    Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretraining tasks, such as Dyck and procedural algorithms, rely on narrow primitives that fail to capture the expressive capacity of natural language. Moreover, prior studies remain restricted to relatively small token budgets, offering limited insight into skill emergence and representational dynamics. To...

    arxiv.org/abs/2608.03930 · PDF

  3. 03

    MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

    Martin Böckling, Elizaveta Nosova, Heiko Paulheim, Andreea Iana

    cs.CL · cs.AI · cs.IR

    Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logistics, yet large language models (LLMs) struggle with the required geometric and topological computation despite storing considerable geographic knowledge. Existing benchmarks localize these failures only partially: they are synthetic or smallscale, largely monolingual, and offer limited control...

    arxiv.org/abs/2608.03882 · PDF

  4. 04

    SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG

    Kaysarul Anas Apurba, Md. Hasibul Hasan, Rofiqul Alam Shehab, Asab Azad

    cs.CL · cs.AI · cs.IR · cs.PF

    We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific RAG pipeline across three corpus scales: 1,034 chunks (1K papers), 5,160 chunks (5K papers), and 15,480 chunks (15K papers). The pipeline combines sentence-window chunking, BM25, BGE-M3 dense retrieval, reciprocal rank fusion, optional...

    arxiv.org/abs/2608.03860 · PDF

  5. 05

    Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking

    Peijia Guo, Wenxuan Xie, ZiGuang Li, Ming Li

    cs.CL · cs.AI

    Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexplored and largely unresolved challenge. Prior work on LLM-generated-text detection targets AI involvement, which may be permissible, rather than source reuse, while similarity-based methods struggle after extensive rewriting and multi-source synthesis. Motivated by the description-length view of...

    arxiv.org/abs/2608.03859 · PDF

  6. 06

    Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

    Nathan Labiosa, David Buff, Ena Nayak, Erica Donno

    cs.CL · cs.LG

    When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate. Across a five-model panel we identify two...

    arxiv.org/abs/2608.03842 · PDF

  7. 07

    VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

    Andrei Chetvergov, Alexander Evseev, Timofei Sivoraksha, Stepan Ukolov, Mikhail Solovev, Danil Sazanakov, Sergey Bolovtsov

    cs.CL · cs.AI

    Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable. Existing work captures parts of this space through sentiment, favorability, and emotion benchmarks, but none combines...

    arxiv.org/abs/2608.03810 · PDF

  8. 08

    M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models

    Tomáš Burkert, Angelika Peljak-Łapińska, David Zelený

    cs.CL · cs.LG

    Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency. We introduce M-GATE (Multilingual Grammar, Accuracy in Translation, and Efficiency), a benchmark of linguistic proficiency spanning 30 typologically diverse languages from high- to low-resource. M-GATE...

    arxiv.org/abs/2608.03803 · PDF

  9. 09

    Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

    Bakbergen Ryskulov, Iker García-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene,...

    cs.CL · cs.AI · cs.LG

    Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD). This recovery step largely decides the final quality, yet it is expensive. We present a practitioner's study of how to make distillation training efficient, organised around two systems contributions. First,...

    arxiv.org/abs/2608.03796 · PDF

  10. 10

    MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models

    Tong Ling, Hang Lei, Feng Xiao, Changhui Sun, Jiahang Xie, Hao Liu, Lu Liu, Yanlong Du

    cs.CL · cs.AI

    Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundamentally from that of autoregressive (AR) models. Whereas AR decoding exposes a contiguous prefix, MDLM denoising produces dynamic, non-contiguous configurations of revealed and masked tokens. Conventional positional encodings such as RoPE capture sequence order and pairwise displacement but remain...

    arxiv.org/abs/2608.03769 · PDF

  11. 11

    GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models

    Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski

    cs.CL · cs.AI · cs.DB

    Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess no representation of entities, leading to duplicate entries as well as conflations. We propose GPTKB 2.0, a methodology for constructing disambiguated KBs directly from LLMs. GPTKB 2.0 incorporates...

    arxiv.org/abs/2608.03729 · PDF

  12. 12

    How Closely Do LLM Reviews Align with Human Peer Review?

    Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez

    cs.CL · cs.AI

    Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 topic-matched ICLR 2026 submissions,...

    arxiv.org/abs/2608.03659 · PDF

  13. 13

    Decoupling Generation and Selection for Budget-Constrained Faithful Summarization

    Zeyu Wang, Guanghua Wang, Meng Xu

    cs.CL · cs.AI

    Abstractive summarization models remain vulnerable to factual inconsistency, redundancy, and weak length control. We propose a modular generation-and-selection framework for sentence-budget-constrained summarization. A pretrained generator produces multiple candidate summaries, which are decomposed into sentence-level candidates. A combinatorial selector then constructs the final summary by balancing relevance, factuality, and redundancy...

    arxiv.org/abs/2608.03655 · PDF

  14. 14

    SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

    Kejian Zhu, Zhuoran Jin, Shangqing Tu, Hongbang Yuan, Yushi Bai, Kang Liu, Juanzi Li, Jun Zhao

    cs.CL · cs.LG

    Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage training, whereas RL enables stable coexistence across diverse tasks. Empirically, we trace this to the parameter level, observing that RL induces sparse and...

    arxiv.org/abs/2608.03573 · PDF

  15. 15

    ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels

    Gagan Bhatia, Julian Schlenker, Simone Paolo Ponzetto, Steffen Eger

    cs.CL · cs.AI

    Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages. We address this problem by asking how the magnitude and direction of change vary across linguistic levels, languages, and historical periods within a single analytical space. We introduce...

    arxiv.org/abs/2608.03507 · PDF

  16. 16

    Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension

    Raviraj Joshi, Utkarsh Vaidya, Sanjay Singh Chauhan, Niranjan Wartikar

    cs.CL · cs.LG

    Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token embeddings can strongly affect continued pre-training (CPT) efficiency. We present a systematic study of more than 20 initialization strategies for Hindi vocabulary extension in Nemotron-3-Nano-30B-A3B. Our comparison spans vocabulary-averaging baselines; external and learned initialization...

    arxiv.org/abs/2608.03494 · PDF

  17. 17

    Dynamically Allocating Evaluation Effort for Model Ranking

    Vilém Zouhar, Julia Kreutzer, Alon Lavie, Tom Kocmi, Matt Post, Ondřej Bojar, Mrinmaya Sachan

    cs.CL · cs.LG

    While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation protocols waste effort by exhaustively evaluating all models on the entire benchmark, a safe but inefficient approach. In this work, we formalize multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms,...

    arxiv.org/abs/2608.03437 · PDF

  18. 18

    FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact

    Alex Kwon

    cs.CL · cs.AI

    AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing away what made it checkable, who said it, how sure they were, when it held. We call that failure factwashing, and release factwash, an open-source write-time gate that catches it deterministically, with named flags and evidence rather than an LLM judge. Building it answers a practical question:...

    arxiv.org/abs/2608.03372 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.