cs.CL · 2026-07-21 · No. 60

Computation and Language, 2026-07-21.

8 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

8 entries
  1. 01

    Automated Discovery Has No Universally Superior Harness

    Akshat Gupta, Jermaine Lei, Alexander Lu, Gopala Anumanchipalli, Leshem Choshen

    cs.CL · cs.AI

    Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these are composite systems combining several design choices about archives, parent selection, exploration, and budget allocation into a single recipe. Because discovery runs are expensive and inherently stochastic, existing harnesses are often compared using too few independent trials to distinguish key...

    arxiv.org/abs/2607.18235 · PDF

  2. 02

    PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning

    Hang Zhang, Warren J. Gross

    cs.CL · cs.LG

    Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstream performance. Many existing data selection methods rely on indirect heuristics, such as data quality, diversity or reasoning trace length. However, the effectiveness of these fixed criteria is task-dependent and difficult to generalize across diverse downstream...

    arxiv.org/abs/2607.18199 · PDF

  3. 03

    How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

    Prakhar Gupta, Terry Jingchen Zhang, Florent Draye, Bernhard Schölkopf, Zhijing Jin

    cs.CL · cs.AI · cs.LG

    Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model. Across five model families and seven BCT bias types, we extract a per-bias direction from hidden states...

    arxiv.org/abs/2607.18114 · PDF

  4. 04

    DeLIVeR: Decomposed Learning for Information-grounded Veracity Recognition via Reinforced Knowledge Graph Exploration

    Cong Hoan Nguyen, Thomas Hoang, Hieu Minh Duong, Long Nguyen

    cs.CL · cs.AI

    Automated fact-checking remains a challenge for Large Language Models (LLMs) due to "query brittleness" in traditional retrieval systems. We propose DeLIVeR (Decomposed Learning for Information-grounded Veracity Recognition), a framework that treats evidence retrieval as a reinforced strategic exploration task. DeLIVeR utilizes a Planner LLM to decompose complex claims into targeted question sets, which are used to traverse structured...

    arxiv.org/abs/2607.17935 · PDF

  5. 05

    Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI

    Bogdan Raduta, Horia Velicu, Alexandru Preda, Serban Chiricescu

    cs.CL · cs.AI

    Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output that is simply not true. The common response is to wait for a model that does not hallucinate. We argue that this is the wrong target. Large language models are, by construction, capable of generating unsupported text, and no amount of scale removes the possibility; a faithfulness judge bolted onto a raw...

    arxiv.org/abs/2607.17883 · PDF

  6. 06

    Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration

    Jie Hu

    cs.CL · cs.AI

    Test-time collaboration, including self-consistency, best-of-N selection, critic models, and verifier pipelines, is often credited with broadly improving LLM reasoning, yet its gains are uneven and sometimes negative. We ask when training-free collaboration should be expected to help. For a fixed candidate pool, we decompose a selector or verifier's net gain into measurable factors: recoverable mass, verification-signal coverage, conditional...

    arxiv.org/abs/2607.17531 · PDF

  7. 07

    Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

    Zitong Huang, Gustavo Lucas Carvalho, Deqing Fu, Robin Jia

    cs.CL · cs.LG

    We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, while avoiding the pitfalls that come with directly training the model to generate off-policy tokens. Experiments on document...

    arxiv.org/abs/2607.17524 · PDF

  8. 08

    Multilingual Sentence Embeddings for Linguistic-Integrated Reliability Audit

    Ummugul Bezirhan, Ji Yoon Jung, Matthias von Davier

    cs.CL · cs.AI

    Multilingual assessment systems commonly rely on translation for scoring and quality-control processes. We evaluate whether multilingual sentence embeddings can replace translated English input for Linguistic-Integrated Reliability Auditing (LiRA) across 11 PIRLS constructed-response items and three embedding models. Native-language embeddings reproduced translation-based reliability estimates closely while recovering responses excluded after...

    arxiv.org/abs/2607.17466 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.