cs.CL · 2026-07-11 · No. 50

Computation and Language, 2026-07-11.

22 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

22 entries
  1. 01

    Validity of LLMs as data annotators: AMALIA on authority

    Manuel Pita

    cs.CL · cs.AI · cs.CY

    A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size. Yet agreement is reliability, not validity....

    arxiv.org/abs/2607.08731 · PDF

  2. 02

    WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search

    Xiaoshuai Song, Liancheng Zhang, Kangzhi Zhao, Yutao Zhu, Zhongyuan Wang, Guanting Dong, Jinghan Yang, Han Li, Kun...

    cs.CL · cs.AI · cs.MA

    Large language model (LLM)-based web search agents are transforming information seeking from simple factoid question answering into complex, deep-and-wide search and research-oriented tasks. A single ReAct-style agent is constrained by one long trajectory and limited context, making it difficult to handle depth and coverage simultaneously. Existing multi-agent systems improve search coverage through parallel execution and aggregation, but...

    arxiv.org/abs/2607.08662 · PDF

  3. 03

    UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing

    Xinlong Zhao, Dongsheng Liu, Hengyu Zhao, Zixuan Fu, Zheng Wang, Jie Cai, Jie Zhou, Qiang Ma, Xuanhe Zhou, Xu Han,...

    cs.CL · cs.AI

    As available training data approaches its physical limit, gains from Scaling Laws have begun to diminish. Consequently, improving Large Language Models (LLMs) now depends less on data expansion and more on higher-quality data utilization. However, in the context of large-scale corpora, existing refinement methodologies face significant limitations in quality, efficiency, and reliability: Rule-based approaches are constrained by fixed...

    arxiv.org/abs/2607.08646 · PDF

  4. 04

    When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

    Zongyou Yang, Yinghan Hou, Xiaokun Yang

    cs.CL · cs.AI

    An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evaluator-replacement ambiguity as a measurement-validity problem. Across four judgment datasets, we compare two upgrade paths available in practice: scaling Qwen3 dense judges from 1.7B to 32B parameters and moving across MiniMax M2-M2.7 released APIs. The main pattern is that judge upgrades are not...

    arxiv.org/abs/2607.08535 · PDF

  5. 05

    Two Axes of LLM Abstention: Answer Correctness and Question Answerability

    Benedikt J. Wagner

    cs.CL · cs.AI

    A model should refuse two different things: answers it would get wrong, and questions it should not answer at all, such as unanswerable ones or ones resting on a false premise. The usual recipe thresholds a single confidence score, which cannot tell these apart. Across five instruction-tuned models from three families (2B to 14B), we find they are separate axes. Ordinary answer-confidence tracks whether an answer is right but is nearly blind...

    arxiv.org/abs/2607.08456 · PDF

  6. 06

    When Synthetic Speech Is All You Have: Better Call GRPO

    Shashi Kumar, Yanis Labrak, Hasindri Watawana, Sergio Burdisso, Esaú Villatoro-Tello, Kadri Hacioğlu, Petr Motlicek,...

    cs.CL · cs.AI

    LLM-based ASR adapted to regulated domains such as banking is bottlenecked by privacy: real speech is costly and legally constrained to collect, making synthetic text-to-speech (TTS) an attractive substitute. Yet synthetic speech stays acoustically mismatched with real recordings, and work on this gap has stayed within supervised fine-tuning (SFT). We instead turn to reinforcement learning, and show that Group Relative Policy Optimization...

    arxiv.org/abs/2607.08409 · PDF

  7. 07

    Prompt Compression via Activation Aggregation

    Thibaud Ardoin, Semira Einsele, Evis Bregu, Gerhard Wunder

    cs.CL · cs.LG

    Large language models process prompts by propagating activations through dozens of layers before generating a response. We ask whether the task-relevant information contained in an instruction prompt can be compressed into a single activation vector and re-injected into the model, replacing the original token sequence? We show this is achievable using a learned weighted sum of activations extracted at an intermediate layer and injected at an...

    arxiv.org/abs/2607.08399 · PDF

  8. 08

    Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition

    Jing Jie Tan, Ban-Hoe Kwan, Danny Wee-Kiat Ng, Yan-Chai Hum, Shih-Yu Lo, Po-An Chen, Noriyuki Kawarazaki, Kosuke...

    cs.CL · cs.AI · cs.HC · cs.RO · cs.SI

    Personality recognition has traditionally been constrained by theory-dependent formulations, where models are trained to fit predefined psychological taxonomies rather than uncovering shared underlying behavioral structure. This limits generalization, as personality itself is better understood as theory-invariant, while existing annotations reflect only partial and sometimes inconsistent views of the same latent traits. In this work, we...

    arxiv.org/abs/2607.08374 · PDF

  9. 09

    TypeProbe: Recovering Type Representations from Hidden States of Pre-trained Code Models

    Giuliano Gorgone, Fausto Carcassi

    cs.CL · cs.AI · cs.PL

    State-of-the-art code models achieve impressive performance, yet the extent to which they internally encode type information remains poorly understood. We probe the residual streams of pretrained code models for internal type representations using a parallel dataset of Java and Python code examples. Our results show that cross-lingual type representations emerge even from untyped code. Moreover, we test whether hidden states linearly encode...

    arxiv.org/abs/2607.08339 · PDF

  10. 10

    Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment

    Taehyung Yu, Seongjae Kang

    cs.CL · cs.AI · cs.LG · cs.SD

    Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from $N$ candidates with an automatic speech recognition (ASR) verifier. We identify an underexplored evaluation confound: a verifier's apparent quality depends strongly on which ASR family judges it. On LibriSpeech-PC test-clean~\citep{librispeechpc} with F5-TTS~\citep{f5tts}, verifier rankings reverse across Whisper, wav2vec~2.0, and HuBERT...

    arxiv.org/abs/2607.08256 · PDF

  11. 11

    LEXIC: Lightweight Eye-tracking eXtension via Injected Complexity

    Sumin Lee, Kyeonghun Kim, Subeen Lee, Jiwon Yang, Tien Nguyen, Ken Ying-Kai Liao, Nam-Joon Kim

    cs.CL · cs.AI · cs.HC · cs.LG

    On the recent EyeBench benchmark, predicting reading comprehension from eye movements exposes a stark gap: text-aware models using pretrained language models reach 56--63% AUROC, while gaze-only models operate at chance. We ask how far a gaze-only model can be pushed by lightweight, language-model-free conditioning. Building on the EyeBench AhnCNN baseline, LEXIC-Base, we propose two mechanisms to inject three precomputed word-level...

    arxiv.org/abs/2607.08152 · PDF

  12. 12

    ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents

    Maud Ehrmann, Emanuela Boros, Juri Opitz, Andrianos Michail, Florian Wagner, Simon Clematide

    cs.CL · cs.AI · cs.IR

    We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents. OCR post-correction remains a long-standing challenge in digital heritage: large-scale collections of digitized documents are affected by legacy OCR errors, while re-digitization at scale remains impractical. Large language models (LLMs) offers a major opportunity to revisit this challenge, yet their effectiveness...

    arxiv.org/abs/2607.08143 · PDF

  13. 13

    COBART: Controlled, Optimized, Bidirectional and Auto-Regressive Transformer for Ad Headline Generation

    Yashal Shakti Kanungo, Gyanendra Das, Pooja A, Sumit Negi

    cs.CL · cs.AI · cs.LG

    Online ads are essential to all businesses and ad headlines are one of their core creative component. Existing methods can generate headlines automatically and also optimize their click-through-rate (CTR) and quality. However, evolving ad formats and changing creative requirements make it difficult to generate optimized & customized headlines. We propose a novel method that uses prefix control tokens along with BART fine-tuning. It yields the...

    arxiv.org/abs/2607.08071 · PDF

  14. 14

    Holographic Neural PCFG for Unsupervised Parsing

    Ryosuke Yamaki, Daichi Mochihashi, Nobutaka Shimada, Tadahiro Taniguchi

    cs.CL · cs.LG

    Unsupervised constituency parsing aims to accurately induce latent tree structures from raw text alone. Recent neural parameterizations of PCFGs achieve strong performance in both supervised and unsupervised parsing, yet rely on high-capacity black-box networks for rule scoring -- as exemplified by the Neural PCFG family -- leaving rule probabilities without an interpretable mathematical form. In this paper, we propose Holographic Neural PCFG...

    arxiv.org/abs/2607.08063 · PDF

  15. 15

    What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

    Raphaël Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana, Christopher J. Earls, Srikar Varadaraj, Eric Ho

    cs.CL · cs.AI

    Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast. We ask whether internal representations offer a more direct window into both. Working with Eternis-Forecaster 8B on OpenForesight, we train representation-pooling probes on intermediate activations and find they achieve substantially better calibration; a...

    arxiv.org/abs/2607.08046 · PDF

  16. 16

    PLURAL: A Global Dataset for Value Alignment

    Dhruv Agarwal, Anya Shukla, Tanya Goyal, Aditya Vashistha

    cs.CL · cs.AI · cs.CY

    Large language models (LLMs) are used worldwide, yet disproportionately reflect Western values, limiting their ability to represent diverse value systems. We introduce PLURAL, a large-scale, value-focused preference dataset grounded in the Integrated Values Survey (IVS), a nationally representative survey spanning 92 countries. Using a two-stage generation pipeline, we transform survey responses into synthetic preference triplets that...

    arxiv.org/abs/2607.08034 · PDF

  17. 17

    Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention

    Ryota Kobayashi, Tsubasa Hirakawa, Takayoshi Yamashita, Hironobu Fujiyoshi, Yasunori Ishii, Tomoyuki Okuno, Kazuki Kozuka

    cs.CL · cs.AI

    This paper proposes an improved structured pruning method for large language models (LLMs) that addresses key challenges in adapting Adaptive Feature Retention (AFR), an unstructured pruning technique, to structured pruning. When applying AFR to structured pruning, three major problems arise: distribution mismatch between heterogeneous pruning scores, loss of sign information indicating optimization direction consistency, and influence of...

    arxiv.org/abs/2607.08027 · PDF

  18. 18

    Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework

    Riccardo Revalor, Jalees Rehman, Debjit Pal

    cs.CL · cs.AI

    Large-Language Models (LLMs) can be prone to flawed and unfaithful reasoning that decoding strategies like Self-Consistency (SC) fail to detect as they evaluate only final-answer agreement while ignoring the logical validity of intermediate steps. This raises three fundamental questions: How can we reliably quantify uncertainty in LLM reasoning? Can semantic, structural, and causal awareness select more faithful reasoning compared to naïve...

    arxiv.org/abs/2607.08017 · PDF

  19. 19

    Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems

    Kalle Kujanpää, Ning Liu, Shahnawaz Alam, Yeshwanth Reddy Sura, Tianyu Yang, Kristina Klinkner, Shervin Malmasi

    cs.CL · cs.LG · cs.SE

    Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference-time coding loop with an agentic tool-making pipeline that compiles repeated SOP steps into validated, versioned tools before deployment. The tool-maker grounds synthesis in the live environment as it collects execution traces, observes backend schemas and values, generates candidate tools,...

    arxiv.org/abs/2607.08010 · PDF

  20. 20

    Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator

    Shiping Yang, Shining Liang, Weihao Liu, Wenbiao Ding, Linjun Shou, Lu Cheng, Angel X. Chang

    cs.CL · cs.LG

    Identifying faithfulness hallucinations in LLM-generated outputs remains challenging due to the scarcity of high-quality annotated data. Recent work relies on advanced LLMs to synthesize training data, including rationales, labels, and hallucinated claims. However, these methods treat the generator as a static component, limiting iterative improvement of the detector. To address this limitation, we introduce Hallucination Self-Play (HSP), a...

    arxiv.org/abs/2607.07993 · PDF

  21. 21

    A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

    A. Sayyad, J. Emmons, S. Jones, T. Lin, H. Krishnan

    cs.CL · cs.AI · cs.SD · eess.AS

    We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.5 Flash, 3.5 Flash, and 3.1 Pro. Our primary evidence base uses Gemini 2.5 Flash as the ground-truth model, validated against three calibrated human raters on 209 stereo sessions, scored on 8 production dimensions: 152 full-duplex conversations...

    arxiv.org/abs/2607.07985 · PDF

  22. 22

    When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

    Xiuyi Lou, Zicheng Xu, Yu-Neng Chuang, Hoang Anh Duy Le, Zhaozhuo Xu, Guanchu Wang, Vladimir Braverman

    cs.CL · cs.AI · cs.LG

    Reinforcement learning (RL) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs). However, widely used critic-free RL methods rely on uniform credit assignment, broadcasting the same advantage to all tokens regardless of their differences. We identify a critical failure mode of this design, which we refer to as Positive-Credit Contamination: low-probability tail tokens that are contextually...

    arxiv.org/abs/2607.07976 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.