cs.CL · 2026-05-28 · No. 11

Computation and Language, 2026-05-28.

19 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

19 entries
  1. 01

    OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration

    Xinchen Zhang, Bowei Liu, Jiale Liu, Chufan Shi, Yizhen Zhang, Junhong Liu, Youliang Zhang, Zhiheng Li, Yujiu Yang, Ling Yang

    cs.CL · cs.AI · cs.CV · cs.LG

    Visual outcomes are increasingly central to multimodal large language models, making reliable and fine-grained verification essential for scaling generalist foundation models. In this work, we investigate multimodal meta-verification, which leverages verifier-generated rationales rather than decision-only signals, and explore how to effectively incorporate meta-verification feedback into multimodal verifier training. We identify two key...

    arxiv.org/abs/2605.28805 · PDF

  2. 02

    Skill-Conditioned Gated Self-Distillation for LLM Reasoning

    Jiazhen Huang, Xiao Chen, Xiao Luo, Yong Dai, Senkang Hu, Yuzhi Zhao

    cs.CL · cs.AI

    On-policy self-distillation (SD) improves LLM reasoning by using teacher-side privileged information (PI) to turn sparse verifier outcomes into dense token-level supervision. Existing methods usually assume trusted PI, such as reference answers or successful traces. We ask whether PI can instead come from an experience-derived skill bank, where retrieved skills are compact and reusable but may also be irrelevant or misleading. We propose...

    arxiv.org/abs/2605.28791 · PDF

  3. 03

    Rethinking Memory as Continuously Evolving Connectivity

    Jizhan Fang, Buqiang Xu, Zhixian Wang, Haoliang Cao, Xinle Deng, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu,...

    cs.CL · cs.AI · cs.LG · cs.MA · cs.MM

    Existing memory-augmented LLM agents often treat memory as a static repository with pre-defined representations and fixed retrieval pipelines, which is brittle in dynamic agentic environments where feedback, task variation, and heterogeneous signals continuously reshape what should be remembered and how it should be connected. To address this, we propose FluxMem, a connectivity-evolving memory framework that models memory as a heterogeneous...

    arxiv.org/abs/2605.28773 · PDF

  4. 04

    Reverse Probing: Supervised Token-level Uncertainty Quantification for Large Language Models in Clinical Text

    Bushi Xiao, Sarvesh Soni, Daisy Zhe Wang

    cs.CL · cs.AI

    As large language models are increasingly deployed for clinical text, ensuring they can reliably signal their own uncertainty becomes critical. Most existing uncertainty quantification (UQ) methods are designed for open-domain generation and cannot localize uncertainty at the token or span level in long clinical text. We propose Reverse Probing, the first UQ framework specialized for clinical summarization, which estimates token-level...

    arxiv.org/abs/2605.28740 · PDF

  5. 05

    MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems

    Xinle Deng, Ruobin Zhong, Hujin Peng, Xiaoben Lu, Yanzhe Wu, Guang Li, Buqiang Xu, Yunzhi Yao, Jizhan Fang, Haoliang...

    cs.CL · cs.AI · cs.LG

    Memory is essential for enabling large language models to support long-horizon reasoning, yet existing memory systems remain unreliable and difficult to debug. Tracing memory's dynamic evolution is crucial to understand how information is synthesized, propagated, or corrupted over time. In this work, we study the new problem of error tracing and attribution in LLM memory systems. We propose a novel framework that transforms memory pipelines...

    arxiv.org/abs/2605.28732 · PDF

  6. 06

    IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents

    Michael Galarnyk, Siddharth Lohani, Vidhyakshaya Kannan, Sagnik Nandi, Aman Patel, Liqin Ye, Arnav Hiray, Rutwik...

    cs.CL · cs.AI

    An Initial Public Offering (IPO) filing is a document released when a private firm goes public, allowing individual (retail) investors to purchase its shares. These filings describe a firm's business, financials, and risks and are long, multimodal documents with narrative text and images. Despite their importance to financial markets, there is no large-scale, standardized dataset or benchmark for studying IPO filings with modern language and...

    arxiv.org/abs/2605.28714 · PDF

  7. 07

    Towards Reliable Multilingual LLMs-as-a-Judge: An Empirical Study

    Irune Zubiaga, Aitor Soroa, Rodrigo Agerri

    cs.CL · cs.AI

    Large language models (LLMs) are increasingly used for the automatic evaluation of generated text, yet most prior work focuses on English. Despite the growing demand for multilingual evaluation, extending LLM-based evaluators to multilingual settings remains challenging, particularly for low-resource languages and scenarios where in-domain data is scarce. This work explores several strategies for developing multilingual LLMs-as-a-judge,...

    arxiv.org/abs/2605.28710 · PDF

  8. 08

    Sense Representations Are Inducible Interfaces

    Jan Christian Blaise Cruz, Alham Fikri Aji

    cs.CL · cs.AI

    Sense representations (explicit, per-token meaning decompositions) are useful for disambiguation, steering, and cross-lingual alignment, but existing approaches require models to be pretrained with sense structure baked in. We introduce ACROS, which induces an explicit sense pathway into a frozen pretrained decoder LM through a gated residual addition. On SmolLM2-360M, ACROS preserves base LM quality while supporting three uses of the same...

    arxiv.org/abs/2605.28669 · PDF

  9. 09

    The Attentional White Bear Effect in Transformer Language Models

    Rebecca Ramnauth, Brian Scassellati

    cs.CL · cs.AI

    Instruction-based suppression is widely used to prevent language models from generating prohibited content, yet it remains unclear whether suppression reduces internal representation or merely suppresses expression. We investigate this question through representational probing, attention analysis, and behavioral semantic leakage experiments across multiple transformer models. We find that prohibited concepts remain highly recoverable from...

    arxiv.org/abs/2605.28639 · PDF

  10. 10

    Measuring Form and Function in Language Models

    Héctor Javier Vázquez Martínez, Charles Yang

    cs.CL · cs.AI

    We introduce quantitative metrics for child language acquisition to evaluate language models. Our focus is on the formal syntactic and functional discourse properties of determiners in English, which young children acquire early and accurately. We propose Contextual Alternative Choice (CAC), a new prompting method which provides targeted tests for both syntactic and discourse knowledge of language. The method enables direct comparison of...

    arxiv.org/abs/2605.28616 · PDF

  11. 11

    Evaluating the Realism of LLM-powered Social Agents: A Case Study of Reactions to Spanish Online News

    Alejandro Buitrago López, Alberto Ortega Pastor, Javier Pastor-Galindo, José A. Ruipérez-Valiente

    cs.CL · cs.AI

    LLM-powered social agents are increasingly used to simulate online social behavior, yet their realism remains difficult to validate. Existing work has largely relied on general-purpose benchmarks, while less attention has been paid to short, reactive discourse such as audience replies to online news. In this paper, we evaluate whether LLM-generated reactions to Spanish online news reproduce measurable properties of real audience discourse....

    arxiv.org/abs/2605.28598 · PDF

  12. 12

    Models That Know How Evaluations Are Designed Score Safer

    Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi

    cs.CL · cs.AI

    The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such as hypothetical scenarios, as a source of verbalized evaluation awareness and subsequent behavioral shift. In this paper, we investigate a potential explanation of this phenomenon: evaluation meta-knowledge, defined as parametric knowledge about the structural traits...

    arxiv.org/abs/2605.28591 · PDF

  13. 13

    Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards

    Saurabh Dash, Pierre Clavier, John Dang, Matthias Galle, Marzieh Fadaee, Ahmet Üstün, Beyza Ermis

    cs.CL · cs.LG

    Reinforcement Learning from Verifiable Rewards (RLVR) has improved language models in domains such as mathematics and code, where correctness can be checked automatically. However, many important tasks are only partially verifiable: prompts contain multiple requirements, responses may satisfy some but not all of them, or no single reference answer might exist. We introduce Soft-RLVR, a framework for reinforcement learning from decomposed,...

    arxiv.org/abs/2605.28561 · PDF

  14. 14

    Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification

    Dylan Bouchard, Mohit Singh Chauhan, Zeya Ahmad, Ho-Kyeong Ra

    cs.CL · cs.AI · cs.LG

    Large language models have shown impressive capabilities in code generation, yet they often produce functionally incorrect code. Uncertainty quantification (UQ) methods have emerged as a promising approach for detecting hallucinations in natural language generation, but their effectiveness for code generation tasks remains underexplored. We systematically evaluate how UQ techniques transfer to code generation across three programming...

    arxiv.org/abs/2605.28500 · PDF

  15. 15

    The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment

    Junyu Lu, Qi Wei, Peishuo Zheng, Jie Zhang, Hui Huang, Qianru Wang, Chuan Xiao, Jianbin Qin, Shuyuan Zheng

    cs.CL · cs.AI

    Legal Judgment Prediction (LJP) has become a core benchmark for evaluating AI in the criminal legal domain, but it only sees criminal cases that have already passed prosecutorial review and been formally indicted. As a result, LJP leaves a substantial blind spot in assessing criminal liability, overlooking cases involving insufficient evidence, no criminal liability, or guilt exempted from punishment. To fill this gap, we propose...

    arxiv.org/abs/2605.28464 · PDF

  16. 16

    AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates

    Shaolong Chen, Madalina Ciobanu, Qingqing Mao, Ritankar Das

    cs.CL · cs.LG

    DPO has become a widely adopted alternative to RLHF for aligning LLMs with human preferences, eliminating the need for a separate reward model or RL loop. Recent theoretical analysis uncovers an asymmetric gradient behavior in DPO: the loss suppresses dispreferred responses substantially faster than it promotes preferred ones, causing the model to learn to avoid bad answers rather than to generate good ones. We propose AdaDPO, a Self-Adaptive...

    arxiv.org/abs/2605.28440 · PDF

  17. 17

    Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models

    Guanzhi Deng, Kuan Wu, Haibo Wang, Shing Yin Wong, Sichun Luo, Linqi Song

    cs.CL · cs.AI

    Mixture-of-Experts (MoE) models have emerged as a dominant paradigm for efficient LLM scaling, yet adapting them to non-English downstream tasks remains challenging. Existing fine-tuning approaches treat MoE models as monolithic learners, ignoring the heterogeneous routing structure that develops during pretraining. We validate across multiple MoE models and downstream tasks that middle layers form a language-universal alignment zone where...

    arxiv.org/abs/2605.28306 · PDF

  18. 18

    Revisiting Anthropomorphic Reflection Markers in Large Language Model Reasoning

    Yahan Yu, Noa Nakanishi, Fei Cheng

    cs.CL · cs.AI

    Large Language Models (LLMs) often produce explicit reflective traces during complex reasoning, accompanied by anthropomorphic markers such as wait, hmm, and alternatively. Although these markers are commonly used as visible indicators of reflection, their mechanisms remain unclear, which leaves the risk of overthinking associated with redundant and repetitive reflection markers. In this work, we revisit anthropomorphic reflection markers,...

    arxiv.org/abs/2605.28305 · PDF

  19. 19

    PrunePath: Towards Highly Structured Sparse Language Models

    Zhexuan Gu, Zixun Fu, Yancheng Yuan

    cs.CL · cs.AI

    Feed-forward networks (FFNs) dominate the parameter count and computation of modern language models, yet existing pruning methods often struggle to convert sparsity into hardware-friendly inference efficiency gains. We introduce \textbf{PrunePath}, a budget-adaptive structured sparsification framework for FFN layers. Built on MoEfication, PrunePath replaces independent expert-wise thresholding with a softmax-normalized routing distribution...

    arxiv.org/abs/2605.28283 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.