cs.CL · 2026-05-22 · No. 7

Computation and Language, 2026-05-22.

15 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

15 entries
  1. 01

    Reducing Political Manipulation with Consistency Training

    Long Phan, Devin Kim, Alexander Pan, Alice Blair, Adam Khoja, Dan Hendrycks

    cs.CL · cs.AI

    Large language models (LLMs) exhibit systematic political bias across a variety of sensitive contexts. We find that LLMs handle counterpart topics from opposing political sides asymmetrically. We refer to this phenomenon as covert political bias and identify 7 categories of techniques through which it operates. We propose two metrics for covert bias: Sentiment Consistency measures symmetry in rhetoric and framing across paired political...

    arxiv.org/abs/2605.22771 · PDF

  2. 02

    Understanding Data Temporality Impact on Large Language Models Pre-training

    Pilchen Hippolyte, Fabre Romain, Signe Talla Franck, Perez Patrick, Grave Edouard

    cs.CL · cs.AI

    Large language models (LLMs) are typically trained on shuffled corpora, yielding models whose knowledge is frozen at train time and whose temporal grounding remains poorly understood. In this work, we study the impact of pre-training dynamics on the acquisition of time-sensitive factual knowledge, focusing specifically on data ordering. Our main contributions are twofold. First, we introduce a comprehensive benchmark of over 7,000 temporally...

    arxiv.org/abs/2605.22769 · PDF

  3. 03

    Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora

    Maciej Skorski

    cs.CL · cs.AI

    Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages. Idiomatic expressions, slang, and cultural references introduce hard-to-avoid translation artifacts. Yet automated moral values classification depends on language-specific annotated corpora that exist almost exclusively in English. We investigate whether LLM-based translation can bridge this gap, taking Polish as a test case. Using...

    arxiv.org/abs/2605.22660 · PDF

  4. 04

    More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts

    Víctor Yeste, Paolo Rosso

    cs.CL · cs.AI · cs.LG

    Detecting Schwartz values in political text is difficult because implicit cues often depend on surrounding arguments and fine-grained distinctions between neighboring values. We study when context and explicit moral knowledge help sentence-level value detection. Using the ValuesML/Touch{é} ValueEval format, we compare sentence, window, and full-document inputs; no-RAG and retrieval-augmented settings with a curated moral knowledge base;...

    arxiv.org/abs/2605.22641 · PDF

  5. 05

    Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents

    Asaf Yehudai, Lilach Eden, Michal Shmueli-Scheuer

    cs.CL · cs.AI

    Agentic systems are becoming more capable: agents define strategies, take actions, and interact with different environments. This autonomy poses serious challenges for overseeing and assessing agent behavior. Most current tools are limited, focusing on observability with basic evaluation capabilities or imposing static, hand-crafted error taxonomies that cannot adapt to new domains. To address this gap, we present Agentic CLEAR, an automatic,...

    arxiv.org/abs/2605.22608 · PDF

  6. 06

    Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion

    Meimingwei Li, Yuanhao Ding, Esteban Garces Arias, Christian Heumann

    cs.CL · cs.AI · stat.ML

    Recent work has identified a counterintuitive phenomenon termed "Hyperfitting", where fine-tuning Large Language Models (LLMs) to near-zero training loss on small datasets surprisingly enhances open-ended generation quality and mitigates repetition in greedy decoding. While effective, the underlying mechanism remains poorly understood, with the extremely low-entropy output distributions suggesting a potential equivalence to simple temperature...

    arxiv.org/abs/2605.22579 · PDF

  7. 07

    BeLink: Biomedical Entity Linking Meets Generative Re-Ranking

    Darya Shlyk, Stefano Montanelli, Lawrence Hunter

    cs.CL · cs.AI · cs.IR

    Despite recent progress, Biomedical Entity Linking (BEL) with large language models (LLMs) remains computationally inefficient and challenging to deploy in practical settings. In this work, we demonstrate that instruction-tuning of open-source generative models can offer an effective solution when applied at the re-ranking stage of the BEL pipeline. We propose a set-wise instruction-tuning formulation that enables fast and accurate candidate...

    arxiv.org/abs/2605.22501 · PDF

  8. 08

    From Correlation to Cause: A Five-Stage Methodology for Feature Analysis in Transformer Language Models

    Caleb Munigety

    cs.CL · cs.AI

    We propose a five-stage methodology for causal feature analysis in transformer language models (probe design, feature extraction, causal validation, robustness testing, and deployment integration) and demonstrate it end-to-end on GPT-2 small performing the Indirect Object Identification (IOI) task. Activation patching recovers the canonical IOI circuit (layer-9 head 9 alone gives recovery +1.02). A sparse autoencoder recovers per-name...

    arxiv.org/abs/2605.22462 · PDF

  9. 09

    DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QA

    Jianing Yin, Tan Tang

    cs.CL · cs.AI · cs.LG

    Large language model (LLM) agents still struggle with long-term memory question answering, where answer-supporting evidence is often scattered across long conversational histories and buried in substantial irrelevant content. Existing memory systems typically process memory before future queries are known, then retrieve the resulting units based on similarity rather than their utility for answering the query. This workflow leaves downstream...

    arxiv.org/abs/2605.22411 · PDF

  10. 10

    TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation

    Hanyu Guo, Jiedong Yang, Chao Chen, Longfei Xu, Kaikui Liu, Xiangxiang Chu

    cs.CL · cs.AI · cs.LG

    Public transit route planning traditionally depends on structured map infrastructure and complex routing engines, and no existing dataset supports training models to bypass this dependency. We present TransitLM, a large-scale dataset of over 13 million transit route planning records from four Chinese cities covering 120,845 stations and 13,666 lines, released as a continual pre-training corpus and benchmark data for three evaluation tasks...

    arxiv.org/abs/2605.22355 · PDF

  11. 11

    From TF-IDF to Transformers: A Comparative and Ensemble Approach to Sentiment Classification

    Dip Biswas Shanto, Mitali Yadav, Prajwal Panth, Suresh Chandra Satapathy

    cs.CL · cs.AI · cs.IR · cs.LG

    Sentiment analysis, also referred to as opinion mining, primarily tries to extract opinion from any text-based data. In the context of movie reviews and critics, sentimental analysis can be a helpful tool to predict whether a movie review is generally positive or negative. It can be difficult for the ML models to understand the context or metaphysical sentiment accurately, as ML models rely largely on statistical word representations. The...

    arxiv.org/abs/2605.22003 · PDF

  12. 12

    ACC: Compiling Agent Trajectories for Long-Context Training

    Qisheng Su, Zhen Fang, Shiting Huang, Yu Zeng, Yiming Zhao, Kou Shi, Ziao Zhang, Lin Chen, Zehui Chen, Lijun Wu, Feng Zhao

    cs.CL · cs.AI

    Recent development of agents has renewed demand for long-context reasoning capacity of LLMs. However, training LLMs for this capacity requires costly long-document curation or heuristic context synthesis. We observe that agents produce massive trajectories when solving problems, invoking tools and receiving environment observations across many turns. The evidence needed to answer the original question is thus scattered throughout these turns,...

    arxiv.org/abs/2605.21850 · PDF

  13. 13

    Comparing LLM and Fine-Tuned Model Performance on NVDRS Circumstance Extraction with Varying Prompt Complexity

    Geoffrey Martin, Xuan Zhong Feng, Yifan Peng

    cs.CL · cs.AI

    Suicide is a leading cause of death in the United States, and understanding the circumstances that precede it requires extracting structured information from death investigation narratives. Many of these circumstances require semantic inference beyond simple keyword matching. We develop a ``Complexity Score'' algorithm that analyzes coding manual structure to predict when detailed prompts with full coding guidelines improve over name-only...

    arxiv.org/abs/2605.21845 · PDF

  14. 14

    Does Slightly Mean Somewhat? Measuring Vague Intensity Words in LLM Numeric Actions

    Daniel Tabach

    cs.CL · cs.AI

    Do language models preserve the ordinal meaning of intensity words when those words must produce numeric actions? I study a researcher-constructed scale of 10 English degree modifiers, from slightly to drastically, informed by the Quirk et al. degree-modifier taxonomy, in a controlled resource-allocation environment where Claude Haiku receives a natural-language instruction, produces a numeric allocation, and a deterministic backend converts...

    arxiv.org/abs/2605.21827 · PDF

  15. 15

    Residual Skill Optimization for Text-to-SQL Ensembles

    Jiongli Zhu, Haoquan Guan, Parjanya Prajakta Prashant, Nikki Lijing Kuang, Seyedeh Baharan Khatami, Canwen Xu,...

    cs.CL · cs.AI · cs.DB · cs.LG

    Text-to-SQL ensembles improve over single-candidate generation by drawing multiple SQL candidates and selecting one, but their effectiveness is bounded by Pass@K, the probability that at least one of K candidates is correct. Existing methods source diversity heuristically through stochastic decoding or prompt variants, leaving candidate sets dominated by correlated failures. We present DivSkill-SQL, a residual skill optimization framework...

    arxiv.org/abs/2605.21792 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.