cs.CL · 2026-07-25 · No. 64

Computation and Language, 2026-07-25.

12 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

12 entries
  1. 01

    Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it

    Federico Boggia

    cs.CL · cs.AI

    A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language models: epanorthosis, the self-correction of the specimen «This is not a course. It is a journey of transformation». This essay argues that the overuse is a trained disposition, driven mainly by a training distribution rich in promotional prose and by preference tuning (RLHF) that rewards confident,...

    arxiv.org/abs/2607.21498 · PDF

  2. 02

    What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations

    Piotr Wilam

    cs.CL · cs.LG

    Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-circuit extraction method to a 2x2 design -- Python and Rust crossed with Qwen2.5-Coder-7B and DeepSeek-Coder-V1-6.7B -- and measuring a complete inventory of grammatical concepts (58 Python, 57 Rust) identically in all four cells: the smallest design that separates what depends on the task,...

    arxiv.org/abs/2607.21491 · PDF

  3. 03

    RUMBA: Russian User Memory Benchmark

    Elizaveta Shevtsova, Inna Glebkina, Mark Baushenko, Pavel Gulyaev, Alena Fenogenova

    cs.CL · cs.AI

    The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely on aggregate retrieval metrics, failing to capture interactions between long-range context, temporal information, and reasoning. To address this, we introduce RUMBA (Russian User Memory BenchmArk) - a new benchmark for long-term conversational memory that provides a fine-grained taxonomy of memory-centric...

    arxiv.org/abs/2607.21447 · PDF

  4. 04

    Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models

    Renuka Oladri, Niveda Jawahar, Abdirisak Mohamed

    cs.CL · cs.AI · cs.LG

    Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate within a token budget (converged) or exhaust it without reaching a conclusion (non-converged). We characterize this phenomenon empirically, showing that converged generations achieve 90.3% accuracy on AIME 1983-2024 while non-converged ones achieve only 6.6%, with an overall convergence rate of 62.0%. We...

    arxiv.org/abs/2607.21433 · PDF

  5. 05

    Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

    Zhiheng Qian, Aini Li, Hai Hu, Liang Zhao

    cs.CL · cs.AI

    Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary. We trained a text-dependent GMM-HMM model (Chengdu-MFA) and fine-tuned a pretrained audio encoder on frame classification with Chengdu-MFA's...

    arxiv.org/abs/2607.21332 · PDF

  6. 06

    GRADRAG: Cross-Component Prompt Adaptation for Coordinated Multi-Agent RAG

    Paolo Pedinotti, Enrico Santus

    cs.CL · cs.AI

    Retrieval-Augmented Generation (RAG) systems increasingly employ multiple LLM agents. Yet, most prior work optimizes components in isolation rather than coordinating improvements across the pipeline. We introduce GRADRAG, a framework for cross-component prompt adaptation that models the RAG pipeline as a computational graph and propagates structured evaluation feedback to update upstream agents. An Evaluator critiques downstream answers and...

    arxiv.org/abs/2607.21324 · PDF

  7. 07

    Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs

    Yidu Wu, Xiang Wang, Kejie Zhao, Zhangchi Wang, Qinghai Guo, Xiaoying Tang

    cs.CL · cs.LG

    Large language models (LLMs) achieve strong generation and reasoning performance, but the Transformer architecture incurs high inference cost. Existing acceleration methods often rely on task-specific fine-tuning or training from scratch, increasing adaptation cost and limiting cross-task usability. We present an Adaptive Depth Sparse Framework (AdaDSF) that converts off-the-shelf pre-trained LLMs into depth-sparse models without full...

    arxiv.org/abs/2607.21291 · PDF

  8. 08

    A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset

    Katerina Papantoniou, Panagiotis Papadakos, Theodore Patkos, Dimitris Garefalakis, Nikos Vardakis, Dimitris Plexousakis

    cs.CL · cs.AI

    We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance judgments. We evaluate sparse (BM25), dense (sentence-transformers), hybrid, and LLM-assisted retrieval methods in this book-search setting. Multilingual embeddings outperform Greek-specific models, while hybrid retrieval performs best overall. A query-level analysis shows that BM25 excels at named-entity...

    arxiv.org/abs/2607.21274 · PDF

  9. 09

    slang.gr as a Large-Scale Crowdsourced Resource for Non-Standard Greek

    Panagiotis Papadakos, Katerina Papantoniou, Dimitris Plexousakis

    cs.CL · cs.AI

    Slang is a central component of everyday language, reflecting linguistic creativity, social identity, and cultural change, yet its dy- namic and non-standard nature makes it difficult to model computationally. We present the first large-scale computational study of slang.gr, a crowdsourced lexicon of Greek non-standard language, combining lexical content, user-generated tags, and interaction data. To enable the systematic analysis, we map...

    arxiv.org/abs/2607.21255 · PDF

  10. 10

    One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies

    Minh Ngoc Ta, My Anh Tran Nguyen, Duong D. Nguyen, Yuxia Wang, Preslav Nakov

    cs.CL · cs.AI

    Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to ask, when to stop, and when to answer. We introduce RegretBench, a multi-turn benchmark that evaluates clarification as policy behavior rather than isolated question quality. RegretBench provides a hidden-intent formulation of ambiguity, supports free-form interaction grounded in semantic-state...

    arxiv.org/abs/2607.21143 · PDF

  11. 11

    The Geometry of Personality: Activation Steering with Jungian Cognitive Functions

    Liu Zai, Yumeng Wang, Junchen Fu, Joemon M. Jose

    cs.CL · cs.AI · cs.LG

    Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as the Big Five. We investigate whether personality can instead be represented and controlled as a set of cognitive processes using the eight Jungian Cognitive Functions. To this end, we introduce a framework comprising a Jungian evaluation protocol and a dataset of over 2,100 role-playing...

    arxiv.org/abs/2607.20803 · PDF

  12. 12

    Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles

    Donghwan Kim

    cs.CL · cs.AI · cs.LG

    Majority voting over LLMs is widely assumed to benefit from diversity, and diversity measures are used to choose which models to combine. We ask whether five such measures track diversity or mainly re-express capability, auditing them as predictors of majority-vote gain over the best member across 31,900 subsets of 30 LLMs on MMLU-Pro (29 on TruthfulQA) under explicit capability controls. Three findings emerge. First, latent complementarity...

    arxiv.org/abs/2607.20768 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.