cs.CL · 2026-07-20 · No. 59

Computation and Language, 2026-07-20.

17 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

17 entries
  1. 01

    ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning

    Binglin Zhou, Peng Shi, Ryo Kamoi, Nan Zhang, Rui Zhang

    cs.CL · cs.AI

    Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and textual context. However, existing methods often fail because they struggle to locate decisive visual evidence, accurately read structured scientific visuals, and integrate multimodal observations into reliable reasoning. We introduce ToolSciVer, the first...

    arxiv.org/abs/2607.16131 · PDF

  2. 02

    Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

    Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani, Mitch Weiss

    cs.CL · cs.AI

    Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under...

    arxiv.org/abs/2607.16057 · PDF

  3. 03

    Loop the Loopies!

    Zitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang, Ran Tao, Joey Zhou, Bryan Dai

    cs.CL · cs.AI

    We present Loopie, the most powerful looped Transformer to date. The Loopie series consists of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6Bparameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N-fold increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N times. Loopie addresses this...

    arxiv.org/abs/2607.16051 · PDF

  4. 04

    Candidate Attended Dialogue State Tracking Using BERT

    Junyuan Zheng, Onkar Salvi, John Chan

    cs.CL · cs.AI

    Dialogue state tracking (DST) is one of the core components in task-oriented dialogue systems. At each turn in a conversation, DST estimates the user belief or dialogue state, which is used as input for downstream modules to predict system actions and generate responses. The increasingly popular dialogue system applications like Google Assistant, Siri and Alexa need to support a large number of services and APIs, resulting in growing...

    arxiv.org/abs/2607.16021 · PDF

  5. 05

    Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models

    Andy Catruna, Emilian Radoi

    cs.CL · cs.AI · cs.LG

    While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is known about diffusion language models (DLMs), an emerging alternative that generates text by iterative denoising. In this work, we study how DLMs implement induction, a mechanism behind in-context learning in which the model finds a repeated context and copies the token that followed it. Our analysis compares attention-only AR models...

    arxiv.org/abs/2607.15893 · PDF

  6. 06

    DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods

    Jens Frankenreiter

    cs.CL · cs.AI · cs.CY

    Much empirical legal research depends on translating unstructured text into structured variables. In corporate governance research as elsewhere, this translation has traditionally relied on human coding of documents such as charters and bylaws, a process that is costly, difficult to scale, and often opaque. This paper introduces DECODEM, a set of benchmark datasets for evaluating the automated extraction of corporate governance variables from...

    arxiv.org/abs/2607.15879 · PDF

  7. 07

    Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection

    Indraveni Chebolu, Rohan Singh, Arnab Mallick, Harmesh Rana

    cs.CL · cs.AI

    Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch. We study the \emph{conditional reliability} of toxicity priors in Indian multilingual and code-mixed short text: English toxicity, Indic abuse, and rule-based severity cues can be useful evidence, but only in some linguistic and abuse-severity contexts. We propose ToxGate, a...

    arxiv.org/abs/2607.15861 · PDF

  8. 08

    CAMMAR: Culture-Aware Matryoshka for Metaphorical Arabic Representations

    Suzan Awinat, Alfonso Ortega del Puente

    cs.CL · cs.AI

    Metaphor in Arabic is a culturally grounded mechanism for constructing meaning, encoding cultural knowledge that shapes interpretation. Yet current Arabic language models typically collapse lexical, cultural, and metaphorical information into a single representational space, a phenomenon we term "semantic smearing". We introduce CAMMAR (Culture-Aware Matryoshka for Metaphorical Arabic Representations), a representation learning framework that...

    arxiv.org/abs/2607.15847 · PDF

  9. 09

    Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment

    Haowei Hua

    cs.CL · cs.LG

    Automated essay scoring (AES) enables scalable assessment and timely feedback but remains challenged by transformer input-length limitations, which can cause information loss when processing long essays. This study proposes a generative AI-assisted summarization framework to improve long-form essay representation while maintaining scoring reliability. Using the ASAP 2.0 dataset, we generate controlled-length summaries with three GPT-5...

    arxiv.org/abs/2607.15829 · PDF

  10. 10

    Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models

    Yingqian Cui, Wei Deng, Lantao Mei, Hang Li, Charu C. Aggarwal, Hui Liu, Yue Xing

    cs.CL · cs.LG

    Masked diffusion language models (DLMs) enable parallel text generation by iteratively refining masked tokens, offering a promising alternative to autoregressive decoding. Recent lookahead-based decoding methods improve the accuracy--efficiency trade-off by exploring future decoding states before committing token updates. However, existing approaches mainly rely on shallow one-step lookahead, which optimizes immediate information gain but can...

    arxiv.org/abs/2607.15655 · PDF

  11. 11

    On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous Levels

    Taiga Mori, Koji Inoue, Divesh Lala, Tatsuya Kawahara

    cs.CL · cs.AI · cs.HC

    In multi-party dialogues between a dialogue system and multiple users, identifying to whom an utterance is addressed is a key challenge. Prior work has typically treated addressee detection as a multi-class classification task, selecting a single label representing an individual participant or the group. This formulation assumes that address is inherently discrete and has primarily been used for predicting turn-taking. In this paper, we...

    arxiv.org/abs/2607.15648 · PDF

  12. 12

    Process Reward Informed Tree Rollout for Effective Multi-Turn RL

    Xintong Li, Sha Li, Yuwei Zhang, Changlong Yu, Rongmei Lin, Hongye Jin, Shuyi Guan, Xin Liu, Linwei Li, Qingyu Yin,...

    cs.CL · cs.AI · cs.LG

    Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple independently sampled complete trajectories for advantage estimation. In long-horizon agentic tasks, such a uniform rollout strategy can waste budget on uninformative dead-end attempts, while promising intermediate states do not receive sufficient exploration. The multi-turn structure of agentic trajectories,...

    arxiv.org/abs/2607.15610 · PDF

  13. 13

    VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs

    Shahrzad Esmat, Dhawal Shah, Ali Jannesari

    cs.CL · cs.LG

    The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score importance from an observation window and evict low-scoring tokens, but eviction is irreversible -- so when the importance signal degrades under query-agnostic reuse, accuracy collapses by 11-15 points; uniform low-rank...

    arxiv.org/abs/2607.15498 · PDF

  14. 14

    Verbalizable Representations Form a Global Workspace in Language Models

    Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul...

    cs.CL · cs.AI · cs.LG

    Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its...

    arxiv.org/abs/2607.15495 · PDF

  15. 15

    Large Language Models as Unified Multimodal Learners for Clinical Prediction

    Ajay Madhavan Ravichandran, Bilgin Osmandoja, Klemens Budde, Klaus Netter, Tobias Strapatsas, Aljoscha Burchardt,...

    cs.CL · cs.AI

    Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities. Yet most clinical prediction systems still rely on task-specific fusion architectures, pairing dedicated encoders for each modality with learned combination mechanisms that must be re-engineered for every new task and clinical setting. We propose a simpler alternative: convert all patient...

    arxiv.org/abs/2607.15380 · PDF

  16. 16

    SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

    Yasheng Sun, Zezi Zeng, Yifan Yang, Chong Luo, Wenyi Wang, Ziwei Liu, Jürgen Schmidhuber

    cs.CL · cs.AI

    Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise their manuscripts. Automating this editing workflow under a natural-language instruction, however, is challenging, because a scientific figure is a dense infographic in which heterogeneous visual elements such as schematics, plots, photos, captions, and...

    arxiv.org/abs/2607.15272 · PDF

  17. 17

    In-Place Tokenizer Expansion for Pre-trained LLMs

    Jimmy T. H. Smith, Tarek Dakhran, Alberto Cabrera, Simon S. Lee, Paul Pak, Aditya Tadimeti, Tim Seyde, Maxime...

    cs.CL · cs.AI · cs.LG

    A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflecting the deployment priorities at that time. When those priorities shift, languages added later are split into many more tokens per word, which can raise latency, compute, and energy consumption for users of those languages. Cloud models can afford a broad vocabulary because the embedding and LM-head matrices are a small...

    arxiv.org/abs/2607.15232 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.