cs.CL · 2026-09-04 · No. 105

Computation and Language, 2026-09-04.

19 new papers in cs.CL. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

19 entries
  1. 01

    Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

    Yuntian Deng, Pengyu Nie, Stuart Shieber

    cs.CL · cs.AI · cs.LG

    Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task-specific examples that are used to train a small adapter for a compact interpreter....

    arxiv.org/abs/2609.04199 · PDF

  2. 02

    ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

    Lihao Liu, Peng Tang, Kunwar Yashraj Singh, Shabnam Ghadar

    cs.CL · cs.AI

    Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, producing prompts up to 3$\times$ longer yet no more accurate. We trace this to three deficiencies - incomplete error observation, limited search diversity, and unreliable selection - and propose ESPO (Error-Structured Prompt Optimization), which decomposes prompt optimization into three phases: Diagnose clusters all training errors...

    arxiv.org/abs/2609.04197 · PDF

  3. 03

    Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

    Kevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli

    cs.CL · cs.LG

    Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode...

    arxiv.org/abs/2609.04194 · PDF

  4. 04

    Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

    Joseph Lee, Yidi Huang, Dokyoon Kim, Shu Yang, Li Shen

    cs.CL · cs.AI

    Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from...

    arxiv.org/abs/2609.04180 · PDF

  5. 05

    Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR

    Boyan Li, Bingsen Chen, Chenghao Yang, Ping Nie, Chen Zhao, Xi Ye

    cs.CL · cs.AI · cs.LG

    Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL reward, fusing the two signals within a single step: either as a \emph{weighted-additive combination} or a \emph{teacher-modulated rescaling} of the RL advantage. In this paper, we show that a simple...

    arxiv.org/abs/2609.04108 · PDF

  6. 06

    Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation

    Hasan Alkhder, Mohammad Abboush, Igor Tchappi, Ahmet Zengin, Amro Najjar

    cs.CL · cs.AI

    Neural machine translation (NMT) systems typically produce a single output per input, obscuring the alternative decision trajectories implicitly available within multilingual decoding. This opacity becomes particularly problematic in low-resource dialect settings, where multiple linguistically valid realizations may differ in lexical authenticity, register, and structural stability. We propose reframing translation as a structured decision...

    arxiv.org/abs/2609.04048 · PDF

  7. 07

    Representational alignment yields generalizable safety in language models

    Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu

    cs.CL · cs.AI

    Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their...

    arxiv.org/abs/2609.04022 · PDF

  8. 08

    Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes

    Revathy Venkataramanan, Aditya Luthra, Venkatesan Nadimuthu, Amit Sheth

    cs.CL · cs.AI

    Several studies have evaluated the ability of Large Language Models (LLMs) for meal planning, yielding positive outcomes. These models can process natural language inputs and leverage learned knowledge from their pretraining to generate meal plans. In this work, we investigate the ability of LLMs to analyze the suitability of given recipes for diabetes. The primary challenge for LLMs is to retrieve relevant dietary guidelines for diabetes,...

    arxiv.org/abs/2609.03967 · PDF

  9. 09

    Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs

    Jiacheng Xu, Wentao Zhang, Zhiyi Lyu, Fuxiang Zhang, Chaojie Wang, Yang Liu, Bo An

    cs.CL · cs.LG

    Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) through executable feedback. The feedback for coding problems mainly comes from specific test cases, where high-quality test cases are often scarce since they should be both sound and discriminative. We thus turn to study the auto-generation of test cases using the learned model. We find this is naturally an adversarial RL problem: the...

    arxiv.org/abs/2609.03955 · PDF

  10. 10

    IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

    Saikat Mondal, Mamta, Deeksha Varshney, Oana Cocarascu, Asif Ekbal

    cs.CL · cs.AI

    Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a persuasion-based jailbreak evaluation framework for Indian languages. Our benchmark combines ten safety critical content categories with six human-like persuasive...

    arxiv.org/abs/2609.03781 · PDF

  11. 11

    Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks

    Oline Ranum, Edward Fish, Simon Hadfield, Richard Bowden

    cs.CL · cs.AI

    BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit spurious correlations and spoken-language priors, rather than learning stronger sign representations. In this paper, we evaluate the relationship between spatio-temporal understanding and BLEU-4 across six SLT models...

    arxiv.org/abs/2609.03734 · PDF

  12. 12

    Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements

    Arianna Miola, Bruno Spaccavento, Lorenzo Silotto, Marco Bianchetti, Luca Cagliero

    cs.CL · cs.AI · cs.CE · cs.IR

    The comparative analysis of banks' financial statements poses significant challenges for automated question answering systems due to their complexity, substantial length, technical language, and inhomogeneity of both textual and numerical content across different jurisdictions and institutions. We introduce FinRAG-QA, a novel benchmark dataset for financial question answering, which comprises 999 practitioner-curated questions on 10...

    arxiv.org/abs/2609.03654 · PDF

  13. 13

    </think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination

    Seunghee Koh, Sungjae Choi, Minchan Kwon, Sunghyun Baek, Junmo Kim

    cs.CL · cs.AI

    Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to stop reasoning. We study one such strategy that injects an end-of-think token (EoT, </think>) at this point to trigger the reasoning-to-answering transition, and find that the injected EoT does not always induce a...

    arxiv.org/abs/2609.03633 · PDF

  14. 14

    Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation

    Xuanfa Jin, Zhijian Ma, Yongcheng Zeng, Xinyu Cui, Haifeng Zhang, Jun Wang

    cs.CL · cs.AI

    Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as shared misconception: when a majority of agents initially converge on an incorrect answer, the debate process tends to amplify rather than correct the error. Existing methods primarily address peer skew but leave the...

    arxiv.org/abs/2609.03619 · PDF

  15. 15

    How Far Can Synthetic Data Take Thai OCR?

    Kunat Pipatanakul

    cs.CL · cs.AI · cs.CV

    We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to build Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages. Synthetic data provide exact labels at scale, but "realism" conflates source domain, page context, typography, spatial structure, and glyph variation. We disentangle these factors with a controlled document-reconstruction...

    arxiv.org/abs/2609.03595 · PDF

  16. 16

    Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

    Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek...

    cs.CL · cs.AI

    In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (e.g., 15 seconds) into a compact fixed-voice student trained entirely on synthetic speech. This setting makes pipeline design...

    arxiv.org/abs/2609.03502 · PDF

  17. 17

    Pattern Over-Generalization of Knowledge Graph Embedding

    Junsik Kim, Kangil Kim

    cs.CL · cs.AI

    Knowledge graph embedding (KGE) demonstrates its effectiveness for predicting missing links in knowledge graphs (KGs) by projecting entities and relations into a low-dimensional vector space. It is crucial for KGE models to effectively capture inference patterns (patterns) inherent in KGs, such as symmetry/antisymmetry, inversion and composition. Although recent KGE models exhibit strong capabilities in modeling such diverse patterns, they...

    arxiv.org/abs/2609.03487 · PDF

  18. 18

    When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

    Wen-Yu Chang, Yun-Nung Chen

    cs.CL · cs.AI

    Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conversational usage. We introduce LOCOMO-CONV, a conversa- tional memory benchmark derived from Lo- CoMo with four query styles: dialog, implicit, counterfactual, and composed. Across five rep-...

    arxiv.org/abs/2609.03467 · PDF

  19. 19

    TabScope: Question-Adaptive Scope Selection for Table Question Answering

    Yuxiang Wang, Junhao Gan, Jianzhong Qi

    cs.CL · cs.AI

    Large Language Models (LLMs) have shown strong performance on table question answering, yet their accuracy often degrades as table size increases. We find that this degradation is not uniform across question types. Localization-sensitive questions are particularly affected by irrelevant table content, while questions requiring broader evidence may still benefit from full-table reasoning. Based on this observation, we propose a...

    arxiv.org/abs/2609.03395 · PDF

This edition is part of The Daily Abstract — cs.CL archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.