cs.AI · 2026-09-28 · No. 127

Artificial Intelligence, 2026-09-28.

52 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

52 entries
  1. 01

    Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

    Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe, Chenrui Fan, Sourya Basu, Genta Indra Winata, Anirban Das,...

    cs.AI · cs.CL · cs.LG

    Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision:...

    arxiv.org/abs/2609.31619 · PDF

  2. 02

    DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education

    Quang Nguyen, Hieu Nguyen, Hien Hoang, Toan Pham, Cong Tran, Nam Vu

    cs.AI

    AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is...

    arxiv.org/abs/2609.31568 · PDF

  3. 03

    Multi-agent Scaling Across Disjunctive and Compensatory Tasks

    Carolina Fortuna, Blaz Bertalanic

    cs.AI · cs.MA

    Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits:...

    arxiv.org/abs/2609.31563 · PDF

  4. 04

    A Flow Matching Framework for Neural Representational Dissimilarity

    Zeyuan Ye, Xue-Xin Wei

    cs.AI · cs.IT · cs.LG

    Neural representational dissimilarity quantifies differences between neural response distributions, and is essential for comparing neural codes across stimuli, brain areas, tasks, and models. Commonly used distance metrics involve different assumptions and are estimated with separate methods. Here, we show that a variety of distance metrics can be unified under a flow matching framework developed in deep generative models. That is, these...

    arxiv.org/abs/2609.31544 · PDF

  5. 05

    Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity

    Marius F. R. Juston, Kevin A. Karim, Jonathan Gao, Kevin C. Li, Rudhi Bashambu

    cs.AI

    Despite the growing capabilities of large language models (LLMs), prompt design remains largely heuristic and ad hoc. This project will explore $\textit{prompt minimization}$, the process of reducing prompts to their smallest, most information-dense form while preserving output fidelity. Practically, shorter prompts reduce computational overhead and inference latency, especially when large contexts, such as entire documents or codebases, are...

    arxiv.org/abs/2609.31505 · PDF

  6. 06

    UQ-LOB: Uncertainty-Aware Limit Order Book Mid-Price Forecasting

    Derrick Gilchrist Edward Manoharan, Eljas Linna, Kestutis Baltakys, Hao Dong, Juho Kanniainen

    cs.AI

    Forecasting short-horizon mid-price movements from limit order book (LOB) data is central to algorithmic trading, yet most deep LOB forecasters are point predictors: they output a direction or a displacement, but never indicate which of their forecasts can be trusted. We introduce UQ-LOB, a lightweight, encoder-agnostic uncertainty quantification module that attaches to any pretrained LOB encoder and, in the spirit of attentive neural...

    arxiv.org/abs/2609.31491 · PDF

  7. 07

    "AI is (not) the new...": A Diagnostic Analogy Framework for Generative AI's Cultural Impacts

    Rida Qadri, Vinodkumar Prabhakaran, Remi Denton

    cs.AI

    Generative AI is reshaping the cultural infrastructures through which knowledge is found, synthesized, and held accountable. To make sense of this shift, scholars and policymakers reach for historical analogies of technologies such as the printing press, steam power or electricity. But these comparisons are typically imprecise about which property of the technology carries the comparison, and imprecise analogies produce imprecise governance...

    arxiv.org/abs/2609.31482 · PDF

  8. 08

    Game Arena: Strategic LLM Evaluation in Competitive Environments

    Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang, Timothy Chung, Martyna Plomecka, John Schultz, Jon...

    cs.AI

    We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three...

    arxiv.org/abs/2609.31473 · PDF

  9. 09

    Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency

    Myeongjun Erik Jang, Antonios Georgiadis, Sae Young Moon, Fran Silavong

    cs.AI

    Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that prompt LLMs to generate topics then assign the topics to documents, producing more natural and human-readable topics than conventional topic modeling algorithms. However, the nature of...

    arxiv.org/abs/2609.31460 · PDF

  10. 10

    Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

    Zhensheng Zou, Guoqing Wang, Dan Hao

    cs.AI

    Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representations can compromise their original behavior. To reduce context while preserving action-critical information and agent behavior, we combine Latent Observations, Hard Actions (LOHA), a...

    arxiv.org/abs/2609.31430 · PDF

  11. 11

    Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection

    Guangzhe Zhang

    cs.AI

    Context projection replaces older tool observations with compact, addressable excerpts, reducing repeated input while potentially adding evidence-retrieval turns. We study this trade-off in ReVerPi, a Pi extension with archived observations and matched full/projected continuations. In an 86-run source-reading campaign with 641 model requests, the 15 completed pairs show identical success: 12/15 per arm. Twelve further boundary runs stop, with...

    arxiv.org/abs/2609.31381 · PDF

  12. 12

    Programs-of-Layers in LLMs through the Lens of Cortical Areas

    Justus Westerhoff, Stephan Olbrich, Hatem Oraby, Matthew Evan Larkum, Felix Alexander Gers

    cs.AI

    Inference in LLMs is conventionally a fixed-depth, fixed-order forward pass through every layer, regardless of how difficult the input is. The human brain does not work this way: using the thalamus as a central hub, it routes information flexibly to all regions of the cortex according to demand. Li et al. (2026) recently showed, with a system they call program-of-layers (PoLar), that transformers can be given an analogous flexibility if their...

    arxiv.org/abs/2609.31360 · PDF

  13. 13

    Mutable Transcripts: Mitigating Context Pollution through Editable Conversation State

    Dan Barry, Andrew Hines

    cs.AI

    Contemporary large language model (LLM) chat systems treat conversation history as an immutable sequence of turns that defines the model's working context. However, user intent in real interactions is not static: it evolves through correction, refinement, and shifting constraints. This mismatch between dynamic intent and static transcripts can result in context pollution, where outdated or irrelevant information persists and continues to...

    arxiv.org/abs/2609.31354 · PDF

  14. 14

    The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models

    Christoph Walser, Mauricio Fadel Argerich, Jonathan Fürst

    cs.AI · cs.CL

    Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ parameter) text-only and vision--language models, evaluated on both accuracy and energy over a design space spanning input...

    arxiv.org/abs/2609.31341 · PDF

  15. 15

    G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

    Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu

    cs.AI

    Offline multi-agent reinforcement learning (MARL) learns cooperative policies from fixed datasets without further environment interaction and a learned policy is frozen at deployment. Such a frozen policy typically proposes a single joint action and executes it directly at deployment time. However, this one-shot deployment often commits to a suboptimal proposal, even when better nearby alternatives remain consistent with the behavior data. To...

    arxiv.org/abs/2609.31286 · PDF

  16. 16

    MA-WAM: Multi-Agent World-Action Model for Test-Time Planning

    Guowei Zou, Haitao Wang, Guoxin Wang, Beiwen Zhang, Zhiquan Chen, Guojie Wang, Hejun Wu

    cs.AI

    Multi-agent cooperative tasks require different agents to execute a joint action simultaneously, and each agent's action affects both the observations and responses of the other agents. Hence, a world model is needed to predict the team return resulting from the joint actions of all agents. A naive extension directly applies a single-agent world model to each agent's action when predicting the team return step by step. However, such an...

    arxiv.org/abs/2609.31281 · PDF

  17. 17

    Purin: A Biology-inspired Mechanism for Artificial Neural Networks

    Zishu Liu, Chunbo Luo, Christos Grecos

    cs.AI · cs.NE

    Artificial neural networks (ANNs) usually represent neural transmission with fixed trainable weights during a training batch, which omits short-term changes in synaptic efficacy. In addition, the discrete time-step simulation requires additional temporal processing that many conventional ANN architectures do not use. To overcome these challenges, we propose Purin, a biology-inspired and ANN-compatible mechanism, that introduces synaptic...

    arxiv.org/abs/2609.31235 · PDF

  18. 18

    DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration

    Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du

    cs.AI · stat.AP · stat.ME · stat.ML

    Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified framework that combines abundant LLM comparisons with limited human comparisons to separate judge-specific position effects, learn shared structure in position-debiased LLM preferences, and...

    arxiv.org/abs/2609.31215 · PDF

  19. 19

    Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution

    Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun

    cs.AI · cs.LG

    Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a more fundamental source of disagreement is specification mismatch: influence depends on the behavior being attributed, the intervention applied to each training example, and the...

    arxiv.org/abs/2609.31214 · PDF

  20. 20

    Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture

    Christian Schiffer, Mathis Bode, Thomas Lippert, Katrin Amunts, Timo Dickscheid

    cs.AI

    Scaling studies typically represent training data by a single count of samples. For hierarchically and spatially structured data, however, the same number of samples can be drawn from few or many sources and distributed differently across the underlying domain. We therefore study data scaling as an allocation problem, separating unique sample count, source diversity, and spatial coverage. We study this decomposition in microscopic whole-brain...

    arxiv.org/abs/2609.31201 · PDF

  21. 21

    Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation

    Chang Gong, Jingping Bi, Di Yao, Xinjian Liang, Chao Xiang, Ruijie Guo

    cs.AI

    Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development. As AI begins to participate in its own improvement, from model training and experience accumulation to agent evolution and automated AI development, the prospect of recursive self-improvement (RSI) is becoming increasingly relevant. This transition raises a...

    arxiv.org/abs/2609.31186 · PDF

  22. 22

    Accounting for Bias Enables Sustainable LLM Evaluation

    Harshita Katoch, David Antony Selby, Gerrit Großmann, Sebastian Vollmer

    cs.AI · cs.LG · stat.AP · stat.ME

    LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge...

    arxiv.org/abs/2609.31184 · PDF

  23. 23

    SPO: Discovering Adaptive Large Neighborhood Search Operators via Stackelberg Program Optimization

    Xinyi Ke, Kai Li, Junliang Xing, Yifan Zhang, Jian Cheng

    cs.AI

    Large neighborhood search (LNS) relies critically on destroy and repair operators, whose effectiveness depends on both adaptation to the evolving LNS state and interaction between the two roles. We introduce Stackelberg Program Optimization (SPO), an LLM-based framework for discovering adaptive executable destroy-repair programs. SPO conditions operator decisions on a compact LNS state, allowing state-dependent behavior to emerge through...

    arxiv.org/abs/2609.31179 · PDF

  24. 24

    Semantic Navigation for Issue Localization in Code Repository

    Yunxiang Wei, Zhenyu Lei, Jundong Li

    cs.AI

    Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue. LLM agents approach this task iteratively: they identify a set of potentially relevant locations, inspect the corresponding code, and revise their judgments about these candidates as new evidence is acquired. Existing environments, however, provide limited support for this loop: agents must search for unresolved...

    arxiv.org/abs/2609.31176 · PDF

  25. 25

    Neural State Prediction: Obstructing Shortcut Learning in EEG Foundation Models

    Kieren Yu, Ziyang Liu, Chang Huang, Jintai Chen, Kaishun Wu

    cs.AI

    EEG foundation models increasingly use masked prediction to learn from unlabeled recordings, but optimizing this objective does not ensure transferable neural representations. A central challenge is that stable positional cues and local correlations can make masked regions predictable without integrating distributed neural context. To reduce this reliance on low-information prediction paths, we introduce Neural State Prediction (NSP), a...

    arxiv.org/abs/2609.31167 · PDF

  26. 26

    Momentum-Guided Federated Split Distillation for Personalized Temporal Edge Intelligence

    Ahmed-Rafik Baahmed, Jean-François Dollinger, Amine Brahmia, Mourad Zghal

    cs.AI · cs.DC

    We propose a momentum-guided federated split distillation framework for personalized, efficient, and autonomous temporal edge intelligence. We introduce TeRR-SAtt, our novel temporal reservoir student attention design that combines fixed reservoir representations, a lightweight temporal student, and personalized output modules. We also present AMGF, our anticipatory momentum-guided fusion mechanism that clusters clients through learning...

    arxiv.org/abs/2609.31159 · PDF

  27. 27

    Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

    Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu, Shuxia Lin, Xu Yang

    cs.AI

    Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base LLM. While the base LLM retains usable reasoning after scaling, the aligned VLM itself cannot reliably access this ability. Therefore, recovering the degraded reasoning capability in VLMs...

    arxiv.org/abs/2609.31140 · PDF

  28. 28

    Toward AI-Augmented Cooperative Engineering Workflows: Requirements and Architecture the European Rover Challenge

    Ahmed R. Sadik, Frank Joublin, Mariusz Bujny, Antonello Ceravola, Joan Smith

    cs.AI

    The growing availability of Artificial Intelligence (AI) tools creates new opportunities to support engineering design processes, yet their current use often remains limited to isolated tasks such as coding, documentation, or information retrieval. Less attention has been given to how AI can support cooperative engineering workflows at the process level, where teams must coordinate requirements, tasks, communication, knowledge transfer, and...

    arxiv.org/abs/2609.31136 · PDF

  29. 29

    AtomWorld-Mem: Memory-Restored World States for Long-Horizon Atomistic Evolution

    Tian Luo, Ruge Zhang, Haozhi Han, Yifrng Chen, Yunquan Zhang, Yunxin Liu, Ting Cao, Kun Li

    cs.AI · cond-mat.mtrl-sci

    High-fidelity atomistic evolution over long timescales requires more than observing the current crystal configuration. Instantaneous atomistic snapshots are often incomplete: locally similar configurations can correspond to different hidden dynamical contexts, future event preferences, and waiting-time scales. We argue that this snapshot ambiguity makes long-horizon atomistic evolution fundamentally a memory-based world-state restoration...

    arxiv.org/abs/2609.31133 · PDF

  30. 30

    Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning

    Julian Schulz

    cs.AI

    Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning before models act. A key concern is encoded reasoning, where models hide their true reasoning in ways that monitors and humans cannot interpret. Optimization pressure from CoT monitors during reinforcement learning is considered a likely driver of such behavior. We investigate this by training reasoning models to...

    arxiv.org/abs/2609.31121 · PDF

  31. 31

    OmouAI: Argumentative Human-AI Policy Deliberation with Simulated Personas

    Stylianos Loukas Vasileiou, Antonio Rago, William Yeoh, Georgina Curto

    cs.AI

    Debates amongst agents driven by large language models (LLMs) have demonstrated vast potential in various applications, but when these interactions include humans and take place in high-stakes environments, e.g., in public policy deliberations, they are beset with issues such as sycophancy and a lack of faithful explanations. To tackle these issues, we present OmouAI, an interactive and inclusive deliberation system that uses LLMs in...

    arxiv.org/abs/2609.31078 · PDF

  32. 32

    Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

    Bartłomiej Cupiał, Jens Tuyls, Maciej Wołczyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Miłoś,...

    cs.AI

    Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions. The code handles recurring local decisions, while the language model decides which skills to use and how to combine them. Yet abstractions are leaky, and situations beyond a skill's...

    arxiv.org/abs/2609.31076 · PDF

  33. 33

    Externalized CPDAG Summaries Improve LLM Causal Deduction

    Wentao Sun, João Paulo Nogueira, Dominique Verchere, Mathieu Acher, Alonso Silva

    cs.AI

    Corr2Cause asks whether a causal claim holds in every DAG compatible with observed correlations and conditional independencies. We frame this as latent-object reasoning: the label is defined by a CPDAG query, but free-form chain-of-thought often collapses the Markov-equivalence-class problem into local pattern matching. We propose Structured Thinking, a two-turn pipeline that first externalizes a typed, schema-constrained CPDAG summary and...

    arxiv.org/abs/2609.31071 · PDF

  34. 34

    Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders

    Itai Zehavi, Fanny Jourdan, Ulrich Aivodji

    cs.AI

    Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral forgetting alone may be insufficient, motivating interventions directly on internal representations. However, standard mechanistic-interpretability extractors are poorly selective for...

    arxiv.org/abs/2609.31056 · PDF

  35. 35

    Cheap, open agents make LLM pollution harder to mitigate

    Raluca Rilla, Anne-Marie Nussberger, Rui Mata, Dirk U. Wulff

    cs.AI

    Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones....

    arxiv.org/abs/2609.31054 · PDF

  36. 36

    Governed Deduction: Policy-Grounded Premise Authorization Beyond Relevance

    Wesley Shu, Hsi-Ching Lin

    cs.AI

    Reasoning systems usually treat premise use as a question of relevance: if a fact is available and useful, it may be selected for inference. Authorization imposes a different constraint: a premise may be represented and logically usable but not permitted for a particular local transition. We formalize this distinction as Governed Deduction (GD), with a transition-local admission predicate admit(p, tau, S). From an independently produced...

    arxiv.org/abs/2609.31029 · PDF

  37. 37

    Same Text, Different Numbers: The Divergence of LLM-Based Measures

    Hamid Boustanifar, Sasan Mansouri

    cs.AI · cs.CL · q-fin.GN · q-fin.RM

    Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these...

    arxiv.org/abs/2609.31013 · PDF

  38. 38

    Factorized axis convolutional gated recurrent unit with dynamic adaptive pooling for remaining useful life prediction of rolling bearings

    Hanbyeol Park, Jungho Choo, Hyerim Bae

    cs.AI

    Convolutional neural networks (CNN) are widely used to predict the remaining useful life (RUL) of rolling bearings from time-frequency representations (TFRs) of vibration signals. However, during degradation, characteristic structures in TFRs align predominantly along the frequency or time axis, making it challenging for conventional CNN isotropic kernels to capture directional structure. Furthermore, global average pooling (GAP) averages...

    arxiv.org/abs/2609.30972 · PDF

  39. 39

    SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents

    Maokai Qin, Chuan Qin, Qi Zhang, Dianyu Liu, Zirui Liu, Hongting Niu, Yuanchun Zhou, Hengshu Zhu

    cs.AI

    Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale. To address this challenge, we introduce...

    arxiv.org/abs/2609.30971 · PDF

  40. 40

    MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

    Subhojyoti Mukherjee, Md Mehrab Tanjim

    cs.AI

    Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inherently multi-objective: an accurate harness that refuses no unsafe request, or that consumes an order of magnitude more tokens, is not a good harness. We present Meta-Harness, a system...

    arxiv.org/abs/2609.30967 · PDF

  41. 41

    FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models

    Arjun Pillai, Christian Hoang, Anjelo Laroza

    cs.AI

    Large language models operating in multilingual contexts must resolve target response languages early in generation, yet the causal circuitry governing first-token language identity decisions remains poorly mapped. We present an end-to-end structural circuit analysis across six model architectures spanning four families: GPT-2, BLOOM-560M, Pythia-1B/2.8B, and Qwen2.5-1.5B Base/Instruct. Using Edge Attribution Patching (EAP) with FP16 active...

    arxiv.org/abs/2609.30954 · PDF

  42. 42

    LogicTree-RAG: Logic Tree-guided Retrieval-Augmented Generation for Long-form Patent Drafting

    Jiaqi Zhu, Naili Xing, Hexiang Pan, Haotian Gao, Jianwei Yin, Xiaokui Xiao, Beng Chin Ooi

    cs.AI

    Long-form technical text generation underpins knowledge-intensive workflows, yet remains challenging for large language models (LLMs) due to the need for globally consistent logical structuring and faithful technical reasoning beyond local coherence. Patent drafting is a canonical instance of this challenge, demanding holistic generation of a legally compliant and technically exhaustive document through sustained multi-expert collaboration....

    arxiv.org/abs/2609.30943 · PDF

  43. 43

    Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms

    Zhenhao Fu, Ruipeng Xu, Qibing Ren

    cs.AI · q-fin.GN

    Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, but also at the level of the systems they jointly create. We study this problem with FRAIL, a controlled experimental framework that places LLM agents in three dynamic financial...

    arxiv.org/abs/2609.30940 · PDF

  44. 44

    MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory

    De Jiang, Shuo Zhang, Weiwei Liao, Jianying Zhang, Chuanhui Yu, Hongen Liao, Kehong Yuan

    cs.AI

    Cognitive behavioral therapy (CBT) is an evidence-based first-line treatment for depression, yet its scale is constrained by the time clinicians spend on pre-session preparation, post-session documentation, and longitudinal cognitive-pathology tracking. We present a clinician-facing AI decision-support system that combines a multi-agent CBT framework (MACBT) with a CBT-specific longitudinal memory module (CD Memory). MACBT encodes the...

    arxiv.org/abs/2609.30939 · PDF

  45. 45

    Self-Play Search Distillation for Large Language Model Reasoning

    Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni, Luca Ragazzi, Gianluca Moro, Pavlos Vougiouklis, Jeff Z. Pan,...

    cs.AI

    Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games. SPSD uses...

    arxiv.org/abs/2609.30936 · PDF

  46. 46

    JevSoup: System-One Routing for Training-Free LoRA Composition

    Xiuying Wang, Jiahua Cheng, Shuotian Li, Yufan Cheng, Bowen Deng, Zhexuan Bai, Yichen Li

    cs.AI

    Building adaptable AI systems requires effective coordination of specialized capabilities across diverse tasks. Low-rank adaptation (LoRA) enables modular expertise, but existing routing approaches may require auxiliary data, additional training, or autoregressive decoding. We propose JevSoup, a training-free framework separating System One expert routing from System Two execution. Using only the input and expert descriptions, Jev selects two...

    arxiv.org/abs/2609.30922 · PDF

  47. 47

    Training Graph Foundation Models on The Web Graph

    Ryoma Sato

    cs.AI · cs.LG

    We introduce Acacia, a graph foundation model, trained on the web graph. Acacia (i) supports arbitrary feature dimensionalities and semantics without additional training, (ii) supports a wide range of tasks, including node classification, link prediction, node clustering, and graph generation, without additional training, (iii) has in-context learning capabilities, and (iv) does not rely on pretrained LLMs. In particular, existing graph...

    arxiv.org/abs/2609.30894 · PDF

  48. 48

    From Tapping to Hopping: Augmenting Mobile GUI Agents with App-Native Deeplinks

    Yuchen Sun, Chenglin Cai, Gongjie Zhang, Tianyu Xia, Quyu Kong, Panrong Tong, Zhengwen Zeng, Long Chen, Steven Hoi,...

    cs.AI

    Mobile GUI agents complete tasks using GUI actions like taps and swipes. These actions are broadly applicable across applications, but reaching a navigation interface. A single deeplink call can replace a sequence of screen-by-screen GUI actions. We therefore introduce hybrid interaction, using deeplinks for direct navigation and GUI actions for other on-screen operations and fallback. To enable this, we discover candidate deeplinks through...

    arxiv.org/abs/2609.30887 · PDF

  49. 49

    EXAONE Demand 1.0: A Time Series Foundation Model for Demand Forecasting

    Seunghan Lee, Sangjun Han, Jun Seo, Junhyeok Kang, Jaehoon Lee, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Minjae Kim,...

    cs.AI · cs.LG

    Time series foundation models (TSFMs) are pretrained on series from diverse domains, where demand series make up only a small fraction. Demand data has properties that such corpora rarely contain: Short histories, frequent zeros, censoring by stock-outs, and exogenous events that the series does not record. To this end, we propose EXAONE Demand, built on 1) a demand-specific corpus and 2) a demand-aware adapter. For the corpus, we assemble...

    arxiv.org/abs/2609.30880 · PDF

  50. 50

    TISD: On-Policy Self-Distillation with Trajectory Intervention

    Taeckyung Lee, Rinat Amankos, Jeonghye Kim, Hyungjun Yoon, Woogyeol Jin, Sung-Ju Lee

    cs.AI · cs.LG

    On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but cannot supervise the successor contexts induced by that action unless the student samples it. This creates a training-time data-collection bottleneck and suggests a different role for...

    arxiv.org/abs/2609.30878 · PDF

  51. 51

    SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting

    Guanyu Nie, Fangzhou Zhu, Shixiong Kai, Xiongwei Han, Tao Zhong, Mingxuan Yuan

    cs.AI

    Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task-specific instructions, while new updates can disrupt behavior that previously worked. We study this problem as skill-evolution overfitting and introduce SkillEvoReg, a general regularization framework for skill...

    arxiv.org/abs/2609.30861 · PDF

  52. 52

    Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis

    Thong Bach, Dung Nguyen, Thao Minh Le, Truyen Tran

    cs.AI

    Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscape: a well-aligned model routes harmful queries toward safe outputs through an energy barrier that separates the two regions. Current jailbreak attacks reduce to two strategies for...

    arxiv.org/abs/2609.30841 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.