cs.AI · 2026-10-01 · No. 130

Artificial Intelligence, 2026-10-01.

30 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

30 entries
  1. 01

    Turbo Harness: Instance-Adaptive Harness Optimization

    Tunyu Zhang, Hao Wang, Kai Xu, Dimitris N. Metaxas

    cs.AI

    Automating the search for effective harnesses is an important step toward enabling agents to recursively self-improve. Existing harness optimizations typically produce a single global harness that is applied uniformly across task instances. However, a harness that works well on average may not be optimal for every instance. We introduce Turbo Harness, a framework that can adapt a globally optimized harness to each instance by reusing...

    arxiv.org/abs/2609.40330 · PDF

  2. 02

    WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

    Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu, Qiucheng Wu, Tommi Jaakkola, Yang Zhang, Shiyu Chang

    cs.AI

    As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However,...

    arxiv.org/abs/2609.40325 · PDF

  3. 03

    Cogentic: Multi-Agent Orchestration for Automated Proof Discovery

    Yang Cai, Vineet Gupta, Yanchen Jiang, Christopher Liaw, Aranyak Mehta, Grigoris Velegkas, Di Wang

    cs.AI · cs.GT

    We present Cogentic, a multi-agent harness for automated proof discovery on open research problems. While frontier language models can generate strong mathematical ideas in a single shot, single-shot generation is often insufficient for open problems that require exploring multiple competing conjectures, overcoming subtle technical obstructions, and retaining intermediate progress over a long horizon. Cogentic addresses these challenges...

    arxiv.org/abs/2609.40324 · PDF

  4. 04

    How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

    Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel Abbé

    cs.AI

    Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more primitive but improved...

    arxiv.org/abs/2609.40303 · PDF

  5. 05

    PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

    Yinghui He, Yapei Chang, Khushi Bhardwaj, Daniele Molinari, Tugrul Konuk, Jan Kautz, Ali Hatamizadeh

    cs.AI

    On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action...

    arxiv.org/abs/2609.40285 · PDF

  6. 06

    Belief-Aware Multi-Agent Path Finding under Map Uncertainty

    Viraj Parimi, Shao-Hung Chan, Han Zhang, Jingkai Chen, Brian Williams

    cs.AI · cs.MA · cs.RO

    Multi-Agent Path Finding (MAPF) aims to find collision-free paths for multiple agents in a shared environment. Classical MAPF assumes that all static obstacles are known in advance, but real-world environments can change unexpectedly due to fallen objects, spills, or other local disturbances. When such changes are spatially correlated, an observation can inform traversability estimates beyond the observed location. Prior approaches address...

    arxiv.org/abs/2609.40269 · PDF

  7. 07

    Learning from Research: Toward Lifelong Agent Harness Evolution

    Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Yaar Harari, Evgeniy Gabrilovich, Shiyu Chang

    cs.AI

    Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed. Recent methods automate this process by using a meta coding agent to modify the harness based on execution feedback. However, relying on that...

    arxiv.org/abs/2609.40169 · PDF

  8. 08

    Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR

    Chandak Chakma, Syed Nazmus Sakib, Nafiul Haque, Shifat E. Arman

    cs.AI

    Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study...

    arxiv.org/abs/2609.40115 · PDF

  9. 09

    Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

    Kunlun Zhu, Xuyan Ye, Yibo Li, Cheng Qian, Beibin Li, Heng Ji

    cs.AI · cs.CL

    An unsuccessful LLM agent rollout contains more information than its final reward: the observations available to the agent, the actions it chose, and the environment's responses. Reusing this experience for learning requires identifying a decision to revise and testing a concrete alternative. We introduce the Agent Error Dataset (AED), comprising 50,228 error-diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families,...

    arxiv.org/abs/2609.40111 · PDF

  10. 10

    PTNO: Training Neural Operators with Noisy Monte Carlo Estimates for Particle Transport Problems

    Yubo Cao, Xi Deng, Mengqi Xia, Vignesh Gopakumar, Ander Gray, Anima Anandkumar

    cs.AI · cs.LG

    Particle transport under multiple scattering is central to radiative transfer and plasma physics, yet high-fidelity Monte Carlo (MC) simulations must trace prohibitively many particles. Learning-based surrogates can amortize this cost, but typically train on expensive, well-converged MC solutions. We propose the Particle Transport Neural Operator (PTNO), a neural operator that learns particle transport surrogates directly from noisy, low-cost...

    arxiv.org/abs/2609.40090 · PDF

  11. 11

    Who Verifies the Graph? Misspecification Attacks on Causal Action Verification for Language Agents

    Fabio Rovai

    cs.AI · cs.LG

    Causal action verifiers gate an agent's state-changing tool calls by checking whether each proposed intervention is identifiable against a committed action-state graph, and they issue a certificate that carries the identification argument and a one-sided lower confidence bound. One such verifier, CIVeX, reports zero false executions on a confounded tool-use benchmark. We red-team it by corrupting only the committed graph. Omitting a single...

    arxiv.org/abs/2609.40027 · PDF

  12. 12

    What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents

    Shuyang Zhang, Jianshuo Chang

    cs.AI

    Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is needed to assess its task-level benefit and the contribution of local decision quality. Methods. This critical scoping review maps 348 studies and examines 90 comparison records: 88 from 40 included studies and two from supplementary studies....

    arxiv.org/abs/2609.39989 · PDF

  13. 13

    AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC

    Yijie Bian, Kai Zhang, Wei Guo, Zixin Wang, Shenghui Song, Jun Zhang, Khaled B. Letaief

    cs.AI · cs.MA

    Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. Data-driven multi-modal ISAC models depend heavily on annotated real-world data to learn relationships across sensing and wireless observations, thereby constraining scalable deployment. Although synthetic data generation reduces the burden, adapting existing simulation pipelines to a target...

    arxiv.org/abs/2609.39964 · PDF

  14. 14

    Better Deck or Different Judge? Evaluating Agentic Harness Gains in Corporate and Investment Banking

    Ludovic Gibert, Matis Despujols, Andre-Louis Rochet

    cs.AI

    Corporate and investment banking teams use presentations to support credit decisions and advise clients on financing and transactions. Producing these decks requires reconciling financial data, tracing sources and turning analysis into a recommendation. We retrospectively study the development of an agentic harness combining a 27B language model, financial calculations, narrative templates and validation checks. LLM judges guide engineering...

    arxiv.org/abs/2609.39958 · PDF

  15. 15

    Coverage Before Control: Route-Instruction Grounding and Steering for Controllable Retrosynthesis

    Xuemin Chen, Xiaozhuang Song, Xinjian Zhao, Yaoyao Xu, Tianshu Yu

    cs.AI · cs.LG

    Single-step retrosynthesis models are commonly evaluated by their ability to recover recorded reactions. In practice, chemists may need to choose among several precursor sets for the same product, for example to preserve a particular motif. Recovering a recorded answer alone does not establish this ability to follow a preference. Satisfying such requests requires both coverage of relevant alternatives and control over which alternatives are...

    arxiv.org/abs/2609.39955 · PDF

  16. 16

    ConflictGuide: AutoResearch Improves When Competing Behaviors Are Made Visible

    Binqian Xu, Qiran Zou, Xiangbo Shu, Dianbo Liu

    cs.AI · cs.LG

    When designing machine learning models, desirable properties are often in tension: improving one behavior can impair another, so task progress can depend on alleviating the conflict. LLM-based AutoResearch systems, which iteratively edit model code and retain edits based on scalar task-performance feedback, have largely ignored this trade-off. We find that scalar feedback supports broad exploration early in search, but it does not reveal how...

    arxiv.org/abs/2609.39933 · PDF

  17. 17

    OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

    Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, Oliver Sun, Eric Wang,...

    cs.AI

    Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for...

    arxiv.org/abs/2609.39903 · PDF

  18. 18

    GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning

    Gabriele Tuccio, Antonino Furnari, Aldo Gangemi, Misael Mongiovì

    cs.AI · cs.CL · cs.LG

    Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when the prompt is underspecified or the model has limited instruction-following ability. Beam search can partially mitigate these failures by exploring multiple valid sequences, but its computational cost grows with...

    arxiv.org/abs/2609.39869 · PDF

  19. 19

    Completion-Aware Cross-Fidelity Offline-to-Online Reinforcement Learning for Multi-Line Bus Holding

    Yifan Zhang, Qifan Zhang, Liang Zheng

    cs.AI · cs.LG

    Exploratory reinforcement learning (RL) on an operating bus fleet is impractical,while policies trained only from historical data cannot acquire new experience. Hybrid Offline-and-Online (H2O) RL combines fixed target replay with simulator interaction, but the inexpensive online simulator can differ from the target in transition and event-duration dynamics. We study this cross-fidelity problem for multi-line bus holding and address a failure...

    arxiv.org/abs/2609.39868 · PDF

  20. 20

    FIGS: Evaluating Multi-Turn Sycophancy Without Penalizing Empathy

    Sidharth Pulipaka, Ruta Binkyte, Ivaxi Sheth, Sahar Abdelnabi

    cs.AI · cs.CL

    Large language models frequently fail to balance staying truthful with being supportive. They often exhibit sycophancy in responses to users, agreeing with false claims, offering unwarranted flattery, and giving advice skewed toward users' expressed views. In reality, sycophancy rarely happens in a single exchange; it may emerge organically as users repeatedly insist or subtly steer the dialogue over time. Current evaluations, however, rely...

    arxiv.org/abs/2609.39863 · PDF

  21. 21

    Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard

    Julian Schulz, Lukas Fülle, Rieke Fruengel

    cs.AI · cs.CL · cs.LG

    Chain-of-thought monitoring as an approach for AI oversight and control is threatened by the possibility of steganographic reasoning, where LLMs conceal their reasoning inside innocuous-looking text. Two neighbouring capabilities, steganographic messaging (passing a concealed message) and encoded reasoning (reasoning in an illegible but unconcealed format), have already been shown to emerge under training pressures that occur in real...

    arxiv.org/abs/2609.39838 · PDF

  22. 22

    Safety of Latent Communication in Multi-Agent Systems

    Muhammad Huzaifa, Sina Mavali, Thorsten Eisenhofer

    cs.AI · cs.LG · cs.MA

    Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication...

    arxiv.org/abs/2609.39788 · PDF

  23. 23

    OverForge: Reasoning Through Strategies and Tactics Helps Cooperative Lifelong Adaptation

    Oana Madalina Fron, Ojas Shirekar, Chirag Raman

    cs.AI · cs.CL · cs.MA

    Cooperative language-model agents must coordinate over long horizons and adapt to changing environments and to partners with unfamiliar conventions, yet existing agents map observations to actions without separating persistent coordination strategies from their tactical execution. We introduce OverForge, a training-free hierarchical architecture that separates strategic reasoning over roles and divisions of labour from tactical reasoning over...

    arxiv.org/abs/2609.39727 · PDF

  24. 24

    Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

    Serhii Zabolotnii

    cs.AI · cs.CY · cs.SE

    Benchmarks, audits, and agent protocols describe performance, permissions, and repair, but not how observed evidence should change an agent's authority during a consequential task. We call this the assurance-transition gap. We propose a Runtime Assurance Contract (RAC), a policy-level formal schema binding autonomy boundaries, component eligibility, evidence state, transition policy, human-review capacity, and non-compensatory gates. Under...

    arxiv.org/abs/2609.39717 · PDF

  25. 25

    ArchitectureIQ: On the Measure of Training Intuition

    Zirui Ren, Shaoyang Guo, Chencheng Tang, Jinxin Wang, Chengyu Xiong, Shanbin Yu, Peihang Li, Yidi Wu, Bangzhe Huang,...

    cs.AI

    Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuition of LLMs and humans, we introduce the ArchitectureIQ benchmark. Each question presents a synthetic dataset and several training recipes, and the test-taker is asked to predict the recipe yielding the best test metric. Overall, we find that LLMs' model intuition is good but has four...

    arxiv.org/abs/2609.39714 · PDF

  26. 26

    A helps B while B hurts A: directed transfer in instruction-tuning mixture

    Nima H. Siboni, Vahid Rostami

    cs.AI · cs.CL · cs.LG

    Adapting a language model to a specialized corpus means choosing which instruction-tuning tasks to train on under a fixed budget, and testing one choice costs a fine-tuning run. Common heuristics add more source tasks or pick sources similar to the target. The first assumes transfer is never negative; the second, that it is symmetric. We show that both assumptions fail: task $A$ can help task $B$ while $B$ hurts $A$, so helpfulness is a...

    arxiv.org/abs/2609.39702 · PDF

  27. 27

    Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering

    Jiale Dai, Hongcan Deng, Liuxian Ma, Xiaoke Niu, Guojie Song

    cs.AI · cs.CY · cs.LG

    Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and...

    arxiv.org/abs/2609.39701 · PDF

  28. 28

    ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning

    Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari, Koushil Sreenath, Marc...

    cs.AI

    Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that...

    arxiv.org/abs/2609.39665 · PDF

  29. 29

    Why Do Conventional World Models Fail to Learn Cellular Automata?

    Shaoyang Guo, Ziming Liu

    cs.AI

    Although conventional world models - auto-regressive or diffusion models based on transformers or convolutional networks - may learn surface statistics of world dynamics, can they learn the exact world dynamics from its observed history? Leveraging cellular automata as a simple testbed, we find the answer to be no in many cases. Conventional architectures predict most pixels correctly yet rarely complete a rollout: a CNN predicts 96.3% of...

    arxiv.org/abs/2609.39604 · PDF

  30. 30

    AVERT-VLN: Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation

    Minrui Liu, Jingke Wang, Yuehao Huang, Hao Su, Jiajun Lv, Yukai Ma, Yong Liu

    cs.AI

    Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relying on continuous human supervision, a practical strategy is to selectively request corrective guidance, recover the ongoing task, and reuse corrective interactions to improve subsequent navigation. We propose Abstention-aware Visual Error...

    arxiv.org/abs/2609.39579 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.