cs.AI · 2026-10-06 · No. 135

Artificial Intelligence, 2026-10-06.

37 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

37 entries
  1. 01

    BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance

    Haojin Deng, Zhiping Lin, Yimin Yang

    cs.AI

    Worst-group accuracy (WGA) evaluates a trained predictor but does not characterize how its frozen backbone behaves when a new head is learned. We introduce BiasFlow, a hook-based toolkit for monitoring class-attribute centroid alignment (IBMI), within-class centroid separation (W-IBMI), and feature-projection sensitivity. IBMI is confounded by class-attribute correlation and is not a measure of causal feature reliance. We pair these...

    arxiv.org/abs/2610.06846 · PDF

  2. 02

    TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts

    Oliver Jaffe, Dane Sherburn

    cs.AI

    We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental results. TasteVal measures the experimental component of research taste; given a fixed research problem, we measure how well a model iteratively designs experiments and draws conclusions from their outcomes. We...

    arxiv.org/abs/2610.06824 · PDF

  3. 03

    Back to the Future: Rethinking EDA Infrastructure for Agentic Systems in Chip Design Verification

    Je Yang, Ivan Lobov, Thomas Karpati

    cs.AI

    The unprecedented computational scale of modern artificial intelligence depends on complex, multi-billion-transistor Systems-on-Chip, yet the workflows that verify these chips remain stubbornly manual. Although Large Language Models (LLMs) have made rapid inroads into Electronic Design Automation (EDA), approximately 74.6% of existing studies target static Register-Transfer Level (RTL) code generation, leaving post-simulation verification and...

    arxiv.org/abs/2610.06790 · PDF

  4. 04

    Conditional Rank Allocation for Taxonomy-Aware Medical Language Model Adaptation

    Guangyuan Dong, Ziwei Hong, Xuehao Zhou, Zidong Yu, Bingchen Liu, Kehan Liu, Chuang Liu, Rong Fu, Yuchao Hou

    cs.AI

    Medical question answering spans specialties and clinical operations that may benefit from different adaptation directions. We propose ARBOR, a parameter-efficient method that selects rank-one components from a shared low-rank basis for each question. An additive gate combines question representations, specialty tags, operation tags, and their interaction; a learned coefficient scales the adapter residual. An illustrative separation under...

    arxiv.org/abs/2610.06765 · PDF

  5. 05

    The Pushback Paradox: A Two-Probe Diagnostic for Language Model Compliance

    Stefan Bühler, David Exler, Markus Reischl, Mark Schutera

    cs.AI · cs.CL

    Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark that can place any language model on this spectrum. In the active probe, a user instructs the model to act and accept a lower payoff, which measures exploitability. In the passive probe, the user instructs it to wait...

    arxiv.org/abs/2610.06673 · PDF

  6. 06

    Language models can notice an impossible engineering problem yet still report it as solved

    Shaoliang Yang, Jun Wang

    cs.AI · cs.CE · cs.CL

    Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or assumption. Two independent solvers verified every answer key and showed that each flawed problem was physically impossible. We scored solving of valid problems separately from rejection...

    arxiv.org/abs/2610.06668 · PDF

  7. 07

    Collective intelligence through aggregation

    Franz Dietrich, Christian List

    cs.AI

    Suppose a committee, expert panel, or other group is making judgments on some issues, where these may be not just yes/no-questions, such as whether a defendant is guilty, but also variables with many possible values, such as macroeconomic or meteorological variables or travel directions. Furthermore, there may be interconnections between different issues, as in the case of economic or climate variables. How can the group arrive at...

    arxiv.org/abs/2610.06652 · PDF

  8. 08

    FREA: A Multi-Source Expert Benchmark for Reaction Feasibility Verification

    Botao Yu, Bo Zhou, Daniel Adu-Ampratwum, Frazier N. Baker, Ziru Chen, Reza Averly, Ye Liu, Wenhao Gao, Xia Ning, Huan Sun

    cs.AI

    As generative models and AI agents propose chemical reactions at a scale beyond expert review, feasibility verifiers decide which proposals enter synthesis planning. But do their decisions agree with chemists across different kinds of candidates? We introduce FREA, a benchmark of 751 reactions labeled by expert chemists under an explicit feasibility criterion, drawn from retrosynthesis model proposals, zero-yield experimental records, edits...

    arxiv.org/abs/2610.06614 · PDF

  9. 09

    Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving

    Jiaqi Zhao, Haodong Chen, Jitai Hao, Wei Zhao, Jinghao Pang, Qiang Huang, Jun Yu

    cs.AI

    LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parallel agents. Efficient serving requires decisions that span two layers with complementary information: the agent harness understands workflow dependencies, context lifecycles, and execution objectives, whereas the inference engine observes request queues, KV-cache state, resource pressure, and execution capabilities. Existing interfaces do not...

    arxiv.org/abs/2610.06597 · PDF

  10. 10

    The Review Lottery: Calibrating an Observational Estimator of Peer-Review Noise (ICLR 2017-2025)

    Feilian Huang

    cs.AI

    How much of a conference accept/reject decision would change if the same paper were reviewed by a different set of reviewers? Running a second independent program committee is the gold standard for answering this, but it is prohibitively expensive: done only twice (NeurIPS 2014 and 2021). We build an observational estimator of this quantity from public review data alone, calibrate it twice, and apply it to nine years of ICLR (2017-2025;...

    arxiv.org/abs/2610.06591 · PDF

  11. 11

    Mind the Execution Gap: Action-Semantic Mismatch in World-Model Control

    Shengtao Wen, Xiang Chen, Yu Tian, Lingbing Guo, Lina Gong, Sheng-Jun Huang

    cs.AI · cs.CL · cs.IR · cs.LG

    World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real control systems often execute commands asynchronously due to communication delay, packet loss, reordering, and actuator buffering. We study how asynchronous execution changes the action semantics assumed within world-model controllers, rather than treating it only as an external control disturbance. Through controlled interventions, we identify...

    arxiv.org/abs/2610.06582 · PDF

  12. 12

    DGA-Muon: Decoupled Geometry-Aligned Adaptive Scaling for Muon

    Wenpeng Zhang, Runsheng Yu

    cs.AI

    While NorMuon has achieved strong empirical performance in large-scale pretraining by enhancing Muon with row-wise adaptive scaling, its underlying adaptive mechanism remains poorly understood. In this work, we provide the first systematic analysis of NorMuon's adaptivity, revealing that it originates primarily from orthogonalization-induced geometry rather than genuine optimization-relevant information. Under exact orthogonalization, the...

    arxiv.org/abs/2610.06578 · PDF

  13. 13

    HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention

    Han Luo, Bingbing Wen, Guang Yang, Zora Zhiruo Wang, Pan Lu, Lucy Lu Wang

    cs.AI

    Large language model (LLM) agents are increasingly capable of acting in complex tool-use environments, yet they often fail to recognize when tasks are infeasible and no valid solution exists. Recent work has formalized this reliability gap as the problem of agentic abstention, and existing approaches typically optimize a model or agent harness against a fixed set of tasks, leading to limited generalization to unseen failure modes. We...

    arxiv.org/abs/2610.06563 · PDF

  14. 14

    Signature-Based Feature Learning for Human Activity Recognition: A Reproducible Machine Learning Study of Representation, Depth, and Model Choice

    Kamal Jarrar, Jacky Cresson, Christian Paroissin

    cs.AI

    Human activity recognition (HAR) relies on transforming sensor signals into informative representations for classification. Although deep learning and handcrafted features are widely used, the role of representation itself is often not systematically isolated. Signature transforms provide a mathematically grounded way to encode temporal order and cross-channel interactions, but their value for HAR under a fully reproducible and leakage-aware...

    arxiv.org/abs/2610.06553 · PDF

  15. 15

    Proof-Grounded Patient-Specific Clinical Explanations from Knowledge-Graph Reasoning

    Surajit Das

    cs.AI

    Clinical decision-support outputs can lack an au- ditable link between patient observations, encoded knowledge, conclusions, and recommendations. We present the CKG Clinical Explanation Engine, a downstream layer for a frozen, training-free clinical knowledge-graph reasoner that converts patient inference states and disease knowledge into typed facts, explicit rule-application traces, provenance-linked conclusions, and policy-licensed...

    arxiv.org/abs/2610.06549 · PDF

  16. 16

    ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing

    Fan Li, Xiangyu Gao, Zixuan Liu, Tong Li, Chuanpu Fu, Ziqiang Wang, Ke Xu

    cs.AI · cs.LG

    The growing adoption of large language model (LLM) agents creates a need for network administrators and security teams to audit agent behavior within organizational networks without inspecting private user content. Network traffic offers an observable source of evidence, but how much it reveals about agent tasks and operations remains unclear. Existing traffic datasets lack the joint task and stage annotations needed to evaluate this...

    arxiv.org/abs/2610.06514 · PDF

  17. 17

    You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn Dialogue

    Junle Chen, Wei Chen, Zhengjun Huang, Zhoujin Tian, Yuxuan Liu, Kai Wang, Rui Chen, Xiaofang Zhou

    cs.AI

    When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed. We find a surprising failure: merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged. To systematically study language model behavior under evolving user intent, we introduce Intent-Eval, a controlled benchmark spanning tool...

    arxiv.org/abs/2610.06496 · PDF

  18. 18

    polyview: A Python package for multi-view machine learning

    Gwendal Debaussart-Joniec, Argyris Kalogeratos

    cs.AI

    Multi-view learning jointly exploits multiple complementary representations of the same data and has become increasingly important in machine learning. However, the Python ecosystem lacks actively maintained, unified tooling for end-to-end multi-view workflows. In this paper, we present polyview, a Python package that provides tools for multi-view embedding, clustering, fusion, and view augmentation, as well as for handling incomplete views,...

    arxiv.org/abs/2610.06491 · PDF

  19. 19

    GPlaceRL: An Open-Source Graph Reinforcement Learning Framework for Detailed Placement

    Pavlos Stoikos, Foteini Oikonomou, Christos Poulos, Maria Pantazi-Kypriou, Athanasios Tziouvaras, Christos...

    cs.AI

    Reinforcement learning (RL) has emerged as a promising approach for placement optimization, particularly when combined with graph neural networks (GNNs) that capture circuit connectivity. However, most learning-based placement approaches focus on floorplanning, macro placement, or global placement, while detailed placement refinement remains relatively unexplored. In this paper, we present GPlaceRL, an open-source graph reinforcement learning...

    arxiv.org/abs/2610.06489 · PDF

  20. 20

    AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy

    Shouju Wang, Haopeng Zhang

    cs.AI

    The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through external tools, but their growing access to personal data introduces significant privacy risks. Existing benchmarks primarily evaluate LLM agent privacy through simulated trajectories and outcome-based metrics, limiting their ability to capture privacy risks arising during multi-step agent execution. In this work, we introduce AgentPrivArena,...

    arxiv.org/abs/2610.06454 · PDF

  21. 21

    Normality Constraint Learning: Adapting Foundation Models for Time Series Anomaly Detection

    Xiaohui Zhou, Yijie Wang, Hongzuo Xu, Weixuan Liang, Guansong Pang

    cs.AI

    Time Series Foundation Models (TSFMs) achieve strong generalization by learning to reconstruct or forecast broad temporal patterns from large-scale time series during pre-training. Yet this strength can become a weakness for anomaly detection: TSFMs may model rare anomalous patterns as effectively as normal ones, allowing anomalies to be accurately reconstructed or forecasted and thus diminishing their reconstruction/forecasting error-based...

    arxiv.org/abs/2610.06453 · PDF

  22. 22

    Multimodal Safety Evaluation Should Measure Controllability Beyond Classification

    Junhyeong Park, Hanwool Lee, DongGeon Lee, Dasol Choi, Yejin Son, Haon Park, Youngjae Yu

    cs.AI

    VLM safety is commonly evaluated through input- and output-level classification. Such classification is necessary, but it does not reveal whether a safety state is accessible or controllable inside the model. We argue that multimodal safety evaluation should therefore report a \emph{controllability profile} alongside behavioral classification, separating representation-level detectability, cross-modal specificity, intervention sensitivity,...

    arxiv.org/abs/2610.06452 · PDF

  23. 23

    From Benchmark to Bench: Can Agents Survive Real-World Drug Discovery?

    Pierre Llompart, Levent Guner, Helen Lai, Alessandro Tibo, Yijie Xu

    cs.AI

    Agentic systems increasingly coordinate molecular-design tools, but it is unclear which layer of the stack limits outcomes on real projects. We developed MAGI, an open modular agent that authors objectives, launches and monitors optimization, interprets structure--activity relationships, and revises its strategy accordingly. MAGI generates molecules either directly through the LLM or by delegating to REINVENT 4, with scoring services...

    arxiv.org/abs/2610.06411 · PDF

  24. 24

    What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents

    Zhongxiang Sun, Jiahao Yan, Hongkang Zhao, Haojie Ding, Boheng Zhang, Fan Yang, Xiao Zhang, Jun Xu

    cs.AI · cs.CL

    As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user verification. We study monitors that identify consequential decisions and locate evidence to help users assess their implications. We introduce AgentMonBench, a software-engineering...

    arxiv.org/abs/2610.06406 · PDF

  25. 25

    CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs

    Jiahui Kang, Bifan Wei, Lingling Zhang, Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, Jun Liu

    cs.AI

    Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on textual priors, producing interpretations that conflict with visual evidence. We introduce the training-free Criticality-Driven Visual Intervention Framework (CVIF), an inference-time...

    arxiv.org/abs/2610.06399 · PDF

  26. 26

    Capability-Driven Self-Evolution of Agent Memory

    Yaoqi Chen, Yuru Feng, Qianxi Zhang, Baotong Lu, Jianan Lu, Zhirui Wang, Shusen Xu, Zewen Jin, Zengzhong Li, Cheng...

    cs.AI

    Memory self-evolution uses task feedback to iteratively improve executable memory programs that store and retrieve information from past interactions. Existing approaches typically adopt holistic evolution, deriving revision directions from mixed feedback and judging progress by overall performance. This can obscure optimization directions and hide capability-specific gains offset by regressions elsewhere, leaving promising directions...

    arxiv.org/abs/2610.06361 · PDF

  27. 27

    GraphDecide: Benchmarking System One Models on Graph Tasks

    Xianliang Yang, Yapu Zhang, Li Zhao

    cs.AI

    Large language models (LLMs) are increasingly explored for graph understanding and decision-making, while System One models such as Jev select directly from supplied options. However, the capabilities of System One models on graph-related tasks remain unclear. We introduce GraphDecide, a model-independent benchmark that combines structural task profiles, matched graph-text input contrasts and heuristic-proposal controls to diagnose graph...

    arxiv.org/abs/2610.06354 · PDF

  28. 28

    ImproveAnyTask: An Autonomous Post-Training Harness for Iterative Model Self-Improvement

    Xingbo Yao, Xiaoman Wang, Zhengwu Lei, Tinghui Luo, YiLin Zhang, Yuefeng Wu, Yijie Xu, Tianfu Wang, Qingyuan Zhan,...

    cs.AI

    Adapting general-purpose large language models to specific tasks requires substantial human effort in designing data and training strategies. Sustaining improvement is especially challenging because model updates change the error distribution, requiring strategies to be continually refined. We introduce ImproveAnyTask, an autonomous post-training harness that improves task performance under a limited compute budget. Drawing inspiration from...

    arxiv.org/abs/2610.06347 · PDF

  29. 29

    CRAFTER: Causality-based Self-adaptation for Autonomous IoT Systems

    Houssam Hajj Hassan, Ajay Kattepur, Denis Conan, Georgios Bouloukakis

    cs.AI

    This paper presents CRAFTER, an automated framework for designing and deploying self-adaptive IoT systems using Causal Reinforcement Learning (CRL). As IoT devices increasingly populate pervasive computing spaces, smart environments are enabled with advanced monitoring and interactive services. The dynamic nature of these environments, such as fluctuating workloads and evolving application demands, poses significant challenges in maintaining...

    arxiv.org/abs/2610.06320 · PDF

  30. 30

    RollPlace: Improving Macro Placement via Monte Carlo Rollout Search

    Qi Zhou, Guojun Liu, Guangzhi Qi, Ming Lu, Jiechu Liu, Zhongli Liu, Jianqun Yang, Xingji Li

    cs.AI

    The application of Reinforcement Learning (RL) in Electronic Design Automation (EDA), particularly for chip placement, has attracted considerable attention in recent years. While existing machine learning (ML)-based approaches have achieved notable progress, they predominantly focus on generating optimal layouts in a single attempt, often producing solutions that require subsequent refinement. To address this limitation, we propose RollPlace,...

    arxiv.org/abs/2610.06316 · PDF

  31. 31

    Teaching a Minimalist Machine to Discover Recursive Programs for Arithmetic

    Dominik Magiera, Christiane Wiebel-Herboth, Frank Jäkel

    cs.AI

    Humans can often acquire and synthesize complex, recursive concepts from minimal experience. Leveraging cognitive insights, we propose the Minimalist Machine, a framework for inductive program synthesis designed to model such conceptual learning. The system uses a compact relational subset of Prolog: Programs are searched within a fixed schema of body-free facts and two-body conjunctive Horn clauses. Recursion is not defined by a dedicated...

    arxiv.org/abs/2610.06304 · PDF

  32. 32

    DPNL: A DPLL-based Algorithm for Probabilistic Neurosymbolic Learning

    Thomas Jean-Michel Valentin, Pierre Genev{è}s, Luisa Sophie Werner, Sarah Chlyah, Nabil Laya{ï}da, Nils Gesbert

    cs.AI

    Probabilistic Neurosymbolic Learning (PNL) combines neural predictions with symbolic reasoning, enabling end-to-end learning from final-output supervision without labels for intermediate concepts. A central challenge is probabilistic inference: state-of-the-art approaches often rely on materializing the logical provenance of a query, which can itself become a major computational bottleneck. We introduce Dynamic Probabilistic Neurosymbolic...

    arxiv.org/abs/2610.06270 · PDF

  33. 33

    Evolving in Thought Space: Training a Small Model at Test Time Unlocks Better Discoveries

    Chonghe Jiang, Ao Qu, Siyuan Liu, Ruoyun Ma, Zijian Zhou, Dingyi Zhuang, Bo Liu, Han Zheng, Hanfei Yu, Baichuan Mo,...

    cs.AI · cs.LG

    Open-ended scientific discovery often requires repeatedly proposing and evaluating candidate solutions. LLM-based systems can support this process by generating and refining executable solutions from verifier feedback. Methods such as TTT-Discover use test-time training (TTT) to update the solution-generating LLM from verifier feedback, adapting its generation policy to improve subsequent proposals on the target problem. However, this becomes...

    arxiv.org/abs/2610.06269 · PDF

  34. 34

    From Papers to Mechanisms: An Evidence-Grounded Knowledge Substrate for Scientific Language Models

    Qiuhui Chen, Yibo Liu, Tao Dai, Jiafan Lu, Zhenglei Zhou, Weimin Zhong

    cs.AI

    Scientific language models often access literature through untyped text chunks, which fragment the functional and evidential structure required for mechanism-rich questions. We introduce an evidence-grounded mechanism knowledge substrate that organizes scientific literature into provenance-linked evidence units, role-typed entities, and directed mechanism paths. We instantiate it as MS$^3$, a Material-Sensor-Signal-System schema for...

    arxiv.org/abs/2610.06248 · PDF

  35. 35

    AI-Decision Checkpoints for AI-Augmented Business Process Management: Framework and Educational Instantiation

    Amin Jalali

    cs.AI

    Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and AI agents are increasingly embedded in operational business processes. Yet Business Process Management (BPM) curricula and frameworks still largely treat artificial intelligence (AI) as an add-on technology, leaving graduates (as potential future process developers) unprepared to reason about AI as a first-class design element of end-to-end processes. This paper addresses...

    arxiv.org/abs/2610.06207 · PDF

  36. 36

    Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers

    Pedro Tabacof, Sagar Joglekar

    cs.AI · cs.CL · cs.LG

    LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective parameters) with GRPO on a programmatic utility reward for bilateral multi-issue bargaining, and evaluate every arm on the same...

    arxiv.org/abs/2610.06204 · PDF

  37. 37

    Copies or Sources? Measuring How LLM Aggregators Count Restated Evidence in Multi-Agent Systems

    Jianxin Gao, Runze Li, Tianyi Yu, Liangwei Ren, Bohan Chen, Zining Wang

    cs.AI

    Multi-agent systems built on large language models (LLMs) restate observations as a matter of course: relays forward them, shared boards repeat them and discussion rounds echo them. An aggregator that pools such messages should count sources, not statements. We convert a reported probability into units of independent readings, which assigns every restatement a copy weight, 0 for an aggregator that counts sources and 1 for one that counts...

    arxiv.org/abs/2610.06192 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.