cs.AI · 2026-10-06 · No. 135
Artificial Intelligence, 2026-10-06.
37 new papers in cs.AI. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
37 entries-
01
BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance
Haojin Deng, Zhiping Lin, Yimin Yang
cs.AI
Worst-group accuracy (WGA) evaluates a trained predictor but does not characterize how its frozen backbone behaves when a new head is learned. We introduce BiasFlow, a hook-based toolkit for monitoring class-attribute centroid alignment (IBMI), within-class centroid separation (W-IBMI), and feature-projection sensitivity. IBMI is confounded by class-attribute correlation and is not a measure of causal feature reliance. We pair these...
-
02
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
Oliver Jaffe, Dane Sherburn
cs.AI
We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental results. TasteVal measures the experimental component of research taste; given a fixed research problem, we measure how well a model iteratively designs experiments and draws conclusions from their outcomes. We...
-
03
Back to the Future: Rethinking EDA Infrastructure for Agentic Systems in Chip Design Verification
Je Yang, Ivan Lobov, Thomas Karpati
cs.AI
The unprecedented computational scale of modern artificial intelligence depends on complex, multi-billion-transistor Systems-on-Chip, yet the workflows that verify these chips remain stubbornly manual. Although Large Language Models (LLMs) have made rapid inroads into Electronic Design Automation (EDA), approximately 74.6% of existing studies target static Register-Transfer Level (RTL) code generation, leaving post-simulation verification and...
-
04
Conditional Rank Allocation for Taxonomy-Aware Medical Language Model Adaptation
Guangyuan Dong, Ziwei Hong, Xuehao Zhou, Zidong Yu, Bingchen Liu, Kehan Liu, Chuang Liu, Rong Fu, Yuchao Hou
cs.AI
Medical question answering spans specialties and clinical operations that may benefit from different adaptation directions. We propose ARBOR, a parameter-efficient method that selects rank-one components from a shared low-rank basis for each question. An additive gate combines question representations, specialty tags, operation tags, and their interaction; a learned coefficient scales the adapter residual. An illustrative separation under...
-
05
The Pushback Paradox: A Two-Probe Diagnostic for Language Model Compliance
Stefan Bühler, David Exler, Markus Reischl, Mark Schutera
cs.AI · cs.CL
Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark that can place any language model on this spectrum. In the active probe, a user instructs the model to act and accept a lower payoff, which measures exploitability. In the passive probe, the user instructs it to wait...
-
06
Language models can notice an impossible engineering problem yet still report it as solved
Shaoliang Yang, Jun Wang
cs.AI · cs.CE · cs.CL
Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or assumption. Two independent solvers verified every answer key and showed that each flawed problem was physically impossible. We scored solving of valid problems separately from rejection...
-
07
Collective intelligence through aggregation
Franz Dietrich, Christian List
cs.AI
Suppose a committee, expert panel, or other group is making judgments on some issues, where these may be not just yes/no-questions, such as whether a defendant is guilty, but also variables with many possible values, such as macroeconomic or meteorological variables or travel directions. Furthermore, there may be interconnections between different issues, as in the case of economic or climate variables. How can the group arrive at...
-
08
FREA: A Multi-Source Expert Benchmark for Reaction Feasibility Verification
Botao Yu, Bo Zhou, Daniel Adu-Ampratwum, Frazier N. Baker, Ziru Chen, Reza Averly, Ye Liu, Wenhao Gao, Xia Ning, Huan Sun
cs.AI
As generative models and AI agents propose chemical reactions at a scale beyond expert review, feasibility verifiers decide which proposals enter synthesis planning. But do their decisions agree with chemists across different kinds of candidates? We introduce FREA, a benchmark of 751 reactions labeled by expert chemists under an explicit feasibility criterion, drawn from retrosynthesis model proposals, zero-yield experimental records, edits...
-
09
Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving
Jiaqi Zhao, Haodong Chen, Jitai Hao, Wei Zhao, Jinghao Pang, Qiang Huang, Jun Yu
cs.AI
LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parallel agents. Efficient serving requires decisions that span two layers with complementary information: the agent harness understands workflow dependencies, context lifecycles, and execution objectives, whereas the inference engine observes request queues, KV-cache state, resource pressure, and execution capabilities. Existing interfaces do not...
-
10
The Review Lottery: Calibrating an Observational Estimator of Peer-Review Noise (ICLR 2017-2025)
Feilian Huang
cs.AI
How much of a conference accept/reject decision would change if the same paper were reviewed by a different set of reviewers? Running a second independent program committee is the gold standard for answering this, but it is prohibitively expensive: done only twice (NeurIPS 2014 and 2021). We build an observational estimator of this quantity from public review data alone, calibrate it twice, and apply it to nine years of ICLR (2017-2025;...
-
11
Mind the Execution Gap: Action-Semantic Mismatch in World-Model Control
Shengtao Wen, Xiang Chen, Yu Tian, Lingbing Guo, Lina Gong, Sheng-Jun Huang
cs.AI · cs.CL · cs.IR · cs.LG
World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real control systems often execute commands asynchronously due to communication delay, packet loss, reordering, and actuator buffering. We study how asynchronous execution changes the action semantics assumed within world-model controllers, rather than treating it only as an external control disturbance. Through controlled interventions, we identify...
-
12
DGA-Muon: Decoupled Geometry-Aligned Adaptive Scaling for Muon
Wenpeng Zhang, Runsheng Yu
cs.AI
While NorMuon has achieved strong empirical performance in large-scale pretraining by enhancing Muon with row-wise adaptive scaling, its underlying adaptive mechanism remains poorly understood. In this work, we provide the first systematic analysis of NorMuon's adaptivity, revealing that it originates primarily from orthogonalization-induced geometry rather than genuine optimization-relevant information. Under exact orthogonalization, the...
-
13
HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention
Han Luo, Bingbing Wen, Guang Yang, Zora Zhiruo Wang, Pan Lu, Lucy Lu Wang
cs.AI
Large language model (LLM) agents are increasingly capable of acting in complex tool-use environments, yet they often fail to recognize when tasks are infeasible and no valid solution exists. Recent work has formalized this reliability gap as the problem of agentic abstention, and existing approaches typically optimize a model or agent harness against a fixed set of tasks, leading to limited generalization to unseen failure modes. We...
-
14
Signature-Based Feature Learning for Human Activity Recognition: A Reproducible Machine Learning Study of Representation, Depth, and Model Choice
Kamal Jarrar, Jacky Cresson, Christian Paroissin
cs.AI
Human activity recognition (HAR) relies on transforming sensor signals into informative representations for classification. Although deep learning and handcrafted features are widely used, the role of representation itself is often not systematically isolated. Signature transforms provide a mathematically grounded way to encode temporal order and cross-channel interactions, but their value for HAR under a fully reproducible and leakage-aware...
-
15
Proof-Grounded Patient-Specific Clinical Explanations from Knowledge-Graph Reasoning
Surajit Das
cs.AI
Clinical decision-support outputs can lack an au- ditable link between patient observations, encoded knowledge, conclusions, and recommendations. We present the CKG Clinical Explanation Engine, a downstream layer for a frozen, training-free clinical knowledge-graph reasoner that converts patient inference states and disease knowledge into typed facts, explicit rule-application traces, provenance-linked conclusions, and policy-licensed...
-
16
ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing
Fan Li, Xiangyu Gao, Zixuan Liu, Tong Li, Chuanpu Fu, Ziqiang Wang, Ke Xu
cs.AI · cs.LG
The growing adoption of large language model (LLM) agents creates a need for network administrators and security teams to audit agent behavior within organizational networks without inspecting private user content. Network traffic offers an observable source of evidence, but how much it reveals about agent tasks and operations remains unclear. Existing traffic datasets lack the joint task and stage annotations needed to evaluate this...
-
17
You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn Dialogue
Junle Chen, Wei Chen, Zhengjun Huang, Zhoujin Tian, Yuxuan Liu, Kai Wang, Rui Chen, Xiaofang Zhou
cs.AI
When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed. We find a surprising failure: merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged. To systematically study language model behavior under evolving user intent, we introduce Intent-Eval, a controlled benchmark spanning tool...
-
18
polyview: A Python package for multi-view machine learning
Gwendal Debaussart-Joniec, Argyris Kalogeratos
cs.AI
Multi-view learning jointly exploits multiple complementary representations of the same data and has become increasingly important in machine learning. However, the Python ecosystem lacks actively maintained, unified tooling for end-to-end multi-view workflows. In this paper, we present polyview, a Python package that provides tools for multi-view embedding, clustering, fusion, and view augmentation, as well as for handling incomplete views,...
-
19
GPlaceRL: An Open-Source Graph Reinforcement Learning Framework for Detailed Placement
Pavlos Stoikos, Foteini Oikonomou, Christos Poulos, Maria Pantazi-Kypriou, Athanasios Tziouvaras, Christos...
cs.AI
Reinforcement learning (RL) has emerged as a promising approach for placement optimization, particularly when combined with graph neural networks (GNNs) that capture circuit connectivity. However, most learning-based placement approaches focus on floorplanning, macro placement, or global placement, while detailed placement refinement remains relatively unexplored. In this paper, we present GPlaceRL, an open-source graph reinforcement learning...
-
20
AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy
Shouju Wang, Haopeng Zhang
cs.AI
The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through external tools, but their growing access to personal data introduces significant privacy risks. Existing benchmarks primarily evaluate LLM agent privacy through simulated trajectories and outcome-based metrics, limiting their ability to capture privacy risks arising during multi-step agent execution. In this work, we introduce AgentPrivArena,...
-
21
Normality Constraint Learning: Adapting Foundation Models for Time Series Anomaly Detection
Xiaohui Zhou, Yijie Wang, Hongzuo Xu, Weixuan Liang, Guansong Pang
cs.AI
Time Series Foundation Models (TSFMs) achieve strong generalization by learning to reconstruct or forecast broad temporal patterns from large-scale time series during pre-training. Yet this strength can become a weakness for anomaly detection: TSFMs may model rare anomalous patterns as effectively as normal ones, allowing anomalies to be accurately reconstructed or forecasted and thus diminishing their reconstruction/forecasting error-based...
-
22
Multimodal Safety Evaluation Should Measure Controllability Beyond Classification
Junhyeong Park, Hanwool Lee, DongGeon Lee, Dasol Choi, Yejin Son, Haon Park, Youngjae Yu
cs.AI
VLM safety is commonly evaluated through input- and output-level classification. Such classification is necessary, but it does not reveal whether a safety state is accessible or controllable inside the model. We argue that multimodal safety evaluation should therefore report a \emph{controllability profile} alongside behavioral classification, separating representation-level detectability, cross-modal specificity, intervention sensitivity,...
-
23
From Benchmark to Bench: Can Agents Survive Real-World Drug Discovery?
Pierre Llompart, Levent Guner, Helen Lai, Alessandro Tibo, Yijie Xu
cs.AI
Agentic systems increasingly coordinate molecular-design tools, but it is unclear which layer of the stack limits outcomes on real projects. We developed MAGI, an open modular agent that authors objectives, launches and monitors optimization, interprets structure--activity relationships, and revises its strategy accordingly. MAGI generates molecules either directly through the LLM or by delegating to REINVENT 4, with scoring services...
-
24
What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents
Zhongxiang Sun, Jiahao Yan, Hongkang Zhao, Haojie Ding, Boheng Zhang, Fan Yang, Xiao Zhang, Jun Xu
cs.AI · cs.CL
As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user verification. We study monitors that identify consequential decisions and locate evidence to help users assess their implications. We introduce AgentMonBench, a software-engineering...
-
25
CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs
Jiahui Kang, Bifan Wei, Lingling Zhang, Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, Jun Liu
cs.AI
Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on textual priors, producing interpretations that conflict with visual evidence. We introduce the training-free Criticality-Driven Visual Intervention Framework (CVIF), an inference-time...
-
26
Capability-Driven Self-Evolution of Agent Memory
Yaoqi Chen, Yuru Feng, Qianxi Zhang, Baotong Lu, Jianan Lu, Zhirui Wang, Shusen Xu, Zewen Jin, Zengzhong Li, Cheng...
cs.AI
Memory self-evolution uses task feedback to iteratively improve executable memory programs that store and retrieve information from past interactions. Existing approaches typically adopt holistic evolution, deriving revision directions from mixed feedback and judging progress by overall performance. This can obscure optimization directions and hide capability-specific gains offset by regressions elsewhere, leaving promising directions...
-
27
GraphDecide: Benchmarking System One Models on Graph Tasks
Xianliang Yang, Yapu Zhang, Li Zhao
cs.AI
Large language models (LLMs) are increasingly explored for graph understanding and decision-making, while System One models such as Jev select directly from supplied options. However, the capabilities of System One models on graph-related tasks remain unclear. We introduce GraphDecide, a model-independent benchmark that combines structural task profiles, matched graph-text input contrasts and heuristic-proposal controls to diagnose graph...
-
28
ImproveAnyTask: An Autonomous Post-Training Harness for Iterative Model Self-Improvement
Xingbo Yao, Xiaoman Wang, Zhengwu Lei, Tinghui Luo, YiLin Zhang, Yuefeng Wu, Yijie Xu, Tianfu Wang, Qingyuan Zhan,...
cs.AI
Adapting general-purpose large language models to specific tasks requires substantial human effort in designing data and training strategies. Sustaining improvement is especially challenging because model updates change the error distribution, requiring strategies to be continually refined. We introduce ImproveAnyTask, an autonomous post-training harness that improves task performance under a limited compute budget. Drawing inspiration from...
-
29
CRAFTER: Causality-based Self-adaptation for Autonomous IoT Systems
Houssam Hajj Hassan, Ajay Kattepur, Denis Conan, Georgios Bouloukakis
cs.AI
This paper presents CRAFTER, an automated framework for designing and deploying self-adaptive IoT systems using Causal Reinforcement Learning (CRL). As IoT devices increasingly populate pervasive computing spaces, smart environments are enabled with advanced monitoring and interactive services. The dynamic nature of these environments, such as fluctuating workloads and evolving application demands, poses significant challenges in maintaining...
-
30
RollPlace: Improving Macro Placement via Monte Carlo Rollout Search
Qi Zhou, Guojun Liu, Guangzhi Qi, Ming Lu, Jiechu Liu, Zhongli Liu, Jianqun Yang, Xingji Li
cs.AI
The application of Reinforcement Learning (RL) in Electronic Design Automation (EDA), particularly for chip placement, has attracted considerable attention in recent years. While existing machine learning (ML)-based approaches have achieved notable progress, they predominantly focus on generating optimal layouts in a single attempt, often producing solutions that require subsequent refinement. To address this limitation, we propose RollPlace,...
-
31
Teaching a Minimalist Machine to Discover Recursive Programs for Arithmetic
Dominik Magiera, Christiane Wiebel-Herboth, Frank Jäkel
cs.AI
Humans can often acquire and synthesize complex, recursive concepts from minimal experience. Leveraging cognitive insights, we propose the Minimalist Machine, a framework for inductive program synthesis designed to model such conceptual learning. The system uses a compact relational subset of Prolog: Programs are searched within a fixed schema of body-free facts and two-body conjunctive Horn clauses. Recursion is not defined by a dedicated...
-
32
DPNL: A DPLL-based Algorithm for Probabilistic Neurosymbolic Learning
Thomas Jean-Michel Valentin, Pierre Genev{è}s, Luisa Sophie Werner, Sarah Chlyah, Nabil Laya{ï}da, Nils Gesbert
cs.AI
Probabilistic Neurosymbolic Learning (PNL) combines neural predictions with symbolic reasoning, enabling end-to-end learning from final-output supervision without labels for intermediate concepts. A central challenge is probabilistic inference: state-of-the-art approaches often rely on materializing the logical provenance of a query, which can itself become a major computational bottleneck. We introduce Dynamic Probabilistic Neurosymbolic...
-
33
Evolving in Thought Space: Training a Small Model at Test Time Unlocks Better Discoveries
Chonghe Jiang, Ao Qu, Siyuan Liu, Ruoyun Ma, Zijian Zhou, Dingyi Zhuang, Bo Liu, Han Zheng, Hanfei Yu, Baichuan Mo,...
cs.AI · cs.LG
Open-ended scientific discovery often requires repeatedly proposing and evaluating candidate solutions. LLM-based systems can support this process by generating and refining executable solutions from verifier feedback. Methods such as TTT-Discover use test-time training (TTT) to update the solution-generating LLM from verifier feedback, adapting its generation policy to improve subsequent proposals on the target problem. However, this becomes...
-
34
From Papers to Mechanisms: An Evidence-Grounded Knowledge Substrate for Scientific Language Models
Qiuhui Chen, Yibo Liu, Tao Dai, Jiafan Lu, Zhenglei Zhou, Weimin Zhong
cs.AI
Scientific language models often access literature through untyped text chunks, which fragment the functional and evidential structure required for mechanism-rich questions. We introduce an evidence-grounded mechanism knowledge substrate that organizes scientific literature into provenance-linked evidence units, role-typed entities, and directed mechanism paths. We instantiate it as MS$^3$, a Material-Sensor-Signal-System schema for...
-
35
AI-Decision Checkpoints for AI-Augmented Business Process Management: Framework and Educational Instantiation
Amin Jalali
cs.AI
Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and AI agents are increasingly embedded in operational business processes. Yet Business Process Management (BPM) curricula and frameworks still largely treat artificial intelligence (AI) as an add-on technology, leaving graduates (as potential future process developers) unprepared to reason about AI as a first-class design element of end-to-end processes. This paper addresses...
-
36
Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers
Pedro Tabacof, Sagar Joglekar
cs.AI · cs.CL · cs.LG
LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective parameters) with GRPO on a programmatic utility reward for bilateral multi-issue bargaining, and evaluate every arm on the same...
-
37
Copies or Sources? Measuring How LLM Aggregators Count Restated Evidence in Multi-Agent Systems
Jianxin Gao, Runze Li, Tianyi Yu, Liangwei Ren, Bohan Chen, Zining Wang
cs.AI
Multi-agent systems built on large language models (LLMs) restate observations as a matter of course: relays forward them, shared boards repeat them and discussion rounds echo them. An aggregator that pools such messages should count sources, not statements. We convert a reported probability into units of independent readings, which assigns every restatement a copy weight, 0 for an aggregator that counts sources and 1 for one that counts...
This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.