cs.AI · 2026-10-02 · No. 131
Artificial Intelligence, 2026-10-02.
38 new papers in cs.AI. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
38 entries-
01
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha...
cs.AI · cs.CL · cs.IR
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author...
-
02
VISTA: A Visual Harness for Reasoning in an Interactive World
Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He
cs.AI · cs.CV
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations...
-
03
A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia
cs.AI · cs.CL
A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic...
-
04
Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints
Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir
cs.AI · cs.CR
Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applying FHE to reinforcement learning (RL) requires replacing non-linear operations with polynomial approximations, which diverge catastrophically due to...
-
05
PyPottery: an AI-powered end-to-end suite for pottery processing and publication
Lorenzo Cardarelli
cs.AI
The study of ceramic materials constitutes a cornerstone of archaeological research, yet the post-production workflow for pottery documentation remains labor-intensive and creates significant publication bottlenecks. This paper presents PyPottery, an open-source, AI-powered suite designed to semi-automate the complete ceramic documentation pipeline. The suite comprises four integrated modules: PyPotteryScan for automated image extraction and...
-
06
Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval
Arman Behnam, Binghui Wang
cs.AI
Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories. When a memory is never retrieved, store-level interventions produce identical outcomes, leaving its utility unidentified. This is a retrieval-level positivity violation, invisible to diagnostics that examine only memory...
-
07
External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing
Kingshuk Gupta, Davide Buscaldi
cs.AI
As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary classification task, failing to capture the structured, sequential boundaries of semantic drift. Here, we introduce an internal...
-
08
HydroJEV: A one-second, training-free screen for cyber-attack and fault attribution in water distribution networks
Tianwei Mu, Shengyan Jiang, Mingzhe Yuan, Qing Luo, Min Xiao, Wenhong Wang, Jun Li, Manhong Huang
cs.AI
When a SCADA alarm is raised in a water distribution network, operators must decide quickly whether it reflects a cyberattack, a physical fault, a normal transient or a faulty sensor. Supervised classifiers need labelled incidents that utilities rarely have, and frontier large language models (LLMs) take tens of seconds per decision. We tested whether Jev, a training-free model that returns class probabilities in about one second, can serve...
-
09
Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control
Yimeng Liu, Mi Zhang, Younsuk Dong, Zhichao Cao
cs.AI
Large language model (LLM) agents increasingly combine reasoning, tool use, and action, but most evidence comes from episodic tasks with relatively immediate feedback and reset failures. Long-running physical control operates in a different regime: actions alter future states, errors compound across decisions, and an agent must improve from experience without being allowed to rewrite the physical rules that make execution safe. We study this...
-
10
Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration
Xin Heng
cs.AI
AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence. Our Observation-Aliasing Impossibility Theorem gives the exact boundary. A policy can guarantee a valid action exactly when all worlds producing the same observation share an admissible action. If k indistinguishable worlds require pairwise-disjoint...
-
11
SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL
Hyeonmin Lee, Zheng Wei, Kyungmin Kwon, Jumin Seo, Jiwon Park, Hayoung Oh
cs.AI · cs.HC
While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated synthesis into continuous human-AI co-creation. SPHERE extracts persistent spatial preferences from natural multimodal interactions (speech and...
-
12
Atoms to Processes: The Role of Artificial Intelligence and Machine Learning in Chemical Engineering
Michael Baldea, Linda J. Broadbelt, Marianthi G. Ierapetritou, Akhilesh Jain, Ankur Kumar, Thomas A. Kwan, Fèlix...
cs.AI
The rapid maturation of artificial intelligence (AI) and machine learning (ML) has catalyzed a profound shift in how chemical engineering problems are formulated, analyzed, and solved. Advances in computing, data availability, and learning algorithms have enabled AI/ML methods to impact applications spanning atomic-scale simulations, materials and catalyst discovery, transport and thermodynamics, separations, process systems engineering, and...
-
13
Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries
Ionel Eduard Stan, Paolo Napoletano
cs.AI
A multi-LLM \emph{council} lets several large language models (LLMs) deliberate on a question and return an answer together with a confidence estimate. As these systems become increasingly used for reasoning, that confidence should represent a calibrated \emph{probability of being correct}, and the decision should remain robust when some agents are persistently unreliable. Existing \emph{council aggregation} methods fail on both fronts: their...
-
14
Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks
Hao Wang, Ting Huang
cs.AI · cs.CL
Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned. We present evidence, from a controlled single-machine comparison and one third-party benchmark, that a substantial share of these failures is attributable to the harness rather than the model. We...
-
15
Can AI Oversight Be Zero Knowledge?
Alessandro Chiesa, Ziyi Guan, Burcu Yildiz
cs.AI · cs.CC · cs.CR
AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on...
-
16
Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage
Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
cs.AI · cs.CL · cs.LG
Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted...
-
17
Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning
Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR
cs.AI
Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG...
-
18
AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes
Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley
cs.AI · cs.SD · eess.AS
Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset. To construct the dataset,...
-
19
Temporal-Difference Learning for Dragonchess
Jim O'Connor, Annika Hoag, Sarah Goyette, Gary B. Parker
cs.AI
Our research investigates how two adaptive AI methods, evolutionary transfer learning and TD(lambda), perform in the three-dimensional chess environment Dragonchess. The game challenges players with its unique board structure and computational load, making it an ideal setting to study how adaptive methods can update evaluation heuristics in novel environments. In this work we re-implement the Dragonchess engine, changing it from a PyGame...
-
20
On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models
Haochen Zhang, Jiaheng Guo, Zhen Xu, Zachary Plotkin, Nicholas Konz, Zhen Tan, Tianlong Chen
cs.AI
A time series world model (TSWM) predicts a controlled system's state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but their responses to changed plans remain untested. We ask which design choices matter and whether accurate forecasters respond to...
-
21
Code Owns the Simulation, Jev Owns the Evaluation
Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang
cs.AI · cs.LG
Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the...
-
22
Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills
Ngoc Phuoc An Vo, Aarya Doshi, Vadim Sheinin
cs.AI · cs.SE
Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determination skill in an enterprise Value Aware Resiliency system. The framework independently computes per-run ground truth,...
-
23
AI-assisted mitotic counting improves reproducibility and efficiency across multiple tumour types
Simon Graham, Mostafa Jahanifar, Quoc Dang Vu, Vygante Maskoliunaite, Donatas Petroska, Ruta Barbora Valkiuniene,...
cs.AI
Mitotic counting is an important component of tumour grading, diagnosis and prognostic assessment across several tumour types, but manual assessment is time-consuming and subject to inter-pathologist variability. To help address these challenges, we developed MitPro, an AI tool designed to improve consistency and efficiency by directing pathologists towards regions with the highest predicted mitotic activity and highlighting mitotic figures...
-
24
LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification
Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen
cs.AI
Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were designed for other modalities and other tasks, and they transfer poorly to open-ended generation in the time series domain. We address this by...
-
25
Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents
Beining Wu, Zihao Ding, Jun Huang
cs.AI · cs.GR
Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of experience: a trajectory bundles items with different properties, so a conclusion about the bundle depends on its mix. To address this, (i) we introduce component routing, which splits the experience into locators,...
-
26
Q-Learning for Reachability in MEC-Free MDPs
Lu-Chin Chang, Suguman Bansal
cs.AI · cs.LG · cs.LO
Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of...
-
27
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
Arman Behnam, Sunglyoung Kim, Liangwei Yang
cs.AI
A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person's record, and such records are private, so benchmarks generate the person and the questions and settle in advance what matters. We release \bench, ten real relationships with an AI companion: 27,218 messages over up...
-
28
VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding
Bingjun Luo, Yuhuan Fan, Jialin Guo, Siqi Li
cs.AI · cs.CV
Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions are refined. Manually refining these harnesses requires diagnosing grounding failures and coordinating changes to both agent workflows and instructions. We introduce VideoEvolve, a framework that automatically...
-
29
TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference
Mukund Agarwalla, Chih-Jen Lin
cs.AI
Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same...
-
30
vFedProtoQNAS: Prototype-Guided Personalized Quantum Neural Architecture Search for Virtual Federated Learning
Seok Bin Son, Samuel Yen-Chi Chen, Soohyun Park, Joongheon Kim
cs.AI
Quantum federated learning (QFL) has emerged as a promising approach for collaboratively training compact quantum neural networks (QNNs) over distributed private data on resource-constrained devices. However, differences in device capabilities make a single shared QNN architecture unsuitable for all clients. While personalized quantum neural architecture search (QNAS) allows each client to select a device-specific QNN, averaging parameters...
-
31
CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement
Dongwei Sun, Yujie Zhang, Bowen Yao, Pei Liu, Jing Yao, Xiangyong Cao
cs.AI · cs.CV
Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect...
-
32
Measuring the Stability Assumption Behind Action Chunking
Aryan Goyal
cs.AI
Action chunking improves the performance of policies learned by behavioural cloning, and several mechanisms have been proposed to explain why, including temporal consistency, horizon reduction, representation learning, and reduced error compounding. We instead study what happens to an action error once it enters the system. At each state, we inject a small action error and measure how fast it grows or shrinks under two execution regimes:...
-
33
FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection
Junkang Liu
cs.AI
Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create a problem we term \emph{subspace fragmentation}: local projections interact with data heterogeneity to bias aggregated...
-
34
Agents Are Systems, Not Models: Rethinking Agentic Evaluation
Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata
cs.AI
Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific...
-
35
Evaluating Physical Consistency and Plausibility in Generative Scenario Models for Autonomous Driving
Manasa Mariam Mammen, Zafer Kayatas, Stefan Wagner
cs.AI
Generative AI models are increasingly used for scenario generation in autonomous driving. While they can generate realistic-looking scenarios, they often provide limited transparency into learned representations and consistency with real-world vehicle dynamics. This lack of formal assurance limits their use in safety-critical validation and certification workflows. To address this aspect, we introduce a layered evaluation protocol that...
-
36
The AI Assessment Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes
Alessio Buscemi, German Castignani, Daniele Pagani, Maxime Cordy, Jordi Cabot
cs.AI
The EU's Artificial Intelligence Act requires all Member States to establish AI Regulatory Sandboxes (AIRS) by August 2027: supervised environments bringing together national Competent Authorities, technical experts, and the organisations under assessment. When AIRS engagements include structured technical testing, running such testing at scale demands dedicated infrastructure, yet the tooling ecosystem remains structurally fragmented, with...
-
37
Neither Black nor White: Balancing Semantic and Collaborative Signals with Graph-Informed Semantic IDs (GrIS)
Aleksei Medvedev, Alejandro Ariza-Casabona, Steven Derby, Gonzalo Fiz Pontiveros, Xinyang Shao, Florian Spiess
cs.AI · cs.IR · cs.LG
Existing work on Semantic IDs (SIDs) for generative recommendation treats SID construction as a representation learning problem: encode items into a quantised latent space and read off codes. We argue this view is incidental. SID construction is, at heart, a recursive clustering problem, and once stated this way the natural object to cluster is a graph whose nodes carry semantic content and whose edges carry collaborative signal; SID...
-
38
Towards Reliable Vision-Language Models for Autonomous Driving
Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner
cs.AI · cs.CV
Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visual inputs may be degraded by sensor imperfections and environmental conditions, potentially affecting both model predictions...
This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.