cs.AI · 2026-10-02 · No. 131

Artificial Intelligence, 2026-10-02.

38 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

38 entries
  1. 01

    ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

    Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha...

    cs.AI · cs.CL · cs.IR

    What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author...

    arxiv.org/abs/2610.02202 · PDF

  2. 02

    VISTA: A Visual Harness for Reasoning in an Interactive World

    Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He

    cs.AI · cs.CV

    We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations...

    arxiv.org/abs/2610.02200 · PDF

  3. 03

    A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

    Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia

    cs.AI · cs.CL

    A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic...

    arxiv.org/abs/2610.02116 · PDF

  4. 04

    Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

    Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir

    cs.AI · cs.CR

    Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applying FHE to reinforcement learning (RL) requires replacing non-linear operations with polynomial approximations, which diverge catastrophically due to...

    arxiv.org/abs/2610.02074 · PDF

  5. 05

    PyPottery: an AI-powered end-to-end suite for pottery processing and publication

    Lorenzo Cardarelli

    cs.AI

    The study of ceramic materials constitutes a cornerstone of archaeological research, yet the post-production workflow for pottery documentation remains labor-intensive and creates significant publication bottlenecks. This paper presents PyPottery, an open-source, AI-powered suite designed to semi-automate the complete ceramic documentation pipeline. The suite comprises four integrated modules: PyPotteryScan for automated image extraction and...

    arxiv.org/abs/2610.02072 · PDF

  6. 06

    Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

    Arman Behnam, Binghui Wang

    cs.AI

    Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories. When a memory is never retrieved, store-level interventions produce identical outcomes, leaving its utility unidentified. This is a retrieval-level positivity violation, invisible to diagnostics that examine only memory...

    arxiv.org/abs/2610.02070 · PDF

  7. 07

    External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing

    Kingshuk Gupta, Davide Buscaldi

    cs.AI

    As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary classification task, failing to capture the structured, sequential boundaries of semantic drift. Here, we introduce an internal...

    arxiv.org/abs/2610.02066 · PDF

  8. 08

    HydroJEV: A one-second, training-free screen for cyber-attack and fault attribution in water distribution networks

    Tianwei Mu, Shengyan Jiang, Mingzhe Yuan, Qing Luo, Min Xiao, Wenhong Wang, Jun Li, Manhong Huang

    cs.AI

    When a SCADA alarm is raised in a water distribution network, operators must decide quickly whether it reflects a cyberattack, a physical fault, a normal transient or a faulty sensor. Supervised classifiers need labelled incidents that utilities rarely have, and frontier large language models (LLMs) take tens of seconds per decision. We tested whether Jev, a training-free model that returns class probabilities in about one second, can serve...

    arxiv.org/abs/2610.02048 · PDF

  9. 09

    Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control

    Yimeng Liu, Mi Zhang, Younsuk Dong, Zhichao Cao

    cs.AI

    Large language model (LLM) agents increasingly combine reasoning, tool use, and action, but most evidence comes from episodic tasks with relatively immediate feedback and reset failures. Long-running physical control operates in a different regime: actions alter future states, errors compound across decisions, and an agent must improve from experience without being allowed to rewrite the physical rules that make execution safe. We study this...

    arxiv.org/abs/2610.02038 · PDF

  10. 10

    Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration

    Xin Heng

    cs.AI

    AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence. Our Observation-Aliasing Impossibility Theorem gives the exact boundary. A policy can guarantee a valid action exactly when all worlds producing the same observation share an admissible action. If k indistinguishable worlds require pairwise-disjoint...

    arxiv.org/abs/2610.02036 · PDF

  11. 11

    SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL

    Hyeonmin Lee, Zheng Wei, Kyungmin Kwon, Jumin Seo, Jiwon Park, Hayoung Oh

    cs.AI · cs.HC

    While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated synthesis into continuous human-AI co-creation. SPHERE extracts persistent spatial preferences from natural multimodal interactions (speech and...

    arxiv.org/abs/2610.02023 · PDF

  12. 12

    Atoms to Processes: The Role of Artificial Intelligence and Machine Learning in Chemical Engineering

    Michael Baldea, Linda J. Broadbelt, Marianthi G. Ierapetritou, Akhilesh Jain, Ankur Kumar, Thomas A. Kwan, Fèlix...

    cs.AI

    The rapid maturation of artificial intelligence (AI) and machine learning (ML) has catalyzed a profound shift in how chemical engineering problems are formulated, analyzed, and solved. Advances in computing, data availability, and learning algorithms have enabled AI/ML methods to impact applications spanning atomic-scale simulations, materials and catalyst discovery, transport and thermodynamics, separations, process systems engineering, and...

    arxiv.org/abs/2610.02014 · PDF

  13. 13

    Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

    Ionel Eduard Stan, Paolo Napoletano

    cs.AI

    A multi-LLM \emph{council} lets several large language models (LLMs) deliberate on a question and return an answer together with a confidence estimate. As these systems become increasingly used for reasoning, that confidence should represent a calibrated \emph{probability of being correct}, and the decision should remain robust when some agents are persistently unreliable. Existing \emph{council aggregation} methods fail on both fronts: their...

    arxiv.org/abs/2610.02005 · PDF

  14. 14

    Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

    Hao Wang, Ting Huang

    cs.AI · cs.CL

    Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned. We present evidence, from a controlled single-machine comparison and one third-party benchmark, that a substantial share of these failures is attributable to the harness rather than the model. We...

    arxiv.org/abs/2610.02001 · PDF

  15. 15

    Can AI Oversight Be Zero Knowledge?

    Alessandro Chiesa, Ziyi Guan, Burcu Yildiz

    cs.AI · cs.CC · cs.CR

    AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on...

    arxiv.org/abs/2610.01995 · PDF

  16. 16

    Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

    Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang

    cs.AI · cs.CL · cs.LG

    Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted...

    arxiv.org/abs/2610.01963 · PDF

  17. 17

    Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning

    Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR

    cs.AI

    Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG...

    arxiv.org/abs/2610.01936 · PDF

  18. 18

    AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes

    Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley

    cs.AI · cs.SD · eess.AS

    Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset. To construct the dataset,...

    arxiv.org/abs/2610.01861 · PDF

  19. 19

    Temporal-Difference Learning for Dragonchess

    Jim O'Connor, Annika Hoag, Sarah Goyette, Gary B. Parker

    cs.AI

    Our research investigates how two adaptive AI methods, evolutionary transfer learning and TD(lambda), perform in the three-dimensional chess environment Dragonchess. The game challenges players with its unique board structure and computational load, making it an ideal setting to study how adaptive methods can update evaluation heuristics in novel environments. In this work we re-implement the Dragonchess engine, changing it from a PyGame...

    arxiv.org/abs/2610.01845 · PDF

  20. 20

    On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models

    Haochen Zhang, Jiaheng Guo, Zhen Xu, Zachary Plotkin, Nicholas Konz, Zhen Tan, Tianlong Chen

    cs.AI

    A time series world model (TSWM) predicts a controlled system's state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but their responses to changed plans remain untested. We ask which design choices matter and whether accurate forecasters respond to...

    arxiv.org/abs/2610.01842 · PDF

  21. 21

    Code Owns the Simulation, Jev Owns the Evaluation

    Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang

    cs.AI · cs.LG

    Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the...

    arxiv.org/abs/2610.01834 · PDF

  22. 22

    Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

    Ngoc Phuoc An Vo, Aarya Doshi, Vadim Sheinin

    cs.AI · cs.SE

    Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determination skill in an enterprise Value Aware Resiliency system. The framework independently computes per-run ground truth,...

    arxiv.org/abs/2610.01833 · PDF

  23. 23

    AI-assisted mitotic counting improves reproducibility and efficiency across multiple tumour types

    Simon Graham, Mostafa Jahanifar, Quoc Dang Vu, Vygante Maskoliunaite, Donatas Petroska, Ruta Barbora Valkiuniene,...

    cs.AI

    Mitotic counting is an important component of tumour grading, diagnosis and prognostic assessment across several tumour types, but manual assessment is time-consuming and subject to inter-pathologist variability. To help address these challenges, we developed MitPro, an AI tool designed to improve consistency and efficiency by directing pathologists towards regions with the highest predicted mitotic activity and highlighting mitotic figures...

    arxiv.org/abs/2610.01813 · PDF

  24. 24

    LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification

    Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen

    cs.AI

    Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were designed for other modalities and other tasks, and they transfer poorly to open-ended generation in the time series domain. We address this by...

    arxiv.org/abs/2610.01800 · PDF

  25. 25

    Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents

    Beining Wu, Zihao Ding, Jun Huang

    cs.AI · cs.GR

    Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of experience: a trajectory bundles items with different properties, so a conclusion about the bundle depends on its mix. To address this, (i) we introduce component routing, which splits the experience into locators,...

    arxiv.org/abs/2610.01787 · PDF

  26. 26

    Q-Learning for Reachability in MEC-Free MDPs

    Lu-Chin Chang, Suguman Bansal

    cs.AI · cs.LG · cs.LO

    Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of...

    arxiv.org/abs/2610.01781 · PDF

  27. 27

    RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

    Arman Behnam, Sunglyoung Kim, Liangwei Yang

    cs.AI

    A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person's record, and such records are private, so benchmarks generate the person and the questions and settle in advance what matters. We release \bench, ten real relationships with an AI companion: 27,218 messages over up...

    arxiv.org/abs/2610.01780 · PDF

  28. 28

    VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding

    Bingjun Luo, Yuhuan Fan, Jialin Guo, Siqi Li

    cs.AI · cs.CV

    Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions are refined. Manually refining these harnesses requires diagnosing grounding failures and coordinating changes to both agent workflows and instructions. We introduce VideoEvolve, a framework that automatically...

    arxiv.org/abs/2610.01766 · PDF

  29. 29

    TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference

    Mukund Agarwalla, Chih-Jen Lin

    cs.AI

    Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same...

    arxiv.org/abs/2610.01763 · PDF

  30. 30

    vFedProtoQNAS: Prototype-Guided Personalized Quantum Neural Architecture Search for Virtual Federated Learning

    Seok Bin Son, Samuel Yen-Chi Chen, Soohyun Park, Joongheon Kim

    cs.AI

    Quantum federated learning (QFL) has emerged as a promising approach for collaboratively training compact quantum neural networks (QNNs) over distributed private data on resource-constrained devices. However, differences in device capabilities make a single shared QNN architecture unsuitable for all clients. While personalized quantum neural architecture search (QNAS) allows each client to select a device-specific QNN, averaging parameters...

    arxiv.org/abs/2610.01718 · PDF

  31. 31

    CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement

    Dongwei Sun, Yujie Zhang, Bowen Yao, Pei Liu, Jing Yao, Xiangyong Cao

    cs.AI · cs.CV

    Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect...

    arxiv.org/abs/2610.01710 · PDF

  32. 32

    Measuring the Stability Assumption Behind Action Chunking

    Aryan Goyal

    cs.AI

    Action chunking improves the performance of policies learned by behavioural cloning, and several mechanisms have been proposed to explain why, including temporal consistency, horizon reduction, representation learning, and reduced error compounding. We instead study what happens to an action error once it enters the system. At each state, we inject a small action error and measure how fast it grows or shrinks under two execution regimes:...

    arxiv.org/abs/2610.01626 · PDF

  33. 33

    FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection

    Junkang Liu

    cs.AI

    Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create a problem we term \emph{subspace fragmentation}: local projections interact with data heterogeneity to bias aggregated...

    arxiv.org/abs/2610.01620 · PDF

  34. 34

    Agents Are Systems, Not Models: Rethinking Agentic Evaluation

    Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata

    cs.AI

    Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific...

    arxiv.org/abs/2610.01618 · PDF

  35. 35

    Evaluating Physical Consistency and Plausibility in Generative Scenario Models for Autonomous Driving

    Manasa Mariam Mammen, Zafer Kayatas, Stefan Wagner

    cs.AI

    Generative AI models are increasingly used for scenario generation in autonomous driving. While they can generate realistic-looking scenarios, they often provide limited transparency into learned representations and consistency with real-world vehicle dynamics. This lack of formal assurance limits their use in safety-critical validation and certification workflows. To address this aspect, we introduce a layered evaluation protocol that...

    arxiv.org/abs/2610.01581 · PDF

  36. 36

    The AI Assessment Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes

    Alessio Buscemi, German Castignani, Daniele Pagani, Maxime Cordy, Jordi Cabot

    cs.AI

    The EU's Artificial Intelligence Act requires all Member States to establish AI Regulatory Sandboxes (AIRS) by August 2027: supervised environments bringing together national Competent Authorities, technical experts, and the organisations under assessment. When AIRS engagements include structured technical testing, running such testing at scale demands dedicated infrastructure, yet the tooling ecosystem remains structurally fragmented, with...

    arxiv.org/abs/2610.01539 · PDF

  37. 37

    Neither Black nor White: Balancing Semantic and Collaborative Signals with Graph-Informed Semantic IDs (GrIS)

    Aleksei Medvedev, Alejandro Ariza-Casabona, Steven Derby, Gonzalo Fiz Pontiveros, Xinyang Shao, Florian Spiess

    cs.AI · cs.IR · cs.LG

    Existing work on Semantic IDs (SIDs) for generative recommendation treats SID construction as a representation learning problem: encode items into a quantised latent space and read off codes. We argue this view is incidental. SID construction is, at heart, a recursive clustering problem, and once stated this way the natural object to cluster is a graph whose nodes carry semantic content and whose edges carry collaborative signal; SID...

    arxiv.org/abs/2610.01533 · PDF

  38. 38

    Towards Reliable Vision-Language Models for Autonomous Driving

    Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner

    cs.AI · cs.CV

    Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visual inputs may be degraded by sensor imperfections and environmental conditions, potentially affecting both model predictions...

    arxiv.org/abs/2610.01531 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.