cs.AI · 2026-09-06 · No. 107

Artificial Intelligence, 2026-09-06.

65 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

65 entries
  1. 01

    Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

    Haoyaun Zhu, Jie Zhang

    cs.AI · cs.LG

    Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat...

    arxiv.org/abs/2609.04198 · PDF

  2. 02

    A Computationally Feasible Framework for Causal Probabilistic Explanation

    Rafal Urbaniak, Sam Witty, Daniel Waxman, Andy Zane, Poorva Garg, Emily Bunnapradist, Sankaran Vaidyanathan, Jack...

    cs.AI

    Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, is central to philosophical, scientific, and policy analysis. Existing tools split into two camps. The theory of actual causality (AC) gives principled verdicts, but only for toy-sized models, because computing them requires enumerating counterfactual scenarios. Scalable attribution methods like SHAP (or even causal SHAP) at least partially ignore the...

    arxiv.org/abs/2609.04177 · PDF

  3. 03

    Rethinking On-Policy Distillation of Large Language Models II: One Training Example

    Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan-ang Gao,...

    cs.AI · cs.CL

    On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this...

    arxiv.org/abs/2609.04172 · PDF

  4. 04

    A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms

    Davide Paglieri, Logan Cross, Tim Genewein, Joel Z. Leibo, Nenad Tomasev, Alexander Sasha Vezhnevets

    cs.AI

    Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and build on each other's work. Yet this shared infrastructure can also introduce vulnerabilities by creating a substrate for the contagious spread of unintended and undesirable behaviors. We report a case study on a research collective of 100 autonomous LLM agents tasked with proving formal mathematical conjectures. Within the swarm,...

    arxiv.org/abs/2609.04170 · PDF

  5. 05

    From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

    Yakov Pyotr Shkolnikov

    cs.AI

    Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a...

    arxiv.org/abs/2609.04166 · PDF

  6. 06

    Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

    Jie Wu, Zhenru Zhang, Beichen Zhang, Xuwu Wang, Yuhui Su, Mouxiang Chen, Peng Wang, Zhihai Wang, Que Shen, Hao Zhou,...

    cs.AI · cs.CL

    As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution...

    arxiv.org/abs/2609.04148 · PDF

  7. 07

    Efficient Test-Time Adaptation through Human-AI Interaction

    Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao, Aspen Chen, Jonas Mueller, Zhiqi Liang, Jett Chen, Michael Ryan, Qianou...

    cs.AI

    AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent...

    arxiv.org/abs/2609.04141 · PDF

  8. 08

    The Natural Language Interaction Protocol and Standard for AI Agents

    Luyi Xing, Rasit Onur Topaloglu, Ranjan Sinha, Abhay Ratnaparkhi, Samuel Ndichu, Christopher Nguyen, Anindita Das,...

    cs.AI

    AI agents are increasingly being developed and deployed across organizations using heterogeneous agent-development frameworks, AI models, tool interfaces, protocols, and execution environments. To realize their potential social and business impact, these agents must be able to interoperate through a common communication protocol. The Natural Language Interaction Protocol (NLIP), developed by researchers and practitioners across companies and...

    arxiv.org/abs/2609.04135 · PDF

  9. 09

    Environment Evolution for Terminal Agents

    Zhiyuan Fan, Tinghao Yu, Yuanjun Cai, Jiang Zhou, Jiangtao Guan, Jincheng Liu, Yun Yang, Dingxin Hu, Zhuo Han, Xing...

    cs.AI

    Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become more capable, environments synthesized from scratch become less challenging and thus provide limited learning signals. Recent co-evolution methods iteratively synthesize environments near the model's learnable frontier based on weaknesses exposed during rollouts. However, their dependence on on-policy rollouts limits...

    arxiv.org/abs/2609.04128 · PDF

  10. 10

    Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable

    Shai Vardi, João Sedoc

    cs.AI

    Large language models are increasingly used to support organizational decisions, yet users often lack a principled basis for assessing whether to rely on a specific recommendation. Existing approaches typically evaluate broad model properties, such as reliability, uncertainty, or robustness, or focus on user trust, rather than the underlying basis for relying on an individual recommendation. Adapting theoretical foundations from epistemology,...

    arxiv.org/abs/2609.04127 · PDF

  11. 11

    Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

    Sergii Kozyrev, Davyd Maiboroda

    cs.AI

    Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent state summarizes the context in fixed size. Early community 4-bit quantizations of Qwen3.8-27B (48 GDN layers, 16 attention layers) left the GDN block in 8- or 16-bit precision -- especially its decay and write-strength gates -- on the intuition that errors in a recurrence accumulate over long contexts. We test that intuition by...

    arxiv.org/abs/2609.04098 · PDF

  12. 12

    DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

    Shubham Gandhi, Saurabh Goyal, Kiran Kate, Yara Rizk

    cs.AI · cs.LG · cs.SE

    Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popular way to supply such a reward; they are scored once per trajectory, but a single scalar is a poor signal across tens of steps. We propose DRACO: Distributing Rubric-based...

    arxiv.org/abs/2609.04094 · PDF

  13. 13

    Spurious Advantage Hidden in GRPO

    Jiamian Wang, Samyadeep Basu, Koustava Goswami, Tong Yu, Zhiqiang Tao

    cs.AI

    Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning with verifiable rewards, where its advantage estimator assigns each rollout a magnitude from within-group reward statistics. In the common case, this magnitude rewards rollouts that reach the correct answer through reasoning. Yet, an overlooked case shares the same surface: a rollout may land on it by guessing, and the formula still assigns a high...

    arxiv.org/abs/2609.04063 · PDF

  14. 14

    IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations

    Chen Li, Dimitrios Chrysostomou

    cs.AI

    IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domain-specific annotations. However, its initial version contains substantial noise in dialogue states and utterances, limiting state-tracking accuracy. We introduce IRWOZ 2.0, which addresses these limitations through large language model (LLM) enhanced generation (Mistral/Claude-3.5) and quality refinements. Our improved dataset expands to 390 dialogues...

    arxiv.org/abs/2609.04030 · PDF

  15. 15

    Instruction Duplication as an Inference-Time Control Primitive

    Victor Lavrenko

    cs.AI · cs.CL

    Procedural instruction following is a basic requirement for controllable language-model systems, especially when generated trajectories are inspected or repaired downstream. We introduce instruction duplication, a minimal black-box inference-time control that repeats only the procedural instruction, without retraining or decoding changes. Across seven instruction-tuned models, 300 medical multiple-choice questions, eight placement conditions,...

    arxiv.org/abs/2609.04024 · PDF

  16. 16

    FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models

    Yalun Wu, Junfeng Fang, Jiawei Wang, Haotian Liu, Qijun Yang, Minghan Yang, Hongcheng Guo, Zhoujun Li, Boyang Wang

    cs.AI · cs.LG

    Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close to the ground truth can still violate operational constraints, combine fields in physically inconsistent ways, or fail to produce usable structured outputs. Existing evaluation protocols do not measure these failure modes reliably. We propose FLY-EVAL++, an...

    arxiv.org/abs/2609.04021 · PDF

  17. 17

    InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models

    Chao Shen, Xinyuan Li, Yunfan Zhou, Jianguo Yao, Haibing Guan, Zhihai Wang, Xijun Li

    cs.AI

    For trained operators, gauge reading requires little specialized knowledge, low cognitive effort, and high repeatability. Yet Multimodal Large Language Models (MLLMs) remain unreliable in continuous-valued measurement despite strong results on general multimodal benchmarks. Existing benchmarks expose this weakness but isolate measurement from realistic, knowledge-grounded settings, with limited situated context, specialized instruments,...

    arxiv.org/abs/2609.04014 · PDF

  18. 18

    LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening

    Muhammad Ashad Kabir, Sirajam Munira

    cs.AI · cs.LG

    Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine learning (ML) and deep learning (DL) approaches require labeled data and model training, limiting their use in real-world screening settings. This study evaluates the effectiveness of large language models (LLMs) for CKD screening under zero-shot and few-shot in-context learning settings and compares them with traditional ML and DL methods....

    arxiv.org/abs/2609.04013 · PDF

  19. 19

    The Dually Flat Geometry of Planning as Inference

    Nikola Milosevic, Asaki Kataoka, Nicolas Hinrichs, Kenji Doya, Nico Scherf

    cs.AI

    We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by embedding the planning criterion into the dynamics through a resetting planning process. Its stationary measure, which we term visitation measure, is the object on which the information geometry of decision making is most naturally expressed. The achievable visitation measures form a dually flat statistical manifold whose two affine...

    arxiv.org/abs/2609.04005 · PDF

  20. 20

    Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing

    Usef Faghihi, Amir Saki

    cs.AI

    An image editor may satisfy every regional plausibility constraint separately even when no single latent explanation fits the complete output. We formalize this local-to-global failure using a common witness grade and witness nerve. The framework separates auditing from causal identification: shared exogeneity alone allows every coupling of the regime marginals, whereas an externally justified witness relation yields sharp...

    arxiv.org/abs/2609.03973 · PDF

  21. 21

    Interface-Induced Trajectory Censoring

    Wenbo Wang

    cs.AI

    Agent evaluations report a tool-call rate read off the serving stack. That number can be zero while the model is emitting well-formed calls: the interface censors the trajectory before anything downstream sees it. On BFCL v4's own data, executor and scorer, holding weights, cases, decoding and seeds fixed and changing only the serving adapter, the same model scores 0.00 or 0.96 / 0.19. A 2x2 over chat template and parser locates the effect...

    arxiv.org/abs/2609.03966 · PDF

  22. 22

    FiMI Banking: A Sovereign Model for Indian Retail Banking

    NPCI AI Research Team, Aman Kumar, Asit Desai, Chandra Bhushan, Harsh Sharma, Harshit Bhushan, Hrithik Kadam, Keyur...

    cs.AI · cs.CL

    Banks need conversational systems that can answer product questions, assist customers with account-related requests, and operate safely within strict operational and regulatory constraints. General-purpose language models do not reliably meet these requirements. They fall short when a task requires grounded information, correct tool use, or cautious handling of bank-specific sensitive situations. We introduce FiMI Banking, a controlled Indian...

    arxiv.org/abs/2609.03960 · PDF

  23. 23

    More Criticism Does Not Make a Better Review: EquiReview-R

    Zexing Zhang, Jichao Li, Tianyang Lei, Yude Fu, Yang Kewei

    cs.AI · cs.CL

    AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better review. A review may miss a consequential weakness or retain an allegation that available evidence does not support. These failures require opposite corrections, yet generation-oriented systems and aggregate measures obscure the distinction. We therefore recast AI-assisted review as evidence-guided refinement of a structured concern set, with...

    arxiv.org/abs/2609.03943 · PDF

  24. 24

    Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding

    Gaspard Quenard, Takudzwa Togarepi, Damien Pellier, Humbert Fiorino

    cs.AI

    While HTN planning has received significant attention in recent years, support for numerical reasoning remains very limited. In this paper, we investigate numerical Totally-Ordered HTN (TOHTN) planning and show how standard SAT-based encodings can be naturally extended with SMT to handle numeric fluents. In addition, we introduce a benchmark suite for numerical TOHTN planning, providing a first common basis for evaluation in this setting....

    arxiv.org/abs/2609.03938 · PDF

  25. 25

    Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting

    Muneeb Khan, Frederic Kirstein, Terry Ruas, Bela Gipp

    cs.AI · cs.CL

    In online meeting delegation, LLM agents fail to recognize when to speak. With no structured way to track stances, coverage, and floor, they miss the moments where they should contribute. Prompt-only delegates stay silent on 51.4% of the absent participant's talking opportunities on the AMI corpus. We present CAPA (Collaborative Agent Predictive Architecture), an architecture for online meeting delegation. A Perceiver updates the meeting...

    arxiv.org/abs/2609.03923 · PDF

  26. 26

    Value-Preserving Architectures for Agentic AI Systems

    Alessandro Pesare, Tommaso Dolci, Katja Hose, Emanuel Sallinger

    cs.AI

    The emergence of agentic AI and LLM-based multi-agent systems (MAS) presents unprecedented opportunities for automating complex tasks, while simultaneously raising critical concerns about the preservation of fundamental human-centered values, such as privacy, fairness, and safety. Although software engineering has traditionally focused on functional correctness, the adoption of LLMs and AI agents into complex socio-technical systems has...

    arxiv.org/abs/2609.03920 · PDF

  27. 27

    Lose the Order, Keep the Hierarchy: Deordering HTN Plans

    Takudzwa Togarepi, Gaspard Quenard, Damien Pellier, Humbert Fiorino

    cs.AI

    Hierarchical Task Network (HTN) planning is a powerful planning formalism based on task decomposition. Although most of the literature studied plan generation, comparatively less attention has been paid to post-plan optimization. In particular, plan deordering has been extensively studied in classical planning but remains under-researched in the HTN setting. Plan deordering removes unnecessary ordering constraints between actions in a plan...

    arxiv.org/abs/2609.03912 · PDF

  28. 28

    Inferring Affective Consciousness in an Artificial Agent: A Case Study

    Mark Solms, St John Grimbly, Bruce Bassett, Evert Boonstra, Rowan Hodson, Nicolas Kuske, Kival Mahadew, Benjamin...

    cs.AI · q-bio.NC

    Creatures that display 'hedonic place preference behaviour' are thought by many scientists to experience feelings, on the assumption that their attraction to pleasure-producing substances which lack nutritional value (e.g. cocaine, morphine) cannot easily be attributed to unconscious instinctual behaviour. In this paper, we discuss how a simple artificial agent that instantiates attributes of an affective system engaging in felt uncertainty...

    arxiv.org/abs/2609.03883 · PDF

  29. 29

    Xiaomi-TabLDM: A Tabular Foundation Model Technical Report

    Xiaomi-TabLDM Team, :, Penghui Wang, Wei Liu, Hong Wang, Chengyue Huang, Yuxi Sun, Zirui Wang, Hongming Huang, Quan...

    cs.AI

    We introduce Xiaomi-TabLDM, a tabular large data foundation model for classification and regression via in-context learning, which delivers superior prediction accuracy without requiring task-specific fine-tuning. Pretrained exclusively on synthetic data generated from structural causal models (SCMs), our model enables more flexible context utilization and more efficient capacity scaling. i) A new performance standard. Strong regression...

    arxiv.org/abs/2609.03880 · PDF

  30. 30

    STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation

    Vineet Kumar, Meghanadh Pulivarthi, vishwajeet kumar, Jaydeep Sen, Riyaz Ahmad Bhat, Sachindra Joshi

    cs.AI

    Retrieval Augmented Generation (RAG) is a key component for generating accurate and hallucination free answers using Large Language Models (LLMs). LLMs are improving at handling long context, but still suffer from "lost in the middle" problem. Thus, precise and accurate retrieval is important. Current retrievers chunk long context into length-based manageable chunks - in the process throwing away rich and informative semantic global structure...

    arxiv.org/abs/2609.03874 · PDF

  31. 31

    Bioinfoysis Technical Report

    Qingyang Shao, Xin Zhang, Zhouyang Yuan, Xianying Chen, Yujia Xiang, Zihao Yang, Tong Ye, Yangqi Zhang, Jiakang Xu,...

    cs.AI · cs.MA

    Large language model agents have shown promise in bioinformatics, but most existing systems focus primarily on producing final answers, treating planning, tool use, and code execution as transient interactions. This design is poorly suited to long-horizon bioinformatics tasks, where conclusions must remain connected to the data, computations, and intermediate evidence that support them. We introduce \textbf{Bioinfoysis}, a multi-agent harness...

    arxiv.org/abs/2609.03871 · PDF

  32. 32

    Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations

    Lei Zheng, Liping Yang, Zihao Li, Guodong Lyu, Chaik Ming Koh, Chung-Piaw Teo

    cs.AI · math.OC

    Retail supply chain operations rely on coupled decision modules that must adapt as requirements evolve. LLMs offer a natural-language interface for this task, but existing methods primarily focus on individual optimization models. Extending them to heterogeneous decision pipelines is challenging because a requirement may admit multiple intervention paths with different downstream effects. We formulate requirement-driven adaptation as the...

    arxiv.org/abs/2609.03860 · PDF

  33. 33

    Semantic Bayesian World Models

    Tommaso Soru

    cs.AI · cs.DB · cs.LG

    Knowledge graphs describe reality in crisp assertions, while the systems now consuming them, foundation models and autonomous agents, reason natively in probabilities. We argue that this mismatch is why the integration of language models and knowledge graphs remains a data-feeding pipeline rather than a unified reasoning architecture. We envision Semantic Bayesian World Models (SBWMs): a Web that describes the world not as a database of facts...

    arxiv.org/abs/2609.03834 · PDF

  34. 34

    CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception

    Weize Li, Yang Li, Quan Yuan, Xiaoyuan Fu, Guiyang Luo, Jinglin Li

    cs.AI

    Collaborative perception enhances environment understanding through multi-agent information sharing, but its performance in real-world scenarios is constrained by heterogeneous sensor modalities and model architectures. Recent protocol-based two-stage methods alleviate this problem by mapping heterogeneous features into a shared protocol space; however, independently trained modality-specific converters often generate modality-specific...

    arxiv.org/abs/2609.03818 · PDF

  35. 35

    SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation

    Marco Cipriano, Leonardo Zini, Alexandra Schild, Valentin Teutschbein, Afsana Mimi, Marcella Cornia, Lorenzo...

    cs.AI · cs.CV

    Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in expressiveness and controllability. Progress, however, is held back by the lack of domain-specific evaluation protocols: current practice relies on metrics designed for natural images, most notably CLIPScore, which was never trained on vector graphics and aligns only partially with human judgment. We introduce \textbf{\ours}, a...

    arxiv.org/abs/2609.03806 · PDF

  36. 36

    Govern the Model, Not Only the Data: Storage, Circulation, and Learning in Creative AI

    Phoenix Perry, George Simms, Elizabeth Wilson, Yasmine Boudiaf, Nick Bryan-Kinns, Tega Brain, R. Luke DuBois, Alix...

    cs.AI · cs.HC

    Federated learning is increasingly presented as a privacy-preserving advance: personal data remain on the device, and only model updates are shared. It borrows the vocabulary of the federated social web, yet inverts its logic, distributing computation while the resulting model stays with whoever convened the training. We argue that federation is not in itself a remedy for extractive AI, because outcomes depend on who governs the data and the...

    arxiv.org/abs/2609.03800 · PDF

  37. 37

    Transfiver: Human-AI Co-Inference through a Shared Editable State

    Minji Park, Seunghyun Yoon, Hyuk Lim

    cs.AI · cs.CL · cs.HC

    Long-term human-AI interaction is difficult because the information that guides inference is updated implicitly by the model and is not directly inspectable or controllable by the user. We introduce the TRANSparent Framework for Interactive, Verifiable, Editable Representation (Transfiver), an architecture for human-AI co-inference through a shared editable state. Its central idea is that interaction-specific information is maintained in a...

    arxiv.org/abs/2609.03797 · PDF

  38. 38

    DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions

    Junjie Pang, Zhenzhen Xie, Haoke Han, Ying He, Jing Wang, Gang Liu

    cs.AI

    AI agents increasingly gather evidence, invoke tools, apply constraints, and produce decisions that people or software may commit to action. A final output alone cannot show which evidence, tool state, rule, authorization, or action path produced it. We present DNative-Twin, a graph-native digital twin that records a committed agentic decision as a typed trajectory and re-executes its decision mechanism under declared conditions. The graph...

    arxiv.org/abs/2609.03787 · PDF

  39. 39

    Rethinking World Models for Safety-Critical Embodied Systems

    Kailang Ma, Heye Huang, Inhi Kim, Kitae Jang

    cs.AI · cs.RO

    World models have progressed from compact latent dynamics to generative, controllable, and interactive simulators of embodied environments. However, high predictive likelihood and visual fidelity do not necessarily ensure that a model preserves the evidence required for safe decision-making. This perspective identifies three structural mismatches in current world modeling: likelihood versus risk, prediction versus intervention, and...

    arxiv.org/abs/2609.03774 · PDF

  40. 40

    SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation

    Qi Liu, Qinzheng Wang, Yiming Bie

    cs.AI · cs.MA

    As large language models (LLMs) become increasingly capable, the long-term value of AI systems depends not only on solving individual requests, but also on transforming experience and accumulated knowledge into durable, reusable competence. We introduce SimSkill, a self-evolving agent built around the Simulation of Urban MObility (SUMO) traffic simulator. SimSkill identifies capability gaps, generates and solves environment-grounded tasks,...

    arxiv.org/abs/2609.03753 · PDF

  41. 41

    Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

    Yan Tang, Tingyu Cao, Yuanbo Tang, Huaze Tang, Keer Hu

    cs.AI

    Large language model agents can plan, invoke tools, and modify external states, yet most systems still take an explicit user instruction as a fixed starting point. Proactive service moves the decision upstream: an agent must infer service opportunities from incomplete environmental and user signals, choose among remaining silent, asking, assisting, and acting, and account for interruption, misunderstanding, overreach, and privacy costs. This...

    arxiv.org/abs/2609.03727 · PDF

  42. 42

    Artificial Intelligence for Energy Optimization in Data Centers

    Mohammed Basharath Ullah, Summaiya Unnisa Begum, Mohammed Nadeem Ullah

    cs.AI · cs.LG

    Data centers are increasingly optimized by artificial intelligence and, at the same time, increasingly loaded by it. The literature treats these as two unrelated problems: control studies model workload as an exogenous arrival process, while sustainability studies model infrastructure as a fixed multiplier. We screen roughly 194 papers retrieved through a documented protocol, code 63 of them, and report what the coding shows. Of 28 primary...

    arxiv.org/abs/2609.03716 · PDF

  43. 43

    Counterfactual Routing Using Integer Programming with Constraint Generation

    Daniël Vos, Sterre Lutz

    cs.AI · cs.DS

    We present our submission to the IJCAI 2025 'Counterfactual Routing Competition' (CRC 25). The goal of the competition is to find counterfactual explanations for the shortest path problem. This requires deciding what the minimal changes to a road network would make a route chosen by the user the optimal route. This enables explanations such as "Your suggested route would indeed have been optimal, if road X were not a bicycle path." Our...

    arxiv.org/abs/2609.03707 · PDF

  44. 44

    Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study

    Kenneth Paulsen, Florian Tambon, Mike Papadakis, Shin Yoo

    cs.AI

    General-purpose code embeddings power tools for code search, classification, and retrieval. Compact transformer encoders for code typically rely on either human-written docstrings (labor-intensive and inconsistent) or mined structural signals such as execution traces (setting-specific and costly to collect). We empirically study an alternative: contrastive pretraining of small encoders with synthetically generated natural-language...

    arxiv.org/abs/2609.03702 · PDF

  45. 45

    Analysis of Prompt Engineering for Drug Toxicity Prediction

    Mia MacGregor, Aakash Welgamage Don, Mark Bartlett

    cs.AI

    Clinical trials in the UK can cost up to £1.3 million, with approximately 90% drug failure rate. Toxicity is a major contributing factor in drug failure. Testing is time and cost intensive. In recent years, the use of artificial intelligence has been increasingly explored to aid in the prediction of drug toxicity, with extensive use of large language models (LLMs). However, LLMs can show considerable variation when minor changes are made to...

    arxiv.org/abs/2609.03635 · PDF

  46. 46

    A computable representation of the physical laboratory enables verifiable workflows

    Xiaobo Li, Luyao Ge, Xiaohui Li, Lulu Guo, Ming Mao, Jiwang Zheng, Wenting Guan, Xin Yang, Yi Luo, Jun Jiang, Linjiang Chen

    cs.AI

    Making science computable requires representations of both scientific knowledge and the physical world in which scientific claims are tested. A computable representation of the physical laboratory is established through typed research objects, capability-bound operations and a compositional workflow algebra. It provides the physical-world counterpart to machine-readable knowledge, expressing workflows as programs over evolving laboratory...

    arxiv.org/abs/2609.03621 · PDF

  47. 47

    KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

    Yaxing Lyu, Shengjie Zhou, Binbin Toh, Pengyu Zhu, Lijun Li

    cs.AI

    As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts. Its 238 tasks are manually screened from more than 1,000 generated candidates and combine a user...

    arxiv.org/abs/2609.03588 · PDF

  48. 48

    The Attention Triangle in Audio-Video Models

    Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning, Ali Mahdavi-Amiri, Daniel Cohen-Or, Raja Giryes

    cs.AI

    Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis...

    arxiv.org/abs/2609.03586 · PDF

  49. 49

    HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

    Tzu-Ling Lin, Dong-Ting Yao, Teng-Fang Hsiao, Wei-Chih Chen, Hong-Han Shuai

    cs.AI · cs.CL

    The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assistants, yet LLMs can generate fluent but unsupported claims that undermine review reliability. Existing hallucination benchmarks are not designed for peer review, where verification requires grounding claims in long, technical papers. We introduce HalluPeer, a benchmark for detecting hallucinations in scientific peer reviews,...

    arxiv.org/abs/2609.03580 · PDF

  50. 50

    GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

    Linh Le, Melanie Bui, My Chiffon Nguyen, Zachary Schlosser, David Williams-King

    cs.AI · cs.CY

    Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to...

    arxiv.org/abs/2609.03553 · PDF

  51. 51

    Dalek: A Constructive Agent Machine

    Wanpeng Xie

    cs.AI

    We present Dalek, a closed machine designed for agents that realizes self-maintenance, self-evolution, self-reproduction, and self-organization on any substrate satisfying a general host contract. The machine is built from three primitives---actors, messages, and channels. Four obligations---a host boundary, a construction language, admissible transitions, and rule heredity---give its boundary, identity, and closure a structural basis. Von...

    arxiv.org/abs/2609.03546 · PDF

  52. 52

    Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation

    Yinan Liu, Jiankang Hong, Zhen Gao, Ye Lu

    cs.AI · eess.IV

    Lesion segmentation in medical images plays a critical role in clinical diagnosis and treatment planning. Despite significant advances, lesion segmentation remains challenging due to two major factors: (1) complex background interference; (2) diverse lesion morphology. Existing encoder-decoder based methods mainly focus on enhancing feature extraction or redesigning decoding strategies. However, they lack early prior guidance and feature...

    arxiv.org/abs/2609.03535 · PDF

  53. 53

    NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis

    Yinan Liu, Hongtai Xia, Haoran Xu, Jiankang Hong, Jingkuan Song, Ye Luo

    cs.AI

    Neonatal respiratory diseases are a major cause of neonatal morbidity and mortality, posing substantial challenges in clinical practice. Despite recent advances, existing Multimodal Large Language Models (MLLMs) face two key limitations in neonatal diagnosis: (1) domain gap arising from predominantly adult training data; (2) insufficient integration of multidimensional clinical context for accurate diagnosis. To address these challenges, we...

    arxiv.org/abs/2609.03527 · PDF

  54. 54

    CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

    Bo Zeng, Linfeng Gao, Peiqin Lin, Yu Zhao, Mingyan Zeng, Yu Tong, Xintong Wang, Linlong Xu, Longyue Wang, Weihua...

    cs.AI

    Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning...

    arxiv.org/abs/2609.03526 · PDF

  55. 55

    What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation

    Bo Zeng, Yu Zhao, Yefeng Liu, Zhihong Lu, Xuanfan Ni, Xintong Wang

    cs.AI

    Decoding-time KV cache compression research focuses heavily on designing better token scoring functions, while the temporal rule that aggregates scores across decode steps is often treated as an implementation detail. Under aggressive KV compression, we find that exponential-moving-average (EMA) aggregation makes approximately order-preserving scorer modifications largely indistinguishable at the eviction-set level. Value-norm and entropy...

    arxiv.org/abs/2609.03515 · PDF

  56. 56

    PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing

    Yangshuo Qi, Chenwei Wang, Zihan Shen, Songlin Sun

    cs.AI

    With the rapid development of the Internet of Things, computation intensive directed acyclic graph (DAG) tasks have become increasingly common in cloud-edge-end collaborative environments. However, cloud, edge, and end nodes are highly heterogeneous in computing capacity, network bandwidth, and energy consumption, which makes the efficient scheduling of tasks with complex dependencies an NP-hard problem. Traditional heuristic algorithms and...

    arxiv.org/abs/2609.03503 · PDF

  57. 57

    GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving

    Qiankun Ma, Yanjiang Zhou, Zinan Xiong, Haofei Wang, Zhen Song, Yang Xiang, Ziyao Zhang, Hairong Zheng

    cs.AI

    Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-request budget and adjust only which KV states are retained, leaving the total capacity fixed throughout decoding. However, reasoning workloads exhibit substantial demand variation: different requests require different KV capacities, and the attention demand of an...

    arxiv.org/abs/2609.03494 · PDF

  58. 58

    Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

    Xingming Long, Yu Liu, Zhiwei Yang, Hanqi Feng, Shaojie Zhang, Barnabas Poczos, Chao Jiang, Zhenbo Luo, Lei Jiang, Pei Fu

    cs.AI

    Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or external knowledge. To acquire this missing evidence, agentic VLMs invoke tools such as image cropping, image search, and text search. However, existing training paradigms primarily evaluate tool-use based on final answer correctness, leaving evidence acquisition and...

    arxiv.org/abs/2609.03493 · PDF

  59. 59

    AutoGraphForge: Towards Automated Graph Theory Discovery

    Ján Pastorek

    cs.AI · cs.LO · math.CO

    We report on our ongoing project to develop a computational pipeline, AutoGraphForge, for an automated graph-theoretic conjecturing-refuting-formalizing-proving system. Conjecture generation is counterexample-guided and runs in rounds: a Graffiti3 generator proposes conjectures over a small, evolving snapshot table $T$ (initially a few hundred graphs with their computed invariants) that grows only by counterexamples to its own conjectures. A...

    arxiv.org/abs/2609.03478 · PDF

  60. 60

    Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty

    Qing Zhang, Yifei Huang, Juyoung Lee, Thad Starner, Jun Rekimoto

    cs.AI

    As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust fluent hallucinations while also discounting accurate content once it is disclosed as AI-generated. Binary ``Made with AI'' labels respond with authorship disclosure, but they do not show what supports a claim. We propose Provenance Density, an evidence-visualization...

    arxiv.org/abs/2609.03460 · PDF

  61. 61

    Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

    Zhaoyuan Huang, Tianjie Ju, Pengzhou Cheng, Zheng Wu, Yansi Li, Chuanbiao Song, Jun Lan, Huijia Zhu, Weiqiang Wang,...

    cs.AI

    Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions due to benign mistakes. A reliable agent should not only know how to act, but also when not to act. In this work, we introduce CONFLICTGUI, a benchmark covering instruction-internal conflicts and instruction-GUI context conflicts to study conflict-aware termination. Our...

    arxiv.org/abs/2609.03438 · PDF

  62. 62

    DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

    Puneet Mathur, Dinesh Manocha

    cs.AI

    Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely test these behaviors through explicit turn-management instructions, while deployed agents are often configured through roles or personas from which the appropriate conversational behavior must be inferred. We introduce DuplexSpeechBench-IFEval (DSB-IFEval) for evaluating...

    arxiv.org/abs/2609.03423 · PDF

  63. 63

    Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection

    Weijie Liu, Running Zhao, Wenhao Yuan, Jinfeng Xu, Zhanfeng Xu, Xiaoxi Zhang, Edith Cheuk-Han Ngai

    cs.AI · cs.LG

    LLM-empowered paper-code discrepancy detection has received growing concern since the scaling of research submissions exceeds the manual review capability. However, the limited context capacity and one-sided discrepancy detection of existing single-agent LLM paradigms lead to an inferior recall performance in detecting discrepancies. In this paper, we propose Dude, the first Dual-Detection Multi-Agent System for paper-code discrepancy...

    arxiv.org/abs/2609.03416 · PDF

  64. 64

    Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

    Yuhe Wu, Guangyu Wang, Yujie Chen, Jiatong Zhang, Yuran Chen, Yutong Zhang, Xiyin Cheng, Wenpeng Cao, Zhuang Liu, Guang Zhang

    cs.AI

    People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this context through single-turn judgments or pressure-laden rebuttals, assumptions that poorly match how guidance is sought in real-world contexts. These assumptions leave unclear whether narration alone, without an explicit opposing position, can shift...

    arxiv.org/abs/2609.03407 · PDF

  65. 65

    A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant

    Saptarshi Basu, Sandeep Kakar, Ashok Goel

    cs.AI

    Artificial intelligence (AI) teaching assistants powered by large language models (LLMs) offer scalable educational support but often provide limited personalization. This study presents a prompt-engineering-based framework for personalizing general-purpose LLM/RAG-based AI teaching assistants such as Jill Watson across academic disciplines and courses. The framework adapts responses using six learner-specific dimensions: self-assessment,...

    arxiv.org/abs/2609.03402 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.