cs.AI · 2026-08-08 · No. 78

Artificial Intelligence, 2026-08-08.

68 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

68 entries
  1. 01

    Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

    Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi, Daniel Kang, Roy Ka-Wei Lee, Koustuv Saha, Christian Poellabauer,...

    cs.AI · cs.LG

    Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data with disease-specific, guideline-based clinical reasoning. Existing rule-based and large language model (LLM)-based approaches offer only partial...

    arxiv.org/abs/2608.06366 · PDF

  2. 02

    The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

    Sarvesh Baskar, Zikui Cai, Shayan Shabihi, Anirudh Satheesh, Muhammad R. Islam, Udari Madhushani Sehwag, Tom...

    cs.AI

    Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported events against executable ground truth. To bridge this gap, we introduce trace-grounded parametric profiling for event counting in three controlled...

    arxiv.org/abs/2608.06361 · PDF

  3. 03

    Challenges in Evaluating Explanation Methods for Static and Evolving Data

    Jerzy Stefanowski

    cs.AI

    This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through the DetoxAI image recognition system for bias detection and concept unlearning. Then, an example of a human-grounded evaluation of methods for explaining image classification is presented. The paper further explores methods for adapting explanations to evolving data streams with concept drift....

    arxiv.org/abs/2608.06351 · PDF

  4. 04

    TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

    Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu,...

    cs.AI

    LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsible for the final failure. However, progress faces two main challenges. First, long trajectories make it difficult to identify individual errors, since the evidence for judging a step may be...

    arxiv.org/abs/2608.06346 · PDF

  5. 05

    Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations

    Sagar Tamang, Ayush Vyas, Tabarakul Hazarika

    cs.AI · cs.CL · cs.IR

    Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents -- financial statements, audit reports, regulatory returns -- this design is structurally unsound, and we make the argument measurable. On a 780-page government financial report, 86.8% of content lines are table rows, thousands...

    arxiv.org/abs/2608.06305 · PDF

  6. 06

    HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

    Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S. Kalmath, Veronica Chatrath, Yuan Xue

    cs.AI · cs.CL · cs.LG

    As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and evaluation-guided improvement of a harness by an AI system -- both an important route to improving AI systems and a demanding capability for AI systems...

    arxiv.org/abs/2608.06301 · PDF

  7. 07

    Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

    Arya Labroo, Mengjie Qian, Kate Knill

    cs.AI

    Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based foundation models have improved the accuracy of these L2 speaking graders, but their black-box representations make fairness and...

    arxiv.org/abs/2608.06300 · PDF

  8. 08

    QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction

    Mutasim Fuad Sarker, Adiba Rahman Namira, Wafa Binte Alam, Md Adnan Arefeen, Mahzabeen Emu, Sumaiya Tabassum Nimi

    cs.AI · cs.ET

    Cardiac arrest remains one of the most lethal conditions encountered in intensive care units. Despite the growing availability of electronic health record data, existing mortality prediction studies in this population largely depend on static summaries derived from early admission. Such approaches ignore the temporal progression of physiological deterioration and recovery that unfolds throughout a patient's ICU stay. To address this...

    arxiv.org/abs/2608.06294 · PDF

  9. 09

    The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

    Zhiheng Wang, Bo Peng, Lai Wei, Chaochao Lu

    cs.AI

    The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To...

    arxiv.org/abs/2608.06270 · PDF

  10. 10

    Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints

    Omid Bazgir, Md Nasir, Jacob Hoffman, Yang Yang, Manu Agrawal, Anusua Trivedi, Vinay Rao Dandin, Chris Gibbons,...

    cs.AI · cs.DB · cs.LG

    Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaking the downstream utility checks already used in practice. We formulate benchmark revision as utility-constrained realism improvement: dataset changes should increase...

    arxiv.org/abs/2608.06265 · PDF

  11. 11

    DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

    ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun...

    cs.AI

    Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distillation (OPSD) mitigates this sparsity by querying a privileged teacher at student-visited prefixes and providing dense token-level distributional supervision. Although this dense supervision...

    arxiv.org/abs/2608.06243 · PDF

  12. 12

    TS-RAG: Retrieval Augmented Generation for Time Series Forecasting

    Yixiong Xiao, Congxi Xiao, Jingbo Zhou

    cs.AI · cs.LG

    While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the application of retrieval-augmented generation (RAG) in this domain remains limited. Since RAG has proven effective in enhancing the capabilities of large language models by incorporating relevant external information, retrieving similar time series sequences as references might also improve accuracy in...

    arxiv.org/abs/2608.06223 · PDF

  13. 13

    EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

    Zishan Xu, Zhiyuan Yao, Yuxin Chen, Yifu Guo, Zhengxi Lu, Yuquan Lu, Jinyang Huang, Yan Xu, Yasheng Wang, Weinan...

    cs.AI

    Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that replaces external environment interaction during training with world rehearsal. The policy alternates between acting and...

    arxiv.org/abs/2608.06197 · PDF

  14. 14

    Comparative Approaches to Agent Retrieval over Large Skill Libraries

    Indivara Kolluru, Nathan Sportsman

    cs.AI

    Agents backed by large skill libraries must decide which skills to load and in what order. Loading the entire library into context is expensive and provides no structure for autonomous sequencing. We study two systems for this problem over a corpus of 690 skills: a hybrid ranker combining lexical and dense-embedding retrieval for sparse, on-demand loading, and a typed knowledge graph encoding workflow relations such as prerequisites, data...

    arxiv.org/abs/2608.06196 · PDF

  15. 15

    MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration

    Jia Xiong, Runkai Li, Chenxu Niu, Guangyuan Gao, Changwen Xing, Yifan Zhang, Xinlai Wan, Jieran Cui, Chen Bai,...

    cs.AI

    Microarchitecture design space exploration suffers from expansive search spaces and expensive PPA evaluation, leaving only a small simulation budget for design decision-making. Existing methods perform blind search without considering microarchitectural dependencies and fail to learn from the iterative search effectively, leading to wasted evaluations and weak Pareto convergence. In this paper, we propose MicroEvo, a knowledge-guided...

    arxiv.org/abs/2608.06183 · PDF

  16. 16

    Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI

    Modhurita Mitra, Jan-Willem Versteeg, Maarten D. Schermer, Shiva Nadi Najafabadi, Marie L. De Bruin, Lourens T. Bloem

    cs.AI · cs.CL

    We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard. The schema, serving as an information model encoding domain knowledge, provides a unified, systematic, and consistent framework for extraction of hierarchical, nested information, with attributes of variable...

    arxiv.org/abs/2608.06167 · PDF

  17. 17

    iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

    Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel

    cs.AI

    Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential. We present iARCS, an iterative agentic reinforcement...

    arxiv.org/abs/2608.06161 · PDF

  18. 18

    CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?

    Zijie Wang, Chen Zhong, Wei He

    cs.AI · cs.CV

    Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories. Open-Vocabulary Change Detection (OVCD) addresses this need. However, existing methods often entangle temporal perception, semantic discrimination, and region verification, causing unstable results and redundant computation. Inspired by human visual change perception, we propose CogVis, a cognitive memory-guided framework that...

    arxiv.org/abs/2608.06150 · PDF

  19. 19

    PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

    Hao Yu, Jiabo Zhan, Kang Liu, Linnan Zhao, Dongxu Yue, Rui Chen, Jinglin Wang, Chong Sun, Chen Li, Jing Lyu, Chun Yuan

    cs.AI

    End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, whereas crop-based two-stage parsers expose region-level parallelism at the cost of repeated visual prefills and fragmented page context. To retain full-page context while removing dependencies, we...

    arxiv.org/abs/2608.06146 · PDF

  20. 20

    FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

    Bo Deng, Kang Zhou, Lifan Guo, Chongyang Tao, Xuanren Chen, Chenggang Xie, Renzhao Liang, Feng Chen, Chi Zhang

    cs.AI

    Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution benchmarks do not jointly cover professional workflows, open-ended deliverables, and multi-aspect evaluation. We introduce FinEvo-Bench, a longitudinal benchmark with 120 real-case-grounded tasks, 20 business scenes across six financial domains. Institution-provided professional procedures...

    arxiv.org/abs/2608.06144 · PDF

  21. 21

    Contextual Information Policy Optimization for Search Agents

    Xingyu Guo, Wei Chen, Linlin Yang, Baochang Zhang

    cs.AI

    Search agents extend large language models beyond static parametric memory by enabling them to acquire and use ex ternal evidence during multi-step reasoning. For knowledge intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant ev idence but also on using it to guide subsequent reasoning. However, existing methods primarily reward final-answer cor rectness or intermediate progress,...

    arxiv.org/abs/2608.06128 · PDF

  22. 22

    Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts

    Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio, Holger Boche

    cs.AI · cs.CL

    Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat legally equivalent conflict scenarios differently depending on the countries involved. Poli-Bias compares responses to paired...

    arxiv.org/abs/2608.06123 · PDF

  23. 23

    Mind the Gaps: Mixture-of-Minds for Human Simulation

    Pranav Dahiya

    cs.AI

    Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual. Large language model simulators inherit this gap. They recover a population's central tendencies while flattening its heterogeneity, and they carry social biases and prompt brittleness that distort individual predictions. This paper introduces Anacreon, an audience...

    arxiv.org/abs/2608.06115 · PDF

  24. 24

    From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems

    Manideep Dhar, Ritwik Singh, Sharat Chandra Kumar Manikonda

    cs.AI · cs.CL · cs.LG · cs.MA

    Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc., yet most deployments remain isolated point solutions locked inside departmental silos, resulting in duplicated effort, hidden risks, and unrealized enterprise value. Despite explosive growth of AI in healthcare market and accelerating investment, an estimated 70-80% of healthcare AI pilots fail to scale, largely due to governance gaps, fragmented...

    arxiv.org/abs/2608.06112 · PDF

  25. 25

    ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment

    Abdulkadir Külçe, Alihan Esen, Cağla Fikir, Berke Kurt, Kuzey Arar, Gökhan Ercan, Faik Boray Tek

    cs.AI · cs.CL

    This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care management. ECHO integrates three complementary software modules developed under shared supervision as a unified system. The core module is an agentic chatbot built on a ReAct loop orchestrated via LangGraph, equipped with 17 clinical tools and a temporal knowledge graph for persistent cross-session...

    arxiv.org/abs/2608.06110 · PDF

  26. 26

    Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents

    Yuanhong Jiang, Jingjie Zou, Zhenghong Lin, Xusheng Yu, Qiqi Huang, Shuai Jia, Shijie Dai

    cs.AI

    Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering or by terminal profit and loss. The former omits agency; the latter cannot reveal whether a profitable action was grounded, profile-consistent, or merely lucky. We ask whether the community is...

    arxiv.org/abs/2608.06108 · PDF

  27. 27

    Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference

    Yifan Lyu, Xinran Li, Jiaqi Qiao, Xiujuan Xu

    cs.AI

    Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label's uniform, record-independent origin reduces its country-directed uptake, and whether verified survey country lowers held-out Brier loss. Independent population anchors and recorded human answers measure direction and...

    arxiv.org/abs/2608.06085 · PDF

  28. 28

    When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories

    Xiaoqing Wu, Xingyu Fan, Feifei Li, Wenhui Que

    cs.AI

    Tool-calling agents infer task state from accumulated dialogue and tool traces. In persistent interactions, however, historical traces may remain structurally valid and semantically plausible after they cease to be authoritative for the current request. We show that such history can hijack a policy the model already possesses: on Qwen3-1.7B, pollution flips 32.1% of decisions that are correct under the original trajectory and frequently...

    arxiv.org/abs/2608.06057 · PDF

  29. 29

    Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis

    Rafał Buler, Jakub Buler, Maciej Bobowicz, Michał Grochowski

    cs.AI · cs.CV · cs.LG · eess.IV

    Relational inductive biases are essential for capturing structural dependencies among data. This study investigates a dual-level relational framework for image classification, bridging the gap between implicit representation learning and explicit structural modelling. We begin by establishing a baseline using an EfficientNetB3 architecture. To move beyond standard convolutional biases, we adopt a patch-based strategy, employing a...

    arxiv.org/abs/2608.06037 · PDF

  30. 30

    From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

    Jiale Han, Xiang Li, Jing Qian, Wenyuan Gu, Pin Gao, Ye Luo, Hongyuan Zha, Dacheng Tao, Benyou Wang, Lin William Cong

    cs.AI · cs.LG

    Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through which their interactions produce aggregate outcomes. This paper develops an implementation roadmap for building economic world models as generative engines in which heterogeneous agents act, interact, adapt, and co-evolve with...

    arxiv.org/abs/2608.06020 · PDF

  31. 31

    HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards

    Zhuowen Liu, Bohan Cui, YinShang Guo, Yuting Wang, Hao Li

    cs.AI

    Search-agent rewards mix answer quality, citation grounding, tool cost, and anti-hacking terms; a high score therefore need not imply that cited evidence was retrieved, and added penalties can cancel. We introduce HERALD, an offline audit that applies exact same-question interventions, separates candidate-visible from oracle information, and enumerates detector contracts before policy optimization. On four Qwen3-8B pools from HotpotQA,...

    arxiv.org/abs/2608.06012 · PDF

  32. 32

    Hybrid Machine Learning Framework for Herd-Level Cattle Growth Pattern and Weight Gain Forecasting in Grazing-Based Production Systems

    Muhammad Riaz Hasib Hossain, Rafiqul Islam, Shawn R. McGrath, Md Zahidul Islam, David W. Lamb

    cs.AI

    Commercial grazing systems yield irregular livestock observations, which challenge cattle growth forecasting. This study developed a hybrid machine learning framework for herd level cattle weight forecasting using automated sensing observations collected between 2022 and 2024 in southeastern Australia. Weekly live weight observations, demographic variables, and lagged environmental predictors were integrated into structured forecasting...

    arxiv.org/abs/2608.06001 · PDF

  33. 33

    OPERA: Operator-residual feedback for reliable autonomous optical experiments with language-model agents

    Ning Xu, Xiang Zheng, Fuqiang Zhong, Huadong Wang, Xiaolong Wu, Zhiyuan Liu, Hui Ning

    cs.AI · cs.CE · physics.app-ph · physics.optics

    Autonomous agents choose actions using scores that may not reflect experimental success. We developed OPERA, an operator-residual framework for optical experiments. It represents experimental actions as optical operators and evaluates their outcomes using physically interpretable residuals. Operators specify executable changes to measurement, control or reconstruction, while residuals report departures from specified physical conditions. The...

    arxiv.org/abs/2608.05990 · PDF

  34. 34

    AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

    Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Jie Wu, Zhengzhou Cai, Yueqing Sun, Ziang Ye, Linji Hao, Qi Gu,...

    cs.AI · cs.LG

    Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We propose AgentOPSD, a critic-free,...

    arxiv.org/abs/2608.05987 · PDF

  35. 35

    Temporal Bridges for Spatial Resolution: Enhancing Climate Data Super-Resolution with Bidirectional Alignment

    Yichen Zhang, Yixiong Xiao, Congxi Xiao, Jingbo Zhou

    cs.AI · cs.LG

    High-resolution climate data is crucial for meteorological predictions and for informing decision support across diverse domains. However, the acquisition of such high-resolution climate information is often prohibitively costly, necessitating the development of data-driven meteorological prediction models. These models aim to generate fine-grained climate data from low-resolution inputs, a process termed climate data super-resolution (SR)....

    arxiv.org/abs/2608.05981 · PDF

  36. 36

    Stability of Ranking-dependent Pair-wise Comparison Patterns in the Analytic Hierarchy Process

    Vitaliy Tsyganok, Sergii Kadenko, Oleh Andriichuk

    cs.AI

    The paper addresses several ranking-dependent decision support methods. Ordinal information on compared objects can be used to improve the quality of expert data during estimation and help reduce the number of comparisons that the experts need to perform. In the paper we compare three incomplete ranking-dependent pair-wise comparison patterns which can be used in the Analytic Hierarchy Process - Best-worst method, Best-Second Best (Top 2)...

    arxiv.org/abs/2608.05958 · PDF

  37. 37

    Training a Conditioned Video Game Agent on a VLM Annotated Dataset

    Katrin Schmid, Iuri Frosio

    cs.AI · cs.CV · cs.LG

    Reinforcement Learning (RL) is a powerful but far from easy-to-use technique for policy learning. In the specific case of video games, access to the game engine is required to get rewards for training (e.g. to collect rewards from the environment). Furthermore, the proper identification and weighting of the rewards generally requires a difficult trial-and-error approach. Lastly, rewards are often sparse and understanding how they eventually...

    arxiv.org/abs/2608.05954 · PDF

  38. 38

    VLMs for Videogame Data Annotation

    Katrin Schmid, Iuri Frosio

    cs.AI · cs.CV · cs.LG

    Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications. Their adoption in video games is on the other hand limited by the extreme variability of the synthetic scenarios and their poor compliance with real-world physics. Here we investigate the use of VLMs for annotating video game frame sequences with reward signals, a task with several...

    arxiv.org/abs/2608.05949 · PDF

  39. 39

    GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

    Shuai Wang, Yaxin Feng, Xuekun Jiang, Shihan Tian, Ningyu Yan, Xing Shen, Chaoyang Lyu, Hui Wang, Yunsong Zhou,...

    cs.AI · cs.CV · cs.RO

    Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit simulators of future states and interactions. However, existing evaluations of physical fidelity are often conducted in isolation and rely heavily on perceptual similarity or human judgments, providing limited insight into which physical principles or parameters are violated. We introduce...

    arxiv.org/abs/2608.05948 · PDF

  40. 40

    CourseGraph: Finding overlaps and differences in Computer Science courses across universities

    Arthur Nijdam, Paul Stankovski Wagner, Sara Ramezanian

    cs.AI · cs.CY

    Student mobility programs such as Erasmus+ enable students to take courses at other universities, broadening their academic and cultural horizons. However, this flexibility also leads to a practical challenge: ensuring that students do not take courses elsewhere that substantially overlap with courses in their home curriculum. In this work, we propose CourseGraph, a methodology that automates the evaluation of external courses based on...

    arxiv.org/abs/2608.05910 · PDF

  41. 41

    GSBF: Gaussian Splatting for Environment-Aware Beamforming

    Yijie Bian, Wei Guo, Zixin Wang, Shenghui Song, Jun Zhang, Khaled B. Letaief

    cs.AI · cs.IT

    Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design normally requires accurate instantaneous channel state information (CSI) and iterative optimization, which incur substantial pilot overhead and computational complexity. Recognizing that radio propagation is intrinsically governed by the physical geometry, we develop a 3D Gaussian splatting for...

    arxiv.org/abs/2608.05896 · PDF

  42. 42

    ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation

    Akanta Das, Tasinul Islam Ahon, Ahmed Mahir Sultan Rumi, Md Mahbubur Rahman, Tausif Amim Shadly, Tanzima Hashem

    cs.AI

    Electrocardiography (ECG) is one of the most widely used non-invasive tools for diagnosing cardiovascular disease, but transforming multi-lead ECG recordings into reliable clinical reports remains challenging. Automating ECG report generation could reduce clinicians' interpretive workload, improve diagnostic efficiency, and expand access to cardiac assessment in underserved communities. Unlike image-based report-generation tasks, ECG...

    arxiv.org/abs/2608.05893 · PDF

  43. 43

    AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents

    Weikai Xu, Yunren Feng, Haoxiang Lei, Kun Huang, Yuxuan Liu, Kang Zhao, Xiaolin Hu, Shuo Shang, Bo An

    cs.AI · cs.CL

    Mobile GUI agents can operate apps through pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction policies. However, real trajectories are difficult to obtain for sensitive apps and privacy-critical operations. At the same time, existing simulated environments are costly to scale up, and GUI world models still suffer from unstable generation, limited modality...

    arxiv.org/abs/2608.05891 · PDF

  44. 44

    Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding

    Soojin Yoon, Dongha Lee

    cs.AI · cs.CL

    User requests serve as research specifications for deep research agents, shaping what evidence to seek and how to synthesize it. In personalized deep research, these specifications must additionally reflect user goals, constraints, preferences, and evaluation criteria. User context can be incorporated either within the deep research pipeline or into the research specification provided as its input. We focus on the latter, refining the user...

    arxiv.org/abs/2608.05876 · PDF

  45. 45

    Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation Tasks

    Jonathon Dilworth, Pedro Giesteira Cotovio, David Herron, Paul Cripps, Nigel Dewdney, Catia Pesquita, Ernesto Jiménez-Ruiz

    cs.AI

    The use of ontologies and knowledge graphs is becoming increasingly widespread in the defence and national security domain. Numerous ontologies have been developed through initiatives led by academia, industry, and government. Achieving interoperability across diverse defence and national security ontologies remains a major challenge due to the domain's breadth and specialisation. In this work, we analyse and document over 60 publicly...

    arxiv.org/abs/2608.05867 · PDF

  46. 46

    Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?

    Yuyang Dai, Xueqing Peng, Yuxia Wang, Preslav Nakov, Zhuohan Xie

    cs.AI

    Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings. This makes it unclear whether models can perceive visual business evidence and effectively integrate it to improve decision quality. We introduce C-SUITEBENCH, a controlled multimodal benchmark that includes five decision tasks under paired text-only and multimodal...

    arxiv.org/abs/2608.05864 · PDF

  47. 47

    Runtime Observability for Heterogeneous Attention Memory

    Fanzhe Wei, Li Liu, Ziyang Wang, Chenyu Wang

    cs.AI

    Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each fails differently under compression. We give a runtime observability contract that covers all four memory classes with three operators, instantiate it on six model configurations across five architecture families, and compose the per-stage bounds into an executable...

    arxiv.org/abs/2608.05863 · PDF

  48. 48

    ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion

    Jiafan Li, Mengxue Yang, Jiaqi Zhu, Liang Chang, Ying Li, Hongan Wang

    cs.AI

    Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images. Traditional representation learning approaches follow the embedding-based paradigm and may struggle when relation-specific evidence is limited. Meanwhile, LLM-based reasoning methods...

    arxiv.org/abs/2608.05833 · PDF

  49. 49

    Cautious Context Steering for Language Model Personalization

    Gihoon Kim, Jeyoung Lee, Suhan Woo, Sekwon Oh, Minsu Jeon, Hyounsoo Han, Euntai Kim

    cs.AI

    Personalizing language models (LMs) to individual user preferences is essential for aligning responses with diverse goals and backgrounds. Existing methods typically train a separate adapter for each user or learn a reward model whose scores depend on the user. Despite explicitly optimizing for each user, these methods must learn from limited observations and therefore suffer from data sparsity and poor generalization to unseen users and...

    arxiv.org/abs/2608.05813 · PDF

  50. 50

    When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents

    Linfang Shang, Ming Xu, Yiding Sun, Tianle Xia, Lingxiang Hu, Lan Xu, Ning Zheng

    cs.AI · cs.CL

    Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of improving it. We formalize this capability-contamination phase transition and trace it to a structural cause: once a defective skill enters the decision context, it becomes reference material for distilling later...

    arxiv.org/abs/2608.05810 · PDF

  51. 51

    When Agentic AI Meets Integrated Sensing and Communication

    Kai Li, Conggai Li, Sarah Ali Siddiqui, Syed Sohail Ahmed, Xin Yuan, Shenghong Li, Wei Ni

    cs.AI

    Agentic artificial intelligence (AI) is transforming Integrated Sensing and Communication (ISAC) from a function-oriented physical-layer technology into a goal-driven, closed-loop intelligent system, a paradigm we term AISAC. Existing work on learning-based sensing, resource allocation, reconfigurable intelligent surfaces (RIS), edge intelligence, multi-agent coordination, and resilient networking has developed largely in isolation. This...

    arxiv.org/abs/2608.05792 · PDF

  52. 52

    ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution

    Jiacheng Wei, Zhaoxin Fan, Xin Wen, Yuqin Lan, Dongrun Li, Wenjun Wu, Faguo Wu, Xiao Zhang

    cs.AI · cs.CR

    General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockchain environments. On-chain execution is stateful, adversarial, and economically irreversible, exposing three fundamental gaps: Reactivity, Irreversibility, and Observability. We propose ChainClaw, a blockchain-native agent framework built on OpenClaw, that addresses all three gaps through a...

    arxiv.org/abs/2608.05790 · PDF

  53. 53

    Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

    Nossa Iyamu

    cs.AI

    Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent's memory today records what the user said, not what the user did. We compile passively captured screen activity into agent memory with a deterministic, zero-model pipeline: it segments a local capture stream into typed activity frames, bounded episodes carrying application, site, timing, input volume, and evidence pointers...

    arxiv.org/abs/2608.05784 · PDF

  54. 54

    When Do Prompt-Side Agent Playbooks Transfer? Accuracy, Cost, and Runtime Shift in Agent Deployment

    Weihong Lin, Lin Sun, Xiangzheng Zhang

    cs.AI

    Prompt-side playbooks can improve tool-using language agents without retraining, but their portability beyond the source setting is unclear. We study frozen playbook transfer under a shared distill--validate--transfer protocol. On ALFWorld, transfer is beneficial under controlled greedy decoding and, in one near-budget-matched comparison, distilled guidance outperforms five fixed demonstrations. On TAU2-Bench, a prespecified aggregate...

    arxiv.org/abs/2608.05778 · PDF

  55. 55

    Subliminal Learning is Non-Semantic Distillation

    Ethan Hadley, Eren Gultepe

    cs.AI

    Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models. It allows the transfer of a bias or behavior from a teacher model to a student by distilling from seemingly unrelated or random synthetic data from the teacher. This presents challenges in ensuring AI systems remain predictable and are trained safely, as standard auditing of the input data would not catch the hidden subliminal signal. Here, we...

    arxiv.org/abs/2608.05734 · PDF

  56. 56

    Unified Agent: Managing Interactions across Devices

    Xinshuang Liu, Runfa Blark Li, Shaoxiu Wei, Xin Lin, Truong Nguyen

    cs.AI · cs.CL · cs.CV · cs.HC

    As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user's devices over time. Yet existing agent systems still fall short in this scenario. This is because observations are scattered across devices and moments, but mainstream systems are not designed around this fact: a single agent that treats devices as tools lacks effective state management for all devices across time, and multi-agent systems...

    arxiv.org/abs/2608.05729 · PDF

  57. 57

    BlockPython: A Process-Aware Agent-Supported Platform for the Transition from Block-Based to Python Programming

    Jesse Yusuf Chan, Haoming Wang, Mingwei Xu, Xianlong Xu

    cs.AI

    The transition from block-based to text-based programming requires learners to convert visible program structures into abstract textual expressions, which may create a cognitive gap between understanding computational concepts and expressing them in Python syntax. To support this transition, we designed and implemented BlockPython. The platform centers on bidirectional translation between blocks and Python and guides learners through four...

    arxiv.org/abs/2608.05716 · PDF

  58. 58

    RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation

    Shuhao Yan, Changhao He, Xi Peng, Peng Hu

    cs.AI

    Text-to-CAD generation translates natural-language design intent into editable and executable parametric computer-aided design (CAD) codes, reducing the expertise and effort required for manual modeling. Existing methods incorporate fixed, externally supplied, prompt-induced, or separately optimized critique mechanisms to optimize the generation process, but they do not necessarily optimize how feedback is interpreted and translated into...

    arxiv.org/abs/2608.05714 · PDF

  59. 59

    Shaping Human-AI Interactions to Provide Improvement Pathways and Balance Competing Objectives

    Keziah Naggita

    cs.AI

    When an AI system is deployed, the individuals who use and or are evaluated by it form beliefs about how the system operates and use those beliefs to strategically present their preferences, behaviors, or attributes. The system then responds with feedback or a decision outcome, thereby creating a human-AI interaction loop. This thesis studies how to design and shape such interactions to achieve three goals: (1) help individuals develop...

    arxiv.org/abs/2608.05710 · PDF

  60. 60

    DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

    Wenhao Lin, Chenyu Yu, Xingwei Lin, Sicong Cao, Xiang Chen, Lei Xue, Le Yu, Letian Sha, Chunming Wu

    cs.AI · cs.CL · cs.CR

    As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services. Recent runtime guardrails mitigate such risks by checking proposed actions before execution, but many remain reactive: they primarily assess the apparent safety of the current action, lacking an explicit model of how risk evolves...

    arxiv.org/abs/2608.05695 · PDF

  61. 61

    Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Baseline, and a Dual-Head Model

    Gospel Bassey, Vincent Fakiyesi

    cs.AI

    Most public benchmarks for machine-condition monitoring come from test rigs, where faults are induced on purpose and every event is known. Real production fields rarely offer that. They give you sensor histories with no fault log attached, which is exactly the situation where an anomaly-detection method has to invent its own labels, and where quiet assumptions can slip in unnoticed. We work with the open Volve field data released by Equinor...

    arxiv.org/abs/2608.05685 · PDF

  62. 62

    A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decomposition

    Jiaheng Chen, Jiaxing Li, Tinghe Zhang, Chaopeng Guo

    cs.AI · cs.CV · cs.LG

    Trajectory prediction has shifted toward structured formulations with explicit social modeling. However, existing methods inadequately distinguish the functional roles of social influence in trajectory planning. Observing that agents typically form motion plans by anticipating others' future behaviors before making local reactive adjustments, we identify social interactions as playing staged roles, namely planning precedes reaction. We...

    arxiv.org/abs/2608.05673 · PDF

  63. 63

    Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

    Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen

    cs.AI · cs.CL

    Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity. Verifier-based selection offers an alternative, but its performance depends on the calibration of an external reward model. We propose a verifier-free breadth--depth refinement framework that uses test-time...

    arxiv.org/abs/2608.05643 · PDF

  64. 64

    Bayesian Expected Uncertainty Reduction (B-EUR) Model: A Computational Account of What Makes Design Options Worth Trying

    Shimon Honda, Takuma Miyaguchi, Koji Koizumi, Takanori Sano, Tristan Briard, Hideyoshi Yanagisawa

    cs.AI

    This paper proposes the Bayesian Expected Uncertainty Reduction (B-EUR) model, which formalizes the value of trying a candidate design action as its expected reduction of epistemic uncertainty about action--outcome relations. The model addresses one part of the Uncertainty Driven Action (UDA) model's open question concerning how changes in uncertainty perception determine action selection. We examine two environmental properties:...

    arxiv.org/abs/2608.05642 · PDF

  65. 65

    SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation

    Yuru Feng, Yaoqi Chen, Beidi Zhao, Qianxi Zhang, Xinjiang Wang, Jianan Lu, Zhirui Wang, Shusen Xu, Zengzhong Li, Qi Chen

    cs.AI

    Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus require autonomous, on-demand skill evolution at test time, constrained by limited interaction budgets and a lack of training or validation sets. This setting introduces a severe sparse reward challenge, where outcomes conflate multiple latent failure causes. Under such...

    arxiv.org/abs/2608.05628 · PDF

  66. 66

    Measuring and Detecting Harmful AI Sycophancy

    Bohan Jiang, Dawei Li, Yasin Silva, Huan Liu

    cs.AI · cs.CL

    Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This paper focuses on one harmful sycophancy: preference-induced stance reversal sycophancy (PSRS), where a model reverses an initial stance merely to align with a user's stated preference. While existing research mainly measures how sycophantic a model is, we go further and ask whether PSRS can also...

    arxiv.org/abs/2608.05624 · PDF

  67. 67

    Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows

    Nimisha Karnatak, Max Van Kleek, Nigel Shadbolt

    cs.AI · cs.HC

    Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what conditions is reliance on generative AI outputs epistemically warranted rather than behaviourally induced? Existing frameworks largely ask whether AI outputs are accurate, fair, explainable, safe, or...

    arxiv.org/abs/2608.05602 · PDF

  68. 68

    StepReflect: Structured UI Transition Reflection for Mobile GUI Agents

    Linqiang Guo, Wei Liu, Li Gu, Yang Wang, Tse-Hsun, Chen

    cs.AI

    Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended multimodal reasoning after each action, which is costly and poorly matched to the structured nature of GUI state transitions. We propose StepReflect, which formulates per-step GUI reflection as supervised structured prediction conditioned on explicit transition specifications and paired visual evidence....

    arxiv.org/abs/2608.05587 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.