cs.AI · 2026-08-03 · No. 73

Artificial Intelligence, 2026-08-03.

29 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

29 entries
  1. 01

    ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

    Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo

    cs.AI

    Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together. The...

    arxiv.org/abs/2607.29677 · PDF

  2. 02

    Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics

    Yimin Chen, Brian Fricke, Bo Shen, Jamie Lian, Mingkan Zhang, James Lo, Yun Zhang, Shi Ye, Jiajing Huang, Han Hu,...

    cs.AI

    Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency, and maintenance effectiveness. However, effective deployment of FDD solutions in buildings requires structured domain knowledge that can bridge heterogeneous data sources, diverse equipment types, and varied diagnostic outputs. Limited data interpretability and interoperability within the FDD domain have led to fragmented...

    arxiv.org/abs/2607.29657 · PDF

  3. 03

    AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

    Tianyu Huai, Tingshuo Fan, Xinchi Chen, Yining Zheng, Yuxin Wang, Shuang Chen, Jie Zhou, Xuanjing Huang

    cs.AI

    As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper replication, or final answer correctness, but do not directly assess whether agents can interpret experimental evidence and use it to guide subsequent hyperparameter decisions. To address this gap, we introduce...

    arxiv.org/abs/2607.29626 · PDF

  4. 04

    DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

    Ismayil Ismayilov, Atakan Kara, Kaan Oktay

    cs.AI · cs.LG

    Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once. We introduce DungeonBench, a benchmark for tactical reasoning in Dungeons & Dragons combat, built to cover the vast majority of combat-relevant 2014 System Reference...

    arxiv.org/abs/2607.29577 · PDF

  5. 05

    LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

    Manith Adikari, Bei Peng, Samuele Vinanzi, Angelo Cangelosi

    cs.AI · cs.RO

    Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decision-making tasks often involve multiple, competing objectives, such as performance versus efficiency, where ground-truth reward functions are difficult to specify or inaccessible. While Multi-Objective RL (MORL) addresses such trade-offs by modeling rewards as vectors, existing approaches typically assume...

    arxiv.org/abs/2607.29559 · PDF

  6. 06

    COntExt: Towards Context-Aware Ontology Extension from Operational Metrics

    Hussain Hussain, Stefan Schöberl, Angelika Schneider, Verena Geist

    cs.AI

    Organizations increasingly define operational metrics in structured, machine-readable formats to monitor systems, processes, and compliance. These metric definitions implicitly encode domain knowledge, such as referencing concepts, properties, and relationships, that often extends what is captured in formal ontologies. Yet the connection between operational metric catalogues and ontological knowledge remains manual, ad-hoc, and...

    arxiv.org/abs/2607.29553 · PDF

  7. 07

    AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

    Rui Zou, Yutao Zhu, Mengqi Wei, Ji-Rong Wen

    cs.AI

    Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challenging. Existing representative methods mainly revise outputs through natural-language reflection or assist verification by directly generating verification programs; the former may not reliably support exact computation, whereas the latter prematurely couples mathematical modeling with...

    arxiv.org/abs/2607.29549 · PDF

  8. 08

    Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

    Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Haoyu Wu, Minghui Wu, Chenxu Zhao, Ante Wang, Guannan He, Changwei Wang

    cs.AI

    Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gradients yet do not explicitly shape future practice. External skill memories preserve procedural experience but are typically learned from fixed task distributions. We introduce \textbf{SESA} (Self-Evolving Skill-Augmented Agent), which makes procedural memory an evolving state of...

    arxiv.org/abs/2607.29468 · PDF

  9. 09

    Beyond Retrieval: Analytic Memory for Multimodal Agents

    Zhoujin Tian, Yao Tian, Hao Zhang, Cheng Chen, Yakun Li, Lei Zhang, Xiaofang Zhou

    cs.AI

    Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizing interaction histories through summaries and indexes to return query-relevant information at multiple granularities, from high-level abstractions to underlying records. In this paper, we formulate \emph{analytic memory} as a...

    arxiv.org/abs/2607.29440 · PDF

  10. 10

    ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

    Penglin Zhu, Jungang Xu

    cs.AI

    Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree. We present ModelEquivBench, a certifying, multi-relational evaluation system...

    arxiv.org/abs/2607.29431 · PDF

  11. 11

    Beyond Component Testing: Validating Agentic AI Systems

    Fabio Orazio Mirto, Luca D'Agati, Giuseppe Tricomi, Stefano Silvestri, Francesco Longo, Antonio Puliafito, Giovanni Merlino

    cs.AI · cs.MA · cs.SE

    Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches validation practice beyond component testing and one-shot input--output evaluation, because acceptable system behavior now depends on how decisions unfold over time and under changing environmental conditions. This survey synthesizes 257 papers spanning agent evaluation, software assurance,...

    arxiv.org/abs/2607.29405 · PDF

  12. 12

    MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation

    Hang Yan, Zhangxuan GU, Beitong Zhou, Jiaxuan Chen, Runze Li, Yusong Hu, Shuheng Shen, Changhua Meng

    cs.AI

    Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments. However, existing agents are typically domain-specific, limiting the deployment and user experience. This motivates the consolidation of specialized models into a single cross-environment policy. Weight merging directly merges domain-specific experts but can corrupt executable actions under expert...

    arxiv.org/abs/2607.29320 · PDF

  13. 13

    Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

    Minghui Pan, Jiayuxuan Yang, Yuanyuan Yuan, Yu Jiang, Zhenpeng Chen

    cs.AI

    AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions. Yet LLMs often become substantially less safe when deployed as agents, and the source of this degradation remains poorly understood. In this paper, we identify schema-formatted tool specifications as a primary source of agent safety degradation and show, through white-box...

    arxiv.org/abs/2607.29254 · PDF

  14. 14

    Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

    Ruiming Liang, Yi Zhong, Yizhen Yuan, Yinan Zheng, Tianyi Tan, Tianyue Wang, Haiyun Guo, Jinqiao Wang, Xianyuan Zhan

    cs.AI

    Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinforcement learning (RL) has become an increasingly important problem for LLMs, where each reward captures a different aspect of desired behavior. However, optimizing with multiple rewards suffers from a more severe alignment tax issue, where different optimization...

    arxiv.org/abs/2607.29246 · PDF

  15. 15

    MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft

    Jianxin Gao, Beini Hu, Runze Li, Wanli Peng, Ruohan Lei, Jinyuan Zhang, Linna Deng, Tianyi Yu, Zining Wang

    cs.AI

    With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unfortunately, most existing benchmarks evaluate them under fixed game mechanics. High performance in these settings does not show whether an agent can continue making progress when familiar recipes, drops, and other rules change. In this paper, we introduce MirrorCraft, a paired benchmark for evaluating...

    arxiv.org/abs/2607.29218 · PDF

  16. 16

    CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

    Blaise Delattre, Cong Wang, Yang Cao

    cs.AI

    Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates generally authorize the observed return and action, leaving the decision unprotected against small errors in how the return was bound to its source. We ask whether a candidate action stays authorized over a declared neighborhood of plausible correctly bound returns: one admissible binding fault...

    arxiv.org/abs/2607.29190 · PDF

  17. 17

    Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration

    Yanbin Fang, Xuan Wei, Wei Chen

    cs.AI

    Large language models (LLMs) are increasingly deployed in enterprise settings, yet individual models remain bounded by model-specific capability limitations. These heterogeneous boundaries pose a deployment challenge, but also create an opportunity: strategically coordinating multiple LLMs may unlock collective intelligence exceeding any single model. Existing approaches fix how models are combined in advance, overlooking the dynamic,...

    arxiv.org/abs/2607.29087 · PDF

  18. 18

    A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation

    Keita Kinjo

    cs.AI · stat.ML

    Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input required to obtain a desired output. Although CEs are conventionally formulated as a distance-minimization problem, the theoretical basis of this formulation has received limited attention. We show that a distance-minimization-based CE is mathematically equivalent to the maximum a posteriori (MAP) estimate...

    arxiv.org/abs/2607.29077 · PDF

  19. 19

    On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness

    Matthew Nguyen, Kyle Cox, Austin Meek, Iván Arcuschin

    cs.AI

    Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is possible to monitor it. However, in some cases, models do not verbalize important steps in their reasoning process. For example, models prompted with a cue suggesting the incorrect answer may fail to acknowledge that cue, even when it appears instrumental to their...

    arxiv.org/abs/2607.29062 · PDF

  20. 20

    Evidence-Grounded Constraint Checking in Construction Documents

    Rashid Mushkani, Hugo Berard, Shin Koseki

    cs.AI

    Professional-document review is a constraint-checking problem in which decisions depend on relations among text, geometry, pages, and document revisions. We present an evidence-grounded pipeline that normalizes extracted facts, executes four-state rules deterministically, retains source spans, and escalates unresolved cases. We evaluate its PDF evidence allocator on 160 reference-based tasks from 29 construction projects using a repeated...

    arxiv.org/abs/2607.29058 · PDF

  21. 21

    MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents

    Zeying Hao, Hao Guo, Mengtao Xu, Yimin Hu, Yuheng Song, Zesheng Zhou, Jinsong Lan, Xiaoyong Zhu

    cs.AI

    Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone. However, existing benchmarks largely rely on text-only or synthetic requests, underrepresenting complex real-world shopping requirements jointly expressed through images and language. We introduce MMShopBench, the first real-log benchmark for multimodal,...

    arxiv.org/abs/2607.29002 · PDF

  22. 22

    Scaling Scientific Discovery Environments for Turn-Level Agentic RL

    Yucheng Xu, Keyi Zhang, Yuyang Yu, Min Zhang, Shiyuan Meng, Pei Chu, Zhongying Tu

    cs.AI

    Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remains constrained by the lack of process supervised environments over real-world scientific data. This paper introduces SciDisco, a scalable framework for training Scientific Discovery agents in process-verifiable...

    arxiv.org/abs/2607.28990 · PDF

  23. 23

    MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

    Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue...

    cs.AI

    Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback...

    arxiv.org/abs/2607.28956 · PDF

  24. 24

    NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability

    Duo Xu, Faramarz Fekri

    cs.AI

    Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scientific discovery. In these settings, agents must act based on limited observations rather than full environmental states, leading to partial observability. This introduces several key challenges: belief state inference, task objective misalignment, and planning under...

    arxiv.org/abs/2607.28942 · PDF

  25. 25

    Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design

    Manisha Dubey, Rimvydas Rubavicius, N. Siddharth, Subramanian Ramamoorthy

    cs.AI · stat.ML

    Computational cognitive modeling seeks to infer latent cognitive mechanisms underlying observed behavior. Bayesian inverse planning provides a principled framework for such inference, but its success depends critically on the experimental environment. Existing approaches typically treat environments as fixed, leaving open the question of which cognitive experiments are most informative for cognition parameter inference. We formulate the...

    arxiv.org/abs/2607.28894 · PDF

  26. 26

    Fragility of Value under Imperfect Alignment

    Winter Cross

    cs.AI

    As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its...

    arxiv.org/abs/2607.28881 · PDF

  27. 27

    Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

    Pranav Narayanan Venkit, Akshara Prabhakar, Yu Li, Daniel Lee, Chien-Sheng Wu

    cs.AI · cs.CL

    As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that...

    arxiv.org/abs/2607.28818 · PDF

  28. 28

    Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

    Harsh Raj, Vipul Gupta, Anas Mahmoud, Razvan-Gabriel Dumitru, Darvin Yi, Aakash Sabharwal, Yunzhong He

    cs.AI

    Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training, harness engineering, environment redesign, or benchmark repair depending on its source. Because agent behavior emerges from interactions among models, harnesses, users, tools,...

    arxiv.org/abs/2607.28802 · PDF

  29. 29

    EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses

    Jiahui Li, Ruili Fang, Zishuai Liu, Yutong Guo, Nan Yang, Wenzhan Song, Jin Lu, Fei Dou

    cs.AI

    Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks are poorly suited to this setting: they restrict prediction to closed code sets, exclude free-text notes, and supervise with discharge diagnoses that incorporate the full inpatient course. We introduce EarlyDx, a large-scale benchmark for open-ended early diagnosis, built from 154,834 emergency department...

    arxiv.org/abs/2607.28788 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.