cs.AI · 2026-08-04 · No. 74

Artificial Intelligence, 2026-08-04.

58 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

58 entries
  1. 01

    AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies

    Qiushi Lin, Chaojie Zhang, Íñigo Goiri, Aditya Akella, Ricardo Bianchini, Jovan Stojkovic

    cs.AI · cs.DC · cs.OS

    The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent, and prototyping a single policy takes months. Agentic AI promises to automate this search. Off the shelf, however, it falls short on three fronts. It is not formal: with no structured, searchable statement of the problem, the search has little...

    arxiv.org/abs/2608.02569 · PDF

  2. 02

    A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI

    Taye Akinrele, Sindhuja Penchala, Noorbakhsh Amiri Golilarz, Sudip Mittal, Shahram Rahimi

    cs.AI

    Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustained reasoning, adaptive behavior, persistent memory, and self-regulation. While generative and agentic AI have demonstrated impressive capabilities across a wide range of tasks, many fundamental cognitive functions remain fragmented or weakly developed, limiting reliable operation over extended time horizons. This paper presents...

    arxiv.org/abs/2608.02553 · PDF

  3. 03

    Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation

    Natalie Isak, Matthew Dressman

    cs.AI · cs.CY

    The most capable AI deployments are not single models but ensembles of specialized agents that delegate and act in coordination. This architecture unlocks powerful new capabilities, and it also introduces risks that existing frameworks for monitoring, detection, and mitigation were not designed to address. Most state-of-the-art AI abuse detection literature focuses on single-turn or multi-turn (single-session) threat models. This leaves a...

    arxiv.org/abs/2608.02518 · PDF

  4. 04

    Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies

    Sterre Lutz, Daniël Vos, Matthijs T. J. Spaan, Anna Lukina

    cs.AI · cs.LG

    Sequential decision-making in real-world applications often involves uncertainty about the environment's model. Uncertain Markov decision processes (UMDPs) represent the possible environments as a set of MDPs with shared states and actions but potentially different transition probabilities and rewards. Optimizing a single policy across all possible MDPs may sacrifice performance, while preparing an individually optimized policy for every MDP...

    arxiv.org/abs/2608.02509 · PDF

  5. 05

    Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation

    Michael Farmer

    cs.AI · cs.CV · cs.IR

    Can scientific abduction occur without continuous sensorimotor embodiment? Recent arguments in AI and philosophy of science hold that genuine hypothesis generation requires an agent continuously coupled to the physical world. We defend a narrower claim: online embodiment is not necessary for every abductive scientific act. Our focus is identity abduction: the inference that two independently developed structures are one object under an...

    arxiv.org/abs/2608.02505 · PDF

  6. 06

    CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization

    Chuyan Chen, Peng Sun, Kun Yuan

    cs.AI

    Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence. In this paper, we identify the root cause of this bottleneck: standard DiT architectures fuse...

    arxiv.org/abs/2608.02502 · PDF

  7. 07

    Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

    Nicole Mitchell, Dhruv Agarwal, Maty Bohacek, Remi Denton, Roma Patel

    cs.AI

    Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features can introduce longitudinal risks---cognitive, developmental and socio-affective changes in humans---that might not surface in short-term interactions, but can have lasting long-term effects on users. This forms the basis of a critical new mission for NLP: to pivot...

    arxiv.org/abs/2608.02491 · PDF

  8. 08

    Real-Time Detection and Repair of LLM Agent Failures

    Sunny Dubey

    cs.AI · cs.LG · cs.SE

    LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself. We ask how much detection is achievable from observable step telemetry alone, using monitors costing microseconds per step and trained only on healthy runs. On 2,823 committed agent episodes across three...

    arxiv.org/abs/2608.02464 · PDF

  9. 09

    Infinite Trace Objectives with Finite Trace Techniques: Translating LTL to LTLf+

    Christoph Weinhuber, Maximilian Prokop, Giuseppe De Giacomo, Moshe Y. Vardi

    cs.AI · cs.FL · cs.LO

    Linear Temporal Logic (LTL) is one of the most widely adopted languages for specifying temporal extended objectives in AI, with applications ranging from reactive synthesis to stochastic planning in Markov decision processes and reinforcement learning. Traditionally, solving any of these problems requires translating the LTL specification to a nondeterministic automata on infinite words and then determinizing it, a step that is notoriously...

    arxiv.org/abs/2608.02454 · PDF

  10. 10

    ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision

    Wei-Jung Huang, Bonan Shen

    cs.AI

    LLM-agent evaluations often produce task outcomes long before the full benchmark run is complete. A partial score is tempting to report, but it does not show whether the observed tasks support the same conclusion as the completed evaluation. Early tasks can omit important parts of a benchmark, running cheaper tasks first can distort the observed sample, and a rule that decides only easy pairs can appear accurate while leaving many comparisons...

    arxiv.org/abs/2608.02444 · PDF

  11. 11

    Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

    Xuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Yihua Zhu, Hu Wei, Bing Zhao

    cs.AI · cs.CL

    Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasoning capability targeted by the problem. We identify Solution Hacking, a failure mode in which an LLM reaches the correct answer through invalid shortcuts, such as numerical search, enumeration, guessing, or answer-first verification, without providing a valid...

    arxiv.org/abs/2608.02442 · PDF

  12. 12

    Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce

    Shicheng Fan, Mingdai Yang, Duohao Wang, Canyu Chen, Yongfeng Zhang, Hua Wei, Manling Li, Julian McAuley, Kun Zhang,...

    cs.AI

    In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to agents. Commerce, however, requires independently controlled Buyer and Merchant agents to interact in a shared market while preserving their private objectives and distinct authority. We introduce Agentic...

    arxiv.org/abs/2608.02441 · PDF

  13. 13

    xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding

    Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim, Mudhakar Srivatsa, Raghu K. Ganti, Minjia Zhang, Naigang Wang

    cs.AI

    Block-diffusion drafters like dFlash generate an entire block of draft tokens in a single forward pass, drastically reducing the overhead of multiple-token drafting in speculative decoding. The crucial final step of the single-pass discrete denoising process involves using the logit distribution at each position to sample conditionally independent tokens. The resulting draft is thus a set of per-position marginals, rather than a joint...

    arxiv.org/abs/2608.02438 · PDF

  14. 14

    MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models

    Victor Ojewale, Ro Encarnación, Suresh Venkatasubramanian, Danaé Metaxa

    cs.AI · cs.CY

    Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms record satisfaction without task purpose. Together, they leave a critical gap in LLM evaluation: no existing infrastructure routinely links interaction trajectories to user-defined outcomes. We introduce MonitrLLM, open-source infrastructure for community-centered...

    arxiv.org/abs/2608.02409 · PDF

  15. 15

    Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training

    Zhiyuan Wang, Shengcai Liu, Jiahao Wu, Ning Lu, Hui Ouyang, Shaofeng Zhang, Haoze Lv, Ke Tang

    cs.AI · cs.LG

    Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution strategies (ES) enable memory-efficient full-parameter post-training without backpropagation and can eventually match the performance of gradient-based reinforcement learning (RL). However, resource-constrained settings typically offer only a few GPUs, so the high GPU-hour requirements of ES...

    arxiv.org/abs/2608.02391 · PDF

  16. 16

    Chess on Ice: Curling Tactical Decision-Making via Backward Induction and Deep Reinforcement Learning

    Patrick Oberlin, Matteo Cederle, Aren Karapetyan, Saverio Bolognani, Gian Antonio Susto, Florian Dörfler

    cs.AI

    Curling is often referred to as "Chess on Ice", owing to the tactical complexity of its decision-making process. Yet unlike chess, curling remains largely underexplored from a machine learning perspective, with prior work confined mainly to statistical approaches. We propose a reinforcement learning framework capable of quantitatively evaluating and comparing tactical options in curling. The game poses several modeling challenges: continuous...

    arxiv.org/abs/2608.02379 · PDF

  17. 17

    Faster-WAM: Do World Action Models Need Deep Action Modules?

    Liheng Ma, Rui Heng Yang, Zhanguang Zhang, Mateo Clemente, Ziwen Hu, Tongtong Cao, Yingxue Zhang

    cs.AI · cs.LG · cs.RO

    World Action Models (WAMs) couple robot action prediction with video world models. Existing WAMs with shared-backbone and Mixture-of-Transformers designs generally tie the depth of the action module to that of the video backbone, resulting in substantial computational overhead and high inference latency. To address this limitation, we introduce Dock of Transformer (DoT), a video-centric design principle that treats a pretrained video...

    arxiv.org/abs/2608.02365 · PDF

  18. 18

    SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents

    Yue Yao, Shengyuan Wang, Xin Chen, Minke Zhang, Jia He, Bingjun Luo, Tom Gedeon

    cs.AI

    Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challenge is not merely to retrieve individually relevant skills, but to identify a complete and executable skill composition. In this paper, we argue that this problem can be solved in a graph with three levels: compositional relations among skill queries, similarity between queries and candidates in the skill...

    arxiv.org/abs/2608.02356 · PDF

  19. 19

    KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model Improvement

    Gusseppe Bravo-Rocca, Jordi Guitart, Ajay Dholakia, David Ellison, Puneet Jain

    cs.AI

    Data drift poses significant challenges for machine learning systems in production, requiring continuous model updates to maintain performance. We present KC-Agent, a dual-process cognitive architecture for automated ML model improvement that combines fast pattern recognition (System 1) with deliberate incremental updates (System 2). Our approach implements structured memory systems enabling System 1 to leverage successful solutions...

    arxiv.org/abs/2608.02351 · PDF

  20. 20

    Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling

    Qinwen Wang, Jieping Luo, Aoxiang Qin, Ruoyu Zhao, Jianxiong Tang, Wei Zhang, Zhichao Lu, Luziwei Leng

    cs.AI

    Recurrent linear attention models (RLAs) such as Mamba offer efficient linear-time sequence modeling as an alternative to Transformers, yet their fixed-capacity recurrent states limit long-sequence modeling. Drawing inspiration from hierarchical human memory, we propose Hierarchical Memory Mamba (HMM) to address this limitation. Building upon a pre-trained Mamba backbone, HMM integrates a lightweight working memory that extracts slow...

    arxiv.org/abs/2608.02347 · PDF

  21. 21

    Hard Constraints, Smooth Gradients: Learning Feasible Inventory Policies via Differentiable Projection

    Patrick Helm, Jan-Niklas Doerr, Joren Gijsbrechts, Stefan Minner

    cs.AI · cs.LG

    Many operational problems are constrained sequential decision processes with large, combinatorial action spaces and interdependent feasibility constraints. Mixed-integer linear programs (MILPs) handle such constraints flexibly but scale poorly in stochastic environments. Deep reinforcement learning (DRL) promises scalable decision rules, but existing methods either penalize constraints rather than enforce them, or rely on feasibility...

    arxiv.org/abs/2608.02343 · PDF

  22. 22

    Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit

    Jingxi Wei

    cs.AI · cs.LG · cs.SE

    Long-horizon coding-agent trajectories are poorly matched to the credit units available to train on: a single action has no stable value, an episode label merges productive exploration with abandoned directions, and a fixed window cuts where the logging mechanics fall. We introduce collection-time semantic self-segmentation, in which a declarative contract has the acting agent expose its own boundaries while the trajectory is generated....

    arxiv.org/abs/2608.02302 · PDF

  23. 23

    MechGeo: Autoformalizing and Proving Euclidean Geometry in Lean 4

    Hao Shen, Junyu Guo, Tian Cui, Yuxuan Xiao, Lihong Zhi

    cs.AI · cs.LO

    We present MechGeo, a Mathlib native agentic framework that jointly addresses faithful autoformalization and certified proof construction for Euclidean geometry. In this framework, GeoFormalizer represents informal problems in GeoIR, deterministically translates them into Lean 4, and iteratively repairs candidate statements using structural diagnostics and semantic evaluation. GeoProver constructs geometric proof plans, derives intermediate...

    arxiv.org/abs/2608.02295 · PDF

  24. 24

    Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning

    Yiqing Liu, Zihao Wang, Hantao Yao, Wu Liu, Yongdong Zhang

    cs.AI

    Multi-agent reasoning (MAR) improves reasoning reliability through iterative solution exchange and refinement. Existing adaptive MAR methods typically learn routing decisions from query-level labels or trajectory-level returns, but such coarse supervision cannot accurately estimate the state-conditioned utility of individual operators in multi-step collaboration. We propose TreeCredit, a shared-prefix credit assignment framework for efficient...

    arxiv.org/abs/2608.02291 · PDF

  25. 25

    SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

    Zelin Tan, Yiqun Zhang, Hao Li, Zhiyao Cui, Hejia Geng, Shao Zhang, Hangfan Zhang, Yang Chen, Xiaosong Wang, Lilong...

    cs.AI

    Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that current models can effectively identify, apply, and coordinate them. To improve skill-use capabilities, we introduce SKT, a verified data synthesis pipeline that constructs skill-grounded tasks and executable trajectories from large collections of agent skills. SKT...

    arxiv.org/abs/2608.02287 · PDF

  26. 26

    Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories

    Shuai Shao, Kangning Zhang, Qingyao Li, Shijian Wang, Hao Wang, Wenxiang Jiao, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang

    cs.AI

    Agents built around large language models continually accumulate interaction trajectories during deployment, yet their behavior typically remains fixed. Beyond updating model weights, these trajectories can improve the agent harness that constructs context, mediates tools, validates actions, and recovers execution. We introduce Harness-R1, the first method, to our knowledge, that makes failure-conditioned, lifecycle-wide editing of an...

    arxiv.org/abs/2608.02276 · PDF

  27. 27

    Self-Certification of Representation Adequacy: Sequential Certification at Minimum Task Loss

    Zijie Huang

    cs.AI · cs.LG

    Agents that act on a compressed representation of their history face a structural risk: if the representation aliases histories with different optimal actions, no rule measurable with respect to the representation can avoid an irreducible per-round loss, and the agent may be unable to detect this from its own transcript. This paper develops a four-layer theory of self-certification of representation adequacy. The static layer defines...

    arxiv.org/abs/2608.02267 · PDF

  28. 28

    Homebot: A Personal AI Agent for Conversational Home Assistance and Automation

    Shengyuan Ye, Yixin Zhang, Han Liang, Liekang Zeng, Jiangsu Du, Mu Yuan

    cs.AI

    \texttt{Homebot} is a locally deployable AI agent for conversational household assistance and automation. It accepts voice and instant-messaging requests through a shared runtime that combines language-model responses with registered tools and task-specific skills. The design separates common request processing from session ownership: messaging history remains scoped to a channel and chat, whereas voice interaction is bounded by wake-word...

    arxiv.org/abs/2608.02254 · PDF

  29. 29

    Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability

    Abdullah Mamun, Shovito Barua Soumma, Hassan Ghasemzadeh

    cs.AI · cs.LG

    Ensuring trust in AI systems is essential for the safe and ethical integration of machine learning systems into high-stakes domains such as digital health. Key dimensions, including robustness, explainability, fairness, accountability, and privacy, need to be addressed throughout the AI lifecycle, from problem formulation and data collection to model deployment and human interaction. While various contributions address different aspects of...

    arxiv.org/abs/2608.02238 · PDF

  30. 30

    PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs

    Haojie Hu, Chenhao Dang, Yaojia Liu, Hengrui Kang, Conghui He, Weijia Li

    cs.AI

    Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs; direct image generation is not element-editable, while coding-agent workflows are costly. PosterMELD is a template-conditioned multi-agent pipeline: capacity-aware slots guide writing before rendering, and deterministic gates plus vision-language model (VLM) review...

    arxiv.org/abs/2608.02218 · PDF

  31. 31

    MEGRAG: Multi-Granular Evidence Graphs for Answer-Aware Multi-Hop RAG

    Weidong Bao, Yingying Sun, Jun Yang, Yilin Wang, Zili Wei, Yubin Bao, Fangling Leng, Minghe Yu, Tiancheng Zhang, Ge Yu

    cs.AI

    Multi-hop question answering is a fundamental challenge in retrieval-augmented generation (RAG), because deriving an answer requires integrating dispersed evidence. Iterative RAG (iRAG) is widely used for this challenge, but existing methods have two limitations. First, most methods still support each reasoning step with single-granularity evidence, making it difficult to balance information density and contextual noise. Second, existing...

    arxiv.org/abs/2608.02195 · PDF

  32. 32

    PAC Approximation and DIRECT Optimization for Parametric Markov Models

    Zhiming Chi, Ying Liu, Andrea Turrini, Lijun Zhang, David N. Jansen

    cs.AI · cs.FL · cs.LO

    In this paper, we consider the parameter synthesis and optimization problem for parametric Markov decision processes (pMDPs), the extension of classical MDPs where exact probability values are replaced by parametric expressions. Computing the rational function $f_{\lsf}$ that maps parameter valuations to the satisfaction value of a PRCTL property $\lsf$ is a computationally expensive task, particularly for pMDPs where the optimal policy may...

    arxiv.org/abs/2608.02184 · PDF

  33. 33

    From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents

    Jiajia Song, Bobo Li, Haiwen Yi, Zibo Ji, Meishan Zhang, Hao Fei, Min Zhang, Mong-Li Lee, Wynne Hsu

    cs.AI

    Large Language Models have enabled increasingly capable autonomous agents, yet personalization remains critical for making such agents practically useful. Recent benchmarks have begun evaluating personalization in agents, but they largely rely on static preference snapshots, fixed interaction logs, or question answering over predefined user profiles. Such designs fail to capture the complexity of evolving user preferences and neglect...

    arxiv.org/abs/2608.02171 · PDF

  34. 34

    From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution

    Can Wang, Haoran Chen, Haowen Gao, Hao Ding, Zhaoyang Liu, Zhiying Tu

    cs.AI

    Deep research benchmarks require expert-level tasks and reliable evaluation grounded in task-specific knowledge. Existing benchmarks rely heavily on expert authoring or pre-existing human-authored materials, while fully automatic construction struggles to ensure consistent and traceable verification. To address this gap, we introduce a verifiable benchmark of 500 deep research tasks spanning 31 topics and 10 major categories, with three query...

    arxiv.org/abs/2608.02163 · PDF

  35. 35

    Auditing Data Provenance in LLM Fine-tuning via Intrinsic Distributional Fingerprints

    Zirui Huang, Yunlong Mao, Wei Tong, Tingting Wu, Xin Ge, Sheng Zhong

    cs.AI

    The proliferation of customized Large Language Models (LLMs) poses critical risks of Data Intellectual Property (Data IP) infringement via unauthorized fine-tuning on proprietary data. Existing audit techniques are limited, as they require intervention during data preparation or training and remain fragile under malicious obfuscations such as data paraphrasing and knowledge distillation. We propose \textit{Distribution Provenance Audit...

    arxiv.org/abs/2608.02154 · PDF

  36. 36

    Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning

    Yijun Zhang, Yule Xie, Jiaxin Ding, Xin Ding, Fan Xu, Haoxiang Zhang, Luoyi Fu

    cs.AI

    Reinforcement learning has become a central paradigm for improving the reasoning capabilities of large language models. Existing methods generally aim to reduce the failure probabilities induced across problems. In this paper, we introduce a moment-based perspective on policy optimization for LLM reasoning by treating the failure probability of a randomly sampled problem as a random variable and characterizing optimization objectives through...

    arxiv.org/abs/2608.02149 · PDF

  37. 37

    Beyond Solution-Centric Search: Adaptive Inquiry and Knowledge Revision for Autonomous ML Engineering

    Shaokang Fu, Yulong Tao, Linbo Jin, Jiarong Zhao, Qiming Shi, Tianjun Pan, Haonan Li, Chengyu Wang, Jia Wu, Chengfu Huo

    cs.AI

    Long-horizon autonomous research tasks such as machine learning engineering require systems to make interdependent decisions under a limited budget. Existing LLM-based agents typically organize candidate-solution improvement through tree, graph, or chain structures, meaning that the search process determines how information is acquired and managed. We call this design solution-centric search and propose instead the information paradigm, in...

    arxiv.org/abs/2608.02143 · PDF

  38. 38

    MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents

    Jiajun Dong, Yutao Hu, Fengrui Fan, Shihan Dou, Yueming Wu, Deqing Zou

    cs.AI

    Large language model (LLM) agents must retain and use cross-step information to act coherently in long-horizon tasks. Existing methods improve memory accessibility, yet action-relevant information may still fail to guide the current decision because it is poorly formed, organized, prioritized, or presented. We call this post-access failure the Memory-Action Gap. We propose MemArbiter, a function-aware memory arbitration framework that...

    arxiv.org/abs/2608.02113 · PDF

  39. 39

    Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents

    Qi Liu, Yiqun Chen, Zidan Chen, Yan Gao, Yi Wu, Yao Hu, Jiaxin Mao, Fengbin Zhu, Tat-Seng Chua

    cs.AI · cs.IR

    Search agents now answer questions that take dozens of searches to settle, yet how such an agent reads a page has drawn far less attention than how it finds one. Nearly all of them use one of two document interfaces, and both tie a page to the moment it is opened. \emph{Visit-and-read} injects a reading of the page into the message history at fetch time, fixing that reading before the agent knows which fact it will need. Stateful...

    arxiv.org/abs/2608.02097 · PDF

  40. 40

    Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation

    Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein

    cs.AI · cs.CL · cs.LG

    Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introduces challenges distinct from classical RL, with implications for inducing exploration. New methods are required that leverage the broad knowledge and flexibility of pre-trained LLMs to deliberately generate diverse experience at training time. We propose...

    arxiv.org/abs/2608.02087 · PDF

  41. 41

    Cross-Fitted Residual Utility for Primary-Preserving Cognitive Decision Correction in Automatic Modulation Classification

    Linzhuo Han, Zongyong Cui, Houbiao Li

    cs.AI

    Automatic modulation classification research has largely emphasized representation accuracy, but a cognitive receiver must also decide when heterogeneous evidence justifies overriding a trusted default prediction. We study this post-inference problem through cross-fitted residual utility and a primary-preserving cognitive decision policy. A structured KAN-Fourier classifier supplies the default probability, while neural and non-neural...

    arxiv.org/abs/2608.02063 · PDF

  42. 42

    HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning

    Runchuan Zhu, Hongbin Lai, Bowen Jiang, Junrui Zhang, Zhangheng LI, Ostap Kilbasovych, Junyuan Hong

    cs.AI

    Reflection is a powerful mechanism for LLM reasoning, yet its effectiveness hinges on accurately attributing failures to specific reasoning steps, a capability that current models notably lack. Existing failure attribution methods either require expensive step-by-step counterfactual testing that scales poorly with trajectory length, or treat reasoning traces as flat sequences that ignore the inherent non-linear logical dependencies. We...

    arxiv.org/abs/2608.02026 · PDF

  43. 43

    EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers

    Junyeong Park, Jieun Han, Haneul Yoo, So-Yeon Ahn, Jinsung Yoon, Alice Oh

    cs.AI · cs.CY

    Large language models (LLMs) are increasingly used across diverse tasks in K-12 education, yet existing safety evaluations rarely examine how harmful or inappropriate content appears in interactions between LLMs and students or teachers. To address this, we present EduZone, an evaluation framework for LLM safety across diverse educational scenarios. Our framework systematically combines (1) student- and teacher-facing LLM usage contexts, (2)...

    arxiv.org/abs/2608.02024 · PDF

  44. 44

    Before Reasoning Fails: Pre-Evidence Procedural Failures in Agentic RAG

    Daeyoung Roh, Donghee Han

    cs.AI

    Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them. We study this failure mode as a procedural property of the agent trajectory, decomposing wrong answers into pre-evidence discipline failures and post-gold-read failures using saved tool-call traces, retrieved evidence, read passages, and final answers....

    arxiv.org/abs/2608.02011 · PDF

  45. 45

    HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents

    Daeyoung Roh, Donghee Han

    cs.AI

    Retrieval-augmented search agents answer multi-hop questions by repeatedly issuing search queries and accumulating evidence. This creates a stopping problem: after the necessary evidence has appeared, further retrieval often adds cost, latency, and distracting context rather than useful information. We frame stopping as evidence coverage rather than generator confidence, and introduce HALT, a lightweight verification-aware policy that leaves...

    arxiv.org/abs/2608.02009 · PDF

  46. 46

    Evolving in the Agent Jungle via History-Informed Opponent Awareness

    Zhaofeng Zhang, Linhan Xia, Rui Liu, Yihao Wang, Binrui Shen, Shengxin Zhu

    cs.AI

    Learning to adapt strategies through interaction is a key step toward more general and autonomous LLM agents. Existing approaches typically achieve behavioral adaptation by revising skill libraries. However, in multi-agent environments, opponents may simultaneously update their strategies, causing the environment itself to evolve continuously. Applying skill-revision methods designed for static environments in such settings therefore amounts...

    arxiv.org/abs/2608.02005 · PDF

  47. 47

    Long-Horizon Autonomous Architecture Research with a Language-Model Agent: A Behavioural Case Study

    Aon Safdar, Mohamed Saadeldin

    cs.AI

    We study what happens when a single general-purpose large language model acts as the sole researcher on a long-horizon neural architecture design problem. The agent receives a scientific question, an initial hypothesis and motivation, a compute budget, and research affordances (source and experiment management, experiment tracking, literature access, and persistent memory), then autonomously proposes, implements, evaluates, and records...

    arxiv.org/abs/2608.01995 · PDF

  48. 48

    A Contractualist Argumentation Framework for Moral Decision-Making

    Luis Marcos-Vidal, Giulio Antonio Abbo, Tony Belpaeme

    cs.AI · cs.CY

    Autonomous agents operating in shared environments must make decisions that affect multiple individuals with potentially conflicting interests. We propose a formal framework for moral decision-making grounded in Scanlon's contractualism, an ethical theory that evaluates the permissibility of actions in terms of principles that no one could reasonably reject. To operationalise contractualist reasoning, we use ASPIC+, a structured argumentation...

    arxiv.org/abs/2608.01937 · PDF

  49. 49

    ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching

    Zihan Liu, Yuzhe Zhuang, Yuanzu Li, Wanshuang Gou, Jiahong Liu, Min Zhou, Menglin Yang

    cs.AI

    JEPA-style visual world models offer an effective paradigm for visual goal planning by predicting future latent representations. Existing methods typically learn local transition consistency through next-step representation prediction. However, in long-horizon tasks, accurate local prediction alone need not ensure sustained progress toward the goal. First, multi-step rollouts can remain locally plausible while drifting away from goal-relevant...

    arxiv.org/abs/2608.01926 · PDF

  50. 50

    Diagnosing Search Behavior and Failure Modes in Long-Horizon Search Agents

    Qi Liu, Jiaxin Mao, Fengbin Zhu, Tat-Seng Chua

    cs.AI · cs.CL · cs.IR

    Deep search agents answer difficult information-seeking questions by iteratively issuing search queries to gather supporting evidence, but it remains unclear whether and how greater search effort leads to better answers. We study these questions through a trajectory-level diagnosis of long-horizon search agents. Using human-annotated document-level relevance judgments, we evaluate the evidence retrieved at each search step and separate two...

    arxiv.org/abs/2608.01913 · PDF

  51. 51

    CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents

    Zhaoyang Li, Zenghuang Fu, Qiuyuan Ai, Ping Jiang, Haoyu Wu, Minghui Wu, Chenxu Zhao, Jie Song, Guannan He

    cs.AI

    Large language models can improve with reinforcement learning for search agents, yet existing self play agents repeatedly generate tasks while discarding the knowledge gained during successful searches. We introduce CoEvoKG, a framework that turns a knowledge graph into both a source of verifiable training tasks and a persistent evidence memory for agent evolution. CoEvoKG jointly trains a task generator and a search agent: the generator...

    arxiv.org/abs/2608.01904 · PDF

  52. 52

    ReasonCast: Towards Explainable Time Series Forecasting with Reasoning

    Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan...

    cs.AI · cs.LG

    Most time series (TS) models are specialized for a single task, either understanding (i.e., returning text answers about a TS) or generation (i.e., returning a numeric forecast). Only recently have unified models begun to handle the two within a single architecture. Even these models, however, produce the two outputs as task-separated paths and cannot predict a series and explain why that prediction arises within a single coherent response....

    arxiv.org/abs/2608.01875 · PDF

  53. 53

    Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge

    Xiaofeng Shi, Xiaosong Qiu, Wenxin Ma, Qian Kou, Yiming Pan, Longbin Yu, Ying Liu, Haiping Wang, Hua Zhou

    cs.AI

    Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On the 707-question WnuanBench, the primary 32B route raises acceptable-answer rate (AAR) from 52.76% before...

    arxiv.org/abs/2608.01862 · PDF

  54. 54

    EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning

    Dongwei Sun, Bowen Yao, Yujie Zhang, Pei Liu, Jing Yao, Xiangyong Cao

    cs.AI

    Bi-temporal remote-sensing disaster change captioning often needs to identify sparse and spatially localized changes across large pre- and post-event scenes and then translate them into coherent, factual descriptions. However, existing change captioning methods always follow an autoregressive decoding paradigm to generate the change description and thus an early misinterpretation of the changed object, event, or spatial relation becomes an...

    arxiv.org/abs/2608.01856 · PDF

  55. 55

    Physics-Informed Neural Networks for Complex Eigenfrequency Identification and Mode Structure Reconstruction of the Ground-State ITG Branch

    Dengdi Sun, Bingbing Zhang, Xiao Wang, Zikang Yan, Yuqiang Tao, Qingquan Yang, Guosheng Xu, Jin Tang

    cs.AI

    Physics-informed neural networks (PINNs) combine sparse observations with physical equations, providing an important approach for modeling complex plasma processes and inferring unknown physical quantities. The steep-gradient pedestal of high-confinement-mode tokamaks is closely linked to plasma confinement and edge transport. Analyzing ion-temperature-gradient (ITG) drift waves in this region requires jointly identifying complex...

    arxiv.org/abs/2608.01850 · PDF

  56. 56

    Exploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models

    Junxiang You, Junkai Chen, Yuhao He, Ruiqi Liu, Zhetao Guo, Shu Wu

    cs.AI

    Machine unlearning offers a promising approach to remove unsafe content from Multimodal Large Language Models (MLLMs), yet ensuring the precision of unlearning remains a persistent challenge. One reason is that current MLLM unlearning evaluation paradigms suffer from a critical blind spot: they assess model utility through benchmarks whose representations are distant from the forget set, failing to capture knowledge holes---severe degradation...

    arxiv.org/abs/2608.01849 · PDF

  57. 57

    FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling

    Xianglong Yan, Hong Liu, Chengzhu Bao, Tianao Zhang, Guanghua Yu, Jianchen Zhu, Yulun Zhang

    cs.AI

    Large language models (LLMs) achieve remarkable performance but are expensive to deploy due to their enormous size. FP4 quantization, with formats such as MXFP4 and NVFP4, offers an appealing solution with native hardware support on modern accelerators. However, maintaining accuracy under FP4 precision remains difficult. A key bottleneck lies in scale optimization: existing methods tightly couple the quantization and dequantization scales,...

    arxiv.org/abs/2608.01847 · PDF

  58. 58

    PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

    Chunji Lv, Yangguang Wei, Junlin Liu, Yang Gao, Ming Liu, Xinming Wang, Jinyang Wu, Guoren Wang, Changsheng Li

    cs.AI

    Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-turn trajectory may receive only a single outcome-level signal. On-policy self-distillation (OPSD) provides dense token-level supervision from a privileged teacher, but the teacher may not be reliable at every position. Existing methods commonly rely on isolated...

    arxiv.org/abs/2608.01837 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.