cs.AI · 2026-07-02 · No. 41

Artificial Intelligence, 2026-07-02.

20 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

20 entries
  1. 01

    AutoMem: Automated Learning of Memory as a Cognitive Skill

    Shengguang Wu, Hao Zhu, Yuhui Zhang, Xiaohan Wang, Serena Yeung-Levy

    cs.AI · cs.CL · cs.MA

    Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capacity known in cognitive science as metamemory. We bring this perspective to LLMs by treating memory management as a trainable skill. We promote file-system operations to first-class memory actions alongside task actions, letting the model itself decide how to manage its memory. This memory skill improves along two axes: the...

    arxiv.org/abs/2607.01224 · PDF

  2. 02

    Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

    Ben Slivinski, Michael Saldivar

    cs.AI · cs.CL · cs.LG · cs.LO · cs.SE

    When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; scalar LLM judges offer coverage but produce opaque scores that cannot be audited after the fact and are subject to the same coherence issues as any LLM. We present Theoria, a verification architecture that closes this gap. A candidate solution is rewritten into a sequence of typed state transitions, each...

    arxiv.org/abs/2607.01223 · PDF

  3. 03

    Optimal Resource Utilization for Autonomous Laboratory Orchestrators

    Austin McDannald, Julia Tisaranni, Howie Joress

    cs.AI · cond-mat.mtrl-sci

    In autonomous laboratories, AI agents suggest the next batch of experiments to do. However, planning and executing those tasks taking full advantage of the available resources is a completely different question. This can be challenging when dealing with real-world hardware constraints, especially so when there are multiple instruments with different capacities and throughputs. Here we demonstrate a 2-step method to address resource...

    arxiv.org/abs/2607.01188 · PDF

  4. 04

    Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

    Song-Lin Lv, Weiming Wu, Rui Zhu, Zi-Jian Cheng, Lan-Zhe Guo

    cs.AI

    While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, tool sets, and interaction dynamics. To address this generalization gap, we formalize OpenAgent (Tool-Use Agent in Open-World), a problem setting characterized by distributional shifts across query, action, observation, and domain dimensions. To systematically...

    arxiv.org/abs/2607.01084 · PDF

  5. 05

    Agentic generation of verifiable rules for deterministic, self-expanding reaction classification

    Daniel Armstrong, Maarten Dobbelaere, Valentas Olikauskas, Helena Avila, Octavian Susanu, Jérôme Waser, Philippe Schwaller

    cs.AI · cs.CL

    Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign each transformation a deterministic, interpretable label. But chemistry is long-tailed, making manual encoding intractable, and existing tools rely on fixed rulesets that cannot adapt to new chemistries. Here we present a fully automated pipeline in which a multi-agent framework of large language models...

    arxiv.org/abs/2607.01061 · PDF

  6. 06

    PedNStream: Scalable Network Flow Simulation for Pedestrian Traffic Management

    Weiming Mai, Dorine Duives, Serge Hoogendoorn

    cs.AI

    Large-scale crowd management requires pedestrian simulations that are both computationally efficient and compatible with feedback-based control. However, most open-source tools are either microscopic or not designed for network-scale closed-loop evaluation. This paper presents PedNStream (Pedestrian Network Flow Simulation), an open-source, Python-native simulator for macroscopic pedestrian network loading based on the Link Transmission Model...

    arxiv.org/abs/2607.01021 · PDF

  7. 07

    Bayesian Uncertainty Propagation for Agentic RAG Pipelines: A Proof-of-Concept Study on Multi-Hop Question Answering

    Louis Donaldson, Connor Walker, Koorosh Aslansefat, Yiannis Papadopoulos

    cs.AI

    Trustworthy deployment of Agentic Retrieval-Augmented Generation (RAG) systems requires mechanisms for estimating when multi-stage reasoning pipelines may fail. This paper presents an uncertainty-aware Agentic Retrieval-Augmented Generation (RAG) framework in which planner, evaluator and generator stages produce uncertainty signals derived from semantic divergence and generator self-evaluation. These signals are propagated through a Bayesian...

    arxiv.org/abs/2607.00972 · PDF

  8. 08

    Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

    Subhadeep Pal, Shashwat Sourav, Tirthankar Ghosal, Markus J. Buehler

    cs.AI · cond-mat.mtrl-sci · cs.CL · cs.LG

    Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-step, domain-grounded reasoning. Standard large language models often produce fluent but weakly traceable responses to open-ended materials design problems, making it difficult to determine whether final answers are supported by coherent intermediate reasoning. We develop Graph-PRefLexOR, a family of graph-native reasoning...

    arxiv.org/abs/2607.00924 · PDF

  9. 09

    Two AI Metrics Diverged: Will it Make All the Difference?

    Alex Fogelson, Zachary A. Brown, Hans Gundlach, Jayson Lynch, Neil Thompson

    cs.AI

    As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget? Or will capabilities converge, with "meek models inheriting the earth"? Building on Gundlach et al. (2025b), we show that the answer depends on how we value and measure AI capabilities. We discuss conventional performance measures and show that, while validation loss shows a shrinking gap, on...

    arxiv.org/abs/2607.00913 · PDF

  10. 10

    Self-Evolving Agents with Anytime-Valid Certificates

    Biswa Sengupta

    cs.AI · cs.CL

    Self-evolving agents violate the assumption behind most learning-theoretic guarantees: the data, evaluator, components, and hypothesis space are produced by the policy being updated. We present \textbf{SEA}, an architecture that confines self-modification to a small steering adapter and a versioned harness around a \emph{frozen} base model and admits each modification only through an anytime-valid gate that emits an auditable certificate...

    arxiv.org/abs/2607.00871 · PDF

  11. 11

    Self-GC: Self-Governing Context for Long-Horizon LLM Agents

    Xubin Hao, Hongjin Meng, Xin Yin, Jiawei Zhu, Chenpeng Cao

    cs.AI

    Long-horizon LLM agents accumulate tool results, files, plans, and user constraints that are too structured to be treated as a disposable text suffix. Current systems mostly rely on in-run heuristics such as chronological pruning and tool-output masking, or on final self-summary near a context limit. Heuristics are cheap but blind to future dependencies; summaries preserve narrative state but often hide exact evidence, locators, and editable...

    arxiv.org/abs/2607.00692 · PDF

  12. 12

    Coachable agents for interactive gameplay

    Roberto Capobianco, Harm van Seijen, Nolan D. Bard, Neil Burch, Fatima Davelouis, Josh Davidson, Alisa Devlic,...

    cs.AI · cs.LG

    Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models. Through trial-and-error, these AI systems typically learn one, near-optimal behavior to solve their tasks. However, there are many use cases in which one would like to assert some level of control, preferably in real time, over how the task is solved. We...

    arxiv.org/abs/2607.00642 · PDF

  13. 13

    AGI Maze as a Benchmark Framework for World-Modeling Agents

    Alexey Potapov

    cs.AI

    Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a static context - does not reliably produce persistent, manipulable representations of an external world. Many tasks that look like "reasoning" in text become substantially harder once the environment is partially observable, stateful, and requires memory and structured hypotheses about hidden state. AGI...

    arxiv.org/abs/2607.00627 · PDF

  14. 14

    HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

    Shei Pern Chua, Fangzhao Wu

    cs.AI · cs.CR

    Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or...

    arxiv.org/abs/2607.00572 · PDF

  15. 15

    AI Native Games: A Survey and Roadmap

    Zhiyue Xu, Fandi Meng, Kaijie Xu, Clark Verbrugge, Simon Lucas, Jian Zhao

    cs.AI

    Generative AI now enables games to produce dialogue, quests, characters, images, and worlds at runtime. Yet generation alone does not make a game AI-native, nor does it guarantee playability. This paper defines AI-native games by whether runtime generative AI is constitutive of the core loop: if the AI component were removed or trivially replaced, the central form of play would collapse or become fundamentally different. This counterfactual...

    arxiv.org/abs/2607.00527 · PDF

  16. 16

    Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments

    Jinwoo Jang, Daniel J. Rho, Sihyung Yoon, Hyunsuk Cho, Honguk Woo

    cs.AI

    Embodied agents operating in the real world require multi-scale reasoning and knowledge adaptation as conditions change. We identify two challenges in applying Mixture of Experts (MoE) to this setting: routing lacks an explicit notion of scale, preventing targeted updates at specific scales, and a uniform update policy cannot accommodate the different rates at which knowledge at each scale becomes outdated. We present MuSix, a framework that...

    arxiv.org/abs/2607.00457 · PDF

  17. 17

    Agri-SAGE: Simulation-Grounded Multi-Agent LLM for Context-Aware Agricultural Advisory Generation

    Vedant Balasubramaniam, Geetha Charan, Manojkumar Patil, Rohit P Suresh, V Priyanka, Kodur Sai Vinay Sathvik, Y. Narahari

    cs.AI · cs.MA

    Agricultural advisory systems face a fundamental tension: static agronomic guidelines offer consistent, evidence-based recommendations, yet remain blind to in-season variability and dynamic uncertainties. Recent advisory systems powered by LLMs are liable for a different risk of generating recommendations that are agronomically credible but physiologically unconvincing. Agri-SAGE is a closed-loop framework designed to resolve the above two...

    arxiv.org/abs/2607.00454 · PDF

  18. 18

    PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

    Ke Zhang, Sahchit Chundur, Mohammad Javad Qomi, Maziar Raissi

    cs.AI

    Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely more complex. We introduce PHREEQC-MCQ-200, a benchmark for evaluating tool-augmented agents on deterministic aqueous-geochemistry simulations. The benchmark contains 200 multiple-choice questions derived from 21 validated PHREEQC scenarios, requiring agents to...

    arxiv.org/abs/2607.00436 · PDF

  19. 19

    Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising

    Tianci Liu, Zihan Dong, Linjun Zhang, Haoyu Wang, jing Gao, Emre Kiciman, Ranveer Chandra, Wei-Ting Chen

    cs.AI

    Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design. Solely relying on prespecified templates or user verbose instructions, they fail to capture latent design intents, leaving Page-level Slide Personalization (PSP) unresolved. To close this gap, this work formulates PSP as an inverse planning problem. We propose to learn a design intent...

    arxiv.org/abs/2607.00407 · PDF

  20. 20

    Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems

    Srini Ramaswamy, Wang Miaosheng

    cs.AI

    Autonomous agents, whether LLM-driven software agents or robotic physical agents, face a common class of failure modes when operating without continuous human oversight: safety violations from unverified actions, behavioral instability from unconstrained loops, and continuity loss from unhandled error states. We develop \system{}, a discrete-time control system that combines five execution gears (\Gobs{}, \Gsug{}, \Gplan{}, \Gexec{}, \Gint{})...

    arxiv.org/abs/2607.00334 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.