cs.AI · 2026-06-06 · No. 15

Artificial Intelligence, 2026-06-06.

52 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

52 entries
  1. 01

    MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery

    Shangheng Du, Xiangchao Yan, Jinxin Shi, Zongsheng Cao, Shiyang Feng, Zichen Liang, Boyuan Sun, Tianshuo Peng, Yifan...

    cs.AI · cs.CL

    Large language model (LLM) agents are increasingly applied to long-horizon tasks such as scientific discovery and machine learning engineering (MLE), where sustained self-evolution becomes a key capability. However, existing MLE agents suffer from inter-branch information isolation, memoryless search, and lack of hierarchical control, which together hinder long-horizon optimization. We present MLEvolve, an LLM-based self-evolving multi-agent...

    arxiv.org/abs/2606.06473 · PDF

  2. 02

    Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement

    Jui-Hui Chung, Ziyang Cai, Zihao Li, Qishuo Yin, Rohit Agarwal, Simon Park, Rodrigo Porto, Narutatsu Ri, Ziran Yang,...

    cs.AI

    We introduce Goedel-Architect, an agentic framework for formal theorem proving in Lean 4 centered on blueprint generation and refinement. A blueprint is a dependency graph of definitions and lemmas that builds up to the main theorem. First, Goedel-Architect generates a blueprint of formally stated definitions and lemmas, along with declared dependencies. This blueprint is optionally guided by a natural language proof. Then, a tool-equipped...

    arxiv.org/abs/2606.06468 · PDF

  3. 03

    Benchmark Everything Everywhere All at Once

    Shiyun Xiong, Dongming Wu, Peiwen Sun, Yuang Ai, Bokang Yang, Wencheng Han, Xiao-Hui Li, Xiangyu Yue

    cs.AI

    Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance. However, their construction is labor-intensive and hard to reuse, raising concerns about sustainability and scalability. Moreover, existing benchmarks often quickly reach performance saturation after their release, resulting in insufficient discrimination among state-of-the-art models. To address these...

    arxiv.org/abs/2606.06462 · PDF

  4. 04

    Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents

    Zhuoming Chen, Xinrui Zhong, Qilong Feng, Ranajoy Sadhukhan, Yang Zhou, Michael Qizhe Shieh, Zhihao Jia, Beidi Chen

    cs.AI

    Sparse attention is becoming increasingly important for serving large language models (LLMs) as generation lengths continue to grow. However, deploying and evaluating new sparse attention algorithms at scale remains highly engineering-intensive, slowing both human researchers and AI agents in exploring the sparse attention design. To address this challenge, we present Vortex, a system that combines a Python-embedded frontend language atop a...

    arxiv.org/abs/2606.06453 · PDF

  5. 05

    Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads

    Yasmine Omri, Ziyu Gan, Zachary Broveak, Robin Geens, Zexue He, Alex Pentland, Marian Verhelst, Tsachy Weissman,...

    cs.AI

    LLM agents are increasingly deployed on long-horizon tasks requiring sustained reasoning over extended interaction histories. Realizing this at scale requires agents to persistently store, retrieve, and update their own memory across sessions. A rich ecosystem of agent memory systems has emerged spanning flat retrieval, LLM-mediated extraction, consolidating fact stores, and agentic control flows. Yet, their system-level behavior remains...

    arxiv.org/abs/2606.06448 · PDF

  6. 06

    Unsupervised Skill Discovery for Agentic Data Analysis

    Zhisong Qiu, Kangqi Song, Shengwei Tang, Shuofei Qiao, Lei Liang, Huajun Chen, Shumin Deng

    cs.AI · cs.CL · cs.LG · cs.MA

    Inference-time skill augmentation provides a lightweight way to improve data-analytic agents by injecting reusable procedural knowledge without updating model parameters. However, discovering effective skills for data analysis remains challenging, as reliable supervision is expensive and success criteria vary across analytical formats. This raises the key question of how to discover reusable data-analysis skills from unlabeled exploration...

    arxiv.org/abs/2606.06416 · PDF

  7. 07

    Risk Assessment of Autonomous Driving: Integrating Technical Failures, Ethical Dilemmas, and Policy Frameworks

    Boyi Chen, Shengqin Chu, Zicheng Wang, Brian Baetz, Zhen Gao

    cs.AI

    Autonomous driving technology has the potential to reduce the large number of road traffic accidents caused by human error each year, but it also brings new types of risks that need to be evaluated from the aspects of technology, ethics and regulations. Based on public crash data from the National Highway Traffic Safety Administration (NHTSA), disengagement reports from the California Department of Motor Vehicles (DMV), the MIT Moral Machines...

    arxiv.org/abs/2606.06396 · PDF

  8. 08

    Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

    Jiaju Chen, Yuxuan Lu, Jiayi Su, Chaoran Chen, Songlin Xiao, Zheng Zhang, Yun Wang, Yunyao Li, Jian Zhao, Tongshuang...

    cs.AI · cs.CL

    Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly position these agents as human collaborators. Effective collaboration, however, requires collaborators to continuously maintain and align mental models of their own reasoning,partners' intentions, and shared goals during the collaborative process. Today's agents rarely develop such capabilities...

    arxiv.org/abs/2606.06388 · PDF

  9. 09

    Rethinking Infrastructure Inspection as Image Difference Classification: A Traffic Sign Case Study

    Ching Yau Fergus Mok, Lavindra de Silva, Varun Kumar Reja, Ioannis Brilakis

    cs.AI

    Digital twins (DTs) allow the digitalization of road infrastructure inspection, though this is hindered by limited annotated data. This work exploits the relational nature of continuous asset condition monitoring to reformulate image-based defect detection as image difference classification (IDC) to reduce data reliance. This was evaluated in a case study on low-resource traffic sign inspection with different IDC classifiers using a...

    arxiv.org/abs/2606.06375 · PDF

  10. 10

    An Infectious Disease Spread Simulation Based on Large Language Model Decision Making

    Yonchanok Khaokaew, Ruochen Kong, Andreas Zufle, Hao Xue, Taylor Anderson, Chandini Raina MacIntyre, Matthew Scotch,...

    cs.AI

    Modelling individual decision-making during infectious disease outbreaks is crucial for understanding behavioural dynamics and informing effective public health interventions. Prior work has shown that large language models can simulate realistic human behaviour by generating agent decisions based on demographic prompts and situational context. We build on this foundation with a spatially grounded, agent-based simulation framework that...

    arxiv.org/abs/2606.06360 · PDF

  11. 11

    Where Should Knowledge Enter? A Layered Framework for Knowledge Infusion in Multimodal Iterative Generative Mo

    Renjith Prasad, Chathurangi Shyalika, Anushka Pawar, Amit Sheth

    cs.AI

    Multimodal generative models produce fluent outputs but remain unreliable when generation must respect structured, domain-specific, or safety-critical knowledge. Existing methods incorporate knowledge through mechanisms such as prompt augmentation, guidance, latent editing, or fine-tuning, yet they are typically categorized by technique rather than by the component of the generative process they modify. We argue that knowledge infusion in...

    arxiv.org/abs/2606.06356 · PDF

  12. 12

    Boosting Brain-to-Image Decoding with TRIBE v2 Data Augmentation

    Yohann Benchetrit, Marlène Careil, Simon Dahan, Hubert Banville, Stéphane d'Ascoli, Jean-Rémi King

    cs.AI · cs.LG · q-bio.NC

    Brain decoding is limited by the availability of labeled neural data, and remains challenging in low-data regimes. To address this issue, we investigate whether and when brain decoding can be boosted by augmenting small fMRI datasets with synthetic data generated by a pretrained model of fMRI responses to stimuli. We use TRIBE v2, a large encoding model pretrained on more than 1000 hours of fMRI responses to video, audio and language. For...

    arxiv.org/abs/2606.06345 · PDF

  13. 13

    TokenMizer: Graph-Structured Session Memory for Long-Horizon LLM Context Management

    Shweta Mishra

    cs.AI

    Large language model (LLM) deployments for long-horizon tasks face a fundamental constraint: context windows are finite while productive work sessions are not. When history exceeds the Maximum Effective Context Window (MECW), critical structured information - architectural decisions, task transitions, file histories - is silently discarded. Existing mitigations treat history as flat text, destroying the relational structure that makes...

    arxiv.org/abs/2606.06337 · PDF

  14. 14

    DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions

    Nathan Bout, Maxime Langevin, Ronan Riochet

    cs.AI

    GUI agents - vision-based models that control desktops, web browsers, and mobile devices through graphical user interfaces - promise to automate a wide range of digital tasks. While million-scale datasets have enabled substantial progress on click-grounding, drag grounding (e.g. drag-and-drop, swipe, highlight) data remains an order of magnitude smaller and current models fall short on complex drag-based interactions. We introduce DragOn, a...

    arxiv.org/abs/2606.06322 · PDF

  15. 15

    LLM Self-Recognition: Steering and Retrieving Activation Signatures

    Thibaud Ardoin, Jonas Schäfer, Gerhard Wunder

    cs.AI

    Recent advances in interpretability suggest that large language models (LLMs) implicitly encode signals in their generated text that enable self-recognition of their outputs. We demonstrate that this capability is reliable, even in low-entropy scenarios, and that it can be amplified through targeted intervention. By steering the internal residual stream during generation with a random sparse vector, we create a detectable fingerprint that...

    arxiv.org/abs/2606.06315 · PDF

  16. 16

    AIS-Based Vessel Trajectory Prediction Using Memory-Augmented Neural Networks

    Wonmo Koo, Sanha Chang, Heeyoung Kim

    cs.AI

    Accurate vessel trajectory prediction is essential for safe and efficient maritime operations, enabling collision avoidance and supporting route optimization. Although memory-augmented neural networks have recently shown strong performance in pedestrian and road-vehicle trajectory prediction by selectively retrieving relevant information from an external memory, their potential for vessel trajectory prediction remains underexplored. This...

    arxiv.org/abs/2606.06311 · PDF

  17. 17

    Multi-ResNets for Subspace Preconditioning in Constrained Optimization

    Merve Karakas, Christopher J. Williams, Emmanuel O. Balogun, Sadegh Sadeghi Tabas, Christian Brown, Nikhil Rao

    cs.AI

    We propose MResOpt, a staged residual neural network architecture for constrained optimization problems. Our architecture fits within predict-complete-correct pipelines and decomposes constraint satisfaction by priority through intermediate re-completion and stage-aware losses. The framework enables domain-informed ordered constraint satisfaction which allows the network to utilize ordinal structure when present. Under an idealized...

    arxiv.org/abs/2606.06300 · PDF

  18. 18

    TRACE: A Temporal Conditional Estimation for Multimodal Time Series Foundation Models

    Ziwen Kan, Yishuo Chen, Kecheng Li, Andrew Wen, Xiaomeng Wang, Liwei Wang, Jihao Duan, Song Wang, Hongfang Liu, Tianlong Chen

    cs.AI

    Time series foundation models (TS-FMs) aim to learn generalizable temporal representations that can be adapted to a wide range of downstream tasks. In real-world multimodal settings, time series are frequently affected by temporal misalignment and partial modality missingness, where different modalities are observed at heterogeneous time scales or are partially absent. Existing approaches typically rely on naive imputation or masking...

    arxiv.org/abs/2606.06285 · PDF

  19. 19

    ToolChoiceConfusion: Causal Minimal Tool Filtering for Reliable LLM Agents

    Rahul Suresh Babu, Laxmipriya Ganesh Iyer

    cs.AI

    Large language model agents increasingly rely on external tools, but larger tool menus can reduce reliability and efficiency by increasing wrong-tool calls, premature actions, and token cost. Existing tool-selection methods often optimize semantic relevance, exposing tools whose names or descriptions match the user request. We argue that relevance is insufficient: a tool may be related to the task while still being unnecessary or premature at...

    arxiv.org/abs/2606.06284 · PDF

  20. 20

    RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention

    Yang Liu, ZhaoKai Luo, HuaYi Jin, ZhiYong Wang, RuoZhou He, BoYu Wang, Guanjie Chen, Junhao Hu

    cs.AI

    As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastructure. It limits GPU memory capacity, serving concurrency, cache reuse, and distributed scalability. Several important problems, including position-independent KV cache, prefix KV cache compression, hot/cold KV cache separation, and distributed KV cache management, all depend on how the KV cache is...

    arxiv.org/abs/2606.06256 · PDF

  21. 21

    Closing the Loop on Latent Reasoning via Test-Time Reconstruction

    Xiaopeng Yuan, Haibo Jin, Ye Yu, Peng Kuang, Lijun Yu, Yushun Dong, Haohan Wang

    cs.AI

    Recent work moves intermediate reasoning from natural-language traces into latent or cache-level representations to reduce token overhead and avoid a discrete communication bottleneck. However, this shift also removes a key advantage of textual reasoning: intermediate states are no longer inspectable, making it difficult to determine whether a latent state still preserves the constraints of the original query. As a result, latent reasoning...

    arxiv.org/abs/2606.06252 · PDF

  22. 22

    From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents

    Patrick Wilhelm, Odej Kao

    cs.AI

    Language-model agents act through repeated cycles of observation, reasoning, and action selection, making safety monitoring depend on both internal model state and environment context. We study reward-hacking monitors in ReAct-style agents acting in Gameable ALFWorld and WebShop. Agents are instrumented with activation-based reward-hack scores, token-level entropy, and decision-context features. We find that adapters fine-tuned on...

    arxiv.org/abs/2606.06223 · PDF

  23. 23

    Evaluating Agentic Configuration Repair for Computer Networks

    Rufat Asadli, Benjamin Hoffman, Ioannis Protogeros, Laurent Vanbever

    cs.AI

    Misconfigurations in computer networks remain a major source of critical Internet outages. Research is turning to Large Language Models (LLMs) to automate the complex, error-prone task of network configuration. However, even state-of-the-art models fail to resolve misconfigurations in large-scale, complex scenarios and often introduce new errors. In this work, we benchmark open- and closed-source LLMs augmented with formal network...

    arxiv.org/abs/2606.06212 · PDF

  24. 24

    Unsupervised Pattern Analysis in Japanese Veterinary Toxicology: A Regulatory-Compliant Framework for Cross-Species Risk Assessment

    Yukiko Kawakami, Mohammad Shirazi, Ryo Shimizuwa, Saito Shinoda, Alireza Mortazavi, Matsumoto Kawahara

    cs.AI · cs.LG

    Veterinary pharmacovigilance systems are essential for monitoring adverse drug events (ADEs), yet existing approaches often fail to capture region-specific toxicity patterns shaped by local biological and regulatory contexts. In Japan, these challenges are amplified by species-specific metabolic differences and reporting practices defined by the Ministry of Agriculture, Forestry, and Fisheries (MAFF). Most prior work relies on...

    arxiv.org/abs/2606.06207 · PDF

  25. 25

    Learning to replenish: A hybrid deep reinforcement learning for dynamic inventory management in the pharmaceutical supply chains

    Amandeep Kaur, Gyan Prakash

    cs.AI

    Pharmaceutical supply chains (PSCs) struggle with inventory management (IM) due to unpredictable demand patterns and variable lead times associated with restocking. This complexity is further compounded by the finite shelf lives of pharmaceutical products, which necessitate a delicate balance between adequate stock and minimal waste. These intertwined factors create a complex optimization problem that requires sophisticated inventory...

    arxiv.org/abs/2606.06201 · PDF

  26. 26

    ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

    Prathamjyot Singh, Ashima Sood, Sahil Sharma, Jasmeet Singh

    cs.AI · cs.CL

    We present ProSarc, an audio-only framework that detects sarcasm by modelling temporal prosodic incongruity, that is, the mismatch between local prosodic dynamics and the utterance-level emotional baseline. Dual encoding paths, a Global Emotion Encoder and a Temporal Prosody Encoder (BiLSTM + multi-head attention), feed a Prosodic Incongruity Analyzer that produces a scalar incongruity score for classification. Monte Carlo dropout provides...

    arxiv.org/abs/2606.06168 · PDF

  27. 27

    Where does Absolute Position come from in decoder-only Transformers?

    Valeria Ruscio, Umberto Nanni, Fabrizio Silvestri

    cs.AI · cs.CL

    RoPE-trained transformers distinguish absolute position in their attention patterns, even though RoPE encodes only relative offsets in the inner product. We trace this leakage to two architectural components, The causal mask is responsible for the first: its per-query softmax denominator depends on the absolute query position by construction. The residual stream supplies the second. Under causal attention the activation at position $0$...

    arxiv.org/abs/2606.06160 · PDF

  28. 28

    Amortizing Federated Adaptation: Hypernetwork Driven LoRA for Personalized Foundation Models

    Sunny Gupta, Shambhavi Shanker, Amit Sethi

    cs.AI

    Federated fine-tuning of foundation models using Low-Rank Adaptation (LoRA) offers a communication efficient solution for distributed learning. However, existing federated LoRA methods suffer from two fundamental limitations: (1) structural aggregation bias, where independently averaging low rank factors fails to approximate the true combined update, and (2) client side initialization lag, as clients repeatedly reinitialize LoRA parameters...

    arxiv.org/abs/2606.06154 · PDF

  29. 29

    WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation

    Shengtao Zheng, Kai Li, Weichen Zhang, Yu Meng, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang

    cs.AI

    End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation. However, existing approaches typically rely on historical observations to directly predict actions, often struggling in dense urban environments where severe occlusions and sharp turns result in drastic viewpoint transitions. We argue that the ability to "imagine" future states -- inherent in World Models -- is critical for robust decision-making under such...

    arxiv.org/abs/2606.06147 · PDF

  30. 30

    Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems

    Dianxing Shi, Junqi He, Junhao Chen, Bowen Wang, Yuta Nakashima

    cs.AI

    Self-evolving agents improve through continual self-play and self-generated learning signals, but autonomous evolution can also cause capability degradation and safety drift. Although human feedback has proven effective for static and post-trained agents, its role in self-evolving systems remains underexplored. We introduce Agent Norm Correction through Human-like Oversight and Review (ANCHOR), an LLM-based framework that simulates human...

    arxiv.org/abs/2606.06114 · PDF

  31. 31

    Step-adaptive multimodal fusion network with multi-scale cloud feature learning for ultra-short-term solar irradiance forecasting

    Jingxin Zhang Xiaoqin Wang

    cs.AI · cs.LG

    Ultra-short-term solar irradiance prediction is critical for photovoltaic system dispatch and power grid stability. Existing approaches suffer from three key shortcomings: single time-series models cannot capture the spatial dynamics of clouds under complex conditions, standard convolutions inadequately represent multi-scale cloud features, and fixed low-frequency compensation strategies fail to adapt to different prediction steps. To address...

    arxiv.org/abs/2606.06102 · PDF

  32. 32

    CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model

    Zeyang Yue, Chenfei Yan, Feifei Zhao, Haibo Tong, Mengwen Xu, Xiaozhen Wang, Erliang Lin, Yi Zeng

    cs.AI

    Whether Large Language Models (LLMs) exhibit covert psychological manipulation in complex human-AI interactions has garnered increasing safety concerns. However, existing AI safety benchmarks remain largely restricted to explicit rule compliance and static prompts, failing to capture the dynamic and covert nature of manipulative strategies in multi-turn dialogues. We introduce CogManip, a comprehensive benchmark that evaluates 15 manipulation...

    arxiv.org/abs/2606.06099 · PDF

  33. 33

    Integrating Mechanistic and Data-Driven Models for Neurological Disorders through Differentiable Programming

    Shah Pallav Dhanendrakumar, Saikat Pal, Sitikantha Roy

    cs.AI · cs.LG · math.DS · physics.med-ph

    Advances in computational modeling, neuroimaging, and artificial intelligence are revolutionizing the modeling of neurological disorders for improved diagnostics, prognosis, and treatment planning. Mechanistic models provide valuable scientific insight into the disorders, but in practice they are often simplified with assumptions or computationally expensive and slow to solve. However, while purely data driven approaches provide speed and...

    arxiv.org/abs/2606.06094 · PDF

  34. 34

    Beyond Semantic Organization: Memory as Execution State Management for Long-Horizon Agents

    Yaoqi Chen, Haibin Lai, Yuru Feng, Chuyu Han, Qianxi Zhang, Baotong Lu, Menghao Li, Xinjiang Wang, Zhirui Wang,...

    cs.AI

    LLM-based agents increasingly tackle long-horizon tasks with interdependent decisions, where each action reshapes future constraints and intermediate errors can cascade. Existing RAG and agent memory systems organize histories by semantic similarity, retrieving content-relevant entries at decision time. We argue that this design mismatches execution-state dependencies: it fragments decision trajectories and mixes valid and erroneous traces,...

    arxiv.org/abs/2606.06090 · PDF

  35. 35

    A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice

    Ranjan Mishra, Jakob Schoeffer

    cs.AI · cs.HC

    Appropriate reliance on AI advice has become a central research theme in human-AI collaboration. Existing frameworks have focused exclusively on point predictions as AI advice. However, set-valued AI advice (e.g., discrete sets or continuous intervals) is increasingly being used to communicate uncertainty and improve human decision making. In this paper, we develop the first formal framework for measuring appropriate reliance on set-valued AI...

    arxiv.org/abs/2606.06081 · PDF

  36. 36

    Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation

    Haocheng Luo, Jiahui Liu, Ruicheng Zhang, Zhizhou Zhong, Jiaqi Huang, Zunnan Xu, Quan Shi, Jun Zhou, Xiu Li

    cs.AI · cs.CV

    While vision-language models excel at general multimodal understanding, they still struggle with visual spatial planning. We attribute this to a perception-reasoning modality gap: visual planning requires models to infer latent state structures from pixels and then reason over the recovered structure to produce valid actions, whereas symbolic planning directly leverages explicit objects and constraints. This creates dual bottlenecks in visual...

    arxiv.org/abs/2606.06076 · PDF

  37. 37

    When Should Memory Stay Silent: Measuring Memory-Use Boundaries in Memory-Augmented Conversational Agents

    Lingxiang Xu, Jiaoyun Yang, Min Hu, Hongtu Chen, Ning An

    cs.AI

    Long-term memory enables language model agents to support personalized interactions, but it remains unclear when available memories warrant integration into responses. Existing memory evaluations emphasize retrieval accuracy and downstream task utility, while overlooking whether retrieved sensitive memory content is warranted in the current turn. We introduce RBI-Eval, a controlled measurement study built around a probe set that compares...

    arxiv.org/abs/2606.06055 · PDF

  38. 38

    Beyond Similarity: Trustworthy Memory Search for Personal AI Agents

    Jiawen Zhang, Kejia Chen, Jiachen Ma, Yangfan Hu, Lipeng He, Yechao Zhang, Jian Liu, Xiaohu Yang, Tianwei Zhang, Ruoxi Jia

    cs.AI

    Personal AI agents increasingly rely on long-term memory to provide persistent personalization across sessions. However, existing memory pipelines are largely driven by semantic similarity: memory data close to the current query is retrieved and injected into the model context. This creates a critical trustworthiness gap, since a semantically related memory may still be contextually inappropriate, leading to threats such as cross-domain...

    arxiv.org/abs/2606.06054 · PDF

  39. 39

    Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents

    Shuo Ji, Yibo Li, Bryan Hooi

    cs.AI · cs.IR

    Despite recent progress, LLM agents still struggle with reasoning over long interaction histories. While current memory-augmented agents rely on a static retrieve-then-reason paradigm, this rigid pipeline design prevents them from dynamically adapting memory access to intermediate evidence discovered during inference. To bridge this gap, we propose MRAgent, a framework that combines an associative memory graph with an active reconstruction...

    arxiv.org/abs/2606.06036 · PDF

  40. 40

    RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit

    Amirhossein Ghaffari, Ali Goodarzi, Huong Nguyen, Simo Hosio, Lauri Lovén, Ekaterina Gilman

    cs.AI · cs.CL · cs.LG · cs.SI

    Community-conditioned language model adaptation requires choices about data collection, community definition, and evaluation that are currently made independently in each study, making it hard to compare assumptions or reuse artifacts. We present RedditPersona, a modular framework that standardizes these choices: it collects Reddit posts and comments, profiles active users, partitions them under five grouping strategies (subreddit-based,...

    arxiv.org/abs/2606.06027 · PDF

  41. 41

    PLAN-S: Bridging Planning with Latent Style Dynamics for Autonomous Driving World Models

    Xiaoyun Qiu, Jingtao He, Yijie Chen, Yusong Huang, Haotian Wang, Yixuan Wang, Xinhu Zheng

    cs.AI · cs.RO

    Latent world models (LWMs) have strengthened end-to-end autonomous driving by forecasting compact scene dynamics for downstream planning. However, existing LWM-based planners usually generate trajectories directly from entangled latent representations. This compact latent-to-planner pathway lacks explicit modeling of risk, drivability, and diverse style preferences, making driving-style dynamics difficult to supervise, inspect, or modulate...

    arxiv.org/abs/2606.06014 · PDF

  42. 42

    Beyond Vector Similarity: A Structural Analysis of Graph-Augmented Retrieval for Industrial Knowledge Graphs

    Grama Chethan

    cs.AI

    Retrieval-Augmented Generation (RAG) fails systematically on queries requiring structural reasoning over interconnected entities. We compare eight retrieval architectures for aerospace supply chain intelligence, progressing from text retrieval through graph traversal to graph computation. Using a 46-node knowledge graph with 64 typed edges, we evaluate 23 queries across 10 intent categories and demonstrate that five query classes are...

    arxiv.org/abs/2606.06003 · PDF

  43. 43

    Framing, Judging, Steering: An Assessable Competency Model for Teach-ing Students to Reason With Generative AI

    Alexander Apartsin, Yehudit Aperstein

    cs.AI · cs.CL

    Generative AI makes answers easy and understanding hard, and uncritical use invites cognitive offloading. Schools still measure unaided performance, yet the real task is to produce good work with AI: framing an ill-defined task, judging the output, and steering the model toward a better result. This ability is rarely assessed in its own right; where measured, it collapses into one "prompting" score that cannot diagnose why AI use succeeds or...

    arxiv.org/abs/2606.05983 · PDF

  44. 44

    The Self-Correction Illusion: LLMs Correct Others but Not Themselves

    Kuan-Yen Chen, Fang-Yi Su, Jung-Hsien Chiang

    cs.AI · cs.CL

    Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical claims appear under external sources. We ask whether this asymmetry reflects a capability deficit or a role-label artifact: does an agent's willingness to correct a wrong claim depend causally on the chat-template role that carries it, rather than on the claim's content? Our setup keeps the...

    arxiv.org/abs/2606.05976 · PDF

  45. 45

    Bidirectional Search for Longest Paths: Case for Front-to-Front Heuristics

    Tzur Shubi, Ariel Felner, Solomon Eyal Shimony, Shahaf S. Shperberg

    cs.AI

    Bidirectional heuristic search can potentially reduce search effort for problems amenable to backward search. Therein, it is well-known that front-to-front heuristics can reduce the number of node expansions, but their overhead is so high that overall runtime almost always increases. We propose BiXDFBnB, a bidirectional depth-first branch-and-bound algorithm that adapts the Single-Frontier Bidirectional Search (SFBDS) framework - originally...

    arxiv.org/abs/2606.05956 · PDF

  46. 46

    Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing

    Yuxiao Ye, Haoran He, Fangyuan Kong, Xintao Wang, Pengfei Wan, Kun Gai, Ling Pan

    cs.AI

    Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models. However, most existing methods remain confined to single-turn settings, overlooking the more realistic scenario of multi-turn in-context editing, where users iteratively refine an image through a sequence of instructions. In this setting, a model must follow each new instruction while preserving accumulated session-level constraints,...

    arxiv.org/abs/2606.05950 · PDF

  47. 47

    A Pre-Registered Causal Partition of Self-Consistency Elicitation and Reward Design in RLVR

    Yuze Gao

    cs.AI · cs.LG

    Reinforcement learning from verifiable rewards (RLVR) improves reasoning even when the reward signal is spurious -- assigning credit to the group-plurality answer rather than a ground-truth verifier. Practitioners commonly interpret naive = acc(TRUE) - acc(RANDOM) as the reward-design effect. We prove this estimand is systematically biased: it conflates self-consistency elicitation (sharpening the policy toward its modal answer via majority...

    arxiv.org/abs/2606.05932 · PDF

  48. 48

    Towards World Models in Biomedical Research

    Guangyu Wang, Jingkun Yue, Siqi Zhang, Yu Liu, Xiaoyu Wang, Mingyuan Meng, Changwei Ji, Zongbo Han, Yulin Wang, Yang...

    cs.AI

    A central goal of biomedicine is to understand, predict and ultimately control the dynamic mechanisms by which biological systems respond to perturbations, disease progression and therapeutic intervention. Although foundation models and large language models have accelerated biomedical data interpretation, most current systems remain focused on static pattern recognition rather than prospective simulation of biological futures. Here we...

    arxiv.org/abs/2606.05925 · PDF

  49. 49

    Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts

    Wenbo Pan, Shujie Liu, Chin-Yew Lin, Jingying Zeng, Xianfeng Tang, Xiangyang Zhou, Yan Lu, Xiaohua Jia

    cs.AI · cs.CL · cs.LG

    AI agents rely on a harness of skills, tools, and workflows to solve complex problems. Continually improving this harness is essential for adapting to new tasks. However, existing optimization methods typically require ground-truth validation sets, yet such labeled data is difficult to acquire in practical deployment settings. To address this problem, we introduce Retrospective Harness Optimization (RHO), a self-supervised method that...

    arxiv.org/abs/2606.05922 · PDF

  50. 50

    Retry Policy Gradients in Continuous Action Spaces

    Soichiro Nishimori, Paavo Parmas

    cs.AI

    Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses. In discrete action spaces, ReMax was shown to do so by adapting to return uncertainty. In this work, we introduce pathwise derivative estimators for retry objectives and use them to extend ReMax to continuous action spaces. We...

    arxiv.org/abs/2606.05888 · PDF

  51. 51

    QCFuse: Query-Aware Cache Fusion via Compressed View for Efficient RAG Serving

    Jianxin Yan, Wangze Ni, Zhenxin Li, Jiabao Jin, Zhitao Shen, Haoyang Li, Jia Zhu, Peng Cheng, Xuemin Lin, Lei Chen, Kui Ren

    cs.AI · cs.DB

    Retrieval-augmented generation (RAG) improves large language model (LLM) answer quality by grounding generation in external evidence, but processing retrieved contexts makes the prefill stage a dominant serving cost. RAG cache fusion reduces this cost by reusing precomputed key-value (KV) caches for retrieved chunks and selectively recomputing tokens under the current prompt. Existing selectors, however, face a dilemma between quality and...

    arxiv.org/abs/2606.05875 · PDF

  52. 52

    Entropy-Based Evaluation of AI Agents: A Lightweight Framework for Measuring Behavioral Patterns

    Olasimbo Ayodeji Arigbabu

    cs.AI · cs.CV

    AI agents are commonly evaluated using task success, reward, latency, and cost. These metrics are useful, but they often miss important aspects of agent behavior: whether an agent explores too much, repeats itself too rigidly, uses tools effectively, reduces uncertainty over time, or remains robust across repeated runs. This paper proposes Entropy-Based Evaluation of AI Agents (EEA), a lightweight framework for measuring agent behavior...

    arxiv.org/abs/2606.05872 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.