cs.AI · 2026-05-27 · No. 10

Artificial Intelligence, 2026-05-27.

43 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

43 entries
  1. 01

    MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation

    Huawei Lin, Peng Li, Jie Song, Fuxin Jiang, Tieying Zhang

    cs.AI · cs.CL · cs.LG · cs.MA

    Large language model (LLM) agents rely on reusable skills to solve complex tasks. However, existing skill creation approaches treat skills as isolated and static artifacts, limiting their reusability, reliability, and long-term improvement. We propose MUSE-Autoskill Agent (Memory-Utilizing Skill Evolution), a skill-centric agent framework that lets agents continuously improve their task-solving capability by creating, reusing, and refining...

    arxiv.org/abs/2605.27366 · PDF

  2. 02

    Natural Language Query to Configuration for Retrieval Agents

    Melissa Z. Pan, Negar Arabzadeh, Mathew Jacob, Fiodar Kazhamiaka, Esha Choukse, Matei Zaharia

    cs.AI · eess.SY

    Modern retrieval agents expose many configuration choices -- LLM, retriever, number of documents, number of hops, and synthesis strategy -- each shaping both answer quality and serving cost. Today, these pipelines are typically hand-tuned once per workload, leaving substantial per-query optimization untapped. We formulate the problem: given a natural-language query and either an accuracy or a budget target, select from a predefined pipeline...

    arxiv.org/abs/2605.27361 · PDF

  3. 03

    Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

    Dongyoon Hahm, Dylan Hadfield-Menell, Kimin Lee

    cs.AI · cs.CL · cs.LG

    Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tampering, a potential vulnerability where the LLM undergoing alignment influences the preference dataset, causing RLHF to amplify undesired behaviors. This arises from core limitations of RLHF: (1) preference datasets are constructed from the LLM's own outputs, allowing it...

    arxiv.org/abs/2605.27355 · PDF

  4. 04

    2-ASP(Q) programs with weak constraints: Complexity and efficient implementation

    Andrea Cuteri, Giuseppe Mazzotta, Francesco Ricca

    cs.AI · cs.CC · cs.CL · cs.LO

    ASP(Q) extends Answer Set Programming (ASP) with Quantifiers over answer sets. In this paper we focus on the class of ASP(Q) programs with two quantifiers and weak constraints, denoted as 2-ASP(Q)^w. 2-ASP(Q)^w is a practically relevant fragment of ASP(Q) that is expressive enough to capture optimization problems up to the class Delta_3^P. On the theoretical side, we provide a complete complexity characterization of the main computational...

    arxiv.org/abs/2605.27338 · PDF

  5. 05

    Maat: The Agentic Legal Research Assistant for Competition Protection

    Basant Mounir, Farida Madkour, Amira Abdelaziz, Asmaa Sami

    cs.AI

    Competition law experts conducting legal research must review extensive volumes of cases, decisions, and judicial reports to identify precedents and assess key elements in competition and merger cases. Although general research assistants such as Claude and ChatGPT and legal assistants such as SaulLM-7B and LegalGPT are increasingly used to assist legal research, they remain inadequate for competition law analysis: they lack specialized...

    arxiv.org/abs/2605.27331 · PDF

  6. 06

    Modeling Agentic Technical Debt and Stochastic Tax: A Standalone Framework for Measurement, Simulation, and Dashboarding

    Muhammad Zia Hydari, Raja Iqbal, Narayan Ramasubbu

    cs.AI · cs.CY · econ.GN

    Agentic AI systems combine probabilistic reasoning with delegated action through tools, context, memory, orchestration, and external workflow integration. This note develops a formal and managerially usable model that distinguishes Agentic Technical Debt from Stochastic Tax. Agentic Technical Debt is a stock of accumulated design and governance liability. Stochastic Tax is a recurring flow of operating burden that arises when stochastic...

    arxiv.org/abs/2605.27320 · PDF

  7. 07

    SIA: Self Improving AI with Harness & Weight Updates

    Prannay Hebbar, Yogendra Manawat, Samuel Verboomen, Alesia Ivanova, Selvam Palanimalai, Kunal Bhatia, Vignesh Baskaran

    cs.AI · cs.CL

    Humans are the bottleneck in building and improving AI. Both the models and the agents that wrap them are written, tuned, and corrected by people. The long-horizon goal of an AI that can figure out how to improve itself remains open. Two largely disjoint research lines attack this bottleneck. The harness-update school has a meta-agent rewrite the scaffold of a task-specific agent (its tools, prompts, retry logic, and search procedure) while...

    arxiv.org/abs/2605.27276 · PDF

  8. 08

    Gumbel Machine: Counterfactual Student Writing Generation via Gumbel Noise Steering

    Hunter McNichols, Alexander Scarlatos, Mihai Dascalu, Danielle McNamara, Andrew Lan

    cs.AI · cs.CL

    An effective method of teaching across disciplines is to provide examples of high-quality work. However, an example may be significantly different from a student's current work, making it challenging for them to emulate. An ideal learning demonstration is a counterfactual version of the student work, an improved version that is still similar to their own. Existing automated approaches for counterfactual text generation using Large Language...

    arxiv.org/abs/2605.27249 · PDF

  9. 09

    Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments

    Yuxin Chen, Xiaodong Cai, Junfeng Fang, Zhuowen Han, Yu Wang, Yaorui Shi, Yi Zhang, Qi Gu, Xunliang Cai, Xiang Wang,...

    cs.AI

    Recent advances in large language models (LLMs) have facilitated the widespread deployment of LLMs as interactive agents capable of reasoning, planning, and tool use. Despite strong performance on existing benchmarks, such agents often exhibit notable degradation when deployed in real-world settings, where environments are inherently stochastic and imperfect. We argue that this discrepancy arises from a fundamental mismatch between idealized...

    arxiv.org/abs/2605.27209 · PDF

  10. 10

    The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypothesis Generation?

    Shashwat Sourav, Viktoriia Baibakova, Sanjay Das, Ran Elgedawy, Maria Mahbub, Emily Herron, Tirthankar Ghosal

    cs.AI

    Knowledge graphs (KGs) can provide structured scientific context to language models, but it remains unclear which graph facts actually shape the generated hypotheses. We study KG-guided hypothesis generation for battery materials across Mistral-7B, Llama-3.1-70B, and Gemini 2.5 Flash. We perturb local KGs by varying density, ontology richness, topology, and control structure, and evaluate outputs with both provided-graph and fixed-reference...

    arxiv.org/abs/2605.27176 · PDF

  11. 11

    Query Symbolically or Retrieve Semantically? A Dataset and Method for Semi-Structured Question Answering

    Mateusz Czyżnikiewicz, Ryszard Tuora, Adam Kozakiewicz, Tomasz Ziętkiewicz, Mateusz Galiński, Michał Godziszewski,...

    cs.AI

    Retrieval-Augmented Generation (RAG) systems for question answering typically retrieve evidence by semantic similarity between the query and document chunks. While effective for unstructured text, this approach is less reliable on semi-structured corpora where answering may require exact filtering, aggregation, or exhaustive retrieval over structured attributes across multiple documents. Symbolic approaches support such operations, but they...

    arxiv.org/abs/2605.27164 · PDF

  12. 12

    Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs

    Zhe Yu, Wenpeng Xing, Chen Ye, Xuyang Teng, Bo Yang, Changting Lin, Meng Han

    cs.AI

    Retrieval-augmented LLMs are deployed for tasks where evidence quality determines action safety, yet evaluation protocols assume that single-turn robustness predicts robustness when evidence accumulates across turns. We show this assumption is fundamentally incorrect. Models exhibit a monitoring-control gap: they readily acknowledge contradictory evidence, yet this awareness fails to constrain their final recommendations - detecting epistemic...

    arxiv.org/abs/2605.27157 · PDF

  13. 13

    VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

    Yuxin Chen, Yi Zhang, Zhengzhou Cai, Yaorui Shi, Zhiyuan Yao, Chenhang Cui, Jingnan Zheng, Yaqi Huo, Xi Su, Qi Gu,...

    cs.AI

    Large language models (LLMs) have evolved into interactive agents that collaborate with users in real-world tasks. Effective collaboration in such settings increasingly depends on understanding the user beyond what is explicitly stated, as user intent is often reflected in fragmented daily interactions and requires both personalized modeling and proactive interaction. However, existing agent benchmarks primarily evaluate reasoning and tool...

    arxiv.org/abs/2605.27141 · PDF

  14. 14

    StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

    Yanfei Zhang, Xu Lin, Chenglin Wu

    cs.AI

    Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hinges on a few local decisions. Existing online policy distillation (OPD) provides denser token-level supervision, but typically treats heterogeneous agent trajectories as monolithic strings rather than causal interaction units. We present StepOPSD, a post-rollout preference self-distillation...

    arxiv.org/abs/2605.27140 · PDF

  15. 15

    ICCU: In-Context Continual Unlearning via Pattern-Induced Refusal Rules

    Ruihao Pan, Suhang Wang

    cs.AI

    Machine unlearning aims to remove the influence of specific data from trained language models. In real-world deployments, unlearning requests often arrive sequentially, which challenges existing fine-tuning-based methods: fine-tuning each request is costly, accumulates utility loss, and may cause cross-request interference. To address these issues, we propose ICCU (In-Context Continual Unlearning), an in-context continual unlearning framework...

    arxiv.org/abs/2605.27138 · PDF

  16. 16

    Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation

    Heng Qu, Yike Liu, Renren Jin, Wenzong Zhang, Pengzhi Gao, Wei Liu, Jian Luan

    cs.AI

    Vision-Language Models (VLMs) have shown rapid progress in mobile GUI navigation. This paper presents a systematic study of data scaling, benchmarking, and reasoning for VLM-based agents in this domain. To facilitate rigorous evaluation, we introduce HyperTrack, a large-scale dataset with over 16000 real-world tasks across more than 650 Chinese mobile applications, along with GUIEvalKit, an open-source toolkit for unified benchmarking of VLMs...

    arxiv.org/abs/2605.27134 · PDF

  17. 17

    Position: AI Safety Requires Effective Controllability

    Yige Li, Yunhao Feng, Jun Sun

    cs.AI

    AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing has improved the behavior of modern language models, but aligned behavior does not by itself guarantee that a deployed agent can be stopped, overridden, or constrained once it operates in open-ended, interactive, and tool-using environments. A system may be safe in expectation and still fail to...

    arxiv.org/abs/2605.27117 · PDF

  18. 18

    Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation

    Tianlei Chen, Jiao Ou, Ziyuan Liu, Ruiming Tang, Jian Liang, Han Li

    cs.AI

    Domain specialization can improve LLM behavior in vertical domains, but often weakens the general capabilities inherited from the original model. Recent Multi-Teacher On-Policy Distillation (MOPD) pipelines recover model capabilities by supervising student-generated trajectories with teacher feedback, but typically assume teacher-aligned prompt coverage, requiring prompts to match the teachers' training distributions. This assumption is...

    arxiv.org/abs/2605.27115 · PDF

  19. 19

    Can Broad Biomedical Knowledge be Contextualized into Scenario-Grounded Propositions?

    Qingyuan Zeng, Ziyang Chen, Pengxiang Cai, Zixin Guan, Anglin Liu, Lang Qin, Xinyao Lai, Jintai Chen

    cs.AI

    Biomedical discovery often requires connecting broad biomedical knowledge with specific experimental or clinical data. Background knowledge suggests relevant mechanisms but is usually too general to map directly onto dataset variables, while data-driven patterns can be dataset-specific and hard to interpret mechanistically. We study this missing link as knowledge contextualization: transforming broad biomedical knowledge into...

    arxiv.org/abs/2605.27082 · PDF

  20. 20

    Traceable Knowledge Graph Reasoning Enables LLM-Assisted Decision Support for Industrial VOCs in the Steel Industry

    Changqing Su, Yu Ding, Zuhong Lin, Hongyu Liu, Xi He, Zheng Zeng, Liqing Li

    cs.AI

    Key knowledge for steel-industry volatile organic compounds (VOCs) governance is scattered across unstructured scientific literature, making it difficult to integrate process, pollutant, and control-technology evidence and increasing the risk of hallucination when general large language models (LLMs) answer low-frequency industrial questions. Here we developed Chat-ISV, a knowledge graph (KG) enhanced multi-agent Q&A system that parses a...

    arxiv.org/abs/2605.27071 · PDF

  21. 21

    BatteryMFormer: Multi-level Learning for Battery Degradation Trajectory Forecasting

    Ruifeng Tan, Jintao Dong, Weixiang Hong, Jia Li, Jiaqiang Huang, Tong-Yi Zhang

    cs.AI

    Early battery degradation trajectory forecasting (BDTF), which predicts the full-life state-of-health trajectory from early operational data, is critical for battery optimization, manufacturing, and deployment. Battery degradation data exhibit two key characteristics. First, degradation data present a multi-level structure, including regularities shared within aging conditions and trajectory patterns shared across batteries. Second,...

    arxiv.org/abs/2605.27044 · PDF

  22. 22

    Boosting Knowledge Graph Foundation Models via Enhanced Negative Sampling

    Yinan Liu, Wenjin Xu, Zhiyuan Zha, Xiaochun Yang, Bin Wang

    cs.AI

    Knowledge graphs (KGs) have become the core backbone of numerous downstream tasks such as question answering and recommender systems. However, despite all this, KGs are often very incomplete. To perform zero-shot knowledge graph completion in unseen KGs, which have different relational vocabularies from those used for pre-training, KG foundation models (KGFMs) receive a wide range of attention. Existing KGFMs often perform training using...

    arxiv.org/abs/2605.27023 · PDF

  23. 23

    ORCA: An End-to-End Interactive Copilot for Optimized Root Cause Analysis

    Phi Nguyen Xuan, Nicholas Tagliapietra, Lavdim Halilaj, Kristian Kersting, Juergen Luettin

    cs.AI

    Causal analysis is a crucial task in many domains, including manufacturing, social science, and medicine. However, despite recent progress, the conceptual and methodological complexity of causal methods makes them largely inaccessible to domain experts. This gap prevents experts from leveraging these advances and hinders researchers who lack access to real-world data for validation. To bridge this divide, we introduce ORCA, a copilot for...

    arxiv.org/abs/2605.27022 · PDF

  24. 24

    Generating Robust Portfolios of Optimization Models using Large Language Models

    Eleni Straitouri, Cheol Woo Kim, Milind Tambe

    cs.AI

    Mathematical optimization is a powerful tool for structured decision-making across domains such as resource allocation and planning. Formulating optimization models faithful to reality, though, remains a significant bottleneck as it typically demands both domain expertise and optimization knowledge that are often scarce. Recent advances in large language models (LLMs) promise to bridge this gap, enabling the generation of candidate...

    arxiv.org/abs/2605.27013 · PDF

  25. 25

    LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation

    Samy Haffoudhi, Nikola Dobričić, Fabian Suchanek, Nils Holzenberger

    cs.AI · cs.CL

    Entity linking is a key component of many downstream NLP systems, yet existing approaches are often tied to the specific target knowledge bases and domains, limiting their real world application. In this paper, we extend LELA, a modular and domain-agnostic LLM-based entity disambiguation method, into a practical Python library that integrates zero-shot Named Entity Recognition (NER) -thereby providing a complete end-toend pipeline for...

    arxiv.org/abs/2605.26956 · PDF

  26. 26

    Neuro-Symbolic Verification of LLM Outputs for Data-Sensitive Domains (extended preprint)

    Paul Sigloch, Christoph Benzmüller

    cs.AI · cs.LO · cs.SE

    LLMs deployed in high-stakes domains face fundamental reliability challenges: hallucinations, inconsistencies, and privacy vulnerabilities introduce unacceptable risks where errors carry legal, financial, or safety consequences. This paper presents a hybrid verification architecture combining formal symbolic methods with neural semantic analysis to provide complementary guarantees for LLM-generated content. This architecture employs logical...

    arxiv.org/abs/2605.26942 · PDF

  27. 27

    Developing a Totally Unimodular Linear Program for Optimal Conformance Checking: When and Why It Complements A*

    Izack Cohen

    cs.AI · math.OC

    Alignment-based conformance checking is the state-of-the-art approach for comparing observed process executions with normative process models. The standard exact solution relies on an A*-based heuristic search, which can exhibit exponential runtime in the presence of long traces or substantial deviations. This paper introduces a reformulation of alignment-based conformance checking as a totally unimodular linear program (LP) defined on the...

    arxiv.org/abs/2605.26938 · PDF

  28. 28

    From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation

    Youssef Al Mouatamid, Marie Bonnin, Jihad Zahir

    cs.AI

    Computing legal indicators from normative texts is a key task in legal monitoring and policy evaluation, but presents significant challenges due to the complexity, scale, and interpretive nature of legal language, as well as the variability in available document quality. Existing natural language processing techniques and generative models can assist in legal analysis, but often suffer from high risk of hallucinations and lack the...

    arxiv.org/abs/2605.26926 · PDF

  29. 29

    TADDLE: A Tool-Augmented Agent for Detecting Deficient LLM-Generated Peer Reviews

    Hanqi Duan, Xiang Li

    cs.AI

    LLM-generated peer reviews are increasingly common at major venues, yet their deficiencies are hard to detect because they are uniformly fluent and well-structured. Existing work either classifies authorship without judging quality, or scores quality with features designed for human-written reviews; no prior system detects deficiencies in LLM-generated reviews at the level of individual defect types. To bridge the gap, we introduce TADDLE, a...

    arxiv.org/abs/2605.26911 · PDF

  30. 30

    On the Detection of Commutative Factors in Factor Graphs: Necessary and Sufficient Conditions

    Malte Luttermann, Ralf Möller, Marcel Gehrke

    cs.AI · cs.DS · cs.LG

    Exploiting the indistinguishability of objects in a probabilistic graphical model such as a factor graph is key to lifted probabilistic inference algorithms and allows for tractable probabilistic inference problems with respect to domain sizes. A central building block for the exploitation of indistinguishable objects in factor graphs is the identification of commutative factors, i.e., factors whose output values are invariant under...

    arxiv.org/abs/2605.26908 · PDF

  31. 31

    Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation

    Lulu Zheng, Wenjin Yang, Xiangwen Zhang, Rong Yin, Yulan Hu, Zheng Pan, Xin Li

    cs.AI

    Multi-stakeholder tasks require one output to satisfy users with conflicting preferences. Holistic LLM judges conflate utility estimation and utility aggregation, yielding unstable implicit weights. We show empirically and theoretically that this aggregation-specific \emph{weighting noise} can create large score shifts when stakeholder satisfaction is dispersed; in our experiments, these weight-induced shifts also increase with stakeholder...

    arxiv.org/abs/2605.26878 · PDF

  32. 32

    Helicase: Uncertainty-Guided Supply Chain Knowledge Graph Construction with Autonomous Multi-Agent LLMs

    Yunbo Long, Haolang Zhao, Ge Zheng, Alexandra Brintrup

    cs.AI

    LLM-based multi-agent systems have been widely adopted for knowledge retrieval and report generation, synthesizing known information through web search and textual reasoning. However, many critical information tasks in supply chains are not simple one-shot queries: they are structural inference problems requiring multi-hop reasoning across complex, fragmented web resources. Questions such as \textit{``Which Tesla components use lithium from...

    arxiv.org/abs/2605.26835 · PDF

  33. 33

    What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global Derivation

    Xiang Wang, Wei Wei

    cs.AI

    Chain-of-thought (CoT) prompting reliably improves language-model accuracy, but which properties of a rationale text drive the improvement is poorly understood. Prior work has largely studied generation-time behavior. We instead ask a probe-time question: given a fixed rationale in context, what in that text changes the answer? We identify two complementary sources of the gain. First, even a globally word-shuffled rationale substantially...

    arxiv.org/abs/2605.26795 · PDF

  34. 34

    Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning

    Zhe Yu, Wenpeng Xing, Yunzhao Wei, Jie Chen, Hongzhi Wang, Xuyang Teng, Meng Han

    cs.AI

    Post-training is routinely evaluated through aggregate benchmark scores that treat multi-hop reasoning as a single capability -- as if a model that answers more questions correctly must be better at assembling facts. We show that this assumption can be misleading: recipes with statistically indistinguishable atomic knowledge produce composition behaviour separated by over 40 percentage points, a phenomenon we call composition collapse: the...

    arxiv.org/abs/2605.26789 · PDF

  35. 35

    LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?

    Xiaohan Wang, Mingze Yin, Yilin Zhao, Gang Liu, Dian Li

    cs.AI · cs.MM

    Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors. Realizing this potential requires models to navigate real-world examinations effectively, yet most existing benchmarks fail to capture the complexity of authentic testing environments. Specifically, most datasets are static, prone to data contamination, and are often confined to restricted...

    arxiv.org/abs/2605.26781 · PDF

  36. 36

    The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context

    Zhe Yu, Wenpeng Xing, Yunzhao Wei, Bo Yang, Chen Ye, Gaolei Li, Meng Han

    cs.AI

    Retrieval-augmented generation promises to ground language model outputs in external evidence, yet the field has no reliable way to verify whether retrieved context actually governs generation -- a prerequisite for any high-stakes deployment. The standard assumption, that context-consistent output implies context-governed output, breaks when the retrieved document overlaps with the model's pretraining data: the model can produce...

    arxiv.org/abs/2605.26778 · PDF

  37. 37

    Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

    Kia-Jüng Yang, Dominik Meier, Jiachen Zhao, Terry Ruas, Bela Gipp

    cs.AI

    Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may complicate control mechanisms such as refusal. Unlike instruction-tuned LLMs, where refusal is mediated by a single directional subspace, refusal in large reasoning models (LRMs) additionally depends on the CoT. In DeepSeek-R1-Distill-LLaMA-8B, activation steering reverses refusal in only 39% of...

    arxiv.org/abs/2605.26772 · PDF

  38. 38

    A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

    Heriberto Cuayahuitl, Grace Jang

    cs.AI

    Large Language Models (LLMs) have brought huge improvements to Artificial Intelligence (AI), which can be applied to general-purpose tasks. However, their application to textual or spoken medical consultations is still an open research problem. This paper proposes MeDial-Speech, a novel speech dataset for training and evaluating Med-AIs that can carry out consultations with patients. It was collected in realistic environments from...

    arxiv.org/abs/2605.26747 · PDF

  39. 39

    It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers

    Yong-eun Cho

    cs.AI · cs.CL

    A prevalent assumption in LLM agent deployment holds that more structured harnesses universally improve reliability, and that higher-capability models need proportionally less structural guidance -- together implying a monotone inverse relationship between model capability tier and optimal harness complexity. We test this hypothesis through a controlled 432-run experiment crossing six models across four capability tiers with three harness...

    arxiv.org/abs/2605.26731 · PDF

  40. 40

    Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation

    Yee Hin Chong, Jiaming Wu, Youhui Zhang, Peng Qu

    cs.AI

    Large language models (LLMs) have shown strong empirical gains as self-evolving agents for CUDA kernel generation, driven by feedback-conditioned planning across generations. However, how planning decisions attribute and combine heterogeneous feedback signals remains opaque. Standard end-to-end ablations fail to resolve this question, as iterative planning amplifies early perturbations and conflates feedback effects with trajectory-dependent...

    arxiv.org/abs/2605.26720 · PDF

  41. 41

    Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents

    Yunhui Gan, Tan Pan, Kaiyu Guo, Limei Han, Weimiao Yu, Guangnan Ye, Chen Jiang, Yuan Cheng

    cs.AI

    Medical AI agents increasingly use external tools for diagnosis, treatment recommendation, and evidence retrieval, yet most existing approaches assume that task-appropriate tools are reliable within their intended scope. This assumption is fragile in real clinical settings, where even relevant tools may fail on challenging instances and lead to unsafe downstream decisions. To address this issue, we study medical tool use under imperfect-tool...

    arxiv.org/abs/2605.26691 · PDF

  42. 42

    MemFail: Stress-Testing Failure Modes of LLM Memory Systems

    Ishir Garg, Neel Kolhe, Dawn Song, Xuandong Zhao

    cs.AI · cs.LG

    Large language model (LLM) agents increasingly rely on external memory systems to remain consistent across long-horizon interactions, but little empirical work has been done to understand the specific failure modes and design choices that these systems present. Existing benchmarks report aggregate question-answering accuracy and treat memory systems as black boxes, making it impossible to attribute an incorrect answer to a particular failure...

    arxiv.org/abs/2605.26667 · PDF

  43. 43

    Completion vs Optimality: Policy Gradient in Long-Horizon Cumulative-Damage Problems

    Wolfgang Maass, Sabine Janzen

    cs.AI

    Long-horizon decision problems with cumulative damage couple locally attractive actions to globally adverse outcomes. We identify two orthogonal failure modes for policy-gradient methods on this class and propose a decomposition that separates them: \emph{completion} (reaching the terminal horizon rather than exiting via an implicit terminal constraint) and \emph{optimality} (matching the dynamic-programming reference given completion). Under...

    arxiv.org/abs/2605.26657 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.