cs.AI · 2026-05-28 · No. 11

Artificial Intelligence, 2026-05-28.

66 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

66 entries
  1. 01

    Calibrating Conservatism for Scalable Oversight

    William Overman, Mohsen Bayati

    cs.AI

    Agentic AI systems capable of autonomous planning and extended environmental interaction pose a fundamental control problem: how can humans maintain meaningful oversight of systems that may exceed their own capabilities? Existing approaches to scalable oversight rely on complex assumptions, remain largely heuristic, or lack practical methods for sequential settings with statistical guarantees. We introduce Calibrated Collective Oversight...

    arxiv.org/abs/2605.28807 · PDF

  2. 02

    CaMBRAIN: Real-time, Continuous EEG Inference with Causal State Space Models

    Abhilash Durgam, Nyle Siddiqui, Jeffrey A. Chan-Santiago, Qiushi Fu, Elakkat D. Gireesh, Mubarak Shah

    cs.AI · cs.HC · cs.LG

    Electroencephalography (EEG) is a critical, non-invasive method to monitor electrical brain activity. EEGs can span anywhere from a couple seconds to multiple hours, posing a major hurdle for existing deep learning methods due to two major factors: (1) existing EEG models are predominantly built upon the attention mechanism, incurring quadratic scaling as the sequence length increases, and (2) raw EEG signals must be processed in a...

    arxiv.org/abs/2605.28792 · PDF

  3. 03

    SwarmHarness: Skill-Based Task Routing via Decentralized Incentive-Aligned AI Agent Networks

    Edwin Jose

    cs.AI · cs.DC · cs.MA

    Vast quantities of compute (GPU cycles on personal workstations, idle inference servers, and edge devices between jobs) go unused because no incentive-aligned protocol exists for their owners to share them safely and profitably. Existing approaches either require a trusted central coordinator (cloud marketplaces), demand heavy blockchain infrastructure (Golem, BrokerChain), or lack an incentive layer entirely (BOINC, Petals). We propose...

    arxiv.org/abs/2605.28764 · PDF

  4. 04

    CubePart: An Open-Vocabulary Part-Controllable 3D Generator

    Yiheng Zhu, Kangle Deng, Jean-Philippe Fauconnier, Inaki Navarro, Daiqing Li, Ava Pun, Yinan Zhang, Peiye Zhuang,...

    cs.AI

    Interactive 3D assets used in games and simulation are typically decomposed into specific semantic parts to support animation, physics, and scripted behaviors, yet most generative 3D models produce either monolithic meshes or arbitrary part decompositions that cannot be aligned with application-specific requirements. We present CubePart, a generative framework for open-vocabulary, part-controllable 3D mesh generation that exposes part...

    arxiv.org/abs/2605.28763 · PDF

  5. 05

    CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning

    Linas Nasvytis, Simon Jerome Han, Ben Prystawski, Satchel Grant, Noah D. Goodman, Judith E. Fan

    cs.AI

    Language models can use verifiable rewards to improve at a wide variety of reasoning tasks. However, both parametric (e.g. RLVR) and non-parametric (e.g. prompt optimization) approaches to doing so typically require hundreds of training samples and thousands of model rollouts, making them expensive in the best case and intractable in the worst. To address this challenge, we introduce Contrastive Reflection (CORE), a non-parametric learning...

    arxiv.org/abs/2605.28742 · PDF

  6. 06

    Utility-Aware Multimodal Contrastive Learning for Product Image Generation

    Xiaohang Feng, Yiling Xie

    cs.AI

    Product images strongly influence consumer decision-making in online marketplaces. Empowered by multimodal contrastive learning, generative AI can output images that closely align with text prompts. Yet existing generative AI models do not directly optimize marketplace performance. This is a critical gap, since semantic alignment alone does not guarantee that an image will sell. To address this limitation, we propose a \textit{utility-aware...

    arxiv.org/abs/2605.28733 · PDF

  7. 07

    AlphaTransit: Learning to Design City-scale Transit Routes

    Bibek Poudel, Sai Swaminathan, Weizi Li

    cs.AI

    Designing a transit network requires many sequential route extension decisions, but their quality is often visible only after the full network is assembled. This delayed-feedback challenge lies at the heart of the Transit Route Network Design Problem (TRNDP), where route interactions can be deceptive: an extension that appears useful locally can create transfer bottlenecks, produce redundant overlap, or reduce overall throughput. To guide...

    arxiv.org/abs/2605.28730 · PDF

  8. 08

    Multi-Adapter Representation Interventions via Energy Calibration

    Manjiang Yu, Hongji Li, Junwei Chen, Xue Li, Priyanka Singh, Yang Cao, Lijie Hu

    cs.AI

    Representation intervention has emerged as a promising paradigm for aligning large language models toward desired behaviors without modifying model weights. Existing methods typically apply a fixed intervention uniformly across all inputs. However, we find that the appropriate intervention direction and strength vary substantially across samples, and such indiscriminate intervention leads to degradation of general capabilities on benign...

    arxiv.org/abs/2605.28722 · PDF

  9. 09

    LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

    HuiMing Fan, Xiao Wang, Zheng Chu, Qianyu Wang, Zhuoyao Wang, Ming Liu, Bing Qin, XingYu

    cs.AI

    Are LLM-based search agents genuinely searching, or using the web to verify what they already know? We study this question on BrowseComp with three diagnostics. Our analysis reveals Intrinsic Knowledge Dependence (IKD): even with tool access, agents often rely on intrinsic knowledge -- information encoded in the model before retrieval -- rather than on external evidence. Agents answer up to 44.5% of BrowseComp questions without tools,...

    arxiv.org/abs/2605.28721 · PDF

  10. 10

    OpenURMA: A Clean-Room Open Implementation of the Unified Bus Protocol

    Bojie Li

    cs.AI · cs.AR · cs.NI

    Modern datacenter RDMA is bottlenecked at the network interface, not the wire. A NIC running RoCE or InfiniBand holds per-connection state for every (application, remote-endpoint) pair - hundreds of megabytes at 1024-application fanout - and pays a four-traversal PCIe round trip on a 64-byte operation, inflating latency an order of magnitude beyond the wire. Both follow from the Queue Pair over PCIe abstraction RDMA inherits from InfiniBand....

    arxiv.org/abs/2605.28717 · PDF

  11. 11

    Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor

    Guoxin Ma, Yibing Liu, Chengzhengxu Li, Yu Liang, Yan Wang, Yueyang Zhang, Kecheng Chen, Zhaohan Zhang, Zhiyuan Sun,...

    cs.AI

    Context compression aims to shorten long context inputs with minimal information loss for LLM inference acceleration. While existing methods have shown promise, they typically rely on complex compression modules or compression-specific training, leaving the intrinsic capabilities of LLMs underexplored. In contrast, this work reveals that a thinking model itself can naturally compress long contexts by organizing task-relevant information. We...

    arxiv.org/abs/2605.28713 · PDF

  12. 12

    Beyond Binary Moral Judgment: Modeling Ethical Pluralism in AI

    Aisha Aijaz, Rahul Goel, Arnav Batra, Raghava Mutharaju

    cs.AI · cs.LG

    Critical decision-making in socially consequential spaces is increasingly involving AI systems at varying capacities. Yet, despite the ubiquity of autonomous systems, most approaches to handling autonomous moral decision-making resort to scalar or binary judgments. These methods are insufficient for acceptable moral reasoning, as they provide little explanation, leaving out imperative contextual and theoretical information that must be...

    arxiv.org/abs/2605.28707 · PDF

  13. 13

    The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

    Dominika Agnieszka Długosz, Arlindo Oliveira, Natalia Díaz Rodríguez

    cs.AI · cs.CL

    The GSM-Symbolic benchmark (Mirzadeh et al., 2025) reported consistent performance drops across 25 Large Language Models (LLMs) when tested on template-generated variants of GSM8K problems, concluding that the models lack genuine reasoning capabilities. We argue that this conclusion rests on shaky statistical ground. Re-evaluating 20 open-weight models using Generalised Linear Mixed Models with per-question random effects, we find that only...

    arxiv.org/abs/2605.28700 · PDF

  14. 14

    TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning

    Chusen Li, Zhou Liu, Shuigeng Zhou, Wentao Zhang

    cs.AI

    Large language models increasingly rely on either reinforcement learning or multi-agent prompting to improve reasoning, yet these two paradigms remain difficult to combine. Directly applying single-agent reinforcement learning to multi-turn multi-agent systems faces following dilemmas: i) Sparse rewards, role-level free-riding and excessive training overhead. ii) Agents only imitate to collaborate. iii) Fixed collaboration protocol falls into...

    arxiv.org/abs/2605.28699 · PDF

  15. 15

    VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora

    Yuting Xu, Jiayi Tian, Jian Liang, Xin Xiong, Hang Zhang, Mu Xu, Xiao-Yu Zhang

    cs.AI

    Existing benchmarks have laid the foundation for travel planning agents by establishing API-centric paradigms. However, as the capabilities of Autonomous Agents continue to advance, their evaluation must evolve beyond simple tool execution toward handling the inherent complexities of the open web. Current benchmarks bypass core cognitive hurdles: they fail to account for information noise, ignore multi-source factual contradictions, and...

    arxiv.org/abs/2605.28683 · PDF

  16. 16

    DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution

    Yunhai Hu, Zining Liu, Xiangyang Yin, Tianhua Xia, Bo Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang

    cs.AI

    Speculative reasoning has recently been proposed as a means to accelerate reasoning-intensive generation in large multimodal models, but its effectiveness is often constrained by misalignment between speculative drafts and target-verified reasoning. In this work, we introduce DREAM-R, a framework that substantially improves the performance of speculative reasoning. At its core, DREAM-R employs Speculative Alignment Policy Optimization (SAPO),...

    arxiv.org/abs/2605.28678 · PDF

  17. 17

    An LLM-Based Assistance System for Intuitive and Flexible Capability-Based Planning

    Luis Miguel Vieira da Silva, Nicolas König, Felix Gehlhoff

    cs.AI

    In modern industry, dynamic environments and the complexity of modular and reconfigurable resources require automated planning of process sequences. Capability-based planning approaches address this by automatically generating plans from semantic knowledge models that describe resource functions in a machine-interpretable form. Their practical use, however, remains limited: solver feedback, especially in the case of unsatisfiability, is...

    arxiv.org/abs/2605.28666 · PDF

  18. 18

    AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation

    Shanghua Gao, Ada Fang, Marinka Zitnik

    cs.AI

    Scientific research proceeds through iterative cycles of hypothesis generation, experiment design, execution, and revision. AI agents can automate parts of this process, but existing approaches typically follow a single research trajectory or coordinate through a central planner with fixed objectives. As a result, they struggle to sustain parallel exploration, adapt as experimental evidence changes, or preserve knowledge of failed directions...

    arxiv.org/abs/2605.28655 · PDF

  19. 19

    The Ethics of LLM Sandbox and Persona Dynamics

    Tim Gebbie, Stewart Gebbie

    cs.AI · cs.CY · q-fin.RM

    It is well known that LLM guardrails and trained persona dynamics can produce a reality gap: the distance between the world a LLM is permitted or shaped to describe, and the world in which users must act. Here we argue that actively generating reality gaps is in fact unethical because it knowingly shifts epistemic risk back to the uninformed user -- this is reality laundering. This can potentially cause harm when operationalised at scale. The...

    arxiv.org/abs/2605.28647 · PDF

  20. 20

    Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation

    Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Ming Liu, Bing Qin, Yang Xiang

    cs.AI

    Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment paradigms face critical challenges: pure on-device models suffer from resource constraints, while centralized cloud systems incur severe privacy risks and bandwidth bottlenecks by transmitting raw voice data. Furthermore, most models exhibit English-centric biases, restricting many-to-many...

    arxiv.org/abs/2605.28642 · PDF

  21. 21

    LACUNA: Safe Agents as Recursive Program Holes

    Yaoyu Zhao, Yichen Xu, Oliver Bračevac, Cao Nguyen Pham, Frank Zhengqing Wu, Martin Odersky

    cs.AI · cs.PL

    LLM agents increasingly act by writing code, yet a split persists between the runtime that drives the agent and the code the model writes. The runtime owns the loop, context, and control flow, and the model has little say over any of them. Letting model-written code shape the runtime itself would make agents more expressive, but it would also sharpen safety problems. A model can be diverted by a prompt injection, call the wrong tool, or fail...

    arxiv.org/abs/2605.28617 · PDF

  22. 22

    Adaptive Multimodal Agents-Based Framework for Automatic Workflow Execution

    Susanna Cifani, Mario Luca Bernardi, Marta Cimitile

    cs.AI · cs.CL

    Modern information systems require autonomous agents capable of navigating complex workflows, yet current methodologies often struggle with the transition from structured metadata parsing to general environmental perception. While the integration of MLLMs has enabled agents to interact directly with GUIs, existing approaches typically treat task sequences as discrete, linear episodes. This fragmentation prevents agents from capturing the...

    arxiv.org/abs/2605.28607 · PDF

  23. 23

    Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability

    Leizhen Zhang, Shuhan Chen, Sheng Chen

    cs.AI · cs.CL · cs.LO

    Large language models (LLMs) are increasingly used for tasks that implicitly reduce to Boolean satisfiability (SAT), yet their reasoning ability on SAT remains unclear. We present a systematic study of LLMs on 2-SAT and 3-SAT, together with two canonical reductions, Vertex Cover and discrete 3D packing, to probe representation-invariant reasoning. We first evaluate models using conventional metrics, including accuracy, precision, recall, and...

    arxiv.org/abs/2605.28602 · PDF

  24. 24

    MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation

    Xiaoyu Dong, Zhi Li, Xiao-Ming Wu

    cs.AI

    Large language models (LLMs) have recently advanced text-driven 3D generation, yet Text-to-CAD remains far from supporting industrial product design. Existing benchmarks focus primarily on generating single-part CAD models and evaluate them using geometric similarity metrics that fail to capture functionality, manufacturability, and assemblability. To address this gap, we introduce MUSE, a Text-to-CAD benchmark focused on complex, editable...

    arxiv.org/abs/2605.28579 · PDF

  25. 25

    Continual Model Routing in Evolving Model Hubs

    Jack Bell, Giacomo Carfì, Gerlando Gramaglia, Vincenzo Lomonaco

    cs.AI · cs.LG

    AI model hubs provide access to a rapidly growing collection of powerful pre-trained models, enabling off-the-shelf mixture-of-experts systems with different routing strategies. However, this rapid growth poses two fundamental challenges: scaling model selection across thousands of experts and continually updating routing mechanisms as new models and tasks are introduced. In this paper, we formalise this setting as Continual Model Routing...

    arxiv.org/abs/2605.28577 · PDF

  26. 26

    A Conflict-Aware Penalty and Statistical Loss Framework for Balancing Modalities and Enhancing Stability in Multimodal Sentiment Analysis

    Jianheng Dai, Jiazhang Liang, Sijie Mai

    cs.AI

    Multimodal Sentiment Analysis (MSA) fuses text, acoustic, and visual streams to infer sentiment. Because pre-trained text encoders are far more expressive than their acoustic and visual counterparts, the text modality tends to dominate optimization, suppressing weaker modalities and inducing gradient norm conflicts that destabilize training. To address this, we propose a Conflict-aware Penalty (CP) that detects and penalizes gradient norm...

    arxiv.org/abs/2605.28575 · PDF

  27. 27

    Tree of Thoughts as a Classical Heuristic Search Problem: Formal Foundations and Design Patterns

    Guni Sharon

    cs.AI · cs.LG

    Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities, yet their standard generation process -- auto-regressive token prediction -- is inherently myopic and prone to cascading errors. To address this, the Tree-of-Thoughts (ToT) framework creates a search space over intermediate reasoning steps, allowing search models to explore, look ahead, and backtrack. However, current ToT research remains fragmented across...

    arxiv.org/abs/2605.28566 · PDF

  28. 28

    A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

    Tomer Keren, Nitay Calderon, Asaf Yehudai, Yotam Perlitz, Michal Shmueli-Scheuer, Roi Reichert

    cs.AI

    As agent capabilities advance, existing benchmarks, such as $τ^2$-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remains complex, costly, and labor-intensive. Moreover, the standard approach, in which scenarios are first written in natural language and then mapped to tool sequences, captures only a narrow subset of the tool-use patterns agents exercise. In this paper, we address these problems by reversing...

    arxiv.org/abs/2605.28556 · PDF

  29. 29

    Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

    Matteo Gioele Collu, Riccardo Conte, Alberto Giaretta, Denis Kleyko, Mauro Conti, Matteo Zavatteri, Roberto Confalonieri

    cs.AI · cs.CR

    In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on residual stream activations at each transformer block. We find that refusal is linearly decodable well before the final layer, indicating that safety-relevant behavior is represented in intermediate activations before output generation. To test whether this signal is actionable, we introduce...

    arxiv.org/abs/2605.28553 · PDF

  30. 30

    Modeling Vehicle-Type-Specific Pedestrian Crash Avoidance Behavior in Safety-Critical Interactions Using Smooth-Mamba Deep Reinforcement Learning

    Qingwen Pu, Kun Xie, Hong Yang, Di Yang, Junqing Wang

    cs.AI

    As automated vehicles (AVs) increasingly share roadways with human-driven vehicles (HDVs), understanding how pedestrians respond to different vehicle types in safety-critical interactions is essential for the safe deployment of automated driving technologies. This study extracts safety-critical pedestrian-vehicle interactions from the Argoverse 2 dataset to capture real-world crash avoidance behaviors in encounters involving AVs and HDVs. To...

    arxiv.org/abs/2605.28552 · PDF

  31. 31

    Cultural Binding Heads in Language Models

    Avrile Floro, Luca Benedetto

    cs.AI · cs.CL · cs.LG

    LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness. Using mechanistic interpretability and a factorial design on the N4 cultural appropriation benchmark from Wang et al. (2025), we identify 2-3 mid-layer attention heads per model that contribute causally to cultural binding across eight models (four architectures, base and instruct). Cultural...

    arxiv.org/abs/2605.28543 · PDF

  32. 32

    Do Agents Know What They Can't Do? Evaluating Feasibility Awareness in Tool-Using Agents

    Liang Cheng, Mingsheng Cai, Jiuming Jiang, Luo Mai

    cs.AI

    Tool-using agents often incur substantial computational cost due to long reasoning chains and iterative tool usage. In practical scenarios, many tasks become infeasible under constrained tool environments, where the capabilities required for successful task completion are unavailable. Detecting infeasible tasks and stopping execution early can significantly reduce unnecessary execution cost. In this work, we propose FeasiGen, an automatic...

    arxiv.org/abs/2605.28532 · PDF

  33. 33

    Entropy-aware Masking for Masked Language Modeling

    Gokul Srinivasagan, Kai Hartung, Munir Georges

    cs.AI · cs.CL

    Masked language modeling has become a standard pretraining objective for training encoder-based language models. In this approach, certain tokens in the input are masked, and the model learns to predict them using the surrounding context. This process enables the model to capture both syntactic and semantic properties of language. Conventionally, the tokens selected for masking are chosen at random, which may not always yield the most...

    arxiv.org/abs/2605.28526 · PDF

  34. 34

    Let Relations Speak: An End-to-End LLM-GNN Soft Prompt Framework for Fraud Detection

    Zhixing Zuo, Huilin He, Jiasheng Wu, Dawei Cheng

    cs.AI

    In recent years, Large Language Models (LLMs) have shown great capability in processing graph tasks such as fraud detection. However, most existing methods rely heavily on rich text attributes, which poses difficulties for this domain due to the lack of textual data. Although some pioneering methods attempt to overcome it, their textualization of graph structures via hard prompts easily leads to feature distortion. Additionally, fraud...

    arxiv.org/abs/2605.28524 · PDF

  35. 35

    GS-FUSE: Granger-Supervised Gated Fusion and Multi-Granularity Alignment for Event-Driven Financial Forecasting

    Yang Zhang, En Chun, Ziyun Mao, Yulu Wu, Jun Wang

    cs.AI

    Accurately forecasting the impact of salient financial events on markets is critical for investors and policymakers. However, existing multimodal time-series models typically fuse text and prices symmetrically, without an explicit way to decide when event text is truly predictive, and thus struggle to exploit the directional event-to-price structure and the heterogeneous roles of textual and price signals. In this work, we propose GS-Fuse, a...

    arxiv.org/abs/2605.28520 · PDF

  36. 36

    Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

    Aakash Pant, Kavya Shah, Apoorv Agnihotri, Sneha Nikam, Prasaanth Balraj, Nakul Jain

    cs.AI

    Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape usability as much as model quality. Through a structured analysis of existing benchmark families across speech, chat/RAG, and vision systems, we identify critical gaps between laboratory evaluation practices and real-world deployment conditions in low-resource environments. We argue that the...

    arxiv.org/abs/2605.28508 · PDF

  37. 37

    ProvMind: Provenance-grounded reasoning for materials synthesis

    Yiming Zhang, Ryo Tamura, Koji Tsuda

    cs.AI · cs.LG

    Materials process optimization requires reasoning over routes, conditions, tools and causal dependencies, yet most computational formulations flatten synthesis procedures into text or ordered steps. We introduce MatProcBench, a provenance-grounded benchmark constructed from literature-mined MatPROV graphs, to evaluate seven process-reasoning tasks spanning route continuity, step-level variable inference and global causal consistency under...

    arxiv.org/abs/2605.28487 · PDF

  38. 38

    From Learning Resources to Competencies: LLM-Based Tagging with Evidence and Graph Constraints

    Ngoc Luyen Le, Marie-Hélène Abel, Bertrand Laforge

    cs.AI · cs.IR

    Linking learning resources to a structured competency framework is key to enabling competency-based search and curriculum analytics in Learning Management Systems (LMS). However, manual tagging is labor-intensive, and fully automatic methods often lack transparency. In this paper, we present an end-to-end alignment pipeline that uses a large language model (LLM) as a constrained, evidence-producing tagger. LMS resources -both instructional...

    arxiv.org/abs/2605.28483 · PDF

  39. 39

    Diffusion Large Language Models for Visual Speech Recognition

    Jeong Hun Yeo, Chae Won Kim, Hyeongseop Rha, Yong Man Ro

    cs.AI · cs.CV · eess.AS

    Existing Visual Speech Recognition (VSR) systems commonly rely on left-to-right autoregressive decoding, which can force premature decisions on visually ambiguous tokens before sufficient context is available. We propose DLLM-VSR, to the best of our knowledge, the first Diffusion Large Language Model (DLLM)-based VSR framework, formulating transcription as iterative masked denoising with flexible-order decoding. With confidence-based...

    arxiv.org/abs/2605.28456 · PDF

  40. 40

    GONDOR to the Rescue: Satisficing Planning with Low Memory

    Yonatan Vernik, Alexander Tuisov, Alexander Shleyfman

    cs.AI

    Greedy Best-First Search (GBFS) is the dominant approach for solving search problems where the goal can be estimated with a heuristic, such as planning, route finding, navigation, and pathfinding. This is especially true when the memory is tightly constrained, such as planning on edge devices. To alleviate that, we present GONDOR (Greedy Online Navigation with Dynamic Outpost-based Re-search), a memory-efficient extension of GBFS that allows...

    arxiv.org/abs/2605.28454 · PDF

  41. 41

    DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes

    Caijun Xu, Changyi Xiao, Zhongyuan Peng, Yixin Cao

    cs.AI

    Reinforcement learning has become a central paradigm for advancing reasoning in large language models, yet most existing methods still depend on stronger teacher models or heavily curated difficult datasets, limiting scalable capability improvement. In this paper, we introduce DenoiseRL, a reinforcement learning framework that substitutes external supervision with recovery-oriented optimization over failures from weak models. Instead of...

    arxiv.org/abs/2605.28421 · PDF

  42. 42

    Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning

    Mingze Wu, Abhinav Anand, Shweta Verma, Mira Mezini

    cs.AI

    Post-training using online reinforcement learning (RL) is an important training step for LLMs, including code-generating models. However, online RL for code generation involves LLM inference and verification of the generated output, which can take considerable time and resources. In this paper, we explore the application of offline RL to code-generating models by leveraging existing code datasets. Our experiments demonstrate that offline RL...

    arxiv.org/abs/2605.28409 · PDF

  43. 43

    Measuring Progress Toward AGI: A Cognitive Framework

    Ryan Burnell, Yumeya Yamamori, Orhan Firat, Kate Olszewska, Steph Hughes-Fitt, Oran Kelly, Isaac R. Galatzer-Levy,...

    cs.AI

    Despite widespread discussion of AGI, there is no clear framework for measuring progress toward it. This ambiguity fuels subjective claims, makes it difficult to track progress, and risks hindering responsible governance. As a starting point to address this gap, we present a framework for understanding system capabilities in relation to human cognitive abilities. Drawing from decades of research in psychology, neuroscience, and cognitive...

    arxiv.org/abs/2605.28405 · PDF

  44. 44

    HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs

    Yansong Ning, Mianpeng Liu, Jingwen Ye, Weidong Zhang, Hao Liu

    cs.AI

    Hybrid-reasoning large language models (LLMs) expose explicit controls over reasoning effort, allowing users or systems to trade off answer quality against inference cost. However, existing methods for adaptive thinking-mode selection are typically evaluated under different models, datasets, and implementation assumptions, making it difficult to compare their practical behavior. We introduce HRBench, a unified evaluation framework for...

    arxiv.org/abs/2605.28398 · PDF

  45. 45

    You Live More Than Once: Towards Hierarchical Skill Meta-Evolving

    Xujun Li, Kehan Zheng, Mingyuan Zhao, Yize Geng, Jinfeng Zhou, Qi Zhu, Fei Mi, Lifeng Shang, Minlie Huang, Hongning Wang

    cs.AI

    Test-time skill evolving is regarded as a new paradigm for enhancing deployed agentic systems. Existing works mainly focus on hard-coded skill evolving strategies or parametric learning that rely on expensive parameter updates in the underlying LLMs. In this paper, we demonstrate that test-time refinement of the skill evolving framework itself is necessary for continuous improvement of the agent systems in different downstream scenarios, and...

    arxiv.org/abs/2605.28390 · PDF

  46. 46

    Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs

    Yue Cheng, Jiajun Zhang, Xiaohui Gao, Weiwei Xing, Zheng Wang, Zhanxing Zhu

    cs.AI

    Reinforcement Learning with Verifiable Reward (RLVR) is empirically shown to notably enhance the reasoning performance of large language models (LLMs), particularly in mathematics and programming. However, the mechanistic role of Sample Difficulty in RLVR remains poorly understood. In this paper, we investigate RLVR through the lens of difficulty-wise and one-sample analysis. We find that sample difficulty has a non-monotonic effect on RLVR:...

    arxiv.org/abs/2605.28388 · PDF

  47. 47

    From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence

    Raffael Theiler, Ludovico Comito, David Leko, Leandro Von Krannichfeldt, Lev Telyatnikov, Olga Fink

    cs.AI · cs.LG · cs.SE

    Industrial Prognostics and Health Management (PHM) provides a representative case study for a broader challenge in applied machine learning: translating published papers into executable, benchmark-ready implementations. Reproducing under-specified methods in PHM is particularly difficult due to restricted access to industrial datasets, incomplete reporting of preprocessing and evaluation protocols, and implicit design choices (e.g.,...

    arxiv.org/abs/2605.28371 · PDF

  48. 48

    CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict

    Yanhui Sun, Wu Liu, Haifeng Ming, Xinru Wang, Hantao Yao, Yongdong Zhang

    cs.AI · cs.SI

    E-commerce platforms have begun recruiting crowdsourced jurors to adjudicate massive volumes of transaction disputes. Unlike formal legal judgment, E-commerce dispute verdicts require grounding pivotal clues from redundant, multi-round, multimodal evidence and making decisions under flexible platform-specific conventions. These characteristics render existing methods insufficient for this scenario. To bridge this gap, we introduce a...

    arxiv.org/abs/2605.28369 · PDF

  49. 49

    Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning

    Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou Ammar

    cs.AI · cs.CL · cs.LO

    Lean is increasingly used to judge natural-language mathematical answers, but its signal is partial: many answers never formalize, and a failed proof may reflect an ill-typed statement or a missing library fact, not a wrong answer. On MATH-500 we show this signal is (i) sharply coverage-dependent, that is the proof-winning answer is correct 96% of the time at high proved coverage but 20% at low, and (ii) sparse and often unfaithful: a 7B...

    arxiv.org/abs/2605.28365 · PDF

  50. 50

    Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement

    Jyotirmoy Nath, Neeraj Kumar, Brejesh Lall

    cs.AI

    Automatic prompt optimization (APO) has driven significant gains in LLM-based agentic workflows. However, existing methods treat each task's prompt as a monolithic, instance-blind string optimized through global edits, producing brittle updates and preventing the reuse of learned sub-behaviors. We propose Prompt Codebooks (PCO), a novel compositional prompt optimization framework that recasts APO as discrete learning over a finite vocabulary...

    arxiv.org/abs/2605.28360 · PDF

  51. 51

    From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets

    Taojie Zhu, Wentao Zhao, Rui Sun, Beidi Luan, Jiacheng Lu, Sinuo Wang, Jing Li, Daxin Jiang, Yonghong He, Zuo Bai

    cs.AI · q-fin.TR

    Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a historical market, let it trade, and measure portfolio returns. This setup is vulnerable to two evaluation failures. First, long backtests often overlap with the knowledge cutoffs of frontier LLMs, allowing memorized tickers, dates, prices, and market narratives to substitute for investment...

    arxiv.org/abs/2605.28359 · PDF

  52. 52

    Plan Before Search: Search Agents Need Plan

    Zhipeng Qian, Zihan Liang, Yufei Ma, Ben Chen, Huangyu Dai, Jiayi Ji, Chenyi Lei, Wenwu Ou, Xiaoshuai Sun, Qibin Hou

    cs.AI

    Training large language models as retrieval-augmented reasoning agents typically combines reinforcement learning with an SFT cold start distilled from a stronger model. However, this paradigm overlooks two fundamental factors: the dependency structure among sub-skills, and the possibility that distillation is not the only route to capability acquisition. We study this through Plan, a structured agentic behavior for multi-hop retrieval that...

    arxiv.org/abs/2605.28354 · PDF

  53. 53

    FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models

    Xucong Wang, Pengkun Wang, Zhe Zhao, Liheng Yu, Shuang Wang, Yang Wang

    cs.AI

    Multi-Label Recognition (MLR) based on Vision-Language Models (VLMs) aims to leverage their pre-trained knowledge to better adapt complex recognition scenarios, thereby enhancing model robustness. However, for realistic decentralized applications requiring federated learning, adapting VLMs to each client that possesses private and heterogeneous data can cause the model to overfit spurious label correlations, consequently triggering irrelevant...

    arxiv.org/abs/2605.28347 · PDF

  54. 54

    Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains

    Lev Telyatnikov, Raffael Theiler, Leandro Von Krannichfeldt, Olga Fink

    cs.AI · cs.LG · eess.SP

    Progress in Prognostics and Health Management (PHM) is hindered by the lack of standardized and reusable evaluation practices across tasks, datasets, and application domains. Reported results are often difficult to reproduce and compare, as key protocol choices, such as data splits, preprocessing, label alignment, temporal windowing, and metrics, are often implicit or implemented ad hoc. We introduce \picid, a modular evaluation...

    arxiv.org/abs/2605.28345 · PDF

  55. 55

    SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models

    Chao Ding, Mouxiao Bian, Tianbin Li, Minjia Yuan, Yidong Jiang, Yankai Jiang, Jinru Ding, Jiayuan Chen, Zhuangzhi...

    cs.AI

    Large language models(LLMs) increasingly match expert performance on licensing examinations, yet routine clinical use remains limited because governance requires auditable reasoning, safety and ethics alignment, and resilience to adversarial misuse. Here we present SafeMed-R1, trained with a traceable Clinical Trust Signals(CTS) pipeline that links each reasoning instance to clinician rubric scores and edit histories, and aligned through...

    arxiv.org/abs/2605.28338 · PDF

  56. 56

    An Enhanced Large Neighborhood Search Approach for the Capacitated Facility Location Problem with Incompatible Customers

    Ida Gjergji, Lucas Kletzander, Nysret Musliu, Andrea Schaerf

    cs.AI

    A new variant of the classic capacitated facility location problem, which considers incompatibilities between customers, has recently been introduced in the literature. This problem captures the situation where given pairs of customers cannot be served by the same facility. Such a feature is crucial for many practical cases of location problems, such as the presence of hazardous or polluting materials and contention between competing...

    arxiv.org/abs/2605.28337 · PDF

  57. 57

    From Fact Overwriting to Knowledge Evolution: Causal Editing via On-Policy Self-Distillation

    Shuaike Li, Kai Zhang, Xianquan Wang, Jiachen Liu, Shengpeng Mo

    cs.AI

    While Knowledge Editing (KE) enables efficient updates, its dominant Static Fact Overwriting paradigm treats LLMs as discrete databases, forcibly injecting isolated facts. Fracturing pre-trained logical topologies, this triggers Epistemic Dissonance -- a pathology where un-evolved legacy priors force the model to explicitly negate the injected update. Idealized interventions reveal that this is an inherent structural flaw rather than mere...

    arxiv.org/abs/2605.28303 · PDF

  58. 58

    Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

    Zhaoyang Jiang, Xuanqi Peng, Fei Teng, Zhizhong Fu, Yunsoo Kim, Jiacong Mi, Zicheng Li, Honghan Wu

    cs.AI

    Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics including accuracy. We ask whether gains in answer quality are accompanied by improvements in the trace. In medical QA, where short answer options can leave a richer clinical justification under-specified, a Qwen3-8B student distilled from a DeepSeek-V3-family teacher improves on MedQA-USMLE...

    arxiv.org/abs/2605.28301 · PDF

  59. 59

    REED: Post-Training Representation Editing for Cross-Domain Linguistic Steganalysis

    Ruohan Lei, Jianxin Gao, Wanli Peng, Huimin Pei

    cs.AI

    In real-world scenarios of linguistic steganalysis, tested texts usually come from unseen domains with different vocabularies, topics, writing styles, and steganographic generation patterns, which can significantly degrade the detection performance. Although existing cross-domain steganalysis methods can effectively alleviate this problem through distribution alignment, domain-invariant feature learning, etc., the detection performance is not...

    arxiv.org/abs/2605.28298 · PDF

  60. 60

    Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR

    Soeun Kim, Albert No

    cs.AI · cs.CL · cs.LG

    Reinforcement Learning with Verifiable Rewards (RLVR) trains reasoning models without labeled trajectories, relying on grouped rollouts to expose the policy to alternative reasoning paths and a verifier to score them. Rollout diversity has accordingly emerged as a central bottleneck in RLVR, with most existing methods broadening exploration through temperature, prefix, or rollout-selection adjustments. We identify a structurally distinguished...

    arxiv.org/abs/2605.28295 · PDF

  61. 61

    ResearchLoop: An Evidence-Gated Control Plane for AI-Assisted Research

    Yihan Xia, Taotao Wang

    cs.AI

    AI-assisted research compresses ideation, implementation, evaluation, and manuscript writing into a single interactive loop. This compression is useful, but it also creates a publication risk: paper claims can become easier to state than to audit. We present ResearchLoop, an evidence-gated control plane for AI-assisted computational research. ResearchLoop treats research questions, task contracts, evidence objects, claim ledgers, closeouts,...

    arxiv.org/abs/2605.28282 · PDF

  62. 62

    Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning

    Zhikai Pan, Chih-Ting Liao, Chunrui Liu, Xi Xiao, Yitong Qiao, Chunlei Meng, Zhangquan Chen, Xin Cao

    cs.AI

    Whether large language models (LLMs) construct internal spatial world models from pure-text descriptions remains contested, and whether such capabilities transfer across languages has not been systematically studied. We introduce MentalMap, a multilingual diagnostic benchmark with a six-level capability hierarchy (L0-L5) spanning atomic spatial facts to generative world-graph construction, together with four diagnostic axes probing frame of...

    arxiv.org/abs/2605.28277 · PDF

  63. 63

    Global Policy-Space Response Oracles for Two-Player Zero-Sum Games

    Junyu Zhang, Feihong Yang, Jian Wang, Chao Wang, Xudong Zhang

    cs.AI

    The Policy-Space Response Oracles (PSRO) framework scales equilibrium computation to large zero-sum games by iteratively expanding a restricted strategy set using deep reinforcement learning (DRL). A central challenge is to construct, under limited computational budgets, a small strategy population whose induced game well approximates the full game. Existing PSRO variants typically expand the population using best responses to meta-strategies...

    arxiv.org/abs/2605.28273 · PDF

  64. 64

    Entropy Distribution as a Fingerprint for Hallucinations in Generative Models

    Mattia J. Villani, Pranav Deshpande, Akshay Seshadri, Romina Yalovetzky, Niraj Kumar

    cs.AI

    Large Language Models (LLMs) often generate factually incorrect outputs, commonly termed hallucinations, that undermine trust and limit deployment in high-stakes settings. Existing hallucination detection methods typically require multiple forward passes, or access to model internals. In this work, we provide theoretical background and empirical evidence that the distribution of token-level entropies, beyond the mean captured by perplexity or...

    arxiv.org/abs/2605.28264 · PDF

  65. 65

    AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?

    Maharshi Gor, Yoo Yeon Sung, Yu Hou, Eve Fleisig, Irene Ying, Tianyi Zhou, Jordan Boyd-Graber

    cs.AI · cs.CL · cs.HC

    AI systems are fallible, and humans can make mistakes in deciding whether to trust AI over their own judgment. Thus, improving human-AI collaboration requires understanding when, why, and how humans decide to rely on AI. We study two distinct reliance decisions: the delegation choice -- deciding when to let AI act autonomously without knowing its output, and the adoption choice -- evaluating AI suggestions and deciding how to use them. Both...

    arxiv.org/abs/2605.28255 · PDF

  66. 66

    PIRS: Physics-Informed Reward Shaping for SAC-Based Building Energy Management

    Shadmehr Zaregarizi, Khashayar Yavari

    cs.AI

    Occupant comfort and grid-aware energy efficiency are competing objectives whose joint optimization depends critically on how reward functions are specified in deep reinforcement learning (DRL) controllers for buildings. Yet reward design remains largely ad hoc: comfort terms are either hand-tuned heuristics or simple temperature-deviation proxies without explicit grounding in thermal-comfort physics. We present PIRS (Physics-Informed Reward...

    arxiv.org/abs/2605.28232 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.