cs.AI · 2026-07-27 · No. 66

Artificial Intelligence, 2026-07-27.

31 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

31 entries
  1. 01

    Explainable Reinforcement Learning for assisting Air Traffic Controllers

    Anduel Mehmeti, Gabriella Gigante, Salvatore Venticinque

    cs.AI

    To effectively integrate AI into high-stakes, critical environments such as healthcare, autonomous driving, and aviation--and to advance toward higher levels of automation and seamless human-AI collaboration--building trust in AI-driven solutions is essential. Trust, in turn, is closely linked to the explainability of AI systems. The rapid advancements in AI across various domains have underscored the challenges of establishing trust, raising...

    arxiv.org/abs/2607.22525 · PDF

  2. 02

    The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents

    Darshan Tank, Baran Nama

    cs.AI

    Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this metric hides an important cost: skills can also make agents worse. We measure both sides by comparing agents with and without skills across nearly 6,000 runs spanning two office automation benchmarks and three model harness stacks. This allows us to distinguish two outcomes. A regression is a task solved without skills but...

    arxiv.org/abs/2607.22520 · PDF

  3. 03

    TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI

    Ritik Raj, Souvik Kundu, Sarbartha Banerjee, Dheemanth Joshi, Ishita Vohra, Tushar Krishna

    cs.AI · cs.LG · cs.MA

    Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, agentic applications execute as long-horizon workflows whose quality is determined only by a delayed, task-level outcome. This mismatch prevents per-call routers from correctly attributing feedback to...

    arxiv.org/abs/2607.22465 · PDF

  4. 04

    Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture

    Halil Burak Noyan

    cs.AI · cs.CR

    Enterprise AI agents are typically granted static credential sets at configuration time, holding every tool the role might need for every task they perform. This persistent over-privilege expands the attack surface. We argue that capability scoping must follow a dynamic least-privilege principle and be treated as a prevention mechanism before a detection one. A credential that does not exist in an agent's context cannot be misused regardless...

    arxiv.org/abs/2607.22445 · PDF

  5. 05

    SceneActBench: Can Agents Act on the 3D Scenes They See?

    Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang, Pu Jian, Huanjin Yao, Jiarui Yao, Haowei Lin, Chunchao Guo, Zhuo...

    cs.AI · cs.CV

    Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where...

    arxiv.org/abs/2607.22393 · PDF

  6. 06

    Agentic Root Cause Analysis through Evidence-Grounded Reasoning

    Amaury Wei, Olga Fink

    cs.AI · cs.LG

    Diagnosing the root cause of anomalies is essential for safe industrial operation. Despite extensive sensor instrumentation, formulating hypotheses and gathering evidence remains a manual process, creating a major operational bottleneck. While existing data-driven approaches aim to automate this, two critical limitations restrict their deployment: their operate as black boxes unable to justify their diagnosis, and they require scarce labeled...

    arxiv.org/abs/2607.22385 · PDF

  7. 07

    IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation

    Varun Gumma, Navonil Majumder, Soumitra Sinhahajari, Soujanya Poria

    cs.AI

    Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize ideas independently for either Quality or Diversity. This often leads to the generation of ideas in close proximity to one another or to a large set of trivial, unsound, or unclear concepts. In this work, we instead argue that research ideation...

    arxiv.org/abs/2607.22375 · PDF

  8. 08

    Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI

    Jiaqi Shao, Hanck Chen, Wei Zhang, Maxm Pan, Bing Luo

    cs.AI

    Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can instead recover public solutions, read evaluation artifacts, infer generator structure, manipulate feedback, or benefit from...

    arxiv.org/abs/2607.22368 · PDF

  9. 09

    Learning Structural Convergence: A Neuro-Symbolic Benchmark for Temporal Reasoning

    Michael Romei De Socio, Gian Luca Pozzato, Alessio Merlo

    cs.AI

    High-complexity operational environments require methods that detect and anticipate temporally distributed patterns rather than classify isolated events. This paper introduces TRACTA (Temporal Reasoning and Capability-Trajectory Analysis), a controlled synthetic benchmark for temporal structural reasoning in high-complexity event-driven systems, instantiated through Multi-Domain Operations (MDO)-like scenarios. The benchmark includes three...

    arxiv.org/abs/2607.22365 · PDF

  10. 10

    A Roadmap to Impactful Pluralistic Alignment Research

    Elinor Poole-Dayan, Jillian Fisher, Atoosa Kasirzadeh, Jacob Andreas, Mitchell Gordon, Michiel A. Bakker

    cs.AI

    Pluralistic value alignment---the goal of building AI systems that represent and serve diverse human values and perspectives---has emerged as an active research agenda. Yet, there's no public evidence that it has shaped the training or evaluation of the AI systems people actually use. We audit the public behavior documents and evaluations of frontier labs, finding none name pluralism as a goal, and as of this writing, no clear indication that...

    arxiv.org/abs/2607.22305 · PDF

  11. 11

    AI4PLE: A Methodology for Integrating AI into Product Line Engineering

    Bedir Tekinerdogan

    cs.AI

    Reuse-based development has become increasingly important in the creation of complex systems, offering significant opportunities to reduce costs, improve quality, and accelerate time-to-market. Product Line Engineering (PLE) provides a systematic approach to realizing this potential by enabling the efficient creation, management, and customization of product families by reusing shared assets and capabilities. PLE involves addressing numerous...

    arxiv.org/abs/2607.22260 · PDF

  12. 12

    Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

    Guanqun Zhao, Zijun Xie, Binbin Zheng, Enlei Gong, Jiafeng Lu, Yehan Yang, Aoqi Hu, Zeyu Chen

    cs.AI

    Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause policy collapse. Existing methods typically retain or discard tokens based solely on the magnitude of their importance ratios, applying the same threshold uniformly across token positions. In this...

    arxiv.org/abs/2607.22186 · PDF

  13. 13

    Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents

    Valentin Tablan, Scott Taylor, Kristoffer Bernhem

    cs.AI

    AI agents encounter learning opportunities in every episode they run, and discard nearly all of them: the underlying models are frozen at deployment, so an agent that resolves a difficult request today starts from zero when it recurs tomorrow. Yet ordinary operation already produces feedback, in the form of outcome verdicts and after-the-fact corrections. We show that this feedback is a sufficient signal for continual learning when the frozen...

    arxiv.org/abs/2607.22157 · PDF

  14. 14

    Industrial Tokenization for LLM-Based Health Intelligence: A Federated Architecture for Industrial Evidence Integration

    Deshui Li, Xiao-Ming Yuan, Zishun Wang

    cs.AI · cs.LG

    Industrial health management increasingly relies on heterogeneous information sources, including condition monitoring systems, supervisory control and data acquisition systems, maintenance records, inspection results, and prognostic models. Although large language models provide new opportunities for cross-source reasoning, industrial data and analytical outputs differ substantially in structure, temporal resolution, physical meaning, and...

    arxiv.org/abs/2607.22153 · PDF

  15. 15

    Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models

    Junlin Fang, Do Nguyen-Thanh, Xiaogang Xu, Zhen Fang, Sean Du

    cs.AI · cs.LG

    Large reasoning models (LRMs) generate long reasoning traces before producing final answers. While these traces may contain useful signals for hallucination detection, harnessing them is non-trivial because long trajectories often include noisy steps that obscure the cues relevant to truthfulness assessment. In this paper, we identify two prevalent forms of reasoning noises, i.e., irrelevant steps and repetitive steps, and show that both...

    arxiv.org/abs/2607.22098 · PDF

  16. 16

    Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode

    Nanbeige Lab, :, Chen Yang, Chengrui Huang, Fufeng Lan, Hanhui Chen, Hao Zhou, Huatong Song, Jiaqi Cao, Jiaying Zhu,...

    cs.AI · cs.CL

    We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT...

    arxiv.org/abs/2607.22083 · PDF

  17. 17

    Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

    Suman Navaratnarajah, Taehyoung Kim, Jona Ruthardt, Ishaan Bhimwal, Ryousuke Yamada, Yannik Blei, Wolfram Burgard,...

    cs.AI · cs.CL · cs.CV · cs.RO

    Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level instruction. We introduce MissionBench, a benchmark for mission-level evaluation of MLLMs in aerial 3D environments. It comprises 120 missions across five simulated 3D environments and four task families. Agents must...

    arxiv.org/abs/2607.22014 · PDF

  18. 18

    Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning

    Heyang Jiang, Henry Liu, Baharan Mirzasoleiman

    cs.AI

    Reinforcement learning with verifiable rewards (RLVR) has emerged as a highly effective framework for improving LLM reasoning, with methods such as GRPO among its most successful instantiations. However, GRPO relies on repeated generation of long chain-of-thought rollouts. Training time scales with the number of rollouts, a large fraction of which are uninformative. Thus, GRPO is computationally expensive and unstable. To mitigate this,...

    arxiv.org/abs/2607.22002 · PDF

  19. 19

    Semiotic logical hexagon theory for LLM logical reasoning

    Yunyao Zhang, Xinglang Zhang, Zeliang Chen, Junqing Yu, Zikai Song

    cs.AI

    Large language models (LLMs) have become powerful tools for language understanding and logical reasoning. However, they still make mistakes when a problem requires both understanding meaning and following logic. A key reason is that natural-language statements often carry implicit semantic relations before any formal reasoning begins. If these hidden meanings are not properly organized, the model may reach incorrect conclusions even when the...

    arxiv.org/abs/2607.21933 · PDF

  20. 20

    TRW: TRACE-RealWorld---An Auditable Consistency Contract for World Models as Materialized Views

    Edward Y. Chang

    cs.AI · cs.DB

    TRACE-RealWorld addresses a core data-management problem: maintaining an actionable materialized view over a continuously changing physical world when reads of the base state are priced, delayed, heterogeneous, and fallible. Its data-management contributions are a commitment-level validity abstraction for materialized predictions; consequence-conditioned adaptive view maintenance; transaction-style, dependency-scoped compensation for...

    arxiv.org/abs/2607.21910 · PDF

  21. 21

    Multi-Agent System-driven Digital Twins for predictive maintenance: architectures, technologies and open research challenges

    Korota Arsène Coulibaly, Mohamed Hamlich

    cs.AI

    Digital twins have emerged as a foundational technology within the context of Industry 4.0, offering a paradigm for the real-time virtual representation of physical systems. However, managing their growing complexity, particularly in distributed industrial environments, requires intelligent architectures capable of autonomous decision-making, dynamic adaptability, and inter-agent coordination. This systematic review explores the intersection...

    arxiv.org/abs/2607.21873 · PDF

  22. 22

    When Is a Learned Command Adapter Worth It? Closed-Loop Identification and Counterfactual Auditing of Frozen Locomotion Policies

    Zongtan Li

    cs.AI · eess.SY

    Adding a learned adapter to a frozen, command-conditioned locomotion policy is worthwhile only if the interface exposes improvements that are both real and recoverable from deployment-time observations. We introduce an adapter necessity audit that separates global operating-point gain,same-state counterfactual headroom, deployment gain over a cross-fitted fixed action, and state-allocation gain over a frequency-matched randomized policy....

    arxiv.org/abs/2607.21867 · PDF

  23. 23

    DAGForge: Auditable Causal DAG Authoring with Biomedical Literature

    Yi-han Sheu, Michael R. Steigman, Yu Zhou, Bo Wang, Fan-Yu Yen, Jordan W. Smoller

    cs.AI

    Constructing causal directed acyclic graphs (DAGs) is a core step in biomedical causal analysis, yet it remains a largely manual process. Analysts must connect study variables to prior literature, evaluate uncertain causal claims, and preserve sufficient provenance for expert review. We present DAGForge, a browser-based system for authoring causal DAGs as auditable, evidence-linked artifacts. Given free-text descriptions of study concepts,...

    arxiv.org/abs/2607.21859 · PDF

  24. 24

    QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization

    Siwei Chen, Siqi Chen, Xupeng Miao, Bin Cui

    cs.AI

    Recent large reasoning models often develop long chain-of-thought responses during reinforcement learning (RL), resulting in high inference latency and deployment cost. Existing methods for response length control typically rely on explicit length penalties or additional control modules, which require careful tuning and may compromise reasoning quality. We propose Quadrant-weighted Sampling for Length-aware Policy Optimization (QLPO), a...

    arxiv.org/abs/2607.21793 · PDF

  25. 25

    From Seasonality to Semantics: Benchmarking a Hybrid Probabilistic Forecasting System for Roadblocks in Bolivia

    Rodrigo Vargas Sainz, Christian Berón Curti

    cs.AI · cs.CL · cs.LG

    Roadblocks in Bolivia are a social conflict phenomenon with devastating economic impacts, estimated at losses equivalent to 4% of the national Gross Domestic Product. Despite their recurrence and impact, there is a lack of local predictive systems to anticipate these events for logistical decision-making. This paper presents a hybrid probabilistic forecasting system that integrates time series decomposition (Prophet) with natural language...

    arxiv.org/abs/2607.21785 · PDF

  26. 26

    What AI Red-Team Evaluations Can and Cannot Prove

    Bandana Kaur

    cs.AI · cs.CR

    Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a...

    arxiv.org/abs/2607.21735 · PDF

  27. 27

    Unsupervised Consensus-Based Anomaly Detection for Spatiotemporal Malaria Incidence in Ghana

    T. Ansah-Narh, Y. Asare Afrane

    cs.AI · cs.CE · cs.ET · stat.AP · stat.ML

    A consensus anomaly detection framework was applied to monthly malaria surveillance data from Ghana (2014-2023) to identify atypical transmission patterns. Anomalies were highly structured in space and time. Ashanti and Northern Regions accounted for most recurrent anomalies, with persistent hotspots at Tamale, Kumasi, and Accra. A key finding was the spatial distinction between anomaly burden (cumulative cases during anomalous periods) and...

    arxiv.org/abs/2607.21559 · PDF

  28. 28

    Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

    Baihui Wang, Bernard Koch

    cs.AI

    Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode. Models must distinguish when to incorporate others' perspectives from when to maintain a well-grounded moral judgment. We study the broader resistance-compliance process governing this distinction. Across three studies, we show that models' judgment revision...

    arxiv.org/abs/2607.21558 · PDF

  29. 29

    OpenForgeRL: Train Harness-native Agents in Any Environment

    Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou, Qianhui Wu, Hao Cheng, Wenlin Yao, Nikhil Singh, Zhou Yu, Jianfeng Gao

    cs.AI · cs.CL

    Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training...

    arxiv.org/abs/2607.21557 · PDF

  30. 30

    MIRROR: Learning from the Other View for Multi-Modal Reasoning

    Wen Ye, Yuxiao Qu, Aviral Kumar, Xuezhe Ma

    cs.AI · cs.LG

    Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit different behaviors: a model may solve a problem from text but fail on the corresponding diagram, or succeed visually while failing textually. This inconsistency suggests...

    arxiv.org/abs/2607.21552 · PDF

  31. 31

    The Boundaries of Automation: A Theory of Persistent Human Participation

    Fares Fourati, Hinrich Schütze, Eyke Hüllermeier, Iryna Gurevych

    cs.AI · cs.CL · cs.ET · cs.LG · cs.MA

    The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with algorithms wherever possible. Implicit in this pursuit is the assumption that humans remain in the loop only because current AI systems are not yet sufficiently capable. This paper challenges that assumption. Rather than asking how far automation can extend, we ask where its conceptual limits lie and argue that human...

    arxiv.org/abs/2607.21547 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.