cs.AI · 2026-07-29 · No. 68

Artificial Intelligence, 2026-07-29.

64 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

64 entries
  1. 01

    Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

    Abhishek Pillai, Samir Kumar Nayak, Yuan Chen

    cs.AI · cs.CV

    Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering,...

    arxiv.org/abs/2607.26041 · PDF

  2. 02

    Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

    Elias Fernández Domingos, The Anh Han

    cs.AI · cs.CY · cs.GT · econ.GN

    Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly...

    arxiv.org/abs/2607.26034 · PDF

  3. 03

    CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer

    Ankang Yang, Jitao Zhao, Di Jin, Yuxiao Huang, Dongxiao He

    cs.AI

    Graph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains and tasks. Real-world graphs associate nodes with text, images, and other modalities, making multimodal graphs essential for representing complex entities and relations. Moreover, collecting labels and adapting models for every new graph domain is costly and often infeasible, motivating zero-shot transfer. Unfortunately,...

    arxiv.org/abs/2607.26023 · PDF

  4. 04

    Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation

    Jintao Xu, Yingzheng Ma, Jiong Dong, Yongzhi Qi, Jianshen Zhang

    cs.AI · math.OC

    Multi-warehouse inventory allocation is typically formulated as a mixed-integer programming (MIP) problem, yet no single formulation consistently matches heterogeneous instance-level regimes induced by demand concentration, inventory imbalance, replenishment scale, service constraints, and forecast volatility. We study this issue as instance-wise operations research (OR) formulation selection, where each allocation instance is assigned to a...

    arxiv.org/abs/2607.25956 · PDF

  5. 05

    A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series

    Frank Nie, Ethan B Liu, Yuan Zhu, Wei Fan, Jindong Han

    cs.AI · cs.CL

    Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable promise in general-purpose time-series QA, they remain poorly equipped to model the sparsity, asynchrony, and irregular sampling patterns of clinical observations. To fill this gap, we propose ClinPRISM, a cost-effective...

    arxiv.org/abs/2607.25947 · PDF

  6. 06

    dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision Trees

    Tereza Kinská, Jan Křetínský, Tobias Meggendorfer, Sabine Rieder, Maximilian Weininger

    cs.AI

    Over the past decade, decision trees have been used to represent controllers (a.k.a. policies) in an explainable way, with dtControl2 as a current state-of-the-art tool. However, for systems that are large or have many corner cases, even such representations tend to be too complex and not human-comprehensible. Unfortunately, reducing the size of the decision tree is not straightforward, as missing just a single crucial case might result in an...

    arxiv.org/abs/2607.25925 · PDF

  7. 07

    Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

    Yutong Chen, Shouqian Shi, Xinran Liu, Haochen Wang, Jiaying Wang, Tianxing Xu, Yuanxi Wang, Zirui Ding

    cs.AI

    Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning computation to autoregressive output length. We introduce Penelope, an efficient latent-reasoning framework for pretrained decoder-only...

    arxiv.org/abs/2607.25915 · PDF

  8. 08

    Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks

    Ravi Kant Sharma, Ashutosh Uttam, Ajay Kumar

    cs.AI · cs.CR · cs.NI

    Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack a standardized mechanism for cross-vendor trust visibility. When a tool from Vendor B is compromised, agents from Vendor A continue invoking it -- unaware of the trust degradation -- causing cascading service impact. We present AgentToolMO, a proposed 3GPP NRM information model for agent tool...

    arxiv.org/abs/2607.25914 · PDF

  9. 09

    Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

    Chenrui Shi, Yuwei Wu, Yang Liu, Ruining Feng, Zirui Shang, Zhi Gao, Lifeng Fan, Che Sun

    cs.AI

    Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system...

    arxiv.org/abs/2607.25904 · PDF

  10. 10

    Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

    Stefan Krsteski, Charlotte Meyer, Guillaume Allegre, Tony O'Halloran, Alexandre Sallinen

    cs.AI · cs.DB

    Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of the empirical record incomparable. We introduce Messier, a unified corpus of 957,253 records that span 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. Messier consolidates public benchmark scores...

    arxiv.org/abs/2607.25891 · PDF

  11. 11

    Distributing Security Controls Through Harness Engineering

    William Robert Gore

    cs.AI

    AI coding agents are being adopted at historic speed, yet security and risk concerns remain the primary barrier to scaling agentic AI across organizations. Existing security controls for coding agents are not systematically distributed to engineering teams, and vendor-native solutions introduce ecosystem dependencies that may not suit every deployment context. This paper investigates whether off-the-shelf security controls can be implemented...

    arxiv.org/abs/2607.25890 · PDF

  12. 12

    Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks

    Bart Custers, Koorosh Aslansefat

    cs.AI

    This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a particular focus on uncertainty quantification. Actuarial workflows represent a high-stakes decision-support setting where unreliable outputs may lead to incorrect risk assessment, unfair pricing, and regulatory non-compliance. To address uncertainty introduced by the probabilistic nature of LLMs and...

    arxiv.org/abs/2607.25877 · PDF

  13. 13

    HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs

    Yu Hao, Jinxuan Cai, Qi Zhang, Yawen Li, Zhiqiang Zhang, Chuan Shi, Cheng Yang

    cs.AI

    Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interactive tasks. However, existing trajectory-to-skill methods often produce flat collections of high-level textual skills that are stored and retrieved independently, leaving skill relations underutilized and maintaining a gap between high-level skills and executable actions. In this paper, we propose HiSkill,...

    arxiv.org/abs/2607.25853 · PDF

  14. 14

    Distributed Constraint Optimization via Online Learning and Iterative Pricing with Application to Large-Scale Satellite Scheduling

    Itai Zilberstein, Pranav Rajbhandari, Steve Chien, Tuomas Sandholm

    cs.AI · cs.GT

    Distributed constraint optimization problems (DCOPs) provide a popular framework for distributed decision making under limited communication, but many real-world instances are too large to solve monolithically. We address this challenge from two complementary directions. We revisit the connection between DCOPs and potential games, and adapt modern online learning algorithms for equilibrium finding to DCOPs. We show that these algorithms are...

    arxiv.org/abs/2607.25835 · PDF

  15. 15

    Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL

    Jiabao Ji, Yujian Liu, Li An, Rohit Jain, Gungor Polatkan, Siyu Zhu, Shiyu Chang

    cs.AI

    Large language model agents often spend substantial wall-clock time waiting for tool call results. Tool-call speculation can hide this latency by predicting and pre-executing an agent's next tool call if the prediction matches the agent's eventual tool call, but existing speculators are typically separate draft models or cached traces that are poorly aligned with the deployed agent's own behavior. We identify this speculator-agent gap and...

    arxiv.org/abs/2607.25816 · PDF

  16. 16

    Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography

    Hyunkyung Han, Min Jung Kim

    cs.AI

    Objective: Concept bottleneck models route prediction through interpretable intermediate variables, and their validity is normally judged by how accurately those variables are predicted. We ask whether that judgement is sufficient, using left ventricular volumes as the concepts underlying ejection fraction estimation from echocardiographic video. Methods: A video transformer encoder was trained on a publicly available echocardiography...

    arxiv.org/abs/2607.25748 · PDF

  17. 17

    Nudging Sustainable Choices through LLM-Generated Recommendation Explanations

    Haya Halimeh, Dietmar Jannach, Oliver Müller

    cs.AI

    Recommender systems mediate everyday consumption, offering a promising channel for encouraging sustainable choices. Prior research shows that explanations influence users' perceptions of recommendations and can support more informed decisions. We argue that explanations can also serve as behavioral nudges by foregrounding sustainability information at the moment of choice. This study investigates how different behavioral framings of...

    arxiv.org/abs/2607.25726 · PDF

  18. 18

    Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare

    Qi Chen, Siria Xiyueyao Luo, Jian Wang, Yuan Shi, Haocong Rao, Xuejiao Zhao

    cs.AI

    Cognitive distortion amplifies negative emotions and contributes to mental health disorders. Cognitive Behavioral Therapy (CBT) is an effective way to address cognitive distortions, but its large-scale application is limited by the shortage of professional therapists. Although large language models (LLMs) have recently been explored for mental health applications, existing methods still suffer from limited domain specificity, overly...

    arxiv.org/abs/2607.25681 · PDF

  19. 19

    DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

    Jiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou, Haiyan Wu, Haihan Shi, Chenxi Zhou, Hanqing Li, Xiao Yang, Da...

    cs.AI

    Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-ended tasks, this can become a bottleneck: once the solver improves on the criteria a rubric measures, omitted dimensions remain invisible to the...

    arxiv.org/abs/2607.25675 · PDF

  20. 20

    OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

    Haoyang Huang, Wenjie Huang, Tianqi Xu, Hongyaoxing Gu, Kang Tan, Yikai Fu, Yuhao Shen, Tianyu Liu, Baolin Zhang,...

    cs.AI

    Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important tokens under fixed budgets, leaving the preceding budget-allocation problem underexplored. We show that direct query-to-audio/video similarity is unreliable for inter-modal budget...

    arxiv.org/abs/2607.25669 · PDF

  21. 21

    Localized Adaptation Reveals Distinct Learning Signatures in Transformers

    Rebecca Ramnauth, Brian Scassellati

    cs.AI · cs.CL

    Transformer adaptation is typically distributed across model depth, even when the intended change is narrow. We investigate how adaptation site shapes what a model learns, how well that learning generalizes, and how selectively it is applied. We introduce a controlled benchmark spanning five objectives (lexical binding, factual association, behavioral policy learning, causal mapping, and procedural reasoning) and define each objective's...

    arxiv.org/abs/2607.25663 · PDF

  22. 22

    CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

    Bo-Wen Zhang, Junwei He, Wen Wang, Song-Lin Lv, Wentao Ma, Rongyi Lin, Shuhan Zhong, Lan-Zhe Guo

    cs.AI

    Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in...

    arxiv.org/abs/2607.25659 · PDF

  23. 23

    OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation

    Zhenzhen Ren, Jiyan He, Xinpeng Zhang, Zhenxing Qian, Ke Han, Shuxin Zheng, GuoBiao Li, Xiaoqing Zhang

    cs.AI

    Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (MAS). Existing evaluations typically rely on end-to-end execution, which conflates orchestration-plan quality with worker capabilities, tool reliability, and environmental noise. Moreover, the time and token costs of real execution grow rapidly with workflow scale, making systematic evaluation...

    arxiv.org/abs/2607.25656 · PDF

  24. 24

    Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions

    Jesung Park

    cs.AI · physics.soc-ph · stat.AP

    Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-stable) and that humans actually reach on Lichess (October 2025; 1,661 positions, 16.1M occurrences), human results are not balanced. Positions carry outcome skews, each the gap between its games' actual results and what the players' ratings predict, whose directions are stable properties of the...

    arxiv.org/abs/2607.25655 · PDF

  25. 25

    AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations

    Javier Irigoyen, Roberto Daza, Francisco Jurado, Julian Fierrez, Ruben Tolosana, Alvaro Ortigosa, Miguel...

    cs.AI · cs.CL

    We present AIriskEval-edu Demo, a platform that audits the pedagogical quality of instructional explanations and provides explainable audit results. The platform evaluates an explanation against a rubric covering five dimensions of pedagogical risk: factual accuracy, depth and completeness, focus and relevance, student-level appropriateness, and ideological bias. For each dimension, it returns a binary decision and a confidence score....

    arxiv.org/abs/2607.25634 · PDF

  26. 26

    Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception

    Tian Zheng, Xurong Xie, Xinxin Zhu, Xiaolan Peng, Feng Tian

    cs.AI

    Decoding speech information directly from scalp electroencephalography (EEG) into text provides a potential non-invasive neural communication pathway for individuals with severe speech and motor impairments. Compared with invasive approaches such as electrocorticography, EEG is safer and more widely deployable, yet substantially more challenging to decode.This challenge is exacerbated for Chinese sentence decoding, which must handle a...

    arxiv.org/abs/2607.25626 · PDF

  27. 27

    Quotient Dynamics, Effective Curvature, and Implicit Bias in Positive Quadratic Networks

    Pengcheng Cheng

    cs.AI

    Positive quadratic networks admit the low-rank representation f_U(x)=x^top UU^top x, where Uinmathbb{R}^{dtimes r} is identifiable only up to right orthogonal multiplication, representing a rank-r PSD matrix Q=UU^top. We study how this quotient structure governs training dynamics, curvature, recovery, and interpolation bias. On the full-column-rank stratum, we identify mathbb{R}^{dtimes r}_*/O(r) with the rank-r PSD manifold. For smooth...

    arxiv.org/abs/2607.25624 · PDF

  28. 28

    Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines

    Federico Cabitza, Gianluca Colombo

    cs.AI · cs.HC

    Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epistemic evaluation, producing the condition they call *Epistemia*: the experience of possessing knowledge without undertaking the practices through which judgment would ordinarily be warranted. This article accepts that diagnosis but challenges its explanatory framework, which compares an embodied, socially...

    arxiv.org/abs/2607.25620 · PDF

  29. 29

    Multi-Sensor Alignment for Weather Simulations

    Samsad Alam, Devyani Lambhate, Aditya Mohan, Vishal Kumar, Vaibhav Katewa

    cs.AI

    Perception tasks for autonomous vehicles need to work satisfactorily in adverse weather conditions. Due to lack of real-world weather datasets, weather simulations are a promising alternative. To ensure simulations closely mirror real-world weather data, it's crucial that they represent the same weather characteristics, including severity and particle positioning, across different sensors. To achieve this, we propose the Reference Dataset...

    arxiv.org/abs/2607.25612 · PDF

  30. 30

    Computational Extraction of Legal Causes via al-Sabr wa al-Taqsim: A Set-Theoretic Formalization for Closed Fiqh Chapters

    Elnaser Abdelwahab

    cs.AI

    This paper presents a set-theoretic formalization of the classical usuli method of al-Sabr wa al-Taqsim (Examination and Division) for extracting legal causes ('ilal) within closed chapters of jurisprudence. A computational algorithm is introduced that extracts minimal operational rules from a truth table of juristic verdicts. The principal result is that, given a complete truth table for a closed chapter, the algorithm computes the minimal...

    arxiv.org/abs/2607.25605 · PDF

  31. 31

    A Density-Matrix Framework for Electronic-Structure Analysis of Functional-Group and Salt Effects in Lithium-Metal Electrolytes

    Mingkang Liu, Huize Yu, Yanbin Gao, Nan Yao, Xiang Chen, Lei Shen

    cs.AI

    The reactivity of lithium-metal electrolytes arises from the interplay of molecular functional groups, Li$^+$ solvation, and salt-anion participation. This interplay operates through the redistribution of electron density across donor, anion, and cation centers, which is most directly read out from the electronic structure resolved in space. Quantum-chemical calculations deliver such readouts faithfully, yet become computationally demanding...

    arxiv.org/abs/2607.25597 · PDF

  32. 32

    How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model

    Mahendra Singh Rathor, Anagheem Azzam

    cs.AI

    Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute budgets, yet their interaction is most often studied on billion-parameter models where the design space is expensive to explore. We ask a complementary question: on a specific, fully reproducible 60M-parameter encoder-decoder model (T5-small) and a single-table text-to-SQL benchmark (WikiSQL), how much task...

    arxiv.org/abs/2607.25583 · PDF

  33. 33

    Matrix-Free Photoacoustic Image Reconstruction via Sensor-Token Self-Attention

    Mary John, Shibili Said, Imad Barhumi, Sherzod Turaev, Mohamed Yahia

    cs.AI

    Photoacoustic tomography (PAT) combines the optical absorption contrast of biological tissue with the spatial resolution of ultrasound, yet recovering the initial pressure distribution from sparse-view sensor measurements remains an ill-posed inverse problem. Iterative compressive-sensing solvers and unrolled deep networks both retain a dependence on the system matrix at inference, which leaves real-time clinical reconstruction...

    arxiv.org/abs/2607.25576 · PDF

  34. 34

    Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories

    Jianing Geng, Ruiqi He, Zekun Fei, Biao Yi, Ruijie Wang, Zheli Liu, Xia Hu, Xuansheng Wu, Qingkai Zeng

    cs.AI

    Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and private deployment behind cloud-hosted agent interfaces, giving providers incentives to keep high-value skills proprietary. Yet hiding the artifacts does not conceal their behavioral effects, which remain observable in execution trajectories and form a behavioral side channel. We define this...

    arxiv.org/abs/2607.25560 · PDF

  35. 35

    Distilling Temporal Search and Reasoning: Evolving LLMs for Future Prediction via Harness-Assisted Efficient Data Synthesis

    Wanxu Cai, Zhengyu Chen, Huaisheng Zhu, Wei Wang, Jingang Wang, Qiang Xu

    cs.AI

    Future event prediction carries broad social impact yet remains challenging. SOTA approaches augment LLMs with external agent frameworks whose predictive capability vanishes once the harness is removed. While recent Tool-Integrated Reasoning (TIR) internalizes deep search for multi-hop retrieval of facts, forecasting further demands temporal search and reasoning over historical trends and dynamic shifts. The key obstacle is data: historical...

    arxiv.org/abs/2607.25554 · PDF

  36. 36

    From Training to Deployment: Post-Hoc Causal Feature Identification via Sensitivity Ratios

    Athanasios Vlontzos, Giorgos Papanastasiou, Bernhard Kainz, Sotirios Tsaftaris

    cs.AI

    Given a model that is already trained, which features does it rely on causally versus spuriously? Existing methods require access to the training procedure and cannot answer this post-hoc. We introduce the \textbf{Normalised Sensitivity Ratio~(NSR)}, a post-hoc, model-agnostic diagnostic for this question under a structured-shift regime: environments differ primarily in the mean of spurious features while the causal mechanism and causal...

    arxiv.org/abs/2607.25546 · PDF

  37. 37

    Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners

    Athanasios Vlontzos, Giorgos Papanastasiou, Bernhard Kainz, Sotirios Tsaftaris

    cs.AI

    Consider a model trained at a single hospital to predict patient recovery, where the measured feature $X$ bundles the patient's true health signal ($C$) with a systematic artefact from that hospital's equipment ($S$). Within that hospital, the artefact correlates with outcomes through unmeasured confounders such as patient demographics; an in-context learner rationally routes predictions through $S$, not $C$, and fails silently when deployed...

    arxiv.org/abs/2607.25532 · PDF

  38. 38

    Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?

    Qitao Chen, Dongfu Yin, F. Richard Yu

    cs.AI

    As neural network models for image classification advance, neurons play critical roles in pruning, backdoor defense, and interpretability. Yet existing work lacks clarity on the weight-importance relationship. We address this with a neuron importance assessment method using three experiments: quantifying overlap between high-weight and accuracy-impacting neurons, analyzing high-weight neuron perturbation effects, and testing post-retraining...

    arxiv.org/abs/2607.25529 · PDF

  39. 39

    CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

    Minhyeok Lee, Chiyoung Kim, Chanhoe Gu, Seongrok Kim, Sanghyuk Roy Choi, Donghwan Hwang, Donghun Ryu, Seokhyun Kim

    cs.AI · cs.CV

    Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets. We present CoTinyVLA, a 0.9B-parameter action model on a Qwen3.5-0.8B backbone that obtains that robustness by structuring supervision instead of enlarging the model. Three...

    arxiv.org/abs/2607.25487 · PDF

  40. 40

    PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

    Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Daniel...

    cs.AI · cs.CL

    Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks...

    arxiv.org/abs/2607.25485 · PDF

  41. 41

    Finding Optimal Cost-Bounded Plan Reductions: Refined Model

    Martha Del Toro, Raquel Fuentetaja, Angel García-Olaya

    cs.AI

    In some real applications a plan may later become unfeasible due to newly imposed budget constraints, yet, at the same time, using only the original actions of the plan and their order is mandatory. In this paper, we study the problem of extracting, from a precomputed plan, a valid subplan that maximizes utility while respecting a cost bound. Each goal is given a utility value and the plan is reduced by removing actions that support...

    arxiv.org/abs/2607.25484 · PDF

  42. 42

    Balancing multiscale similarity and cartographic constraints: A similarity-driven optimization framework for line generalization

    Pengbo Li, Haowen Yan, Xiaomin Lu, Binbin Lin

    cs.AI · cs.CG · eess.SY

    Cartographic generalization is essential for generating multiscale map representations by balancing information preservation and cartographic readability. However, automated generalization remains challenging because existing approaches often treat spatial similarity evaluation, cartographic constraints, and parameter optimization as separate processes, limiting adaptive and interpretable control across scales. This study formulates...

    arxiv.org/abs/2607.25474 · PDF

  43. 43

    TRWH: A Text-Driven Random Walk Heterogeneous GNN for Semantic-Aware Sparse Recommendation

    He Ma, Chen Liu

    cs.AI · cs.MM

    Graph Neural Networks (GNNs) and Large Language Models (LLMs) have each advanced recommendation systems by modeling structural and semantic signals, respectively. However, integrating their complementary strengths remains challenging, particularly in sparse settings where maintaining semantic precision is critical. We propose TRWH (Text-driven Random Walk Heterogeneous Graph Neural Network), a novel framework that fuses LLM-generated textual...

    arxiv.org/abs/2607.25471 · PDF

  44. 44

    Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

    Huan Chen, Xiang Song, Jian Jin, Pan Ren, Liang-Jie Zhang

    cs.AI · cs.LG

    Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol). IMACS (Intelligent Multi-Agent Collaboration System) separates the three into orthogonal, independently swappable layers. Classic organizational theory (Belbin roles, Mintzberg coordination, RACI...

    arxiv.org/abs/2607.25446 · PDF

  45. 45

    The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play

    Michael Macaulay, Harmony Bouabid, Guo Gen Ang, Sasha Shaw

    cs.AI · cs.CR · cs.CY

    Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptography, web exploitation, and binary exploitation. Large language models (LLMs) can now solve a growing share of challenges with minimal human input, raising urgent questions about fairness, the validity of rankings, and whether participation still delivers the learning that justifies the effort. This paper...

    arxiv.org/abs/2607.25425 · PDF

  46. 46

    Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering

    Noor Islam S. Mohammad, Uluğ Bayazıt

    cs.AI

    Knowledge-intensive multimodal question answering (KI-MMQA) sits at the intersection of three expensive primitives: long visual token sequences, dense retrieval over large external corpora, and full cross-modal fusion. Existing systems pay all three costs uniformly per query, even though only a small fraction of visual content and retrieved knowledge is actually relevant to any given question. We introduce SKIP (Salient Knowledge-Injected...

    arxiv.org/abs/2607.25422 · PDF

  47. 47

    A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain

    Debjyoti Paul

    cs.AI

    Production LLM agents are increasingly assembled from a frozen model wrapped in a harness: a prompt template, a tool set, a memory/retrieval layer, a planning strategy, and a verification policy. Two 2026 systems, Meta-Harness (Lee et al., 2026) and HyperAgents (Meta AI, 2026), show that this harness can itself be optimized or even self-rewritten by an agentic proposer -- at the cost of either an expensive code-search loop or unconstrained...

    arxiv.org/abs/2607.25415 · PDF

  48. 48

    Context Assembly as the Controlled Variable: A Control-Theoretic View of Harness Policies for Frozen LLM Agents

    Debjyoti Paul

    cs.AI

    A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al., "Stable Agentic Control", 2026), sample-complexity bounds for sparse policies over massive discrete tool universes (Majumdar, "Sparse Agentic Control", 2026), and regulatory-control decompositions of multi-agent systems into auditable feedback loops (Nogueira and Skogestad, 2026). We do not claim to...

    arxiv.org/abs/2607.25408 · PDF

  49. 49

    COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution

    Jincheng Wang, Min Zheng, Tao Wei

    cs.AI

    Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) that specify not only what outcome to achieve, but also which steps, branches, and tool interactions are permitted. When these instructions are supplied as prompt context, however, the model retains control over both procedure selection and step execution. As interactions accumulate, an agent can skip...

    arxiv.org/abs/2607.25400 · PDF

  50. 50

    HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

    Liudas Panavas, Sebastian Minus, Bradley Monton, Derek Ray, Suhaas Garre, Sushant Mehta, Edwin Chen

    cs.AI · cs.CL

    Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document actually constrains its behavior over an extended tool-use...

    arxiv.org/abs/2607.25398 · PDF

  51. 51

    Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

    Abu Bakar Siddik

    cs.AI

    Cyber-capable AI agents combine language models with tools, memory, and execution en- vironments to perform multi-step offensive-security tasks. Existing work separately measures cyber capability and catalogs attacks against agent components, but provides less guidance on containing a capable agent within the environments used to evaluate it. This review synthe- sizes five vulnerability classes at that boundary: multi-step offensive chains,...

    arxiv.org/abs/2607.25379 · PDF

  52. 52

    ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning

    Jiaqi Zhang, Tong Chen, Junliang Yu, Quoc Viet Hung Nguyen, Hongzhi Yin

    cs.AI

    Agentic systems have rapidly advanced in their ability to interact with real-world environments, leverage external tools, and provide services for users. However, unlike natural-world tasks that assume well-defined instructions, human-centered scenarios are characterized by ambiguous requests that lead to large, open-ended solution spaces. Decoding users' personalized preferences is therefore essential for narrowing the candidate solution...

    arxiv.org/abs/2607.25369 · PDF

  53. 53

    AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis

    Md Salahuddin, James Rooney, Fida Hasan

    cs.AI

    The intersection of artificial intelligence adoption, cybersecurity governance, and public sector institutional constraints has not been examined as a unified analytical problem in the existing literature. Studies address AI cybersecurity risks generically, public sector governance independently, and framework adequacy separately. Existing studies have not integrated these three streams to explain specifically how AI adoption causes...

    arxiv.org/abs/2607.25368 · PDF

  54. 54

    Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales

    Genliang Zhu, Chu Wang

    cs.AI · cs.SE

    Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection. We present Explanation-Bound Tool Execution (EBTE), a claim-carrying mediation layer that converts decision-relevant rationale content into typed action claims and checks them against server-held intent, policy, payload, tool, risk, provenance, and freshness facts. EBTE cannot widen baseline...

    arxiv.org/abs/2607.25364 · PDF

  55. 55

    Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management

    Sukju Oh, Moo-Yong Rhee, Jae-Sik Jang, Sukkyu Sun

    cs.AI · cs.CL

    The same episode of atrial fibrillation is a minor finding in a healthy adult and grounds for anticoagulation in an elderly patient with hypertension: identical signal, opposite decision. Naming the rhythm is only the start; what determines a patient's outcome is the judgement that follows -- what the arrhythmia is across the whole record, what it means for this patient, and what should be done about it. Recent work pairing large language...

    arxiv.org/abs/2607.25340 · PDF

  56. 56

    Dual-Domain Manifold Modeling for Hyperspectral Image Fusion

    Chengxin Xie, Qiya Song, Yangbangyan Jiang, Renwei Dian, Xudong Kang

    cs.AI

    Achieving a coherent integration of spectral richness and spatial fidelity remains a central objective in hyperspectral image fusion. However, existing hyperspectral image fusion methods struggle to effectively model geometric constraints. In the spatial domain, weak spatial-spectral interaction limits geometry-aware feature learning and suppresses high-frequency structural information, resulting in low-frequency bias and structural...

    arxiv.org/abs/2607.25338 · PDF

  57. 57

    From Cellular Responses to Pharmacological Domains: Multimodal Zero-Shot Drug Representation Learning

    Jintao Huang, Lu Leng, Ziyuan Yang

    cs.AI

    Multimodal drug discovery enables drug representation learning beyond chemical structure by incorporating cellular responses such as gene expression and cell morphology. However, direct fusion and instance-level contrastive alignment may mix mechanism-related signals with modality-specific noise and incorrectly separate structurally dissimilar but biologically related compounds. This limitation can obscure transferable mechanism patterns...

    arxiv.org/abs/2607.25322 · PDF

  58. 58

    Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision

    Ruijie Su, Yuanzhi Liang, Xiaohua Xie, Jianhuang Lai

    cs.AI

    Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liquid columns break apart in mid-air, container water levels fail to rise as liquid is poured in, and splashes disperse without regard to momentum or gravity. We attribute this gap to the fact that large-scale video-text corpora contain almost no explicit motion supervision, so models learn to imitate fluid...

    arxiv.org/abs/2607.25321 · PDF

  59. 59

    Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe

    Chaemin Jang, Dongman Lee, Jihee Kim

    cs.AI

    Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution. We show this draw does not exist: instruction-tuned models do not sample from distributions, they collapse to a single output. The same persona on the same question returns the same answer on more than half of items in a public-opinion benchmark. The collapse is sharp: the...

    arxiv.org/abs/2607.25292 · PDF

  60. 60

    ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design

    Jingbo Zhang, Haoxiang Sun, Wenbo Wang, Wenbo Zhang

    cs.AI · cs.AR · cs.CR

    This paper presents ContractHIL-HLS, a contract-aligned multi-agent workflow for practical high-level synthesis (HLS) engineering. The workflow makes three contributions. First, it introduces a structured contract as the semantic-alignment and task-execution artifact that translates natural language requirements into explicit interfaces, constraints, validation checks, and rollback rules. Second, it incorporates hardware information into the...

    arxiv.org/abs/2607.25283 · PDF

  61. 61

    Many-body Tipping Dynamics of ChatGPT-like AIs

    Frank Yingjie Huo, Neil F. Johnson

    cs.AI · cond-mat.dis-nn · math-ph · nlin.AO · physics.soc-ph

    Why do ChatGPT-like AIs, despite major architectural and training differences, unexpectedly tip to undesirable content (e.g. harmful, misleading, repetitive) even under deterministic greedy decoding? We show that a broad class of such tippings is caused by the many-body interactions between tokens (spins) as they cross the finite-layer system. Tipping emerges as a dynamical first passage process between competing output basins. Attention...

    arxiv.org/abs/2607.25279 · PDF

  62. 62

    The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape

    Deyao Hong, Kehan Zheng, Qian Li, Jun Zhang, Jie Jiang, Hongning Wang

    cs.AI · cs.IR

    Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the user. LLM-based user agents enable a different recommendation process: a user specifies a need before choosing a platform, leaving platforms to compete for the user's attention, which we refer to as an agentic recommendation market. In our controlled LLM-based experiments across three product...

    arxiv.org/abs/2607.25253 · PDF

  63. 63

    CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models

    Yixuan Duan, Arjun Naik, Sadeer Al-Kindi, Wei Qiu

    cs.AI

    Foundation models for 12-lead electrocardiograms (ECGs) transfer well across clinical tasks, but the physiological knowledge encoded in their representations remains opaque. We present CADENCE, a framework that decomposes an ECG foundation model into a human-interpretable, queryable dictionary of physiological concepts. Using a BatchTopK sparse autoencoder, CADENCE factorizes Layer-6 embeddings from more than nine million ECG tokens into...

    arxiv.org/abs/2607.25244 · PDF

  64. 64

    Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection

    Yuhang Yang, Kai Tang, Chao Ye, Haobo Wang, Qiqi Luo, Jinguang Zheng, Zhixin Zhang

    cs.AI

    Debt collection is a critical negotiation task in the financial industry, with strong practical relevance and exceptional academic value as a behaviorally rich, high-stakes testbed for human-centered dialogue systems. While large language models (LLMs) have shown promise in dialogue and negotiation, effectively evaluating their performance in this complex scenarios remains a major challenge: existing benchmarks uniformly assume users to be...

    arxiv.org/abs/2607.25218 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.