cs.AI · 2026-08-16 · No. 86

Artificial Intelligence, 2026-08-16.

54 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

54 entries
  1. 01

    OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

    Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu

    cs.AI · cs.CL

    Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal,...

    arxiv.org/abs/2608.13558 · PDF

  2. 02

    QuoteBench: How Matched Scores Can Hide Command-Path Failures

    Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang

    cs.AI · cs.SE

    LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately...

    arxiv.org/abs/2608.13547 · PDF

  3. 03

    AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)

    AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan...

    cs.AI

    This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the previous release, we substantially revise how conditioning signals are represented and integrated into the model. The new design is guided by a simple principle: conditioning signals should match the generated content as closely as possible in both latent...

    arxiv.org/abs/2608.13492 · PDF

  4. 04

    MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

    Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem, Jacinta Arnold, Shahriar Faghani,...

    cs.AI · cs.CL

    We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution. We additionally introduce a Decomposer module...

    arxiv.org/abs/2608.13476 · PDF

  5. 05

    A Unifying Perspective on Causal World Models: From Observations to Representations to Structure

    Avinash Kori, Fabrizio Russo

    cs.AI · cs.CV

    World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics. We argue that useful WMs must go beyond generative capabilities alone: they...

    arxiv.org/abs/2608.13456 · PDF

  6. 06

    Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension

    Alison R. Panisson, Maria Eduarda W. M. Vianna, Italo Firmino da Silva, Heitor Henrique da Silva, Rafaela Fernandes...

    cs.AI

    Academic leagues have become important mechanisms for promoting extracurricular education and strengthening the integration between universities and society. This paper presents the organizational framework adopted by the Academic League of Artificial Intelligence (LIA) at the Federal University of Santa Catarina (UFSC), designed to integrate teaching, research, and university extension through a student-centered, project-based approach. The...

    arxiv.org/abs/2608.13447 · PDF

  7. 07

    RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level

    Juan Irving Vasquez, Juan Terven, Laura-Ivoone Garay-Jimenez

    cs.AI

    Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to apply automatically: the adaptation of Technology Readiness Levels to AI lacks AI-specific gating criteria, the Machine Learning Technology Readiness Levels presuppose access to internal process artifacts, and AI/data readiness...

    arxiv.org/abs/2608.13428 · PDF

  8. 08

    Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes

    Aimilios Hadjiliasi, Louis Nisiotis

    cs.AI

    Embodied intelligent virtual agents are expected to operate as persistent, adaptive, and context-aware entities within complex virtual and Metaverse worlds. However, implementing cognitively capable agents in such environments is conceptually and technologically challenging. Among a range of blueprints and development approaches, the Cognitive Embodied Agent Architecture (CEAA) has been developed as an implementation-oriented framework for...

    arxiv.org/abs/2608.13420 · PDF

  9. 09

    Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

    Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu,...

    cs.AI

    Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36...

    arxiv.org/abs/2608.13417 · PDF

  10. 10

    Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings

    Mirko Tritella, Riccardo Pozzi, Matteo Palmonari

    cs.AI

    Parliamentary proceedings are a primary record of democratic deliberation, yet their volume and fragmentation make multi-perspective access difficult for citizens, journalists, and researchers. Applying Retrieval-Augmented Generation (RAG) to parliamentary transcripts introduces three specific risks: dominance of the most frequent speakers, inability to weight speakers according to topical expertise, and citation misattribution in politically...

    arxiv.org/abs/2608.13410 · PDF

  11. 11

    Jointly Predicting Courses and Grades Using a Transformer-Based Model

    Paul Savala

    cs.AI

    Existing predictive models in learning analytics often treat student academic history as a simple sequence, overlooking the concurrent nature of courses taken within a semester. This simplification can lead to inaccurate performance predictions, particularly for students with heavy or challenging course loads. This paper introduces a TRansformer for Academic Course-grade Estimation (TRACE) that addresses this limitation by jointly predicting...

    arxiv.org/abs/2608.13409 · PDF

  12. 12

    TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies

    Xiaokang Qu, Jianliang Ma, Zao Fan, Tianshu Chu, Tianlong Fan, Linyuan Lü

    cs.AI · cs.CR · cs.NI

    Enterprise security topology design requires translating business intent, regulatory requirements, and risk assumptions into zones, boundary devices, inter-zone paths, and access-control policies. Existing NetOps automation tools mainly operate after this design is fixed, providing limited support for generating structured security topologies from underspecified natural-language requirements. We present TopoIntent, a system that compiles...

    arxiv.org/abs/2608.13389 · PDF

  13. 13

    Rules or Character? Scaling Laws for AI Safety Design

    Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu, Ryuji Hamamoto

    cs.AI

    Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, safety classifiers), which blocks harmful outputs at inference time, yet little formal analysis exists on how their optimal balance should change as deployment scales increase. We introduce a...

    arxiv.org/abs/2608.13345 · PDF

  14. 14

    LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

    Yupan Ding, Jing Xiao, Zhenyuan Zhang, Chaofeng Chen, Liang Liao, Gui-Song Xia, Mi Wang

    cs.AI

    Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing...

    arxiv.org/abs/2608.13344 · PDF

  15. 15

    LLM-Guided Graph Generation for Structure-Based Local Improvement Methods

    Hai Xia, Vaidyanathan Peruvemba Ramaswamy, Stefan Szeider

    cs.AI

    Large neighborhood search normally selects a random subset of decision variables for iterative optimization. For efficiently solving different problems, researchers tend to design variable selection strategies by taking into account structural features from different domains. In this paper, we build an automatic pipeline that is problem-agnostic to all problems in the MiniZinc format. By prompting an LLM with our semantic guidelines, we guide...

    arxiv.org/abs/2608.13333 · PDF

  16. 16

    StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems

    Yanwen Peng, Delvin Ce Zhang, Xi Wang, Nikolaos Aletras

    cs.AI

    Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discrete bottleneck. Converting the sender's continuous hidden states into discrete tokens discards information that token identities alone cannot capture. Recent work proposes latent communication as an alternative, where agents transmit hidden representations directly without converting them to text. However,...

    arxiv.org/abs/2608.13317 · PDF

  17. 17

    NAS-Driven Hardware Accelerator Exploration for Edge AI and Quantization Effects on the Pareto Space

    Eleftherios Mylonas, Angelos Kouprizas, Michael Birbas, Alexios Birbas

    cs.AI

    Edge AI deployment demands neural architectures that are simultaneously accurate, computationally efficient, and hardware-deployable - a challenge addressed by hardware-aware Neural Architecture Search (NAS). While recent works incorporate quantization directly into the NAS loop, these approaches expand search complexity and tightly couple architecture and quantization design. The simpler post-search quantization strategy has received little...

    arxiv.org/abs/2608.13293 · PDF

  18. 18

    Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision

    Vayalet Stefanova, Diwas Lamsal, Margot Genbrugge, Maxim Yudayev, Christian Schlenstedt, Moran Gilat, Bart...

    cs.AI

    Understanding motion in daily living requires context beyond kinematics, because similar inertial patterns during activities of daily living (ADLs) can reflect intentional stopping, object interaction, or pathological movement impairment. Egocentric vision provides task-related context that may help disambiguate these cases. We investigate this challenge through freezing of gait (FOG) detection in Parkinson's disease (PD), a symptom strongly...

    arxiv.org/abs/2608.13283 · PDF

  19. 19

    Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability

    Alan Woodward, Andrew Rogoyski

    cs.AI · cs.CR

    A small number of firms based in two states produce the most capable frontier AI models. The governments of those states have shown both the legal power and the political will to decide which other countries may use these systems. In June 2026 the United States required a leading developer to obtain licences before releasing its most advanced models to any foreign person, including foreign nationals resident in the United States. The affected...

    arxiv.org/abs/2608.13272 · PDF

  20. 20

    vToken: Token-Level Virtualization for Reclaimable KV Caches

    Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li

    cs.AI · cs.DC · cs.OS

    Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragmentation, but recent KV eviction algorithms operate at a token granularity finer than block-level management. This mismatch causes intra-block fragmentation, leaving a large fraction of allocated KV memory unreclaimable. We present vToken, a...

    arxiv.org/abs/2608.13263 · PDF

  21. 21

    Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test

    Saveliy Batruin

    cs.AI · cs.LG

    Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. We model this failure with a finite \emph{capability sheaf}: stalks encode typed behavior signatures, restriction maps retain shared fields, and accepted runs are useful global sections. An exact finite constraint-satisfaction problem (CSP) defines acceptance, while a linearized relative cohomology...

    arxiv.org/abs/2608.13228 · PDF

  22. 22

    TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems

    Shunwen Bai, Ziping Ma, Chaoyang Zhang, Yarong Wang, Jiale Liu, Zhen Qin, Qingpei Guo

    cs.AI

    The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search. Prior process-level methods focus on the coherence and redundancy of chain-of-thought (CoT), and most benchmark tasks have a single objective solvable by static capabilities such as derivation and tool...

    arxiv.org/abs/2608.13221 · PDF

  23. 23

    Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents

    Zechuan Wang, Siyuan Lu, Hongxuan Zhang, Linjian Mo, Chenyi Zhuang, Leilei Gan

    cs.AI

    Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signal. On-policy distillation provides dense per-token supervision but is either teacher-bounded or prone to gradient concentration collapse. We introduce $\textbf{CrEST}$, a hierarchical credit...

    arxiv.org/abs/2608.13179 · PDF

  24. 24

    SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents

    Chang Liu, Yuqi Zhang, Yiman Zhong, Boyi Liu, Hengjun Wang, Shuyue Wei

    cs.AI

    Agent skills are crucial external instructions that enable language agents to execute long procedural tasks such as coding or document processing. Existing agent skills are primarily created through human manual crafting or agent execution traces, with limited understanding of how each step contributes to overall skill performance on specific tasks; i.e., there remains an open problem in quantifying the contribution of individual steps within...

    arxiv.org/abs/2608.13173 · PDF

  25. 25

    Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing

    Sheng Ren, Yadong Wang, Naiqiang Tan, Jiangang Kong, Jun Fang, Rui Liu, Jun Wang, Kai Chen, Lipeng Liang, Xiang Chen

    cs.AI

    Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and...

    arxiv.org/abs/2608.13156 · PDF

  26. 26

    Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement

    Aoxin Ni

    cs.AI

    Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain unreliable on elementary numerical tasks, including magnitude comparison, large-integer arithmetic, fractions, and scientific notation. This survey examines basic numerical understanding as a capability distinct from high-level mathematical reasoning. We propose the Numerical Grounding Framework (NGF), which decomposes numeracy into...

    arxiv.org/abs/2608.13129 · PDF

  27. 27

    SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

    Qianxi Yan, Chunrong Chen, Jiuzhou Zhao, Min Zhang, Yongzhou Xu, Xiaochuan Xu

    cs.AI

    Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the...

    arxiv.org/abs/2608.13120 · PDF

  28. 28

    Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting

    Huiyu Li, Weibo Liu, Xinru Xu, Dongchen Gao, Meng Zhang, Junhua Hu

    cs.AI

    Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding incomplete evaluations, and current fusion methods evaluate evidence sources exclusively through instantaneous comparisns without exploiting their long-term reliability across diverse decision contexts. This paper proposes a unified...

    arxiv.org/abs/2608.13108 · PDF

  29. 29

    Multi-Layer Context Camouflaging: A Semantic Superposition and Contextual Lamination Framework for Malpractice-Resilient Online Assessment

    Gupta Lovi Raj, Kaur Kamalpreet, Dama Sri Ram, Parani Prajithaa

    cs.AI · cs.CY · cs.HC

    Contemporary online assessment systems rely primarily on browser lockdown, webcam monitoring, and behavioural analytics, yet remain vulnerable to attacks that extract the assessment content itself through screenshots, screen sharing, optical character recognition, and automated scraping. This paper extends the Multi-dimensional Spatio-Temporal Context Camouflaging Model (MSCCM) within the MARS (Multi-modal Assessment Resilience Suite) by...

    arxiv.org/abs/2608.13100 · PDF

  30. 30

    SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference

    Divya Jyoti Bajpai, Kishan Kumar Upadhyay, Manjesh Kumar Hanawal

    cs.AI

    Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands. Deploying smaller LLMs directly on the edge can circumvent this, but with degraded accuracy. Deploying smaller cloud-based big LLMs preserves performance, but at the cost of expensive per-token computation. We present a distributed inference framework, \our{}, that...

    arxiv.org/abs/2608.13076 · PDF

  31. 31

    EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding

    Shuailei Zhang, Muyun Jiang, Wei Zhang, Jinbo Chen, Zhiwei Guo, Yong Li, Yi Ding, Cuntai Guan

    cs.AI

    Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. We propose EEG-PRIME, a two-stage EEG foundation model for cross-dataset multi-task decoding. EEG-PRIME combines masked pretraining with prototype-aligned instruction tuning to enable instruction-aware and subject-invariant decoding across diverse BCI paradigms. During...

    arxiv.org/abs/2608.13072 · PDF

  32. 32

    Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds

    Lucia Malíčková

    cs.AI

    Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming. Our objective is to induce a proactive, Socratic conversational framework, characterized by high-frequency question generation under strictly constrained high-performance...

    arxiv.org/abs/2608.13069 · PDF

  33. 33

    Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)

    Sam Mao

    cs.AI · cs.CL · cs.LG

    Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self-reported confidence - change as failure grows asymptotically rarer? We built a local, zero-cost harness on three open-weight models (qwen3:8b, llama3.1:8b, mistral:7b) running a repeated...

    arxiv.org/abs/2608.13063 · PDF

  34. 34

    Uniform Herding: Exemplar Replay with Representation Refresh

    Krishna Subedi

    cs.AI

    As the feature representation changes, replay must preserve the earlier classes. However, only a bounded active exemplar set can be replayed. We propose Uniform Herding, which allocates the current active set across observed classes and uses a bounded candidate pool to refresh their chosen exemplars in the current representation. On CIFAR-100 with ten class-incremental tasks, a ResNet-18 backbone, active budget $M=2{,}000$, retrieval budget...

    arxiv.org/abs/2608.13061 · PDF

  35. 35

    VALG: An Agentic System for ML Theory Research

    Dechen Zhang, Xuan Tang, Xinxiang Yin, Xingwu Chen, Jian Qian, Difan Zou

    cs.AI · cs.LG · math.OC · stat.ML

    Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. Solving an open problem therefore requires the problem formulation, theorem target, and proof mechanism to be developed in concert. Researchers formulate hypotheses, test them through preliminary theoretical or empirical...

    arxiv.org/abs/2608.13060 · PDF

  36. 36

    DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition

    Amogh Joshi, Animesh Mukherjee, Sergey Utyuzhnikov

    cs.AI

    In this work, we introduce DMDIntel which uses dynamic mode decomposition (DMD) to make the predictions made by LLMs in a classification task interpretable. It develops an input attribution pipeline, that first decomposes the hidden states of an LLM into prominent patterns, also known as modes, and then associates ranks to the input tokens based on the projection values on those modes. Rigorous experiments across three datasets and three...

    arxiv.org/abs/2608.13048 · PDF

  37. 37

    BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs

    Sanjeev Manivannan

    cs.AI · cs.CE · cs.ET

    Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-based multi-agent systems, humans typically provide an initial problem, agents deliberate internally, and the system returns a final response. BoardroomAI instead treats the human as a persistent participant who can intervene by challenging assumptions, modifying constraints, changing priorities, introducing...

    arxiv.org/abs/2608.13046 · PDF

  38. 38

    From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

    Xichen Ye, Yifan Wu, Zhikang Xie, Xiangyu Yue, Cheng Jin, Weizhong Zhang

    cs.AI · cs.CV · cs.LG

    Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we identify as being significantly misaligned with final generation quality. This discrepancy stems from the non-uniform propagation and accumulation of errors along the denoising trajectory. To...

    arxiv.org/abs/2608.13043 · PDF

  39. 39

    Foundations of MT-PDCL: Measure-Theoretic Probabilistic Definite Clause Logic

    Costin Bădică, Amelia Bădică

    cs.AI

    Standard probabilistic logic programming frameworks typically rely on grounding logic programs into discrete propositional representations. This operational requirement restricts exact inference to finite domains and discrete probability distributions. In this paper, we introduce Measure-Theoretic Probabilistic Definite Clause Logic (MT-PDCL), a generalized foundational framework that eliminates this finite-domain restriction. By explicitly...

    arxiv.org/abs/2608.13018 · PDF

  40. 40

    OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways

    Mao Jiayang, Wang Lanfeng, Peng Zhao-Han

    cs.AI · cs.MA

    Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constraints. This paper proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that is decoupled from a specific MARL algorithm. OGR-MARL integrates shared evader belief, role-conditioned option targets, adaptive rule penalties, and residual policy learning, allowing different...

    arxiv.org/abs/2608.12995 · PDF

  41. 41

    Moose: Latent concept learning with reasoning-shortcut awareness in $\mathcal{EL}^{++}$

    Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho, Robert Hoehndorf

    cs.AI

    The OWL 2 EL profile is used in some of the largest production ontologies, including the Gene Ontology and SNOMED CT. Existing neuro-symbolic (NeSy) learning methods accept propositional theories or Datalog, and reasoning-shortcut (RS) awareness has not been investigated in ontology settings. We present Moose, a method that compiles an $\mathcal{EL}^{++}$ TBox and finite ABox to a Sentential Decision Diagram (SDD). The SDD acts as a...

    arxiv.org/abs/2608.12961 · PDF

  42. 42

    Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses

    Lei You

    cs.AI · cs.LG

    Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint. We therefore track how the contrast develops as paired inputs are progressively revealed,...

    arxiv.org/abs/2608.12935 · PDF

  43. 43

    FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

    Zekai Li, Yihao Liang, Hongfei Zhang, Jian Chen, Yesheng Liang, Zhijian Liu

    cs.AI

    Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are...

    arxiv.org/abs/2608.12932 · PDF

  44. 44

    Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

    Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa

    cs.AI

    We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanning diverse medical specialties and visual domains, together with a text-only question answering (QA) control set. We evaluate Polish-oriented, general-purpose open-weight, and...

    arxiv.org/abs/2608.12928 · PDF

  45. 45

    Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence

    Varun Pratap Bhardwaj, Garima Singh, Arun Pratap Bhardwaj

    cs.AI · cs.MA

    Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0% of the missions on which either fails (log OR 6.66, 95% CI [6.38, 7.00]; phi 0.916), in a preregistered evaluation of 18,000 missions scored by deterministic code with no...

    arxiv.org/abs/2608.12895 · PDF

  46. 46

    Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals

    Jinhao Jing, Tian Zeyu, Lucas Qingyang Fang, Zhisheng Chen, Shuang Chen, Yuhao Luo, Qiannian Zhao

    cs.AI

    Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid intervention path as the predictive object of memory localization. PML separates random-calibrated target movement from semantic-neighbor and capability damage, and compares static localization...

    arxiv.org/abs/2608.12892 · PDF

  47. 47

    ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification

    Runze Zhao, Zixin Tang, Xiaoshuai Hao, Leyuan Chang, Xiaopeng Fu, Boyu Qiao, Dongyang Zhang

    cs.AI

    Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging. Recent methods primarily rely on multi-agent collaboration to decompose fact verification into specialized subtasks. However, these methods face two critical limitations: (1) agents may perform individual subtasks without sufficient awareness of the global...

    arxiv.org/abs/2608.12877 · PDF

  48. 48

    AI and Consumer Rights in India Working Paper

    Omir Kumar, Sriya Sridhar, Vibhav Mithal, Balaraman Ravindran

    cs.AI

    As AI systems proliferate in consumer facing applications, questions about liability for AI related harms remain unresolved. This working paper examines whether India's Consumer Protection Act, 2019, adequately addresses harm caused by defective AI products and services, and whether it proportionately allocates liability across the AI value chain. The Act's broad definitions of product liability, harm, and deficiency appear technology...

    arxiv.org/abs/2608.12863 · PDF

  49. 49

    Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

    Xutao Mao, Liangjie Zhao, Xiang Zheng, Cong Wang

    cs.AI

    Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by distilling operational trajectories into executable, transferable, and inspectable procedures. Because evolution optimizes task outcomes rather than procedure safety, compromised experience can cause skill...

    arxiv.org/abs/2608.12851 · PDF

  50. 50

    Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories

    Yifei Li, Heng Wang, Lingling Zhang, Muye Huang, Xinyu Zhang, Jiashuai Liu, Hang Yan, Rongman Xu

    cs.AI · cs.CL

    Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memory and formulate an evaluation framework that holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying the...

    arxiv.org/abs/2608.12847 · PDF

  51. 51

    CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation

    Yuchen Liu, Zongzhen Yang, Binhang Qi, Hailong Sun, Xiang Gao

    cs.AI

    Model merging has recently attracted significant attention as a promising paradigm for constructing unified multi-task models without requiring additional retraining. However, parameter conflicts and knowledge interference across tasks often degrade merged-model performance. Prior work introduced Conflict-Aware and Balanced Sparsification (CABS), which reduces parameter interference through structured pruning and sequential masking. However,...

    arxiv.org/abs/2608.12842 · PDF

  52. 52

    ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

    Jiale Cui, Yueyao Yuan, Kaixi Zhong, Xiaogang Xu, Jiafei Wu, Zhe Liu

    cs.AI

    The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes....

    arxiv.org/abs/2608.12788 · PDF

  53. 53

    PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs

    Sadat Shahriyar, Shareef Ahmed, Abdullah Al Arafat

    cs.AI

    Schedulability analysis is essential for certifying real-time systems, but existing tests are often developed through pen-and-paper proofs that are difficult to scale, validate, and maintain. Mechanized verification in PROSA/ROCQ offers a rigorous alternative, yet manually constructing such proofs requires substantial domain expertise and proof-engineering effort. Recent successes of large language models (LLMs) across a wide range of tasks...

    arxiv.org/abs/2608.12762 · PDF

  54. 54

    Correct Is Not Governed: Provenance Integrity in Agentic Workflows

    Jesus Salas

    cs.AI · cs.CR

    Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a correct action may rely on the wrong authority, an unsupported completion claim, or work made stale by a later change. We define governed execution as work whose decisions, completion, and response to change are supported by inspectable provenance. We present Matrix, a deterministic causal-state layer...

    arxiv.org/abs/2608.12761 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.