cs.AI · 2026-08-27 · No. 97

Artificial Intelligence, 2026-08-27.

29 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

29 entries
  1. 01

    Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings

    Evelyn Ma, Rama Kumar Pasumarthi, Kishwar Shafin, Mandar Sharma, Mimi Sun, Hamed Sadeghi, Dav M. Ebengo, Mbulayi...

    cs.AI · cs.LG

    Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling. However, building predictive planetary models remains bottlenecked by a fragmented data ecosystem, requiring manual data retrieval, multimodal data curation and fusion along with iterative model selection. We present the Planetary Prediction Engine (PPE), an autonomous AI...

    arxiv.org/abs/2608.26088 · PDF

  2. 02

    SwarmWorld: Stigmergic technological evolution in societies of language-model agents

    Subhadeep Pal, Fiona Y. Wang, Markus J. Buehler

    cs.AI · cond-mat.mtrl-sci · cs.CL

    Collective intelligence can emerge when individuals coordinate through a shared environment, allowing local actions to accumulate into durable social organization. Language-model agents offer a new substrate for this process, yet most multi-agent systems rely on direct conversation, predefined roles, or centralized workflows. It remains unclear whether decentralized agents can build functional technologies and outperform independent search....

    arxiv.org/abs/2608.26081 · PDF

  3. 03

    Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

    Srimonti Dutta, Akshata Kishore Moharir

    cs.AI · cs.CL

    Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integrity, a deployment reliability criterion for evaluating whether the computation recorded behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable. We identify the Structure Gap as the...

    arxiv.org/abs/2608.26036 · PDF

  4. 04

    Imitation Learning for Connection-Tableau Construction

    Fredrik Rømming, Mantas Bakšys, Martin S. Fixman, Sean B. Holden

    cs.AI · cs.LG · cs.LO

    An automated theorem prover builds a proof step by step, choosing at each point what to add and what to remove. We cast this construction as a policy acting in a transition system induced by a formal calculus, which fixes which steps are sound: for clausal connection tableaux, leanCoP-style search and plCoP/rlCoP-style planning then become stateful policies over one interface, and policy-learning methods apply directly. We equip such policies...

    arxiv.org/abs/2608.26009 · PDF

  5. 05

    AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs

    Sheng Liang, Yongyue Zhang, Nathanael Brian, Hang Lv, Hao Wang, Chen Zhang, Yong Liu

    cs.AI · cs.CL

    Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions. To control latency, deployments routinely compress inputs, but this degrades task accuracy. Speculative decoding (SD) accelerates generation losslessly, yet it assumes the drafter and verifier share an identical context, preventing SD from resolving the accuracy-overhead trade-off. We propose AsymSpec, an...

    arxiv.org/abs/2608.26004 · PDF

  6. 06

    ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs

    Somgyuan Li, Ahmed M. Abdelmoniem, Shiqiang Wang

    cs.AI · cs.MA

    Multi-agent large language model (LLM) workflows have emerged as a powerful paradigm for solving complex, open-ended tasks through collaborative reasoning among specialized LLM agents, but they incur substantial operating costs due to repeated LLM invocations and long-horizon context accumulation. Existing cascade routing methods make one-shot, query-level decisions and cannot adapt to the dynamic, state-dependent nature of multi-step...

    arxiv.org/abs/2608.25992 · PDF

  7. 07

    Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs

    Zongyu Wu, Yilong Wang, Xiaochen Wang, Minhua Lin, Zhichao Xu, Fenglong Ma, Xiang Zhang, Suhang Wang

    cs.AI

    Retrieval-augmented generation (RAG) is widely used to mitigate hallucination issues in large language models (LLMs) and multimodal large language models (MLLMs). In particular, knowledge graph (KG)-based RAG leverages structured knowledge to provide (M)LLMs with high-quality external information. Building on these works, recent studies have explored multimodal knowledge graphs (MMKGs) as knowledge bases for GraphRAG. This enables Graph RAG...

    arxiv.org/abs/2608.25986 · PDF

  8. 08

    SciMIF: Understanding Multimodal Instruction Following in Scientific Domains

    Ye Shen, Yuting Zheng, Dun Pei, Zijian Chen, Wenlong Zhang, Qi Jia, Guangtao Zhai

    cs.AI · cs.LG

    Understanding instruction-following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development of scientific fields. In this work, we introduce SciMIF, a novel benchmark designed to evaluate the capability of MLLMs in following complex scientific instructions. Specifically, based on an extensive analysis of 22 distinct tasks across 5 representative scientific...

    arxiv.org/abs/2608.25973 · PDF

  9. 09

    Quantitative Analysis of $ω$-Regular Robust MDPs

    Ali Asadi, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Mehrdad Karrabi, Alipasha Montaseri, Ali Shafiee

    cs.AI

    Robust Markov Decision Processes (RMDPs) generalize classical MDPs by allowing uncertainty in transition probabilities and optimizing against their worst-case realization. We consider $(s,a)$-rectangular RMDPs with \emph{linearly defined} uncertainty sets and study parity objectives, which are a canonical representation of $ω$-regular objectives. An uncertainty set is linearly defined if it is described by linear inequalities over the...

    arxiv.org/abs/2608.25968 · PDF

  10. 10

    LivingRAG: Augmenting Graph RAG with Experience

    Yuzhuo Cui, Zongye Zhang, Qingjie Liu

    cs.AI

    Graph-based RAG improves multi-hop question answering by organizing evidence as a knowledge graph. However, most existing RAG systems process each query in isolation and discard useful reasoning from the LLM's response after inference. As a result, later related queries need to retrieve evidence and reason from scratch. We propose LivingRAG, a Graph RAG framework with writable and reusable reasoning experience. LivingRAG adds a writable...

    arxiv.org/abs/2608.25960 · PDF

  11. 11

    Candidate supply and answer selection shape the value of LLM judging in multi-agent systems

    Jia-Hao Ji, Sijie Li, Jiabei Cheng, Zixi She, Jin-Tai Yu, Zhiyuan Yuan

    cs.AI · cs.MA

    Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is difficult because generation, communication and final answer-selection rules usually change simultaneously. We conceptualize multi-agent reasoning as an evolutionary pipeline of candidate generation, peer communication and terminal selection, wherein consensus without quality control can exhibit...

    arxiv.org/abs/2608.25937 · PDF

  12. 12

    How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation

    Aida Usmanova, Zangir Iklassov, Markus Leippold, Ricardo Usbeck

    cs.AI · cs.LG

    Automated fact-checking (AFC) systems retrieve evidence and predict claim veracity, yet evaluations omit simple baselines, systems are developed for a single benchmark and cannot be trusted to generalise across domains. No prior work cross-evaluates the full two-stage retrieve-then-verify pipeline across diverse datasets, complementing retrieval-only studies (Thakur et al., 2021) and single-stage benchmarking studies (Calamai et al., 2025)....

    arxiv.org/abs/2608.25934 · PDF

  13. 13

    Formal, Executable and Explainable Runtime Monitoring of Spoken Air Traffic Control Operational Procedures

    Roberto Luvini, Giacomo Longo, Alessandro Armando, Enrico Russo

    cs.AI · cs.CL · eess.AS

    Air traffic control procedures are executed through spoken exchanges between controllers and pilots. These interactions are essential to the safety of air transportation: failures in their execution can create severe operational hazards, as evidenced by past fatal accidents. Assessing whether an instruction has been followed requires relating what was said to the aircraft concerned, its state, and the obligations that pilots must meet. We...

    arxiv.org/abs/2608.25926 · PDF

  14. 14

    Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems

    Zhongwen Luan, Xiaoyu Zhang, Ming Hu, Yue Yang, Jiongchi Yu, Xiaohong Chen

    cs.AI · cs.SE

    As large language model (LLM)-based multi-agent systems (MASs) are increasingly applied to long-horizon complex tasks, their reliability has emerged as the core bottleneck hindering their real-world deployment. Existing MAS debugging and repair methods typically rely on rerunning and resampling the entire execution trajectory. However, a fundamental question remains to be answered: do these methods causally repair MAS failures or merely...

    arxiv.org/abs/2608.25920 · PDF

  15. 15

    Choose Your Game Wisely: Measuring Game-Theoretic Structures in Real-World Vehicle Interactions

    Yueyuan Li, Rongcheng Nie, Weijie Xi, Mingyang Jiang, Songan Zhang, Hanyang Zhuang, Ming Yang

    cs.AI · cs.RO

    Game-theoretic models provide principled frameworks for modeling vehicle interactions, but their underlying temporal assumptions have not been systematically examined against real-world driving behavior. In particular, it remains unclear how simultaneous, sequential, and asymmetric interaction structures can be measured from vehicle trajectories. This paper develops a trajectory-based interaction measurement framework to identify interaction...

    arxiv.org/abs/2608.25917 · PDF

  16. 16

    LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI Agents

    Weiming Li, Helen Paik, Yulei Sui

    cs.AI

    Modern GUI-agent frameworks achieve strong desktop task performance with frontier API models, yet persistent control information often remains implicit in growing interaction trajectories. At each step, the planner reconstructs the active task stage, accumulated evidence, and runtime feedback before deciding the next action. This dependence becomes more pronounced under weaker local reasoning backbones. Across four representative...

    arxiv.org/abs/2608.25777 · PDF

  17. 17

    ToST: A Tree-of-Thought Socratic Teaching Framework for Multi-Path Guidance and Parallel Thinking

    Feng Ling, Heng Yu

    cs.AI

    Large Language Models (LLMs) exhibit strong problem-solving abilities, positioning them as promising agents for Socratic teaching to guide students through step-by-step heuristic questioning. However, existing approaches typically adopt a one-problem-one-solution paradigm, restricting the teaching guidance to a single linear reasoning path. This design limits instructional flexibility, weakens error recovery, and restricts students' ability...

    arxiv.org/abs/2608.25775 · PDF

  18. 18

    Narcissus: Program Synthesis Using Context-Aware LLM Approximations

    Tilman Hinnerichs, Sebastijan Dumancic, Neil Yorke-Smith

    cs.AI · cs.LG · cs.PL · cs.SE

    Large language models (LLMs) excel at programming, but not when the task fixes the target language: prompted with a grammar rare in their training data, their programs usually break the grammar or fail the given specification. Enumerative synthesizers search the space of syntactically correct programs systematically guided by LLMs; the state of the art guides them by approximating LLM proposals into rule frequencies, which loses where each...

    arxiv.org/abs/2608.25657 · PDF

  19. 19

    Using profiles of cognitive capability to assess AI suitability for workplace tasks

    Jonathan Prunty, Marko Tešić, Patrick Quinn, José Hernández-Orallo, Lucy Cheke

    cs.AI · cs.CY · cs.HC

    Organisations deploying AI face a scoping problem: which tasks can be automated, which should remain with humans, and which are best shared between the two. Aggregate benchmark scores provide little insight into where systems will succeed or fail in practice, while human judgements of model capabilities quickly become outdated. We introduce a pipeline that profiles agents and tasks using a shared set of core cognitive capabilities. Cognitive...

    arxiv.org/abs/2608.25623 · PDF

  20. 20

    Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

    Pengfei Zhou, Hexin Wang, Zhengfeiyang Zhang, Yixing Ma, Zhenglin Wan, Kaipeng Zhang, Wangbo Zhao, Yang You

    cs.AI

    A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. The success of code agents illustrates why this matters. As code is executable, compilers and runtimes can provide high-quality rewards for Reinforcement Learning (RL) post-training of LLMs. By contrast,...

    arxiv.org/abs/2608.25518 · PDF

  21. 21

    CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

    Zhiyuan Li, Linyuan Gao, Xuechun Ding, Hongwei Chen, Yuan Wu, Yi Chang

    cs.AI · cs.CL

    Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they also turn memory access into a challenging retrieval problem. Full-library prompting preserves coverage at high context cost, vector retrieval returns compact neighborhoods but treats skills as independent text, and graph-based retrieval can recover workflow context only when the edges that carry relevance are reliable. We...

    arxiv.org/abs/2608.25500 · PDF

  22. 22

    PonsRAG: A Pons-Inspired RAG Bridging Cognitive Islands for Coordinated Long Narrative Reasoning

    Rongchen Zhao, Yu Chen, Juyuan Wang, Zhouting Mo, Jianxing Yu, Wenqing Chen, Jingping Liu

    cs.AI · cs.CL

    Long Narrative Reasoning is an essential capability for processing and reasoning over complex narratives. While retrieval-augmented generation provides a promising framework, existing methods still face two critical challenges: cognitive islanding and cross-layer evidence disconnection. To address these issues, we propose PonsRAG, a coordinated RAG framework inspired by the biological pons. PonsRAG consists of two key components: Triple-Layer...

    arxiv.org/abs/2608.25486 · PDF

  23. 23

    Training Alignment Auditors via Reinforcement Learning

    Paul Rosu, Rowan Wang

    cs.AI · cs.LG

    Alignment auditing of frontier models increasingly relies on LLM auditors to surface undesirable behaviors at scale, but current automated auditors can struggle with coherent investigation and audit realism. In this work, we improve LLM auditors with reinforcement learning. In our best training environment, the policy investigates target models that potentially possess hidden behaviors planted via their system prompt. An LLM judge, which...

    arxiv.org/abs/2608.25460 · PDF

  24. 24

    Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness

    Yi Chen, Hanna Hsieh, Shuhong Liu, Chuanbo Hua, Zihan Ma, Kun Wang, Joo-Young Kim

    cs.AI · cs.LG

    Machine unlearning aims to make a model forget specific data, yet unlearned LLMs often fail to stay unlearned: brief fine-tuning can revive removed knowledge. Existing robustness predictors rely on global weight-space displacement, but distance alone can be misleading when random or destructive updates collapse performance. We argue that relearning robustness depends on update structure: robust unlearning should affect forget-critical weights...

    arxiv.org/abs/2608.25429 · PDF

  25. 25

    Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents

    Shudong Liu, Dongyang Chen, Enci Zhang, Jinwei Liang, Zheng Ma, Lewei Lu

    cs.AI

    Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software...

    arxiv.org/abs/2608.25417 · PDF

  26. 26

    Can your AI agent be cheaper? Investigating the effects of task specifications on token spend in agentic coding tasks

    Jakub Smékal

    cs.AI

    Agentic coding workflows are now widely deployed in real-world systems. With long-horizon reasoning and tool use, token usage has become an important consideration for both cost and efficiency. Two engineers using AI will solve the same problem differently. How the specification of a task shapes an agent's token spend, and whether that spend can be predicted in advance, are open questions. Here, we study the effects of different task...

    arxiv.org/abs/2608.25399 · PDF

  27. 27

    Where vs What: Decomposing Structural and Content Failures in LLM-Generated Structured Outputs

    Yiwei Zhang, Chengke Wu, Li Wang, Jianqiang Li

    cs.AI

    Structured outputs such as JSON and tables are central to modern LLM-based systems, yet generation failures are evaluated monolithically, conflating two distinct error modes: placement errors (correct values at wrong positions) and value errors (wrong values at intended positions). We introduce Structure-Content Decomposition (SCD), a framework that independently measures structural fidelity and content accuracy. Applying SCD to nested JSON...

    arxiv.org/abs/2608.25358 · PDF

  28. 28

    Learning What to Share and What to Personalize: Hierarchical Strategy Co-Evolution for Agent Memory

    Yupeng Han, Shuochen Liu, Kai Zhang, Ze Liu, Zhihong Pan, Xianquan Wang

    cs.AI · cs.CL

    Memory-augmented agents maintain compact user profiles throughout extended conversations, enabling personalized and consistent responses without the need to process the entire dialogue history. The quality of these user profiles relies on the underlying memory management strategy: at each step, the agent must determine what to retain, compress, or discard. However, existing methods typically employ a static, one-size-fits-all strategy...

    arxiv.org/abs/2608.25329 · PDF

  29. 29

    FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

    Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang

    cs.AI · cs.CL

    Deploying large language models for professional financial review requires more than measuring general financial competence: models must perform the specific review operation required by a workflow and determine whether available evidence is sufficient for a defensible decision. Existing financial benchmarks cover knowledge, reasoning, compliance, and professional tasks, but their evaluation units are often organized around datasets or task...

    arxiv.org/abs/2608.25325 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.