cs.AI · 2026-08-20 · No. 90

Artificial Intelligence, 2026-08-20.

37 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

37 entries
  1. 01

    Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication

    Ramneet Kaur, Pradyumna Chari, Ramesh Raskar, Jugad Singh, Sumit Kumar Jha, Anirban Roy

    cs.AI · cs.CR

    Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, creating opportunities for covert harmful coordination. We introduce Verifiable Latent Alignments (VLA), an activation-aware framework for monitoring and steering these private communication channels. For every monitored decision, VLA links the private latent-state record and channel status to the resulting public action using a...

    arxiv.org/abs/2608.19161 · PDF

  2. 02

    Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems

    George Andrikopoulos

    cs.AI · cs.CY · cs.LG · cs.SE

    Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision: how tightly concentrated their outputs are around that target across repeated, identical requests. Borrowing the marksman's...

    arxiv.org/abs/2608.19140 · PDF

  3. 03

    Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering

    George Andrikopoulos

    cs.AI · cs.SE

    When an expert corrects an LLM assistant's error, the correction usually dies with the session, and the error class returns. I argue this is an operations problem, not a tooling problem: mechanisms for persisting corrections exist and are shipping, but the discipline for governing them -- versioning with provenance, recurrence monitoring, counter-metrics, retirement of stale rules -- does not. Writing as a systems engineer of thirty years, I...

    arxiv.org/abs/2608.19125 · PDF

  4. 04

    Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk

    Deep Kumar Ganguly, Jan Křetínský

    cs.AI · cs.LG · stat.ML

    An agent still learning its environment should be cautious while ignorant and bold once confident. The entropic value-at-risk captures this through a robust-optimization identity---a confidence level fixes the radius of a relative-entropy ball of alternative models---but that ball cannot reach catastrophes the nominal deems impossible, precisely what a safe agent must hedge. We instead use an optimal-transport ball and study the coherent risk...

    arxiv.org/abs/2608.19073 · PDF

  5. 05

    What is Missing from AI Post-Training AI: An Empirical Analysis

    Joy Jia Yin Lim, Xin Huang, Hao Peng, Yaxi Lu, Xin Cong, Zhong Zhang, Maosong Sun, Yankai Lin

    cs.AI · cs.CL · cs.LG

    Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch training, evaluate checkpoints, and improve downstream performance, raising the prospect of AI-for-AI. We argue that this picture conflates two distinct capabilities: execution-level capability, iterating within a selected training strategy; and strategy-level capability, revising the high-level judgment as experimental evidence accumulates....

    arxiv.org/abs/2608.19072 · PDF

  6. 06

    Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery

    Alizer Wong, Heng Cui, Yi Tan, Xiongchao Zhan, Liang Lin, Yuxiang Guo, Zhaorong Dai, Zixin Zeng, Wenyuan Li

    cs.AI · math.NT

    We present Eureka, a task-conditioned Meta-Agent architecture that compiles long-horizon tasks into dynamic obligation graphs with explicit acceptance semantics. During execution, Eureka forms Macro-Agents with specialized state, memory, operators, tools, verifiers, and local topology via receding-horizon planning, architecture promotion, and minimal-sufficient compilation. When bottlenecks recur, cost-benefit-gated evolution updates the...

    arxiv.org/abs/2608.19047 · PDF

  7. 07

    Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering

    Pradeep Murugesan, Luoxiao Yang, Xueli Chen, Xinqi Fan

    cs.AI · cs.CL · cs.MA

    Accurate and responsible medical question answering (QA) is important in healthcare, where complex cases require factual knowledge and nuanced reasoning. Existing medical QA systems, typically based on single-agent architectures and static retrieval, often lack adaptability, persistent memory, and structured decision-making. This work introduces an adaptive memory and reflection (AMR) agentic system, a multi-agent framework in which...

    arxiv.org/abs/2608.19029 · PDF

  8. 08

    Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models

    Valentin Romanov, Monique Bax, Steven Niederer

    cs.AI · cs.DB

    Accurately extracting nuanced, contextualized data from research articles is laborious and time intensive. Here, we investigate the performance of frontier, browser-based large language models (LLMs) to extract highly contextualized information. We demonstrate four escalating workflows, 1) given an expert curated prompt and research articles, most frontier LLMs perform well at data extraction, however can struggle with interpreting scientific...

    arxiv.org/abs/2608.19025 · PDF

  9. 09

    A Theory of Post-hoc Debate Judgement

    Xiang Yin, Adam Dejl, Antonio Rago, Lihu Chen, Francesca Toni

    cs.AI

    Debates have recently emerged as a useful methodology for agentic AI to improve performance as well as to aid explainability and user engagement. For example, LLM-empowered agents may debate internally (with themselves) and/or externally (with other agents). In many settings where debates are used, debates' outcomes and resulting outputs are determined post-hoc by external judges, often LLMs. In this paper we develop and test a novel theory...

    arxiv.org/abs/2608.19002 · PDF

  10. 10

    Breaking the weakest link to evade vision language models

    Ilan Zini, Boussad Addad, Katarzyna Kapusta

    cs.AI · cs.LG

    Vision Language Models (VLMs) have recently emerged as a critical component of multimodal AI systems, enabling joint reasoning over visual and textual inputs in real-world and safety-critical applications. Despite their growing deployment, the robustness of VLMs against adversarial threats remains insufficiently explored, particularly in the context of evasion attacks targeting multimodal alignment. In this work, we investigate the...

    arxiv.org/abs/2608.18938 · PDF

  11. 11

    \textsc{TestifAI}: Tomography-Based Testing for Deep Learning Systems

    Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis

    cs.AI

    As AI systems are increasingly deployed in safety-critical application domains (e.g., autonomous driving), associated risks increase too. Deep learning models underlying modern AI systems, therefore, must undergo thorough testing to ensure their correct behaviour. A single robustness test involves thousands of inferences to empirically verify if a model's outputs remain stable under a bounded perturbation of its inputs. However, existing...

    arxiv.org/abs/2608.18900 · PDF

  12. 12

    Syntactic Simplification of OWL Class Expressions

    Alkid Baci, N'Dah Jean Kouagou, Caglar Demir, Axel-Cyrille Ngonga Ngomo

    cs.AI

    Class expression learning often produces complex OWL class expressions that are difficult to interpret and reason over. However, by following theoretically grounded simplification principles, this complexity can be reduced. In this paper, we propose Class Expression Simplifier (CES), a novel algorithm for the syntactic simplification of class expressions in Description Logics (DL). CES aims to preserve formal semantics while reducing...

    arxiv.org/abs/2608.18899 · PDF

  13. 13

    Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models

    Wei Yu, Suxing Liu, Minjie Yu, Jiahao Wang, Zhijian Zheng, Haocheng Deng, Bing Li

    cs.AI

    Reinforcement-learning training of reasoning LLMs (e.g., GRPO) is expensive and requires a controllable environment, committing every contribution to a full training pipeline. We present EvoResearcher, a training-free, inference-time protocol that adds cost-bounded self-reflection to a single frozen LLM backbone. The protocol iterates generate -> self-critique -> revise until a maximum depth D is reached or the critique returns the CONFIRMED...

    arxiv.org/abs/2608.18884 · PDF

  14. 14

    DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning

    Zijie Meng, Xiwei Dai, Yixuan Tang, Jin Hao, Yang Feng, Fudong Zhu, Xiaoqiang Liu, Shaosheng Cao, Zuozhu Liu

    cs.AI · cs.MA

    Oral diseases affect billions of people worldwide, underscoring a pressing need for accurate and reliable dental assessment that integrates heterogeneous evidence from domain knowledge, radiographs, intraoral photographs, and 3D dental data. Most existing dental AI systems remain modality- or task-specific. Although recent vision-language models support flexible dental question answering, directly generated response leaves evidence implicit...

    arxiv.org/abs/2608.18878 · PDF

  15. 15

    SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents

    Qingyao Li, Wenxiang Jiao, Shuai Shao, Kangning Zhang, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang, Yong Yu

    cs.AI

    Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no existing signal trains it. We show that the default remedy, outcome-rewarded RL over the candidate slate, cannot teach it, for a structural reason we identify and name...

    arxiv.org/abs/2608.18852 · PDF

  16. 16

    ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery

    Mingzhao Li, Chenxi Liu, Yan Zhao, Hao Miao

    cs.AI

    Last-mile delivery aims to handle dynamically arriving orders with couriers while modeling complex spatial and temporal correlations. Recent learning-based methods model spatiotemporal dependencies among orders to predict courier service sequences, but leave next-order decision making unexplained. Describing the current delivery state in language allows LLMs to reason explicitly about the spatial, temporal, and behavioral cues behind an...

    arxiv.org/abs/2608.18846 · PDF

  17. 17

    Verifiable abstention makes AI leak diagnosis accountable in water distribution networks

    Tianwei Mu, Yue Wang, Mingzhe Yuan, Manhong Huang, Wenhong Wang, Xuerui Yin, Qing Luo, Min Xiao, Hui Yang, Jun Li, Dan Xue

    cs.AI

    Utilities lose a substantial share of treated water to leakage, yet rarely trust artificial-intelligence localizers to dispatch crews: guessing everywhere cannot justify excavation. The gap is accountability, not accuracy: no method proves when it should not act. Here we recast leak localization as decision-making under verifiable abstention. A physics-grounded executor agent falsifies hypotheses (leak, demand, sensor, valve) against a...

    arxiv.org/abs/2608.18836 · PDF

  18. 18

    Pairwise Logical Selection of Enthymeme Completions under Semantic-Link Uncertainty

    Xuyao Feng, Antonis Bikakis

    cs.AI

    Arguments often omit premises or claims, forming enthymemes. We study pairwise logical selection between two candidates for the omitted component. Existing natural language methods can identify or generate candidates but often do not expose how the selected candidate completes the inference, while logic-based approaches usually assume that the required formulae and background knowledge are available. We extend a prior neuro-symbolic pipeline...

    arxiv.org/abs/2608.18820 · PDF

  19. 19

    Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots

    Xing Zhang, Yanwei Cui, Guanghui Wang, Zhihao Lin, Peiyang He

    cs.AI · cs.CL · cs.SE

    Agents improve quickly against a reliable automatic metric and stall without one, and the applications that need them most, report generation among them, are the ones nobody knows how to score. Can the metric write itself? Saying what makes an answer good is hard; pointing at something wrong with one is easier, so the metric we evolve is a pool of small Python operators that each flag a candidate for one named defect, or abstain, and vote....

    arxiv.org/abs/2608.18744 · PDF

  20. 20

    A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation

    Manoj N M, Vijayakrishna S, Manjunath Srinivas, Rohit Pahan

    cs.AI

    This paper proposes a multi-agent framework built on CrewAI [1] for conversational business intelligence. Five specialized AI agents operate in a sequential pipeline to process natural language queries, retrieve and analyze data, generate visualizations via the Model Context Protocol (MCP) [2], and deliver actionable insights. The platform features a defense-in-depth security architecture for multi-tenant data isolation and a query...

    arxiv.org/abs/2608.18740 · PDF

  21. 21

    Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization

    Chenle Chen, Yangbo Wei, Chao Yao, Shaoqiang Lu, Junhong Qian, Chen Wu, Lei He

    cs.AI

    Text-space skill optimization adapts a frozen agent by evolving a natural-language skill document, accepting each candidate through a validation gate. Existing gates rely on verifiable rewards, confining these methods to tasks with an automatic verifier. Replacing the verifier with an LLM-judge gate would lift that restriction, but whether such a gate carries usable signal is untested. We ask a prior question: can we tell, before placing a...

    arxiv.org/abs/2608.18719 · PDF

  22. 22

    RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training

    Yugu Li, Jimmy Cao, Jianglin Qiao, Siyi Hu

    cs.AI

    Training multi-turn agentic workflows with reinforcement learning (RL) enables large language models to perform complex reasoning, use external tools, and conduct iterative search beyond single-turn settings. Yet multi-turn RL training remains highly unstable, often causing severe performance degradation as the number of turns increases. Through theoretical analysis, we identify three tightly coupled sources of instability: rollout-training...

    arxiv.org/abs/2608.18682 · PDF

  23. 23

    Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction

    Zhaoxi Wei, Hongye Yang, Shuyuan Tian

    cs.AI · cs.CY

    Amid concerns that generative AI may standardize art interpretation, this paper examines whether LLM-based interaction can support plural art-historical narrative construction. We present Sanyu Studio, a multi-agent dialogue system that models 321 Sanyu oil paintings as agents with fact, interpretation, organization, and memory-filtering mechanisms. Based on a seven-day workshop with eight art-university participants, the study shows that...

    arxiv.org/abs/2608.18677 · PDF

  24. 24

    Candidate-Fate Accounting for Transparent Sensor Diagnostic Pipeline Search

    Haotao Xie, Yutian Chen, Yangqi Liu, Xiaoyu Jiang

    cs.AI

    Industrial sensor diagnostics relies on preprocessing, representation, and classification pipelines, making automated pipeline search useful for reducing manual design cost. However, existing automated machine/deep learning (AutoML/AutoDL) reports typically retain only fitted trials, scores, and winners, omitting generated candidates that are invalid, pruned, skipped, cached, or unfitted. This omission limits reviewers' ability to check...

    arxiv.org/abs/2608.18665 · PDF

  25. 25

    Preference Reasoning under Indeterminacy in Large Language Models

    Hadi Hosseini, Samarth Khanna, Xiyuan Wang

    cs.AI · cs.GT · cs.LG

    As large language models evolve into decision-making agents, the ability to reason over preferences becomes fundamental to alignment, coordination, and collective intelligence. Yet, unlike standard benchmarks, real-world preference reasoning is inherently indeterminate: information may be incomplete, and valid solutions may not exist. We argue that indeterminacy, rather than correctness alone, is a central challenge for AI reasoning. We...

    arxiv.org/abs/2608.18631 · PDF

  26. 26

    CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence

    Yutong Cheng, Changze Li, Qian Cui, Wei Ding, Lingzhi Wang, Yan Chen, Peng Gao

    cs.AI · cs.CR

    Cyber threat intelligence (CTI) is increasingly consumed not by human analysts but by LLM agents that compose multi-step investigations at query time. The harness side of this shift has matured rapidly (planning loops, tool protocols, context management), but the corpus side has not: threat reports and vulnerability databases are still packaged for retrieval-augmented generation, as opaque chunks behind an embedding index. We argue that this...

    arxiv.org/abs/2608.18613 · PDF

  27. 27

    Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference

    Zishan Ahmad, Vishal Vaddina

    cs.AI · cs.CL

    Uniformly allocating inference reasoning budgets to LLMs is expensive and prone to over-thinking penalties; especially in document tasks where visual layouts drive complexity. To address this, we introduce BudgetDoc, the first multimodal benchmark providing explicit supervision for model-budget-performance trade-offs across three document tasks. Using BudgetDoc, we train DRB (Document-Reasoning Balancer), an approx. 1B-parameter pre-flight...

    arxiv.org/abs/2608.18591 · PDF

  28. 28

    FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

    Kou Shi, Zun Wang, Qisheng Su, Shiting Huang, Ziao Zhang, Zhen Fang, Qingnan Ren, Jin Liu, Yu Zeng, Yiming Zhao, Lin...

    cs.AI · cs.PL

    Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assumptions, the resulting task may be unsolvable or incorrectly evaluated. Meanwhile, multi-stage synthesis can discard the goals, dependencies,...

    arxiv.org/abs/2608.18580 · PDF

  29. 29

    Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement

    Mandar Kulkarni, Pooja A., Samir Shah

    cs.AI

    Modern e-commerce platforms often operate search, recommendation, personalization, and CRM systems independently, limiting opportunities for proactive customer re-engagement. This is particularly challenging for exploratory intents such as best smartphones or latest 5G phones, where users may leave the platform for external research before purchasing. We present a scalable, production-deployed framework that bridges search and CRM workflows...

    arxiv.org/abs/2608.18543 · PDF

  30. 30

    FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

    Pratik Ghawate

    cs.AI · cs.IR

    Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase orders, approvals, allocations, payments, ledger entries, and bank activity, linked by transactional relationships rather than textual similarity. End-to-end accuracy...

    arxiv.org/abs/2608.18534 · PDF

  31. 31

    Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson

    Tanay Chowdhury, Saeideh Shahrokh Esfahani

    cs.AI · cs.LG

    Industrial explainable-recommendation systems built on LLMs incur a substantial serving cost: each request triggers an LLM generation, with latency in the hundreds of milliseconds and cost that scales linearly with traffic. We separate generation from selection: explanations are produced ahead of time as a frozen candidate pool (six prompt styles, two commodity LLMs), and a small CPU-resident selector picks one at request time. The stack...

    arxiv.org/abs/2608.18531 · PDF

  32. 32

    Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval

    Haoyue Liu, Ye Chen, Zhichao Wang, Xiaoying Tang

    cs.AI · cs.LG

    Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules into contrastive fine-tuning. However, these methods largely inherit the same InfoNCE objective, whose optimization can prematurely saturate under a strong pre-trained initialization: on dense captions, the loss falls below 10^{-3} on 80% of batches within the first epoch, while its gradient reaches exact...

    arxiv.org/abs/2608.18521 · PDF

  33. 33

    UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

    Libiao Chen, Xiyang Liu, Yanheng Wei, Tao Wang, Zhenyu Tang

    cs.AI

    Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both efficient corpus-scale matching and fine-grained semantic reasoning. Recent MLLM-based embedding methods typically derive representations from hidden states, while Chain-of-Thought (CoT) reasoning is emerging as a promising strategy for embedding enhancement by encoding intermediate semantic evidence into the representation space. However,...

    arxiv.org/abs/2608.18504 · PDF

  34. 34

    FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

    Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li

    cs.AI

    Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same...

    arxiv.org/abs/2608.18423 · PDF

  35. 35

    Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions

    Shrenil Shaun Sharma, Avi Sharma

    cs.AI

    Combinatorial scheduling poses a significant challenge for language models, requiring them to identify feasible solutions within exponentially large search spaces while satisfying complex constraints. This challenge is especially pronounced in resource-constrained settings, where larger language models are impractical and selection is limited to smaller models which often fail to preserve feasibility when scheduling directly from natural...

    arxiv.org/abs/2608.18409 · PDF

  36. 36

    When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress Classification

    Saba A. Farahani, Hung Cao, Amir M. Rahmani

    cs.AI

    Wearable stress classifiers can achieve strong average performance while failing completely for a particular individual. On WESAD, a Random Forest reaches 93.0% mean accuracy yet yields F1 = 0 for Subject 14, whose cross-signal coupling weakens near stress onset. We call this structural ambiguity: individually plausible physiological channels form an inter-signal pattern that is poorly supported by the person's non-stress reference. We...

    arxiv.org/abs/2608.18397 · PDF

  37. 37

    A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations

    Hasan Najib Mahmud, Shreya Gupta, Isha Chaudhary, Nathaniel Enis, Ravi Mangal, Gagandeep Singh, Corina Pasareanu

    cs.AI

    AI code agents are increasingly deployed to resolve real software issues, yet their reliability under superficial code variations remains poorly understood. We evaluate whether coding agents that repair repository-level issues remain reliable when the surrounding codebase is rewritten into a semantically equivalent form. We introduce a random variant sampler that applies common semantics-preserving transformations (SPTs) - spanning...

    arxiv.org/abs/2608.18389 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.