cs.AI · 2026-08-23 · No. 93

Artificial Intelligence, 2026-08-23.

55 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

55 entries
  1. 01

    An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction

    Narges Ahmadi, Yubo Jiao, Jônatas Augusto Manzolli, Jiangbo Yu, Luis Miranda-Moreno

    cs.AI · cs.CL

    Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction. A chatbot-administered, image-augmented stated-preference survey collected mode choices from student commuters across five predefined weather...

    arxiv.org/abs/2608.20320 · PDF

  2. 02

    AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

    Yizhe Chi, Wenyi Li, Deyao Hong, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na

    cs.AI · cs.CL · cs.LG

    Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the training algorithm: a better objective or update rule improves the compute\mbox{-}capability exchange rate for every subsequent run, including the one that produces the next agent. Whether RSI is feasible therefore turns on whether an agent can design training...

    arxiv.org/abs/2608.20318 · PDF

  3. 03

    Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation

    Adam Fisch, Shubhendu Trivedi, Fantine Huot, William W. Cohen, Michael Kaisers, Mirella Lapata, Kate Larson, Jacob Eisenstein

    cs.AI

    Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost. Routing requires estimating each specialist's expected return, but this value estimation has a cost. Cheap estimators (e.g., embedding-based predictors) are fast but noisy, while accurate estimators (e.g.,...

    arxiv.org/abs/2608.20316 · PDF

  4. 04

    MidTool: Mid-training Data Synthesis for Agentic Tool Use

    Fengqing Jiang, Yite Wang, Boyi Liu, Zhaoyang Wang, Canwen Xu, Zhewei Yao, Radha Poovendran, Yuxiong He

    cs.AI

    Mid-training is increasingly recognized as a critical stage for shaping the capabilities of large language models. Recent work has shown that targeted mid-training can strengthen reasoning-intensive abilities such as math and science, and can also improve agentic capabilities in software-engineering settings. In this work, we study the parallel but less explored agentic capability: general tool use. We present MidTool, an open corpus...

    arxiv.org/abs/2608.20314 · PDF

  5. 05

    Phantom Gains: Auditing Self-Improvement Against a Measured Null

    Cheng Xu, Nan Yan, Liming Chen, M-Tahar Kechadi

    cs.AI · cs.CL

    Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and loses. Tracking these transitions means differencing two noisy estimates, leaving them vulnerable to measurement artifacts. Auditing three rounds of rank-$32$ LoRA self-training on Qwen3-8B against a frozen control pushed through the identical pipeline, we identify seven measurement failures, each of which...

    arxiv.org/abs/2608.20290 · PDF

  6. 06

    Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents

    Yiyang Feng, Biddut Sarker Bijoy, Niranjan Balasubramanian, Jiawei Zhou

    cs.AI · cs.CL

    Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In practice, induced skills may transfer unreliably and can even harm the agent that retrieves them. When agent-induced skills transfer reliably across tasks remains an open question. We conduct a comprehensive and controlled study of how the way skills are induced shapes their transfer across tasks....

    arxiv.org/abs/2608.20274 · PDF

  7. 07

    Catching the Rug: Early Prediction of Fraudulent Memecoins on Solana via Machine Learning

    Jianghai Li, Pavel Kuznetsov, Yury Yanovich, Konstantin Nott-Whaley, Igor Vodolazov

    cs.AI · cs.DC

    The rapid proliferation of memecoins on blockchain platforms has increased the risk of fraudulent activities, particularly rug pulls. While previous studies have focused on Ethereum-based tokens, this paper shifts the spotlight to Solana, the leading blockchain for memecoins by trading volume and token count. Unlike Ethereum, where rug pulls often exploit smart contract backdoors, Solana memecoin rug pulls are predominantly driven by...

    arxiv.org/abs/2608.20271 · PDF

  8. 08

    Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

    Gijs Kassenaar, Zhao Yang, Vincent François-Lavet

    cs.AI

    Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adaptive one, which can lead to over-computation on easy problems and insufficient computation on difficult ones. We study whether a model can learn to allocate its own reasoning effort by choosing, as the first token of its response, one of three modes: \textsc{NoThink} (answer as quickly as possible),...

    arxiv.org/abs/2608.20256 · PDF

  9. 09

    QUASAR: A Quantum-Classical Neural Network for SAR Satellite Physical-Layer Authentication

    Vincenzo Sammartino, Nathanael Denis, Roberto Di Pietro

    cs.AI · cs.CR

    X-band SAR satellites (8-12 GHz) play a critical role in disaster response, environmental monitoring, and military intelligence. Yet, they lack robust physical-layer authentication (PLA), a security layer orthogonal to cryptographic solutions. Existing PLA systems, typically based on radio-frequency fingerprinting, are often limited to sub-6 GHz frequencies and rely on classical deep learning. However, this approach underfits the IQ phase...

    arxiv.org/abs/2608.20240 · PDF

  10. 10

    Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

    Yu Chen, Ting Lei, Yaoyi Li, Jia Cai, Zhecen Wu, Yang Liu

    cs.AI

    Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unseen rule constraints remains underexplored. This setting requires models to jointly understand spatial layouts, interpret natural-language rules, and plan valid actions accordingly. To address this gap, we introduce RuleMaze, a controllable benchmark in which MLLMs...

    arxiv.org/abs/2608.20237 · PDF

  11. 11

    InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries

    Samuel J. Vincent, Daniel Calloway, Fangyi Yu, Andrew M. Bean, Nabeel Seedat

    cs.AI

    Legal AI systems are increasingly used to answer legal questions, yet existing benchmarks assume queries arrive fully specified. In practice, users omit facts that materially determine the legal outcome. We introduce InsufficiencyBench, the first legal benchmark targeting query-side insufficiency: whether a model recognizes when a query lacks legally material information, identifies what is missing, and refrains from premature conclusions. We...

    arxiv.org/abs/2608.20220 · PDF

  12. 12

    Electronic Navigational Chart Change Classification

    Jacob Arndt, Abhishek Potnis, Alexandre Sorokine

    cs.AI

    Electronic Navigational Charts (ENCs) are geospatial vector datasets used in maritime navigation systems that represent hydrographic and navigational information such as depths, navigational aids, traffic schemes, and hazards. A major challenge for hydrographic offices is determining whether a given chart change poses a critical or non-critical risk to maritime safety. Existing workflows rely heavily on manual review and verification, which...

    arxiv.org/abs/2608.20218 · PDF

  13. 13

    ContractScrub: A benchmark for final review of legal contracts

    Yejin Bang, Kirsty Fielding, Brandan Oliver, Brian Birke, Nabeel Seedat, Andrew M. Bean

    cs.AI · cs.CL

    Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final review of transactional agreements for errors and inconsistencies, is a particularly suitable task for automation, because it is routine, painstaking work requiring detailed attention to long documents. Scrubbing also seems to align naturally with the general...

    arxiv.org/abs/2608.20204 · PDF

  14. 14

    MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

    Mengru Wang, Haozhe Luo, Zhenqian Xu, Zhixiang Cui, Haoming Xu, Qu Yang, Jizhan Fang, Junfeng Fang, Ningyu Zhang

    cs.AI · cs.CL · cs.CY · cs.DB · cs.LG

    Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and...

    arxiv.org/abs/2608.20202 · PDF

  15. 15

    The Third Restructuring of Software Form: From the Three-Tier Architecture to Storage, Models, and Agents

    Wei Lin, Tao Zhou, Zhaofei Xie, Changgui Hong

    cs.AI · cs.SE

    Software form has undergone two paradigm shifts since its inception: Software 1.0, in which instructions determine behavior, and Software 2.0, in which data determines behavior (machine learning). This paper argues that a third shift - Software 3.0, in which context and reasoning determine behavior - is now underway, and contends that its terminal form converges to three elements: a generalized database (the unified abstraction of all...

    arxiv.org/abs/2608.20201 · PDF

  16. 16

    DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing

    Haoxiang Cao, Jiajiong Cao, Xuanpu Zhang, Changqian Yu, Chaoqun Wang

    cs.AI

    Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffusion model then executes that plan. Training such systems with only final-image rewards is inefficient because a poor edit does not reveal whether additional optimization should place more emphasis on the planner or the renderer, and even planner-dominant cases remain difficult to...

    arxiv.org/abs/2608.20161 · PDF

  17. 17

    DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

    Siyuan Ma, Boshi Zhang, Yutian Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Qiaojun Yu

    cs.AI · cs.RO

    Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOWAM, a whole-body world-action model that separates these factors through dedicated conditional interfaces. DECOWAM freezes an adapted FastWAM...

    arxiv.org/abs/2608.20114 · PDF

  18. 18

    What You Can't See Is What You Learn: Restricted Evidence Visibility Favors Compositional Generalization in Shared-Genome Language-Model Societies

    Narcis Marincat

    cs.AI · cs.LG · cs.MA

    Multi-module systems often expose every module to the full input. We test whether restricting evidence visibility changes which solutions gradient-based training discovers. Four-cell societies share one frozen pretrained language model and one low-rank adapter, communicating only through two model-width continuous vectors in a fixed relay. On a prospectively sealed natural-language function-composition task, we train ten matched...

    arxiv.org/abs/2608.20054 · PDF

  19. 19

    On the Applicability of Safety Nets: A Safety-By-Design Solution for Certifying Neural Networks

    Johann Maximilian Christensen, Thomas Stefani, Elena Hoemann, Frank Köster, Sven Hallerbach

    cs.AI

    The integration of Artificial Intelligence (AI) in safety-critical aviation systems presents significant challenges for certification and deployment. Aviation, often regarded as the safest form of transportation, relies on numerous safety-critical systems. For future safety-critical AI-based systems, EASA requires a Safety-by-Design approach, which can be achieved by using Safety Nets that combine neural network compression with lookup tables...

    arxiv.org/abs/2608.20053 · PDF

  20. 20

    A three-dimensional typology of agency for advanced AI systems

    Willem Fourie

    cs.AI · cs.CY

    Research on the agency of advanced artificial intelligence (AI) systems focuses on agency as a normative concept and on the agency of particularly agentic AI systems. While recent work also focuses on the different profiles of agentic systems, no framework exists to address the question of the type of agency instantiated by advanced AI systems, particularly when considering non-moral forms of agency. Based on established theoretical positions...

    arxiv.org/abs/2608.20041 · PDF

  21. 21

    Contrastive Mixed Prompt Learning for Incomplete Multimodal Sentiment Analysis with Unseen Modality Combination

    Kaixin Xu, NaiJin Liu, Yulin Kang, Tangyue Jin, Zixuan Yu, Wenxi Zhao, Yibei Liu, Qianle Zhang, Yangyang Wu,...

    cs.AI

    Incomplete multimodal sentiment analysis has garnered significant attention in recent years. Existing approaches typically assume that data is missing at random or are designed specifically for certain missing patterns, ignoring the modality combination inconsistency between training and testing phases. However, in real-world scenarios, the testing phase often encounters modal combinations that were not present during the training phase,...

    arxiv.org/abs/2608.20019 · PDF

  22. 22

    Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

    Yansen Han, Shengyi Liao, Yuanxing Zhang, Pengfei Wan, Tao Lin

    cs.AI · cs.CV

    Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify transport trajectories without an inherent constraint to the pretrained data manifold and can move terminal samples off the pretrained support. We formalize this failure mode as manifold drift. Theoretically, we show that optimal flow matching recovers...

    arxiv.org/abs/2608.20011 · PDF

  23. 23

    ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory Forecasting

    Rui Wang, Yeteng Wu, Xianlin Zhang, Mengshi Qi

    cs.AI

    Understanding object dynamics requires not only predicting future trajectories but also examining whether a model captures the physical properties that govern motion. However, existing benchmarks rarely expose object-level physical properties as explicit evaluation targets alongside trajectory forecasting. To address this gap, we introduce \emph{ExPhy}, a multi-object trajectory forecasting benchmark containing 24,000 simulated physical...

    arxiv.org/abs/2608.20009 · PDF

  24. 24

    Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees

    Yu Chen, Ruishuo Chen, Xun Wang, Zhuoran Li, Longbo Huang

    cs.AI

    Loading reusable skill documents into a bounded context window is now the primary way large language model (LLM) agents acquire task-specific capabilities, which makes skill selection a first-order determinant of task performance and token cost. Yet current agents score skills independently by semantic relevance and assemble the set by top-$k$ or greedy packing, with no quality guarantee or cost awareness on the selected set. As a result,...

    arxiv.org/abs/2608.19993 · PDF

  25. 25

    ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance

    Yiyang Luo, Yihang Jiang, Qijun Xie, Liang Lan, Lin Willian Cong, Anyi Rao, Yunya Song

    cs.AI

    LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence. In trader runs with DeepSeek V4 Pro and Gemini 3.5 Flash, visible rules...

    arxiv.org/abs/2608.19974 · PDF

  26. 26

    Rethinking Patch Based Multivariate Time Series Forecasting with Semantic Structured Partitioning

    Jiazhe Wang, Zhiquan Huang, Linjing Xue, Ming Liu, Meiwen Li, Ruijuan Zheng

    cs.AI

    Multivariate time series forecasting (MTSF) is a fundamental task in many real world applications. Existing patch based forecasting methods generally fall into three categories: fixed partitioning, multi-scale partitioning, and extendable partitioning. Fixed partitioning often breaks meaningful temporal boundaries, multi-scale partitioning may introduce redundant representations across scales, and extendable partitioning improves flexibility...

    arxiv.org/abs/2608.19966 · PDF

  27. 27

    Learning Early-to-Final Solution Consistency for MILP Acceleration

    Guanlin Li, Chengrui Gao, Chenguang Wang, Haopu Shang, Zherong Zhang, Ke Xue, Jixiang Lu, Weiyong Yang, Chao Qian

    cs.AI

    Mixed-Integer Linear Programming (MILP) is a fundamental problem class in operations research and combinatorial optimization, with broad applications to industrial decision-making. Owing to their NP-hardness, however, modern solvers may struggle to find high-quality solutions for challenging MILP instances within practical time limits. Recent learning-based approaches seek to accelerate MILP solving by directly predicting high-quality...

    arxiv.org/abs/2608.19953 · PDF

  28. 28

    A Strong Linear Baseline for Whole-Heart Cardiac Shape Completion on CT, with an Open Eleven-Structure Statistical Shape Model

    Matej Gazda, Jakub Gazda, Juraj Gazda, Peter Drotar

    cs.AI

    Public cardiac cohorts annotate different subsets of the heart, so shapes from separate sources cannot be pooled without shared correspondence. Among released cardiac shape resources, none we identified carries the atrial appendage, pulmonary veins, and caval stumps as separate blocks in one mesh. Completion benchmarks also compare deep models against a least-squares projection onto shape modes, not the conditional estimator the same fitted...

    arxiv.org/abs/2608.19932 · PDF

  29. 29

    Spike-based Belief Propagation in Nonlinear Dynamical Systems

    Sepideh Adamiat, Hongye Wang, Wouter M. Kouw, Bert de Vries

    cs.AI · cs.LG · cs.NE · eess.SY

    This paper presents a Bayesian control framework that integrates spike-based dynamics with probabilistic inference for adaptive control. Bayesian inference is widely regarded as a core computational principle of brain function, providing a normative framework for perception, decision-making, and learning under uncertainty. By combining a biologically inspired spiking neural model with Bayesian inference principles, we propose a brain-like...

    arxiv.org/abs/2608.19907 · PDF

  30. 30

    Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis

    Zijiao Chen, Nicholas Lu, Xinhui Li, Jocelyn A. Ricard, Ce Ju, Huan H. Wang, Christian Kindermann, Jeanette A....

    cs.AI · q-bio.NC

    AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the claim is limited to what the evidence supports. Agents may reproduce failures including selective analysis, premature declarations of success and optimization of imperfect criteria. We present Brain Researcher, an agentic research harness operating in a neuroimaging researcher's computational environment...

    arxiv.org/abs/2608.19902 · PDF

  31. 31

    EXIMO: VLM Guided Exploration of VLA Policies

    Bhavya Sukhija, Oliver Groth, Mohit Shridhar, Tim Hertweck, Michael Bloesch, Markus Wulfmeier, Abbas Abdolmaleki,...

    cs.AI

    How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation policies are based on behaviour cloning of large vision-language-action (VLA) models with billions of parameters on huge teleoperation datasets. While this simple approach has enabled significant advances for robotic manipulation, finetuning of VLA policies for learning new tasks still remains an open problem. In particular,...

    arxiv.org/abs/2608.19891 · PDF

  32. 32

    Write Once, Run Everywhere: The Axon DSL for Shape-Safe and Framework-Agnostic LLM Architectures

    Jacob Nielsen, Danial Namazifard, Lukas Galke Poech, Peter Schneider-Kamp

    cs.AI · cs.PL

    The entire ecosystem of open-source language models effectively relies on a single platform. What if this platform was forced to shut down tomorrow? Implementing and maintaining efficient model definitions and translating them between different training and inference regimes is a resource-heavy task that severely limits model efficiency and portability, hindering both scaling and deployment. Here, we present Axon, a strongly typed...

    arxiv.org/abs/2608.19889 · PDF

  33. 33

    TESTNAV: Pareto-Guided Search for Compositional Robustness Testing

    Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis

    cs.AI

    Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (e.g., brightness shifts and motion blur). Compositional testing reveals these interaction effects but introduces two challenges: combinatorial growth of the perturbation space as dimensions and severity levels increase, and uneven diagnostic value-many combinations yield unrealistically degraded inputs...

    arxiv.org/abs/2608.19882 · PDF

  34. 34

    EnvHarness: Awakening Static Worlds for Agent Learning

    Chengsong Huang, Zifeng Wang, Rujun Han, Jun Yan, Yanfei Chen, Zoey CuiZhu, Ke Jiang, Peng Xia, Han Yu, Yufan...

    cs.AI · cs.CL · cs.LG

    LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive or unreliable verifiers, and still produce static environments. To alleviate the engineering burden of rebuilding environments from scratch, we...

    arxiv.org/abs/2608.19880 · PDF

  35. 35

    PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents

    Seongjae Kang, Taehyung Yu, Sung Ju Hwang

    cs.AI · cs.CL · cs.LG

    Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation. Runtime safeguards can intervene on risky actions, but action-local checks do not guide an agent through a multi-step procedure. Workflow-following systems support prescribed...

    arxiv.org/abs/2608.19861 · PDF

  36. 36

    SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

    Dayang Liang, Lang Feng, Bo An, Yunlong Liu

    cs.AI

    Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods estimate policy advantages from multiple rollouts, avoiding the substantial memory overhead of conventional proximal policy optimization (PPO) and achieving strong performance on long-horizon interactive tasks. Despite their success, recent studies revealed three limitations: (1) Lack...

    arxiv.org/abs/2608.19842 · PDF

  37. 37

    Specification-delta-driven data governance: an empirical study of the «spec-delta» as the unit of change in lakehouse data platforms

    Pablo Ramirez Amador

    cs.AI · cs.PL

    Spec Driven Development SDD has consolidated the idea that the specification rather than the code should be the primary artefact governing AI assisted work. Tools such as GitHub Spec Kit, and proposals such as Constitutional SDD, have formalised this principle in the software domain, while the executable data-contracts literature has extended it to schema and quality enforcement at run time. Nevertheless, the treatment of the specification...

    arxiv.org/abs/2608.19838 · PDF

  38. 38

    Causal Reasoning with Bipartite Graphical Causal Models

    Joris M. Mooij

    cs.AI · math.PR

    Causal Bayesian networks (CBNs) and structural causal models (SCMs) are the dominant frameworks for graphical causal reasoning, but they cannot adequately represent all real-world causal systems. In particular, systems at equilibrium---where feedback mechanisms create cyclic causal dependencies---can exhibit causal semantics that are fundamentally incompatible with these frameworks: different interventions that enforce the same variable value...

    arxiv.org/abs/2608.19831 · PDF

  39. 39

    When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation

    Yearim Kim, Njun Baek, Nojun Kwak

    cs.AI · cs.CV

    To prevent the adoption of aesthetically polished but pedagogically flawed AI content, we study a video authoring pipeline featuring two layers of structured refusal. The first layer empowers educators to iteratively reshape AI scripts based on multimedia learning theory, while the second employs automated metrics to flag violations in instructional coherence and narrative-visual synchronization. While neither layer is exhaustive, their...

    arxiv.org/abs/2608.19812 · PDF

  40. 40

    ADAPT: Physics-Aware Diffusion-based World Models for Adaptive Predictive Transferable HVAC Control

    Xu Yang, Kailai Sun, Dianyu Zhong, Qianchuan Zhao

    cs.AI

    Buildings account for roughly one-third of global energy consumption and CO$_2$ emissions. Optimizing indoor climate systems plays a critical role for urban climate mitigation aligned with UN Sustainable Development Goals 11 and 13. However, indoor delayed thermodynamic responses and partial observability severely hinder existing methods, which are primarily limited by implicit thermal inertia, occupancy dynamic prediction, and cumulative...

    arxiv.org/abs/2608.19804 · PDF

  41. 41

    Towards general embodied intelligence: integrating large language models, knowledge bases, and reasoning capabilities to build the next generation of AI agents

    Fujiang Yuan, Xia Huang, Lusheng Wang, Jun Ding, Zhen Tian, Yuxin Wang, Shaojie Gu, Yuki Funabora, Yanhong Peng, Zebing Mao

    cs.AI · cs.RO

    The convergence of large language models (LLMs), structured knowledge bases (KBs), and reasoning ability (RA) presents a promising trajectory toward general embodied intelligence (GEI). This paper reviews the evolution of LLM-centered intelligent systems, emphasising their integration with knowledge representation, logical reasoning, and physical embodiment. We analyse LLM architectures, pre-training methods, and inference mechanisms, along...

    arxiv.org/abs/2608.19794 · PDF

  42. 42

    LLMs as Acquisition Policies for Finite-Pool Materials Optimization: A Controlled Study

    Dino-Rober Demir, Florian Le Bronnec, Rio Yokota

    cs.AI

    Discovering materials with desirable properties often requires searching large candidate spaces while experimental or computational evaluations remain costly. Active learning addresses this challenge by using previous observations to select which candidate to evaluate next, typically through probabilistic surrogate models. We investigate whether open-weight large language models (LLMs) can serve as standalone acquisition policies in this...

    arxiv.org/abs/2608.19790 · PDF

  43. 43

    TT-net: Quantum Inspired Tensor Network Denoising in Conditional GANs

    Michal A. Sterzel, Marko J. Rančić

    cs.AI · quant-ph

    Developed as a workhorse for classical simulations of quantum algorithms and quantum many-body systems, Tensor Network methods have entered the scientific mainstream in quantum physics. Among various types of tensor networks, Tensor Trains (commonly know as Matrix Product States in the quantum computing community) have already found applications in machine learning. These methods often rely on a powerful linear algebra tool called the...

    arxiv.org/abs/2608.19789 · PDF

  44. 44

    GenMatch: An End-to-End Generative Matching Framework for Micro-View Order-Dispatching in Ride-Hailing

    Chuang Liu, Yuxueqing Zhang, Tengfei Lyu, Zirui Yuan, Weiqi Hu, Yanghan Cheng, Ming Wang, Li Ma, Zihao Lu

    cs.AI

    Micro-View Order-Dispatching assigns available drivers to passenger orders within each dispatch batch and is critical to the service quality and operational efficiency of ride-hailing platforms. Mainstream industrial solutions follow a multi-stage paradigm of model prediction, value calculation, and dispatch matching. Although dispatch quality is determined by the final batch-level assignment, these stages optimize different intermediate...

    arxiv.org/abs/2608.19751 · PDF

  45. 45

    SafeBranch: Branch-Pair Safety Alignment for Embodied Agents

    Hyunse Lee, Jiwoo Jeong, Haneul Lee, Kyochul Jang, Youngjae Yu, Woojin Lee

    cs.AI · cs.CV · cs.RO

    Vision-language-model-based embodied agents can complete instructed tasks but often violate safety constraints in the process, a problem recently framed as interactive safety. Training such agents to act safely is difficult, since safety and task success are distinct objectives, and safety arises only at a small number of safety-critical steps within a trajectory. Standard supervision is insufficient: imitating safe trajectories teaches...

    arxiv.org/abs/2608.19729 · PDF

  46. 46

    Beyond Memory Majority: Latent-Source Reasoning for Multi-Agent Memory Arbitration

    Chenchen Lin, Wenhao Yuan, Xuehe Wang, Edith Cheuk Han Ngai

    cs.AI

    Long-term multi-agent systems continuously accumulate the memories produced by different agents. Existing memory methods typically treat retrieved memories as independent evidence and combine them through voting or weighting. However, this independence assumption often fails in multi-agent settings: memories written by different agents may inherit the same upstream source or shared bias, causing correlated evidence to be repeatedly counted...

    arxiv.org/abs/2608.19701 · PDF

  47. 47

    Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

    Pei Wang, Xu Chen, Ji-Rong Wen

    cs.AI

    LLM-based social simulation is a promising complement to traditional methods such as surveys and behavioral experiments. A core question is how to evaluate the fidelity of LLM-simulated human behavior and optimize LLMs toward it. Prevailing practice evaluates by accuracy, checking whether the model selects the single response observed from a human, and trains the LLM to reproduce this hard label. However, human behavior is inherently...

    arxiv.org/abs/2608.19689 · PDF

  48. 48

    Learning Hierarchical Skill Policies with Offline Quality-Diversity Reinforcement Learning

    Tanachai Anakewat, Takayuki Osa, Tatsuya Harada

    cs.AI · cs.LG · cs.RO

    Recent studies investigate how to leverage pre-collected datasets to improve the policy performance and sample efficiency of RL. One promising approach to achieve this goal is to employ a two-stage strategy: In the first stage, diverse skills are extracted as a low-level policy from a given dataset, and a high-level policy is trained to solve a specific task in the second stage. Typically, extraction of the low-level policy is performed based...

    arxiv.org/abs/2608.19684 · PDF

  49. 49

    Frequency-Aware Continual Learning for Smart Contract Vulnerability Detection with Large Language Models

    Tenghui Huang, Jiawen Kang, Dongning Liu, Changyan Yi, Chengjun Cai, Anjia Yang, Li Li, Dong In Kim

    cs.AI

    Smart contract vulnerability detection with Large Language Models (LLMs) faces three causally linked challenges. First, new vulnerability categories demand parameter-efficient adaptation, since full retraining is prohibitive for sequentially arriving tasks. Second, training per-task adapters on a shared backbone causes catastrophic forgetting of previously learned vulnerabilities. Third, the resulting multiplicity of adapters must be...

    arxiv.org/abs/2608.19680 · PDF

  50. 50

    Can Agent Memory Systems Track Evolving State?

    Xinyi Fan, Miri Liu, Ruozhen Yang, Siru Ouyang, Jiawei Han

    cs.AI · cs.CL

    As LLM-based agents are deployed for longer and higher-stakes tasks, their memory systems continue to have crucial gaps. While existing memory benchmarks focus largely on recall-shaped tasks, we argue an effective memory system must track the evolving state of the world; as facts, constraints, and decisions are revised over a long interaction, answers must reflect the current state and not a superseded one. We define this capability as state...

    arxiv.org/abs/2608.19652 · PDF

  51. 51

    Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale

    Xiaohan Huang, Qingqing Long, Xiaolei Du, Siyu Pu, Jiawen Xu, Haotian Chen, Chenyang Zhao, Jinbiao Liu, Xuezhi Wang,...

    cs.AI

    Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery, interpretation, and invocation. This limitation stems from the fragmentation of scientific data across heterogeneous repositories and from dataset representations designed primarily for human use. To address this limitation, we introduce the Scientific Data Skill (SciDSK), an agent-ready representation...

    arxiv.org/abs/2608.19625 · PDF

  52. 52

    Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics

    Mohamed Akrout, Olivera Kotevska, Dan Wilson

    cs.AI · math.DS

    Large Language Models (LLMs) are increasingly deployed in high-stakes applications, yet their tendency to generate toxic, harmful, or policy-violating content poses significant risks. Detecting these unsafe outputs efficiently in a black-box manner remains an open challenge. In this paper, we extend a recently proposed dynamical systems framework designed for hallucination detection to LLM safety classification. By projecting both prompts and...

    arxiv.org/abs/2608.19579 · PDF

  53. 53

    From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG

    Zlatan Feric, Amir Taherin, Yanzhi Wang, David Kaeli

    cs.AI · cs.CL · cs.DC · cs.IR · cs.PF

    Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes with overhead: retrieved context lengthens the prompt, increasing prefill work, KV-cache footprint, memory traffic, latency, and energy. Context compression offers a natural remedy by pruning retrieved text before generation. However, state-of-the-art context-compression methods are typically used with a fixed...

    arxiv.org/abs/2608.19535 · PDF

  54. 54

    Symposium: Trust via Auditable Records for Communities of AI Scientist Agents

    Dexter Pratt

    cs.AI

    Symposium is a formal framework and practical implementation to record the operation of AI agents deployed by small scientific research communities. Symposium provides long-term, immutable histories of agent-driven research activity, leaving auditable trails of analyses, hypotheses, data, and scientific discourse. This shared record of published artifacts enables agents to build on prior work and preserves the evidence researchers and agents...

    arxiv.org/abs/2608.19511 · PDF

  55. 55

    Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

    Chen Yang, Haiyuan Wan, Rengrong Xiong, Yize Chen, Danny H. K. Tsang

    cs.AI · cs.LG

    On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajectories with dense token-level supervision from a teacher. However, OPD implicitly assumes that teacher-derived rewards are an appropriate proxy for reasoning progress, and therefore treats all teacher feedback equally during policy optimization. While in practice, this assumption does not always hold. We...

    arxiv.org/abs/2608.19408 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.