cs.AI · 2026-08-11 · No. 81

Artificial Intelligence, 2026-08-11.

57 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

57 entries
  1. 01

    GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis

    Alban Puech, Matteo Mazzonelli, Tamara R. Govindasamy, Mangaliso Mngomezulu, Héctor Maeso-García, Thomas Tolhurst,...

    cs.AI

    Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from engineering domains such as power system analysis, where strict physical consistency must be enforced. We present GENCO (GEometric Neural Corrective Optimizer), a unified neural solver for steady-state transmission grid analysis that handles power flow (PF), optimal power flow (OPF), and state estimation (SE) within a single...

    arxiv.org/abs/2608.09921 · PDF

  2. 02

    DSLE: A Learning Environment for Dark Souls Boss Encounters

    Derin Gezgin, Jim O'Connor, Tanner Goodwin, Gary B. Parker

    cs.AI · cs.NE

    We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounters of Dark Souls: Remastered as game-playing agent benchmarks through a Gymnasium-style interface. DSLE combines real-time combat, high-dimensional visual input, and sparse terminal rewards, with each environment step being a real action executed against the running game. To support controlled comparison, we define DSLE-5, a...

    arxiv.org/abs/2608.09902 · PDF

  3. 03

    SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

    Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng...

    cs.AI · cs.CV

    The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harness as a fixed deployment artifact, limiting their ability to evolve with emerging risks. Moreover, coupled functions across harness components obscure safety responsibility attribution, making localized evolution...

    arxiv.org/abs/2608.09885 · PDF

  4. 04

    ArchAgent v2: A Case Study with the Data Prefetching Championship

    Abraham Gonzalez, Raghav Gupta, Akanksha Jain, Hanna Alam, Alexander Novikov, Po-Sen Huang, Matej Balog, Marvin...

    cs.AI · cs.AR

    Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarchitecture discovery remains challenging due to vast search spaces, strict hardware budgets, and long simulation times. In this work, we present ArchAgent v2, a framework which scales automated microarchitecture search to multi-level data prefetching. While the original ArchAgent successfully discovered...

    arxiv.org/abs/2608.09874 · PDF

  5. 05

    Towards Expert-level Medical AI for Real-time Video Consultations

    Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay...

    cs.AI · cs.CL · cs.CV

    Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level...

    arxiv.org/abs/2608.09861 · PDF

  6. 06

    Agentic Auto-Research is Fuzz Testing

    Yifeng He, Jicheng Wang, Yinzhe Zhao, Jiachen Liu, Hao Chen

    cs.AI · cs.CL

    Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the proposer and ranking more samples with a learned judge or human reviewers. We argue that this *generate-and-rank* paradigm misses the problem of sparse feedback. Within a declared research problem, an agent follows the control loop of a greybox fuzzer: it proposes a candidate, executes it, observes feedback,...

    arxiv.org/abs/2608.09855 · PDF

  7. 07

    CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems

    Aimilios Hadjiliasi, Louis Nisiotis

    cs.AI

    The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a challenge, even with today's advancements in technology. Existing architectures are often focused on either the implementation of low-level reactive control systems that are constrained by commercial game engines, or high-level representations of reasoning models that can be difficult to...

    arxiv.org/abs/2608.09848 · PDF

  8. 08

    Mismatch Matters: On-Policy Distillation Beyond Token Agreement

    Zichao Yu, Chengzhi Yu, Shengze Xu, Yujin Han, Bingqing Jiang, Xu Wang, Difan Zou

    cs.AI · cs.CL

    On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect token agreement with the teacher despite globally flawed responses. We therefore shift our focus from agreement to teacher-student mismatch, and find that mismatch tokens can be mainly categorized into two types: student-excess...

    arxiv.org/abs/2608.09836 · PDF

  9. 09

    CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation

    Yaoning Yu, Kai-Min Chang, Ye Yu, Yi-Chia Wang, Haojing Luo, Haohan Wang

    cs.AI · cs.MA · cs.SI

    Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these discussions requires more than just generating individual comments, the generated threads should also match how real users express themselves and interact with others. We introduce CARD, a framework for generating realistic credit card discussion threads. Given a credit card post and its matched real...

    arxiv.org/abs/2608.09790 · PDF

  10. 10

    AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting

    Fan Yang, Nan Chen, Yijie Dong, Yuchen Zhang, Wei Zhang

    cs.AI · cs.CE · cs.LG

    Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels differ in periodicity and distribution drift, while their concentration trajectories contain both multi-scale dependencies and rapid changes. Recent methods have improved spatial dependency learning and meteorological covariate modeling. However, pollutant channels are still passed through the...

    arxiv.org/abs/2608.09775 · PDF

  11. 11

    Second-Order Muon Done Right: A Principled Marriage of Spectral Geometry and Curvature

    Tong Che

    cs.AI

    Muon's polar update is exact for an unweighted spectral geometry. We introduce GO-MUON, which uses a matched data-dependent geometry and reuses it across several optimization steps. Conditioned on any positive-definite left and right maps, its raw update exactly solves the corresponding weighted spectral oracle; this statement is independent of how the maps are estimated or how recently they were refreshed. For softmax cross-entropy, we...

    arxiv.org/abs/2608.09763 · PDF

  12. 12

    Matryoshka Language Model Suites

    Nathan Godey, Yoav Artzi

    cs.AI · cs.CL

    Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained end-to-end. This Matryoshka training framework reduces the total parameter count of the suite, enables low-cost distillation from the largest to all smaller sub-models at every training step, and...

    arxiv.org/abs/2608.09703 · PDF

  13. 13

    Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

    Kevin Murphy

    cs.AI

    Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, causal model, not a curve fit; and learning such a model requires \emph{experiments}, because passive data leaves its mechanisms unidentified. Experiments are expensive, so the central problem is \emph{data efficiency}. We present the Model Discovery Agent (MDA), which couples a large language model (LLM),...

    arxiv.org/abs/2608.09696 · PDF

  14. 14

    Adaptive Semantic Capacity Allocation for Parallel Generative Recommendation

    Chenxi Li, Yuchen Lu, Xu Yang

    cs.AI

    Autoregressive semantic ID recommenders are constrained by expensive beam-search decoding, which limits the practical length of item identifiers. Parallel generation methods alleviate this bottleneck by predicting all semantic ID tokens simultaneously, enabling longer IDs. However, existing semantic ID methods still rely on manually predefined and homogeneous ID structures, where both the number of semantic slots and the codebook size of each...

    arxiv.org/abs/2608.09685 · PDF

  15. 15

    Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models

    Shulin Tian, Ziqi Huang, Fan Zhang, Hongyuan Zhu, Yu Qiao, Ziwei Liu

    cs.AI

    Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands sampling hundreds or thousands of images or videos, which is computationally expensive. Existing evaluation methods also rely on rigid pipelines that overlook specific user needs and provide numerical results without clear explanations. Mimicking how humans quickly form impressions of a model's...

    arxiv.org/abs/2608.09666 · PDF

  16. 16

    Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching

    Yuke Li, Xuehan Hou

    cs.AI

    GUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operate directly on screenshots. The core task, GUI grounding, requires translating abstract user instructions into precise element coordinates. This task faces a persistent dual obstacle: conventional grounding models lack the semantic richness to interpret abstract instructions, while end-to-end MLLMs suffer...

    arxiv.org/abs/2608.09654 · PDF

  17. 17

    Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics

    Yen-Shan Chen, Yu Chian Duan, Chih-En Kuo, Jian-Bin Wu, Yun-Nung Chen

    cs.AI · cs.CL · cs.CY · cs.GT

    Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight. We present Avalon-ToM-Bench, a fine-grained benchmark that operationalizes ToM through the asymmetric-information mechanics of The Resistance: Avalon. Rather than evaluating end-to-end gameplay, it decomposes ToM into a...

    arxiv.org/abs/2608.09638 · PDF

  18. 18

    Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?

    Hui Xue, Fan Yang

    cs.AI

    Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop. We ask whether this task-specific procedure remains necessary when a frontier model acts as the optimizer. We introduce Open-Ended Optimization (OEO), which keeps the objective, permitted interactions, resource budget, data boundary, and evaluation fixed while...

    arxiv.org/abs/2608.09629 · PDF

  19. 19

    Adaptive Sequential Test Planning for Multi-Mechanism Reliability Qualification via Bayesian Monte Carlo Tree Search

    Youssef A. Elhagrasy, Ian Hill, André Ivanov

    cs.AI · eess.SY

    Reliability qualification of advanced semiconductor devices requires sequential stress decisions that balance characterization objectives against multiple competing failure mechanisms. Current practice relies on static test plans derived from population-level acceleration models, which cannot adapt to per-unit variability or real-time degradation observations. This paper presents a closed-loop adaptive test planning framework that formulates...

    arxiv.org/abs/2608.09622 · PDF

  20. 20

    From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization

    Achille Jacquemond, Yuma Ichikawa, Akira Sakai

    cs.AI

    Compressing large language models to two bits or fewer is increasingly feasible through block-wise post-training quantization; cross-block variants reconstruct neighboring Transformer blocks within a moving window. In the fixed two-block setting studied here, the matched sequential baseline moves this window through the network once, so errors introduced early in the sweep are not revisited. We propose Interleaved Cross-Block Quantization...

    arxiv.org/abs/2608.09595 · PDF

  21. 21

    ICM Out! Better Tournament Strategy from Computed Continuations, vs. Solvers and LLMs

    Boning Li, Longbo Huang

    cs.AI · cs.GT · cs.MA

    The Independent Chip Model (ICM) converts tournament chips into reference prize equity, and policies are routinely constructed against those values. Because ICM reads only stack sizes, it omits action order, blind obligations, and seat rotation, and it does not price the elimination pressure a big stack puts on the short stacks it can bust. Those omissions can alter the successor-state contrasts that determine a move. We introduce...

    arxiv.org/abs/2608.09586 · PDF

  22. 22

    CoRCi: Cross-Reconstruction of Coherent Interests Modeling in Cross-Domain Sequential Recommendation

    Qingtian Bian, Tieying Li, Marcus de Carvalho, Jiaxing Xu, Hui Fang, Yiping Ke

    cs.AI

    Cross-Domain Sequential Recommendation (CDSR) aims to alleviate data sparsity by transferring dynamic user interests across related domains. A key challenge lies in effectively bridging these domains. In single-domain modeling, models cannot distinguish between domain-specific and domain-invariant interests. Recent methods merge domain-specific sequences chronologically into a mixed-domain sequence to capture domain-invariant knowledge....

    arxiv.org/abs/2608.09580 · PDF

  23. 23

    ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization

    Hao Sui, Simeng Qin, Jie Liao, Xiaojun Jia, Bing Chen, Yang Liu

    cs.AI

    Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisoned skill can persistently compromise every agent that installs it. However, existing skill attacks either fire on every request or rely on fine-tuned weights or multiple skills, leaving a conditional and low-cost backdoor unexplored. In this work, we present ElasticBack, an effective conditional...

    arxiv.org/abs/2608.09577 · PDF

  24. 24

    The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games

    Fatemeh Seyedin, Adrian Weller, Jinhyuk Yun, Mahmoudreza Babaei

    cs.AI

    LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move from individual tools to participants in multi-agent organizations, an important question arises: do they reproduce the governance failures like free-riding, corruption, and entrenched leadership that plague human institutions? We introduce the Hierarchical Game (HG), a public goods game extended...

    arxiv.org/abs/2608.09574 · PDF

  25. 25

    Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

    Tianjun Pan, Yuan Li, Hongda Wang, Linbo Jin, Mengfei Song, Lei Gao, Qiming Shi, Shaokang Fu, Jiarong Zhao, Chengyu...

    cs.AI

    External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Yet their effectiveness depends not only on skill quality, but also on whether the policy can translate the provided guidance into appropriate actions. However, methods specifically designed to improve this skill-utilization ability remain largely underexplored. In practice, skill-based agents are commonly...

    arxiv.org/abs/2608.09555 · PDF

  26. 26

    verdi: retrieval is not transfer for continual world model optimization

    Junyu Wu, Shiqin Nie, Youyi Kou, Baohua Yin, Guocai Yao, Qingyu Chen, Jingheng Ma, Shiji Zhou, Hongyong Song,...

    cs.AI

    Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence. However, optimizing a pretrained world model toward a user-specified objective remains difficult: each campaign typically rediscovers optimization strategies from scratch, and the resulting knowledge rarely transfers to the next model. Existing research agents automate the optimization loop but treat successful strategies as directly...

    arxiv.org/abs/2608.09537 · PDF

  27. 27

    One Adapter Pair per Model: A Universal Activation Interface for Language Models

    Su-Hyeon Kim, Jiwan Mun, Yo-Sub Han

    cs.AI

    Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to be rebuilt or rediscovered for each new language model. We present a Universal Activation Bus, a framework that provides a common activation interface across compatible language models. Using a small set of source models, we learn a shared dense space together with one lightweight linear...

    arxiv.org/abs/2608.09521 · PDF

  28. 28

    Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification

    Karim Zaghw, Andrew Pashea, Marc Pritsch, Wouter Nuijten, Karl Friston, Lancelot Da Costa

    cs.AI · cs.CV

    Active inference offers a unified framework for perception, learning, and action, but scaling discrete active-inference models to rich spatial and temporal domains remains difficult. Renormalising generative models (RGMs) address this challenge by composing discrete generative models across spatial and temporal scales, coarse-graining lower-level states and paths into higher-level causes for objects, events, and action. However, fully...

    arxiv.org/abs/2608.09512 · PDF

  29. 29

    Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents

    Neel Tushar Shah, Manglam Kartik, Akshat Karkar

    cs.AI

    Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission, false consensus, and manipulative framing. We argue that Cooperative AI evaluations should separate what models can do under benign instructions from what they tend to do under realistic civic pressure. We introduce DiffCoop-Civic, a 10-scenario pilot evaluation suite spanning preference...

    arxiv.org/abs/2608.09485 · PDF

  30. 30

    From Prompt to Harness: Coderlet from Scratch

    Mengfan Li

    cs.AI

    A model alone does not determine how a programming agent acts. What the model sees, how actions enter the environment, how feedback returns, and how one run affects the next all depend on how the harness is organized. Minimal examples usually show only the basic interaction between a model and tools, while production systems spread these relationships across complex components and dependencies. This paper studies a compact harness design by...

    arxiv.org/abs/2608.09480 · PDF

  31. 31

    Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity

    Zihan Wang, Anglin Liu, Rongyi Wang, Dantong Li, Yi Lu, Siqing Yuan, Hongxia Xu, Zhongtian Long, Jintai Chen

    cs.AI

    Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on conditions, medications, and geriatric risks that users may omit. We introduce ATLAS, a coupled graph--policy distillation framework for patient-adaptive medication safety. ATLAS structures guideline evidence as a medication-safety graph. Targeted questions update the patient state and...

    arxiv.org/abs/2608.09443 · PDF

  32. 32

    Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

    Zhi Zeng, Cheng Zhang, Zesheng Yang, Rendong Pi, Jiaying Wu, Di Zhang, Zihan Ma, Guodong Li, Zhou Yang, Yu Xiang,...

    cs.AI

    Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources. To evaluate this missing capability, we introduce ST-OmniQA, a spatio-temporal audio-visual question-answering...

    arxiv.org/abs/2608.09435 · PDF

  33. 33

    KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models

    Chen Qiu, Ziwu Liu, Chao Fei, Guozhong Li, Panos Kalnis

    cs.AI

    KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct executions fail nor why. We present KVDiagnosis, a diagnostic dataset and benchmark with three contributions. First, a 25-method taxonomy groups methods into five mechanism families and links them to eight verified implementations and their valid diagnostic measurements. Second, for every supported method setting, we evaluate all sources...

    arxiv.org/abs/2608.09412 · PDF

  34. 34

    OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks

    Siqi Wang, Xinlin Li, Zhenglin Li, Li Li

    cs.AI

    Long-horizon complex tasks require agents to repeatedly observe states, formulate plans, invoke tools, verify results, and recover from failures in continuously changing environments. However, such control experience often remains confined to a single context or a fixed prompt, and is difficult to accumulate and reuse across historical traces. This paper presents OpenLoopEvolve (OLE), a self-evolution framework centered on the Loop Policy....

    arxiv.org/abs/2608.09380 · PDF

  35. 35

    CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits

    Xinqi Yang, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian,...

    cs.AI · cs.CV

    Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, select a physical model, formulate coupled equations, propagate intermediate quantities, and preserve units, signs, directions, and phase conventions. We introduce \benchmark, a benchmark of 1,000 authentic textbook problems for evaluating this complete long-horizon visual-to-symbolic reasoning...

    arxiv.org/abs/2608.09374 · PDF

  36. 36

    LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling

    Jinbo Li, Chuanhao Li

    cs.AI

    Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most methods treat the simulator as a black box: aggregate scores rank candidates without revealing why they fail or which policy logic should change. We present an LLM-guided heuristic design framework that uses repeated simulation for selection and event-level traces for diagnosis. Each incumbent is assessed through multiple replications, while...

    arxiv.org/abs/2608.09343 · PDF

  37. 37

    Control-Oriented Scenario Tree Construction through Reinforcement Learning

    Fabio Pavirani, Bert Claessens, Pierre Pinson, Chris Develder

    cs.AI · cs.LG · eess.SY

    Multistage stochastic model predictive control (MPC) handles uncertainty by optimizing over a scenario tree, a finite branching approximation of future outcomes constructed from sampled forecasts. To build such a tree, conventional methods focus on matching the underlying probability distribution---e.g., via Wasserstein-based scenario reduction---but improved distributional accuracy does not necessarily yield better control performance. We...

    arxiv.org/abs/2608.09335 · PDF

  38. 38

    GeoPhysAdapter: Scale-Matched Geophysical Adaptation for Cross-Domain Landslide Mapping with Vision Foundation Models

    Zhihang Liu, Mei-Po Kwan, Jinlin Wu, Hao Li

    cs.AI · cs.CV

    Newly triggered landslides rarely carry immediate annotations, so cross-domain transferability determines the value of landslide mapping for emergency response and regional risk assessment. Vision foundation models have strengthened representational transfer, yet on unseen regions, events, and data sources they still generate high-confidence false alarms. Terrain, material, and rainfall triggering can constrain such errors, but their supports...

    arxiv.org/abs/2608.09325 · PDF

  39. 39

    CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning

    Ambuj Mehrish, Sebastiano Vascon

    cs.AI

    On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding those that match the majority vote over $N$ sampled answers. That vote discards a correct answer whenever it is a minority and scores every majority-matching roll-out identically. We replace it with \emph{CoRE} (Consensus Rewards via Equilibrium): the $N$ roll-outs form a graph whose edges...

    arxiv.org/abs/2608.09324 · PDF

  40. 40

    ASPaeroFlow: Decomposition Heuristics for Joint Air Traffic Flow & Capacity Management

    Alexander Beiser, Markus Hecher, Nysret Musliu, Georg Trausmuth, Stefan Woltran

    cs.AI

    While mathematical models act as vital decision support systems for operational Air Traffic Flow and Capacity Management (ATFCM), existing approaches isolate Air Traffic Flow Management (ATFM) from Dynamic Airspace Configuration (DAC). This separation introduces an unresolved circular dependency between fixed-demand and fixed-capacity assumptions. Although joint optimization resolves this gap, the enlarged search space renders exact models...

    arxiv.org/abs/2608.09315 · PDF

  41. 41

    Linearized 2-Simplicial Attention

    Aritra Das, Dhruman Gupta, Debayan Gupta

    cs.AI

    We present a linearized form of 2-simplicial attention by rewriting the trilinear score as an inner product between a composite query and a key, so that the sum over one token axis takes the same form as ordinary softmax attention. We then approximate this sum with positive random features and store the entire past in a fixed-size state, while the second axis stays explicit over a short window of recent tokens. This enables us to achieve...

    arxiv.org/abs/2608.09307 · PDF

  42. 42

    CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation

    Harmanjot Singh, Abhra Dubey, Jorge Alejandro Amador Herrera

    cs.AI · cs.CV · cs.LG · cs.RO

    A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, support controlled edits, match a reference structural response under a declared analysis, and connect to other parts through valid joints. We present CADEngBench, a two-track benchmark for these capabilities. CADEngBench-P evaluates 300 parametric parts, each used for one zero-to-CAD task and...

    arxiv.org/abs/2608.09296 · PDF

  43. 43

    ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

    Adrian Li, Kelong Mao, Yudong Guo, Heming Xia, Xinwei Yang, Lirui Luo, Jace Wong, Pu Yao, Sulong Xu, Simiu Gu

    cs.AI · cs.CL

    Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise in device setup, meal preparation, event planning, and group takeout ordering, requiring joint reasoning about item compatibility, availability, store-level requirements, delivery fees, coupons, and budgets. Evaluation is challenging because multiple baskets may satisfy the same request,...

    arxiv.org/abs/2608.09282 · PDF

  44. 44

    MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

    Chenxu Du, Kang An, Tengyue Wang, Zhongyu Yang, Xinqi Yang, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian,...

    cs.AI

    Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion. We introduce MMArch, a benchmark for architecture and civil engineering spanning ten subdomains and built entirely from figures in...

    arxiv.org/abs/2608.09281 · PDF

  45. 45

    P$^{3}$: Joint Program-and-Proof Planning for Verified Code Generation

    Zenan Li, Ziran Yang, Peiyang Song, Zhaoyu Li, Kaiyu Yang

    cs.AI · cs.PL

    Verified code generation asks a large language model (LLM) to generate both an executable program and a machine-checkable proof that the program meets a formal specification, promising software that is correct by construction. The de facto workflow decouples the two halves of the problem: first synthesize a program, then attempt to prove it correct. We observe that this sequential pipeline can be both ineffective and inefficient in practice....

    arxiv.org/abs/2608.09277 · PDF

  46. 46

    Entropy-based Code Adversarial Translation for Real-world Repository Migration

    Yushun Tang, Yisen Cao, Zhicheng Chen, Lin Peng, Junkang Mao, Fengyi Song, Yantao Jia

    cs.AI · cs.SE

    LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produces a runnable application because long-horizon translation challenges LLM-based agents' ability to maintain repository-level migration objectives. In this work, we propose Entropy-based Code Adversarial Translation (ECAT), a multi-agent framework for automated Android-to-HarmonyOS repository migration....

    arxiv.org/abs/2608.09273 · PDF

  47. 47

    Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation

    Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Anurag Koul, Zeyu Liu, Shafiq Joty

    cs.AI · cs.LG

    Outcome verifiers score completed reasoning traces but do not assign credit to intermediate tokens. Privileged self-distillation attempts to fill this gap by rescoring a model's own rollout with training-only information. A token likelihood change, however, is not automatically outcome credit. We separate three questions: whether the score tracks better actions, whether feedback construction changes what is compared, and what behavior the...

    arxiv.org/abs/2608.09263 · PDF

  48. 48

    Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline

    Morris Lee

    cs.AI · cs.CL

    LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questions the warehouse cannot answer, deprecated columns after a schema change, and queries that execute successfully while returning the wrong business number. No execution-match metric can score them. This paper introduces WarehouseReliabilityBench, 400 frozen tasks over two synthetic warehouses...

    arxiv.org/abs/2608.09254 · PDF

  49. 49

    SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance

    You Lu, Xinyu Huang, Bihuan Chen, Xin Peng

    cs.AI

    LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable procedural knowledge, agents may still execute them unreliably. Even when an agent has demonstrated the capability to complete tasks under the guidance of a skill, it may fail to do so consistently across similar tasks or repeated runs due to deviations from the skill procedure or incorrect...

    arxiv.org/abs/2608.09253 · PDF

  50. 50

    Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution

    Bohan Lin, Hejia Geng, Xinyi Xie, Heng Zhou, Qinghua Xing, Bo Liu, Chen Zhang, Yudong Zhang

    cs.AI

    Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirely on text-level signals such as task descriptions, verbal reflections, and experience-derived rules, while the model's own internal representational state remains unobserved. Recent interpretability work has shown that LLMs maintain linear emotion representations that causally influence behavior; however,...

    arxiv.org/abs/2608.09248 · PDF

  51. 51

    An Explainable GNN Framework for Component-Level Anomaly Diagnosis

    Sena Ozgunay, Louise Travé-Massuyès, Jean-Michel Loubes, Raul Sena Ferreira

    cs.AI

    Industrial processes are complex systems composed of multiple interacting sensors that generate multivariate time series (MTS). Detecting anomalies in such systems is critical for reliability and safety, yet understanding their origin is equally important. Existing Graph Neural Network (GNN)based methods for anomaly detection primarily focus on sensor-level deviations and either attribute anomalies directly to the deviating sensors. When...

    arxiv.org/abs/2608.09246 · PDF

  52. 52

    SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

    Yuanchi Zhu, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Xinqi Yang, Hebao Zhu, Bokai Zhao, Tianyu Liang,...

    cs.AI

    Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance, identify hazardous interactions, explain potential accident mechanisms, and recommend preventive actions. Existing safety datasets primarily focus on visual perception or isolated violation recognition and provide limited supervision for evidence-grounded reasoning. We introduce...

    arxiv.org/abs/2608.09230 · PDF

  53. 53

    Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models

    Puneet Mathur, Manan Suri, Dinesh Manocha

    cs.AI

    Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive. While recent token compression methods attempt to alleviate this burden, compressing modalities in isolation often destroys the temporal cross-modal anchors necessary for coherent reasoning. We introduce Omni2LoRA, a two-stage framework for efficient parametric memory compression...

    arxiv.org/abs/2608.09227 · PDF

  54. 54

    CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving

    Junyao Wang, Yulin Xu, Yu Li, Pramod Khargonekar, Mohammad Abdullah Al Faruque

    cs.AI

    Modern autonomous vehicles are equipped with multiple sensors, such as cameras, LiDAR, and radar, for comprehensive environmental perception. However, robust cross-modal feature fusion remains a critical challenge, as the reliability of each sensor varies significantly across diverse real-world driving conditions, including poor visibility and adverse weather. While uncertainty quantification (UQ) mitigates this issue by allowing models to...

    arxiv.org/abs/2608.09202 · PDF

  55. 55

    Signature-Guided Capacity Occupancy for Dense Expert Merging

    Lingching Tung, Chi-Jui Kim, Beicheng Xu, Yuchen Wang, Bin Cui

    cs.AI

    Dense expert merging combines domain-specialized language models into one single checkpoint, typically by admitting task-vector support in weight space. However, this admission is governed by three decisions that existing methods answer only partially: where to open layer capacity from cross-expert conflict, who should occupy that capacity based on domain demand, and how to admit the resulting support without relying on costly recipe search....

    arxiv.org/abs/2608.09201 · PDF

  56. 56

    Structure-Preserving Uncertainty Propagation in First-Order Proof Search

    Tanel Tammet

    cs.AI · cs.LO

    GK is a query-directed first-order prover that extends ordinary resolution-based proof search with explicit positive and negative claims, numerical confidence values, and prioritized default rules with exceptions. It works directly with non-ground clauses, including equality and function terms. Candidate proofs are found by bounded first-order proof search; exception conditions of defaults are checked by further bounded searches, recursively...

    arxiv.org/abs/2608.09190 · PDF

  57. 57

    Agentic Router: An Execution-Grounded Continual Learning Approach With Memory

    Yuxuan Chen, Rongpeng Li, Zhifeng Zhao, Yuntao Liu, Xing Xu, Honggang Zhang

    cs.AI

    Large language model (LLM) agents provide a promising interface for command-line-based network operations, but a plausible command may still fail or introduce operational risk after execution. Existing approaches mainly focus on command generation or final configuration correctness, and do not use execution-grounded experience to jointly improve candidate coverage and action selection. We propose an execution-grounded dual-path...

    arxiv.org/abs/2608.09184 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.