cs.AI · 2026-07-01 · No. 40

Artificial Intelligence, 2026-07-01.

49 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

49 entries
  1. 01

    AxDafny: Agentic Verified Code Generation in Dafny

    Benjamin Breen, Austin Letson, Borja Requena Pozo, Leopoldo Sarra

    cs.AI

    We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny, a verifier-guided repair framework that iteratively generates implementations, invariants, assertions, and termination arguments. We also introduce LiveCodeBench-Pro-Dafny (LCB-Pro-Dafny), a benchmark of 250 competition-style programming problems translated into Dafny with formal...

    arxiv.org/abs/2606.32007 · PDF

  2. 02

    PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines

    Sameer Malik, Ayush Singh, Amar Prakash Azad

    cs.AI · cs.LG · cs.LO · cs.SC

    Policy-grounded document review requires determining whether a target document complies with organization-specific policies, guidelines, or playbooks. While large language models can assist with policy interpretation and document analysis, end-to-end prompting leaves the applied policy logic implicit, making compliance decisions difficult to inspect, update, and test. We present PolicyGuard, a neuro-symbolic framework for policy-grounded...

    arxiv.org/abs/2606.32004 · PDF

  3. 03

    Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA

    Ekaterina Alimaskina, Denis Shveykin, Gleb Molodtsov, Igor Shalygin, Alexey Kadeishvili, Aleksandr Beznosikov

    cs.AI · cs.LG

    Language models are increasingly taught from synthetic question--answer (QA) supervision: a model generates questions about a document, answers them from the same text, and the resulting pairs are used to fine-tune, distill, or compress knowledge into another model. We show that this generation step is not neutral preprocessing. It is an implicit policy that both selects which evidence becomes training signal and decides how that evidence is...

    arxiv.org/abs/2606.32002 · PDF

  4. 04

    TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models

    Shiyi Chen, Nicholas Saban, Collin Hargreaves, Huiqi Wang

    cs.AI · cs.MA

    Human-labeled data are widely used as reference annotations in ML, despite known variability across annotators in many expert-driven domains. In addition, expert annotation is slow, inconsistent, and remains a major bottleneck for scaling tasks like tree height bias classification in forestry remote sensing. We propose a multi-agent system (MAS) that orchestrates expert decision trees with Vision-Language Models (VLMs), treating the decision...

    arxiv.org/abs/2606.31976 · PDF

  5. 05

    Harnessing Textual Refusal Directions for Multimodal Safety

    Moreno D'Incà, Massimiliano Mancini, Nicu Sebe

    cs.AI · cs.CV · cs.LG

    To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. Both strategies are less feasible in Multimodal LLMs (MLLMs) as they require unsafe multimodal data, harder to collect than their unimodal counterpart. In this work, we relax this constraint and investigate whether textual refusal directions, extracted directly from the LLM backbone, generalize...

    arxiv.org/abs/2606.31876 · PDF

  6. 06

    An Agentic AI Framework to Accelerate Scientific Discovery in Plant Phenotyping

    Renan Souza, Daniel Rosendo, Kelsey Carter, John Lagergren, Frédéric Suter, Shelaine L. Curd, Gerald A. Tuskan,...

    cs.AI

    High-throughput plant phenotyping now generates image derived datasets far faster than scientists can analyze them. At Oak Ridge National Laboratory's Advanced Plant Phenotyping Laboratory (APPL), automated stations image hundreds of plants daily across multiple remote sensing modalities; yet, trait extraction and interpretation remain manual, expert-bound, and strictly post-hoc, making analysis, not acquisition, the binding constraint on...

    arxiv.org/abs/2606.31831 · PDF

  7. 07

    Adaptive Cluster-First Route-Second Decomposition for Industrial-Scale Vehicle Routing

    Oguzhan Karaahmetoglu, Hyong Kim

    cs.AI

    Large-scale capacitated vehicle routing problems (CVRPs) are commonly addressed using cluster-first route-second (CFRS) approaches that split a routing instance into smaller, computationally tractable subproblems. Existing splitting methods typically rely on fixed partitioning rules, predefined optimization objectives, or learned policies, which may perform inconsistently across instances exhibiting different spatial, demand, and operational...

    arxiv.org/abs/2606.31820 · PDF

  8. 08

    Creating Intelligence: A Computational Foundation for AGI

    Peter Overmann

    cs.AI

    This work introduces a new computational theory of mind grounded in set theory and hyperdimensional computing. Whereas traditional neural networks rely on continuous weights and matrix multiplication, this framework works with sparse binary data. It represents information as discrete sets, directly modeling biological neural population codes. I demonstrate that associative memory emerges naturally from network topologies featuring a...

    arxiv.org/abs/2606.31819 · PDF

  9. 09

    Large Databases Need Small, Open-Weight Language Models

    Parker Glenn, Alfy Samuel

    cs.AI · cs.DB

    Language model systems built around proprietary APIs often operate on a token-based cost model. This becomes prohibitively expensive in the context of large databases, where LM-enhanced relational operators can incur costs exceeding $10,000 for a single set of experiments, hindering thorough research and practical deployment. In this paper, we demonstrate that quantized, open-weight models running locally on just 16GB of VRAM can match or...

    arxiv.org/abs/2606.31808 · PDF

  10. 10

    RAISE: LLM-based Automated Heuristic Design with Robust Adversary Instance Search

    Fei Liu, Alessio Figalli, Patrick Owen, Nicola Serra

    cs.AI

    Automated Heuristic Design (AHD) with Large Language Models (LLMs) has shown remarkable progress in discovering high-quality heuristics. However, existing LLM-based AHD methods optimize heuristics for a fixed training instance set and may fail catastrophically when deployed under real-world distributional shifts. We propose Robust Adversary Instance Search (RAISE), a framework that integrates constrained worst-case instance search within a...

    arxiv.org/abs/2606.31801 · PDF

  11. 11

    Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision

    Xianda Zheng, Huan Gao, Meng-Fen Chiang, Michael Witbrock, Kaiqi Zhao, Shangyang Li

    cs.AI

    Despite recent progress, the reasoning capabilities of large multimodal language models (MLLMs) remain fundamentally constrained by static supervision, where fixed prompts, rules, or reward models provide non-adaptive guidance throughout training. Such static signals are often sufficient to enforce output formats, but fail to shape the underlying reasoning process, leading to brittle generalization and performance saturation in complex...

    arxiv.org/abs/2606.31800 · PDF

  12. 12

    A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols

    Yankai Jiang, Weiting Tang, Haoran Sun, Zhenyu Tang, Yuejie Hou, Yingnan Han, Rubo Wang, Yueyuxiao Yang, Cheng...

    cs.AI

    Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constraints and experimental feedback must remain aligned from protocol and SOP design to code and physical execution. We developed ProtoPilot, a self-evolving multi-agent system, together with an expert-grounded benchmark and evaluation framework for testing this conversion as an experimental automation problem....

    arxiv.org/abs/2606.31763 · PDF

  13. 13

    Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

    Yuanhao Ban, Tong Xie, Sohyun An, Yunqi Hong, Evan Frick, I-Hung Hsu, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh

    cs.AI

    Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness benchmarks, however, rely on simple atomic instructions, on which top-tier systems already achieve near-perfect scores. As T2I models enter creative workflows, users issue multi-faceted requests combining intricate spatial relationships, stylistic constraints, and...

    arxiv.org/abs/2606.31711 · PDF

  14. 14

    FARS: A Fully Automated Research System Deployed at Scale

    Qiong Tang, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao

    cs.AI

    Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks. We present FARS (Fully Automated Research System), a fully automated AI-for-AI research system designed to operate across research topics at scale. FARS autonomously generates and advances...

    arxiv.org/abs/2606.31651 · PDF

  15. 15

    Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents

    Utsav Garg, Sungjin Hong, Jason Jung, Justin Lee, Shaan Desai, Joon Hee Kim, Anirudh Shrinivason, Edmond Wen, Susie Park

    cs.AI · cs.LG

    We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-English enterprise agents under practical memory and serving constraints. The model trains from Cohere's fully post-trained Command A model rather than a new pretraining run, and uses preamble conditioning to switch between concise non-reasoning behavior and longer tool-oriented reasoning. We study four...

    arxiv.org/abs/2606.31648 · PDF

  16. 16

    Scientific Explanations in Health Sciences: Causality, Trust, and Epistemic Adequacy

    Martina Mattioli, Marcello Pelillo

    cs.AI

    Medical Artificial Intelligence (AI) is widely expected to transform clinical practice, yet the decision-making processes of many Machine Learning (ML) models remain opaque. Explainability has been advanced as a partial remedy to clarify why AI generates predictions, particularly in high-stakes contexts. Despite ongoing efforts, debates on what constitutes an adequate medical explanation remain unsettled. Yet, explanation has long been a...

    arxiv.org/abs/2606.31616 · PDF

  17. 17

    Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index

    Outongyi Lv, Yanzhao Zheng, Yuanwei Zhang, Zhenghao Huang, Xingjun Wang, Baohua Dong, Hangcheng Zhu, Yingda Chen

    cs.AI

    Reinforcement learning (RL) has become a powerful tool for propelling Large Language Models (LLMs) beyond imitation-based training towards more robust reasoning capabilities. Among existing approaches, RL with Verifiable Rewards (RLVR) has emerged as a pivotal paradigm for advancing LLM reasoning. Despite its empirical success, recent studies have offered different insights. One line of inquiry advocates prioritizing high-entropy token...

    arxiv.org/abs/2606.31575 · PDF

  18. 18

    ACE: Pluggable Adaptive Context Elasticizer across Agents

    Ning Liao, Zihao Long, Xiaoxing Wang, Xue Yang, Yaoming Wang, Ziyuan Zhuang, Xunliang Cai, Rongxiang Weng, Junchi Yan

    cs.AI

    The increasing complexity of agentic tasks has led to rapidly growing trajectory lengths, which poses significant challenges for large language model (LLM) based agents with fixed context windows. Existing context management techniques, such as truncation and summarization, suffer from inherent inflexibility and irreversibility: once information is discarded or compressed, it cannot be recovered even when it becomes critically relevant in...

    arxiv.org/abs/2606.31564 · PDF

  19. 19

    Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2

    Johan Land

    cs.AI · cs.CL · cs.LG

    Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making selection among candidates, not just generation, the central challenge. I present a solver for ARC-AGI-2, a few-shot visual reasoning benchmark, built around two principles: (i) treating reasoning modalities as search operators, generating diverse candidates independently across text, image,...

    arxiv.org/abs/2606.31543 · PDF

  20. 20

    A time-series classification framework for individual-level absenteeism prediction under severe class imbalance

    Kwong Ho Li, Matthew Roughan, Wathsala Karunarathne

    cs.AI

    Staff absenteeism imposes substantial operational costs in high-demand work environments such as healthcare, emergency services, meat processing, construction, and courier and delivery services, where proactive workforce planning depends on reliable individual-level absence prediction. Existing regression and classification approaches share a structural limitation; they map features observed at time t to labels at the same time t, reproducing...

    arxiv.org/abs/2606.31532 · PDF

  21. 21

    Design and Implementation of Agentic Orchestrations and Orchestration of Agents

    Stefanie Rinderle-Ma, Juergen Mangler, Johannes Loebbecke, Dominik Voigt, Nataliia Klievtsova, Matthias Ehrendorfer

    cs.AI

    Agentic Business Process Management has gained momentum recently. The prospect is that the autonomy of AI agents, i.e., predominantly LLM-based agents, can be balanced with a certain level of robustness, tractability, and traceability through a combination with process technology. In this paper, we provide a classification framework for agentic orchestration options along properties such as task specificity, traceability and tractability,...

    arxiv.org/abs/2606.31518 · PDF

  22. 22

    Surprise as a Signal for Plasticity and Metacognition

    Louis Mouchon

    cs.AI · cs.LG

    We study a single idea across two settings: that a prediction-error signal, computed by a small predictor over the latent space of a frozen encoder, can serve both as a gate on plasticity and as a substrate for metacognition. In the first system, a non-parametric episodic memory writes a new concept only when this surprise is high, and a periodic offline replay phase consolidates recent traces into a slow linear readout. On a continual stream...

    arxiv.org/abs/2606.31495 · PDF

  23. 23

    One Reflection Is Not Enough: Self-Correcting Autonomous Research via Multi-Hypothesis Failure Attribution

    Jie Ma, Binfei Chu, Jie Gao, Jinlu Zhang, Yiwei Ma, Yi Tan, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji

    cs.AI · cs.CV

    Autonomous research agents can now draft hypotheses, write code, run experiments, and produce papers, but they remain brittle when experiments fail. Under the prevailing paradigm, failure recovery is usually delegated to a single free-form reflection: a rich trajectory of metrics, logs, and design choices is compressed into one verbal critique, which often leads either to localized trial-and-error or to hard pivots that discard useful...

    arxiv.org/abs/2606.31478 · PDF

  24. 24

    CLOUDADV: Decision-Aligned Instance Sizing with Zero-Shot Foundation Models under Drift

    Jack Bell, Giacomo Carfi, Gerlando Gramaglia, Andrea Simioni, Daniele Fontani, Vincenzo Lomonaco

    cs.AI

    Cloud virtual machines are often overprovisioned, creating avoidable cost and operational inefficiency. We present CLOUDADV, an interactive engineer-facing advisory system for cloud instance sizing under workload drift. The system combines zero-shot time-series forecasting with bounded recommendation generation across day-, week-, and month-scale planning horizons. For each query, CLOUDADV constructs a structured decision context from...

    arxiv.org/abs/2606.31470 · PDF

  25. 25

    CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market

    Yao Shi, Kingfung Luo, Nan Tang, Yuyu Luo

    cs.AI · cs.CE

    Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and platform rules. These properties make them hard for traditional quantitative models, but provide an ideal testbed for studying how large language models (LLMs) turn unstructured text into trading actions. We present CSTrader, a multi-agent framework for language-grounded trading in the CS2 skin market. The...

    arxiv.org/abs/2606.31461 · PDF

  26. 26

    Who Determines the Meaning of an Emotion? Affective Sovereignty as an Epistemic Consequence of Measurement Limits

    Keito Inoshita

    cs.AI

    Emotion-sensing AI is rapidly becoming embedded in vehicles, home appliances, dialogue agents, and social infrastructure, giving rise to a sphere in which emotion is no longer confined to individual experience but is instead observed and computed at a societal scale, a domain we term the Affectosphere. Yet a central normative question in this domain has remained underexplored: who has the final authority to determine the meaning of one's own...

    arxiv.org/abs/2606.31442 · PDF

  27. 27

    CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes

    Yuchen Huang, Xiang Li, Zhenqing Ling, Sijia Li, Qianli Shen, Daoyuan Chen, Yi R. Fung, Yaliang Li

    cs.AI · cs.CL

    Data refinement involves executing multi-step recipes over evolving text states, where both composition and execution order of processing operators determine the outcome. While existing benchmarks either isolate text editing or entangle it with code and tool execution, it remains unclear whether LLMs can directly and faithfully execute these compositional, order-sensitive data refinement recipes. To fill this gap, we introduce CDR-Bench, a...

    arxiv.org/abs/2606.31435 · PDF

  28. 28

    Ask the World Before Acting: Budgeted Environment Probing for World-Model Calibration

    Xinyuan Song, Zekun Cai

    cs.AI

    Long-horizon language agents do not only choose actions; they carry a private model of the world from one decision to the next. When that model drifts, a later failure can be decided before the failing action is ever taken. We study a direct repair mechanism: before committing to the next task action, an agent may ask the environment about one belief field and write the answer back into its world model. This makes environment interaction a...

    arxiv.org/abs/2606.31422 · PDF

  29. 29

    BP-TTA: Balanced and Prototype-Guided Test-Time Adaptation in Dynamic Scenarios

    Shaoyang Huang, Yashi Zhu, Yichen Yu, Lei Zhang, Zhang Yi, Tao He

    cs.AI

    Test-Time Adaptation (TTA) enables models trained on a source domain to adapt online to unlabeled test data under distribution shifts. While recent TTA methods have moved beyond static settings and begun to consider continual domain shifts, they primarily address distribution drift and fail to account for class imbalance in dynamic scenarios. In real-world test-time streams, class imbalance and continual domain shifts often occur at the same...

    arxiv.org/abs/2606.31420 · PDF

  30. 30

    Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs

    Seyed Alireza Molavi, Zhan Su, Yan Hu, Peyman Sheikholharam Mashhadi, Stefan Byttner, Prayag Tiwari

    cs.AI · cs.LG

    Composing independently trained LoRA adapters into a single large language model is useful for multi-domain adaptation, especially when the original training data cannot be shared. A common approach is to use MoE-style routing over LoRA experts, but for frozen pretrained adapters, soft weighted combinations can change the unit-scale additive update under which each LoRA module was originally trained. We propose \textbf{Hard-Routed MoR-LoRA},...

    arxiv.org/abs/2606.31413 · PDF

  31. 31

    Xiaomi-GUI-0 Technical Report

    Wanxia Cao, Chengzhen Duan, Pei Fu, Pengzhi Gao, Niu Lian, Fazhan Liu, Hui Liu, Heng Qu, Qinzhuo Wu, Zhehao Yu,...

    cs.AI

    Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigation. However, existing GUI agents are trained and evaluated largely on offline trajectories, simulated environments, and standardized benchmarks. These differ substantially from real applications in interface layout, interaction logic, and...

    arxiv.org/abs/2606.31410 · PDF

  32. 32

    Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models

    Justin Brenne, Christian Meske

    cs.AI

    Human swarm intelligence demonstrates remarkable collective accuracy but faces scalability constraints in cost, coordination, and time. We investigate whether large language models (LLMs) can approximate swarm intelligence effects through artificial swarms, addressing a critical gap in understanding AI-based aggregation mechanisms. We conducted a controlled experiment with 960 manually executed prompts across three proprietary models (GPT-5,...

    arxiv.org/abs/2606.31404 · PDF

  33. 33

    World-Model Collapse as a Phase Transition

    Xinyuan Song, Zekun Cai

    cs.AI

    Water looks unchanged as it warms, then at a critical point it boils. We ask whether long-horizon language agents show an analogous transition in their implicit world models. In some parameter settings, changing state load by a small amount, or adding a single step of horizon, leaves behavior nearly unchanged; near a critical boundary, the same small change causes a sudden world collapse. We study this effect in a deterministic task family...

    arxiv.org/abs/2606.31399 · PDF

  34. 34

    ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents

    Binjie Zhang, Mike Zheng Shou

    cs.AI

    Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in practice. Existing works have two common gaps. Supervised fine-tuning (SFT) is built mostly on successful trajectories and offers little signal for recovery after tool failures, while sparse trajectory-level RL rewards provide limited guidance on which step failed and how to repair it. We introduce ReGRPO...

    arxiv.org/abs/2606.31392 · PDF

  35. 35

    Smart charging of large fleets of Electric Vehicles: Independent Multi-Agent Reinforcement Learning approaches

    Xavier Rate, Eloann Le Guern, Raphaël Féraud, Fatma Salem, Melissa Chiknoun, Eymeric Giabicani, Mehdi Feki, Patrick...

    cs.AI

    The electrification of transportation through electric vehicles introduces new challenges for power grid management, such as increased peak demand, voltage fluctuations, line overloads, and the integration of variable renewable energy sources. To enable efficient integration of EVs while minimizing costs for users and avoiding network overloads, implicit coordination between EVs is required. This work compares two independent multi-agent...

    arxiv.org/abs/2606.31347 · PDF

  36. 36

    Optimization Algorithms for Joint OFDM Waveform Design and RIS Configuration in 6G Networks: From Convex Relaxation to Foundation Models

    Ahmet Kaplan

    cs.AI

    Joint OFDM-RIS optimization for 6G is a mixed-integer nonlinear programming (MINLP) problem covering sum-rate maximization, energy efficiency, max-min fairness, and peak-to-average power ratio (PAPR)-constrained objectives. Seventy-eight joint OFDM-RIS optimization works published between 2021 and 2026 are surveyed. No standardized benchmark exists, and cross-paper comparisons remain infeasible. This survey classifies these works into four...

    arxiv.org/abs/2606.31334 · PDF

  37. 37

    CryoACE: An Atom-centric Framework for Accurate and Automated Model Building in Cryo-EM

    Minzhang Li, Mingrui Li, Weichen Qin, Qihe Chen, Sixian Shen, Yuan Pei, Jiakai Zhang, Jingyi Yu

    cs.AI

    Protein automodeling from cryo-EM density maps faces unique challenges in enforcing physicochemical validity and managing conformational heterogeneity. Current solvers are often limited to static predictions or require computationally intensive heuristic searches. We present CryoACE, an end-to-end framework that reconstructs precise atomic graphs for both homogeneous and heterogeneous structures. Our method features two key innovations: an...

    arxiv.org/abs/2606.31332 · PDF

  38. 38

    HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)

    Aurélien Pellet, Julien Perez, Marie Puren

    cs.AI

    We present HistoriQA-ThirdRepublic: a French-language dataset of multi-hop historical questions derived from parliamentary debates and newspapers of the French Third Republic. Designed in collaboration with a historian, the corpus captures complex reasoning patterns typical of historical inquiry, including cross-source synthesis, temporal reasoning, and the integration of sparse evidence. The dataset is made of 1782 questions and emphasizes...

    arxiv.org/abs/2606.31325 · PDF

  39. 39

    Benchmarking Large Language Models on Floating-Point Error Classification

    Lisa Taldir, Muhammad Ahmad Saeed, David Defour, Pablo de Oliveira Castro, Eric Petit

    cs.AI

    This paper investigates the capability of Large Language Models (LLMs) to detect and classify floating-point errors statically in software code. We introduce InterFLOPBench, a benchmark of 90 C kernels with 1 130 test samples designed to evaluate LLMs across six categories of floating-point error: cancellation, comparison, division by zero, overflow, underflow and NaN, compared across 14 LLMs. The evaluation framework treats floating-point...

    arxiv.org/abs/2606.31308 · PDF

  40. 40

    Spatial Reasoning via Modality Switching Between Language and Symbolic Representation

    Shreya Rajpal, Tanawan Premsri, Parisa Kordjamshidi

    cs.AI

    Human reasoning is inherently multimodal: when problems become difficult, we rarely think in words alone. We often externalize our reasoning by sketching diagrams or drawing grids to understand the underlying conceptual structure and avoid mistakes. Building on this premise, our research investigates: (a) whether grounding multi-hop textual-spatial stories into geometry-aware modalities, such as layouts or grids, improves reasoning compared...

    arxiv.org/abs/2606.31285 · PDF

  41. 41

    Embodied CAD: Solver-Grounded LLM Agents for Parametric B-Rep Assembly Modeling

    Fumin Liu, Haoyu Zhou, Fei Hao, Lin Yang

    cs.AI

    Large language models can write plausible CAD scripts, but reliable industrial CAD modeling requires more than syntactically valid code: every feature, placement, and assembly relation must be accepted by an exact geometric kernel while remaining editable as parametric boundary representation geometry. We present Embodied CAD, solver-grounded LLM agents for parametric B-Rep assembly modeling. Instead of generating a complete script in one...

    arxiv.org/abs/2606.31252 · PDF

  42. 42

    Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

    Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Tianyu Zong, Hongzhu Yi, Guoqing Chao,...

    cs.AI

    Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations. We propose Delta-JEPA, an end-to-end reconstruction-free world model that augments latent forward prediction with a Latent Difference Action Decoder (LDAD). Unlike inverse decoders that infer actions from concatenated endpoint...

    arxiv.org/abs/2606.31232 · PDF

  43. 43

    Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents

    Keyu Zhao, Lingyan Kong, Fengli Xu, Yong Li

    cs.AI

    Ideation plays a pivotal role in scientific discovery. Recent LLM, especially AI Scientist systems, show promising potential for automated ideation. However, existing approaches predominantly rely on pre-defined agentic workflows. This constraint severely limits the flexibility required to navigate the vast search space of scientific literature and the complex action space of research reasoning. Recently, training Agentic LLMs has emerged as...

    arxiv.org/abs/2606.31229 · PDF

  44. 44

    Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism

    Gunho Jung, Jeong-Woo Park, Seon Bin Kim, Seong-Whan Lee

    cs.AI

    Composed image retrieval requires identifying a target image from a gallery by integrating a reference image with a textual modification instruction. In a training-free zero-shot setting, this task relies on constructing a retrieval-oriented textual query within a frozen vision--language embedding space at inference time. Existing approaches predominantly rely on a single-pass generation strategy that fuses the reference context and...

    arxiv.org/abs/2606.31222 · PDF

  45. 45

    Long-term Traffic Simulation via Structured Autoregressive Modeling

    Lingyu Xiao, Zexin Feng, Xintao Yan

    cs.AI · cs.RO

    Interactive traffic simulation is a vital world model for autonomous driving. A central challenge in long-horizon simulation is modeling sustained multi-agent interactions, which is further exacerbated by dynamic token cardinality as agents continuously enter and exit the scene. In this work, we propose that the solution lies in the synergy between the architectural inductive biases and statistical priors of large-scale sequence models, e.g.,...

    arxiv.org/abs/2606.31209 · PDF

  46. 46

    Towards Inclusive Mobility Modeling: Characterizing and Evaluating Elderly Trajectory Patterns in Urban Systems

    Zhengxuan Wang, Haohan He, Mengying Zhou

    cs.AI · cs.CY

    The rapid advance of smart cities increasingly depends on trajectory data mining, yet underrepresented demographic groups, particularly the elderly, are often sparsely represented in public mobility datasets. This underrepresentation can introduce systematic bias into mobility modeling and downstream urban planning. Using the 2016-2020 Jersey City subset of the Citi Bike System Data, this study quantitatively examines how the absence of...

    arxiv.org/abs/2606.31207 · PDF

  47. 47

    Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping

    Tao Chen, Lizheng Liu, Jiaxu Wang, Ziyue Jiang, Ruiqi Tian, JiGuang Huo, Zhongxue Gan

    cs.AI

    Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based methods rely on visual similarity for object matching, neglecting physical affordances such as handle graspability and material fragility, and operate open-loop without spatial reasoning or failure recovery, limiting their effectiveness when objects are densely packed or physically diverse. We...

    arxiv.org/abs/2606.31200 · PDF

  48. 48

    AI-Assisted Discovery of Convex Relaxations via Dual Agents

    Sungyoon Kim, Mert Pilanci

    cs.AI

    Recent work shows that LLM agents can improve sharp-constant inequalities by searching for extremal constructions, which yield upper bounds. We address the complementary side: a lower bound holds for every admissible function and follows from a convex relaxation of the nonconvex problem, with tighter relaxations giving stronger bounds. We instantiate the autoresearch paradigm to discover such relaxations: a coding agent proposes valid...

    arxiv.org/abs/2606.31182 · PDF

  49. 49

    HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

    Qianchu Liu, Sheng Zhang, Guanghui Qin, Jeya Maria Jose Valanarasu, Maximilian Rokuss, Mingyu Lu, Timothy Ossowski,...

    cs.AI · cs.CL · cs.CV

    As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories each with its unique environment. The benchmark suite spans diverse workflows throughout the patient journey and a broad range of modalities. Each task is designed to...

    arxiv.org/abs/2606.31179 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.