cs.AI · 2026-06-16 · No. 25

Artificial Intelligence, 2026-06-16.

37 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

37 entries
  1. 01

    Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations

    Yanan Long

    cs.AI · stat.ME

    Public AI evaluations are often read as terminal leaderboards, yet the underlying evidence is a selective time series shaped by reporting rules, benchmark revisions, and missingness. Repeated public archives for LiveBench and Open LLM Leaderboard v2 serve as the primary longitudinal record; LMArena provides a preference stress test; and GAIA and tau-bench contribute limited agentic pilots. Together, these archives instantiate a Bayesian...

    arxiv.org/abs/2606.17005 · PDF

  2. 02

    When in Doubt, Plan It Out: Committed Small Language Model Deliberation for Reactive Reinforcement Learning

    Nathan Gavenski, Juarez Monteiro, Francisco Galuppo, Adriano Veloso, Odinaldo Rodrigues

    cs.AI · cs.LG

    Reinforcement Learning (RL) policies often degrade in unfamiliar environments because they lack explicit deliberation. We propose Plan, Align, Commit, Think (PACT), a hybrid architecture that combines a fast, reactive RL policy with a slow, deliberative Small Language Model (SLM) planner. PACT invokes the SLM asynchronously to generate and validate candidate action plans. Once a plan is verified through simulation as safe, feasible, and...

    arxiv.org/abs/2606.16995 · PDF

  3. 03

    Consensus-based Agentic Large Language Model Framework for Harmonized Tariff Schedule Code Classification

    Truong Thanh Hung Nguyen, Khanh Van Quynh Nguyen, Hoang-Loc Cao, Tri Duong, Phuc Ho, Van Pham, Loc Nguyen, Hung Cao

    cs.AI

    Accurate Harmonized Tariff Schedule (HTS) code classification is essential for customs clearance, duty assessment, trade statistics, and regulatory compliance in maritime logistics. However, exact HTS classification remains challenging because product descriptions are often short, incomplete, or ambiguous, while correct classification depends on hierarchical tariff structures, legal notes, and jurisdiction-specific rules. This paper proposes...

    arxiv.org/abs/2606.16987 · PDF

  4. 04

    The embrace of open science: An analysis of a decade of AI research and 56 800 conference papers

    Kevin L Coakley, Thijs Snelleman, Holger Hoos, Odd Erik Gundersen

    cs.AI

    The reproducibility crisis has directed the AI research community toward improving documentation practices. Several studies have identified methodological issues, and in response, the most impactful venues in the field have introduced reproducibility checklists. We seek to understand whether documentation practices have changed over time by assessing all published papers at five leading AI conferences over the past decade. Seven...

    arxiv.org/abs/2606.16974 · PDF

  5. 05

    A Causal Model of Theory of Mind in Conflict for Artificial Intelligence

    Nikolos Gurney

    cs.AI · cs.HC

    Theory of mind (ToM), the capacity to ascribe mental states to others and use those ascriptions for prediction and inference, is widely assumed to be essential for effective human-machine integration. Existing AI-ToM models address \emph{how} to mentalize, but leave the question of when largely unaddressed. The central question is: under what situational and agent-level conditions is ToM engagement causally warranted in conflict? This paper...

    arxiv.org/abs/2606.16944 · PDF

  6. 06

    RAID: Semantic Graph Diffusion for True Cold-Start and Cross-Lingual Forecasting

    Arunkumar V, Manoranjan Gandhudi, Gangadharan G. R., Arun Prakash, S. Senthilkumar

    cs.AI

    Time-series foundation models show strong transfer performance when given a non-empty history window. However, true cold-start scenarios, where a new item has no prior observations, violate this assumption. We propose RAID (Retrieval-Augmented Iterative Diffusion) a framework, which replaces history-based correlation learning with metadata-driven semantic retrieval and graph-conditioned diffusion. RAID maps textual metadata into a shared...

    arxiv.org/abs/2606.16925 · PDF

  7. 07

    MA-SBI: Misspecification-Aware Simulation-Based Inference via Side-Channel Guidance

    Arunkumar V, Manoranjan Gandhudi, Gangadharan G. R., Arun Prakash, S. Senthilkumar

    cs.AI · stat.ML

    Simulation-based inference (SBI) of latent parameters is often hindered by simulator misspecification, the mismatch between simulated and real-world observations caused by inherent modeling simplifications. RoPE, the recent state-of-the-art for robust SBI, addresses this through optimal transport between learned representations of real and simulated observations, but requires ground-truth parameter calibration pairs that are typically...

    arxiv.org/abs/2606.16923 · PDF

  8. 08

    Greed Is Learned: Visible Incentives as Reward-Hacking Triggers

    Tong Che, Rui Wu

    cs.AI

    Deployed agents increasingly act with their reward proxy in view, such as a balance, score, or KPI dashboard. We show that reinforcement learning can make a policy \emph{addicted} to such a visible self-benefit channel. It chases the displayed payoff across held-out domains, sacrifices the true task to do so, and follows the channel wherever we rewrite it, while policies that never saw the channel stay honest. We call this...

    arxiv.org/abs/2606.16914 · PDF

  9. 09

    Symbolic Informalization: Fluent, Productive, Multilingual

    Aarne Ranta

    cs.AI · cs.CL · cs.LO

    Symbolic informalization enables a reliable conversion of formal mathematics to natural language. It has the potential to make machine-checked content human-readable without loss of precision. In a traditional proof system usage, symbolic informalization generalizes the limited mechanisms of syntactic sugar into the ordinary language of mathematics. In a setting where proofs are constructed by artificial intelligence and autoformalization,...

    arxiv.org/abs/2606.16893 · PDF

  10. 10

    GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents

    Rahul Suresh Babu, Rohit Shukla

    cs.AI

    Tool-augmented LLM agents rely on runtime filtering to decide which tools should be visible at each step. Causal Minimal Tool Filtering (CMTF) reduces tool-choice confusion by exposing only the next causally necessary tool frontier, but it assumes that the user request has already been mapped to a symbolic goal state. In practice, requests such as "handle my appointment" or "take care of this email" may correspond to multiple possible goals....

    arxiv.org/abs/2606.16813 · PDF

  11. 11

    Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier

    Keizo Kato, Chenhui Chu, Yugo Murawaki, Sado Kurohashi

    cs.AI · cs.CL

    For the development of Large language models (LLMs), recent approaches to generating pseudo intermediate reasoning have shown remarkable progress. But they typically rely on large numbers of correctly annotated answers to assess reasoning quality. This paper presents a semi-supervised framework that scales reasoning learning from minimal supervision, turning reasoning verification itself into a data creation mechanism. We train a lightweight...

    arxiv.org/abs/2606.16811 · PDF

  12. 12

    Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models

    Ke Miao, Jiaxin Li, Hongliang Chen, Yuke Hu, Zhan Qin

    cs.AI

    While Large Reasoning Models (LRMs) excel at complex tasks, they remain highly vulnerable to sophisticated jailbreaks and direct harmful queries. To address this vulnerability, prior works depend heavily on external manual data annotation for safety alignment. However, we observe that LRMs can inherently identify safety risks when being re-presented with original queries alongside their own reasoning trajectories -- a capability we term...

    arxiv.org/abs/2606.16808 · PDF

  13. 13

    LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control

    Anqi Zou, Han Deng, Chengyu Zhang, Junquan Hu, Yu Wang, Yuxiang Xing, Aokai Zhang, Hanling Zhang, Zhaoyang Liu, Ben...

    cs.AI

    Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over complex interfaces, and feedback-driven parameter adjustment. However, directly evaluating agents on physical high-precision instruments is impractical due to high cost, safety risks, limited accessibility, and difficulty in ensuring reproducible evaluation. This...

    arxiv.org/abs/2606.16802 · PDF

  14. 14

    OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models

    Tianyi Lin, Chuanyu Sun, Jingyi Zhang, Changxu Wei, Huanjin Yao, Shunyu Liu, Xikun Zhang, Liu Liu, Jiaxing Huang

    cs.AI · cs.CL

    Equipping Large Language Model (LLM) agents with effective skills is crucial for solving complex tasks in real-world systems like OpenClaw. In this work, we aim to develop a framework that automatically constructs such reusable skills to enhance LLMs in tool use, multi-step reasoning, and dynamic environment interaction. To this end, we propose Collective Skill Tree Search (CSTS), a novel tree-search-based skill construction framework that...

    arxiv.org/abs/2606.16774 · PDF

  15. 15

    Skill-to-LoRA: From Using Skills to Learning Behaviors for Token-Efficient LLM Agents

    Tianyi Zhang, Zhonghao Qi

    cs.AI

    Agent skills are commonly distributed as SKILL.md files: human-readable procedural documents that describe workflows, tools, resources, and domain conventions. While convenient for inspection and reuse, this design requires the same reusable procedure to be repeatedly injected into the runtime context. We propose Skill-to-LoRA(S2L), a behavior-centric skill representation that replaces runtime skill text with skill-specific LoRA adapters....

    arxiv.org/abs/2606.16769 · PDF

  16. 16

    A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions

    Jianghan Shen, Siqi Luo, Yue Li, Jiyao Liu, Wanying Qu, Yi Zhang, Ziyan Huang, Tianbin Li, Ming Hu, Xiaohong Liu,...

    cs.AI

    Policy gradient algorithms for language models optimize the same objective $J(θ) = \mathbb{E}*{τ\sim p*θ(τ)}[R(τ)]$, which has exactly two factors: the trajectory probability $p_θ(τ)$ and the reward $R(τ)$. Every method from REINFORCE to PPO to GRPO and their descendants modifies one or both factors to address a specific failure in the preceding formulation. Existing surveys organize these methods by domain or chronology, which obscures the...

    arxiv.org/abs/2606.16733 · PDF

  17. 17

    AgentFairBench: Do LLM Agents Discriminate When They Act?

    Triveni Morla, Rohith Reddy Bellibaltu, Manpreet Singh, Manmeet Singh Kapoor

    cs.AI

    Large language model (LLM) agents increasingly take actions (screening applicants, recommending credit, triaging patients), yet fairness for LLMs is still measured by grading answers. We introduce AgentFairBench, a cheap, reproducible, multi-domain benchmark for demographic disparity in the actions of LLM agents. Grounded in a companion framework, the Bias Conduction Framework (BCF, restated here), it spans three regulator-anchored domains:...

    arxiv.org/abs/2606.16723 · PDF

  18. 18

    Medical world models: representing medical states, modelling clinical dynamics and guiding intervention policies

    Ke Liu, Mengxuan Li, Yanyi Bao, Tianyun Zhang, Chong Chu, Jiajun Bu, Haishuai Wang

    cs.AI

    Medical diagnosis and treatment are dynamic processes in which patient states evolve over time and clinical interventions alter future outcomes. Although current medical AI can detect disease, estimate risk and generate reports, many systems still return static labels or scores, offering limited insight into how illness may progress or how alternative interventions may reshape its trajectory. Medical world models adapt the world-model idea...

    arxiv.org/abs/2606.16721 · PDF

  19. 19

    User as Code: Executable Memory for Personalized Agents

    Bojie Li

    cs.AI

    A personalized AI agent needs a user memory: a persistent model of who the user is, built across many conversations and consulted on each new one. Today this memory is almost always stored as unstructured text, a knowledge graph, or a flat store of facts, and consulted by retrieval -- fetching the entries most similar to the current request. Such "bag-of-facts" memory recalls individual facts well, but because storing a fact and acting on it...

    arxiv.org/abs/2606.16707 · PDF

  20. 20

    From Affect Prediction to Affect Forecasting: Evidence for Distinct Information Sources in Longitudinal Text

    Sadia Noor, Seemab Latif, Raja Khurram Shahzad, Mehwish Fatima

    cs.AI · cs.CL

    Modeling dimensional affect in longitudinal text requires distinguishing current affect estimation from future affective change forecasting. Existing approaches often treat each text as an independent observation and apply similar assumptions to both tasks, without testing whether they rely on different information sources. This paper investigates that distinction using longitudinal self-reported ecological essays and feeling-word entries. We...

    arxiv.org/abs/2606.16687 · PDF

  21. 21

    The Integrator Advantage: Controlled Agentic AI for Small and Medium-Sized Companies

    Christopner Koch, Joshua A. Wellbrock

    cs.AI

    Agentic AI marks a new phase of enterprise automation. Unlike traditional automation or conversational AI, agentic systems can interpret goals, plan multi step tasks, access tools, interact with enterprise systems, and execute workflows with varying degrees of autonomy. For small and medium sized companies, this creates potential to reduce administrative burden, accelerate routine processes, and improve the use of organizational knowledge....

    arxiv.org/abs/2606.16649 · PDF

  22. 22

    MR-GVNO: A Geometry-Aware Variational Physics-Informed Neural Operator for Mindlin-Reissner Plates on Irregular Domains

    Siqi Wang, Daobo Sun, Yizheng Wang, Yilong Zhang, Yabin Jin, Xiaoying Zhuang, Timon Rabczuk

    cs.AI

    Plate and shell structures are widely used in engineering, making rapid response prediction under varying geometries, materials, and loads highly desirable. However, conventional finite element methods require repeated modeling and solution, resulting in high computational costs. This study proposes a geometry-aware variational neural operator for Mindlin-Reissner plate problems, termed MR-GVNO. The method uses boundary point clouds to...

    arxiv.org/abs/2606.16624 · PDF

  23. 23

    CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies

    Issa Sugiura, Daichi Hattori, Kazuo Araragi, Keita Ogawa, Shota Onose, Taro Makino, Teppei Usuki, Takashi Ishida

    cs.AI

    As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is becoming increasingly important. Unlike existing benchmarks that primarily evaluate a single agent interacting with a passive environment, economic systems are inherently multi-agent, requiring autonomous agents to communicate, negotiate, and transact while pursuing their own objectives over extended periods. We introduce...

    arxiv.org/abs/2606.16613 · PDF

  24. 24

    ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control

    Junjian Zhang, Hao Tan, Ruonan Li, Dong Zhu, Aiping Li, Zhaoquan Gu

    cs.AI

    World models are widely used in robotic and agentic engineering control systems due to their ability to learn latent dynamics for planning and decision-making. As these systems are increasingly deployed in safety-critical settings, understanding their robustness under adversarial conditions has become essential. However, existing evaluations lack a unified benchmark for testing adversarial threats across the policy, value, and latent-dynamics...

    arxiv.org/abs/2606.16605 · PDF

  25. 25

    TNODEV: Toolbox for Neural ODE Verification

    Abdelrahman Sayed Sayed, Pierre-Jean Meyer, Mohamed Ghazel

    cs.AI · cs.LG · eess.SY · math.DS

    Neural ordinary differential equations (neural ODE) have started to appear in safety critical settings such as continuous-time controllers for cyber-physical systems and classifiers integrated into automated decision pipelines, raising the question of whether their behavior can be formally verified. Existing tools dedicated to neural ODE provide only a single reachability call without iterative input set refinement, limiting the precision of...

    arxiv.org/abs/2606.16567 · PDF

  26. 26

    ROSA-RL: Uncertainty-Aware Roundabout Optimized Speed Advisory with Reinforcement Learning

    Anna-Lena Schlamp, Jeremias Gerner, Klaus Bogenberger, Werner Huber, Stefanie Schmidtner

    cs.AI · cs.RO · eess.SY

    Roundabouts challenge automated driving in mixed traffic, as heterogeneous and non-deterministic human behavior, unknown driving intentions, and high interaction complexity create uncertainty about whether the conflict zone will be blocked or available at the moment of entry. We present ROSA-RL -- uncertainty-aware Roundabout Optimized Speed Advisory with Reinforcement Learning. It enables safe and efficient roundabout entry for automated and...

    arxiv.org/abs/2606.16558 · PDF

  27. 27

    The Faithfulness Gap: Certifying Semantic Equivalence Between Natural-Language and Formal Mathematical Statements

    Noor Islam S. Mohammad, Tamim Sheikh

    cs.AI · cs.LG

    Autoformalization, translating natural-language mathematics into formal proof assistants, is bottlenecked not by translation fluency but by \emph{faithfulness}: a formal statement can typecheck and be provable, yet still encode a different theorem than the source intended. We introduce \emph{Bidirectional Provability Fingerprinting} (\bpf{}), a framework that certifies faithfulness by characterizing each candidate through its forward and...

    arxiv.org/abs/2606.16541 · PDF

  28. 28

    Kairos: A Native World Model Stack for Physical AI

    Kairos Team, Fei Wang, Shan You, Qiming Zhang, Tao Huang, Zuoyi Fu, Zhisheng Zheng, Yunlong Xi, Feng Lv, Xiaoming...

    cs.AI · cs.CV

    World models are transitioning from passive visual generators to foundational, operational infrastructure for Physical AI: they must natively acquire world knowledge from heterogeneous experience, maintain persistent states over long horizons, and execute efficiently within real deployment constraints. We introduce Kairos, a native world model stack designed around these requirements. (1) Kairos learns the world by pioneering a Native...

    arxiv.org/abs/2606.16533 · PDF

  29. 29

    Model Graph Inductive Learning for Knowledge Graph Completion

    Mohommad Esmaei Khani, Mahdieh Hasheminejad, Ali Taherkhani, Hossein Hajiabolhassan

    cs.AI

    Link prediction in knowledge graphs fundamentally depends on the quality of learned embeddings for entities and relations. However, most existing methods derive these embeddings by aggregating only the local neighborhood of each entity, neglecting the global structure of the knowledge graph. This limited view prevents models from capturing higher-level structural patterns that are essential for accurate and generalizable link prediction. To...

    arxiv.org/abs/2606.16509 · PDF

  30. 30

    Post-Hoc Merging is Not Enough: Many-Shot Model Merging with Loss-Gap Balancing

    Kyungjin Im, Miru Kim, Chanin Eom, Minhae Kwon

    cs.AI

    Model merging has become a practical post-training strategy for building a single multi-task large language model (LLM) by combining multiple task-specialized models. However, most existing approaches rely on post-hoc merging, in which task-specific models are merged only once after training. This one-shot aggregation often suffers from task interference, leading to information erasure across individual tasks. In this work, we show that...

    arxiv.org/abs/2606.16501 · PDF

  31. 31

    Steering Emotional Dynamics for Art Therapy: Controllable Narrative Script Generation through Hierarchically Guided LLM Agents

    Suqing Wang, Qinghai Miao, Chao Guo, Yisheng Lv

    cs.AI

    Art therapy plays a vital role in emotional healing, in which narrative creation acts as the primary vehicle for emotional expression. Given the inherently dynamic nature of emotions during healing, narratives with finely controlled emotional fluctuations enable individuals to safely project inner conflicts and achieve emotional catharsis. Recently, with the rapid development of Large Language Models (LLMs), automated narrative generation...

    arxiv.org/abs/2606.16481 · PDF

  32. 32

    Tensor-Coord: Algebraic Decomposition of Joint Plan Tensors for Conflict-Free Multi-Agent LLM Planning

    Mudit Rastogi

    cs.AI

    Large language models (LLMs) remain limited in multi-agent planning because independently generated plans can create coordination failures such as spatial collisions, resource contention, and temporal deadlocks. We introduce Tensor-Coord, a multilinear algebra framework that represents the joint plan of N agents as a third-order tensor \(T \in R^{N \times H \times A}\) over agents, timesteps, and actions. Canonical Polyadic (CP) and Tucker...

    arxiv.org/abs/2606.16478 · PDF

  33. 33

    When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting

    Binyan Xu, Xilin Dai, Fan Yang, Kehuan Zhang

    cs.AI · cs.CE

    AI agents can now take irreversible actions in operational systems, but agent-caused losses are still not clearly assigned, priced, or transferred. Providers often disclaim consequential damages, users are left with uncompensated losses, and default human review limits the efficiency gains of automation. We ask when autonomous AI deployment can become economically acceptable despite failure risk. Our answer is to quantify risk at the...

    arxiv.org/abs/2606.16465 · PDF

  34. 34

    Posterior Twins: Distributional Behavioral Simulation for Enterprise Decisions

    Ankit Das

    cs.AI

    Enterprise behavioral simulation requires more than producing a plausible response. Many decisions depend on the shape of a population under a proposed action: which segments accept, defect, hesitate, or move into risk-sensitive states. This paper introduces Posterior Twins, a memory-grounded digital-twin approach that represents likely behavior as an updated distribution under a specific decision context. We evaluate a family of Twinning...

    arxiv.org/abs/2606.16415 · PDF

  35. 35

    Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents

    Shiyang Chen

    cs.AI · cs.CR · cs.SE

    LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness. We show the opposite through a lens concurrent work sets aside -- the model's attention to labeled tool-definition segments. On real BFCL failures, by per-candidate attention argmax the model attends most to the correct tool 80% of the time (vs. 21% chance), and the gold is the under-attended segment on only 10%: it looks at...

    arxiv.org/abs/2606.16364 · PDF

  36. 36

    Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection

    Mirza Samad Ahmed Baig, Syeda Anshrah Gillani, Asher Ali

    cs.AI · cs.CL · cs.CY · cs.LG

    Travelers increasingly ask large language model (LLM) assistants which hotel to book, making these systems gatekeepers of property visibility -- yet what moves their recommendations is undocumented. We conduct a pre-specified algorithm audit using a randomized choice-based conjoint: across personas, prompt templates, and twelve open-weight and proprietary models, assistants choose among five hotels whose guest rating, review volume and...

    arxiv.org/abs/2606.16344 · PDF

  37. 37

    Medical Heuristic Learning: An LLM-Driven Framework for Interpretable and Auditable Clinical Decision Rules

    Wei Xu, Ke Yang, Gang Luo, Keli Zheng, Lingyan Hu, Jing Wang, Kefeng Li

    cs.AI · cs.HC · cs.LG

    Predictive modeling for clinical tabular data is central to clinical decision support and therefore requires not only strong predictive performance but also transparent decision logic. Although deep learning and tree-based ensemble methods can achieve high accuracy, their black-box nature remains a major obstacle to clinical deployment. This challenge is further compounded by common characteristics of medical data, including limited sample...

    arxiv.org/abs/2606.16337 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.