cs.AI · 2026-06-30 · No. 39

Artificial Intelligence, 2026-06-30.

41 new papers in cs.AI. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

41 entries
  1. 01

    Self-Evolving World Models for LLM Agent Planning

    Xuan Zhang, Wenxuan Zhang, See-Kiong Ng, Yang Deng

    cs.AI · cs.CL

    World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequences before execution. However, unreliable foresight can be ignored, misused, or even degrade downstream decision-making. In this paper, we introduce WorldEvolver, a self-evolving world model framework that revises its deployment-time context while keeping the downstream agent and all model parameters frozen. WorldEvolver...

    arxiv.org/abs/2606.30639 · PDF

  2. 02

    DOPD: Dual On-policy Distillation

    Xinlei Yu, Gen Li, Qingyi Si, Guibin Zhang, Yuqi Xu, Congcong Wang, Shuai Dong, Kaiwen Tuo, Xiangyu Zeng, Kaituo...

    cs.AI

    On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sources and thereby elevate the performance frontier of distillation, an intuitive direction is to infuse privileged information to either teacher or student itself. However, this additional input induces a potential failure mode we dub privilege illusion: a pattern that...

    arxiv.org/abs/2606.30626 · PDF

  3. 03

    The Human Creativity Benchmark

    Aspen Hopkins, Allison Nulty, Alexandria Minetti, Anoop Pakki, Angad Singh

    cs.AI · cs.CV · cs.HC

    Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflects genuine differences in taste, not measurement error. We argue that evaluating creative AI requires preserving two distinct signals: convergence, where professionals align around shared best practices, and divergence, where individual taste legitimately varies. We present the Human Creativity Benchmark...

    arxiv.org/abs/2606.30561 · PDF

  4. 04

    Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing

    Dvir Alsheich, Adar Peleg, Ben Hagag, Rom Himelstein, Amit Levi, Avi Mendelson

    cs.AI · cs.MA

    The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents collaborate to execute complex workflows. Effective orchestration in these environments requires robust routing mechanisms to efficiently allocate tasks to the most suitable agent. However, existing routers fundamentally rely on unverified proxies, ranging from textual self-descriptions to static surrogate...

    arxiv.org/abs/2606.30555 · PDF

  5. 05

    Latent Actions from Factorized Transition Effects under Agent Ambiguity

    Heejeong Nam, Chandradithya S Jonnalagadda, Harshit Aggarwal, Eric Xu, Randall Balestriero

    cs.AI

    Latent Action Models (LAMs) learn action-like proxies from observation transitions. However, in multi-object or distractor-rich scenes, these visual effects mix agent motion with distractors, camera dynamics, and background changes, making the underlying action source ambiguous without supervision. Structuring this mixture as reusable transition effects provides an intermediate representation from which action-like latents can be more...

    arxiv.org/abs/2606.30544 · PDF

  6. 06

    Entity Binding Failures in Tool-Augmented Agents

    Rahul Suresh Babu, Shashank Indukuri

    cs.AI

    Tool-augmented language-model agents are often evaluated by whether they select the correct tool, produce valid API arguments, and complete the requested task. However, an agent may choose the right tool and still act on the wrong external entity. For example, a request to "email Alex about the launch" may lead the agent to contact the wrong Alex, attach the wrong launch document, reply in the wrong thread, or update the wrong customer...

    arxiv.org/abs/2606.30531 · PDF

  7. 07

    The FIL Hypothesis: Inductive Biases Help with Kernel Engineering

    Nikolai Rozanov, Subhabrata Dutta, Preslav Nakov, Iryna Gurevych

    cs.AI

    The Bitter Lesson, which posits that general-purpose methods that scale with computation and data ultimately outperform those with built-in human knowledge, has become a dominant paradigm in the era of Large Language Models. We revisit this principle by observing a new and critical scaling dimension: the duration of the Feedback Information Loop (FIL), the time required for a system to receive a verification signal after generating a...

    arxiv.org/abs/2606.30442 · PDF

  8. 08

    ENC-ODE: Event-level Neurodegenerative Modeling in Continuous Time with Neural ODEs

    Yujee Song, Seunghun Baek, Guorong Wu, Won Hwa Kim

    cs.AI · cs.IR · cs.LG

    Accurately predicting the temporal evolution of clinical biomarkers is crucial for the early diagnosis and management of neurodegenerative diseases such as Alzheimer's disease. However, this relies on longitudinal data to capture biomarker changes over time, which is often sparse and irregular due to the high cost, labor-intensive nature, and patient burden. To address these challenges, we propose ENC-ODE, an Event-level Neurodegenerative...

    arxiv.org/abs/2606.30398 · PDF

  9. 09

    Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents

    Bojie Li, Noah Shi

    cs.AI

    A rapidly growing class of LLM agents is multi-party: the agent acts for a principal (who briefs it, sends follow-ups, and receives results) while also conversing in a separate channel with a counterparty whose interests may diverge (negotiating with a vendor, screening inbound requests, or mediating between employees). Here "help whoever you are talking to" is the wrong objective. The agent must stay loyal to the principal it represents...

    arxiv.org/abs/2606.30383 · PDF

  10. 10

    Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data

    Haobo Yang

    cs.AI · cs.CY · cs.HC

    Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject to sampling bias. Here we show that pretrained large language models induce risk-equivalent estimators of conditional expectations under squared loss, establishing restricted functional risk equivalence: under squared loss, the LLM induces an estimator whose risk matches the Bayes optimal risk for...

    arxiv.org/abs/2606.30372 · PDF

  11. 11

    Sequential Fairness Auditing with Limited Output Access

    Ioannis Pitsiorlas, Martha V. Sourla, Marios Kountouris

    cs.AI

    External evaluations are becoming increasingly central to the governance of AI systems. In practice, however, independent auditors often have limited access to deployed models and must rely on query-based interactions. Most existing fairness evaluation methods assume static datasets and fixed-sample statistical tests, making them poorly suited to real-world auditing scenarios in which evidence must be collected sequentially under query...

    arxiv.org/abs/2606.30338 · PDF

  12. 12

    BayesEvolve: Explicit Belief States for Autonomous Scientific Discovery

    Xuening Wu, Shan Yu, Qianya Xu, Shenqin Yin

    cs.AI

    Autonomous scientific discovery systems increasingly use large language models (LLMs) to propose new hypotheses, but many such systems condition primarily on experimental memory: archives of high-scoring candidates or heuristic summaries of recent trials. We argue that discovery agents should instead maintain explicit, uncertainty-aware beliefs about hypothesis quality. We introduce BayesEvolve, a belief-guided discovery framework that...

    arxiv.org/abs/2606.30335 · PDF

  13. 13

    ManimAgent: Self-Evolving Multimodal Agents for Visual Education

    Wenjia Jiang, Zongyuan Cai, Yuanhang Shao, Chenru Wang, Boyan Han, Zhixue Song, Keyu Chen, Shengwei An, Xu Yang, Zhou Yang

    cs.AI

    Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isolated episode: lessons learned across many reflection rounds on one task are discarded before the next begins. We study this gap on a code-generation task: from a scientific paper section, the agent writes Python in the open-source Manim library to render a mathematical animation. We present ManimAgent, a...

    arxiv.org/abs/2606.30296 · PDF

  14. 14

    Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering

    Rahul Khedar, Mayank Malhotra, Avinash Karn, Mouli V, Prakhar Mehrotra

    cs.AI · cs.HC · cs.SE

    Live product demonstrations are a recurring, high-cost activity in software organizations: a human presenter must select features, dispatch the corresponding interactions on a running application, narrate them coherently, and answer questions in real time. Existing automation addresses only fragments -- generalist browser agents target instruction-conditioned task completion, and demo-video tools produce fixed MP4 artifacts that cannot be...

    arxiv.org/abs/2606.30294 · PDF

  15. 15

    PromptGNN-sim: Deep Fusion and Alignment of GNN and LLMs for Text-Attributed Graph Learning

    Zhifei Hu, Alexandra I. Cristea

    cs.AI

    Text-Attributed Graphs (TAGs) combine textual semantics with graph structure and are central to many graph learning tasks. However, existing fusion methods often treat text and structure as separate inputs in a shallow, one-way pipeline, which limits deep interaction between modalities and weakens performance under sparse connectivity or cross-graph generalisation. To address this issue, we propose PromptGNN-sim, a bi-directional...

    arxiv.org/abs/2606.30291 · PDF

  16. 16

    EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots

    Camilo Chacón Sartori

    cs.AI · cs.CY

    Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure. For emotional-support chatbots, that bargain hides precisely where safety failures emerge: across a multilingual, multi-turn crisis conversation. We present EMPATH, a benchmark for safety evaluation of emotional-support chatbots. An auditor model role-plays help-seeking users, generating multi-turn conversations from 140 seed instructions and...

    arxiv.org/abs/2606.30256 · PDF

  17. 17

    Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors

    Maxime Riché, Daniel Tan, Vili Kohonen, Niels Warncke

    cs.AI

    Inoculation prompting is a selective generalization technique used against Emergent Misalignment. We introduce inoculation adapters (IA), which similarly diminish the optimization pressure to learn undesired traits by strengthening the trait at train time. Inoculation adapters are LoRAs that are trained and used over three steps: 1) trained on undesired traits; 2) attached frozen while a separate task adapter is trained on data exhibiting...

    arxiv.org/abs/2606.30252 · PDF

  18. 18

    Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration

    Zihan Guo, Zeyi Chen, Zhiyu Chen, Zicai Cui, Shuai Shao, Bo Huang, Zhi Han, Yuanyi Song, Yuan Yuan, Chenxi Zeng,...

    cs.AI · cs.CY · cs.MA

    Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant task or a closed workflow. Therefore, autonomous science needs a collaboration infrastructure that coordinates projects, agents, and digital and physical resources. We identify this as a shift from code-centered execution loops to research-oriented collaboration processes, where questions,...

    arxiv.org/abs/2606.30246 · PDF

  19. 19

    EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

    Buğra Alperen Uluırmak, Rifat Kurban

    cs.AI · cs.CL · cs.LG · cs.SE

    LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight...

    arxiv.org/abs/2606.30219 · PDF

  20. 20

    The Many-Body Problem of the Data Centre

    Marcin Korecki, Cesare Carissimo

    cs.AI · cs.CY

    Modern Artificial Intelligence is often framed as limited by its own disembodiment, as if giving it a body would unlock its true potential. We argue to the contrary that it is the Data Centre that is, in many cases, the body of the AI. At the same time, the Data Centre is part of the labouring body of Capital and possesses staggering organismic qualities when seen through a biological lens. We elucidate the organic analogy and identify the...

    arxiv.org/abs/2606.30206 · PDF

  21. 21

    Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data

    Hyunwoo Park, Sang-Hyun Lee

    cs.AI

    Sim-to-real transfer remains a major obstacle for reinforcement learning (RL), especially for vision-based control where image observations exacerbate the state-distribution shift between simulation and the real world. Domain adaptation (DA) is a promising remedy for this challenge. Prior sim-to-real DA works have demonstrated encouraging results, yet these approaches typically assume substantially more target data, which is not available in...

    arxiv.org/abs/2606.30192 · PDF

  22. 22

    From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent

    Haoliang Han

    cs.AI · cs.LG · cs.NE

    How does an agent that can tell self from world come to be durably shaped by that distinction? Recent work shows that a predictive system can detect its own agency (Ye, 2026), but detecting agency does not explain durable, self-shaped behavior. We show that agency-gated slow credit -- a conjunctive term Own*Agency*Salience driving a slow parameter update -- produces post-unload behavioral residue: on a spiking substrate (Nengo LIF/PES), a...

    arxiv.org/abs/2606.30191 · PDF

  23. 23

    Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

    Yutao Sun, Yanting Miao, Hao-Xuan Ma, Mengyu Zhou, Mingshuai Chen, Tiancheng Zhao, Dexin Wang, Lei Lv, Li Xu, Xiaoxi...

    cs.AI

    Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adapts a frozen VLM without any weight updates. On a small labeled training subset, the agent inspects its own correct and incorrect attempts and evolves two complementary capabilities: reusable reasoning skills for cognitive bottlenecks, and executable visual tools for...

    arxiv.org/abs/2606.30185 · PDF

  24. 24

    MirrorCode: AI can rebuild entire programs from behavior alone

    Tom Adamczewski, David Owen, David Rein, Florian Brand, Giles Edkins, Allen Hart, Daniel O'Connell

    cs.AI

    AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C compiler. However, existing coding benchmarks tend to focus on shorter tasks, and one-off demonstrations are hard to compare systematically because they often have some human guidance, and are not standardized or repeated across models. To address these challenges, we introduce MirrorCode, a long-horizon...

    arxiv.org/abs/2606.30182 · PDF

  25. 25

    FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars

    Habin Lim, Jae-Ho Lee, Hah Min Lew, Ji-Su Kang, Gyeong-Moon Park

    cs.AI · cs.CV · cs.LG

    Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion. Existing systems only partially address this problem: speech-only full-duplex models can generate speech in real time but do not produce facial motion, while audio-driven facial motion models animate a face from already available audio rather than jointly generating speech and motion online. To bridge this gap, we first formalize...

    arxiv.org/abs/2606.30145 · PDF

  26. 26

    Relevance Is Not Permission: Warranted Attention for Value Contributions

    Minwoo Yu, Young-guk Ha

    cs.AI

    Relevance is not permission. Attention lets a model read key-value items related to the current query, but it does not guarantee that the value contribution of such an item becomes prediction evidence. A retrieved passage may be relevant to a question without being supporting evidence, and a historical fact or temporal neighbor may even blur true-tail ranking or the current edge score. This paper formalizes this gap as a permission problem...

    arxiv.org/abs/2606.30139 · PDF

  27. 27

    Does Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, Matters

    Wenlong Wang, Fergal Reid

    cs.AI · cs.CL

    Chain-of-thought (CoT) prompting improves LLM reasoning, but the source is contested: do the intermediate steps help because they carry useful semantic content, or because conditioning on more tokens buys extra computation before the model commits to an answer? We bring two lines of evidence to bear. First, in distribution: we repeatedly sample each model on the same question and pair a shorter with a longer of its own natural generations...

    arxiv.org/abs/2606.30128 · PDF

  28. 28

    Open Problems in Constitutional Preference Reconstruction

    Eleanor Clifford, Michael Amir, Arduin Findeis, Aaron Zhao, Robert Mullins

    cs.AI

    Pairwise preference data is widely used for training and evaluating language models (e.g., RLHF), but each datapoint records a \emph{choice}, not the rationale behind it. Methods such as Inverse Constitutional AI (ICAI) attempt to improve interpretability by compressing datasets into short ``constitutions'' of natural-language principles. We argue this framing is under-specified: a flat list of principles is not yet an executable decision...

    arxiv.org/abs/2606.30116 · PDF

  29. 29

    Structural Certification for Reliable Physical Design with Language Models

    Nakul Vyas, Iliya D. Stoev

    cs.AI · cs.LG

    An unreliable language model can be made to produce reliable physical designs if the authority to assert is moved out of the model: the model proposes, and a deterministic engine alone certifies, returning certified, impossible, or unknown. We introduce Physics-Anchored Certification (PHACT), a propose-certify loop spanning five scientific domains, and identify what makes such a certificate trustworthy. A checker that accepts a model-supplied...

    arxiv.org/abs/2606.30107 · PDF

  30. 30

    Propagation of~Interval Belief Structures and~Imprecise Copulas for~Neural Network Verification

    Francesc Pifarre-Esquerda, Eric Goubault, Sylvie Putot

    cs.AI · cs.LO · math.PR

    Quantitative verification of neural networks requires reasoning about probabilities under substantial uncertainty in both input distributions and their dependence structure. In realistic settings, this information is often only partially specified, and assuming precise probabilistic models can lead to unreliable results. We propose a sound framework for quantitative verification under imprecise probabilistic information, combining interval...

    arxiv.org/abs/2606.30105 · PDF

  31. 31

    Temporal Feature Extractors in EEG Foundation Models: A Controlled Comparison Including a Pretrained Time-Series Model

    Ayşe Betül Yüce, Chris Joey Leffler, Sarun Varghese, Myra Spiliopoulou, Sebastian Stober

    cs.AI

    Electroencephalography (EEG) foundation models aim to learn generalizable representations from large-scale brain recordings. However, the role of temporal feature extractors and whether pretrained time-series foundation models (TSFMs) can be effectively transferred to this setting remains underexplored. We conduct a controlled comparison of three temporal feature extraction strategies, including a linear baseline, a convolutional encoder, and...

    arxiv.org/abs/2606.30104 · PDF

  32. 32

    Hierarchical Reinforcement Learning in StarCraft Micromanagement with Influence Maps and Cluster-based Scripts

    Chunhui Bai, Changhe Li, Dequan Li, Xinye Cai, Shengxiang Yang

    cs.AI

    Real-time strategy (RTS) games present significant AI challenges, characterized by expansive state-action spaces arising from multi-unit coordination in continuous battlefields, and sparse delayed rewards stemming from final win/lose signals. Existing approaches face a trade-off between managing the dimensionality explosion of joint actions and maintaining the interpretability of complex state representations. This complexity is further...

    arxiv.org/abs/2606.30092 · PDF

  33. 33

    SAT-RTS: A systematic framework for tactical knowledge extraction and visualization-based analysis in real-time strategy games

    Chunhui Bai, Changhe Li, Yuqiang Li, Lei Liu, Shoufei Han

    cs.AI

    Efficient tactical knowledge extraction and analysis in real-time strategy (RTS) games micromanagement are constrained by the high-dimensional coupled state-action sequential data and the black-box decision-making process. Current research rarely provides a hierarchical visualization-based attribution analysis from the perspective of data decoupling and abstraction. To facilitate interpretable tactical knowledge extraction and...

    arxiv.org/abs/2606.30090 · PDF

  34. 34

    ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning

    Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara, Scott Sanner, Pascal Poupart, Jongmin Lee, Kee-Eung Kim

    cs.AI

    Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centralized Training with Decentralized Execution (CTDE) paradigm, policy gradients have remained difficult to compute directly. Prior methods largely follow two approaches: independent factorized updates with centralized critics, which lack general joint-improvement guarantees without value decomposition...

    arxiv.org/abs/2606.30072 · PDF

  35. 35

    AlgoSkill: Learning to Design Algorithms by Scheduling Human-Like Skills

    Xinyuan Song, Zekun Cai, Liang Zhao

    cs.AI

    Designing an algorithm from a natural-language problem statement requires identifying the problem structure, reading constraints, choosing a suitable paradigm, checking correctness, and refining complexity. Existing large language model (LLM) methods often rely on direct generation or generic self-refinement, leaving these steps implicit. We propose AlgoSkill, which models algorithm design as sequential decision-making over a typed library of...

    arxiv.org/abs/2606.29999 · PDF

  36. 36

    Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning

    Peng, Lee, Yin Zhang, Yanglin Zhang, Haonan Wu, Zishan Liu, Ruoxi Zang, Xin Zhu, Jiayin Zheng, Jian Yao, Zefeng Ji, Fei Ma

    cs.AI

    Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs). However, directly applying RL to rollout multimodal reasoning can lead to instability, due to the exploitation of language priors, the neglect of visual evidence, and the generation of reasoning traces that are fluent yet not visually grounded. The question arises: Can initially steer the policy toward visually...

    arxiv.org/abs/2606.29984 · PDF

  37. 37

    Exploration and Online Transfer with Behavioral Foundation Models

    Louis Bagot, Mathieu Lefort, Laëtitia Matignon

    cs.AI · cs.LG

    Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, without additional learning at transfer time, while training only on reward-free trajectories. For their generality over tasks, such models are sometimes called ``Behavioral Foundation Models'' (BFMs). While they have shown strong performances and improvements in recent years, the current framework and...

    arxiv.org/abs/2606.29980 · PDF

  38. 38

    First-Order Temporal Logic Tensor Networks

    Luca Boscarato, Ivan Donadello, Alessandro Artale, Marco Montali, Fabrizio Maria Maggi

    cs.AI · cs.LG · cs.LO

    Most of the existing neuro-symbolic AI methods focus on the scenario of static knowledge where objects do not change according to a temporal dimension. Temporal neuro-symbolic works are still under explored and are mainly developed for time-interval logic or propositional linear temporal logic. There is a lack of models studying linear temporal logics with predicates that deal with objects whose properties and relations change through the...

    arxiv.org/abs/2606.29972 · PDF

  39. 39

    SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon CivRealm Strategy Planning

    Tianyu Jin, Shuo Chen, Yida Wang, Liuyu Xiang, Yingzhuo Liu, Zhiyao Jiang, Yexin Li, Zhaofeng He

    cs.AI

    Long-horizon strategic planning in complex strategy games demands concurrent reasoning across multiple decision domains under imperfect information and sparse reward. Existing LLM-based agents suffer from three systematic failures: scene blindness from raw tile coordinates, context overflow and domain coupling from monolithic state dumps, and shallow cross-game learning that treats each episode in isolation. We present SAGA, an LLM...

    arxiv.org/abs/2606.29932 · PDF

  40. 40

    HippoSpark: An On-Demand Experience System for LLM Reasoning

    Jingyao Liu, Danling Meng, Chen Huang, Yukun Yan, Zhenghao Liu, Wenqiang Lei, See-Kiong Ng, Maosong Sun

    cs.AI

    Distilling historical trajectories into reusable experience to enhance future problem-solving has become a focal point of recent LLM research. However, existing methods predominantly operate at the task level, leveraging general summaries or rules under the assumption that analogous tasks share universal solution patterns. This approach often fails in complex reasoning, which typically falters at local bottlenecks that require precise,...

    arxiv.org/abs/2606.29929 · PDF

  41. 41

    A causal modeling perspective on decision theory

    Arvid Sjölander

    cs.AI · stat.ME

    Decision theory provides a formal framework for how agents should make choices under uncertainty, drawing on ideas from philosophy, probability, and causality. Despite significant progress, the field still lacks a unified modeling language, and key concepts - such as the distinction between subjective and objective elements, or what it means for a decision theory to perform well - are often left implicit. This can make it difficult to...

    arxiv.org/abs/2606.29911 · PDF

This edition is part of The Daily Abstract — cs.AI archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.