stat.ML · 2026-06-06 · No. 15

Machine Learning, 2026-06-06.

11 new papers in stat.ML. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

11 entries
  1. 01

    Conformal Risk Sharing: Certified Cost Allocation with Participation Guarantees

    Ieva Kazlauskaite

    stat.ML · cs.LG

    Sharing the financial impact of rare adverse events across a group can soften extreme individual burdens, but any participant made worse off by the arrangement has reason to leave. A credible mechanism must therefore provide each agent with a trustworthy cap on their future obligation and should be deployed only if the aggregate harm across participants is bounded. We formalise this as the Certified Allocation Problem: from finite data and...

    arxiv.org/abs/2606.06391 · PDF

  2. 02

    Function-Space Priors for Bayesian Neural ODEs with Application to Vessel Trajectory Prediction

    Jaeyeong Lee, Wonmo Koo, Heeyoung Kim

    stat.ML · cs.LG

    Vessel trajectory prediction from Automatic Identification System (AIS) data is essential for maritime situational awareness, yet it remains challenging due to irregular sampling, missing reports, and complex dynamics. Beyond accurate point forecasts, maritime applications also demand well-calibrated uncertainty estimates for reliable decision-making. Bayesian Neural Ordinary Differential Equations (ODEs) offer a principled framework for...

    arxiv.org/abs/2606.06351 · PDF

  3. 03

    Symmetric Divergence and Normalized Similarity: A Unified Topological Framework for Representation Analysis

    Yan Wang, Tianyang Hu

    stat.ML · cs.LG

    Topological Data Analysis (TDA) offers a principled, intrinsic lens for comparing neural representations. However, existing paired topological divergences (e.g., RTD) are limited by heuristic asymmetry and, more critically, unbounded scores that depend on sample size, hindering reliable cross-scenario benchmarking. To address these challenges, we develop a unified topological toolkit serving two complementary needs: fine-grained structural...

    arxiv.org/abs/2606.06342 · PDF

  4. 04

    Discrete Causal Representations from Heterogeneous Domains: A Bayesian Approach with Social Survey Applications

    Ankur Garg, Michael Stettler, Aaron Schein, Julius von Kügelgen

    stat.ML · cs.LG

    Causal representation learning aims to infer the high-level latent causal concepts that give rise to observed low-level measurements. This is particularly relevant for heterogeneous data from different environments or domains since distribution shifts often arise through sparse, localized changes in some of the underlying causal mechanisms, while other parts of the generative process remain unchanged. Whereas identifiability of causal...

    arxiv.org/abs/2606.06288 · PDF

  5. 05

    Anchor PCA

    Benedikt Seiter, Anya Fries, Julius von Kügelgen, Jonas Peters

    stat.ML · cs.LG · stat.ME

    Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques. We study PCA for data from multiple related domains. Since principal components generally differ across domains, one way to obtain a shared low-rank embedding is to perform PCA on the pooled data. However, this approach can focus on spurious directions that exhibit high variation in only a few domains. To find a robust embedding that...

    arxiv.org/abs/2606.06233 · PDF

  6. 06

    Diffusion Models Observe Only Gradients: A Geometric Perspective on Score Matching Errors

    Naïl B. Khelifa, Richard E. Turner, Ramji Venkataramanan

    stat.ML · cs.LG

    Score-based diffusion models are typically trained by minimizing the $L^2$ score matching error, and standard theoretical analyses rely on this quantity to bound the sampling discrepancy between the learned and target distributions. We show the $L^2$ score error is not the right intrinsic measure of marginal distributional quality: a learned diffusion model can incur arbitrarily large $L^2$ score error while perfectly matching the target...

    arxiv.org/abs/2606.06179 · PDF

  7. 07

    Effective Dimensionality as an Operator Invariant for Physics-Preserving Constraint Adaptation in Physics-Informed Neural Networks

    Cornelius Otchere, Michael Shields

    stat.ML · cs.LG · math.NA · physics.comp-ph

    Physics-Informed Neural Networks inherently suffer from task interference because they rely on a shared parameter space to satisfy both governing differential equations and boundary conditions. We analyze this structural conflict using the Fisher Information Matrix to quantify the effective degrees of freedom ($d_{eff}$) in a physics-constrained model. Unlike the classical $d_{eff}$ which measures how many parameter directions are informed by...

    arxiv.org/abs/2606.06171 · PDF

  8. 08

    Adaptive Learning Rates with Surrogate Probability for Follow-the-Perturbed-Leader

    Jongyeong Lee, Junya Honda, Shinji Ito, Chansoo Kim

    stat.ML · cs.LG

    Follow-the-regularized-leader framework has shown effectiveness and flexibility in online learning problems, where the choice of learning rates are known to be crucial. Recently, adaptive learning rates defined in terms of the arm-selection probabilities, obtained by solving convex optimization, have achieved improved best-of-both-worlds (BOBW) guarantees in various bandit problems. In contrast, BOBW guarantees for its computationally...

    arxiv.org/abs/2606.06043 · PDF

  9. 09

    Fast and Robust Convergence Rate for TD(0) with Linear Function Approximation, Universal Learning Steps and I.I.D. Samples

    Ziad Kobeissi, Éloïse Berthier

    stat.ML · cs.LG

    In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA). We consider on-policy independent and identically distributed (i.i.d.) samples, a constant learning step, and the Polyak-Juditsky averaging method. We establish a new convergence rate, for the Mean-Square Error (MSE) on the approximated function, that is (i) fast in the sense that it admits an optimal dependency...

    arxiv.org/abs/2606.05967 · PDF

  10. 10

    EML-CD: Causal Mechanism Recovery via EML Symbolic Trees in Structure Learning

    Sota Asanuma

    stat.ML · cs.LG

    Neural network (NN)-based nonlinear causal discovery methods recover DAG structure but leave each causal mechanism as a black box. Waxman et al. argued that extracting causal mechanisms from NN weights is ill-posed. We propose EML-CD, a framework that integrates the EML operator (capable of composing elementary functions from a single binary operator) into causal structure learning, with interpretable mechanism recovery as the primary...

    arxiv.org/abs/2606.05942 · PDF

  11. 11

    Finding Most Influential Sets

    Lucas D. Konrad, Nikolas Kuschnig

    stat.ML · cs.LG · econ.EM · stat.CO

    Identifying most influential sets (MIS) - size-$k$ subsets whose removal maximally changes a target estimand - is typically infeasible because it requires searching over $\binom{n}{k}$ subsets. For estimands with linear-fractional leave-set-out effects, we show that MIS selection reduces to a one-parameter sequence of top-$k$ problems. Dinkelbach's method yields an algorithm with $\mathcal{O}(n)$ cost per iteration and finite termination. For...

    arxiv.org/abs/2606.05919 · PDF

This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.