stat.ML · 2026-06-06 · No. 15
Machine Learning, 2026-06-06.
11 new papers in stat.ML. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
11 entries-
01
Conformal Risk Sharing: Certified Cost Allocation with Participation Guarantees
Ieva Kazlauskaite
stat.ML · cs.LG
Sharing the financial impact of rare adverse events across a group can soften extreme individual burdens, but any participant made worse off by the arrangement has reason to leave. A credible mechanism must therefore provide each agent with a trustworthy cap on their future obligation and should be deployed only if the aggregate harm across participants is bounded. We formalise this as the Certified Allocation Problem: from finite data and...
-
02
Function-Space Priors for Bayesian Neural ODEs with Application to Vessel Trajectory Prediction
Jaeyeong Lee, Wonmo Koo, Heeyoung Kim
stat.ML · cs.LG
Vessel trajectory prediction from Automatic Identification System (AIS) data is essential for maritime situational awareness, yet it remains challenging due to irregular sampling, missing reports, and complex dynamics. Beyond accurate point forecasts, maritime applications also demand well-calibrated uncertainty estimates for reliable decision-making. Bayesian Neural Ordinary Differential Equations (ODEs) offer a principled framework for...
-
03
Symmetric Divergence and Normalized Similarity: A Unified Topological Framework for Representation Analysis
Yan Wang, Tianyang Hu
stat.ML · cs.LG
Topological Data Analysis (TDA) offers a principled, intrinsic lens for comparing neural representations. However, existing paired topological divergences (e.g., RTD) are limited by heuristic asymmetry and, more critically, unbounded scores that depend on sample size, hindering reliable cross-scenario benchmarking. To address these challenges, we develop a unified topological toolkit serving two complementary needs: fine-grained structural...
-
04
Discrete Causal Representations from Heterogeneous Domains: A Bayesian Approach with Social Survey Applications
Ankur Garg, Michael Stettler, Aaron Schein, Julius von Kügelgen
stat.ML · cs.LG
Causal representation learning aims to infer the high-level latent causal concepts that give rise to observed low-level measurements. This is particularly relevant for heterogeneous data from different environments or domains since distribution shifts often arise through sparse, localized changes in some of the underlying causal mechanisms, while other parts of the generative process remain unchanged. Whereas identifiability of causal...
-
05
Anchor PCA
Benedikt Seiter, Anya Fries, Julius von Kügelgen, Jonas Peters
stat.ML · cs.LG · stat.ME
Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques. We study PCA for data from multiple related domains. Since principal components generally differ across domains, one way to obtain a shared low-rank embedding is to perform PCA on the pooled data. However, this approach can focus on spurious directions that exhibit high variation in only a few domains. To find a robust embedding that...
-
06
Diffusion Models Observe Only Gradients: A Geometric Perspective on Score Matching Errors
Naïl B. Khelifa, Richard E. Turner, Ramji Venkataramanan
stat.ML · cs.LG
Score-based diffusion models are typically trained by minimizing the $L^2$ score matching error, and standard theoretical analyses rely on this quantity to bound the sampling discrepancy between the learned and target distributions. We show the $L^2$ score error is not the right intrinsic measure of marginal distributional quality: a learned diffusion model can incur arbitrarily large $L^2$ score error while perfectly matching the target...
-
07
Effective Dimensionality as an Operator Invariant for Physics-Preserving Constraint Adaptation in Physics-Informed Neural Networks
Cornelius Otchere, Michael Shields
stat.ML · cs.LG · math.NA · physics.comp-ph
Physics-Informed Neural Networks inherently suffer from task interference because they rely on a shared parameter space to satisfy both governing differential equations and boundary conditions. We analyze this structural conflict using the Fisher Information Matrix to quantify the effective degrees of freedom ($d_{eff}$) in a physics-constrained model. Unlike the classical $d_{eff}$ which measures how many parameter directions are informed by...
-
08
Adaptive Learning Rates with Surrogate Probability for Follow-the-Perturbed-Leader
Jongyeong Lee, Junya Honda, Shinji Ito, Chansoo Kim
stat.ML · cs.LG
Follow-the-regularized-leader framework has shown effectiveness and flexibility in online learning problems, where the choice of learning rates are known to be crucial. Recently, adaptive learning rates defined in terms of the arm-selection probabilities, obtained by solving convex optimization, have achieved improved best-of-both-worlds (BOBW) guarantees in various bandit problems. In contrast, BOBW guarantees for its computationally...
-
09
Fast and Robust Convergence Rate for TD(0) with Linear Function Approximation, Universal Learning Steps and I.I.D. Samples
Ziad Kobeissi, Éloïse Berthier
stat.ML · cs.LG
In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA). We consider on-policy independent and identically distributed (i.i.d.) samples, a constant learning step, and the Polyak-Juditsky averaging method. We establish a new convergence rate, for the Mean-Square Error (MSE) on the approximated function, that is (i) fast in the sense that it admits an optimal dependency...
-
10
EML-CD: Causal Mechanism Recovery via EML Symbolic Trees in Structure Learning
Sota Asanuma
stat.ML · cs.LG
Neural network (NN)-based nonlinear causal discovery methods recover DAG structure but leave each causal mechanism as a black box. Waxman et al. argued that extracting causal mechanisms from NN weights is ill-posed. We propose EML-CD, a framework that integrates the EML operator (capable of composing elementary functions from a single binary operator) into causal structure learning, with interpretable mechanism recovery as the primary...
-
11
Finding Most Influential Sets
Lucas D. Konrad, Nikolas Kuschnig
stat.ML · cs.LG · econ.EM · stat.CO
Identifying most influential sets (MIS) - size-$k$ subsets whose removal maximally changes a target estimand - is typically infeasible because it requires searching over $\binom{n}{k}$ subsets. For estimands with linear-fractional leave-set-out effects, we show that MIS selection reduces to a one-parameter sequence of top-$k$ problems. Dinkelbach's method yields an algorithm with $\mathcal{O}(n)$ cost per iteration and finite termination. For...
This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.