stat.ML · 2026-06-02 · No. 13
Machine Learning, 2026-06-02.
9 new papers in stat.ML. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
9 entries-
01
Doing well with less! On Sampling Techniques for Empirical Pairwise Loss Estimation/Minimization
Louise Davy, Stephan Clémençon, Charlotte Laclau
stat.ML · cs.LG
Many machine learning problems, including similarity learning, ranking, and clustering, rely on empirical pairwise loss functions whose quadratic computational cost quickly becomes prohibitive at scale. We demonstrate how a frugal approach that retains only a fraction of the available information on pairs can achieve estimation or optimization performance comparable to that obtained by using all pairs, by leveraging survey sampling...
-
02
ShaplEIG: Bayesian Experimental Design for Shapley Value Estimation
David Rundel, Fabian Fumagalli, Maximilian Muschalik, Bernd Bischl, Matthias Feurer
stat.ML · cs.LG
Shapley values are a principled attribution measure widely used in interpretable machine learning, but their exact computation scales exponentially with the number of players, motivating a wide range of approximation methods based on value function evaluations of sampled coalitions. This raises the question of whether approximation accuracy can be improved by adaptively selecting coalitions for evaluation based on previous evaluations. This...
-
03
Identifiable Markov Switching Models with Instantaneous Effects and Exponential Families
Roel Hulsman, Carles Balsells-Rodas, Sara Magliacane
stat.ML · cs.LG · stat.ME
Temporal systems often exhibit non-stationary behaviour, such as seasonal climate variation or glucose fluctuations in patients with type-1 diabetes. One way to model non-stationarity is through discrete latent regimes, i.e., stationary segments of time. Such systems induce a Markov Switching Model (MSM), a class of Hidden Markov Models with autoregressive dependencies among latent regimes and observed variables. Identifying latent regimes is...
-
04
Bayesian meta-learning for modeling Alzheimer's disease progression
Clara Hoffmann, Nadja Klein
stat.ML · cs.CV · cs.LG
Predicting whether an individual with Alzheimer's disease will experience mild or severe disease progression is essential for personalized treatment. Typically, practitioners seek to predict the distribution of a discrete disease score, conditional on an individual's current MRI volume and their historical disease trajectory. Classical statistical regression models and single-task neural networks are not well-suited for this purpose because...
-
05
ProbRes: Volatility Learning for Probabilistic Time-Series Forecasting
Tingting Wang, Yunyi Zhang, Benyou Wang
stat.ML · cs.LG · stat.ME
Probabilistic time series forecasting has attracted increasing attention in financial applications due to the need to quantify risk and uncertainty in future observations. We propose ProbRes, a post-hoc probabilistic calibration method that explicitly learns and incorporates volatility dynamics into probabilistic forecasting, enabling effective handling of heteroskedastic data. During training, ProbRes employs two architecture-agnostic...
-
06
Error Bounds for a Diffusion Model-Based Drift Estimator
Ioar Casado-Telletxea, Omar Rivasplata
stat.ML · cs.LG
Parameter estimation in stochastic differential equations is a classical statistical problem of much importance in many scientific fields. Recent work of Tapia Costa et al. (2026) introduced a novel technique for estimating the drift when the diffusion parameter is known, using discrete samples from multiple trajectories. Their method treats drift estimation as a denoising problem, and leverages tools from (conditional) score-matching...
-
07
It does what it says on the tin: safe synthetic data from coarsened margins
Gillian M Raab
stat.ML · cs.LG · stat.AP
This paper proposes a method of creating synthetic data (SD) that will have two important advantages for the user compared to other methods currently available. The first is transparency; unlike other methods, the person in receipt of the SD will know which of the relationships between variables in the original data will be approximately maintained in the SD. The second is a guarantee that the SD is derived from information that has already...
-
08
Convex Distance Operator Transport: A Convex and Geometry-Preserving Formulation
Junhyoung Chung, Euijong Song, Won Hwa Kim, Gunwoong Park
stat.ML · cs.LG · math.ST · stat.ME
We introduce Convex Distance Operator Transport (CDOT), the first convex optimal transport framework that aligns distributions across heterogeneous domains by jointly preserving feature correspondence and intrinsic geometric structure. Specifically, CDOT employs an operator-based regularization that aligns aggregated distance structures by introducing distance and conditional expectation operators. Consequently, the proposed regularization...
-
09
Provable Data Scaling Law for Meta Learning via Complexity Minimization
Kazuto Fukuchi, Ryuichiro Hataya, Kota Matsui
stat.ML · cs.LG
Pre-training has become a fundamental paradigm in modern machine learning, with one of its key empirical benefits being reduced downstream sample complexity as the scale of pre-training data increases. However, existing theoretical frameworks for pre-training do not fully explain this phenomenon. In this paper, we introduce complexity minimization, a novel meta-representation learning framework designed to enable theoretical analysis of this...
This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.