stat.ML · 2026-06-30 · No. 39
Machine Learning, 2026-06-30.
9 new papers in stat.ML. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
9 entries-
01
Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms
Ziwei Su, Junyu Ren, Victor Veitch
stat.ML · cs.AI · cs.LG · math.OC
Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosine similarity, effectively ignoring embedding magnitudes. However, surprisingly, empirical studies reveal that despite this, these "discarded" norms seem to correlate with semantic properties such as concept specificity, token frequency, and human uncertainty. In this work, we provide a formal theoretical framework explaining...
-
02
Doubly Robust Adaptive Conformal Inference for Causal Effects Under Temporal Dependence
Andreas Koukorinis, Ricardo Silva
stat.ML · cs.LG · stat.CO
We propose doubly robust adaptive conformal inference (DR-ACI), which constructs prediction intervals for doubly robust pseudo-outcomes under temporal dependence.
-
03
Factorizable Normalizing Flows for parameter-dependent density morphing
Davide Valsecchi, Mauro Donegà, Rainer Wallny
stat.ML · cs.LG · hep-ex · hep-th · physics.data-an
Normalizing Flows excel at modeling a single fixed density, yet many problems across the sciences, such as high energy physics, instead require modeling how that density deforms as a function of continuous parameters: the strength of a physical effect, a calibration constant, or a source of systematic uncertainty. Learning a separate flow for every parameter configuration quickly becomes intractable, since the number of joint settings grows...
-
04
Non-parametric recovery of causal diffusion mechanisms from steady-state observations
Richard Schwank, Mathias Drton
stat.ML · cs.LG
We consider sparse multivariate stochastic systems that evolve in continuous time according to a causal mechanism and present methodology to recover the system's time-infinitesimal transition mechanism from mere cross-sectional data. This observational paradigm is motivated by applications such as gene expression analysis, where destructive experimental techniques may only allow recording data once over a cell's lifetime. Precisely, we assume...
-
05
SGD Provably Prioritizes a Shortcut Spurious Feature in the XOR Model
Tyler LaBonte, Vidya Muthukumar
stat.ML · cs.LG
Neural networks are known to be susceptible to over-reliance on spurious correlations. However, the precise mechanism by which models exploit shortcut features is not fully understood, and algorithms to mitigate this behavior rely on as yet unjustified assumptions about the learned representations. In this work, we provide the first end-to-end theoretical characterization of spurious feature learning for two-layer ReLU neural networks trained...
-
06
A Stochastic--Geometric Theory of Scaling Laws in Grokking
Róisín Luo, Christian Gagné, Jonas Ngnawé, Ihsan Ullah, Karyn Morrissey
stat.ML · cs.AI · cs.LG
Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged delay, often through an abrupt transition. Despite extensive empirical study, its underlying mechanism remains poorly understood. In this work, we first theoretically characterize a shell--core topological configuration of the reachable solution space induced by...
-
07
Extrapolating from Regularised Solutions for Solving Ill-Conditioned Linear Systems in Machine Learning
Disha Hegde, Jon Cockayne, Chris. J. Oates
stat.ML · cs.LG · math.NA
Rapid prototyping of algorithms is a critical step in modern machine learning. Most algorithms exploit linear algebra, creating a need for lightweight numerical routines which -- while potentially sub-optimal for the task at hand -- can be rapidly implemented. For the numerical solution of ill-conditioned linear systems of equations, the standard solution for prototyping is Tikhonov-regularised inversion using a nugget. However, selection of...
-
08
Highly Data Parallelizable Estimation of the Sliced-Wasserstein Distance Using Cumulative Distribution Functions
Christophe Vauthier, Quentin Mérigot, Anna Korba
stat.ML · cs.LG
The Sliced Wasserstein (SW) distance has emerged as a computationally attractive alternative to the Wasserstein distance by leveraging one-dimensional optimal transport along random projections. Standard estimators of the SW distance rely on Monte Carlo averages of one-dimensional Wasserstein distances computed via quantile functions, which require sorting projected samples and access to full datasets. In this work, we introduce a new class...
-
09
Notes on generative modeling: flow matching, diffusion, optimal transport and Schr{ö}dinger bridge
Titouan Vayer
stat.ML · cs.LG
These notes recapitulate the high level mathematical principles behind different techniques for generative modeling. I show the connections between optimal transport and standard techniques such as Schr{ö}dinger bridge and flow matching.
This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.