stat.ML · 2026-07-23 · No. 62

Machine Learning, 2026-07-23.

10 new papers in stat.ML. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

10 entries
  1. 01

    Adaptive deep nonparametric regression from dependent data under covariate shift

    William Kengne, Ehud Mossa Ockegna

    stat.ML · cs.LG

    Covariate shift often occurs because, in many real applications, the source and the target observations may be generated from different distributions. In this case, the standard metric under the source distribution is not appropriate. This paper considers deep neural network estimators for nonparametric quantile and Huber regression under covariate shift and from dependent observations. We deal with a generalized Bernstein-type inequality...

    arxiv.org/abs/2607.20309 · PDF

  2. 02

    Adaptive Bayesian Online Learning via Expert Aggregation

    Jungbin Jun, Ilsang Ohn

    stat.ML · cs.LG

    Bayesian online learning promises uncertainty-aware prediction on data streams, but its performance hinges on inferential choices, including learning rates, prior distributions and variational families, which are usually fixed before seeing the stream. We address this by treating Bayesian update rules as experts and aggregating the Bayesian experts according to sequential predictive losses. We prove that the resulting aggregate competes with...

    arxiv.org/abs/2607.20239 · PDF

  3. 03

    Statistical Inference for Rank Allocation in Low-Rank Adaptation

    Yihang Gao, Vincent Y. F. Tan

    stat.ML · cs.LG · math.ST

    Low-rank adaptation (LoRA) has become a widely used parameter-efficient fine-tuning method for large language models. Since different modules and layers may contribute unequally to downstream adaptation, allocating rank resources under a fixed parameter budget is an important problem for balancing efficiency, expressiveness, and generalization. Existing adaptive rank methods address this problem mainly through carefully designed importance...

    arxiv.org/abs/2607.20205 · PDF

  4. 04

    Directional Kernel Mean Difference: A Fast Signed Statistic for Univariate Distribution Comparison

    Shijie Zhong, Jiangfeng Fu

    stat.ML · cs.LG

    We introduce the Directional Kernel Mean Difference (DKMD), a signed statistic for univariate distribution comparison that preserves the direction of distributional shifts. Unlike the squared Maximum Mean Discrepancy (MMD), which discards directional information by squaring the RKHS distance, DKMD integrates the difference of kernel mean embeddings against a fixed odd weighting function. This construction yields three structural properties:...

    arxiv.org/abs/2607.20119 · PDF

  5. 05

    Non--negative matrix factorization using the \textit{R} package \textsf{nnmf}

    Volkan Sevinç, Nikolas Kontemeniotis, Theodoros Perdikis, Michail Tsagris

    stat.ML · cs.LG

    Non--negative matrix factorization (NMF) has become an established dimensionality reduction technique for extracting latent structures from non--negative data and has found widespread applications in fields such as bioinformatics, text mining, image analysis, and recommender systems. As the popularity of NMF has increased, numerous \textit{R} packages implementing different optimization strategies and computational frameworks have been...

    arxiv.org/abs/2607.20084 · PDF

  6. 06

    Data-Poisoning Audits for Causal Effect Estimation

    Kwangho Kim

    stat.ML · cs.LG · stat.ME

    Observational causal analyses increasingly pool records across sites, vendors, and collection systems, creating vulnerability to append-only attacks in which plausible records are strategically selected to alter a reported treatment effect. We develop a data-poisoning audit for augmented inverse-probability-weighted estimation. The analyst specifies a finite catalog of feasible records, an append budget, and nested source capacities, and the...

    arxiv.org/abs/2607.19692 · PDF

  7. 07

    Optimal Recalibration of an Online Predictor

    Lunjia Hu, Kevin Tian, Chutong Yang

    stat.ML · cs.DS · cs.LG

    We study the problem of recalibrating an online predictor [KE17, OKS24]: given an arbitrary "hint" sequence of forecasts, the learner must output new predictions that are calibrated while incurring small excess error relative to the original forecasts, under a proper loss. We give an online algorithm that achieves $(\varepsilon, \varepsilon^2)$-recalibration for Lipschitz proper losses in $T \approx \varepsilon^{-3}$ rounds, using an...

    arxiv.org/abs/2607.19689 · PDF

  8. 08

    RELTA-SGLD: Relative-Growth Localized Taming for Nonconvex Stochastic-Gradient Langevin Learning

    Yiwei Zhou, Ziheng Chen

    stat.ML · cs.LG · math.OC

    We introduce RELTA-SGLD, a taming scheme that stabilizes superlinear stochastic-gradient updates while reducing unnecessary suppression of the original learning drift. A threshold determines where the taming turns on, while a relative-growth principle derived from the one-step Lyapunov stability condition determines the required taming strength. Together, they produce a lighter $λ$-scale denominator and preserve a nonvanishing far-tail...

    arxiv.org/abs/2607.19544 · PDF

  9. 09

    Boltzmann-Expected Molecular Design with Decoupled Annealing Flows

    Selma Moqvist, Richard Beckmann, Ross Irwin, Rocío Mercado, Simon Olsson

    stat.ML · cs.LG

    Most 3D properties relevant to molecular design, including free energies and shape descriptors, are $\textit{expectations}$ over the Boltzmann distribution over 3D configurations of a molecular graph. However, existing property-guided generative models tie each property to a single structure, ignoring the underlying ensemble. We recast 3D molecular design as $\textbf{Boltzmann-expected design}$ and realise it with $\textbf{DECAF}$ (Decoupled...

    arxiv.org/abs/2607.19519 · PDF

  10. 10

    A Bayesian Framework for Built-in Input Dimension Reduction for Gaussian Process Modeling

    Eric Herrison Gyamfi, Emily L. Kang, Bledar A. Konomi, Guang Lin

    stat.ML · cs.LG · math.PR · stat.AP · stat.CO · stat.ME

    Gaussian process (GP) modeling is widely used in computational science and engineering. However, fitting a GP to high-dimensional inputs remains challenging due to the curse of dimensionality. While various methods have been proposed to reduce input dimensionality, they typically follow a two-stage approach, performing dimension reduction and GP fitting separately. We introduce a Bayesian framework that seamlessly integrates dimensionality...

    arxiv.org/abs/2607.19498 · PDF

This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.