stat.ML · 2026-08-16 · No. 86

Machine Learning, 2026-08-16.

6 new papers in stat.ML. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

6 entries
  1. 01

    Bagging Robustly Learns VC Classes with Linear Sample Complexity

    Omar Montasser

    stat.ML · cs.DS · cs.LG

    We revisit the problem of learning predictors robust to adversarial examples at test-time. We prove that VC classes are adversarially robustly learnable with sample complexity linear in the VC dimension $d$, providing an exponential improvement over the previous upper bound of Montasser, Hanneke, and Srebro (2019). Remarkably, this result is achieved with a simple improper algorithm that combines the classic heuristic bagging (bootstrap...

    arxiv.org/abs/2608.13514 · PDF

  2. 02

    Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

    Yikai Xu, Zhao Chen, Jian Huang

    stat.ML · cs.LG

    Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data. The core insight is to select a subset of samples whose empirical distribution maximizes...

    arxiv.org/abs/2608.13418 · PDF

  3. 03

    Sinkhorn Linearization and the Spectral Proxy: Unifying the Statistical and Algorithmic Theory of Feature-Parameterized Inverse Optimal Transport via a Single Spectral Sandwich

    Han Dong, Jiaming Li, Yongqiang Gong, Ruixi Li, Yin Liu

    stat.ML · cs.LG · math.OC · math.ST

    We develop the statistical and algorithmic theory of inverse optimal transport (IOT) under the feature-parameterized cost C_theta(i,j) = -theta^T phi(i,j). The core technical contribution is the Sinkhorn linearization -- the implicit-function sensitivity of the entropic OT plan to the cost -- together with its spectral proxy, a formula that is spectrally exact yet geometrically transparent. The restricted Hessian on the tangent space...

    arxiv.org/abs/2608.13201 · PDF

  4. 04

    High-dimensional networks and mean squared error for possibly misspecified models

    Lourens Waldorp

    stat.ML · cs.LG

    To avoid missing important variables and their connections in networks, more and more variables are included in network analysis. Here we show that in a setting with many more parameters than observations (high-dimensional) it is possible to get a conservative (i.e., low false positive rate) estimate of the neighbourhood for each node (which connections are in the network). A neighbourhood is often estimated with a linear model, and this...

    arxiv.org/abs/2608.13171 · PDF

  5. 05

    Statistical Properties of Robust Learning under Distributional Shifts

    Zhiyi Li, Xiaojie Mao, Yunbei Xu, Ruohan Zhan

    stat.ML · cs.LG

    Distributional shifts arise when the target deployment environment differs from the source environment that generated the training data. Robust learning frameworks such as Distributionally Robust Optimization (DRO) and Robust Satisficing (RS) aim to address this challenge, yet their finite-sample guarantees under such shifts, and their systematic comparison, remain underexplored: existing analyses typically establish guarantees either in the...

    arxiv.org/abs/2608.13133 · PDF

  6. 06

    Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

    Zijie Cheng, Yang Peng, Zhihua Zhang

    stat.ML · cs.LG

    In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming access to a generative model, we first establish functional central limit theorems for both synchronous and asynchronous QTD, which show that the averaged iterates of QTD converge weakly to a rescaled Brownian motion. We next provide online inference methods. Based on random scaling,...

    arxiv.org/abs/2608.12973 · PDF

This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.