stat.ML · 2026-08-03 · No. 73

Machine Learning, 2026-08-03.

6 new papers in stat.ML. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

6 entries
  1. 01

    Analytical and Bootstrap Confidence Intervals of Double Machine Learning: Simulation studies and an application to rural-urban difference in obesity prevalence

    Haozheng Xu, Siyuan Ma, Qingyan Xiang

    stat.ML · cs.LG · stat.AP

    Double Machine Learning (DML) is a popular approach for treatment effect estimation in various settings, which allows a wide range of flexible machine learning methods to be used for nuisance parameter estimation while preserving valid inference. In practice, however, applied researchers must choose among many machine learning algorithms for nuisance models, and the impact of this choice on the variance estimation of DML is not well...

    arxiv.org/abs/2607.29456 · PDF

  2. 02

    The Greedy Advantage in Finite-Horizon Bandits

    Kai Zhou, Michael Lingzhi Li, Kai Wang

    stat.ML · cs.LG

    Organizations increasingly rely on sequential experimentation to improve decision-making. While the multi-armed bandit literature has developed algorithms with strong asymptotic regret guarantees, many practical applications operate over finite and externally imposed horizons. Motivated by the finite-horizon setting, we develop a class of regularized greedy algorithms for multi-armed Bernoulli bandits. We derive the first finite-horizon...

    arxiv.org/abs/2607.29375 · PDF

  3. 03

    Simple-regret rates and minimax optimality of fixed-prior expected improvement in Matérn and squared-exponential RKHSs

    Emmanuel Vazquez, Sébastien Petit

    stat.ML · cs.LG · math.NA · math.ST

    We study the expected improvement (EI) policy for minimizing a deterministic objective function $f$ on a nonempty compact set $\mathcal X \subset\mathbb R^d$. We assume that $f$ belongs to the RKHS $\mathcal H_k$ of a continuous positive-semidefinite kernel $k$ on $\mathcal X$. Function values are observed exactly, and EI is computed from a fixed zero-mean Gaussian-process model with covariance $σ^2k$. After an initial design, the policy...

    arxiv.org/abs/2607.29245 · PDF

  4. 04

    Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization

    Tyler Ashoff, Jordan Rodu

    stat.ML · cs.LG

    Modern opaque AI models prize performance over interpretability, which makes testing difficult. However, formal statistical tests conducted on a model's embedding space can provide robust characterizations of semantic structure, concept separation, and knowledge graph alignment. Model developers would benefit from a model comparison technique that leverages human-curated knowledge structures to test alignment. The scale of the input space for...

    arxiv.org/abs/2607.29008 · PDF

  5. 05

    Structured Neural Chaos: An Adaptive Surrogate Modeling Framework for Functional Uncertainty Quantification and Global Sensitivity Analysis

    Isabel Corona Guevara, Yeping Hu

    stat.ML · cs.LG

    Variance-based global sensitivity analysis (GSA) plays a key role in uncertainty quantification by identifying the contributions of uncertain inputs to the variability of the model response. The repeated model evaluations required for these tasks are often prohibitively expensive; surrogate models provide an efficient alternative by constructing inexpensive approximations of the underlying system response. Constructing surrogate models that...

    arxiv.org/abs/2607.28903 · PDF

  6. 06

    Conditioning Tree-Based Diffusions and Flows for Probabilistic Tabular Regression

    Silas Koemen

    stat.ML · cs.LG

    Tree-based diffusion models fit flexible conditional predictive distributions for tabular regression without a neural density estimator, but they inherit their design defaults---noising path, parameterization, training distribution, features, sampler---from the neural setting. We show these defaults are the binding constraint: what a gradient-boosted ensemble actually solves is a supervised regression problem whose conditioning they...

    arxiv.org/abs/2607.28864 · PDF

This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.