stat.ML · 2026-09-23 · No. 122

Machine Learning, 2026-09-23.

11 new papers in stat.ML. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

11 entries
  1. 01

    On Basis Function Selection for Sparse Gaussian Process Regression

    Marnix Van Soom, Ivan De Boi

    stat.ML · cs.LG

    Sparse Gaussian processes achieve $O(N)$ inference by replacing the kernel with an appropriate expansion in a fixed basis $\{φ_j\}$ on the input space. Given a compute budget $M \ll N$, practitioners conventionally truncate the basis to its first $M$ entries. Nothing in the formalism, however, prevents one from selecting only those $M$ basis functions that matter for the data at hand. This would avoid spending budget on basis functions where...

    arxiv.org/abs/2609.26624 · PDF

  2. 02

    A Practical Guide on Graphical Model Validation

    Mario V. Wüthrich

    stat.ML · cs.LG · q-fin.RM

    This manuscript formalizes the most popular model validation tools used in general insurance actuarial modeling. These include graphical tools like calibration plots, actual-vs-expected plots, lift charts, Murphy diagrams, as well as classical statistical tools such as Bregman losses, deviance losses, elementary losses, Murphy's decomposition and Gini scores. Particular emphasis is placed on whether calibration and discrimination are studied...

    arxiv.org/abs/2609.26445 · PDF

  3. 03

    SuperPCA: subspace analysis and an efficient algorithm for high-dimensional PCA

    Irina-Beatrice Haas, Maike Meier, Yuji Nakatsukasa, Taejun Park

    stat.ML · cs.LG · math.NA · stat.CO

    Principal component analysis (PCA) is a fundamental tool to reduce the dimensionality of the data in many applications. PCA finds a few signal directions that contain most of the variability of the data by computing the eigenvectors of the sample covariance matrix. In this work, we focus on the spiked covariance model, in which the data vectors are defined by a few orthogonal signals plus an isotropic Gaussian noise, and our goal is to...

    arxiv.org/abs/2609.26406 · PDF

  4. 04

    Error Bounds for Statistical Estimators in BTL Model with Parametric Multivariate Utility Functions

    Yicheng Li, Huifu Xu

    stat.ML · cs.LG

    We study preference elicitation under the Bradley-Terry-Luce (BTL) model where the true partworth vector is unknown and has to be estimated as a parameter with elicited preference information. The set of selected pairwise queries is non-uniform, deterministic, and arbitrary over a collection of alternatives, provided that it satisfies a joint identifiability condition. We focus on understanding when the canonical maximum likelihood estimator...

    arxiv.org/abs/2609.26326 · PDF

  5. 05

    Learning to Fluctuate: Statistical Foundations for Causal Tabular Pretraining

    Zhiheng Zhang

    stat.ML · cs.LG

    Causal tabular foundation models amortize effect estimation across synthetic mechanisms, but latent-effect supervision rewards posterior shrinkage instead of directly encoding the repeated-sample response needed in a fixed deployment population. We introduce fluctuation-supervised pretraining (FSP): each synthetic table is labeled by its average treatment effect plus its efficient influence-function fluctuation, while deployment remains a...

    arxiv.org/abs/2609.26290 · PDF

  6. 06

    Conditional Tensor Diffusion: Distributional Counterfactual Learning and Inference

    Xinbing Kong, Zeyu Li, Junfan Mao, Bin Wu

    stat.ML · cs.LG · econ.EM

    Causal inference guides operational and managerial decisions but remains challenging in high-dimensional panel or tensor settings, where decisions may depend on the joint conditional distribution of missing control outcomes. We develop \emph{Counterfactual Tucker Diffusion} (\CFTDiff), which integrates the treatment mask and latent Tucker structure into conditional diffusion to recover this distribution given observed control outcomes through...

    arxiv.org/abs/2609.25924 · PDF

  7. 07

    Statistical Gains from Looped Estimation under Parameter Budgets

    Xinyu Tian, Xiaotong Shen

    stat.ML · cs.LG · math.ST

    Growing memory demands in artificial intelligence motivate learning with fewer trainable parameters. We ask whether a looped estimator, which repeatedly applies one fitted operator with parameters shared across iterations, can improve statistical accuracy under a common parameter budget. Its conventional untied counterpart uses separate parameters at each iteration. For general likelihood models, we establish an upper bound on squared...

    arxiv.org/abs/2609.25778 · PDF

  8. 08

    Optimal Tradeoffs Between Network Size and Parameter Magnitude in Neural Approximation and Minimax Regression

    Baicheng Li, Zuowei Shen, Haizhao Yang, Shijun Zhang

    stat.ML · cs.LG

    The statistical accuracy of neural networks depends on both their approximation power and the complexity of the class fitted from data. While increasing network size is a natural way to improve approximation, parameter magnitude provides another resource whose role must be quantified in both respects. We establish a sharp width--magnitude tradeoff at fixed depth using one elementary bounded $1$-Lipschitz Dyadic--Triangular Activation. For the...

    arxiv.org/abs/2609.25710 · PDF

  9. 09

    On the Gradient Heterogeneity Dynamics of Adversarially Robust Federated Regression

    Leonardo F. Toso, James Anderson, Nirupam Gupta, Rafael Pinot

    stat.ML · cs.LG

    Federated learning (FL) is intrinsically heterogeneous: honest clients may have different data-generating models. On top of that, adversarial clients can make heterogeneity even more pronounced by sharing arbitrary updates. Existing analyses typically control the interaction between statistical heterogeneity and adversarial behavior through gradient-dissimilarity conditions. However, the underlying bound is imposed a priori and may yield...

    arxiv.org/abs/2609.25705 · PDF

  10. 10

    Generalized Deep Regression for Repeated Measurements

    Kexuan Li

    stat.ML · cs.LG

    In this paper, we study the estimation of a marginal regression function from independent units with repeated binary, count, or continuous responses using ReLU deep neural networks. In the model, we assume that the dependence is generated by an unobserved random mean function within each unit. We then fit a neural network with a convex generalized regression loss. We show an oracle inequality by separating conditional measurement variation...

    arxiv.org/abs/2609.25605 · PDF

  11. 11

    Scalable Minimum-Volume Simplex Estimation with Non-asymptotic Analysis

    Jun LI, Yanlong Guo, Zhaozhao Zeng

    stat.ML · cs.LG

    We study the estimation of a $K$-dimensional simplex from $N$ i.i.d.\ points sampled uniformly from its interior; the observations are convex combinations of $K+1$ unknown prototypes. Existing polynomial-time estimators need cubic per-sample work or $O(NK)$ storage and are impractical at $N\sim 10^6$--$10^8$. We propose DeepMVSA, which re-expresses the minimum-volume principle in neural implicit form: a lightweight coordinate network...

    arxiv.org/abs/2609.25576 · PDF

This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.