stat.ML · 2026-09-02 · No. 103

Machine Learning, 2026-09-02.

5 new papers in stat.ML. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

5 entries
  1. 01

    Variable Selection for Feature-Based Newsvendor

    Zhaoliang Yuan, Jie Wang

    stat.ML · cs.LG

    Feature-based newsvendor models use observable covariates to tailor inventory decisions, aiming to balance holding and shortage costs under demand uncertainty. However, high-dimensional feature sets often hinder interpretability and inflate data collection and implementation costs. This paper studies variable selection for the feature-based newsvendor problem under a hard cardinality constraint on the number of selected features. We formulate...

    arxiv.org/abs/2609.01544 · PDF

  2. 02

    On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study

    Chathurika S Abeykoon, Mathias Nthiani Muia, Mallory Goldstein

    stat.ML · cs.LG

    Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstream generalization remains poorly understood. In this work, we develop a statistical framework for conditional generative augmentation and analyze its impact on classification risk. We formalize augmentation as a distribution-mixing process and show that the resulting risk distortion is controlled by both the augmentation strength and...

    arxiv.org/abs/2609.01410 · PDF

  3. 03

    Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity

    Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate

    stat.ML · cs.AI · cs.LG

    The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictions for the same inputs (predictive multiplicity). Existing work primarily focuses on multiplicity within individual models, but in more complex decision systems, the impact of the Rashomon effect is less well understood. In this work, we study multiplicity from the perspective of auditing incorrect ensemble predictions, where the...

    arxiv.org/abs/2609.01397 · PDF

  4. 04

    Matched Queries for Curvature and Density at Branching Junctions

    Ziqi Zhao, Qingjian Ni

    stat.ML · cs.LG

    At a junction, a score field can reveal weighted tangent rays, yet these first-order quantities do not determine how individual branches bend or how their densities change away from the center. Recovering this missing information is necessary for describing local continuation beyond a single point, but finite observations must separate branchwise second-order effects while allowing error in the estimated center. We address this inverse...

    arxiv.org/abs/2609.01319 · PDF

  5. 05

    Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches

    Marco Simnacher, Georg Keilbar, Benjamin König, Christoph Lippert, Sonja Greven

    stat.ML · cs.AI · cs.LG · math.ST · stat.ME

    Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ given a third random object $Z$. Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. However, we show that such tests are of interest for large language model (LLM) outputs, where we test whether an output $X$ generated from a source text $Z$ carries information about an attribute...

    arxiv.org/abs/2609.00946 · PDF

This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.