stat.ML · 2026-09-02 · No. 103
Machine Learning, 2026-09-02.
5 new papers in stat.ML. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
5 entries-
01
Variable Selection for Feature-Based Newsvendor
Zhaoliang Yuan, Jie Wang
stat.ML · cs.LG
Feature-based newsvendor models use observable covariates to tailor inventory decisions, aiming to balance holding and shortage costs under demand uncertainty. However, high-dimensional feature sets often hinder interpretability and inflate data collection and implementation costs. This paper studies variable selection for the feature-based newsvendor problem under a hard cardinality constraint on the number of selected features. We formulate...
-
02
On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study
Chathurika S Abeykoon, Mathias Nthiani Muia, Mallory Goldstein
stat.ML · cs.LG
Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstream generalization remains poorly understood. In this work, we develop a statistical framework for conditional generative augmentation and analyze its impact on classification risk. We formalize augmentation as a distribution-mixing process and show that the resulting risk distortion is controlled by both the augmentation strength and...
-
03
Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity
Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate
stat.ML · cs.AI · cs.LG
The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictions for the same inputs (predictive multiplicity). Existing work primarily focuses on multiplicity within individual models, but in more complex decision systems, the impact of the Rashomon effect is less well understood. In this work, we study multiplicity from the perspective of auditing incorrect ensemble predictions, where the...
-
04
Matched Queries for Curvature and Density at Branching Junctions
Ziqi Zhao, Qingjian Ni
stat.ML · cs.LG
At a junction, a score field can reveal weighted tangent rays, yet these first-order quantities do not determine how individual branches bend or how their densities change away from the center. Recovering this missing information is necessary for describing local continuation beyond a single point, but finite observations must separate branchwise second-order effects while allowing error in the estimated center. We address this inverse...
-
05
Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches
Marco Simnacher, Georg Keilbar, Benjamin König, Christoph Lippert, Sonja Greven
stat.ML · cs.AI · cs.LG · math.ST · stat.ME
Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ given a third random object $Z$. Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. However, we show that such tests are of interest for large language model (LLM) outputs, where we test whether an output $X$ generated from a source text $Z$ carries information about an attribute...
This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.