stat.ML · 2026-08-26 · No. 96

Machine Learning, 2026-08-26.

6 new papers in stat.ML. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

6 entries
  1. 01

    What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation

    Hao Chen

    stat.ML · cs.LG

    Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation. FID's moment restriction has concrete consequences: on ImageNet, visually unrecognizable images optimized only to match the reference Inception mean and covariance obtain FID...

    arxiv.org/abs/2608.24881 · PDF

  2. 02

    $\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions

    Victor Medina-Olivares, Stefan Lessmann, Jonathan Crook

    stat.ML · cs.AI · cs.LG · q-fin.RM

    Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. Logistic regression remains attractive because its coefficients are easy to interpret, but it can miss nonlinear structure. Flexible models can improve prediction, but their explanations are often post-hoc and may not describe the decision rule itself. We introduce $\texttt{findr}$, short for flexible,...

    arxiv.org/abs/2608.24582 · PDF

  3. 03

    Scalable and Versatile Identification for Hierarchical Structural Causal Models: A New Look at Project STAR

    Janis Aiad, Aghiles Drali, Aymen El Ouadrhiri, Anass Ettahiri, Yasser Oufqir, Simon Patry, David Cortes, Marianne...

    stat.ML · cs.AI

    The STAR (Student-Teacher Achievement Ratio) experiment (1985, Tennessee, USA) is a landmark hierarchical dataset designed to assess the impact of class size on student outcomes, with observations nested within classes. To encode class-level interventions in such hierarchical settings, we develop a complete, scalable, open-source pipeline for Hierarchical Structural Causal Models (HSCM) that bridges symbolic identification and practical...

    arxiv.org/abs/2608.24500 · PDF

  4. 04

    Sequential operator learning under dependent data

    Rafael Oliveira

    stat.ML · cs.LG

    Learning operators from sequentially collected data arises in adaptive experimental design, Bayesian optimization, and dynamical-system modelling, where observations may be dependent, and future inputs or sensing operators may depend on preceding data. We derive time-uniform self-normalized concentration bounds for stochastic processes in Hilbert spaces with vector-valued noise. We use these bounds to obtain regression-error guarantees for...

    arxiv.org/abs/2608.24426 · PDF

  5. 05

    A Heterogeneous Mixture of Experts Framework for Interpretable Machine Learning

    Soham Chatterjee, Rwitobroto Dey, Smarajit Bose

    stat.ML · cs.LG

    Mixture-of-Experts (MoE) models provide a flexible framework for partitioning complex prediction problems into simpler local learning tasks through an input-dependent gating mechanism. Existing interpretable MoE approaches, such as Mixture of Decision Trees (MoDT), achieve transparency by employing homogeneous decision-tree experts, but this restricts the model to a single inductive bias across all regions of the feature space. We extend the...

    arxiv.org/abs/2608.24195 · PDF

  6. 06

    qshap: Fast Shapley Decomposition of $R^2$ for Gradient-Boosted Trees

    Zhongli Jiang, Min Zhang, Dabao Zhang

    stat.ML · cs.LG

    Numerous methods have been developed to quantify feature attributions in individual predictions for tree ensembles. However, many applications require global measures of feature contributions to overall model performance. Although local attribution scores can be aggregated to characterize feature importance, such summaries do not directly decompose measures of predictive performance, such as $R^2$. This article introduces qshap, available in...

    arxiv.org/abs/2608.24104 · PDF

This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.