stat.ML · 2026-07-07 · No. 46

Machine Learning, 2026-07-07.

9 new papers in stat.ML. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

9 entries
  1. 01

    Fitted Occupancy-Ratio Evaluation without Bellman Completeness

    Lars van der Laan, Nathan Kallus

    stat.ML · cs.LG

    Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class. We propose fitted occupancy-ratio evaluation (FORE), a fitted fixed-point method that characterizes the discounted occupancy ratio through an adjoint Bellman recursion. At each iteration, FORE...

    arxiv.org/abs/2607.05375 · PDF

  2. 02

    msPCA: An R Package for Sparse PCA with Multiple Components

    Ryan Cory-Wright, Jean Pauphilet

    stat.ML · cs.LG · stat.ME

    We present msPCA: an open-source R package for sparse principal component analysis with multiple components. It implements an alternating maximization algorithm to generate a set of sparse loading vectors that collectively explain a large fraction of the variance in a dataset, while remaining non-redundant. The algorithm supports two definitions of non-redundancy: either orthogonality of the loading vectors or zero pairwise correlation...

    arxiv.org/abs/2607.05229 · PDF

  3. 03

    Geometric Causal Models

    Eli N. Weinstein, David M. Blei

    stat.ML · cs.LG · q-bio.BM

    Scientists often seek to draw causal inferences from structured data that is not independently and identically distributed, such as spatial data, network data, or molecular data. We develop geometric causal models (GCMs), a framework for causal inference from dependent data that exploits underlying symmetries of the data generating process. For example, in spatial data, we consider processes that are symmetric under translations, or in graph...

    arxiv.org/abs/2607.05153 · PDF

  4. 04

    Context-Constrained Transfer Learning for Tabular Foundation Models via Data Distillation

    Yijun Lin, Sai Li

    stat.ML · cs.LG

    Tabular Foundation Models (TFMs) have demonstrated strong empirical performance as black-box inference engines through in-context learning. However, their use in transfer learning is limited by two obstacles: strict context-size constraints and sensitivity to distribution shifts between source and target tasks. Directly pooling heterogeneous source data can therefore lead to negative transfer. To address these challenges, we propose...

    arxiv.org/abs/2607.04809 · PDF

  5. 05

    Non-Asymptotic Error Bounds for SMC with Biased Proposals: Application to Conditional Diffusion Sampling

    Stanislas Strasman, Gabriel Victorino Cardoso, Sylvain Le Corff, Vincent Lemaire, Antonio Ocello

    stat.ML · cs.LG

    Sequential Monte Carlo (SMC) methods are a natural tool for post-hoc conditioning of pretrained generative models, but in many applications the mutation kernels used by the particle system are biased approximations of an ideal Feynman--Kac flow. This paper develops a non-asymptotic error analysis for such SMC samplers. Under forward-smoothing forgetting conditions, we decompose the total error into a kernel bias, measuring the effect of...

    arxiv.org/abs/2607.04780 · PDF

  6. 06

    Non-asymptotic Convergence of Stochastic Gradient Descent in Score-based Generative Models

    Stanislas Strasman, Sobihan Surendran, Sylvain Le Corff

    stat.ML · cs.LG

    Score-based Generative Models (SGMs) have achieved impressive performance in data generation across a wide range of applications. While the statistical properties of their sampling procedures are increasingly well understood, the optimization dynamics underlying their training remain less explored. SGMs are typically trained by minimizing a weighted denoising scorematching objective, yet optimization guarantees with stochastic gradients...

    arxiv.org/abs/2607.04775 · PDF

  7. 07

    Wasserstein Residuals: Learning Gradient Flows from Population Dynamics

    Markus Heinonen, Yair Shenfeld, Ricardo Baptista, Daniel Waxman, Dmitry Batenkov, Tim Cooijmans, Eli Bingham

    stat.ML · cs.AI · cs.LG

    Reconstructing population dynamics is a central problem in the physical and data sciences. Often, the dynamics are modeled as a Wasserstein gradient flow (WGF): a curve of distributions driven by an energy functional. Though there are multiple mathematical characterizations of a WGF, the dominant algorithmic approach relies on the Jordan--Kinderlehrer--Otto (JKO) scheme. JKO-based methods are inflexible to time discretisation and require...

    arxiv.org/abs/2607.04738 · PDF

  8. 08

    Decomposition for Bayesian Networks: Local and Parallel Inference

    Pei Heng, Xinyi Hu, Yi Sun

    stat.ML · cs.LG

    Probabilistic inference in high-dimensional Bayesian networks is difficult because exact manipulation of the joint distribution scales exponentially with network size. We propose a decomposition framework based on directed convex subgraphs and introduce a minimal d-decomposition tree. Together, they provide a principled alternative to classical junction-tree constructions. The proposed framework represents the joint distribution by...

    arxiv.org/abs/2607.04650 · PDF

  9. 09

    Integrating Neural Encoders in Bayesian Generalized Linear Mixed Models for Multimodal Data

    Yuankang Zhao, Youngsoo Baek, Felipe A. Medeiros, Samuel Berchuck, Matthew M. Engelhard

    stat.ML · cs.LG · stat.CO

    Scalable Bayesian inference for generalized linear mixed models (GLMMs) provides uncertainty-aware analysis of correlated longitudinal data, but existing scalable approaches largely assume low-dimensional tabular predictors and do not directly accommodate high-dimensional modalities such as images and text. We address this limitation by learning one or more modality-specific neural encoders jointly with a GLMM objective, then performing...

    arxiv.org/abs/2607.04647 · PDF

This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.