stat.ML · 2026-05-27 · No. 10

Machine Learning, 2026-05-27.

6 new papers in stat.ML. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

6 entries
  1. 01

    Gaussian Process-based learning with new MCMC-based implementation of Wishart prior on correlation matrix

    Kane Warrior, Dalia Chakrabarty

    stat.ML · cs.LG

    In probabilstic supervised learning of an input-output relationship - as a sample function of a Gaussian Process (GP) - priors are typically specified for the hyperparameters of the kernel that parametrises the covariance function of the GP, where the induced covariance matrix of the (resulting multivariate Normal) likelihood, governs the learning and prediction. When the sought function is highly multivariate, multiple lengthscale parameters...

    arxiv.org/abs/2605.27093 · PDF

  2. 02

    Causal Representation Learning for Generalisable Recommendation

    Yorgos Felekis, Michael O'Riordan, Oriol Corcoll, Ciarán M. Gilligan-Lee

    stat.ML · cs.LG · stat.ME

    Predictive models trained on observational data often fail to generalise to the distributions they encounter when deployed, especially when the training data is a product of the system being optimised. Recommender systems are a canonical example: they are trained on interaction logs confounded by the deployed policy, past user behaviour, and platform filtering. As a result, the training distribution differs substantially from the candidate...

    arxiv.org/abs/2605.27043 · PDF

  3. 03

    Constrained Bayesian Experimental Design via Online Planning

    Yujia Guo, Daolang Huang, Xinyu Zhang, Sammie Katt, Samuel Kaski, Ayush Bharti

    stat.ML · cs.LG

    Bayesian experimental design (BED) is a principled framework for data-efficient design of sequential experiments. However, existing BED methods are unable to adapt to dynamic constraints inherent in real-world tasks due to budget limitations, varying costs, or physical constraints that restrict how designs evolve over time. In this paper, we introduce a novel approach to BED that enables constrained optimization of experimental designs by...

    arxiv.org/abs/2605.26990 · PDF

  4. 04

    Signal-to-Noise Ratio and Sample Size Govern Representational Alignment in Neural Networks

    Ali Hussaini Umar, Alessandro Laio

    stat.ML · cond-mat.dis-nn · cs.LG · cs.NE · q-bio.NC

    Neural networks are known to develop latent representations that are $aligned$, namely structurally similar across networks trained with different architectures, training protocols, or training datasets. We study this phenomenon in a controlled setting, where we train an ensemble of networks on regression and classification tasks using training sets perturbed by independent realizations of a noise process. We show that the signal-to-noise...

    arxiv.org/abs/2605.26973 · PDF

  5. 05

    Transformers Can Learn Posterior Predictive Distributions In-Context

    Gyeonghun Kang, Changwoo J. Lee, Xiang Cheng

    stat.ML · cs.LG

    Prior-data fitted networks (PFNs) have recently emerged as a powerful approach for Bayesian prediction tasks, approximating the posterior predictive distribution (PPD) through in-context learning. Despite their strong empirical performance and ability to go beyond point predictions, theoretical understandings of the algorithmic capability of transformers to learn distributions in context are still lacking. Focusing on Gaussian process...

    arxiv.org/abs/2605.26713 · PDF

  6. 06

    CART Random Forests as Sequential Allocation over Random Opportunity Sets: A Stochastic-Control Theory of Ensemble Risk

    Tianxing Mei, Yingying Fan, Mingming Leng, Jinchi Lv

    stat.ML · cs.LG

    CART random forests are among the most widely used modern predictive methods, with well-documented empirical success. Yet, at the mechanistic level, the algorithm is often treated as a black box because of its complexity. In this paper, we develop a stochastic-control perspective on feature-subsampled CART random forests, named CART random opportunity-set allocation (CART-ROSA). At each node, the random subset of features is interpreted as a...

    arxiv.org/abs/2605.26675 · PDF

This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.