stat.ML · 2026-06-11 · No. 20

Machine Learning, 2026-06-11.

4 new papers in stat.ML. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

4 entries
  1. 01

    Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence

    Itay Lavie, Kirsten Fischer, Andrey Lekov, Frederic Van Maele, Zohar Ringel, Moritz Helias

    stat.ML · cond-mat.dis-nn · cs.LG

    Attention is the key mechanism underlying in-context learning in transformers, and attention patterns have been observed empirically to emerge abruptly during training. We present a Bayesian theory of feature learning in attention; we then focus on how the copy subcircuit in the first layer of an induction head is learned by analyzing a single-layer softmax attention network trained on a copy task. We derive a closed-form posterior over the...

    arxiv.org/abs/2606.12058 · PDF

  2. 02

    From Persistence to Survival: Hypothesis Testing, Effect Sizes and Vectorisation for Topological Features

    Juliette Murris, Bernadette Stolz, Karsten Borgwardt

    stat.ML · cs.LG · math.AT

    Persistence diagrams are common representations in topological data analysis, but they do not naturally live in a vector space, and the statistical tools developed for comparing them have largely evolved separately from those used for downstream prediction. We introduce STRAND (Survival Topological Representation ANalysis of Diagrams), which treats (collections of) PDs as survival data: each topological feature with persistence value $p = d -...

    arxiv.org/abs/2606.11911 · PDF

  3. 03

    Conformal Bayes under Label Shift: Post-Hoc Calibration vs. In-Training Adaptation

    Seungjin Choi

    stat.ML · cs.LG

    Conformal Bayes combines Bayesian posterior predictives with conformal calibration to produce prediction sets that are both statistically valid and geometrically efficient. We study conformal Bayes under label shift from a unified perspective, identifying two complementary approaches that restore nominal target-domain coverage through importance-weighted conformal calibration but operate through independent mechanisms. \emph{Post-hoc...

    arxiv.org/abs/2606.11865 · PDF

  4. 04

    Renewable Lasso without Batch-Number Constraints: A Gradient-Enhanced Approach

    Junzhuo Gao, Ling Peng, Xu Guo, Heng Lian

    stat.ML · cs.LG

    We study online estimation for high-dimensional generalized linear models with streaming data. First, for the non-distributed setting, we propose a gradient-enhanced surrogate loss that approximates the cumulative loss using only historical summaries, which modifies and improves upon the existing renewable estimation approach for the same model in the high-dimensional setting, and removes the batch-number constraint in previous studies. We...

    arxiv.org/abs/2606.11738 · PDF

This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.