stat.ML · 2026-07-27 · No. 66
Machine Learning, 2026-07-27.
11 new papers in stat.ML. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
11 entries-
01
CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
Jiyuan Tan, Vasilis Syrgkanis
stat.ML · cs.AI · cs.LG · econ.EM
Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large language model (LLM) reviewer. However, such reviewers remain empirically unreliable: they may accept fabricated papers and detect them at rates close to chance (Bad Scientist, 2025). We present CausalForge, a framework for automated theoretical...
-
02
Graph-Based Correlation Matrix Generation: A Convex Optimization Approach
Ali Fakhar, K{é}vin Polisano, Ir{è}ne Gannaz, Sophie Achard
stat.ML · cs.LG
This work addresses the generation of theoretical correlation matrices with prescribed sparsity patterns associated to graph structures. We propose a novel convex optimization framework in which an initial matrix is projected onto an elliptope under a positive semidefiniteness constraint. Several numerical schemes are implemented and compared. The problem falls within the broader class of matrix completion, where off-diagonal entries...
-
03
Learning Ergodic Dynamical Systems from a Finite Trajectory
Oleksii Kachaiev, Silvia Villa, Lorenzo Rosasco
stat.ML · cs.LG
We consider the problem of learning from a single finite trajectory of an ergodic stochastic dynamical system. More precisely, we study discrete-time autonomous stochastic systems defining time-homogeneous Markov processes. We first focus on estimating the optimal one-step prediction function by nonlinear least squares, and derive high-probability guarantees measured with respect to the invariant measure of the process. These results make...
-
04
Learning Bidirectional Causal Interactions with Heteroscedastic Neural Networks
Masahiro Tanaka
stat.ML · cs.LG · stat.ME
Estimating contemporaneous bidirectional interactions from observational data is difficult because each outcome is endogenous to the other, while flexible regressions may capture only reduced-form dependence. This paper proposes SEM-DNN, a heteroscedastic neural simultaneous-equation estimator that learns reciprocal structural interactions without external instruments. Identification exploits conditional covariance diagonalization: when...
-
05
Hopformer: Homogeneity-Pursuit Transformer for Time Series Forecasting
Wan Zhang, Qinjie Lin, Chan Lee, Weijian Li, Han Liu, Kai Zhang
stat.ML · cs.LG
Forecasting multiple time-series with high-dimensional covariates presents a core challenge: unifying common temporal patterns while retaining meaningful series-specific information. We introduce Hopformer (Homogeneity-Pursuit Transformer), a two-stage framework that addresses this challenge. In the first stage, we perform a Sparsity Pattern Aggregation (SPA) scheme extracting a common low-variance trend that incorporates the covariates. This...
-
06
General Value Functions for Remaining Useful Life and Failure-Mode Prediction
Hao Yan, Ali Sarabi, Qing Zou, Boyang Xu
stat.ML · cs.LG · stat.AP
Remaining useful life (RUL) prediction and failure-mode classification are central tasks in predictive maintenance. Many data-driven pipelines use fixed-window supervised learning with complete terminal labels; such routes do not naturally encode the temporal recursion linking successive degradation-state predictions when observations are partial or unit identities are unavailable. We formulate prognostics as vector General Value Function...
-
07
Variational Low-rank Tensor Decomposition for Multisubject Spatiotemporal Data Analysis
Laura M. Montaldo, Ricardo A. Borsoi, Sebastian Miron, Tulay Adali
stat.ML · cs.LG · eess.SP
Modeling shared and subject-specific structure in multisubject spatiotemporal data remains challenging, particularly in neuroimaging, where both spatial and temporal patterns exhibit rich variability across subjects. Existing matrix and tensor decompositions provide interpretable factorizations, but rely on fixed multilinear structures or coupling schemes that may limit their flexibility in capturing complex variability. In this work, we...
-
08
Convergence analysis of a family of Zermelo-type iterations for the Bradley--Terry model
Ruijian Han, Ding Lu, Yiming Xu
stat.ML · cs.LG · math.NA · stat.CO
Zermelo's algorithm is a classical method for computing the maximum likelihood estimator in the Bradley--Terry (BT) model, but its convergence can be slow in practice. To accelerate computation, Newman introduced a family of Zermelo-type fixed-point iterations parameterized by $α$, with Zermelo's algorithm recovered at $α=1$. Empirical evidence suggests that the choice $α=0$ often converges substantially faster, making it a promising...
-
09
Efficient Online LLM Watermark Detection via Rao-Blackwellized E-Processes
Lu Luo, Dandan Mo, Chengdong Xu, Ting Li, Jinhan Xie, Huiqiong Li, Niansheng Tang
stat.ML · cs.LG
As large language models (LLMs) are increasingly deployed, reliable and efficient mechanisms for distinguishing AI-generated text from human-written content have become essential. Statistical watermarking has emerged as a promising solution, yet most existing methods are typically fixed-horizon procedures, precluding valid early stopping in streaming generation. In this paper, we develop an efficient online watermark detection framework with...
-
10
Simulation-Based Empirical Bayes
Xinwei Shen, Diana Cai, Cheng Zhang, David M. Blei
stat.ML · cs.LG
Empirical Bayes (EB) performs simultaneous inference across many related latent variables. Classical EB assumes that the likelihood p(x | z) is tractable. In many scientific applications, however, the likelihood is available only through a simulator. This paper develops EB for such implicit likelihoods. We introduce simulation-based empirical Bayes (SBEB), which connects nonparametric EB to simulation-based inference (SBI). SBEB computes EB...
-
11
Prior laundering: learned priors with inherited, undetectable overconfidence
Ali Siahkoohi, Sina Alemohammad
stat.ML · cs.LG
Learned generative priors are increasingly used for ill-posed Bayesian inverse problems, their posterior uncertainty treated as earned from data. But training one requires truths, scarce in seismic and medical imaging, so the recourse is an archive of legacy reconstructions---prior laundering. Where the measurements are uninformative the posterior reverts to the prior, so the uncertainty reported there is the archive's, not the data's, and...
This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.