stat.ML · 2026-06-18 · No. 27

Machine Learning, 2026-06-18.

7 new papers in stat.ML. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

7 entries
  1. 01

    Generalised Eigenvalue Geometry of Semantic Adversarial Attacks

    Martin Anthony, Kaveh Salehzadeh Nobari

    stat.ML · cs.LG

    Recent empirical work shows that semantically equivalent paraphrases can fool financial sentiment classifiers: although a paraphrase remains close to the original under a strong reference embedding, it may shift the target model's representation enough to change the predicted class. Existing robustness theory either assumes a single-model threat model or focuses mainly on empirical attack algorithms. We develop a continuous local model of...

    arxiv.org/abs/2606.19212 · PDF

  2. 02

    On Local Population-Risk Certificates

    Mingzhi Song

    stat.ML · cs.LG · math.ST

    This paper develops local certificates for population-risk increments around a current model. For a local candidate set \(\mathcal D\), the certificate is a two-sided confidence band for \(P({\ell_{θ+v}-\ell_θ})\) over \(v\in\mathcal D\). As an application, the upper endpoint of this band yields a risk-controlled update rule: an update is accepted only when its certified upper endpoint is nonpositive; otherwise the current model is retained.

    arxiv.org/abs/2606.19147 · PDF

  3. 03

    Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning

    Zilong Zhang, Yi-Ting Hung, Lei Ding, Chi-Kuang Yeh

    stat.ML · cs.LG · stat.CO · stat.ME

    Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic quality, most notably verbosity bias. Meanwhile, human supervision is costly and typically selective, yielding reliable positive judgments but leaving most outputs unlabelled and potentially mixed in quality. We formulate LLM evaluation under selective human...

    arxiv.org/abs/2606.19057 · PDF

  4. 04

    Sequential Kernel-based Conditional Independence Testing via Adaptive Betting

    Zheng He, Danica J. Sutherland

    stat.ML · cs.LG · stat.ME

    Testing conditional independence is fundamental yet intrinsically difficult: without additional assumptions, Type I error control is impossible in general. The "Model-X'' paradigm addresses this difficulty by assuming exact knowledge of a relevant conditional distribution. While small deviations from this assumption can sometimes be tolerated in classical one-shot testing, existing sequential conditional independence tests typically require...

    arxiv.org/abs/2606.18993 · PDF

  5. 05

    FOSC-X: An Extended Framework for Optimal Local Cuts and Non-Horizontal Cluster Selection from Clustering Hierarchies

    Connor Simpson, Ricardo J. G. B. Campello

    stat.ML · cs.LG

    Extracting a flat clustering solution from a hierarchy is a common task in practical cluster analysis and can be formulated as an optimisation problem. Existing approaches focus on finding a single optimal solution. We introduce FOSC-X, a framework for extracting the top-M globally optimal flat clusterings from local, non-horizontal cuts of a hierarchical cluster tree, while optionally enforcing constraints on the number of clusters. This...

    arxiv.org/abs/2606.18972 · PDF

  6. 06

    Kernel of Partition Paths: A Unified Representation for Tree Ensembles

    Nicolas Mahler

    stat.ML · cs.LG

    A recent line of work has reframed individual decision trees as linear models on engineered features associated with their splits, opening routes for oracle inequalities and feature-importance reinterpretation, but leaving open the question of what unified geometric object a forest induces when one indexes its feature map by nodes rather than by splits. The present paper studies that object. KPP indexes the feature map by the nodes of the...

    arxiv.org/abs/2606.18853 · PDF

  7. 07

    TimeLAVA: Learning-Agnostic Data Valuation for Time Series

    Wenqin Liu, Weizhi Quan, Aoqi Zuo, Erdun Gao, Vu Nguyen, Dino Sejdinovic, Howard Bondell, Mingming Gong

    stat.ML · cs.LG

    Data valuation quantifies the intrinsic quality of individual samples to enable principled data curation, quality control, and robust learning. For time series in critical domains such as healthcare, finance, and industrial monitoring, effective valuation methods are essential yet fundamentally lacking. Existing approaches are either model-dependent, limiting their generalizability, or designed for i.i.d. data and thus fail to capture...

    arxiv.org/abs/2606.18729 · PDF

This edition is part of The Daily Abstract — stat.ML archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.