cs.LG · 2026-06-09 · No. 18

Machine Learning, 2026-06-09.

64 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

64 entries
  1. 01

    An Agency-Transferring Model-Free Policy Enhancement Technique

    Anton Bolychev, Georgiy Malaniya, Sinan Ibrahim, Pavel Osinenko

    cs.LG · cs.AI · eess.SY · math.OC

    Training reinforcement learning (RL) policies from scratch is costly: it requires careful reward and environment design, extensive tuning, and substantial computation. Yet many control problems already have a functional but suboptimal policy available as a baseline. This paper proposes a method for embedding such a baseline into the RL training process, simultaneously improving training efficiency relative to from-scratch methods and...

    arxiv.org/abs/2606.09825 · PDF

  2. 02

    Rethinking the Divergence Regularization in LLM RL

    Jiarui Yao, Xiangxin Zhou, Penghui Qi, Wee Sun Lee, Liefeng Bo, Tianyu Pang

    cs.LG

    Reinforcement learning (RL) has become a key component of post-training large language models (LLMs). In practice, LLM RL is often off-policy because of training-inference mismatch and policy staleness, making trust-region control essential for stable optimization. Mainstream methods such as PPO and GRPO approximate this control with a ratio-clipping mechanism, but the importance ratio can be a poor proxy for distributional shift in...

    arxiv.org/abs/2606.09821 · PDF

  3. 03

    Topological Neural Operators

    Lennart Bastian, Samuel Leventhal, Mustafa Hajij, Tolga Birdal

    cs.LG · cs.AI

    We introduce Topological Neural Operators (TNOs), a principled framework for operator learning on cell complexes that lifts neural operators (NOs) from functions on points and/or edges to topological domains. TNOs represent data as features defined on cells of varying dimension and model their interactions through Discrete Exterior Calculus, enabling explicit cross-dimensional coupling via gradient-, curl-, and divergence-type operators. The...

    arxiv.org/abs/2606.09806 · PDF

  4. 04

    Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

    Udvas Das, Waris Radji, Debabrota Basu, Odalric-Ambrym Maillard

    cs.LG · cs.AI · stat.ML

    We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, each having its personalized preference vector, and in the presence of context distributions that are drifting over time. Under practitioner-friendly assumptions, we reduce this setting to linear bandit with stationary mean but heteroskedastic and non-stationary noise. We further study the case...

    arxiv.org/abs/2606.09802 · PDF

  5. 05

    Zero Touch Predictive Orchestration: Automating Time-Series Models for the Cloud-Edge Continuum

    Abd Elghani Meliani, Arora Sagar, Adlen Ksentini, Raymond Knopp

    cs.LG · cs.NI

    The Cloud-Edge Continuum (CEC) enables latency-critical applications by distributing resources to the far edge, but its extreme volatility makes proactive Zero Touch Management via time-series forecasting essential. However, orchestrators face a severe "cold start" problem: newly discovered nodes lack the historical data required to train localized predictive models, while generalized models fail to capture unique hardware and microservice...

    arxiv.org/abs/2606.09787 · PDF

  6. 06

    iOSWorld: A Benchmark for Personally Intelligent Phone Agents

    Lawrence Keunho Jang, Mareks Woodside, Geronimo Carom, Andrew Keunwoo Jang, Jing Yu Koh, Ruslan Salakhutdinov

    cs.LG · cs.CL

    A useful phone agent needs to be personally intelligent. It should reason over a user's identity, history, and preferences as they exist on the device, not just follow isolated instructions in an impersonal sandbox. Existing mobile agent benchmarks lack this kind of personalization. We introduce iOSWorld, the first interactive native iOS simulator benchmark built around a persistent user identity spanning 26 newly built iOS apps. These apps...

    arxiv.org/abs/2606.09764 · PDF

  7. 07

    Preserving Plasticity in Continual Learning via Dynamical Isometry

    Andries Rosseau, Robert Müller, Ann Nowé

    cs.LG · cs.AI

    Continual training of deep neural networks under non-stationarity often leads to a progressive loss of plasticity, eventually limiting further learning. We relate plasticity to the empirical Neural Tangent Kernel, and identify dynamical isometry (the condition that layer-wise Jacobian singular values remain close to one) as a key mechanism for preserving plasticity in continual learning. We revisit a class of networks that are...

    arxiv.org/abs/2606.09762 · PDF

  8. 08

    Perturbative Contrastive Physical Learning

    Kyungeun Kim, Amanuel Anteneh, Israel Klich, Olivier Pfister, J. M. Schwarz

    cs.LG · cond-mat.dis-nn

    Responses to perturbations are key to understanding physical systems. The ability to contrast such responses by comparing how a system reacts under slightly different conditions provides a mechanism for learning. Here, we introduce Perturbative Contrastive Physical Learning (PCPL), a general framework in which learning emerges from measurable contrasts between physical states produced by controlled changes to inputs, boundary conditions,...

    arxiv.org/abs/2606.09756 · PDF

  9. 09

    Learning Dynamics Reveal a Hierarchy of Weight-Induced Layerwise Gram Metrics

    Claudio Nordio

    cs.LG · cond-mat.dis-nn

    We study feed-forward ReLU networks with fixed readout and quadratic loss. The aim is to rewrite gradient descent not primarily as a dynamics in weight space, but as a collective dynamics closed in terms of fields defined on the training-set space. For a single hidden layer, the weight variables can be eliminated from the activation dynamics, yielding a closed equation for the residuals governed by a collective kernel that factorizes into an...

    arxiv.org/abs/2606.09744 · PDF

  10. 10

    Tight Sample Complexity of Transformers

    Chenxiao Yang, Nathan Srebro, Zhiyuan Li

    cs.LG

    We tightly characterize the VC dimension of depth-$L$ Transformers with a total of $W$ parameters, mapping an input sequence of length $T$ to a single output, establishing an upper bound of $O(L W \log (T W))$ and a nearly matching lower bound of $Ω(L W \log (T W / L))$. We further tightly characterize the sample complexity of chain-of-thought learning using such a Transformer, showing teacher forcing (i.e. selecting a predictor consistent...

    arxiv.org/abs/2606.09731 · PDF

  11. 11

    Disentanglement with Holographic Reduced Representations

    Jhonny J. Velasquez Olivera, Christo K. Thomas, Walid Saad

    cs.LG

    Disentanglement, the separation of factors of variation in data using neural networks, remains a long-standing challenge in machine learning. Prior work has addressed this problem with variational autoencoders and generative adversarial networks that incorporate ideas from variational inference and information-theoretic constraints. In contrast to methods that rely on continuous representations, we propose a design that treats disentangled...

    arxiv.org/abs/2606.09725 · PDF

  12. 12

    Evaluating the Representation Space of Diffusion Models via Self-Supervised Principles

    Xiao Li, Yixuan Jia, Zekai Zhang, Xiang Li, Lianghe Shi, Jinxin Zhou, Zhihui Zhu, Liyue Shen, Qing Qu

    cs.LG · cs.CV

    Diffusion models have demonstrated remarkable generative capabilities and have also emerged as powerful self-supervised representation learners, yet the connection between these two abilities remains less explored. Drawing inspiration from self-supervised learning (SSL), we introduce a framework for jointly evaluating the representation and generation capabilities of diffusion models. Specifically, we decompose features into invariant and...

    arxiv.org/abs/2606.09718 · PDF

  13. 13

    BrainSurgery: Reproducible and Reliable Declarative Weight Manipulations for Model Editing and Upcycling

    Gianluca Barmina, Annemette Broch Pirchert, Andrea Blasi Núñez, Lukas Galke Poech, Peter Schneider-Kamp

    cs.LG · cs.CL

    As deep learning models scale, managing, inspecting, and modifying large checkpoints has become increasingly challenging. Researchers often need to alter model weights for layer restructuring, precision casting, low-rank factorization, and architectural debugging, yet these workflows often rely on fragile ad-hoc Python scripts. Here, we introduce BrainSurgery, a tool for robust and reproducible "tensor surgery" on neural network checkpoints,...

    arxiv.org/abs/2606.09707 · PDF

  14. 14

    When Do Local Score Models Extrapolate Across Size? A Diagnostic Theory and Benchmark

    Wenjie Xi

    cs.LG · cond-mat.stat-mech

    Scientific generative modeling often requires size transfer, where models trained on small systems are evaluated on larger ones. While translation-invariant architectures enable this evaluation, we show that architectural locality alone does not guarantee stable size extrapolation. Instead, stable extrapolation is governed by the quasi-locality of the Gaussian-smoothed score. Through Tweedie's formula, far-away perturbations can influence...

    arxiv.org/abs/2606.09705 · PDF

  15. 15

    AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis

    Jaber Jaber, Osama Jaber

    cs.LG · cs.DC · cs.PF

    AutoMegaKernel (AMK) compiles a HuggingFace Llama-family model into a single persistent cooperative CUDA kernel that runs the whole forward pass in one launch, with no per-model hand-written CUDA. The contribution is the system, not raw speed. A frozen schedule-IR validator statically certifies deadlock-freedom and race-freedom via static graph checks (not a mechanized proof), so an unsafe agent-proposed schedule is rejected before launch:...

    arxiv.org/abs/2606.09682 · PDF

  16. 16

    Transition-Based Digital Twin Modelling for Alzheimer's Disease under Sparse Longitudinal Data

    Yinyu Huang, Yilin Zhang, Sofia Michopoulou, Christopher Kipps, Rahman Attar

    cs.LG · cs.AI

    Alzheimer's disease (AD) progression is highly heterogeneous and is typically observed through sparse and irregular longitudinal data, posing challenges for prediction and personalised monitoring. Existing machine learning approaches have improved AD prediction using multimodal data, yet often focus on static classification or cohort-level risk estimation, providing limited support for subject-specific modelling and uncertainty-aware...

    arxiv.org/abs/2606.09671 · PDF

  17. 17

    Algorithm for Contextual Queueing Bandits with Rate-Optimal Queue Length Regret

    Seoungbin Bae, Dabeen Lee

    cs.LG

    Contextual queueing bandits provide a framework for learning to schedule heterogeneous jobs under unknown context-dependent service rates. Under stochastic contexts, existing algorithms achieve $\widetilde{\mathcal{O}}(T^{-1/4})$ queue length regret, defined as the expected difference between the learner's and oracle's queue lengths at horizon $T$. In this paper, we improve this rate to $\widetilde{\mathcal{O}}(T^{-1/2})$. The key observation...

    arxiv.org/abs/2606.09668 · PDF

  18. 18

    In-Context Learning for Latent Space Bayesian Optimization

    Tuan A. Vu, Harri Lähdesmäki, Julien Martinelli

    cs.LG · stat.ML

    Bayesian optimization (BO) is a central tool for sample-efficient design, and latent-space Bayesian optimization (LSBO) extends it to structured objects such as molecules and proteins. In parallel, tabular foundation models such as TabPFN and TabICL now achieve state-of-the-art regression performance and are increasingly used as BO surrogates. Because their Bayesian behavior is induced by large synthetic pretraining collections, the...

    arxiv.org/abs/2606.09664 · PDF

  19. 19

    Muon Learns More Robust and Transferable Features than Adam

    Tianyu Ruan, Fengzhuo Zhang, Shuche Wang, Shihua Zhang

    cs.LG · cs.AI

    Muon has recently emerged as a state-of-the-art optimizer for pretraining Large Language Models (LLMs) and vision classifiers. Despite its efficiency advantage over Adam and SGD, the feature-learning advantage of Muon remains unclear. This paper investigates Muon's feature-learning advantage through the lens of robustness and transferability. First, by evaluating pretrained models on corrupted images and texts, we show that features learned...

    arxiv.org/abs/2606.09658 · PDF

  20. 20

    A Unifying Framework for Concept-Based Representational Similarity

    Grégoire Dhimoïla, Victor Boutin, Agustin Martin Picard, Thomas Fel, Thomas Serre

    cs.LG

    Learned representations across models and modalities often exhibit striking structural similarities, suggesting shared underlying concept decompositions. However, concept alignment remains poorly defined: existing approaches optimize different objectives under the same terminology, obscuring what is actually aligned. We propose a unifying framework that decomposes alignment along two axes: what is aligned (representations vs. concepts) and at...

    arxiv.org/abs/2606.09653 · PDF

  21. 21

    Data-driven discovery of governing differential equations across physical systems

    Siyu Lou, Hao Xu, Wenguan Wang, Lu Lu, Hao Sun, Yang Liu, Linfeng Zhang, Dongxiao Zhang, Yuntian Chen

    cs.LG · cs.SC · math-ph · physics.comp-ph · stat.AP

    Differential equations play a critical role in scientific discovery because they provide a mathematical framework to describe the behaviour of physical phenomena. As a promising alternative to traditional first principles, data-driven differential equation discovery has attracted increasing attention for its ability to infer governing laws directly from experimental or simulated data, especially when the underlying physics is unclear....

    arxiv.org/abs/2606.09638 · PDF

  22. 22

    Constrained user-item allocation for e-commerce marketing campaigns

    Maja Lindström, Natalija Glisovic, Jan von Pichowski, Tommy Löfstedt, Martin Rosvall

    cs.LG

    When running marketing campaigns, retailers must decide which products to promote and which users to target. These decisions are inherently coupled: effective campaigns match users and items with strong mutual affinity into non-overlapping groups of predefined sizes. However, existing approaches assume predefined campaign structure or decouple item selection from user assignment, and cannot discover campaign groupings directly from joint...

    arxiv.org/abs/2606.09623 · PDF

  23. 23

    Closure-Validated Circuit Discovery in Attention Heads: Co-activation Proposes, Ablation Disposes

    Yongzhong Xu

    cs.LG · cs.AI

    Interpretability increasingly treats groups of components, not individual units, as the basic object, and proposes to find them by clustering co-activation statistics. We ask whether such a cheap signal actually identifies an attention-head circuit. Adapting a sparse-autoencoder clustering recipe to attention heads -- but validating by causal ablation rather than reconstruction -- we cluster heads and then run a closure test: ablate the...

    arxiv.org/abs/2606.09607 · PDF

  24. 24

    Assessing Sample Quality in Conditional Generation under Compositional Shift

    Berker Demirel, Valentino Maiorca, Marco Fumero, Theofanis Karaletsos, Francesco Locatello

    cs.LG

    Conditional generators provide a natural tool for controllable generation, including settings where the desired condition is a new composition of observed attributes or experimental factors. In many applications, especially in scientific domains, such models are attractive to explore conditions for which real samples are rare, expensive, or not yet observed. However, this creates a circularity for evaluation: standard conditional quality...

    arxiv.org/abs/2606.09601 · PDF

  25. 25

    On Choosing the $μ$ Parameter in Gaussian Differential Privacy

    Bogdan Kulynych, Antti Honkela

    cs.LG · stat.ML

    Recent work argues for using Gaussian differential privacy (GDP) to report the privacy guarantees in privacy-preserving machine learning. We provide principled mappings from pure-DP $\varepsilon$ to GDP $μ$ by matching the worst-case success of a strong-adversary membership inference attack in terms of three metrics: multiplicative advantage at fixed FPR, precision at fixed recall, and the standard privacy profile. We tabulate $μ$ values...

    arxiv.org/abs/2606.09582 · PDF

  26. 26

    Safe-RULE: Safe Reinforcement UnLEarning

    Shixiong Jiang, Taozheng Zhu, Fanxin Kong

    cs.LG · cs.AI · cs.CR · cs.RO

    Offline safe reinforcement learning (Safe RL) enables policy learning without online interactions, making it suitable for safety-critical systems such as robotics systems. However, its reliance on static datasets exposes offline Safe RL to data poisoning attacks, where adversaries inject malicious samples that compromise safety and induce unsafe policy behavior. In this work, we propose a new learning paradigm, named safe reinforcement...

    arxiv.org/abs/2606.09559 · PDF

  27. 27

    Efficient Traffic Prediction at Scale: A Systematic Study of STGCN Architectural Depth

    Soban Nasir Lone, Mohamed Abouelela, Taeyoung Yu, Jiwon Kim, Constantinos Antoniou

    cs.LG

    Spatio-temporal graph neural networks (STGNNs) have become the dominant approach for traffic prediction, yet their computational requirements pose challenges for practical deployment in intelligent transportation systems (ITS). While recent work has proposed efficient alternatives to STGNNs, a fundamental question remains unexplored: are these architectures themselves over-parameterised? We examine this question using the Spatio-Temporal...

    arxiv.org/abs/2606.09539 · PDF

  28. 28

    Investigating Calibration Challenges in Probabilistic Electricity Price Forecasting

    Jan Niklas Lettner, Hadeer El Ashhab, Benjamin Schäfer

    cs.LG

    As renewable energy integration increases market volatility, probabilistic electricity price forecasting has become essential for effective risk management. However, current-proper-scoring rules often prioritize forecast sharpness at the expense of calibration, leading to overconfident and statistically unreliable uncertainty estimates. This work highlights the critical gap between theoretical scoring and practical calibration, demonstrating...

    arxiv.org/abs/2606.09517 · PDF

  29. 29

    BUDDY: BUdget-Driven DYnamic Depth Routing for Adaptive Large Language Model Inference

    Yuhua Zhou, Shaoqi Yu, Shichao Weng, Changhai Zhou, Mingze Yin, Fei Yang, Aimin Pan

    cs.LG

    Large language models (LLMs) incur high inference cost due to their depth and parameter scale. Depth pruning can reduce latency by skipping redundant Transformer blocks, but existing methods (i) provide limited control under user-specific compute budgets and (ii) typically fix the routing path, failing to adapt as the context grows during decoding. We propose Buddy, a budget-driven dynamic depth routing framework. Buddy uses a lightweight...

    arxiv.org/abs/2606.09514 · PDF

  30. 30

    Loss-Guided Adaptive Scale Refinement for Molecular Force Prediction

    Limin Yu

    cs.LG

    Molecular systems involve interactions across multiple spatial scales, from local coordination and short-range perturbations to long-range electrostatic and solvent-mediated effects. However, most molecular representation learning methods rely on manually predefined scales, and the task-optimal modeling scale may not coincide with these fixed levels. This study introduces a loss-guided adaptive scale refinement framework for molecular force...

    arxiv.org/abs/2606.09480 · PDF

  31. 31

    Escaping the KL Agreement Trap in On-Policy Distillation

    Haoran Xin, Anhao Zhao, Ying Sun, Jin Li, Xiaoyu Shen, Hui Xiong

    cs.LG · cs.CL

    On-policy distillation (OPD) provides dense token-level supervision by asking a teacher to score student-generated rollouts. However, when the student drifts into an unrecoverable prefix, the teacher may locally agree with the degraded state, producing low reverse KL but little corrective training signal. We identify this persistent regime as a low-KL agreement trap. Further analyses show that tokens during and after such traps produce less...

    arxiv.org/abs/2606.09471 · PDF

  32. 32

    Breaking the Tokenizer Barrier: On-Policy Distillation across Model Families

    Yifan Niu, Han Xiao, Dongyi Liu, Zelong Wang, Dihong Gong, Yasheng Wang, Jia Li

    cs.LG

    On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain experts to student models. However, existing OPD distillation methods require teacher and student models to share the same tokenizer, restricting the applicability of OPD within the model series. Current mainstream practice typically employs Supervised Fine-Tuning (SFT) on teacher-generated...

    arxiv.org/abs/2606.09456 · PDF

  33. 33

    Operator learning for solving Fokker-Planck equations with various initial conditions

    Li Zeng, Xiaoliang Wan, Yaobin Wang, Fabio Nobile, Tao Zhou

    cs.LG

    The Fokker-Planck equation (FPE) plays a pivotal role in describing the time evolution of probability density functions (PDFs) for systems governed by stochastic dynamics. In this work, we propose a conditional normalizing flow-based physics-informed neural network (PINN) framework for efficiently approximating the solution operator of the FPE for a whole range of initial conditions. Leveraging the Chapman-Kolmogorov equation for Markovian...

    arxiv.org/abs/2606.09434 · PDF

  34. 34

    Graph Mamba Operator: A Latent Simulator for Interacting Particle Systems

    Karn Tiwari, Niladri Dutta, N M Anoop Krishnan, Prathosh A P

    cs.LG

    Modeling interacting dynamical systems requires capturing spatial interactions alongside long-range temporal dependencies. Graph neural networks (GNNs) provide a natural representation but typically rely on autoregressive rollouts and treat spatial and temporal dynamics separately, leading to error accumulation over long horizons. Existing approaches also focus on local interactions and short temporal contexts, limiting their ability to...

    arxiv.org/abs/2606.09432 · PDF

  35. 35

    LargeMonitor: Monitoring Online Task-Free Continual Learning via Large Pretrained Models

    Mingqi Yuan, Xiaoquan Sun, Shihao Luo, Jiayu Chen

    cs.LG · cs.AI

    Online task-free continual learning (TFCL) requires intelligent agents to sequentially accumulate knowledge from an unbounded, non-stationary data stream under strict single-pass constraints and without any explicit task identifiers. Existing online TFCL paradigms primarily rely on parameter-efficient prompt tuning or dynamic structure expansion driven by training-coupled optimization dynamics, such as empirical loss fluctuations or evolving...

    arxiv.org/abs/2606.09430 · PDF

  36. 36

    Benchmarking Empirical Privacy Protection for Adaptations of Large Language Models

    Bartłomiej Marek, Lorenzo Rossi, Vincent Hanke, Xun Wang, Michael Backes, Franziska Boenisch, Adam Dziedzic

    cs.LG · cs.CR

    Recent work has applied differential privacy (DP) to adapt large language models (LLMs) for sensitive applications, offering theoretical guarantees. However, its practical effectiveness remains unclear, partly due to LLM pretraining, where overlaps and interdependencies with adaptation data can undermine privacy despite DP efforts. To analyze this issue in practice, we investigate privacy risks under DP adaptations in LLMs using...

    arxiv.org/abs/2606.09401 · PDF

  37. 37

    Distilling Safe LLM Systems via Soft Prompts for On Device Settings

    Motasem Alfarra, Cristina Pinneri, Dana Kianfar, Mohammed Almousa, Christos Louizos

    cs.LG

    Deploying safe large language models (LLMs) on resource-constrained edge devices presents a critical challenge: while dual-model systems combining LLMs with guard models provide effective safety guarantees, their substantial memory and computational demands make them prohibitively expensive for on-device deployment. This paper presents a comprehensive study of parameter-efficient safety alignment methods for resource-constrained settings....

    arxiv.org/abs/2606.09388 · PDF

  38. 38

    Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short

    Han Zhou, Adam X. Yang, Laurence Aitchison, Anna Korhonen, Albert Q. Jiang

    cs.LG · cs.AI · cs.CL

    Reinforcement learning with verifiable rewards (RLVR) has become a leading paradigm for improving the reasoning ability of large language models through outcome-based supervision. However, verifiable rewards frequently become uninformative at the group level: when all sampled traces of a given prompt receive identical rewards, group-relative advantage estimation provides no gradient signal, even though the traces may differ substantially in...

    arxiv.org/abs/2606.09380 · PDF

  39. 39

    Scaling Neural Network Verification with Tensor Parallelism and Fully Sharded Data Parallelism

    Sergei Vorobyov, Eugene Ilyushin

    cs.LG · cs.AI

    Formal neural network verification -- proving that a network satisfies safety properties for \emph{all} inputs in a specified domain -- is bounded in practice by GPU memory: standard implementations of bound-propagation algorithms (IBP, CROWN, $α$-CROWN) require weight and relaxation-coefficient matrices to reside entirely on one accelerator. We adapt two parallelism techniques originally developed for large-scale model training to the...

    arxiv.org/abs/2606.09377 · PDF

  40. 40

    PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment

    Yang Tian, Rui Wang, Xumeng Wen, Junjie Li, Shizhao Sun, Lei Song, Jiang Bian, Bo Zhao

    cs.LG · cs.CL

    Long-horizon agentic tasks pose a fundamental credit assignment challenge for outcome-base reinforcement learning: trajectory-level rewards verify final correctness but provide limited guidance on which intermediate reasoning steps or tool interactions contribute to the outcome. The difficulty is especially pronounced in multi-turn search agents, where successful trajectories may contain misleading actions and failed trajectories may contain...

    arxiv.org/abs/2606.09348 · PDF

  41. 41

    Thresholded Local Hyper-Flow Diffusion

    Meher Chaitanya, Sebastian Dalleiger, Luana Ruiz

    cs.LG

    Local Hyper-Flow Diffusion (HFD) gives an edge-size-independent Cheeger-type guarantee for seeded clustering in general submodular hypergraphs, but existing HFD solvers do not keep intermediate computation local at every iteration. We introduce Thresholded Local HFD (TL-HFD), a first-order method that maintains an active region around the seeds, performs projected subgradient updates on that region and its immediate boundary, and expands via...

    arxiv.org/abs/2606.09340 · PDF

  42. 42

    A Universal Dense Football Event Representation Based on TabTransformer

    Weiran Yang, Daniel Memmert, Maximilian Klemp-Weins

    cs.LG · cs.AI

    Football event data constitute a rich spatiotemporal source for quantitative analysis of player actions in team sports. These datasets contain heterogeneous features, combining continuous location coordinates with categorical variables such as action type, action outcome, and body part. Such data have been applied in sports analytics for match outcome forecasting, player evaluation, and tactical pattern recognition. However, existing...

    arxiv.org/abs/2606.09327 · PDF

  43. 43

    Machine-Learning Emulation of Satellite Greenhouse Gas Retrievals: Stability over Time

    Nugzar Gognadze, Motonobu Kanagawa, Yu Someya, Hisashi Yashiro

    cs.LG · stat.AP

    Retrieval algorithms are used to estimate atmospheric concentrations of greenhouse gases (GHGs), such as carbon dioxide (CO2) and methane (CH4), by solving inverse problems from high-spectral-resolution satellite radiance measurements. However, these algorithms are computationally expensive, which makes real-time estimation at scale difficult. Machine-learning models have therefore been proposed as fast emulators of retrieval algorithms. Most...

    arxiv.org/abs/2606.09313 · PDF

  44. 44

    Toward Compiler World Models: Learning Latent Dynamics for Efficient Tensor Program Search

    Haolin Pan, Lianghong Huang, Xvlin Zhou, Mingjie Xing, Yanjun Wu

    cs.LG · cs.PL

    Tensor program optimization is essential for modern machine learning systems, but its search space is enormous. Existing auto-schedulers reduce measurement cost with learned cost models, yet they usually evaluate each candidate as a static code snapshot, ignoring the schedule trajectory that produced it. This makes them insensitive to action dependencies and vulnerable to superficial code variations. We propose a \emph{world-model-inspired}...

    arxiv.org/abs/2606.09312 · PDF

  45. 45

    PRISM: Topology-Aware Cross-Modal Imputation for Modality-Deficient Federated Graph Learning

    Zekai Chen, Miao Zhang, Jiayang Xing, Xunkai Li, Xun Wu, Rong-Hua Li, Guoren Wang

    cs.LG

    Multimodal federated graph learning (MM-FGL) aims to collaboratively learn from decentralized graphs with text and images. However, real-world clients may not share a common modality basis: a visual-search client may contain image--interaction graphs but no seller descriptions, while a catalog client may provide text but no product images. We refer to this practical setting as client-level modality deficiency. Unlike random instance-wise...

    arxiv.org/abs/2606.09301 · PDF

  46. 46

    Intention Driven Identification of In-Possession Match Phases in Association Football through Temporal Graph Learning

    Yuesen Li, Daniel Link

    cs.LG

    Understanding tactical organisation of association football, hereafter referred to as football, requires identifying distinct match phases. Yet in-possession phases are rarely directly observable and are shaped by evolving tactical intentions, rather than spatial patterns alone. This study proposes a data-driven framework for identifying in-possession match phases from spatiotemporal tracking data. Seven German Bundesliga matches recorded at...

    arxiv.org/abs/2606.09289 · PDF

  47. 47

    Trajectory Geometry of Transformer Representations Across Layers

    Vishal Pandey, Gopal Singh

    cs.LG

    Understanding how transformer representations evolve across layers, not merely what they encode, remains an open problem in mechanistic interpretability. We recast the transformer forward pass as a discrete population trajectory through a high-dimensional representation manifold, drawing on geometric tools from computational neuroscience. Rather than probing for pre-specified features, we characterize trajectory geometry using five metrics...

    arxiv.org/abs/2606.09287 · PDF

  48. 48

    Internalizing Geometric Law: Learning from Solver Residuals for Precision-Critical Generation

    Rafael Cabral, Pang Zixi, Ziyi Shou, Shen Xin

    cs.LG · cs.AI

    Large Language Models frequently hallucinate in precision-critical domains such as technical diagramming and mechanical design, where outputs must satisfy strict geometric constraints. We study open-ended geometric synthesis from natural language: translating free-form descriptions into precise constructions whose entities must simultaneously satisfy dozens of interacting constraints. To make this tractable, we release PyGeoX, a programmable...

    arxiv.org/abs/2606.09278 · PDF

  49. 49

    ERBench: A Benchmark and Testsuite for Equation Discovery Algorithms

    Paul Kahlmeyer, Henrik Voigt, Michael Habeck, Joachim Giesen

    cs.LG

    Equation discovery aims to automate the discovery of scientific models in the form of mathematical equations from data. Technically, equation discovery is implemented by symbolic regression algorithms. Performance of symbolic regression for equation discovery is measured along two dimensions: Prediction accuracy on test data, and recovery of known groundtruth formulas. For standard regression, accuracy is typically measured on in-domain test...

    arxiv.org/abs/2606.09276 · PDF

  50. 50

    BSTabDiff: Block-Subunit Diffusion Priors for High-Dimensional Tabular Data Generation

    Al Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Gyawali, Gianfranco Doretto, Donald A. Adjeroh

    cs.LG · cs.AI · stat.ML

    High-Dimensional Low-Sample Size (HDLSS) tabular domains (e.g., omics) are characterized by $n \ll m$, where $n$ = number of samples, and $m$ = number of features. Such domains often exhibit strong local correlation groups, sparse cross-group dependencies, heavy-tailed non-Gaussian marginals, heteroscedastic noise, and structured missingness, making direct density learning in $\mathbb{R}^m$ ill-conditioned since $n \ll m$. We propose...

    arxiv.org/abs/2606.09257 · PDF

  51. 51

    Orange Lab: Lowering Barriers to Data Mining through Embedded Interactive Workflows

    Matej Bevec, Aleš Erjavec, Vesna Tanko, Lena Trnovec, Lan Žagar, Ana Farič, Janez Demšar, Blaž Zupan

    cs.LG · cs.HC

    While visual programming of data analysis workflows has become an important vehicle for the democratization of data science, such systems remain largely confined to standalone applications and offer limited support for transitioning their visual analytics solutions into interactive web environments. As a result, data analysis pipelines are difficult to share, embed, and adapt into user-facing analytical tools. We present Orange Lab, a...

    arxiv.org/abs/2606.09239 · PDF

  52. 52

    The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection

    Hyunseok Paeng

    cs.LG · cs.CL · cs.CR

    We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injection-free baseline. In safety-trained Claude models, documents containing prompt injections suffer a sharp drop in recommendation rate, and this suppression propagates beyond the injected...

    arxiv.org/abs/2606.09204 · PDF

  53. 53

    Asymptotic Optimality of Thompson Sampling for Risk-Averse Bandits with Sub-Gaussian Rewards

    Joel Q. L. Chang

    cs.LG · stat.ML

    We prove that $ρ\text{-}\mathrm{NPTS}_{\mathrm{SG}}$, an anchor-free nonparametric Thompson Sampling algorithm for risk-averse bandits, achieves regret matching the instance-dependent lower bound to leading order in $\log n$, establishing it as asymptotically optimal for any continuous risk functional $ρ$ (CVaR, mean-variance, Sharpe ratio, distortion risk measures, and more) on the class of distributions with bounded density and sub-Gaussian...

    arxiv.org/abs/2606.09191 · PDF

  54. 54

    CANS: Accelerating Multiuser Collaborative Edge Inference via Cooperative Autodidactic NeuroSurgeon

    Zheshun Wu, Ziyang Zhang, Changyao Lin, Zenglin Xu, Jie Liu

    cs.LG · cs.AI · cs.DC

    Recently, mobile edge computing (MEC)-enabled collaborative deep neural network (DNN) inference has emerged as a promising approach for delivering intelligent services to resource-constrained mobile devices. A representative scenario is multi-user collaborative edge inference, where distinct devices independently partition their DNN models and offload backend computation to a common edge server over wireless networks. However, determining the...

    arxiv.org/abs/2606.09175 · PDF

  55. 55

    Crop Recommendation and Agricultural Query Answering System Using Spatio-Temporal Graph Neural Networks and Hybrid Retrieval Augmentation

    Prajwal Thapa, Yagya Raj Pandeya

    cs.LG · cs.AI

    This paper presents a unified system designed to support precision agriculture by integrating advanced weather prediction, crop recommendation, and a question-answering tool for farmers. We propose two deep learning models -- a Transformer-based Graph Neural Network and a Spatio-Temporal Graph Convolutional Network (STGCN) -- to forecast weather conditions for the next 30 days using data from 1,359 locations in Nepal. The STGCN outperforms...

    arxiv.org/abs/2606.09160 · PDF

  56. 56

    Improved Convergence Analysis of Topology Dependence in Decentralized SGD

    Yuki Takezawa, Anastasia Koloskova, Sebastian U. Stich

    cs.LG

    Decentralized SGD is a fundamental algorithm in decentralized learning, although the influence of an underlying network topology on its convergence behavior is not yet fully understood. Existing convergence analyses have shown that topologies with a small spectral gap significantly deteriorate the convergence rate of Decentralized SGD in both homogeneous and heterogeneous cases. However, many prior papers have reported that indeed the choice...

    arxiv.org/abs/2606.09154 · PDF

  57. 57

    Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning

    Daoyu Wang, Mingyue Cheng, Qingchuan Li, Shuo Yu, Jie Ouyang, Qi Liu

    cs.LG · cs.CL

    Agentic reinforcement learning (RL) has become an important post-training paradigm for turning LLMs from static chatbots into interactive agents, giving rise to representative applications such as OpenClaw. Existing work mainly focuses on policy optimization algorithms and training frameworks, but pays less attention to the full data lifecycle of agent-environment interactions, from data production to training consumption. To bridge this gap,...

    arxiv.org/abs/2606.09138 · PDF

  58. 58

    Optimizing Energy-based Neural Network Training with Coherent Ising Machine

    Chen-Rui Fan, Bo Lu, Zhi-Hong Zhang, Run-Qing Zhang, Jing-Wei Wen, Chuan Wang

    cs.LG · cs.AI

    While Ising machines serve as advanced physical solvers for the Ising model,enabling applications in combinatorial optimization and neural network training,their scalability for large-scale neural networks remains constrained by hardware connectivity limitations and suboptimal training methodologies. In this work,we leverage a Coherent Ising Machine (CIM) to train an energy-based neural network using Equilibrium Propagation, achieving...

    arxiv.org/abs/2606.09117 · PDF

  59. 59

    Counterfactual Transport Flows for Offline Conservative Trajectory Refinement

    Lena Krieger, Xuan Zhao, Zhuo Cao, Qin Wang, Hanno Scharr, Ira Assent

    cs.LG

    Offline reinforcement learning (RL) offers a path to policy improvement from logged data alone, using historical returns or other measurable outcomes as world feedback. A key difficulty is improving observed behavior without extrapolating beyond what the offline data supports. We propose \emph{counterfactual transport flows}, a source-conditioned trajectory refinement framework for offline decision-making guided by world feedback. Given a...

    arxiv.org/abs/2606.09115 · PDF

  60. 60

    Hybridizing Equilibrium Propagation with Ising Machines for Efficient Energy-Based Learning

    Chen-Rui Fan, Bo Lu, Xing-Yu Wu, Tie-Jun Wang, Chuan Wang

    cs.LG · cs.AI

    The rapid evolution of artificial intelligence has led to substantial advances in deep neural networks. Nonetheless, conventional GPU-based training remains highly energy-demanding, motivating the exploration of physical dynamics and compatible energy-based learning schemes, such as equilibrium propagation (EP). EP-based training, however, frequently suffers from convergence to local minima due to phase-space contraction. Here we introduce an...

    arxiv.org/abs/2606.09112 · PDF

  61. 61

    Addressing Market Regime Changes and Heavy-Tailed Returns in Portfolio Optimization via Bayesian VAR and Elliptical Black-Litterman

    Daniil Mikriukov, Ruoyu Sun, Angelos Stefanidis, Jionglong Su, Zhengyong Jiang

    cs.LG · cs.AI · q-fin.PM

    Deep reinforcement learning (DRL) frameworks for portfolio optimization have shown promise for their ability to learn allocation rules dynamically from market data. However, these models fail to account for fat-tailed returns, which characterize actual market behavior with more frequent extreme events. Furthermore, historical data is treated homogeneously, without accounting for temporal importance, leading models to fail during regime...

    arxiv.org/abs/2606.09104 · PDF

  62. 62

    From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning

    Jike Zhong, Yuxiang Lai, Ming Li, Yuheng Li, Wuao Liu, Behzad Dariush, Konstantinos Psounis, Shao-Yuan Lo

    cs.LG

    Theory of Mind (ToM) is a must-acquire skill for modern foundation model systems to operate effectively and safely in the real world. Recent works have explored honing ToM via post-training; however, we show that such progress is confounded by a pervasive "shortcut" issue: tasks can reach up to 99% accuracy by simply exploiting spurious causal correlations, leading to a false sense of ToM. Motivated by this, we first develop a framework to...

    arxiv.org/abs/2606.09092 · PDF

  63. 63

    Stabilizing On-Policy Distillation for MLLM Reasoning with Global Normalization

    Dongze Hao, Zhiwei Jin, Chen Chen, Haonan Lu

    cs.LG · cs.CV

    On-policy distillation (OPD) has recently emerged as an important post-training paradigm. By using a stronger teacher model to provide dense, fine-grained supervision for sampled trajectories, OPD offers a clear advantage over reinforcement learning with verifiable rewards (RLVR), which typically depends on sparse binary or outcome-based environmental feedback. However, naive token-level distillation can suffer from gradient instability, due...

    arxiv.org/abs/2606.09091 · PDF

  64. 64

    Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy

    Haozhe Hu, Hao Wu, Anhao Zhao, Longwei Ding, Peiran Yin, Yunpu Ma, Xiaoyu Shen

    cs.LG · cs.CL

    Pruning has emerged as a dominant paradigm for accelerating large language model (LLM) inference, spanning a broad spectrum of methods that remove computation across tokens, layers, heads, dimensions, and attention patterns. Despite sharing the same objective, these pruning approaches induce fundamentally different execution behaviors, causing realized speedups to depend heavily on hardware and kernel implementations. Consequently, the...

    arxiv.org/abs/2606.09080 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.