cs.LG · 2026-06-18 · No. 27

Machine Learning, 2026-06-18.

71 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

71 entries
  1. 01

    UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning

    Mohamed Nabail, Leo Cheng, Jingmin Wang, Nicholas Rhinehart

    cs.LG · cs.AI · cs.RO

    Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing the need for explicit reward design. However, existing methods typically rely on passive data collection and suffer from poor sample efficiency, especially during the early stages of learning. We introduce a model-based approach that actively directs exploration by jointly reasoning over uncertainties in the reward, dynamics,...

    arxiv.org/abs/2606.19328 · PDF

  2. 02

    Explaining Attention with Program Synthesis

    Amiri Hayes, Belinda Li, Jacob Andreas

    cs.LG · cs.AI

    A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions. In this paper, we propose an approach for approximating the behavior of components of deep networks with executable programs. We focus on attention heads in transformer language models. For a given head, we first compute its associated attention matrices on a collection of randomly selected...

    arxiv.org/abs/2606.19317 · PDF

  3. 03

    Diffusion-Proof: Recipe for Formal Theorem Proving Beyond Auto-Regressive Generation

    Ruida Wang, Rui Pan, Pengcheng Wang, Shizhe Diao, Tong Zhang

    cs.LG

    Enhancing the formal math reasoning capabilities of Large Language Models (LLMs) has become a key focus in both mathematical and computer science communities in recent years. While significant progress has been made in using state-of-the-art Auto-Regressive (AR) LLMs for formal theorem proving, these models suffer from inherent limitations. Their next-token prediction generation methods may yield suboptimal performance due to the challenges...

    arxiv.org/abs/2606.19315 · PDF

  4. 04

    P-K-GCN: Physics-augmented Koopman-enhanced Graph Convolutional Network for Deep Spatiotemporal Super-resolution

    Xizhuo, Zhang, Zekai Wang, Fei Liu, Bing Yao

    cs.LG

    High-fidelity simulation of spatiotemporal dynamics is computationally prohibitive, necessitating efficient super-resolution techniques to reconstruct high-resolution data from coarse-grained inputs. Traditional data-driven methods often lack physical constraints, and simple physics-informed learning struggles with irregular spatial geometries and intricately evolving temporal dynamics. To tackle these challenges, we propose a...

    arxiv.org/abs/2606.19303 · PDF

  5. 05

    Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

    Nikita Kachaev, Andrey Moskalenko, Matvey Skripkin, Nikita Kurlaev, Daria Pugacheva, Albina Burlova, Mikhail...

    cs.LG · cs.RO

    Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on robotics data, yet it is unclear how much commonsense and factual knowledge they retain after adaptation. Failures on knowledge-sensitive tasks are ambiguous, conflating missing knowledge with poor generalization of low-level control. We introduce Act2Answer, a lightweight protocol that adapts VLM knowledge benchmarks to VLA...

    arxiv.org/abs/2606.19297 · PDF

  6. 06

    Risk Stratification for ICU Delirium using Pervasive Ambient Sensing Information

    Jiaqing Zhang, Sabyasachi Bandyopadhyay, Miguel Contreras, Jessica Sena, Yuanfang Ren, Andrea Davidson, Ziyuan Guan,...

    cs.LG

    Delirium is a common and serious complication in the Intensive Care Unit (ICU), associated with increased morbidity, prolonged hospital stays, and higher healthcare costs. Despite its prevalence, early prediction and prevention remain challenging. Environmental factors such as ambient sound and light may influence the onset of delirium, yet they are often overlooked in risk assessments. In this study, we examined whether light intensity and...

    arxiv.org/abs/2606.19292 · PDF

  7. 07

    Structured Inference with Large Language Gibbs

    Sanghyeok Choi, Henry Gouk, Esmeralda S. Whitammer

    cs.LG · cs.CL

    The knowledge encoded in large language models (LLMs) can serve as a substrate for structured reasoning over variables describing a complex world, but accessing this knowledge in a probabilistically coherent manner poses a difficult inference problem. We propose Large Language Gibbs, a scheme for structured probabilistic inference that uses conditional distributions of an LLM as transition operators. Rather than sampling structured objects...

    arxiv.org/abs/2606.19264 · PDF

  8. 08

    Detecting Hidden ML Training With Zero-Overhead Telemetry

    Robi Rahman, Sabiha Tajdari

    cs.LG

    Hardware-enabled monitoring of GPU workloads underpins many proposals for AI compute governance, but if developers can defeat monitoring mechanisms, such schemes are unworkable. We evaluate the adversarial robustness of GPU workload classification using only zero-overhead, privacy-preserving NVML telemetry: content-agnostic signals that observe physical effects of computation without accessing model weights, training data, or hyperparameters....

    arxiv.org/abs/2606.19262 · PDF

  9. 09

    SCAN: Enhance Time Series Anomaly Detection via Multi-Scale Neighborhood-Centered Clustering

    Xingze Zheng, Hanyin Cheng, Siyuan Wang, Yiting Hao, Peng Chen, Yuan Jun, Yang Shu

    cs.LG

    Time series anomaly detection plays a crucial role in a wide range of real-world applications. Reconstruction-based methods have become the mainstream paradigm, but they suffer from over-generalization and under-generalization problems, which are challenging to balance. To address this, we introduce multi-scale clustering to enhance reconstruction-based methods. At the representation level, we integrate the cluster center representations of...

    arxiv.org/abs/2606.19255 · PDF

  10. 10

    STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

    Haipeng Luo, Qingfeng Sun, Songli Wu, Can Xu, Wenfeng Deng, Han Hu, Yansong Tang

    cs.LG · cs.AI · cs.CL

    Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from policy entropy collapse during training. We conduct a first-order gradient analysis of token-level entropy dynamics under GRPO and identify a token-level credit assignment mismatch: the per-token entropy variation decomposes into the product of the trajectory-level...

    arxiv.org/abs/2606.19236 · PDF

  11. 11

    A Human-in-the-Loop Bayesian Optimization Framework for Constraint-Aware Bioprocess Development

    Samuel Stricker, Claus Wirnsperger, Alessandro Butté, Laura Helleckes, Gonzalo Guillén Gosálbez, Antonio del Rio...

    cs.LG · cs.HC · stat.ML

    This work presents an extension to Pareto Front Guided Sampling (PFGS), a Human-in-the-Loop (HitL) Bayesian Optimization (BO) framework in which Gaussian process (GP) surrogate-derived quantities are reformulated as objectives of a multi-objective optimization problem, and the resulting Pareto front is exposed to a domain expert for interactive candidate selection rather than returning a single automated recommendation. The framework is...

    arxiv.org/abs/2606.19230 · PDF

  12. 12

    Mechanism-Guided Selective Unlearning for RLVR-Induced Reasoning

    Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou

    cs.LG · cs.AI

    We propose MAST (Mechanism-Aligned Selective Targeting), a mechanism-guided method for unlearning RLVR-induced reasoning with substantially lower collateral damage than standard full-parameter updates. In matched SFT/RLVR checkpoints on Qwen2.5-Math-1.5B and Qwen3-1.7B-Base, the SFT-to-RLVR increment differs sharply from the SFT update in token-level delta-log-probability, and full-parameter gradient ascent forgets only by damaging retain...

    arxiv.org/abs/2606.19222 · PDF

  13. 13

    Machine Unlearning for the XGBoost Model with Network Intrusion Datasets

    Diana Magalhães, Eva Maia, João Vitorino, Isabel Praça

    cs.LG · cs.AI

    Machine Unlearning (MU) has emerged as an important technique for removing specific data points from trained models without requiring full retraining. However, most existing MU research focuses on deep learning and image data, leaving a gap in the domain of network intrusion detection, which relies heavily on tabular data. This work introduces XGBoost-Forget, an unlearning approach for the XGBoost model, to address this gap. The approach is...

    arxiv.org/abs/2606.19220 · PDF

  14. 14

    Forecasting what Matters: Decision-Focused RL for Controlled EV Charging with Unknown Departure Times

    Giuseppe Gabriele, Fabio Pavirani, Seyed Soroush Karimi Madahi, Chris Develder

    cs.LG · cs.AI

    The recent growth of EV adoption poses challenges for power systems, including increased peak demand and potential grid instability. Smart control of EV charging -- e.g., based on reinforcement learning (RL) -- can alleviate these issues by learning temporal and contextual patterns from historical data. Yet, in real-world scenarios, key features, such as departure time, often are unavailable. This, in turn, makes it harder for an RL agent to...

    arxiv.org/abs/2606.19199 · PDF

  15. 15

    AGDN: Learning to Solve Traveling Salesman Problem with Anisotropic Graph Diffusion Network

    Bolin Shen, Ziwei Huang, Zhiguang Cao, Yushun Dong

    cs.LG

    The Traveling Salesman Problem (TSP) is a cornerstone of combinatorial optimization and arises in many practical scenarios. Although graph-based learning approaches have been explored for TSP, the question of how to exploit graph structure more effectively remains open. We present the Anisotropic Graph Diffusion Network (AGDN), a new Graph Neural Network framework designed to solve TSP. Our method tackles two central difficulties: (1) the...

    arxiv.org/abs/2606.19185 · PDF

  16. 16

    Compute Efficiency and Serial Runtime Tradeoffs for Stochastic Momentum Methods

    Depen Morwani, Alexandru Meterez, Pranav Nair, Sham Kakade

    cs.LG · cs.AI · math.OC · stat.ML

    Stochastic momentum methods such as heavy ball (HB), Nesterov momentum, and variants of Accelerated SGD (ASGD) [Kidambi et al., 2018] are widely used in modern training, but their stochastic benefits depend on two distinct quantities: serial runtime, the number of iterations needed to reach a target accuracy, and compute efficiency (CE), the inverse total gradient-query or FLOP cost. Larger batches reduce serial runtime without hurting CE...

    arxiv.org/abs/2606.19179 · PDF

  17. 17

    Essential Subspace Merging for Multi-Task Learning

    Longhua Li, Lei Qi, Xin Geng, Qi Tian

    cs.LG · cs.AI

    Model merging aims to enable multi-task learning by integrating the capabilities of multiple models fine-tuned from the same pre-trained checkpoint into a single model. Its core challenge is inter-task interference among task-specific parameter updates. In this paper, we analyze the output shifts induced by task updates and observe that their energy is concentrated in a small number of principal directions. We call the subspace spanned by...

    arxiv.org/abs/2606.19164 · PDF

  18. 18

    The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

    Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng, Reyhane Askari-Hemmat, Xiaochuang Han, Adriana...

    cs.LG · cs.CV

    Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering properties such as visual realism and coherent object structure that matching-based training is intended to learn from the data itself. We argue that this reflects a structural mismatch. Matching losses measure $\ell_2$ regression error on the velocity or score field under...

    arxiv.org/abs/2606.19162 · PDF

  19. 19

    Complementary Attention Head Pruning for Efficient Transformers

    Yaniv Livertovsky, Shahar Somin, Gonen Singer

    cs.LG

    The remarkable success of Transformer-based models in natural language processing stems from architectural scaling, which leads to a large number of parameters and hinders deployment in resource-constrained environments. While structured pruning offers a pathway to compression, existing state-of-the-art methods often rely on gradient-based importance ranking or stochastic gating, which suffer from instability, structural degeneration, and the...

    arxiv.org/abs/2606.19150 · PDF

  20. 20

    OrthoReg: Orthogonal Regularization for Hybrid Symbolic-Neural Dynamical Systems

    Till Richter, Niki Kilbertus

    cs.LG · cs.AI · eess.SY

    Dynamical systems are fundamental to modeling the natural world, yet modeling them involves a persistent trade-off: manually prescribed mechanistic models are interpretable by design but often overly simplistic and misspecified; in contrast, flexible data-driven neural methods lack physical insight. Hybrid modeling aims for the best of both worlds by combining a prescribed or symbolic, physics-based component with a flexible neural network. A...

    arxiv.org/abs/2606.19145 · PDF

  21. 21

    ChronoSurv: A Clinical Pathway-Guided Graph Framework for Multimodal Survival Analysis

    Hugo Miccinilli, Theo Di Piazza

    cs.LG

    Accurate survival prediction is essential for personalized treatment planning in head and neck cancer, yet remains challenging due to the heterogeneous and high-dimensional nature of multimodal clinical data. While deep survival models have improved predictive performance over classical statistical approaches, existing methods typically rely on static fusion strategies or temporally agnostic modeling, limiting their ability to capture...

    arxiv.org/abs/2606.19140 · PDF

  22. 22

    INDEQS: Informed Neural controlled Differential EQuationS

    Michael Detzel, Gabriel Nobis, Kristiyan Blagov, Juri Schubert, Jackie Ma, Wojciech Samek

    cs.LG · stat.ML

    Neural Controlled Differential Equations (NCDE) provide a powerful continuous-time framework for forecasting time series, but standard graph-based extensions typically learn spatial structure purely from data, even in settings where a directed graph structure is known a priori. We introduce Informed Neural controlled Differential EQuationS (INDEQS), a graph-based NCDE forecasting method that incorporates prior knowledge of a directed graph at...

    arxiv.org/abs/2606.19138 · PDF

  23. 23

    Pareto Q-Learning with Reward Machines

    Arnaud Lequen, Clément Legrand-Lixon, Léo Saulières

    cs.LG · cs.AI

    We present Pareto Q-Learning with Reward Machines (PQLRM), a multi-objective reinforcement learning algorithm for tasks whose reward structure is specified by a set of reward machines (RMs). PQLRM combines Pareto Q-Learning (PQL), which maintains sets of vector-valued Q-estimates to approximate the Pareto front, with enhancements from Q-Learning with Reward Machines (QRM), which exploits the factored automaton structure of the reward signal....

    arxiv.org/abs/2606.19134 · PDF

  24. 24

    Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

    Sihan Wang, Xiyao Liu, Lianqing Liu, Zhi Han

    cs.LG · cs.CV

    On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned on a reference target. This works well for LLM reasoning, but a direct extension to multimodal large language models (MLLMs) can create a shortcut: the privileged target may guide tokens mainly based on the text reference target rather than the image. We propose ViGOS, a visually grounded OPSD framework...

    arxiv.org/abs/2606.19120 · PDF

  25. 25

    JourneyFormer: Encoding Airbnb Guest Journey with Sequence Modeling

    Daochen Zha, Chun How Tan, Xin Liu, Bin Xu, Han Zhao, Xiaowei Liu, Tracy Yu, Hui Gao, Huiji Gao, Liwei He, Stephanie...

    cs.LG

    Sequence modeling has become increasingly popular in recommendation and ranking algorithms, owing to its capacity to model users' historical behaviors and infer user intentions. Despite its theoretical simplicity, the practical deployment of a sequence model in production is non-trivial due to complexity of the sequence and sparse labels. For example, in Airbnb, guest sequences are often long, exploratory and complex, and we focus on booking...

    arxiv.org/abs/2606.19108 · PDF

  26. 26

    Smoothness-Based Derandomization of PAC-Bayes Bounds

    Alexandre Lemire Paquin, Brahim Chaib-Draa, Philippe Giguère

    cs.LG · stat.ML

    We study PAC-Bayes derandomization for smooth loss functions. Our goal is to obtain generalization bounds that hold with high probability for deterministic predictors by exploiting smoothness properties of both the loss and the predictor class. We show that passing from the Gibbs predictor to the deterministic predictor at the posterior mean has a precise cost, given by the generalization gap of the Jensen gap class. We control this class...

    arxiv.org/abs/2606.19105 · PDF

  27. 27

    Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts

    Tho Tran Huu, Huu-Tuan Nguyen, Thien-Hai Nguyen, Nhat-Tri Ho, Viet-Hoang Tran, Tho Quan, Tan Minh Nguyen

    cs.LG

    Sparse Mixture-of-Experts (SMoE) architectures are now widely deployed in state-of-the-art language and vision models, where conditional routing allows scaling to very large networks. However, this very Top-$k$ expert selection that enables conditional routing also renders the SMoE map inherently discontinuous. In the vicinity of these discontinuity surfaces, even inputs that are arbitrarily close may activate substantially different sets of...

    arxiv.org/abs/2606.19036 · PDF

  28. 28

    A Hybrid LSTM--Vision Transformer Architecture for Predicting HRRR Forecast Errors

    David Aaron Evans, Jay C. Rothenberger, Kara J. Sulia, Nick P. Bassill, Chris D. Thorncroft

    cs.LG · cs.AI · physics.ao-ph

    Forecast errors in high-resolution numerical weather prediction (NWP) systems are often linked to unresolved planetary boundary layer (PBL) processes, convection, terrain-induced circulations, and other vertically structured atmospheric phenomena. Previous work demonstrated that Long Short-Term Memory (LSTM) networks can successfully predict forecast errors in the High-Resolution Rapid Refresh (HRRR) model using mesonet observations, but we...

    arxiv.org/abs/2606.19026 · PDF

  29. 29

    FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs

    Lorenzo Sani, Zeyu Cao, Meghdad Kurmanji, Alex Iacob, Andrej Jovanovic, Yan Gao, Wanru Zhao, Nicholas D. Lane

    cs.LG · cs.AI · cs.DC · eess.SY

    Pre-training Large Language Models (LLMs) typically demands large-scale infrastructure with tightly coupled hardware accelerators. While increasing model and dataset scale remains the dominant driver of performance, Mixture-of-Experts (MoEs) architectures have recently achieved state-of-the-art results by decoupling parameter count from computational cost. This efficiency enables training massive models on constrained compute budgets, yet it...

    arxiv.org/abs/2606.19025 · PDF

  30. 30

    DIPHINE: Diffusion-based $Φ$-ID Neural Estimator

    Simon Pedro Galeano Munoz, Mustapha Bounoua, Giulio Franzese, Pietro Michiardi, Maurizio Filippone

    cs.LG

    Uncovering the true informational architecture of real-world complex systems requires disentangling how their components uniquely store, redundantly share, and synergistically integrate information over time. Integrated Information Decomposition ($Φ$ID) is a framework for decomposing the information dynamics of multivariate systems into sixteen non-overlapping atoms that characterize redundant, unique, and synergistic modes of information...

    arxiv.org/abs/2606.18997 · PDF

  31. 31

    A Controlled Benchmark of Quantum-Latent GAN Augmentation for Brain MRI

    Syed Mujtaba Haider, Silvia Figini

    cs.LG · cs.AI · cs.CV

    Medical image classification is often constrained by limited labeled data, motivating generative augmentation; recently, quantum generative models have been proposed for this purpose, frequently reporting accuracy gains. However, such claims are typically based on single training runs, do not match the parameter budgets of the quantum and classical generators, and do not characterize the data regime in which any benefit appears. We present a...

    arxiv.org/abs/2606.18970 · PDF

  32. 32

    EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts

    Minseo Kim, Minjae Lee, Seunghyuk Oh, Kevin Galim, Donghoon Kim, Coleman Hooper, Harman Singh, Amir Gholami, Hyung...

    cs.LG

    Reinforcement learning (RL) has become a representative post-training paradigm for LLMs, enabling strong reasoning and agentic capabilities. However, rollout generation remains a dominant latency bottleneck because autoregressive sampling decodes responses sequentially and a small number of long-tailed generations often determine completion time. Speculative decoding (SD) offers a natural way to address this bottleneck, as it is a...

    arxiv.org/abs/2606.18967 · PDF

  33. 33

    Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards

    Zirong Li

    cs.LG

    We study online reward-punishment learning when the environment provides no scalar reward or evaluative label. At each step the agent receives only a fixed-channel perceptual packet, and quantities such as pain, energy, contact, damage, or cognitive error are treated as perceptual dimensions whose valence must be inferred from transition consequences. OHIRL separates four roles: M_psi learns next-packet prediction, D_omega models residual...

    arxiv.org/abs/2606.18963 · PDF

  34. 34

    Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization

    Lanqing Li, Shentong Mo, Yang Yu, Pheng-Ann Heng

    cs.LG

    Protein language models (PLMs) have emerged as powerful tools for controllable biomolecular design, yet their post-training adaptation typically relies on costly wet-lab validation or curated preference datasets. To overcome this supervision bottleneck, we introduce unsupervised reward optimization of PLMs, a comprehensive framework for steerable protein generation without ground-truth labels. Our key insight is that task-agnostic rewards,...

    arxiv.org/abs/2606.18961 · PDF

  35. 35

    Zero-Shot Active Feature Acquisition via LLM-Elicitation

    Binyamin Perets, Natalie Mendelson, Shiran Vainberg, Yehuda Chowers, Shai Shen-Orr, Shie Mannor

    cs.LG · cs.IR · stat.ME

    Active feature acquisition (AFA) sequentially selects which features to observe to reach a classification or ranking decision. Its central limitation is reliance on large amount of labeled data to fit probabilistic models guiding acquisition. Large language models (LLMs) supply unsupervised domain knowledge, but are poor sequential planners. Asking one to both know and decide conflates capabilities best kept separate. Here, we develop a...

    arxiv.org/abs/2606.18933 · PDF

  36. 36

    GrapNet: A Programmable Dynamic-Architecture Neural Graph Substrate

    Zirong Li

    cs.LG

    Programmability is a missing first-class interface in fixed-tensor neural networks: editing a relation, freezing a subgraph, auditing a local function, or changing the execution backend should be an operation on the neural program rather than ad-hoc parameter surgery. GrapNet studies this graph-as-network setting. The graph is the architecture and executable program, not an input data graph. Each compute node owns its next-layer child...

    arxiv.org/abs/2606.18923 · PDF

  37. 37

    Some Complexity Results for Robustness Verification for Binarized Neural Networks

    Harshit Goyal, Sudakshina Dutta

    cs.LG · cs.CC

    This paper studies the computational complexity of verification problems for Binarized Neural Networks (BNNs), where activations (and sometimes weights) are binary. We analyze two problems: satisfiability and robustness under uniform image occlusion. We show that BNN satisfiability is NP-complete via a reduction from Boolean satisfiability problem (SAT), and that uniform occlusion induces a piecewise-constant structure in the network output,...

    arxiv.org/abs/2606.18918 · PDF

  38. 38

    REVES: REvision and VErification--Augmented Training for Test-Time Scaling

    Yuanxin Liu, Ruida Zhou, Xinyan Zhao, Amr Sharaf, Hongzhou Lin, Arijit Biswas, Mohammad Ghavamzadeh, Zhaoran Wang,...

    cs.LG · cs.CL

    Test-time scaling via sequential revision has emerged as a powerful paradigm for enhancing Large Language Model (LLM) reasoning. However, standard post-training methods primarily optimize single-shot objectives, creating a fundamental misalignment with multi-step inference dynamics. While recent work treats this as multi-turn reinforcement learning (RL), conventional approaches optimize over the multi-step trajectories directly, failing to...

    arxiv.org/abs/2606.18910 · PDF

  39. 39

    Anomaly Detection for Sparse and Irregular Multivariate Time Series with Latent SDEs

    Martin Uray, Dominik Geng, Florian Graf, Stefan Huber, Roland Kwitt

    cs.LG

    Multivariate time series anomaly detection (MTSAD) is critical for a wide range of application areas, such as industrial monitoring, cybersecurity, or healthcare. Real-world data is often sparse, irregularly sampled or partially observed, yet existing methods assume uniformly sampled time series. We propose a generative approach based on Latent SDEs that projects the observed time series on a continuous-time stochastic dynamical system,...

    arxiv.org/abs/2606.18898 · PDF

  40. 40

    Domain-Shift Aware Neural Networks for Unbalance Characterization in Rotating Systems

    Bernardo Feijó Junqueira, Claudio Kiyoshi Umezu, Bruno Bilhar Karaziack, Tomaz Junior, Daniel Alves Castello

    cs.LG · cs.AI · eess.SP

    This work investigates the application of a domain-shift aware neural network for regression tasks aimed at estimating unbalance masses in rotating shafts under varying operating conditions. Experimental data were collected from a test rig in which a primary shaft, equipped with a flange carrying unbalanced masses, was driven at different rotational speeds, while a secondary shaft could be optionally activated to introduce domain discrepancy....

    arxiv.org/abs/2606.18882 · PDF

  41. 41

    Strategic Feature Selection

    Jivat Neet Kaur, Pratik Patil, Divya Shanmugam, Emma Pierson, Michael I. Jordan, Nika Haghtalab, Meena Jagadeesan,...

    cs.LG · cs.CY · stat.ML

    When algorithmic predictors inform resource allocation in high-stakes domains such as healthcare, these predictors must account for strategic manipulation of input features. The typical solution is to redesign the predictor itself to explicitly account for strategic interactions. In practice, however, decision makers are often constrained to adjusting coarser levers within existing prediction pipelines. For example, healthcare organizations...

    arxiv.org/abs/2606.18867 · PDF

  42. 42

    Scaling Learning-based AEB with Massive Unlabeled Data

    Xiangyu Wang, Yang Zhan, Mengxiang Hao, Chuanchuan Zhong, Yansong Jia, Junjie Zhang, Yu Han, Xin Jiang, Zhen Cao,...

    cs.LG · cs.AI

    This paper studies how to scale learning-based automatic emergency braking (AEB) with massive unlabeled fleet data under production constraints. Our approach is based on meta-feedback semi-supervised learning (MF-SSL), where a teacher generates pseudo labels for unlabeled driving data and is updated using a small labeled anchor set as safety-critical feedback. In production, anchor ambiguity and labeled-unlabeled mismatch can amplify...

    arxiv.org/abs/2606.18864 · PDF

  43. 43

    Investigating Inductive Biases for Machine Learning Emulation of Sudden Stratospheric Warmings in Idealised Isca Simulations

    Oskar Bohn Lassen, Simon Driscoll, Stephen I. Thomson, Sebastian Schemm, Francisco C. Pereira

    cs.LG · physics.ao-ph

    Machine-learning emulators are increasingly used for weather prediction and have the potential to extend skill on subseasonal-to-seasonal timescales by learning dynamically important sources of predictability. A key challenge is whether the models can exploit predictability anchors, such as stratospheric variability, that influence tropospheric circulation beyond short lead times. We test how architectural inductive bias affects emulation of...

    arxiv.org/abs/2606.18857 · PDF

  44. 44

    Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation

    Zhilin Huang, Hang Gao, Ziqiang Dong, Yuan Chen, Yifeng Luo, Chujun Qin, Jingyi Wang, Yang Yang, Guanjun Jiang

    cs.LG

    Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes KL divergence toward a privileged target distribution. However, because this supervision is generated via uncontrolled sampling, it provides no diagnostic insight into the model's specific errors or corrective guidance for its individual failure patterns....

    arxiv.org/abs/2606.18844 · PDF

  45. 45

    Semantic Robustness Certification for Vision-Language Models

    Peiyu Yang, Paul Montague, Feng Liu, Andrew C. Cullen, Amardeep Kaur, Christopher Leckie, Sarah M. Erfani

    cs.LG · cs.CV

    Vision-language models (VLMs) are now widely used in downstream tasks. However, real-world applications often expose VLMs to distribution shifts induced by semantic variation (e.g., shape, size, and style). Robustness certification determines if a model's prediction changes when transformations are applied to its input. While most certification frameworks study geometric or pixel-level transformations over inputs, this work proposes a novel...

    arxiv.org/abs/2606.18839 · PDF

  46. 46

    Identifying Structural Biases from Causal Mechanism Shifts

    Praharsh Nanavati, Jilles Vreeken, David Kaltenpoth

    cs.LG

    Causal discovery methods commonly assume that all data is independently and identically distributed (i.i.d.) and that there are no unmeasured variables affecting the system. In practice, these assumptions are often violated, leading to inaccurate inference. In this paper, we study how to identify hidden confounding and selection biases from causal mechanism shifts. In particular, we show that structural biases lead to dependent mechanism...

    arxiv.org/abs/2606.18834 · PDF

  47. 47

    Seed-Guided Semi-Supervised Clustering by A-Contrario Anomaly Detection

    Nassir Mohammad

    cs.LG

    This paper introduces a semi-supervised clustering framework grounded in the statistical duality between grouping principles and anomaly detection. We address the challenge of robust cluster definition in noisy environments -- a task where partitioning algorithms often over-assign outliers and density-based methods remain sensitive to heuristic global parameters. Drawing on \textit{a-contrario} statistical reasoning and Gestalt proximity...

    arxiv.org/abs/2606.18833 · PDF

  48. 48

    Target-confidence Recourse Using tSeTlin machines: TRUST

    K. Darshana Abeyrathna, Sara El Mekkaoui, Nils Enric Canut Taugbøl, Anuja Vats

    cs.LG · cs.AI

    Counterfactual explanations are widely used to provide algorithmic recourse in high-stakes decision-making systems. Most existing methods seek the smallest change to an input that flips a model's decision. However, decision-makers often rely not only on predicted labels but also on confidence thresholds and risk margins. Counterfactuals that barely cross a decision boundary can be fragile and unstable under noise or model variation. In this...

    arxiv.org/abs/2606.18832 · PDF

  49. 49

    GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents

    Zhe Ren, Yibo Yang, Yimeng Chen, Zijun Zhao, Benshuo Fu, Zhihao Shu, Bingjie Zhang, Yangyang Xu, Dandan Guo, Shuicheng Yan

    cs.LG · cs.CL

    Memory benchmarks for LLM agents largely assume single-user settings, leaving shared assistants for hospitals, workplaces, campuses, and households understudied. In these deployments, multiple principals write to a common memory pool and query it under different roles, scopes, and relationships, so memory quality requires governance as well as recall. We introduce GateMem, a benchmark for multi-principal shared-memory agents. GateMem jointly...

    arxiv.org/abs/2606.18829 · PDF

  50. 50

    Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets

    Jiaxi Liu, Aiping Yang, Yuhang Yang, Shuqi Zhang, Zewei Dong, Jiangming Yang, Xuebin Chen

    cs.LG · cs.AI

    Sequential decision problems often exhibit an asymmetric evolution of information and decision flexibility: as a decision cycle unfolds, the agent receives richer information while feasible actions expire due to operational cutoffs, commitments, or resource constraints. Standard MDP formulations typically flatten this structure into stage-dependent state descriptions and action masks, thereby obscuring the nested information--action asymmetry...

    arxiv.org/abs/2606.18820 · PDF

  51. 51

    Reinforcement Learning Foundation Models Should Already Be A Thing

    Abdelrahman Zighem, Jill-Jênn Vie

    cs.LG · cs.AI

    Foundation models for language and vision are powered by internet-scale data, while structured domains (tabular prediction, time-series forecasting, graph learning, reinforcement learning) are not. The substitute is synthetic data, which shifts the burden from collection to prior design. Such priors already exist for many structured tasks: TabPFN and its successors solve tabular classification with a transformer pretrained on a synthetic...

    arxiv.org/abs/2606.18812 · PDF

  52. 52

    Learning from Own Solutions: Self-Conditioned Credit Assignment for Reinforcement Learning with Verifiable Rewards

    Yingyu Shan, Yuhang Guo, Zihao Cheng, Zeming Liu, Xiangrong Zhu, Xinyi Wang, Jiashu Yao, Wei Lin, Hongru Wang, Heyan Huang

    cs.LG · cs.AI

    Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in training LLMs for reasoning tasks, but representative methods such as GRPO assign uniform credit across all tokens, wasting gradient on routine tokens while under-crediting pivotal reasoning steps. Existing token-level credit assignment methods require resources beyond the model's own rollouts. GRPO variants rely on process reward models or ground-truth...

    arxiv.org/abs/2606.18810 · PDF

  53. 53

    Bayesian Anytime Pareto Set Identification for Multi-Objective Multi-Armed Bandits

    Lennert Saerens, Bram Silue, Eleni Litsa, Peter Vrancx, Pieter Libin

    cs.LG · cs.AI

    Identifying Pareto optimal solutions is critical to support multi-objective decision-making. We introduce the first anytime Multi-Objective Multi-Armed Bandit algorithm for the Pareto Set Identification problem, taking a Bayesian approach: Top-Two Pareto Front Thompson Sampling (TTPFTS). We benchmark TTPFTS against state-of-the-art fixed-budget Pareto Set Identification algorithms on synthetic environments. Next, we demonstrate its practical...

    arxiv.org/abs/2606.18785 · PDF

  54. 54

    Online Distributional Prediction via Latent Cluster Geometry Under Drift and Corruption

    Navyansh Mahla, Prateek Chanda, Ganesh Ramakrishnan

    cs.LG · stat.ML

    Online learning in non-stationary streams is often formulated as tracking a point estimate, but many applications require predicting the full data-generating distribution. We study online distributional prediction under drift and adversarial corruption. Our approach represents each candidate law through a latent cluster geometry: a variable-size configuration of centers that organizes probability mass and induces a predictive distribution. A...

    arxiv.org/abs/2606.18778 · PDF

  55. 55

    RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing

    Guannan Lai, Haoran Hu, Han-Jia Ye

    cs.LG

    We present RouteJudge, an online pairwise preference evaluation framework for LLM routing systems, with a public platform available at https://routejudge.cn. Different from model-level response evaluation, RouteJudge focuses on router-level decision quality. For each user query, multiple routing strategies independently recommend candidate models under the same model pool and budget constraints. The selected model responses are then presented...

    arxiv.org/abs/2606.18774 · PDF

  56. 56

    Private Learning with Public Feature Conditioning

    Shuli Jiang, Walid Krichene, Nicolas Mayoraz

    cs.LG · cs.AI

    We study differentially private (DP) regression in settings where each data sample includes public, non-sensitive features -- common in applications such as recommendation and advertising systems. While such label-DP or semi-sensitive-feature settings have been primarily explored in the context of classification, effective approaches for regression remain underexplored. We introduce Cond-DP, a conditioned variant of DPSGD that leverages the...

    arxiv.org/abs/2606.18773 · PDF

  57. 57

    Low-Cost Neuromorphic Fall Detection Using Synthetic Event Data and Hybrid SNNs

    Guillermo Rojas, Gonzalo Soto, Daniel Yunge

    cs.LG · cs.CV

    This work presents the development of hybrid models that integrate spiking neural networks (SNNs) with components of convolutional neural networks (CNNs) to learn from simulated event-based camera data (Dynamic Vision Sensor, DVS) generated from conventional smartphone videos. Aimed primarily at human fall detection, the approach leverages the energy efficiency and spatio-temporal processing capabilities of SNNs by converting video frames...

    arxiv.org/abs/2606.18732 · PDF

  58. 58

    Graph Grounded Cross Attention Transformer Neural Network for Structurally Constrained Full Event Sequence Generation in Predictive Process Monitoring

    Fang Wang, Ernesto Damiani

    cs.LG · cs.AI

    Structurally constrained event sequence generation remains challenging because generated paths must preserve transition feasibility, temporal order, termination, and attribute consistency. In predictive process monitoring (PPM), this challenge appears as full event sequence generation, whereas existing work mainly addresses component tasks such as next activity, remaining time, outcome, and attribute prediction. This paper proposes the Graph...

    arxiv.org/abs/2606.18726 · PDF

  59. 59

    Trainable Photonic Measurement for Physics-Informed PDE Learning

    Jiale Linghu, Hao Dong, Yangshuai Wang

    cs.LG · physics.comp-ph

    Photonic quantum machine learning offers a route to trainable physical representations built from phase, interference and measurement. However, its role in scientific machine learning remains largely unexplored. Physics-informed neural fields provide a natural setting, because differential equations require trial spaces that preserve phase, frequency and derivative structure. Here we introduce a photonic quantum neural field in which...

    arxiv.org/abs/2606.18713 · PDF

  60. 60

    Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment

    Yanjun Shao, Yundi Chen, Yashvi Patel, Aurelien Pelissier, María Rodríguez Martínez

    cs.LG · q-bio.QM

    Pretrained biological language models expose per-token probability distributions through masked-token prediction, providing the likelihood interface central to sequence design, variant scoring, and mechanistic interpretation. Yet these distributions are learned from broad unlabeled corpora and are not naturally conditioned on task-specific biological contexts such as interaction partners, cellular environments, or therapeutic interventions....

    arxiv.org/abs/2606.18703 · PDF

  61. 61

    Stealthy World Model Manipulation via Data Poisoning

    Yibin Hu, Xiaolin Sun, Zizhan Zheng

    cs.LG · cs.CR · cs.RO

    Model-based learning agents use learned world models to predict future states, plan actions, and adapt to new environments. However, the process of updating world models from collected experience creates a training-time attack surface: adversarially poisoned fine-tuning trajectories can manipulate the learned dynamics and thereby corrupt downstream planning. In this paper, we propose SWAAP, the first two-stage data poisoning framework for...

    arxiv.org/abs/2606.18697 · PDF

  62. 62

    Attention as Frustrated Synchronization

    Joshua Nunley

    cs.LG · cond-mat.dis-nn · cs.CL · cs.NE · nlin.AO

    A network of oscillators that synchronizes perfectly computes nothing further, so an attention architecture built from synchronization must locate its computation in structured departures from agreement. We introduce the Frustrated Synchronization Network (FSN), whose token states are phases on a torus and whose entire value pathway is one learned complex coupling kernel over harmonics and a one-step delay. Each component of the kernel is a...

    arxiv.org/abs/2606.18694 · PDF

  63. 63

    Robust and Interpretable Adaptation of Equivariant Materials Foundation Models via Sparsity-promoting Fine-tuning

    Youngwoo Cho, Seunghoon Yi, Wooil Yang, Sungmo Kang, Young-woo Son, Jaegul Choo, Joonseok Lee, Soo Kyung Kim, Hongkee Yoon

    cs.LG · cond-mat.mtrl-sci

    Pre-trained materials foundation models, or machine learning interatomic potentials, leverage general physicochemical knowledge to effectively approximate potential energy surfaces. However, they often require domain-specific calibration due to physicochemical diversity as well as mismatches between practical computational settings and those used in constructing the pre-training data. To address this, we propose a sparsity-promoting...

    arxiv.org/abs/2606.18691 · PDF

  64. 64

    Dual-Channel Grounded World Modeling (DCGWM): Structural Prevention of Objective Interference Collapse via Heterogeneous External Grounding with Inward-Only Gradient Flow

    Akshay Hazare

    cs.LG · cs.AI

    Joint Embedding Predictive Architectures (JEPAs) are a leading approach to world model representation learning. We identify a failure mode in JEPA-based world models grounded against two qualitatively distinct external signals: physical dynamics (sparse, high-magnitude, constraint-satisfying gradient corrections) and social-behavioral dynamics (diffuse, distribution-matching corrections). We term this Objective Interference Collapse (OIC): we...

    arxiv.org/abs/2606.18688 · PDF

  65. 65

    Bounded Context Management for Tabular Foundation Models on Stream Learning

    Jinmo Lee, Doyun Choi, Moongi Choi, Jaemin Yoo

    cs.LG · cs.AI

    Tabular stream learning requires predictions on sequentially arriving examples under distribution shift. While standard methods adapt by updating model states, tabular foundation models (TFMs) make predictions conditioned on a labeled context in an in-context manner, making them a natural alternative for stream learning. This shifts the challenge from how to update the model to how to manage the context. We propose a future information view...

    arxiv.org/abs/2606.18677 · PDF

  66. 66

    InTrain: Intrinsic Trainability for Zero-Cost Neural Architecture Search

    Qinqin Zhou, Fuhai Chen, Jipeng Wu, Zhiwei Chen, Zhikai Hu, Weiwei Cai

    cs.LG · cs.CV

    Training-free neural architecture search promises efficient discovery of high-performance networks without costly training. However, existing zero-cost proxies rely on fragmented heuristics that fail to capture the fundamental question: what makes an architecture trainable? This paper introduces Intrinsic Trainability (InTrain), a unified theoretical proxy that formalizes trainability as an architectural invariant emerging from two...

    arxiv.org/abs/2606.18676 · PDF

  67. 67

    scGTN: Deep Siamese Graph Transformer Network for Single-cell RNA Sequencing Clustering

    Jinke Wu, Yifan Wang, Siyu Yi, Caiyang Yu, Ziyue Qiao, Nan Yin, Jiancheng Lv, Wei Ju

    cs.LG · cs.AI · q-bio.GN

    Single-cell RNA sequencing (scRNA-seq) serves a pivotal role in characterizing gene expression at the cellular level, enabling the identification of cell types and advancing the understanding of cellular heterogeneity. Despite the significant progress in scRNA-seq data clustering, we argue that current methods always ignore the sparsity and noise, as well as the complex intercellular structural information inherent in scRNA-seq data. Toward...

    arxiv.org/abs/2606.18672 · PDF

  68. 68

    BLADE: Scalable Bi-level Adaptive Data Selection for LLM Training

    Jiaxing Wang, Deping Xiang, Jin Xu, Zirui Liu, Zicheng Zhang, Guoqiang Gong, Jun Fang, Chao Liu, Pengzhang Liu,...

    cs.LG

    As Large Language Model (LLM) datasets scale to trillions of tokens, data selection has emerged as a critical frontier to filter out uninformative noise and construct adaptive learning trajectories. Beyond static heuristic filtering, advanced data selection methods for LLM training largely follow two paradigms, each with fundamental limitations. Influence-based methods provide principled bi-level objectives but require intractable...

    arxiv.org/abs/2606.18650 · PDF

  69. 69

    MetaboNet-Bench: A Multi-modal Benchmark for Glucose Forecasting in Type 1 Diabetes

    Nathaniel Jeffries, Miriam Wolff, Sam Royston, Elizabeth Healey, Caleb Mayer, David Klonoff, Michael Snyder, Tao Wang

    cs.LG · q-bio.QM

    Glucose forecasting algorithms are an important aspect of glycemic control management in type 1 diabetes. So far, the research community has developed numerous algorithms and models for forecasting. However, it is well-recognized that the lack of standardized model performance evaluation benchmarks makes fair comparison difficult and hinders further innovation, and thus benchmark standardization is in urgent need. Furthermore, many published...

    arxiv.org/abs/2606.18640 · PDF

  70. 70

    PACT: Preserving Anchored Cores in Task-vectors for Model Merging

    Ningyuan Shi, Zhipeng Zhou, Hao Wang, Chunyan Miao, Peilin Zhao

    cs.LG

    Model merging has emerged as a training-free alternative to multi-task learning, aiming to combine multiple task-specific fine-tuned models into a single multi-task model. Most existing model merging approaches follow the Task Arithmetic paradigm, which decomposes fine-tuned weights into pre-trained parameters and task vectors, and performs merging exclusively in the task-vector space. The effectiveness of this paradigm implicitly relies on...

    arxiv.org/abs/2606.18627 · PDF

  71. 71

    Towards Anomaly Detection on Relational Data

    Shiyuan Li, Yunfeng Zhao, Yue Tan, Qingfeng Chen, Yixin Liu, Shirui Pan

    cs.LG

    Relational databases are widely used for managing structured data in real-world systems. Detecting anomalies from such relational data is crucial for identifying fraud, risks, and abnormal behaviors, yet remains under-explored. The key challenges lie in the intrinsic complexity of relational data: multi-table attributes are high-dimensional and heterogeneous, making sparse abnormal clues easy to overwhelm by normal or irrelevant information;...

    arxiv.org/abs/2606.18621 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.