cs.LG · 2026-05-28 · No. 11

Machine Learning, 2026-05-28.

65 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

65 entries
  1. 01

    PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective

    Yangyi Huang, Ruotian Peng, Zeju Qiu, Jiale Kang, Yandong Wen, Bernhard Schölkopf, Weiyang Liu

    cs.LG · cs.CL

    Parameter-efficient finetuning (PEFT) has become the standard approach for adapting large language models, yet evaluations largely emphasize downstream accuracy while overlooking the retention of pretrained capabilities. We argue that PEFT should be assessed through the stability-plasticity dilemma: the trade-off between target-task adaptation and resistance to forgetting. We introduce PEFT-Arena, a benchmark that jointly measures downstream...

    arxiv.org/abs/2605.28819 · PDF

  2. 02

    Affective Music Recommendation: A Rollout-Based World Model for Offline Preference Optimization

    Audrey Chan, Aaron Labbé, Jacob Lavoie, Jordan Bannister, Arsène Fansi Tchango, Guillaume Lajoie, Laurent Charlin

    cs.LG · cs.IR · cs.SD

    Functional music applications, from consumer focus and sleep aids to clinical interventions, share a distinctive recommendation problem: success is defined by the listener's affective state, but online experimentation on emotion is ethically constrained, particularly for clinical populations who cannot reliably skip a song or report distress. We describe AMRS, the Affective Music Recommendation System deployed on LUCID's health-and-wellness...

    arxiv.org/abs/2605.28810 · PDF

  3. 03

    Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents

    Suji Kim, Kangsan Kim, Sung Ju Hwang

    cs.LG · cs.AI · cs.CL

    Computer-use agents (CUAs) have recently made substantial progress, but deploying a separate large expert for each software domain remains expensive. Small open computer-use agents are more practical specialization targets, but they remain substantially weaker and exhibit uneven domain-specific failures. A straightforward remedy is to synthesize large-scale training data for the target domain, yet we find that this naive approach yields only...

    arxiv.org/abs/2605.28775 · PDF

  4. 04

    Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations

    Kevin Y. Li, Asher Trockman, Ananda Theertha Suresh, Ziteng Sun

    cs.LG

    Softmax attention is the cornerstone of modern large language models, but its memory scales linearly and compute quadratically with sequence length. Linear recurrent models, such as linear attention and state space models, have become widely studied as alternatives to attention due to their linear compute and constant memory. While these sub-quadratic token mixing methods, or mixers, achieve promising efficiency gains and competitive results...

    arxiv.org/abs/2605.28769 · PDF

  5. 05

    Principled Algorithms for Optimizing Generalized Metrics in Multi-Label Learning

    Mehryar Mohri, Yutao Zhong

    cs.LG · stat.ML

    Many real-world classification tasks require predicting multiple labels per instance, necessitating the optimization of complex evaluation metrics such as the $F$-measure and Jaccard index. While the Empirical Utility Maximization (EUM) framework is natural for these population-level metrics, existing theoretical results are largely limited to asymptotic Bayes-consistency. In this paper, we develop principled learning algorithms for...

    arxiv.org/abs/2605.28767 · PDF

  6. 06

    LLM Zeroth-Order Fine-Tuning is an Inference Workload

    Zelin Li, Caiwen Ding

    cs.LG

    Zeroth-order (ZO) fine-tuning is attractive for large language models because it replaces backpropagation with forward objective evaluations. Existing implementations nevertheless execute ZO algorithms inside conventional training loops, even though their dominant work is repeated scoring under nearby parameter states. This creates a workload-runtime mismatch: the algorithm asks for structured inference-style scoring, while the system exposes...

    arxiv.org/abs/2605.28760 · PDF

  7. 07

    Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL

    Kunhao Zheng, Pierre Chambon, Juliette Decugis, Jonas Gehring, Taco Cohen, Benjamin Negrevergne, Gabriel Synnaeve

    cs.LG · cs.AI · cs.CL

    Linear interpolation between fine-tuned checkpoints has been shown to trace the Pareto front between competing objectives, but whether extrapolative weight averaging can extend such frontiers to new checkpoints useful at inference time, without additional RL training, remains unclear. We study this question in RL for competitive programming, where hidden unit tests under time and memory limits enforce both functional correctness and...

    arxiv.org/abs/2605.28751 · PDF

  8. 08

    BIRDNet: Mining and Encoding Boolean Implication Knowledge Graphs as Interpretable Deep Neural Networks

    Tirtharaj Dash

    cs.LG · cs.AI · cs.NE · q-bio.QM

    Tabular data in knowledge-rich domains often carries a latent prior in the form of Boolean implication relationships (BIRs) between pairs of features. We mine such relationships with a sparse-exception binomial test. The mined implications form a typed directed graph, equivalent to a propositional rule base of 2-literal clauses. We encode this graph as the connectivity of a layered neural network, called BIRDNet, in which each hidden unit...

    arxiv.org/abs/2605.28739 · PDF

  9. 09

    Stage-wise Distortion-Perception Traversal in Zero-shot Inverse Problems with Diffusion Models

    Jiawei Zhang, Ziyuan Liu, Leon Yan, Zhenyu Xiao, Yuantao Gu

    cs.LG

    The distortion-perception (D-P) tradeoff is a fundamental phenomenon of Bayesian inverse problems, which characterizes the inherent tension between distortion performance and perceptual quality. Enabling flexible traversal of the D-P tradeoff at inference time is crucial for practical applications. Despite the recent success of diffusion models in zero-shot inverse problem solving, efficient and principled strategies for D-P traversal in...

    arxiv.org/abs/2605.28711 · PDF

  10. 10

    Understanding Generalization and Forgetting in In-Context Continual Learning

    Guangyu Li, Meng Ding, Lijie Hu

    cs.LG

    In-context learning (ICL) derives its power from enabling Large Language Models to adapt to new tasks via prompt-based reasoning alone, entirely bypassing the need for parameter updates. Existing theories primarily study ICL in single-task settings, while real-world prompts often contain sequences of heterogeneous tasks, leaving a gap in understanding whether Large Language Models implicitly perform continual learning during inference. To...

    arxiv.org/abs/2605.28705 · PDF

  11. 11

    Expressive Power of Floating-Point Neural Networks with Arbitrary Reduction Orders and Inexact Activation Implementations

    Yeachan Park, Geonho Hwang, Wonyeol Lee, Sejun Park

    cs.LG

    Most existing expressivity theories for neural networks assume exact real arithmetic, whereas practical neural networks are executed under finite-precision floating-point arithmetic with implementation-dependent execution semantics. Recent works have begun studying the expressive power of floating-point neural networks, but existing results are limited to highly restricted activation functions and idealized assumptions such as fixed...

    arxiv.org/abs/2605.28704 · PDF

  12. 12

    History-aware adaptive reduced-order models via incremental singular value decomposition

    Amirpasha Hedayat, Ali Mohaghegh, Laura Balzano, Cheng Huang, Karthik Duraisamy

    cs.LG · cs.CE · math.NA · physics.comp-ph

    Reduced-order models (ROMs) can accelerate high-dimensional dynamical simulations, but their accuracy often deteriorates when online dynamics leave the regime represented by offline training data. We develop a projection-based adaptive ROM framework based on incremental singular value decomposition (iSVD), in which occasional full-order operator evaluations provide correction snapshots for online basis updates. The intrusive ROMs considered...

    arxiv.org/abs/2605.28684 · PDF

  13. 13

    Optimal ridge regularization revisited

    Jack Timmermans, Sergio A. Alvarez

    cs.LG · stat.ML

    We consider $L^2$-regularized linear (ridge) regression over a finite data sample $X$ with bounded covariance and linear prediction targets $y$ with additive isotropic noise of finite variance. We present an iterative procedure to compute the optimal regularization strength numerically from the generative parameters in the fixed-$X$ setting and prove its convergence at limited noise levels. Our experimental evaluation over synthetic data...

    arxiv.org/abs/2605.28679 · PDF

  14. 14

    Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective

    Mingjie Hu, Jian-Qiang Hu, Enlu Zhou

    cs.LG

    Data acquisition efficiency is a central challenge in deploying reinforcement learning in business and healthcare operations, where interactions are costly, slow, and often involve humans in the loop. This paper develops a unified large deviations framework for data acquisition in infinite-horizon reinforcement learning. We introduce the exponential decay rate of the policy-selection error probability as a principled efficiency metric and...

    arxiv.org/abs/2605.28675 · PDF

  15. 15

    Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection

    Vijeta Deshpande, Tootiya Giyahchi, Veena Padmanabhan, Leman Akoglu, Anna Rumshisky

    cs.LG · cs.CL

    Safety detection models require examples of HHH (Helpful, Harmless, Honest)-violating outputs for robust generalization, however such examples are scarce. Activation Steering (AS) has emerged as a data-efficient method for generating target-concept-aligned responses. We investigate whether AS can generate high-quality training datasets for downstream classifiers, a question that remains untested. We present a two-fold study with intrinsic and...

    arxiv.org/abs/2605.28664 · PDF

  16. 16

    Applications of temporal graph learning for predicting the dynamics of biological systems

    Manuel Dileo, Andrea Sottoriva

    cs.LG

    Biological foundation models have shown strong performance in single-cell representation learning by applying transformer architectures directly to gene-expression matrices. However, these approaches predominantly operate in static settings and do not explicitly model the temporal evolution of developmental programs in the cell. Modeling such dynamics is important for understanding how cellular states progressively emerge, differentiate, and...

    arxiv.org/abs/2605.28659 · PDF

  17. 17

    Interpretability-Guided Layer Selection over Subspace Projection: SAEs as Stethoscopes, Not Scalpels, for Raw Task Vector Model Editing

    Li Lei, Madalina Ciobanu, Qingqing Mao, Ritankar Das

    cs.LG · cs.CL

    LLMs increasingly require surgical model editing to enhance domain-specific capabilities without incurring the computational cost or catastrophic forgetting associated with full fine-tuning. Sparse Autoencoders (SAEs) have emerged as a promising tool in this setting, in principle allowing for feature-level identification of where to intervene. In this work, we rigorously evaluate an SAE-guided editing pipeline for mathematical reasoning on...

    arxiv.org/abs/2605.28649 · PDF

  18. 18

    Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity

    Xiuying Wei, Caglar Gulcehre

    cs.LG

    Efficient inference is critical for long-context language models, where attention computation and KV-cache access dominate the cost. Recent work RAT+, introduces a recurrence-augmented attention backbone that enables flexible dilated attention at inference time. In this paper, we investigate whether this exponentially decaying memory can also improve existing query-aware sparse inference methods. Using representative methods including Quest,...

    arxiv.org/abs/2605.28640 · PDF

  19. 19

    Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection

    Jianghao Wu, Jianfei Cai, Weiqiang Wang, Jin Ye, Daniel F. Schmidt, Yasmeen George

    cs.LG

    Reinforcement learning with verifiable rewards (RLVR) can yield large reasoning gains from very few training instances, yet its strong sensitivity to which instances are used makes data selection a central bottleneck. Most existing selection pipelines rely on training-time optimization signals and/or require access to verifiable rewards or ground-truth answers over large candidate pools, which is costly and often infeasible in specialized...

    arxiv.org/abs/2605.28631 · PDF

  20. 20

    When Interpretability Is Unequally Distributed: Fairness in Hybrid Interpretable Models

    Ziba Jabbar Zare, Ulrich Aïvodji, Julien Ferry, Thibaut Vidal

    cs.LG

    Hybrid interpretable models combine a transparent component with a black-box model by assigning some examples to the former and deferring the rest to the latter. While this design enables flexible tradeoffs between accuracy and interpretability, it also raises a distinct procedural fairness concern: some demographic groups may systematically receive interpretable decisions, while others are disproportionately routed to a black box. We...

    arxiv.org/abs/2605.28626 · PDF

  21. 21

    Random Process Flow Matching: Generative Implicit Representations of Multivariate Random Fields

    Julien Lalanne, David Picard, Lionel Boillot, Lina-María Guayacán-Carrillo, Leon Barens, Jean-Michel Pereira

    cs.LG

    Generative modeling provides a powerful framework for learning data distributions. These models initially relied on probabilistic methods such as Gaussian Processes (GP) for uncertainty-aware predictions and shifted towards larger trainable models to learn more complex distributions. In this work, we introduce Random Process (RP) Flow, a Flow Matching-based framework that represents the vector field as a neural implicit function. Unlike...

    arxiv.org/abs/2605.28625 · PDF

  22. 22

    Learning High-Dimensional Parity Functions with Product Networks using Gradient Descent

    Guillaume Larue, Louis-Adrien Dufrène, Quentin Lampin, Hadi Ghauch, Ghaya Rekaya

    cs.LG

    Parity functions are fundamental Boolean operations with critical applications across machine learning, cryptography, and error correction. Yet, learning high-dimensional parity functions poses significant challenges: in a general setting, standard neural network architectures typically require exponential sample complexity, making gradient-based optimization intractable for large number of inputs $N$. We demonstrate that compact...

    arxiv.org/abs/2605.28612 · PDF

  23. 23

    Online Irregular Multivariate Time Series Forecasting via Uncertainty-Driven Dual-Expert Calibration

    Haonan Wen, Hanyang Chen, Songhe Feng

    cs.LG · cs.AI

    Irregular multivariate time series forecasting is critical in many real-world applications, where time series are irregularly sampled and exhibit dynamically evolving missingness patterns. Although existing methods perform well in offline settings, they often suffer from significant performance degradation when deployed online due to dynamic shifts in data distribution. Maintaining forecasting capability in such dynamic scenarios typically...

    arxiv.org/abs/2605.28603 · PDF

  24. 24

    Transformers Provably Learn to Internalize Chain-of-Thought

    Yixiao Huang, Hanlin Zhu, Zixuan Wang, Jiantao Jiao, Stuart Russell, Somayeh Sojoudi, Song Mei

    cs.LG

    Chain-of-Thought (CoT) prompting substantially improves the sample efficiency of transformers, reducing the complexity of tasks like parity learning from exponential to polynomial in the input length. However, generating explicit reasoning steps at inference is computationally expensive. Implicit Chain-of-Thought (ICoT) has emerged as a promising empirical remedy that trains models to internalize intermediate steps within their hidden states,...

    arxiv.org/abs/2605.28600 · PDF

  25. 25

    PLS in the Mirror of Self-Attention

    Jiangsheng, You

    cs.LG

    This note provides an interesting observation on casting partial least square (PLS) as a linearized self-attention so that PLS may be studied within the neural network paradigm. On the other hand, the dimensionality reduction and selection of predictors in PLS may indicate that self-attention includes certain degree of dimensionality normalization toward improved learning.

    arxiv.org/abs/2605.28592 · PDF

  26. 26

    Thinned Mean Field Langevin Dynamics

    Zonghao Chen, Heishiro Kanagawa, François-Xavier Briol, Chris J. Oates, Lester Mackey

    cs.LG

    Several important learning tasks can be formulated as minimizing an entropy-regularized objective over an appropriate space of probability distributions. Mean-field Langevin dynamics (MFLD) facilitate computation in this general context, casting the minimizer as the invariant distribution of a McKean--Vlasov process, which can be numerically discretized using $N$ particles and thus simulated. However, simulating this interacting particle...

    arxiv.org/abs/2605.28589 · PDF

  27. 27

    Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization

    Kristi Topollai, Allan Ma, Tolga Dimlioglu, Sui Jiet Tay, Anna Choromanska

    cs.LG

    Communication-efficient distributed optimizers such as DiLoCo reduce synchronization costs by letting workers perform many local updates before aggregating their progress with an outer momentum optimizer. Recent theory suggests that the outer optimizer acts on an effective spectrum induced by the inner optimization loop, and that the choice of outer momentum controls how progress from local updates is accumulated across communication rounds....

    arxiv.org/abs/2605.28585 · PDF

  28. 28

    A Generalized Tikhonov Layer for Interpretable-by-design Graph Neural Networks

    Nicolas Tremblay, Benjamin Ricaud, Filippo Maria Bianchi

    cs.LG

    We propose the Tikhonov layer, a graph neural network layer that is interpretable by design: once trained, its learned parameters directly reveal which node features and which aspects of the graph topology were leveraged for prediction. In practice, the layer's propagation matrix takes the closed-form $R = (p(L)+Q)^{-1} Q$, where $L$ is the normalized graph Laplacian, $Q = diag(q_1,...,q_n)$ a learnable diagonal matrix of positive...

    arxiv.org/abs/2605.28578 · PDF

  29. 29

    Efficient Pre-Training of LLMs through Truncated SVD Layers

    Kaivan Kamali, Kajetan Schweighofer, Hormoz Shahrzad, Olivier Francon, Babak Hodjat, Risto Miikkulainen

    cs.LG · cs.AI

    The massive scaling of Large Language Models (LLMs) has made pretraining increasingly cost-prohibitive. While low-rank representation and orthonormal weight matrices could in principle reduce parameter counts and computational overhead, most existing methods rely on static rank selection and do not enforce weight orthonormality due to high computational cost. This paper introduces TSVD, a framework that maintains low rank and strict...

    arxiv.org/abs/2605.28573 · PDF

  30. 30

    Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression

    Tue M. Cao, Nguyen Do, My T. Thai

    cs.LG · cs.AI

    Sparse autoencoders (SAEs) have become a central tool for interpreting language models. However, two key SAE analyses that remain difficult to scale are (1) matching semantically similar features across multi-layers and (2) compressing large feature circuits into interpretable supernodes. Although these have been treated as separate problems, we show that both are instances of a more fundamental challenge, which we frame as the estimation of...

    arxiv.org/abs/2605.28567 · PDF

  31. 31

    A Multi-dimensional Framework for Evaluating Generalization in EEG Foundation Models

    Aditya Kommineni, Emily Zhou, Kleanthis Avramidis, Tiantian Feng, Shrikanth Narayanan

    cs.LG · cs.AI

    Evaluating foundation models under appropriate adaptation settings is essential for understanding the quality and transferability of the learned representations. Recent EEG foundation models have demonstrated promising transfer capabilities across tasks and datasets, motivating their growing use in neurotechnology and clinical applications. However, these models are typically evaluated under full fine-tuning on well-curated downstream...

    arxiv.org/abs/2605.28563 · PDF

  32. 32

    High Performance, Low Reliability: Uncertainty Benchmarking for Tabular Foundation Models

    José Lucas De Melo Costa, Fabrice Popineau, Arpad Rimmel, Bich-Liên Doan

    cs.LG

    Recent Tabular Foundation Models (TFMs) have demonstrated state-of-the-art predictive performance, often surpassing Gradient-Boosted Decision Trees (GBDTs). However, the trustworthiness of these models, particularly their uncertainty quantification, has been largely overlooked. We investigate this gap through an extensive study comparing TFMs, GBDTs, and classical baselines on the 112 datasets of the TALENT benchmark. Our results reveal a...

    arxiv.org/abs/2605.28554 · PDF

  33. 33

    Semi-Supervised Hypothesis Testing by Betting on Predictions

    Yaniv Tenzer, Elad Tolochinsky, Yaniv Romano

    cs.LG

    We introduce a testing-by-betting framework that leverages predictions on unlabeled data to enhance the power of sequential hypothesis testing. Given limited samples from the joint distribution of $(X,Y)$, and additional unlabeled samples from the marginal of $X$, we ask how unlabeled data can be used to hypothesize about the distribution of $Y$, and the conditional distribution of $Y\mid X$. We introduce an e-statistic and use it to...

    arxiv.org/abs/2605.28533 · PDF

  34. 34

    Stabilizing distribution-free probabilistic forecasts

    Jente Van Belle, Honglin Wen, Wouter Verbeke, Pierre Pinson

    cs.LG

    Multi-step-ahead forecasts are often updated as new observations become available, since shorter forecast horizons typically improve forecast quality. However, such improvements come at the cost of forecast instability, i.e., variability in forecasts for the same target period. This instability can trigger costly changes to plans formulated based on the forecasts and may erode trust in the forecasting system. In this work, we integrate...

    arxiv.org/abs/2605.28531 · PDF

  35. 35

    Stochastic Gradient Descent with Momentum is Algorithmically Stable

    Yunwen Lei, Zimeng Wang, Xiaoming Yuan

    cs.LG · cs.AI

    Stochastic gradient descent with momentum (SGDM) is one of the most widely used optimization algorithms in machine learning. While optimization properties of SGDM have been extensively studied in the literature, it remains insufficiently understood whether and when SGDM can generalize well to unseen data. In particular, it has been conjectured that while momentum accelerates training, it may degrade generalization. In this paper, we close...

    arxiv.org/abs/2605.28517 · PDF

  36. 36

    Learning Theory of the SVRG: Generalization and Convergence Analysis

    Yunwen Lei, Zimeng Wang, Xiaoming Yuan

    cs.LG · cs.AI

    Variance reduction (VR) methods employ stochastic gradients with decreasing variance, and they have been widely applied to solve large-scale optimization problems in machine learning because of their efficiency. Existing theoretical studies of VR methods are mainly focused on the convergence analysis, leaving the generalization behavior largely unexplored. In this paper, we bridge this gap by developing the first non-vacuous generalization...

    arxiv.org/abs/2605.28513 · PDF

  37. 37

    Universal Time Series Generation with Neural Controlled Differential Equations

    Torben Berndt, Elyes Farjallah, Leif Seute, Raeid Saqur, Benjamin Walker, Jan Stühmer

    cs.LG

    Recent work on the sequence universality of State Space Models (SSMs) has introduced efficient, maximally expressive continuous-time approaches for time-series modelling. While these works focus on discriminative settings, we extend this perspective to generative time-series modelling by proving that maximally expressive Structured Linear Controlled Differential Equations (SLiCEs) are universal time-series generators, in the sense that they...

    arxiv.org/abs/2605.28507 · PDF

  38. 38

    Fitting Unknown Number of Hyperplanes with Manifold Optimization

    Zhiqin Cheng, Yu Zhan, Mingjin Zhang, Lingbo Liu, Liang Lin

    cs.LG

    Fitting an unknown number of hyperplanes to data is a fundamental yet challenging problem in machine learning, characterized by its non-convexity, non-differentiability, and unknown model order. Existing approaches often struggle with local optima or lack geometric consistency. To address these limitations, we propose a novel framework based on Manifold Optimization. We reformulate the problem as an unsupervised learning task on the unit...

    arxiv.org/abs/2605.28501 · PDF

  39. 39

    Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training

    Avidan Shah, Jannik Brinkmann, Rico Angell

    cs.LG

    As LLMs gain stronger reasoning capabilities, their extended chain-of-thought introduces new degrees of complexity for defending against adversarial jailbreaks and prompt injection. We study consistency training, a family of fine-tuning objectives that enforce identical behavior on clean prompts and adversarial rewrites, and evaluate its two main variants, output-level (BCT) and activation-level (ACT), across five reasoning models. We...

    arxiv.org/abs/2605.28467 · PDF

  40. 40

    Bilinear Coordinate Alignment for Training-Free Task-Vector Transfer

    Jungyong Son, Jinwook Jung, Minhee Park, Sungyong Baik

    cs.LG

    Fine-tuning large-scale pre-trained models is a recent prevalent paradigm for adapting general representations to specialized tasks. However, when a new version of a pre-trained model becomes available, expertise acquired through fine-tuning cannot be directly reused because it is tied to the parameterization of the original model, requiring another costly fine-tuning. To address this inefficiency, recent work uses task vectors, defined as...

    arxiv.org/abs/2605.28444 · PDF

  41. 41

    Latent Diffusion for Missing Data

    Alberte Heering Estad, Ignacio Peis, Jes Frellsen

    cs.LG · stat.ML

    Diffusion models have emerged as powerful generative approaches for missing-data imputation, yet most existing methods operate directly in data space and degrade when training data are heavily incomplete. We investigate whether shifting diffusion to a learned latent representation improves robustness under missing-completely-at-random (MCAR) corruption. To this end, we propose a two-stage framework: a robust VAE-based imputer first learns...

    arxiv.org/abs/2605.28427 · PDF

  42. 42

    Conveyance: A Versatile Framework for Learning in Structured Class Spaces

    Yasser Taha, Grégoire Montavon, Nils Körber

    cs.LG

    While machine learning (ML) architectures have evolved rapidly to account for complex data, loss functions like cross-entropy remain mostly structure-agnostic in many real-world applications. However, the `class-symmetric' nature of these standard losses fundamentally limits the ability of ML models to exploit structural relationships between classes, particularly when facing structured noise. We propose \textsc{Conveyance}, a new...

    arxiv.org/abs/2605.28420 · PDF

  43. 43

    Revisiting Metafeatures to Explain Model Differences on Tabular Data

    Markus Herre, Andrej Tschalzev, Sascha Marton, Christian Bartelt

    cs.LG

    With the rise of tabular foundation models alongside traditional models still performing well on many tasks, choosing the right model for a tabular dataset remains difficult. We investigate whether dataset meta-features can explain performance gaps between model families on tabular prediction tasks. Using the TabArena benchmark results, we analyze dataset-level performance gaps and relate them to model-agnostic dataset descriptors. After...

    arxiv.org/abs/2605.28418 · PDF

  44. 44

    ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation

    Kun Liang, Chenming Tang, Clive Bai, Weijie Liu, Saiyong Yang, Yunfang Wu

    cs.LG · cs.AI

    On-policy distillation (OPD) transfers reasoning behavior by training a student on teacher feedback along student-generated trajectories, but standard full-rollout training ties every update to a costly completion and can over-allocate supervision to late positions with low marginal value for the current student. We revisit this assumption through the useful supervision horizon: student-induced rollouts can drift from teacher-preferred...

    arxiv.org/abs/2605.28396 · PDF

  45. 45

    CLANE: Continual Learning of Actions on Neuromorphic Hardware from Event Cameras

    Elvin Hajizada, Michael Neumeier, Edward Paxon Frady, Yulia Sandamirskaya, Axel von Arnim, Bing Li, Eyke Hüllermeier

    cs.LG · cs.AI · cs.NE

    Recognizing and continuously learning novel human actions without forgetting prior classes is a requirement for emerging AR/VR and robotics applications. For these applications, both on-device processing and learning are essential for privacy and low-latency adaptation. Event cameras address the efficiency of visual sensing with sparse, asynchronous output that is naturally compatible with neuromorphic processing. Yet no prior system has...

    arxiv.org/abs/2605.28387 · PDF

  46. 46

    Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference

    Alan Ferrari

    cs.LG

    Standard transformer architectures apply a single attention mechanism uniformly across all tokens and sequence positions, irrespective of local context or computational budget. We propose Meta-Attention, a framework that dynamically routes each token to the most appropriate attention strategy -- full softmax attention, linear (kernel) attention, or sliding-window local attention -- via a Bayesian Meta-Controller. Unlike prior routing...

    arxiv.org/abs/2605.28384 · PDF

  47. 47

    Teacher-Student Representational Alignment for Reinforcement Learning-Driven Imitation Learning

    Meraj Mammadov, Pedro Zuidberg Dos Martires, Johannes Andreas Stork

    cs.LG · cs.RO

    Imitation learning (IL) from a state-based reinforcement learning (RL) policy is a common approach to overcome the curse of dimensionality in complex and high-dimensional observation spaces prevalent in robotics. This paper addresses the irreducible imitation gap that emerges when teacher and student are learned in isolation, and the teacher policy has the liberty to rely on privileged state information that the student cannot infer from its...

    arxiv.org/abs/2605.28372 · PDF

  48. 48

    LEIA: Learned Environment for Interactive Architected Materials

    Haiqian Yang, Yuan Cao, Markus J. Buehler

    cs.LG · cond-mat.mtrl-sci · physics.app-ph

    World models have enabled interactive exploration of game environments and robotic manipulation, but physical engineering remains beyond their reach: real materials exhibit nonlinear constitutive laws, carry history-dependent internal state, undergo inertial dynamics, and may possess hierarchical structures spanning multiple length scales. We present LEIA (Learned Environment for Interactive Architected materials), a world model that lets...

    arxiv.org/abs/2605.28368 · PDF

  49. 49

    Score Based Error Correcting Code Decoder

    Alon Helvits, Eliya Nachmani

    cs.LG · cs.AI · cs.IT

    Error-correcting codes enable reliable communication, yet practical soft decoding remains challenging across code families and block lengths. We propose SB-ECC, a score-based decoder that casts decoding as continuous-time denoising. A neural denoiser defines a probability-flow ordinary differential equation (ODE) that iteratively updates the noisy channel observation toward a valid codeword, guided by parity constraints. The model is trained...

    arxiv.org/abs/2605.28358 · PDF

  50. 50

    Detecting Diffusion-Generated Time Series Under Generator Shift

    Zhi Wen Soi, Aditya Shankar, Gert Lek, Abele Mălan, Daniel Neider, Jian-Jia Chen, Lydia Chen

    cs.LG

    The boundary between real and diffusion-generated time series is becoming increasingly difficult to draw, yet detection in this domain remains underexplored, especially when the generator is unknown. We compare white-box detection, which requires access to the generator, against black-box detection, which operates on the raw signal alone. The white-box approach, a reconstruction-based detector adapted from the image domain, works well in...

    arxiv.org/abs/2605.28355 · PDF

  51. 51

    Dimensionality Reduction for Robust Federated Learning: A Theoretical Analysis and Convergence Guarantee

    Shiyuan Zuo, Jiashuo Li, Rongfei Fan, Han Hu, Jie Xu

    cs.LG

    Federated Learning (FL) enables multiple clients to collaboratively train models without sharing raw data, but it is highly vulnerable to Byzantine attacks. Existing robust approaches can neutralize these threats but incur substantial computational overhead during high-dimensional gradient aggregation, an overhead that scales poorly with model size and increasingly dominates the training cost as modern models grow larger. To address this...

    arxiv.org/abs/2605.28335 · PDF

  52. 52

    Learning the Error Patterns of Language Models

    Jinwoo Kim, Taylor Berg-KirkPatrick, Loris D'Antoni

    cs.LG · cs.AI

    When generating outputs for domains with specific validity constraints (e.g., a program should compile), LLMs often fail in a small number of focused ways: for example, by using Python function names when generating TypeScript. We observe that these error patterns can be represented using a small number of constraints that can be learned in practice. We propose \emph{prefix filters}, which are per-domain-and-LLM symbolic functions, as objects...

    arxiv.org/abs/2605.28328 · PDF

  53. 53

    Hybrid Neural World Models

    Pranav Lakshmanan, Paras Chopra

    cs.LG · cs.AI · math.NA · physics.comp-ph

    Neural surrogates promise large speedups over classical solvers for physical dynamics but fail silently at sharp dynamical events such as shocks, fronts, and contact. We present hybrid neural world models for physical dynamics: a recipe for training and deploying multi-horizon surrogates in physical state space, where a single network with continuous horizon conditioning is trained with direct supervision against textbook reference solvers to...

    arxiv.org/abs/2605.28317 · PDF

  54. 54

    Learning to Assess the Reliability of Number-of-Runs Estimation in Stochastic Optimization

    Sara Gjorgjieva, Eva Tuba, Tome Eftimov

    cs.LG · cs.NE

    In large-scale benchmarking of stochastic optimization algorithms, the key challenge is no longer whether repeated runs are needed for reliability, but how to determine when sufficient evidence has been collected without incurring unnecessary computational cost. We study a learning-based extension of a recent empirical online heuristic that adaptively estimates the required number of runs using outlier handling and skewness-based symmetry...

    arxiv.org/abs/2605.28309 · PDF

  55. 55

    Compositional Generalization in Autoregressive Models via Logit Composition

    Aakash Kumar, Maria Sofia Bucarelli, Emanuele Natale

    cs.LG

    Composing autoregressive models remains a core challenge in understanding how large language models can combine behaviors or skills learned across tasks. We introduce a new and principled composition strategy for autoregressive systems, inspired by composition methods developed for diffusion models. Under a factorized-conditionals assumption, we show that the resulting composition is projective: each component model preserves control over its...

    arxiv.org/abs/2605.28304 · PDF

  56. 56

    How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving

    Hanjiang Wu, Abhimanyu Rajeshkumar Bambhaniya, Sarbartha Banerjee, Tuhin Khare, Sudarshan Srinivasan, Suvinay...

    cs.LG · cs.AI · cs.DC

    Modern large language model (LLM) inference has progressively disaggregated to keep pace with growing model sizes and tight TTFT and TPOT service-level objectives: from chunked-prefill aggregation, to prefill-decode (P/D) disaggregation, and most recently to operator-level Attention-FFN Disaggregation (AFD). This trend is especially important for mixture-of-experts (MoE) models, where memory-bound attention, compute-intensive expert FFNs, and...

    arxiv.org/abs/2605.28302 · PDF

  57. 57

    T-GINEE: A Tensor-Based Multilayer Graph Representation Learning

    Maolin Wang, Ziting Mai, Xuhui Chen, Zhiqi Li, Tianshuo Wei, Yutian Xiao, Wenlin Zhang, Wanyu Wang, Ruocheng Guo,...

    cs.LG

    Traditional network analysis focuses on single-layer networks, real-world systems often form multilayer networks with multiple relationship types. However, existing methods typically fail to capture complex inter-layer dependencies by treating layers independently or aggregating them. To address this, we propose T-GINEE (Tensor-Based Generalized Multilayer-graph Estimating Equation), a statistical regularization framework combining...

    arxiv.org/abs/2605.28300 · PDF

  58. 58

    Machine Learning methods for event classification and vertex reconstruction of the 12C + 12C reaction with the MATE-TPC

    Minghui Zhang, Xiaobin Li, Jie Chen, Ningtao Zhang, Fenhua Lu, Junrui Ma, Jiazhen Yan, Wanqin Tu, Xiaodong Tang,...

    cs.LG · nucl-ex · physics.ins-det

    In modern nuclear physics experiments, identifying events of interest is challenging for nuclear reaction studies with the active target Time Projection Chamber (TPC). In this work, machine learning techniques are employed to analyze the complex data of the 12C + 12C fusion reaction from a TPC named MATE (multi-purpose active-target time projection chamber for nuclear experiments). Specifically, we successfully applied Residual Neural Network...

    arxiv.org/abs/2605.28296 · PDF

  59. 59

    ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation

    Hongru Hou, Tiehua Mei, Denghui Geng, Jinhui Huang, Ao Xu, Hengrui Chen, Jiaqing Liang, Deqing Yang

    cs.LG · cs.AI

    Proactive Recommender Systems (PRSs) aim to guide user preference shift toward target items by generating paths of intermediate recommendations. Reinforcement learning (RL) provides a principled framework for optimizing such sequential decision tasks, as path rewards can naturally capture both short-term acceptance and long-term guidance effectiveness. However, naively applying policy gradients to PRS results in deficient gradient estimation....

    arxiv.org/abs/2605.28293 · PDF

  60. 60

    Adaptive Bandit Algorithms for Contextual Matching Markets

    Shiyun Lin, Simon Mauras, Vianney Perchet, Nadav Merlis

    cs.LG · cs.GT · stat.ML

    We study bandit learning in matching markets, where players and arms constitute the two market sides, and the players' utilities are linear in the arm contexts. In each round, new arms arrive with observable contexts. Then, the algorithm matches them to players, aiming to minimize each player's regret against a stable matching benchmark. This contextual structure creates significant complexity: subtle context shifts can slightly alter one...

    arxiv.org/abs/2605.28290 · PDF

  61. 61

    AtomComposer: Discovering Chemical Space from First Principles with Reinforcement Learning

    Bjarke Hastrup, Francois Cornet, Tejs Vegge, Arghya Bhowmik

    cs.LG · cond-mat.mtrl-sci

    Discovering novel stable molecules without training data remains a grand scientific challenge. Current molecular generative models are trained on large, pre-curated datasets, which introduce biases and limit exploration of novel chemistry. In contrast, we propose a new paradigm: autonomous, generalized agents capable of mapping vast, unknown chemical spaces without any pretraining. For the first time, we present AtomComposer, a self-guided...

    arxiv.org/abs/2605.28287 · PDF

  62. 62

    Commit to the Bit: Reactive Reinforcement Learning Done Right

    Onno Eberhard, Claire Vernade, Michael Muehlebach

    cs.LG

    Reinforcement learning algorithms are commonly analyzed (and designed) under the Markov assumption. This is unrealistic, as most environments encountered in practice are either partially observable, or require function approximation that restricts the agent to access non-Markovian state features. We consider the problem of learning an optimal reactive policy in a finite environment with deterministic observations (or equivalently, hard state...

    arxiv.org/abs/2605.28276 · PDF

  63. 63

    Dynamic Topic Modeling with a Higher-Order Hypergraphical Representation

    Hanjia Gao, Hanwen Ye, Qing Nie, Annie Qu

    cs.LG · stat.ME

    Dynamic topic modeling is widely used to analyze evolving trends in scientific literature, medical records, and social media. Traditional topic models represent each topic through a single probability vector on the multinomial simplex and implicitly couple word occurrence and repetition within one probabilistic mechanism. However, this formulation restricts the dependence structure among words and overlooks informative higher-order...

    arxiv.org/abs/2605.28269 · PDF

  64. 64

    Parameter-Efficient Generative Modeling with Controlled Vector Fields

    Peyman Morteza

    cs.LG · stat.ML

    We introduce a continuous-time generative modeling framework, motivated by the Chow-Rashevskii theorem, that builds expressive flows from a small set of fixed vector fields and learned scalar controls. Instead of learning an unconstrained high-dimensional vector field, our framework constructs the velocity by modulating fixed vector fields with learned scalar control functions. When the fixed fields are bracket-generating, their Lie algebra...

    arxiv.org/abs/2605.28267 · PDF

  65. 65

    IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage

    Yuhan Li, Mingxu Zhang, Dazhong Shen, Ying Sun

    cs.LG · cs.AI

    Reinforcement learning with verifiable rewards (RLVR) has become a key technique for en- hancing LLM reasoning, yet its data ineffi- ciency remains a major bottleneck. Existing methods address this problem only partially, each missing at least one of subset-level cov- erage, verifier signal use, or interpretability. To address this gap, we present IRDS (Inter- pretable RLVR Data Selection), which selects RLVR training instances on a sparse...

    arxiv.org/abs/2605.28247 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.