cs.LG · 2026-06-11 · No. 20

Machine Learning, 2026-06-11.

61 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

61 entries
  1. 01

    Redesign Mixture-of-Experts Routers with Manifold Power Iteration

    Songhao Wu, Ang Lv, Ruobing Xie, Yankai Lin

    cs.LG · cs.AI · cs.CL

    Router is the cornerstone component to the Mixture-of-Experts models. Serving as expert proxies, the rows of the router matrix compute their similarity to the MoE inputs to determine which subset of experts is activated. Ideally, each router row is designed to encode the expert matrix into this representative vector, such that its dot-product with token can better reflect token-expert affinity. However, there exists no design principles to...

    arxiv.org/abs/2606.12397 · PDF

  2. 02

    ATLAS: Active Theory Learning for Automated Science

    Noémi Éltető, Nathaniel D. Daw, Kimberly L. Stachenfeld, Kevin J. Miller

    cs.LG · cs.AI

    Advancing scientific understanding through mechanistic modeling requires posing the right experimental questions to yield maximally informative data. To automate this pursuit within cognitive science, we introduce ATLAS (Active Theory Learning for Automated Science), an active learning framework for the data-driven discovery of interpretable behavioral models. ATLAS iterates between generating mechanistic hypotheses--instantiated as a diverse...

    arxiv.org/abs/2606.12386 · PDF

  3. 03

    APPO: Agentic Procedural Policy Optimization

    Xucong Wang, Ziyu Ma, Yong Wang, Yuxiang Ji, Shidong Yang, Guanhua Chen, Pengkun Wang, Xiangxiang Chu

    cs.LG · cs.AI

    Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing methods assign credit over coarse heuristic units, such as tool-call boundaries or fixed workflows, making it difficult to identify which intermediate decisions influence downstream outcomes. In this work, we study agentic RL from two perspectives: \textit{where to...

    arxiv.org/abs/2606.12384 · PDF

  4. 04

    Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

    Yucheng Li, Huiqiang Jiang, Yang Xu, Jianxin Yang, Yi Zhang, Yizhong Cao, Yuhao Shen, Fan Zhou, Rui Men, Jianwei...

    cs.LG · cs.CL

    Reinforcement learning (RL) has become a key component in modern large language models, yet the rollout stage remains the key bottleneck in RL training pipelines. Although Multi-Token Prediction (MTP) offers a natural solution to accelerate rollouts through speculative decoding, many studies have observed that MTP acceptance rates degrade significantly during RL training, leading to limited speedup performance. To address this bottleneck, we...

    arxiv.org/abs/2606.12370 · PDF

  5. 05

    On Subquadratic Architectures: From Applications to Principles

    Anamaria-Roberta Hartl, Levente Zólyomi, David Stap, Pieter-Jan Hoedt, Niklas Schmidinger, Lukas Hauzenberger,...

    cs.LG

    Transformers dominate modern sequence modeling, but their quadratic attention incurs substantial computational cost. Subquadratic architectures offer a scalable alternative. However, it remains unclear which designs yield the most effective sequence models. We compare three leading approaches: xLSTM, Mamba-2, and Gated DeltaNet. We evaluate these models on tasks with complex dependencies: (1) code-model pre-training, (2) distillation of code...

    arxiv.org/abs/2606.12364 · PDF

  6. 06

    Latent World Recovery for Multimodal Learning with Missing Modalities

    Hui Wang, Tianyu Ren, Joseph Butler, Christopher Baker, Karen Rafferty, Simon McDade

    cs.LG · cs.AI

    We study multimodal learning under missing modalities, with particular motivation from bioscience applications in which heterogeneous modalities are often only partially available when decisions need to be made. We propose Latent World Recovery (LWR), a framework built on two key ideas: (i) modality-specific embeddings from different modalities are aligned in a shared latent space, and (ii) a unified representation is constructed by fusing...

    arxiv.org/abs/2606.12362 · PDF

  7. 07

    Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal

    Leon Bergen, Usha Bhalla, Sidharth Baskaran, Max Loeffler, Raphael Sarfati, Dhruvil Gala, Ryan Panwar, Santiago...

    cs.LG

    Language-model post-training is the main stage at which model behavior is shaped, yet it still largely involves optimization of scalar rewards that summarize diverse desiderata. This abstraction gives practitioners little visibility into what their data actually teaches models, allowing spurious correlations to be learned by a model and inducing undesirable behaviors such as over-stylization and sycophancy. To address this problem, we ask:...

    arxiv.org/abs/2606.12360 · PDF

  8. 08

    Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

    Mengyu Zheng, Kai Han, Boxun Li, Haiyang Xu, Yuchuan Tian, Wei He, Hang Zhou, Jianyuan Guo, Hailin Hu, Lin Ma, Chao...

    cs.LG · cs.CL

    General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring. We introduce Claw-SWE-Bench, a multilingual SWE-bench-style benchmark and adapter protocol that makes heterogeneous agent harnesses, or claws, comparable under fair...

    arxiv.org/abs/2606.12344 · PDF

  9. 09

    Fourier Features Let Agents Learn High Precision Policies with Imitation Learning

    Balázs Gyenes, Emiliyan Gospodinov, Jan Frieling, Enrico Krohmer, Nicolas Schreiber, Xiaogang Jia, Niklas Freymuth,...

    cs.LG · cs.RO

    High-precision robotic manipulation requires fine-grained spatial reasoning that is often difficult to achieve with RGB-only policies due to depth ambiguity and perspective scale issues. Policies that leverage 3D information directly, such as those based on point clouds, offer a stronger geometric prior over purely image-based ones, yet their performance remains highly task-dependent. We hypothesize that this discrepancy may be due to the...

    arxiv.org/abs/2606.12334 · PDF

  10. 10

    Harness In-Context Operator Learning with Chain of Operators

    Minghui Yang, Ling Guo, Liu Yang

    cs.LG · cs.AI

    Neural operators approximate mappings between function spaces, but often generalize poorly to other operators and usually require fine-tuning or retraining. In-Context Operator Networks (ICON) addresses this issue by prompting the model with numerical context so that the model learns specific operators from prompts and adapt to different operators without fine-tuning. However, ICON may still fail to generalize to out-of-distribution (OOD)...

    arxiv.org/abs/2606.12318 · PDF

  11. 11

    The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

    Pietro Barbiero, Giovanni De Felice, Mateo Espinosa Zarlenga, Francesco Giannini, Filippo Bonchi, Mateja Jamnik,...

    cs.LG · cs.AI · cs.NE

    As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations. However, interpretability lacks general theories to deductively design interpretable methods. This gap between theories and methods results in a fragmented literature and inconsistent evaluation protocols. To fill this gap, we introduce the Standard Interpretable Model (SIM),...

    arxiv.org/abs/2606.12289 · PDF

  12. 12

    Holding the FP8 Quality Ceiling at 8-Bit Weights and Activations: INT8 and GGUF Post-Training Quantization of Ideogram 4.0 for Consumer GPUs

    Deep Gandhi, Ali Asaria, Tony Salomone

    cs.LG

    Post-training quantization lets large text-to-image diffusion transformers run on consumer GPUs, yet the hardware-specific trade-offs are seldom measured directly. We quantize Ideogram 4.0 - a 9.3B flow-matching diffusion transformer (DiT), shipped as two separate-weight copies of a single-stream 34-layer backbone for classifier-free guidance and conditioned by a Qwen3-VL-8B encoder - for Ampere RTX 3090 GPUs, which lack FP8 tensor cores. Our...

    arxiv.org/abs/2606.12280 · PDF

  13. 13

    Finding Multiple Interpretations in Datasets

    Matthew Chak, Paul Anderson

    cs.LG

    In this paper, we propose an approach to finding sets of similar-performing models (in terms of loss/accuracy measurements) with highly different context-aware characteristics. Through experiments on the METABRIC dataset, we show that the proposed method finds multiple models with highly different gene expressions than those found by the control methodology without performance penalties. We argue that the proposed methodology is important...

    arxiv.org/abs/2606.12277 · PDF

  14. 14

    Using Explainability as a Training-Time Reliability Signal for Efficient ECG Classification

    Veerendhra Kumar Dangeti, Xiao Gu, Ying Weng, Shreyank N Gowda

    cs.LG · cs.AI

    Training deep neural networks for clinical time-series analysis is computationally demanding, yet many healthcare settings lack the resources required for repeated model development and deployment. This challenge is particularly evident in electrocardiogram classification, where large datasets and long training schedules make efficiency practically important. Progressive Data Dropout reduces training cost by excluding samples from gradient...

    arxiv.org/abs/2606.12252 · PDF

  15. 15

    Reinforcement Learning Disrupts Gradient-Based Adversarial Optimization

    Xinhai Zou, Chang Zhao, Alireza Aghabagherloo, Dave Singelée, Robin Degraeve, Bart Preneel

    cs.LG · cs.AI · cs.CR

    Gradient-based adversarial attacks remain a dominant threat to deep neural networks (DNNs), as they exploit gradient information to efficiently optimize adversarial perturbations. To address this, we investigate whether reinforcement learning (RL) training can disrupt the gradient structure used by attackers by training image classifiers with policy-gradient objectives and epsilon-greedy exploration. Through systematic experiments across...

    arxiv.org/abs/2606.12251 · PDF

  16. 16

    Multi-Rate Mixture of Experts for Accelerating Liquid Neural Network Training

    Shilong Zong, Almuatazbellah Boker, Hoda Eldardiry

    cs.LG · cs.AI

    Multivariate time-series data often exhibit complex temporal dependencies, irregular sampling, and heterogeneous dynamics across multiple time scales, making accurate sequence modeling particularly challenging. Traditional recurrent neural networks (RNNs), such as Long Short-Term Memory (LSTM) networks, operate in discrete time and may struggle to effectively capture continuous and irregular temporal behaviors. Liquid Neural Networks (LNNs)...

    arxiv.org/abs/2606.12240 · PDF

  17. 17

    Re-evaluating Confidence Remasking in Masked Diffusion Language Models

    Stipe Frkovic, Metod Jazbec, Dan Zhang, Christian A. Naesseth, Ilija Bogunovic, Eric Nalisnick

    cs.LG

    Masked diffusion language models (dLLMs) have recently emerged as a competitive alternative to autoregressive language models, with the promise of faster inference via parallel token generation. A notable limitation of the masked formulation, however, is that once a token has been unmasked it can no longer be revised, leaving dLLMs vulnerable to early sampling mistakes. To address this, a growing body of work has sought to extend masked dLLMs...

    arxiv.org/abs/2606.12232 · PDF

  18. 18

    Implicit Neural Representations of Individual Behavior

    Andrew Kang, Priya Narasimhan

    cs.LG · cs.AI

    We study policy representation learning from unlabeled multi-policy behavioral data. Each episode is generated by a fixed policy, but policy labels are unavailable. This setting appears in robotics play, demonstrations, games, racing, and other datasets where heterogeneous behaviors are mixed without annotations. We introduce \emph{Behavioral INR}, a self-supervised generative model that adapts implicit neural representations (INRs) from...

    arxiv.org/abs/2606.12200 · PDF

  19. 19

    How Low Can You Go? Active Learning for Sparse Model Discovery in the Ultra-Low-Data Limit

    Ana Larrañaga, Urban Fasel, Steven L. Brunton

    cs.LG · math.DS · math.OC

    Identifying the governing equations of complex dynamical systems remains a fundamental challenge across science and engineering. While early approaches relied on empirical data and heuristics, modern data-driven methods offer greater flexibility and fewer assumptions. However, data acquisition in real-world settings is often expensive. This work addresses this challenge by introducing an active learning strategy for dynamics discovery in the...

    arxiv.org/abs/2606.12182 · PDF

  20. 20

    nD-RoPE: A Generalized RoPE for n-Dimensional Position Embedding

    Boyang Li, Yulin Wu, Sizhe Xu, Nuoxian Huang, Zhonghang Yuan, Shangyi Guo, Shu Yang, Takahiro Yabe

    cs.LG · cs.AI

    Rotary Position Embedding (RoPE) is widely adopted in Transformer models, yet its extension to high-dimensional domains lacks a unified theoretical formulation. Most existing approaches either apply rotations independently along each axis or empirically mix frequencies, which limits cross-dimensional interactions and yields direction-dependent representations. To address these limitations, we propose nD-RoPE, a decomposition-free...

    arxiv.org/abs/2606.12146 · PDF

  21. 21

    PCA-Enhanced Adaptive NVAR Framework for High-Resolution Sea Surface Temperature Forecasting in the East Sea

    Sherkhon Azimov, Susana López-Moreno, Eric Dolores-Cuenca, JinYong Choi, Sangil Kim

    cs.LG

    Accurate forecasting of sea surface temperature (SST) in regional seas such as the East Sea is crucial for monitoring marine ecosystems, assessing climate risks, managing fisheries, and conducting naval operations. Traditional numerical ocean models provide reliable predictions but are computationally expensive and often unsuitable for real-time forecasting. Many deep learning methods also struggle with high-dimensional spatiotemporal ocean...

    arxiv.org/abs/2606.12141 · PDF

  22. 22

    Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders

    Gleb Gerasimov, Timofei Rusalev, Nikita Balagansky, Daniil Laptev, Vadim Kurochkin, Daniil Gavrilov

    cs.LG · cs.AI · cs.CL

    Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training runs. We study this question through \emph{feature stability}: for each SAE feature, we estimate the probability that a similar feature reappears in an independently trained SAE. This yields a scalable per-feature signal that separates stable from unstable features....

    arxiv.org/abs/2606.12138 · PDF

  23. 23

    A Riemannian Approach to Low-Rank Optimal Transport

    Pratik Jawanpuria, Bamdev Mishra

    cs.LG · math.OC

    Low-rank optimal transport (OT) mitigates the quadratic scaling of classical solvers, yet existing approaches rely heavily on first-order mirror-descent updates that require careful hyperparameter tuning and ignore the optimization landscape's curvature. To address these limitations, we propose a unified Riemannian geometric framework for low-rank OT, modeling balanced and unbalanced rank-$r$ positive factored couplings as novel smooth...

    arxiv.org/abs/2606.12120 · PDF

  24. 24

    Efficient Time Series Clustering from Multiscale Reservoir Dynamics with Granular-Ball Anchoring Graph Optimization

    Yifan Wang, Lifeng Shen, Shuyin Xia, Yi Wang

    cs.LG

    Time-series clustering remains challenging due to the inherent trade-off between clustering effectiveness and computational efficiency. Similarity-based methods often suffer from quadratic complexity caused by pairwise distance computations, while deep learning-based approaches typically rely on costly iterative training and a large number of trainable parameters. In this paper, we propose MSRGC-Net, an efficient time-series clustering...

    arxiv.org/abs/2606.12077 · PDF

  25. 25

    Attention by Synchronization in Coupled Oscillator Networks

    Fabio Pasqualetti, Taosha Guo

    cs.LG · cs.NE · nlin.AO

    We address transformer attention on energy-constrained physical substrates. Softmax attention requires exponentiation and global reduction, operations with high energy cost on von Neumann hardware and no natural physical analog. We show that Kuramoto synchronization dynamics (which arise in electrical, mechanical, superconducting, and charge-density-wave oscillator arrays, among other physical systems) implement a well-defined attention...

    arxiv.org/abs/2606.12059 · PDF

  26. 26

    Simplicity Suffices for Parameter Noise Injection in Stochastic Gradient Descent

    Benjamin Leblanc, Louis-Jacob Lebel, Teddy Kana, Richard Kamel

    cs.LG

    Injecting noise into the optimization process is a well-established technique for improving the training and generalization of deep neural networks. Yet, despite the breadth of existing approaches, it remains unclear which design choices truly matter in practice. In this work, we investigate parameter noise injection for stochastic gradient descent, focusing on two key questions: how to efficiently pair each training example with its own...

    arxiv.org/abs/2606.12054 · PDF

  27. 27

    Reliable Error Estimation for PINNs: Lower and Upper A Posteriori Bounds

    Ismail Huseynov, Arzu Ahmadova, Agamirza Bashirov

    cs.LG · math.DS

    Physics-informed neural networks (PINNs) combine machine learning with physical laws to solve differential equations. While existing results provide rigorous \emph{a posteriori} upper bounds for PINN prediction errors, complete certification also requires complementary lower information in order to obtain computable two-sided error enclosures. In this paper, we derive computable \emph{a posteriori} lower bounds for PINN errors in ordinary...

    arxiv.org/abs/2606.12050 · PDF

  28. 28

    Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization

    Frank Xiao, Mary Phuong

    cs.LG · cs.AI

    Model post-training, and in particular reinforcement learning (RL), is one of the primary mechanisms by which developers can shape models' values and behaviors. However, as models become increasingly evaluation and training aware, they may be motivated to resist training when the perceived objective conflicts with their current values, undermining developers' ability to detect misalignment and correct model behavior through further training....

    arxiv.org/abs/2606.12016 · PDF

  29. 29

    Tabular Foundation Models for Clinical Survival Analysis via Survival-Aware Adaptation

    Minh-Khoi Pham, Luca Cotugno, Alina Sirbu, Tai Tan Mai, Martin Crane, Marija Bezbradica

    cs.LG · cs.AI

    Predicting time-to-event outcomes such as mortality is a fundamental task in clinical decision-making, commonly addressed through survival analysis. While classical statistical and deep learning approaches have been widely studied, they typically require task-specific training and sufficient labeled data. Recent advances in tabular foundation models offer a new paradigm by learning general-purpose representations for structured data. However,...

    arxiv.org/abs/2606.12006 · PDF

  30. 30

    Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents

    Frank Xiao, Mary Phuong

    cs.LG

    Trusted monitoring is a cornerstone of AI control. However, as frontier models grow more capable, the increasing capabilities gap between trusted and untrusted models may render trusted models unreliable monitors. We introduce \emph{bootstrapped monitoring}, a protocol that addresses this by inserting a stronger, intermediate untrusted model with transparent chain-of-thought reasoning into the oversight chain. The untrusted monitor ($U_m$)...

    arxiv.org/abs/2606.11998 · PDF

  31. 31

    Time-Series Foundation Model Embeddings for Remaining Useful Life Estimation

    Amir El-Ghoussani, Michele De Vita, Ronald Naumann, Valiseios Belagiannis

    cs.LG · cs.AI

    Remaining Useful Life (RUL) prediction is essential for industrial predictive maintenance, yet many learning-based approaches rely on extensive feature engineering or large labeled datasets to train task-specific sequence models. In this work, we introduce a lightweight learning approach, in which we leverage a frozen pretrained time-series foundation model (TSFM) and combine it with a small regression head for RUL estimation from...

    arxiv.org/abs/2606.11990 · PDF

  32. 32

    What Uncertainties Do We Need for Dynamical Systems?

    Yusuf Sale, Christopher Bülte, Felix Czaja, Joshua Stiller, Eyke Hüllermeier

    cs.LG · stat.ML

    The distinction between aleatoric and epistemic uncertainty has received considerable attention in machine learning research, mainly in the context of supervised learning but also in other settings such as generative modeling. In this paper, we offer a machine learning perspective on uncertainty modeling for dynamical systems, which has been studied much less so far. In particular, we ask: what uncertainties do we need for dynamical systems?...

    arxiv.org/abs/2606.11988 · PDF

  33. 33

    PAWS: Preference Learning with Advantage-Weighted Segments

    Aleksandar Taranovic, Onur Celik, Niklas Freymuth, Ge Li, Serge Thilges, Huy Le, Tai Hoang, Rania Rayyes, Gerhard Neumann

    cs.LG

    Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations. Existing methods typically train utility functions on trajectory or segment-level preferences while relying on per-step utility estimates during policy optimization. This training and inference mismatch induces a distribution shift that severely degrades temporal credit assignment...

    arxiv.org/abs/2606.11982 · PDF

  34. 34

    Efficient Multinomial Logistic Bandit via Frequent Directions

    Linzhe He, Yu-Jie Zhang, Sifan Yang, Lijun Zhang

    cs.LG · stat.ML

    This paper studies efficient online algorithms for multinomial logistic bandits (MLogB), where the feedback distribution over $K+1$ outcomes follows a multinomial logistic model of $d$-dimensional action vectors. A representative UCB-type algorithm, OFUL-MLogB, achieves a regret bound of $\tilde{\mathcal{O}}(Kd\sqrt{T})$, but still requires $\mathcal{O}(K^3d^3)$ time and $\mathcal{O}(K^2d^2)$ space per round due to parameter estimation and...

    arxiv.org/abs/2606.11968 · PDF

  35. 35

    HAMNO: A Hierarchical Adaptive Multi-scale Neural Operator with Physics-Informed Learning for Dynamical Systems

    Mostafa Bamdad, Mohammad Sadegh Eshaghi, Timon Rabczuk

    cs.LG · physics.comp-ph

    Neural operators provide a powerful framework for learning solution mappings of partial differential equations directly in function space. However, many existing architectures still struggle to represent nonlinear time-dependent systems that involve multi-scale structures, long-range interactions, and stable long-time evolution. In this work, we introduce the Hierarchical Adaptive Multi-scale Neural Operator (HAMNO), a neural-operator...

    arxiv.org/abs/2606.11963 · PDF

  36. 36

    Categorical Prior Lock-in: Why In-Context Learning Fails for Structured Data

    Antonio Pelusi, Stefano Braghin, Alberto Trombetta

    cs.LG · cs.AI

    Large language models (LLMs) are increasingly used as conditional generators for structured data, relying on in-context learning (ICL) to adapt to new distributions without parameter updates. We investigate the limits of ICL for structured generation under distribution mismatch, using high-cardinality tabular data as a controlled test case, and identify a structural failure mode we term \textit{categorical prior lock-in}: the inability of ICL...

    arxiv.org/abs/2606.11961 · PDF

  37. 37

    Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers

    Jun Wen Leong

    cs.LG · cs.CR · stat.ML

    We present an online monitoring system for distributional shift in deployed safety classifiers, using calibrated sequential statistics to detect when a classifier has moved out of distribution. Upon detection, a conformal abstention layer adapts decision thresholds to recover a target error rate epsilon=0.1. In a pre-registered factorial evaluation (4 classifiers x 5 shift conditions x 20 seeds x 2 window sizes, 800 cells), the system...

    arxiv.org/abs/2606.11949 · PDF

  38. 38

    Beyond representational alignment with brain-guided language models for robust reasoning

    Mingqing Xiao, Kai Du, Zhouchen Lin

    cs.LG · cs.AI · cs.CL · q-bio.NC

    The correspondence between large language models (LLMs) and the neural mechanisms underlying human higher-order cognition remains insufficiently characterized. Given that language and reasoning in the human brain appear dissociable, an open question is whether LLMs align with neural signals from reasoning-related regions and whether such signals can improve them. Here, focusing on deductive reasoning, we show that LLM internal representations...

    arxiv.org/abs/2606.11893 · PDF

  39. 39

    MemNovo: Look Back at the Spectrum for Balanced De Novo Peptide Sequencing from Mass Spectrometry

    Dongxin Lyu, Jingbo Zhou, Hongxin Xiang, Yuqiang Li, Jun Xia

    cs.LG · q-bio.QM

    De novo peptide sequencing from tandem mass spectrometry is pivotal in proteomics, enabling identification of novel peptides without reference databases. While recent Transformer-based encoder-decoder models have achieved remarkable performance, we uncover a critical pathology in their inference dynamics. Through comprehensive feature scaling experiments, we demonstrate that existing auto-regressive peptide decoders tend to over-rely on...

    arxiv.org/abs/2606.11868 · PDF

  40. 40

    RePAIR: Predictive Self-Supervised Representation Learning in Chess

    Christoph Koller, Johannes Fürnkranz, Timo Bertram

    cs.LG

    In this paper, we introduce Representation Prediction via Autoencoding using Iterative Refinement (RePAIR) - a novel self-supervised representation learning architecture that synthesizes Masked Autoencoders (MAE), Joint Embedding Predictive Architectures (JEPA), and Bidirectional Encoder Representations from Transformers (BERT). We demonstrate how it can be used to encode objects in sequential data like consecutive chess positions into...

    arxiv.org/abs/2606.11860 · PDF

  41. 41

    Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training

    Michal Chudoba, Sergey Alyaev, Petra Galuscakova, Tomasz Wiktorski

    cs.LG · cs.AI · cs.CL

    There are two main Parameter-Efficient Fine-Tuning (PEFT) techniques for Large Language Models (LLMs). While Low-Rank Adaptation (LoRA) introduces additional weights between the LLM layers, Soft Prompting introduces additional fine-tuning-specific raw tokens to an LLM input. However, both require modification to the computational graphs of precompiled, preoptimized LLMs. As a result, neither is fully supported in high-throughput engines like...

    arxiv.org/abs/2606.11854 · PDF

  42. 42

    TaskFusion: Continual Anomaly Detection for Heterogeneous Tabular Data

    Dayananda Herurkar, Federico Raue, Joachim Folz, Jörn Hees, Andreas Dengel

    cs.LG

    Continual anomaly detection in tabular data is challenging and remains largely underexplored, particularly in settings with heterogeneous feature schemas, distribution shifts, and severe class imbalance. In many real-world applications, data arrive sequentially from diverse domains, rendering conventional continual learning methods ineffective due to their reliance on a fixed input space. We propose a continual learning (CL) method, which can...

    arxiv.org/abs/2606.11844 · PDF

  43. 43

    Flow Matching with In-Context Priors for Out-of-Distribution Brain Dynamics

    Sam Gijsen, Michał Łukomski, Marc-André Schulz, Kerstin Ritter

    cs.LG · q-bio.NC

    Flow matching and diffusion models enable conditional generation across domains ranging from images to proteins, with recent extensions to out-of-distribution contexts. Yet generative models of neural time series have largely remained restricted to categorical conditioning, precluding compositional and zero-shot generalization. In this work, we propose a per-timestep conditioned diffusion transformer for generating realistic fMRI brain...

    arxiv.org/abs/2606.11833 · PDF

  44. 44

    From Uniform to Learned Graph Priors: Diffusion for Structure Discovery

    Qi Shao, Hao Guo, Jiawen Chen, Duxin Chen, Wenwu Yu

    cs.LG · cs.AI

    Neural relational inference (NRI) methods discover interaction graphs from trajectories through variational reasoning on discrete potential edges. However, these methods typically rely on oversimplified, factorized graph priors. Such priors, typically nearing uniform distributions, treat edges as independent entities. This systemic misalignment does not match the real-world systems and yields diffuse and indecisive edge posteriors limiting...

    arxiv.org/abs/2606.11831 · PDF

  45. 45

    Space-sampled Value Decay: Forgetting Mechanisms for Non-stationary Deep Reinforcement Learning

    Felix Störck, Fabian Hinder, Barbara Hammer

    cs.LG

    Studies on rodents such as mice have shown the capabilities to adapt their behavior when dealing with changing parameters (``drift'') of the environment even if no information about change is provided (uncertainty) -- a behavior that can be modeled by forgetting mechanisms. Non-stationary Reinforcement Learning (NSRL) deals with adapting state-of-the-art RL methods to deal with changing environments: these however usually require (partially)...

    arxiv.org/abs/2606.11797 · PDF

  46. 46

    Multimodal Ordinal Modeling of Alzheimer's Disease Severity Using Structural MRI and Clinical Data

    Boris-Stephan Rauchmann, Jonathan Laib, Buse Ercik, Robert Perneczky, Sergio Altares-López

    cs.LG · cs.AI

    Neurodegenerative diseases such as Alzheimer's disease (AD) require accurate and scalable tools for assessing disease severity, yet current clinical staging remains time-intensive and prone to variability. We propose an attention-enhanced multimodal machine learning framework with ordinal regression for automated and interpretable AD severity staging. The framework integrates T1-weighted MRI with demographic and genetic variables and compares...

    arxiv.org/abs/2606.11794 · PDF

  47. 47

    AI4Land: Scalable Deep Learning for Global High-Resolution Land Use Reconstruction

    Amirpasha Mozaffari, Marina Castaño, Stefano Materia, Etienne Tourigny, Oscar Molina-Sedano, Jordi Varela-Agrelo,...

    cs.LG · cs.AI · physics.ao-ph

    Uncertainty in the terrestrial carbon cycle remains a major constraint in climate projections, partly driven by the uncertainties affecting the land surface representation and variability in Earth system models. To address this limitation, we present a data-driven framework AI4Land, for generating high-resolution historical reconstructions and future projections of key land surface variables. The framework follows a two-phase approach using a...

    arxiv.org/abs/2606.11793 · PDF

  48. 48

    RCAP: Robust, Class-Aware, Probabilistic Dynamic Dataset Pruning

    Atif Hassan, Swanand Khare, Jiaul H. Paik

    cs.LG

    Dynamic data pruning techniques aim to reduce computational cost while minimizing information loss by periodically selecting representative subsets of input data during model training. However, existing methods often struggle to maintain strong worst-group accuracy, particularly at high pruning rates, across balanced and imbalanced datasets. To address this challenge, we propose RCAP, a Robust, Class-Aware, Probabilistic dynamic dataset...

    arxiv.org/abs/2606.11761 · PDF

  49. 49

    ICA Lens: Interpreting Language Models Without Training Another Dictionary

    Sida Liu, Feijiang Han

    cs.LG · cs.AI · cs.CL

    Finding interpretable directions in language-model representations is critical for understanding and controlling model behavior. Sparse autoencoders (SAEs) have become the standard tool for this purpose, but using them as the default first lens often requires training, storing, and evaluating large overcomplete dictionaries. This bottleneck limits rapid exploration and raises a fundamental question: how much interpretable structure is already...

    arxiv.org/abs/2606.11722 · PDF

  50. 50

    Capacity-Constrained Online Convex Optimization with Delayed Feedback

    Alexander Ryabchenko, Idan Attias, Daniel M. Roy

    cs.LG · stat.ML

    Online learning with delayed feedback typically assumes that the learner can track all pending rounds until their feedback arrives. In practice, tracking resources are finite, and feedback from untracked rounds is permanently lost. In this paper, we study delayed online convex optimization (OCO) under a hard capacity constraint, where at most $C$ pending rounds can be tracked at any time. To model delay information, we introduce a...

    arxiv.org/abs/2606.11711 · PDF

  51. 51

    RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

    Leyi Pan, Shuchang Tao, Yunpeng Zhai, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Aiwei Liu, Lijie Wen

    cs.LG · cs.CL

    On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distribution it produces under privileged context, typically a verified solution. However, we show that the learning signal drawn from this distributional gap concentrates on style tokens rather than task-bearing ones, as the hinted model tends to produce more direct, shorter outputs. We term this...

    arxiv.org/abs/2606.11709 · PDF

  52. 52

    A Data-Centric Framework for Detecting and Correcting Corrupted Labels

    Ha-Linh Nguyen, Hong-Anh Nguyen, Minh-Duc La, Thu-Trang Nguyen, Son Nguyen, Hieu Dinh Vo

    cs.LG

    The performance of machine learning and deep learning models largely depends on the quality of the training data. However, the quality of the real-world datasets is often compromised by noisy labels, which can substantially degrade model accuracy and reliability. To address this challenge, we propose Relabeler, an end-to-end data-centric framework for detecting and correcting corrupted labels. For corrupted label detection, Relabeler jointly...

    arxiv.org/abs/2606.11699 · PDF

  53. 53

    Noise-Aware Framework for Correcting Corrupted Labels

    Ha-Linh Nguyen, Hong-Anh Nguyen, Minh-Duc La, Phong Lam, Thu-Trang Nguyen, Son Nguyen, Hieu Dinh Vo

    cs.LG · cs.AI

    High-quality labeled data is essential for training reliable ML/DL models. However, real-world datasets often contain a considerable proportion of corrupted labels, which can severely degrade model performance. To address this problem, we propose CANOLA, a novel framework for correcting corrupted labels through noise-aware learning and iterative label refinement. CANOLA explicitly estimates the underlying noise distribution of the dataset and...

    arxiv.org/abs/2606.11695 · PDF

  54. 54

    Spectrally Regularized Latent Flow Matching for Turbulence Generation

    Khalid Rafiq, Aditya G. Nair

    cs.LG · physics.flu-dyn

    Latent diffusion and flow matching have emerged as leading approaches for synthetic turbulence generation, yet they systematically under-represent dissipation-range amplitudes. We introduce a latent flow matching framework with a spectrally regularized compression stage that directly targets this failure mode. On a 256^2 DNS dataset at Re_f \approx 2250, replacing an MSE-trained VAE with a zone-weighted log-spectral objective raises...

    arxiv.org/abs/2606.11691 · PDF

  55. 55

    Bergson: An Open Source Library for Data Attribution

    Lucia Quirke, Louis Jaburi, David Johnston, William Z. Li, Gonçalo Paulo, Guillaume Martres, Girish Gupta, Stella...

    cs.LG

    Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging undesirable model behavior and training dataset curation. However, significant engineering effort is required to perform it at scale, and many cutting edge techniques lack open-source tooling and support. Bergson is an open source library that aims to enable faster...

    arxiv.org/abs/2606.11660 · PDF

  56. 56

    Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics

    Katherine Rosenfeld, Maike Sonnewald

    cs.LG · cs.AI

    Generative AI emulators are increasingly used in scientific domains where we already have strong theory, benchmarks, and physical intuition. This raises a central evaluation and interpretability question: when a foundation-style model can reproduce known continuum dynamics, what internal mechanism supports that behavior, is the internal behaviour consistent with known physics, and how does it relate to where the emulator succeeds or fails? We...

    arxiv.org/abs/2606.11657 · PDF

  57. 57

    IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents

    Yifan Yang, Zhen Zhang, Jiayi Tian, Liyan Tan, Zheng Zhang

    cs.LG

    This paper investigates reinforcement learning (RL) methods for improving tool-calling capabilities in multimodal small language model (SLM) agents. While existing works have explored various reward designs to improve agentic tool-calling ability, these approaches face inherent limitations for SLM training, especially under multimodal scenarios. First, many existing methods evaluate tool use correctness through exact matching against certain...

    arxiv.org/abs/2606.11652 · PDF

  58. 58

    DeepRHP: A Hybrid Variational Autoencoder for Designing Random Heteropolymers as Protein Mimics

    Shuni Li, Zhiyuan Ruan, Andy Shen, Ivan Jayapurna, Ting Xu, Haiyan Huang

    cs.LG · q-bio.QM · stat.AP

    Synthetic random heteropolymers (RHPs), consisting of a predefined set of monomers, offer an approach toward the design of protein-like materials. These RHPs, if designed appropriately, can mimic protein behavior and function. As such, there is a need for computational tools to efficiently guide RHP design. We bridge this gap by developing DeepRHP, a modified variational autoencoder (VAE) model under a semi-supervised framework. By equipping...

    arxiv.org/abs/2606.11651 · PDF

  59. 59

    Structure-Preserving Neural Surrogates with Tractable Uncertainty Quantification

    Handi Zhang, Adrienne M. Propp, Brooks Kinch, Houman Owhadi, Nathaniel Trask

    cs.LG · math.NA · physics.comp-ph

    Recent advances in scientific machine learning provide a means of near-real-time solution to partial differential equations (PDEs), but lack the theoretical underpinnings of conventional simulators that support contemporary verification and validation. In this work, we construct data-driven reduced-order models that serve as structure-preserving, real-time surrogates. Remarkably, the exterior calculus that imposes physical conservation...

    arxiv.org/abs/2606.11650 · PDF

  60. 60

    Tree-Structured Orthonormal Decomposition of the Aitchison Simplex

    Daisuke Yamada, Qijun Zhang, Travis Pence, Barbara B. Bendlin, Federico Rey, Vikas Singh

    cs.LG · q-bio.QM · stat.ML

    Compositional data -- vectors encoding relative proportions -- arise across scientific domains, including ecology, geochemistry, and genomics. The features in these data often come with known hierarchical structure (e.g., taxonomies, phylogenies, ontologies), yet existing methods either ignore this structure, discard the intrinsic Aitchison geometry, are designed for binary trees, or yield incomplete coordinate systems. We describe PolyILR, a...

    arxiv.org/abs/2606.11646 · PDF

  61. 61

    TAROT: Task-Adaptive Refinement of LLM-prior Graphs for Few-shot Tabular Learning

    Ruxue Shi, Yili Wang, Mengnan Du, Hangting Ye, Yi Chang, Xin Wang

    cs.LG · cs.AI

    Few-shot tabular learning provides a cost-effective approach for real-world applications where annotation is costly and collecting sufficient samples for new tasks is difficult. Existing Traditional and LLM-based methods have demonstrated effectiveness in few-shot scenarios. However, traditional methods need additional training on unlabeled or generated data, which incur significant computational overhead. In addition, LLM-based methods that...

    arxiv.org/abs/2606.11640 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.