cs.LG · 2026-06-24 · No. 33

Machine Learning, 2026-06-24.

33 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

33 entries
  1. 01

    Real vs. Complex Spectral Bases for Neural Operators: The Role of Green's Function Alignment

    Jason Sulskis, Sathya Ravi

    cs.LG

    Fourier Neural Operators (FNO) learn solution operators of partial differential equations by parameterizing global convolutions in the complex Fourier domain. For real-valued PDE solutions, the complex FFT carries representational redundancy through conjugate symmetry. We introduce the Hartley Neural Operator (HNO), the exact real-valued mirror of FNO: it replaces the FFT with the purely real Discrete Hartley Transform and learns a single...

    arxiv.org/abs/2606.24851 · PDF

  2. 02

    Grad Detect: Gradient-Based Hallucination Detection in LLMs

    Anand Kamat, Daniel Blake, Brent M. Werness

    cs.LG · cs.AI

    Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to generating hallucinations. Detecting these hallucinations is critical for deploying LLMs reliably in high-stakes applications. We present Grad Detect, a gradient-based approach for predicting hallucinations by analyzing layer-wise gradient patterns from a single forward-backward pass during inference. Our method shows that the...

    arxiv.org/abs/2606.24790 · PDF

  3. 03

    FlowPipe: LLM-Enhanced Conditional Generative Flow Networks for Data Preparation Pipeline Construction

    Kunyu Ni, Lei Cao, Jie He, Xiaotong Zhang, Jianfeng Jin, Junyu Dong, Yanwei Yu

    cs.LG · cs.AI

    Data preparation pipelines improve data quality in machine learning by transforming raw tables into learning-ready data through sequential cleaning and feature transformation operators. However, automatically constructing such pipelines is computationally difficult because operator sequences are combinatorial and end-to-end evaluation is expensive. Existing state-of-the-art (SOTA) Multi-DQN methods still face three key limitations: decoupled...

    arxiv.org/abs/2606.24679 · PDF

  4. 04

    QC-SMOTE: Quality-Controlled SMOTE for Imbalanced Classification

    Parth Upman, Shreyank N Gowda

    cs.LG

    Class imbalance poses a significant challenge in classification, where existing methods such as SMOTE often generate low-quality synthetic samples in regions with noise or class overlap. We propose QC-SMOTE, a quality-controlled oversampling framework that estimates minority sample reliability using a composite neighbourhood trustworthiness score combining local density, safe-level, and isolation from the majority class. Synthetic candidates...

    arxiv.org/abs/2606.24625 · PDF

  5. 05

    Reasoning as Attractor Dynamics: Latent Memory Retrieval via Gibbs-Weighted Energy Minimization

    Kanishk Awadhiya

    cs.LG

    Large Language Models (LLMs) are traditionally viewed as autoregressive generators. However, from the perspective of collective computation, they function as high-dimensional Dense Associative Memories that store complex reasoning patterns as latent attractors. In this work, we investigate the energy landscape of mathematical reasoning. We posit that correct reasoning chains correspond to deep, wide attractor basins ("flat minima") in the...

    arxiv.org/abs/2606.24543 · PDF

  6. 06

    A Fair Evaluation of Graph Foundation Models for Node Property Prediction

    Oleg Platonov, Gleb Bazhenov, Dmitry Eremeev, Liudmila Prokhorenkova

    cs.LG · cs.AI · cs.SI

    Due to the wide use of graph-structured data in different fields of industry and science, the development of Graph Foundation Models (GFMs) has recently attracted a lot of attention. While many different types of models are called GFMs, particular interest has been paid to GFMs designed for node property prediction tasks, which is one of the most popular settings in Graph ML with lots of real-world applications from fraud detection in...

    arxiv.org/abs/2606.24509 · PDF

  7. 07

    An LLM-based Two-Stage Transformer Framework for Cross-Domain Bearing Fault Diagnosis with Limited Data

    Jinghan Wang, Feng Cheng, Wentao Wu, Hang Li, Gaoliang Peng, Tianchen Liu

    cs.LG · cs.CL

    Bearing fault diagnosis faces critical challenges when dataset heterogeneity, operating condition variations, and limited labeled data occur simultaneously in industrial environments. Existing approaches address these issues in isolation and rely on implicit feature alignment, limiting effectiveness under concurrent challenges. This paper proposes a knowledge-guided two-stage transfer learning framework that employs a lightweight GPT-2-style...

    arxiv.org/abs/2606.24459 · PDF

  8. 08

    Data Augmentation: A Fourier Analysis Perspective

    Behrooz Tahmasebi, Melanie Weber, Stefanie Jegelka

    cs.LG · stat.ML

    Data augmentation is a simple and model-agnostic approach for exploiting known invariances in learning problems. Given a group acting on the input space, one augments the training set with transformed copies of each sample. Because it exploits symmetries without modifying the underlying learning algorithm, data augmentation can be applied broadly across learning methods. However, this universality comes at a computational cost: when the group...

    arxiv.org/abs/2606.24418 · PDF

  9. 09

    Natural Identifiers for Privacy and Data Audits in Large Language Models

    Lorenzo Rossi, Bartłomiej Marek, Franziska Boenisch, Adam Dziedzic

    cs.LG

    Assessing the privacy of large language models (LLMs) presents significant challenges. In particular, most existing methods for auditing differential privacy require the insertion of specially crafted canary data during training, making them impractical for auditing already-trained models without costly retraining. Additionally, dataset inference, which audits whether a suspect dataset was used to train a model, is infeasible without access...

    arxiv.org/abs/2606.24408 · PDF

  10. 10

    Parallel Manifold Steering: Efficient Adaptation of Large Associative Memories via Residual Energy Shaping

    Kanishk Awadhiya

    cs.LG

    Large Transformer models function as Dense Associative Memories (DAMs), retrieving knowledge via high-dimensional attractor dynamics driven by the self-attention mechanism \citep{ramsauer2020hopfield, wu2024attention}. However, adapting these frozen memory systems to new tasks presents a fundamental ``Plasticity-Stability'' dilemma. Current methods either risk catastrophic interference by modifying synaptic weights directly (e.g., LoRA)...

    arxiv.org/abs/2606.24396 · PDF

  11. 11

    Managing Task Execution for Unknown Workloads in Batteryless IoT: A Hardware-Agnostic Evaluation

    Samer Nasser, Henrique Duarte Moura, Ritesh Kumar Singh, Maarten Weyn, Jeroen Famaey

    cs.LG

    In recent years, the Internet of Things (IoT) paradigm has been shifting toward batteryless, energy-harvesting architectures. Sustaining reliable operation in these systems requires intelligent management of highly volatile stored energy. As edge applications grow in complexity, traditional energy-aware schedulers struggle with unpredictable workloads due to their reliance on static execution thresholds or pre-measured, hardware-specific task...

    arxiv.org/abs/2606.24340 · PDF

  12. 12

    Project Ariadne: Prompt-Conditioned Route Generation for Synthesis Planning

    Anton Morgunov, Victor S. Batista

    cs.LG

    Retrosynthetic planning seeks to connect a target molecule to commercially available starting materials through a multistep route. Classical planners construct such routes by iteratively applying single-step reaction models within a search procedure; constrained variants often require specialized algorithms or architectural changes. Direct route generation reframes retrosynthesis as sequence generation, but existing direct-generation methods...

    arxiv.org/abs/2606.24184 · PDF

  13. 13

    Lightweight Transformer Models for On-Device Fault Detection: A Benchmark Study on Resource-Constrained Deployment

    Disha Patel

    cs.LG · cs.AI

    On-device fault detection enables real-time diagnostics without cloud dependency, but deploying machine learning models on resource-constrained hardware demands careful tradeoffs between accuracy, latency, and model size. We present a benchmark comparing traditional ML methods (Random Forest, XGBoost, SVM, Logistic Regression) against lightweight transformer architectures (DistilBERT, TinyBERT-6L, TinyBERT-4L, MobileBERT) for binary fault...

    arxiv.org/abs/2606.24173 · PDF

  14. 14

    AsyncOPD: How Stale Can On-Policy Distillation Be?

    Wonjun Kang, Kevin Galim, Seunghyuk Oh, Minjun Kang, Sanghyun Park, Donghoon Kim, Minjae Lee, Minseo Kim, Rishabh...

    cs.LG

    On-policy distillation (OPD) trains a student on its own rollouts guided by teacher feedback and is becoming increasingly important for large language model (LLM) post-training. Like reinforcement learning (RL), however, OPD faces an on-policy systems bottleneck, as rollouts can dominate training time for reasoning workloads. Asynchronous training pipelines can alleviate this bottleneck by decoupling rollout generation from learner updates,...

    arxiv.org/abs/2606.24143 · PDF

  15. 15

    A Time-Reparameterized Cumulative Intensity Extrapolation Sampler for Discrete Flow Matching

    Feiyang Fu, Hehe Fan

    cs.LG

    Discrete flow matching (DFM) provides a principled framework for generative modeling on discrete state spaces via continuous-time Markov chain dynamics. In practice, sampling for DFM commonly employs discretizations such as $τ$-leaping, yet efficient sampling methods under a limited number of function evaluations (NFE) remain less studied. To address this gap, we propose the Time-Reparameterized Cumulative Intensity Extrapolation (TR-CIE)...

    arxiv.org/abs/2606.24140 · PDF

  16. 16

    Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning

    Chenhao Dang, Jing Ma, Mingjie Liao

    cs.LG · cs.CL

    The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of Large Language Model (LLM) pre-training. Online Data Mixing (ODM), the technique of adaptively adjusting data mixtures during training, has emerged as a promising direction to improve efficiency. However, existing methods are constrained by their reliance on a singular optimization perspective, which fundamentally overlooks...

    arxiv.org/abs/2606.24133 · PDF

  17. 17

    When Top-1 Fails: Calibrating LoRA Monitors for Masked Diffusion LMs

    Lucky Verma, Pratik Yadav

    cs.LG · cs.CL

    Discrete diffusion language model (DLM) fine-tuning inherits inexpensive diagnostics from denoising-time confidence monitors, but their PEFT-training meaning is untested. We test top-1 argmax concentration as a collapse warning. Across 816 LoRA/PEFT configurations from three DLM families, the warning fires for every configuration while logs record 0/816 actual collapses at the 200 step horizon, giving zero precision. The cause is...

    arxiv.org/abs/2606.24119 · PDF

  18. 18

    FedUP: One-Shot Federated Unlearning via Centroid-Guided Plug-in Filters

    Feihong Nan, Zhengyi Zhong, Pan Wang, Weidong Bao, Xiongtao Zhang, Quan Wen, Ji Wang

    cs.LG

    Federated unlearning (FU) is critical for complying with legal mandates like the right to be forgotten in decentralized systems, yet current methods face a persistent dilemma between non-target knowledge loss and high request latency. To resolve these issues, we propose FedUP, a one-shot federated unlearning framework utilizing lightweight pluggable filters that act as a "knowledge funnel" to screen out target data while preserving original...

    arxiv.org/abs/2606.24113 · PDF

  19. 19

    NeuroSonic: Conditional Flow Matching for EEG-to-Speech Reconstruction

    Wenhao Gao, Yifan Wang, Yijia Ma, Carl Yang, Wen Li, Chenyu You

    cs.LG

    Reconstructing continuous speech from scalp electroencephalography (EEG) remains fundamentally challenging. EEG provides a weak, spatially diffuse, and highly variable measurement of distributed cortical activity, whereas speech is organized as a coherent acoustic trajectory with strong harmonic and temporal structure. The resulting mismatch makes waveform regression unstable and causes stochastic multi-step generation to be sensitive to...

    arxiv.org/abs/2606.24087 · PDF

  20. 20

    Blockwise Policy-Drift Gating for On-Policy Distillation

    Liwen Zheng, Haiyun Jiang

    cs.LG · cs.AI · cs.CL

    On-policy distillation (OPD) trains a student policy using teacher signals computed on trajectories sampled by the student itself. Recent work shows that sampled-token OPD can be fragile on long-horizon reasoning tasks and that local teacher-support matching is a simple and effective repair. This paper introduces blockwise policy-drift gating, a lightweight student-only old-current drift controller for OPD under rollout reuse. The method...

    arxiv.org/abs/2606.24084 · PDF

  21. 21

    RAVEN: A Regime-Aware Variable-context Expert Network for Financial Time Series Forecasting

    Cheng He, Zhenyu Guan, Xijie Liang, Defu Lian, Jiajia Li, Enhong Chen, Patrick P. C. Lee, Geng Hu, Zehao Chen

    cs.LG · cs.AI

    Financial time series forecasting presents structural challenges absent from standard benchmarks. Log-returns are non-stationary, exhibit exceptionally low signal-to-noise (SNR) ratios, and are governed by regime-dependent temporal dependencies. We identify a key limitation of state-of-the-art (SOTA) time series models in financial settings. A fixed context window is mismatched to the time-varying optimal look-back of non-stationary price...

    arxiv.org/abs/2606.24062 · PDF

  22. 22

    Rapid FinFET Modelling Using an Autoencoder

    Amit Sarkar Suman Sau, Swagata Mandal

    cs.LG · cs.AI · physics.app-ph

    This work presents a machine learning framework that leverages an autoencoder (AE) for the efficient modeling of FinFET. We first calibrated a BSIM-CMG model to generate a dataset of current-voltage (ID-VG) characteristics. This data was used to train an autoencoder that compresses full I-V curves into a low-dimensional latent space, which intrinsically encodes key device physics. A key innovation is the explicit incorporation of parameter...

    arxiv.org/abs/2606.24046 · PDF

  23. 23

    RoPE-Aware Bit Allocation for KV-Cache Quantization

    Fengfeng Liang, Yuechen Zhang, Jiaya Jia

    cs.LG · cs.CL

    Existing low-bit KV-cache quantizers often treat each cached key as a flat vector. Under RoPE, however, a key's contribution to a future attention logit decomposes into a position-dependent sum over two-dimensional frequency blocks. This makes key-cache quantization a block-wise bit-allocation problem: high-energy RoPE blocks are more sensitive to quantization error and should receive more bits. We introduce Block-GTQ, a RoPE-aware bit...

    arxiv.org/abs/2606.24033 · PDF

  24. 24

    Information-Theoretic Classifier-Free Guidance with Adaptive Schedule Optimization

    Haobo Chen, Xiangxiang Xu, Yuheng Bu

    cs.LG

    Diffusion models have achieved strong performance in image, text-to-image, and video generation, where conditional generation is often controlled by classifier-free guidance (CFG). CFG improves condition consistency by increasing a guidance weight, but stronger guidance typically reduces diversity and distributional coverage. It remains unclear how this consistency-coverage trade-off should be controlled across the reverse trajectory, since...

    arxiv.org/abs/2606.24025 · PDF

  25. 25

    You Don't Need to Run Every Eval

    Yuchen Zeng, Dimitris Papailiopoulos

    cs.LG

    A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare design choices, and select the checkpoint for the release. But do we need to run every eval? We compile a public score matrix of 84 frontier models on 133 benchmarks (2,604 cells, 23.3% filled) and find it is approximately rank-2: a model's scores across all 133 benchmarks are largely...

    arxiv.org/abs/2606.24020 · PDF

  26. 26

    Fast and Slow Variational Continual Learning

    Subarnaduti Paul, Yohan Jung, Mohammad Emtiyaz Khan, Siddharth Swaroop, Thomas Möllenhoff, Martin Mundt

    cs.LG · cs.AI

    Continual learning remains a major challenge for modern deep networks, partly because commonly used optimizers lack inherent mechanisms for continual adaptation. One such natural mechanism is fast and slow adaptation to balance stability and plasticity. This mechanism has deep roots in neuroscience and biology, but there is no consensus on how to best incorporate it in commonly used optimizers. Here, we show that this can be easily done via...

    arxiv.org/abs/2606.24007 · PDF

  27. 27

    Cyclic Denoising Reveals Ultrastable Memories in Diffusion Models

    Rishabh Sharma, Stefano Martiniani

    cs.LG · cond-mat.dis-nn · cs.CR · cs.CV

    We introduce cyclic denoising -- repeated forward and reverse diffusion at controlled noise amplitudes -- as an extraction attack for image diffusion models. Inspired by random organization in disordered solids, cyclic denoising exposes regions of the learned distribution that are largely inaccessible to standard sampling. The dynamics drive samples toward attractors with a broad stability spectrum. The deepest attractors are ultrastable:...

    arxiv.org/abs/2606.24000 · PDF

  28. 28

    EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

    Tristan Maidment, JB Lanier, Chase McDonald, Nathan Tsang, Eugene Vinitsky, Roy Fox, Albert Wang, Wesley N. Kerr

    cs.LG · cs.AI · cs.GT · cs.MA

    Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized game-theoretic algorithms for solving two-player zero-sum imperfect-information games. The uniform distribution has emerged as a strong policy regularization target for this purpose, but it regularizes equally toward all actions regardless of their viability. We introduce EMAgnet, which instead regularizes...

    arxiv.org/abs/2606.23995 · PDF

  29. 29

    Learning to Trigger: Reinforcement Learning at the Large Hadron Collider

    Zixin Ding, Shaghayegh Emam, Giovanna Salvi, Cecilia Tosciri, Abhijith Gandrakota, Jennifer Ngadiuba, Nhan Tran,...

    cs.LG · cs.AI · hep-ex

    High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (\textit{triggering}) under tight constraints on bandwidth, latency, and storage. In practice, trigger menus are largely static and hand-tuned and can become suboptimal as detector conditions, pileup, and background composition drift over time. We cast online threshold tuning as a sequential decision-making problem: a reinforcement...

    arxiv.org/abs/2606.23993 · PDF

  30. 30

    Offline Reinforcement Learning for Warehouse SLAM Throughput Control

    Tina Dongxu Li, Mouhacine Benosman, Rajat Kumar, Kevin Tan, Ken Meszaros, Trevor Dardik

    cs.LG · cs.AI

    We present an offline reinforcement learning (RL) framework for optimizing SLAM throughput control in a warehouse fulfillment environment. SLAM (Scan/Label/Apply/Manifest) throughput directly influences system congestion and operational efficiency. Our RL-based control approach dynamically recommends SLAM throughput settings that adaptively balance throughput maximization with downstream stability through intelligent adjustment of throttling...

    arxiv.org/abs/2606.23978 · PDF

  31. 31

    A Comparative Study of Bayesian Contextual Bandits for Real-Time Warehouse Sorter Optimization

    Tina Dongxu Li, Mouhacine Benosman, Ken Meszaros, Trevor Dardik

    cs.LG · eess.SY

    Efficient sorter diversion control of automated material handling systems (MHS) is critical for optimizing operational efficiency in large-scale warehouse environments. In this study, we use an inbound receiving sorter at a high-volume e-commerce warehouse as our primary use case, where the sorter diversion system relies on cost functions with static weight configurations that fail to adapt to highly dynamic system contexts, such as volume...

    arxiv.org/abs/2606.23977 · PDF

  32. 32

    3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy

    Amirhossein Kardoost, Lion Gleiter, Tingying Peng, Carsten Marr

    cs.LG · cs.CV · q-bio.QM

    Self-supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three-dimensional nature of cells. We present a systematic comparison of 2D and 3D masked autoencoders (MAE-2D vs. MAE-3D) on volumetric microscopy data. Under matched architectures and training protocols, MAE-3D consistently outperforms 2D max-projection and slice-based variants on downstream single-cell tasks. We further align visual...

    arxiv.org/abs/2606.23964 · PDF

  33. 33

    Forget Without Compromise: Nexus Sampling for Streaming KV-Cache Eviction Under Fixed Budgets

    Duc Duong, Hoang Anh Duy Le, Jianwen Xie, Anshumali Shrivastava, Zhaozhuo Xu

    cs.LG

    Long-context and agentic LLM workloads push the KV cache past any fixed memory budget, forcing the inference stack to permanently evict tokens at every step of a continuous-inference stream. Existing methods all share the same template, a per-step direct-attention score followed by deterministic top-$K$ selection, which converts a single below-cutoff step into an irreversible verdict and permanently erases any subtly important token that...

    arxiv.org/abs/2606.23961 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.