cs.LG · 2026-06-23 · No. 32

Machine Learning, 2026-06-23.

76 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

76 entries
  1. 01

    Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?

    Dingzhi Yu, Hongyi Tao, Yuanyu Wan, Luo Luo, Lijun Zhang

    cs.LG · cs.AI · math.OC · stat.ML

    AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostly in finite-variance regimes. This is increasingly unsatisfying, as empirical evidence indicates that stochastic gradient noise in LLM pretraining is typically heavy-tailed. Recent work shows that sign-based optimizers such as Lion and Muon achieve sharp heavy-tailed rates, and that AdaGrad can also converge under heavy-tailed...

    arxiv.org/abs/2606.23676 · PDF

  2. 02

    Tapered Language Models

    Reza Bayat, Ali Behrouz, Aaron Courville

    cs.LG · cs.AI · cs.CL

    Modern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack of identical layers in which parameters are allocated uniformly across depth. This is a default inherited from the original transformer and largely unchanged since, yet a growing body of evidence suggests that layers contribute non-uniformly to the final output, with later layers refining the residual stream rather than...

    arxiv.org/abs/2606.23670 · PDF

  3. 03

    On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners

    David Mguni, Julian Ma, Jun Wang

    cs.LG

    Large Language Models (LLMs) are frequently portrayed as general-purpose solvers capable of solving arbitrary tasks. We argue that this view overlooks a fundamental constraint: language is a compressed and capacity-limited interface for conveying task information. Modelling User--System interaction as a bilevel \emph{cheap-talk} game, we analyse how latent tasks are encoded into prompts and reinterpreted under alignment and safety...

    arxiv.org/abs/2606.23668 · PDF

  4. 04

    MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems?

    Juyang Bai, Laixi Shi

    cs.LG · cs.MA

    Multi-agent systems (MAS) offer a scalable path forward for agentic AI, comprising multiple LLM-based agents, each assigned a system prompt and a position within a workflow that governs inter-agent coordination and output aggregation. System prompts thus form a critical and accessible optimization surface: they specify agents' roles and behaviors, enabling system-level improvements without model finetuning. Although prompt optimization has...

    arxiv.org/abs/2606.23664 · PDF

  5. 05

    Dynamic estimation of slowly varying sequences

    Prashant Gokhale, Mikhail Khodak, Sandeep Silwal

    cs.LG · cs.DS

    We consider the problem of sequentially approximating functions of each element in a slowly-varying sequence, i.e. one where the magnitude $α_i$ of the difference between the elements at positions $i$ and $i-1$ is small. Recent work on implicit trace estimation shows that when $α_t$ is small, reusing queries to past sequence elements can reduce the overall cost [Dharangutte \& Musco, NeurIPS~2021; Woodruff et al., NeurIPS~2022]. We introduce...

    arxiv.org/abs/2606.23655 · PDF

  6. 06

    Learning Process Rewards via Success Visitation Matching for Efficient RL

    Raymond Tsao, Andrew Wagenmaker, Sergey Levine

    cs.LG · cs.AI · cs.RO · stat.ML

    In many modern applications of reinforcement learning (RL), the natural reward for a task of interest is inherently sparse: a reward of 0 is given everywhere except when the task is completed, when a reward of +1 is given. Training a policy to maximize such a sparse reward requires solving a challenging credit assignment problem, leading to slow or ineffective RL improvement. We propose a simple approach to transform a sparse outcome reward...

    arxiv.org/abs/2606.23640 · PDF

  7. 07

    Muown Implicitly Performs Angular Step-size Decay

    Florian Hübler, Kai Lion, Antonio Orvieto, Niao He

    cs.LG · math.OC

    Matrix-aware optimizers such as Muon and Muown have recently shown strong empirical performance for pre-training Transformers. In particular, Muown separates each weight matrix into row magnitudes and an un-normalized direction variable, updating the former with Adam and the latter with Muon. We show that the directional update of Muown is equivalent to a Riemannian step on the normalized directions, while the magnitude of the un-normalized...

    arxiv.org/abs/2606.23637 · PDF

  8. 08

    DiT-Reward: Generative Representations for Text-to-Image Reward Modeling

    Yuanming Yang, Guoqing Ma, Bo Wang, Yuan Zhang, Wei Tang, Chenyi Li, Haoyang Huang, Nan Duan

    cs.LG · cs.AI

    Can representations learned for image generation also support the evaluation of generated images? We study text-to-image reward prediction as a downstream task of generative representation learning. To this end, we introduce DiT-Reward, which converts a pretrained text-to-image Diffusion Transformer into a reward model by processing near-clean image latents and aggregating text-conditioned image representations across transformer layers....

    arxiv.org/abs/2606.23626 · PDF

  9. 09

    Discovering Latent Groups for Robust Classification

    Ankur Garg, Ulrich Aïvodji, Samira Ebrahimi Kahou, Vincent Michalski

    cs.LG · cs.AI · cs.CV

    Machine learning models exploit spurious correlations, achieving high average accuracy but failing disproportionately on underrepresented subgroups. Existing methods address this by adjusting network parameters, guided either by subgroup annotations or inferred pseudo-group labels. Yet at inference, these methods produce only a class prediction, with no insight into a sample's latent subgroup. We propose neural classification trees (NCT), a...

    arxiv.org/abs/2606.23609 · PDF

  10. 10

    Scaling Linear Mode Connectivity and Merging to Billion Parameter Pretrained Transformers

    Tianyi Li, Zhiqiang Shen

    cs.LG · cs.AI

    Linear mode connectivity (LMC) provides a promising foundation for understanding and merging independently trained neural networks, but existing methods typically optimize the interpolation path from only one model endpoint, limiting their scalability and effectiveness for large pretrained transformers. We propose a novel and scalable framework for enabling LMC-based model merging to {\em billion-parameter pretrained transformers}. Our method...

    arxiv.org/abs/2606.23607 · PDF

  11. 11

    MORL-A2C: Multi-Objective Reinforcement Learning Reranker for Optimizing Healthiness in MOPI-HFRS

    Aarya Vasantlal, Joshua Zolla

    cs.LG

    Unhealthy dietary behavior continues to be a persistent public health issue in the United States, exacerbated by recommendation systems that prioritize user preference without considering nutritional health. The Multi-Objective Personalized Interpretable Health-aware Food Recommendation System (MOPI-HFRS), from which this work extends, addresses this by jointly optimizing preference, health, and diversity through Pareto-based optimization....

    arxiv.org/abs/2606.23603 · PDF

  12. 12

    Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior

    Christopher J. Anders, Henrique Da Silva Gameiro, Nico Daheim, Mohammad Emtiyaz Khan

    cs.LG

    One way to understand LLM behavior is to trace its output back to the training data. Two types of measures are commonly used for output tracing: data-similarity and data-influence. The former is cheaper while the latter is believed to be more accurate. Even though many works have compared them for ground-truth tasks, no such comparisons exist for output tracing. Here, we fill this gap and precisely quantify the commonalities and differences...

    arxiv.org/abs/2606.23591 · PDF

  13. 13

    It's Much Easier for Neural Networks to learn Game of Life Dynamics with the Right Activation Function: Polynomial Kolmogorov-Arnold Networks

    Tashin Ahmed, Q. Tyrell Davis

    cs.LG · cs.NE · nlin.CG

    Previous work has found a gap between the scale of neural networks that reliably learn Conway's Game of Life, and minimal networks capable of representing the classic cellular automaton with hard-coded parameter values. Viewing neural network learning as a search process suggests a dependence on networks large enough to contain sub-networks with lucky initializations (sometimes known as 'winning tickets') that actually learn the task. In this...

    arxiv.org/abs/2606.23587 · PDF

  14. 14

    Solve for the Hyperparameter, Skip the Search: Kolmogorov-Optimal Scaling Laws for Spline Regression

    Yong Yi Bay, Kathleen A. Yearick

    cs.LG · cs.AI · math.NA · stat.ML

    Hyperparameter tuning almost always means search: fit the model at every value on a grid, score each by cross-validation, and keep the winner. For spline regression that search is unnecessary. The optimal resolution can be solved for in closed form, to the accuracy an exhaustive search reaches, at a fraction of the compute. Three ingredients make this possible: classical approximation theory pins the squared bias to a known power of the...

    arxiv.org/abs/2606.23575 · PDF

  15. 15

    A Spectral Theory of Normalized Corrected GNN Propagation

    Qihan Chen, Wei Li, Meng Qin, Jianfeng Hou

    cs.LG

    We develop a spectral theory for \emph{normalized corrected GNN propagation}. The object of study is the symmetric normalized adjacency with its degree-stationary component removed, matching the normalization used by standard GCN-style models while isolating the stationary direction most directly tied to oversmoothing. The central theoretical question is whether this corrected normalized operator preserves class-discriminative signal after...

    arxiv.org/abs/2606.23572 · PDF

  16. 16

    Patient-Aware Contrastive Learning Preserves Per-Patient Structure in RR-Interval Representations

    Yasantha Niroshana, Weijith Wimalasiri, Chathuranga Hettiarachchi

    cs.LG

    Contrastive representation learning struggles on physiological signals when each subject contributes a distinct baseline pattern. If class differences overlap with subject differences,class-level objectives such as supervised contrastive learning tend to merge per-subject structure into a single per-class cluster,removing the individual variation that a model needs to generalize to unseen patients. We study this problem in the setting of...

    arxiv.org/abs/2606.23570 · PDF

  17. 17

    SVD-Surgeon: Optimal Singular-Value Surgery for Large Language Model Compression

    Mahmoud Safari, Frank Hutter

    cs.LG · cs.CL

    Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their deployment is constrained by substantial memory and compute requirements. Low-rank compression via singular value decomposition (SVD) is an effective remedy, but existing methods focus on how to factorize and which components to keep. We introduce SVD-Surgeon, a training-free method that brings the Optimal Brain Surgeon (OBS) framework to the...

    arxiv.org/abs/2606.23568 · PDF

  18. 18

    Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models

    Jiawei Xu, Minghui Liu, Aakriti Agrawal, Yifan Chen, Furong Huang

    cs.LG · cs.AI

    Masked diffusion language models decode by iteratively unmasking tokens, where the unmasking order defines an "order of thought" that strongly influences generation quality yet is typically chosen heuristically. We derive a tractable upper bound on the sequential decoding mismatch, measured by the Kullback-Leibler divergence and expressed in terms of the model's pathwise log-likelihood, with tightness under sufficient model expressivity. This...

    arxiv.org/abs/2606.23567 · PDF

  19. 19

    The Energy Consumption of Transformer Fine-Tuning: A Roofline-Inspired Scaling Model

    Mansour Zoubeirou a Mayaki

    cs.LG · cs.AI · cs.AR · cs.CL · cs.DC

    Transformer-based models underpin modern natural language processing but incur rapidly growing computational and energy costs. As training scales in both model size and parallelism, accurately predicting energy consumption has become critical for sustainable and cost-aware system design. We present a framework for modeling the energy consumption of Transformer training on multiple GPUs. Using controlled architectural sweeps of BERT models, we...

    arxiv.org/abs/2606.23546 · PDF

  20. 20

    Simulation-Free Estimation of Traffic Flows from Sparse Count Data

    Davide Guastella, Gianluca Bontempi

    cs.LG

    We propose a method for estimating time-varying traffic flow patterns from sparse aggregated vehicle counts. The method partitions the study area into spatial regions, constructs a set of feasible region-to-region routes, and solves a weighted least-squares optimization problem to determine the number of vehicles to allocate on each route. A weighted contribution matrix encodes sensor coverage, steering the optimizer toward flow...

    arxiv.org/abs/2606.23536 · PDF

  21. 21

    Collapsed Effective Operators for Higher-order Structures

    Maximilian Krahn, Lennart Bastian, Vikas Garg, Björn Schuller, Tolga Birdal

    cs.LG · stat.ML

    Higher-order structures are powerful relational modeling tools, yet existing spectral operators decompose the topology into separate ranks, leaving practitioners to fuse the information back to vertices through ad hoc choices. We introduce Collapsed Effective Operators, which condense higher-order degrees of freedom into a single vertex-level operator via Schur complementation of a graded Laplacian. This yields a (generally dense) operator...

    arxiv.org/abs/2606.23517 · PDF

  22. 22

    TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization

    Matan Ben-Tov, Mahmood Sharif

    cs.LG · cs.CR

    Discrete text-trigger optimization -- searching for text sequences that, when ingested by a model, steer it toward a specified objective -- underpins model red-teaming (e.g., LLM jailbreaks), as well as auditing and interpretability. However, the current state of discrete optimizers hinders their adoption and progress. First, existing optimizers, when open-sourced at all, are scattered across research codebases tied to specific models,...

    arxiv.org/abs/2606.23496 · PDF

  23. 23

    Do Location Encoders Capture Spatial Effects? A GeoShapley Benchmark Across Scales

    Daniel Kiv, Shaowen Wang

    cs.LG

    Location encoders transform geographic coordinates into high dimensional embeddings for downstream machine learning, but it is unclear how well these representations capture interpretable spatial effects. We benchmark whether GeoShapley, a game-theoretic explainer that treats all location features as a single joint player, can recover spatially varying coefficients from models built on location-encoder embeddings. Eleven encoders from the...

    arxiv.org/abs/2606.23453 · PDF

  24. 24

    Selective Time Series Forecasting via Metalearning

    Ricardo Inácio, Vitor Cerqueira, Marília Barandas, Carlos Soares

    cs.LG

    Deep learning methods have achieved state-of-the-art in time series forecasting, yet their accuracy varies considerably across samples, as some instances remain inherently difficult to predict. Reject option mechanisms, which allow models to abstain from high-risk predictions, are well established in classification and regression but underexplored in forecasting. Existing abstention strategies typically rely on proxies, such as the width of...

    arxiv.org/abs/2606.23448 · PDF

  25. 25

    What Does a Chemical Language Model Know About Molecules?

    Christian Kenneth, Etowah Adams, Liam Bai, Gerard JP van Westen

    cs.LG · cs.AI · physics.chem-ph

    Chemical language models (cLMs) are widely assumed to learn surface-level syntactic patterns rather than learning meaningful molecular semantics. Here, we apply sparse autoencoders (SAEs) to MolFormer, an encoder-only cLM, to mechanistically examine how molecular representations are built across layers. We discover that early layers rely on position-tracking latents to parse molecular grammar, while later layers encode atom-in-substructure...

    arxiv.org/abs/2606.23443 · PDF

  26. 26

    Interpretable Kolmogorov-Arnold Network with Feature-Isolated Temporal Attention Mechanism for Electricity Load Forecasting

    Jinhao Li, Hao Wang

    cs.LG · cs.CE · cs.SC

    Accurate electricity load forecasting is a crucial prerequisite for stable power system operations. While prevalent deep learning models present competitive performance, they often operate as black boxes and lack interpretability. While the Kolmogorov-Arnold network (KAN) has emerged as a promising alternative because of its learnable activation function design, its direct application to time-series forecasting faces challenges in modeling...

    arxiv.org/abs/2606.23425 · PDF

  27. 27

    GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation

    Jette Oberländer, Jan Finkbeiner, Catherine M. Schöfmann, Emre Neftci

    cs.LG · cs.AI

    Autoregressive decoding with LLMs is primarily bottlenecked by GPU memory bandwidth, especially in edge-computing settings. While quantization is essential for mitigating this bottleneck, most existing methods treat inference as a uniform process and fail to account for the asymmetry between the compute-bound prefill stage and the memory-bound decoding stage. We propose GRINQH (GRaded INput-based Quantization Hierarchy), a weight-only...

    arxiv.org/abs/2606.23419 · PDF

  28. 28

    Leveraging Similarities in Multi-Armed Bandits

    Khaled Eldowa, Thibaud Rahier, Augustin Cablant, Panayotis Mertikopoulos, Pierre Gaillard

    cs.LG

    In many online learning and bandit problems, the actions we consider possess inherent similarities--for instance because they share latent traits, tags, or hierarchical structure. We study online learning with a similarity-structured action set, encoded by a rooted tree whose leaves are the actions and whose levels quantify how closely two actions are related. The loss sequence is assumed tree-compatible: losses of similar actions are...

    arxiv.org/abs/2606.23414 · PDF

  29. 29

    Differential Spectral Damping Gap Adaptive Regularization for Ill-Conditioned Kernel Methods

    Praveg Vashishtha

    cs.LG

    Kernel methods requiring matrix inversion -- particularly Least-Squares Twin Support Vector Machines (LSTSVM) -- suffer from exponential eigenvalue decay in their system matrices, producing severely ill-conditioned problems where standard Tikhonov regularization applies uniform damping regardless of eigenvector reliability. We propose Differential Spectral Damping (DSD), a regularization formula that adapts its penalty to localized eigengap...

    arxiv.org/abs/2606.23407 · PDF

  30. 30

    HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models

    Yuval Domb, Hadar Sackstein, Tomer Solberg

    cs.LG · cs.AI

    We present HyperQuant (Hadamard, optimallY Packing, Entropy Rice-coding), a unified post-training quantization pipeline for the weights and the KV cache of large language and diffusion transformers. Across a suite of self-contained experiments (Table 1), HyperQuant outperforms the recent HIGGS scheme at every operating point from 3 to 5 bits per scalar (bps) on weights, and beats both TurboQuant and OCTOPUS on KV quantization down to 1.7 bps....

    arxiv.org/abs/2606.23406 · PDF

  31. 31

    Physics-Informed Modeling for Wood Thermal Analysis and Prediction

    Jingren Xie, Alex John Buckthal, Ryan Anthony O'Connor, Isak Worre Foged, Dim P. Papadopoulos

    cs.LG

    Wood materials exhibit complex, spatially varying thermal properties that challenge traditional architectural assumptions of material homogeneity. Although data-driven approaches can directly map wood RGB images to their corresponding thermal responses, they operate as uninterpretable black boxes that prioritize statistical correlation and may absorb experimental noise rather than thermodynamic plausibility. To address these limitations, we...

    arxiv.org/abs/2606.23402 · PDF

  32. 32

    Distribution-Aware Diffusion-LLM for Robust Ultra-Long-Term Time Series Forecasting

    Falguni Ghosh, Vahid Hashemi, Bernhard Kainz

    cs.LG · cs.AI

    Time series forecasting is a fundamental machine learning task. Recent work has explored Large Language Models (LLMs) for this purpose due to their strong generalization, pattern recognition, and zero-shot or few-shot capabilities. Despite their suitability for long-context learning, LLMs face challenges in multimodal settings: they lack calibrated probabilistic modeling for non-text data and struggle to align heterogeneous representations....

    arxiv.org/abs/2606.23391 · PDF

  33. 33

    Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime

    Yuqing Wang

    cs.LG · math.DS · math.OC

    Training dynamics is central to understanding neural networks, yet its theoretical analysis remains difficult even for simple architectures and becomes substantially more challenging for general modern architectures. In this paper, we propose a convergence framework for analyzing gradient descent (GD) dynamics under a broad family of neural network architectures and datasets beyond the neural tangent kernel (NTK) regime. The framework is...

    arxiv.org/abs/2606.23364 · PDF

  34. 34

    Rethinking Molecular Graph Backdoors under Chemistry-aware Admission

    Thinh T. H. Nguyen, Sze Jue Yang, Khoa D. Doan, Chee Seng Chan, Kok-Seng Wong

    cs.LG · cs.AI

    Backdoor attacks on molecular graph neural networks (GNNs) are typically evaluated as abstract graph edits, but real molecular learning pipelines do not train on arbitrary graphs. Molecular records must first survive parsing, sanitization, canonicalization, and graph-string consistency checks. We formalize this overlooked admission stage as ChemGuard, an operational protocol for testing whether a submitted molecular record can enter a...

    arxiv.org/abs/2606.23361 · PDF

  35. 35

    Adaptive Hard-Soft Physics-Informed Neural Networks for Robust Boundary-Constrained PDE Solving

    Duc Tien Nguyen, Trinh Minh Tuan, Nguyen Duc Manh, Vu Linh Nguyen, Dinh Gia Ninh

    cs.LG · cs.AI

    Physics-informed neural networks (PINNs) provide an effective way to solve partial differential equations (PDEs) by embedding physical principles into the learning process. However, the conventional PINN formulation, in which all constraints are imposed as soft penalty terms within a composite loss, often exhibits slow convergence, sensitivity to loss weight scaling, and inaccurate boundary enforcement due to poor conditioning of the...

    arxiv.org/abs/2606.23359 · PDF

  36. 36

    SOAP-Bubbles: Structured Weight Uncertainty for Neural Networks

    Adrian Robert Minut, Nico Daheim, Marco Miani, Mohammad Emtiyaz Khan, Wu Lin, Thomas Möllenhoff

    cs.LG

    Structured weight-uncertainty can improve many aspects of deep learning, but it remains costly to estimate and difficult to implement. Here, we show that these issues can be addressed by adapting the SOAP optimizer. Our key idea is to run IVON, an existing diagonal-covariance variational method, in the eigenspace of SOAP's preconditioner and then use the preconditioner to transform the diagonal estimate into a non-diagonal covariance. The...

    arxiv.org/abs/2606.23357 · PDF

  37. 37

    Superhuman AI for Generals.io Using Self-Play Reinforcement Learning

    Matej Straka, Viliam Lisý, Martin Schmid

    cs.LG

    We present a superhuman AI agent for Generals.io, a real-time strategy game that requires both long-horizon planning and short-term tactics under strong imperfect information. Trained for four days on 4x NVIDIA H200 GPUs, our agent reaches #1 on the public 1v1 leaderboard of over 5,000 human players, leading the second-ranked player by the same margin that separates second place from 25th, and beats the two top-ranked humans head-to-head with...

    arxiv.org/abs/2606.23348 · PDF

  38. 38

    GRIMIP: A General Framework for Instance-Specific Configuration of MIP Solvers Using LLMs

    Yidong Luo, Xuemin Chen, Chenguang Wang, Fangzhou Zhu, Tao Zhong, Tianshu Yu

    cs.LG

    Configuring the hyperparameters of Mixed-integer programming (MIP) solvers is a high-dimensional, instance-dependent optimization problem where suboptimal settings can degrade solving time by orders of magnitude. Default configurations are often suboptimal, while traditional tuning methods either suffer from the ``cold-start'' problem and inefficient search or heavily rely on expert experience. This paper introduces \textbf{GRIMIP}...

    arxiv.org/abs/2606.23299 · PDF

  39. 39

    Non-asymptotic estimates of the minimal risk in statistical learning

    Liming Wu, Sen Yang

    cs.LG · math.PR · stat.ML

    In this paper we prove some concentration inequalities for two types of error probabilities in the Empirical Risk Principle (ERP) in statistical learning, which provide a lower bound and an upper bound for the minimal risk (in terms of the minimal empirical risk) with non-asymptotic high confidence. The usual boundedness condition of the empirical risk function is relaxed to the Gaussian or exponential integrability condition. The confidence...

    arxiv.org/abs/2606.23295 · PDF

  40. 40

    Exposing the Illusion of Erasure in Knowledge Editing for LLMs

    Advik Raj Basani, Anshuman Chhabra

    cs.LG · cs.AI · cs.CR

    Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understood. In this work, we examine KE from an adversarial elicitation perspective, revealing that edited knowledge is often not fully erased and continues to surface, with consistent failures observed across diverse model architectures. To explain this behavior, we...

    arxiv.org/abs/2606.23276 · PDF

  41. 41

    Dynamic multi-agent deep reinforcement learning-based pricing and incentivization approach in multimodal transportation networks

    Khadidja Kadem, Mostafa Ameli, Carlos Lima Azevedo, Mahdi Zargayouna, Latifa Oukhellou

    cs.LG · cs.AI · math.OC

    In multimodal transportation systems, shared mobility services (SMSs) are promoted for their potential to enhance flexibility and reduce congestion. However, SMS demand is often concentrated in high-density areas, which can limit the effectiveness and accessibility for various commuter groups. This uneven integration challenges transportation system efficiency, especially in terms of emissions and spatial equity. Addressing these issues...

    arxiv.org/abs/2606.23257 · PDF

  42. 42

    Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio

    Ran Piao, Tsai-Ning Wang, Martijn den Dekker, Linda Moonen, Hareld Kemps, Yuan Lu, Aaqib Saeed

    cs.LG · cs.SD

    Clinical audio diagnosis in low-resource settings requires models that identify conditions from minimal examples without large annotated corpora. We propose Federated Self-Contextualization (FSC), a multimodal language model framework for in-context clinical audio diagnosis across federated hospital clients. FSC constructs pseudo-label episodes via unsupervised clustering of audio representations, bypassing scarce real diagnostic labels, and...

    arxiv.org/abs/2606.23243 · PDF

  43. 43

    Deep learning-based detection of cessation of breathing in pre-term infants

    Dineo Serame, Lionel Tarassenko, Mauricio Villarroel

    cs.LG

    Apnoea of prematurity is characterised by recurrent episodes of cessation of breathing and remains difficult to detect reliably using routinely monitored physiological signals in the Neonatal Intensive Care Unit (NICU). Existing bedside monitors rely primarily on respiratory rate and oxygen saturation thresholds, often generating high false-positive alarm rates and missing short or irregular events. Improving automated detection using...

    arxiv.org/abs/2606.23213 · PDF

  44. 44

    Efficient Network Inference via Hardware-Aware Architecture Search, Model Pruning & Quantization

    Lucas Heublein, Mark Deutel, Axel Plinge, Felix Ott

    cs.LG · cs.DC · eess.SP

    Embedded global navigation satellite system (GNSS) interference monitoring requires fast and memory-efficient inference to process large volumes of raw in-phase and quadrature (IQ) samples in real time. At the same time, increasingly expressive deep neural networks (DNNs) are needed for robust interference classification and characterization across diverse signal conditions. This creates a fundamental tension between predictive performance...

    arxiv.org/abs/2606.23210 · PDF

  45. 45

    Leveraging AutoML for Sustainable Deep Learning: A Multi-Objective HPO Approach on Deep Shift Neural Networks

    Leona Hennig, Marius Lindauer

    cs.LG

    Deep Learning (DL) has advanced various fields by extracting complex patterns from large datasets. However, the computational demands of DL models pose environmental and resource challenges. Deep Shift Neural Networks (DSNNs) present a solution by leveraging shift operations to reduce computational complexity at inference. Compared to common DNNs, DSNNs are still less well understood and less well optimized. By leveraging AutoML techniques,...

    arxiv.org/abs/2606.23208 · PDF

  46. 46

    Bridge the Gaps: Heterogeneous Attributed Graph Clustering via Quaternion Representation Learning

    Xinxi Chen, Junyang Chen, Yiqun Zhang, Chuangming Qiu, Xiang Zhang

    cs.LG

    Attributed graph clustering partitions nodes by jointly exploiting node attributes and graph topology. It remains challenging due to attribute heterogeneity and representation degradation during graph learning. Real-world datasets often contain heterogeneous attributes, i.e., numerical and categorical attributes, complicating unified representation learning. This challenge becomes more complex in attributed graphs, where constructing a...

    arxiv.org/abs/2606.23199 · PDF

  47. 47

    Memory Contagion: Cross-Temporal Propagation of Evaluator Bias via Agent Memory

    Zewen Liu

    cs.LG · cs.AI · cs.CL

    Large Language Model (LLM) agents increasingly rely on memory systems to maintain long-term coherence. Recent work shows that agent memories degrade during continuous consolidation. However, existing research assumes memories are derived from unbiased experiences. In this work, we identify and formalize a novel phenomenon: Memory Contagion -- the cross-temporal propagation of evaluator bias through agent memory. We show that when agents are...

    arxiv.org/abs/2606.23195 · PDF

  48. 48

    Stage-dependent integer-binary encoding in factorization-machine black-box optimization

    Ryo Ogawa, Mayumi Nakano, Yuya Seki, Shu Tanaka

    cs.LG

    Black-box optimization (BBO) deals with problems where objective functions lack explicit analytical forms and are expensive to evaluate. Factorization machine with quadratic-optimization annealing (FMQA) constructs a surrogate model using a factorization machine (FM) and optimizes it with an Ising machine. Conventional FMQA applies a single integer-binary encoding throughout the optimization process, although the encoding best suited to...

    arxiv.org/abs/2606.23188 · PDF

  49. 49

    EML Trees Are Universal Approximators

    Joe Germany, Elie Abdo, Joseph Bakarji

    cs.LG · cs.NE · cs.SC · math.NA

    The recently introduced EML (Exp-Minus-Log) function acts as continuous analogue of NAND gates, providing a compositional building block capable of representing elementary functions. In this work, we study the expressive power of tree-structured compositions of EML functions. We show that such trees enjoy a universal approximation property for functions in $W^{k, \infty}$ for $k \in \mathbb N$, drawing on classical neural network...

    arxiv.org/abs/2606.23179 · PDF

  50. 50

    Position: Correct Answer, Wrong Mechanism -- When AI Scientists Defend General Claims Their Own Data Contradicts

    Steven Young Eulig

    cs.LG

    AI scientist systems are described as tools, coauthors, or founders, but we evaluate them as if only the final answer matters. This position paper argues that outcome-only evaluation is insufficient, and that task outcome, mechanism fidelity, and epistemic honesty must be measured separately. Our evidence comes from 28 episodes of a coding agent attempting to rediscover a known particle identification observable in a Geant4 simulation,...

    arxiv.org/abs/2606.23175 · PDF

  51. 51

    Substitution-Based Analysis of Structural Novelty for Generative Models of Materials

    Masahiro Negishi, Aron Walsh

    cs.LG · cond-mat.mtrl-sci

    There has been rapid progress in generative artificial intelligence (AI) models for inorganic crystal design, which can efficiently generate large numbers of candidate compounds after being trained on databases of known crystals. However, it remains unclear whether they genuinely expand the accessible materials search space beyond conventional strategies such as elemental substitution within known structure types. We address this question by...

    arxiv.org/abs/2606.23166 · PDF

  52. 52

    Weighted Score-Oriented Losses for Temporally Localized Event Prediction

    Edoardo Legnaro, Sabrina Guastavino, Francesco Marchetti

    cs.LG

    Operational event-detection systems are rarely assessed by pointwise accuracy alone. In anomaly detection, changepoint detection, and warning systems, the utility of an alarm depends on its temporal position relative to an event. This produces a score-loss mismatch. Neural networks are commonly trained with classical loss functions, such as cross-entropy, whereas deployment decisions are obtained by thresholding network predictions, merging...

    arxiv.org/abs/2606.23145 · PDF

  53. 53

    The Fractal Neural Operator: Overcoming Spectral Bias in Chaotic Attractors via Prime-Harmonic Weierstrass Encodings

    Kanishk Awadhiya

    cs.LG · math.AP

    Deep learning models, particularly Transformers and Neural Operators, exhibit a well-documented "spectral bias," effectively acting as low-pass filters that smooth out high-frequency information. While benign in fluid dynamics, this bias is catastrophic for Chaotic Dynamical Systems, where the underlying strange attractor is characterized by fractal geometry and infinite spectral density. We introduce the Fractal Neural Operator (FNO), a...

    arxiv.org/abs/2606.23123 · PDF

  54. 54

    Temporal-Spectral Alignment with Frequency Adaptation for Source-Free Time-Series Adaptation

    Shichang Meng, Linquan Wu, Xuan Ai, Linqi Song

    cs.LG

    The goal of source-free domain adaptation (SFDA) for time-series data is to transfer knowledge from a pre-trained source model to an unlabeled target domain without requiring access to source data, while addressing feature shift and temporal drift inherent in the signals. Although existing approaches have explored temporal dynamics in unsupervised source-free adaptation, they largely overlook spectral shifts in time-series data. Towards this...

    arxiv.org/abs/2606.23120 · PDF

  55. 55

    Self-Evolution for Multi-Turn Tool-Calling Agents via Divergence-Point Preference Learning

    Jiaqiang Tang

    cs.LG · cs.AI · cs.CL

    Multi-turn tool-using agents must coordinate long-horizon tool sequences while tracking dialogue state and policy constraints. Existing approaches often separate inference-time orchestration from parameter-level learning, leaving tool selection weakly structured and preference updates vulnerable to train--deployment prompt mismatch. For within-benchmark self-improvement, ToolGraph combines schema-derived topology, transition weights estimated...

    arxiv.org/abs/2606.23112 · PDF

  56. 56

    ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation

    Chen Lin, Kedi Chen, Wei Zhang

    cs.LG · cs.AI

    On-policy distillation (OPD) improves LLM reasoning by training a student model on its own generated outputs, but standard OPD treats all student-generated outputs (SGOs) equally regardless of their informativeness. We observe a consistent asymmetry in controlled filtering experiments: in both OPD and on-policy self distillation (OPSD), training only on incorrect SGOs outperforms training only on correct ones. Our further analysis suggests...

    arxiv.org/abs/2606.23104 · PDF

  57. 57

    Minimax Quantile Lower Bounds for Interactive Statistical Decision Making with Privacy

    Raghav Bongole, Amirreza Zamani, Tobias J. Oechtering, Mikael Skoglund

    cs.LG · cs.IT

    Minimax risk and regret are expectation-based criteria and do not capture rare but consequential failures. To address this concern, we develop a $δ$-explicit minimax-quantile theory for interactive statistical decision making (ISDM). We first provide structural relations between minimax quantiles, lower minimax quantiles, and minimax risk. This includes a quantile-to-expectation conversion and an equivalence between strict and lower minimax...

    arxiv.org/abs/2606.23096 · PDF

  58. 58

    FLFL: Federated Latent Factor Learning for Private Recovery of Spatio-Temporal Signals

    Chengjun Yu, Di Wu, Yi He, Jia Chen

    cs.LG · cs.AI

    Wireless sensor network (WSNs) stands out as a burgeoning and promising domain in intelligent sensing. Owing to various factors such as sudden sensor malfunctions or deliberate shutdown of partial nodes to save energy, the collected sensing signals from WSNs commonly have massive missing data, leading to adverse effects on subsequent analysis or decision-making. Latent factor learning (LFL) has proven to be highly effective in recovering the...

    arxiv.org/abs/2606.23091 · PDF

  59. 59

    FlowTrain: Flow-Based Decoupled Training for Industrial-Grade Vision-Language Models

    Zhida Jiang, Zhaolong Xing, Yang Pei, Xiaolong Chen, Yuanhang Xiao, Chengzhi Huang, Xiyu Liu, Haopeng Liu, Qingyuan...

    cs.LG

    Industrial-grade distributed training of vision-language models (VLMs) remains far less efficient than that of unimodal LLMs. Existing solutions either follow a monolithic design that assigns uniform parallelism to heterogeneous modules or adopt a disaggregated deployment that separates modules while executing them as a batch-synchronized pipeline. In this paper, we highlight that the above solutions are still not sufficient, and VLM training...

    arxiv.org/abs/2606.23087 · PDF

  60. 60

    PeLAP-A: Adaptive Latent Pruning for Lightweight Latent Diffusion Models

    Kissa Zahra, Zaib Un Nisa

    cs.LG

    Latent diffusion models achieve strong generative performance by operating in a compressed latent space produced by a variational autoencoder (VAE). However, it remains unclear whether all latent channels contribute equally to the diffusion process, or whether significant redundancy exists. We introduce PeLAP-A (Adaptive Latent Pruning for Diffusion), a lightweight framework that augments a standard latent diffusion pipeline with a learnable...

    arxiv.org/abs/2606.23086 · PDF

  61. 61

    Prime Fourier Embeddings: A Principled Basis for Modular Arithmetic

    Hyunsang Hwang, Suhyun Bae, Donghun Lee

    cs.LG · cs.AI

    Numbers have algebraic structure that standard neural embeddings often fail to expose. We introduce Prime Fourier Embeddings (PFE), which encode integers as prime-indexed (cos, sin) pairs derived from the harmonic analysis of Q, providing a pre-structured representation in which modular arithmetic reduces to selecting the relevant prime channel rather than discovering algebraic structure from scratch. We prove that any linear map equivariant...

    arxiv.org/abs/2606.23044 · PDF

  62. 62

    EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning

    Hongxin Ding, Baixiang Huang, Yue Fang, Weibin Liao, Zheng Li, Jinyang Zhang, Zhijing Wu, Junfeng Zhao, Yasha Wang

    cs.LG · cs.AI

    Rubric-based rewards offer interpretable and fine-grained optimization signals for reinforcement learning in open-ended tasks where verifiable answers are unavailable. However, pre-constructed rubrics remain static throughout training, creating a fundamental mismatch with the evolving policy: fixed criteria gradually lose discriminative power as the model improves, leading to reward saturation and potential hacking. Recent dynamic rubric...

    arxiv.org/abs/2606.23038 · PDF

  63. 63

    Counterfactual learning of new adaptive instructional policies using logged data

    Samuel Girard, Sein Minn, Amel Bouzeghoub, Jill-Jênn Vie

    cs.LG

    Optimizing instructional policies in Intelligent Tutoring Systems (ITS) typically requires costly online experimentation or student simulators that may fail to capture real-world dynamics. This paper introduces an offline contextual bandit framework that learns new adaptive policies directly from logged interaction data. By mapping student-item interactions onto a continuous latent proficiency-difficulty scale using a Rasch model, we cast the...

    arxiv.org/abs/2606.23015 · PDF

  64. 64

    A Novel Approach to Temporal QoS Estimation via Extended Kalman Filter-Incorporated Latent Feature Analysis

    Ye Yuan, Song Wang, Hongxun Zhou, Ling Wang, Xin Luo

    cs.LG

    Predicting temporal Quality of Service (QoS) data is critical for optimizing network services and rationalizing resource allocation in cloud computing and service-oriented systems. Existing mainstream methods have achieved promising predictive performance. However, their purely data-driven manner limits their ability to capture non-stationary temporal patterns, thereby leading to accuracy degradation when temporal QoS data exhibits...

    arxiv.org/abs/2606.23010 · PDF

  65. 65

    Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning

    Yunan Wang, Minghui Song, Zihan Zhang, Shaohan Huang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang

    cs.LG · cs.AI · cs.CL

    Group-based Reinforcement Learning (RL) has significantly enhanced Large Language Models (LLMs) in agentic scenarios. To achieve finer-grained policy updates, recent agentic RL frameworks have shifted from trajectory-level to step-level training. However, long-horizon agentic RL suffers from severe reward sparsity and delay, as feedback is often deferred for dozens of interaction steps. While existing step-level frameworks refine training...

    arxiv.org/abs/2606.22995 · PDF

  66. 66

    Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?

    Nils Grandien, David Steinmann, Felix Friedrich, Kristian Kersting

    cs.LG

    Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable, recent approaches have begun imposing hierarchical structure, either explicitly or as an implicit effect of training constraints, yet rigorous comparison remains difficult. There are no agreed-upon requirements for what a meaningful feature hierarchy should...

    arxiv.org/abs/2606.22994 · PDF

  67. 67

    Neural Architecture Search of Sample Reweighting Networks for Complex Distribution Shift

    Keisuke Sugawara, Kento Uchida, Shinichi Shirakawa

    cs.LG · cs.AI

    Sample reweighting is a major approach to addressing distribution shifts, such as label noise and class imbalance. Meta-Weight-Net (MW-Net) is a promising sample reweighting network that computes weights based on classification loss. Although MW-Net improves prediction performance under a single type of distribution shift using a simple neural network, its performance degrades when facing both label noise and class imbalance, where it is hard...

    arxiv.org/abs/2606.22991 · PDF

  68. 68

    Understanding Parallel Samplers in Masked Diffusion via Random Walks on Graphs

    Vansh Bansal, Cho Cholyeon, Syamantak Kumar, Sujay Sanghavi, Purnamrita Sarkar

    cs.LG · cs.AI · cs.CL

    In this paper, we propose using random walks on graphs as a verifiable sandbox to study different parallel sampling strategies in masked diffusion models (MDMs). We train an MDM on random walk samples from a fixed graph. The graph or the transition kernel is never shown to the model explicitly and plays the role of latent structure in the sequences, albeit one that is controllable and can be used for quantitative evaluation. Thus, this...

    arxiv.org/abs/2606.22976 · PDF

  69. 69

    TaLK: Text-attributed Graph Dataset Distillation via Coupling Language Model with Graph-Aware Kernel

    Yeongho Kim, Yeonje Choi, Kijung Shin

    cs.LG

    Text-attributed graphs (TAGs) are widely used in many real-world domains, and learning on TAGs requires jointly modeling text semantics and graph structure. A standard approach for modeling TAGs is to combine a language model (LM) and a graph neural network (GNN), but joint training is computationally expensive and difficult to scale. Dataset distillation is a promising way to reduce training costs, but existing methods are not well suited to...

    arxiv.org/abs/2606.22975 · PDF

  70. 70

    Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction

    Georg Trede, Charlotte Ricarda Doll, Elias Weber, Daniel Durstewitz

    cs.LG · math.DS · nlin.CD

    Predicting the behavior of dynamical systems (DS) beyond the dynamical and parameter regimes observed in training is a pivotal and essentially unresolved problem in scientific ML. It is central to any good scientific theory, which we expect to be able to make predictions about regimes not covered by currently available data. Recent hierarchical and hyper-network guided approaches for DS reconstruction (DSR) enable training on many DS...

    arxiv.org/abs/2606.22969 · PDF

  71. 71

    Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models

    Linghan Chen, Kaiyan Ji, Minyu Guo

    cs.LG · cs.AI · cs.CR

    Many recent vision-language-action (VLA) policies adopt an imagine-then-act design. A world-action model (WAM) first imagines a short future as a latent trajectory z~, on which the action is then conditioned. We identify this trusted imagination, rather than the reactive policy, as the exposed attack surface. A downstream oracle, such as a safety gate, a visual model-predictive-control (MPC) planner, or an imagine-then-check verifier,...

    arxiv.org/abs/2606.22966 · PDF

  72. 72

    PG-MAP: Joint MAP Optimization for Inference-Time Alignment of Diffusion and Flow-Matching Models

    Ruolan Sun, Pawel Polak

    cs.LG · cs.CV

    Inference-time alignment of pretrained text-to-image models is typically performed along a single control axis, such as classifier-free guidance, attention editing, or reward-based latent perturbations. This limitation prevents modeling joint dependencies between conditioning and latent variables and hinders transfer across generative transports. We propose PG-MAP, a training-free framework that formulates inference-time alignment as a...

    arxiv.org/abs/2606.22958 · PDF

  73. 73

    DT-GOL: Dual-Track Geometric Online Learning in Nonstationary Environment with Label Delay

    Yulin Wang, Yi He, Dianlong You, Di Wu

    cs.LG

    Online learning is crucial for handling complex data streams in big data applications. Recent research has begun to focus on dynamic scenarios, i.e., non-stationary environments. However, a crucial yet often overlooked aspect is label latency, where new data may not receive labels in time due to the slow and expensive labeling process, thus hindering rapid adaptation to dynamic environments. To resolve this impasse, we propose Dual-Track...

    arxiv.org/abs/2606.22950 · PDF

  74. 74

    Neural Operator Processes for Probabilistic Operator Learning under Partial Observations

    Jose Miguel Lara-Rangel, Serge Guillas

    cs.LG

    Neural operators learn mappings between function spaces, but are typically developed with dense input-output training fields and fully observed inputs at inference. Many scientific problems require instead predicting solution fields from sparse, irregular, or partial observations under uncertainty. We introduce Neural Operator Processes (NOPs), a framework that unifies neural-process conditioning with neural-operator decoding to predict full...

    arxiv.org/abs/2606.22946 · PDF

  75. 75

    Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently

    Stanley Wei, Juno Kim

    cs.LG · cs.AI

    Recent advances in large language models (LLMs) have demonstrated that reinforcement fine-tuning of pretrained base models can lead to significant gains in reasoning performance at inference time. In this work, we theoretically analyze why reinforcement fine-tuning induces better reasoning ability than purely supervised fine-tuning (SFT) methods. We model chain-of-thought (CoT) reasoning as a pathfinding problem on graphs and compare the...

    arxiv.org/abs/2606.22938 · PDF

  76. 76

    FORGE: Fused On-Register Gradient Elimination for Memory-Efficient LLM Training

    Dikshant Kukreja, Kritarth Prasad, Avinash Anand, Zhengkui Wang, Erik Cambria, Timothy Liu, Aik Beng Ng, Simon See,...

    cs.LG

    Reverse-mode differentiation computes every weight gradient, writes it to memory, and only then lets the optimizer read it back. This two-phase schedule sets the memory ceiling of modern training: at the seam between the phases, every layer's gradient is live at once. We argue that this materialized gradient is an artifact of how differentiation is staged, not a quantity that learning requires -- and we eliminate it. FORGE folds the optimizer...

    arxiv.org/abs/2606.22932 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.