cs.LG · 2026-07-20 · No. 59

Machine Learning, 2026-07-20.

71 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

71 entries
  1. 01

    PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization

    Yuchen Yang, Yifan Zhao, Anisha Dasgupta, Sasa Misailovic

    cs.LG

    Mixture-of-Experts (MoE) is a popular class of large language models (LLMs), offering high efficiency and accuracy. However, in KV-cache-intensive serving scenarios, MoEs often exhibit a tension between the GPU memory requirements of the model weights and the growing KV cache. We propose PagedWeight, a novel management method for MoE LLM serving that dynamically quantizes MoE model's weights at runtime and balances expert-weight precision...

    arxiv.org/abs/2607.16184 · PDF

  2. 02

    A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing

    Owen Lockwood, Jérémy Béjanin, Joost Bus, Christopher Chamberland, Patrick Huembeli, Frank Schäfer, Guillaume Verdon

    cs.LG · cs.ET · physics.app-ph

    To address the escalating energy and latency demands of machine-learning workloads, we introduce a blueprint for an energy-efficient and fast thermodynamic computing stack that leverages stochastic analog processes in physical hardware. In this work, we focus on energy-based thermodynamic computing where the stochastic process is well described by Langevin dynamics with tunable energy potentials. The implementation of such potentials in...

    arxiv.org/abs/2607.16183 · PDF

  3. 03

    Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems

    Matteo Tomasetto, Nicolò Botteghi, Gabriele Bruni, Andrea Manzoni

    cs.LG · math.OC

    Reinforcement learning (RL) has recently emerged as a promising feedback control strategy for nonlinear and complex dynamical systems. However, RL algorithms are sample inefficient and require a large number of interaction with the environment to synthesize optimal control strategies. Consequently, applications of RL are typically limited to sparse sensors and actuators due to the curse of dimensionality entailed by the...

    arxiv.org/abs/2607.16177 · PDF

  4. 04

    When Does Muon Help Agentic Reinforcement Learning?

    Kai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun

    cs.LG · cs.AI

    Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld using Qwen2.5-0.5B-Instruct. Under Group-in-Group Policy Optimization (GiGPO), applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%);...

    arxiv.org/abs/2607.16169 · PDF

  5. 05

    Behaviour-Conditioned Neural Processes for Adaptive Residential Short-Term Load Forecasting

    Ramin Soleimani, Andrea Visentin, Dirk Pesch

    cs.LG

    Residential short-term load forecasting (STLF) is challenging because household demand is heterogeneous, temporally variable, and shaped by diverse behavioural routines. This work investigates whether inferred behavioural structure can be embedded within the forecasting mechanism of a Neural Process-based probabilistic model, rather than used only as an external grouping signal, for context-conditioned residential STLF. We propose a...

    arxiv.org/abs/2607.16168 · PDF

  6. 06

    PRISA: Proactive Infrastructure LiDAR Framework for Intersection Safety Assessment

    Tam Bang, Hussam Abubakr, Emiliano de la Garza Villarreal, Truc Phuong Nguyen, Austin Harris, Toru Hirano, Mina...

    cs.LG · eess.SY

    Urban intersections are among the most hazardous locations in road networks, posing significant risks to vehicles and vulnerable road users (VRUs) such as pedestrians and cyclists. The complexity of multi-agent interactions demands continuous, real-time monitoring systems capable of anticipating conflicts before they escalate into crashes. We present PRISA, a modular infrastructure LiDAR framework leveraging privacy-preserving,...

    arxiv.org/abs/2607.16156 · PDF

  7. 07

    Improving Improved Kernel PLS

    Ole-Christian Galbo Engstrøm

    cs.LG · cs.DS

    Improved Kernel Partial Least Squares (IKPLS) algorithms 1 and 2 are among the fastest PLS calibration algorithms. This article focuses on two shared steps, the computation of the $\mathbf{X}$ rotations, $\mathbf{R}$, and the $\mathbf{Y}$ loadings, $\mathbf{Q}$, and accelerates both. For $\mathbf{R}$, term-by-term accumulation is replaced by a direct evaluation strategy that requires the same number of multiplications but parallelizes better...

    arxiv.org/abs/2607.16138 · PDF

  8. 08

    When Do Multi-Agent Systems Help? An Information Bottleneck Perspective

    Wendi Yu, Lianhao Zhou, Xiangjue Dong, Sai Sudarshan Barath, Declan Staunton, Byung-Jun Yoon, Xiaoning Qian, James...

    cs.LG · cs.AI

    LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with performance varying inconsistently across settings. Here, we provide an information bottleneck perspective on elucidating the differences between MAS and SAS. Specifically, our key observation is that a SAS accumulates its full reasoning trace in one shared context, while...

    arxiv.org/abs/2607.16133 · PDF

  9. 09

    The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic Infrastructure

    Jun He, Deying Yu

    cs.LG · cs.DC · cs.MA

    State machine replication (SMR) and Byzantine fault-tolerant (BFT) consensus guarantee agreement despite a bounded number of arbitrary, colluding faulty participants. However, these guarantees rely on participants outside this set correctly executing the protocol's transition semantics. Agentic validators expose a weaker boundary: an authenticated, responsive, non-equivocating, and protocol-compliant reasoning participant may still endorse a...

    arxiv.org/abs/2607.16109 · PDF

  10. 10

    Understanding Reasoning from Pretraining to Post-Training

    Jingyan Shen, Ang Li, Salman Rahman, Yifan Sun, Micah Goldblum, Matus Telgarsky, Pavel Izmailov

    cs.LG · cs.AI · cs.CL

    Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM...

    arxiv.org/abs/2607.16097 · PDF

  11. 11

    DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

    Hanyang Chen, Anirudh Satheesh, Longchao Da, Hua Wei

    cs.LG · cs.AI

    Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains. In this paper, we consider the setting of online dynamics adaptation, where policies are trained in the source domain with sufficient data, while only limited interactions with the target domain are allowed. There are a few existing works that address the dynamics mismatch by employing...

    arxiv.org/abs/2607.16090 · PDF

  12. 12

    Neural spectroscopy of AlphaFold2 reveals encoded protein conformational landscapes

    Kaustav Mehta

    cs.LG · q-bio.BM

    AlphaFold2's 93 million parameters, shaped by the evolutionary record of protein structure encoded in the Protein Data Bank and in sequence alignments, are conventionally treated only as machinery for converting sequence to structure. We propose they are also a scientific object that can be analyzed directly: a learned encoding of protein conformational organization that can be probed and characterized. By smoothing the Evoformer's weight...

    arxiv.org/abs/2607.16087 · PDF

  13. 13

    Physics-Based Deep Spatiotemporal Hyperlocal Radar Nowcasting with a Multi-Variable U-Net for High-Resolution Precipitation Forecasting

    Akshay Sunil, Muhammed Rashid, Raja Sekhar Sivaraju, Sushma Nair, Subimal Ghosh

    cs.LG · eess.IV

    Precipitation nowcasting over the immediate 10-90 min period is important for flood management and real-time decision-making in urban regions. Conventional short-range forecasting with high-resolution numerical weather prediction requires frequent data assimilation, model initialization, and spin-up, introducing computational latency. Machine learning provides an alternative by learning storm evolution directly from high-frequency...

    arxiv.org/abs/2607.16080 · PDF

  14. 14

    When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis

    S. Aaron McClendon

    cs.LG · cs.AI

    Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is essentially never tested against the baseline it claims to replace: methods merge independently released agents precisely because a joint model is unavailable. We build the missing comparison. Training difficulty-1 and difficulty-2 Qwen3-8B specialists on the AppWorld agent benchmark with LOOP, we merge them...

    arxiv.org/abs/2607.16062 · PDF

  15. 15

    DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model Embeddings

    Yuya Kawakami, Daniel Cayan, Dongyu Liu, Kwan-Liu Ma, Tom Corringham

    cs.LG

    Pluvial (rainfall-driven) flooding accounts for 45% of National Flood Insurance Program (NFIP) claims in the United States and is harder to predict than its riverine and coastal counterparts, with existing approaches limited to coarse resolution, regional domains, or computationally intensive process-based models unsuitable for daily continental-scale use. We present DELUGE, a multimodal deep learning framework for daily pluvial flood damage...

    arxiv.org/abs/2607.16050 · PDF

  16. 16

    Revisiting data-driven dynamic security assessment with a tabular foundation model

    Olayiwola Arowolo, Maosheng Yang, Jochen Cremer

    cs.LG · cs.AI

    Data-driven pre-fault dynamic security assessment (DSA) rapidly evaluates the dynamic risk of credible contingencies on a power system using machine learning. Existing approaches face two limitations. First, they require a large labelled database for training, with a separate model trained, tuned, and maintained for each contingency in a potentially long list of credible contingencies. Second, the trained models generalize poorly to unseen...

    arxiv.org/abs/2607.16031 · PDF

  17. 17

    CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward Approach

    Andrei Neagu, Eeham Khan, Leila Kosseim

    cs.LG

    This paper presents our system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires daily long, flat, or short trading decisions for Bitcoin (BTC) and Tesla (TSLA) using news and historical market data. We formulate the problem as a discrete-action Markov Decision Process and compare four deep reinforcement learning algorithms: Policy Gradient (PG), Proximal Policy Optimization (PPO), Deep Q-Learning (DQL), and Deep Deterministic Policy...

    arxiv.org/abs/2607.16028 · PDF

  18. 18

    Constrained Hebbian Learning Supports Efficient Representational Allocation under Structural Constraints

    Patrick Inoue, Florian Röhrbein, Andreas Knoblauch

    cs.LG · cs.NE

    Introduction: Biological systems face anatomical and metabolic constraints, including costly synaptic maintenance and limited connectivity. These constraints favor neural codes that compress behaviorally relevant information into low-redundancy patterns. We test whether an excitatory competitive Hebbian rule can support synaptic resource allocation under such constraints and whether the resulting representations occupy a more favorable...

    arxiv.org/abs/2607.16027 · PDF

  19. 19

    Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation

    Zhaoyang Jiang, Zhizhong Fu, Zicheng Li, Yunsoo Kim, Jiacong Mi, Xuanqi Peng, Fei Teng, Honghan Wu

    cs.LG

    AI systems increasingly retrieve from records that revise themselves: issue threads, encyclopedic histories, policy logs, and long conversations. The challenge is not only finding relevant evidence, but deciding which claims remain in force, which were superseded, and when to abstain. Structured memories promise to solve this with typed edges, temporal updates, and conflict status, yet evaluations often change mechanism and prompt...

    arxiv.org/abs/2607.16019 · PDF

  20. 20

    DebrisTracer: Reliable Tracking in Hypervelocity Impact Fast Imaging

    Théophane Loloum, Fabien Vivodtzev, David Hébert, Baptiste Reynier, Michel Arrigoni, Julien Tierny

    cs.LG · cs.CV · cs.GR · eess.IV

    This application paper presents DebrisTracer, a framework for the reliable tracking of debris in hypervelocity impact fast imaging. These noisy and highly specific datasets capture the ejection of a large number of debris fragments after the impact of a projectile launched at hypervelocity into a target material. The reliable estimation of debris mass and speed distributions is of major importance in aerospace applications. We document how to...

    arxiv.org/abs/2607.15986 · PDF

  21. 21

    An Exploratory Study of Single Channel Surface Electromyography for Hand Gesture Classification

    Daanish Hindustani

    cs.LG

    Accurate hand gesture recognition using surface electromyography (sEMG) typically relies on multichannel sensor arrays and computationally intensive models. This limits practical deployment in low-power and embedded systems. This study investigates the feasibility of classifying ten hand gestures using a single sEMG channel combined with lightweight machine learning architectures. Raw sEMG signals were transformed into a comprehensive...

    arxiv.org/abs/2607.15972 · PDF

  22. 22

    Knowledge-Guided Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Label-Conditioned Contrastive Alignment

    Xinran Liu, Yuwen Li, Hongxiang Gao, Heyang Xu, Jianqing Li, Zongmin Wang, Chengyu Liu

    cs.LG

    Adult and pediatric electrocardiogram (ECG) interpretation relies on age-sensitive criteria, and models pretrained mainly on adult ECGs often transfer poorly to pediatric populations when pediatric labels are scarce. Existing multimodal ECG--text methods typically align waveforms and text at the global sample level, entangling evidence from co-occurring diagnoses and limiting transfer under this gap. We propose Pediatric-Adult ECG Alignment...

    arxiv.org/abs/2607.15928 · PDF

  23. 23

    On the Failure of Boundary-Seeking Distillation in Bottlenecked Generative Architectures

    Mohamed Amine Kina

    cs.LG · cs.AI

    Data-free knowledge distillation transfers the knowledge encoded in a teacher model to a student model without access to the original training data. Prior work such as Contrastive Abductive Knowledge Extraction (CAKE) achieves this for classifiers by synthesizing samples near the teacher's decision boundary. In this work, we investigate whether this boundary-seeking principle extends to autoencoder distillation through experiments on the...

    arxiv.org/abs/2607.15919 · PDF

  24. 24

    (MPO)$^2$: Multivariate Polynomial Optimization based on Matrix Product Operators

    Niccolò Ciolli, Anders Vestergaard Nørskov, Michael Kastoryano, Petr Taborsky, Morten Mørup

    cs.LG

    Central to machine learning and signal processing is the ability to perform universal function approximation and learn complex input-output relationships from limited numbers of observations. Multivariate polynomial models offer a natural way to express such relationships through multiplicative feature interactions, but their coefficient tensors grow exponentially in size with the polynomial degree. Existing tensorized polynomial models...

    arxiv.org/abs/2607.15916 · PDF

  25. 25

    A Semiparametric Framework for Stochastic Fundamental Diagram Modeling

    Pengnan Chi, Xiaoliang Ma, Magnus Jansson, Magnus Nordenvaad

    cs.LG

    The stochastic fundamental diagram (SFD) provides a probabilistic description of the relationship between traffic density and flow or speed, enabling uncertainty-aware traffic modeling. However, existing stochastic models frequently struggle to accommodate rigorous physical constraints while retaining sufficient flexibility to capture complex nonlinear patterns. To address this, we propose a novel semiparametric SFD modeling framework by...

    arxiv.org/abs/2607.15907 · PDF

  26. 26

    ContinuityBench: A Benchmark and Systems Study of Stateful Failover in Multi-Provider LLM Routing

    Vishal Pandey, Gopal Singh

    cs.LG

    In production large language model (LLM) deployments, high API availability guarantees do not equate to conversational continuity. When a primary provider experiences an outage or strict rate-limiting, naive stateless failover mechanisms successfully maintain uptime but silently discard conversation history, severely disrupting the user experience. To rigorously quantify and resolve this failure mode, we introduce two novel metrics:...

    arxiv.org/abs/2607.15899 · PDF

  27. 27

    Data-Native Global Optimization for Big Data K-means Clustering

    Ravil Mussabayev, Rustam Mussabayev, Zukhra Yerdaliyeva, Kuldeyev Nursultan

    cs.LG

    Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids. We target arbitrarily tall data: a fixed feature space may contain arbitrarily many, possibly infinitely many, observations, while the algorithm accesses only finite random samples. We propose Big-means++, an algorithm...

    arxiv.org/abs/2607.15835 · PDF

  28. 28

    In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention

    Katsuyuki Hagiwara

    cs.LG · cs.AI

    In-context learning is a remarkable property of transformers and has recently received a lot of interest. In many studies of in-context learning, it has been shown that transformers are capable of implementing solver for linear and non-linear regression problems, in which the most of them implement gradient descent algorithm. However, it is still unclear whether those implementations have actually been acquired through training. In this...

    arxiv.org/abs/2607.15819 · PDF

  29. 29

    Graph Coloring Approach to Solving Sudoku with Oscillatory Neural Networks

    Filip Sabo, Aida Todri-Sanial

    cs.LG

    Oscillatory Neural Networks (ONNs) present an attractive physics-based computing paradigm rooted in the dynamics of a network of typically fully coupled oscillators aiming to minimize an underlying energy function. In this paper, we propose an ONN-based solver for one well-known constrained combinatorial optimization problem, namely a Sudoku, by formulating the problem as a Graph Coloring problem. By modifying the already existing Graph...

    arxiv.org/abs/2607.15814 · PDF

  30. 30

    QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides

    Zhengyang Zhuge, Hao Yu, Xin Wang, Zheng Li, Yizhong Cao, Dayiheng Liu, Jianwei Zhang

    cs.LG

    Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP8. As an emerging low-precision format, NVFP4 combines fine-grained scaling for accuracy preservation with native W4A4 FP4 GEMMs for higher throughput than FP8. However, we find that directly applying NVFP4 to MoE RL rollout is impractical. NVFP4 rollout with BF16...

    arxiv.org/abs/2607.15810 · PDF

  31. 31

    Knowledge-Assisted Multi-Graph Dependency Learning for Multivariate Time Series Anomaly Detection in Multi-Stage Industrial Processes

    Jaeyeong Lee, Taeseong Yoon, Wonmo Koo, Heeyoung Kim

    cs.LG · cs.AI

    Industrial processes often generate complex, interdependent time-series data from multiple sensors across multiple stages, forming complex dependencies among variables and process stages. Effective monitoring and timely anomaly detection of these time series through multivariate time series anomaly detection (MTAD) is crucial for preventing failures and ensuring the reliability of automated systems. Graph neural networks (GNNs) have advanced...

    arxiv.org/abs/2607.15799 · PDF

  32. 32

    AquaAugmentor: A Novel Feature Augmentation Algorithm for Water Potability Prediction

    Muntasir Tabasum, Al Zadid Sultan Bin Habib, Tanpia Tasnim, Md. Ekramul Islam, Md Younus Ahamed, Md Asif Bin Syed

    cs.LG · cs.AI · cs.CE

    Access to potable water is crucial for health, economic development, and sustainability. However, accurately classifying water quality remains a significant challenge due to the complexity and variability of water source data. This paper addresses the challenge of predicting water potability through machine learning and deep learning algorithms. It introduces a novel feature augmentation algorithm, AquaAugmentor, to enhance the predictive...

    arxiv.org/abs/2607.15775 · PDF

  33. 33

    Scaling Time Series Classification via XAI-Driven Data Reduction

    Davide Italo Serramazza, Thach Le Nguyen, Georgiana Ifrim

    cs.LG · cs.AI

    Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI, a novel methodology that repurposes XAI attribution methods for effective data reduction in Time Series Classification (TSC). The core challenge in modern TSC is scalability; state-of-the-art models, such as...

    arxiv.org/abs/2607.15774 · PDF

  34. 34

    From Diffusion to Reaction-Diffusion: A Dynamical-Systems View of Oversmoothing in Hypergraph Neural Networks

    Zhiheng Zhou, Mengyao Zhou, Yancheng Chen, Dengyi Zhao, Xingqin Qi, Guiying Yan

    cs.LG

    Higher-order couplings enhance the expressive power of hypergraph neural networks (HGNNs), but they also intensify representation collapse in deep propagation due to strong multi-way feature mixing. This work investigates hypergraph oversmoothing from a dynamical-systems perspective and develops a reaction--diffusion framework for depth-resistant hypergraph learning. By defining hypergraph gradient and divergence operators, we interpret...

    arxiv.org/abs/2607.15773 · PDF

  35. 35

    CoG-Guided Weight Correction for Fault-Tolerant Deep Neural Networks

    Bahram Parchekani, Samira Nazari, Ali Azarpeyvand, Mohammad Hasan Ahmadilivani, Tara Ghasempouri, Jaan Raik

    cs.LG · cs.AR · math.NA

    Deep Neural Networks (DNNs) used in safety-critical applications are vulnerable to hardware and memory faults that corrupt network weights and degrade reliability. In this paper, we propose a Center of Gravity (CoG) guided weight correction method that restores faulty weights based on their spatial characteristics within each layer. The proposed approach detects and corrects weight faults using distance-aware correction rules, eliminating the...

    arxiv.org/abs/2607.15753 · PDF

  36. 36

    Trainable Spline Representations for Physics-Informed Learning

    Giovanni Canali, Nicola Demo, Gianluigi Rozza

    cs.LG · math.NA

    This work introduces Physics-Informed Splines (PI-Splines), a structured spline-based architecture for physics-informed learning. Instead of representing the solution of a differential equation with a neural network, PI-Splines directly parametrize the unknown field through a tensor-product B-spline expansion with trainable control coefficients. This formulation preserves the residual-based training paradigm of Physics-Informed Neural...

    arxiv.org/abs/2607.15751 · PDF

  37. 37

    Learning Faster without Deeper Networks: A*-Inspired Batch Selection for Efficient CNN Training

    Anxhelo Shehu, Enes Stastoli, Arben Cela

    cs.LG

    Common practice when training Convolutional Neural Networks (CNNs) is to use randomly shuffled mini-batches. This creates two limitations: slower convergence, and a diminishing learning signal, since many samples are quickly classified as easy during training. We address these inefficiencies with A*-Inspired Batch Selection (A*-BS), a lightweight, model-agnostic strategy that formulates mini-batch scheduling as a heuristic search problem....

    arxiv.org/abs/2607.15745 · PDF

  38. 38

    CardioMeta: Calibrated Multi-Task Prediction of Diabetes, Hypertension, and Cardiovascular Disease Across Population and EHR Data

    S M Asif Hossain, Ruksat Khan Shayoni, M. F. Mridha, Jungpil Shin

    cs.LG

    Cardiometabolic diseases remain among the most persistent drivers of preventable morbidity because diabetes, hypertension, and cardiovascular disease frequently co-occur and share metabolic, vascular, demographic, and behavioral determinants. Existing machine learning studies for chronic disease prediction often emphasize discrimination on a single dataset, while underreporting label leakage, calibration, temporal robustness, external...

    arxiv.org/abs/2607.15721 · PDF

  39. 39

    A Benchmark for Electrical Load Forecasting Across Grid Levels: Time-Series Transformers Outperform Established Methods

    Matthias Hertel, Sebastian Pütz, Jonathan Kolar, Benjamin Schäfer, Ralf Mikut, Veit Hagenmeyer

    cs.LG

    Accurate load forecasting at multiple grid levels is essential for future smart grids, ranging from aggregated control area forecasts for balancing supply and demand to forecasts of individual end-consumer loads for demand-side management and energy management systems. We present a comprehensive benchmark for load forecasting across grid levels, comprising three datasets that represent a transmission system operator control area, low-voltage...

    arxiv.org/abs/2607.15705 · PDF

  40. 40

    Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework

    Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang

    cs.LG

    Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than single-modality graphs. In practice, such graphs are fragmented across privacy-restricted silos owned by different platforms and institutions, so learning a broadly transferable model over them...

    arxiv.org/abs/2607.15687 · PDF

  41. 41

    Neural Non-Equilibrium Hamiltonian Monte Carlo for Corrected Boltzmann Sampling

    Moxian Qian

    cs.LG · cond-mat.stat-mech · hep-lat

    Sampling from an unnormalized Boltzmann density requires proposals that move probability mass globally while retaining enough path-probability information for statistical correction. We introduce Neural Non-Equilibrium Hamiltonian Monte Carlo (NHMC), a train-then-correct learned Hamiltonian sampler. Starting from a tractable base distribution, NHMC learns stochastic Hamiltonian-style paths toward the target. Once training is complete, the...

    arxiv.org/abs/2607.15682 · PDF

  42. 42

    ASK-NN: An Asymmetric Nearest-Neighbor Test that detects Distribution Drifts in Natural Language

    Sergey Zakharov, Rodion Oblovatny, Alexey Zaytsev

    cs.LG · stat.ML

    Hallucinations and artificial text in LLM-generated outputs often appear as distributional deviations between prompt and response hidden-state distributions. Since prompts or retrieved contexts typically serve as reference samples and responses as query samples, with major differences in length, these asymmetries motivate the use of change test statistics that treat the two samples differently. We consider an asymmetric two-sample test ASK-NN...

    arxiv.org/abs/2607.15607 · PDF

  43. 43

    Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data

    Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee

    cs.LG · stat.ML

    Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing, yet a generator can reproduce every marginal and every foreign-key relationship while emitting timestamps that run backwards or repeat, and while sending entities along paths that no real entity followed. Conventional tabular evaluation, which pools records into static distributions, is blind to such failures. We present a taxonomy-guided evaluation...

    arxiv.org/abs/2607.15606 · PDF

  44. 44

    Field-Aware RankMixer with Dual-Stream Bilinear Fusion for the Tencent UNI-REC Challenge

    Yufeng Zhang, Zhengqi Xu, Jiajun Cui

    cs.LG · cs.AI

    This paper presents our solution to the KDD Cup 2026 Tencent UNIREC Challenge. The task requires joint modeling of multi-domain user behavior sequences and non-sequential multi-field features for target-ad pCVR prediction. We develop a Field-Aware RankMixer (FA-RankMixer) with dual-stream bilinear fusion. The model first applies target-aware DIN modules to extract user interests from multiple behavior domains. It also models recent and...

    arxiv.org/abs/2607.15590 · PDF

  45. 45

    Rethinking Transfer in Continual Learning: A Replay-Based Realisation

    Yang Meng, Zhenya Liu, Zhuokai Zhao, Yuxin Chen

    cs.LG

    Continual learning studies how deployed language models can continually acquire new tasks without expensive retraining from scratch. Existing methods, whether rehearsal-based (replaying stored past data) or rehearsal-free (regularising or isolating parameters), overwhelmingly target one objective: preventing catastrophic forgetting. Forward transfer, the past helping the future, has meanwhile been pursued almost exclusively through parameter...

    arxiv.org/abs/2607.15587 · PDF

  46. 46

    Information-Directed Sampling for Causal Bandits

    Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi

    cs.LG · cs.AI

    Causal bandits exploit structural relationships among variables to share information across interventions and accelerate the identification of high-reward decisions. In many applications, however, some variables cannot be directly manipulated, even though they influence the reward and provide useful information about the underlying causal system. We study contextual causal bandits with non-manipulable variables, where context variables are...

    arxiv.org/abs/2607.15577 · PDF

  47. 47

    Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation

    Himel Dev, Madhusudan Basak, Tanmoy Sen, Paromita Shome, Bashima Islam

    cs.LG · cs.AI

    Packing for air travel is recurring and error-prone: the checklist must be personal and context-aware, yet feasible under safety rules, item dependencies, and luggage limits. Existing packing assistants are template-driven and generic, or recommendation-driven but unconstrained, leaving users to manually patch regulatory and capacity violations. We propose a reasoning-guided learning framework with three stages: (1) a symbolic engine that...

    arxiv.org/abs/2607.15562 · PDF

  48. 48

    From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation

    Himel Dev, Tanmoy Sen, Madhusudan Basak, Bashima Islam

    cs.LG · cs.AI

    Generating personalized trip itineraries is a complex planning task and involves a tension between hard combinatorial feasibility and soft latent desirability. Classical optimization enforces constraints but fails to capture subjective traveler preferences. While learning-based approaches model preferences, they cannot guarantee feasibility. Mobile deployment imposes additional resource constraints on both. To address this, we propose Plan,...

    arxiv.org/abs/2607.15552 · PDF

  49. 49

    Publicly-Verifiable Certificates for Statistical Algorithms

    Michael Ngo, Michael P. Kim

    cs.LG · cs.CR · cs.DS

    Following Goldwasser, Rothblum, Shafer, and Yehudayoff, who defined a framework for interactive proofs of learning [ITCS'21], we initiate the study of non-interactive proofs of learning. We define and study a new notion: Publicly-Verifiable Certificates of Statistical Validity (pvCSVs), which allow for public, distributionally-robust certification that the result of a learning algorithm is valid. In a pvCSV, a learner publishes a hypothesis...

    arxiv.org/abs/2607.15528 · PDF

  50. 50

    Kolmogorov--Arnold Networks for Small Language Models

    Felippe Alves, Renato Vicente

    cs.LG · cs.AI

    Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks. We test these claims separately. In a six-layer, 10M-parameter B-spline KAN, we reconstruct all 884,736 feed-forward edges: 87.8\% exceed (NLS>0.1) and 0.4\% are inactive. Pruning the lowest-activity 20--25\% causes...

    arxiv.org/abs/2607.15525 · PDF

  51. 51

    Recursive Harness Self-Improvement

    Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Matei Zaharia, Yujin Tang

    cs.LG · cs.AI

    Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate...

    arxiv.org/abs/2607.15524 · PDF

  52. 52

    Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching

    Yan Song

    cs.LG · cs.AI

    Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent). The compression literature has standardized on query-aware methods that produce a different compressed prefix per query, mechanically invalidating the prefix-strict cache on every call. We characterize this cost empirically on Anthropic's Sonnet 4.6 API and find caching is...

    arxiv.org/abs/2607.15516 · PDF

  53. 53

    An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism

    Mobina Kashaniyan, Mehrdad Ashtiani, Amirhossein Ghassemi

    cs.LG · cs.AI · cs.DC · cs.PF

    Serverless computing provides automatic resource management and pay-per-use execution, but effective autoscaling remains challenging because of dynamic workloads, cold-start latency, and dependencies among functions. We present a dependency-aware autoscaling framework that integrates graph-based bottleneck identification, short-term workload forecasting, multi-model consensus, and cost-aware scaling control. Serverless applications are...

    arxiv.org/abs/2607.15511 · PDF

  54. 54

    Diffusion models recover accurate mixture weights despite score function insensitivity

    Andrew Dennehy, Ramchandran Muthukumar, Rebecca Willett, Nisha Chandramoorthy

    cs.LG · math.PR · stat.ML

    Score-based generative models exhibit a puzzling behavior: they often appear to cover all modes of a target multimodal distribution and yet may fail to learn the correct relative mode amplitudes, which can be interpreted as mixture weights. We resolve this apparent paradox by relating the diffusion score matching (DSM) loss to the error in estimating mixture weights from generated samples. We show that, even when the target score is...

    arxiv.org/abs/2607.15485 · PDF

  55. 55

    Inpainting Insights: Elevating Visual XAI with Photorealistic Perturbations

    Josef Lindl, Mariana Chaves, Damien Garreau

    cs.LG

    The increasing complexity of state-of-the-art machine learning models has made their behavior progressively harder to interpret, spurring rapid advancements in the field of eXplainable Artificial Intelligence (XAI). Among many methods proposed, perturbation-based approaches play a major role. By systematically altering (perturbing) input features, these approaches measure the impact on the model's predictions. For image data, traditional...

    arxiv.org/abs/2607.15482 · PDF

  56. 56

    Deep Learning Approaches for Sleep Apnea Classification from Polysomnographic EEG Signals

    Shashank Manjunath, Mukesh Cheemakurthi, Aarti Sathyanarayana

    cs.LG

    Sleep apnea diagnosis via polysomnography remains resource intensive and relies on time consuming manual data analysis and scoring. Recent work has demonstrated that central nervous system effects of sleep apnea events can be detected through electroencephalogram (EEG) signals. However, most work uses a single feature type on various datasets combined with different classification algorithms. In this work, we present a comprehensive...

    arxiv.org/abs/2607.15477 · PDF

  57. 57

    ADS-C: Antidistillation Sampling for Classification

    Khawaja Abaid Ullah, Mohammad Javad Khojasteh

    cs.LG · cs.CR

    Knowledge distillation enables an adversary to replicate a proprietary classifier by querying its prediction interface and training a surrogate on the returned probability vectors. Antidistillation sampling, proposed for large language models, counters this threat with an input-dependent, gradient-directed perturbation of the served distribution; its transfer to classification has not been studied. Adapting the defense to classification, we...

    arxiv.org/abs/2607.15467 · PDF

  58. 58

    Robust Peak-cost Constrained Reinforcement Learning

    Shilpa Mukhopadhyay, Sourav Ganguly, Santosh Mohan Rajkumar, Honghao Wei, Debdipta Goswami, Arnob Ghosh

    cs.LG

    We study robust peak-cost constrained reinforcement learning (RP-CRL), where the objective is to maximize expected reward while controlling the maximum cost encountered along a trajectory. This setting is motivated by safety-critical applications in which a single large violation can be catastrophic and therefore cannot be adequately captured by the standard CMDP framework based on expected cumulative cost. Existing reachability-constrained...

    arxiv.org/abs/2607.15457 · PDF

  59. 59

    Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers

    James O' Neill, Fergal Reid

    cs.LG · cs.CL

    Looped, weight-tied Transformers reduce parameters by reusing a block, but decoding still stores a separate K/V cache for every recurrence step. We show that this loop-indexed cache is highly structured. For a fixed token, layer and head, K/V vectors trace a short low-rank trajectory across loops, while the head and layer axes remain much flatter. We introduce Looped Latent Attention (LLA), a post-training cache codec that stores compact K...

    arxiv.org/abs/2607.15456 · PDF

  60. 60

    Relevant and Irrelevant: A Renormalization Group Analysis of Transformer Attention

    Parviz Haggi-Mani, Irina Rish

    cs.LG

    Using the language of Wilsonian renormalization group theory (RG), we treat the Transformer's attention mechanism as a perturbation of the trained MLP residual-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operator. We derive a fixed-point shift formula and obtain four testable predictions for the fixed-point geometry, effective rank profile, layer specificity, and perturbation decay spectrum. Testing...

    arxiv.org/abs/2607.15449 · PDF

  61. 61

    LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models

    Jingteng Li, Alexander Capstick, Louise Rigny, Iona Biggart, Neil J Sebire, Payam Barnaghi

    cs.LG

    Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models to foundation models, utilising modern representation learning methods. Here, foundation models are pre-trained on mixtures of complex clinical data modalities, useful for various downstream tasks. Existing works often utilise Electronic Health Records (EHR) to provide rich and diverse patient...

    arxiv.org/abs/2607.15447 · PDF

  62. 62

    Who Became Financially Vulnerable After COVID-19? A Population-Level Machine Learning Analysis Using MEPS Data

    Alexey Kresin, Zien Cheng, Ammar Ahad, Ebiyomare Kelvin, Manish Sivaratri, Prabhjeet Singh, Omar Aljawfi, Olabisi...

    cs.LG

    The cost of healthcare remains a concern in the United States and may have been influenced by disruptions associated with the COVID-19 pandemic. This study examines healthcare financial vulnerability before and after the pandemic using Medical Expenditure Panel Survey (MEPS) data from 2019 and 2021. High financial burden was defined as out-of-pocket healthcare expenditures exceeding 10% of family income. Survey-weighted subgroup analyses were...

    arxiv.org/abs/2607.15446 · PDF

  63. 63

    Stochastic Reset Pathfinding: Path-Level Regret for Cascading Bandits over Graph Paths

    Guni Sharon, Wei Zhang

    cs.LG

    We introduce Stochastic Reset Pathfinding (SRP), an episodic learning problem on a known directed graph with unknown stationary edge success probabilities. In each episode, the agent commits to a source-to-goal path, and any edge failure during execution resets it to the source. SRP captures settings such as entanglement distribution in quantum repeater networks, payment routing on the Lightning Network, and delivery in unreliable mesh...

    arxiv.org/abs/2607.15440 · PDF

  64. 64

    From hyperplanes to hyperellipsoids: characterizing the inherent interpretability of linear and single-qubit mixed-state binary classification models

    Kaitlin Gili

    cs.LG · quant-ph

    We characterize and compare the inherent interpretability offerings of a standard linear model with a single qubit mixed state model for the task of supervised binary classification. A side by side comparison reveals that a single qubit mixed state model for binary classification is just the ``ellipsoid version" of standard linear model classification. More precisely, rather than learning a hyperplane to classify data, we learn a...

    arxiv.org/abs/2607.15433 · PDF

  65. 65

    qZACH-ViT: Quantization-Aware Intrinsic Explanations with Recursive Attribution-Stabilized Optimization

    Athanasios Angelakis

    cs.LG · cs.CV

    Compact medical-image classifiers need efficiency and interpretable evidence, yet these goals are often addressed separately. We introduce qZACH-ViT, a quantization-aware extension of the zero-token (CLS-token-free), position-free ZACH-ViT backbone with recursive intrinsic patch-level class evidence. We also introduce Recursive Attribution-Stabilized Optimization (RASO), which norm-matches classification and attribution gradients and removes...

    arxiv.org/abs/2607.15421 · PDF

  66. 66

    AI Trading: Evaluating Large Language Models for Technical Market Analysis

    Geofrey Ntale

    cs.LG · cs.AI · q-fin.CP

    Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial markets. This paper presents a systematic, comparative evaluation of five prominent LLMs: GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro, Llama 3 70B, and the domain-specialized FinGPT, with respect to their capacity for technical market analysis. The evaluation spans four structured tasks: candlestick pattern...

    arxiv.org/abs/2607.15414 · PDF

  67. 67

    Regularity-Aware Stochastic MGDA with Adaptive Conflict-Avoidant Update Direction Control

    Chentong Huang, Lisha Chen

    cs.LG · math.OC

    Multi-objective learning (MOL) aims to optimize multiple objectives simultaneously. The multi-gradient descent algorithm (MGDA) is a workhorse that iteratively updates along a common descent or conflict-avoidant (CA) direction across objectives. In stochastic settings, however, the vanilla stochastic MGDA method, SMG, lacks a fast convergence rate because mini-batch sampling introduces noise in the gradients. This causes bias in the update...

    arxiv.org/abs/2607.15412 · PDF

  68. 68

    A Transportable Threshold-Based Framework for Interpretable Classification of Medical Data

    Antony Garcia, Adrian Noriega, Gabrielle Britton, Xinming Huang

    cs.LG

    Black-box models limit the adoption of artificial intelligence in medicine due to their lack of interpretability and reproducibility. We introduce a statistically grounded framework that provides fully interpretable, rule-based clinical classification using the Bernoulli Naïve Bayes (BNB) model. The method applies supervised $χ^2$-guided statistical binarization to continuous variables, identifying thresholds that maximize association with...

    arxiv.org/abs/2607.15394 · PDF

  69. 69

    Decoding Market Emotion from Blockchain Activity: A Data-Driven Sentiment Classifier

    Arthur G. Bubolz, Abreu Quevedo, Giancarlo Lucca, Rafael A. Berri, Eduardo Borges, Bruno L. Dalmazo

    cs.LG · cs.CE

    The growing use of Bitcoin as a decentralized digital asset and investment tool has sparked strong interest in understanding its market behavior. This study presents a new approach to analyze Bitcoin market sentiment by combining on-chain and financial data with social media posts. Unlike models that aim to predict prices, this work focuses on explaining market sentiment using blockchain transactions, historical price data of Bitcoin, and...

    arxiv.org/abs/2607.15258 · PDF

  70. 70

    Mutable Low-Rank Sketches for Retrain-Free Recommendation

    Hector J. Garcia, Nick Clayton

    cs.LG

    A common bottleneck in two-stage recommendation is embedding staleness: when a user rates a new item, their embedding remains fixed until the next retrain cycle. We propose mutable sketches, which store each user's preferences in a KP-tree (a sparse segment tree with sum aggregation), fit a low-rank projection once, and recompute embeddings on-the-fly as ratings arrive. We prove that each new observation monotonically tightens the prediction...

    arxiv.org/abs/2607.15242 · PDF

  71. 71

    Data Driven Block Replacement Scheduling

    Aniruddhan Ganesaraman, VIdyadhar Kulkarni

    cs.LG · math.OC · stat.AP · stat.ML

    We develop data-driven algorithms for maintaining $N$ independent identical machines under a \textit{block replacement policy}, in which each machine is replaced upon failure and all machines are jointly replaced at regular intervals of length $k$. The goal is to learn the cost-minimizing interval $k^*$ from operational data when the lifetime distribution is unknown. At each decision epoch, the operator selects $k \in \{1, 2, \ldots, K\}$,...

    arxiv.org/abs/2607.15229 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.