cs.LG · 2026-08-20 · No. 90

Machine Learning, 2026-08-20.

73 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

73 entries
  1. 01

    Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

    Zhu Zhang, Jixun Wang, Xiaoang Xu, Xiaorong Wang, Zihan Zhou, Zhiyuan Wang, Shuo Wang, Chaojun Xiao, Yuezhi Zhou

    cs.LG · cs.AI · cs.CL

    On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible responses that omit evidence distributed across the input or violate global task constraints. Task-specific verifiers, in contrast, evaluate task completion at the response level and may return graded rewards that reflect partial...

    arxiv.org/abs/2608.19181 · PDF

  2. 02

    Lévy Attention: Single-Pass Predictive Uncertainty for Continuous-Time Attention

    Sotirios P. Chatzis, Loukas Papadoulas

    cs.LG

    Deep models for irregularly-sampled time series answer queries at arbitrary continuous timestamps, yet report nothing about how far each answer should be trusted. We show the attention layer itself can close that gap: with the right stochastic formulation, the pass that makes each prediction also reports, in closed form and at no extra cost, how far it should be trusted. We introduce Lévy Attention, a cross-attention operator whose output is...

    arxiv.org/abs/2608.19171 · PDF

  3. 03

    Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training

    Zachary Speck, Asa Shepard

    cs.LG

    A single training example's contribution to a finished model is normally estimated rather than measured, because measuring it takes two expensive full pre-training runs that differ in one row of one batch. We ran that counterfactual 24 times at a small scale. We trained 32 GPT-2 models at 124M parameters from scratch on OpenWebText, over four conditions and eight seeds. At step 200 of 9,536, at peak learning rate, we replaced one row of a...

    arxiv.org/abs/2608.19168 · PDF

  4. 04

    Continuous-Time Reinforcement Learning for Controlled Hawkes Jump-Diffusions

    Tomasz R. Bielecki, Thibaut Mastrolia, Haoze Yan

    cs.LG · math.OC · stat.ML

    We study stochastic control of multivariate Hawkes-driven stochastic differential equations with machine learning algorithms in a non-Markovian setting. Due to the path dependence of the memory of the Hawkes intensity, this problem does not fall within classical stochastic control theory outside particular Markovian kernels. We first develop a finite-dimensional Markovianization procedure and algorithm to approximate multivariate Hawkes...

    arxiv.org/abs/2608.19151 · PDF

  5. 05

    SCORE: Subject Coordinate Recovery for Label-Free Cross-Subject EEG-to-Image Retrieval

    Zhenyao Cui, Siyuan Kan, Siyang Li, Ziwei Wang, Dongrui Wu

    cs.LG

    Accurate visual decoding can reveal how the brain represents visual information and recover perceived content from neural signals such as electroencephalography (EEG), with potential for neural communication. However, current EEG-to-image retrieval methods perform far below their within-subject counterparts for new users without labeled calibration, limiting real-world deployment. To understand this gap, we analyze EEG features across...

    arxiv.org/abs/2608.19134 · PDF

  6. 06

    Beyond Trial Averaging: Anchoring Neural and Visual Representations for Few-Repetition Brain-to-Image Retrieval

    Zhenyao Cui, Siyuan Kan, Dingkun Liu, Dongrui Wu

    cs.LG

    Decoding visual information from brain signals probes neural representations and enables neuro-rehabilitation and dream decoding. Recent brain-to-image retrieval approaches have achieved promising performance, typically by averaging many (up to 80) neural trials per image, requiring repeated stimulus presentation that increases latency, cost, and user burden. When only one or a few repetitions are available, the retrieval accuracy drops...

    arxiv.org/abs/2608.19128 · PDF

  7. 07

    Leaf Values as Coordinates: Exact Contrastive Explanation for Gradient-Boosted Ensembles

    Emanuele Luzio

    cs.LG · cs.AI · cs.CY

    A gradient-boosted ensemble predicts by summing one leaf value per tree. Read those values as coordinates rather than as intermediate results, and every instance becomes a point in R^M on which the model acts linearly: the score is the sum of the coordinates. This small change of view makes contrastive explanation exact. The difference between two instances is a vector that is identically zero wherever they share a leaf, so the gap between a...

    arxiv.org/abs/2608.19127 · PDF

  8. 08

    PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints

    Boqiao Zhang, Godbless James, Sai Krishna Gottipati, Andrew Fitzgibbon

    cs.LG · cs.AI

    Improving molecular properties, such as drug-likeness or binding affinity, is a recurring task in early-stage drug discovery. However, molecules optimized in an unconstrained chemical space have limited practical value if they cannot be synthesized. Policy Gradient for Forward Synthesis (PGFS) is a synthesis-aware reinforcement learning method for molecular improvement, but its use of reactant embedding prediction makes reactant selection...

    arxiv.org/abs/2608.19121 · PDF

  9. 09

    Discretizing Continuous Time Series for Imputation with Masked Diffusion Training

    Dongbin Kim, Seungyun Lee, Geonwoo Shin, Jaewook Lee

    cs.LG · cs.AI

    Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the complex temporal dynamics and noise of real-world data. Existing approaches, however, exhibit two limitations: missing and observed values are embedded within the same representation space without explicit structural separation, and continuous diffusion-based methods are trained to predict added noise rather than the original...

    arxiv.org/abs/2608.19119 · PDF

  10. 10

    Enhancing EBSD throughput of battery electrode materials using super-resolution generative adversarial networks

    John Mangum, Andrew Glaws, Francois Usseglio-Viretta, Steven Spurgeon, Donal Finegan

    cs.LG · cond-mat.mtrl-sci

    Quantitative microstructural characterization of Li-ion battery electrode materials using electron backscatter diffraction (EBSD) has been proven as a critical method for optimizing cell performance. However, the inherently slow nature of EBSD can hinder the throughput of analyses needed for statistical representation of a material microstructure being developed. This work demonstrates a machine learning super-resolution framework using a...

    arxiv.org/abs/2608.19117 · PDF

  11. 11

    Pretraining Reusable Inference Across Views with Synthetic Task Priors

    Jielong Lu, Zhihao Wu, Jiajun Yu, Zhaoliang Chen, Haishuai Wang

    cs.LG · cs.MM

    Modern pretrained encoders make representations from heterogeneous views increasingly reusable, but the procedure that determines view utility and combines evidence is still relearned for each downstream task. Consequently, knowledge about view relevance, complementarity, reliability, and missingness is repeatedly discarded rather than transferred across tasks. We therefore reformulate multi-view learning as learning a reusable,...

    arxiv.org/abs/2608.19115 · PDF

  12. 12

    Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

    Huan-ang Gao, Haohan Chi, Yong Yan, Shiyuan Feng, Hanlin Wu, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou

    cs.LG · cs.AI · cs.CL

    Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist student via dense, token-level reward supervision. Despite its practical success, the optimization dynamics governing multi-teacher capability integration remain poorly understood, and open, rigorously reproducible recipes are conspicuously lacking. In this work, we...

    arxiv.org/abs/2608.19098 · PDF

  13. 13

    Does Mapping Non-Maximal Probabilities to GMM Components Matter for S-JEPA Encoder Representations?

    Wenxuan He, Yunpeng Li, Shan Liang

    cs.LG · cs.SD

    S-JEPA uses soft Gaussian mixture model (GMM) posteriors instead of hard cluster labels to preserve uncertainty. It remains unclear whether the probability values alone are sufficient, or whether it also matters which GMM components receive the non-maximal probabilities. We test this with two matched controls. FIXED-RANDPERM keeps the top-1 component and probability together with the multiset of non-maximal probability values, but reassigns...

    arxiv.org/abs/2608.19084 · PDF

  14. 14

    Multi-Agent Off-Policy Deep Reinforcement Learning for Smart Campus Coverage

    Omar Rady, Mohamed Ayman, Ali Arafa, Mohamed Shalma

    cs.LG · eess.SP

    Deep reinforcement learning (DRL) has recently gained a great attention due to its real-time adaptation and effectiveness in complex optimization problems. This paper investigates the optimal deployment of millimeter-wave (mmWave) base stations (BSs) in a realistic, non-convex campus topology. The optimization problem is NP-hard, due to the non-convex, non-smooth nature of the max-min fairness objective. To overcome these constraints, we...

    arxiv.org/abs/2608.19049 · PDF

  15. 15

    Harness Continual Learning: Continual Adaptation Beyond Model Parameters

    Borui Kang, Jinrui Gu, Junhan Lv, Wenbin Li, Lei Wang, Yang Gao

    cs.LG · cs.AI

    Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. Modern agents can also adapt through a harness of prompts, memories, tools, skills, and routing rules. Because these contents jointly shape later execution, a harness update can disrupt previously reliable behavior even when the model is frozen. This raises a new question: how can an agent continually improve its...

    arxiv.org/abs/2608.19013 · PDF

  16. 16

    Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

    Blazej Banaszewski, Andrew W. Fitzgibbon

    cs.LG · q-bio.BM · q-bio.QM

    Bioassay activity prediction is often data-limited because drug-discovery datasets rely on time-consuming and expensive wet-lab experiments for data generation and evaluation. This challenge has inspired recent research into molecular foundation models (MFMs), which aim to encode general-purpose chemical knowledge into molecular representations that generalize well in data-constrained scenarios. This paper presents Monroe, a new MFM with...

    arxiv.org/abs/2608.18982 · PDF

  17. 17

    Fuzzy Accuracy Compensates for Label Subjectivity in Classification of Skin Tone Using Wearable Photoplethysmography Signals

    Padmini Krishnadas, Urs Hackstein, Alen Bosnjakovic, Philip J. Aston

    cs.LG

    We consider the problem of classification of skin tone using photoplethysmography (PPG) signals with labels of the ordinal six-class Fitzpatrick skin tones. A typical accuracy for this task is a poor 40-55 %. However, the labels are subjectively determined by comparing the skin with a colour chart, and hence contain widespread small-scale inaccuracies. By working with a "fuzzy accuracy", which deems a prediction of skin tone class to be...

    arxiv.org/abs/2608.18969 · PDF

  18. 18

    Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

    Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper,...

    cs.LG · cs.AI · cs.CE · cs.CL

    Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified...

    arxiv.org/abs/2608.18940 · PDF

  19. 19

    Graphical Design of Interpretable Architectures

    Pietro Barbiero

    cs.LG · cs.AI · cs.NE

    Designing, implementing, and comparing interpretable architectures requires a formal language to represent them. The most common representations fall short in one of two ways. Symbolic equations give no global view of an architecture at a glance. Probabilistic graphical models and flowcharts do not describe actual tensor manipulations, thus hiding key insights and limiting reproducibility. To close this gap, we introduce a graphical notation...

    arxiv.org/abs/2608.18936 · PDF

  20. 20

    Transportable Causal Effect Estimation across Networks under Interference

    Xiaojing Du, Jiuyong Li, Lin Liu, Debo Cheng, Jixue Liu, Thuc Duy Le

    cs.LG

    Estimating causal effects under network interference typically assumes that the network used for training and the network used for deployment coincide. In practice, an intervention is run on one population while the question of interest concerns a different population, and the two generally differ in topology, node-covariate composition, and spillover pathways. Transporting a causal effect across networks is therefore a data-fusion problem...

    arxiv.org/abs/2608.18932 · PDF

  21. 21

    Lost in Aggregation: How Benchmarks Overlook Irreplaceable Model Strengths

    Andrej Tschalzev, Stefan Lüdtke, Heiner Stuckenschmidt, Christian Bartelt

    cs.LG

    Tabular machine learning benchmarks typically summarize performance by averaging scores, ranks, or pairwise wins across datasets. Such aggregates are useful for selecting robust default models, but they can obscure a different question: which models are necessary to attain peak performance on particular datasets? We argue that benchmark evaluation should also consider the data-centric peak performance frontier, defined by the best...

    arxiv.org/abs/2608.18919 · PDF

  22. 22

    Score the Algebra, Not the Span: Dimension Reduction for Transfer Operator Models of Dynamical Systems

    Mark Kozdoba, Shie Mannor

    cs.LG

    Dimension reduction for dynamical systems is standard practice, and the standard route is spectral: model the transfer (Koopman) operator by its leading modes. We show that on systems assembled from several weakly interacting components --- a structure common in physical and biological settings --- this may either require an exponential number of modes, or drop an entire component: the component is absent from the model rather than modeled...

    arxiv.org/abs/2608.18918 · PDF

  23. 23

    Converting Expert Deliberation into Financial Signals Through A Context-Aware NLP Pipeline

    Vivek Batra, Kristin Chen, Sanjiv Das, Samuel Judge, Harshad Khadilkar, Sukrit Mittal, Amir Nasrollahzadeh, Daniel...

    cs.LG · cs.CE

    We introduce the CDSP (context-conditional deliberation signal pipeline), converting an investment committee's meeting transcripts into structured predictive features. CDSP segments the meeting transcripts into topical chunks, assigns asset-class context labels using a large language model (LLM), maps financial keywords to a pre-determined taxonomy of labels, and constructs complementary features: sentiment polarity and mention frequency....

    arxiv.org/abs/2608.18911 · PDF

  24. 24

    On the Slow Convergence to Trivial Solutions of Algorithms for Hard Optimization Problems

    Ali Hussaini Umar, Jean Barbier, Matthieu Jonckheere, Manuel Sáenz

    cs.LG · cond-mat.dis-nn · cs.DM · math.PR

    Hard combinatorial optimization problems, many of which are NP-hard, present fundamental algorithmic challenges. Average-case analysis on random instances has emerged as a powerful framework for understanding typical algorithmic performance beyond worst-case guarantees. A substantial body of work has established negative results: for sufficiently hard instances (often controlled by the underlying graph connectivity/constraints density), no...

    arxiv.org/abs/2608.18910 · PDF

  25. 25

    A FEM-Based Surrogate Modelling and Optimization Framework for Physics-Constrained Electromagnetic Coil Design

    Yucheng Liu

    cs.LG

    This work evaluates surrogate-assisted optimization of a seven-parameter current-excited coil--core benchmark subject to geometric, manufacturing, and separate core and copper mass constraints. A Python--MPh--COMSOL workflow couples a two-dimensional axisymmetric finite-element method (FEM) model to a Matern 5/2 Gaussian-process (GP) probabilistic surrogate. Here, physics-constrained denotes a design problem evaluated by a governing-equation...

    arxiv.org/abs/2608.18903 · PDF

  26. 26

    Graph-Based Approaches to Learning Epileptogenic Zone Localization Using Stereo-EEG Recordings

    Daniel Wendelken, Brian Ervin, Ravindra Arya, Ali A. Minai

    cs.LG

    The epileptogenic zone (EZ) is the brain region that generates seizures in an individual, and is the target of epilepsy surgery. Localizing the EZ from stereo-EEG (sEEG) recordings supports surgical planning, but manual interpretation is time-consuming and focuses on seizure recordings. Graphical learning models of resting-state functional connectivity among the recorded brain regions are an attractive alternative, but depend crucially on the...

    arxiv.org/abs/2608.18887 · PDF

  27. 27

    Multi-stage neural operator learning with application for convolutions

    Zhiping Mao, Zhenye Wen, Yong Zhang, Xiaofei Zhao

    cs.LG · math.NA

    Convolution integrals widely exist in applications, and to enable fast and accurate computations, this paper introduces two general multi-stage neural operator learning frameworks. The first, Deep Collocation Neural Operator (DCNO), is a supervised approach that iteratively refines the operator approximation by learning residuals from input-output data pairs. The second, Deep Galerkin Neural Operator (DGNO), is an unsupervised framework...

    arxiv.org/abs/2608.18851 · PDF

  28. 28

    GEAR: Generative Expansion and Real Anchoring for Two-Stage Distillation of Tabular Foundation Models

    Qi Qin, Jiajie Zhu, Dali Chen, Yuzhao Zhang, Jia-Xing Han, Yu Su, Peng Zhang, Ying Yan, Yifan Sun

    cs.LG · stat.ME · stat.ML

    Tabular foundation models (TFMs) achieve strong performance through in-context learning, but context-dependent inference imposes substantial latency and memory costs, hindering large-scale deployment. We propose GEAR (\emph{Generative Expansion and Real Anchoring}), a modular two-stage framework that distills TFMs into lightweight MLP or tree-based predictors that can be deployed on commodity CPUs. Stage 1 uses synthetic covariates solely as...

    arxiv.org/abs/2608.18849 · PDF

  29. 29

    MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models

    Chenglin Liu, Xun Wang, Ruishuo Chen, Zhuoran Li, Longbo Huang

    cs.LG · cs.AI · cs.CL

    Reward function design remains a bottleneck in reinforcement learning. While large language models (LLMs) have enabled automated reward generation, existing methods generate and revise reward functions as monolithic programs, making it difficult to reliably preserve and reuse effective components discovered in earlier iterations, leading to unstable performance across iterations. To address this, we propose Module Level Reward Evolution...

    arxiv.org/abs/2608.18827 · PDF

  30. 30

    A Unifying Relational Perspective on Expressive Lottery Tickets

    Lorenz Kummer, Samir Moustafa, Anatol Ehrlich, Franka Bause, Marco Nennstiel, Przemysław Andrzej Wałȩga, Nils Morten Kriege

    cs.LG · stat.ML

    Graph neural networks (GNNs) are widely used, but how parameter sparsity affects the expressivity of relational (RGNNs) and temporal (TGNNs) variants is poorly understood. The Strong Expressive Lottery Ticket Hypothesis (SELTH) posits the existence of sparse GNNs that preserve Weisfeiler-Leman (WL) expressivity on static graphs. We generalize this existence result to a probabilistic statement for multi-relational and temporal domains via the...

    arxiv.org/abs/2608.18819 · PDF

  31. 31

    Many Optimizers But Only One Training Path: Repeated Resampling for Adaptive Optimizer Selection

    Ronald Richman, Mario V. Wüthrich

    cs.LG

    An optimizer is usually chosen before training a deep neural network and then kept fixed. Treating optimizer choice as a hyperparameter could boost performance, but it requires several complete training runs and discards all but the winner. Repeated Optimizer Resampling (ROR) instead searches during one evolving run. Every $b$ epochs, each candidate optimizer scouts from the current model weights for $s$ epochs. The best scout continues for...

    arxiv.org/abs/2608.18810 · PDF

  32. 32

    Tensor Field Models

    Alexander Strunk, Roland Assam

    cs.LG · math.DG

    This paper introduces Tensor Field Models (TFMs), realization-level Mathematical Structures in which a learned Operator maps a product of admissible component-section families to a prescribed family of time-dependent tangent sections on a Generative State Manifold. Analytic and dynamical restrictions are encoded through the choice of admissible families rather than imposed by the root definition. Constructed, component-separable, and Tensor...

    arxiv.org/abs/2608.18808 · PDF

  33. 33

    Forgetting, plasticity, and co-observation: a third facet of continual learning

    Timm Hess, Abhishek Jha, Gido M. van de Ven, Tinne Tuytelaars

    cs.LG · cs.AI

    Efficient continual learning remains a fundamental challenge for deep neural networks. While catastrophic forgetting and loss of plasticity are widely considered the primary obstacles to overcome, we show that these two issues cannot fully explain the performance gap between naive sequential training and offline joint training. In this paper, we highlight data co-observation as a distinct factor influencing continual learning performance. By...

    arxiv.org/abs/2608.18803 · PDF

  34. 34

    A Real-Time Tsetlin Machine-based Non-intrusive Load Monitoring System on MCUs

    Tianhang Tan, Han Wu, Tousif Rahman, Shengyu Duan, Alex Yakovlev, Rishad Shafik

    cs.LG

    Non-Intrusive Load Monitoring (NILM) systems estimate individual appliance energy consumption from a single aggregate meter, without requiring separate sensors for each device. By installing a single meter that measures a building's total electricity consumption, NILM algorithms can determine the active status of each appliance. However, traditional NILM systems use computationally intensive optimization algorithms to process offline data,...

    arxiv.org/abs/2608.18780 · PDF

  35. 35

    GraphK: Variable-Size Graph Generation with Efficient Edge Construction

    Resul Tugay, Eren Oluğ, Elif Ak, Sule Gunduz Oguducu

    cs.LG

    Graph generation models have advanced significantly with deep learning, yet they remain limited in scalability, flexibility, and ability to model underlying structures. We present GraphK, a novel encoder-sampler-decoder framework for graph generation that overcomes these challenges through structural flexibility and computational efficiency. Unlike autoregressive approaches constrained by vocabulary size (i.e. number of nodes in graph...

    arxiv.org/abs/2608.18777 · PDF

  36. 36

    To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

    Taehyung Kim, Jongeun Choi

    cs.LG · cs.RO

    Learning a reward model from human feedback and optimizing a policy against it is one approach to aligning AI systems with individual users. From a fairness perspective, existing work improves such alignment by developing data-efficient and accurate reward models that capture minority preferences despite scarce data. We push this line of inquiry one step further and argue that data-efficient and accurate per-user reward models are not...

    arxiv.org/abs/2608.18770 · PDF

  37. 37

    Enhancing Distance-Based Graph Autoencoders with Structural Penalties for Dynamic Graph Embedding

    Aleksandar Tomčić, Miloš Savić, Miloš Radovanović

    cs.LG · cs.ET

    Graph autoencoders (GAEs) are widely used for learning representations of dynamic graphs. However, their optimisation objectives typically do not take structural heterogeneity across nodes into account. We propose three distance-based GAE variants that incorporate structural penalties into the reconstruction loss. All variants share a two-layer Graph Convolutional Network encoder and a Euclidean-distance decoder trained with distance-based...

    arxiv.org/abs/2608.18762 · PDF

  38. 38

    Beyond Predictive Fairness: Quantifying Attribution Consistency Across Demographic Groups in Diabetic Retinopathy Screening

    Kerol Djoumessi, Philipp Berens

    cs.LG · cs.AI

    Fairness in medical imaging is commonly evaluated through subgroup performance metrics, yet it remains unclear whether models rely on consistent visual evidence across demographic groups. This work introduces the Explanation Consistency Score (ECS), a fairness-aware metric based on Jensen-Shannon divergence that quantifies the similarity of attribution maps across subgroups. Using diabetic retinopathy screening as a case study, ECS is...

    arxiv.org/abs/2608.18759 · PDF

  39. 39

    Geometric Data Perturbation with Noisy-Anchor Alignment for Privacy-Preserving Collaborative Learning

    Keiyu Nosaka, Yamato Suetake, Yuichi Takano, Yukihiko Okada, Akiko Yoshise

    cs.LG · cs.CR

    Geometric Data Perturbation (GDP) enables one-shot, privacy-preserving collaborative learning: each participant applies a distance-preserving transformation to its private data and uploads only the resulting representation to a central analyst. We study GDP under analyst-participant collusion, in which the analyst combines all uploaded representations with the private data and transformations disclosed by colluding participants to recover a...

    arxiv.org/abs/2608.18749 · PDF

  40. 40

    Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

    Jiawei Wang, Ke Rui, Yushen Zuo, Yichun Feng, Minglei Li

    cs.LG · cs.CV

    JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (MPC). Strong decoding of task variables, however, does not guarantee that this particular cost ranks candidate action sequences by real task progress. We call the latter property \emph{decision-metric alignment}. We introduce Plan-Real Spearman, which measures latent--real rank agreement on random plans, and CEM-stage Spearman,...

    arxiv.org/abs/2608.18746 · PDF

  41. 41

    FedLNS: Leverage LayerNorm Signature Modeling to Mitigate Adversarial Manipulation in Federated LLMs

    Kai Li, Jong-Ik Park, Carlee Joe-Wong, Wei Ni, Falko Dressler

    cs.LG · cs.CR · cs.DC

    Federated training enables language models to learn from distributed private text, but the server cannot directly verify the local supervision or optimization process that produces each client update. A malicious client can therefore train on corrupted targets, introduce incorrect context-token associations, and degrade the global model through repeated aggregation. Such degradation can also increase the risk of unreliable or hallucinatory...

    arxiv.org/abs/2608.18736 · PDF

  42. 42

    Visual-Aware Representation of Web Pages for Machine Learning Applications

    Radek Burget, Radek Hranický

    cs.LG · cs.IR

    Applying machine learning to web pages is challenging due to the need to interpret HTML together with associated resources and perform rendering to obtain a meaningful visual and layout-aware representation. As a result, machine learning over web content remains comparatively underexplored. In this paper, we present a platform for visual-aware representation and machine learning over web pages based on the open-source rendering tool...

    arxiv.org/abs/2608.18727 · PDF

  43. 43

    Multi-Class Electrical and Mechanical Fault Classification Using Random Convolutional Kernels

    Mouhamadou Mansour Lo, Mouad Talbaoui, Gildas Morvan, Mathieu Rossi, Fabrice Morganti, David Mercier

    cs.LG

    Diagnosing faults in rotating machinery is essential for ensuring the reliability of industrial processes. Random convolutional kernel-based Time Series Classification (TSC) methods, such as ROCKET and its variants, provide an attractive trade-off between predictive performance and computational efficiency. In this work, we evaluate SelF-Rocket for the multi-class diagnosis of both mechanical and electrical faults and introduce, as a new...

    arxiv.org/abs/2608.18716 · PDF

  44. 44

    Europe's Climate Ambition Under Scrutiny: Evidence from Deep Learning Emission Projections

    Jacopo Ghirri, Carlos Rodriguez-Pardo, Lara Aleluia Reis, Massimo Tavoni

    cs.LG · cs.AI · econ.GN

    The European Union has committed to reducing greenhouse gas emissions 55% below 1990 levels by 2030, but whether current trends are compatible with this ambition remains uncertain. We apply deep learning to high-resolution socioeconomic and sectoral data across EU27 member states till 2023 to project sectoral CO$_2$ trajectories under current trends, extrapolating observed sectoral momentum without assuming changes in the pace or...

    arxiv.org/abs/2608.18690 · PDF

  45. 45

    Transforming Heart Disease Prediction with Advanced Machine Learning Techniques

    Sami Ullah, Muhammad Mohsin Khan

    cs.LG

    Heart disease remains the leading cause of mortality globally, necessitating early and accurate detection to improve patient outcomes. This research focuses on the predictive analysis of heart disease using machine learning (ML) techniques, comparing the performance of multiple classifiers to identify the most accurate and least error-prone method. Two datasets from UCI and Kaggle repositories were utilized, each containing 14 attributes...

    arxiv.org/abs/2608.18687 · PDF

  46. 46

    An Empirical Benchmark of Deep Time-Series Models for Smart Meter Energy Forecasting

    Behnaz Kavoosighafi, Maria Eidenskog, Wiktoria Glad, Katerina Vrotsou

    cs.LG

    Accurate forecasting of energy consumption is important for the efficient operation of power systems, with direct implications for operational costs, energy management, and system maintenance. Due to the availability of extensive high-resolution consumption data from smart meters, data-driven methods have been used for short-term and long-term forecasting. However, their comparative performance on real-world smart meter data is still not well...

    arxiv.org/abs/2608.18675 · PDF

  47. 47

    Reinforced Planning with Latent World Models

    Armin Sommer, Jannik Schilling

    cs.LG

    Humans solve complex problems by constructing plans and mentally simulating their outcomes with an internal model of the world. Machine learning has produced world models that similarly predict the outcomes of action sequences, but the improvement of candidate plans still isn't fully learned. Current planners are either hand-designed, distilled from a hand-designed optimizer, or learned only to inform an amortized policy rather than to revise...

    arxiv.org/abs/2608.18669 · PDF

  48. 48

    Computational Measurement of Team-Process Phase Dynamics in Collaborative Virtual Reality

    Qing Huang, Jianing Zhang, Pooja Pol

    cs.LG

    Collaborative virtual reality (VR) environments make team communication observable as it unfolds, but conventional transcript analyses often summarize entire trials or divide them into fixed temporal windows. Such approaches can obscure changes in team communication and coordination over time. This article presents a computational framework for detecting and interpreting dynamic team-process phases from timestamped dialogue in a collaborative...

    arxiv.org/abs/2608.18660 · PDF

  49. 49

    FlashAttention for Scalable Vector Architectures

    Sonia Rani Gupta, Nikela Papadopoulou, Miquel Pericàs

    cs.LG · cs.PF

    Inference with transformer models on CPUs is increasingly important, especially for Small Language Models (SLMs), where vector architectures are emerging as a promising execution substrate. The attention module is a major bottleneck due to high memory bandwidth requirements; FlashAttention mitigates this by fusing operations to improve data locality and reduce intermediate memory traffic. In this paper, we present FlashAttention-V, a blocked...

    arxiv.org/abs/2608.18656 · PDF

  50. 50

    ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets

    Dipesh Tharu Mahato, Pramod Dhungana

    cs.LG · stat.ME · stat.ML

    Researchers often choose a proxy dataset from many releases, transformations, or seeds. Search can make an invalid release appear adequate, while one adequate release does not establish that its generator is reliable. ProxyGuard controls both errors using prespecified bounded risks and a sealed target set. Named-release mode corrects for multiplicity and certifies specific releases. Direct shared-target mode evaluates independent mechanism...

    arxiv.org/abs/2608.18643 · PDF

  51. 51

    Coordination on a Budget: Federated Active Learning with Few Labels

    Liam Mohr, Daphna Weinshall

    cs.LG

    Federated Active Learning (FAL) addresses the dual challenges of data privacy and label scarcity, where the absence of a global data view introduces additional hurdles for coordinated query selection. We study cross-silo FAL in the low-budget regime, where annotation decisions are most critical. We characterize, both theoretically and empirically, a heterogeneity reversal: in low-budget settings, homogeneous (IID) data requires stronger...

    arxiv.org/abs/2608.18634 · PDF

  52. 52

    Scalable Geospatial Machine Learning for Power-Line Asset Risk: Integrating Remote Sensing for Lightning and Vegetation Risk Modelling

    Artur Sokolovsky, Bhavik Merai, Moe Jafari, Muen Chen

    cs.LG

    Electric power networks are increasingly exposed to weather-sensitive failure mechanisms that require asset-level, spatially explicit risk modelling for effective intervention planning. This study contributes a modular, robust, and explainable probability-of-failure (PoF) modelling framework for utility asset management. The central contribution is an asset-level architecture that can be scaled to new environmental data sources and additional...

    arxiv.org/abs/2608.18611 · PDF

  53. 53

    Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings

    Yubo Wang, Shujie Cui, James Bailey, Hongzhi Yin, Wenyu Liang, Min Tang, Shiyue Qin, Weiqing Wang

    cs.LG · cs.AI · cs.CR

    Dense text embeddings are widely used in data mining, retrieval, and downstream machine learning systems due to their compact and semantically rich representations, but recent embedding inversion attacks have shown that they can expose substantial information about the original text, leading to serious privacy leakage risks. A common defense is to release perturbed embeddings by adding Gaussian noise, which is simple yet effective against...

    arxiv.org/abs/2608.18610 · PDF

  54. 54

    Off-Manifold Collapse in Guided Protein Language Models

    Shuibai Zhang, Xinchi Liu, Fred Zhangzhi Peng, Zhihan Yang, Shutong Wu, Yingzi Ma, Jiawei Zhang

    cs.LG

    Protein language models are widely used priors for protein sequence design, and a growing body of work controls them at inference time as an alternative to fine-tuning. Such guidance faces a dilemma: mild enough to preserve natural activation statistics, it barely moves the property; strong enough to move it, the generations become progressively harder to fold. We show the failure has a specific and cheaply detectable signature, an...

    arxiv.org/abs/2608.18597 · PDF

  55. 55

    Infrared Universality of Collective Dynamics across Transformer and State-Space Architectures

    Byung Gyu Chae

    cs.LG

    Whether distinct neural architectures develop common collective dynamics remains an open question. Recent analysis of Transformer language models revealed a nearly flat, weakly infrared-enhanced time-scale density of states (TDOS) associated with near-marginal long-memory dynamics. Here we test whether a closely related organization emerges in Mamba, whose selective state-space dynamics provides a fundamentally different microscopic...

    arxiv.org/abs/2608.18592 · PDF

  56. 56

    Beyond receptive fields: sequence-pooled normalization can supply most of a sequence labeler's context

    Qing Tian

    cs.LG · cs.NE

    A convolutional sequence labeler's receptive field is routinely treated as the extent of the model's usable context: it sets dilation schedules, bounds streaming horizons, and underwrites locality claims. However, we show that this can be false: when a normalization layer computes statistics from the current input along the sequence at inference, those statistics open a sequence-spanning path that bypasses the convolutional receptive field to...

    arxiv.org/abs/2608.18576 · PDF

  57. 57

    Continual Reasoning Gym: Diagnosing and Harnessing Shared Reasoning in Continual RLVR

    Lirui Luo, Guoxi Zhang, Hongming Xu, Rongqing Li, Cong Fang, Lifeng Fan

    cs.LG

    Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study continual RLVR, which updates the existing model as each task arrives. The central question is whether a model updated this way can perform as well as a jointly trained model. To answer this question, we introduce...

    arxiv.org/abs/2608.18574 · PDF

  58. 58

    NanoSleep: A Parameter-Efficient Hybrid Temporal Convolutional Network for Single-Channel Sleep Stage Classification

    S M Asif Hossain, Shruti Kshirsagar

    cs.LG

    Sleep stage classification from single-channel electroencephalography (EEG) is essential for wearable and home-based sleep monitoring. However, many deep learning models achieve high accuracy at the cost of large model sizes, which limits their deployment on resource-constrained devices. In this work, we present NanoSleep, a compact hybrid temporal convolutional network for automatic sleep stage classification. NanoSleep combines a learnable...

    arxiv.org/abs/2608.18571 · PDF

  59. 59

    MorphoGP: A Nonparametric Framework for Predicting Equilibrium Beach Profiles Under Tidal Influence

    Xi Wu, Yanqing Wei, Hang Yin, Pengze Li, Hongshuai Qi, Xi Chen

    cs.LG · cs.AI · physics.geo-ph

    The prediction of equilibrium beach profiles under tidal influence is of fundamental importance for sustainable coastal development, informing shoreline protection strategies and managing coastal ecosystems under changing environmental conditions. However, it remains challenging due to the highly nonlinear interactions among wave, tide, and sedimentary processes. Traditional empirical and numerical models often exhibit limited adaptability...

    arxiv.org/abs/2608.18558 · PDF

  60. 60

    Performance Drift Detection in Machine Learning as a Service (MLaaS) for IoT Environments

    Deepak Kanneganti, Sajib Mistry, Sheik Mohammad Mostakim Fattah, Erik Elmroth, Aneesh Krishna, Monowar Bhuyan

    cs.LG · cs.AI

    Machine Learning as a Service (MLaaS) is a powerful cloud paradigm enabling data-driven intelligent applications in Internet of Things (IoT) environments, widely adopted across healthcare, smart homes, and industry due to its cost-effectiveness. However, the dynamic nature of IoT frequently alters data distributions, affecting MLaaS stability, while periodic MLaaS updates further introduce performance drift. Unlike traditional ML systems,...

    arxiv.org/abs/2608.18555 · PDF

  61. 61

    MARCUS: Missing-Aware Region Representation with Contextual Urban Signals for Rent Prediction

    Chenya Huang, Bin Liang, Zhidong Li, Yuxi Lu, Kunqi Li, Justin Wang, Fang Chen

    cs.LG

    Multimodal urban data has expanded the applications of urban region representation learning, such as functional zone identification and real estate appraisal, but also introduces challenges caused by data incompleteness. Existing studies usually handle missing data through imputation, treating missingness as noise while ignoring its potential semantic value. To address this issue, we propose MARCUS, a missing-aware region representation model...

    arxiv.org/abs/2608.18546 · PDF

  62. 62

    Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions

    Ruiyang Qin, Qingzhuo Wang, Tian Wang, Zhihua Wei, Wen Shen

    cs.LG · cs.AI · cs.CL

    The remarkable capabilities of large language models (LLMs) are often undermined by their instability. Even subtle and semantically irrelevant changes in prompts can cause dramatic fluctuations in performance, a phenomenon known as prompt sensitivity. Previous studies typically evaluate prompt sensitivity by comparing the LLM's final outputs when prompts change. However, such coarse-grained metrics fail to explain the internal reasons for...

    arxiv.org/abs/2608.18539 · PDF

  63. 63

    LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

    Hanzhao Wang, Jingxuan Wu, Yumeng Li, Yu Pan, Guanting Chen

    cs.LG

    The growing demand for AI-driven workloads, particularly from Large Language Models (LLMs), has raised concerns about the significant energy and resource consumption in data centers. This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers. Our system utilizes an LLM to predict key metrics such as execution time and energy consumption...

    arxiv.org/abs/2608.18503 · PDF

  64. 64

    Tianmu-TC: Physics-constraints Generative Artificial Intelligence for Global Tropical Cyclone Forecasting

    Shiqi Zhang, Pan Mu, Cheng Huang, Hanting Yan, Yuchao Zhu, Jinglin Zhang, Shengyong Chen, Shoujuan Shu, Cong Bai

    cs.LG

    Tropical cyclones (TCs) pose severe risks from strong winds and heavy rainfall. However, forecasting their track and intensity remains challenging due to chaotic atmosphere and the rapid amplification of initial condition errors, leading to growing forecast uncertainty. While numerical weather prediction (NWP) and deep learning models have made progress, they remain computationally demanding and often fail under complex meteorological...

    arxiv.org/abs/2608.18500 · PDF

  65. 65

    Physics-Unrolled Neural Operator for Wireless Field Modeling

    Rafid Umayer Murshed, Saif Ur Rahman, Mingyue Tang, Elahe Soltanaghai

    cs.LG · cs.AI

    Radio maps are essential for wireless decision-making tasks such as access-point placement, coverage planning, and localization, but their fine spatial details are governed by complex propagation effects and are costly to simulate accurately. Machine learning offers a path to high-fidelity radio-map prediction without running expensive high-fidelity simulations for every scene. However, generating high-quality training labels at scale is also...

    arxiv.org/abs/2608.18495 · PDF

  66. 66

    ERASE: EaRly bAckpropagation SchEdule for Faster Training of Modern Recommendation Systems

    Ergan Shang, Flavio Sales Truzzi

    cs.LG · cs.AI

    Lightweight proxy models enable rapid experimentation without repeatedly training frontier-scale systems, but their small kernels often leave modern accelerators underutilized. Conventional training compounds this inefficiency by scheduling the forward and backward passes as disjoint phases, so spare capacity in one cannot be filled by work from the other. We reinterpret the detachment mechanism of Forward-Forward (FF) as a scheduling...

    arxiv.org/abs/2608.18469 · PDF

  67. 67

    Atrial Fibrillation Detection with Arbitrary Leads via a Codebook-Based Reconstruction-Classification Framework

    Hongtao Li, Jia Wei, Guoyao Li, Yuchen Lei, Guangnian Ma, Jia Xiao, Yuanjun Lai, Shuzhen Lv, Xueqiang Ouyang

    cs.LG · q-bio.QM

    \textbf{Background and Objective}: Reliable atrial fibrillation (AF) detection from electrocardiogram (ECG) signals remains challenging in real-world clinical settings due to variable lead configurations, cross-dataset domain shifts, and pervasive physiological and technical artifacts. So we develop a robust and generalizable deep learning model for accurate AF detection.\\ \textbf{Methods}: We propose the Dual-Codebook Graph Collaborative...

    arxiv.org/abs/2608.18451 · PDF

  68. 68

    Adaptive Multi-Agent Feature Selection for Personalized Fall Risk Prevention

    Chang Liu, Ladda Thiamwong, Yanjie Fu, Rui Xie

    cs.LG

    Falls among older adults represent a major public health challenge driven by complex, time-varying interactions across multiple risk domains. Effective fall risk factor identification requires learning from heterogeneous longitudinal data while accounting for sparse and delayed fall-related outcome events. However, existing approaches are largely static and fail to adaptively model evolving, individualized risk factors across modalities and...

    arxiv.org/abs/2608.18450 · PDF

  69. 69

    Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B

    Rahul Chowdhury, Timothy A Rupprecht, Senhao Cao, Jiahao Liu, Octavia Camps, David Bau, Pu Zhao, Yanzhi Wang

    cs.LG · cs.AI

    Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-series prediction. While LLMs display in-context learning capabilities, the mechanisms with which they accomplish time-series prediction remain unclear. Specifically, whether they truly understand the underlying structure, which at a minimum requires reasoning over first differences in the sequence of...

    arxiv.org/abs/2608.18419 · PDF

  70. 70

    The Road Taken: The Role of Optimizers at the Edge of Stability

    Jaerin Lee, Kyoung Mu Lee

    cs.LG

    The edge of stability refers to a phenomenon in deep learning with gradient-based optimizers where the Hessian eigenvalues of the loss remain stable above a threshold that the classical descent lemma predicts to be unstable. Previous works formulate the edge of stability with respect to the maximum Hessian eigenvalue and the learning rate. However, we observe that many first-order methods, including gradient descent, significantly violate the...

    arxiv.org/abs/2608.18415 · PDF

  71. 71

    Role-Conditioned Sub-Token Routing for Efficient Vision-Language-Action Policies

    Wei Jiang, Wei Wang

    cs.LG

    Vision-Language-Action (VLA) models process long multimodal token sequences, making inference expensive in both memory and computation. Existing efficiency methods mainly reduce visual tokens, but aggressive token pruning becomes fragile because removing a token discards its entire representation. Sub-token compression provides a complementary alternative by retaining more tokens while reducing their value width. However, directly applying...

    arxiv.org/abs/2608.18410 · PDF

  72. 72

    Vector Symbolic Policy Gradient

    Ryozo Masukawa, Sanggeon Yun, SungHeon Jeong, Hyunwoo Oh, Raheeb Hassan, Pietro Mercati, Nathaniel D. Bastian, Mahdi...

    cs.LG · cs.AI · cs.SC

    We answer this question with Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hypervector and scores it by similarity to the encoded state. Under the standard softmax policy-gradient surrogate, we prove that its update is exactly advantage-weighted hypervector bundling followed by normalization, and therefore supports standard advantage estimators. We further show that each trained...

    arxiv.org/abs/2608.18404 · PDF

  73. 73

    Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation

    Guiv Farmanfarmaian

    cs.LG · cs.AI · cs.CL

    When every candidate is wrong, correct-candidate selection is unavailable, yet the aggregation call can still solve the problem afresh. A correct aggregate answer may therefore reflect recombination, fresh solving, or both. For efficient test-time reasoning, the relevant question is whether candidate context adds value beyond the additional generation pass. We introduce the missing candidate-free control under the same maximum output-token...

    arxiv.org/abs/2608.18379 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.