cs.LG · 2026-09-28 · No. 127

Machine Learning, 2026-09-28.

57 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

57 entries
  1. 01

    User Model Extraction via Belief Self-Distillation

    Ali Holmov, Yiran Huang, Kirill Bykov, Zeynep Akata

    cs.LG · cs.CL

    Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling...

    arxiv.org/abs/2609.31603 · PDF

  2. 02

    New LoRA Skills Should Read but Never Write

    Zeyan Li, Panqi Yang, Qirong Guo, Shengda Zhuo, SIyuan Qiu, Hu Xu, Chun Li, Jianfeng Xu

    cs.LG

    Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly....

    arxiv.org/abs/2609.31600 · PDF

  3. 03

    Common-Mode Collapse and Recovery in Direct Feedback Alignment

    Varun Reddy, Bernardo L. Sabatini, Houman Safaai

    cs.LG · cs.NE

    Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching...

    arxiv.org/abs/2609.31589 · PDF

  4. 04

    Trust Guided Decision Transformer

    Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath

    cs.LG

    Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each...

    arxiv.org/abs/2609.31586 · PDF

  5. 05

    Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights

    Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodolà

    cs.LG

    We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a...

    arxiv.org/abs/2609.31564 · PDF

  6. 06

    Generalization behavior of OPTQ and the role of regularization

    Erin George, Rayan Saab

    cs.LG

    Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dataset is as small as possible. We study the performance of OPTQ and a variant algorithm, stochastic OPTQ, in a generalization setting and derive...

    arxiv.org/abs/2609.31560 · PDF

  7. 07

    Online Learning via Learned Latent Bayesian Tracking

    Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone

    cs.LG · eess.SP

    Online learning in non-stationary environments requires models to adapt rapidly from streaming data under strict computational constraints. A principled approach casts online learning as Bayesian state tracking, where model parameters are updated sequentially via Bayesian filtering. However, applying Bayesian filters directly to modern deep models is computationally prohibitive due to the high dimensionality of parameter space, forcing...

    arxiv.org/abs/2609.31559 · PDF

  8. 08

    BeatGraph: Self-Supervised Heartbeat Graphs for Infant ECG Representations from the Home Environment

    Mohammad Nur Hossain Khan, M. S. Krafczyk, Beverly G. Bolster, Nancy McElwain, Mark A. Hasegawa-Johnson, Bashima Islam

    cs.LG

    Electrocardiogram (ECG) foundation models typically tokenize the signal into fixed-length patches that ignore cardiac structure, so a patch may split a heartbeat and the number of beats in each patch shifts with heart rate. This matters most for infants, whose heart rates are higher and whose ECG differs from the adult, clinic-recorded 12-lead data these models are built on. A model for infant ECG should therefore reason about heartbeats...

    arxiv.org/abs/2609.31546 · PDF

  9. 09

    NEXT: Physics-Informed Neuro-Spectral Exponential Time Differencing Architectures

    Márcio Marques, Leonardo Mendonça, Leonardo M. Moreira, Christian Júnior de Oliveira, Vitor Balestro, Tiago Novello,...

    cs.LG · math.NA

    Physics-Informed Neural Networks (PINNs) build neural representations of time-dependent PDE solutions, naturally incorporating physics knowledge and observational data, which makes them well suited to both forward and inverse PDE problems. PINNs, however, are known to suffer from spectral bias and lack of causality. Neuro-Spectral Architectures (NeuSA), a recently proposed alternative to PINNs, mitigate both issues, but their numerical...

    arxiv.org/abs/2609.31539 · PDF

  10. 10

    HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning

    Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng, Weiqiang Zhu, Zhenhai Ji, Zhengning Wang

    cs.LG

    Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the mapping from agents and coalitions to value components as interactions or active agents evolve. We refer to this...

    arxiv.org/abs/2609.31531 · PDF

  11. 11

    Scaffold: Support Graph Theory Based Sparsification for Graph Neural Networks

    Siddhartha Shankar Das, Sai Karthik Navuluru, S M Ferdous, Ryan A. Rossi, Baris Coskunuzer, Lakshman Tamil, Edoardo...

    cs.LG

    Graph neural networks (GNNs) rely on message passing over graph edges, making their computational and memory costs strongly dependent on graph density. Graph sparsification offers a natural way to reduce these costs, but removing edges indiscriminately can distort important communication structure and degrade predictive performance. We introduce Scaffold, a topology-based, unsupervised graph sparsification framework derived from support graph...

    arxiv.org/abs/2609.31466 · PDF

  12. 12

    Uncertainty-Aware Federated Learning for Infant Movement Analysis

    Edmond S. L. Ho

    cs.LG · cs.AI · cs.CV

    Infant movement analysis provides valuable biomarkers for the early identification of neurodevelopmental disorders. Recent advances in deep learning have enabled automated analysis of infant movements from video-derived skeletal representations, achieving performance comparable to expert assessment for tasks such as General Movement Assessment (GMA). However, most existing approaches rely on centralized training, requiring data from multiple...

    arxiv.org/abs/2609.31463 · PDF

  13. 13

    Different Corruptions, Different Signals: Uncertainty and Loss in Federated Data Quality

    Bradley Scott, Zeqi Luo, Edmond S. L. Ho

    cs.LG · cs.AI · cs.CV

    Federated learning (FL) data corruption can affect either inputs or labels, but it remains unclear whether input-conditional uncertainty and prediction-label loss expose these corruption modes equally. This paper compares two corruption-detection signals in FL: input-conditional uncertainty and prediction-label loss. The uncertainty signal is characterised using a learned aleatoric variance estimate together with Monte Carlo (MC) dropout...

    arxiv.org/abs/2609.31454 · PDF

  14. 14

    Evaluating the accuracy of KV cache reuse techniques

    Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona

    cs.LG

    Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fail to faithfully capture the loss of accuracy attributable to reuse, often artificially inflating the reported effectiveness. We also show that existing datasets do not exhibit the reuse dynamics needed to...

    arxiv.org/abs/2609.31415 · PDF

  15. 15

    Decodable In-Context State and Model Output Across Training

    Manas Venkata Sai Ravulapalli, Samrath Singh Chadha

    cs.LG

    Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. Probe accuracy rises during Pythia pretraining, while probe-guided steering moves from negligible all-trial benefit to a larger benefit at two model sizes. Saved scores distinguish...

    arxiv.org/abs/2609.31401 · PDF

  16. 16

    Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition

    Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen

    cs.LG · eess.SP

    Multimodal emotion recognition (MER) increasingly pairs EEG with speech, treating internal neural signals and external vocal expression as informative views of affect. In practice, naive fusion underperforms the stronger single modality, because EEG artifacts inject noise that corrupts the shared representation. We introduce EmoSpeechBrain, a multimodal framework built on the insight that noise suppression is a precondition for effective...

    arxiv.org/abs/2609.31399 · PDF

  17. 17

    Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness

    Bin Li, Dongdong Wang, Siyang Lu

    cs.LG · cs.CR

    Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate these factors, we conduct a systematic empirical analysis across three public log datasets, examining different adaptation strategies, model architectures, parameter scales, and...

    arxiv.org/abs/2609.31371 · PDF

  18. 18

    Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning

    Alireza Abdollahpoorrostam, Ehsan Sharifian, Buse Şen, Marco Cuturi, Daniel Kuhn

    cs.LG · math.OC · stat.ML

    Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulation in which the adversary may choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. We show that the adversary's problem can be...

    arxiv.org/abs/2609.31363 · PDF

  19. 19

    Progressive Memory Transformer: Memory-Aware Attention for Time-Series

    Tord Sture Stangeland, Andreas Köhler, Steffen Mæland, Adín Ramíres Rivera

    cs.LG

    Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning...

    arxiv.org/abs/2609.31351 · PDF

  20. 20

    Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design

    Fu Feng, Ruixiao Shi, Yucheng Xie, Jing Wang, Xin Geng

    cs.LG

    Morphology--control co-design jointly optimizes an agent's body structure and control policy as an integrated embodied system. However, existing methods typically model morphology design and control with separate networks coupled only indirectly through a shared task objective, limiting explicit high-level coordination. Inspired by natural genes that coordinate biological development, we introduce \textbf{Morphogene}, a compact latent...

    arxiv.org/abs/2609.31329 · PDF

  21. 21

    More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting

    Lewei Xie, Haoyu Zhang, Jiajun Zhou, Yulong Chen, Guanxing Chen, Yu-An Huang, Hau-San Wong, Yifan Zhang, Zhi-An Huang

    cs.LG

    Continual spatio-temporal forecasting supports traffic management and environmental monitoring under evolving dynamics and expanding sensor networks. However, conventional graph-based continual learning methods tie forecasting representations to the current sensor layout, so sensor expansion can alter the representation of learned spatial relationships. Our key insight is that sensor expansion changes the evidence available about a process...

    arxiv.org/abs/2609.31325 · PDF

  22. 22

    LUCID: Learning Under Confounding for Inference and Discovery in Time Series

    Mohammad Fesanghary

    cs.LG · stat.ML

    Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direct edges. We propose LUCID (Learning Under Confounding for Inference and Discovery, a regime-adaptive deconfounding layer that first estimates the confounding regime from data using a Marčenko--Pastur spectral router, then applies a deconfounding strategy matched to that regime. When the spectrum...

    arxiv.org/abs/2609.31315 · PDF

  23. 23

    Benchmarking Attention for Tabular Foundation Models

    Maximilian Schambach, Clemens Biehl, Sam Thelin

    cs.LG

    Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden...

    arxiv.org/abs/2609.31306 · PDF

  24. 24

    Softmax Reparameterization for Output-Head Quantization

    Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng, Lan Yan, Priya Shanmugasundaram, Tracy Holloway King

    cs.LG · cs.AI

    Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. This...

    arxiv.org/abs/2609.31291 · PDF

  25. 25

    Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization

    Mahesh Reddy Pagadala

    cs.LG · cs.DC

    We report fine-grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory-constrained DNN training, measured on the reference DTR simulator (simrd) using public execution traces. On an LSTM trace, memory budgets differing by 0.10% of unconstrained peak memory select fast and slow execution regimes whose overheads differ by as much as 7.3x; the slow regime is driven by broadly repeated...

    arxiv.org/abs/2609.31250 · PDF

  26. 26

    Budgeted Quotient-Residual Guidance for Frozen Pocket-Conditioned Molecular Diffusion

    Xinyu Wang, Jinbo Bi, Minghu Song

    cs.LG

    Pocket-conditioned molecular diffusion updates ambient atom coordinates, but many lead-optimization objectives are expressed on quotient features such as distances, contacts, and anchored substructures. We introduce budgeted quotient-residual guidance (QRG), an inference-time correction that makes these quotient objectives active without retraining the molecular generator. QRG lifts quotient covectors to metric-horizontal ambient directions...

    arxiv.org/abs/2609.31222 · PDF

  27. 27

    Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra

    Matthieu Le Lain, Gaël Cessateur, Sébastien Lefèvre

    cs.LG · physics.space-ph

    Auroral spectrographs such as the Auroral Spectrograph In Skibotn (ASIS) record hundreds of thousands of emission spectra, but only a few hundred can be labelled by an expert. To exploit the rest, we pretrain a 1D Vision Transformer with a masked autoencoder on 223,000 unlabelled spectra. Without labels, its representation recovers the emission-line intensity ratios that physicists use to diagnose the precipitating particles (R^2 0.91 vs....

    arxiv.org/abs/2609.31206 · PDF

  28. 28

    ALF: An Active Learning Framework for Scientific Discovery

    Shikha Surana, Alex Hawkins-Hooker, Olivia Gallup, Christoph Brunken, Jules Tilly, Paul Duckworth

    cs.LG

    Machine learning for scientific discovery is almost systematically data bound. Producing relevant high quality data, under budget constraints, is amongst the most promising ways to advance the field. Active learning (AL) offers promise wherever labelling requires expensive experiment, measurement, or simulation. Most existing tools cover only part of the data acquisition loop, and typically focus on either offline benchmarking or online...

    arxiv.org/abs/2609.31197 · PDF

  29. 29

    Audio emotion recognition for atypical hearing

    Ulysse Roussel

    cs.LG

    My doctoral work aims to explore Audio Emotion Recognition (AER) in the context of atypical listening. This research focuses on auditory hypersensitivity in people with autism, a phenomenon that is often difficult to evaluate and unique to each individual. Our core idea is to leverage our understanding of affect from acoustic traits, relying on the possibility of generalizing affective responses from a small amount of annotated data. As a...

    arxiv.org/abs/2609.31168 · PDF

  30. 30

    WorldTS: World Modeling for Multimodal Covariate-aware Time Series Forecasting

    Yuhan Zhu, Xiangfei Qiu, Hanyin Cheng, Wangmeng Shen, Chenjuan Guo, Bin Yang, Jilin Hu, Christian S. Jensen

    cs.LG

    Time series forecasting is typically framed as learning a direct mapping from historical to future observations in the observation space. However, sequences of observations generally provide only a partial view of the dynamics of the underlying system, with future observations being shaped by latent dynamics. Recent latent-space forecasting methods thus achieve improved performance by predicting future observations from latent-space...

    arxiv.org/abs/2609.31162 · PDF

  31. 31

    I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?

    Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi

    cs.LG

    Recent empirical and theoretical advances suggest that joint-embedding predictive architectures (JEPAs) may learn meaningful representations for action-conditioned prediction of future outcomes, thus becoming one of the foundational structures for world models. However, accurate prediction does not, in general, necessarily imply recovery of underlying causal states that give rise to the observed dynamics. This work investigates when and how...

    arxiv.org/abs/2609.31161 · PDF

  32. 32

    Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection

    Jianan Liu, Chunguang Li

    cs.LG

    Multi-dimensional time series, inherently tensorial, are common in practice. Despite great progress in time series anomaly detection, most existing methods are confined to uni-/multi-variate time series. When handling multi-dimensional time series using these methods, reshaping operations are required, which inevitably break the intrinsic correlations and thus lead to performance degradation. In uni-/multi-variate time series anomaly...

    arxiv.org/abs/2609.31157 · PDF

  33. 33

    Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift

    Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong

    cs.LG · cs.AI

    Compressing a trained model yields a family of deployment candidates, and under domain shift the most compressed one need not be the one to deploy. We study selection over such a family, with candidates and teacher fixed and target labels absent or scarce. Two findings organize the label-free case. Minimum teacher distortion behaves almost as a constant rule, selecting the same eight-bit, per-channel, unclipped configuration in every run,...

    arxiv.org/abs/2609.31155 · PDF

  34. 34

    CRNDiff: Count-Native Diffusion Framework via Chemical Reaction Networks

    Yuxuan Qiu, Praful Gagrani, Tetsuya J Kobayashi

    cs.LG

    Scientific measurements such as single-cell RNA (scRNA) sequencing often take the form of nonnegative integer counts, whereas continuous-state diffusion models approximate this discrete structure using continuous coordinates. Building on stochastic chemical reaction networks (CRNs), a class of count-native Markov jump processes, we introduce CRNDiff, a structured framework that combines count-space diffusion with inference-time conditioning...

    arxiv.org/abs/2609.31149 · PDF

  35. 35

    Frame the adversary: a structure-aware attack methodology

    Vicky Kouni, Stelios Perrakis, Francis Bach, Pascal Frossard, Yann Chevaleyre

    cs.LG

    Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical and security-sensitive applications. Yet, existing approaches are typically not derived as solutions to an optimization problem that explicitly...

    arxiv.org/abs/2609.31128 · PDF

  36. 36

    From Shortcut Learning to Discrete Neural Insertion Sort

    Konstantinos Mylonas, Thrasyvoulos Spyropoulos

    cs.LG · cs.AI

    Neural algorithmic reasoning aims to train neural networks to follow known algorithms and generalize beyond the input sizes seen during training. However, correct final outputs and intermediate supervision do not necessarily show that a model follows the intended execution. We study this problem using insertion sort. Our analysis of the CLRS30 baseline NAR shows that the hint objective is weakly optimized and that hint accuracy remains low....

    arxiv.org/abs/2609.31114 · PDF

  37. 37

    Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods

    Saksham Kiroriwal, Julius Pfrommer, Jürgen Beyerer

    cs.LG · cs.AI · cs.IT

    We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and...

    arxiv.org/abs/2609.31107 · PDF

  38. 38

    The Residual Stream's Effective Depth

    Barak Gahtan, Ido Galil, Alex M. Bronstein

    cs.LG

    We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only language models, $\Deff$ separates a structural consequence of residual accumulation from an empirical one: even maximally diverse orthogonal updates...

    arxiv.org/abs/2609.31098 · PDF

  39. 39

    Block Sparse Attention with Log-Linear Complexity

    Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu

    cs.LG

    Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query-block pairs and therefore remains quadratic in sequence length. To address this issue, we propose PISA, a block-sparse attention mechanism that employs a pyramid Top-$K$ selection...

    arxiv.org/abs/2609.31093 · PDF

  40. 40

    SAGE: A sampling-aware global evaluation benchmark for species distribution modeling

    Emilia Arens, Nina van Tiel, Robin Zbinden, Damien Robert, Lukas Drees, Chiara Vanalli, Benjamin Kellenberger,...

    cs.LG · q-bio.PE · stat.ML

    Knowing where species occur is fundamental for biodiversity research and conservation. Species distribution models (SDMs) link species observations to environmental conditions to estimate their spatial distribution. However, accuracy varies with the underlying data and models, making it essential to know for which species models can be trusted. Deep-learning-based SDMs ("DeepSDMs") now jointly model thousands of species, drawing on hundreds...

    arxiv.org/abs/2609.31082 · PDF

  41. 41

    Distributed Learning as a Service: The Developer's Perspective

    Tianyue Chu, Filippo Vannella, Dimitra Tsigkari, Paula Delgado-Santos, Fernando López, Pablo Gomez Guerrero,...

    cs.LG · cs.DC · cs.NI

    Application developers of distributed learning services face challenges that a typical federated learning loop does not address. Specifically, the model updates can still leak private data, devices might not be able to participate in the training due to limited resources, a single aggregator might not be able to scale, and the transmissions of model weights induce a considerable bandwidth cost. This paper demonstrates DLaaS (Distributed...

    arxiv.org/abs/2609.31061 · PDF

  42. 42

    Aurora-X: Built for Extreme Time Series Forecasting

    Xingjian Wu, Chenjuan Guo, Xiangfei Qiu, Zhigang Hu, Hanyin Cheng, Peng Chen, Yang Shu, Jilin Hu, Bin Yang

    cs.LG

    Time series foundation models (TSFMs) enable cross-domain forecasting, but their development as general-purpose forecasters remains constrained by underexplored training potential and limited architectural versatility. To address these challenges, we introduce Aurora-X, a billion-scale TSFM with a progressive curriculum and a unified architecture. We first use channel-independent pretraining to learn temporal patterns, then introduce...

    arxiv.org/abs/2609.31038 · PDF

  43. 43

    Robust Graph Clustering Network for Multiple Missing Data

    Keyuan Qiu, Renda Han, Zhen Tang, Qiang He, Xingwei Wang, Wenxin Zhang, Guangzhen Yao, Junxin Chen, Qingjian Ni

    cs.LG

    Clustering on graphs where both node attributes and structural links are partially missing remains a challenging task. Existing methods typically rely on imputation-then-clustering on single-view missingness incomplete graphs, which are vulnerable to cross-view error propagation and cluster-boundary blurring under simultaneous attribute and structure missingness. To address these limitations, we propose a Robust Graph Clustering Network for...

    arxiv.org/abs/2609.31033 · PDF

  44. 44

    Metacognitive Selective Ensemble for Mobile Systems

    Sungmin Lee, Kichang Lee, Joonhee Lee, JaeYeon Park, Songkuk Kim, JeongGil Ko

    cs.LG

    Deep ensembles improve robustness in mobile sensing, but repeatedly executing many models over continuous sensor streams is costly. Selecting only a few members reduces this cost, yet adaptive selection often requires additional model execution to obtain reliable evidence about inactive candidates. We present MetaSE, an active ensemble framework that exploits short-term persistence in per-model reliability. MetaSE maintains a small active set...

    arxiv.org/abs/2609.31031 · PDF

  45. 45

    Robust Successor Features

    Erik Nikulski, Yamen Habib, Vicenç Gomez, Anders Jonsson, Rubén Moreno-Bote, Javier Segovia-Aguas

    cs.LG

    Generalization in Reinforcement Learning (RL) refers to the ability to execute close-to-optimal policies in unseen tasks after the agent has been trained on a different set of tasks. Building on the seminal work of the successor representation and further adaptations with function approximation, Transfer in RL has traditionally focused on generalizing to tasks that only differ in the reward function. A decade after the introduction of the...

    arxiv.org/abs/2609.31016 · PDF

  46. 46

    The Linear Representation Hypothesis for Vision-Language-Action Models

    Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han

    cs.LG · cs.AI

    The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vision-language-action (VLA) models, but the dynamical nature of embodied interaction introduces an additional challenge. Unlike semantic attributes commonly studied in LLMs, such as gender...

    arxiv.org/abs/2609.30996 · PDF

  47. 47

    Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models

    Shan Zhao, Ilija Trajkovic, Julia Kaltenborn, Yaniv Gurwicz, Peer Nowack, David Rolnick, Julien Boussard

    cs.LG

    Machine learning (ML) emulators provide a fast and cost-effective method to simulate climate change scenarios after being trained on Earth System Models projections. However, the black-box nature of those data-driven approaches limit the usability and trustworthiness of their outputs and in particular their use as causal attribution tools. Here, we develop a hierarchical causal representation learning framework applied to sea surface...

    arxiv.org/abs/2609.30995 · PDF

  48. 48

    Does Uniform Discrete Diffusion Need Time?

    Chunsan Hong, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Jong Chul Ye, Yuki Mitsufuji

    cs.LG · cs.AI · cs.CL

    Uniform discrete diffusion models (UDMs) commonly use explicit time conditioning, but we find that it can often be unnecessary in practice. In this paper, we first show that the population-optimal UDM predictor generally depends on time: time controls how much the model should trust the observed context. We then show that this dependence can become negligible in finite-data settings relevant to language. When a corrupted training sequence...

    arxiv.org/abs/2609.30977 · PDF

  49. 49

    LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers

    Natsuki Yoshino, Ren Uchida, Kazuki Matsumoto, Kohei Yatabe

    cs.LG · eess.SY

    Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforcing Lipschitz continuity requires each layer of a DNN to be Lipschitz continuous, thereby guaranteeing overall Lipschitz continuity. However, this layer-wise approach typically imposes...

    arxiv.org/abs/2609.30973 · PDF

  50. 50

    Gradient Surgery for Physics-Informed Neural Networks

    Thomas Borsani, Giuseppe Di Fatta

    cs.LG

    Physics-Informed Neural Networks (PINNs) are trained by optimising a composite objective that combines data fitting with physics-based constraints, typically resulting in a highly imbalanced multi-task optimisation problem. Under these conditions, existing optimisation strategies are affected by conflicting task gradients, leading to slow convergence and unstable training, particularly for stiff and high-frequency partial differential...

    arxiv.org/abs/2609.30966 · PDF

  51. 51

    Towards Understanding Momentum Acceleration in River-Valley Loss Landscape

    Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen, Tianhao Wang, Zhiyuan Li

    cs.LG · math.OC

    The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features a low-loss manifold (river) flanked by sharp orthogonal directions with higher loss (mountains). In the long term, the optimization progress is...

    arxiv.org/abs/2609.30957 · PDF

  52. 52

    Low-Bit Recurrent States in Hybrid Language Models

    Hongren Chen, Jiayang He

    cs.LG

    Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-precision bit allocation, without calibration data, rotation, or training. We also quantize decay rates logarithmically. With per-token state...

    arxiv.org/abs/2609.30950 · PDF

  53. 53

    PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem

    Mateo Toro Diz, Jonathan Hoss, Noah Klarmann

    cs.LG · cs.AI

    The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining with offline fine-tuning on production-specific data. Reinforcement learning through online interaction enables exploration of general scheduling strategies, but typically relies on...

    arxiv.org/abs/2609.30948 · PDF

  54. 54

    EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting

    Takumi Fujimoto, Hiroaki Nishi

    cs.LG

    Completed multi-horizon forecasts provide residual feedback for a fixed forecaster, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order discrete cosine transform (DCT) coefficients and the final value of the preceding residual block. Within each channel, the endpoint is shared across component-wise online ridge regressions...

    arxiv.org/abs/2609.30929 · PDF

  55. 55

    Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations

    Marcin Kostrzewa, Maciej Zięba

    cs.LG · cs.AI

    Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We...

    arxiv.org/abs/2609.30918 · PDF

  56. 56

    CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters

    Yuhang Cao, Yanzhou Mu, Chunrong Fang, Zhenyu Chen

    cs.LG

    Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current model outputs, while complete affected suffix recomputation restores fidelity at substantial cost. We seek minimal recomputation that recovers current adapter behavior. Existing systems track token, context, or...

    arxiv.org/abs/2609.30884 · PDF

  57. 57

    Learning Chance-Constrained MDPs with Bellman Distributional Certificates

    Chenbei Lu, Hongyu Yi

    cs.LG

    Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this...

    arxiv.org/abs/2609.30856 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.