cs.LG · 2026-09-30 · No. 129

Machine Learning, 2026-09-30.

69 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

69 entries
  1. 01

    Breakdown of Local Denoising as Semantic Speciation

    Guangkuo Liu, Mert Okyay, Yifan F. Zhang, Fangjun Hu, Rahul Nandkishore, Xun Gao

    cs.LG · cond-mat.dis-nn · cond-mat.stat-mech

    The dynamics of generative models exhibit two apparently distinct temporal windows: a speciation window, in which a sample commits to a semantic class, and a nonlocality window, in which local context windows become insufficient for generation. Motivated by evidence of their near-concurrence in a variety of frontier models, we investigate their relationship through the spatial distribution of semantic information. Under a "common cause"...

    arxiv.org/abs/2609.38176 · PDF

  2. 02

    LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

    Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci,...

    cs.LG · cs.AI

    Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can...

    arxiv.org/abs/2609.38166 · PDF

  3. 03

    A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization

    Jianru Shen

    cs.LG · cs.DM

    Evaluations of graph reconstruction by language models typically report a single aggregate distance between the original and the reconstructed graph. We prove that for the Wasserstein distance between Laplacian spectra such a summary is bracketed by two edge counts, the net change in edge number from below and the symmetric difference from above, each scaled by $2/n$ where $n$ is the number of vertices. The bracket is sharp: its two ends...

    arxiv.org/abs/2609.38161 · PDF

  4. 04

    Multi-Agent Flow Matching with Decoupled Generative Guidance

    Ruoyu Lin, Magnus Egerstedt, Fabio Pasqualetti

    cs.LG · cs.MA · cs.RO · math.OC

    Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general, come with formal guarantees that the generated objects satisfy hard constraints or requirements. In multi-agent generation, this problem becomes more challenging because a hard requirement can depend on multiple agents, while each agent may need to determine its own guidance input without...

    arxiv.org/abs/2609.38133 · PDF

  5. 05

    Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs

    Yige Hong, Xiangcheng Zhang, Qiaomin Xie, Yudong Chen, Weina Wang

    cs.LG · math.OC · math.PR

    We study average-reward weakly-coupled Markov decision processes (WCMDPs), where a WCMDP consists of $N$ smaller MDPs, called arms, that share multiple per-step budget constraints. We consider the setting where the arms have identical model parameters, multiple actions, and state- and action-dependent costs. For restless bandits (RBs), a well-studied special case of WCMDPs, prior work has developed policies that achieve an $O(1/\sqrt{N})$...

    arxiv.org/abs/2609.38132 · PDF

  6. 06

    WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

    Jiale Chen, Vage Egiazarian, Eldar Kurtić, Torsten Hoefler, Dan Alistarh

    cs.LG

    KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with...

    arxiv.org/abs/2609.38121 · PDF

  7. 07

    Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling

    Panagiotis Theodoropoulos, Nan Jiang, Xintong Duan, Ali Hasan, Yuriy Nevmyvaka, Evangelos A. Theodorou, Wei Deng

    cs.LG

    Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged generalization of RL. However, this approach faces a fundamental exploration--exploitation trade-off, as % strong sharpening...

    arxiv.org/abs/2609.38104 · PDF

  8. 08

    Tail-Influence Sampling for CVaR Policy Evaluation

    Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou-Ammar

    cs.LG

    Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to estimate a fixed policy's CVaR most accurately. We derive a tail influence for each queryable conditional law that aggregates...

    arxiv.org/abs/2609.38096 · PDF

  9. 09

    Probe-Space Preconditioning for Fast and Stable Zero-Order Training

    Francois Chaubard, Mykel J. Kochenderfer, Chris Ré

    cs.LG

    Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU memory (assuming batch size 8 and sequence length 2048). Alternatively, zero-order optimization (ZOO) trains in inference-mode (requiring only $\approx$ 60GB for the same model): no stored activations, no gradients, and no optimizer states. However, ZOO convergence has lagged behind BP. In this...

    arxiv.org/abs/2609.38095 · PDF

  10. 10

    Dimensionally consistent surrogate modelling through dimensional analysis and harmonic expansions

    Ernest Tarrus, Hector Gisbert

    cs.LG · hep-ph

    Dimensional homogeneity is a fundamental constraint on physically meaningful models, requiring invariance under changes of units. We present a data-driven method for constructing surrogate models that satisfy this constraint at the level of the hypothesis class. Starting from a dimension matrix of measured variables, the method derives Buckingham $Π$-groups, constructs admissible dimensional prefactors, and approximates the remaining...

    arxiv.org/abs/2609.38094 · PDF

  11. 11

    Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging

    Sanjali Yadav, Bahar Asgari

    cs.LG

    Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-constrained, single-GPU systems. However, this benefit is difficult to realize because expert parameters dominate memory, and token-level routing is dynamic, unpredictable, and skewed. Prior work using offloading and caching remains fundamentally reactive, as systems wait for router outputs...

    arxiv.org/abs/2609.38090 · PDF

  12. 12

    Traversing the solution space of neural networks with Hessian Null Space Continuation

    Ann Huang, Mitchell Ostrow, Zhouyang Lu, William T. Redman, Leo Kozachkov, Kanaka Rajan

    cs.LG · q-bio.NC · stat.ML

    On a single task, deep networks can learn many solutions, depending on their optimizer, training data, architecture, and hyperparameters. Many of these solutions are mode-connected: rather than isolated points in weight space, they are connected by low-loss regions. Yet how their internal computation varies within these regions is unknown. A parallel line of work has identified the degeneracy of neural representations: many networks reach...

    arxiv.org/abs/2609.38081 · PDF

  13. 13

    A foundation model for energy and radiation systems built on heterogeneous scientific interfaces

    Samrendra Roy, Tapas Tripura, Yoon Pyo Lee, Souvik Chakraborty, Syed Bahauddin Alam

    cs.LG · physics.comp-ph

    Scientific foundation models are commonly evaluated after heterogeneous physical problems have already been translated into a compatible gridded, tokenized or symbolic representation. This leaves the scientific interface outside both the pretrained model and the audit of what is actually reused. We study the complementary setting in which boundary histories, sparse monitor records and loading histories retain their native inference classes...

    arxiv.org/abs/2609.38067 · PDF

  14. 14

    Alpha Diffusion Language Models: Factorization Alone Is Not the Problem

    Nikita Gushchin, Dmitry Baranchuk, Alexander Korotin

    cs.LG

    Discrete diffusion language models can generate multiple tokens in parallel, but reducing the number of denoising steps can lead to inconsistent predictions. Standard cross-entropy training fits conditional token marginals, whereas parallel generation requires consistent joint predictions. We introduce Alpha Diffusion Language Models (AlphaDLM), trained with a sequence-level alpha loss that recovers cross-entropy in the limit of vanishing...

    arxiv.org/abs/2609.38066 · PDF

  15. 15

    Jaxolotl: A Unified High-Performance Benchmark Suite for LTL-Based Multi-Task RL

    Mathias Jackermeier, Jacques Cloete, Alessandro Abate

    cs.LG · cs.AI

    Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully adopted for training generalist multi-task policies. However, differences in implementations, task distributions, and evaluation protocols make existing methods difficult to compare, while high...

    arxiv.org/abs/2609.38065 · PDF

  16. 16

    Improving Function Space Flow Matching with Kernel Optimal Transport

    Fred Xu, Thomas Markovich, Barbora Barancikova, Yizhou Sun

    cs.LG

    Generative models for function-valued data, such as time series and solutions of partial differential equations, must learn distributions over infinite-dimensional spaces. Functional Flow Matching (FFM) extends Flow Matching to this setting, learning a velocity field whose flow transports a Gaussian prior to the data distribution, but it inherits the independent endpoint pairing of standard Flow Matching: in each batch, prior and data samples...

    arxiv.org/abs/2609.38049 · PDF

  17. 17

    Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models

    Zhenyu Wang, Tianze Wang, Linjun Zhang, Yifan Hu

    cs.LG · cs.AI · cs.CL · math.OC

    On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect...

    arxiv.org/abs/2609.38025 · PDF

  18. 18

    Prompts Live on an Arc: Gaussian Curricula in Fisher--Rao Coordinates for Rollout-Efficient GRPO

    Mei Okonkwo, Pixel Nomand, Julian Berg, Elena Voss, Lena Park, Marcus Hale, Adrian Cho, Sofia Reyes

    cs.LG

    Group relative policy optimization (GRPO) learns only from prompts whose sampled responses disagree: a group that is entirely correct or entirely incorrect has zero reward variance, contributes no gradient, and still consumes its rollouts. Prompt-selection methods reduce this waste by steering sampling toward intermediate pass rates, but they choose the target, its width, and the uncertainty model heuristically, in raw pass-rate or logit...

    arxiv.org/abs/2609.38018 · PDF

  19. 19

    When do data mixtures improve scaling laws? Insights from high-dimensional regression

    Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli

    cs.LG · stat.ML

    Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downstream performance. Despite an extensive literature on data mixing and reweighting, existing work is largely empirical and it remains unclear when auxiliary data genuinely improves scaling laws rather than merely providing more samples. To gain insight into this question, we study a...

    arxiv.org/abs/2609.38011 · PDF

  20. 20

    No Scale Left Behind: Multi-Scale Autoencoder with Bi-directional Attention for Time Series Anomaly Detection

    Jiaheng Guo, Haochen Zhang, Yu-Chao Huang, Jinhao Duan, Nicholas Konz, Tianlong Chen

    cs.LG · cs.AI

    Time series anomaly detection (TSAD) plays a crucial role in healthcare, finance, industrial monitoring, and other sectors. Within and between these settings, anomalies span vastly different temporal scales, from sub-second point spikes to multi-hour drift patterns. However, most existing TSAD methods commit to a single temporal granularity, and multi-scale designs either analyze different scales in isolation or are constrained to a...

    arxiv.org/abs/2609.38004 · PDF

  21. 21

    TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models

    Deqing Fu, Huangyuan Su, Rajat Sen, Taman Narayan, Sujay Sanghavi, Abhimanyu Das, Weihao Kong

    cs.LG

    Tabular foundation models achieve strong zero-shot accuracy on structured data by pretraining on synthetic tables, but they ignore the column names, task descriptions, and auxiliary files that carry dataset semantics. Meanwhile, self-evolving machine learning engineering (MLE) agents train models from scratch on each dataset, yet jointly searching over features, architectures, and hyperparameters is noisy and prone to overfitting. We...

    arxiv.org/abs/2609.37989 · PDF

  22. 22

    $S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient

    Hongbo Ma, Sansheng Cao, Jiajun Fan, Bangji Yang, Ge Liu

    cs.LG · cs.AI · cs.CL · cs.CV

    LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during...

    arxiv.org/abs/2609.37976 · PDF

  23. 23

    On Trajectory-Aware Training for Masked Diffusion Language Models

    Manuel Madeira, Amitis Shidani, Alice Bizeul, Victor Turrisi, Louis Béthune, Bhavika Devnani, Dan Busbridge, Pierre...

    cs.LG · cs.AI · cs.CL

    Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model's own predictions. Additionally, each step has no access to what the previous one computed. Recent methods narrow these limitations from separate angles, leaving open how these choices...

    arxiv.org/abs/2609.37974 · PDF

  24. 24

    TabFM: A Zero-Shot Foundation Model for Tabular Data

    Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

    cs.LG

    Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. Trained entirely on synthetic tables generated from structural...

    arxiv.org/abs/2609.37959 · PDF

  25. 25

    Kolmogorov-Arnold Classifier Systems as Universal Approximators

    Hiroki Shiraishi, Hisao Ishibuchi, Masaya Nakata

    cs.LG · cs.NE

    As the input dimension $n$ grows, rule-based machine learning, such as Learning Classifier Systems (LCSs), faces a fundamental scalability bottleneck for function approximation: both rule count and parameter count grow exponentially with $n$. Traditional LCSs partition the $n$-dimensional input space directly, requiring $\mathcal{O}(m^n)$ rules for adequate coverage, where $m$ is the per-variable resolution. This article breaks from this...

    arxiv.org/abs/2609.37958 · PDF

  26. 26

    Scene-Consistent Illumination Transfer for Inserted Advertising Graphics

    Rameshwar Mishra, Bishshoy Das, A. V. Subramanyam, Guan-Ming Su

    cs.LG

    Replacing a visible advertisement in a broadcast frame is geometrically straightforward but photometrically delicate. A pasted graphic can have the correct perspective and still appear detached when its brightness, shading, or shadow disagrees with the surface beneath it. This paper presents Ad-Relight, an inference-only procedure for transferring scene illumination to a supplied advertising graphic without collecting a banner-specific...

    arxiv.org/abs/2609.37951 · PDF

  27. 27

    An Efficient Machine Learning Approach for Degradation Forecasting in AEM Water Electrolysis

    Marco Veneriano, Ani Gjergji, Sebastiano Bellani, Andrea Riva, Vito Paolo Pastore, Matteo Santacesaria

    cs.LG · eess.SY · physics.app-ph

    This study provides a data-driven analysis of a novel dataset of single-cell Anion Exchange Membrane water electrolyzers (AEMWE), operated under constant current load across multiple heterogeneous experimental campaigns. We train and evaluate a range of machine learning models with different complexity, including linear baselines, LSTMs and CNNs, to perform medium-term forecasting of the cell voltage degradation curve. The models are assessed...

    arxiv.org/abs/2609.37941 · PDF

  28. 28

    Learning When to Update: A Near-Optimal Timing Bandit Approach

    Qiulin Lin, Junyan Su, Liyuan Wang, Minghua Chen

    cs.LG

    Systems operating in dynamic environments require timely updates to sustain performance. For resource-intensive systems such as machine learning models and digital twins, strategically timing updates is essential. Updating too frequently wastes resources, while updating too infrequently leads to costly performance degradation. The problem is particularly challenging when the system's degradation pattern is unknown a priori, as is common in...

    arxiv.org/abs/2609.37932 · PDF

  29. 29

    Time-Anchored Diffusion Language Models: Latent-Space Caching for Fast Generation

    Joel Anto Paul, Litu Rout, Aditya Akella, Sanjay Shakkottai

    cs.LG · cs.CL

    Recent work on anchored diffusion language models improves denoising by shaping an intermediate latent space with supervised important-token targets. In this work, we introduce time-based (self-supervised) anchoring, which learns and reuses latent anchors without requiring such targets. Our key observation is that anchors encode persistent properties of the clean sequence, such as its semantic intent, global structure, or intermediate plan....

    arxiv.org/abs/2609.37924 · PDF

  30. 30

    Pattern Formation in Transformers

    Erkan Turan, Gaspard Abel, Maks Ovsjanikov

    cs.LG

    What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Transformers escape from rank collapse or demonstrates that self-attention drives tokens toward cluster patterns. The latter view arises from an elegant dynamical systems perspective, but relies on simplified architectural assumptions, and does not explain the rich structures observed in...

    arxiv.org/abs/2609.37921 · PDF

  31. 31

    Search Dimension in Unlabeled Projection Pursuit: A Scaling Law for Subspace Restriction

    Rares Dimitrie Grozavescu, Mark Girolami

    cs.LG · math.ST · stat.ML

    Projection pursuit searches for a direction along which the data look least Gaussian. When the observation space contains a large Gaussian complement, the empirical objective can be minimized by a direction that carries no signal, with empirical kurtosis as low as at the truth. Sample splitting exposes rather than repairs this failure. Appending coordinates independent of the latent regime degrades the search while leaving Bayes...

    arxiv.org/abs/2609.37917 · PDF

  32. 32

    Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

    Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni

    cs.LG · cs.CL

    On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factorial analysis and find that scaffold correctness has a stronger effect on downstream accuracy than context correctness....

    arxiv.org/abs/2609.37915 · PDF

  33. 33

    Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction

    Shivang Chopra, Fotis Iliopoulos, Zsolt Kira, Gaurav Menghani

    cs.LG · cs.AI

    Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come from increasing the interaction capacity of a single predictor through deeper cross networks and more expressive cross operators. We revisit whether continually increasing interaction...

    arxiv.org/abs/2609.37905 · PDF

  34. 34

    Scaling Zero-Order Pretraining through Model Sharding

    Francois Chaubard, Mykel J. Kochenderfer, Chris Ré

    cs.LG

    Zero-order optimization (ZO) trains without backpropagation, making it relevant to forward-only hardware and non-differentiable loss, but its gradient variance grows with perturbed dimension, inhibiting large-model training. Sharded Optimization Mixture of Assemblies (SOMA) trains LSTM experts independently on $N$ data clusters using simultaneous perturbation stochastic approximation (SPSA), without exchanging gradients, activations or...

    arxiv.org/abs/2609.37899 · PDF

  35. 35

    Behavioral Capacity Certificates for Quantized Language Models

    Arian Eamaz, Mojtaba Soltanalian

    cs.LG

    Activation and key-value cache precision change what a quantized language model computes without altering its stored weights. Direct weight-code bounds, however, assign identical complexity to deployments that behave differently and charge separately for weight codes that behave identically. Behavioral Capacity Certificates (BCC) charge for behavior using the aggregate prior mass of complete implementations---weights, scales, activation and...

    arxiv.org/abs/2609.37887 · PDF

  36. 36

    TopoEmbedX: A General Framework for Representation Learning on Topological Domains

    Florian Frantzen, Ibrahem AlJabea, Ines Henriques-Cadby, Theodore Papamarkou, Mustafa Hajij, Michael T. Schaub

    cs.LG

    Topological structures such as simplicial complexes, hypergraphs, and cell complexes extend standard graph models by modeling higher-order relationships. These structures appear in many modern datasets and require specialized methods for generating meaningful embeddings. In this paper, we introduce TopoEmbedX, a unified framework for embedding a wide range of topological domains into Euclidean spaces. The package brings together several...

    arxiv.org/abs/2609.37884 · PDF

  37. 37

    Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

    Doohyuk Jang, Yoonsik Park, Gyouk Chu, Sihwan Park, Eunho Yang

    cs.LG · cs.AI · cs.CL

    Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models...

    arxiv.org/abs/2609.37868 · PDF

  38. 38

    Strict-Saddle Landscapes and Multi-Rank Geometry in Low-Tubal-Rank Tensor Sensing

    Eugene Agyei-Kodie, Longxiu Huang, Shuang Li, Xiao Liang

    cs.LG · math.OC

    We study the optimization landscape of low-tubal-rank tensor sensing through a balanced factorization. Under a tubal restricted isometry condition, we establish a quantitative strict-saddle landscape with no spurious local minima for arbitrary Fourier multi-rank profiles. We further show that the local geometry depends on the Fourier-slice ranks rather than the tubal rank alone. Uniform ranks yield quadratic growth transverse to the solution...

    arxiv.org/abs/2609.37865 · PDF

  39. 39

    Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

    Tianhao Qian, Ziming Hong, Chongyang Gao, Kezhen Chen, Lixu Wang

    cs.LG · cs.CL

    Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve...

    arxiv.org/abs/2609.37858 · PDF

  40. 40

    Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs

    Haozhan Tang, Hao Kang, Han Cai, Song Han, Chenyan Xiong

    cs.LG

    Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases and downstream degradation at 1.67B and 5.29B. QK normalization, NoPE (no positional encoding), and lower-learning-rate context...

    arxiv.org/abs/2609.37852 · PDF

  41. 41

    Scaling Influence Functions in LLMs through Eigenbasis-Corrected One-Bit Gradient Projection

    Jaeseung Heo, J Rosser, Dongwoo Kim

    cs.LG · cs.AI

    Influence functions estimate how individual training examples affect the behavior of large language models (LLMs). Analyzing how training data influence different behaviors of an LLM involves repeated influence computation. Reusing stored training gradients reduces the computational cost, but storing full gradients is prohibitively expensive at LLM scale. We study how to compress these gradients while preserving influence estimates for future...

    arxiv.org/abs/2609.37842 · PDF

  42. 42

    Counterfactual Probing for Parallel Unmasking with Hidden Forest Structure

    Ryotaro Kawata, Satoshi Hayakawa, Taiji Suzuki

    cs.LG · stat.ML

    Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies are unknown, finding safe batches also costs model evaluations. We study whether total evaluations, including discovery, can be sublinear in sequence length $N$; sublinear sequential depth then follows. We consider discrete distributions with hidden forest structure, accessed through a fixed...

    arxiv.org/abs/2609.37841 · PDF

  43. 43

    Behavioral Convergence Without Representational Convergence: Persistent Training-History Dependence in Neural Networks

    Ertuğrul Mutlu

    cs.LG

    Neural networks trained toward the same final objective can reach similar predictive performance while retaining internal representations shaped by earlier training history. We study this effect using controlled sequential-training experiments in which paired convolutional networks start from identical weights, experience reversed task orders, and then receive the same deterministic common-relaxation distribution. Across 20 paired MNIST runs,...

    arxiv.org/abs/2609.37836 · PDF

  44. 44

    Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR

    Kun Liang, Chenming Tang, Clive Bai, Weijie Liu, Zeyuan Liu, Qingyang Zhang, Saiyong Yang, Yunfang Wu

    cs.LG · cs.AI

    Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges on reliable value estimation, a difficult task requiring the critic to both assess progress toward a correct solution and...

    arxiv.org/abs/2609.37825 · PDF

  45. 45

    Feedback-Calibrated Protein Optimization with Batch-Aligned Tail Arbitration

    Zefeng Lin, Xianyong Fang, Tianfan Fu, Xiaohua Xu

    cs.LG

    Protein optimization aims to discover high-fitness sequences under a limited experimental budget. Existing machine-learning methods use task-specific predictors, biological priors, or ranking-aware objectives to guide which variants are tested in the next experimental round. However, these methods cannot adapt to shifts in the reliability of predictive evidence as measurements accumulate and ensure the correct ranking of key high-fitness...

    arxiv.org/abs/2609.37808 · PDF

  46. 46

    Challenges and Solutions for Bandits in the Wild: Warm-Started Mixture Bandits for Cross-Cohort Slate Recommendation

    Serafima Lebedeva, Sumantrak Mukherjee, Ali Arshad Sadal, Ilias Ekşi, Rahul Sharma, Julia Mueller, Theresa...

    cs.LG · cs.AI

    Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two challenges: learning user preferences quickly from limited feedback and sustaining useful recommendations when each user has a finite catalog that can become repetitive or depleted over time. We propose CohortMix-TS, a warm-started mixture bandit that learns latent user groups from earlier cohorts...

    arxiv.org/abs/2609.37800 · PDF

  47. 47

    Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance

    Fabian A. Mikulasch, Friedemann Zenke

    cs.LG · cs.AI

    Self-supervised learning (SSL) by predicting in latent space, without generating the input data itself, learns highly abstract, useful representations. Intuitively, this success is often attributed to its ability to discard nuisance information that is irrelevant to prediction. However, this poses a conundrum: both stochastic variation in a prediction-relevant latent signal and true nuisance make observations partly unpredictable; how could...

    arxiv.org/abs/2609.37789 · PDF

  48. 48

    Optimizer-dependent training dynamics converge to the same one-third optimal data scaling

    Hyunseok Lee, Mihir Basil, Yizhou Liu, Jeff Gore

    cs.LG

    Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a $1/3$ exponent emerges from learning peaked distributions. That account describes SGD, but models in practice are trained with adaptive optimizers. Here we separate two exponents the $1/3$ account does not distinguish: how fast the loss falls with training steps along a single run, and how fast the optimally...

    arxiv.org/abs/2609.37745 · PDF

  49. 49

    HyDI: A hybrid Deep Learning-Inductive Logic Programming ensemble for multi-label classification

    Simon Flügel, Till Mossakowski

    cs.LG · cs.LO

    While attaining remarkable results for many applications, Deep Learning models are notoriously difficult to explain. This work introduces HyDI, a hybrid ensemble architecture for hierarchical multi-label classification. It combines a Deep Learning (DL) model with rule-based classifiers generated by Inductive Logic Programming (ILP). For leaf classes of the label hierarchy, the rule-based classifiers replace the DL model, leading to more...

    arxiv.org/abs/2609.37740 · PDF

  50. 50

    Weights Read and Write Features: Scalable Parameter Decomposition Grounded in Activation Space

    Tue M. Cao, Lisiane Pruinelli, My T. Thai

    cs.LG

    Activation space and parameter space provide complementary views of model computation. Activations represent information, while weights read, transform, and write that information. Yet existing interpretability methods largely study the two spaces separately, leaving the connection between represented information and parameter-level computation underexplored. We introduce Activation-Supported Parameter Decomposition (ASPD), which jointly...

    arxiv.org/abs/2609.37731 · PDF

  51. 51

    Predictive Geometry of Hidden Trajectories in Transformers

    Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu State

    cs.LG · cs.CL

    Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state through the fixed downstream computation. We formalize this constraint by studying layerwise loss-to-go functions: the terminal loss obtained by continuing a candidate hidden state through the remaining transformer blocks. Around successful validation trajectories, we show that the local...

    arxiv.org/abs/2609.37717 · PDF

  52. 52

    Volatility-Clustering Adaptation for Financial Time Series

    Manh Nguyen, Minh Hoang Nguyen, Huu Hiep Nguyen, Van Dai Do, Hung Le

    cs.LG

    Time-series foundation models are increasingly adapted to new domains through fine-tuning on target data, under the implicit assumption that more target data yields better forecasts. We show that this assumption can fail in financial forecasting, where individual price changes are difficult to predict, but large moves tend to cluster, creating alternating calm and turbulent periods. Using financial foundation models trained on price bars of...

    arxiv.org/abs/2609.37715 · PDF

  53. 53

    Width Expansion as a Method for Class Incremental Learning

    A. L. S. Conde, Y. Elkhatib, C. M. Ranieri

    cs.LG · cs.AI

    Class Incremental Learning (Class-IL) requires models to learn new classes over time while preserving previously acquired knowledge without access to past data or task identity. This setting intensifies the stability-plasticity dilemma and makes catastrophic forgetting a central challenge. Existing approaches include regularization, knowledge distillation, replay, and architectural expansion. However, many expansion methods rely on explicit...

    arxiv.org/abs/2609.37702 · PDF

  54. 54

    GARDiff: Graph-Aligned Residual Diffusion for Probabilistic Multivariate Time-Series Forecasting

    Rui Han, Min Yang, Xu Zhang, Xinghao Yang, Wei Liu, Yongshun Gong

    cs.LG · cs.AI

    Diffusion models have recently shown strong potential for probabilistic multivariate time-series forecasting by modeling complex conditional distributions. Recent decoupled diffusion frameworks further separate forecasting into deterministic prediction and stochastic residual generation, making it natural to derive dependency graphs from deterministic representations and use them to guide residual diffusion. However, we show that this direct...

    arxiv.org/abs/2609.37694 · PDF

  55. 55

    When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task

    Sai Sumedh R. Hindupur, Hadas Orgad, Thomas Fel, Demba Ba

    cs.LG · cs.AI · cs.CL

    One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literature (e.g. numbers encoded on helices, days of the week on a circle, ...), with structure believed to reflect properties of data...

    arxiv.org/abs/2609.37680 · PDF

  56. 56

    LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling

    Biswajit Banerjee, Claudia Alvarez Carreno, Anton S. Petrov

    cs.LG · q-bio.BM

    Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and structure-based tasks, yet the potential of tokenization remains underexploited. Unlike human language, proteins preserve structure despite extensive sequence variation a property standard tokenization strategies fundamentally fail to capture. We introduce ZEST (Zoned Encoding of Sequence Traits),...

    arxiv.org/abs/2609.37675 · PDF

  57. 57

    Where Privacy Belongs: Placement Diagnosis and Certified Selection for Private Counterfactual Explanations on Graphs

    Yuxiang Yao, Zijun Zhao

    cs.LG

    Counterfactual explanations for graph neural networks (GNNs) find the minimal intervention that flips a node's prediction--but computing one requires reading sensitive graph structure, and releasing it discloses that structure. Both existing placements fail. Privatizing the graph before explaining corrupts the target on exactly the borderline nodes needing recourse, manufacturing spurious flips that flip the privatized graph but not the true...

    arxiv.org/abs/2609.37667 · PDF

  58. 58

    Learning Causal Normalizing Flows from Incomplete Data via Observed-Data Likelihood

    Trung-Dung Hoang, Alceu Bissoto, Tim Flühmann, David Herzig, Christos Nakas, Lia Bally, Lisa M. Koch

    cs.LG

    Causal Normalizing Flows (CNFs) enable causal inference from observational data given the causal structure, but they assume fully observed training data. We introduce MissCNF, which trains CNFs directly on incomplete data by maximizing the marginal likelihood of each partially observed sample, without discarding rows or constructing a completed dataset. Thanks to the causal structure encoded in the autoregressive factorization of CNFs, only...

    arxiv.org/abs/2609.37664 · PDF

  59. 59

    Nonpreemptive Scheduling While Learning Context-Dependent Service Rates

    Wansoo Choi, Seoungbin Bae, Dabeen Lee

    cs.LG

    We study nonpreemptive contextual queueing bandits in a single-server system. Each job is represented by a $d$-dimensional context vector; in each round, a job may arrive with its context drawn from an unknown distribution $\mathcal{D}$, and its departure probability is determined by a logistic model of that context vector with an unknown parameter $θ^*$. The server learns from service outcomes while deciding which waiting job to serve and...

    arxiv.org/abs/2609.37660 · PDF

  60. 60

    RACE: Relation-Level Counterfactual Explanations for Heterogeneous Graph Neural Networks

    Yuxiang Yao, Zijun Zhao

    cs.LG

    Counterfactual explanations of graph neural networks identify edge deletions that flip a prediction. On heterogeneous graphs, however, existing methods first collapse the graph into untyped edges, so they cannot answer the question a domain expert actually asks: which relation type drives this prediction? We present RACE (Relation-Aware Counterfactual Explanations), which gives this question an exact, per-instance answer. For every explained...

    arxiv.org/abs/2609.37650 · PDF

  61. 61

    RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

    Michael Kirchhof, Eleonora Gualdoni, Andrew Szot, Khashayar Gatmiry, Aryo Lotfi, Abbas Kazerouni, Omar Attia, Sanjoy...

    cs.LG · cs.AI · cs.CL · stat.ML

    The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed...

    arxiv.org/abs/2609.37633 · PDF

  62. 62

    ProCTI: Prototype-Refined Global Conditioning for Diffusion-Based Time Series Imputation

    Fariza Rashid, Duc Van Le, Rahat Masood, Gustavo Batista, Aruna Seneviratne, Suranga Seneviratne

    cs.LG · cs.AI

    Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent performance. Existing diffusion-based methods typically condition the reverse process using local contextual information from the current or neighbouring windows. Meanwhile, global dataset-level structure often remains implicit, limiting performance when local observations are sparse, noisy, or...

    arxiv.org/abs/2609.37632 · PDF

  63. 63

    Procedural Core: A Compact Recurrent Initialization for Vision Transformers

    Zachary Shinnick, Christian Internò, Hemanth Saratchandran, Anton van den Hengel, Damien Teney

    cs.LG · cs.CV

    Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work showed that a small amount of abstract procedurally generated data can help acquire generic inductive structure at low cost. However, this adds a pretraining stage that must be repeated for every target model. We propose Procedural Core, an initialization strategy that captures this generic...

    arxiv.org/abs/2609.37631 · PDF

  64. 64

    Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable

    Abhinav Rajeev Kumar, Paras Chopra

    cs.LG · cs.CL

    Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval results, tool outputs, and grounded-search content often present information. We measure this gap across five open-weight...

    arxiv.org/abs/2609.37616 · PDF

  65. 65

    PHASE: Multi-Regime Modeling of Incompressible Magnetohydrodynamics

    Radhika Achikanath Chirakkara, Rajdeep Haldar, Zezheng Song, Jiequn Han

    cs.LG · physics.comp-ph · physics.plasm-ph

    Magnetohydrodynamics (MHD) is central to plasma modeling in astrophysics, space science, fusion, and engineering, but resolving multiscale MHD dynamics is computationally expensive. Machine-learning surrogates enable fast inference by learning reusable solution operators, yet existing models require separate training for each physical regime, limiting generalization across varying parameter settings. We introduce PHASE, a PHysics-Adaptive...

    arxiv.org/abs/2609.37609 · PDF

  66. 66

    GraphVQ: Structure-Aware Autoregressive Decoding over Context-Quantized Graph Tokens

    Yuxiang Yao, Zijun Zhao

    cs.LG

    Graph foundation models need a discrete token representation, but casting a graph as a generatable token sequence faces a structural obstacle: edges spanning beyond the serialization window cannot be emitted in one pass--so one-pass autoregressive generators systematically under-produce cycles--and a single global condition cannot tell candidate edges apart. GraphVQ removes both obstacles: node contexts--features plus a local edge mask under...

    arxiv.org/abs/2609.37604 · PDF

  67. 67

    ReLMem: Learning Recurrent Memory for Longitudinal EHR Modeling

    Zijie Meng, Xiwei Dai, Yingying Zhang, Jian Wu, Xian Wu, Zuozhu Liu

    cs.LG · cs.AI

    Longitudinal electronic health record (EHR) modeling requires integrating new visits with an expanding patient history. Yet the continual accumulation of clinical information imposes increasing computational and memory costs on large language models (LLMs) when they process and retain complete patient histories. A practical alternative is visit-wise recurrent compression, which incorporates each incoming visit into a compact, continually...

    arxiv.org/abs/2609.37587 · PDF

  68. 68

    A Model-Agnostic Physics-Guided Adapter for Few-Shot Transfer of Coastal Flood Prediction Models to Unseen Regions

    Bilal Hassan, Areg Karapetyan, Samer Madanat

    cs.LG

    Deep learning surrogates can produce high-resolution coastal flood maps orders of magnitude faster than physics-based hydrodynamic simulators, yet transferring them to new coastal regions remains costly, since generating target-region data for fine-tuning typically requires numerous time-consuming simulations. To tackle this bottleneck, we introduce the Physics Adapter (PA), a compact, architecture-agnostic adaptation interface that enables...

    arxiv.org/abs/2609.37565 · PDF

  69. 69

    Benchmarking graph-based models for in-silico toxicity prediction in drug discovery

    Noel Suarez-Barro, Manuel Lama, Juan C. Vidal

    cs.LG · q-bio.BM · q-bio.QM

    Drug discovery is a costly and high-risk process, where toxicity-related failures remain a major cause of attrition in both preclinical and clinical stages. As a result, accurate early prediction of chemical toxicity is essential to reduce downstream costs and improve compound prioritization. In this context, graph deep learning (GDL) has emerged as a powerful paradigm for toxicity prediction, leveraging molecular graph representations to...

    arxiv.org/abs/2609.37555 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.