cs.LG · 2026-09-30 · No. 129
Machine Learning, 2026-09-30.
69 new papers in cs.LG. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
69 entries-
01
Breakdown of Local Denoising as Semantic Speciation
Guangkuo Liu, Mert Okyay, Yifan F. Zhang, Fangjun Hu, Rahul Nandkishore, Xun Gao
cs.LG · cond-mat.dis-nn · cond-mat.stat-mech
The dynamics of generative models exhibit two apparently distinct temporal windows: a speciation window, in which a sample commits to a semantic class, and a nonlocality window, in which local context windows become insufficient for generation. Motivated by evidence of their near-concurrence in a variety of frontier models, we investigate their relationship through the spatial distribution of semantic information. Under a "common cause"...
-
02
LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci,...
cs.LG · cs.AI
Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can...
-
03
A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization
Jianru Shen
cs.LG · cs.DM
Evaluations of graph reconstruction by language models typically report a single aggregate distance between the original and the reconstructed graph. We prove that for the Wasserstein distance between Laplacian spectra such a summary is bracketed by two edge counts, the net change in edge number from below and the symmetric difference from above, each scaled by $2/n$ where $n$ is the number of vertices. The bracket is sharp: its two ends...
-
04
Multi-Agent Flow Matching with Decoupled Generative Guidance
Ruoyu Lin, Magnus Egerstedt, Fabio Pasqualetti
cs.LG · cs.MA · cs.RO · math.OC
Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general, come with formal guarantees that the generated objects satisfy hard constraints or requirements. In multi-agent generation, this problem becomes more challenging because a hard requirement can depend on multiple agents, while each agent may need to determine its own guidance input without...
-
05
Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs
Yige Hong, Xiangcheng Zhang, Qiaomin Xie, Yudong Chen, Weina Wang
cs.LG · math.OC · math.PR
We study average-reward weakly-coupled Markov decision processes (WCMDPs), where a WCMDP consists of $N$ smaller MDPs, called arms, that share multiple per-step budget constraints. We consider the setting where the arms have identical model parameters, multiple actions, and state- and action-dependent costs. For restless bandits (RBs), a well-studied special case of WCMDPs, prior work has developed policies that achieve an $O(1/\sqrt{N})$...
-
06
WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
Jiale Chen, Vage Egiazarian, Eldar Kurtić, Torsten Hoefler, Dan Alistarh
cs.LG
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with...
-
07
Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
Panagiotis Theodoropoulos, Nan Jiang, Xintong Duan, Ali Hasan, Yuriy Nevmyvaka, Evangelos A. Theodorou, Wei Deng
cs.LG
Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged generalization of RL. However, this approach faces a fundamental exploration--exploitation trade-off, as % strong sharpening...
-
08
Tail-Influence Sampling for CVaR Policy Evaluation
Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou-Ammar
cs.LG
Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to estimate a fixed policy's CVaR most accurately. We derive a tail influence for each queryable conditional law that aggregates...
-
09
Probe-Space Preconditioning for Fast and Stable Zero-Order Training
Francois Chaubard, Mykel J. Kochenderfer, Chris Ré
cs.LG
Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU memory (assuming batch size 8 and sequence length 2048). Alternatively, zero-order optimization (ZOO) trains in inference-mode (requiring only $\approx$ 60GB for the same model): no stored activations, no gradients, and no optimizer states. However, ZOO convergence has lagged behind BP. In this...
-
10
Dimensionally consistent surrogate modelling through dimensional analysis and harmonic expansions
Ernest Tarrus, Hector Gisbert
cs.LG · hep-ph
Dimensional homogeneity is a fundamental constraint on physically meaningful models, requiring invariance under changes of units. We present a data-driven method for constructing surrogate models that satisfy this constraint at the level of the hypothesis class. Starting from a dimension matrix of measured variables, the method derives Buckingham $Π$-groups, constructs admissible dimensional prefactors, and approximates the remaining...
-
11
Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging
Sanjali Yadav, Bahar Asgari
cs.LG
Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-constrained, single-GPU systems. However, this benefit is difficult to realize because expert parameters dominate memory, and token-level routing is dynamic, unpredictable, and skewed. Prior work using offloading and caching remains fundamentally reactive, as systems wait for router outputs...
-
12
Traversing the solution space of neural networks with Hessian Null Space Continuation
Ann Huang, Mitchell Ostrow, Zhouyang Lu, William T. Redman, Leo Kozachkov, Kanaka Rajan
cs.LG · q-bio.NC · stat.ML
On a single task, deep networks can learn many solutions, depending on their optimizer, training data, architecture, and hyperparameters. Many of these solutions are mode-connected: rather than isolated points in weight space, they are connected by low-loss regions. Yet how their internal computation varies within these regions is unknown. A parallel line of work has identified the degeneracy of neural representations: many networks reach...
-
13
A foundation model for energy and radiation systems built on heterogeneous scientific interfaces
Samrendra Roy, Tapas Tripura, Yoon Pyo Lee, Souvik Chakraborty, Syed Bahauddin Alam
cs.LG · physics.comp-ph
Scientific foundation models are commonly evaluated after heterogeneous physical problems have already been translated into a compatible gridded, tokenized or symbolic representation. This leaves the scientific interface outside both the pretrained model and the audit of what is actually reused. We study the complementary setting in which boundary histories, sparse monitor records and loading histories retain their native inference classes...
-
14
Alpha Diffusion Language Models: Factorization Alone Is Not the Problem
Nikita Gushchin, Dmitry Baranchuk, Alexander Korotin
cs.LG
Discrete diffusion language models can generate multiple tokens in parallel, but reducing the number of denoising steps can lead to inconsistent predictions. Standard cross-entropy training fits conditional token marginals, whereas parallel generation requires consistent joint predictions. We introduce Alpha Diffusion Language Models (AlphaDLM), trained with a sequence-level alpha loss that recovers cross-entropy in the limit of vanishing...
-
15
Jaxolotl: A Unified High-Performance Benchmark Suite for LTL-Based Multi-Task RL
Mathias Jackermeier, Jacques Cloete, Alessandro Abate
cs.LG · cs.AI
Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully adopted for training generalist multi-task policies. However, differences in implementations, task distributions, and evaluation protocols make existing methods difficult to compare, while high...
-
16
Improving Function Space Flow Matching with Kernel Optimal Transport
Fred Xu, Thomas Markovich, Barbora Barancikova, Yizhou Sun
cs.LG
Generative models for function-valued data, such as time series and solutions of partial differential equations, must learn distributions over infinite-dimensional spaces. Functional Flow Matching (FFM) extends Flow Matching to this setting, learning a velocity field whose flow transports a Gaussian prior to the data distribution, but it inherits the independent endpoint pairing of standard Flow Matching: in each batch, prior and data samples...
-
17
Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models
Zhenyu Wang, Tianze Wang, Linjun Zhang, Yifan Hu
cs.LG · cs.AI · cs.CL · math.OC
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect...
-
18
Prompts Live on an Arc: Gaussian Curricula in Fisher--Rao Coordinates for Rollout-Efficient GRPO
Mei Okonkwo, Pixel Nomand, Julian Berg, Elena Voss, Lena Park, Marcus Hale, Adrian Cho, Sofia Reyes
cs.LG
Group relative policy optimization (GRPO) learns only from prompts whose sampled responses disagree: a group that is entirely correct or entirely incorrect has zero reward variance, contributes no gradient, and still consumes its rollouts. Prompt-selection methods reduce this waste by steering sampling toward intermediate pass rates, but they choose the target, its width, and the uncertainty model heuristically, in raw pass-rate or logit...
-
19
When do data mixtures improve scaling laws? Insights from high-dimensional regression
Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli
cs.LG · stat.ML
Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downstream performance. Despite an extensive literature on data mixing and reweighting, existing work is largely empirical and it remains unclear when auxiliary data genuinely improves scaling laws rather than merely providing more samples. To gain insight into this question, we study a...
-
20
No Scale Left Behind: Multi-Scale Autoencoder with Bi-directional Attention for Time Series Anomaly Detection
Jiaheng Guo, Haochen Zhang, Yu-Chao Huang, Jinhao Duan, Nicholas Konz, Tianlong Chen
cs.LG · cs.AI
Time series anomaly detection (TSAD) plays a crucial role in healthcare, finance, industrial monitoring, and other sectors. Within and between these settings, anomalies span vastly different temporal scales, from sub-second point spikes to multi-hour drift patterns. However, most existing TSAD methods commit to a single temporal granularity, and multi-scale designs either analyze different scales in isolation or are constrained to a...
-
21
TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models
Deqing Fu, Huangyuan Su, Rajat Sen, Taman Narayan, Sujay Sanghavi, Abhimanyu Das, Weihao Kong
cs.LG
Tabular foundation models achieve strong zero-shot accuracy on structured data by pretraining on synthetic tables, but they ignore the column names, task descriptions, and auxiliary files that carry dataset semantics. Meanwhile, self-evolving machine learning engineering (MLE) agents train models from scratch on each dataset, yet jointly searching over features, architectures, and hyperparameters is noisy and prone to overfitting. We...
-
22
$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient
Hongbo Ma, Sansheng Cao, Jiajun Fan, Bangji Yang, Ge Liu
cs.LG · cs.AI · cs.CL · cs.CV
LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during...
-
23
On Trajectory-Aware Training for Masked Diffusion Language Models
Manuel Madeira, Amitis Shidani, Alice Bizeul, Victor Turrisi, Louis Béthune, Bhavika Devnani, Dan Busbridge, Pierre...
cs.LG · cs.AI · cs.CL
Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model's own predictions. Additionally, each step has no access to what the previous one computed. Recent methods narrow these limitations from separate angles, leaving open how these choices...
-
24
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
cs.LG
Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. Trained entirely on synthetic tables generated from structural...
-
25
Kolmogorov-Arnold Classifier Systems as Universal Approximators
Hiroki Shiraishi, Hisao Ishibuchi, Masaya Nakata
cs.LG · cs.NE
As the input dimension $n$ grows, rule-based machine learning, such as Learning Classifier Systems (LCSs), faces a fundamental scalability bottleneck for function approximation: both rule count and parameter count grow exponentially with $n$. Traditional LCSs partition the $n$-dimensional input space directly, requiring $\mathcal{O}(m^n)$ rules for adequate coverage, where $m$ is the per-variable resolution. This article breaks from this...
-
26
Scene-Consistent Illumination Transfer for Inserted Advertising Graphics
Rameshwar Mishra, Bishshoy Das, A. V. Subramanyam, Guan-Ming Su
cs.LG
Replacing a visible advertisement in a broadcast frame is geometrically straightforward but photometrically delicate. A pasted graphic can have the correct perspective and still appear detached when its brightness, shading, or shadow disagrees with the surface beneath it. This paper presents Ad-Relight, an inference-only procedure for transferring scene illumination to a supplied advertising graphic without collecting a banner-specific...
-
27
An Efficient Machine Learning Approach for Degradation Forecasting in AEM Water Electrolysis
Marco Veneriano, Ani Gjergji, Sebastiano Bellani, Andrea Riva, Vito Paolo Pastore, Matteo Santacesaria
cs.LG · eess.SY · physics.app-ph
This study provides a data-driven analysis of a novel dataset of single-cell Anion Exchange Membrane water electrolyzers (AEMWE), operated under constant current load across multiple heterogeneous experimental campaigns. We train and evaluate a range of machine learning models with different complexity, including linear baselines, LSTMs and CNNs, to perform medium-term forecasting of the cell voltage degradation curve. The models are assessed...
-
28
Learning When to Update: A Near-Optimal Timing Bandit Approach
Qiulin Lin, Junyan Su, Liyuan Wang, Minghua Chen
cs.LG
Systems operating in dynamic environments require timely updates to sustain performance. For resource-intensive systems such as machine learning models and digital twins, strategically timing updates is essential. Updating too frequently wastes resources, while updating too infrequently leads to costly performance degradation. The problem is particularly challenging when the system's degradation pattern is unknown a priori, as is common in...
-
29
Time-Anchored Diffusion Language Models: Latent-Space Caching for Fast Generation
Joel Anto Paul, Litu Rout, Aditya Akella, Sanjay Shakkottai
cs.LG · cs.CL
Recent work on anchored diffusion language models improves denoising by shaping an intermediate latent space with supervised important-token targets. In this work, we introduce time-based (self-supervised) anchoring, which learns and reuses latent anchors without requiring such targets. Our key observation is that anchors encode persistent properties of the clean sequence, such as its semantic intent, global structure, or intermediate plan....
-
30
Pattern Formation in Transformers
Erkan Turan, Gaspard Abel, Maks Ovsjanikov
cs.LG
What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Transformers escape from rank collapse or demonstrates that self-attention drives tokens toward cluster patterns. The latter view arises from an elegant dynamical systems perspective, but relies on simplified architectural assumptions, and does not explain the rich structures observed in...
-
31
Search Dimension in Unlabeled Projection Pursuit: A Scaling Law for Subspace Restriction
Rares Dimitrie Grozavescu, Mark Girolami
cs.LG · math.ST · stat.ML
Projection pursuit searches for a direction along which the data look least Gaussian. When the observation space contains a large Gaussian complement, the empirical objective can be minimized by a direction that carries no signal, with empirical kurtosis as low as at the truth. Sample splitting exposes rather than repairs this failure. Appending coordinates independent of the latent regime degrades the search while leaving Bayes...
-
32
Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning
Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni
cs.LG · cs.CL
On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factorial analysis and find that scaffold correctness has a stronger effect on downstream accuracy than context correctness....
-
33
Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction
Shivang Chopra, Fotis Iliopoulos, Zsolt Kira, Gaurav Menghani
cs.LG · cs.AI
Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come from increasing the interaction capacity of a single predictor through deeper cross networks and more expressive cross operators. We revisit whether continually increasing interaction...
-
34
Scaling Zero-Order Pretraining through Model Sharding
Francois Chaubard, Mykel J. Kochenderfer, Chris Ré
cs.LG
Zero-order optimization (ZO) trains without backpropagation, making it relevant to forward-only hardware and non-differentiable loss, but its gradient variance grows with perturbed dimension, inhibiting large-model training. Sharded Optimization Mixture of Assemblies (SOMA) trains LSTM experts independently on $N$ data clusters using simultaneous perturbation stochastic approximation (SPSA), without exchanging gradients, activations or...
-
35
Behavioral Capacity Certificates for Quantized Language Models
Arian Eamaz, Mojtaba Soltanalian
cs.LG
Activation and key-value cache precision change what a quantized language model computes without altering its stored weights. Direct weight-code bounds, however, assign identical complexity to deployments that behave differently and charge separately for weight codes that behave identically. Behavioral Capacity Certificates (BCC) charge for behavior using the aggregate prior mass of complete implementations---weights, scales, activation and...
-
36
TopoEmbedX: A General Framework for Representation Learning on Topological Domains
Florian Frantzen, Ibrahem AlJabea, Ines Henriques-Cadby, Theodore Papamarkou, Mustafa Hajij, Michael T. Schaub
cs.LG
Topological structures such as simplicial complexes, hypergraphs, and cell complexes extend standard graph models by modeling higher-order relationships. These structures appear in many modern datasets and require specialized methods for generating meaningful embeddings. In this paper, we introduce TopoEmbedX, a unified framework for embedding a wide range of topological domains into Euclidean spaces. The package brings together several...
-
37
Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
Doohyuk Jang, Yoonsik Park, Gyouk Chu, Sihwan Park, Eunho Yang
cs.LG · cs.AI · cs.CL
Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models...
-
38
Strict-Saddle Landscapes and Multi-Rank Geometry in Low-Tubal-Rank Tensor Sensing
Eugene Agyei-Kodie, Longxiu Huang, Shuang Li, Xiao Liang
cs.LG · math.OC
We study the optimization landscape of low-tubal-rank tensor sensing through a balanced factorization. Under a tubal restricted isometry condition, we establish a quantitative strict-saddle landscape with no spurious local minima for arbitrary Fourier multi-rank profiles. We further show that the local geometry depends on the Fourier-slice ranks rather than the tubal rank alone. Uniform ranks yield quadratic growth transverse to the solution...
-
39
Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning
Tianhao Qian, Ziming Hong, Chongyang Gao, Kezhen Chen, Lixu Wang
cs.LG · cs.CL
Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve...
-
40
Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs
Haozhan Tang, Hao Kang, Han Cai, Song Han, Chenyan Xiong
cs.LG
Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases and downstream degradation at 1.67B and 5.29B. QK normalization, NoPE (no positional encoding), and lower-learning-rate context...
-
41
Scaling Influence Functions in LLMs through Eigenbasis-Corrected One-Bit Gradient Projection
Jaeseung Heo, J Rosser, Dongwoo Kim
cs.LG · cs.AI
Influence functions estimate how individual training examples affect the behavior of large language models (LLMs). Analyzing how training data influence different behaviors of an LLM involves repeated influence computation. Reusing stored training gradients reduces the computational cost, but storing full gradients is prohibitively expensive at LLM scale. We study how to compress these gradients while preserving influence estimates for future...
-
42
Counterfactual Probing for Parallel Unmasking with Hidden Forest Structure
Ryotaro Kawata, Satoshi Hayakawa, Taiji Suzuki
cs.LG · stat.ML
Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies are unknown, finding safe batches also costs model evaluations. We study whether total evaluations, including discovery, can be sublinear in sequence length $N$; sublinear sequential depth then follows. We consider discrete distributions with hidden forest structure, accessed through a fixed...
-
43
Behavioral Convergence Without Representational Convergence: Persistent Training-History Dependence in Neural Networks
Ertuğrul Mutlu
cs.LG
Neural networks trained toward the same final objective can reach similar predictive performance while retaining internal representations shaped by earlier training history. We study this effect using controlled sequential-training experiments in which paired convolutional networks start from identical weights, experience reversed task orders, and then receive the same deterministic common-relaxation distribution. Across 20 paired MNIST runs,...
-
44
Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR
Kun Liang, Chenming Tang, Clive Bai, Weijie Liu, Zeyuan Liu, Qingyang Zhang, Saiyong Yang, Yunfang Wu
cs.LG · cs.AI
Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges on reliable value estimation, a difficult task requiring the critic to both assess progress toward a correct solution and...
-
45
Feedback-Calibrated Protein Optimization with Batch-Aligned Tail Arbitration
Zefeng Lin, Xianyong Fang, Tianfan Fu, Xiaohua Xu
cs.LG
Protein optimization aims to discover high-fitness sequences under a limited experimental budget. Existing machine-learning methods use task-specific predictors, biological priors, or ranking-aware objectives to guide which variants are tested in the next experimental round. However, these methods cannot adapt to shifts in the reliability of predictive evidence as measurements accumulate and ensure the correct ranking of key high-fitness...
-
46
Challenges and Solutions for Bandits in the Wild: Warm-Started Mixture Bandits for Cross-Cohort Slate Recommendation
Serafima Lebedeva, Sumantrak Mukherjee, Ali Arshad Sadal, Ilias Ekşi, Rahul Sharma, Julia Mueller, Theresa...
cs.LG · cs.AI
Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two challenges: learning user preferences quickly from limited feedback and sustaining useful recommendations when each user has a finite catalog that can become repetitive or depleted over time. We propose CohortMix-TS, a warm-started mixture bandit that learns latent user groups from earlier cohorts...
-
47
Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance
Fabian A. Mikulasch, Friedemann Zenke
cs.LG · cs.AI
Self-supervised learning (SSL) by predicting in latent space, without generating the input data itself, learns highly abstract, useful representations. Intuitively, this success is often attributed to its ability to discard nuisance information that is irrelevant to prediction. However, this poses a conundrum: both stochastic variation in a prediction-relevant latent signal and true nuisance make observations partly unpredictable; how could...
-
48
Optimizer-dependent training dynamics converge to the same one-third optimal data scaling
Hyunseok Lee, Mihir Basil, Yizhou Liu, Jeff Gore
cs.LG
Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a $1/3$ exponent emerges from learning peaked distributions. That account describes SGD, but models in practice are trained with adaptive optimizers. Here we separate two exponents the $1/3$ account does not distinguish: how fast the loss falls with training steps along a single run, and how fast the optimally...
-
49
HyDI: A hybrid Deep Learning-Inductive Logic Programming ensemble for multi-label classification
Simon Flügel, Till Mossakowski
cs.LG · cs.LO
While attaining remarkable results for many applications, Deep Learning models are notoriously difficult to explain. This work introduces HyDI, a hybrid ensemble architecture for hierarchical multi-label classification. It combines a Deep Learning (DL) model with rule-based classifiers generated by Inductive Logic Programming (ILP). For leaf classes of the label hierarchy, the rule-based classifiers replace the DL model, leading to more...
-
50
Weights Read and Write Features: Scalable Parameter Decomposition Grounded in Activation Space
Tue M. Cao, Lisiane Pruinelli, My T. Thai
cs.LG
Activation space and parameter space provide complementary views of model computation. Activations represent information, while weights read, transform, and write that information. Yet existing interpretability methods largely study the two spaces separately, leaving the connection between represented information and parameter-level computation underexplored. We introduce Activation-Supported Parameter Decomposition (ASPD), which jointly...
-
51
Predictive Geometry of Hidden Trajectories in Transformers
Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu State
cs.LG · cs.CL
Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state through the fixed downstream computation. We formalize this constraint by studying layerwise loss-to-go functions: the terminal loss obtained by continuing a candidate hidden state through the remaining transformer blocks. Around successful validation trajectories, we show that the local...
-
52
Volatility-Clustering Adaptation for Financial Time Series
Manh Nguyen, Minh Hoang Nguyen, Huu Hiep Nguyen, Van Dai Do, Hung Le
cs.LG
Time-series foundation models are increasingly adapted to new domains through fine-tuning on target data, under the implicit assumption that more target data yields better forecasts. We show that this assumption can fail in financial forecasting, where individual price changes are difficult to predict, but large moves tend to cluster, creating alternating calm and turbulent periods. Using financial foundation models trained on price bars of...
-
53
Width Expansion as a Method for Class Incremental Learning
A. L. S. Conde, Y. Elkhatib, C. M. Ranieri
cs.LG · cs.AI
Class Incremental Learning (Class-IL) requires models to learn new classes over time while preserving previously acquired knowledge without access to past data or task identity. This setting intensifies the stability-plasticity dilemma and makes catastrophic forgetting a central challenge. Existing approaches include regularization, knowledge distillation, replay, and architectural expansion. However, many expansion methods rely on explicit...
-
54
GARDiff: Graph-Aligned Residual Diffusion for Probabilistic Multivariate Time-Series Forecasting
Rui Han, Min Yang, Xu Zhang, Xinghao Yang, Wei Liu, Yongshun Gong
cs.LG · cs.AI
Diffusion models have recently shown strong potential for probabilistic multivariate time-series forecasting by modeling complex conditional distributions. Recent decoupled diffusion frameworks further separate forecasting into deterministic prediction and stochastic residual generation, making it natural to derive dependency graphs from deterministic representations and use them to guide residual diffusion. However, we show that this direct...
-
55
When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task
Sai Sumedh R. Hindupur, Hadas Orgad, Thomas Fel, Demba Ba
cs.LG · cs.AI · cs.CL
One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literature (e.g. numbers encoded on helices, days of the week on a circle, ...), with structure believed to reflect properties of data...
-
56
LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling
Biswajit Banerjee, Claudia Alvarez Carreno, Anton S. Petrov
cs.LG · q-bio.BM
Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and structure-based tasks, yet the potential of tokenization remains underexploited. Unlike human language, proteins preserve structure despite extensive sequence variation a property standard tokenization strategies fundamentally fail to capture. We introduce ZEST (Zoned Encoding of Sequence Traits),...
-
57
Where Privacy Belongs: Placement Diagnosis and Certified Selection for Private Counterfactual Explanations on Graphs
Yuxiang Yao, Zijun Zhao
cs.LG
Counterfactual explanations for graph neural networks (GNNs) find the minimal intervention that flips a node's prediction--but computing one requires reading sensitive graph structure, and releasing it discloses that structure. Both existing placements fail. Privatizing the graph before explaining corrupts the target on exactly the borderline nodes needing recourse, manufacturing spurious flips that flip the privatized graph but not the true...
-
58
Learning Causal Normalizing Flows from Incomplete Data via Observed-Data Likelihood
Trung-Dung Hoang, Alceu Bissoto, Tim Flühmann, David Herzig, Christos Nakas, Lia Bally, Lisa M. Koch
cs.LG
Causal Normalizing Flows (CNFs) enable causal inference from observational data given the causal structure, but they assume fully observed training data. We introduce MissCNF, which trains CNFs directly on incomplete data by maximizing the marginal likelihood of each partially observed sample, without discarding rows or constructing a completed dataset. Thanks to the causal structure encoded in the autoregressive factorization of CNFs, only...
-
59
Nonpreemptive Scheduling While Learning Context-Dependent Service Rates
Wansoo Choi, Seoungbin Bae, Dabeen Lee
cs.LG
We study nonpreemptive contextual queueing bandits in a single-server system. Each job is represented by a $d$-dimensional context vector; in each round, a job may arrive with its context drawn from an unknown distribution $\mathcal{D}$, and its departure probability is determined by a logistic model of that context vector with an unknown parameter $θ^*$. The server learns from service outcomes while deciding which waiting job to serve and...
-
60
RACE: Relation-Level Counterfactual Explanations for Heterogeneous Graph Neural Networks
Yuxiang Yao, Zijun Zhao
cs.LG
Counterfactual explanations of graph neural networks identify edge deletions that flip a prediction. On heterogeneous graphs, however, existing methods first collapse the graph into untyped edges, so they cannot answer the question a domain expert actually asks: which relation type drives this prediction? We present RACE (Relation-Aware Counterfactual Explanations), which gives this question an exact, per-instance answer. For every explained...
-
61
RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
Michael Kirchhof, Eleonora Gualdoni, Andrew Szot, Khashayar Gatmiry, Aryo Lotfi, Abbas Kazerouni, Omar Attia, Sanjoy...
cs.LG · cs.AI · cs.CL · stat.ML
The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed...
-
62
ProCTI: Prototype-Refined Global Conditioning for Diffusion-Based Time Series Imputation
Fariza Rashid, Duc Van Le, Rahat Masood, Gustavo Batista, Aruna Seneviratne, Suranga Seneviratne
cs.LG · cs.AI
Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent performance. Existing diffusion-based methods typically condition the reverse process using local contextual information from the current or neighbouring windows. Meanwhile, global dataset-level structure often remains implicit, limiting performance when local observations are sparse, noisy, or...
-
63
Procedural Core: A Compact Recurrent Initialization for Vision Transformers
Zachary Shinnick, Christian Internò, Hemanth Saratchandran, Anton van den Hengel, Damien Teney
cs.LG · cs.CV
Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work showed that a small amount of abstract procedurally generated data can help acquire generic inductive structure at low cost. However, this adds a pretraining stage that must be repeated for every target model. We propose Procedural Core, an initialization strategy that captures this generic...
-
64
Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable
Abhinav Rajeev Kumar, Paras Chopra
cs.LG · cs.CL
Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval results, tool outputs, and grounded-search content often present information. We measure this gap across five open-weight...
-
65
PHASE: Multi-Regime Modeling of Incompressible Magnetohydrodynamics
Radhika Achikanath Chirakkara, Rajdeep Haldar, Zezheng Song, Jiequn Han
cs.LG · physics.comp-ph · physics.plasm-ph
Magnetohydrodynamics (MHD) is central to plasma modeling in astrophysics, space science, fusion, and engineering, but resolving multiscale MHD dynamics is computationally expensive. Machine-learning surrogates enable fast inference by learning reusable solution operators, yet existing models require separate training for each physical regime, limiting generalization across varying parameter settings. We introduce PHASE, a PHysics-Adaptive...
-
66
GraphVQ: Structure-Aware Autoregressive Decoding over Context-Quantized Graph Tokens
Yuxiang Yao, Zijun Zhao
cs.LG
Graph foundation models need a discrete token representation, but casting a graph as a generatable token sequence faces a structural obstacle: edges spanning beyond the serialization window cannot be emitted in one pass--so one-pass autoregressive generators systematically under-produce cycles--and a single global condition cannot tell candidate edges apart. GraphVQ removes both obstacles: node contexts--features plus a local edge mask under...
-
67
ReLMem: Learning Recurrent Memory for Longitudinal EHR Modeling
Zijie Meng, Xiwei Dai, Yingying Zhang, Jian Wu, Xian Wu, Zuozhu Liu
cs.LG · cs.AI
Longitudinal electronic health record (EHR) modeling requires integrating new visits with an expanding patient history. Yet the continual accumulation of clinical information imposes increasing computational and memory costs on large language models (LLMs) when they process and retain complete patient histories. A practical alternative is visit-wise recurrent compression, which incorporates each incoming visit into a compact, continually...
-
68
A Model-Agnostic Physics-Guided Adapter for Few-Shot Transfer of Coastal Flood Prediction Models to Unseen Regions
Bilal Hassan, Areg Karapetyan, Samer Madanat
cs.LG
Deep learning surrogates can produce high-resolution coastal flood maps orders of magnitude faster than physics-based hydrodynamic simulators, yet transferring them to new coastal regions remains costly, since generating target-region data for fine-tuning typically requires numerous time-consuming simulations. To tackle this bottleneck, we introduce the Physics Adapter (PA), a compact, architecture-agnostic adaptation interface that enables...
-
69
Benchmarking graph-based models for in-silico toxicity prediction in drug discovery
Noel Suarez-Barro, Manuel Lama, Juan C. Vidal
cs.LG · q-bio.BM · q-bio.QM
Drug discovery is a costly and high-risk process, where toxicity-related failures remain a major cause of attrition in both preclinical and clinical stages. As a result, accurate early prediction of chemical toxicity is essential to reduce downstream costs and improve compound prioritization. In this context, graph deep learning (GDL) has emerged as a powerful paradigm for toxicity prediction, leveraging molecular graph representations to...
This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.