cs.LG · 2026-09-29 · No. 128
Machine Learning, 2026-09-29.
84 new papers in cs.LG. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
84 entries-
01
Unifying Distributional Training for One-Step Visual Generation
Chi Zhang, Haoyang Shi, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu,...
cs.LG
\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are...
-
02
TokenCast: Forecasting Token Consumption During LLM Agent Execution
Chaoqian Ouyang, Ling Yue, Libin Zheng, Huanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, Shimin Di
cs.LG · cs.AI · cs.SE
When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this...
-
03
Neural Harmonic Measure Operator
Jinjin He, Sinan Wang, Yuchen Sun, Bo Zhu
cs.LG · math.NA
We introduce Neural Harmonic Measure Operator (NHMO), a neural solver for elliptic PDE problems on variable-shape domains. The harmonic measure of a domain is the boundary probability distribution that, integrated against any boundary data, returns the Dirichlet Laplace solution. It depends only on the geometry, not on the boundary data. NHMO parameterizes the density of this measure as a transformer-based boundary kernel supervised by...
-
04
How to Loop MoE: Flatten the Experts, Untie the Attention
Shouren Wang, Chuang Ma, Mohsen Hariri, Debargha Ganguly, Wang Yang, Xiaoqing Tong, Qianying Liu, Xiaotian Han,...
cs.LG · cs.AI · cs.CL
Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil....
-
05
KV-streams for Efficient Compaction in Agentic Reinforcement Learning
Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri,...
cs.LG · cs.AI
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we...
-
06
X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets
Prithwish Dan, Chenyang Ma, Wei Zhan
cs.LG · cs.AI · cs.RO
Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or...
-
07
ScAn-Bench: Evaluating Scaling Analysis Methodology
Artin Sermaxhaj, Nastaran Alipour, Donat Sinani, Johannes Hog, Neeratyoy Mallik, Jenia Jitsev, Danny Stoll
cs.LG
Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types. To shed light on this crucial blind spot and facilitate future...
-
08
A Unified Uncertainty Representation for Graph Neural Networks via Doubly-Spectral Stochastic Expansion
Fred Xu, Thomas Markovich, Florence Regol, Yizhou Sun
cs.LG · cs.AI
Reliable deployment of graph neural networks requires calibration, out-of-distribution (OOD) detection, and robustness to distribution shift, yet existing methods address these needs with separate models and objectives. We model uncertain node embeddings as random graph signals: graph Fourier filters capture structural variation, and a scalar orthogonal-polynomial chaos coordinate captures latent stochastic variation. The resulting...
-
09
MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining
Chang-Wei Shi, Xu Wang, Wu-Jun Li
cs.LG
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose...
-
10
Distillation Defenses Easily Break After Reinforcement Learning
Shidan Javaheri, Alexander Panfilov, Oliver Britton, Yarin Gal, Yonatan Gideoni
cs.LG · cs.AI · cs.CR
Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming...
-
11
Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning
Hanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang
cs.LG
We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components of empirically successful methods such as GAIL and LS-IQ, yet their finite-sample benefits remain underexplored. We establish fast rates for...
-
12
Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models
Qiyao Ma, Junshan Zhang, Zhe Zhao
cs.LG · cs.CL
Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, empirical studies are first used to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment. We demonstrate that personalized generation is uniquely suited for test-time scaling methods like Best-of-N (BoN)...
-
13
Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?
Li Zhang, Chuqin Geng, Mark Zhang, Chen Yang, Luke Zhang, Haolin Ye, Xujie Si
cs.LG · cs.AI
Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its...
-
14
Transferable Mass Spectrum Prediction via Reference-Guided Test-time Specialization
Yunhua Zhong, Runting Li, Yifan Li, Pan Liu, Zhiwen Yang, Zikun Wang, Yixuan Tang, Jun Xia
cs.LG
Tandem mass spectrum prediction supports compound identification across metabolomics, natural-product discovery, and environmental analysis. However, pretrained predictors often degrade under shifts in chemical space and acquisition conditions, while retraining domain-specific models from scratch is costly. We introduce SPARC, a retrieval-guided test-time specialization framework that adapts a pretrained predictor using a spectral reference...
-
15
Bounding Retraining Equivalence and the Deletion Floor in Materials Machine Unlearning
Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban
cs.LG · cond-mat.mtrl-sci · quant-ph
In materials machine learning, closely related retained structures can sustain accurate property predictions even after removing a specific record, rendering post-deletion prediction error an ambiguous metric for machine unlearning. To resolve this ambiguity, we define the deletion floor as the expected target loss under a specified retraining procedure at the deleted request. Standard indistinguishability constraints yield a sharp interval...
-
16
DR-net-Mamba: Selective State-Space Modeling for Long-Range ECG Time-Series Denoising
Basile Morel, Samuel Ruiperez-Campillo, Andreas P. Streich, Julia E. Vogt, Thomas Hofmann
cs.LG · cs.AI
Electrocardiogram (ECG) recordings are corrupted by non-stationary noise sources that degrade diagnostic reliability, particularly in ambulatory and long-duration recordings. Deep learning denoisers exist, but convolutional architectures are limited by their receptive field, transformer-based models scale quadratically with sequence length, and diffusion-based approaches incur prohibitive inference cost. We propose a Mamba-augmented model...
-
17
SANTA++: Sampling Attention through Representative Keys
Kyle Lee, Christian Z. Pratt, Ruoyu Fang, Heekyung Lee, Avinash Lohitsa, Ryan Modafe, Kerem Y. Camsari
cs.LG · cs.CL
Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to...
-
18
Arbitrary-Accuracy Neural Approximation with Optimal Neuron Count and Near-Optimal Bit Complexity
Zilan Cheng, Li-Lian Wang, Zhongjian Wang
cs.LG · cs.NE · stat.ML
We study the minimum number of hidden neurons required for arbitrary-accuracy approximation of multivariate Hölder-continuous functions on $[0,1]^d$ and the associated encoding complexity. For $d\geq 2$, we construct a fixed, explicitly defined activation function for which a closed-form network with two hidden layers of widths $d$ and $1$ achieves arbitrary accuracy in the uniform norm. We prove that $d+1$ is the exact minimum total number...
-
19
Cartridges++: KV Cache Compression without Off-Context Derailment
Sonia Laguna, Joao Monteiro, Marco Cuturi, Pierre Ablin, Eleonora Gualdoni
cs.LG
Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference time. Methods to obtain CKVs range from drop mechanisms that reduce their number of columns, to learned approaches. Among the...
-
20
Attention Graphons: A Graph Limit Perspective on Graph Transformers
Caio F. Deberaldini Netto, Moshe Eliasof, Luana Ruiz
cs.LG
Graph Transformers produce, for each attention head, a dense $n\times n$ matrix of learned pairwise interactions. We ask a fundamental question: do these attention-induced graphs converge to a stable limit object as $n$ grows, or does the learned interaction pattern remain unstructured and size-dependent? We answer this using dense graph limit theory, treating each attention matrix as a finite sample from an underlying kernel---an...
-
21
Behavioral Foundation Models for Quality Diversity
Nazim Bendib, Nicolas Perrin-Gilbert, Olivier Sigaud
cs.LG · cs.AI
Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting the structure of a latent space. In this work, we investigate whether the latent behavioral space induced by BFMs can serve as an effective search...
-
22
EvE: An Alternate Optimizer to Adam
Shashank Raj, Kalyanmoy Deb
cs.LG · cs.NE
Adam and its variants dominate neural network training, but a single run only reveals whether a configuration works well after most of its budget is spent, a poor fit for hyperparameter or architecture search, where configurations must be ranked cheaply and pruned early. We introduce EvE (Evolutionary Explorer), a steady-state, population-of-four differential evolution (DE) optimizer with a targeted Adam fallback: each iteration proposes one...
-
23
Twist, Don't Tilt: Trajectory-Exact Constrained Decoding for Masked Diffusion Models
Aditya Thimmaiah, Lara Marinov, Jayanth Srinivasa, Haris Vikalo, Junyi Jessy Li, Milos Gligoric
cs.LG · cs.AI · cs.CL
Constrained decoding for Masked Diffusion Language Models (MDLMs) aims to ensure that generated outputs satisfy a specified structure or syntax constraint. MDLMs generate outputs by repeatedly unmasking masked positions present in their current state. Recent strategies for constrained decoding constrain the model's per-step mean-field posterior (which factorizes over masked positions) by enforcing the desired constraint with an automaton. The...
-
24
Control-Geometry Straightening for Sampling-Based Latent Planning
Ziang Fu, Ning Ning
cs.LG · stat.ML
Joint-embedding predictive architectures enable planning with latent world models, but accurate transition prediction alone does not ensure that the planning objective is easy to optimize. We introduce Control-Geometry Straightening (CGS), a single auxiliary loss that learns planner-friendly representations by directly straightening control geometry for sampling-efficient planning. CGS matches pairwise cosine similarities among actions to...
-
25
Hardware-Aware Features for CUTLASS Kernel Selection
Shriram Chandran, Dominic Rinderer, Yakup Budanaz, Alexandru Calotoiu, Marcin Copik, Torsten Hoefler
cs.LG · cs.PF
GPU libraries such as CUTLASS expose tens of thousands of semantically equivalent kernels for a single operation, making exhaustive autotuning expensive and execution-free selection difficult. Existing analytical selectors require hand-designed performance rules, while learned selectors operate on raw configuration parameters and must infer hardware consequences from data. We introduce a hardware-aware representation for CUTLASS kernel...
-
26
Output-aware Residual Stream Pruning for Large Language Models
Chayne Thrash, Kevin Chen, Soheil Kolouri
cs.LG
Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directions as equally important, ignoring the sensitivity of downstream layers. We introduce a sensitivity-aware approach to residual-stream pruning that directly accounts for this...
-
27
From Experience to Expertise: Adoption-Aware Memory Learning for Data-Scarce NPU Kernel Synthesis
Longxiao Fan, Tao Zhang, Han Yan, Jiajun Li, Mingcong Song, Guoping Long, Hongjie Si, Weiwei Sun
cs.LG
High-performance kernels underpin efficient accelerator execution but require expert tuning and lengthy manual optimization cycles. LLM coding agents promise automation, yet their CUDA knowledge transfers poorly to data-scarce domain-specific architectures (DSAs) such as NPUs, whose execution models and memory hierarchies differ substantially from those of GPUs. To address this transfer gap, post-training methods adapt LLMs to NPU programming...
-
28
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
cs.LG · stat.ML
Diffusion models have revolutionized generative modeling for continuous data through the gradual refinement of a belief state. This iterative refinement has not yet carried over to discrete diffusion models, which discard uncertainty at intermediate steps through categorical sampling (information collapse). We propose Simplex Diffusion Models (SDMs), a framework that lifts the diffusion process to the probability simplex to represent beliefs...
-
29
Graph World Models for Constrained Epidemic Policy Planning
Yiqi Su, Rashed Shelim, Lingyi Wang, Walid Saad, Naren Ramakrishnan
cs.LG · cs.AI
Epidemic policy planning often requires coordination between geographical regions, taking into account mobility-driven spillovers and how to make use of limited resources. Existing methods either lack action-conditioned models of coupled dynamics or cannot guarantee per-period feasibility. We present EpiMind, a graph world model framework for constrained epidemic policy planning across regions. A graph-factored recurrent state-space model...
-
30
Learning the Robustness Mechanism with Bilevel Optimization
Yiyang Shen, Qihang Lin, Weiran Wang
cs.LG
We propose a distributionally robust learning framework where parameters defining the robustness mechanism are learned from held-out data instead of extensively tuned. Using bilevel optimization with both upper and lower level minimax problems, we create two instances of our framework to tackle setups with and without group labels in the training set. Theoretically, we provide sample complexity analysis for our robustness mechanism learning...
-
31
Optimal Networks for Agentic Information Aggregation
MohammadHossein Bateni, Zahra Hadizadeh, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz, Shayan Taherijam
cs.LG · cs.GT · econ.TH
We study information aggregation in the networked learning model introduced by Kearns, Roth, and Ryu (SODA 2026). There is a fixed distribution over $d$ features and a common label. Agents learn in topological order on a directed acyclic graph. Each observes a subset of the features and its parents' predictions, fits a linear predictor to minimize mean squared error, and passes only its prediction forward. The global predictor is the best...
-
32
Let the Neurons Die: Exploiting ReLU-Induced Model Degradation
Kexin Li, Wenjun Qiu, Joshua Abraham, Aditi Maheshwari, David Lie
cs.LG · cs.AI
Rectified linear unit (ReLU) networks can suffer from dying neurons, where units with persistently negative pre-activations produce zero outputs, blocking gradients through their activations. To exploit this failure mode, we present three training-time availability attacks based on data ordering and poisoning. We begin with the basic dynamic data-ordering attack (DOA), which greedily constructs a training prefix by selecting the next example...
-
33
GeoGAE: Scalable Graph-Level Autoencoding via Hyperball Cloud Representations
Radosław Nowak, Anna Bielawska, Bogusz Stefańczyk, Maciej Sanocki, Paweł Wawrzyński
cs.LG
Embedding structured objects into Euclidean spaces has enabled a wide range of successful machine learning applications. Such objects include words, documents, image patches, time series, and graph nodes. In contrast, embedding entire graphs remains a challenging problem. Existing methods either sustain the original order of the graph nodes or match the output nodes to the input ones, both of which create scalability issues. In this work, we...
-
34
Deep Epistemic Value Functions for Optimistic Exploration
Leander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas Krause
cs.LG
Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central challenge is therefore to scale these ideas robustly. We conduct a systematic empirical study of how epistemic uncertainty is represented, propagated,...
-
35
Reward-Aligned Reweighting for On-Policy Distillation
Haofeng Xu, Junwei Su, Lansong Diao, Wenchao Zhou, Chuan Wu
cs.LG
On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however, depends on how the student completes the subsequent reasoning. This mismatch can cause imitation to suppress viable student...
-
36
One Proposal for Every Margin: Zero-Shot Amortized Sequential Importance Sampling for Binary Matrices
Ruishuo Chen, Weijia Li, Xun Wang, Yu Chen, Leheng Cai, Longbo Huang
cs.LG · stat.CO
In ecology, psychometrics, and the analysis of social and financial networks, binary matrices are often analyzed conditional on their observed row and column sums, which restricts the problem to a finite sample space of matrices with the same margins. Two fundamental problems are to count this space and to sample uniformly from it. Sequential importance sampling (SIS) addresses both with independent weighted samples and an unbiased count...
-
37
Improving Generative Model Self-Training with Geometrically Modified Outputs
Patrick Batsell, Thomas Walker, Richard Baraniuk
cs.LG · cs.AI
Self-training generative models - the continued improvement of a model using its own outputs - is becoming increasingly important as high-quality training data becomes scarce. However, naively finetuning on model-generated samples leads to degradation through model collapse and the model autophagy disorder. Negative-guidance self-training methods turn this degradation into a useful signal, using a model finetuned on its own outputs to guide...
-
38
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao, Xiaoyun Wang, Taylor W. Killian, Weitong Zhang
cs.LG
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through...
-
39
Structured Latent Modeling for Supervised Multimodal Information Decomposition
Wanting Huang, Sanvesh Srivastava, Weiran Wang
cs.LG
Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these target-relative contributions within learned continuous representations. We introduce a framework that applies contrastive or...
-
40
Physics-Guided Conditional Diffusion Model for Rare Event Synthesis and Diagnosis for the Water-Gas Shift Reaction
Md Abrar Rafid Siddique, Bibek Aryal, Qiugang Lu
cs.LG
As the world moves towards sustainable energy sources, hydrogen (H2) can be treated as an eco-friendly alternative to fossil fuels due to its high energy density and zero carbon emissions. The water-gas shift (WGS) reaction is a widely used industrial process for hydrogen production by converting carbon monoxide and steam into hydrogen and carbon dioxide. However, occurrences like severe fouling, catalyst deterioration, and thermal runaway...
-
41
Universal Approximation of Measure-to-Measure Operators by Pushforwards
Takashi Furuya, Nicholas H. Nelsen, Frank Cole
cs.LG · math.NA · stat.ML
Many learning tasks map an input distribution to an output distribution. A natural way to model such an operator is to transform each input sample using a continuous function that may depend on the entire input distribution, and then take the distribution of the transformed samples. This defines a measure-dependent pushforward model and includes measure-theoretic formulations of transformers. We ask when such models can approximate arbitrary...
-
42
An analysis of Mirror-Descent Soft Actor-Critic
Denis Zorba, Michal Valko
cs.LG · math.OC
Soft Actor-Critic (SAC) is widely used for entropy-regularised reinforcement learning with continuous action spaces, and practical implementations perform only a few actor steps towards an evolving target. In this work, we prove convergence guarantees when the target policy arises from policy mirror descent and compare it with the classical Gibbs target. We derive sufficient conditions for the strong convexity and smoothness of the actor...
-
43
Tetra: Serving Leech-Lattice Quantized LLMs at 2.7 Bits per Parameter
Pier-Jean Malandrino
cs.LG
Leech-lattice quantization gives good quality at two bits per weight, but its codebooks hold more than 10^14 points, too many for a lookup table. Our earlier kernel expanded the codes at load time and read 4.804 bits per weight from GPU memory for 2 bits of code. We present Tetra, a new codebook on the same lattice. A 24-weight block still takes 48 bits, most of which index a 64-state trellis of the Golay code and one shared 16 KiB table. The...
-
44
CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations
Yusuf Kesmen, Aniruddha Mukherjee, Yena Chang, David Sasu, Trevor Brokowski, Alexandra V. Kulinkina, Kristina...
cs.LG · cs.AI · cs.CL
Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label diagnosis. We introduce CLIMB, a benchmark in which a doctor model interviews a simulated patient to recover a ground truth set of co-occurring clinical conditions. Cases are...
-
45
Manifold-Stable Flow Matching
Amirhossein Nazerian, Ali Pezeshki, Jianguo Zhao
cs.LG · cs.RO · eess.SY
Flow matching (FM) learns generative dynamics through velocity regression. Geometric FM variants commonly assume a prior supported on the data manifold, requiring geometric knowledge that is often unavailable. Without such knowledge, low regression error alone does not guarantee manifold adherence. Adherence keeps generated samples within valid configurations and is empirically associated with better task performance. We introduce...
-
46
NeuronSifter: Intervention Planning in CNS Microenvironments
Haowei Xu, Wanyi Fu, Hongbin Han, Zhaoheng Xie
cs.LG · q-bio.NC
Prioritizing central nervous system (CNS) interventions requires predicting how a dose, route, and schedule act on a partially observed microenvironment, then choosing the measurement that would change the decision. Action-conditioned predictors reduce a regimen to an identity token or a scalar exposure, discarding where and when the target is engaged; handing a point estimate to a separate planner then discards the joint uncertainty that...
-
47
Riccati State Space Models: Non-iterative Parallelization for Nonlinear Sequence Modeling
Mónika Farsang, Ramin Hasani, Daniela Rus, Radu Grosu
cs.LG · cs.AI
State space models (SSMs) achieve efficient sequence processing because their affine state updates are closed under composition and can therefore be evaluated with an associative parallel scan. Nonlinear recurrent models can provide richer, state-dependent dynamics, but generally lose this compositional structure: parallel evaluation then requires iterative methods that repeatedly linearize and scan the recurrence. We ask, what...
-
48
SOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local Learning
Bojian Yin, Shurong Wang, Yuqi Pan, Guoqi Li
cs.LG · cs.DC
Large language models are trained with backpropagation, whose global gradient coordinates all layers but forces each to hold its activations and wait for the gradient to pass back through every deeper layer. Conventional local learning removes this update locking by training each module to predict the target through its own readout, but has not scaled to billion-parameter pretraining. We identify these private readouts as a key weakness,...
-
49
ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
Yihang Chen, Yuanhao Ban, Cho-Jui Hsieh
cs.LG · cs.AI
Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate...
-
50
LLMs are General Asynchronous Agents
George Yakushev, Denis Mazur, Vladimir Bartenev, Vyacheslav Zhdanovskiy, Timofey Byzov, Vladimir Kaurkin, Vadim Pastushenko
cs.LG · cs.CL
Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous...
-
51
Frontier Learning: Training LLM Reasoners at the Edge of Capability
Robin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic, Aurelien Lucchi
cs.LG · cs.AI · cs.CL
Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal arises only when policy rollouts mix successes and failures, causing the useful portion of any fixed pool to quickly become stale...
-
52
Persistent Partners Raise Prices Among Learning Agents
Paul-Peter Arslan, Yubin Kim, Xiao Xiao
cs.LG · cs.MA
When pricing agents meet repeatedly on a platform, the platform decides who faces whom. We ask whether that choice moves the prices the agents learn, and whether a rise comes with learned punishment. In a pre-registered randomised experiment in the Bertrand duopoly of Calvano et al., each agent's price is set by a tabular Q-learning module, not by the small language model attached to it, and we randomise whether each agent keeps its partner,...
-
53
The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics
Yihe Zhou, Tongtian Zhu, Yingxiao Huo, Satya Prakash Dash, Can Wang, Samuel Kaski, Mingfei Sun
cs.LG · cs.AI
Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-$β$ regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically,...
-
54
Inductive Feedback for Mixed-Policy Distillation
Amir Moeini, Huaijiang Zhu, Daniel Havir, Shangtong Zhang
cs.LG
Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition the teacher in on-policy distillation, which trains the student to match the teacher's predictions on student-generated rollouts. However, this approach can transfer teacher preferences...
-
55
Identifying Neural Source Dynamics from Unknown Local Interventions
Ayana Mussabayeva, Jiaqi Sun, Anuar Aimoldin, Olivier Oullier, Kun Zhang
cs.LG
Electroencephalography (EEG) records mixtures of brain-source activity. Even with a known anatomical forward model, experiments that excite only part of the source-state space leave the dynamics unidentified, and repetition cannot resolve the ambiguity. We show that unknown local mechanism changes can supply the missing information. We consider linear dynamics among fixed anatomical sources with known source-state initialization patterns....
-
56
First Learn, Then Memorize: The Spectral Bias of Diffusion Models
Raphaël Urfin, Tony Bonnaire, Giulio Biroli, Marc Mézard
cs.LG · cond-mat.dis-nn
Diffusion models trained on a finite dataset first learn to generate novel, high-quality samples and only much later collapse onto their training set. We identify the mechanism behind this separation of timescales and the object that probes it. The training dynamics of the score function are governed---exactly, and at any width---by the Gram matrix of the Neural Tangent Kernel (NTK) evaluated on the noisy training data, so the timescales of...
-
57
Do Temporal Link Predictors Need Learned Memory? A Smoothed-Count Baseline with a Handful of Parameters
Lisi Qarkaxhija, Ingo Scholtes
cs.LG · cs.AI
Many temporal link predictors summarize past interactions through learned node representations. We examine whether simple counts of recurring interaction patterns can provide competitive predictions without learning these representations. We propose a temporal link predictor based on statistical language modelling. It pools transition and co-occurrence counts across sources to predict links that a source has never formed. We smooth sparse...
-
58
d-OPD: Future-Aware On-Policy Distillation for Block Diffusion Language Models
Ruitao Liu, Qinghao Hu, Song Han
cs.LG · cs.AI
Large language models (LLMs) typically generate text autoregressively (AR), predicting one token at a time. Block diffusion language models (dLLMs) instead generate blocks sequentially while denoising multiple tokens in parallel within each block, offering a promising way to accelerate generation. Rather than training such models from scratch, recent work adapts strong pretrained AR models into block dLLMs through distillation. On-policy...
-
59
Fiona: Accelerating FHE Inference with Packing-Aware Ternary Weights
Yiteng Peng, Zhibo Liu, Dongwei Xiao, Shuai Wang
cs.LG
Fully homomorphic encryption (FHE) enables neural network inference directly on encrypted inputs, but it remains orders of magnitude slower than plaintext in- ference. Applying the server's plaintext weights to encrypted activations involves plaintext-ciphertext multiplications (PMult) and accounts for more than half of inference time in recent systems. Ternary quantization can replace these multipli- cations with additions and subtractions,...
-
60
Interference Beyond Geometry in Concept Extraction
Valérie Costa, Bahareh Tolooshams
cs.LG
Interference is commonly treated as geometric overlap between learned features. We introduce effective interference, which combines feature geometry and code statistics to capture realized interactions, distinguishing constructive from destructive interference and frequent weak interactions from rare strong ones. Under local fixed-support assumptions, we characterize how architectural constraints shape interference through four mechanisms:...
-
61
Quasi Linear Kernel Attention with Infinite Capacity
Nicolaj Rux, Johannes Hertrich, Sebastian Neumayer
cs.LG · math.NA
The evaluation cost of transformers with softmax attention scales quadratically with sequence length. Kernel attention addresses this by replacing softmax with a more general kernel function. In this paper, we aim to identify kernels that retain the expressivity of attention while enabling quasi linear computation. To quantify expressivity, we introduce a capacity for each kernel, measuring the maximum sequence length for which the attention...
-
62
From Data to Program: Fast & Direct Generative Program Inference from Empirical Data
Simon Klüttermann, Xueying Ding, Leman Akoglu
cs.LG · cs.AI
Estimating probability densities from a finite set of samples typically requires dataset-specific model fitting. We introduce PRODiGI, a pretrained data-to-program model that infers an explicit, executable generative program in a single forward pass. Pretrained on synthetic datasets paired with their ground-truth programs, PRODiGI accommodates diverse generative families and data dimensionalities through template prediction and...
-
63
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Xin Li, Hao Jiang, Xin Gao, Annan Wang, Yuchen Xie, Jinghao Guo, Xingwei Qu, Yichi Zhang, Chau Yuen
cs.LG
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides...
-
64
NeuronDiscover: Agent-in-Twin for Mechanistic Discovery in Neuronal Microenvironments with World Action Models
Haowei Xu, Wanyi Fu, Hongbin Han, Zhaoheng Xie
cs.LG · q-bio.NC
Mechanistic discovery in neuronal microenvironments requires interventions and measurements that separate competing explanations of solute transport and neuronal response. Predictive accuracy cannot settle the question: a real mechanistic change and an error in the computational twin leave the same signature in sparse observations. We formalize this twin confounding and reason over a joint mechanism--discrepancy belief, designing experiments...
-
65
Large Language Models for Automated Cross-Domain Machine Learning Task Type Identification: A Benchmark Dataset and Evaluation
Petros Tsialis, Steffen Limmer, Tobias Rodemann, Martin Heckmann
cs.LG · cs.AI
Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level information when only the target feature is provided by the user. Together with our LLM-based system we also release an annotated benchmark comprising 625...
-
66
Scalable In-Context Reinforcement Learning with Recurrent Algorithm Distillation
Yuanqing Ma, Zhenrui Zheng, Chenjun Xiao
cs.LG
Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur prohibitive memory costs and limit scalability in complex, long-horizon tasks. To address this bottleneck, we propose Recurrent Algorithm Distillation (RAD). RAD employs a...
-
67
Weighting Schedules Govern What and When Score-Based Generative Models Learn from Multimodal Data
Jérémie Klinger, Raphaël Urfin, Giulio Biroli, Marylou Gabrié
cs.LG · cond-mat.dis-nn
Score-based generative models generate new samples by integrating a time-dependent drift that carries Gaussian noise onto the target distribution. In practice this drift is modeled by a neural network, trained on a loss integrated over time $t$ with a weighting schedule $w(t)$. Along the backward dynamics, and for multi-modal distributions, trajectories commit to modes of the target within a narrow time window, the \textit{speciation time}....
-
68
Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents
Tong Zhang, Zhou Liu, Yihao Liu, Jiahua Bao, Xuchen Li, Honglin Lin, Tao Cheng, Zhihan Yu, Kai Tang, Xiaoxi Jiang,...
cs.LG · cs.AI
On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linking larger gaps to a greater need for correction. Yet, our empirical analysis reveals a supervision-benefit mismatch: large gaps can be benign, while small gaps can be outcome-critical....
-
69
Collaborative Principle Evolution via Evidence Transfer for Scientific Discovery
Yingming Pu, Hongyu Chen, Tao Lin
cs.LG
Large Language Model (LLM)-based agents promise to automate scientific discovery, yet exploring the vast hypothesis space remains costly. Existing principle-evolution methods accelerate this loop, but operate sequentially, which caps exploration breadth and wastes wall-clock time on challenging problems. To address this, we formulate collaborative scientific discovery as evidence transfer between parallel principle-evolution branches. We...
-
70
When Should a Satellite Estimate Be Changed? Stress-Testing Neural Corrections for Evapotranspiration
Marco Trotta
cs.LG
Neural residuals can improve satellite evapotranspiration (ET) estimates, but selectors must predict when a correction helps and reject unsupported inputs. We evaluate ten-member models on 16,366 flux-tower observations from 151 stations paired with OpenET, across nine rolling years and five spatial folds. At one held-out station, Gain accepted corrections on all 32 physically invalid records: it predicted a mean benefit of 0.83 mm/day, but...
-
71
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
Arman Bolatov, Artem Riabinin, Nikita Kornilov, Andrey Veprikov, Samuel Horváth, Martin Takáč, Aleksandr Beznosikov
cs.LG
Pretraining a language model takes enormous compute, and the right optimizer can save a good part of it. Muon's spectral step gives a stronger direction than a sign step, but it is expensive. Every step runs Newton-Schulz iterations on the full matrix and, in distributed training, an extra all-reduce. Sign steps, as in Lion and Signum, are cheap and stay local to each device. We propose LionMuon, which takes one Muon step every $P$ iterations...
-
72
Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment
Shunchang Liu, Lukas Fluri, Xin Chen, Francesco Croce
cs.LG · cs.CV
Modern AI models are aligned through post-training to adapt them to downstream tasks. Recent work shows that fine-tuning language models on narrow tasks can induce emergent misalignment (EM), causing broadly harmful behaviors beyond the training task. However, EM has been studied almost entirely in text-only tasks, leaving its manifestation in multimodal models unclear. In this paper, we define and analyze EM in the context of vision-language...
-
73
$λ$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning
Berker Demirel, Clémentine Dominé, Valentino Maiorca, Marco Fumero, Marco Mondelli, Francesco Locatello
cs.LG · cs.CV
Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use the backbone representation before the projector. We find that this mismatch does not necessarily prevent dimensional collapse in the backbone, which can retain low effective rank and...
-
74
Multi-Attractor GNNs: Set-Valued Expressivity Beyond Unique Equilibria
Jialin Liu
cs.LG · cs.AI
Recurrent and equilibrium graph neural networks (GNNs) often enforce a unique fixed point or use one training target per graph. Yet many combinatorial and scientific problems admit multiple valid solutions, with no preferred one. A designated target can then impose an arbitrary selection rule. For tasks invariant to node relabeling, a symmetric graph may have a symmetric solution set but no symmetric solution. We show that multiple equilibria...
-
75
eval-unlearn: Benchmarking unlearning in Text-to-Image Diffusion Models
Mansi, Nikhil Raghavan, Zixia Huang, Kai Sheng Ong, Ji Shen Lim, Brandon Siao Xiang Ling, Francesco Leofante
cs.LG · cs.AI · cs.CV
The rising number of concept unlearning techniques for text-to-image (T2I) diffusion models has produced a fragmented evaluation landscape. Methods are assessed under heterogeneous experimental conditions making principled cross-method comparison difficult. We present eval-unlearn, an open-source Python library providing a unified, reproducible benchmarking framework for concept unlearning in T2I Diffusion models. eval-unlearn integrates...
-
76
SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards
Yingchao Yu, Pengfei Sun, Wenxuan Pan, Wei Chen, Yitian Hong, Kuangrong Hao, Yaochu Jin
cs.LG
Reinforcement learning (RL) with sparse rewards is challenging because delayed outcomes provide little guidance about which intermediate computations caused success or failure. We argue that reliable credit assignment requires policy dynamics that preserve and expose credit-relevant information over time, a role we formalize as Temporal Credit Carriers (TCCs) and that spiking neural networks (SNNs) naturally fulfill through graded membrane...
-
77
Latency and accuracy tradeoffs in Spiking Neural Networks
Zhanglu Yan, Zixuan Zhu, Kaiwen Tang, Yuyang Cai, Qianhui Liu, Weng-Fai Wong
cs.LG
Spiking neural networks are attractive for low-power speech command recognition, yet their latency has received far less attention than their energy efficiency, and their multi-timestep execution is widely assumed to make them slower than quantized neural networks. This paper challenges the assumption that more local timesteps necessarily imply higher network latency. By overlapping computation across adjacent layers at the timestep level,...
-
78
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
Julianna Piskorz, Antonin Berthon, Mihaela van der Schaar
cs.LG
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study the effect of rollout policy in a controlled strong-to-weak distillation setting, by independently varying rollout...
-
79
AIM-ZO: Activation-Informed Subspace Maintenance for Zeroth-Order LLM Fine-Tuning
Yue Xie, Zhi Zheng, Yunpeng Ba, Xuyang Wu, Xialiang Tong, Zhichao Lu, Tao Zhong, Zhenkun Wang
cs.LG
Zeroth-order (ZO) optimization offers a memory-efficient alternative for LLM fine-tuning by estimating updates only from forward evaluations of perturbed parameters, without backpropagation or activation storage. However, in billion-parameter LLMs, isotropic perturbations often waste many forward evaluations on weakly informative directions. To make these evaluations more informative, existing ZO methods restrict perturbations to...
-
80
Disentangling Lung-Cancer CT/LDCT AI: A Systematic Evidence Map of Clinical Tasks, Evidence Chains, and Translational Gaps
Surajit Das
cs.LG
Artificial-intelligence studies using computed tomography (CT) for lung cancer are often broadly labelled "prediction" despite addressing clinically distinct tasks. We systematically mapped CT/low-dose CT (LDCT)-centered lung-cancer AI using five-database retrieval, full-text eligibility assessment, role-aware modality/omics extraction, clinical-task classification, and a Multi-Tier Evidence Graph (MTEG). The final corpus comprised 293...
-
81
Long-Horizon Scaling: How Model Capabilities Shape the Returns to Computation
Haoyu Zheng, Zhengyu Chen, Huaisheng Zhu, Ruishan Fang, Teng Xiao, Yiwei Li, Jingang Wang, Wenqiao Zhang
cs.LG
Long-horizon agents improve solutions through sustained interaction, execution, and task feedback. Scaling studies relate performance to resources and capabilities, yet how existing capabilities shape returns to extended interaction remains less understood. To address this gap, we analyze AutoLab and EdgeBench, two long-horizon benchmarks. We find that starting performance and subsequent growth are associated with different capabilities:...
-
82
TANGO: Watermarking Masked Diffusion Language Models in Token Pairs
Kasra Arabi, Nir Weinberger, Micah Goldblum, Niv Cohen
cs.LG · cs.CL · cs.CR
Masked-diffusion language models fill in masked positions in parallel and in no fixed order. Most practical text watermarks assume left-to-right generation. They key each token to the tokens before it, and in a diffusion model those tokens may still be masked. A fixed green list needs no such context, but it favors the same tokens at every position, so these tokens appear more often in watermarked text. An attacker who compares token...
-
83
Temporal Heterogeneous Graph Pretraining for Relational Deep Learning
Yixin Peng, Er Jin, Diego Collarana, Stefan Decker
cs.LG
Relational deep learning models database rows and foreign-key links as a heterogeneous graph for prediction from record attributes and relational context. These graphs contain two distinct temporal signals: record age changes with the prediction cutoff, while intervals between observed records remain fixed. Prior work often treats time as a single signal or studies temporal representation and pretraining separately. We investigate how...
-
84
Adversarial Consistency-Guided Representation Learning for Multi-view Clustering
Yuchen Lin, Kunpeng Xu, Ying Fang, Lifei Chen
cs.LG
Multi-view clustering aims to capture cross-view consistency while exploiting view-specific information. However, shared representations learned to capture cross-view consistency may still retain view-identifying information, potentially compromising the consistency of cross-view clustering structures. To address this issue, we propose ACGRL, an adversarial consistency-guided representation learning framework for multi-view clustering. ACGRL...
This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.