cs.LG · 2026-10-05 · No. 134
Machine Learning, 2026-10-05.
75 new papers in cs.LG. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
75 entries-
01
What Should World Models Forget? Stratified Retention for Continual Adaptation
Nishit Anand, Ramani Duraiswami, Dinesh Manocha
cs.LG · cs.AI · cs.CV · eess.IV · eess.SP
Continual learning treats degradation on previously seen data as evidence of failure, a convention inherited from settings with a stationary prediction target, where a correct label remains correct indefinitely. World models do not satisfy this condition. Their prediction target is the environment, which changes, so knowledge that was accurate when acquired may later become false, and discarding it is required behavior rather than a defect....
-
02
RNADyn: A Benchmark for Generating and Understanding RNA Dynamics
Yiming Huang, Lennart Bastian, Hanqun Cao, Luis Vollmers, Tolga Birdal
cs.LG · physics.bio-ph · q-bio.BM
Ribonucleic acid (RNA) functions through conformational changes that are not fully captured by static structures. However, large-scale standardized RNA dynamics data remain limited, and existing approaches typically treat trajectory generation and dynamics understanding as separate objectives. Here, we introduce RNADynBench, a standardized RNA molecular dynamics (MD) benchmark with 2585 quality-controlled 100-ns all-atom trajectories and...
-
03
LESSER: Post-Training Data Selection with Output-Layer Gradients
Lyuxin David Zhang, Eric Wong, Surbhi Goel, Anton Xue
cs.LG
The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question:...
-
04
Simulation-Free Learning of Population Dynamics with Wasserstein Lagrangian Residuals
Fedor Sergeev, Markus Heinonen, Daniel Waxman, Tim Cooijmans, Ricardo Baptista, Dmitry Batenkov, Eli Bingham
cs.LG · stat.ME · stat.ML
The dynamics of cells, organisms, and fluids are often modeled as probability distributions evolving over time. Reconstructing and extrapolating this evolution from unpaired snapshots requires assumptions about the underlying process. Wasserstein gradient flows are a common choice, but they cannot describe conservative or periodic dynamics. Lagrangian mechanics in Wasserstein space covers both, but existing methods for learning it are...
-
05
Planning to Learn
Ian Osband
cs.LG · math.OC · stat.ML
Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has none of them. A classifier is a policy whose expected reward, its \emph{expected accuracy}, is the probability it assigns to the correct label, and because that label is known, the policy gradient is exact and smooth. Yet exact...
-
06
Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao, Se-Young Yun, Rahul G. Krishnan
cs.LG · cs.CL
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on...
-
07
Forecasting from Counterfactual Simulator Rollouts: A Sim2Real Evaluation
Angel Wang, Dominique Perrault-Joncas, Alvaro Maggiar, Dean Foster, Carson Eisenach
cs.LG
Deploying a new decision policy creates a cold-start problem for prediction models whose targets depend on the policy's actions: historical observations reflect earlier policies, while real observations under the new policy are not yet available. Simulation offers a way to address this gap by rolling out the target policy across counterfactual scenarios and using the resulting trajectories to learn how the system responds to those controls....
-
08
When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity
Mayand Gulati, Kerong Wang, WeiChen Au
cs.LG · stat.ML
Stationarity rewards memory, but after a change the same history can mislead. We ask when forgetting should be permitted. E-process-authorized Thompson sampling (e-ATS) gives each arm full-history and discounted Beta states. An anytime-valid e-process first authorizes the discounted state, then a reversible relevance score controls its influence. Before authorization, e-ATS exactly follows optimistic Thompson sampling (OTS). Under a...
-
09
On the Convergence of Success Conditioning for Policy Optimization
Matthew Brun, Xu Andy Sun
cs.LG · math.OC
Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a...
-
10
IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models
Vladislav Gromadskii, David Li, Samson Gourevitch, Yazid Janati, Eric Moulines, Maxim Panov, Alexander Korotin
cs.LG
Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation...
-
11
Broken scale symmetries in undercomplete linear autoencoders
Farhad Pashakhanloo, Jacob A. Zavatone-Veth
cs.LG · q-bio.NC · stat.ML
Neural network loss landscapes have many symmetries, which are preserved by gradient flow but broken by finite-stepsize stochastic gradient descent (SGD). A canonical example of such a symmetry is scale in homogeneous networks: one can scale up the parameters in one layer and down in the next without changing the network output. Previous work has documented cases in which SGD breaks this symmetry in favor of balancing gradient noise or...
-
12
UniIntervene++: An Adaptive Intervention Agent for Efficient Real-World Reinforcement Learning
Yudong Lin, Haoyuan Deng, Zhuoxuan Yuan, Zaijia Yang, Yuanjiang Xue, Ziwei Wang
cs.LG · cs.RO
Online reinforcement learning (RL) enables robot policies to improve through physical interaction, but the assistance they require changes as their competence evolves. Existing intervention strategies based on offline estimates or fixed decision rules can therefore become mismatched to the current policy. To address this, we propose UniIntervene++, an adaptive intervention agent that learns to allocate control between autonomous execution and...
-
13
Mastering Atari 2600 Games with Discovered Options
Erik M. Lintunen, Marlos C. Machado
cs.LG
Temporal abstractions, often instantiated as options, have long been regarded as a mechanism for accelerating credit assignment, facilitating exploration, and enabling generalisation in reinforcement learning (RL). However, developing general option discovery methods that are effective in large-scale, high-dimensional domains remains a fundamental challenge. Existing option discovery methods are either confined to relatively simple domains,...
-
14
A Path Integral Surrogate for Multi-Step Gradient Inversion in Federated Learning
Agnivo Ghosh, Saumik Bhattacharya
cs.LG · cs.DC
Federated learning lets many clients train a shared model together without ever sending their private data to a central server. Each client shares only a model update, and this update should reveal far less about the client than its raw training examples would. This premise is what protects the privacy of the clients. Gradient inversion attacks challenge it directly by trying to reconstruct a client's private input images from the single...
-
15
Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain
Antoine Collas, Louis Jalouzot, Géraud Ilinca, Corentin Caris, Romain Valabrègue, Ahmed Hassayoune, David Goncalves,...
cs.LG · cs.AI
Cephalonauts One is a whole-brain 3 Tesla (3T) functional magnetic resonance imaging (fMRI) dataset recorded while subjects listened to audio podcasts. Three healthy subjects underwent multiple scanning sessions, each consisting of five 15-minute runs, while listening to podcasts in their native language. With 30 hours of fMRI data per subject, the current release is the deepest available fMRI dataset using naturalistic speech stimuli. The...
-
16
Get a GRIP, this will be a long TRIP: A Quantifiable Long-Range Framework for Verifying Over-squashing
Ferran Hernandez Caralt, Simon Heilig, Adrián Bazaga, Asja Fischer, Moshe Eliasof, Pietro Liò
cs.LG
Empirical claims about the connection between over-squashing and long-range interactions in GNNs, can only be trusted if the benchmarks used to validate them genuinely require long-range interactions. The de-facto standard, the Long Range Graph Benchmark, has been repeatedly shown to be saturated by tuned short-range models, with existing synthetic alternatives being tied to specific topologies. As such, there is a lack of principled...
-
17
Objects Without Morphisms: What LLMs for Mathematics Do Not Represent
Yanli Wang, Suijin Wang, Xiaopeng Yuan, Haohan Wang
cs.LG
Large language models (LLMs) have reached expert-level performance on competition mathematics largely through the volume of search placed around them: candidate solutions are sampled in quantity and retained only when an external criterion accepts them. Such a procedure improves the outcome that survives it while leaving untouched what the model represents. We examine that question where no external criterion exists: translating statements...
-
18
ZeroMAG: Zero-Shot Multimodal Adapter Generation for Plug-and-Play EEG Foundation Models
Yubo Wang, Jingying Ma, Xinliang Zhou, Yangxuan Zhou, Jiquan Wang, Sha Zhao, Yiyuan Yang, Yi Ding, Ziyu Jia, Chenyu...
cs.LG
EEG foundation models (EFMs) capture reusable knowledge from large-scale EEG data, while many EEG recordings also include companion physiological signals that provide complementary information beyond the EEG-only interface. The challenge is to preserve this pretrained knowledge while extending the EFM to heterogeneous multimodal recordings through an adaptation inferred from unlabeled target data. We introduce ZeroMAG, a zero-shot multimodal...
-
19
Divergence controls entropy in distillation
Nicolas Zucchet, Scott W. Linderman
cs.LG · cs.CL
Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, studying how the entropy of the student depends on the data and the divergence that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we...
-
20
Beyond Trained Models: Compiling GNNs for a Sound Explainer Benchmark
Steve Azzolin, Francesco Paolo Nerini, Stefano Teso, Francesco Bonchi, Bruno Lepri, André Panisson, Andrea Passerini
cs.LG · cs.AI
Explainers for Graph Neural Networks (GNNs) are commonly evaluated by their plausibility, i.e., how well their explanations recover a predefined ground truth, such as a motif planted in the data. This protocol implicitly assumes that a GNN trained on such data relies on the intended motif. Although prior work has questioned this assumption, plausibility remains widespread. First, we show that the assumption is violated on several widely used...
-
21
An Automated and Reproducible Workflow for Crack Identification and Damage Assessment of Fusion Materials
Rinkle Juneja, Viktor Reshniak, Richard K. Archibald, John W. Duggan, Gregory R. Watson, Cory D. Hauck, Gary M. Staebler
cs.LG · cond-mat.mtrl-sci
Post-exposure microscopy is central to qualification of fusion materials. However, manual analysis does not scale to the volume, heterogeneity, and multiresolution character of modern fusion-materials campaigns. To address this challenge, we present a reproducible workflow, implemented in the Galaxy scientific workflow environment, for automated crack identification and quantitative damage assessment from scanning electron microscopy images....
-
22
Getting Your Guidance Weights Right in diffusion and flow-matching posterior sampling
Liam Moroy, Jean-François Giovannelli, Yoann Altmann, Steve McLaughlin, Frédéric Champagnat, Guillaume Bourmaud
cs.LG
Training-free posterior sampling methods, also known as Plug-and-Play methods, leverage pretrained unconditional diffusion or flow-matching models to solve inverse problems. Most existing approaches rely on guidance weights to balance, at each time step, prior information from the unconditional score or velocity network with measurement consistency, yet the tuning of these weights is often not discussed and is largely left to heuristics. We...
-
23
Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation
Md Sazid Uddin, Md. Khairul Alam Mazumder, M. F. Mridha
cs.LG · cs.AI · cs.CR
Mechanistic edits (ablations, weight edits, activation steering) are the standard tools for unlearning a harmful capability from a neural network while preserving useful ones. Current approaches validate their effects only by testing, which can never cover an entire continuous region of inputs. Prior work at the interpretability-verification boundary certifies descriptions of a model: what a circuit computes, or whether it faithfully explains...
-
24
Below what training size do deep tabular generators stop beating trivial baselines? A preregistered benchmark on a size ladder of clinical and standard datasets
Shivam Shrivastava
cs.LG · stat.ML
Deep tabular generative models are benchmarked on datasets with tens of thousands of rows; clinical datasets have hundreds. We preregistered and ran a size-ladder benchmark to find where the two regimes diverge: 8 public datasets subsampled from 200 to 20,000 training rows, seven generators (independent marginals, Gaussian copula, SMOTE, unconditional SMOTE, CTGAN, TVAE, TabDDPM) with a fixed 20-trial tuning budget and 5 evaluation seeds,...
-
25
Most-Recent Anchoring with Recurrent Ordering for Time Series Forecasting
Jung Min Choi, Ngoc Son Le, Ibram Abdelmalak, Vijaya Krishna Yalavarthi, Lars Schmidt-Thieme
cs.LG
Long-term forecasting models commonly process all patches in a look-back window using the same fixed stack. Older contextual patches and recent evidence therefore receive the same computational depth. Yet the information closest to the forecast and the more distant context do not contribute equally. Uniform processing leaves this distinction unexpressed in the architecture. We propose MARO, a Most-Recent Anchoring with Recurrent Ordering...
-
26
Dual-Context Analog Retrieval for Time Series Forecasting
Jung Min Choi, Ngoc Son Le, Ibram Abdelmalak, Vijaya Krishna Yalavarthi, Lars Schmidt-Thieme
cs.LG
Most long-term time-series forecasting models map the look-back window directly to the full horizon in a single pass. While efficient, this design does not explicitly identify which historical states are most relevant to different future segments or exploit what followed those states. Analog forecasting addresses this by retrieving past states similar to the present and using their observed continuations, but single nearest matches can be...
-
27
Metropolis-Hastings Dominates Importance Resampling for Policy Composition
Alexey Kurennoy, Ramil Yarullin, Fergal Reid
cs.LG
Post-training a large language model (LLM) often requires exploring trade-offs between multiple rewards, but retraining for each trade-off is expensive. Decoding-time policy composition allows these trade-offs to be adjusted by combining reward-specific policies at inference time. This composition targets a weighted product of the policies' probabilities over complete responses, but standard implementations combine their next-token...
-
28
Single or Multiple Policies for Phase-Structured Reinforcement Learning?
Guilhem Loussouarn, Nancy Nayak, Kin K. Leung
cs.LG · cs.AI · eess.SY
Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution augments the state with information to satisfy the Markovian property and applies standard RL techniques. However, prior work finds that the multi-policy approach for different phases can outperform a single...
-
29
Beyond Random Splits: Evaluating Drug-Target Affinity Models Under Chemically and Biologically Motivated Distribution Shifts Copy
Minjae Chung, Clara Li, Malar Paavai Muthukumaran, Shaunna Wang, Aniket Ramkrishnan Iyer, Shaun Qien Yeau Tan,...
cs.LG
Drug-target affinity (DTA) prediction is widely used to prioritize candidate compounds before costly experimental screening. DTA models are often compared under a single data split, even though deployment may require extrapolation to new chemical series, new protein targets, or both. We ask whether the distribution shift used for evaluation changes which architecture appears best. We curate 718,800 unique drug-protein pairs from the ChEMBL...
-
30
Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition
Eeshaan Jain, Linus Bleistein, Bart Deplancke, Charlotte Bunne
cs.LG · cs.AI
Recent progress in multimodal, high-dimensional learning has enabled foundation models to process heterogeneous, large-scale data. However, at test time, acquiring all features or modalities can be prohibitively costly and often redundant. Sequentially selecting informative modalities is therefore critical, yet challenging when the downstream task or prediction target is unknown. To this end, we introduce ECHO-$k$, a task-agnostic and...
-
31
Causal Representation Learning with Instantaneous and Lagged Relations via Nonstationarity
Tatsuya Yamada, Hiroshi Morioka, Yoshinobu Kawahara
cs.LG
Causal representation learning for time-series data aims to identify latent states and their causal relations from observations. In this setting, an important challenge is to model both lagged causal relations across observation intervals and faster causal effects that appear as instantaneous relations within an interval, while accounting for nonstationarity in time-series data. However, methods that jointly handle these causal relations and...
-
32
OptiSelect: How does the Optimizer Shape Data Curriculum?
Simin Fan, Alireza Abdollahpoorrostam, Martin Jaggi
cs.LG · cs.AI
Online data selection has demonstrated substantial efficiency gains for LLM pretraining by training on the most valuable candidates within each batch. Since a candidate's value is realized through its effective model update, principled selection should account for the optimizer step, which reshapes the raw gradient before it updates model parameters. We formalize this optimizer-aware selection paradigm as OptiSelect and present the first...
-
33
Deep Bayesian REFoCUS
Simon Penninga, Ruud van Sloun
cs.LG
In this work we formulate ultrasound multistatic recovery from arbitrary transmit sequences as a Bayesian inference problem. To that end, we train a deep generative prior on multistatic data sets to tackle the rank-deficient regime in which classical linear REFoCUS decoders fail. This appproach, which we term Deep Bayesian REFoCUS, outperforms the linear baselines for all regimes of rank-deficiency and noise levels, and regresses to linear...
-
34
Rethinking Epistemic Uncertainty in Node Classification through Information Growth
Emma Meneghini, Francesco Ferrini, Bruno Lepri, Andrea Passerini, Veronica Lachi
cs.LG · cs.AI
Epistemic uncertainty should decrease as additional information about the data-generating process (DGP) becomes available to the predictor. Yet, existing graph evidential deep learning (EDL) methods for node classification typically construct epistemic uncertainty from graph-specific properties and evaluate it on downstream tasks such as out-of-distribution detection, which do not test its reducibility as information about the DGP increases....
-
35
AIBL: Augmented Instance-Based Learning with Structured Memory and Neural Embeddings
Radha Poovendran, Andrea Stocco, Linda Bushnell
cs.LG
Sequential learning systems often make decisions from accumulated experience while receiving high-dimensional inputs whose distribution may change over time. Instance-Based Learning Theory (IBLT) provides a principled case-based framework for such settings through stored situation-decision-utility instances, partial matching, activation, and blending. IBLT relies on symbolic knowledge representation in dictionary-like formats, but text,...
-
36
16-bit Precision of Convolutional Neural Networks on Microcontroller Units for 8-bit Costs
Rui Liu, Benjamin Paaßen
cs.LG
To deploy deep neural networks on edge hardware, highly efficient inference schemes are necessary that retain high accuracy. This work presents W16A16, a high precision (16-bit), fast speed, low energy quantization method. On a widely applied microcontroller architecture Armv7E-M, our proposed approach achieves faster speed and lower energy consumption on layer- and model-level compared to alternative quantization schemes. We analyze the...
-
37
A Unified Framework for Bayesian Data Assimilation with Generative Models and Observation Interpolants
Nikolaj T. Mücke, Benjamin Sanderse
cs.LG · cs.CE
Bayesian data assimilation combines model forecasts with noisy observations, but sampling high-dimensional, non-Gaussian posteriors remains challenging. We introduce an observation-interpolant framework that turns pretrained stochastic interpolant, flow matching, and diffusion models into posterior samplers without retraining. Conditioning the interpolant path on observations yields a shared likelihood-score correction to the drift or...
-
38
Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning
Juri Pfammatter, Kaixian Qu, Clemens Schwarke, Victor Klemm, Marco Hutter
cs.LG · cs.RO
Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, hand-designed curricula, and shaped rewards supply this signal but require demonstrations or task-specific engineering; automatic start-state and goal curricula avoid this but typically expand from one side only, so...
-
39
Operator-informed initialization for Fourier features physics-informed neural networks
Juan Molina, Paris Perdikaris, Mircea Petrache, Matías Courdurier, Francisco Sahli Costabal
cs.LG
Physics-Informed Neural Networks (PINNs) typically exhibit spectral bias, where some frequencies of the target function converge more slowly than others. In this work, we analyze the training dynamics of Fourier Feature PINNs in the Neural Tangent Kernel regime to address this limitation. We derive an explicit evolution equation to estimate the residual error in the frequency domain, demonstrating that the convergence rate of specific...
-
40
SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents
Shangyang Wu, Shuai Zhao, Ziyue Zhu, Jinyang Wu, Anh Tuan Luu, Haoran Luo
cs.LG
Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as student-generated histories grow. To address this problem, we introduce SCAD, which organizes interactions into planning and bounded subtask execution, distills execution in local...
-
41
Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT
Joery Ariën de Vries, Neil David Lawrence, Zhenwen Dai
cs.LG · cs.AI
Critic-free reinforcement fine-tuning (RFT) for agentic large language models is often done through GRPO-style methods, which compute a group baseline over repeated rollouts to reduce target variance. However, this setup is ill-suited to agents acting in stateful environments such as live services or security sandboxes, where repeated rollouts are impractical to obtain and aggressive updates entrench the noise of long, sparsely verified...
-
42
Cordial Learning: Distributed Training with Correlated Data
Sarah Shitrit, Ilai Bistritz
cs.LG · cs.AI
We consider a distributed learning task with agents that have correlated data. Specifically, the label of an agent depends on the input of other agents for the same sample, and these inputs are also correlated. Correlated data is the reality when agents share the same environment. Existing decentralized methods, such as federated learning, ignore the structure of the problem and perform poorly on correlated data. On the other hand,...
-
43
Training-Loss Guarantees for Muon with Finite-Step Newton--Schulz Orthogonalization
Amartya Roy, Souvik Chakraborty
cs.LG · cs.AI · math.OC
Existing convergence analyses of Muon either assume exact orthogonalization or analyze classical Newton--Schulz polynomials, and guarantee only stationarity, so it is unresolved what Muon's five tuned Newton--Schulz steps preserve and whether that suffices to reach a prescribed neural-network training loss. We establish a finite-time training guarantee that accounts for both momentum accumulation before orthogonalization and the tuned...
-
44
S$^{2}$-PINN: Stochastic Separable Physics-Informed Neural Networks
Zhendong Li, Akwum Onwunta
cs.LG · math-ph
Uncertainty quantification (UQ) for random partial differential equations (PDEs) is ubiquitous in computational science and engineering. However, classical spectral solvers for this class of problems face the curse of dimensionality, and existing neural solvers often ignore the stochastic structure that makes moments and calibration tractable. We introduce a stochastic separable physics-informed neural network, dubbed S$^{2}$-PINN, that...
-
45
Architecture-Dependent Fusion Pathways in MLLMs
Hebao Zhu, Dongxia Wu
cs.LG
Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures. We conduct three progressively connected analyses: alignment decoupling identifies...
-
46
SPEAR: A Spectral-Disentangled MoE Neural Operator with Knowledge-Guided Expert Aggregation for Large-Scale PDE Pretraining
Dengdi Sun, Xiaoya Zhou, Xiao Wang, Wanli Lyu, Jin Tang, Bin Luo
cs.LG · cs.AI
Large-scale pre-training has improved the generalization of neural operators across diverse PDEs. However, existing PDE foundation models still struggle with heterogeneous dynamics, where shared representations may cause knowledge interference, while mixture-of-experts (MoE) architectures suffer from increasing expert redundancy. We propose SPEAR, a spectral-disentangled MoE neural operator with knowledge-guided expert aggregation for...
-
47
PaMIR: Open Benchmark of Public Credit-Default Datasets
Mikhail Liashkov, Ilyas Varshavskiy, Shuhratjon Khalilbekov, Azizjon Azimi, Bonu Boboeva
cs.LG
We release PaMIR (Public Arrival-ordered Measurement for Inference in Risk), an open benchmark for credit-default prediction when labels are scarce and arrive late. The field's reference benchmark studies use eight datasets each, only two or four of them public. PaMIR brings together 19 public datasets with binary default labels -- 1.24M loans, firms and card accounts from nine countries -- rebuilt from pinned source snapshots by one...
-
48
Mapping and Advancing the Scalability-Accuracy Frontier of Nonlinear Causal Discovery
Hendrik Suhr, Sascha Xu, Jilles Vreeken
cs.LG · cs.AI
Scalable nonlinear causal discovery requires methods that combine flexible mechanism estimators with efficient search over large graph spaces. Several algorithmic families have been proposed to address this challenge, yet their accuracy-runtime trade-offs remain poorly understood. We empirically compare the four major approaches: differentiable structure learning, amortized structure learning, score-matching, and combinatorial search. Our...
-
49
Cross-cohort TB classification using clinical data gathered in Uganda and South Africa
Joshua M. Jansen van Vüren, Devendra S. Parihar, Daphne Naidoo, Marisa Klopper, Frank Cobelens, Lutz Kolbe, Kimsey...
cs.LG
We present a first evaluation of machine learning applied to patient clinical and demographic data gathered in two different countries for the purpose of tuberculosis (TB) screening to identify people who would benefit from expensive molecular testing. Experiments are based on the recently-compiled CAGE-TB dataset, which includes sub-cohorts of people with presumptive TB presenting at community health care centres in South Africa and Uganda....
-
50
D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?
Daifeng Li, Huiqiang Jiang, Chengruidong Zhang, Wei Wu, Xudong Guo, Jianhong Tu, Jianwei Zhang, Binhang Yuan, Dayiheng Liu
cs.LG · cs.AI · cs.DC
GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents translate expert design guidance into efficient GPU kernels. The guidance covers L1: high-level algorithmic insights, L2:...
-
51
The Neuro-Physical Inverter: A Modular Framework for Magnetotelluric Inversion Coupling Ensemble Conditioning with Residual Learning
Jae Deok Kim, Sai Ravela, Rob. L. Evans
cs.LG
We present the Neuro-Physical Inverter (NPI), a modular, uncertainty-aware framework for geophysical inversion that couples ensemble-based conditioning with constrained residual learning, demonstrated in the 1D magnetotelluric (MT) setting as a controlled testbed. The framework operates in two stages. An Ensemble-Conditional Gaussian Process (EnsCGP) conditions a prior ensemble of resistivity models on the observed response, producing a...
-
52
AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning
Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan
cs.LG · cs.CL
Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained supervision, but its estimates can be unreliable because observed returns also depend on subsequent actions, environment transitions, and trajectory length. We propose AdaStep, an Adaptive Step-credit weighting...
-
53
Kernel Singular Value Decomposition with Extension to Multiple Data Sources
Xinjie Zeng, Qinghua Tao, Johan Suykens
cs.LG
Kernel Singular Value Decomposition (KSVD) learns a pair of singular vectors w.r.t. an asymmetric kernel matrix, which can be induced by two data sources, e.g., the queries and keys in self-attention or the rows and columns of a given matrix. In this work, we extend KSVD to multiple data sources, namely eKSVD, which conducts joint nonlinear feature learning upon asymmetric kernels. In the primal formulation, the projections associated with...
-
54
HyperFuse: Fast Self-Supervised Node Embeddings for Attributed Hypergraphs
Megha P, Harshit Kumar, Srajan Agarwal, Anirban Banerjee, Olaf Wolkenhauer, Saptarshi Bej
cs.LG
Self-supervised hypergraph representation learning can produce informative node embeddings, but existing methods often require deep encoders trained for hundreds of epochs, making embedding generation costly even for hypergraphs with a few thousand nodes. This limits applications requiring embeddings for many or evolving hypergraphs. We present HyperFuse, a label-free pipeline for fast hypergraph representation learning. HyperFuse (i)...
-
55
Predicting and Repairing Merge Collapse in Large Language Models
Jungseob Lee, Seungyoon Lee, Sugyeong Eo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim
cs.LG · cs.CL
Large language models fine-tuned from a shared base can be merged by averaging their task vectors, but some merges collapse far below the base model, and common merge operators give no warning before evaluation. We show that one statistic of the specialists' task vectors both predicts this collapse and calibrates its repair. The power that averaging removes equals the variance of the task vectors across specialists, our measure of...
-
56
Sample complexity of variance-reduced policy gradient: weaker assumptions and lower bounds
Gabor Paczolay, Matteo Papini, Alberto Maria Metelli, Istvan Harmati, Marcello Restelli
cs.LG
Several variance-reduced versions of REINFORCE based on importance sampling achieve an improved $O(ε^{-3})$ sample complexity to find an $ε$-stationary point, under an unrealistic assumption on the variance of the importance weights. In this paper, we propose the \algo (Defensive Policy Gradient) algorithm, based on defensive importance sampling, which achieves the same rate without any assumption on the variance of ordinary importance...
-
57
Page-EntroKV: Hardware-Aligned, Entropy-Weighted KV-Cache Eviction under Grouped-Query Attention
Inbasekaran S
cs.LG
Serving long-context autoregressive language models is constrained by the key-value (KV) cache. Most dynamic eviction methods score token importance per query head and choose tokens independently. This fits poorly with grouped-query attention (GQA), where several query heads share one physical KV buffer: divergent per-head selections force the serving engine to retain the union of their choices - inflating the cache by up to the group ratio r...
-
58
Coverage You Can Steer: Online Conformal Calibration for RL-Driven Hardware-Aware NAS
Pedro Brandimarte, Nerea Aranjuelo, Marcos Nieto, Oihana Otaegui
cs.LG
Hardware-aware neural architecture search (NAS) is dominated by evaluation cost: every architecture must be trained before its reward is known. Conformal-prediction filters cut this cost by pruning candidates whose predicted-reward upper bound misses a threshold, with a distribution-free guarantee that at most a fraction $δ$ are wrongly discarded. That guarantee assumes exchangeability between calibration and test candidates, which the...
-
59
How to Find and Reuse Policies for Continuous Adaptation in Lifelong Reinforcement Learning
Saptarshi Nath, Inish M. D'Souza, Antonio Carta, Soheil Kolouri, Andrea Soltoggio
cs.LG · cs.AI · cs.RO
In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as the learner acquires experience. One hypothesis is that task similarity can be effectively used in a continual learning setting to find and combine previously learned policies. To test it, Adaptive Mask Selection and...
-
60
Exploring the Trade-Off Between Structured Pruning and Fault Tolerance in Deep Neural Networks for Space Applications
Toon Vinck, Naïn Jonckers, Jaro De Roose, Jeffrey Prinzie, Peter Karsmakers
cs.LG
Deep Neural Networks (DNNs) inherently exhibit a degree of robustness to bit-level faults due to their distributed representation of information. As a model increases in width, this information becomes more dispersed, theoretically reducing the impact of any single bit fault. In this paper, we empirically investigate the relationship between model width and robustness to Single Event Upsets (SEUs). We conduct a comprehensive experiment in...
-
61
LS-AR: Future-Predictive Latent Steering in Autoregressive LLMs
Anubha Gupta, Eduardo Pignatelli
cs.LG
Standard autoregressive (AR) models process high-level task instructions, state history, and transient tokens within a single shared sequence of tokens. Consequently, they lack the architectural mechanisms needed to isolate macro-objectives from context noise. To overcome this single-channel limitation, we introduce Latent-Steered Autoregressive (LS-AR), a dual-channel architecture that decouples continuous goal steering from discrete token...
-
62
ULTRADISCOVERY: Abductive Exploration in an Interconnected, Epistemically Open Universe
Weihan Li, Tianshi Zheng, Yangqiu Song, Ginny Y. Wong, Simon See
cs.LG · cs.AI
Scientific discovery often begins when scattered clues call for a new way of describing the world. Such abductive exploration can require constructing the representation in which an explanation is stated, when the world is epistemically open, and composing evidence scattered across contexts, when it is structurally interconnected. Existing benchmarks rarely separate these two demands or control them independently. We introduce ULTRADISCOVERY,...
-
63
Zephon: Elastic Determinism for Online, Stateful Foundation Model Data Loading Pipelines
Maximilian Böther, Josh Wills, Ties Robroek, Sonnet Xu, Paul Burstein, Daniel Zayas, Cody Blakeney, Siddharth Joshi,...
cs.LG · cs.AI · cs.DB
Deterministic data loading is important for foundation model development: model researchers need confidence that differences they observe across costly ablations are caused by the parameter they changed rather than non-determinism in the training data sequence. The data loader must provide elastic determinism, i.e., a deterministic sequence of global training data batches despite changes to the GPU topology across runs (e.g., due to GPU...
-
64
Light Entropic Optimal Transport on Riemannian Manifolds
Xavier Aramayo-Carrasco, Petr Mokrov, Alexander Korotin
cs.LG
Entropic Optimal Transport (EOT) has become a practical framework for learning stochastic couplings between complex distributions, with applications in generative modeling and domain adaptation. However, most EOT solvers are designed for Euclidean spaces, while manifold extensions remain limited and often rely on costly iterative methods, simulated dynamics, or generic neural models that do not fully exploit the underlying geometry. We...
-
65
Smart Sensing for Safer Bridges: From Sensor Signals to AI-Driven Anomaly Detection
Rahul Jaiswal, Joakim Hellum, Halvor Heiberg
cs.LG
Bridges contribute significantly to transportation connectivity and urban development. Therefore, reliable bridge monitoring is crucial for protecting public safety and detecting anomalous behavior in bridge sensor data that may provide early indications of abnormal structural conditions. This paper investigates anomaly detection in real-world bridge sensor data using two different complementary approaches, namely signal processing and the...
-
66
Learn Feasibility Once, Optimize All Objectives: Derivative-Free Diffusion Models for Chance-Constrained Programming
Ziwen Liu, Yan Liu, Congying Han, Tiande Guo, Yao Yan, Weichen Zhao
cs.LG · math.OC
Chance-constrained programs (CCPs) optimize decisions under uncertainty by limiting the probability of constraint violation. Despite advances in traditional and learning-based approaches, optimizing non-convex or non-smooth objectives and adapting to different objectives under fixed chance constraints remain challenging. In this paper, we propose a \textbf{D}erivative-free \textbf{D}iffusion-based framework that \textbf{D}isentangles...
-
67
Learning Transferable Policies from Action-free Time Series Through Dynamical Embeddings
Niklas Emonds, Georgia Koppe
cs.LG
Learning control from action-free recordings is challenging because intervention effects are unobserved and policies may exploit errors in reconstructed dynamics. We present a hierarchical model-based reinforcement learning framework that uses shared structure across related systems to learn system-specific control policies from action-free recordings. A hierarchical dynamical system reconstruction model captures shared dynamics and...
-
68
When Does Synthetic Relational Data Teach Models to Use Relations? Tracing Predictive Structure from Pretraining Data to Model Behavior
Shivam Dubey, Mohamed Bouadi, Nassim Bouarour, Varun Kulkarni, Aditya Tanna, Vinay Kumar Sankarapu
cs.LG
Relational foundation models are increasingly pretrained on synthetic databases, yet downstream benchmarks reveal little about why one synthetic corpus produces a better model than another. In particular, strong performance may arise from realistic row-level statistics without the model ever learning to use relational structure. We study this as a data-attribution problem: which property of synthetic pretraining data induces relational...
-
69
RIPPLE in Still Water: Zero-Shot Clustering in Federated Learning with Wavelet Scattering Transform
Alessandro Licciardi
cs.LG
Clustered Federated Learning (FL) partitions a client population into groups of similar local distributions and trains one specialized model per cluster, mitigating client drift that degrades single-model methods under non-IID data. Prior methods discover cluster structure inside the training loop through gradient similarity, loss evaluation, or EM-style updates, thus increasing communication overhead, exposing gradients to inversion attacks,...
-
70
WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites
Jiangang Han
cs.LG · cs.AI · cs.CV
We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site's own interface and return a verifiable answer. A capable multimodal large language model (LLM) is necessary for this, but...
-
71
Balancing Multimodal Learning via Functional Progress
Zhongjing Gu, Fengqiang Wan, Yiming Cui, Yufa Feng, Yang Yang
cs.LG
Multimodal learning often suffers from modality imbalance, where the joint optimization process is dominated by a single modality. Existing methods typically estimate modality imbalance from score disparities derived from prediction uncertainty or optimization statistics. However, due to distinct prediction uncertainty and learning dynamics across modalities, direct comparison of such scores may misinterpret intrinsic modality differences as...
-
72
Adaptive Second-Order Solvers for Fast Stochastic Diffusion Sampling
Ella Kemperman, Luca Ambrogioni
cs.LG · cs.CL · cs.CV
Diffusion models rely on numerical solvers requiring time-discretization, which has a large influence on the tradeoff between sampling cost and quality. However, the computational difficulty of the reverse process varies along the sampling trajectory and across data distributions, making the choice of discretization important. We adapt proportional-integral (PI) step-size control to diffusion, using our diffusion noise-normalised error...
-
73
Tailoring the Quantization Space for 1-Bit KV Cache Compression
Minsoo Cheong, Donghyun Son, Sungjoo Yoo
cs.LG · cs.AI · cs.CL
The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade substantially in the 1-bit regime. At such extreme compression, each codebook must represent a larger group of channels with a...
-
74
Signal Simplification Is Not Predictive Simplification: Diagnosing Residual Neural Forecasting in Short-Horizon Volatility
Bingqi Lian, Linfeng Cheng, Mei Lu, Jerry Wu
cs.LG
Hybrid statistical-neural pipelines often assume that a successful statistical first stage leaves a cleaner and more learnable residual target. We examine that assumption in short-horizon volatility forecasting through a signal-forecast-system diagnostic framework. Across five liquid U.S. assets, a volatility-aligned HAR-style model outperforms AR, MA, and ARIMA. Within expanding training windows, the pre-standardization fitted residual...
-
75
AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning
Han Yu, Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Hejian Sang, Han Shi, Menglin Zhou, Xuanzhao Dong, Minzhou Huang, Rui...
cs.LG · cs.AI
Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: low-attention entries can carry large value payloads whose removal changes future predictions, and newly generated states can...
This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.