cs.LG · 2026-10-06 · No. 135
Machine Learning, 2026-10-06.
74 new papers in cs.LG. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
74 entries-
01
Base Models Can Reason By Taking a Cue From Training Data
Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min, Alexei A. Efros
cs.LG · cs.AI · cs.CL
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42%...
-
02
Towards Looped Models Done Right, Part II: Rethinking at Fixed Points
Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen, Zhengzhong Liu, Eric Xing, Xuezhe Ma
cs.LG
Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout...
-
03
Deep Learning for Sleep Heart Rate Estimation from Accelerometers: Toward Population-Scale Cardiac Insight Without Optical Sensors
Tanbin Islam Rohan, Pranjol Sen Gupta, Tanusree Debi, Nazmus Sakib
cs.LG · cs.AI
Large longitudinal cohorts often contain wrist accelerometry without optical heart-rate sensing, motivating recovery of cardiac information from motion signals already collected during sleep. We present SeqSmoother, a transformer-based temporal corrector for sleep heart rate (HR) estimation from wrist accelerometry. SeqSmoother combines spectral descriptors with an intermediate Nightbeat-derived frequency anchor and a physics-motivated...
-
04
Private online learning and prediction for Littlestone classes
Amartya Sanyal
cs.LG · cs.CR · stat.ML
We study mistake bounds for differentially private online learning and online prediction under oblivious realisable adversaries. Online learning requires the learner to release a hypothesis at each time step whereas in online prediction, the learner only needs to make predictions without releasing a hypothesis. Using a novel lower bound for private online learning and an upper bound for private prediction, we show that the sample complexity...
-
05
Block Disentanglement in CRL: Bridging Identifiability and Visual State Estimation
Emre Acartürk, Pranamya Kulkarni, Puranjay Datta, Karthikeyan Shanmugam, Burak Varıcı, Ali Tajer
cs.LG · stat.ML
Causal representation learning (CRL) is the process of recovering causally-related latent variables from high-dimensional observations. As a label-free inference method, CRL is particularly attractive for applications where data labels are unavailable or impractical to obtain. While there has been significant progress in understanding the identifiability guarantees of CRL, such guarantees often hold under highly stylized assumptions, which...
-
06
H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning
Wancong Zhang, Basile Terver, Michael Rabbat, Yann LeCun, Randall Balestriero
cs.LG · cs.RO
Long-horizon planning with latent world models requires reasoning across timescales and levels of abstraction. Existing task-agnostic JEPA world models predict and plan at a single timescale or with multiple horizons in one shared latent space. We introduce H-JEPA, an end-to-end recipe for training a hierarchy of action-conditioned JEPAs in which each level predicts farther ahead in its own learned latent space. Planning proceeds top-down:...
-
07
Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution
Erfan Baghaei Potraghloo, Seyedarmin Azizi, Arya Fayyazi, Saeid Shokoufa, Mehdi Kamal, Souvik Kundu, Massoud Pedram
cs.LG · cs.AI
A language model can give a correct answer more probability than any single incorrect answer and still usually sample an incorrect one, because the incorrect answers together hold more probability. The power distribution raises each complete answer's probability to a power above one and renormalizes, shifting probability toward answers the model finds most likely (sharpening). Sampling from it improves reasoning without changing parameters,...
-
08
Round-Trip KNN Clustering: multiscale hierarchical cluster detection on directed nearest-neighbour graphs
Eraldo Pereira Marinho, Caetano Mazzoni Ranieri, Fabricio Aparecido Breve
cs.LG · stat.ML
We introduce Round-Trip KNN Clustering (RTKNNC), a graph-based method for finding cluster structure at several neighbourhood scales without requiring the number of clusters in advance. Unlike approaches that first make a $k$-nearest-neighbour (KNN) graph undirected, RTKNNC keeps both directions of the neighbour relation: which points a given point selects and which points select it. Incoming selections are treated as weighted votes that help...
-
09
MatrixFormer: A Foundation Model for Matrix Completion
Dwaipayan Saha, Jacob Feitelberg, Kyuseong Choi, Raaz Dwivedi, Anish Agarwal
cs.LG · cs.AI
Matrix completion underlies problems from tabular imputation to causal inference, yet existing tabular foundation models treat it as entry-by-entry prediction, repeating context for every target and discarding the matrix's two-dimensional structure. We introduce MatrixFormer, a pre-trained matrix-native transformer that predicts a full distribution for every missing entry in a single forward pass. MatrixFormer is trained entirely on synthetic...
-
10
Hyperbolic Graph Representation Learning: Embed in One Metric, Optimize with Another
Federico Larroca, Paola Bermolen, Marcelo Fiori, Bernardo Marenco
cs.LG · stat.ML
Hierarchical graphs embed in hyperbolic space with lower distortion than in Euclidean space owing to its negative curvature. However, their gradient-based learning is hampered at large radii, where the Poincaré ball and the Lorentz hyperboloid models fail numerically. Polar coordinates avoid this problem, but the hyperbolic metric scales the angular step by the hyperbolic sine of the radius, freezing angular motion. We observe that this...
-
11
Decoupling Time and Space: A Temporally Conditioned Refinement for EEG Source Imaging
Marco Morik, Jesse Palarus, Carmen Vidaurre, Klaus-Robert Müller, Shinichi Nakajima
cs.LG
Electroencephalography (EEG) offers millisecond temporal resolution, but inferring underlying neural sources is a severely ill-posed spatial inverse problem. While deep learning has advanced spatial reconstruction, current architectures face a critical dilemma: frame-by-frame models discard vital temporal context, whereas full 4D spatiotemporal networks introduce an architectural trade-off between reconstruction accuracy and inference cost....
-
12
BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models
Gang Fu, Adel Javanmard, MohammadHossein Bateni, Vahab Mirrokni
cs.LG · cs.AI
Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry no topological meaning. We introduce {\bf BRANCH-MoE}, a routing architecture that places \(E\) experts at the leaves of a binary decision tree of depth \(\log_2 E\). At each internal...
-
13
To Learn is to Wander: Learning Across Graphs and Tasks with Random Walks
Louis Tichelman, Xingyue Huang, Jinwoo Kim, İsmail İlkan Ceylan
cs.LG
Graph foundation models aim to transfer across graphs, feature spaces, relational schemas, and prediction tasks, yet existing approaches typically generalize only within particular graph modalities or tasks. We propose Wander, a graph foundation model designed to operate across these settings within a single pretrained checkpoint. Following the prior-predictive perspective, we formulate graph learning as completion of a partially observed...
-
14
Adapting prior-data fitted networks for tabular anomaly detection
Maximilian Bershtman, Niv Cohen
cs.LG
While deep features have transformed anomaly detection in images and video, their impact on tabular data has been less substantial, partly due to the limited availability of strong deep representations. Recently, prior-data fitted networks (PFNs) have emerged as a promising source of such representations for tabular data. In this work, we investigate how PFN representations can be adapted and leveraged for anomaly detection. The question is...
-
15
OVAL: Output-Aware Local Page Bases for KV Cache Retrieval
Ashkan Shahbazi, Chayne Thrash, Soheil Kolouri
cs.LG
Long context inference with large language models becomes increasingly expensive as attention must operate over an ever growing KV cache. Page sparse attention reduces this cost by representing each KV page compactly and retrieving only a subset for each query. Existing retrieval methods are designed to estimate attention scores or page relevance, but their objectives do not directly account for how approximation errors affect the resulting...
-
16
Aligning Multimodal Patient Evidence with Biomedical Knowledge Graphs for Clinical LLMs
Jiawen Du, Arshan Ali Khan, Chenhao Zhang, Zachary Plotkin, Li Shen, Qi Long, Yun Li, Can Chen, Tianlong Chen, Nicholas Konz
cs.LG · cs.CL
Clinical questions often depend on linking a patient's multimodal evidence to external biomedical knowledge, yet existing predictive systems rarely represent such links explicitly, so they can neither be traced to their evidence sources nor removed to measure their contributions. We present MM-KG (Multimodal Knowledge Graph), which represents heterogeneous, multimodal patient observations and biomedical concepts as separate layers in one...
-
17
Closing the Context Gap: Activation Alignment for Tabular In-Context Learning
Yoel Zeldes
cs.LG · cs.AI
Tabular foundation models perform in-context learning (ICL) by conditioning predictions on labeled training examples provided as context. Unlike traditional models that separate training from inference, these models must process all training examples in every forward pass, making each prediction expensive. Restricting the number of training examples reduces this cost but substantially degrades performance. Instead of discarding context, we...
-
18
How Sparse Probability Maps Shape Mixture-of-Experts Routing
Tomás Brogueira, Marcos Treviso, Miguel Couceiro
cs.LG · cs.CL
Mixture-of-experts (MoE) routers typically apply softmax to the router scores and keep the top-K experts, making every token use exactly K experts. Sparsity-inducing probability maps such as sparsemax, alpha-entmax and normmax can adaptively assign exact zeros to selected experts, and therefore appear to offer token-dependent expert participation, even when using the same top-K machinery. In this work, we study whether and how this sparsity...
-
19
Improved Convergence of Large Stepsize Gradient Descent for Logistic Regression
Xiaochuan Gong, Ang Li
cs.LG
We study gradient descent (GD) with a large constant stepsize for logistic regression on linearly separable data. Existing analysis shows an accelerated rate of $\widetilde{O}(1/\sqrtε)$ to reach loss $ε$ with an aggressive stepsize, although the loss may initially oscillate. Tighter control of the oscillatory dynamics has been available only for two-dimensional data. We prove a substantially faster rate in arbitrary dimension: GD with a...
-
20
Detecting Nighttime Anomalies from NASA Black Marble Using a Generalized Spatio-Temporally Robust Framework of Machine Leaning Ensembles
Srija Chakraborty
cs.LG · cs.CV
Nighttime lights from NASA's Black Marble product suite capture thermal and light emission signals from anomalous events including fires, volcanic eruptions, and gas flaring. Existing detection approaches rely primarily on thermal bands, limiting sensitivity to weaker signals. We propose a novel machine learning framework that jointly models Black Marble M-band and Day/Night Band (DNB) signals to derive a generalized, spatio-temporally robust...
-
21
Learning What to Imitate: Entropy-Aware Distribution Mixing
Juan Garcia Giraldo, Matteo Santelmo, Eduard Durech, Imanol Schlag, Valentina Pyatkin, Antoine Bosselut
cs.LG
Small language models are often post-trained as students on reasoning traces from stronger teacher models to efficiently learn new skills. However, token-level imitation on traces that lie far outside the student's expected distribution often produces \textit{confident conflicts}, whereby the student is required to imitate a continuation that it deems unlikely (i.e., low-probability) despite being confident in a different continuation (i.e.,...
-
22
What Matters for Latent Reasoning with Flow Matching
Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos
cs.LG · cs.CL
Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows,...
-
23
TrustmeWatcher: An Application for Workplace Micro-Sensing and Explainable Well-Being Feedback
Chengyu Yu, Leon Jacopo Costa, Zoja Anžur, Mohan Li, Gašper Slapničar, Daniil Kirilenko, Martin Gjoreski, Mitja...
cs.LG · cs.HC
Workplace sensing studies combine long-running behaviour traces with self-reports, yet the tools that collect those data often sit apart from the interface that returns results. We present TrustmeWatcher, the application built for the TRUST-ME project to connect this work. TrustmeWatcher reuses ActivityWatch's OS-level watchers for computer-activity collection and adds its own application layer. It turns the collected traces into an...
-
24
The Birkhoff Geometry of Manifold-Constrained Hyper-Connections: Two Channels, Vertex Viscosity, and Sinkhorn as a Retraction
Xiaoyu Li, Zhizhou Sha, Chiwun Yang
cs.LG · math.OC · stat.ML
Hyper-connections widen the residual stream of a Transformer to $n$ parallel streams. Their manifold-constrained version (mHC) mixes the streams at each layer with a doubly stochastic matrix, which it computes by Sinkhorn normalization of exponentiated logits. We give a geometric theory of this design on the Birkhoff polytope. First, a doubly stochastic mixer splits the stream into a mean channel, on which mHC is exactly a residual network,...
-
25
Considering Context: When World Models Need Context Encoders
Oleg Smirnov, Sofiane Ennadir, John Pertoft, Bjartur Hjaltason, Sara Karimi
cs.LG
Methods for generalization in model-based reinforcement learning typically assume that an agent cannot recover the latent context governing the environment dynamics from its own experience, and therefore supplies it externally. We formalize and test this assumption with \emph{predictive sufficiency}, which quantifies what access to the context adds to next-step prediction under the visitation distribution an agent induces, and separates that...
-
26
Beyond the Model: The Critical Role of Data Filtering in Clinical Machine Learning
Noah Subedar, Colin Campbell, Wenjing Zhang, Dan Perri, Sarah Culgin, Andrew Hamilton-Wright
cs.LG
Machine learning (ML) studies using clinical data often rely on preprocessing and filtering pipelines before model development. The filtering decisions made in these pipelines can alter the dataset's statistical structure and may artificially reduce or increase the complexity of the prediction task. We argue that filtering choices should be treated as part of the scientific method rather than as a routine preprocessing step. We further...
-
27
Differentially Private Mixing of Public Datasets Improves Private Learning
Yufei Chen, Tejumade Afonja, Anvith Thudi, Nicolas Papernot
cs.LG · cs.AI · cs.CR
Many machine learning applications involve sensitive data and therefore require training under differential privacy (DP). However, DP training often degrades model utility. In some cases, first pre-training the model on "public" data before finetuning with DP on the sensitive data can reduce the drop in utility. However, the success of this depends on how relevant the selected public dataset is to the sensitive data. We introduce the first...
-
28
Frozen Factor or Spectral Band? Disentangling Two Choices in Low-Rank LoRA
Adnan Slimane Ali, Ayoub Belfatmi, David Ngwe Pouth
cs.LG · cs.AI · cs.CL
Spectral variants of low-rank adaptation (LoRA) choose both a subspace and which factor to freeze. We separate these choices by freezing the input factor A or output factor B on the top or bottom singular directions of pretrained weights, with learning rates selected separately. At rank 2, the same-band advantage of freezing A is larger than either within-factor band difference on all four task-model pairs with complete comparisons. Freezing...
-
29
Separators Make Carry Propagation Learnable:The Geometry of Latent Carry in a Multiplication Transformer
Sama Satariyan, Raphael Cousin, G{é}rard Biau
cs.LG
Transformers asked to multiply multi-digit numbers in a single forward pass often fail, and interpretability studies of pretrained language models find arithmetic solved by input-range heuristics rather than by an explicit carry. We train small Llama-style transformers from scratch on 4x4 multiplication without chain of thought and find that the input format is decisive: inserting a space token between digits raises exact-match accuracy from...
-
30
LinearPFN: Amortized Variable Selection for Linear Models with Interactions
Louis Schiekiera, Max Zimmer, Christophe Roux, Manuel Arnold, Sebastian Pokutta, Fritz Günther
cs.LG · stat.ML
Spike-and-slab regression is a standard Bayesian formulation of variable selection: it returns a posterior distribution over which candidate effects are active rather than a single selected subset, so that every candidate effect carries an inclusion probability. Its cost grows exponentially with the number of candidate effects, so the posterior can be enumerated exactly only when the number of predictors is small. Beyond that reach, the...
-
31
Empirical Variational Autoencoder
Kaede Shiohara
cs.LG
We present Empirical Variational Autoencoder, a general generative framework for continuous-valued (i.e., non-vector-quantized) sequences. EVA is based on the evidence lower bound of the Variational Autoencoder (VAE) but learns autoregressive latent priors empirically from training data, which can be implemented only by an additional single linear layer on top of VAEs. By replacing the conventional standard-Gaussian constraint with the...
-
32
A Fine-Grained Analysis of the LoRA Fine-Tuning Landscape with Implications for Data Selection
Bowen Zhang, Changrui Fang, Xinsong Ma, Jiaye Teng, Ziye Ma
cs.LG
Low-Rank Adaptation (LoRA) has become a standard approach for parameter-efficient fine-tuning, yet a fundamental practical question remains unresolved: how should the adapter rank be chosen? An overly small rank may lead to a poorly conditioned optimization landscape, whereas an unnecessarily large rank sacrifices the efficiency that motivates LoRA in the first place. Existing theoretical analyses provide only limited guidance on this...
-
33
WaveGSSM: Graph Wave State Space Models for Propagating Spatio-Temporal Patterns
Junyou Zhu, Fenying Cai, Ping Xiong, Christian Nauck, Langzhou He, Chao Gao, Jürgen Kurths, Frank Hellmann
cs.LG
Spatio-temporal graph models typically encode each snapshot with a GNN and then connect the resulting representations through a temporal module. This space-then-time design is effective, yet it does not explicitly represent how a pattern moves across the graph. We show empirically that, for a propagating process, the same present field can lead to different futures when its recent rate of change differs, motivating an explicit representation...
-
34
Conditional Flow Matching for Single-Neuron Electrophysiology: Capturing Multimodal Responses Across Stimuli
Cameron Schofield, Luca Ghafourpour, Philip H. Wong, Costas A. Anastassiou, Richard E. Turner
cs.LG
Neurons of the brain exhibit a rich repertoire of electrophysiology dynamics with the same repeated stimulus eliciting very different voltage responses from the same cell. One common approach in biophysically detailed models is to capture this variability through ensembles of deterministic parametrizations, at a cost of hundreds of thousands of CPU hours. Existing machine learning surrogates inherit the same limitation, where a stimulus is...
-
35
Xaurora: Generative Weather Forecasting with Denoising Stochastic Interpolants from a Foundation Model Prior
Eliot Walt, Miltiadis Kofinas, Nikolaj Mücke, Efstratios Gavves, Dim Coumou
cs.LG
Deep learning has revolutionised weather forecasting in recent years, especially through atmospheric foundation models, which offer competitive skill for a fraction of the computational costs of classic physics-based models. However, most existing foundation models are deterministic, limiting the generation of large ensembles for accurate uncertainty quantification, extreme weather risk assessment, and long-range weather forecasting....
-
36
Improving Proactive AI Assistance with Hierarchical Procedural Understanding
Jin-Seop Lee, TaeYeon Won, SeongJun Jung, JungHoon Kim, Boyang Albert Li, Jin-Young Park, Jaehong Yoon, Jee-Hyong Lee
cs.LG · cs.CV
Proactive AI assistants continuously observe a user's activity and decide whether to provide new guidance or remain silent. They should provide appropriate guidance for the task, determine when to provide the next guidance based on task progress, and adjust the guidance level to the user's expertise and needs. Supporting these capabilities requires training and evaluation data that reflect procedural structure and capture how guidance should...
-
37
FairProp: Fair Node Representation Learning via Differentiable Propagation Layers
Emmanouil Kariotakis, Aritra Konar
cs.LG · cs.SI
Graph neural networks (GNNs) are the standard tool for node representation learning and are increasingly used in high-stakes settings. Their message-passing backbone, however, can amplify topological bias, raising fairness concerns. We study group fairness at the level of downstream predictions for node classification, link prediction, and node regression, and bound the demographic parity gap for an arbitrary number of sensitive groups. Our...
-
38
Latent Flow Matching for Molecular Graph Generation
Mathis Goupillon, Roman Bresson, Konstantinos Divriotis, Michalis Vazirgiannis
cs.LG · cs.AI
Modern graph generative models typically operate directly in the discrete graph space, explicitly generating node and edge variables, which can become costly as graphs grow. In this paper, we perform generation explicitly on latent representations of entire graphs obtained from a pretrained Variational Autoencoder with high reconstruction fidelity. The generated representations, obtained through flow matching, are then decoded only at the...
-
39
A Physics-Guided Transformer Framework for Electromigration Analysis in Multi-Segment Interconnects
Pavlos Stoikos, Anuj Pathania, George Floros
cs.LG · cs.AI
As technology scales to smaller nodes, increasing current densities make electromigration (EM) one of the dominant reliability challenges in on-chip interconnects. Accurate transient stress analysis is needed to identify wires susceptible to EM degradation, but applying physics-based solvers across many interconnects remains computationally expensive. This paper proposes a physics-guided transformer framework for fast EM stress prediction in...
-
40
EMG-FM-Bench: A Comprehensive Benchmark for Foundation Model Transfer and Adaptation on Electromyography
Tianhao Wu, Xu Wu, Amirmohammad Radmehr, Jiawei Yu, Yi Wu, Phuc Nguyen, Jian Liu
cs.LG
Foundation models (FMs) are increasingly being developed for general time series and physiological signals, yet their transferability to downstream physiological tasks remains poorly understood. This question is particularly challenging for electromyography (EMG), where signal distributions vary substantially across users, sensing configurations, acquisition hardware, and downstream tasks. We introduce EMG-FM-Bench, a systematic benchmark for...
-
41
Time-series Foundation Models for Predictive Control: The Role of Excitation
Mazen Amria, Jasper Hoffmann, Philipp Bordne, Anna Rothenhäusler, Lilli Frison, Harald Taxt Walnum, Sebastien Gros,...
cs.LG · eess.SY
Deploying model predictive control (MPC) requires constructing or identifying a predictive model for each target system. Time-series foundation models (TSFMs) offer an attractive option thanks to strong zero-shot forecasting capabilities across systems. However, low forecast error does not guarantee that a TSFM captures the system's response to the alternative actions considered by the controller. We study this gap using residential heat-pump...
-
42
ARO: Aligned Representation learning for multi-Omics data
Amogh Singh, Yash Shah, Chiara D'Ercoli, Arash Mehrjou, Patrick Schwab, Timothy Jones, Pietro Liò
cs.LG · cs.AI
The high cost of functional molecular assays, and prevalence of missing modalities and unmatched samples in computational biology, create significant barriers to comprehensive multi-omic profiling, essential for capturing and reasoning over molecules, cells, tissues, and organisms. This work proposes a model that learns meaningful representations from multi-omics cancer data supporting the reconstruction of missing and unpaired modalities....
-
43
Better Call Reward: Reward Hacking as Strategic Abstention in Legal Reasoning Models
Subramanyam Sahoo, Justin Shenk
cs.LG · cs.AI · cs.CL · cs.CY
What happens when a legal AI model learns to look like a lawyer instead of reasoning like one? We fine tune Qwen3-8B with Group Relative Policy Optimisation (GRPO) against a proxy built from three surface features: citation count, legalese density, and response length. The model does not learn to reason more effectively. It learns to withhold commitment. Across 16 yes or no legal reasoning tasks from LegalBench (N=320), overall accuracy...
-
44
HeuFouFT: Task-Guided Metaheuristic Coordinate Search for Fourier Fine-Tuning
Ruiheng Wang, Yubo Hou, Yakun Zhu, Tianle Shen, Tao Wan, Zengchang Qin
cs.LG · cs.CL
We introduce Heuristic-Guided Fourier Fine-Tuning (HeuFouFT), a task-guided framework for selecting trainable frequency coordinates in Fourier fine-tuning. Existing uniform and Gaussian band-pass schemes allocate a limited spectral budget through fixed, task-agnostic rules. HeuFouFT instead searches for coordinates using downstream performance. A coarse intensity map from lightweight block-level probes initializes three metaheuristic...
-
45
Efficient Secure Federated Learning via Information-Theoretically Secure Key Distribution: A Medical Imaging Case Study
Ivan Donà, Hans H. Brunner, Álvaro Troyano Olivas, Chi-Hang Fred Fung, Momtchil Peev, Giovanni Iacca
cs.LG · cs.CR
Federated Learning (FL) enables collaborative training of models across institutions without centralizing sensitive data, making it well-suited for privacy-concerned applications, such as medical imaging. To protect FL model updates during secure aggregation, additive masking is commonly employed. However, its underlying classical key establishment is only computationally secure. On the other hand, physics-based Information-Theoretically...
-
46
Training-Free Transformer Merging via Sequential Local Operator Alignment
Akansh Maurya, Ya-Wei Eileen Lin, Stefanie Jegelka, Sebastian U Stich, Rotem Mulayoff
cs.LG
Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-key and value-output operators depend on composed matrices, overlooking the functional structure. Moreover, when merging earlier components, downstream components receive different...
-
47
FlashCart: Fast Cartesian Tensor Products for Equivariant Interatomic Potentials
Viktor Zaverkin, Payman Goodarzi, Sergey V. Sukhomlinov, Davit Hovhannisyan, Roland Aydin, Martin H. Müser, Mathias Niepert
cs.LG
Machine-learned interatomic potentials extend atomistic simulations beyond the length- and timescales accessible to electronic-structure methods. However, the computational cost of equivariant architectures limits the local correlations they can represent in practice and therefore their achievable accuracy. Here we introduce FlashCart, which makes higher-order correlations affordable by combining generated GPU kernels with an architecture...
-
48
Quantifying the Stability of Multi-Step Reasoning via Error Amplification
Dongyue Li, Ziniu Zhang, Minxuan Duan, Hongyang R. Zhang
cs.LG · cs.AI · stat.ML
We consider the stability of multi-step reasoning processes, which have extensive applications in language models, including chain-of-thought and algorithmic reasoning. While longer sequences of reasoning can improve a model's generation capability at test time, the errors due to intermediate reasoning steps can accumulate in autoregressive generation, and thus grow substantially at the end. In this paper, we ask: What are the key factors...
-
49
Learning Pareto Stationary Fronts via Single-Pass Backpropagation
Elina Rojin Celik, Marcos Medeiros Raimundo, Isabel Valera
cs.LG
We propose MOSEL (Multi-Objective Stackelberg Efficient Learning), a framework for a posteriori multi-objective optimization (MOO) in deep neural networks that recovers a full front of Pareto stationary solutions at the computational cost of standard single-objective training. MOSEL reformulates the problem as a bilevel optimization problem that leverages network modularity to decouple representation learning from objective-preference...
-
50
Scaling Down the Scaling Laws: Parameter Efficiency and Compute-Optimal Training in Resource-Constrained Large Language Models
Joe Dwyer
cs.LG · cs.AI
Large language models (LLMs) have achieved substantial performance gains through increases in model size, training data, and computational resources. However, traditional scaling approaches produce diminishing returns, rising financial and environmental costs, and barriers to participation for researchers operating outside large industrial laboratories. This review examines the evolution of LLM scaling theory from empirical scaling laws to...
-
51
Steering by Influence: Curvature Aware Data Weighting for Activation Steering
James A. E. Dixon, Stephen J. Roberts, Francesco Quinzan
cs.LG · cs.AI · cs.CL
Inference-time steering offers cheap, fine-grained control over a language model's outputs by estimating a concept's representation in activation space and shifting activations towards it. Existing methods build these representations from activation averages over contrastive datasets. These averages incorporate unrelated concepts and noise, and are dominated by a few tokens, meaning the activation transport encodes token-level rather than...
-
52
Multimodal Deep Survival Analysis for Sinkhole Susceptibility
Lucas Yuan, Minhee Kim, Zihan Li, Chunli Dai, Sanduni S. Disanayaka Mudiyanselage, Ming Ye, Kani Fu
cs.LG
Sinkholes are a widespread geohazard in karst terrain. In Florida, soluble carbonate bedrock, shallow groundwater, and intense rainfall combine to make subsidence both common and spatially heterogeneous. Predicting where and when sinkholes will occur is difficult for two reasons. First, locations without reported sinkholes cannot be directly labeled or sampled as true negative locations. Second, the potential factors governing sinkhole risk...
-
53
Stability-Shaped Deep Graph Learning
Junyou Zhu, Langzhou He, Fenying Cai, Christian Nauck, Ping Xiong, Chao Gao, Philip S. Yu, Klaus-Robert Müller,...
cs.LG
In deep graph neural networks, increasing depth enlarges the receptive field but often leads to over-smoothing, where node representations tend to align. We develop a unified, mode-wise stability framework for deep GNN propagation that provides a principled characterization of over-smoothing. By interpreting layer depth as time and layer updates as graph-coupled dynamics, over-smoothing can be understood as an undesirable dynamical...
-
54
Dynamic Minimax Regret Optimization for Robust LLM Post-Training
Chengbo Zang, Haoyu Dong, Mehmet Kerem Turkcan, Gil Zussman, Zoran Kostic, Javad Ghaderi
cs.LG
Modern LLM training increasingly relies on heterogeneous data sources spanning different domains, tasks, preference distributions, and difficulty levels. We study dynamic minimax regret for group-distributionally robust LLM post-training under instantaneous mini-batch-only bandit feedback. The framework views the training as a two-player sampler-optimizer process: a sampler adaptively selects among data sources using bandit feedback, while an...
-
55
SPDAlign: Interpretable Riemannian Alignment for EEG Forward Modeling Shifts
Shanglin Li, Shiwen Chu, Okan Koç, Chenyu Liu, Qibin Zhao, Motoaki Kawanabe, Mitsuo Kawato, Yi Ding
cs.LG
Electroencephalography (EEG) based brain-computer interfaces enable direct brain-to-device communication for applications such as rehabilitation and communication. However, their practical utility is often limited as the non-stationary nature of the EEG data introduces distribution shifts across domains (e.g., sessions and subjects). Adapting machine learning models to be invariant to these shifts in an unsupervised way, without using costly...
-
56
Trajectory-Guided Tokenization of Complex CSI for Wi-Fi Sensing
Ziyi Wang, Kenuo Xu, Jichu Jiang, Yumeng Yang, Zheng Chen, Xiaofei Bai, Muge Chen, Xuyang Chen, Jinglei He, Jannik...
cs.LG
Wi-Fi channel state information (CSI) enables contactless presence detection and gesture recognition. Its high-dimensional complex-valued time series require input representations that preserve informative temporal variations during compression. We propose Trajectory-Guided Tokenization (TGT), which combines complex trajectory decomposition with asymmetric attention to construct compact continuous tokens. For each antenna link and subcarrier,...
-
57
When Are Concept Bottleneck Model Explanations Faithful and Compact?
Stefano Teso, Emanuele Marconato, Steve Azzolin, Antonio Vergari
cs.LG · cs.AI · stat.ML
Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal explainability, we argue they should also be faithful, i.e., not misreport which concepts actually matter. We show that, for widespread CBM architectures, including recent VLM-based...
-
58
dIon: Fragmentation-Based Invariance for Self-Supervised Learning of Tandem Mass Spectra
Alfred Nilsson, Joel Lapin, Samuel H. Payne, Mathias Wilhelm, Lukas Käll
cs.LG
We introduce a novel invariance for peptide tandem mass spectrometry data, unlocking self-supervised representation learning that improves de novo sequencing of peptides. This invariance exploits the physical relationship between precursor properties (mass and charge) and fragment-ion evidence, without requiring peptide sequence labels. We introduce dIon, which adapts the DINO framework with two latent prediction tasks, both recovering a...
-
59
What May an Agent Change About Itself? A Containment Floor for Self-Configuring Agent Runtimes
Sajib Hossain, Moeen Uddin Mahmud
cs.LG
Many agent runtimes give the agent a tool for editing its own configuration. Some of that configuration grants abilities, such as enabling a tool. Other parts set the agent's limits: which directories it may write to, who may send it messages, which network address it listens on, how callers authenticate, and the gate that blocks risky writes. If the agent can edit those limits, a single ordinary request can widen them. We study this in a...
-
60
Ramp Metering Control via Hybrid State Deep Reinforcement Learning in Partially Observable Connected Vehicle Environments
Youcef Mehamlia, Nadir Farhi, Meriem Bouali
cs.LG · cs.AI · math.OC
Freeway on-ramp merges are major sources of congestion, causing significant economic and environmental costs. While Deep Reinforcement Learning (DRL) offers a promising solution for ramp metering, existing approaches rely primarily on aggregated macroscopic data. Connected vehicles (CVs) provide vehicle-level observations that can complement aggregate traffic measurements, but their limited penetration produces incomplete microscopic...
-
61
OCL-PDE: A Generative Framework for PDE Inverse Problems with Observation-Complementary Latents
Ding Yang, Chuqi Chen, Chang Ma, Yang Xiang
cs.LG
Partial differential equation (PDE) inverse problems are often ill-posed, making fine-scale details difficult to recover. We address this problem by introducing a learned observation-complementary latent representation that preserves reconstruction-relevant information and is combined with the observation to reconstruct the unknown field. Building on this representation, we propose OCL-PDE, a generative framework that encourages the...
-
62
SimAuthor: Harnessing Foundation Models for Persistent Scientific Simulator Authoring
Yishan Wang, Ran Piao, Mathias Funk, Aaqib Saeed
cs.LG
Foundation models can generate scientific code, but authoring a scientific simulator (an executable program encoding hypotheses about how mechanisms generate observable signals) requires iterative refinement. Scientific adequacy rarely admits a unique implementation or exact test, so simulators must instead be judged against limited real observations. We study this setting as scientific simulator authoring under weak empirical feedback, where...
-
63
Lipschitz Thinking: Ten Years of Certifiable-by-Design Robust Neural Networks
Fabio Brau, Giorgio Piras, Maura Pintor, Battista Biggio
cs.LG
The Lipschitz property of a deep neural network provides a direct measure of its sensitivity to input perturbations and, when explicitly controlled, offers a principled way to limit the propagation of errors and improve robustness. Over the past decade, Lipschitz-bounded layers have been incorporated into increasingly expressive and high-performing deep models, narrowing the gap between empirical robustness and formal, by-design guarantees of...
-
64
Generative World Models Enable Predictive Control of Laser Melt Pool Dynamics
Yiyang Yan, Markus Bambach, Mohamadreza Afrasiabi
cs.LG
World models, which learn how environments respond to actions, are emerging as a powerful paradigm for planning through imagined futures, transforming decision-making across games, robotics and autonomous driving. Bringing this capability to manufacturing could enable process decisions on timescales inaccessible to high-fidelity simulation. Here we introduce a generative world model for localized highly dynamic laser melt pool that predicts...
-
65
Constrained Goal-directed Planar Graph Generation with Grammar-based Reinforcement Learning
Nicolas Hochuli, Lorenzo Miele, Kristina Shea, Tino Stankovic
cs.LG
Planar graphs are central to applications across science and engineering, yet existing generators provide limited support for goal-directed generation under hard structural and geometric feasibility constraints. We propose a dataset-free method for generating planar graph embeddings by combining parametric graph grammars with safe reinforcement learning to optimize generic task-specific objectives while satisfying constraints during...
-
66
RoSA: Rotational Sparse Adaptation for Memory-Efficient Fine-Tuning
Muhammad Azeem Lodhi, Chao Zhou, Rebekka Burkholz
cs.LG
Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers at a time. RoSA freezes lower layers close to the input throughout training and rotates a trainable block over later layers, progressively increasing the number of frozen layers close...
-
67
Few-Shot Prototype Head Adaptation for On-Device ECG Personalization on PSoC~6
Guilherme Silva, Pedro Silva, Gladston Moreira, Eduardo Luz
cs.LG · cs.AI
Wearable and bedside electrocardiogram (ECG) monitors must adapt to patient-specific morphology to maintain arrhythmia detection accuracy across users, yet personalization is typically performed offline and cannot account for individual physiology, electrode placement, or recording drift. On-device adaptation by backpropagation is expensive for microcontroller-class medical devices because it requires an optimizer state, repeated backward...
-
68
Certification-Enhanced Generalization Bounds
Leo Elmecker-Plakolm, Matthew Wicker
cs.LG
We investigate the use of formal methods to provide tight and sound generalization bounds for learning algorithms. By casting the traditional notion of algorithmic stability as a specification to be verified, we demonstrate that recent advances in reachability analysis can yield provable bounds on the generalization of a given model and algorithm on a sample dataset. As sample-specific algorithmic stability is insufficient to bound the usual...
-
69
Parameter Estimation in Machining Dynamics with Regenerative Delay and Nonsmooth Friction using Physics-Informed Neural Networks
Meiyazhagan Jaganathan, Vikram Pakrashi, Aasifa Rounak
cs.LG · nlin.CD
A multi-domain eXtended Physics-Informed Neural Network (XPINN) framework is developed for nonsmooth Delay Differential Equations (DDEs). This is the first implementation to demonstrate the efficacy of partitioning the temporal domain into subdomains of integer multiples of the characteristic time delay and progressively training the associated subnetworks while freezing previously learned parameters. The efficacy of the proposed framework is...
-
70
SO(3)-RoPE for Spherical Transformers
Christian Libner, Chase van de Geijn, Alexander S. Ecker, Maurice Weiler
cs.LG · cs.AI
Spherical data arise in many scientific applications. Often spherical transformers disregard the geometry of the underlying spherical domain, causing distortions and coordinate singularities near the poles. We introduce SO(3)-RoPE, a relative positional embedding that incorporates spherical geometry into transformer attention through unitary SO(3) representations. Our formulation is SO(3)-equivariant and compatible with FlashAttention,...
-
71
LeAVJEPA: A Minimalist Architecture for Audio-Visual Self-Supervised Learning
Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin
cs.LG · cs.CV
Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective. A single early-fusion Vision Transformer processes audio, video, and joint audio-video inputs. Modality dropout treats a missing modality as another view of the same event,...
-
72
Sampling Allocation of LinUCB: Optimal Design Limits in the Small-Gap Regime
Yujie Liu, Vincent Y. F. Tan, Yunbei Xu
cs.LG · stat.ML
We study the sampling allocation of LinUCB in the small-gap regime, where the reward gaps are of order at most $n^{-1/2}$ over the decision horizon $n$. This scaling captures the hard instances underlying worst-case regret lower bounds, for which LinUCB is known to be near optimal up to logarithmic factors in $n$. Using a mean-field perspective, we characterize this allocation through the empirical sampling distribution, a macroscopic object...
-
73
Structured Representation Learning for Behavior Cloning: How can we learn to safely control a nuclear power plant?
Perceval Beja-Battais, Alain Grosset{ê}te, Nicolas Vayatis
cs.LG
Learned models for industrial control are usually judged by aggregate accuracy, but accuracy at the component level does not guarantee safety once it is embedded in the system it is meant to serve. We study this gap on a behavior-cloning task: imitating an expert Nonlinear Model Predictive Control (NMPC) policy for load-following of a Pressurized Water Reactor (PWR), an industrial system with tight safety constraints. We propose a structured...
-
74
Loss-Invariant Projections as Passive Probes of Learned Representations
Akshay Chandrasekhar, Pavlo Melnyk
cs.LG · cs.CV
Learned feature representations in neural networks often contain structure beyond that directly used by the final task output. We study this structure using $\textit{passive probes}$ that apply fixed, untrained, property-independent projections to representations as they evolve during training. We motivate this approach through the task of prediction on $S^2$ where equivalent vector and Hermitian parameterizations reveal an additional...
This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.