cs.LG · 2026-10-06 · No. 135

Machine Learning, 2026-10-06.

74 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

74 entries
  1. 01

    Base Models Can Reason By Taking a Cue From Training Data

    Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min, Alexei A. Efros

    cs.LG · cs.AI · cs.CL

    In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42%...

    arxiv.org/abs/2610.06851 · PDF

  2. 02

    Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

    Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen, Zhengzhong Liu, Eric Xing, Xuezhe Ma

    cs.LG

    Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout...

    arxiv.org/abs/2610.06833 · PDF

  3. 03

    Deep Learning for Sleep Heart Rate Estimation from Accelerometers: Toward Population-Scale Cardiac Insight Without Optical Sensors

    Tanbin Islam Rohan, Pranjol Sen Gupta, Tanusree Debi, Nazmus Sakib

    cs.LG · cs.AI

    Large longitudinal cohorts often contain wrist accelerometry without optical heart-rate sensing, motivating recovery of cardiac information from motion signals already collected during sleep. We present SeqSmoother, a transformer-based temporal corrector for sleep heart rate (HR) estimation from wrist accelerometry. SeqSmoother combines spectral descriptors with an intermediate Nightbeat-derived frequency anchor and a physics-motivated...

    arxiv.org/abs/2610.06823 · PDF

  4. 04

    Private online learning and prediction for Littlestone classes

    Amartya Sanyal

    cs.LG · cs.CR · stat.ML

    We study mistake bounds for differentially private online learning and online prediction under oblivious realisable adversaries. Online learning requires the learner to release a hypothesis at each time step whereas in online prediction, the learner only needs to make predictions without releasing a hypothesis. Using a novel lower bound for private online learning and an upper bound for private prediction, we show that the sample complexity...

    arxiv.org/abs/2610.06822 · PDF

  5. 05

    Block Disentanglement in CRL: Bridging Identifiability and Visual State Estimation

    Emre Acartürk, Pranamya Kulkarni, Puranjay Datta, Karthikeyan Shanmugam, Burak Varıcı, Ali Tajer

    cs.LG · stat.ML

    Causal representation learning (CRL) is the process of recovering causally-related latent variables from high-dimensional observations. As a label-free inference method, CRL is particularly attractive for applications where data labels are unavailable or impractical to obtain. While there has been significant progress in understanding the identifiability guarantees of CRL, such guarantees often hold under highly stylized assumptions, which...

    arxiv.org/abs/2610.06809 · PDF

  6. 06

    H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning

    Wancong Zhang, Basile Terver, Michael Rabbat, Yann LeCun, Randall Balestriero

    cs.LG · cs.RO

    Long-horizon planning with latent world models requires reasoning across timescales and levels of abstraction. Existing task-agnostic JEPA world models predict and plan at a single timescale or with multiple horizons in one shared latent space. We introduce H-JEPA, an end-to-end recipe for training a hierarchy of action-conditioned JEPAs in which each level predicts farther ahead in its own learned latent space. Planning proceeds top-down:...

    arxiv.org/abs/2610.06805 · PDF

  7. 07

    Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution

    Erfan Baghaei Potraghloo, Seyedarmin Azizi, Arya Fayyazi, Saeid Shokoufa, Mehdi Kamal, Souvik Kundu, Massoud Pedram

    cs.LG · cs.AI

    A language model can give a correct answer more probability than any single incorrect answer and still usually sample an incorrect one, because the incorrect answers together hold more probability. The power distribution raises each complete answer's probability to a power above one and renormalizes, shifting probability toward answers the model finds most likely (sharpening). Sampling from it improves reasoning without changing parameters,...

    arxiv.org/abs/2610.06804 · PDF

  8. 08

    Round-Trip KNN Clustering: multiscale hierarchical cluster detection on directed nearest-neighbour graphs

    Eraldo Pereira Marinho, Caetano Mazzoni Ranieri, Fabricio Aparecido Breve

    cs.LG · stat.ML

    We introduce Round-Trip KNN Clustering (RTKNNC), a graph-based method for finding cluster structure at several neighbourhood scales without requiring the number of clusters in advance. Unlike approaches that first make a $k$-nearest-neighbour (KNN) graph undirected, RTKNNC keeps both directions of the neighbour relation: which points a given point selects and which points select it. Incoming selections are treated as weighted votes that help...

    arxiv.org/abs/2610.06795 · PDF

  9. 09

    MatrixFormer: A Foundation Model for Matrix Completion

    Dwaipayan Saha, Jacob Feitelberg, Kyuseong Choi, Raaz Dwivedi, Anish Agarwal

    cs.LG · cs.AI

    Matrix completion underlies problems from tabular imputation to causal inference, yet existing tabular foundation models treat it as entry-by-entry prediction, repeating context for every target and discarding the matrix's two-dimensional structure. We introduce MatrixFormer, a pre-trained matrix-native transformer that predicts a full distribution for every missing entry in a single forward pass. MatrixFormer is trained entirely on synthetic...

    arxiv.org/abs/2610.06751 · PDF

  10. 10

    Hyperbolic Graph Representation Learning: Embed in One Metric, Optimize with Another

    Federico Larroca, Paola Bermolen, Marcelo Fiori, Bernardo Marenco

    cs.LG · stat.ML

    Hierarchical graphs embed in hyperbolic space with lower distortion than in Euclidean space owing to its negative curvature. However, their gradient-based learning is hampered at large radii, where the Poincaré ball and the Lorentz hyperboloid models fail numerically. Polar coordinates avoid this problem, but the hyperbolic metric scales the angular step by the hyperbolic sine of the radius, freezing angular motion. We observe that this...

    arxiv.org/abs/2610.06745 · PDF

  11. 11

    Decoupling Time and Space: A Temporally Conditioned Refinement for EEG Source Imaging

    Marco Morik, Jesse Palarus, Carmen Vidaurre, Klaus-Robert Müller, Shinichi Nakajima

    cs.LG

    Electroencephalography (EEG) offers millisecond temporal resolution, but inferring underlying neural sources is a severely ill-posed spatial inverse problem. While deep learning has advanced spatial reconstruction, current architectures face a critical dilemma: frame-by-frame models discard vital temporal context, whereas full 4D spatiotemporal networks introduce an architectural trade-off between reconstruction accuracy and inference cost....

    arxiv.org/abs/2610.06726 · PDF

  12. 12

    BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models

    Gang Fu, Adel Javanmard, MohammadHossein Bateni, Vahab Mirrokni

    cs.LG · cs.AI

    Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry no topological meaning. We introduce {\bf BRANCH-MoE}, a routing architecture that places \(E\) experts at the leaves of a binary decision tree of depth \(\log_2 E\). At each internal...

    arxiv.org/abs/2610.06725 · PDF

  13. 13

    To Learn is to Wander: Learning Across Graphs and Tasks with Random Walks

    Louis Tichelman, Xingyue Huang, Jinwoo Kim, İsmail İlkan Ceylan

    cs.LG

    Graph foundation models aim to transfer across graphs, feature spaces, relational schemas, and prediction tasks, yet existing approaches typically generalize only within particular graph modalities or tasks. We propose Wander, a graph foundation model designed to operate across these settings within a single pretrained checkpoint. Following the prior-predictive perspective, we formulate graph learning as completion of a partially observed...

    arxiv.org/abs/2610.06694 · PDF

  14. 14

    Adapting prior-data fitted networks for tabular anomaly detection

    Maximilian Bershtman, Niv Cohen

    cs.LG

    While deep features have transformed anomaly detection in images and video, their impact on tabular data has been less substantial, partly due to the limited availability of strong deep representations. Recently, prior-data fitted networks (PFNs) have emerged as a promising source of such representations for tabular data. In this work, we investigate how PFN representations can be adapted and leveraged for anomaly detection. The question is...

    arxiv.org/abs/2610.06693 · PDF

  15. 15

    OVAL: Output-Aware Local Page Bases for KV Cache Retrieval

    Ashkan Shahbazi, Chayne Thrash, Soheil Kolouri

    cs.LG

    Long context inference with large language models becomes increasingly expensive as attention must operate over an ever growing KV cache. Page sparse attention reduces this cost by representing each KV page compactly and retrieving only a subset for each query. Existing retrieval methods are designed to estimate attention scores or page relevance, but their objectives do not directly account for how approximation errors affect the resulting...

    arxiv.org/abs/2610.06686 · PDF

  16. 16

    Aligning Multimodal Patient Evidence with Biomedical Knowledge Graphs for Clinical LLMs

    Jiawen Du, Arshan Ali Khan, Chenhao Zhang, Zachary Plotkin, Li Shen, Qi Long, Yun Li, Can Chen, Tianlong Chen, Nicholas Konz

    cs.LG · cs.CL

    Clinical questions often depend on linking a patient's multimodal evidence to external biomedical knowledge, yet existing predictive systems rarely represent such links explicitly, so they can neither be traced to their evidence sources nor removed to measure their contributions. We present MM-KG (Multimodal Knowledge Graph), which represents heterogeneous, multimodal patient observations and biomedical concepts as separate layers in one...

    arxiv.org/abs/2610.06685 · PDF

  17. 17

    Closing the Context Gap: Activation Alignment for Tabular In-Context Learning

    Yoel Zeldes

    cs.LG · cs.AI

    Tabular foundation models perform in-context learning (ICL) by conditioning predictions on labeled training examples provided as context. Unlike traditional models that separate training from inference, these models must process all training examples in every forward pass, making each prediction expensive. Restricting the number of training examples reduces this cost but substantially degrades performance. Instead of discarding context, we...

    arxiv.org/abs/2610.06679 · PDF

  18. 18

    How Sparse Probability Maps Shape Mixture-of-Experts Routing

    Tomás Brogueira, Marcos Treviso, Miguel Couceiro

    cs.LG · cs.CL

    Mixture-of-experts (MoE) routers typically apply softmax to the router scores and keep the top-K experts, making every token use exactly K experts. Sparsity-inducing probability maps such as sparsemax, alpha-entmax and normmax can adaptively assign exact zeros to selected experts, and therefore appear to offer token-dependent expert participation, even when using the same top-K machinery. In this work, we study whether and how this sparsity...

    arxiv.org/abs/2610.06677 · PDF

  19. 19

    Improved Convergence of Large Stepsize Gradient Descent for Logistic Regression

    Xiaochuan Gong, Ang Li

    cs.LG

    We study gradient descent (GD) with a large constant stepsize for logistic regression on linearly separable data. Existing analysis shows an accelerated rate of $\widetilde{O}(1/\sqrtε)$ to reach loss $ε$ with an aggressive stepsize, although the loss may initially oscillate. Tighter control of the oscillatory dynamics has been available only for two-dimensional data. We prove a substantially faster rate in arbitrary dimension: GD with a...

    arxiv.org/abs/2610.06675 · PDF

  20. 20

    Detecting Nighttime Anomalies from NASA Black Marble Using a Generalized Spatio-Temporally Robust Framework of Machine Leaning Ensembles

    Srija Chakraborty

    cs.LG · cs.CV

    Nighttime lights from NASA's Black Marble product suite capture thermal and light emission signals from anomalous events including fires, volcanic eruptions, and gas flaring. Existing detection approaches rely primarily on thermal bands, limiting sensitivity to weaker signals. We propose a novel machine learning framework that jointly models Black Marble M-band and Day/Night Band (DNB) signals to derive a generalized, spatio-temporally robust...

    arxiv.org/abs/2610.06674 · PDF

  21. 21

    Learning What to Imitate: Entropy-Aware Distribution Mixing

    Juan Garcia Giraldo, Matteo Santelmo, Eduard Durech, Imanol Schlag, Valentina Pyatkin, Antoine Bosselut

    cs.LG

    Small language models are often post-trained as students on reasoning traces from stronger teacher models to efficiently learn new skills. However, token-level imitation on traces that lie far outside the student's expected distribution often produces \textit{confident conflicts}, whereby the student is required to imitate a continuation that it deems unlikely (i.e., low-probability) despite being confident in a different continuation (i.e.,...

    arxiv.org/abs/2610.06671 · PDF

  22. 22

    What Matters for Latent Reasoning with Flow Matching

    Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

    cs.LG · cs.CL

    Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows,...

    arxiv.org/abs/2610.06666 · PDF

  23. 23

    TrustmeWatcher: An Application for Workplace Micro-Sensing and Explainable Well-Being Feedback

    Chengyu Yu, Leon Jacopo Costa, Zoja Anžur, Mohan Li, Gašper Slapničar, Daniil Kirilenko, Martin Gjoreski, Mitja...

    cs.LG · cs.HC

    Workplace sensing studies combine long-running behaviour traces with self-reports, yet the tools that collect those data often sit apart from the interface that returns results. We present TrustmeWatcher, the application built for the TRUST-ME project to connect this work. TrustmeWatcher reuses ActivityWatch's OS-level watchers for computer-activity collection and adds its own application layer. It turns the collected traces into an...

    arxiv.org/abs/2610.06657 · PDF

  24. 24

    The Birkhoff Geometry of Manifold-Constrained Hyper-Connections: Two Channels, Vertex Viscosity, and Sinkhorn as a Retraction

    Xiaoyu Li, Zhizhou Sha, Chiwun Yang

    cs.LG · math.OC · stat.ML

    Hyper-connections widen the residual stream of a Transformer to $n$ parallel streams. Their manifold-constrained version (mHC) mixes the streams at each layer with a doubly stochastic matrix, which it computes by Sinkhorn normalization of exponentiated logits. We give a geometric theory of this design on the Birkhoff polytope. First, a doubly stochastic mixer splits the stream into a mean channel, on which mHC is exactly a residual network,...

    arxiv.org/abs/2610.06653 · PDF

  25. 25

    Considering Context: When World Models Need Context Encoders

    Oleg Smirnov, Sofiane Ennadir, John Pertoft, Bjartur Hjaltason, Sara Karimi

    cs.LG

    Methods for generalization in model-based reinforcement learning typically assume that an agent cannot recover the latent context governing the environment dynamics from its own experience, and therefore supplies it externally. We formalize and test this assumption with \emph{predictive sufficiency}, which quantifies what access to the context adds to next-step prediction under the visitation distribution an agent induces, and separates that...

    arxiv.org/abs/2610.06651 · PDF

  26. 26

    Beyond the Model: The Critical Role of Data Filtering in Clinical Machine Learning

    Noah Subedar, Colin Campbell, Wenjing Zhang, Dan Perri, Sarah Culgin, Andrew Hamilton-Wright

    cs.LG

    Machine learning (ML) studies using clinical data often rely on preprocessing and filtering pipelines before model development. The filtering decisions made in these pipelines can alter the dataset's statistical structure and may artificially reduce or increase the complexity of the prediction task. We argue that filtering choices should be treated as part of the scientific method rather than as a routine preprocessing step. We further...

    arxiv.org/abs/2610.06640 · PDF

  27. 27

    Differentially Private Mixing of Public Datasets Improves Private Learning

    Yufei Chen, Tejumade Afonja, Anvith Thudi, Nicolas Papernot

    cs.LG · cs.AI · cs.CR

    Many machine learning applications involve sensitive data and therefore require training under differential privacy (DP). However, DP training often degrades model utility. In some cases, first pre-training the model on "public" data before finetuning with DP on the sensitive data can reduce the drop in utility. However, the success of this depends on how relevant the selected public dataset is to the sensitive data. We introduce the first...

    arxiv.org/abs/2610.06636 · PDF

  28. 28

    Frozen Factor or Spectral Band? Disentangling Two Choices in Low-Rank LoRA

    Adnan Slimane Ali, Ayoub Belfatmi, David Ngwe Pouth

    cs.LG · cs.AI · cs.CL

    Spectral variants of low-rank adaptation (LoRA) choose both a subspace and which factor to freeze. We separate these choices by freezing the input factor A or output factor B on the top or bottom singular directions of pretrained weights, with learning rates selected separately. At rank 2, the same-band advantage of freezing A is larger than either within-factor band difference on all four task-model pairs with complete comparisons. Freezing...

    arxiv.org/abs/2610.06621 · PDF

  29. 29

    Separators Make Carry Propagation Learnable:The Geometry of Latent Carry in a Multiplication Transformer

    Sama Satariyan, Raphael Cousin, G{é}rard Biau

    cs.LG

    Transformers asked to multiply multi-digit numbers in a single forward pass often fail, and interpretability studies of pretrained language models find arithmetic solved by input-range heuristics rather than by an explicit carry. We train small Llama-style transformers from scratch on 4x4 multiplication without chain of thought and find that the input format is decisive: inserting a space token between digits raises exact-match accuracy from...

    arxiv.org/abs/2610.06605 · PDF

  30. 30

    LinearPFN: Amortized Variable Selection for Linear Models with Interactions

    Louis Schiekiera, Max Zimmer, Christophe Roux, Manuel Arnold, Sebastian Pokutta, Fritz Günther

    cs.LG · stat.ML

    Spike-and-slab regression is a standard Bayesian formulation of variable selection: it returns a posterior distribution over which candidate effects are active rather than a single selected subset, so that every candidate effect carries an inclusion probability. Its cost grows exponentially with the number of candidate effects, so the posterior can be enumerated exactly only when the number of predictors is small. Beyond that reach, the...

    arxiv.org/abs/2610.06580 · PDF

  31. 31

    Empirical Variational Autoencoder

    Kaede Shiohara

    cs.LG

    We present Empirical Variational Autoencoder, a general generative framework for continuous-valued (i.e., non-vector-quantized) sequences. EVA is based on the evidence lower bound of the Variational Autoencoder (VAE) but learns autoregressive latent priors empirically from training data, which can be implemented only by an additional single linear layer on top of VAEs. By replacing the conventional standard-Gaussian constraint with the...

    arxiv.org/abs/2610.06545 · PDF

  32. 32

    A Fine-Grained Analysis of the LoRA Fine-Tuning Landscape with Implications for Data Selection

    Bowen Zhang, Changrui Fang, Xinsong Ma, Jiaye Teng, Ziye Ma

    cs.LG

    Low-Rank Adaptation (LoRA) has become a standard approach for parameter-efficient fine-tuning, yet a fundamental practical question remains unresolved: how should the adapter rank be chosen? An overly small rank may lead to a poorly conditioned optimization landscape, whereas an unnecessarily large rank sacrifices the efficiency that motivates LoRA in the first place. Existing theoretical analyses provide only limited guidance on this...

    arxiv.org/abs/2610.06542 · PDF

  33. 33

    WaveGSSM: Graph Wave State Space Models for Propagating Spatio-Temporal Patterns

    Junyou Zhu, Fenying Cai, Ping Xiong, Christian Nauck, Langzhou He, Chao Gao, Jürgen Kurths, Frank Hellmann

    cs.LG

    Spatio-temporal graph models typically encode each snapshot with a GNN and then connect the resulting representations through a temporal module. This space-then-time design is effective, yet it does not explicitly represent how a pattern moves across the graph. We show empirically that, for a propagating process, the same present field can lead to different futures when its recent rate of change differs, motivating an explicit representation...

    arxiv.org/abs/2610.06540 · PDF

  34. 34

    Conditional Flow Matching for Single-Neuron Electrophysiology: Capturing Multimodal Responses Across Stimuli

    Cameron Schofield, Luca Ghafourpour, Philip H. Wong, Costas A. Anastassiou, Richard E. Turner

    cs.LG

    Neurons of the brain exhibit a rich repertoire of electrophysiology dynamics with the same repeated stimulus eliciting very different voltage responses from the same cell. One common approach in biophysically detailed models is to capture this variability through ensembles of deterministic parametrizations, at a cost of hundreds of thousands of CPU hours. Existing machine learning surrogates inherit the same limitation, where a stimulus is...

    arxiv.org/abs/2610.06520 · PDF

  35. 35

    Xaurora: Generative Weather Forecasting with Denoising Stochastic Interpolants from a Foundation Model Prior

    Eliot Walt, Miltiadis Kofinas, Nikolaj Mücke, Efstratios Gavves, Dim Coumou

    cs.LG

    Deep learning has revolutionised weather forecasting in recent years, especially through atmospheric foundation models, which offer competitive skill for a fraction of the computational costs of classic physics-based models. However, most existing foundation models are deterministic, limiting the generation of large ensembles for accurate uncertainty quantification, extreme weather risk assessment, and long-range weather forecasting....

    arxiv.org/abs/2610.06509 · PDF

  36. 36

    Improving Proactive AI Assistance with Hierarchical Procedural Understanding

    Jin-Seop Lee, TaeYeon Won, SeongJun Jung, JungHoon Kim, Boyang Albert Li, Jin-Young Park, Jaehong Yoon, Jee-Hyong Lee

    cs.LG · cs.CV

    Proactive AI assistants continuously observe a user's activity and decide whether to provide new guidance or remain silent. They should provide appropriate guidance for the task, determine when to provide the next guidance based on task progress, and adjust the guidance level to the user's expertise and needs. Supporting these capabilities requires training and evaluation data that reflect procedural structure and capture how guidance should...

    arxiv.org/abs/2610.06505 · PDF

  37. 37

    FairProp: Fair Node Representation Learning via Differentiable Propagation Layers

    Emmanouil Kariotakis, Aritra Konar

    cs.LG · cs.SI

    Graph neural networks (GNNs) are the standard tool for node representation learning and are increasingly used in high-stakes settings. Their message-passing backbone, however, can amplify topological bias, raising fairness concerns. We study group fairness at the level of downstream predictions for node classification, link prediction, and node regression, and bound the demographic parity gap for an arbitrary number of sensitive groups. Our...

    arxiv.org/abs/2610.06484 · PDF

  38. 38

    Latent Flow Matching for Molecular Graph Generation

    Mathis Goupillon, Roman Bresson, Konstantinos Divriotis, Michalis Vazirgiannis

    cs.LG · cs.AI

    Modern graph generative models typically operate directly in the discrete graph space, explicitly generating node and edge variables, which can become costly as graphs grow. In this paper, we perform generation explicitly on latent representations of entire graphs obtained from a pretrained Variational Autoencoder with high reconstruction fidelity. The generated representations, obtained through flow matching, are then decoded only at the...

    arxiv.org/abs/2610.06468 · PDF

  39. 39

    A Physics-Guided Transformer Framework for Electromigration Analysis in Multi-Segment Interconnects

    Pavlos Stoikos, Anuj Pathania, George Floros

    cs.LG · cs.AI

    As technology scales to smaller nodes, increasing current densities make electromigration (EM) one of the dominant reliability challenges in on-chip interconnects. Accurate transient stress analysis is needed to identify wires susceptible to EM degradation, but applying physics-based solvers across many interconnects remains computationally expensive. This paper proposes a physics-guided transformer framework for fast EM stress prediction in...

    arxiv.org/abs/2610.06464 · PDF

  40. 40

    EMG-FM-Bench: A Comprehensive Benchmark for Foundation Model Transfer and Adaptation on Electromyography

    Tianhao Wu, Xu Wu, Amirmohammad Radmehr, Jiawei Yu, Yi Wu, Phuc Nguyen, Jian Liu

    cs.LG

    Foundation models (FMs) are increasingly being developed for general time series and physiological signals, yet their transferability to downstream physiological tasks remains poorly understood. This question is particularly challenging for electromyography (EMG), where signal distributions vary substantially across users, sensing configurations, acquisition hardware, and downstream tasks. We introduce EMG-FM-Bench, a systematic benchmark for...

    arxiv.org/abs/2610.06450 · PDF

  41. 41

    Time-series Foundation Models for Predictive Control: The Role of Excitation

    Mazen Amria, Jasper Hoffmann, Philipp Bordne, Anna Rothenhäusler, Lilli Frison, Harald Taxt Walnum, Sebastien Gros,...

    cs.LG · eess.SY

    Deploying model predictive control (MPC) requires constructing or identifying a predictive model for each target system. Time-series foundation models (TSFMs) offer an attractive option thanks to strong zero-shot forecasting capabilities across systems. However, low forecast error does not guarantee that a TSFM captures the system's response to the alternative actions considered by the controller. We study this gap using residential heat-pump...

    arxiv.org/abs/2610.06447 · PDF

  42. 42

    ARO: Aligned Representation learning for multi-Omics data

    Amogh Singh, Yash Shah, Chiara D'Ercoli, Arash Mehrjou, Patrick Schwab, Timothy Jones, Pietro Liò

    cs.LG · cs.AI

    The high cost of functional molecular assays, and prevalence of missing modalities and unmatched samples in computational biology, create significant barriers to comprehensive multi-omic profiling, essential for capturing and reasoning over molecules, cells, tissues, and organisms. This work proposes a model that learns meaningful representations from multi-omics cancer data supporting the reconstruction of missing and unpaired modalities....

    arxiv.org/abs/2610.06443 · PDF

  43. 43

    Better Call Reward: Reward Hacking as Strategic Abstention in Legal Reasoning Models

    Subramanyam Sahoo, Justin Shenk

    cs.LG · cs.AI · cs.CL · cs.CY

    What happens when a legal AI model learns to look like a lawyer instead of reasoning like one? We fine tune Qwen3-8B with Group Relative Policy Optimisation (GRPO) against a proxy built from three surface features: citation count, legalese density, and response length. The model does not learn to reason more effectively. It learns to withhold commitment. Across 16 yes or no legal reasoning tasks from LegalBench (N=320), overall accuracy...

    arxiv.org/abs/2610.06439 · PDF

  44. 44

    HeuFouFT: Task-Guided Metaheuristic Coordinate Search for Fourier Fine-Tuning

    Ruiheng Wang, Yubo Hou, Yakun Zhu, Tianle Shen, Tao Wan, Zengchang Qin

    cs.LG · cs.CL

    We introduce Heuristic-Guided Fourier Fine-Tuning (HeuFouFT), a task-guided framework for selecting trainable frequency coordinates in Fourier fine-tuning. Existing uniform and Gaussian band-pass schemes allocate a limited spectral budget through fixed, task-agnostic rules. HeuFouFT instead searches for coordinates using downstream performance. A coarse intensity map from lightweight block-level probes initializes three metaheuristic...

    arxiv.org/abs/2610.06437 · PDF

  45. 45

    Efficient Secure Federated Learning via Information-Theoretically Secure Key Distribution: A Medical Imaging Case Study

    Ivan Donà, Hans H. Brunner, Álvaro Troyano Olivas, Chi-Hang Fred Fung, Momtchil Peev, Giovanni Iacca

    cs.LG · cs.CR

    Federated Learning (FL) enables collaborative training of models across institutions without centralizing sensitive data, making it well-suited for privacy-concerned applications, such as medical imaging. To protect FL model updates during secure aggregation, additive masking is commonly employed. However, its underlying classical key establishment is only computationally secure. On the other hand, physics-based Information-Theoretically...

    arxiv.org/abs/2610.06420 · PDF

  46. 46

    Training-Free Transformer Merging via Sequential Local Operator Alignment

    Akansh Maurya, Ya-Wei Eileen Lin, Stefanie Jegelka, Sebastian U Stich, Rotem Mulayoff

    cs.LG

    Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-key and value-output operators depend on composed matrices, overlooking the functional structure. Moreover, when merging earlier components, downstream components receive different...

    arxiv.org/abs/2610.06415 · PDF

  47. 47

    FlashCart: Fast Cartesian Tensor Products for Equivariant Interatomic Potentials

    Viktor Zaverkin, Payman Goodarzi, Sergey V. Sukhomlinov, Davit Hovhannisyan, Roland Aydin, Martin H. Müser, Mathias Niepert

    cs.LG

    Machine-learned interatomic potentials extend atomistic simulations beyond the length- and timescales accessible to electronic-structure methods. However, the computational cost of equivariant architectures limits the local correlations they can represent in practice and therefore their achievable accuracy. Here we introduce FlashCart, which makes higher-order correlations affordable by combining generated GPU kernels with an architecture...

    arxiv.org/abs/2610.06409 · PDF

  48. 48

    Quantifying the Stability of Multi-Step Reasoning via Error Amplification

    Dongyue Li, Ziniu Zhang, Minxuan Duan, Hongyang R. Zhang

    cs.LG · cs.AI · stat.ML

    We consider the stability of multi-step reasoning processes, which have extensive applications in language models, including chain-of-thought and algorithmic reasoning. While longer sequences of reasoning can improve a model's generation capability at test time, the errors due to intermediate reasoning steps can accumulate in autoregressive generation, and thus grow substantially at the end. In this paper, we ask: What are the key factors...

    arxiv.org/abs/2610.06404 · PDF

  49. 49

    Learning Pareto Stationary Fronts via Single-Pass Backpropagation

    Elina Rojin Celik, Marcos Medeiros Raimundo, Isabel Valera

    cs.LG

    We propose MOSEL (Multi-Objective Stackelberg Efficient Learning), a framework for a posteriori multi-objective optimization (MOO) in deep neural networks that recovers a full front of Pareto stationary solutions at the computational cost of standard single-objective training. MOSEL reformulates the problem as a bilevel optimization problem that leverages network modularity to decouple representation learning from objective-preference...

    arxiv.org/abs/2610.06397 · PDF

  50. 50

    Scaling Down the Scaling Laws: Parameter Efficiency and Compute-Optimal Training in Resource-Constrained Large Language Models

    Joe Dwyer

    cs.LG · cs.AI

    Large language models (LLMs) have achieved substantial performance gains through increases in model size, training data, and computational resources. However, traditional scaling approaches produce diminishing returns, rising financial and environmental costs, and barriers to participation for researchers operating outside large industrial laboratories. This review examines the evolution of LLM scaling theory from empirical scaling laws to...

    arxiv.org/abs/2610.06387 · PDF

  51. 51

    Steering by Influence: Curvature Aware Data Weighting for Activation Steering

    James A. E. Dixon, Stephen J. Roberts, Francesco Quinzan

    cs.LG · cs.AI · cs.CL

    Inference-time steering offers cheap, fine-grained control over a language model's outputs by estimating a concept's representation in activation space and shifting activations towards it. Existing methods build these representations from activation averages over contrastive datasets. These averages incorporate unrelated concepts and noise, and are dominated by a few tokens, meaning the activation transport encodes token-level rather than...

    arxiv.org/abs/2610.06383 · PDF

  52. 52

    Multimodal Deep Survival Analysis for Sinkhole Susceptibility

    Lucas Yuan, Minhee Kim, Zihan Li, Chunli Dai, Sanduni S. Disanayaka Mudiyanselage, Ming Ye, Kani Fu

    cs.LG

    Sinkholes are a widespread geohazard in karst terrain. In Florida, soluble carbonate bedrock, shallow groundwater, and intense rainfall combine to make subsidence both common and spatially heterogeneous. Predicting where and when sinkholes will occur is difficult for two reasons. First, locations without reported sinkholes cannot be directly labeled or sampled as true negative locations. Second, the potential factors governing sinkhole risk...

    arxiv.org/abs/2610.06365 · PDF

  53. 53

    Stability-Shaped Deep Graph Learning

    Junyou Zhu, Langzhou He, Fenying Cai, Christian Nauck, Ping Xiong, Chao Gao, Philip S. Yu, Klaus-Robert Müller,...

    cs.LG

    In deep graph neural networks, increasing depth enlarges the receptive field but often leads to over-smoothing, where node representations tend to align. We develop a unified, mode-wise stability framework for deep GNN propagation that provides a principled characterization of over-smoothing. By interpreting layer depth as time and layer updates as graph-coupled dynamics, over-smoothing can be understood as an undesirable dynamical...

    arxiv.org/abs/2610.06344 · PDF

  54. 54

    Dynamic Minimax Regret Optimization for Robust LLM Post-Training

    Chengbo Zang, Haoyu Dong, Mehmet Kerem Turkcan, Gil Zussman, Zoran Kostic, Javad Ghaderi

    cs.LG

    Modern LLM training increasingly relies on heterogeneous data sources spanning different domains, tasks, preference distributions, and difficulty levels. We study dynamic minimax regret for group-distributionally robust LLM post-training under instantaneous mini-batch-only bandit feedback. The framework views the training as a two-player sampler-optimizer process: a sampler adaptively selects among data sources using bandit feedback, while an...

    arxiv.org/abs/2610.06329 · PDF

  55. 55

    SPDAlign: Interpretable Riemannian Alignment for EEG Forward Modeling Shifts

    Shanglin Li, Shiwen Chu, Okan Koç, Chenyu Liu, Qibin Zhao, Motoaki Kawanabe, Mitsuo Kawato, Yi Ding

    cs.LG

    Electroencephalography (EEG) based brain-computer interfaces enable direct brain-to-device communication for applications such as rehabilitation and communication. However, their practical utility is often limited as the non-stationary nature of the EEG data introduces distribution shifts across domains (e.g., sessions and subjects). Adapting machine learning models to be invariant to these shifts in an unsupervised way, without using costly...

    arxiv.org/abs/2610.06315 · PDF

  56. 56

    Trajectory-Guided Tokenization of Complex CSI for Wi-Fi Sensing

    Ziyi Wang, Kenuo Xu, Jichu Jiang, Yumeng Yang, Zheng Chen, Xiaofei Bai, Muge Chen, Xuyang Chen, Jinglei He, Jannik...

    cs.LG

    Wi-Fi channel state information (CSI) enables contactless presence detection and gesture recognition. Its high-dimensional complex-valued time series require input representations that preserve informative temporal variations during compression. We propose Trajectory-Guided Tokenization (TGT), which combines complex trajectory decomposition with asymmetric attention to construct compact continuous tokens. For each antenna link and subcarrier,...

    arxiv.org/abs/2610.06288 · PDF

  57. 57

    When Are Concept Bottleneck Model Explanations Faithful and Compact?

    Stefano Teso, Emanuele Marconato, Steve Azzolin, Antonio Vergari

    cs.LG · cs.AI · stat.ML

    Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal explainability, we argue they should also be faithful, i.e., not misreport which concepts actually matter. We show that, for widespread CBM architectures, including recent VLM-based...

    arxiv.org/abs/2610.06285 · PDF

  58. 58

    dIon: Fragmentation-Based Invariance for Self-Supervised Learning of Tandem Mass Spectra

    Alfred Nilsson, Joel Lapin, Samuel H. Payne, Mathias Wilhelm, Lukas Käll

    cs.LG

    We introduce a novel invariance for peptide tandem mass spectrometry data, unlocking self-supervised representation learning that improves de novo sequencing of peptides. This invariance exploits the physical relationship between precursor properties (mass and charge) and fragment-ion evidence, without requiring peptide sequence labels. We introduce dIon, which adapts the DINO framework with two latent prediction tasks, both recovering a...

    arxiv.org/abs/2610.06282 · PDF

  59. 59

    What May an Agent Change About Itself? A Containment Floor for Self-Configuring Agent Runtimes

    Sajib Hossain, Moeen Uddin Mahmud

    cs.LG

    Many agent runtimes give the agent a tool for editing its own configuration. Some of that configuration grants abilities, such as enabling a tool. Other parts set the agent's limits: which directories it may write to, who may send it messages, which network address it listens on, how callers authenticate, and the gate that blocks risky writes. If the agent can edit those limits, a single ordinary request can widen them. We study this in a...

    arxiv.org/abs/2610.06274 · PDF

  60. 60

    Ramp Metering Control via Hybrid State Deep Reinforcement Learning in Partially Observable Connected Vehicle Environments

    Youcef Mehamlia, Nadir Farhi, Meriem Bouali

    cs.LG · cs.AI · math.OC

    Freeway on-ramp merges are major sources of congestion, causing significant economic and environmental costs. While Deep Reinforcement Learning (DRL) offers a promising solution for ramp metering, existing approaches rely primarily on aggregated macroscopic data. Connected vehicles (CVs) provide vehicle-level observations that can complement aggregate traffic measurements, but their limited penetration produces incomplete microscopic...

    arxiv.org/abs/2610.06266 · PDF

  61. 61

    OCL-PDE: A Generative Framework for PDE Inverse Problems with Observation-Complementary Latents

    Ding Yang, Chuqi Chen, Chang Ma, Yang Xiang

    cs.LG

    Partial differential equation (PDE) inverse problems are often ill-posed, making fine-scale details difficult to recover. We address this problem by introducing a learned observation-complementary latent representation that preserves reconstruction-relevant information and is combined with the observation to reconstruct the unknown field. Building on this representation, we propose OCL-PDE, a generative framework that encourages the...

    arxiv.org/abs/2610.06259 · PDF

  62. 62

    SimAuthor: Harnessing Foundation Models for Persistent Scientific Simulator Authoring

    Yishan Wang, Ran Piao, Mathias Funk, Aaqib Saeed

    cs.LG

    Foundation models can generate scientific code, but authoring a scientific simulator (an executable program encoding hypotheses about how mechanisms generate observable signals) requires iterative refinement. Scientific adequacy rarely admits a unique implementation or exact test, so simulators must instead be judged against limited real observations. We study this setting as scientific simulator authoring under weak empirical feedback, where...

    arxiv.org/abs/2610.06257 · PDF

  63. 63

    Lipschitz Thinking: Ten Years of Certifiable-by-Design Robust Neural Networks

    Fabio Brau, Giorgio Piras, Maura Pintor, Battista Biggio

    cs.LG

    The Lipschitz property of a deep neural network provides a direct measure of its sensitivity to input perturbations and, when explicitly controlled, offers a principled way to limit the propagation of errors and improve robustness. Over the past decade, Lipschitz-bounded layers have been incorporated into increasingly expressive and high-performing deep models, narrowing the gap between empirical robustness and formal, by-design guarantees of...

    arxiv.org/abs/2610.06252 · PDF

  64. 64

    Generative World Models Enable Predictive Control of Laser Melt Pool Dynamics

    Yiyang Yan, Markus Bambach, Mohamadreza Afrasiabi

    cs.LG

    World models, which learn how environments respond to actions, are emerging as a powerful paradigm for planning through imagined futures, transforming decision-making across games, robotics and autonomous driving. Bringing this capability to manufacturing could enable process decisions on timescales inaccessible to high-fidelity simulation. Here we introduce a generative world model for localized highly dynamic laser melt pool that predicts...

    arxiv.org/abs/2610.06250 · PDF

  65. 65

    Constrained Goal-directed Planar Graph Generation with Grammar-based Reinforcement Learning

    Nicolas Hochuli, Lorenzo Miele, Kristina Shea, Tino Stankovic

    cs.LG

    Planar graphs are central to applications across science and engineering, yet existing generators provide limited support for goal-directed generation under hard structural and geometric feasibility constraints. We propose a dataset-free method for generating planar graph embeddings by combining parametric graph grammars with safe reinforcement learning to optimize generic task-specific objectives while satisfying constraints during...

    arxiv.org/abs/2610.06244 · PDF

  66. 66

    RoSA: Rotational Sparse Adaptation for Memory-Efficient Fine-Tuning

    Muhammad Azeem Lodhi, Chao Zhou, Rebekka Burkholz

    cs.LG

    Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers at a time. RoSA freezes lower layers close to the input throughout training and rotates a trainable block over later layers, progressively increasing the number of frozen layers close...

    arxiv.org/abs/2610.06243 · PDF

  67. 67

    Few-Shot Prototype Head Adaptation for On-Device ECG Personalization on PSoC~6

    Guilherme Silva, Pedro Silva, Gladston Moreira, Eduardo Luz

    cs.LG · cs.AI

    Wearable and bedside electrocardiogram (ECG) monitors must adapt to patient-specific morphology to maintain arrhythmia detection accuracy across users, yet personalization is typically performed offline and cannot account for individual physiology, electrode placement, or recording drift. On-device adaptation by backpropagation is expensive for microcontroller-class medical devices because it requires an optimizer state, repeated backward...

    arxiv.org/abs/2610.06241 · PDF

  68. 68

    Certification-Enhanced Generalization Bounds

    Leo Elmecker-Plakolm, Matthew Wicker

    cs.LG

    We investigate the use of formal methods to provide tight and sound generalization bounds for learning algorithms. By casting the traditional notion of algorithmic stability as a specification to be verified, we demonstrate that recent advances in reachability analysis can yield provable bounds on the generalization of a given model and algorithm on a sample dataset. As sample-specific algorithmic stability is insufficient to bound the usual...

    arxiv.org/abs/2610.06238 · PDF

  69. 69

    Parameter Estimation in Machining Dynamics with Regenerative Delay and Nonsmooth Friction using Physics-Informed Neural Networks

    Meiyazhagan Jaganathan, Vikram Pakrashi, Aasifa Rounak

    cs.LG · nlin.CD

    A multi-domain eXtended Physics-Informed Neural Network (XPINN) framework is developed for nonsmooth Delay Differential Equations (DDEs). This is the first implementation to demonstrate the efficacy of partitioning the temporal domain into subdomains of integer multiples of the characteristic time delay and progressively training the associated subnetworks while freezing previously learned parameters. The efficacy of the proposed framework is...

    arxiv.org/abs/2610.06230 · PDF

  70. 70

    SO(3)-RoPE for Spherical Transformers

    Christian Libner, Chase van de Geijn, Alexander S. Ecker, Maurice Weiler

    cs.LG · cs.AI

    Spherical data arise in many scientific applications. Often spherical transformers disregard the geometry of the underlying spherical domain, causing distortions and coordinate singularities near the poles. We introduce SO(3)-RoPE, a relative positional embedding that incorporates spherical geometry into transformer attention through unitary SO(3) representations. Our formulation is SO(3)-equivariant and compatible with FlashAttention,...

    arxiv.org/abs/2610.06229 · PDF

  71. 71

    LeAVJEPA: A Minimalist Architecture for Audio-Visual Self-Supervised Learning

    Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin

    cs.LG · cs.CV

    Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective. A single early-fusion Vision Transformer processes audio, video, and joint audio-video inputs. Modality dropout treats a missing modality as another view of the same event,...

    arxiv.org/abs/2610.06226 · PDF

  72. 72

    Sampling Allocation of LinUCB: Optimal Design Limits in the Small-Gap Regime

    Yujie Liu, Vincent Y. F. Tan, Yunbei Xu

    cs.LG · stat.ML

    We study the sampling allocation of LinUCB in the small-gap regime, where the reward gaps are of order at most $n^{-1/2}$ over the decision horizon $n$. This scaling captures the hard instances underlying worst-case regret lower bounds, for which LinUCB is known to be near optimal up to logarithmic factors in $n$. Using a mean-field perspective, we characterize this allocation through the empirical sampling distribution, a macroscopic object...

    arxiv.org/abs/2610.06213 · PDF

  73. 73

    Structured Representation Learning for Behavior Cloning: How can we learn to safely control a nuclear power plant?

    Perceval Beja-Battais, Alain Grosset{ê}te, Nicolas Vayatis

    cs.LG

    Learned models for industrial control are usually judged by aggregate accuracy, but accuracy at the component level does not guarantee safety once it is embedded in the system it is meant to serve. We study this gap on a behavior-cloning task: imitating an expert Nonlinear Model Predictive Control (NMPC) policy for load-following of a Pressurized Water Reactor (PWR), an industrial system with tight safety constraints. We propose a structured...

    arxiv.org/abs/2610.06211 · PDF

  74. 74

    Loss-Invariant Projections as Passive Probes of Learned Representations

    Akshay Chandrasekhar, Pavlo Melnyk

    cs.LG · cs.CV

    Learned feature representations in neural networks often contain structure beyond that directly used by the final task output. We study this structure using $\textit{passive probes}$ that apply fixed, untrained, property-independent projections to representations as they evolve during training. We motivate this approach through the task of prediction on $S^2$ where equivalent vector and Hermitian parameterizations reveal an additional...

    arxiv.org/abs/2610.06195 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.