cs.LG · 2026-10-01 · No. 130

Machine Learning, 2026-10-01.

83 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

83 entries
  1. 01

    Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis

    Tian Xia, Minghao Liu, Yiqing Liang, Laixi Shi, Jiayun Wang

    cs.LG · cs.CL · cs.CV

    Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus...

    arxiv.org/abs/2609.40361 · PDF

  2. 02

    Semifactual Credit-Augmented Policy Optimization

    Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding, Zifan Cheng, Wenqi Shao, Qiaosheng Zhang, Yue Zhang

    cs.LG · cs.AI · cs.CL

    Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token...

    arxiv.org/abs/2609.40360 · PDF

  3. 03

    Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text

    Dulhan Jayalath, Oiwi Parker Jones

    cs.LG · q-bio.NC

    We find that major reported improvements in decoding words from non-invasive brain recordings are largely reproducible without any brain data. In the influential work of d'Ascoli et al. (2025), time series of brain activity from subjects perceiving continuous speech are segmented into fixed-length windows starting at each word. A neural network then generates predictions for all of the words in a sentence together. Neighbouring windows...

    arxiv.org/abs/2609.40359 · PDF

  4. 04

    Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

    Razan El Mais, Ali Chehab, Ibrahim Issa, Razane Tajeddine

    cs.LG

    Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains...

    arxiv.org/abs/2609.40335 · PDF

  5. 05

    Scaling Laws for Looped Mixture of Experts

    Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

    cs.LG · cs.AI · cs.CL

    Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its...

    arxiv.org/abs/2609.40316 · PDF

  6. 06

    Compression Footprints as Security Signals for Model-Poisoning Defense in Federated Learning

    Sachi Shome, William Eiers

    cs.LG

    Lossy compression is widely used in Federated Learning (FL) but is generally treated as an error source, while conventional poisoning defenses inspect update geometry. In this work, we instead treat the compressor's response as a security signal: the input-dependent distortion and payload behavior induced by lossy compression can expose differences between honest and attack-generated updates. We introduce the concept of a \emph{compression...

    arxiv.org/abs/2609.40312 · PDF

  7. 07

    Disentangling Computation in Multi-Task Neural Networks with the Green's Operator

    James Hazelden

    cs.LG · q-bio.NC

    How is computation organized and reused across tasks and time in a trained recurrent network? Most analyses emphasize the geometry of neural activity, dynamical motifs, or local perturbation growth. We instead study the network's global first-order perturbation response. The finite-horizon Green's operator maps perturbations at each source along a trajectory to their downstream state-space responses and therefore directly represents...

    arxiv.org/abs/2609.40292 · PDF

  8. 08

    PMosFM: Preconditioned Manifold Matching for One-Step Physics-Constrained Generation

    Zhangyong Liang, Haibin Ling

    cs.LG

    Physics-constrained generative models aim to generate physical fields that match a target distribution and satisfy prescribed constraints. However, enforcing these constraints often increases sampling costs through iterative corrections or training costs through residual optimization and trajectory unrolling. To address this issue, we introduce \textbf{P}reconditioned \textbf{M}anifold \textbf{o}ne-\textbf{s}tep \textbf{F}low...

    arxiv.org/abs/2609.40287 · PDF

  9. 09

    cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

    Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh

    cs.LG · cs.AI · cs.CL

    Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their...

    arxiv.org/abs/2609.40284 · PDF

  10. 10

    OpenTSLM TeeMoE: A Unified Time-Series Language Model for Forecasting, Contextual Prediction, and Reasoning

    Tony Chen, Timo Stoffregen, Maxwell Xu, Thomas Kaar, Martin Maritsch, Geremia Pompei, Nicolas Zumarraga, Robert...

    cs.LG

    Real-world time-series applications increasingly require models that can handle time series forecasting, context-conditioned prediction, and language-based temporal reasoning. Yet current time-series foundation models remain fragmented across these capabilities: numerical specialists often provide the strongest forecasts, while language-based models offer broader contextual understanding and analysis. A central challenge is to unify these...

    arxiv.org/abs/2609.40265 · PDF

  11. 11

    Distribution Matching Distillation for Continuous Diffusion Language Models

    Paul Le Van Kiem, Dario Shariatian, Umut Simsekli, Alain Durmus

    cs.LG · cs.CL · stat.ML

    Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student's output parameterization to the resulting gradient estimators and yields two methods with the same student architecture and...

    arxiv.org/abs/2609.40235 · PDF

  12. 12

    PhantomEnvironments: Training LLM Agents in Fictional Worlds

    Anmol Kabra, Swathi Saravana Selvam, Albert Gong, Chao Wan, Christian Belardi, Dongyoung Go, Katie Z. Luo, Kilian Q....

    cs.LG · cs.AI · cs.CL

    Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules,...

    arxiv.org/abs/2609.40221 · PDF

  13. 13

    Near-Linear Accuracy Bounds for Moreau--Yosida Unadjusted Langevin Sampling

    Yuchen Xin, Zhihua Zhang

    cs.LG

    We establish near-linear accuracy bounds for the classical Moreau--Yosida unadjusted Langevin algorithm (MYULA). The target is $π\propto e^{-f-g}$, where $f\in C^2(\mathbb{R}^d)$ is $m$-strongly convex with Lipschitz gradient and $g$ is convex and globally Lipschitz. Under an explicit parameter-dependent step-size condition, we bound the invariant-measure bias relative to the Moreau-smoothed target by $\widetilde O(h)$, with only logarithmic...

    arxiv.org/abs/2609.40193 · PDF

  14. 14

    Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves

    Sohail, Sarkar, Shakuntala Baichoo

    cs.LG · cs.CL · math.ST · stat.ML

    Sampling several answers and keeping the one a verifier scores highest is one of the simplest ways to buy accuracy at test time. Its effect is reported as a scaling curve: accuracy against the number $k$ of sampled answers. The curve is cheap to draw and expensive to trust. A budget read off it is chosen after looking at every point, so only a band that covers all budgets at once protects the choice, and on a 100-question benchmark a fixed...

    arxiv.org/abs/2609.40190 · PDF

  15. 15

    MANET-GNN: Learned Decentralized Optimization of Power Allocation in Multi-Channel MANETs

    Tomer Alter, Nir Shlezinger, Michael Segal

    cs.LG

    MANETs enable flexible infrastructure-less wireless connectivity in dynamic and resource-constrained environments. As modern MANETs exploit multiple frequency channels and support heterogeneous traffic patterns, decentralized transmit-power allocation becomes increasingly challenging. We develop a unified learned optimization framework for decentralized power allocation in dynamic multi-hop, multi-channel MANETs. We formulate a constrained...

    arxiv.org/abs/2609.40170 · PDF

  16. 16

    Role-Adaptive Policy Optimization for Offline Reinforcement Learning

    Seonvin Cho, Soohyun Choi, Songnam Hong

    cs.LG

    Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic bootstrapping, yet methods such as TD3+BC couple these roles through a shared policy. We propose Role-Adaptive Policy Optimization (RAPO), which adapts policy-update coefficients according to their roles in value...

    arxiv.org/abs/2609.40149 · PDF

  17. 17

    From Spectra to Joint Schedules in LLM Pre-training: 3+3(+2) Scaling-Law Regimes

    Yichen Wang, Fanghui Liu, Yudong Chen

    cs.LG

    Power-law learning curves are often treated as fixed properties of a model and its data, although learning-rate and batch-size schedules can change the observed loss. We study this dependence in noisy online SGD with linear random features. Conditional on the representation, an exact Volterra equation separates two response components: a forcing term that propagates unresolved target error and a memory kernel that propagates stochastic-error...

    arxiv.org/abs/2609.40148 · PDF

  18. 18

    Policy Iteration Is Not Strongly Polynomial for Deterministic Markov Decision Processes: The Price of Algorithmic Anarchy

    Han Zhong, Yinyu Ye

    cs.LG · cs.DS · math.OC

    We establish an exponential iteration lower bound in the number of states for Howard's policy iteration on deterministic discounted Markov decision processes, with at most two actions per state. This rules out strong polynomiality of Howard's policy iteration when the discount factor is part of the input and yields an exponential separation from the simplex method with Dantzig's pivoting rule, which is proved to be strongly polynomial on this...

    arxiv.org/abs/2609.40147 · PDF

  19. 19

    From DNA Design to DNA Slimming: Auditable Agentic Discovery of a Deletion-Only Designer

    Joel Shor

    cs.LG · cs.NE · q-bio.GN

    Compact regulatory DNA can free up space in vector payloads, reduce synthesis and assay burden, and expose which sequence features drive predicted activity. Yet most model-based nucleic-acid designers optimize fixed-length sequences through substitutions; they do not ask which bases of an existing functional element can be removed while retaining predicted activity. We define the task of sequence slimming as selecting an exact-length,...

    arxiv.org/abs/2609.40143 · PDF

  20. 20

    Game-Guided Skill Discovery through Self-Play for Playable Agent Control

    Seungeun Rho, Jeonghwan Kim, Xue Bin Peng, Sehoon Ha

    cs.LG · cs.AI · cs.RO

    We present Game-Guided Skill Discovery (GGSD), a framework that uses self-play in games to discover motor skills that are directly playable by humans. Playable skills provide a compact abstraction for controlling embodied agents through a small set of learned behaviors rather than low-level actions. To be effective, these skills should be semantically distinct, interpretable, and expressive; properties that existing unsupervised...

    arxiv.org/abs/2609.40137 · PDF

  21. 21

    Prototype-Rule Neurosymbolic Regularization for Rank-Constrained Tensor Neural Networks under Label Scarcity

    Eftychios Protopapadakis, Konstantinos Makantasis, Konstantinos M. Giannoutakis

    cs.LG · cs.CV

    Rank-constrained tensor neural networks reduce the parameterization of high-order inputs, but they do not explicitly constrain class geometry in the learned representation. This study investigates whether a differentiable prototype-rule can provide a complementary inductive bias for Rank-R tensor learning under limited supervision. The proposed framework augments the Rank-R objective with prototype-based regularization and optionally fuses...

    arxiv.org/abs/2609.40131 · PDF

  22. 22

    Learning Functional Subspaces for Neural Network Compression

    Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder, Ole Winther, Yann LeCun, Ravid Shwartz-Ziv, Zeynep Akata

    cs.LG · cs.CL

    Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keeping the matrices dense, and thus efficient on standard hardware. Existing methods, however, choose the subspace to remove from each weight matrix with local closed-form criteria: activation energy, layer-wise reconstruction error, or a quadratic approximation of the loss. These criteria ignore how...

    arxiv.org/abs/2609.40127 · PDF

  23. 23

    Scalable Cox Regression via Grouped Risk Sets and Sharper LogSumExp Rates

    Elizaveta Iashchinskaia, Egor Gladin

    cs.LG · math.OC

    Motivated by the computational challenges of large-scale Cox regression, we study stochastic minimization of LogSumExp objectives over large sets. Mini-batch normalizer estimates generally yield biased gradients. We instead use a softplus surrogate that introduces one auxiliary scalar per normalizer and admits unbiased single-sample gradients. For smooth convex LogSumExp objectives, we prove an $O(T^{-1/2})$ averaged objective bound,...

    arxiv.org/abs/2609.40120 · PDF

  24. 24

    Beyond Model Ranking: Regime Diagnosis for Distributional-Statistical Misspecification in Industrial Time-Series Forecasting

    Pengyu Nie, Chenglang Xu, Yaoshi Chen, Chaogan Ren, Wei Hu, Chao Yang, Jiangong Zhang

    cs.LG

    Time-series forecasting models achieve strong benchmark performance but exhibit severe systematic bias in industrial deployments. This train--deploy gap is conventionally attributed to temporal-structural errors or distribution shifts. We characterize a complementary source that these explanations overlook: canonical losses embed fixed statistical priors, while industrial demand mixes benign and pathological regimes---zero-inflation,...

    arxiv.org/abs/2609.40117 · PDF

  25. 25

    Replay on Demand: An Emergent Curriculum for Balancing Adaptation and Forgetting in Continued Pretraining

    Lukas Thede, Shengzhuang Chen, Stefan Winzeck, Matthias Bethge, Zeynep Akata, Jonathan Richard Schwarz

    cs.LG

    Continued pretraining enables language models to adapt to new domains and knowledge, but often at the cost of forgetting previously acquired capabilities. Replay can mitigate this trade-off, but fixed replay mixtures allocate training independently of the model's actual retention needs. We introduce Replay on Demand (RoD), which instead derives the replay allocation from the model's learning dynamics. RoD jointly prioritizes adaptation...

    arxiv.org/abs/2609.40089 · PDF

  26. 26

    Accelerated Algorithm for Sparse Regularized Partial Optimal Transport

    Khoa Nguyen, Dung T. Nguyen, Thong Huynh, Hoang-Hiep Nguyen-Mau, Anh Nguyen, Minh Ngoc Dinh, Juho Kannala

    cs.LG

    Partial Optimal Transport (POT) extends the classical optimal transport problem by relaxing the strict mass conservation constraint, enabling its use in a wide range of real-world applications. In many of these settings, sparse transport plans are preferred for their interpretability and computational benefits. While smooth and strongly convex regularizers - such as quadratic or elastic net - have been vastly used in various machine learning...

    arxiv.org/abs/2609.40075 · PDF

  27. 27

    Inference Auctions

    Keegan Harris, Siddharth Prasad, Asher Trockman, Nika Haghtalab, Michael I. Jordan

    cs.LG · cs.AI · cs.GT

    When inference demand exceeds available compute capacity, model providers must decide which requests should be served first. Users have different tolerances for delay from an LLM API, but current priority pricing schemes compress these differences into coarse fixed-price service tiers. We design an inference auction that allows users to bid for faster service. Our auction allocates priority in an economically efficient way without sacrificing...

    arxiv.org/abs/2609.40070 · PDF

  28. 28

    LARC: Low-Rank Adaptive Residual Connections for Learning in Frozen Models

    Junyi Zou, Avrova Donz

    cs.LG · cs.CL

    Low-Rank Adaptive Residual Connections (LARC) give a frozen model a compact numerical state that can learn from feedback. The map $h+BAh$ adds a low-rank correction to a hidden representation. A slow state $ρ$ learns starting factors across tasks; a private fast state $Φ$ copies them, changes with feedback, and resets to the trained initialization. This report specifies an input-side realization of the numerical policy carrier in...

    arxiv.org/abs/2609.40063 · PDF

  29. 29

    Gromov-Wasserstein Distillation for Inductive Multi-View Embedding

    Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, Charles Casimiro Cavalcante

    cs.LG · stat.ML

    Gromov-Wasserstein multidimensional scaling (GW-MDS) learns low-dimensional representations from relational data but remains transductive, providing no explicit mapping for unseen samples. We introduce an inductive framework based on barycentric distillation. A GW-MDS teacher learns a latent support and an optimal transport plan from the training data, and barycentric projection converts the resulting coupling into sample-aligned targets. A...

    arxiv.org/abs/2609.40047 · PDF

  30. 30

    Efficient Active Auditing of Multi-Group Fairness with Bias Probes

    Ayoub Ajarra, Debabrota Basu

    cs.LG · cs.AI · cs.CY · stat.AP · stat.ML

    Over the past decade, Machine Learning (ML) has been trained under dual objectives: minimizing prediction error via Empirical Risk Minimization (ERM) while controlling unfairness bias. In practice, however, fairness-aware training often yields limited improvements over standard ERM, making reliable post hoc auditing essential. Existing auditing approaches for black-box models either rely on model reconstruction --exposing systems to...

    arxiv.org/abs/2609.40034 · PDF

  31. 31

    Fenchel Tilting: Weighted Correction for Efficient Finetuning of Generative Models

    Maksim Bobrin, Maksim Zhdanov, Dmitry Dylov

    cs.LG · cs.AI

    Adapting a pretrained generative model to an arbitrary preference expressed as a utility function underlies reward alignment, guided design, and constraint satisfaction, enabling diverse applications. Existing fine-tuning methods trade off generality against computational cost: they either restrict the family class of supported preferences to keep optimization simple or preserve generality at the expense of efficiency. We introduce Fenchel...

    arxiv.org/abs/2609.40030 · PDF

  32. 32

    DashVMC: Real-Time Discrete World Model Control in Geometry Dash

    Florent Tariolle, Florian Yger

    cs.LG

    World-model agents are usually evaluated in simulators that can wait for the policy; live games impose the opposite constraint, requiring capture, prediction, and action before the next frame. We present DashVMC, which learns a compact, action-conditioned world model from approximately two hours of recorded Geometry Dash gameplay. To test whether the learned dynamics are actionable, a controller is initialized by behavioural cloning (BC) and...

    arxiv.org/abs/2609.40003 · PDF

  33. 33

    Learning to Explain While Planning: Rule-Aligned Diffusion Planning for Autonomous Driving

    Jiaxi Ye

    cs.LG

    Diffusion planners exhibit strong capabilities in generating multimodal trajectories. However, existing methods primarily rely on expert demonstrations to fit trajectory distributions, learning statistical correlations among scenes, behaviors, and trajectories without explicitly modeling driving rules. In long-tail scenarios where expert data are scarce, the lack of behaviors to imitate may lead to trajectories that violate safety or...

    arxiv.org/abs/2609.39995 · PDF

  34. 34

    What Limits Recursive Reasoning Models: Optimization, Architecture and Test-Time Scaling

    Yuliana Shakhvalieva, Dmitrii Kharchev, Viacheslav Bezrukov, Inessa Fedorova, Dmitry Bocharov, Ivan Oseledets,...

    cs.LG · cs.AI

    Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few parameters and makes them strong on algorithmic tasks. Such compact solvers are natural candidates for tools that an LLM can call on narrow algorithmic subproblems. However, existing models such as HRM, TRM and URM differ in architecture, gradient propagation and training procedure...

    arxiv.org/abs/2609.39967 · PDF

  35. 35

    Reliability-Aware Checkpoint Selection for Domain Generalization

    Jinshi Liu, Jiahao Li, Pan Liu, Yanfeng Li, Rui Qian, Zhao Tong, Yue Sun, Tao Tan

    cs.LG · cs.CV

    Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable probabilities on unseen target domains. Source-target distribution shifts can alter accuracy rankings, while accuracy alone does not measure predictive probability quality. We identify an empirical selection opportunity within fixed training trajectories: reselecting among checkpoints with...

    arxiv.org/abs/2609.39934 · PDF

  36. 36

    RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures

    Yuyang Wu, Yufeng Du, Hao Peng

    cs.LG · cs.CL

    Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns...

    arxiv.org/abs/2609.39929 · PDF

  37. 37

    TRACE: Trajectory Selection for Parallel Scaling of Search Agents

    Qisheng Zhou, Zhen Xiong, Qiaoyu Tan

    cs.LG · cs.AI

    Parallel search may generate a correct answer that final-answer voting fails to select. We formulate this consolidation stage as trajectory selection and introduce TRACE (Trajectory Ranking with Aggregated Cross-Rollout Evidence), a lightweight learned selector that ranks completed trajectories using the search evidence behind their answers. TRACE preserves individual query and evidence occurrences, connects rollouts through shared content or...

    arxiv.org/abs/2609.39912 · PDF

  38. 38

    Patient-Centered Treatment Planning for Chronic Multimorbidity: A Hierarchical Reinforcement Learning Framework for Preference Modeling

    Nafiseh Payani, Soham Das, G. Anthony Wilson, Anahita Khojandi

    cs.LG

    Patient preference, defined as a patient's demonstrated willingness and capacity to adhere to clinical recommendations, is a primary determinant of therapeutic effect yet remains structurally absent from existing computational treatment planning models. We address this gap by presenting patient-centered factored-action hierarchical option-critic (FAHOC), a hierarchical reinforcement learning (HRL) framework that jointly learns high-level...

    arxiv.org/abs/2609.39911 · PDF

  39. 39

    Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning?

    Syed Nazmus Sakib, Abdul Monaf Chowdhury, Nafiul Haque, Shifat E Arman, Md Mehedi Hasan

    cs.LG · cs.AI

    Goal-conditioned reinforcement learning (GCRL) relies heavily on how target goals are represented to the policy. While recent methods encode goals via temporal distance, occupancy, or controllability, it remains unclear how much downstream performance actually depends on representation quality. We study this in offline GCRL by constructing an exact temporal-distance goal representation in deterministic mazes. We then systematically corrupt...

    arxiv.org/abs/2609.39901 · PDF

  40. 40

    Shared Weights, Selected Computations: How Looped Transformers Route What Each Loop Does

    Jiaju Wu, Yi Hu, Muhan Zhang

    cs.LG

    Looped Transformers repeatedly apply the same set of Transformer layers, giving them a recurrent architecture for latent computation. Their strong performance on iterative reasoning and length-generalization tasks suggests an appealing explanation: recurrence may provide an inductive bias that lets the model reuse a learned algorithm across loops. However, weight sharing alone does not imply that every loop performs the same operation. This...

    arxiv.org/abs/2609.39892 · PDF

  41. 41

    Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dynamics Manifolds

    Kevin Yu, Tao Guo, Constantinos Antoniou, Panagiotis Angeloudis

    cs.LG · cs.RO · eess.SY

    Neural trajectory predictors can reach low prediction error while violating dynamics, actuator limits, or state constraints, especially when controls are unobserved and dynamics are partially specified. We introduce the Markovian Dynamics Enforcer (MaDE), a time-invariant post-hoc operator mapping state-transition proposals onto a learned feasible dynamics manifold, trained on feasible states without ground-truth controls. For each transition...

    arxiv.org/abs/2609.39888 · PDF

  42. 42

    Algorithmic Recourse Under Competition

    Shahin Jabbari

    cs.LG · cs.AI

    Algorithmic recourse provides individuals who have received undesirable outcomes from machine learning models with suggestions for minimum-cost improvements to achieve the desired outcome. A central assumption when computing recourse is that the decision rule remains fixed throughout the recourse implementation phase. We challenge this assumption in settings where individuals compete for limited resources. In such settings, widespread...

    arxiv.org/abs/2609.39877 · PDF

  43. 43

    Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing

    Kemou Li, Qizhou Wang, Yue Wang, Fengpeng Li, Zhuan Shi, Negar Rostamzadeh, Golnoosh Farnadi, Masashi Sugiyama, Jiantao Zhou

    cs.LG

    Open-weight LLMs are released not only as fixed products but also as substrates for downstream fine-tuning. This openness, however, creates legal and ethical risks because users may misuse fine-tuning to instill illicit knowledge or enable hostile operations. Model providers therefore need apre-release defense against such acquisition, motivating the problem of preemptive unlearning. Unlike retrospective unlearning, which removes capabilities...

    arxiv.org/abs/2609.39866 · PDF

  44. 44

    Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models

    Stipe Frković, Metod Jazbec, Christian A. Naesseth

    cs.LG

    Masked diffusion language models (dLLMs) have shown strong potential for faster inference through parallel token generation when combined with confidence-based samplers. However, recent work has shown that such methods can defer unmasking high-entropy fork positions at which multiple plausible continuations exist. This results in reduced generation diversity, as shown by worse pass@k scaling, and limits gains obtainable from RL post-training....

    arxiv.org/abs/2609.39859 · PDF

  45. 45

    Dimension-Free Rank Lifting from Random Hyperplane Arrangements

    Luca Becchetti, Matteo Russo, Ruben Skorupinski

    cs.LG · math.PR

    We study the width required for a randomly initialized hidden layer of a neural network to achieve rank lifting. Namely, given a dataset $X \in \mathbb{R}^{m \times d}$ of $m$, $d$-dimensional input vectors separated by an angle of at least $θ$, we consider the random feature matrix $σ(XR)$, where $R$ is standard Gaussian. For positively homogeneous nonpolynomial activations, which include sign, Heaviside, ReLU, and ReLU powers among others,...

    arxiv.org/abs/2609.39855 · PDF

  46. 46

    Predicting Multi-View Rashomon Representation: Can We Learn Where Models Disagree?

    Mingyue Ma, Zongbo Han, Changqing Zhang, Guangyu Wang

    cs.LG

    Foundation models are increasingly adopted across a wide range of applications, often serving as core blocks within AI systems. Yet different foundation models may encode the same input from multiple different views, leading to substantial representation disagreement, which we term Rashomon Representation. Such disagreement often signals inputs that a given model encodes in a way inconsistent with other models, offering a valuable yet...

    arxiv.org/abs/2609.39848 · PDF

  47. 47

    Dynamic LoRA-Experts and Prototype-Ensemble Matching for Class-Incremental Learning

    Hongwei Zhao, Rui Liu, Yansong Liu

    cs.LG

    Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. Parameter-efficient fine-tuning with pre-trained models reduces parameter overhead but can suffer from cumulative interference and suboptimal alignment between inference samples and specialized modules. We propose Dynamic LoRA-Experts and Prototype-Ensemble Matching (DLEPEM), a two-stage rehearsal-free framework. DLEPEM...

    arxiv.org/abs/2609.39839 · PDF

  48. 48

    Fast Regularized Policy Mirror Descent with One-Step TD Updates

    Qipei Chen, Wenye Li, Yule Sun, Ke Wei

    cs.LG · math.OC

    Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or increasingly accurate policy evaluation. We analyze PMD coupled with a persistent critic advanced by one temporal-difference (TD) update. For finite discounted MDPs, we establish global linear convergence in value for exact coordinate-wise Bellman updates, with any positive constant actor stepsize...

    arxiv.org/abs/2609.39837 · PDF

  49. 49

    RainAtlas: A Multi-Continental Dataset for Precipitation Downscaling

    Pierre-Louis Lemaire, Luca Schmidt, Wietze Suijker, Alex Hernandez-Garcia, David Rolnick

    cs.LG

    Extreme rainfall events are increasing in intensity and frequency as climate change accelerates. While kilometer-scale precipitation forecasts are critical for supporting local decision-making, the limited availability of high-resolution precipitation observations hinders their accuracy, especially in under-resourced regions. Machine learning models are widely used to downscale precipitation data to km-scale, but their application to unseen...

    arxiv.org/abs/2609.39833 · PDF

  50. 50

    Should I stay or should I show? Learning to selectively disclose information

    Carlotta Giacchetta, Alessando Bogani, Cesare Barbera, Giovanni De Toni, Michele Caprio, Andrea Pugnana, Andrea Passerini

    cs.LG

    In many high-stakes settings, human decision-makers can acquire support information before making a decision. However, acquiring information is costly, and disclosure may fail to improve human decisions or may even impair them. We tackle this problem by studying selective disclosure, i.e., the problem of learning when to reveal support information to a human decision-maker under a budget constraint. We first show that the optimal policy is a...

    arxiv.org/abs/2609.39818 · PDF

  51. 51

    Beyond Accuracy: Prefix-Invariant Realizations of Low-Precision Fast Matrix Multiplication

    Shuxiao Xie, Shuyang Xie, Yuan Cao, Dezhi Ran, Wei Yang, Tao Xie

    cs.LG

    Fast matrix multiplication saves multiplications through exact cancellation, but rounding sums that mix token rows can leave contributions from later tokens in earlier language model outputs. This threatens prefix invariance, which multiple-choice likelihood scoring relies on: a scored likelihood must depend only on its allowed prefix. On Qwen2.5-14B-Instruct, two fast FP8 realizations repaired to ordinary-looking accuracy still change the...

    arxiv.org/abs/2609.39816 · PDF

  52. 52

    Backward-State Policy Is Part of the Learning Algorithm

    Shuxiao Xie, Shuyang Xie, Dezhi Ran, Wei Yang, Tao Xie

    cs.LG

    Low-precision training rounds tensors that the backward pass reads again, often for several gradients; each use can read the forward's rounded value, the original, or a new random rounding. This backward-state policy looks like a memory and precision detail, settled by copy accuracy and final loss. We argue that it is part of the learning algorithm, and that neither check shows whether it is right. Copy accuracy does not decide the outcome:...

    arxiv.org/abs/2609.39813 · PDF

  53. 53

    A Comprehensive Benchmark of Source-Free Universal Domain Adaptation on Time Series Representations

    Romain Mussard, Fannia Pacheco, Maxime Berar, Paul Honeine, Gilles Gasso

    cs.LG

    Source-Free Universal Domain Adaptation (SF-UniDA) extends Universal Domain Adaptation by removing access to source data at adaptation time while still handling label-set mismatches between domains. Despite growing interest in this setting for image data, no benchmark exists for time series, which are more challenging. We present the first SF-UniDA benchmark on time series. In addition, we provide the first study of pretrained foundation...

    arxiv.org/abs/2609.39810 · PDF

  54. 54

    RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models

    Chengzhu Bao, Xianglong Yan, Tianao Zhang, Jiaqi Chen, Shaoqiu Zhang, Yulun Zhang

    cs.LG

    Post-training quantization (PTQ) has become a widely adopted technique for reducing the memory footprint and inference cost of large language models (LLMs). However, recent studies reveal that when applied to reasoning models, PTQ not only degrades reasoning performance but also exacerbates overthinking, leading to longer reasoning trajectories. These issues may offset the efficiency gains expected from lower-precision inference. Existing...

    arxiv.org/abs/2609.39801 · PDF

  55. 55

    Finite-Horizon Fisher Memory in Two-Sided Power-Bounded Recurrent Systems

    Jeonghoon Lee

    cs.LG

    We analyse allocation, admission and post-write retention in finite-horizon linear-Gaussian noisy recurrent memories. At every horizon, the directional Fisher memory $M_n$ satisfies $\operatorname{tr}M_n=N$: non-normality redistributes information but cannot raise its spherical average, while normal carriers satisfy $M_n=I$. For bi-power-bounded carriers, we derive uniform $1/n$ lag bounds, identify the limit of $M_n$ with the inverse of the...

    arxiv.org/abs/2609.39800 · PDF

  56. 56

    Probabilistic Adversarial Training

    Andi Zhang, Xingyu Zhao, Siddartha Khastgir

    cs.LG · cs.AI · stat.ML

    Building on a probabilistic perspective in which adversarial examples arise from the overlap between a distance-based distribution $p_{\mathrm{dis}}$ and a victim-classifier-induced distribution $p_{\mathrm{vic}}$, we start from a simple intuition: adversarial examples become harder to generate when these two distributions are pushed apart, as their overlap becomes smaller, thereby increasing robustness. This intuition naturally motivates a...

    arxiv.org/abs/2609.39798 · PDF

  57. 57

    TopTimeNet: Topologically-assisted time-series classification model

    Sharareh Sayyad, Sophia Bazzi

    cs.LG · cs.CG · math.DS · nlin.CD · physics.app-ph

    Distinguishing periodic from chaotic dynamics in a time series is a fundamental challenge in both physics and engineering. Yet, end-to-end learned architectures must discover both a representation and a decision boundary from data, at substantial cost. We introduce TopTimeNet, which decouples these tasks: a fixed, non-learned stage extracts a $42$-dimensional geometric and topological descriptor from Takens delay embeddings and persistent...

    arxiv.org/abs/2609.39792 · PDF

  58. 58

    Pseudo-Label-Triggered Retraining from Forecast Errors for Online Time Series Forecasting

    Yeryeong Kwak, Yoo-Min Jung, Jonghun Park

    cs.LG · cs.AI

    Real-world time series forecasting systems operate under non-stationary data streams, where forecasting performance may degrade over time. Although retraining can recover the performance, it incurs non-trivial computational and operational costs. Under limited deployment resources, the key challenge is therefore not only how to retrain but also when to retrain. While existing retraining policies often rely on indirect indicators such as drift...

    arxiv.org/abs/2609.39789 · PDF

  59. 59

    CORD: Learning Reusable Degradation Representations Across Heterogeneous Physical Systems

    Haibo Li, Zhiguo Zeng

    cs.LG

    Can heterogeneous physical degradation systems benefit from joint pretraining and move beyond system-specific prognostics toward reusable cross-system representation learning? CORD combines type-specific observation interfaces with a shared degradation backbone. Its two self-supervised objectives learn at complementary scales: Intra-Observation Structure Modeling (ISM) captures structure within observations, while Inter-Observation Dynamics...

    arxiv.org/abs/2609.39784 · PDF

  60. 60

    GraphMAS: A Systematic Benchmark of Multi-Agent Coordination for Graph Learning

    Jiayi Yang, Yifang Chen, Yuanfu Sun, Xinyan Ge, Qiaoyu Tan

    cs.LG

    LLM-based multi-agent systems coordinate specialized reasoning through aggregation, interaction, and adaptive control, yet their potential for graph learning remains unexplored. Graph learning is a natural setting for such systems because useful evidence may arise from heterogeneous local, long-range, global structural, and semantic perspectives whose relevance varies across instances. Existing LLM-based graph learning approaches primarily...

    arxiv.org/abs/2609.39777 · PDF

  61. 61

    Riemannian Flow Models with Reinforcement Learning for Molecular Crystal Structure Prediction

    Thomas Egg, Harry Winston Sullivan, Maya M. Martirossyan, Philipp Höllmer, Cheng Zeng, Adrian Roitberg, Mingjie Liu,...

    cs.LG · physics.comp-ph

    Crystal structure governs material properties, making crystal structure prediction (CSP) a fundamental problem in materials science. Generative models are a promising approach for solving this problem, but the prevalence of polymorphism, coupled with large unit cells and complex packing geometry, makes the molecular CSP task challenging for existing models. To address this, we introduce Coarse-Grained Open Materials Generation (CG-OMatG), an...

    arxiv.org/abs/2609.39773 · PDF

  62. 62

    How Does Local Landscape Geometry Evolve in Language Model Pre-Training?

    Zhanpeng Zhou, Yuhan Sun, Bingrui Li, Jinbo Wang, Huaijin Wu, Lei Wu, Junchi Yan

    cs.LG · cs.AI

    The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape...

    arxiv.org/abs/2609.39767 · PDF

  63. 63

    Revisiting On-policy Adversarial Black-Box Distillation: Calibrating Groupwise Reward Geometry for Effective Advantage Construction

    Xiao Cui, Mo Zhu, Yulei Qin, Yuze Wu, Wengang Zhou, Houqiang Li

    cs.LG · cs.CV

    Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the critic provides rewards for GRPO-based student policy optimization over the student's sampled responses. However, GRPO computes...

    arxiv.org/abs/2609.39757 · PDF

  64. 64

    Validity-Preserving Hierarchical RL for Joint Routing and Switch Placement in EDA

    Dorian Gailhard, Ugo Lecerf, Enzo Tartaglione, Donatello Conte, Jhony H. Giraldo

    cs.LG

    Routing and switch placement are fundamental combinatorial optimization problems in chip design, requiring the joint optimization of routing topology and physical placement under strict structural, geometric and logical constraints. Existing approaches typically rely on carefully engineered heuristics that incorporate strong problem-specific biases to navigate the enormous space of possible designs. In this work, we introduce a hierarchical...

    arxiv.org/abs/2609.39749 · PDF

  65. 65

    The Nixtlaverse: An Open-Source Ecosystem for Forecasting

    Olivier Sprangers, Max Mergenthaler Canseco, Marco Peixeiro, Saul Caballero Ramirez, Mariana Menchero García,...

    cs.LG

    Large forecasting applications often combine statistical, machine-learning, and neural models. These families solve the same problem but differ in fitted state, training procedures, and how they parallelize work. Forecasting software must therefore either hide these differences behind a single estimator interface, or keep the families in separate packages, forcing users to rewrite data preparation and evaluation for every package. We present...

    arxiv.org/abs/2609.39741 · PDF

  66. 66

    Stable Transformers for Graph Generation

    Luca Miglior, Alessio Gravina, Davide Bacciu

    cs.LG

    Graph generative models increasingly rely on Graph Transformers (GT) to capture complex dependencies among nodes and edges. While deeper architectures should provide greater expressive capacity and a broader receptive field, their effectiveness can decline with depth: repeated self-attention progressively contracts node representations, impeding information flow and gradient propagation. We analyse this phenomenon from a dynamical systems...

    arxiv.org/abs/2609.39739 · PDF

  67. 67

    CNCGEN: A Dataset and Framework for Machining Process Planning and Toolpath Generation from B-rep Models

    Xiaolei Zhou, Boyi Lin, Yuchao Feng, Jianwei Zheng

    cs.LG

    Learning to generate machining process plans and toolpaths from B-rep CAD requires coupling discrete operation decisions with continuous tool motion as the workpiece evolves. Correctly predicting an operation sequence does not by itself ensure correct material removal, because each toolpath acts on the stock left by preceding cuts. We formulate this problem around persistent manufacturing objects: object identity determines the target of an...

    arxiv.org/abs/2609.39738 · PDF

  68. 68

    A library for differentiable signal processing and machine learning on the sphere

    Thorsten Kurth, Max Rietmann, Mauro Bisson, Andrea Paris, Alberto Carpentieri, Jean Kossaifi, Anima Anandkumar,...

    cs.LG · physics.ao-ph · physics.geo-ph

    The two-dimensional sphere embedded in three-dimensional Euclidean space S2, plays a central role in a variety of scientific and engineering domains, including geophysics, planetary science, geodesy, atmospheric physics, quantum chemistry, cosmology, and virtual reality, among many others. As machine learning increasingly permeates these fields, the demand grows for robust tools that process and model functions on the sphere, while respecting...

    arxiv.org/abs/2609.39737 · PDF

  69. 69

    GFD-OPD: Guidance-Folded On-Policy Distillation of Diffusion Models Across Scales

    Zhenxing Zhang, Jiayan Teng, Wenxu Wu, Zhuoyi Yang, Jiazheng Xu, Wendi Zheng, Jie Tang, Dan Guo, Meng Wang

    cs.LG · cs.AI

    On-policy distillation (OPD) has demonstrated two important capabilities in language models: compressing large teachers into smaller students and merging expert models into a single model. Existing diffusion OPD, however, mostly focus on the latter, with teachers and students sharing the same backbone and scale. We investigate large-to-small diffusion opd from large teachers to a small student and find that the standard recipe fails. To find...

    arxiv.org/abs/2609.39692 · PDF

  70. 70

    NodeGround: A Node Classification Benchmark in the Graph Foundation Model Era

    Jinmo Lee, Dooho Lee, Minho Jeong, Jaemin Yoo

    cs.LG

    Can a pretrained graph model replace training and tuning a separate predictor for each dataset? Answering this requires evaluating prediction quality alongside computational cost. We present NodeGround, a node classification benchmark that puts graph foundation models (GFMs) and dataset-specific supervised learning under a common evaluation framework. The benchmark spans 51 datasets and evaluates six GFMs alongside 15 supervised methods under...

    arxiv.org/abs/2609.39673 · PDF

  71. 71

    Graph Residual Conjugate Diffusion: SNR-Equalized Heat Flow for Graph Signals

    Jinwei Li, Daniel Tenbrinck

    cs.LG

    Diffusion models generate data by reversing a forward corruption process that typically approaches a simple Gaussian prior. Recent work has extended this framework to signals supported on fixed graphs, e.g., road-network traffic and sensor-network measurements. Many graph signals have nonuniform spectral energy, whereas isotropic corruption adds the same conditional noise variance to every graph-frequency mode. Driving all modes to near-zero...

    arxiv.org/abs/2609.39658 · PDF

  72. 72

    From Modes to Memories: Characterizing the Scale-Space Dynamics of Diffusion Models

    Cristina López Amado, Marco Fumero, Francesco Locatello

    cs.LG · cs.AI · cs.CV

    Diffusion models are typically viewed as stochastic processes that transform noise into data. We take a complementary perspective: a diffusion model defines a family of deterministic dynamical systems indexed by noise scale. At each fixed scale $σ$, we treat the denoiser as a self-map and study its dynamics. For an exact denoiser, fixed points correspond to critical points of the smoothed data density, while attractors correspond to its...

    arxiv.org/abs/2609.39648 · PDF

  73. 73

    Beyond Uniform Compression: Budgeted Transmission Allocation for Extreme Federated Learning

    Pengfei Li, Mohammad Khalil

    cs.LG

    Federated learning faces severe communication bottlenecks when clients upload high-dimensional model updates. Existing methods often compress these updates uniformly across all layers. This uniform approach ignores the heterogeneous value of different parameter blocks and wastes limited bandwidth on insensitive layers. To address this issue, we propose Layer-wise Budgeted Adaptive Transmission (LBAT). LBAT reframes federated communication...

    arxiv.org/abs/2609.39646 · PDF

  74. 74

    RiboUnmix: Learning Shared Translational Dynamics from Biased and Noisy Ribo-seq Measurements

    Gabriele Martino, Denis Skibinski, Ivo L. Hofacker, Sebastian Tschiatschek

    cs.LG

    Ribosome profiling (Ribo-seq) measures ribosome distributions along mRNAs, but observed occupancy profiles also contain experiment-specific distortions and stochastic variability. Consequently, models that accurately predict measured profiles may reproduce technical effects rather than recover the underlying biology. We ask whether jointly modeling datasets collected under different experimental conditions can reveal shared,...

    arxiv.org/abs/2609.39644 · PDF

  75. 75

    Free Everywhere, Exact on Trees: PPO's Dropped Correction Buys Sample Efficiency Under Aggressive Reuse

    Nima H. Siboni

    cs.LG · cs.AI

    Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribution rather than the improved policy's own. The substitution makes the objective estimable from the behavioral policy's rollouts but adds a bias growing with policy divergence, hence the trust region or clip, and hence no reuse of a batch far off-policy. We show that under history-injective...

    arxiv.org/abs/2609.39634 · PDF

  76. 76

    Towards Better Exploration in Sequential Test-Time Scaling

    Joseph Rance, Fabio Pizzati, Juil Sock, Woody Bayliss, Marc Górriz Blanch, Philip Torr, Adel Bibi

    cs.LG

    Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to solve in a single attempt. In contrast, sequential methods build on previous answers to access new ideas, yet so far have not...

    arxiv.org/abs/2609.39632 · PDF

  77. 77

    PEG-Tab: Sampling-Time Record Repair and Release Control for Tabular Synthesis

    Pengfei Li, QinYi Liu, Mohammad Khalil

    cs.LG

    Pretrained tabular generators can reproduce training records even when aggregate utility remains high. When retraining is unavailable or too costly, sampling and release are the remaining intervention points. We present PEG-Tab (Post-Training Energy Guidance for Tabular Synthesis), a post-training repair and release-control framework for frozen tabular generators. For each generated row, a generator-native operator creates two alternatives. A...

    arxiv.org/abs/2609.39630 · PDF

  78. 78

    Certification-Based Differentially Private Learning

    Mihnea Ghitu, Matthew Wicker

    cs.LG

    Differential privacy (DP) in machine learning is typically achieved by adding noise to model parameters (private learning) or to model outputs (private prediction). Recent work uses formal methods, namely abstract interpretation, to provide tighter privacy guarantees, but only for private prediction in classification settings. In this work, we investigate the use of formal methods as a general tool for tighter privacy analysis. First, we...

    arxiv.org/abs/2609.39629 · PDF

  79. 79

    MIND: Marginal-Invariant Neural Dependency Diffusion for Mixed-Type Tabular Generation

    Pengfei Li, Mohammad Khalil

    cs.LG

    This paper proposes MIND, a marginal-invariant neural dependency diffusion model for mixed-type tabular data. MIND does not directly learn the joint distribution in the original heterogeneous feature space. Instead, it first maps different variable types into a unified latent dependency space via column-wise marginal transport. A conditional diffusion model then learns cross-column relationships. Copula-tangent denoising separates known...

    arxiv.org/abs/2609.39628 · PDF

  80. 80

    Parameterization method of reservoir properties for ensemble-based data assimilation using intermediate latent space of StyleGAN

    Marcio A. Sampaio, Paulo H. Ranazzi, Martin J. Blunt

    cs.LG · cs.AI

    Ensemble smoothers are the most successful and efficient techniques currently available for history matching. However, because these methods rely on Gaussian assumptions, their performance is severely degraded when the prior geology is described in terms of complex facies distributions (non-Gaussian). In this way, for these methods, we need to apply efficient parameterization techniques. Currently, the most efficient methods for performing...

    arxiv.org/abs/2609.39626 · PDF

  81. 81

    Hybrid Methods for Robust Tabular Data Imputation

    Jinwei Li, Michelle Bruch, Daniel Tenbrinck

    cs.LG

    Missing data are a fundamental challenge in statistical analysis and machine learning, as the choice of imputation method substantially impacts downstream inference. In this work, we propose two hybrid imputation methods called NuclearForest and SoftForest, which combine nuclear-norm-based low-rank initialization using Singular Value Thresholding (SVT) and SoftImpute, respectively, with a non-iterative Random Forest refinement. For the...

    arxiv.org/abs/2609.39613 · PDF

  82. 82

    Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum

    Hanyng Peng, Hui Wang, Yue Yu

    cs.LG

    Practical Muon maintains momentum and performs a small, fixed number of Newton--Schulz iterations separately for each parameter matrix, often with a Nesterov correction. We analyze these layer-wise finite-step updates jointly on a coupled nonconvex objective, rather than replacing them by exact polar factors or one global orthogonalization. Under gradient-dependent $(\mathcal L_0,\mathcal L_1,q)$-smoothness and conditionally unbiased...

    arxiv.org/abs/2609.39595 · PDF

  83. 83

    Robust Transfer Learning for Paper ECG Recognition

    Yinghao Xie, Zhenbang Dai, Haojun Wang, Jinyu Cai, Fabio Bonassi, Hongwu Chen, Johan Sundström, Jiawei Li, Antônio H. Ribeiro

    cs.LG · cs.AI

    Paper ECG recognition is challenging because real-world ECG images vary in layout, physical artifacts, and label availability. We introduce RobECG-CL, a rank-aware contrastive learning framework for robust paper ECG representation learning. Starting from standard 12-lead ECG recordings, we construct progressively degraded paper ECG views with heterogeneous layouts and train the model to balance same-recording invariance with degradation-aware...

    arxiv.org/abs/2609.39581 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.