cs.LG · 2026-09-09 · No. 110

Machine Learning, 2026-09-09.

57 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

57 entries
  1. 01

    Learning Length-Extrapolatable Recurrent Models

    Hanwen Jiang

    cs.LG · cs.CL

    Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their training horizon. Classical analyses emphasize gradients that vanish or explode along temporal paths. However, dense per-token losses can still train a shared recurrent rule despite severe decay, showing that decay alone does not determine whether learning fails. We instead study state credit:...

    arxiv.org/abs/2609.09157 · PDF

  2. 02

    NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

    Tobias Susetzky, Raphael Rehms, Dmitrii Seletkov, Özgün Turgut, Michelle Espranita Liman, Lisa Steinhelfer, Rickmer...

    cs.LG · cs.AI

    The digitization of healthcare has generated vast, longitudinal, and multimodal patient records over a lifetime, yet fully exploiting these data to represent and predict patient state trajectories remains a critical challenge. Current AI models often struggle to capture the complex, irregular temporal dynamics and inherent stochasticity of real-world multimodal patient data. Existing AI approaches for modeling longitudinal patient records are...

    arxiv.org/abs/2609.09140 · PDF

  3. 03

    Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation

    Jiacheng Xu, Feng Chen, Xiuneng Xu, Bo An

    cs.LG · cs.CL

    Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on unlabeled test-time tasks with canonical answers, but this breaks down for code generation because programs cannot be compared by surface form and therefore do not directly provide a usable training signal. To make TTRL applicable to code generation, we propose probe-driven TTRL, which constructs output-free probe inputs from the...

    arxiv.org/abs/2609.09135 · PDF

  4. 04

    Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks

    Xiaoyu Li, Zhizhou Sha, Jiaojiao Jiang, Junbin Gao, Andi Han

    cs.LG · stat.ML

    An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input-dependent sparsity in the one-hidden-layer ReLU model of Awasthi et al. (COLT 2024). For width $s$, at most $k$ active units per input, and effective weight and bias bounds $W,B$, every size-$m$ sample in the class's fixed radius-$R$ input domain satisfies $\mathcal{R}(S)\le...

    arxiv.org/abs/2609.09130 · PDF

  5. 05

    When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay

    Hasan Amin, Wei-Kai Chang, Rajiv Khanna

    cs.LG

    Normalization renders large parts of neural networks effectively scale invariant, inducing a hidden feedback loop in which learning-rate schedules and weight decay interact through the parameter norm to control the effective step taken by the optimizer. We show that this interaction is governed by an exact discrete-time law: a single scalar quantity captures all schedule and decay forcing, while norm growth induces an opposing geometric...

    arxiv.org/abs/2609.09116 · PDF

  6. 06

    Curriculum Learning as Transport: Understanding Curricula with Wasserstein Geodesics

    Changho Shin, David Alvarez-Melis

    cs.LG

    Curriculum learning is governed by several coupled design choices---how difficulty is defined, how examples are ordered, how much exposure each level receives, and how quickly training moves across levels---making it hard to isolate what actually helps. We present Wasserstein curriculum paths, a simple transport-based framework that decouples these factors by representing curricula as trajectories of training distributions over discrete...

    arxiv.org/abs/2609.09099 · PDF

  7. 07

    ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR

    Tommy Sha, Skylar Zhai, Siqi Zhao

    cs.LG · cs.AI

    In reinforcement learning with verifiable rewards (RLVR) trained with group relative policy optimization (GRPO), the KL-free reward-advantage term studied here depends on within-group reward variation. If all rollouts in a group are correct or all are wrong, their group-relative advantages are identically zero; these zero-advantage silent groups provide no reward-advantage gradient, yet uniform sampling spends 39% of a run's rollouts on them....

    arxiv.org/abs/2609.09075 · PDF

  8. 08

    Multi-Task Learning for Sparsely-Labeled Time Series: A Case Study on Cold-Hardiness Modeling

    Aseem Saxena, Paola Pesántez-Cabrera, Jonathan Magby, Markus Keller, Alan Fern

    cs.LG

    We present a real-world case study of multi-task learning (MTL) for temporal process modeling from limited data with temporally sparse labels. Specifically, we investigate multi-task learning for the important agricultural problem of predicting grape cold hardiness, which is the temperature at which lethal freezing occurs. Cold hardiness changes in response to weather and is difficult to measure directly in the field. Thus, growers rely on...

    arxiv.org/abs/2609.09062 · PDF

  9. 09

    PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games

    Ryan Truong, Lance Ying, Samuel J. Gershman, Kazuki Irie

    cs.LG

    While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existing ones to support new features, has been a laborious process requiring extensive hand-coding. Here we present PlayTrain, an RL framework that combines the abilities of large language models (LLMs) to robustly generate JavaScript (JS) games from a minimal human prompt, and an efficient pipeline...

    arxiv.org/abs/2609.09059 · PDF

  10. 10

    Training-Free Task Vectors for LLM Behavioral Control

    Gabriel J. Perin, Lucas Boscaini, André Araujo, Nina S. T. Hirata

    cs.LG · cs.AI

    Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. However, this reliance on fine-tuning makes discovering such directions costly and limits the practicality of post-training model editing. To address this limitation, we introduce Training-Free Task Vectors (TFTVs), a novel method...

    arxiv.org/abs/2609.09054 · PDF

  11. 11

    Do Reasoning Representations Help Humans Evaluate LLM Outputs?

    Jaewoo Lim, Sungbok Shin, Sanghyun Hong

    cs.LG · cs.HC

    Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations as human-facing interfaces rather than proxies for model reasoning ability. We conduct a controlled human study of six...

    arxiv.org/abs/2609.09038 · PDF

  12. 12

    Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks

    Mariia Drozdova, Stéphane Liem Nguyen, François Fleuret

    cs.LG · cs.AI

    Denoising Diffusion Probabilistic Models (DDPMs) generate samples by starting from noise and repeatedly denoising while keeping each update close to the current noisy state. This behavior is effective in many continuous domains, but its role is less clear for globally constrained discrete tasks, such as Sudoku, graph connectivity, Latin squares, and N-queens. In such settings, early discrete errors can be difficult to undo. As a result,...

    arxiv.org/abs/2609.09009 · PDF

  13. 13

    Physics-Informed Deep Learning for False Ventricular Tachycardia Alarm Reduction in the ICU

    Athanasios Papastathopoulos-Katsaros, Alexandra Stavrianidi, Zhandong Liu

    cs.LG · eess.SP

    False ventricular tachycardia (VT) alarms are a leading contributor to alarm fatigue in intensive care units. We propose a deep learning framework combining a 1D SE-ResNet with ICU-realistic data augmentations and a physics-informed auxiliary reconstruction task based on the three-element Windkessel hemodynamic model, implemented as a differentiable forward simulation. By requiring the network's latent representation to produce...

    arxiv.org/abs/2609.08992 · PDF

  14. 14

    Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling

    Arman Adibi, Alireza Jafari, Mohammad Ghavamzadeh, Hadi Daneshmand

    cs.LG · cs.AI · stat.AP · stat.CO · stat.ML

    A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates. Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression. We prove that in-context learning extends further to \emph{data generation}:...

    arxiv.org/abs/2609.08981 · PDF

  15. 15

    GraphFAS: A Distributed System for Automated Graph Feature Generation and Selection in Industrial Transaction Networks

    Yice Luo, Yun Zhu, Xi Chen, Yongchao Liu, Xintan Zeng, Chengying Huan, Kai Zhang, Jinrui Zhang, Juelu Zhang, Jiajun Zheng

    cs.LG · cs.AI · cs.DC

    Industrial fraud detection often relies on costly expert-crafted features that overlook graph-structured relational signals, while GNNs often do not meet the interpretability and deployment requirements of financial risk control. We propose GraphFAS (Graph Feature Automated Selection), a distributed feature selection procedure based on Boruta that bridges this gap through: (1) a non-parametric graph feature generation module that constructs...

    arxiv.org/abs/2609.08970 · PDF

  16. 16

    The BatchNorm Illusion: Diagnosing Normalization Artifacts in Machine Unlearning Evaluation

    Aaryaman Kalani, Murari Mandal, Dhruv Kumar, Mohan Kankanhalli, Yash Sinha

    cs.LG

    Approximate machine unlearning aims to remove the influence of specific training data from a trained model without retraining from scratch. We identify a previously undocumented confound in how unlearning is evaluated on BatchNorm-based architectures: a single forward pass over retain data, an operation that modifies no weight, can deterministically rewrite the model's normalization state and reverse the apparent surface-metric forgetting. We...

    arxiv.org/abs/2609.08901 · PDF

  17. 17

    Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems

    Zhihao Wang, Ruichen Wang, Ruohan Li, Lei Ma, George Hurtt, Xiaowei Jia, Gengchen Mai, Shaowen Wang, Yiqun Xie

    cs.LG · cs.AI

    Machine learning emulators have become essential for accelerating expensive Earth-system simulations, but most existing approaches remain passive forecasters: they reproduce simulator trajectories under prescribed forcings without an explicit interaction mechanism for user-specified interventions. This limits their use in interactive scientific workflows and Earth-system digital twins, where users often need to explore how a system would...

    arxiv.org/abs/2609.08855 · PDF

  18. 18

    Length Generalization for Transformers via Compression

    Georg Zetzsche, Hongjian Jiang, Andy Yang, Pascal Bergsträßer, Marco Sälzer, David Chiang, Anthony W. Lin

    cs.LG · cs.FL · cs.LO

    Recent advancements in transformer length generalization theory enable us to reliably predict when a transformer can learn to solve a task. In particular, the C-RASP hypothesis (a formalized version of the so-called RASP-l conjecture) posits that transformers length-generalize on a task if and only if a solution is expressible in the C-RASP language. While this hypothesis has strong empirical validation, theoretical problems arise from the...

    arxiv.org/abs/2609.08851 · PDF

  19. 19

    Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

    Youngrok Park, Sangmin Bae, Hojung Jung, Jongwoo Ko, Yunseon Choi, Young Jin Kim, Pashmina Cameron, Aaron Courville,...

    cs.LG · cs.CL

    Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student....

    arxiv.org/abs/2609.08798 · PDF

  20. 20

    Adaptive Anisotropic Attention for Axis-Structured Signals

    Mahir Jain, Parshva Runwal, Aditya Ray Mishra, Arvasu Kulkarni, Sandeep Singh, Siddharth Panwar

    cs.LG · cs.AI

    Dense self-attention treats all token pairs as equally plausible before learning, an interaction-isotropic prior that can be mismatched to structured signals. For structured, low signal-to-noise ratio (SNR) signals such as EEG, dependencies are organized along the electrode and time axes, and this uniform prior exposes each token to many irrelevant interactions. We introduce Adaptive Anisotropic Attention (AAA), which splits attention into...

    arxiv.org/abs/2609.08788 · PDF

  21. 21

    PAC-Bayesian Bounds for Learning Partially Observed Stochastic Linear Time-Invariant State-Space Systems with Inputs and Sub-Gaussian Noise

    Mihaly Petreczky, Mohamad Al Ahdab, John Leth

    cs.LG · math.OC

    In this paper we derive a Probably Approximately Correct (PAC)-Bayesian error bound for partially observed linear time-invariant (LTI) stochastic dynamical systems in state-space form with inputs and sub-Gaussian noise. Such bounds are widespread in machine learning, and they are useful for characterizing the predictive power of models learned from finitely many data points. The bound derived in this paper relates the expectation of...

    arxiv.org/abs/2609.08740 · PDF

  22. 22

    BAFF: Bid-Aware Filter Family for Mitigating Training Data Interference in RTB A/B Tests

    Jeonglyul Oh, Ikkyu Choi, Inseop Youn, Youngjae Kim

    cs.LG

    In online A/B tests for real-time bidding (RTB), control and treatment models are typically trained on a shared serving log that includes data generated by the counterpart model. This shared-log training biases each model's training data through two channels: the counterpart model may have selected a different ad from the ad-candidate pool (ad-ranking disagreement) and may have bid a different price (bid-pricing disagreement), potentially...

    arxiv.org/abs/2609.08725 · PDF

  23. 23

    Chimaera: A Mixture-of-Graph-Experts Architecture for Cross-Task and Cross-Dataset Graph Learning

    Jonathan Frank, David Richerby, Ansgar Scherp

    cs.LG

    Designing foundation models for graphs is challenging due to the irregular structure of graphs and the different sizes and characteristics of embeddings. Chimaera integrates mixture-of-experts with graph foundation models (GFM). It integrates different GFM architectures, such as graph prompts and linear GNN models. Large language models are used to generate embeddings, and experts can be trained and combined following different strategies,...

    arxiv.org/abs/2609.08709 · PDF

  24. 24

    Hyperparameter Scaling Laws Across MoE Sparsity

    Changxin Tian, Kunlong Chen, Jia Liu, Ziqi Liu, Zhiqiang Zhang, Jun Zhou

    cs.LG · cs.AI · cs.CL

    Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyperparameter transfer challenging. In this work, we show that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count...

    arxiv.org/abs/2609.08690 · PDF

  25. 25

    HOPE: Heterophily-Aware Open-Set Node Classification with Pseudo-Extrapolation

    Yumeng Dai, Yue Tan, Yixin Liu, Chenxu Wang, Pinghui Wang, Tao Qin

    cs.LG

    Standard open-set node classification methods rely on the homophily assumption, where connected nodes share labels. However, real-world graphs are often heterophilic, exposing the limitations of current methods and posing new challenges to open-set node classification. On the one hand, cross-class connectivity causes representations from different known or unknown classes to become intertwined after aggregation, undermining their...

    arxiv.org/abs/2609.08685 · PDF

  26. 26

    Neither Adversarial Training Nor Purification: Emergent Adversarial Robustness from Oscillatory Predictive Learning

    Mohammed-Yassine Habibi, Klea Ziu, Martin Takáč, Makoto Yamada

    cs.LG · cs.AI

    Adversarial robustness in computer vision is still largely achieved through adversarial training or test-time adversarial purification, both of which introduce significant computational overhead by generating adversarial examples during training or performing iterative denoising at test time. We study whether empirical robustness can instead emerge from architectural and representation-learning inductive biases. We introduce Oscillatory...

    arxiv.org/abs/2609.08683 · PDF

  27. 27

    MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models

    Xuanming Cui, Shlok Kumar Mishra, Wentao Bao, Aashu Singh, Zihao Wang, Xiangjun Fan, Jun Xiao, Ser-Nam Lim, Jianpeng Cheng

    cs.LG · cs.AI

    Universal multimodal embedding (UME) increasingly demands encoder's capacity for handling a broad range of tasks and modalities with increased complexity. Prior scaling methods either increase the representation size, retrieval effort, or scales the encoder into a heavy multimodal LLM. Recent works, such as Think-Then-Embed (TTE), explore scaling via reasoning tokens. However, embedding models are hard to scale up: increasing parameters...

    arxiv.org/abs/2609.08663 · PDF

  28. 28

    Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR

    Youngjun Yu, Sanghwan Jang, Hwanjo Yu

    cs.LG · cs.AI · cs.CL

    Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the model's intrinsic reasoning coverage (pass@k) due to limited exploration during training. To address this, we optimize the structural design of train-time rollouts to enhance pass@k. Our analysis identifies three key design...

    arxiv.org/abs/2609.08650 · PDF

  29. 29

    SUN: Reaching for Novelty in Reinforcement Learning

    Wenyan Yang, Arsenii Mustafin, Dominik Baumann, Joni Pajarinen, Simone Parisi

    cs.LG · cs.AI · cs.RO

    Exploration in reinforcement learning (RL) remains a fundamental challenge. Recent goal-conditioned RL strategies (which select goals to encourage broader state coverage) have shown promising results, but none scores a goal by novelty and reachability jointly: the two signals are traded off by hand, applied in sequence, or one is neglected outright. In this paper, we introduce a reachability-aware goal-selection framework that explicitly...

    arxiv.org/abs/2609.08642 · PDF

  30. 30

    Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

    Oleksandr Cherednichenko, Roman Klypa

    cs.LG · cs.AI

    Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. While proprietary systems exhibit reliable safety controls, their underlying methodologies and trade-offs remain largely undisclosed. Achieving comparable security in open-weight models remains a persistent challenge, as post-trained variants frequently suffer from over-refusal and degraded general quality. To...

    arxiv.org/abs/2609.08634 · PDF

  31. 31

    Leveraging contextual events on structure-aware next activity prediction

    Alessandro Mele, Claudia Diamantini, Domenico Potena

    cs.LG · cs.AI

    Predictive process monitoring aims at forecasting various aspects of running processes. Among the different tasks, next activity prediction represents the most extensively investigated. However, only a limited number of existing approaches explicitly encode contextual information, i.e., the environmental conditions in which the process is executed, typically modeled through event log attributes or aggregated measures. In this paper, an...

    arxiv.org/abs/2609.08622 · PDF

  32. 32

    Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families

    Zhongxuan Liu, Sicheng Zhou, Hongzhi Wang

    cs.LG

    Benchmark scores describe what a checkpoint can do now, but they do not determine how it will respond to the next training episode. We measure this missing state by branching four short, standardized, target-independent micro-interventions from the same checkpoint and recording their effects in a common capability space. Together with current capability, these responses form L-State; its pulse block supports a flexible direct readout and a...

    arxiv.org/abs/2609.08618 · PDF

  33. 33

    Why shared attention vectors fail: a case for outcome-indexed tuning

    Lenard Dome

    cs.LG · cs.NE · q-bio.NC

    Dimensional attention in learning is often implemented as a globally shared attention vector, where each stimulus dimension corresponds to a single scalar. These scalars are learned by models through gradient-descent on error, where predictive features acquire more salience. We show that under multi-outcome learning, where models predict more than one outcome, this shared vector becomes unstable; it collapses to its bounds and prevents the...

    arxiv.org/abs/2609.08615 · PDF

  34. 34

    Multi-Level-Set-Based Physics-Driven Neural Network to Solve 3-D Inverse Scattering Problems

    Yutong Du, Zicheng Liu, Bo Qi, Yali Zong, Peixian Han

    cs.LG · physics.comp-ph

    This paper proposes a level-set-based physics-driven neural network solver (LSPDNN) for 3-D electromagnetic inverse scattering. To mitigate boundary blurring and reconstruction artifacts in voxel-wise contrast reconstruction, the proposed solver exploits the piecewise homogeneity of practical scatterers by representing unknown targets with multiple coordinate-dependent neural level-set components. Specifically, a soft-union multi-material...

    arxiv.org/abs/2609.08594 · PDF

  35. 35

    Leveraging Cardiac Imaging to Improve ECG-Based Detection of Chagas Disease in Resource-Constrained Settings

    Laura Alvarez-Florez, Daniel Uyterlinde, Samuel Ruipérez-Campillo, Lukas P. A. Arts, Folkert W. Asselbergs, Fleur V. Y. Tjong

    cs.LG · cs.AI · eess.IV

    Chagas disease is a major cause of cardiomyopathy in Latin America. Cardiac magnetic resonance (CMR) imaging can characterize its structural abnormalities, but scanners and expert readers remain scarce in endemic regions. Electrocardiography (ECG) is inexpensive and widely available, yet structural disease must be inferred indirectly from electrical signals. We propose to transfer CMR-derived structural knowledge to ECG through contrastive...

    arxiv.org/abs/2609.08582 · PDF

  36. 36

    AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery

    Sayan Dhan, Selvaraju Natarajan

    cs.LG · q-fin.CP · stat.ML

    Formulaic alpha discovery is a pool-dependent symbolic search problem in which informative feedback is observed primarily when a complete expression is evaluated. This delayed feedback creates two coupled difficulties: the retained alpha pool does not preserve the full history of realized evaluation feedback, and the value of an intermediate construction action is uncertain because its consequence depends on the formula eventually completed....

    arxiv.org/abs/2609.08581 · PDF

  37. 37

    Certified Topological Interaction in Neural Representations: Class Disentanglement Is Mostly Pairwise

    Sushovan Majhi

    cs.LG · math.AT

    Class disentanglement (the separation of a representation's class-conditional point clouds along depth and over training) is usually read off descriptive curves. We measure it as certified topological interaction between labeled point clouds, using the recently introduced Intersection Euler Characteristic Profile: the Euler characteristic of the overlap of the clouds' ball unions as a function of scale, computed by one Alpha-complex sweep...

    arxiv.org/abs/2609.08561 · PDF

  38. 38

    Not All Variables Agree: Reliability-Aware Variable-Wise Gradient Surgery for Multivariate Time-Series Forecasting

    Jinwoo Park, Hyeongwon Kang, Pilsung Kang

    cs.LG

    In data-driven training, multivariate time-series forecasting is usually optimized with a scalar loss averaged over samples, variables, and horizons. This averaging is convenient, but the optimizer sees only the aggregated gradient, which does not reveal whether the variable-wise contributions align or oppose one another. To quantify how often this disagreement arises, we measure the variable-wise gradients directly and find that 30.6% of...

    arxiv.org/abs/2609.08554 · PDF

  39. 39

    Topological Fraud Detection in Latent Transaction Spaces

    Avraham Bourla

    cs.LG · cs.CR

    Working entirely on topologically anonymized embeddings, we perform fraud detection using iterative rounds of unsupervised filtering followed by supervised sniping. The result is an ultra-low latency privacy--preserving triage that allows institutions to flag suspicious activity without compromising Personally Identifiable Information.

    arxiv.org/abs/2609.08445 · PDF

  40. 40

    Stochastically Perturbed Weights: Ensembles from Deterministic Machine-Learning Weather Models

    Simon Adamov, Oliver Fuhrer, Reto Knutti, Sebastian Schemm

    cs.LG · physics.ao-ph

    Machine-learning weather models (MLWMs) now match or outperform operational numerical weather prediction (NWP) at global medium-range forecasting, at far lower inference cost. Many deployed MLWMs are deterministic, producing a single forecast with no estimate of its own uncertainty, whereas a growing family of trained-probabilistic models generate calibrated ensembles directly, at the price of a dedicated training run. We ask instead how much...

    arxiv.org/abs/2609.08412 · PDF

  41. 41

    Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

    Hongbang Yuan, Zhuoran Jin, Yixin Cao

    cs.LG · cs.AI · cs.CL

    Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. While conventional \textit{agent-side warming} up via supervised fine-tuning (SFT) can alleviate this, it is frequently limited by data scarcity and constrained exploration. To address this, we propose a paradigm shift to...

    arxiv.org/abs/2609.08404 · PDF

  42. 42

    MLIP Detective: Active Failure Mode Discovery Beyond Benchmark Scores for Machine-Learning Interatomic Potentials

    Ryuhei Okuno, Nontawat Charoenphakdee, Kaoru Hisama, Yuta Tsuboi

    cs.LG · cond-mat.mtrl-sci · physics.chem-ph

    Universal machine-learning interatomic potentials (u-MLIPs) aim to generalize across diverse configurations. Benchmarks enable reproducible evaluation but may not expose failures outside their predefined scope. Here, we show that physics-informed search can complement benchmark-based evaluation by uncovering hidden failure modes. We introduce MLIP Detective, an agentic framework for active failure mode discovery. Starting from benchmark...

    arxiv.org/abs/2609.08399 · PDF

  43. 43

    Equivariance Breaks the Learning Rate

    Andrei Manolache, Mathias Niepert

    cs.LG · cs.AI

    Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on these architectures without explaining why. We identify one source of this difference inside equivariant linear layers. Each irrep block learns a channel-mixing matrix $W_l$ shared across its $2l+1$ components, giving the expanded map $W_l \otimes I_{2l+1}$. For a single application of the layer,...

    arxiv.org/abs/2609.08381 · PDF

  44. 44

    Geographically Regularized AUC-Maximizing Personalized Federated Learning

    Mayu Hiraishi, Kensuke Tanioka, Toshio Shimokawa

    cs.LG · stat.ME

    Accurate diagnostic and risk-prediction models are important for supporting clinical decision-making during infectious disease outbreaks. However, privacy and governance requirements may restrict patient-level data sharing across healthcare institutions, and data distributions often vary. Moreover, AUC is widely used to evaluate discriminative performance, motivating its direct optimization in model development. We propose geographically...

    arxiv.org/abs/2609.08379 · PDF

  45. 45

    IPM-FM: A Foundation Model with Consensus Feature Selection for Industrial Process Monitoring

    Liang Cao, Weide Liu, Yan Qin, Jun Cheng, Weisi Lin, Bhushan Gopaluni

    cs.LG · cs.AI

    Industrial process monitoring is fundamental to the safety and economic performance of modern process plants. Current practice remains a one-task-one-model paradigm that is label-inefficient and prone to degradation under operating drift. Foundation models have reshaped language, vision, and generic time-series forecasting, but it has not been adapted to industrial process monitoring. This setting poses domain-specific challenges, including...

    arxiv.org/abs/2609.08375 · PDF

  46. 46

    Miles v0.1: Production-Level Post-Training

    RadixArk, :, Tom Chen, Mao Cheng, Shi Dong, Kangrui Du, Yanbin Jiang, Jiajun Li, Yiming Li, Tao Lin, Yusheng Su,...

    cs.LG · cs.CL

    We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable. With accuracy, efficiency, reliability, and scalability as first-class goals, Miles aims to make frontier-scale RL accessible to researchers and enterprises...

    arxiv.org/abs/2609.08368 · PDF

  47. 47

    Geometry-Aware Bayesian Parameter-Efficient Fine-Tuning on the Stiefel Manifold via Stein Variational Gradient Descent

    Quang-Duy Tran, Trung Le, Bao Duong, Phuoc Nguyen, Thin Nguyen

    cs.LG

    Several geometry-aware approaches to low-rank adaptation have emerged for parameter-efficient fine-tuning of large pre-trained models. These methods aim to take full advantage of the geometric structure of low-rank manifolds for improving the efficiency in subspace utilization and reducing redundancy by enforcing orthogonality constraints during optimization. The strong empirical results of these techniques have motivated further study into...

    arxiv.org/abs/2609.08354 · PDF

  48. 48

    TV-Regulated OPD: Direction Matters in On-Policy Distillation

    Han Xiao, Yifan Niu, Dongyi Liu, Chang Luo, Jia Li

    cs.LG

    On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). However, the supervision signals in mainstream OPD methods suffer from high variance and noise which is generally instable during training. In this work, we systematically investigated what really matters to the performance and the fundamental mechanisms behind the instability during...

    arxiv.org/abs/2609.08341 · PDF

  49. 49

    Distillation as Probability Transport: Routed On-Policy Distillation

    Tianle Xia, Lingxiang Hu, Yiding Sun, Linfang Shang, Ming Xu, Lan Xu, Ning Zheng, Wei Xu, Jie Jiang

    cs.LG · cs.CL

    On-policy distillation (OPD) transfers teacher knowledge on student-generated trajectories, but efficient sampled objectives reduce the teacher distribution to scalar credit on individual tokens. Such credit indicates whether a token should gain or lose probability, yet leaves the corresponding redistribution unspecified. We recast OPD as teacher-guided probability transport and propose RouteOPD (Routed On-Policy Distillation), which...

    arxiv.org/abs/2609.08337 · PDF

  50. 50

    EMBLEM: Enhancing Multi-script Table Detection through Masking

    Dhruv Kudale, Udhay Brahmi, Ganesh Ramakrishnan

    cs.LG

    Table detection is a core task in document analysis, supporting downstream applications such as information retrieval, document reconstruction, and visual question answering. While existing deep learning models perform well on English and Chinese documents, they struggle with multilingual, multi-script documents due to script diversity and the limited availability of labeled data. To address this challenge, we introduce MANDALA (Multi-script...

    arxiv.org/abs/2609.08330 · PDF

  51. 51

    HypLTSF: A Hyperbolic Geometric View of Multi-Scale Hierarchies for Long-Term Time Series Forecasting

    Namwoo Kim, Hyungryul Baik, Yoonjin Yoon

    cs.LG

    Multi-scale modeling has become an effective approach for long-term time series forecasting, capturing temporal patterns that range from fine-grained local dynamics to coarse global trends. Representations across these temporal scales are inherently hierarchical, with coarser scales abstracting and aggregating information from finer ones. While existing approaches readily exchange information across these scales, the hierarchy itself is...

    arxiv.org/abs/2609.08286 · PDF

  52. 52

    Adaptively Incorporating Directional Hints into Zeroth-Order Optimization

    Alexander Ryabchenko, Jian Qian, Wenlong Mou

    cs.LG · math.OC

    We study zeroth-order optimization of non-convex functions with the aid of directional hints, which are cheap but potentially inaccurate approximations of the true gradient direction, given by linear subspaces at each iteration. To leverage these hints adaptively while maintaining robustness to their quality, we introduce Control-Variate Zeroth-Order Descent (CV-ZOD), a new framework that refines the classical zeroth-order gradient estimator...

    arxiv.org/abs/2609.08277 · PDF

  53. 53

    Online Signature Verification Using Augmented Path Signature and T-Mamba

    Ruiling Li, Danyu Yang

    cs.LG · cs.CV

    Handwritten signature verification is vital for personal authentication across commercial and financial applications. Although deep learning methods are widely adopted for online signature verification (OSV), they often struggle with capturing highly discriminative features and modelling long-range dependencies. To address these issues, we propose a novel framework that integrates the augmented path signature (APS) descriptor with the T-Mamba...

    arxiv.org/abs/2609.08276 · PDF

  54. 54

    Synergistic Fusion of Topological Structure and Temporal Semantics of Mobility for Urban Region Embedding

    Namwoo Kim, Jeeyun Chang, Kanghoon Lee, Yoonjin Yoon

    cs.LG · cs.AI

    Urban region embeddings have shown promising results in diverse urban sensing tasks such as crime, income, and service-call prediction. Recent methods improve representation quality by integrating mobility data with auxiliary modalities, using cross-view attention or contrastive objectives to align heterogeneous features into a unified region representation. However, leveraging the temporal dynamics of human mobility remains under-explored....

    arxiv.org/abs/2609.08268 · PDF

  55. 55

    Revisiting Spectral Representations in Generative Diffusion Models

    Yuehao Wang, Peihao Wang, Hanwen Jiang, Ziyi Yang, Qixing Huang, Zhangyang Wang

    cs.LG

    Diffusion models have shown remarkable performance on diverse generation tasks. Recent work finds that imposing representation alignment on the hidden states of diffusion networks can both facilitate training convergence and enhance sampling quality, yet the mechanism driving this synergy remains insufficiently understood. In this paper, we investigate the connection between self-supervised spectral representation learning and diffusion...

    arxiv.org/abs/2609.08253 · PDF

  56. 56

    CUNO: Curriculum and Preference Optimization for Stable Graph Unlearning under Mass Deletion

    Chenhan Zhang, Ali Braytee, Madhushi Bandara, Xin Hao, Paul J. Kennedy, Massimo Piccardi, Raymond Owen

    cs.LG · cs.AI

    Graph unlearning removes the influence of designated training data from a trained graph model without retraining from scratch. However, existing methods suffer a sharp drop in model utility under large deletion ratios (mass deletion), a phenomenon we refer to as catastrophic unlearning. We find that a key cause is the uniform treatment of all deleted samples, which is particularly damaging in graph learning: structural dependencies cause...

    arxiv.org/abs/2609.08244 · PDF

  57. 57

    SIM: Subspace Interaction-based Method for Token-Level Text Anomaly Detection

    Kehan Yan, Yue Tan, Qingfeng Chen, Shiyuan Li, Yu Zheng, Yixin Liu

    cs.LG

    Token-level text anomaly detection, as an emerging trend of text anomaly detection, moves beyond coarse-grained document-level detection by localizing anomalous tokens within text. By providing fine-grained abnormality prediction, token-level text anomaly detection plays a critical role in various real-world applications, such as spam filtering and fake news detection. However, existing methods still rely on the global distance calculation...

    arxiv.org/abs/2609.08200 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.