cs.LG · 2026-09-15 · No. 114

Machine Learning, 2026-09-15.

61 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

61 entries
  1. 01

    Bellman Policy Optimization

    Zhuoqing Song, Haotian Xu, Xikun Zhang, Lidong Bing

    cs.LG · cs.CL · math.OC

    Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD). For autoregressive generation with terminal rewards, BPO uses the Bellman equations to reformulate PMD as a trajectory-level objective. The reformulation avoids estimating state values at intermediate states. We...

    arxiv.org/abs/2609.15987 · PDF

  2. 02

    The Router Within: Eliciting Native Skill Routing from a Frozen LLM

    Ruishuo Chen, Xun Wang, Yu Chen, Zhuoran Li, Longbo Huang

    cs.LG · cs.AI · cs.CL

    Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by preloading every skill's metadata into the context, which disperses the agent's attention and caps the library size. Retrieval pipelines move the selection out of the context, but also out of the agent's capability. We show that the frozen agent LLM already carries the routing signal in its own...

    arxiv.org/abs/2609.15982 · PDF

  3. 03

    A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models

    Xingyun Wang, Haomin Zheng, Man Yuan, Leqian Yang, Ziming Liu

    cs.LG · cs.CV

    When a video model generates physically incorrect motion, did it fail to learn the correct motion, or did it learn it but fail to use it? We show the latter: the correct motion remains available inside the model and can still be made to control the generated video. We train on videos where red masses oscillate slowly and blue masses oscillate quickly, then test a red mass with fast observed motion. Even when the model generates slow motion in...

    arxiv.org/abs/2609.15980 · PDF

  4. 04

    Privacy-Aligned Personalized Federated Learning with Compact Adaptation and Variable-Length Gaussian Communication

    Yilin Xu, Chun Hei Michael Shiu, Chih Wei Ling, Linqi Song

    cs.LG

    Record-level differential privacy exposes a structural misalignment in personalized federated learning when client-specific variation is low-dimensional while training repeatedly releases high-dimensional updates. In this paper, we address this misalignment by releasing a private client context once and confining repeated adaptation to a fixed coefficient space. Beyond dimensionality reduction, the factorized generator induces an adaptive...

    arxiv.org/abs/2609.15950 · PDF

  5. 05

    Safe Meta-Reinforcement Learning via Information Space Reachability

    Zeyang Li, Sunbochen Tang, Navid Azizan

    cs.LG · eess.SY

    Meta-reinforcement learning (meta-RL) enables agents to adapt to unseen tasks with limited experience. Despite its promise, the application of meta-RL in real-world tasks is hindered by safety requirements, which have been underexplored in prior work. In this paper, we propose a safe meta-RL framework that explicitly accounts for safety during adaptation. Our key insight is to reason about safety in the information space, which captures both...

    arxiv.org/abs/2609.15915 · PDF

  6. 06

    Discrete Beckmann Transport Models for One-Step Language Modeling and Reasoning

    Sophia Tang, Shiyi Wang

    cs.LG

    Discrete diffusion and flow models are a promising alternative to autoregressive language models, but compressing many-step sampling into fewer steps typically requires distilling a pretrained teacher model. This caps the student at the teacher's quality and requires a costly two-stage training pipeline. We introduce Discrete Beckmann Transport Models (DBTM), built on a time-independent flow whose autonomous transport map provably carries any...

    arxiv.org/abs/2609.15903 · PDF

  7. 07

    Privacy-enhanced federated learning via asynchronous aggregation and local differential perturbation

    Zhen Zhong, Shini Yang, Liesheng Wei

    cs.LG · cs.AI · cs.CE · cs.DB

    This study proposes a privacy-enhanced federated learning framework to address secure collaborative training in distributed data environments. The framework integrates Dynamic Differential Privacy (DDP), lightweight Homomorphic Encryption (HE), and Local Differential Privacy (LDP) mechanisms to ensure data privacy protection during model training. Additionally, the framework employs an asynchronous aggregation strategy with version control to...

    arxiv.org/abs/2609.15885 · PDF

  8. 08

    Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport

    Jaehun Shon, Jinha Choi, Jongwook Jeon, Jongmin Lee

    cs.LG · cs.AI

    Offline reinforcement learning aims to learn a policy solely from fixed datasets, which often contain multimodal action distributions. Flow policies can naturally represent such multimodal behaviors, but learning an efficient one-step flow policy remains challenging: standard value guidance often leads to mode collapse or exploits overestimation bias in out-of-distribution regions. To address this, we introduce One-step Flow policy via...

    arxiv.org/abs/2609.15883 · PDF

  9. 09

    LLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys

    Md Khalid Syfullah, Alvi Ataur Khalil

    cs.LG · cs.AI · cs.CR

    Rising societal and lifestyle complexity has been linked to a growing prevalence of mental distress worldwide. Educational institutions, workplaces, clinics, etc. collect large volumes of mental health survey data to understand and reduce this burden. Collaborative analysis of such data could yield effective generalizable predictive models. Privacy constraints and varied survey designs (i.e., different questions, scales, and formats) hinder...

    arxiv.org/abs/2609.15871 · PDF

  10. 10

    Task-Directed Residual AddUNet:Perfect-Reconstruction Routing for Full-Rate Representations

    Vikram R. Lakkavalli

    cs.LG · eess.SP

    This paper establishes a perfect-reconstruction (PR) interpretation of AddUNet and its full-rate realization, and introduces a Residual Full-Rate PR architecture for task-directed representation learning. The survivor--skip structure of a constrained additive U-Net is shown to be exactly equivalent to a critically sampled multirate PR filter bank. The full-rate formulation removes the complementary-subband restrictions of the critically...

    arxiv.org/abs/2609.15857 · PDF

  11. 11

    Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression

    Huicheng Zhang, Xiyao Feng, Ze-Tong Li, Chengkai Zhu, Xiao Shi, Xiwei Pan, Jinguo Liu, Ge Bai, Xin Wang

    cs.LG · cs.AI

    Per-matrix singular value decomposition (SVD) truncation is Eckart-Young optimal in the whitened Frobenius norm, but errors from independently compressed matrices compound through the block's nonlinear forward pass. Inspired in part by hierarchical variational optimization in quantum many-body methods, we introduce a three-level chain that widens optimization scope from individual matrices to Transformer blocks to the full model: whitened...

    arxiv.org/abs/2609.15838 · PDF

  12. 12

    Sharp Rates and a One-Line Correction for Spectral Representation Learning

    Dier Tang, Jing Yee Tan, Guangyue Han

    cs.LG · cs.IT · stat.ML

    A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's question is when the off-the-shelf features are good enough and when they need fixing. Canonical correlation analysis, HGR maximal correlation, and the population optimum of the spectral contrastive loss all return the top-$k$ singular subspace of a cross-view dependence operator, justified by...

    arxiv.org/abs/2609.15825 · PDF

  13. 13

    MoveBench: A Benchmark for Global-Scale Wildlife Movement Forecasting

    Justin Kay, Shir Bar, Ellen O. Aikens, Martin Becker, Francesca Cagnacci, Juliet Cohen, Scott W. Forrest, Jessica...

    cs.LG

    Understanding and predicting wildlife movement is critical for ecology and conservation. While trajectory forecasting has advanced for human and vehicle movement, wildlife trajectories present distinct challenges: they are unconstrained in space, highly stochastic, and influenced by environmental conditions. We introduce MoveBench, the first large-scale benchmark for probabilistic wildlife movement forecasting, containing 2.6M GPS locations...

    arxiv.org/abs/2609.15780 · PDF

  14. 14

    Transfer Learning for Socioeconomic Estimation in Forced-Displacement Settings

    Steven Ndung'u, Adel Daoud, Ismael Yacoubou Djima, Hai-Anh H. Dang, Patrick Michael Brock

    cs.LG · cs.AI

    Progress in inclusive household surveys has strengthened socioeconomic evidence for forcibly displaced populations, providing indispensable benchmarks on living conditions and welfare. However, these surveys remain resource-intensive and periodic, while conditions can change between rounds, particularly in settings affected by fragility, conflict, and violence. More frequently updated, spatially granular complementary evidence is therefore...

    arxiv.org/abs/2609.15773 · PDF

  15. 15

    Sylvas: Synergistic Learning Value based Device Scheduling in Federated Continual Learning

    Yuxuan Sun, Yuxuan Bai, Tan Chen, Sheng Zhou, Zhisheng Niu

    cs.LG · cs.AI

    Federated continual learning (FCL) enables shared global models to continuously adapt to distributed and non-stationary data streams, making it important for Internet of Things applications such as intelligent transportation, industrial monitoring, and unmanned systems. Under spatio-temporal data distribution dynamics and label scarcity, a key challenge is how to quantify the contribution of each edge device to global learning performance and...

    arxiv.org/abs/2609.15763 · PDF

  16. 16

    A Language-Guided Multimodal Foundation Model for Zero-Shot and Multi-Task Brain Signal Analysis

    Mingzhi Chen, Yiyu Gui, Guibo Luo, Yuchao Yang

    cs.LG · cs.AI

    Brain signal analysis is essential for both neuroscience research and clinical diagnostics, yet current approaches face critical limitations. End-to-end models require task-specific retraining and exhibit limited generalization, while pre-trained models lack semantic depth and still depend on extensive fine-tuning. Meanwhile, general-purpose multimodal foundation models, though powerful in other domains, struggle to interpret brain signals...

    arxiv.org/abs/2609.15740 · PDF

  17. 17

    Solving Finite-sum Coupled Compositional Optimization via Multi-block-Single-probe Estimator

    Wei Jiang, Sifan Yang, Yibo Wang, Lijun Zhang, Zechao Li

    cs.LG · math.OC

    Traditional variance reduction methods (e.g., SPIDER, SARAH, STORM) have been extensively investigated for improving the convergence rates of stochastic optimization. These techniques typically maintain a sequence of estimators for a single function (or gradient) across iterations. However, what if we need to track multiple functions, but can only access stochastic samples of $\mathcal{O}(1)$ functions at each iteration? This scenario arises...

    arxiv.org/abs/2609.15723 · PDF

  18. 18

    Assembling the CREW: A Collaborative Multi-agent Reinforcement Learning Framework for Automated Related Work Generation

    Hai-Dang Dang, Bao-Yen Pham, Bao Nguyen, Tran Thi Huong, Huynh Thi Thanh Binh

    cs.LG

    Automatic Related Work Generation (RWG) significantly reduces the human time and effort required to author the Related Work Section (RWS) of a research paper. However, prior methods leveraging multi-agent Large Language Models (LLMs) typically rely on a predefined workflow, where each agent is responsible for a specific step in the entire process. This rigid, static inter-agent coordination limits the adaptive collaboration required to...

    arxiv.org/abs/2609.15721 · PDF

  19. 19

    Knowledge-Enriched Structured EHR Features for 30-Day Hospital Readmission Prediction on MIMIC-IV

    Mohamad Najafi, Hongyun Fu, Mathias Brochhausen, Jian Wu, Yaohang Li

    cs.LG · q-bio.QM

    Recent approaches to 30-day hospital readmission prediction rely on pre-trained language models applied to discharge summaries. Although these methods achieve strong performance, they depend on the availability of clinical notes, incur substantial computational costs, and yield representations that lack interpretability. We propose a knowledge-enriched feature representation that augments structured Electronic Health Record (EHR) data with...

    arxiv.org/abs/2609.15713 · PDF

  20. 20

    Backward SDEs-based Diffusion for Physics-Constrained Generation

    Zihao Wang

    cs.LG

    Pretrained score-based diffusion models provide strong unconditional priors, yet enforcing measurement or physics consistency in inverse problems is often handled by heuristic guidance, intermittent projections, or task-specific conditional training, with limited guarantees of feasibility at the end of inference. We propose terminal-conditioned inversion for score-based SDE priors. Given a frozen Score-SDE prior and a task-defined terminal...

    arxiv.org/abs/2609.15702 · PDF

  21. 21

    Principal-timestep Restricted Init via Sparse Matrix-decomposition in Flow-matching

    Jiayang Gu, Zheng Fang, Lichaun Xiang, Fanghui Liu, Xu Cai, Hongkai Wen

    cs.LG

    Flow-matching diffusion models have recently emerged as a strong paradigm for high-fidelity visual generation. However, their prohibitively high fine-tuning cost limits scalability to downstream tasks. While Low-Rank Adaptation (LoRA) combined with spectral initialization has demonstrated accelerated convergence and improved performance in autoregressive language models by better aligning gradient directions, we find that it fails to deliver...

    arxiv.org/abs/2609.15643 · PDF

  22. 22

    FedLTLib: A Comprehensive Benchmark for Federated Long-Tail Learning

    Changkun Lin, Junxiao Wang

    cs.LG · cs.AI

    Driven by the escalating demand for privacy-preserving computing, Federated Learning (FL) has witnessed remarkable progress, becoming a cornerstone technology for bridging distributed data silos in mobile edge networks. However, in real-world mobile computing environments, data is generated by heterogeneous mobile devices with varying user behaviors, leading to a significant Long-Tail Distribution. Unlike idealized balanced datasets, data in...

    arxiv.org/abs/2609.15625 · PDF

  23. 23

    Where to Compute and How to Interact: Operator-Readable Adaptation with Gauge-Aware Transport

    Zixuan Shen, Quanxu Wan, Bingchuan Wang, Zhi Wang, Biao Luo

    cs.LG

    Adaptive meshes enable neural operators for partial differential equations (PDEs) to allocate spatial samples and computation according to local physical structures. Existing approaches, however, mainly address where to compute, with less attention to how information should interact after node relocation. Mesh adaptation changes local sampling scales, neighborhood structures, and geometric contexts, so representations formed at different...

    arxiv.org/abs/2609.15620 · PDF

  24. 24

    Multi-View Molecular Representation Learning with Hierarchical Graphs and Contextualized Fingerprints

    Gwang-Hyeon Yun, Jong-Hoon Park, Bing Hu, Helen Chen, Anita Layton, Young-Rae Cho

    cs.LG · cs.AI

    Molecular property prediction requires representations that generalize from limited labeled data to structurally novel compounds. Existing molecular pretraining methods often rely on a single view: graph-based approaches model atom-bond topology but provide limited fragment-level supervision, whereas fingerprint descriptors encode chemical patterns but are typically used as fixed auxiliary features. We propose HiFi-Mol, a multi-view framework...

    arxiv.org/abs/2609.15611 · PDF

  25. 25

    Bayesian Optimisation Using Product-of-Experts Gaussian Process Models with Uncertainty Calibration

    Yean Hoon Ong

    cs.LG

    Bayesian optimisation (BO) typically relies on a single global Gaussian process (GP) model as its surrogate model. However, GP regression has cubic computational complexity in the number of training data points, limiting its applicability to large-scale optimisation problems. The product-of-experts Gaussian process model with uncertainty calibration (GP-pro-c) mitigates this limitation by combining multiple local GP experts, enabling improved...

    arxiv.org/abs/2609.15555 · PDF

  26. 26

    The Token Before the Value Is the Key: How Hybrid Architectures Organize Induction Circuits

    Ke Cheng, Xin Xu, Yixiao Chen, Lei Xin, Jianbo Zhao, Fanhu Zeng, Yue Liu, Jun Zhang, Jie Jiang

    cs.LG

    Hybrid language models can improve capability as well as efficiency, raising the question of how architectural complementarity becomes learned computation. We examine the established induction roles of Carrying predecessor information, Matching a source by content, and Copying its value. How are these position-sensitive and content-based computations allocated across heterogeneous layers? We introduce layer-type-agnostic paired probes that...

    arxiv.org/abs/2609.15545 · PDF

  27. 27

    Specifying Reward Functions for RL Without Environment Sampling

    Stephane Hatgis-Kessell, W. Bradley Knox, Emma Brunskill

    cs.LG · cs.AI

    Enabling human stakeholders to specify reward functions that lead to their desired outcomes is a key challenge in deploying reinforcement learning agents. Preference-based methods such as online RLHF can reduce the burden of manual reward design, but they require repeatedly training policies, sampling trajectories from the real world, and eliciting feedback, making them impractical in settings where environment interaction is computationally...

    arxiv.org/abs/2609.15544 · PDF

  28. 28

    The Misery of Mechanistic Interpretability: A Formal Perspective

    Tobias Ladner, Matthias Althoff

    cs.LG · cs.AI

    Mechanistic interpretability has become the dominant lens for understanding frontier language models, as their inner workings are complex and inherently black boxes. To gain insights into these models, interpretable replacement networks (IRNs) are trained at all layers, exposing interpretable features through sparsely activated neurons. However, the faithfulness of an IRN is usually evaluated only empirically on clean data, and we show that...

    arxiv.org/abs/2609.15533 · PDF

  29. 29

    Beyond Noise: Understanding and Overcoming Temperature Effects in Analog DNN Inference

    Niklas Summ, Xiao Wang, Hendrik Borras, Bernhard Klein, Holger Fröning

    cs.LG

    The energy efficiency of analog computing makes it one of the most promising candidates for deploying resource-intensive machine learning workloads on constrained platforms such as mobile and embedded devices. However, analog accelerators are inherently susceptible to noise and non-idealities arising from physical component variations, whose behavior is further sensitive to environmental factors. These effects can significantly degrade...

    arxiv.org/abs/2609.15527 · PDF

  30. 30

    End-to-End Verifiable and Robust Federated Learning

    Doryan Lesaignoux, Enrique Mármol Campos, Gabriele Spini, José L. Hernández-Ramos, Stephan Krenn

    cs.LG

    Federated learning enables multiple parties to train a shared model without centralizing raw data with the help of an aggregator, but introduces integrity risks once participants or infrastructure are not fully trustworthy. Two requirements are particularly important: robustness to poisoned or Byzantine client updates, and verifiability of the aggregator so that clients or third parties can audit the reported aggregation without learning...

    arxiv.org/abs/2609.15521 · PDF

  31. 31

    Same path, different: a mechanistic comparison of looped and stacked transformer encoders on 12-lead ECG

    Pawel Olszowiec, Michal Byra, Grzegorz Gruszczynski, Grzegorz Stefanski, Alberto Presta

    cs.LG · math.DS

    Recurrent Transformers reusing their weights rather than stacking $L$ distinct layers are becoming widely adopted due to their parameter efficiency [1,2,3]. However, the exact representational and dynamical differences between looped and stacked architectures remain uncharacterized. This paper presents a controlled study on the example of bViT model [1] applying one weight-tied block $L$ times. We train two models: bViT and standard ViT [4]...

    arxiv.org/abs/2609.15498 · PDF

  32. 32

    Rotation-Based Subspace Tracking for Robust Kernel PCA on Streaming Data

    Kris Lokere, John Fossaceca

    cs.LG

    Machine learning models process large amounts of data, and Principal Component Analysis (PCA) is a widely used technique to reduce the dimensionality of the data and extract useful features. In practice, datasets often change over time (data drift) and/or arrive one sample at a time (streaming data), making it infeasible to process the entire dataset at once in batch mode. Real-world data also often contains nonlinear patterns, which...

    arxiv.org/abs/2609.15488 · PDF

  33. 33

    Data-driven Prediction of Satellite-observed Avalanche Activity from Snowpack Simulations

    Jakob Grah, Filippo Maria Bianchi, Bert Kruyt, Karsten Müller

    cs.LG

    Avalanche forecasting requires knowledge of snowpack conditions and recent avalanche activity, but field observations are sparse across large mountain regions. We explore whether SNOWPACK simulations can predict avalanche activity mapped by synthetic aperture radar (SAR). We compiled five winters of Sentinel-1 avalanche detections across Norway and parts of Sweden, alongside SNOWPACK simulations forced by numerical weather predictions on a 20...

    arxiv.org/abs/2609.15485 · PDF

  34. 34

    GSLAD: Prototype-Regularized Graph Structure Learning for Multivariate Time Series Anomaly Detection

    Zepeng Zhang, Fuad Khuri, Keivan Faghih Niresi, Olga Fink

    cs.LG

    Unsupervised multivariate time series anomaly detection methods typically identify anomalies through forecasting, reconstruction, or representation discrepancies. However, industrial faults may first alter inter-variable structural patterns while individual trajectories remain close to normal, resulting in weak anomaly signals. In this paper, we propose GSLAD, a prototype-regularized graph structure learning framework that uses structural...

    arxiv.org/abs/2609.15483 · PDF

  35. 35

    On the role of the tokenizer in ECG transformer models

    Jiawei Li, Fabio Bonassi, Johan Sundström, Thomas B. Schön, Antônio H. Ribeiro

    cs.LG · cs.AI

    Tokenization determines both the physiological content presented to an ECG Transformer and the sequence over which attention operates. We compare eight tokenization strategies across Transformer, Informer, Reformer, and FEDformer on the nine-label CPSC2018 classification task. The input projection and principal backbone capacity are controlled to isolate the effect of token construction. Median-beat and HeartLang tokenization achieve mean...

    arxiv.org/abs/2609.15433 · PDF

  36. 36

    Single-condition neural solvers encode transferable response spaces for parametric differential equations

    Wenbo Cao, Weiwei Zhang

    cs.LG · physics.comp-ph

    Operator learning for parametric partial differential equations (PDEs) typically builds global models over prescribed domains, requiring cross-condition data or costly physics-constrained training. Here we show that the output Jacobian of a neural solution model trained at one condition defines a reusable response space for cross-condition solution variations. We introduce Linearized Subspace Transfer (LST) to exploit this space and recover...

    arxiv.org/abs/2609.15432 · PDF

  37. 37

    CodeTS: Verifiable Text-to-Time Series Generation via Executable Code

    Xudong Yuan, Shunyu Liu, Tongya Zheng, Huiping Zhuang, Mingli Song, Kaixuan Chen

    cs.LG · cs.AI

    Text-to-Time Series Generation (Text-to-TS) provides a promising paradigm for synthesizing time series from natural language, enabling scenario-specific generation when real observations are scarce or costly to acquire. However, existing methods typically lack an explicit mechanism for deriving generation logic from textual descriptions to guide time series synthesis. In this paper, we propose CodeTS, a verifiable framework that uses code as...

    arxiv.org/abs/2609.15393 · PDF

  38. 38

    Representing Clinical Conditions on Vital Signs from Healthy Individuals using Latent Modeling

    Rafael Pina, Varuna De Silva, Mindula Illeperuma

    cs.LG

    Machine learning can be crucial to help scale complex signal processing applications in scenarios such as healthcare. However, these machine learning models need rich datasets to be trained and there are often cases where it is not possible to access representative datasets. In this paper, we propose a deep generative model based on conditional variational autoencoders with the objective of augmenting the vital signs of healthy individuals in...

    arxiv.org/abs/2609.15379 · PDF

  39. 39

    Robust and Efficient Communication for Multi-Agent Learning

    Rafael Pina, Varuna De Silva, Corentin Artaud

    cs.LG · cs.AI · cs.MA

    Effective communication is a cornerstone of distributed intelligence in Multi-Agent Reinforcement Learning (MARL), yet ensuring that generated messages are both informative and robust to physical constraints remains a significant challenge. This paper introduces Multi-Agent Regularized Communication (MARC), a novel framework inspired by information-theoretic principles of conditional mutual information. MARC employs an attention-based...

    arxiv.org/abs/2609.15361 · PDF

  40. 40

    The Universe of Universes: Benefit Yield Functions, Implosion Thresholds, and Infrastructure-Aware Optimization in Multi-LLM Systems

    Danielle Franklin, Vasu Raj Jain

    cs.LG · cs.AI · cs.MA

    We introduce the Universe of Universes (UoU) framework, which treats the full ecosystem of major large language models (LLMs) as a structured retrieval corpus and proposes a compositional Automated Reasoning (AR) and Machine Learning (ML) architecture for cross-model retrieval-augmented generation. The central contribution is the formal characterization of the Benefit Yield Function (BYF), the marginal performance gain per additional model...

    arxiv.org/abs/2609.15314 · PDF

  41. 41

    When Correlations Mislead: Confounder-Aware Multi-View Urban Region Representation Learning

    Sean Bin Yang, Ying Sun, Zongyi Xu, Tung Kieu, Jilin Hu, Bin Yang, Kristian Torp, Hua Lu, Torben Bach Pedersen

    cs.LG · cs.AI

    Urban region representation learning commonly combines heterogeneous data sources, such as mobility flows, points of interest, and land-use information, to support tasks including mobility analysis, public safety forecasting, and service demand estimation. Existing multi-view methods typically improve region embeddings by strengthening interactions across views. However, such methods often overlook view-specific regional structures and may...

    arxiv.org/abs/2609.15305 · PDF

  42. 42

    Admissable: Training Reinforcement Learning Agents against Adversarial Missingness

    Paul Stahlhofen, Luca Hermes, Tim Kochs, Markus Vieth, Barbara Hammer

    cs.LG

    In order to make Reinforcement Learning algorithms applicable in real world scenarios, safety must be ensured even under adverse operating conditions. In this work, we consider the challenge of adversarial feature missingness: a scenario in which an adversary occludes features from the agent's observation in order to reduce performance as much as possible. We formally define adversarial missingness for Reinforcement Learning and compare it to...

    arxiv.org/abs/2609.15297 · PDF

  43. 43

    Impute-EM: Native Mixed-State Diffusion Models for Heterogeneous Data Imputation

    Sergei Kholkin, Kirill Sokolov, Dmitry Baranchuk, Evgeny Burnaev, Alexander Korotin

    cs.LG

    Missing values are ubiquitous in heterogeneous data mining, where numerical, categorical, and binary variables often coexist. Many imputation methods, especially diffusion-based ones, treat discrete variables through continuous surrogates such as one-hot relaxations rather than modeling them natively. This creates a mismatch between the model state space and the mixed discrete and continuous structure of the data. We propose Impute-EM, an...

    arxiv.org/abs/2609.15284 · PDF

  44. 44

    Draining Fictitious Knots: Restoring Distance-Awareness Guarantees for High-Dimensional Spline Networks

    Masoud Ataei, Mohammad Javad Khojasteh, Vikas Dhiman

    cs.LG

    Kolmogorov-Arnold Networks (KANs) with spline activations have recently shown promise for interpretable function approximation. Distance-Aware Error for Kolmogorov Networks (DAREK) introduces a computationally efficient bottom-up approach to uncertainty quantification by equipping KANs with distance-aware error bounds; yet, in high-dimensional settings, the theoretical guarantees can be weakened by the emergence of fictitious knots. Inspired...

    arxiv.org/abs/2609.15274 · PDF

  45. 45

    Learning CNF Formulas from Uniform Random Solutions: Near-Tight Sample Complexity for Valiant's Algorithm

    Weiming Feng, Yixiao Yu, Yiyao Zhang

    cs.LG · cs.DS

    We revisit Valiant's algorithm (Commun. ACM'84) for learning $n$-variable CNF formulas with clause size $k$ and variable degree $d$ from i.i.d. uniform random solutions in the local lemma regime. For fixed $t\geq1$, under $k\gtrsim(1+1/t)\log d$, Valiant's algorithm achieves total variation error $\varepsilon$ with $\widetilde{O}(n^{\lceil t \rceil}/\varepsilon)$ sample complexity. For $t>1$, we prove a matching lower bound for Valiant's...

    arxiv.org/abs/2609.15268 · PDF

  46. 46

    BioDCASE: Active Learning for Bioacoustics

    Ben McEwen, Rupa Kurinchi-Vendhan, Shiqi Zhang, Lukas Rauch, Marek Herde, Sara Beery

    cs.LG · cs.SD

    Ecological monitoring increasingly relies on machine learning models, whose performance depends on the quality and quantity of labelled data. However, obtaining these labels is costly, particularly in passive acoustic monitoring, where vast amounts of data are collected but only a small proportion can feasibly be annotated. Active learning addresses this bottleneck by prioritizing which samples should be labelled. However, progress is...

    arxiv.org/abs/2609.15255 · PDF

  47. 47

    Bandits with Probing: Optimal Regret and the Limits of Winner Feedback

    Yongjie Guan

    cs.LG · cs.DS · stat.ML

    A learner probes at most $k$ of $n$ arms each round, receives the maximum of their rewards in $[0,1]$, and competes with the best fixed arm. When does the probing advantage pay for learning? We determine two minimax laws. Under independent stochastic rewards with winner feedback (the maximum and a winning label), or on arbitrary fixed sequences given a single signed contrast between block maxima, the minimax regret has order...

    arxiv.org/abs/2609.15248 · PDF

  48. 48

    ProtoGuide: Prototype-Driven Guidance for Class-Conditional Graph Generation

    Salvatore Romano, Marco Grassia, Pietro Liò, Giuseppe Mangioni

    cs.LG · cs.AI · physics.soc-ph

    Discrete diffusion models are a prominent family for graph generation, but standard class-conditional mechanisms embed the class signal in the denoiser during training, tying the conditioning mechanism to the trained model. Classifier guidance avoids this coupling in continuous domains by steering a frozen model with a classifier's gradient, but discrete graph diffusion samples discrete edge states, so gradients cannot propagate through the...

    arxiv.org/abs/2609.15239 · PDF

  49. 49

    Convergence rates for generative drifting flows: fixed-scale obstructions and multihead acceleration

    Arthur Stéphanovitch, Eddie Aamari

    cs.LG

    Drifting models offer a promising route to faster generative AI: they perform gradual transport during training, while generating new samples in a single step. This paper asks whether the underlying drifting process can converge rapidly to a target distribution under ideal conditions, before finite-data or optimization effects are introduced. We show that its convergence rate depends critically on how it handles spatial scale. With a single...

    arxiv.org/abs/2609.15193 · PDF

  50. 50

    Rethinking Correctness for Uncertainty Estimation in Clinical Prediction with Vision-Language Models

    Mingcheng Zhu, Jinning Liang, Tingting Zhu

    cs.LG

    Vision-language models are increasingly explored for clinical prediction from electronic health records and medical images, where identifying unreliable predictions is important for safe deployment. Uncertainty estimation (UE) enables detecting such predictions, but its evaluation depends on a correctness criterion that determines whether each model output is correct. If this criterion disagrees with human judgement or distorts downstream UE...

    arxiv.org/abs/2609.15180 · PDF

  51. 51

    Temporal Self-Distillation: Faster Inference in Discrete Diffusion Language Models

    Shijian Xu, Andrea Miele, Metod Jazbec, Volker Roth, Eric Nalisnick, Ilija Bogunovic

    cs.LG

    Diffusion language models (dLLMs) promise fast inference by generating multiple tokens in parallel, but suffer severe performance degradation when parallel decoding is pushed too aggressively. We introduce Temporal Self-Distillation (TSD), a simple on-policy method that trains dLLMs for fast inference by distilling predictions across time. Specifically, TSD distills the model's denoising distribution at earlier timesteps toward its...

    arxiv.org/abs/2609.15177 · PDF

  52. 52

    Nearly Minimax-Optimal Regret for Linear Contextual Bandits with Arbitrary Adaptive Action Sets

    Tianyuan Jin

    cs.LG · cs.GT

    We study stochastic linear contextual bandits with arbitrary action menus that may depend on the fixed parameter and the interaction history. We establish matching upper and lower bounds, up to logarithmic factors. Let $d$ be the dimension, $K$ be the menu size, and $T$ the time horizon. For $2\le K\le d$, we prove an upper bound $\widetilde O(K^{1/4}\sqrt{dT})$. When $T\ge d^2$, we further prove a lower bound $Ω(K^{1/4}\sqrt{dT})$. Thus, for...

    arxiv.org/abs/2609.15170 · PDF

  53. 53

    Multi-source Transfer Learning of Time Series with a Shapelet-based Distance Measure

    Jiseok Lee, Brian Kenji Iwana

    cs.LG

    Transfer learning is an effective technique for addressing data scarcity in deep learning for time series classification, but its success depends on the selection of source datasets. Conventional transferability estimation methods are often computationally expensive, as they require fully pre-training a model on each potential source dataset to assess its suitability. This paper introduces a novel, training-free source selection method named...

    arxiv.org/abs/2609.15148 · PDF

  54. 54

    Omni-Streaming Thinking

    Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li, Yiwen Guo, Yongqi Zhang, Difan Zou

    cs.LG

    Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If that interpretation enters memory as a fact, later reasoning can keep relaying it even after audio contradicts it. We call this failure premature cross-modal commitment. We propose Omni-Streaming Thinking (OST), which...

    arxiv.org/abs/2609.15128 · PDF

  55. 55

    Refinement-based Flow Policy Optimization

    Bumgeun Park, Hyukjun Yang, Donghwan Lee

    cs.LG · cs.AI

    Flow-based policies offer an expressive representation for online reinforcement learning, but conventional flow matching requires samples drawn from the distribution to be modeled. This poses a challenge when the desired action distribution is defined only implicitly by a Q-function, since directly sampling actions from the resulting distribution is generally intractable. We propose Refinement-Based Flow Policy Optimization (RFPO), a novel...

    arxiv.org/abs/2609.15123 · PDF

  56. 56

    Beyond Numerical Time Series: A Unified Benchmark for Multimodal Forecasting with Heterogeneous Context

    Peng Chen, Zhihao Zhuang, Hongzhou Chen, Junhao Huang, Aiping Yang, Mengsen Wu, Yiding Liu, Xilin Dai, Zewei Dong

    cs.LG · cs.AI

    Most time series forecasting benchmarks remain numerical-centric and provide limited support for evaluating contextual information that shapes real-world temporal dynamics. Existing multimodal benchmarks also suffer from limited data and context coverage, fragmented evaluation settings, and overreliance on aggregate evaluation. In this paper, we propose \textbf{MUSE-Bench}, a unified benchmark for multimodal time series forecasting with...

    arxiv.org/abs/2609.15087 · PDF

  57. 57

    $\mathbb{SL}(n)$ Representation Learning: An Intrinsic Mixed-Curvature Space with Higher Curvature Capacities and Deeper Order-Aware Composition

    Xingrun Li, Yusuke Mukuta, Xin Yang, Yinyu Ye, Tatsuya Harada

    cs.LG

    Mixed-curvature representation learning seeks to capture rich geometric structures that cannot be adequately modeled by a single curvature regime. Existing approaches largely rely on product manifolds, which require manually specifying how different curvature spaces are combined and separate their curvature contributions across factors. We introduce the $\mathbb{SL}(n)$ space, a representation geometry defined by the simple $\det(A)=1$...

    arxiv.org/abs/2609.15083 · PDF

  58. 58

    Ensemble-Conditioned Molecular Design

    Ross Irwin, Alessandro Tibo, Jon Paul Janet, Simon Olsson

    cs.LG · cs.NE

    Molecular design is typically approached as a problem of finding molecules which can adopt a single bioactive conformation. In reality, molecules occupy a distribution over conformations, and many of the properties which determine whether a candidate is viable depend on that distribution rather than on any single conformer. We reframe molecular design as an optimisation of both the modes and properties of molecules' conformational ensembles,...

    arxiv.org/abs/2609.15077 · PDF

  59. 59

    Branched Optimal Transport Amortization

    Semyon Semenov, Viktor Kovalchuk, Meir Roketlishvili, Albert Baichorov, Fakhri Karray, Martin Takac, Arip Asadulaev

    cs.LG · cs.AI

    Methods of Branched Optimal Transport (BOT) mimic the economy and efficiency of natural tree-like structures, such as those found in rivers and biological systems. These methods are widely applicable for designing efficient networks in society, from river basins and blood vessels to mail and gas distribution systems. However, they remain understudied in the context of designing deep generative models, particularly at a large scale. Standard...

    arxiv.org/abs/2609.15072 · PDF

  60. 60

    Sensory Precision Inference for Multimodal Arbitration under Uncertainty

    Tin Mišić, Takato Horii

    cs.LG · cs.NE

    Autonomous agents operating on multisensory data cannot assume that all sensory modalities remain consistently informative. In real environments, sensory streams are frequently corrupted by noise, missing data, or inter-modal incongruence, requiring adaptive arbitration between competing sensory hypotheses. While active inference provides a principled framework for uncertainty-guided inference, the role of dynamically inferred sensory...

    arxiv.org/abs/2609.15065 · PDF

  61. 61

    What Does an LLM Learn from Reinforcement Learning? A Mechanistic Interpretability Perspective with Fixed-SAE Track

    Lingheng Du, Yiming Tang, Xufeng Duan, Dianbo Liu

    cs.LG

    Reinforcement learning (RL) is widely utilized in large language model training to improve targeted capabilities, yet how RL reshapes a model remains poorly understood. Prior attempts to explain how RL works largely offer behavioral perspectives, leaving open what RL gives a model at the representation level: can RL create genuinely novel features, and which existing features does it enhance or suppress? Recent developments in mechanistic...

    arxiv.org/abs/2609.15064 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.