cs.LG · 2026-07-02 · No. 41

Machine Learning, 2026-07-02.

68 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

68 entries
  1. 01

    Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training

    Zijian Zhang, Rizhen Hu, Athanasios Glentis, Dawei Li, Chung-Yiu Yau, Hongzhou Lin, Mingyi Hong

    cs.LG · cs.CL

    Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little is understood about how RL adaptation is distributed across transformer layers. Existing approaches typically update all model parameters uniformly, implicitly assuming that every layer contributes similarly to the gains obtained during RL post-training. In this work, we challenge this assumption through a systematic layer-wise...

    arxiv.org/abs/2607.01232 · PDF

  2. 02

    Language-Critique Imitation Learning from Suboptimal Demonstrations

    Chih-Han Yang, Dai-Jie Wu, Yun-Ping Huang, Ping-Chun Hsieh, Kenneth Marino, Shao-Hua Sun

    cs.LG · cs.AI

    Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estimates, discriminator scores, or importance weights. These scalar signals are inherently limited, as they cannot explicitly express intermediate reasoning about task progress, failure modes, or corrective actions. We propose a language-critique framework for imitation learning from suboptimal demonstrations...

    arxiv.org/abs/2607.01225 · PDF

  3. 03

    TiRex-2: Generalizing TiRex to Multivariate Data and Streaming

    Patrick Podest, Marco Pichler, Elias Bürger, Levente Zólyomi, Bernhard Voggenberger, Wilhelm Berghammer, Daniel...

    cs.LG

    We introduce TiRex-2, a recurrent xLSTM-based time series foundation model that generalizes the univariate TiRex to multivariate forecasting with both past and future covariates. Real-world forecasting is inherently sequential: observations arrive continuously, variables evolve jointly, and a subset of covariates is known ahead of time. Existing Transformer-based time series foundation models capture cross-variate dependencies but incur...

    arxiv.org/abs/2607.01204 · PDF

  4. 04

    Quantum vs. Classical Machine Learning: A Unified Empirical Comparison

    Chuanming Yu, Jiaming Liu, Zihao Ge, Xiongfei Wu, Lulu Zhu, Pengzhan Zhao, Jianjun Zhao

    cs.LG

    Quantum computing has emerged as a promising computational paradigm for machine learning (ML), with the potential to offer computational advantages over classical approaches. At this stage, the evidence supporting the performance and advantages of quantum machine learning (QML) models relative to classical models is insufficient.To address this gap, this paper presents an empirical study on the performance of QML models and their classical...

    arxiv.org/abs/2607.01197 · PDF

  5. 05

    Neural Certificate Pricing for Combinatorial Optimization Problems

    Jingyi Chen, Xinyuan Zhang, Xinwu Qian

    cs.LG

    Combinatorial optimization (CO) problems are difficult because certifiable discrete structure induces exponential search. One needs to search over the set exponentially many candidates to certify optimality, however, the structural feasibility of a path, packing, or cover can be verified in polynomial time once supplied. In this study, we introduce Neural Certificate Pricing (NCP) that exploits this asymmetry under an unsupervised learning...

    arxiv.org/abs/2607.01185 · PDF

  6. 06

    Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations

    Mehul Damani, Isha Puri, Idan Shenfeld, Jacob Andreas

    cs.LG · cs.AI · cs.CL

    RL with verifiable rewards (RLVR) has emerged as a powerful paradigm for training LMs on tasks with well-defined success metrics, such as code generation and mathematical reasoning. However, current RLVR methods optimize only what can be objectively scored, often neglecting subjective, non-verifiable aspects of human-like outputs, such as style and structure. This limitation leads to well-documented failure modes such as diversity collapse,...

    arxiv.org/abs/2607.01181 · PDF

  7. 07

    QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling

    Michael Y. Li, Anthony Zhan, Kanishk Gandhi, Noah D. Goodman, Emily B. Fox

    cs.LG · cs.CL

    Scaling inference compute, by generating many parallel attempts per problem, is a costly but reliable lever for improving language model capabilities. By default these attempts are generated independently, wasting inference compute on redundant solutions. This waste seems unavoidable. After all, independence is what makes parallel sampling trivial to scale. However, this tradeoff is not fundamental: there is a rich design space of samplers...

    arxiv.org/abs/2607.01179 · PDF

  8. 08

    Decision-Aware Training for Sample-Based Generative Models

    Kornelius Raeth, Nicole Ludwig

    cs.LG · stat.ML

    Sample-based generative models are increasingly used for probabilistic forecasting in high-stakes decision settings, yet their training objectives are blind to the decision maker's cost structure. These models are commonly trained with strictly proper scoring rules, such as the energy score, which allocate their training signal in proportion to data density, with no awareness of where forecast errors are most costly for downstream decisions....

    arxiv.org/abs/2607.01171 · PDF

  9. 09

    Efficient Compression of Structured and Unstructured Volumes via Learned 3D Gaussian Representation

    Landon Dyken, Sharmistha Chakrabarti, Nathan Debardeleben, Steve Petruzza, Qi Wu, Will Usher, Sidharth Kumar

    cs.LG

    Recent work has shown that implicit neural representations (INRs) can be trained to effectively compress structured and unstructured volume data, allowing for direct data querying with a reduced memory footprint. However, as existing INRs for unstructured volumes do not encode geometry, they require partial mesh storage for later sampling, limiting achievable compression. At the same time, novel view synthesis methods have shown that explicit...

    arxiv.org/abs/2607.01164 · PDF

  10. 10

    A Lightweight Self-Supervised Learning Framework for Multivariate Time Series using Hierarchical-JEPA on ECG Data

    Siwon Kim

    cs.LG · eess.SP

    Data analysis in the medical domain often encounters scenarios involving a limited target dataset and a large, unannotated dataset with a general distribution. Under such circumstances, self-supervised learning (SSL) methods are highly effective for utilizing large datasets, making them a popular choice for electrocardiogram (ECG) analysis. This work presents the Event Reconstruction Joint-Embedding Predictive Architecture (ER-JEPA), a...

    arxiv.org/abs/2607.01145 · PDF

  11. 11

    Sequentially-Controlled Interactive Multi-Particle Flow-Maps for Online Feedback-Driven Search

    Binglin Ji, Anindya Sarkar, Hengchang Lu, Jens Sjölund, Yevgeniy Vorobeychik

    cs.LG · cs.AI · cs.CE

    While generative models have enabled training-free reward alignment, current methods typically excel in local exploration within narrow regions of the underlying distribution. These approaches struggle when preferences are unknown a priori and only revealed through sequential feedback-a scenario demanding broad exploration to uncover high-utility regions. To address this, we propose Sequentially-Controlled Interactive Multi-Particle Flow-Maps...

    arxiv.org/abs/2607.01144 · PDF

  12. 12

    GAIA: Geometry-Adaptive Operator Learning for Forward and Inverse Problems

    Meenakshi Krishnan, Pranav Pulijala, Ke Chen, Haizhao Yang, Ramani Duraiswami

    cs.LG · math.NA

    Operator learning for partial differential equations (PDEs) on arbitrary geometries builds fast neural surrogates for large-scale simulation. Although recent geometry-adaptive neural operators have made substantial progress, they are mainly designed for forward problems in which inputs and outputs share the same spatial domain. This limits their applicability for boundary value problems (BVPs) and inverse problems, where inputs and outputs...

    arxiv.org/abs/2607.01128 · PDF

  13. 13

    ZO-Act: Efficient Zeroth-Order Fine-Tuning via One-Shot Activation-Informed Low-Rank Subspaces

    Xun Dong, Yibo Xu, Naigang Wang, Xin Li, Penghang Yin, Zi Yang

    cs.LG · math.OC

    Zeroth-order (ZO) optimization enables fine-tuning large language models when backpropagation is unavailable or memory-prohibitive, but existing methods often perturb full model weights or randomly constructed low-dimensional subspaces, yielding high-variance estimates and limited performance. We propose ZO-Act, an activation-informed ZO fine-tuning method that restricts perturbations to a fixed low-rank subspace derived from input...

    arxiv.org/abs/2607.01125 · PDF

  14. 14

    Muon as a Residual Connection

    Hao Huang

    cs.LG · cs.AI

    Muon has recently emerged as one of the most effective optimizers for training large neural networks, yet its empirical success has been explained from several different perspectives. In this paper, we propose a simple mechanistic interpretation: Muon can be understood as an implicit residual connection during training. Specifically, orthogonalizing the update can sacrifice some immediate gradient fidelity while improving representation...

    arxiv.org/abs/2607.01124 · PDF

  15. 15

    SynLaD: Latent Diffusion for Generating Synthesizable Molecules Conditioned on 3D Pharmacophore Profiles

    Miruna Cretu, John Bradshaw, Patricia Suriana, Saeed Saremi, Omar Mahmood, Kirill Shmilovich, Kangway Chuang, Vishnu...

    cs.LG

    We present SynLaD, a latent diffusion framework for small-molecule generation that unifies ligand-based drug design objectives (what to make) with synthetic accessibility (how to make it). Current models typically optimize one objective at the expense of the other, creating a bottleneck for discovering high-scoring and synthesizable molecules. SynLaD combines reaction-constrained generation with pharmacophore-conditioned 3D design by learning...

    arxiv.org/abs/2607.01105 · PDF

  16. 16

    CausalMix: Data Mixture as Causal Inference for Language Model Training

    Zinan Tang, Yukun Zhang, Shaomian Zheng, Zhuoshi Pan, Qizhi Pei, Dingnan Jin, Jun Zhou, Yujun Wang, Biqing Huang

    cs.LG · cs.AI · cs.CL

    In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely on the assumption of static data distributions. As a result, when the underlying data pool shifts, these methods require costly retraining from scratch. This limitation restricts their ability to scale seamlessly from small settings to larger data pools and model...

    arxiv.org/abs/2607.01104 · PDF

  17. 17

    Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

    Jingwei Song, Haofeng Xu, Jie Xiao, Chengke Bao, Jingwei Shi, Pengbin Feng, Weixun Wang, Yuhang Han, Chuan Wu,...

    cs.LG · cs.AI

    High-throughput RLHF systems often decouple rollout generation from policy optimization, leading to the use of stale rollouts during learner updates. In this work, we study the effect of such staleness in asynchronous GRPO. We make the behavior policy explicit in the GRPO surrogate objective and distinguish between the surrogate-gradient mapping used by the learner and the true total derivative of a distribution-dependent population...

    arxiv.org/abs/2607.01083 · PDF

  18. 18

    When Context Compensates for Sparse Event History: AlphaEarth for Spatio-Temporal Point-Process Forecasting

    Yahya Aalaila, Mouad Elhamdi, Gerrit Großmann, Daniel Jenson, Elizaveta Semenova, Sebastian Vollmer

    cs.LG

    Spatio-temporal point-process models must often generalise across space when local event histories are sparse. We study whether exogenous spatial context can compensate in such regimes. Using a fixed log-Gaussian Cox process backbone, we compare an event-only model with the same model augmented by AlphaEarth embeddings as linear spatial context. We evaluate spatial transfer on emergency medical services (EMS) forecasting across eight held-out...

    arxiv.org/abs/2607.01082 · PDF

  19. 19

    Balancing Expressivity and Learnability in Quantum Kernel Bandit Optimization

    Yuqi Huang, Vincent Y. F. Tan, Sharu Theresa Jose

    cs.LG · cs.IT

    We investigate Gaussian process (GP) bandit optimization with quantum kernels, assuming the mean reward function lies in the reproducing kernel Hilbert space (RKHS) induced by the quantum kernel. This setting is motivated by NISQ-era tasks such as quantum control, state preparation and variational quantum algorithms. While quantum kernels can offer a `quantum advantage' via domain-specific inductive biases, naïvely using full,...

    arxiv.org/abs/2607.01080 · PDF

  20. 20

    GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache

    Soosung Kim, Minjae Park, Eui-Young Chung, Jaeyong Chung

    cs.LG

    The deployment of Large Language Models (LLMs) with extended context windows is increasingly constrained by the linear growth of Key-Value (KV) cache memory. Vector Quantization (VQ), particularly Residual Quantization (RQ), is a promising approach for pushing KV cache storage toward the sub-1-bit regime by progressively encoding residuals with small codebooks. However, most VQ methods still rely on standard $\ell_2$ $K$-means as the core...

    arxiv.org/abs/2607.01065 · PDF

  21. 21

    The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

    Andrzej Szablewski, Gabriel Konar-Steenberg, Raffaello Fornasiere, Nikita Menon, Stefan Heimersheim

    cs.LG

    Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box interpretability techniques. Current MOs are typically constructed via post-hoc supervised fine-tuning (SFT) on behavioural transcripts or synthetic documents. Prior research has shown that interpretability methods can easily identify hidden behaviours in these MOs. However, recent work...

    arxiv.org/abs/2607.01033 · PDF

  22. 22

    Seahorse: A Unified Benchmarking Framework for Spatiotemporal Event Modeling

    Yahya Aalaila, Gerrit Großmann, Sebastian Vollmer

    cs.LG

    Spatiotemporal point processes (STPPs) model event data in continuous time and space, with applications in mobility, epidemiology, and public safety. Recent neural STPPs span expressive intensity models, conditional density models, continuous-time latent dynamics, normalizing-flow spatial decoders, and score-based generative mechanisms. Yet comparison remains fragile because implementations differ in preprocessing, coordinate normalization,...

    arxiv.org/abs/2607.01022 · PDF

  23. 23

    Generative Model Proposal based Particle Filtering for Data Assimilation

    Chandni Nagda, Mayank Shrivastavam Gudrun Thorkelsdottir, Gan Zhang, Morteza Mardani, Arindam Banerjee

    cs.LG

    Data assimilation models state dynamics conditioned on sequential observations, and has wide-ranging scientific applications. In the filtering setting, the goal is to model the posterior over the current state given all observations so far. Classical solutions typically make simplifying distributional or functional assumptions, e.g., linear-Gaussian systems, which can be inaccurate in many scenarios. In principle, particle filters (PFs)...

    arxiv.org/abs/2607.01012 · PDF

  24. 24

    Automatic Detection of Stress from Speech in the Trier Social Stress Test

    Hanna Drimalla, Wieland R. Cremer, Christine Kraus, Oliver T. Wolf

    cs.LG

    Automatically detecting stress in speech provides an unobtrusive way to gain insights relevant to behavioral research or clinical assessment. This study investigates the automatic differentiation between a stressful and non-stressful situation, and the prediction of physiological and affective stress responses. Speech data was collected from 50 participants who either completed the Trier Social Stress Test (TSST) or a non-stressful control...

    arxiv.org/abs/2607.00986 · PDF

  25. 25

    LeNEPA: No-Augmentation Next-Latent Prediction for Time-Series Representation Learning

    Alexander Chemeris, Ming Jin, Randall Balestriero

    cs.LG

    Time series are central to modern data mining applications, from industrial telemetry and server metrics to finance and physiology, yet time-series self-supervised learning often depends on view and augmentation choices that encode domain-specific invariances. We study how an SSL recipe behaves when its method-specific configuration is reused unchanged after the pretraining signal family changes, framing this as a fixed-recipe stress test...

    arxiv.org/abs/2607.00958 · PDF

  26. 26

    Aionoscope: Debugging Latent-State Accessibility in Time-Series Representations

    Alexander Chemeris, Ming Jin, Randall Balestriero

    cs.LG · cs.AI

    Time-series models are often evaluated by what they can forecast or classify, but those scores do not show whether their representations preserve the process state a user may want to inspect: event timing, phase, amplitude, frequency, or regime variables. We introduce Aionoscope, a generator-based diagnostic tool for debugging latent-state accessibility in frozen time-series representations. Aionoscope separates process generation from...

    arxiv.org/abs/2607.00956 · PDF

  27. 27

    Diffeomorphic Optimization

    Ludwig Winkler, Andrew Leaver-Fay, Joseph Kleinhenz, Pan Kessel

    cs.LG

    Generative models learn data distributions that reside on a low-dimensional manifold within a higher-dimensional ambient space. Optimizing differentiable objectives on this manifold is challenging: the ambient loss landscape is high-dimensional, rugged, and non-convex. Direct gradient descent, blind to the manifold's geometry, quickly drifts off it. Diffeomorphic optimization starts from the observation that diffusion and flow models provide...

    arxiv.org/abs/2607.00947 · PDF

  28. 28

    Explainable AI for Cancer Drug Response Prediction: Beyond Univariate Feature Attributions

    Martino Ciaperoni, Margherita Lalli, Simone Piaggesi, Martina Varisco, Francesco Carli, Riccardo Guidotti, Dino...

    cs.LG

    Predicting cancer drug response from transcriptomic profiles is a cornerstone of precision oncology, yet the scientific value of machine learning models hinges not solely on predictive accuracy, but also on their capacity to generate reliable biological insights. Current explainability approaches in this setting are computationally costly, lack robustness, and reduce complex drug response to univariate gene importance scores, overlooking the...

    arxiv.org/abs/2607.00931 · PDF

  29. 29

    Human-Machine Collaboration on Generative Meta-Learning: Model and Algorithm

    Midhun Parakkal Unni, Samuel Kaski

    cs.LG · cs.AI

    Generalizing machine learning models to environments that differ from their training distribution remains a critical hurdle, particularly when data from the target domain is entirely or partially unavailable. We propose Generative Meta-Learning with Human Feedback (GMHF), a novel framework that bridges this domain gap by leveraging expert intuition to guide data synthesis. Grounded in a theoretical analysis of generalization error, we derive...

    arxiv.org/abs/2607.00926 · PDF

  30. 30

    Valdi: Value Diffusion World Models

    Christopher Lindenberg, Kashyap Chitta

    cs.LG · cs.AI

    World models can enable Model Predictive Control (MPC), but this requires dynamics prediction that is both fast enough for online use and expressive enough to represent uncertain futures. Diffusion models offer a natural mechanism for modeling uncertain dynamics, yet their iterative inference procedure makes them difficult to use for low-latency latent planning. We bridge this gap with Value Diffusion World Models (Valdi), combining...

    arxiv.org/abs/2607.00917 · PDF

  31. 31

    Beyond Activation Alignment:The Alignment-Diversity Tradeoff in Task-Aware LLM Quantization

    Fei Wang, Chao Xue, Taoran Liu, Li Shen, Ye Liu, ChangXing Ding

    cs.LG

    Mixed-precision quantization (MPQ) has become a key technique for deploying large language models under stringent memory and compute constraints. We first identify a phenomenon that we term the Perplexity Illusion: layers ranked as important by perplexity-based sensitivity show little rank correlation with those that are most influential for complex reasoning performance, with Kendall $τ\approx 0$ in our analysis. We further reveal an...

    arxiv.org/abs/2607.00908 · PDF

  32. 32

    Constrained Bayesian Optimisation with Multiple Information Sources

    Hauke Maathuis, Roeland De Breuker, Saullo Castro, Maike Osborne

    cs.LG

    Bayesian Optimisation (BO) under unknown constraints is particularly challenging when feasible regions are small. In such settings, existing methods that typically rely solely on evaluations of the true objective and constraints struggle to efficiently explore the design space. However, many real-world applications offer auxiliary data sources (e.g. surrogate models or simplified simulations) that can support early exploration. Despite this...

    arxiv.org/abs/2607.00865 · PDF

  33. 33

    Spectroscopy Analysis with Machine Learning Regression for the Quantification of Carbon and Nitrogen Contents in Inceptisol and Oxisol Soil Types: Comparing Different Preprocessing and Validation methods as well as Feature Importance

    Vinicius Herique Kieling, Guilherme Macedo Baggio, Felipe Augusto Bueno Rossi, Marco Antonio de Castro Barbosa,...

    cs.LG

    Near-Infrared (NIR) spectroscopy has emerged as a promising alternative to traditional soil analysis methods, offering advantages such as speed, low cost, and non-destructive testing. This work proposes a machine learning (ML) approach to calibrate predictive models for carbon (C) and nitrogen (N) content in Oxisols and Inceptisols, utilizing NIR spectral data acquired with a portable MyNIR device. Various preprocessing methods were...

    arxiv.org/abs/2607.00834 · PDF

  34. 34

    From Pixels to Temporal Correlations: Learning Informative Representations for Reinforcement Learning Pre-training

    Jinwen Wang, Youfang Lin, Xiaobo Hu, Siyu Yang, Sheng Han, Shuo Wang, Kai Lv

    cs.LG

    Unsupervised pre-training on large-scale datasets has demonstrated significant potential for improving the sample efficiency and performance of Reinforcement Learning (RL). Given the large-scale action-free internet videos, existing methods utilize single-step transition prediction and image reconstruction to learn representations. However, these methods prefer to preserve large-proportion stationary information in the pixel space, neglecting...

    arxiv.org/abs/2607.00811 · PDF

  35. 35

    Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos

    Jinwen Wang, Youfang Lin, Xiaobo Hu, Shuo Wang, Kai Lv

    cs.LG

    Pre-training on large-scale videos to improve reinforcement learning efficiency is promising yet remains challenging. Existing methods typically treat the agent as an indivisible entity, modeling motion patterns globally. Such global modeling is tightly coupled with the morphology, hindering transfer across domains. In contrast, despite the vast disparity in global motions, the local components exhibit similar motion patterns across different...

    arxiv.org/abs/2607.00808 · PDF

  36. 36

    Task-Relevant Representation Decoupling for Visual Reinforcement Learning Generalization

    Jinwen Wang, Youfang Lin, Xiaobo Hu, Qian Xu, Shuo Wang, Zhuo Chen, Kai Lv

    cs.LG

    Visual Reinforcement Learning (VRL) has achieved considerable success in solving control tasks. However, generalizing learned policies to new environments remains a major challenge, as agents often overfit to task-irrelevant features in the training environment. To solve this problem, we introduce the concept of decoupling observations into task-relevant and task-irrelevant representations. Building on this idea, we propose a self-supervised...

    arxiv.org/abs/2607.00796 · PDF

  37. 37

    Which Metric Reflects the Spelling Rate Accuracy in Event-Related Potential-Based Brain-Computer Interfaces?

    Okba Bekhelifi, Naoual El Djouher Mebtouche

    cs.LG · eess.SP

    For predictive models, the often-reported performance metrics are the loss and accuracy. In synchronous Brain- Computer Interface (BCI) systems, these metrics are informative for most BCI paradigms; however, for Event-Related Potential (ERP) applications the spelling rate, which measures the number of characters correctly selected is more important as it influences the estimation of information transfer rate (ITR) and any related metric...

    arxiv.org/abs/2607.00794 · PDF

  38. 38

    Accelerating Discrete Diffusion Models with Parallel-In-Time Sampling

    Yu Yao, Huanjian Zhou, Andi Han, Wei Huang, Masashi Sugiyama

    cs.LG · cs.DC · cs.DS · math.NA

    Discrete diffusion models are widely used for learning and generating discrete distributions. As the generation process is inherently sequential, the acceleration of sampling is of significant importance. In this work, we parallelize the mainstream $τ$-leaping algorithm for absorbing discrete diffusion in a Continuous-Time Markov Chain (CTMC) framework. By leveraging the continuous-time stochastic integral form of the $τ$-leaping algorithm...

    arxiv.org/abs/2607.00773 · PDF

  39. 39

    MosaicKV: Serving Long-Context LLM with Dynamic Two-D KV Cache Compression

    Sheng Qiang, Ruiwei Chen, Yinpeng Wu, Jinyu Gu, Zhichao Hua, Yubin Xia, Binyu Zang, Haibo Chen

    cs.LG · cs.DC

    Long-context LLM services now sustain prompts with hundreds of thousands to millions of tokens, making the key-value (KV) cache a first-order serving cost. Because the cache grows linearly with context length, it can exhaust GPU memory, force smaller batches, and reduce serving throughput. Prior KV cache compression techniques typically target only the sequence dimension or only the channel dimension, which leaves limited headroom as context...

    arxiv.org/abs/2607.00760 · PDF

  40. 40

    LLM-Guided ODE Discovery and Parameter Inference from Small-Cohort Aggregate Data

    Hanning Yang, Meropi Karakioulaki, Lennart Purucker, Tim Litwin, Cristina Has, Moritz Hess

    cs.LG · cs.AI

    Mechanistic modeling via ordinary differential equations (ODEs) provides interpretable descriptions of complex dynamics and enables inference of underlying mechanisms, which is particularly valuable in clinical settings. However, in rare diseases, both the structure and parameters of the model are typically unknown, while individual-level data is scarce, noisy, heterogeneous, and subject to privacy constraints. In such settings,...

    arxiv.org/abs/2607.00733 · PDF

  41. 41

    Detecting the Undetectable: Enhancing Unsupervised time series Anomaly Detection via Active Learning

    Seung Hun Han, Hyeongwon Kang, Jinwoo Park, Pilsung Kang

    cs.LG · cs.AI

    Despite the increasing sophistication of industrial AI systems, the ability to reliably detect subtle and noisy anomalies in complex time series data remains a critical yet unresolved challenge. In large-scale industrial applications, labeling time series data is often prohibitively expensive and time-consuming, making unsupervised learning a practical and widely adopted approach. However, existing unsupervised methods frequently struggle to...

    arxiv.org/abs/2607.00720 · PDF

  42. 42

    Generative Refinement for Low-Budget Black-Box Optimization

    Edouard R. Dufour, Pascal Fua

    cs.LG

    Black-box optimization is a fundamental science and engineering tool that makes it possible to optimize objectives without gradient information. Unfortunately, as it often requires many function evaluations, it can be challenging when each one is costly. This is especially true when the evaluation function is noisy or failure-prone, and when high-performing solutions are confined to thin, curved, or disconnected regions of the search space....

    arxiv.org/abs/2607.00691 · PDF

  43. 43

    AdaBoosting Text Prompts for Vision-Language Models

    Seokhee Jin, Changhwan Sung, Sunung Mun, Hoyoung Kim, Jungseul Ok

    cs.LG

    The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts. Handcrafted templates and Large Language Model (LLM)-generated descriptions not only make predictions more interpretable, but also enable reuse of the same prompts across heterogeneous VLMs. Recent works construct task-adapted text prompts with a small number of labeled images. However, existing few-shot text prompting methods do...

    arxiv.org/abs/2607.00684 · PDF

  44. 44

    Distributed Online Bandit Submodular Maximization with Bounded Sampling Violations

    Bin Du, Chang Liu, Dingqi Zhu, Lintao Ye, Dengfeng Sun

    cs.LG

    We study distributed online submodular maximization under partition matroid constraints, in which multiple agents select a limited number of actions from their own subsets sequentially to maximize the cumulative value of a sequence of objective functions. We develop a unified algorithmic framework that accommodates full-information and bandit feedback models. For both feedback models, we prove that the proposed algorithms achieve sublinear...

    arxiv.org/abs/2607.00680 · PDF

  45. 45

    Multi-Label Node Classification with Label Influence Propagation

    Yifei Sun, Zemin Liu, Bryan Hooi, Yang Yang, Rizal Fathony, Jia Chen, Bingsheng He

    cs.LG · cs.AI

    Graphs are a complex and versatile data structure used across various domains, with possibly multi-label nodes playing a particularly crucial role. Examples include proteins in PPI networks with multiple functions and users in social or e-commerce networks exhibiting diverse interests. Tackling multi-label node classification (MLNC) on graphs has led to the development of various approaches. Some methods leverage graph neural networks (GNNs)...

    arxiv.org/abs/2607.00671 · PDF

  46. 46

    Loss Smoothing for Stable Adaptation Under Distribution Shift

    Darshan Patil, Ekaterina Lobacheva, Razvan Pascanu, Sarath Chandar

    cs.LG · cs.AI

    In settings such as fine-tuning and reinforcement learning, neural networks are often adapted under distribution shift. Standard adaptation methods typically optimize the target objective directly, inducing an abrupt change from the source training objective. This abrupt transition can distort learned representations, including features that may still be useful for the new task. We investigate whether a more gradual transition can improve...

    arxiv.org/abs/2607.00634 · PDF

  47. 47

    Measuring Dead Directions: Decomposing and Classifying Singular Structure off Canonical Alignment

    Tejas Pradeep Shirodkar

    cs.LG

    We give a descent-free, alignment-free measurement of singular structure on trained networks. At a single frozen checkpoint the read recovers the order $k$ of each dead direction from the directional-Fisher rate, the master invariant from which the per-direction learning coefficient $1/(2k)$ follows exactly, in whatever basis the optimizer left. The same read classifies each direction, separating a genuine singularity, whose order the...

    arxiv.org/abs/2607.00603 · PDF

  48. 48

    Decision-focused Sparse Tangent Portfolio Optimization

    Haeun Jeon, Seunghoon Choi, Hyunglip Bae, Yongjae Lee, Woo Chang Kim

    cs.LG

    Sparse tangent portfolio optimization aims to learn an interpretable, low-cardinality portfolio in the tangency direction of the mean-variance frontier. However, the associated cardinality-constrained formulation is NP-hard, and standard predict-then-optimize pipelines often misalign forecasting accuracy with downstream portfolio quality. We propose an end-to-end decision-focused learning framework that reformulates Sharpe ratio maximization...

    arxiv.org/abs/2607.00581 · PDF

  49. 49

    Group-Equivariant Poincaré Convolutional Networks

    Aiden Durrant, Rahul Baburajan, Georgios Leontidis

    cs.LG · cs.AI

    While recent advancements like the Poincaré ResNet have demonstrated the potential of learning visual representations directly in hyperbolic space, their optimisation remains hampered by the computationally intensive nature of Riemannian gradients and the strict boundaries of the manifold. Furthermore, standard hyperbolic networks treat spatial transformations of the same object as distinct hierarchical concepts, leading to redundant...

    arxiv.org/abs/2607.00556 · PDF

  50. 50

    Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition

    Zhiqi Li, Wen Zhang, Bo Zhu

    cs.LG · cs.AI · cs.CV

    Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by directly learning long-range transport maps between noise and data. However, these models are typically deterministic, which makes them difficult to optimize with reinforcement learning (RL) post-training methods that require stochastic trajectories and well-defined likelihood ratios. Existing SDE-based stochasticization techniques are designed for...

    arxiv.org/abs/2607.00535 · PDF

  51. 51

    Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization

    Xuefeng Liu, Mingxuan Cao, Qinan Huang, Thomas Brettin, Rick Stevens, Le Cong

    cs.LG · cs.AI · q-bio.BM · stat.ML

    Scientific reasoning is an increasingly important capability of large language models, yet improving the robustness and efficiency of training such reasoning remains a key open challenge. We study this problem in instruction-based molecular optimization, where answer-only supervised fine-tuning (SFT) collapses multi-step reasoning and reinforcement learning with verifiable rewards (RLVR) suffers from sparse feedback. Reference-guided Policy...

    arxiv.org/abs/2607.00531 · PDF

  52. 52

    From Structural Equation Modelling to Double Machine Learning: Robustness Analysis for Survey-Based Research

    Ka Ching Chan, Qiana Liu, Sanjib Tiwari, Ranga Chimhundu

    cs.LG · stat.ML

    Structural equation modelling (SEM) is widely used in survey-based business and information systems research to assess latent constructs and theory-driven structural relationships. However, SEM path significance is obtained within a particular model specification and may not show whether findings remain stable under alternative estimation frameworks. This study develops and demonstrates a staged robustness analysis framework that connects...

    arxiv.org/abs/2607.00512 · PDF

  53. 53

    Prototype Language Models

    Dan Ley, Giang Nguyen, Himabindu Lakkaraju, Julius Adebayo

    cs.LG · stat.ML

    Knowing which training examples drive outputs is fundamental to auditing, correcting, and understanding language models, yet for modern LLMs this remains expensive, approximate, and largely post-hoc. Standard language models generate tokens through a dense network pathway, causing training data's influence to be distributed across parameters rather than organized along explicit, traceable components. We introduce a prototype language model...

    arxiv.org/abs/2607.00510 · PDF

  54. 54

    PAPA: Online Personalized Active Preference Alignment

    Anindya Sarkar, Nasik Muhammad Nafi, Isaac Lyngaas, Muralikrishnan Gopalakrishnan Meena, Yevgeniy Vorobeychik

    cs.LG · cs.AI · cs.CV

    Diffusion models are highly effective at modeling complex data distributions, including images and text. However, in applications like personalized recommender systems, the objective often shifts to modeling specific regions of the distribution that maximize user preferences-initially unknown but gradually uncovered through interactive feedback. This can naturally be framed as a reinforcement learning problem, where the goal is to fine-tune a...

    arxiv.org/abs/2607.00486 · PDF

  55. 55

    Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization

    Peilin Liu, Ding-Xuan Zhou

    cs.LG · stat.ML

    Transformer-based large models have demonstrated remarkable generalization abilities across different tasks by leveraging a context-aware attention module for in-context learning. With richer context, transformers adapt more effectively to the current use case without any parameter updates. However, the quadratic computational and memory complexity with respect to context length significantly slows data processing in softmax transformers....

    arxiv.org/abs/2607.00479 · PDF

  56. 56

    Interpretable vs Learned Encoders for High-Cardinality Fraud Detection

    Xiao Han, Jingjing Liu, Moxuan Zheng, Zhen Zhang, Chenyu Wu

    cs.LG · cs.CE

    A total of seven categorical encoding methods were tested on the IEEE-CIS fraud benchmark dataset (590,540 records, 3.5% positives, 8 high-cardinality columns). The encoders were evaluated using a stratified 5-fold cross-validation (CV) with three repetitions. Five of the encoders had identical frozen LightGBM learners in the downstream phase, allowing for controlled comparisons of their performance to each other. CatBoost and TabNet were...

    arxiv.org/abs/2607.00477 · PDF

  57. 57

    How Early Is Early Enough? Design-Dependent Observation-Window Sufficiency in Subscription Churn Prediction

    Xiao Han, Yao Xiao, Chenyu Wu, Tongchen Zhang

    cs.LG

    How many days of early behavior suffice for subscription churn prediction? In the public KKBox dataset, the early indicator of churn is typically an indicator of someone's contract status; however, when looking in the heavily churned manual-renewal segment, having access to early behavior creates a substantial increase in prediction for that specific segment (PR +0.10 at 120 days). A nine-window sufficiency curve shows a diminishing-returns...

    arxiv.org/abs/2607.00473 · PDF

  58. 58

    MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules

    Tong Xu, Xinzhe Cao, Zhihui Zhu, Keyan Ding, Huajun Chen

    cs.LG · cs.CL

    Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern: the potential safety risks of AI-generated molecules. In practice, many generative models may produce molecules with toxic, reactive, or otherwise hazardous characteristics - posing hidden dangers that remain insufficiently addressed. To address this gap, we introduce MolSafeEval, a benchmark...

    arxiv.org/abs/2607.00464 · PDF

  59. 59

    Gauging, Measuring, and Controlling Critic Complexity in Actor-Critic Reinforcement Learning

    Konstantin Garbers

    cs.LG · cs.AI

    Actor-critic methods depend on learned critics, but critic quality is often evaluated only indirectly through return, temporal-difference error, or value loss. Critic complexity is introduced as an additional diagnostic and intervention dimension for actor-critic reinforcement learning. The analysis uses spectral effective-rank entropy, a rank-like summary of the singular-value distributions of critic weight matrices, to assess critic model...

    arxiv.org/abs/2607.00452 · PDF

  60. 60

    Timesynth: A Temporal Fidelity Framework for Health Signal Digital Twins

    Md Rakibul Haque, Shireen Elhabian, Warren Woodrich Pettine

    cs.LG

    Forecasting models for health-signal digital twins must preserve the oscillatory, frequency, phase, and state-transition dynamics of physiological signals, yet the pointwise metrics used to benchmark them cannot detect when these fundamental properties are lost. We show that this blind spot misranks models: across 11 architectures, models with comparable pointwise error diverge by up to 53° in phase accuracy, equivalent to roughly 123 ms for...

    arxiv.org/abs/2607.00431 · PDF

  61. 61

    Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL

    Jongchan Park, Seungjun Oh, Seungho Baek, Yusung Kim

    cs.LG · cs.AI

    Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks. Despite recent progress, we argue that current off-policy URL methods are limited by two critical, overlooked bottlenecks: (1) non-stationary skill semantics and (2) brittle generalization. To address these challenges, we propose GenDa (Generalizable Data-efficient...

    arxiv.org/abs/2607.00392 · PDF

  62. 62

    SAOT: Self-Supervised Continual Graph Learning with Structure-Aware Optimal Transport

    Yuting Zhang, Yanbei Liu, Zhitao Xiao, Lei Geng, Yanwei Pang, Xiao Wang

    cs.LG · cs.SI

    Self-supervised Continual Graph Learning (CGL) aims to successively learn from a graph sequence with different tasks without label supervision - a paradigm that has attracted widespread attention. Most existing self-supervised CGL methods rely on instance-level consistency objectives that enforce stability of individual node (or node-pair) embeddings. Due to optimizing nodes in isolation, these methods fail to maintain global relational...

    arxiv.org/abs/2607.00377 · PDF

  63. 63

    PRISM: Prioritized Channel Importance with Semi-supervised Domain Adaptation for Cross-Subject EEG Emotion Recognition

    Xin Zhou, Xiang Zhang, Hao Deng, Lijun Yin

    cs.LG

    Electroencephalogram (EEG) captures endogenous brain activity with high temporal fidelity and holds substantial promise for precise emotion decoding. However, channel redundancy and pronounced inter-subject variability remain key obstacles to scalable generalization. To address these limitations, we propose a novel framework termed PRioritized channel Importance with Semi-supervised doMain adaptation (PRISM), enabling label-efficient...

    arxiv.org/abs/2607.00358 · PDF

  64. 64

    K-Inverse-RFM: A Modified RFM that Bridges the Gap to Neural Networks for Data-Corrupted Mathematical Tasks

    Gil Pasternak

    cs.LG · cs.AI

    Recursive Feature Machines (RFMs) are a class of kernel machines that utilize the Average Gradient Outer Product (AGOP) as a mechanism for feature learning. They have been shown to effectively replicate the learning dynamics and feature representations of Feedforward Neural Networks (FNNs) across various settings. However, despite comparable capacity for feature learning and the similarities in the features they acquire, RFMs exhibit...

    arxiv.org/abs/2607.00329 · PDF

  65. 65

    Watermarking for Proprietary Dataset Protection

    John Kirchenbauer, Brian R. Bartoldson, Bhavya Kailkhura, Tom Goldstein

    cs.LG · cs.CL

    A growing body of literature suggests that training data membership inference problems are fundamentally hard tasks in modern language modeling settings. We argue that output watermarking techniques are the right gadget to make training membership tests for generative models more tractable, based on prior results showing that language models exhibit residual watermark "radioactivity" under partially watermarked training datasets. We pit a...

    arxiv.org/abs/2607.00325 · PDF

  66. 66

    Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions

    Zewen Liu

    cs.LG · cs.AI · cs.CL

    The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be simultaneously optimized at fixed sample size N. Prior evidence rests on n=5 conditions with complete metrics from a single study. We expand the empirical base to 11 conditions, measuring gamma and H for all 11 (nine...

    arxiv.org/abs/2607.00304 · PDF

  67. 67

    Generative Modeling of Quantum Distribution with Functional Flow Matching

    Jaehoon Hahm, Tak Hur, Joonseok Lee, Daniel K. Park

    cs.LG · quant-ph

    The emergence of powerful deep generative models based on diffusion and flow matching has enabled the learning and modeling of complex distributions. Learning quantum distributions, however, remains challenging due to the inherent difficulty of accurately modeling the meaningful physical properties of quantum states. We propose Quantum Flow Matching (QFM), a novel generative model designed to learn quantum distribution by utilizing spin...

    arxiv.org/abs/2607.00301 · PDF

  68. 68

    EPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems

    Zewen Liu

    cs.LG · cs.CL

    When LLM agents use evaluator feedback to adapt their behavior in closed loops, evaluator biases propagate through the agent's strategy distribution -- a phenomenon known as evaluator preference coupling. Prior work has documented coupling across multiple evaluator families and model versions, but the field lacks a standardized protocol that enables third-party researchers to (i) reproduce coupling measurements, (ii) compare results across...

    arxiv.org/abs/2607.00297 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.