cs.LG · 2026-06-29 · No. 38

Machine Learning, 2026-06-29.

73 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

73 entries
  1. 01

    VGB for Masked Diffusion Model: Efficient Test-time Scaling for Reward Satisfaction and Sample Editing

    Kijung Jeon, Thuy-Duong Vuong, Molei Tao

    cs.LG · cs.DS · math.NA · math.PR · stat.ML

    Inference-time scaling is a promising paradigm to improve generative models, especially when outputs must satisfy structural constraints or optimize downstream rewards. We consider Masked Diffusion Model (MDM) and introduce MDM-VGB, a discrete diffusion sampler that augments unmasking generation with theoretically principled reward-guided remasking. Inspired by the recent success of the classical Jerrum-Sinclair backtracking Markov chain in...

    arxiv.org/abs/2606.28301 · PDF

  2. 02

    Democratic ICAI: Debating Our Way to Steering Principles from Preferences

    Kevin Kingslin, Anish Natekar, Ashutosh Ranjan, Vivek Srivastava, Savita Bhat, Shirish Karande

    cs.LG · cs.MA

    Preference-based alignment often struggles to capture the reasoning that underlies human judgments. Many evaluations rely on multiple interacting criteria, yet pairwise labels reveal only the final choice rather than the considerations that shape preferences. Inverse Constitutional AI (ICAI) improves interpretability in decision making by summarizing preferences into natural-language principles, but its single-pass explanations miss much of...

    arxiv.org/abs/2606.28294 · PDF

  3. 03

    Towards Automating Scientific Review with Google's Paper Assistant Tool

    Rajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes, Yossi Matias, Vahab Mirrokni, Vincent Cohen-Addad

    cs.LG · cs.AI · cs.CL · cs.CY

    Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem proving. However, this rapid acceleration is creating a systemic challenge: traditional human peer review cannot scale to match the influx of AI-assisted science. Ultimately, to resolve this tension, we must also deploy AI to accelerate the verification and review process itself. To frame the...

    arxiv.org/abs/2606.28277 · PDF

  4. 04

    Parameter Efficient Hybrid Transformer (PEHT) for Network Traffic Prediction via Dynamic Urban Congestion Integration

    Abdolazim Rezaei, Mehdi Sookhak, Mahboobeh Haghparast

    cs.LG · cs.AI

    Accurate network traffic prediction is a critical element for efficient resource allocation in dynamic urban cellular networks. However, prediction remains challenging because network demand is influenced by complex mobility patterns, congestion dynamics, and heterogeneous user behavior. This paper introduces the Parameter-Efficient Hybrid Transformer (PEHT), a network traffic prediction framework that integrates urban mobility and congestion...

    arxiv.org/abs/2606.28274 · PDF

  5. 05

    How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks

    Julius Girardin, Emanuele Troiani, Yizhou Xu, Vittorio Erba, Florent Krzakala, Lenka Zdeborová

    cs.LG · cond-mat.dis-nn · cs.AI · stat.ML

    Understanding how performance scales jointly with model size and data is a central problem in modern machine learning. Existing theoretical works on scaling laws typically describe generalization as a function of data or compute, often in fixed-feature or infinite-width regimes and for online SGD. Here, we instead study how generalization scales with the number of trainable parameters and the number of samples in a feature-learning model. We...

    arxiv.org/abs/2606.28242 · PDF

  6. 06

    Disentangling Continuous-Time Latent Dynamics: Identifiability of Latent SDEs via Diffusion Shifts

    Yuanyuan Wang, Wenjie Wang, Haoxuan Li, Mingming Gong, Kun Zhang

    cs.LG · stat.ML

    Causal representation learning for time series has developed strong identifiability results in discrete-time latent causal models, but identifiability in continuous-time latent stochastic differential equation (SDE) models remains largely open. We address this gap using environment-induced shifts in diffusion covariance. We study additive-noise latent SDEs observed through an unknown nonlinear diffeomorphism, with shared drift but...

    arxiv.org/abs/2606.28228 · PDF

  7. 07

    Estimation--Prediction Tradeoff in Causal Probabilistic Temporal Graphs

    Aniq Ur Rahman

    cs.LG · cs.IT · cs.MA · cs.SI · eess.SY

    Temporal link prediction is usually evaluated by predictive performance on unseen edges, but in probabilistic temporal graphs this criterion can conflate model error with irreducible uncertainty. We study this issue by characterising an inherent estimation--prediction tradeoff in binary logistic models where regimes that maximise Fisher information and improve parameter recoverability are also those with the highest entropy, making individual...

    arxiv.org/abs/2606.28225 · PDF

  8. 08

    Physics-Informed Neural Network with Transfer Learning for State Estimation in Lithium-Ion Batteries using the Single Particle Model with Electrolyte

    Gift Modekwe, Qiugang Lu

    cs.LG

    Physics-informed neural networks (PINNs) have emerged as a powerful tool for solving nonlinear partial differential equations (PDEs), including battery electrochemical models. They typically en-force conservation laws within the loss function to ensure physically consistent solutions. Tradi-tional numerical methods such as finite difference, finite volume, and finite element techniques, re-ly on discretization and can be computationally...

    arxiv.org/abs/2606.28220 · PDF

  9. 09

    Towards Value-Constrained Credit Assignment in Fully Delegated AI Cooperatives

    Young Yoon, Jimin Kim, Soyeon Park

    cs.LG · cs.AI · cs.DC · cs.MA

    We propose a framework for reward allocation in fully delegated AI cooperatives where humans are represented by agents that contribute data and participate in model updates under heterogeneous value constraints. The key idea is to credit only those updates that remain admissible after screening them against each principal's value profile. We formulate value-conditioned gradient filtering, online marginal contribution signals, and cumulative...

    arxiv.org/abs/2606.28217 · PDF

  10. 10

    COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-Negatives

    David Steinmann, Antonia Wüst, Kristian Kersting, Wolfgang Stammer

    cs.LG

    While interpretable models such as concept bottleneck models (CBMs) and program synthesis methods enable verification of model decisions, their evaluation is typically limited to simple tasks, leaving complex reasoning on real-world images largely unexplored. We introduce COCOLogic-V2, an object-centric dataset for visual inductive reasoning on real-world images covering a broad subset of first-order logic. By categorizing samples into...

    arxiv.org/abs/2606.28194 · PDF

  11. 11

    The Remittance Blueprint: Data-driven Intelligence for Sri Lanka

    Dhinanjaya Fernando, Dinura Ginige, Kalana Lakshan, Chanupa Gurusinghe, Lasana Pahanga, Subavarshana Arumugam,...

    cs.LG · cs.AI

    This study analyzes Sri Lankan migration and remittances over 32 years (1994-2025). Using a 384-month harmonized dataset, we apply exploratory data analysis, stationarity corrected time-series modeling (ADF, Johansen, VAR/VECM), and supervised learning. Results reveal remittance inflows are primarily driven by external macroeconomic variables, specifically exchange rate dynamics and global oil prices, rather than domestic indicators. Impulse...

    arxiv.org/abs/2606.28190 · PDF

  12. 12

    LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior

    Qinhong Zhou, Chuang Gan, Anoop Cherian

    cs.LG · cs.AI · cs.CV · cs.RO

    Embodied agents operating in decentralized and partially observable environments have attracted growing attention in recent years. However, existing large language model (LLM)-based agents often exhibit behaviors that are misaligned with their partners or inconsistent with the environment state, leading to inefficient cooperation and poor task success. To address this challenge, we propose a novel framework, Learning Laws of Cooperation...

    arxiv.org/abs/2606.28182 · PDF

  13. 13

    CPAgents: Agentic Composite Phenotype Generation for Cardiac Disease Association

    Zuoou Li, Wenlong Zhao, Kelly Yu, Weitong Zhang, Paul M. Matthews, Wenjia Bai, Bernhard Kainz, Mengyun Qiao

    cs.LG · cs.AI

    Identifying robust associations between cardiac imaging phenotypes and clinical diseases is fundamental to population-scale cardiovascular research and reliable risk stratification. However, current phenome-wide association studies rely on pre-defined, single-variable phenotypes or expert-crafted features, which limits their ability to capture clinically meaningful non-linear effects and cross-phenotype interactions. To address this, we...

    arxiv.org/abs/2606.28179 · PDF

  14. 14

    Recovering Sharp Conductivity Features in the Finite-Data Calderón Problem with Physics-Informed Neural Networks

    Ali AlHadi Kalout, Pablo Tejerina-Pérez, Konstantin Karchev, Pedro Tarancón-Álvarez, Leonid Sarieddine, Raul...

    cs.LG · eess.IV · math.NA

    Physics-informed neural networks (PINNs) have recently emerged as a promising framework for addressing the Calderón inverse problem from limited boundary data. In this work, we revisit neural Calderón inversion by introducing multiscale boundary excitations based on randomized wavelet functions and investigating the role of Fourier-feature encoding (FFE) for representing sharp conductivity variations. We propose a physics-informed...

    arxiv.org/abs/2606.28158 · PDF

  15. 15

    Regularized Reward-Punishment Reinforcement Learning

    Jiexin Wang, Eiji Uchibe

    cs.LG · cs.RO

    We propose KL-Coupled Policy Regularization (KCPR), a policy coordination framework for Reward-Punishment Reinforcement Learning (RPRL). Based on KCPR, we derive KL-Coupled Soft Optimality (KCSO) and develop its deep realization, klDMP. Unlike existing RPRL approaches that optimize reward-seeking and punishment-related policies largely independently, KCPR enables direct interactions between companion policies by treating each as a dynamically...

    arxiv.org/abs/2606.28152 · PDF

  16. 16

    Autoencoder Architectures for Athlete Performance Scoring from Wearable Telemetry

    Mateusz Kubita, Jan Zubalewicz, Krzysztof Siwek

    cs.LG

    Wearable devices produce large, high dimensional training logs for everyday runners, and interpretation rather than data collection is now the limiting step. This paper evaluates five dimensionality reduction models, three autoencoder variants, PCA, and a Variational Autoencoder, on their ability to compress nine sensor runner profiles into a single scalar performance indicator, the latent score. Because the setting is fully unsupervised,...

    arxiv.org/abs/2606.28145 · PDF

  17. 17

    MixTTA: Low-Rank Cross-Channel Mixing for Reliable Test-Time Adaptation

    Mansoo Jung, Youngwook Kim, Jungwoo Lee

    cs.LG

    Test-Time Adaptation (TTA) methods commonly update the affine parameters of normalization layers to adapt deployed models under distribution shifts. However, per-channel affine parameters perform axis-aligned scaling and shifting, making them geometrically incapable of correcting cross-channel structural changes induced by distribution shift. To address this limitation, we propose MixTTA, a lightweight plug-in module that equips normalization...

    arxiv.org/abs/2606.28142 · PDF

  18. 18

    Beyond Sparse Supervision: Diffusion-Guided Learning for Few-Shot Graph Fraud Detection

    Liming Liu, Chao Hu, Mingfei Lu, Yiwei Ge, Xingle Li, Heyuan Shi

    cs.LG · cs.AI

    Graph-based fraud detection is essential for safeguarding large-scale transaction systems, where undetected anomalies may lead to substantial financial losses and security risks. Real-world fraud graphs pose two coupled challenges: sparse and imbalanced supervision, where verified fraudulent labels are scarce and heavily skewed toward benign accounts, and representation dilution, where spatial message passing may oversmooth camouflaged...

    arxiv.org/abs/2606.28134 · PDF

  19. 19

    Dangerous Liaisons of Convex Learning and Non-Affine Aggregation

    Thomas Boudou, Batiste Le Bars, Nirupam Gupta, Aurélien Bellet

    cs.LG · math.OC · stat.ML

    Last-iterate convergence and generalization guarantees in first-order convex learning hinge on the monotonicity of the update operator. While linear averaging preserves the monotonicity of gradient updates, this property is often violated when gradients are aggregated non-affinely, as in modern pipelines enforcing constraints like adaptivity, privacy, robustness or fairness. Whether it is possible to design non-affine aggregation rules that...

    arxiv.org/abs/2606.28123 · PDF

  20. 20

    When One Adapter Speaks for Many: Discovering Low-Rank Redundancy in Continual Fine-Tuning

    Tanguy Dieudonné, Giulia Lanzillotta, Enis Simsar, Louis Barinka, Thomas Hofmann

    cs.LG

    Low-Rank Adaptation (LoRA) has become the standard tool for parameter-efficient fine-tuning of large pretrained models. When applied sequentially across tasks in Continual Learning (CL), the standard assumption is that each new task requires a dedicated low-rank adapter. In this work, we challenge this assumption empirically and structurally. We show that task-specific LoRA adapters in CL exhibit significant low-rank redundancy: the subspaces...

    arxiv.org/abs/2606.28117 · PDF

  21. 21

    Fair Classification with Efficient and Post-hoc Controllable Fairness-Accuracy Trade-off

    Maaya Sakata, Kazuto Fukuchi

    cs.LG

    Post-hoc controllability of fair machine learning models, the ability to control the trade-off between fairness and accuracy after training, is valuable for practical deployment. Existing post-processing methods provide such post-hoc controllability but often suffer from significant accuracy degradation, whereas in-processing methods achieve efficient trade-offs but require computationally expensive retraining for each change in trade-off...

    arxiv.org/abs/2606.28097 · PDF

  22. 22

    OperatorSHAP: Fast and Accurate Shapley Value Estimation for Neural Operators

    Joshua Stiller, Santo M. A. R. Thies, Felix Czaja, Eyke Hüllermeier

    cs.LG · cs.AI

    Understanding model predictions is essential for physical applications, where outputs often inform safety-critical decisions, such as structural load assessment, weather warnings, and clinical diagnosis. Shapley values satisfy many desirable properties as an attribution method, but their computational cost during inference hinders their practical use. Current amortized explainers, such as FastSHAP, are limited to homogeneous inputs, which is...

    arxiv.org/abs/2606.28065 · PDF

  23. 23

    Benchmarking on Tasks That Matter: Dataset Selection for Preserving Model Rankings

    Rostislav Gusev, Alexey Zaytsev

    cs.LG · stat.ML

    Benchmarks of machine learning models often include many datasets, making evaluation expensive. For efficiency, it is preferable to perform evaluations on small, representative datasets instead. The selection of such subsets typically relies on heuristics and is rarely analyzed for the robustness of the resulting model rankings. We introduce a framework to perform the task of selecting datasets subsets with an evaluation of how different...

    arxiv.org/abs/2606.27997 · PDF

  24. 24

    Dual-Learning based Penalized Multi-Align Clustering for Multi-View Incomplete and Disorderly Data

    Liang Zhao, Shubin Ma, Bo Xu, Qingchen Zhang

    cs.LG

    Multimodal feature fusion can effectively capture complex patterns in real-world data by integrating complementary information from different modalities. However, in many applications, such as boiler combustion monitoring, equipment failure, inconsistent sensor sampling frequencies, and network delays often cause missing modalities and temporal asynchrony. These issues lead to incomplete and disorderly multimodal data. To address them,...

    arxiv.org/abs/2606.27984 · PDF

  25. 25

    RECAST: Model Reconstruction via Counterfactual-Aware Wasserstein Geometry under Limited Data

    Xuan Zhao, Lena Krieger, Zhuo Cao, Arya Bangun, Hanno Scharr, Ira Assent

    cs.LG

    Counterfactual explanations (CFs) help understand machine learning models by identifying minimal input changes that would lead to alternative model outcomes. Recent work demonstrates their utility for reconstructing black-box models, enabling third-party auditing of opaque decision systems for fairness and accountability. Still, CF-based reconstruction may suffer from decision boundary shifts, overfitting, and restrictive assumptions...

    arxiv.org/abs/2606.27948 · PDF

  26. 26

    Two-Stage Fine-Tuning for Protein Sequence Generation with Targeted Amino-Acid Composition

    Violeta Basten-Romero, Rubén Muñoz-Tafalla, Anna María Díaz-Rovira, Bertran Miquel-Oliver, Isaac Filella-Merce,...

    cs.LG · cs.AI · q-bio.BM · q-bio.GN

    Protein language models are standard priors for biological sequence generation, but steering them toward explicit distributional design targets remains largely unexplored. We study a constrained protein generation problem in which sequences must match a desired amino-acid (AA) composition profile while preserving plausible sequence statistics and diversity. The motivating application is synthetic feed protein design, where the AA composition...

    arxiv.org/abs/2606.27939 · PDF

  27. 27

    Graph Dimensionality Reduction for Contextual Bandits: Structure-Specific Regret Bounds under Approximate Smoothness and Noisy Eigenspaces

    Joyanta Jyoti Mondal, Ibne Farabi Shihab, Anuj Sharma

    cs.LG

    Contextual bandits with graph-structured arms arise in recommendation, citation retrieval, and social advertising, where arms connected on a graph tend to share reward signal. Standard dimensionality reduction ignores this structure, inflating exploration cost by a factor of $d/k$. We propose GraphDR-LinUCB, which projects arm features onto the graph's low-frequency spectral subspace and runs linear UCB in the resulting $k$-dimensional space....

    arxiv.org/abs/2606.27917 · PDF

  28. 28

    TA-SparseMG: Trend-Aware Sparse Forecasting via Multi-Scale Gating for Long-Term Time Series

    Wenchao Liu, Hongbing Wang, Youji Zhu, Xiaodong Liu, Xiangguang Xiong

    cs.LG

    Long-term time series forecasting finds extensive applications in domains such as power demand, traffic flow, meteorological observation, and renewable energy dispatch. Forecasting dynamically varying long-term time series poses inherent challenges, including statistical nonstationarity, local high-frequency disturbances, and coupled cross-period dependencies, which make it difficult for lightweight models to balance parameter efficiency and...

    arxiv.org/abs/2606.27908 · PDF

  29. 29

    A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset

    Ahmed Mohamady, Robin Burchard, Kristof Van Laerhoven

    cs.LG

    Recent advances in Human Activity Recognition (HAR) from wearable sensors have shown that multi-modal deep learning models consistently outperform their uni-modal counterparts. Modalities can include IMUs, RGB cameras, audio signals, and others. One important aspect of multi-modal deep learning is the sensor fusion approach we apply. Over recent years, multiple fusion paradigms have been proposed for multi-modal HAR. However, to the best of...

    arxiv.org/abs/2606.27886 · PDF

  30. 30

    FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models

    Fan Mo, Yuxuan Han, Geng Zhang, Wangbo Zhao, Yang You

    cs.LG

    Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models. However, sparse activation does not remove the deployment burden of storing and serving all experts, and the available deployment budget can vary substantially across devices, users, and workloads. Existing MoE compression methods are still largely fixed-budget, typically optimizing...

    arxiv.org/abs/2606.27866 · PDF

  31. 31

    GNBAN: Graph Neural Basis Attention Networks for Long-Horizon Forecasting over Large Entity Sets

    Janak M. Patel, Anirudh Deodhar, Dagnachew Birru

    cs.LG · cs.AI

    Demand forecasting at the bottom of a retail hierarchy requires predicting tens of thousands of correlated long-horizon series across products, stores, and regions. Modern systems must scale across massive catalogs, capture shared demand dynamics, and remain interpretable enough to be trusted. Classical statistical methods need a separate model per series and are hard to manage at scale; deep autoregressive models struggle as the joint state...

    arxiv.org/abs/2606.27863 · PDF

  32. 32

    Applicability of memorization indicators for early spotting of overfitting while recalibrating sEMG-decoders on low sample sizes

    Stephan J. Lehmler, Tobias Glasmachers, Ioannis Iossifidis

    cs.LG · cs.AI

    Deep learning models for surface electromyography (sEMG) can benefit substantially from subject-specific (re-)calibration, since no sufficiently large and diverse datasets are available to train fully generic decoders. However, for user acceptance, the number of repetitions that can realistically be collected during calibration is severely limited, which increases the risk of overfitting and, in extreme cases, can even degrade performance...

    arxiv.org/abs/2606.27855 · PDF

  33. 33

    WattLayer: Get Layers Right to Estimate Inference Energy of Neural Networks

    Adrien Sardi, Marie-Line Alberi Morel, Sara Alouf, Frédéric Giroire, Joanna Moulierac

    cs.LG · cs.AI

    The widespread adoption of Artificial Intelligence (AI) has led to increasing concerns about energy consumption, yet there is a lack of standardized methodologies to accurately estimate AI inference energy consumption, particularly across various tasks and architectures. In this study, we propose a task independent, layer-wise energy estimation model for AI architectures. Our model is evaluated on a large dataset of more than 100,000 layers...

    arxiv.org/abs/2606.27841 · PDF

  34. 34

    USAD: Uncertainty-aware Statistical Adversarial Detection

    Zhijian Zhou, Xunye Tian, Jiacheng Zhang, Zesheng Ye, Yiyi Guo, Donghao Zhang, Liuhua Peng, Feng Liu

    cs.LG

    Statistical adversarial detection (SAD) treats detection as a two-sample test. Given a reference set of clean examples (CEs) and a batch of queries, potentially containing an unknown mixture of CEs and adversarial examples (AEs), SAD decides whether the query distribution drifts away from the CE distribution while controlling the false-alarm rate. Existing SAD-based methods mainly use maximum mean discrepancy (MMD) to measure the...

    arxiv.org/abs/2606.27832 · PDF

  35. 35

    Pepti-drift: Toxicity-Repulsive Drifting for Antigen-Conditioned Discrete Peptide Generation

    Takashi Fujiwara, Hikaru Shindo, Kaushalya Madhawa, Jun Jin Choong, Keisuke Ozawa

    cs.LG · cs.AI

    Peptides are a promising therapeutic modality that combine the chemical tunability of small molecules with the target specificity of macromolecular therapeutics. However, designing antigen-specific binding peptides while avoiding toxicity remains a major challenge for therapeutic peptide discovery. Here, we present Pepti-drift, a toxicity-aware latent refinement framework that generates peptide candidates through a single antigen-conditioned...

    arxiv.org/abs/2606.27824 · PDF

  36. 36

    Accelerating Hierarchical Sparse Predictive Coding with Hybrid Amortized Inference

    Kazuhisa Fujita

    cs.LG

    Hierarchical predictive coding provides an interpretable framework for perception as error-driven inference in multi-layer generative models, while sparse coding imposes parsimonious latent representations through explicit sparsity constraints. Their combination yields hierarchical sparse predictive coding models with appealing computational and neuroscientific properties, but practical use is often limited by the cost of iterative latent...

    arxiv.org/abs/2606.27802 · PDF

  37. 37

    NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

    Tianlin Pan, Lianyu Pang, Cheng Da, Huan Yang, Changqian Yu, Kun Gai, Wenhan Luo

    cs.LG · cs.CV

    Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that are not captured by the reward proxy. We identify a simple structural signature of this drift: across three post-training methods (NFT, AWM, DPO), RL fine-tuning inflates the per-step velocity norm $\|v_θ\|$ by $5\%$ to $15\%$ relative to the reference. A form of norm inflation has been studied...

    arxiv.org/abs/2606.27771 · PDF

  38. 38

    Difference of Convex Programming in the Wasserstein Space with Applications to MMD Optimization

    Clément Bonet, Pierre-Cyril Aubin-Frankowski, Youssef Mroueh

    cs.LG · math.OC

    Optimizing functionals over the space of probability measures is now ubiquitous in machine learning. A widely used approach is to perform the optimization directly over the Wasserstein space, but many objective functionals of practical interest are non-convex along Wasserstein geodesics, making the analysis of standard first-order methods challenging. In this work, we study a class of objectives over the Wasserstein space that admit a...

    arxiv.org/abs/2606.27767 · PDF

  39. 39

    RS-Diffuser: Risk-Sensitive Diffusion Planning with Distributional Value Guidance

    Shiqiang Gong

    cs.LG · cs.AI · cs.RO

    Offline reinforcement learning enables policy learning from fixed datasets without additional environment interaction, making it appealing for safety-critical applications where online exploration is costly or unsafe. Diffusion-based decision-making methods have recently achieved strong performance in offline RL by modeling rich, multimodal trajectory distributions. However, existing diffusion planners are typically risk-neutral and therefore...

    arxiv.org/abs/2606.27766 · PDF

  40. 40

    Layerwise Progressive Freezing: A Training Scaffold for Depth-Scalable Binary Networks

    Evan Gibson Smith, Bashima Islam

    cs.LG

    Training binary neural networks (BNNs) from scratch is dominated by the straight-through estimator (STE), whose forward/backward mismatch produces severe accuracy degradation as networks deepen. We study an orthogonal axis: when and where binarization is enforced during training. We introduce StoMPP (Stochastic Masked Partial Progressive Binarization), which gradually replaces clipped weights and activations with their hard binary...

    arxiv.org/abs/2606.27759 · PDF

  41. 41

    PerturbCellRL: Verifier-Guided Reinforcement Learning for Single-Cell Perturbation Prediction

    Dongxia Wu, Mingyu Li, Yuhui Zhang, Anurendra Kumar, Emma Lundberg, Serena Yeung-Levy, Emily B. Fox

    cs.LG

    Single-cell perturbation models can reduce costly wet-lab screening by predicting how cells respond transcriptionally to interventions. While recent generative models improve population-level prediction, individual generated cells are not explicitly checked for biological consistency. We introduce PerturbCellRL, a reinforcement learning (RL) framework that post-trains a pretrained single-cell transcriptomic generator using a suite of...

    arxiv.org/abs/2606.27752 · PDF

  42. 42

    Flexformer: Flexible Linear Transformer with Learnable Attention Kernel

    Haoran Zhang, Feng Zhou

    cs.LG · cs.AI

    Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences. Kernel-based linear attention reduces this complexity but typically relies on fixed or weakly learnable kernels, restricting expressiveness and performance. In this work, we propose Flexformer, a flexible linear Transformer that learns attention kernels in a fully data-driven...

    arxiv.org/abs/2606.27748 · PDF

  43. 43

    The Weakest Link Tells It All: Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment

    Tianyu Jia, Yue Fang, Hongxin Ding, Rihong Qiu, Zhibang Yang, Zhijing Wu, Xu Chu, Junfeng Zhao, Yasha Wang

    cs.LG

    Process reward models (PRMs) enhance the reasoning capabilities of large language models (LLMs) by providing fine-grained feedback, yet training PRMs typically requires expensive stepwise annotations. Outcome-supervised PRMs offer a scalable alternative by learning from final-answer correctness alone, but this introduces a fundamental *credit assignment* challenge, i.e., attributing outcomes to responsible reasoning steps. Existing approaches...

    arxiv.org/abs/2606.27739 · PDF

  44. 44

    Reduction of Probabilistic Chemical Reaction Networks

    Mauricio Montes, Gregoire Sergeant-Perthuis

    cs.LG · math.CT

    Programming adaptive behaviors at the cellular level is a long-standing goal that raises the question of how probabilistic computation can be implemented in biochemical systems. Chemical reaction networks (CRNs) provide such a substrate and have been shown to realize probabilistic models, including hidden Markov models and factor graphs, with dynamics reproducing Bayesian inference and belief propagation. However, encoding these algorithms...

    arxiv.org/abs/2606.27737 · PDF

  45. 45

    Learning to Reason with Curriculum II: Compositional Generalization

    Nived Rajaraman, Audrey Huang, Miroslav Dudik, Robert Schapire, Dylan Foster, Akshay Krishnamurthy

    cs.LG

    Compositional generalization, the ability to solve complex problems by combining solutions to simpler sub-problems, is a fundamental capability of both natural and artificial intelligence, and a key mechanism underlying chain-of-thought reasoning. However, the theoretical underpinnings of compositional generalization remain poorly understood: when and why does decomposing a problem into parts yield more efficient learning than solving it...

    arxiv.org/abs/2606.27721 · PDF

  46. 46

    Aurora: A Leverage-Aware Spectral Optimizer

    Alec Dewulf, Dhruv Pai, Li Yang, Ashley Zhang, Ben Keigwin

    cs.LG

    We show that for tall matrix parameters, like projection matrices in the MLP layers, the Muon update can have row norms that are arbitrarily non-uniform. This can lead to a self-reinforcing feedback loop whereby neurons receive persistently small updates and eventually do not contribute meaningfully to network outputs. This problem is effectively mitigated by an additional row normalization step, but current methods do this in a way that...

    arxiv.org/abs/2606.27715 · PDF

  47. 47

    The Simulacrum: Decision-Theoretic Pretraining for Near-Optimal Time-Series Forecasting and Inference

    Pablo Montero-Manso, Marcel Scharth

    cs.LG · cs.AI · stat.CO

    We introduce a neural network-based framework for learning time series estimators through a process we term decision-theoretic pretraining. Analysts specify a generative world, a distribution over data-generating processes, and a target decision objective. A neural network trained on stratified simulations from this world approximates the corresponding optimal decision rule, yielding a neural estimator that provides forecasts, parameter...

    arxiv.org/abs/2606.27711 · PDF

  48. 48

    What Was That Again? Certified Robustness for Automatic Speech Recognition

    Andrew C. Cullen, Neil Marchant, Jiani Xie, Paul Montague, Benjamin I. P. Rubinstein

    cs.LG · cs.AI · cs.CR · cs.SD

    Automatic Speech Recognition systems are notoriously both sensitive to adversarial and benign perturbations. While this has been repeatedly demonstrated using reference datasets, detecting such behaviors in deployed systems is incredibly challenging, due to the absence of oracle knowledge of the true transcription. We demonstrate that employing a certification-inspired mechanism can significantly decrease WER, increase recall, and decrease...

    arxiv.org/abs/2606.27698 · PDF

  49. 49

    Class-frequency Guided Noise Schedule for Diffusion Models

    Jiequan Cui, Beier Zhu, Qingshan Xu, Xiaojuan Qi, Bei Yu, Hanwang Zhang

    cs.LG · cs.AI · cs.CV

    In this paper, we are the first to examine the correlations between class frequency and the multi-scale noise schedule within diffusion models. For score-based generative models, low-density regions often lead to inaccurately estimated scores, thereby compromising the generation quality. Although the multi-scale noise schedule can alleviate this issue during the diffusion process, low-frequency classes still face the challenge of large...

    arxiv.org/abs/2606.27696 · PDF

  50. 50

    Halt Fast! Early Stopping for Certified Robustness

    Andrew C. Cullen, Paul Montague, Benjamin I. P. Rubinstein

    cs.LG · cs.AI · cs.CR

    Randomized Smoothing (RS) provides rigorous robustness guarantees for neural networks without architectural constraints, yet its adoption is limited by extreme computational costs. Standard RS requires tens of thousands of model evaluations per input and forces practitioners to commit to fixed sample sizes a priori. In this work, we present a novel meta-learning framework for anytime-valid certified robustness that adaptively deploys...

    arxiv.org/abs/2606.27694 · PDF

  51. 51

    Deployment-Side Adaptiveness in Multi-Horizon Volatility Forecasting

    Riku Green, Zahraa S. Abdallah, Telmo M Silva Filho

    cs.LG · cs.AI

    In financial forecasting, predictive performance depends not only on which model is trained, but also on how the trained model is deployed. We study this issue in multi-horizon volatility forecasting. Our starting point is that a trained multi-output (MIMO) forecaster does not define a single deployable predictor: by changing the inference-time rollout rule, the same trained model induces a family of forecasts with different accuracy and cost...

    arxiv.org/abs/2606.27688 · PDF

  52. 52

    CBD: API-Only LLM Black-Box Unlearning through Controlled Behavioral Divergence

    Zhiqiang Xie, Yijing Lin, Zhipeng Gao, Dong In Kim

    cs.LG · cs.AI

    Edge devices increasingly invoke large language models (LLMs) through API services for context aware edge intelligence, while edge generated data may be collected to improve LLMs and may introduce sensitive, copyrighted, harmful, or outdated information into model behavior. Machine unlearning offers a practical way to remove the influence of undesired data without retraining LLMs. However, existing methods still face two gaps. The first is...

    arxiv.org/abs/2606.27683 · PDF

  53. 53

    Textual Belief States for World Models: Identifiable Representation Learning Under Strict Mediation

    Xiang Gao, Kaiwen Dong, Yuguang Yao, Padmaja Jonnalagedda, Kamalika Das

    cs.LG · cs.CL

    World models in partially observed environments rely on latent representations that summarize interaction history, but in many modern LLM-based architectures predictive performance fails to reflect representation quality due to history bypass, rendering the latent state unidentifiable. Strict latent state mediation, requiring predictions to depend only on the latent state and action, is a classical principle that resolves this, but enforcing...

    arxiv.org/abs/2606.27681 · PDF

  54. 54

    Are Time-Series Foundation Models Ready for E-Nose Data? An Empirical Assessment of Their Embeddings

    Taeyeong Choi, Mohammed Kamruzzaman

    cs.LG

    Inspired by advances in natural language processing and computer vision, "time-series foundation models" (TSFMs) have recently been introduced with the promise of strong generalization across diverse time-series tasks, including forecasting, classification, and anomaly detection, as well as across domains such as healthcare, climate science, and manufacturing. However, their utility for gas-sensing data remains largely unexplored. To address...

    arxiv.org/abs/2606.27672 · PDF

  55. 55

    TeRoR: Decoupled Temporal Rotation with Relational Circular Region for Temporal Knowledge Graph Embedding

    Peijia Xie, Yike Liu, Chao He, Huiling Zhu

    cs.LG

    In recent years, with the emergence of Temporal Knowledge Graphs (TKGs), research on learning entity and relation representations in TKGs has attracted increasing attention, giving rise to a large number of TKG embedding methods. TeRo is a simple and efficient temporal knowledge graph embedding approach. However, TeRo does not do well in modeling the mapping properties of various relations, such as one-to-many, many-to-one, and many-to-many....

    arxiv.org/abs/2606.27651 · PDF

  56. 56

    Continual Learning for Sequential Personalization of Small Language Models: A Stability Monitoring Analysis

    Thomas S. Paula, Lucas S. Kupssinskü, Rodrigo C. Barros

    cs.LG

    Small Language Models (SLMs) are increasingly being considered for deployment on edge devices such as laptops, enabling private, low-latency, and locally personalized applications. However, personalization requires models to adapt over time to evolving user- or task-specific data, placing them in a continual learning setting. This creates the risk of catastrophic forgetting, where learning new information degrades performance on previously...

    arxiv.org/abs/2606.27634 · PDF

  57. 57

    HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models

    Artem Ploujnikov, Francesco Verdini, Samir Sadok, Mirco Ravanelli

    cs.LG · cs.AI

    Discrete audio representations have become increasingly popular for building multimodal text-audio systems and integrating audio capabilities into Large Language Models (LLMs). However, numerous studies report performance degradation on various downstream tasks due to information loss during discretization. To address this, we propose a novel approach combining temporally compressed discrete tokens with dimensionality-reduced continuous...

    arxiv.org/abs/2606.27627 · PDF

  58. 58

    FoggyTrust: Robust Federated Learning with Hierarchical Trust Networks

    Emmanuel Rassou, Tomas Gonzalez

    cs.LG · cs.DC

    Byzantine-robust federated learning seeks to protect distributed model training from malicious or corrupted clients without requiring access to their private data. FLTrust addresses this challenge by introducing a trusted server-side root dataset that assigns trust scores to client updates for more robust aggregation. In this work, we propose FOGGYTRUST, a hierarchical extension of FLTrust that localizes trust computation to fog nodes,...

    arxiv.org/abs/2606.27622 · PDF

  59. 59

    COOPA: A Modular LLM Agent Architecture for Operations Research Problems

    Chuanhao Li, Xiaoan Xu, Dirk Bergemann, Ethan X. Fang, Yehua Wei, Zhuoran Yang

    cs.LG

    Operations Research (OR) provides a rigorous framework for high-stakes decision-making, but effective OR modeling requires substantial domain knowledge, mathematical abstraction, and solver expertise. Recent LLM-based systems automate parts of this pipeline, yet remain limited by low accuracy on complex problems, opaque outputs, and narrow solver support. We propose COOPA (COoperative OPerations Agent), a modular LLM-agent architecture for...

    arxiv.org/abs/2606.27611 · PDF

  60. 60

    Training Observable Control Policies to Expose Agent State Through Actions

    Andres Enriquez Fernandez, John J. Bird

    cs.LG · eess.SY

    Physical or operational constraints often impose communications limitations on autonomous agents. Such limitations complicate monitoring or multiagent coordination. Even when strong communications are absent, some information may still be available. The remainder of the relevant agent state may be reconstructed via estimation. The actions taken by an agent are a potential source of information -- as the agent interacts with the environment,...

    arxiv.org/abs/2606.27609 · PDF

  61. 61

    Global Explanations for Multivariate Time Series Forecasting Models via $K$-Order Markov Approximations

    Amadeo Tunyi

    cs.LG · cs.AI

    While many explainable AI (XAI) methods have been proposed, most are not designed for time-series forecasting models and often rely on the implicit assumption that timestamp features are independent. This assumption ignores the fundamental property of temporal dependence and can lead to explanations that violate the sequential and causal structure of the data. We introduce \textsc{KARMA}, a method for explaining time-series predictors by...

    arxiv.org/abs/2606.27599 · PDF

  62. 62

    Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF

    Arnav Raj

    cs.LG · cs.AI

    Reinforcement learning from human feedback (RLHF) in production does not always have a synchronous reward signal. Code-execution verifiers, slow judge ensembles, and queued human review can return several gradient steps after the rollout that produced them, breaking the synchronous-reward assumption underlying standard PPO. We address this gap with Retroactive Advantage Correction (RAC): each pending slow completion is queued, aged through a...

    arxiv.org/abs/2606.27580 · PDF

  63. 63

    PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration

    Arnav Raj

    cs.LG · cs.AI

    Reward models for Reinforcement Learning from Human Feedback (RLHF) pool preferences across thousands of annotators and fit one global affine calibrator, collapsing raters with systematically different rating-scale offsets and slopes into a single average-rater fit that does not match any individual annotator. PEBS is a per-rater empirical-Bayes shrinkage estimator: it fits per-rater affine calibrators on a held-out slice of each annotator's...

    arxiv.org/abs/2606.27578 · PDF

  64. 64

    hia-gat: A Heterogeneous Interaction-Aware Graph Attention Network For Frame-Level Traffic Conflict Risk Prediction On Freeways

    Mahshid Malazizi, Seyedmehdi Khaleghian, Mina Sartipi, Toru Hirano, Yunfei Xu, Hoang H. Nguyen

    cs.LG · cs.AI

    This paper formulates frame-level freeway risk assessment as a multi-agent scene graph-level binary classification problem, where each video or trajectory frame is labeled risky if any TTC- or PET-based conflict violates a specified severity threshold. We construct a relation-aware graph per frame with vehicles as nodes and two interaction types as edges: same-lane (longitudinal) and adjacent-lane (lateral), augmented with physics-informed...

    arxiv.org/abs/2606.27577 · PDF

  65. 65

    Quantum Generative Diffusion Model for Real-World Time Series

    Jack Waller, Filippo Caruso, Dimitrios Makris, Rajagopal Nilavalan, Xing Liang

    cs.LG · quant-ph

    Generative models have achieved remarkable success in data synthesis, though recent advances driven by increasing model scale have introduced challenges in computational cost and efficiency. Quantum machine learning offers a promising alternative, representing complex data distributions using compact, highly expressive models. Here, we propose QDiffusion-TS, the first quantum generative diffusion model for time series synthesis, and validate...

    arxiv.org/abs/2606.27561 · PDF

  66. 66

    Productionized Fairness Measurement Under Privacy Constraints

    Osonde A. Osoba, Yuzi He, Saikrishna Badrinarayanan, Varun Mithal, Sakshi Jain, Natesh S. Pillai

    cs.LG · cs.CR

    Fairness measurements in the form of disaggregated evaluations often rely on demographic signals that are legally constrained or culturally sensitive. Race and ethnicity signals are among the more difficult signals to curate and use for this task. This paper presents Privacy-Preserving Probabilistic Race/Ethnicity Estimation (PPRE) as a method for enabling fairness measurements with respect to race/ethnicity for U.S.\ LinkedIn members in a...

    arxiv.org/abs/2606.27558 · PDF

  67. 67

    Boundary condition fidelity for bottom-hole pressure and CO2 plume prediction in geological carbon storage

    Romal Ramadhan, Seyyed A. Hosseini, Larry W. Lake

    cs.LG · physics.geo-ph

    Accurate prediction of bottom-hole pressure (BHP) and CO2 plume migration is essential for safe geological carbon storage, yet practical simulations often rely on truncated domains where artificial boundaries distort pressure diffusion and CO2 saturation footprints. In this study, we evaluate how boundary-condition fidelity affects BHP and CO2 plume prediction by comparing ten reduced-domain boundary treatments against full-domain reference...

    arxiv.org/abs/2606.27515 · PDF

  68. 68

    The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

    Sankaran Vaidyanathan, David Arbour, Aaron Mueller, Scott Niekum, David Jensen

    cs.LG · cs.CL

    Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating its natural indirect effect (NIE). Re-deriving the activation patching estimand from causal mediation analysis, we find that the NIE does not solely capture the causal effect through the specific component. It also contains interaction effects (INT) that measure...

    arxiv.org/abs/2606.27510 · PDF

  69. 69

    Operator Learning for Cubic Nonlinear Schrödinger Equation on Periodic Domains

    Emmanuel E. Oguadimma, Victory C. Obieke, Xueying Yu

    cs.LG · math.AP · math.NA

    We consider the cubic nonlinear Schrödinger (NLS) equation on two-dimensional flat tori with varying aspect ratios. In this formulation, the choice of aspect ratio governs the Fourier resonance structure, so rational and irrational geometries can exhibit different high-frequency cascade behaviors. We present a geometry-conditioned Fourier neural operator (FNO) for the cubic defocusing NLS equation, where the input consists of the real and...

    arxiv.org/abs/2606.27459 · PDF

  70. 70

    Prism Transformer: Progressive Head Schedules for Hierarchical Attention Processing

    Shubham Aggarwal

    cs.LG

    Multi-head attention conventionally partitions the hidden dimension equally across all heads at every layer, enforcing an identical representational subspace dimension (dh = dmodel/h) throughout the models depth. In this work, we identify this uniform allocation as a fundamental structural bottleneck: due to their restricted dimensional space, early-layer heads are unable to faithfully capture complex, high-dimensional contextual patterns. To...

    arxiv.org/abs/2606.27449 · PDF

  71. 71

    Learning in Markovian bandits with non-observable states and constrained decision epochs

    Thomas Hira, Victor Boone, Urtzi Ayesta, Ina Maria Verloop

    cs.LG

    This paper studies the problem of regret minimization in Markovian bandits with \emph{non-observable states} and possibly \emph{constrained} decision epochs. The focus is restricted to a ``pure'' regret benchmark, that compares the performance of the learning algorithm to the best \emph{pure policy} which -- akin to optimal policies of stochastic bandits -- picks the optimal arm from start to finish without ever switching. We introduce a...

    arxiv.org/abs/2606.27448 · PDF

  72. 72

    PairSAE: Mechanistic Interpretability from Pair Representations in Protein Co-Folding

    Giosue Migliorini, Aristofanis Rontogiannis, Grigori Guitchounts, Nicholas Franklin, Axel Elaldi, Olivia Viessmann

    cs.LG

    Foundation models for structural biology have achieved remarkable performance in predicting biomolecular structure and show promise for the design of proteins and small molecules. Yet understanding which internal features drive their outputs remains challenging. Standard sparse autoencoders (SAEs), effective on transformer-style sequence embeddings, do not transfer cleanly to pairformer-like architectures: naively operating on pairwise...

    arxiv.org/abs/2606.27440 · PDF

  73. 73

    Unified Zero-Shot Time Series Forecasting: A Darts Foundation

    Zhihao Dai, Dennis Bader, Alain Gysi

    cs.LG

    Since its initial release in 2020, Darts has become a widely used open-source Python library for time series analysis. A series of foundation models have recently claimed accuracy improvements in zero-shot forecasting, promising a paradigm shift from training custom models to harnessing pre-trained general-purpose forecasters. Foundation models, however, are often released as isolated packages with fragmented interfaces and limited...

    arxiv.org/abs/2606.27438 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.