cs.LG · 2026-07-13 · No. 52

Machine Learning, 2026-07-13.

69 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

69 entries
  1. 01

    Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection

    Cláudio Lúcio do Val Lopes, Lucca Machado da Silva

    cs.LG · cs.AI

    Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collapse'', defaulting to the majority class and failing to balance anomaly interdiction with customer friction. To overcome this without distortive data resampling, we propose the Semantic Pareto-DQN, a multi-objective reinforcement learning framework. Our approach synthesizes heterogeneous transaction features...

    arxiv.org/abs/2607.09641 · PDF

  2. 02

    Graph-Regularized Low-Rank Matrix Completion by Variable Projection

    Benoît Loucheur, P. -A. Absil, Michel Journée

    cs.LG · math.NA · math.OC

    We address the low-rank matrix completion problem by incorporating graph regularization into the existing Riemannian Trust-Region Matrix Completion (RTRMC) framework. The latter uses the geometry of the low-rank constraint to remodel the problem as an unconstrained optimization problem on a single Grassmann manifold. Our approach, named Graph-Regularized RTRMC (GR-RTRMC), exploits the inherent relationships between rows and columns of the...

    arxiv.org/abs/2607.09546 · PDF

  3. 03

    CoCoT-EEG: Contrastive-Pretrained Multiscale Convolutional Transformer for EEG Decoding

    Gabriel Mahuas, Victoria Shevchenko, Ugo Tanielian, Yassir Bendou, Richard Gao

    cs.LG · q-bio.NC

    Self-supervised pretrained foundation models (FM) have shown early promise for non-invasive electroencephalogram (EEG) decoding applications. Many recent large-scale models converged on the approach of tokenizing raw EEG followed by masked reconstruction pretraining. However, this recipe has been shown to be suboptimal for data, like EEG, with high noise amplitude and information confined to limited dimensions such as narrow frequency bands....

    arxiv.org/abs/2607.09543 · PDF

  4. 04

    GatedLinear: Adaptive Routing of Complementary Linear Bases for Time Series Forecasting

    Qitai Tan, Ruiwen Gu, Yilin Su, Mo Li, Xu Lin, Xiao-Ping Zhang

    cs.LG

    Time series forecasting requires models to capture diverse, often mutually exclusive, temporal dynamics, from smooth trend continuation to nonstationary drift and strict phase-aligned recurrence. While recent deep learning models have improved accuracy, they typically force these diverse patterns through a single computational backbone governed by fixed algorithmic inductive biases (e.g., self-attention or spectral filtering). This...

    arxiv.org/abs/2607.09537 · PDF

  5. 05

    Statistically Undetectable Backdoors in Deep Neural Networks

    Andrej Bogdanov, Alon Rosen, Neekon Vafa

    cs.LG · cs.CR · stat.ML

    We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input,...

    arxiv.org/abs/2607.09532 · PDF

  6. 06

    TSAI-MetaFraud: A Benchmark Dataset for Financial Fraud Transaction and Behavioral Risk Detection in Metaverse Ecosystems

    Refat Ishrak Hemel, Ehsan Hallaji, Roozbeh Razavi-Far

    cs.LG · cs.CR · cs.DB · cs.SI

    The emergence of metaverse platforms has created virtual economies that introduce new challenges related to fraud, bot activity, and illicit financial behavior. Despite growing interest in trustworthy metaverse analytics, existing datasets typically focus on user behavior, authentication, or financial transactions in isolation, limiting the development and reproducible evaluation of multimodal fraud detection methods. To address this gap, we...

    arxiv.org/abs/2607.09528 · PDF

  7. 07

    All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models

    Pan Li

    cs.LG · cs.AI · cs.IR

    Explaining machine-learning models is increasingly important for decision-making and consumer trust, yet it is widely believed to come at a cost: existing Explainable AI (XAI) methods suffer from a persistent accuracy-explainability trade-off. We argue that this trade-off is not fundamental, but an artifact of treating explanation and prediction as separate objectives; when properly coupled, they become complementary, so that equipping a...

    arxiv.org/abs/2607.09502 · PDF

  8. 08

    Neural Collapse Is Forbidden: Information Floors in Language Models

    Bruno Abrahao

    cs.LG · cs.CL · stat.ML

    Within-class variance in language-model representations is commonly read as incomplete neural collapse. We argue it is allocated information storage, and that the allocation obeys a law. A one-line centering identity voids a family of simplex equiangular-tight-frame claims, including our own earlier ones; in dimensionless variance shares across 14 models, macro-category structure carries only 4-12% of representational variance and...

    arxiv.org/abs/2607.09487 · PDF

  9. 09

    Active rejection enables reliable generalization of universal machine-learning interatomic potentials

    Mingxiang Luo, Xinnan Mao, Lu Wang, Lei Bai, Feng Ding, Yuqiang Li

    cs.LG

    Universal machine learning interatomic potentials (uMLIPs) bridge quantum-mechanical accuracy and large-scale molecular dynamics, but the cost of high-accuracy calculations such as r$^2$SCAN limits training to datasets that remain small relative to the open materials space. Strong average benchmark performance also does not guarantee reliable energy--force predictions for every structure. We propose Adaptive Multi-Teacher Routing (ATR), which...

    arxiv.org/abs/2607.09456 · PDF

  10. 10

    Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning

    Edwin De Nicolo, Rahul Marchand, Cornelius Carlsson, Pranav Vaidhyanathan, Natalia Ares

    cs.LG · cond-mat.mes-hall

    Cooperative multi-agent reinforcement learning is well suited to problems with large parameter spaces and exploitable local structure, such as the tuning of electrostatically-defined quantum-dot arrays. However, if parameter cross-talk is strong, a non-stationary environment from the perspective of any individual agent can destabilize learning - the same effect that plagues manual tuning of such systems. We propose using a factored...

    arxiv.org/abs/2607.09422 · PDF

  11. 11

    Similarity search generalisation in contrastive learning with InfoNCE loss

    Nick Whiteley

    cs.LG · stat.ML

    Similarity search is a primary application of embedding models trained by contrastive learning. For one of the most popular contrastive learning loss functions, InfoNCE, we show that the population risk with $k$ negative samples is $O(1/k)$ close to an expected cross-entropy which quantifies deviation between i) a softmax similarity search over unseen data using the learned embedding function, and ii) an idealised softmax search over the same...

    arxiv.org/abs/2607.09405 · PDF

  12. 12

    SYNRARE: Synthetic Rare Disease EHR Generation for ML Benchmarking

    Nicolai Dinh Khang Truong, Richard Röttger

    cs.LG

    Motivation: Rare disease (RD) diagnosis is frequently delayed due to the similarities in symptoms to common disease variants. Machine Learning Algorithms applied to Electronic Health Records show promise for accelerating the diagnosis; however, legal and privacy concerns pose significant barriers. To address these issues, Synthetic Data Generation is an alternative method for obtaining Electronic Health Records and can be applied with any...

    arxiv.org/abs/2607.09404 · PDF

  13. 13

    Data-Efficient Deep Learning: Empirical Guidelines for Training Set Size Estimation in Inertial Sensor Classification

    Ofir Kruzel, Itzik Klien

    cs.LG

    Deep learning models dependency on large-scale inertial datasets presents a significant bottleneck in inertial sensor-based classification tasks, such as human activity recognition and smartphone location recognition. In these domains, data collection requires massive recording campaigns that are complex, time-consuming, and difficult to scale. Currently, data-driven guidelines for determining the minimum sample size required to reach a...

    arxiv.org/abs/2607.09402 · PDF

  14. 14

    On-Device Adaptive Battery Power Prediction for Electric Vehicles

    Avik Bhatnagar, Anton Paule, Tobias Schuermann, Sebastian Reiter, Oliver Bringmann

    cs.LG · cs.AI · cs.AR · cs.PF

    Adaptive power management in Electric Vehicles (EVs) requires accurate power prediction. Although deep learning models have emerged as highly effective for time-series forecasting in this domain, their performance is prone to degradation when exposed to data with distributions different from the training data. We introduce a novel approach that enables on-device learning in resource-constrained EV systems to continuously adapt pretrained...

    arxiv.org/abs/2607.09400 · PDF

  15. 15

    Fully Trainable Deep Differentiable Logic Gate Networks and Lookup Table Networks

    Wout Mommen, Lars Keuninckx, Matthias Hartmann, Werner Van Leekwijck, Piet Wambacq

    cs.LG · cs.AI

    We introduce a novel method for both partial and full optimization of the connections in deep differentiable logic gate networks (LGNs) and lookup table networks (LUTNs). Our training method utilizes a probability distribution over a set of connections per gate/lookup table (LUT) input pin, selecting the connection with highest merit, all whilst the optimal gate types or LUT-entries are learned in parallel. We show that the...

    arxiv.org/abs/2607.09399 · PDF

  16. 16

    Learning Physics-Informed Surrogate Model of Linear Elastic Displacement Fields from Geometry

    Rodolphe Barlogis, Ferhat Tamssaouet, Quentin Falcoz, Stéphane Grieu

    cs.LG

    This work aims to develop a fast and physically consistent surrogate model for real-time structural health monitoring of fractured elastic domains. We propose a physics-informed DeepONet framework that predicts displacement fields from both boundary conditions and fracture geometry, using a dedicated encoding strategy for the latter and without relying on finite-element-generated training data. The traction-free condition on the fracture...

    arxiv.org/abs/2607.09382 · PDF

  17. 17

    Mach-Mind-4-Flash Technical Report

    Foundation Model Team

    cs.LG · cs.CL

    We present Mach-Mind-4-Flash, a 35B-parameter Mixture-of-Experts (MoE) agentic model with 3B activated parameters. Through post-training optimization alone without scaling pre-training compute, the model achieves performance on par with or surpassing that of 100B-parameter-class models. By introducing scalable agentic interaction environments for large-scale reinforcement learning, the model attains significant performance gains on real-world...

    arxiv.org/abs/2607.09375 · PDF

  18. 18

    Graph Neural Networks for Scalable and Transferable Node Centrality Approximation

    Samra Sana, Giorgio Mantica, Saul Imbrici

    cs.LG

    Graph Neural Networks (GNNs) provide a learning-based framework for approximating graph quantities that are expensive to compute exactly. This paper investigates GNNs for scalable approximation of betweenness and closeness centrality, formulated as a node-ranking problem. Exact centrality values are used as supervision, and ranking quality is evaluated using Kendall's tau rank correlation. We study whether message-passing GNNs can learn...

    arxiv.org/abs/2607.09372 · PDF

  19. 19

    Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning

    Guanquan Wang, Yoshimasa Tsuruoka

    cs.LG · cs.AI · cs.RO

    Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling steps, yet they typically rely on a two-stage teacher--student distillation pipeline that increases training cost and may introduce instability. We propose Shortcut Trajectory Planning (STP), an offline model-based...

    arxiv.org/abs/2607.09336 · PDF

  20. 20

    Risk-Aware General-Utility Markov Decision Processes

    Pedro P. Santos, Fábio Vital, Alberto Sardinha, Francisco S. Melo

    cs.LG · cs.AI

    We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives. In this framework, an agent aims to optimize a risk measure of the distribution of objective values, where the objective function depends on the frequency of visitation of states induced by the agent's policy. First, we motivate, propose, and formalize risk-aware GUMDPs, which enable agents and decision makers to trade off expected performance by risk...

    arxiv.org/abs/2607.09298 · PDF

  21. 21

    Super-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuning

    Ivan Ilin, Philip Zmushko, Peter Richtárik

    cs.LG · cs.CL

    Large language models (LLMs) remain expensive to fine-tune because full-parameter updates require substantial memory, compute, and per-task storage. We study whether saliency signals originally developed for pruning can be reused to choose where a model should adapt. We propose Super, a sparse parameter-efficient fine-tuning (PEFT) method that fixes a small trainable support using a Wanda-style activation-weighted magnitude score [Sun et al.,...

    arxiv.org/abs/2607.09287 · PDF

  22. 22

    Autoregressive latent diffusion for 3D molecule generation

    Federico Ottomano, Gaopeng Ren, Yingzhen Li, Kim E. Jelfs, Alex M. Ganose

    cs.LG

    Three-dimensional (3D) molecule generation has been dominated by diffusion models, which achieve strong generation quality but typically require the molecular size to be specified a priori. Recent autoregressive approaches have substantially narrowed the performance gap while naturally supporting variable-length generation and conditioning on partial molecular context. However, balancing unconditional and context-conditioned generation...

    arxiv.org/abs/2607.09277 · PDF

  23. 23

    LionVote: Per-Layer Learning Rate Adaptation for Lion

    Kris Atallah

    cs.LG

    Per-layer diagnostics reveal that, at the prescribed learning rate, Lion's effective scale is 2.6-2.8x too high for attention and MLP parameters and ~2x too high for normalization layers on ViT-Tiny/CIFAR-100; this 32% cross-layer-type disparity cannot be reproduced by a single global rate. The measurement comes from LionVote, a per-layer learning rate mechanism in which each parameter tensor maintains a compound level, a persistent integer...

    arxiv.org/abs/2607.09266 · PDF

  24. 24

    Forget Narrowly, Retain Broadly: Unlearning as an Asymmetric Generalization Problem

    Amit Peleg, Naman Deep Singh, Naama Pearl, Bibhabasu Mohapatra, Matthias Hein

    cs.LG

    Machine unlearning in LLMs is the targeted removal of specific knowledge while preserving all other capabilities, critical for privacy and safety. Yet existing benchmarks measure it unreliably. They miss knowledge that resurfaces under paraphrased or indirect queries, a failure we call under-forgetting, and lack the semantic, syntactic, and lexical probes needed to verify that unrelated knowledge is preserved, a failure we call...

    arxiv.org/abs/2607.09236 · PDF

  25. 25

    All you need is SAMPAT

    Jayadeva, Madhur Aswani

    cs.LG · cs.AI · cs.CV · math.FA

    The current state of the art in AI/ML rests on deep neural architectures, which, in general, suffer from a lack of interpretability. Interpretability is crucial to gleaning insights while analyzing experimental data, where quantitative predictions may not be adequate for a scientist. We present a three layer neural architecture, SAMPAT (Smooth Approximation via Multivariate Polynomials and Analytic Transformations), that can provably learn a...

    arxiv.org/abs/2607.09235 · PDF

  26. 26

    Temporal Knowledge Graph Forecasting under Distribution Shifts: A Synthetic Evaluation

    Konrad Özdemir, Julia Gastinger, Lukas Kirchdorfer, Heiner Stuckenschmidt

    cs.LG

    Temporal knowledge graphs (TKGs) represent evolving relational systems, whose underlying data-generating processes often change over time. Yet, TKG forecasting models are commonly evaluated only on empirical benchmark datasets that provide limited insight into the models' robustness to such distribution shifts. Recognising this issue, we study TKG forecasting under controlled shift environments using a synthetic TKG generator that encodes...

    arxiv.org/abs/2607.09232 · PDF

  27. 27

    Interference and Retention in Continual Learning

    Julius Störk

    cs.LG · cs.AI · cs.NE

    Continual learning commonly relies on post-hoc mechanisms such as replay, elastic regularization, or distillation. This work argues that forgetting should instead be modeled directly as interference between tasks. In the frozen-feature regime, forgetting from learning a new task is exactly the interference energy induced on the old task. In deep networks, the same quantity is recovered through path-averaged curvature with minimal additional...

    arxiv.org/abs/2607.09202 · PDF

  28. 28

    Application of machine learning to monster level prediction in tabletop RPG game design

    Jolanta Śliwa, Jakub Adamczyk

    cs.LG

    Designing balanced adversaries is a central but labor-intensive task in tabletop role-playing game (TTRPG) development. In systems such as Pathfinder, each monster is described by many numerical attributes that jointly determine its power, summarized as an ordinal level. We investigate whether machine learning can support designers by predicting this level from a monster's attributes, framing the task as tabular ordinal regression. We...

    arxiv.org/abs/2607.09196 · PDF

  29. 29

    Understanding Schedule-Free Methods in Nonconvex Optimization: Rate Guarantees and Escaping Saddles

    Jiseok Chae, Donghwan Kim

    cs.LG · math.OC

    Schedule-Free methods have attracted growing interest for alleviating the burden of designing and tuning a learning rate scheduler, while matching and sometimes even outperforming optimizers with tuned schedulers. Despite their strong empirical results, their convergence theory in nonconvex optimization, where modern machine learning objectives typically arise, has remained largely unexplored. In this paper, we provide worst-case analyses of...

    arxiv.org/abs/2607.09167 · PDF

  30. 30

    COAST: Context-Aware Differential Learning for Gene Expression Prediction in Spatial Transcriptomics

    Keunho Byeon, Sunhong Park, Jeewoo Lim, Jin Tae Kwak

    cs.LG

    Spatial transcriptomics enables profiling of spatial gene expression but is limited by high cost and low throughput, motivating prediction from H&E histopathology images. Existing context-aware methods mainly supervise absolute expression, while relative expression relationships between spots are rarely used explicitly. We propose COAST, a context-aware differential learning framework for spatial gene expression prediction. COAST conditions...

    arxiv.org/abs/2607.09166 · PDF

  31. 31

    A Personalized Computational Framework for Assessing the Sufficiency of Partially Observed Data in Healthcare AI models

    Qingchu Jin, Felistas Mazhude, Jamie B. Rabb, Robert S. Kramer, Douglas B. Sawyer, Raimond L. Winslow

    cs.LG · cs.AI

    Achieving early and timely diagnosis and treatment for disease is a major challenge. Recent applications of machine learning (ML) algorithms trained on patient data have shown promise in many different settings for predicting the patient health state. A challenge often faced when applying these ML algorithms is that at any given time, not all clinical variables (features) needed as input to perform prediction tasks are available. We define...

    arxiv.org/abs/2607.09165 · PDF

  32. 32

    Present but Rescaled: Chat-to-Agent Transfer of Additive Activation Steering

    Lucas Pinto

    cs.LG

    Additive activation steering (injecting a scaled residual-stream direction during generation) is calibrated almost entirely in single-turn chat, yet the models it targets are increasingly deployed as tool-using ReAct agents. We present the first systematic chat-to-agent transfer study of additive steering, coupling behavioral measurement with a representation read-out in a matched-information design: the same items rendered as plain chat or...

    arxiv.org/abs/2607.09156 · PDF

  33. 33

    Power Flow Feasibility Assessment Using Variational Graph Autoencoders

    Ferran Bohigas-Daranas, Hamid Latif-Martinez, Eduardo Prieto-Araujo, Pere Barlet-Ros, Oriol Gomis-Bellmunt

    cs.LG · eess.SY

    Data-driven methods, including graph neural networks, have been studied for accelerating power flow calculations in recent years, but very little attention has been paid to the solution feasibility, which can be obtained by traditional solvers. This paper presents a Variational Graph Autoencoder (VGAE) that detects the power flow solution feasibility, using the IEEE 118-bus case, to assess the validity of the solutions provided by AI-driven solvers.

    arxiv.org/abs/2607.09122 · PDF

  34. 34

    Quantum Circuits in Diffusion Models: A Fair-Comparison Study and a Mechanistic Analysis of Angle-Embedding Failures

    Jaeuk Kim, Sanghoon Yoo

    cs.LG

    We study the integration of variational quantum circuits (VQCs) into diffusion models through a squeeze-and-excitation (SE) channel-modulation scaffold that isolates the quantum contribution. Using a role-matched classical control and multi-seed significance testing across DDPM and latent diffusion on MNIST and CIFAR-10, with a score-based NCSN study on MNIST, we find that quantum cores achieve comparable mean FID to the classical control...

    arxiv.org/abs/2607.09108 · PDF

  35. 35

    EXHOLD: Experience-Aware Real-Time Hold Control for Large-Scale Ride-Hailing Matching at DiDi

    Xu Liu, Kai Wan, Zihao Lu

    cs.LG

    In large-scale ride-hailing, hold control is a critical mechanism for improving passenger-driver experience. By selectively delaying certain driver-order pairs, the system waits for better opportunities, reduces cancellations, and mitigates wasted driver effort. However, existing industrial hold strategies often rely on heuristic thresholding over multiple predictive models, which can be brittle under non-stationary traffic and hard to...

    arxiv.org/abs/2607.09090 · PDF

  36. 36

    A Survey on the Green Development of Large Models: From Resource-Efficient Architectures to Hardware-Software Co-Design

    Linhui Xiao, Guiping Cao, Mingyue Guo, Xianchao Guan, Fan Yang, Ming Tao, Xin Li, Yuxin Peng, Yaowei Wang

    cs.LG · cs.CY

    The rapid expansion of large-scale AI models has led to significant performance breakthroughs across diverse domains, yet it has also raised critical concerns regarding computational costs, energy consumption, and environmental sustainability. This survey provides a comprehensive overview of the green development of large models, emphasizing resource-efficient architectures and full-stack hardware-software co-design. We systematically review...

    arxiv.org/abs/2607.09084 · PDF

  37. 37

    Pitfalls and Remedies for Multi-Task Bayesian Optimization

    Carl Hvarfner, Sam Daulton, Max Balandat, Eytan Bakshy

    cs.LG

    Bayesian optimization routinely warm-starts a target experiment with data from related source tasks, and the multi-task Gaussian process is the textbook surrogate for the job. We revisit this default in a controlled setting and find that it misestimates the cross-task correlation even in the simplest non-trivial case, affinely related source and target tasks, where a working transfer learning method should obviously succeed. We trace the...

    arxiv.org/abs/2607.09073 · PDF

  38. 38

    EvoLP: Self-Evolving Latency Predictor for Model Compression in Real-Time Edge Systems

    Shuo Huai, Hao Kong, Shiqing Li, Xiangzhong Luo, Ravi Subramaniam, Christian Makaya, Qian Lin, Weichen Liu

    cs.LG

    Edge devices are increasingly utilized for deploying deep learning applications on embedded systems. The real-time nature of many applications and the limited resources of edge devices necessitate latency-targeted neural network compression. However, measuring latency on real devices is challenging and expensive. Therefore, this letter presents a novel and efficient framework, named EvoLP, to accurately predict the inference latency of models...

    arxiv.org/abs/2607.09063 · PDF

  39. 39

    COBS: Cumulant Order Block Sparse Attention

    Alexander Tian, Aditya Ghai, Sanjit Neelam, Zaal Vasania, Akshay Mishra

    cs.LG

    Block sparse attention is a hardware friendly way to alleviate the key-value (KV) cache read bottleneck in large language models (LLMs). However, it is not prevalent among leading open-weight LLMs, which rely instead on dense attention or fine-grained selection, thereby motivating our analysis. We study DeepSeek's Native Sparse Attention (NSA) as a representative method, whose three-branch design lets us isolate block selection, the most...

    arxiv.org/abs/2607.09052 · PDF

  40. 40

    Learning More from Less: Reinforcement Learning from Hindsight

    Iris Xu, Sunshine Jiang, John Marangola, Nitish Dashora, Richard Li, Thomas Liu, Zexue He, Yuheng Zhi, Alex...

    cs.LG

    Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern. Manipulation tasks typically provide only sparse rewards, so a weak policy fails almost every rollout early in training and has little to learn from, even when those failures execute coherent behavior. Such a failure,...

    arxiv.org/abs/2607.09042 · PDF

  41. 41

    Variable-Length Generative Protein Design via Generalized Poisson Flow

    Chaoran Cheng, Zhanghan Ni, Yanru Qu, Yuxin Chen, Ruihan Guo, Jiajun Fan, Ge Liu

    cs.LG · q-bio.QM

    The ability to generate variable-length proteins is crucial in protein design, where the optimal length is often unknown and tightly coupled to designability. Current diffusion- and flow-based generative models typically require the protein length to be specified before sampling, limiting their flexibility in exploring the feasible design space. To address this limitation, we introduce Generalized Poisson Flow (GPFlow), a variable-length...

    arxiv.org/abs/2607.09039 · PDF

  42. 42

    Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing

    Ajay Narayanan Sridhar, Ronak Singh, Mehrdad Mahdavi, Vijaykrishnan Narayanan

    cs.LG · cs.AI

    We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large language model (LLM) routing. Unlike classical contextual bandits that rely solely on bandit feedback and assume conditional independence across arms, our setting allows context-dependent inter-arm correlations and auxiliary reward information that may be noisy or...

    arxiv.org/abs/2607.09015 · PDF

  43. 43

    Model Agnostic Graph Prompt Learning for Crystal Property Prediction

    Shrimon Mukherjee, Kishalay Das, Partha Basuchowdhuri, Pawan Goyal, Niloy Ganguly

    cs.LG · cs.AI

    Graph Neural Networks have emerged as a powerful tool for the fast and accurate prediction of various crystal properties. These models often encode domain-specific knowledge into their graph encoding modules, which increases their parameter size and makes their performance heavily dependent on domain expertise. Added to this, explicitly incorporating all chemical and structural features, that might influence a specific crystal property into...

    arxiv.org/abs/2607.08996 · PDF

  44. 44

    Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models

    Bishmoy Paul, Youngmin Yi, Hoeseok Yang

    cs.LG · cs.CL

    Efficient inference in Large Language Models (LLMs) requires deciding where computation can be reduced while preserving model quality. We study this problem through multilayer perceptron (MLP) activation sparsification and token-level conditional routing. We first propose Sensitivity-Aware Thresholding for Sparsity (SATS), a threshold calibration method to choose layerwise gate thresholds using a local MLP output sensitivity proxy rather than...

    arxiv.org/abs/2607.08991 · PDF

  45. 45

    Group Invariant Spectral Embedding

    Yeari Vigder, Paulina Hoyos, David Thong, Joakim andén, Joe Kileel, Amit Moscovich

    cs.LG · math.NA · math.ST

    Spectral embedding methods are widely used for dimensionality reduction and clustering of high-dimensional datasets with intrinsic low-dimensional structures. Although many datasets of practical interest exhibit invariance under symmetries such as rotations, standard spectral embedding methods do not account for this, treating symmetry-related data points as unrelated. Our approach to this problem is to incorporate the symmetries directly...

    arxiv.org/abs/2607.08987 · PDF

  46. 46

    AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision

    Brent Kong, Tejas Ram, Tony Yue Yu

    cs.LG · cs.AI · cs.GT · math.CO

    AlphaZero has demonstrated that a neural-guided Monte Carlo Tree Search can achieve superhuman performance, but strong play does not necessarily imply perfect play. We study this gap in two oracle-evaluable domains with contrasting structure: Connect Four, a solved partisan game with exact game-theoretic values, and Chomp, an impartial game whose optimal play is governed by Grundy-number structure. Under a unified self-play $+$ MCTS pipeline,...

    arxiv.org/abs/2607.08984 · PDF

  47. 47

    Optimal Top-$k$ Identification from Pairwise Comparisons

    Motti Goldberger, Nils Rudi

    cs.LG · stat.AP · stat.ML · stat.OT

    We study the active learning problem of fixed-confidence top-$k$ identification from noisy pairwise comparisons. In this problem, an algorithm sequentially chooses pairs of items to compare, observes the outcomes, and stops when it can return the set of top-$k$ items with error probability at most $δ$. The objective is to design such a $δ$-correct procedure that minimizes the expected number of comparisons (the sample complexity). This...

    arxiv.org/abs/2607.08979 · PDF

  48. 48

    Federated Low-Rank Koopman Learning for Multivariate Time-Series Anomaly Detection in IoT Systems

    Tung-Anh Nguyen, Van-Phuc Bui, Anh Tuyen Le, Kim Hue Ta, Minh Thuy Le, J. Andrew Zhang, Xiaojing Huang

    cs.LG · eess.SP

    Distributed IoT systems generate multivariate time-series streams for monitoring physical assets, servers, and embedded sensing platforms. Detecting abnormal temporal behavior is critical for fault diagnosis, predictive maintenance, and security. However, practical IoT anomaly detection is hindered by decentralized and non-IID data, limited bandwidth, and the constrained computation and memory of edge devices. This paper proposes FedKAD, a...

    arxiv.org/abs/2607.08978 · PDF

  49. 49

    Stochastic Linear Bandits with Partially Observed Actions

    Gautam Dasarathy, Vineet Gattani, Lalit Jain

    cs.LG · math.ST · stat.ML

    The stochastic linear bandit, where actions are represented as vectors and rewards are linear, is a central paradigm for sequential decision making. We study a partially observed variant of this problem in which the learning agent only sees a random subset of coordinates for each action. Such partial observability arises naturally in settings like recommendation and healthcare, where full action descriptions can be expensive or even...

    arxiv.org/abs/2607.08971 · PDF

  50. 50

    NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision

    Berkay Anahtarci

    cs.LG · cs.AI · math.ST

    Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional labels reduce sampling error without resolving the resulting identification problem. We introduce Natural Language PAC (NL-PAC), a framework that uses a fixed model's thresholded decoding law to define...

    arxiv.org/abs/2607.08961 · PDF

  51. 51

    Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution

    Ning Liu, Kalle Kujanpää, Zhaoxuan Zhu, P Aditya Sreekar, Kaiwen Liu, Chuanneng Sun, Jorge Marchena Menendez,...

    cs.LG · cs.AI

    Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be executed reliably under strict time constraints, yet LLM agents lack mechanisms to enforce procedural compliance and degrade under the context overload full SOP specifications introduce. We present Eluna, a production-deployed agentic system for reliable SOP execution. Eluna is a graph-guided, multi-agent...

    arxiv.org/abs/2607.08960 · PDF

  52. 52

    FairSelect: A Systematic Evaluation of Multi-Level and Intersectional Algorithmic Fairness

    Nick Souligne, Isabella Mixton-Garcia, Vignesh Subbian

    cs.LG

    Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes. This limits practical guidance for selecting fairness strategies, where disparities may arise across intersectional subgroups and across multiple stages of the modeling lifecycle. This work presents FairSelect, a toolkit for systematically evaluating...

    arxiv.org/abs/2607.08953 · PDF

  53. 53

    Training, Reading, and Editing Legible Transformers

    Mark Oskin

    cs.LG · cs.CL

    A transformer can be built from operators that are legible by construction -- bounded, named units that read as fuzzy set operations rather than dense activations -- but legibility must be pressed for during training, and the pressure has a failure mode. A crispness penalty meant to sharpen a bounded operator into a decisive detector instead collapses it into a dead constant. An identity, E[v(1-v)] = mu(1-mu) - var, shows why -- the penalty...

    arxiv.org/abs/2607.08946 · PDF

  54. 54

    TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning

    Fangxu Yu, Tao Feng, Dehai Min, Lu Cheng, Ge Liu, Tianyi Zhou

    cs.LG

    Time series reasoning is essential for real-world problem-solving. While both Large Language Models (LLMs) and Vision-Language Models (VLMs) can reason about time-series data, their capabilities are complementary: LLMs process time series as text sequences and thus preserve exact numerical understanding, but struggle with global patterns, whereas VLMs efficiently capture these patterns by visualizing time series but may lose fine-grained...

    arxiv.org/abs/2607.08940 · PDF

  55. 55

    BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving

    Yuanjie Zhu, Liangwei Yang, Ke Xu, Weizhi Zhang, Shanghao Li, Zihe Song, Philip S. Yu

    cs.LG

    Efficient serving of diffusion large language models (dLLMs) is hindered by convergence heterogeneity: when batching multiple requests, different sequences converge at different rates, causing faster requests to stall behind slower stragglers and introducing compute bubbles and tail latency. We present BlockServe, a continuous batching framework that integrates block-grained scheduling -- immediately evicting completed requests at block...

    arxiv.org/abs/2607.08930 · PDF

  56. 56

    SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

    Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin

    cs.LG

    Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do. A standard mitigation hands control to a separate recovery policy whenever the agent leaves a...

    arxiv.org/abs/2607.08925 · PDF

  57. 57

    A Machine Learning Surrogate for Component Criticality Ranking in Interdependent Power-Communication Networks

    Sohini Roy, Xheni Hylviu

    cs.LG

    Cyber-physical power systems are vulnerable to cascading failures caused by tight interdependencies between power and communication infrastructures. Evaluating these failures over large N-k contingency sets with a high-fidelity simulator is computationally prohibitive for resilience planning. Using the previously published Modified Implicative Interdependency Model (MIIM) as the ground-truth cascade simulator, this paper develops a...

    arxiv.org/abs/2607.08918 · PDF

  58. 58

    Pattern-Aware Graph Neural Networks for Handling Missing Data

    Minett Tran, Taehee Jeong

    cs.LG

    Missing data is ubiquitous in real-world datasets. Traditional methods either discard incomplete samples or apply imputation techniques that ignore potentially informative missingness patterns, implicitly assuming that missingness occurs randomly. However, missingness patterns might provide additional information. We propose pattern-aware graph neural networks that explicitly encode which features are missing alongside observed values. We...

    arxiv.org/abs/2607.08915 · PDF

  59. 59

    Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal

    Ege Çakar, Hannah Guan, Kayden Kehe

    cs.LG

    Behavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directions in activation space. This raises questions about how such representations are structured, localized, and accessed by optimization. We study adversarial suffix attacks as a probe of representational alignment. We introduce Activation-Guided GCG, which replaces...

    arxiv.org/abs/2607.08883 · PDF

  60. 60

    How are linear representations learned? Exact solutions to the dynamics of abstraction

    William W. Yang, Andrew M. Saxe, Peter E. Latham

    cs.LG

    In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins many interpretability and control methods based on linear probes, from concept detection to activation steering. Yet while prior work has studied whether such directions should exist $\textit{after}$ training, the dynamics of...

    arxiv.org/abs/2607.08843 · PDF

  61. 61

    Prompt-Driven Exploration

    Sunshine Jiang, John Marangola, David Zhang, Raghuram Kowdeed, Ruiyang Luo, Nitish Dashora, Richard Li, Pulkit...

    cs.LG · cs.AI

    Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject stochasticity in the action space, but such jitter only yields rollouts close to the original. Escaping a weak policy often requires global perturbations that action noise cannot produce. Large language models (LLMs) and vision-language-action (VLA) models offer a pathway: they condition the policy on a...

    arxiv.org/abs/2607.08837 · PDF

  62. 62

    SLORR: Simple and Efficient In-Training Low-Rank Regularization

    David González-Martínez, Shiwei Liu

    cs.LG · cs.AI

    Low-rank factorization is widely used to compress neural networks, but modern models are often not naturally amenable to aggressive factorization without significant accuracy loss. Existing training-time low-rank regularizers can improve compressibility, but they often require SVDs of large weight matrices, modify the model architecture (introducing additional trainable parameters), or rely on stateful cached quantities. To address these...

    arxiv.org/abs/2607.08754 · PDF

  63. 63

    Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph

    Duen Horng Chau, Donghao Ren, Fred Hohman, Dominik Moritz

    cs.LG · cs.AI · cs.DS · cs.HC

    While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally. This graph encodes the data manifold in its original high-dimensional space, before the distortion that UMAP's 2D projection introduces. We demonstrate the untapped potential of this internal representation, showing how standard...

    arxiv.org/abs/2607.08746 · PDF

  64. 64

    Super Weights in LLMs and the Failure of Selective Training

    Shreyas Subramanian, Adewale Akinfaderin, Akarsha Sehwag

    cs.LG

    Recent work identified Super Weights, individual parameters whose removal degrades model performance by orders of magnitude. We show that this degradation due to pruning Super Weights does not universally apply to all LLMs. Furthermore, if these parameters are so important, Super Weight-aware training should be effective. We show the opposite. Training Super Weights in isolation (100 to 8,192 parameters) drops accuracy to random-guessing...

    arxiv.org/abs/2607.08733 · PDF

  65. 65

    Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference

    Chuning Zhu, Eva Xu, Jose Barreiros, Krishnan Srinivasan, Paarth Shah, Abhishek Gupta

    cs.LG · cs.RO

    Human decision-making is highly flexible -- some actions are taken immediately; others require longer deliberation. Language models have exhibited a similar capacity for adaptive "reasoning." However, transferring this capability to continuous control policies has been challenging, as directly reasoning in language space may lack the granularity for spatial understanding and precise motions. In this work, we show that reasoning for control...

    arxiv.org/abs/2607.08724 · PDF

  66. 66

    Deep Learning for Joint Narrowband Interference Cancellation and Soft Demodulation in OFDM Systems

    Emmanouil Kavvousanos, Francky Catthoor, Vassilis Paliouras

    cs.LG · eess.SP

    Narrowband interference (NBI) severely degrades orthogonal frequency-division multiplexing (OFDM) systems by corrupting subcarriers and rendering classical soft demodulation ineffective. Conventional compressed-sensing (CS) mitigation exhibits high sequential latency and leaves structured, non-Gaussian residuals that cause log-likelihood ratio (LLR) unreliability, decoder saturation, and severe error floors when employing classical Gaussian...

    arxiv.org/abs/2607.08717 · PDF

  67. 67

    MPFlow: Learning Budgeted Max-Flow Optimization on the Lightning Network with Deep Graph Reinforcement Learning

    Harrison Rush, Vincent Davis, Simone Antonelli, Vikash Singh, Jesse Shrader, Emanuele Rossi

    cs.LG

    We address liquidity placement in the Bitcoin Lightning Network (LN): given a fixed budget, which channels should a node open to maximize its routing capacity? We cast this as a budget-constrained combinatorial optimization problem on graphs, selecting $k$ edge additions that maximize $s$--$t$ max-flow, a theory-grounded measure of routing capacity, and solve it with graph reinforcement learning. Our lightweight agent combines a...

    arxiv.org/abs/2607.08703 · PDF

  68. 68

    A Practical Investigation of Training-free Relaxed Speculative Decoding

    Guoxuan Xia, Luka Ribar, Paul Balanca

    cs.LG · cs.AI

    Speculative decoding accelerates sampling from an autoregressive LLM by using a faster auxiliary model to draft tokens which are then verified in parallel by the LLM. Standard speculative decoding is lossless: its rejection and resampling steps exactly preserve the LLM's sampling distribution. Recent work argues that relaxing this strict guarantee can yield further speed-ups, controlled capability-speed trade-offs, or even capability gains....

    arxiv.org/abs/2607.08690 · PDF

  69. 69

    Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models

    Teng-Ruei Chen

    cs.LG

    Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle. Recent analysis shows that test-time resampling can recover per-instance selection headroom that no single-commit router captures; however, that guarantee holds only under an idealized oracle equipped with correctness labels and an unconstrained budget, neither of which a...

    arxiv.org/abs/2607.08665 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.