cs.LG · 2026-07-22 · No. 61

Machine Learning, 2026-07-22.

73 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

73 entries
  1. 01

    Provable diffusion-based posterior sampling for linear inverse problems via DDIM

    Yuchen Jiao, Na Li, Changxiao Cai, Yuxin Chen, Gen Li

    cs.LG · cs.AI · stat.ML

    Diffusion-based methods have achieved remarkable empirical success in solving inverse problems. However, many existing posterior samplers either lack rigorous theoretical guarantees or incur substantial computational overhead. We propose a simple and efficient algorithm, called \pddim, for solving linear inverse problems with diffusion priors via a DDIM-type sampler. Our method requires only lightweight, coordinate-wise modifications to the...

    arxiv.org/abs/2607.19333 · PDF

  2. 02

    ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling

    Chirag Vashist, Ke Li

    cs.LG · cs.CV

    Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching. Along the way, the underlying techniques have become more complicated and various beliefs about what drives strong empirical performance have taken hold. Due to the success of diffusion models and flow matching, one of the more common beliefs is the importance of transforming the noise distribution to the data distribution gradually...

    arxiv.org/abs/2607.19332 · PDF

  3. 03

    ISO: An RLVR-Native Optimization Stack

    Hanqing Zhu, Wenyan Cong, Zhizhou Sha, Sagnik Mukherjee, Xinyuan Song, David González-Martínez, Xiaoxia Wu, Yuandong...

    cs.LG · cs.AI

    Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through the singular structure of model weights and identify spectral inheritance: RLVR can reuse the base model's weight spectra while...

    arxiv.org/abs/2607.19331 · PDF

  4. 04

    CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability

    Pratinav Seth, Hem Gosalia, Aditya Kasliwal, Vinay Kumar Sankarapu

    cs.LG · cs.CL · cs.ET

    Circuit analysis can support not only model explanation but also downstream interventions such as pruning, editing, steering, and selective fine-tuning. However, conducting such analyses currently requires stitching together separate implementations for discovery, evaluation, and intervention, as well as hand-authoring the contrastive prompts required by many discovery methods. This fragmentation makes methods difficult to compare and limits...

    arxiv.org/abs/2607.19317 · PDF

  5. 05

    Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

    Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou, Jalaj Bhandari, Kavosh Asadi, Daniel Jiang, Aditya Modi

    cs.LG · cs.AI

    Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{zero} learning signal. Providing privileged guidance during training, such as solution prefixes, can help overcome this learning cliff by steering the model towards {correct solutions with non-zero reward}. {We call...

    arxiv.org/abs/2607.19313 · PDF

  6. 06

    Staypoint Detection from Noisy Trajectory Data [Experiment Paper]

    Lance Kennedy, Hossein Amiri, Yueyang Liu, Riyang Bao, Hanqi Chen, Mohammad Hashemi, Ruochen Kong, Xiaotong Liu,...

    cs.LG · cs.CG

    Detecting staypoints from raw trajectory data is fundamental to numerous spatial computing applications. This process transforms raw numeric sequences of geolocations into semantically meaningful locations, such as homes, workplaces, or restaurants. Despite its importance for semantic trajectory analysis, staypoint detection lacks standard benchmarks, and existing algorithms have never been systematically evaluated. This gap persists because...

    arxiv.org/abs/2607.19312 · PDF

  7. 07

    Riemannian Deep Learning:Modules, Networks, and Geometries

    Chen Ziheng

    cs.LG · cs.AI · math.DG

    Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operations. This thesis develops a unified framework for Riemannian deep learning from three complementary perspectives: reusable neural modules, manifold-specific network architectures, and the design of...

    arxiv.org/abs/2607.19305 · PDF

  8. 08

    Real-time optimal control with shallow recurrent decoder networks

    Matteo Tomasetto, Francesco Braghin, J. Nathan Kutz, Andrea Manzoni

    cs.LG · math.OC

    Controlling dynamical systems in real-time across multiple scenarios is critical to enabling adaptive control strategies, ensuring stability and efficiency. However, to tailor control actions in response to varying scenarios, traditional optimal control problems typically require several system simulations, which are often computationally demanding due to the high-dimensionality of the underlying spatio-temporal dynamics. In this work, we...

    arxiv.org/abs/2607.19302 · PDF

  9. 09

    A Reinforcement-Learning-Augmented Liquid-Fueled Reactor Network Model for Predicting Lean Blowout in Gas Turbine Combustors

    Philip John, Eloghosa Ikponmwoba, Pinaki Pal, Opeoluwa Owoyele

    cs.LG

    This study introduces a reinforcement learning (RL) framework for generating optimal liquid-fueled reactors to improve lean blowout (LBO) predictions in gas turbine combustors. Existing approaches for determining cluster boundaries rely on manual heuristics or distance-based metrics in the input space. In contrast, the proposed method is goal-oriented, explicitly accounting for the target metric (e.g., LBO prediction accuracy) during cluster...

    arxiv.org/abs/2607.19281 · PDF

  10. 10

    GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models

    Alessandro Scalese, Santhanakrishnan Narayanan, Constantinos Antoniou

    cs.LG · cs.AI

    The Traffic Assignment Problem is a fundamental but computationally expensive component of transportation planning. While Graph Neural Networks have emerged as fast, data-driven surrogates, their practical deployment is severely constrained by a spatial generalization gap. Standard models rely on transductive feature initializations that tie travel demand to fixed network topologies, preventing seamless transfer to new urban environments. To...

    arxiv.org/abs/2607.19270 · PDF

  11. 11

    Toward Auditable Fraud Detection: Combining Graph Features, Model Explanations, and Agentic Case Investigation

    Rahil Sharma

    cs.LG · cs.AI

    Fraud detection systems must scale with rising transaction volume while remaining explainable and reviewable. We study a layered pipeline on the PaySim dataset that combines a gradient-boosted classifier, graph-derived structural features, an autoencoder-based anomaly signal, TreeSHAP explanations, and a bounded LLM investigation agent applied to cases the classifier scores uncertainly. Before any model comparison, we identify and remove a...

    arxiv.org/abs/2607.19266 · PDF

  12. 12

    Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks

    Guy Stephane Waffo Dzuyo, Gaël Guibon, Christophe Cerisara, Luis Belmar-Letelier

    cs.LG · cs.AI

    Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and under-utilized textual data in financial reports. Existing methods often rely on random data splits, leading to overoptimistic performance estimates that do not reflect real-world generalization to new companies or future periods. To address this recurring problem with the state of the art, we propose a...

    arxiv.org/abs/2607.19259 · PDF

  13. 13

    Thermodynamics-Informed Input Reparameterization for Neural Prediction of Real-Fluid Thermodynamic Properties in Supercritical Combustion

    Haoze Zhang, Han Li, Ke Xiao, Yangchen Xu, Runze Mao, Zhi X. Chen

    cs.LG · physics.comp-ph · physics.flu-dyn

    Real-fluid thermodynamic property evaluation is a major computational cost in supercritical combustion simulations. In the enthalpy-based pressure-correction formulation, the closure evaluates temperature T, density $ρ$, and compressibility coefficient $ψ$ from the solver state (h,p,Y) through enthalpy-temperature inversion and repeated real-fluid equation-of-state evaluations. Neural-network surrogates offer fixed-cost inference, but direct...

    arxiv.org/abs/2607.19241 · PDF

  14. 14

    DBMol: Design of High-Affinity, Target-Specific Small Molecules through Structure Prediction Models

    Yiming Qin, Kai Yi, Miruna Cretu, Sjors H. W. Scheres, Pietro Liò, Pascal Frossard

    cs.LG

    Designing small molecule ligands that bind with high affinity to specific protein pockets is a fundamental goal in drug discovery, as small molecules constitute a major fraction of approved therapeutics. Recent breakthroughs in structure prediction, such as AlphaFold-3 and Boltz-2, enable accurate biomolecular interaction prediction and show promise as foundation models for downstream tasks, including binding affinity prediction. We propose...

    arxiv.org/abs/2607.19237 · PDF

  15. 15

    In-Context Time Series Classification with Random Convolutional Features

    Joscha Cüppers, Jilles Vreeken

    cs.LG

    Time series classification is central to domains like medical signal analysis, industrial monitoring, and sensor-based activity recognition, where class information manifests as localized shapes, specific frequencies, temporal shifts, or complex cross-channel interactions. Random convolutional transforms efficiently map these sequences to fixed-dimensional tabular features but are traditionally paired with simple linear classifiers. We...

    arxiv.org/abs/2607.19234 · PDF

  16. 16

    S3: Stable Subgoal Selection by Constraining Uncertainty of Coarse Dynamics in Hierarchical Reinforcement Learning

    Kshitij Kumar Srivastava, Kshitij Jerath

    cs.LG · cs.MA

    Hierarchical Reinforcement Learning (HRL) intends to separate strategic planning from primitive execution. It has been widely successful in solving long-horizon and complex tasks, where flat-RL algorithms have difficulty in learning. However, while the low-level agent in HRL benefits from dense feedback and abundant trial opportunities, the high-level agent receives sparse, delayed feedback from the environment and its performance depends on...

    arxiv.org/abs/2607.19232 · PDF

  17. 17

    AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

    Yu-Yang Qian, Hao-Cong Wu, Chen Chen, Jiacheng Sun, Zhenhua Dong, Peng Zhao, Zhi-Hua Zhou

    cs.LG · cs.CL

    Speculative decoding, in which a lightweight draft model first generates a draft sequence that is then verified in parallel by the target model, has become a prevalent paradigm for accelerating large language model inference. Recent work such as DFlash further boosts drafting efficiency by leveraging diffusion drafters, whose parallel denoising mechanism enables draft generation in a single forward pass. In this work, we uncover a central...

    arxiv.org/abs/2607.19223 · PDF

  18. 18

    Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation

    Li-Rong Zhou, Qin-Wen Luo, Sheng-Jun Huang

    cs.LG

    Offline reinforcement learning (RL) aims to learn an effective policy from a static dataset, but its performance is fundamentally limited by dataset coverage. Action preference queries leverage expert feedback without additional environment interaction, enabling policy improvement during offline training. However, existing methods still face two key challenges: selecting informative preference queries and effectively exploiting the collected...

    arxiv.org/abs/2607.19199 · PDF

  19. 19

    Neural Kolmogorov Equations: Parallelizable Learning of Stochastic Dynamics under General Noise

    Arthur Bizzi, Olga Fink

    cs.LG

    Neural stochastic differential equations (SDEs) have emerged as powerful tools for learning noisy or stochastic dynamics directly from data; however, existing approaches largely assume uncoupled and continuous noise, limiting their applicability to realistic stochastic drivers, and often scale poorly in time, requiring expensive autoregressive training. To address these limitations, we propose Neural Kolmogorov Equations (NKEs), a...

    arxiv.org/abs/2607.19173 · PDF

  20. 20

    Breaking the Homogeneity Assumption: Specialized Multi-Generator Adversarial Learning for Rare Failure Detection in Predictive Maintenance

    Alexis Lazanas, Georgios Kampouropoulos

    cs.LG · cs.AI

    Supervised learning models in the predictive maintenance field are regularly trained on highly imbalanced industrial datasets: machine failures occur rarely but have a disproportionate effect on operations. In addition to the clear class disparity, failure data are typically non-homogeneous, with different failure modes arising from distinct physical processes and exhibiting a multimodal distribution across minorities and classes. Traditional...

    arxiv.org/abs/2607.19153 · PDF

  21. 21

    Incomplete Observations Boost Evolutionary Performance in Ocean Modeling

    Yangyang Kong, Yutong Jiang, Yanhai Gan, Junyu Dong, Feng Gao, Xiaopei Lin

    cs.LG · cs.AI

    Data-driven methods have revolutionized ocean modeling, yet current approaches rely heavily on complete reanalysis datasets, imposing computational constraints and limiting model performance to that of the training data. Here, we present a generative state-space model and an optimization framework that enable learning directly from sparse and noisy observations. The model is essentially a hidden Markov model with a continuous state space,...

    arxiv.org/abs/2607.19147 · PDF

  22. 22

    One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models

    Jiayi Yang, Yifang Chen, Yuanfu Sun, Jiajin Liu, Qiaoyu Tan

    cs.LG

    Vision-language models (VLMs) provide a unified representation space for textual and visual information, yet their potential as general-purpose backbones for graph-structured data remains largely unexplored. In practice, attributed graphs exhibit substantial modality heterogeneity: some graphs contain only textual node attributes, others only visual attributes, while still others provide both. Existing graph learning approaches are typically...

    arxiv.org/abs/2607.19128 · PDF

  23. 23

    Parallel Noising in Neural Markov Logic Networks

    Peter Jung, Giuseppe Marra, Ondrej Kuzelka

    cs.LG · cs.AI

    Neural Markov Logic Networks (NMLNs) are a flexible neurosymbolic relational model. Previous work has shown that, although NMLNs achieve strong performance as generative models for small relational structures, they underperform diffusion-based generative graph models on larger structures. In this paper, we strengthen NMLNs along two main dimensions: (i) we increase the expressive capacity of their potential functions using graph neural...

    arxiv.org/abs/2607.19126 · PDF

  24. 24

    Predicting Activities in Aqueous Electrolyte Solutions with Hybrid Machine Learning

    Zeno Romero, Maximilian Kohns, Fabian Jirasek

    cs.LG · physics.chem-ph

    Activities in aqueous electrolyte solutions, usually described by ionic activity and osmotic coefficients, are important properties for modeling many processes in industry and nature. Established activity models, such as those of Pitzer or Bromley, require fitting to experimental data for each electrolyte of interest and thus cannot predict properties for unstudied systems. While some predictive approaches exist, they are typically limited in...

    arxiv.org/abs/2607.19114 · PDF

  25. 25

    An unsupervised clustering analysis of breast cancer data derived from electronic health records enhanced through UMAP dimensionality reduction

    Davide Chicco, Nicoletta Benvenuto

    cs.LG · q-bio.QM

    Breast cancer is one of the most widespread types of cancer, affecting approximately 8 million women worldwide. Electronic health records of patients diagnosed with this disease can serve as valuable datasets for computational analyses, enabling the discovery of new insights about the pathology. Unsupervised clustering, in particular, can identify groups of patients with medically significant features, revealing data trends that might...

    arxiv.org/abs/2607.19089 · PDF

  26. 26

    GEqTrain: A Configuration-Driven Framework for Retargeting Equivariant Graph Neural Networks Across 3D Scientific Tasks

    Daniele Angioletti, Marco Nobile, Vittorio Limongelli

    cs.LG · q-bio.BM

    Equivariant graph neural networks provide a powerful modeling language for three-dimensional scientific data, but their reuse is often limited by implementations tied to specific tasks, outputs, and training regimes. We present GEqTrain, a configuration-driven framework that separates dataset semantics, model composition, and training objectives. Raw data are mapped to typed node-, edge-, and graph-level fields, while model stacks, losses,...

    arxiv.org/abs/2607.19083 · PDF

  27. 27

    Deep learning-based prediction of time-resolved adhesive forces in viscoelastic Hertzian contacts

    Ali Maghami, Merten Stender, Michele Ciavarella, Antonio Papangelo

    cs.LG · cond-mat.soft · cs.AI · physics.data-an · stat.ML

    Fast prediction of the response of adhesive soft viscoelastic contacts represents a current challenge in soft robotics and for gripping and manipulation tasks. Determining the complete time-resolved force trajectory requires full numerical simulations, whose computational cost is strongly parameter-dependent, making them impractical for real-time application or design-optimization loops. In this work, we overcome this limitation by training a...

    arxiv.org/abs/2607.19060 · PDF

  28. 28

    Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training

    Nuemaan Malik

    cs.LG · cs.AI

    Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on a 6.78B-parameter MoE language model, AdamW keeps 50.6 GB of first and second moments to update 12.6 GB of bfloat16 weights. We study SkewAdam, an optimizer built on the observation that the three parameter populations of an MoE - the dense backbone, the experts, and the router - differ enough in size and gradient statistics that they...

    arxiv.org/abs/2607.19058 · PDF

  29. 29

    Probabilistic Physics-Aware Machine Learning Predictions of Electric Truck Energy Consumption with Field Data

    Hannes Nilsson, Rafael Basso, Balázs Kulcsár, Morteza Haghir Chehreghani

    cs.LG

    In this work, we incorporate first principle physics into the construction of data-driven methods by considering a model that accounts for the different sources of energy losses during vehicle operations. Our results show that Bayesian linear regression based on this physics-aware model can improve the reliability of the expected energy consumption, as compared with standard linear regression. Further, it is shown that more complex machine...

    arxiv.org/abs/2607.19054 · PDF

  30. 30

    Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation

    Mingxuan Ouyang, Hao Lan, Wanyu Lin

    cs.LG

    Leveraging large language models (LLMs) for molecular generation has shown remarkable potential in chemical and drug design. Current methods primarily rely on supervised training or fine-tuning with limited datasets, which are insufficient to capture complex molecular design objectives. While some approaches attempt to guide generation toward specific goals, they often lack direct optimization mechanisms, making it difficult to align...

    arxiv.org/abs/2607.19044 · PDF

  31. 31

    Spectral Higher-Order Neural Networks Have Sharp Expressivity Bounds

    Gianluca Peri, Diego Febbe, Duccio Fanelli

    cs.LG · cs.AI

    Neural hypergraphs are a natural generalization of neural networks, the reference models in modern machine learning. Yet, their deployment has proven demanding: the number of weighted hyperedges required leads to an intractable parameter explosion. However, a novel parametrization that leverages spectral attributes for neural hypergraphs has been recently proposed, that enables to recycle parameters via a weight sharing scheme and...

    arxiv.org/abs/2607.19042 · PDF

  32. 32

    Unsupervised Multi-kernel Learning for Automated Algorithm Selection

    Yihang Lu, Tome Eftimov, Carola Doerr

    cs.LG

    Automated algorithm selection in black-box optimization typically relies on supervised models that map landscape features to algorithm performance labels. Such models are costly to train, benchmark-dependent, and often fail to generalize to unseen problem classes. We study an unsupervised alternative: multi-kernel clustering over heterogeneous landscape representations, in which problem instances are grouped without using performance labels...

    arxiv.org/abs/2607.19031 · PDF

  33. 33

    Biological Amnesia in ICU Time-Series Prediction: A Drift-Adaptive Two-Stream Architecture with Temporal Retrieval

    Fatema Ferdous Tamanna, K. M. Merajul Arefin, Md. Abdul Masud

    cs.LG · cs.AI · cs.IR · q-bio.QM

    Background: Clinical decision support systems degrade silently as treatment protocols evolve, yet standard adaptation methods treat models as monolithic blocks, unable to distinguish stable patient physiology from shifting institutional practice. Methods: We propose an adaptive clinical intelligence architecture for ICU intervention prediction that structurally decouples physiological from treatment representations, confining parameter...

    arxiv.org/abs/2607.19020 · PDF

  34. 34

    Subject-Conditioned Glucose Forecasting in Type-1 Diabetes

    Giorgia Rigamonti, Mirko Paolo Barbato, Davide Marelli, Paolo Napoletano

    cs.LG · q-bio.QM

    Accurate forecasting of blood glucose concentration is key in the management of Type 1 Diabetes, facilitating early detection of adverse glycemic events and supporting timely therapeutic interventions. Despite recent advances in glucose prediction, most existing approaches rely on population-level representations or implicit personalization strategies that fail to deliver effective subject-specific forecasts. In this work, we propose...

    arxiv.org/abs/2607.19006 · PDF

  35. 35

    Variational meta-learning inference for low dimensional neural system identification

    Matteo Rufolo, Dario Piga, Marco Forgione

    cs.LG · cs.AI · eess.SY

    Deep learning has proven highly effective for nonlinear system identification, but heavily parameterized neural networks are prone to overfitting in low-data regimes and lack reliable uncertainty quantification. The recently developed manifold meta-learning framework addresses the data efficiency problem by restricting the model parameters to a meta-learned low-dimensional manifold. However, that method is purely deterministic. We propose a...

    arxiv.org/abs/2607.18965 · PDF

  36. 36

    SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement

    Arther Tian, Alex Ding, Simon Wu, Aaron Chan

    cs.LG · cs.AI · cs.CR

    Procuring supervised fine-tuning (SFT) data forces a buyer to decide, before any downstream training, whether a candidate corpus is worth acquiring. We present \sys{}, a statistics-first gating architecture that treats procurement as a cost-aware routing problem over three intrinsic quality axes -- diversity, utility, and redundancy. Cheap blind measurements are summarised into per-axis estimates with confidence intervals; a gate accepts a...

    arxiv.org/abs/2607.18960 · PDF

  37. 37

    H$^2$SD: Hybrid Hindsight Self-Distillation

    Qiye Cai, Yichuan Ma, Linyang Li, Peiji Li, Yongkang Chen, Qipeng Guo, Yicheng Zou, Tao Gui, Xiaocheng Feng, Bing Qin

    cs.LG · cs.CL

    Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation. However, most RLVR methods assign a scalar outcome reward to an entire trajectory, resulting in sparse supervision and limited token-level credit assignment. On-policy distillation (OPD) provides denser supervision by distilling token-level...

    arxiv.org/abs/2607.18955 · PDF

  38. 38

    Functional Equivalence and Geometric Diversity in Neural Network Approximations: An Empirical Characterization

    Anuragine S A, Prem Jagadeesan

    cs.LG · cs.AI

    The Universal Approximation Theorem states that a neural network with a single hidden layer is sufficient to approximate any continuous univariate function on a compact domain to arbitrary error. However, the uniqueness of such neural network representations is not guaranteed, raising questions about practical identifiability. In this work, we address this concern by analyzing functional equivalence and geometric diversity of neural network...

    arxiv.org/abs/2607.18930 · PDF

  39. 39

    Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning

    Stella Ho, Joel Villalobos, Joseph West, Jingyang Liu, Weijie Qi, Haruhiko Kishima, Ryohei Fukuma, Takufumi...

    cs.LG · q-bio.NC

    ECoG-based visual semantic decoding enables inference of semantic interpretation of visual perception from complex, noisy brain activity. This study examines the feasibility of visual semantic decoding using an end-to-end deep learning framework using electrocorticography (ECoG). Specifically, the decoding task is to predict visual categories from video stimuli using time-series neural inputs. A previously collected ECoG dataset from...

    arxiv.org/abs/2607.18923 · PDF

  40. 40

    Circuit Claims Depend on What Is Extracted and How It Is Compared

    Yang Sheng, Jie Fu

    cs.LG · cs.AI

    Circuit extraction identifies a small set of model components whose presence preserves a target behavior under ablation, and the resulting circuit is often read as the mechanism behind that behavior. We argue that this reading is under-determined: preserving behavior does not single out one circuit, because the claim it supports depends on which circuit is reported and how two circuits are compared. We make this concrete in a synthetic Lean...

    arxiv.org/abs/2607.18921 · PDF

  41. 41

    Breaking Feedback-Blindness: Utility-Augmented Transformer for Sequential Decision Making

    Yuyang Shen, Shan Dai, Daimin Chen

    cs.LG

    Sequential decision making in non-stationary and partially observable environments requires rapid adaptation to latent regime changes. However, existing Transformer decision models face a structural bottleneck in the retrieval mechanism: even when reward is used for training or exposed as an input token, attention retrieval remains primarily driven by observation-derived similarity. We formalize this limitation as feedback-blind retrieval,...

    arxiv.org/abs/2607.18910 · PDF

  42. 42

    KALE: Kernel Alignment with Loss Equilibration for Stable CLIP-DINOv2 Alignment at Web Scale

    Michał Pawłowicz

    cs.LG

    Kernel-based alignment of CLIP toward a vision centric teacher such as DINOv2 (KUEA) improves CLIP's visual representations while preserving text-encoder compatibility, using a fixed trade-off weight tuned on curated ImageNet-1K. We ask whether this transfers to noisy, web-scale data (CC12M) and find that it does not: the alignment term's weighted contribution falls to about 0.2% of the clean term, so under any fixed weight its gradient is...

    arxiv.org/abs/2607.18885 · PDF

  43. 43

    RAMP: Recognition parametrisation by Amortised Message Passing

    Lior Fox, Kai Biegun, James Heald, Samo Hromadka, Arielle Rosinski, Maneesh Sahani

    cs.LG · cs.AI

    A central aim of unsupervised learning is to uncover latent factors that explain dependencies among observations. Probabilistic models typically achieve this by introducing multiple latent variables linked through a graph of conditional relationships, with distributional parameters and their dependence learnt from data. Learning relies either on distributional choices that allow tractable belief propagation, or on approximations that scale...

    arxiv.org/abs/2607.18883 · PDF

  44. 44

    Physics-Informed Super-Resolution of Atmospheric Data

    Chang Xu, Gencer Sumbul, Hugo Porta, Manon Béchaz, Sebastian Schemm, Devis Tuia

    cs.LG

    In the context of global warming, extreme events have become more frequent and intense, making their trustworthy detection and forecasting more important than ever. Yet, atmospheric observations lack sufficient spatial resolution, motivating atmospheric data downscaling as a way to reconstruct high-resolution data from coarse observations. This task is now being formulated as a super-resolution (SR) problem with machine learning methods...

    arxiv.org/abs/2607.18877 · PDF

  45. 45

    Reinforcement Learning for Delivery Drone-Based Participatory Sensing in Dynamic Environments

    Xin Ouyang, Songxin Lei, Xusen Guo, Yutian Jiang, Sijie Ruan, Yuxuan Liang

    cs.LG · cs.CY

    Using Unmanned Aerial Vehicle (UAV) for urban sensing has emerged as a powerful paradigm to monitor the status of the city, e.g., air quality and noise levels, through agile aerial crowdsourcing. Despite this potential, existing UAV-based sensing approaches overlook environmental disturbances like wind that drastically impact drone velocity and energy efficiency. Consequently, directly applying existing methods to this joint delivery and...

    arxiv.org/abs/2607.18874 · PDF

  46. 46

    HindsightBench: A Black-Box Behavioral Audit Protocol for Parametric Hindsight in Time-Indexed LLM Decision Tasks

    Haozhe Jia

    cs.LG · cs.CL

    Large language models leak parametric knowledge of realized outcomes into historical financial decision tasks. Existence is settled; what users lack is a cheap way to audit a given model for it. We present HindsightBench, a black-box behavioral audit protocol that profiles parametric hindsight in any time-indexed LLM decision task at probe-level cost (no backtests, no logprobs, no corpus access). The protocol chains a four-arm...

    arxiv.org/abs/2607.18867 · PDF

  47. 47

    Regime-Aware Physics-Guided Early Warning of Lithium-Ion Battery Thermal Runaway Using Thermo-Mechanical Signals

    Syed Sajid Ullah, Muhammad Zunair Zamir, Salman Khan

    cs.LG · cs.AI

    Thermal runaway in lithium-ion batteries poses a major safety risk to electric vehicles and energy storage systems. Current early-warning methods depend mainly on temperature and may therefore miss mechanical precursors that emerge before rapid heating. We introduce a regime-aware, physics-guided framework that integrates temperature, voltage, force, deformation, and state-of-charge measurements for early warning under controlled mechanical...

    arxiv.org/abs/2607.18860 · PDF

  48. 48

    ABOPD: Antibody CDR Design via On-Policy Distillation

    Zhuo Yang, Jiaying He, Jiaqing Xie, Daolang Wang, Xipeng Qiu, Yuxin Wang, Tianfan Fu, Beilun Wang

    cs.LG · cs.AI

    Antibodies are essential therapeutic molecules, and their complementarity-determining regions (CDRs) form the primary antigen-recognition interface. Recent protein generative models have demonstrated broad capabilities in biomolecular design, yet post-training strategies for downstream objectives remain limited. Standard denoising training operates on noisy states obtained by perturbing native structures, whereas recursive generation proceeds...

    arxiv.org/abs/2607.18835 · PDF

  49. 49

    From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning

    Garvit Singla, Uma Maheswari Natarajan, Raghuram Bharadwaj Diddigi

    cs.LG · cs.AI

    Model-Agnostic Meta-Learning (MAML) is a widely used framework for reinforcement learning (RL) that enables efficient transfer by learning global policy parameters that can be rapidly adapted to new tasks. MAML training proceeds in two loops: an inner loop where the global parameters are adapted to task-specific parameters, and an outer loop where these task-specific parameters are evaluated and losses are back-propagated to improve the...

    arxiv.org/abs/2607.18830 · PDF

  50. 50

    Countercurrent Multiplier Networks: A Renal-Inspired Iterative Operator with Provably Bounded Fixed-Point Dynamics

    Snigdha Chandan Khilar

    cs.LG · math-ph

    The mammalian kidney concentrates urine using a mechanism with no analogue in current neural architectures: the countercurrent multiplier. Two anti-parallel flows joined at a hairpin recirculate a weak magnitude-bounded local pump into a large axial gradient achieving a four-fold concentration increase from a single-effect gradient that never exceeds 200 mOsm at any point. We formalize this mechanism as a differentiable sequence operator the...

    arxiv.org/abs/2607.18829 · PDF

  51. 51

    Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior

    Garrett Baker, Vinayak Pathak, Daniel Murfet, Susan Wei

    cs.LG · stat.ML

    In the \emph{latent posterior model} of transformer behavior, the next-token distribution arises from a posterior over latent predictive models conditioned on the context, mixed to generate continuations. We exploit this model in settings where it is exact, namely Bayes-filtered transformers (BFTs) meta-learned on sequences from a hierarchical prior, to introduce \textbf{Posterior Prefix Tuning (PPT)}, a new method for \emph{eliciting}...

    arxiv.org/abs/2607.18804 · PDF

  52. 52

    QScheduler: Adaptive Gradient Sampling for Zeroth-Order On-Device Training on INT8 NPUs

    Victor Felipe Domingues Do Amaral, Pierre Demaj, Erwan Libessart, Laurent Folliot, Anthony Kolar, Philippe Bénabès

    cs.LG

    Zeroth-Order (ZO) optimization enables On-Device Learning (ODL) on NPU-equipped microcontrollers by estimating gradients through forward passes alone, bypassing the need for backpropagation primitives and reducing memory requirements. The number of gradient samples q critically affects training: insufficient samples produce noisy gradients that plateau early, while excessive samples consume more computational resources. However, finding an...

    arxiv.org/abs/2607.18802 · PDF

  53. 53

    PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects

    Dongkwan Kim, Yiming Gao, Yining Yang, Yang Shen

    cs.LG · q-bio.MN

    Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under realistic distribution shifts. We introduce PertReason, a knowledge-grounded benchmark and framework suite for cell-state--conditioned reasoning about perturbation effects. At its core, PertReasonQA is a benchmark that tests whether models can generate mechanistically faithful explanations while remaining robust to complex...

    arxiv.org/abs/2607.18777 · PDF

  54. 54

    Formulation-Level Auto-Tuning for QUBO-Based Machine Learning: A Case Study Across Multiple Quantum-Inspired Annealers

    Naoya Mizuki, Takahiro Katagiri, Daichi Mukunoki, Tetsuya Hoshino

    cs.LG · cs.PF

    This paper presents an Optuna-based formulation-level auto-tuning framework for support vector machines (SVMs) implemented on multiple quantum-inspired annealers. In an annealing-based SVM, continuous dual variables are discretized and converted into a quadratic unconstrained binary optimization (QUBO) model. This transformation introduces three coupled classes of parameters: representation parameters-the encoding base B and bit depth K-which...

    arxiv.org/abs/2607.18774 · PDF

  55. 55

    Relative Positions Generalize, Absolute Positions Memorize: An Implicit-Bias Account of Length Generalization in Attention

    Subham Singh, Ashutosh Mishra, Subha Raut

    cs.LG

    Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not. This is a robust empirical regularity, and the explanations offered for it so far are chiefly about expressivity, that is, about whether a length-generalizing solution exists. We give an optimization explanation. On a minimal fixed-offset retrieval task...

    arxiv.org/abs/2607.18759 · PDF

  56. 56

    Decafs: Disentangled Conditional adversarial Flows

    Anirudh jain, Sakshi Varshney, Samuel Kaski, Vikas Garg

    cs.LG

    Flow-based models have established state-of-the-art performance in generative modeling across domains, but are hard to interpret due to their complex latent embeddings. In particular, the entanglement of generative factors in the latent space hinders controlled generation. We circumvent this issue by appealing to a novel conditional generator based on Lie groups that disentangles an alternative latent space, which is aligned closely with the...

    arxiv.org/abs/2607.18755 · PDF

  57. 57

    Is EEG-to-Text Feasible in Real-World Scenarios? An In-Depth Analysis Using a Neuropsychology-Inspired Benchmark

    Zihan Zhang, Yu Bao, Xiao Ding, Tianyi Jiang, Kai Xiong

    cs.LG · cs.CL · cs.ET · q-bio.NC

    Translating brain signals into text could restore communication for people with severe paralysis, yet practically usable systems to date rely on invasive electrocorticography (ECoG). Electroencephalography (EEG) offers a non-invasive alternative, and EEG-to-text (EEG2Text) has been widely explored. Interestingly, however, EEG2Text models generally rely on teacher-forcing evaluation; without it, they fail to generate meaningful decoding. This...

    arxiv.org/abs/2607.18749 · PDF

  58. 58

    ConceptCF: Concept-based Counterfactuals for the Explainability of Time Series

    Annemarie Jutte, Faizan Ahmed, Jeroen Linssen, Maurice van Keulen

    cs.LG · cs.AI

    This paper proposes ConceptCF, a method for counterfactual generation that operates on human-interpretable concepts. In high-stakes domains such as healthcare and predictive maintenance, artificial intelligence models can increase efficiency and safety. Explainability is key to ensure these models rely on causal relationships rather than spurious correlations. Counterfactual explanations identify minimal modifications that would change a...

    arxiv.org/abs/2607.18748 · PDF

  59. 59

    Contraction-Gauge Preconditioning for Quantized Matrix Multiplication

    Piyush Sao, Narasinga Miniskar, Pedro Valero-Lara, Keita Teranishi, Sudip Seal

    cs.LG · cs.IT · math.NA

    We study low-precision computation of C=AB with both factors quantized. We derive an exact finite-dimensional identity for the expected squared product error under independent, zero-mean entrywise errors with known variance fields; it holds exactly for non-overloading subtractive dither and for independent stochastic rounding, and we empirically assess deterministic round-to-nearest (RTN). Using the product-preserving equivalence...

    arxiv.org/abs/2607.18745 · PDF

  60. 60

    Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

    Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang

    cs.LG · cs.CL

    Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs approximation error in finite-horizon bounds, whereas PPO clipping only gates sampled outward updates, acting as a sampled...

    arxiv.org/abs/2607.18722 · PDF

  61. 61

    Exposure-Based Reinforcement Learning to Rank

    Harrie Oosterhuis, Rolf Jagerman, Zhen Qin, Xuanhui Wang

    cs.LG · cs.IR

    Reinforcement learning (RL) methods for learning-to-rank (LTR) can optimize (almost) any ranking goal, e.g., from precision or discounted cumulative gain to fairness-of-exposure or ranking distillation. However, standard RL is ineffective and computationally costly due to the enormous action space in LTR settings. Existing methods reach computational efficiency through custom gradient computation algorithms, but they are very complex to...

    arxiv.org/abs/2607.18689 · PDF

  62. 62

    Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator

    Yuxiang Ji

    cs.LG · cs.SE

    Mined code corpora are abundant but uncontrolled: a snippet's semantics, surface "messiness," and difficulty are whatever the wild contained; there is no known-optimal reference to grade against; and any public sample may already sit in a model's training set. We present Spaghetti Architect, a tool that mints code datasets with the control such corpora lack. An anti-optimization transpiler maps a clean, language-agnostic JSON intermediate...

    arxiv.org/abs/2607.18642 · PDF

  63. 63

    Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs

    Seunghyun Lee, Dongyoon Han, Sangdoo Yun

    cs.LG · cs.CL

    Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal training); both pay a tax in adjacent-domain competence or over-refusal. We argue that the right operation is conditioning, not reduction: we show that hazardous knowledge can be retained in the model and behaviorally gated by a privileged control token. Our...

    arxiv.org/abs/2607.18639 · PDF

  64. 64

    Graph Neural Network-based Algorithm Selection for the Traveling Salesman Problem: A Systematic Study of Cost and Rank Losses under Distinct Budget Regimes

    Zhaoxuan Li, Jiale Yang, Yifei Lu, Mustafa Misir

    cs.LG

    Automated Algorithm Selection (AS) aims to improve problem-solving performance by selecting, for each problem instance, the most suitable algorithm from a predefined portfolio. This is particularly relevant to the Traveling Salesman Problem (TSP), where solver performance is strongly instance-dependent. We introduce GNNAS-TSP, a Graph Neural Network (GNN)-based AS framework that learns TSP instance representations directly from raw graph...

    arxiv.org/abs/2607.18632 · PDF

  65. 65

    BRIDGE: Bottleneck-Aware Regulator-Set Inference and Diagnosis for Cooperative Gene Regulatory Recovery

    Maryam Rahimimovassagh, Clayton Thomas Barham, Ivan Garibay, Niloofar Yousefi

    cs.LG

    Cooperative gene regulation often depends on groups of regulators acting jointly, but most gene regulatory network (GRN) inference methods output pairwise regulator-target rankings. We introduce Bottleneck-Aware Regulator-Set Inference and Diagnosis (BRIDGE), a framework for complete regulator-set recovery, and Targeted Recovery Attribution for Cooperative Evaluation (TRACE), a diagnostic suite that attributes failures to retrieval, set-level...

    arxiv.org/abs/2607.18602 · PDF

  66. 66

    A Self-Evolving Default Action for Cooperative Tasks with Continuous Action Space

    Shuangyao Huang

    cs.LG · cs.MA

    Counterfactual credit assignment has proven effective in multi-agent reinforcement learning (MARL) for discrete action spaces, yet its extension to continuous-action cooperative tasks remains challenging. Existing methods that approximate the counterfactual baseline via Monte Carlo sampling often introduce bias into policy gradients and fail to guarantee convergence to local optima, as the sampled actions may not have been sufficiently...

    arxiv.org/abs/2607.18597 · PDF

  67. 67

    Planning as Emergent Behavior in Reinforcement Learning with Relational Hidden States

    Armin Sommer

    cs.LG · cs.AI

    Reinforcement learning is conventionally divided into model-based and model-free methods. In this taxonomy, model-based methods perform lookahead planning over a learned world model, whereas model-free methods learn a reactive state-action mapping. Recent work, however, has shown that planning can emerge from model-free reinforcement learning alone. The conditions under which this behavior emerges from a pure reward-maximization objective...

    arxiv.org/abs/2607.18589 · PDF

  68. 68

    On the Diverse Dynamical Behaviors Arising in Deep Linear Transformers

    Sixu Li, Thomas Jacob Maranzatto, Jan Peszek, Trevor Teolis, Semih Akkoc, Konstantin Riedl, Sennur Ulukus, Nicolás...

    cs.LG · math.DS

    We study the inference-time behavior of deep linear encoder-only transformers through the lens of interacting particle systems. In this perspective, tokens are modeled as particles that interact dynamically through successive linear self-attention layers. We show that in embedding dimension two, for any key, query, and value matrices, the dynamics can be reformulated as a generalized Kuramoto-type model with pure second-harmonic coupling....

    arxiv.org/abs/2607.18584 · PDF

  69. 69

    Conditioned Direct Feedback Alignment via Activity and Error Geometry

    Houman Safaai, Varun Reddy, Bernardo L. Sabatini

    cs.LG · cs.NE · q-bio.NC

    Direct feedback alignment (DFA) trains hidden layers with fixed random projections of the output error, avoiding the transposed-weight backward pass of backpropagation (BP). We study a failure mode of DFA training that is distinct from feedback quality: the local weight update is calculated by an outer product, so anisotropy can enter through either its presynaptic-activity factor or its local-error factor. Our analyses with controlled...

    arxiv.org/abs/2607.18574 · PDF

  70. 70

    AMICA-Python: Adaptive Mixture Independent Component Analysis with Anderson Acceleration

    Scott Huberty, Christian O'Reilly

    cs.LG

    Adaptive Mixture Independent Component Analysis (AMICA) is widely used in EEG research and has long been associated with strong empirical performance for blind source separation. Despite its impact, practical use has historically depended on a single Fortran implementation, accessed via the EEGLAB toolbox for MATLAB, limiting its accessibility for analytical pipelines not designed within the MATLAB ecosystem. Here we present AMICA-Python, a...

    arxiv.org/abs/2607.18568 · PDF

  71. 71

    Robust Multi-View Classification under Noisy Supervision via Global Anchor Consensus

    Yuliang Yang, Hongzhe Zhang, Huiru Wang

    cs.LG · cs.CV

    In recent years, multi-view learning has attracted increasing attention, as it integrates the complementary information of heterogeneous views. Most existing multi-view classification methods rely on accurate annotations to guarantee performance. However, noisy labels are ubiquitous in practice due to imperfect annotation, and the refinement signals that existing methods derive from models trained on such noisy supervision can gradually lose...

    arxiv.org/abs/2607.18561 · PDF

  72. 72

    Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary

    Jan Kirin

    cs.LG · cs.AI

    Can a language model read the quality of ongoing computation, and can an external intervention turn that readout into better outcomes? We test both questions in a frozen 2.6B looped transformer, Ouro-RLTT. On GSM8K, a strict pre-answer probe excludes the answer region and gold value yet predicts eventual success: hidden states plus length and log-probability shortcuts reach AUROC 0.797, versus 0.731 for the shortcuts alone (incremental...

    arxiv.org/abs/2607.18553 · PDF

  73. 73

    Censoring-Aware In-Context Learning for Generalized Supplier Lead Time Estimation in Supply Chain Planning

    Christopher Wang, Sebastien Ouellet, Behrouz Haji Soleimani, Ali Etemad

    cs.LG · cs.AI

    Supplier lead time forecasting is a central input to material requirements planning, inventory optimization, and supply chain risk management. However, many industrial lead time datasets are naturally right-censored: at the time forecasts are required, some orders have not yet arrived. Standard regression and classification approaches discard this information, while conventional survival models require task-specific modeling. We propose...

    arxiv.org/abs/2607.18530 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.