cs.LG · 2026-06-27 · No. 36

Machine Learning, 2026-06-27.

58 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

58 entries
  1. 01

    Reinforcement Learning without Ground-Truth Solutions can Improve LLMs

    Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang, Xunpeng Huang, Kun Zhou, Tongtong Liang, Zhewei Yao, Yi-An Ma, Yuxiong He

    cs.LG

    Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown. We introduce a \textbf{R}anking-\textbf{i}nduced \textbf{VER}ifiable framework (RiVER) that trains LLMs on score-based optimization tasks without ground-truth solutions, using deterministic execution feedback as continuous-valued...

    arxiv.org/abs/2606.27369 · PDF

  2. 02

    Autoregressive Boltzmann Generators

    Danyal Rehman, Charlie B. Tan, Yoshua Bengio, Avishek Joey Bose, Alexander Tong

    cs.LG · cs.AI

    Efficient sampling of molecular systems at thermodynamic equilibrium is a hallmark challenge in statistical physics. This challenge has driven the development of Boltzmann Generators (BGs), which allow rapid generation of uncorrelated equilibrium samples by combining a generative model with exact likelihoods and an importance sampling correction. However, modern BGs predominantly rely on normalizing flows (NFs), which either suffer from...

    arxiv.org/abs/2606.27361 · PDF

  3. 03

    Error-Conditioned Neural Solvers

    Haina Jiang, Liam Wang, Peng-Chen Chen, Min Seop Kwak, Seungryong Kim, Brian Bell, Jeong Joon Park

    cs.LG · cs.AI · cs.CV · math.NA

    Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat solving as a purely statistical task: once trained, they struggle to correct their own constraint violations and extrapolate beyond the training distribution. Recent hybrid methods promote physical correctness by targeting the PDE residual via gradient descent or Gauss--Newton steps, but inherit the compute cost and instability...

    arxiv.org/abs/2606.27354 · PDF

  4. 04

    Hallucination in World Models is Predictable and Preventable

    Nicklas Hansen, Xiaolong Wang

    cs.LG · cs.CV · cs.RO

    Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fluent while drifting from the ground-truth dynamics. We hypothesize that hallucination concentrates in low-coverage regions of the state-action space, where lightweight data-centric signals can both detect it and guide mitigation. To test this, we introduce MMBench2, a 427-hour, 210-task dataset...

    arxiv.org/abs/2606.27326 · PDF

  5. 05

    Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

    Nathanaël Jacquier, Maria Vakalopoulou, Mahdi S. Hosseini

    cs.LG · cs.AI

    Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation models, decomposing their polysemantic activations into a larger set of sparse, more monosemantic features. The Top-$k$ SAE, a now-standard variant, enforces sparsity architecturally through its activation function, retaining only the $k$ most active latents per input. Because it was designed precisely to avoid the $\ell_1$ penalty...

    arxiv.org/abs/2606.27321 · PDF

  6. 06

    Blackwell Approachability and Gradient Equilibrium are Equivalent

    Brian W. Lee, Nika Haghtalab, Michael I. Jordan, Ryan J. Tibshirani

    cs.LG

    Gradient equilibrium (GEQ) is a recently introduced online optimization framework that generalizes first-order stationarity from offline optimization and abstracts problems like online conformal prediction. While GEQ has curious similarities with known online learning frameworks, namely regret minimization, prior work has shown that GEQ error and regret are incomparable objectives, leaving open a precise understanding of how GEQ fits into the...

    arxiv.org/abs/2606.27315 · PDF

  7. 07

    A Multi-Fidelity Convolutional Autoencoder-Transfer Learning Framework for Guided-Wave-Based Damage Diagnosis Using Large Simulated and Limited Experimental Datasets

    Santosh Kapuria, Abhishek

    cs.LG

    Guided wave-based structural health monitoring (GWSHM) with onboard transducers offers significant potential for the early diagnosis of damage in engineering structures. However, the practical deployment of deep learning models is often hindered by the limited availability of labelled experimental data and the high computational cost of generating large-scale high-fidelity simulation datasets. This study presents a multifidelity transfer...

    arxiv.org/abs/2606.27304 · PDF

  8. 08

    Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search

    Ping Liu, Qianqi Shen, Jianqiang Shen, Wenqiong Liu, Rajat Arora, Yunxiang Ren, Chunnan Yao, Dan Xu, Baofen Zheng,...

    cs.LG

    Job-search platforms rely on low-bandwidth query interfaces that often fail to capture the high-dimensional complexity of candidate profiles. We present an end-to-end RLAIF (Reinforcement Learning from AI Feedback) framework to generate \emph{portable} job search queries, terms that abstract away seeker-specific identifiers while preserving generalizable qualifications. This task introduces a highly adversarial reward surface where policy...

    arxiv.org/abs/2606.27291 · PDF

  9. 09

    Recovering Governing Equations from Solution Data: Identifiability Bounds for Linear and Nonlinear ODEs

    Yang Pan, Helmut Bölcskei

    cs.LG · cs.IT · math.CA · math.DS

    Learning governing equations from observed solution data is a fundamental challenge in scientific machine learning \cite{bruntonDiscoveringGoverningEquations2016,kovachkiNeuralOperatorLearning2023,longPDENetLearningPDEs2018,rudyDatadrivenDiscoveryPartial2017,raonicConvolutionalNeuralOperators2023}, yet the theoretical conditions under which a ground-truth ODE can be uniquely and stably identified from multiple solution observations remain...

    arxiv.org/abs/2606.27285 · PDF

  10. 10

    How Good Can Linear Models Be for Time-Series Forecasting?

    Lang Huang, Jinglue Xu, Luke Darlow

    cs.LG

    Time-series forecasting research has been moving steadily toward larger architectures, from specialized transformers to general-purpose foundation models, on the assumption that capacity is what unlocks accuracy. We take the opposite position: most of the gap can be closed at far lower cost by tuning preprocessing rather than scaling models. We use Ridge regression as the testbed, since it has a closed-form solution and interpretable weights,...

    arxiv.org/abs/2606.27282 · PDF

  11. 11

    BetXplain: An Explanation-Annotated Dataset for Detecting Manipulative Betting Advertisements on Social Media

    MSVPJ Sathvik, Parmitha Vangapadu, Nishit Rane, Sathwik Narkedimilli, Mark Lee, Akrati Saxena

    cs.LG

    The promotion of betting applications on social media platforms has increased significantly in recent years. Many of these advertisements use persuasive techniques that may mislead users, encourage risky behavior, and potentially influence users' mental well-being. However, research on the automated detection of manipulative and deceptive betting advertisements remains limited due to the lack of publicly available annotated datasets. In this...

    arxiv.org/abs/2606.27274 · PDF

  12. 12

    RSPC: A Benchmark for Modeling Stress and Psychiatric Conditions in Digitally Mediated Relationships using Psychiatrist Annotations

    Parmitha Vangapandu, Sai Ganesh Mokkapati, Sathwik Narkedimilli, MSVPJ Sathvik, Timothy Liu, Simon See, Johannes C....

    cs.LG

    In NLP, mental health conditions are often modeled as isolated phenomena, without interpersonal context. We use Reddit posts about long-distance relationships to capture both mental health distress and associated relational triggers. We introduce the Relational Stress and Psychiatry Corpus (RSPC) containing 1,799 Reddit posts annotated by psychiatrists for diagnostic categories, including the most prevalent mood disorders (anxiety and...

    arxiv.org/abs/2606.27247 · PDF

  13. 13

    Effective Covariance Dynamics in Solvable High-Dimensional GANs

    Andrew Bond, Zafer Doğan

    cs.LG

    We study a solvable high-dimensional model of generative adversarial network (GAN) training in which a linear generator learns a low-dimensional subspace from data with structured latent covariance. Prior solvable GAN analyses assume unconditional signals with diagonal latent covariance; we extend the multi-feature discriminator setting to class-dependent, correlated, and non-zero-mean latent structure. For the quadratic energy discriminator,...

    arxiv.org/abs/2606.27246 · PDF

  14. 14

    The Geometry of Updates: Fisher Alignment at Vocabulary Scale

    John Sweeney

    cs.LG · cs.CL · stat.ML

    Training-free source selection for LLM families with shared vocabularies arises in scientific string domains such as SMILES, protein, and genomic sequences, where candidate corpora share a tokenizer but differ in prediction targets. This creates an activation-dark regime: representation-similarity metrics can be uninformative without assumptions about label-conditioned error geometry, while classical update-geometry metrics are...

    arxiv.org/abs/2606.27242 · PDF

  15. 15

    Graph Neural Networks Applications Across Domains: All Insights You Need

    Abderaouf Bahi

    cs.LG

    Graph neural networks have moved from a niche representation-learning technique to the default model class wherever data carry relational structure. The interesting question is no longer whether message passing helps on a given dataset, but where graph structure earns its computational cost and where it does not. This survey organises the field around a single design space, derives the spectral and spatial formulations from shared first...

    arxiv.org/abs/2606.27202 · PDF

  16. 16

    Explaining Temporal Graph Neural Networks via Feature-induced Information Flow

    Ping Xiong, Thomas Schnake, Klaus-Robert Müller, Shinichi Nakajima

    cs.LG

    Event-based Temporal Graph Neural Networks (ETGNNs) have demonstrated strong performance across a wide range of applications, including social network analysis, epidemic tracing, recommender systems, and political event forecasting. However, their increasing complexity poses significant challenges for explainability. Existing explanation methods focus only on a subset of the information flow within ETGNNs, typically tracing contributions from...

    arxiv.org/abs/2606.27201 · PDF

  17. 17

    Automating Potential-based Reward Shaping with Vision Language Model Guidance

    Henrik Müller, Daniel Kudenko

    cs.LG · cs.AI · cs.RO

    Sparse rewards are inherently challenging for reinforcement learning agents as they lack intermediate feedback to guide exploration and to correctly attribute the sparse success rewards to relevant parts of the trajectory. Naive reward shaping can induce reward hacking, yielding policies that exploit auxiliary signals instead of solving the intended task. Potential-based reward shaping (PBRS) guarantees preservation of the optimal policy set,...

    arxiv.org/abs/2606.27180 · PDF

  18. 18

    RecallRisk-BERT: A Multi-Task Framework for Post-Report Medical Device Recall Triage

    Ali Semih Atalay, Sevgi Yigit-Sert

    cs.LG

    Medical device recalls are a critical regulatory mechanism for protecting patient safety. The growing volume of FDA recall records presents challenges in post-report recall triage, severity assessment, and root-cause interpretation. Existing studies mostly address recall occurrence prediction or root-cause analysis separately, while joint modeling of recall severity and root-cause categories has received limited attention. We develop an...

    arxiv.org/abs/2606.27174 · PDF

  19. 19

    Stochastic Gradient Optimization with Model-Assisted Sampling

    Jonne Pohjankukka, Jukka Heikkonen

    cs.LG · stat.ME

    This work addresses the problem of variance in stochastic gradient estimation for machine learning optimization. Deep learning relies on mini-batch methods such as stochastic gradient descent, which approximate full gradients but introduce noise, creating trade-offs between convergence stability, speed, and generalization. Existing methods, including variance reduction techniques (e.g., SVRG and SAG) and adaptive optimizers, aim to mitigate...

    arxiv.org/abs/2606.27171 · PDF

  20. 20

    fTNN: a tensor neural network for fractional PDEs

    Qingkui Ma, Hehu Xie, Xiaobo Yin

    cs.LG

    We develop the fTNN, a deterministic tensor neural network subspace method for problems involving the fractional Laplacian on bounded domains, taking the fractional Poisson equation and time-dependent fractional advection-diffusion equation as typical representatives. The work employs a geometry-adapted integration split featuring a spatially dependent near-field radius, which decomposes the fractional Laplacian into three contributions: a...

    arxiv.org/abs/2606.27140 · PDF

  21. 21

    Kolmogorov Arnold networks (KAN) for aerodynamic prediction: a comparison with MLPs and GNNs

    Miguel Jaraiz, Fermin Gutierrez, Pablo Yeste, Miguel Sánchez-Domínguez, Eusebio Valero, Gonzalo Rubio, Lucas Lacasa

    cs.LG · physics.data-an · physics.flu-dyn

    Kolmogorov Arnold networks (KAN) have recently been introduced as a (deep) neural network architecture whose trainable parameters adapt the activation functions, instead of the coefficients of the affine transformations at the core of traditional architectures such as deep multilayer perceptrons (MLPs). This architecture builds on the Kolmogorov-Arnold theorem, which endows it with universal approximation properties. While the advent of KANs...

    arxiv.org/abs/2606.27126 · PDF

  22. 22

    Cross-Head Attention Uplift Network with Inverse Propensity Score under Unobserved Confounding

    Haoran Zhang, Chuanpu Li, Yuxin Fu, Bin Tong, Guan Wang, Bo Zheng, Feng Zhou

    cs.LG

    Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discriminative power and debiasing under unobserved confounding scenarios. In this paper, we propose the Cross-Head Attention Uplift Network (CHAUN) and Robust Adversarial Inverse Propensity Score (RA-IPS) method to address these limitations. CHAUN employs shared feature embeddings and...

    arxiv.org/abs/2606.27114 · PDF

  23. 23

    Heavy-Ball Q-Learning with Residual Weighting Correction

    Donghwan Lee

    cs.LG · cs.AI

    This paper proposes a corrected heavy-ball Q-learning method for reinforcement learning (RL) and establishes its convergence. It also identifies conditions under which the method is theoretically guaranteed to converge faster than standard Q-learning. The same construction is then extended to Q-learning with linear function approximation, where analogous convergence and acceleration statements are derived. The analysis is based on a switched...

    arxiv.org/abs/2606.27112 · PDF

  24. 24

    Transformer-Based Classification of Bacterial Raman Spectra with LOOCV

    Jamile Mohammad Jafari, Thomas Bocklitz

    cs.LG

    Transformer-based models have recently attracted increasing attention for Raman spectral classification. In this study, a transformer-based approach was systematically evaluated using a nested leave-one-replicate-out cross-validation framework and compared with conventional machine-learning pipelines combining PCA or ICA with LDA, SVM, and Random Forest classifiers. A bacterial Raman dataset comprising 5,417 single-cell spectra from six...

    arxiv.org/abs/2606.27096 · PDF

  25. 25

    Data-Free Reservoir Features for Efficient Long-Horizon Cold-Start Continual Learning

    Augustinas Jučas, Yangchen Pan

    cs.LG · cs.AI

    Cold-start exemplar-free class-incremental learning requires learning a growing set of classes without replay, external pretraining, or a large initial task. Existing cold-start methods typically either train the backbone throughout the stream and compensate for semantic drift, or freeze a backbone after the first task, producing features biased toward the initial classes. These choices also create a computational tension: drift-compensation...

    arxiv.org/abs/2606.27095 · PDF

  26. 26

    Finding Stationary Points by Comparisons

    Helin Wang, Chenyi Zhang, Xiwen Tao, Yexin Zhang, Tongyang Li

    cs.LG · cs.DS · math.OC · quant-ph

    We study the problem of finding stationary points of non-convex functions when access to the objective is provided only through a comparison oracle that, given two points, outputs which has the larger function value. For a twice differentiable $f\colon\mathbb R^n\to\mathbb R$ with Lipschitz gradient and Hessian, we develop an algorithm that visits an $ε$-stationary point using $\widetilde O(n^2/ε^{1.5})$ queries. Our approach uses a...

    arxiv.org/abs/2606.27082 · PDF

  27. 27

    State Representation Matters in Deep Reinforcement Learning: Application to Energy Trading

    Jesper Klicks, Sander Vržina, Vincent François-Lavet

    cs.LG · cs.AI

    Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constraints. This makes the state representation given to a reinforcement learning agent an important design choice. We study this in HydroDam, a pumped-storage arbitrage environment, using a fixed Double DQN agent. The environment, action space, reward function, network, and training protocol are kept fixed; only...

    arxiv.org/abs/2606.27032 · PDF

  28. 28

    Symplectic Neural Networks for learning Generalized Hamiltonians

    Harsh Choudhary, Vyacheslav Kungurtsev, Chandan Gupta, Melvin Leok, Georgios Korpas

    cs.LG

    Hamiltonian Neural Networks (HNNs) integrate physical priors into neural models by learning a system's Hamiltonian, improving generalization and sample efficiency. Identifying the system Hamiltonian from noisy observations of state variables is a challenging task. For simulations to faithfully reflect the long-term behavior of Hamiltonian systems, especially energy conservation, it is essential to use symplectic integrators, which preserve...

    arxiv.org/abs/2606.27029 · PDF

  29. 29

    Just how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQA

    Eren Senoglu, Federico Toschi, Nicolo Brunello, Andrea Sassella, Mark James Carman

    cs.LG · cs.CL · cs.CV

    Multimodal large language models (MLLMs) applied to Medical Visual Question Answering (VQA) tend to produce overconfident outputs regardless of actual correctness, and existing verbalized confidence calibration methods, developed primarily for text only LLMs, do not account for the multimodal nature of medical image understanding. This work proposes a training based framework that finetunes MLLMs to improve their calibration using a composite...

    arxiv.org/abs/2606.27023 · PDF

  30. 30

    A Generalization Theory for JEPA-Based World Models

    Jingyi Cui, Qi Zhang, Hongwei Wen, Yisen Wang

    cs.LG

    Joint Embedding Predictive Architectures (JEPAs) have recently emerged as a promising paradigm for world modeling by learning predictive dynamics in a latent space rather than generating future observations at the input level. Despite their empirical success, the theoretical understanding of JEPA-based world models remains limited. In this paper, we develop the first generalization theory for JEPA-based world models. We formulate JEPA...

    arxiv.org/abs/2606.27014 · PDF

  31. 31

    Uncertainty quantification via conformal prediction in data assimilation

    Catherine George, Alireza Javanmardi, Tijana Janjić, Eyke Hüllermeier

    cs.LG

    Quantifying the evolution of uncertainty is critical to both probabilistic forecasting and data assimilation in numerical weather prediction. In this study, we investigate the applicability of conformal prediction (CP), a recent machine learning (ML) method, to quantify uncertainty in a controlled, idealized setting. We use the one dimensional modified shallow water model, designed to mimic the convective process. CP provides a set of...

    arxiv.org/abs/2606.27001 · PDF

  32. 32

    Decision-Aligned Evaluation of Uncertainty Quantification

    Annika Schneider, Tommy Rochussen, Joshua Stiller, Vincent Fortuin

    cs.LG · cs.AI · stat.ML

    Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions. We introduce decision-alignment, a criterion that reveals which evaluation metrics meaningfully align with downstream utilities. Applying this framework, we show that many widely used...

    arxiv.org/abs/2606.26990 · PDF

  33. 33

    GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning

    Ting Zhou, Zhenqing Ling, Yiyang Zhao, Ying Shen, Daoyuan Chen

    cs.LG · cs.AI

    Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inconsistency: within a batch, a small set of high-reward rollouts induces representation-space preference directions that sharply disagree with the batch majority, resulting in high-variance and destabilizing updates. We propose...

    arxiv.org/abs/2606.26917 · PDF

  34. 34

    Asymptotically Optimal Learning for Parametric Prophet Inequalities

    Jung-hun Kim, Anna Grebennikova, Vianney Perchet

    cs.LG · stat.ML

    We study learning in prophet inequalities with i.i.d. rewards drawn from an exponential-type parametric family with an unknown parameter $θ$, a class that includes exponential, Pareto, and bounded-support power-family distributions. We first characterize the optimal full-information asymptotic competitive ratio for this family. In the unbounded-support case, the limit is $ {\left(θ/({θ-c_+})\right)^{c_+/θ}}/ {Γ(1-c_+/θ)},$ while in the...

    arxiv.org/abs/2606.26893 · PDF

  35. 35

    Quantization in Federated Learning: Methods, Challenges and Future Directions

    Farwa Ikram, Dipanwita Thakur, Antonella Guzzo, Giancarlo Fortino

    cs.LG

    Federated Learning (FL) has become a foundational paradigm for privacy-preserving distributed intelligence, yet its scalability remains fundamentally constrained by communication bottlenecks, device heterogeneity, and the challenges of training under statistically non-IID data. Quantization is one of the most effective mechanisms for mitigating these limitations, reducing both uplink/downlink payloads and on-device computation. This paper...

    arxiv.org/abs/2606.26822 · PDF

  36. 36

    Reasoning Quality Emerges Early: Data Curation for Reasoning Models

    Hongyi Henry Jin, Wenhan Yang, Meysam Ghaffari, Carlos Morato, Baharan Mirzasoleiman

    cs.LG

    Supervised fine-tuning (SFT) on a small, high-quality set of long reasoning traces is an effective approach for eliciting strong reasoning capabilities in Large Language Models (LLMs). However, existing methods for curating high-quality SFT data rely heavily on strong reasoning models to filter examples based on diversity and difficulty, making the curation process costly while often yielding suboptimal data quality. In this work, we show...

    arxiv.org/abs/2606.26797 · PDF

  37. 37

    AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing

    Chennan Ma, Yanning Zhang, Siqi Hong, Xiuchong Wang, Fei Xiao, Keping Yang

    cs.LG · cs.AI · cs.CL

    Traditional dynamic pricing models in large-scale e-commerce suffer from limited interpretability, poor utilization of unstructured information, and misalignment with long-term business objectives such as cumulative Gross Merchandise Value (GMV), Return on Investment (ROI) and milestone achievement. We propose AIGP, a novel framework that leverages a Large Language Model (LLM) prompted with domain knowledge, structured data and textual...

    arxiv.org/abs/2606.26787 · PDF

  38. 38

    Reproducibility Study of "AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models"

    Ananth K S, Arya Hariharan

    cs.LG · cs.CL

    Fang et al. (2025) introduced a null-space constrained projection, named AlphaEdit, for locate-then-edit knowledge editing methods, theoretically guaranteeing that edits do not disrupt previously preserved knowledge, and reports substantial gains over existing editing methods on LLaMA3, GPT2-XL, and GPT-J. In this work, we present a reproducibility study of AlphaEdit, reproducing its reported results under the original experimental setup and...

    arxiv.org/abs/2606.26783 · PDF

  39. 39

    Escaping Iterative Parameter-Space Noise: Differentially Private Learning with a Hypernetwork

    Naoki Nishikawa, Shokichi Takakura, Satoshi Hasegawa

    cs.LG · stat.ML

    Differentially private (DP) training of neural networks is often hindered by the large amount of noise required by gradient-based methods such as DP-SGD, which repeatedly inject high-dimensional noise in parameter space throughout training. In this paper, we propose a new framework for DP learning that avoids iterative optimization in parameter space. Instead of updating the target model using privatized gradients, we employ a hypernetwork...

    arxiv.org/abs/2606.26772 · PDF

  40. 40

    Batch-Invariant Spectral Intelligence for Robust and Explainable Insect Authentication

    Majharulislam Babor, Giacomo Rossi, Annalisa Altavilla, Oliver Schlüter, Marina M. -C. Höhne

    cs.LG

    Edible insects offer an efficient source of alternative protein, requiring less land, water and emitting less greenhouse gas than conventional livestock. However, their successful integration into the food supply chain demands reliable species authentication to control allergen exposure, prevent adulteration, and meet regulatory standards. Near-infrared spectroscopy provides a rapid analytical tool, but its performance drops when applied to...

    arxiv.org/abs/2606.26757 · PDF

  41. 41

    Structure Before Collapse: Transient semantic geometry in next-token prediction

    Yize Zhao, Isabel Papadimitriou, Christos Thrampoulidis

    cs.LG · cs.CL

    Neural Collapse predicts that balanced one-hot classification pushes model representations to be equally far from each other; a symmetric configuration that depends only on the output label and ignores any semantic similarity in the inputs. This creates a puzzle: next-token prediction language models are trained predominantly (as context length increases) with one-hot labels: the same context is very unlikely to appear twice in training with...

    arxiv.org/abs/2606.26749 · PDF

  42. 42

    HyperDFlash: MHC-Aligned Block Speculative Decoding with Gated Residual Reduction

    Luxi Lin, Shuang Peng, Rui Ma, Junhao Hua, Shuwei Fan, Zhengda Qin, Qiang Wang, Hongjian Sun, Fangmin Chen, Songwei Liu

    cs.LG · cs.CL

    We present HyperDFlash, a block-parallel speculative decoding framework tailored to the novel multi-hyper-connection (MHC) architecture proposed by DeepSeek-V4. Despite the strong initial-token drafting performance of the native Multi-Token Prediction (MTP) module in DeepSeek-V4, its draft accuracy degrades sharply at later positions, as error accumulation from unverified intermediate tokens harms acceptance rates. Although the original...

    arxiv.org/abs/2606.26744 · PDF

  43. 43

    Algorithmic Foundations of Deep Learning: Complexity-Theoretic Rates and a Characterization of Universal Approximation

    Anastasis Kratsios, Simone Brugiapaglia, Bum Jun Kim, Gregory Cousins, Haitz Sáez de Ocáriz Borde

    cs.LG · cs.AI · cs.LO · math.NA

    Feedforward neural network (NN) expressivity is typically studied by emulating optimal basis-expansion schemes. While powerful, this perspective is incomplete: it primarily captures complexity through regularity, and therefore does not distinguish intuitively simple and complicated objects with comparable regularity, such as the square-root function and a typical Brownian path. The guiding message is that neural networks should be viewed not...

    arxiv.org/abs/2606.26705 · PDF

  44. 44

    PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs

    Muhammad Ahmed

    cs.LG

    Autoregressive large language model (LLM) serving is increasingly limited by key-value (KV) cache movement rather than dense matrix multiplication. Modern paged-attention systems reduce KV-cache fragmentation and mature kernels such as FlashInfer provide highly optimized native-paged decode attention. However, the best single-kernel implementation is not always the best serving schedule: low-active long-context decode can under-utilize...

    arxiv.org/abs/2606.26666 · PDF

  45. 45

    Zero-Shot Size Transfer for Neural ODEs on Sparse Random Graphs: Graphon Limits and Adjoint Convergence

    Mingsong Yan, Zhida Wang, Sui Tang

    cs.LG · cs.AI · math.DS · math.NA · math.OC

    Graph Neural Differential Equations (GNDEs) model continuous-time graph dynamics by parameterizing Neural ODE velocity fields with Graph Neural Networks. Their local, size-independent filters suggest a zero-shot size-transfer principle: train on a small graph and deploy on larger, similar graphs without retraining. We develop a quantitative theory for this principle on sparse random graphs sampled from graphons. We consider Graphon Neural...

    arxiv.org/abs/2606.26662 · PDF

  46. 46

    Target-Aware Bandit Allocation for Scalable Surrogate Optimization in Chemical Space

    Mohammad Haddadnia, Yuvan Chali, Abhilash Jayaraj, Constance Kraay, Joana Reis, Felix Strieth-Kalthoff, Haribabu Arthanari

    cs.LG

    Identifying high-utility candidates from massive discrete spaces under expensive evaluations is a recurring challenge across the sciences, with structure-based drug discovery as a prominent example. While surrogate-based optimization can increase sample efficiency by reducing the number of expensive evaluations, modern molecular libraries have reached billions to trillions of compounds, making full-library surrogate inference itself a major...

    arxiv.org/abs/2606.26657 · PDF

  47. 47

    From Weights to Features: SAE-Guided Activation Regularization for LLM Continual Learning

    Evan Ning, Wei Xue, Dong Lou, Yike Guo

    cs.LG · cs.CL

    Weight-space regularization methods such as Elastic Weight Consolidation (EWC) are the standard approach to catastrophic forgetting in continual learning. However, those methods tend to underperform when applied to large language models. We argue that such underperformance can be partly explained by the ``polysemantic'' nature of large language models: per-weight importance estimates utilized by EWC-style regularization are too coarse and...

    arxiv.org/abs/2606.26629 · PDF

  48. 48

    Discovering Millions of Interpretable Features with Sparse Autoencoders

    XinYang He, Wei Wang, Bing Zhao, Xuan Ren, WenBo Li, WeiXu Qiao, Hu Wei, Lin Qu

    cs.LG · cs.AI

    Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs is computationally expensive, and available open-source SAE models remain limited. In this work, we introduce \textbf{Qwen3-Instruct SAE}, a comprehensive suite of SAEs trained on the Qwen3 instruction-tuned model family, covering Qwen3-1.7B, Qwen3-4B, and Qwen3-8B....

    arxiv.org/abs/2606.26620 · PDF

  49. 49

    Sketched Linear Contrastive Learning: Approximation, Optimization, and Statistical Scaling

    Ziyan Chen, Zhongzhu Zhou, Ding-Xuan Zhou

    cs.LG

    Scaling laws describe how learning performance varies with model size, data size, and compute. While recent theoretical work has established scaling laws for sketched linear regression, much less is understood for contrastive representation learning. In this paper, we study a sketched linear model for contrastive learning under a paired Gaussian latent-variable setup. The learner observes only sketched views of two correlated variables and...

    arxiv.org/abs/2606.26617 · PDF

  50. 50

    Empirical Software Engineering TerraProbe: A Layered-Oracle Framework for Detecting Deceptive Fixes in LLM-Assisted Terraform

    Manar Alsaid, Chimdumebi Nebolisa, Faris Abbas

    cs.LG · cs.CR

    Security misconfigurations in Terraform Infrastructure-as-Code are a growing risk in cloud deployments, and large language models are increasingly used as automated repair agents. Existing evaluations often treat a repair as successful when the targeted static-analysis finding disappears, without checking planning validity, behavioral change, or security intent. This paper presents TerraProbe, a five-layer oracle framework for evaluating...

    arxiv.org/abs/2606.26590 · PDF

  51. 51

    SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference

    Haoqian Meng, Yilun Luo, Yafei Zhao, Wenyuan Liu, Huaqing Zheng, Xindian Ma, Peng Zhang

    cs.LG · cs.AI

    Low-bit floating-point formats and semi-structured sparsity are increasingly supported by modern accelerators, yet combining them for LLM activation compression remains challenging: activations contain input-dependent outliers that dominate block scales in FP4 quantization, and directly applying N:M sparsity masks discards moderate values, coupling sparsification loss with quantization error. We introduce SharQ, a training-free inference...

    arxiv.org/abs/2606.26587 · PDF

  52. 52

    Revisiting Action Factorization for Complex Action Spaces

    Timothy Flavin, Sandip Sen

    cs.LG

    Many real-world control problems involve hybrid discrete-continuous action spaces. For example, steering and signaling in autonomous driving, and aiming and firing in robotics or video-games. Despite real-world hybrid factorization and reinforcement learning framework support for complex action spaces (e.g., Gymnasium, PettingZoo, TorchRL, SeedRL, Mujoco, etc), the default environments within those frameworks often implement uniform action...

    arxiv.org/abs/2606.26574 · PDF

  53. 53

    Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication

    Jerome Marston, Tino Kreutzer, Salomé Garnier, Ella Boone, Phuong N Pham, Patrick Vinck

    cs.LG · cs.CY

    Data from affected populations are crucial for informing humanitarian response, but their value depends on timely and consistent interpretation of nuanced accounts of need. Humanitarian organizations often lack the staff, time, and specialist expertise required to analyze this information at scale. Large language models (LLMs) may expand this capacity, but their reliability for coding qualitative humanitarian data has not been directly...

    arxiv.org/abs/2606.26541 · PDF

  54. 54

    CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry

    Huzama Ahmad, Cao Viet Hai Nam, Se-Young Yun

    cs.LG · cs.AI

    Deep Transformers are composed of uniformly stacked residual blocks, yet their deepest layers often add little value. We present two efficiency methods that exploit this asymmetry. CascadeFormer tapers width with depth to match the uneven information flow across layers, achieving comparable perplexity to a uniform baseline at the same training budget while reducing latency by 8.6% and increasing throughput by 9.4%. CascadeFlow Pruning removes...

    arxiv.org/abs/2606.26538 · PDF

  55. 55

    Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy

    Wenjie Huang, Yang Li, Jingjia Teng, Mingwei Jin, Kai Song, Yougang Bian, Yongfu Li, Qisong Yang, Helai Huang

    cs.LG

    Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient autonomous highway lane changing decision-making. Existing methods frequently encounter transfer mismatch induced by distribution shifts between source and target domains, leading to training oscillation and performance decline. Besides, target domain adaptation depends on exploratory interactions,...

    arxiv.org/abs/2606.26527 · PDF

  56. 56

    Theory-Scale Auto-Formalization of Logics for Computer Science

    Yuming Feng, Frederick Pu, One An, Osbert Bastani, Li Zhang, Jiani Huang, Xujie Si, Ziyang Li

    cs.LG · cs.LO · cs.PL

    Auto-formalization is critical for scalable formal verification, but existing progress largely focuses on isolated statements, while theory-scale auto-formalization, which coherently translates hundreds of interdependent definitions, lemmas, and theorems, remains open due to challenges in consistency, faithfulness, scalability, and correctness. In this paper, we introduce LCS-Bench, a stand-alone, theory-scale benchmark based on Logics for...

    arxiv.org/abs/2606.26525 · PDF

  57. 57

    Multipath Adaptive Gated Bottleneck Latent ODE with Raman Data Fusion for Cell Culture Process Forecasting

    Johnny Peng, Thanh Tung Khuat, Ellen Otte, Katarzyna Musial, Bogdan Gabrys

    cs.LG · cs.AI · cs.CE

    Mammalian cell-culture processes underpin the manufacture of many biopharmaceuticals, yet keeping a run on track is hard: critical process parameters drift over days, and an off-specification trend is often confirmed too late to intervene. Early-stage, multi-day forecasts could enable timely adjustment of feeding, sampling, and control, but bioprocess forecasting is challenging because measurements are sparse and irregularly sampled,...

    arxiv.org/abs/2606.26520 · PDF

  58. 58

    Learning Probabilistic Filters with Strictly Proper Scoring Rules

    Eviatar Bach, Ricardo Baptista, Jochen Bröcker, Bohan Chen, Andrew Stuart

    cs.LG · math.DS · stat.ML

    Bayesian filtering of partially and noisily observed dynamical systems seeks to infer the evolving conditional distribution of the state of a dynamical system, given observations, in an online fashion. This Bayesian filtering distribution is the natural object for uncertainty quantification, but it is rarely available as a supervised learning target. However, one can often use the forecast model to generate synthetic system trajectories,...

    arxiv.org/abs/2606.26497 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.