cs.LG · 2026-06-15 · No. 24

Machine Learning, 2026-06-15.

72 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

72 entries
  1. 01

    Persona-Pruner: Sculpting Lightweight Models for Role-Playing

    Jinsu Kim, Jihoon Tack, Noah Lee, Jongheon Jeong

    cs.LG · cs.CL

    Language Models (LMs) have shown remarkable potential as role-playing chatbots, delivering consistent, stylized interactions when given a specification of a character or user persona. However, applying these capabilities to real-world applications (e.g., ecosystems with numerous NPCs interacting simultaneously) exposes a critical inefficiency due to the excessive computational cost. In this paper, we question the necessity of dedicating a...

    arxiv.org/abs/2606.14695 · PDF

  2. 02

    A Complexity Measure for Active Learning in Multi-group Mean Estimation

    Abdellah Aznag, Rachel Cummings, Adam N. Elmachtoub

    cs.LG · cs.IT

    We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}σ_k^2/n_k$, where $σ_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled. We develop a local minimax framework and prove the first general...

    arxiv.org/abs/2606.14690 · PDF

  3. 03

    Flood and Harvest: The Provable Necessity of Trivia for Generating Valuable Mathematics via the Lens of Language Generation in the Limit

    Xiaoyu Li, Andi Han, Dai Shi, Zheng Gao, Jiaojiao Jiang, Junbin Gao

    cs.LG · cs.AI · cs.CL · cs.DS

    AI systems coupled to proof assistants now generate formal mathematics at scale, and the gap between what a checker can verify and what a mathematician would value has become the binding constraint. We model the generation of valuable mathematics as nested language generation in the limit: a verifiable formal language $F$, accessed through a membership oracle (the proof checker), contains an unknown valuable language $H \in \mathcal{H}$...

    arxiv.org/abs/2606.14688 · PDF

  4. 04

    Optimal Hidden-Target Learning for Online Inventory Optimization on General Convex Sets

    Anthony Pineci, Yunzong Xu

    cs.LG · eess.SY · math.OC · stat.ML

    Online inventory optimization (OIO) is online convex optimization with physical memory: inventory carryover makes the feasible action set depend on the past. A natural principle, used in stochastic inventory learning and recently in OIO under a single linear capacity constraint, is to maintain a hidden target chosen by an online learner and implement its projection onto the currently feasible order-up-to set. We prove that this simple...

    arxiv.org/abs/2606.14679 · PDF

  5. 05

    Compressed Computation is (probably) not Computation in Superposition

    Jai Bhagat, Sara Molas-Medina, Giorgi Giglemiani, Stefan Heimersheim

    cs.LG

    We study whether the Compressed Computation (CC) toy model (Braun et al., 2025) is an instance of computation in superposition. The CC model appears to compute 100 ReLU functions with just 50 neurons, achieving a better loss than expected from only representing 50 ReLU functions. We show that the model mixes inputs via its noisy residual stream, corresponding to an unintended mixing matrix in the labels. Splitting the training objective into...

    arxiv.org/abs/2606.14673 · PDF

  6. 06

    When to Write and When to Suppress: Route-Specialized Dual Adapters for Memory-Assisted Knowledge Editing

    Yining Huang

    cs.LG

    Knowledge editing systems must update selected facts while preserving nearby but irrelevant behavior. This paper studies this problem in a memory-assisted setting where an edit memory is retrieved at inference time and a parameter-efficient adapter corrects the model's object preference. We argue that the central design question is not only how to write an edit, but also when to suppress it. We introduce \method{}, a route-specialized...

    arxiv.org/abs/2606.14668 · PDF

  7. 07

    Beyond task performance: Decoding bioacoustic embeddings with speech features

    Ines Nolasco, Jules Cauzinille, Marius Miron, Gagan Narula, Milad Alizadeh, Emmanuel Fernandez, Matthieu Geist,...

    cs.LG · cs.SD

    Pretrained audio embeddings are standard in bioacoustics, yet little is known about which acoustic features these models encode, nor which are useful for a given task. This hinders transparency and limits extension to rare species or data-scarce domains. Here we reveal which speech-like features are encoded in bioacoustic representations. Using the 88~eGeMAPS features across six taxonomic groups, we apply linear and nonlinear regression...

    arxiv.org/abs/2606.14662 · PDF

  8. 08

    Graph Structured Combinatorial Semi-Bandit with Nonlinear Reward Associations through Separable Signals

    Christoph Bauschmann, Setareh Maghsudi

    cs.LG

    The identification of optimal structures within vast arrays of interconnected data necessitates significant sampling- and computational effort. Learning and leveraging underlying signal dependencies can improve efficiency and predictive capabilities considerably, but the ubiquity of nonlinear statistical relations amplifies the complexity of such undertakings. In this paper, we develop novel generic and adaptive strategies equipped with...

    arxiv.org/abs/2606.14650 · PDF

  9. 09

    Which Directions Matter? Sparse Design for Affine Robust Optimization

    Pedro Chumpitaz-Flores, My Duong, Juan S. Borrero, Kaixun Hua

    cs.LG · math.OC

    Robust machine learning and optimization rely on the uncertainty model choice. We investigate which uncertainty directions a model must cover when defined by a finite dictionary and a budget constraint. Selecting a subset forms an atomic uncertainty set with a closed form support function, yielding tractable robust programs for affine objectives. We propose a data driven selection rule based on a coverage objective over evaluation directions,...

    arxiv.org/abs/2606.14648 · PDF

  10. 10

    Online Convex Optimization with Sublinear Noisy Probes

    Simone Di Gregorio, Anupam Gupta, Stefano Leonardi, Matteo Russo

    cs.LG · cs.DS

    We study Online Convex Optimization (OCO) over a convex set $K\subseteq \mathbb R^d$, where in each round $t$ the learner selects $x_t\in K$ and then observes a convex loss $f_t:K\to[0,1]$, with the goal of minimizing regret to the best fixed decision in hindsight. We introduce a unified probing model that generalizes two recent lines of work: sublinear best-expert queries in the experts setting, and pairwise (comparison-based) feedback...

    arxiv.org/abs/2606.14640 · PDF

  11. 11

    Graph Diffusion Residuals for Control-Function Instrumental Variables

    Rui Wu, Zongyuan Chen, Hong Xie, Defu Lian, Enhong Chen

    cs.LG

    Control-function instrumental variable estimators need a first-stage residual, not merely a first-stage prediction. High-capacity first stages can interpolate treatment and leave too little residual information for the outcome equation. We study Adaptive Anisotropic Instrumental Heat Flow (A-IHF), a deterministic graph-diffusion residual extractor for flexible control functions. A-IHF treats treatment as a signal on a graph of first-stage...

    arxiv.org/abs/2606.14636 · PDF

  12. 12

    Neither Parallel Nor Sequential: How DiffusionGemma Actually Commits Tokens

    Ali Asaria, Tony Salomone, Deep Gandhi

    cs.LG

    Open diffusion language models are marketed as parallel, non-autoregressive decoders, yet the order in which a shipped checkpoint actually commits its tokens is almost never measured. We instrument DiffusionGemma 26B, a masked discrete-diffusion mixture-of-experts model built on Gemma 4, hooking its sampler's accept step to record which canvas positions commit, when, and at what confidence. Across a 686-prompt, six-regime probe suite we find...

    arxiv.org/abs/2606.14620 · PDF

  13. 13

    Expert-Driven Survival Machines: Improving Stratification and Interpretability in Multiple Clinical Cohorts

    Farica Zhuang, Zixuan Wen, Christos Davatzikos, Li Shen

    cs.LG · cs.AI

    Survival prediction plays a central role for healthcare providers and clinical researchers. Accurate risk stratification enables early intervention and improved patient management. Most existing deep survival models learn one common feature representation for all patients, which may hide important differences between patient subgroups. In contrast, a Mixture-of-Experts (MoE) framework allows different parts of the model to focus on different...

    arxiv.org/abs/2606.14608 · PDF

  14. 14

    A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health

    Pavlos Nicolaou, Kleanthis Malialis, Artemis Kontou, Panayiotis Kolios

    cs.LG · cs.AI

    Wearable devices and smartphones generate rich behavioural time series that can support proactive health interventions, yet systematic comparisons of modern forecasting architectures for these data are lacking. In particular, it remains unclear how models generalise across populations, how different architectures respond to participant-level fine-tuning and how forecasting accuracy degrades across multi-day horizons. We benchmark six deep...

    arxiv.org/abs/2606.14604 · PDF

  15. 15

    A Statistical and Machine Learning Framework for Operational Threshold Detection and Deployable Dispatch Controller Development in Hydrogen Multi-Energy Systems

    Shadi Heenatigala, Hasanika Samarasinghe

    cs.LG · eess.SY · math.OC · stat.CO

    This study presents a statistical and machine learning framework for characterizing a hydrogen-based multi-energy system (H-MES) using one year of high-resolution operational data. Statistical analysis revealed a binary operation driven by renewable surplus, with solar irradiance explaining 45.7% of rank-based variance in hydrogen production, a large effect by conventional standards. Only high-irradiance periods triggered meaningful...

    arxiv.org/abs/2606.14601 · PDF

  16. 16

    Realizing Native INT8 Compute for Diffusion Transformers on Consumer GPUs: A Fused INT8 GEMM Kernel for Ideogram 4.0

    Ali Asaria, Tony Salomone, Deep Gandhi

    cs.LG

    Post-training INT8 (W8A8) quantization of diffusion transformers is widely deployed as a speed optimization, yet on consumer Ampere GPUs it is frequently slower than the FP8 and NF4 alternatives it is meant to beat. We trace this to a software artifact: the production "INT8" forward quantizes weights and activations only to immediately dequantize them back to bf16 and run a bf16 matrix multiply, never engaging the GPU's INT8 tensor cores, so...

    arxiv.org/abs/2606.14598 · PDF

  17. 17

    Zero-shot generalization of transformer neural operators to larger domains

    Armand de Villeroché, Sibo Cheng, Vincent Le Guen, Marc Bocquet, Rem-Sophia Mouradi, Patrick Armand, Alban Farchi,...

    cs.LG

    Transformer-based neural operators have shown remarkable performance for approximating solution operators of partial differential equations on complex geometries. However, existing approaches implicitly assume a fixed domain size, which limits their ability to generalize at inference. In this work, we investigate domain extension, namely zero-shot inference on spatial domains that are significantly larger than those encountered during...

    arxiv.org/abs/2606.14597 · PDF

  18. 18

    CARE: Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation

    Guanyu Liu, Weiyi Kong, Zeyu Wang, Boer Zhang, Baiqing Li, Peiyu Zhang, Tianyu Shi

    cs.LG · cs.AI

    Granting LLMs direct control over costly, irreversible scientific experiments leads to unsafe exploration and unstable performance, but discarding LLM creativity entirely sacrifices significant optimization potential. We introduce CARE (Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation), an auditable controller for high-throughput experimentation (HTE) optimization that keeps a non-LLM...

    arxiv.org/abs/2606.14581 · PDF

  19. 19

    Provably Safe, Yet Scalable Reinforcement Learning

    Kai S. Yun, Zeyang Li, Navid Azizan

    cs.LG · cs.RO · eess.SY

    Safe reinforcement learning (RL) aims to learn policies that optimize rewards while satisfying constraints. Predominant approaches rely on soft-constrained policy optimization, which has achieved empirical success but does not provide formal safety guarantees for the learned policy. In contrast, methods with strict guarantees typically rely on explicit certificate functions, whose construction requires the direct synthesis and verification of...

    arxiv.org/abs/2606.14536 · PDF

  20. 20

    The Risk Shadow of Principal Component Analysis: When 99.9999% Variance Preservation Causes Catastrophic Decision Errors

    Hamidou Tembine

    cs.LG · cs.GT

    Principal Component Analysis (PCA) preserves variance, not the information needed to detect rare catastrophic events. This paper proves the existence of a {\it Risk Shadow}: PCA can retain over 99.9999 percent of total variance while completely erasing all signal about rare, high-impact failures. When this happens, even the best possible classifier operating on the PCA representation reduces to a constant predictor. The root cause is a...

    arxiv.org/abs/2606.14533 · PDF

  21. 21

    Code Correctness Signals in LLM Hidden States: Pre-Generation Probing and Repair Geometry

    Carlo Di Cicco

    cs.LG

    Large language models encode rich information in their hidden states. This work asks whether code correctness is legible in the hidden states of Qwen3-4B-Instruct-2507, before it generates and as it repairs a failed attempt, studied on 444 LiveCodeBench tasks. It reports two findings connected by a single confound-control tool: residualization. First, the correctness of the model's first-attempt code is linearly decodable from the...

    arxiv.org/abs/2606.14530 · PDF

  22. 22

    Behavioral Audit of Machine Unlearning Has a Privacy Cost

    Liou Tang, James Joshi, Ashish Kundu

    cs.LG

    The removal of learned data from Machine Learning models through Machine Unlearning (MU) has been widely studied; however, there has yet to be an agreed-upon scheme for auditing MU. Existing work has shown that a dishonest model owner can falsify evidence to avoid executing MU, while curious auditors (and adversaries) can infer the privacy-sensitive properties of the model and its training data even with limited access. Yet auditing of MU...

    arxiv.org/abs/2606.14518 · PDF

  23. 23

    PepALD: Macrocyclic Peptide Generation via Autoregressive Latent Diffusion

    Junming Zhang, Siyu Yi, Wei Ju, Zhonghui Gu

    cs.LG

    Macrocyclic peptides are promising therapeutic candidates for intracellular targets, but their design requires simultaneous control over non-natural monomer chemistry, ring topology, membrane permeability, and target binding. Existing SMILES- or HELM-string generative models either operate in long atom-level sequence spaces or treat monomers as symbolic tokens with limited chemical grounding. We introduce PepALD, an Autoregressive Latent...

    arxiv.org/abs/2606.14510 · PDF

  24. 24

    Recipe-Controlled Decoder Audit for Structural Knowledge-Graph Completion

    Xihang Shan, Ye Luo

    cs.LG

    We present a recipe-controlled decoder audit (RCDA) for structural transductive knowledge-graph completion (KGC). The audit asks a simple reporting question: before attributing gains to an encoder or training recipe, what changes when the decoder is swapped under the same recipe? Using ComplEx and DistMult as the primary controlled pair, with targeted RotatE/TransE spot-checks, we evaluate seven benchmarks. On five standard KGs,...

    arxiv.org/abs/2606.14492 · PDF

  25. 25

    EM-NeSy: Expectation Maximization for Neurosymbolic Learning

    Annegret Seibt, Luc De Raedt, Giuseppe Marra

    cs.LG

    Neurosymbolic (NeSy) models integrate neural networks and symbolic reasoning for robust and interpretable AI. State-of-the-art NeSy models require that the symbolic component is expressed in a differentiable way, often complicating the use of approximate inference. We propose EM-NeSy which casts probabilistic NeSy learning as an instance of the Expectation-Maximization (EM) algorithm. In the expectation step, we compute the posterior over the...

    arxiv.org/abs/2606.14463 · PDF

  26. 26

    Federated Learning for Feature Generalization with Convex Constraints

    Dongwon Kim, Donghee Kim, Sung Kuk Shyn, Kwangsu Kim

    cs.LG · stat.ML

    Federated learning (FL) often struggles with generalization due to heterogeneous client data. Local models are prone to overfitting their local data distributions, and even transferable features can be distorted during aggregation. To address these challenges, we propose FedCONST, an approach that adaptively modulates update magnitudes based on the parameter strength of the global model. This prevents over-emphasizing well-learned parameters...

    arxiv.org/abs/2606.14416 · PDF

  27. 27

    A theoretical model for task routing in mixture-of-expert transformers

    Yongli Xiang, Vinoth Nandakumar, Yunzhi Yao, Peike Li, Tongliang Liu

    cs.LG

    Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed. While task-expert specialization has been observed in empirical studies of frontier MoE transformer models, existing theoretical work analyzes this using continuous mixture models that cannot be used to model natural language effectively. An important open question is to \textit{theoretically explain task-expert specialization...

    arxiv.org/abs/2606.14398 · PDF

  28. 28

    Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

    Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel, Michal Zakrzewski, Sebastian Montagna, Damian Rynczak, Shreyansh...

    cs.LG

    As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular applications with relatively simple tasks and focus on a narrow set of capabilities while overlooking broader dimensions, resulting in saturated performance on modern agents and failing to probe their limitations. To this end, we...

    arxiv.org/abs/2606.14397 · PDF

  29. 29

    A Low-Rank Subspace Analysis of LLM Interventions

    Angira Sharma, Christian Schroeder de Witt, Philip Torr, Anisoara Calinescu, Jialin Yu

    cs.LG

    Interventions designed to modify a particular behavior in LLMs, such as refusal or sycophancy, often produce unintended changes in other behaviors. This lack of targeted control makes it difficult to design and implement reliable safety controls. To understand these side-effects, we introduce a diagnostic framework for analyzing interacting behaviors in LLMs. We model behaviors as low-rank subspaces in activation space, and study how...

    arxiv.org/abs/2606.14388 · PDF

  30. 30

    Discovery under Hypothesis Redundancy: A Geometric Theory of Discovery Bottlenecks

    Li Xia, Baoxun Wang

    cs.LG · cs.AI · q-fin.PM

    Scientific discovery saturates when new hypotheses cease to provide independent information, even if the nominal hypothesis space remains large. We study hybrid discovery systems that combine structured local search with LLM-generated non-local proposals and pose the Search Compression Hypothesis: non-local exploration helps only when three geometric conditions co-occur: spectral compression, orthogonal escape from the explored span, and...

    arxiv.org/abs/2606.14386 · PDF

  31. 31

    Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback

    Woohyeon Byeon, Jiwon Jeon, Jeonghye Kim, Youngchul Sung

    cs.LG · cs.CL

    We study multi-domain LLM training in which two models, each stronger in a different domain, co-evolve by tutoring each other through on-policy feedback. Unlike one-way distillation or single-model fine-tuning, our goal is mutual Pareto improvement: each model improves across domains without losing its original strength. To this end, we propose On-Policy Co-Distillation (OPCoD), where each student's self-distillation is conditioned on its own...

    arxiv.org/abs/2606.14368 · PDF

  32. 32

    SemPiper: Interactive Code Synthesis for Semantic Operators in Machine Learning Pipelines

    Olga Ovcharenko, Luciano Duarte, Sebastian Schelter

    cs.LG · cs.DB

    Machine learning (ML) pipelines require extensive data preparation, feature engineering, and integration across heterogeneous sources, making them tedious and error-prone to develop. While large language models (LLMs) have recently shown promise for assisting programming tasks, chat-based interfaces provide limited control over pipeline behavior and often produce code that is difficult to optimize or integrate into production systems. We...

    arxiv.org/abs/2606.14361 · PDF

  33. 33

    MUFFLe: Efficient Model Update Compression via Generalized Deduplication for Federated Learning

    Xiaobo Zhao, Daniel E. Lucani

    cs.LG

    Federated learning is well suited to edge environments but is often limited by the uplink cost of transmitting model updates. This Work-in-Progress paper presents MUFFLe, a communication-efficient update compression scheme that integrates generalized deduplication (GD) into the FedAvg pipeline. MUFFLe deduplicates repeated patterns across the update vector, yielding a fixed-rate, variable-count compression scheme. Preliminary experiments on...

    arxiv.org/abs/2606.14354 · PDF

  34. 34

    Can Deep Neural Networks Improve Compression of Very Large Scientific Data?

    Muhannad Alhumaidi, Guozhong Li, Spiros Skiadopoulos, Panos Kalnis

    cs.LG

    Error-bounded lossy compression is a fundamental technique for managing the rapidly growing volumes of scientific data produced by modern simulations and observational instruments. Most state-of-the-art-compressors follow a prediction-residual paradigm, where compression effectiveness depends on the quality of the predictor: more accurate predictions generate smaller residuals that are easier to compress. This observation raises a question:...

    arxiv.org/abs/2606.14353 · PDF

  35. 35

    When Language Representations Interact: Separability and Cross-Lingual Effects in LLMs

    Boris Marinov, Angira Sharma, Christian Schroeder de Witt, Philip Torr, Anisoara Calinescu, Jialin Yu

    cs.LG

    Large language models exhibit strong multilingual capabilities, however, their internal representations are difficult to interpret. Understanding these interactions is important for ensuring reliable behavior in multilingual systems. Recent work has shown that causal-geometric structure can explain how certain concepts are encoded as approximately linear and separable directions, but whether this framework extends to multilingual models,...

    arxiv.org/abs/2606.14347 · PDF

  36. 36

    Squeeze-Release: Iterative Pruning with Exact Structural Minimization

    Roman Denkin, Ida Akerholm, Prashant Singh, Ida-Maria Sintorn

    cs.LG · cs.AI

    Unstructured pruning produces sparse weight tensors, but the standard implementation keeps tensor shapes unchanged so the deployed model is no smaller than before pruning. We present an exact structural rewrite, which we call minimization, that converts a masked network into a smaller dense network with the same forward function up to floating-point rounding. The Squeeze-Release cycle iterates pruning and minimization with an intermediate...

    arxiv.org/abs/2606.14346 · PDF

  37. 37

    More with LESS -- Local Scene Representations for Tactile Imaging

    Zohar Rimon, Elisei Shafer, Tal Tepper, Daniel Kozin, Alon Malka, Roy Holland, Aviv Tamar

    cs.LG

    Tactile imaging seeks to reconstruct the internal structure of soft objects through touch sensing, with applications in medical diagnosis and robotic manipulation. Recent self-supervised learning approaches have shown promising results, but rely on global, unstructured representations and robot-controlled sensing, limiting generalization and practical use. We propose Local Encoder for Spatial Sensing (LESS), an object-centric tactile...

    arxiv.org/abs/2606.14344 · PDF

  38. 38

    Riemannian Metric Matching for Scalable Geometric Modeling of Distributions

    Jacob Bamberger, Adam Gosztolai, Pierre Vandergheynst, Michael Bronstein, Iolo Jones

    cs.LG · math.DG

    High-dimensional datasets often concentrate near low-dimensional structures, but estimating their geometry from samples typically relies on graphs and kernels that scale poorly with dataset size and dimension. We propose Riemannian metric matching: a denoising probabilistic framework for learning the Riemannian geometry of data using neural networks. Specifically, we learn the carré du champ operator, which, using diffusion geometry, gives us...

    arxiv.org/abs/2606.14334 · PDF

  39. 39

    Hierarchical ODE: Learning Continuous-Time Physical Prototypes for Early Link Failure Detection

    Jiaen Lv, Leran Qi, Shaowei Wang

    cs.LG · cs.AI

    Time series prototype learning is fundamentally challenged by observational ambiguity. Discrete architectures fail to resolve this, as they lack the capacity to decouple stochastic noise from continuous dynamics. Furthermore, rigid closed-set assumptions fail to capture unseen diversity. To address these limitations, we propose a hierarchical ordinary differential equation clustering network, which utilizes neural ordinary differential...

    arxiv.org/abs/2606.14284 · PDF

  40. 40

    DIFF-ERO: A Conformance-Aware Loss for Deep Learning in Process Mining

    Johannes De Smedt, Jari Peeperkorn, Artem Polyvyanyy, Jochen De Weerdt

    cs.LG · cs.AI

    Deep learning has driven many recent advances in process analytics, especially for predictive and prescriptive monitoring. However, standard objectives such as cross-entropy optimize local next-step likelihoods and only implicitly capture control-flow structure. As a result, models can achieve high token-level accuracy while permitting imprecise global behaviour. We introduce DIFF-ERO, a conformance-aware loss function for deep learning...

    arxiv.org/abs/2606.14283 · PDF

  41. 41

    Beyond a Single Explanation of the Adam--SGD Gap

    Chenxiang Zhang, Rustem Islamov, Enea Monzio Compagnoni, Jun Pang, Aurelien Lucchi, Antonio Orvieto

    cs.LG

    Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties. Yet these explanations are often studied in isolation, leaving their relative importance unclear. In this work, we revisit these hypotheses through a controlled empirical study across vision, language, genomics, and graph tasks, spanning modern and classical...

    arxiv.org/abs/2606.14259 · PDF

  42. 42

    Where Black-box Drug-Target Interaction Prediction Models Look: Cross-Method Explainability

    Ali Vefghi, Zahed Rahmati, Mohammad Akbari

    cs.LG

    Drug-target interaction (DTI) and affinity (DTA) predictors increasingly achieve strong benchmark scores, yet their internal use of sequence, fingerprint, and graph features often remains opaque. We present an interpretability audit of BridgeDPI architecture on three different datasets including Gao, Human, and C.elegans. This study combines gradient-based attributions -- integrated gradients, saliency, layer-wise relevance propagation,...

    arxiv.org/abs/2606.14245 · PDF

  43. 43

    Implicit Variational Rejection Sampling

    Jian Xu, Shigui Li, Wei Chen, Jiacheng Li, Zhiqi Lin, Delu Zeng, Xinghao Ding, John Paisley, Qibin Zhao

    cs.LG

    Variational Inference (VI) is a fundamental inference technique in Bayesian machine learning for approximating complex posterior distributions. Traditional VI often relies on the mean-field factorization, which can inadequately capture true posterior complexity. Recent advancements have leveraged neural networks to model implicit distributions, offering increased flexibility. However, the practical constraints of neural network architectures...

    arxiv.org/abs/2606.14235 · PDF

  44. 44

    Learning the Context of Errors: Black-Box Online Adaptation of Time Series Foundation Models

    Xilin Dai, Yiding Liu, Hongjie Xia, Yifan Hu, Zewei Dong, Jiang-Ming Yang, Qiang Xu

    cs.LG

    The rapid evolution of Time Series Foundation Models (TSFMs) has advanced zero-shot forecasting across diverse domains. Inspired by the current form of Large Language Models, future TSFMs may be offered as commercialized, closed-source API services. However, many existing online adaptation methods still rely on white-box access for parameter fine-tuning or gradient backpropagation. This paradigm mismatch raises a question: In black-box online...

    arxiv.org/abs/2606.14222 · PDF

  45. 45

    Curvature-Informed Potential Energy Surface for Protein-Ligand Binding Affinity Prediction

    Peng-Fei Sun, Chuan-Xian Ren, Hong Yan

    cs.LG · q-bio.BM

    Accurate prediction of protein-ligand binding affinity is essential for structure-based drug discovery. Recent geometric deep learning methods have achieved promising performance by representing protein-ligand complexes as three-dimensional graphs. However, most existing approaches mainly rely on static interaction geometry from a single bound conformation, while neglecting molecular flexibility and binding-induced conformational changes. To...

    arxiv.org/abs/2606.14217 · PDF

  46. 46

    LapidaryEngine: Fully Conversational Crystal Generation

    Yusei Ito, Yuta Suzuki, Tomoya Murata, Masaki Adachi

    cs.LG

    The emergence of Large Language Models (LLMs) has inspired the vision of generating bespoke crystal materials directly from natural-language instructions, enabling users to design materials through intuitive, conversational interaction. Existing text-to-crystal generative models represent important early steps toward this goal, but they suffer from two critical limitations: (i) restricted input formats that require highly structured...

    arxiv.org/abs/2606.14215 · PDF

  47. 47

    Structured Noise Adaptation for Sequential Bayesian Filtering with Embedded Latent Transfer Operators

    Naichang Ke, Pongpisit Thanasutives, Yoshinobu Kawahara

    cs.LG · math.OC

    Kalman filters based on the Embedded Latent Transfer Operators (ELTO) emerge as novel statistical tools for sequential state estimation. However, a critical limitation stems from their use of simplified noise models, which fail to dynamically adapt to non-stationary processes. To address this limitation, we introduce an ELTO-based Bayesian filtering approach with a new structured parameterization for the filter's noise model. This...

    arxiv.org/abs/2606.14195 · PDF

  48. 48

    DRIVE: Distributional and Retrieval-Augmented Bidding with Value Evaluation

    Miduo Cui, Haochen Wang, Shangqin Mao, Xun Yang, Qianlong Xie, Xingxing Wang, Xuri Ge, Ying Zhou, Zhiwei Xu

    cs.LG

    Auto-bidding is a core component of real-time advertising systems, where decisions must optimize long-term performance under budget and cost constraints, while online exploration is prohibitively risky. Offline reinforcement learning and, more recently, Transformer-based sequence modeling have shown promise for learning bidding policies from logged data, but their unimodal and purely parametric formulations often collapse multiple effective...

    arxiv.org/abs/2606.14192 · PDF

  49. 49

    Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning

    Kaiwen Chen, Shuhai Zhang, Qiuwu Chen, Zimo Liu, Linxiao Li, Ying Sun, Yuchen Li, Yifan Zhang, Bo Han, Mingkui Tan

    cs.LG

    Large-scale neural network training increasingly relies on matrix-aware optimizers that exploit the structure of weight parameters beyond element-wise adaptation. However, existing matrix-aware methods such as Muon have an underappreciated vulnerability: their core operation, Newton-Schulz iteration, depends critically on input conditioning, yet the raw momentum matrices exhibit severe coordinate-wise scale heterogeneity. In this paper, we...

    arxiv.org/abs/2606.14187 · PDF

  50. 50

    Context-aware Modality-Topology Co-Alignment for Multimodal Attributed Graphs

    Sirui Zhang, Xu Wang, Zhengyu Wu, Xunkai Li, Hongchao Qin

    cs.LG · cs.CV

    Multimodal Attributed Graphs (MAGs) model real-world entities by coupling graph topology with heterogeneous attributes such as text and images. They support graph-centric tasks requiring structural and class-discriminative representations, and modality-centric tasks requiring fine-grained cross-modal correspondence. However, existing MAG methods often rely on fixed graph contexts or uniformly fused representations, causing task-agnostic...

    arxiv.org/abs/2606.14172 · PDF

  51. 51

    Machine Learning for Biomedical Raman Spectroscopy: From Spectral Acquisition to Clinical Translation

    Bogdan Oancea, Ana Maria Seciu-Grama, Nicoleta Siminea, Laura Mihaela Stefan, Alice Stoica, Joel Sjoberg, Marian...

    cs.LG

    Raman spectroscopy provides label-free, chemically specific characterization of biological systems and has become an important tool for cancer diagnosis, molecular subtyping, microbiological identification, and intraoperative decision support. Biomedical Raman spectra are, however, high-dimensional, noisy, and affected by fluorescence background, acquisition variability, and biological heterogeneity, making robust computational analysis...

    arxiv.org/abs/2606.14169 · PDF

  52. 52

    Curvature-Guided Geometric Representation for Protein-Ligand Binding Affinity Prediction

    Shuai Li, Chuan-Xian Ren, Yuhao Li, Ziqi Huang, Yue Pan, Mingzhe Tang, Hong Yan

    cs.LG · q-bio.BM

    Protein-ligand binding affinity (PLA) prediction is critical in drug discovery. Despite the notable advancements in machine learning-based approaches, existing methods struggle to jointly characterize local geometric organization and globally coordinated cross-molecular interactions, limiting their ability to model complex binding mechanisms. Here, we propose RicciBind, a geometric representation framework that integrates curvature-guided...

    arxiv.org/abs/2606.14159 · PDF

  53. 53

    Learning Urban Access Costs from Origin-Destination Flows via Inverse Optimal Transport

    Paula Joy B. Martinez

    cs.LG · cs.AI

    Cities deliver basic services through mixed public-private facility networks, including schools, clinics, transit providers, and subsidized service points. In these systems, planners often observe where households go, but not the latent cost function through which they trade off factors such as distance, price, and institutional access. We study this urban problem through school choice in the Philippines, where the country's largest national...

    arxiv.org/abs/2606.14157 · PDF

  54. 54

    Learning High Coverage Discriminative Parsimonious Rulesets

    Mariamma Antony, Raman Sankaran, Chiranjib Bhattacharyya, Uma Satya Ranjan

    cs.LG · cs.AI

    Learning systems based on IF-THEN rule representations readily offer interpretability, making them a crucial focus in contemporary AI research. A key objective for such rule sets is to achieve both high discriminative power and interpretability. While existing state-of-the-art algorithms implicitly prioritize predictive accuracy, they often fall short on one or more quality metrics that ensure interpretability, such as coverage and parsimony...

    arxiv.org/abs/2606.14156 · PDF

  55. 55

    Graph-based Target Back-Propagation for Context Adaptation in Multi-LLM Agentic Systems

    Tan Zhu, Tong Yao, Kananart Kuwaranancharoen, Amit Singh, Yushang Lai, Deepa Mohan, Shankara Bhargava

    cs.LG · cs.CL

    Context adaptation automates prompt engineering in LLM-based systems by iteratively revising tunable prompts from task feedback, without modifying model weights. Extending this paradigm to multi-LLM agentic systems is crucial: existing methods suffer from inaccurate credit assignment and lack convergence guarantees. We propose \textbf{G}raph-based \textbf{T}arget \textbf{B}ack-\textbf{P}ropagation (GTBP), a context adaptation framework for...

    arxiv.org/abs/2606.14155 · PDF

  56. 56

    Small LLMs: Pruning vs. Training from Scratch

    Yufeng Xu, Taiming Lu, Kunjun Li, Jiachen Zhu, Mingjie Sun, Zhuang Liu

    cs.LG · cs.CL

    Pruning promises a shortcut to strong small language models. In this work, we examine this promise by pruning Llama-3.1-8B at pruning ratios of 0.5--0.8 with six methods spanning depth, width, and sparse granularities, under two controlled token-matched settings. (1) With the same training token budget, pruned initialization consistently outperforms random initialization. This shows that the parent model provides a strong starting point,...

    arxiv.org/abs/2606.14150 · PDF

  57. 57

    Trust but Verify: Mitigating Medical Hallucinations via Post-Hoc Adversarial Auditing and Multi-Agent Feedback Loops

    Muhammad Osama, Maheera Amjad, Zartasha Mustansar, Arslan Shaukat, Muhammad U. S. Khan

    cs.LG

    Large Language Models (LLMs) are increasingly deployed in healthcare settings, yet their tendency to hallucinate poses risks when clinical decisions are involved. This study examine whether LLMs recommend recently banned or withdrawn pharmaceuticals when answering clinical questions and tests an agent-based method for reducing such errors. We developed a five-agent "Trust but Verify" system using a single LLM backbone. To measure regulatory...

    arxiv.org/abs/2606.14149 · PDF

  58. 58

    Decoupled Latent Optimization of Diffusion Models for Full Waveform Inversion

    Chen Min, Zheng Ma

    cs.LG

    Full waveform inversion (FWI) recovers subsurface velocity from seismic recordings by solving a severely ill-posed, nonconvex PDE-constrained optimization. Classical regularizers stabilize the inversion but fail to reproduce realistic geological structures; recent diffusion-prior methods improve realism at the cost of a fragile trade-off between data fidelity and prior consistency. We propose Decoupled Latent Optimization (DLO), which relaxes...

    arxiv.org/abs/2606.14139 · PDF

  59. 59

    Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning

    Omar Adalat, Edwin Hamel-De le Court, Francesco Belardinelli

    cs.LG · cs.MA

    Safe coordination problems surface in multi-agent reinforcement learning when global safety cannot be enforced by any agent unilaterally: the admissibility of one agent's action may depend on the dynamics of other agents. Decentralised shields can enforce safety at runtime, but purely factorised permissions often exclude optimal team behaviour that is safe only through coordination. We study deterministic safety guarantees for agents trained...

    arxiv.org/abs/2606.14130 · PDF

  60. 60

    Recovering Stranded Discrimination in Knowledge Tracing: Per-Item Bias Correction via Empirical-Bayes Shrinkage

    Xiaoran Yan, Cheng Tang, Atsushi Shimada

    cs.LG · cs.AI

    Deployed knowledge-tracing models are typically frozen after training, yet systematic per-item logit bias arises, from limited per-item expressivity in backbone architectures and from post-deployment shifts in item properties, degrading prediction quality. Global post-hoc calibrators such as Platt scaling, temperature scaling, and isotonic regression improve probability estimates but leave discriminative ability, as measured by AUC,...

    arxiv.org/abs/2606.14123 · PDF

  61. 61

    DTVEM-RE: A Hierarchical Random-Effects Extension of the Differential Time-Varying Effect Model for Person-Specific Multi-Lag Estimation in Intensive Longitudinal Data

    Amartya Bhattacharya

    cs.LG · stat.ME

    The Differential Time-Varying Effect Model (DTVEM) of Jacobson et al. (2019) is a popular tool for finding the best time lag in intensive longitudinal data, but it assumes everyone shares the same lag structure. The original authors named fixing this as future work, and it clashes with the premise of modern clinical research, which is that people differ. We present DTVEM-RE, an extension that lets each person have their own lag coefficients,...

    arxiv.org/abs/2606.14116 · PDF

  62. 62

    Numbers Already Carry Their Own Embeddings

    Suhyun Bae, Donghun Lee

    cs.LG · cs.AI

    We introduce Adelic operation-preserved embeddings (AOE), a training-free representation that captures both a number's real value and its modular (p-adic) signatures. This construction preserves additive and multiplicative structure by design, turning numerical input into embeddings that "speak in the language of mathematics." Unlike prior approaches that rely on task-specific retraining, AOE is plug-and-play and drops seamlessly into...

    arxiv.org/abs/2606.14108 · PDF

  63. 63

    Lyapunov-Based Sample Complexity Analysis for Weakly-Coupled MDPs

    Tianhao Wu, Matthew Zurek, Weina Wang, Qiaomin Xie

    cs.LG · math.OC · math.PR · stat.ML

    We study the sample complexity of learning in average-reward weakly-coupled Markov decision processes (WCMDPs) and Restless Bandits (RBs) under a generative model. Naive reduction to a tabular MDP leads to high complexity bounds as the state-action space is exponentially large in the number of arms $N$. By exploiting the weakly coupled structure, we show that near-optimal policies can be learned with sample and computational complexities that...

    arxiv.org/abs/2606.14095 · PDF

  64. 64

    Deep Spectral Learning of Embedded Latent Transfer Operators for Stochastic Dynamical Systems

    Ryogo Tanaka, Yoshinobu Kawahara

    cs.LG

    We propose a spectral learning method for stochastic nonlinear dynamical systems represented with embedded latent transfer operators in deep feature spaces. We instantiate the method as Deep Spectral Encoder (DSE), an operator-based latent state-space model in which a time-invariant neural encoder implements learnable nonlinear feature maps from observations, and these features define Markovian latent states whose temporal evolution and...

    arxiv.org/abs/2606.14079 · PDF

  65. 65

    Rethinking Backdoor Adversarial Unlearning through the Lens of Catastrophic Forgetting in Continual Learning

    Zhenqian Zhu, Yamin Hu, Yujiang Liu, Luping Wei, Wenbo Hou, Bin Li, Haodong Li, Wenjian Luo

    cs.LG · cs.AI

    Existing studies reveal that current backdoor defenses exhibit limited robustness and often fail against specific types of attacks. More concerningly, prevailing safety tuning strategies tend to provide only superficial safety protection, as they fall short of completely eliminating the backdoor effects. In this work, we present a novel formulation of backdoor learning and unlearning as a sequential, three-stage process from a continual...

    arxiv.org/abs/2606.14078 · PDF

  66. 66

    Non-Parametric Machine Text Detection via Multi-View Gaussian Processes

    Aleem Khan, Nicholas Andrews

    cs.LG · cs.CL

    Adversarial conditions such as paraphrasing and targeted style transfer sharply degrade the accuracy of machine text detectors. A document, however, carries multiple complementary signals (e.g., stylistic features, likelihood and rank-order features, and structural features), and an attack that suppresses one may leave others intact. While a parametric classifier can learn to combine these features given sufficient supervision, classifiers...

    arxiv.org/abs/2606.14060 · PDF

  67. 67

    Decompose Sparsely Where You Should, Absorb Densely Where You Should No

    Ruixuan Deng, Zehao Jin, Zekun Wang, Zihan Dong

    cs.LG

    Sparse autoencoders (SAEs) are typically trained to reconstruct the \textbf{entire} residual stream through a sparse dictionary, implicitly assuming that all activation content is amenable to sparse, monosemantic decomposition. We question this assumption and hypothesize that activations contain a low-rank, dense component that is computationally important to the model yet inherently unsuitable for sparse representation, which serves as a...

    arxiv.org/abs/2606.14040 · PDF

  68. 68

    Utility-Constrained Policy Optimization

    Mehrdad Moghimi, Bernardo Avila Pires

    cs.LG

    Constrained MDPs (CMDPs) are a widely adopted framework for incorporating safety into RL agents; however, the framework does not support risk-sensitive constraints. This can be problematic: For example, CMDPs allow for optimal solutions that, in order to satisfy the risk-neutral constraints, mix infrequent catastrophic behaviors and frequent, overly conservative ones. Moreover, prior empirical results suggest that enforcing stricter,...

    arxiv.org/abs/2606.14029 · PDF

  69. 69

    PostDeg: Placement Beats Parameterization in LayerNorm GNNs

    Yash Tomar, Aryav Das

    cs.LG

    LayerNorm-based GNNs routinely erase the topology signals (degree, centrality, $k$-core) that node-selection policies should depend on, but the literature has not located where in the residual block the erasure happens. We answer that question: a positive per-node scalar inserted before LayerNorm is divided out up to a stabilizer term, while the same scalar inserted after LayerNorm reaches the score head as representation magnitude. The...

    arxiv.org/abs/2606.14022 · PDF

  70. 70

    Can Machine Learning Forecast Rice Yields in Data-Constrained Settings? Satellite Climate Data, National Crop Statistics, and Lessons from Sierra Leone

    Ibrahim Denis Fofanah

    cs.LG

    Sierra Leone's agriculture operates with almost no data-driven decision support, and no published machine learning study has examined the country's crop yields. We ask whether rice yield can be forecast from data Sierra Leone currently has. Using 25 years of FAOSTAT production data (2000-2024) for nine major crops, we train XGBoost, Gradient Boosting, and Random Forest under a strict anti-leakage protocol with expanding-window walk-forward...

    arxiv.org/abs/2606.13959 · PDF

  71. 71

    Smoothing Dark Areas in Molecular Latent Diffusion

    Xi Wang, Jiahan Li, Yuxuan Xia, Yingcheng Wu, Shaoyi Zheng, Shengjie Wang

    cs.LG

    Latent diffusion is a promising framework for scalable 3D molecular generation, but it requires a latent space that remains smooth, valid, and navigable beyond posterior samples. Existing molecular VAEs, however, are typically learned through reconstruction-based objectives, which do not guarantee such a latent space. We show that this leads to dark areas: regions of latent space that are reachable during diffusion sampling but decode to...

    arxiv.org/abs/2606.13955 · PDF

  72. 72

    SpikF-GO: Spiking Fourier Graph Operators for Multivariate Time Series Forecasting

    Jafar Bakhshaliyev, Niels Landwehr

    cs.LG · cs.NE

    Spiking Neural Networks (SNNs) have emerged as an energy-efficient alternative to conventional neural networks, demonstrating strong performance in computer vision and robotics. More recently, SNNs have been applied to time series forecasting (TSF), with methods exploring spiking temporal backbones, spike-compatible positional encodings, Fourier-domain processing, and redesigned neuron dynamics. However, existing SNN forecasting approaches...

    arxiv.org/abs/2606.13901 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.