cs.LG · 2026-07-27 · No. 66

Machine Learning, 2026-07-27.

64 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

64 entries
  1. 01

    Dysphagia Risk Stratification in Head and Neck Cancer via Two-Stage PRO-Clinical Stacking

    Siyuan Zhao, Eric Ababio Anyimadu, Zachary G. Brumm, Yue Ma, Clifton David Fuller, Xinhua Zhang, G. Elisabeta Marai,...

    cs.LG

    Dysphagia is a debilitating late effect of head and neck cancer (HNC) treatment, yet timely identification of at-risk patients remains challenging in survivorship care. Definitive assessment relies on videofluoroscopic imaging, as captured by the Dynamic Imaging Grade of Swallowing Toxicity (CTCAE-DIGEST), which, while validated, requires specialized equipment, trained personnel, and significant patient burden, limiting its routine use in...

    arxiv.org/abs/2607.22514 · PDF

  2. 02

    Interpretable EEG biomarkers with bag-of-waves: Spatial and temporal waveform dictionaries for low-data regimes

    Athanasios Papastathopoulos-Katsaros, Steven T. Lee, Lin Yao, Ajay Thomas, Junseok Park, Matthew J. McGinley, Zhandong Liu

    cs.LG · eess.SP

    Electroencephalography (EEG) is widely used to diagnose neurological conditions, but its analysis usually relies on either predefined spectral features or deep neural networks. Predefined features carry a strong bias, since they fix in advance what counts as informative, while deep neural networks and foundation models are hard to interpret and need large amounts of data and compute. We present bag-of-waves, an interpretable framework that...

    arxiv.org/abs/2607.22508 · PDF

  3. 03

    Susceptible Reservoir Architectures for Regime-Conditional Volatility Forecasting

    Aliaksei Kaliutau

    cs.LG

    Volatility forecasting is dominated by persistence and measurement noise, leaving limited residual structure for nonlinear models to exploit. We introduce Susceptible Architectures (SUSA), a reservoir-design principle for volatility forecasting, and its two concrete implementations, based on complex-valued open-chain and periodic reservoirs and regime-conditioned experts to interpret reservoir features across calm, onset, recovery, and...

    arxiv.org/abs/2607.22491 · PDF

  4. 04

    \k{appa}-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating

    Jianghui Wang, Silong Yong, Francesco Orabona, Marco Canini, Katia P. Sycara, Yaqi Xie

    cs.LG · cs.AI

    Low-Rank Adaptation (LoRA) has become a widely adopted technique for efficient neural network fine-tuning, decomposing model updates into low-rank matrices. However, LoRA remains computationally costly because it updates all matrices uniformly, regardless of their actual contribution to adaptation. This cost is especially prohibitive for large-scale models with billions of parameters and for resource-constrained settings such as edge...

    arxiv.org/abs/2607.22489 · PDF

  5. 05

    Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent

    Peng Zhao

    cs.LG · math.ST · stat.ML

    In overparameterized linear regression, many weak spectral directions act like a ridge penalty on the signal-bearing spectrum; negative ridge is the natural correction, pushing filters above one. The stable negative-ridge endpoint, however, is structurally limited: its pole must stay below the smallest nonzero empirical eigenvalue, and it anti-shrinks smaller eigenvalues more than larger ones. Early-stopped negative-shifted gradient descent...

    arxiv.org/abs/2607.22474 · PDF

  6. 06

    Complexity Bounds and Approaches to Learning Projected Gradient Descent Solver Iterates

    Anjian Li, Ryne Beeson

    cs.LG

    Data scarcity poses a fundamental challenge in training generative models to produce initial guesses for parametric optimization problems that are otherwise numerically expensive to solve. We therefore study a $k$-neighborhood data collection strategy that augments datasets of converged solutions with intermediate solver iterates, increasing the amount of training data without additional solver runs. To understand the benefits of this...

    arxiv.org/abs/2607.22467 · PDF

  7. 07

    Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining

    Víctor Rincón Yepes

    cs.LG · cs.AI

    Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio models (AST, CLAP, BEATs-bio and BirdNET) with a downstream task none of them saw during training: recovering phylogenetic distance from species vocalizations. If the geometry of the embedding space tracks the tree of life, the representation is picking up something deeper than the labels the model was optimized for. We run...

    arxiv.org/abs/2607.22458 · PDF

  8. 08

    Hyperball May Not Be a Free Lunch

    Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai

    cs.LG · cs.AI

    For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing updates. However, the source of their advantage remains unclear. Starting from the angular displacement between consecutive parameter states, we derive an angular effective learning rate that accounts for the parameter-update angle, parameter norm, and update norm....

    arxiv.org/abs/2607.22444 · PDF

  9. 09

    On the Identifiability of Controlled World Models

    Xiangteng Zhang, Yang Guan, Bo Zhang, Ya-Qin Zhang, Shengbo Eben Li

    cs.LG

    Learning world models that infer environment dynamics from high-dimensional observations and predict outcomes under candidate actions is central to planning and control. Joint-Embedding Predictive Architectures (JEPAs) provide a compelling framework for learning such models in representation space. Recent action-conditioned extensions perform promisingly in visual control and latent-space planning, but leave a fundamental question unresolved:...

    arxiv.org/abs/2607.22430 · PDF

  10. 10

    LunarFM: A Shared Multimodal Representation of the Moon's Surface

    Marc Girona-Mata, Jakob Gawlikowski, Sumit Goski, Gautier Bardi de Fourtou, Valentin T. Bickel, Ben Moseley, Abigail...

    cs.LG

    The renewed global focus on lunar exploration, driven by the prospect of in-situ resource utilization and a sustained human presence on the Moon, has created growing demand for accurate, large-scale characterization of the lunar surface. Although vast quantities of orbital remote-sensing data have been collected, scientific analysis and resource mapping remain fragmented by heterogeneous multiinstrument observations, sparse labels, and...

    arxiv.org/abs/2607.22408 · PDF

  11. 11

    Local-Global Geometric Insights for Graph Neural Networks via Entropic Curvature

    Rachid Caich, Yassine Abbahaddou

    cs.LG · cs.SI · stat.ML

    Curvature notions on graphs, particularly Ollivier-Ricci and Forman, have emerged as powerful tools for addressing fundamental issues in Graph Neural Networks (GNNs) such as oversmoothing and oversquashing, but rely almost exclusively on local edge-level comparisons and therefore fail to certify how information actually propagates over long distances. We introduce Entropic Curvature, a global, transport-based curvature obtained by extending...

    arxiv.org/abs/2607.22381 · PDF

  12. 12

    Interior interpretability with attention rollout: contraction and propagation profiles in Transformers

    Umberto Biccari, Qian Huang, Enrique Zuazua

    cs.LG · cs.AI

    Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers. We introduce \emph{interior interpretability}, a propagation-based perspective on internal model organization, and instantiate it for tabular Transformers using attention rollout. We interpret rollout as a row-stochastic operator...

    arxiv.org/abs/2607.22367 · PDF

  13. 13

    Indexing: the Beginning and the End

    Alexander Kozachinskiy, Vicente Opazo, Felipe Urrutia

    cs.LG · cs.AI

    We study information bottlenecks in modern deep-learning architectures -- RNNs, softmax transformers, linear-attention transformers and state-space models -- through the lens of the indexing primitive. In this primitive, the input consists of $n$ bits and one integer $i$ from $1$ to $n$ called the index, and the output equals the value of the $i$-th bit. We introduce causal complexity for masked architectures. We show that architectures with...

    arxiv.org/abs/2607.22361 · PDF

  14. 14

    Integrated Order Dispatching and Routing for Last-Mile Pickup via Deep Reinforcement Learning

    Yida Xu, Zhaofang Mao, Yuheng Miao, Jiaxin Zhang, Yiting Sun

    cs.LG

    In recent years, the growing complexity of last-mile pickup operations has increased the need for fast and accurate decision-making on logistics platforms. This challenge is fundamentally driven by two key and tightly coupled decision-making processes: order dispatching and routing. Solving them separately overlooks their interdependence, while fully end-to-end learning can be unstable and costly on large, variable-scale instances due to...

    arxiv.org/abs/2607.22356 · PDF

  15. 15

    IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data

    Masashi Sode, Gianmarco Pinton

    cs.LG · physics.med-ph

    The speed of sound in tissue is a prerequisite for well-focused imaging and has diagnostic value, but recovering it from raw pulse-echo channel data is fundamentally a nonlinear inverse problem. Learned solvers are fast yet label hungry. Simulated sound-speed labels are expensive, while abundant real channel data is unlabeled. We propose IQ-JEPA to exploit both data types. An encoder is pretrained without labels to predict the latent...

    arxiv.org/abs/2607.22351 · PDF

  16. 16

    Beyond Binary Rooftop Mapping: A Four-Class Deep Learning Framework for Green Roof Potential Assessment from Open Swiss Geospatial Data

    Htet Yamin Ko Ko

    cs.LG

    The development of effective urban climate adaptation strategies requires comprehensive spatial information on rooftops and buildings, since such information underpins the assessment of ecosystem services provided by green infrastructure, particularly for urban heat island (UHI) mitigation. Although green roofs are widely acknowledged as a promising measure for improving urban thermal comfort, most existing research maps either current green...

    arxiv.org/abs/2607.22342 · PDF

  17. 17

    Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization

    Hao Wang, Kun Yuan, Wenlin Zhong, Minglei Zhang, Han Xiao, Ming Sun, Honggang Qi

    cs.LG · cs.AI · cs.CL

    Open-weight language models from different families exhibit complementary capabilities, motivating their consolidation into a compact student through on-policy distillation (OPD). However, full-vocabulary OPD typically assumes a shared tokenizer, while existing cross-tokenizer methods may discard teacher probability mass or assign it to student tokens with unrelated content. We introduce Byte-Prefix Marginalization (BPM), which re-expresses...

    arxiv.org/abs/2607.22334 · PDF

  18. 18

    Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs

    Zhangzhi Xiong, Minzhang Li, Haotian Yu, Sixian Shen, Kexin Zhang, Mingrui Li, Jie Zheng, Kewei Tu, Jingyi Yu

    cs.LG

    Multiple Sequence Alignments (MSAs) provide protein language models with explicit evolutionary context, but their large depth makes subsampling unavoidable under limited token budgets. Existing strategies, including random selection, identity-based filtering, and diversity-driven sampling, are effective heuristics, yet provide limited control over the evolutionary signals retained in the subset. In this work, we recast MSA subsampling as an...

    arxiv.org/abs/2607.22314 · PDF

  19. 19

    Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

    Roseline Polle, Owen Parsons, George Fairs, Luis Miguel San Martin Fernandez, Cole Looney, Xiaoliang Wu, Alexandra...

    cs.LG · cs.SD

    Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled data is expensive and some patient groups are underrepresented. Voice cloning is one such augmentation approach, but is typically evaluated on speech intelligibility (WER) or speaker similarity (SS) rather than on downstream performance, and it remains unclear...

    arxiv.org/abs/2607.22304 · PDF

  20. 20

    Efficient Recommendations via Graph Coarsening and Label Propagation

    Alessandro Sbandi, Federico Siciliano, Fabrizio Silvestri

    cs.LG

    Graph-based recommendations are widely adopted in real-world industrial applications. However, graphs in these systems often reach a massive scale, posing notable scalability and efficiency challenges. This requires techniques that can effectively balance predictive quality with computational cost. One promising approach is graph coarsening, an adaptive graph reduction technique that offers a way to systematically construct smaller, yet...

    arxiv.org/abs/2607.22287 · PDF

  21. 21

    An Insight on Evaluation Metrics Under the Imbalanced Case of Anomaly Detection

    Romain Hermary, Nesryne Mejri, Djamila Aouada

    cs.LG · stat.ML

    Anomaly detection is inherently characterised by severe class imbalance, making the interpretation of evaluation metrics challenging. Although metrics such as AUROC, AUPR, F1-score, and MCC are widely used, their values convey different meanings depending on the anomaly ratio. In this work, we analyse the behaviour of those four common anomaly detection metrics under varying levels of imbalance. We focus on the study of metric landscapes,...

    arxiv.org/abs/2607.22286 · PDF

  22. 22

    Autoregressive EHR Foundation Models with Multimodal Inputs

    Yuxuan Liu, Joshua Placidi, Jinpei Han, Alfred John Balston, Marek Rei, A. Aldo Faisal

    cs.LG

    Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way. We present a framework for conditioning such models on auxiliary clinical modalities, including ECG waveforms, chest X-ray images, and clinical notes, using modality-specific latent compression and...

    arxiv.org/abs/2607.22264 · PDF

  23. 23

    Class-Balanced Softmax: A Bayes Theory-Based Method for Long-Tailed Recognition

    Yi-Hang Zhu, Rajeev Raman, Shiqi Su, Jianyuan Sun, Xinyu Yang, Nan Xing, Huiyu Zhou

    cs.LG · cs.CV

    Deep learning models using traditional softmax classifiers have achieved remarkable success in various classification tasks. However, their performance degrades significantly on imbalanced datasets. Although Balanced Softmax is widely adopted as a state-of-the-art rebalancing method, it possesses inherent limitations, such as yielding disproportionately lower testing accuracy for tail classes. To mitigate these shortcomings, we propose the...

    arxiv.org/abs/2607.22258 · PDF

  24. 24

    IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning

    Wei Zhang, Xinwu Liu, Yihang Cheng

    cs.LG · cs.AI

    Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning method for large language models, but its performance depends strongly on how a fixed rank budget is distributed across Transformer modules. Existing adaptive-rank methods usually rely on local gradient statistics collected during training, which introduces extra memory and computation and overlooks task-conditioned global information flow. We propose IFCLoRA, a...

    arxiv.org/abs/2607.22251 · PDF

  25. 25

    Optimization of time-consuming experimental conditions using pseudo-experimental data guided by adaptive polynomial regression

    Hirotaka Sugawara, Yujin Taguchi, Kei Minagawa, Yusuke Hiki, Takashi Morikura, Akira Funahashi

    cs.LG · cs.AI

    Bayesian optimization (BO) is an optimization method that sequentially proposes the next candidate explainable variables for optimizing target variables by balancing exploration and exploitation. BO is often used under a limited evaluation budget, such as hyperparameter tuning of deep learning. Despite its effectiveness, conventional BO may have poor convergence in practical experimental science where each evaluation is often costly and...

    arxiv.org/abs/2607.22238 · PDF

  26. 26

    Latent PDE mapping for efficient physics-informed learning across geometries with limited data

    Ingvild Askim Adde, Mary M. Maleckar, Gabriel Balaban

    cs.LG · physics.comp-ph

    In this study, we introduce latent PDE mapping, a broadly applicable physics-informed learning technique designed to enable efficient geometric generalization with sparse training data. Latent PDE mapping pulls back geometry-specific PDE residuals and boundary conditions to a predefined latent geometry via the deformation gradient, thereby enabling the automated calculation of geometry-consistent shape gradients that are missing in...

    arxiv.org/abs/2607.22215 · PDF

  27. 27

    From Score Approximation to Distribution Approximation in Score-Based Diffusion Models

    Lan V. Truong

    cs.LG · cs.IT · stat.ML

    Score-based diffusion models have achieved remarkable empirical success in generative modeling, yet their approximation-theoretic foundations remain incomplete. In particular, although classical universal approximation theorems guarantee that neural networks can approximate score functions, it remains unclear whether such approximation guarantees translate into approximation of the probability distributions generated by reverse diffusion...

    arxiv.org/abs/2607.22199 · PDF

  28. 28

    Unbiased Open World Regularization for Fair Self-Supervised Learning

    L{é}o Nicollier, Marc Pic, Pablo Mus{é}, Enric Meinhardt-Llopis, Gabriele Facciolo

    cs.LG

    Despite recent advances, self-supervised learning (SSL) models and Joint-Embedding Predictive Architectures (JEPAs) remain susceptible to learning spurious biases in the dataset. These techniques rely on regularization, which prevents representation collapse by enforcing a global target distribution such as a multivariate Gaussian or a uniform distribution on the sphere. However, these global constraints are insufficient to prevent bias...

    arxiv.org/abs/2607.22149 · PDF

  29. 29

    TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex

    Yuliang Yan, Shuo Yan, Haochun Tang, Yiqin Sun, Enyan Dai

    cs.LG · cs.AI

    Molecular glue degraders have emerged as a promising strategy for targeted protein degradation by inducing ternary complex formation between an E3 ubiquitin ligase and a target protein. Despite their therapeutic potential, computational design of molecular glues remains largely unexplored. Unlike conventional structure-based drug design, molecular glue design is governed by the unknown protein-protein interface and requires the simultaneous...

    arxiv.org/abs/2607.22143 · PDF

  30. 30

    Pretraining EHR Foundation Models with Patient-Aware Sampling

    Joshua Placidi, Yuxuan Liu, Jinpei Han, Marek Rei, A. Aldo Faisal

    cs.LG

    Autoregressive foundation models for electronic health records (EHRs) typically inherit pretraining methods from language modeling, where patient trajectories are concatenated into a single token stream and windows are sampled from that stream. In EHR data, this choice is consequential: windows may mix multiple patients, and patients with longer records contribute more optimization updates, potentially introducing bias. We propose Patient...

    arxiv.org/abs/2607.22114 · PDF

  31. 31

    A Leakage-Free Stacked Ensemble Method for Multiclass Classification

    S. P. Sharmila, Aruna Tiwari

    cs.LG · cs.AI

    Multiclass classification is a fundamental problem across a wide range of domains. It is still challenging due to possession of high inter-class similarity, class imbalance datasets, and variability in data distributions. Rule-based classifiers such as XGBoost often achieve stronger performance on structured features, but they are limited in capturing smooth functional relationships among variables. Similarly, neural network models can...

    arxiv.org/abs/2607.22081 · PDF

  32. 32

    CEL: Comprehensive Counterfactual Explanations Library and Benchmark

    Oleksii Furman, Łukasz Lenkiewicz, Marcel Musiałek, Maciej Zięba

    cs.LG · cs.AI

    Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome. While early methods primarily focused on minimal feature changes, recent work incorporates additional properties such as sparsity, actionability and plausibility. Despite this progress, fair and systematic evaluation remains challenging....

    arxiv.org/abs/2607.22045 · PDF

  33. 33

    DCS: A Unified Conditional Sensitivity Framework for Cross-Modal Copyright Infringement Detection

    Xiafeng Man

    cs.LG

    Currently, most foundation models can reproduce or strongly depend on copyrighted training content, but output similarity alone is insufficient for infringement detection, because similar outputs may also arise from public-domain concepts, common stylistic conventions, or ordinary statistical generalization. In this paper, we develops a unified post-hoc detection framework that treats copyright infringement evidence as a counterfactual...

    arxiv.org/abs/2607.22035 · PDF

  34. 34

    Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits

    Yuta Natsubori, Masataka Ushiku, Yuta Saito

    cs.LG

    Off-Policy Evaluation and Learning (OPE/L) in contextual bandits is rapidly gaining popularity in real systems because new policies can be evaluated and learned securely using only historical logged data. However, existing methods in OPE/L cannot handle many challenging but prevalent scenarios such as few-shot data, deterministic logging policies, and new actions. In many applications, such as personalized medicine, content recommendations,...

    arxiv.org/abs/2607.22012 · PDF

  35. 35

    Energy Manifold Natural Gradient Descent: Riemannian Optimization for Neural PDE Solvers

    Zhangyong Liang, Huanhuan Gao

    cs.LG

    Energy natural gradient descent (ENGD) aligns parameter updates with the curvature of an underlying function-space energy, but existing formulations assume an unconstrained Euclidean parameter domain. We introduce \EMNGDfull{}, a manifold optimization framework for physics-informed and variational neural PDE solvers whose parameters lie on a Riemannian manifold. EMNGD restricts the energy-induced quadratic model to feasible tangent directions...

    arxiv.org/abs/2607.22004 · PDF

  36. 36

    From Perturbation Correction to Geometry-Aware Sampling: Sharpness-Guided Equilibrium Sampling for Balanced Flat Minima in Long-Tailed Learning

    Jiaxin Deng, Junbiao Pang

    cs.LG · cs.AI

    Long-tailed learning couples two sources of poor generalization: head classes dominate training exposure, while under-represented classes often converge to sharper regions of the loss landscape. Conventional re-sampling addresses the former without considering geometry, whereas existing long-tailed sharpness-aware minimization (SAM) methods modify losses or perturbations only after biased mini-batches have been drawn. We introduce...

    arxiv.org/abs/2607.21999 · PDF

  37. 37

    On the Convergence of Stochastic Low-Rank Adaptation

    Ru Wang, Chengchang Liu, John C. S. Lui

    cs.LG · math.OC

    Low-rank adaptation (LoRA) optimizes $J(B,A)=\mathcal L(W_\mathrm{base}+sBA)$ over two adapters $B \in \mathbb{R}^{m \times r}$ and $A \in \mathbb{R}^{r \times n}$ that form a low-rank update to a frozen pretrained weight matrix $W_\mathrm{base} \in \mathbb{R}^{m \times n}$. The prior analysis shows LoRA-GD takes $\exp\{\mathcal{O}(ε^{-2})\}$ oracle calls to find an $ε$-stationary point such that $\|\nabla J(B,A)\|\leq ε$ in the deterministic...

    arxiv.org/abs/2607.21975 · PDF

  38. 38

    Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning

    Shujin Wu, Cheng Qian, Xiusi Chen, Heng Ji

    cs.LG · cs.AI · cs.CL

    Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We hypothesize that the success of such evolution frameworks hinges on meta-skills, such as self-reflection with environment feedback, that enable effective multi-round refinement, yet are largely neglected by traditional post-training. To bridge this gap, we present MetaEvolve, a framework designed...

    arxiv.org/abs/2607.21971 · PDF

  39. 39

    MA-DAR: Manifold-Aligned Dynamic Adaptive Routing for Continual Temporal Knowledge Graph Reasoning

    Xiangjun Shi, Chong Mu, Jinchuan Zhang, Lizong Zhang, Yuefeng He, Shang Liu

    cs.LG · cs.AI

    Continual temporal knowledge graph (TKG) reasoning aims to continuously incorporate newly emerging facts while preserving previously acquired knowledge. Replay-based continual learning has achieved promising performance by revisiting historical representations. However, existing methods primarily focus on what to replay, while largely overlooking how replayed representations should be integrated with current ones. Such direct integration...

    arxiv.org/abs/2607.21949 · PDF

  40. 40

    Multi-Agent Debate and Visual Information Extraction for SeePhys Pro: A 1st-Place Technical Report from ICML 2026 AI4Math Track 3 Challenge

    Jiseok Kwak, Suhyeon Jo, Taewoo Kim, Yeongmin Kim, Byeonghu Na, Il-chul Moon

    cs.LG

    This technical report presents our approach to Challenge Track~3: SeePhys Pro at the 3rd AI for Math Workshop, where the task is to answer college-level physics questions whose statement and figure may be given partly or entirely as an image. Visual physics problems become substantially harder for large language models when the decisive information resides in a figure rather than in the text, and this modality gap widens as more of the...

    arxiv.org/abs/2607.21946 · PDF

  41. 41

    LatentFlow: Visual Analytics for Latent Space Analysis in Molecular Graph Neural Networks

    Shiyi Liu, Jiaqing Chen, Nicholas Hadler, Rostyslav Hnatyshyn, Michael W. Mahoney, Talita Perciano, John F. Hartwig,...

    cs.LG · cs.HC

    Chemists and materials scientists increasingly use machine learning models, such as graph neural networks (GNNs), to predict properties of molecules and the outcomes of their reactions. Beyond predictive performance, understanding how these models organize chemical information internally in their latent spaces, i.e., the embeddings of the molecules, is critical. Analyzing latent spaces helps diagnose model behavior and assess whether the...

    arxiv.org/abs/2607.21941 · PDF

  42. 42

    RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention

    Anderson R. Santos

    cs.LG

    Full self-attention in large language models scales as O(N^2), which limits long-context document analysis to 65,536 tokens and requires costly GPU clusters. The Reduced Interaction Sampling (RIS) inference engine addresses this constraint as a model-agnostic architecture. Without modifying weights, RIS reduces self-attention complexity to O(N log N) using sparse stochastic geometry that fits within commodity memory limits. We validate RIS on...

    arxiv.org/abs/2607.21927 · PDF

  43. 43

    MissHyper: Restoring Clinical Synchronicity in Missingness-Guided Hypergraph Forecasting

    Mingyi Ma, Qingxiong Tan

    cs.LG

    Clinical irregular multivariate time series are shaped not only by physiological dynamics but also by the measurement process that determines when and what to observe. In event-centric models, however, co-timestamp structure can be flattened too early: measurements acquired at the same timestamp are embedded as isolated nodes, leaving local patient-state context unavailable until later message-passing layers. We study this pre-propagation...

    arxiv.org/abs/2607.21922 · PDF

  44. 44

    Remedying Coarsening-Based GNN Training under Heterophily via Adaptive Complementary Enhancement

    Guoming Li, Jian Yang, Xukun Wang, Zixiao Wang, Shangsong Liang, Yifan Chen

    cs.LG · eess.SP · math.NA

    Coarsening-based training for graph neural networks (GNNs), i.e.\ training on coarsened graphs rather than the original large ones, has become a promising direction for scaling GNNs to massive graphs. However, prior work has been evaluated almost exclusively on \textit{homophilic} graphs, leaving the more challenging \textit{heterophilic} settings underexplored. We show, both empirically and theoretically, that existing coarsening-based...

    arxiv.org/abs/2607.21885 · PDF

  45. 45

    Variance-Reduced Q-Learning over Static and Time-Varying Networks

    Sreejeet Maity, Feng Zhu, Aritra Mitra, Robert W. Heath

    cs.LG · eess.SY

    We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The agents can exchange information over a network to collectively learn the optimal state-action value function. For this setting, we introduce a novel epoch-based distributed $Q$-learning algorithm called VRDQ, where within each epoch, agents locally estimate the Bellman optimality operator and...

    arxiv.org/abs/2607.21876 · PDF

  46. 46

    Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study

    Kaihua Ding

    cs.LG · stat.ML

    Prior classical-ML learning-curve work fits power laws to tree, linear, and kernel models on tabular data, but at small scale: typically one curve, one team, a handful of cells. We present a distributed classroom-scale replication: 127 students each ran a fixed protocol on 3 assigned datasets, drawn from 18 tabular classification and regression datasets and 6 model families (Boosting, Random Forest, SVM, Linear/Logistic, Ridge, Lasso),...

    arxiv.org/abs/2607.21866 · PDF

  47. 47

    LeAct: Learning to Reason from Expert Actions

    Ziran Yang, Chengshuai Shi, Raj Ghugare, Benjamin Eysenbach, Karthik Narasimhan, Chi Jin

    cs.LG · cs.AI · cs.CL

    Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines, classical planners, theorem provers), which routinely produce near-optimal actions across diverse domains. But these experts are silent: they commit to an action without writing down the chain of thought (CoT) behind it....

    arxiv.org/abs/2607.21856 · PDF

  48. 48

    Searching the Space of Feed-Forward Neural-Network Weight-Update Rules with Fixed Depth Symbolic Regression

    Charles Brum, Edward Finkelstein

    cs.LG

    We investigate whether symbolic regression can discover explicit neural network weight-update rules that outperform standard hand-designed optimizers on small symbolic regression benchmarks. Candidate update rules are represented as fixed-depth symbolic expressions over operands derived from common optimizers, including gradient, momentum, adaptive-gradient, and moment-estimate quantities. Across 30 benchmark/neural network combinations, the...

    arxiv.org/abs/2607.21855 · PDF

  49. 49

    A Graph-Based Control Interface for Traffic Signals on Heterogeneous Road Networks

    Bertil Braun

    cs.LG · eess.SY

    We present a traffic-signal control interface in which a shared graph neural network assigns scores to individual traffic movements. Each junction converts these scores into its own variable-sized set of legal signal phases using a deterministic incidence matrix. Directed corridor nodes provide traffic context, while movement nodes represent controlled input-to-output paths through junctions. Typed mean aggregation produces one scalar per...

    arxiv.org/abs/2607.21831 · PDF

  50. 50

    Data eccentricity, asymptotics of Gaussian RBF reproducing kernel Hilbert space, and kernel PCA

    Sergio A. Alvarez

    cs.LG

    We show that, up to isotropic scaling, the Gaussian RBF reproducing kernel Hilbert space (RKHS) is asymptotically isometric to Euclidean space in the large bandwidth limit. This strongly suggests that kernel-based constructions reliant on metric properties of the RKHS will yield results for Gaussian RBF kernels that similarly approach those of linear kernels for large bandwidths. The asymptotic behavior of Gaussian CKA can be understood in...

    arxiv.org/abs/2607.21823 · PDF

  51. 51

    Bounding the Causal Impact of ML-assisted Decision-Making via Counterfactual Correctness

    Jonathan Zhang, Erik Skalnes, Jacob Chen, Michael Oberst

    cs.LG

    Predictive machine learning (ML) models are increasingly used to aid human decision-makers across various high-risk domains such as healthcare and criminal justice. There is a growing recognition of the need to evaluate the causal impact of deploying these systems on downstream outcomes, such as patient survival or crime recidivism. Randomized control trials (RCTs) can provide high-quality evidence on the impact of a deployed model, but they...

    arxiv.org/abs/2607.21806 · PDF

  52. 52

    Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection

    Othmane Harraq, Tamer Aldwairi

    cs.LG · cs.CR · cs.CV · cs.MM

    Talking-face (TF) deepfake generation synthesizes photore- alistic facial video from a static source image and an au- dio signal, producing forgeries that current image-based detectors consistently fail to identify. Unlike face-swap ma- nipulation, TF synthesis has no underlying real video from which to inherit physiological characteristics, making re- mote photoplethysmography (rPPG) a uniquely motivated detection modality for this forgery...

    arxiv.org/abs/2607.21776 · PDF

  53. 53

    Smart predict-then-robustly-optimize

    Aakil Caunhye, Xuefei Lu, Belen Martin-Barragan

    cs.LG · math.OC

    In this paper, we propose and study a robust variant of the smart predict-then-optimize approach that accounts for prediction shifts due to disturbance in the covariate feature space. While traditional integrated-learning-and-optimization models assume that side information is perfectly revealed, empirical data-driven features are frequently corrupted or noisy at the time of decision-making, leading to fragile operational policies. To bridge...

    arxiv.org/abs/2607.21773 · PDF

  54. 54

    Parameter-free Adaptive Sparse Attention via Compression-Based Content Selection

    Debarshi Kundu, Swaroop Ghosh, Vasant Honavar

    cs.LG

    Data-adaptive sparse attention masks substantially outperform fixed patterns (e.g., BigBird and Longformer) and can even exceed dense attention on long sequences. Existing adaptive approaches---including SBM-Transformer, Dynamic Mask Attention, and NSA---typically require additional learnable parameters, custom gradient estimators, or specialized CUDA kernels. We show that classical data compression provides an effective masking signal with...

    arxiv.org/abs/2607.21752 · PDF

  55. 55

    RED-PIM: Reducing Data Movement for Transformers using Processing-in-Memory

    Zahra Yousefijamarani, Alaa Alameldeen

    cs.LG · cs.AR

    Transformers are widely used across many domains, including natural language processing, computer vision, web search, and DNA sequence analysis. Given their broad applicability, improving the performance of transformer models is critical. However, the high volume of data movement between processing units and memory during attention operations significantly limits their efficiency. Processing-In-Memory (PIM) mitigates this issue by performing...

    arxiv.org/abs/2607.21731 · PDF

  56. 56

    A Defense of the Quadratic Model

    Alexandru Meterez, Pranav Ajit Nair, Depen Morwani, Cengiz Pehlevan, Sham Kakade, Alex Damian

    cs.LG · cs.AI · math.OC · stat.ML

    Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimization dynamics. In this work, we stress test the simplest possible model of optimization -- the quadratic model -- and show that it can be surprisingly predictive in an LLM setting with 150M...

    arxiv.org/abs/2607.21716 · PDF

  57. 57

    An Introduction to Bayesian and Frequentist Simulation-Based Inference with Machine Learning

    Maximilian Dax, Theo Heimel, Gilles Louppe

    cs.LG · astro-ph.CO · astro-ph.GA · hep-ex · hep-ph · stat.ML

    Simulation-based inference (SBI) with machine learning is an increasingly important tool for solving inverse problems in science and engineering, including parameter inference and the inversion of detector effects. We provide an overview of the Bayesian and frequentist statistical frameworks, describe how machine-learning-based SBI methods, such as neural posterior estimation and neural likelihood estimation, can be used for parameter...

    arxiv.org/abs/2607.21702 · PDF

  58. 58

    Expanding Flow Maps

    Sophia Tang, Pranam Chatterjee

    cs.LG

    Flow-based generative models have enabled remarkable progress in fast and controllable generation across continuous and discrete state spaces, yet existing parameterizations are constrained to fixed dimensions or fixed sequence lengths. Here, we introduce Expanding Generative Flows (EFlows), which define flows between distributions of increasing dimensionality along an expanding interpolant that grows the state by augmenting it with...

    arxiv.org/abs/2607.21585 · PDF

  59. 59

    Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity

    Hongnan Ma, Yiwei Shi, Mengyue Yang, Weiru Liu

    cs.LG · cs.AI

    Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserve a black-box model's prediction, but also necessary for maintaining it. However, existing sufficiency-oriented methods can assign high importance to spurious subsequences that support the prediction without being essential to the model's decision. We introduce \textbf{TimePNS}, a necessity-aware framework for time-series...

    arxiv.org/abs/2607.21573 · PDF

  60. 60

    Graph Learning on Ensembles of Cyclic Peptides: An Investigation of Molecular Ensemble Modeling

    Aaron Feller, Kris Deibler, Maxim Secor

    cs.LG · q-bio.BM

    Molecular property prediction from structure often uses a single representative conformation, even though many molecules exist as conformational ensembles in solution. We introduce EnsembleEGNN, a molecular ensemble foundation model that encodes an ensemble by first encoding each conformer with shared Equivariant Graph Neural Network (EGNN) layers, then pooling the resulting conformer representations with a Set Attention Block. We pretrain...

    arxiv.org/abs/2607.21561 · PDF

  61. 61

    X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

    Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin

    cs.LG

    While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-OPD, a cross-modal on-policy distillation framework that transfers reasoning capabilities from a powerful text teacher to an audio-language student. During training,...

    arxiv.org/abs/2607.21550 · PDF

  62. 62

    Zero-Flow Two-Sample Tests

    Yakun Wang, Leyang Wang, Song Liu, Taiji Suzuki

    cs.LG · stat.ML

    We propose a new approach to two-sample testing for deciding whether two sets of samples are drawn from the same distribution. The test is built on a statistical discrepancy based on the zero-flow criterion, termed zero-flow discrepancy (ZFD). We prove the validity of ZFD and propose a practical testing procedure, termed the zero-flow two-sample test (ZF2ST). The key idea is to learn how samples from the two distributions are locally...

    arxiv.org/abs/2607.21542 · PDF

  63. 63

    Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

    Alagappan Valliappan

    cs.LG · cs.CL · cs.PF

    Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models increasingly ship a built-in Multi-Token-Prediction (MTP/NEXTN) draft head under the assumption that the draft is negligibly cheap. At million-token context this breaks: an MTP draft head typically runs full attention over the entire KV cache at every draft step, so its read grows linearly with...

    arxiv.org/abs/2607.21535 · PDF

  64. 64

    Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets

    Jim Allchin

    cs.LG · cs.CL

    Sparse attention reduces the cost of long contexts by allowing each query to read only selected parts of the input. These selectors are often trained by distilling the attention patterns of a dense teacher, assuming that attention reveals which context the teacher actually uses. We test that assumption on retrieval tasks where the evidence for each answer is known exactly. By masking parts of the context and measuring whether the answer...

    arxiv.org/abs/2607.21692 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.