cs.LG · 2026-08-27 · No. 97

Machine Learning, 2026-08-27.

59 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

59 entries
  1. 01

    Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher's Role

    Ahmad Khan, Akram Bin Sediq, Sara Azadegi Naeini, Raviraj S. Adve

    cs.LG · cs.IT · eess.SY

    Designing machine learning algorithms for wireless resource management is labour-intensive: the architecture, the loss function and the training recipe are all specified by hand. We demonstrate that this design layer can be surrendered to an autonomous agent in its entirety. We adopt the autoresearch protocol, in which an AI coding agent edits a training script, runs a fixed-budget experiment, and retains or discards the change according to a...

    arxiv.org/abs/2608.26093 · PDF

  2. 02

    TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

    Jiarui Yan, Weiwei Sun, Sijie Li, Wenhan Li, Yiming Yang

    cs.LG · cs.AI

    Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning development, where an agent must revise data pipelines, models, and validation over hours of feedback, and on most competitions still finishes below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We...

    arxiv.org/abs/2608.26086 · PDF

  3. 03

    ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

    Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer, Marc-Andre Schulz, Nys Tjade Siegel, Maximilian Dreyer,...

    cs.LG · cs.AI · cs.CV · stat.ML

    Deep neural networks often exploit spurious associations in their training data, a failure known as shortcut learning. Concept-based explainability methods screen for shortcuts by testing whether concepts such as a patient's sex or scanner settings can be decoded from a network layer. Because each concept is evaluated in isolation, these methods can mistake correlations between concepts as evidence that the model uses them. We introduce ICON...

    arxiv.org/abs/2608.26083 · PDF

  4. 04

    Group-Shared Low-Rank Approximation for Mobile-Efficient Pointwise Convolutions in Large-Kernel CNNs

    Hao Luo, Yiting Yang, Wenyi Zhao, Man Jiang, Zhijun Lin, Ghulam Mohiuddin, Ting Jiang, Kunming Luo, Zihao Zhang,...

    cs.LG

    Large-kernel Convolutional Neural Networks (CNNs) deliver remarkable performance in vision tasks by significantly expanding receptive fields, yet their quadratic parameter growth critically impedes storage-efficient edge deployment. While existing efficient architectures adopt parameter-efficient depthwise separable convolution backbones that leverage techniques like low-rank approximation and weight sharing to compress depthwise...

    arxiv.org/abs/2608.26069 · PDF

  5. 05

    How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention

    Gerard Conangla Planes

    cs.LG · cs.AI · cs.CL

    Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task. In this paper, we provide a task-dependent theory of the approximation error achievable at each LoRA rank for Transformer attention. We fix a pretrained attention head, a target attention function, and a distribution over inputs from the downstream task, and bound the smallest expected Kullback--Leibler (KL) error achievable by a rank-$r$ query LoRA update....

    arxiv.org/abs/2608.26052 · PDF

  6. 06

    Robust CurveMoE: Multi-Norm Adversarial Defense for Mixture-of-Experts Models via Mode Connectivity

    Xu Zhang, Ren Wang

    cs.LG

    Multi-norm adversarial defense aims to protect neural networks against perturbations defined by different norm constraints, but existing methods typically optimize competing robustness objectives within a single parameter configuration, leading to substantial training cost and unfavorable robustness trade-offs. We propose Robust CurveMoE, an efficient mixture-of-experts framework that connects models specialized for different perturbation...

    arxiv.org/abs/2608.26043 · PDF

  7. 07

    DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation

    Yutong Chen, Guangfu Guo, Zhichao Xu, Kunpeng Liu

    cs.LG · cs.AI

    On-policy self-distillation (OPSD) uses a privileged copy of the student model to provide dense supervision without an external teacher. OPSD keeps this privileged teacher fixed, even though the student distribution and output style change during training. We propose DualOPSD, an asymmetric alternating framework that adapts both policies. The student first learns from the privileged teacher. The teacher then moves toward the updated student...

    arxiv.org/abs/2608.26019 · PDF

  8. 08

    Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

    Xiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland

    cs.LG

    Orthogonal optimisers such as Muon can substantially accelerate large language model pretraining relative to Adam, yet the mechanism remains incompletely understood. We investigate this through an out-of-sample spectral probing analysis of Transformer loss landscapes. At checkpoints along real training trajectories, we decompose each momentum buffer into its singular directions and estimate the loss-optimal step size along each direction on...

    arxiv.org/abs/2608.25990 · PDF

  9. 09

    When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs

    Suchit Gupte, Xueru Zhang, Mohammad Mahdi Khalili

    cs.LG

    Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood. We present a systematic study of how pruning affects SAE behavior and theoretically show that, for a fixed SAE, its impact is governed by perturbation energy, a covariance-weighted norm. This perspective exposes a key limitation of magnitude...

    arxiv.org/abs/2608.25941 · PDF

  10. 10

    One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

    Justin Robert, Raheel Qader

    cs.LG · cs.AI · cs.CL

    On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test...

    arxiv.org/abs/2608.25936 · PDF

  11. 11

    Quantum-Inspired Modeling of Driving Behavior

    Mohammad Elayan, Omid Armantalab, Wissam Kontar

    cs.LG · eess.SY · stat.ML

    Driver behavior is heterogeneous, context-dependent, and changes over time, and these properties shape the traffic phenomena we observe. Most models, however, fix in advance which behavioral variables interact and how. Behavior outside that form is absorbed as noise, while models flexible enough to capture it tend to lose interpretability. We introduce a quantum-inspired representation of driver behavior that combines properties usually...

    arxiv.org/abs/2608.25907 · PDF

  12. 12

    Forecasting Multiple Observables with SCROLL: Score-Trained Uncertainty for Stochastic Dynamics

    Pavel Prochazka

    cs.LG

    Forecasting a stochastic dynamical system rarely means a single number: one wants several observables---future state, threshold event, regime label---each with its own likelihood. Standard multi-task recipes balance per-task losses, tuned or learned. We instead compose the observables' likelihoods in per-task free-routed last-layer beliefs on a shared backbone; this absorbs unit-dependent loss scaling into likelihood parameters learned in the...

    arxiv.org/abs/2608.25898 · PDF

  13. 13

    Towards A Unified Information Bottleneck Framework for Time Series Explanations

    Xu Zheng, Zichuan Liu, Zhuomin Chen, Mayur Akewar, Janki Bhimani, Jason Liu, Mo Sha, Jingchao Ni, Wei Cheng, Dongsheng Luo

    cs.LG · cs.AI

    Explaining deep learning models operating on time series data is crucial in various applications that require transparent and interpretable insights into model behavior. {Existing explanation methods generally fall into two categories: attribution-based explanations, which identify the temporal regions most responsible for a prediction, and counterfactual explanations, which reveal how an input should be modified to alter the model's...

    arxiv.org/abs/2608.25897 · PDF

  14. 14

    A General-Purpose Molecular Foundation Model Transfers Across Diverse Olfactory Tasks

    Yikun Han, Yi Wang, Neil Mankodi, Stephen Yang, Ambuj Tewari

    cs.LG

    Foundation models have transformed molecular property prediction, yet it remains unclear whether a molecular foundation model, fine-tuned on a single canonical olfactory prediction task, can learn representations that transfer across diverse machine olfaction problems. We investigate this question by fine-tuning Uni-Mol2 on the GS-LF benchmark for multi-label odor descriptor prediction and evaluating the resulting model, without additional...

    arxiv.org/abs/2608.25893 · PDF

  15. 15

    How Edge of Stability Hinders SCAFFOLD in Federated Optimization

    Anant Khandelwal, Michael Crawshaw, Mingrui Liu

    cs.LG

    In federated learning, it is well known that heterogeneous data can (in theory) slow down optimization, and much effort has been directed at designing optimization algorithms that are unaffected by data heterogeneity, such as the SCAFFOLD algorithm. Yet, despite strong theoretical guarantees, SCAFFOLD does not usually outperform the much simpler FedAvg in practice. In this work, we propose that this gap is due to the presence of Edge of...

    arxiv.org/abs/2608.25873 · PDF

  16. 16

    CEDAR: Controlled and Event-Driven Demand Forecasting via Residual Decomposition

    Junjie Meng, Ranxu Zhang, Zi-an Zhang, Shujun Liu, Xiaoning Qi, Xiaozhou Xu, Yanyong Zhang, Hui Xiong, Chao Wang

    cs.LG

    Forecasting in large-scale e-commerce marketplaces is increasingly required to support planning: merchants need to evaluate sales outcomes under future action sequences such as budget schedules, rather than passively predicting what happens next. However, most existing time series forecasting (TSF) approaches remain inherently passive. Even when incorporating operational decisions as auxiliary covariates, they typically optimize for...

    arxiv.org/abs/2608.25871 · PDF

  17. 17

    VINCENT: Validated Interaction Network for Cross-drug Explanation of Therapeutics

    Fan-Sheng Chuang, Xuchen Li, Yujing Bian, Kaixiong Zhou

    cs.LG · cs.AI

    Drug synergy prediction estimates whether two drugs produce a stronger joint effect than expected from their individual activities. For drug combination discovery, a single synergy score is often not enough: researchers also need to know which molecular regions jointly drive the prediction. We study motif-pair synergy explanation, which identifies pairs of chemically coherent regions, one from each drug, that jointly contribute to predicted...

    arxiv.org/abs/2608.25841 · PDF

  18. 18

    Learning Continuous Regional Temperature Fields with Lead-Time and Resolution Queries

    Chunlei Shi, Jiong Wang, Yi-Lin Wei, Junming Hou, Jinjin Liu, Yecheng Zhang, Dan Niu

    cs.LG · cs.MM

    Accurate regional near-surface temperature forecasting is fundamental to short-range weather services and downstream risk assessment. Existing deep learning-based regional forecasters commonly produce a fixed set of future frames on a prescribed grid, limiting their use when forecast products must be evaluated at query-dependent lead times or display resolutions. To overcome these fixed-output constraints, we formulate regional T2M...

    arxiv.org/abs/2608.25823 · PDF

  19. 19

    Canalization Before Generalization: Grokking as a Dynamical Probe

    Yiming Lin

    cs.LG

    For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. We sweep short, fixed-duration weight-decay (WD) pulses across this plateau and measure how they shift later generalization time. Across three grokking tasks, these...

    arxiv.org/abs/2608.25813 · PDF

  20. 20

    Geometry-Constrained Kolmogorov-Arnold Networks: Learning Edge Geometry via Banach Duality

    K S Sesh Kumar

    cs.LG · stat.ML

    Kolmogorov-Arnold Networks (KANs) replace fixed activations in deep architectures with learnable univariate edge functions, making the choice of edge parametrisation central. Existing variants rely on fixed bases such as splines, polynomials, or Fourier features, which impose a function-space geometry before data are observed. We introduce geometry-constrained KANs, a family of edge activations derived from Banach duality maps in which the...

    arxiv.org/abs/2608.25807 · PDF

  21. 21

    Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data

    Rene Glitza, Luca Becker, Rainer Martin

    cs.LG · cs.DC · cs.SD · eess.AS · eess.SP

    Federated Learning (FL) enables distributed training of machine learning models while preserving data privacy. However, FL struggles with heterogeneous, non-IID client data distributions, resulting in sub-optimal and biased global models. In this paper, we propose pFedMARL, a novel approach leveraging Multi-Agent Reinforcement Learning (MARL) with Twin Delayed Deep Deterministic Policy Gradient (TD3) to dynamically adapt aggregation...

    arxiv.org/abs/2608.25794 · PDF

  22. 22

    EXAONE Tabular 1.0 : Technical Report

    Moonjung Eo, Min-Kook Suh, Hye-Seung Cho, Jiwon Kim, Seoyoon Kim, Sangjun Nam, Soonyoung Lee

    cs.LG

    EXAONE Tabular is a compact tabular foundation model family for classification and regression via in-context learning, producing predictions without dataset-specific gradient updates. Pretrained exclusively on a synthetic structural-causal-model (SCM) prior, its central contribution is an architecture-centered redesign of tabular in-context learning. Rather than compressing features into a fixed row embedding before a separate row-level...

    arxiv.org/abs/2608.25774 · PDF

  23. 23

    Drift-Aware Multimodal User Representation Learning via Multi-Scale Temporal Modeling and Sparse Mixture-of-Experts

    Ziqing Qian, Haohang Chen, Shengqi Dang, Yuhan Xiong, Canyu Shen, Jiaying Lei, Nan Cao

    cs.LG

    Understanding user preferences from noisy and temporally evolving social media behaviors is fundamentally challenging due to interest drift, where user preferences shift across time and exhibit both multi-scale temporal patterns and diverse co-existing interests. To address this, we propose DUMoE, a unified framework for drift-aware multimodal user representation learning. Our model consists of (i) a temporal dynamics-aware backbone that...

    arxiv.org/abs/2608.25773 · PDF

  24. 24

    Learning from waste: Machine Learning for health risk prediction and computer vision-based sorting in Ghana

    Hilda Adwubi Osei, Catherine Tenewaa Osei, Desdemona Yaa Asobayire

    cs.LG · cs.CV · cs.CY

    The inappropriate disposal of solid waste remains a significant public health and environmental concern worldwide, including in Ghana. Poor sanitation and improper waste management practices contribute to substantial economic costs and avoidable deaths annually. In 2022, a field study in Atonsu, Kumasi, Ghana, reported a community-perceived relationship between household waste disposal and illness patterns, but only through descriptive...

    arxiv.org/abs/2608.25759 · PDF

  25. 25

    TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

    Sadhika Malladi, Samy Jelassi, Dylan Foster, Jordan T. Ash, Akshay Krishnamurthy

    cs.LG · cs.AI

    Reinforcement learning post-training drives reasoning and agentic capabilities in modern AI systems, yet a growing body of work shows that it is most effective when used to fine-tune an already capable base model. We question whether existing pipelines yield models that are most suitable for reinforcement learning. Building on prior work highlighting the role of coverage and pass@K as predictors of post-RL performance, we design a simple...

    arxiv.org/abs/2608.25756 · PDF

  26. 26

    Comparing Corrupted Constrained Learning Problems

    Laura Iacovissi, Rabanus Derr, Robert C. Williamson

    cs.LG · math.ST · stat.ML

    A key result in statistics is the data processing inequality, originally proved by Blackwell (1951) and later refined by DeGroot (1962) in terms of statistical uncertainty. It states that the Bayes risk of a statistical experiment obtained by stochastically modifying another experiment cannot be lower than the Bayes risk of the original experiment, regardless of the loss function or prior chosen. In machine learning, this result underlies...

    arxiv.org/abs/2608.25745 · PDF

  27. 27

    A Constitutive Markov Physics-Informed Neural Operator (MPNO) for Autoregressive Stability in Transient Dynamics

    Wenpu Du, Peng Zhou, Yunlong Xia, Sinuo Xin, Congcong Zhang, Boyang Zhang, Yi Zhang, Wenzheng Xu

    cs.LG

    Neural operators applied to transient-dynamics PDEs with strong discontinuities exhibit autoregressive instability: in concrete-penetration stress-field prediction, the wavelet neural operator (WNO) diverges in autoregressive rollout, while MeshGraphNets collapse to zero predictions. WNO's instability stems from the lack of a structural constraint on the spectral radius of its propagation operator; the Fourier neural operator (FNO) is stable...

    arxiv.org/abs/2608.25744 · PDF

  28. 28

    Why Does Graph Learning Fail to Fully Benefit from a Text Teacher?

    Fumiaki Kimino, Ryoma Sato

    cs.LG · cs.CL

    Graph neural networks (GNNs) are widely used to represent complex interactions and relationships among entities. We investigate a multimodal model that combines two complementary ideas: a self-supervised method that enables a GNN encoder pretrained on one dataset to operate directly on another dataset with a different node-feature dimensionality, without rebuilding the model or realigning the data; and an alternating optimization method that...

    arxiv.org/abs/2608.25741 · PDF

  29. 29

    Are LLM-Enhanced GNNs Privacy-Safe?

    Longzhu He, Zelang Wen, Chaozhuo Li, Sen Su

    cs.LG · cs.CR

    Large language models (LLMs) have recently advanced graph neural networks (GNNs) by enriching node representations with semantic information, giving rise to LLM-enhanced GNNs that achieve substantial performance gains. However, their vulnerability to privacy attacks, in which adversaries infer sensitive information from model outputs, remains largely underexplored. To bridge this gap, we present a systematic evaluation of privacy risks in...

    arxiv.org/abs/2608.25727 · PDF

  30. 30

    It's a matter of timescale: non-linear utility in successor features and multi-objective planning and learning

    Liam P. H. Mertens, Lucas N. Alegre, Florent Delgrange, Diederik M. Roijers, Ann Nowé, Peter Vamplew

    cs.LG · cs.AI

    Time is of the essence when dealing with multiple reward signals and non-linear utility. In this paper we argue that the current main approaches in multi-objectiveRL (SER and ESR), and successor features, are insufficient. While each approach deals with non-linear effects on user utility on different timescales, none of them take into account that different effects happening on different timescales can happen within the same decision problem....

    arxiv.org/abs/2608.25723 · PDF

  31. 31

    Fairness-Aware Test-Time Prompt Tuning

    Yoann Launay, Parameswaran Kamalaruban, Tom Kempton, Stuart Burrell, David Sutton

    cs.LG

    Vision-language models have displayed remarkable capabilities in multi-modal understanding and are increasingly used in critical applications where economic and practical deployment constraints prohibit re-training or fine-tuning. However, these models can also exhibit systematic biases that disproportionately affect protected demographic groups and existing approaches to addressing these biases require extensive model retraining and access...

    arxiv.org/abs/2608.25707 · PDF

  32. 32

    Tropospheric temperature and humidity profile retrieval from Meteosat Flexible Combined Imager based on deep learning

    Alejandro Salgueiro, Johannes Rausch, Julie Thérèse Villinger, Angela Meyer

    cs.LG · physics.ao-ph

    The Meteosat Third Generation (MTG) Flexible Combined Imager (FCI) offers new opportunities for tropospheric temperature and humidity profiling, at higher spatio-temporal resolutions and expanded spectral coverage relative to its predecessor. Vertically resolved retrievals from broadband imagers are inherently challenging, and operational retrieval algorithms typically rely on numerical weather prediction (NWP) background fields to compensate...

    arxiv.org/abs/2608.25700 · PDF

  33. 33

    Modeling spatio-temporal locality in multi-step forecasting of geo-referenced time series

    Annunziata D'Aversa, Gianvito Pio, Michelangelo Ceci

    cs.LG

    Forecasting future measurements from geographically distributed sensors is essential across many domains. However, the spatial distribution of these sensors raises multiple challenges, primarily due to spatial autocorrelation phenomena, that introduce inter-dependencies among nearby locations, that cannot therefore be treated independently. While some existing approaches can capture such phenomena, they generally model the spatial dimension...

    arxiv.org/abs/2608.25698 · PDF

  34. 34

    Adversarial Training of Linear Models under Stealthy Attacks

    Lovisa Eriksson, Dave Zachariah, André M. H. Teixeira

    cs.LG · cs.CR · eess.SY

    Predictive models are widely used in many fields, but are vulnerable to false data injection attacks. To address this, detection schemes and adversarial training have been proposed, but such approaches lack guarantees against stealthy attacks. We therefore propose a detector-based switched model, in which optimal attack strategies are stealthy. For linear prediction models, we derive a convex formulation of the resulting adversarial risk. The...

    arxiv.org/abs/2608.25681 · PDF

  35. 35

    LDAC-Net: A Learnable Multi-Lag Differencing Attention-Convolution Network for Drift-Robust Recognition with Low-Cost MOX Gas Sensors

    Xin Zhang, Liangxiu Han, Yue Shi, Tam Sobeih

    cs.LG · cs.AI

    Portable electronic-nose systems based on low-cost metal-oxide (MOX) gas sensors offer a practical solution for gas and odour recognition, but their signals are affected by slow chemical transients, drifting sensor offsets, scale variation, and cross-channel correlations. Existing pipelines commonly use fixed first-order temporal differencing (FOTD), which requires a manually selected lag and may discard useful response information. We...

    arxiv.org/abs/2608.25646 · PDF

  36. 36

    A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation

    Bing Shao, Jiazheng Zhang, Long Ma, Yujiong Shen, Senjie Jin, Xin Guo, Yuming Yang, Mingxu Chai, Zhiheng Xi, Tao...

    cs.LG · cs.CL

    On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens remains poorly understood. We analyze the gradient of the per-token K2 estimator of reverse KL with respect to the student logits. The $\ell_1$ norm of this gradient factorizes into the absolute teacher--student log-probability gap and a student-side softmax factor...

    arxiv.org/abs/2608.25643 · PDF

  37. 37

    DCEO: Direct Causal Effect Optimization for Long-Term User Value Modeling in E-commerce Search

    Junzhao Zhang, Tao Zhang, Liren Yu, Feiyi Dong, Zhixuan Zhang, Dan Ou, Haihong Tang

    cs.LG · cs.IR

    Industrial e-commerce search systems ultimately aim to optimize the user-level long-term objective, such as n-day cumulative purchases or gross merchandise value (GMV) per user. However, such objectives are defined at the user level, whereas search ranking is based on item-level scores within each request. Existing methods typically bridge this granularity gap through manually designed multi-objective fusion, where predictions of multiple...

    arxiv.org/abs/2608.25635 · PDF

  38. 38

    Frequency-aware forecasting for short-term typhoon gust prediction

    Xuefei Wang, Tingyi Liu, Heng Zhang, Shengjun Zhang

    cs.LG

    Accurate gust forecasting under typhoon conditions remains challenging due to the highly non-stationary and multi-scale characteristics of extreme wind fluctuations. Existing deep learning models often struggle to simultaneously capture long-term trends and rapid local variations, resulting in degraded performance during extreme events. We propose WDANet, a frequency-aware forecasting framework that integrates stationary wavelet...

    arxiv.org/abs/2608.25604 · PDF

  39. 39

    M-Fibration Theory with Applications to Neural Network Compression

    Paolo Boldi

    cs.LG

    The purpose of this paper is to provide a general, comprehensive, theoretical framework that allows one to deal with fibrations on graphs labelled on a commutative monoid. This is a genuine extension of the theory of graph fibrations (as introduced in "Fibrations of Graphs" [Discrete Math., vol. 243, pp. 21-66, 2002]), that makes it possible to deal with weighted graphs, and also graphs labelled with other algebraic structures. The derived...

    arxiv.org/abs/2608.25598 · PDF

  40. 40

    Individual Fairness in Hierarchical Clustering

    Binita Maity, Shrutimoy Das

    cs.LG

    Hierarchical clustering produces ultrametric representations that impose strong global geometric constraints and may distort local similarities in ways that disproportionately affect individual data points. We study hierarchical clustering under an individual fairness requirement that bounds relative distortion within local $k$-nearest neighborhoods. We formulate this requirement as a feasibility problem over dominated ultrametrics and...

    arxiv.org/abs/2608.25586 · PDF

  41. 41

    Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory

    Siyuan Chen, Runlin Hou, Shenxiu Wu, Yansong Sun, Junming Cao, Yiyu Zhang, Shudi Shao, Junhao Qiu, Zhichao Lu, Qingfu Zhang

    cs.LG · cs.MA

    Hardware kernel optimization requires repeated compilation, correctness testing, profiling, and revision. LLM agents can automate parts of this process, and stronger foundation models, longer context windows, and longer execution horizons have improved optimization within individual tasks. These advances alone do not enable an agent to learn from completed optimization runs. Existing kernel-optimization agents seldom preserve a decision, its...

    arxiv.org/abs/2608.25570 · PDF

  42. 42

    Physics-Informed Foresight Pruning for Sparse PINN Solvers of Nonlinear PDEs

    Ahmad Ishaque Karimi, Uvini Balasuriya Mudiyanselage, Kookjin Lee

    cs.LG · cs.AI

    Physics-informed neural networks (PINNs) often rely on over-parameterized models to optimize coupled solution and differential-residual objectives, leaving unclear how much capacity is necessary and what pruning should preserve. We study foresight pruning at initialization for sparse PirateNet PDE solvers. Standard neural tangent kernel spectrum-aware pruning (NTK-SAP) aims to preserve output-side training dynamics but may overlook parameters...

    arxiv.org/abs/2608.25564 · PDF

  43. 43

    Beyond Optimal Rates in Stochastic Optimization: Trajectory-Adaptive Stopping Rules

    Liviu Aolaritei, Lucas Lévy, Francis Bach, Michael I. Jordan

    cs.LG · math.OC · math.ST · stat.ML

    Stochastic gradient descent (SGD) is typically analyzed at a deterministic horizon chosen before the algorithm is run, even though practical stopping decisions are made adaptively by inspecting the evolving trajectory. This mismatch creates a fundamental certification problem: fixed-time guarantees do not generally remain valid at data-dependent stopping times, while deterministic horizons derived from worst-case bounds can be highly...

    arxiv.org/abs/2608.25551 · PDF

  44. 44

    Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction

    Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer, Marco Simnacher, Jordan F. Safer, Sumaiya Iqbal,...

    cs.LG · q-bio.QM

    Recently, there has been a growing adoption of protein language models (PLMs) in biomedical science. Their embeddings provide a rich numerical representation of protein sequences which achieve state-of-the-art performance on several downstream tasks including protein fitness prediction. However, PLM embeddings are not directly interpretable and, thereby, it remains unclear what features they encode. To gain insight into which biochemical...

    arxiv.org/abs/2608.25548 · PDF

  45. 45

    Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference

    Jiarui Hu, Zhiyuan Wen, Xiaoyun Liu, Jiaxing Shen, Yu Yang

    cs.LG · cs.CL

    Large reasoning models often produce reasoning traces with verification, revision, and backtracking. When reflection merely re-checks established results, it wastes reasoning tokens and increases latency. Most existing reflection steering methods add a label-derived mean-difference direction across preset layers, but its entanglement with reasoning and length signals destabilizes the accuracy-efficiency trade-off. In this paper, we propose...

    arxiv.org/abs/2608.25542 · PDF

  46. 46

    Resilient Decentralized Wireless Federated Learning via Gradient Tracking with AdamW

    Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Vu Nguyen Ha, Symeon Chatzinotas

    cs.LG · cs.DC · cs.NI

    Wireless Internet-of-Things (IoT) edge networks require decentralized learning (DecL) methods that can operate reliably under both heterogeneous local data and communication-constrained wireless links. However, existing decentralized optimization schemes often incur substantial communication overhead and degraded performance when transmissions are constrained by strict airtime budgets, fading channels, and packet losses. This paper proposes...

    arxiv.org/abs/2608.25535 · PDF

  47. 47

    FedQoS: Federated QoS-Risk Learning for Heterogeneous Indoor-Outdoor Access Selection

    Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Zerihun Huruy, Vu Nguyen Ha, Symeon Chatzinotas

    cs.LG

    Reliable access selection in dynamic and heterogeneous indoor-outdoor environments is challenging because instantaneous radio measurements alone cannot capture future QoS degradation caused by mobility, blockage, traffic load, and resource competition. This paper proposes FedQoS, a federated QoS-risk learning framework for predicting the future reliability of candidate access links and supporting access-node selection without centralizing...

    arxiv.org/abs/2608.25496 · PDF

  48. 48

    A Storage-Retrieval Gap in Parametric Knowledge Graph Memory

    Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov, Volker Tresp

    cs.LG · cs.CL · cs.IR

    Graph retrieval-augmented generation places retrieved subgraphs into the model's context window at query time, paying a recurring token cost and exposing source data on every call. We study an alternative: compiling a knowledge graph offline into a bank of LoRA adapters, one per entity, that serve as a parametric knowledge layer queried by injecting weights rather than text, at zero query-time context cost. On the MetaQA dataset, we find that...

    arxiv.org/abs/2608.25489 · PDF

  49. 49

    Resolving Multi-Modal Regression by Difference-Quotient-Based Clustering:Fast Coarse Conditional-Label Assignment

    Huang Weiquan

    cs.LG

    Multimodal regression suffers from the mean-collapse pathology: under squared loss, an unconstrained regressor converges to the conditional mean, which for K > 1 lies away from all modes. We attribute this failure to pairwise contradictions--samples with nearly identical inputs but distant outputs--and propose Difference-Quotient Clustering (DQC), which partitions data to minimize intra-cluster output-vs-input discrepancy. Each sample is...

    arxiv.org/abs/2608.25467 · PDF

  50. 50

    Joint Initialization of Flux Networks and Effective Multiplication Factor for Physics-Informed Neural Networks Solving Neutron Diffusion Problems

    Qin Hang, Yangdi Yi, Jiayi Li, Xu Wang, Heng Zhang

    cs.LG

    Efficient determination of the effective multiplication factor (keff) is an important computational task in reactor core neutronics analysis. Physics-informed neural networks (PINNs) incorporate neutron diffusion equations and boundary conditions into network training to efficiently determine the neutron flux distribution and keff. To further improve the efficiency of keff calculations using PINNs, a Joint Initialization Physics-Informed...

    arxiv.org/abs/2608.25443 · PDF

  51. 51

    Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks

    Andrey Labunets

    cs.LG · cs.AI · cs.CR

    Refusal training protects AI models from jailbreaks by training models to decline unsafe queries, reducing the risk of misuse. Recent work finds that refusal behavior in aligned language models can be mediated by a single activation direction or a low-dimensional refusal subspace shared across harmful prompts: ablating those directions suppresses refusals while largely preserves other model capabilities. Yet it remains unclear why...

    arxiv.org/abs/2608.25390 · PDF

  52. 52

    PaSta: Noisy Node Classification with Partial Label Learning

    Yujing Liu, Yixin Liu, Yu Zheng, Yue Tan, Alan Wee-Chung Liew, Shirui Pan

    cs.LG

    Noisy node classification problem is a fundamental yet challenging task for real-world graph-related web services, where node labels are often corrupted or unreliable due to weak supervision or automatic annotation. However, existing methods typically train models based on one-hot labels, which not only makes models susceptible to overfitting on noisy labels, but also leads to error accumulation after pseudo-label-guided enhancement. In this...

    arxiv.org/abs/2608.25365 · PDF

  53. 53

    Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement

    Yihang Zhang, Shengke Sun, Junjie Wen, Feng Zeng

    cs.LG · cs.AI · cs.CL

    Model merging provides an efficient way to construct multi-task generalist models without additional training, but its performance often degrades under severe task interference. Task interference in model merging primarily stems from \textit{superposition}, where task-specific features become entangled within the parameter space. This entanglement renders conventional decomposition methods insufficient for effectively isolating useful task...

    arxiv.org/abs/2608.25354 · PDF

  54. 54

    Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning

    Srivalli Katkuri, Maxwell Kawada, Juan Wachs

    cs.LG · cs.RO

    Vision-language models (VLMs) have emerged as a powerful source of supervision for reinforcement learning, enabling agents to leverage rich semantic knowledge during training. Inspired by the success of preference-based reward learning (PbRL) in reinforcement learning from human feedback (RLHF), vision-language model generated image-based preferences provide an effective source for learning reward functions. This can be done by visually...

    arxiv.org/abs/2608.25350 · PDF

  55. 55

    Neither Precision Nor Architecture Alone: Controlled Tests of Failure Remedies for Physics-Informed Neural Networks

    Jinyuan Zhang, Peng He, He Hu, Yin Yuan, ShengShuo Jiao

    cs.LG · cs.AI

    Physics-Informed Neural Networks (PINNs) frequently fail on stiff or advection-dominated PDEs, and two recent accounts offer competing remedies: switching from FP32 to FP64 to repair an L-BFGS stopping artifact, or replacing the MLP with a state-space-model (SSM) backbone plus sub-sequence alignment to counter architectural simplicity bias. We test both under matched, seed-paired controls in a pre-registered 144-run study spanning convection,...

    arxiv.org/abs/2608.25327 · PDF

  56. 56

    Two Dimensions Govern Agnostic Multiclass Transductive Learning

    Pahan Dewasurendra

    cs.LG

    In transductive classification, an adversary fixes a labeled population, one label is hidden uniformly, and the learner sees all remaining labels. For binary classes, agnostic transductive and PAC learning have the same minimax rate. Whether this extends to multiclass learning was open, especially for unbounded label spaces where uniform convergence can fail. We resolve the question up to logarithmic factors. For every multiclass class...

    arxiv.org/abs/2608.25326 · PDF

  57. 57

    Activation-Space Order-Swap Geometry: A Site-Asymmetry Audit

    Anqi Peter Li

    cs.LG

    Order-dependent activation statistics are often interpreted as evidence of interaction, but that interpretation can be confounded by where interventions enter the network. We introduce a no-fit site-asymmetry audit. For a twice-differentiable readout, the open-path order-swap decomposes into a canonical additive response measured by single interventions and an antisymmetrized second difference free of first-order and pure self-curvature terms...

    arxiv.org/abs/2608.25315 · PDF

  58. 58

    Prefix-Denoising Consistency: Test-Time Verification for Diffusion Language Models

    Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar, Junpei Komiyama

    cs.LG

    Diffusion Language Models (DLMs) have recently become increasingly competitive with autoregressive (AR) models, and even outperform them on certain tasks. Unlike AR models, DLMs produce output through iterative denoising without a left-to-right order. To further improve the performance of DLMs, we introduce PDC (\emph{Prefix-Denoising Consistency}), a test-time self-verification method for DLMs. PDC exploits a distinctive test-time signal in...

    arxiv.org/abs/2608.25311 · PDF

  59. 59

    InsightSR: Refining Symbolic Regression Search Spaces via Parallel Semantic and Structural LLM Guidance

    Yating Ling, Wenjing Cun, Zhitang Chen

    cs.LG · cs.AI

    Symbolic regression (SR) seeks to discover parsimonious mathematical laws from observational data, yet conventional approaches often struggle with the vast combinatorial search space of physically meaningful expressions. We present InsightSR, a framework that embeds Large Language Models (LLMs) as a guiding layer around the PySR genetic programming engine. Rather than relying on LLMs to generate expressions directly, InsightSR uses LLMs to...

    arxiv.org/abs/2608.25291 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.