cs.LG · 2026-07-31 · No. 70

Machine Learning, 2026-07-31.

48 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

48 entries
  1. 01

    KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models

    Sparsh Roy, Samuel Girmachew, Nishita Chavan

    cs.LG · q-bio.QM

    Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pipelines have been proposed to catch this, but their components are rarely stress-tested, so it is unclear which parts of an audit can be trusted and under what conditions. We present KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism...

    arxiv.org/abs/2607.28608 · PDF

  2. 02

    $β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

    Jiawei Xu, Minghui Liu, Juzheng Zhang, Tom Goldstein, Furong Huang

    cs.LG

    On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is precisely the $β=1$ member of a broader policy-optimization family, where $β$ weights the KL penalty anchoring the student to a reference policy. This equivalence turns $β$...

    arxiv.org/abs/2607.28582 · PDF

  3. 03

    APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

    Shentong Mo, Yatao Bian

    cs.LG · cs.AI · cs.MA

    Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While flow-matching models (, FlowDPO) have recently shown promise in this domain, their performance relies heavily on alignment with ground-truth coordinates via supervised preference learning. However, obtaining experimental labels for novel crystal phases or de novo proteins is prohibitively expensive, creating a bottleneck for...

    arxiv.org/abs/2607.28553 · PDF

  4. 04

    Same Graph Cross-Task Transfer in GNNs: Protocols and Predictors

    Neelam Akula, Surbhi Kumar, Murat Kantarcioglu, Baris Coskunuzer

    cs.LG · cs.SI

    Many real-world graphs support multiple predictive tasks over the same underlying structure, creating an opportunity to reuse supervision across node classification (NC) and link prediction (LP). However, existing evaluations often rely on incompatible splits, observed-graph assumptions, and negative sampling rules, making conclusions about same-graph cross-task transfer unreliable. We formalize same-graph NC-LP transfer and propose a...

    arxiv.org/abs/2607.28525 · PDF

  5. 05

    The Role of Causality in Algorithmic Recourse

    Srikanth Avasarala, Varun Gupta, Shahin Jabbari, Saber Salehkaleybar, Juba Ziani

    cs.LG · cs.CY · cs.GT

    Algorithmic recourse aims to provide individuals with actionable changes to improve their predicted outcomes in high-stakes classification settings, such as loan and mortgage applications. However, most existing approaches focus only on flipping a model's prediction, without accounting for whether the recommended changes lead to genuine improvement in an individual's true qualifications or merely enable strategic gaming of the classifier....

    arxiv.org/abs/2607.28497 · PDF

  6. 06

    Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation

    Alexander Boesgaard Lorup

    cs.LG · cs.CL

    Stage-replay diagnostics reconstruct intermediate token prefixes and treat fresh-prefill continuation as continuation from the decoder state that originally reached the prefix. We audit that assumption at a whole reasoning-stage boundary in a Qwen2.5-derived system. A matched 200-item experiment compares retained live cache with one-shot prefill of identical integer tokens and places an exact replica on both sides. In BF16, replicas remain...

    arxiv.org/abs/2607.28495 · PDF

  7. 07

    Cybersecurity Detection Classification with Reasoning-enabled Language Models

    Amol Khanna, Manu Nandan, Cristian Viorel Popa, Joan Pujol-Roig, Diana Bolocan, Laura Vasilie, Alexandru Apostu,...

    cs.LG · cs.CR

    A major issue in Security Operations Centers (SOCs) is alert fatigue, as the number of detections reported is more than staff can triage in a given day. Prior work prompts or fine-tunes large language models (LLMs) to emit a triage label directly, but does not train them to reason about whether a detection is a genuine threat. We train a chain-of-thought (CoT) reasoning-enabled triage classifier on real, human-labeled Windows endpoint...

    arxiv.org/abs/2607.28460 · PDF

  8. 08

    Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory

    Jiannan Yang, Veronika Thost, Xiang Ling, Tengfei Ma

    cs.LG

    Molecular optimization is commonly performed under a limited oracle budget, which makes deciding what to evaluate as important as deciding what to generate. We introduce short-term graph memory, a plug-in module that preserves the generator architecture and native update rule while learning from previously evaluated molecules to prioritize subsequent oracle queries. The module maintains an online graph neural surrogate that pre-screens each...

    arxiv.org/abs/2607.28437 · PDF

  9. 09

    Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification

    V. S. Usatyuk, D. A. Sapozhnikov, S. I. Egorov

    cs.LG · cs.CV · cs.IT

    We introduce Kohn--Sham Spectral Embedding (KSSE), a physics-inspired energy-based model replacing dense CNN classifiers with a sparse-graph spectral embedding evaluated at the Nishimori temperature of an associated Random-Bond Ising Model. By mapping pre-trained features onto quasi-cyclic low-density parity-check graphs and constructing a regularized Laplacian acting as a Kohn--Sham Hamiltonian, we solve $D$ independent channel spectral...

    arxiv.org/abs/2607.28428 · PDF

  10. 10

    QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction

    Ran Miao, Rui Luo, Xiaohan Shan, Xiaoming Sun

    cs.LG

    Fault-tolerant quantum computing (FTQC) relies on quantum error correction to suppress physical errors and preserve logical information at scale. In practice, however, performance is constrained not only by physical noise but also by the latency of classical decoders processing rapidly generated syndrome data. This challenge is exacerbated by hardware noise that is strong, heterogeneous, and nonstationary, as well as by the...

    arxiv.org/abs/2607.28422 · PDF

  11. 11

    QQWorld: Quantile-Quantile Matching for World Model Regularization

    Zhoushun Yu, Xiaoyu Hu, Xiangyu Xu

    cs.LG · cs.AI · cs.CV · cs.MM · cs.RO

    Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its latents toward an isotropic Gaussian using the Epps-Pulley (EP) objective. We show that the corrective gradients of EP rapidly vanish for isolated tail samples, leaving heavy-tailed deviations insufficiently...

    arxiv.org/abs/2607.28415 · PDF

  12. 12

    On-Policy and Off-Policy Learning for Large Action Spaces

    Imad Aouali

    cs.LG · cs.AI · math.ST · stat.ML

    This thesis studies policy learning in interactive systems where an agent observes a context, selects an action from a very large set, and receives partial feedback. The main framework is contextual bandits, with two paradigms: on-policy learning, where the agent interacts sequentially with the environment and minimizes regret, and off-policy learning, where it learns from logged data collected by a logging policy. In large action spaces,...

    arxiv.org/abs/2607.28408 · PDF

  13. 13

    Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees

    Zihan Dong, Rui Qian, Qishi Zhan, Dongshen Peng, Kaixin Li, Yu Li

    cs.LG

    Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We identify the main cause as expensive autoregressive decoding on the decision-time critical path. We propose Adaptive Anticipatory Policy Trees (AAPT), which eliminates this delay without modifying the underlying model. During idle screen periods, the same frozen multimodal model constructs a...

    arxiv.org/abs/2607.28399 · PDF

  14. 14

    Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs

    Ankur Naskar, Vaneet Aggarwal

    cs.LG

    Constrained Markov Decision Processes (CMDPs) provide a natural framework for reinforcement learning in safety-critical applications, where agents maximize long-term reward while satisfying long-term constraints. Although primal-dual actor-critic methods with linear critics are well understood, extending order-optimal convergence guarantees to neural critics in average-reward CMDPs has remained open. The main challenge is a fundamental...

    arxiv.org/abs/2607.28390 · PDF

  15. 15

    LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

    Enjun Du, Hange Zhou, Chenxu Du, Siyi Liu, Zirong Chen, Ziyu Zheng, Yongqi Zhang

    cs.LG

    Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state...

    arxiv.org/abs/2607.28374 · PDF

  16. 16

    Encryption-Compatible Clustered Federated Learning via Distributed Expectation-Maximization over Metadata

    Michael Ben Ali, Imen Megdiche, André Péninou, Olivier Teste

    cs.LG · cs.CR · cs.DC · stat.ML

    Clustered Federated Learning (CFL) addresses data heterogeneity in federated settings by grouping clients with similar data distributions to enable effective training. Existing methods face a trade-off between privacy preservation, communication cost, and computational efficiency. We formalize this as the CFL trilemma, according to which improving two of these dimensions comes at the expense of the third. A prominent paradigm relies on...

    arxiv.org/abs/2607.28338 · PDF

  17. 17

    Measuring Distortion in the Empty Regions of Dimensionality Reduction Scatterplots with the Gap Index

    Jaume Ros, Alessio Arleo, Fernando Paulovich

    cs.LG

    Quality metrics play a crucial role in the proper use of dimensionality reduction projections for visual analysis of high-dimensional data. They quantify the degree of distortion of a projection compared to the high-dimensional data and provide a reliable indication of how confident users can be in the structures they see in the resulting layouts. However, most popular metrics focus on capturing direct relationships between points (e.g.,...

    arxiv.org/abs/2607.28324 · PDF

  18. 18

    Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing

    Huiyuan Tian, Bonan Xu, Shijian Li

    cs.LG

    Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction. We distinguish these quantities using an Expert Subspace Separation Index (ESSI), matched-route residuals, and a...

    arxiv.org/abs/2607.28308 · PDF

  19. 19

    Semi-Supervised Learning for Molecular Graphs via Ensemble Consensus

    Rasmus Tirsgaard, Laurits Fredsgaard, Marisa Wodrich, Mikkel Jordahn, Mikkel N. Schmidt

    cs.LG

    Machine learning is transforming molecular sciences by accelerating property prediction, simulation, and the discovery of new molecules and materials. Acquiring labeled data in these domains is often costly and time-consuming, whereas large collections of unlabeled molecular data are readily available. Standard semi-supervised learning methods often rely on label-preserving augmentations, which are challenging to design in the molecular...

    arxiv.org/abs/2607.28304 · PDF

  20. 20

    HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks

    Tiangang Li, Xiangbo Tian

    cs.LG

    Supervised fine-tuning (SFT) can equip large language models (LLMs) with domain knowledge for high-performance computing (HPC) tasks such as data race detection and benchmark question answering. However, knowledge alone does not guarantee task-appropriate behavior: the same SFT model that correctly classifies 88.65\% of C/C++ data race samples produces verbose, imprecise answers to factual queries, with 65.9\% of MLPerf responses exceeding 40...

    arxiv.org/abs/2607.28301 · PDF

  21. 21

    TopoFormer: Topology Meets Attention for Graph Learning

    Md Joshem Uddin, Astrit Tola, Cuneyt Gurcan Akcora, Baris Coskunuzer

    cs.LG · math.AT

    We introduce Topoformer, a lightweight and scalable framework for graph representation learning that encodes topological structure into attention-friendly sequences. At the core of our method is Topo-Scan, a novel module that decomposes a graph into a short, ordered sequence of topological tokens by slicing over node or edge filtrations. These sequences capture multi-scale structural patterns, from local motifs to global organization, and are...

    arxiv.org/abs/2607.28259 · PDF

  22. 22

    Persistent Gaussian Perturbations Prevent Oversmoothing in Recurrent Graph Neural Networks

    Mostafa Haghir Chehreghani

    cs.LG · cs.AI · cs.IT

    Oversmoothing is a fundamental limitation of deep graph neural networks (GNNs), where repeated message passing causes node representations to become increasingly similar, eventually collapsing toward a low-dimensional subspace. This phenomenon limits the effective depth of message-passing architectures and motivates the search for mechanisms that preserve representation diversity. In this paper, we study a recurrent graph neural network in...

    arxiv.org/abs/2607.28185 · PDF

  23. 23

    Multi-channel Uplift Policy Learning

    Changjian Liu, Tianyu Wang, Xiaoxuan Deng, WenTao Zhu, Yuwei Xu, Jungqi Jin, Yong Gao, Chuan Yu, Jian Xu, Bo Zheng

    cs.LG

    E-commerce platforms must allocate fixed marketing budgets across multiple channels to maximize business utility. However, standard predict-then-optimize (PTO) paradigms fail in this compositional space due to observational confounding and severe extrapolation. We formulate this challenge as a simplex-constrained uplift decision problem and propose ReAlloc, a fast-slow causal framework. Specifically, an agile Orthogonal Teacher extracts...

    arxiv.org/abs/2607.28182 · PDF

  24. 24

    Search Strategies for Optimal Classification and Regression Trees

    Jacobus G. M. van der Linden, Mim van den Bos, Emir Demirović

    cs.LG · cs.AI

    Optimal decision trees (ODTs) are compact, interpretable machine learning models that globally optimize a given objective, but their scalability remains challenging. While recent work has proposed a variety of search strategies to improve scalability, the precise contribution of each strategy remains unclear. To address this gap, we introduce a general algorithmic framework for ODTs that instantiates previously used search strategies and...

    arxiv.org/abs/2607.28170 · PDF

  25. 25

    LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning

    Mohand Mezmaz, Grégoire Danoy

    cs.LG

    Machine learning for combinatorial optimization typically relies on neural constructors trained via reinforcement learning on large offline datasets for a fixed problem class-incurring high pretraining costs and generalizing poorly outside the training distribution. We propose an alternative: a metaheuristic framework that reformulates the randomized constructive phase of GRASP as an online imitation learning task, trained from scratch on...

    arxiv.org/abs/2607.28135 · PDF

  26. 26

    Information Bottleneck Learning for Faithful Time Series Forecasting Explanations

    Xu Zheng, Wei Cheng, Zhuomin Chen, Mo Sha, Jingchao Ni, Dongsheng Luo

    cs.LG · cs.AI

    As forecasts increasingly drive decisions in fields such as energy, transportation, and healthcare, understanding the historical data behind these predictions has become as crucial as the predictions themselves. Although existing interpretable-by-design forecasters reveal their internal structures, they offer no guarantee that these structures faithfully reflect the underlying evidence driving the predictions. In contrast, while...

    arxiv.org/abs/2607.28124 · PDF

  27. 27

    From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE Inference

    Tianyang Zhu

    cs.LG

    Mathematically equivalent expert-reduction orders can produce observably different sparse-MoE executions. We isolate this effect in native DeepSeek-V4-Flash by freezing local MoE state and varying only aggregation semantics. Four schemes separate operand representation from accumulator precision. At one layer-5 fork, 720 A-mode orders yield 10 continuation basins; 720 B-mode orders form 360 exact structural classes and 11 basins. Under one...

    arxiv.org/abs/2607.28097 · PDF

  28. 28

    Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction

    Tianyou Bai, Huan Wang, Mingchen Gao, Fangyue Lin, Pinze Ren, Zhenlin Zhao, Siming Dong

    cs.LG · cs.AI

    Chemical property prediction plays a critical role in accelerating scientific discovery in chemistry, materials science, and drug development. However, existing benchmarks often suffer from limited task diversity, fragmented datasets, and inconsistent evaluation protocols, making it challenging to systematically assess the reliability and generalization of AI models. In this work, we introduce Chem World, a comprehensive benchmark for...

    arxiv.org/abs/2607.28079 · PDF

  29. 29

    GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios

    Yiming Xu, Jihua Kang, Chunsai Du, Qifan Zhang, Wangqiu Zhou, Yiting Wu, Tianqi Li, Qi Song

    cs.LG · cs.CV

    In demanding professional environments and meeting review scenarios, lengthy text often imposes a high cognitive load. To facilitate efficient information communication, transforming verbose text into logically clear diagrams is essential. Scalable Vector Graphics (SVG) provide an effective representation for this purpose due to their editability and resolution independence. However, current research on Text-to-SVG generation remains hindered...

    arxiv.org/abs/2607.28073 · PDF

  30. 30

    ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

    Xingjian Wu, Xuhang Zhu, Xingchen Liu, Junlin Liu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai

    cs.LG

    As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute failures to specific process deficiencies, hindering attribution in long-horizon tasks. In this work, we present ClawTrack, a dual-assessment benchmark that simultaneously measures what an agent achieves (Task...

    arxiv.org/abs/2607.28037 · PDF

  31. 31

    Learning features from Newton's algorithm: a way to accelerate nonlinear parametrized PDE solvers

    Rémy Vallot, Florian de Vuyst, Thibault Dairay, Mathilde Mougeot

    cs.LG · math.AP · math.NA

    It is well known that Newton's method converges faster when the initial guess is closer to a root of a system of nonlinear equations. In this paper, a two-stage Newton initial guess strategy is proposed by learning features from a parameter-space sampling and a database of precomputed solutions. The method uses discrete Newton trajectories to construct two complementary reduced spaces: a solution feature space, built from converged states,...

    arxiv.org/abs/2607.28036 · PDF

  32. 32

    Enhancing Irregular Time Series Forecasting with Continuous-Time Modeling Framework

    Tianen Shen, Zhengyu Li, Yutong Li, Xiangfei Qiu, Xingjian Wu, Bin Yang, Jilin Hu

    cs.LG

    Irregular multivariate time series are widely encountered in applications such as healthcare monitoring, human activity recognition, and environmental sensing. Their core challenges stem from asynchronous observations, non-uniform sampling intervals, and the fact that temporal patterns themselves carry critical dynamic information. Existing approaches either rely on discretization-based preprocessing (e.g., interpolation, imputation, or...

    arxiv.org/abs/2607.28035 · PDF

  33. 33

    Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

    Xingjian Wu, Junlin Liu, Xingchen Liu, Xuhang Zhu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai

    cs.LG

    Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). While OPSD provides dense, logit-level supervision, it inherently suffers from exposure bias due to the privileged information of the self-teacher. In multi-turn agentic settings, this leads to reasoning route convergence and the loss of clear optimization directions....

    arxiv.org/abs/2607.28026 · PDF

  34. 34

    Flux-OPD: On-Policy Distillation with Evolving Contexts

    Yuran Wang, Zekun Wang, Bohan Zeng, Ruixu Zhang, Wenxuan Liu, Liu Yang, Yifan Dai, Yang Shi, Bozhou Li, Chengzhuo...

    cs.LG · cs.AI

    Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and...

    arxiv.org/abs/2607.28022 · PDF

  35. 35

    Building a User Foundation Model for the Open Web

    Solal Vernier, Ivan Can Arisoy, Merwan Barlier, Blaž Škrlj

    cs.LG

    User foundation models have demonstrated strong results in e-commerce and social recommendation, but most industrial deployments assume environments where user identity is stable and persistent. Open-web real-time bidding (RTB) operates on a structurally different data distribution: user identity is fragmented and non-persistent across browsing sessions, and the availability of browsing history depends on user privacy choices. Consequently, a...

    arxiv.org/abs/2607.28019 · PDF

  36. 36

    It's All Just Vectorization: einx, a Universal Notation for Tensor Operations

    Florian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens

    cs.LG

    Tensor operations represent a cornerstone of modern scientific computing. However, the Numpy-like notation adopted by predominant tensor frameworks is often difficult to read and write and prone to so-called shape errors, i.a., due to following inconsistent rules across a large, complex collection of operations. Alternatives like einsum and einops have gained popularity, but are inherently restricted to few operations and lack the generality...

    arxiv.org/abs/2607.27987 · PDF

  37. 37

    Generalization Bounds on Optimal Control for Transformer Training and Wasserstein Distributional Robustness

    Kağan Akman, Naci Saldi, Serdar Yüksel

    cs.LG · math.OC

    We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions. Building on the doubly lifted, measure-valued formulation of Transformer dynamics, we view data sets as probability laws on pairs of empirical input-output measures, allowing us to interpret the training problem as a finite-horizon Markovian control problem. We then analyze a quantized model, derived by quantizing the state, action, and...

    arxiv.org/abs/2607.27975 · PDF

  38. 38

    TAPO: Transition-Aware Policy Optimization for LLM Agents

    Cong Li, Peixi Peng, Yisen Zhao, Xinyu Hu, Shudong Liu, Zhan Su, Zhuojian Li

    cs.LG · cs.AI

    Recently, Reinforcement Learning (RL) has emerged as a crucial paradigm for the post-training of Large Language Model (LLM) agents. However, existing methods predominantly rely on sparse task rewards for policy optimization, failing to fully exploit another class of inherently dense supervisory signals naturally present during online interaction: environmental feedback following action execution. Recent theoretical studies suggest that...

    arxiv.org/abs/2607.27973 · PDF

  39. 39

    Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning

    Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj, Gjergji Kasneci

    cs.LG

    Machine unlearning seeks to selectively remove specific knowledge from trained language models without full retraining, a growing necessity under privacy regulations such as GDPR and the EU AI Act. Recent work has reformulated unlearning as a Reinforcement Learning with Verifiable Rewards (RLVR) problem, where models are optimized against verifiable rewards computed directly from their outputs. However, existing methods rely on sparse binary...

    arxiv.org/abs/2607.27968 · PDF

  40. 40

    What Makes Graph Unified? Principles and Generative Sliding-Window Transformer for Graph Foundation Models

    Dongxiao He, Siqi Liu, Jitao Zhao, Yawen Li, Yi Wang, Di Jin

    cs.LG

    Graph Foundation Models (GFMs) have recently emerged as a promising paradigm for general-purpose graph learning, aiming to learn reusable knowledge that generalizes across diverse graph domains and downstream tasks, reducing the need for specific model development. Achieving this goal requires reconciling the substantial heterogeneity in node features, graph structures, and semantic information across domains. Among them, heterogeneous node...

    arxiv.org/abs/2607.27966 · PDF

  41. 41

    AutoPref: Automatic Discovery of Task-Specific Preference Objectives for Neural Combinatorial Optimization

    Shengda Gu, Kai Li, Xinyi Ke, Haobo Fu, Yifan Zhang, Jian Cheng

    cs.LG

    Combinatorial optimization problems (COPs) underpin many real-world decisions, but their exponentially large search spaces make high-quality solutions costly to obtain. Neural combinatorial optimization (NCO) learns fast construction policies, typically with reinforcement learning (RL), while preference-based NCO improves sample efficiency by learning from relative solution quality. However, existing preference objectives combine two distinct...

    arxiv.org/abs/2607.27953 · PDF

  42. 42

    TriShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement

    Cheng Wei

    cs.LG · cs.CL

    Federated fine-tuning of large language models (LLMs) enables collaborative training without exposing raw data. However, a recent attack, NeuroImprint [1] (arXiv:2606.20553), demonstrates that a malicious parameter server can corrupt a PEFT adapter into a privacy backdoor: by assigning a dedicated memorization neuron to each training sample and ensuring each neuron updates at most once, the server can analytically reconstruct 59\%--79\% of...

    arxiv.org/abs/2607.27940 · PDF

  43. 43

    Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting

    Xiang Yuan, Kaiqing Lei, Zhenyu Jin, Jun Shu, Deyu Meng, Zongben Xu

    cs.LG

    The performance of Large Language Models (LLMs) is fundamentally influenced by the distributional composition of multi-domain pre-training data. While manual heuristics were prevalent in early models, they increasingly fail to capture the intricate synergies between domains as data complexity grows. To overcome the issue, a dominant approach seeks to fit a proxy function mapping between domain weights and their corresponding validation...

    arxiv.org/abs/2607.27928 · PDF

  44. 44

    ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

    Dongxiu Liu, Haoyi Niu, Peng Cheng, Yuan Gao, Xirui Kang, Sangli Teng, Koushil Sreenath, Xianyuan Zhan

    cs.LG · cs.CV · cs.RO

    In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are largely confined to discrete-time prediction, thereby exhibiting significant inefficiency in capturing the dynamics of physical world. We introduce Physical-Time Flow (\textbf{PT-Flow}), a novel approach that learns a continuous latent velocity field operating in physical time. Crucially, the...

    arxiv.org/abs/2607.27924 · PDF

  45. 45

    Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control

    Takumi Shioda, Kohei Terashima, Tatsuo Nagai

    cs.LG · eess.SY

    Multi-zone variable-air-volume control must balance thermal comfort, indoor air quality, and electricity use across several continuous actuators. Model predictive control and reinforcement learning are widely studied, but deployment typically requires building-specific modeling or training, limiting scalability. We first test whether a frontier reasoning model (an LLM trained to use additional inference-time computation) can achieve...

    arxiv.org/abs/2607.27914 · PDF

  46. 46

    S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring

    Glenn Anta Bucagu, Thorir Mar Ingolfsson, Yawei Li, Luca Benini

    cs.LG

    Foundation models offer a promising paradigm for Electroencephalography (EEG) analysis, leveraging generalizable representations from vast unlabeled datasets. Yet, Transformer-based architectures face a critical bottleneck: global attention mechanisms couple the attention memory state to the signal duration, causing memory overflow during continuous monitoring. To address this, we introduce S-CEReBrO (Streaming CEReBrO), an evolution of the...

    arxiv.org/abs/2607.27913 · PDF

  47. 47

    Class-Aware Reinforcement Learning for Counterfactual Explanation Generation

    Muhammad Adil Saleem, Syed Ali Raza, Mary-Anne Williams

    cs.LG · cs.AI

    Counterfactual explanations (CFEs) enhance the interpretability of black-box models by generating alternative instances with adjusted feature values that achieve a contrastive outcome. Reinforcement learning (RL) offers a promising approach for CFE generation, enabling efficient exploration of counterfactual instances while ensuring control over key metrics like validity, sparsity, and proximity. Previous studies have formulated RL states...

    arxiv.org/abs/2607.27905 · PDF

  48. 48

    Contrastive Concept Importance: Explaining Pairwise Class Decisions Through Automatically Extracted Concept Representations

    Roel Visser, Isaac Roberts, Barbara Hammer

    cs.LG

    Concept-based explanations are a prevalent way to explain the decisions of complex black-box methods through semantically meaningful, human-interpretable concepts. To attribute the contribution of such concepts to a model's decisions, feature attribution methods are used to quantify how strongly each concept contributes to a model output. These attributions are typically computed for a single output class and therefore answer a...

    arxiv.org/abs/2607.27904 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.