cs.LG · 2026-07-30 · No. 69

Machine Learning, 2026-07-30.

68 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

68 entries
  1. 01

    Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

    Perry Dong, Ron Polonsky, Dorsa Sadigh, Chelsea Fin

    cs.LG

    Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized Q-function can result in highly performant and reliable policies without...

    arxiv.org/abs/2607.27203 · PDF

  2. 02

    From Classification to Regression: Using a Fruitfly to Solve Equations

    Shady E. Ahmed, Panos Stinis

    cs.LG · math.NA

    We present a novel approach to regression tasks using classification which is motivated by the mechanism used by fruitflies to sense their environment. Specifically, we formulate a general framework for learning nonlinear input-output relationships by replacing complex global surrogate models with a finite library of representative local patterns. Since scientific data often occupy limited and recurring regions of the input space, we generate...

    arxiv.org/abs/2607.27196 · PDF

  3. 03

    Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes

    Lennon J. Shikhman, Michael Galarnyk, Aadi Dash, Nicholas A. Welsh

    cs.LG · q-fin.CP · q-fin.PR · q-fin.ST

    Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A controlled benchmark exposes simulator-truth densities for latent evaluation, while a chronological NIFTY benchmark tests only held-out market prices. A two-component lognormal mixture has the lowest aggregate price, $L^1$, Wasserstein, and fixed-tail errors on the synthetic benchmark....

    arxiv.org/abs/2607.27188 · PDF

  4. 04

    When Do Learned Diffusion Proposals Help Constraint Solving? A Controlled Study on Continuous Algebraic Systems

    Quang Bui, Sparsh Roy, Akash Gundimeda, Davin Yin

    cs.LG

    Solving a continuous algebraic constraint system requires two decisions: which values satisfy the constraints, and which structural augmentation renders an unsolvable system solvable. Classical solvers answer the first well and the second only by enumeration. On that discrete decision, a candidate-conditioned repair ranker choosing among K augmentations reaches the exhaustive-search ceiling at a fraction of the calls, outperforming random...

    arxiv.org/abs/2607.27169 · PDF

  5. 05

    Skillful forecasting of offshore winds from satellite scatterometer constellations

    Francesco Pinto, Luca Lanzilao, Paco Lopez Dekker, Angela Meyer

    cs.LG

    Accurate intraday forecasts of offshore wind are becoming increasingly important for power system operation and the integration of growing shares of offshore wind energy. Operational forecasts rely predominantly on numerical weather prediction (NWP), which is not optimized for lead times of minutes to hours, where initial-condition accuracy dominates forecast skill. Although satellite scatterometer observations are routinely assimilated into...

    arxiv.org/abs/2607.27152 · PDF

  6. 06

    Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark

    Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal

    cs.LG · cs.AI

    High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage guarantees; however, we show that it severely under-covers rare, costly minority classes, with minority-class coverage dropping to as low as 0.5% on certain datasets. To...

    arxiv.org/abs/2607.27143 · PDF

  7. 07

    Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes

    Zuyuan Zhang, Yongshan Chen, Mahdi Imani, Tian Lan

    cs.LG

    An agent acting under partial observability must retain a recursively updateable statistic of history that restores the Markov property, but the smallest such statistic is generally unknown. We characterize this minimal Markov sufficient statistic for holonomy-cover decision processes, a structured POMDP class in which the visible dynamics are Markov and every realized visible transition applies a fixed permutation to a hidden mode. In...

    arxiv.org/abs/2607.27132 · PDF

  8. 08

    Voronoi Histograms for Adaptive Vectorization of Expected Persistence Diagrams

    Kaifeng Zhang, Kai Ming Ting

    cs.LG

    Persistence Diagram (PD) is known to capture point cloud topology effectively, but its computation has high time complexity. Expected Persistence Diagram (EPD) has been developed to reduce the time cost by studying the topology of multiple subsets of a point cloud and it serves as a distribution of topological features. Existing EPD vectorizations often rely on predefined point transformations, such as Gaussian or landscape functions. We...

    arxiv.org/abs/2607.27126 · PDF

  9. 09

    Hierarchical Spatio-Temporal Transformer for Coherent Emergency Department Forecasting

    Filipa Lino, Bárbara Tavares, Carlos Santiago, Cláudia Soares, Manuel Marques

    cs.LG

    Emergency Departments (EDs) are critical access points in healthcare systems, yet they face persistent pressure from unpredictable patient demand, seasonal surges, and non-urgent visits. Effective ED planning requires forecasts at multiple decision-making levels: hospitals need local demand estimates for staffing and bed management, regions require forecasts to coordinate healthcare units, and national authorities need system-wide projections...

    arxiv.org/abs/2607.27106 · PDF

  10. 10

    Sky sphere representation in language models

    Aleksandr Berdnikov, Yevgeny Liokumovich

    cs.LG

    We analyze whether language models of size ~100B have a representation of the night sky map that is decodable from their residual stream. We find that most of the considered open-source models do have such a representation, and it often even surfaces to the top principal components on prompts that ask questions like ``what is close to this object in the night sky''. In all but one model this representation showed significant scores in LOO...

    arxiv.org/abs/2607.27092 · PDF

  11. 11

    Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

    Yicheng Feng, Yan Zhang, Yan Cheng, Wei Qi

    cs.LG · cs.AI

    As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds cost, context load, and privacy exposure. Routers and retrievers can rank candidate tools by relevance, but a ranking alone does not determine how many are worth selecting. Existing approaches...

    arxiv.org/abs/2607.27083 · PDF

  12. 12

    Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering

    Nicolas Béreux, Aurélien Decelle, Cyril Furtlehner, Beatriz Seoane

    cs.LG

    Energy-Based Models (EBMs) provide an interpretable framework for generative modeling of scientific data, but poor Markov Chain Monte Carlo mixing often limits their reliability. We introduce a training algorithm based on Parallel Trajectory Tempering (PTT), which exploits the continuity of the optimization path to maintain equilibrium sampling throughout learning. This enables stable and fast training on highly multimodal and data-scarce...

    arxiv.org/abs/2607.27077 · PDF

  13. 13

    Single-Beat Cuffless Blood Pressure Estimation Using Ear-PPG and ECG with a Lightweight Hybrid Learning Framework

    Kindeep K. Dhatt, Tengyue Wu, Hanbang Hua, Yayun Du

    cs.LG · eess.SP · eess.SY

    Continuous cuffless blood pressure (BP) monitoring remains challenging due to motion artifacts, physiological variability, and the limited robustness of conventional pulse transit time (PTT) models under dynamic conditions. Many prior approaches rely on multi-second windows to stabilize estimation, an assumption that is frequently violated during real-world monitoring with intermittent signal corruption. Here, we show that discriminative...

    arxiv.org/abs/2607.27076 · PDF

  14. 14

    Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise

    Vaneet Aggarwal

    cs.LG · cs.AI · math.OC

    We study online convex optimization (OCO) in non-stationary environments under heavy-tailed noise, where the stochastic gradient oracle admits only a finite $p$-th central moment for some $p \in (1, 2]$. While static regret is well-understood, achieving universal dynamic regret in a parameter-free manner remains an open challenge. We resolve this by proposing \textbf{HT-PAder}, a parameter-free algorithm combining restarted AdaGrad experts...

    arxiv.org/abs/2607.27073 · PDF

  15. 15

    CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation

    Fengming Yu, Haiwei Pan, Kejia Zhang, Chunling Chen, Jian Guan, Baoying Ma

    cs.LG · cs.AI

    Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences in architectural inductive biases between the teacher and student models often result in substantial representation discrepancies, limiting the effectiveness of direct...

    arxiv.org/abs/2607.27054 · PDF

  16. 16

    Lottery Tickets Are Not Deployment Tickets

    Bum Jun Kim

    cs.LG · cs.CV

    Reports on how sparsification, compression, and lottery tickets change model behavior have been mixed in the prior literature, with beneficial effects observed in some studies and adverse effects in others. Moreover, prior work has not considered actual deployment conditions, where decision logic is already fixed for the incumbent. To assess these mixed findings from a practical standpoint, we study the production-replacement question at the...

    arxiv.org/abs/2607.27031 · PDF

  17. 17

    TreeCCA: Canonical Correlation Analysis via Gradient-Boosted Trees

    James Chapman

    cs.LG

    Gradient-boosted trees dominate tabular machine learning, yet canonical correlation analysis has always relied on linear or neural encoders. We propose \textbf{TreeCCA}, the first method to train gradient-boosted tree ensembles end-to-end as CCA encoders, inheriting their plug-and-play reliability: no architecture design, familiar hyperparameters, and strong performance with defaults. The technical enabler is the Eckart-Young (EY) loss, which...

    arxiv.org/abs/2607.27027 · PDF

  18. 18

    BayesAME: Bayesian Active Model Evaluation

    Paula Cordero Encinar, Taylan Cemgil, Arnaud Doucet, Virginia Aglietti, Silvia Chiappa

    cs.LG · cs.AI · stat.ML

    Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Current literature mostly requires the practitioner to input a coreset size. However, when reliable performance estimation takes priority over efficiency, an evaluation method should also be capable...

    arxiv.org/abs/2607.27023 · PDF

  19. 19

    What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

    Kaizhen Tan, Xin Xu, Siru Tao, Hanzhe Hong, Yang Feng, Heqing Du

    cs.LG · cs.RO

    A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a trained latent actually contain, and what decides this? We answer with controlled interventions in POKEWORLD, an interactive environment whose visually identical objects hide mass, drag, and contact stiffness. A certificate-gated protocol first certifies each parameter...

    arxiv.org/abs/2607.27017 · PDF

  20. 20

    Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark

    Peter Lorenz, Anjith George, Sébastien Marcel

    cs.LG

    Face presentation attack detection (PAD) remains challenging under cross-dataset evaluation, where domain shift degrades models trained on a single dataset. The scarcity of large-scale labeled data motivates adapting pretrained vision models rather than training task-specific architectures from scratch, raising a fundamental question: do general-purpose vision foundation models encode PAD-relevant information accessible with minimal...

    arxiv.org/abs/2607.26993 · PDF

  21. 21

    Surrogate assisted diversity estimation in neural ensemble search

    Alexandr Udeneev, Petr Babkin, Oleg Bakhteev

    cs.LG

    Ensembles are a standard way to improve the performance and robustness of deep neural networks, but their effectiveness crucially depends on both the quality and the diversity of individual models. Most neural architecture search (NAS) methods are computationally expensive. Extending them to neural ensemble search (NES), which requires joint optimization of individual architectures and their ensemble composition, leads to an exponential...

    arxiv.org/abs/2607.26940 · PDF

  22. 22

    Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method

    Chang Liu, Fei Suo, Yanzhou Jin, Yusuke Iwasawa, Yutaka Matsuo, Yaonan Zhu

    cs.LG · cs.RO

    Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning from pixels by regularizing the latent marginal distribution toward an isotropic Gaussian, thereby preventing representation collapse. While effective and elegant in single-task settings, this recipe does not extend reliably to multi-task training, leading to substantially worse downstream...

    arxiv.org/abs/2607.26924 · PDF

  23. 23

    Two Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local Language Models

    Ashish Prajapati, Om Mohite

    cs.LG

    Multi-agent LLM pipeline systems break down the task among multiple roles for better reasoning, but are benchmarked mainly with large-scale commercial models. In this study, we investigate Parishad, a structured multi-agent system involving five roles, by deploying it on Qwen2.5-7B-Instruct, a local model, on two datasets: GSM8K (500 questions) and HumanEval (164 questions), compared with prompting directly and two-call self-refinement. The...

    arxiv.org/abs/2607.26922 · PDF

  24. 24

    Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing

    Brandon Gower-Winter, Georg Krempl

    cs.LG · cs.AI

    In many domains such as Palliative Care, Credit Assignment and Recommender Systems, predictions may causally influence the outcomes they predict. This phenomena is known as Outcome Performativity. This paper formalises an approach for detecting Outcome Performativity using prediction intervention called Outcome Performativity A/B Detection (OPAB). OPAB enables the detection of Outcome Performativity by assessing the dissimilarity in outcome...

    arxiv.org/abs/2607.26908 · PDF

  25. 25

    ReCo: Reweighting GRPO Against Distributional Concentration

    Junoh Park, Junseo Hwang, Wonguk Cho, Taesup Kim

    cs.LG · cs.AI

    Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths. We find that this reduction is associated with GRPO concentrating on responses that the base model already generates with high probability. We...

    arxiv.org/abs/2607.26862 · PDF

  26. 26

    Amortized Moment Matching for Visual Generation

    Wenze Liu, Xintao Wang, Pengfei Wan, Xiangyu Yue

    cs.LG

    We propose amortized moment matching, utilizing neural networks to learn data moments as distributional training signals. By casting diffusion denoisers through polynomial projections, we establish a general framework for moment amortization, revealing that an $n$-th degree projection explicitly identifies data moments up to order $n+1$. Derived from the tractable affine case, we instantiate the Amortized Fréchet Distance (AMFD) loss. Unlike...

    arxiv.org/abs/2607.26860 · PDF

  27. 27

    TREA-Net: A Transferable Residual Epidemiological Adaptation Network for Dengue Incidence Forecasting

    Inesh Shukla, Madhurima Panja, Tanujit Chakraborty, Chittaranjan Hens

    cs.LG · q-bio.QM

    Accurate multi-week dengue forecasting supports timely vector-control interventions, outbreak preparedness, and healthcare resource allocation. However, newly established surveillance systems often lack the historical data needed to train reliable neural forecasting models. Although pretrained time-series models offer promising zero-shot forecasts, their cross-domain training may not capture local epidemiological dynamics. We propose...

    arxiv.org/abs/2607.26854 · PDF

  28. 28

    Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models

    Hua-Dong Xiong, Xinyuan Yan, Ji-An Li, Jingming Xue, Marcelo G. Mattar, Robert C. Wilson

    cs.LG

    Inference-time thinking improves the performance of large language models, but aggregate outcomes do not reveal whether models use available evidence more effectively or seek information that could improve future decisions. We distinguish these responses by measuring action preference, thinking length, and reported confidence under matched uncertainty. Ten open-weight models completed matched horizon-style two-armed bandit trials in thinking...

    arxiv.org/abs/2607.26845 · PDF

  29. 29

    Tight Generalization Bound for AdaBoost

    Mikael Møller Høgsgaard

    cs.LG

    In this paper we show that the generalization error of AdaBoost is $Θ\big(\tfrac{d\ln(nγ^{2}/d)}{nγ^2}+\tfrac{\ln(1/δ)}{n}\big)$, where $γ$ is the advantage guaranteed by the weak learner, $d$ is the VC-dimension of the class containing the weak hypotheses, $n$ is the sample size, and $δ$ is the confidence parameter. The contribution of this paper is the upper bound; the matching lower bound follows from prior work. The upper bound proof...

    arxiv.org/abs/2607.26838 · PDF

  30. 30

    Kairos: Numerically Robust News Recommendation under Item Cold-Start via Cholesky-based LinUCB

    Finn Hertsch

    cs.LG · cs.IR

    Algorithmic news personalization in regional markets often fails because modern deep learning models require massive interaction data while real-world news has a short Time-to-Live (TTL < 48 h) and shallow article pools. This structural item cold-start deprives collaborative filtering of the data needed for robust modeling. This paper presents Project Kairos, a framework that bridges this data scarcity through a contextual online learning...

    arxiv.org/abs/2607.26832 · PDF

  31. 31

    Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility

    Yansen Zhang, Yilu Liu, Tianyu Liu, Jiamin Chen, Xiaokun Zhang, Kai Xie, Xue Liu, Chen Ma, Yiyan Qi

    cs.LG · cs.AI

    Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls cause search actions to incur different token costs. We prove that cost-blind credit can forfeit all but a vanishing fraction of attainable quality as frontiers multiply...

    arxiv.org/abs/2607.26828 · PDF

  32. 32

    Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

    Shi Lin, Peng Qian, Dinghao Liu, Renjie Sun, Sifan Wu, Dezhang Kong, Chenpei Wang, Xun Wang

    cs.LG · cs.CR

    As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories. In multi-turn interactions, malicious intent can be decomposed across seemingly harmless turns and gradually reconstructed through interaction trajectories, eventually resulting in safety failures. Existing...

    arxiv.org/abs/2607.26820 · PDF

  33. 33

    FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning

    Zhaoyang Ma, Zhihao Wu, Xin Gao, Lipo Wang, Youfang Lin, Jing Wang

    cs.LG · cs.AI

    Federated learning (FL) enables collaborative learning over decentralized data silos without centralizing raw data. However, heterogeneous local architectures often induce non-aligned representation spaces, making it difficult to transfer global knowledge across silos. Existing paradigms share this knowledge as model parameters, distilled predictions, or class prototypes, yet all encode it in an absolute space that must be aligned across...

    arxiv.org/abs/2607.26801 · PDF

  34. 34

    SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

    Zhiyuan Yao, Yuxin Chen, Zhengxi Lu, Zishan Xu, Yueqing Sun, Yifu Guo, Yuquan Lu, Zhengzhou Cai, Kangning Zhang,...

    cs.LG · cs.AI

    Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution. We introduce SkillRise, a unified reinforcement learning framework for...

    arxiv.org/abs/2607.26784 · PDF

  35. 35

    Journey Operators for Structured Multi-Axis Composition

    Mahesh Godavarti

    cs.LG · cs.AI

    Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cells in a 3D volume. Along one axis, order matters: "the dog bit the man" is different from "the man bit the dog." Across independent axes, however, neither composition nor movement should depend on the order of axes: in an image, composing right then down should give the same result as composing down then...

    arxiv.org/abs/2607.26775 · PDF

  36. 36

    CalTwin: Towards Calibrated, Shift-Robust Medical World Models via Fisher-Information Regularisation

    Behraj Khan, Shabir Ahmad, Syed Ahmad Chan Bukhari, Tahir Qasim Syed

    cs.LG

    Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under interventions, supporting downstream tasks from imaging-based diagnosis to digital-twin treatment planning. Two failure modes threaten the reliability of such models in clinical deployment: (i)~\emph{covariate shift}, because training data are fragmented across hospitals, scanners, and time, so...

    arxiv.org/abs/2607.26752 · PDF

  37. 37

    Domain adaptation for handwriting trajectory reconstruction from IMU sensors

    Florent Imbert, Romain Tavenard, Yann Soullard, Eric Anquetil

    cs.LG

    Digital pens are commonly used to write on digital devices, providing the handwriting trace and enhancing human-computer interation. This study focuses on a digital pen equipped with kinematic sensors, allowing users to write on any surface while simultaneously preserving a digital trajectory of handwriting. This technology holds significant potential as a valuable educational tool, particularly in classrooms where it can facilitate the...

    arxiv.org/abs/2607.26736 · PDF

  38. 38

    PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems

    Kaiwen Jiang, Siya Xu, Ziyue Zhu, Chao Yang, Anh Tuan Luu, Haoran Luo

    cs.LG

    The rapid growth of AI workloads is turning data centers into large-scale, volatile, yet spatiotemporally flexible grid loads, creating an urgent need for coordinated electricity-computing scheduling. Under stringent grid constraints, schedules from general-purpose large language models (LLMs) are often infeasible, causing line-flow violations and unserved load. We present PowerAtlas, an LLM-agent framework for electricity-computing...

    arxiv.org/abs/2607.26710 · PDF

  39. 39

    Mixture-of-experts for handwriting trajectory reconstruction from IMU sensors

    Florent Imbert, Eric Anquetil, Yann Soullard, Romain Tavenard

    cs.LG

    The use of digital pens for online handwriting trajectory reconstruction is a prevalent method for human-computer interaction. In this study, we focus on a digital pen equipped with sensors where we aim at reconstructing the online handwriting trajectory. This pen enables writing on any surface and preserving the digital trace of handwriting. This type of pen could be used as an aid to learning to write in classroom. In this paper, we propose...

    arxiv.org/abs/2607.26708 · PDF

  40. 40

    Universality and Approximation Rates of Graph Neural Networks with Random Features

    Lukas Gonon, Thilo Meyer-Brandis, Niklas Weber

    cs.LG · stat.ML

    We investigate message-passing graph neural networks with random node features. Random node features are known to enhance the expressiveness of graph neural networks (GNNs) both theoretically and empirically. Here, we establish a novel universality result focusing on permutation-equivariant neural networks (PENNs), a class of GNNs built from feedforward neural network components that subsumes many prominent GNN architectures. We show that...

    arxiv.org/abs/2607.26699 · PDF

  41. 41

    Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL

    Mingxuan Che, Tsung-Yuan Tseng, Theresa Eimer, Marius Lindauer, Alexander von Rohr

    cs.LG · cs.AI

    Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks. However, RL outcomes can be highly stochastic, and both expected performance and variability often depend on hyperparameter (HP) configurations. We propose efficient and risk-averse heteroscedastic Bayesian Optimization (ERAHBO), a Bayesian optimization method that models both the mean and variance of learning outcomes as functions of the HP...

    arxiv.org/abs/2607.26680 · PDF

  42. 42

    AIGen: Automating AI Bill of Materials Generation Through Hybrid MLOps Integration

    Federica Pepe, Daniele Bifolco, Costantino Martignetti, Aureliano D'Amici, Fabiano Izzo, Damian A. Tamburri,...

    cs.LG

    The responsible development and deployment of artificial intelligence (AI) systems requires rigorous documentation of their constituent artifacts, e.g., datasets, model weights, training pipelines, and runtime dependencies. Although the Software Package Data Exchange (SPDX) 3.0 standard introduced native support for AI and dataset profiles, practical tooling capable of generating standards-compliant AI Bills of Materials (AIBoMs) in an...

    arxiv.org/abs/2607.26652 · PDF

  43. 43

    RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment

    Hansi Karunarathna, Nirhoshan Sivaroopan, Chamara Madarasingha, Anura Jayasumana, Kanchana Thilakarathna

    cs.LG · cs.IR

    Human Activity Recognition (HAR) from wearable sensors supports applications in healthcare, rehabilitation, fitness tracking, and smart environments. Yet, existing deep learning approaches require dataset-specific training, large labeled corpora, and repeated adaptation to new sensor settings or activity taxonomies. Retrieval-Augmented Generation for Human Activity Recognition (RAG-HAR) addresses this by framing HAR as a training-free,...

    arxiv.org/abs/2607.26631 · PDF

  44. 44

    Understanding Context Sampling in TabPFN on Small Tabular Datasets

    Mohammed Abdullah

    cs.LG · cs.AI

    TabPFN performs classification through in-context learning: it conditions on a set of labeled training rows (the context, or prototypes) and predicts test labels without gradient updates. On small tabular datasets, practitioners must still choose the context size and which rows constitute the context. We study how these choices affect prediction stability, accuracy, and selection cost using repeated context sampling on 15 OpenML datasets....

    arxiv.org/abs/2607.26628 · PDF

  45. 45

    Enhancing Automated Machine Learning via Homogeneous Train-Test Splitting Methods

    Yearn Tan Yin Tze, Charles Grellois

    cs.LG

    Accurate model evaluation in machine learning depends critically on how datasets are split into training and testing subsets. Standard random splitting assumes that both partitions share the same underlying distribution, an assumption often violated in datasets with class imbalance, natural clustering, or spatial autocorrelation. This paper investigates the role of statistical similarity in train-test splitting and its consequences for AutoML...

    arxiv.org/abs/2607.26625 · PDF

  46. 46

    FedWeave: Rethinking the Unit of Specialization in Heterogeneous Federated MoE-LoRA

    Donghang Duan, Xu Zheng, Lizong Zhang, Chong Mu, Meng Han

    cs.LG · cs.CL

    Federated PEFT enables LLMs to collaboratively adapt to decentralized private data without sharing raw examples. However, task heterogeneity across clients can cause cross-task interference and gradient conflicts during aggregation. Federated MoE-LoRA addresses this challenge through specialized LoRA experts and conditional routing. Yet existing methods typically specialize at client granularity, implicitly assuming task-coherent clients. Our...

    arxiv.org/abs/2607.26618 · PDF

  47. 47

    Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation

    Jialu Xu, Mengkun Liang, Guannan Liu, Xiaojie Mao, Junjie Wu

    cs.LG

    Estimating heterogeneous treatment effects is central to targeted interventions, such as personalized promotions and precision medicine. We focus on the conditional average treatment effect (CATE), a standard estimand for characterizing such heterogeneity. Even under standard identification conditions, finite-sample CATE estimation requires learning the nuisance structure for covariate adjustment and treatment-effect heterogeneity, often...

    arxiv.org/abs/2607.26599 · PDF

  48. 48

    Benchmarking ConvLSTM for One-Day-Ahead IMDAA Rainfall-Field Prediction across Four Indian Cities

    Tanmay Ghosh, Shaurabh Anand, Rakesh Gomaji Nannewar, Nithin Nagaraj

    cs.LG

    Convolutional long short-term memory networks (ConvLSTMs) are widely used for precipitation forecasting, but most evidence for their performance comes from dense, high-frequency radar sequences. This study tests whether convolutional recurrence improves one-day-ahead rainfall-field prediction on small daily reanalysis grids. Indian Monsoon Data Assimilation and Analysis (IMDAA) fields for June-September 1998-2020 were analysed for Bengaluru,...

    arxiv.org/abs/2607.26581 · PDF

  49. 49

    Simultaneous Coverage and Efficiency Guarantee in Online Conformal Prediction

    Rahul Vaze

    cs.LG · cs.DS

    Adaptive conformal inference (ACI) of Gibbs and Cand{è}s and its variants are the standard approach to online conformal prediction under distribution shift, but they suffer from three fundamental limitations. First, their guarantees control only the \emph{signed} long-run coverage error: persistent miscoverage in one direction can be masked by compensating errors later, so a method can satisfy the theoretical guarantee while being badly wrong...

    arxiv.org/abs/2607.26577 · PDF

  50. 50

    From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs

    Tina Vartziotis, Rodopi Kosteli, Elli Vartziotis, George Dasoulas, Michael Keckeisen, Konstantinos Skianis, Sotirios...

    cs.LG · cs.SE

    The operational energy consumption of large language model (LLM) inference is becoming an increasingly important component of the environmental footprint of deployed AI systems. However, direct measurement of inference energy often requires hardware telemetry, power instrumentation, or infrastructure-specific monitoring, limiting its applicability in comparative studies, early-stage system design, and sustainability reporting. This report...

    arxiv.org/abs/2607.26571 · PDF

  51. 51

    AgentGFM: A Graph Foundation Model with Node-Agent Information-Flow Control

    Jingbo Cui, Jitao Zhao, Di Jin, Dongxiao He

    cs.LG · cs.AI

    Graph Foundation Models (GFMs) aim to learn transferable knowledge from multi-domain graphs and adapt to unseen scenarios. As a fundamental source of relational semantics in graphs, the transferability of topological patterns has long been central to GFM research. However, local structural patterns may vary across graphs and even among nodes within the same graph. Despite such structural variation, most existing GFMs rely on manually designed...

    arxiv.org/abs/2607.26533 · PDF

  52. 52

    The Art of Not Forgetting A Local Learning Architecture for Continual Learning

    Ashmith Atmuri, Yashaswini Rao Bhogarajula

    cs.LG · cs.AI

    We introduce CMP (Cognitive Memory Primitive), a continual-learning architecture that repre?sents inputs as sparse relational codes, stores them in a two-tier competitive memory, and learns through local updates without end-to-end backpropagation through its feature-generating system. We investigate whether combining sparse representations, local learning, and persistent memory can reduce catastrophic forgetting relative to conventional...

    arxiv.org/abs/2607.26523 · PDF

  53. 53

    From Unsupervised Subgroups to Hypothetical State-Intervention Policies: An Evaluation of Selected Subgrouping Methods in Observational Health Data

    Vasundhara Acharya, Bulent Yener

    cs.LG

    Conventional subgroup analyses can yield unstable and difficult-to-interpret conclusions, especially in observational biomedical data where each individual is observed under only one exposure state, true individual treatment effects are unavailable, and causal structure is uncertain. We investigate whether subgroups constructed from pretreatment characteristics, without using exposure, outcome, or estimated treatment-effect information, can...

    arxiv.org/abs/2607.26521 · PDF

  54. 54

    HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

    Hei Yi Mak, Shadan Golestan, Hoang Le, Mehran Taghian Jazi, Yunke Peng, Yaoyuan Wang, Yao Wang, Junsong Wang,...

    cs.LG · cs.AI

    We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. A systematic study reveals that the dominant source of degradation in FP4 RL is not training-side quantization error but rollout activation quantization: outliers stretch the dynamic range so far that a large number of activation values underflow to...

    arxiv.org/abs/2607.26515 · PDF

  55. 55

    Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning

    Gong Gao, Xiao Lai, Ziqi Xie, Guojie Chen, Xianhui Liu, Weidong Zhao

    cs.LG · cs.AI

    Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement. However, temporal-difference (TD) learning introduces noisy targets, resulting in non-stationary optimization, while greedy policy updates amplify early-stage estimation errors. The recursive propagation of such errors leads to persistent overestimation bias and degraded training stability...

    arxiv.org/abs/2607.26509 · PDF

  56. 56

    From Interface to Inference: Eliciting Any-Order Inference from Any-Order Models

    Seunggeun Kim, Jaeyeon Kim, Taekyun Lee, Yuyuan Chen, Yilun Du, Sham Kakade, Sitan Chen

    cs.LG

    Many discrete reasoning tasks, such as code generation, are inherently non-causal: programmers move between high-level structure and local details, a process we call any-order inference. For autoregressive language models, which lack a native any-order interface, non-causal abilities such as infilling and next-edit prediction require hand-designed mechanisms. Can we instead design models that natively support any-order inference? Masked...

    arxiv.org/abs/2607.26504 · PDF

  57. 57

    From Conceptual Hydrologic Models to Conceptually Interpretable Neural Networks: A Snow-Water Mass-Conserving-Perceptron Framework for Discovering Catchment-Scale Precipitation-Storage-Runoff Representations

    Yuan-Heng Wang, Hoshin V. Gupta

    cs.LG

    The Mass-Conserving Perceptron (MCP) establishes a modeling paradigm in which conceptual hydrologic models can be reformulated as physically constrained, conceptually interpretable neural networks. Here, we develop a snow-water MCP network framework and evaluate it across 513 CAMELS-US basins. We first recast a coupled two-state SOIL-MCP and SNOWMCP conceptual model as a mass-conserving neural network and show that the hydrologic-model and...

    arxiv.org/abs/2607.26492 · PDF

  58. 58

    Conformal Changepoint Localization and Root Cause Analysis with Corrupted Observations

    Seunghun Yu, Meiyi Zhu, Petar Popovski, Joonhyuk Kang, Osvaldo Simeone

    cs.LG · eess.SP

    Detecting when the statistical behavior of an engineered system changes, and identifying which component is responsible, are core problems in the monitoring of telecommunication networks, robotic platforms, security infrastructure, and multi-agent systems. In safety- and mission-critical deployments, such decisions must be accompanied by statistical reliability guarantees rather than by point estimates alone. Conformal changepoint...

    arxiv.org/abs/2607.26481 · PDF

  59. 59

    Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement

    Haifeng Wu

    cs.LG · cs.CL

    Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on explicit preference supervision such as pairwise comparisons or demographic attributes, limiting their applicability in natural interaction settings. We propose IRIS, a framework that learns dynamic user personas directly from implicit interaction streams by extracting behavioral signals from...

    arxiv.org/abs/2607.26473 · PDF

  60. 60

    Neural Architecture Search for Traffic Prediction: A Survey of Methods, Challenges, and Future Directions

    Truong Giang Vu, Li Yang, Richard W. Pazzi

    cs.LG · cs.NE

    Traffic prediction is a core task in intelligent transportation systems, supporting applications such as adaptive signal control, route guidance, and ride-hailing dispatch. Deep learning models, including graph convolutional networks, recurrent networks, and Transformers, achieve strong results on standard benchmarks, but their architectures are designed by hand, requiring significant expert effort and producing models that often generalize...

    arxiv.org/abs/2607.26467 · PDF

  61. 61

    DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning

    Shuhang Wang, Ziming Li, Hui Cheng

    cs.LG

    Reinforcement learning is a natural post-training paradigm for code-oriented large language models because generated programs can be evaluated through parsing, execution, unit tests, and structural analysis.However, existing methods often rely on sparse outcome rewards or statically combine heterogeneous dense signals, even though syntax validity, executability, functional correctness, and structural organization describe different and...

    arxiv.org/abs/2607.26457 · PDF

  62. 62

    Existence-Field Diffusion Model for Spatial Point Processes with Variable Cardinality

    Xiaoyin Pan, Christian R. Shelton, Rakshith Mahishi, Chengkuan Hong

    cs.LG · stat.ML

    We study generative modeling of spatial point processes (SPP), where both the number of points and their spatial configuration are governed by a joint distribution. While diffusion models have achieved strong performance in modeling complex distributions, extending them to variable-cardinality SPP remains challenging. Existing approaches either decouple the modeling of cardinality and spatial structure, or rely on discrete trans-dimensional...

    arxiv.org/abs/2607.26428 · PDF

  63. 63

    SCOUT: Per-Context Reset Curricula for Sparse-Reward Reinforcement Learning

    Siddharth Aphale, Ayushman Singh

    cs.LG

    Sparse-reward reinforcement learning often fails because rollouts from the unassisted evaluation start rarely reach later task stages. Reset curricula address this by starting some training rollouts from easier intermediate states, called scaffolds. Such a curriculum faces two decisions: scaffold access, obtaining informative starts, and scaffold allocation, deciding how quickly that assistance is removed. Most prior curricula pace removal on...

    arxiv.org/abs/2607.26417 · PDF

  64. 64

    Examining the Efficacy of Graph Neural Network Message-Passing in Regression Contexts

    Keith G. Mills, Aedan J. DeFrates, Joong Ho Kim

    cs.LG

    Graph Neural Networks (GNN) facilitate effective prediction on graph data such as molecules, media networks and neural network blueprints. GNNs facilitate prediction through message passing techniques which define how information flows from a node to its neighbors. Due to the ubiquity of the graph data type, the development of newer and better GNNs has garnered much interest in the machine learning community. However, GNN evaluation and...

    arxiv.org/abs/2607.26404 · PDF

  65. 65

    Flow Map Learning via Nongradient Vector Flow

    Mark Goldstein, Anshuk Uppal, Raghav Singhal, Aahlad Puli, Rajesh Ranganath

    cs.LG

    Diffusion and flow-based models benefit from simple regression losses, but inference incurs significant overhead because sampling requires integration. Consistency models address this by directly learning the flow maps along the ODE trajectory, opening a design space between one-step and many-step approaches. However, existing methods face computational challenges such as requiring model inverses or backpropagation through iterated model...

    arxiv.org/abs/2607.26398 · PDF

  66. 66

    Q-Steer: Action-Value Guidance for Molecular Policy Optimization

    Xinyu Wang, Jinbo Bi, Minghu Song

    cs.LG · q-bio.BM

    Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions. This delayed-feedback interface makes molecular policy optimization myopic: an optimizer can learn that a molecule was good without knowing which intermediate actions made it good. We introduce Q-Steer, a rollout-time action-value steering primitive for molecular language models. Q-Steer...

    arxiv.org/abs/2607.26391 · PDF

  67. 67

    ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling

    Yiwen Chen, Joshua Ainslie, Krzysztof Choromanski, Xiang Gao, Su-Lin Wu, Yiping Yuan, Qian Sun

    cs.LG

    Rotary Position Embedding (RoPE) has been widely adopted in transformer-based large language models. However, its log-linear frequency schedule, originally designed to produce long-term attention decay, limits its adoption in domains with more complex distance-correlation patterns, such as temporal periodicity in sequential recommendation. We investigate the expressiveness of general query/key rotations and find that any normalized continuous...

    arxiv.org/abs/2607.26369 · PDF

  68. 68

    Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning

    Keegan Harris, Brian W. Lee, Ian Waudby-Smith, Philip Amortila, Nika Haghtalab, Michael I. Jordan

    cs.LG · cs.AI · cs.GT

    Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient. In practice, the coefficient is typically chosen heuristically or via hyperparameter...

    arxiv.org/abs/2607.26358 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.