cs.LG · 2026-08-19 · No. 89

Machine Learning, 2026-08-19.

59 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

59 entries
  1. 01

    The concentration game: Bayesian updating, regret, and information

    Akshay Balsubramani

    cs.LG · cs.GT · math.PR · math.ST

    We give a two-player zero-sum repeated game between a learner and nature whose value identity generates Bayesian updating and an exact accounting of exponential-weights regret at once, and supplies the comparator-class variational form that a wide class of concentration phenomena share. The terminal payoff is the most a comparator can gain at fixed relative entropy from the prior, and the one-step constraint is an information budget on...

    arxiv.org/abs/2608.18061 · PDF

  2. 02

    Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization

    Travis Zhang, Christian Belardi, Justin Lovelace, Jin Peng Zhou, Saebyeol Shin, Carla P. Gomes, Kilian Q. Weinberger

    cs.LG · cs.CV

    Sampling from a diffusion model typically requires many forward passes through a large neural network, making generation computationally expensive. While much work has focused on efficient solvers and samplers, comparatively little attention has been paid to selecting the sampling timesteps themselves. A recent line of work optimizes theoretically derived surrogates for sample quality rather than the quality metric itself. We propose...

    arxiv.org/abs/2608.18040 · PDF

  3. 03

    TabNSM: Neural Sparse Mixer for Tabular Regression

    Ali Eslamian, Qiang Cheng

    cs.LG · cs.CE

    Large-scale, high-dimensional tabular regression remains challenging: tree-based models are robust but lack end-to-end representation learning, while deep models enable flexible feature learning but often incur costly interaction modeling and sensitivity to noisy or redundant features. We propose TabNSM, a scalable regression framework that extends our earlier sparse-attention and mixer architectures. At its core, the Adaptive Sparse...

    arxiv.org/abs/2608.18026 · PDF

  4. 04

    Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System

    Yi Wang

    cs.LG · cs.AI · cs.SD

    GPT-style models achieve strong performance by representing language with finite vocabularies of reusable discrete tokens. This success has motivated symbolic music tokenizations to treat recurring musical structures, such as chords, motifs, and phrases, as reusable units analogous to linguistic tokens. However, tokenization derives its advantage not from reusable combinations alone, but from compression: effective compression requires...

    arxiv.org/abs/2608.18025 · PDF

  5. 05

    Revisiting WEASEL 2.0: Reproduction, Sensitivity, and an Adaptive Ensemble-Size Rule

    Cian Higgins, Gerard Carrigan, Pinar Sungu Isiacik, Georgiana Ifrim

    cs.LG

    WEASEL 2.0 is a dictionary-based time series classifier that combines dilated sliding windows with a randomised hyperparameter ensemble and a fixed-size dense feature representation. Two of its hyperparameter choices, the maximum ensemble size and the maximum window size, are specified by simple thresholding rules whose chosen thresholds are not empirically justified in the original paper. In this work we reproduce WEASEL 2.0 on 114 UCR...

    arxiv.org/abs/2608.18021 · PDF

  6. 06

    Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

    Christophe D. Hounwanou, John Emeka Eze, Yaé U. Gaba

    cs.LG · cs.AI

    Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. We formalize the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and show that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the...

    arxiv.org/abs/2608.18008 · PDF

  7. 07

    Composing Flow-Matching Energies with Known Physics: Generation, OOD Detection, and Inversion on PDE Fields

    Yixuan Sun, Anirban Samaddar, Sandeep Madireddy

    cs.LG · physics.comp-ph

    Probabilistic modeling of physical fields benefits from both a data-driven prior and known physical structure such as the governing equations. Energy-based models (EBMs) are a natural fit since energies compose additively, which enables augmenting physics information during inference. However, EBMs have been difficult to train and sample from due to the intractable partition function. We show in this work that flow matching models with a...

    arxiv.org/abs/2608.18004 · PDF

  8. 08

    Recirculation

    Michael C. Mozer, Shoaib Ahmed Siddiqui, Danny Sawyer, Sunny Sanyal, Rosanne Liu

    cs.LG

    We describe an inference-time architectural enhancement for off-the-shelf foundation models that markedly reduces perplexity and boosts accuracy across generation and reasoning tasks. Our approach incurs essentially no additional latency during generation, though it requires serial processing in the prefill phase. Motivated by the fundamental limitation that state updates in feedforward transformers are bounded by model depth, our technique,...

    arxiv.org/abs/2608.17981 · PDF

  9. 09

    Evaluating and improving crop-yield forecasting methods during extreme drought

    Shrey Gupta, Yi Ming, George Mohler

    cs.LG

    The impact of climate variability on food production has led to the creation of various forecasting models that uses machine learning (ML), numerical weather predictors (NWP) or a hybrid of ML-NWP models to identify structural and physical relationships between meteorological drivers and crop growth, in order to predict crop yield. Droughts, for example the 2012 Midwestern US (Corn Belt) drought, are extreme events that affect crop production...

    arxiv.org/abs/2608.17971 · PDF

  10. 10

    Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

    Bin Li, Dongdong Wang, Siyang Lu

    cs.LG · cs.AI · cs.SE

    Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous...

    arxiv.org/abs/2608.17965 · PDF

  11. 11

    Understanding the Surprising Generalization Properties of Tabular Foundation Models

    Nour Shaheen, Junwei Ma, Alex Labach, Frank Hutter, Valentin Thomas, Anthony L. Caterini

    cs.LG

    Tabular Foundation Models (TFMs) increasingly rely on in-context learning, where a model receives labelled examples at inference time and predicts labels for new inputs without updating its weights. Existing TFMs are typically trained on either massive synthetic corpora or very large collections of real datasets. In contrast, we show that surprisingly strong transfer can emerge from self-supervised pre-training on just a single real table. In...

    arxiv.org/abs/2608.17957 · PDF

  12. 12

    An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

    Javier Aguilar Martín

    cs.LG · cs.AI · eess.SY

    In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and the model is accepted when it reproduces sampled transitions. We ask what that acceptance certifies in continuous control. We define the pipeline's danger as an expected risk and isolate its exact factor: the probability that N i.i.d. gate rollouts all miss a critical event of probability r is exactly (1-r)^N; an independent...

    arxiv.org/abs/2608.17956 · PDF

  13. 13

    SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE

    Xuan Zheng, Kento Uchida, Shinichi Shirakawa

    cs.LG · cs.AI

    Recent research has leveraged Large Language Models (LLMs) to enhance Automated Feature Engineering (AutoFE) through semantic descriptions and trajectory-based prompting. However, there exist two challenges that limit their applicability and scalability in long-horizon optimization: (1) semantic metadata is unavailable in many practical settings, and (2) trajectory accumulation increases the risk of exceeding the context window, while without...

    arxiv.org/abs/2608.17948 · PDF

  14. 14

    Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation

    Zhizhao Liu, Zhiliang Tian, Xi Wang, Zhihua Wen, Yihang Xiong, Zhiquan Lai, Dongsheng Li

    cs.LG · cs.AI · cs.CL

    Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same exploration budget to samples with different difficulty levels is inefficient: easy samples may receive redundant rollouts, whereas difficult but learnable samples may receive too little exploration. Existing adaptive schedulers address this mismatch through...

    arxiv.org/abs/2608.17941 · PDF

  15. 15

    Hybrid ML for Lightweight Pre-Route Delay Estimation in Open-Source IC Design

    Marvin Castro Castro, Erick Carvajal Barboza

    cs.LG

    Static Timing Analysis (STA) is a critical step in the design flow of digital integrated circuits, however, obtaining accurate delay estimations can represent a challenge when limited information regarding physical design is available. In response, this work presents a hybrid and light-weight machine learning (ML) based approach that combines a decision tree with linear regression to improve pre-routing delay estimations generated by the...

    arxiv.org/abs/2608.17914 · PDF

  16. 16

    Dynamic Compression in Recurrent Networks

    Jyothish Pari, Ryan Bahlous-Boldi, Pulkit Agrawal

    cs.LG

    Recurrent models process long contexts efficiently by compressing their history into a fixed-size state, but modern architectures typically do so in a single causal pass over the sequence. Each input must therefore be compressed before the model knows how it will later be used, forcing a limited state to compromise across possible future demands. We introduce dynamic compression, which allows a recurrent model to selectively revisit past...

    arxiv.org/abs/2608.17896 · PDF

  17. 17

    Efficient Resource Optimization for Split Federated Learning

    Wei Wei, Xianhao Chen

    cs.LG

    Split federated learning (SFL) has emerged as a powerful paradigm for model training at the edge. However, SFL inherently involves discrete decision variables for model splitting and resource allocation, resulting in a challenging mixed-integer problem. Consequently, prior optimization schemes for SFL are either \textit{heuristic} or \textit{computationally inefficient}, which cannot handle large-scale user populations. To address this...

    arxiv.org/abs/2608.17849 · PDF

  18. 18

    MoRAX: Mobility-based Representation Augmentation for Geospatial Foundation Models

    Ya Wen, Jixuan Cai, Yulun Zhou, Alec Kirkley

    cs.LG · cs.SI

    Geospatial Foundation Models (GFMs) are emerging as a powerful paradigm for learning semantically rich and geographically consistent visual and physical representations. However, their reliance on Earth-observation (EO) data leaves information about human activity largely underrepresented. Human mobility data reveals the functional and relational structure between regions that is missing from EO data, but is often limited only to the city...

    arxiv.org/abs/2608.17848 · PDF

  19. 19

    Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs

    Roman Maksimov, Vladimir Aletov, Vladimir Solodkin, Dmitry Bylinkin, Daniil Medyakov, Aleksandr Beznosikov

    cs.LG

    As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a novel white-box attack inspired by locate-then-edit approaches from the field of Knowledge Editing. Our choice is motivated by the observation that models edited with such schemes tend to assign unusually high prediction probabilities to the edit target, a property that is particularly...

    arxiv.org/abs/2608.17836 · PDF

  20. 20

    MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure

    Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil, Subasish Das

    cs.LG · cs.AI · cs.HC

    Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. To address this gap, we introduce a large-scale dataset of over 129,000 labeled multivariate time-series sequences from 153 simulator rides by 51 participants under No, Low, and High TP, capturing 64 features across vehicle dynamics, control inputs,...

    arxiv.org/abs/2608.17823 · PDF

  21. 21

    An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

    Rubén Balbastre, Juan Manuel Orduña, Mariano Pérez

    cs.LG · cs.CL

    Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting,...

    arxiv.org/abs/2608.17804 · PDF

  22. 22

    Fourth-Moment Geometry of Rademacher Sums

    Peigan Gao, Jian Qian

    cs.LG · math.PR

    Let $\varepsilon_1,\ldots,\varepsilon_n$ be independent Rademacher signs and let $a=(a_1,\ldots,a_n)\in\R^n$ satisfy the normalization below. For the normalized Rademacher sum, we determine how its higher moments depend on the fourth-order mass. Combining a sharp fixed-q moment envelope with a separate argument below the convexity threshold gives the Gaussian stability inequality for the full range $p\geq4$ of this linear-in-q bound. The same...

    arxiv.org/abs/2608.17802 · PDF

  23. 23

    Debate Training Reduces Reward Hacking in RLAIF

    Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards,...

    cs.LG

    We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking compared to a reinforcement learning from AI feedback (RLAIF) baseline. Reward hacking is a central obstacle in RLAIF: as training progresses, the policy learns to exploit systematic errors in its AI judge, degrading task performance, a problem that worsens precisely...

    arxiv.org/abs/2608.17776 · PDF

  24. 24

    Training-Free Human-in-the-Loop Anomaly Detection via Memory Bank Correction

    Ayusha Abbas, Saram Abbas, Kabita Adhikari

    cs.LG

    Anomaly detectors are hardest to deploy exactly where training data is scarcest: a newly commissioned production line has a handful of verified "golden" samples and no machine-learning engineer on the factory floor. We present a training-free human-in-the-loop framework in which a domain expert corrects a PatchCore detector by direct memory bank editing: no retraining, no gradients, no original training data. A false-positive correction...

    arxiv.org/abs/2608.17775 · PDF

  25. 25

    MAGPIE-Net: Predicting short-duration heavy-rainfall events in station neighborhoods from multitemporal FY-4A AGRI observations

    Xiang Lin, Yunying Li, Chengzhi Ye, Zitong Chen, Jing Sun

    cs.LG

    Short-duration heavy-rainfall warning determines whether 1 h rainfall will exceed a threshold within a target-station neighborhood over the next few hours. Multitemporal infrared and water-vapor observations from the Fengyun-4A Advanced Geostationary Radiation Imager (FY-4A AGRI) capture cloud-top cooling, moisture evolution, and cloud expansion before substantial surface rainfall develops. However, most deep-learning nowcasting methods...

    arxiv.org/abs/2608.17753 · PDF

  26. 26

    Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment

    Zhen Zhang, Ahmad Hafez, Amr Alanwar

    cs.LG

    Agent evaluations and trace-based learning often compare outputs across transformed views through a post-response correspondence treated as neutral preprocessing. We show that this correspondence is a measurement intervention: omitting it can manufacture sensitivity, an over-aggressive map can manufacture invariance, and multiple optimal correspondences can leave mechanism labels and signed learning credit unidentified. We develop a validity...

    arxiv.org/abs/2608.17713 · PDF

  27. 27

    Conformal Prediction for Molecular Properties under Label Shift

    Hyeonsu Lee, Juyeon Kim, Erkhembayar Jadamba, Seungjin Choi, Hyunjin Shin

    cs.LG

    Drug discovery and development underpins healthcare but remains costly and failure-prone. A critical bottleneck lies in predicting molecular properties such as solubility, potency, and toxicity, which directly determine whether a candidate can advance from preclinical to clinical trials. Artificial Intelligence (AI) has accelerated this process, yet its reliability is often undermined by distribution shift, as experimental conditions...

    arxiv.org/abs/2608.17678 · PDF

  28. 28

    Picard Proximal Monte Carlo for Parallel Bayesian Imaging with Score-Based Generative Priors

    Deliang Wei, Evan Bell, Wenhan Guo, Yifan Chen, Yu Sun

    cs.LG

    Bayesian imaging inverse problems often require sampling from high-dimensional posterior distributions. While recent score-based and diffusion models provide expressive Bayesian priors, their sampling procedures remain inherently sequential and computationally expensive for large-scale imaging applications. We propose PiX-MC, a time-parallel posterior sampling framework based on proximal Langevin dynamics and Picard iteration. The...

    arxiv.org/abs/2608.17666 · PDF

  29. 29

    Elimination Geometry

    Mian Huang, Xueqin Wang

    cs.LG

    This monograph develops elimination geometry (EG), a typed, native-loss, audit-oriented framework for studying when locally optimal objects can be realized by a shared deployment rule. Elimination and compression may erase distinctions required by prediction, inference, control, or representation. EG asks which distinctions are lost, whether the induced defect is visible to the declared task, and whether changing information, architecture,...

    arxiv.org/abs/2608.17646 · PDF

  30. 30

    rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment

    Lars Simon Zehnder

    cs.LG · cs.DC · cs.PF

    We present rl-triton, an open-source library of high-performance GPU kernels for reinforcement learning credit assignment, implemented in Triton. The core contribution is a unified associative scan framework that recasts seven distinct RL estimation algorithms - Generalized Advantage Estimation (GAE), V-Trace, Retrace($λ$), TD($λ$) returns, discounted returns, eligibility traces, and episodic prefix sums - as instances of a single first-order...

    arxiv.org/abs/2608.17641 · PDF

  31. 31

    OOD Detection for EEG-based Machine Learning in High-Risk Environments

    Philipp Bomatter, Henry Gouk

    cs.LG

    Machine learning models for electroencephalography (EEG) analysis show great promise across a wide range of applications, but their deployment in high-risk domains is hindered by their vulnerability to distribution shifts. Encountering out-of-distribution (OOD) data can lead to catastrophic, overconfident predictive failures. While OOD detection methods can mitigate these risks, they remain heavily under-explored for EEG. Moreover,...

    arxiv.org/abs/2608.17620 · PDF

  32. 32

    Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries

    Henrik Wille, Luis-Finley Schütz, Felix Strieth-Kalthoff

    cs.LG · cs.AI · cs.CL

    Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear. Herein, we systematically benchmark four molecular language models across six virtual molecular libraries spanning drug discovery, organic materials, and catalysis. Native molecular language model...

    arxiv.org/abs/2608.17567 · PDF

  33. 33

    No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

    Jack Boylan, Chris Hokamp

    cs.LG · cs.AI

    Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting future embeddings, but the objective admits a trivial solution of a constant encoder, so every practical system adds an anti-collapse mechanism (LeCun, 2022; Assran et al., 2023; Bardes et al., 2022; 2024). LeWorldModel (LeWM) prevents collapse with SIGReg, a regularizer that forces the latent distribution to match an isotropic Gaussian: the representation is...

    arxiv.org/abs/2608.17542 · PDF

  34. 34

    Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents

    Ram Rachum, Yotam Amitai, Bálint Gyevnár, Reuth Mirsky, Cameron Allen

    cs.LG

    This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like faithfulness and compactness, and on human-grounded proxies like subjective ratings or prediction accuracy. We suggest evaluating XRL methods by how effectively their generated explanations help to diagnose and fix malfunctioning reinforcement learning (RL) agents....

    arxiv.org/abs/2608.17524 · PDF

  35. 35

    Causal Local States: Scalable Simultaneous Causal Network Inference and Forecasting for Dynamical Systems

    Jonas Braun, Fabian Fischbach, Daniel Köglmayr, Sebastian Baur, Christoph Räth

    cs.LG

    Machine learning methods predict many real-world systems with remarkable accuracy, but they are typically treated as black boxes that offer no insight into which interactions drive the dynamics. Causal discovery methods reconstruct the interaction network from observational data, but without regard to whether the inferred structure supports prediction. Existing approaches combining both tasks rely on a single global hyperparameter, such as a...

    arxiv.org/abs/2608.17452 · PDF

  36. 36

    General Semantic Knowledge Infusion for Spatio-Temporal Traffic Forecasting

    Mattis thor Straten, Yannick Wolker, Steffen Strohm, Prathvish Mithare, Ralf Krestel, Matthias Renz

    cs.LG

    Although Graph Neural Networks (GNNs) have made significant advances in spatio-temporal traffic forecasting, their performance is limited when they rely solely on sensor proximity or road-network topology. This paper presents a spatio-temporal prediction framework, developed to incorporate knowledge in various forms. This framework aims to improve sensor-level, contextual understanding of the environment. A general-purpose knowledge graph...

    arxiv.org/abs/2608.17440 · PDF

  37. 37

    GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models

    Peizheng Guo, Jianqi Zhang, Xingyu Zhang, Yun Fan, Jiahuan Zhou, Changwen Zheng, Wenwen Qiang

    cs.LG

    Group Relative Policy Optimization (GRPO) has become a widely used approach for post-training Large Language Models (LLMs) for reasoning. In GRPO, the group gradients induced by different queries within the same mini-batch are directly averaged to form the policy update. However, these group gradients can point in conflicting directions. Our empirical analysis suggests that group-gradient conflicts tend to be associated with less effective...

    arxiv.org/abs/2608.17411 · PDF

  38. 38

    Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

    Hoda Yamani, Henry Williams, Bruce A. MacDonald

    cs.LG · cs.AI

    Sample efficiency is a central challenge in reinforcement learning (RL), particularly in image-based domains where agents must learn from high-dimensional visual inputs. Traditional sampling often relies on random or suboptimal experience selection, leading to redundant updates and slow learning. Improving efficiency requires mechanisms that prioritize informative experiences while also encouraging effective exploration. Prioritized...

    arxiv.org/abs/2608.17373 · PDF

  39. 39

    Pathology Transport: Optimal-Transport Explanations for Clinical Data, and When Their Heatmaps (Fail to) Localize Disease

    Lalit Kumar

    cs.LG

    Generative models promise a route to explainable clinical AI: rather than probe a classifier, model the distributions of healthy and diseased patients and read explanations off the geometry between them. We build such a system - an optimal-transport rectified flow trained between two clinical distributions - and use it to ask a pointed question the field too rarely tests: do the resulting explanation heatmaps actually localize disease? On...

    arxiv.org/abs/2608.17370 · PDF

  40. 40

    CORAM: Coherent Orthogonal Rotation for Model Merging

    Xinyi Sui, Ziran Liu, Nam Ling, Wei Wang, Wei Jiang

    cs.LG

    Merging finetuned models combines specialized capabilities without joint training or access to the original data. Most methods operate by linear arithmetic in Euclidean weight space, which cannot carry the geometry of the update. Orthogonal Model Merging (OrthoMerge) uses a single orthogonal transform for each weight matrix, but such a transform cannot change singular values. We propose CORAM, which partitions each target matrix into row...

    arxiv.org/abs/2608.17366 · PDF

  41. 41

    Repetition as Reinforcement: Enhancing Sample Efficiency via Instant Episode Repetition in Reinforcement Learning

    Hoda Yamani, Yuning Xing, Koen van Rijnsoever, Bruce A. MacDonald, Henry Williams

    cs.LG · cs.RO

    Repetition is a fundamental mechanism in human learning, where revisiting successful experiences strengthens memory, consolidates skills, and improves future performance. Motivated by this biological principle, we introduce Instant Episode Repetition (IER), a simple and novel mechanism that improves sample efficiency by immediately repeating action sequences from successful episodes during environment interaction. Unlike conventional...

    arxiv.org/abs/2608.17347 · PDF

  42. 42

    Tight Bounds for Data-driven Multiple Hyper-parameter Tuning with Structured Loss Function

    Anh Tuan Nguyen, Viet Anh Nguyen

    cs.LG · stat.ML

    Data-driven algorithm design frames hyperparameter tuning as a statistical learning problem, but establishing generalization guarantees remains challenging due to the implicit, non-smooth dependence of model performance on hyperparameters. Existing multi-dimensional bounds under piecewise-polynomial assumptions remain theoretically loose and lack comprehensive lower bounds. We resolve this by establishing tight pseudo-dimension bounds for...

    arxiv.org/abs/2608.17343 · PDF

  43. 43

    MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting

    Bowen Liu, Mingming Sun

    cs.LG · cs.AI

    Forecasting cryptocurrency prices remains a formidable challenge due to inherent non-stationarity, abrupt regime shifts, and multi-scale stochastic dependencies. Conventional deep learning models often struggle to capture complex underlying dynamics, frequently resulting in persistent phase-lagged predictions. To address these limitations, we propose MoFE, a novel deep learning framework that integrates Fourier Neural Operators (FNOs) within...

    arxiv.org/abs/2608.17342 · PDF

  44. 44

    Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

    Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee

    cs.LG

    Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs, and longer-horizon trajectories make credit assignment in RL substantially harder. This paper argues that evolution...

    arxiv.org/abs/2608.17310 · PDF

  45. 45

    Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting

    Rongwen Li, Haixin Xie, Xiao Wang, Changjian Chen

    cs.LG · cs.AI

    Existing research on irregular time-series forecasting has primarily focused on model design, while evaluation metrics remain insufficiently studied. Existing benchmarks typically use mean squared error (MSE) as the evaluation metric. We show that, in irregular forecasting, MSE is determined not only by the model prediction but also by the sample-specific timestamp sampling distributions, leading to a biased assessment of the models'...

    arxiv.org/abs/2608.17293 · PDF

  46. 46

    Abra: Scaling Diffusion Image Training

    Kyle Chickering, Wei-An Lin, Swayam Bhanded, Dan Saunders, Akshat Tripathi, Jiaming Song, Shyamal Buch, Xinchen Yan

    cs.LG

    Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled family of flow-matching transformers trained across three orders of magnitude worth of compute ($10^{19}$ to $10^{22}$ FLOPs), reaching significantly larger compute budgets than previous works. We demonstrate that...

    arxiv.org/abs/2608.17286 · PDF

  47. 47

    Rethinking Irregular Time Series Forecasting from the Perspective of Basis Functions

    Rongwen Li, Changjian Chen

    cs.LG · cs.AI

    Irregular time series forecasting is crucial in many domains, such as healthcare and meteorological observation. However, due to the inherent characteristics of irregular time series, including sparse observations and non-uniform sampling, accurately predicting future dynamics remains challenging. In light of these two characteristics, many existing methods aggregate irregular observations into fixed-dimensional estimated response...

    arxiv.org/abs/2608.17284 · PDF

  48. 48

    Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics

    Zhikai Ding, Ziyi Ye

    cs.LG · cs.AI

    Curriculum learning has been widely adopted in the post-training of large language models by organizing training data from easy to hard. However, its effectiveness varies substantially across reasoning tasks, suggesting that no single curriculum is universally optimal and raising a fundamental question: what determines when curriculum learning works? In this paper, we answer this question by analyzing the optimization dynamics induced by...

    arxiv.org/abs/2608.17268 · PDF

  49. 49

    Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

    Yunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu, Tianjin Huang, Yuanyuan Shi, Ziang Xiao, Nuno Vasconcelos, Yijiang Li

    cs.LG · cs.AI · cs.CV

    Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). Such annotations are costly to obtain and become increasingly scarce as reasoning capabilities advance beyond what humans can reliably evaluate. Self-rewarding RL reduces this dependence by enabling models to derive...

    arxiv.org/abs/2608.17253 · PDF

  50. 50

    Physics-Informed and Hybrid Machine Learning in Additive Manufacturing: Application to Fused Filament Fabrication

    Berkcan Kapusuzoglu, Sankaran Mahadevan

    cs.LG · cs.CE · stat.CO

    This article investigates several physics-informed and hybrid machine learning strategies that incorporate physics knowledge in experimental data-driven deep-learning models for predicting the bond quality and porosity of fused filament fabrication (FFF) parts. Three types of strategies are explored to incorporate physics constraints and multi-physics FFF simulation results into a deep neural network (DNN), thus ensuring consistency with...

    arxiv.org/abs/2608.17246 · PDF

  51. 51

    Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer's Disease Detection

    Chanwoo Park, Chanwoo Kim

    cs.LG · cs.AI

    Low-cost, scalable screening for dementia remains an open problem. Imaging-based diagnosis is costly and hard to deploy widely. Electroencephalography (EEG) is portable and inexpensive, but its recordings are noisy, vary widely across subjects, and carry few clinical labels. We tackle this with Delta2Gamma, a self-supervised framework that learns EEG representations from unlabeled data by contrasting augmented views of each signal. Rather...

    arxiv.org/abs/2608.17231 · PDF

  52. 52

    Pessimistic Meta-Induction and Its Limits: Lessons from Frequentist Statistics and Machine Learning Theory

    Hanti Lin

    cs.LG · stat.ME

    This paper challenges the pessimistic meta-inductive argument against scientific realism by undermining its inductive step rather than its historical premise. Although related challenges already exist, I develop a new one. Drawing on a general epistemology of scientific inference developed in frequentist statistics, machine learning, and formal epistemology, I evaluate induction in terms of convergence to the truth. I argue that ordinary...

    arxiv.org/abs/2608.17213 · PDF

  53. 53

    How smoothing the affinity matrix affects neighborhood preservation in t-SNE

    Shirin Mohebi, Guillaume Bied, Jefrey Lijffijt

    cs.LG · cs.CV

    Dimensionality reduction methods are instrumental to visualize high-dimensional data, and t-SNE stands as one of the most widely used methods due to its emphasis on local neighborhood preservation. A central component of t-SNE is the affinity matrix, which expresses pairwise similarities in the form of symmetrized probabilities, over which the optimization problem of t-SNE is defined. We study how the sharpness of this probability...

    arxiv.org/abs/2608.17190 · PDF

  54. 54

    Reinforcement Learning as (Discrete) Potential Theory

    Christopher Connolly

    cs.LG · cs.GT

    Reinforcement learning (RL) theory fundamentally depends on probability theory through the Markov chain. There is a deep connection between probability theory and potential theory. This paper reviews that connection and explores the potential-theoretic viewpoint for core reinforcement learning representations and algorithms under a fixed-policy assumption. This viewpoint may offer a path for improved sample efficiency and formal constraints...

    arxiv.org/abs/2608.17181 · PDF

  55. 55

    Task Specialization Fine-Tuning for Contextual Reinforcement Learning

    Jianan Zhou, Jung-Hoon Cho, Tianyue Zhou, Han Zheng, Jie Zhang, Roy Dong, Yining Ma, Cathy Wu

    cs.LG · cs.AI

    Contextual Reinforcement Learning (CRL) seeks to generalize classical RL by maximizing task coverage across a context space of related tasks. While prior works often train from scratch and rely on either multi-task learning for a single policy or strategically training multiple policies, we advocate for a unified alternative: pretraining a single policy with good initial performance, followed by fine-tuning multiple policies for task...

    arxiv.org/abs/2608.17180 · PDF

  56. 56

    Population Health-Based Machine Learning Reveals Associations Between Psychosocial Factors and Chronic Kidney Disease

    Md. Atik Shams, David Eisenberg, Sumaiya Fatema, Asma Sultana, D. M Hasibul Islam, Junnatul Mawa, Anindita Datta,...

    cs.LG

    Chronic kidney disease (CKD) progresses silently and severely undermines quality of life, making early detection critical for improving patient outcomes. We present a two-part study that combines large-scale telehealth data with advanced machine learning to both classify self-reported CKD status and identify key drivers of disease. Using selected features from the Behavioral Risk Factor Surveillance System (BRFSS 2021: 438,693 samples; BRFSS...

    arxiv.org/abs/2608.17174 · PDF

  57. 57

    SCENARIODIFF: A Scenario-level Guidance Framework for Multimodal Time Series Forecasting--Extended Version

    Tuan-Binh Tran, Dat Nguyen Cong, Duc-Trong Le, Thanh Trung Huynh, Tung Kieu

    cs.LG

    Textual context such as news, reports, and logs can provide valuable signals for time series forecasting, especially when future dynamics are driven by external events that are not yet visible in historical values. Existing multimodal forecasting methods often either ask large language models (LLMs) to predict numerical values directly or fuse text and time series implicitly, making contextual influence difficult to interpret and control. We...

    arxiv.org/abs/2608.17164 · PDF

  58. 58

    Q-Learning With World Models

    Perry Dong, Yueru Jia, Chelsea Finn, Dorsa Sadigh

    cs.LG · cs.AI

    Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Language-Action models into reliable, high-performing policies. World models offer a further lever for sample efficiency, as they predict state changes rather than actions alone, but their success has largely been confined to supervised policy learning. Prior model-based RL methods often optimize the policy...

    arxiv.org/abs/2608.17163 · PDF

  59. 59

    OraclePhys: A Systematic Framework for LLM Fine-Tuning on Structural Mechanics

    Mingyu Li, Guorui Song, Jing Lin, Haoqian Wang

    cs.LG

    What a language model internalizes from fine-tuning is usually diagnosed after the fact. We make it an experimental variable. OraclePhys is a systematic fine-tuning framework with three components: OraclePhys-Bench, an exactly-graded structural-mechanics benchmark whose finite-element oracle scores every answer and counterfactual edit -- no human labels, no LLM judging; OraclePhys-30K, a supervision dataset of seven answer forms over...

    arxiv.org/abs/2608.17162 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.