cs.LG · 2026-06-12 · No. 21

Machine Learning, 2026-06-12.

62 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

62 entries
  1. 01

    Understanding Truncated Positional Encodings for Graph Neural Networks

    James Flora, Mitchell Black, Weng-Keen Wong, Amir Nayyeri

    cs.LG

    Positional encodings (PEs) enhance the power of graph neural networks (GNNs), both theoretically and empirically. Two of the most popular families of PEs - spectral (e.g., Laplacian eigenspaces, effective resistance) and walk-based (polynomials of the adjacency matrix) - are theoretically equivalent in expressive power, with expressivity between the 1-WL and 3-WL tests. However, this equivalence assumes the GNN uses the "complete" version of...

    arxiv.org/abs/2606.13671 · PDF

  2. 02

    Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation

    Guo Yu, Wenlin Liu, Yulan Hu, Hao-Xuan Ma, Jun-Peng Jiang, Han-Jia Ye

    cs.LG

    On-policy distillation (\textsc{OPD}) has recently become a prominent post-training recipe as it combines two desirable ingredients: on-policy student trajectories and dense teacher supervision, yet how this hybrid changes a model's parameters remains unclear. Across several language and vision-language model pairs and use cases, our analysis yields two main findings. On sparsity, \textsc{OPD}-style updates are small and coordinate-sparse....

    arxiv.org/abs/2606.13657 · PDF

  3. 03

    The Stable Recovery Manifold: Geometric Principles Governing Recoverability in Continual Learning

    Ayushman Trivedi, Bhavika Melwani

    cs.LG

    Catastrophic forgetting is often viewed as the destruction of previously learned knowledge during sequential learning. Building on the Accessibility Collapse framework, we investigate the geometric structure of recoverability in continual learning. Using Split CIFAR-100 and a sequentially trained ResNet-18, we analyze recoverability, representational drift, and recovery complexity across ten tasks. We introduce Recovery Subspace...

    arxiv.org/abs/2606.13637 · PDF

  4. 04

    Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models

    Daniel Scalena, Sara Candussio, Luca Bortolussi, Elisabetta Fersini, Malvina Nissim, Gabriele Sarti

    cs.LG · cs.AI · cs.CL

    Chain-of-thought (CoT) reasoning is the dominant paradigm for inference-time scaling in language models, yet the causal influence of individual steps on the final answer poorly understood. We estimate each step's causal importance via early exit and use this measure to study how answers form across the reasoning traces of several model families. Across diverse tasks, we find that reasoning typically crosses a \emph{commitment boundary} -- a...

    arxiv.org/abs/2606.13603 · PDF

  5. 05

    Simplex-Constrained Sparse Bagging: Transitioning from Uniform Priors to Sparse Posteriors in Ensemble Learning

    Meher Sai Preetam, Meher Bhaskar

    cs.LG

    We present Simplex-Constrained Sparse Bagging (SCSB), a mathematically rigorous framework for post-training compression and probability calibration of bootstrap-based bagging ensembles. Standard bagging ensembles (such as Random Forests, Bagged SVMs, and Bagged Neural Networks) assign uniform voting power to all constituent estimators. However, this naive uniform prior ignores the varying local competence of base estimators and contributes to...

    arxiv.org/abs/2606.13589 · PDF

  6. 06

    Learning with Simulators: No Regret in a Computationally Bounded World

    Sasha Voitovych, Abhishek Shetty, Noah Golowich, Alexander Rakhlin

    cs.LG · cs.CC · cs.DS · stat.ML

    Understanding the minimal assumptions necessary for generalization is the fundamental question in learning theory. Unfortunately, most results rely heavily on independence (or some proxy thereof) of the data-generating process, while results for strongly dependent data are far more limited. Towards addressing this gap, we introduce the framework of simulatable processes, where the learner has access to a simulator that approximates the...

    arxiv.org/abs/2606.13576 · PDF

  7. 07

    Existence Precedes Value: Joint Modeling of Observational Existence and Evolving States in Time Series Forecasting

    Yifan Hu, Hongzhou Chen, Peiyuan Liu, Yiding Liu, Zewei Dong, Jiang-Ming Yang

    cs.LG · cs.AI

    Real-world time series are often highly incomplete and irregular due to sensor dormancy, transmission delays, and event-driven sampling, making reliable forecasting fundamentally challenging. Existing methods have evolved from impute-then-forecast pipelines to continuous-time models such as Neural ODEs and continuous-time graph networks. While these approaches improve the modeling of historical irregularity, they still rely on an implicit...

    arxiv.org/abs/2606.13571 · PDF

  8. 08

    Adjusted Cup-Product Neural Layer

    Snigdha Chandan Khilar

    cs.LG · math-ph

    Many important observables in physics and geometry are cup products of cochains. The adjusted cup product neural layer has been introduced in this paper. It is a neural primitive that hard wires the cup product with an adjustment term from higher gauge theory. This creates a readout that is gauge invariant by design. Their main theoretical result shows that on a closed cycle the output relies entirely on the adjustment coefficient. Setting...

    arxiv.org/abs/2606.13568 · PDF

  9. 09

    A2D2: Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding

    Sophia Tang, Yuchen Zhu, Molei Tao, Pranam Chatterjee

    cs.LG

    Discrete diffusion models offer a simple and stable likelihood-based framework for sequence generation, recently extended to any-length settings via token insertion. Principled reward-guided fine-tuning for any-length discrete diffusion, however, remains largely unexplored. We introduce Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding (A2D2), a unified framework for reward-guided fine-tuning of any-length discrete diffusion...

    arxiv.org/abs/2606.13565 · PDF

  10. 10

    CRAFTIIF: Cross-Resolution Analytic Four-Type Interpretable Isolation Forest for Multivariate Time Series Anomaly Detection

    William Smits

    cs.LG · cs.AI

    Anomaly detection in multivariate time series is challenged by four structurally distinct anomaly types -- point (isolated spikes), distributional (level shifts), temporal (rhythm changes), and collective (inter-sensor correlation breakdowns) -- each requiring different feature representations. Most unsupervised methods target only one or two types and provide limited interpretability. We present CRAFTIIF (Cross-Resolution Analytic Four-Type...

    arxiv.org/abs/2606.13486 · PDF

  11. 11

    SupraBench: A Benchmark for Supramolecular Chemistry

    Tianyi Ma, Yijun Ma, Zehong Wang, Weixiang Sun, Ziming Li, Connor R. Schmidt, Chuxu Zhang, Matthew J. Webber, Yanfang Ye

    cs.LG · cs.AI · cs.CL

    Supramolecular chemistry, which includes the study of non-covalent host-guest assemblies, has advanced various applications. However, designing host-guest systems remains time-consuming, requiring days of dry-lab verification per candidate pair. Although LLMs have emerged as a fast alternative with strong performance on molecular binding tasks, no benchmark currently systematically evaluates LLMs for host-guest reasoning across fundamental...

    arxiv.org/abs/2606.13477 · PDF

  12. 12

    MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

    Jiacheng Chen, Xinyu Zhang, Shunkai Zhang, Yanmohan Wang, Lin Li, Tiancheng Qin, Qin Wang, Zhengmao Zhu, Tianle Li,...

    cs.LG · cs.AI · cs.CL

    We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabilities -- proof generation, proof verification, and critique-conditioned proof repair -- using a defense-in-depth generative verifier engineered for low false-positive rate. These capabilities are merged into a single released M3 model. At test time, MaxProof treats...

    arxiv.org/abs/2606.13473 · PDF

  13. 13

    Reinforcement Learning for Neural Model Editing

    Shaivi Malik

    cs.LG · cs.CV

    Editing pretrained neural networks requires specialized algorithms tailored to specific objectives. Designing such algorithms is often time-consuming and demands significant effort. We present an exploratory framework that formulates neural model editing as a reinforcement learning problem, where agents modify models using reward feedback. We introduce two environments: MaskWorld, where agents scale weights multiplicatively, and ShiftWorld,...

    arxiv.org/abs/2606.13461 · PDF

  14. 14

    Uncertainty Estimation for Molecular Diffusion Models

    Paul Seij, Christian A. Naesseth, Stephan Mandt, Metod Jazbec

    cs.LG

    Diffusion models have seen wide adoption for 3D molecular generation, yet they offer no principled signal of when a generated molecule is likely to be of low quality. We propose a post-hoc method for estimating per-sample uncertainty in pretrained molecular diffusion models. Building on a Laplace approximation of the denoising network, we measure the variability of the noise prediction across the generation trajectory. Empirically, we show...

    arxiv.org/abs/2606.13451 · PDF

  15. 15

    Clustering Node Attributed Networks with Graph Neural Networks and Self Learning

    Rodrigo de Sapienza Luna, Daniel Ratton Figueiredo

    cs.LG

    Graph clustering - partitioning the node set of a graph into disjoint subsets that reflect some latent information - is a fundamental problem as it finds applications in a myriad of different scenarios. While this classic problem has been tackled for decades by different communities, a recent variation of the problem driven by real data considers the scenario where nodes have attributes that are also informative. This has triggered novel...

    arxiv.org/abs/2606.13444 · PDF

  16. 16

    How Much Memory Do We Need? Adaptive Memory Gate for Neural Operators

    Jihyeon Hur, Yongseok Kwon, Min-Gi Jo, Jeongwhan Choi, Noseong Park

    cs.LG

    Neural operators have emerged as a powerful data-driven approach for solving time-dependent PDEs. Among recent advances, memory-augmented neural operators explicitly incorporate past states and have achieved remarkable performance under low-resolution observation settings. However, existing approaches apply a fixed memory weight regardless of observation conditions, such as resolution or physical parameters, limiting their adaptability. Our...

    arxiv.org/abs/2606.13443 · PDF

  17. 17

    Accelerating Speculative Diffusions via Block Verification

    Alexander Soen, Hisham Husain, Valentin De Bortoli, Arnaud Doucet

    cs.LG · stat.ML

    Speculative decoding speeds up LLM inference by using a draft model to generate tokens, with an acceptance-rejection scheme that ensures that the output matches the target distribution. Adapting this to continuous diffusions is difficult because speculative sampling requires drawing from a residual distribution. While straightforward in discrete spaces, efficiently sampling this residual in continuous space is non-trivial. Consequently,...

    arxiv.org/abs/2606.13426 · PDF

  18. 18

    PolyFlow: Safe and Efficient Polytope-Constrained Flow Matching with Constraint Embedding and Projection-free Update

    Jianming Ma, Qiyue Yang, Yang Zhang, Liyun Yan, Zhanxiang Cao, Yazhou Zhang, Yue Gao

    cs.LG · cs.AI · cs.RO

    While flow-based generative models have demonstrated strong performance across a wide range of domains, deploying them in safety-critical physical systems remains challenging due to strict constraint requirements. Existing approaches typically enforce safety through post-hoc corrections, which incur substantial computational overhead and may distort the learned distribution. We propose PolyFlow, a polytope-constrained flow matching framework...

    arxiv.org/abs/2606.13400 · PDF

  19. 19

    Hölder++: Improving the Quality-Coherence Trade-off in Multimodal VAEs

    Huyen Vo, María Martínez-García, Isabel Valera

    cs.LG

    Existing approaches for multimodal variational autoencoders (VAEs) face a trade-off between generative quality and coherence-i.e., they struggle to generate realistic and diverse samples that, at the same time, are semantically consistent across modalities. A recent work shows that using a simple approximation to Hölder pooling as an aggregation method improves coherence over the SOTA MMVAE+, despite assuming a single shared representation...

    arxiv.org/abs/2606.13381 · PDF

  20. 20

    Positional Encoding in the Context of Memristor-Based Analog Computation for Automatic Speech Recognition

    Benedikt Hilmes, Nick Rossenbach, Ralf Schlüter

    cs.LG · cs.AR · cs.ET

    Memristors provide a new chance for resource-efficient computation of neural models for natural language processing by enabling analog execution of vector-matrix-multiplication. Yet, computations on these devices are currently subject to larger distortion, both in weight programming and execution. In this work, we identify large output values of transformed positional encodings to cause major degradation within analog-to-digital conversion...

    arxiv.org/abs/2606.13379 · PDF

  21. 21

    VideoMDM: Towards 3D Human Motion Generation From 2D Supervision

    Amir Mann, Gal Michael Harari, Merav Keidar, Or Litany

    cs.LG · cs.CV

    We introduce VideoMDM, a diffusion-based framework that trains 3D human motion priors directly from accurate 2D poses extracted from monocular videos, without any 3D ground truth. A pretrained 2D-to-3D lifter provides approximate 3D pose sequences that serve as a noisy teacher: these are diffused, denoised by the model in 3D, and supervised in 2D by reprojecting the prediction and comparing against accurate keypoints. We show that, under mild...

    arxiv.org/abs/2606.13364 · PDF

  22. 22

    Enhanced Low-Density Region Exploration in Classifier-Guided Diffusion Models Through Modified Reverse Diffusion Sampling

    Jagriti Singh, Shekhar Verma, Muneendra Ojha

    cs.LG

    Diffusion models have emerged as state-of-the-art generative models for high-fidelity image synthesis, particularly in their classifier-free guided and classifier-guided forms. However, standard classifier guidance concentrates probability mass around high-density class mean, leading to poor coverage of rare samples in the tails of the class-conditional distributions. Recent work on diffusion-based tail sampling mitigates this by training an...

    arxiv.org/abs/2606.13347 · PDF

  23. 23

    Navigating the Safety-Fidelity Trade-off: Massive-Variate Time Series Forecasting for Power Systems via Probabilistic Scenarios

    Kaijie Xu, Anqi Wang, Xilin Dai

    cs.LG

    Probabilistic forecasting models are increasingly deployed on multivariate systems with distinct channel physics and operational constraints, but existing benchmarks evaluate neither property at scale. Public canonical multivariate benchmarks cap out at 2,000 channels, while power-system benchmarks either lack temporal structure or probabilistic evaluation. We introduce PowerPhase, a probabilistic forecasting benchmark built on six...

    arxiv.org/abs/2606.13338 · PDF

  24. 24

    Rarity-Gated Context Conditioning for Offline Imitation Learning-Based Maritime Anomaly Detection

    Yongmin Kim, ByeongHoon Jeon, Sungil Kim

    cs.LG · cs.AI

    Contextual anomaly detection aims to identify abnormal behavior conditional on context variables, but practical deployments often face highly imbalanced context distributions where rare regimes can be critical information. Under such frequency bias, context-conditioned models can produce unstable decisions and excessive false alarms in rare contexts. We propose Rarity-Gated Feature-wise Linear Modulation (RGFiLM), a rarity-aware conditioning...

    arxiv.org/abs/2606.13311 · PDF

  25. 25

    Quantizing Time-Series Models As Dynamical Systems: Trajectory-Based Quantization Sensitivity Score

    Mariya Pavlova, Harrison Bo Hua Zhu, Elizsveta Semenova, Yingzhen Li

    cs.LG

    We introduce the Trajectory-based Quantization Sensitivity Score (TQS), a metric that reframes post-training quantization (PTQ) through the lens of dynamical-systems stability. By modeling the network's rollout as a discrete-time dynamical system, TQS characterizes how quantization-induced errors propagate and amplify over the rollout horizon. Unlike conventional PTQ methods, where sensitivity analysis is often coupled to the quantization...

    arxiv.org/abs/2606.13300 · PDF

  26. 26

    Clipping Makes Distributed and Federated Asynchronous SGD Robust to Stragglers

    Samuel Erickson, Mikael Johansson

    cs.LG · cs.DC · math.OC

    In modern machine learning, parallelization of training is an important strategy for increasing scale. Asynchronous stochastic gradient descent (ASGD), which maximizes the utilization of available hardware by avoiding waiting for slow workers. However, with constant step sizes, the convergence of ASGD is nonetheless affected negatively by slow workers due to large delays in updates. At the same time, it has been empirically observed in...

    arxiv.org/abs/2606.13287 · PDF

  27. 27

    Once-for-All: Scalable Simultaneous Forecasting via Equilibrium State Estimation

    Beinan Xu, Andy Song, Jiti Gao, Feng Liu

    cs.LG · cs.AI

    We introduce Equilibrium State Estimation (ESE), a novel paradigm for simultaneous prediction, where multiple interacting systems require separate yet coordinated forecasts. Such scenarios often arise in real-world settings such as economics and healthcare modeling. Unlike existing approaches that predict one system at a time, ESE forecasts all systems in a single pass. It first estimates the equilibrium state across systems, then generates...

    arxiv.org/abs/2606.13285 · PDF

  28. 28

    Different Layers, Different Manifolds: Module-Wise Weight-Space Geometry in Transformer Optimization

    Kirato Yoshihara

    cs.LG · cs.AI

    Weight-space geometry plays a central role in neural network optimization, yet manifold constraints are often applied uniformly across all weight matrices. In this work, we ask whether different transformer modules prefer different manifold geometries. We study Manifold Muon for GPT-2 pretraining and compare layer-wise assignments of Stiefel and DGram constraints across attention and MLP blocks. Our results show a clear asymmetry:...

    arxiv.org/abs/2606.13276 · PDF

  29. 29

    Extracting Governing Equations from Latent Dynamics via Multi-View Contrastive Learning

    Paolo Muratore, Mackenzie Weygandt Mathis

    cs.LG · q-bio.NC

    Identifying latent dynamical systems from noisy, high-dimensional measurements is a central problem at the intersection of representation learning, system identification, and scientific discovery. We present DYSCO, a multi-view temporal contrastive learning algorithm that jointly recovers latent trajectories and the governing dynamics from such observations, by leveraging multiple independent noisy views of the same underlying process to...

    arxiv.org/abs/2606.13260 · PDF

  30. 30

    To GAN or Not To GAN: Segmentation Analysis on Mars DEM

    Douglas Dziedzorm Agbeve, Aditya V. Handrale, Salim Fares, Seif E. Idani

    cs.LG

    To better understand Martian Surface, which is needed to enable Rovers navigate Mars with ease, it is necessary to be able to determine the location of mounds. Detecting and studying these morphologies can also help us find evidence of extraterrestrial life, in this case, more specifically, water or signs of life conducive environments. Detection of mounds was done by manually mapping morphological parameters onto Digital Elevation Models....

    arxiv.org/abs/2606.13252 · PDF

  31. 31

    Towards More General Control of Diffusion Models Using Jeffrey Guidance

    Raphaël Razafindralambo, Rémy Sun, Frédéric Precioso, Jes Frellsen, Pierre-Alexandre Mattei

    cs.LG · cs.AI · cs.CV · stat.ME · stat.ML

    A key strength of diffusion models lies in their flexibility, since their outputs can be controlled at sampling time through guidance. However, beyond simple cases such as conditional sampling, the target distribution is often left implicit, defined only through a sampling rule or a heuristic energy function. To address this, we propose Jeffrey guidance, a principled framework that extends diffusion-model control to applications beyond what...

    arxiv.org/abs/2606.13240 · PDF

  32. 32

    Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier

    Olga Isupova, Danil Kuzin, Ella Browning, Tom Mills, Steven Reece

    cs.LG · cs.AI · cs.SD · stat.AP

    Passive acoustic monitoring holds great promise for ecological inference, yet existing automated tools are typically narrowly trained and non-transferable. We address these limitations with PULSE, a semi-supervised, multi-task framework for Orthoptera bioacoustics, combining weakly-supervised species classification, self-supervised learning on unlabelled field audio, and knowledge distillation from a general-purpose bioacoustic model. Our...

    arxiv.org/abs/2606.13236 · PDF

  33. 33

    ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling

    Sihwa Lee, Janghwan Lee, Donghoon Yoo, Jae Gon Kim, Hanyul Ryu, Soojung Ryu, Jungwook Choi

    cs.LG · cs.AI

    Large reasoning models (LRMs) improve complex problem-solving by generating long intermediate reasoning traces, but this substantially increases inference costs. NVFP4 inference offers a promising approach to reduce both computational and memory costs through hardware-supported low-precision execution. However, directly applying NVFP4 to LRMs introduces two practical limitations: reasoning accuracy degrades under quantization, and existing...

    arxiv.org/abs/2606.13233 · PDF

  34. 34

    Distributional Loss for Robust Classification

    Kathleen Anderson, Thomas Martinetz

    cs.LG · cs.CV

    This paper proposes a novel loss concept for supervised classification tasks. Rather than enforcing a direct mapping from each input sample to a single assigned label, we define an optimization objective over all classifier outputs as a bimodal Gaussian distribution. This softer target formulation implicitly captures class ambiguity, mitigates overfitting, and encourages the learning of more robust decision boundaries, all without requiring...

    arxiv.org/abs/2606.13223 · PDF

  35. 35

    From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation

    Bora Kargi, David Salinas

    cs.LG

    Evaluating new large language models typically requires costly human annotation campaigns at scale. LLM-as-a-judge offers a cheaper alternative, but judge scores carry systematic errors - such as position bias, self-preference, or intransitivity - that can strongly miscalibrate the resulting rankings. We quantify the resulting judge-human disagreement at two complementary levels. At the local level, we estimate per-battle uncertainty from the...

    arxiv.org/abs/2606.13221 · PDF

  36. 36

    Understanding helpfulness and harmless tension in reward models

    Eshaan Tanwar, Pepa Atanasova

    cs.LG · cs.CL

    Reward models are a key component of reinforcement learning from human feedback (RLHF), aligning language models toward both helpful and harmless behaviour. However, the internal mechanisms underlying these objectives and their conflicts remain poorly understood. We study alignment tension in reward models trained under helpfulness-only, harmlessness-only, and mixed-objective settings. We find that mixed-objective models often underperform...

    arxiv.org/abs/2606.13209 · PDF

  37. 37

    WHAR Arena: Benchmarking the State of the Art in Efficient Wearable Human Activity Recognition

    Maximilian Burzer, Tobias King, Till Riedel, Michael Beigl, Tobias Röddiger

    cs.LG

    Deep learning has become the dominant paradigm in Wearable Human Activity Recognition (WHAR), yet progress is obscured by a comparability crisis. Results are often reported using inconsistent datasets, custom data processing, and varying evaluation protocols, making state-of-the-art claims fragile. We address this with a large-scale, open-source benchmark that integrates 30 diverse datasets under standardized processing, unified model...

    arxiv.org/abs/2606.13194 · PDF

  38. 38

    The Geometry of Phase Transitions in Generative Dynamics via Projection Caustics

    Ryosuke Sakamoto, Kotaro Sakamoto

    cs.LG

    Continuous-state generative samplers, including diffusion and flow-matching models, evolve through continuous reverse-time dynamics, yet their samples often undergo abrupt qualitative changes: trajectories commit to modes, semantic alternatives collapse, and small perturbations in narrow time windows can produce large downstream effects. This paper develops a geometric account of such phase-transition-like behaviour. We view denoising as...

    arxiv.org/abs/2606.13191 · PDF

  39. 39

    Loss-Shift Transfer via Bayes Quotients

    Vasileios Sevetlidis

    cs.LG

    Transfer learning is usually studied as a consequence of distribution shift. This paper identifies an orthogonal failure mode in which the data distribution is fixed and the loss changes. This setting is called \emph{loss shift}. A loss determines which information in \(X\) is Bayes-relevant, and two losses may therefore require different representations even under the same joint law \(P(X,Y)\). The idea is formalized using Bayes quotients,...

    arxiv.org/abs/2606.13178 · PDF

  40. 40

    Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents

    Yujun Zhou, Kehan Guo, Haomin Zhuang, Xiangqi Wang, Yue Huang, Zhenwen Liang, Pin-Yu Chen, Tian Gao, Nuno Moniz,...

    cs.LG · cs.CL

    Interactive LLM agents are becoming part of daily work, but they do not reliably become easier to work with over time: a correction remembered in one session may still be violated in the next. We study this gap between preference access and preference compliance. In tasks derived from anonymized real-user friction cases, Mem0 memory still leaves 57.5% of applicable preference checks violated. We introduce Test-time Rule Acquisition and...

    arxiv.org/abs/2606.13174 · PDF

  41. 41

    Detecting Explanatory Insufficiency in Learned Representations: A Framework for Representational Vigilance

    Jacques Raynal, Pierre Slangen, Elsa Raynal, Jacques Margerit

    cs.LG

    Learned representations are central to modern machine learning and are commonly evaluated through predictive performance, robustness, uncertainty estimation, or generalization. However, a learned representation may remain operationally successful while progressively failing to organize persistent residual structures that are not fully captured by conventional evaluation metrics. This article introduces VER, the Vigilant Evaluator of...

    arxiv.org/abs/2606.13172 · PDF

  42. 42

    When Does Routing Become Interpretable? Causal Probes on Block Attention Residuals

    Aydin Javadov

    cs.LG

    Block Attention Residuals (Block AttnRes) by replace fixed additive residuals with a learned softmax over earlier depth-source representations, surfacing cross-layer routing as an inspectable tensor in the forward pass. This is a tempting interpretability target: information flow normally inferred indirectly is now directly observable. We ask whether such exposure suffices for mechanistic interpretation. We probe two same-scale ($0.6$B) Block...

    arxiv.org/abs/2606.13168 · PDF

  43. 43

    MiniPIC: Flexible Position-Independent Caching in <100LOC

    Nathan Ordonez, Thomas Parnell

    cs.LG · cs.AI · cs.CL

    Retrieval-augmented and agentic workloads repeatedly prefill recurring predictable structured inputs (which we call "spans") such as documents and code files. Yet, prefix caching in engines such as vLLM cannot reuse their KV entries unless they share identical prefixes with another request, while Position-Independent Caching (PIC) implementations within production-grade inference servers typically either require substantial server code...

    arxiv.org/abs/2606.13126 · PDF

  44. 44

    Select and Improve: Understanding the Mechanics of Post-Training for Reasoning

    Akshay Krishnamurthy, Audrey Huang, Nived Rajaraman

    cs.LG · cs.AI

    Reinforcement learning has rapidly emerged as a key component in the training of reasoning and coding models, yet it remains poorly understood from a mechanistic perspective. We study how and through what underlying processes capabilities are acquired or enhanced via reinforcement learning post-training. Our analysis, based on controlled math reasoning experiments with Qwen-2.5-1.5B, reveals two core mechanisms: strategy selection and...

    arxiv.org/abs/2606.13125 · PDF

  45. 45

    MP3: Multi-Period Pattern Pre-training forSpatio-Temporal Forecasting

    Lilan Peng, Yandi Liu, Qingren Yao, Chongshou Li, Tianrui Li

    cs.LG · cs.AI · cs.NE

    Spatio-Temporal forecasting is crucial in diverse fields, such as transportation, climate, and energy. Urban spatio-temporal data exhibits temporal mirage: similar short-window inputs have divergent future trends, and vice versa. Existing spatio-temporal graph neural networks (STGNNs) cannot effectively identify such mirages. We argue that the core reason lies in the short-window inputs that have incomplete period observation, heterogeneous...

    arxiv.org/abs/2606.13119 · PDF

  46. 46

    Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning

    Jiayu Yang, Chao Chen, Shengen Wu, Yinhong Liu, Yuxuan Fan, Lujundong Li, Songning Lai, Chengwei Qin, Zhijiang Guo

    cs.LG · cs.CL

    Latent chain-of-thought compresses reasoning by replacing visible reasoning traces with continuous hidden-state recurrence, but existing formulations are difficult to optimize with standard on-policy reinforcement learning (RL) and hard to interpret causally. Our key insight is that a single pair of explicit boundary tokens can address both issues at once: discrete entry and exit anchors make the latent block compatible with standard...

    arxiv.org/abs/2606.13106 · PDF

  47. 47

    Disparate Impact in Synthetic Data Generation

    Paul Andrey, Michaël Perrot, Batiste Le Bars, Marc Tommasi

    cs.LG

    We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is the same across sensitive groups. Our approach departs from existing work on fair SDG, that address the problem of correcting for undue biases in the observed distribution, hence redefining SDG as learning a distribution that is not that of the real data. By contrast, non-disparate impact is notably...

    arxiv.org/abs/2606.13105 · PDF

  48. 48

    Authority, Truth, and Citation Bias: A Large-Scale Multi-Domain Benchmark for Studying Epistemic Susceptibility in Large Language Models

    Aryan Khurana, Aravind Ramana RN, Dhruv Kumar

    cs.LG

    Large language models are increasingly deployed in citation-augmented settings, yet the effect of citation presence on model behavior independent of factual content remains poorly understood. We introduce AuthorityBench, a 220,564-prompt multi-domain benchmark that isolates how citation-based authority signals influence epistemic behavior in LLMs. The benchmark uses a fully balanced 2x2 factorial design crossing claim veracity with citation...

    arxiv.org/abs/2606.13104 · PDF

  49. 49

    Scale Buys Interpolation, Structure Buys a Horizon: Certified Predictability for Equivariant World Models

    Hongbo Wang

    cs.LG · cs.RO · math.DS

    Scale buys interpolation; structure buys a certified horizon. A world model's average error says nothing about whether a particular prediction can be trusted, or for how long. For equivariant latent world models we give a computable, multi-step certificate of the predictable horizon: $T$-step rollout error is provably constant over each symmetry orbit (Theorem A) and stratified channel-by-channel by the predictor's Lyapunov spectrum,...

    arxiv.org/abs/2606.13092 · PDF

  50. 50

    Emotional regulation improves deep learning-based image classification

    Riccardo Emanuele Landi, João M. F. Rodrigues, Marta Chinnici

    cs.LG · cs.AI

    Emotion significantly influences cognition, enhancing memory and learning under certain conditions. Drawing on this principle, emotion-augmented deep learning investigates how affective states can improve neural network architectures and learning paradigms, achieving better generalization than non-emotional models. However, existing methods often rely solely on objective neurophysiological factors, neglecting the role of subjectivity in...

    arxiv.org/abs/2606.13081 · PDF

  51. 51

    Limits of spectral learning under noise

    Sabin Roman, Ljupco Todorovski, Saso Dzeroski, Marta Sales-Pardo, Roger Guimera

    cs.LG

    Learning functional relationships from noisy data is a central problem in scientific inference. Spectral methods approximate unknown functions by expanding them in a basis and estimating the corresponding coefficients from data, but the stability of these coefficients under noise remains poorly understood. Here we study supervised regression with additive label noise using sparse spectral representations across multiple bases and dimensions....

    arxiv.org/abs/2606.13067 · PDF

  52. 52

    A green solvent screening tool for emerging materials via uncertainty aware, transformer enhanced transfer learning

    Ioannis Kouroudis, Simon Ternes, Zhaosu Gu, Gohar Ali Siddiqui, Marina Ustinova, Angelo Lembo, Alessio Gagliardi,...

    cs.LG

    Accurate prediction of solubility remains a central challenge across materials science and sustainable chemistry. In particular due to emerging technologies like organic and hybrid photovoltaics, batteries, and catalysis, solvent usage is expected to increase significantly within the coming years. Therefore, substituting solvents with greener alternatives is vital. This is where machine learning can have substantial impact. However, the...

    arxiv.org/abs/2606.13060 · PDF

  53. 53

    TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization

    Zhixiong Zhao, Zukang Xu, Zhixuan Chen, Xing Hu, Zhe Jiang, Dawei Yang

    cs.LG · cs.AI

    Large language models (LLMs) exhibit exceptional general language processing capabilities, but their memory and compute costs hinder deployment. Ternarization has emerged as a promising compression technique, offering significant reductions in model size and inference complexity. However, existing methods struggle with heavy-tailed activation distributions and therefore keep activations in high precision, fundamentally limiting end-to-end...

    arxiv.org/abs/2606.13054 · PDF

  54. 54

    CausalMoE: A Billion-Scale Multimodal Foundation Model for Granger Causal Discovery with Pattern-Routed Heterogeneous Experts

    Bo Liu, Di Dai, Jingwei Liu, Jiarui Jin, Xiaocheng Fang, Guangkun Nie, Hongyan Li, Shenda Hong

    cs.LG · cs.AI

    Granger Causal Discovery (GCD) is fundamental for analyzing temporal dependencies in complex systems. However, existing neural GCD methods predominantly rely on a "one-size-fits-all" paradigm, struggling to capture distribution shifts and dynamic regime changes inherent in real-world time series. This often leads to entangled representations and spurious causal graphs. In this paper, we propose CausalMoE, a billion-scale multimodal Granger...

    arxiv.org/abs/2606.13024 · PDF

  55. 55

    scLLM-DSC: LLM-Knowledge Enhanced Cross-Modal Deep Structural Clustering for Single-Cell RNA Sequencing

    Ping Xu, Pengjiang Li, Tian Du, Zaitian Wang, Jiawei Gu, Ziyue Qiao, Pengfei Wang, Yuanchun Zhou

    cs.LG · cs.AI

    Clustering is fundamental to scRNA-seq analysis, serving as a cornerstone for identifying cell populations and resolving tissue heterogeneity. However, existing methods focus on mining numerical statistical patterns, suffering from semantic agnosticism by neglecting the intrinsic biological functions encoded by genes. While Large Language Models (LLMs) offer promising semantic capabilities, their direct adaptation to cell clustering is...

    arxiv.org/abs/2606.13007 · PDF

  56. 56

    Reliability of Probabilistic Emulation of Physical Systems

    Sam F. Greenbury, Radka Jersakova, Paolo Conti, Marjan Famili, Christopher Iliffe Sprague, Edwin Brown, Jason D. McEwen

    cs.LG · stat.ML

    Two dominant approaches have emerged for generating probabilistic forecasts of physical systems: generative models, such as diffusion or flow matching; and ensembles of deterministic models with stochasticity injected, trained using the continuous ranked probability score (CRPS) loss. While both approaches have demonstrated strong predictive accuracy, the reliability of their uncertainties has not been systematically assessed. We address this...

    arxiv.org/abs/2606.12997 · PDF

  57. 57

    DeepJEB++: Foundation Model-Driven Large-Scale 3D Engineering Dataset via 2D Latent Space Augmentation

    Soyoung Yoo, Leekyo Jeong, Jinsu Ra, Dongeon Lee, Sunwoong Yang, Hyogu Jeong, Namwoo Kang

    cs.LG · cs.CE

    Data-driven engineering design is constrained by the lack of large-scale 3D datasets that pair geometry with physics-based performance labels. In particular, existing 3D data augmentation techniques have limitations in preserving subtle and diverse geometric variations, and it remains difficult to automate the subsequent simulation-labeling process, where boundary conditions vary depending on the generated geometry. We present DeepJEB++, a...

    arxiv.org/abs/2606.12994 · PDF

  58. 58

    Exposure Bias as Epistemic Underidentification in Recursive Forecasting

    Riku Green, Zahraa S. Abdallah, Telmo M Silva Filho

    cs.LG

    Recursive multi-step forecasting is usually framed as distribution shift: models are trained on observed histories but deployed on their own predictions. We show this framing is incomplete by proving that, under partial observability or state truncation, recursive rollout is also an epistemic underidentification problem. Even with deterministic latent dynamics, one-step Bayes supervision identifies behavior only on observed contexts and need...

    arxiv.org/abs/2606.12990 · PDF

  59. 59

    EPM-JEPA: Operator-Side Experience Modulation in JEPA-Family World Models

    Vedant Pandya

    cs.LG

    JEPA-family world models use a static predictor whose weights do not adapt when test-time dynamics diverge from training. We compare two mechanisms for incorporating accumulated experience into a JEPA predictor under distribution shift: operand-side injection, where a compressed experience representation is added as a residual to the predictor's hidden state (EI-JEPA), and operator-side modulation, where the same representation generates...

    arxiv.org/abs/2606.12979 · PDF

  60. 60

    Predicting Cognitive Load from Speech and Interaction Dynamics in Dyadic Conversations

    Tahiya Chowdhury

    cs.LG

    Estimating cognitive load from speech has largely been studied in controlled laboratory settings, with limited understanding of its reliability in natural collaborative conversations. We investigate whether speech and interaction dynamics predict perceived cognitive load during dyadic conversations. We analyze audio from 53 dyads performing nine collaborative tasks and extract static acoustic, dynamic, and interaction features to train a...

    arxiv.org/abs/2606.12971 · PDF

  61. 61

    Circuit Synchronization Precedes Generalization: Causal Evidence from Fourier Structure in Grokking Transformers

    Achyuthan Sivasankar

    cs.LG · cs.NE

    Grokking -- where a transformer on modular arithmetic suddenly transitions from near-chance to near-perfect validation accuracy -- is attributed to a Fourier circuit, but its timing, causal structure, and controllability remain poorly understood. We introduce the Frequency Synchronization Degree (FSD), a normalised, permutation-tested metric for Fourier circuit synchronisation requiring no prior circuit knowledge. Across nine modular addition...

    arxiv.org/abs/2606.12966 · PDF

  62. 62

    Is Spurious Correlation Removal Always Learnable?

    Yibo Zhou, Bo Li, Hai-Miao Hu, Hanzi Wang, Xiaokang Zhang, Ruifan Zhang

    cs.LG

    Invariant learning can fail even when the invariant structure is statistically identifiable. We show a conditional computational barrier: under a black-box samplable supervised sparse recovery primitive motivated by average-case sparse-recovery reductions, there exist \emph{samplable} multi-environment instances with a one-dimensional predictive invariant subspace ($k=1$) that are learnable with polynomial samples by exhaustive search, while...

    arxiv.org/abs/2606.12930 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.