cs.LG · 2026-09-28 · No. 127
Machine Learning, 2026-09-28.
57 new papers in cs.LG. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
57 entries-
01
User Model Extraction via Belief Self-Distillation
Ali Holmov, Yiran Huang, Kirill Bykov, Zeynep Akata
cs.LG · cs.CL
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling...
-
02
New LoRA Skills Should Read but Never Write
Zeyan Li, Panqi Yang, Qirong Guo, Shengda Zhuo, SIyuan Qiu, Hu Xu, Chun Li, Jianfeng Xu
cs.LG
Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly....
-
03
Common-Mode Collapse and Recovery in Direct Feedback Alignment
Varun Reddy, Bernardo L. Sabatini, Houman Safaai
cs.LG · cs.NE
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching...
-
04
Trust Guided Decision Transformer
Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath
cs.LG
Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each...
-
05
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodolà
cs.LG
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a...
-
06
Generalization behavior of OPTQ and the role of regularization
Erin George, Rayan Saab
cs.LG
Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dataset is as small as possible. We study the performance of OPTQ and a variant algorithm, stochastic OPTQ, in a generalization setting and derive...
-
07
Online Learning via Learned Latent Bayesian Tracking
Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone
cs.LG · eess.SP
Online learning in non-stationary environments requires models to adapt rapidly from streaming data under strict computational constraints. A principled approach casts online learning as Bayesian state tracking, where model parameters are updated sequentially via Bayesian filtering. However, applying Bayesian filters directly to modern deep models is computationally prohibitive due to the high dimensionality of parameter space, forcing...
-
08
BeatGraph: Self-Supervised Heartbeat Graphs for Infant ECG Representations from the Home Environment
Mohammad Nur Hossain Khan, M. S. Krafczyk, Beverly G. Bolster, Nancy McElwain, Mark A. Hasegawa-Johnson, Bashima Islam
cs.LG
Electrocardiogram (ECG) foundation models typically tokenize the signal into fixed-length patches that ignore cardiac structure, so a patch may split a heartbeat and the number of beats in each patch shifts with heart rate. This matters most for infants, whose heart rates are higher and whose ECG differs from the adult, clinic-recorded 12-lead data these models are built on. A model for infant ECG should therefore reason about heartbeats...
-
09
NEXT: Physics-Informed Neuro-Spectral Exponential Time Differencing Architectures
Márcio Marques, Leonardo Mendonça, Leonardo M. Moreira, Christian Júnior de Oliveira, Vitor Balestro, Tiago Novello,...
cs.LG · math.NA
Physics-Informed Neural Networks (PINNs) build neural representations of time-dependent PDE solutions, naturally incorporating physics knowledge and observational data, which makes them well suited to both forward and inverse PDE problems. PINNs, however, are known to suffer from spectral bias and lack of causality. Neuro-Spectral Architectures (NeuSA), a recently proposed alternative to PINNs, mitigate both issues, but their numerical...
-
10
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng, Weiqiang Zhu, Zhenhai Ji, Zhengning Wang
cs.LG
Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the mapping from agents and coalitions to value components as interactions or active agents evolve. We refer to this...
-
11
Scaffold: Support Graph Theory Based Sparsification for Graph Neural Networks
Siddhartha Shankar Das, Sai Karthik Navuluru, S M Ferdous, Ryan A. Rossi, Baris Coskunuzer, Lakshman Tamil, Edoardo...
cs.LG
Graph neural networks (GNNs) rely on message passing over graph edges, making their computational and memory costs strongly dependent on graph density. Graph sparsification offers a natural way to reduce these costs, but removing edges indiscriminately can distort important communication structure and degrade predictive performance. We introduce Scaffold, a topology-based, unsupervised graph sparsification framework derived from support graph...
-
12
Uncertainty-Aware Federated Learning for Infant Movement Analysis
Edmond S. L. Ho
cs.LG · cs.AI · cs.CV
Infant movement analysis provides valuable biomarkers for the early identification of neurodevelopmental disorders. Recent advances in deep learning have enabled automated analysis of infant movements from video-derived skeletal representations, achieving performance comparable to expert assessment for tasks such as General Movement Assessment (GMA). However, most existing approaches rely on centralized training, requiring data from multiple...
-
13
Different Corruptions, Different Signals: Uncertainty and Loss in Federated Data Quality
Bradley Scott, Zeqi Luo, Edmond S. L. Ho
cs.LG · cs.AI · cs.CV
Federated learning (FL) data corruption can affect either inputs or labels, but it remains unclear whether input-conditional uncertainty and prediction-label loss expose these corruption modes equally. This paper compares two corruption-detection signals in FL: input-conditional uncertainty and prediction-label loss. The uncertainty signal is characterised using a learned aleatoric variance estimate together with Monte Carlo (MC) dropout...
-
14
Evaluating the accuracy of KV cache reuse techniques
Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona
cs.LG
Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fail to faithfully capture the loss of accuracy attributable to reuse, often artificially inflating the reported effectiveness. We also show that existing datasets do not exhibit the reuse dynamics needed to...
-
15
Decodable In-Context State and Model Output Across Training
Manas Venkata Sai Ravulapalli, Samrath Singh Chadha
cs.LG
Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. Probe accuracy rises during Pythia pretraining, while probe-guided steering moves from negligible all-trial benefit to a larger benefit at two model sizes. Saved scores distinguish...
-
16
Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition
Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen
cs.LG · eess.SP
Multimodal emotion recognition (MER) increasingly pairs EEG with speech, treating internal neural signals and external vocal expression as informative views of affect. In practice, naive fusion underperforms the stronger single modality, because EEG artifacts inject noise that corrupts the shared representation. We introduce EmoSpeechBrain, a multimodal framework built on the insight that noise suppression is a precondition for effective...
-
17
Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness
Bin Li, Dongdong Wang, Siyang Lu
cs.LG · cs.CR
Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate these factors, we conduct a systematic empirical analysis across three public log datasets, examining different adaptation strategies, model architectures, parameter scales, and...
-
18
Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
Alireza Abdollahpoorrostam, Ehsan Sharifian, Buse Şen, Marco Cuturi, Daniel Kuhn
cs.LG · math.OC · stat.ML
Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulation in which the adversary may choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. We show that the adversary's problem can be...
-
19
Progressive Memory Transformer: Memory-Aware Attention for Time-Series
Tord Sture Stangeland, Andreas Köhler, Steffen Mæland, Adín Ramíres Rivera
cs.LG
Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning...
-
20
Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design
Fu Feng, Ruixiao Shi, Yucheng Xie, Jing Wang, Xin Geng
cs.LG
Morphology--control co-design jointly optimizes an agent's body structure and control policy as an integrated embodied system. However, existing methods typically model morphology design and control with separate networks coupled only indirectly through a shared task objective, limiting explicit high-level coordination. Inspired by natural genes that coordinate biological development, we introduce \textbf{Morphogene}, a compact latent...
-
21
More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting
Lewei Xie, Haoyu Zhang, Jiajun Zhou, Yulong Chen, Guanxing Chen, Yu-An Huang, Hau-San Wong, Yifan Zhang, Zhi-An Huang
cs.LG
Continual spatio-temporal forecasting supports traffic management and environmental monitoring under evolving dynamics and expanding sensor networks. However, conventional graph-based continual learning methods tie forecasting representations to the current sensor layout, so sensor expansion can alter the representation of learned spatial relationships. Our key insight is that sensor expansion changes the evidence available about a process...
-
22
LUCID: Learning Under Confounding for Inference and Discovery in Time Series
Mohammad Fesanghary
cs.LG · stat.ML
Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direct edges. We propose LUCID (Learning Under Confounding for Inference and Discovery, a regime-adaptive deconfounding layer that first estimates the confounding regime from data using a Marčenko--Pastur spectral router, then applies a deconfounding strategy matched to that regime. When the spectrum...
-
23
Benchmarking Attention for Tabular Foundation Models
Maximilian Schambach, Clemens Biehl, Sam Thelin
cs.LG
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden...
-
24
Softmax Reparameterization for Output-Head Quantization
Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng, Lan Yan, Priya Shanmugasundaram, Tracy Holloway King
cs.LG · cs.AI
Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. This...
-
25
Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization
Mahesh Reddy Pagadala
cs.LG · cs.DC
We report fine-grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory-constrained DNN training, measured on the reference DTR simulator (simrd) using public execution traces. On an LSTM trace, memory budgets differing by 0.10% of unconstrained peak memory select fast and slow execution regimes whose overheads differ by as much as 7.3x; the slow regime is driven by broadly repeated...
-
26
Budgeted Quotient-Residual Guidance for Frozen Pocket-Conditioned Molecular Diffusion
Xinyu Wang, Jinbo Bi, Minghu Song
cs.LG
Pocket-conditioned molecular diffusion updates ambient atom coordinates, but many lead-optimization objectives are expressed on quotient features such as distances, contacts, and anchored substructures. We introduce budgeted quotient-residual guidance (QRG), an inference-time correction that makes these quotient objectives active without retraining the molecular generator. QRG lifts quotient covectors to metric-horizontal ambient directions...
-
27
Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra
Matthieu Le Lain, Gaël Cessateur, Sébastien Lefèvre
cs.LG · physics.space-ph
Auroral spectrographs such as the Auroral Spectrograph In Skibotn (ASIS) record hundreds of thousands of emission spectra, but only a few hundred can be labelled by an expert. To exploit the rest, we pretrain a 1D Vision Transformer with a masked autoencoder on 223,000 unlabelled spectra. Without labels, its representation recovers the emission-line intensity ratios that physicists use to diagnose the precipitating particles (R^2 0.91 vs....
-
28
ALF: An Active Learning Framework for Scientific Discovery
Shikha Surana, Alex Hawkins-Hooker, Olivia Gallup, Christoph Brunken, Jules Tilly, Paul Duckworth
cs.LG
Machine learning for scientific discovery is almost systematically data bound. Producing relevant high quality data, under budget constraints, is amongst the most promising ways to advance the field. Active learning (AL) offers promise wherever labelling requires expensive experiment, measurement, or simulation. Most existing tools cover only part of the data acquisition loop, and typically focus on either offline benchmarking or online...
-
29
Audio emotion recognition for atypical hearing
Ulysse Roussel
cs.LG
My doctoral work aims to explore Audio Emotion Recognition (AER) in the context of atypical listening. This research focuses on auditory hypersensitivity in people with autism, a phenomenon that is often difficult to evaluate and unique to each individual. Our core idea is to leverage our understanding of affect from acoustic traits, relying on the possibility of generalizing affective responses from a small amount of annotated data. As a...
-
30
WorldTS: World Modeling for Multimodal Covariate-aware Time Series Forecasting
Yuhan Zhu, Xiangfei Qiu, Hanyin Cheng, Wangmeng Shen, Chenjuan Guo, Bin Yang, Jilin Hu, Christian S. Jensen
cs.LG
Time series forecasting is typically framed as learning a direct mapping from historical to future observations in the observation space. However, sequences of observations generally provide only a partial view of the dynamics of the underlying system, with future observations being shaped by latent dynamics. Recent latent-space forecasting methods thus achieve improved performance by predicting future observations from latent-space...
-
31
I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?
Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
cs.LG
Recent empirical and theoretical advances suggest that joint-embedding predictive architectures (JEPAs) may learn meaningful representations for action-conditioned prediction of future outcomes, thus becoming one of the foundational structures for world models. However, accurate prediction does not, in general, necessarily imply recovery of underlying causal states that give rise to the observed dynamics. This work investigates when and how...
-
32
Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection
Jianan Liu, Chunguang Li
cs.LG
Multi-dimensional time series, inherently tensorial, are common in practice. Despite great progress in time series anomaly detection, most existing methods are confined to uni-/multi-variate time series. When handling multi-dimensional time series using these methods, reshaping operations are required, which inevitably break the intrinsic correlations and thus lead to performance degradation. In uni-/multi-variate time series anomaly...
-
33
Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift
Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
cs.LG · cs.AI
Compressing a trained model yields a family of deployment candidates, and under domain shift the most compressed one need not be the one to deploy. We study selection over such a family, with candidates and teacher fixed and target labels absent or scarce. Two findings organize the label-free case. Minimum teacher distortion behaves almost as a constant rule, selecting the same eight-bit, per-channel, unclipped configuration in every run,...
-
34
CRNDiff: Count-Native Diffusion Framework via Chemical Reaction Networks
Yuxuan Qiu, Praful Gagrani, Tetsuya J Kobayashi
cs.LG
Scientific measurements such as single-cell RNA (scRNA) sequencing often take the form of nonnegative integer counts, whereas continuous-state diffusion models approximate this discrete structure using continuous coordinates. Building on stochastic chemical reaction networks (CRNs), a class of count-native Markov jump processes, we introduce CRNDiff, a structured framework that combines count-space diffusion with inference-time conditioning...
-
35
Frame the adversary: a structure-aware attack methodology
Vicky Kouni, Stelios Perrakis, Francis Bach, Pascal Frossard, Yann Chevaleyre
cs.LG
Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical and security-sensitive applications. Yet, existing approaches are typically not derived as solutions to an optimization problem that explicitly...
-
36
From Shortcut Learning to Discrete Neural Insertion Sort
Konstantinos Mylonas, Thrasyvoulos Spyropoulos
cs.LG · cs.AI
Neural algorithmic reasoning aims to train neural networks to follow known algorithms and generalize beyond the input sizes seen during training. However, correct final outputs and intermediate supervision do not necessarily show that a model follows the intended execution. We study this problem using insertion sort. Our analysis of the CLRS30 baseline NAR shows that the hint objective is weakly optimized and that hint accuracy remains low....
-
37
Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods
Saksham Kiroriwal, Julius Pfrommer, Jürgen Beyerer
cs.LG · cs.AI · cs.IT
We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and...
-
38
The Residual Stream's Effective Depth
Barak Gahtan, Ido Galil, Alex M. Bronstein
cs.LG
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only language models, $\Deff$ separates a structural consequence of residual accumulation from an empirical one: even maximally diverse orthogonal updates...
-
39
Block Sparse Attention with Log-Linear Complexity
Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu
cs.LG
Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query-block pairs and therefore remains quadratic in sequence length. To address this issue, we propose PISA, a block-sparse attention mechanism that employs a pyramid Top-$K$ selection...
-
40
SAGE: A sampling-aware global evaluation benchmark for species distribution modeling
Emilia Arens, Nina van Tiel, Robin Zbinden, Damien Robert, Lukas Drees, Chiara Vanalli, Benjamin Kellenberger,...
cs.LG · q-bio.PE · stat.ML
Knowing where species occur is fundamental for biodiversity research and conservation. Species distribution models (SDMs) link species observations to environmental conditions to estimate their spatial distribution. However, accuracy varies with the underlying data and models, making it essential to know for which species models can be trusted. Deep-learning-based SDMs ("DeepSDMs") now jointly model thousands of species, drawing on hundreds...
-
41
Distributed Learning as a Service: The Developer's Perspective
Tianyue Chu, Filippo Vannella, Dimitra Tsigkari, Paula Delgado-Santos, Fernando López, Pablo Gomez Guerrero,...
cs.LG · cs.DC · cs.NI
Application developers of distributed learning services face challenges that a typical federated learning loop does not address. Specifically, the model updates can still leak private data, devices might not be able to participate in the training due to limited resources, a single aggregator might not be able to scale, and the transmissions of model weights induce a considerable bandwidth cost. This paper demonstrates DLaaS (Distributed...
-
42
Aurora-X: Built for Extreme Time Series Forecasting
Xingjian Wu, Chenjuan Guo, Xiangfei Qiu, Zhigang Hu, Hanyin Cheng, Peng Chen, Yang Shu, Jilin Hu, Bin Yang
cs.LG
Time series foundation models (TSFMs) enable cross-domain forecasting, but their development as general-purpose forecasters remains constrained by underexplored training potential and limited architectural versatility. To address these challenges, we introduce Aurora-X, a billion-scale TSFM with a progressive curriculum and a unified architecture. We first use channel-independent pretraining to learn temporal patterns, then introduce...
-
43
Robust Graph Clustering Network for Multiple Missing Data
Keyuan Qiu, Renda Han, Zhen Tang, Qiang He, Xingwei Wang, Wenxin Zhang, Guangzhen Yao, Junxin Chen, Qingjian Ni
cs.LG
Clustering on graphs where both node attributes and structural links are partially missing remains a challenging task. Existing methods typically rely on imputation-then-clustering on single-view missingness incomplete graphs, which are vulnerable to cross-view error propagation and cluster-boundary blurring under simultaneous attribute and structure missingness. To address these limitations, we propose a Robust Graph Clustering Network for...
-
44
Metacognitive Selective Ensemble for Mobile Systems
Sungmin Lee, Kichang Lee, Joonhee Lee, JaeYeon Park, Songkuk Kim, JeongGil Ko
cs.LG
Deep ensembles improve robustness in mobile sensing, but repeatedly executing many models over continuous sensor streams is costly. Selecting only a few members reduces this cost, yet adaptive selection often requires additional model execution to obtain reliable evidence about inactive candidates. We present MetaSE, an active ensemble framework that exploits short-term persistence in per-model reliability. MetaSE maintains a small active set...
-
45
Robust Successor Features
Erik Nikulski, Yamen Habib, Vicenç Gomez, Anders Jonsson, Rubén Moreno-Bote, Javier Segovia-Aguas
cs.LG
Generalization in Reinforcement Learning (RL) refers to the ability to execute close-to-optimal policies in unseen tasks after the agent has been trained on a different set of tasks. Building on the seminal work of the successor representation and further adaptations with function approximation, Transfer in RL has traditionally focused on generalizing to tasks that only differ in the reward function. A decade after the introduction of the...
-
46
The Linear Representation Hypothesis for Vision-Language-Action Models
Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han
cs.LG · cs.AI
The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vision-language-action (VLA) models, but the dynamical nature of embodied interaction introduces an additional challenge. Unlike semantic attributes commonly studied in LLMs, such as gender...
-
47
Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models
Shan Zhao, Ilija Trajkovic, Julia Kaltenborn, Yaniv Gurwicz, Peer Nowack, David Rolnick, Julien Boussard
cs.LG
Machine learning (ML) emulators provide a fast and cost-effective method to simulate climate change scenarios after being trained on Earth System Models projections. However, the black-box nature of those data-driven approaches limit the usability and trustworthiness of their outputs and in particular their use as causal attribution tools. Here, we develop a hierarchical causal representation learning framework applied to sea surface...
-
48
Does Uniform Discrete Diffusion Need Time?
Chunsan Hong, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Jong Chul Ye, Yuki Mitsufuji
cs.LG · cs.AI · cs.CL
Uniform discrete diffusion models (UDMs) commonly use explicit time conditioning, but we find that it can often be unnecessary in practice. In this paper, we first show that the population-optimal UDM predictor generally depends on time: time controls how much the model should trust the observed context. We then show that this dependence can become negligible in finite-data settings relevant to language. When a corrupted training sequence...
-
49
LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers
Natsuki Yoshino, Ren Uchida, Kazuki Matsumoto, Kohei Yatabe
cs.LG · eess.SY
Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforcing Lipschitz continuity requires each layer of a DNN to be Lipschitz continuous, thereby guaranteeing overall Lipschitz continuity. However, this layer-wise approach typically imposes...
-
50
Gradient Surgery for Physics-Informed Neural Networks
Thomas Borsani, Giuseppe Di Fatta
cs.LG
Physics-Informed Neural Networks (PINNs) are trained by optimising a composite objective that combines data fitting with physics-based constraints, typically resulting in a highly imbalanced multi-task optimisation problem. Under these conditions, existing optimisation strategies are affected by conflicting task gradients, leading to slow convergence and unstable training, particularly for stiff and high-frequency partial differential...
-
51
Towards Understanding Momentum Acceleration in River-Valley Loss Landscape
Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen, Tianhao Wang, Zhiyuan Li
cs.LG · math.OC
The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features a low-loss manifold (river) flanked by sharp orthogonal directions with higher loss (mountains). In the long term, the optimization progress is...
-
52
Low-Bit Recurrent States in Hybrid Language Models
Hongren Chen, Jiayang He
cs.LG
Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-precision bit allocation, without calibration data, rotation, or training. We also quantize decay rates logarithmically. With per-token state...
-
53
PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem
Mateo Toro Diz, Jonathan Hoss, Noah Klarmann
cs.LG · cs.AI
The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining with offline fine-tuning on production-specific data. Reinforcement learning through online interaction enables exploration of general scheduling strategies, but typically relies on...
-
54
EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting
Takumi Fujimoto, Hiroaki Nishi
cs.LG
Completed multi-horizon forecasts provide residual feedback for a fixed forecaster, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order discrete cosine transform (DCT) coefficients and the final value of the preceding residual block. Within each channel, the endpoint is shared across component-wise online ridge regressions...
-
55
Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
Marcin Kostrzewa, Maciej Zięba
cs.LG · cs.AI
Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We...
-
56
CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters
Yuhang Cao, Yanzhou Mu, Chunrong Fang, Zhenyu Chen
cs.LG
Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current model outputs, while complete affected suffix recomputation restores fidelity at substantial cost. We seek minimal recomputation that recovers current adapter behavior. Existing systems track token, context, or...
-
57
Learning Chance-Constrained MDPs with Bellman Distributional Certificates
Chenbei Lu, Hongyu Yi
cs.LG
Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this...
This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.