cs.LG · 2026-09-10 · No. 111

Machine Learning, 2026-09-10.

60 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

60 entries
  1. 01

    A positive resolution of the gap-entropy conjecture

    P. M. Aronow, Nathan Kallus, Patrick Lopatto

    cs.LG · stat.ML

    We prove the gap-entropy conjecture for fixed-confidence best-arm identification with independent unit-variance Gaussian arms, means in $[0,1]$, and a unique optimal arm. For each suboptimal arm $i$, let $Δ_i=μ_*-μ_i$ be its gap from the optimal mean, and write $H=\sum_{i\ne *}Δ_i^{-2}$. Let $p_r$ be the fraction of $H$ contributed by arms with $2^{-(r+1)}<Δ_i\le2^{-r}$, and let $\mathrm{Ent}(I)=\sum_{r:p_r>0} p_r\log(1/p_r)$. Among all...

    arxiv.org/abs/2609.10529 · PDF

  2. 02

    Quantum Feature Engineering for Credit Default Prediction: When and Why IQP Circuits Help Linear Classifiers

    Menachem Finkelstein, Diana Legziel Levy, Zohar Yakhini, Sarel Cohen

    cs.LG · quant-ph

    Credit default prediction is a tabular classification problem in which modest gains in F1 translate directly into reduced financial exposure. We ask whether Instantaneous Quantum Polynomial-time (IQP) circuits can produce features that improve a classifier over both its raw classical baseline and Kernel PCA - the strongest unsupervised classical non-linear alternative - at an equal feature budget. The dataset provides 23 financial attributes...

    arxiv.org/abs/2609.10505 · PDF

  3. 03

    Learning with Covariance Matrices: Principal Component Analysis Meets Learning with Graphs

    Saurabh Sihag, Andrea Cavallo, Elvin Isufi, Gonzalo Mateos, Alejandro Ribeiro

    cs.LG · eess.SP

    This feature article provides an overview of the theoretical foundations for coVariance neural networks (VNNs), i.e., graph neural networks (GNNs) operating on covariance matrices as graphs. Covariance matrices are ubiquitous across domains, and hence, the deployment of GNNs often leverages graphs of pairwise statistical dependencies. Existing theoretical contributions on GNNs consider abstract graph representations and cannot accommodate the...

    arxiv.org/abs/2609.10490 · PDF

  4. 04

    Nonmaximal sums of maximally monotone operators under Rockafellar's constraint qualification

    Weifeng Yang

    cs.LG · math.FA

    We construct counterexamples to Rockafellar's sum conjecture in which two maximally monotone operators satisfy the interior-domain condition but their sum is not maximally monotone. We give one counterexample on $c_0$ and another on $\ell^1$ with its usual norm. We establish a general construction theorem that computes the entire monotone polar of a class of graphs, gives a necessary and sufficient condition for their maximal monotonicity,...

    arxiv.org/abs/2609.10487 · PDF

  5. 05

    Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization

    Andy Zeyi Liu, Haoran Sun, Lucas Baker, Randall Balestriero, John Sous

    cs.LG · cs.AI · cs.CV

    Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning, but their capability to learn physics and generate physically realistic dynamics remains hitherto untested. In this work, we introduce SemiGroup-JEPA (SG-JEPA), which extends the LeWorldModel framework by supplying the parameter governing the physics to the temporal model via action-conditioning...

    arxiv.org/abs/2609.10464 · PDF

  6. 06

    Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

    Ravi Ranjan, Olivera Kotevska, Agoritsa Polyzou

    cs.LG · cs.AI

    Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable training content, creating privacy, safety, and regulatory concerns. Machine unlearning offers a practical alternative to full retraining, but many existing methods apply broad or fixed parameter updates that can degrade utility and remain brittle under deployment changes such as post-training quantization, where forgotten knowledge may...

    arxiv.org/abs/2609.10439 · PDF

  7. 07

    OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis

    Ayush Debnath, Ruelia Saha, Sudip Misra

    cs.LG · cs.AI

    Simultaneous assessment of medical imaging and patient records is often required in clinical diagnosis. However, standard machine learning algorithms cannot analyze these data types together. Meanwhile, compliance with HIPAA and GDPR can constrain centralized aggregation of sensitive patient data. This leaves a crucial void of secure fusion of visual and textual context across distant networks. Thus, we present OmniMed-FL, a controlled...

    arxiv.org/abs/2609.10364 · PDF

  8. 08

    A Later Test Set Is Not a New Domain: Pretraining Familiarity Survives a Contamination-Free Hold-Out

    Mahdi Naser Moghadasi, Faezeh Ghaderi

    cs.LG

    Time-series foundation models are evaluated almost exclusively on public archives that predate them, so a strong score cannot be separated from having seen the test set during pretraining. The obvious remedy is a hold-out that postdates the models. We build one: thirteen forecasters -- four classical, three trained per dataset, six pretrained -- on seven groups drawn from five domains, every observation published after the last model was...

    arxiv.org/abs/2609.10357 · PDF

  9. 09

    One Loop, Two Gains: Can Active Learning win the Lottery for Free?

    Benedikt Tscheschner, Eduardo Veas, Marc Masana

    cs.LG · cs.AI · cs.CV

    The lottery ticket hypothesis posits the existence of winning tickets: sparse subnetworks that, when trained in isolation from their original initialization, match the accuracy of the full dense network. The predominant method for discovering such tickets, iterative magnitude pruning, alternates pruning with full retraining from scratch until convergence over many cycles. Similarly, deep active learning also retrains a model from scratch...

    arxiv.org/abs/2609.10311 · PDF

  10. 10

    View-Structured Conformal Prediction for 3D Gaussian Splatting

    Junzheng Chu, Bin Pan, Zhenwei Shi

    cs.LG · cs.CV

    3D Gaussian Splatting (3DGS) renders novel views in real time, but an uncertainty heatmap does not certify that a rendered view meets a certain prediction coverage. We treat novel-view synthesis as structured regression and ask that, with probability at least $1-α$, RGB prediction boxes cover at least a $1-β$ fraction of pixels in a new view. We propose View-Structured Conformal Prediction (VSCP). It splits the pre-calibration scale into a...

    arxiv.org/abs/2609.10307 · PDF

  11. 11

    A Dominant Diffuse Phase in the Sparse Autoencoder Phase Diagram

    Alexis D. Plascencia

    cs.LG

    Sparse autoencoders (SAEs) are increasingly used to recover interpretable features from neural-network activations, yet systematic feature co-occurrence can cause distinct features to be absorbed or merged. The MAIS-O43 open problem proposes a controlled experiment to characterize when recovery of a true synthetic dictionary gives way to feature merging as the nesting fraction $γ$, sparsity penalty $λ$, and dictionary size $M$ vary. We...

    arxiv.org/abs/2609.10299 · PDF

  12. 12

    Training Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers

    Zonglin Yang, Ziming Zhao, Wei Tang, Xunyu Jiang, Yihong Liu, Tailin Chen, Zifu Yu, Jiayu Liu

    cs.LG · cs.NE

    Soft positional priors can help small Transformers learn retrieval circuits, but it is unclear whether the resulting circuits remain functional once the prior is removed. We test this with an annealable soft-prior Transformer whose attention biases can be learned, faded, or zeroed during training and evaluation. On associative recall, unforced models perform well with the prior active ($0.772 \pm 0.020$) but collapse at zero gate ($0.095 \pm...

    arxiv.org/abs/2609.10287 · PDF

  13. 13

    Hierarchical and Permutation-Invariant Feature Transformation Learning via Policy-Guided Embedding Search

    Rui Liu, Tao Zhe, Yanyong Huang, Sankha Narayan Guria, Xiao Luo, Wei Fan, Yanjie Fu, Dongjie Wang

    cs.LG · cs.AI

    Feature transformation improves predictive performance on tabular data by constructing informative abstractions from raw features. Recent generative approaches encode transformation knowledge into continuous embedding spaces for efficient exploration of candidate strategies, but face three key limitations: (1) overlooking hierarchical relationships between low-level features, operations, and high-level abstractions; (2) enforcing...

    arxiv.org/abs/2609.10225 · PDF

  14. 14

    Robust Beam Prediction for V2X Networks with Multi-Modal Sensing

    Chen Shang, Dinh Thai Hoang, Diep N. Nguyen, Jiadong Yu

    cs.LG

    Integrated sensing and communication (ISAC) provides a promising foundation for beam prediction in future vehicle-to-everything (V2X) networks. However, existing sensing-assisted beamforming methods still rely heavily on radio-frequency sensing, which may become unreliable in complex vehicular environments. Meanwhile, the growing availability of heterogeneous sensors, such as cameras and LiDAR, offers new opportunities to improve beam...

    arxiv.org/abs/2609.10200 · PDF

  15. 15

    An Exponential Deterministic--Randomized Gap in ERM-Oracle Complexity for Thresholds on an Unknown Order

    Xuan Li

    cs.LG · stat.ML

    Attias, Hanneke and Ramaswami (NeurIPS 2025) asked whether randomization provably reduces the oracle calls needed for online learning when the class is accessible only through an oracle. We study the instance they singled out: transductive online learning of thresholds on an unknown total order of T instances, with a consistency-type ERM oracle that returns a full concept consistent with a queried labeled set (or reports non-realizability)....

    arxiv.org/abs/2609.10196 · PDF

  16. 16

    CoGe-GCD: Reframing Generalized Category Discovery with Compositional Generalization

    Luyao Tang, Jiewei Zheng, Kunze Huang, Chaoqi Chen, Yue Huang, Cheng Chen

    cs.LG

    Generalized Category Discovery (GCD) assigns unlabeled instances, mixed with labeled data, to known or novel categories, requiring human-like compositional reasoning: reusing primitives learned from known classes and deciding when new combinations imply new categories. Existing GCD methods operate on unstructured token features and struggle to extrapolate to novel compositions. We propose CoGe-GCD, which rethinks GCD through compositional...

    arxiv.org/abs/2609.10158 · PDF

  17. 17

    CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts

    Naibin Gu, Qingyi Si, Chenxu Yang, Chuanyu Qin, Junhao Zhou, Peng Fu, Zheng Lin, Weiping Wang

    cs.LG

    On-policy distillation (OPD) provides dense token-level supervision on student-generated trajectories. Although OPD performs strongly when teacher and student belong to the same model family, we find that its effectiveness degrades in cross-family settings even after tokenizer alignment, with substantially stronger external teachers offering little additional improvement. To understand this disconnect, we decompose the cross-family OPD signal...

    arxiv.org/abs/2609.10154 · PDF

  18. 18

    Storage-Scalable Progressive Semantic Communication via Knowledge-Base Reuse

    Heng Zhu, Ye Liu, Kun Zhu, Feifei Song

    cs.LG · cs.NI · eess.IV

    Existing knowledge-base-assisted semantic communication schemes commonly adopt either single knowledge-base quantization (SKBQ) or multi-knowledge-base residual quantization (MKBQ). SKBQ incurs limited storage overhead but has restricted quantization capacity, whereas MKBQ supports progressive refinement by assigning an independent knowledge base (KB) to each stage, causing the KB storage to grow linearly with the transmission depth. To...

    arxiv.org/abs/2609.10112 · PDF

  19. 19

    A Trust-Network-Based Federated Learning Framework for Multi-Center Aging Clock Prediction

    Chunxu Zhang, Bo Li, Wenliang Wang, Yang Liu, Di Jiang, Yuan Huang, Yo-ichi Nabeshima, Akinori Yamamura, Bo Yang, Qiang Yang

    cs.LG · cs.AI

    Aging clocks quantify biological aging and help characterize individual health status. What protein interactions are important for accurate aging clocks, and are they zeroth-order or higher-order? Addressing these questions requires learning from large molecular datasets distributed across medical centers, where privacy constraints prevent centralized data sharing. Federated learning offers a natural solution but faces four challenges in this...

    arxiv.org/abs/2609.10108 · PDF

  20. 20

    A Systematic Evaluation of Molecule Generation Models for De Novo Drug Design: From Benchmarks to Practical Insights

    Xinrui Xu, Xueer Wang, Dan Luo, Sisi Yuan, Xuan Lin

    cs.LG

    Molecule generation has emerged as a powerful computational tool for de novo drug design, enabling the exploration of chemical space beyond the limits of conventional virtual screening. The field has progressed rapidly, driven by advances in molecular representations, generative architectures, and target-aware modeling strategies. However, existing reviews typically address specific model families or application scenarios in isolation, rather...

    arxiv.org/abs/2609.10099 · PDF

  21. 21

    Hybrid Quantum-Classical NLP Classification with Compact Semantic Representations: An Experimental Analysis of Representation Compression

    Ali Hassan, Zijia Zhao, Maha A. Metawei

    cs.LG

    Large language and sentence-embedding models provide rich semantic representations, but their high dimensionality poses a challenge for near-term quantum machine learning (QML), where quantum circuits can process only a limited number of input features. We investigate a hybrid quantum-classical pipeline that transforms high-dimensional sentence embeddings into compact representations for variational quantum classification. The workflow...

    arxiv.org/abs/2609.10089 · PDF

  22. 22

    Field-level prediction of mid-plane stress tensor fields in concrete target penetration: a cross-velocity graph neural operator surrogate

    Wenpu Du, Peng Zhou, Yunlong Xia, Sinuo Xin, Congcong Zhang, Boyang Zhang, Yi Zhang, Wenzheng Xu

    cs.LG

    Although the impact resistance of concrete has been studied extensively, a framework linking mesoscale heterogeneity to full-field stress-tensor prediction has been lacking. Data were generated with a full-scale aggregate-resolved LS-DYNA model (projectile diameter 45 mm, mass 2.13 kg, target diameter 500 mm x thickness 200 mm, mesh 10 mm), verified against published penetration experiments (Frew 2006, Hanchak 1992, Forrestal 1996) by...

    arxiv.org/abs/2609.10032 · PDF

  23. 23

    Beyond Contact Sensors: Deep learning with Pseudo-Labeling for remote Photoplethysmography

    Bhargav Acharya, Barbara Hammer, Hanna Drimalla

    cs.LG

    Heart rate is a critical biomarker of health, and remote photoplethysmography (rPPG) enables its contactless estimation from video data for telemedicine applications. Recent advancements in deep learning based rPPG methods achieve state-of-the-art results, outperforming classical signal-processing methods in complex scenarios. However, deep learning methods depend on datasets with precise synchronization between videos and ground truth...

    arxiv.org/abs/2609.10026 · PDF

  24. 24

    Structure-Aware Unsupervised Anomaly Detection for Spacecraft Telemetry with Adaptive EVT Thresholding

    Óscar Alcarria, Rafael Sánchez, Javier Sempere, Pablo Torrijos, Juan C. Alfaro, Juan M. Auñón, José A. Gámez, José M. Puerta

    cs.LG

    Operational anomaly detection in spacecraft telemetry typically requires labeled historical anomalies or extended warm-up periods. These requirements are rarely met in practice. We propose an unsupervised, deployment-ready framework that produces predictions from the second month of operation without any labels, prior fault knowledge, or mission-specific tuning. The approach combines incremental monthly retraining, statistical model...

    arxiv.org/abs/2609.10017 · PDF

  25. 25

    MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

    Remco Hendriks

    cs.LG · cs.AI · cs.CL

    We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote...

    arxiv.org/abs/2609.10016 · PDF

  26. 26

    An Explainable Machine Learning Framework for Predicting Blood-Brain Barrier Permeability Using Molecular Descriptors

    Fatemeh Mahmoudi

    cs.LG · cond-mat.mtrl-sci

    Blood-brain barrier (BBB) permeability is a critical determinant in the development of central nervous system therapeutics because it directly influences the ability of drug candidates to reach their target sites within the brain. In this study, an explainable machine learning framework was developed to predict BBB permeability using molecular descriptors generated from the MoleculeNet BBBP dataset with the RDKit cheminformatics toolkit....

    arxiv.org/abs/2609.10012 · PDF

  27. 27

    Deep Neural Networks for Learning Intent from sEMG Signals to Support Hardware Devices for Post-Stroke Neurorehabilitation

    Zakariyya Brewster, Divy Wadhwani, Emily Yan, Aidan Wang, Karma Namgyal, Shuting Xie, Markiyan Konyk, Tala Abdelmaguid

    cs.LG

    Finger-specific motor intent is a clinically meaningful control signal for post-stroke neurorehabilitation, where residual muscle activity may remain measurable despite weak or incomplete movement. We study five-finger multilabel intent decoding from impaired-arm high-density surface electromyography (sEMG) in PhysioMio, a bilateral longitudinal dataset collected from stroke patients. A common processing protocol aligns movement labels,...

    arxiv.org/abs/2609.09971 · PDF

  28. 28

    Adversarial Training for Tabular Credit Scoring: A Multi-Attack Robustness Evaluation in P2P Lending

    Gijs A. F. Niewzwaag, Marijn G. S. Veth, Manuele Massei, Marcos R. Machado

    cs.LG · q-fin.RM · stat.ML

    Machine learning-based credit scoring is increasingly central to Peer-to-Peer (P2P) lending, yet its resilience to adversarial manipulation, where applicants strategically alter self-reported inputs to secure favourable decisions, remains poorly understood. Most adversarial-robustness evidence comes from image and text domains and evaluates a single attack against a matching defence, offering little guidance on how defences generalise across...

    arxiv.org/abs/2609.09945 · PDF

  29. 29

    Multi-Pass, Multi-View Blended Learning for High-Fidelity Volumetric CT Synthesis from Chest X-Rays

    Ozer Can Devecioglu, Serkan Kiranyaz, Rashid Mazhar, Tahir Hamid, Muhammad Chowdhury, Moncef Gabbouj

    cs.LG

    Reconstructing volumetric Computed Tomography (CT) from a single 2D chest radiograph (CXR) is an ill-posed inverse problem, further complicated by the scarcity of paired CXR-CT training data. Prior approaches address this by training on Digitally Reconstructed Radiographs (DRRs), which are synthetic projections derived from CT volumes. However, the domain gap between DRRs and real CXRs limits generalization, often resulting in coarse or...

    arxiv.org/abs/2609.09920 · PDF

  30. 30

    Development and Validation of a Physics-Guided Machine Learning Extrapolation Framework Using a Classical Transient Diffusion Benchmark

    Ashutosh Yadav, Alok Dubey, Prodyut Ranjan Chakraborty, Harshal Akolekar

    cs.LG

    Machine learning models used in engineering are typically trained within limited operating ranges, yet reliable predictions are often required beyond these domains. Consequently, the primary challenge is extrapolation rather than interpolation. Rigorous validation is hindered by the scarcity of data outside the training range. To address this limitation, a novel extrapolation framework is integrated with established machine learning...

    arxiv.org/abs/2609.09912 · PDF

  31. 31

    A Kernel-Based Modular Discriminant Analysis Framework for Small-Sample Learning

    Lingxiao Qu, Yan Pei

    cs.LG

    The small-sample-size (SSS) problem remains a fundamental challenge in machine learning when labeled data are scarce due to cost, accessibility, or ethical constraints. While numerous approaches have been proposed, existing methods often struggle to maintain stable and discriminative representations under high-dimensional and limited-data conditions. Kernelized Linear Principal Component Discriminant Analysis (KLPCDA), a recently proposed...

    arxiv.org/abs/2609.09910 · PDF

  32. 32

    Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits

    Hao Li, Jie Xu, Zheng Xie

    cs.LG · math.OC

    Meta-learning has emerged as an effective paradigm for transferring knowledge across sequential bandit tasks. While substantial progress has been made for stochastic bandits and non-contextual adversarial bandits, meta-learning for adversarial linear contextual bandits (ALCBs) with random action sets remains largely unexplored. To address this problem, we propose Meta-LinEXP3, an online-within-online algorithm that constructs a predictable...

    arxiv.org/abs/2609.09907 · PDF

  33. 33

    Beyond Conventional Federated Learning via High-Order Regularization

    Alireza Kabgani, Masoud Ahookhosh

    cs.LG · stat.ML

    Federated clients that perform several local optimization steps can return parameter displacements with widely different magnitudes. The quadratic regularization of FedProx grows linearly with displacement and therefore offers limited control over the contrast between ordinary and unusually large client movements. We here introduce HiFedProx, which replaces the quadratic penalty with a scale-matched power-type regularizer indexed by $p\geq2$....

    arxiv.org/abs/2609.09904 · PDF

  34. 34

    Strangers to Themselves: What Language Models Say About Themselves Is Generic

    Phil Blandfort, Urja Pawar

    cs.LG · cs.AI · cs.CL · cs.CV · cs.CY

    Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question....

    arxiv.org/abs/2609.09899 · PDF

  35. 35

    ProMeta: Few-shot PROTAC-targeted degradation prediction across E3 ligases

    Yuansheng Liu, Yufei Ye, Tao Tang, Jiawei Luo, Wen Tao, Xiao Luo

    cs.LG · q-bio.QM

    Proteolysis-targeting chimeras (PROTACs) have emerged as a transformative therapeutic strategy that selectively degrades historically ''undruggable'' targets via the ubiquitin-proteasome system. Despite growing efforts to develop computational predictors of PROTAC degradation activity, existing supervised approaches remain severely challenged by data scarcity and imbalance across E3 ligases, limiting their ability to generalize beyond...

    arxiv.org/abs/2609.09891 · PDF

  36. 36

    Forward-Free LLM Depth Pruning via Weight Redundancy

    Vincent-Daniel Yun, Woosang Lim

    cs.LG · cs.AI · cs.PF

    Depth pruning reduces large language model (LLM) inference cost by removing complete Transformer blocks. Activation-based methods collect hidden states through forward passes on calibration data, while existing forward-free methods score each Transformer block separately without measuring similarity between blocks. We propose Weight-Redundancy Pruning (WRP), a forward-free depth-pruning method that estimates inter-layer redundancy from...

    arxiv.org/abs/2609.09883 · PDF

  37. 37

    Exact Degeneracy Under Balanced k-Shot Sampling:Consequences for Small-Sample Discriminant Analysis on LLM Embeddings

    Lingxiao Qu

    cs.LG

    Balanced k-shot sampling draws exactly k labeled examples per class. We show that it induces an exact, provable degeneracy in a family of small-sample discriminant estimators. Under balanced sampling, the within-class scatter operator of Kernelized Linear Principal Component Discriminant Analysis (KLPCDA) is not merely rank-deficient but exactly a scaled orthogonal projector. We derive the consequences in closed form: two of KLPCDA's seven...

    arxiv.org/abs/2609.09860 · PDF

  38. 38

    TempTPI: Informer-Based trajectory prediction for maritime vessels

    Kevin Ferneding, Veronika Lietavcova, Aleksandra M. Blachowiak, Peder Heiselberg

    cs.LG

    Accurate long-term trajectory prediction for maritime vessels is essential for safety and logistical efficiency. While deep learning models, particularly Transformers, have shown promise in processing Automatic Identification System (AIS) data, they often struggle with the quadratic computational complexity of self-attention and the loss of accuracy over extended forecasting horizons. This study proposes TempTPI, a novel prediction framework...

    arxiv.org/abs/2609.09840 · PDF

  39. 39

    In Medical Claims Data, Enhancing Predictive Performance for Major Adverse Cardiovascular Events Using Cross Attention

    Yuhei Fujioka, Daitaro Misawa, Tatsuyoshi Ikenoue, Shingo Fukuma

    cs.LG

    Medical claims data comprise the financial details, including the expenses and billing information, as well as the clinical information, such as the diagnoses and treatments, of patients visiting medical facilities. Recently, it has been acknowledged that large databases can be constructed from medical claims data for medical research purposes. However, the clinical information within these datasets is often medically unstructured, limiting...

    arxiv.org/abs/2609.09824 · PDF

  40. 40

    Online Inverse Integer Linear Optimization via Small-Gradient Skipping: Constant Regret and Finite Mistakes

    Akira Kitaoka

    cs.LG · cs.DS · math.OC

    In online inverse linear optimization, the learner predicts a weight at each round, observes the optimal action of the agent, and updates its prediction. In the general setting, the gap of $\log T$ between the regret upper bound $O(d \log T)$ and the lower bound $Ω(d)$ is unresolved (here $T$ is the total number of rounds and $d$ is the dimension). When the action set is M-convex, the regret is known to be bounded by $O(d \log d)$, but the...

    arxiv.org/abs/2609.09809 · PDF

  41. 41

    A practical DIRECT-type algorithm for medium-scale black-box global optimization

    Linas Stripinis, Remigijus Paulavičius

    cs.LG · math.NA

    The DIRECT algorithm is a deterministic global optimization method known for its versatility and balanced exploration-exploitation strategy. However, DIRECT-type algorithms are primarily effective for low-dimensional problems and often exhibit slow convergence as dimensionality increases, limiting their applicability to more complex optimization tasks. To address this limitation, this paper introduces X-DTC-GL, a novel DIRECT-type algorithm...

    arxiv.org/abs/2609.09796 · PDF

  42. 42

    Privacy-Preserving Split Learning for Federated LLM Fine-Tuning

    Heng Jin, Chaoyu Zhang, Hexuan Yu, Wenjing Lou, Y. Thomas Hou

    cs.LG

    Fine-tuning large language models (LLMs) on domain-specific data is essential for downstream adaptation. In many deployments, a participant cannot hold the complete model locally. This happens because the model owner keeps the full model proprietary, or because the participant lacks sufficient compute resources. Split Learning (SL) addresses this by partitioning the model between the participant and a server so that only a small portion runs...

    arxiv.org/abs/2609.09794 · PDF

  43. 43

    Evaluating Model Retraining under Drift: Paired Comparisons of Cumulative Subgroup Disparity

    Aaron Ceross

    cs.LG

    Choosing when to retrain a deployed classifier requires assessing subgroup error rates across the sequence of models used, including periods between updates. We compare complete scheduled, loss-triggered, and subgroup-gap-triggered policies with retaining the initial model on the same observations and delayed labels. For true-positive and false-positive rates separately, the outcome is the paired difference in absolute subgroup gaps summed...

    arxiv.org/abs/2609.09788 · PDF

  44. 44

    NEXUS-MI: Communication-Aware Federated Personalization for Gateway-Coordinated Motor-Imagery Brain-Computer Interfaces

    Daniel Adu Worae, Aarthy Nagarajan

    cs.LG · cs.NI

    Electroencephalography (EEG)-based motor-imagery brain-computer interfaces (MI-BCIs) vary across subjects and sessions, complicating personalization from limited calibration data. Federated learning can exploit shared representations without centralizing raw EEG, but existing federated MI studies largely assume regular synchronization. We introduce NEXUS-MI, a gateway-coordinated federated personalization framework that treats synchronization...

    arxiv.org/abs/2609.09786 · PDF

  45. 45

    BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL

    Guanqun Zhao, Zijun Xie, Binbin Zheng, Jiafeng Lu, Enlei Gong, Zeyu Chen

    cs.LG · cs.AI

    Asynchronous reinforcement learning has become the standard way to scale training for language models, but the resulting policy lag biases the critic toward the stale behavior policy. Existing work on asynchronous LLM training corrects the actor and leaves this bias unaddressed, while the off-policy value correction of classical RL does not carry over to long-horizon agentic tasks, since a short correction horizon leaves the regression target...

    arxiv.org/abs/2609.09783 · PDF

  46. 46

    Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?

    Fumihiko Tachibana, Daisuke Miyashita, Jun Deguchi

    cs.LG · cs.AI · cs.CL

    In Retrieval-Augmented Generation (RAG) systems, a large number of retrieved chunks are concatenated to form the input context so that users can receive high-quality responses based on external knowledge. As a result, the input context length increases substantially, leading to a larger prefill workload and, in turn, a longer time to first token (TTFT). While previous works that reuse precomputed key-value (KV) caches effectively reduce TTFT...

    arxiv.org/abs/2609.09768 · PDF

  47. 47

    EEGBind: Detecting Source-Level Interictal Epileptiform Discharges via EEG-Centric Multimodal Binding

    Muchen Li, Anglin Liu, Xuetian Gao, Ruijian Xu, Jintai Chen

    cs.LG · cs.MM · q-bio.NC

    Source-level analysis of interictal epileptiform discharges (IEDs) is relevant to presurgical evaluation and treatment planning because it helps characterize where epileptiform activity is likely to arise. Beyond detecting whether an IED is present, this setting requires assigning IED-positive activity to clinically meaningful brain-region categories. This setting is challenging because source-region evidence in short electroencephalography...

    arxiv.org/abs/2609.09728 · PDF

  48. 48

    EFQ-Softmax: Exp-Free Quantization for Softmax

    Haohui Han, Yuming Wan, Hongni Wang, Pengcheng Xie, Xiaodong Yan, Runqi You, Wencong Zhang

    cs.LG

    Low-bit attention accelerates Transformer inference by moving the $QK^\top$ and $PV$ matrix multiplications to FP8 or FP4 matrix engines. However, the softmax path often evaluates shifted-score exponentials in higher precision, forms a temporary probability block, and quantizes it before low-bit $PV$ multiplication. This exp-then-quantize path creates a mismatch between a high-precision probability producer and a low-bit matrix consumer. We...

    arxiv.org/abs/2609.09721 · PDF

  49. 49

    Kernel-Complexity Edge Sanitization for Training-Free Defense against Structural Graph Attacks

    Yaning Jia, Shenyang Deng, Yaoqing Yang, Chiyu Ma, Wenxuan Xu, Soroush Vosoughi

    cs.LG · cs.AI

    Graph Neural Networks (GNNs) have achieved remarkable success across diverse applications, yet they remain highly vulnerable to adversarial attacks that maliciously perturb graph structure. Existing defenses often lack rigorous theoretical grounding, rely on attack-specific heuristics, or require costly retraining procedures such as adversarial training. To address these limitations, we propose Kernel-Complexity Edge Sanitization (KCES), a...

    arxiv.org/abs/2609.09698 · PDF

  50. 50

    ALIGN-HOLD: Experience Alignment for Real-Time Hold Control in Large-Scale Ride-Hailing Matching at DiDi

    Zuhao Zhang, Xu Liu, Kai Wan, Zihao Lu, Li Ma, Shuai Li

    cs.LG

    Real-time hold control is a high-leverage mechanism in large-scale ride-hailing systems: by selectively deferring driver-order pairs, the platform can wait for better matching opportunities and improve end-to-end passenger-driver experience. Existing production systems such as EXHOLD learn bandit-based hold policies from handcrafted combinations of trip completion, cancellations, waiting time, and driver effort. However, designing such...

    arxiv.org/abs/2609.09685 · PDF

  51. 51

    Settling: Equilibrium Inference for Non-Convex Validity Sets

    Lyes Saad Saoud

    cs.LG

    Many learning systems return a single point estimate even when admissible outputs form disconnected or non-convex sets. Under squared loss, an ambiguous conditional distribution can therefore have a Bayes-optimal conditional mean that is invalid. We formalize this failure as conditional mean collapse and introduce Settling, an equilibrium-based inference operator that separates proposal generation, consistency evaluation, and test-time...

    arxiv.org/abs/2609.09682 · PDF

  52. 52

    Muon-C: Operator-Aligned Muon for Convolutional Kernels

    Jiaxin Qing, Lexin Li

    cs.LG · stat.ML

    Muon replaces matrix momentum with an approximately orthogonal polar direction, but its geometry depends on the matrix representation. For convolution, standard unfolding describes a local patch map rather than the convolution operator. We introduce Muon-C, an operator-aligned optimizer that represents kernel momentum as frequency-wise channel-transfer matrices, polarizes these blocks independently, and uses a critical Fourier grid to return...

    arxiv.org/abs/2609.09676 · PDF

  53. 53

    PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling

    Weisi Yang, Stephen Xia

    cs.LG · cs.CL · cs.OS · cs.PF

    Deploying Large Language Models (LLMs) directly on mobile platforms at the edge is gaining traction due to a myriad of benefits, such as increased privacy, personalization, and reduced latency. However, LLMs have heavy computational requirements, which are difficult for resource-constrained mobile and edge platforms to fulfill. In addition to limited compute resources, mobile and edge systems often have a compact form factor and lack physical...

    arxiv.org/abs/2609.09662 · PDF

  54. 54

    Cascading Gradient Inversion via LT-Code Inspired Peeling in Federated Learning

    Saeed Shariati, Mohsen Alambardar Meybodi

    cs.LG · cs.AI · cs.CR

    Federated learning shares model updates rather than raw data, yet these updates can be inverted to reconstruct the clients' training data. Analytic reconstruction attacks, which invert a gradient in closed form, degrade as the batch grows: prior single-round attacks recover only about half of a batch of size $100$ even when the attacker fully controls the network parameters, and known upper bounds limit what any such method can recover. We...

    arxiv.org/abs/2609.09659 · PDF

  55. 55

    Teacher Geometry Shapes Learnability in Teacher-Student Networks

    Kai J. Sandbrink, Flavio Martinelli, Alexander van Meegen, Wulfram Gerstner, Johanni Brea

    cs.LG · cs.AI · cs.NE

    Teacher-student systems, in which a teacher neural network generates training labels so that a student neural network can learn to implement the same function, are widely used as an abstract setting to study learning. However, the structure of the teachers is often overlooked by assuming randomly-generated, normally-distributed parameters. This hides substantial variation in how learnable different teachers are. We formalize learnability as...

    arxiv.org/abs/2609.09595 · PDF

  56. 56

    Positional task conditioning for scalable defect detection across product families in large product catalogs

    Soham Satyadharma, Gabriel Roccabruna, Suleiman A. Khan

    cs.LG

    Product families in large product catalogs suffer from inconsistencies such as duplicates and unit mismatches that degrade customer experience. Detecting these requires reasoning over multiple error types across lengthy product listings, where LLM classification quality degrades due to long-context limitations. We address this by decomposing detection into focused sub-tasks that reduce context and isolate error types, improving F1 from 52\%...

    arxiv.org/abs/2609.09567 · PDF

  57. 57

    Robust Industrial Cyber Physical Classification Using Neuromorphic Temporal Embeddings and Hybrid SNN XGBoost Under Machine Unlearning Attacks

    Ammar Kamoona, Sajad Koushkbaghi, Mahdi Jalili, Peter McTaggart, Xinghuo Yu

    cs.LG · cs.CR · cs.NE

    The digitalisation of electrical distribution networks has increased the exposure of power-grid infrastructure to cyber attacks. Existing intrusion detection systems (IDSs), however, often rely on computationally expensive deep learning models that are difficult to deploy at the edge. Periodic retraining also exposes these systems to machine unlearning attacks, where selective data removal can degrade detection performance. We propose a...

    arxiv.org/abs/2609.09564 · PDF

  58. 58

    A Statistical Approach to Estimating Sample Size of Machine Learning Models

    Dat Phan-Trong, Sunil Gupta, Svetha Venkatesh

    cs.LG · cs.AI

    Sample size determination for machine learning (ML) prediction models is challenging because conventional power analysis typically requires the predictor-outcome relationship and effect structure to be specified a priori. Nonlinear ML models learn complex prediction surfaces that do not admit straightforward analytical power calculations. We propose a framework that approximates nonlinear ML models with localized linear representations and...

    arxiv.org/abs/2609.09547 · PDF

  59. 59

    Efficient Fairness Auditing Across Guidance Scales in Text-to-Image Diffusion Models via Causal Abstraction

    Nabila Tasfiha Rahman, Rajatsubhra Chakraborty, Depeng Xu, Lu Zhang

    cs.LG · cs.CV

    Fairness auditing of text-to-image diffusion models often requires generating large numbers of images across sampling configurations, making comprehensive evaluation computationally expensive. We propose a causal-abstraction-based audit instrument for efficiently evaluating fairness under interventions on the classifier-free guidance scale. Given a fixed prompt and a target feature function, we represent the diffusion process as a low-level...

    arxiv.org/abs/2609.09486 · PDF

  60. 60

    Unthrottling the Tanh Jacobian in SAC: A Negative Result on Bang-Bang Control and MetaDrive

    Faiq Shamass

    cs.LG · cs.RO

    Soft Actor-Critic (SAC) represents a continuous policy as an unbounded Gaussian that is squashed by tanh. The Jacobian of that map is $\partial a/\partial u = 1-a^2$, which vanishes as $|a|\to 1$. A natural concern is that this throttle starves the actor of critic signal exactly where extreme actions (full brake, full throttle) are optimal. We test a minimal intervention that restores the missing signal: one extra term in the actor loss whose...

    arxiv.org/abs/2609.09478 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.