cs.LG · 2026-09-11 · No. 112

Machine Learning, 2026-09-11.

45 new papers in cs.LG. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

45 entries
  1. 01

    General Quantification of Covariate and Concept Shifts

    Hongbo Chen, Li Charlie Xia

    cs.LG · cs.AI · stat.ML

    Generalization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-estimable from samples. In this paper, we bridge the gap between theory and practical applications. We first show that existing definition of concept shift breaks when the source and target supports mismatch. Leveraging entropic optimal transport, we propose a key...

    arxiv.org/abs/2609.11918 · PDF

  2. 02

    Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data

    Atindra Jha, Margaret Li, Jure Leskovec, Percy Liang, Luke Zettlemoyer

    cs.LG · cs.CL

    As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data...

    arxiv.org/abs/2609.11917 · PDF

  3. 03

    From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good

    Nitesh V. Chawla, Paulo Benanti

    cs.LG

    Artificial Intelligence does more than create a governance problem. It can also reveal where institutions have already failed to provide responsiveness, belonging, care, and accountability. Once deployed, AI becomes an intervention in those conditions. It can repair, compound, substitute for, or conceal the failures it encounters. Responsible AI must therefore evaluate both the system and the institutional rupture into which it is introduced....

    arxiv.org/abs/2609.11910 · PDF

  4. 04

    TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription

    Akshaj Gupta, Hwi Joo Park, Andrea Guzman, Shamak Gowda, Samhita Konduri, Jiachen Lian, Robin Netzorg, Gopala Anumanchipalli

    cs.LG

    Automatic Music Transcription (AMT) for guitar remains limited by three challenges: existing systems often fail to capture expressive techniques such as slides, bends, and percussive hits; they often assign notes to incorrect string-fret combinations; and they are typically trained on clean recordings, limiting their generalization to noisy real-world audio. To address these challenges, we propose TART, a modular four-stage audio-to-tablature...

    arxiv.org/abs/2609.11904 · PDF

  5. 05

    CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

    Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye

    cs.LG

    Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a causal graph together with the mechanisms that generate data, yet existing studies differ substantially in graph families, mechanisms, and evaluation protocols. The emergence of causal discovery foundation models (CDFMs)...

    arxiv.org/abs/2609.11897 · PDF

  6. 06

    CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search

    Yifan Yang, Zhaoyan Wang, Zheng Gao, Xiaoyu Li, Jiaojiao Jiang

    cs.LG · cs.CV

    Zero-cost proxies rank architectures cheaply, but their reliability varies across search spaces. We introduce CoRA-NAS (COarse Ranking + Anchor-residual), a two-stage framework combining a static ranking prior with low-cost learning-curve refinement. CoRA-Rank aggregates capacity and structure-at-initialization proxies through an equal-weight log-rank consensus and a target-free consensus gate. CoRA-Refine samples anchors across this prior,...

    arxiv.org/abs/2609.11884 · PDF

  7. 07

    The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

    Yi Duan, Ying Liu, Zirui Tang, Haodong Chen, Jun Zhou, Yumou Liu, Bangrui Xu, Yukai Wu, Sidi Chen, Yuhan Zhou, Haoyu...

    cs.LG · cs.AI · cs.CL

    Recursive self-improvement (RSI) enables AI systems to turn experience and feedback into persistent changes that improve both their capabilities and the process of future improvement. We first use the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, then introduce the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and...

    arxiv.org/abs/2609.11873 · PDF

  8. 08

    AdamX: Cosine similarity meets gradient descent

    Francisco Caldas, Ruben Belo, Cláudia Soares

    cs.LG · math.OC

    We introduce AdamX, a first-order optimizer that incorporates cosine similarity as an adaptive mechanism for controlling update magnitudes. The proposed method is scalable, model-agnostic, and straightforward to integrate into existing training pipelines. We further introduce a variance rectification scheme that promotes smoother optimization during the early stages of training. Overall, we provide empirical evidence that AdamX achieves...

    arxiv.org/abs/2609.11867 · PDF

  9. 09

    Model-Aware Schedules Improve Generation via Fiberwise Optimal Transport

    Luyi Jia, Boyan Zhang, Yilun Liu, Steffen Rulands

    cs.LG · cs.AI

    Diffusion and flow-matching schedules control the signal and noise coefficients that mix data and noise along affine probability paths. Minimizing a kinetic action defined on coefficient paths, motivated by optimal transport, helps explain strong baselines but remains model-agnostic and ignores prediction error. Here we introduce a model-aware schedule construction based on fiberwise optimal transport. At a fixed time and state on the...

    arxiv.org/abs/2609.11842 · PDF

  10. 10

    Thinking with Looped Flows

    Ayhan Suleymanzade, Chanhyuk Lee, Floor Eijkelboom, Nicholas M. Boffi, İsmail İlkan Ceylan, Jinwoo Kim

    cs.LG · cs.AI

    Humans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a hidden state. In practice, however, their training backpropagates through only one or a few updates, making it hard to train early updates to support future ones. We propose looped flows, an approach that sidesteps this issue by training the recurrence with local...

    arxiv.org/abs/2609.11801 · PDF

  11. 11

    Dynamic language model representations for multi-objective reaction optimisation

    Joshua W. Sin, David Ming Segura, Bojana Ranković, Siu Lun Chau, Marius D. R. Lutz, Andrea Anelli, Ryan P. Burwood,...

    cs.LG

    Optimising chemical reactions across multiple objectives, such as yield, selectivity, and safety, is central to chemical synthesis, and model-driven approaches depend critically on how reaction components are represented. Established featurisations are either chemically uninformative, as with one-hot encodings, or, as with molecular descriptors, do not readily extend across chemically distinct components. For structurally and functionally...

    arxiv.org/abs/2609.11790 · PDF

  12. 12

    Predicting Privacy Leakage from Weight Spectral Density

    Richard J. Preen, Jim Smith

    cs.LG · cs.CR · cs.NE

    Membership inference attacks (MIAs) are widely used to audit the privacy disclosure risk of machine learning models, however current state-of-the-art attacks require training computationally expensive shadow models, making large-scale privacy evaluation impractical. In this work, we investigate whether inexpensive spectral metrics derived from the heavy-tailed self-regularisation framework can serve as proxies for MIA vulnerability. We...

    arxiv.org/abs/2609.11780 · PDF

  13. 13

    Why Does Post-Training Quantization Work?

    Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen

    cs.LG · cs.CL

    Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and corrupt next-token prediction; randomly initialized models accumulate these discrepancies rapidly, whereas quantized pretrained models accumulate much less hidden-state error and largely maintain downstream...

    arxiv.org/abs/2609.11716 · PDF

  14. 14

    Learnware and AI Model Management System

    Zhi-Hua Zhou

    cs.LG

    The transition from file storage to database management systems transformed stored data into managed resources. AI now faces an analogous transition from AI model storage to AI model management. Existing model pools essentially serve as \textit{AI model storage systems}. What is needed instead are \textit{AI model management systems} that enable models trained by different developers, for different tasks, with different data, and under...

    arxiv.org/abs/2609.11656 · PDF

  15. 15

    Musec: MomentUm SpEctral Clipping for Stable Muon-type Training

    Zhuanghua Liu, Menglian Wang, Luo Luo

    cs.LG

    Muon has emerged as a highly effective optimizer for large language model training, often achieving superior convergence and performance compared with the widely adopted Adam and AdamW optimizers. Nevertheless, Muon is prone to training instability due to its spectral flattening, manifested by loss spikes and unbounded growth of model weights. Existing approaches primarily rely on weight or attention-logit clipping, which require...

    arxiv.org/abs/2609.11655 · PDF

  16. 16

    RDDMPI: Residual Denoising Diffusion Model for Probabilistic Multivariate Time Series Imputation

    Ramiro Valdes Jara, David Chapman, Adam Meyers

    cs.LG · stat.ML

    Multivariate time series imputation (MTSI) aims to recover missing values in temporal data composed of multiple interdependent variables. This problem is central to real-world applications such as healthcare monitoring, traffic networks, and energy systems. Recent diffusion-based approaches have shown strong potential for probabilistic imputation by learning to generate missing values through iterative denoising. However, most existing...

    arxiv.org/abs/2609.11648 · PDF

  17. 17

    LoaDiff: Conditional Generation of Electricity Consumption Time Series for Energy Analytics

    Mariia Baranova, Adrien Petralia, Etienne Le Naour, Nathan Etourneau, Guillaume Hofmann, Themis Palpanas

    cs.LG · cs.AI · eess.SP

    The energy transition is reshaping residential electricity consumption through the increasing adoption of distributed generation, electrified appliances, and demand-response programs. Understanding these evolving behaviors requires access to granular smart-meter data for applications such as load forecasting, appliance detection, and demand-side flexibility analysis. However, such data are subject to strict access restrictions and...

    arxiv.org/abs/2609.11639 · PDF

  18. 18

    A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph

    Ruben Cartuyvels, Karim Douch, Gabriele Bertoli, Mounia El Baz, Artemis Vrettou, Sébastien Lefèvre, Diego Fernandez Prieto

    cs.LG

    Continuous monitoring of water surface elevation across river networks is critical for flood forecasting, water resource management, and understanding the global water cycle. Yet, the scarcity of in situ gauges across much of the globe constrains the development of reliable modeling frameworks. Satellite altimetry has the potential to alleviate this problem but its use is currently hindered by sparse temporal coverage. To this end, we...

    arxiv.org/abs/2609.11580 · PDF

  19. 19

    Particle GFlowNets: Rethinking Generative Marginalization Models

    Tiago da Silva, Diego Mesquita, Salem Lahlou

    cs.LG

    Generative Marginalization Models (MaMs) have been recently introduced as efficient neural sampling models for any-order autoregressive modelling of discrete distributions. By learning both the marginal and conditional probabilities of a persistent-block Gibbs sampler, MaMs enable fast posterior evaluation with a single neural network forward pass. While prior work has considered MaMs to be distinct from Generative Flow Networks (GFlowNets),...

    arxiv.org/abs/2609.11538 · PDF

  20. 20

    Generalized Score Matching for Parameter Estimation on Convex Domains

    Nishanth Shetty, Saisuchith Mahajan, Chandra Sekhar Seelamantula

    cs.LG · stat.ML

    Maximum likelihood (ML) estimation is a principled and statistically efficient approach for learning probabilistic models. However, for unnormalized models, ML estimation requires evaluating the partition function and differentiating through it, which may not always be tractable. Score matching provides a practically viable alternative that circumvents this obstacle by fitting the score in a way that eliminates dependence on the normalizing...

    arxiv.org/abs/2609.11521 · PDF

  21. 21

    DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis

    Abhinav Rajeev Kumar, Harshit Arora, Varun Singh, Manikandan Nanjappan

    cs.LG · cs.SE

    A structurally valid DeFi workflow can still authorize a costly trade. We introduce DeFiFlowBench, a benchmark of 207 team-authored prompts for natural-language DeFi workflow synthesis. It measures graph coverage, configuration completeness, and declared safety predicates, then tests supported trade configurations on a local EVM. Direct, constrained, and few-shot prompting produce 14-19 unsafe held-out executions per configuration under a...

    arxiv.org/abs/2609.11504 · PDF

  22. 22

    Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study

    Yan Hon Michael Chung, Hanlin Wang

    cs.LG

    Manchu, now critically endangered, was one of the principal languages of the Qing empire (1636-1912), and its extensive archival record is increasingly digitized but remains difficult to search and analyze at scale. Previous work showed that vision-language models (VLMs) trained only on synthetic Manchu word images can reach 87.4% word accuracy on real Qing manuscripts and prints, leaving a substantial synthetic-to-real gap. This study...

    arxiv.org/abs/2609.11495 · PDF

  23. 23

    Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets

    Jia Huang, Yankai Wan, Yangjun Ou

    cs.LG · cs.AI

    Many ML datasets are constructed by running a detector, heuristic, or model over candidate pools; accepted items become labels. Dataset precision is then governed by true-positive prevalence in each pool via Bayes, not solely by detector quality. Using one instrument and period, we hold a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom. One detector, three pools yield phantom...

    arxiv.org/abs/2609.11449 · PDF

  24. 24

    Local Robustness Quantification for Naive Bayes Classifiers and Generative Forests: a General Approach

    Adrián Detavernier, Jasper De Bock

    cs.LG

    We provide methods for calculating the robustness of the predictions of two types of generative classifiers whose underlying distribution is a Probabilistic Graphical Model (PGM): naive Bayes classifiers and generative forests (a probabilistic extension of random forests). Following the paradigm of robustness quantification, we define the robustness of a prediction as the extent to which the distribution of the classifier can be perturbed...

    arxiv.org/abs/2609.11366 · PDF

  25. 25

    Reification as a Transferable Vocabulary: Zero-Shot Link Prediction with Vanilla GNNs

    Camille Pradel

    cs.LG · cs.AI

    Knowledge graph foundation models such as ULTRA achieve zero-shot link prediction on unseen graphs through dedicated architectures that hard-code a transfer mechanism. In this work we move that mechanism out of the architecture and into the representation, by \emph{reifying} the input graph: every fact becomes a node, connected to its subject, object, and relation type through a fixed vocabulary of six meta-relations, with relation types as...

    arxiv.org/abs/2609.11347 · PDF

  26. 26

    Estimating Inconsistency Response Surfaces under Uncertainty in Cyber-Physical System Development

    Johannes Mäkelburg, Tim Schwabe, Maribel Acosta

    cs.LG · cs.SE · eess.SY

    Cyber-Physical Systems (CPS) are commonly represented through multiple interconnected models. During development, CPS consistency requires that shared model elements remain compatible across these models. Uncertainty, for example, due to sensor noise or model abstraction, changes the admissible values of model elements and can introduce inconsistencies, i.e., situations in which models can no longer be jointly satisfied. While existing...

    arxiv.org/abs/2609.11331 · PDF

  27. 27

    A Dynamic Fusion Large Language Model for Traffic Flow Prediction

    Xue Qiu, Jianli Xiao

    cs.LG

    Traffic flow prediction is a core supporting technology for intelligent transportation systems. It uses historical data to infer future traffic dynamics in specific areas, thereby helping to alleviate congestion and improve resource allocation efficiency. Traditional neural networks struggle to break through accuracy limits due to their reliance on singular feature modeling, while large language models (LLMs) suffer from insufficient capture...

    arxiv.org/abs/2609.11314 · PDF

  28. 28

    MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions

    Antoine Saillenfest

    cs.LG · cs.CL

    Erasing concept-specific information from representations has been proven useful for mitigating bias or interpreting model decisions. The joint objective is to transform the original representations such that the target concept becomes unpredictable, while maximally preserving concept-unrelated information. In this work, we revisit the optimal bounds of concept erasure to derive a novel class of erasure functions that naturally induce a...

    arxiv.org/abs/2609.11253 · PDF

  29. 29

    Solving Few-Shot Multiobjective Multitask Optimization via Iterative Sequential Transfer

    Tingyang Wei, Haofeng Wu, Ananda Phan Iman, Zhao Wei, Jiao Liu, Yew-Soon Ong

    cs.LG · cs.AI · cs.NE

    Applying knowledge transfer across multiple optimization tasks, multitask optimization (MTO) emerges as a promising approach to solving synergistic optimization tasks simultaneously. However, the development of effective knowledge transfer mechanisms in MTO fundamentally relies on aligning elite solution distributions across tasks. This dependency creates a critical bottleneck in few-shot optimization regimes, as restricted evaluation budgets...

    arxiv.org/abs/2609.11228 · PDF

  30. 30

    Polyhedral Geometry of Time-to-First-Spike Neural Networks

    Manjot Singh, Guido Montúfar, Gitta Kutyniok

    cs.LG · math.CO

    We study the expressivity of spiking neural networks, which provide a natural framework for asynchronous, event-driven computation complementary to conventional feedforward neural networks. We consider the time-to-first-spike model in a setting for which the input-output map is continuous and piecewise linear, with affine pieces governed by causal feasibility constraints that determine which presynaptic spikes occur before a neuron fires. We...

    arxiv.org/abs/2609.11227 · PDF

  31. 31

    Legible Failures: Detecting and Repairing In-Context Binding Errors

    Manas Venkata Sai Ravulapalli, Samrath Singh Chadha, Abhinav M. Hari

    cs.LG

    A wrong answer does not show whether the model lacked the needed information or held it and failed to use it. On an entity-obligation binding task, a language model can emit an incorrect prompt-supplied binding while a linear probe can recover the correct one from its frozen hidden state. We measure how often this occurs across 16 public checkpoints, each evaluated with three seeds. We fit a probe on a training fold, select its layer on a...

    arxiv.org/abs/2609.11216 · PDF

  32. 32

    REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

    Tuan Nguyen, Qiran Hu, Banruo Liu, Khoa D. Doan, Kok-Seng Wong, Fan Lai

    cs.LG · cs.CL · cs.IR

    Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and token cost. Post-retrieval compression can reduce this cost, yet existing compressors often operate independently for each query, rely on auxiliary models or rewriting, and introduce online overhead that can offset the...

    arxiv.org/abs/2609.11209 · PDF

  33. 33

    Convex Optimization with Nested Evolving Feasible Sets (CONES) under Time-Varying Loss Functions

    Rahul Vaze

    cs.LG · cs.DS · math.OC

    Convex Optimization with Nested Evolving Feasible Sets (CONES)} was introduced in \cite{CONESVaze} where the objective function \(f\) remains fixed but the feasible region evolves over time as a nested sequence \(S_1 \supseteq S_2 \supseteq \cdots \supseteq S_T\). The goal of an online algorithm is to simultaneously minimize the regret with respect to hindsight static optimal benchmark and the total movement cost $M_\cA(T)$ while ensuring...

    arxiv.org/abs/2609.11207 · PDF

  34. 34

    Hierarchical Clustering Can Jointly Satisfy Richness, Consistency, and Scale Invariance

    Daichi Kuroda, Maximilien Dreveton, Matthias Grossglauser, Patrick Thiran

    cs.LG · stat.ME · stat.ML

    Despite its ubiquity, clustering lacks a universally accepted definition of what is a cluster. Kleinberg's Impossibility Theorem formalizes this difficulty by showing that no flat clustering method can simultaneously satisfy three natural axioms: scale invariance, richness, and consistency. In this paper, we ask whether this impossibility persists when the output is a hierarchy rather than a single partition. We show that, in contrast to the...

    arxiv.org/abs/2609.11173 · PDF

  35. 35

    Semi-Tensor Product-Based Multi-Term Randomized T-SVD and Its Visual Applications

    Xingchen Xiao, Feng Zhang, Wenjin Qin, Jianjun Wang

    cs.LG

    Tensor singular value decomposition (T-SVD), which is built upon the tensor-tensor product (t-product), has emerged as a powerful tool for processing high-dimensional visual data such as color images and videos. However, the standard t-product imposes strict dimensional compatibility constraints. Although extensions based on the semi-tensor product (STP) relax this restriction, their single-term formulations still suffer from limited...

    arxiv.org/abs/2609.11168 · PDF

  36. 36

    When does a spectral prior help graph learning? Connectivity-loss estimation under road-network disruptions

    Van-Truong Le

    cs.LG

    Rapid evaluation of many simultaneous road-link disruptions requires a practical compromise between exact spectral recomputation and local approximation. We estimate relative algebraic-connectivity loss after multi-edge deletion using graph neural networks (GNNs) that learn a bounded correction to a first-order Fiedler sensitivity. The study considers independent, spatially clustered, and edge-betweenness-targeted failures, with...

    arxiv.org/abs/2609.11166 · PDF

  37. 37

    LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

    Sankar Behera, Dhruv Singh, Anshika Agnihotri, Raj Kumar Choudhary, Satyadev Ahlawat, Yamuna Prasad

    cs.LG · cs.CL

    Structured pruning of large language models (LLMs) offers hardware-efficient compression, yet existing methods require calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (\emph{Latent-Informed Layer Analysis}) scores neuron importance via the Kolmogorov--Smirnov (KS) distance between empirical singular value distributions of the full and neuron-ablated feed-forward network (FFN) weight matrix,...

    arxiv.org/abs/2609.11163 · PDF

  38. 38

    Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models

    Ken Chen, Maneesha Perera, Wei Wang, Sachith Seneviratne, Hansani Weeratunge, Saman Halgamuge

    cs.LG

    Short-term photovoltaic (PV) power and global horizontal irradiance (GHI) forecasts are essential for effective dispatch, reserve scheduling, and grid operations. At these forecasting horizons, errors are predominantly driven by cloud induced ramps: relying solely on historical numerical data may struggle to anticipate an incoming cloud, making ground-based sky images a crucial complementary physical signal. Furthermore, forecast performance...

    arxiv.org/abs/2609.11135 · PDF

  39. 39

    Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving

    Jae Gon Kim, Donghoon Yoo, Hanyul Ryu, Sungho Ha, Juyeon Lee, Soojung Ryu

    cs.LG · cs.DC

    Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference profile on a disaggregated B200 system, we found its realized gain modest (+8.6% tokens/J), model-dependent, and carrying a mean end-to-end latency cost (+5.2%) that throughput-only evaluation does not surface; the profile also applies one setting to prefill and...

    arxiv.org/abs/2609.11133 · PDF

  40. 40

    How Wrong Can a Good Predictor Be? Diverging Updates with Vanishing Predictive KL

    Qifu Wen, Shuaijun Liu, Zihan Zhou, Xi Zeng, Ningxin Su

    cs.LG · stat.ML

    Accurate posterior prediction need not require accurate approximation of Bayesian updates. We prove that an unbounded gap between the update maps can coexist with vanishing predictive KL for every fixed finite $K\ge2$ in a stationary symmetric Gaussian HMM. Exact Bayesian mixing and an explicit deterministic radial filter act on the same $K-1$ belief coordinates. As $q\to0^+$, their separation in centered logits in the worst case grows at...

    arxiv.org/abs/2609.11132 · PDF

  41. 41

    HERALD: High-Fidelity Exemplar Retrieval with Adaptive Landmark Distillation for Heterophily-Aware Graph Condensation

    Sujan Chakraborty, Priyanka Saha, Saptarshi Bej

    cs.LG

    Graph condensation aims to produce a small surrogate graph that preserves the downstream node-classification performance of a much larger original graph. Existing methods rely on Weisfeiler-Lehman neighbourhood aggregation or gradient-based distribution matching, both of which assume that adjacent nodes share the same label, an assumption that breaks down under heterophily. We propose HERALD (High-fidelity Exemplar Retrieval with Adaptive...

    arxiv.org/abs/2609.11123 · PDF

  42. 42

    Beyond Solver Verdicts: Generative Reward Models for Autoformalization

    Vikash Singh, Debargha Ganguly, Aman Goel, Ali Torkamani, Xiaoxue Han, Joseph Lilien, Ferhat Erata, Vipin Chaudhary

    cs.LG · cs.CL

    Neurosymbolic systems rely on mathematical solvers to guarantee reasoning correctness, yet solvers are fundamentally blind to whether a formal translation maintains strict reference-equivalence to a designated formalization. We formalize this vulnerability as Verdict-Preserving-Unfaithfulness (VPU): a failure mode where an incorrect encoding executes successfully and matches the expected verdict. We theoretically prove that structural,...

    arxiv.org/abs/2609.11085 · PDF

  43. 43

    The information geometry of large language models is shared, learned, and controllable

    Dario Picozzi

    cs.LG · cs.CL

    Large language models learn similar behaviours, yet it remains unclear what structure they share or how to change one behaviour without disturbing others. The Fisher-Rao geometry of next-token probabilities connects these questions: behaviour determines this geometry up to output-preserving symmetries, whereas activation geometry depends on coordinates. Across transformer, state-space and recurrent models, output geometries agree more...

    arxiv.org/abs/2609.11063 · PDF

  44. 44

    EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression

    Motahare Mounesan, Irfan Khan

    cs.LG · cs.DC

    Recent advances in multimodal large language mod- els (MLLMs) have opened new opportunities for edge intelligence by enabling reasoning across heterogeneous sensor modalities, such as vision, text, and telemetry data. However, deploying these capabilities on resource-constrained edge platforms remains challenging due to the substantial computational, memory, and communication demands of modern MLLMs. Rather than transmitting raw sensor...

    arxiv.org/abs/2609.11058 · PDF

  45. 45

    T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

    Junyao Yang, Yucheng Shi, Zhongzhi Li, Ruhan Wang, Zongxia Li, Haitao Mi, Leowei Liang

    cs.LG · cs.AI

    Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier. We provide a comprehensive recipe: First, an aggressively warm-started to...

    arxiv.org/abs/2609.11042 · PDF

This edition is part of The Daily Abstract — cs.LG archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.