cs.DC · 2026-10-06 · No. 135
Distributed, Parallel, and Cluster Computing, 2026-10-06.
8 new papers in cs.DC. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
8 entries-
01
One Global Beam Across Many GPUs: High-Throughput Beam Search at Billion-Record Frontier Scale
Ivan Litvak
cs.DC
Beam search repeatedly makes many children, removes duplicates, and keeps the best $B$. We show how many GPUs can perform these steps as one search even when the retained set does not fit on one device. States and candidates stay on the GPUs; the CPU receives only small control and ancestry records. We prove that an abstract distributed pipeline returns the monolithic reduced-key top-$B$ for the same complete candidates, integer scores,...
-
02
OrigaMIG: MIG-Aware VM Placement with a Neighborhood-Restricted BILP and Live Migration
Ahmad Siavashi, Mahmoud Momtazpour
cs.DC
The extensive use of GPUs in cloud computing, accelerated by the spread of large language model (LLM) services, and the growing need for multitenancy have driven the development of innovative solutions for efficient GPU resource management. Multi-Instance GPU (MIG) technology from NVIDIA enables shared GPU usage in cloud data centers by providing isolated instances, which are offered as MIG-backed virtual GPUs (vGPUs). However, MIG placement...
-
03
GPU-Initiated Discrete Simulated Bifurcation: Low-Latency Requests and Streaming Dense Couplings
Yaocheng Chen
cs.DC · physics.comp-ph · quant-ph
GPU-based optimization faces two communication bottlenecks: coordinating frequent requests and delivering dense models that exceed device memory. We present a discrete simulated bifurcation (dSB) architecture that addresses both through NVIDIA DOCA GPUNetIO. For resident models, a persistent service receives field updates, executes each solve within one GPU thread block, and returns the result. Exact integer coupling sums, GPU work queues,...
-
04
DIALER: A Case for Improving Rare-Class Accuracy in Retraining-Free Edge Video Analytics
Dongyoon Ryu, Sungho Jeon, Xinyue Ma, Di Wang, Jonghyun Choi, Minjia Zhang, Myeongjae Jeon
cs.DC
Edge video analytics with lightweight models is prone to accuracy degradation due to persistent distributional shifts in live video streams. While continuous learning (CL) addresses such data drift, it heavily strains the limited compute resources of edge servers originally provisioned for inference. Our empirical study reveals that emerging vision foundation models (VFMs) offer a practical, retraining-free alternative that delivers high...
-
05
From a Last-Week Add-On to First-Day Practice: Three Years of Integrating MPI Performance Analysis into HPC Education with EduMPI Suite
Anna-Lena Roth, Jonas Posner
cs.DC
Performance analysis is essential for scalable MPI programs, yet tool and workflow complexity often relegates it to the end of Parallel Programming courses. EduMPI Suite lowers this entry barrier by automating cluster execution, measurement, and near-real-time analysis while visualizing MPI communication from the first program. This paper presents a three-year experience report on its integration into HPC education. We trace the transition of...
-
06
Fine-Tuning a 3B-Parameter LLM on a Smartphone: Characterizing Sustained Training
Andrew Geyko, Marius Mosbach, André Brinkmann
cs.DC · cs.AI · cs.PF
Multi-billion-parameter LLMs now run on phones for inference, and training them on the device would personalize them without user data leaving the phone. Prior work has measured individual training steps of such models on phones, but not complete training runs, and not whether adapters trained on the device improve personalization. We present the first systematic characterization of a multi-billion-parameter LLM fine-tuned on a mobile device,...
-
07
Acceleration of Data Analytics on Heterogeneous Supercloud Systems
Georgios Zacharopoulos, Ilias Bournias, Lukas Cavigelli
cs.DC · cs.MS · cs.PF
Heterogeneous Supercloud systems are transforming data analytics by enabling scalable and efficient task distribution across diverse resources. This paper presents computational models and performance estimation techniques tailored for accelerating Dense Cholesky (CH) Decomposition and Singular Value Decomposition (SVD)-based analytics on Heterogeneous Superclouds. Our ReDSEa tool-chain automates mapping, load balancing, scheduling,...
-
08
Serve Now or Improve Later? Scheduling Self-Evolution in Online Agent Systems
Yangbo Wei, Junhong Qian, Zhen Huang, Zhenyu Su, Qifan Wang, Shaoqiang Lu, Rumin Zhang, Chen Wu, Lei He
cs.DC
Online agents can improve future service by constructing reusable tools, guidance, or model states, but this work competes with current requests for the same GPUs. Exploiting idle compute for self-evolution faces a fundamental systems constraint: benefits arrive only after an artifact is published and used, while pausing evolution leaves service capacity waiting for memory release and runtime recovery. An investment worth completing may...
This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.