cs.DC · 2026-09-29 · No. 128
Distributed, Parallel, and Cluster Computing, 2026-09-29.
5 new papers in cs.DC. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
5 entries-
01
Dynamic Wakeup under Costly Collisions
Umesh Biswas, Maxwell Young
cs.DC
The wakeup problem captures a fundamental symmetry-breaking challenge among devices sharing a communication channel. We study the dynamic setting, where packets become active at arbitrary times on a time-slotted multiple access channel. In each slot, a transmission succeeds if and only if exactly one packet transmits; two or more simultaneous transmissions cause a collision. The goal is to obtain a successful transmission quickly. Prior work...
-
02
GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation
Yuchen Sun, Jinjin He, Sinan Wang, Bo Zhu
cs.DC · cs.AI
Writing fast GPU code for physical simulation is difficult: implementations must preserve numerical accuracy while handling irregular data access, synchronization, and iterative solvers. We introduce GPUPhysBench, a benchmark of 50 tasks testing whether coding agents can meet these demands. Tasks cover fluids, deformable solids, and granular materials, from individual simulation operators to complete simulators. Agents write, compile, test,...
-
03
TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training
Jiacheng Zhu, Xie Zhao, Gongming Zhao, Hongli Xu, Yao Fei, Jin Fang
cs.DC · cs.LG
Dynamic routing creates severe load imbalance in large-scale expert-parallel Mixture-of-Experts (MoE) training, turning GPUs that host hot experts into stragglers. As each MoE layer waits for its slowest rank, these stragglers prolong the expert-parallel stage and reduce overall training efficiency. Existing expert-parallelism load-balancing (EPLB) systems commonly compute load-balancing plans on the CPU, incurring device--host data transfers...
-
04
Weaver: A System for AI-RAN Compute Sharing with Foundation Model Training
Leyang Xue, Tianxin Wang, Xin Zhe Khooi, Jiaxun Yang, Dheeraj Mahendiran, Yufeng Xia, Mun Choon Chan, Myungjin Lee,...
cs.DC · cs.NI
The emergence of AI-RAN infrastructure, which equips cell sites with GPU-accelerated hardware, creates an opportunity to colocate non-RAN workloads with primary RAN processing. We explore using this spare capacity for decentralized training of foundation models (FMs), one of the most compute-intensive AI workloads. We present the first characterization of spare GPU capacity in AI-RAN systems at both micro-scale--across transmission slots...
-
05
WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse
Aaryam Sharma
cs.DC · cs.AI
Pipeline parallelism can improve prefill throughput by processing multiple request chunks concurrently across different stages of the model. However, keeping the pipeline fully utilized requires efficient scheduling and request preparation. In systems where stages retain and evict cache state independently, a local cache hit does not guarantee that the same prefix can be reused across the pipeline. Here, coordination overhead can impede...
This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.