cs.DC · 2026-06-18 · No. 27

Distributed, Parallel, and Cluster Computing, 2026-06-18.

14 new papers in cs.DC. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

14 entries
  1. 01

    TurboServe: Serving Streaming Video Generation Efficiently and Economically

    Youhe Jiang, Haoxu Wang, Haotong Bao, Kai Jiang, Jianfei Chen, Jun Zhu, Fangcheng Fu, Jintao Zhang

    cs.DC

    Streaming video generation is emerging as a new serving workload in which users interact with long-lived sessions that generate video progressively, chunk by chunk. Unlike offline video generation or typical LLM serving, streaming video generation must preserve session state across active and idle periods, repeatedly schedule ongoing sessions, and deliver each chunk under a tight latency target. This creates two key serving challenges in...

    arxiv.org/abs/2606.19271 · PDF

  2. 02

    Pulse: Training Acceleration for Large Diffusion Models with Automatic Pipeline Parallelism

    Boran Sun, Guoyong Jiang, Lin Zhang, Chen Chen, Yuechen Tao, Zhishu Che, Jieling Yu, Shan Chang, Huaxi Gu, Fangming...

    cs.DC

    Diffusion models are now a dominant approach for high-fidelity image and video generation, yet scaling their training across GPU clusters remains challenging. Unlike transformer-only architectures, diffusion backbones commonly adopt UNet-style encoder-decoder structures with heterogeneous layers and long-range skip connections. Under conventional pipeline parallelism, these non-local dependencies force large skip activations and their...

    arxiv.org/abs/2606.19163 · PDF

  3. 03

    Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training

    Ruiqi Lai, Dakai An, Wei Gao, Ju Huang, Siran Yang, Jiamang Wang, Lin Qu, Dmitrii Ustiugov, Wei Wang

    cs.DC · cs.AI · cs.LG

    Reinforcement learning (RL) post-training of Diffusion Transformers (DiTs) is prohibitively expensive, requiring thousands of high-end GPUs. Existing works explore two directions to reduce cost: seed exploration improves training convergence by selecting high-contrast samples, yet adds compute to the critical path; spot GPUs offer 69--77\% lower cost, yet sit idle during training because DiT rollouts finish nearly simultaneously, which...

    arxiv.org/abs/2606.19004 · PDF

  4. 04

    On the Notions of Bounded Bypass, and How to Make any Deadlock-Free MUTEX Protocol Satisfy One of Them

    Rob van Glabbeek, Daniele Gorla, Myrthe Spronck

    cs.DC

    In the literature on mutual exclusion, bounded bypass has been used for a long time as a strengthening of starvation-freedom, but, to the best of our knowledge, it still lacks a satisfying definition as a liveness property on its own. Moreover, we have encountered MUTEX protocols for which this notion needs to be slightly weakened in order to be met. To solve these issues, we first provide a formal definition of bounded bypass (that also...

    arxiv.org/abs/2606.19003 · PDF

  5. 05

    A Composable CRDT Layer for Byzantine-Resilient Deterministic Reconstruction

    Amos Brocco

    cs.DC · cs.CR

    Conflict-free Replicated Data Types (CRDTs) ensure Strong Eventual Consistency without coordination, but typically assume benign participants and rely on validation or exclusion to handle Byzantine behavior. We address this problem through deterministic state reconstruction: rather than deciding which updates are admissible, all accepted updates are incorporated, while only a subset contributes to the reconstructed state. We instantiate this...

    arxiv.org/abs/2606.18966 · PDF

  6. 06

    LiveStack: OS Support for Cluster-Scale Full-Stack Live Simulation

    Yiliang Wan, Haifeng Sun, Yihan Yang, Jonas Kaufmann, Antoine Kaufmann, Jialin Li

    cs.DC · cs.OS

    Cluster-scale full-stack simulation is essential for evaluating distributed software stacks and emerging hardware components before deployment. Such simulation must achieve both full-stack fidelity for the unmodified production stack and the simulation performance required for iterative configuration exploration. However, no existing method achieves both. We present LiveStack, an OS-level approach to cluster-scale full-stack simulation built...

    arxiv.org/abs/2606.18958 · PDF

  7. 07

    Urban Limits as Design Constraints: Identifying Suitable Locations for Distributed, Photovoltaic-Powered Servers

    Justin Chikhaoui, Thomas Leduc, Daniel Siret, Abdoulaye Gamatie

    cs.DC

    Urban territories face growing tensions between increasing digital demand, limited resources, and socially constrained built environments. Although distributed computing paradigms such as edge and fog computing are widely presented as solutions for reducing latency and energy dissipation, the scientific literature largely overlooks where such infrastructures can be physically and socially deployed within cities, and typically neglects urban...

    arxiv.org/abs/2606.18940 · PDF

  8. 08

    Compressed-Resident Genomics: Full-Pipeline Device-Resident GPU LZ77 Decode with Position-Invariant Random Access

    Yakiv Shavidze

    cs.DC

    Genomic archives grow faster than decompression keeps up: the European Nucleotide Archive holds tens of petabytes of fastq.gz, and gzip is fundamentally sequential. GPU decompressors (nvCOMP DEFLATE at ~50GB/s on A100) decode whole files with no random access; CPU genomic tools (CRAM, samtools) support region seeks but only at CPU speed. We extend ACEAPEX, an absolute-offset parallel LZ77 codec included in the official lzbench 2.3 release,...

    arxiv.org/abs/2606.18900 · PDF

  9. 09

    ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving

    Haipeng Yuan, Kaining Zheng, Yongshu Bai, Yuchen Zhang, Yunquan Zhang, Baodong Wu, Xiang Gao, Daning Cheng

    cs.DC

    Current large language model (LLM) inference systems universally deploy ultra-large-scale models using a combination of Tensor Parallelism (TP) and Pipeline Parallelism (PP). However, existing systems treat the model parallelism topology as a static configuration that cannot be flexibly adjusted at runtime. This rigid design creates a fundamental contradiction with the dynamically changing inference workloads in real-world scenarios....

    arxiv.org/abs/2606.18741 · PDF

  10. 10

    Closed-Form and Constant-Time New-Source Selection for Fault-Tolerant Broadcasting in Dense Gaussian Networks

    Bader Albader

    cs.DC · cs.IT · cs.NI

    Fault-tolerant broadcasting in dense Gaussian networks is recovered by re-rooting the broadcast at a new source at maximum graph distance from the faulty nodes. This paper extends the re-rooting framework by replacing its boundary-search source-selection step with a quotient-lattice-aware algebraic construction. The first contribution is a constant-time counting method for valid new sources, formulated as an intersection of two diameter-$k$...

    arxiv.org/abs/2606.18715 · PDF

  11. 11

    Closed-Form and Constant-Time New-Source Selection for Fault-Tolerant Broadcasting in Dense Eisenstein--Jacobi Networks

    Bader Albader

    cs.DC · cs.IT · cs.NI

    Fault-tolerant broadcasting in dense Eisenstein--Jacobi networks requires efficient recovery when faulty nodes disrupt the original broadcast structure. A re-rooting-based method guarantees that, for any two faulty nodes, a valid new source exists at maximum graph distance from both faults. However, identifying such a source without scanning the network or testing all boundary candidates remains an open practical problem. This paper presents...

    arxiv.org/abs/2606.18714 · PDF

  12. 12

    Re-Rooting-Based Fault-Tolerant One-to-All Broadcasting in Dense Eisenstein--Jacobi Networks

    Bader Albader

    cs.DC · cs.IT · cs.NI

    Dense Eisenstein--Jacobi networks are degree-six algebraic interconnection topologies with regular structure, vertex symmetry, small diameter, and efficient communication algorithms. These properties make them suitable for parallel and on-chip communication systems in which collective operations such as one-to-all broadcasting are frequent. Existing optimal broadcasting algorithms for dense hexagonal/Eisenstein--Jacobi networks assume...

    arxiv.org/abs/2606.18712 · PDF

  13. 13

    HI-HCQC: A Tightly-Coupled Hardware Interface with High-Efficiency Communication for Hybrid Classical-Quantum Computing

    Shibo Liang, Junchao Wang, Zeyuan Wang, Feng Wang, Xiaoyu Li, Lei Li, FuDong Liu, Zheng Shan

    cs.DC

    Hybrid classical-quantum computing requires frequent data exchange between classical processors and quantum control hardware. However, existing superconducting quantum control systems are commonly connected through loosely coupled interfaces such as Ethernet, resulting in high communication latency and limited task throughput. To address this issue, we present HI-HCQC, an RFSoC-based hardware interface for tightly coupled hybrid...

    arxiv.org/abs/2606.18642 · PDF

  14. 14

    ShuntServe: Cost-Efficient LLM Serving on Heterogeneous Spot GPU Clusters

    Seungwoo Jeong, Moohyun Song, Juhyun Park, Kyungyong Lee

    cs.DC

    As large language model (LLM) services become widely adopted, the cost of GPU resources for serving these models in cloud environments has emerged as a critical concern. Spot instances offer up to 90% cost savings over on-demand instances, but their frequent interruptions and limited availability pose significant challenges for continuous LLM serving. GPU spot instances, in particular, exhibit lower and more volatile availability than...

    arxiv.org/abs/2606.18600 · PDF

This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.