cs.DC · 2026-09-30 · No. 129
Distributed, Parallel, and Cluster Computing, 2026-09-30.
5 new papers in cs.DC. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
5 entries-
01
RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust
Eugene Hauptmann, Nataliya Kosmyna
cs.DC · cs.AI
Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallbacks hard to reason about end-to-end. RLX addresses this gap with a single Rust codebase that combines compiler and runtime roles around one primitive-level, three-level intermediate representation (IR), plus a transparent dispatch contract that...
-
02
Byzantine Causal Reliable Broadcast with Constant Metadata Overhead
Purv Patel, Ajay D. Kshemkalyani
cs.DC · cs.DS
Asynchronous Byzantine Reliable Broadcast (BRB) is a fundamental primitive that guarantees agreement and validity in distributed systems subject to Byzantine faults, but it lacks ordering guarantees. Causal message ordering is important for many applications such as blockchain and social networking. Existing solutions for Byzantine Causal Reliable Broadcast (BCRB) have several drawbacks. Such protocols typically append vector clocks or...
-
03
Joint Effects of GPU Server Topology, Parallelism, and Congestion Control on MoE Inference: A Controlled Simulation Study
Kaikai Yuan, Rui Xi, Yu Liu
cs.DC
Mixture-of-experts (MoE) models expand capacity via sparse activation, but inference across GPUs introduces tensor-parallel (TP) collectives and expert-parallel (EP) dispatch and combine operations. Completion time depends not just on communication volume but on how logical groups map onto intra-server interconnects, GPU--NIC connections, and the inter-node network. Using ASTRA-sim with the NS-3 discrete-event backend, we build a controlled...
-
04
FP64 Is All You Want, INT8 Is All You Need, FP4/6/8 Is All You Have
Pratyai Mazumder, Alexandru Calotoiu, Torsten Hoefler
cs.DC
Ozaki scheme II emulates FP64 matrix products with INT8 ones through residues modulo pairwise coprime moduli, and variants for FP8 and FP4 have followed. We treat these schemes as one family and pose the choice of a scheme as a combinatorial program that minimizes the number of low-precision GEMMs. Given, for each modulus, a finite set of ways to compute products modulo it from low-precision GEMMs, we find the choice of moduli and ways with...
-
05
SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving
Chuan Liu, Shuoming Zhang, Zhicheng Li, Qianqi Sun, Ruiyuan Xu, Qiuchu Yu, Xiyu Shi, Huimin Cui, Jiacheng Zhao
cs.DC · cs.AI
No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice untenable: a batch that begins as many short requests ends as a few very long ones, so the best layout changes while the same requests run....
This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.