cs.DC · 2026-06-16 · No. 25
Distributed, Parallel, and Cluster Computing, 2026-06-16.
10 new papers in cs.DC. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
10 entries-
01
Re-Rooting-Based Fault-Tolerant Broadcasting in Dense Gaussian Networks
Bader Albader, Mohamed R. Al-Mulla, Galal Hassan
cs.DC · cs.IT · cs.NI
Dense Gaussian networks provide degree-4 interconnection topologies with small diameter and regular structure, making them suitable for efficient one-to-all broadcasting. However, node failures can disrupt the broadcast process when faulty nodes occupy internal forwarding positions. This paper proposes a lightweight fault-tolerant broadcasting method based on dynamic source relocation, or re-rooting. Instead of constructing redundant spanning...
-
02
Tangram: Hiding GPU Heterogeneity for Efficient LLM Parallelization
Yanda Tao, Pedro F. Silvestre, Marcel Wagenländer, Peter Pietzuch
cs.DC
The scale of LLM training jobs requires parallelization planning over large GPU clusters. Due to different GPU types and interconnects added over time, these GPU clusters are increasingly heterogeneous. Automatic LLM parallelizers can search for parallelization plans but face an exploding search space with heterogeneous GPUs. To make search tractable in heterogeneous GPU clusters, parallelizers often omit types of parallelism (e.g., expert...
-
03
A Unified Constant-Time Switch Rule for Constructing Edge-Disjoint Hamiltonian Cycles in Gaussian Networks
Bader Albader
cs.DC · cs.DM · cs.IT · cs.NI
Gaussian networks are degree-four symmetric interconnection networks defined over residue classes of Gaussian integers. Earlier work showed that when the generator $α=a+bi$ satisfies $\gcd(a,b)=1$, the real and imaginary dimensions directly form two edge-disjoint Hamiltonian cycles. A later construction extended the result to the non-coprime case $\gcd(a,b)=d>1$, but its proof used long node-sequence tables and separate odd/even cases for...
-
04
CacheWise: Understanding Workloads and Optimizing KVCache Management for Efficiently Serving LLM Coding Agents
Shubham Tiwari, Tapan Chugh, Nash Rickert, Simon Peter, Ratul Mahajan, Haiying Shen
cs.DC · cs.OS
Coding agents are a fast-growing LLM application, executing as long-running closed-loop sessions in which LLM generations alternate with external tool calls. Yet, unlike chat workloads, their serving behavior has not been studied extensively. We address this gap by collecting a dataset of real-world coding assistant traces. Our analysis shows that coding agent sessions repeatedly reuse large prefixes and create sustained KVCache pressure that...
-
05
From the NYU Ultracomputer to Modern Exascale: A Historical and Architectural Survey of In-Network Computing and Scalable Synchronization
Lars Warren Ericson
cs.DC · cs.AR · cs.GL
This paper presents a historical and technical survey of the hardware architectures, interconnection networks, and synchronization primitives that have shaped massively parallel systems over the past four decades. We examine the design of the NYU Ultracomputer and the IBM Research Parallel Processor Prototype (RP3), focusing on the hardware implementation of the Fetch-and-Add primitive in multistage interconnection networks. We contrast these...
-
06
Robust and Automated Reconfiguration of Byzantine Wide-Area Replication
Rowdy Chotkan, Bulat Nasrulin, Johan Pouwelse, Jérémie Decouchant
cs.DC · cs.CR · cs.NI
Distributed systems handle adversarial nodes through redundancy, which imposes a significant performance overhead. In blockchain systems, Byzantine fault-tolerant state-machine replication (BFT-SMR) is the replicated service that totally orders client transactions before execution. While prior research has primarily focused on designing novel consensus algorithms with improved performance, recent studies have shown that further gains can be...
-
07
DRIFT: Risk-Constrained Diffusion with Imitation Priors for Mixed-Autonomy Traffic Generation
Yaoshen Yu, Minghui Liwang, Wenbo Zhu, Xinlei Yi, Yiguang Hong, Yuhan Su, Seyyedali Hosseinalipour
cs.DC
Future intelligent transportation systems are envisioned to evolve toward a long-term mixed-autonomy paradigm, where human-driven vehicles (HVs) and autonomous vehicles (AVs) coexist within highly coupled traffic ecosystems. Such coexistence introduces pronounced heterogeneity, amplified uncertainty, and increasingly intricate interaction dynamics. In this context, it remains fundamentally challenging to simultaneously capture the...
-
08
Incentives and Evidence in Learned Service Orchestration
Syed Izhan Khilji, Alireza Furutanpey, Schahram Dustdar
cs.DC · cs.LG
Reinforcement learning for service orchestration has been the subject of sustained research for over a decade, yet it is not used in production at scale. The usual explanation is that learned controllers degrade under delayed and noisy telemetry, workload shifts, and uncontrolled tenants. We test whether existing evidence supports that explanation. We evaluate three highly influential RL-based orchestration systems spanning resource...
-
09
Generated, Parallel, Scalable? A Study of Agentic AI-Generated Julia Code on Supercomputers
Linus Bantel, Anna-Lena Roth, Jonas Posner, Dirk Pflüger
cs.DC
Julia is increasingly used in hpc as a single-language alternative to combining high-level scripting with low-level systems languages, but achieving scalable performance still requires expertise in parallel programming. llms are increasingly used for code generation and are advancing rapidly with each new version. Yet, existing studies focus on single-shot prompting rather than agentic settings, in which an llm autonomously plans, generates,...
-
10
SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix Extensions
Feiyang Chen, Haibo Chen
cs.DC · cs.AI · cs.PF
Modern CPUs increasingly integrate matrix extensions, such as Arm Scalable Matrix Extension (SME), that provide high-throughput matrix execution within the CPU. For LLM inference, however, these units are not a universal replacement for conventional CPU cores: prefill, decode, attention, and KV-cache operations expose different arithmetic intensities, vector behavior, and layout requirements, while SME units and CPU cores still compete for...
This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.