cs.DC · 2026-10-01 · No. 130
Distributed, Parallel, and Cluster Computing, 2026-10-01.
8 new papers in cs.DC. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
8 entries-
01
Reinforcement Learning-Guided Graph Transformations for SpTRSV Optimization
Buse Yılmaz
cs.DC · cs.LG
Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit the available parallelism and make efficient workload distribution challenging. Recent graph transformation techniques address these limitations by modifying the dependency graph of the input matrix to improve parallel execution. Existing graph...
-
02
Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs
Jaehwan Lee, Sangmin Lee, Chaewon Kim, Junsik Shin, Jaejin Lee
cs.DC · cs.LG
Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer. As contemporary MoE models activate more experts per token, this communication accounts for a growing fraction of inference time. The cost becomes particularly pronounced on PCIe-based consumer GPU systems, where all inter-GPU transfers...
-
03
LatencyLab: A DPDK-Based P4 Pipeline Latency Measurement Framework for FPGA SmartNICs
Pavani Kuppili, Zhaoyang Han, Yicheng Qian, Suranga Handagala, Michael Zink, Miriam Leeser, Robert Ricci
cs.DC
P4-programmable FPGA SmartNICs place packet processing directly on the wire, but open FPGA P4 toolflows do not expose timestamping at the pipeline boundary, so the latency a P4 program adds on the target FPGA is rarely measured. This paper presents LatencyLab, a DPDK-based measurement framework for FPGA P4 pipeline latency that needs neither PHC/PTP support on the datapath nor clock synchronization. The FPGA's two ports share a network...
-
04
EPR Count for Runtime Prediction in Distributed Quantum Computing
Fatih E. Bilgen, Ozgur B. Akan
cs.DC · quant-ph
EPR-pair consumption is commonly used as a communication-cost objective in distributed quantum computing, but minimizing EPR cost does not necessarily minimize distributed execution time. Despite its widespread use, the reliability of EPR count as a runtime surrogate has received limited direct characterization across different workloads and communication conditions. This work addresses this gap by systematically evaluating the relationship...
-
05
From Pilots to Production: Lessons in Cross-Institutional Federated Training and Artificial Intelligence for Science
Olivera Kotevska, Max Carlson, Yan Gao, Francis Jeanson, Yijiang Li, William Lindskog, Mohammad Naseri, Minseok Ryu,...
cs.DC
Many of the most valuable scientific datasets cannot be centralized: they are proprietary, export-controlled, classified, or bound by data-sovereignty restrictions. This inverts the usual paradigm: the model must move to the data, making federated artificial intelligence (AI) core infrastructure for open science. We synthesize lessons from U.S. Department of Energy national laboratories, industry deployments, and the open-source community...
-
06
Darpan: A Digital Twin Framework for the Next-Generation Computing Continuum
Zhiyu Wang, Rajkumar Buyya
cs.DC
Computing-continuum applications distribute work across devices, edge systems, fog resources, and clouds. While a placement, scheduling, or recovery decision is being made, resource availability, network conditions, and application progress may change, so the decision can be invalid by the time it is executed. Existing runtimes enact predetermined decisions, whereas simulation tools compare alternatives in preconfigured environments; neither...
-
07
Exploring Adaptive Byzantine Quorum Systems to Improve Latency in the WAN
Linus Gnan, Rüdiger Kapitza, Christian Berger
cs.DC
Quorum systems enforce strict consistency in Byzantine fault-tolerant (BFT) state machine replication: Before a value is decided, a subset of replicas (called quorum) must exchange votes for the value. In wide-area networks, the size and composition of a quorum determines the speed at which replicas can make progress and thus impacts the latency perceived by clients. A variety of quorum constructions has been proposed, e.g., threshold,...
-
08
Kirin: Cloud-native WebAssembly Service Orchestration
Joshua Bauer, Sebastian Werner, Maria C. Borges
cs.DC
Modern cloud computing infrastructure relies heavily on virtualization to provide workload isolation and resource efficiency. While containers have become the dominant deployment primitive due to their fast orchestration compared to traditional virtual machines, they still introduce non-trivial overhead. WebAssembly (Wasm) has emerged as an alternative isolation technology, offering a lightweight execution model. Because of compatibility...
This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.