cs.DC · 2026-09-04 · No. 105

Distributed, Parallel, and Cluster Computing, 2026-09-04.

5 new papers in cs.DC. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

5 entries
  1. 01

    Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

    Yujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra

    cs.DC · cs.LG · cs.PF

    As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining techniques distributing the computation across different on-chip processing units, while effective for throughput, do not address the latency demands posed by modern neural networks with complex interdependencies and extensive operator parallelism. There is a potential...

    arxiv.org/abs/2609.04168 · PDF

  2. 02

    Barnacle: Adaptive Multi-Leader Scheduling for DAG-Based Consensus

    Zeno De Angeli, Alexandru Ianov Vitanov, Philipp Jovanovic, Lefteris Kokoris-Kogias, Alberto Sonnino, Pasindu...

    cs.DC · cs.CR

    In DAG-based consensus, all validators propose blocks concurrently, and designated leader blocks drive transaction commit. Having multiple leader slots per round cuts queuing latency, yet production deployments run a single leader because of head-of-line blocking: a slow leader stalls the pipeline for at least one leader timeout, and for several waves when its slot must wait for the fallback indirect decision rule. This risk grows with the...

    arxiv.org/abs/2609.03978 · PDF

  3. 03

    Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in Einsummable

    Zhimin Ding, Chen-Kuan Liao, Chima Adiole, Brianna Barrow, Fangzhou Du, Yu Hsiao, Ge Huang, Yicheng Jin, Ismail...

    cs.DC

    Distributing an AI computation across the GPUs of a multi-GPU server is one of the central problems in systems-for-AI. We present Einsummable, a prototype system that accepts a PyTorch-like description of an AI computation and automatically distributes it across a multi-GPU server, with no device assignments, sharding annotations, or communication operations written by the programmer. Einsummable models every operation as a relational join...

    arxiv.org/abs/2609.03905 · PDF

  4. 04

    JuPyLive: Seamless Migration of Jupyter Notebook Resources from Laptop to HPC

    Sima Attar-Khorasani, Matthias Lieber, Siavash Ghiasvand

    cs.DC

    This work introduces JuPyLive, a migration mechanism that enables seamless transition of Jupyter notebooks between local resources of user's workstation and remote resources of high-performance computing~(HPC) environments, while preserving the user experience. JuPyLive eliminates the underlying complexities of migration process, enabling users to freely choose among available local and remote resources, directly within the familiar Jupyter...

    arxiv.org/abs/2609.03562 · PDF

  5. 05

    FlowTT: Exploiting Computation Flow Reuse in Irregular Tensor-Train Embedding

    Jongmin Seok, Chae Eun Rhee

    cs.DC · cs.AR

    Tensor-Train (TT) decomposition effectively compresses large embedding tables in recommendation models, but TT-based embedding lookup remains inefficient because partially shared computation flows across input indices are not fully reused and intermediate results are repeatedly materialized off-chip between sequential TT-core contractions. We present FlowTT, a flow-aware GPU execution framework that reformulates TT gather as a set of...

    arxiv.org/abs/2609.03459 · PDF

This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.