cs.DC · 2026-06-17 · No. 26

Distributed, Parallel, and Cluster Computing, 2026-06-17.

9 new papers in cs.DC. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

9 entries
  1. 01

    Latency Prediction for LLM Inference on NPU Systems

    Juhyun Park, Seungwoo Jeong, Jingyu Lee, Kyungyong Lee

    cs.DC

    Deploying Large Language Models (LLMs) requires exploring a large configuration space spanning parallelization strategies, batching techniques, and scheduling policies. Exhaustive measurement across this space is impractical, making latency prediction essential for system optimization. While NPUs have emerged as accelerators designed for LLM inference, no prediction methodology has been established for them. Specifically, applying prior work...

    arxiv.org/abs/2606.18042 · PDF

  2. 02

    RouteBalance: Fused Model Routing and Load Balancing for Heterogeneous LLM Serving

    Wei Da, Evangelia Kalyvianaki

    cs.DC

    Heterogeneous LLM serving stacks split scheduling into two layers that optimize in isolation: model routers pick a model from quality and cost signals while ignoring instance load, and serving load balancers optimize queues while ignoring quality. We present RouteBalance, a serving-aware scheduling layer that fuses both into a single online assignment over concrete model instances, jointly trading off quality, latency, and cost. A batched...

    arxiv.org/abs/2606.17949 · PDF

  3. 03

    An Epistemic Analysis of Random Coordinated Attack

    Sophia Knight, David Lehnherr, Sergio Rajsbaum

    cs.DC · cs.LO

    The coordinated attack problem models the challenge of coordinating a joint action within a bounded time by communicating over unreliable links. It was the first distributed computing problem proven unsolvable. Its analysis also revealed the importance of common knowledge, a central concept in epistemic logic. However, the randomized version of coordinated attack, which is solvable, has not, to the best of our knowledge, been studied through...

    arxiv.org/abs/2606.17860 · PDF

  4. 04

    LUMEN: Coordinated Failure Recovery for Distributed LLM Serving

    Zhang Cao, Shujie Han, Juncheng Zhang, Yuanming Ren, Yongkun Li, Patrick P. C. Lee

    cs.DC

    Modern large language model (LLM) serving clusters distribute inference requests across multiple worker processes on different GPUs, but failures are prevalent at scale. When a worker fails, the cluster simultaneously loses the failed worker's GPU-resident key-value (KV) caches and serving capacity, leaving surviving workers to absorb the redirected traffic while re-running interrupted requests from scratch. Existing fault-tolerant systems...

    arxiv.org/abs/2606.17787 · PDF

  5. 05

    From GPU to Microcontroller: Online Ridge Regression for Edge-Deployable Traffic Prediction

    Suresh Purini, Archit Narwadkar, Deepak Gangadharan

    cs.DC

    State-of-the-art traffic flow forecasting models, including Graph Convolutional Networks and graph-less MLPs, require centralized GPU training across all sensors, making them impractical for resource-constrained intelligent transportation deployments. We show that much of this complexity is unnecessary. A parametric analysis of the recent graph-less model GLMST reveals that reducing its internal embedding dimension from 64 to 4 degrades MAPE...

    arxiv.org/abs/2606.17613 · PDF

  6. 06

    AoiZora: Topology-Aware Auto-Parallel Optimization for Inference of Diffusion Transformers

    Kaijian Wang, Yuanyuan Xu, Fanjiang Ye, Ye Cao, Jingwei Zuo, T. S. Eugene Ng, Yarong Mu, Yuke Wang

    cs.DC · cs.LG

    Video diffusion has quickly grown into a key generative serving workload, yet producing each clip demands many denoising iterations over large spatio-temporal latents, which puts low-latency inference out of reach on a single device. A denoising step is therefore typically distributed across multiple accelerators, and TPU sub-slices have become an attractive and practical fabric for doing so. Current auto-parallel systems, however, search...

    arxiv.org/abs/2606.17566 · PDF

  7. 07

    Multi-Orientation Edge-Minimum Repair for Non-Redundant Fault-Tolerant Broadcasting in Dense Gaussian Networks

    Bader Albader

    cs.DC · cs.IT · cs.NI

    Dense Gaussian networks are degree-four algebraic interconnection networks with compact diameter and simple modular routing. This paper studies non-redundant one-to-all broadcast repair in the dense Gaussian network generated by $α=k+(k+1)i$. We propose multi-orientation edge-minimum repair (MOEM), which evaluates a constant-size family of Gaussian broadcast-tree orientations, selects a fault-aware orientation, contracts the fault-pruned tree...

    arxiv.org/abs/2606.17528 · PDF

  8. 08

    Local Fault Repair of Perfect Resource Placements in Dense Gaussian Networks

    Bader Albader

    cs.DC · cs.IT · cs.NI

    Perfect resource placement in dense Gaussian networks partitions the network into Lee balls centered at resource nodes. The fault-free placement problem is already classified; this paper studies the complementary post-deployment problem of repairing such placements after resource faults. The paper gives exact local repair theorems for the dense Gaussian placement generated by $t+(t+1)i$; by conjugation and rotation symmetry, the same results...

    arxiv.org/abs/2606.17527 · PDF

  9. 09

    SpecGen: Accelerating Agentic Kernel Optimization with Speculative Generation

    Jihu Guo, Sitian Lu, Tenghui Ma, Wei Gao, Zhisheng Ye, Xingcheng Zhang, Dahua Lin

    cs.DC

    Agentic kernel optimization automates manual GPU kernel tuning via iterative generation, validation, and profiling with reasoning LLMs, casting the optimization task as feedback-guided search. However, our workload characterization reveals three system-level inefficiencies that limit search efficiency: (1) long generation latency due to LLM reasoning, (2) insufficient profiling feedback, and (3) underutilized validation/profiling resources....

    arxiv.org/abs/2606.17518 · PDF

This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.