cs.DC · 2026-06-30 · No. 39

Distributed, Parallel, and Cluster Computing, 2026-06-30.

7 new papers in cs.DC. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

7 entries
  1. 01

    Data Replication Meets Function Scheduling in the Edge-Cloud Continuum

    Matteo Cenzato, Dario d'Abate, Arianna Dragoni, Matteo Briscini, Alessandro Margara

    cs.DC

    Serverless computing is an appealing model for the edge-cloud continuum, but its stateless assumption breaks down once functions need persistent data: fetching state from a distant cloud store erases the latency benefit of running at the edge. Keeping data close means replicating it, and replication forces a placement decision that is coupled with where functions execute and with the consistency each application demands. We study this joint...

    arxiv.org/abs/2606.30563 · PDF

  2. 02

    Spandana: Reconciling Strict SLOs with Low Cost under Fine-Grained Load Fluctuations

    Dilina Dehigama, Shyam Jesalpura, Zeyu Xu, Marton Nemeth, Shengda Zhu, Marios Kogias, Boris Grot

    cs.DC

    Cloud-based online services face significant sub-second load fluctuations while needing to meet strict Service Level Objectives (SLOs). Cluster operators often over-provision resources to protect SLOs, sacrificing utilization and cost efficiency. Existing reactive and proactive autoscalers, serverless (FaaS) deployments, and VM/FaaS hybrid systems fail to reconcile strict SLO compliance with low cost and high utilization under fine-grained...

    arxiv.org/abs/2606.30533 · PDF

  3. 03

    GPU Parallelization Strategies for Forward and Backward Propagation in Shallow Neural Networks: A CUDA-Based Comparative Study

    Rania Zitouni, Nadine Bousdjira, Sarah Hasnaoui, Amel Sadoun, Fatma Salhi

    cs.DC · cs.LG

    We present a comparative study of CUDA optimization strategies applied to forward and backward propagation in a shallow neural network. Three stacked optimizations are evaluated: (1) tiled shared memory with bank-conflict elimination via +1-column padding, (2) pre-transposed weight matrices for coalesced global memory access, and (3) a fused MatMul+ReLU kernel that eliminates intermediate global-memory round-trips. Experiments on an NVIDIA...

    arxiv.org/abs/2606.30497 · PDF

  4. 04

    Analyzing Linearizability in Relativistic Distributed Systems

    Kahbod Aeini, Wojciech Golab

    cs.DC

    Einstein's theory of relativity correctly predicted that time is relative, and subject to both kinematic and gravitational dilation. Therefore, executions of distributed systems cannot always be modeled as sequences of events totally ordered according to wall clock time. To address this fundamental problem, Gilbert and Golab formulated a generalization of Herlihy and Wing's linearizability property for shared objects, which they called...

    arxiv.org/abs/2606.30419 · PDF

  5. 05

    Energy-Aware Scheduling for Serverless LLM Serving on Shared GPUs

    Tianyu Wang, Gourav Rattihalli, Aditya Dhakal, Longfei Shangguan, Dejan Milojicic

    cs.DC

    As LLM inference becomes a major cloud workload, its growing energy footprint makes cluster-wide energy optimization increasingly important. Serverless LLM serving helps platforms absorb traffic volatility by elastically sharing GPU resources across models, but this sharing also makes energy optimization difficult. Multiple co-resident models run under one device-wide operating point, while their resource demands and latency slack change...

    arxiv.org/abs/2606.30391 · PDF

  6. 06

    FBench: A Flexible Benchmark for CFG-Based What-If Exploration of HPC I/O Patterns

    Zhaobin Zhu, Chen Wang, Kathryn Mohror, Sarah Neuwirth

    cs.DC · cs.PF

    The I/O performance of large-scale HPC applications depends on a complex interplay of access patterns, middleware optimizations, and file system configurations. To systematically explore these effects without repeatedly rerunning full applications, we introduce FBench, a flexible and code-transparent benchmarking tool for what-if analysis and I/O performance exploration. FBench leverages context-free grammars (CFGs) derived from Recorder...

    arxiv.org/abs/2606.30197 · PDF

  7. 07

    Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference

    Hui Zang, Pengfei Xia, Hong Liu, Jiajia Chu, Tuo Hao, Minghao Chen, Rui Zhang, Ziyang Zhang

    cs.DC

    Mixture-of-Experts (MoE) architectures enable language models to achieve unprecedented scale via sparse activation. However, their inference performance is often limited by data movement bottlenecks. Two coupled challenges exacerbate this limtation: (1) Importance-Agnostic Cost: Low-contribution experts incur nearly uniform memory and transfer costs, resulting in a low cost-to-benefit ratio and wasting critical bandwidth; (2) System-Level...

    arxiv.org/abs/2606.29982 · PDF

This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.