cs.DC · 2026-08-04 · No. 74

Distributed, Parallel, and Cluster Computing, 2026-08-04.

10 new papers in cs.DC. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

10 entries
  1. 01

    Analyzing GPU Performance in Virtualized Environments: A~Case Study

    Adel Belkhiri, Michel Dagenais

    cs.DC · cs.PF

    The graphics processing unit (GPU) plays a crucial role in boosting application performance and enhancing computational tasks. Thanks to its parallel architecture and energy efficiency, the GPU has become essential in many computing scenarios. On the other hand, the advent of GPU virtualization has been a significant breakthrough, as it provides scalable and adaptable GPU resources for virtual machines. However, this technology faces...

    arxiv.org/abs/2608.02414 · PDF

  2. 02

    Greedy-Like Defective Coloring: Distributed Algorithms and Applications

    Marc Fuchs, Fabian Kuhn

    cs.DC

    A $d$-defective $c$-coloring of a graph $G=(V,E)$ is a coloring of the nodes $V$ with $c$ colors such that every node has at most $d$ neighbors of the same color. Distributed algorithms for computing different variants of defective coloring are at the core of most deterministic state-of-the-art distributed coloring algorithms, and they are also an important tool in many other distributed graph algorithms. In several cases, the overall...

    arxiv.org/abs/2608.02386 · PDF

  3. 03

    Epico: Long-Lived WebAssembly Components for High-Performance Serverless Stream Processing

    Matteo Della Bartola, Valerio Besozzi, Patrizio Dazzi, Marco Danelutto

    cs.DC

    While serverless computing is popular, its dominant Function-as-a-Service (FaaS) model is ill-suited for stream processing because its stateless, centrally orchestrated functions cannot efficiently handle continuous, low-latency event flows. We introduce Epico, a serverless runtime explicitly designed to resolve these inefficiencies at the runtime level. Epico executes pipeline stages as persistent WebAssembly components, enabling...

    arxiv.org/abs/2608.02361 · PDF

  4. 04

    Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling

    Dayi Yao, Zijie Zhou

    cs.DC · eess.SY

    This paper studies a resource-allocation inefficiency in batched large language model (LLM) serving: heterogeneous requests that share a decode batch impose max-driven computational costs on one another. Because the wall-clock cost of a batch step is largely governed by the largest active KV-cache footprint, a short request co-batched with a long request can experience latency and GPU-resource consumption disproportionate to its own token...

    arxiv.org/abs/2608.02244 · PDF

  5. 05

    DEFT: Joint Task Placement and DVFS for Energy-Efficient Multi-GPU Runtimes

    Jing Chen, Miquel Pericas

    cs.DC

    Energy efficiency has become a first-order concern in modern high-performance computing systems, as it directly determines achievable throughput under fixed power budgets. Although Dynamic Voltage and Frequency Scaling (DVFS) provides an effective mechanism for reducing GPU energy consumption, existing runtime systems decouple DVFS from task placement and inter-GPU communication, focus on single-GPU execution, or cannot adapt frequency to...

    arxiv.org/abs/2608.02122 · PDF

  6. 06

    Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPO

    Ngoc Hung Nguyen, Bjorn Landfeldt

    cs.DC · cs.LG · cs.NI

    This paper investigates collaborative mobile edge computing (MEC) servers for large language model (LLM) inference under soft deadline constraints. In this system, to improve the quality of service, computations are expected to be completed within their deadlines. However, due to dependencies among tasks or subtasks, any missed deadline can lead to catastrophic consequences for the entire request. In this context, this work proposes an...

    arxiv.org/abs/2608.02031 · PDF

  7. 07

    TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous Driving

    Mengmeng Zhu, Yuxuan Sun, Wei Chen, Bo Ai

    cs.DC · cs.AI · cs.NI

    The deployment of Vision-Language Models (VLMs) in autonomous driving (AD) systems is constrained by on-board computing power, restricting vehicles to small VLMs (SVLMs) with limited perception and reasoning capabilities. Infrastructure-assisted AD alleviates this resource constraint by enabling collaboration with large VLMs (LVLMs) at edge servers. However, in dynamic vehicular environments, the utility of sensory data for downstream tasks...

    arxiv.org/abs/2608.01998 · PDF

  8. 08

    Diagnosing High-Performance BFT Consensus via Mixture Modeling of Block Time Distributions

    Hongru He, Akihiro Fujihara

    cs.DC · cs.CE · cs.CR · cs.PF

    High-performance Byzantine Fault Tolerant (BFT) blockchains are designed to achieve high throughput and low latency, yet their observed block time distributions often reveal complex behaviors arising from networking, pipelining, and deployment heterogeneity. In this paper, we diagnose HotStuff-based high-performance BFT consensus by modeling block times through a quorum-based multicast framework that links each block interval to quorum...

    arxiv.org/abs/2608.01934 · PDF

  9. 09

    Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling

    Cunchen Hu, Liangliang Xu, Tian Liu, Min Lyu, Yongkun Li, Sa Wang, Shuo Quan, Yanan Yang, Wenda Tang, Yiduo Wang, Fu...

    cs.DC · cs.AI

    Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing energy consumption. Existing energy-management approaches adapt GPU frequencies only at the request or inference-phase level, overlooking operator-level differences in frequency sensitivity between Attention and feed-forward networks (FFNs). We find that the...

    arxiv.org/abs/2608.01891 · PDF

  10. 10

    FedJigsaw: Multi-Agent Collaborative Model Reassembly for Decentralized Heterogeneous Federated Learning

    Jifeng Chen, Haibo Zhang, Yawen Chen

    cs.DC

    Model Heterogeneous Federated Learning (MHFL) addresses client-level resource heterogeneity by allowing each participant to train a personalized model architecture under a shared training objective. A prevalent paradigm, Partial Training (PT), achieves this by allowing each client to train a subnetwork of the global model. However, existing PT methods typically rely on predefined architectural templates or over-parameterized supernets,...

    arxiv.org/abs/2608.01861 · PDF

This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.