cs.DC · 2026-07-14 · No. 53

Distributed, Parallel, and Cluster Computing, 2026-07-14.

8 new papers in cs.DC. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

8 entries
  1. 01

    MemExchange: Cloud-Scale Memory Trading

    AmirHossein Seyri, Abhisek Pan, Balajee Vamanan

    cs.DC

    To handle unpredictable workloads, cloud providers typically over-provision memory to meet peak demand, resulting in substantial underutilization across datacenter clusters. At the same time, memory-constrained tenants may suffer elevated cache miss rates, even when idle capacity remains stranded elsewhere in the infrastructure. MemExchange is a cluster-wide, multi-tenant memory management system that dynamically right-sizes in-memory caching...

    arxiv.org/abs/2607.11579 · PDF

  2. 02

    Stage-Level Executor Allocation in Apache Spark with Cost-Performance Trade-offs

    Miriam Rateike, Isaac Waweru Wambugu, Celia Cintas, Michael Kaufmann, Ioana Giurgiu, Skyler Speakman

    cs.DC

    Allocating executors (i.e. compute resources) to distributed processing systems must balance resource costs of scaling-out unnecessarily against artificial, performance-limiting bottlenecks. Naive approaches may allocate executors at the application level, which have predictable costs and performance but are almost guaranteed to be sub-optimal for each of the thousands of diverse, individual stages executed by the application. Users may also...

    arxiv.org/abs/2607.11415 · PDF

  3. 03

    Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUs

    Weijia Han, Lisha Qu

    cs.DC · cs.LG · cs.PF

    Reported serving speedups from quantized kernels typically bundle the weight format, the kernel, and the inference runtime into one number. We present an attribution study on four NVIDIA RTX A5000 GPUs, 24 GiB each, on a single host with NVLink-bridged pairs. A matched intermediate stack that keeps the faster runtime without the quantized kernel splits the full speedup into a runtime part and a kernel and quantization part. Under matched...

    arxiv.org/abs/2607.11368 · PDF

  4. 04

    GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design

    Yitong Ding, Jiawei Huang, Renyang Guan, Yangjie Zhou, Zihan Liu, Yu Feng, Shixuan Sun, Mingyi Guo, Jingwen Leng, Jian Weng

    cs.DC

    Modern LLM (large language model) workloads increasingly rely on optimized GPU kernels through hardware-software co-design. These kernels achieve high-performance through fine-grained dependency scheduling and computation-memory overlap. As such, they incur new challenges on existing GPU performance models. Instruction-driven simulators are costly to adapt to evolving architectures, while analytical models are too coarse to capture kernels'...

    arxiv.org/abs/2607.11262 · PDF

  5. 05

    Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration

    Xueze Kang, Guangyu Xiang, Suyi Li, Yuxin Wang, Shaohuai Shi, Lin Zhang, Xiaowen Chu

    cs.DC

    Diffusion models are increasingly deployed as production visual-generation services, where serving high-resolution image and long video generation is often limited by GPU memory. Popular memory-saving techniques such as weight offloading, sharding, and VAE slicing are often not practical because they tend to introduce significant performance overhead. In this paper, we present Xema, a memory-efficient diffusion serving system that exploits...

    arxiv.org/abs/2607.11136 · PDF

  6. 06

    Energy Calculus: A Compositional Algebra of Energy in Computational Systems

    Mosharaf Chowdhury, Jae-Won Chung, Jeff J. Ma, Nishil Talati, Ruofan Wu

    cs.DC · cs.PF

    Energy is a binding constraint for AI scaling, yet it lacks the formal treatment that computation, communication, and learning have long enjoyed. Recent systems demonstrate large energy savings, but each targets a specific granularity and structure; one cannot combine frequency scaling from one system with critical-path analysis from another and reason about their joint effect on total energy. Energy remains a monolithic scalar that is...

    arxiv.org/abs/2607.11087 · PDF

  7. 07

    Domain Extension of Lock-Freedom and Wait-Freedom for Group Computations

    Raaghav Ravishankar, Sandeep Kulkarni, Sathya Peri, Gokarna Sharma, Manaswini Piduguralla, Yashi Rastogi

    cs.DC

    A domain extension of a definition refers to broadening the scope of a definition so that it applies to a larger set of cases than originally specified. The notion of lock-free and wait-free computation is designed for the domain of tasks that are completed by a single thread (in competition with other threads). The goal of this paper is to extend the definition of lock-freedom and wait-freedom to group-computations (denoted by gl-freedom and...

    arxiv.org/abs/2607.11014 · PDF

  8. 08

    [AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent Workflows

    Arup Kumar Sarker, Alexander James Halpern, Mills Staylor, Aymen Alsaadi, Gregor von Laszewski, Yue Cheng, Shantenu...

    cs.DC

    Multi-agent LLM systems increasingly integrate retrieval, planning, and reasoning, but remain fundamentally text-centric, requiring agents to repeatedly recompute shared context through expensive prefill. Although single-request inference is known to be accelerated by KV-cache management, it is usually restricted to local serving scopes. We introduce AAFLOW+, a stateful extension of agentic workflow operators that makes KV cache a first-class...

    arxiv.org/abs/2607.10987 · PDF

This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.