cs.DC · 2026-07-12 · No. 51

Distributed, Parallel, and Cluster Computing, 2026-07-12.

8 new papers in cs.DC. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

8 entries
  1. 01

    SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling

    Jiahao Wang, Kaizhan Lin, Kaixi Zhang, Jinbo Han, Xingda Wei, Sijie Shen, Chenguang Fang, Wenyuan Yu, Rong Chen, Haibo Chen

    cs.DC · cs.AI

    LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by agents instead of humans. This shifts the workload in two ways: (1) agents act only on complete responses, making the cluster's tokens per second (TPS) the primary goal and relaxing--not eliminating--per-token latency requirements; and (2) requests share much of their KV\$-reuse exceeds 80% of request tokens...

    arxiv.org/abs/2607.08565 · PDF

  2. 02

    Coded Task Offloading for Fluid Computing: A Privacy-Aware Approach under D2D Networks

    Diego Cajaraville-Aboy, Manuel Fernández-Veiga, Ana Fernández-Vilas, Rebeca P. Díaz-Redondo

    cs.DC

    Fluid Computing aims to support distributed applications execution across heterogeneous cloud, edge, and device resources, motivating task execution mechanisms that adapt to dynamic and privacy-sensitive environments under runtime conditions. In this context, current task offloading schemes rarely address privacy risks and information leakage under adversarial execution settings; furthermore, most coded computing proposals focus on straggler...

    arxiv.org/abs/2607.08440 · PDF

  3. 03

    Computing in Anonymous Dynamic Networks with One-Bit Communications

    Thibaut Blanc, Giuseppe Antonio Di Luna, Giovanni Viglietta

    cs.DC

    We initiate the study of deterministic computation in anonymous dynamic networks where each agent broadcasts one bit per round and receives only the number of neighbors broadcasting each bit value. Despite this severe restriction, surprisingly rich global computation is possible. With a unique leader and a known upper bound $U$ on the network size $n$, we give a terminating algorithm for any computable function of the input multiset in...

    arxiv.org/abs/2607.08358 · PDF

  4. 04

    Adaptive Row Selection Meets Asynchrony in Randomized Kaczmarz

    Evan Coleman

    cs.DC · math.NA

    Randomized Kaczmarz is a natural fit for large sparse least-squares and tomographic reconstruction, and adaptive row selection can reduce iteration counts. However, deploying adaptive selection on a shared-memory machine means sampling from a residual that lock-free workers are concurrently modifying, often using stale data. We present the first systematic study of this regime: residual-weighted and greedy Kaczmarz under asynchronous...

    arxiv.org/abs/2607.08313 · PDF

  5. 05

    Empirical Analysis of GPU Frequency Behavior Under ML Workloads

    Truong-Thanh Le, Hoang-Loc La, Amir Taherkordi, Frank Eliassen, Phuong Hoai Ha, Peiyuan Guan

    cs.DC

    This work presents ongoing research on the frequency scaling behavior of NVIDIA GPUs when executing ML/AI workloads. Our preliminary findings show that, on lower-performance GPUs, the operating frequency is strongly affected by the recent workload history, typically within an 80ms window. This behavior challenges a common assumption underlying several state-of-the-art ML latency-prediction techniques, which treat individual GPU kernel...

    arxiv.org/abs/2607.08307 · PDF

  6. 06

    Self-Stabilizing Algorithms in the Uniform Port Model

    Liam Brinker, Yuval Emek, Oren Louidor

    cs.DC

    We introduce a distributed computational model referred to as the \emph{uniform port} model. An algorithm operating in this model is defined by means of local automata associated with the ports (a.k.a.\ half-edges) of the input graph. The crux of the uniform port model is that a single constant-size finite automaton is hosted by every port of every graph, making the model \emph{truly uniform}. Moreover, since the new model explicitly supports...

    arxiv.org/abs/2607.08244 · PDF

  7. 07

    On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend

    Zheng Yu

    cs.DC

    Non-GPU AI accelerators are increasingly adopted as alternatives to general-purpose GPUs for large-model inference, but the real engineering cost of migrating demanding workloads beyond CUDA remains poorly documented. We present a field study of deploying two large inference workloads on a 16-device Huawei Ascend 910 system using CANN and vLLM-Ascend: an LLM-as-a-judge safety and alignment evaluation pipeline based on a W8A8 MoE judge model,...

    arxiv.org/abs/2607.08215 · PDF

  8. 08

    Toward a Unified GPU-Aware OpenSHMEM Specification

    Naveen Ravi, Nathan Wichmann, Md. Wasi-ur- Rahman, Aurelien Bouteiller, Yıltan Hassan Temuçin, Avinash Kethineedi,...

    cs.DC

    Leadership-class HPC systems are now accelerator-centric, with GPUs providing most floating-point throughput and memory bandwidth. As next-generation systems increasingly integrate accelerators through high-speed memory fabrics and system interconnects, exposing larger tightly coupled device domains, \ac{PGAS} models such as OpenSHMEM provide a natural abstraction for expressing fine-grained remote memory operations across these devices....

    arxiv.org/abs/2607.08006 · PDF

This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.