cs.DC · 2026-09-02 · No. 103

Distributed, Parallel, and Cluster Computing, 2026-09-02.

9 new papers in cs.DC. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

9 entries
  1. 01

    Faster Convergence of Multidimensional Approximate Agreement via Smallest Enclosing Balls

    Darya Melnyk

    cs.DC

    This work considers the multidimensional approximate agreement problem. In this problem, $n$ parties in a distributed system, up to $t$ of which may be corrupted by a Byzantine adversary, need to output vectors that are close to each other and that lie inside the convex hull of all non-corrupted input vectors. We assume that nodes communicate in a fully-connected authenticated network and analyze synchronous and asynchronous communication...

    arxiv.org/abs/2609.01490 · PDF

  2. 02

    Just Talk Once: Communication-Efficient Split Federated LLM Fine-Tuning on Edge Devices

    Jiaxiang Geng, Xianhao Chen, Bing Luo

    cs.DC

    Large language model (LLM) fine-tuning is increasingly shifting toward data generated on edge devices, where memory, computation, bandwidth, and connectivity constraints make conventional federated learning difficult to sustain. Split federated fine-tuning (SFT) improves client-side efficiency by offloading most model parameters and computation to the server but requires step-by-step bidirectional communication loop across the split interface...

    arxiv.org/abs/2609.01457 · PDF

  3. 03

    Update for Decisions, Not Freshness: Goal-Oriented Status Updating and Selective Offloading at the Network Edge

    Jianpeng Qi, Qiyang Zhang, Chao Liu, Jing Sun, Yimei Liu, Yanwei Yu, Yingjie Wang, Wei Ni

    cs.DC · cs.MA · cs.NI

    In an edge--cloud collaborative edge-computing environment, an edge node (EN) must decide whether each user task should be executed locally, forwarded to a remote service (or cloud) node (SN), or rejected. The EN observes its local state directly but receives the SN state only through an intermittently refreshed cache. Status updating and task control therefore form an asynchronous closed loop under partial observability. Freshness-driven...

    arxiv.org/abs/2609.01082 · PDF

  4. 04

    MakoXC: Rearchitecting DFT Exchange-Correlation with Matrix-Aligned and Knowledge-Organized Sparsity

    Haozhi Han, Fusong Ju, Jing Bai, Ruge Zhang, Xiang Zhao, Liang Yuan, Yunquan Zhang, Ting Cao, Liu Yunxin, Yifeng Chen, Kun Li

    cs.DC · physics.chem-ph

    Density Functional Theory (DFT) is indispensable for materials science and drug discovery, yet the exchange--correlation (XC) evaluation remains a major bottleneck due to its cubic scaling. Although linear-scaling methods exploit electronic nearsightedness to reduce asymptotic complexity, they produce irregular sparse workloads that hide implicit sparsity and prevent efficient use of modern AI accelerators. We present MakoXC, a modular...

    arxiv.org/abs/2609.01025 · PDF

  5. 05

    AInfer-PD: Communication-Safe In-Place Prefill-Decode Multiplexing for Distributed MoE Rollouts

    Guowei Wang, Chaokun Yang, Zhenxuan Pan, Yuhong Guo, Minghua Zhu, Zhechuan Zhang, Shuo Wan, Xiaowei Zhu

    cs.DC

    Rollout inference often dominates the wall-clock time of large-scale reinforcement learning (RL). In agentic RL, each trajectory alternates between model generation and environment interaction over multiple turns. Asynchronous trajectories consequently introduce new prefill (P) work while other trajectories remain in decode (D), making P/D coexistence a persistent property of the rollout rather than a one-time prompt-ingestion event. On...

    arxiv.org/abs/2609.00993 · PDF

  6. 06

    Prediction-Robust Service Deployment with Capacity-Aware Edge Admission

    Hailiang Zhao, Ziqi Wang, Yifei Zhang, Mingyi Liu, Xinkui Zhao, Kingsum Chow, Shuiguang Deng

    cs.DC

    Edge platforms instantiate executable services close to users to reduce request-serving cost, but each instance incurs a one-time deployment cost and remains useful only for a finite time-to-live (TTL). The resulting online decision is both prediction-sensitive and capacity-coupled: an optimistic forecast can waste deployment cost, whereas a delayed decision misses the burst it is intended to serve. We study this problem under a common TTL...

    arxiv.org/abs/2609.00877 · PDF

  7. 07

    Shared-Memory Range-Tiled CDF Sort for Small-Range Integer Keys on GPUs

    Kento Ando, Kaito Takase, Noriyuki Fujimoto, Koichi Wada

    cs.DC

    We study unstable integer sorting on GPUs for arrays whose elements lie in a known integer range. Focusing on counting-sort-based methods that determine the output interval of each value from its frequency and the prefix sums of the frequencies, we propose and evaluate Range-Tiled CDF sort (RT-CDF), which partitions the possible value range into small intervals, called tiles, that fit in shared memory. For each tile, RT-CDF constructs a...

    arxiv.org/abs/2609.00843 · PDF

  8. 08

    Breaking Cycles for Scalable Fair Ordering in Blockchain Systems

    Jinchun He, Wangjie Qiu, Yizhong Liu, Shengda Zhuo, Kwok-Yan Lam

    cs.DC

    In blockchain systems, transaction order directly determines financial outcomes: unfair ordering enables front-running and sandwich attacks that have extracted over \$686M from Ethereum users. Current fair-ordering protocols aggregate pairwise receive-order evidence from replicas. Under contention or adversarial manipulation, however, Condorcet cycles force them into global strongly connected component (SCC) condensation, causing delays,...

    arxiv.org/abs/2609.00837 · PDF

  9. 09

    Characterizing the Scalability and Performance of Large-Scale AI Training Under Multi-Tenancy

    Jacopo Raffi, Thomas Pasquali, Lorenzo Piarulli, Filippo Spiga, Marco Faltelli, Andreas Herten, Domenico Siracusa,...

    cs.DC

    Characterising AI workload performance on modern HPC systems requires understanding both their scalability in isolation and their behaviour under concurrent execution. However, the interplay among parallelisation strategies, network congestion, compute capability, and interconnect technologies remains poorly understood. This work investigates the performance and scalability of AI models up to 2400 GPUs. We quantify the communication overheads...

    arxiv.org/abs/2609.00817 · PDF

This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.