cs.DC · 2026-09-15 · No. 114

Distributed, Parallel, and Cluster Computing, 2026-09-15.

6 new papers in cs.DC. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

6 entries
  1. 01

    Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility Twin

    Weqaar Janjua, Eoin OConnell, Mihai Penica

    cs.DC · cs.AR · cs.NI

    Modern AI and HPC systems integrate accelerators, high-speed networks, and management controllers at rack scale. Developing software for this infrastructure typically requires access to scarce, costly hardware, while software abstractions can obscure how workloads depend on resources across servers and accelerators. This paper presents Cnuas, an open-source, experimental rack-scale emulation platform whose baseline architecture follows the...

    arxiv.org/abs/2609.15889 · PDF

  2. 02

    Scalability and Performance Evaluation of Federated Learning Frameworks: A Comparative Analysis

    Bassel Soudan, Sohail Abbas, Ahmed Kubba, Manar Wasif Abu Talib, Qassim Nasir

    cs.DC · cs.AI

    This paper presents a systematic examination and experimental comparison of the prominent Federated Learning (FL) frameworks FedML, Flower, Substra, and OpenFL. The frameworks are evaluated experimentally by implementing Federated Learning over a varying number of clients, emphasizing a thorough analysis of scalability and key performance metrics. The study assesses the impact of increasing client counts on total training time, loss and...

    arxiv.org/abs/2609.15681 · PDF

  3. 03

    CIDERS: Cloud-Edge LLM Collaborative Learning via Accelerating Personalized Bilevel Optimization

    Victor H. Chen, Hairui Yu, Stella K. Chung, Hong Yan

    cs.DC · cs.AI

    Amid the rapid advancement of physical-world intelligence, cloud-edge collaborative large language models (LLMs) have emerged as a promising roadmap for practical LLM deployment. However, existing cloud-edge paradigms struggle to balance global consensus with local personalization, which fails to satisfy the need for a unified knowledge foundation on the cloud and domain-specific adaptation at the edge. To address this, we introduce, for the...

    arxiv.org/abs/2609.15664 · PDF

  4. 04

    DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance Engineering

    Siming Huang

    cs.DC · cs.AR · cs.PF

    We present the enablement, correctness recovery, and performance engineering of DeepSeek-V4-Flash inference on AMD Instinct MI250 GPUs using the gfx90a/CDNA2 architecture. The system integrates native safetensors loading, tensor and expert parallelism, FP4 routed mixture-of-experts computation, FP8 dense projections, sparse attention, HIP graph execution, and OpenAI-compatible serving within SGLang. An initially fast execution path was found...

    arxiv.org/abs/2609.15627 · PDF

  5. 05

    ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI Datacenters

    Rui Lu, Rui Ge, Huanghuang Liang, Xiaobo Zhou, Dan Wang

    cs.DC · cs.PF · eess.SY

    Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and carbon, but also shrinks thermal headroom, induces GPU throttling, and leads to Service-Level-Objective (SLO) violations. In this paper, we study joint cooling--computing control for LLM inference: minimizing per-job GPU-plus-cooling energy while...

    arxiv.org/abs/2609.15230 · PDF

  6. 06

    Accelerating the Solving of Many Tiny General Linear Systems on GPUs: Application to Constitutive Laws

    Tristan Chenaille, Francesca Cuteri, Rapha{ë}l Prat, Guillaume Latu, Thomas Helfer

    cs.DC

    Many applications require solving large numbers of independent linear systems on GPUs. While this need is well addressed for small to large systems, tiny ones, understood here as systems of dimension below 32, remain challenging. This is especially relevant in constitutive law evaluation, where millions of integration points are handled independently, and where each constitutive update generally relies on a Newton iterative method. Each...

    arxiv.org/abs/2609.15217 · PDF

This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.