cs.DC · 2026-09-22 · No. 121
Distributed, Parallel, and Cluster Computing, 2026-09-22.
9 new papers in cs.DC. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
9 entries-
01
Who Pays for the KV Cache? Attributing Shared AI Inference Spend Across Kubernetes and LLM Provider Bills
Timothy Urista
cs.DC · cs.PF
Organizations pay for AI through disconnected ledgers: Kubernetes allocations for self-hosted inference, gateway logs, and per-token bills from API providers. We present unalloc, an open-source tool that joins OpenCost, LiteLLM, OpenAI and Anthropic cost data into one exact ledger and reports the share of spend with no owner, and use it to study where attribution breaks at the seams between these systems. Five case studies run inference for...
-
02
Tiga: Compiling Graph Message Passing at Scale
Mingyuan Chi
cs.DC
Graph message passing offers a common way to express learning algorithms, physical simulations, and numerical solvers. Efficient execution depends on interaction structure and data movement, which can be obscured when a program is expressed as a sequence of tensor operations. On memory-constrained systems such as laptops, materializing connectivity and intermediate messages can also exhaust device memory. We present Tiga, a just-in-time...
-
03
Mitigating Front-Running Attacks through Fair and Resilient Transaction Dissemination
Wassim Yahyaoui, Joachim Bruneau-Queyreix, Jérémie Decouchant, Marcus Völp
cs.DC
In modern blockchains, efficient, fair, and faulttolerant information dissemination is critical for performance and security. Several stages of the transaction lifecycle are affected, from the creation and dissemination of transactions to the dissemination of blocks in the consensus layer. Mempool protocols, such as L{Ø}, already address some of modern blockchains' security threats. However, others remain unless fairness is embedded most...
-
04
Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI Inference
Mingyuan Yan, Haiyu Wang, Linxuan Biao, H. Jonathan Chao, Sai Qian Zhang, Wenqi Cui
cs.DC · cs.PF · eess.SY
Power availability increasingly constrains the operation of AI inference fleets, creating a need for provisioning methods that jointly consider serving capacity and power consumption. Prefill--decode (PD) disaggregation has emerged as a prevalent architecture for large-scale inference serving. However, determining the appropriate numbers of prefill and decode instances is challenging because serving capacity depends jointly on workload...
-
05
Bridging the Vendor Gap: Enabling AMD GPU Support for Awkward Array via ROCm/HIP for the HL-LHC Era
Ianna Osborne, Maxym Naumchyk, Tai Sakuma, Andres Rios-Tascon, Peter Elmer
cs.DC · cs.PL
The High-Luminosity LHC (HL-LHC) will demand order-of-magnitude gains in analysis throughput, and increasingly those gains must come from GPUs that are not made by a single vendor. Leadership-class systems such as El Capitan, Frontier and LUMI are built on AMD accelerators, yet the Scikit-HEP analysis stack---and Awkward Array in particular---has grown up CUDA-first. We report on $rawkward$, a Rust-backed kernel engine that adds a ROCm/HIP...
-
06
Conduit: An Experience Data Plane for Distributed Reinforcement Learning
Sitong Zhang, Tuo Shi, Mario Di Francesco, Zeke Wang, Bo Zhao
cs.DC · cs.AI
Distributed reinforcement learning (RL) scales training by parallelizing actors and learners around an Experience Buffer. As RL workloads grow, however, the buffer becomes more than a replay queue: it is the storage substrate of a large-capacity, latency-critical experience path that every iteration traverses to move, transform, sample, and batch experiences before learner updates can begin. Existing RL systems embed this path inside...
-
07
Toward GPU-Resident Climate Models: A Feasibility Study on Lossy Compression for the Spherical Harmonic Transform's Communication Bottleneck
Lorenzo Breschi, Flavio Vella
cs.DC · physics.ao-ph
Operational pseudospectral atmospheric models such as the ECMWF Integrated Forecasting System (IFS) run today almost exclusively on CPUs; GPU ports are under active development but not yet used in production. These models rely on the Spherical Harmonic Transform (SHT). Each time-step requires forward and inverse SHTs, and both passes depend on global pencil transposition that redistribute multi-dimensional arrays across compute nodes. At...
-
08
A principled approach for energy-efficient training via phase-aware GPU frequency tuning
Miguel Braga, Júlio Pinto, Rahma Nouaji, Olivier Michaud, Bettina Kemme, Oana Balmau, Cláudia Brito, Ricardo Macedo
cs.DC · cs.LG
Modern AI model training imposes unprecedented computational demands, making it a key contributor to datacenter energy consumption. Yet a significant fraction of the energy consumed during training does not translate to useful computation due to bottlenecks throughout the training pipeline. We present PAFT, a phase-aware, dynamically adaptable GPU frequency tuning system that reduces energy consumption of training workloads with minimal...
-
09
MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents
Demetris Paschalides, Moysis Symeonides, George Pallis, Marios D. Dikaiakos
cs.DC · cs.AI · cs.LG
As LLM agents increasingly interact with external tools through standardized protocols such as MCP, tool-interface design becomes a critical yet underexplored factor. How funψtionality is decomposed into tools affects whether an agent can select the right tool and construct valid arguments. This choice is especially consequential at the edge, where resource constraints limit which models can run locally and scaling up is often not an option....
This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.