cs.DC · 2026-09-25 · No. 124

Distributed, Parallel, and Cluster Computing, 2026-09-25.

8 new papers in cs.DC. Titles, authors, abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →

01 — The papers

8 entries
  1. 01

    Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG

    George Danezis, Zeno De Angeli, Philipp Jovanovic, Lefteris Kokoris-Kogias, Markus Legner, Alberto Sonnino

    cs.DC · cs.CR

    Dual-mode consensus protocols are fast when the network is partially synchronous and remain live under asynchrony. We introduce Steelhead, a dual-mode mechanism that composes a partially synchronous and an asynchronous commit rule over one DAG: every k-th round is decided by the asynchronous rule, whose leader a common coin reveals after the votes, and all other rounds by the partially synchronous rule. Every interval, validators replay the...

    arxiv.org/abs/2609.30163 · PDF

  2. 02

    KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization

    Aheli Poddar, Sanskar Prasad, Arindam Samanta, Subha Chakraborty, Vishal Goyal, Rohit Singh Rathaur

    cs.DC · cs.AI · cs.LG

    Deep learning inference and training performance depends critically on GPU kernel efficiency. Modern compilers such as PyTorch Inductor automatically generate GPU kernels from high-level model code, but frequently underperform expert-written implementations by wide margins. Recent LLM-assisted kernel optimizers can close this gap for standalone kernels, yet treat compiled models as black boxes, generally optimizing individual standalone...

    arxiv.org/abs/2609.30059 · PDF

  3. 03

    KREX: Concurrent Kernel Benchmarking on Shared GPUs via Region-Granular Exclusivity

    Tianyu Feng, Haoxuan Yu, Tianyuan Wu, Lingyun Yang, Daocheng Ying, Yuxiao Wang, Ruibo Fan, Yinghao Yu, Guodong Yang,...

    cs.DC · cs.OS · cs.PF

    LLM agents automate GPU kernel optimization by repeatedly composing candidates and measuring their duration on real GPUs. Existing systems preserve measurement fidelity by reserving a GPU for an entire agent session or benchmarking command. However, this results in poor utilization because only a small fraction of command execution requires exclusive GPU access. Sharing GPUs could recover this idle capacity, but introduces contention that...

    arxiv.org/abs/2609.30057 · PDF

  4. 04

    A Block Decomposed QUBO Workflow for Chromosome-Y Phylogeny Reconstruction

    Giuliana Siddi Moreau, Riccardo Berutti, Manuela Profir, Lorenzo Pisani, Maria Laura Clemente, Lidia Leoni

    cs.DC · quant-ph

    This paper sets out a computational workflow that reconstructs the phylogeny of human Y-chromosome populations from a Variant Call Format (VCF) file of biallelic Single Nucleotide Polymorphisms (SNP). The two classical phylogenetic decisions - topology selection and root placement - are cast as Quadratic Unconstrained Binary Optimisation (QUBO) problems. The workflow combines two QUBO formulations with an Alternating Direction Method of...

    arxiv.org/abs/2609.29856 · PDF

  5. 05

    Reusing Spare Vehicle Computing Capacity: Is It Viable, Profitable and Sustainable?

    Rosario Patanè, Nadjib Achir, Andrea Araldo, Lila Boukhatem

    cs.DC

    Vehicular Cloud Computing (VCC) exploits computing hardware already embedded in vehicles for purposes unrelated to offloading and puts its idle cycles to work executing end-users' offloaded tasks, avoiding the deployment of new computation infrastructure. Despite its conceptual appeal, adoption is hindered by the lack of quantitative evidence that sharing spare vehicular capacity is viable, profitable and sustainable. A management scheme for...

    arxiv.org/abs/2609.29731 · PDF

  6. 06

    Resource-Aware Model Selection for Scalable Indoor Localization on HPC Platforms

    Fukuharu Tanaka, Hamada Rizk, Moustafa Youssef, Hirozumi Yamaguchi

    cs.DC

    Large-scale indoor localization is increasingly needed in campuses, smart buildings, factories, and digital-twin infrastructures, where wireless conditions, access-point deployments, and spatial layouts evolve over time. Such systems must be accurate, extendable, and maintainable, allowing new buildings, floors, rooms, and service areas to be added without retraining a monolithic model. Modular learning-based localization supports this goal...

    arxiv.org/abs/2609.29402 · PDF

  7. 07

    Concurrent Split Learning Through Stable Client Clustering

    Mohammad Kohankhaki, Valentin Rentschler, Anke Schmeink

    cs.DC · cs.LG

    Training with a fixed global batch limits how many distributed clients can provide examples in any one step. We examine a way to use additional server workers without increasing the batch processed by an individual workload. Global Clustered Parallel Split Learning (GCPSL) assigns clients to fixed clusters, executes a Parallel Split Learning with Global Sampling (GPSL) workload for each cluster concurrently, and periodically fuses the client...

    arxiv.org/abs/2609.29395 · PDF

  8. 08

    TrafficFab: An Autonomic Edge-Cloud Testbed Fabric forAI-Driven Traffic Management

    Mayank Arya, Pranjal Naman, Priyanshu Pansari, Roopkatha Banerjee, Daksh Mehta, Manjil Nepal, Akash Sharma, Yogesh Simmhan

    cs.DC

    Traffic management in emerging megacities requires real-time analytics over thousands of CCTV video streams under latency, bandwidth, compute and energy constraints. We present TrafficFab, an autonomic edge--cloud testbed for AI-driven traffic management, designed to validate a representative slice of a megacity deployment. TrafficFab combines RTSP stream emulation, heterogeneous edge inference using DNNs, cloud-based nowcasting and...

    arxiv.org/abs/2609.29223 · PDF

This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.

Colophon Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber #D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.