cs.DC · 2026-09-25 · No. 124
Distributed, Parallel, and Cluster Computing, 2026-09-25.
8 new papers in cs.DC. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
8 entries-
01
Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
George Danezis, Zeno De Angeli, Philipp Jovanovic, Lefteris Kokoris-Kogias, Markus Legner, Alberto Sonnino
cs.DC · cs.CR
Dual-mode consensus protocols are fast when the network is partially synchronous and remain live under asynchrony. We introduce Steelhead, a dual-mode mechanism that composes a partially synchronous and an asynchronous commit rule over one DAG: every k-th round is decided by the asynchronous rule, whose leader a common coin reveals after the votes, and all other rounds by the partially synchronous rule. Every interval, validators replay the...
-
02
KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization
Aheli Poddar, Sanskar Prasad, Arindam Samanta, Subha Chakraborty, Vishal Goyal, Rohit Singh Rathaur
cs.DC · cs.AI · cs.LG
Deep learning inference and training performance depends critically on GPU kernel efficiency. Modern compilers such as PyTorch Inductor automatically generate GPU kernels from high-level model code, but frequently underperform expert-written implementations by wide margins. Recent LLM-assisted kernel optimizers can close this gap for standalone kernels, yet treat compiled models as black boxes, generally optimizing individual standalone...
-
03
KREX: Concurrent Kernel Benchmarking on Shared GPUs via Region-Granular Exclusivity
Tianyu Feng, Haoxuan Yu, Tianyuan Wu, Lingyun Yang, Daocheng Ying, Yuxiao Wang, Ruibo Fan, Yinghao Yu, Guodong Yang,...
cs.DC · cs.OS · cs.PF
LLM agents automate GPU kernel optimization by repeatedly composing candidates and measuring their duration on real GPUs. Existing systems preserve measurement fidelity by reserving a GPU for an entire agent session or benchmarking command. However, this results in poor utilization because only a small fraction of command execution requires exclusive GPU access. Sharing GPUs could recover this idle capacity, but introduces contention that...
-
04
A Block Decomposed QUBO Workflow for Chromosome-Y Phylogeny Reconstruction
Giuliana Siddi Moreau, Riccardo Berutti, Manuela Profir, Lorenzo Pisani, Maria Laura Clemente, Lidia Leoni
cs.DC · quant-ph
This paper sets out a computational workflow that reconstructs the phylogeny of human Y-chromosome populations from a Variant Call Format (VCF) file of biallelic Single Nucleotide Polymorphisms (SNP). The two classical phylogenetic decisions - topology selection and root placement - are cast as Quadratic Unconstrained Binary Optimisation (QUBO) problems. The workflow combines two QUBO formulations with an Alternating Direction Method of...
-
05
Reusing Spare Vehicle Computing Capacity: Is It Viable, Profitable and Sustainable?
Rosario Patanè, Nadjib Achir, Andrea Araldo, Lila Boukhatem
cs.DC
Vehicular Cloud Computing (VCC) exploits computing hardware already embedded in vehicles for purposes unrelated to offloading and puts its idle cycles to work executing end-users' offloaded tasks, avoiding the deployment of new computation infrastructure. Despite its conceptual appeal, adoption is hindered by the lack of quantitative evidence that sharing spare vehicular capacity is viable, profitable and sustainable. A management scheme for...
-
06
Resource-Aware Model Selection for Scalable Indoor Localization on HPC Platforms
Fukuharu Tanaka, Hamada Rizk, Moustafa Youssef, Hirozumi Yamaguchi
cs.DC
Large-scale indoor localization is increasingly needed in campuses, smart buildings, factories, and digital-twin infrastructures, where wireless conditions, access-point deployments, and spatial layouts evolve over time. Such systems must be accurate, extendable, and maintainable, allowing new buildings, floors, rooms, and service areas to be added without retraining a monolithic model. Modular learning-based localization supports this goal...
-
07
Concurrent Split Learning Through Stable Client Clustering
Mohammad Kohankhaki, Valentin Rentschler, Anke Schmeink
cs.DC · cs.LG
Training with a fixed global batch limits how many distributed clients can provide examples in any one step. We examine a way to use additional server workers without increasing the batch processed by an individual workload. Global Clustered Parallel Split Learning (GCPSL) assigns clients to fixed clusters, executes a Parallel Split Learning with Global Sampling (GPSL) workload for each cluster concurrently, and periodically fuses the client...
-
08
TrafficFab: An Autonomic Edge-Cloud Testbed Fabric forAI-Driven Traffic Management
Mayank Arya, Pranjal Naman, Priyanshu Pansari, Roopkatha Banerjee, Daksh Mehta, Manjil Nepal, Akash Sharma, Yogesh Simmhan
cs.DC
Traffic management in emerging megacities requires real-time analytics over thousands of CCTV video streams under latency, bandwidth, compute and energy constraints. We present TrafficFab, an autonomic edge--cloud testbed for AI-driven traffic management, designed to validate a representative slice of a megacity deployment. TrafficFab combines RTSP stream emulation, heterogeneous edge inference using DNNs, cloud-based nowcasting and...
This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.