cs.DC · 2026-07-15 · No. 54
Distributed, Parallel, and Cluster Computing, 2026-07-15.
10 new papers in cs.DC. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
10 entries-
01
Proceedings of HLPP 2026: 19th International Symposium on High-Level Parallel Programming and Applications
Chong Li, Corinne Ancourt, Gaétan Hains
cs.DC · cs.PL
This volume contains the ten peer-reviewed papers presented at HLPP 2026, the 19th International Symposium on High-Level Parallel Programming and Applications, held on 9-10 July 2026 at the Institut Henri Poincare in Paris, France. The symposium covers high-level approaches to parallel programming: programming models, languages, libraries, algorithmic skeletons, compilers, and runtime systems for multi-core, GPU, and distributed platforms....
-
02
HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference
Gregory Hyegang Jun, Wesley Pang, Eddie Richter, Mehdi Saeedi, Aporva Amarnath, Pallavi Ferrao, Deming Chen
cs.DC · cs.AR
Modern edge system-on-chips (SoCs) combine CPUs, integrated GPUs (iGPUs), and neural processing units (NPUs), yet existing LLM runtimes typically make coarse device-level decisions or optimize operators in isolation. As a result, they underutilize heterogeneous resources, particularly on unified-memory platforms where performance depends on both device placement and task-graph coordination. We present HeteroMosaic, a heterogeneity-first...
-
03
Scaling Synthetic-Image Pre-Training for Federated Fine-Tuning of Large Vision Models
Qianpiao Ma, Xiaozhu Song, Junlong Zhou, Yue Zeng, Jianchun Liu, Huaqing Tu
cs.DC
Federated fine-tuning (FedFT) enables adapting pre-trained large vision models (LVMs) on distributed, privacy-sensitive devices, while its practical deployment is hindered by three critical challenges: resource constraints, system heterogeneity, and non-IID data. While prior studies partially address these issues, e.g., by pre-training initial models on synthetic images to mitigate the adverse effects of non-IID data, or leveraging...
-
04
Parallelizing Legacy Mesh Generation Software: Lessons Learned from a Pseudo-Constrained Parallel Data Refinement Approach for Advancing Front Local Reconnection
Kevin Garner, David Marcum, Nikos Chrisochoides
cs.DC
This paper presents lessons learned from parallelizing the legacy software known as Advancing Front Local Reconnection (AFLR) as a black box. The parallel procedure utilizes (i) a data decomposition scheme where each subdomain is refined in parallel using the sequential mesh generation code and (ii) a runtime system for work-load balancing. Results on the mesh refinement operation show that the parallel method's stability (output mesh...
-
05
Profiling and Scheduling Complex O-RAN Applications Across the 5G Edge and Cloud
Yoonjae Hwang, Bhaskar Krishnamachari
cs.DC
The O-RAN paradigm decomposes intelligent RAN control into pipelines of interdependent AI/ML functions, including traffic prediction, signal quality estimation, and slice scheduling, that must execute across a dispersed continuum of far-edge, near-edge, and cloud resources under heterogeneous latency and bandwidth constraints. Despite the natural expression of these pipelines as Directed Acyclic Graphs (DAGs), no integrated methodology exists...
-
06
Overcoming Orchestration Bottlenecks at Exascale: A Decentralized, Policy-Driven Approach for Sim-AI Ensembles
Harikrishna Tummalapalli, Christine M. Simpson, Riccardo Balin, Vitali A. Morozov, Thang D. Pham, Murat Keceli,...
cs.DC · cs.CE
Scientific computing is increasingly shifting from monolithic applications to coupled simulation-AI workflows composed of highly heterogeneous tasks with diverse hardware, scale, and runtime requirements. As these workflows scale to leadership-class systems, the resulting extreme ensemble sizes and task variability can create orchestration bottlenecks. System-level schedulers are often configured for limited throughput, while workflow tools...
-
07
Decentralized Gradient Descent: Bottleneck Regimes and Budget Complexity
Nicolò Michelusi
cs.DC · cs.LG · eess.SP
Decentralized gradient descent (DGD) is widely used for solving distributed optimization problems over networks of agents. While its convergence properties are well understood, less is known about the communication and computation resources required to attain a prescribed accuracy. In this paper, we study DGD from a resource-aware perspective and characterize the communication-computation budget required to attain a target error level. We...
-
08
FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving
Yaqi Qiao, Ping He, Songrun Xie, Ayush Barik, Chensong Zhang, Zhengzhong Tu, Fan Lai
cs.DC · cs.LG
Diffusion models have become the central backbone for modern image, video, and audio generation, but their efficient service remains a challenge. Unlike autoregressive decoding, diffusion inference repeatedly updates high-dimensional spatial or temporal latents over many denoising steps. This all-region execution pattern makes generation latency high and limits serving throughput. Existing multi-GPU parallelization methods can reduce per-step...
-
09
Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap
Rafael Ferreira da Silva, Milad Abolhasani, Peter Beaucage, Laura Biven, Michael Bussmann, Kyle Chard, Ryan Coffee,...
cs.DC · cs.AI
One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around five critical dimensions. The field has since moved faster than anticipated. Multi-agent systems have produced experimentally validated hypotheses, self-driving laboratories have grown more interoperable and orchestrated, reasoning-trained and domain foundation models have raised the capability...
-
10
HARP-ME: Closure-Driven Exact Induced Motif Enumeration on GPUs
Ashwina Kumar, Rupesh Nasre
cs.DC
Exact induced motif enumeration is a fundamental operation in graph mining, but it remains challenging on GPUs because candidate expansion is irregular, repeated set intersections dominate execution, and induced counting must consider both the presence and absence of edges. We present HARP-ME, which stands for Hierarchical Anchor-Reuse Partitioned Motif Enumeration. It is a GPU framework for the exact enumeration of connected induced...
This edition is part of The Daily Abstract — cs.DC archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.