12 papers selected.
Enhancing OmpSs-2 Suspendable Tasks by Combining Operating System and User-Level Threads with C++ Coroutines
Arnau Cinca, Aleix Roca, Kevin Sala, Raúl Peñacoba Veigas et al.
TL;DR — Extends the OmpSs-2 task-based runtime with C++ coroutines to implement suspendable tasks that can yield while blocked on I/O or communication without stalling the OS thread.
Why notable — Suspendable tasks are a key missing primitive for overlapping computation and communication in task-graph runtimes; the hybrid OS/user-level thread design avoids the overhead of full context switches while remaining portable, with broad implications for OpenMP-style programming on modern heterogeneous nodes.
PISA: An Adversarial Approach to Comparing Task Graph Scheduling Algorithms
Jared Coleman, Bhaskar Krishnamachari
TL;DR — Introduces an adversarial instance-generation framework that automatically synthesizes task graphs that expose worst-case performance gaps between competing scheduling heuristics.
Why notable — Benchmark-driven comparison of DAG schedulers is notoriously biased toward whoever designed the benchmark; PISA’s adversarial synthesis provides a principled, algorithm-agnostic methodology that could become a standard evaluation tool for the task-scheduling community.
Parallel Scheduling of Task Graphs with Minimal Memory Requirements
Pascal Fradet, Alain Girault, Alexandre Honorat
TL;DR — Proves tight bounds and provides scheduling algorithms for task graphs that minimize peak memory usage while preserving parallelism, targeting memory-constrained accelerators.
Why notable — Peak memory is increasingly the binding constraint on accelerators with fixed HBM capacity; deriving schedules that are jointly memory-optimal and parallel addresses a theoretically hard trade-off that directly impacts large-scale neural network and sparse-solver pipelines.
FATHOM: Fast Attention Through Optimizing Memory
Elliott Binder, Arvind Sudarsanam, Ravi Sunkavalli, Tze Meng Low
TL;DR — Redesigns the attention kernel memory access pattern to maximize reuse across the Q, K, and V tiles, achieving significant throughput gains over FlashAttention on modern GPUs.
Why notable — Attention is the dominant compute bottleneck in transformer inference and training; FATHOM’s memory-centric reformulation improves on the widely deployed FlashAttention baseline and demonstrates that analytical cache-tile reasoning still yields practical gains at scale.
Fast and Effective Lossy Compression on GPUs and CPUs with Guaranteed Error Bounds
Alex Fallin, Noushin Azami, Sheng Di, Franck Cappello et al.
TL;DR — Presents a portable error-bounded lossy compressor that runs natively on both GPU and CPU, delivering competitive compression ratios with strict point-wise error guarantees and low latency.
Why notable — Scientific simulations increasingly need in-situ compression on the same GPU that runs the simulation; providing tight error bounds alongside GPU portability bridges a critical gap between scientific fidelity requirements and storage bandwidth constraints at exascale.
Phase-Based Frequency Scaling for Energy-Efficient Heterogeneous Computing
Lorenzo Carpentieri, Antonio De Caro, Majid Salimi Beni, Kaijie Fan et al.
TL;DR — Dynamically scales CPU and GPU frequencies based on detected application phases to reduce energy consumption while preserving performance on heterogeneous nodes.
Why notable — Energy efficiency is a first-class constraint at exascale; phase-aware DVFS that jointly controls both CPU and GPU frequencies avoids the over-provisioning inherent in static policies, with measured gains directly applicable to production HPC clusters.
HiCCL: A Hierarchical Collective Communication Library
Mert Hidayetoglu, Simon Garcia de Gonzalo, Elliott Slaughter, Pinku Surana et al.
TL;DR — Implements a multi-level collective communication library that decomposes all-reduce and other collectives into hierarchical sub-collectives matched to node-local, intra-rack, and inter-rack bandwidth tiers.
Why notable — Flat NCCL/MPI collectives ignore the hierarchical bandwidth structure of modern GPU clusters; HiCCL’s topology-aware decomposition achieves substantial throughput improvements on large distributed training jobs and is designed to be backend-agnostic.
Unified Designs of Multi-Rail-Aware MPI Allreduce and Alltoall Operations Across Diverse GPU and Interconnect Systems
Chen-Chun Chen, Jinghan Yao, Lang Xu, Hari Subramoni et al.
TL;DR — Develops unified multi-rail-aware algorithms for MPI Allreduce and Alltoall that exploit all available NIC ports simultaneously across heterogeneous interconnect systems.
Why notable — Multi-rail configurations are increasingly common in HPC clusters yet most MPI libraries treat them as a single logical link; this work systematically closes the performance gap and the unified design enables deployment without per-system hand-tuning.
CoRD: Converged RDMA Dataplane
Maksym Planeta, Jan Bierbaum, Michael Roitzsch, Hermann Härtig
TL;DR — Proposes a converged RDMA dataplane that unifies one-sided and two-sided communication semantics over a single kernel-bypass path, reducing software overhead and improving scalability.
Why notable — RDMA stacks remain split between one-sided verbs and two-sided message passing, forcing application writers to choose and preventing optimal use of NIC offload capabilities; CoRD’s converged abstraction enables a new class of communication patterns at near-wire speed.
FlexRLHF: A Flexible Placement and Parallelism Framework for Efficient RLHF Training
Youshao Xiao, Zhenglei Zhou, Fagui Mao, Weichang Wu et al.
TL;DR — Introduces a placement and parallelism co-optimization framework for Reinforcement Learning from Human Feedback (RLHF) training that jointly schedules the actor, critic, and reward models to maximize GPU utilization.
Why notable — RLHF is the dominant fine-tuning paradigm for large language models but its multi-model, heterogeneous-workload structure makes naive data/model parallelism highly inefficient; FlexRLHF’s co-placement approach delivers meaningful throughput gains and establishes a design template for future RLHF infrastructure.
GuardianOMP: A Framework for Highly Productive Fault Tolerance Via OpenMP Task-Level Replication
Adrian Munera, Eduardo Quiñones, Sara Royuela
TL;DR — Adds transparent task-level redundant execution to OpenMP applications, enabling automatic detection and recovery from silent data corruption without application source changes.
Why notable — Silent data corruption is an escalating concern as DRAM and compute elements scale into billions of transistors; GuardianOMP’s integration at the OpenMP runtime level makes resilience accessible to the broad HPC community that already uses OpenMP without requiring manual checkpoint/restart logic.
Tera-Scale Multilevel Graph Partitioning
Daniel Salwasser, Daniel Seemaier, Lars Gottesbüren, Peter Sanders 0001
TL;DR — Scales multilevel graph partitioning to trillion-edge graphs through a distributed coarsening and refinement pipeline that maintains partition quality competitive with state-of-the-art tools on billion-edge benchmarks.
Why notable — Graph partitioning is a prerequisite for almost every distributed graph workload; reaching the tera-scale regime with near-optimal quality is a significant algorithmic and engineering milestone that directly enables graph-parallel simulation at the largest current HPC scales.