Skip to content

RCCL

RCCL (pronounced "Rickle") is a stand-alone library of standard collective communication routines for GPUs, implementing all-reduce, all-gather, reduce, broadcast, reduce-scatter, gather, scatter, and all-to-all. There is also initial support for direct GPU-to-GPU send and receive operations. It has been optimized to achieve high bandwidth on platforms using PCIe, xGMI as well as networking using InfiniBand Verbs or TCP/IP sockets. RCCL supports an arbitrary number of GPUs installed in a single node or multiple nodes, and can be used in either single- or multi-process (e.g., MPI) applications.

homepage: https://github.com/ROCm/RCCL

Available installations

RCCL version Supported CPU targets Supported GPU targets EESSI version Module
2.22.3 generic: x86_64
Arm:
AMD: zen2, zen3, zen4, zen5
Intel: haswell, skylake_avx512, sapphirerapids, icelake, cascadelake
AMD: gfx1030, gfx1100, gfx1101, gfx1200, gfx1201, gfx908, gfx90a, gfx942
2025.06 RCCL/2.22.3-rocm-compilers-19.0.0-ROCm-6.4.1