UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
See the code
Blog | Join Slack | Twitter/X | Roadmap | Quick Start | Open Letter
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., IBGDA), with two key focuses:
An UCCL overview can be found in this slide deck with the following components:
UCCL-collective (UCCL-Tran) serves as a drop-in replacement for NCCL/RCCL (e.g., requiring no changes to application code), and significantly outperforms them in both latency and throughput across various settings.
g4dn.8xlarge instances with 1x50G ENA NICs and 1xT4 GPUs within the same cluster placement group, UCCL-collective outperforms NCCL by up to 3.7x for AllReduce:
UCCL-P2P provides NIXL-style initiator-target transfer APIs. UCCL-P2P is purposely designed for the next-gen 800Gbps NICs with efficient multi-threaded transfer engines.
UCCL-EP allows running DeepEP atop of heterogeneous hardware platforms, including AMD and Nvidia GPUs, and any RDMA NICs such as AWS EFA NICs and Broadcom NICs, while achieving IBGDA-level performance.
UCCL has been adopted as part of the AMD TheRock ecosystem.
More UCCL features are under development in this repo, currently including:
The easiest way to use UCCL is to first build based on your platform. The build script will automatically detect the py_version of your current environment. If you need to compile UCCL for a specific python version, please specify the py_version, such as 3.10.
git clone https://github.com/uccl-project/uccl.git && cd uccl
# Eg, bash build.sh cu12 ep --install
bash build.sh [cu12|cu13|roc7|roc6|therock] [all|ccl_rdma|ccl_efa|p2p|ep] \
[py_version] [rocm_index_url] --install
Note:
- By default,
build.sh cu12targets CUDA 12.8 andbuild.sh roc7targets ROCm 7.1, but you can also specifycu13|roc6to target CUDA 13.0 or ROCm 6.4.- UCCL uses nanobind for C++/Python bindings. On Python 3.12+, wheels are tagged
cp312-abi3(stable ABI, one wheel for all 3.12+ interpreters); on older Pythons, wheels are CPython-version-specific.- When building for ROCm with python packaging through TheRock, please specify your ROCm index url; the default is
https://rocm.prereleases.amd.com/whl/gfx94X-dcgpuand it may not be what you want. When installing UCCL wheels for TheRock, please provide pip with the index url and add the optional extra[rocm]to the wheel, e.g.,pip install --extra-index-url https://rocm.prereleases.amd.com/whl/gfx94X-dcgpu wheelhouse-therock/uccl-0.0.1.post4-py3-none-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl[rocm].
Then, when running your PyTorch applications, set the environment variable accordingly:
# NCCL over IB/RoCE on x86 or GH200 ARM hosts
NCCL_NET_PLUGIN=`python -c "import uccl; print(uccl.nccl_plugin_path())"`
# RCCL over IB/RoCE on x86 hosts
NCCL_NET_PLUGIN=`python -c "import uccl; print(uccl.rccl_plugin_path())"`
# NCCL over AWS EFA NICs (p4d and p4de only)
LD_PRELOAD=`python -c "import uccl; print(uccl.efa_nccl_path())"`
NCCL_NET_PLUGIN=`python -c "import uccl; print(uccl.efa_plugin_path())"`
Now, you can just run your PyTorch applications and enjoy UCCL performance benefits!
First clone the UCCL repo and init submodules:
git clone https://github.com/uccl-project/uccl.git
export UCCL_HOME=$(pwd)/uccl
To build UCCL for development, you need to install some common dependencies:
# Note if you are using docker+wheel build, there is no need to install the following dependencies.
sudo apt update
sudo apt install linux-tools-$(uname -r) clang llvm cmake m4 build-essential \
net-tools libgtest-dev libgflags-dev \
libelf-dev libpcap-dev libc6-dev-i386 libpci-dev \
libopenmpi-dev libibverbs-dev clang-format -y
# Install and activate Miniconda (you can choose any recent versions)
wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash ./Miniconda3-latest-Linux-x86_64.sh -b
source ~/miniconda3/bin/activate
source ~/.bashrc # or .zshrc and others
conda init
# Install python ssh lib and more
pip install paramiko intervaltree pybind11 nanobind
# Upgrade conda glic to modern ones
conda install -c conda-forge "libstdcxx-ng>=12" "libgcc-ng>=12"
Alternatively, use uv for a faster, conda-free dev setup:
source scripts/bootstrap.sh
This single command installs uv if missing, creates a .venv virtualenv with Python 3.12, pins all non-CUDA dev tools (black, clang-format, pytest, paramiko, etc.) from uv.lock via uv sync --group dev, and then runs ep/install_deps.sh to install CUDA-specific packages (torch, etc.) with automatic hardware detection.
For quick installation with docker, you can directly dive into:
UCCL-Collective RDMA: Collectives for Nvidia/AMD GPUs + IB/RoCE RDMA NICs (currently support Nvidia and Broadcom NICs)
UCCL-Collective EFA: Collectives for AWS EFA NIC (currently support p4d.24xlarge)
On p5/p5e/p5en/p6, the offical aws-ofi-nccl NCCL plugin with proper env variables already makes NCCL perform excellent
UCCL-Collective AFXDP: Collectives for Non-RDMA NICs (currently support AWS ENA NICs and IBM VirtIO NICs)
UCCL-P2P: P2P for RDMA NICs and GPU IPCs (currently support Nvidia/AMD GPUs and Nvidia/Broadcom NICs)
UCCL-EP: EP for MoE training and inference with DeepEP-compatible APIs (currently support Nvidia/AMD GPUs and Nvidia/Broadcom/EFA NICs)
The code in this repository is mostly described in the papers below. Please consider citing this work if you find the repository helpful.
@article{uccl_tran,
title={UCCL-Tran: An Extensible Software Transport Layer for GPU Networking},
author={Zhou, Yang and Chen, Zhongjie and Mao, Ziming and Lao, ChonLam and Yang, Shuo and Kannan, Pravein Govindan and Gao, Jiaqi and Zhao, Yilong and Wu, Yongji and You, Kaichao and Ren, Fengyuan and Xu, Zhiying and Raiciu, Costin and Stoica, Ion},
journal={USENIX OSDI},
year={2026}
}
@article{uccl_ep,
title={UCCL-EP: Portable Expert-Parallel Communication},
author={Mao, Ziming and Zhang, Yihan and Cui, Chihan and You, Kaichao and Chen, Zhongjie and Xu, Zhiying and Shenker, Scott and Raiciu, Costin and Zhou, Yang and Stoica, Ion},
journal={USENIX OSDI},
year={2026}
}
UCCL is being actively developed at UC Berkeley Sky Computing Lab and UC Davis ArtSy lab. We enthusiastically welcome open-source developers joining us!
UCCL is generously supported by (in alphabetical order): AMD, AWS, Broadcom, CloudLab, Google Cloud, IBM, Lambda, Mibura.
3,104 followers · starred Aug 2025
534 followers · starred Jun 2025
325 followers · starred Jun 2025
820 followers · starred Apr 2026
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
See the code
Blog | Join Slack | Twitter/X | Roadmap | Quick Start | Open Letter
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., IBGDA), with two key focuses:
An UCCL overview can be found in this slide deck with the following components:
UCCL-collective (UCCL-Tran) serves as a drop-in replacement for NCCL/RCCL (e.g., requiring no changes to application code), and significantly outperforms them in both latency and throughput across various settings.
g4dn.8xlarge instances with 1x50G ENA NICs and 1xT4 GPUs within the same cluster placement group, UCCL-collective outperforms NCCL by up to 3.7x for AllReduce:
UCCL-P2P provides NIXL-style initiator-target transfer APIs. UCCL-P2P is purposely designed for the next-gen 800Gbps NICs with efficient multi-threaded transfer engines.
UCCL-EP allows running DeepEP atop of heterogeneous hardware platforms, including AMD and Nvidia GPUs, and any RDMA NICs such as AWS EFA NICs and Broadcom NICs, while achieving IBGDA-level performance.
UCCL has been adopted as part of the AMD TheRock ecosystem.
More UCCL features are under development in this repo, currently including:
The easiest way to use UCCL is to first build based on your platform. The build script will automatically detect the py_version of your current environment. If you need to compile UCCL for a specific python version, please specify the py_version, such as 3.10.
git clone https://github.com/uccl-project/uccl.git && cd uccl
# Eg, bash build.sh cu12 ep --install
bash build.sh [cu12|cu13|roc7|roc6|therock] [all|ccl_rdma|ccl_efa|p2p|ep] \
[py_version] [rocm_index_url] --install
Note:
- By default,
build.sh cu12targets CUDA 12.8 andbuild.sh roc7targets ROCm 7.1, but you can also specifycu13|roc6to target CUDA 13.0 or ROCm 6.4.- UCCL uses nanobind for C++/Python bindings. On Python 3.12+, wheels are tagged
cp312-abi3(stable ABI, one wheel for all 3.12+ interpreters); on older Pythons, wheels are CPython-version-specific.- When building for ROCm with python packaging through TheRock, please specify your ROCm index url; the default is
https://rocm.prereleases.amd.com/whl/gfx94X-dcgpuand it may not be what you want. When installing UCCL wheels for TheRock, please provide pip with the index url and add the optional extra[rocm]to the wheel, e.g.,pip install --extra-index-url https://rocm.prereleases.amd.com/whl/gfx94X-dcgpu wheelhouse-therock/uccl-0.0.1.post4-py3-none-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl[rocm].
Then, when running your PyTorch applications, set the environment variable accordingly:
# NCCL over IB/RoCE on x86 or GH200 ARM hosts
NCCL_NET_PLUGIN=`python -c "import uccl; print(uccl.nccl_plugin_path())"`
# RCCL over IB/RoCE on x86 hosts
NCCL_NET_PLUGIN=`python -c "import uccl; print(uccl.rccl_plugin_path())"`
# NCCL over AWS EFA NICs (p4d and p4de only)
LD_PRELOAD=`python -c "import uccl; print(uccl.efa_nccl_path())"`
NCCL_NET_PLUGIN=`python -c "import uccl; print(uccl.efa_plugin_path())"`
Now, you can just run your PyTorch applications and enjoy UCCL performance benefits!
First clone the UCCL repo and init submodules:
git clone https://github.com/uccl-project/uccl.git
export UCCL_HOME=$(pwd)/uccl
To build UCCL for development, you need to install some common dependencies:
# Note if you are using docker+wheel build, there is no need to install the following dependencies.
sudo apt update
sudo apt install linux-tools-$(uname -r) clang llvm cmake m4 build-essential \
net-tools libgtest-dev libgflags-dev \
libelf-dev libpcap-dev libc6-dev-i386 libpci-dev \
libopenmpi-dev libibverbs-dev clang-format -y
# Install and activate Miniconda (you can choose any recent versions)
wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash ./Miniconda3-latest-Linux-x86_64.sh -b
source ~/miniconda3/bin/activate
source ~/.bashrc # or .zshrc and others
conda init
# Install python ssh lib and more
pip install paramiko intervaltree pybind11 nanobind
# Upgrade conda glic to modern ones
conda install -c conda-forge "libstdcxx-ng>=12" "libgcc-ng>=12"
Alternatively, use uv for a faster, conda-free dev setup:
source scripts/bootstrap.sh
This single command installs uv if missing, creates a .venv virtualenv with Python 3.12, pins all non-CUDA dev tools (black, clang-format, pytest, paramiko, etc.) from uv.lock via uv sync --group dev, and then runs ep/install_deps.sh to install CUDA-specific packages (torch, etc.) with automatic hardware detection.
For quick installation with docker, you can directly dive into:
UCCL-Collective RDMA: Collectives for Nvidia/AMD GPUs + IB/RoCE RDMA NICs (currently support Nvidia and Broadcom NICs)
UCCL-Collective EFA: Collectives for AWS EFA NIC (currently support p4d.24xlarge)
On p5/p5e/p5en/p6, the offical aws-ofi-nccl NCCL plugin with proper env variables already makes NCCL perform excellent
UCCL-Collective AFXDP: Collectives for Non-RDMA NICs (currently support AWS ENA NICs and IBM VirtIO NICs)
UCCL-P2P: P2P for RDMA NICs and GPU IPCs (currently support Nvidia/AMD GPUs and Nvidia/Broadcom NICs)
UCCL-EP: EP for MoE training and inference with DeepEP-compatible APIs (currently support Nvidia/AMD GPUs and Nvidia/Broadcom/EFA NICs)
The code in this repository is mostly described in the papers below. Please consider citing this work if you find the repository helpful.
@article{uccl_tran,
title={UCCL-Tran: An Extensible Software Transport Layer for GPU Networking},
author={Zhou, Yang and Chen, Zhongjie and Mao, Ziming and Lao, ChonLam and Yang, Shuo and Kannan, Pravein Govindan and Gao, Jiaqi and Zhao, Yilong and Wu, Yongji and You, Kaichao and Ren, Fengyuan and Xu, Zhiying and Raiciu, Costin and Stoica, Ion},
journal={USENIX OSDI},
year={2026}
}
@article{uccl_ep,
title={UCCL-EP: Portable Expert-Parallel Communication},
author={Mao, Ziming and Zhang, Yihan and Cui, Chihan and You, Kaichao and Chen, Zhongjie and Xu, Zhiying and Shenker, Scott and Raiciu, Costin and Zhou, Yang and Stoica, Ion},
journal={USENIX OSDI},
year={2026}
}
UCCL is being actively developed at UC Berkeley Sky Computing Lab and UC Davis ArtSy lab. We enthusiastically welcome open-source developers joining us!
UCCL is generously supported by (in alphabetical order): AMD, AWS, Broadcom, CloudLab, Google Cloud, IBM, Lambda, Mibura.
3,104 followers · starred Aug 2025
534 followers · starred Jun 2025
325 followers · starred Jun 2025
820 followers · starred Apr 2026