Rewsr/unawesome-ai-fabric-engineering

Curated resources on RDMA, collectives, and AI cluster fabrics for GPU performance engineers

Python

30

0 commits

updated Oct 8, 2026

See the code

See what people are saying

README

Unawesome AI Fabric Engineering

Networking that moves data between GPUs in AI clusters: RDMA, GPU-to-NIC data paths, collectives, transports, and cluster fabrics.

Written for GPU performance engineers moving into the network.

It assumes you already know GPU kernels, profiling, inference engines, and the basics of distributed inference. The list is ordered from one NIC to one GPU-NIC path, collectives, inference transfer, transports, and whole fabrics. Read Start here first. After that, use it as a reference.

Every resource assumes real NICs, switches, and GPUs. See Footnotes.

Contents

Start here: the minimum mental model

Read these in order.

1. RDMA fundamentals

Verbs and memory registration

NIC behavior

Measurement

2. The GPU-to-NIC data path

GPUDirect and PCIe

  • Linux dma-buf - The kernel buffer-sharing mechanism used to register GPU memory without a peer-memory module.
  • PCI peer-to-peer DMA - Kernel support and constraints for device-to-device DMA across PCIe.
  • NCCL troubleshooting - PCI ACS, IOMMU, and platform settings that silently break or slow GPU-NIC paths.
  • NCCL environment variables - Topology files, NIC selection, GPUDirect levels, and InfiniBand settings.

GPU-initiated networking

  • NVSHMEM and GPUDirect Async - InfiniBand GPUDirect Async: GPU threads post RDMA work to the NIC directly.
  • NVSHMEM - The GPU-initiated communication library behind it.
  • DeepEP - Expert dispatch and combine over NVLink and RDMA, including GPU-initiated low-latency kernels.
  • DOCA GPUNetIO - GPU-controlled packet I/O on ConnectX NICs.
  • Device Memory TCP - TCP that receives payloads directly into device memory.

3. Collectives

Algorithms

Libraries

4. Point-to-point transfer for inference

5. Transport and congestion control

  • DCQCN - ECN-based congestion control for RoCEv2.
  • TIMELY - RTT-based congestion control in the datacenter.
  • HPCC - Congestion control from in-network telemetry.
  • Swift - Delay-based congestion control in production.
  • Revisiting Network Support for RDMA - RDMA without a lossless fabric.
  • SRD - AWS's multipath, out-of-order reliable datagram transport behind EFA.
  • Falcon - Google's hardware-offloaded reliable transport, published through OCP.
  • Ultra Ethernet 1.0.3 Specification - The scale-out transport specification.

6. Fabric design for AI clusters

7. Scale-up fabrics

8. Cloud fabrics

9. Operating fabrics at scale

Frontier

Verified on 2026-10-05. Kept separate from the core list because the evidence changes quickly.

Watchlist

  • Ultra Ethernet NICs and switches in production, pending measured deployments.
  • Adaptive routing and packet spraying on Ethernet AI fabrics, pending reproducible measurements.
  • Co-packaged optics switches in deployed GPU clusters.
  • UALink and Ethernet-based scale-up silicon, pending shipped systems.
  • GPU-initiated networking on non-NVIDIA NICs.
  • KV-cache transfer across heterogeneous GPU fleets with measured tail latency.

Footnotes

Hardware policy

Every resource assumes real NICs, switches, and GPUs. Software RDMA emulation and network simulators are excluded: their numbers do not transfer, and they cannot exercise GPUDirect, NIC offloads, congestion control, or adaptive routing.

Source policy

A core source must be one of the following:

  • the paper that introduced the mechanism;
  • the specification or official documentation that defines it;
  • the repository that implements it;
  • an operator or implementer report with hardware, measurements, and enough detail to reproduce the result.

Performance claims need the NIC, switch, topology, message sizes, software versions, and baseline. Otherwise the number is omitted.

Contributing

See CONTRIBUTING.md before proposing a resource. python3 scripts/check_links.py checks every link.

gpu
nccl
networking
rdma

Rewsr/unawesome-ai-fabric-engineering

Curated resources on RDMA, collectives, and AI cluster fabrics for GPU performance engineers

Python

30

0 commits

updated Oct 8, 2026

See the code

See what people are saying

README

Unawesome AI Fabric Engineering

Networking that moves data between GPUs in AI clusters: RDMA, GPU-to-NIC data paths, collectives, transports, and cluster fabrics.

Written for GPU performance engineers moving into the network.

It assumes you already know GPU kernels, profiling, inference engines, and the basics of distributed inference. The list is ordered from one NIC to one GPU-NIC path, collectives, inference transfer, transports, and whole fabrics. Read Start here first. After that, use it as a reference.

Every resource assumes real NICs, switches, and GPUs. See Footnotes.

Contents

Start here: the minimum mental model

Read these in order.

1. RDMA fundamentals

Verbs and memory registration

NIC behavior

Measurement

2. The GPU-to-NIC data path

GPUDirect and PCIe

  • Linux dma-buf - The kernel buffer-sharing mechanism used to register GPU memory without a peer-memory module.
  • PCI peer-to-peer DMA - Kernel support and constraints for device-to-device DMA across PCIe.
  • NCCL troubleshooting - PCI ACS, IOMMU, and platform settings that silently break or slow GPU-NIC paths.
  • NCCL environment variables - Topology files, NIC selection, GPUDirect levels, and InfiniBand settings.

GPU-initiated networking

  • NVSHMEM and GPUDirect Async - InfiniBand GPUDirect Async: GPU threads post RDMA work to the NIC directly.
  • NVSHMEM - The GPU-initiated communication library behind it.
  • DeepEP - Expert dispatch and combine over NVLink and RDMA, including GPU-initiated low-latency kernels.
  • DOCA GPUNetIO - GPU-controlled packet I/O on ConnectX NICs.
  • Device Memory TCP - TCP that receives payloads directly into device memory.

3. Collectives

Algorithms

Libraries

4. Point-to-point transfer for inference

5. Transport and congestion control

  • DCQCN - ECN-based congestion control for RoCEv2.
  • TIMELY - RTT-based congestion control in the datacenter.
  • HPCC - Congestion control from in-network telemetry.
  • Swift - Delay-based congestion control in production.
  • Revisiting Network Support for RDMA - RDMA without a lossless fabric.
  • SRD - AWS's multipath, out-of-order reliable datagram transport behind EFA.
  • Falcon - Google's hardware-offloaded reliable transport, published through OCP.
  • Ultra Ethernet 1.0.3 Specification - The scale-out transport specification.

6. Fabric design for AI clusters

7. Scale-up fabrics

8. Cloud fabrics

9. Operating fabrics at scale

Frontier

Verified on 2026-10-05. Kept separate from the core list because the evidence changes quickly.

Watchlist

  • Ultra Ethernet NICs and switches in production, pending measured deployments.
  • Adaptive routing and packet spraying on Ethernet AI fabrics, pending reproducible measurements.
  • Co-packaged optics switches in deployed GPU clusters.
  • UALink and Ethernet-based scale-up silicon, pending shipped systems.
  • GPU-initiated networking on non-NVIDIA NICs.
  • KV-cache transfer across heterogeneous GPU fleets with measured tail latency.

Footnotes

Hardware policy

Every resource assumes real NICs, switches, and GPUs. Software RDMA emulation and network simulators are excluded: their numbers do not transfer, and they cannot exercise GPUDirect, NIC offloads, congestion control, or adaptive routing.

Source policy

A core source must be one of the following:

  • the paper that introduced the mechanism;
  • the specification or official documentation that defines it;
  • the repository that implements it;
  • an operator or implementer report with hardware, measurements, and enough detail to reproduce the result.

Performance claims need the NIC, switch, topology, message sizes, software versions, and baseline. Otherwise the number is omitted.

Contributing

See CONTRIBUTING.md before proposing a resource. python3 scripts/check_links.py checks every link.

gpu
nccl
networking
rdma