Any-Winter-4079/DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training

Python

1

3 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

DDR4/PCIe4 vs DDR5/PCIe5 for LLMs- I benchmarked them for pre-training. What are your thoughts? (r/LocalLLaMA)

Hello everyone. I've recently ran some experiments comparing DDR4/PCIe4 and DDR5/PCIe5 for AI workstations on a pre-training run, and would like to hear yours thoughts. First of all, and as a summary of my results ( code here:…

12

Oct 1, 2026

README

DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training

This repository compares training on DDR4/PCIe 4 and DDR5/PCIe 5 workstations using RTX PRO 6000 WS GPUs. train_DDP.py supports one or two GPUs; train_PP.py uses two pipeline stages. The plot reports cumulative training-step time to validation loss ≤ 3.28, excluding validation and warm-up.

workstation-training-comparisons

Running the benchmarks

  1. Start up a container with the given image (PyTorch 2.10.0 + CUDA 12.8):

https://cloud.vast.ai/?ref_id=195884&creator_id=195884&name=anywinter4079%2Fpytorch%3A2.10.0-cu128

  1. Clone and cd into this repository:
git clone https://github.com/Any-Winter-4079/DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training.git && \
cd DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training
  1. Download the dataset:
python data/fineweb-npy.py
  1. Run, e.g., for pipeline parallelism:
torchrun --standalone --nproc_per_node=2 train_PP.py

Any-Winter-4079/DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training

Python

1

3 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

DDR4/PCIe4 vs DDR5/PCIe5 for LLMs- I benchmarked them for pre-training. What are your thoughts? (r/LocalLLaMA)

Hello everyone. I've recently ran some experiments comparing DDR4/PCIe4 and DDR5/PCIe5 for AI workstations on a pre-training run, and would like to hear yours thoughts. First of all, and as a summary of my results ( code here:…

12

Oct 1, 2026

README

DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training

This repository compares training on DDR4/PCIe 4 and DDR5/PCIe 5 workstations using RTX PRO 6000 WS GPUs. train_DDP.py supports one or two GPUs; train_PP.py uses two pipeline stages. The plot reports cumulative training-step time to validation loss ≤ 3.28, excluding validation and warm-up.

workstation-training-comparisons

Running the benchmarks

  1. Start up a container with the given image (PyTorch 2.10.0 + CUDA 12.8):

https://cloud.vast.ai/?ref_id=195884&creator_id=195884&name=anywinter4079%2Fpytorch%3A2.10.0-cu128

  1. Clone and cd into this repository:
git clone https://github.com/Any-Winter-4079/DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training.git && \
cd DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training
  1. Download the dataset:
python data/fineweb-npy.py
  1. Run, e.g., for pipeline parallelism:
torchrun --standalone --nproc_per_node=2 train_PP.py

Languages

Python

100.0%