Python
1
3 commits
updated Oct 1, 2026
This repository compares training on DDR4/PCIe 4 and DDR5/PCIe 5 workstations using RTX PRO 6000 WS GPUs. train_DDP.py supports one or two GPUs; train_PP.py uses two pipeline stages. The plot reports cumulative training-step time to validation loss ≤ 3.28, excluding validation and warm-up.
https://cloud.vast.ai/?ref_id=195884&creator_id=195884&name=anywinter4079%2Fpytorch%3A2.10.0-cu128
cd into this repository:git clone https://github.com/Any-Winter-4079/DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training.git && \
cd DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training
python data/fineweb-npy.py
torchrun --standalone --nproc_per_node=2 train_PP.py
Python
100.0%
Python
1
3 commits
updated Oct 1, 2026
This repository compares training on DDR4/PCIe 4 and DDR5/PCIe 5 workstations using RTX PRO 6000 WS GPUs. train_DDP.py supports one or two GPUs; train_PP.py uses two pipeline stages. The plot reports cumulative training-step time to validation loss ≤ 3.28, excluding validation and warm-up.
https://cloud.vast.ai/?ref_id=195884&creator_id=195884&name=anywinter4079%2Fpytorch%3A2.10.0-cu128
cd into this repository:git clone https://github.com/Any-Winter-4079/DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training.git && \
cd DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training
python data/fineweb-npy.py
torchrun --standalone --nproc_per_node=2 train_PP.py
Python
100.0%