Tiny byte-level GPT training in CUDA C++. MIT.
Windows, Visual Studio 2022 C++ tools, and CUDA 13.4:
.\build.ps1 -CudaArch 86
CudaArch is the GPU compute capability from NVIDIA's CUDA GPU list.
.\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4
The run stores its checkpoint at
runs/ctx4/ctx4.pt, containing model, optimizer, scheduler, and trainer state.
Use --load CHECKPOINT.pt to resume it.
runs/ctx4/report.json is derived from the log directory. Logs are TensorBoard-compatible:
one report per batch, capped at 8Mi reports, and flushed with periodic or final checkpoints.
python tests\verify.py
C++
41.1%
Cuda
38.0%
Python
14.7%
PowerShell
6.2%
Tiny byte-level GPT training in CUDA C++. MIT.
Windows, Visual Studio 2022 C++ tools, and CUDA 13.4:
.\build.ps1 -CudaArch 86
CudaArch is the GPU compute capability from NVIDIA's CUDA GPU list.
.\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4
The run stores its checkpoint at
runs/ctx4/ctx4.pt, containing model, optimizer, scheduler, and trainer state.
Use --load CHECKPOINT.pt to resume it.
runs/ctx4/report.json is derived from the log directory. Logs are TensorBoard-compatible:
one report per batch, capped at 8Mi reports, and flushed with periodic or final checkpoints.
python tests\verify.py
C++
41.1%
Cuda
38.0%
Python
14.7%
PowerShell
6.2%