lostmsu/TurboGPT

Train a tiny GPT in under a minute (CUDA only)

C++

3

4 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

Show HN: TurboGPT: train 22KiB transformer in 13s

29

Sep 29, 2026

README

turboGPT

Tiny byte-level GPT training in CUDA C++. MIT.

Build

Windows, Visual Studio 2022 C++ tools, and CUDA 13.4:

.\build.ps1 -CudaArch 86

CudaArch is the GPU compute capability from NVIDIA's CUDA GPU list.

Run

.\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4

The run stores its checkpoint at runs/ctx4/ctx4.pt, containing model, optimizer, scheduler, and trainer state. Use --load CHECKPOINT.pt to resume it.

runs/ctx4/report.json is derived from the log directory. Logs are TensorBoard-compatible: one report per batch, capped at 8Mi reports, and flushed with periodic or final checkpoints.

Result

  • hn1g after 1.5G training tokens: 2.52435 BPB.

Tests

python tests\verify.py

lostmsu/TurboGPT

Train a tiny GPT in under a minute (CUDA only)

C++

3

4 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

Show HN: TurboGPT: train 22KiB transformer in 13s

29

Sep 29, 2026

README

turboGPT

Tiny byte-level GPT training in CUDA C++. MIT.

Build

Windows, Visual Studio 2022 C++ tools, and CUDA 13.4:

.\build.ps1 -CudaArch 86

CudaArch is the GPU compute capability from NVIDIA's CUDA GPU list.

Run

.\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4

The run stores its checkpoint at runs/ctx4/ctx4.pt, containing model, optimizer, scheduler, and trainer state. Use --load CHECKPOINT.pt to resume it.

runs/ctx4/report.json is derived from the log directory. Logs are TensorBoard-compatible: one report per batch, capped at 8Mi reports, and flushed with periodic or final checkpoints.

Result

  • hn1g after 1.5G training tokens: 2.52435 BPB.

Tests

python tests\verify.py

Languages

C++

41.1%

Cuda

38.0%

Python

14.7%

PowerShell

6.2%