Reproducible local ML training, operator, and LLM inference results from NVIDIA
DGX Station systems with one NVIDIA GB300 selected per station. Multi-node
experiments use generic node0 and node1 roles. These are GB300 systems, not
DGX Spark/GB10.

The two-station direct-connect test setup.
Start with the DGX Station GB300 field guide for measured system specifications, photos, two-node networking, container setup, performance tuning, runtime quirks, benchmarking practice, and safe recovery.
| Experiment | Checkpoint / precision | Headline result |
|---|---|---|
| GLM-5.3 | Current recipe: local-inference-lab/GLM-5.3-NVFP4 (unmeasured); published 464.8 GB results: incoai/GLM-5.3-NVFP4, SGLang TP2 and patched vLLM PP2 DFlash2 on 2× GB300 | TP2 C1: 165.5 code / 107.4 prose tok/s; PP2 K7: 742.0 C16, 1,093.8 C64; PP2/AR prefill: 16,425 at 8K, 25,854 at 64K, 25,249 at 128K prompt tok/s |
| Qwen3.8-Flash-Next | local-inference-lab NVFP4-4p89/SGLang on 1× DGX Station GB300; 4× RTX PRO 6000 comparison | 1× TP1/MTP3 + ReplaySSM: 354.6 tok/s C1, 2,927.8 C64; TP1/AR: 4,090.4 C64, 38,653 tok/s 64K prefill |
| Qwen3.8-27B | BF16 plus unofficial Huginn FP8 and NVFP4A16 targets; BF16 KV/Mamba state | DFlash2: 265.8 tok/s C1; MTP: 6,348.8 C128. Quant AR C128: FP8 5,494.4 (+8.7% vs BF16), NVFP4A16 3,607.4 (−28.6%) |
| Qwen2.5-72B LoRA FSDP training | BF16 LoRA SFT; FSDP2 over 2× GB300; packed UltraChat 10K at 2,048 tokens | 4,453.19 tokens/s; 29.433 s/optimizer step; global batch 131,072 tokens |
| RF-DETR Large training | BF16 fine-tuning on the 1.17M-image MLPerf OpenImages subset; 2× GB300 DDP | Optimized global-batch-128 epoch: 220.45 images/s and 1:46:52 end to end (2.44× faster than control) |
| MLPerf RetinaNet training | Unofficial MLPerf Training v4.0 ssd reproduction; RetinaNet/ResNeXt-50 on OpenImages; 2× GB300 | Reached mAP 0.34759 in 5,365.422 s (1:29:25), versus a 2,159.003 s median for the published 8× H100 reference |
| nanoGPT training | modded-nanogpt FineWeb time-to-loss plus classic GPT-2 124M; 1× and 2× GB300 | modded 2×: 225.081 s to loss 3.2764; classic 2×: 1.838M tok/s; 95.0% / 98.95% scaling efficiency |
| GDN2 vs Mamba-3 vs Transformer Engine | Matched full training at ~1B parameters and 2,048 tokens, plus a separate GDN2 operator control | TE delayed FP8: 413.3k / 817.3k tokens/s; GDN2 BF16: 147.5k / 288.2k; Mamba-3 SISO BF16: 85.5k / 169.1k |
| DeepSeek-V4-Flash-0731 | 304B/13B-active native mixed FP4-expert/FP8-dense checkpoint | DSpark: 345.8 output tok/s at C1; C128 raw, capacity-limited: 6,511.1 aggregate output tok/s |
| DeepSeek-V4.1-Flash | Official native FP8-dense/FP4-expert checkpoint; 2× DGX Station only: SGLang TP2+EP2 and vLLM PP2/TP2 with DSpark (PP2+DSpark via a local five-file overlay) over Data Direct RDMA (1× not attempted) | Prefill: vLLM PP2 55,992 prompt tok/s for one 128K request, 65,966 aggregate at 64K C16; SGLang +SWA replay 39,549 for one 16K request, 37,693 aggregate at 128K C16 (exact full prefill 26,316 / 24,419); decode: vLLM PP2 DSpark (local overlay) 252.9 output tok/s per user at C1; vLLM TP2 DSpark 201.0 per user at C1 and 3,401.6 aggregate at C64; SGLang DSpark C1 180.0 output tok/s |
| Ornith-1.5-397B | Official ModelOpt NVFP4 W4A4 checkpoint; 1× TP1 and 2× PP2/TP2+EP | 1× C1: 129.8 output tok/s; 2× PP2 stable, capacity-limited C128: 3,799.6 aggregate tok/s |
| GLM-5.2 | Official NVIDIA NVFP4 checkpoint; 2× TP2+EP (1× does not fit) | C1: 68.0 output tok/s; shared-prefix C128: 2,012.4 aggregate tok/s |
| GLM-5.3-Flash | Official native FP8/vLLM plus LibertAIDAI/GLM-5.3-Flash-NVFP4/SGLang on 1× and 2× GB300; DFlash2 speculative decoding uses that NVFP4 base; 4× RTX PRO 6000 reference data | 1×: DFlash2 187.1 tok/s C1, AR 1,005.1 C64; 2×: DFlash2 198.0 C1 and 1,738.6 C64, AR 2,100.4 C64 |
| Hy3-FP8 | Official FP8 checkpoint; 2× PP2 and TP2+EP (1× does not fit) | MTP2 C1: 141.9 output tok/s; MTP1 C64: 2,563.7; no-spec C128: 3,078.5 aggregate tok/s |
| MiniMax H3 video | Official BF16 FL2VA checkpoint; resident 1× GB300, no offload | Official 5 s: 116.86 s mean; official 15 s: 719.29 s; experimental patched 30 s: 2,454.44 s |
| MiniMax M3 | Official NVIDIA NVFP4 (1×) plus official MiniMax MXFP8 (2× PP2) | NVFP4: 152.6 C1, 1,595.8 C16; MXFP8 PP2: 998.3 C32; 128K prefill: 35,683 tok/s; WikiText-2 PPL: 5.7120 / 5.4323 |
| Dual-station networking | ConnectX-8 400GbE RoCE with GB300 Data Direct | 392.1 Gb/s one-way raw GPUDirect; 389.8 Gb/s tuned NCCL all-reduce bus bandwidth |


Full GLM-5.3 results, proposal sweep, and recipes →
Each experiment folder contains:
recipes/ directory with pinned setup and reproduction commands| Component | Configuration |
|---|---|
| Hosts | One DGX Station, plus an identical peer where noted |
| Inference GPU | 1× NVIDIA GB300, 256,703 MiB reported HBM, 1,300 W power limit |
| CPU | NVIDIA Grace, 72 Arm Neoverse-V2 cores |
| System memory | 744 GiB |
| NVIDIA driver | 595.84 |
The display GPU was excluded from inference. Results are direct measurements from the tested systems, not vendor projections.
The throughput measurements use llm-inference-bench v0.4.29 at commit 0b4185b5b435e948b199c9077a00b084864aa963. Qwen and DeepSeek use its finite-request layer:
5 × concurrency measured requests after concurrency warm-up requestsQuality was tested separately with EOS respected, experiment-specific natural or mixed prompts, canonical WikiText-2 perplexity, and automated repetition audits. See each experiment README before comparing numbers; prompt construction, cache precision, and model architecture differ.
DeepSeek-V4.1-Flash prefill does not use llm-inference-bench. Its prefill
tables come from that section's own bench_prefill.py client against each
engine's native endpoint: unique random-token prompts of exactly 16K, 32K,
64K, and 128K tokens, one generated token, temperature 0, a fixed 1, 4, or 16
requests held in flight, a cache flush before every point, and aggregate
prompt tokens divided by the wave's wall time. Those multi-request aggregate
values are not interchangeable with the single-request cold-prefill cells of
the llm-inference-bench sections. Its decode rows use the finite-request
layer above with the same pinned client.
MiniMax H3 uses a separate fixed-seed video-and-audio methodology: one full 50-step warmup followed by three measured 1344×768, 124-frame requests. Its folder also measures native 10- and 15-second samples plus one explicitly unsupported patched 30-second attempt. It reports end-to-end latency, stage timing, peak HBM, media integrity, and manual non-degeneracy review; its results are not comparable to LLM token rates.
Ornith, GLM-5.2, and Hy3 instead use the benchmark's fixed-duration
sustained-decode layer: offered concurrency is maintained during a 30-second
measurement window, with separately labeled 60-second stability cells where
present. Boundary requests may remain in flight when the window closes.
Sustained-window throughput is not directly interchangeable with the finite
5 × concurrency results above; each experiment README identifies its layer
and retains completion and scheduler-residency fields.
.
├── qwen3.8-flash-next/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── qwen3.8-27b/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── nanogpt-training/
│ ├── README.md
│ ├── data/
│ └── recipes/
├── mlperf-retinanet-training/
│ ├── README.md
│ ├── data/
│ └── recipes/
├── gdn2-linear-attention/
│ ├── README.md
│ ├── data/
│ └── recipes/
├── gdn2-mamba3-te-comparison/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── transformer-engine-training/
│ ├── README.md
│ ├── data/
│ └── recipes/
├── mamba3-training/
│ ├── README.md
│ ├── data/
│ └── recipes/
├── dgx-station-guide/
│ ├── README.md
│ └── photos/
├── deepseek-v4-flash-0731/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── deepseek-v4.1-flash/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ ├── notes/
│ └── recipes/
├── ornith-1.5-397b/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── glm-5.2/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── glm-5.3/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── glm-5.3-flash/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── hy3/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── minimax-h3-video/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── minimax-m3/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
└── gb300-networking/
├── README.md
├── charts/
├── data/
└── recipes/
Measured August–September 2026.
Python
84.2%
Shell
15.6%
Reproducible local ML training, operator, and LLM inference results from NVIDIA
DGX Station systems with one NVIDIA GB300 selected per station. Multi-node
experiments use generic node0 and node1 roles. These are GB300 systems, not
DGX Spark/GB10.

The two-station direct-connect test setup.
Start with the DGX Station GB300 field guide for measured system specifications, photos, two-node networking, container setup, performance tuning, runtime quirks, benchmarking practice, and safe recovery.
| Experiment | Checkpoint / precision | Headline result |
|---|---|---|
| GLM-5.3 | Current recipe: local-inference-lab/GLM-5.3-NVFP4 (unmeasured); published 464.8 GB results: incoai/GLM-5.3-NVFP4, SGLang TP2 and patched vLLM PP2 DFlash2 on 2× GB300 | TP2 C1: 165.5 code / 107.4 prose tok/s; PP2 K7: 742.0 C16, 1,093.8 C64; PP2/AR prefill: 16,425 at 8K, 25,854 at 64K, 25,249 at 128K prompt tok/s |
| Qwen3.8-Flash-Next | local-inference-lab NVFP4-4p89/SGLang on 1× DGX Station GB300; 4× RTX PRO 6000 comparison | 1× TP1/MTP3 + ReplaySSM: 354.6 tok/s C1, 2,927.8 C64; TP1/AR: 4,090.4 C64, 38,653 tok/s 64K prefill |
| Qwen3.8-27B | BF16 plus unofficial Huginn FP8 and NVFP4A16 targets; BF16 KV/Mamba state | DFlash2: 265.8 tok/s C1; MTP: 6,348.8 C128. Quant AR C128: FP8 5,494.4 (+8.7% vs BF16), NVFP4A16 3,607.4 (−28.6%) |
| Qwen2.5-72B LoRA FSDP training | BF16 LoRA SFT; FSDP2 over 2× GB300; packed UltraChat 10K at 2,048 tokens | 4,453.19 tokens/s; 29.433 s/optimizer step; global batch 131,072 tokens |
| RF-DETR Large training | BF16 fine-tuning on the 1.17M-image MLPerf OpenImages subset; 2× GB300 DDP | Optimized global-batch-128 epoch: 220.45 images/s and 1:46:52 end to end (2.44× faster than control) |
| MLPerf RetinaNet training | Unofficial MLPerf Training v4.0 ssd reproduction; RetinaNet/ResNeXt-50 on OpenImages; 2× GB300 | Reached mAP 0.34759 in 5,365.422 s (1:29:25), versus a 2,159.003 s median for the published 8× H100 reference |
| nanoGPT training | modded-nanogpt FineWeb time-to-loss plus classic GPT-2 124M; 1× and 2× GB300 | modded 2×: 225.081 s to loss 3.2764; classic 2×: 1.838M tok/s; 95.0% / 98.95% scaling efficiency |
| GDN2 vs Mamba-3 vs Transformer Engine | Matched full training at ~1B parameters and 2,048 tokens, plus a separate GDN2 operator control | TE delayed FP8: 413.3k / 817.3k tokens/s; GDN2 BF16: 147.5k / 288.2k; Mamba-3 SISO BF16: 85.5k / 169.1k |
| DeepSeek-V4-Flash-0731 | 304B/13B-active native mixed FP4-expert/FP8-dense checkpoint | DSpark: 345.8 output tok/s at C1; C128 raw, capacity-limited: 6,511.1 aggregate output tok/s |
| DeepSeek-V4.1-Flash | Official native FP8-dense/FP4-expert checkpoint; 2× DGX Station only: SGLang TP2+EP2 and vLLM PP2/TP2 with DSpark (PP2+DSpark via a local five-file overlay) over Data Direct RDMA (1× not attempted) | Prefill: vLLM PP2 55,992 prompt tok/s for one 128K request, 65,966 aggregate at 64K C16; SGLang +SWA replay 39,549 for one 16K request, 37,693 aggregate at 128K C16 (exact full prefill 26,316 / 24,419); decode: vLLM PP2 DSpark (local overlay) 252.9 output tok/s per user at C1; vLLM TP2 DSpark 201.0 per user at C1 and 3,401.6 aggregate at C64; SGLang DSpark C1 180.0 output tok/s |
| Ornith-1.5-397B | Official ModelOpt NVFP4 W4A4 checkpoint; 1× TP1 and 2× PP2/TP2+EP | 1× C1: 129.8 output tok/s; 2× PP2 stable, capacity-limited C128: 3,799.6 aggregate tok/s |
| GLM-5.2 | Official NVIDIA NVFP4 checkpoint; 2× TP2+EP (1× does not fit) | C1: 68.0 output tok/s; shared-prefix C128: 2,012.4 aggregate tok/s |
| GLM-5.3-Flash | Official native FP8/vLLM plus LibertAIDAI/GLM-5.3-Flash-NVFP4/SGLang on 1× and 2× GB300; DFlash2 speculative decoding uses that NVFP4 base; 4× RTX PRO 6000 reference data | 1×: DFlash2 187.1 tok/s C1, AR 1,005.1 C64; 2×: DFlash2 198.0 C1 and 1,738.6 C64, AR 2,100.4 C64 |
| Hy3-FP8 | Official FP8 checkpoint; 2× PP2 and TP2+EP (1× does not fit) | MTP2 C1: 141.9 output tok/s; MTP1 C64: 2,563.7; no-spec C128: 3,078.5 aggregate tok/s |
| MiniMax H3 video | Official BF16 FL2VA checkpoint; resident 1× GB300, no offload | Official 5 s: 116.86 s mean; official 15 s: 719.29 s; experimental patched 30 s: 2,454.44 s |
| MiniMax M3 | Official NVIDIA NVFP4 (1×) plus official MiniMax MXFP8 (2× PP2) | NVFP4: 152.6 C1, 1,595.8 C16; MXFP8 PP2: 998.3 C32; 128K prefill: 35,683 tok/s; WikiText-2 PPL: 5.7120 / 5.4323 |
| Dual-station networking | ConnectX-8 400GbE RoCE with GB300 Data Direct | 392.1 Gb/s one-way raw GPUDirect; 389.8 Gb/s tuned NCCL all-reduce bus bandwidth |


Full GLM-5.3 results, proposal sweep, and recipes →
Each experiment folder contains:
recipes/ directory with pinned setup and reproduction commands| Component | Configuration |
|---|---|
| Hosts | One DGX Station, plus an identical peer where noted |
| Inference GPU | 1× NVIDIA GB300, 256,703 MiB reported HBM, 1,300 W power limit |
| CPU | NVIDIA Grace, 72 Arm Neoverse-V2 cores |
| System memory | 744 GiB |
| NVIDIA driver | 595.84 |
The display GPU was excluded from inference. Results are direct measurements from the tested systems, not vendor projections.
The throughput measurements use llm-inference-bench v0.4.29 at commit 0b4185b5b435e948b199c9077a00b084864aa963. Qwen and DeepSeek use its finite-request layer:
5 × concurrency measured requests after concurrency warm-up requestsQuality was tested separately with EOS respected, experiment-specific natural or mixed prompts, canonical WikiText-2 perplexity, and automated repetition audits. See each experiment README before comparing numbers; prompt construction, cache precision, and model architecture differ.
DeepSeek-V4.1-Flash prefill does not use llm-inference-bench. Its prefill
tables come from that section's own bench_prefill.py client against each
engine's native endpoint: unique random-token prompts of exactly 16K, 32K,
64K, and 128K tokens, one generated token, temperature 0, a fixed 1, 4, or 16
requests held in flight, a cache flush before every point, and aggregate
prompt tokens divided by the wave's wall time. Those multi-request aggregate
values are not interchangeable with the single-request cold-prefill cells of
the llm-inference-bench sections. Its decode rows use the finite-request
layer above with the same pinned client.
MiniMax H3 uses a separate fixed-seed video-and-audio methodology: one full 50-step warmup followed by three measured 1344×768, 124-frame requests. Its folder also measures native 10- and 15-second samples plus one explicitly unsupported patched 30-second attempt. It reports end-to-end latency, stage timing, peak HBM, media integrity, and manual non-degeneracy review; its results are not comparable to LLM token rates.
Ornith, GLM-5.2, and Hy3 instead use the benchmark's fixed-duration
sustained-decode layer: offered concurrency is maintained during a 30-second
measurement window, with separately labeled 60-second stability cells where
present. Boundary requests may remain in flight when the window closes.
Sustained-window throughput is not directly interchangeable with the finite
5 × concurrency results above; each experiment README identifies its layer
and retains completion and scheduler-residency fields.
.
├── qwen3.8-flash-next/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── qwen3.8-27b/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── nanogpt-training/
│ ├── README.md
│ ├── data/
│ └── recipes/
├── mlperf-retinanet-training/
│ ├── README.md
│ ├── data/
│ └── recipes/
├── gdn2-linear-attention/
│ ├── README.md
│ ├── data/
│ └── recipes/
├── gdn2-mamba3-te-comparison/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── transformer-engine-training/
│ ├── README.md
│ ├── data/
│ └── recipes/
├── mamba3-training/
│ ├── README.md
│ ├── data/
│ └── recipes/
├── dgx-station-guide/
│ ├── README.md
│ └── photos/
├── deepseek-v4-flash-0731/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── deepseek-v4.1-flash/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ ├── notes/
│ └── recipes/
├── ornith-1.5-397b/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── glm-5.2/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── glm-5.3/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── glm-5.3-flash/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── hy3/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── minimax-h3-video/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
├── minimax-m3/
│ ├── README.md
│ ├── charts/
│ ├── data/
│ └── recipes/
└── gb300-networking/
├── README.md
├── charts/
├── data/
└── recipes/
Measured August–September 2026.
Python
84.2%
Shell
15.6%