vLLM inference solutions for NVIDIA GB10 (DGX Spark, sm121) — KIMI-K3 B12X_MLA + DSpark, GLM-5.2
Python
48
63 commits
updated Sep 28, 2026
vLLM inference solutions for NVIDIA DGX Spark / DGX Spark+ (GB10) — Blackwell SM121.
This repo is a model-agnostic platform: a common GB10/SM121 vLLM stack (B12X_MLA attention, speculative decoding, PP/TP/EP) plus per-model subtrees for building images, deploying recipes, and applying runtime mods.
| Model | Subtree | Highlights |
|---|---|---|
| KIMI-K3 (v5, RoCE collectives + fused verify + baked perf stack) | kimi-k3/v5 | Recommended — TP16+DCP8, RoCEnante TP/DCP collectives, fused verify TILE8+nst6, _C rebuild (#54896/#55180), adaptive draft depth recipe |
| KIMI-K3 (v3, MoE fusion + DCP8 + RedHat DSpark) | kimi-k3/v3 | Superseded by v5 — TP16+DCP8, 2M total ctx capacity, >270K-context loop fixed, no runtime mods |
| KIMI-K3 (v2, sm121 image + RedHat DSpark) | kimi-k3/v2 | Superseded by v3 — TP16 (~500K ctx) / TP8+PP2 (~1M ctx) |
| KIMI-K3 (v1, B12X_MLA + DSpark) | kimi-k3/v1 | Historic artifact — kept for reference only |
| GLM-5.3 (Int4-Int8Mix) | glm-5.3 | v19 + v19-vision — same arch as 5.2, runs on the proven v19 image unchanged; vision via the same zero-copy Baseten composite |
| GLM-5.2 (Int4-Int8) | glm-5.2 | v19-vision (current) — DCP1 query split (+10% prefill @ 200K), deterministic MoE, vision, adaptive MTP 2/4/5 |
Each subtree is self-contained: a README.md model page, build script + Dockerfile, deploy
recipes (*.yaml), and runtime mods. Open the model page for its full guide.
TORCH_CUDA_ARCH_LIST=12.1a, CUTE_DSL_ARCH=sm_121a)nvidia/cuda:13.2.0-devel-ubuntu24.04, Python 3.12, uvlocal-inference-lab/vllm fork (B12X_MLA, DSpark, PP, KDA)zyang-dev/nccl (dgxspark-3node-ring)| Technology | What it does |
|---|---|
| B12X_MLA | Dense & sparse MLA attention via CuTe DSL (sparkinfer/b12x) |
| DSpark | Speculative decoding (draft model + acceptance) |
| MTP | Multi-token prediction drafting (GLM-5.2) |
| PP / TP / EP | Pipeline / tensor / expert parallelism across nodes |
| InstantTensor | Fast weights loading from NFS |
| Marlin MoE | Quantized expert parallel backend |
| Mods | Runtime patches applied at container start (mods/<name>/run.sh) |
gb10-vllm/
├── kimi-k3/v1/ KIMI-K3 v1 (historic artifact) — see kimi-k3/v1/README.md
├── kimi-k3/v2/ KIMI-K3 v2 (superseded) — sm121 image + RedHat DSpark
├── kimi-k3/v3/ KIMI-K3 v3 (recommended) — MoE fusion + DCP8 — see kimi-k3/v3/README.md
├── glm-5.2/ GLM-5.2 v16/v18/v18-vision/v19-vision — see glm-5.2/README.md
├── glm-5.3/ GLM-5.3 (Tech2wild Int4-Int8Mix) v19/v19-vision — see glm-5.3/README.md
├── README.md this file (platform overview + model index)
├── ATTRIBUTION.md upstream credits
└── CREDITS.md component licenses
Every model subtree follows one convention:
my-model/
├── README.md model page (quick start, build, bench, known issues)
├── build.sh image build (→ GHCR or local tag)
├── Dockerfile multi-stage SM121 build
├── recipes/*.yaml run-recipe deployables (recipe-schema YAML)
├── mods/<name> runtime mods (run.sh applied at container start)
└── (optional) wheels/ NOT used — vLLM/FlashInfer compiled in-image
Each model's deploy config is a set of YAML recipes run with the
eugr spark-vllm-docker harness (GLM-5.2 and KIMI-K3 both use ./run-recipe.sh <name>.yaml from
eugr). There are no per-model launcher scripts. Recipes encode model path, image tag, mods,
env vars, and the full vllm serve command.
./run-recipe.sh kimi-k3/v1/recipes/kimi-k3-full-hh-b12x-pp2.yaml # KIMI-K3 TP8+PP2
./run-recipe.sh glm-5.2/v18-vision/recipes/glm52-int4int8-v18.1-vision.yaml # GLM-5.2 v18.1 (fp8)
See each model page for the exact recipe list and pre-stop requirements.
./kimi-k3/v2/build.sh # KIMI-K3 v2 local tag (vllm-node-kimi3-sm121)
./kimi-k3/v2/build.sh --push # → ghcr.io/ciprianveg/gb10-vllm/kimi-k3:v2-sm121
./kimi-k3/v1/build.sh --push # → ghcr.io/ciprianveg/gb10-vllm/kimi-k3:latest (v1, historic)
./glm-5.2/v18-vision/build-nvfp4.sh --push # → ghcr.io/ciprianveg/gb10-glm-5.2:v18.1-vision
Prebuilt images:
ghcr.io/ciprianveg/gb10-vllm/kimi-k3:v2-sm121 (recommended)ghcr.io/ciprianveg/gb10-vllm/kimi-k3:latest (v1, historic)ghcr.io/ciprianveg/gb10-glm-5.2:v18.1-vision (and other v18 tags)Recipes set the same per-image env knobs (see each model page). Common ones:
TORCH_CUDA_ARCH_LIST=12.1a, CUTE_DSL_ARCH=sm_121aVLLM_USE_V2_MODEL_RUNNER=1HF_HOME=/cache/huggingface, offline modeNCCL_NET=IB, NCCL_IB_GID_INDEX=3, NCCL_BUFFSIZE=16777216, NCCL_MAX_NCHANNELS=8Model-specific credits and full attribution: see ATTRIBUTION.md and the model READMEs. Upstream: vLLM (vllm-project/vllm), local-inference-lab/vllm (B12X_MLA/DSpark/PP), b12x/sparkinfer, NCCL, FlashInfer, Marlin, InstantTensor.
Apache-2.0 (this repo). Model weights are subject to the respective providers' licenses — see each model page and CREDITS.
16 followers · starred Aug 2026
Python
96.0%
Shell
2.4%
Cuda
1.1%
vLLM inference solutions for NVIDIA GB10 (DGX Spark, sm121) — KIMI-K3 B12X_MLA + DSpark, GLM-5.2
Python
48
63 commits
updated Sep 28, 2026
vLLM inference solutions for NVIDIA DGX Spark / DGX Spark+ (GB10) — Blackwell SM121.
This repo is a model-agnostic platform: a common GB10/SM121 vLLM stack (B12X_MLA attention, speculative decoding, PP/TP/EP) plus per-model subtrees for building images, deploying recipes, and applying runtime mods.
| Model | Subtree | Highlights |
|---|---|---|
| KIMI-K3 (v5, RoCE collectives + fused verify + baked perf stack) | kimi-k3/v5 | Recommended — TP16+DCP8, RoCEnante TP/DCP collectives, fused verify TILE8+nst6, _C rebuild (#54896/#55180), adaptive draft depth recipe |
| KIMI-K3 (v3, MoE fusion + DCP8 + RedHat DSpark) | kimi-k3/v3 | Superseded by v5 — TP16+DCP8, 2M total ctx capacity, >270K-context loop fixed, no runtime mods |
| KIMI-K3 (v2, sm121 image + RedHat DSpark) | kimi-k3/v2 | Superseded by v3 — TP16 (~500K ctx) / TP8+PP2 (~1M ctx) |
| KIMI-K3 (v1, B12X_MLA + DSpark) | kimi-k3/v1 | Historic artifact — kept for reference only |
| GLM-5.3 (Int4-Int8Mix) | glm-5.3 | v19 + v19-vision — same arch as 5.2, runs on the proven v19 image unchanged; vision via the same zero-copy Baseten composite |
| GLM-5.2 (Int4-Int8) | glm-5.2 | v19-vision (current) — DCP1 query split (+10% prefill @ 200K), deterministic MoE, vision, adaptive MTP 2/4/5 |
Each subtree is self-contained: a README.md model page, build script + Dockerfile, deploy
recipes (*.yaml), and runtime mods. Open the model page for its full guide.
TORCH_CUDA_ARCH_LIST=12.1a, CUTE_DSL_ARCH=sm_121a)nvidia/cuda:13.2.0-devel-ubuntu24.04, Python 3.12, uvlocal-inference-lab/vllm fork (B12X_MLA, DSpark, PP, KDA)zyang-dev/nccl (dgxspark-3node-ring)| Technology | What it does |
|---|---|
| B12X_MLA | Dense & sparse MLA attention via CuTe DSL (sparkinfer/b12x) |
| DSpark | Speculative decoding (draft model + acceptance) |
| MTP | Multi-token prediction drafting (GLM-5.2) |
| PP / TP / EP | Pipeline / tensor / expert parallelism across nodes |
| InstantTensor | Fast weights loading from NFS |
| Marlin MoE | Quantized expert parallel backend |
| Mods | Runtime patches applied at container start (mods/<name>/run.sh) |
gb10-vllm/
├── kimi-k3/v1/ KIMI-K3 v1 (historic artifact) — see kimi-k3/v1/README.md
├── kimi-k3/v2/ KIMI-K3 v2 (superseded) — sm121 image + RedHat DSpark
├── kimi-k3/v3/ KIMI-K3 v3 (recommended) — MoE fusion + DCP8 — see kimi-k3/v3/README.md
├── glm-5.2/ GLM-5.2 v16/v18/v18-vision/v19-vision — see glm-5.2/README.md
├── glm-5.3/ GLM-5.3 (Tech2wild Int4-Int8Mix) v19/v19-vision — see glm-5.3/README.md
├── README.md this file (platform overview + model index)
├── ATTRIBUTION.md upstream credits
└── CREDITS.md component licenses
Every model subtree follows one convention:
my-model/
├── README.md model page (quick start, build, bench, known issues)
├── build.sh image build (→ GHCR or local tag)
├── Dockerfile multi-stage SM121 build
├── recipes/*.yaml run-recipe deployables (recipe-schema YAML)
├── mods/<name> runtime mods (run.sh applied at container start)
└── (optional) wheels/ NOT used — vLLM/FlashInfer compiled in-image
Each model's deploy config is a set of YAML recipes run with the
eugr spark-vllm-docker harness (GLM-5.2 and KIMI-K3 both use ./run-recipe.sh <name>.yaml from
eugr). There are no per-model launcher scripts. Recipes encode model path, image tag, mods,
env vars, and the full vllm serve command.
./run-recipe.sh kimi-k3/v1/recipes/kimi-k3-full-hh-b12x-pp2.yaml # KIMI-K3 TP8+PP2
./run-recipe.sh glm-5.2/v18-vision/recipes/glm52-int4int8-v18.1-vision.yaml # GLM-5.2 v18.1 (fp8)
See each model page for the exact recipe list and pre-stop requirements.
./kimi-k3/v2/build.sh # KIMI-K3 v2 local tag (vllm-node-kimi3-sm121)
./kimi-k3/v2/build.sh --push # → ghcr.io/ciprianveg/gb10-vllm/kimi-k3:v2-sm121
./kimi-k3/v1/build.sh --push # → ghcr.io/ciprianveg/gb10-vllm/kimi-k3:latest (v1, historic)
./glm-5.2/v18-vision/build-nvfp4.sh --push # → ghcr.io/ciprianveg/gb10-glm-5.2:v18.1-vision
Prebuilt images:
ghcr.io/ciprianveg/gb10-vllm/kimi-k3:v2-sm121 (recommended)ghcr.io/ciprianveg/gb10-vllm/kimi-k3:latest (v1, historic)ghcr.io/ciprianveg/gb10-glm-5.2:v18.1-vision (and other v18 tags)Recipes set the same per-image env knobs (see each model page). Common ones:
TORCH_CUDA_ARCH_LIST=12.1a, CUTE_DSL_ARCH=sm_121aVLLM_USE_V2_MODEL_RUNNER=1HF_HOME=/cache/huggingface, offline modeNCCL_NET=IB, NCCL_IB_GID_INDEX=3, NCCL_BUFFSIZE=16777216, NCCL_MAX_NCHANNELS=8Model-specific credits and full attribution: see ATTRIBUTION.md and the model READMEs. Upstream: vLLM (vllm-project/vllm), local-inference-lab/vllm (B12X_MLA/DSpark/PP), b12x/sparkinfer, NCCL, FlashInfer, Marlin, InstantTensor.
Apache-2.0 (this repo). Model weights are subject to the respective providers' licenses — see each model page and CREDITS.
16 followers · starred Aug 2026
Python
96.0%
Shell
2.4%
Cuda
1.1%