ciprianveg/gb10-vllm

vLLM inference solutions for NVIDIA GB10 (DGX Spark, sm121) — KIMI-K3 B12X_MLA + DSpark, GLM-5.2

Python

48

63 commits

updated Sep 28, 2026

See the code

See what people are saying

SourceMessageScoreDate

Do you need some extra memory on your DGX Spark? (r/LocalLLaMA)

​ I created this repo to help the DGX Spark users that have a spare 10-24 GB GPU at home to squeeze some extra memory out of a single Spark or a Sparks cluster. It moves the spec-decode draft model off your Sparks onto that GPU: the freed GB of memory can be used for extra context, or…

2

Sep 29, 2026

README

gb10-vllm

vLLM inference solutions for NVIDIA DGX Spark / DGX Spark+ (GB10) — Blackwell SM121.

This repo is a model-agnostic platform: a common GB10/SM121 vLLM stack (B12X_MLA attention, speculative decoding, PP/TP/EP) plus per-model subtrees for building images, deploying recipes, and applying runtime mods.

Models index

ModelSubtreeHighlights
KIMI-K3 (v5, RoCE collectives + fused verify + baked perf stack)kimi-k3/v5Recommended — TP16+DCP8, RoCEnante TP/DCP collectives, fused verify TILE8+nst6, _C rebuild (#54896/#55180), adaptive draft depth recipe
KIMI-K3 (v3, MoE fusion + DCP8 + RedHat DSpark)kimi-k3/v3Superseded by v5 — TP16+DCP8, 2M total ctx capacity, >270K-context loop fixed, no runtime mods
KIMI-K3 (v2, sm121 image + RedHat DSpark)kimi-k3/v2Superseded by v3 — TP16 (~500K ctx) / TP8+PP2 (~1M ctx)
KIMI-K3 (v1, B12X_MLA + DSpark)kimi-k3/v1Historic artifact — kept for reference only
GLM-5.3 (Int4-Int8Mix)glm-5.3v19 + v19-vision — same arch as 5.2, runs on the proven v19 image unchanged; vision via the same zero-copy Baseten composite
GLM-5.2 (Int4-Int8)glm-5.2v19-vision (current) — DCP1 query split (+10% prefill @ 200K), deterministic MoE, vision, adaptive MTP 2/4/5

Each subtree is self-contained: a README.md model page, build script + Dockerfile, deploy recipes (*.yaml), and runtime mods. Open the model page for its full guide.

Architecture

Hardware

  • GPU: 1× NVIDIA GB10 (DGX Spark / DGX Spark+) per node — Blackwell SM121
  • Cluster: 16 nodes (2×8), RoCE v2 (ConnectX-7, 100 Gbit) + NVLink-C2C, unified memory
  • Target arch: sm_121 (TORCH_CUDA_ARCH_LIST=12.1a, CUTE_DSL_ARCH=sm_121a)

Software stack

  • Base: nvidia/cuda:13.2.0-devel-ubuntu24.04, Python 3.12, uv
  • PyTorch: 2.11.0 + cu130
  • vLLM: local-inference-lab/vllm fork (B12X_MLA, DSpark, PP, KDA)
  • b12x/sparkinfer: CuTe DSL kernels for SM120/SM121 MLA
  • NCCL: zyang-dev/nccl (dgxspark-3node-ring)
  • FlashInfer: cu130/torch2.11 builds

Key technologies

TechnologyWhat it does
B12X_MLADense & sparse MLA attention via CuTe DSL (sparkinfer/b12x)
DSparkSpeculative decoding (draft model + acceptance)
MTPMulti-token prediction drafting (GLM-5.2)
PP / TP / EPPipeline / tensor / expert parallelism across nodes
InstantTensorFast weights loading from NFS
Marlin MoEQuantized expert parallel backend
ModsRuntime patches applied at container start (mods/<name>/run.sh)

Repository structure

gb10-vllm/
├── kimi-k3/v1/               KIMI-K3 v1 (historic artifact) — see kimi-k3/v1/README.md
├── kimi-k3/v2/               KIMI-K3 v2 (superseded) — sm121 image + RedHat DSpark
├── kimi-k3/v3/               KIMI-K3 v3 (recommended) — MoE fusion + DCP8 — see kimi-k3/v3/README.md
├── glm-5.2/                  GLM-5.2 v16/v18/v18-vision/v19-vision — see glm-5.2/README.md
├── glm-5.3/                  GLM-5.3 (Tech2wild Int4-Int8Mix) v19/v19-vision — see glm-5.3/README.md
├── README.md                 this file (platform overview + model index)
├── ATTRIBUTION.md            upstream credits
└── CREDITS.md                component licenses

Every model subtree follows one convention:

my-model/
├── README.md             model page (quick start, build, bench, known issues)
├── build.sh              image build (→ GHCR or local tag)
├── Dockerfile            multi-stage SM121 build
├── recipes/*.yaml        run-recipe deployables (recipe-schema YAML)
├── mods/<name>           runtime mods (run.sh applied at container start)
└── (optional) wheels/    NOT used — vLLM/FlashInfer compiled in-image

Recipes & the run-recipe harness

Each model's deploy config is a set of YAML recipes run with the eugr spark-vllm-docker harness (GLM-5.2 and KIMI-K3 both use ./run-recipe.sh <name>.yaml from eugr). There are no per-model launcher scripts. Recipes encode model path, image tag, mods, env vars, and the full vllm serve command.

./run-recipe.sh kimi-k3/v1/recipes/kimi-k3-full-hh-b12x-pp2.yaml                    # KIMI-K3 TP8+PP2
./run-recipe.sh glm-5.2/v18-vision/recipes/glm52-int4int8-v18.1-vision.yaml         # GLM-5.2 v18.1 (fp8)

See each model page for the exact recipe list and pre-stop requirements.

Building & publishing

./kimi-k3/v2/build.sh             # KIMI-K3 v2 local tag (vllm-node-kimi3-sm121)
./kimi-k3/v2/build.sh --push     # → ghcr.io/ciprianveg/gb10-vllm/kimi-k3:v2-sm121
./kimi-k3/v1/build.sh --push      # → ghcr.io/ciprianveg/gb10-vllm/kimi-k3:latest (v1, historic)
./glm-5.2/v18-vision/build-nvfp4.sh --push   # → ghcr.io/ciprianveg/gb10-glm-5.2:v18.1-vision

Prebuilt images:

  • ghcr.io/ciprianveg/gb10-vllm/kimi-k3:v2-sm121 (recommended)
  • ghcr.io/ciprianveg/gb10-vllm/kimi-k3:latest (v1, historic)
  • ghcr.io/ciprianveg/gb10-glm-5.2:v18.1-vision (and other v18 tags)

Environment conventions

Recipes set the same per-image env knobs (see each model page). Common ones:

  • TORCH_CUDA_ARCH_LIST=12.1a, CUTE_DSL_ARCH=sm_121a
  • VLLM_USE_V2_MODEL_RUNNER=1
  • HF_HOME=/cache/huggingface, offline mode
  • NCCL: NCCL_NET=IB, NCCL_IB_GID_INDEX=3, NCCL_BUFFSIZE=16777216, NCCL_MAX_NCHANNELS=8

Credits

Model-specific credits and full attribution: see ATTRIBUTION.md and the model READMEs. Upstream: vLLM (vllm-project/vllm), local-inference-lab/vllm (B12X_MLA/DSpark/PP), b12x/sparkinfer, NCCL, FlashInfer, Marlin, InstantTensor.

License

Apache-2.0 (this repo). Model weights are subject to the respective providers' licenses — see each model page and CREDITS.

Significant stargazers

Ian Levesque

16 followers · starred Aug 2026

ciprianveg/gb10-vllm

vLLM inference solutions for NVIDIA GB10 (DGX Spark, sm121) — KIMI-K3 B12X_MLA + DSpark, GLM-5.2

Python

48

63 commits

updated Sep 28, 2026

See the code

See what people are saying

SourceMessageScoreDate

Do you need some extra memory on your DGX Spark? (r/LocalLLaMA)

&amp;#x200B; I created this repo to help the DGX Spark users that have a spare 10-24 GB GPU at home to squeeze some extra memory out of a single Spark or a Sparks cluster. It moves the spec-decode draft model off your Sparks onto that GPU: the freed GB of memory can be used for extra context, or…

2

Sep 29, 2026

README

gb10-vllm

vLLM inference solutions for NVIDIA DGX Spark / DGX Spark+ (GB10) — Blackwell SM121.

This repo is a model-agnostic platform: a common GB10/SM121 vLLM stack (B12X_MLA attention, speculative decoding, PP/TP/EP) plus per-model subtrees for building images, deploying recipes, and applying runtime mods.

Models index

ModelSubtreeHighlights
KIMI-K3 (v5, RoCE collectives + fused verify + baked perf stack)kimi-k3/v5Recommended — TP16+DCP8, RoCEnante TP/DCP collectives, fused verify TILE8+nst6, _C rebuild (#54896/#55180), adaptive draft depth recipe
KIMI-K3 (v3, MoE fusion + DCP8 + RedHat DSpark)kimi-k3/v3Superseded by v5 — TP16+DCP8, 2M total ctx capacity, >270K-context loop fixed, no runtime mods
KIMI-K3 (v2, sm121 image + RedHat DSpark)kimi-k3/v2Superseded by v3 — TP16 (~500K ctx) / TP8+PP2 (~1M ctx)
KIMI-K3 (v1, B12X_MLA + DSpark)kimi-k3/v1Historic artifact — kept for reference only
GLM-5.3 (Int4-Int8Mix)glm-5.3v19 + v19-vision — same arch as 5.2, runs on the proven v19 image unchanged; vision via the same zero-copy Baseten composite
GLM-5.2 (Int4-Int8)glm-5.2v19-vision (current) — DCP1 query split (+10% prefill @ 200K), deterministic MoE, vision, adaptive MTP 2/4/5

Each subtree is self-contained: a README.md model page, build script + Dockerfile, deploy recipes (*.yaml), and runtime mods. Open the model page for its full guide.

Architecture

Hardware

  • GPU: 1× NVIDIA GB10 (DGX Spark / DGX Spark+) per node — Blackwell SM121
  • Cluster: 16 nodes (2×8), RoCE v2 (ConnectX-7, 100 Gbit) + NVLink-C2C, unified memory
  • Target arch: sm_121 (TORCH_CUDA_ARCH_LIST=12.1a, CUTE_DSL_ARCH=sm_121a)

Software stack

  • Base: nvidia/cuda:13.2.0-devel-ubuntu24.04, Python 3.12, uv
  • PyTorch: 2.11.0 + cu130
  • vLLM: local-inference-lab/vllm fork (B12X_MLA, DSpark, PP, KDA)
  • b12x/sparkinfer: CuTe DSL kernels for SM120/SM121 MLA
  • NCCL: zyang-dev/nccl (dgxspark-3node-ring)
  • FlashInfer: cu130/torch2.11 builds

Key technologies

TechnologyWhat it does
B12X_MLADense & sparse MLA attention via CuTe DSL (sparkinfer/b12x)
DSparkSpeculative decoding (draft model + acceptance)
MTPMulti-token prediction drafting (GLM-5.2)
PP / TP / EPPipeline / tensor / expert parallelism across nodes
InstantTensorFast weights loading from NFS
Marlin MoEQuantized expert parallel backend
ModsRuntime patches applied at container start (mods/<name>/run.sh)

Repository structure

gb10-vllm/
├── kimi-k3/v1/               KIMI-K3 v1 (historic artifact) — see kimi-k3/v1/README.md
├── kimi-k3/v2/               KIMI-K3 v2 (superseded) — sm121 image + RedHat DSpark
├── kimi-k3/v3/               KIMI-K3 v3 (recommended) — MoE fusion + DCP8 — see kimi-k3/v3/README.md
├── glm-5.2/                  GLM-5.2 v16/v18/v18-vision/v19-vision — see glm-5.2/README.md
├── glm-5.3/                  GLM-5.3 (Tech2wild Int4-Int8Mix) v19/v19-vision — see glm-5.3/README.md
├── README.md                 this file (platform overview + model index)
├── ATTRIBUTION.md            upstream credits
└── CREDITS.md                component licenses

Every model subtree follows one convention:

my-model/
├── README.md             model page (quick start, build, bench, known issues)
├── build.sh              image build (→ GHCR or local tag)
├── Dockerfile            multi-stage SM121 build
├── recipes/*.yaml        run-recipe deployables (recipe-schema YAML)
├── mods/<name>           runtime mods (run.sh applied at container start)
└── (optional) wheels/    NOT used — vLLM/FlashInfer compiled in-image

Recipes & the run-recipe harness

Each model's deploy config is a set of YAML recipes run with the eugr spark-vllm-docker harness (GLM-5.2 and KIMI-K3 both use ./run-recipe.sh <name>.yaml from eugr). There are no per-model launcher scripts. Recipes encode model path, image tag, mods, env vars, and the full vllm serve command.

./run-recipe.sh kimi-k3/v1/recipes/kimi-k3-full-hh-b12x-pp2.yaml                    # KIMI-K3 TP8+PP2
./run-recipe.sh glm-5.2/v18-vision/recipes/glm52-int4int8-v18.1-vision.yaml         # GLM-5.2 v18.1 (fp8)

See each model page for the exact recipe list and pre-stop requirements.

Building & publishing

./kimi-k3/v2/build.sh             # KIMI-K3 v2 local tag (vllm-node-kimi3-sm121)
./kimi-k3/v2/build.sh --push     # → ghcr.io/ciprianveg/gb10-vllm/kimi-k3:v2-sm121
./kimi-k3/v1/build.sh --push      # → ghcr.io/ciprianveg/gb10-vllm/kimi-k3:latest (v1, historic)
./glm-5.2/v18-vision/build-nvfp4.sh --push   # → ghcr.io/ciprianveg/gb10-glm-5.2:v18.1-vision

Prebuilt images:

  • ghcr.io/ciprianveg/gb10-vllm/kimi-k3:v2-sm121 (recommended)
  • ghcr.io/ciprianveg/gb10-vllm/kimi-k3:latest (v1, historic)
  • ghcr.io/ciprianveg/gb10-glm-5.2:v18.1-vision (and other v18 tags)

Environment conventions

Recipes set the same per-image env knobs (see each model page). Common ones:

  • TORCH_CUDA_ARCH_LIST=12.1a, CUTE_DSL_ARCH=sm_121a
  • VLLM_USE_V2_MODEL_RUNNER=1
  • HF_HOME=/cache/huggingface, offline mode
  • NCCL: NCCL_NET=IB, NCCL_IB_GID_INDEX=3, NCCL_BUFFSIZE=16777216, NCCL_MAX_NCHANNELS=8

Credits

Model-specific credits and full attribution: see ATTRIBUTION.md and the model READMEs. Upstream: vLLM (vllm-project/vllm), local-inference-lab/vllm (B12X_MLA/DSpark/PP), b12x/sparkinfer, NCCL, FlashInfer, Marlin, InstantTensor.

License

Apache-2.0 (this repo). Model weights are subject to the respective providers' licenses — see each model page and CREDITS.

Significant stargazers

Ian Levesque

16 followers · starred Aug 2026

Languages

Python

96.0%

Shell

2.4%

Cuda

1.1%