LibertAI Labs: running GLM-5.3-Flash-NVFP4 under vLLM on 2x GB10 (sm_121). Hand-written sparse-MLA CUDA kernel for NoPE MLA, plus a fix for vLLM's uninitialised NVFP4 MoE activation scale.
Python
6
7 commits
updated Aug 30, 2026
A LibertAI Labs artefact.
Running LibertAIDAI/GLM-5.3-Flash-NVFP4
under vLLM on Blackwell parts that vLLM did not support: sm_121 (GB10 /
DGX Spark class) and sm_120 (RTX PRO 6000 Blackwell, RTX 5090). Before this
work vLLM loaded the model, allocated KV, served an API, and then emitted
"locklocklocklock..." for every prompt. SGLang was the only engine that ran it.
It now serves correctly, on both. Two independent faults had to be fixed, and either one alone leaves the model degenerate. That is why the failure resisted diagnosis for so long: swapping attention backends never changed the output, because the MoE was zeroed regardless, which reads as "attention is exonerated" when it is not.
| fault | fix | |
|---|---|---|
| 1 | No MLA path for NoPE on sm_120/121. Every MLA prefill backend rejects (256, 0, 256) at capability 12.x, and the only sparse decode backend requires fp8_ds_mla, whose cache kernel asserts pe_dim == 64. | a hand-written sparse-MLA CUDA kernel plus vLLM backend (kernel/) |
| 2 | Uninitialised activation scale. ModelOptNvFp4FusedMoE registers w13_input_scale as PerTensorScaleParameter(data=torch.empty(...)), and a weight-only NVFP4 checkpoint never fills it. Observed value 0.0, so g1_alphas = weight_scale_2 * 0 = 0 and every expert output is multiplied by zero. | kernel/glm53_sparse_mla/moe_fix.py |
Fault 2 is not GB10-specific and not sm_12x-specific. The community workaround for it — "FLASHINFER_CUTLASS MoE produces garbage, use marlin" — works only because marlin dequantises to bf16 and never consumes an activation scale at all. The trigger is the checkpoint, not the GPU.
The full evidence chain, including the two hypotheses that turned out to be wrong, is in probe/FINDINGS.md.
The repository name says
gb10because that is where the work started. The kernel and the recipe now cover sm_120 as well, and the 4x RTX PRO 6000 lane below is the one carrying production traffic.
Shipped as two vllm.general_plugins entry points, so no vLLM file is
patched:
VLLM_GLM53_CUDA_SPARSE_MLA=1 — the sparse-MLA backendVLLM_GLM53_MOE_INPUT_SCALE=1.0 — supplies the missing activation scale| axis | support |
|---|---|
| arch | sm_120 and sm_121 |
| heads per rank | 8, 16, 32 — that is TP=8, 4, 2 on a 64-head model |
| KV dtype | bfloat16, fp8_e4m3 |
| prefill / decode | one code path, mixed batches in a single launch |
| CUDA graphs | AttentionCGSupport.UNIFORM_BATCH |
Verified on GB10 (sm_121), RTX PRO 6000 Blackwell (sm_120) and RTX 5090 (sm_120): 6 of 6 across TP and KV dtype, worst relative error 4.1e-3, cosine at or above 0.999998.
Design points worth keeping:
index >= 0
sentinel, which makes the index axis layout-agnostic and lets vLLM's paged slot
ids pass straight through. It is also why prefill and decode share one launch.mma.sync.m16n8k16 plus ldmatrix.-1 sentinel, and an all-NaN result for a fully masked token."The capital of France is" -> " Paris. It is the largest city in France and is known for
its iconic landmarks such as the Eiffel Tower..."
"Q: What is 17 times 23?" -> 391
| case | prompt tokens | result |
|---|---|---|
| three primes between 10 and 30, summed | 32 | 11 + 13 + 17 = 41, correct |
| capital of France | 25 | correct |
| long-context summarisation | 5,745 | correct |
The last row is the one that matters technically. At 5,745 tokens the prompt
exceeds index_topk = 2048, so the Lightning Indexer genuinely selects a subset
and the sparse kernel is exercised for real rather than degenerating to dense.
Both run the same checkpoint, the same kernel and the same two plugins. Almost everything else differs, because the two boxes have opposite bottlenecks.
| A. 2x GB10 | B. 4x RTX PRO 6000 Blackwell | |
|---|---|---|
| arch | sm_121, 2 nodes over RoCE | sm_120, 1 node, PCIe |
| memory | 121.69 GiB unified per node, shared with the OS | 96 GB per GPU, 384 GB total |
| tensor parallel | 2 (32 heads/rank) | 4 (16 heads/rank) |
| KV dtype | fp8 | fp8 |
| CUDA graphs | off (eager) | on |
| context | 65,536 | 262,144 |
--max-num-seqs | 2 | 256 |
gpu-memory-utilization | 0.80 | 0.94 |
| KV tokens | 88,790 | 5.3M - 6.3M |
| tok/s at c=1 | 24.2 | 129.8 |
| peak aggregate | not a batching box | 1,691 at c=128 |
| role | personal, publicly served | larger multi-GPU lane |
| lane | tok/s | KV tokens |
|---|---|---|
| this recipe, no speculation, eager | 14.46 | 127,291 @ 8K |
MTP num_speculative_tokens=3, eager | 24.04 | |
| MTP + CUDA graphs (breakable, vLLM's default) | 23.69 | 27,852 |
MTP + CUDA graphs (VLLM_USE_BREAKABLE_CUDAGRAPH=0) | 24.38 | 37,273 |
| MTP + eager + fp8 KV (deployed) | 24.21 | 88,790 @ 64K |
| SGLang, no speculation (reference) | 14.95 | |
| SGLang with NEXTN MTP | about 25 |
MTP acceptance rises with the config: mean length 2.37 eager, 2.46 with breakable graphs, 2.57 with non-breakable graphs.
Prefill and decode, measured under live traffic:
| prompt tokens | TTFT | prefill tok/s | decode tok/s |
|---|---|---|---|
| 299 | 0.70 s | 428 | 22.27 |
| 2,093 | 2.08 s | 1,005 | 25.07 |
| 8,303 | 6.47 s | 1,283 | 22.14 |
Decode is flat across context because index_topk = 2048 caps the attention cost.
The losses, stated next to the wins. Without speculation the lane sits at 14.46, which matches the SGLang baseline and is the memory-bandwidth wall for an 18B-active model on this hardware, so the attention kernel is not the limiter at concurrency 1. The whole stack reaches parity with SGLang rather than beating it; what it buys is vLLM's serving stack, MTP, and a second engine that can serve the model at all.
| concurrency | aggregate tok/s | per-request tok/s |
|---|---|---|
| 1 | 129.8 | 135.8 |
| 32 | 658.9 | 35.9 |
| 128 | 1,691.1 | 14.3 |
| 256 | 1,353.8 | 10.4 |
Aggregate peaks at 128 and falls back at 256.
Against a 4x B200 GLM-5.2 lane measured separately: it wins single-stream (130 versus about 100), matches context at 262K, and vastly exceeds KV capacity (5.3M versus 831K tokens). It loses on aggregate, roughly 1,700-2,200 against the B200's ~5,000 at c=256, about 2.5x below. Interactive traffic favours this box; batch-heavy traffic still favours the B200. That 5,000 figure is for a different model on a different vLLM and has not been re-measured with this harness, so treat the ratio as indicative.
Serving chain, matching the B200 stack:
https://<ip>:<external 8000>/v1 -> Caddy (TLS, AUTH_EXCLUDE="8000") ->
libertai-models uvicorn on 127.0.0.1:9000 (LibertAI signature auth) ->
vLLM on 127.0.0.1:18000.
# one-time: build the kernel inside the serving image and sync it to both ranks
export WORKER_HOST=10.0.0.2
./deploy/build-kernel.sh
cp deploy/env.glm53.example ~/glm53-serve/.env.glm53
cp deploy/docker-compose.glm53.yml ~/glm53-serve/
./deploy/start-glm53.sh
The settings that carry the fixes:
VLLM_GLM53_CUDA_SPARSE_MLA=1 # enable the sparse-MLA backend
VLLM_GLM53_MOE_INPUT_SCALE=1.0 # supply the scale vLLM reads but never loads
KV_CACHE_DTYPE=fp8 # bfloat16 also supported by the kernel
VLLM_GLM53_NOPE_PE_PAD= # off, head_size 512 is native now
MOE_BACKEND=flashinfer_cutlass # correct once the input scale is supplied
SPEC_CONFIG='{"method":"mtp","num_speculative_tokens":3}'
git clone https://github.com/Libertai/glm53-flash-vllm-gb10.git
cd glm53-flash-vllm-gb10
./deploy/rtx6000-4x/provision.sh
About 20 minutes end to end, most of it the 181 GiB download and the first engine
start. It is idempotent enough to re-run, and it does eight things: fetch the
weights, lift the glm5_next python tree out of the per-model vLLM image, build
and install the kernel for sm_120a, write the serve script, register it under
supervisor with the image's own vLLM disabled, install the libertai-models
gateway, repoint the host's external port 8000 at that gateway, and wait for
readiness.
The launch flags it writes:
--tensor-parallel-size 4
--kv-cache-dtype fp8 --block-size 256
--max-model-len 262144 --max-num-seqs 256 --max-num-batched-tokens 8192
--gpu-memory-utilization 0.94
--moe-backend flashinfer_cutlass
--tool-call-parser glm47 --reasoning-parser deepseek_r1 --enable-auto-tool-choice
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
with VLLM_USE_BREAKABLE_CUDAGRAPH=0, NCCL_MIN_NCHANNELS=32 and
NCCL_P2P_LEVEL=PXB.
Why the container lane extracts a python tree instead of running the image. A rented
instance is an unprivileged container with no Docker-in-Docker, so the per-model
vllm/vllm-openai:glm53-flash-x86_64-cu130 image that carries glm5_next cannot
be run directly. deploy/rtx6000-4x/pull_vllm.py pulls it from the registry and
lifts usr/local/lib/python3.12/dist-packages out of the layers, which works
because the instance ships the same interpreter and torch build (python 3.12,
torch 2.13.0+cu130). If those ever diverge from the image, this breaks, and the
symptom will be an ABI error rather than anything about GLM.
Build the kernel for the right arch. GLM53_ARCHS=120a for RTX PRO 6000 and
5090, 121a for GB10. The .so links libtorch, so it must be built against the
same CUDA and torch as the runtime that loads it.
The single most transferable lesson here, and the one that cost the most time:
CUDA graphs are worth about 1% on GB10 and about 500% on 4x RTX PRO 6000 (roughly 20 to 120 tok/s). GB10 is memory-bandwidth-bound, so kernel launch overhead hides behind the memory traffic and graphs buy nothing. TP=4 across PCIe is latency-bound over 45 layers of collectives, and graphs collapse exactly that. Do not carry a config lever across boxes with different bottlenecks; measure it again on each one.
The same asymmetry explains the rest of the table. GB10 runs eager at
--max-num-seqs 2 because its unified 121.69 GiB pool is shared with the OS and
three vLLM host processes, and because there is no throughput to win at
concurrency; the 4x box runs graphs at --max-num-seqs 256 and
gpu-memory-utilization 0.94 because it has 384 GB of dedicated VRAM and its
whole job is concurrency.
fp8 KV is slower than bf16, 0.88x to 0.96x, because the per-element dequantise costs more than the halved gather traffic saves. It is deployed on both anyway: on GB10 the doubled capacity buys 8x the context, and on the 4x box it is what puts KV in the 5-6M token range.
--kv-offloading-size and --kv-offloading-backend from the B200 GLM-5.2
recipe do not transfer. They fail with tokens_per_block=4 not divisible by tokens_per_hash=2304, because GLM-5.3-Flash is a hybrid linear-attention model
and GLM-5.2 is not.
Each of these cost real time, so they are recorded rather than left for the next person to rediscover.
source eats JSON quotes. start-glm53.sh runs set -a; source .env.glm53,
so an unquoted SPEC_CONFIG={"method":"mtp",...} reaches vLLM as
{method:mtp,...} and is rejected. Single-quote the whole value..so. A rank mismatch inside an attention
kernel is silently wrong output rather than a crash, so build-kernel.sh
checksums both..so rsynced from a dev box shadows the installed package and fails
with undefined symbol, which silently disables the plugins and resurfaces as
the original pe_dim must be 64 for fp8_ds_mla. The provision script deletes
any stale .so before building..so links libtorch, so cu129 against
cu130 is an ABI mismatch.ATen/cuda/CUDAContext.h does not compile in the serving image, because it
pulls in cusparse.h and that CUDA toolkit omits the cuSPARSE headers. Use
c10/cuda/* instead.--use_fast_math must stay off. It rewrites exp2f and breaks the numeric
match against the reference kernel.supervisorctl stop orphans workers that keep holding VRAM. Kill the
numeric PIDs from nvidia-smi --query-compute-apps=pid and verify the GPUs are
clear before relaunching.pkill -f 'vllm.entrypoints...' self-matches its own SSH command line and
kills the shell before the launch runs. Kill by PID instead.--reasoning-parser glm45 silently discards every reply. The model card's
SGLang recipe uses it, but on vLLM the name resolves to
Glm47MoeParserReasoningAdapter, whose state machine expects an opening
<think> in the output. This model's chat template emits <think> as the last
prompt token and the model only ever produces the closing </think>, so the
engine never opens a valid span and drops the text: completion_tokens: 400 with
content and reasoning_content both empty, on /v1/chat/completions and
/v1/messages alike. It reads as a hung or "lobotomised" model, not as a parser
fault. Use --reasoning-parser deepseek_r1, which terminates on a bare
</think>: verified to split cleanly into a thinking + text block pair on the
Anthropic path, with no </think> leaking into content.reasoning, not reasoning_content. Same rename as
the GB10 DeepSeek lane. A client reading reasoning_content sees null and may
conclude the trace was lost when it is simply under another key.sed a flag out of the compose command: block. It is a folded YAML
scalar that becomes a single shell string, so deleting only the flag text
leaves a whitespace-only line that folds into a literal newline. The shell then
splits the command in two and vLLM starts without every flag after that point
— in our case --nnodes 2, which failed as World size (2) is larger than the number of available GPUs (1) on both ranks and looked like a cluster/fabric
fault. Delete the whole line or replace it with a real flag, then verify with
yaml.safe_load that the command contains no interior newline.sm_scale and the epilogue, to remove a multiply
per element: exactly zero speedup, and it pushed two cases past tolerance
because acc_o then accumulates unscaled values up to fp8 max.docs/upstream-vllm-issue.md is the report filed as
vllm-project/vllm#54189. The
defect is an inconsistency: the NVFP4 linear methods refuse a weight-only
checkpoint loudly, and there is a ModelOptNvFp4W4A16LinearMethod, while
ModelOptNvFp4FusedMoE accepts the same checkpoint and computes with
uninitialised memory. There is no W4A16 equivalent for MoE.
kernel/ sparse-MLA CUDA kernel, vLLM backend, MoE fix, correctness harness
probe/ the standalone MoE harness that isolated fault 2, plus FINDINGS.md
deploy/ compose, env, launch and stop scripts, in-container kernel build
deploy/rtx6000-4x/ one-shot provisioning for the 4x RTX PRO 6000 lane
docs/ upstream bug report
Apache-2.0.
Python
62.0%
Cuda
34.3%
Shell
3.8%
LibertAI Labs: running GLM-5.3-Flash-NVFP4 under vLLM on 2x GB10 (sm_121). Hand-written sparse-MLA CUDA kernel for NoPE MLA, plus a fix for vLLM's uninitialised NVFP4 MoE activation scale.
Python
6
7 commits
updated Aug 30, 2026
A LibertAI Labs artefact.
Running LibertAIDAI/GLM-5.3-Flash-NVFP4
under vLLM on Blackwell parts that vLLM did not support: sm_121 (GB10 /
DGX Spark class) and sm_120 (RTX PRO 6000 Blackwell, RTX 5090). Before this
work vLLM loaded the model, allocated KV, served an API, and then emitted
"locklocklocklock..." for every prompt. SGLang was the only engine that ran it.
It now serves correctly, on both. Two independent faults had to be fixed, and either one alone leaves the model degenerate. That is why the failure resisted diagnosis for so long: swapping attention backends never changed the output, because the MoE was zeroed regardless, which reads as "attention is exonerated" when it is not.
| fault | fix | |
|---|---|---|
| 1 | No MLA path for NoPE on sm_120/121. Every MLA prefill backend rejects (256, 0, 256) at capability 12.x, and the only sparse decode backend requires fp8_ds_mla, whose cache kernel asserts pe_dim == 64. | a hand-written sparse-MLA CUDA kernel plus vLLM backend (kernel/) |
| 2 | Uninitialised activation scale. ModelOptNvFp4FusedMoE registers w13_input_scale as PerTensorScaleParameter(data=torch.empty(...)), and a weight-only NVFP4 checkpoint never fills it. Observed value 0.0, so g1_alphas = weight_scale_2 * 0 = 0 and every expert output is multiplied by zero. | kernel/glm53_sparse_mla/moe_fix.py |
Fault 2 is not GB10-specific and not sm_12x-specific. The community workaround for it — "FLASHINFER_CUTLASS MoE produces garbage, use marlin" — works only because marlin dequantises to bf16 and never consumes an activation scale at all. The trigger is the checkpoint, not the GPU.
The full evidence chain, including the two hypotheses that turned out to be wrong, is in probe/FINDINGS.md.
The repository name says
gb10because that is where the work started. The kernel and the recipe now cover sm_120 as well, and the 4x RTX PRO 6000 lane below is the one carrying production traffic.
Shipped as two vllm.general_plugins entry points, so no vLLM file is
patched:
VLLM_GLM53_CUDA_SPARSE_MLA=1 — the sparse-MLA backendVLLM_GLM53_MOE_INPUT_SCALE=1.0 — supplies the missing activation scale| axis | support |
|---|---|
| arch | sm_120 and sm_121 |
| heads per rank | 8, 16, 32 — that is TP=8, 4, 2 on a 64-head model |
| KV dtype | bfloat16, fp8_e4m3 |
| prefill / decode | one code path, mixed batches in a single launch |
| CUDA graphs | AttentionCGSupport.UNIFORM_BATCH |
Verified on GB10 (sm_121), RTX PRO 6000 Blackwell (sm_120) and RTX 5090 (sm_120): 6 of 6 across TP and KV dtype, worst relative error 4.1e-3, cosine at or above 0.999998.
Design points worth keeping:
index >= 0
sentinel, which makes the index axis layout-agnostic and lets vLLM's paged slot
ids pass straight through. It is also why prefill and decode share one launch.mma.sync.m16n8k16 plus ldmatrix.-1 sentinel, and an all-NaN result for a fully masked token."The capital of France is" -> " Paris. It is the largest city in France and is known for
its iconic landmarks such as the Eiffel Tower..."
"Q: What is 17 times 23?" -> 391
| case | prompt tokens | result |
|---|---|---|
| three primes between 10 and 30, summed | 32 | 11 + 13 + 17 = 41, correct |
| capital of France | 25 | correct |
| long-context summarisation | 5,745 | correct |
The last row is the one that matters technically. At 5,745 tokens the prompt
exceeds index_topk = 2048, so the Lightning Indexer genuinely selects a subset
and the sparse kernel is exercised for real rather than degenerating to dense.
Both run the same checkpoint, the same kernel and the same two plugins. Almost everything else differs, because the two boxes have opposite bottlenecks.
| A. 2x GB10 | B. 4x RTX PRO 6000 Blackwell | |
|---|---|---|
| arch | sm_121, 2 nodes over RoCE | sm_120, 1 node, PCIe |
| memory | 121.69 GiB unified per node, shared with the OS | 96 GB per GPU, 384 GB total |
| tensor parallel | 2 (32 heads/rank) | 4 (16 heads/rank) |
| KV dtype | fp8 | fp8 |
| CUDA graphs | off (eager) | on |
| context | 65,536 | 262,144 |
--max-num-seqs | 2 | 256 |
gpu-memory-utilization | 0.80 | 0.94 |
| KV tokens | 88,790 | 5.3M - 6.3M |
| tok/s at c=1 | 24.2 | 129.8 |
| peak aggregate | not a batching box | 1,691 at c=128 |
| role | personal, publicly served | larger multi-GPU lane |
| lane | tok/s | KV tokens |
|---|---|---|
| this recipe, no speculation, eager | 14.46 | 127,291 @ 8K |
MTP num_speculative_tokens=3, eager | 24.04 | |
| MTP + CUDA graphs (breakable, vLLM's default) | 23.69 | 27,852 |
MTP + CUDA graphs (VLLM_USE_BREAKABLE_CUDAGRAPH=0) | 24.38 | 37,273 |
| MTP + eager + fp8 KV (deployed) | 24.21 | 88,790 @ 64K |
| SGLang, no speculation (reference) | 14.95 | |
| SGLang with NEXTN MTP | about 25 |
MTP acceptance rises with the config: mean length 2.37 eager, 2.46 with breakable graphs, 2.57 with non-breakable graphs.
Prefill and decode, measured under live traffic:
| prompt tokens | TTFT | prefill tok/s | decode tok/s |
|---|---|---|---|
| 299 | 0.70 s | 428 | 22.27 |
| 2,093 | 2.08 s | 1,005 | 25.07 |
| 8,303 | 6.47 s | 1,283 | 22.14 |
Decode is flat across context because index_topk = 2048 caps the attention cost.
The losses, stated next to the wins. Without speculation the lane sits at 14.46, which matches the SGLang baseline and is the memory-bandwidth wall for an 18B-active model on this hardware, so the attention kernel is not the limiter at concurrency 1. The whole stack reaches parity with SGLang rather than beating it; what it buys is vLLM's serving stack, MTP, and a second engine that can serve the model at all.
| concurrency | aggregate tok/s | per-request tok/s |
|---|---|---|
| 1 | 129.8 | 135.8 |
| 32 | 658.9 | 35.9 |
| 128 | 1,691.1 | 14.3 |
| 256 | 1,353.8 | 10.4 |
Aggregate peaks at 128 and falls back at 256.
Against a 4x B200 GLM-5.2 lane measured separately: it wins single-stream (130 versus about 100), matches context at 262K, and vastly exceeds KV capacity (5.3M versus 831K tokens). It loses on aggregate, roughly 1,700-2,200 against the B200's ~5,000 at c=256, about 2.5x below. Interactive traffic favours this box; batch-heavy traffic still favours the B200. That 5,000 figure is for a different model on a different vLLM and has not been re-measured with this harness, so treat the ratio as indicative.
Serving chain, matching the B200 stack:
https://<ip>:<external 8000>/v1 -> Caddy (TLS, AUTH_EXCLUDE="8000") ->
libertai-models uvicorn on 127.0.0.1:9000 (LibertAI signature auth) ->
vLLM on 127.0.0.1:18000.
# one-time: build the kernel inside the serving image and sync it to both ranks
export WORKER_HOST=10.0.0.2
./deploy/build-kernel.sh
cp deploy/env.glm53.example ~/glm53-serve/.env.glm53
cp deploy/docker-compose.glm53.yml ~/glm53-serve/
./deploy/start-glm53.sh
The settings that carry the fixes:
VLLM_GLM53_CUDA_SPARSE_MLA=1 # enable the sparse-MLA backend
VLLM_GLM53_MOE_INPUT_SCALE=1.0 # supply the scale vLLM reads but never loads
KV_CACHE_DTYPE=fp8 # bfloat16 also supported by the kernel
VLLM_GLM53_NOPE_PE_PAD= # off, head_size 512 is native now
MOE_BACKEND=flashinfer_cutlass # correct once the input scale is supplied
SPEC_CONFIG='{"method":"mtp","num_speculative_tokens":3}'
git clone https://github.com/Libertai/glm53-flash-vllm-gb10.git
cd glm53-flash-vllm-gb10
./deploy/rtx6000-4x/provision.sh
About 20 minutes end to end, most of it the 181 GiB download and the first engine
start. It is idempotent enough to re-run, and it does eight things: fetch the
weights, lift the glm5_next python tree out of the per-model vLLM image, build
and install the kernel for sm_120a, write the serve script, register it under
supervisor with the image's own vLLM disabled, install the libertai-models
gateway, repoint the host's external port 8000 at that gateway, and wait for
readiness.
The launch flags it writes:
--tensor-parallel-size 4
--kv-cache-dtype fp8 --block-size 256
--max-model-len 262144 --max-num-seqs 256 --max-num-batched-tokens 8192
--gpu-memory-utilization 0.94
--moe-backend flashinfer_cutlass
--tool-call-parser glm47 --reasoning-parser deepseek_r1 --enable-auto-tool-choice
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
with VLLM_USE_BREAKABLE_CUDAGRAPH=0, NCCL_MIN_NCHANNELS=32 and
NCCL_P2P_LEVEL=PXB.
Why the container lane extracts a python tree instead of running the image. A rented
instance is an unprivileged container with no Docker-in-Docker, so the per-model
vllm/vllm-openai:glm53-flash-x86_64-cu130 image that carries glm5_next cannot
be run directly. deploy/rtx6000-4x/pull_vllm.py pulls it from the registry and
lifts usr/local/lib/python3.12/dist-packages out of the layers, which works
because the instance ships the same interpreter and torch build (python 3.12,
torch 2.13.0+cu130). If those ever diverge from the image, this breaks, and the
symptom will be an ABI error rather than anything about GLM.
Build the kernel for the right arch. GLM53_ARCHS=120a for RTX PRO 6000 and
5090, 121a for GB10. The .so links libtorch, so it must be built against the
same CUDA and torch as the runtime that loads it.
The single most transferable lesson here, and the one that cost the most time:
CUDA graphs are worth about 1% on GB10 and about 500% on 4x RTX PRO 6000 (roughly 20 to 120 tok/s). GB10 is memory-bandwidth-bound, so kernel launch overhead hides behind the memory traffic and graphs buy nothing. TP=4 across PCIe is latency-bound over 45 layers of collectives, and graphs collapse exactly that. Do not carry a config lever across boxes with different bottlenecks; measure it again on each one.
The same asymmetry explains the rest of the table. GB10 runs eager at
--max-num-seqs 2 because its unified 121.69 GiB pool is shared with the OS and
three vLLM host processes, and because there is no throughput to win at
concurrency; the 4x box runs graphs at --max-num-seqs 256 and
gpu-memory-utilization 0.94 because it has 384 GB of dedicated VRAM and its
whole job is concurrency.
fp8 KV is slower than bf16, 0.88x to 0.96x, because the per-element dequantise costs more than the halved gather traffic saves. It is deployed on both anyway: on GB10 the doubled capacity buys 8x the context, and on the 4x box it is what puts KV in the 5-6M token range.
--kv-offloading-size and --kv-offloading-backend from the B200 GLM-5.2
recipe do not transfer. They fail with tokens_per_block=4 not divisible by tokens_per_hash=2304, because GLM-5.3-Flash is a hybrid linear-attention model
and GLM-5.2 is not.
Each of these cost real time, so they are recorded rather than left for the next person to rediscover.
source eats JSON quotes. start-glm53.sh runs set -a; source .env.glm53,
so an unquoted SPEC_CONFIG={"method":"mtp",...} reaches vLLM as
{method:mtp,...} and is rejected. Single-quote the whole value..so. A rank mismatch inside an attention
kernel is silently wrong output rather than a crash, so build-kernel.sh
checksums both..so rsynced from a dev box shadows the installed package and fails
with undefined symbol, which silently disables the plugins and resurfaces as
the original pe_dim must be 64 for fp8_ds_mla. The provision script deletes
any stale .so before building..so links libtorch, so cu129 against
cu130 is an ABI mismatch.ATen/cuda/CUDAContext.h does not compile in the serving image, because it
pulls in cusparse.h and that CUDA toolkit omits the cuSPARSE headers. Use
c10/cuda/* instead.--use_fast_math must stay off. It rewrites exp2f and breaks the numeric
match against the reference kernel.supervisorctl stop orphans workers that keep holding VRAM. Kill the
numeric PIDs from nvidia-smi --query-compute-apps=pid and verify the GPUs are
clear before relaunching.pkill -f 'vllm.entrypoints...' self-matches its own SSH command line and
kills the shell before the launch runs. Kill by PID instead.--reasoning-parser glm45 silently discards every reply. The model card's
SGLang recipe uses it, but on vLLM the name resolves to
Glm47MoeParserReasoningAdapter, whose state machine expects an opening
<think> in the output. This model's chat template emits <think> as the last
prompt token and the model only ever produces the closing </think>, so the
engine never opens a valid span and drops the text: completion_tokens: 400 with
content and reasoning_content both empty, on /v1/chat/completions and
/v1/messages alike. It reads as a hung or "lobotomised" model, not as a parser
fault. Use --reasoning-parser deepseek_r1, which terminates on a bare
</think>: verified to split cleanly into a thinking + text block pair on the
Anthropic path, with no </think> leaking into content.reasoning, not reasoning_content. Same rename as
the GB10 DeepSeek lane. A client reading reasoning_content sees null and may
conclude the trace was lost when it is simply under another key.sed a flag out of the compose command: block. It is a folded YAML
scalar that becomes a single shell string, so deleting only the flag text
leaves a whitespace-only line that folds into a literal newline. The shell then
splits the command in two and vLLM starts without every flag after that point
— in our case --nnodes 2, which failed as World size (2) is larger than the number of available GPUs (1) on both ranks and looked like a cluster/fabric
fault. Delete the whole line or replace it with a real flag, then verify with
yaml.safe_load that the command contains no interior newline.sm_scale and the epilogue, to remove a multiply
per element: exactly zero speedup, and it pushed two cases past tolerance
because acc_o then accumulates unscaled values up to fp8 max.docs/upstream-vllm-issue.md is the report filed as
vllm-project/vllm#54189. The
defect is an inconsistency: the NVFP4 linear methods refuse a weight-only
checkpoint loudly, and there is a ModelOptNvFp4W4A16LinearMethod, while
ModelOptNvFp4FusedMoE accepts the same checkpoint and computes with
uninitialised memory. There is no W4A16 equivalent for MoE.
kernel/ sparse-MLA CUDA kernel, vLLM backend, MoE fix, correctness harness
probe/ the standalone MoE harness that isolated fault 2, plus FINDINGS.md
deploy/ compose, env, launch and stop scripts, in-container kernel build
deploy/rtx6000-4x/ one-shot provisioning for the 4x RTX PRO 6000 lane
docs/ upstream bug report
Apache-2.0.
Python
62.0%
Cuda
34.3%
Shell
3.8%