GLM-5.3-Flash EXL3 3.05bpw full-VRAM serving on 2x 64GB GPUs (CMP 170HX) with DFlash2 K7 / MTP d2 — no P2P required
Python
1
2 commits
updated Oct 7, 2026
Full-VRAM serving stack for GLM-5.3-Flash with DFlash2 K7 speculative decoding, tuned for two 64 GB GPUs with no P2P support (e.g. CMP 170HX mining cards). It also runs stock MTP d2 as an alternative speculation mode. ExLlamaV3 sources are not modified; every engine adjustment here is a process-local monkeypatch that is enabled only while this server runs.
English README (this file) · 한국어 README
| Component | Path | Notes |
|---|---|---|
| OpenAI-compatible API server | glm/glm_api.py | chat/completions + SSE, native PP/TG timings, tool calls |
| DFlash2 integration patches | glm/glm_dflash.py | drafter pinning, 8K GPU ring cache, fused-export preservation |
| Responsive async generator | glm/glm_async.py | GPU steps off the HTTP event loop, shielded cancellation |
| Q8 memory layout patches | glm/glm_q8_layout.py, glm/glm_q8_staging.py | lm_head on GPU 0, bounded staging, batch-1 decode workspace |
| GLM tool-call translation | glm/glm_tools.py | native XML tool calls → OpenAI tool_calls |
| mHC fused decode kernels | kernels/k_hcfuse.py | 0xSero k_hcfuse, ExLlamaV3 1.5.4 compatible |
| DFlash2 BF16 → EXL3 6bpw conversion | dflash2/convert_drafter.py, dflash2/package_native.py | 36 quantized linears, ~0.96 GiB drafter |
| Previous-request prefix reuse | glm/glm_dflash_prefix.py | opt-in cache-hit for repeated prompts, enabled by default via serve.sh |
| Setup / serve launchers | scripts/setup.sh, scripts/serve.sh | auto-download from Hugging Face on first run |
Not included (out of scope): the benchmark harness and raw measurement data from the experiment campaign, and the dashboard integration that this server was originally embedded in.
| Item | Value |
|---|---|
| Target model | turboderp/GLM-5.3-Flash-exl3 @ 3.05bpw (pinned revision) |
| Engine | ExLlamaV3 1.5.4 (SM80 native build) |
| Drafter | incoai/GLM-5.3-Flash-DFlash2 BF16 → self-converted EXL3 6bpw (~0.96 GiB) |
| Speculation | DFlash2 K7 (default) or MTP d2 |
| KV cache | Q8 (MLA latent), FP16 profile also available |
| Service context | 384K (Q8) / 320K / 256K / 192K (FP16) |
| Max request | input + output ≈ 392,960 tokens (384K profile) |
| GPUs | 2× 64 GB, SM80, no P2P (host-bounce D2D) |
| Split budget | -gs 62,63 (each number is a per-component load budget, not a VRAM split) |
All weights stay resident in HBM: no CPU expert tier, no SSD paging, no runtime weight transfer. This is a Strata-style HBM-first layout rebuilt on ExLlamaV3: the memory that grows with context (KV cache, drafter state) is kept small instead of offloading weights.
# 1. Environment (venv + ExLlamaV3 1.5.4 built for SM80 + API deps)
scripts/setup.sh
# 2. Serve. First run downloads the target (~125 GB) and the BF16 drafter
# (~2 GB) from Hugging Face, then converts the drafter to EXL3 6bpw
# (one-time, tens of minutes). Later runs reuse ./models.
scripts/serve.sh --profile q8_384k
# 3. Use it
curl http://127.0.0.1:8012/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"GLM-5.3-Flash","messages":[{"role":"user","content":"Hello"}],"max_tokens":512}'
Requirements: CUDA toolkit with nvcc (compute_80), Linux, ~140 GB free disk
for models + conversion workspace, and both GPUs free of other workloads.
The server refuses to start if either GPU has less than ~64 GB free.
HF_TOKEN is not required for these public repos, but set it if you hit rate
limits during the first download.
| Profile | Cache | Context limit (input+output) | Loader margin | Notes |
|---|---|---|---|---|
q8_384k | Q8, 393,216 tokens | 392,960 | 128 MiB | lm_head reserved on GPU 0; tightest layout |
q8_320k | Q8, 327,680 tokens | 327,424 | 256 MiB | |
q8_fast | Q8, 262,144 tokens | 261,888 | 192 MiB | |
fp16 | FP16, 196,608 tokens | 196,352 | 256 MiB | stock speed-first profile |
mtp | same as fp16 | 196,352 | 256 MiB | stock MTP d2 instead of DFlash2 |
-gs is ExLlamaV3's per-component load budget, not a VRAM split. For the
MTP profile the effective split is 59,63 because the MTP head is loaded onto
GPU 0 first; see glm/glm_profile.py for the validated values.
http://127.0.0.1:8012/v1 (model id GLM-5.3-Flash, no API key)POST /v1/chat/completions, POST /v1/completions, SSE streamingreasoning_content, the answer as contenttools / tool_choice / role: tool messages are supportedGET /health reports cache format, context, speculation mode, draft ring size
and prefix-cache stats (prompt_cache: hits/misses/restored tokens).GET /v1/models reports the active context limitreasoning_effort=high, temperature=1.0, top_p=0.95,
max_tokens=32768. One request runs at a time; the rest wait up to 120 s.timings exposes ExLlamaV3-native prompt_ms / predicted_ms plus draft
accepted/rejected counts and cached_tokens (also in
usage.prompt_tokens_details.cached_tokens), so PP/TG/acceptance and cache
hits can be measured from the client without trusting SSE arrival times.The 384K Q8 profile is the largest layout that fits on 2× 64 GB with this checkpoint:
--q8-head-gpu0)EXL3_AUTOSPLIT_MARGIN_MB)448K and 512K fail during the final transformer-layer load with this checkpoint and layout; they are not offered as profiles.
The drafter uses a fixed 2048-token sliding window. Its logical cache still matches the target context, but the physical GPU storage is an 8,192-token ring that recycles pages; absolute RoPE positions and the attention window are unchanged. This is what lets the drafter serve a 384K request without a 384K-token draft cache. The ring is verified for equivalence against the unbounded cache (bit-exact outputs over repeated wraps); see the experiment notes in the blog post this repo accompanies.
dflash2/convert_drafter.py wraps the stock ExLlamaV3 converter with two
process-local adjustments (the ExLlamaV3 tree itself is never modified):
k_proj/v_proj stay BF16 — serving builds a fused context-KV weight from
the qkv projection's K/V rows, which must remain a plain tensor.DFlash2DynConv.kernel_projection Linears are given a qmap so they enter
the quantization budget (env DRAFT_QUANT_KERNEL_PROJ=0 to keep them BF16).dflash2/package_native.py then restores every unquantized tensor byte-for-byte
from the BF16 source and writes native quantization_config / tensor_storage
metadata so ExLlamaV3 1.5.4 loads the result directly. Result: 36 quantized
linears at 6-bit trellis, ~0.96 GiB drafter weights.
Conversion uses synthetic-Hessian calibration only (DFlash2's
uncalibrated_quantize capability); no target-model forwards are involved.
k_hcfuse fuses the four mHC decode launches per sublayer site into two,
bit-exact against the stock kernels (shadow-check mode: GLM53_K_HCFUSE_CHECK=1000).
It JIT-compiles on first decode (SM80, nvcc); subsequent runs reuse the build cache.
The async generator runs native GPU steps on a dedicated worker thread so the HTTP loop stays responsive; disconnect cancellation is shielded until the in-flight GPU step and native job cleanup finish.
Q8 staging bounds packed-pool gather copies to 32K-token tiles and pins the quantized-append workspace to two fixed buffers per GPU, so repeated requests do not accumulate VRAM.
Prompt KV reuse is opt-in in the DFlash2 mode and enabled by default when
launched through scripts/serve.sh (--dflash-prefix-cache; disable with
GLM53_PREFIX_CACHE=0). The drafter ring has no snapshot of previous
requests on its own, so the prefix manager pairs each target recurrent
checkpoint with a host-RAM snapshot of the last 2048 drafter tokens and
restores both on a cache hit. A prefix miss, a mismatched checkpoint, a
cancelled request or an engine error all fall back to cold prefill.
Measured on the validated host (temperature 0, seed 42, 1K/8K/32K repeats): first-token latency for a repeated 1K prompt dropped 1.50 s → 0.65 s, an 8K conversation follow-up 7.86 s → 1.11 s (93.98% tokens reused), and a 32K follow-up 28.49 s → 1.12 s (98.42% reused). The host snapshot costs ~43 MiB of CPU RAM per stored checkpoint; GPU VRAM usage was unchanged within measurement noise. Greedy outputs for the repeated inputs matched the cold-run outputs, but longer real-text follow-ups at 8K/32K did not fully reproduce the full-reprocess outputs in the native GPU comparison, so the defect is documented rather than fixed — treat the reuse path as validated for the measured repeat cases, not as proven lossless in general.
Warm prefill throughput and the reuse-time saving are separate metrics.
This repo does not redistribute any model weights; everything is downloaded from the original Hugging Face repos at their pinned revisions.
MIT — see LICENSE.
GLM-5.3-Flash EXL3 3.05bpw full-VRAM serving on 2x 64GB GPUs (CMP 170HX) with DFlash2 K7 / MTP d2 — no P2P required
Python
1
2 commits
updated Oct 7, 2026
Full-VRAM serving stack for GLM-5.3-Flash with DFlash2 K7 speculative decoding, tuned for two 64 GB GPUs with no P2P support (e.g. CMP 170HX mining cards). It also runs stock MTP d2 as an alternative speculation mode. ExLlamaV3 sources are not modified; every engine adjustment here is a process-local monkeypatch that is enabled only while this server runs.
English README (this file) · 한국어 README
| Component | Path | Notes |
|---|---|---|
| OpenAI-compatible API server | glm/glm_api.py | chat/completions + SSE, native PP/TG timings, tool calls |
| DFlash2 integration patches | glm/glm_dflash.py | drafter pinning, 8K GPU ring cache, fused-export preservation |
| Responsive async generator | glm/glm_async.py | GPU steps off the HTTP event loop, shielded cancellation |
| Q8 memory layout patches | glm/glm_q8_layout.py, glm/glm_q8_staging.py | lm_head on GPU 0, bounded staging, batch-1 decode workspace |
| GLM tool-call translation | glm/glm_tools.py | native XML tool calls → OpenAI tool_calls |
| mHC fused decode kernels | kernels/k_hcfuse.py | 0xSero k_hcfuse, ExLlamaV3 1.5.4 compatible |
| DFlash2 BF16 → EXL3 6bpw conversion | dflash2/convert_drafter.py, dflash2/package_native.py | 36 quantized linears, ~0.96 GiB drafter |
| Previous-request prefix reuse | glm/glm_dflash_prefix.py | opt-in cache-hit for repeated prompts, enabled by default via serve.sh |
| Setup / serve launchers | scripts/setup.sh, scripts/serve.sh | auto-download from Hugging Face on first run |
Not included (out of scope): the benchmark harness and raw measurement data from the experiment campaign, and the dashboard integration that this server was originally embedded in.
| Item | Value |
|---|---|
| Target model | turboderp/GLM-5.3-Flash-exl3 @ 3.05bpw (pinned revision) |
| Engine | ExLlamaV3 1.5.4 (SM80 native build) |
| Drafter | incoai/GLM-5.3-Flash-DFlash2 BF16 → self-converted EXL3 6bpw (~0.96 GiB) |
| Speculation | DFlash2 K7 (default) or MTP d2 |
| KV cache | Q8 (MLA latent), FP16 profile also available |
| Service context | 384K (Q8) / 320K / 256K / 192K (FP16) |
| Max request | input + output ≈ 392,960 tokens (384K profile) |
| GPUs | 2× 64 GB, SM80, no P2P (host-bounce D2D) |
| Split budget | -gs 62,63 (each number is a per-component load budget, not a VRAM split) |
All weights stay resident in HBM: no CPU expert tier, no SSD paging, no runtime weight transfer. This is a Strata-style HBM-first layout rebuilt on ExLlamaV3: the memory that grows with context (KV cache, drafter state) is kept small instead of offloading weights.
# 1. Environment (venv + ExLlamaV3 1.5.4 built for SM80 + API deps)
scripts/setup.sh
# 2. Serve. First run downloads the target (~125 GB) and the BF16 drafter
# (~2 GB) from Hugging Face, then converts the drafter to EXL3 6bpw
# (one-time, tens of minutes). Later runs reuse ./models.
scripts/serve.sh --profile q8_384k
# 3. Use it
curl http://127.0.0.1:8012/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"GLM-5.3-Flash","messages":[{"role":"user","content":"Hello"}],"max_tokens":512}'
Requirements: CUDA toolkit with nvcc (compute_80), Linux, ~140 GB free disk
for models + conversion workspace, and both GPUs free of other workloads.
The server refuses to start if either GPU has less than ~64 GB free.
HF_TOKEN is not required for these public repos, but set it if you hit rate
limits during the first download.
| Profile | Cache | Context limit (input+output) | Loader margin | Notes |
|---|---|---|---|---|
q8_384k | Q8, 393,216 tokens | 392,960 | 128 MiB | lm_head reserved on GPU 0; tightest layout |
q8_320k | Q8, 327,680 tokens | 327,424 | 256 MiB | |
q8_fast | Q8, 262,144 tokens | 261,888 | 192 MiB | |
fp16 | FP16, 196,608 tokens | 196,352 | 256 MiB | stock speed-first profile |
mtp | same as fp16 | 196,352 | 256 MiB | stock MTP d2 instead of DFlash2 |
-gs is ExLlamaV3's per-component load budget, not a VRAM split. For the
MTP profile the effective split is 59,63 because the MTP head is loaded onto
GPU 0 first; see glm/glm_profile.py for the validated values.
http://127.0.0.1:8012/v1 (model id GLM-5.3-Flash, no API key)POST /v1/chat/completions, POST /v1/completions, SSE streamingreasoning_content, the answer as contenttools / tool_choice / role: tool messages are supportedGET /health reports cache format, context, speculation mode, draft ring size
and prefix-cache stats (prompt_cache: hits/misses/restored tokens).GET /v1/models reports the active context limitreasoning_effort=high, temperature=1.0, top_p=0.95,
max_tokens=32768. One request runs at a time; the rest wait up to 120 s.timings exposes ExLlamaV3-native prompt_ms / predicted_ms plus draft
accepted/rejected counts and cached_tokens (also in
usage.prompt_tokens_details.cached_tokens), so PP/TG/acceptance and cache
hits can be measured from the client without trusting SSE arrival times.The 384K Q8 profile is the largest layout that fits on 2× 64 GB with this checkpoint:
--q8-head-gpu0)EXL3_AUTOSPLIT_MARGIN_MB)448K and 512K fail during the final transformer-layer load with this checkpoint and layout; they are not offered as profiles.
The drafter uses a fixed 2048-token sliding window. Its logical cache still matches the target context, but the physical GPU storage is an 8,192-token ring that recycles pages; absolute RoPE positions and the attention window are unchanged. This is what lets the drafter serve a 384K request without a 384K-token draft cache. The ring is verified for equivalence against the unbounded cache (bit-exact outputs over repeated wraps); see the experiment notes in the blog post this repo accompanies.
dflash2/convert_drafter.py wraps the stock ExLlamaV3 converter with two
process-local adjustments (the ExLlamaV3 tree itself is never modified):
k_proj/v_proj stay BF16 — serving builds a fused context-KV weight from
the qkv projection's K/V rows, which must remain a plain tensor.DFlash2DynConv.kernel_projection Linears are given a qmap so they enter
the quantization budget (env DRAFT_QUANT_KERNEL_PROJ=0 to keep them BF16).dflash2/package_native.py then restores every unquantized tensor byte-for-byte
from the BF16 source and writes native quantization_config / tensor_storage
metadata so ExLlamaV3 1.5.4 loads the result directly. Result: 36 quantized
linears at 6-bit trellis, ~0.96 GiB drafter weights.
Conversion uses synthetic-Hessian calibration only (DFlash2's
uncalibrated_quantize capability); no target-model forwards are involved.
k_hcfuse fuses the four mHC decode launches per sublayer site into two,
bit-exact against the stock kernels (shadow-check mode: GLM53_K_HCFUSE_CHECK=1000).
It JIT-compiles on first decode (SM80, nvcc); subsequent runs reuse the build cache.
The async generator runs native GPU steps on a dedicated worker thread so the HTTP loop stays responsive; disconnect cancellation is shielded until the in-flight GPU step and native job cleanup finish.
Q8 staging bounds packed-pool gather copies to 32K-token tiles and pins the quantized-append workspace to two fixed buffers per GPU, so repeated requests do not accumulate VRAM.
Prompt KV reuse is opt-in in the DFlash2 mode and enabled by default when
launched through scripts/serve.sh (--dflash-prefix-cache; disable with
GLM53_PREFIX_CACHE=0). The drafter ring has no snapshot of previous
requests on its own, so the prefix manager pairs each target recurrent
checkpoint with a host-RAM snapshot of the last 2048 drafter tokens and
restores both on a cache hit. A prefix miss, a mismatched checkpoint, a
cancelled request or an engine error all fall back to cold prefill.
Measured on the validated host (temperature 0, seed 42, 1K/8K/32K repeats): first-token latency for a repeated 1K prompt dropped 1.50 s → 0.65 s, an 8K conversation follow-up 7.86 s → 1.11 s (93.98% tokens reused), and a 32K follow-up 28.49 s → 1.12 s (98.42% reused). The host snapshot costs ~43 MiB of CPU RAM per stored checkpoint; GPU VRAM usage was unchanged within measurement noise. Greedy outputs for the repeated inputs matched the cold-run outputs, but longer real-text follow-ups at 8K/32K did not fully reproduce the full-reprocess outputs in the native GPU comparison, so the defect is documented rather than fixed — treat the reuse path as validated for the measured repeat cases, not as proven lossless in general.
Warm prefill throughput and the reuse-time saving are separate metrics.
This repo does not redistribute any model weights; everything is downloaded from the original Hugging Face repos at their pinned revisions.
MIT — see LICENSE.