malaiwah/qwen38-27b-exl3

EXL3 mixed-precision quantization of Qwen3.8-27B: measured recipes, runtime contract, KLD receipts, iteration log

1

stars

559

commits

Python

primary language

Aug 26, 2026

updated

README

Status 2026-08-16: the collection is hardware-qualified, the comparator set is measured, and the method is now reproducible from published artifacts. The current checkpoint is malaiwah/Qwen3.8-27B-EXL3-K5K6: mean KLD 0.003210 (95 % CI [0.002982, 0.003480]) body-only at 20.32 GiB resident weights, versus official Qwen/Qwen3.8-27B-FP8 at 0.00529439 % lower divergence at 71 % of its resident weight, paired -0.002084 [-0.002249, -0.001942] winning 5,105 of 5,120 held-out contexts. The best profile in the collection is the hydrated build at 0.002760, paired -0.002534 on 5,118/5,120. The headline suite is v5: 5,120 contexts x 2,047 positions = 10,480,640 scored positions over 842 source clusters, calibration- and v4-token-disjoint.

Landed since: the context edition is qualified on a physical RTX 5090 (all seven gates at --gpu-memory-utilization 0.955 with expandable_segments, 265,122 KV tokens at 262,144 — and a second physical 5090 needed 0.956, so utilisation is a per-card measurement, not a constant). unsloth's NVFP4 is now on our suite at 0.031059 over 10.48 M positions, losing 5,120 of 5,120 contexts to both the context edition and FP8. The external llama-perplexity protocol was run for cross-citation and reproduces our ordering. Capture and replay are bit-reproducible, and a fresh conversion of the same recipe — a sibling, 97.6 % of quantized modules differing in bytes — scores indistinguishable from the published checkpoint, so the recipe determines fidelity while the published bytes remain the artifact. The suite, the shared BF16 reference capture and every per-shard report are published as a dataset, with archival mirrors of the third-party artifacts our numbers cite.

Two more artifacts published 2026-08-16, both from pre-registered conversions. malaiwah/Qwen3.8-27B-EXL3-K6-parity uses nearly the same file-byte budget as GGUF Q6_K by promoting gate_proj/up_proj K5 → K6. It measures 0.001634 [0.001541, 0.001742] under vLLM versus Q6_K at 0.002035 under llama.cpp, while beating hydrated on 511/512 contexts (-39.5 % for +1.348 GiB) and carrying 2.31 GiB less transformer body than Q6_K. The cross-engine control proves these are complete-pipeline results; it cannot be subtracted to establish format parity or prove that the earlier gap was caused by bytes. malaiwah/Qwen3.8-27B-EXL3-S16-V-research is the opposite: the rejected sub-4-bit 16 GB candidate at 0.045374 [0.041959, 0.049351], 1.5x its pre-registered NO threshold, losing 512/512 contexts to K4 (4.39x) and to the context edition (13.31x). It is published because it failed - a NO nobody can audit is an assertion - and it is the empirical anchor for the K3 rung of the per-bit law (docs/37 §3.3). Both payloads were predicted to the byte by the affine law before conversion, and both registered intervals are reported with their misses: S16-V's point estimate was 1.52x too high, K6-parity's interval fell 2.0 % short.

The earlier headline — 0.007945 body-only on 136 v3 contexts / 278,392 positions (docs/22) — stands exactly as measured and is superseded only as the headline; v5 absolute KLD is not comparable to v3 (the two suites differ by a measured 2.46-3.08x with identical ordering), so quote paired differences across suites, never means. Figures labelled K4 or v1 belong to iteration 1 and are superseded. One earlier control was withdrawn: the "CUDA-graph parity 0.000000" receipt captured a prefill forward, which FULL_DECODE_ONLY never captures, so it could not have measured decode; the decode probe that replaces it is docs/27. Open items are tracked in docs/29.

Fixed optimization objective (frozen 2026-08-21)

The primary control is the exact PROFILE=fidelity tuple, not PROFILE=throughput and not a memory-modified derivative. The objective is to recover enough PP/TTFT from that control to pass the absolute candidate gates without sacrificing its paired KLD/tails, TG, context, or capability envelope. The fidelity control is allowed to fail the 7,000 tok/s PP candidate threshold; that threshold governs promotion of a new candidate, not acceptance of the control.

receipts/frontier-fidelity-control.json pins the model/profile identity, INT6 embedding overlay, exact evidence hashes, reference metrics, candidate rules, and the prohibition on substituting the high-KLD throughput profile. Every INT8/FP4/FP6, K-map, topology, kernel-route, KV, MTP, graph, or profile change is a candidate variable and requires preregistered paired comparison against this exact control.

Frontier G02 pre-candidate controls (measured 2026-08-21)

The exact control and method-neutral harness are ready. Clean packed INT6 was integrated onto runtime b19029d, matched the historical encoder byte-for-byte on real BF16 rows, and passed three independent AIBoss boots. Aggregate PP was 2,941.9 tok/s, fox/essay TG 226.9/104.3, context 238,400, with text, vision, MTP, 200k context and rollback passing on every boot.

The clean-runtime fidelity recapture is exactly equal to the historical control: mean KLD 0.0034052972, p99 0.0348892, paired delta and CI 0.0, with all 512 per-context rows and the full tail histogram identical. Its 513-file, 10.73 GB hidden-state capture has an independently verified private-bucket copy.

The repaired v3 harness uses the actual EXL3 Viterbi quantizer—not the invalidated uniform proxy—on four real 128×128 Qwen slices with matched BF16/quant-flow activations and Hessians. RTN and both LDLQ controls carry the same exact 10,756-byte / 5.251953-bpw payload. Quant-flow-H LDLQ reduced mean excess running-output error from RTN's 0.0011858 to 0.0008865 at those same bytes; this is a control result, not a method promotion. Whole-block and end-logit gates remain mandatory and intentionally pending for researcher-proposed methods.

receipts/frontier-g02/manifest.json is validated by python3 tools/verify_frontier_g02.py --campaign receipts/frontier-g02 as pre_candidate_ready. Candidate slots remain unopened, E-final-v7 remains sealed, and rental authorization remains false.

Final Frontier G0/G1 campaign (measured 2026-08-20)

The preregistered campaign in docs/61-final-frontier-requant-stack-plan.md ended at its local RTX 5090 gate with the terminal disposition three_candidate_no_go. The complete hashed record is receipts/frontier-g01/manifest.json; python3 tools/verify_frontier_g01.py --campaign receipts/frontier-g01 validates every referenced G0 artifact, preregistration, result and terminal decision and exits 0.

G0 passed: the immutable BF16 payload was fully censused (1,199 logical tensors), converter/runtime environments were locked, the fidelity-derived incumbent-fidelity-int8 clean runtime was qualified, and every AIBoss maintenance transaction proved service restoration. That historical G0 arm used INT8 embeddings and did not remeasure the exact PROFILE=fidelity KLD; it is runtime/rollback evidence, not a substitute for the frozen primary control. G1 then consumed the three frozen candidate opportunities:

candidatemeasured resultfrozen Gate A failure
full gate/up FP6startup rejected the 199,104-token configuration; estimated maximum 192,576context below 238,400
full gate/up FP4PP 4,393.6 tok/s; fox/essay TG 211.6/98.8; context and capability checks passedPP below 7,000
gate/up FP4, language layers 0–31PP 3,529.5 tok/s; fox/essay TG 215.9/96.0; context and capability checks passedPP below 7,000

The preregistered hard-gate rule stopped KLD scoring after each deterministic runtime failure. Supporting experiments also closed without a promotable candidate: independent seeded converter A/A runs had 24 payload mismatches; the pinned AutoRound consumer could not build; ResComp improved one actual 16×16 weight-tile SSE by 21.1% but lacked a valid whole-block implementation; and the hybrid QKV producer/consumer pilot passed its component checks without the mandatory full-checkpoint, startup, graph and logit gates. No paid rental or new model release is authorized by this campaign; the measured serving profiles below remain current.

Serving on one RTX 5090 (measured 2026-08-19)

Speed vs fidelity

What each profile delivers

Three profiles, selected by PROFILE= in patches/run-qwen38-27b.sh. All three pass tools/verify-profile.sh (exit 0, 8/8 checks each, including a 200k-token prompt). PP/TG are the tools/bench-profile.sh harness at n=3 boots; KLD is 512 contexts of shard-0000 against the BF16 reference.

PROFILE=throughputPROFILE=fidelityPROFILE=balanced
weightsall-FP4all-trellis (K5K6 as shipped)trellis + gate_up MXFP6
PP, 2051-tok9638.9 ± 18.3 tok/s2987.7 ± 4.4 tok/s3925.2 ± 13.1 tok/s
TG fox / essay187.4 ± 0.6 / 94.3 ± 0.0228.3 ± 0.4 / 104.1 ± 0.1215.6 ± 0.2 / 103.7 ± 0.1
MTP acceptance fox / essay0.930 / 0.2981.000 / 0.3041.000 / 0.324
KLD mean0.0637590.003405 [0.003166, 0.003672]0.005672 [0.005302, 0.006087]
KLD p990.70100.0348890.059908
max context249,600238,400199,104
vision + MTPpasspasspass
criteria met4/6 (fails KLD)5/6 (fails PP)5/6 (fails ctx)

fidelity serves the checkpoint at KLD 0.003405 — within 26% of this collection's own published trellis fidelity (0.002700) — with TG 228.3 tok/s and full 238,400 context; the residual over the checkpoint is the int6 embedding table (~0.0007), not any GEMM approximation. It costs ~3.2x prefill. throughput is the only profile above 7000 tok/s prefill.

Long-context needle retrieval on the fidelity profile (fixed harness, single needle per context): 8/8 at 2k, 8/8 at 100k, 8/8 at 195k — 24/24. This is an easy proxy — finding a planted needle does not establish unimpaired long-context reasoning, only that the attention window and KV cache are intact to 195k.

No single profile meets all six north-star criteria. Concretely: throughput fails KLD (0.063759, well above the 0.012 budget); fidelity fails the 7000 tok/s prefill criterion (2987.7); balanced fails the context criterion (199,104 tokens, below the 238,400 that fidelity achieves). The other two profiles each clear five of six.

receipts/frontier-2026-08-19.md shows why from three directions: prefill-grade throughput needs the MLP resident in a GEMM-ready format, trellis-grade fidelity needs weights that are decoded per prefill chunk, and GEMM-resident FP8 for all 24.3e9 quantized parameters is 22.6 GiB — which cannot coexist with a 238,400-token KV cache in 31.4 GiB. The blocker is memory, not kernel quality.

Fixes landed while getting here, with receipts: an engine-fatal OOM on any prompt over ~4k tokens (max_num_batched_tokens 8192 → 3072, which also freed 0.93 GiB of KV), +32% TG by discovering the MTP draft loop ran eager on the V1 model runner, and the first profile of this stack (prefill is 92% GPU-bound; decode was 2%). Upstream: vllm-project/vllm#52871, #52872, local-inference-lab/vllm#439, #440, b12x #232/#233/#234.

Reproducible container image

A baked serving image is published at docker.io/malaiwah/qwen38-27b-exl3-gg:r34-p2-41a5d16 (also tagged :latest). It layers every patch this project bind-mounts on top of the digest-pinned Gilded Gnosis r34 base (voipmonitor/vllm:gilded-gnosis-v20-vllm4d006a4-b12xcd3ce19-fi1ac6942-cu132-20260810-r34, sha256 820181fbbc975cd5291c411cda9771d58fecee1636d916f508f47230df20592b), so the repo is not required at runtimepodman run plus a Hugging Face cache mount is sufficient. All 10 patches (7 vLLM source overlays and 3 conversion modules) are baked in by docker/Containerfile, which also bakes the CUDA lib64 symlink the B12X JIT needs as real links and carries OCI provenance labels (org.opencontainers.image.source, .revision, .description, ai.malaiwah.base-digest).

The image was gated 9/9 mount-free (NO_PATCH_MOUNTS=1, every patch bind-mount disabled): PP 2933.0, fox 229.1, essay 103.7, ctx 238,400, 200k-token prompt OK, vision OK, MTP acceptance_fox 1.000 — confirming the image is self-contained.

PROFILE=throughput|balanced|fidelity selects the serving profile; NO_PATCH_MOUNTS=1 is what makes the baked image run mount-free. The full flag set is in patches/run-qwen38-27b.sh. A standalone fidelity invocation — every flag below is taken from that script — is:

podman run --replace -d \
  --name qwen38-27b \
  --device nvidia.com/gpu=all --ipc=host --network host \
  --tmpfs /usr/local/cuda-13.2/lib64:rw,size=16m \
  -e HF_HUB_OFFLINE=1 \
  -e CUTE_DSL_ARCH=sm_120a -e FLASHINFER_CUDA_ARCH_LIST=12.0f \
  -e OMP_NUM_THREADS=8 -e CUDA_DEVICE_MAX_CONNECTIONS=32 \
  -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True -e SAFETENSORS_FAST_GPU=1 \
  -e VLLM_WORKER_MULTIPROC_METHOD=spawn \
  -e VLLM_USE_V2_MODEL_RUNNER=1 -e VLLM_NO_USAGE_STATS=1 -e DO_NOT_TRACK=1 \
  -e XDG_CACHE_HOME=/cache/jit -e CUDA_CACHE_PATH=/cache/jit \
  -e TRITON_CACHE_DIR=/cache/jit/triton -e TORCHINDUCTOR_CACHE_DIR=/cache/jit/torchinductor \
  -e FLASHINFER_WORKSPACE_BASE=/cache/jit/flashinfer \
  -e VLLM_EXL3_MULTIPRECISION=1 -e VLLM_EXL3_GRAPH_DECODE=1 \
  -e VLLM_EXL3_FP4_TRITON_DECODE=0 -e VLLM_EXL3_FP4_PER_ROW_GS=0 -e VLLM_EXL3_FP4_DRAFT_HEAD=0 \
  -e VLLM_EXL3_ONLINE_TRELLIS_BITS=6 \
  -e VLLM_EXL3_ONLINE_CACHE_DIR=/cache/jit/exl3-online -e VLLM_EXL3_ONLINE_CACHE_MODE=readwrite \
  -e VLLM_EXL3_EMBED_ONLINE_BITS=6 -e B12X_PACKED_B_MIN_N=1024 \
  -e VLLM_EXL3_B12X_ANY_BITS=1 -e VLLM_EXL3_B12X_MIN_M=128 -e VLLM_EXL3_SKIP_TRELLIS_PREP=0 \
  -e VLLM_EXL3_PREFILL_RECONSTRUCT_M=1 -e VLLM_EXL3_PREFILL_RECONSTRUCT_MAX_MB=512 \
  -e VLLM_EXL3_PREFILL_RECONSTRUCT_CACHE=0 -e VLLM_EXL3_FOLD_FP32_BUDGET_MB=48 \
  -e VLLM_EXL3_FP4_LAYERS=, \
  -e VLLM_EXL3_EXT_PATH=/opt/exllamav3 \
  -e HF_HOME=/root/.cache/huggingface \
  -v ~/.cache/huggingface/hub:/root/.cache/huggingface:ro \
  -v ~/.cache/jit:/cache/jit \
  --entrypoint /bin/bash \
  docker.io/malaiwah/qwen38-27b-exl3-gg:r34-p2-41a5d16 \
  -lc 'set -euo pipefail; \
       ln -sf /usr/local/cuda-13.2/targets/x86_64-linux/lib/* /usr/local/cuda-13.2/lib64/ 2>/dev/null || true; \
       exec vllm serve \
         /root/.cache/huggingface/models--malaiwah--Qwen3.8-27B-EXL3-K5K6-hydrated/snapshots/ab3a91a13813df8096cb4c1d560ed3669035d0cf \
         --served-model-name Qwen3.8-27B --trust-remote-code \
         --host 0.0.0.0 --port 8000 \
         --quantization exl3 \
         --quantization-config "{\"linear\":{\"weight\":\"mxfp8\"},\"ignore\":[\"re:.*visual\\..*\",\"re:.*in_proj_a$\",\"re:.*in_proj_b$\",\"re:.*in_proj_ba$\",\"re:.*mtp\\..*\",\"lm_head\"]}" \
         --attention-backend TRITON_ATTN \
         --gpu-memory-utilization 0.945 --kv-cache-dtype fp8_e4m3 \
         --max-model-len 238400 --max-num-seqs 4 --max-num-batched-tokens 3072 \
         --compilation-config "{\"mode\":\"NONE\",\"cudagraph_mode\":\"FULL_DECODE_ONLY\"}" \
         --mm-processor-kwargs "{\"truncation\":false}" --mm-processor-cache-type shm \
         --default-chat-template-kwargs "{\"preserve_thinking\": true}" \
         --enable-chunked-prefill --reasoning-parser qwen3 \
         --enable-auto-tool-choice --tool-call-parser qwen3_coder \
         --speculative-config "{\"method\":\"mtp\",\"num_speculative_tokens\":6}"'

Flags taken from run-qwen38-27b.sh: --tmpfs + ln -sf symlink (B12X JIT workaround, lines 225/296 — the baked image already carries these symlinks, so the --tmpfs is redundant with the baked image but harmless), --device / --ipc / --network (line 226), the -e environment block (lines 231–280), the HF cache mount -v …/huggingface:ro (line 281), the JIT cache volume -v …/cache/jit:/cache/jit (line 288), and the vllm serve argument list (lines 303–322). For balanced or throughput, change the profile env vars per the case block in the script (lines 83–154).

Qwen3.8-27B EXL3 mixed-precision quants (K4, EXL3-K5K6)

Research materials and progress log for building a dense EXL3 quant of Qwen/Qwen3.8-27B that is inspired by NVIDIA's NVFP4 recipe for the previous generation, but keeps the attention projections BF16 on disk for runtime K6 encoding by the Gilded Gnosis vLLM fork, and serializes the MLP as EXL3 K5 gate / K5 up / K6 down with a K6 mcg head and a quantized MTP draft. Iteration 1 serialized the whole MLP at K4.

Design goal: NVFP4-class VRAM footprint, lower KLD. Met on fidelity, memory, decode and speculative decode; prefill is a measured structural deficit of a 4-bit-class trellis format in this runtime (docs/25, docs/26).

Evidence at a glance

Held-out v5 suite, receipts/kld5-suite-manifest.json (schema qwen38-distribution-fidelity/6, suite_token_sha256 510541f6…09482b88): 5,120 contexts x 2,047 positions = 10,480,640 scored positions, 842 source clusters, 1,024 contexts per stratum. 44 of 941 corpus documents were excluded whole before window selection for any all-position 12-token overlap with exllamav3 calibration data, leaving 897 eligible, so contamination hits are 0 by construction; all 160 v4 context token hashes were seeded as exclusions and 0 were reachable. Every figure below is body-only — both operands scored through one shared BF16 LM head — with a 95 % CI from a 10,000-resample bootstrap over the 842 source clusters.

candidatemean KLD95 % CItop-1paired vs FP8contexts wonreceipt
hydrated0.002760[0.002540, 0.003020]97.70 %-0.0025345,118 / 5,120receipts/kld5-10M-hyd.json
online K5/K60.003210[0.002982, 0.003480]97.52 %-0.0020845,105 / 5,120receipts/kld5-10M-k5k6.json
context0.003509[0.003220, 0.003852]97.44 %-0.0017855,109 / 5,120receipts/kld5-10M-ctx.json
official FP80.005294[0.004927, 0.005728]96.79 %receipts/kld5-10M-fp8.json
K4 (iteration 1)0.010604[0.009640, 0.011746]95.76 %+0.0053107 / 5,120receipts/kld5-10M-k4.json

Paired intervals and win counts come from receipts/kld5-10M-paired.json (10,000 resamples, seed 1); hydrated also beats the runtime overlay directly, -0.000450 [-0.000469, -0.000433] on 4,922/5,120. Exact worst single position over the whole run: hydrated 8.258, online 22.241, context 5.557, FP8 10.714, K4 14.283. The cumulative hydrated mean at the 1M / 2M / 5M / 10M ladder checkpoints is 0.002700 / 0.002759 / 0.002699 / 0.002760 (shards[] of receipts/kld5-10M-hyd.json).

The tail, not just the mean. The ten-shard run predates the histogram, so the tail was measured by re-running shard 0 of the same v5 suite with the qwen38-fidelity-report/2 harness: 512 contexts x 2,047 positions = 1,048,064 scored positions, the same contexts every candidate saw. Quantiles are bin-bounded — 560 log-spaced bins from 1e-12 to 1e2 nats, one bin ~5.6 % wide, and every receipt carries lower / upper / estimate per quantile — while the maxima and the exceedance counts are exact. The two columns worth reading are p99.9 and the share of positions above 0.1 nats.

candidatemeanp50p95p99p99.9p99.99exact maxabove 0.1above 1.0receipt
hydrated0.0027000.001090.00820.02760.13190.4633.7350.1534 %0.00219 %receipts/kld5-1M-tail-hyd.json
online K5/K60.0031410.001280.00990.03210.14460.4985.5070.1820 %0.00200 %receipts/kld5-1M-tail-k5k6.json
context0.0034090.001350.01070.03570.16420.5873.7490.2287 %0.00305 %receipts/kld5-1M-tail-ctx.json
official FP80.0051970.002020.01670.05310.24380.8125.2960.3912 %0.00592 %receipts/kld5-1M-tail-fp8.json
K40.0103450.003200.03320.11940.55551.8707.5651.2604 %0.03807 %receipts/kld5-1M-tail-k4.json

The ordering at p50, p95, p99, p99.9 and p99.99 is the same as the ordering of the means, so for these candidates the mean is not hiding a worse tail: every EXL3 K5/K6-class build has a lighter tail than official FP8 at every measured quantile, and K4 is worse than FP8 at every quantile. Receipts are schema qwen38-kld-ladder-cumulative/2, welded by tools/kld_aggregate.py from /2 replay reports.

Against GGUF, measured, including the part that goes against us. The standing critique is that official FP8 is a throughput format whose quality is Q4-to-Q5 class, so beating it is a weak claim, and that Q8_0 and Q6_K are the honest bar. That is now measured rather than argued. Three GGUFs from unsloth/Qwen3.8-27B-GGUF@f1bfb127c64f7072bdd2cad55f258b9c8b2910fe were captured under llama.cpp pinned at commit ece963f41b0b02d7a0d61436ae365762c073a4c8 with tools/gguf_capture.cpp, which reads the post-final-norm state — the same mathematical point our vLLM hook takes, with bf16 rounding verified bit-identical to torch on 2,012,449 probe values — and scored through the same shared BF16 head on shard 0 of the v5 suite: the same 512 contexts and the same 1,048,064 positions every candidate above saw (tools/gguf_manifest.py, tools/build_llamacpp.sh; receipt receipts/cross-engine-comparator.json, per-candidate reports receipts/gguf-report-{q8_0,q6_k,q5_k_xl}.json).

candidateenginemeasured mean KLDtop-1p99.9serialized
GGUF Q8_0llama.cpp0.00108798.53 %0.035127.05 GiB
GGUF Q6_Kllama.cpp0.00203597.98 %0.079421.31 GiB
hydrated EXL3vLLM0.00270097.80 %0.131320.12 GiB payload
online K5/K6 EXL3vLLM0.00314197.61 %0.1447
context EXL3vLLM0.00340997.55 %0.163219.27 GiB payload
GGUF UD-Q5_K_XLllama.cpp0.00444497.20 %0.214418.83 GiB
official FP8vLLM0.00519796.92 %0.244028.51 GiB resident
K4 EXL3vLLM0.01034595.91 %0.5576
unsloth/Qwen3.8-27B-NVFP4vLLM0.03011593.16 %1.622822.57 GB checkpoint

The cross-engine BF16 control is measured the same way, not assumed: llama.cpp against the vLLM BF16 reference on identical tokens, the shared head and the same 512 contexts reads 0.000507 mean, 99.07 % top-1 and p99.9 0.0113 (receipts/gguf-report-engine-floor.json). It proves the engine is a confounder. KL is neither additive nor a metric, so this value cannot be subtracted from a candidate KL and provides no quantization-only upper or lower bound. The two cells are builds that ship BF16 attention for the runtime to encode at load, so their disk bytes are not a like-for-like payload; payload figures are immutable_payload_bytes from receipts/collection-index.json (hydrated 21,610,916,123 B = 20.127 GiB, context 20,696,033,532 B = 19.275 GiB; the table truncates both to two decimals) and are serialized bytes, never VRAM. The FP8 figure is resident weights and is labelled as such.

The p99.9 column is each report's exact shard-0 p99.9 as the comparator receipt read them; the tail table above quotes the bin-bounded cumulative estimate from the 560-bin histogram (bins about 5.6 % wide), which is why hydrated reads 0.1319 there and 0.1313 here — each exact value lies inside the bin its estimate names. The two differ by construction, not by measurement.

What the table supports is complete-pipeline ordering, not format attribution:

  • At the nominal 6-bit point, llama.cpp Q6_K measures 0.002035 and vLLM hydrated measures 0.002700 on the same contexts.
  • At the nominal 5-bit point, vLLM context measures 0.003409 and llama.cpp UD-Q5_K_XL measures 0.004444.

Both comparisons are cross-engine. They are evidence about the tested artifact-plus-engine pipelines; neither isolates the quantization format.

Correction, 2026-08-16 — the byte axis in the table above is not one axis, and every mixed comparison flattered us. A GGUF row is the whole file of a text-only artifact; our row is tensor payload of a multimodal tree that also carries an MTP draft. Read from each artifact's own tensor table, without downloading any payload (cross-candidate-byte-accounting.json):

candidatefiletensor totaltoken_embdoutputtransformer bodymultimodal deployed
GGUF Q8_027.0527.041.258 (Q8_0)1.258 (Q8_0)24.52627.92 (+ mmproj-BF16 0.867)
GGUF Q6_K21.3121.300.971 (Q6_K)0.971 (Q6_K)19.36022.18 (+ mmproj-BF16 0.867)
GGUF UD-Q5_K_XL18.8318.820.814 (Q5_K)0.971 (Q6_K)17.03419.70 (+ mmproj-BF16 0.867)
online K5/K628.5028.472.368 (BF16)0.889 (K6)24.11928.50, vision inside
K426.3726.372.368 (BF16)0.593 (K4)21.46326.37, vision inside
hydrated K5/K620.1320.102.368 (BF16)0.889 (K6)15.72620.13, vision inside
context edition19.2719.252.368 (BF16)0.889 (K6)14.88619.27, vision inside

All figures GiB. The embedding and head widths are not uniform across GGUF tiers, and the vision encoder is absent from every GGUF text file - it ships separately as mmproj-BF16.gguf, which no earlier comparison of ours counted. What that does to the four published claims, two against us and two for us:

  1. 6 bits, Q6_K against hydrated: our sentence understated their byte spend roughly threefold. "+1.186 GiB more file" is +1.198 on tensors and +3.634 GiB of transformer body (23.1 % more than ours). Their lower complete-pipeline KL remains measured; the engine-confounded comparison does not establish a format-only fidelity win.
  2. 6 bits, deployed: a multimodal Q6_K deployment is 22.18 GiB against our 20.13 GiB whole tree, so ours is 2.053 GiB smaller and needs no second file.
  3. 5 bits, UD-Q5_K_XL against the context edition: the old byte claim was wrong against us. We do not pay "0.445 GiB more" for the lower complete-pipeline KL — on transformer body they carry 2.148 GiB more (+14.4 %), and our deployed multimodal artifact is 0.422 GiB smaller. The KL comparison remains engine-confounded.
  4. 8 bits, Q8_0 against online K5/K6: the bodies agree to 1.7 % (24.526 against 24.119), the most byte-comparable pair on the table, and the Q8_0 complete pipeline measures lower KL. Format attribution still requires a same-engine capture.

Rule from here on, and it is printed rather than footnoted: "at equal file bytes", "at equal tensor bytes", "at equal transformer body" and "as a deployed multimodal artifact" are four different claims, and whichever one a sentence means is written into the sentence. No fidelity number changes.

  • Q8_0 has the lowest measured complete-pipeline KL, 0.001087 at 27.05 GiB. Its engine differs from every vLLM row, so no quantization-only ordering is inferred from the 0.000507 BF16 control.
  • Every GGUF point at or above 5 bits measures lower complete-pipeline KL than official FP8. This makes our "34-48 % below FP8" headline a weaker achievement than it sounds, while remaining a cross-engine observation.
  • K4 is the weakest of our own builds, worse than every GGUF measured here and worse than official FP8. The one row it beats is unsloth's NVFP4.
  • The other 4-bit-weight artifact is far behind. unsloth/Qwen3.8-27B-NVFP4 reads 0.030115 with 93.16 % top-1 on this same shard, under the same engine so with no cross-engine term: 2.9x K4 at the same weight class, 8.8x the context edition, 27.7x Q8_0. It loses 512 of 512 contexts to both the context edition and official FP8 with zero ties, and over the full ten shards reads 0.031059 [0.027916, 0.034795] at 92.90 % top-1, losing 5,120 of 5,120 (receipts/kld5-1M-nvfp4.json, receipts/kld5-10M-nvfp4.json, receipts/kld5-1M-paired-nvfp4.json). The revision we measured, 9c73e2da, no longer resolves upstream after a 2026-08-15 history squash; the weights at current HEAD are byte-identical and we keep the reviewed revision alive as malaiwah/Qwen3.8-27B-NVFP4-archival-9c73e2da.

What it does not settle: this is text-only teacher-forced fidelity on one shard — one tenth of the suite, though a close tenth: over all 10,480,640 positions the five vLLM means read 0.002760 / 0.003210 / 0.003509 / 0.005294 / 0.010604, 1.9-2.9 % above these shard-0 values with the ordering unchanged (receipts/kld5-10M-*.json), and the GGUFs have no ten-shard equivalent. It says nothing about serving 262,144 tokens with vision and MTP on a 32 GB card, which is where these artifacts actually differ, and llama.cpp KV-quant behaviour, prefill and decode speed are separate axes that were not measured. A separate control bounds the protocol objection instead of arguing it: restricting scoring to positions with at least 256 tokens of left context — the floor llama-perplexity enforces — lowers every candidate's mean by 1.3-2.1 %, and second-half-only scoring by 3.9-4.9 %, uniformly enough to change no ordering (receipts/scored-window-offset.json), so the external protocol's scoring floor explains at most about 5 % of any cross-protocol gap. Protocol identity and the remaining open comparators are in docs/35 and docs/29 F2.

Public capability, all six models on one pinned MMLU-Pro subset. 70 questions, 14 official categories x 5, official five-shot prefixes, TIGER-Lab/MMLU-Pro@b189ec765aa7ed75c8acfea42df31fdae71f97be, greedy, thinking at low effort, 5,120-token completion cap, every candidate scored item-paired against the BF16 control. This is a first public, licence-compatible, item-paired benchmark, not a leaderboard claim.

modelabsoluteWilson 95 %BF16-pass retentionWilson lowerregressionsimprovementscompletion-cap failuresreceipt
BF16 Qwen/Qwen3.8-27B57/70 (81.4 %)[70.8 %, 88.8 %]reference4receipts/public-capability-bf16.json
context edition58/70 (82.9 %)[72.4 %, 89.9 %]56/5790.7 %123receipts/public-capability-ctx.json
K457/70 (81.4 %)[70.8 %, 88.8 %]55/5788.1 %224receipts/public-capability-k4.json
hydrated56/70 (80.0 %)[69.2 %, 87.7 %]54/5785.6 %324receipts/public-capability-hyd.json
official FP8 Qwen/Qwen3.8-27B-FP856/70 (80.0 %)[69.2 %, 87.7 %]55/5788.1 %214receipts/public-capability-fp8.json
online K5/K655/70 (78.6 %)[67.6 %, 86.6 %]54/5785.6 %314receipts/public-capability-k5k6.json

The bar, and who clears it. receipts/public-capability-plan.json pre-registered two conditions before any candidate ran: BF16-pass retention with a Wilson 95 % lower bound at or above 0.90, and no category losing more than two BF16 passes. All five candidates satisfy the second — the worst per-category loss is two passes (hydrated and online K5/K6, philosophy). Only the context edition satisfies the first, at 90.7 %. K4 and official FP8 read 88.1 %; hydrated and online K5/K6 read 85.6 %. Four of five candidates therefore miss a bar written down in advance, and that is published rather than restated as a pass.

Why that is a power limit, not a verdict. With 57 BF16 passes as the paired denominator, 56/57 is the smallest count whose Wilson lower bound clears 0.90 — a single paired miss already fails it — so a 70-item suite has too few items to certify this bar at all, and every candidate interval overlaps every other. The matrix separates nothing at this size. It applies equally to official FP8, which misses the same bar; none of this is evidence that any build is broken. Exact-output agreement is 0/70 for every EXL3 candidate and 1/70 for official FP8 (a single 113-token math answer) because long chains of thought differ token-wise, so the pass/fail outcome is the only meaningful pairing. Harness tools/public_capability.py, sweep runner tools/run_public_capability.sh, superseded 2,048-cap control receipts/public-capability-bf16-superseded-cap2048.json. The honest next step is more items and more task types, the plan's own P1 — HumanEval+/MBPP-style executable cases, IFEval-style constraints, tool schemas and a larger MMLU-Pro draw — not another sweep of the same 70 questions; no capability claim graduates before that.

Three limits stated up front: v5 absolute KLD is not comparable to v3 (K4 reads 0.029679 there and 0.010604 here — only within-suite ordering and paired differences transfer); this run has no cumulative percentiles over all ten shards, only per-shard p95/p99/p999 and the exact global maximum, because its shard reports predate the qwen38-fidelity-report/2 KLD histogram — the tail table above is the shard-0 rerun, 1,048,064 of those 10,480,640 positions; and its hidden states were deleted shard by shard to fit 135 GB of scratch, so it is reproducible from the pinned corpus fetch log and suite manifest plus a GPU rather than from published captures. See PROGRESS.md for the full session record.

File map

DocContents
docs/01-nvfp4-composition.mdMeasured tensor-level composition of nvidia/Qwen3.6-27B-NVFP4 and unsloth/Qwen3.8-27B-NVFP4, plus the BF16 parameter census
docs/02-recipe-k4.mdThe proposed recipe, footprint arithmetic, and the headroom exchange rates
docs/03-gg-runtime-contract.mdWhat the Gilded Gnosis EXL3 loader requires, and what it does not support
docs/04-exllamav3-toolchain.mdexllamav3 conversion flags, the missing per-module override, and the splice route
docs/05-kld-protocol.mdSuperseded iteration-1 protocol: the single-window, full-vocabulary logits KLD inherited from the GG r34 runner. Kept for provenance; the current method is docs/42
docs/06-baseline-validation.mdRunning the official GG image with no container runtime, and the two proven baselines
docs/07-serving-recommendations.mdServing-guide differences across upstream / NVIDIA / Unsloth cards, cross-checked against shipped configs
docs/08-upstream-cards-digest.mdPer-card digest: declared recipe, benchmarks, harnesses, limitations
docs/09-variant-publication.mdHow iterative variants are published for independent re-measurement
docs/14-fidelity-protocol-v2.mdHidden-state replay protocol, supersedes the single-window KLD
docs/18-results-fidelity-v3.mdHeld-out re-measurement, the later offset-independent contamination correction, per-stratum means, and controls
docs/22-results-iteration-2.mdGate K5 / up K5 / down K6: corrected mean KLD 0.007945, 38 % below official FP8 at 71 % of its resident weight
docs/15-results-fidelity-v2.mdSuperseded v2 run: 151,478 positions, 74/74 paired wins — measured on a corpus that overlapped our calibration data
docs/16-head-attribution.mdIs lm_head the sensitive tensor? Measured: not at K6
docs/11-kld-external-comparison.mdPublished KLD data for this family; why "FP8 = 0.5" is wrong
docs/13-upstream-contributions.mdUpstream issue + verified PR
docs/19-cuda-graphs-patch.mdThe autotune-priming patch behind PR #314, deliverables and verification
docs/20-context-extension-and-k3-gap.mdHow far context can be extended, and the distance to the reference protocol
docs/12-iteration-2-plan.mdWhere to invest next, ranked
docs/17-next-iteration-shopping-list.mdIteration-2 shopping list, each item with its closing test
docs/10-results-iteration-1.mdIteration 1: build, serve, KLD, and the upstream defect
docs/21-independent-review-response.mdIndependent review, per finding: fixed, fixed-in-v2, or still open
docs/24-p0-results.mdP0 done: prefill +113 %, and why fp32 replay was a negative result
docs/23-next-attack-list.mdRanked plan for iteration 3, with evidence, cost and acceptance per item
docs/25-goal-pareto-dominate-fp8.mdthe goal as a gate, and the iteration-3 verdict per axis
docs/26-prefill-attribution.mdwhere prefill time goes: MLP 2.13-2.26x, attention overlay 1.05-1.11x, hgemm at cuBLAS parity
docs/27-graph-decode-drift-control.mdeager-vs-graph drift is ambient: BF16 drifts the same as EXL3
PROGRESS.mdChronological work log
docs/28-external-validation-and-corrections.mdexternal RTX 5090 validation and the corrected capacity boundary
docs/30-iteration-4-context-edition.mdserialized-K5 context build, receipts and rejected vision quant
docs/31-frozen-qualification.mdsource-disjoint frozen v4 qualification
docs/32-native-context-embedding-overlay.mdint8 input table, native MTP-3 plus 8.4 MP engine-budget proof (since superseded by the physical 5090 qualification at utilisation 0.955), and corrected draft accounting
docs/33-evidence-volume-and-intervals.mdthe 10,480,640-position rerun, the measured tail, and why ten times the positions did not narrow the interval
docs/29-plan-and-loose-ends.mdranked open work and acceptance gates, including the evidence-volume objection (F1) this session closed
docs/35-external-protocol-comparability.mdwhy our KLD and Unsloth's are not interchangeable, the twelve deltas, and the two of them now measured: the scoring-window offset and the cross-engine floor
docs/34-vram-class-profiles.mdthe 24 GB and 16 GB class verdicts, the measured affine KV law, and a capped-budget proxy qualification
docs/36-performance-levers-5090.md11 configurations on the physical 5090: what moved throughput, what is a no-op, and the GDN gate reachability rate
docs/37-error-driven-allocation.mdthe error-driven allocation experiment that lost, the 3.73x-per-bit law it produced, and why the objective is not a surrogate for KLD
docs/39-chat-template-audit.mdevery chat template in this model family captured with revision and sha256, diffed by named construct, and the verdict that ours is byte-identical to official; the community "fixed" template measured losing tool-call arguments under qwen3_coder, and the loudest complaint traced to the 2.4T flagship's template rather than the 27B's
docs/42-kld-method.mdthe KLD method of record: the metric, the shared BF16 head and why, the capture point, the suite and its contamination policy, the ladder, the bootstrap and the tail histogram, the harness-revision pin, how to re-derive a published number on a CPU in seconds, every floor that bounds an absolute value, and the claim boundary
docs/14-paro-assessment.mdParoQuant (Pairwise Rotation Quantization, arXiv:2511.10645) assessed: a pre-quantization transform orthogonal to EXL3, not a comparator; the checkpoint is INT4-linear and protocol-incomparable, the learned rotation could replace our fixed Hadamard — follow-up in docs/13
docs/57-eda-allocation-revisit.mdthe error-driven allocation revisit, including the later re-solve that refuted sqrt_energy and left a compounding-aware objective as the next defensible in-family method
docs/58-qwen36-quant-prior-art.mdprior art of per-module bitrate attribution in the Qwen 3.5/3.6 generation: llama.cpp use_more_bits U-shaped layer heuristic (first+last 1/8 promoted, attn_v special-cased as "more sensitive"), EXL2 Frobenius-norm simulated annealing vs EXL3 direct-KLD greedy optimizer with U-shaped allocation.py layer term, GDN-at-higher-precision as standard NVFP4 practice (driven by correctness bug not measurement), Unsloth finding attn sensitive for hybrid archs (direction matches our physics, mechanism not diagnosed), no published early-vs-late KLD experiment, and five concrete inspirations with cost and falsification

Tooling in tools/ is what produced the evidence: an unprivileged OCI image puller and a proot-based runner for it, the BF16 attention splice and checkpoint finaliser, the fidelity harness (fidelity.py, suite3.py) that builds the suite and replays captures, the decode-parity probe (decode_parity.py), the kernel microbenchmarks (prefill_micro*.py, gemm_cmp.py) and the upstream patches as standalone files. The v5 10.48 M-position run, the public-benchmark harness and the collection index are these files:

ToolWhat it doesReceipts
tools/fetch_corpus_v5.pyfetches the five-stratum v5 corpus (941 documents / 70,348,971 bytes) and pins every document by URL and sha256receipts/kld5-corpus-fetch-log.json
tools/suite3.pybuilds the frozen suite: exact-advance non-overlapping windows, whole-document calibration-overlap pre-exclusion (44 of 941), prior-suite token exclusion (160 v4 hashes, 0 reachable)receipts/kld5-suite-manifest.json (qwen38-distribution-fidelity/6)
tools/kld_ladder.shwalks the ladder one 512-context shard at a time — capture six models, replay five candidates, verify, delete the shard's ~64 GB of hidden states — because scratch is ~135 GBten per-shard reports, listed in every cumulative receipt
tools/kld_aggregate.pywelds verified per-shard reports into cumulative means, cluster bootstraps, paired comparisons and (for qwen38-fidelity-report/2 inputs) bin-bounded cumulative quantilesreceipts/kld5-10M-{hyd,k5k6,ctx,fp8,k4}.json, receipts/kld5-10M-paired.json
tools/fidelity.pycapture/replay harness; now also counts every scored position into a fixed log-spaced kld_tail histogram, bumping reports to qwen38-fidelity-report/2per-shard and per-candidate reports
tools/public_capability.pypaired MMLU-Pro run against a pinned dataset revision, official five-shot prefixes, greedy, with Wilson intervals and BF16-pass retentionreceipts/public-capability-{plan,suite-mmlupro-70,bf16,ctx,k4,hyd,fp8,k5k6}.json, superseded 2,048-cap control kept as receipts/public-capability-bf16-superseded-cap2048.json
tools/run_public_capability.shdrives the six-model sweep one server at a time against the frozen BF16 reference, with the per-candidate environment each build needs (every EXL3 candidate requires VLLM_EXL3_GRAPH_DECODE=1)the five candidate receipts above
tools/collection_index.pyone immutable row per published checkpoint, every field carrying its receipt path, sha256 and RFC 6901 pointer; disk bytes and resident_weights kept strictly distinctreceipts/collection-index.json (qwen38-collection-index/1)
tools/gguf_capture.cppllama.cpp-side capture of the post-final-norm state for a GGUF at pinned commit ece963f4, so a GGUF can be replayed through our shared BF16 head; bf16 rounding verified bit-identical to torch on 2,012,449 probe valuesreceipts/gguf-report-{q8_0,q6_k,q5_k_xl,engine-floor}.json, receipts/cross-engine-comparator.json
tools/gguf_manifest.pypins each capture to its GGUF blob digest and the llama.cpp identity that produced itthe four receipts/gguf-report-*.json
tools/build_llamacpp.shbuilds the capture binary against the pinned llama.cpp commitengine_identity in every GGUF report
tools/chat_template_matrix.pyrenders every candidate chat template over 34 fixed cases inside the pinned image, reconstructing transformers' Jinja environment exactly, and reports token counts and sha256 per rendered prompt; CPU onlyreceipts/chat-template-audit.json, receipts/chat-templates/renders.json
tools/chat_template_roundtrip.pyfeeds each template's rendered tool call back through the image's own vllm.parser.qwen3._qwen3_arg_converter and reports whether --tool-call-parser qwen3_coder recovers the arguments losslesslytool_call_roundtrip in the audit receipt
tools/chat_template_prefix.pyfirst-divergent-token between the generated token stream and the re-rendered history, per client reasoning-echo behaviour, i.e. the point past which no prefix-cache block is reusableprompt_growth_and_prefix_stability in the audit receipt
tools/fetch_chat_templates.sh, tools/gguf_chat_template.pyfetch every family template with its revision and digest, including off-default-branch refs and the one that exists only inside GGUF metadata (read over HTTP range requests, no tensor bytes)receipts/chat-templates/{SHA256SUMS,raw/}

Status

  • Both upstream Qwen3.8-27B artifacts proven runnable under the official GG image
  • Iteration 1 (K4 MLP + BF16 attention splice) served under GG with ONLINE_QUANT=exl3-b6, 19.21 GB resident, vision verified, published as malaiwah/Qwen3.8-27B-K4
  • Held-out v3 fidelity suite with a contamination scan, plus the dataset that lets anyone recompute it without a GPU
  • Iteration 2 built, measured and published as malaiwah/Qwen3.8-27B-EXL3-K5K6: 20.32 GiB resident, overlap-corrected mean KLD 0.007945 body-only / 0.008078 as served, top-1 96.86 %
  • CUDA graphs (PR #314), row-count prefill dispatch (PR #316), B12X prefill routing (PR #318) and quantized-embedding construction (PR #319) filed against the fork and exercised in the tested runtime
  • Prefill attributed and bounded: the MLP kernel is the whole story (2.13x / 2.26x), the online-K6 attention overlay is not the bottleneck (1.05x / 1.11x), ext.hgemm is already at cuBLAS parity — FP8 prefill parity needs a fused dequant-in-epilogue kernel, not tuning
  • Attention-overlay width is a runtime knob, measured on the same suite and real RTX 5090: K6 0.007945 at 20.32 GiB and ~185k context; K5 0.011801 at 19.82 GiB and 206,400 retrieval-verified; K4 0.026619 at 19.05 GiB
  • Serializing attention offline instead of encoding it at load — the "hydrated" build — measures 0.007172, i.e. calibrated offline K6 beats the runtime overlay by 0.000773 on the overlap-corrected subset
  • Graph-vs-eager decode drift measured properly and traced to the build, not to EXL3 (docs/27)
  • Frozen v4 suite: 160 source-disjoint contexts, 100 documents, zero token/document/content overlap; all five candidate capture sets and qualification receipts published (2,708 files / 51.0 GB)
  • Context edition plus per-row int8 input table starts native 262,144 with MTP-3, decode graphs and an 8.4 MP image cap under a 30.24 GiB engine budget; 266,612 KV tokens, exact retrieval at 261,794 text tokens and in a 236,824-token seven-megapixel request — that engine-budget run is now superseded by the physical-card qualification below
  • Hard-limit qualification on a physical RTX 5090 closed (receipts/qualification-5090-context.json, per-process server logs receipts/qualification-5090-context-server-{B3,B4,C,D,E,F}.log): the context edition serves native 262,144 with MTP-3, the full 8,388,608-pixel image ceiling and fp8 KV on one 32,607 MiB card — but at --gpu-memory-utilization 0.955 with PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, not the 0.97 the cards used to print. Measured at startup: engine budget 29.98 GiB, free 30.9 of 31.4 GiB usable, usage 18.19 weight + 1.78 peak activation + 0.27 non-torch + 0.45 CUDAGraph = 20.69 GiB, available KV 9.28 GiB = 265,122 KV tokens, 1.01x maximum concurrency at 262,144, attention block size forced to 1600 tokens so the attention page is at least the mamba page, 3 padding layers at at most 6.25 % KV waste, startup 55.7 s, model load 18.19 GiB in 3.99 s. All seven gates PASS: (1) startup allocates a native-length request inside the utilisation ceiling; (2) the 261,794-token needle retrieved exactly; (3) a combined 236,824-token plus 7,077,888-pixel request returns 1376346594 | red, blue exactly; (4) the 30-case image suite at 24/30; (5) three warmed 256-token concurrency-1 decode runs; (6) a second native-length request after release in the same process; (7) receipt identity complete
  • 0.97 published as a bounded negative, with its mechanism: at 0.97 the combined text-plus-image request dies with torch.OutOfMemoryError inside vllm/v1/attention/ops/vit_attn_wrappers.py wanting 62.00 MiB with 26.50 MiB free, and expandable_segments does not rescue it because vLLM spends the freed bytes on more KV (272,570 → 280,017 tokens); lowering max_pixels to 4,194,304 at 0.97 is strictly worse (profiled peak activation falls to 1.35 GiB, KV grows to 291,933 tokens, OOM with 6.56 MiB free). The knob is utilisation, not the image ceiling — seven megapixels is not a hard 32 GB ceiling, since the identical request succeeds at 0.955 with the full 8,388,608-pixel ceiling
  • Task-retention smoke: BF16, all four EXL3 profiles and official FP8 each pass 40/40 deterministic paired tasks with zero regressions; a hardened rescore leaves every pass unchanged and corrects exact-final-answer agreement to 32/40–35/40 (receipts/task-retention-v2-summary.json, receipts/task-retention-v2-strict-rescore.json)
  • Evidence volume objection (F1) closed: fidelity re-measured on the held-out v5 suite at 10,480,640 scored positions / 5,120 contexts / 842 source clusters, ten verified shards welded by tools/kld_aggregate.py — hydrated 0.002760, online K5/K6 0.003210, context 0.003509, official FP8 0.005294, K4 0.010604, all body-only through one shared BF16 head (receipts/kld5-10M-*.json)
  • Paired at 5,120 contexts: hydrated -0.002534 (5,118 wins), online K5/K6 -0.002084 (5,105), context -0.001785 (5,109) against official FP8; hydrated beats the runtime overlay by -0.000450 (4,922); K4 loses to FP8 on 5,113 of 5,120 (receipts/kld5-10M-paired.json)
  • Ladder stability shown, not assumed: cumulative hydrated mean 0.002700 / 0.002759 / 0.002699 / 0.002760 at the 1M / 2M / 5M / 10M checkpoints
  • Exact tail now aggregable for future runs: a fixed log-spaced KLD histogram in tools/fidelity.py (qwen38-fidelity-report/2) plus bin-bounded cumulative quantiles in tools/kld_aggregate.py; this run kept only per-shard percentiles and the exact global maximum, and says so in its receipts
  • First public-benchmark matrix, all six models: paired MMLU-Pro (70 questions, pinned TIGER-Lab/MMLU-Pro@b189ec76) — BF16 57/70, context 58/70 at 56/57 retention (Wilson lower 90.7 %, the only candidate clearing the pre-registered 0.90 bar), K4 57/70 and official FP8 56/70 at 55/57 (88.1 %), hydrated 56/70 and online K5/K6 55/70 at 54/57 (85.6 %); the four shortfalls are a 70-item power limit — 56/57 is the smallest count that can clear 0.90 — not a broken build, so the next step is item volume and task diversity
  • Collection index published: receipts/collection-index.json from tools/collection_index.py, four immutable rows, receipt-traced fields, one disclosed divergence left unfixed rather than rewriting a published receipt
  • Comparator breadth (F2) closed by measurement: GGUF Q8_0 0.001087, Q6_K 0.002035 and UD-Q5_K_XL 0.004444 were scored on shard 0 of the v5 suite through the shared BF16 head. The unquantized llama.cpp-versus-vLLM control is 0.000507 and is diagnostic, not subtractable; cross-engine rows therefore compare complete pipelines and do not isolate quantization format (receipts/cross-engine-comparator.json)
  • The largest cross-protocol objection bounded instead of argued: restricting scoring to their 256-token left-context floor lowers every candidate's mean by 1.3-2.1 % and second-half-only by 3.9-4.9 %, uniformly, changing no ordering (receipts/scored-window-offset.json)
  • Their protocol run for cross-citation: llama-perplexity --kl-divergence on wikitext-2 raw at ctx 512, 147,900 scored positions, base PPL 6.950230 ± 0.044933 — Q8_0 0.000926, Q6_K 0.002286, UD-Q5_K_XL 0.004426, with the same ordering as our suite. The absolute values and ratios are protocol-specific; their harness floor was measured on our hardware at 5.6e-5 to 8.0e-5 and tokenisation was bit-identical over 297,194 tokens. PPL does not reproduce the KLD ordering (receipts/wikitext-kld-run-a.json, docs/35)
  • unsloth/Qwen3.8-27B-NVFP4 measured on our suite, the comparator readers ask for most: 0.031059 [0.027916, 0.034795] at 92.90 % top-1 over 10,480,640 positions, losing 5,120 of 5,120 contexts to both the context edition and official FP8; 2.9x K4 at the same 4-bit weight class (receipts/kld5-10M-nvfp4.json, receipts/kld5-1M-paired-nvfp4.json)
  • Capture and replay are bit-reproducible: 5,240,320 scored positions identical across independent runs, separate model loads and three harness generations, under eager single-sequence capture — which makes every published KLD number re-derivable rather than merely re-runnable (receipts/capture-determinism.json)
  • The converter is not deterministic, and it does not matter for fidelity. A fresh conversion of the published hydrated recipe returns 13 of 16 pinned files identical and three safetensors shards differing in tensor bytes only — 399 of 409 quantized modules, 97.6 %, at 41-92 % of bytes each — with identical headers, offsets, shard sizes and per-role byte totals. That sibling then scored -3.755e-06 paired against the published checkpoint, 95 % [-2.854e-05, +2.062e-05] bracketing zero, 257 contexts to 255. So the recipe determines fidelity to within this protocol's resolution and the published bytes are the artifact (receipts/converter-determinism.json, receipts/sibling-rebuild-fidelity.json)
  • Performance levers measured on the physical 5090, 11 configurations: MTP depth is a concurrency decision — depth 3 wins at one stream (83.31 vs 74.28 per-request tok/s) and num_speculative_tokens 1 wins at eight (+30.67 %, 313.28 → 409.35 tok/s aggregate, plus 10,911 more KV tokens). Closed as no-ops or losses: explicit FLASHINFER (already auto-selected on SM120), TRITON_ATTN, custom_ops:["all"], and dynamic speculative decoding (refuses to start under FULL_DECODE_ONLY). Prefill did not move on any lever (docs/36)
  • Prefix caching shipped, scoped by measurement: on the three 8,192-token recipes it is worth 11.6x and 29.3x warm TTFT at 32 k and 131 k prefixes (whole schedule 144.4 s → 84.0 s); at the context edition's native 262,144 it is declined because align-mode block rounding makes the window unservable three different ways, with a measured 256,000-token option priced at 6,144 tokens. The four-module image gg-r34-patched-apc is the release unit (receipts/apc-poison-repro.json, receipts/qualification-5090-apc.json)
  • VRAM classes decided: 24 GB is a GO as a serving profile over the published context edition at 24,576 tokens with MTP-3 or 45,056 with MTP off, validated by a capped-budget proxy with the physical-board gate left open; 16 GB is a NO-GO as a SKU and published as a design study. The per-token KV model was replaced by the measured affine law a*L + M (docs/34)
  • A user-reported defect investigated to a measured negative: 266 requests across seven freshly started servers reproduced no corruption under the reporter's exact condition, which decomposed his four-variable control and left LMCache as the sole suspect; the two absent upstream fixes were requested at the fork (issue #392, PR #393) and a net-new admission livelock was filed upstream as vllm-project/vllm#52520
  • Materials published for third-party replay: the v5 suite, all ten shard views, the shard-0 BF16 reference capture and 79 per-shard reports as malaiwah/qwen38-27b-fidelity-suite-v5, the losing error-driven candidate as a documented negative, and archival mirrors of the NVFP4 and GGUF revisions our published numbers cite — because one of them stopped resolving upstream mid-session
  • Still open in F2: stock uniform-bitrate EXL3 controls, and the reader-suggested ik_llama (IQ4_KS_KT) and NVFP4-NInfer comparators — neither of which has any KLD number today (docs/29 F2)

Contributors

malaiwah

559 commits

malaiwah/qwen38-27b-exl3

EXL3 mixed-precision quantization of Qwen3.8-27B: measured recipes, runtime contract, KLD receipts, iteration log

1

stars

559

commits

Python

primary language

Aug 26, 2026

updated

README

Status 2026-08-16: the collection is hardware-qualified, the comparator set is measured, and the method is now reproducible from published artifacts. The current checkpoint is malaiwah/Qwen3.8-27B-EXL3-K5K6: mean KLD 0.003210 (95 % CI [0.002982, 0.003480]) body-only at 20.32 GiB resident weights, versus official Qwen/Qwen3.8-27B-FP8 at 0.00529439 % lower divergence at 71 % of its resident weight, paired -0.002084 [-0.002249, -0.001942] winning 5,105 of 5,120 held-out contexts. The best profile in the collection is the hydrated build at 0.002760, paired -0.002534 on 5,118/5,120. The headline suite is v5: 5,120 contexts x 2,047 positions = 10,480,640 scored positions over 842 source clusters, calibration- and v4-token-disjoint.

Landed since: the context edition is qualified on a physical RTX 5090 (all seven gates at --gpu-memory-utilization 0.955 with expandable_segments, 265,122 KV tokens at 262,144 — and a second physical 5090 needed 0.956, so utilisation is a per-card measurement, not a constant). unsloth's NVFP4 is now on our suite at 0.031059 over 10.48 M positions, losing 5,120 of 5,120 contexts to both the context edition and FP8. The external llama-perplexity protocol was run for cross-citation and reproduces our ordering. Capture and replay are bit-reproducible, and a fresh conversion of the same recipe — a sibling, 97.6 % of quantized modules differing in bytes — scores indistinguishable from the published checkpoint, so the recipe determines fidelity while the published bytes remain the artifact. The suite, the shared BF16 reference capture and every per-shard report are published as a dataset, with archival mirrors of the third-party artifacts our numbers cite.

Two more artifacts published 2026-08-16, both from pre-registered conversions. malaiwah/Qwen3.8-27B-EXL3-K6-parity uses nearly the same file-byte budget as GGUF Q6_K by promoting gate_proj/up_proj K5 → K6. It measures 0.001634 [0.001541, 0.001742] under vLLM versus Q6_K at 0.002035 under llama.cpp, while beating hydrated on 511/512 contexts (-39.5 % for +1.348 GiB) and carrying 2.31 GiB less transformer body than Q6_K. The cross-engine control proves these are complete-pipeline results; it cannot be subtracted to establish format parity or prove that the earlier gap was caused by bytes. malaiwah/Qwen3.8-27B-EXL3-S16-V-research is the opposite: the rejected sub-4-bit 16 GB candidate at 0.045374 [0.041959, 0.049351], 1.5x its pre-registered NO threshold, losing 512/512 contexts to K4 (4.39x) and to the context edition (13.31x). It is published because it failed - a NO nobody can audit is an assertion - and it is the empirical anchor for the K3 rung of the per-bit law (docs/37 §3.3). Both payloads were predicted to the byte by the affine law before conversion, and both registered intervals are reported with their misses: S16-V's point estimate was 1.52x too high, K6-parity's interval fell 2.0 % short.

The earlier headline — 0.007945 body-only on 136 v3 contexts / 278,392 positions (docs/22) — stands exactly as measured and is superseded only as the headline; v5 absolute KLD is not comparable to v3 (the two suites differ by a measured 2.46-3.08x with identical ordering), so quote paired differences across suites, never means. Figures labelled K4 or v1 belong to iteration 1 and are superseded. One earlier control was withdrawn: the "CUDA-graph parity 0.000000" receipt captured a prefill forward, which FULL_DECODE_ONLY never captures, so it could not have measured decode; the decode probe that replaces it is docs/27. Open items are tracked in docs/29.

Fixed optimization objective (frozen 2026-08-21)

The primary control is the exact PROFILE=fidelity tuple, not PROFILE=throughput and not a memory-modified derivative. The objective is to recover enough PP/TTFT from that control to pass the absolute candidate gates without sacrificing its paired KLD/tails, TG, context, or capability envelope. The fidelity control is allowed to fail the 7,000 tok/s PP candidate threshold; that threshold governs promotion of a new candidate, not acceptance of the control.

receipts/frontier-fidelity-control.json pins the model/profile identity, INT6 embedding overlay, exact evidence hashes, reference metrics, candidate rules, and the prohibition on substituting the high-KLD throughput profile. Every INT8/FP4/FP6, K-map, topology, kernel-route, KV, MTP, graph, or profile change is a candidate variable and requires preregistered paired comparison against this exact control.

Frontier G02 pre-candidate controls (measured 2026-08-21)

The exact control and method-neutral harness are ready. Clean packed INT6 was integrated onto runtime b19029d, matched the historical encoder byte-for-byte on real BF16 rows, and passed three independent AIBoss boots. Aggregate PP was 2,941.9 tok/s, fox/essay TG 226.9/104.3, context 238,400, with text, vision, MTP, 200k context and rollback passing on every boot.

The clean-runtime fidelity recapture is exactly equal to the historical control: mean KLD 0.0034052972, p99 0.0348892, paired delta and CI 0.0, with all 512 per-context rows and the full tail histogram identical. Its 513-file, 10.73 GB hidden-state capture has an independently verified private-bucket copy.

The repaired v3 harness uses the actual EXL3 Viterbi quantizer—not the invalidated uniform proxy—on four real 128×128 Qwen slices with matched BF16/quant-flow activations and Hessians. RTN and both LDLQ controls carry the same exact 10,756-byte / 5.251953-bpw payload. Quant-flow-H LDLQ reduced mean excess running-output error from RTN's 0.0011858 to 0.0008865 at those same bytes; this is a control result, not a method promotion. Whole-block and end-logit gates remain mandatory and intentionally pending for researcher-proposed methods.

receipts/frontier-g02/manifest.json is validated by python3 tools/verify_frontier_g02.py --campaign receipts/frontier-g02 as pre_candidate_ready. Candidate slots remain unopened, E-final-v7 remains sealed, and rental authorization remains false.

Final Frontier G0/G1 campaign (measured 2026-08-20)

The preregistered campaign in docs/61-final-frontier-requant-stack-plan.md ended at its local RTX 5090 gate with the terminal disposition three_candidate_no_go. The complete hashed record is receipts/frontier-g01/manifest.json; python3 tools/verify_frontier_g01.py --campaign receipts/frontier-g01 validates every referenced G0 artifact, preregistration, result and terminal decision and exits 0.

G0 passed: the immutable BF16 payload was fully censused (1,199 logical tensors), converter/runtime environments were locked, the fidelity-derived incumbent-fidelity-int8 clean runtime was qualified, and every AIBoss maintenance transaction proved service restoration. That historical G0 arm used INT8 embeddings and did not remeasure the exact PROFILE=fidelity KLD; it is runtime/rollback evidence, not a substitute for the frozen primary control. G1 then consumed the three frozen candidate opportunities:

candidatemeasured resultfrozen Gate A failure
full gate/up FP6startup rejected the 199,104-token configuration; estimated maximum 192,576context below 238,400
full gate/up FP4PP 4,393.6 tok/s; fox/essay TG 211.6/98.8; context and capability checks passedPP below 7,000
gate/up FP4, language layers 0–31PP 3,529.5 tok/s; fox/essay TG 215.9/96.0; context and capability checks passedPP below 7,000

The preregistered hard-gate rule stopped KLD scoring after each deterministic runtime failure. Supporting experiments also closed without a promotable candidate: independent seeded converter A/A runs had 24 payload mismatches; the pinned AutoRound consumer could not build; ResComp improved one actual 16×16 weight-tile SSE by 21.1% but lacked a valid whole-block implementation; and the hybrid QKV producer/consumer pilot passed its component checks without the mandatory full-checkpoint, startup, graph and logit gates. No paid rental or new model release is authorized by this campaign; the measured serving profiles below remain current.

Serving on one RTX 5090 (measured 2026-08-19)

Speed vs fidelity

What each profile delivers

Three profiles, selected by PROFILE= in patches/run-qwen38-27b.sh. All three pass tools/verify-profile.sh (exit 0, 8/8 checks each, including a 200k-token prompt). PP/TG are the tools/bench-profile.sh harness at n=3 boots; KLD is 512 contexts of shard-0000 against the BF16 reference.

PROFILE=throughputPROFILE=fidelityPROFILE=balanced
weightsall-FP4all-trellis (K5K6 as shipped)trellis + gate_up MXFP6
PP, 2051-tok9638.9 ± 18.3 tok/s2987.7 ± 4.4 tok/s3925.2 ± 13.1 tok/s
TG fox / essay187.4 ± 0.6 / 94.3 ± 0.0228.3 ± 0.4 / 104.1 ± 0.1215.6 ± 0.2 / 103.7 ± 0.1
MTP acceptance fox / essay0.930 / 0.2981.000 / 0.3041.000 / 0.324
KLD mean0.0637590.003405 [0.003166, 0.003672]0.005672 [0.005302, 0.006087]
KLD p990.70100.0348890.059908
max context249,600238,400199,104
vision + MTPpasspasspass
criteria met4/6 (fails KLD)5/6 (fails PP)5/6 (fails ctx)

fidelity serves the checkpoint at KLD 0.003405 — within 26% of this collection's own published trellis fidelity (0.002700) — with TG 228.3 tok/s and full 238,400 context; the residual over the checkpoint is the int6 embedding table (~0.0007), not any GEMM approximation. It costs ~3.2x prefill. throughput is the only profile above 7000 tok/s prefill.

Long-context needle retrieval on the fidelity profile (fixed harness, single needle per context): 8/8 at 2k, 8/8 at 100k, 8/8 at 195k — 24/24. This is an easy proxy — finding a planted needle does not establish unimpaired long-context reasoning, only that the attention window and KV cache are intact to 195k.

No single profile meets all six north-star criteria. Concretely: throughput fails KLD (0.063759, well above the 0.012 budget); fidelity fails the 7000 tok/s prefill criterion (2987.7); balanced fails the context criterion (199,104 tokens, below the 238,400 that fidelity achieves). The other two profiles each clear five of six.

receipts/frontier-2026-08-19.md shows why from three directions: prefill-grade throughput needs the MLP resident in a GEMM-ready format, trellis-grade fidelity needs weights that are decoded per prefill chunk, and GEMM-resident FP8 for all 24.3e9 quantized parameters is 22.6 GiB — which cannot coexist with a 238,400-token KV cache in 31.4 GiB. The blocker is memory, not kernel quality.

Fixes landed while getting here, with receipts: an engine-fatal OOM on any prompt over ~4k tokens (max_num_batched_tokens 8192 → 3072, which also freed 0.93 GiB of KV), +32% TG by discovering the MTP draft loop ran eager on the V1 model runner, and the first profile of this stack (prefill is 92% GPU-bound; decode was 2%). Upstream: vllm-project/vllm#52871, #52872, local-inference-lab/vllm#439, #440, b12x #232/#233/#234.

Reproducible container image

A baked serving image is published at docker.io/malaiwah/qwen38-27b-exl3-gg:r34-p2-41a5d16 (also tagged :latest). It layers every patch this project bind-mounts on top of the digest-pinned Gilded Gnosis r34 base (voipmonitor/vllm:gilded-gnosis-v20-vllm4d006a4-b12xcd3ce19-fi1ac6942-cu132-20260810-r34, sha256 820181fbbc975cd5291c411cda9771d58fecee1636d916f508f47230df20592b), so the repo is not required at runtimepodman run plus a Hugging Face cache mount is sufficient. All 10 patches (7 vLLM source overlays and 3 conversion modules) are baked in by docker/Containerfile, which also bakes the CUDA lib64 symlink the B12X JIT needs as real links and carries OCI provenance labels (org.opencontainers.image.source, .revision, .description, ai.malaiwah.base-digest).

The image was gated 9/9 mount-free (NO_PATCH_MOUNTS=1, every patch bind-mount disabled): PP 2933.0, fox 229.1, essay 103.7, ctx 238,400, 200k-token prompt OK, vision OK, MTP acceptance_fox 1.000 — confirming the image is self-contained.

PROFILE=throughput|balanced|fidelity selects the serving profile; NO_PATCH_MOUNTS=1 is what makes the baked image run mount-free. The full flag set is in patches/run-qwen38-27b.sh. A standalone fidelity invocation — every flag below is taken from that script — is:

podman run --replace -d \
  --name qwen38-27b \
  --device nvidia.com/gpu=all --ipc=host --network host \
  --tmpfs /usr/local/cuda-13.2/lib64:rw,size=16m \
  -e HF_HUB_OFFLINE=1 \
  -e CUTE_DSL_ARCH=sm_120a -e FLASHINFER_CUDA_ARCH_LIST=12.0f \
  -e OMP_NUM_THREADS=8 -e CUDA_DEVICE_MAX_CONNECTIONS=32 \
  -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True -e SAFETENSORS_FAST_GPU=1 \
  -e VLLM_WORKER_MULTIPROC_METHOD=spawn \
  -e VLLM_USE_V2_MODEL_RUNNER=1 -e VLLM_NO_USAGE_STATS=1 -e DO_NOT_TRACK=1 \
  -e XDG_CACHE_HOME=/cache/jit -e CUDA_CACHE_PATH=/cache/jit \
  -e TRITON_CACHE_DIR=/cache/jit/triton -e TORCHINDUCTOR_CACHE_DIR=/cache/jit/torchinductor \
  -e FLASHINFER_WORKSPACE_BASE=/cache/jit/flashinfer \
  -e VLLM_EXL3_MULTIPRECISION=1 -e VLLM_EXL3_GRAPH_DECODE=1 \
  -e VLLM_EXL3_FP4_TRITON_DECODE=0 -e VLLM_EXL3_FP4_PER_ROW_GS=0 -e VLLM_EXL3_FP4_DRAFT_HEAD=0 \
  -e VLLM_EXL3_ONLINE_TRELLIS_BITS=6 \
  -e VLLM_EXL3_ONLINE_CACHE_DIR=/cache/jit/exl3-online -e VLLM_EXL3_ONLINE_CACHE_MODE=readwrite \
  -e VLLM_EXL3_EMBED_ONLINE_BITS=6 -e B12X_PACKED_B_MIN_N=1024 \
  -e VLLM_EXL3_B12X_ANY_BITS=1 -e VLLM_EXL3_B12X_MIN_M=128 -e VLLM_EXL3_SKIP_TRELLIS_PREP=0 \
  -e VLLM_EXL3_PREFILL_RECONSTRUCT_M=1 -e VLLM_EXL3_PREFILL_RECONSTRUCT_MAX_MB=512 \
  -e VLLM_EXL3_PREFILL_RECONSTRUCT_CACHE=0 -e VLLM_EXL3_FOLD_FP32_BUDGET_MB=48 \
  -e VLLM_EXL3_FP4_LAYERS=, \
  -e VLLM_EXL3_EXT_PATH=/opt/exllamav3 \
  -e HF_HOME=/root/.cache/huggingface \
  -v ~/.cache/huggingface/hub:/root/.cache/huggingface:ro \
  -v ~/.cache/jit:/cache/jit \
  --entrypoint /bin/bash \
  docker.io/malaiwah/qwen38-27b-exl3-gg:r34-p2-41a5d16 \
  -lc 'set -euo pipefail; \
       ln -sf /usr/local/cuda-13.2/targets/x86_64-linux/lib/* /usr/local/cuda-13.2/lib64/ 2>/dev/null || true; \
       exec vllm serve \
         /root/.cache/huggingface/models--malaiwah--Qwen3.8-27B-EXL3-K5K6-hydrated/snapshots/ab3a91a13813df8096cb4c1d560ed3669035d0cf \
         --served-model-name Qwen3.8-27B --trust-remote-code \
         --host 0.0.0.0 --port 8000 \
         --quantization exl3 \
         --quantization-config "{\"linear\":{\"weight\":\"mxfp8\"},\"ignore\":[\"re:.*visual\\..*\",\"re:.*in_proj_a$\",\"re:.*in_proj_b$\",\"re:.*in_proj_ba$\",\"re:.*mtp\\..*\",\"lm_head\"]}" \
         --attention-backend TRITON_ATTN \
         --gpu-memory-utilization 0.945 --kv-cache-dtype fp8_e4m3 \
         --max-model-len 238400 --max-num-seqs 4 --max-num-batched-tokens 3072 \
         --compilation-config "{\"mode\":\"NONE\",\"cudagraph_mode\":\"FULL_DECODE_ONLY\"}" \
         --mm-processor-kwargs "{\"truncation\":false}" --mm-processor-cache-type shm \
         --default-chat-template-kwargs "{\"preserve_thinking\": true}" \
         --enable-chunked-prefill --reasoning-parser qwen3 \
         --enable-auto-tool-choice --tool-call-parser qwen3_coder \
         --speculative-config "{\"method\":\"mtp\",\"num_speculative_tokens\":6}"'

Flags taken from run-qwen38-27b.sh: --tmpfs + ln -sf symlink (B12X JIT workaround, lines 225/296 — the baked image already carries these symlinks, so the --tmpfs is redundant with the baked image but harmless), --device / --ipc / --network (line 226), the -e environment block (lines 231–280), the HF cache mount -v …/huggingface:ro (line 281), the JIT cache volume -v …/cache/jit:/cache/jit (line 288), and the vllm serve argument list (lines 303–322). For balanced or throughput, change the profile env vars per the case block in the script (lines 83–154).

Qwen3.8-27B EXL3 mixed-precision quants (K4, EXL3-K5K6)

Research materials and progress log for building a dense EXL3 quant of Qwen/Qwen3.8-27B that is inspired by NVIDIA's NVFP4 recipe for the previous generation, but keeps the attention projections BF16 on disk for runtime K6 encoding by the Gilded Gnosis vLLM fork, and serializes the MLP as EXL3 K5 gate / K5 up / K6 down with a K6 mcg head and a quantized MTP draft. Iteration 1 serialized the whole MLP at K4.

Design goal: NVFP4-class VRAM footprint, lower KLD. Met on fidelity, memory, decode and speculative decode; prefill is a measured structural deficit of a 4-bit-class trellis format in this runtime (docs/25, docs/26).

Evidence at a glance

Held-out v5 suite, receipts/kld5-suite-manifest.json (schema qwen38-distribution-fidelity/6, suite_token_sha256 510541f6…09482b88): 5,120 contexts x 2,047 positions = 10,480,640 scored positions, 842 source clusters, 1,024 contexts per stratum. 44 of 941 corpus documents were excluded whole before window selection for any all-position 12-token overlap with exllamav3 calibration data, leaving 897 eligible, so contamination hits are 0 by construction; all 160 v4 context token hashes were seeded as exclusions and 0 were reachable. Every figure below is body-only — both operands scored through one shared BF16 LM head — with a 95 % CI from a 10,000-resample bootstrap over the 842 source clusters.

candidatemean KLD95 % CItop-1paired vs FP8contexts wonreceipt
hydrated0.002760[0.002540, 0.003020]97.70 %-0.0025345,118 / 5,120receipts/kld5-10M-hyd.json
online K5/K60.003210[0.002982, 0.003480]97.52 %-0.0020845,105 / 5,120receipts/kld5-10M-k5k6.json
context0.003509[0.003220, 0.003852]97.44 %-0.0017855,109 / 5,120receipts/kld5-10M-ctx.json
official FP80.005294[0.004927, 0.005728]96.79 %receipts/kld5-10M-fp8.json
K4 (iteration 1)0.010604[0.009640, 0.011746]95.76 %+0.0053107 / 5,120receipts/kld5-10M-k4.json

Paired intervals and win counts come from receipts/kld5-10M-paired.json (10,000 resamples, seed 1); hydrated also beats the runtime overlay directly, -0.000450 [-0.000469, -0.000433] on 4,922/5,120. Exact worst single position over the whole run: hydrated 8.258, online 22.241, context 5.557, FP8 10.714, K4 14.283. The cumulative hydrated mean at the 1M / 2M / 5M / 10M ladder checkpoints is 0.002700 / 0.002759 / 0.002699 / 0.002760 (shards[] of receipts/kld5-10M-hyd.json).

The tail, not just the mean. The ten-shard run predates the histogram, so the tail was measured by re-running shard 0 of the same v5 suite with the qwen38-fidelity-report/2 harness: 512 contexts x 2,047 positions = 1,048,064 scored positions, the same contexts every candidate saw. Quantiles are bin-bounded — 560 log-spaced bins from 1e-12 to 1e2 nats, one bin ~5.6 % wide, and every receipt carries lower / upper / estimate per quantile — while the maxima and the exceedance counts are exact. The two columns worth reading are p99.9 and the share of positions above 0.1 nats.

candidatemeanp50p95p99p99.9p99.99exact maxabove 0.1above 1.0receipt
hydrated0.0027000.001090.00820.02760.13190.4633.7350.1534 %0.00219 %receipts/kld5-1M-tail-hyd.json
online K5/K60.0031410.001280.00990.03210.14460.4985.5070.1820 %0.00200 %receipts/kld5-1M-tail-k5k6.json
context0.0034090.001350.01070.03570.16420.5873.7490.2287 %0.00305 %receipts/kld5-1M-tail-ctx.json
official FP80.0051970.002020.01670.05310.24380.8125.2960.3912 %0.00592 %receipts/kld5-1M-tail-fp8.json
K40.0103450.003200.03320.11940.55551.8707.5651.2604 %0.03807 %receipts/kld5-1M-tail-k4.json

The ordering at p50, p95, p99, p99.9 and p99.99 is the same as the ordering of the means, so for these candidates the mean is not hiding a worse tail: every EXL3 K5/K6-class build has a lighter tail than official FP8 at every measured quantile, and K4 is worse than FP8 at every quantile. Receipts are schema qwen38-kld-ladder-cumulative/2, welded by tools/kld_aggregate.py from /2 replay reports.

Against GGUF, measured, including the part that goes against us. The standing critique is that official FP8 is a throughput format whose quality is Q4-to-Q5 class, so beating it is a weak claim, and that Q8_0 and Q6_K are the honest bar. That is now measured rather than argued. Three GGUFs from unsloth/Qwen3.8-27B-GGUF@f1bfb127c64f7072bdd2cad55f258b9c8b2910fe were captured under llama.cpp pinned at commit ece963f41b0b02d7a0d61436ae365762c073a4c8 with tools/gguf_capture.cpp, which reads the post-final-norm state — the same mathematical point our vLLM hook takes, with bf16 rounding verified bit-identical to torch on 2,012,449 probe values — and scored through the same shared BF16 head on shard 0 of the v5 suite: the same 512 contexts and the same 1,048,064 positions every candidate above saw (tools/gguf_manifest.py, tools/build_llamacpp.sh; receipt receipts/cross-engine-comparator.json, per-candidate reports receipts/gguf-report-{q8_0,q6_k,q5_k_xl}.json).

candidateenginemeasured mean KLDtop-1p99.9serialized
GGUF Q8_0llama.cpp0.00108798.53 %0.035127.05 GiB
GGUF Q6_Kllama.cpp0.00203597.98 %0.079421.31 GiB
hydrated EXL3vLLM0.00270097.80 %0.131320.12 GiB payload
online K5/K6 EXL3vLLM0.00314197.61 %0.1447
context EXL3vLLM0.00340997.55 %0.163219.27 GiB payload
GGUF UD-Q5_K_XLllama.cpp0.00444497.20 %0.214418.83 GiB
official FP8vLLM0.00519796.92 %0.244028.51 GiB resident
K4 EXL3vLLM0.01034595.91 %0.5576
unsloth/Qwen3.8-27B-NVFP4vLLM0.03011593.16 %1.622822.57 GB checkpoint

The cross-engine BF16 control is measured the same way, not assumed: llama.cpp against the vLLM BF16 reference on identical tokens, the shared head and the same 512 contexts reads 0.000507 mean, 99.07 % top-1 and p99.9 0.0113 (receipts/gguf-report-engine-floor.json). It proves the engine is a confounder. KL is neither additive nor a metric, so this value cannot be subtracted from a candidate KL and provides no quantization-only upper or lower bound. The two cells are builds that ship BF16 attention for the runtime to encode at load, so their disk bytes are not a like-for-like payload; payload figures are immutable_payload_bytes from receipts/collection-index.json (hydrated 21,610,916,123 B = 20.127 GiB, context 20,696,033,532 B = 19.275 GiB; the table truncates both to two decimals) and are serialized bytes, never VRAM. The FP8 figure is resident weights and is labelled as such.

The p99.9 column is each report's exact shard-0 p99.9 as the comparator receipt read them; the tail table above quotes the bin-bounded cumulative estimate from the 560-bin histogram (bins about 5.6 % wide), which is why hydrated reads 0.1319 there and 0.1313 here — each exact value lies inside the bin its estimate names. The two differ by construction, not by measurement.

What the table supports is complete-pipeline ordering, not format attribution:

  • At the nominal 6-bit point, llama.cpp Q6_K measures 0.002035 and vLLM hydrated measures 0.002700 on the same contexts.
  • At the nominal 5-bit point, vLLM context measures 0.003409 and llama.cpp UD-Q5_K_XL measures 0.004444.

Both comparisons are cross-engine. They are evidence about the tested artifact-plus-engine pipelines; neither isolates the quantization format.

Correction, 2026-08-16 — the byte axis in the table above is not one axis, and every mixed comparison flattered us. A GGUF row is the whole file of a text-only artifact; our row is tensor payload of a multimodal tree that also carries an MTP draft. Read from each artifact's own tensor table, without downloading any payload (cross-candidate-byte-accounting.json):

candidatefiletensor totaltoken_embdoutputtransformer bodymultimodal deployed
GGUF Q8_027.0527.041.258 (Q8_0)1.258 (Q8_0)24.52627.92 (+ mmproj-BF16 0.867)
GGUF Q6_K21.3121.300.971 (Q6_K)0.971 (Q6_K)19.36022.18 (+ mmproj-BF16 0.867)
GGUF UD-Q5_K_XL18.8318.820.814 (Q5_K)0.971 (Q6_K)17.03419.70 (+ mmproj-BF16 0.867)
online K5/K628.5028.472.368 (BF16)0.889 (K6)24.11928.50, vision inside
K426.3726.372.368 (BF16)0.593 (K4)21.46326.37, vision inside
hydrated K5/K620.1320.102.368 (BF16)0.889 (K6)15.72620.13, vision inside
context edition19.2719.252.368 (BF16)0.889 (K6)14.88619.27, vision inside

All figures GiB. The embedding and head widths are not uniform across GGUF tiers, and the vision encoder is absent from every GGUF text file - it ships separately as mmproj-BF16.gguf, which no earlier comparison of ours counted. What that does to the four published claims, two against us and two for us:

  1. 6 bits, Q6_K against hydrated: our sentence understated their byte spend roughly threefold. "+1.186 GiB more file" is +1.198 on tensors and +3.634 GiB of transformer body (23.1 % more than ours). Their lower complete-pipeline KL remains measured; the engine-confounded comparison does not establish a format-only fidelity win.
  2. 6 bits, deployed: a multimodal Q6_K deployment is 22.18 GiB against our 20.13 GiB whole tree, so ours is 2.053 GiB smaller and needs no second file.
  3. 5 bits, UD-Q5_K_XL against the context edition: the old byte claim was wrong against us. We do not pay "0.445 GiB more" for the lower complete-pipeline KL — on transformer body they carry 2.148 GiB more (+14.4 %), and our deployed multimodal artifact is 0.422 GiB smaller. The KL comparison remains engine-confounded.
  4. 8 bits, Q8_0 against online K5/K6: the bodies agree to 1.7 % (24.526 against 24.119), the most byte-comparable pair on the table, and the Q8_0 complete pipeline measures lower KL. Format attribution still requires a same-engine capture.

Rule from here on, and it is printed rather than footnoted: "at equal file bytes", "at equal tensor bytes", "at equal transformer body" and "as a deployed multimodal artifact" are four different claims, and whichever one a sentence means is written into the sentence. No fidelity number changes.

  • Q8_0 has the lowest measured complete-pipeline KL, 0.001087 at 27.05 GiB. Its engine differs from every vLLM row, so no quantization-only ordering is inferred from the 0.000507 BF16 control.
  • Every GGUF point at or above 5 bits measures lower complete-pipeline KL than official FP8. This makes our "34-48 % below FP8" headline a weaker achievement than it sounds, while remaining a cross-engine observation.
  • K4 is the weakest of our own builds, worse than every GGUF measured here and worse than official FP8. The one row it beats is unsloth's NVFP4.
  • The other 4-bit-weight artifact is far behind. unsloth/Qwen3.8-27B-NVFP4 reads 0.030115 with 93.16 % top-1 on this same shard, under the same engine so with no cross-engine term: 2.9x K4 at the same weight class, 8.8x the context edition, 27.7x Q8_0. It loses 512 of 512 contexts to both the context edition and official FP8 with zero ties, and over the full ten shards reads 0.031059 [0.027916, 0.034795] at 92.90 % top-1, losing 5,120 of 5,120 (receipts/kld5-1M-nvfp4.json, receipts/kld5-10M-nvfp4.json, receipts/kld5-1M-paired-nvfp4.json). The revision we measured, 9c73e2da, no longer resolves upstream after a 2026-08-15 history squash; the weights at current HEAD are byte-identical and we keep the reviewed revision alive as malaiwah/Qwen3.8-27B-NVFP4-archival-9c73e2da.

What it does not settle: this is text-only teacher-forced fidelity on one shard — one tenth of the suite, though a close tenth: over all 10,480,640 positions the five vLLM means read 0.002760 / 0.003210 / 0.003509 / 0.005294 / 0.010604, 1.9-2.9 % above these shard-0 values with the ordering unchanged (receipts/kld5-10M-*.json), and the GGUFs have no ten-shard equivalent. It says nothing about serving 262,144 tokens with vision and MTP on a 32 GB card, which is where these artifacts actually differ, and llama.cpp KV-quant behaviour, prefill and decode speed are separate axes that were not measured. A separate control bounds the protocol objection instead of arguing it: restricting scoring to positions with at least 256 tokens of left context — the floor llama-perplexity enforces — lowers every candidate's mean by 1.3-2.1 %, and second-half-only scoring by 3.9-4.9 %, uniformly enough to change no ordering (receipts/scored-window-offset.json), so the external protocol's scoring floor explains at most about 5 % of any cross-protocol gap. Protocol identity and the remaining open comparators are in docs/35 and docs/29 F2.

Public capability, all six models on one pinned MMLU-Pro subset. 70 questions, 14 official categories x 5, official five-shot prefixes, TIGER-Lab/MMLU-Pro@b189ec765aa7ed75c8acfea42df31fdae71f97be, greedy, thinking at low effort, 5,120-token completion cap, every candidate scored item-paired against the BF16 control. This is a first public, licence-compatible, item-paired benchmark, not a leaderboard claim.

modelabsoluteWilson 95 %BF16-pass retentionWilson lowerregressionsimprovementscompletion-cap failuresreceipt
BF16 Qwen/Qwen3.8-27B57/70 (81.4 %)[70.8 %, 88.8 %]reference4receipts/public-capability-bf16.json
context edition58/70 (82.9 %)[72.4 %, 89.9 %]56/5790.7 %123receipts/public-capability-ctx.json
K457/70 (81.4 %)[70.8 %, 88.8 %]55/5788.1 %224receipts/public-capability-k4.json
hydrated56/70 (80.0 %)[69.2 %, 87.7 %]54/5785.6 %324receipts/public-capability-hyd.json
official FP8 Qwen/Qwen3.8-27B-FP856/70 (80.0 %)[69.2 %, 87.7 %]55/5788.1 %214receipts/public-capability-fp8.json
online K5/K655/70 (78.6 %)[67.6 %, 86.6 %]54/5785.6 %314receipts/public-capability-k5k6.json

The bar, and who clears it. receipts/public-capability-plan.json pre-registered two conditions before any candidate ran: BF16-pass retention with a Wilson 95 % lower bound at or above 0.90, and no category losing more than two BF16 passes. All five candidates satisfy the second — the worst per-category loss is two passes (hydrated and online K5/K6, philosophy). Only the context edition satisfies the first, at 90.7 %. K4 and official FP8 read 88.1 %; hydrated and online K5/K6 read 85.6 %. Four of five candidates therefore miss a bar written down in advance, and that is published rather than restated as a pass.

Why that is a power limit, not a verdict. With 57 BF16 passes as the paired denominator, 56/57 is the smallest count whose Wilson lower bound clears 0.90 — a single paired miss already fails it — so a 70-item suite has too few items to certify this bar at all, and every candidate interval overlaps every other. The matrix separates nothing at this size. It applies equally to official FP8, which misses the same bar; none of this is evidence that any build is broken. Exact-output agreement is 0/70 for every EXL3 candidate and 1/70 for official FP8 (a single 113-token math answer) because long chains of thought differ token-wise, so the pass/fail outcome is the only meaningful pairing. Harness tools/public_capability.py, sweep runner tools/run_public_capability.sh, superseded 2,048-cap control receipts/public-capability-bf16-superseded-cap2048.json. The honest next step is more items and more task types, the plan's own P1 — HumanEval+/MBPP-style executable cases, IFEval-style constraints, tool schemas and a larger MMLU-Pro draw — not another sweep of the same 70 questions; no capability claim graduates before that.

Three limits stated up front: v5 absolute KLD is not comparable to v3 (K4 reads 0.029679 there and 0.010604 here — only within-suite ordering and paired differences transfer); this run has no cumulative percentiles over all ten shards, only per-shard p95/p99/p999 and the exact global maximum, because its shard reports predate the qwen38-fidelity-report/2 KLD histogram — the tail table above is the shard-0 rerun, 1,048,064 of those 10,480,640 positions; and its hidden states were deleted shard by shard to fit 135 GB of scratch, so it is reproducible from the pinned corpus fetch log and suite manifest plus a GPU rather than from published captures. See PROGRESS.md for the full session record.

File map

DocContents
docs/01-nvfp4-composition.mdMeasured tensor-level composition of nvidia/Qwen3.6-27B-NVFP4 and unsloth/Qwen3.8-27B-NVFP4, plus the BF16 parameter census
docs/02-recipe-k4.mdThe proposed recipe, footprint arithmetic, and the headroom exchange rates
docs/03-gg-runtime-contract.mdWhat the Gilded Gnosis EXL3 loader requires, and what it does not support
docs/04-exllamav3-toolchain.mdexllamav3 conversion flags, the missing per-module override, and the splice route
docs/05-kld-protocol.mdSuperseded iteration-1 protocol: the single-window, full-vocabulary logits KLD inherited from the GG r34 runner. Kept for provenance; the current method is docs/42
docs/06-baseline-validation.mdRunning the official GG image with no container runtime, and the two proven baselines
docs/07-serving-recommendations.mdServing-guide differences across upstream / NVIDIA / Unsloth cards, cross-checked against shipped configs
docs/08-upstream-cards-digest.mdPer-card digest: declared recipe, benchmarks, harnesses, limitations
docs/09-variant-publication.mdHow iterative variants are published for independent re-measurement
docs/14-fidelity-protocol-v2.mdHidden-state replay protocol, supersedes the single-window KLD
docs/18-results-fidelity-v3.mdHeld-out re-measurement, the later offset-independent contamination correction, per-stratum means, and controls
docs/22-results-iteration-2.mdGate K5 / up K5 / down K6: corrected mean KLD 0.007945, 38 % below official FP8 at 71 % of its resident weight
docs/15-results-fidelity-v2.mdSuperseded v2 run: 151,478 positions, 74/74 paired wins — measured on a corpus that overlapped our calibration data
docs/16-head-attribution.mdIs lm_head the sensitive tensor? Measured: not at K6
docs/11-kld-external-comparison.mdPublished KLD data for this family; why "FP8 = 0.5" is wrong
docs/13-upstream-contributions.mdUpstream issue + verified PR
docs/19-cuda-graphs-patch.mdThe autotune-priming patch behind PR #314, deliverables and verification
docs/20-context-extension-and-k3-gap.mdHow far context can be extended, and the distance to the reference protocol
docs/12-iteration-2-plan.mdWhere to invest next, ranked
docs/17-next-iteration-shopping-list.mdIteration-2 shopping list, each item with its closing test
docs/10-results-iteration-1.mdIteration 1: build, serve, KLD, and the upstream defect
docs/21-independent-review-response.mdIndependent review, per finding: fixed, fixed-in-v2, or still open
docs/24-p0-results.mdP0 done: prefill +113 %, and why fp32 replay was a negative result
docs/23-next-attack-list.mdRanked plan for iteration 3, with evidence, cost and acceptance per item
docs/25-goal-pareto-dominate-fp8.mdthe goal as a gate, and the iteration-3 verdict per axis
docs/26-prefill-attribution.mdwhere prefill time goes: MLP 2.13-2.26x, attention overlay 1.05-1.11x, hgemm at cuBLAS parity
docs/27-graph-decode-drift-control.mdeager-vs-graph drift is ambient: BF16 drifts the same as EXL3
PROGRESS.mdChronological work log
docs/28-external-validation-and-corrections.mdexternal RTX 5090 validation and the corrected capacity boundary
docs/30-iteration-4-context-edition.mdserialized-K5 context build, receipts and rejected vision quant
docs/31-frozen-qualification.mdsource-disjoint frozen v4 qualification
docs/32-native-context-embedding-overlay.mdint8 input table, native MTP-3 plus 8.4 MP engine-budget proof (since superseded by the physical 5090 qualification at utilisation 0.955), and corrected draft accounting
docs/33-evidence-volume-and-intervals.mdthe 10,480,640-position rerun, the measured tail, and why ten times the positions did not narrow the interval
docs/29-plan-and-loose-ends.mdranked open work and acceptance gates, including the evidence-volume objection (F1) this session closed
docs/35-external-protocol-comparability.mdwhy our KLD and Unsloth's are not interchangeable, the twelve deltas, and the two of them now measured: the scoring-window offset and the cross-engine floor
docs/34-vram-class-profiles.mdthe 24 GB and 16 GB class verdicts, the measured affine KV law, and a capped-budget proxy qualification
docs/36-performance-levers-5090.md11 configurations on the physical 5090: what moved throughput, what is a no-op, and the GDN gate reachability rate
docs/37-error-driven-allocation.mdthe error-driven allocation experiment that lost, the 3.73x-per-bit law it produced, and why the objective is not a surrogate for KLD
docs/39-chat-template-audit.mdevery chat template in this model family captured with revision and sha256, diffed by named construct, and the verdict that ours is byte-identical to official; the community "fixed" template measured losing tool-call arguments under qwen3_coder, and the loudest complaint traced to the 2.4T flagship's template rather than the 27B's
docs/42-kld-method.mdthe KLD method of record: the metric, the shared BF16 head and why, the capture point, the suite and its contamination policy, the ladder, the bootstrap and the tail histogram, the harness-revision pin, how to re-derive a published number on a CPU in seconds, every floor that bounds an absolute value, and the claim boundary
docs/14-paro-assessment.mdParoQuant (Pairwise Rotation Quantization, arXiv:2511.10645) assessed: a pre-quantization transform orthogonal to EXL3, not a comparator; the checkpoint is INT4-linear and protocol-incomparable, the learned rotation could replace our fixed Hadamard — follow-up in docs/13
docs/57-eda-allocation-revisit.mdthe error-driven allocation revisit, including the later re-solve that refuted sqrt_energy and left a compounding-aware objective as the next defensible in-family method
docs/58-qwen36-quant-prior-art.mdprior art of per-module bitrate attribution in the Qwen 3.5/3.6 generation: llama.cpp use_more_bits U-shaped layer heuristic (first+last 1/8 promoted, attn_v special-cased as "more sensitive"), EXL2 Frobenius-norm simulated annealing vs EXL3 direct-KLD greedy optimizer with U-shaped allocation.py layer term, GDN-at-higher-precision as standard NVFP4 practice (driven by correctness bug not measurement), Unsloth finding attn sensitive for hybrid archs (direction matches our physics, mechanism not diagnosed), no published early-vs-late KLD experiment, and five concrete inspirations with cost and falsification

Tooling in tools/ is what produced the evidence: an unprivileged OCI image puller and a proot-based runner for it, the BF16 attention splice and checkpoint finaliser, the fidelity harness (fidelity.py, suite3.py) that builds the suite and replays captures, the decode-parity probe (decode_parity.py), the kernel microbenchmarks (prefill_micro*.py, gemm_cmp.py) and the upstream patches as standalone files. The v5 10.48 M-position run, the public-benchmark harness and the collection index are these files:

ToolWhat it doesReceipts
tools/fetch_corpus_v5.pyfetches the five-stratum v5 corpus (941 documents / 70,348,971 bytes) and pins every document by URL and sha256receipts/kld5-corpus-fetch-log.json
tools/suite3.pybuilds the frozen suite: exact-advance non-overlapping windows, whole-document calibration-overlap pre-exclusion (44 of 941), prior-suite token exclusion (160 v4 hashes, 0 reachable)receipts/kld5-suite-manifest.json (qwen38-distribution-fidelity/6)
tools/kld_ladder.shwalks the ladder one 512-context shard at a time — capture six models, replay five candidates, verify, delete the shard's ~64 GB of hidden states — because scratch is ~135 GBten per-shard reports, listed in every cumulative receipt
tools/kld_aggregate.pywelds verified per-shard reports into cumulative means, cluster bootstraps, paired comparisons and (for qwen38-fidelity-report/2 inputs) bin-bounded cumulative quantilesreceipts/kld5-10M-{hyd,k5k6,ctx,fp8,k4}.json, receipts/kld5-10M-paired.json
tools/fidelity.pycapture/replay harness; now also counts every scored position into a fixed log-spaced kld_tail histogram, bumping reports to qwen38-fidelity-report/2per-shard and per-candidate reports
tools/public_capability.pypaired MMLU-Pro run against a pinned dataset revision, official five-shot prefixes, greedy, with Wilson intervals and BF16-pass retentionreceipts/public-capability-{plan,suite-mmlupro-70,bf16,ctx,k4,hyd,fp8,k5k6}.json, superseded 2,048-cap control kept as receipts/public-capability-bf16-superseded-cap2048.json
tools/run_public_capability.shdrives the six-model sweep one server at a time against the frozen BF16 reference, with the per-candidate environment each build needs (every EXL3 candidate requires VLLM_EXL3_GRAPH_DECODE=1)the five candidate receipts above
tools/collection_index.pyone immutable row per published checkpoint, every field carrying its receipt path, sha256 and RFC 6901 pointer; disk bytes and resident_weights kept strictly distinctreceipts/collection-index.json (qwen38-collection-index/1)
tools/gguf_capture.cppllama.cpp-side capture of the post-final-norm state for a GGUF at pinned commit ece963f4, so a GGUF can be replayed through our shared BF16 head; bf16 rounding verified bit-identical to torch on 2,012,449 probe valuesreceipts/gguf-report-{q8_0,q6_k,q5_k_xl,engine-floor}.json, receipts/cross-engine-comparator.json
tools/gguf_manifest.pypins each capture to its GGUF blob digest and the llama.cpp identity that produced itthe four receipts/gguf-report-*.json
tools/build_llamacpp.shbuilds the capture binary against the pinned llama.cpp commitengine_identity in every GGUF report
tools/chat_template_matrix.pyrenders every candidate chat template over 34 fixed cases inside the pinned image, reconstructing transformers' Jinja environment exactly, and reports token counts and sha256 per rendered prompt; CPU onlyreceipts/chat-template-audit.json, receipts/chat-templates/renders.json
tools/chat_template_roundtrip.pyfeeds each template's rendered tool call back through the image's own vllm.parser.qwen3._qwen3_arg_converter and reports whether --tool-call-parser qwen3_coder recovers the arguments losslesslytool_call_roundtrip in the audit receipt
tools/chat_template_prefix.pyfirst-divergent-token between the generated token stream and the re-rendered history, per client reasoning-echo behaviour, i.e. the point past which no prefix-cache block is reusableprompt_growth_and_prefix_stability in the audit receipt
tools/fetch_chat_templates.sh, tools/gguf_chat_template.pyfetch every family template with its revision and digest, including off-default-branch refs and the one that exists only inside GGUF metadata (read over HTTP range requests, no tensor bytes)receipts/chat-templates/{SHA256SUMS,raw/}

Status

  • Both upstream Qwen3.8-27B artifacts proven runnable under the official GG image
  • Iteration 1 (K4 MLP + BF16 attention splice) served under GG with ONLINE_QUANT=exl3-b6, 19.21 GB resident, vision verified, published as malaiwah/Qwen3.8-27B-K4
  • Held-out v3 fidelity suite with a contamination scan, plus the dataset that lets anyone recompute it without a GPU
  • Iteration 2 built, measured and published as malaiwah/Qwen3.8-27B-EXL3-K5K6: 20.32 GiB resident, overlap-corrected mean KLD 0.007945 body-only / 0.008078 as served, top-1 96.86 %
  • CUDA graphs (PR #314), row-count prefill dispatch (PR #316), B12X prefill routing (PR #318) and quantized-embedding construction (PR #319) filed against the fork and exercised in the tested runtime
  • Prefill attributed and bounded: the MLP kernel is the whole story (2.13x / 2.26x), the online-K6 attention overlay is not the bottleneck (1.05x / 1.11x), ext.hgemm is already at cuBLAS parity — FP8 prefill parity needs a fused dequant-in-epilogue kernel, not tuning
  • Attention-overlay width is a runtime knob, measured on the same suite and real RTX 5090: K6 0.007945 at 20.32 GiB and ~185k context; K5 0.011801 at 19.82 GiB and 206,400 retrieval-verified; K4 0.026619 at 19.05 GiB
  • Serializing attention offline instead of encoding it at load — the "hydrated" build — measures 0.007172, i.e. calibrated offline K6 beats the runtime overlay by 0.000773 on the overlap-corrected subset
  • Graph-vs-eager decode drift measured properly and traced to the build, not to EXL3 (docs/27)
  • Frozen v4 suite: 160 source-disjoint contexts, 100 documents, zero token/document/content overlap; all five candidate capture sets and qualification receipts published (2,708 files / 51.0 GB)
  • Context edition plus per-row int8 input table starts native 262,144 with MTP-3, decode graphs and an 8.4 MP image cap under a 30.24 GiB engine budget; 266,612 KV tokens, exact retrieval at 261,794 text tokens and in a 236,824-token seven-megapixel request — that engine-budget run is now superseded by the physical-card qualification below
  • Hard-limit qualification on a physical RTX 5090 closed (receipts/qualification-5090-context.json, per-process server logs receipts/qualification-5090-context-server-{B3,B4,C,D,E,F}.log): the context edition serves native 262,144 with MTP-3, the full 8,388,608-pixel image ceiling and fp8 KV on one 32,607 MiB card — but at --gpu-memory-utilization 0.955 with PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, not the 0.97 the cards used to print. Measured at startup: engine budget 29.98 GiB, free 30.9 of 31.4 GiB usable, usage 18.19 weight + 1.78 peak activation + 0.27 non-torch + 0.45 CUDAGraph = 20.69 GiB, available KV 9.28 GiB = 265,122 KV tokens, 1.01x maximum concurrency at 262,144, attention block size forced to 1600 tokens so the attention page is at least the mamba page, 3 padding layers at at most 6.25 % KV waste, startup 55.7 s, model load 18.19 GiB in 3.99 s. All seven gates PASS: (1) startup allocates a native-length request inside the utilisation ceiling; (2) the 261,794-token needle retrieved exactly; (3) a combined 236,824-token plus 7,077,888-pixel request returns 1376346594 | red, blue exactly; (4) the 30-case image suite at 24/30; (5) three warmed 256-token concurrency-1 decode runs; (6) a second native-length request after release in the same process; (7) receipt identity complete
  • 0.97 published as a bounded negative, with its mechanism: at 0.97 the combined text-plus-image request dies with torch.OutOfMemoryError inside vllm/v1/attention/ops/vit_attn_wrappers.py wanting 62.00 MiB with 26.50 MiB free, and expandable_segments does not rescue it because vLLM spends the freed bytes on more KV (272,570 → 280,017 tokens); lowering max_pixels to 4,194,304 at 0.97 is strictly worse (profiled peak activation falls to 1.35 GiB, KV grows to 291,933 tokens, OOM with 6.56 MiB free). The knob is utilisation, not the image ceiling — seven megapixels is not a hard 32 GB ceiling, since the identical request succeeds at 0.955 with the full 8,388,608-pixel ceiling
  • Task-retention smoke: BF16, all four EXL3 profiles and official FP8 each pass 40/40 deterministic paired tasks with zero regressions; a hardened rescore leaves every pass unchanged and corrects exact-final-answer agreement to 32/40–35/40 (receipts/task-retention-v2-summary.json, receipts/task-retention-v2-strict-rescore.json)
  • Evidence volume objection (F1) closed: fidelity re-measured on the held-out v5 suite at 10,480,640 scored positions / 5,120 contexts / 842 source clusters, ten verified shards welded by tools/kld_aggregate.py — hydrated 0.002760, online K5/K6 0.003210, context 0.003509, official FP8 0.005294, K4 0.010604, all body-only through one shared BF16 head (receipts/kld5-10M-*.json)
  • Paired at 5,120 contexts: hydrated -0.002534 (5,118 wins), online K5/K6 -0.002084 (5,105), context -0.001785 (5,109) against official FP8; hydrated beats the runtime overlay by -0.000450 (4,922); K4 loses to FP8 on 5,113 of 5,120 (receipts/kld5-10M-paired.json)
  • Ladder stability shown, not assumed: cumulative hydrated mean 0.002700 / 0.002759 / 0.002699 / 0.002760 at the 1M / 2M / 5M / 10M checkpoints
  • Exact tail now aggregable for future runs: a fixed log-spaced KLD histogram in tools/fidelity.py (qwen38-fidelity-report/2) plus bin-bounded cumulative quantiles in tools/kld_aggregate.py; this run kept only per-shard percentiles and the exact global maximum, and says so in its receipts
  • First public-benchmark matrix, all six models: paired MMLU-Pro (70 questions, pinned TIGER-Lab/MMLU-Pro@b189ec76) — BF16 57/70, context 58/70 at 56/57 retention (Wilson lower 90.7 %, the only candidate clearing the pre-registered 0.90 bar), K4 57/70 and official FP8 56/70 at 55/57 (88.1 %), hydrated 56/70 and online K5/K6 55/70 at 54/57 (85.6 %); the four shortfalls are a 70-item power limit — 56/57 is the smallest count that can clear 0.90 — not a broken build, so the next step is item volume and task diversity
  • Collection index published: receipts/collection-index.json from tools/collection_index.py, four immutable rows, receipt-traced fields, one disclosed divergence left unfixed rather than rewriting a published receipt
  • Comparator breadth (F2) closed by measurement: GGUF Q8_0 0.001087, Q6_K 0.002035 and UD-Q5_K_XL 0.004444 were scored on shard 0 of the v5 suite through the shared BF16 head. The unquantized llama.cpp-versus-vLLM control is 0.000507 and is diagnostic, not subtractable; cross-engine rows therefore compare complete pipelines and do not isolate quantization format (receipts/cross-engine-comparator.json)
  • The largest cross-protocol objection bounded instead of argued: restricting scoring to their 256-token left-context floor lowers every candidate's mean by 1.3-2.1 % and second-half-only by 3.9-4.9 %, uniformly, changing no ordering (receipts/scored-window-offset.json)
  • Their protocol run for cross-citation: llama-perplexity --kl-divergence on wikitext-2 raw at ctx 512, 147,900 scored positions, base PPL 6.950230 ± 0.044933 — Q8_0 0.000926, Q6_K 0.002286, UD-Q5_K_XL 0.004426, with the same ordering as our suite. The absolute values and ratios are protocol-specific; their harness floor was measured on our hardware at 5.6e-5 to 8.0e-5 and tokenisation was bit-identical over 297,194 tokens. PPL does not reproduce the KLD ordering (receipts/wikitext-kld-run-a.json, docs/35)
  • unsloth/Qwen3.8-27B-NVFP4 measured on our suite, the comparator readers ask for most: 0.031059 [0.027916, 0.034795] at 92.90 % top-1 over 10,480,640 positions, losing 5,120 of 5,120 contexts to both the context edition and official FP8; 2.9x K4 at the same 4-bit weight class (receipts/kld5-10M-nvfp4.json, receipts/kld5-1M-paired-nvfp4.json)
  • Capture and replay are bit-reproducible: 5,240,320 scored positions identical across independent runs, separate model loads and three harness generations, under eager single-sequence capture — which makes every published KLD number re-derivable rather than merely re-runnable (receipts/capture-determinism.json)
  • The converter is not deterministic, and it does not matter for fidelity. A fresh conversion of the published hydrated recipe returns 13 of 16 pinned files identical and three safetensors shards differing in tensor bytes only — 399 of 409 quantized modules, 97.6 %, at 41-92 % of bytes each — with identical headers, offsets, shard sizes and per-role byte totals. That sibling then scored -3.755e-06 paired against the published checkpoint, 95 % [-2.854e-05, +2.062e-05] bracketing zero, 257 contexts to 255. So the recipe determines fidelity to within this protocol's resolution and the published bytes are the artifact (receipts/converter-determinism.json, receipts/sibling-rebuild-fidelity.json)
  • Performance levers measured on the physical 5090, 11 configurations: MTP depth is a concurrency decision — depth 3 wins at one stream (83.31 vs 74.28 per-request tok/s) and num_speculative_tokens 1 wins at eight (+30.67 %, 313.28 → 409.35 tok/s aggregate, plus 10,911 more KV tokens). Closed as no-ops or losses: explicit FLASHINFER (already auto-selected on SM120), TRITON_ATTN, custom_ops:["all"], and dynamic speculative decoding (refuses to start under FULL_DECODE_ONLY). Prefill did not move on any lever (docs/36)
  • Prefix caching shipped, scoped by measurement: on the three 8,192-token recipes it is worth 11.6x and 29.3x warm TTFT at 32 k and 131 k prefixes (whole schedule 144.4 s → 84.0 s); at the context edition's native 262,144 it is declined because align-mode block rounding makes the window unservable three different ways, with a measured 256,000-token option priced at 6,144 tokens. The four-module image gg-r34-patched-apc is the release unit (receipts/apc-poison-repro.json, receipts/qualification-5090-apc.json)
  • VRAM classes decided: 24 GB is a GO as a serving profile over the published context edition at 24,576 tokens with MTP-3 or 45,056 with MTP off, validated by a capped-budget proxy with the physical-board gate left open; 16 GB is a NO-GO as a SKU and published as a design study. The per-token KV model was replaced by the measured affine law a*L + M (docs/34)
  • A user-reported defect investigated to a measured negative: 266 requests across seven freshly started servers reproduced no corruption under the reporter's exact condition, which decomposed his four-variable control and left LMCache as the sole suspect; the two absent upstream fixes were requested at the fork (issue #392, PR #393) and a net-new admission livelock was filed upstream as vllm-project/vllm#52520
  • Materials published for third-party replay: the v5 suite, all ten shard views, the shard-0 BF16 reference capture and 79 per-shard reports as malaiwah/qwen38-27b-fidelity-suite-v5, the losing error-driven candidate as a documented negative, and archival mirrors of the NVFP4 and GGUF revisions our published numbers cite — because one of them stopped resolving upstream mid-session
  • Still open in F2: stock uniform-bitrate EXL3 controls, and the reader-suggested ik_llama (IQ4_KS_KT) and NVFP4-NInfer comparators — neither of which has any KLD number today (docs/29 F2)

Contributors

malaiwah

559 commits

Languages

Python

85.2%

Shell

9.6%

Jinja

2.5%

Cuda

2.3%