Status 2026-08-16: the collection is hardware-qualified, the comparator set is measured, and the method is now reproducible from published artifacts. The current checkpoint is
malaiwah/Qwen3.8-27B-EXL3-K5K6: mean KLD 0.003210 (95 % CI [0.002982, 0.003480]) body-only at 20.32 GiB resident weights, versus officialQwen/Qwen3.8-27B-FP8at 0.005294 — 39 % lower divergence at 71 % of its resident weight, paired -0.002084 [-0.002249, -0.001942] winning 5,105 of 5,120 held-out contexts. The best profile in the collection is the hydrated build at 0.002760, paired -0.002534 on 5,118/5,120. The headline suite is v5: 5,120 contexts x 2,047 positions = 10,480,640 scored positions over 842 source clusters, calibration- and v4-token-disjoint.Landed since: the context edition is qualified on a physical RTX 5090 (all seven gates at
--gpu-memory-utilization 0.955withexpandable_segments, 265,122 KV tokens at 262,144 — and a second physical 5090 needed 0.956, so utilisation is a per-card measurement, not a constant).unsloth's NVFP4 is now on our suite at 0.031059 over 10.48 M positions, losing 5,120 of 5,120 contexts to both the context edition and FP8. The externalllama-perplexityprotocol was run for cross-citation and reproduces our ordering. Capture and replay are bit-reproducible, and a fresh conversion of the same recipe — a sibling, 97.6 % of quantized modules differing in bytes — scores indistinguishable from the published checkpoint, so the recipe determines fidelity while the published bytes remain the artifact. The suite, the shared BF16 reference capture and every per-shard report are published as a dataset, with archival mirrors of the third-party artifacts our numbers cite.Two more artifacts published 2026-08-16, both from pre-registered conversions.
malaiwah/Qwen3.8-27B-EXL3-K6-parityuses nearly the same file-byte budget as GGUFQ6_Kby promotinggate_proj/up_projK5 → K6. It measures 0.001634 [0.001541, 0.001742] under vLLM versusQ6_Kat 0.002035 under llama.cpp, while beating hydrated on 511/512 contexts (-39.5 % for +1.348 GiB) and carrying 2.31 GiB less transformer body thanQ6_K. The cross-engine control proves these are complete-pipeline results; it cannot be subtracted to establish format parity or prove that the earlier gap was caused by bytes.malaiwah/Qwen3.8-27B-EXL3-S16-V-researchis the opposite: the rejected sub-4-bit 16 GB candidate at 0.045374 [0.041959, 0.049351], 1.5x its pre-registered NO threshold, losing 512/512 contexts to K4 (4.39x) and to the context edition (13.31x). It is published because it failed - a NO nobody can audit is an assertion - and it is the empirical anchor for the K3 rung of the per-bit law (docs/37 §3.3). Both payloads were predicted to the byte by the affine law before conversion, and both registered intervals are reported with their misses: S16-V's point estimate was 1.52x too high, K6-parity's interval fell 2.0 % short.The earlier headline — 0.007945 body-only on 136 v3 contexts / 278,392 positions (docs/22) — stands exactly as measured and is superseded only as the headline; v5 absolute KLD is not comparable to v3 (the two suites differ by a measured 2.46-3.08x with identical ordering), so quote paired differences across suites, never means. Figures labelled K4 or v1 belong to iteration 1 and are superseded. One earlier control was withdrawn: the "CUDA-graph parity 0.000000" receipt captured a prefill forward, which
FULL_DECODE_ONLYnever captures, so it could not have measured decode; the decode probe that replaces it is docs/27. Open items are tracked in docs/29.
The primary control is the exact PROFILE=fidelity tuple, not
PROFILE=throughput and not a memory-modified derivative. The objective is to
recover enough PP/TTFT from that control to pass the absolute candidate gates
without sacrificing its paired KLD/tails, TG, context, or capability envelope.
The fidelity control is allowed to fail the 7,000 tok/s PP candidate threshold;
that threshold governs promotion of a new candidate, not acceptance of the
control.
receipts/frontier-fidelity-control.json
pins the model/profile identity, INT6 embedding overlay, exact evidence hashes,
reference metrics, candidate rules, and the prohibition on substituting the
high-KLD throughput profile. Every INT8/FP4/FP6, K-map, topology, kernel-route,
KV, MTP, graph, or profile change is a candidate variable and requires
preregistered paired comparison against this exact control.
The exact control and method-neutral harness are ready. Clean packed INT6 was
integrated onto runtime b19029d, matched the historical encoder byte-for-byte
on real BF16 rows, and passed three independent AIBoss boots. Aggregate PP was
2,941.9 tok/s, fox/essay TG 226.9/104.3, context 238,400, with text,
vision, MTP, 200k context and rollback passing on every boot.
The clean-runtime fidelity recapture is exactly equal to the historical control: mean KLD 0.0034052972, p99 0.0348892, paired delta and CI 0.0, with all 512 per-context rows and the full tail histogram identical. Its 513-file, 10.73 GB hidden-state capture has an independently verified private-bucket copy.
The repaired v3 harness uses the actual EXL3 Viterbi quantizer—not the invalidated
uniform proxy—on four real 128×128 Qwen slices with matched BF16/quant-flow
activations and Hessians. RTN and both LDLQ controls carry the same exact
10,756-byte / 5.251953-bpw payload. Quant-flow-H LDLQ reduced mean excess
running-output error from RTN's 0.0011858 to 0.0008865 at those same bytes;
this is a control result, not a method promotion. Whole-block and end-logit gates
remain mandatory and intentionally pending for researcher-proposed methods.
receipts/frontier-g02/manifest.json is
validated by python3 tools/verify_frontier_g02.py --campaign receipts/frontier-g02 as pre_candidate_ready. Candidate slots remain unopened,
E-final-v7 remains sealed, and rental authorization remains false.
The preregistered campaign in
docs/61-final-frontier-requant-stack-plan.md
ended at its local RTX 5090 gate with the terminal disposition
three_candidate_no_go. The complete hashed record is
receipts/frontier-g01/manifest.json;
python3 tools/verify_frontier_g01.py --campaign receipts/frontier-g01
validates every referenced G0 artifact, preregistration, result and terminal
decision and exits 0.
G0 passed: the immutable BF16 payload was fully censused (1,199 logical
tensors), converter/runtime environments were locked, the fidelity-derived
incumbent-fidelity-int8 clean runtime was qualified, and every AIBoss
maintenance transaction proved service restoration. That historical G0 arm
used INT8 embeddings and did not remeasure the exact PROFILE=fidelity KLD; it
is runtime/rollback evidence, not a substitute for the frozen primary control.
G1 then consumed the three frozen candidate opportunities:
| candidate | measured result | frozen Gate A failure |
|---|---|---|
| full gate/up FP6 | startup rejected the 199,104-token configuration; estimated maximum 192,576 | context below 238,400 |
| full gate/up FP4 | PP 4,393.6 tok/s; fox/essay TG 211.6/98.8; context and capability checks passed | PP below 7,000 |
| gate/up FP4, language layers 0–31 | PP 3,529.5 tok/s; fox/essay TG 215.9/96.0; context and capability checks passed | PP below 7,000 |
The preregistered hard-gate rule stopped KLD scoring after each deterministic runtime failure. Supporting experiments also closed without a promotable candidate: independent seeded converter A/A runs had 24 payload mismatches; the pinned AutoRound consumer could not build; ResComp improved one actual 16×16 weight-tile SSE by 21.1% but lacked a valid whole-block implementation; and the hybrid QKV producer/consumer pilot passed its component checks without the mandatory full-checkpoint, startup, graph and logit gates. No paid rental or new model release is authorized by this campaign; the measured serving profiles below remain current.


Three profiles, selected by PROFILE= in
patches/run-qwen38-27b.sh. All three pass
tools/verify-profile.sh (exit 0, 8/8 checks each, including a 200k-token
prompt). PP/TG are the tools/bench-profile.sh harness at n=3 boots;
KLD is 512 contexts of shard-0000 against the BF16 reference.
PROFILE=throughput | PROFILE=fidelity | PROFILE=balanced | |
|---|---|---|---|
| weights | all-FP4 | all-trellis (K5K6 as shipped) | trellis + gate_up MXFP6 |
| PP, 2051-tok | 9638.9 ± 18.3 tok/s | 2987.7 ± 4.4 tok/s | 3925.2 ± 13.1 tok/s |
| TG fox / essay | 187.4 ± 0.6 / 94.3 ± 0.0 | 228.3 ± 0.4 / 104.1 ± 0.1 | 215.6 ± 0.2 / 103.7 ± 0.1 |
| MTP acceptance fox / essay | 0.930 / 0.298 | 1.000 / 0.304 | 1.000 / 0.324 |
| KLD mean | 0.063759 | 0.003405 [0.003166, 0.003672] | 0.005672 [0.005302, 0.006087] |
| KLD p99 | 0.7010 | 0.034889 | 0.059908 |
| max context | 249,600 | 238,400 | 199,104 |
| vision + MTP | pass | pass | pass |
| criteria met | 4/6 (fails KLD) | 5/6 (fails PP) | 5/6 (fails ctx) |
fidelity serves the checkpoint at KLD 0.003405 — within 26% of this
collection's own published trellis fidelity (0.002700) — with TG 228.3 tok/s
and full 238,400 context; the residual over the checkpoint is the int6 embedding
table (~0.0007), not any GEMM approximation. It costs ~3.2x prefill.
throughput is the only profile above 7000 tok/s prefill.
Long-context needle retrieval on the fidelity profile (fixed harness, single
needle per context): 8/8 at 2k, 8/8 at 100k, 8/8 at 195k — 24/24. This is
an easy proxy — finding a planted needle does not establish unimpaired
long-context reasoning, only that the attention window and KV cache are intact
to 195k.
No single profile meets all six north-star criteria. Concretely: throughput
fails KLD (0.063759, well above the 0.012 budget); fidelity fails the 7000 tok/s
prefill criterion (2987.7); balanced fails the context criterion (199,104
tokens, below the 238,400 that fidelity achieves). The other two profiles each
clear five of six.
receipts/frontier-2026-08-19.md shows why
from three directions: prefill-grade throughput needs the MLP resident in a
GEMM-ready format, trellis-grade fidelity needs weights that are decoded per
prefill chunk, and GEMM-resident FP8 for all 24.3e9 quantized parameters is
22.6 GiB — which cannot coexist with a 238,400-token KV cache in 31.4 GiB. The
blocker is memory, not kernel quality.
Fixes landed while getting here, with receipts: an engine-fatal OOM on any
prompt over ~4k tokens (max_num_batched_tokens 8192 → 3072, which also
freed 0.93 GiB of KV), +32% TG by discovering the MTP draft loop ran eager
on the V1 model runner, and the first profile of this stack (prefill is 92%
GPU-bound; decode was 2%). Upstream: vllm-project/vllm#52871, #52872,
local-inference-lab/vllm#439, #440, b12x #232/#233/#234.
A baked serving image is published at
docker.io/malaiwah/qwen38-27b-exl3-gg:r34-p2-41a5d16 (also tagged :latest).
It layers every patch this project bind-mounts on top of the digest-pinned
Gilded Gnosis r34 base (voipmonitor/vllm:gilded-gnosis-v20-vllm4d006a4-b12xcd3ce19-fi1ac6942-cu132-20260810-r34,
sha256 820181fbbc975cd5291c411cda9771d58fecee1636d916f508f47230df20592b),
so the repo is not required at runtime — podman run plus a Hugging Face
cache mount is sufficient. All 10 patches (7 vLLM source overlays and 3
conversion modules) are baked in by docker/Containerfile,
which also bakes the CUDA lib64 symlink the B12X JIT needs as real links and
carries OCI provenance labels (org.opencontainers.image.source, .revision,
.description, ai.malaiwah.base-digest).
The image was gated 9/9 mount-free (NO_PATCH_MOUNTS=1, every patch
bind-mount disabled): PP 2933.0, fox 229.1, essay 103.7, ctx 238,400, 200k-token
prompt OK, vision OK, MTP acceptance_fox 1.000 — confirming the image is
self-contained.
PROFILE=throughput|balanced|fidelity selects the serving profile;
NO_PATCH_MOUNTS=1 is what makes the baked image run mount-free. The full flag
set is in patches/run-qwen38-27b.sh. A
standalone fidelity invocation — every flag below is taken from that script —
is:
podman run --replace -d \
--name qwen38-27b \
--device nvidia.com/gpu=all --ipc=host --network host \
--tmpfs /usr/local/cuda-13.2/lib64:rw,size=16m \
-e HF_HUB_OFFLINE=1 \
-e CUTE_DSL_ARCH=sm_120a -e FLASHINFER_CUDA_ARCH_LIST=12.0f \
-e OMP_NUM_THREADS=8 -e CUDA_DEVICE_MAX_CONNECTIONS=32 \
-e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True -e SAFETENSORS_FAST_GPU=1 \
-e VLLM_WORKER_MULTIPROC_METHOD=spawn \
-e VLLM_USE_V2_MODEL_RUNNER=1 -e VLLM_NO_USAGE_STATS=1 -e DO_NOT_TRACK=1 \
-e XDG_CACHE_HOME=/cache/jit -e CUDA_CACHE_PATH=/cache/jit \
-e TRITON_CACHE_DIR=/cache/jit/triton -e TORCHINDUCTOR_CACHE_DIR=/cache/jit/torchinductor \
-e FLASHINFER_WORKSPACE_BASE=/cache/jit/flashinfer \
-e VLLM_EXL3_MULTIPRECISION=1 -e VLLM_EXL3_GRAPH_DECODE=1 \
-e VLLM_EXL3_FP4_TRITON_DECODE=0 -e VLLM_EXL3_FP4_PER_ROW_GS=0 -e VLLM_EXL3_FP4_DRAFT_HEAD=0 \
-e VLLM_EXL3_ONLINE_TRELLIS_BITS=6 \
-e VLLM_EXL3_ONLINE_CACHE_DIR=/cache/jit/exl3-online -e VLLM_EXL3_ONLINE_CACHE_MODE=readwrite \
-e VLLM_EXL3_EMBED_ONLINE_BITS=6 -e B12X_PACKED_B_MIN_N=1024 \
-e VLLM_EXL3_B12X_ANY_BITS=1 -e VLLM_EXL3_B12X_MIN_M=128 -e VLLM_EXL3_SKIP_TRELLIS_PREP=0 \
-e VLLM_EXL3_PREFILL_RECONSTRUCT_M=1 -e VLLM_EXL3_PREFILL_RECONSTRUCT_MAX_MB=512 \
-e VLLM_EXL3_PREFILL_RECONSTRUCT_CACHE=0 -e VLLM_EXL3_FOLD_FP32_BUDGET_MB=48 \
-e VLLM_EXL3_FP4_LAYERS=, \
-e VLLM_EXL3_EXT_PATH=/opt/exllamav3 \
-e HF_HOME=/root/.cache/huggingface \
-v ~/.cache/huggingface/hub:/root/.cache/huggingface:ro \
-v ~/.cache/jit:/cache/jit \
--entrypoint /bin/bash \
docker.io/malaiwah/qwen38-27b-exl3-gg:r34-p2-41a5d16 \
-lc 'set -euo pipefail; \
ln -sf /usr/local/cuda-13.2/targets/x86_64-linux/lib/* /usr/local/cuda-13.2/lib64/ 2>/dev/null || true; \
exec vllm serve \
/root/.cache/huggingface/models--malaiwah--Qwen3.8-27B-EXL3-K5K6-hydrated/snapshots/ab3a91a13813df8096cb4c1d560ed3669035d0cf \
--served-model-name Qwen3.8-27B --trust-remote-code \
--host 0.0.0.0 --port 8000 \
--quantization exl3 \
--quantization-config "{\"linear\":{\"weight\":\"mxfp8\"},\"ignore\":[\"re:.*visual\\..*\",\"re:.*in_proj_a$\",\"re:.*in_proj_b$\",\"re:.*in_proj_ba$\",\"re:.*mtp\\..*\",\"lm_head\"]}" \
--attention-backend TRITON_ATTN \
--gpu-memory-utilization 0.945 --kv-cache-dtype fp8_e4m3 \
--max-model-len 238400 --max-num-seqs 4 --max-num-batched-tokens 3072 \
--compilation-config "{\"mode\":\"NONE\",\"cudagraph_mode\":\"FULL_DECODE_ONLY\"}" \
--mm-processor-kwargs "{\"truncation\":false}" --mm-processor-cache-type shm \
--default-chat-template-kwargs "{\"preserve_thinking\": true}" \
--enable-chunked-prefill --reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_coder \
--speculative-config "{\"method\":\"mtp\",\"num_speculative_tokens\":6}"'
Flags taken from run-qwen38-27b.sh: --tmpfs + ln -sf symlink (B12X JIT
workaround, lines 225/296 — the baked image already carries these symlinks, so
the --tmpfs is redundant with the baked image but harmless), --device /
--ipc / --network (line 226), the -e environment block (lines 231–280),
the HF cache mount -v …/huggingface:ro (line 281), the JIT cache volume
-v …/cache/jit:/cache/jit (line 288), and the vllm serve argument list
(lines 303–322). For balanced or throughput, change the profile env vars
per the case block in the script (lines 83–154).
K4, EXL3-K5K6)Research materials and progress log for building a dense EXL3 quant of
Qwen/Qwen3.8-27B that is inspired by
NVIDIA's NVFP4 recipe for the previous generation, but keeps the attention projections
BF16 on disk for runtime K6 encoding by the Gilded Gnosis vLLM fork, and
serializes the MLP as EXL3 K5 gate / K5 up / K6 down with a K6 mcg head and a
quantized MTP draft. Iteration 1 serialized the whole MLP at K4.
Design goal: NVFP4-class VRAM footprint, lower KLD. Met on fidelity, memory, decode and speculative decode; prefill is a measured structural deficit of a 4-bit-class trellis format in this runtime (docs/25, docs/26).
Held-out v5 suite, receipts/kld5-suite-manifest.json (schema
qwen38-distribution-fidelity/6, suite_token_sha256 510541f6…09482b88): 5,120 contexts
x 2,047 positions = 10,480,640 scored positions, 842 source clusters, 1,024 contexts per
stratum. 44 of 941 corpus documents were excluded whole before window selection for any
all-position 12-token overlap with exllamav3 calibration data, leaving 897 eligible, so
contamination hits are 0 by construction; all 160 v4 context token hashes were seeded as
exclusions and 0 were reachable. Every figure below is body-only — both operands scored
through one shared BF16 LM head — with a 95 % CI from a 10,000-resample bootstrap over the 842
source clusters.
| candidate | mean KLD | 95 % CI | top-1 | paired vs FP8 | contexts won | receipt |
|---|---|---|---|---|---|---|
| hydrated | 0.002760 | [0.002540, 0.003020] | 97.70 % | -0.002534 | 5,118 / 5,120 | receipts/kld5-10M-hyd.json |
| online K5/K6 | 0.003210 | [0.002982, 0.003480] | 97.52 % | -0.002084 | 5,105 / 5,120 | receipts/kld5-10M-k5k6.json |
| context | 0.003509 | [0.003220, 0.003852] | 97.44 % | -0.001785 | 5,109 / 5,120 | receipts/kld5-10M-ctx.json |
| official FP8 | 0.005294 | [0.004927, 0.005728] | 96.79 % | — | — | receipts/kld5-10M-fp8.json |
| K4 (iteration 1) | 0.010604 | [0.009640, 0.011746] | 95.76 % | +0.005310 | 7 / 5,120 | receipts/kld5-10M-k4.json |
Paired intervals and win counts come from receipts/kld5-10M-paired.json (10,000 resamples,
seed 1); hydrated also beats the runtime overlay directly, -0.000450 [-0.000469,
-0.000433] on 4,922/5,120. Exact worst single position over the whole run: hydrated 8.258,
online 22.241, context 5.557, FP8 10.714, K4 14.283. The cumulative hydrated mean at the
1M / 2M / 5M / 10M ladder checkpoints is 0.002700 / 0.002759 / 0.002699 / 0.002760
(shards[] of receipts/kld5-10M-hyd.json).
The tail, not just the mean. The ten-shard run predates the histogram, so the tail was
measured by re-running shard 0 of the same v5 suite with the qwen38-fidelity-report/2
harness: 512 contexts x 2,047 positions = 1,048,064 scored positions, the same contexts
every candidate saw. Quantiles are bin-bounded — 560 log-spaced bins from 1e-12 to 1e2
nats, one bin ~5.6 % wide, and every receipt carries lower / upper / estimate per
quantile — while the maxima and the exceedance counts are exact. The two columns worth
reading are p99.9 and the share of positions above 0.1 nats.
| candidate | mean | p50 | p95 | p99 | p99.9 | p99.99 | exact max | above 0.1 | above 1.0 | receipt |
|---|---|---|---|---|---|---|---|---|---|---|
| hydrated | 0.002700 | 0.00109 | 0.0082 | 0.0276 | 0.1319 | 0.463 | 3.735 | 0.1534 % | 0.00219 % | receipts/kld5-1M-tail-hyd.json |
| online K5/K6 | 0.003141 | 0.00128 | 0.0099 | 0.0321 | 0.1446 | 0.498 | 5.507 | 0.1820 % | 0.00200 % | receipts/kld5-1M-tail-k5k6.json |
| context | 0.003409 | 0.00135 | 0.0107 | 0.0357 | 0.1642 | 0.587 | 3.749 | 0.2287 % | 0.00305 % | receipts/kld5-1M-tail-ctx.json |
| official FP8 | 0.005197 | 0.00202 | 0.0167 | 0.0531 | 0.2438 | 0.812 | 5.296 | 0.3912 % | 0.00592 % | receipts/kld5-1M-tail-fp8.json |
| K4 | 0.010345 | 0.00320 | 0.0332 | 0.1194 | 0.5555 | 1.870 | 7.565 | 1.2604 % | 0.03807 % | receipts/kld5-1M-tail-k4.json |
The ordering at p50, p95, p99, p99.9 and p99.99 is the same as the ordering of the means, so
for these candidates the mean is not hiding a worse tail: every EXL3 K5/K6-class build has a
lighter tail than official FP8 at every measured quantile, and K4 is worse than FP8 at
every quantile. Receipts are schema qwen38-kld-ladder-cumulative/2, welded by
tools/kld_aggregate.py from /2 replay reports.
Against GGUF, measured, including the part that goes against us. The standing critique is
that official FP8 is a throughput format whose quality is Q4-to-Q5 class, so beating it is a weak
claim, and that Q8_0 and Q6_K are the honest bar. That is now measured rather than argued.
Three GGUFs from unsloth/Qwen3.8-27B-GGUF@f1bfb127c64f7072bdd2cad55f258b9c8b2910fe were
captured under llama.cpp pinned at commit ece963f41b0b02d7a0d61436ae365762c073a4c8 with
tools/gguf_capture.cpp, which reads the post-final-norm state — the same mathematical point our
vLLM hook takes, with bf16 rounding verified bit-identical to torch on 2,012,449 probe values —
and scored through the same shared BF16 head on shard 0 of the v5 suite: the same 512
contexts and the same 1,048,064 positions every candidate above saw (tools/gguf_manifest.py,
tools/build_llamacpp.sh; receipt receipts/cross-engine-comparator.json, per-candidate reports
receipts/gguf-report-{q8_0,q6_k,q5_k_xl}.json).
| candidate | engine | measured mean KLD | top-1 | p99.9 | serialized |
|---|---|---|---|---|---|
GGUF Q8_0 | llama.cpp | 0.001087 | 98.53 % | 0.0351 | 27.05 GiB |
GGUF Q6_K | llama.cpp | 0.002035 | 97.98 % | 0.0794 | 21.31 GiB |
| hydrated EXL3 | vLLM | 0.002700 | 97.80 % | 0.1313 | 20.12 GiB payload |
| online K5/K6 EXL3 | vLLM | 0.003141 | 97.61 % | 0.1447 | — |
| context EXL3 | vLLM | 0.003409 | 97.55 % | 0.1632 | 19.27 GiB payload |
GGUF UD-Q5_K_XL | llama.cpp | 0.004444 | 97.20 % | 0.2144 | 18.83 GiB |
| official FP8 | vLLM | 0.005197 | 96.92 % | 0.2440 | 28.51 GiB resident |
| K4 EXL3 | vLLM | 0.010345 | 95.91 % | 0.5576 | — |
unsloth/Qwen3.8-27B-NVFP4 | vLLM | 0.030115 | 93.16 % | 1.6228 | 22.57 GB checkpoint |
The cross-engine BF16 control is measured the same way, not assumed:
llama.cpp against the vLLM BF16 reference on identical tokens, the shared head
and the same 512 contexts reads 0.000507 mean, 99.07 % top-1 and p99.9
0.0113 (receipts/gguf-report-engine-floor.json). It proves the engine is a
confounder. KL is neither additive nor a metric, so this value cannot be
subtracted from a candidate KL and provides no quantization-only upper or lower
bound. The two — cells are builds that ship
BF16 attention for the runtime to encode at load, so their disk bytes are not a like-for-like
payload; payload figures are immutable_payload_bytes from receipts/collection-index.json
(hydrated 21,610,916,123 B = 20.127 GiB, context 20,696,033,532 B = 19.275 GiB; the table truncates
both to two decimals) and are serialized bytes, never VRAM. The
FP8 figure is resident weights and is labelled as such.
The p99.9 column is each report's exact shard-0 p99.9 as the comparator receipt read them; the tail table above quotes the bin-bounded cumulative estimate from the 560-bin histogram (bins about 5.6 % wide), which is why hydrated reads 0.1319 there and 0.1313 here — each exact value lies inside the bin its estimate names. The two differ by construction, not by measurement.
What the table supports is complete-pipeline ordering, not format attribution:
Q6_K measures 0.002035 and vLLM
hydrated measures 0.002700 on the same contexts.UD-Q5_K_XL measures 0.004444.Both comparisons are cross-engine. They are evidence about the tested artifact-plus-engine pipelines; neither isolates the quantization format.
Correction, 2026-08-16 — the byte axis in the table above is not one axis, and every mixed
comparison flattered us. A GGUF row is the whole file of a text-only artifact; our row is
tensor payload of a multimodal tree that also carries an MTP draft. Read from each artifact's
own tensor table, without downloading any payload
(cross-candidate-byte-accounting.json):
| candidate | file | tensor total | token_embd | output | transformer body | multimodal deployed |
|---|---|---|---|---|---|---|
GGUF Q8_0 | 27.05 | 27.04 | 1.258 (Q8_0) | 1.258 (Q8_0) | 24.526 | 27.92 (+ mmproj-BF16 0.867) |
GGUF Q6_K | 21.31 | 21.30 | 0.971 (Q6_K) | 0.971 (Q6_K) | 19.360 | 22.18 (+ mmproj-BF16 0.867) |
GGUF UD-Q5_K_XL | 18.83 | 18.82 | 0.814 (Q5_K) | 0.971 (Q6_K) | 17.034 | 19.70 (+ mmproj-BF16 0.867) |
| online K5/K6 | 28.50 | 28.47 | 2.368 (BF16) | 0.889 (K6) | 24.119 | 28.50, vision inside |
| K4 | 26.37 | 26.37 | 2.368 (BF16) | 0.593 (K4) | 21.463 | 26.37, vision inside |
| hydrated K5/K6 | 20.13 | 20.10 | 2.368 (BF16) | 0.889 (K6) | 15.726 | 20.13, vision inside |
| context edition | 19.27 | 19.25 | 2.368 (BF16) | 0.889 (K6) | 14.886 | 19.27, vision inside |
All figures GiB. The embedding and head widths are not uniform across GGUF tiers, and the vision
encoder is absent from every GGUF text file - it ships separately as mmproj-BF16.gguf, which no
earlier comparison of ours counted. What that does to the four published claims, two against us and
two for us:
Q6_K against hydrated: our sentence understated their byte spend roughly threefold.
"+1.186 GiB more file" is +1.198 on tensors and +3.634 GiB of transformer body (23.1 % more than
ours). Their lower complete-pipeline KL remains measured; the engine-confounded
comparison does not establish a format-only fidelity win.Q6_K deployment is 22.18 GiB against our 20.13 GiB whole tree,
so ours is 2.053 GiB smaller and needs no second file.UD-Q5_K_XL against the context edition: the old byte claim was wrong against us.
We do not pay "0.445 GiB more" for the lower complete-pipeline KL — on transformer body
they carry 2.148 GiB more (+14.4 %), and our deployed multimodal artifact is
0.422 GiB smaller. The KL comparison remains engine-confounded.Q8_0 against online K5/K6: the bodies agree to 1.7 % (24.526 against 24.119), the
most byte-comparable pair on the table, and the Q8_0 complete pipeline measures lower KL.
Format attribution still requires a same-engine capture.Rule from here on, and it is printed rather than footnoted: "at equal file bytes", "at equal tensor bytes", "at equal transformer body" and "as a deployed multimodal artifact" are four different claims, and whichever one a sentence means is written into the sentence. No fidelity number changes.
Q8_0 has the lowest measured complete-pipeline KL, 0.001087 at 27.05 GiB.
Its engine differs from every vLLM row, so no quantization-only ordering is
inferred from the 0.000507 BF16 control.unsloth's NVFP4.unsloth/Qwen3.8-27B-NVFP4 reads
0.030115 with 93.16 % top-1 on this same shard, under the same engine so with no
cross-engine term: 2.9x K4 at the same weight class, 8.8x the context edition, 27.7x Q8_0.
It loses 512 of 512 contexts to both the context edition and official FP8 with zero ties,
and over the full ten shards reads 0.031059 [0.027916, 0.034795] at 92.90 % top-1, losing
5,120 of 5,120 (receipts/kld5-1M-nvfp4.json, receipts/kld5-10M-nvfp4.json,
receipts/kld5-1M-paired-nvfp4.json). The revision we measured, 9c73e2da, no longer
resolves upstream after a 2026-08-15 history squash; the weights at current HEAD are
byte-identical and we keep the reviewed revision alive as
malaiwah/Qwen3.8-27B-NVFP4-archival-9c73e2da.What it does not settle: this is text-only teacher-forced fidelity on one shard — one tenth of the
suite, though a close tenth: over all 10,480,640 positions the five vLLM means read 0.002760 / 0.003210 /
0.003509 / 0.005294 / 0.010604, 1.9-2.9 % above these shard-0 values with the ordering unchanged
(receipts/kld5-10M-*.json), and the GGUFs have no ten-shard equivalent. It says
nothing about serving 262,144 tokens with vision and MTP on a 32 GB card, which is where these
artifacts actually differ, and llama.cpp KV-quant behaviour, prefill and decode speed are separate
axes that were not measured. A separate control bounds the protocol objection instead of arguing
it: restricting scoring to positions with at least 256 tokens of left context — the floor
llama-perplexity enforces — lowers every candidate's mean by 1.3-2.1 %, and second-half-only
scoring by 3.9-4.9 %, uniformly enough to change no ordering
(receipts/scored-window-offset.json), so the external protocol's scoring floor explains at most
about 5 % of any cross-protocol gap. Protocol identity and the remaining open comparators are in
docs/35 and
docs/29 F2.
Public capability, all six models on one pinned MMLU-Pro subset. 70 questions, 14 official
categories x 5, official five-shot prefixes,
TIGER-Lab/MMLU-Pro@b189ec765aa7ed75c8acfea42df31fdae71f97be, greedy, thinking at low effort,
5,120-token completion cap, every candidate scored item-paired against the BF16 control.
This is a first public, licence-compatible, item-paired benchmark, not a leaderboard claim.
| model | absolute | Wilson 95 % | BF16-pass retention | Wilson lower | regressions | improvements | completion-cap failures | receipt |
|---|---|---|---|---|---|---|---|---|
BF16 Qwen/Qwen3.8-27B | 57/70 (81.4 %) | [70.8 %, 88.8 %] | reference | — | — | — | 4 | receipts/public-capability-bf16.json |
| context edition | 58/70 (82.9 %) | [72.4 %, 89.9 %] | 56/57 | 90.7 % | 1 | 2 | 3 | receipts/public-capability-ctx.json |
| K4 | 57/70 (81.4 %) | [70.8 %, 88.8 %] | 55/57 | 88.1 % | 2 | 2 | 4 | receipts/public-capability-k4.json |
| hydrated | 56/70 (80.0 %) | [69.2 %, 87.7 %] | 54/57 | 85.6 % | 3 | 2 | 4 | receipts/public-capability-hyd.json |
official FP8 Qwen/Qwen3.8-27B-FP8 | 56/70 (80.0 %) | [69.2 %, 87.7 %] | 55/57 | 88.1 % | 2 | 1 | 4 | receipts/public-capability-fp8.json |
| online K5/K6 | 55/70 (78.6 %) | [67.6 %, 86.6 %] | 54/57 | 85.6 % | 3 | 1 | 4 | receipts/public-capability-k5k6.json |
The bar, and who clears it. receipts/public-capability-plan.json pre-registered two
conditions before any candidate ran: BF16-pass retention with a Wilson 95 % lower bound at or
above 0.90, and no category losing more than two BF16 passes. All five candidates satisfy
the second — the worst per-category loss is two passes (hydrated and online K5/K6, philosophy).
Only the context edition satisfies the first, at 90.7 %. K4 and official FP8 read
88.1 %; hydrated and online K5/K6 read 85.6 %. Four of five candidates therefore miss a
bar written down in advance, and that is published rather than restated as a pass.
Why that is a power limit, not a verdict. With 57 BF16 passes as the paired denominator,
56/57 is the smallest count whose Wilson lower bound clears 0.90 — a single paired miss
already fails it — so a 70-item suite has too few items to certify this bar at all, and every
candidate interval overlaps every other. The matrix separates nothing at this size. It applies
equally to official FP8, which misses the same bar; none of this is evidence that any build is
broken. Exact-output agreement is 0/70 for every EXL3 candidate and 1/70 for official
FP8 (a single 113-token math answer) because long chains of thought differ token-wise, so the
pass/fail outcome is the only meaningful pairing. Harness tools/public_capability.py, sweep
runner tools/run_public_capability.sh, superseded 2,048-cap control
receipts/public-capability-bf16-superseded-cap2048.json. The honest next step is more items
and more task types, the plan's own P1 — HumanEval+/MBPP-style executable cases, IFEval-style
constraints, tool schemas and a larger MMLU-Pro draw — not another sweep of the same 70
questions; no capability claim graduates before that.
Three limits stated up front: v5 absolute KLD is not comparable to v3 (K4 reads 0.029679
there and 0.010604 here — only within-suite ordering and paired differences transfer); this run
has no cumulative percentiles over all ten shards, only per-shard p95/p99/p999 and the
exact global maximum, because its shard reports predate the qwen38-fidelity-report/2 KLD
histogram — the tail table above is the shard-0 rerun, 1,048,064 of those 10,480,640
positions; and its hidden
states were deleted shard by shard to fit 135 GB of scratch, so it is reproducible from the
pinned corpus fetch log and suite manifest plus a GPU rather than from published captures.
See PROGRESS.md for the full session record.
| Doc | Contents |
|---|---|
| docs/01-nvfp4-composition.md | Measured tensor-level composition of nvidia/Qwen3.6-27B-NVFP4 and unsloth/Qwen3.8-27B-NVFP4, plus the BF16 parameter census |
| docs/02-recipe-k4.md | The proposed recipe, footprint arithmetic, and the headroom exchange rates |
| docs/03-gg-runtime-contract.md | What the Gilded Gnosis EXL3 loader requires, and what it does not support |
| docs/04-exllamav3-toolchain.md | exllamav3 conversion flags, the missing per-module override, and the splice route |
| docs/05-kld-protocol.md | Superseded iteration-1 protocol: the single-window, full-vocabulary logits KLD inherited from the GG r34 runner. Kept for provenance; the current method is docs/42 |
| docs/06-baseline-validation.md | Running the official GG image with no container runtime, and the two proven baselines |
| docs/07-serving-recommendations.md | Serving-guide differences across upstream / NVIDIA / Unsloth cards, cross-checked against shipped configs |
| docs/08-upstream-cards-digest.md | Per-card digest: declared recipe, benchmarks, harnesses, limitations |
| docs/09-variant-publication.md | How iterative variants are published for independent re-measurement |
| docs/14-fidelity-protocol-v2.md | Hidden-state replay protocol, supersedes the single-window KLD |
| docs/18-results-fidelity-v3.md | Held-out re-measurement, the later offset-independent contamination correction, per-stratum means, and controls |
| docs/22-results-iteration-2.md | Gate K5 / up K5 / down K6: corrected mean KLD 0.007945, 38 % below official FP8 at 71 % of its resident weight |
| docs/15-results-fidelity-v2.md | Superseded v2 run: 151,478 positions, 74/74 paired wins — measured on a corpus that overlapped our calibration data |
| docs/16-head-attribution.md | Is lm_head the sensitive tensor? Measured: not at K6 |
| docs/11-kld-external-comparison.md | Published KLD data for this family; why "FP8 = 0.5" is wrong |
| docs/13-upstream-contributions.md | Upstream issue + verified PR |
| docs/19-cuda-graphs-patch.md | The autotune-priming patch behind PR #314, deliverables and verification |
| docs/20-context-extension-and-k3-gap.md | How far context can be extended, and the distance to the reference protocol |
| docs/12-iteration-2-plan.md | Where to invest next, ranked |
| docs/17-next-iteration-shopping-list.md | Iteration-2 shopping list, each item with its closing test |
| docs/10-results-iteration-1.md | Iteration 1: build, serve, KLD, and the upstream defect |
| docs/21-independent-review-response.md | Independent review, per finding: fixed, fixed-in-v2, or still open |
| docs/24-p0-results.md | P0 done: prefill +113 %, and why fp32 replay was a negative result |
| docs/23-next-attack-list.md | Ranked plan for iteration 3, with evidence, cost and acceptance per item |
| docs/25-goal-pareto-dominate-fp8.md | the goal as a gate, and the iteration-3 verdict per axis |
| docs/26-prefill-attribution.md | where prefill time goes: MLP 2.13-2.26x, attention overlay 1.05-1.11x, hgemm at cuBLAS parity |
| docs/27-graph-decode-drift-control.md | eager-vs-graph drift is ambient: BF16 drifts the same as EXL3 |
| PROGRESS.md | Chronological work log |
| docs/28-external-validation-and-corrections.md | external RTX 5090 validation and the corrected capacity boundary |
| docs/30-iteration-4-context-edition.md | serialized-K5 context build, receipts and rejected vision quant |
| docs/31-frozen-qualification.md | source-disjoint frozen v4 qualification |
| docs/32-native-context-embedding-overlay.md | int8 input table, native MTP-3 plus 8.4 MP engine-budget proof (since superseded by the physical 5090 qualification at utilisation 0.955), and corrected draft accounting |
| docs/33-evidence-volume-and-intervals.md | the 10,480,640-position rerun, the measured tail, and why ten times the positions did not narrow the interval |
| docs/29-plan-and-loose-ends.md | ranked open work and acceptance gates, including the evidence-volume objection (F1) this session closed |
| docs/35-external-protocol-comparability.md | why our KLD and Unsloth's are not interchangeable, the twelve deltas, and the two of them now measured: the scoring-window offset and the cross-engine floor |
| docs/34-vram-class-profiles.md | the 24 GB and 16 GB class verdicts, the measured affine KV law, and a capped-budget proxy qualification |
| docs/36-performance-levers-5090.md | 11 configurations on the physical 5090: what moved throughput, what is a no-op, and the GDN gate reachability rate |
| docs/37-error-driven-allocation.md | the error-driven allocation experiment that lost, the 3.73x-per-bit law it produced, and why the objective is not a surrogate for KLD |
| docs/39-chat-template-audit.md | every chat template in this model family captured with revision and sha256, diffed by named construct, and the verdict that ours is byte-identical to official; the community "fixed" template measured losing tool-call arguments under qwen3_coder, and the loudest complaint traced to the 2.4T flagship's template rather than the 27B's |
| docs/42-kld-method.md | the KLD method of record: the metric, the shared BF16 head and why, the capture point, the suite and its contamination policy, the ladder, the bootstrap and the tail histogram, the harness-revision pin, how to re-derive a published number on a CPU in seconds, every floor that bounds an absolute value, and the claim boundary |
| docs/14-paro-assessment.md | ParoQuant (Pairwise Rotation Quantization, arXiv:2511.10645) assessed: a pre-quantization transform orthogonal to EXL3, not a comparator; the checkpoint is INT4-linear and protocol-incomparable, the learned rotation could replace our fixed Hadamard — follow-up in docs/13 |
| docs/57-eda-allocation-revisit.md | the error-driven allocation revisit, including the later re-solve that refuted sqrt_energy and left a compounding-aware objective as the next defensible in-family method |
| docs/58-qwen36-quant-prior-art.md | prior art of per-module bitrate attribution in the Qwen 3.5/3.6 generation: llama.cpp use_more_bits U-shaped layer heuristic (first+last 1/8 promoted, attn_v special-cased as "more sensitive"), EXL2 Frobenius-norm simulated annealing vs EXL3 direct-KLD greedy optimizer with U-shaped allocation.py layer term, GDN-at-higher-precision as standard NVFP4 practice (driven by correctness bug not measurement), Unsloth finding attn sensitive for hybrid archs (direction matches our physics, mechanism not diagnosed), no published early-vs-late KLD experiment, and five concrete inspirations with cost and falsification |
Tooling in tools/ is what produced the evidence: an unprivileged OCI image
puller and a proot-based runner for it, the BF16 attention splice and checkpoint
finaliser, the fidelity harness (fidelity.py, suite3.py) that builds the suite and
replays captures, the decode-parity probe (decode_parity.py), the kernel
microbenchmarks (prefill_micro*.py, gemm_cmp.py) and the upstream patches as
standalone files. The v5 10.48 M-position run, the public-benchmark harness and the
collection index are these files:
| Tool | What it does | Receipts |
|---|---|---|
| tools/fetch_corpus_v5.py | fetches the five-stratum v5 corpus (941 documents / 70,348,971 bytes) and pins every document by URL and sha256 | receipts/kld5-corpus-fetch-log.json |
| tools/suite3.py | builds the frozen suite: exact-advance non-overlapping windows, whole-document calibration-overlap pre-exclusion (44 of 941), prior-suite token exclusion (160 v4 hashes, 0 reachable) | receipts/kld5-suite-manifest.json (qwen38-distribution-fidelity/6) |
| tools/kld_ladder.sh | walks the ladder one 512-context shard at a time — capture six models, replay five candidates, verify, delete the shard's ~64 GB of hidden states — because scratch is ~135 GB | ten per-shard reports, listed in every cumulative receipt |
| tools/kld_aggregate.py | welds verified per-shard reports into cumulative means, cluster bootstraps, paired comparisons and (for qwen38-fidelity-report/2 inputs) bin-bounded cumulative quantiles | receipts/kld5-10M-{hyd,k5k6,ctx,fp8,k4}.json, receipts/kld5-10M-paired.json |
| tools/fidelity.py | capture/replay harness; now also counts every scored position into a fixed log-spaced kld_tail histogram, bumping reports to qwen38-fidelity-report/2 | per-shard and per-candidate reports |
| tools/public_capability.py | paired MMLU-Pro run against a pinned dataset revision, official five-shot prefixes, greedy, with Wilson intervals and BF16-pass retention | receipts/public-capability-{plan,suite-mmlupro-70,bf16,ctx,k4,hyd,fp8,k5k6}.json, superseded 2,048-cap control kept as receipts/public-capability-bf16-superseded-cap2048.json |
| tools/run_public_capability.sh | drives the six-model sweep one server at a time against the frozen BF16 reference, with the per-candidate environment each build needs (every EXL3 candidate requires VLLM_EXL3_GRAPH_DECODE=1) | the five candidate receipts above |
| tools/collection_index.py | one immutable row per published checkpoint, every field carrying its receipt path, sha256 and RFC 6901 pointer; disk bytes and resident_weights kept strictly distinct | receipts/collection-index.json (qwen38-collection-index/1) |
| tools/gguf_capture.cpp | llama.cpp-side capture of the post-final-norm state for a GGUF at pinned commit ece963f4, so a GGUF can be replayed through our shared BF16 head; bf16 rounding verified bit-identical to torch on 2,012,449 probe values | receipts/gguf-report-{q8_0,q6_k,q5_k_xl,engine-floor}.json, receipts/cross-engine-comparator.json |
| tools/gguf_manifest.py | pins each capture to its GGUF blob digest and the llama.cpp identity that produced it | the four receipts/gguf-report-*.json |
| tools/build_llamacpp.sh | builds the capture binary against the pinned llama.cpp commit | engine_identity in every GGUF report |
| tools/chat_template_matrix.py | renders every candidate chat template over 34 fixed cases inside the pinned image, reconstructing transformers' Jinja environment exactly, and reports token counts and sha256 per rendered prompt; CPU only | receipts/chat-template-audit.json, receipts/chat-templates/renders.json |
| tools/chat_template_roundtrip.py | feeds each template's rendered tool call back through the image's own vllm.parser.qwen3._qwen3_arg_converter and reports whether --tool-call-parser qwen3_coder recovers the arguments losslessly | tool_call_roundtrip in the audit receipt |
| tools/chat_template_prefix.py | first-divergent-token between the generated token stream and the re-rendered history, per client reasoning-echo behaviour, i.e. the point past which no prefix-cache block is reusable | prompt_growth_and_prefix_stability in the audit receipt |
| tools/fetch_chat_templates.sh, tools/gguf_chat_template.py | fetch every family template with its revision and digest, including off-default-branch refs and the one that exists only inside GGUF metadata (read over HTTP range requests, no tensor bytes) | receipts/chat-templates/{SHA256SUMS,raw/} |
ONLINE_QUANT=exl3-b6, 19.21 GB resident, vision verified, published as malaiwah/Qwen3.8-27B-K4malaiwah/Qwen3.8-27B-EXL3-K5K6: 20.32 GiB resident, overlap-corrected mean KLD 0.007945 body-only / 0.008078 as served, top-1 96.86 %ext.hgemm is already at cuBLAS parity — FP8 prefill parity needs a fused dequant-in-epilogue kernel, not tuningreceipts/qualification-5090-context.json, per-process server logs receipts/qualification-5090-context-server-{B3,B4,C,D,E,F}.log): the context edition serves native 262,144 with MTP-3, the full 8,388,608-pixel image ceiling and fp8 KV on one 32,607 MiB card — but at --gpu-memory-utilization 0.955 with PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, not the 0.97 the cards used to print. Measured at startup: engine budget 29.98 GiB, free 30.9 of 31.4 GiB usable, usage 18.19 weight + 1.78 peak activation + 0.27 non-torch + 0.45 CUDAGraph = 20.69 GiB, available KV 9.28 GiB = 265,122 KV tokens, 1.01x maximum concurrency at 262,144, attention block size forced to 1600 tokens so the attention page is at least the mamba page, 3 padding layers at at most 6.25 % KV waste, startup 55.7 s, model load 18.19 GiB in 3.99 s. All seven gates PASS: (1) startup allocates a native-length request inside the utilisation ceiling; (2) the 261,794-token needle retrieved exactly; (3) a combined 236,824-token plus 7,077,888-pixel request returns 1376346594 | red, blue exactly; (4) the 30-case image suite at 24/30; (5) three warmed 256-token concurrency-1 decode runs; (6) a second native-length request after release in the same process; (7) receipt identity completetorch.OutOfMemoryError inside vllm/v1/attention/ops/vit_attn_wrappers.py wanting 62.00 MiB with 26.50 MiB free, and expandable_segments does not rescue it because vLLM spends the freed bytes on more KV (272,570 → 280,017 tokens); lowering max_pixels to 4,194,304 at 0.97 is strictly worse (profiled peak activation falls to 1.35 GiB, KV grows to 291,933 tokens, OOM with 6.56 MiB free). The knob is utilisation, not the image ceiling — seven megapixels is not a hard 32 GB ceiling, since the identical request succeeds at 0.955 with the full 8,388,608-pixel ceilingreceipts/task-retention-v2-summary.json, receipts/task-retention-v2-strict-rescore.json)tools/kld_aggregate.py — hydrated 0.002760, online K5/K6 0.003210, context 0.003509, official FP8 0.005294, K4 0.010604, all body-only through one shared BF16 head (receipts/kld5-10M-*.json)receipts/kld5-10M-paired.json)tools/fidelity.py (qwen38-fidelity-report/2) plus bin-bounded cumulative quantiles in tools/kld_aggregate.py; this run kept only per-shard percentiles and the exact global maximum, and says so in its receiptsTIGER-Lab/MMLU-Pro@b189ec76) — BF16 57/70, context 58/70 at 56/57 retention (Wilson lower 90.7 %, the only candidate clearing the pre-registered 0.90 bar), K4 57/70 and official FP8 56/70 at 55/57 (88.1 %), hydrated 56/70 and online K5/K6 55/70 at 54/57 (85.6 %); the four shortfalls are a 70-item power limit — 56/57 is the smallest count that can clear 0.90 — not a broken build, so the next step is item volume and task diversityreceipts/collection-index.json from tools/collection_index.py, four immutable rows, receipt-traced fields, one disclosed divergence left unfixed rather than rewriting a published receiptQ8_0 0.001087, Q6_K 0.002035 and UD-Q5_K_XL 0.004444 were scored on shard 0 of the v5 suite through the shared BF16 head. The unquantized llama.cpp-versus-vLLM control is 0.000507 and is diagnostic, not subtractable; cross-engine rows therefore compare complete pipelines and do not isolate quantization format (receipts/cross-engine-comparator.json)receipts/scored-window-offset.json)llama-perplexity --kl-divergence on wikitext-2 raw at ctx 512, 147,900 scored positions, base PPL 6.950230 ± 0.044933 — Q8_0 0.000926, Q6_K 0.002286, UD-Q5_K_XL 0.004426, with the same ordering as our suite. The absolute values and ratios are protocol-specific; their harness floor was measured on our hardware at 5.6e-5 to 8.0e-5 and tokenisation was bit-identical over 297,194 tokens. PPL does not reproduce the KLD ordering (receipts/wikitext-kld-run-a.json, docs/35)unsloth/Qwen3.8-27B-NVFP4 measured on our suite, the comparator readers ask for most: 0.031059 [0.027916, 0.034795] at 92.90 % top-1 over 10,480,640 positions, losing 5,120 of 5,120 contexts to both the context edition and official FP8; 2.9x K4 at the same 4-bit weight class (receipts/kld5-10M-nvfp4.json, receipts/kld5-1M-paired-nvfp4.json)receipts/capture-determinism.json)receipts/converter-determinism.json, receipts/sibling-rebuild-fidelity.json)num_speculative_tokens 1 wins at eight (+30.67 %, 313.28 → 409.35 tok/s aggregate, plus 10,911 more KV tokens). Closed as no-ops or losses: explicit FLASHINFER (already auto-selected on SM120), TRITON_ATTN, custom_ops:["all"], and dynamic speculative decoding (refuses to start under FULL_DECODE_ONLY). Prefill did not move on any lever (docs/36)gg-r34-patched-apc is the release unit (receipts/apc-poison-repro.json, receipts/qualification-5090-apc.json)a*L + M (docs/34)malaiwah/qwen38-27b-fidelity-suite-v5, the losing error-driven candidate as a documented negative, and archival mirrors of the NVFP4 and GGUF revisions our published numbers cite — because one of them stopped resolving upstream mid-sessionIQ4_KS_KT) and NVFP4-NInfer comparators — neither of which has any KLD number today (docs/29 F2)559 commits
Python
85.2%
Shell
9.6%
Jinja
2.5%
Cuda
2.3%
Status 2026-08-16: the collection is hardware-qualified, the comparator set is measured, and the method is now reproducible from published artifacts. The current checkpoint is
malaiwah/Qwen3.8-27B-EXL3-K5K6: mean KLD 0.003210 (95 % CI [0.002982, 0.003480]) body-only at 20.32 GiB resident weights, versus officialQwen/Qwen3.8-27B-FP8at 0.005294 — 39 % lower divergence at 71 % of its resident weight, paired -0.002084 [-0.002249, -0.001942] winning 5,105 of 5,120 held-out contexts. The best profile in the collection is the hydrated build at 0.002760, paired -0.002534 on 5,118/5,120. The headline suite is v5: 5,120 contexts x 2,047 positions = 10,480,640 scored positions over 842 source clusters, calibration- and v4-token-disjoint.Landed since: the context edition is qualified on a physical RTX 5090 (all seven gates at
--gpu-memory-utilization 0.955withexpandable_segments, 265,122 KV tokens at 262,144 — and a second physical 5090 needed 0.956, so utilisation is a per-card measurement, not a constant).unsloth's NVFP4 is now on our suite at 0.031059 over 10.48 M positions, losing 5,120 of 5,120 contexts to both the context edition and FP8. The externalllama-perplexityprotocol was run for cross-citation and reproduces our ordering. Capture and replay are bit-reproducible, and a fresh conversion of the same recipe — a sibling, 97.6 % of quantized modules differing in bytes — scores indistinguishable from the published checkpoint, so the recipe determines fidelity while the published bytes remain the artifact. The suite, the shared BF16 reference capture and every per-shard report are published as a dataset, with archival mirrors of the third-party artifacts our numbers cite.Two more artifacts published 2026-08-16, both from pre-registered conversions.
malaiwah/Qwen3.8-27B-EXL3-K6-parityuses nearly the same file-byte budget as GGUFQ6_Kby promotinggate_proj/up_projK5 → K6. It measures 0.001634 [0.001541, 0.001742] under vLLM versusQ6_Kat 0.002035 under llama.cpp, while beating hydrated on 511/512 contexts (-39.5 % for +1.348 GiB) and carrying 2.31 GiB less transformer body thanQ6_K. The cross-engine control proves these are complete-pipeline results; it cannot be subtracted to establish format parity or prove that the earlier gap was caused by bytes.malaiwah/Qwen3.8-27B-EXL3-S16-V-researchis the opposite: the rejected sub-4-bit 16 GB candidate at 0.045374 [0.041959, 0.049351], 1.5x its pre-registered NO threshold, losing 512/512 contexts to K4 (4.39x) and to the context edition (13.31x). It is published because it failed - a NO nobody can audit is an assertion - and it is the empirical anchor for the K3 rung of the per-bit law (docs/37 §3.3). Both payloads were predicted to the byte by the affine law before conversion, and both registered intervals are reported with their misses: S16-V's point estimate was 1.52x too high, K6-parity's interval fell 2.0 % short.The earlier headline — 0.007945 body-only on 136 v3 contexts / 278,392 positions (docs/22) — stands exactly as measured and is superseded only as the headline; v5 absolute KLD is not comparable to v3 (the two suites differ by a measured 2.46-3.08x with identical ordering), so quote paired differences across suites, never means. Figures labelled K4 or v1 belong to iteration 1 and are superseded. One earlier control was withdrawn: the "CUDA-graph parity 0.000000" receipt captured a prefill forward, which
FULL_DECODE_ONLYnever captures, so it could not have measured decode; the decode probe that replaces it is docs/27. Open items are tracked in docs/29.
The primary control is the exact PROFILE=fidelity tuple, not
PROFILE=throughput and not a memory-modified derivative. The objective is to
recover enough PP/TTFT from that control to pass the absolute candidate gates
without sacrificing its paired KLD/tails, TG, context, or capability envelope.
The fidelity control is allowed to fail the 7,000 tok/s PP candidate threshold;
that threshold governs promotion of a new candidate, not acceptance of the
control.
receipts/frontier-fidelity-control.json
pins the model/profile identity, INT6 embedding overlay, exact evidence hashes,
reference metrics, candidate rules, and the prohibition on substituting the
high-KLD throughput profile. Every INT8/FP4/FP6, K-map, topology, kernel-route,
KV, MTP, graph, or profile change is a candidate variable and requires
preregistered paired comparison against this exact control.
The exact control and method-neutral harness are ready. Clean packed INT6 was
integrated onto runtime b19029d, matched the historical encoder byte-for-byte
on real BF16 rows, and passed three independent AIBoss boots. Aggregate PP was
2,941.9 tok/s, fox/essay TG 226.9/104.3, context 238,400, with text,
vision, MTP, 200k context and rollback passing on every boot.
The clean-runtime fidelity recapture is exactly equal to the historical control: mean KLD 0.0034052972, p99 0.0348892, paired delta and CI 0.0, with all 512 per-context rows and the full tail histogram identical. Its 513-file, 10.73 GB hidden-state capture has an independently verified private-bucket copy.
The repaired v3 harness uses the actual EXL3 Viterbi quantizer—not the invalidated
uniform proxy—on four real 128×128 Qwen slices with matched BF16/quant-flow
activations and Hessians. RTN and both LDLQ controls carry the same exact
10,756-byte / 5.251953-bpw payload. Quant-flow-H LDLQ reduced mean excess
running-output error from RTN's 0.0011858 to 0.0008865 at those same bytes;
this is a control result, not a method promotion. Whole-block and end-logit gates
remain mandatory and intentionally pending for researcher-proposed methods.
receipts/frontier-g02/manifest.json is
validated by python3 tools/verify_frontier_g02.py --campaign receipts/frontier-g02 as pre_candidate_ready. Candidate slots remain unopened,
E-final-v7 remains sealed, and rental authorization remains false.
The preregistered campaign in
docs/61-final-frontier-requant-stack-plan.md
ended at its local RTX 5090 gate with the terminal disposition
three_candidate_no_go. The complete hashed record is
receipts/frontier-g01/manifest.json;
python3 tools/verify_frontier_g01.py --campaign receipts/frontier-g01
validates every referenced G0 artifact, preregistration, result and terminal
decision and exits 0.
G0 passed: the immutable BF16 payload was fully censused (1,199 logical
tensors), converter/runtime environments were locked, the fidelity-derived
incumbent-fidelity-int8 clean runtime was qualified, and every AIBoss
maintenance transaction proved service restoration. That historical G0 arm
used INT8 embeddings and did not remeasure the exact PROFILE=fidelity KLD; it
is runtime/rollback evidence, not a substitute for the frozen primary control.
G1 then consumed the three frozen candidate opportunities:
| candidate | measured result | frozen Gate A failure |
|---|---|---|
| full gate/up FP6 | startup rejected the 199,104-token configuration; estimated maximum 192,576 | context below 238,400 |
| full gate/up FP4 | PP 4,393.6 tok/s; fox/essay TG 211.6/98.8; context and capability checks passed | PP below 7,000 |
| gate/up FP4, language layers 0–31 | PP 3,529.5 tok/s; fox/essay TG 215.9/96.0; context and capability checks passed | PP below 7,000 |
The preregistered hard-gate rule stopped KLD scoring after each deterministic runtime failure. Supporting experiments also closed without a promotable candidate: independent seeded converter A/A runs had 24 payload mismatches; the pinned AutoRound consumer could not build; ResComp improved one actual 16×16 weight-tile SSE by 21.1% but lacked a valid whole-block implementation; and the hybrid QKV producer/consumer pilot passed its component checks without the mandatory full-checkpoint, startup, graph and logit gates. No paid rental or new model release is authorized by this campaign; the measured serving profiles below remain current.


Three profiles, selected by PROFILE= in
patches/run-qwen38-27b.sh. All three pass
tools/verify-profile.sh (exit 0, 8/8 checks each, including a 200k-token
prompt). PP/TG are the tools/bench-profile.sh harness at n=3 boots;
KLD is 512 contexts of shard-0000 against the BF16 reference.
PROFILE=throughput | PROFILE=fidelity | PROFILE=balanced | |
|---|---|---|---|
| weights | all-FP4 | all-trellis (K5K6 as shipped) | trellis + gate_up MXFP6 |
| PP, 2051-tok | 9638.9 ± 18.3 tok/s | 2987.7 ± 4.4 tok/s | 3925.2 ± 13.1 tok/s |
| TG fox / essay | 187.4 ± 0.6 / 94.3 ± 0.0 | 228.3 ± 0.4 / 104.1 ± 0.1 | 215.6 ± 0.2 / 103.7 ± 0.1 |
| MTP acceptance fox / essay | 0.930 / 0.298 | 1.000 / 0.304 | 1.000 / 0.324 |
| KLD mean | 0.063759 | 0.003405 [0.003166, 0.003672] | 0.005672 [0.005302, 0.006087] |
| KLD p99 | 0.7010 | 0.034889 | 0.059908 |
| max context | 249,600 | 238,400 | 199,104 |
| vision + MTP | pass | pass | pass |
| criteria met | 4/6 (fails KLD) | 5/6 (fails PP) | 5/6 (fails ctx) |
fidelity serves the checkpoint at KLD 0.003405 — within 26% of this
collection's own published trellis fidelity (0.002700) — with TG 228.3 tok/s
and full 238,400 context; the residual over the checkpoint is the int6 embedding
table (~0.0007), not any GEMM approximation. It costs ~3.2x prefill.
throughput is the only profile above 7000 tok/s prefill.
Long-context needle retrieval on the fidelity profile (fixed harness, single
needle per context): 8/8 at 2k, 8/8 at 100k, 8/8 at 195k — 24/24. This is
an easy proxy — finding a planted needle does not establish unimpaired
long-context reasoning, only that the attention window and KV cache are intact
to 195k.
No single profile meets all six north-star criteria. Concretely: throughput
fails KLD (0.063759, well above the 0.012 budget); fidelity fails the 7000 tok/s
prefill criterion (2987.7); balanced fails the context criterion (199,104
tokens, below the 238,400 that fidelity achieves). The other two profiles each
clear five of six.
receipts/frontier-2026-08-19.md shows why
from three directions: prefill-grade throughput needs the MLP resident in a
GEMM-ready format, trellis-grade fidelity needs weights that are decoded per
prefill chunk, and GEMM-resident FP8 for all 24.3e9 quantized parameters is
22.6 GiB — which cannot coexist with a 238,400-token KV cache in 31.4 GiB. The
blocker is memory, not kernel quality.
Fixes landed while getting here, with receipts: an engine-fatal OOM on any
prompt over ~4k tokens (max_num_batched_tokens 8192 → 3072, which also
freed 0.93 GiB of KV), +32% TG by discovering the MTP draft loop ran eager
on the V1 model runner, and the first profile of this stack (prefill is 92%
GPU-bound; decode was 2%). Upstream: vllm-project/vllm#52871, #52872,
local-inference-lab/vllm#439, #440, b12x #232/#233/#234.
A baked serving image is published at
docker.io/malaiwah/qwen38-27b-exl3-gg:r34-p2-41a5d16 (also tagged :latest).
It layers every patch this project bind-mounts on top of the digest-pinned
Gilded Gnosis r34 base (voipmonitor/vllm:gilded-gnosis-v20-vllm4d006a4-b12xcd3ce19-fi1ac6942-cu132-20260810-r34,
sha256 820181fbbc975cd5291c411cda9771d58fecee1636d916f508f47230df20592b),
so the repo is not required at runtime — podman run plus a Hugging Face
cache mount is sufficient. All 10 patches (7 vLLM source overlays and 3
conversion modules) are baked in by docker/Containerfile,
which also bakes the CUDA lib64 symlink the B12X JIT needs as real links and
carries OCI provenance labels (org.opencontainers.image.source, .revision,
.description, ai.malaiwah.base-digest).
The image was gated 9/9 mount-free (NO_PATCH_MOUNTS=1, every patch
bind-mount disabled): PP 2933.0, fox 229.1, essay 103.7, ctx 238,400, 200k-token
prompt OK, vision OK, MTP acceptance_fox 1.000 — confirming the image is
self-contained.
PROFILE=throughput|balanced|fidelity selects the serving profile;
NO_PATCH_MOUNTS=1 is what makes the baked image run mount-free. The full flag
set is in patches/run-qwen38-27b.sh. A
standalone fidelity invocation — every flag below is taken from that script —
is:
podman run --replace -d \
--name qwen38-27b \
--device nvidia.com/gpu=all --ipc=host --network host \
--tmpfs /usr/local/cuda-13.2/lib64:rw,size=16m \
-e HF_HUB_OFFLINE=1 \
-e CUTE_DSL_ARCH=sm_120a -e FLASHINFER_CUDA_ARCH_LIST=12.0f \
-e OMP_NUM_THREADS=8 -e CUDA_DEVICE_MAX_CONNECTIONS=32 \
-e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True -e SAFETENSORS_FAST_GPU=1 \
-e VLLM_WORKER_MULTIPROC_METHOD=spawn \
-e VLLM_USE_V2_MODEL_RUNNER=1 -e VLLM_NO_USAGE_STATS=1 -e DO_NOT_TRACK=1 \
-e XDG_CACHE_HOME=/cache/jit -e CUDA_CACHE_PATH=/cache/jit \
-e TRITON_CACHE_DIR=/cache/jit/triton -e TORCHINDUCTOR_CACHE_DIR=/cache/jit/torchinductor \
-e FLASHINFER_WORKSPACE_BASE=/cache/jit/flashinfer \
-e VLLM_EXL3_MULTIPRECISION=1 -e VLLM_EXL3_GRAPH_DECODE=1 \
-e VLLM_EXL3_FP4_TRITON_DECODE=0 -e VLLM_EXL3_FP4_PER_ROW_GS=0 -e VLLM_EXL3_FP4_DRAFT_HEAD=0 \
-e VLLM_EXL3_ONLINE_TRELLIS_BITS=6 \
-e VLLM_EXL3_ONLINE_CACHE_DIR=/cache/jit/exl3-online -e VLLM_EXL3_ONLINE_CACHE_MODE=readwrite \
-e VLLM_EXL3_EMBED_ONLINE_BITS=6 -e B12X_PACKED_B_MIN_N=1024 \
-e VLLM_EXL3_B12X_ANY_BITS=1 -e VLLM_EXL3_B12X_MIN_M=128 -e VLLM_EXL3_SKIP_TRELLIS_PREP=0 \
-e VLLM_EXL3_PREFILL_RECONSTRUCT_M=1 -e VLLM_EXL3_PREFILL_RECONSTRUCT_MAX_MB=512 \
-e VLLM_EXL3_PREFILL_RECONSTRUCT_CACHE=0 -e VLLM_EXL3_FOLD_FP32_BUDGET_MB=48 \
-e VLLM_EXL3_FP4_LAYERS=, \
-e VLLM_EXL3_EXT_PATH=/opt/exllamav3 \
-e HF_HOME=/root/.cache/huggingface \
-v ~/.cache/huggingface/hub:/root/.cache/huggingface:ro \
-v ~/.cache/jit:/cache/jit \
--entrypoint /bin/bash \
docker.io/malaiwah/qwen38-27b-exl3-gg:r34-p2-41a5d16 \
-lc 'set -euo pipefail; \
ln -sf /usr/local/cuda-13.2/targets/x86_64-linux/lib/* /usr/local/cuda-13.2/lib64/ 2>/dev/null || true; \
exec vllm serve \
/root/.cache/huggingface/models--malaiwah--Qwen3.8-27B-EXL3-K5K6-hydrated/snapshots/ab3a91a13813df8096cb4c1d560ed3669035d0cf \
--served-model-name Qwen3.8-27B --trust-remote-code \
--host 0.0.0.0 --port 8000 \
--quantization exl3 \
--quantization-config "{\"linear\":{\"weight\":\"mxfp8\"},\"ignore\":[\"re:.*visual\\..*\",\"re:.*in_proj_a$\",\"re:.*in_proj_b$\",\"re:.*in_proj_ba$\",\"re:.*mtp\\..*\",\"lm_head\"]}" \
--attention-backend TRITON_ATTN \
--gpu-memory-utilization 0.945 --kv-cache-dtype fp8_e4m3 \
--max-model-len 238400 --max-num-seqs 4 --max-num-batched-tokens 3072 \
--compilation-config "{\"mode\":\"NONE\",\"cudagraph_mode\":\"FULL_DECODE_ONLY\"}" \
--mm-processor-kwargs "{\"truncation\":false}" --mm-processor-cache-type shm \
--default-chat-template-kwargs "{\"preserve_thinking\": true}" \
--enable-chunked-prefill --reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_coder \
--speculative-config "{\"method\":\"mtp\",\"num_speculative_tokens\":6}"'
Flags taken from run-qwen38-27b.sh: --tmpfs + ln -sf symlink (B12X JIT
workaround, lines 225/296 — the baked image already carries these symlinks, so
the --tmpfs is redundant with the baked image but harmless), --device /
--ipc / --network (line 226), the -e environment block (lines 231–280),
the HF cache mount -v …/huggingface:ro (line 281), the JIT cache volume
-v …/cache/jit:/cache/jit (line 288), and the vllm serve argument list
(lines 303–322). For balanced or throughput, change the profile env vars
per the case block in the script (lines 83–154).
K4, EXL3-K5K6)Research materials and progress log for building a dense EXL3 quant of
Qwen/Qwen3.8-27B that is inspired by
NVIDIA's NVFP4 recipe for the previous generation, but keeps the attention projections
BF16 on disk for runtime K6 encoding by the Gilded Gnosis vLLM fork, and
serializes the MLP as EXL3 K5 gate / K5 up / K6 down with a K6 mcg head and a
quantized MTP draft. Iteration 1 serialized the whole MLP at K4.
Design goal: NVFP4-class VRAM footprint, lower KLD. Met on fidelity, memory, decode and speculative decode; prefill is a measured structural deficit of a 4-bit-class trellis format in this runtime (docs/25, docs/26).
Held-out v5 suite, receipts/kld5-suite-manifest.json (schema
qwen38-distribution-fidelity/6, suite_token_sha256 510541f6…09482b88): 5,120 contexts
x 2,047 positions = 10,480,640 scored positions, 842 source clusters, 1,024 contexts per
stratum. 44 of 941 corpus documents were excluded whole before window selection for any
all-position 12-token overlap with exllamav3 calibration data, leaving 897 eligible, so
contamination hits are 0 by construction; all 160 v4 context token hashes were seeded as
exclusions and 0 were reachable. Every figure below is body-only — both operands scored
through one shared BF16 LM head — with a 95 % CI from a 10,000-resample bootstrap over the 842
source clusters.
| candidate | mean KLD | 95 % CI | top-1 | paired vs FP8 | contexts won | receipt |
|---|---|---|---|---|---|---|
| hydrated | 0.002760 | [0.002540, 0.003020] | 97.70 % | -0.002534 | 5,118 / 5,120 | receipts/kld5-10M-hyd.json |
| online K5/K6 | 0.003210 | [0.002982, 0.003480] | 97.52 % | -0.002084 | 5,105 / 5,120 | receipts/kld5-10M-k5k6.json |
| context | 0.003509 | [0.003220, 0.003852] | 97.44 % | -0.001785 | 5,109 / 5,120 | receipts/kld5-10M-ctx.json |
| official FP8 | 0.005294 | [0.004927, 0.005728] | 96.79 % | — | — | receipts/kld5-10M-fp8.json |
| K4 (iteration 1) | 0.010604 | [0.009640, 0.011746] | 95.76 % | +0.005310 | 7 / 5,120 | receipts/kld5-10M-k4.json |
Paired intervals and win counts come from receipts/kld5-10M-paired.json (10,000 resamples,
seed 1); hydrated also beats the runtime overlay directly, -0.000450 [-0.000469,
-0.000433] on 4,922/5,120. Exact worst single position over the whole run: hydrated 8.258,
online 22.241, context 5.557, FP8 10.714, K4 14.283. The cumulative hydrated mean at the
1M / 2M / 5M / 10M ladder checkpoints is 0.002700 / 0.002759 / 0.002699 / 0.002760
(shards[] of receipts/kld5-10M-hyd.json).
The tail, not just the mean. The ten-shard run predates the histogram, so the tail was
measured by re-running shard 0 of the same v5 suite with the qwen38-fidelity-report/2
harness: 512 contexts x 2,047 positions = 1,048,064 scored positions, the same contexts
every candidate saw. Quantiles are bin-bounded — 560 log-spaced bins from 1e-12 to 1e2
nats, one bin ~5.6 % wide, and every receipt carries lower / upper / estimate per
quantile — while the maxima and the exceedance counts are exact. The two columns worth
reading are p99.9 and the share of positions above 0.1 nats.
| candidate | mean | p50 | p95 | p99 | p99.9 | p99.99 | exact max | above 0.1 | above 1.0 | receipt |
|---|---|---|---|---|---|---|---|---|---|---|
| hydrated | 0.002700 | 0.00109 | 0.0082 | 0.0276 | 0.1319 | 0.463 | 3.735 | 0.1534 % | 0.00219 % | receipts/kld5-1M-tail-hyd.json |
| online K5/K6 | 0.003141 | 0.00128 | 0.0099 | 0.0321 | 0.1446 | 0.498 | 5.507 | 0.1820 % | 0.00200 % | receipts/kld5-1M-tail-k5k6.json |
| context | 0.003409 | 0.00135 | 0.0107 | 0.0357 | 0.1642 | 0.587 | 3.749 | 0.2287 % | 0.00305 % | receipts/kld5-1M-tail-ctx.json |
| official FP8 | 0.005197 | 0.00202 | 0.0167 | 0.0531 | 0.2438 | 0.812 | 5.296 | 0.3912 % | 0.00592 % | receipts/kld5-1M-tail-fp8.json |
| K4 | 0.010345 | 0.00320 | 0.0332 | 0.1194 | 0.5555 | 1.870 | 7.565 | 1.2604 % | 0.03807 % | receipts/kld5-1M-tail-k4.json |
The ordering at p50, p95, p99, p99.9 and p99.99 is the same as the ordering of the means, so
for these candidates the mean is not hiding a worse tail: every EXL3 K5/K6-class build has a
lighter tail than official FP8 at every measured quantile, and K4 is worse than FP8 at
every quantile. Receipts are schema qwen38-kld-ladder-cumulative/2, welded by
tools/kld_aggregate.py from /2 replay reports.
Against GGUF, measured, including the part that goes against us. The standing critique is
that official FP8 is a throughput format whose quality is Q4-to-Q5 class, so beating it is a weak
claim, and that Q8_0 and Q6_K are the honest bar. That is now measured rather than argued.
Three GGUFs from unsloth/Qwen3.8-27B-GGUF@f1bfb127c64f7072bdd2cad55f258b9c8b2910fe were
captured under llama.cpp pinned at commit ece963f41b0b02d7a0d61436ae365762c073a4c8 with
tools/gguf_capture.cpp, which reads the post-final-norm state — the same mathematical point our
vLLM hook takes, with bf16 rounding verified bit-identical to torch on 2,012,449 probe values —
and scored through the same shared BF16 head on shard 0 of the v5 suite: the same 512
contexts and the same 1,048,064 positions every candidate above saw (tools/gguf_manifest.py,
tools/build_llamacpp.sh; receipt receipts/cross-engine-comparator.json, per-candidate reports
receipts/gguf-report-{q8_0,q6_k,q5_k_xl}.json).
| candidate | engine | measured mean KLD | top-1 | p99.9 | serialized |
|---|---|---|---|---|---|
GGUF Q8_0 | llama.cpp | 0.001087 | 98.53 % | 0.0351 | 27.05 GiB |
GGUF Q6_K | llama.cpp | 0.002035 | 97.98 % | 0.0794 | 21.31 GiB |
| hydrated EXL3 | vLLM | 0.002700 | 97.80 % | 0.1313 | 20.12 GiB payload |
| online K5/K6 EXL3 | vLLM | 0.003141 | 97.61 % | 0.1447 | — |
| context EXL3 | vLLM | 0.003409 | 97.55 % | 0.1632 | 19.27 GiB payload |
GGUF UD-Q5_K_XL | llama.cpp | 0.004444 | 97.20 % | 0.2144 | 18.83 GiB |
| official FP8 | vLLM | 0.005197 | 96.92 % | 0.2440 | 28.51 GiB resident |
| K4 EXL3 | vLLM | 0.010345 | 95.91 % | 0.5576 | — |
unsloth/Qwen3.8-27B-NVFP4 | vLLM | 0.030115 | 93.16 % | 1.6228 | 22.57 GB checkpoint |
The cross-engine BF16 control is measured the same way, not assumed:
llama.cpp against the vLLM BF16 reference on identical tokens, the shared head
and the same 512 contexts reads 0.000507 mean, 99.07 % top-1 and p99.9
0.0113 (receipts/gguf-report-engine-floor.json). It proves the engine is a
confounder. KL is neither additive nor a metric, so this value cannot be
subtracted from a candidate KL and provides no quantization-only upper or lower
bound. The two — cells are builds that ship
BF16 attention for the runtime to encode at load, so their disk bytes are not a like-for-like
payload; payload figures are immutable_payload_bytes from receipts/collection-index.json
(hydrated 21,610,916,123 B = 20.127 GiB, context 20,696,033,532 B = 19.275 GiB; the table truncates
both to two decimals) and are serialized bytes, never VRAM. The
FP8 figure is resident weights and is labelled as such.
The p99.9 column is each report's exact shard-0 p99.9 as the comparator receipt read them; the tail table above quotes the bin-bounded cumulative estimate from the 560-bin histogram (bins about 5.6 % wide), which is why hydrated reads 0.1319 there and 0.1313 here — each exact value lies inside the bin its estimate names. The two differ by construction, not by measurement.
What the table supports is complete-pipeline ordering, not format attribution:
Q6_K measures 0.002035 and vLLM
hydrated measures 0.002700 on the same contexts.UD-Q5_K_XL measures 0.004444.Both comparisons are cross-engine. They are evidence about the tested artifact-plus-engine pipelines; neither isolates the quantization format.
Correction, 2026-08-16 — the byte axis in the table above is not one axis, and every mixed
comparison flattered us. A GGUF row is the whole file of a text-only artifact; our row is
tensor payload of a multimodal tree that also carries an MTP draft. Read from each artifact's
own tensor table, without downloading any payload
(cross-candidate-byte-accounting.json):
| candidate | file | tensor total | token_embd | output | transformer body | multimodal deployed |
|---|---|---|---|---|---|---|
GGUF Q8_0 | 27.05 | 27.04 | 1.258 (Q8_0) | 1.258 (Q8_0) | 24.526 | 27.92 (+ mmproj-BF16 0.867) |
GGUF Q6_K | 21.31 | 21.30 | 0.971 (Q6_K) | 0.971 (Q6_K) | 19.360 | 22.18 (+ mmproj-BF16 0.867) |
GGUF UD-Q5_K_XL | 18.83 | 18.82 | 0.814 (Q5_K) | 0.971 (Q6_K) | 17.034 | 19.70 (+ mmproj-BF16 0.867) |
| online K5/K6 | 28.50 | 28.47 | 2.368 (BF16) | 0.889 (K6) | 24.119 | 28.50, vision inside |
| K4 | 26.37 | 26.37 | 2.368 (BF16) | 0.593 (K4) | 21.463 | 26.37, vision inside |
| hydrated K5/K6 | 20.13 | 20.10 | 2.368 (BF16) | 0.889 (K6) | 15.726 | 20.13, vision inside |
| context edition | 19.27 | 19.25 | 2.368 (BF16) | 0.889 (K6) | 14.886 | 19.27, vision inside |
All figures GiB. The embedding and head widths are not uniform across GGUF tiers, and the vision
encoder is absent from every GGUF text file - it ships separately as mmproj-BF16.gguf, which no
earlier comparison of ours counted. What that does to the four published claims, two against us and
two for us:
Q6_K against hydrated: our sentence understated their byte spend roughly threefold.
"+1.186 GiB more file" is +1.198 on tensors and +3.634 GiB of transformer body (23.1 % more than
ours). Their lower complete-pipeline KL remains measured; the engine-confounded
comparison does not establish a format-only fidelity win.Q6_K deployment is 22.18 GiB against our 20.13 GiB whole tree,
so ours is 2.053 GiB smaller and needs no second file.UD-Q5_K_XL against the context edition: the old byte claim was wrong against us.
We do not pay "0.445 GiB more" for the lower complete-pipeline KL — on transformer body
they carry 2.148 GiB more (+14.4 %), and our deployed multimodal artifact is
0.422 GiB smaller. The KL comparison remains engine-confounded.Q8_0 against online K5/K6: the bodies agree to 1.7 % (24.526 against 24.119), the
most byte-comparable pair on the table, and the Q8_0 complete pipeline measures lower KL.
Format attribution still requires a same-engine capture.Rule from here on, and it is printed rather than footnoted: "at equal file bytes", "at equal tensor bytes", "at equal transformer body" and "as a deployed multimodal artifact" are four different claims, and whichever one a sentence means is written into the sentence. No fidelity number changes.
Q8_0 has the lowest measured complete-pipeline KL, 0.001087 at 27.05 GiB.
Its engine differs from every vLLM row, so no quantization-only ordering is
inferred from the 0.000507 BF16 control.unsloth's NVFP4.unsloth/Qwen3.8-27B-NVFP4 reads
0.030115 with 93.16 % top-1 on this same shard, under the same engine so with no
cross-engine term: 2.9x K4 at the same weight class, 8.8x the context edition, 27.7x Q8_0.
It loses 512 of 512 contexts to both the context edition and official FP8 with zero ties,
and over the full ten shards reads 0.031059 [0.027916, 0.034795] at 92.90 % top-1, losing
5,120 of 5,120 (receipts/kld5-1M-nvfp4.json, receipts/kld5-10M-nvfp4.json,
receipts/kld5-1M-paired-nvfp4.json). The revision we measured, 9c73e2da, no longer
resolves upstream after a 2026-08-15 history squash; the weights at current HEAD are
byte-identical and we keep the reviewed revision alive as
malaiwah/Qwen3.8-27B-NVFP4-archival-9c73e2da.What it does not settle: this is text-only teacher-forced fidelity on one shard — one tenth of the
suite, though a close tenth: over all 10,480,640 positions the five vLLM means read 0.002760 / 0.003210 /
0.003509 / 0.005294 / 0.010604, 1.9-2.9 % above these shard-0 values with the ordering unchanged
(receipts/kld5-10M-*.json), and the GGUFs have no ten-shard equivalent. It says
nothing about serving 262,144 tokens with vision and MTP on a 32 GB card, which is where these
artifacts actually differ, and llama.cpp KV-quant behaviour, prefill and decode speed are separate
axes that were not measured. A separate control bounds the protocol objection instead of arguing
it: restricting scoring to positions with at least 256 tokens of left context — the floor
llama-perplexity enforces — lowers every candidate's mean by 1.3-2.1 %, and second-half-only
scoring by 3.9-4.9 %, uniformly enough to change no ordering
(receipts/scored-window-offset.json), so the external protocol's scoring floor explains at most
about 5 % of any cross-protocol gap. Protocol identity and the remaining open comparators are in
docs/35 and
docs/29 F2.
Public capability, all six models on one pinned MMLU-Pro subset. 70 questions, 14 official
categories x 5, official five-shot prefixes,
TIGER-Lab/MMLU-Pro@b189ec765aa7ed75c8acfea42df31fdae71f97be, greedy, thinking at low effort,
5,120-token completion cap, every candidate scored item-paired against the BF16 control.
This is a first public, licence-compatible, item-paired benchmark, not a leaderboard claim.
| model | absolute | Wilson 95 % | BF16-pass retention | Wilson lower | regressions | improvements | completion-cap failures | receipt |
|---|---|---|---|---|---|---|---|---|
BF16 Qwen/Qwen3.8-27B | 57/70 (81.4 %) | [70.8 %, 88.8 %] | reference | — | — | — | 4 | receipts/public-capability-bf16.json |
| context edition | 58/70 (82.9 %) | [72.4 %, 89.9 %] | 56/57 | 90.7 % | 1 | 2 | 3 | receipts/public-capability-ctx.json |
| K4 | 57/70 (81.4 %) | [70.8 %, 88.8 %] | 55/57 | 88.1 % | 2 | 2 | 4 | receipts/public-capability-k4.json |
| hydrated | 56/70 (80.0 %) | [69.2 %, 87.7 %] | 54/57 | 85.6 % | 3 | 2 | 4 | receipts/public-capability-hyd.json |
official FP8 Qwen/Qwen3.8-27B-FP8 | 56/70 (80.0 %) | [69.2 %, 87.7 %] | 55/57 | 88.1 % | 2 | 1 | 4 | receipts/public-capability-fp8.json |
| online K5/K6 | 55/70 (78.6 %) | [67.6 %, 86.6 %] | 54/57 | 85.6 % | 3 | 1 | 4 | receipts/public-capability-k5k6.json |
The bar, and who clears it. receipts/public-capability-plan.json pre-registered two
conditions before any candidate ran: BF16-pass retention with a Wilson 95 % lower bound at or
above 0.90, and no category losing more than two BF16 passes. All five candidates satisfy
the second — the worst per-category loss is two passes (hydrated and online K5/K6, philosophy).
Only the context edition satisfies the first, at 90.7 %. K4 and official FP8 read
88.1 %; hydrated and online K5/K6 read 85.6 %. Four of five candidates therefore miss a
bar written down in advance, and that is published rather than restated as a pass.
Why that is a power limit, not a verdict. With 57 BF16 passes as the paired denominator,
56/57 is the smallest count whose Wilson lower bound clears 0.90 — a single paired miss
already fails it — so a 70-item suite has too few items to certify this bar at all, and every
candidate interval overlaps every other. The matrix separates nothing at this size. It applies
equally to official FP8, which misses the same bar; none of this is evidence that any build is
broken. Exact-output agreement is 0/70 for every EXL3 candidate and 1/70 for official
FP8 (a single 113-token math answer) because long chains of thought differ token-wise, so the
pass/fail outcome is the only meaningful pairing. Harness tools/public_capability.py, sweep
runner tools/run_public_capability.sh, superseded 2,048-cap control
receipts/public-capability-bf16-superseded-cap2048.json. The honest next step is more items
and more task types, the plan's own P1 — HumanEval+/MBPP-style executable cases, IFEval-style
constraints, tool schemas and a larger MMLU-Pro draw — not another sweep of the same 70
questions; no capability claim graduates before that.
Three limits stated up front: v5 absolute KLD is not comparable to v3 (K4 reads 0.029679
there and 0.010604 here — only within-suite ordering and paired differences transfer); this run
has no cumulative percentiles over all ten shards, only per-shard p95/p99/p999 and the
exact global maximum, because its shard reports predate the qwen38-fidelity-report/2 KLD
histogram — the tail table above is the shard-0 rerun, 1,048,064 of those 10,480,640
positions; and its hidden
states were deleted shard by shard to fit 135 GB of scratch, so it is reproducible from the
pinned corpus fetch log and suite manifest plus a GPU rather than from published captures.
See PROGRESS.md for the full session record.
| Doc | Contents |
|---|---|
| docs/01-nvfp4-composition.md | Measured tensor-level composition of nvidia/Qwen3.6-27B-NVFP4 and unsloth/Qwen3.8-27B-NVFP4, plus the BF16 parameter census |
| docs/02-recipe-k4.md | The proposed recipe, footprint arithmetic, and the headroom exchange rates |
| docs/03-gg-runtime-contract.md | What the Gilded Gnosis EXL3 loader requires, and what it does not support |
| docs/04-exllamav3-toolchain.md | exllamav3 conversion flags, the missing per-module override, and the splice route |
| docs/05-kld-protocol.md | Superseded iteration-1 protocol: the single-window, full-vocabulary logits KLD inherited from the GG r34 runner. Kept for provenance; the current method is docs/42 |
| docs/06-baseline-validation.md | Running the official GG image with no container runtime, and the two proven baselines |
| docs/07-serving-recommendations.md | Serving-guide differences across upstream / NVIDIA / Unsloth cards, cross-checked against shipped configs |
| docs/08-upstream-cards-digest.md | Per-card digest: declared recipe, benchmarks, harnesses, limitations |
| docs/09-variant-publication.md | How iterative variants are published for independent re-measurement |
| docs/14-fidelity-protocol-v2.md | Hidden-state replay protocol, supersedes the single-window KLD |
| docs/18-results-fidelity-v3.md | Held-out re-measurement, the later offset-independent contamination correction, per-stratum means, and controls |
| docs/22-results-iteration-2.md | Gate K5 / up K5 / down K6: corrected mean KLD 0.007945, 38 % below official FP8 at 71 % of its resident weight |
| docs/15-results-fidelity-v2.md | Superseded v2 run: 151,478 positions, 74/74 paired wins — measured on a corpus that overlapped our calibration data |
| docs/16-head-attribution.md | Is lm_head the sensitive tensor? Measured: not at K6 |
| docs/11-kld-external-comparison.md | Published KLD data for this family; why "FP8 = 0.5" is wrong |
| docs/13-upstream-contributions.md | Upstream issue + verified PR |
| docs/19-cuda-graphs-patch.md | The autotune-priming patch behind PR #314, deliverables and verification |
| docs/20-context-extension-and-k3-gap.md | How far context can be extended, and the distance to the reference protocol |
| docs/12-iteration-2-plan.md | Where to invest next, ranked |
| docs/17-next-iteration-shopping-list.md | Iteration-2 shopping list, each item with its closing test |
| docs/10-results-iteration-1.md | Iteration 1: build, serve, KLD, and the upstream defect |
| docs/21-independent-review-response.md | Independent review, per finding: fixed, fixed-in-v2, or still open |
| docs/24-p0-results.md | P0 done: prefill +113 %, and why fp32 replay was a negative result |
| docs/23-next-attack-list.md | Ranked plan for iteration 3, with evidence, cost and acceptance per item |
| docs/25-goal-pareto-dominate-fp8.md | the goal as a gate, and the iteration-3 verdict per axis |
| docs/26-prefill-attribution.md | where prefill time goes: MLP 2.13-2.26x, attention overlay 1.05-1.11x, hgemm at cuBLAS parity |
| docs/27-graph-decode-drift-control.md | eager-vs-graph drift is ambient: BF16 drifts the same as EXL3 |
| PROGRESS.md | Chronological work log |
| docs/28-external-validation-and-corrections.md | external RTX 5090 validation and the corrected capacity boundary |
| docs/30-iteration-4-context-edition.md | serialized-K5 context build, receipts and rejected vision quant |
| docs/31-frozen-qualification.md | source-disjoint frozen v4 qualification |
| docs/32-native-context-embedding-overlay.md | int8 input table, native MTP-3 plus 8.4 MP engine-budget proof (since superseded by the physical 5090 qualification at utilisation 0.955), and corrected draft accounting |
| docs/33-evidence-volume-and-intervals.md | the 10,480,640-position rerun, the measured tail, and why ten times the positions did not narrow the interval |
| docs/29-plan-and-loose-ends.md | ranked open work and acceptance gates, including the evidence-volume objection (F1) this session closed |
| docs/35-external-protocol-comparability.md | why our KLD and Unsloth's are not interchangeable, the twelve deltas, and the two of them now measured: the scoring-window offset and the cross-engine floor |
| docs/34-vram-class-profiles.md | the 24 GB and 16 GB class verdicts, the measured affine KV law, and a capped-budget proxy qualification |
| docs/36-performance-levers-5090.md | 11 configurations on the physical 5090: what moved throughput, what is a no-op, and the GDN gate reachability rate |
| docs/37-error-driven-allocation.md | the error-driven allocation experiment that lost, the 3.73x-per-bit law it produced, and why the objective is not a surrogate for KLD |
| docs/39-chat-template-audit.md | every chat template in this model family captured with revision and sha256, diffed by named construct, and the verdict that ours is byte-identical to official; the community "fixed" template measured losing tool-call arguments under qwen3_coder, and the loudest complaint traced to the 2.4T flagship's template rather than the 27B's |
| docs/42-kld-method.md | the KLD method of record: the metric, the shared BF16 head and why, the capture point, the suite and its contamination policy, the ladder, the bootstrap and the tail histogram, the harness-revision pin, how to re-derive a published number on a CPU in seconds, every floor that bounds an absolute value, and the claim boundary |
| docs/14-paro-assessment.md | ParoQuant (Pairwise Rotation Quantization, arXiv:2511.10645) assessed: a pre-quantization transform orthogonal to EXL3, not a comparator; the checkpoint is INT4-linear and protocol-incomparable, the learned rotation could replace our fixed Hadamard — follow-up in docs/13 |
| docs/57-eda-allocation-revisit.md | the error-driven allocation revisit, including the later re-solve that refuted sqrt_energy and left a compounding-aware objective as the next defensible in-family method |
| docs/58-qwen36-quant-prior-art.md | prior art of per-module bitrate attribution in the Qwen 3.5/3.6 generation: llama.cpp use_more_bits U-shaped layer heuristic (first+last 1/8 promoted, attn_v special-cased as "more sensitive"), EXL2 Frobenius-norm simulated annealing vs EXL3 direct-KLD greedy optimizer with U-shaped allocation.py layer term, GDN-at-higher-precision as standard NVFP4 practice (driven by correctness bug not measurement), Unsloth finding attn sensitive for hybrid archs (direction matches our physics, mechanism not diagnosed), no published early-vs-late KLD experiment, and five concrete inspirations with cost and falsification |
Tooling in tools/ is what produced the evidence: an unprivileged OCI image
puller and a proot-based runner for it, the BF16 attention splice and checkpoint
finaliser, the fidelity harness (fidelity.py, suite3.py) that builds the suite and
replays captures, the decode-parity probe (decode_parity.py), the kernel
microbenchmarks (prefill_micro*.py, gemm_cmp.py) and the upstream patches as
standalone files. The v5 10.48 M-position run, the public-benchmark harness and the
collection index are these files:
| Tool | What it does | Receipts |
|---|---|---|
| tools/fetch_corpus_v5.py | fetches the five-stratum v5 corpus (941 documents / 70,348,971 bytes) and pins every document by URL and sha256 | receipts/kld5-corpus-fetch-log.json |
| tools/suite3.py | builds the frozen suite: exact-advance non-overlapping windows, whole-document calibration-overlap pre-exclusion (44 of 941), prior-suite token exclusion (160 v4 hashes, 0 reachable) | receipts/kld5-suite-manifest.json (qwen38-distribution-fidelity/6) |
| tools/kld_ladder.sh | walks the ladder one 512-context shard at a time — capture six models, replay five candidates, verify, delete the shard's ~64 GB of hidden states — because scratch is ~135 GB | ten per-shard reports, listed in every cumulative receipt |
| tools/kld_aggregate.py | welds verified per-shard reports into cumulative means, cluster bootstraps, paired comparisons and (for qwen38-fidelity-report/2 inputs) bin-bounded cumulative quantiles | receipts/kld5-10M-{hyd,k5k6,ctx,fp8,k4}.json, receipts/kld5-10M-paired.json |
| tools/fidelity.py | capture/replay harness; now also counts every scored position into a fixed log-spaced kld_tail histogram, bumping reports to qwen38-fidelity-report/2 | per-shard and per-candidate reports |
| tools/public_capability.py | paired MMLU-Pro run against a pinned dataset revision, official five-shot prefixes, greedy, with Wilson intervals and BF16-pass retention | receipts/public-capability-{plan,suite-mmlupro-70,bf16,ctx,k4,hyd,fp8,k5k6}.json, superseded 2,048-cap control kept as receipts/public-capability-bf16-superseded-cap2048.json |
| tools/run_public_capability.sh | drives the six-model sweep one server at a time against the frozen BF16 reference, with the per-candidate environment each build needs (every EXL3 candidate requires VLLM_EXL3_GRAPH_DECODE=1) | the five candidate receipts above |
| tools/collection_index.py | one immutable row per published checkpoint, every field carrying its receipt path, sha256 and RFC 6901 pointer; disk bytes and resident_weights kept strictly distinct | receipts/collection-index.json (qwen38-collection-index/1) |
| tools/gguf_capture.cpp | llama.cpp-side capture of the post-final-norm state for a GGUF at pinned commit ece963f4, so a GGUF can be replayed through our shared BF16 head; bf16 rounding verified bit-identical to torch on 2,012,449 probe values | receipts/gguf-report-{q8_0,q6_k,q5_k_xl,engine-floor}.json, receipts/cross-engine-comparator.json |
| tools/gguf_manifest.py | pins each capture to its GGUF blob digest and the llama.cpp identity that produced it | the four receipts/gguf-report-*.json |
| tools/build_llamacpp.sh | builds the capture binary against the pinned llama.cpp commit | engine_identity in every GGUF report |
| tools/chat_template_matrix.py | renders every candidate chat template over 34 fixed cases inside the pinned image, reconstructing transformers' Jinja environment exactly, and reports token counts and sha256 per rendered prompt; CPU only | receipts/chat-template-audit.json, receipts/chat-templates/renders.json |
| tools/chat_template_roundtrip.py | feeds each template's rendered tool call back through the image's own vllm.parser.qwen3._qwen3_arg_converter and reports whether --tool-call-parser qwen3_coder recovers the arguments losslessly | tool_call_roundtrip in the audit receipt |
| tools/chat_template_prefix.py | first-divergent-token between the generated token stream and the re-rendered history, per client reasoning-echo behaviour, i.e. the point past which no prefix-cache block is reusable | prompt_growth_and_prefix_stability in the audit receipt |
| tools/fetch_chat_templates.sh, tools/gguf_chat_template.py | fetch every family template with its revision and digest, including off-default-branch refs and the one that exists only inside GGUF metadata (read over HTTP range requests, no tensor bytes) | receipts/chat-templates/{SHA256SUMS,raw/} |
ONLINE_QUANT=exl3-b6, 19.21 GB resident, vision verified, published as malaiwah/Qwen3.8-27B-K4malaiwah/Qwen3.8-27B-EXL3-K5K6: 20.32 GiB resident, overlap-corrected mean KLD 0.007945 body-only / 0.008078 as served, top-1 96.86 %ext.hgemm is already at cuBLAS parity — FP8 prefill parity needs a fused dequant-in-epilogue kernel, not tuningreceipts/qualification-5090-context.json, per-process server logs receipts/qualification-5090-context-server-{B3,B4,C,D,E,F}.log): the context edition serves native 262,144 with MTP-3, the full 8,388,608-pixel image ceiling and fp8 KV on one 32,607 MiB card — but at --gpu-memory-utilization 0.955 with PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, not the 0.97 the cards used to print. Measured at startup: engine budget 29.98 GiB, free 30.9 of 31.4 GiB usable, usage 18.19 weight + 1.78 peak activation + 0.27 non-torch + 0.45 CUDAGraph = 20.69 GiB, available KV 9.28 GiB = 265,122 KV tokens, 1.01x maximum concurrency at 262,144, attention block size forced to 1600 tokens so the attention page is at least the mamba page, 3 padding layers at at most 6.25 % KV waste, startup 55.7 s, model load 18.19 GiB in 3.99 s. All seven gates PASS: (1) startup allocates a native-length request inside the utilisation ceiling; (2) the 261,794-token needle retrieved exactly; (3) a combined 236,824-token plus 7,077,888-pixel request returns 1376346594 | red, blue exactly; (4) the 30-case image suite at 24/30; (5) three warmed 256-token concurrency-1 decode runs; (6) a second native-length request after release in the same process; (7) receipt identity completetorch.OutOfMemoryError inside vllm/v1/attention/ops/vit_attn_wrappers.py wanting 62.00 MiB with 26.50 MiB free, and expandable_segments does not rescue it because vLLM spends the freed bytes on more KV (272,570 → 280,017 tokens); lowering max_pixels to 4,194,304 at 0.97 is strictly worse (profiled peak activation falls to 1.35 GiB, KV grows to 291,933 tokens, OOM with 6.56 MiB free). The knob is utilisation, not the image ceiling — seven megapixels is not a hard 32 GB ceiling, since the identical request succeeds at 0.955 with the full 8,388,608-pixel ceilingreceipts/task-retention-v2-summary.json, receipts/task-retention-v2-strict-rescore.json)tools/kld_aggregate.py — hydrated 0.002760, online K5/K6 0.003210, context 0.003509, official FP8 0.005294, K4 0.010604, all body-only through one shared BF16 head (receipts/kld5-10M-*.json)receipts/kld5-10M-paired.json)tools/fidelity.py (qwen38-fidelity-report/2) plus bin-bounded cumulative quantiles in tools/kld_aggregate.py; this run kept only per-shard percentiles and the exact global maximum, and says so in its receiptsTIGER-Lab/MMLU-Pro@b189ec76) — BF16 57/70, context 58/70 at 56/57 retention (Wilson lower 90.7 %, the only candidate clearing the pre-registered 0.90 bar), K4 57/70 and official FP8 56/70 at 55/57 (88.1 %), hydrated 56/70 and online K5/K6 55/70 at 54/57 (85.6 %); the four shortfalls are a 70-item power limit — 56/57 is the smallest count that can clear 0.90 — not a broken build, so the next step is item volume and task diversityreceipts/collection-index.json from tools/collection_index.py, four immutable rows, receipt-traced fields, one disclosed divergence left unfixed rather than rewriting a published receiptQ8_0 0.001087, Q6_K 0.002035 and UD-Q5_K_XL 0.004444 were scored on shard 0 of the v5 suite through the shared BF16 head. The unquantized llama.cpp-versus-vLLM control is 0.000507 and is diagnostic, not subtractable; cross-engine rows therefore compare complete pipelines and do not isolate quantization format (receipts/cross-engine-comparator.json)receipts/scored-window-offset.json)llama-perplexity --kl-divergence on wikitext-2 raw at ctx 512, 147,900 scored positions, base PPL 6.950230 ± 0.044933 — Q8_0 0.000926, Q6_K 0.002286, UD-Q5_K_XL 0.004426, with the same ordering as our suite. The absolute values and ratios are protocol-specific; their harness floor was measured on our hardware at 5.6e-5 to 8.0e-5 and tokenisation was bit-identical over 297,194 tokens. PPL does not reproduce the KLD ordering (receipts/wikitext-kld-run-a.json, docs/35)unsloth/Qwen3.8-27B-NVFP4 measured on our suite, the comparator readers ask for most: 0.031059 [0.027916, 0.034795] at 92.90 % top-1 over 10,480,640 positions, losing 5,120 of 5,120 contexts to both the context edition and official FP8; 2.9x K4 at the same 4-bit weight class (receipts/kld5-10M-nvfp4.json, receipts/kld5-1M-paired-nvfp4.json)receipts/capture-determinism.json)receipts/converter-determinism.json, receipts/sibling-rebuild-fidelity.json)num_speculative_tokens 1 wins at eight (+30.67 %, 313.28 → 409.35 tok/s aggregate, plus 10,911 more KV tokens). Closed as no-ops or losses: explicit FLASHINFER (already auto-selected on SM120), TRITON_ATTN, custom_ops:["all"], and dynamic speculative decoding (refuses to start under FULL_DECODE_ONLY). Prefill did not move on any lever (docs/36)gg-r34-patched-apc is the release unit (receipts/apc-poison-repro.json, receipts/qualification-5090-apc.json)a*L + M (docs/34)malaiwah/qwen38-27b-fidelity-suite-v5, the losing error-driven candidate as a documented negative, and archival mirrors of the NVFP4 and GGUF revisions our published numbers cite — because one of them stopped resolving upstream mid-sessionIQ4_KS_KT) and NVFP4-NInfer comparators — neither of which has any KLD number today (docs/29 F2)559 commits
Python
85.2%
Shell
9.6%
Jinja
2.5%
Cuda
2.3%