Strata fork: Qwen3.8-Flash-Next 125B MoE in ModelOpt NVFP4 on one RTX 20/30/40/50 card (12 GB+, measured on an RTX 5090) + 64 GB of RAM or more. W4A8 prompt path, AVX-512 CPU experts, a low-RAM tier, pictures on the CPU, Cyrillic MTP drafts. Windows / CUDA 13.
C++
18
717 commits
updated Oct 4, 2026
Qwen3.8-Flash-Next (125B hybrid MoE) in NVFP4 on one RTX 20, 30, 40 or 50 card (12 GB of VRAM or more; built and measured on an RTX 5090) + 64 GB of RAM or more, text and pictures. A fork of Niko1221/Strata that runs a ModelOpt NVFP4 checkpoint — here OrcaRouter's abliterated Flash-Next — instead of the Q2/Q3 quants Strata ships for. NVFP4 keeps the experts at 4.5 bits with calibrated scales, which is why this fork exists: the 2-3 bit quants were noticeably less accurate on the same model.
Strata itself keeps the routed experts in RAM (63 GiB here; with less than 96 GB of RAM this fork keeps only the ones outside VRAM and reads the rest from the SSD), caches the most-used ones in VRAM, computes the misses on the CPU and over PCIe in parallel with the GPU, and decodes with an MTP draft head. Everything about that design is upstream's; the original README is kept as README.upstream.md.
Runs an NVFP4 checkpoint (ModelOpt, 4.5-bit experts with calibrated scales) instead of Strata's Q2/Q3 quants: lossless repack, per-expert FP32 scales kept and applied to each projection's output.
Nothing is rounded coarser than the checkpoint where it can be avoided: the n-gram (PLE) table in its shipped FP8 (upstream reads it as IQ4_NL, 8% off per row), the token embedding in its shipped BF16, RoPE angles from a float64 table in every kernel (upstream's fast-math angles were ~0.02 rad off at 262K), FP32-exact inputs to the prompt path's BF16 projections (router, indexer, gates), int8 KV behind a Hadamard rotation.
An int8 (W4A8) prompt path for Blackwell: llama.cpp builds NVFP4 MMQ as FP4 x FP4 on sm_120; this keeps 8-bit activations - 16x smaller error per product - at 86% of its speed.
AVX-512 NVFP4 kernels for the CPU share of the experts (1.8-3.7x ggml-cpu).
Faster start: experts read unbuffered into the pinned arena on their own thread, registered with CUDA one layer ahead of the readers, beside everything else - the arena is in at ~7 s from a PCIe 5 drive.
The K/V grows with the context: at a 262K window upstream allocates the whole context's K/V at start (3.35 GiB
here with the draft layer's) - room for ~1,200 more experts in VRAM. Here the K/V takes VRAM only as the requests
reach further, from the expert cache, and gives it back after them (CUDA virtual memory: no address moves, no graph
is captured again). The window is never smaller; --no-kv-grow allocates it whole.
Faster prompt reading: an MMQ group's experts gathered in one launch after one wait (under WDDM every streamed expert's wait and event had left ~10 us of GPU idle, 226 ms of a 32K prompt), the n-gram rows read 256 at a time and beside layer 0, upstream's split hyper-connection kernels on.
An adaptive VRAM tier with a longer memory: it re-ranks the experts every 2 rounds, up to 192 swaps, and the
routing counts fade x0.92 per pass. Upstream re-ranks every 4 rounds, up to 96, x0.7. In fixed 1,000-token runs on
the 5090 it missed 30-40% fewer experts and the round did not change measurably. About half of those runs' tokens
came after the answer had ended, though, where the model loops over a few experts, so the gain on an answer alone
is not measured yet. It applies to a cache holding 20-60% of the experts with every expert in RAM. Elsewhere
upstream's settings stay; --adapt-every, --adapt-swaps and --adapt-decay override.
Tuned for NVFP4's larger experts (PCIe share, prompt chunks up to 32K, fused scale passes, a verify commit that overlaps the draft) and the fine-tune's own abliterated MTP draft head.
A draft vocabulary with Cyrillic: the MTP draft head proposes only tokens of its subset, and upstream's held
142 of the vocabulary's 18,580 Cyrillic tokens - an answer in Cyrillic decoded at 83 tokens/s with 1.4 tokens a round;
with the whole Cyrillic script (tools/draft_vocab.py --add cyrillic), 109 and 2.1. English is unchanged.
Upstream's CJK subset is one --add cjk away (data/draft_vocab_en.bin is the English/code one).
64 GB of RAM is enough: with less than 96 GB installed the engine runs upstream's file tier
(--mmap-experts --resident-budget-gib) with a budget of the free RAM less 6 GiB: the experts outside VRAM,
hottest first, pinned; the rest is read from experts.bin when needed - unbuffered, in merged requests (this
fork's addition, upstream since 0.1.38 as #362). The whole 63 GiB arena needs 96 GB. On an RTX 5090 + 64 GB a 32K
prompt waits 9.3 s instead of 5.7, and the answer after it runs at 97 tokens/s.
Every RTX 20, 30, 40 and 50 card with 12 GB or more: the NVFP4 path needed Blackwell only for the optional
FP4 x FP4 prompt path, which falls back; the release carries code for all four generations, each one's own path
tested on the 5090 (built as its PTX, the host answering as that card: STRATA_EMULATE_CC).
Images on the CPU, from the checkpoint's own vision tower: the encoder (strata-vision, FP32 weights) runs in
RAM, so the expert cache keeps all its VRAM and text decodes as fast as without images; a picture takes 2-6 s.
Upstream's GPU encoder took ~1.6 GB of VRAM (~600 cached experts, 10-20% of decode speed in served runs here) and, measured
against an FP32 reference, was up to 11% off (ggml-cuda's FP16 flash attention); this one is 0.1% off.
A picture an agent opens itself (Claude Code's Read) reaches the encoder too: it arrives inside a tool result,
which the server used to turn into text only (fixed here, sent upstream as #529).
Fixes: a scale fold that left NVFP4 hidden activations in FP16's subnormals (2-12% expert error), and the batched verify path skipping the query rotation of rotated KV caches.
A more accurate experts pack, re-quantized from the BF16 checkpoint (optional, tools/requant.py):
On upstream Strata 0.1.39. Since 0.1.38 upstream added:
"parallel": N, #465);--vram-elastic, #533);This fork's elastic K/V is off beside --vram-elastic and parallel: each of those manages the cache's VRAM its
own way. #554 (the vision markers in a message's text) now uses 0.1.39's own way of keeping a quoted <think> as
text, and #567's incremental tokenizing sits behind 0.1.39's prompt encoder. Upstream's byte-sized prompt ring
(#583) is kept: on this pack it reads at the same speed and borrows 1.2 GiB less. Against 0.1.38-nvfp4.2, with the
same expert cache, the logits and tokens are identical after 2K on the GPTQ + Q8_0-down pack, and decode
rounds are 6.9% shorter (15.7 against 16.9 ms, upstream's #646).
0.1.39-nvfp4.2: a long Claude Code conversation turned into "!" mid-reply, and every later request of it answered "!" until the engine was restarted by hand. The server now starts a fresh engine after such a reply (one bad reply instead of a broken session). The shared expert's fused SwiGLU quantizer keeps its fp16 scale finite, as 0.1.39's other two do (sent upstream as #838; not this incident's cause). What corrupted the state is not found yet: docs/NVFP4.md lists what was ruled out. Logits and tokens are identical to 0.1.39-nvfp4.1 on both packs.
0.1.38-nvfp4.2: the VRAM reserve is 700 MiB again on Windows (1500 no longer avoided a stall on 0.1.38, only cost expert slots). From other people's upstream pull requests: a prompt tokenised from the last shared prefix (#567: 6.8 ms instead of 128 ms at 125K tokens), token accounting across reasoning continuations (#615), and no lost 401/403 on Windows (#594). docs/NVFP4.md lists what else was measured and why it was not taken.
Fixes from a code review (0.1.37-nvfp4.2): the prompt path borrows 1.25 GiB less VRAM at a 32K chunk; the elastic K/V no longer runs past its mapped cells when it cannot lend slots; pictures Claude Code reads that the server cannot read become a note instead of a 400; safer image sources. Each upstream bug went up as a pull request (#546, #547, #550, #553, #554, #555; 0.1.38 fixed #546 its own way); docs/NVFP4.md has the list.
Each change was measured - first-token KL against a reference, and interleaved speed A/B runs; docs/NVFP4.md has the numbers, and everything that was tried and dropped.
GPU: an RTX 20, 30, 40 or 50 card with 12 GB of VRAM or more (the release has code for sm_75, 86, 89 and
120, and the optional FP4 x FP4 prompt unit for 120a). Built and measured on an RTX 5090, 32 GB: dense weights and
the 262K KV cache first, then ~19 GB of cached experts (~7,200). The other generations were tested on the 5090
through their own code paths (docs/NVFP4.md, "Other GPUs"). A card with less VRAM caches fewer experts and decodes
slower; lower --max-context with it (measured with 0.1.28-nvfp4.4):
| card's VRAM | --max-context | expert slots | decode, measured* |
|---|---|---|---|
| 32 GB (RTX 5090) | 262144 | 7,352 | 111 tok/s |
| 24 GB (RTX 3090 / 4090) | 131072 | 5,158 | 95 tok/s |
| 16 GB (RTX 4080 / 5080 / 4060 Ti 16 GB) | 65536 | 2,422 | 67 tok/s |
| 12 GB (RTX 3060 12 GB / 4070) | 32768 | 1,055 | 59 tok/s |
* On the RTX 5090 with the smaller card's VRAM budget (--vram-reserve-mib); a real card's own compute and PCIe
make it slower. 12 GB at 262144 and 8 GB cards at any context stop with no VRAM is left for the expert cache.
RAM: 64 GB minimum. With 96 GB or more the engine pins all 63 GiB of experts (~69 GiB of physical RAM in
all, measured with 128 GB). With less, the low-RAM mode starts by itself (--low-ram; --no-low-ram turns it
off; --ram-budget GIB caps it): the experts outside VRAM, hottest first, pinned up to the free RAM minus 6 GiB;
the rest is read from the file when needed. With 64 GB and an RTX 5090 all 44 GiB of them fit:
| 64 GB of RAM (measured*) with an RTX 5090 | |
|---|---|
| a 32K prompt: to the first token | 9.3 s (96+ GB: 5.7 s) |
| decode after it | 97 tok/s (0.1.28-nvfp4.4: 72) |
| start, to the session | ~16 s (the profile fill 4.3 s, the 44 GiB budget 9.5 s, read unbuffered) |
* On the RTX 5090 + 128 GB PC with 60 GiB of RAM locked away by a large-page ballast (58 GiB stay available, as on a 64 GB PC whose Windows uses 6 GB). The smaller cards' 64 GB figures (0.1.28-nvfp4.4: 24 GB 86-90 tok/s, 16 GB 54-56) were not measured again on this release.
Pagefile: Windows lets all processes together commit at most RAM + pagefile, and the engine commits ~98 GiB with all experts pinned, ~70-85 GiB in the low-RAM mode (the pinned RAM plus what WDDM reserves for the GPU's allocations; nothing of the model is ever paged out). With Windows and the usual apps on top: a pagefile of at least 32 GB with 128 GB of RAM, 64 GB with 96 GB, 48 GB with 64 GB. Set a fixed minimum rather than relying on a system-managed file to grow in time. Too small, and the start fails with an allocation error.
Disk: ~200 GB for the model files: GGUF 74 GB, expert pack 70 GB, n-gram table 51 GB, image encoder 1.8 GB, embedding 1.3 GB, MTP head 0.8 GB. ~340 GB while preparing them (the 135 GB checkpoint and the MTP intermediates can go afterwards). Use the fastest NVMe drive you have: every start reads 63 GiB.
RTX 5090 (32 GB, PCIe 5 x16), Ryzen 9 9950X3D, 128 GB DDR5-5600, Samsung 9100 PRO, Windows 11, CUDA 13.3. 262,144-token context, KV cache int8, large pages on. The machine is also a desktop: runs vary by ~5%.
| Writes a chat answer (~520 tokens, to its end) | 138 tokens/s on average, 125-146 in 6 runs; 18.2 ms a round |
| Writes a ~1,000-token answer after a 32K prompt | 132 tokens/s on average, 130-133 in 4 runs; 20.0 ms a round |
| The same with the re-quantized pack (GPTQ + Q8_0 down in 27 layers) | chat 126 tokens/s (-12%), after a 32K prompt 116 (-15%), the prompt read as fast |
| Reads a 32K prompt | 6,740-6,930 tokens/s (0.1.31-nvfp4.1: 5,400-5,900) |
| Start: the expert arena loaded | ~7 s (63 GiB of experts read at 10-11 GiB/s) |
0.1.31-nvfp4.3 measured the same in the same interleaved runs: 18.1 and 20.0 ms a round. Its notes said 161-173 tokens/s after a 32K prompt, but that prompt's answer ends after 22 tokens. Those runs went on to a fixed 256 tokens with no end-of-turn stop, so that was the speed of the loop after the answer, not of an answer. The rates vary with the draft acceptance of the path the tokens take (docs/NVFP4.md, "On upstream 0.1.32"). With 64 GB of RAM (the low-RAM mode): 9.3 s to the first token of a 32K prompt instead of 5.7, then 97 tokens/s.
Where precision was still being lost, first-token KL divergence from the more exact variant (8 prompts of 1K-8K tokens plus one of 32K; two runs of the same configuration differ by a median 0.000007):
| source | KL mean | at 32K | now |
|---|---|---|---|
| RoPE: fast-math float angles | - | 0.0063 | float64 table |
| token embedding stored as Q8_0 | 0.0025 | 0.0041 | BF16 as shipped (--embd-gguf) |
| prompt path: BF16-rounded activations into BF16 projections | 0.0023 | 0.0093 | hi + lo split (0.0029 left: not the hyper-connection) |
| n-gram table as IQ4_NL | 0.0026 | - | FP8 as shipped |
| int8 KV without rotation (vs fp16 KV) | 0.0017 | 0.0068 | rotated: 0.0011 / 0.0038 |
| prompt path W4A8 (vs FP16 activations) | 0.0018 | - | default; fp16 is 24% slower |
Needs Visual Studio 2022 Build Tools, CUDA 13+ (sm_120 wants 13), CMake, Ninja, Python 3.11+. From an x64 Native Tools prompt:
git clone https://github.com/ggml-org/llama.cpp third_party\llama.cpp
git -C third_party\llama.cpp checkout 3cf03257f219afbe7334045ff7c6a06ac68c627d
cmake -G Ninja -S . -B build -DCMAKE_BUILD_TYPE=Release -DSTRATA_ENABLE_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=120 ^
-DSTRATA_GGML_DIR=%CD%\third_party\llama.cpp
cmake --build build --target strata
Without -DSTRATA_GGML_DIR CMake fetches the same llama.cpp commit itself. 120 is enough: CMake builds the one
FP4 x FP4 unit for 120a by itself, and the rest runs on sm_121 too.
python -m venv .venv
.venv\Scripts\python -m pip install numpy torch safetensors transformers sentencepiece
set PYTHONPATH=third_party\llama.cpp\gguf-py
:: 1. the checkpoint (126 GiB)
hf download jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 --local-dir models\orca-nvfp4
:: 2. the n-gram (PLE) table, 51.2 GB: its FP8 bytes copied as they are (read from the SSD, never loaded into RAM)
.venv\Scripts\python tools\ple_fp8_pack.py --model models\orca-nvfp4 --out models\ple-fp8.gguf
:: 3. the token embedding, 1.3 GB: BF16 as shipped (the GGUF below stores it as Q8_0)
.venv\Scripts\python tools\embd_bf16_pack.py --model models\orca-nvfp4 --out models\token-embd-bf16.gguf
:: 4. GGUF (NVFP4 experts, Q8_0/BF16 dense) and the pack (experts.bin 63 GiB, tokenizer)
.venv\Scripts\python tools\nvfp4_convert.py --model models\orca-nvfp4 --outfile models\orca-nvfp4.gguf
.venv\Scripts\python tools\iq_pack.py --gguf models\orca-nvfp4.gguf --out packs\orca-nvfp4
:: 5. the fine-tune's own MTP draft head
.venv\Scripts\python tools\mtp_extract.py --model models\orca-nvfp4 --out mtp-orca
.venv\Scripts\python tools\mtp_pack.py --src mtp-orca --experts q2_0 --out mtp-orca\mtp-q2_0.gguf
.venv\Scripts\python tools\mtp_rt.py --gguf mtp-orca\mtp-q2_0.gguf --out mtp-orca\rt
copy data\draft_vocab.bin mtp-orca\rt\
Put packs\ and the GGUF on the fastest drive you have: the start is a 63 GiB read.
Optional: the more accurate experts pack. It needs the BF16 checkpoint (360 GB) and ~2 hours on an RTX 5090. The pack from step 4 stays the dense, index and tokenizer source.
hf download orcarouter/Qwen3.8-Flash-Next-Uncensored --local-dir models\orca-bf16
:: calibration: the engine's own MoE inputs of every layer for the four token files in data\requant_calib (uk, en, code, chat)
set STRATA_DUMP_MOE_INPUT=calib\d
set STRATA_DUMP_MOE_LAYER=all
build\strata.exe <the usual arguments> --tokens-file data\requant_calib\calib_uk.txt --max-new 1 (and en, code, chat)
:: errors per layer and method, then the plan for a size budget
.venv\Scripts\python tools\requant.py analyze --bf16 models\orca-bf16 --calib calib\d --out calib\errors.jsonl
.venv\Scripts\python tools\requant_plan.py calib\errors.jsonl calib 74
:: the pack: GPTQ NVFP4, Q8_0 down in the planned layers
.venv\Scripts\python tools\requant.py pack --bf16 models\orca-bf16 --calib calib\d --plan calib\plan_d8_74.json --base packs\orca-nvfp4 --out packs\orca-nvfp4-gptq-q8d
Then run with --pack packs\orca-nvfp4-gptq-q8d. The engine reads mixed expert formats from 0.1.32-nvfp4.2 on.
One-shot:
build\strata.exe --pack packs\orca-nvfp4 --native models\orca-nvfp4.gguf --native-dense-gguf models\orca-nvfp4.gguf ^
--ple-gguf models\ple-fp8.gguf --embd-gguf models\token-embd-bf16.gguf ^
--mtp mtp-orca\rt --spec 4 --spec-min-p 0.5 --prefill auto ^
--expert-profile data\expert-profile.bin --expert-cache auto ^
--max-context 262144 --kv int8 --tokens-file prompt.txt --max-new 256
OpenAI-compatible server: save the same arguments as a config ({"exe": "build/strata.exe", "args": [...], "tokenizer": "packs/orca-nvfp4/tokenizer", "host": "127.0.0.1", "port": 8097}) and start
python -m serve.server --engine strata --config that.json.
--native and --native-dense-gguf both point at the GGUF: --native alone would take the PLE file for a second
shard of the same model.
Images: the image encoder comes from the checkpoint (convert_hf_to_gguf.py <checkpoint> --mmproj --outtype f32,
1.8 GB) and is built with release\build-vision.cmd (CPU only). The engine takes --vision, the config a
"vision": {"exe": "build-vision-cpu/bin/strata-vision.exe", "mmproj": "models/mmproj-f32.gguf", "model": "models/orca-nvfp4.gguf", "gpu": false, "max_tokens": 1024} entry; then OpenAI image_url parts and Anthropic
image blocks (a screenshot pasted into Claude Code) work. The encoder uses one thread per core while it runs; the
engine is idle then.
set ANTHROPIC_BASE_URL=http://127.0.0.1:8097
set ANTHROPIC_API_KEY=local
set ANTHROPIC_MODEL=strata-nvfp4
set ANTHROPIC_SMALL_FAST_MODEL=strata-nvfp4
claude
The model is hybrid (Gated DeltaNet layers keep a recurrent state that cannot be cut back to a position), so an
agent turn that differs from the last one a few tokens in would re-read everything without checkpoints. The server
keeps one at every turn boundary. Measured with a real claude -p session doing three tool turns: the first request
read its 21,964-token system prompt once (5.8 s), the next ones 209 and 106 new tokens (0.55 s and 0.4 s). An edit
in the middle of a history falls back to the checkpoint just before it. /v1/messages/count_tokens is served, and
a request that asks for no thinking (Claude Code's small helper calls) gets none.
This is not more memory, it is bigger pages. Windows maps memory in 4 KB pages, so the 63 GiB expert arena is 16.5 million of them; the CPU part of every token reads experts from it at DRAM speed, and each page needs a TLB entry. With 2 MB pages it is 32 thousand, and the CPU pool holds steady: 7.3-7.7 ms per round on this machine against 7.4-11 ms with 4 KB pages. Decode itself is GPU-bound here, so the rate moves little; the point is the steadiness (and a slightly faster start and exit).
Windows only gives large pages to an account that holds Lock pages in memory (SeLockMemoryPrivilege), and it puts a privilege into a sign-in's token only when that sign-in starts. So:
powershell -ExecutionPolicy Bypass -File tools\enable-large-pages.ps1 — asks for admin (UAC) and grants it to
the current user through secedit (works on Windows Home, which has no secpol.msc); the policy as it was is
saved to %LOCALAPPDATA%\strata-large-pages\before.inf, and -Revoke takes it back.-Check, or the engine's start line expert arena: ... large pages (2097152 B).Without the privilege the engine says large pages refused ... VirtualAlloc error 1314 and runs on 4 KB pages.
Error 1450 instead means the privilege is there but Windows found no 32 thousand free 2 MB blocks (memory
fragmented after a long uptime): it falls back the same way, and a reboot clears it. The arena is locked in RAM
either way - CUDA pins it for the GPU's copies.
STRATA_PREFILL_NVFP4=w4a8|w4a4|fp16 | prompt path precision (default w4a8) |
--embd-gguf PATH | the token embedding from this GGUF (BF16 from tools/embd_bf16_pack.py) |
STRATA_PREFILL_BF16X2=2|1|0 | exact inputs to the prompt path's BF16 projections: all but the hyper-connection (default), all (~9% slower prompt reading), off |
STRATA_KV_ROT=0 | int8 K/V without the Hadamard rotation (A/B) |
STRATA_ROPE_TABLE=0 | the fast-math RoPE angles instead of the float64 table (upstream's default; A/B) |
--prefill auto:8192|16384 | cap --prefill auto's chunk (default 32768 here, 8192 upstream) |
config "anthropic_thinking": "on_request"|"model" | an Anthropic request that does not ask for thinking renders without it (the bundle's default) / thinks as the template does (upstream's default) |
STRATA_COMMIT_SYNC=1 | the verify commit waits for its graph again (A/B) |
--pcie-frac F | share of cache misses fetched over PCIe (NVFP4 default 0.25) |
STRATA_NO_LARGEPAGES=1 | 4 KB pages even when large pages are allowed (A/B) |
STRATA_NO_NVFP4_512=1 | CPU pool on ggml-cpu's NVFP4 dot instead of the AVX-512 rows |
STRATA_UNBUFFERED_LOAD=1|0 | force the expert reads unbuffered / through the file cache (default: unbuffered only when the cache cannot keep the files) |
STRATA_DEFERRED_REGISTER=0, STRATA_ARENA_SYNC=1 | A/B: register the arena before the load / load it after the dense weights |
STRATA_ADAPT_WAIT=1 | each decode window waits for the adaptive tier's copies, so greedy decode repeats exactly (upstream's default since 0.1.38; ~11% slower with this fork's tier) |
STRATA_VERIFY_ARENA=1 | print a checksum of the loaded arena |
STRATA_DUMP_FIRST_LOGITS=file | write the first generated token's logits (compare prompt paths) |
STRATA_DUMP_MOE_INPUT=file, STRATA_DUMP_MOE_LAYER=l | dump one layer's real MoE input rows |
STRATA_REQUEST_LINES=1 | serve.server echoes one summary line per request to stdout |
STRATA_EMULATE_CC=75|86|89 | tests: answer as that generation (with an engine built as its PTX, 86-virtual) |
STRATA_QSA_WARP=1|select|attn | the pre-sm_80 QSA kernels on any card, as RTX 20 runs them (A/B) |
--vram-reserve-mib N | VRAM left unused (default 700, as upstream; this fork used 1500 on Windows until 0.1.38, where 700 measured 0.8% faster with no stalls); a smaller card's budget on a bigger one |
--low-ram / --no-low-ram | the low-RAM mode on / off (default: on with less than 96 GB installed) |
--ram-budget GIB | the low-RAM mode with at most GIB of pinned expert copies (upstream's --resident-budget-gib) |
STRATA_RESIDENT_HEADROOM_GIB=6 | RAM the low-RAM mode leaves free |
--no-kv-grow, STRATA_KV_GROW=0 | the K/V allocated for the whole context at start (before 0.1.31-nvfp4.2) |
STRATA_KV_GROW_INIT / _STEP | the K/V's cells at start (16384) and its growth step (8192) |
STRATA_PREFILL_GROUP_GATHER=0 | the prompt path gathers, waits and releases one expert at a time (A/B) |
STRATA_GR_V3=0 | upstream's default hyper-connection read (0.1.32: #315's staged variant, STRATA_HC_SPLIT) instead of the split V3 kernels (A/B: 18.49 vs 18.23 ms a round) |
--ple-inflight N | outstanding n-gram row reads (default 256; 64 before 0.1.31-nvfp4.2) |
--adapt-every N, --adapt-swaps N, --adapt-decay F | the adaptive VRAM tier: re-rank every N rounds, up to N swaps, counts x F after each (default 2 / 192 / 0.92 with a cache of 20-60% of the experts and every expert in RAM, else upstream's 4 / 96 / 0.7) |
| executable | checks |
|---|---|
nvfp4_avx512_parity [experts.bin] | AVX-512 rows vs ggml-cpu (--expert: one whole expert vs FP64; --bw: DRAM rate) |
nvfp4_expert_gpu_parity experts.bin | the GPU decode path's experts vs FP64, one per sampled layer |
mmq_nvfp4_parity experts.bin | MMQ products vs FP64 (--group, --layers, --real dump layer) |
python release\make_windows_bundle.py :: clean tree only; builds build-release\ itself, zips, SHA-256
git push origin HEAD:main
python release\publish.py --title "Strata NVFP4 v... - what changed" --notes notes.md :: --dry-run first
publish.py names this repository in every gh call (a clone's gh default can point at upstream) and refuses a
zip whose engine\BUILD.json is not the pushed HEAD, this version and a clean tree.
MIT, as upstream (LICENSE, copyright Niko1221 and the Strata contributors). llama.cpp / ggml code compiled into the build is MIT as well.
Strata fork: Qwen3.8-Flash-Next 125B MoE in ModelOpt NVFP4 on one RTX 20/30/40/50 card (12 GB+, measured on an RTX 5090) + 64 GB of RAM or more. W4A8 prompt path, AVX-512 CPU experts, a low-RAM tier, pictures on the CPU, Cyrillic MTP drafts. Windows / CUDA 13.
C++
18
717 commits
updated Oct 4, 2026
Qwen3.8-Flash-Next (125B hybrid MoE) in NVFP4 on one RTX 20, 30, 40 or 50 card (12 GB of VRAM or more; built and measured on an RTX 5090) + 64 GB of RAM or more, text and pictures. A fork of Niko1221/Strata that runs a ModelOpt NVFP4 checkpoint — here OrcaRouter's abliterated Flash-Next — instead of the Q2/Q3 quants Strata ships for. NVFP4 keeps the experts at 4.5 bits with calibrated scales, which is why this fork exists: the 2-3 bit quants were noticeably less accurate on the same model.
Strata itself keeps the routed experts in RAM (63 GiB here; with less than 96 GB of RAM this fork keeps only the ones outside VRAM and reads the rest from the SSD), caches the most-used ones in VRAM, computes the misses on the CPU and over PCIe in parallel with the GPU, and decodes with an MTP draft head. Everything about that design is upstream's; the original README is kept as README.upstream.md.
Runs an NVFP4 checkpoint (ModelOpt, 4.5-bit experts with calibrated scales) instead of Strata's Q2/Q3 quants: lossless repack, per-expert FP32 scales kept and applied to each projection's output.
Nothing is rounded coarser than the checkpoint where it can be avoided: the n-gram (PLE) table in its shipped FP8 (upstream reads it as IQ4_NL, 8% off per row), the token embedding in its shipped BF16, RoPE angles from a float64 table in every kernel (upstream's fast-math angles were ~0.02 rad off at 262K), FP32-exact inputs to the prompt path's BF16 projections (router, indexer, gates), int8 KV behind a Hadamard rotation.
An int8 (W4A8) prompt path for Blackwell: llama.cpp builds NVFP4 MMQ as FP4 x FP4 on sm_120; this keeps 8-bit activations - 16x smaller error per product - at 86% of its speed.
AVX-512 NVFP4 kernels for the CPU share of the experts (1.8-3.7x ggml-cpu).
Faster start: experts read unbuffered into the pinned arena on their own thread, registered with CUDA one layer ahead of the readers, beside everything else - the arena is in at ~7 s from a PCIe 5 drive.
The K/V grows with the context: at a 262K window upstream allocates the whole context's K/V at start (3.35 GiB
here with the draft layer's) - room for ~1,200 more experts in VRAM. Here the K/V takes VRAM only as the requests
reach further, from the expert cache, and gives it back after them (CUDA virtual memory: no address moves, no graph
is captured again). The window is never smaller; --no-kv-grow allocates it whole.
Faster prompt reading: an MMQ group's experts gathered in one launch after one wait (under WDDM every streamed expert's wait and event had left ~10 us of GPU idle, 226 ms of a 32K prompt), the n-gram rows read 256 at a time and beside layer 0, upstream's split hyper-connection kernels on.
An adaptive VRAM tier with a longer memory: it re-ranks the experts every 2 rounds, up to 192 swaps, and the
routing counts fade x0.92 per pass. Upstream re-ranks every 4 rounds, up to 96, x0.7. In fixed 1,000-token runs on
the 5090 it missed 30-40% fewer experts and the round did not change measurably. About half of those runs' tokens
came after the answer had ended, though, where the model loops over a few experts, so the gain on an answer alone
is not measured yet. It applies to a cache holding 20-60% of the experts with every expert in RAM. Elsewhere
upstream's settings stay; --adapt-every, --adapt-swaps and --adapt-decay override.
Tuned for NVFP4's larger experts (PCIe share, prompt chunks up to 32K, fused scale passes, a verify commit that overlaps the draft) and the fine-tune's own abliterated MTP draft head.
A draft vocabulary with Cyrillic: the MTP draft head proposes only tokens of its subset, and upstream's held
142 of the vocabulary's 18,580 Cyrillic tokens - an answer in Cyrillic decoded at 83 tokens/s with 1.4 tokens a round;
with the whole Cyrillic script (tools/draft_vocab.py --add cyrillic), 109 and 2.1. English is unchanged.
Upstream's CJK subset is one --add cjk away (data/draft_vocab_en.bin is the English/code one).
64 GB of RAM is enough: with less than 96 GB installed the engine runs upstream's file tier
(--mmap-experts --resident-budget-gib) with a budget of the free RAM less 6 GiB: the experts outside VRAM,
hottest first, pinned; the rest is read from experts.bin when needed - unbuffered, in merged requests (this
fork's addition, upstream since 0.1.38 as #362). The whole 63 GiB arena needs 96 GB. On an RTX 5090 + 64 GB a 32K
prompt waits 9.3 s instead of 5.7, and the answer after it runs at 97 tokens/s.
Every RTX 20, 30, 40 and 50 card with 12 GB or more: the NVFP4 path needed Blackwell only for the optional
FP4 x FP4 prompt path, which falls back; the release carries code for all four generations, each one's own path
tested on the 5090 (built as its PTX, the host answering as that card: STRATA_EMULATE_CC).
Images on the CPU, from the checkpoint's own vision tower: the encoder (strata-vision, FP32 weights) runs in
RAM, so the expert cache keeps all its VRAM and text decodes as fast as without images; a picture takes 2-6 s.
Upstream's GPU encoder took ~1.6 GB of VRAM (~600 cached experts, 10-20% of decode speed in served runs here) and, measured
against an FP32 reference, was up to 11% off (ggml-cuda's FP16 flash attention); this one is 0.1% off.
A picture an agent opens itself (Claude Code's Read) reaches the encoder too: it arrives inside a tool result,
which the server used to turn into text only (fixed here, sent upstream as #529).
Fixes: a scale fold that left NVFP4 hidden activations in FP16's subnormals (2-12% expert error), and the batched verify path skipping the query rotation of rotated KV caches.
A more accurate experts pack, re-quantized from the BF16 checkpoint (optional, tools/requant.py):
On upstream Strata 0.1.39. Since 0.1.38 upstream added:
"parallel": N, #465);--vram-elastic, #533);This fork's elastic K/V is off beside --vram-elastic and parallel: each of those manages the cache's VRAM its
own way. #554 (the vision markers in a message's text) now uses 0.1.39's own way of keeping a quoted <think> as
text, and #567's incremental tokenizing sits behind 0.1.39's prompt encoder. Upstream's byte-sized prompt ring
(#583) is kept: on this pack it reads at the same speed and borrows 1.2 GiB less. Against 0.1.38-nvfp4.2, with the
same expert cache, the logits and tokens are identical after 2K on the GPTQ + Q8_0-down pack, and decode
rounds are 6.9% shorter (15.7 against 16.9 ms, upstream's #646).
0.1.39-nvfp4.2: a long Claude Code conversation turned into "!" mid-reply, and every later request of it answered "!" until the engine was restarted by hand. The server now starts a fresh engine after such a reply (one bad reply instead of a broken session). The shared expert's fused SwiGLU quantizer keeps its fp16 scale finite, as 0.1.39's other two do (sent upstream as #838; not this incident's cause). What corrupted the state is not found yet: docs/NVFP4.md lists what was ruled out. Logits and tokens are identical to 0.1.39-nvfp4.1 on both packs.
0.1.38-nvfp4.2: the VRAM reserve is 700 MiB again on Windows (1500 no longer avoided a stall on 0.1.38, only cost expert slots). From other people's upstream pull requests: a prompt tokenised from the last shared prefix (#567: 6.8 ms instead of 128 ms at 125K tokens), token accounting across reasoning continuations (#615), and no lost 401/403 on Windows (#594). docs/NVFP4.md lists what else was measured and why it was not taken.
Fixes from a code review (0.1.37-nvfp4.2): the prompt path borrows 1.25 GiB less VRAM at a 32K chunk; the elastic K/V no longer runs past its mapped cells when it cannot lend slots; pictures Claude Code reads that the server cannot read become a note instead of a 400; safer image sources. Each upstream bug went up as a pull request (#546, #547, #550, #553, #554, #555; 0.1.38 fixed #546 its own way); docs/NVFP4.md has the list.
Each change was measured - first-token KL against a reference, and interleaved speed A/B runs; docs/NVFP4.md has the numbers, and everything that was tried and dropped.
GPU: an RTX 20, 30, 40 or 50 card with 12 GB of VRAM or more (the release has code for sm_75, 86, 89 and
120, and the optional FP4 x FP4 prompt unit for 120a). Built and measured on an RTX 5090, 32 GB: dense weights and
the 262K KV cache first, then ~19 GB of cached experts (~7,200). The other generations were tested on the 5090
through their own code paths (docs/NVFP4.md, "Other GPUs"). A card with less VRAM caches fewer experts and decodes
slower; lower --max-context with it (measured with 0.1.28-nvfp4.4):
| card's VRAM | --max-context | expert slots | decode, measured* |
|---|---|---|---|
| 32 GB (RTX 5090) | 262144 | 7,352 | 111 tok/s |
| 24 GB (RTX 3090 / 4090) | 131072 | 5,158 | 95 tok/s |
| 16 GB (RTX 4080 / 5080 / 4060 Ti 16 GB) | 65536 | 2,422 | 67 tok/s |
| 12 GB (RTX 3060 12 GB / 4070) | 32768 | 1,055 | 59 tok/s |
* On the RTX 5090 with the smaller card's VRAM budget (--vram-reserve-mib); a real card's own compute and PCIe
make it slower. 12 GB at 262144 and 8 GB cards at any context stop with no VRAM is left for the expert cache.
RAM: 64 GB minimum. With 96 GB or more the engine pins all 63 GiB of experts (~69 GiB of physical RAM in
all, measured with 128 GB). With less, the low-RAM mode starts by itself (--low-ram; --no-low-ram turns it
off; --ram-budget GIB caps it): the experts outside VRAM, hottest first, pinned up to the free RAM minus 6 GiB;
the rest is read from the file when needed. With 64 GB and an RTX 5090 all 44 GiB of them fit:
| 64 GB of RAM (measured*) with an RTX 5090 | |
|---|---|
| a 32K prompt: to the first token | 9.3 s (96+ GB: 5.7 s) |
| decode after it | 97 tok/s (0.1.28-nvfp4.4: 72) |
| start, to the session | ~16 s (the profile fill 4.3 s, the 44 GiB budget 9.5 s, read unbuffered) |
* On the RTX 5090 + 128 GB PC with 60 GiB of RAM locked away by a large-page ballast (58 GiB stay available, as on a 64 GB PC whose Windows uses 6 GB). The smaller cards' 64 GB figures (0.1.28-nvfp4.4: 24 GB 86-90 tok/s, 16 GB 54-56) were not measured again on this release.
Pagefile: Windows lets all processes together commit at most RAM + pagefile, and the engine commits ~98 GiB with all experts pinned, ~70-85 GiB in the low-RAM mode (the pinned RAM plus what WDDM reserves for the GPU's allocations; nothing of the model is ever paged out). With Windows and the usual apps on top: a pagefile of at least 32 GB with 128 GB of RAM, 64 GB with 96 GB, 48 GB with 64 GB. Set a fixed minimum rather than relying on a system-managed file to grow in time. Too small, and the start fails with an allocation error.
Disk: ~200 GB for the model files: GGUF 74 GB, expert pack 70 GB, n-gram table 51 GB, image encoder 1.8 GB, embedding 1.3 GB, MTP head 0.8 GB. ~340 GB while preparing them (the 135 GB checkpoint and the MTP intermediates can go afterwards). Use the fastest NVMe drive you have: every start reads 63 GiB.
RTX 5090 (32 GB, PCIe 5 x16), Ryzen 9 9950X3D, 128 GB DDR5-5600, Samsung 9100 PRO, Windows 11, CUDA 13.3. 262,144-token context, KV cache int8, large pages on. The machine is also a desktop: runs vary by ~5%.
| Writes a chat answer (~520 tokens, to its end) | 138 tokens/s on average, 125-146 in 6 runs; 18.2 ms a round |
| Writes a ~1,000-token answer after a 32K prompt | 132 tokens/s on average, 130-133 in 4 runs; 20.0 ms a round |
| The same with the re-quantized pack (GPTQ + Q8_0 down in 27 layers) | chat 126 tokens/s (-12%), after a 32K prompt 116 (-15%), the prompt read as fast |
| Reads a 32K prompt | 6,740-6,930 tokens/s (0.1.31-nvfp4.1: 5,400-5,900) |
| Start: the expert arena loaded | ~7 s (63 GiB of experts read at 10-11 GiB/s) |
0.1.31-nvfp4.3 measured the same in the same interleaved runs: 18.1 and 20.0 ms a round. Its notes said 161-173 tokens/s after a 32K prompt, but that prompt's answer ends after 22 tokens. Those runs went on to a fixed 256 tokens with no end-of-turn stop, so that was the speed of the loop after the answer, not of an answer. The rates vary with the draft acceptance of the path the tokens take (docs/NVFP4.md, "On upstream 0.1.32"). With 64 GB of RAM (the low-RAM mode): 9.3 s to the first token of a 32K prompt instead of 5.7, then 97 tokens/s.
Where precision was still being lost, first-token KL divergence from the more exact variant (8 prompts of 1K-8K tokens plus one of 32K; two runs of the same configuration differ by a median 0.000007):
| source | KL mean | at 32K | now |
|---|---|---|---|
| RoPE: fast-math float angles | - | 0.0063 | float64 table |
| token embedding stored as Q8_0 | 0.0025 | 0.0041 | BF16 as shipped (--embd-gguf) |
| prompt path: BF16-rounded activations into BF16 projections | 0.0023 | 0.0093 | hi + lo split (0.0029 left: not the hyper-connection) |
| n-gram table as IQ4_NL | 0.0026 | - | FP8 as shipped |
| int8 KV without rotation (vs fp16 KV) | 0.0017 | 0.0068 | rotated: 0.0011 / 0.0038 |
| prompt path W4A8 (vs FP16 activations) | 0.0018 | - | default; fp16 is 24% slower |
Needs Visual Studio 2022 Build Tools, CUDA 13+ (sm_120 wants 13), CMake, Ninja, Python 3.11+. From an x64 Native Tools prompt:
git clone https://github.com/ggml-org/llama.cpp third_party\llama.cpp
git -C third_party\llama.cpp checkout 3cf03257f219afbe7334045ff7c6a06ac68c627d
cmake -G Ninja -S . -B build -DCMAKE_BUILD_TYPE=Release -DSTRATA_ENABLE_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=120 ^
-DSTRATA_GGML_DIR=%CD%\third_party\llama.cpp
cmake --build build --target strata
Without -DSTRATA_GGML_DIR CMake fetches the same llama.cpp commit itself. 120 is enough: CMake builds the one
FP4 x FP4 unit for 120a by itself, and the rest runs on sm_121 too.
python -m venv .venv
.venv\Scripts\python -m pip install numpy torch safetensors transformers sentencepiece
set PYTHONPATH=third_party\llama.cpp\gguf-py
:: 1. the checkpoint (126 GiB)
hf download jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 --local-dir models\orca-nvfp4
:: 2. the n-gram (PLE) table, 51.2 GB: its FP8 bytes copied as they are (read from the SSD, never loaded into RAM)
.venv\Scripts\python tools\ple_fp8_pack.py --model models\orca-nvfp4 --out models\ple-fp8.gguf
:: 3. the token embedding, 1.3 GB: BF16 as shipped (the GGUF below stores it as Q8_0)
.venv\Scripts\python tools\embd_bf16_pack.py --model models\orca-nvfp4 --out models\token-embd-bf16.gguf
:: 4. GGUF (NVFP4 experts, Q8_0/BF16 dense) and the pack (experts.bin 63 GiB, tokenizer)
.venv\Scripts\python tools\nvfp4_convert.py --model models\orca-nvfp4 --outfile models\orca-nvfp4.gguf
.venv\Scripts\python tools\iq_pack.py --gguf models\orca-nvfp4.gguf --out packs\orca-nvfp4
:: 5. the fine-tune's own MTP draft head
.venv\Scripts\python tools\mtp_extract.py --model models\orca-nvfp4 --out mtp-orca
.venv\Scripts\python tools\mtp_pack.py --src mtp-orca --experts q2_0 --out mtp-orca\mtp-q2_0.gguf
.venv\Scripts\python tools\mtp_rt.py --gguf mtp-orca\mtp-q2_0.gguf --out mtp-orca\rt
copy data\draft_vocab.bin mtp-orca\rt\
Put packs\ and the GGUF on the fastest drive you have: the start is a 63 GiB read.
Optional: the more accurate experts pack. It needs the BF16 checkpoint (360 GB) and ~2 hours on an RTX 5090. The pack from step 4 stays the dense, index and tokenizer source.
hf download orcarouter/Qwen3.8-Flash-Next-Uncensored --local-dir models\orca-bf16
:: calibration: the engine's own MoE inputs of every layer for the four token files in data\requant_calib (uk, en, code, chat)
set STRATA_DUMP_MOE_INPUT=calib\d
set STRATA_DUMP_MOE_LAYER=all
build\strata.exe <the usual arguments> --tokens-file data\requant_calib\calib_uk.txt --max-new 1 (and en, code, chat)
:: errors per layer and method, then the plan for a size budget
.venv\Scripts\python tools\requant.py analyze --bf16 models\orca-bf16 --calib calib\d --out calib\errors.jsonl
.venv\Scripts\python tools\requant_plan.py calib\errors.jsonl calib 74
:: the pack: GPTQ NVFP4, Q8_0 down in the planned layers
.venv\Scripts\python tools\requant.py pack --bf16 models\orca-bf16 --calib calib\d --plan calib\plan_d8_74.json --base packs\orca-nvfp4 --out packs\orca-nvfp4-gptq-q8d
Then run with --pack packs\orca-nvfp4-gptq-q8d. The engine reads mixed expert formats from 0.1.32-nvfp4.2 on.
One-shot:
build\strata.exe --pack packs\orca-nvfp4 --native models\orca-nvfp4.gguf --native-dense-gguf models\orca-nvfp4.gguf ^
--ple-gguf models\ple-fp8.gguf --embd-gguf models\token-embd-bf16.gguf ^
--mtp mtp-orca\rt --spec 4 --spec-min-p 0.5 --prefill auto ^
--expert-profile data\expert-profile.bin --expert-cache auto ^
--max-context 262144 --kv int8 --tokens-file prompt.txt --max-new 256
OpenAI-compatible server: save the same arguments as a config ({"exe": "build/strata.exe", "args": [...], "tokenizer": "packs/orca-nvfp4/tokenizer", "host": "127.0.0.1", "port": 8097}) and start
python -m serve.server --engine strata --config that.json.
--native and --native-dense-gguf both point at the GGUF: --native alone would take the PLE file for a second
shard of the same model.
Images: the image encoder comes from the checkpoint (convert_hf_to_gguf.py <checkpoint> --mmproj --outtype f32,
1.8 GB) and is built with release\build-vision.cmd (CPU only). The engine takes --vision, the config a
"vision": {"exe": "build-vision-cpu/bin/strata-vision.exe", "mmproj": "models/mmproj-f32.gguf", "model": "models/orca-nvfp4.gguf", "gpu": false, "max_tokens": 1024} entry; then OpenAI image_url parts and Anthropic
image blocks (a screenshot pasted into Claude Code) work. The encoder uses one thread per core while it runs; the
engine is idle then.
set ANTHROPIC_BASE_URL=http://127.0.0.1:8097
set ANTHROPIC_API_KEY=local
set ANTHROPIC_MODEL=strata-nvfp4
set ANTHROPIC_SMALL_FAST_MODEL=strata-nvfp4
claude
The model is hybrid (Gated DeltaNet layers keep a recurrent state that cannot be cut back to a position), so an
agent turn that differs from the last one a few tokens in would re-read everything without checkpoints. The server
keeps one at every turn boundary. Measured with a real claude -p session doing three tool turns: the first request
read its 21,964-token system prompt once (5.8 s), the next ones 209 and 106 new tokens (0.55 s and 0.4 s). An edit
in the middle of a history falls back to the checkpoint just before it. /v1/messages/count_tokens is served, and
a request that asks for no thinking (Claude Code's small helper calls) gets none.
This is not more memory, it is bigger pages. Windows maps memory in 4 KB pages, so the 63 GiB expert arena is 16.5 million of them; the CPU part of every token reads experts from it at DRAM speed, and each page needs a TLB entry. With 2 MB pages it is 32 thousand, and the CPU pool holds steady: 7.3-7.7 ms per round on this machine against 7.4-11 ms with 4 KB pages. Decode itself is GPU-bound here, so the rate moves little; the point is the steadiness (and a slightly faster start and exit).
Windows only gives large pages to an account that holds Lock pages in memory (SeLockMemoryPrivilege), and it puts a privilege into a sign-in's token only when that sign-in starts. So:
powershell -ExecutionPolicy Bypass -File tools\enable-large-pages.ps1 — asks for admin (UAC) and grants it to
the current user through secedit (works on Windows Home, which has no secpol.msc); the policy as it was is
saved to %LOCALAPPDATA%\strata-large-pages\before.inf, and -Revoke takes it back.-Check, or the engine's start line expert arena: ... large pages (2097152 B).Without the privilege the engine says large pages refused ... VirtualAlloc error 1314 and runs on 4 KB pages.
Error 1450 instead means the privilege is there but Windows found no 32 thousand free 2 MB blocks (memory
fragmented after a long uptime): it falls back the same way, and a reboot clears it. The arena is locked in RAM
either way - CUDA pins it for the GPU's copies.
STRATA_PREFILL_NVFP4=w4a8|w4a4|fp16 | prompt path precision (default w4a8) |
--embd-gguf PATH | the token embedding from this GGUF (BF16 from tools/embd_bf16_pack.py) |
STRATA_PREFILL_BF16X2=2|1|0 | exact inputs to the prompt path's BF16 projections: all but the hyper-connection (default), all (~9% slower prompt reading), off |
STRATA_KV_ROT=0 | int8 K/V without the Hadamard rotation (A/B) |
STRATA_ROPE_TABLE=0 | the fast-math RoPE angles instead of the float64 table (upstream's default; A/B) |
--prefill auto:8192|16384 | cap --prefill auto's chunk (default 32768 here, 8192 upstream) |
config "anthropic_thinking": "on_request"|"model" | an Anthropic request that does not ask for thinking renders without it (the bundle's default) / thinks as the template does (upstream's default) |
STRATA_COMMIT_SYNC=1 | the verify commit waits for its graph again (A/B) |
--pcie-frac F | share of cache misses fetched over PCIe (NVFP4 default 0.25) |
STRATA_NO_LARGEPAGES=1 | 4 KB pages even when large pages are allowed (A/B) |
STRATA_NO_NVFP4_512=1 | CPU pool on ggml-cpu's NVFP4 dot instead of the AVX-512 rows |
STRATA_UNBUFFERED_LOAD=1|0 | force the expert reads unbuffered / through the file cache (default: unbuffered only when the cache cannot keep the files) |
STRATA_DEFERRED_REGISTER=0, STRATA_ARENA_SYNC=1 | A/B: register the arena before the load / load it after the dense weights |
STRATA_ADAPT_WAIT=1 | each decode window waits for the adaptive tier's copies, so greedy decode repeats exactly (upstream's default since 0.1.38; ~11% slower with this fork's tier) |
STRATA_VERIFY_ARENA=1 | print a checksum of the loaded arena |
STRATA_DUMP_FIRST_LOGITS=file | write the first generated token's logits (compare prompt paths) |
STRATA_DUMP_MOE_INPUT=file, STRATA_DUMP_MOE_LAYER=l | dump one layer's real MoE input rows |
STRATA_REQUEST_LINES=1 | serve.server echoes one summary line per request to stdout |
STRATA_EMULATE_CC=75|86|89 | tests: answer as that generation (with an engine built as its PTX, 86-virtual) |
STRATA_QSA_WARP=1|select|attn | the pre-sm_80 QSA kernels on any card, as RTX 20 runs them (A/B) |
--vram-reserve-mib N | VRAM left unused (default 700, as upstream; this fork used 1500 on Windows until 0.1.38, where 700 measured 0.8% faster with no stalls); a smaller card's budget on a bigger one |
--low-ram / --no-low-ram | the low-RAM mode on / off (default: on with less than 96 GB installed) |
--ram-budget GIB | the low-RAM mode with at most GIB of pinned expert copies (upstream's --resident-budget-gib) |
STRATA_RESIDENT_HEADROOM_GIB=6 | RAM the low-RAM mode leaves free |
--no-kv-grow, STRATA_KV_GROW=0 | the K/V allocated for the whole context at start (before 0.1.31-nvfp4.2) |
STRATA_KV_GROW_INIT / _STEP | the K/V's cells at start (16384) and its growth step (8192) |
STRATA_PREFILL_GROUP_GATHER=0 | the prompt path gathers, waits and releases one expert at a time (A/B) |
STRATA_GR_V3=0 | upstream's default hyper-connection read (0.1.32: #315's staged variant, STRATA_HC_SPLIT) instead of the split V3 kernels (A/B: 18.49 vs 18.23 ms a round) |
--ple-inflight N | outstanding n-gram row reads (default 256; 64 before 0.1.31-nvfp4.2) |
--adapt-every N, --adapt-swaps N, --adapt-decay F | the adaptive VRAM tier: re-rank every N rounds, up to N swaps, counts x F after each (default 2 / 192 / 0.92 with a cache of 20-60% of the experts and every expert in RAM, else upstream's 4 / 96 / 0.7) |
| executable | checks |
|---|---|
nvfp4_avx512_parity [experts.bin] | AVX-512 rows vs ggml-cpu (--expert: one whole expert vs FP64; --bw: DRAM rate) |
nvfp4_expert_gpu_parity experts.bin | the GPU decode path's experts vs FP64, one per sampled layer |
mmq_nvfp4_parity experts.bin | MMQ products vs FP64 (--group, --layers, --real dump layer) |
python release\make_windows_bundle.py :: clean tree only; builds build-release\ itself, zips, SHA-256
git push origin HEAD:main
python release\publish.py --title "Strata NVFP4 v... - what changed" --notes notes.md :: --dry-run first
publish.py names this repository in every gh call (a clone's gh default can point at upstream) and refuses a
zip whose engine\BUILD.json is not the pushed HEAD, this version and a clean tree.
MIT, as upstream (LICENSE, copyright Niko1221 and the Strata contributors). llama.cpp / ggml code compiled into the build is MIT as well.