lkarlslund/ninfer6000

High-performance single-GPU inference for selected model checkpoints and GPUs.

C++

0

1,161 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s (r/LocalLLaMA)

I've tinkered some more with my fork of NInfer for the Qwen 3.8 Flash Next model on RTX6000, and I've just hit 400tg/s under ideal conditions with it, so I thought it was worth a share. # Decode |Mode|Context|16-bit|8-bit|Change| |:-|:-|:-|:-|:-| |No speculative decoding|512|118.0 tok/s|172.0…

3

Oct 6, 2026

README

NInfer 6000

NInfer 6000 runs Qwen3.8 Flash-Next 125B-A6B on one NVIDIA RTX PRO 6000 Blackwell GPU.

It is a fork of NInfer, a C++/CUDA inference engine for single Blackwell GPUs. This fork adds the following Flash-Next features:

  • an 8-bit weight option;
  • an 8-bit PLE table option;
  • kernel optimizations for decode and MTP speculative decoding.

For build requirements, CLI and HTTP usage, Docker and the general architecture, refer to the NInfer README and the documentation index. These instructions also apply to this fork.

Performance

These measurements use one request, BF16 KV, CUDA Graphs, greedy sampling and an 8,192-token prefill chunk. The model is Swift 1.5. The GPU power limit is 600 W. Both variants use the same build.

Decode

ModeContext16-bit8-bitChange
No speculative decoding512118.0 tok/s172.0 tok/s+46%
No speculative decoding8K117.8 tok/s170.6 tok/s+45%
MTP3512171.1 tok/s256.2 tok/s+50%
MTP38K268.2 tok/s381.5 tok/s+42%
MTP3 with --lm-head-draft512197.1 tok/s274.8 tok/s+39%
MTP3 with --lm-head-draft8K303.1 tok/s401.3 tok/s+32%

Prefill

Prompt length16-bit8-bitChange
512 tokens6,709 tok/s5,905 tok/s-12%
8,192 tokens13,903 tok/s13,908 tok/s0%

The 16-bit variant has BF16 non-expert weights and a BF16 PLE table. The 8-bit variant has FP8 non-expert weights and an FP8 PLE table. Both variants use NVFP4 routed experts.

--lm-head-draft makes the MTP drafter use a smaller, quantized proposal head. It increases MTP3 decode by 5% to 15%.

MTP3 throughput depends on the text. The 8K benchmark text accepts 3.9 tokens per verification round. The 512-token text accepts 2.5 to 2.6.

Power limit

The 8-bit variant at 450 W and at 600 W:

Measurement450 W600 W
Decode without speculative decoding, 8K171.6 tok/s170.6 tok/s
Decode with MTP3, 8K390.8 tok/s401.3 tok/s
Prefill, 8,192 tokens12,104 tok/s13,908 tok/s

Decode is limited by memory bandwidth, so the power limit has almost no effect. Prefill is faster at 600 W.

Benchmark command

./build/bench/ninfer_bench \
  --weights out/v3/swift_1_5_qwen3_8_flash_next_nvfp4_fp8ple_fp8proj.ninfer \
  --corpus bench/fixtures/qwen3_8_flash_next_context.ids \
  -pg "512,256;8192,256" --max-ctx 9216 --prefill-chunk 8192 --kv-dtype bf16 \
  --spec mtp --draft-tokens 3 --lm-head-draft --warmup 1 -r 2

The benchmark binary is in the dev preset. Remove the --spec, --draft-tokens and --lm-head-draft options to measure decode without speculative decoding.

Target system

ItemValue
GPURTX PRO 6000 Blackwell Workstation Edition, 96 GB
GPU power limit600 W
Host memory64 to 128 GB
CUDA13.x, sm_120a
Operating system64-bit Linux

The engine reads the PLE table directly from the artifact file. Use fast local storage. Keep the table in the page cache for the best prefill and decode speed. An 8-bit PLE table is 51 GB. A 128 GB host keeps all of it in the page cache. A 64 GB host keeps part of it, and the other rows are read from storage when they are used.

Model variants

The converter accepts two source checkpoints. Both store the routed experts as NVFP4.

Source profileHugging Face checkpointSource PLE table
radixark (default)RadixArk/Qwen3.8-Flash-Next-NVFP4FP8, 51 GB
swiftukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4BF16, 102 GB

Conversion options

The converter has two optional 8-bit formats:

OptionEffect
--ple-format fp8_e4m3fnRe-encodes a BF16 PLE table as FP8 E4M3 with one BF16 scale. Applies to swift only.
--projection-format fp8_e4m3fn_row_bf16Stores 510 non-expert matrices as FP8 E4M3 with one BF16 scale per row. Activations stay BF16.

The FP8 projection option converts these matrices:

  • attention q, k, v and o projections;
  • GDN in_proj_qkv, in_proj_z and out_proj projections;
  • the MoE router and the shared-expert gate and up projections;
  • the HyperConnection Down and Up projections;
  • the output head;
  • the MTP drafter projections.

The routed experts stay NVFP4. The indexer, the GDN control weights, the shared-expert down projection and the MTP experts stay BF16.

The output file name is fixed for each set of options. The converter rejects other names.

ProfileOptionsOutput fileSize
radixarknoneqwen3_8_flash_next_125b_a6b_nvfp4.ninfer134.8 GB
swiftnoneswift_1_5_qwen3_8_flash_next_nvfp4.ninfer186.0 GB
swiftboth 8-bit optionsswift_1_5_qwen3_8_flash_next_nvfp4_fp8ple_fp8proj.ninfer130.5 GB

The radixark profile with --projection-format adds _fp8proj to the name. The swift profile with only --ple-format adds _fp8ple. We measured only the Swift 8-bit combination. The other combinations are not measured.

Convert Swift 1.5 with both 8-bit options:

python3 -m tools.convert.qwen3_8_flash_next_125b_a6b.convert \
  --model /path/to/Swift-1.5-Qwen3.8-Flash-Next-NVFP4 \
  --source-profile swift \
  --ple-format fp8_e4m3fn \
  --projection-format fp8_e4m3fn_row_bf16 \
  --out out/v3/swift_1_5_qwen3_8_flash_next_nvfp4_fp8ple_fp8proj.ninfer \
  --device cuda

The conversion takes approximately 7 minutes. The writer does not overwrite an existing artifact. Keep the entry file and all .part-NNNN files together. The artifact reference gives the complete inventory and binding rules.

Quality of the 8-bit variant

We compared the 8-bit Swift artifact with the BF16 Swift artifact on six oracle prompts:

  • The mean log-probability change of the generated tokens is 0.003 to 0.086.
  • The perplexity changes by -2.4% to +1.7%.
  • Greedy generation changes only at near-tie tokens.

Build

git clone https://github.com/lkarlslund/ninfer6000.git
cd ninfer6000
cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

Serve

This command starts an OpenAI- and Anthropic-compatible server with MTP3 and Vision:

./build/apps/ninfer-serve out/v3/swift_1_5_qwen3_8_flash_next_nvfp4_fp8ple_fp8proj.ninfer \
  --host 0.0.0.0 --port 8003 \
  --model-id Qwen/Qwen3.8-Flash-Next \
  --max-context 196608 --kv-capacity auto --kv-dtype bf16 \
  --max-concurrency 2 \
  --device-state-slots 2 --host-state-slots 8 --host-kv-mib 16384 \
  --spec mtp --draft-tokens 3 --lm-head-draft \
  --vision --preserve-thinking \
  --prefill-chunk 8192

--kv-capacity auto gives all free GPU memory to the KV cache. The cache size is limited to --max-context multiplied by --max-concurrency. Thus, smaller weights give more KV capacity but do not decrease the total GPU memory use.

License

NInfer is licensed under the Apache License 2.0. The model weights have their own licenses. Refer to the source checkpoints on Hugging Face. Vendored dependencies keep their license files under third_party/.

lkarlslund/ninfer6000

High-performance single-GPU inference for selected model checkpoints and GPUs.

C++

0

1,161 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s (r/LocalLLaMA)

I've tinkered some more with my fork of NInfer for the Qwen 3.8 Flash Next model on RTX6000, and I've just hit 400tg/s under ideal conditions with it, so I thought it was worth a share. # Decode |Mode|Context|16-bit|8-bit|Change| |:-|:-|:-|:-|:-| |No speculative decoding|512|118.0 tok/s|172.0…

3

Oct 6, 2026

README

NInfer 6000

NInfer 6000 runs Qwen3.8 Flash-Next 125B-A6B on one NVIDIA RTX PRO 6000 Blackwell GPU.

It is a fork of NInfer, a C++/CUDA inference engine for single Blackwell GPUs. This fork adds the following Flash-Next features:

  • an 8-bit weight option;
  • an 8-bit PLE table option;
  • kernel optimizations for decode and MTP speculative decoding.

For build requirements, CLI and HTTP usage, Docker and the general architecture, refer to the NInfer README and the documentation index. These instructions also apply to this fork.

Performance

These measurements use one request, BF16 KV, CUDA Graphs, greedy sampling and an 8,192-token prefill chunk. The model is Swift 1.5. The GPU power limit is 600 W. Both variants use the same build.

Decode

ModeContext16-bit8-bitChange
No speculative decoding512118.0 tok/s172.0 tok/s+46%
No speculative decoding8K117.8 tok/s170.6 tok/s+45%
MTP3512171.1 tok/s256.2 tok/s+50%
MTP38K268.2 tok/s381.5 tok/s+42%
MTP3 with --lm-head-draft512197.1 tok/s274.8 tok/s+39%
MTP3 with --lm-head-draft8K303.1 tok/s401.3 tok/s+32%

Prefill

Prompt length16-bit8-bitChange
512 tokens6,709 tok/s5,905 tok/s-12%
8,192 tokens13,903 tok/s13,908 tok/s0%

The 16-bit variant has BF16 non-expert weights and a BF16 PLE table. The 8-bit variant has FP8 non-expert weights and an FP8 PLE table. Both variants use NVFP4 routed experts.

--lm-head-draft makes the MTP drafter use a smaller, quantized proposal head. It increases MTP3 decode by 5% to 15%.

MTP3 throughput depends on the text. The 8K benchmark text accepts 3.9 tokens per verification round. The 512-token text accepts 2.5 to 2.6.

Power limit

The 8-bit variant at 450 W and at 600 W:

Measurement450 W600 W
Decode without speculative decoding, 8K171.6 tok/s170.6 tok/s
Decode with MTP3, 8K390.8 tok/s401.3 tok/s
Prefill, 8,192 tokens12,104 tok/s13,908 tok/s

Decode is limited by memory bandwidth, so the power limit has almost no effect. Prefill is faster at 600 W.

Benchmark command

./build/bench/ninfer_bench \
  --weights out/v3/swift_1_5_qwen3_8_flash_next_nvfp4_fp8ple_fp8proj.ninfer \
  --corpus bench/fixtures/qwen3_8_flash_next_context.ids \
  -pg "512,256;8192,256" --max-ctx 9216 --prefill-chunk 8192 --kv-dtype bf16 \
  --spec mtp --draft-tokens 3 --lm-head-draft --warmup 1 -r 2

The benchmark binary is in the dev preset. Remove the --spec, --draft-tokens and --lm-head-draft options to measure decode without speculative decoding.

Target system

ItemValue
GPURTX PRO 6000 Blackwell Workstation Edition, 96 GB
GPU power limit600 W
Host memory64 to 128 GB
CUDA13.x, sm_120a
Operating system64-bit Linux

The engine reads the PLE table directly from the artifact file. Use fast local storage. Keep the table in the page cache for the best prefill and decode speed. An 8-bit PLE table is 51 GB. A 128 GB host keeps all of it in the page cache. A 64 GB host keeps part of it, and the other rows are read from storage when they are used.

Model variants

The converter accepts two source checkpoints. Both store the routed experts as NVFP4.

Source profileHugging Face checkpointSource PLE table
radixark (default)RadixArk/Qwen3.8-Flash-Next-NVFP4FP8, 51 GB
swiftukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4BF16, 102 GB

Conversion options

The converter has two optional 8-bit formats:

OptionEffect
--ple-format fp8_e4m3fnRe-encodes a BF16 PLE table as FP8 E4M3 with one BF16 scale. Applies to swift only.
--projection-format fp8_e4m3fn_row_bf16Stores 510 non-expert matrices as FP8 E4M3 with one BF16 scale per row. Activations stay BF16.

The FP8 projection option converts these matrices:

  • attention q, k, v and o projections;
  • GDN in_proj_qkv, in_proj_z and out_proj projections;
  • the MoE router and the shared-expert gate and up projections;
  • the HyperConnection Down and Up projections;
  • the output head;
  • the MTP drafter projections.

The routed experts stay NVFP4. The indexer, the GDN control weights, the shared-expert down projection and the MTP experts stay BF16.

The output file name is fixed for each set of options. The converter rejects other names.

ProfileOptionsOutput fileSize
radixarknoneqwen3_8_flash_next_125b_a6b_nvfp4.ninfer134.8 GB
swiftnoneswift_1_5_qwen3_8_flash_next_nvfp4.ninfer186.0 GB
swiftboth 8-bit optionsswift_1_5_qwen3_8_flash_next_nvfp4_fp8ple_fp8proj.ninfer130.5 GB

The radixark profile with --projection-format adds _fp8proj to the name. The swift profile with only --ple-format adds _fp8ple. We measured only the Swift 8-bit combination. The other combinations are not measured.

Convert Swift 1.5 with both 8-bit options:

python3 -m tools.convert.qwen3_8_flash_next_125b_a6b.convert \
  --model /path/to/Swift-1.5-Qwen3.8-Flash-Next-NVFP4 \
  --source-profile swift \
  --ple-format fp8_e4m3fn \
  --projection-format fp8_e4m3fn_row_bf16 \
  --out out/v3/swift_1_5_qwen3_8_flash_next_nvfp4_fp8ple_fp8proj.ninfer \
  --device cuda

The conversion takes approximately 7 minutes. The writer does not overwrite an existing artifact. Keep the entry file and all .part-NNNN files together. The artifact reference gives the complete inventory and binding rules.

Quality of the 8-bit variant

We compared the 8-bit Swift artifact with the BF16 Swift artifact on six oracle prompts:

  • The mean log-probability change of the generated tokens is 0.003 to 0.086.
  • The perplexity changes by -2.4% to +1.7%.
  • Greedy generation changes only at near-tie tokens.

Build

git clone https://github.com/lkarlslund/ninfer6000.git
cd ninfer6000
cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

Serve

This command starts an OpenAI- and Anthropic-compatible server with MTP3 and Vision:

./build/apps/ninfer-serve out/v3/swift_1_5_qwen3_8_flash_next_nvfp4_fp8ple_fp8proj.ninfer \
  --host 0.0.0.0 --port 8003 \
  --model-id Qwen/Qwen3.8-Flash-Next \
  --max-context 196608 --kv-capacity auto --kv-dtype bf16 \
  --max-concurrency 2 \
  --device-state-slots 2 --host-state-slots 8 --host-kv-mib 16384 \
  --spec mtp --draft-tokens 3 --lm-head-draft \
  --vision --preserve-thinking \
  --prefill-chunk 8192

--kv-capacity auto gives all free GPU memory to the KV cache. The cache size is limited to --max-context multiplied by --max-concurrency. Thus, smaller weights give more KV capacity but do not decrease the total GPU memory use.

License

NInfer is licensed under the Apache License 2.0. The model weights have their own licenses. Refer to the source checkpoints on Hugging Face. Vendored dependencies keep their license files under third_party/.