Chida82/sf-q3-8flash

C

3

735 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

I took antirez's ds4, stripped it down to Qwen3.8 Flash Next on Metal, ported a bunch of improvements, and it's now ~10% faster with bit-exact output (r/LocalLLaMA)

I've had one pull request merged into ds4 (DwarfStar), a tiny one. There are a few more still waiting in the queue. I’m not complaining. Antirez says it clearly in the README: with coding agents everyone can tune the engine for their hardware and model and he can’t review everything. That made me…

0

Oct 4, 2026

README

sf-q3-8flash

sf-q3-8flash is a specialized fork of ds4 / DwarfStar by Salvatore Sanfilippo and contributors, reduced to Qwen3.8 Flash Next on Apple Metal. The upstream commit this fork sits on is not written here: ask git, which cannot go stale -- git describe --tags --match 'sync-*' --abbrev=0 for the last sync, or git merge-base HEAD upstream/main for the base itself. Everything that works here works because of ds4, llama.cpp and GGML; see LICENSE and the acknowledgements below.

Acknowledgements to llama.cpp and GGML

ds4.c does not link against GGML, but it exists thanks to the path opened by the llama.cpp project and the kernels, quantization formats, GGUF ecosystem, and hard-won engineering knowledge developed there. We are thankful and indebted to llama.cpp and its contributors. Their implementation, kernels, tests, and design choices were an essential reference while building this DeepSeek V4 specific inference path. Some source-level pieces are retained or adapted here under the MIT license: GGUF quant layouts and tables, CPU quant/dot logic, and certain kernels. For this reason, and because we are genuinely grateful, we keep the GGML authors copyright notice in our LICENSE file.

Why this fork exists

ds4 is built around a few models rather than as a general GGUF runner, and it is meant to be read and changed with a coding agent: a working template to adapt to your model and hardware, not a product that covers every setup. This fork pushes both ideas to the end: one model, one backend, and nothing else in the tree. Code for other models, other GPU backends and the bundled agent is deleted, not hidden behind flags. The result is a source tree small enough that a person, or an LLM, can load it whole and see how this one model actually runs, which makes it cheap to try an idea, measure it and keep or drop it.

Metal is the only GPU backend because the only hardware this fork is developed and tested on is an Apple M5 Max; kernels are tuned for that chip.

The smaller tree is also what made the rest of this work possible: open pull requests on ds4 were analysed against this one model and ported where they held up (docs/upstream-prs.md records the verdicts), and further improvements were investigated for Apple Silicon. Every change is held to the token-identity and speed checks below.

Scope

This repository intentionally supports one model and one production backend: Qwen3.8 Flash Next (qwen4exp) on Apple Metal. The CPU implementation remains for reference and tests. Qwen vision, built-in MTP, directional steering, the HTTP server, disk KV cache, and the inherited TP/RDMA/pipeline plumbing are kept. There is no bundled agent binary; connect an external client to the server.

The model combines gated delta-net and gated GQA layers, block-sparse attention, hyper-connections, MoE, built-in MTP, and a large BF16 n-gram table read directly from its GGUF. It is not a general GGUF runner and rejects other architectures.

Build

Requirements: Apple Silicon, macOS, Xcode command-line tools, and enough memory for the selected quantization.

make -j8
./download.sh q2                 # or q4
./sf-q3-8flash --ctx 8192 --prefill-chunk 1024

Q2 is about 137.10 GiB on disk with 41.73 GiB of resident weights. Q4 is about 165.11 GiB on disk with 69.74 GiB resident. Context and runtime buffers require additional RAM. Both files include the original BF16 n-grams and MTP weights.

./sf-q3-8flash --mtp -p "Explain mmap in C"
./sf-q3-8flash-server --ctx 8192
curl http://127.0.0.1:8004/v1/models

The server supports OpenAI chat/completions and Responses APIs, Anthropic Messages, request streaming, batching, tool calls, and optional disk KV cache:

./sf-q3-8flash-server --ctx 8192 \
  --kv-disk-dir ~/.sf/q3-8flash/kv --kv-disk-space-mb 8192

See docs/SERVER.md and docs/CLIENTS.md.

Vision

./download.sh vision
./sf-q3-8flash --vision gguf/mmproj-Qwen3.8-Flash-Next-Q8_0.gguf

In the interactive CLI, use /read image.png. API clients can send image content through the supported chat endpoints. See docs/QWEN38_FLASH_NEXT.md.

Directional steering

Directional steering is retained for all 48 trunk layers:

./sf-q3-8flash --dir-steering-file vectors.bin \
  --dir-steering-ffn 1.0 --dir-steering-attn 1.0

The interactive /steer command can change scales. See dir-steering/README.md for vector creation and format details.

Development and verification

Read AGENTS.md before changing code. The normal local loop is:

make -j8
make test -j8
make test-qwen4-kernels
python3 tests/test_model_download.py

Model-backed MTP, vision, evaluation, benchmark, and upstream parity checks are documented in AGENTS.md and docs/TESTING.md. The default model file is qwen3.8-flash-next.gguf; -m FILE overrides it.

Quality and performance

Hardware. Every performance number in this repository was measured on one machine: an Apple M5 Max with 128 GB of unified memory. Other Macs will give different absolute numbers.

Token quality. Performance work here must not change what the model writes. Every change is checked in five ways:

  • the StarForge parity oracle (tools/parity-check.sh) runs ten prompts greedily on this child and on upstream ds4 at the child's merge-base, with the same GGUF, and requires token-identical output;
  • the A/B harness requires identical tokens against the previous build, and bit-identical logits (--bitwise) when a change claims it;
  • kernel tests compare every optimized kernel with the CPU reference or with the kernel it replaces, byte for byte at their edge sizes;
  • tests/test_qwen4_mtp_limits.py checks MTP's draft depth and rollback limits.
  • tests/test_qwen4_mtp_identity.py checks that greedy output with MTP equals plain greedy output, over 12 prompts at every draft depth.

Performance. speed-bench/ab_bench.py alternates the two builds in ABBA pairs for 600 s, drops pairs whose GPU clock sagged, and reports medians with bootstrap 95% intervals for plain decode, prefill at three shapes, and MTP on code and prose. A step is kept only when its target gains and no metric clearly loses. Each change appends a row to speed-bench/perf-record.md against a fixed start commit. See speed-bench/README.md.

Against ds4. Measured on 2026-09-27 using upstream's own methods. The builds compared are ds4 at the merge-base 0aaea5a and this child with 81-q2-prefill-tails; the two MTP rows were measured again on 2026-09-29 with 82-mtp-greedy-divergence, ds4 and this child paired afresh.

  • The sweep rows use ds4-bench on I Promessi Sposi: 2048-token intervals up to 65536, and 128 generated tokens per context size.
  • Past the 1 GiB snapshot limit, both benches replay the prefix.
  • The CLI rows use upstream's three Qwen cases (--ctx 8192 --temp 0 --nothink) and give the mean generation speed.
  • Each value is the mean of two runs per build, in the order ds4, sf, sf, ds4.

Q2

Measurementds4 t/ssf t/ssf vs ds4
prefill, context 20481361.01427.8+4.9%
prefill, context 163841293.21372.2+6.1%
prefill, context 327681131.51264.8+11.8%
prefill, context 65536916.4999.6+9.1%
generation, context 204851.956.7+9.3%
generation, context 1638451.756.3+9.1%
generation, context 3276848.654.7+12.6%
generation, context 6553643.047.3+10.1%
CLI generation, no MTP54.259.1+9.0%
CLI generation, MTP75.886.7+14.4%

Q4

Measurementds4 t/ssf t/ssf vs ds4
prefill, context 20481365.91396.5+2.2%
prefill, context 163841261.41343.3+6.5%
prefill, context 327681107.91234.9+11.5%
prefill, context 65536907.4994.0+9.5%
generation, context 204854.154.4+0.5%
generation, context 1638453.554.0+0.9%
generation, context 3276848.552.4+8.1%
generation, context 6553641.744.1+5.9%
CLI generation, no MTP56.557.5+1.9%
CLI generation, MTP77.885.9+10.4%

Without MTP, the text of all six CLI cases is identical between ds4 and this child. With MTP, this child's text equals its plain text in all six cases. 82-mtp-greedy-divergence made every MTP verify row compute as a single decoded token would; before it, Q2 on the networking prompt diverged from plain greedy output under MTP. The fix changes no plain row, which stays bitwise identical.

Licence

MIT. LICENSE is inherited unchanged from ds4 and includes the GGML notice.

Chida82/sf-q3-8flash

C

3

735 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

I took antirez's ds4, stripped it down to Qwen3.8 Flash Next on Metal, ported a bunch of improvements, and it's now ~10% faster with bit-exact output (r/LocalLLaMA)

I've had one pull request merged into ds4 (DwarfStar), a tiny one. There are a few more still waiting in the queue. I’m not complaining. Antirez says it clearly in the README: with coding agents everyone can tune the engine for their hardware and model and he can’t review everything. That made me…

0

Oct 4, 2026

README

sf-q3-8flash

sf-q3-8flash is a specialized fork of ds4 / DwarfStar by Salvatore Sanfilippo and contributors, reduced to Qwen3.8 Flash Next on Apple Metal. The upstream commit this fork sits on is not written here: ask git, which cannot go stale -- git describe --tags --match 'sync-*' --abbrev=0 for the last sync, or git merge-base HEAD upstream/main for the base itself. Everything that works here works because of ds4, llama.cpp and GGML; see LICENSE and the acknowledgements below.

Acknowledgements to llama.cpp and GGML

ds4.c does not link against GGML, but it exists thanks to the path opened by the llama.cpp project and the kernels, quantization formats, GGUF ecosystem, and hard-won engineering knowledge developed there. We are thankful and indebted to llama.cpp and its contributors. Their implementation, kernels, tests, and design choices were an essential reference while building this DeepSeek V4 specific inference path. Some source-level pieces are retained or adapted here under the MIT license: GGUF quant layouts and tables, CPU quant/dot logic, and certain kernels. For this reason, and because we are genuinely grateful, we keep the GGML authors copyright notice in our LICENSE file.

Why this fork exists

ds4 is built around a few models rather than as a general GGUF runner, and it is meant to be read and changed with a coding agent: a working template to adapt to your model and hardware, not a product that covers every setup. This fork pushes both ideas to the end: one model, one backend, and nothing else in the tree. Code for other models, other GPU backends and the bundled agent is deleted, not hidden behind flags. The result is a source tree small enough that a person, or an LLM, can load it whole and see how this one model actually runs, which makes it cheap to try an idea, measure it and keep or drop it.

Metal is the only GPU backend because the only hardware this fork is developed and tested on is an Apple M5 Max; kernels are tuned for that chip.

The smaller tree is also what made the rest of this work possible: open pull requests on ds4 were analysed against this one model and ported where they held up (docs/upstream-prs.md records the verdicts), and further improvements were investigated for Apple Silicon. Every change is held to the token-identity and speed checks below.

Scope

This repository intentionally supports one model and one production backend: Qwen3.8 Flash Next (qwen4exp) on Apple Metal. The CPU implementation remains for reference and tests. Qwen vision, built-in MTP, directional steering, the HTTP server, disk KV cache, and the inherited TP/RDMA/pipeline plumbing are kept. There is no bundled agent binary; connect an external client to the server.

The model combines gated delta-net and gated GQA layers, block-sparse attention, hyper-connections, MoE, built-in MTP, and a large BF16 n-gram table read directly from its GGUF. It is not a general GGUF runner and rejects other architectures.

Build

Requirements: Apple Silicon, macOS, Xcode command-line tools, and enough memory for the selected quantization.

make -j8
./download.sh q2                 # or q4
./sf-q3-8flash --ctx 8192 --prefill-chunk 1024

Q2 is about 137.10 GiB on disk with 41.73 GiB of resident weights. Q4 is about 165.11 GiB on disk with 69.74 GiB resident. Context and runtime buffers require additional RAM. Both files include the original BF16 n-grams and MTP weights.

./sf-q3-8flash --mtp -p "Explain mmap in C"
./sf-q3-8flash-server --ctx 8192
curl http://127.0.0.1:8004/v1/models

The server supports OpenAI chat/completions and Responses APIs, Anthropic Messages, request streaming, batching, tool calls, and optional disk KV cache:

./sf-q3-8flash-server --ctx 8192 \
  --kv-disk-dir ~/.sf/q3-8flash/kv --kv-disk-space-mb 8192

See docs/SERVER.md and docs/CLIENTS.md.

Vision

./download.sh vision
./sf-q3-8flash --vision gguf/mmproj-Qwen3.8-Flash-Next-Q8_0.gguf

In the interactive CLI, use /read image.png. API clients can send image content through the supported chat endpoints. See docs/QWEN38_FLASH_NEXT.md.

Directional steering

Directional steering is retained for all 48 trunk layers:

./sf-q3-8flash --dir-steering-file vectors.bin \
  --dir-steering-ffn 1.0 --dir-steering-attn 1.0

The interactive /steer command can change scales. See dir-steering/README.md for vector creation and format details.

Development and verification

Read AGENTS.md before changing code. The normal local loop is:

make -j8
make test -j8
make test-qwen4-kernels
python3 tests/test_model_download.py

Model-backed MTP, vision, evaluation, benchmark, and upstream parity checks are documented in AGENTS.md and docs/TESTING.md. The default model file is qwen3.8-flash-next.gguf; -m FILE overrides it.

Quality and performance

Hardware. Every performance number in this repository was measured on one machine: an Apple M5 Max with 128 GB of unified memory. Other Macs will give different absolute numbers.

Token quality. Performance work here must not change what the model writes. Every change is checked in five ways:

  • the StarForge parity oracle (tools/parity-check.sh) runs ten prompts greedily on this child and on upstream ds4 at the child's merge-base, with the same GGUF, and requires token-identical output;
  • the A/B harness requires identical tokens against the previous build, and bit-identical logits (--bitwise) when a change claims it;
  • kernel tests compare every optimized kernel with the CPU reference or with the kernel it replaces, byte for byte at their edge sizes;
  • tests/test_qwen4_mtp_limits.py checks MTP's draft depth and rollback limits.
  • tests/test_qwen4_mtp_identity.py checks that greedy output with MTP equals plain greedy output, over 12 prompts at every draft depth.

Performance. speed-bench/ab_bench.py alternates the two builds in ABBA pairs for 600 s, drops pairs whose GPU clock sagged, and reports medians with bootstrap 95% intervals for plain decode, prefill at three shapes, and MTP on code and prose. A step is kept only when its target gains and no metric clearly loses. Each change appends a row to speed-bench/perf-record.md against a fixed start commit. See speed-bench/README.md.

Against ds4. Measured on 2026-09-27 using upstream's own methods. The builds compared are ds4 at the merge-base 0aaea5a and this child with 81-q2-prefill-tails; the two MTP rows were measured again on 2026-09-29 with 82-mtp-greedy-divergence, ds4 and this child paired afresh.

  • The sweep rows use ds4-bench on I Promessi Sposi: 2048-token intervals up to 65536, and 128 generated tokens per context size.
  • Past the 1 GiB snapshot limit, both benches replay the prefix.
  • The CLI rows use upstream's three Qwen cases (--ctx 8192 --temp 0 --nothink) and give the mean generation speed.
  • Each value is the mean of two runs per build, in the order ds4, sf, sf, ds4.

Q2

Measurementds4 t/ssf t/ssf vs ds4
prefill, context 20481361.01427.8+4.9%
prefill, context 163841293.21372.2+6.1%
prefill, context 327681131.51264.8+11.8%
prefill, context 65536916.4999.6+9.1%
generation, context 204851.956.7+9.3%
generation, context 1638451.756.3+9.1%
generation, context 3276848.654.7+12.6%
generation, context 6553643.047.3+10.1%
CLI generation, no MTP54.259.1+9.0%
CLI generation, MTP75.886.7+14.4%

Q4

Measurementds4 t/ssf t/ssf vs ds4
prefill, context 20481365.91396.5+2.2%
prefill, context 163841261.41343.3+6.5%
prefill, context 327681107.91234.9+11.5%
prefill, context 65536907.4994.0+9.5%
generation, context 204854.154.4+0.5%
generation, context 1638453.554.0+0.9%
generation, context 3276848.552.4+8.1%
generation, context 6553641.744.1+5.9%
CLI generation, no MTP56.557.5+1.9%
CLI generation, MTP77.885.9+10.4%

Without MTP, the text of all six CLI cases is identical between ds4 and this child. With MTP, this child's text equals its plain text in all six cases. 82-mtp-greedy-divergence made every MTP verify row compute as a single decoded token would; before it, Q2 on the networking prompt diverged from plain greedy output under MTP. The fix changes no plain row, which stays bitwise identical.

Licence

MIT. LICENSE is inherited unchanged from ds4 and includes the GGML notice.

Languages

C

58.5%

Objective-C

22.5%

Metal

15.2%

Python

3.3%