sf-q3-8flash is a specialized fork of ds4 / DwarfStar
by Salvatore Sanfilippo and contributors, reduced to Qwen3.8 Flash Next on
Apple Metal. The upstream commit this fork sits on is not written here:
ask git, which cannot go stale --
git describe --tags --match 'sync-*' --abbrev=0 for the last sync, or
git merge-base HEAD upstream/main for the base itself.
Everything that works here works because of ds4, llama.cpp and GGML; see
LICENSE and the acknowledgements below.
ds4.c does not link against GGML, but it exists thanks to the path opened by the
llama.cpp project and the kernels, quantization formats, GGUF ecosystem, and hard-won
engineering knowledge developed there.
We are thankful and indebted to llama.cpp
and its contributors. Their implementation, kernels, tests, and design choices were
an essential reference while building this DeepSeek V4 specific inference path.
Some source-level pieces are retained or adapted here under the MIT license: GGUF
quant layouts and tables, CPU quant/dot logic, and certain kernels. For this
reason, and because we are genuinely grateful, we keep the GGML authors copyright
notice in our LICENSE file.
ds4 is built around a few models rather than as a general GGUF runner, and it is meant to be read and changed with a coding agent: a working template to adapt to your model and hardware, not a product that covers every setup. This fork pushes both ideas to the end: one model, one backend, and nothing else in the tree. Code for other models, other GPU backends and the bundled agent is deleted, not hidden behind flags. The result is a source tree small enough that a person, or an LLM, can load it whole and see how this one model actually runs, which makes it cheap to try an idea, measure it and keep or drop it.
Metal is the only GPU backend because the only hardware this fork is developed and tested on is an Apple M5 Max; kernels are tuned for that chip.
The smaller tree is also what made the rest of this work possible: open pull
requests on ds4 were analysed against this one model and ported where they
held up (docs/upstream-prs.md records the verdicts), and further
improvements were investigated for Apple Silicon. Every change is held to the
token-identity and speed checks below.
This repository intentionally supports one model and one production backend:
Qwen3.8 Flash Next (qwen4exp) on Apple Metal. The CPU implementation remains
for reference and tests. Qwen vision, built-in MTP, directional steering, the
HTTP server, disk KV cache, and the inherited TP/RDMA/pipeline plumbing are kept.
There is no bundled agent binary; connect an external client to the server.
The model combines gated delta-net and gated GQA layers, block-sparse attention, hyper-connections, MoE, built-in MTP, and a large BF16 n-gram table read directly from its GGUF. It is not a general GGUF runner and rejects other architectures.
Requirements: Apple Silicon, macOS, Xcode command-line tools, and enough memory for the selected quantization.
make -j8
./download.sh q2 # or q4
./sf-q3-8flash --ctx 8192 --prefill-chunk 1024
Q2 is about 137.10 GiB on disk with 41.73 GiB of resident weights. Q4 is about 165.11 GiB on disk with 69.74 GiB resident. Context and runtime buffers require additional RAM. Both files include the original BF16 n-grams and MTP weights.
./sf-q3-8flash --mtp -p "Explain mmap in C"
./sf-q3-8flash-server --ctx 8192
curl http://127.0.0.1:8004/v1/models
The server supports OpenAI chat/completions and Responses APIs, Anthropic Messages, request streaming, batching, tool calls, and optional disk KV cache:
./sf-q3-8flash-server --ctx 8192 \
--kv-disk-dir ~/.sf/q3-8flash/kv --kv-disk-space-mb 8192
See docs/SERVER.md and docs/CLIENTS.md.
./download.sh vision
./sf-q3-8flash --vision gguf/mmproj-Qwen3.8-Flash-Next-Q8_0.gguf
In the interactive CLI, use /read image.png. API clients can send image
content through the supported chat endpoints. See docs/QWEN38_FLASH_NEXT.md.
Directional steering is retained for all 48 trunk layers:
./sf-q3-8flash --dir-steering-file vectors.bin \
--dir-steering-ffn 1.0 --dir-steering-attn 1.0
The interactive /steer command can change scales. See
dir-steering/README.md for vector creation and format details.
Read AGENTS.md before changing code. The normal local loop is:
make -j8
make test -j8
make test-qwen4-kernels
python3 tests/test_model_download.py
Model-backed MTP, vision, evaluation, benchmark, and upstream parity checks are
documented in AGENTS.md and docs/TESTING.md. The default model file is
qwen3.8-flash-next.gguf; -m FILE overrides it.
Hardware. Every performance number in this repository was measured on one machine: an Apple M5 Max with 128 GB of unified memory. Other Macs will give different absolute numbers.
Token quality. Performance work here must not change what the model writes. Every change is checked in five ways:
tools/parity-check.sh) runs ten prompts
greedily on this child and on upstream ds4 at the child's merge-base, with
the same GGUF, and requires token-identical output;--bitwise) when a change claims it;tests/test_qwen4_mtp_limits.py checks MTP's draft depth and rollback
limits.tests/test_qwen4_mtp_identity.py checks that greedy output with MTP equals
plain greedy output, over 12 prompts at every draft depth.Performance. speed-bench/ab_bench.py alternates the two builds in ABBA
pairs for 600 s, drops pairs whose GPU clock sagged, and reports medians with
bootstrap 95% intervals for plain decode, prefill at three shapes, and MTP on
code and prose. A step is kept only when its target gains and no metric
clearly loses. Each change appends a row to speed-bench/perf-record.md
against a fixed start commit. See speed-bench/README.md.
Against ds4. Measured on 2026-09-27 using upstream's own methods. The
builds compared are ds4 at the merge-base 0aaea5a and this child with
81-q2-prefill-tails; the two MTP rows were measured again on 2026-09-29 with
82-mtp-greedy-divergence, ds4 and this child paired afresh.
ds4-bench on I Promessi Sposi: 2048-token intervals up
to 65536, and 128 generated tokens per context size.--ctx 8192 --temp 0 --nothink) and give the mean generation speed.Q2
| Measurement | ds4 t/s | sf t/s | sf vs ds4 |
|---|---|---|---|
| prefill, context 2048 | 1361.0 | 1427.8 | +4.9% |
| prefill, context 16384 | 1293.2 | 1372.2 | +6.1% |
| prefill, context 32768 | 1131.5 | 1264.8 | +11.8% |
| prefill, context 65536 | 916.4 | 999.6 | +9.1% |
| generation, context 2048 | 51.9 | 56.7 | +9.3% |
| generation, context 16384 | 51.7 | 56.3 | +9.1% |
| generation, context 32768 | 48.6 | 54.7 | +12.6% |
| generation, context 65536 | 43.0 | 47.3 | +10.1% |
| CLI generation, no MTP | 54.2 | 59.1 | +9.0% |
| CLI generation, MTP | 75.8 | 86.7 | +14.4% |
Q4
| Measurement | ds4 t/s | sf t/s | sf vs ds4 |
|---|---|---|---|
| prefill, context 2048 | 1365.9 | 1396.5 | +2.2% |
| prefill, context 16384 | 1261.4 | 1343.3 | +6.5% |
| prefill, context 32768 | 1107.9 | 1234.9 | +11.5% |
| prefill, context 65536 | 907.4 | 994.0 | +9.5% |
| generation, context 2048 | 54.1 | 54.4 | +0.5% |
| generation, context 16384 | 53.5 | 54.0 | +0.9% |
| generation, context 32768 | 48.5 | 52.4 | +8.1% |
| generation, context 65536 | 41.7 | 44.1 | +5.9% |
| CLI generation, no MTP | 56.5 | 57.5 | +1.9% |
| CLI generation, MTP | 77.8 | 85.9 | +10.4% |
Without MTP, the text of all six CLI cases is identical between ds4 and this
child. With MTP, this child's text equals its plain text in all six cases.
82-mtp-greedy-divergence made every MTP verify row compute as a single
decoded token would; before it, Q2 on the networking prompt diverged from plain
greedy output under MTP. The fix changes no plain row, which stays bitwise
identical.
MIT. LICENSE is inherited unchanged from ds4 and includes the GGML notice.
C
58.5%
Objective-C
22.5%
Metal
15.2%
Python
3.3%
sf-q3-8flash is a specialized fork of ds4 / DwarfStar
by Salvatore Sanfilippo and contributors, reduced to Qwen3.8 Flash Next on
Apple Metal. The upstream commit this fork sits on is not written here:
ask git, which cannot go stale --
git describe --tags --match 'sync-*' --abbrev=0 for the last sync, or
git merge-base HEAD upstream/main for the base itself.
Everything that works here works because of ds4, llama.cpp and GGML; see
LICENSE and the acknowledgements below.
ds4.c does not link against GGML, but it exists thanks to the path opened by the
llama.cpp project and the kernels, quantization formats, GGUF ecosystem, and hard-won
engineering knowledge developed there.
We are thankful and indebted to llama.cpp
and its contributors. Their implementation, kernels, tests, and design choices were
an essential reference while building this DeepSeek V4 specific inference path.
Some source-level pieces are retained or adapted here under the MIT license: GGUF
quant layouts and tables, CPU quant/dot logic, and certain kernels. For this
reason, and because we are genuinely grateful, we keep the GGML authors copyright
notice in our LICENSE file.
ds4 is built around a few models rather than as a general GGUF runner, and it is meant to be read and changed with a coding agent: a working template to adapt to your model and hardware, not a product that covers every setup. This fork pushes both ideas to the end: one model, one backend, and nothing else in the tree. Code for other models, other GPU backends and the bundled agent is deleted, not hidden behind flags. The result is a source tree small enough that a person, or an LLM, can load it whole and see how this one model actually runs, which makes it cheap to try an idea, measure it and keep or drop it.
Metal is the only GPU backend because the only hardware this fork is developed and tested on is an Apple M5 Max; kernels are tuned for that chip.
The smaller tree is also what made the rest of this work possible: open pull
requests on ds4 were analysed against this one model and ported where they
held up (docs/upstream-prs.md records the verdicts), and further
improvements were investigated for Apple Silicon. Every change is held to the
token-identity and speed checks below.
This repository intentionally supports one model and one production backend:
Qwen3.8 Flash Next (qwen4exp) on Apple Metal. The CPU implementation remains
for reference and tests. Qwen vision, built-in MTP, directional steering, the
HTTP server, disk KV cache, and the inherited TP/RDMA/pipeline plumbing are kept.
There is no bundled agent binary; connect an external client to the server.
The model combines gated delta-net and gated GQA layers, block-sparse attention, hyper-connections, MoE, built-in MTP, and a large BF16 n-gram table read directly from its GGUF. It is not a general GGUF runner and rejects other architectures.
Requirements: Apple Silicon, macOS, Xcode command-line tools, and enough memory for the selected quantization.
make -j8
./download.sh q2 # or q4
./sf-q3-8flash --ctx 8192 --prefill-chunk 1024
Q2 is about 137.10 GiB on disk with 41.73 GiB of resident weights. Q4 is about 165.11 GiB on disk with 69.74 GiB resident. Context and runtime buffers require additional RAM. Both files include the original BF16 n-grams and MTP weights.
./sf-q3-8flash --mtp -p "Explain mmap in C"
./sf-q3-8flash-server --ctx 8192
curl http://127.0.0.1:8004/v1/models
The server supports OpenAI chat/completions and Responses APIs, Anthropic Messages, request streaming, batching, tool calls, and optional disk KV cache:
./sf-q3-8flash-server --ctx 8192 \
--kv-disk-dir ~/.sf/q3-8flash/kv --kv-disk-space-mb 8192
See docs/SERVER.md and docs/CLIENTS.md.
./download.sh vision
./sf-q3-8flash --vision gguf/mmproj-Qwen3.8-Flash-Next-Q8_0.gguf
In the interactive CLI, use /read image.png. API clients can send image
content through the supported chat endpoints. See docs/QWEN38_FLASH_NEXT.md.
Directional steering is retained for all 48 trunk layers:
./sf-q3-8flash --dir-steering-file vectors.bin \
--dir-steering-ffn 1.0 --dir-steering-attn 1.0
The interactive /steer command can change scales. See
dir-steering/README.md for vector creation and format details.
Read AGENTS.md before changing code. The normal local loop is:
make -j8
make test -j8
make test-qwen4-kernels
python3 tests/test_model_download.py
Model-backed MTP, vision, evaluation, benchmark, and upstream parity checks are
documented in AGENTS.md and docs/TESTING.md. The default model file is
qwen3.8-flash-next.gguf; -m FILE overrides it.
Hardware. Every performance number in this repository was measured on one machine: an Apple M5 Max with 128 GB of unified memory. Other Macs will give different absolute numbers.
Token quality. Performance work here must not change what the model writes. Every change is checked in five ways:
tools/parity-check.sh) runs ten prompts
greedily on this child and on upstream ds4 at the child's merge-base, with
the same GGUF, and requires token-identical output;--bitwise) when a change claims it;tests/test_qwen4_mtp_limits.py checks MTP's draft depth and rollback
limits.tests/test_qwen4_mtp_identity.py checks that greedy output with MTP equals
plain greedy output, over 12 prompts at every draft depth.Performance. speed-bench/ab_bench.py alternates the two builds in ABBA
pairs for 600 s, drops pairs whose GPU clock sagged, and reports medians with
bootstrap 95% intervals for plain decode, prefill at three shapes, and MTP on
code and prose. A step is kept only when its target gains and no metric
clearly loses. Each change appends a row to speed-bench/perf-record.md
against a fixed start commit. See speed-bench/README.md.
Against ds4. Measured on 2026-09-27 using upstream's own methods. The
builds compared are ds4 at the merge-base 0aaea5a and this child with
81-q2-prefill-tails; the two MTP rows were measured again on 2026-09-29 with
82-mtp-greedy-divergence, ds4 and this child paired afresh.
ds4-bench on I Promessi Sposi: 2048-token intervals up
to 65536, and 128 generated tokens per context size.--ctx 8192 --temp 0 --nothink) and give the mean generation speed.Q2
| Measurement | ds4 t/s | sf t/s | sf vs ds4 |
|---|---|---|---|
| prefill, context 2048 | 1361.0 | 1427.8 | +4.9% |
| prefill, context 16384 | 1293.2 | 1372.2 | +6.1% |
| prefill, context 32768 | 1131.5 | 1264.8 | +11.8% |
| prefill, context 65536 | 916.4 | 999.6 | +9.1% |
| generation, context 2048 | 51.9 | 56.7 | +9.3% |
| generation, context 16384 | 51.7 | 56.3 | +9.1% |
| generation, context 32768 | 48.6 | 54.7 | +12.6% |
| generation, context 65536 | 43.0 | 47.3 | +10.1% |
| CLI generation, no MTP | 54.2 | 59.1 | +9.0% |
| CLI generation, MTP | 75.8 | 86.7 | +14.4% |
Q4
| Measurement | ds4 t/s | sf t/s | sf vs ds4 |
|---|---|---|---|
| prefill, context 2048 | 1365.9 | 1396.5 | +2.2% |
| prefill, context 16384 | 1261.4 | 1343.3 | +6.5% |
| prefill, context 32768 | 1107.9 | 1234.9 | +11.5% |
| prefill, context 65536 | 907.4 | 994.0 | +9.5% |
| generation, context 2048 | 54.1 | 54.4 | +0.5% |
| generation, context 16384 | 53.5 | 54.0 | +0.9% |
| generation, context 32768 | 48.5 | 52.4 | +8.1% |
| generation, context 65536 | 41.7 | 44.1 | +5.9% |
| CLI generation, no MTP | 56.5 | 57.5 | +1.9% |
| CLI generation, MTP | 77.8 | 85.9 | +10.4% |
Without MTP, the text of all six CLI cases is identical between ds4 and this
child. With MTP, this child's text equals its plain text in all six cases.
82-mtp-greedy-divergence made every MTP verify row compute as a single
decoded token would; before it, Q2 on the networking prompt diverged from plain
greedy output under MTP. The fix changes no plain row, which stays bitwise
identical.
MIT. LICENSE is inherited unchanged from ds4 and includes the GGML notice.
C
58.5%
Objective-C
22.5%
Metal
15.2%
Python
3.3%