Run Qwen3.8-Flash-Next, a 111 GB model, on a 64 GB Mac.
This is a llama.cpp fork for Apple Silicon. The model is a mixture of experts: 48 layers of 512 experts, and each token uses only 10 experts on each layer. So the experts stay on the SSD. The ones a token needs are read in as it is computed, and the ones used most stay in an expert cache in RAM. Everything else stays resident on the GPU.
It was tuned for a Mac mini M5 Pro with 64 GB and a second SSD over Thunderbolt 5. The tuning was measured on real agent conversations, and any timing that swap touched was thrown out.
Measured on a Mac mini M5 Pro (64 GB) with Unsloth's UD-Q4_K_XL quant, a 28 GiB expert cache, two SSDs and the MTP draft head on. Sources and method are in Results.
| Writing (decode) | 17.5 tokens/s across 120 real agent conversations replayed in order: +47% over the study's first setup (11.9) |
| Time to first token | 5.5 s median: −52% (from 11.4 s). A typical first reply reads ~1,500 new tokens; the other ~12,800 come from the prompt cache |
| Reading a prompt (prefill) | 523 / 430 / 391 tokens/s at 4K / 32K / 100K tokens, from cold |
| Quality | 18/20 on a fixed 20-task exam, the same score as the plain setup |
--moe-stream keeps the routed experts on disk and holds a fixed-size expert cache in RAM
(--moe-stream-cache 28 = 28 GiB). Reads skip macOS's file cache (--moe-stream-direct), so the
cache and the rest of the system keep their memory, and many reads are in flight at once. While a
token is being written, the next layer's experts are fetched ahead of time; about 72% of those
guesses turn out right. The cache keeps the experts used most recently and most often. At 28 GiB
about 3 in 4 expert lookups find their expert already in RAM.
A long prompt needs nearly every expert on every layer. The reading room is a small ring buffer (about 1.4 GiB) carved out of the expert cache. The SSDs fill it layer after layer, ahead of the GPU, so prompt reading costs whichever of the two is slower, not both added together. The output is bit-identical to having every expert in RAM. While writing, the buffer is lent back as extra cache, which won back its cost (+5.5% writing). It is on by default.
Put a byte-identical copy of the model on a second SSD and the expert reads are split across both
by expert id (--moe-stream-alt-path, --moe-stream-alt-split 53). The internal SSD (6.5 GB/s)
and a 990 PRO over Thunderbolt 5 (5.7 GB/s) give 12.3 GB/s together. That made writing 15% faster
and cut the time to first token by 31%.
A small draft head guesses up to 5 tokens ahead, and the model checks them all in one pass
(--spec-type draft-mtp-adaptive with LLAMA_SPEC_ADAPTIVE_RATE=1). It picks its own guessing
depth from how often recent guesses were kept. On the agent conversations (tool calls, code) it
made writing 23% faster. On prose it breaks even, because every guess, right or wrong, needs its own experts
fetched. An optional token list (LLAMA_MTP_VOCAB) narrows what it guesses from, for another +2.6%.
The model's sparse attention picks the context blocks each query attends to. Union attention
shares one deduplicated pick list per batch while reading a prompt. At 160K context its working
memory drops from 5.6 GiB to 1.9 GiB, and prompt reading gets 5% faster. It sums in a different
order, so words can differ at near-ties; perplexity is unchanged. It is the default
(LLAMA_QSA_UNION=1).
One server, two users: a person chatting and a coding agent working in the background. With
--chat-window 300, a chat request stops a running background request, which gets a 503 that its
client retries. Background requests then wait until 5 minutes after the last chat reply. Only the
background client needs to send a header (X-Lane: code). It is off by default.
--ple-shelf, +3% writing).You need an Apple Silicon Mac with 64 GB, about 115 GB free on the internal SSD, Xcode's command line tools and CMake. A second fast SSD is optional.
1. Build
git clone https://github.com/skeggsguy/Flash-next-ssd
cd Flash-next-ssd
cmake -B build -DGGML_METAL=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
2. Download the model (111.3 GB, 4 files) and the draft head (1.9 GB)
D=~/models/flashnext && mkdir -p $D
for i in 1 2 3 4; do
curl -fL --retry 5 -C - -o $D/Qwen3.8-Flash-Next-UD-Q4_K_XL-0000$i-of-00004.gguf \
https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/resolve/main/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-0000$i-of-00004.gguf
done
curl -fL --retry 5 -C - -o $D/mtp-shared-Q4_K_M.gguf \
https://huggingface.co/nitinpanj/qwen38-flash-next-v3/resolve/main/MTP/mtp-shared-Q4_K_M.gguf
3. Let the GPU wire enough memory (this resets on every reboot)
sudo sysctl iogpu.wired_limit_mb=59392
4. Run the server
LLAMA_SPEC_ADAPTIVE_RATE=1 LLAMA_MOE_STREAM_ALLOC_CHUNK_MIB=4096 \
GGML_METAL_RESIDENCY_KEEP_ALIVE_S=10000000 \
./build/bin/llama-server \
-m ~/models/flashnext/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -ngl 99 \
--moe-stream --moe-stream-cache 28 --moe-stream-io-threads 8 --moe-stream-direct \
-md ~/models/flashnext/mtp-shared-Q4_K_M.gguf --spec-draft-ngl 99 \
--spec-type draft-mtp-adaptive --spec-draft-n-max 5 --spec-draft-p-min 0.3 --spec-max-prompt 0 \
-c 200000 -b 4096 -ub 4096 -cms 8192 -ctxcp 3 -np 1 -fa on \
--cache-reuse 0 --cache-ram 0 --jinja --reasoning-format deepseek \
--host 127.0.0.1 --port 8080
It loads in seconds; the first request is slower while the expert cache fills. Open
http://127.0.0.1:8080, or point any OpenAI-compatible client at http://127.0.0.1:8080/v1
(the Anthropic-style /v1/messages endpoint works too).
With a second SSD, copy the four model files to it and add
--moe-stream-alt-path /Volumes/<ssd>/flashnext/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf --moe-stream-alt-split 53.
To share the server with a coding agent, add --chat-window 300.
This is the study's everyday command. Close memory-hungry apps while it runs: at a 28 GiB cache
the Mac has little room to spare. If you see swap (sysctl vm.swapusage), use a smaller cache or
context; Memory sizing explains how to choose.
Once it is built and downloaded (steps 1-3 above), double-click one of these in Finder, in the repo's folder:
Start Flash-Next.command: the model on the internal SSD only.Start Flash-Next (two SSDs).command: the internal SSD and a copy on a second SSD
(Getting started explains the copy).A Terminal window opens and runs the server with the everyday command from step 4, plus the chat window (the table below). Keep the window open while you use it; to stop the server, press Control-C or close the window.
Before it starts, it checks that the server is built, the model files are where the settings say,
the second drive is connected (two SSDs only) and no other server is running. If something is
missing, it says what to do and stops. It never starts a second server: two at once would push the
Mac into swap. To see the checks and the exact command without starting anything, run it from
Terminal with --dry-run:
./"Start Flash-Next (two SSDs).command" --dry-run
Your settings live in flashnext.conf, beside the start files. The first start makes it from
flashnext.conf.example; edit it in any text editor. It is yours: git ignores it, so pulling an
update never changes it. A setting you delete takes the example's value.
| Setting | Default | What it is |
|---|---|---|
MODEL | ~/models/flashnext/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf | The model's first file; the other three sit beside it. |
DRAFT_HEAD | ~/models/flashnext/mtp-shared-Q4_K_M.gguf | The MTP draft head. Empty (DRAFT_HEAD=) runs without it. |
SECOND_COPY | /Volumes/ssd/flashnext/...-00001-of-00004.gguf | Two SSDs only: the first file of the copy on the second drive. |
SECOND_SPLIT | 53 | Two SSDs only: the percent of experts read from the internal copy (Settings). |
CACHE_GIB | 28 | The expert cache in GiB, the biggest speed setting (Memory sizing). |
CONTEXT | 200000 | The context in tokens. |
CHAT_WINDOW | 300 | Seconds chat keeps the server after its last reply, when you share it with a coding agent (The chat window); 0 = off. Without such an agent it changes nothing. |
PORT | 8080 | The server's address is http://127.0.0.1:PORT. |
Everything else is fixed at the everyday values (the table in
Getting started says what each is for). To change one of
those, run llama-server yourself with the options in Settings.
To have the server start by itself whenever you log in, install it as a launchd agent:
tools/flashnext/install-autostart.sh # internal SSD only
tools/flashnext/install-autostart.sh two-ssds # or with the second drive
It checks your settings first, then copies a filled-in
tools/flashnext/launchd/com.flashnext.server.plist.template to
~/Library/LaunchAgents/com.flashnext.server.plist and starts the server in the background. Add
--no-load to wait for the next login instead. From then on, at every login it:
flashnext.conf;~/Library/Logs/flashnext/server.log.To stop it until the next login, and to start it again:
launchctl bootout gui/$(id -u)/com.flashnext.server
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.flashnext.server.plist
tools/flashnext/uninstall-autostart.sh stops it and removes the agent.
It needs you logged in. The GPU (Metal) works only in a logged-in session, so the agent starts
at login, not at power-on. For a Mac that should come back by itself after a power cut, turn on
automatic login (System Settings > Users & Groups; macOS offers it only with FileVault off, since
FileVault waits for your password at the unlock screen) and sudo pmset -a autorestart 1 (start up
after a power failure). Install the wired-limit LaunchDaemon from
Getting started too, so the
limit is set at every boot.
The rest of llama.cpp works as upstream documents it: see docs/ and the upstream README. npanj's notes on the model's architecture and kernels are in docs/Qwen3.8-Flash-Next.md, and the README of his fork is kept as docs/npanj-fork-README.md.
This is a fork of a fork of a fork, and it stands on the work underneath it:
--moe-stream), phase-aware batching, the SSD context cache and
MTP rejection sampling. Without that work none of this runs.This fork adds two-drive striping, the reading room, the measured-depth draft head and its token list, union attention, the n-gram shelf, the pinned checkpoint, the cheaper per-layer stops and the chat window, and the study that measured them.
MIT, as upstream: see LICENSE. Bundled third-party code keeps its own licences, listed in licenses/.
Run Qwen3.8-Flash-Next, a 111 GB model, on a 64 GB Mac.
This is a llama.cpp fork for Apple Silicon. The model is a mixture of experts: 48 layers of 512 experts, and each token uses only 10 experts on each layer. So the experts stay on the SSD. The ones a token needs are read in as it is computed, and the ones used most stay in an expert cache in RAM. Everything else stays resident on the GPU.
It was tuned for a Mac mini M5 Pro with 64 GB and a second SSD over Thunderbolt 5. The tuning was measured on real agent conversations, and any timing that swap touched was thrown out.
Measured on a Mac mini M5 Pro (64 GB) with Unsloth's UD-Q4_K_XL quant, a 28 GiB expert cache, two SSDs and the MTP draft head on. Sources and method are in Results.
| Writing (decode) | 17.5 tokens/s across 120 real agent conversations replayed in order: +47% over the study's first setup (11.9) |
| Time to first token | 5.5 s median: −52% (from 11.4 s). A typical first reply reads ~1,500 new tokens; the other ~12,800 come from the prompt cache |
| Reading a prompt (prefill) | 523 / 430 / 391 tokens/s at 4K / 32K / 100K tokens, from cold |
| Quality | 18/20 on a fixed 20-task exam, the same score as the plain setup |
--moe-stream keeps the routed experts on disk and holds a fixed-size expert cache in RAM
(--moe-stream-cache 28 = 28 GiB). Reads skip macOS's file cache (--moe-stream-direct), so the
cache and the rest of the system keep their memory, and many reads are in flight at once. While a
token is being written, the next layer's experts are fetched ahead of time; about 72% of those
guesses turn out right. The cache keeps the experts used most recently and most often. At 28 GiB
about 3 in 4 expert lookups find their expert already in RAM.
A long prompt needs nearly every expert on every layer. The reading room is a small ring buffer (about 1.4 GiB) carved out of the expert cache. The SSDs fill it layer after layer, ahead of the GPU, so prompt reading costs whichever of the two is slower, not both added together. The output is bit-identical to having every expert in RAM. While writing, the buffer is lent back as extra cache, which won back its cost (+5.5% writing). It is on by default.
Put a byte-identical copy of the model on a second SSD and the expert reads are split across both
by expert id (--moe-stream-alt-path, --moe-stream-alt-split 53). The internal SSD (6.5 GB/s)
and a 990 PRO over Thunderbolt 5 (5.7 GB/s) give 12.3 GB/s together. That made writing 15% faster
and cut the time to first token by 31%.
A small draft head guesses up to 5 tokens ahead, and the model checks them all in one pass
(--spec-type draft-mtp-adaptive with LLAMA_SPEC_ADAPTIVE_RATE=1). It picks its own guessing
depth from how often recent guesses were kept. On the agent conversations (tool calls, code) it
made writing 23% faster. On prose it breaks even, because every guess, right or wrong, needs its own experts
fetched. An optional token list (LLAMA_MTP_VOCAB) narrows what it guesses from, for another +2.6%.
The model's sparse attention picks the context blocks each query attends to. Union attention
shares one deduplicated pick list per batch while reading a prompt. At 160K context its working
memory drops from 5.6 GiB to 1.9 GiB, and prompt reading gets 5% faster. It sums in a different
order, so words can differ at near-ties; perplexity is unchanged. It is the default
(LLAMA_QSA_UNION=1).
One server, two users: a person chatting and a coding agent working in the background. With
--chat-window 300, a chat request stops a running background request, which gets a 503 that its
client retries. Background requests then wait until 5 minutes after the last chat reply. Only the
background client needs to send a header (X-Lane: code). It is off by default.
--ple-shelf, +3% writing).You need an Apple Silicon Mac with 64 GB, about 115 GB free on the internal SSD, Xcode's command line tools and CMake. A second fast SSD is optional.
1. Build
git clone https://github.com/skeggsguy/Flash-next-ssd
cd Flash-next-ssd
cmake -B build -DGGML_METAL=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
2. Download the model (111.3 GB, 4 files) and the draft head (1.9 GB)
D=~/models/flashnext && mkdir -p $D
for i in 1 2 3 4; do
curl -fL --retry 5 -C - -o $D/Qwen3.8-Flash-Next-UD-Q4_K_XL-0000$i-of-00004.gguf \
https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/resolve/main/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-0000$i-of-00004.gguf
done
curl -fL --retry 5 -C - -o $D/mtp-shared-Q4_K_M.gguf \
https://huggingface.co/nitinpanj/qwen38-flash-next-v3/resolve/main/MTP/mtp-shared-Q4_K_M.gguf
3. Let the GPU wire enough memory (this resets on every reboot)
sudo sysctl iogpu.wired_limit_mb=59392
4. Run the server
LLAMA_SPEC_ADAPTIVE_RATE=1 LLAMA_MOE_STREAM_ALLOC_CHUNK_MIB=4096 \
GGML_METAL_RESIDENCY_KEEP_ALIVE_S=10000000 \
./build/bin/llama-server \
-m ~/models/flashnext/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -ngl 99 \
--moe-stream --moe-stream-cache 28 --moe-stream-io-threads 8 --moe-stream-direct \
-md ~/models/flashnext/mtp-shared-Q4_K_M.gguf --spec-draft-ngl 99 \
--spec-type draft-mtp-adaptive --spec-draft-n-max 5 --spec-draft-p-min 0.3 --spec-max-prompt 0 \
-c 200000 -b 4096 -ub 4096 -cms 8192 -ctxcp 3 -np 1 -fa on \
--cache-reuse 0 --cache-ram 0 --jinja --reasoning-format deepseek \
--host 127.0.0.1 --port 8080
It loads in seconds; the first request is slower while the expert cache fills. Open
http://127.0.0.1:8080, or point any OpenAI-compatible client at http://127.0.0.1:8080/v1
(the Anthropic-style /v1/messages endpoint works too).
With a second SSD, copy the four model files to it and add
--moe-stream-alt-path /Volumes/<ssd>/flashnext/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf --moe-stream-alt-split 53.
To share the server with a coding agent, add --chat-window 300.
This is the study's everyday command. Close memory-hungry apps while it runs: at a 28 GiB cache
the Mac has little room to spare. If you see swap (sysctl vm.swapusage), use a smaller cache or
context; Memory sizing explains how to choose.
Once it is built and downloaded (steps 1-3 above), double-click one of these in Finder, in the repo's folder:
Start Flash-Next.command: the model on the internal SSD only.Start Flash-Next (two SSDs).command: the internal SSD and a copy on a second SSD
(Getting started explains the copy).A Terminal window opens and runs the server with the everyday command from step 4, plus the chat window (the table below). Keep the window open while you use it; to stop the server, press Control-C or close the window.
Before it starts, it checks that the server is built, the model files are where the settings say,
the second drive is connected (two SSDs only) and no other server is running. If something is
missing, it says what to do and stops. It never starts a second server: two at once would push the
Mac into swap. To see the checks and the exact command without starting anything, run it from
Terminal with --dry-run:
./"Start Flash-Next (two SSDs).command" --dry-run
Your settings live in flashnext.conf, beside the start files. The first start makes it from
flashnext.conf.example; edit it in any text editor. It is yours: git ignores it, so pulling an
update never changes it. A setting you delete takes the example's value.
| Setting | Default | What it is |
|---|---|---|
MODEL | ~/models/flashnext/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf | The model's first file; the other three sit beside it. |
DRAFT_HEAD | ~/models/flashnext/mtp-shared-Q4_K_M.gguf | The MTP draft head. Empty (DRAFT_HEAD=) runs without it. |
SECOND_COPY | /Volumes/ssd/flashnext/...-00001-of-00004.gguf | Two SSDs only: the first file of the copy on the second drive. |
SECOND_SPLIT | 53 | Two SSDs only: the percent of experts read from the internal copy (Settings). |
CACHE_GIB | 28 | The expert cache in GiB, the biggest speed setting (Memory sizing). |
CONTEXT | 200000 | The context in tokens. |
CHAT_WINDOW | 300 | Seconds chat keeps the server after its last reply, when you share it with a coding agent (The chat window); 0 = off. Without such an agent it changes nothing. |
PORT | 8080 | The server's address is http://127.0.0.1:PORT. |
Everything else is fixed at the everyday values (the table in
Getting started says what each is for). To change one of
those, run llama-server yourself with the options in Settings.
To have the server start by itself whenever you log in, install it as a launchd agent:
tools/flashnext/install-autostart.sh # internal SSD only
tools/flashnext/install-autostart.sh two-ssds # or with the second drive
It checks your settings first, then copies a filled-in
tools/flashnext/launchd/com.flashnext.server.plist.template to
~/Library/LaunchAgents/com.flashnext.server.plist and starts the server in the background. Add
--no-load to wait for the next login instead. From then on, at every login it:
flashnext.conf;~/Library/Logs/flashnext/server.log.To stop it until the next login, and to start it again:
launchctl bootout gui/$(id -u)/com.flashnext.server
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.flashnext.server.plist
tools/flashnext/uninstall-autostart.sh stops it and removes the agent.
It needs you logged in. The GPU (Metal) works only in a logged-in session, so the agent starts
at login, not at power-on. For a Mac that should come back by itself after a power cut, turn on
automatic login (System Settings > Users & Groups; macOS offers it only with FileVault off, since
FileVault waits for your password at the unlock screen) and sudo pmset -a autorestart 1 (start up
after a power failure). Install the wired-limit LaunchDaemon from
Getting started too, so the
limit is set at every boot.
The rest of llama.cpp works as upstream documents it: see docs/ and the upstream README. npanj's notes on the model's architecture and kernels are in docs/Qwen3.8-Flash-Next.md, and the README of his fork is kept as docs/npanj-fork-README.md.
This is a fork of a fork of a fork, and it stands on the work underneath it:
--moe-stream), phase-aware batching, the SSD context cache and
MTP rejection sampling. Without that work none of this runs.This fork adds two-drive striping, the reading room, the measured-depth draft head and its token list, union attention, the n-gram shelf, the pinned checkpoint, the cheaper per-layer stops and the chat window, and the study that measured them.
MIT, as upstream: see LICENSE. Bundled third-party code keeps its own licences, listed in licenses/.