paperniuk/ds4

Fork of antirez/ds4 for Apple Silicon, tuned mostly for M1/M2: Qwen3.8-Flash-Next with 262K context on 64 GB and up, 35 tok/s on an M1 Max, MTP and vision, one-line install.

C

1

736 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

Qwen3.8-Flash-Next on a 2021 M1 Max: 44 tok/s, and still 35 tok/s with 400K tokens in the context (r/LocalLLM)

After my [Splash M1 port](https://www.reddit.com/r/LocalLLM/comments/1woq7cd/you_can_now_run_qwen3827b_on_a_2021_m1_max_at_39/) the most common request was Qwen3.8-Flash-Next. Adding its architecture to Splash from scratch (hybrid recurrent layers, indexed sparse attention, n-gram tables, MTP)…

1

Oct 3, 2026

README

ds4 Flash-Next: Qwen3.8-Flash-Next on Apple Silicon

A fork of antirez/ds4 (DwarfStar), tuned mostly for M1/M2, that runs Qwen3.8-Flash-Next on Apple Silicon Macs with 64 GB of memory or more: the full 262K context, MTP speculative decoding and vision, as a local OpenAI/Anthropic compatible server for OpenCode, Claude Code, Pi or any other agent. The same launcher also runs DeepSeek V4 Flash and the other ds4 models, see Other models.

What it adds to stock ds4:

  • The small ISTA-DASLab GSQ-RCO files. Their weights take 35 to 51 GiB, so the model fits a 64 GB Mac with room for the whole context window. Stock ds4 does not open these files; its own take 42 and 70 GiB.
  • One launcher. dstar serve picks the context that fits the memory of the machine, turns on MTP and vision, and prints the one sysctl line it may need.
  • Agent sessions that do not replay. A retried or edited turn on a 31K prompt takes 0.3 s instead of 109 s, and a new session with the same system prompt and tools starts from a checkpoint on disk.
  • Metal kernels for these quants, written and measured on an M1 Max, where stock ds4 decoded at 21 to 23 tokens per second and this fork decodes at 35.
  • DeepSeek V4 Flash on a 64 GB Mac. The 81 GiB Q2 file is larger than the memory, so dstar serve deepseek streams it from the SSD: 10.35 tokens per second on an M1 Max, against 9.4 in stock ds4. See the numbers.

Nearly all of it is the same code on every Apple Silicon chip. What is and what is not specific to M1 is listed in Which Macs gain.

Qwen3.8-Flash-Next: which quant for your Mac

Qwen3.8-Flash-Next is the main model here and what dstar pull downloads. ISTA-DASLab publishes it in three sizes (quants); a larger one answers better and is slower. --quant picks one, in dstar pull and in dstar serve:

Your MacQuant--quantWeightsContext it runs with
64 GBQ2_0, the fastestq2 (default)35 GiB262K; up to 524K
64 GBIQ3_XXS, better answersiq344 GiB131K; 262K with --ctx 262k
96 GB or moreIQ3_S, the best fileiq3s51 GiB262K; up to 524K from 128 GB

The 64 GB rows are measured on an M1 Max. The IQ3_S row comes from the memory plan: on 64 GB that file runs, but only with a 32K context. ISTA's LiveCodeBench scores are 81.1, 86.3 and 86.9; speeds are in The larger quants. dstar doctor prints the largest context each quant can hold on your Mac.

On 64 GB the GPU memory limit has to be raised for the long contexts (above 131K with Q2_0, 262K with IQ3_XXS). macOS lets the GPU use about three quarters of the memory by default, 48 GiB of 64. When a context does not fit, dstar serve takes a smaller one and prints the line to run:

sudo sysctl iogpu.wired_limit_mb=57344    # 61440 for iq3 at 262K or q2 at 524K

The setting is lost at every reboot. With IQ3_XXS at 262K the server holds 56.6 GiB of the 64, so close the browser first.

Below 64 GB nothing has been tried. Qwen3.8 has no SSD streaming in ds4, so the weights must fit in memory: Q2_0 with a 32K context needs 42 GiB.

Flash-Next speed on an M1 Max

M1 Max, 32-core GPU, 64 GB, the Q2_0 file. A prompt of C code with a 300 token greedy answer, --prefill-chunk 2048:

ContextDecodeDecode with MTPPrefill
4K35.5 to 37.7 tok/s44 to 45 tok/s~355 tok/s
128K36.0 tok/s43.4 tok/s~345 tok/s
256K32.9 tok/s37.4 tok/s~290 to 320 tok/s

Prose is a little slower to prefill, about 280 to 320 tokens per second. MTP gains depend on the text: about +20 to 30% on code, about +10% on prose. In one chat grown to 398K tokens, MTP decode went from 44 to 35 tokens per second and prefill from 328 to 292. IQ3_XXS runs at about 0.8 of this speed (35 to 31 tokens per second with MTP up to 259K) and IQ3_S a little below that.

These are the only measured numbers so far. Newer chips have faster GPUs and run the same kernels, but nobody has timed them; if you do, the two commands at the end of Which Macs gain make a useful report.

DeepSeek V4 Flash on a 64 GB Mac

The other model measured here, dstar pull ds4f-q2 and dstar serve deepseek. The Q2 file is 81 GiB, so on 64 GB ds4 keeps the dense weights and a cache of routed experts in memory and reads the rest from the SSD as tokens need them. Measured on an M1 Max, 32K context, greedy decoding:

Before the changesThis fork
Decode, 400 token answer9.4 tok/s10.35 tok/s
Prefill, 5K prompt94 tok/s101 tok/s

Decode reaches 14.4 tokens per second while every expert it needs is in the cache, and falls as the answer wanders: each of the 15 or so experts read from the SSD per token costs about 2 ms. The memory plan is 48.9 GiB, of which 36.4 GiB is the cache, 5526 of the 11008 experts.

Two things to know. DSpark speculative decoding is slower here (6.3 to 7.4 tokens per second), because every drafted token pulls more experts from the disk, so the launcher leaves it off. And a busy performance core slows the GPU by up to 45%: run it on a quiet machine. With 32 GB it is not usable, 3.6 tokens per second in a simulation.

Requirements

  • An Apple Silicon Mac with 64 GB of RAM or more. The fork targets M1 and M2 (Max or Ultra); M3 and later run the same code. Only the M1 Max has been measured so far; reports from other chips are welcome.
  • For Flash-Next, about 67 GB of disk for the model (both shards) and 1.5 GB for the MTP block. A fast internal SSD, since the n-gram table is read from disk on every token.
  • For Flash-Next contexts above 131K on 64 GB, a higher GPU memory limit, see above.
  • 32 and 48 GB Macs are planned but not supported yet. Until then my Splash M1 port with Qwen3.8-27B (21 GB) is the better choice there.

Quick start

The prebuilt release (dstar 1.0), macOS 15 or newer, nothing to compile:

curl -fsSL https://raw.githubusercontent.com/paperniuk/ds4/m1-flash-next/install.sh | bash

It checks the download against its SHA-256, unpacks it into ~/dstar (DSTAR_DIR to change) and links the launcher into ~/.local/bin, so dstar works from any directory. If ~/.local/bin is not on your PATH, it prints the line to add. Then:

dstar doctor     # what your Mac fits: memory, files, the context per quant
dstar pull       # 67 GB: the model, the MTP block, the vision encoder
dstar serve      # server on http://127.0.0.1:8010/v1
dstar opencode   # provider block for OpenCode

Running the installer again updates to the latest release; the models stay where they are.

Or build from source and run the launcher from the repository folder:

git clone https://github.com/paperniuk/ds4.git
cd ds4 && make
./dstar doctor   # the same commands, as ./dstar

./dstar link puts that copy on your PATH instead, if you prefer it to the release.

dstar serve picks the context for you: the full 262K when the memory of the machine and the GPU memory limit allow it, a smaller one otherwise, and it prints the sudo sysctl line that unlocks the larger one. MTP and vision are on when their files are present.

CommandWhat it does
dstar pulldownload what is missing into ~/models/flash-next, resumable; --quant q2/iq3/iq3s picks the quant, --no-vision skips the encoder
dstar servestart the server; --ctx 131k/262k/400k/524k, --quant q2/iq3/iq3s, --port N, --lan, --no-mtp, --no-vision, --power N
dstar serve --dry-runprint the ds4-server command and environment instead of running it
dstar chattalk to the model in the terminal
dstar doctorcheck the machine, the files, the memory limit and which contexts fit
dstar opencodeprint the provider block for ~/.config/opencode/opencode.json
dstar linklink the launcher into ~/.local/bin, so dstar works from any directory
dstar modelslist the other models ds4 runs
dstar pull NAMEdownload one of them into ./gguf
dstar serve NAMEserve another model by a word from its file name (dstar serve deepseek); also dstar chat NAME

Other models

ds4 is not a Flash-Next engine, and neither is the launcher. It runs DeepSeek V4 and V4.1 Flash, GLM 5.2 and 5.3 and the upstream Qwen3.8 files too:

dstar models                  # names and sizes
dstar pull ds4f-q2            # DeepSeek V4 Flash Q2, 81 GiB
dstar serve deepseek          # or a path to the GGUF
dstar chat deepseek

The name is looked up in ./gguf and next to ~/models/flash-next. Without a name, dstar serve and dstar chat run Flash-Next when it is on disk and otherwise the one other model that is; DSTAR_MODEL=deepseek makes another model the default. dstar doctor lists what it found.

Only Flash-Next gets its context sized to the machine. For the others the context defaults to 32768 (--ctx N changes it) and is not checked against the memory. A model larger than the RAM is started with --ssd-streaming. Everything after -- goes to ds4-server unchanged. See docs/MODELS.md for what fits where.

DSTAR_MODELS changes the Flash-Next directory, DSTAR_QUANT the default quant, PORT and HOST the address. See docs/CLIENTS.md for Claude Code and other clients; use port 8010 and the context the server was started with.

--power N keeps the GPU busy N percent of the time, for a cooler and quieter Mac: the engine sleeps after every decoded token and every prefill chunk. Speed drops a little more than in proportion, --power 50 gives 16 tokens per second where the full speed is 35, and the output does not change. In dstar chat the /power N command changes it on the fly.

After a git pull, run make and restart the server. There is nothing to switch on: the speedups are the default path.

Tips for agents

  • Pick the reasoning effort once per session. Qwen3.8 writes it at the top of the system prompt, so switching it mid-session prefills the whole conversation again.
  • The server keeps one live conversation. Switching to another one saves the current state to the disk cache (--kv-disk-dir), and switching back restores it in about a second.

Status

Experimental, one developer, one machine. Parts of this fork that help every ds4 user are being proposed upstream. Everything else in ds4 (other models, CUDA, distributed inference, SSD streaming) is inherited from upstream and should keep working, but is not the focus here.

Model weights are under the Qwen community license; read it before any commercial use. The code is MIT, like upstream.

More

apple-silicon
gguf
llm-inference
metal
qwen

paperniuk/ds4

Fork of antirez/ds4 for Apple Silicon, tuned mostly for M1/M2: Qwen3.8-Flash-Next with 262K context on 64 GB and up, 35 tok/s on an M1 Max, MTP and vision, one-line install.

C

1

736 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

Qwen3.8-Flash-Next on a 2021 M1 Max: 44 tok/s, and still 35 tok/s with 400K tokens in the context (r/LocalLLM)

After my [Splash M1 port](https://www.reddit.com/r/LocalLLM/comments/1woq7cd/you_can_now_run_qwen3827b_on_a_2021_m1_max_at_39/) the most common request was Qwen3.8-Flash-Next. Adding its architecture to Splash from scratch (hybrid recurrent layers, indexed sparse attention, n-gram tables, MTP)…

1

Oct 3, 2026

README

ds4 Flash-Next: Qwen3.8-Flash-Next on Apple Silicon

A fork of antirez/ds4 (DwarfStar), tuned mostly for M1/M2, that runs Qwen3.8-Flash-Next on Apple Silicon Macs with 64 GB of memory or more: the full 262K context, MTP speculative decoding and vision, as a local OpenAI/Anthropic compatible server for OpenCode, Claude Code, Pi or any other agent. The same launcher also runs DeepSeek V4 Flash and the other ds4 models, see Other models.

What it adds to stock ds4:

  • The small ISTA-DASLab GSQ-RCO files. Their weights take 35 to 51 GiB, so the model fits a 64 GB Mac with room for the whole context window. Stock ds4 does not open these files; its own take 42 and 70 GiB.
  • One launcher. dstar serve picks the context that fits the memory of the machine, turns on MTP and vision, and prints the one sysctl line it may need.
  • Agent sessions that do not replay. A retried or edited turn on a 31K prompt takes 0.3 s instead of 109 s, and a new session with the same system prompt and tools starts from a checkpoint on disk.
  • Metal kernels for these quants, written and measured on an M1 Max, where stock ds4 decoded at 21 to 23 tokens per second and this fork decodes at 35.
  • DeepSeek V4 Flash on a 64 GB Mac. The 81 GiB Q2 file is larger than the memory, so dstar serve deepseek streams it from the SSD: 10.35 tokens per second on an M1 Max, against 9.4 in stock ds4. See the numbers.

Nearly all of it is the same code on every Apple Silicon chip. What is and what is not specific to M1 is listed in Which Macs gain.

Qwen3.8-Flash-Next: which quant for your Mac

Qwen3.8-Flash-Next is the main model here and what dstar pull downloads. ISTA-DASLab publishes it in three sizes (quants); a larger one answers better and is slower. --quant picks one, in dstar pull and in dstar serve:

Your MacQuant--quantWeightsContext it runs with
64 GBQ2_0, the fastestq2 (default)35 GiB262K; up to 524K
64 GBIQ3_XXS, better answersiq344 GiB131K; 262K with --ctx 262k
96 GB or moreIQ3_S, the best fileiq3s51 GiB262K; up to 524K from 128 GB

The 64 GB rows are measured on an M1 Max. The IQ3_S row comes from the memory plan: on 64 GB that file runs, but only with a 32K context. ISTA's LiveCodeBench scores are 81.1, 86.3 and 86.9; speeds are in The larger quants. dstar doctor prints the largest context each quant can hold on your Mac.

On 64 GB the GPU memory limit has to be raised for the long contexts (above 131K with Q2_0, 262K with IQ3_XXS). macOS lets the GPU use about three quarters of the memory by default, 48 GiB of 64. When a context does not fit, dstar serve takes a smaller one and prints the line to run:

sudo sysctl iogpu.wired_limit_mb=57344    # 61440 for iq3 at 262K or q2 at 524K

The setting is lost at every reboot. With IQ3_XXS at 262K the server holds 56.6 GiB of the 64, so close the browser first.

Below 64 GB nothing has been tried. Qwen3.8 has no SSD streaming in ds4, so the weights must fit in memory: Q2_0 with a 32K context needs 42 GiB.

Flash-Next speed on an M1 Max

M1 Max, 32-core GPU, 64 GB, the Q2_0 file. A prompt of C code with a 300 token greedy answer, --prefill-chunk 2048:

ContextDecodeDecode with MTPPrefill
4K35.5 to 37.7 tok/s44 to 45 tok/s~355 tok/s
128K36.0 tok/s43.4 tok/s~345 tok/s
256K32.9 tok/s37.4 tok/s~290 to 320 tok/s

Prose is a little slower to prefill, about 280 to 320 tokens per second. MTP gains depend on the text: about +20 to 30% on code, about +10% on prose. In one chat grown to 398K tokens, MTP decode went from 44 to 35 tokens per second and prefill from 328 to 292. IQ3_XXS runs at about 0.8 of this speed (35 to 31 tokens per second with MTP up to 259K) and IQ3_S a little below that.

These are the only measured numbers so far. Newer chips have faster GPUs and run the same kernels, but nobody has timed them; if you do, the two commands at the end of Which Macs gain make a useful report.

DeepSeek V4 Flash on a 64 GB Mac

The other model measured here, dstar pull ds4f-q2 and dstar serve deepseek. The Q2 file is 81 GiB, so on 64 GB ds4 keeps the dense weights and a cache of routed experts in memory and reads the rest from the SSD as tokens need them. Measured on an M1 Max, 32K context, greedy decoding:

Before the changesThis fork
Decode, 400 token answer9.4 tok/s10.35 tok/s
Prefill, 5K prompt94 tok/s101 tok/s

Decode reaches 14.4 tokens per second while every expert it needs is in the cache, and falls as the answer wanders: each of the 15 or so experts read from the SSD per token costs about 2 ms. The memory plan is 48.9 GiB, of which 36.4 GiB is the cache, 5526 of the 11008 experts.

Two things to know. DSpark speculative decoding is slower here (6.3 to 7.4 tokens per second), because every drafted token pulls more experts from the disk, so the launcher leaves it off. And a busy performance core slows the GPU by up to 45%: run it on a quiet machine. With 32 GB it is not usable, 3.6 tokens per second in a simulation.

Requirements

  • An Apple Silicon Mac with 64 GB of RAM or more. The fork targets M1 and M2 (Max or Ultra); M3 and later run the same code. Only the M1 Max has been measured so far; reports from other chips are welcome.
  • For Flash-Next, about 67 GB of disk for the model (both shards) and 1.5 GB for the MTP block. A fast internal SSD, since the n-gram table is read from disk on every token.
  • For Flash-Next contexts above 131K on 64 GB, a higher GPU memory limit, see above.
  • 32 and 48 GB Macs are planned but not supported yet. Until then my Splash M1 port with Qwen3.8-27B (21 GB) is the better choice there.

Quick start

The prebuilt release (dstar 1.0), macOS 15 or newer, nothing to compile:

curl -fsSL https://raw.githubusercontent.com/paperniuk/ds4/m1-flash-next/install.sh | bash

It checks the download against its SHA-256, unpacks it into ~/dstar (DSTAR_DIR to change) and links the launcher into ~/.local/bin, so dstar works from any directory. If ~/.local/bin is not on your PATH, it prints the line to add. Then:

dstar doctor     # what your Mac fits: memory, files, the context per quant
dstar pull       # 67 GB: the model, the MTP block, the vision encoder
dstar serve      # server on http://127.0.0.1:8010/v1
dstar opencode   # provider block for OpenCode

Running the installer again updates to the latest release; the models stay where they are.

Or build from source and run the launcher from the repository folder:

git clone https://github.com/paperniuk/ds4.git
cd ds4 && make
./dstar doctor   # the same commands, as ./dstar

./dstar link puts that copy on your PATH instead, if you prefer it to the release.

dstar serve picks the context for you: the full 262K when the memory of the machine and the GPU memory limit allow it, a smaller one otherwise, and it prints the sudo sysctl line that unlocks the larger one. MTP and vision are on when their files are present.

CommandWhat it does
dstar pulldownload what is missing into ~/models/flash-next, resumable; --quant q2/iq3/iq3s picks the quant, --no-vision skips the encoder
dstar servestart the server; --ctx 131k/262k/400k/524k, --quant q2/iq3/iq3s, --port N, --lan, --no-mtp, --no-vision, --power N
dstar serve --dry-runprint the ds4-server command and environment instead of running it
dstar chattalk to the model in the terminal
dstar doctorcheck the machine, the files, the memory limit and which contexts fit
dstar opencodeprint the provider block for ~/.config/opencode/opencode.json
dstar linklink the launcher into ~/.local/bin, so dstar works from any directory
dstar modelslist the other models ds4 runs
dstar pull NAMEdownload one of them into ./gguf
dstar serve NAMEserve another model by a word from its file name (dstar serve deepseek); also dstar chat NAME

Other models

ds4 is not a Flash-Next engine, and neither is the launcher. It runs DeepSeek V4 and V4.1 Flash, GLM 5.2 and 5.3 and the upstream Qwen3.8 files too:

dstar models                  # names and sizes
dstar pull ds4f-q2            # DeepSeek V4 Flash Q2, 81 GiB
dstar serve deepseek          # or a path to the GGUF
dstar chat deepseek

The name is looked up in ./gguf and next to ~/models/flash-next. Without a name, dstar serve and dstar chat run Flash-Next when it is on disk and otherwise the one other model that is; DSTAR_MODEL=deepseek makes another model the default. dstar doctor lists what it found.

Only Flash-Next gets its context sized to the machine. For the others the context defaults to 32768 (--ctx N changes it) and is not checked against the memory. A model larger than the RAM is started with --ssd-streaming. Everything after -- goes to ds4-server unchanged. See docs/MODELS.md for what fits where.

DSTAR_MODELS changes the Flash-Next directory, DSTAR_QUANT the default quant, PORT and HOST the address. See docs/CLIENTS.md for Claude Code and other clients; use port 8010 and the context the server was started with.

--power N keeps the GPU busy N percent of the time, for a cooler and quieter Mac: the engine sleeps after every decoded token and every prefill chunk. Speed drops a little more than in proportion, --power 50 gives 16 tokens per second where the full speed is 35, and the output does not change. In dstar chat the /power N command changes it on the fly.

After a git pull, run make and restart the server. There is nothing to switch on: the speedups are the default path.

Tips for agents

  • Pick the reasoning effort once per session. Qwen3.8 writes it at the top of the system prompt, so switching it mid-session prefills the whole conversation again.
  • The server keeps one live conversation. Switching to another one saves the current state to the disk cache (--kv-disk-dir), and switching back restores it in about a second.

Status

Experimental, one developer, one machine. Parts of this fork that help every ds4 user are being proposed upstream. Everything else in ds4 (other models, CUDA, distributed inference, SSD streaming) is inherited from upstream and should keep working, but is not the focus here.

Model weights are under the Qwen community license; read it before any commercial use. The code is MIT, like upstream.

More

apple-silicon
gguf
llm-inference
metal
qwen

Languages

C

46.9%

Cuda

24.2%

Objective-C

12.6%

Metal

8.0%

Python

4.7%

C++

2.9%