Run 95.5 GiB Qwen3.8-Flash-Next on a single 64 GB Mac at 41–52 tok/s.
Slipstream is a lean, high-performance C++ and Metal inference engine built specifically for Apple Silicon. It combines SSD expert streaming with predictive read-ahead and Prompt Lookup + MTP speculative drafting to serve frontier-scale models that exceed your Mac's physical RAM.
It serves Qwen3.8-Flash-Next V3 (125.7B parameters, 512 routed experts, 7.3B active per token) and its Swift KV-sparse variant at 1.76x the speed of llama.cpp, with context scaling tested all the way out to 130,000 tokens without decode collapse.
Everything is open source under Apache-2.0.
| Model Variant | HF Checkpoint (GGUF) | Active / Total Weights | Reasoning (Scorecard) | Peak Speed | RAM Needed |
|---|---|---|---|---|---|
| Swift-Flash-Next V3 (New KV-Sparse) | nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF | 7.3B / 125.7B | 70.3% (GPQA 54.3%, MATH 62.9%) | 41–52 tok/s | 64 GB Mac |
| Qwen3.8-Flash-Next V3 (34k+ Downloads) | nitinpanj/qwen38-flash-next-v3 | 7.3B / 125.7B | 67.6% (GPQA 45.7%, MATH 60.0%) | 40–48 tok/s | 64 GB Mac |
xcode-select --install).git clone https://github.com/npanj/slipstream.git
cd slipstream
make -j4
Note: make compiles the native C++ runtime and Metal compute kernels into build/slipstream and build/slipstream.metallib in under a minute.
We recommend the Swift KV-sparse variant for optimal reasoning accuracy and lower KV memory footprint:
# Recommended: Swift-Qwen3.8-Flash-Next V3 (95.5 GiB)
huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
--local-dir ~/models/swift-qwen38-flash-next-v3
# Or download with fast parallel transfer if hf_transfer is installed:
HF_HUB_ENABLE_HF_TRANSFER=1 huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
--local-dir ~/models/swift-qwen38-flash-next-v3
(Alternative: If you prefer the plain dense base model without Swift KV-sparsity:)
huggingface-cli download nitinpanj/qwen38-flash-next-v3 \
--local-dir ~/models/qwen38-flash-next-v3
On 64 GB Macs, macOS defaults the wired GPU limit to ~48 GiB. Raise it to 58 GiB so the SSD expert cache and KV pool have ample headroom:
sudo sysctl iogpu.wired_limit_mb=59392
Point ./slipstream serve directly at the downloaded model directory:
./slipstream serve --model ~/models/swift-qwen38-flash-next-v3 --port 8090
First Run Note: On first launch, Slipstream detects the multi-shard GGUF files and prepares optimized streaming package files into
<model-dir>/prepared/(~5–7 minutes). Subsequent launches load in ~10–15 seconds.
The server provides a standard OpenAI-compatible API on http://127.0.0.1:8090:
curl -s http://127.0.0.1:8090/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "local/swift-qwen38-flash-next-v3",
"messages": [
{"role": "user", "content": "Write a clean, optimal Python function for interval merging."}
],
"temperature": 0.0
}'
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8090/v1", api_key="not-needed")
response = client.chat.completions.create(
model="local/swift-qwen38-flash-next-v3",
messages=[{"role": "user", "content": "Explain multi-head self-attention with linear algebra."}],
temperature=0.0,
)
print(response.choices[0].message.content)
# Launch omp connected directly to Slipstream
omp --model splash-flashnext/local/swift-qwen38-flash-next-v3 \
--tools=read,write,edit,bash,grep,glob,todo \
--thinking=low --approval-mode=yolo
Evaluated on Apple MacBook M5 Pro (64 GB Unified Memory, Temperature 0.0):
| Task Domain | Benchmark / Prompt | llama.cpp Fork | Slipstream | Speedup | llama.cpp TTFT | Slipstream TTFT |
|---|---|---|---|---|---|---|
| Math Reasoning | GSM8K (eggs derivation) | 24.0 tok/s | 43.6 tok/s | 1.82x | 4,024 ms | 2,337 ms |
| Math Derivation | MATH-500 series ($p - q$) | 24.3 tok/s | 43.1 tok/s | 1.77x | 1,655 ms | 1,587 ms |
| Constraint Logic | 3-chair deduction | 25.4 tok/s | 46.0 tok/s | 1.81x | 1,469 ms | 1,042 ms |
| Python Coding | merge_intervals ($O(N \log N)$) | 19.7 tok/s | 35.0 tok/s | 1.77x | 1,507 ms | 1,070 ms |
| Systems Coding | Rust CSV parser | 22.7 tok/s | 37.5 tok/s | 1.65x | 1,257 ms | 859 ms |
| Tech Communication | Multi-head attention | 22.5 tok/s | 39.4 tok/s | 1.75x | 1,267 ms | 843 ms |
| OVERALL AVERAGE | Across all 6 domains | 23.1 tok/s | 40.8 tok/s | 1.76x | 1,863 ms | 1,290 ms |

Evaluated across 145 standardized items under memory guard (Seed 1234, T=0.0):
| Benchmark | Items | Plain Flash-Next V3 | Swift-Flash-Next V3 | Accuracy Delta |
|---|---|---|---|---|
| AIME 2025 | 20 | 45.0% (9/20) | 45.0% (9/20) | 0.0% |
| MATH-500 (L4–5) | 35 | 60.0% (21/35) | 62.9% (22/35) | +2.9% |
| GPQA Diamond | 35 | 45.7% (16/35) | 54.3% (19/35) | +8.6% |
| GSM8K | 25 | 96.0% (24/25) | 96.0% (24/25) | 0.0% |
| HumanEval | 25 | 92.0% (23/25) | 92.0% (23/25) | 0.0% |
| Hard Logic | 5 | 100.0% (5/5) | 100.0% (5/5) | 0.0% |
| OVERALL SCORECARD | 145 | 67.6% (98/145) | 70.3% (102/145) | +2.8% |

Measured across 3,086 live agent requests on Apple Silicon (M5 Pro 64 GB):
| Context Range (Tokens) | Live Runs | Average Decode | Median Decode (p50) | Peak Decode | Avg TTFT |
|---|---|---|---|---|---|
| < 1,000 | 314 | 41.5 tok/s | 41.9 tok/s | 59.8 tok/s | 2.16 s |
| 1k – 4,000 | 21 | 41.0 tok/s | 42.5 tok/s | 64.5 tok/s | 5.26 s |
| 4k – 8,000 | 58 | 43.6 tok/s | 43.2 tok/s | 67.2 tok/s | 7.36 s |
| 8k – 16,000 | 117 | 43.6 tok/s | 44.6 tok/s | 58.2 tok/s | 7.91 s |
| 16k – 32,000 | 562 | 38.2 tok/s | 40.9 tok/s | 58.0 tok/s | 13.59 s |
| 32k – 64,000 | 1,029 | 35.0 tok/s | 37.5 tok/s | 55.6 tok/s | 13.24 s |
| 64k – 96,000 | 650 | 32.4 tok/s | 34.7 tok/s | 53.9 tok/s | 12.81 s |
| 96k – 130,000 | 364 | 32.9 tok/s | 33.3 tok/s | 43.8 tok/s | 7.95 s |

The core primitives implemented in Slipstream:
...were engineered to match the upcoming model generation. Assuming Qwen4 follows Flash-Next's architectural blueprint (hybrid linear recurrence + sparse attention + routed MoE experts), Slipstream can serve as a direct template to run Qwen4 locally on Apple Silicon on day one.
Apache-2.0. See LICENSE.
Python
38.7%
C++
31.2%
Objective-C++
19.6%
Metal
7.6%
Makefile
1.1%
Run 95.5 GiB Qwen3.8-Flash-Next on a single 64 GB Mac at 41–52 tok/s.
Slipstream is a lean, high-performance C++ and Metal inference engine built specifically for Apple Silicon. It combines SSD expert streaming with predictive read-ahead and Prompt Lookup + MTP speculative drafting to serve frontier-scale models that exceed your Mac's physical RAM.
It serves Qwen3.8-Flash-Next V3 (125.7B parameters, 512 routed experts, 7.3B active per token) and its Swift KV-sparse variant at 1.76x the speed of llama.cpp, with context scaling tested all the way out to 130,000 tokens without decode collapse.
Everything is open source under Apache-2.0.
| Model Variant | HF Checkpoint (GGUF) | Active / Total Weights | Reasoning (Scorecard) | Peak Speed | RAM Needed |
|---|---|---|---|---|---|
| Swift-Flash-Next V3 (New KV-Sparse) | nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF | 7.3B / 125.7B | 70.3% (GPQA 54.3%, MATH 62.9%) | 41–52 tok/s | 64 GB Mac |
| Qwen3.8-Flash-Next V3 (34k+ Downloads) | nitinpanj/qwen38-flash-next-v3 | 7.3B / 125.7B | 67.6% (GPQA 45.7%, MATH 60.0%) | 40–48 tok/s | 64 GB Mac |
xcode-select --install).git clone https://github.com/npanj/slipstream.git
cd slipstream
make -j4
Note: make compiles the native C++ runtime and Metal compute kernels into build/slipstream and build/slipstream.metallib in under a minute.
We recommend the Swift KV-sparse variant for optimal reasoning accuracy and lower KV memory footprint:
# Recommended: Swift-Qwen3.8-Flash-Next V3 (95.5 GiB)
huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
--local-dir ~/models/swift-qwen38-flash-next-v3
# Or download with fast parallel transfer if hf_transfer is installed:
HF_HUB_ENABLE_HF_TRANSFER=1 huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
--local-dir ~/models/swift-qwen38-flash-next-v3
(Alternative: If you prefer the plain dense base model without Swift KV-sparsity:)
huggingface-cli download nitinpanj/qwen38-flash-next-v3 \
--local-dir ~/models/qwen38-flash-next-v3
On 64 GB Macs, macOS defaults the wired GPU limit to ~48 GiB. Raise it to 58 GiB so the SSD expert cache and KV pool have ample headroom:
sudo sysctl iogpu.wired_limit_mb=59392
Point ./slipstream serve directly at the downloaded model directory:
./slipstream serve --model ~/models/swift-qwen38-flash-next-v3 --port 8090
First Run Note: On first launch, Slipstream detects the multi-shard GGUF files and prepares optimized streaming package files into
<model-dir>/prepared/(~5–7 minutes). Subsequent launches load in ~10–15 seconds.
The server provides a standard OpenAI-compatible API on http://127.0.0.1:8090:
curl -s http://127.0.0.1:8090/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "local/swift-qwen38-flash-next-v3",
"messages": [
{"role": "user", "content": "Write a clean, optimal Python function for interval merging."}
],
"temperature": 0.0
}'
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8090/v1", api_key="not-needed")
response = client.chat.completions.create(
model="local/swift-qwen38-flash-next-v3",
messages=[{"role": "user", "content": "Explain multi-head self-attention with linear algebra."}],
temperature=0.0,
)
print(response.choices[0].message.content)
# Launch omp connected directly to Slipstream
omp --model splash-flashnext/local/swift-qwen38-flash-next-v3 \
--tools=read,write,edit,bash,grep,glob,todo \
--thinking=low --approval-mode=yolo
Evaluated on Apple MacBook M5 Pro (64 GB Unified Memory, Temperature 0.0):
| Task Domain | Benchmark / Prompt | llama.cpp Fork | Slipstream | Speedup | llama.cpp TTFT | Slipstream TTFT |
|---|---|---|---|---|---|---|
| Math Reasoning | GSM8K (eggs derivation) | 24.0 tok/s | 43.6 tok/s | 1.82x | 4,024 ms | 2,337 ms |
| Math Derivation | MATH-500 series ($p - q$) | 24.3 tok/s | 43.1 tok/s | 1.77x | 1,655 ms | 1,587 ms |
| Constraint Logic | 3-chair deduction | 25.4 tok/s | 46.0 tok/s | 1.81x | 1,469 ms | 1,042 ms |
| Python Coding | merge_intervals ($O(N \log N)$) | 19.7 tok/s | 35.0 tok/s | 1.77x | 1,507 ms | 1,070 ms |
| Systems Coding | Rust CSV parser | 22.7 tok/s | 37.5 tok/s | 1.65x | 1,257 ms | 859 ms |
| Tech Communication | Multi-head attention | 22.5 tok/s | 39.4 tok/s | 1.75x | 1,267 ms | 843 ms |
| OVERALL AVERAGE | Across all 6 domains | 23.1 tok/s | 40.8 tok/s | 1.76x | 1,863 ms | 1,290 ms |

Evaluated across 145 standardized items under memory guard (Seed 1234, T=0.0):
| Benchmark | Items | Plain Flash-Next V3 | Swift-Flash-Next V3 | Accuracy Delta |
|---|---|---|---|---|
| AIME 2025 | 20 | 45.0% (9/20) | 45.0% (9/20) | 0.0% |
| MATH-500 (L4–5) | 35 | 60.0% (21/35) | 62.9% (22/35) | +2.9% |
| GPQA Diamond | 35 | 45.7% (16/35) | 54.3% (19/35) | +8.6% |
| GSM8K | 25 | 96.0% (24/25) | 96.0% (24/25) | 0.0% |
| HumanEval | 25 | 92.0% (23/25) | 92.0% (23/25) | 0.0% |
| Hard Logic | 5 | 100.0% (5/5) | 100.0% (5/5) | 0.0% |
| OVERALL SCORECARD | 145 | 67.6% (98/145) | 70.3% (102/145) | +2.8% |

Measured across 3,086 live agent requests on Apple Silicon (M5 Pro 64 GB):
| Context Range (Tokens) | Live Runs | Average Decode | Median Decode (p50) | Peak Decode | Avg TTFT |
|---|---|---|---|---|---|
| < 1,000 | 314 | 41.5 tok/s | 41.9 tok/s | 59.8 tok/s | 2.16 s |
| 1k – 4,000 | 21 | 41.0 tok/s | 42.5 tok/s | 64.5 tok/s | 5.26 s |
| 4k – 8,000 | 58 | 43.6 tok/s | 43.2 tok/s | 67.2 tok/s | 7.36 s |
| 8k – 16,000 | 117 | 43.6 tok/s | 44.6 tok/s | 58.2 tok/s | 7.91 s |
| 16k – 32,000 | 562 | 38.2 tok/s | 40.9 tok/s | 58.0 tok/s | 13.59 s |
| 32k – 64,000 | 1,029 | 35.0 tok/s | 37.5 tok/s | 55.6 tok/s | 13.24 s |
| 64k – 96,000 | 650 | 32.4 tok/s | 34.7 tok/s | 53.9 tok/s | 12.81 s |
| 96k – 130,000 | 364 | 32.9 tok/s | 33.3 tok/s | 43.8 tok/s | 7.95 s |

The core primitives implemented in Slipstream:
...were engineered to match the upcoming model generation. Assuming Qwen4 follows Flash-Next's architectural blueprint (hybrid linear recurrence + sparse attention + routed MoE experts), Slipstream can serve as a direct template to run Qwen4 locally on Apple Silicon on day one.
Apache-2.0. See LICENSE.
Python
38.7%
C++
31.2%
Objective-C++
19.6%
Metal
7.6%
Makefile
1.1%