npanj/slipstream

Python

1

123 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant (r/LocalLLaMA)

I've been working on getting the 95.5 GiB Qwen3.8-Flash-Next model to run fast on a single 64GB Mac. In my earlier posts, I shared a custom expert-streaming fork of llama.cpp . It worked, but decode capped out around \~23–27 tok/s and slowed down as context grew. Today I'm releasing **Slipstream**:…

5

Oct 1, 2026

README

Slipstream

Run 95.5 GiB Qwen3.8-Flash-Next on a single 64 GB Mac at 41–52 tok/s.

Slipstream is a lean, high-performance C++ and Metal inference engine built specifically for Apple Silicon. It combines SSD expert streaming with predictive read-ahead and Prompt Lookup + MTP speculative drafting to serve frontier-scale models that exceed your Mac's physical RAM.

It serves Qwen3.8-Flash-Next V3 (125.7B parameters, 512 routed experts, 7.3B active per token) and its Swift KV-sparse variant at 1.76x the speed of llama.cpp, with context scaling tested all the way out to 130,000 tokens without decode collapse.

Everything is open source under Apache-2.0.


Why Qwen3.8-Flash-Next V3?

Model VariantHF Checkpoint (GGUF)Active / Total WeightsReasoning (Scorecard)Peak SpeedRAM Needed
Swift-Flash-Next V3 (New KV-Sparse)nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF7.3B / 125.7B70.3% (GPQA 54.3%, MATH 62.9%)41–52 tok/s64 GB Mac
Qwen3.8-Flash-Next V3 (34k+ Downloads)nitinpanj/qwen38-flash-next-v37.3B / 125.7B67.6% (GPQA 45.7%, MATH 60.0%)40–48 tok/s64 GB Mac
  • Frontier Reasoning on a Laptop: Outperforms standard 27B dense models on GPQA Diamond (54.3% vs 45.7%) and MATH-500 (62.9% vs 60.0%), with 92% on HumanEval and 96% on GSM8K.
  • Fast Generation: 41–52 tok/s sustained decode on an M5 Pro (64 GB).
  • No Context Collapse: Maintains 32–43 tok/s out to 130,000 tokens thanks to 48 recurrent linear DeltaNet layers ($O(1)$ state growth) and only 16 full-attention layers.
  • KV-Sparsity (Swift): The Swift variant incorporates KV-sparse attention distilled from Swift-1.5 with spliced Q8 donor backbones, minimizing RAM growth in long agent sessions.

Quickstart (Step-by-Step)

Prerequisites

  • Hardware: Apple Silicon Mac with 64 GB Unified Memory (M2/M3/M4/M5 Pro/Max).
  • Disk: ~100 GB for the multi-shard GGUF weights + ~95 GB working disk space.
  • macOS: macOS 15.0+ with Command Line Tools or Xcode installed (xcode-select --install).

Step 1: Clone & Build Slipstream

git clone https://github.com/npanj/slipstream.git
cd slipstream
make -j4

Note: make compiles the native C++ runtime and Metal compute kernels into build/slipstream and build/slipstream.metallib in under a minute.


Step 2: Download the Model

We recommend the Swift KV-sparse variant for optimal reasoning accuracy and lower KV memory footprint:

# Recommended: Swift-Qwen3.8-Flash-Next V3 (95.5 GiB)
huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
    --local-dir ~/models/swift-qwen38-flash-next-v3

# Or download with fast parallel transfer if hf_transfer is installed:
HF_HUB_ENABLE_HF_TRANSFER=1 huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
    --local-dir ~/models/swift-qwen38-flash-next-v3

(Alternative: If you prefer the plain dense base model without Swift KV-sparsity:)

huggingface-cli download nitinpanj/qwen38-flash-next-v3 \
    --local-dir ~/models/qwen38-flash-next-v3

Step 3: Raise GPU Memory Limit (One-Time per Boot)

On 64 GB Macs, macOS defaults the wired GPU limit to ~48 GiB. Raise it to 58 GiB so the SSD expert cache and KV pool have ample headroom:

sudo sysctl iogpu.wired_limit_mb=59392

Step 4: Serve the Model

Point ./slipstream serve directly at the downloaded model directory:

./slipstream serve --model ~/models/swift-qwen38-flash-next-v3 --port 8090

First Run Note: On first launch, Slipstream detects the multi-shard GGUF files and prepares optimized streaming package files into <model-dir>/prepared/ (~5–7 minutes). Subsequent launches load in ~10–15 seconds.


Step 5: Connect Your Tools & Clients

The server provides a standard OpenAI-compatible API on http://127.0.0.1:8090:

curl

curl -s http://127.0.0.1:8090/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "local/swift-qwen38-flash-next-v3",
    "messages": [
      {"role": "user", "content": "Write a clean, optimal Python function for interval merging."}
    ],
    "temperature": 0.0
  }'

Python (OpenAI SDK)

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8090/v1", api_key="not-needed")

response = client.chat.completions.create(
    model="local/swift-qwen38-flash-next-v3",
    messages=[{"role": "user", "content": "Explain multi-head self-attention with linear algebra."}],
    temperature=0.0,
)
print(response.choices[0].message.content)

Oh My Pi (omp) & Coding Agents

# Launch omp connected directly to Slipstream
omp --model splash-flashnext/local/swift-qwen38-flash-next-v3 \
    --tools=read,write,edit,bash,grep,glob,todo \
    --thinking=low --approval-mode=yolo

Benchmarks & Performance

1. Head-to-Head: llama.cpp Fork vs. Slipstream

Evaluated on Apple MacBook M5 Pro (64 GB Unified Memory, Temperature 0.0):

Task DomainBenchmark / Promptllama.cpp ForkSlipstreamSpeedupllama.cpp TTFTSlipstream TTFT
Math ReasoningGSM8K (eggs derivation)24.0 tok/s43.6 tok/s1.82x4,024 ms2,337 ms
Math DerivationMATH-500 series ($p - q$)24.3 tok/s43.1 tok/s1.77x1,655 ms1,587 ms
Constraint Logic3-chair deduction25.4 tok/s46.0 tok/s1.81x1,469 ms1,042 ms
Python Codingmerge_intervals ($O(N \log N)$)19.7 tok/s35.0 tok/s1.77x1,507 ms1,070 ms
Systems CodingRust CSV parser22.7 tok/s37.5 tok/s1.65x1,257 ms859 ms
Tech CommunicationMulti-head attention22.5 tok/s39.4 tok/s1.75x1,267 ms843 ms
OVERALL AVERAGEAcross all 6 domains23.1 tok/s40.8 tok/s1.76x1,863 ms1,290 ms

Throughput Comparison


2. Reasoning Accuracy: Swift V3 vs. Plain Base V3

Evaluated across 145 standardized items under memory guard (Seed 1234, T=0.0):

BenchmarkItemsPlain Flash-Next V3Swift-Flash-Next V3Accuracy Delta
AIME 20252045.0% (9/20)45.0% (9/20)0.0%
MATH-500 (L4–5)3560.0% (21/35)62.9% (22/35)+2.9%
GPQA Diamond3545.7% (16/35)54.3% (19/35)+8.6%
GSM8K2596.0% (24/25)96.0% (24/25)0.0%
HumanEval2592.0% (23/25)92.0% (23/25)0.0%
Hard Logic5100.0% (5/5)100.0% (5/5)0.0%
OVERALL SCORECARD14567.6% (98/145)70.3% (102/145)+2.8%

Reasoning Benchmarks


3. Context Scaling: Live Telemetry to 130,000 Tokens

Measured across 3,086 live agent requests on Apple Silicon (M5 Pro 64 GB):

Context Range (Tokens)Live RunsAverage DecodeMedian Decode (p50)Peak DecodeAvg TTFT
< 1,00031441.5 tok/s41.9 tok/s59.8 tok/s2.16 s
1k – 4,0002141.0 tok/s42.5 tok/s64.5 tok/s5.26 s
4k – 8,0005843.6 tok/s43.2 tok/s67.2 tok/s7.36 s
8k – 16,00011743.6 tok/s44.6 tok/s58.2 tok/s7.91 s
16k – 32,00056238.2 tok/s40.9 tok/s58.0 tok/s13.59 s
32k – 64,0001,02935.0 tok/s37.5 tok/s55.6 tok/s13.24 s
64k – 96,00065032.4 tok/s34.7 tok/s53.9 tok/s12.81 s
96k – 130,00036432.9 tok/s33.3 tok/s43.8 tok/s7.95 s

Context Scaling


Foundation for Qwen4

The core primitives implemented in Slipstream:

  • 512-route sparse MoE streaming with predictive read-ahead
  • Quasi-Sparse Attention (QSA) indexer and selection kernels
  • Hyper-connection mixing and per-layer embedding gathers
  • Metal GPU-mapped n-gram tables
  • Single-lane speculative verification with Prompt Lookup Decoding (PLD) + Multi-Token Prediction (MTP)

...were engineered to match the upcoming model generation. Assuming Qwen4 follows Flash-Next's architectural blueprint (hybrid linear recurrence + sparse attention + routed MoE experts), Slipstream can serve as a direct template to run Qwen4 locally on Apple Silicon on day one.


Call for Porting Partners: NVIDIA (CUDA) & AMD (ROCm)

  • Current Status: Tested exclusively on an Apple MacBook Pro (M5 Pro, 64 GB Unified Memory).
  • Porting: Because I do not have access to modern NVIDIA (CUDA) or AMD (ROCm) GPU hardware, I cannot build and test those backends myself.
  • Collaboration Offer: If anyone in the community has hardware available and wants to bring these expert-streaming and speculative decoding gains to CUDA or ROCm, I am happy to collaborate and help with the port. Open an issue or reach out!

Credits & Acknowledgments

  • Splash Team (Incoai): Full credit to the creators of Splash (github.com/incoai/splash). Their C++ Metal speculative decoding design and memory architecture provided the foundation for this work. We will prepare a clean PR/patch proposing these Flash-Next and SSD streaming extensions to the Splash upstream repo.
  • ds4 Team: For their valuable insights on Metal router numerical precision (Taylor polynomial softplus expansion) and streaming scheduling designs.
  • Qwen Team: For training Qwen3.8-Flash-Next and open-sourcing the hybrid linear MTP architecture.
  • ukisai: For the Swift-1.5 distillation work enabling KV-sparse reasoning.
  • bartowski & unsloth: For donor quants and quantization tooling.
  • mihailescu2m: For initial expert streaming concepts in llama.cpp.

License

Apache-2.0. See LICENSE.

npanj/slipstream

Python

1

123 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant (r/LocalLLaMA)

I've been working on getting the 95.5 GiB Qwen3.8-Flash-Next model to run fast on a single 64GB Mac. In my earlier posts, I shared a custom expert-streaming fork of llama.cpp . It worked, but decode capped out around \~23–27 tok/s and slowed down as context grew. Today I'm releasing **Slipstream**:…

5

Oct 1, 2026

README

Slipstream

Run 95.5 GiB Qwen3.8-Flash-Next on a single 64 GB Mac at 41–52 tok/s.

Slipstream is a lean, high-performance C++ and Metal inference engine built specifically for Apple Silicon. It combines SSD expert streaming with predictive read-ahead and Prompt Lookup + MTP speculative drafting to serve frontier-scale models that exceed your Mac's physical RAM.

It serves Qwen3.8-Flash-Next V3 (125.7B parameters, 512 routed experts, 7.3B active per token) and its Swift KV-sparse variant at 1.76x the speed of llama.cpp, with context scaling tested all the way out to 130,000 tokens without decode collapse.

Everything is open source under Apache-2.0.


Why Qwen3.8-Flash-Next V3?

Model VariantHF Checkpoint (GGUF)Active / Total WeightsReasoning (Scorecard)Peak SpeedRAM Needed
Swift-Flash-Next V3 (New KV-Sparse)nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF7.3B / 125.7B70.3% (GPQA 54.3%, MATH 62.9%)41–52 tok/s64 GB Mac
Qwen3.8-Flash-Next V3 (34k+ Downloads)nitinpanj/qwen38-flash-next-v37.3B / 125.7B67.6% (GPQA 45.7%, MATH 60.0%)40–48 tok/s64 GB Mac
  • Frontier Reasoning on a Laptop: Outperforms standard 27B dense models on GPQA Diamond (54.3% vs 45.7%) and MATH-500 (62.9% vs 60.0%), with 92% on HumanEval and 96% on GSM8K.
  • Fast Generation: 41–52 tok/s sustained decode on an M5 Pro (64 GB).
  • No Context Collapse: Maintains 32–43 tok/s out to 130,000 tokens thanks to 48 recurrent linear DeltaNet layers ($O(1)$ state growth) and only 16 full-attention layers.
  • KV-Sparsity (Swift): The Swift variant incorporates KV-sparse attention distilled from Swift-1.5 with spliced Q8 donor backbones, minimizing RAM growth in long agent sessions.

Quickstart (Step-by-Step)

Prerequisites

  • Hardware: Apple Silicon Mac with 64 GB Unified Memory (M2/M3/M4/M5 Pro/Max).
  • Disk: ~100 GB for the multi-shard GGUF weights + ~95 GB working disk space.
  • macOS: macOS 15.0+ with Command Line Tools or Xcode installed (xcode-select --install).

Step 1: Clone & Build Slipstream

git clone https://github.com/npanj/slipstream.git
cd slipstream
make -j4

Note: make compiles the native C++ runtime and Metal compute kernels into build/slipstream and build/slipstream.metallib in under a minute.


Step 2: Download the Model

We recommend the Swift KV-sparse variant for optimal reasoning accuracy and lower KV memory footprint:

# Recommended: Swift-Qwen3.8-Flash-Next V3 (95.5 GiB)
huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
    --local-dir ~/models/swift-qwen38-flash-next-v3

# Or download with fast parallel transfer if hf_transfer is installed:
HF_HUB_ENABLE_HF_TRANSFER=1 huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
    --local-dir ~/models/swift-qwen38-flash-next-v3

(Alternative: If you prefer the plain dense base model without Swift KV-sparsity:)

huggingface-cli download nitinpanj/qwen38-flash-next-v3 \
    --local-dir ~/models/qwen38-flash-next-v3

Step 3: Raise GPU Memory Limit (One-Time per Boot)

On 64 GB Macs, macOS defaults the wired GPU limit to ~48 GiB. Raise it to 58 GiB so the SSD expert cache and KV pool have ample headroom:

sudo sysctl iogpu.wired_limit_mb=59392

Step 4: Serve the Model

Point ./slipstream serve directly at the downloaded model directory:

./slipstream serve --model ~/models/swift-qwen38-flash-next-v3 --port 8090

First Run Note: On first launch, Slipstream detects the multi-shard GGUF files and prepares optimized streaming package files into <model-dir>/prepared/ (~5–7 minutes). Subsequent launches load in ~10–15 seconds.


Step 5: Connect Your Tools & Clients

The server provides a standard OpenAI-compatible API on http://127.0.0.1:8090:

curl

curl -s http://127.0.0.1:8090/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "local/swift-qwen38-flash-next-v3",
    "messages": [
      {"role": "user", "content": "Write a clean, optimal Python function for interval merging."}
    ],
    "temperature": 0.0
  }'

Python (OpenAI SDK)

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8090/v1", api_key="not-needed")

response = client.chat.completions.create(
    model="local/swift-qwen38-flash-next-v3",
    messages=[{"role": "user", "content": "Explain multi-head self-attention with linear algebra."}],
    temperature=0.0,
)
print(response.choices[0].message.content)

Oh My Pi (omp) & Coding Agents

# Launch omp connected directly to Slipstream
omp --model splash-flashnext/local/swift-qwen38-flash-next-v3 \
    --tools=read,write,edit,bash,grep,glob,todo \
    --thinking=low --approval-mode=yolo

Benchmarks & Performance

1. Head-to-Head: llama.cpp Fork vs. Slipstream

Evaluated on Apple MacBook M5 Pro (64 GB Unified Memory, Temperature 0.0):

Task DomainBenchmark / Promptllama.cpp ForkSlipstreamSpeedupllama.cpp TTFTSlipstream TTFT
Math ReasoningGSM8K (eggs derivation)24.0 tok/s43.6 tok/s1.82x4,024 ms2,337 ms
Math DerivationMATH-500 series ($p - q$)24.3 tok/s43.1 tok/s1.77x1,655 ms1,587 ms
Constraint Logic3-chair deduction25.4 tok/s46.0 tok/s1.81x1,469 ms1,042 ms
Python Codingmerge_intervals ($O(N \log N)$)19.7 tok/s35.0 tok/s1.77x1,507 ms1,070 ms
Systems CodingRust CSV parser22.7 tok/s37.5 tok/s1.65x1,257 ms859 ms
Tech CommunicationMulti-head attention22.5 tok/s39.4 tok/s1.75x1,267 ms843 ms
OVERALL AVERAGEAcross all 6 domains23.1 tok/s40.8 tok/s1.76x1,863 ms1,290 ms

Throughput Comparison


2. Reasoning Accuracy: Swift V3 vs. Plain Base V3

Evaluated across 145 standardized items under memory guard (Seed 1234, T=0.0):

BenchmarkItemsPlain Flash-Next V3Swift-Flash-Next V3Accuracy Delta
AIME 20252045.0% (9/20)45.0% (9/20)0.0%
MATH-500 (L4–5)3560.0% (21/35)62.9% (22/35)+2.9%
GPQA Diamond3545.7% (16/35)54.3% (19/35)+8.6%
GSM8K2596.0% (24/25)96.0% (24/25)0.0%
HumanEval2592.0% (23/25)92.0% (23/25)0.0%
Hard Logic5100.0% (5/5)100.0% (5/5)0.0%
OVERALL SCORECARD14567.6% (98/145)70.3% (102/145)+2.8%

Reasoning Benchmarks


3. Context Scaling: Live Telemetry to 130,000 Tokens

Measured across 3,086 live agent requests on Apple Silicon (M5 Pro 64 GB):

Context Range (Tokens)Live RunsAverage DecodeMedian Decode (p50)Peak DecodeAvg TTFT
< 1,00031441.5 tok/s41.9 tok/s59.8 tok/s2.16 s
1k – 4,0002141.0 tok/s42.5 tok/s64.5 tok/s5.26 s
4k – 8,0005843.6 tok/s43.2 tok/s67.2 tok/s7.36 s
8k – 16,00011743.6 tok/s44.6 tok/s58.2 tok/s7.91 s
16k – 32,00056238.2 tok/s40.9 tok/s58.0 tok/s13.59 s
32k – 64,0001,02935.0 tok/s37.5 tok/s55.6 tok/s13.24 s
64k – 96,00065032.4 tok/s34.7 tok/s53.9 tok/s12.81 s
96k – 130,00036432.9 tok/s33.3 tok/s43.8 tok/s7.95 s

Context Scaling


Foundation for Qwen4

The core primitives implemented in Slipstream:

  • 512-route sparse MoE streaming with predictive read-ahead
  • Quasi-Sparse Attention (QSA) indexer and selection kernels
  • Hyper-connection mixing and per-layer embedding gathers
  • Metal GPU-mapped n-gram tables
  • Single-lane speculative verification with Prompt Lookup Decoding (PLD) + Multi-Token Prediction (MTP)

...were engineered to match the upcoming model generation. Assuming Qwen4 follows Flash-Next's architectural blueprint (hybrid linear recurrence + sparse attention + routed MoE experts), Slipstream can serve as a direct template to run Qwen4 locally on Apple Silicon on day one.


Call for Porting Partners: NVIDIA (CUDA) & AMD (ROCm)

  • Current Status: Tested exclusively on an Apple MacBook Pro (M5 Pro, 64 GB Unified Memory).
  • Porting: Because I do not have access to modern NVIDIA (CUDA) or AMD (ROCm) GPU hardware, I cannot build and test those backends myself.
  • Collaboration Offer: If anyone in the community has hardware available and wants to bring these expert-streaming and speculative decoding gains to CUDA or ROCm, I am happy to collaborate and help with the port. Open an issue or reach out!

Credits & Acknowledgments

  • Splash Team (Incoai): Full credit to the creators of Splash (github.com/incoai/splash). Their C++ Metal speculative decoding design and memory architecture provided the foundation for this work. We will prepare a clean PR/patch proposing these Flash-Next and SSD streaming extensions to the Splash upstream repo.
  • ds4 Team: For their valuable insights on Metal router numerical precision (Taylor polynomial softplus expansion) and streaming scheduling designs.
  • Qwen Team: For training Qwen3.8-Flash-Next and open-sourcing the hybrid linear MTP architecture.
  • ukisai: For the Swift-1.5 distillation work enabling KV-sparse reasoning.
  • bartowski & unsloth: For donor quants and quantization tooling.
  • mihailescu2m: For initial expert streaming concepts in llama.cpp.

License

Apache-2.0. See LICENSE.

Languages

Python

38.7%

C++

31.2%

Objective-C++

19.6%

Metal

7.6%

Makefile

1.1%