npanj/splash-plus

Empirical benchmarks and quality evaluations of Qwen3.8-27B with Multi-Token Prediction (MTP) across Splash, MTPLX, MLX, and llama.cpp on Apple Silicon

Python

13

13 commits

updated Sep 21, 2026

See the code

See what people are saying

SourceMessageScoreDate

Qwen3.8-Flash-Next on a 2021 M1 Max: 44 tok/s, and still 35 tok/s with 400K tokens in the context (r/LocalLLM)

After my [Splash M1 port](https://www.reddit.com/r/LocalLLM/comments/1woq7cd/you_can_now_run_qwen3827b_on_a_2021_m1_max_at_39/) the most common request was Qwen3.8-Flash-Next. Adding its architecture to Splash from scratch (hybrid recurrent layers, indexed sparse attention, n-gram tables, MTP)…

1

Oct 3, 2026

README

Splash-Plus: Qwen3.8-27B on Apple Silicon (Benchmarking, Evaluation & Q8 Extensions)

Empirical benchmark suite, precision analysis, and quality evaluation comparing Splash-Plus / Splash (Q4, Q8, Mixed, HQ) against MTPLX, MLX, and llama.cpp running Qwen3.8-27B on Apple Silicon with Multi-Token Prediction (MTP).


Executive Summary

  • Throughput (6.1x Peak Speedup): Splash-Q4 hits 60.7 tok/s average (peaking at 83.3 tok/s), outperforming standard autoregressive MLX and llama.cpp baselines (9.9 tok/s) by over 6x.
  • Direct 8-bit MTP Showdown: Evaluating on the exact same 8-bit weight representation, Splash-HQ beats MTPLX-Optimized-Quality by +39% overall (36.9 vs 26.5 tok/s) and +92% on mathematical reasoning (54.8 vs 28.5 tok/s).
  • The "Reasoning Cliff": On competition-grade math (MATH-500 Problem 0: symbolic infinite double summation $\sum_{j=1}^\infty \sum_{k=1}^\infty \frac{1}{(j+k)^3} = p - q$), aggressive 4-bit and compressed 8-bit suffered numerical drift and failed. Upgrading sensitive layers (Splash-Mixed and Splash-HQ) resolved the cliff and solved the problem correctly (p - q).
  • The Precision Speed Paradox: Uncompressed native 8-bit (Splash-HQ) runs faster than compressed 8-bit (Splash-Q8) because higher-precision weights produce lower-entropy logits, increasing draft token acceptance rate and eliminating verification rollbacks.

Grand Throughput Comparison (5 Standardized Domains)

Evaluated at temperature=0.0, max_tokens=250 on identical hardware:

Task / DomainPrompt DescriptionSplash-Q4Splash-HQ (Native 8b)Splash-Q8 (Compressed)Splash-Mixed (Selective 8b)MTPLX-Q8 (MTP D3)MLX-Q8 (Stock AR)llama.cpp (Q8_0 GGUF)
Math & LogicBat & Ball algebraic derivation83.3 t/s54.8 t/s52.7 t/s48.6 t/s28.5 t/s9.9 t/s9.8 t/s
Coding & Algorithmsmerge_intervals $O(N \log N)$75.5 t/s34.7 t/s40.3 t/s35.0 t/s28.8 t/s9.9 t/s9.9 t/s
Constraint Reasoning3-chair spatial permutation59.3 t/s39.3 t/s37.4 t/s36.6 t/s27.7 t/s9.9 t/s9.9 t/s
Domain KnowledgeFlashAttn vs PagedAttn vs SpecDec39.0 t/s21.9 t/s22.8 t/s23.3 t/s23.8 t/s9.9 t/s10.0 t/s
Nuanced WritingMemory bandwidth constraint (3 sentences)46.2 t/s33.7 t/s29.5 t/s27.0 t/s23.4 t/s9.9 t/s10.0 t/s
AVERAGE SPEEDAcross all 5 domains60.7 tok/s36.9 tok/s36.5 tok/s34.1 tok/s26.5 tok/s9.9 tok/s9.9 tok/s
Speedup vs AR Baseline(Baseline = 9.9 tok/s)6.13x3.73x3.69x3.44x2.68x1.00x1.00x

Quality & Reasoning Accuracy Scorecard

Evaluated on standard benchmarks with extended Chain-of-Thought thinking budgets (GPQA Diamond, AIME 2025, MATH-500, GSM8K):

BenchmarkFocusSplash-Q4Splash-Q8Splash-MixedSplash-HQMTPLX-Q8
GSM8KGrade School Arithmetic80.0% (4/5)80.0% (4/5)——80.0% (4/5)
GPQA DiamondPhD-Level Physics & Science33.3% (1/3)66.7% (2/3)33.3% (1/3)33.3% (1/3)20.0% (1/5)
AIME 2025High School Math Olympiad66.7% (2/3)66.7% (2/3)33.3% (1/3)66.7% (2/3)0.0% (0/5)*
MATH-500Advanced Symbolic Math0.0% (0/3)0.0% (0/3)33.3% (1/3)33.3% (1/3)—
OVERALL ACCURACYCumulative Reasoning Score33.3% (3/9)44.4% (4/9)33.3% (3/9)44.4% (4/9)—

*Note: Early MTPLX test was truncated before completion on AIME due to a 400-token ceiling before expanding generation limits.


Published Hugging Face Models

ModelHugging Face RepositoryPrecision ProfileRole
Qwen3.8-27B-Splash-HQnitinpanj/Qwen3.8-27B-Splash-HQNative 8-bit uncompressed (64 layers + heads)Maximum quality, zero quant error, high draft acceptance
Qwen3.8-27B-Splash-Mixednitinpanj/Qwen3.8-27B-Splash-MixedTop 8 sensitive layers (56–63) + heads in 8-bitMemory-efficient hybrid solving the reasoning cliff
Qwen3.8-27B-Splash-Q8nitinpanj/Qwen3.8-27B-Splash-Q8Compressed 8-bit baselineHigh throughput 8-bit reference

Repository Structure

.
├── README.md               # Main benchmark overview and results
├── REPORT.md               # In-depth technical report and architectural analysis
├── scripts/
│   ├── run_all_models_standard_bench.py # 5-prompt grand comparison suite
│   ├── run_quality_benchmarks.py        # GPQA, AIME 2025, MATH-500 evaluator
│   ├── pack_mixed_q8.py                 # LRU-cached mixed-precision converter
│   ├── bench_mtplx_quality.py           # MTPLX MTP runner
│   ├── bench_mlx_q8.py                  # MLX stream generation benchmark
│   └── bench_llamacpp.py                # llama-cli Q8_0 GGUF benchmark
└── results/
    ├── splash_all_variants_5prompts.json
    ├── quality_benchmark_scorecard.json
    ├── mtplx_q8_results.json
    ├── mlx_q8_results.json
    ├── llamacpp_q8_results.json
    └── benchmark_results_q8_q4_mtp.json

How to Run / Serve the Model

To run the native 8-bit model (nitinpanj/Qwen3.8-27B-Splash-HQ) with the high-performance Metal Q8 engine:

# 1. Clone the Q8-enabled Splash fork
git clone https://github.com/npanj/splash.git -b q8
cd splash

# 2. Build the Metal kernels
make -j4

# 3. Serve the model (downloads automatically from Hugging Face on first launch)
./splash serve --model nitinpanj/Qwen3.8-27B-Splash-HQ

The server binds to http://127.0.0.1:8000 with an OpenAI-compatible API (/v1/chat/completions). Connect any coding agent or client:

# Oh My Pi (OMP)
omp --model splash/nitinpanj/Qwen3.8-27B-Splash-HQ

Quickstart & Reproduction

1. Run the 5-Prompt Head-to-Head Comparison

python3 scripts/run_all_models_standard_bench.py

2. Run the Quality Benchmark (GPQA, AIME, MATH-500)

python3 scripts/run_quality_benchmarks.py --limit 3 --max-tokens 800

3. Build a Selective Mixed-Precision Model

python3 scripts/pack_mixed_q8.py \
  --source-safetensors ~/.mtplx/models/Qwen3.8-27B-MTPLX-Optimized-Quality \
  --base-splash-dir install/models/incoai/Qwen3.8-27B-Splash-Q8 \
  --out-dir install/models/incoai/Qwen3.8-27B-Splash-Mixed \
  --upgrade-layers 56,57,58,59,60,61,62,63 \
  --upgrade-embedding \
  --upgrade-head

License

Code and benchmark tools are released under the Apache-2.0 License. Model weights follow the original Qwen License.

npanj/splash-plus

Empirical benchmarks and quality evaluations of Qwen3.8-27B with Multi-Token Prediction (MTP) across Splash, MTPLX, MLX, and llama.cpp on Apple Silicon

Python

13

13 commits

updated Sep 21, 2026

See the code

See what people are saying

SourceMessageScoreDate

Qwen3.8-Flash-Next on a 2021 M1 Max: 44 tok/s, and still 35 tok/s with 400K tokens in the context (r/LocalLLM)

After my [Splash M1 port](https://www.reddit.com/r/LocalLLM/comments/1woq7cd/you_can_now_run_qwen3827b_on_a_2021_m1_max_at_39/) the most common request was Qwen3.8-Flash-Next. Adding its architecture to Splash from scratch (hybrid recurrent layers, indexed sparse attention, n-gram tables, MTP)…

1

Oct 3, 2026

README

Splash-Plus: Qwen3.8-27B on Apple Silicon (Benchmarking, Evaluation & Q8 Extensions)

Empirical benchmark suite, precision analysis, and quality evaluation comparing Splash-Plus / Splash (Q4, Q8, Mixed, HQ) against MTPLX, MLX, and llama.cpp running Qwen3.8-27B on Apple Silicon with Multi-Token Prediction (MTP).


Executive Summary

  • Throughput (6.1x Peak Speedup): Splash-Q4 hits 60.7 tok/s average (peaking at 83.3 tok/s), outperforming standard autoregressive MLX and llama.cpp baselines (9.9 tok/s) by over 6x.
  • Direct 8-bit MTP Showdown: Evaluating on the exact same 8-bit weight representation, Splash-HQ beats MTPLX-Optimized-Quality by +39% overall (36.9 vs 26.5 tok/s) and +92% on mathematical reasoning (54.8 vs 28.5 tok/s).
  • The "Reasoning Cliff": On competition-grade math (MATH-500 Problem 0: symbolic infinite double summation $\sum_{j=1}^\infty \sum_{k=1}^\infty \frac{1}{(j+k)^3} = p - q$), aggressive 4-bit and compressed 8-bit suffered numerical drift and failed. Upgrading sensitive layers (Splash-Mixed and Splash-HQ) resolved the cliff and solved the problem correctly (p - q).
  • The Precision Speed Paradox: Uncompressed native 8-bit (Splash-HQ) runs faster than compressed 8-bit (Splash-Q8) because higher-precision weights produce lower-entropy logits, increasing draft token acceptance rate and eliminating verification rollbacks.

Grand Throughput Comparison (5 Standardized Domains)

Evaluated at temperature=0.0, max_tokens=250 on identical hardware:

Task / DomainPrompt DescriptionSplash-Q4Splash-HQ (Native 8b)Splash-Q8 (Compressed)Splash-Mixed (Selective 8b)MTPLX-Q8 (MTP D3)MLX-Q8 (Stock AR)llama.cpp (Q8_0 GGUF)
Math & LogicBat & Ball algebraic derivation83.3 t/s54.8 t/s52.7 t/s48.6 t/s28.5 t/s9.9 t/s9.8 t/s
Coding & Algorithmsmerge_intervals $O(N \log N)$75.5 t/s34.7 t/s40.3 t/s35.0 t/s28.8 t/s9.9 t/s9.9 t/s
Constraint Reasoning3-chair spatial permutation59.3 t/s39.3 t/s37.4 t/s36.6 t/s27.7 t/s9.9 t/s9.9 t/s
Domain KnowledgeFlashAttn vs PagedAttn vs SpecDec39.0 t/s21.9 t/s22.8 t/s23.3 t/s23.8 t/s9.9 t/s10.0 t/s
Nuanced WritingMemory bandwidth constraint (3 sentences)46.2 t/s33.7 t/s29.5 t/s27.0 t/s23.4 t/s9.9 t/s10.0 t/s
AVERAGE SPEEDAcross all 5 domains60.7 tok/s36.9 tok/s36.5 tok/s34.1 tok/s26.5 tok/s9.9 tok/s9.9 tok/s
Speedup vs AR Baseline(Baseline = 9.9 tok/s)6.13x3.73x3.69x3.44x2.68x1.00x1.00x

Quality & Reasoning Accuracy Scorecard

Evaluated on standard benchmarks with extended Chain-of-Thought thinking budgets (GPQA Diamond, AIME 2025, MATH-500, GSM8K):

BenchmarkFocusSplash-Q4Splash-Q8Splash-MixedSplash-HQMTPLX-Q8
GSM8KGrade School Arithmetic80.0% (4/5)80.0% (4/5)——80.0% (4/5)
GPQA DiamondPhD-Level Physics & Science33.3% (1/3)66.7% (2/3)33.3% (1/3)33.3% (1/3)20.0% (1/5)
AIME 2025High School Math Olympiad66.7% (2/3)66.7% (2/3)33.3% (1/3)66.7% (2/3)0.0% (0/5)*
MATH-500Advanced Symbolic Math0.0% (0/3)0.0% (0/3)33.3% (1/3)33.3% (1/3)—
OVERALL ACCURACYCumulative Reasoning Score33.3% (3/9)44.4% (4/9)33.3% (3/9)44.4% (4/9)—

*Note: Early MTPLX test was truncated before completion on AIME due to a 400-token ceiling before expanding generation limits.


Published Hugging Face Models

ModelHugging Face RepositoryPrecision ProfileRole
Qwen3.8-27B-Splash-HQnitinpanj/Qwen3.8-27B-Splash-HQNative 8-bit uncompressed (64 layers + heads)Maximum quality, zero quant error, high draft acceptance
Qwen3.8-27B-Splash-Mixednitinpanj/Qwen3.8-27B-Splash-MixedTop 8 sensitive layers (56–63) + heads in 8-bitMemory-efficient hybrid solving the reasoning cliff
Qwen3.8-27B-Splash-Q8nitinpanj/Qwen3.8-27B-Splash-Q8Compressed 8-bit baselineHigh throughput 8-bit reference

Repository Structure

.
├── README.md               # Main benchmark overview and results
├── REPORT.md               # In-depth technical report and architectural analysis
├── scripts/
│   ├── run_all_models_standard_bench.py # 5-prompt grand comparison suite
│   ├── run_quality_benchmarks.py        # GPQA, AIME 2025, MATH-500 evaluator
│   ├── pack_mixed_q8.py                 # LRU-cached mixed-precision converter
│   ├── bench_mtplx_quality.py           # MTPLX MTP runner
│   ├── bench_mlx_q8.py                  # MLX stream generation benchmark
│   └── bench_llamacpp.py                # llama-cli Q8_0 GGUF benchmark
└── results/
    ├── splash_all_variants_5prompts.json
    ├── quality_benchmark_scorecard.json
    ├── mtplx_q8_results.json
    ├── mlx_q8_results.json
    ├── llamacpp_q8_results.json
    └── benchmark_results_q8_q4_mtp.json

How to Run / Serve the Model

To run the native 8-bit model (nitinpanj/Qwen3.8-27B-Splash-HQ) with the high-performance Metal Q8 engine:

# 1. Clone the Q8-enabled Splash fork
git clone https://github.com/npanj/splash.git -b q8
cd splash

# 2. Build the Metal kernels
make -j4

# 3. Serve the model (downloads automatically from Hugging Face on first launch)
./splash serve --model nitinpanj/Qwen3.8-27B-Splash-HQ

The server binds to http://127.0.0.1:8000 with an OpenAI-compatible API (/v1/chat/completions). Connect any coding agent or client:

# Oh My Pi (OMP)
omp --model splash/nitinpanj/Qwen3.8-27B-Splash-HQ

Quickstart & Reproduction

1. Run the 5-Prompt Head-to-Head Comparison

python3 scripts/run_all_models_standard_bench.py

2. Run the Quality Benchmark (GPQA, AIME, MATH-500)

python3 scripts/run_quality_benchmarks.py --limit 3 --max-tokens 800

3. Build a Selective Mixed-Precision Model

python3 scripts/pack_mixed_q8.py \
  --source-safetensors ~/.mtplx/models/Qwen3.8-27B-MTPLX-Optimized-Quality \
  --base-splash-dir install/models/incoai/Qwen3.8-27B-Splash-Q8 \
  --out-dir install/models/incoai/Qwen3.8-27B-Splash-Mixed \
  --upgrade-layers 56,57,58,59,60,61,62,63 \
  --upgrade-embedding \
  --upgrade-head

License

Code and benchmark tools are released under the Apache-2.0 License. Model weights follow the original Qwen License.

Languages

Python

100.0%