nitinpanj/Qwen3.8-27B-Splash-Q8

Model

Qwen3.8-27B-Splash-Q8 (Compressed 8-bit Baseline)

1

8 commits

1 linked in READMEs

updated Sep 20, 2026

See the code

README

Qwen3.8-27B-Splash-Q8 (Compressed 8-bit Baseline)

This repository contains the baseline 8-bit compressed weights for Qwen3.8-27B for the Splash inference engine on Apple Silicon.

Performance

  • Decode Speed: 36.5 tok/s average (peaking at 52.7 tok/s).
  • Speedup: 3.69x over standard autoregressive decoding.
  • Reasoning Accuracy: 44.4% across GPQA Diamond and AIME 2025.
8bit
apple-silicon
metal
speculative-decoding
splash
text-generation

nitinpanj/Qwen3.8-27B-Splash-Q8

Model

Qwen3.8-27B-Splash-Q8 (Compressed 8-bit Baseline)

1

8 commits

1 linked in READMEs

updated Sep 20, 2026

See the code

README

Qwen3.8-27B-Splash-Q8 (Compressed 8-bit Baseline)

This repository contains the baseline 8-bit compressed weights for Qwen3.8-27B for the Splash inference engine on Apple Silicon.

Performance

  • Decode Speed: 36.5 tok/s average (peaking at 52.7 tok/s).
  • Speedup: 3.69x over standard autoregressive decoding.
  • Reasoning Accuracy: 44.4% across GPQA Diamond and AIME 2025.
8bit
apple-silicon
metal
speculative-decoding
splash
text-generation