Qwen3.8-27B-Splash-Mixed (Selective 8-bit Sensitive Layers)
1
8 commits
1 linked in READMEs
updated Sep 20, 2026
This repository contains Qwen3.8-27B-Splash-Mixed, an experimental hybrid-precision model for the Splash inference engine on Apple Silicon.
Based on mathematical sensitivity analysis (Hessian trace and singular-value decay), the top 8 deepest transformer layers (Layers 56β63), the token embedding table, and the LM head were selectively upgraded to uncompressed native 8-bit precision, while maintaining baseline compression on the earlier 56 layers.
Qwen3.8-27B-Splash-Mixed (Selective 8-bit Sensitive Layers)
1
8 commits
1 linked in READMEs
updated Sep 20, 2026
This repository contains Qwen3.8-27B-Splash-Mixed, an experimental hybrid-precision model for the Splash inference engine on Apple Silicon.
Based on mathematical sensitivity analysis (Hessian trace and singular-value decay), the top 8 deepest transformer layers (Layers 56β63), the token embedding table, and the LM head were selectively upgraded to uncompressed native 8-bit precision, while maintaining baseline compression on the earlier 56 layers.