This is an MLX conversion of deepseek-ai/DeepSeek-V4-Flash.
deepseek-ai/DeepSeek-V4-Flash6e763230a9d263eca2023f1d4a5ce1bfe126cf48DeepseekV4ForCausalLMdeepseek_v4Thump604/mlx-lm, branch deepseek-v4-support-fixes9c990f4/Volumes/Lexar/mlx_models/DeepSeek-V4-Flash-MLX-Q3-mixed-gs128-affinemixed_3_6affine1283.80828135,346,422,876 bytesThe mixed recipe uses 3-bit affine quantization for lower-risk routed expert paths and 6-bit affine quantization for sensitive paths including embeddings, LM head, attention projections, compressed-attention/indexer components, shared experts, and selected down projections.
DeepSeek V4 support in MLX is still under active development. This artifact was produced with local DeepSeek V4 support fixes, including FP4/FP8 checkpoint handling, F8_E8M0 scale metadata reinterpretation as raw uint8 exponent bytes before sanitizer decode, attention sink dtype handling, and quantized grouped output projection support.
3 commits
This is an MLX conversion of deepseek-ai/DeepSeek-V4-Flash.
deepseek-ai/DeepSeek-V4-Flash6e763230a9d263eca2023f1d4a5ce1bfe126cf48DeepseekV4ForCausalLMdeepseek_v4Thump604/mlx-lm, branch deepseek-v4-support-fixes9c990f4/Volumes/Lexar/mlx_models/DeepSeek-V4-Flash-MLX-Q3-mixed-gs128-affinemixed_3_6affine1283.80828135,346,422,876 bytesThe mixed recipe uses 3-bit affine quantization for lower-risk routed expert paths and 6-bit affine quantization for sensitive paths including embeddings, LM head, attention projections, compressed-attention/indexer components, shared experts, and selected down projections.
DeepSeek V4 support in MLX is still under active development. This artifact was produced with local DeepSeek V4 support fixes, including FP4/FP8 checkpoint handling, F8_E8M0 scale metadata reinterpretation as raw uint8 exponent bytes before sanitizer decode, attention sink dtype handling, and quantized grouped output projection support.
3 commits