AX Gemma 4 26B A4B IT MLX QAT 4-bit Assistant MTP
0
7 commits
1 linked in READMEs
updated Jul 23, 2026
Parameter count: approximately 25.81B logical target parameters (the 26B-A4B model class), with approximately 4B active per token.
4-bitis the target quantization precision, not a 4B model-size claim. The separately packaged assistant is not included in the target count.
This is a self-contained MLX model package for Apple Silicon. It combines the quantization-aware-trained Gemma 4 26B-A4B instruction MoE target with its exact paired 4-bit assistant for AX Engine multi-token prediction (MTP) / speculative decoding.
The target verifies every drafted token. The assistant improves decode speed without replacing the target model. This repository does not contain PyTorch, GGUF, or the unquantized Google QAT weights.
0e3cbab38ce568cf6e23543010d08d03b731910cbb94eae1b70a80dac16cbf959bb4b7d56bd1fb8cQAT means that the upstream checkpoint was optimized during training for its target quantization scheme before the MLX conversion. It is distinct from a post-training-only 4-bit conversion.
hf download AutomatosX/AX-Gemma-4-26B-A4B-IT-MLX-QAT-4bit-Assistant-MTP \
--local-dir ./AX-Gemma-4-26B-A4B-IT-MLX-QAT-4bit-Assistant-MTP
The repository is approximately 15.9 GB.
Install AX Engine, then serve the downloaded repository directory:
ax-engine serve ./AX-Gemma-4-26B-A4B-IT-MLX-QAT-4bit-Assistant-MTP --port 31418
Gemma assistant MTP is enabled by default. AX Engine reads
ax_gemma4_assistant_mtp.json, loads the drafter from assistant/, and uses
the 26B-A4B target to verify proposals. Do not load assistant/ by itself as a
general-purpose chat model.
The root target weights can also be loaded for direct MLX generation, but the nested assistant pairing and acceleration are AX Engine-specific.
Validated on macOS arm64 with AX Engine 6.9.0 on 2026-07-20:
ready, with no issuesexactSee ax_provenance.json for immutable source revisions and SHA-256 values.
Apache License 2.0. Review the Gemma 4 license and the official Google model cards for usage limitations and responsible-use guidance.
7 commits
AX Gemma 4 26B A4B IT MLX QAT 4-bit Assistant MTP
0
7 commits
1 linked in READMEs
updated Jul 23, 2026
Parameter count: approximately 25.81B logical target parameters (the 26B-A4B model class), with approximately 4B active per token.
4-bitis the target quantization precision, not a 4B model-size claim. The separately packaged assistant is not included in the target count.
This is a self-contained MLX model package for Apple Silicon. It combines the quantization-aware-trained Gemma 4 26B-A4B instruction MoE target with its exact paired 4-bit assistant for AX Engine multi-token prediction (MTP) / speculative decoding.
The target verifies every drafted token. The assistant improves decode speed without replacing the target model. This repository does not contain PyTorch, GGUF, or the unquantized Google QAT weights.
0e3cbab38ce568cf6e23543010d08d03b731910cbb94eae1b70a80dac16cbf959bb4b7d56bd1fb8cQAT means that the upstream checkpoint was optimized during training for its target quantization scheme before the MLX conversion. It is distinct from a post-training-only 4-bit conversion.
hf download AutomatosX/AX-Gemma-4-26B-A4B-IT-MLX-QAT-4bit-Assistant-MTP \
--local-dir ./AX-Gemma-4-26B-A4B-IT-MLX-QAT-4bit-Assistant-MTP
The repository is approximately 15.9 GB.
Install AX Engine, then serve the downloaded repository directory:
ax-engine serve ./AX-Gemma-4-26B-A4B-IT-MLX-QAT-4bit-Assistant-MTP --port 31418
Gemma assistant MTP is enabled by default. AX Engine reads
ax_gemma4_assistant_mtp.json, loads the drafter from assistant/, and uses
the 26B-A4B target to verify proposals. Do not load assistant/ by itself as a
general-purpose chat model.
The root target weights can also be loaded for direct MLX generation, but the nested assistant pairing and acceleration are AX Engine-specific.
Validated on macOS arm64 with AX Engine 6.9.0 on 2026-07-20:
ready, with no issuesexactSee ax_provenance.json for immutable source revisions and SHA-256 values.
Apache License 2.0. Review the Gemma 4 license and the official Google model cards for usage limitations and responsible-use guidance.
7 commits