AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

Model

0

stars

11

commits

1

linked in READMEs

Aug 21, 2026

updated

4-bit
4bit
apple-silicon
axq
axquant
certified
checkpoint-tier1
conversational
decode-heavy-mtp-certified
mixed-precision
mlx
mtp
mtp-acceleration-tier2
quality-validated
quantized
qwen3_5
qwen3.6
safetensors
stable-default-runtime
text-generation
v2
vision
Browse cluster: Quantized LLM Inference on Apple Silicon

README

AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved as BF16 sidecars when present.

AXQuant checkpoint Tier 1 certified on df-macbookpro-m5 for Hub commit f44a9eeebec0c488d0f42201c8763db770a1c0a8. Product class is 5p6bpw (mixed AXQ; marketing name 4-bit). Size vs uniform-4 is 1.207× under the class 1.25 budget; agent-coding quality retention 0.993, general 1.000. Certificate.

MTP acceleration Tier 2 certified (scoped). On df-macbookpro-m5 with AX Engine 6.14.0, greedy MTP-off/on streams match and decode-heavy profiles clear ≥1.20× / ≥1.10× (agent-coding 1.301× / 1.104×, long-form general 1.223× / 1.249×). Tier 2 certificate.

Product default remains direct fallback. Short-answer chat is not a universal speed claim. Formal route: Qwen linear MTP exact + certification-candidate opt-in.

Model details

PropertyValue
Base modelQwen/Qwen3.6-27B
Source revision6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Product familyqwen3.6
Source architectureQwen3_5ForConditionalGeneration (dense); text path optimized
Main-model parameters27.36B logical parameters
QuantizerAXQuant 1.2.0
Hub budget class4bit
Artifact editionv2
AXQuant base precision class5p6bpw
Planned storage-adjusted BPW5.5800
Measured main-model BPW5.4183
Measured total BPW, including MTP5.5801
Safetensors weight size19.38 GB
Approximate complete download19.40 GB
Configured maximum context262,144 tokens; practical limits depend on unified memory
MLX-LM compatibilityStandard text inference, compatibility level B
AX Engine native executionTier 1 safe default direct route; scoped Tier 2 MTP certified (opt-in formal contract)
MTP presentTrue
Vision sidecar presentTrue

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models.

SiblingIntended trade-off
4bit siblingLower-storage AXQ budget; check its exact BPW
6bit siblingHigher average precision near the 6-BPW budget

See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP --local-dir ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

Allow at least 19.40 GB of free disk space. Pin the resulting Hub commit in reproducible deployments rather than relying indefinitely on main.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

AX Engine status

AX Engine 6.14.0 on df-macbookpro-m5 loads this checkpoint for formal certification. Product default remains direct fallback (safe Tier 1). Scoped Tier 2 MTP requires the formal Qwen linear MTP exact / certification-candidate contract (see certificate). MLX-LM remains the standard text inference path and does not by itself establish MTP acceleration.

Use the packaged Qwen MTP head with oMLX or MTPLX

Download the complete repository to a writable local directory. In oMLX 0.6.3rc2 or newer, add that directory, open Model Settings, choose Import MTP side-car, and then enable Lightning MTP. The import changes only the local copy so the sidecar tensors become visible through the checkpoint index. This is a text-path compatibility result; it does not certify VLM loading or vision quality.

MTPLX can consume the packaged sidecar directly:

mtplx quickstart \
  --model ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP \
  --profile stable \
  --depth 1 \
  --reasoning off

mtplx_runtime.json declares the canonical qwen3-next-mtp execution contract. This enables strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed certification.

Quantization layout

Main-weight precisionParametersShare
4bit24.35B87.65%
8bit1.27B4.58%
bf162.16B7.77%
  • Quantization methods: affine, bf16.
  • Group sizes used by quantized assignments: 32, 64.
  • MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
  • Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

CheckStatus
Planning evidencearchitecture_prior (Tier 1 still uses measured quality/size on host)
Quality vs uniform-4Certified: agent-coding retention 0.992647; general 1.0
Size vs uniform-41.206999 under 5p6bpw max ratio 1.25
MTP exactness + speed (decode-heavy)Tier 2 certified (scoped) — see certificate
Vision-language qualityNot claimed; vision tensors preserved at BF16
Full M0–M8 flagship campaignSeparate process; not implied by these certificates
CertificatesTier 1 · Tier 2 · index

Modalities (capability-gated)

Text checkpoint Tier 1 does not imply vision or audio quality. Vision present=true on a pack is not a quality pass.

ModalityClaimSupportedReason
Visionpresent-not-certifiedtruevision present sidecar=['vision.safetensors']; mlx-vlm smoke failed on df-macstudio-m2 (see evidence). Text Tier 1 unchanged. Evidence: /Users/akiralam/code/axquant/docs/certifications/evidence/modality-recert-macstudio-m2/results/qwen36-27b-axq4-mtp.json
Audionot-applicablefalseaudio not supported (no tower config and no sidecar weights)

Intended use and limitations

  • Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.

  • MTP acceleration is opt-in under the formal exact contract; default remains direct fallback.

  • Vision weights are byte-preserved at BF16, but this release does not claim validated VLM quality.

  • The configured context window can require substantially more memory as the KV cache grows.

  • AX Engine Tier 1 default is direct fallback; Tier 2 is scoped formal-route only.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. Parallel OptiQ repositories use a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.6-27B model card for license terms, model limitations, and responsible-use guidance.

Contributors

AutomatosX

11 commits

AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

Model

0

stars

11

commits

1

linked in READMEs

Aug 21, 2026

updated

4-bit
4bit
apple-silicon
axq
axquant
certified
checkpoint-tier1
conversational
decode-heavy-mtp-certified
mixed-precision
mlx
mtp
mtp-acceleration-tier2
quality-validated
quantized
qwen3_5
qwen3.6
safetensors
stable-default-runtime
text-generation
v2
vision
Browse cluster: Quantized LLM Inference on Apple Silicon

README

AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved as BF16 sidecars when present.

AXQuant checkpoint Tier 1 certified on df-macbookpro-m5 for Hub commit f44a9eeebec0c488d0f42201c8763db770a1c0a8. Product class is 5p6bpw (mixed AXQ; marketing name 4-bit). Size vs uniform-4 is 1.207× under the class 1.25 budget; agent-coding quality retention 0.993, general 1.000. Certificate.

MTP acceleration Tier 2 certified (scoped). On df-macbookpro-m5 with AX Engine 6.14.0, greedy MTP-off/on streams match and decode-heavy profiles clear ≥1.20× / ≥1.10× (agent-coding 1.301× / 1.104×, long-form general 1.223× / 1.249×). Tier 2 certificate.

Product default remains direct fallback. Short-answer chat is not a universal speed claim. Formal route: Qwen linear MTP exact + certification-candidate opt-in.

Model details

PropertyValue
Base modelQwen/Qwen3.6-27B
Source revision6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Product familyqwen3.6
Source architectureQwen3_5ForConditionalGeneration (dense); text path optimized
Main-model parameters27.36B logical parameters
QuantizerAXQuant 1.2.0
Hub budget class4bit
Artifact editionv2
AXQuant base precision class5p6bpw
Planned storage-adjusted BPW5.5800
Measured main-model BPW5.4183
Measured total BPW, including MTP5.5801
Safetensors weight size19.38 GB
Approximate complete download19.40 GB
Configured maximum context262,144 tokens; practical limits depend on unified memory
MLX-LM compatibilityStandard text inference, compatibility level B
AX Engine native executionTier 1 safe default direct route; scoped Tier 2 MTP certified (opt-in formal contract)
MTP presentTrue
Vision sidecar presentTrue

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models.

SiblingIntended trade-off
4bit siblingLower-storage AXQ budget; check its exact BPW
6bit siblingHigher average precision near the 6-BPW budget

See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP --local-dir ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

Allow at least 19.40 GB of free disk space. Pin the resulting Hub commit in reproducible deployments rather than relying indefinitely on main.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

AX Engine status

AX Engine 6.14.0 on df-macbookpro-m5 loads this checkpoint for formal certification. Product default remains direct fallback (safe Tier 1). Scoped Tier 2 MTP requires the formal Qwen linear MTP exact / certification-candidate contract (see certificate). MLX-LM remains the standard text inference path and does not by itself establish MTP acceleration.

Use the packaged Qwen MTP head with oMLX or MTPLX

Download the complete repository to a writable local directory. In oMLX 0.6.3rc2 or newer, add that directory, open Model Settings, choose Import MTP side-car, and then enable Lightning MTP. The import changes only the local copy so the sidecar tensors become visible through the checkpoint index. This is a text-path compatibility result; it does not certify VLM loading or vision quality.

MTPLX can consume the packaged sidecar directly:

mtplx quickstart \
  --model ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP \
  --profile stable \
  --depth 1 \
  --reasoning off

mtplx_runtime.json declares the canonical qwen3-next-mtp execution contract. This enables strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed certification.

Quantization layout

Main-weight precisionParametersShare
4bit24.35B87.65%
8bit1.27B4.58%
bf162.16B7.77%
  • Quantization methods: affine, bf16.
  • Group sizes used by quantized assignments: 32, 64.
  • MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
  • Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

CheckStatus
Planning evidencearchitecture_prior (Tier 1 still uses measured quality/size on host)
Quality vs uniform-4Certified: agent-coding retention 0.992647; general 1.0
Size vs uniform-41.206999 under 5p6bpw max ratio 1.25
MTP exactness + speed (decode-heavy)Tier 2 certified (scoped) — see certificate
Vision-language qualityNot claimed; vision tensors preserved at BF16
Full M0–M8 flagship campaignSeparate process; not implied by these certificates
CertificatesTier 1 · Tier 2 · index

Modalities (capability-gated)

Text checkpoint Tier 1 does not imply vision or audio quality. Vision present=true on a pack is not a quality pass.

ModalityClaimSupportedReason
Visionpresent-not-certifiedtruevision present sidecar=['vision.safetensors']; mlx-vlm smoke failed on df-macstudio-m2 (see evidence). Text Tier 1 unchanged. Evidence: /Users/akiralam/code/axquant/docs/certifications/evidence/modality-recert-macstudio-m2/results/qwen36-27b-axq4-mtp.json
Audionot-applicablefalseaudio not supported (no tower config and no sidecar weights)

Intended use and limitations

  • Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.

  • MTP acceleration is opt-in under the formal exact contract; default remains direct fallback.

  • Vision weights are byte-preserved at BF16, but this release does not claim validated VLM quality.

  • The configured context window can require substantially more memory as the KV cache grows.

  • AX Engine Tier 1 default is direct fallback; Tier 2 is scoped formal-route only.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. Parallel OptiQ repositories use a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.6-27B model card for license terms, model limitations, and responsible-use guidance.

Contributors

AutomatosX

11 commits