AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP

Model

0

stars

13

commits

1

linked in READMEs

Aug 21, 2026

updated

4-bit
6-bit
6bit
apple-silicon
axq
axquant
certified
checkpoint-tier1
conversational
decode-heavy-mtp-certified
mixed-precision
mlx
mtp
mtp-acceleration-tier2
quality-validated
quantized
qwen3_5
qwen3.6
safetensors
stable-default-runtime
text-generation
v3
vision
Browse cluster: Quantized LLM Inference on Apple Silicon

README

AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved at BF16 in the checkpoint (or a bound sidecar when present).

AXQuant checkpoint Tier 1 certified — artifact edition v3. The exact v3 weight set passed measured size, matched-reference quality, zero-fallback conversion, and safe default-runtime gates on df-macbookpro-m5. Read the certificate and exact hashes.

MTP acceleration Tier 2 certified (scoped). On df-macbookpro-m5 with AX Engine 6.14.0, greedy MTP-off/on streams are identical and decode-heavy authorizing profiles clear ≥1.20× token-weighted and ≥1.10× prompt-median speedup (agent-coding 1.258× / 1.112×, long-form general 1.233× / 1.250×). Tier 2 certificate.

Product default remains direct fallback for the safe Tier 1 route. Short-answer chat is not an authorizing universal speed claim. Use the formal Qwen linear MTP exact / certification-candidate contract to exercise the certified acceleration path.

Model details

PropertyValue
Base modelQwen/Qwen3.6-27B
Source revision6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Product familyqwen3.6
Source architectureQwen3_5ForConditionalGeneration (dense); text path optimized
Main-model parameters27.36B logical parameters
QuantizerAXQuant 1.5.1
Hub budget class6bit
Artifact editionv3
AXQuant base precision class6bit
Planned storage-adjusted BPW5.9616
Measured main-model BPW5.8058
Measured total BPW, including MTP5.9617
Safetensors weight size20.70 GB
Approximate complete download20.73 GB
Configured maximum context262,144 tokens; practical limits depend on unified memory
Primary MLX runtimeMLX-LM
AX Engine native executionNative manifest included; Tier 1 safe default direct route; scoped Tier 2 MTP certified (opt-in formal contract)
MTP presentTrue
Vision presentTrue
Audio presentFalse

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models. When that collapse happens, AutomatosX does not publish a separate misleading 4bit sibling for that base.

SiblingIntended trade-off
4bit siblingLower-storage AXQ budget; check its exact BPW
6bit siblingHigher average precision near the 6-BPW budget

See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP \
  --revision v3 \
  --local-dir ./AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP

Allow at least 20.73 GB of free disk space. The v3 tag identifies this certified artifact; pin the resolved Hub commit for production reproducibility. The previous development checkpoint remains available at v2.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

Serve with AX Engine

After installing AX Engine, download the complete repository (see AXQuant for conversion, certificates, and model-card tooling) and serve the local directory:

ax-engine serve ./AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP --port 31418

AX Engine is the authority for the AXQ runtime contract and native MTP sidecar. Tier 1 default-route smoke (AX Engine 6.13.5) keeps policy 4, MTP inactive, and direct fallback active. Scoped Tier 2 MTP acceleration is certified on AX Engine 6.14.0 under the formal Qwen linear MTP exact / certification-candidate contract (decode-heavy suites only). Product default remains direct fallback. Native model-manifest.json status: included.

Use the packaged Qwen MTP head with oMLX or MTPLX

Download the complete repository to a writable local directory. In oMLX 0.6.3rc2 or newer, add that directory, open Model Settings, choose Import MTP side-car, and then enable Lightning MTP. The import changes only the local copy so the sidecar tensors become visible through the checkpoint index. This is a text-path compatibility result; it does not certify VLM loading or vision quality.

MTPLX can consume the packaged sidecar directly:

mtplx quickstart \
  --model ./AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP \
  --profile stable \
  --depth 1 \
  --reasoning off

mtplx_runtime.json declares the canonical qwen3-next-mtp execution contract. This enables strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed certification.

Quantization layout

Main-weight precisionParametersShare
4bit20.78B74.79%
6bit3.20B11.52%
8bit1.27B4.58%
bf162.53B9.11%
  • Quantization methods: affine, bf16, dwq.
  • Group sizes used by quantized assignments: 64.
  • MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
  • Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
  • Vision weights: protected BF16 sidecar.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

CheckStatus
Planning evidencemeasured
Calibrationrecorded in measured-forward-probe-refinement
Quantizer execution487/487 recorded module conversions succeeded; 0 fallbacks
AX Engine native manifestincluded as model-manifest.json
General quality vs matched uniform-644 tasks; retention 1.011494; perplexity ratio 0.966136; 0/0 errors
Agent-coding quality vs matched uniform-676 tasks; retention 1.007353; perplexity ratio 0.969531; 0/0 errors
Size vs matched uniform-6ratio 0.876234; 12.38% fewer weight bytes
Stable default AX Engine routePass; MTP inactive and direct fallback active
MTP acceptance, exactness, and speedScoped Tier 2 certified on df-macbookpro-m5 / AX Engine 6.14.0: greedy exactness + agent-coding 1.258× / 1.112×, long-form general 1.233× / 1.250× (Tier 2 cert); product default remains direct fallback; short-answer not claimed
AX Engine kernel-speed evidenceNot part of Tier 1
Vision-language qualityNot evaluated or claimed; vision tensors are preserved at BF16
Speech-recognition qualityNot applicable
Long-context quality262,144-token capacity is config metadata, not a validated claim
Release certificationCheckpoint Tier 1 certified + scoped MTP Tier 2 certified (T1 · T2)

Modalities (capability-gated)

Text checkpoint Tier 1 does not imply vision or audio quality. Vision present=true on a pack is not a quality pass.

ModalityClaimSupportedReason
Visionpresent-not-certifiedtruevision present sidecar=['vision.safetensors']; mlx-vlm smoke failed on df-macstudio-m2 (mlx-vlm expects vision_tower.*; sidecar/layout mismatch). Text Tier 1 unchanged. Evidence: docs/certifications/evidence/modality-recert-capability-gated/results/AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP.json
Audionot-applicablefalseaudio not supported (no tower config and no sidecar weights)

Intended use and limitations

  • Intended for local text generation and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Quality evidence is relative to the matched uniform-6 reference on reproducible AXQuant suites; it is not a third-party benchmark, a universal quality guarantee, or a BF16-equivalence claim.

  • “Stable default runtime” means the bound text path passed model loading, inference, and the fail-closed route gate on the certification machine. It is not a promise that every context, application, or third-party runtime is defect-free.

  • MTP acceleration is opt-in under the formal Qwen linear MTP exact / certification-candidate contract (scoped Tier 2 on decode-heavy suites). Product default remains direct fallback. Short-answer chat is not a universal speed claim. Outside AX Engine, use a sidecar-aware runtime; oMLX/MTPLX compatibility does not extend the AX Engine certificate.

  • Vision weights are preserved at BF16, but this release does not claim validated VLM quality.

  • The configured context window can require substantially more memory as the KV cache grows.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. If an OptiQ repository is published separately, it uses a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.6-27B model card for license terms, model limitations, and responsible-use guidance.

Contributors

AutomatosX

13 commits

AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP

Model

0

stars

13

commits

1

linked in READMEs

Aug 21, 2026

updated

4-bit
6-bit
6bit
apple-silicon
axq
axquant
certified
checkpoint-tier1
conversational
decode-heavy-mtp-certified
mixed-precision
mlx
mtp
mtp-acceleration-tier2
quality-validated
quantized
qwen3_5
qwen3.6
safetensors
stable-default-runtime
text-generation
v3
vision
Browse cluster: Quantized LLM Inference on Apple Silicon

README

AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved at BF16 in the checkpoint (or a bound sidecar when present).

AXQuant checkpoint Tier 1 certified — artifact edition v3. The exact v3 weight set passed measured size, matched-reference quality, zero-fallback conversion, and safe default-runtime gates on df-macbookpro-m5. Read the certificate and exact hashes.

MTP acceleration Tier 2 certified (scoped). On df-macbookpro-m5 with AX Engine 6.14.0, greedy MTP-off/on streams are identical and decode-heavy authorizing profiles clear ≥1.20× token-weighted and ≥1.10× prompt-median speedup (agent-coding 1.258× / 1.112×, long-form general 1.233× / 1.250×). Tier 2 certificate.

Product default remains direct fallback for the safe Tier 1 route. Short-answer chat is not an authorizing universal speed claim. Use the formal Qwen linear MTP exact / certification-candidate contract to exercise the certified acceleration path.

Model details

PropertyValue
Base modelQwen/Qwen3.6-27B
Source revision6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Product familyqwen3.6
Source architectureQwen3_5ForConditionalGeneration (dense); text path optimized
Main-model parameters27.36B logical parameters
QuantizerAXQuant 1.5.1
Hub budget class6bit
Artifact editionv3
AXQuant base precision class6bit
Planned storage-adjusted BPW5.9616
Measured main-model BPW5.8058
Measured total BPW, including MTP5.9617
Safetensors weight size20.70 GB
Approximate complete download20.73 GB
Configured maximum context262,144 tokens; practical limits depend on unified memory
Primary MLX runtimeMLX-LM
AX Engine native executionNative manifest included; Tier 1 safe default direct route; scoped Tier 2 MTP certified (opt-in formal contract)
MTP presentTrue
Vision presentTrue
Audio presentFalse

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models. When that collapse happens, AutomatosX does not publish a separate misleading 4bit sibling for that base.

SiblingIntended trade-off
4bit siblingLower-storage AXQ budget; check its exact BPW
6bit siblingHigher average precision near the 6-BPW budget

See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP \
  --revision v3 \
  --local-dir ./AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP

Allow at least 20.73 GB of free disk space. The v3 tag identifies this certified artifact; pin the resolved Hub commit for production reproducibility. The previous development checkpoint remains available at v2.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

Serve with AX Engine

After installing AX Engine, download the complete repository (see AXQuant for conversion, certificates, and model-card tooling) and serve the local directory:

ax-engine serve ./AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP --port 31418

AX Engine is the authority for the AXQ runtime contract and native MTP sidecar. Tier 1 default-route smoke (AX Engine 6.13.5) keeps policy 4, MTP inactive, and direct fallback active. Scoped Tier 2 MTP acceleration is certified on AX Engine 6.14.0 under the formal Qwen linear MTP exact / certification-candidate contract (decode-heavy suites only). Product default remains direct fallback. Native model-manifest.json status: included.

Use the packaged Qwen MTP head with oMLX or MTPLX

Download the complete repository to a writable local directory. In oMLX 0.6.3rc2 or newer, add that directory, open Model Settings, choose Import MTP side-car, and then enable Lightning MTP. The import changes only the local copy so the sidecar tensors become visible through the checkpoint index. This is a text-path compatibility result; it does not certify VLM loading or vision quality.

MTPLX can consume the packaged sidecar directly:

mtplx quickstart \
  --model ./AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP \
  --profile stable \
  --depth 1 \
  --reasoning off

mtplx_runtime.json declares the canonical qwen3-next-mtp execution contract. This enables strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed certification.

Quantization layout

Main-weight precisionParametersShare
4bit20.78B74.79%
6bit3.20B11.52%
8bit1.27B4.58%
bf162.53B9.11%
  • Quantization methods: affine, bf16, dwq.
  • Group sizes used by quantized assignments: 64.
  • MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
  • Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
  • Vision weights: protected BF16 sidecar.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

CheckStatus
Planning evidencemeasured
Calibrationrecorded in measured-forward-probe-refinement
Quantizer execution487/487 recorded module conversions succeeded; 0 fallbacks
AX Engine native manifestincluded as model-manifest.json
General quality vs matched uniform-644 tasks; retention 1.011494; perplexity ratio 0.966136; 0/0 errors
Agent-coding quality vs matched uniform-676 tasks; retention 1.007353; perplexity ratio 0.969531; 0/0 errors
Size vs matched uniform-6ratio 0.876234; 12.38% fewer weight bytes
Stable default AX Engine routePass; MTP inactive and direct fallback active
MTP acceptance, exactness, and speedScoped Tier 2 certified on df-macbookpro-m5 / AX Engine 6.14.0: greedy exactness + agent-coding 1.258× / 1.112×, long-form general 1.233× / 1.250× (Tier 2 cert); product default remains direct fallback; short-answer not claimed
AX Engine kernel-speed evidenceNot part of Tier 1
Vision-language qualityNot evaluated or claimed; vision tensors are preserved at BF16
Speech-recognition qualityNot applicable
Long-context quality262,144-token capacity is config metadata, not a validated claim
Release certificationCheckpoint Tier 1 certified + scoped MTP Tier 2 certified (T1 · T2)

Modalities (capability-gated)

Text checkpoint Tier 1 does not imply vision or audio quality. Vision present=true on a pack is not a quality pass.

ModalityClaimSupportedReason
Visionpresent-not-certifiedtruevision present sidecar=['vision.safetensors']; mlx-vlm smoke failed on df-macstudio-m2 (mlx-vlm expects vision_tower.*; sidecar/layout mismatch). Text Tier 1 unchanged. Evidence: docs/certifications/evidence/modality-recert-capability-gated/results/AX-Qwen3.6-27B-MLX-AXQ-6bit-MTP.json
Audionot-applicablefalseaudio not supported (no tower config and no sidecar weights)

Intended use and limitations

  • Intended for local text generation and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Quality evidence is relative to the matched uniform-6 reference on reproducible AXQuant suites; it is not a third-party benchmark, a universal quality guarantee, or a BF16-equivalence claim.

  • “Stable default runtime” means the bound text path passed model loading, inference, and the fail-closed route gate on the certification machine. It is not a promise that every context, application, or third-party runtime is defect-free.

  • MTP acceleration is opt-in under the formal Qwen linear MTP exact / certification-candidate contract (scoped Tier 2 on decode-heavy suites). Product default remains direct fallback. Short-answer chat is not a universal speed claim. Outside AX Engine, use a sidecar-aware runtime; oMLX/MTPLX compatibility does not extend the AX Engine certificate.

  • Vision weights are preserved at BF16, but this release does not claim validated VLM quality.

  • The configured context window can require substantially more memory as the KV cache grows.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. If an OptiQ repository is published separately, it uses a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.6-27B model card for license terms, model limitations, and responsible-use guidance.

Contributors

AutomatosX

13 commits