AutomatosX/AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP

Model

0

stars

12

commits

1

linked in READMEs

Aug 21, 2026

updated

4-bit
4bit
apple-silicon
axq
axquant
certified
checkpoint-tier1
conversational
mixed-precision
mlx
mtp
quality-validated
quantized
qwen3_5_moe
qwen3.6
safetensors
stable-default-runtime
text-generation
v2
vision
Browse cluster: Quantized LLM Inference on Apple Silicon

README

AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved as BF16 sidecars when present.

AXQuant checkpoint Tier 1 certified on df-macbookpro-m5 for Hub commit a549387d5b812c6f6cbdb0ebde37adb3b3f4a2bc. Agent-coding and general quality retention both 1.000 vs mlx-community/Qwen3.6-35B-A3B-4bit. Size vs uniform-4 is 1.132× under a class budget of 1.15. Certificate.

MTP acceleration Tier 2 is not certified. Formal greedy exactness is achievable with MoE MTP loaded, but decode-heavy speed gates (≥1.20× weighted / ≥1.10× prompt-median) are not met (agent ~0.95× / 0.80×; long ~0.91× / 0.93×). Product default remains direct fallback. Certification index.

Model details

PropertyValue
Base modelQwen/Qwen3.6-35B-A3B
Source revision995ad96eacd98c81ed38be0c5b274b04031597b0
Product familyqwen3.6
Source architectureQwen3_5MoeForConditionalGeneration (mixture of experts (MoE)); text path optimized
Main-model parameters35.11B logical parameters
QuantizerAXQuant 1.2.0
Hub budget class4bit
Artifact editionv2
AXQuant base precision class4bit
Planned storage-adjusted BPW5.1400
Measured main-model BPW4.8788
Measured total BPW, including MTP5.1401
Safetensors weight size23.10 GB
Approximate complete download23.12 GB
Configured maximum context262,144 tokens; practical limits depend on unified memory
MLX-LM compatibilityStandard text inference, compatibility level B
AX Engine native executionNot established; no validated native manifest is included
MTP presentTrue
Vision sidecar presentTrue

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models.

SiblingIntended trade-off
4bit siblingLower-storage AXQ budget; check its exact BPW
6bit siblingHigher average precision near the 6-BPW budget

See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP --local-dir ./AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP

Allow at least 23.12 GB of free disk space. Pin the resulting Hub commit in reproducible deployments rather than relying indefinitely on main.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

AX Engine status

AX Engine on df-macbookpro-m5 loads this checkpoint for Tier 1 certification. Product default remains direct fallback. Do not claim MTP acceleration for this MoE pack until Tier 2 speed gates pass. MoE MTP load requires an engine that accepts mlp.experts.gate_up_proj packing (see ax-engine PR for the load fix).

Use the packaged Qwen MTP head with oMLX or MTPLX

Download the complete repository to a writable local directory. In oMLX 0.6.3rc2 or newer, add that directory, open Model Settings, choose Import MTP side-car, and then enable Lightning MTP. The import changes only the local copy so the sidecar tensors become visible through the checkpoint index. This is a text-path compatibility result; it does not certify VLM loading or vision quality.

MTPLX can consume the packaged sidecar directly:

mtplx quickstart \
  --model ./AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP \
  --profile stable \
  --depth 1 \
  --reasoning off

mtplx_runtime.json declares the canonical qwen3-next-mtp execution contract. This enables strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed certification.

Quantization layout

Main-weight precisionParametersShare
4bit33.62B93.52%
8bit529.61M1.47%
bf161.80B5.01%
  • Quantization methods: affine, bf16.
  • Group sizes used by quantized assignments: 32, 64.
  • MTP sidecar: 19 tensors, 844.64M parameters, 1.69 GB, BF16.
  • Vision sidecar: 333 tensors, 446.57M parameters, 0.89 GB, BF16.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

CheckStatus
Planning evidencearchitecture_prior
Calibrationnone; the allocation is based on architecture priors
Quantizer execution471/471 recorded module conversions succeeded; 0 fallbacks
AX Engine native manifestnot included
Quality versus matched uniform baselineCertified on host — see certificate
MTP acceptance and speedExactness ok in formal A/B; speed not certified (agent ~0.95× / 0.80×; long ~0.91× / 0.93×)
AX Engine kernel evidenceunmeasured
Vision-language qualityNot evaluated or claimed; vision tensors are preserved at BF16
Long-context quality262,144-token capacity is config metadata, not a validated claim
Release certificationCheckpoint Tier 1 certified; MTP Tier 2 not certified

Modalities (capability-gated)

Text checkpoint Tier 1 does not imply vision or audio quality. Vision present=true on a pack is not a quality pass.

ModalityClaimSupportedReason
Visionpresent-not-certifiedtruevision present sidecar=['vision.safetensors']; mlx-vlm smoke failed on df-macstudio-m2 (mlx-vlm expects vision_tower.*; sidecar/layout mismatch). Text Tier 1 unchanged. Evidence: docs/certifications/evidence/modality-recert-capability-gated/results/AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP.json
Audionot-applicablefalseaudio not supported (no tower config and no sidecar weights)

Intended use and limitations

  • Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.

  • MTP requires a sidecar-aware runtime. oMLX/MTPLX discovery compatibility does not establish exactness or speed certification for those runtimes.

  • Vision weights are byte-preserved at BF16, but this release does not claim validated VLM quality.

  • The configured context window can require substantially more memory as the KV cache grows.

  • AX Engine execution is not established because this package has no validated native manifest.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. Parallel OptiQ repositories use a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.6-35B-A3B model card for license terms, model limitations, and responsible-use guidance.

Contributors

AutomatosX

12 commits

AutomatosX/AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP

Model

0

stars

12

commits

1

linked in READMEs

Aug 21, 2026

updated

4-bit
4bit
apple-silicon
axq
axquant
certified
checkpoint-tier1
conversational
mixed-precision
mlx
mtp
quality-validated
quantized
qwen3_5_moe
qwen3.6
safetensors
stable-default-runtime
text-generation
v2
vision
Browse cluster: Quantized LLM Inference on Apple Silicon

README

AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved as BF16 sidecars when present.

AXQuant checkpoint Tier 1 certified on df-macbookpro-m5 for Hub commit a549387d5b812c6f6cbdb0ebde37adb3b3f4a2bc. Agent-coding and general quality retention both 1.000 vs mlx-community/Qwen3.6-35B-A3B-4bit. Size vs uniform-4 is 1.132× under a class budget of 1.15. Certificate.

MTP acceleration Tier 2 is not certified. Formal greedy exactness is achievable with MoE MTP loaded, but decode-heavy speed gates (≥1.20× weighted / ≥1.10× prompt-median) are not met (agent ~0.95× / 0.80×; long ~0.91× / 0.93×). Product default remains direct fallback. Certification index.

Model details

PropertyValue
Base modelQwen/Qwen3.6-35B-A3B
Source revision995ad96eacd98c81ed38be0c5b274b04031597b0
Product familyqwen3.6
Source architectureQwen3_5MoeForConditionalGeneration (mixture of experts (MoE)); text path optimized
Main-model parameters35.11B logical parameters
QuantizerAXQuant 1.2.0
Hub budget class4bit
Artifact editionv2
AXQuant base precision class4bit
Planned storage-adjusted BPW5.1400
Measured main-model BPW4.8788
Measured total BPW, including MTP5.1401
Safetensors weight size23.10 GB
Approximate complete download23.12 GB
Configured maximum context262,144 tokens; practical limits depend on unified memory
MLX-LM compatibilityStandard text inference, compatibility level B
AX Engine native executionNot established; no validated native manifest is included
MTP presentTrue
Vision sidecar presentTrue

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models.

SiblingIntended trade-off
4bit siblingLower-storage AXQ budget; check its exact BPW
6bit siblingHigher average precision near the 6-BPW budget

See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP --local-dir ./AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP

Allow at least 23.12 GB of free disk space. Pin the resulting Hub commit in reproducible deployments rather than relying indefinitely on main.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

AX Engine status

AX Engine on df-macbookpro-m5 loads this checkpoint for Tier 1 certification. Product default remains direct fallback. Do not claim MTP acceleration for this MoE pack until Tier 2 speed gates pass. MoE MTP load requires an engine that accepts mlp.experts.gate_up_proj packing (see ax-engine PR for the load fix).

Use the packaged Qwen MTP head with oMLX or MTPLX

Download the complete repository to a writable local directory. In oMLX 0.6.3rc2 or newer, add that directory, open Model Settings, choose Import MTP side-car, and then enable Lightning MTP. The import changes only the local copy so the sidecar tensors become visible through the checkpoint index. This is a text-path compatibility result; it does not certify VLM loading or vision quality.

MTPLX can consume the packaged sidecar directly:

mtplx quickstart \
  --model ./AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP \
  --profile stable \
  --depth 1 \
  --reasoning off

mtplx_runtime.json declares the canonical qwen3-next-mtp execution contract. This enables strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed certification.

Quantization layout

Main-weight precisionParametersShare
4bit33.62B93.52%
8bit529.61M1.47%
bf161.80B5.01%
  • Quantization methods: affine, bf16.
  • Group sizes used by quantized assignments: 32, 64.
  • MTP sidecar: 19 tensors, 844.64M parameters, 1.69 GB, BF16.
  • Vision sidecar: 333 tensors, 446.57M parameters, 0.89 GB, BF16.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

CheckStatus
Planning evidencearchitecture_prior
Calibrationnone; the allocation is based on architecture priors
Quantizer execution471/471 recorded module conversions succeeded; 0 fallbacks
AX Engine native manifestnot included
Quality versus matched uniform baselineCertified on host — see certificate
MTP acceptance and speedExactness ok in formal A/B; speed not certified (agent ~0.95× / 0.80×; long ~0.91× / 0.93×)
AX Engine kernel evidenceunmeasured
Vision-language qualityNot evaluated or claimed; vision tensors are preserved at BF16
Long-context quality262,144-token capacity is config metadata, not a validated claim
Release certificationCheckpoint Tier 1 certified; MTP Tier 2 not certified

Modalities (capability-gated)

Text checkpoint Tier 1 does not imply vision or audio quality. Vision present=true on a pack is not a quality pass.

ModalityClaimSupportedReason
Visionpresent-not-certifiedtruevision present sidecar=['vision.safetensors']; mlx-vlm smoke failed on df-macstudio-m2 (mlx-vlm expects vision_tower.*; sidecar/layout mismatch). Text Tier 1 unchanged. Evidence: docs/certifications/evidence/modality-recert-capability-gated/results/AX-Qwen3.6-35B-A3B-MLX-AXQ-4bit-MTP.json
Audionot-applicablefalseaudio not supported (no tower config and no sidecar weights)

Intended use and limitations

  • Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.

  • MTP requires a sidecar-aware runtime. oMLX/MTPLX discovery compatibility does not establish exactness or speed certification for those runtimes.

  • Vision weights are byte-preserved at BF16, but this release does not claim validated VLM quality.

  • The configured context window can require substantially more memory as the KV cache grows.

  • AX Engine execution is not established because this package has no validated native manifest.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. Parallel OptiQ repositories use a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.6-35B-A3B model card for license terms, model limitations, and responsible-use guidance.

Contributors

AutomatosX

12 commits