AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP

Model

10

stars

11

commits

1

linked in READMEs

Aug 21, 2026

updated

4-bit
6-bit
6bit
apple-silicon
axq
axquant
conversational
development
mixed-precision
mlx
mtp
quantized
qwen3_5
qwen3.8
safetensors
text-generation
vision
Browse cluster: Quantized LLM Inference on Apple Silicon

README

AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved at BF16 in the checkpoint (or a bound sidecar when present).

Checkpoint Tier 1 certified on df-macbookpro-m3 (2026-08-14) at Hub commit a5a0b700ea7c — measured size against a matched uniform baseline, quality retention, and conversion integrity. Current main preserves that revision's exact Safetensors payloads while allowing metadata-only compatibility fixes. Tier 1 is a checkpoint claim, not a general speed claim: MTP acceleration is certified for the certificate's authorizing profiles only; outside that scope there is no speedup claim. See the checkpoint Tier 1 certificate and Tier 2 MTP acceleration certificate for the bound evidence and thresholds.

Model details

PropertyValue
Base modelQwen/Qwen3.8-27B
Source revision1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
Product familyqwen3.8
Source architectureQwen3_5ForConditionalGeneration (dense); text path optimized
Main-model parameters27.36B logical parameters
QuantizerAXQuant 1.6.2
Hub budget class6bit
AXQuant base precision class6bit
Planned storage-adjusted BPW6.0000
Measured main-model BPW5.8448
Measured total BPW, including MTP6.0001
Safetensors weight size20.84 GB
Approximate complete download20.86 GB
Configured maximum context262,144 tokens; practical limits depend on unified memory
Primary MLX runtimeMLX-LM
AX Engine native executionNative manifest included; execution still requires a runtime check
MTP presentTrue
Vision presentTrue
Audio presentFalse

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models. When that collapse happens, AutomatosX does not publish a separate misleading 4bit sibling for that base.

SiblingIntended trade-off
4bit siblingLower-storage AXQ budget; check its exact BPW
6bit siblingHigher average precision near the 6-BPW budget

See the AutomatosX collections for the family catalog, or the complete index.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP --local-dir ./AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP

Allow at least 20.86 GB of free disk space. Pin the resulting Hub commit in reproducible deployments rather than relying indefinitely on main.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

Serve with AX Engine and MTP

After installing AX Engine, download the complete repository (see AXQuant for conversion, certificates, and model-card tooling) and serve the local directory:

ax-engine serve ./AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP --port 31418

AX Engine is the authority for the AXQ runtime contract and native MTP sidecar. Runtime speedup claims are limited to the authorizing profiles in the linked Tier 2 certificate. The artifact records AX Engine version 6.16.1. Native model-manifest.json status: included as model-manifest.json.

Use with oMLX Lightning MTP

Download the complete repository into a writable local directory and add that directory to oMLX. With an oMLX build that includes the sidecar importer, select Import MTP side-car in Model Settings, then enable Lightning MTP. The bundled mtplx_runtime.json declares the qwen3-next-mtp execution contract required by oMLX's strict importer; this identifier describes the MTP head layout, not the base model's qwen3_5 model type.

The one-time import adds the sidecar tensors to the local checkpoint index. oMLX compatibility does not extend AXQuant's Tier 2 speed claim beyond the certificate's authorizing profiles; benchmark the result in the intended oMLX workload. At the time of the 2026-08-20 smoke test, direct sidecar import was available on oMLX main (0.6.3rc2); stable oMLX 0.5.7 successfully served the already imported text checkpoint with Lightning MTP. Its VLM loader rejected 333 vision parameters and fell back to the LLM path, so this is a text-only compatibility result.

Use with MTPLX

MTPLX 2.5.2 recognizes the bundled depth-1 sidecar and can serve it with the stable profile:

mtplx quickstart \
  --model ./AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP \
  --profile stable \
  --depth 1 \
  --reasoning off

The contract is intentionally marked unverified with a public-release blocker because no MTPLX Forge exactness baseline is bound to this development artifact. It permits compatibility testing without --unsafe-force-unverified; it does not claim MTPLX exactness certification.

Runtime compatibility speed smoke (not certification)

Measured on an Apple M5 Max with 128 GB unified memory and AC power. Each number is the median of three measured batch-1 requests after one warmup, with temperature 0, thinking off, and MTP depth

  1. Prefill prompts use a unique salt; decode uses the same fixed prompt and runs to 256 tokens.
RuntimePrefill inputPrefillDecode input / outputDecode
oMLX 0.5.72,158 uncached tokens838.9 tok/s155 / 256 tokens43.46 tok/s
AX Engine 7.1.52,122 uncached of 2,154 tokens802.5 tok/s148 / 256 tokens42.56 tok/s
MTPLX 2.5.22,157 uncached tokens738.5 tok/s155 / 256 tokens42.89 tok/s

Prefill throughput is uncached prompt tokens divided by local loopback time to first content, so it includes HTTP, scheduling, and first-token overhead. Decode throughput is (completion tokens - 1) divided by the first-to-last-content interval. The three decode results are within about 2.1%; do not generalize this short smoke to concurrency, long context, vision, quality, or another host.

Metadata compatibility update (2026-08-20): mtplx_runtime.json now includes the complete Qwen MTP runtime identity, tensor count, MTPLX compatibility version, recommended profile, and explicit unverified-exactness status. Safetensors payloads are unchanged.

Quantization layout

Main-weight precisionParametersShare
4bit24.34B87.60%
6bit15.16M0.05%
8bit1.27B4.58%
bf162.16B7.77%
  • Quantization methods: affine, bf16.
  • Group sizes used by quantized assignments: 32, 64.
  • MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
  • Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
  • Vision weights: protected BF16 sidecar.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

CheckStatus
Planning evidencearchitecture_prior
Calibrationnone; the allocation is based on architecture priors
Quantizer execution497/497 recorded module conversions succeeded; 0 fallbacks
AX Engine native manifestincluded as model-manifest.json
Quality versus BF16 or uniform baselinesNot published; no quality-retention claim
MTP acceptance and speedcertified for the certificate's authorizing profiles only; outside that scope there is no speedup claim
AX Engine kernel evidenceunmeasured
Vision-language qualityNot evaluated or claimed; vision tensors are preserved at BF16
Speech-recognition qualityNot applicable
Long-context quality262,144-token capacity is config metadata, not a validated claim
Release certificationCheckpoint Tier 1 certified on df-macbookpro-m3 (2026-08-14), Hub commit a5a0b700ea7c; the formal AXQuant M0-M8 release campaign is a separate process and is not implied
Studio recert (df-macstudio-m2, 2026-08-15)Not certified as a replacement T1. Candidate means on prepare-suite v2: agent-coding 0.933 (52), general 0.875 (16). BF16 is not on Ext12T so 0.98 retention was not computed. Historical record is unchanged. See studio evaluation.

Intended use and limitations

  • Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.

  • MTP requires a sidecar-aware runtime. oMLX/MTPLX discovery compatibility does not extend the linked AX Engine certificate to those runtimes.

  • Vision weights are preserved at BF16, but this release does not claim validated VLM quality.

  • The configured context window can require substantially more memory as the KV cache grows.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. If an OptiQ repository is published separately, it uses a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.8-27B model card for license terms, model limitations, and responsible-use guidance.

Contributors

AutomatosX

11 commits

AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP

Model

10

stars

11

commits

1

linked in READMEs

Aug 21, 2026

updated

4-bit
6-bit
6bit
apple-silicon
axq
axquant
conversational
development
mixed-precision
mlx
mtp
quantized
qwen3_5
qwen3.8
safetensors
text-generation
vision
Browse cluster: Quantized LLM Inference on Apple Silicon

README

AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved at BF16 in the checkpoint (or a bound sidecar when present).

Checkpoint Tier 1 certified on df-macbookpro-m3 (2026-08-14) at Hub commit a5a0b700ea7c — measured size against a matched uniform baseline, quality retention, and conversion integrity. Current main preserves that revision's exact Safetensors payloads while allowing metadata-only compatibility fixes. Tier 1 is a checkpoint claim, not a general speed claim: MTP acceleration is certified for the certificate's authorizing profiles only; outside that scope there is no speedup claim. See the checkpoint Tier 1 certificate and Tier 2 MTP acceleration certificate for the bound evidence and thresholds.

Model details

PropertyValue
Base modelQwen/Qwen3.8-27B
Source revision1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
Product familyqwen3.8
Source architectureQwen3_5ForConditionalGeneration (dense); text path optimized
Main-model parameters27.36B logical parameters
QuantizerAXQuant 1.6.2
Hub budget class6bit
AXQuant base precision class6bit
Planned storage-adjusted BPW6.0000
Measured main-model BPW5.8448
Measured total BPW, including MTP6.0001
Safetensors weight size20.84 GB
Approximate complete download20.86 GB
Configured maximum context262,144 tokens; practical limits depend on unified memory
Primary MLX runtimeMLX-LM
AX Engine native executionNative manifest included; execution still requires a runtime check
MTP presentTrue
Vision presentTrue
Audio presentFalse

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models. When that collapse happens, AutomatosX does not publish a separate misleading 4bit sibling for that base.

SiblingIntended trade-off
4bit siblingLower-storage AXQ budget; check its exact BPW
6bit siblingHigher average precision near the 6-BPW budget

See the AutomatosX collections for the family catalog, or the complete index.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP --local-dir ./AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP

Allow at least 20.86 GB of free disk space. Pin the resulting Hub commit in reproducible deployments rather than relying indefinitely on main.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

Serve with AX Engine and MTP

After installing AX Engine, download the complete repository (see AXQuant for conversion, certificates, and model-card tooling) and serve the local directory:

ax-engine serve ./AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP --port 31418

AX Engine is the authority for the AXQ runtime contract and native MTP sidecar. Runtime speedup claims are limited to the authorizing profiles in the linked Tier 2 certificate. The artifact records AX Engine version 6.16.1. Native model-manifest.json status: included as model-manifest.json.

Use with oMLX Lightning MTP

Download the complete repository into a writable local directory and add that directory to oMLX. With an oMLX build that includes the sidecar importer, select Import MTP side-car in Model Settings, then enable Lightning MTP. The bundled mtplx_runtime.json declares the qwen3-next-mtp execution contract required by oMLX's strict importer; this identifier describes the MTP head layout, not the base model's qwen3_5 model type.

The one-time import adds the sidecar tensors to the local checkpoint index. oMLX compatibility does not extend AXQuant's Tier 2 speed claim beyond the certificate's authorizing profiles; benchmark the result in the intended oMLX workload. At the time of the 2026-08-20 smoke test, direct sidecar import was available on oMLX main (0.6.3rc2); stable oMLX 0.5.7 successfully served the already imported text checkpoint with Lightning MTP. Its VLM loader rejected 333 vision parameters and fell back to the LLM path, so this is a text-only compatibility result.

Use with MTPLX

MTPLX 2.5.2 recognizes the bundled depth-1 sidecar and can serve it with the stable profile:

mtplx quickstart \
  --model ./AX-Qwen3.8-27B-MLX-AXQ-6bit-MTP \
  --profile stable \
  --depth 1 \
  --reasoning off

The contract is intentionally marked unverified with a public-release blocker because no MTPLX Forge exactness baseline is bound to this development artifact. It permits compatibility testing without --unsafe-force-unverified; it does not claim MTPLX exactness certification.

Runtime compatibility speed smoke (not certification)

Measured on an Apple M5 Max with 128 GB unified memory and AC power. Each number is the median of three measured batch-1 requests after one warmup, with temperature 0, thinking off, and MTP depth

  1. Prefill prompts use a unique salt; decode uses the same fixed prompt and runs to 256 tokens.
RuntimePrefill inputPrefillDecode input / outputDecode
oMLX 0.5.72,158 uncached tokens838.9 tok/s155 / 256 tokens43.46 tok/s
AX Engine 7.1.52,122 uncached of 2,154 tokens802.5 tok/s148 / 256 tokens42.56 tok/s
MTPLX 2.5.22,157 uncached tokens738.5 tok/s155 / 256 tokens42.89 tok/s

Prefill throughput is uncached prompt tokens divided by local loopback time to first content, so it includes HTTP, scheduling, and first-token overhead. Decode throughput is (completion tokens - 1) divided by the first-to-last-content interval. The three decode results are within about 2.1%; do not generalize this short smoke to concurrency, long context, vision, quality, or another host.

Metadata compatibility update (2026-08-20): mtplx_runtime.json now includes the complete Qwen MTP runtime identity, tensor count, MTPLX compatibility version, recommended profile, and explicit unverified-exactness status. Safetensors payloads are unchanged.

Quantization layout

Main-weight precisionParametersShare
4bit24.34B87.60%
6bit15.16M0.05%
8bit1.27B4.58%
bf162.16B7.77%
  • Quantization methods: affine, bf16.
  • Group sizes used by quantized assignments: 32, 64.
  • MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
  • Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
  • Vision weights: protected BF16 sidecar.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

CheckStatus
Planning evidencearchitecture_prior
Calibrationnone; the allocation is based on architecture priors
Quantizer execution497/497 recorded module conversions succeeded; 0 fallbacks
AX Engine native manifestincluded as model-manifest.json
Quality versus BF16 or uniform baselinesNot published; no quality-retention claim
MTP acceptance and speedcertified for the certificate's authorizing profiles only; outside that scope there is no speedup claim
AX Engine kernel evidenceunmeasured
Vision-language qualityNot evaluated or claimed; vision tensors are preserved at BF16
Speech-recognition qualityNot applicable
Long-context quality262,144-token capacity is config metadata, not a validated claim
Release certificationCheckpoint Tier 1 certified on df-macbookpro-m3 (2026-08-14), Hub commit a5a0b700ea7c; the formal AXQuant M0-M8 release campaign is a separate process and is not implied
Studio recert (df-macstudio-m2, 2026-08-15)Not certified as a replacement T1. Candidate means on prepare-suite v2: agent-coding 0.933 (52), general 0.875 (16). BF16 is not on Ext12T so 0.98 retention was not computed. Historical record is unchanged. See studio evaluation.

Intended use and limitations

  • Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.

  • MTP requires a sidecar-aware runtime. oMLX/MTPLX discovery compatibility does not extend the linked AX Engine certificate to those runtimes.

  • Vision weights are preserved at BF16, but this release does not claim validated VLM quality.

  • The configured context window can require substantially more memory as the KV cache grows.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. If an OptiQ repository is published separately, it uses a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.8-27B model card for license terms, model limitations, and responsible-use guidance.

Contributors

AutomatosX

11 commits