AutomatosX/AX-Qwen3-Coder-Next-MLX-6bit

Model

AX Qwen3 Coder Next MLX 6-bit

1

5 commits

1 linked in READMEs

updated Jul 20, 2026

See the code
6-bit
apple-silicon
automatosx
ax-engine
coding
conversational
mirror
mixture-of-experts
quantized
qwen3-next
qwen3_next
safetensors
text-generation

README

AX Qwen3 Coder Next MLX 6-bit

Parameter count: approximately 79.67B logical parameters (80B total, 3B active per token). 6-bit is the quantization precision, not a 6B model-size claim.

This is a revision-pinned, transparent mirror of mlx-community/Qwen3-Coder-Next-6bit at commit 9d12cc36cc6c386ffd04f7c8f0de6ccb29c5927e.

The model weights, index, configuration, tokenizer, chat template, generation configuration, and tool-parser files are byte-identical to that upstream revision. AutomatosX did not fine-tune, merge, re-quantize, or otherwise alter the model artifacts. We add only mirror documentation, a copy of the declared Apache 2.0 license, and machine-readable provenance.

Model details

  • Base model: Qwen/Qwen3-Coder-Next
  • Format: MLX Safetensors for Apple Silicon
  • Architecture: Qwen3NextForCausalLM, mixture of experts
  • Main quantization: 6-bit affine, group size 64
  • Quantization exceptions: router and shared-expert gate tensors remain 8-bit, as defined by the unchanged upstream config
  • Layers: 48
  • Experts: 512 total, 10 selected per token
  • Configured context limit: 262,144 tokens
  • Weight shards: 13, totaling 64,749,929,465 bytes
  • Upstream conversion tool: mlx-lm 0.30.5

Download

hf download AutomatosX/AX-Qwen3-Coder-Next-MLX-6bit \
  --local-dir ./AX-Qwen3-Coder-Next-MLX-6bit

Use with MLX-LM

pip install mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3-Coder-Next-MLX-6bit \
  --prompt "Write a Python function that merges two sorted lists."

Applications should apply the included chat template for conversational or tool-using prompts.

Serve with AX Engine

You can also serve the downloaded model through the OpenAI-compatible API in AX Engine:

ax-engine serve ./AX-Qwen3-Coder-Next-MLX-6bit --port 31418

Mirror policy and provenance

This release is meant to group a required upstream artifact under the AutomatosX catalog, not to claim a new conversion. UPSTREAM_README.md preserves the original upstream model card. ax_provenance.json pins the source commit and records SHA-256 values and sizes for every mirrored artifact.

This is a standard direct-decoding model. It does not contain an MTP head and no MTP conversion was applied.

License

The upstream model metadata declares Apache License 2.0. See LICENSE, the upstream model card, and the base-model card for limitations and responsible-use guidance.

Contributors

AutomatosX

5 commits

AutomatosX/AX-Qwen3-Coder-Next-MLX-6bit

Model

AX Qwen3 Coder Next MLX 6-bit

1

5 commits

1 linked in READMEs

updated Jul 20, 2026

See the code
6-bit
apple-silicon
automatosx
ax-engine
coding
conversational
mirror
mixture-of-experts
quantized
qwen3-next
qwen3_next
safetensors
text-generation

README

AX Qwen3 Coder Next MLX 6-bit

Parameter count: approximately 79.67B logical parameters (80B total, 3B active per token). 6-bit is the quantization precision, not a 6B model-size claim.

This is a revision-pinned, transparent mirror of mlx-community/Qwen3-Coder-Next-6bit at commit 9d12cc36cc6c386ffd04f7c8f0de6ccb29c5927e.

The model weights, index, configuration, tokenizer, chat template, generation configuration, and tool-parser files are byte-identical to that upstream revision. AutomatosX did not fine-tune, merge, re-quantize, or otherwise alter the model artifacts. We add only mirror documentation, a copy of the declared Apache 2.0 license, and machine-readable provenance.

Model details

  • Base model: Qwen/Qwen3-Coder-Next
  • Format: MLX Safetensors for Apple Silicon
  • Architecture: Qwen3NextForCausalLM, mixture of experts
  • Main quantization: 6-bit affine, group size 64
  • Quantization exceptions: router and shared-expert gate tensors remain 8-bit, as defined by the unchanged upstream config
  • Layers: 48
  • Experts: 512 total, 10 selected per token
  • Configured context limit: 262,144 tokens
  • Weight shards: 13, totaling 64,749,929,465 bytes
  • Upstream conversion tool: mlx-lm 0.30.5

Download

hf download AutomatosX/AX-Qwen3-Coder-Next-MLX-6bit \
  --local-dir ./AX-Qwen3-Coder-Next-MLX-6bit

Use with MLX-LM

pip install mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3-Coder-Next-MLX-6bit \
  --prompt "Write a Python function that merges two sorted lists."

Applications should apply the included chat template for conversational or tool-using prompts.

Serve with AX Engine

You can also serve the downloaded model through the OpenAI-compatible API in AX Engine:

ax-engine serve ./AX-Qwen3-Coder-Next-MLX-6bit --port 31418

Mirror policy and provenance

This release is meant to group a required upstream artifact under the AutomatosX catalog, not to claim a new conversion. UPSTREAM_README.md preserves the original upstream model card. ax_provenance.json pins the source commit and records SHA-256 values and sizes for every mirrored artifact.

This is a standard direct-decoding model. It does not contain an MTP head and no MTP conversion was applied.

License

The upstream model metadata declares Apache License 2.0. See LICENSE, the upstream model card, and the base-model card for limitations and responsible-use guidance.

Contributors

AutomatosX

5 commits