AX Qwen3 Embedding 4B MLX 4-bit DWQ
0
7 commits
1 linked in READMEs
updated Jul 23, 2026
Parameter count: approximately 4.02B logical parameters.
4-bitis the DWQ quantization precision, not the model size.
This is an MLX text-embedding model for Apple Silicon. The weight tensor payload, configuration, and tokenizer are byte-identical to mlx-community/Qwen3-Embedding-4B-4bit-DWQ at revision b5d88f1fe49b50d2ac01b4692ca2d387f14f9c72.
AutomatosX adds an AX Engine native manifest, pinned provenance, tested serving instructions, and this model card. Only Safetensors header metadata was added to report the logical parameter count to the Hub; tensor payload bytes were not changed or re-quantized. This is an embedding-only release and does not use MTP.
hf download AutomatosX/AX-Qwen3-Embedding-4B-MLX-4bit-DWQ \
--local-dir ./AX-Qwen3-Embedding-4B-MLX-4bit-DWQ
Install AX Engine, then serve the downloaded directory:
ax-engine serve ./AX-Qwen3-Embedding-4B-MLX-4bit-DWQ \
--port 31418 -- --model-id qwen3
Create full-size normalized embeddings:
curl http://127.0.0.1:31418/v1/embeddings \
-H 'Content-Type: application/json' \
--data '{
"model": "qwen3",
"input": [
"Instruct: Find passages that answer this query.\nQuery: What is the capital of Canada?",
"Ottawa is the capital city of Canada."
],
"encoding_format": "float",
"pooling": "last",
"normalize": true
}'
For text inputs, AX Engine appends the configured EOS token before last-token pooling. Add an English task instruction to retrieval queries when useful; documents normally remain unprefixed. The endpoint returns 2,560-dimensional vectors. Do not pass the OpenAI dimensions field to this AX Engine version.
The release was validated on macOS arm64 with AX Engine 6.9.0:
See ax_provenance.json for pinned source and SHA-256 values.
Apache License 2.0. See LICENSE and the upstream Qwen model card for model limitations and responsible-use guidance.
7 commits
AX Qwen3 Embedding 4B MLX 4-bit DWQ
0
7 commits
1 linked in READMEs
updated Jul 23, 2026
Parameter count: approximately 4.02B logical parameters.
4-bitis the DWQ quantization precision, not the model size.
This is an MLX text-embedding model for Apple Silicon. The weight tensor payload, configuration, and tokenizer are byte-identical to mlx-community/Qwen3-Embedding-4B-4bit-DWQ at revision b5d88f1fe49b50d2ac01b4692ca2d387f14f9c72.
AutomatosX adds an AX Engine native manifest, pinned provenance, tested serving instructions, and this model card. Only Safetensors header metadata was added to report the logical parameter count to the Hub; tensor payload bytes were not changed or re-quantized. This is an embedding-only release and does not use MTP.
hf download AutomatosX/AX-Qwen3-Embedding-4B-MLX-4bit-DWQ \
--local-dir ./AX-Qwen3-Embedding-4B-MLX-4bit-DWQ
Install AX Engine, then serve the downloaded directory:
ax-engine serve ./AX-Qwen3-Embedding-4B-MLX-4bit-DWQ \
--port 31418 -- --model-id qwen3
Create full-size normalized embeddings:
curl http://127.0.0.1:31418/v1/embeddings \
-H 'Content-Type: application/json' \
--data '{
"model": "qwen3",
"input": [
"Instruct: Find passages that answer this query.\nQuery: What is the capital of Canada?",
"Ottawa is the capital city of Canada."
],
"encoding_format": "float",
"pooling": "last",
"normalize": true
}'
For text inputs, AX Engine appends the configured EOS token before last-token pooling. Add an English task instruction to retrieval queries when useful; documents normally remain unprefixed. The endpoint returns 2,560-dimensional vectors. Do not pass the OpenAI dimensions field to this AX Engine version.
The release was validated on macOS arm64 with AX Engine 6.9.0:
See ax_provenance.json for pinned source and SHA-256 values.
Apache License 2.0. See LICENSE and the upstream Qwen model card for model limitations and responsible-use guidance.
7 commits