Qwen3-ASR speech recognition on Apple Silicon via MLX
Python
215
239 commits
updated Sep 20, 2026
Run Qwen3-ASR — one of the strongest open-source speech recognition models — natively on Apple Silicon.
A ground-up reimplementation of the official PyTorch model using Apple's MLX framework. Same weights, benchmarked against official/reference outputs and ground-truth eval sets, optimized for Mac GPUs via Metal. No PyTorch dependency for core transcription.
Qwen3-ASR is one of the strongest open-source ASR models available, with benchmark results exceeding Whisper-large-v3 across multiple languages and datasets. It supports 30 languages plus 22 Chinese dialects. But the official implementation is PyTorch + NVIDIA CUDA — it doesn't use Apple GPUs.
This project rewrites every layer for MLX so the same model runs natively on M1/M2/M3/M4 hardware. Not a wrapper — a full reimplementation with correct interleaved MRoPE, per-chunk windowed encoder attention, and all the architectural details that matter for output quality.
pyannote integration (--diarize)mlx-qwen3-asr serve exposes the pipeline over HTTP with async jobs, OpenAI API compatibility, and Bearer token authInstall from PyPI:
pip install mlx-qwen3-asr
For video and most non-WAV audio formats, install ffmpeg on your system:
brew install ffmpeg
Install with optional timestamp alignment extras (for Japanese/Korean tokenization parity):
pip install "mlx-qwen3-asr[aligner]"
Install with optional microphone capture support:
pip install "mlx-qwen3-asr[mic]"
Install with HTTP server support:
pip install "mlx-qwen3-asr[serve]"
Install with diarization extras:
pip install "mlx-qwen3-asr[diarize]"
Note: --diarize uses pyannote.audio 4.x and defaults to
pyannote/speaker-diarization-community-1. Accept the model terms on
Hugging Face and set a token:
export PYANNOTE_AUTH_TOKEN=hf_...
Core ASR does not require any Hugging Face token.
For development:
git clone https://github.com/moona3k/mlx-qwen3-asr.git
cd mlx-qwen3-asr
pip install -e ".[dev]"
from mlx_qwen3_asr import transcribe
result = transcribe("audio.wav")
print(result.text)
print(result.language)
By default, transcribe() uses Qwen/Qwen3-ASR-0.6B for fast local usage on Mac. Use Qwen/Qwen3-ASR-1.7B when you want higher accuracy and can afford higher latency/memory.
With options:
result = transcribe(
"meeting.mp3",
model="Qwen/Qwen3-ASR-1.7B",
language="English",
return_chunks=True,
on_progress=lambda e: print(e["event"], e.get("progress", 0.0)),
verbose=True,
)
print(result.text)
print(result.chunks)
The Session object owns model and tokenizer state explicitly — no hidden globals, no cache surprises:
from mlx_qwen3_asr import Session
session = Session(model="Qwen/Qwen3-ASR-0.6B")
# Fast repeated transcription — model stays loaded
for audio_file in audio_files:
result = session.transcribe(audio_file)
print(result.text)
from mlx_qwen3_asr import load_model, load_audio, transcribe
model, config = load_model("Qwen/Qwen3-ASR-0.6B")
audio = load_audio("speech.wav")
result = transcribe(audio, model=model)
mlx-qwen3-asr audio.wav
Specify model, language, and output format:
mlx-qwen3-asr recording.mp3 --model Qwen/Qwen3-ASR-0.6B --language English -f srt -o output/
Word-level timestamps:
mlx-qwen3-asr audio.wav --timestamps
Speaker-labeled output (experimental, offline):
mlx-qwen3-asr meeting.wav --diarize --num-speakers 2 -f json
Multiple files with all output formats:
mlx-qwen3-asr *.wav -f all -o transcripts/ --verbose
Stdout/file behavior:
mlx-qwen3-asr audio.wav --stdout-only # print only (no output file)
mlx-qwen3-asr audio.wav --quiet -o out/ # write files only (no stdout text)
Language discovery:
mlx-qwen3-asr --list-languages
Environment diagnostics (ffmpeg, optional diarization deps, token status):
mlx-qwen3-asr --doctor
Run mlx-qwen3-asr --help for the full list of options.
Serve transcriptions over HTTP. Two endpoint styles: an async job API and an OpenAI-compatible synchronous endpoint.
pip install "mlx-qwen3-asr[serve]"
mlx-qwen3-asr serve --api-key $(openssl rand -hex 16)
Submit audio and poll for results:
# Submit
curl -X POST http://localhost:8765/transcribe \
-H "Authorization: Bearer YOUR_KEY" \
-F "audio=@recording.wav"
# Poll
curl http://localhost:8765/jobs/JOB_ID \
-H "Authorization: Bearer YOUR_KEY"
Or use the OpenAI-compatible endpoint with existing SDK code:
from openai import OpenAI
client = OpenAI(api_key="YOUR_KEY", base_url="http://localhost:8765/v1")
result = client.audio.transcriptions.create(
model="Qwen/Qwen3-ASR-0.6B",
file=open("recording.wav", "rb"),
)
print(result.text)
The async API is better for long audio (no HTTP timeout risk). The OpenAI endpoint blocks until done — simpler for short clips and SDK integration.
The server also implements /v1/models for SDK clients that perform model discovery.
See docs/server/ for the full API spec, deployment guide, and architecture decision record. See examples/ for copy-paste workflows covering the OpenAI-compatible server, subtitles, meetings, scanner/noisy audio, and batch folders.
Measured on Apple M4 Pro (48 GB), macOS 26, v0.4.0. All numbers come from
committed JSON artifacts under docs/benchmarks/; see
docs/BENCHMARKS.md for the full breakdown and
docs/benchmarks/2026-09-07-quality-matrix-refresh.md for the exact commands.
| Configuration | Short clip (~2.5s) | 10s clip | RTF (10s) | vs fp16 (10s) |
|---|---|---|---|---|
| 0.6B fp16 (baseline) | 0.17s | 0.30s | 0.029 | — |
| 0.6B 8-bit (g64) | 0.10s | 0.23s | 0.024 | 1.32x |
| 0.6B 4-bit (g64) | 0.09s | 0.17s | 0.018 | 1.71x |
| 1.7B fp16 | 0.36s | 0.73s | 0.077 | 2.4x slower |
| Model | Subset | WER | CER | Mean Latency | RTF |
|---|---|---|---|---|---|
| 0.6B | test-clean | 2.33% | 0.59% | 0.35s | 0.0393 |
| 0.6B | test-other | 4.30% | 2.11% | 0.40s | 0.0553 |
| 1.7B | test-clean | 1.94% | 0.57% | 0.77s | 0.0862 |
| 1.7B | test-other | 3.45% | 1.48% | 0.66s | 0.0914 |
| Configuration | test-clean WER | test-other WER | Speed vs fp16 (10s clip) |
|---|---|---|---|
| fp16 | 2.33% | 4.30% | — |
| 8-bit (g64) | 2.33% | 4.14% | 1.32x |
| 4-bit (g64) | 2.59% | 5.74% | 1.71x |
8-bit reproduces fp16 output exactly on test-clean. 4-bit trades about +0.3pp (clean) to +1.4pp (other) WER for the lowest latency.
| Model | Primary error rate | Mean latency | Best languages | Weakest |
|---|---|---|---|---|
| 0.6B fp16 | 9.54% | 0.65s | Spanish 3.0%, English 4.6%, Chinese 5.0% | Hindi 16.7%, French 17.3%, Arabic 21.5% |
| 1.7B fp16 | 6.70% | 1.22s | Spanish 0.7%, Japanese 3.6%, French 4.1% | Chinese 8.5%, Arabic 16.0%, Hindi 17.7% |
The 1.7B delivers a 30% relative improvement at 1.9x the latency. Per-language tables are in docs/BENCHMARKS.md.
| Metric | MLX | PyTorch | Delta |
|---|---|---|---|
| Primary error rate | 9.54% | 10.34% | -0.81pp |
| WER | 16.00% | 16.69% | -0.70pp |
| CER | 5.43% | 5.64% | -0.21pp |
68% of clips produce identical text; the rest differ by lexical or numeric surface form (10,000 vs zehntausend) or punctuation, not by quality. On LibriSpeech test-other the two are within 0.11pp WER, and on 80-second clips MLX scores 10.59% vs 17.99% because it chunks at pauses while the reference decodes the whole clip. The PyTorch reference runs on CPU on a Mac, so it is 7x to 20x slower here; that is a platform difference, not a like-for-like GPU comparison.
On the real-world mixed lane (AMI IHM meetings + Earnings22 chunked, n=200, measured February 2026), MLX was within 0.19pp WER of PyTorch (23.23% vs 23.04%).
mx.fast.scaled_dot_product_attention (no explicit K/V head expansion)transformers dependency in runtime transcription pathtranscribe() calls skip reload overheadFull benchmark report: docs/BENCHMARKS.md. Latest refresh snapshot: docs/benchmarks/2026-09-07-quality-matrix-refresh.md. All benchmark artifacts are committed under docs/benchmarks/ for reproducibility.
Word error rates from the Qwen3-ASR technical report compared against current open-source and proprietary leaders (lower is better):
| Benchmark | GPT-4o-Transcribe | Parakeet-TDT-0.6B | Whisper-large-v3 | Qwen3-ASR-0.6B | Qwen3-ASR-1.7B |
|---|---|---|---|---|---|
| LibriSpeech test-clean | 1.39 | 1.93 | 1.51 | 2.11 | 1.63 |
| LibriSpeech test-other | 3.75 | 3.59 | 3.97 | 4.55 | 3.38 |
| FLEURS-en | 2.40 | 4.85 | 4.08 | 4.39 | 3.35 |
| GigaSpeech | 25.50 | — | 9.76 | 8.88 | 8.45 |
| Benchmark | GPT-4o-Transcribe | Whisper-large-v3 | Qwen3-ASR-0.6B | Qwen3-ASR-1.7B |
|---|---|---|---|---|
| WenetSpeech test-net | 15.30 | 9.86 | 5.97 | 4.97 |
| AISHELL-2 test | 4.24 | 5.06 | 3.15 | 2.71 |
| FLEURS (12-lang avg) | — | 5.27 | 7.57 | 4.90 |
| CommonVoice | — | 10.77 | 12.75 | 9.18 |
| Benchmark | GPT-4o-Transcribe | Whisper-large-v3 | Qwen3-ASR-0.6B | Qwen3-ASR-1.7B |
|---|---|---|---|---|
| Accented English | 28.56 | 21.30 | 16.62 | 16.07 |
| Extreme Noise | 36.11 | 63.17 | 17.88 | 16.17 |
| Elders & Kids (Mandarin) | 14.27 | 10.61 | 4.48 | 3.81 |
GPT-4o-Transcribe leads on clean English read speech (1.39 WER). Parakeet-TDT-0.6B is strong on English. But Qwen3-ASR dominates on Chinese, multilingual, noisy, and accented speech — and is the only open-source model competitive across all categories.
Parakeet numbers from model card. All other numbers from the Qwen3-ASR paper. Robustness benchmarks are Qwen3-ASR internal test sets.
This implementation is validated against the official PyTorch model via multiple parity gates:
qwen-asr backend with 100% text match rate, <6ms timing MAE, and 2.64x speed advantage on 50 LibriSpeech samples| Qwen3-ASR-0.6B (default) | Qwen3-ASR-1.7B | |
|---|---|---|
| Parameters | 0.6B | 1.7B |
| Audio encoder layers | 18 | 24 |
| Audio encoder dim | 896 | 1024 |
| Text decoder layers | 28 | 28 |
| Text hidden size | 1024 | 2048 |
| Text attention (Q/KV heads) | GQA (16/8) | GQA (16/8) |
| RoPE theta | 1,000,000 | 1,000,000 |
| HuggingFace | Qwen/Qwen3-ASR-0.6B | Qwen/Qwen3-ASR-1.7B |
Both models use interleaved Multi-dimensional RoPE (MRoPE) with sections [24, 20, 20], 128-bin mel spectrograms, and the same tokenizer (vocabulary size 151,936).
# Default: 0.6B (fast, ~1.2 GB memory)
result = transcribe("audio.wav")
# Accuracy-first: 1.7B (~3.4 GB memory)
result = transcribe("audio.wav", model="Qwen/Qwen3-ASR-1.7B")
Word-level timestamps via forced alignment using a dedicated aligner model (Qwen/Qwen3-ForcedAligner-0.6B). This path is native MLX (no PyTorch backend bridge):
mlx-qwen3-asr audio.wav --timestamps
result = transcribe("audio.wav", return_timestamps=True)
for segment in result.segments:
print(f"{segment['start']:.2f}s - {segment['end']:.2f}s: {segment['text']}")
SRT/VTT outputs are grouped into subtitle-friendly phrase segments (not one word per cue).
When -f srt or -f vtt is requested in offline mode, timestamps are auto-enabled.
Measured parity (LibriSpeech test-clean, n=50):
| Metric | Value |
|---|---|
| Text match rate (MLX vs official) | 100% |
| Timing MAE (all word boundaries) | 5.69 ms |
| MLX aligner mean latency | 0.21s |
| Official backend mean latency | 0.56s |
| Relative speed | 2.64x faster |
The aligner uses O(n log n) LIS-based timestamp correction (Fenwick tree) for monotonicity repair, validated against the legacy O(n^2) implementation via randomized parity tests.
For Japanese/Korean timestamp alignment, install the [aligner] extra so nagisa/soynlp tokenization matches the official path.
Speaker attribution is available as an offline optional path powered by
pyannote.audio:
result = transcribe("meeting.wav", diarize=True)
print(result.speaker_segments)
mlx-qwen3-asr meeting.wav --diarize -f json
Current status:
[diarize] extra).pyannote/speaker-diarization-community-1; accept its
Hugging Face terms and configure PYANNOTE_AUTH_TOKEN (or HF_TOKEN).PYANNOTE_MODEL_ID can point to another pyannote pipeline or a local
offline clone.--diarize auto-enables timestamps and is not supported in --streaming/--mic mode.--diarize-device {auto,cpu,mps,cuda} (Python: diarization_device) selects
where the pyannote pipeline runs. The default auto prefers MPS, then CUDA,
then CPU; on an M-series Mac this cuts the diarization stage from minutes to
seconds with identical output. If the accelerator cannot run the pipeline,
the run warns and falls back to CPU.window/hop controls were
removed (diarization_window_sec, diarization_hop_sec,
--diarization-window-sec, --diarization-hop-sec). Speaker-count controls
remain (--num-speakers, --min-speakers, --max-speakers).pip install "mlx-qwen3-asr[diarize]"
pyannote/speaker-diarization-community-1 model terms
on Hugging Face and set a token:
export PYANNOTE_AUTH_TOKEN=hf_...
mlx-qwen3-asr meeting.wav --diarize -f json
Common errors and fixes:
requires optional dependency 'pyannote.audio': install [diarize] extra.requires PyTorch via pyannote dependencies: reinstall [diarize] extra in the active environment.Failed to initialize pyannote pipeline ...: accept model terms on Hugging Face, set PYANNOTE_AUTH_TOKEN (or HF_TOKEN), and inspect the Root cause: details.--streaming does not support --diarize / --mic does not support --diarize: use offline file transcription mode for diarization.Pre-quantized artifacts, validated with this runtime on 100 speaker-balanced
LibriSpeech test-clean clips (mlx-qwen3-asr >= 0.4.3):
| Model | Download | WER (fp16) | Notes |
|---|---|---|---|
moona3k/mlx-qwen3-asr-0.6b-4bit | 517 MB | 2.37% (2.33%) | 4-bit decoder, 8-bit encoder |
moona3k/mlx-qwen3-asr-0.6b-8bit | 801 MB | 2.33% (2.33%) | identical output to fp16 |
moona3k/mlx-qwen3-asr-1.7b-4bit | 1.2 GB | 1.73% (1.94%) | 4-bit decoder, 8-bit encoder |
moona3k/mlx-qwen3-asr-1.7b-8bit | 2.0 GB | 1.94% (1.94%) | identical output to fp16 |
mlx-qwen3-asr audio.wav --model moona3k/mlx-qwen3-asr-0.6b-4bit
Each model card carries the recipe, the per-sample evaluation reference and a
reproduce command. The mlx-community/Qwen3-ASR-* checkpoints also load
(since 0.4.1 they run in float16 rather than being promoted to float32).
Convert your own:
python scripts/convert.py \
--model Qwen/Qwen3-ASR-0.6B \
--quantize 4 --encoder-bits 8 --group-size 64 \
--output-dir ./qwen3-asr-4bit
mlx-qwen3-asr audio.wav --model ./qwen3-asr-4bit
The audio encoder carries most of the 4-bit quality loss (0.6B all-4-bit:
2.63% WER; with an 8-bit encoder: 2.37%), so --encoder-bits 8 is the
recommended 4-bit recipe; 8-bit throughout is lossless on this lane. Speed:
4-bit is about 1.7x and 8-bit about 1.3x faster than fp16 on a 10 s clip.
Publish to HuggingFace (converts, load-checks and uploads; --from-dir
uploads an already validated directory):
python scripts/publish_quantized.py \
--source-model Qwen/Qwen3-ASR-0.6B \
--repo-id YOUR_USER/mlx-qwen3-asr-0.6b-4bit \
--bits 4
The token is read from HF_TOKEN, else from the gitignored file
.secrets/hf_token at the repo root, else from huggingface-cli login.
mlx-qwen3-asr audio.wav -f txt # plain text
mlx-qwen3-asr audio.wav -f srt -o out/ # SRT subtitles
mlx-qwen3-asr audio.wav -f json # structured JSON
mlx-qwen3-asr audio.wav -f vtt -o out/ # WebVTT
mlx-qwen3-asr *.wav -f all -o out/ # all formats at once
Supported: txt, json, srt, vtt, tsv.
Subtitle formats (srt/vtt) require timestamp segments and are only supported in offline mode.
Qwen3-ASR officially lists 30 core languages:
| Arabic | Cantonese | Chinese | Czech |
| Danish | Dutch | English | Filipino |
| Finnish | French | German | Greek |
| Hindi | Hungarian | Indonesian | Italian |
| Japanese | Korean | Macedonian | Malay |
| Persian | Polish | Portuguese | Romanian |
| Russian | Spanish | Swedish | Thai |
| Turkish | Vietnamese |
Plus 22 Chinese dialects (Sichuan, Shanghai, Cantonese, and others), for 52 total language/dialect variants.
Print CLI-accepted aliases/codes:
mlx-qwen3-asr --list-languages
Uses the 0.6B model as a draft to accelerate 1.7B inference. Currently parity-safe but slower on tested workloads due to draft audio encoder overhead:
mlx-qwen3-asr audio.wav \
--model Qwen/Qwen3-ASR-1.7B \
--draft-model Qwen/Qwen3-ASR-0.6B \
--num-draft-tokens 4
result = transcribe(
"audio.wav",
model="Qwen/Qwen3-ASR-1.7B",
draft_model="Qwen/Qwen3-ASR-0.6B",
num_draft_tokens=4,
)
Status: greedy parity verified, but 0.53-0.55x on short/10s clips. Not enabled by default until benchmark evidence shows net speed wins.
When transcribing specialized audio — earnings calls, medical dictation, legal
proceedings — the model can confuse rare terms with more common homophones.
The context parameter lets you provide a hint: a string of domain-specific
words or phrases that gets injected into the system prompt, nudging the decoder
toward the correct vocabulary.
This matches the official Qwen3-ASR context API. The format is
space-separated terms:
# Finance: avoids "e-bit-da" → "EBITDA", "FX" not "effects", etc.
result = transcribe("earnings-call.wav", context="EBITDA non-GAAP FX hedging")
# Medical
result = transcribe("consult.wav", context="metformin HbA1c nephropathy")
# Also works with streaming
state = init_streaming(context="EBITDA non-GAAP FX hedging")
mlx-qwen3-asr earnings-call.wav --context "EBITDA non-GAAP FX hedging"
For batch transcription, pass a list of per-audio context strings:
results = transcribe_batch(
[audio_en, audio_zh],
context=["EBITDA non-GAAP", "交易 停滞"],
)
When omitted, the system prompt is empty (matching the official default) — no domain bias is applied.
Near-real-time transcription following the official streaming recipe: each chunk decodes the accumulated window (bounded by max_context_sec) with the previous text, minus its last few tokens, forced as a prefix. Encoder output and decoder KV for the parts of the window that cannot change are reused across chunks, and when the window fills it is committed at a pause rather than mid-word. Partial text is stable and the final text tracks offline quality (multilingual-100 primary error 11.3% streaming vs 9.5% offline; long-form 11.9% vs 10.6%).
from mlx_qwen3_asr.streaming import (
init_streaming,
feed_audio,
finish_streaming,
streaming_metrics,
)
state = init_streaming(chunk_size_sec=2.0, max_context_sec=30.0)
for chunk in audio_chunks:
state = feed_audio(chunk, state)
print(state.text)
state = finish_streaming(state)
print(streaming_metrics(state))
CLI:
mlx-qwen3-asr --streaming --stream-finalization-mode accuracy audio.wav
# Optional: speech-aware boundary selection near chunk edges
mlx-qwen3-asr --streaming --stream-endpointing-mode energy audio.wav
Live microphone transcription:
mlx-qwen3-asr --mic
mlx-qwen3-asr --mic --language Japanese
Optional microphone flags: --mic-device, --mic-duration-sec, --mic-sample-rate.
unfixed_token_num tokens, so per-chunk cost is bounded by the
window, not by session length; encoder output and decoder KV for complete
8 s attention windows are reused across chunks (RTF 0.08 on 5-20 s clips,
0.08 on 75 s clips with a 30 s window)unfixed_chunk_num, unfixed_token_num)stable_text is monotonic by design: corrections that would shorten already-stable
prefix text are intentionally not applied to the stable prefix (favoring stability
over maximal editability in partial output)endpointing_mode="energy") that selects
low-energy boundaries near chunk edgesfinalization_mode and enable_tail_refine are accepted for compatibility; the
window re-decode at finish covers what the former tail-refine pass didtranscribe(audio, *, model, draft_model, context, language, return_timestamps, diarize, diarization_num_speakers, diarization_min_speakers, diarization_max_speakers, diarization_device, return_chunks, forced_aligner, dtype, max_new_tokens, num_draft_tokens, verbose, on_progress)Transcribe audio to text. Accepts a file path, numpy array, mx.array, or (array, sample_rate) tuple. Returns a TranscriptionResult.
max_new_tokens=None (default) uses a duration-aware per-chunk decode budget to
avoid runaway generation on noisy inputs that do not emit EOS. Pass an integer
to override the cap explicitly. If you use unusually long custom chunks and see
truncated=True, pass a larger explicit value for that workload.
Additional Python entry points:
transcribe_batch(audios, ...) and transcribe_batch_async(audios, ...)transcribe_async(audio, ...)Session(model, *, dtype, tokenizer_model)Explicit transcription session. Owns model and tokenizer state with no hidden globals.
session.transcribe(audio, ...) with the same parameters as top-level transcribe.await session.transcribe_async(audio, ...).session.init_streaming(...), session.feed_audio(pcm, state), session.finish_streaming(state).session.model_info (model id/path, dtype, vocab size, model-declared language codes).streaming_metrics(state)Return streaming diagnostics for a session state:
partial_stabilityrewrite_ratefinalization_delta_charsload_model(name_or_path, *, dtype)Load a Qwen3-ASR model and config from HuggingFace or local path. Returns (model, config).
load_audio(path_or_url)Load and resample audio to mono 16 kHz. Returns an mx.array.
ForcedAligner(model_path, *, dtype, backend)Word-level forced aligner. Native backend: mlx (default).
TranscriptionResultFrozen dataclass:
text (str) — transcribed textlanguage (str) — detected or forced language (canonicalized names, e.g. English)segments (list[dict] | None) — word-level timestamps when requested: [{"text": "hello", "start": 0.5, "end": 0.8}, ...]chunks (list[dict] | None) — chunk-level transcript and generation metadata when return_chunks=Truespeaker_segments (list[dict] | None) — speaker-attributed spans when diarize=True: [{"speaker": "SPEAKER_00", "start": 0.0, "end": 2.0, "text": "..."}, ...]finish_reason (str | None) — aggregate decode stop reason: eos, repetition, length, or mixedtruncated (bool) — true when any chunk exhausted its token budget before EOS/repetitionThis project enforces parity with the official PyTorch implementation. No optimization lands without passing quality gates and committing benchmark artifacts.
# Unit tests (735 tests)
pytest -q
# Fast quality gate
python scripts/quality_gate.py --mode fast
# Release gate with token-level parity (downloads model weights)
RUN_REFERENCE_PARITY=1 python scripts/quality_gate.py --mode release
# Speaker-balanced WER evaluation (100 samples)
python scripts/eval_librispeech.py --subset test-clean --samples 100 --sampling speaker_round_robin
# Latency benchmark
python scripts/benchmark_asr.py tests/fixtures/test_speech.wav \
--model Qwen/Qwen3-ASR-0.6B --runs 5 \
--json-output docs/benchmarks/latest.json
Additional quality lanes available:
RUN_ALIGNER_PARITY=1 — validates MLX aligner against official backendRUN_REFERENCE_PARITY_SUITE=1 — test-clean, test-other, long mixes, noise variants with Unicode-safe text comparisonscripts/build_multilingual_manifest.py for cross-language validationRUN_STREAMING_MANIFEST_QUALITY_EVAL=1 with STREAMING_MANIFEST_QUALITY_EVAL_JSONL=... — multi-file streaming stability/rewrite/finalization lane via scripts/eval_streaming_manifest.pyRUN_REALWORLD_LONGFORM_EVAL=1 on full-recording Earnings22 manifestsRUN_DIARIZATION_QUALITY_EVAL=1 with DIARIZATION_QUALITY_EVAL_JSONL=... — DER/JER lane via scripts/eval_diarization.pySee docs/QUALITY_GATE.md for full documentation.
Evaluation coverage status and prioritized gaps are tracked in docs/EVAL_GAPS.md.
Audio (16kHz mono)
→ 128-bin log-mel spectrogram (native MLX, Whisper-compatible)
→ Conv2d stem (3 layers, stride 2 each → 8x downsample)
→ Sinusoidal position embeddings
→ Windowed transformer encoder (18 or 24 layers, hybrid dense/segmented attention)
→ LayerNorm + GELU projection → audio features
Chat-template prompt (context is optional domain vocabulary, empty by default):
<|im_start|>system\n{context}<|im_end|>
<|im_start|>user\n<|audio_start|><|audio_pad|>*N<|audio_end|><|im_end|>
<|im_start|>assistant\n
→ Token embedding (151,936 vocab)
→ Replace audio_pad positions with encoded audio features
→ Qwen3 text decoder (28 layers, interleaved MRoPE, SwiGLU, RMSNorm)
→ Autoregressive decode with preallocated KV cache
→ Parse output: "language English<asr_text>transcribed text here"
Key architectural details:
mlx_qwen3_asr/
├── transcribe.py # Public pipeline: transcribe, batch, async, diarization glue
├── session.py # Session API: explicit model/tokenizer ownership
├── streaming.py # Windowed re-decode with text-prefix rollback
├── cli.py # CLI (transcribe, serve, --mic, --doctor)
├── server.py # HTTP server + OpenAI-compatible endpoint
├── audio.py # Audio I/O, WAV fast path, mel spectrogram
├── chunking.py # Energy-based long-audio splitting
├── encoder.py # Audio encoder (Conv2d stem + windowed transformer)
├── decoder.py # Text decoder (GQA, SwiGLU, KV cache)
├── mrope.py # Interleaved MRoPE
├── attention.py # Shared SDPA helper
├── model.py # Qwen3ASRModel: audio-text fusion, prefill/step
├── generate.py # Greedy + speculative decoding
├── forced_aligner.py # Native MLX forced aligner + LIS correction
├── diarization.py # Optional pyannote integration
├── tokenizer.py # Native BPE tokenizer, language aliases, output parsing
├── load_models.py # HF download, weight loading, model cache
├── convert.py # Weight key remapping + Conv2d transpose
├── writers.py # txt/json/srt/vtt/tsv writers, subtitle cue grouping
└── config.py # Dataclass configs
tests/ # 12,672 lines, 735 tests
scripts/ # Benchmarks, evaluation, conversion, publishing
docs/ # Architecture, decisions, benchmarks, roadmap
docs/benchmarks/ # 160+ committed artifacts for reproducibility
git clone https://github.com/moona3k/mlx-qwen3-asr.git
cd mlx-qwen3-asr
pip install -e ".[dev]"
pytest -q # 735 tests
--diarize-device), @Tadanobu0 (server inference-thread fix), @cms42 (GPU memory release between chunks)Apache 2.0. See LICENSE for details.
781 followers · starred Aug 2026
241 followers · starred Jul 2026
188 followers · starred May 2026
Python
100.0%
Qwen3-ASR speech recognition on Apple Silicon via MLX
Python
215
239 commits
updated Sep 20, 2026
Run Qwen3-ASR — one of the strongest open-source speech recognition models — natively on Apple Silicon.
A ground-up reimplementation of the official PyTorch model using Apple's MLX framework. Same weights, benchmarked against official/reference outputs and ground-truth eval sets, optimized for Mac GPUs via Metal. No PyTorch dependency for core transcription.
Qwen3-ASR is one of the strongest open-source ASR models available, with benchmark results exceeding Whisper-large-v3 across multiple languages and datasets. It supports 30 languages plus 22 Chinese dialects. But the official implementation is PyTorch + NVIDIA CUDA — it doesn't use Apple GPUs.
This project rewrites every layer for MLX so the same model runs natively on M1/M2/M3/M4 hardware. Not a wrapper — a full reimplementation with correct interleaved MRoPE, per-chunk windowed encoder attention, and all the architectural details that matter for output quality.
pyannote integration (--diarize)mlx-qwen3-asr serve exposes the pipeline over HTTP with async jobs, OpenAI API compatibility, and Bearer token authInstall from PyPI:
pip install mlx-qwen3-asr
For video and most non-WAV audio formats, install ffmpeg on your system:
brew install ffmpeg
Install with optional timestamp alignment extras (for Japanese/Korean tokenization parity):
pip install "mlx-qwen3-asr[aligner]"
Install with optional microphone capture support:
pip install "mlx-qwen3-asr[mic]"
Install with HTTP server support:
pip install "mlx-qwen3-asr[serve]"
Install with diarization extras:
pip install "mlx-qwen3-asr[diarize]"
Note: --diarize uses pyannote.audio 4.x and defaults to
pyannote/speaker-diarization-community-1. Accept the model terms on
Hugging Face and set a token:
export PYANNOTE_AUTH_TOKEN=hf_...
Core ASR does not require any Hugging Face token.
For development:
git clone https://github.com/moona3k/mlx-qwen3-asr.git
cd mlx-qwen3-asr
pip install -e ".[dev]"
from mlx_qwen3_asr import transcribe
result = transcribe("audio.wav")
print(result.text)
print(result.language)
By default, transcribe() uses Qwen/Qwen3-ASR-0.6B for fast local usage on Mac. Use Qwen/Qwen3-ASR-1.7B when you want higher accuracy and can afford higher latency/memory.
With options:
result = transcribe(
"meeting.mp3",
model="Qwen/Qwen3-ASR-1.7B",
language="English",
return_chunks=True,
on_progress=lambda e: print(e["event"], e.get("progress", 0.0)),
verbose=True,
)
print(result.text)
print(result.chunks)
The Session object owns model and tokenizer state explicitly — no hidden globals, no cache surprises:
from mlx_qwen3_asr import Session
session = Session(model="Qwen/Qwen3-ASR-0.6B")
# Fast repeated transcription — model stays loaded
for audio_file in audio_files:
result = session.transcribe(audio_file)
print(result.text)
from mlx_qwen3_asr import load_model, load_audio, transcribe
model, config = load_model("Qwen/Qwen3-ASR-0.6B")
audio = load_audio("speech.wav")
result = transcribe(audio, model=model)
mlx-qwen3-asr audio.wav
Specify model, language, and output format:
mlx-qwen3-asr recording.mp3 --model Qwen/Qwen3-ASR-0.6B --language English -f srt -o output/
Word-level timestamps:
mlx-qwen3-asr audio.wav --timestamps
Speaker-labeled output (experimental, offline):
mlx-qwen3-asr meeting.wav --diarize --num-speakers 2 -f json
Multiple files with all output formats:
mlx-qwen3-asr *.wav -f all -o transcripts/ --verbose
Stdout/file behavior:
mlx-qwen3-asr audio.wav --stdout-only # print only (no output file)
mlx-qwen3-asr audio.wav --quiet -o out/ # write files only (no stdout text)
Language discovery:
mlx-qwen3-asr --list-languages
Environment diagnostics (ffmpeg, optional diarization deps, token status):
mlx-qwen3-asr --doctor
Run mlx-qwen3-asr --help for the full list of options.
Serve transcriptions over HTTP. Two endpoint styles: an async job API and an OpenAI-compatible synchronous endpoint.
pip install "mlx-qwen3-asr[serve]"
mlx-qwen3-asr serve --api-key $(openssl rand -hex 16)
Submit audio and poll for results:
# Submit
curl -X POST http://localhost:8765/transcribe \
-H "Authorization: Bearer YOUR_KEY" \
-F "audio=@recording.wav"
# Poll
curl http://localhost:8765/jobs/JOB_ID \
-H "Authorization: Bearer YOUR_KEY"
Or use the OpenAI-compatible endpoint with existing SDK code:
from openai import OpenAI
client = OpenAI(api_key="YOUR_KEY", base_url="http://localhost:8765/v1")
result = client.audio.transcriptions.create(
model="Qwen/Qwen3-ASR-0.6B",
file=open("recording.wav", "rb"),
)
print(result.text)
The async API is better for long audio (no HTTP timeout risk). The OpenAI endpoint blocks until done — simpler for short clips and SDK integration.
The server also implements /v1/models for SDK clients that perform model discovery.
See docs/server/ for the full API spec, deployment guide, and architecture decision record. See examples/ for copy-paste workflows covering the OpenAI-compatible server, subtitles, meetings, scanner/noisy audio, and batch folders.
Measured on Apple M4 Pro (48 GB), macOS 26, v0.4.0. All numbers come from
committed JSON artifacts under docs/benchmarks/; see
docs/BENCHMARKS.md for the full breakdown and
docs/benchmarks/2026-09-07-quality-matrix-refresh.md for the exact commands.
| Configuration | Short clip (~2.5s) | 10s clip | RTF (10s) | vs fp16 (10s) |
|---|---|---|---|---|
| 0.6B fp16 (baseline) | 0.17s | 0.30s | 0.029 | — |
| 0.6B 8-bit (g64) | 0.10s | 0.23s | 0.024 | 1.32x |
| 0.6B 4-bit (g64) | 0.09s | 0.17s | 0.018 | 1.71x |
| 1.7B fp16 | 0.36s | 0.73s | 0.077 | 2.4x slower |
| Model | Subset | WER | CER | Mean Latency | RTF |
|---|---|---|---|---|---|
| 0.6B | test-clean | 2.33% | 0.59% | 0.35s | 0.0393 |
| 0.6B | test-other | 4.30% | 2.11% | 0.40s | 0.0553 |
| 1.7B | test-clean | 1.94% | 0.57% | 0.77s | 0.0862 |
| 1.7B | test-other | 3.45% | 1.48% | 0.66s | 0.0914 |
| Configuration | test-clean WER | test-other WER | Speed vs fp16 (10s clip) |
|---|---|---|---|
| fp16 | 2.33% | 4.30% | — |
| 8-bit (g64) | 2.33% | 4.14% | 1.32x |
| 4-bit (g64) | 2.59% | 5.74% | 1.71x |
8-bit reproduces fp16 output exactly on test-clean. 4-bit trades about +0.3pp (clean) to +1.4pp (other) WER for the lowest latency.
| Model | Primary error rate | Mean latency | Best languages | Weakest |
|---|---|---|---|---|
| 0.6B fp16 | 9.54% | 0.65s | Spanish 3.0%, English 4.6%, Chinese 5.0% | Hindi 16.7%, French 17.3%, Arabic 21.5% |
| 1.7B fp16 | 6.70% | 1.22s | Spanish 0.7%, Japanese 3.6%, French 4.1% | Chinese 8.5%, Arabic 16.0%, Hindi 17.7% |
The 1.7B delivers a 30% relative improvement at 1.9x the latency. Per-language tables are in docs/BENCHMARKS.md.
| Metric | MLX | PyTorch | Delta |
|---|---|---|---|
| Primary error rate | 9.54% | 10.34% | -0.81pp |
| WER | 16.00% | 16.69% | -0.70pp |
| CER | 5.43% | 5.64% | -0.21pp |
68% of clips produce identical text; the rest differ by lexical or numeric surface form (10,000 vs zehntausend) or punctuation, not by quality. On LibriSpeech test-other the two are within 0.11pp WER, and on 80-second clips MLX scores 10.59% vs 17.99% because it chunks at pauses while the reference decodes the whole clip. The PyTorch reference runs on CPU on a Mac, so it is 7x to 20x slower here; that is a platform difference, not a like-for-like GPU comparison.
On the real-world mixed lane (AMI IHM meetings + Earnings22 chunked, n=200, measured February 2026), MLX was within 0.19pp WER of PyTorch (23.23% vs 23.04%).
mx.fast.scaled_dot_product_attention (no explicit K/V head expansion)transformers dependency in runtime transcription pathtranscribe() calls skip reload overheadFull benchmark report: docs/BENCHMARKS.md. Latest refresh snapshot: docs/benchmarks/2026-09-07-quality-matrix-refresh.md. All benchmark artifacts are committed under docs/benchmarks/ for reproducibility.
Word error rates from the Qwen3-ASR technical report compared against current open-source and proprietary leaders (lower is better):
| Benchmark | GPT-4o-Transcribe | Parakeet-TDT-0.6B | Whisper-large-v3 | Qwen3-ASR-0.6B | Qwen3-ASR-1.7B |
|---|---|---|---|---|---|
| LibriSpeech test-clean | 1.39 | 1.93 | 1.51 | 2.11 | 1.63 |
| LibriSpeech test-other | 3.75 | 3.59 | 3.97 | 4.55 | 3.38 |
| FLEURS-en | 2.40 | 4.85 | 4.08 | 4.39 | 3.35 |
| GigaSpeech | 25.50 | — | 9.76 | 8.88 | 8.45 |
| Benchmark | GPT-4o-Transcribe | Whisper-large-v3 | Qwen3-ASR-0.6B | Qwen3-ASR-1.7B |
|---|---|---|---|---|
| WenetSpeech test-net | 15.30 | 9.86 | 5.97 | 4.97 |
| AISHELL-2 test | 4.24 | 5.06 | 3.15 | 2.71 |
| FLEURS (12-lang avg) | — | 5.27 | 7.57 | 4.90 |
| CommonVoice | — | 10.77 | 12.75 | 9.18 |
| Benchmark | GPT-4o-Transcribe | Whisper-large-v3 | Qwen3-ASR-0.6B | Qwen3-ASR-1.7B |
|---|---|---|---|---|
| Accented English | 28.56 | 21.30 | 16.62 | 16.07 |
| Extreme Noise | 36.11 | 63.17 | 17.88 | 16.17 |
| Elders & Kids (Mandarin) | 14.27 | 10.61 | 4.48 | 3.81 |
GPT-4o-Transcribe leads on clean English read speech (1.39 WER). Parakeet-TDT-0.6B is strong on English. But Qwen3-ASR dominates on Chinese, multilingual, noisy, and accented speech — and is the only open-source model competitive across all categories.
Parakeet numbers from model card. All other numbers from the Qwen3-ASR paper. Robustness benchmarks are Qwen3-ASR internal test sets.
This implementation is validated against the official PyTorch model via multiple parity gates:
qwen-asr backend with 100% text match rate, <6ms timing MAE, and 2.64x speed advantage on 50 LibriSpeech samples| Qwen3-ASR-0.6B (default) | Qwen3-ASR-1.7B | |
|---|---|---|
| Parameters | 0.6B | 1.7B |
| Audio encoder layers | 18 | 24 |
| Audio encoder dim | 896 | 1024 |
| Text decoder layers | 28 | 28 |
| Text hidden size | 1024 | 2048 |
| Text attention (Q/KV heads) | GQA (16/8) | GQA (16/8) |
| RoPE theta | 1,000,000 | 1,000,000 |
| HuggingFace | Qwen/Qwen3-ASR-0.6B | Qwen/Qwen3-ASR-1.7B |
Both models use interleaved Multi-dimensional RoPE (MRoPE) with sections [24, 20, 20], 128-bin mel spectrograms, and the same tokenizer (vocabulary size 151,936).
# Default: 0.6B (fast, ~1.2 GB memory)
result = transcribe("audio.wav")
# Accuracy-first: 1.7B (~3.4 GB memory)
result = transcribe("audio.wav", model="Qwen/Qwen3-ASR-1.7B")
Word-level timestamps via forced alignment using a dedicated aligner model (Qwen/Qwen3-ForcedAligner-0.6B). This path is native MLX (no PyTorch backend bridge):
mlx-qwen3-asr audio.wav --timestamps
result = transcribe("audio.wav", return_timestamps=True)
for segment in result.segments:
print(f"{segment['start']:.2f}s - {segment['end']:.2f}s: {segment['text']}")
SRT/VTT outputs are grouped into subtitle-friendly phrase segments (not one word per cue).
When -f srt or -f vtt is requested in offline mode, timestamps are auto-enabled.
Measured parity (LibriSpeech test-clean, n=50):
| Metric | Value |
|---|---|
| Text match rate (MLX vs official) | 100% |
| Timing MAE (all word boundaries) | 5.69 ms |
| MLX aligner mean latency | 0.21s |
| Official backend mean latency | 0.56s |
| Relative speed | 2.64x faster |
The aligner uses O(n log n) LIS-based timestamp correction (Fenwick tree) for monotonicity repair, validated against the legacy O(n^2) implementation via randomized parity tests.
For Japanese/Korean timestamp alignment, install the [aligner] extra so nagisa/soynlp tokenization matches the official path.
Speaker attribution is available as an offline optional path powered by
pyannote.audio:
result = transcribe("meeting.wav", diarize=True)
print(result.speaker_segments)
mlx-qwen3-asr meeting.wav --diarize -f json
Current status:
[diarize] extra).pyannote/speaker-diarization-community-1; accept its
Hugging Face terms and configure PYANNOTE_AUTH_TOKEN (or HF_TOKEN).PYANNOTE_MODEL_ID can point to another pyannote pipeline or a local
offline clone.--diarize auto-enables timestamps and is not supported in --streaming/--mic mode.--diarize-device {auto,cpu,mps,cuda} (Python: diarization_device) selects
where the pyannote pipeline runs. The default auto prefers MPS, then CUDA,
then CPU; on an M-series Mac this cuts the diarization stage from minutes to
seconds with identical output. If the accelerator cannot run the pipeline,
the run warns and falls back to CPU.window/hop controls were
removed (diarization_window_sec, diarization_hop_sec,
--diarization-window-sec, --diarization-hop-sec). Speaker-count controls
remain (--num-speakers, --min-speakers, --max-speakers).pip install "mlx-qwen3-asr[diarize]"
pyannote/speaker-diarization-community-1 model terms
on Hugging Face and set a token:
export PYANNOTE_AUTH_TOKEN=hf_...
mlx-qwen3-asr meeting.wav --diarize -f json
Common errors and fixes:
requires optional dependency 'pyannote.audio': install [diarize] extra.requires PyTorch via pyannote dependencies: reinstall [diarize] extra in the active environment.Failed to initialize pyannote pipeline ...: accept model terms on Hugging Face, set PYANNOTE_AUTH_TOKEN (or HF_TOKEN), and inspect the Root cause: details.--streaming does not support --diarize / --mic does not support --diarize: use offline file transcription mode for diarization.Pre-quantized artifacts, validated with this runtime on 100 speaker-balanced
LibriSpeech test-clean clips (mlx-qwen3-asr >= 0.4.3):
| Model | Download | WER (fp16) | Notes |
|---|---|---|---|
moona3k/mlx-qwen3-asr-0.6b-4bit | 517 MB | 2.37% (2.33%) | 4-bit decoder, 8-bit encoder |
moona3k/mlx-qwen3-asr-0.6b-8bit | 801 MB | 2.33% (2.33%) | identical output to fp16 |
moona3k/mlx-qwen3-asr-1.7b-4bit | 1.2 GB | 1.73% (1.94%) | 4-bit decoder, 8-bit encoder |
moona3k/mlx-qwen3-asr-1.7b-8bit | 2.0 GB | 1.94% (1.94%) | identical output to fp16 |
mlx-qwen3-asr audio.wav --model moona3k/mlx-qwen3-asr-0.6b-4bit
Each model card carries the recipe, the per-sample evaluation reference and a
reproduce command. The mlx-community/Qwen3-ASR-* checkpoints also load
(since 0.4.1 they run in float16 rather than being promoted to float32).
Convert your own:
python scripts/convert.py \
--model Qwen/Qwen3-ASR-0.6B \
--quantize 4 --encoder-bits 8 --group-size 64 \
--output-dir ./qwen3-asr-4bit
mlx-qwen3-asr audio.wav --model ./qwen3-asr-4bit
The audio encoder carries most of the 4-bit quality loss (0.6B all-4-bit:
2.63% WER; with an 8-bit encoder: 2.37%), so --encoder-bits 8 is the
recommended 4-bit recipe; 8-bit throughout is lossless on this lane. Speed:
4-bit is about 1.7x and 8-bit about 1.3x faster than fp16 on a 10 s clip.
Publish to HuggingFace (converts, load-checks and uploads; --from-dir
uploads an already validated directory):
python scripts/publish_quantized.py \
--source-model Qwen/Qwen3-ASR-0.6B \
--repo-id YOUR_USER/mlx-qwen3-asr-0.6b-4bit \
--bits 4
The token is read from HF_TOKEN, else from the gitignored file
.secrets/hf_token at the repo root, else from huggingface-cli login.
mlx-qwen3-asr audio.wav -f txt # plain text
mlx-qwen3-asr audio.wav -f srt -o out/ # SRT subtitles
mlx-qwen3-asr audio.wav -f json # structured JSON
mlx-qwen3-asr audio.wav -f vtt -o out/ # WebVTT
mlx-qwen3-asr *.wav -f all -o out/ # all formats at once
Supported: txt, json, srt, vtt, tsv.
Subtitle formats (srt/vtt) require timestamp segments and are only supported in offline mode.
Qwen3-ASR officially lists 30 core languages:
| Arabic | Cantonese | Chinese | Czech |
| Danish | Dutch | English | Filipino |
| Finnish | French | German | Greek |
| Hindi | Hungarian | Indonesian | Italian |
| Japanese | Korean | Macedonian | Malay |
| Persian | Polish | Portuguese | Romanian |
| Russian | Spanish | Swedish | Thai |
| Turkish | Vietnamese |
Plus 22 Chinese dialects (Sichuan, Shanghai, Cantonese, and others), for 52 total language/dialect variants.
Print CLI-accepted aliases/codes:
mlx-qwen3-asr --list-languages
Uses the 0.6B model as a draft to accelerate 1.7B inference. Currently parity-safe but slower on tested workloads due to draft audio encoder overhead:
mlx-qwen3-asr audio.wav \
--model Qwen/Qwen3-ASR-1.7B \
--draft-model Qwen/Qwen3-ASR-0.6B \
--num-draft-tokens 4
result = transcribe(
"audio.wav",
model="Qwen/Qwen3-ASR-1.7B",
draft_model="Qwen/Qwen3-ASR-0.6B",
num_draft_tokens=4,
)
Status: greedy parity verified, but 0.53-0.55x on short/10s clips. Not enabled by default until benchmark evidence shows net speed wins.
When transcribing specialized audio — earnings calls, medical dictation, legal
proceedings — the model can confuse rare terms with more common homophones.
The context parameter lets you provide a hint: a string of domain-specific
words or phrases that gets injected into the system prompt, nudging the decoder
toward the correct vocabulary.
This matches the official Qwen3-ASR context API. The format is
space-separated terms:
# Finance: avoids "e-bit-da" → "EBITDA", "FX" not "effects", etc.
result = transcribe("earnings-call.wav", context="EBITDA non-GAAP FX hedging")
# Medical
result = transcribe("consult.wav", context="metformin HbA1c nephropathy")
# Also works with streaming
state = init_streaming(context="EBITDA non-GAAP FX hedging")
mlx-qwen3-asr earnings-call.wav --context "EBITDA non-GAAP FX hedging"
For batch transcription, pass a list of per-audio context strings:
results = transcribe_batch(
[audio_en, audio_zh],
context=["EBITDA non-GAAP", "交易 停滞"],
)
When omitted, the system prompt is empty (matching the official default) — no domain bias is applied.
Near-real-time transcription following the official streaming recipe: each chunk decodes the accumulated window (bounded by max_context_sec) with the previous text, minus its last few tokens, forced as a prefix. Encoder output and decoder KV for the parts of the window that cannot change are reused across chunks, and when the window fills it is committed at a pause rather than mid-word. Partial text is stable and the final text tracks offline quality (multilingual-100 primary error 11.3% streaming vs 9.5% offline; long-form 11.9% vs 10.6%).
from mlx_qwen3_asr.streaming import (
init_streaming,
feed_audio,
finish_streaming,
streaming_metrics,
)
state = init_streaming(chunk_size_sec=2.0, max_context_sec=30.0)
for chunk in audio_chunks:
state = feed_audio(chunk, state)
print(state.text)
state = finish_streaming(state)
print(streaming_metrics(state))
CLI:
mlx-qwen3-asr --streaming --stream-finalization-mode accuracy audio.wav
# Optional: speech-aware boundary selection near chunk edges
mlx-qwen3-asr --streaming --stream-endpointing-mode energy audio.wav
Live microphone transcription:
mlx-qwen3-asr --mic
mlx-qwen3-asr --mic --language Japanese
Optional microphone flags: --mic-device, --mic-duration-sec, --mic-sample-rate.
unfixed_token_num tokens, so per-chunk cost is bounded by the
window, not by session length; encoder output and decoder KV for complete
8 s attention windows are reused across chunks (RTF 0.08 on 5-20 s clips,
0.08 on 75 s clips with a 30 s window)unfixed_chunk_num, unfixed_token_num)stable_text is monotonic by design: corrections that would shorten already-stable
prefix text are intentionally not applied to the stable prefix (favoring stability
over maximal editability in partial output)endpointing_mode="energy") that selects
low-energy boundaries near chunk edgesfinalization_mode and enable_tail_refine are accepted for compatibility; the
window re-decode at finish covers what the former tail-refine pass didtranscribe(audio, *, model, draft_model, context, language, return_timestamps, diarize, diarization_num_speakers, diarization_min_speakers, diarization_max_speakers, diarization_device, return_chunks, forced_aligner, dtype, max_new_tokens, num_draft_tokens, verbose, on_progress)Transcribe audio to text. Accepts a file path, numpy array, mx.array, or (array, sample_rate) tuple. Returns a TranscriptionResult.
max_new_tokens=None (default) uses a duration-aware per-chunk decode budget to
avoid runaway generation on noisy inputs that do not emit EOS. Pass an integer
to override the cap explicitly. If you use unusually long custom chunks and see
truncated=True, pass a larger explicit value for that workload.
Additional Python entry points:
transcribe_batch(audios, ...) and transcribe_batch_async(audios, ...)transcribe_async(audio, ...)Session(model, *, dtype, tokenizer_model)Explicit transcription session. Owns model and tokenizer state with no hidden globals.
session.transcribe(audio, ...) with the same parameters as top-level transcribe.await session.transcribe_async(audio, ...).session.init_streaming(...), session.feed_audio(pcm, state), session.finish_streaming(state).session.model_info (model id/path, dtype, vocab size, model-declared language codes).streaming_metrics(state)Return streaming diagnostics for a session state:
partial_stabilityrewrite_ratefinalization_delta_charsload_model(name_or_path, *, dtype)Load a Qwen3-ASR model and config from HuggingFace or local path. Returns (model, config).
load_audio(path_or_url)Load and resample audio to mono 16 kHz. Returns an mx.array.
ForcedAligner(model_path, *, dtype, backend)Word-level forced aligner. Native backend: mlx (default).
TranscriptionResultFrozen dataclass:
text (str) — transcribed textlanguage (str) — detected or forced language (canonicalized names, e.g. English)segments (list[dict] | None) — word-level timestamps when requested: [{"text": "hello", "start": 0.5, "end": 0.8}, ...]chunks (list[dict] | None) — chunk-level transcript and generation metadata when return_chunks=Truespeaker_segments (list[dict] | None) — speaker-attributed spans when diarize=True: [{"speaker": "SPEAKER_00", "start": 0.0, "end": 2.0, "text": "..."}, ...]finish_reason (str | None) — aggregate decode stop reason: eos, repetition, length, or mixedtruncated (bool) — true when any chunk exhausted its token budget before EOS/repetitionThis project enforces parity with the official PyTorch implementation. No optimization lands without passing quality gates and committing benchmark artifacts.
# Unit tests (735 tests)
pytest -q
# Fast quality gate
python scripts/quality_gate.py --mode fast
# Release gate with token-level parity (downloads model weights)
RUN_REFERENCE_PARITY=1 python scripts/quality_gate.py --mode release
# Speaker-balanced WER evaluation (100 samples)
python scripts/eval_librispeech.py --subset test-clean --samples 100 --sampling speaker_round_robin
# Latency benchmark
python scripts/benchmark_asr.py tests/fixtures/test_speech.wav \
--model Qwen/Qwen3-ASR-0.6B --runs 5 \
--json-output docs/benchmarks/latest.json
Additional quality lanes available:
RUN_ALIGNER_PARITY=1 — validates MLX aligner against official backendRUN_REFERENCE_PARITY_SUITE=1 — test-clean, test-other, long mixes, noise variants with Unicode-safe text comparisonscripts/build_multilingual_manifest.py for cross-language validationRUN_STREAMING_MANIFEST_QUALITY_EVAL=1 with STREAMING_MANIFEST_QUALITY_EVAL_JSONL=... — multi-file streaming stability/rewrite/finalization lane via scripts/eval_streaming_manifest.pyRUN_REALWORLD_LONGFORM_EVAL=1 on full-recording Earnings22 manifestsRUN_DIARIZATION_QUALITY_EVAL=1 with DIARIZATION_QUALITY_EVAL_JSONL=... — DER/JER lane via scripts/eval_diarization.pySee docs/QUALITY_GATE.md for full documentation.
Evaluation coverage status and prioritized gaps are tracked in docs/EVAL_GAPS.md.
Audio (16kHz mono)
→ 128-bin log-mel spectrogram (native MLX, Whisper-compatible)
→ Conv2d stem (3 layers, stride 2 each → 8x downsample)
→ Sinusoidal position embeddings
→ Windowed transformer encoder (18 or 24 layers, hybrid dense/segmented attention)
→ LayerNorm + GELU projection → audio features
Chat-template prompt (context is optional domain vocabulary, empty by default):
<|im_start|>system\n{context}<|im_end|>
<|im_start|>user\n<|audio_start|><|audio_pad|>*N<|audio_end|><|im_end|>
<|im_start|>assistant\n
→ Token embedding (151,936 vocab)
→ Replace audio_pad positions with encoded audio features
→ Qwen3 text decoder (28 layers, interleaved MRoPE, SwiGLU, RMSNorm)
→ Autoregressive decode with preallocated KV cache
→ Parse output: "language English<asr_text>transcribed text here"
Key architectural details:
mlx_qwen3_asr/
├── transcribe.py # Public pipeline: transcribe, batch, async, diarization glue
├── session.py # Session API: explicit model/tokenizer ownership
├── streaming.py # Windowed re-decode with text-prefix rollback
├── cli.py # CLI (transcribe, serve, --mic, --doctor)
├── server.py # HTTP server + OpenAI-compatible endpoint
├── audio.py # Audio I/O, WAV fast path, mel spectrogram
├── chunking.py # Energy-based long-audio splitting
├── encoder.py # Audio encoder (Conv2d stem + windowed transformer)
├── decoder.py # Text decoder (GQA, SwiGLU, KV cache)
├── mrope.py # Interleaved MRoPE
├── attention.py # Shared SDPA helper
├── model.py # Qwen3ASRModel: audio-text fusion, prefill/step
├── generate.py # Greedy + speculative decoding
├── forced_aligner.py # Native MLX forced aligner + LIS correction
├── diarization.py # Optional pyannote integration
├── tokenizer.py # Native BPE tokenizer, language aliases, output parsing
├── load_models.py # HF download, weight loading, model cache
├── convert.py # Weight key remapping + Conv2d transpose
├── writers.py # txt/json/srt/vtt/tsv writers, subtitle cue grouping
└── config.py # Dataclass configs
tests/ # 12,672 lines, 735 tests
scripts/ # Benchmarks, evaluation, conversion, publishing
docs/ # Architecture, decisions, benchmarks, roadmap
docs/benchmarks/ # 160+ committed artifacts for reproducibility
git clone https://github.com/moona3k/mlx-qwen3-asr.git
cd mlx-qwen3-asr
pip install -e ".[dev]"
pytest -q # 735 tests
--diarize-device), @Tadanobu0 (server inference-thread fix), @cms42 (GPU memory release between chunks)Apache 2.0. See LICENSE for details.
781 followers · starred Aug 2026
241 followers · starred Jul 2026
188 followers · starred May 2026
Python
100.0%