Apple Silicon q8 export of Omi Med STT v1.
This is the default Mac artifact used by the omi-med-stt CLI. It is much
smaller than the full MLX export and keeps very similar benchmark quality.
pip install -U "omi-med-stt[mlx]"
omi-med-stt audio.wav
Explicit selection:
omi-med-stt audio.wav --runtime mlx --model omi-health/omi-med-stt-v1-mlx-q8
The public weights are unchanged. These September 2026 results come from the current documented runtime recipes on the same frozen 1,513-clip, 7.18-hour medical benchmark and scorer. No dictionary, custom vocabulary, contextual bias, reference-aware selection, or transcript correction was used.
| Runtime artifact | Platform | WER | M-WER | Drug M-WER | Medical recall |
|---|---|---|---|---|---|
| NeMo canonical | NVIDIA CUDA (L4, BF16) | 6.54% | 2.23% | 4.75% | 97.77% |
| MLX q8 | Apple Silicon (M4 Max) | 6.65% | 2.12% | 4.52% | 97.88% |
| GGUF q8_0 | Linux/Windows CPU | 7.10% | 2.16% | 4.30% | 97.84% |
Lower is better for the error rates; higher is better for medical recall. MLX q8 remains the Mac default because it combines the strongest canonical M-WER in the runtime matrix with a 0.94 GB artifact and bounded memory use.
Compared with the open-model rows on Omi's standing 30-system board, MLX q8 has the lowest observed WER and the second-lowest observed M-WER. These are positions in this benchmark draw, not a universal ranking.
The previous model-card figures were produced by an older runtime/scorer path. The numbers above supersede them as runtime results; the q8 weights did not change.
This is not a drop-in parakeet-mlx checkpoint. Omi Med STT v1 includes a
medical adapter, and the supported Mac path is the omi-med-stt CLI.
omi-health/omi-med-stt-v1omi-health/omi-med-stt-v1-mlxomi-health/omi-med-stt-v1-ggufOmi-Health/omi-med-stt-runtimeOmi Med STT v1 is speech-to-text only. It is not a diagnostic, triage, prescribing, or clinical decision model, and it is not clinically validated. Transcripts must be reviewed before any clinical use.
Apple Silicon q8 export of Omi Med STT v1.
This is the default Mac artifact used by the omi-med-stt CLI. It is much
smaller than the full MLX export and keeps very similar benchmark quality.
pip install -U "omi-med-stt[mlx]"
omi-med-stt audio.wav
Explicit selection:
omi-med-stt audio.wav --runtime mlx --model omi-health/omi-med-stt-v1-mlx-q8
The public weights are unchanged. These September 2026 results come from the current documented runtime recipes on the same frozen 1,513-clip, 7.18-hour medical benchmark and scorer. No dictionary, custom vocabulary, contextual bias, reference-aware selection, or transcript correction was used.
| Runtime artifact | Platform | WER | M-WER | Drug M-WER | Medical recall |
|---|---|---|---|---|---|
| NeMo canonical | NVIDIA CUDA (L4, BF16) | 6.54% | 2.23% | 4.75% | 97.77% |
| MLX q8 | Apple Silicon (M4 Max) | 6.65% | 2.12% | 4.52% | 97.88% |
| GGUF q8_0 | Linux/Windows CPU | 7.10% | 2.16% | 4.30% | 97.84% |
Lower is better for the error rates; higher is better for medical recall. MLX q8 remains the Mac default because it combines the strongest canonical M-WER in the runtime matrix with a 0.94 GB artifact and bounded memory use.
Compared with the open-model rows on Omi's standing 30-system board, MLX q8 has the lowest observed WER and the second-lowest observed M-WER. These are positions in this benchmark draw, not a universal ranking.
The previous model-card figures were produced by an older runtime/scorer path. The numbers above supersede them as runtime results; the q8 weights did not change.
This is not a drop-in parakeet-mlx checkpoint. Omi Med STT v1 includes a
medical adapter, and the supported Mac path is the omi-med-stt CLI.
omi-health/omi-med-stt-v1omi-health/omi-med-stt-v1-mlxomi-health/omi-med-stt-v1-ggufOmi-Health/omi-med-stt-runtimeOmi Med STT v1 is speech-to-text only. It is not a diagnostic, triage, prescribing, or clinical decision model, and it is not clinically validated. Transcripts must be reviewed before any clinical use.