GGUF export of Omi Med STT v1
for Linux and Windows CPU use through the omi-med-stt CLI.
This is the portability path. If you have Apple Silicon, use the MLX q8 repo. If you have an NVIDIA GPU, use the canonical NeMo checkpoint.
pip install -U omi-med-stt
omi-med-stt install-cpp --cpp-backend cpu
omi-med-stt audio.wav --runtime cpp
| File | Status |
|---|---|
omi-med-stt-v1-q8_0.gguf | Default CPU artifact, benchmarked |
omi-med-stt-v1-f16.gguf | Provided for conversion/experimentation; not independently benchmarked |
The public weights are unchanged. These September 2026 results come from the current documented runtime recipes on the same frozen 1,513-clip, 7.18-hour medical benchmark and scorer. No dictionary, custom vocabulary, contextual bias, reference-aware selection, or transcript correction was used.
| Runtime artifact | Platform | WER | M-WER | Drug M-WER | Medical recall |
|---|---|---|---|---|---|
| NeMo canonical | NVIDIA CUDA (L4, BF16) | 6.54% | 2.23% | 4.75% | 97.77% |
| MLX q8 | Apple Silicon (M4 Max) | 6.65% | 2.12% | 4.52% | 97.88% |
| GGUF q8_0 | Linux/Windows CPU | 7.10% | 2.16% | 4.30% | 97.84% |
The CPU path is the portable fallback. It recorded the lowest occurrence-based medical and drug error counts in this draw, but paired tests do not establish it as clinically better than GPU or MLX. Use CUDA for the best WER and throughput, or MLX q8 on Apple Silicon.
The previous model-card figures were produced by an older runtime/scorer path. The numbers above supersede them as runtime results; the GGUF weights did not change.
These files are not llama.cpp text-model GGUF files. They require a Parakeet ASR runtime. The supported path is:
omi-med-stt audio.wav --runtime cpp
The CLI installs the patched parakeet.cpp runtime needed for Omi Med STT v1.
omi-health/omi-med-stt-v1omi-health/omi-med-stt-v1-mlx-q8Omi-Health/omi-med-stt-runtimemudler/parakeet.cppOmi Med STT v1 is speech-to-text only. It is not a diagnostic, triage, prescribing, or clinical decision model, and it is not clinically validated. Transcripts must be reviewed before any clinical use.
GGUF export of Omi Med STT v1
for Linux and Windows CPU use through the omi-med-stt CLI.
This is the portability path. If you have Apple Silicon, use the MLX q8 repo. If you have an NVIDIA GPU, use the canonical NeMo checkpoint.
pip install -U omi-med-stt
omi-med-stt install-cpp --cpp-backend cpu
omi-med-stt audio.wav --runtime cpp
| File | Status |
|---|---|
omi-med-stt-v1-q8_0.gguf | Default CPU artifact, benchmarked |
omi-med-stt-v1-f16.gguf | Provided for conversion/experimentation; not independently benchmarked |
The public weights are unchanged. These September 2026 results come from the current documented runtime recipes on the same frozen 1,513-clip, 7.18-hour medical benchmark and scorer. No dictionary, custom vocabulary, contextual bias, reference-aware selection, or transcript correction was used.
| Runtime artifact | Platform | WER | M-WER | Drug M-WER | Medical recall |
|---|---|---|---|---|---|
| NeMo canonical | NVIDIA CUDA (L4, BF16) | 6.54% | 2.23% | 4.75% | 97.77% |
| MLX q8 | Apple Silicon (M4 Max) | 6.65% | 2.12% | 4.52% | 97.88% |
| GGUF q8_0 | Linux/Windows CPU | 7.10% | 2.16% | 4.30% | 97.84% |
The CPU path is the portable fallback. It recorded the lowest occurrence-based medical and drug error counts in this draw, but paired tests do not establish it as clinically better than GPU or MLX. Use CUDA for the best WER and throughput, or MLX q8 on Apple Silicon.
The previous model-card figures were produced by an older runtime/scorer path. The numbers above supersede them as runtime results; the GGUF weights did not change.
These files are not llama.cpp text-model GGUF files. They require a Parakeet ASR runtime. The supported path is:
omi-med-stt audio.wav --runtime cpp
The CLI installs the patched parakeet.cpp runtime needed for Omi Med STT v1.
omi-health/omi-med-stt-v1omi-health/omi-med-stt-v1-mlx-q8Omi-Health/omi-med-stt-runtimemudler/parakeet.cppOmi Med STT v1 is speech-to-text only. It is not a diagnostic, triage, prescribing, or clinical decision model, and it is not clinically validated. Transcripts must be reviewed before any clinical use.