Omi-Health/omi-med-stt-runtime

On-device English medical speech-to-text — CLI for Omi Med STT v1 (MLX / NeMo / parakeet.cpp)

Python

8

11 commits

updated Sep 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

Live scribing with Jev-ish utterance gating (r/LocalLLaMA)

Like everyone, I've been following the back-and-forth regarding Jev with interest. Arguments aside about the originality of the idea, the first thing I thought of when all of this came out is "gosh, that could really help my local scribe run realtime loops during a consultation". I pointed my…

3

Oct 7, 2026

README

Omi Med STT Runtime

PyPI Tests License: MIT

Command-line runtime for Omi Med STT v1, an English medical speech-to-text model built from NVIDIA Parakeet TDT 0.6B v2.

The package downloads the right model artifact for your machine and transcribes audio locally.

0.2.1 refreshes the public evaluation text across PyPI and the model cards. The runtime code and published model weights are unchanged from 0.2.0.

Install

pip install -U omi-med-stt

Apple Silicon:

pip install -U "omi-med-stt[mlx]"

NVIDIA CUDA / NeMo:

pip install -U "omi-med-stt[nemo]"

The NVIDIA adapter applies the qualified GPU recipe automatically: NeMo 3.0, BF16, local [256,256] attention, greedy-batch TDT decoding with max_symbols=10, timestamps disabled, and duration-sorted batches capped at eight files or 900 audio-seconds. Inputs are normalized through FFmpeg to mono 16 kHz PCM16. A BF16-capable NVIDIA GPU is required; no inference flags are needed beyond --runtime nemo.

Run

omi-med-stt audio.wav

Useful options:

omi-med-stt audio.wav --json
omi-med-stt audio.wav --runtime mlx
omi-med-stt audio.wav --runtime nemo
omi-med-stt audio.wav --runtime cpp
omi-med-stt check

Audio formats. WAV, FLAC, OGG and other libsndfile formats are read directly. Other inputs — .m4a (iPhone Voice Memos / QuickTime), .mp3, .aac, .mp4, .mov, .wma, .opus, .webm, … — are decoded with ffmpeg, which ships with the package, so there's nothing extra to install. If a system ffmpeg is on your PATH it's used instead (e.g. a newer build). Whatever the input, audio is downmixed to mono and resampled to 16 kHz automatically.

Runtime Choices

PlatformDefault runtimeModel artifact
Apple Siliconmlxomi-health/omi-med-stt-v1-mlx-q8
NVIDIA CUDAnemoomi-health/omi-med-stt-v1
Linux/Windows CPUcppomi-health/omi-med-stt-v1-gguf

The canonical model is the NeMo checkpoint. MLX and GGUF are runtime exports.

CPU setup:

omi-med-stt install-cpp --cpp-backend cpu
omi-med-stt audio.wav --runtime cpp

The CPU path uses a patched parakeet.cpp runtime and downloads the q8_0 GGUF artifact only. It does not download the NeMo or MLX weights. Unknown tokens are rendered as the same U+2047 marker the NeMo and MLX runtimes emit (rendering parity, not transcript correction).

Runtime Quality

ArtifactWERM-WERDrug M-WERMedical Recall
NeMo canonical6.54%2.23%4.75%97.77%
MLX q86.65%2.12%4.52%97.88%
GGUF q8_0 / CPU7.10%2.16%4.30%97.84%

These numbers compare the unchanged runtime artifacts on the same frozen 1,513-clip, 7.18-hour medical benchmark and scorer, using the runtime recipes shipped in this package. No dictionary, custom vocabulary, contextual bias, or transcript correction was used. The CPU row uses the silence-aware long-audio chunking shipped in 0.1.25. Its lower drug-error count in this draw is not a statistically established ranking over GPU or MLX; the GPU remains the best overall WER and throughput path, while MLX q8 is the selected Apple runtime.

Compared with the open-model rows on Omi's standing 30-system board, the CUDA and MLX q8 runtimes have the lowest observed WER, while MLX q8 has the second-lowest observed M-WER. These are positions in this benchmark draw, not a universal ranking.

See the full benchmark and runtime-specific results for the broader evaluation and product context.

Runtime recipes and checks:

The recipe pages include the exact runtime settings, verification commands, and paths to the tests that enforce them.

Model Repositories

If the model repositories are private before launch, authenticate first:

huggingface-cli login

CUDA Note

If --runtime nemo fails with a CUDA driver mismatch, install a PyTorch wheel matching your driver before installing the NeMo extra. For example, on CUDA 12.8 hosts:

pip install torch --index-url https://download.pytorch.org/whl/cu128
pip install -U "omi-med-stt[nemo]"

Development

git clone https://github.com/Omi-Health/omi-med-stt-runtime
cd omi-med-stt-runtime
pip install -e ".[dev]"
python scripts/prepublish_check.py --skip-build
python -m pytest -q tests

Safety

Omi Med STT v1 is speech-to-text only. It is not a diagnostic, triage, prescribing, or clinical decision model, and it is not clinically validated. Transcripts must be reviewed before any clinical use.

License And Attribution

Runtime code is MIT licensed.

Model weights are CC-BY-4.0 and are derived from nvidia/parakeet-tdt-0.6b-v2. Omi Med STT v1 is not an NVIDIA model.

The CPU runtime uses parakeet.cpp.

asr
gguf
medical-asr
mlx
nemo
on-device
parakeet
speech-to-text

Omi-Health/omi-med-stt-runtime

On-device English medical speech-to-text — CLI for Omi Med STT v1 (MLX / NeMo / parakeet.cpp)

Python

8

11 commits

updated Sep 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

Live scribing with Jev-ish utterance gating (r/LocalLLaMA)

Like everyone, I've been following the back-and-forth regarding Jev with interest. Arguments aside about the originality of the idea, the first thing I thought of when all of this came out is "gosh, that could really help my local scribe run realtime loops during a consultation". I pointed my…

3

Oct 7, 2026

README

Omi Med STT Runtime

PyPI Tests License: MIT

Command-line runtime for Omi Med STT v1, an English medical speech-to-text model built from NVIDIA Parakeet TDT 0.6B v2.

The package downloads the right model artifact for your machine and transcribes audio locally.

0.2.1 refreshes the public evaluation text across PyPI and the model cards. The runtime code and published model weights are unchanged from 0.2.0.

Install

pip install -U omi-med-stt

Apple Silicon:

pip install -U "omi-med-stt[mlx]"

NVIDIA CUDA / NeMo:

pip install -U "omi-med-stt[nemo]"

The NVIDIA adapter applies the qualified GPU recipe automatically: NeMo 3.0, BF16, local [256,256] attention, greedy-batch TDT decoding with max_symbols=10, timestamps disabled, and duration-sorted batches capped at eight files or 900 audio-seconds. Inputs are normalized through FFmpeg to mono 16 kHz PCM16. A BF16-capable NVIDIA GPU is required; no inference flags are needed beyond --runtime nemo.

Run

omi-med-stt audio.wav

Useful options:

omi-med-stt audio.wav --json
omi-med-stt audio.wav --runtime mlx
omi-med-stt audio.wav --runtime nemo
omi-med-stt audio.wav --runtime cpp
omi-med-stt check

Audio formats. WAV, FLAC, OGG and other libsndfile formats are read directly. Other inputs — .m4a (iPhone Voice Memos / QuickTime), .mp3, .aac, .mp4, .mov, .wma, .opus, .webm, … — are decoded with ffmpeg, which ships with the package, so there's nothing extra to install. If a system ffmpeg is on your PATH it's used instead (e.g. a newer build). Whatever the input, audio is downmixed to mono and resampled to 16 kHz automatically.

Runtime Choices

PlatformDefault runtimeModel artifact
Apple Siliconmlxomi-health/omi-med-stt-v1-mlx-q8
NVIDIA CUDAnemoomi-health/omi-med-stt-v1
Linux/Windows CPUcppomi-health/omi-med-stt-v1-gguf

The canonical model is the NeMo checkpoint. MLX and GGUF are runtime exports.

CPU setup:

omi-med-stt install-cpp --cpp-backend cpu
omi-med-stt audio.wav --runtime cpp

The CPU path uses a patched parakeet.cpp runtime and downloads the q8_0 GGUF artifact only. It does not download the NeMo or MLX weights. Unknown tokens are rendered as the same U+2047 marker the NeMo and MLX runtimes emit (rendering parity, not transcript correction).

Runtime Quality

ArtifactWERM-WERDrug M-WERMedical Recall
NeMo canonical6.54%2.23%4.75%97.77%
MLX q86.65%2.12%4.52%97.88%
GGUF q8_0 / CPU7.10%2.16%4.30%97.84%

These numbers compare the unchanged runtime artifacts on the same frozen 1,513-clip, 7.18-hour medical benchmark and scorer, using the runtime recipes shipped in this package. No dictionary, custom vocabulary, contextual bias, or transcript correction was used. The CPU row uses the silence-aware long-audio chunking shipped in 0.1.25. Its lower drug-error count in this draw is not a statistically established ranking over GPU or MLX; the GPU remains the best overall WER and throughput path, while MLX q8 is the selected Apple runtime.

Compared with the open-model rows on Omi's standing 30-system board, the CUDA and MLX q8 runtimes have the lowest observed WER, while MLX q8 has the second-lowest observed M-WER. These are positions in this benchmark draw, not a universal ranking.

See the full benchmark and runtime-specific results for the broader evaluation and product context.

Runtime recipes and checks:

The recipe pages include the exact runtime settings, verification commands, and paths to the tests that enforce them.

Model Repositories

If the model repositories are private before launch, authenticate first:

huggingface-cli login

CUDA Note

If --runtime nemo fails with a CUDA driver mismatch, install a PyTorch wheel matching your driver before installing the NeMo extra. For example, on CUDA 12.8 hosts:

pip install torch --index-url https://download.pytorch.org/whl/cu128
pip install -U "omi-med-stt[nemo]"

Development

git clone https://github.com/Omi-Health/omi-med-stt-runtime
cd omi-med-stt-runtime
pip install -e ".[dev]"
python scripts/prepublish_check.py --skip-build
python -m pytest -q tests

Safety

Omi Med STT v1 is speech-to-text only. It is not a diagnostic, triage, prescribing, or clinical decision model, and it is not clinically validated. Transcripts must be reviewed before any clinical use.

License And Attribution

Runtime code is MIT licensed.

Model weights are CC-BY-4.0 and are derived from nvidia/parakeet-tdt-0.6b-v2. Omi Med STT v1 is not an NVIDIA model.

The CPU runtime uses parakeet.cpp.

asr
gguf
medical-asr
mlx
nemo
on-device
parakeet
speech-to-text