FermionResearch/Phonon-1

Model

Phonon-1

9

7 commits

1 linked in READMEs

updated Sep 2, 2026

See the code

README

Phonon-1

Phonon-1 is an open speech recognition model for English. It downloads in 415 MB, runs on a laptop or a datacenter GPU, and transcribes an hour of audio in about two and a half minutes. It was trained at 2.4 bits per weight from the start, and it is the second model in the lab's low-bit lane after Neutrino-1.

Benchmarks

DatasetPhonon-1 (415 MB)Phonon-1 Micro (285 MB)Parakeet-0.6B 4-bit (637 MB)Moonshine base (248 MB)Whisper large-v3-turbo (1,619 MB)Whisper small (967 MB)wav2vec2-large (1,262 MB)Qwen3-ASR teacher (1,569 MB)
LibriSpeech test-clean2.6403.0022.1863.4172.103.4†2.8†2.235
LibriSpeech test-other5.6996.5113.9378.2624.077.6†6.3†4.618
TED-LIUM3.4213.8782.8295.272β€”β€”β€”2.889
SPGISpeech4.1634.8584.1045.7312.79†—13.31†3.074
VoxPopuli8.3949.1776.34510.47011.22†——7.151
GigaSpeech11.39611.8829.61412.1148.52†——9.321
Earnings-2212.57114.77111.19017.87211.07†—36.28†11.188
AMI13.08414.09412.72317.79015.16†——12.560
Macro (eight benchmarks)7.678.526.6210.1β€”β€”β€”6.63

Word error rate, lower is better. Unmarked cells: measured by us β€” full test sets, Whisper English text normalizer, greedy decoding. † = published figure (model card, paper, or the Open ASR Leaderboard). Dash = no comparable measurement.

Median 23.9Γ— realtime across nine corpora on a base M5 MacBook Air. That figure is decode time only, measured over utterances of up to 30 s with the model already loaded; the first run in a new environment compiles the runtime once (about 20 s), after which runs are fast.

Run it

pip install fermion-research
fermion transcribe recording.wav
fermion serve
curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
  -F "file=@recording.wav" \
  -F "model=FermionResearch/Phonon-1"

The same weights run on a Mac (via MLX), on an NVIDIA GPU, or on a plain CPU; the runtimes and Docker images are in the GitHub repo.

License

Apache License 2.0 for the weights and the command line. Base model: Qwen/Qwen3-ASR-0.6B, Apache-2.0.

apple-silicon
automatic-speech-recognition
low-bit
on-device
quantization-aware-training
speech-to-text
streaming
ternary

FermionResearch/Phonon-1

Model

Phonon-1

9

7 commits

1 linked in READMEs

updated Sep 2, 2026

See the code

README

Phonon-1

Phonon-1 is an open speech recognition model for English. It downloads in 415 MB, runs on a laptop or a datacenter GPU, and transcribes an hour of audio in about two and a half minutes. It was trained at 2.4 bits per weight from the start, and it is the second model in the lab's low-bit lane after Neutrino-1.

Benchmarks

DatasetPhonon-1 (415 MB)Phonon-1 Micro (285 MB)Parakeet-0.6B 4-bit (637 MB)Moonshine base (248 MB)Whisper large-v3-turbo (1,619 MB)Whisper small (967 MB)wav2vec2-large (1,262 MB)Qwen3-ASR teacher (1,569 MB)
LibriSpeech test-clean2.6403.0022.1863.4172.103.4†2.8†2.235
LibriSpeech test-other5.6996.5113.9378.2624.077.6†6.3†4.618
TED-LIUM3.4213.8782.8295.272β€”β€”β€”2.889
SPGISpeech4.1634.8584.1045.7312.79†—13.31†3.074
VoxPopuli8.3949.1776.34510.47011.22†——7.151
GigaSpeech11.39611.8829.61412.1148.52†——9.321
Earnings-2212.57114.77111.19017.87211.07†—36.28†11.188
AMI13.08414.09412.72317.79015.16†——12.560
Macro (eight benchmarks)7.678.526.6210.1β€”β€”β€”6.63

Word error rate, lower is better. Unmarked cells: measured by us β€” full test sets, Whisper English text normalizer, greedy decoding. † = published figure (model card, paper, or the Open ASR Leaderboard). Dash = no comparable measurement.

Median 23.9Γ— realtime across nine corpora on a base M5 MacBook Air. That figure is decode time only, measured over utterances of up to 30 s with the model already loaded; the first run in a new environment compiles the runtime once (about 20 s), after which runs are fast.

Run it

pip install fermion-research
fermion transcribe recording.wav
fermion serve
curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
  -F "file=@recording.wav" \
  -F "model=FermionResearch/Phonon-1"

The same weights run on a Mac (via MLX), on an NVIDIA GPU, or on a plain CPU; the runtimes and Docker images are in the GitHub repo.

License

Apache License 2.0 for the weights and the command line. Base model: Qwen/Qwen3-ASR-0.6B, Apache-2.0.

apple-silicon
automatic-speech-recognition
low-bit
on-device
quantization-aware-training
speech-to-text
streaming
ternary