Oído: open-vocabulary speech recognition that fits in a $5 ESP32-S3. 3.7% LibriSpeech WER, no cloud, no NPU.
See the code¡Oído! is what cooks call out in a Spanish kitchen to confirm an order: heard, got it.
Speech-to-text for any English sentence, running entirely on an ESP32-S3 (240 MHz dual-core Xtensa LX7, 8 MB PSRAM, 16 MB flash). No cloud, no command list, no neural accelerator. Built by Lokutor. Model on Hugging Face: lokutor-ai/oido-ctc-small-int8.
Status (30 September 2026). Every transcript below comes from the exact arithmetic of the on-chip engine: the host build is bit-identical to the firmware, and firmware transcripts under Espressif's QEMU emulator match it word for word. Real-time speed is estimated from exact emulator instruction counts. Measurements on physical boards follow in the next days and will be added here.
Word error rate (%) on LibriSpeech, same text normalization for every system.
| System | Runs on | test-clean | test-other | Size |
|---|---|---|---|---|
| Oído: NVIDIA Conformer-CTC Small, int8, greedy (this repo) | ESP32-S3 | 3.7 | 8.2 | 14.0 MB |
| Oído with NVIDIA Conformer-Transducer Small, int8 (weights not included, see below) | ESP32-S3 | 3.0 | 6.7 | 15.5 MB |
| Espressif MultiNet7 (ESP-SR benchmark; its API takes fixed command lists) | ESP32-S3 | 8.5 | 21.3 | 2.9 MB |
| Moonshine tiny, fp32 | laptop | 5.0 | 12.1 | 27 M params |
| Whisper tiny.en, fp32 | laptop | 6.3 | 15.9 | 39 M params |
| Vosk small (Kaldi) | laptop | 9.9 | 21.6 | 40 MB |
Robustness (300 LibriSpeech utterances under real DEMAND noise, babble and room reverb; eval/make_robust.py,
full numbers in results/robustness.json):
| Mean WER over 14 conditions | This repo, CTC int8 (chip) | Transducer int8 (chip) | Whisper tiny.en | Moonshine tiny | Vosk small |
|---|---|---|---|---|---|
| 8.4 | 6.7 | 12.1 | 12.2 | 21.7 |
For the transducer, car and kitchen noise at 5 dB SNR cost under 1 point, and living-room noise about 1.7. Four-talker babble at 5 dB and very reverberant rooms are the hard cases.
| Flash | 14.0 MB model (int8); partition layouts for 16 MB modules in esp32/firmware/partitions_*.csv |
| PSRAM | 2.4 MB working memory peak for a 20 s utterance (measured in QEMU); the rest caches the most-reused weights |
| Compute | ~225 M instructions per second of audio across both cores, ~121 M on the dual-core critical path (exact, QEMU -icount) |
| Real-time factor | estimated 0.7–0.95: 1.3–1.6 cycles per instruction at 240 MHz, plus flash/PSRAM stalls. Not yet measured on silicon |
| Latency | Utterance mode. Text appears after a 0.8 s pause plus compute: about 3 s for a 2–4 s command |
The engine (esp32/components/tinyasr) is new C written for the ESP32-S3's PIE vector unit:
EE.VMULAS.S8.ACCX (16 MACs per instruction), plus int4 outer-product kernels;On a laptop, with the chip's exact arithmetic. Needs Python with numpy, soundfile, sentencepiece and sounddevice.
cd esp32/host && make
./tasr_cli ../../models/nemo8.tnm recording.wav # 16 kHz mono PCM16 wav
python live_demo.py # microphone -> the firmware's VAD + engine, with ESP32 time estimates
On a board. ESP32-S3-DevKitC-1 N16R8, an INMP441 I2S microphone (SCK→GPIO4, WS→GPIO5, SD→GPIO6, L/R→GND), and optionally a 0.96" SSD1306 OLED (SDA→GPIO8, SCL→GPIO9). Needs ESP-IDF v5.5.
esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm # live microphone
TASR_OLED=1 esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm # + transcript on the OLED
TASR_MODE=file esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm clip.wav "reference" # prints measured RTF
In the emulator (Espressif QEMU 9.x): runs the real firmware, then reports the transcript and instruction counts.
esp32/tools/emulate.sh clip.wav
The transducer model. Its weights are NVIDIA's, distributed on NGC under NVIDIA's terms, so they are not included here. You can fetch and convert them yourself:
cd train && python fetch_nemo_small.py ../models/nemo_rnnt --transducer
python export_nemo.py ../models/nemo_rnnt ../models/rnnt8.tnm 8
esp32/components/tinyasr/ on-chip engine: tasr_nemo.c (Conformer CTC/RNN-T), kernels.c (PIE SIMD), tinyasr_lm.c
(GRU LM + beam search), tasr_seg.c (VAD), tinyasr.c (streaming engine)
esp32/firmware/ ESP-IDF app: live I2S microphone or benchmark mode, OLED, partition layouts
esp32/host/ host build of the engine: tasr_cli, live_demo.py, seg_test, eval_engine.py, benchmark.py
esp32/tools/ flash.sh, emulate.sh, run_qemu.sh, bench_latency.py, mkimages.py
train/ PyTorch port of NVIDIA's model (nemo_small.py, rnnt_small.py), exporters, GRU LM training
eval/ WER normalization, robustness benchmark builder, laptop baselines
results/ benchmark outputs behind the numbers above
models/ nemo8.tnm (int8 Conformer-CTC Small) and its tokenizer
LICENSE).COMMERCIAL.md.models/ are derived from NVIDIA's stt_en_conformer_ctc_small and remain under
CC-BY-4.0. See NOTICE.Lokutor also has Spanish and other-language models, an int4 profile with more compute headroom, and an on-device TTS for the same chip. Contact us for these.
C
70.0%
Python
27.0%
Shell
2.3%
Oído: open-vocabulary speech recognition that fits in a $5 ESP32-S3. 3.7% LibriSpeech WER, no cloud, no NPU.
See the code¡Oído! is what cooks call out in a Spanish kitchen to confirm an order: heard, got it.
Speech-to-text for any English sentence, running entirely on an ESP32-S3 (240 MHz dual-core Xtensa LX7, 8 MB PSRAM, 16 MB flash). No cloud, no command list, no neural accelerator. Built by Lokutor. Model on Hugging Face: lokutor-ai/oido-ctc-small-int8.
Status (30 September 2026). Every transcript below comes from the exact arithmetic of the on-chip engine: the host build is bit-identical to the firmware, and firmware transcripts under Espressif's QEMU emulator match it word for word. Real-time speed is estimated from exact emulator instruction counts. Measurements on physical boards follow in the next days and will be added here.
Word error rate (%) on LibriSpeech, same text normalization for every system.
| System | Runs on | test-clean | test-other | Size |
|---|---|---|---|---|
| Oído: NVIDIA Conformer-CTC Small, int8, greedy (this repo) | ESP32-S3 | 3.7 | 8.2 | 14.0 MB |
| Oído with NVIDIA Conformer-Transducer Small, int8 (weights not included, see below) | ESP32-S3 | 3.0 | 6.7 | 15.5 MB |
| Espressif MultiNet7 (ESP-SR benchmark; its API takes fixed command lists) | ESP32-S3 | 8.5 | 21.3 | 2.9 MB |
| Moonshine tiny, fp32 | laptop | 5.0 | 12.1 | 27 M params |
| Whisper tiny.en, fp32 | laptop | 6.3 | 15.9 | 39 M params |
| Vosk small (Kaldi) | laptop | 9.9 | 21.6 | 40 MB |
Robustness (300 LibriSpeech utterances under real DEMAND noise, babble and room reverb; eval/make_robust.py,
full numbers in results/robustness.json):
| Mean WER over 14 conditions | This repo, CTC int8 (chip) | Transducer int8 (chip) | Whisper tiny.en | Moonshine tiny | Vosk small |
|---|---|---|---|---|---|
| 8.4 | 6.7 | 12.1 | 12.2 | 21.7 |
For the transducer, car and kitchen noise at 5 dB SNR cost under 1 point, and living-room noise about 1.7. Four-talker babble at 5 dB and very reverberant rooms are the hard cases.
| Flash | 14.0 MB model (int8); partition layouts for 16 MB modules in esp32/firmware/partitions_*.csv |
| PSRAM | 2.4 MB working memory peak for a 20 s utterance (measured in QEMU); the rest caches the most-reused weights |
| Compute | ~225 M instructions per second of audio across both cores, ~121 M on the dual-core critical path (exact, QEMU -icount) |
| Real-time factor | estimated 0.7–0.95: 1.3–1.6 cycles per instruction at 240 MHz, plus flash/PSRAM stalls. Not yet measured on silicon |
| Latency | Utterance mode. Text appears after a 0.8 s pause plus compute: about 3 s for a 2–4 s command |
The engine (esp32/components/tinyasr) is new C written for the ESP32-S3's PIE vector unit:
EE.VMULAS.S8.ACCX (16 MACs per instruction), plus int4 outer-product kernels;On a laptop, with the chip's exact arithmetic. Needs Python with numpy, soundfile, sentencepiece and sounddevice.
cd esp32/host && make
./tasr_cli ../../models/nemo8.tnm recording.wav # 16 kHz mono PCM16 wav
python live_demo.py # microphone -> the firmware's VAD + engine, with ESP32 time estimates
On a board. ESP32-S3-DevKitC-1 N16R8, an INMP441 I2S microphone (SCK→GPIO4, WS→GPIO5, SD→GPIO6, L/R→GND), and optionally a 0.96" SSD1306 OLED (SDA→GPIO8, SCL→GPIO9). Needs ESP-IDF v5.5.
esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm # live microphone
TASR_OLED=1 esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm # + transcript on the OLED
TASR_MODE=file esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm clip.wav "reference" # prints measured RTF
In the emulator (Espressif QEMU 9.x): runs the real firmware, then reports the transcript and instruction counts.
esp32/tools/emulate.sh clip.wav
The transducer model. Its weights are NVIDIA's, distributed on NGC under NVIDIA's terms, so they are not included here. You can fetch and convert them yourself:
cd train && python fetch_nemo_small.py ../models/nemo_rnnt --transducer
python export_nemo.py ../models/nemo_rnnt ../models/rnnt8.tnm 8
esp32/components/tinyasr/ on-chip engine: tasr_nemo.c (Conformer CTC/RNN-T), kernels.c (PIE SIMD), tinyasr_lm.c
(GRU LM + beam search), tasr_seg.c (VAD), tinyasr.c (streaming engine)
esp32/firmware/ ESP-IDF app: live I2S microphone or benchmark mode, OLED, partition layouts
esp32/host/ host build of the engine: tasr_cli, live_demo.py, seg_test, eval_engine.py, benchmark.py
esp32/tools/ flash.sh, emulate.sh, run_qemu.sh, bench_latency.py, mkimages.py
train/ PyTorch port of NVIDIA's model (nemo_small.py, rnnt_small.py), exporters, GRU LM training
eval/ WER normalization, robustness benchmark builder, laptop baselines
results/ benchmark outputs behind the numbers above
models/ nemo8.tnm (int8 Conformer-CTC Small) and its tokenizer
LICENSE).COMMERCIAL.md.models/ are derived from NVIDIA's stt_en_conformer_ctc_small and remain under
CC-BY-4.0. See NOTICE.Lokutor also has Spanish and other-language models, an int4 profile with more compute headroom, and an on-device TTS for the same chip. Contact us for these.
C
70.0%
Python
27.0%
Shell
2.3%