Oído: Conformer-CTC Small, int8, for the ESP32-S3
1
3 commits
1 linked in READMEs
updated Sep 30, 2026
¡Oído! is Spanish kitchen slang for heard, got it.
This is open-vocabulary English speech recognition that runs entirely on an ESP32-S3 microcontroller (240 MHz
dual-core, 8 MB PSRAM, 16 MB flash), with no cloud and no neural accelerator. The file nemo8.tnm is NVIDIA's
stt_en_conformer_ctc_small (13 M parameters), quantized to
int8 and packed for Lokutor's on-chip engine.
Code, firmware and tools: github.com/lokutor-ai/oido (GPLv3, commercial licenses available).
| LibriSpeech WER (%) | test-clean | test-other |
|---|---|---|
| This model on the ESP32-S3 engine (int8, greedy) | 3.70 | 8.23 |
| Original fp32 model (PyTorch) | 3.68 | 8.11 |
| Espressif MultiNet7 on the same chip (ESP-SR benchmark) | 8.5 | 21.3 |
| Whisper tiny.en, fp32 on a laptop (500-utterance subsets) | 6.3 | 15.9 |
Status (30 September 2026). The WERs are computed with the host build of the firmware engine, whose arithmetic is bit-identical to the chip's. Firmware transcripts under Espressif's QEMU emulator match it word for word. Real-time speed is estimated at 0.7–0.95× real time from exact instruction counts; measurements on physical boards are coming.
git clone https://github.com/lokutor-ai/oido && cd oido
esp32/host/tasr_cli models/nemo8.tnm recording.wav # laptop, chip-exact arithmetic (after `make` in esp32/host)
esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone
nemo8.tnm: int8 weights in the TNM1 format (14.0 MB), produced by train/export_nemo.py in the GitHub repository.tokenizer.model: the original SentencePiece tokenizer (1024 BPE pieces).Derived from NVIDIA's stt_en_conformer_ctc_small, licensed CC-BY-4.0. Changes: re-implemented in PyTorch, quantized
to int8 and repacked. This model is also released under CC-BY-4.0. The engine and firmware are GPLv3, with commercial
licenses from Lokutor.
Oído: Conformer-CTC Small, int8, for the ESP32-S3
1
3 commits
1 linked in READMEs
updated Sep 30, 2026
¡Oído! is Spanish kitchen slang for heard, got it.
This is open-vocabulary English speech recognition that runs entirely on an ESP32-S3 microcontroller (240 MHz
dual-core, 8 MB PSRAM, 16 MB flash), with no cloud and no neural accelerator. The file nemo8.tnm is NVIDIA's
stt_en_conformer_ctc_small (13 M parameters), quantized to
int8 and packed for Lokutor's on-chip engine.
Code, firmware and tools: github.com/lokutor-ai/oido (GPLv3, commercial licenses available).
| LibriSpeech WER (%) | test-clean | test-other |
|---|---|---|
| This model on the ESP32-S3 engine (int8, greedy) | 3.70 | 8.23 |
| Original fp32 model (PyTorch) | 3.68 | 8.11 |
| Espressif MultiNet7 on the same chip (ESP-SR benchmark) | 8.5 | 21.3 |
| Whisper tiny.en, fp32 on a laptop (500-utterance subsets) | 6.3 | 15.9 |
Status (30 September 2026). The WERs are computed with the host build of the firmware engine, whose arithmetic is bit-identical to the chip's. Firmware transcripts under Espressif's QEMU emulator match it word for word. Real-time speed is estimated at 0.7–0.95× real time from exact instruction counts; measurements on physical boards are coming.
git clone https://github.com/lokutor-ai/oido && cd oido
esp32/host/tasr_cli models/nemo8.tnm recording.wav # laptop, chip-exact arithmetic (after `make` in esp32/host)
esp32/tools/flash.sh /dev/ttyUSB0 models/nemo8.tnm # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone
nemo8.tnm: int8 weights in the TNM1 format (14.0 MB), produced by train/export_nemo.py in the GitHub repository.tokenizer.model: the original SentencePiece tokenizer (1024 BPE pieces).Derived from NVIDIA's stt_en_conformer_ctc_small, licensed CC-BY-4.0. Changes: re-implemented in PyTorch, quantized
to int8 and repacked. This model is also released under CC-BY-4.0. The engine and firmware are GPLv3, with commercial
licenses from Lokutor.