Oído int4: 8.3 MB speech recognition for the ESP32-S3
Animated demo: the transcripts are Oído's output (sped up). Footage from a physical board is coming.
¡Oído! is Spanish kitchen slang for heard, got it.
This is the compact profile of Oído: open-vocabulary English speech recognition
that runs entirely on an ESP32-S3, with no cloud and no NPU. At 8.3 MB it leaves a 6 MB app partition free on a 16 MB
flash module for your own application code (esp32/firmware/partitions_nemo4.csv).
| LibriSpeech WER (%) | test-clean | test-other | Size |
|---|---|---|---|
| This model (int4, greedy, on-chip arithmetic) | 4.61 | 9.98 | 8.3 MB |
| Oído int8 | 3.70 | 8.23 | 14.0 MB |
| Espressif MultiNet7 on the same chip (ESP-SR benchmark) | 8.5 | 21.3 | 2.9 MB |
On a board (ESP32-S3-WROOM-1-N16R8, one core at 240 MHz) the int4 model runs at 1.76× real time (1.73–1.83 over 10 clips), about 11% faster than int8 (1.97×). That is not real time yet: the two-core mode is not correct on silicon. On those 10 clips the board's transcript equals the host build's on 8 (the laptop's math library rounds the last bit differently); both score 8.33 % WER on them. Details: github.com/lokutor-ai/oido.
Need lower latency? The streaming variant shows text while you speak; on the board its final text appeared 4.1–11.1 s after clips of 3.6–10.6 s ended (about half the wait of utterance mode), at some cost in accuracy.
Which Oído model?
One model per language; the models that stream also run in full-context (utterance) mode.
| Model | Language | Modes | Size | License |
|---|---|---|---|---|
| oido-ctc-small-int8 | English | utterance only (best accuracy) | 14.0 MB (+1.3 MB LM) | CC-BY-4.0 |
| oido-ctc-small-int4 | English | utterance only (smallest) | 8.3 MB | CC-BY-SA-4.0 |
| oido-ctc-small-stream-int8 | English | streaming + full-context (low latency) | 14.0 MB | CC-BY-SA-4.0 |
| oido-es-ctc-small-int8 | Spanish | streaming + full-context in one file | 14.0 MB (+1.3 MB LM) | CC-BY-4.0 |
Streaming-capable files stream automatically in the firmware's microphone mode and in live_demo.py.
Use
git clone https://github.com/lokutor-ai/oido && cd oido
esp32/tools/flash.sh /dev/ttyUSB0 models/nemo4.tnm # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone
How it was made
We took NVIDIA's stt_en_conformer_ctc_small (CC-BY-4.0)
and fine-tuned it with 4-bit quantization-aware training for 8,000 steps. Linear layers inside the Conformer blocks use
int4 per-channel weights; the front end and output head use int8; activations and attention are int8. The training
used public corpora:
- Common Voice 17 and VoxPopuli (CC0);
- LibriSpeech, MLS English, AMI and VCTK (CC-BY-4.0);
- People's Speech, with transcripts regenerated by NVIDIA parakeet-tdt-0.6b-v2;
- OpenSLR 70 and 83 (CC-BY-SA-4.0).
License
This model is released under CC-BY-SA-4.0: it is derived from NVIDIA's CC-BY-4.0 model and trained on data that includes share-alike sources. The Oído engine and firmware are GPLv3, with commercial licenses available from Lokutor.
Model tree for lokutor-ai/oido-ctc-small-int4
Base model
nvidia/stt_en_conformer_ctc_smallDatasets used to train lokutor-ai/oido-ctc-small-int4
MLCommons/peoples_speech
facebook/voxpopuli
Evaluation results
- WER (on-chip int4 arithmetic, greedy) on LibriSpeech (clean)test set self-reported4.610
- WER (on-chip int4 arithmetic, greedy) on LibriSpeech (other)test set self-reported9.980