Oído int4: 8.3 MB speech recognition for the ESP32-S3

Animated demo: the transcripts are Oído's output (sped up). Footage from a physical board is coming.

¡Oído! is Spanish kitchen slang for heard, got it.

This is the compact profile of Oído: open-vocabulary English speech recognition that runs entirely on an ESP32-S3, with no cloud and no NPU. At 8.3 MB it leaves a 6 MB app partition free on a 16 MB flash module for your own application code (esp32/firmware/partitions_nemo4.csv).

LibriSpeech WER (%) test-clean test-other Size
This model (int4, greedy, on-chip arithmetic) 4.61 9.98 8.3 MB
Oído int8 3.70 8.23 14.0 MB
Espressif MultiNet7 on the same chip (ESP-SR benchmark) 8.5 21.3 2.9 MB

On a board (ESP32-S3-WROOM-1-N16R8, one core at 240 MHz) the int4 model runs at 1.76× real time (1.73–1.83 over 10 clips), about 11% faster than int8 (1.97×). That is not real time yet: the two-core mode is not correct on silicon. On those 10 clips the board's transcript equals the host build's on 8 (the laptop's math library rounds the last bit differently); both score 8.33 % WER on them. Details: github.com/lokutor-ai/oido.

Need lower latency? The streaming variant shows text while you speak; on the board its final text appeared 4.1–11.1 s after clips of 3.6–10.6 s ended (about half the wait of utterance mode), at some cost in accuracy.

Which Oído model?

One model per language; the models that stream also run in full-context (utterance) mode.

Model Language Modes Size License
oido-ctc-small-int8 English utterance only (best accuracy) 14.0 MB (+1.3 MB LM) CC-BY-4.0
oido-ctc-small-int4 English utterance only (smallest) 8.3 MB CC-BY-SA-4.0
oido-ctc-small-stream-int8 English streaming + full-context (low latency) 14.0 MB CC-BY-SA-4.0
oido-es-ctc-small-int8 Spanish streaming + full-context in one file 14.0 MB (+1.3 MB LM) CC-BY-4.0

Streaming-capable files stream automatically in the firmware's microphone mode and in live_demo.py.

Use

git clone https://github.com/lokutor-ai/oido && cd oido
esp32/tools/flash.sh /dev/ttyUSB0 models/nemo4.tnm      # ESP32-S3-DevKitC-1 N16R8 + INMP441 microphone

How it was made

We took NVIDIA's stt_en_conformer_ctc_small (CC-BY-4.0) and fine-tuned it with 4-bit quantization-aware training for 8,000 steps. Linear layers inside the Conformer blocks use int4 per-channel weights; the front end and output head use int8; activations and attention are int8. The training used public corpora:

  • Common Voice 17 and VoxPopuli (CC0);
  • LibriSpeech, MLS English, AMI and VCTK (CC-BY-4.0);
  • People's Speech, with transcripts regenerated by NVIDIA parakeet-tdt-0.6b-v2;
  • OpenSLR 70 and 83 (CC-BY-SA-4.0).

License

This model is released under CC-BY-SA-4.0: it is derived from NVIDIA's CC-BY-4.0 model and trained on data that includes share-alike sources. The Oído engine and firmware are GPLv3, with commercial licenses available from Lokutor.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lokutor-ai/oido-ctc-small-int4

Finetuned
(5)
this model

Datasets used to train lokutor-ai/oido-ctc-small-int4

Evaluation results