The Lokutor team released Oído, a speech‑recognition model that can run on an ESP32‑S3 microcontroller that costs about five dollars. The model is a 13‑million‑parameter NVIDIA Conformer‑CTC Small, quantized to int8 so it fits into the chip’s 8 MB of RAM. In tests on the LibriSpeech benchmark, Oído recorded a word‑error rate (WER) of 3.7 % for the clean set and 8.2 % for the noisy set. Whisper‑tiny, a small open‑source model from OpenAI, scored 6.3 % and 15.9 % on the same data when run on a laptop.
The team also evaluated Oído under real‑world noise conditions – car interior, kitchen, cafeteria, plus background chatter and reverberation. The average WER was 8.4 %, while Whisper‑tiny’s was 12.1 %. A live demo script lets anyone connect a laptop microphone to the ESP32 and try the exact chip arithmetic.
All the code and model weights are released under an open‑source license on GitHub, so hobbyists and developers can download, flash, and experiment with the system themselves.
How does this compare to other on‑device solutions?
Most on‑device speech recognizers need a dedicated neural‑processing unit (NPU) or a more expensive processor. Oído shows that a plain microcontroller with modest RAM can still beat a larger software model running on a full laptop. The result is a tiny, low‑cost device that can understand spoken commands without sending audio to the cloud.
Why it matters
For everyday users, this means cheap devices like smart home gadgets can understand speech locally, keeping data private and reducing reliance on internet connections. Developers can prototype voice‑controlled products without expensive hardware or cloud‑API fees.