Yes—you can classify a small set of spoken words offline on an Arduino Nano 33 BLE Sense. The EloquentTinyML example does this by reducing each short microphone recording to 32 root-mean-square (RMS) sound values, training a small TensorFlow/Keras neural network on a computer, and running the converted model in an Arduino sketch. It is a compact keyword demonstrator, not general dictation or assistant-grade speech recognition.
What this project actually recognizes
You choose a few word classes, collect examples of those words, and train the model to distinguish them. The board does not retain or stream a full recording for cloud processing. After the model is transferred to the board, feature extraction and inference happen locally.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
Nano 33 BLE Sense Rev2 [ABX00069] | $38.70 | Buy on Amazon |
| 2 |
|
Arduino Nano ESP32 with Headers [ABX00083] - ESP32-S3, USB-C, Wi-Fi, Bluetooth, HID Support,... | $19.30 | Buy on Amazon |
That scope matters: the tutorial does not establish continuous-speech recognition, arbitrary vocabulary, speaker-independent performance, or robustness in noisy rooms. It demonstrates classification of the specific words and recording conditions represented in your training data.
How the Nano turns speech into a tiny feature vector
Capture trigger and microphone input
The sampler uses the Nano 33 BLE Sense’s PDM digital microphone. A callback reads microphone data in small batches and calculates an RMS value for each batch. Capture starts when the sound level exceeds a threshold. The tutorial sets that threshold high enough to reduce accidental triggers from room noise and breathing.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute#1 Best Overall
- You can build wearables that use artificial intelligence to recognize movements.
- You can build a room temperature monitoring system that can make suggestions or even make changes to the thermostat settings.
- A gesture or voice recognition device can be created using the microphone or the gesture sensor, taking advantage of the AI capabilities of the card.
Thirty-two values represent one word
Once triggered, the sketch records 32 RMS values at 20 ms intervals. That produces a window of about 640 ms (0.64 seconds), which the tutorial treats as enough time for one spoken word. Each example is therefore a 32-number array, not a stored waveform or spectrogram.
The author says FFT-based experiments were abandoned because the libraries tested caused the sampling program to hang. RMS is less information-rich than frequency-domain features, but it keeps the embedded data path and model simple.
Microphone technique affects the result
For the demonstrated setup, speak close to the microphone and move the board away immediately afterward. That advice is intended to limit breath noise and shows how strongly the trigger and feature values depend on microphone distance and direction.
End-to-end workflow
- Confirm the board and connection. Use a Nano 33 BLE Sense and connect it to a computer with a Micro-B USB cable for programming and power.
- Upload the sampler sketch. The sketch uses the PDM microphone callback, applies the RMS trigger, and prints or otherwise exposes each captured 32-value array.
- Collect labeled examples. Speak each target word under conditions you want the classifier to handle. Keep labels attached to the corresponding arrays; variation in speakers and volume is useful, but the tutorial’s short capture window still assumes one word at a time.
- Train on the computer. A Python script loads the labeled arrays with TensorFlow/Keras and trains a small dense network. The example architecture has dense layers of 32, 12, and 3 units with dropout between layers, for 1,491 total parameters.
- Convert the trained model. The workflow converts the Keras model to TensorFlow Lite and then uses tinymlgen to emit the model as a C array/header suitable for an embedded project. The tutorial’s generated model header is reported as 7,644 bytes.
- Build the classifier sketch. Include the generated header, initialize the TensorFlow Lite runtime used by the tutorial, capture a fresh 32-value RMS sequence, and pass it to the model.
- Run predictions locally. The sketch reports the predicted class after each trigger. No network connection is required during inference.
Model size and what the numbers mean
| Item | Figure in the tutorial | Interpretation |
|---|---|---|
| Input features | 32 RMS values | One value every 20 ms across approximately 640 ms |
| Dense-layer widths | 32, 12, 3 | A small classifier for three output classes in the shown example |
| Total parameters | 1,491 | Model size for this particular architecture |
| Generated model header | 7,644 bytes | Reported embedded C representation for this tutorial model |
These are example-project figures, not limits or benchmarks for every EloquentTinyML voice model. Changing the number of classes, feature window, or network layers changes both memory use and behavior.
How accurate is it?
The tutorial author reports approximately 90% overall accuracy for the collected dataset and setup. That figure is author-reported rather than an independent benchmark, and the author explicitly says it does not account for speaking incorrectly toward the microphone. It should therefore be read as an indication of the demonstration’s performance, not a guarantee across speakers, rooms, distances, or microphones.
An element14 road-test author who followed the project described the Arduino deployment as easy but said installing TensorFlow took time as a first-time user. That report used 60 samples total—20 for each of three words—and is one person’s experience, not a required dataset size or universal setup result.
Rank #2
- Powerful ESP32-S3 Microcontroller: The Arduino Nano ESP32 is powered by the ESP32-S3 chip, featuring a dual-core Xtensa 32-bit LX7 processor running at up to 240 MHz. This high-performance microcontroller offers excellent computational power for IoT, wireless communication, and advanced embedded applications like real-time data processing, voice recognition, and machine learning at the edge.
- Comprehensive Wireless Connectivity: The board supports both Wi-Fi and Bluetooth 5.0, enabling seamless communication with other devices, networks, and cloud platforms. Whether you're building a smart home system, wearable tech, or remote sensors, the Nano ESP32 offers reliable and high-speed connectivity for wireless data transfer and control.
- USB-C for Power and Programming: With the modern USB-C port, the Nano ESP32 ensures faster programming, better power delivery, and a more stable connection compared to traditional micro-USB boards. This makes it easier to work with, especially in development and prototyping stages.
- HID Support for Advanced Applications: The board supports Human Interface Device (HID) profiles, making it ideal for projects that require integration with keyboards, mice, or other HID peripherals. This feature allows you to create custom input devices, virtual controllers, or even USB-based projects that interact directly with computers and other devices.
- MicroPython Compatible: The Arduino Nano ESP32 is compatible with MicroPython, a streamlined version of Python designed for embedded systems. This makes the board perfect for rapid prototyping, educational projects, and developers who prefer Python over C/C++ for ease of use and faster development cycles.
Original Nano 33 BLE Sense versus Rev2
Check the exact hardware revision before copying the tutorial. Arduino marks the original Nano 33 BLE Sense End of Life. Its datasheet identifies an MP34DT05 microphone, a 64 MHz Arm Cortex-M4F processor, omnidirectional sensitivity, and a 64 dB signal-to-noise ratio. The Rev2 datasheet identifies a different microphone, MP34DT06JTR, while retaining a 64 MHz Cortex-M4F.
| Board | Microphone named in the datasheet | Practical implication |
|---|---|---|
| Original Nano 33 BLE Sense | MP34DT05 | The hardware described by the original tutorial; Arduino lists this board as End of Life |
| Nano 33 BLE Sense Rev2 | MP34DT06JTR | Different microphone component; verify the sketch and PDM/library behavior before assuming an unchanged build |
The available documentation does not verify that the original sampler and classifier compile and behave unchanged on Rev2. Treat revision compatibility as a check to perform, not a promise.
RMS features, FFT features, and broader speech recognition
| Approach | Strength | Trade-off |
|---|---|---|
| RMS sequence used here | Very small input and straightforward sampling pipeline | Less acoustic detail; sensitive to volume, timing, and capture geometry |
| FFT or spectrogram features | Represent frequency structure that can separate more sound patterns | More computation, memory, and implementation complexity; the tutorial’s tested FFT libraries caused hangs |
| General speech recognition | Potentially handles larger vocabularies and continuous speech | Requires a substantially different dataset, model, and resource budget than this three-class demonstration |
Troubleshooting the demonstration
The board triggers without a word
- Increase the RMS trigger threshold cautiously so ordinary room noise and breathing remain below it.
- Reduce nearby mechanical or airflow noise.
- Keep the microphone position consistent while collecting examples.
Predictions change when you move the board
- Collect training examples at the distances and angles you expect in use.
- Speak toward the microphone, then move away to reduce breath noise after the word.
- Remember that RMS captures loudness over time, so changes in volume can look like a different class.
The tutorial does not build on Rev2
Check that the selected board in the Arduino IDE, the PDM library, microphone initialization, and any board-specific definitions match Rev2. Because the microphone part differs, do not diagnose this as a model problem until the capture sketch is producing sensible RMS arrays.
TensorFlow setup is difficult
Keep the Python training environment separate from the Arduino project, follow the versions expected by the tutorial’s training script, and verify that the script can load your labeled arrays before attempting model conversion. The reported element14 experience indicates that this computer-side installation can be the slower part for newcomers.
What you need to reproduce it
- An Arduino Nano 33 BLE Sense, with the revision identified before you begin.
- A Micro-B USB cable for connection, programming, and power.
- Arduino software with the board and microphone/PDM support selected.
- A computer with Python, TensorFlow/Keras, and the tinymlgen conversion step used by the tutorial.
- Labeled recordings represented as the sampler’s 32-value RMS arrays.
The Bottom Line
EloquentTinyML makes offline word classification approachable by replacing a full audio pipeline with 32 RMS features and a small dense model. Expect a focused classifier for your chosen words—not general speech recognition—and verify microphone and library compatibility if you use the Nano 33 BLE Sense Rev2.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




