Digital filtering on a microcontroller processes sampled data to change its frequency content. It can attenuate sensor noise, reject mains hum, remove drift, isolate vibration bands or prepare data for control and detection. The correct design is a trade-off among attenuation, bandwidth, phase, latency, CPU time, memory, numerical precision and real-time deadlines.
This guide explains the signal-chain constraints, FIR and IIR mathematics, biquad structures, fixed- and floating-point implementation, streaming strategies and the LPC55S69 PowerQuad as a hardware-specific case study.
Where a digital filter belongs
A typical embedded signal path is:
- Physical signal and sensor
- Analog conditioning
- Analog anti-alias filter
- ADC sampling
- Digital filtering
- Control, detection, logging or communications
- Optional DAC and analog reconstruction filter
Sampling at fs gives a Nyquist frequency of fs/2. Components above that limit can alias into the measured band. Once aliasing occurs, software cannot identify or remove the folded component; an analog anti-alias filter and an appropriate sample rate are therefore required before the ADC. Decimation likewise requires low-pass filtering before samples are discarded. Oversampling increases the available transition band, but does not compensate for a poorly designed analog front end, clock jitter or ADC noise.
Specify the response before choosing an algorithm
A filter’s frequency response describes amplitude and phase versus frequency. Define the passband edge, stopband edge, transition width, required stopband attenuation, allowable passband ripple and acceptable group delay. Also inspect impulse and step responses: a design with excellent magnitude attenuation may still ring, overshoot or respond too slowly for a control loop.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11#1 Best Overall
- 2.4GHz Dual Mode WiFi + Bluetooth Development Board
- Support LWIP protocol, Freertos
- SupportThree Modes: AP, STA, and AP+STA
- Ultra-Low power consumption, Compatible with Arduino IDE
- ESP32 is a safe, reliable, and scalable to a variety of applications
- Low-pass: reduces high-frequency noise on accelerometer, temperature or ADC data.
- High-pass: removes slow drift and baseline movement.
- Band-pass: isolates a vibration or audio band.
- Notch: suppresses a known interference such as mains hum.
Filtering cannot separate signal and noise that occupy the same frequencies without additional information. Cutoff frequency must be selected together with sample rate, not independently.
FIR filters: feed-forward and inherently stable
An N-tap finite impulse response filter is:
y[n] = Σ(k=0 to N−1) b[k] x[n−k]
For three taps, y[n] = b0*x[n] + b1*x[n−1] + b2*x[n−2]. The input history is delayed samples, b[k] are coefficients (taps), and the multiply-accumulate result is the output. With no feedback, a finite-coefficient FIR has no pole-instability problem. Symmetric coefficients can provide exactly linear phase, useful when waveform timing matters.
The cost is roughly one multiply and accumulate per tap per output sample, plus storage for the delay line and coefficients. Sharp transitions or narrow bands may require many taps, increasing CPU use, RAM, flash and group delay. Circular buffers, SIMD instructions and optimized libraries avoid the overhead of a naïve loop. ARM’s CMSIS-DSP provides FIR functions for several numeric formats: CMSIS-DSP FIR documentation.
IIR filters and cascaded biquads
An IIR filter feeds back previous outputs as well as delayed inputs. A common second-order section (biquad) is:
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #2
- 2.4GHz Dual Mode WiFi + Bluetooth Development Board
- Support LWIP protocol, Freertos;ESP32 is a safe, reliable, and scalable to a variety of applications
- SupportThree Modes: AP, STA, and AP+STA
- Ultra-Low power consumption, Compatible with Arduino IDE
- 1PCS 30Pin ESP32 Development Board 2.4GHz WiFi Dual Cores Microcontroller Integrated with Antenna RF Low Noise Amplifiers Filters
y[n] = b0*x[n] + b1*x[n−1] + b2*x[n−2] − a1*y[n−1] − a2*y[n−2]
The feedback sign is a convention, not a universal rule: some APIs store the opposite signs. Copy coefficients only after checking the target library’s equation. The source article’s displayed example repeats a1 for both feedback terms; the second coefficient should be an independent a2.
Several biquads are normally cascaded for a higher-order response. Compared with a similar-magnitude FIR, an IIR often needs fewer coefficients, operations and delay elements, making it attractive for low-latency sensor and control work. Feedback also brings risks: poles outside the unit circle cause instability, coefficient quantization moves poles, internal states can overflow, and fixed-point arithmetic can produce limit cycles (a nonzero output with zero input). Phase is generally nonlinear.
Direct-form structures
Direct Form I
Direct Form I stores separate delayed-input and delayed-output histories. It is straightforward to inspect and can offer useful range behavior in some fixed-point designs, at the cost of more delay storage.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsRank #3
- Powerful ESP-32 Board: Unlock the world of Internet of Things (IoT) and advanced electronics with the heart of this kit: the ESP-32 board. It features a powerful dual-core processor, integrated Wi-Fi and Bluetooth 4.2, making it perfect for building connected, smart devices that communicate with your phone or the cloud. It's fully compatible with the Arduino IDE for easy programming.
- Super Starter Kit: This kit contains over 35 different modules and electronic components, including sensors, displays, motors, and input devices. From LEDs and buttons to an OLED screen, servo motor, and keypad, you have everything needed to explore a vast range of projects in one box.
- Step by Step Online Tutorial: Jump right in with our detailed, beginner-friendly tutorial. Access 30+ projects with complete code, clear circuit diagrams, and step-by-step instructions. Learn the fundamentals of electronics, coding, and how to utilize the ESP-32's unique capabilities without any prior experience.
- Hands-on Learning for All Skill Levels: Perfect for students, makers, engineers, and hobbyists. Start with basic circuits and coding, then progress to intermediate and advanced IoT applications. Build practical projects like weather stations, smart home controllers, remote-controlled devices, and interactive gadgets. The skills you learn are the foundation for real-world innovation.
- Quality & Great Support: Elegoo is committed to quality. We provide a clear, detailed tutorial guide, refined code, and a well-organized component kit. All modules are carefully selected for reliability and ease of use. Our dedicated technical support team and active online community are ready to help you succeed in your learning journey.
Direct Form II
Direct Form II combines histories through an intermediate state, reducing delay elements and mapping efficiently to hardware. Its internal state can have a larger dynamic range and greater finite-precision sensitivity, so it is not universally superior.
Transposed forms
Transposed structures rearrange the same mathematics and may improve fixed-point rounding or state-range behavior. Select the structure after considering scaling, coefficient precision and the processor, then validate the actual implementation.
Floating-point or fixed-point?
| Choice | Benefits | Engineering concerns |
|---|---|---|
| Floating point | Simple coefficient handling; broad dynamic range; convenient reference implementation | Check NaNs, infinities, denormals and overflow; speed depends on the MCU’s floating-point hardware |
| Fixed point (Q15/Q31) | Efficient on processors without a suitable FPU; deterministic integer operations | Requires scaling, headroom, saturation and quantization analysis; internal states may overflow even when output does not |
A stable floating-point design is not guaranteed to remain stable after quantization. Quantize coefficients, recalculate poles or measure the quantized response, and test worst-case amplitude. Saturating only the final output cannot repair an accumulator or state overflow.
Generate and validate coefficients
Start with sample rate, filter type, passband and stopband edges, ripple, attenuation, desired phase or delay, numeric format and maximum signal amplitude. Use a recognized design method or coefficient tool rather than guessing taps. The original article discusses filter “cookbooks” and NXP design tooling; regardless of tool, validate the exported coefficients after quantization against the exact target equation and scaling.
Rank #4
- High-performance foundation line, ARM Cortex-M4 core with DSP and FPU, 512 Kbytes Flash, 180 MHz CPU, ART Accelerator, Dual QSPI
- On-board ST-LINK/V2-1 debugger/programmer with SWD connector
- Can be powered from USB
- Three LEDs, Two Push-buttons
- Support of wide choice of Integrated Development Environments (IDEs) including IAR, ARM Keil, GCC-based IDEs
Streaming implementation: samples, blocks and DMA
Sample-by-sample
A callback or interrupt can process each sample immediately, minimizing apparent latency. Function-call and synchronization overhead may be significant at high rates.
Block processing
Libraries and accelerators often process vectors more efficiently. Buffering adds up to a block of latency, and state must persist across blocks. Resetting a biquad for every block creates discontinuities and repeated startup transients.
DMA and ping-pong buffers
ADC DMA can fill one buffer while the CPU or accelerator processes the other. Confirm buffer alignment, source/destination aliasing rules and worst-case execution time; an average benchmark that misses a deadline is not real-time safe.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.LPC55S69 PowerQuad case study
The LPC55S69 includes the PowerQuad DSP/math accelerator and two biquad engines. This is an NXP-specific example, not a requirement for embedded filtering generally. Current MCUXpresso documentation lists floating-point, fixed-16/Q15 and fixed-32/Q31 operations, direct-form-II and cascade functions, plus FIR support: LPC55S69 MCUXpresso SDK API.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Best Value
- with pre-soldered header Raspberry Pi Pico. RP2040 microcontroller chip designed by Raspberry Pi in the United Kingdom
- Dual-core Arm Cortex M0+ processor, flexible clock running up to 133 MHz. 264KB of SRAM, and 2MB of on-board Flash memory.
- Castellated module allows soldering direct to carrier boards. USB 1.1 with device and host support. Low-power sleep and dormant modes. Drag-and-drop programming using mass storage over USB. 26 × multi-function GPIO pins.
- 2 × SPI, 2 × I2C, 2 × UART, 3 × 12-bit ADC, 16 × controllable PWM channels.Accurate clock and timer on-chip.Temperature sensor.
- Accelerated floating-point libraries on-chip.8 × Programmable I/O (PIO) state machines for custom peripheral support
A representative vector flow is:
pq_biquad_state_t state = {
.param = { .a_1 = a1, .a_2 = a2,
.b_0 = b0, .b_1 = b1, .b_2 = b2 }
};
PQ_BiquadRestoreInternalState(POWERQUAD, 0, &state);
PQ_StartVector(input, output, VECTOR_LEN);
PQ_Vector8BiquadDf2F32();
PQ_EndVector();
This is a schematic API pattern, not a drop-in program. Clocking, PowerQuad initialization, headers, coefficient convention, state initialization, alignment and the selected SDK release must be checked. NXP documentation demonstrates eight-sample vector operations, while other functions expose a general blockSize; do not impose a universal multiple-of-eight rule. See the PowerQuad API reference and NXP application note AN13498.
Offload is not free. The CPU still performs setup, buffer transfers, synchronization and memory traffic. Measure end-to-end latency and CPU occupancy for the actual order, block size, memory placement and sample rate. If transfer overhead dominates, a short CPU implementation may be faster and easier to maintain.
Portable software with CMSIS-DSP
For Arm Cortex-M projects that may move between vendors, CMSIS-DSP offers software FIR and biquad APIs without tying the design to PowerQuad. Its available functions, data types and performance depend on the target and build configuration. Consult the CMSIS-DSP filtering index and benchmark on the intended MCU.
A verification plan that catches real failures
- Run impulse and step tests and compare with a trusted desktop reference.
- Sweep frequency and measure passband, transition, stopband and phase or group delay.
- Compare floating-point output with the quantized target implementation.
- Test zero input, maximum amplitude, reset and startup state.
- Check coefficient-sign conventions, state continuity and block-boundary behavior.
- Measure worst-case execution time, CPU load, DMA contention and buffer-overrun handling.
- Inject realistic noise and verify that the filter improves the required signal metric without unacceptable latency.
Which approach fits?
| Need | Typical choice | Reason |
|---|---|---|
| Very simple smoothing | Moving average or exponential smoother | Minimal code and state; limited response control |
| Median rejection of isolated spikes | Median filter | Handles impulsive outliers; not a substitute for frequency-selective design |
| Linear phase or tightly controlled waveform timing | FIR | Stable feed-forward structure and possible linear phase |
| Low latency and low CPU/RAM cost | IIR biquad cascade | Efficient response with careful stability and scaling checks |
| Portable Arm software | CMSIS-DSP | Vendor-independent library option |
| Demonstrated high-throughput bottleneck on LPC55S69 | PowerQuad | Hardware arithmetic can reduce CPU work when transfer overhead is acceptable |
The Bottom Line
Choose the filter from measured signal requirements and latency budget, then validate the quantized, block-integrated implementation on the target hardware. FIR favors predictable phase and stability; IIR biquads favor efficiency; an accelerator is worthwhile only when end-to-end measurements show a real bottleneck.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




