Audio Reactive LED Design Optimization
Key Takeaways:
- The phyllotaxis spiral math is inexpensive; the audio FFT and the serial LED bus take most of the frame time.
- In Phil Schatzmann’s July 2026 benchmark, the ESP32-S3 runs an N=64 32-bit float FFT in 62.62 microseconds, about 16,000 transforms per second.
- WS2812B pixels take roughly 30 microseconds each to update, which limits how many cells you can refresh at a 60 FPS target.
- Precompute LED positions and divergence angles once at boot; fixed-point approximations reduce floating-point overhead in the hot loop.
- Splitting audio analysis and rendering across cores, or across devices, is the common solution for frame-rate stalls.
Where the Frame Budget Actually Goes
On a WLED discourse thread, a user running 768 WS2812B LEDs across three 16×16 panels reported frame rates dropping to 10 FPS on several audio-reactive effects. The spiral math was not the issue. The serial LED bus had already consumed the frame.

A phyllotaxis pattern rendered once to a file involves looping over a few thousand florets, placing each at a radius proportional to the square root of its index, rotating by the divergence angle, and saving. Doing that at a live display refresh rate on a microcontroller with a few hundred kilobytes of RAM presents a different engineering challenge. The spiral arithmetic requires only a few multiplies per cell. The main cost that limits a live system comes from the audio analysis that drives the pattern and the LED bus that outputs it.
On an N=64 complex FFT, the ESP32-S3 completes a 32-bit float transform in 62.62 microseconds, about 16,000 transforms per second, according to Phil Schatzmann’s July 2026 microcontroller FFT benchmark. That benchmark also measures the STM32H7, a 480 MHz Cortex-M7, at 23.36 microseconds, and the RP2350 at 91.78 microseconds. A 16 MHz Arduino Nano is not practical: the benchmark shows it takes 17,165 microseconds per float FFT, fast enough only for a slow volume meter.
A single N=64 float FFT at 62.62 microseconds uses a small portion of a 16.67 millisecond frame, so the transform alone rarely delays a frame. The problem occurs when you combine transforms with windowing, band splitting, and a per-frame render loop on one core, then add a serial LED bus that cannot run in parallel with your compute.
Divergence Angle Precision vs Processing Load
Phyllotaxis is the arrangement of leaves and florets on a plant stem. In the spiral form, the divergence angle between successive organs is about 137.5 degrees, common across flowering plants because it closely matches natural patterns, as explained on Wikipedia’s Phyllotaxis entry. A Python implementation at GeeksforGeeks gives the polar form most LED code uses: r = c * sqrt(n) and theta = n * 137.508 degrees, where n is the floret index and c is a spread constant. That closed form is the Vogel, or Fermat, lattice, where floret n’s position depends only on n.
Precision matters more than it seems. The exact golden angle is irrational, and 137.5 is only an approximation. Plotting n times an irrational angle means the pattern never repeats a radial line, no matter how high n gets. Using a rational approximation causes successive points to return to the same spokes after a fixed number of steps, creating visible seams. The Wikipedia entry notes that when the divergence fraction is a simple ratio such as 3/8, successive leaves line up in vertical rows, while larger Fibonacci pairs produce a non-repeating pattern. That property lets the same lattice scale from a few dozen cells to tens of thousands.
The cost of precision is where floating-point overhead appears. Computing cosf, sinf, sqrtf, and atan2f for every cell every frame is the most expensive part of a naive render on a microcontroller without a fast transcendental unit. Two fixes remove almost all of it: compute the divergence angle and per-cell coordinates once at boot, and when a per-frame angle must vary, use fixed-point or integer approximations instead of float. The 64-bit performance matters too: in Schatzmann’s measurements, ESP32-S3 double-precision math is emulated in software and runs about 11 times slower than 32-bit float, so keep any double work limited to the one-time table build.
From Naive Loops to Precomputed Tables
The first optimization is to stop recalculating positions every frame. LED positions in a physical installation do not move, so their coordinates belong in a lookup table built once at boot. Jagi Natarajan’s audio-reactive phyllotaxis sculpture uses this approach: a lookup table of floating-point positions for every LED “as if it were on unit circle,” which lets per-frame code read a position and skip the trigonometry. That project drives 89 addressable cells from an STM32 blackpill with an INMP441 I2S microphone, according to his write-up on the build.

The alternative to the closed form is the iterative inhibition model, where each new point is pushed away from its neighbors until it settles. It reproduces the same spirals and mimics how a meristem grows, but it costs time proportional to the number of points already placed. Every added floret must be tested against existing ones. On a desktop that is fine; on a microcontroller rendering every frame it is the wrong trade.
Note: The following code is an illustrative example and has not been verified against official documentation. Please refer to the official docs for production-ready code.
// Build the position LUT ONCE at boot. What remains per frame is a sine evaluation for the travelling wave and a multiply for the audio term, both of which a fixed-point sine table handles without float math.
Feeding Spectrum Data Back Into the Pattern
An audio-reactive pattern needs a small number of control values, not a full spectrum. The standard pipeline splits FFT output into bass, mid, and treble band energies and uses those to modulate brightness, spacing, or radial scale. Whether spectrum data can modify the divergence angle itself in real time depends on the effect you want: shifting the angle away from the golden value during a drop produces a visible “unwinding” of the spirals, but it also creates the rational-lock seams described above, so most installations keep the base angle fixed and drive amplitude and phase instead.
Where the spectrum changes per frame, the cost of recalculating positions from scratch is what to avoid. One open-source ESP32-S3 system samples the microphone at 1024 samples around 20.5 kHz, runs an FFT with Hann windowing, and splits the spectrum into bass, mid, and treble energies on core 0, while a second FreeRTOS task on core 1 handles rendering and a 16×2 LCD, according to the project’s GitHub repository. Analysis and render never compete for the same frame budget. Incremental FFT updates, where only the newest samples enter the transform, are the natural next step when the transform itself becomes the bottleneck.
Efficient LED Update Routines
The LED bus usually limits performance. WS2812B pixels are addressed serially, one bit at 1.25 microseconds, so a 24-bit pixel takes about 30 microseconds to update, according to the Arduino forum thread on maximum WS2812 refresh rate. That thread also calculates for a 256-pixel string: 30 microseconds times 256 pixels gives 7.68 milliseconds for a full refresh.
The 256-pixel figure surprises many. A single refresh of a 256-LED string takes 7.68 milliseconds, about 46 percent of a 16.67 millisecond frame budget, before your own logic runs. At 768 pixels the bus alone takes roughly 23 milliseconds, longer than one frame. That matches the WLED user’s report of 10 FPS on some audio-reactive effects, according to the WLED discourse thread. The LED bus, not the math, limits the frame rate.
| Stage | Platform | Reported cost | Source |
|---|---|---|---|
| N=64 float FFT | ESP32-S3 @ 240 MHz | 62.62 microseconds | Schatzmann, July 2026 |
| N=64 float FFT | STM32H7 @ 480 MHz | 23.36 microseconds | Schatzmann, July 2026 |
| N=64 float FFT | RP2350 @ 150 MHz | 91.78 microseconds | Schatzmann, July 2026 |
| N=64 float FFT | Arduino Nano @ 16 MHz | 17,165 microseconds | Schatzmann, July 2026 |
| LED bus update, per pixel | WS2812B (any host) | 30 microseconds | Arduino forum, WS2812 timing |
| LED bus update, 256 pixels | WS2812B (any host) | 7.68 milliseconds | Arduino forum, WS2812 timing |
Two techniques reduce the transfer time. Batching the whole frame into one DMA-driven write frees the CPU while the bus clocks out, and pushing only changed cells reduces the byte count when most of the display is static. Some installations avoid the serial bus entirely with parallel-output panels, but that changes the wiring and the driver library and is not a drop-in replacement for a WS2812B string.
Microcontroller Choice and Refresh Rate
The FFT numbers make the ESP32-S3 the clear choice for audio analysis. Nearly 16,000 transforms per second provide enough headroom to run a transform far more often than any display refreshes. FFT throughput and refresh rate differ, though. The LED bus serializes regardless of how fast the host computes, so a chip that finishes its math in 62 microseconds still waits 7.68 milliseconds to push 256 pixels. A faster microcontroller improves the analysis stage and has little effect on the transfer stage.
WLED’s documentation clarifies the hardware trade-offs. Audio-reactive support runs across the ESP family with differences per variant: the classic ESP32 supports digital and analog microphones, the ESP32-S3 supports I2S digital and PDM microphones only, the ESP32-S2 supports I2S digital only, and the ESP8266 has no microphone input and can only receive synced audio, according to the WLED audio-reactive documentation. That documentation recommends I2S digital microphones such as the INMP441 for clean signal and warns that the ESP32’s built-in ADC is poorly suited to audio and easily disturbed by power-supply noise.
If your display has fewer than roughly 200 pixels, the bus is not your bottleneck and a faster FFT does not help. If it has more than 500 pixels, no microcontroller upgrade fixes a serial bus that needs more than a frame to finish, and the best solution is fewer pixels, a parallel protocol, or a lower target frame rate.
A 256-LED Setup
A 256-pixel WS2812B grid is a useful test case because the bus cost is known and significant. At 30 microseconds per pixel the raw refresh takes 7.68 milliseconds, leaving about 9 milliseconds of a 16.67 millisecond frame for analysis and rendering. A naive implementation that recalculates every cell’s polar coordinates each frame, running sqrtf and atan2f 256 times, spends most of that remaining budget on trigonometry that never changes. Moving the coordinates into a boot-time LUT and driving per-frame brightness from three band energies reduces the render loop to a table read and a sine evaluation per cell.
Another improvement is the dual-core split. Assigning audio sampling and FFT to one core and rendering to the other removes the scheduling contention that causes periodic frame drops when a transform occurs mid-render. An open-source ESP32-S3 system uses this setup, with the audio task sampling at 1024 samples around 20.5 kHz and the LED task handling rendering and an LCD on the second core, according to the project repository.
A third option removes on-device analysis altogether. PatternFlow, an open-source ESP32-S3 LED synthesizer, captures audio from a browser tab, runs FFT across four frequency bands there, and sends the band values to the device over WebSockets, so any knob-driven pattern becomes audio-reactive without extra firmware code. It has its own limitation: the pattern stops responding if the Wi-Fi drops, and the analysis never runs on the microcontroller, according to the PatternFlow project page. WLED offers the same split officially, multicasting audio data from a sender to receivers at UDP address 239.0.0.1 port 11988 so ESP8266 nodes without microphones can still participate, according to the WLED documentation.
The Hackaday retrospective on a widely copied audio-reactive strip project describes the remaining challenge: the hard problem is interfacing with human perception rather than producing a technically correct visualization, and making a system work well across all musical genres remains unsolved, according to its April 2026 coverage. Frame rate is necessary for a responsive installation but not sufficient. The number that predicts whether a display feels alive is end-to-end audio-to-photon latency, and none of the projects above publish it.
The two reference repositories for the phyllotaxis pattern itself, jagnat/esp32_phyllotaxis and its companion jagnat/phyllotaxis-sketch-sdk, are both small, recent projects, so they serve better as a working reference than a widely used library. Anyone building a phyllotaxis installation today starts from that code, the WLED audio-reactive stack, and a boot-time LUT.
Sources and References
Sources cited while researching and writing this article:
- Microcontroller FFT & IFFT Performance Benchmark (N=64)
- Phyllotaxis – Wikipedia
- Phyllotaxis pattern in Python | A unit of Algorithmic Botany
- Phyllotaxis: An audio-reactive LED display – Jagi Natarajan
- GitHub – adityabhagwani/Embedded-Audio-Visualization-System: Real-time …
- Frame rate issues on WLED 14.0 Audio reactive using ESP32
- Audio Reactive WLED – WLED Project – GitHub
- PatternFlow – Open Source Embedded Project
- Audio Reactive LED Strips Are Hard – Hackaday
- GitHub – jagnat/esp32_phyllotaxis · GitHub
- jagnat/phyllotaxis-sketch-sdk
Rafael
Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...
