Desert Ant Labs

Voz and Clear are faster on Mac

The Voz and Clear model cards, fanned out side by side. Voz shows a person in pale green sunglasses against teal and coral sound waves; Clear shows a woman standing still in a blurred crowd.

Voz, our speech-to-text model, now transcribes audio up to 2.9x faster on Mac, with the same transcript as before. Clear, our audio enhancement model, is faster on every Mac we tested. On an M4 iPad Air, Voz is 20% faster. We launched both models two weeks ago, and this is the first round of speedups.

On an M3 Ultra, Voz transcribes 10 minutes of audio in 0.9s, down from 2.7s, and Clear cleans up those 10 minutes in 1.6s, down from 2.0s. Both updates shipped in version 3.2.0 of the Swift SDK. Update to the latest release, 3.3.0, and you get them with no code changes.

Before this update, Voz ran faster on an iPhone 16 Pro than on an M3 Ultra Mac. The Mac has more processing power and more thermal headroom to keep that power running, but the way we scheduled the work wasn’t using either.

A lot of the work was happening in sequence, so some parts of the chip were waiting while others were busy. We’ve changed how Voz and Clear run on Mac and iPad hardware so more of that work can happen at the same time.

Running more work in parallel

Voz transcribes in two stages. The encoder turns short windows of audio into a representation of the speech. The decoder turns that representation into words and timestamps.

The decoder now runs on its own thread. As soon as the encoder finishes a window, the decoder starts turning that window into words while the encoder moves on to the next window. On M-series chips, decoding runs on the CPU, leaving the Neural Engine available for the encoder. Voz also encodes four windows at a time instead of one.

Clear had a similar constraint. Calls into a Core ML session shared one set of input buffers, so each run had to finish before the next could use them. Each run now gets its own buffers from a pool, so independent runs can use the session at the same time.

The results

We compared the new and previous builds using the same 10-minute speech recording on each machine. The tables show realtime speed: audio duration divided by processing time. Higher is faster.

We also compared Voz’s transcripts with the previous version on full-length recordings. Every transcript in those tests was identical, word for word.

Voz on Mac and iPad, realtime speed on 10 minutes of audio

Chip Before After
M1 223x 228x
M5 357x 457x
M3 Ultra 221x 640x
M4 iPad Air 359x 431x

At 640x realtime, the M3 Ultra transcribes 10 minutes of audio in 0.9s, down from 2.7s.

Clear on Mac, realtime speed on 10 minutes of audio

Chip Before After
M1 208x 245x
M5 413x 425x
M3 Ultra 294x 382x

The M3 Ultra benefits most from the extra parallel work. Voz is 28% faster on the M5 and 2% faster on the M1, while Clear improved in every test. On the iPhone 16 Pro, Voz is unchanged at 315x realtime, or 10 minutes of audio in 1.9s.

Ready for iPhone 18 Pro

Apple’s new A20 Pro chip in the iPhone 18 Pro has a dual 16-core Neural Engine, which Apple says provides twice the compute for on-device AI.

Voz and Clear now spread their work across more of the chip, so we expect similar gains on the iPhone 18 Pro. We’ll publish the numbers once we’ve measured them.

Try Voz and Clear in your own products

Both models are in the Swift SDK. Clear is also in the Kotlin and JavaScript SDKs, and Voz is coming to both. The Voz and Clear model pages have the full specs.

vozclearmacperformanceon-device
All posts