Skip to content
Memory Notch
Memory NotchParakeet vs Whisper, and why we use neither liveMADE FOR YOUR MAC

MEMORY NOTCH BLOG

Parakeet vs Whisper, and why we use neither live

On recorded meetings, Parakeet gets fewer words wrong than Whisper in the vendors’ tests. Both transcribe files; a live transcript needs a streaming model.

Published September 28, 2026 · The Memory Notch team

The short answer

Parakeet makes fewer mistakes than Whisper on recorded meetings, going by each vendor’s own numbers. On a standard meeting test, NVIDIA’s Parakeet TDT v3 (TDT is the name of its decoder design) gets about 11 words in 100 wrong, and OpenAI’s Whisper large-v3 about 16. Moondream’s Parakeet Ultra, released September 22, 2026, cuts errors further. All three transcribe finished files. A transcript that appears during the meeting needs a streaming model instead.

Our verdict, by reader:

Recording in English or another European language? Use Parakeet Ultra.

Speaking a language Parakeet doesn’t cover? Use Whisper.

Short on disk space? Use Parakeet Redux, at 178 megabytes.

Want the words on screen during the call? Use a streaming model.

The table below sums up all four models, checked September 28, 2026.

Parakeet TDT v3 Parakeet Ultra / Redux Whisper large-v3 Nemotron 3.5 streaming speech recognition (ASR)
Made by NVIDIA Moondream, from Parakeet TDT v3 OpenAI NVIDIA
Released August 14, 2025 September 22, 2026 November 6, 2023 June 4, 2026
Size 600M parameters Ultra full size; Redux 178 MB 1.55B parameters 600M parameters
Languages 25 European 25 99 40 language-locales
Built for Whole files Whole files 30-second windows of a file Live audio, in chunks
Licence CC-BY-4.0 CC-BY-4.0 Apache 2.0 OpenMDW-1.1

What changed this week

On September 22, 2026, Moondream released Parakeet Ultra and Parakeet Redux, two speech-to-text models built from Parakeet TDT v3. We call that base model Parakeet v3 from here on. Ultra keeps the full-size weights and is aimed at graphics cards. Redux squeezes the model from 1.2 gigabytes to 178 megabytes so it runs on an ordinary processor, including a laptop’s. Both keep the original Creative Commons licence and the same 25 languages.

Two days later, on September 24, FluidAudio v0.17.3 added both as CoreML models. FluidAudio is an open-source Swift library for running speech models on Apple hardware. With this release, an app can swap either new model in for the old one with one setting. The release notes call Ultra “recommended for new integrations” but keep v3 as the default. Existing apps don’t change until their developers choose to.

All three models are batch models: they transcribe a recording you already have. That matters for the rest of this comparison.

Parakeet vs Whisper: which makes fewer mistakes on meetings?

Parakeet v3 makes about 11 mistakes in every 100 words of recorded meetings, and Whisper about 16, according to each model’s own page. Both figures come from the same meeting test: a public set of recorded business meetings with several people in one room.

The measure is word error rate: the share of words a model gets wrong, leaves out or adds. The lower the rate, the better the transcript.

Test (vendor-reported, from each model card) Parakeet TDT v3 Whisper large-v3
Recorded meetings (AMI) 11.31% 15.95%
Average across the Open ASR Leaderboard’s English test sets 6.34% 7.44%
Multilingual read speech (FLEURS) 11.97% average over 25 languages not listed on the card

A 5-point gap on meeting audio is large. Take an hour-long call of about 9,000 words. Parakeet would get roughly 1,000 of them wrong, and Whisper roughly 1,400. The 1-point gap in the broader average is smaller and could move with how the tests are run. These are the vendors’ numbers, and we have not run our own benchmark of either model.

Two other differences matter before accuracy does.

Languages come first. Parakeet v3 covers 25 European languages, from Bulgarian to Ukrainian, while Whisper covers 99. If your meetings are in Hindi, Japanese or Arabic, only Whisper will do.

Then there is made-up text. Whisper’s own page warns that its output “may include texts that are not actually spoken in the audio input”, and that it can repeat itself.

For more speed, OpenAI offers a turbo version that cuts the decoder from 32 layers to 4. The company says it is “way faster, at the expense of a minor quality degradation.”

Parakeet Ultra and Redux vs Parakeet TDT v3

Ultra makes fewer mistakes than Parakeet v3 on every test Moondream published. Redux trades a little accuracy for a model about a seventh of the size, and it loses the most in background noise.

Test (Moondream-reported) Parakeet TDT v3 Redux Ultra
English, 7 test sets (not named) 6.26% 6.55% 5.80%
25 languages (FLEURS) 11.62% 10.56% 9.55%
Business speech 6.15% 6.96% 5.79%
Background noise 6.72% 9.04% 5.82%
Long recordings (11 TED-LIUM talks) 2.71% 2.51% 1.94%

Moondream’s English figure for the original (6.26%) differs from the 6.34% on the model’s own page because the two use different test sets. Compare within one table, not across them.

For a meeting recorder, the noise row is the one to read. Redux gets about 9 words in every 100 wrong in noisy audio, against about 6 for the other two. A call with a fan, a café or a bad headset microphone sits closer to that row than to clean read speech.

FluidAudio timed its ports on the chip’s machine-learning cores, using clean audiobook recordings. Ultra got 2.13% of words wrong and ran 126.7 times faster than real time. An hour of audio takes about 28 seconds. The model weighs 595 megabytes and needs macOS 14 or later.

Redux got 2.71% wrong at 83.9 times real time, about 43 seconds for an hour. Expect a download of about 220 megabytes and macOS 15 or later, and a first load that spends several minutes compiling. Moondream separately reports Redux on a laptop processor at 38 seconds of audio per second, through its own inference engine. All of these runs used audio much cleaner than a meeting.

Our pick: Ultra if you have the disk space. Redux only if size is the constraint and your audio is clean.

Why does a live meeting transcript need a streaming model?

A live transcript needs a model trained to work on a second or less of audio at a time while remembering what came before. Parakeet and Whisper were trained on whole utterances. Running them live means re-running them over overlapping slices, which wastes work and shifts words as each slice is redone.

Whisper’s page is direct about it: “Whisper models cannot be used for real-time transcription out of the box.” The model reads 30-second windows, and longer files need a chunking algorithm on top. Parakeet v3 was built for files too. Its page says it handles up to 24 minutes in one pass on a data-centre graphics card, and it points to a separate script for streaming.

Some tools build live modes on top of batch models, such as WhisperKit and Moondream’s own engine. They work, but the model underneath still expects to see more audio than it has been given.

A streaming model is built the other way round. Nemotron 3.5, a speech-recognition model from NVIDIA, keeps what it has already heard and processes each new chunk once. The company says this design “eliminates redundant overlapping computations common in traditional ‘buffered’ streaming.”

Your job Model type Examples
Transcribe a recording after the meeting Batch Parakeet TDT v3, Parakeet Ultra or Redux, Whisper large-v3
Read the transcript during the meeting Streaming Nemotron 3.5 ASR Streaming
Both Streaming live, then a second pass See the Memory Notch section below

Nemotron ASR vs Parakeet: what’s the difference?

Nemotron and Parakeet v3 are both 600-million-parameter speech models from NVIDIA, built for different jobs. Nemotron streams live audio in 40 language-locales and detects the language as it goes. Parakeet v3 transcribes finished files in 25 languages.

Nemotron 3.5 ASR Streaming Parakeet TDT v3
Job Live audio, chunk by chunk Finished files
Chunk sizes 80, 160, 320, 560 or 1,120 ms Whole file (up to 24 min in one pass)
Languages 40 language-locales; 19 rated transcription-ready 25 European
Language detection Automatic, per chunk Automatic
English error rate (vendor, FLEURS) 7.91% at 1.12-second chunks FLEURS average 11.97% over 25 languages
Licence OpenMDW-1.1 CC-BY-4.0

The two error rates in that table are not a fair head-to-head. Nemotron’s figure covers one language, and Parakeet’s is an average over 25. Nobody has published both models on the same test in the same mode.

What the model card does publish is capacity. At 1.12-second chunks, it says Nemotron sustains about 6 times as many simultaneous streams as a larger, 1.1-billion-parameter Parakeet, while being roughly half its size. That is a server figure, not a laptop one.

On your own computer the difference is simpler. With Nemotron, words appear a chunk at a time while the person is still talking. With Parakeet, they appear after you stop recording.

Which one should you use?

Pick by when you need the words and which language you speak.

For podcasts, interviews and lectures you already recorded, use Ultra. If your app hasn’t added it yet, Parakeet v3 is close behind.

For recordings in Asian, Middle Eastern or African languages, use Whisper, since Parakeet doesn’t cover them. Read the output for lines nobody said.

On an older or smaller laptop with clean audio, Redux is the practical choice.

For notes you read during the call, or a transcript that is ready the moment the call ends, use a streaming model.

If you are comparing apps rather than models, MacWhisper is built around Whisper for file transcription. Our local AI note taker comparison lists which meeting apps run which models.

What Memory Notch runs, and what it doesn’t do

Our app transcribes with the multilingual streaming version of Nemotron, not Parakeet or Whisper. The model runs on your Mac through FluidAudio, which is bundled in the app. It works in 1.12-second chunks and detects the language for each one. Speaker labels come from a second model, Nemotron 3 Diarization, run once per audio channel.

After a recording, you can ask for an Enhanced transcript. It runs the same model over the saved file in 2.24-second chunks. With no call to keep up with, it can take in more audio at each step.

Transcripts can be wrong, as with every model here, and overlapping voices can be difficult to separate. Speakers are numbered, not named. There are no automatic summaries.

Library search covers titles and metadata, not every word. Recording the computer’s own audio needs screen-recording permission in macOS. And we have not benchmarked any of these models ourselves, so every number in this post is a vendor’s.

Recording, transcription and speaker labels run on your computer. The app makes three network requests, and none carries a recording: a signed list of revoked licences, a once-a-day update check you can turn off, and an update download when you choose one. The offline transcription guide explains what works with no connection. Running the model locally does not replace the consent of the people you record.

You need an Apple Silicon Mac on macOS 26 or later (specs). It costs US$79 once for up to five of your own computers, with a year of updates. You can try it free for recordings of up to 5 minutes, with no account. For why a note taker doesn’t need a bot in the call, see bot-free AI note takers.

Frequently asked questions

Is Parakeet better than Whisper?

On English meeting recordings, Parakeet v3 makes fewer mistakes in the vendors’ published tests. On that meeting test it gets 11.31% of words wrong, against 15.95% for Whisper. Whisper covers 99 languages, so it is still the choice for languages Parakeet doesn’t cover.

Is Parakeet real-time?

Parakeet v3 is a batch model built for finished files. Its maker provides a script to run it in chunks, and apps can build live modes on top, but it was not trained for streaming. Nemotron 3.5 was, with chunks from 80 milliseconds to 1.12 seconds.

What are Parakeet Ultra and Parakeet Redux?

Moondream released both on September 22, 2026, built from Parakeet v3. Ultra is full size and made fewer mistakes than the original on every test Moondream published. Redux is 178 megabytes, made for ordinary processors and laptops, with slightly more errors, most of them in noisy audio.

Parakeet vs WhisperKit: which is faster on a Mac?

WhisperKit is a framework for running Whisper on Apple chips, not a separate model. FluidAudio reports its port of Ultra at 126.7 times real time on the chip’s machine-learning cores. We have not timed the two side by side, so we can’t give you a fair ratio.

What is the difference between Nemotron ASR and Parakeet?

Nemotron and Parakeet are both 600-million-parameter speech models from NVIDIA. Nemotron transcribes live audio in chunks as short as 80 milliseconds, in 40 language-locales. Parakeet v3 transcribes finished files in 25 languages. Memory Notch uses the multilingual streaming version of Nemotron for its live transcripts.

Sources

Related posts

Download Memory Notch · More from the blog