Whisper vs Parakeet: which speech model on a Mac?

By · Updated · 4 min read

Whisper and Parakeet are the two speech model families most local Mac dictation apps use. Whisper comes from OpenAI and Parakeet from NVIDIA. Both are openly published, and both can run on Apple Silicon. Here is how they compare on the numbers their own model cards publish.

Which is more accurate, Whisper or Parakeet?

On the standard English benchmark average, Parakeet has the lower error rate. The Open ASR Leaderboard averages word error rate (WER) across several English test sets. Lower is better.

Model Size Languages Mean WER LibriSpeech clean / other Source
NVIDIA Parakeet TDT 0.6B v2 600 M parameters English 6.05% 1.69% / 3.19% model card
NVIDIA Parakeet TDT 0.6B v3 600 M parameters 25 European 6.34% 1.93% / 3.59% model card
OpenAI Whisper large-v3-turbo 809 M parameters 99 7.83% not on card model card

A few caveats, because these numbers get quoted carelessly:

  • They are English benchmark sets. They say nothing about how either model does in German or Polish.
  • They come from batch inference on GPUs, not from live dictation on a Mac. An app that runs a Core ML conversion of the model will not reproduce them exactly.
  • The Whisper turbo figure sits in the model page's evaluation metadata, sourced from the leaderboard and marked as community-submitted. The Parakeet v3 card's own comparison table shows the same 7.83% figure.
  • A gap of 1.5 points is about one or two extra wrong words per 100. You may or may not notice it.

Where does Parakeet do well and badly?

Parakeet is strongest on clean, prepared speech and weakest on messy meeting audio. The per-dataset numbers on the model cards show the spread.

Test set Parakeet v2 WER Parakeet v3 WER
SPGI (earnings calls, read style) 2.17% 3.97%
TEDLIUM-v3 (talks) 3.38% 2.75%
VoxPopuli (parliament speech) 5.95% 6.14%
GigaSpeech 9.74% 9.59%
Earnings-22 11.15% 11.42%
AMI (meetings) 11.16% 11.31%

Dictation is closer to the top of that table than the bottom. You are one speaker, close to a microphone, talking on purpose.

What is the catch with each model?

Parakeet's catch is language coverage, and Whisper's catch is that it can invent text.

Parakeet. Version 2 is English only. Version 3 covers 25 European languages: Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian, Spanish, Swedish, Russian and Ukrainian (model card). If your language is not on that list, Parakeet is not an option. Both versions add punctuation and capitalisation themselves.

Whisper. It covers 99 languages, but OpenAI's own model card warns that it can produce text that was not actually spoken, that accuracy is uneven across languages and accents, and that it can produce repetitive output. The turbo model is a pruned and fine-tuned large-v3 with the decoding layers cut from 32 to 4, which is where its speed comes from.

Which models does Gabble ship?

Gabble ships two Parakeet models and five Whisper models. Parakeet runs through FluidAudio and Whisper through WhisperKit, both on Core ML. Each model is downloaded once from Hugging Face.

Model in Gabble Family Download Languages Published mean WER
Parakeet v3 (default) Parakeet about 480 MB 25 European 6.34%
Parakeet v2 Parakeet about 480 MB English 6.05%
Whisper Tiny Whisper about 77 MB Multilingual no verified figure
Whisper Base Whisper about 147 MB English no verified figure
Whisper Small Whisper about 216 MB 99 no verified figure
Distil Large v3 Whisper about 595 MB English no verified figure
Whisper Large v3 Turbo Whisper about 646 MB 99 7.83%

The WER column is the published figure for the original model, not a measurement of Gabble. Gabble has no benchmark of its own yet, so none is printed here. Where there is no figure from a checked source, the table says so.

Which one should you pick?

Start with the default and only change it if you have a reason.

Your situation Pick
English, or one of the 25 European languages Parakeet v3
English only, and you want the lowest published WER Parakeet v2
A language outside Parakeet's 25 Whisper Large v3 Turbo
A language outside Parakeet's 25, and little disk space Whisper Small
Very little disk space, quick notes Whisper Tiny
You want the recogniser biased towards your own terms Any Whisper model, with the Dictionary filled in

That last row matters if you dictate a lot of names or jargon. Gabble's Dictionary fixes spellings after transcription with every model. With Whisper models it also biases the recogniser towards your terms. The details are in teach dictation your names and jargon.

How do you switch models?

You pick a model in Gabble and it downloads the first time you select it. After that, switching does not need a connection. The best test is your own voice: dictate the same paragraph with two models and compare the output. It takes two minutes and tells you more than any leaderboard.

If you have not installed it yet, download Gabble. It is free and needs an Apple Silicon Mac on macOS 14 or later. For the offline setup, see offline dictation on Mac.

Frequently asked questions

Is Parakeet more accurate than Whisper?

On the Open ASR Leaderboard English average, yes. Parakeet v2 scores 6.05% word error rate and v3 scores 6.34%, while Whisper large-v3-turbo scores 7.83%. Those are English benchmark sets, so results in other languages and in real dictation will differ.

Which is better for dictation on a Mac, Whisper or Parakeet?

Parakeet v3 is a good default if your language is one of its 25 European languages. Use Whisper Large v3 Turbo or Whisper Small if you need one of the other languages Whisper supports.

How many languages does Parakeet support?

Parakeet v2 is English only. Parakeet v3 supports 25 European languages, including English, German, French, Spanish, Italian, Polish, Russian and Ukrainian.

How many languages does Whisper support?

OpenAI's Whisper large-v3-turbo model card lists 99 languages. Accuracy varies a lot between them.

What is word error rate?

Word error rate, or WER, is the share of words a speech model gets wrong against a reference transcript, counting substitutions, deletions and insertions. A 6% WER means roughly 6 errors per 100 words.

Can I use a custom dictionary with Parakeet?

In Gabble the Dictionary's 'heard as' replacements work with every model because they run on the text afterwards. The extra step of biasing the recogniser towards your terms applies to Whisper models only.

Try Gabble

Free dictation for Mac that runs on your machine. No account.

Download for Mac