Why your dictation app shouldn't need the cloud
Think about what you say to a dictation app over a month. Messages to your partner. A note to HR. Half a contract. The prompt where you paste in the customer's name and the bug they hit. It is the least filtered stream of text you produce, because you are talking, not composing.
My view is simple: that stream should stay on your computer unless there is a good reason for it not to. For a long time there was a good reason. I don't think there is one any more.
Why did dictation go to the cloud in the first place?
It went to the cloud because good speech recognition used to need a data centre. The models were large, laptops were slow, and sending audio to a server was the only way to get accurate text back quickly.
That changed in two steps. Capable speech models were published openly, first OpenAI's Whisper and then NVIDIA's Parakeet. And Apple Silicon Macs turned out to be fast enough to run them.
The numbers are public. On the Open ASR Leaderboard English average, Parakeet v3 scores a 6.34% word error rate and Whisper large-v3-turbo scores 7.83%. The versions Gabble ships are downloads of about 480 MB and 646 MB. That is smaller than a lot of apps.
What does "cloud dictation" mean for your audio?
It means a recording of your voice is uploaded to a company's servers every time you press the key. The vendors say so plainly.
- Wispr Flow's privacy page states that transcription always happens in the cloud, and that audio is sent to its servers over TLS and deleted once transcribed unless you turn on cloud storage.
- Typeless's privacy page says audio is processed in real time on its cloud servers and discarded once the result is returned.
- Apple's Siri and Dictation privacy page says that when Dictation is not processed on your device, what you dictate is sent to and processed on Apple's servers, and is not stored unless you opt in.
I am not accusing any of them of anything. These are reasonable policies, clearly written. My point is about what you are being asked to do: read a policy, trust that it is followed, and keep trusting it as the policy, the company and its suppliers change.
With a local model you skip that whole exercise. There is nothing to audit, because the audio does not go anywhere.
What do you get from keeping it local?
You get four concrete things.
| Cloud dictation | Local dictation | |
|---|---|---|
| Where your audio goes | The vendor's servers | Nowhere |
| Works with no connection | No | Yes, after the model download |
| Account | Usually tied to a plan | Not needed in Gabble |
| Cost per word to the vendor | Real, every time | None |
That last row matters more than it looks. A cloud service pays for server time on every sentence you speak, so it has to meter you or bill you. A local app has no such cost. Your Mac does the work. I wrote up what that means in money in what dictation subscriptions actually cost over three years.
Offline matters too, and not only on planes. An app with no server cannot have a server outage.
What does the cloud still do better?
Quite a lot, and it would be dishonest to skip it.
- Phones and other platforms. Cloud services run the same on Mac, Windows, iOS and Android. Gabble runs on Apple Silicon Macs with macOS 14 or later and nothing else.
- Nothing to download. A cloud app works the moment you sign in. A local app needs a model on disk first.
- Live text. Some tools show words as you speak. Gabble shows the text after you stop.
- Bigger models. A server can run models no laptop can. Vendors do not all publish which models they use, so I cannot tell you how large that gap is.
- Extras. Meeting recording, speaker labels and team features are common in cloud products. Gabble has none of those.
If you dictate mostly on your phone, a local Mac app is the wrong tool. That is a fine reason to choose a cloud service.
Isn't AI cleanup a cloud feature anyway?
It does not have to be. This is where a lot of "local" claims get fuzzy, so here is exactly how Gabble splits it.
Speech recognition is always on the Mac. So is the default cleanup, which removes ums and repeated words and fixes capitals and spacing. No language model is involved in that.
The rewriting modes, such as Email or Polished, do use a language model. That part is optional and off by default. If you turn it on, you choose where the text goes:
- Ollama on the same Mac. The transcript goes to a local server. Nothing leaves the machine. Setup is in clean up dictation with a local LLM.
- Your own API key, for Anthropic or an OpenAI-compatible service. The transcript text goes to that provider under your account. The key is stored in the macOS Keychain.
In neither case is audio sent. And it is a decision you make, not a default you have to find and switch off.
How can you check a "local" claim yourself?
Turn off Wi-Fi and dictate. If text still appears, recognition is local. It is the only test that does not depend on taking anyone's word for it, including mine.
Then look at what the app does with history. Gabble keeps a searchable history on your Mac, and you choose how long: forever, 30 days, 7 days, or don't save.
The practical setup is in offline dictation on Mac. If you want to compare a local and a cloud app directly, see Gabble vs Wispr Flow. And if you just want to try it, Gabble is a free download.
Frequently asked questions
Is local dictation as accurate as cloud dictation?
It can be very good. On the Open ASR Leaderboard English average, NVIDIA's Parakeet v3 scores 6.34% word error rate and it runs on a Mac. There is no public apples-to-apples benchmark against the private models cloud dictation services use, so test with your own voice.
Why do dictation apps send audio to the cloud?
Running the model on a server means the app works on any device, including phones and older computers, with nothing to download. It also lets the vendor use larger models than a laptop can run.
Is cloud dictation safe?
That depends on the vendor's policies and on what you dictate. Reputable services publish what they retain and for how long. The difference with local dictation is that there is no upload to evaluate in the first place.
Does Apple Dictation send audio to Apple?
Sometimes. Apple says Keyboard settings will tell you if audio is processed on your device. Otherwise what you dictate is sent to Apple's servers.
Does Gabble ever send anything to a server?
Speech recognition never does. Gabble downloads its speech model once from Hugging Face. If you turn on the optional AI cleanup with your own API key, the transcript text goes to that provider, and audio still does not.
What do I give up with a local dictation app?
With Gabble you give up phone and Windows versions, live streaming text, meeting recording, and some disk space for the model. It also needs an Apple Silicon Mac.
Try Gabble
Free dictation for Mac that runs on your machine. No account.