← BlogFor developers

Why Cloud Dictation Tools Have Word Limits (And Local Models Don't)

Cloud-based speech-to-text tools cap your usage because they're paying for every inference. That cap isn't a feature, it's a business model. Once you understand the economics, you see why local transcription changes everything for how developers work.

The Economics Behind the Meter

Every time you dictate to a cloud vendor, your audio travels to their server. A model processes it. The output comes back to you. Each inference has a direct cost. If the vendor is paying for compute, they need to manage variable expenses.

A word cap isn't a design choice. It's a protection mechanism against runaway costs. Wispr is $14/month and limits you to 1,500 words. Willow is $12 and caps at 2,000. Superwhisper is $8.49 with a daily limit. Those numbers exist because the vendor is buying cloud inference. If speech-to-text costs them $0.001 per minute, and a power user dictates 500 minutes monthly, that's $500 in costs against a $14 subscription. The cap protects their margin.

This is how most infrastructure SaaS works. Stripe meters API calls. AWS meters compute. Figma meters collaborators. It's not arbitrary. It's rational business logic.

Why Local Transcription Changes Everything

Whisper-large-v3, the model Recitey runs, achieves 96.3% accuracy on LibriSpeech benchmarks. It's also local. Once downloaded to your machine, the cost to transcribe is zero. No per-inference billing. No cloud infrastructure metering.

That's not a compromise. It's a different constraint set entirely. You trade response latency for infinite usage. You trade cloud polish for local accuracy-first.

When you remove the metering incentive, the product logic shifts. No need to explain word limits. No need to upsell you to premium for basic dictation. The free tier runs Whisper locally and uncapped. You dictate as much as you need.

This also changes what the vendor optimizes for. If cloud tools optimize for cost management, local tools optimize for workflow. That's a fundamentally different design direction.

The Workflow That's Changing

Two years ago, developers dictated quick things. Code review comments. Slack messages. Snippets. Voice was an alternative to typing for short bursts.

Now you're dictating intent.

The rise of AI agents changed what writing means. You're not documenting code. You're specifying intent for a model. That means longer-form thinking out loud. Architecture decisions. API contracts. Incident postmortems. Things that used to be typed design docs now start as voice sketches because voice is faster when you're thinking.

Marcus is a backend engineer at a Series B fintech in Stockholm. He writes design docs by voice. Not bullet points; full architectural reasoning at 11pm after an incident lands. He'd used Wispr for longer sessions, but hit the word cap mid-sentence and lost the thread. He'd finish typing manually. Next morning meant cleanup and rewriting.

He switched to Cursor specifically because tab-complete reduces voice-to-intent rewrites. Fewer typed explanations because Cursor autocompletes the intent. But cloud transcription felt wrong for code IP. So he was stuck: slow typing or cloud tools he didn't trust.

The shift is real: the new bottleneck isn't typing speed. It's prompt clarity. Dictating a spec takes longer to think through than to speak. Local Whisper with no cap means you're thinking out loud, not rationing your speaking against a meter. You follow a thought all the way through, not stop at a cap and context-switch to typing.

The Trade-off: Speed for Accuracy

Local transcription is slower than cloud on first token. A few hundred milliseconds of latency instead of sub-100ms. And raw Whisper output isn't polished. It's accurate, but unrefined. You get "we probably want to async this endpoint" instead of "We should consider making this endpoint asynchronous to reduce request latency."

That's where Pro tier lives. The rewrite engine runs in the cloud, using Claude to polish Whisper's output. Five seconds of waiting, not five minutes of manual cleanup. You get polish when you need it. For raw thinking, you get speed.

The free tier is dictation. The paid tier is polish. The architecture acknowledges they're different needs. The place where you do most of your raw thinking, that's local, uncapped, yours.

What It Means If You Care About Your Code

Local means your code doesn't leave your machine. If you're dictating API calls, business logic, database schemas, or security decisions, that stays on your device.

That's a verifiable privacy property, not a trust claim. A consultant on a client project can't afford business logic in some vendor's logs. A fintech team building settlement systems can't send that context to the cloud. Local transcription is a structural guarantee, not a service promise.

The Pricing Model Tells You the Priorities

If a vendor meters by word, they're optimizing for infrastructure cost. If they don't, they're optimizing for your workflow. The architectural choice, local vs. cloud, isn't neutral. It's a statement about what the vendor believes your actual bottleneck is.

A metered tool says: "You're trying to speak fast." An uncapped local tool says: "You're trying to think clearly." Those are different philosophies that lead to different product shapes. Understanding which philosophy a tool embodies tells you whether it's designed for your workflow or built to protect someone else's margin.

More posts
Keep reading

More like this.

  1. For developers

    Stop Hitting Word Caps in Your Design Docs

    Marcus is spending 11pm writing a design doc about payment settlement for a new feature. He's speaking clearly, the tool's...

  2. For developers

    The design doc that never got written

    It's 11 PM. Marcus, a backend engineer at a fintech in Stockholm, is voice-dictating a design doc for a new payment settlement...

  3. For developers

    Local transcription changes what you can dictate

    You wrote the design doc perfectly on voice, then scrolled up and realized 1400 words in, you'd stopped mid-sentence. Not...

All posts →