← BlogFor developers

Why free-tier voice tools cap you at 600 words

Developers aren't typing code anymore. They're writing intent. A design doc for a payment settlement change. A PR description that explains the tradeoffs. A Slack thread walking through a bug. That's long-form writing. Most voice tools were designed when faster typing meant dictating emails. They cap the free tier not because of the model, but because of pricing strategy.

The shift happened quietly. Five years ago, a developer's typing speed was a productivity metric. Wrists hurt, fingers could not keep up with thought. Tools like Dragon NaturallySpeaking promised transcription as speed. They charged accordingly.

But the workflow changed. Now you are typing intent for LLMs. You are telling a code model what to build. You are not describing your existing code. You are specifying new behavior. That demands words. Lots of them. And most existing voice tools crumble under the load.

The workflow shifted, the tools priced for yesterday

When you move from typing code to typing prompts for LLMs, the bottleneck changes. It's no longer how fast your fingers go. It's whether you can keep the thought coherent from voice to Notion to the IDE.

A design doc for a payment settlement change is 1500 words. A PR description that actually explains the tradeoffs is 400 to 600. An incident postmortem thread accumulates over hours. An explanation of a subtle bug in Slack? That's 800 words if you do it right.

Every time you hit a word cap mid-thought, you break flow. You stop dictating. You switch context. You either memorize what you were saying and continue in a new session, or you finish it by typing, which defeats the purpose. The prose fragments. The clarity dissolves.

Marcus, a backend engineer at a Series B fintech, refused to switch to voice dictation for design docs until he found a tool that trusted him. Not because voice is slow. Because he will not route his code or design context through a third-party server. Fintech is regulated. IP matters. Code is IP. A design doc that describes payment settlement logic should not cross a third-party network boundary.

Yet every free tier caps him hard: Wispr at 600 words per month, Superwhisper at 2000, Willow at 2000. He writes a design doc. It is 2400 words. He hits the cap. He either pays for Pro or rewrites the doc in three chunks and stitches it together. Both destroy the writing flow.

Why cloud transcription costs what it costs

Whisper is OpenAI's open speech model, released in September 2022. It has 96% word accuracy on LibriSpeech, a standard audio benchmark. Running it locally on a modern laptop takes under 3 seconds for a minute of audio. There is no variable cost per word. The model file is 140MB. That is the entire technical constraint.

Yet the free tiers meter you like Whisper is scarce. Like the speech model itself is the bottleneck. This is not scarcity. It is pricing strategy.

Cloud services charge for distribution, uptime, and infrastructure. Not for the speech model. When Wispr says "upgrade to Pro for unlimited words," they are not unlocking a different, more sophisticated model. They are unlocking the same model, unmetered. The variable cost to them is near zero. The pricing reflects distribution and business model, not technical scarcity.

This is rational. They have a sales team, a support team, servers running 24/7, compliance costs. Those are real. But they're also not directly tied to whether you speak 600 words or 6000 words into Whisper. The constraint is artificial.

What changes when it runs on your device

Local-first means your code stays on your machine. No third-party server. No transcription logs elsewhere. No word counter. No monthly quota to manage. Your speech-to-text runs in the background, the rough transcript ready in under 3 seconds.

That transcript will be rough. Filler words ("um", "like", "so"). Minor errors where you mumbled. Casual punctuation. Because Whisper captures what you say, not what sounds polished. That is acceptable. That is the trade-off.

For Marcus, this is non-negotiable. He works in Cursor and Claude Code. He is not typing emails. He is building toward products. The thinking he captures in voice is intellectual property. He cannot send it to a cloud server, no matter the privacy policy. Local Whisper solves this structurally. Your machine. Your data. No upload.

The latency matters too. Curl up in Cursor at 11pm, speak a 30-second paragraph about rate limits, get it back in under 3 seconds. That is fast enough to stay in flow. That is fast enough to not interrupt the thinking.

The real trade-off: capture versus rewrite

Local Whisper gives you capture uncapped. It does not rewrite or polish. The first draft is closer to the words that came out of your mouth than to the published prose.

Most cloud tools bundle transcription and polish as one premium feature. They clean up filler, fix common errors, add light punctuation. That polish is genuine LLM work. That requires infrastructure. That is what should actually cost money.

Recitey splits the two:

Free tier: local capture via Whisper on Windows. Uncapped. No metering. Runs on your device. Works across all Windows apps via the clipboard. You get the rough draft fast.

Pro tier: optional cloud rewrite. Send the rough draft, get back polished text in under 2 seconds. Removes filler, fixes errors, adds structure. The thing that actually requires a model on a server.

The first part (capture) costs you almost nothing to run yourself. The second part (rewrite) is what Recitey brings infrastructure for. The pricing should reflect that split. It does not in existing tools.

When this matters and when it doesn't

If you dictate email status updates or Slack messages under 600 words a month, the existing tools are optimized for your case. You spend 3 minutes typing. You will not hit the cap. No point changing.

If you are writing design docs, PR reviews, incident postmortems, or deep technical Slack threads, the cap is friction. You are not trying to dictate faster. You are trying to capture a complete thought in voice because typing 2000 words on a keyboard loses coherence. You lose the narrative thread. You get self-conscious about phrasing. You rewrite it three times.

You need uncapped capture. You probably do not want your code on someone else's infrastructure. You want the draft local, fast, and the rewrite optional when it matters.

The gap is not in the speech model. It is in the pricing architecture and the assumption about who the customer is.

More posts
Keep reading

More like this.

  1. For developers

    The Word Cap That Kills Your Design Doc at Midnight

    You're 30 minutes into a design doc at 11pm. The thought is flowing. Then the app cuts you off: "Free tier limit reached....

  2. For developers

    The Word Cap That Breaks Your Thinking

    Marcus hits the word limit at 2,000 words on his cloud dictation tool. It's mid-design-doc, mid-thought, explaining the...

  3. For developers

    Stop Hitting Word Caps in Your Design Docs

    Marcus is spending 11pm writing a design doc about payment settlement for a new feature. He's speaking clearly, the tool's...

All posts →