← BlogFor developers

The word cap mid-sentence is the worst moment

You're dictating a design doc at 11pm. The payment settlement team needs to understand the edge case you just discovered, and speaking it out is faster than typing, you're three paragraphs in, the logic is flowing, and then you hit the limit. Wispr Flow caps you at 5,000 words on the free tier. Superwhisper stops at 2,000 a day. You're mid-sentence. You copy what you have, open a new session, and the thread breaks. By morning, the doc needs stitching back together, and half the clarity is gone.

The bottleneck shifted

Five years ago, the work was typing code. Today, the work is typing intent. You're writing a PR description that explains what the model should build. You're drafting a design doc that specifies the edge case. You're in Slack explaining a bug investigation so the team doesn't have to spin up the logs themselves.

The throughput demand moved from commands to explanation. Your code-writing speed plateaued years ago (IDE autocomplete solved that). Your explanation-writing speed didn't. Your mouth runs at 150 words a minute. Your fingers run at 60. The difference matters most when you're translating a technical thought into words, you're thinking and speaking at the same time.

Most voice tools were built for the old frame: capture a voice memo, transcribe a meeting. They assume a sentence or two per session. They cap the free tier at a number that sounds generous, 5,000 words!, until you realize that's two design docs or one long Slack thread. Then they ask you to pay.

Cloud dictation has a structural problem

Wispr Flow charges $14 a month. Superwhisper is $8.49. Willow is $12. All of them cap the free tier.

That pricing makes sense if you're paying for cloud compute. Audio travels to a server, gets transcribed, comes back. Each transcription costs money. The business model is: usage-based pricing, and free tier caps as a funnel.

But most developers know what Whisper is. The large model (Whisper-large-v3) runs locally. The marginal cost of transcribing another 10,000 words is electricity. Maybe a penny. Definitely not a dollar a word.

The structural choice, cloud vs local, changes the entire business model. Cloud dictation needs to cap you to stay profitable. Local dictation doesn't.

Recitey runs Whisper locally on your device. No word limit on the free tier. No meter. No "upgrade to continue" interruption. The latency flip matters too. Cloud transcription adds 500ms to a second of round-trip delay. You speak, you wait for the server, text appears. Local transcription finishes while you're still reading what you said.

The moment it clicks

Marcus, a backend engineer at a fintech in Stockholm, refused to use Wispr or Superwhisper because of IP concerns. Code samples, incident postmortems, design decisions, they're not going to a cloud service. He uses Cursor for the autocomplete, not because it's marginally faster to type, but because fewer rewrites means fewer seconds lost in the thinking-to-code loop.

When he switched to local voice, the design doc at 11pm became a single unbroken session. Eighteen minutes start to finish. No cap. No interruption. No next-morning stitching and clarifying what yesterday-Marcus meant.

The accuracy of Whisper-large-v3 on technical language is high enough that Cursor's tab-complete catches the meaning; maybe 8% of sentences need a rewrite. That's less friction than stopping mid-thought to switch modes.

The reframe is: voice isn't a replacement for typing. It's a new input mode for the new workflow. Developers are writing more intent per day and more explanation per sprint. Voice scales that without the word-limit tax.

Who this is actually for

This matters most for people writing specifications. Product specs. Design docs. PR descriptions. Bug investigations. Postmortems. Technical RFCs. The long-form explanation that context-switching kills.

It matters less if you're dictating voice memos or transcribing meetings. Those use cases fit the old tool model fine. Superwhisper and Wispr are good products for their actual audiences.

But for developers who've moved from typing code to typing intent, Cursor, Claude Code, GitHub Copilot, the word limit is a bottleneck pretending to be a feature. It's a pricing artifact, not a real constraint.

The trade-offs

Local Whisper isn't magic. It's not perfect accuracy on technical jargon. You rewrite maybe 8% of sentences. It won't beat a human transcriber for meeting notes.

It won't work if you're offline (Recitey needs an internet connection to activate the session). The transcription itself is local, but the session needs to bootstrap.

But the trade-off is real: accuracy and offline-first for zero word limit and zero variable cost. Most developers take that trade.

More posts
Keep reading

More like this.

  1. For developers

    Stop Hitting Word Caps in Your Design Docs

    Marcus is spending 11pm writing a design doc about payment settlement for a new feature. He's speaking clearly, the tool's...

  2. For developers

    The design doc that never got written

    It's 11 PM. Marcus, a backend engineer at a fintech in Stockholm, is voice-dictating a design doc for a new payment settlement...

  3. For developers

    Local transcription changes what you can dictate

    You wrote the design doc perfectly on voice, then scrolled up and realized 1400 words in, you'd stopped mid-sentence. Not...

All posts →