You explain it perfectly on the call. Then you sit down to write the design doc.
The logic's clear in your head. Complex, but clear. You start voicing it: payment settlement, edge cases, retry logic, the why behind each decision. Words flow because you're thinking out loud.
Two thousand words in. You hit the cap. Whisper stops. The flow breaks. You're left with incomplete thoughts that need stitching together tomorrow morning.
Most voice tools charge you for not hitting word caps. They're pricing the problem as if it's yours to solve.
The bottleneck shifted
Code typing stopped being the bottleneck years ago. Cursor's tab-complete means you're not typing the code anymore. You're typing intent. Prompts. Specifications. Design reasoning. The shape of the work changed, but the tools stayed the same.
Marcus is a backend engineer building payment settlement at a fintech in Stockholm. He writes design docs at 11pm because that's when the thinking is sharpest. He's not drafting press releases. He's explaining to his team why certain retry logic exists, how edge cases are handled, what the failure modes are.
That's not 500 words. It's not 1,500 words. It's as long as the thinking takes.
He switched to Cursor specifically because tab-complete reduces the rewrites he'd otherwise do by voice. One less decision per sentence. But his voice tool caps him. Wispr Flow costs $14/month and gives him 2,000 words a month free. Superwhisper is $8.49 and caps free users harder.
The pricing model assumes you're still dictating the old thing: texts, emails, short form. Not design docs. Not 30 minutes of careful explanation.
Why local actually matters here
Recitey runs Whisper locally on your device. No server round-trip. No word counter. No cap.
The technical reason is straightforward: Whisper model runs on your machine, not the cloud, so there's no variable cost per transcription. Nothing to meter. Nothing to charge by the word.
The practical reason matters more. You're explaining code logic, architectural decisions, edge cases. Marcus refuses cloud transcription because the IP doesn't leave his machine. It's not paranoia. It's the reality of how fintech works.
Local-first isn't a feature. It's structural. Everything downstream, the no-cap policy, the privacy, the latency, flows from that choice.
What actually changes
You finish the design doc in one voice session. No fragmentation. No cleanup in the morning of thoughts you interrupted mid-sentence.
The doc is rough. Recitey's cloud rewrite engine (the paid tier) polishes it. But the thinking is intact. The flow isn't broken by an arbitrary limit.
When you're explaining something that complex at 11pm, that continuity is the whole point. You're not trying to hit a number. You're trying to explain why something matters.
The trade-off isn't one
If you've already committed to local-first tooling (Cursor instead of VS Code, because it's better at your specific workflow), this choice feels consistent. The tool respects how you actually work. It doesn't impose someone else's idea of what voice dictation should be.
Some tools will feel broken to you because they're built for a different problem. Others will feel invisible because they understood the shift first. Recitey sides with the second camp.