It's 11pm. You finish the code, decide the design doc can't wait, and reach for voice dictation. Typing takes too long. Your thoughts fragment.
Twenty minutes in, explaining the payment settlement flow, the tool stops. Word cap hit. You manually transcribe the rest, and the prose breaks.
Next morning, you're rewriting something that should have flowed naturally in one pass.
This was Marcus's pattern for months. He'd switch voice tools looking for a solution. Wispr charges $14/month for unlimited, the free tier is capped. Willow's free tier is capped. Superwhisper ($8.49/month) caps the free plan.
None of them understood the real constraint.
The work shifted, the tools didn't
Three years ago, developers typed code and occasionally spoke.
Now developers speak intent to LLMs, wait for code, speak corrections, write design docs in voice at midnight.
That's more words leaving your mouth. PR descriptions. Slack threads untangling a bug investigation. Design specs while the architecture is still clear. Long-form explanations, often 600 to 800 words, with the thought still forming.
Voice wins when the explanation is long. The bottleneck is not transcription speed. It's keeping the thought coherent across a full session.
Marcus works in Cursor (not VS Code, specifically because tab-complete reduces voice rewrites). He writes design docs for payment settlement code. A word cap would interrupt him mid-thought.
Local is not about speed, it's about trust
Marcus once considered Superwhisper. Looked at the architecture.
Audio goes to an Anthropic cloud service. For code explanations, that's a data-leakage risk. Code stays internal. The audio of him explaining code doesn't.
He didn't want to optimize his words for what's safe to say aloud.
Local Whisper runs on his device. Zero external calls. Zero risk. The technical detail is structural: speech-to-text runs locally with zero variable cost. That's why there's no word cap on the free tier. No metering, no cloud infrastructure bill, no per-word pricing model.
What one draft feels like
Marcus tried Recitey. No cap. No cloud calls.
For the first time, he finished a design doc in one voice session without hitting a boundary.
The prose still needs editing (voice is not clarity; that's separate). The structured thinking still matters. But the interruption was gone.
One draft instead of two. No fragmentation to fix in the morning.
The wrong metric, the right decision
Voice tools get reviewed on WPM. Wispr advertises accuracy. Superwhisper sells indie purity.
Marcus doesn't optimize for any of that. He optimizes for finishing a thought without hitting a meter set by a business model.
The word cap was never a technical limitation. It was a pricing design. And for developers writing long-form intent, it's the wrong one.
Local speech-to-text doesn't cost per word. So Recitey doesn't charge per word.
The thinking continues. The design doc completes. Marcus's midnight workflow stays intact.