The shift from coding to prompting has redrawn what developers actually do. You're not typing code anymore; you're typing intent. The bottleneck moved from keystroke speed to thought coherence. And every time a voice-to-text tool cuts you off mid-sentence, mid-paragraph, mid-design-doc because you've hit the monthly word limit, it's not just inconvenient. It breaks the thinking.
The work changed, but the tools didn't notice
Five years ago, a developer's day looked like: write code, debug code, read code. Voice dictation would have been nice, but typing was already the mode you were in.
Today it's different. You're still in code, but you're also living in Cursor, Claude, GitHub conversations, Slack threads where you're explaining a bug or defending a design decision. The output has expanded. You're not faster at typing code; you're doing much more explaining. Design docs at 11pm because the idea won't leave your head. PR comments that need to walk reviewers through the why, not just the what. Incident postmortems where you're reconstructing 40 minutes of investigation into words.
Voice makes sense here. It's faster than typing for the kind of writing that requires you to hold a complex thought in your head. A 900-word design doc takes 11 minutes to speak. The same 900 words typed takes three times as long because you're stopping to structure, search for the word, second-guess the phrasing.
But voice tools don't charge for speed; they charge for words. When you're halfway through explaining a database migration and the tool tells you that you've hit your monthly word limit, the thinking breaks. You're left with a fragmented doc. Half coherent. Something you'll have to fix tomorrow morning. The benefit of voice (staying in flow, keeping the thought intact) evaporates.
Why word limits are a proxy for bad incentives
Most voice-to-text pricing models are built for a different use case: short-form, high-frequency transcription. Call centers. Legal discovery. People who need a lot of short bursts.
Superwhisper charges $8.49 one-time. Wispr Flow charges $14/month for cloud speech-to-text, capped at 2,000 words per month on the free tier. Willow charges $12/month with similar caps. They all assume you're doing discrete, bounded tasks.
A developer writing a design doc isn't bounded. You don't know in advance how long the thought is. You discover it as you speak. Cut it off, and you've interrupted something that won't stitch back together cleanly.
The word counter itself is a cognitive burden. You're speaking, holding the idea, and in the back of your head is the question: "How many words left?" That's friction that doesn't show up in any feature list. It's the reason some teams fall back to typing their design docs, tho they know voice would be faster. Because at least typing doesn't have a hard ceiling in the middle of the thought.
What changes if the limit goes away
Local Whisper-large-v3 runs on your device with no variable cost per word transcribed. No counter. No reset date. No metering. It's why the incentive structure is different.
When there's no word limit, you don't optimize for brevity during the speak phase. You optimize for clarity. You let the thought expand. You circle back. You talk through a problem from two angles because you want to see which framing lands better. That iteration, that exploratory thinking, is only possible if the tool isn't watching the meter.
Marcus, a backend engineer at a Series B fintech in Stockholm, hit this exact wall with a cloud voice tool. He was drafting settlement logic in a design doc at 11pm. He ran out of words mid-explanation. The next morning, he had to read his fragmented notes and reconstruct what he was thinking. It took 40 minutes to re-explain something he'd already explained to himself the night before.
He switched to local transcription. Same speech engine. Same quality. But because there's no metering, no word counter, no monthly reset, the doc stayed coherent. The thinking stayed intact. The time savings from voice didn't evaporate halfway through.
That's the difference. It's not about speed. It's about the thinking staying in one piece.
The privacy angle, unspoken until you need it
Most developers skip this in conversations because it seems paranoid. But if the speech-to-text is happening in the cloud, every design doc, every code explanation, every incident postmortem you speak is leaving your device.
Legal teams have opinions about this. Clients have opinions about this. And if the tool will not clearly explain where the audio goes and how long it's kept, that's an intentional blur.
Local speech-to-text removes the variable. The audio stays on your device. Whisper runs on your GPU. The transcription stays local. The only thing that leaves is the text, and only if you choose to paste it somewhere.
Marcus specifically chose to use Cursor for his voice drafting workflow partly because Cursor's tab-complete reduces voice rewrites, but also because he wanted to avoid any tool that wasn't transparent about where his code explanations go. For a backend engineer writing about payment settlement logic, that opacity felt like a liability.
The honest trade-off
Cloud voice tools are better at some things. Specialized models for medical transcription. Real-time collaboration. The infrastructure to handle thousands of concurrent users. If you need those things, local doesn't win.
But if you need to draft a design doc without hitting a word counter, and you want your code explanations to stay on your device, the incentive structure of local is clearer. There's no metering. No reason to rush. No monthly billing that punishes coherence.
The developer workflow has shifted to requiring more words, longer thoughts, and more explanation. The tools that charge per word are working against that shift. The tools that don't charge per word, and don't watch the meter, are aligned with it.