You aren't typing less. You're typing differently.
A year ago, the conversation in developer communities was about writing code faster. Typing speed. How many words per minute. Now the question is different. It's about writing intent faster. The shift happened quietly, almost imperceptibly, but it's complete.
You spend your 11pm design doc session not building the thing, but explaining to Claude or Copilot what you want it to build. You write Slack threads describing bugs to your team instead of debugging in isolation. You draft PR descriptions that are longer than your actual code diff. You spend thirty minutes composing the prompt that Copilot will turn into fifty lines of code. The keyboard is still there. The work just moved upstream, to the layer where you explain what should happen.
This is the problem that broke most of the existing dictation tools for developers. They were built for the old world. The world where you talked faster than you could type, so you dictated voicemail or memos. That tool was fine for speeches or quick notes. Now you're writing software specifications, incident postmortems, design rationales, and architectural explanations. You need twenty minutes of uninterrupted voice thinking converted into clean text. Most of the tools that claim to handle this cap you at a word limit.
The Word Cap Problem
Wispr Flow charges $14 a month for their paid tier and their free tier is capped at 1,200 words per day. Willow is $12 per month and also capped. Superwhisper is $8.49 as a one-time purchase and caps free tier to 2,000 words per dictation. The pricing economics make sense from a business perspective. They run speech-to-text processing on the cloud, and cloud processing costs scale with your volume. So they meter you.
Here's what actually happens when you hit that cap mid-flow.
You're twenty minutes into a design doc explaining a payment settlement edge case. You're in the zone. The architecture is coming together in your voice. The system is making sense. And then the app stops recording. You have three options: wait until tomorrow when your word counter resets, pay them monthly, or switch to typing and lose the momentum you just built.
Most developers choose the third option. They go back to typing.
Marcus and the IP Problem
Marcus is a backend engineer at a Series B fintech in Stockholm. His team owns payment settlement. He writes design docs at 11pm when the thinking is clearest, and he wanted to try voice dictation to speed that up. He looked at the cloud transcription options. Every single one of them sends your voice to their servers. Your voice contains your payment algorithm explanation, your edge case handling, your security assumptions, your incident response reasoning. All of it travels to a cloud service and gets processed there.
Marcus refused. His company's code and architecture are competitive advantages. Sending design docs to a cloud service for transcription, even if they claim to delete them, felt like a data spill waiting to happen. The voice memo is temporary, but the content isn't. He uses Cursor, not VS Code, specifically because Cursor's tab-complete reduces the number of voice rewrites he needs to do. He's deliberate about his tools. The cloud transcription option was a non-starter.
The local-first framing matters here. It's not paranoia. It's the worldview that many engineering teams have settled on: cloud processing of code-adjacent data is a convenience trade-off that should be optional, not mandatory. If a tool won't run locally, if it requires sending your thinking to the cloud, you look for another tool. Marcus did.
The Whisper Model Changed the Equation
OpenAI's Whisper model can run locally. It ships as open-source code. You can run the large model on a mid-range GPU, or even on CPU with some patience. The variable cost is zero once the model's downloaded. No per-transcription fees. No cloud services. No data leaving your device.
This changes what "free" actually means in the context of speech-to-text tools. It means you can build a tool that doesn't have to meter you, because the processing cost doesn't scale with your usage. You can make the free tier genuinely uncapped. No word counter. No daily reset. No per-dictation limit. You record for as long as you need to explain the thing, and it transcribes all of it.
The Pro Tier Is the Rewrite
That's where the product design gets interesting. If local speech-to-text is free and unlimited, what's the paid tier for?
It's for the rewrite. After you've dictated the twenty-minute design doc explaining the payment settlement edge case, the raw transcript is rough. It's got false starts, repeated words, conversational filler. It's your thinking made text, not publication-ready prose.
A local rewrite engine (running on the cloud, because the rewrite's more sophisticated than transcription) takes that rough transcript and converts it into publication-ready prose in under two seconds. It cleans up the false starts, removes the repetition, structures the ideas, and leaves you with something you can paste into a design doc or Slack without embarrassment.
That's where the value gets concentrated. The free tier gives you the dictation, uncapped. The paid tier makes it polished. It's a feature strategy, not a metering strategy. The pricing aligns with where the actual work is, not where the infrastructure cost is.
Who This Actually Is For
This isn't for people who want to talk instead of typing. That product already exists, and it's existed for years.
This is for engineers who write a lot now. Who spend their evenings thinking through design decisions in long-form voice. Who have code IP they don't want floating around cloud services. Who understand that Cursor beats VS Code because of what it does with the output, not the input. Who refuse to send their architectural thinking to cloud transcription services, no matter the convenience.
If you're recording five-minute voice memos and expecting them to become polished Slack messages, you already have a tool that works for you.
If you're recording twenty-minute design doc sessions and hitting word caps mid-thought, if you've abandoned cloud transcription because your architecture decisions aren't public, if you know the bottleneck is explaining intent to Claude and not typing the intent itself: this is built for you.
The free tier is local Whisper. It stays uncapped.