← BlogFor developers

Local transcription changes what you can dictate

You wrote the design doc perfectly on voice, then scrolled up and realized 1400 words in, you'd stopped mid-sentence. Not because you ran out of thoughts. Because your free tier word limit cut out mid-flow.

This is what developers encounter when their voice tool meters by cloud cost instead of actual value.

The workflow shifted, but the tools didn't

Three years ago, coding meant typing. A lot. Muscle memory, habits, the whole stack of developer identity wrapped up in keyboard speed.

Then LLMs changed the game. Now coding means specifying intent. You're not typing code. You're explaining the shape of code to a model that will draft it.

The keyboard hasn't gone anywhere. But the work is different. You type longer prompts in Cursor. Longer context in Claude Code conversations. Longer PR descriptions that explain the why, not just the what. Longer async explanations in Linear when a bug deserves more than a one-liner.

The bottleneck is no longer keystroke speed. It's the time between the thought and the text. Speaking faster than typing has always been true. But it was never practical before because developers coded, not dictated specifications.

Now it is. And the tooling caught up at a weird moment.

The word counter becomes a thinking interrupt

Marcus works at a Series B fintech in Stockholm. His job is payment settlement logic. The kind of work that needs design docs, not just code comments. He'll draft a design doc by voice at 11pm when the thinking is clear, when he can explain the tradeoffs without reading code.

The first few times he tried a cloud dictation tool, it worked. Then he'd review his draft and notice the doc ended abruptly. Not cut off mid-word. Cut off mid-thought, right at the word limit.

He'd rewritten that doc three times before he realized the problem wasn't his voice or the transcript quality. The tool had a ceiling: 2000 words on the free tier, another 1500 if you paid $14/month. Willow's free tier caps at 1500 words total. Superwhisper charges $8.49/month for unlimited. They're all capped in some way.

The interruption itself breaks the thing Marcus needed from voice in the first place: the unbroken thinking flow. He'd explain the first tradeoff, the architecture, the why. Then hit a counter. Then restart his thinking the next day with a fragment.

It's not really about the word count. It's that the tool inserted a cost center into a moment where he needed to think clearly.

Local changes what you can afford not to meter

There's a second friction Marcus hit even before the word cap: code IP.

If he's dictating a design doc that explains his payment settlement logic, that's proprietary. His startup's moat lives in that detail. When he dictates into a cloud tool, that audio goes somewhere. It gets transcribed. It might get stored. The terms of service usually allow the vendor to use it for model training. It's in a cloud provider's hands.

That hesitation is not paranoia. It's structural. Cloud dictation tools have variable costs. They absorb cost by offering their usage data as a service. Privacy policies use careful language, but the business model is metering: you speak, they transcribe and monetize the data.

Marcus won't use a cloud voice tool for code-adjacent work for that reason alone. He uses Cursor over VS Code partly because the autocompletion reduces the number of times he has to speak code out loud.

The pricing you see on a voice tool often reflects distribution and monetization, not the actual cost of speech-to-text. OpenAI's Whisper model was released as open source in 2022. Run it locally, and the cost is zero per word. It's a one-time compute cost on your machine. No usage tracking. No variable billing.

Most commercial voice tool costs come from supporting their infrastructure, not running the model.

What local-first actually solves

When your transcription runs locally on your device, three things change.

First, there's no word counter. Whisper doesn't meter words. It processes them. There's no "you're at 87% of your limit" notification. You speak the full thought. You hit save. It's done.

Second, nothing leaves your device by default. Your design doc stays on your machine during transcription. Your code snippets don't flow through a vendor's pipes. If you want to share the transcript, you copy it yourself. That's a different choice from the tool automatically uploading for "cloud benefits."

Third, the rewrite doesn't need to round-trip a cloud API. Recitey runs Whisper locally for the initial transcription, then uses a cloud LLM for the rewrite pass if you want it. But the raw speech-to-text? That's local, uncapped, no metering. The free tier includes it.

The tradeoff is that local models have higher latency on your machine than cloud APIs. Whisper isn't instant. It takes a few seconds depending on your hardware. For a design doc, a five-second transcription delay is fine. You're not waiting for a keystroke response. You're composing a paragraph.

Why this matters for the work developers actually do now

Wispr Flow ($14/month for unlimited) and Willow ($12/month) exist because their business is metering and monetization. They sell limitations to justify pricing. Superwhisper charges by month for local transcription partly because it's one developer building it.

Recitey approaches it differently. The free tier gives you uncapped local transcription because the cost structure is actually zero. The paying tier is for the cloud rewrite capability, not for more speech-to-text.

That's a different bet about what developers actually need. Not faster dictation. Uninterrupted thinking space when you're explaining intent to an LLM. No IP concerns mid-design-doc. No word counter breaking your flow.

It assumes you've already moved from coding to specifying. And it builds the tool around that assumption, not the old one.

For Marcus, it changed when he could work

He started using the tool for his 11pm design docs. The ones where the thinking is sharp but the time zone is awkward and he can't context-switch to a keyboard without losing the flow.

No word limit meant he finished the thought. No cloud transcription of his code logic meant he didn't hesitate before speaking proprietary tradeoff details. The tool worked the way the work actually shaped itself.

That's not productivity theater. That's a tool that understood the workflow changed.

More posts
Keep reading

More like this.

  1. For developers

    Stop Hitting Word Caps in Your Design Docs

    Marcus is spending 11pm writing a design doc about payment settlement for a new feature. He's speaking clearly, the tool's...

  2. For developers

    The design doc that never got written

    It's 11 PM. Marcus, a backend engineer at a fintech in Stockholm, is voice-dictating a design doc for a new payment settlement...

  3. For developers

    Design Docs at Midnight (no word counter telling you to stop)

    The moment you hit the word limit mid-thought is the moment you lose the design doc. For Marcus, a payment-settlement engineer...

All posts →