← BlogFor developers

Dictation without the leash

When you're three paragraphs deep into a design doc at 11pm, explaining how the settlement pipeline should handle retries, the last thing you need is a modal telling you you've hit your word limit for the month. That's when most developers switch tools, not to a better tool, just away from voice altogether. It's not about speed. It's about the specific shape of how developers write now.

The Whisper-to-Copilot Pipeline

The work changed. You used to write code. Now you write intent.

A pull request description used to be two sentences. Now it's a paragraph explaining the intent, the trade-offs, the reason you chose this approach over three others. The same for design docs. The same for incident postmortems. Anywhere you used to leave a comment, you now write a mini-essay.

Marcus, a backend engineer at a fintech in Stockholm, noticed it first in his design docs. A few years ago, a design doc was a bulleted outline: problem, approach, alternatives, decision. Now it's 2,000 words of narrative. He explains the original API design. He walks through why that design created the settlement race condition. He outlines the payoff curve for each fix. He writes the thinking out, not just the conclusion.

When you're writing that much context-heavy prose, especially at 11pm, when you're tired and the explanation is tangled, voice is faster than typing. Not because you type slowly. Because you're tired, and voice lets you think in sentences instead of debugging typos and syntax. The words come out rough, but they come out.

Then he hits the word cap.

Marcus uses Wispr Flow ($14/month, 2,000 words per month on the free tier). It cuts him off mid-explanation on a regular design doc. He's got another 1,500 words to explain the routing logic, but the tool says he's done. He switches to typing. The flow breaks. The prose fragments. He finishes the doc the next morning, re-reading the rough draft and rewriting half of it because it doesn't sound like the thinking that was in his head.

Worse: he's paranoid about where that audio goes. Payment settlement code is IP-sensitive. GDPR-sensitive. The idea of cloud transcription sitting in some service's database for even a few seconds between utterance and text makes him deeply uncomfortable. So he doesn't use cloud dictation for sensitive code contexts at all. He types those instead. Slowly. He switched to Cursor (not VS Code) specifically because Cursor's tab-complete reduces the number of rewrites he needs. But the architectural choice, cloud transcription plus metering, means voice never touches his IP-sensitive work.

Why "Cloud" Is the Wrong Architecture for This

The industry standardized on cloud transcription because it unlocks per-user pricing and metering. Better accuracy, maybe. Continuous training. But cloud transcription creates two specific problems for the developer workflow.

First, it introduces latency. You speak a sentence. It goes up to a server. The server transcribes it. The text comes back. Two to ten seconds, depending on the service and your network. It breaks the feel of typing. Voice is only faster than typing if it feels like thinking, if the latency between thought and text is negligible. Cloud transcription adds just enough delay that you start watching the transformation instead of continuing to think. The cognitive load flips from "expressing the idea" to "waiting for the service."

Second, it creates a metering layer. The business model requires capping. If transcription runs in the cloud, the service has to meter it, price it, limit it. A $14/month plan with 2,000 words per month is not a technical decision. It's a business decision. The tech could handle unlimited words. The pricing says it can't. You become aware of the meter. Awareness breaks flow.

Local speech-to-text, OpenAI's Whisper running on your device, solves both. No latency. No meter. No cloud. The transcription happens on your machine. The audio never leaves. The accuracy is solid: Whisper-large-v3 hits 96.3% word accuracy on standard English speech, which is sufficient for intent capture in design docs and code comments. It's not perfect. It's not meant to be. The next layer fixes the rough edges.

How Recitey Uses Local Whisper

Recitey runs Whisper locally on your device and hands you the rough transcript.

That's the free tier. No word limit. No cap. No meter. No cloud. Just the transcript and a text box. You can voice 10,000 words if you need to. Four hours of continuous speaking. It makes no difference to the system. The architecture doesn't care.

The pro tier is where the rewrite happens. That's cloud-based. That's where Recitey takes the rough voice draft and polishes it into clean prose, fixing grammar, tightening structure, tuning register for the platform. A Slack message sounds different from a GitHub PR comment. That polish takes cloud processing power. That's where the pricing lives. But the transcription, the hard part, the latency-critical part, the privacy-critical part, is local. It's on your machine. It doesn't phone home.

The architecture acknowledges what developers actually need: fast, private speech capture. And then: good prose polish, optionally in the cloud, only when you ask for it.

The Unrestricted Workflow

With no word limit on the transcription side, the workflow shifts.

Marcus doesn't think about conservation anymore. He doesn't ration his explanations or condense mid-thought. He drafts the full arc of the argument into voice. Eight minutes. Four thousand words. The rough transcript is messy, run-on sentences, repeated phrases, words he half-mumbled, but the logic is complete. The structure is there. The thinking is on the page.

Then the rewrite layer polishes it. Two seconds. Four thousand words become clean prose, properly punctuated, tightened where he rambled, broken into readable paragraphs. And the IP stays local. The audio transcript happens offline. The only thing that travels to the cloud is the already-transcribed text, which he can read and delete if it's sensitive.

The time swing is real. A 2,000-word design doc that used to take 45 minutes to type now takes 8 minutes to voice and 2 minutes to polish. The quality is the same or better, because he's thinking in full sentences instead of managing his own word rate. He goes from 2 design docs per week to 8 or 10.

The Trade-Offs

This only works if you accept the rough first draft.

Most voice-to-text marketing tries to make voice sound as polished as typing on the first pass. It can't. The transcript will have errors. Words will be repeated or mumbled. Sentence structure will be tangled. If you expect perfect prose directly from voice, you'll get frustrated.

But if you expect the rough capture, if you're comfortable with a two-stage workflow where the first stage is speed and completeness and the second stage is polish, then the uncapped local transcription changes the economics of how you write.

The other trade-off: you're responsible for accuracy tuning. Whisper is trained on English primarily. If you speak with a heavy accent or your code context includes non-English terms, the transcript will reflect that. It's your job to know your own accent and whether Whisper handles it. It usually does. Sometimes it doesn't. But you own the transcription, not rely on a service to fix it.

Why Developers Should Care About This Architecture

The shift from "developers type code" to "developers write intent for AI" is real. The bottleneck is no longer typing speed. It's the cognitive load of holding a complex explanation in your head while your fingers hunt for keys at midnight.

Local voice capture removes two specific frictions: the latency of cloud round-trips and the anxiety of IP leaking. That's not a speed improvement. That's a workflow unblock. You can write the way you think. The rough output gets polished downstream. The speed comes from the fact that thinking-out-loud is faster than typing-while-editing.

It's a small architectural difference. It's a big workflow difference.

More posts
Keep reading

More like this.

  1. For developers

    Stop Hitting Word Caps in Your Design Docs

    Marcus is spending 11pm writing a design doc about payment settlement for a new feature. He's speaking clearly, the tool's...

  2. For developers

    The design doc that never got written

    It's 11 PM. Marcus, a backend engineer at a fintech in Stockholm, is voice-dictating a design doc for a new payment settlement...

  3. For developers

    Local transcription changes what you can dictate

    You wrote the design doc perfectly on voice, then scrolled up and realized 1400 words in, you'd stopped mid-sentence. Not...

All posts →