← BlogFor developers

Design Docs at Midnight (no word counter telling you to stop)

The moment you hit the word limit mid-thought is the moment you lose the design doc. For Marcus, a payment-settlement engineer at a fintech, this happens at 11pm. He's dictating a complex design doc into Slack, explaining race conditions, database lock patterns, the exact state transitions that matter. His voice is fast. His typing would take forty minutes. His cloud dictation tool caps at 1500 words per day on the free tier. He hits it at minute 23. The doc fragments. The clarity fractures. Tomorrow morning he has to reconstruct what he meant, and the thinking is cold.

This is what's changed for developers in the LLM era. We stopped typing code. We started typing intent. That shift rewired the bottleneck.

The Bottleneck Moved

Five years ago, developer productivity tools focused on typing speed. Type less. Click less. Get your hands out of the mouse. The friction was mechanical.

That was before Cursor and Claude Code made prompt-writing the rate-limiter. Now the problem is different: you have thirty seconds to describe a complex refactor, and your hands can't keep up with your thinking. Typing three paragraphs of spec feels slow. Dictating them feels natural. But then you hit the cap.

Wispr charges $14 a month to remove word limits. Willow does $12 for the same gate. Even the indie tools, Superwhisper at $8.49, put you on a leash. They're cloud-based. Variable costs. They meter it.

Recitey doesn't. The transcription runs locally on your machine using Whisper. No variable cost per word. No word counter. No artificial scarcity. The free tier is actually free, with no monthly reset surprise in January.

What "Local First" Actually Means (and Why It Matters for Code)

Most developers hear "local" and think latency or offline mode. That's real, but it's not the actual reason developers prefer it.

The actual reason: your code never leaves your device.

If you're dictating a PR description that quotes ten lines of a security fix, you don't want that context, even compressed, flying through someone else's servers. If you're narrating a design doc that discusses your authentication model's edge cases, you want that to stay on your machine until you edit it. If you're dictating a debug session that walks through a race condition, you want the full context under your control.

Cloud dictation tools don't advertise that they log what you say. But they might. The terms of service are rarely clear about what happens to code context. Developers who've worked on proprietary models or payments infrastructure don't want to find out mid-deployment that their voice snippets ended up in a training corpus somewhere.

Recitey's architecture avoids this entirely. Whisper-large-v3 has a 96.3% word-error rate on LibriSpeech and runs on your device. What you say stays on your machine. The rewrite pass, the polish that turns rough speech into clean, structured text, is optional and you control when it happens.

For Marcus specifically, this was the blocker. He'd looked at Superwhisper and Wispr and both require choosing: either cap yourself or trust a service with your codebase context. Neither option is great at 11pm when the code is fresh in your head.

The Workflow is Different Now

The moment you start dictating prompts instead of typing code, the shape of the work changes. You're not documenting what you built. You're architecting what you want the model to build. Those are different processes.

Typing feels like writing. It has friction built in, backspace, rewrite, rethink. That friction is sometimes good. It slows you down. You catch mistakes.

Dictating feels like thinking out loud. The friction is lower. Your words come faster than your thoughts clarify. That's actually the point. You want the rough shape, then you refine it in the editor.

The problem with most voice-to-text tools is they assume you're dictating a finished thought. They try to be smart about punctuation and capitalization and paragraph breaks. They try to make it immediately publishable. It never is. Developers end up post-processing anyway, and now they've used up their word cap on a draft they're going to rewrite.

Recitey inverts this. Dictation is rough and uncapped. You get the full thought without worrying about the meter running. Then the polish pass, turning rough speech into clean sentences, is separate and optional.

Cursor Changed This, Too

Marcus switched from VS Code to Cursor specifically because Cursor's tab-complete reduces the number of voice rewrites he needs.

If you're dictating a function signature and Cursor autocompletes the pattern, you accept it with Tab and keep speaking. If you're in VS Code, you have to stop, type the completion, start speaking again. That context switch adds up. Over a 20-minute design doc, it adds maybe four minutes of friction and two moments where your train of thought derails.

Most voice-writing tools don't consider the IDE loop. They assume you dictate and then edit in isolation. They don't understand that developers spend most of their time in Cursor or Claude Code or GitHub's web interface, and the voice tool needs to work inside those flows, not alongside them.

Why the Pricing Matters (Beyond Sticker Price)

This gets at something deeper about how software gets priced. Wispr and Willow aren't expensive because transcription is expensive. They're expensive because they've chosen a distribution model that requires them to be. Cloud infrastructure, variable costs, a per-user SaaS flywheel.

That's a valid business model. It's not the only one.

If the transcription runs locally and costs zero per word, the pricing problem changes. You don't need to meter. You don't need to charge $14 to "remove limits." The limits were never real constraints. They were business decisions.

For developers who've built their own payment metering or consumption tracking, this is obvious. You price based on what's actually scarce. Words aren't scarce when they're local. Attention is scarce. Time is scarce. But not words.

The Identity Question

This is subtle, but it matters. Developers are skeptical of tools that obscure how they work.

If a voice tool won't tell you which model it's using or where processing happens, that's a red flag. If it claims to be powerful without explaining what that means, you're right to be suspicious. The vagueness is intentional. It lets the vendor hide cost structure or lock-in.

Recitey's choice to run Whisper-large-v3 locally and be transparent about when rewriting happens reads as engineering taste to developers. It says: we're not trying to make this magical. We're being clear about what does what.

That clarity is trust. And for developers who've been burned by opaque licensing or hidden word-limits before, clarity is worth more than the price.

More posts
Keep reading

More like this.

  1. For developers

    Stop Hitting Word Caps in Your Design Docs

    Marcus is spending 11pm writing a design doc about payment settlement for a new feature. He's speaking clearly, the tool's...

  2. For developers

    The design doc that never got written

    It's 11 PM. Marcus, a backend engineer at a fintech in Stockholm, is voice-dictating a design doc for a new payment settlement...

  3. For developers

    Local transcription changes what you can dictate

    You wrote the design doc perfectly on voice, then scrolled up and realized 1400 words in, you'd stopped mid-sentence. Not...

All posts →