← BlogFor developers

Local speech-to-text, no counting.

When you're dictating a design document at 11pm and hit a word cap mid-sentence, the thinking stops. You copy what transcribed, paste it into a text editor, open a new session, splice everything back together. By morning, the prose reads like it was written by three different people. The real bottleneck in modern development isn't typing anymore. It's the long-form intent you need to articulate so Cursor, Claude Code, and Copilot can actually build what you mean.

The work has changed

A few years ago, the productivity question was simple: how fast can you type code?

Now it's different. You spend half your time writing specifications, design docs, PR descriptions, and Slack explanations to machines and people. A design doc that would take Marcus (a backend engineer at a Series B fintech) 90 minutes to type takes 12 minutes by voice. The shape of the work shifted. Speaking intent is faster than typing it.

But that only works if the tool doesn't interrupt you. Word caps, latency hiccups, and automatic uploads to unknown servers all break the flow. The moment you stop speaking to troubleshoot the tool, you lose the thought.

Where the cap hits hardest

Marcus writes design documents for settlement architecture in Notion and Slack. The work is technical, detailed, and too long to type comfortably.

A typical architecture doc runs 1400 to 1800 words. By voice, that's a single continuous session. Whisper-large-v3 handles the transcription with 96.3% accuracy on LibriSpeech benchmarks. He opens Cursor (because it tab-completes his technical vocabulary better than VS Code does), starts dictating, and the words flow.

Then he hits the cap on Wispr Flow's free plan. Around 1000 words, maybe 1100 if he's lucky. The transcription stops. He has to manually copy, paste into a separate text editor, start a new session, and splice the pieces back. The train of thought breaks. By the time he picks it up again, the prose has lost its coherence. Three different writing registers. Chunks that don't connect.

This isn't a productivity problem. It's a cognitive problem. The interruption happens exactly when he needs continuity most.

Why local transcription matters more than speed

Marcus doesn't use cloud-based dictation tools for his work. Not because he's paranoid. Because settlement code contains PII, account numbers, and routing logic that shouldn't leave his machine. Architectural decisions sit behind customer data protection.

He's not alone. Any developer working on anything remotely sensitive, payment processing, healthcare, security infrastructure, faces the same constraint.

Cloud transcription services promise encryption in transit and at rest. The promise is sound in theory. But the audio still travels over the internet. It still gets stored in someone else's infrastructure, briefly or not. For IP-sensitive work, that friction point turns into a blocker.

Whisper running locally means the audio never leaves the device. No cloud calls. No data residency questions. No "where did my words go?" The speech-to-text happens locally. Your words stay yours.

The free tier is the uncapped tier

Recitey runs Whisper locally on your device. The free tier does that, forever, with no word limit and no recount.

Other tools position the free tier as a trial. Superwhisper charges $8.49 per month for their free equivalent. Willow caps free users at a metered word count. Wispr Flow offers about 1000 words per month before you hit the paywall. The pricing strategy is consistent: make the free tier good enough to try, cheap enough to upgrade from.

Recitey inverted that. The free tier is Whisper, the same model that powers everything. Local, on-device, uncapped. The paid tier is the cloud rewrite engine, which polishes dictated speech into formal prose (fixing filler words, correcting grammar, adjusting tone). You can use Recitey for free indefinitely and never see a word counter. If you want the cloud rewrite, that's when you pay.

The structural difference: Recitey's variable cost is the cloud rewrite service, not the speech-to-text. Whisper on your device costs zero to run, whether you dictate a Slack message or a 10,000-word technical specification.

The trade-off isn't free

Local processing trades cloud latency for device CPU. Whisper on an M-series Mac, an RTX 4090, or a modern Intel processor is fast enough. On older hardware or resource-constrained devices, it's slower. That's the honest constraint.

Some developers skip the cloud rewrite entirely and edit the Whisper output by hand. Others want the cloud step to fix the rough edges (filler words like "um," repeated phrases, tone correction). The cloud rewrite is optional and separate. You can dictate everything locally and choose whether to polish it in the cloud.

Who needs this

Any developer working in Cursor, Claude Code, GitHub Copilot, or similar LLM-first tools. Anyone writing design documents, PR descriptions, incident postmortems, or long-form Slack threads that define how work should be done.

Indie builders who can't trust their architecture and design details to a third party. Teams where code and systems knowledge stay on-device. People who tried Wispr, Otter, or Superwhisper and hit the word cap mid-thought, then lost hours splicing the pieces back together.

People who believe the tool should work across the application stack they already use: Cursor, Slack, GitHub, browser, email, Notion. Not lock you into a proprietary editor or agent.

The frame is uncapping, not counting

Voice writing only works when the tool gets out of your way. When you're dictating a design doc at 11pm, you don't need the tool to tell you how many words you've used. You need it to listen. To keep pace. To not interrupt.

A local Whisper running uncapped, on your device, no word counter, no cloud dependency, no variable cost, that's the architecture that trusts you to know how much you need to say.

More posts
Keep reading

More like this.

  1. For developers

    Stop Hitting Word Caps in Your Design Docs

    Marcus is spending 11pm writing a design doc about payment settlement for a new feature. He's speaking clearly, the tool's...

  2. For developers

    The design doc that never got written

    It's 11 PM. Marcus, a backend engineer at a fintech in Stockholm, is voice-dictating a design doc for a new payment settlement...

  3. For developers

    Local transcription changes what you can dictate

    You wrote the design doc perfectly on voice, then scrolled up and realized 1400 words in, you'd stopped mid-sentence. Not...

All posts →