← BlogFor developers

Voice notes should not stop mid-thought.

When you work with LLMs, you do not type code anymore; you type intent, and the model builds. This means longer prompts, longer design docs, longer explanations in Slack threads. Voice should make this faster, not faster until you hit an artificial ceiling.

The speech-to-text tax

The bottleneck is not transcription. Whisper (the model most good voice tools use) is already nearly perfect: 96.3% word accuracy on the LibriSpeech benchmark. The bottleneck is the sentence after the model listens to you.

Most voice tools impose a per-month word cap. Wispr Flow caps free users at 2,000 words per month. Superwhisper stops at 2,000 per month for $8.49. Willow does similar. The reasoning is always the same: cloud transcription costs money, so free tier means limited usage.

The problem is not theoretical. Marcus, a backend engineer at a fintech, hit this during a design doc at 11pm. He was explaining a payment settlement system architecture to the model. The tool stopped him at 1,800 words. He either typed the rest manually or stopped halfway. Either way, the thinking broke.

Why local matters for long-form thinking

There are two ways to do speech-to-text. Cloud transcription (send audio to a server, get text back; this is how Wispr, Superwhisper, and Willow work) or local transcription (run the model on your device, audio never leaves). Cloud is cheaper per call in aggregate because the vendor spreads cost across users. Local has zero variable cost, the model runs on your GPU or CPU, once, then it is done.

For developers, local matters more than marketing copy admits. Code snippets in your voice memos do not leave your machine. Architecture decisions you are thinking out loud stay on your device. No telemetry, no data-export agreements, no "we will delete it after 30 days." It is gone when you are done.

Recitey runs Whisper locally on Windows. No word limit. No monthly cap. No cloud gateway. The free tier is the dictation. The paid tier is the rewrite into clean prose.

The asymmetry of pricing and technology

Most SaaS pricing reflects distribution and support costs more than the actual technology cost. Wispr Flow charges $14 per month, not because Whisper costs $14 per month to run (it does not), but because they built a business around cloud infrastructure. That cost structure is real, and the product is good for some users.

But for a developer who already has a GPU (or who does not mind the latency of CPU transcription), that cost structure is paying for something you do not need. You do not need their server. You need your machine to understand your voice.

The secondary problem (and what Pro solves)

The real limit is not transcription. It is prose quality. When you voice a 500-word design doc, the raw transcript is dense, has restarts, includes thinking artifacts ("wait, that does not make sense, let me rewind"). That is when the rewrite matters.

Recitey's Pro tier includes cloud-based polishing. It takes the rough transcript and cleans it into a coherent paragraph in under 2 seconds. But this step is optional. Many developers do not need it. Many would rather edit their own voice memo and keep the thinking intact.

The free tier is for developers who want speed without infrastructure rent.

Who this is for

This matters if you think local-first. If you build in Cursor because the tab-complete reduces your voice rewrites and you refuse cloud transcription because code IP. If you work in Claude Code, GitHub Copilot, or anything that sends prompts to a third party and you want to control one layer. If you have words to say and would rather not pay per month to be heard.

The word cap was designed for a different use case than yours. Remove it, and you find out how much you were holding back.

More posts
Keep reading

More like this.

  1. For developers

    Stop Hitting Word Caps in Your Design Docs

    Marcus is spending 11pm writing a design doc about payment settlement for a new feature. He's speaking clearly, the tool's...

  2. For developers

    The design doc that never got written

    It's 11 PM. Marcus, a backend engineer at a fintech in Stockholm, is voice-dictating a design doc for a new payment settlement...

  3. For developers

    Local transcription changes what you can dictate

    You wrote the design doc perfectly on voice, then scrolled up and realized 1400 words in, you'd stopped mid-sentence. Not...

All posts →