← BlogFor developers

The word counter you're not supposed to notice

You hit it at 11pm Tuesday. Three paragraphs into the payment settlement architecture doc, mid-sentence explaining why the retry logic needed a different state model. The dictation tool stops listening. No warning. No UI change. Just silence. When you click play to review, it ends at "we decided to use exponential backoff because, " and never finishes the thought.

Most voice dictation tools have invisible word caps. Wispr Flow stops at 2000 words/day on free tier. Superwhisper caps at 5000. MacWhisper tops out somewhere around the same. The caps exist because cloud transcription costs money. The company has to meter it, so they add a counter somewhere in the UX, or just let the service drop.

The problem isn't that the cap is low. The problem is that it's invisible until you hit it. You can't see it coming. You're mid-flow, explaining something you've been wrestling with all day, and the tool goes silent. You finish the thought on your keyboard. It feels like the tool is telling you: I've heard enough. Pay us more.

Your workflow changed, but the tools didn't notice

Two years ago, the bottleneck was typing speed. You sat down, opened the IDE, and your hands knew what to do. The limiting factor was how fast you could translate thought into code.

That's not true anymore. The bottleneck shifted. Now you're dictating for intent. You're explaining what you want to build to Claude, Cursor, Copilot, GitHub Copilot. You're writing prompts. You're drafting design specs. You're documenting decisions. And then, maybe, you write code.

The friction isn't in your voice speed. The friction is in the 30 minutes you spend shaping a prompt into something the model will actually understand. Or the design doc that should be one coherent narrative but becomes three fragmented sections because you had to stop halfway through, switch to the keyboard, and finish the thought manually. That's cognitive context lost between capture and completion. That's flow state interrupted.

Voice should help with this. You can explain architecture faster by voice than by typing. You can think while you speak. The tool should just listen and get out of the way.

But if the tool has a word cap, it can't get out of the way. It stops listening, and you lose the narrative thread.

The real cost of word caps: losing the thought

When you're explaining payment settlement logic, or documenting a bug investigation, or writing the rationale for a refactor, you need to finish the thought. You need to say it all before you stop. Voice lets you do that. But only if the tool doesn't decide your explanation has enough words now.

Marcus is a backend engineer at a Series B fintech in Stockholm working on payment settlement systems. He used to split his workflow between voice and keyboard: Slack voice clips for quick explanations, design docs for detailed architecture. But he'd start explaining a bug investigation by voice, hit Wispr's word limit at "the issue was that the transaction state machine didn't account for, ", then finish the rest on the keyboard. The next morning, his notes were scattered: a voice transcript ending mid-sentence, keyboard prose starting abruptly, and him spending 20 minutes stitching it back together to make sense of his own work.

The word cap didn't exist to serve Marcus. It existed to solve Wispr's business model problem. Cloud transcription costs money, so you have to limit usage to control costs. But Marcus doesn't care about Wispr's infrastructure constraints. Marcus cares about finishing his thought.

What changes when the cap disappears

Recitey runs Whisper locally. On your device. No metering, no cloud cost, no word cap. You can dictate a 10,000-word design doc and the tool just keeps listening.

With local dictation and no cap, Marcus speaks the full investigation into Slack. The thread is coherent. No interruption. No rewriting tomorrow morning. The architecture doc gets written in one pass, in voice, with his thinking intact.

This isn't about speed. Marcus isn't twice as fast now. This is about coherence. This is about saying what you mean before the tool interrupts you.

The absence of a word counter means absence of interruption. That's different from most productivity tools, which measure your output and tell you how much more you can do. This tool measures nothing. It just listens until you stop.

The trade-off you're actually making

Local dictation is slower than cloud transcription. Whisper takes a second or two per sentence to process. If you need millisecond-instant feedback while you speak, you'll feel the latency.

Most developers don't need that. Most developers speak, pause, think, then speak again. The second of processing time is invisible in that rhythm.

Whisper-large-v3, the model Recitey uses, hits 96.3% accuracy on LibriSpeech test data. For developer voice, payment systems, API calls, error codes, familiar variable names, accuracy is higher because you're in a known domain. Cloud models don't have that advantage. They train on general speech, not your specific code.

Local execution also means your design docs, your bug explanations, your sensitive code context never leaves your device. For backend engineers working on regulated systems, financial services, or systems where code IP is a concern, that's a real difference. You're not sending your payment settlement architecture to some cloud API endpoint.

You're also not paying per word. There's no metering. There's no "you've used 2000 words this week, upgrade to $14/month to use more." The cost is zero. The limitations are zero.

Comparing the alternatives

If you're using cloud transcription because you want specific features, real-time transcript display, voice cloning, automatic formatting, moderation guardrails, Wispr and Superwhisper are built for that. They've invested in cloud infrastructure to add those features. That's a valid choice.

But if those features don't matter to you, if you just want to dictate your architectural decisions without the tool interrupting you, local dictation is simpler. It works because it doesn't try to do extra things. It just transcribes and gets out of the way.

The choice between them isn't about which is better. It's about which one serves your actual workflow.

The real insight: what got built for money isn't what you need

Premium SaaS pricing usually reflects distribution costs more than technology costs. Wispr has 1000+ reviews, media coverage, marketing spend. Superwhisper is built by an indie developer, so less overhead, but still charges because solo development has real costs.

Recitey runs on a different model. The transcription engine (Whisper) is free and open source. Running it locally means no cloud metering, no per-message billing, no infrastructure costs. The unlimited free tier isn't a loss leader. It's just the actual cost of the technology.

That's worth understanding. Not because Recitey is moral or principled. But because it means your tools can serve your workflow instead of their business model.

You can dictate a 10,000-word design doc and it just keeps listening.

More posts
Keep reading

More like this.

  1. For developers

    Stop Hitting Word Caps in Your Design Docs

    Marcus is spending 11pm writing a design doc about payment settlement for a new feature. He's speaking clearly, the tool's...

  2. For developers

    The design doc that never got written

    It's 11 PM. Marcus, a backend engineer at a fintech in Stockholm, is voice-dictating a design doc for a new payment settlement...

  3. For developers

    Local transcription changes what you can dictate

    You wrote the design doc perfectly on voice, then scrolled up and realized 1400 words in, you'd stopped mid-sentence. Not...

All posts →