← BlogFor developers

The word cap hits at the worst possible time

When you're drafting a design doc at 11pm and you're in the flow, the last thing you need is a tool that stops listening. But that's what happens with cloud-based voice transcription. You hit the monthly cap. The tool goes silent. Your thought fragments. And now you've got a half-finished paragraph that reads like someone interrupted mid-sentence, which is exactly what happened. The next morning you stare at fragmented prose and spend 30 minutes rewording what should have taken 20 minutes to dictate.

The moment Marcus recognizes

This is how Marcus works. Backend engineer at a Series B fintech in Stockholm. He's been writing code for eight years, but the last 18 months changed his workflow fundamentally. He started using Cursor three months ago, specifically for the tab-complete feature that predicts code based on comments and context. VS Code didn't have it, and the time savings were immediate. But more importantly, Cursor reduced how many times he had to rewrite his voice transcriptions. Better predictions meant fewer "wait, let me rephrase that" moments.

But Cursor didn't change everything about his workflow. Design docs still can't be typed. API contracts at 11pm are better thought through out loud. Incident postmortems need narrative flow that dictation captures better than typing. So Marcus drafts on voice. Long form. Multiple thoughts chained together.

That's when he hits the wall. Wispr Flow caps the free tier. Superwhisper caps the free tier. Willow caps the free tier. Each has a monthly limit that made sense when voice transcription cost real money to process. But Marcus's design doc is 2,400 words. His postmortem is 1,800. His API contract spec is 3,100. He's dictated 1,200 words and Wispr goes silent. The cap he didn't think about until that exact moment.

Why this wasn't a problem before

Five years ago, Marcus would've just typed these documents. Slower, but straightforward. The bottleneck was clear: typing is slow compared to thinking.

But the shift to LLM-assisted development changed the math. When you're writing prompts instead of writing code, when half your day is explaining to Cursor what the code should do instead of writing the code yourself, the volume of words you produce went way up. A function used to be 20 lines of code and maybe 2 lines of comments. Now it's 100 lines of prompt context and 20 lines of generated code. The commentary is now the primary artifact.

Design documents exploded. They're not 400-word briefs anymore. They're 2,500-word narratives that spell out context, assumptions, failure modes, edge cases. You're writing a thesis for an LLM to understand. A 400-word doc means the model doesn't have enough context. A 2,500-word doc means you've thought through the problem properly.

Same with code reviews. They used to be "looks good, approved." Now they're 800-word walkthroughs of the architectural decisions, the tradeoffs considered, the future implications. You're not reviewing the code. You're explaining the intent behind it.

Dictating these documents faster than typing them makes sense. But the voice tools still price themselves like transcription was expensive. Because for them, it used to be.

Why cloud voice tools have to cap you

The answer is technical and economic. Most voice transcription tools send your audio to the cloud. Whisper runs on a server somewhere (OpenAI's in most cases). Every minute of audio costs them compute. Storage. Bandwidth. Inference. They can't offer unlimited free transcription. The math doesn't work.

So they cap you. Wispr Flow is $14 a month for unlimited. Willow is $12 a month. Superwhisper is $8.49, indie, still capped on free. The pricing isn't malicious. It's just proportional to their infrastructure cost.

But the moment they do this, they've created a tier system. Free users, people like Marcus, hit the wall and stop using it. Paid users get to use it freely but now you're on a treadmill. You hit the cap, you pay for unlimited, you don't think about it anymore. That's a customer for $14/month.

What happens when transcription is free

What if transcription didn't cost anything? What if the bottleneck moved somewhere else entirely?

That's the shift Recitey made. Whisper runs locally on your device. Your audio never goes to a server. It runs on your GPU (or CPU, slowly). The only cost is what it costs to ship you the software. No per-transcription meter. No cap.

The free tier is uncapped because running Whisper on your machine is essentially free for the company. The Pro tier exists for the rewrite layer, the feature that takes your rough transcript and cleans it into polished prose. That's a stateless operation that runs on your device or a cloud service briefly. It's not metered per transcription. It's a feature difference.

For Marcus, this means he stops watching the word counter. He can dictate his 2,400-word design doc without the tool stopping him at 1,800. The interruption is gone. His thinking doesn't fragment. He finishes the thought.

The engineering reason it matters

There's another reason Marcus specifically likes this. Local transcription means his code stays on his machine. Design docs with proprietary algorithm details, security architecture, internal naming conventions, none of it leaves his laptop. He refuses cloud transcription for payment systems because his company's IP policy is explicit. On-device Whisper solves that problem completely.

That's not a privacy theater argument. That's a realistic constraint for anyone working on financial systems, healthcare data, embedded algorithms, or anything else where the drafting process itself is confidential.

The real tradeoff (and it matters)

Recitey only runs on Windows right now. If you're on a Mac, Willow and Wispr are still the way to go. That's not a minor thing. It's a real constraint.

But if you work on Windows, and your job involves thinking out loud into design docs, specs, postmortems, code comments, and detailed explanations for LLMs, the word cap isn't a minor inconvenience. It's a structural problem with how you work. You've optimized your workflow around speaking longer because speaking longer makes better prompts, better explanations, better documents. The tool that stops you mid-thought isn't saving you time. It's misaligned with your actual work.

More posts
Keep reading

More like this.

  1. For developers

    Stop Hitting Word Caps in Your Design Docs

    Marcus is spending 11pm writing a design doc about payment settlement for a new feature. He's speaking clearly, the tool's...

  2. For developers

    The design doc that never got written

    It's 11 PM. Marcus, a backend engineer at a fintech in Stockholm, is voice-dictating a design doc for a new payment settlement...

  3. For developers

    Local transcription changes what you can dictate

    You wrote the design doc perfectly on voice, then scrolled up and realized 1400 words in, you'd stopped mid-sentence. Not...

All posts →