Marcus is a backend engineer at a Series B fintech in Stockholm. At 11pm on a Tuesday, he's in Cursor documenting a payment settlement redesign, explaining the logic out loud instead of typing. The cloud dictation tool he's been using caps him at 1,500 words per day on the free tier. He hits that cap halfway through the third section. The flow breaks. The architectural reasoning that made sense in his voice becomes fragmented prose across Slack threads and morning cleanup.
This is not a productivity problem. It's a problem of thinking. When the tool interrupts you mid-thought because some vendor's business model requires a paywall, you don't just lose the last sentence. You lose the momentum of the idea.
The Shift From Typing Code to Typing Intent
Five years ago, a backend engineer's day was predictable: open a file, type code, run tests, repeat. The keyboard was the interface between thought and implementation. Voice dictation tools existed, but they solved a niche problem, faster note-taking, maybe dictating emails while driving.
That has changed. Today, an engineer's day involves explaining what the code should do. Explaining to Claude what the algorithm should handle. Explaining to Cursor what file to modify and why. Explaining to Copilot what the edge case is. The keyboard is still the same. The work is not.
The bottleneck did not move from typing to dictating. It moved to explaining. The difference matters. Typing code is a precision task. Explaining intent is a thinking task. The two have different requirements.
When you type code, you can pause, edit, reconsider. The tool is a mirror. When you explain intent, you need flow. The tool should disappear. But most dictation tools cannot disappear, they interrupt. They meter. They prompt for payment.
This is where the word limit becomes a breaking point. A 600-word cap on Wispr Flow's free tier, or 1,500 words daily across most free tiers, forces you to choose: split the thought across multiple sessions, or type the long-form explanation yourself. Neither preserves the momentum of the thinking.
The Word Limit Is Not an Accident
Cloud dictation runs on servers. Every word transcribed costs something: storage, bandwidth, licensing fees from whoever built the transcription model. Wispr Flow charges $14 a month for unlimited words; the free tier caps at 600. Willow charges $12 and caps free tier words. Superwhisper charges $8.49 and applies limits. The meter is not a technical limitation. It is a business model.
Vendors are transparent about this. The cost of running transcription at scale is not zero. So they meter. The meter exists to funnel you toward a subscription the moment you outgrow the free tier.
The problem is timing. You hit the limit in the middle of the thought. Not before. Not after. Mid-thought. The interruption is the damage.
Local Whisper Changes the Economics
Recitey runs speech-to-text on your device using Whisper, the same model cloud vendors license from OpenAI. There is no server cost per word. No bandwidth meter. No API count. Which means no reason to cap.
The free tier dictates everything you speak. No countdown timer. No "you have reached your daily limit" notification mid-sentence. You are only charged for what goes up to the cloud, the rewrite that polishes rough draft into clean prose, turning "and then the uh the settlement uh like if the bank rejects it we need to" into "If the bank rejects the settlement, we initiate a reversal and log the event for manual review."
This is structural. It is not free as a marketing statement. It is free because there is no infrastructure cost to meter. The economics of local transcription are different from cloud transcription.
Why This Matters to Developers Working With LLMs
When Marcus sits down to spec a payment settlement algorithm, he is not describing what the code should be. He is describing what the code should decide. The constraints. The edge cases. The STP retry logic. The handling of partial reversals. Why a certain decision was made. What could break. All of that is faster out loud than typed.
But only if the tool does not interrupt.
Most developers who use voice dictation right now have hit the same wall. Halfway through explaining the bug, the word limit shows up. Halfway through the design doc, the meter pings. The thought gets split. By the time you finish explaining, momentum is gone. You have to switch into editing mode. The flow breaks.
Marcus uses Cursor, not VS Code, specifically because Cursor's tab-complete reduces the number of times he has to rephrase voice-dictated explanations. It is not just about speed. It is about flow. He refuses cloud transcription for a more fundamental reason: a payment settlement algorithm is proprietary. Design docs are proprietary. Code details are company secrets. The moment you send that audio to someone's API, you have made a choice about where your work lives. You have made a choice about whether you trust that vendor with your company's details.
Local transcription means the audio never leaves your machine. Your explanation stays yours. Your algorithm stays yours. The vendor cannot accidentally leak it. The vendor cannot use it to train their next model. It is genuinely local.
The Workflow After
Marcus opens Recitey in Cursor. He speaks for 9 minutes straight, explaining how the settlement engine should handle edge cases around failed STP transfers. No word counter in his head. No calculation of "am I at 1,200 words yet?" The rough draft appears in his buffer. The moment he stops, the optional rewrite polishes it into coherent prose, proper punctuation, sentence structure, clarity. He pastes it into the design doc. The thinking is preserved. The flow was never interrupted.
This is not about dictating faster. It is about thinking without a meter in the room.
He also uses it for PR review comments when he's explaining a subtle decision to a teammate, or Slack threads when he's investigating an incident that happened at 3am. The moment the explanation gets long, most dictation tools start asking for money. Recitey does not. He speaks until the thought is done. Full stop.
The design docs end up more coherent. The PR comments are more thorough. The Slack explanations actually explain, instead of getting cut short and then having to type out the rest in fragments.
The Trade-off You Accept
The free tier includes local dictation, unlimited words, but no cloud rewrite. If you want the polish, the automatic sentence restructuring, the punctuation sweep, the sense-check, that is Pro. The trade-off is real.
But it matters less than you think. Most of the time, the rough draft from Whisper is clean enough to paste directly. The audio quality matters; the algorithm is solid. And for the moments when the rough draft needs work, Pro is not $14 a month. You pay for what you use.
The real difference is permission. Permission to think out loud. Permission to explain for as long as the thought needs explaining. Permission to not have someone's server hit a wall and interrupt you.
That permission breaks the meter.