You explained the payment settlement flow perfectly on the call. You're sitting at your desk at 11pm, still thinking through the edge cases, and you open Cursor to write the design doc.
Your voice tool starts transcribing. You're talking through the retry logic, the idempotency keys, the webhook signatures. Seven minutes in, mid-explanation, it stops.
"Word limit reached on free tier."
You've written 847 words. The cap was 500. You switch to typing. The thinking stops. The prose fragments. You finish in pieces the next morning, and it shows.
This is the shape of developer work now. You're not typing code. You're typing specs, design docs, PR descriptions, Slack threads that explain what the model should build. Long-form intent. Voice is faster than your fingers for that. But most voice tools were built for the old shape: quick notes, voice memos, transcribed meetings.
They priced and capped for that world.
Most voice tools treat free as a trial, not a tier
Wispr charges $14 a month. Willow charges $12. Superwhisper is $8.49 and indie. All three cap their free tiers. Wispr at 500 words per month. Willow at 600. They're not being mean. They're managing infrastructure costs.
But here's the thing: they're managing cloud infrastructure costs. If transcription is happening on your machine, not theirs, there's no per-word variable cost. The economics are structurally different.
Whisper, the underlying model, runs locally. It's been open-source since 2022. If the speech-to-text isn't leaving your device, and it's not consuming their compute, why does the free tier have a cap?
The answer is: it doesn't have to.
The architecture shift changes what you can offer
Recitey runs Whisper locally on your Windows machine. The transcription happens on your hardware. No API calls. No metering. No word counter because there's nothing to measure. You can dictate a 4000-word design doc in a single session and hit no limit.
This is not a marketing move. It's a consequence of the design.
The Pro tier handles the rewrite: takes your rough dictation (um, like, the the the) and polishes it in the cloud. That costs something. That should have a cap or be paid. But transcription itself? It's running on your machine. There's no rational reason to restrict it.
For Marcus, this meant something specific. He was losing context mid-doc because the free tier cut him off. He switched to Cursor partly because Cursor's tab-complete reduced the rewrites he'd need to do anyway. But the real shift was this: he could dictate the whole design doc without stopping, without losing the thread, without the artificial break that made him switch to typing and fragmented the thinking.
The trust problem with cloud-first voice tools
There's another reason developers adopt local-first voice tools: code is data. If you're dictating a design doc that talks through your API surface, your database schema, your error handling, you're being asked to send that to someone else's server.
"It won't be logged," they promise.
"Your data is safe," they say.
Most of them probably mean it. But the architecture is causal. If it's in the cloud, it can be logged. It can be indexed. It can be used to train a model. The policy might be "we don't do this," but the possibility exists.
If the transcription happens on your machine, the possibility doesn't exist. There's no server to send it to.
This isn't paranoia. It's just engineers thinking through dependencies.
What you actually give up
Running speech-to-text locally means Whisper's latency. It's slower than Wispr's cloud model because it's running on your GPU, not a farm of TPUs. A long design doc might take 30 to 45 seconds to transcribe. That's fine for 11pm writing. It's not fine for real-time customer conversations.
That's why the Pro tier exists. The cloud rewrite is faster. If you're writing a Slack message to a customer, you'll want the 2-second polish, not the 30-second transcription. You'll pay for that speed.
But for the work that's actually bottlenecked by thinking, not transcription, local is plenty fast. And uncapped.
The frame
The word counter on a free voice tier is not a feature. It's a breakpoint. It assumes you want to try the tool with a small task, then upgrade to the paid version for real work.
But if the tool understands the new workflow shape, the assumption is wrong. Your real work is the long-form thinking. Your trial is the quick Slack message. The tiers should reflect that.
Recitey does, because it's structured around what actually costs something: the cloud rewrite, not the local dictation.