Driving AI coding agents means typing long, natural-language prompts — and research suggests most people speak far faster than they type. Mosaic Voice lets you say your prompt out loud and drops it straight into the agent pane you have focused. It runs locally, with no per-word credits, and comes with Mosaic Pro.
Available to Mosaic Pro members · Windows · your voice stays local
Spoken aloud · transcribed on-device · dropped into the focused pane
Speaking-vs-typing figures reflect general research on dictation speed (people tend to type around 40 wpm and speak around 150), not measurements of Mosaic Voice. Your mileage varies.
Driving an agent means writing paragraphs of natural-language instructions. Typing them out is the bottleneck — and it breaks your train of thought mid-sentence.
Run several sessions at once and you spend the day hopping between panes, re-focusing, and retyping. The friction adds up fast.
Hours of typing long prompts every day takes a physical toll. Speaking most of them instead is easier to sustain.
A plain dictation tool types wherever the operating system cursor sits — Mosaic Voice does that too, in any app on your machine. Inside Mosaic Terminal it goes one step further: it knows which agent pane you have focused, so what you say lands in that session even when several agents are running side by side. You pick the pane; Voice makes sure your words land there.
Say your prompt the way you would explain it to a teammate — no need to slow down and type it.
Whisper turns your speech into text locally. Nothing is metered and nothing leaves your machine.
The text drops straight into the agent you have focused in Mosaic, ready to run.
Transcription happens on-device with Whisper. No round-trip to a server, no per-word credits, no metering. Your voice never has to leave your computer.
Works in any app out of the box — and inside Mosaic, dictation can target the agent pane you have focused, not just wherever the OS cursor sits. You pick the pane; your words land there.
Want a cloud model for a tricky accent or a language? Add your own ElevenLabs key. It stays on your machine, and we never mark it up or meter it.
Many cloud dictation tools stream your voice to a server, meter usage, and cost a separate subscription. Mosaic Voice runs the model on your machine and treats cloud as an optional extra you control.
Comparison describes common cloud-dictation patterns, not any specific product. Individual tools vary.
Voice input for your agents is core to how Mosaic is meant to be used, so it ships inside Mosaic Pro rather than as an add-on with its own recurring charge. Because the default engine runs on your machine, there is nothing for us to meter — you speak as much as you like at no per-word cost. If you already have Pro, Voice is included.
Yes. It comes with Pro at no extra cost.
Yes. It types into any input on your machine — Cursor, VS Code, a terminal, a browser tab. Inside Mosaic Terminal it can also target the agent pane you have focused instead of the OS-focused window.
Yes. The default engine is a local Whisper model. Cloud transcription is optional, with your own key.
No. The local engine has nothing to meter. If you enable cloud transcription with your own key, you pay that provider directly.
English works best. Other languages depend on the available open models — quality varies. You can switch models in settings.
It is live. Pro members download it at mosaicterminal.dev/download; it activates with your Pro licence in the app.
A Windows or Linux machine and Mosaic Pro. Everything runs on your machine; the optional cloud features take your own API key. The Linux build is new — X11 works best.
Speak them instead. Mosaic Voice ships to Mosaic Pro members at no extra cost — get Pro, and it arrives with your plan.
See Mosaic Pro — $149