Mosaic Voice
Included with Mosaic Pro

Vibe code with
your voice.

Driving AI coding agents means typing long, natural-language prompts — and research suggests most people speak far faster than they type. Mosaic Voice lets you say your prompt out loud and drops it straight into the agent pane you have focused. It runs locally, with no per-word credits, and comes with Mosaic Pro.

Available to Mosaic Pro members · Windows · your voice stays local

claude · agent pane
refactor the auth guard to fail closed and add a regression test

Spoken aloud · transcribed on-device · dropped into the focused pane

~150words/min speaking
~40words/min typing
~3×faster than typing
On-deviceno credits, no metering

Speaking-vs-typing figures reflect general research on dictation speed (people tend to type around 40 wpm and speak around 150), not measurements of Mosaic Voice. Your mileage varies.

The problem

Typing is the bottleneck when you drive agents.

Typing long prompts is slow

Driving an agent means writing paragraphs of natural-language instructions. Typing them out is the bottleneck — and it breaks your train of thought mid-sentence.

Juggling agents is context-switching

Run several sessions at once and you spend the day hopping between panes, re-focusing, and retyping. The friction adds up fast.

Heavy prompting is hard on the wrists

Hours of typing long prompts every day takes a physical toll. Speaking most of them instead is easier to sustain.

The voice-to-pane wedge

Speak it, and it lands in the right agent.

A plain dictation tool types wherever the operating system cursor sits — Mosaic Voice does that too, in any app on your machine. Inside Mosaic Terminal it goes one step further: it knows which agent pane you have focused, so what you say lands in that session even when several agents are running side by side. You pick the pane; Voice makes sure your words land there.

01

Speak

Say your prompt the way you would explain it to a teammate — no need to slow down and type it.

02

Transcribe on-device

Whisper turns your speech into text locally. Nothing is metered and nothing leaves your machine.

03

Lands in the focused pane

The text drops straight into the agent you have focused in Mosaic, ready to run.

Local-first

Runs offline, on your machine

Transcription happens on-device with Whisper. No round-trip to a server, no per-word credits, no metering. Your voice never has to leave your computer.

Voice-to-pane

Speaks straight to the right agent

Works in any app out of the box — and inside Mosaic, dictation can target the agent pane you have focused, not just wherever the OS cursor sits. You pick the pane; your words land there.

Your keys, your call

Optional cloud, bring your own key

Want a cloud model for a tricky accent or a language? Add your own ElevenLabs key. It stays on your machine, and we never mark it up or meter it.

Local-first, by default

Many cloud dictation tools stream your voice to a server, meter usage, and cost a separate subscription. Mosaic Voice runs the model on your machine and treats cloud as an optional extra you control.

Mosaic Voice

  • Transcribes on-device with Whisper — works offline
  • No credits and no per-word metering
  • Your voice stays on your machine by default
  • Optional cloud, using your own key, never marked up
  • Types into your focused Mosaic agent pane
  • Included with Mosaic Pro

Typical cloud dictation

  • Sends your audio to a remote server to transcribe
  • Usually metered — credits, minutes, or a monthly cap
  • Needs a connection to work at all
  • Types wherever the OS cursor happens to be
  • Often a separate subscription on top of your tools

Comparison describes common cloud-dictation patterns, not any specific product. Individual tools vary.

No separate subscription

It comes with Pro, not a monthly bill.

Voice input for your agents is core to how Mosaic is meant to be used, so it ships inside Mosaic Pro rather than as an add-on with its own recurring charge. Because the default engine runs on your machine, there is nothing for us to meter — you speak as much as you like at no per-word cost. If you already have Pro, Voice is included.

Questions

+Is Mosaic Voice included in Mosaic Pro?

Yes. It comes with Pro at no extra cost.

+Is it good for vibe coding and driving agents like Claude or Cursor?

Yes. It types into any input on your machine — Cursor, VS Code, a terminal, a browser tab. Inside Mosaic Terminal it can also target the agent pane you have focused instead of the OS-focused window.

+Does it work offline?

Yes. The default engine is a local Whisper model. Cloud transcription is optional, with your own key.

+Do I pay per word or per minute?

No. The local engine has nothing to meter. If you enable cloud transcription with your own key, you pay that provider directly.

+Which languages are supported?

English works best. Other languages depend on the available open models — quality varies. You can switch models in settings.

+When can I download it?

It is live. Pro members download it at mosaicterminal.dev/download; it activates with your Pro licence in the app.

+What do I need to run it?

A Windows or Linux machine and Mosaic Pro. Everything runs on your machine; the optional cloud features take your own API key. The Linux build is new — X11 works best.

Stop typing your prompts.

Speak them instead. Mosaic Voice ships to Mosaic Pro members at no extra cost — get Pro, and it arrives with your plan.

See Mosaic Pro — $149