Skip to main content
Vox picks up your voice from very close range, which is why it can hear things a laptop microphone cannot. That changes what counts as good technique.

You can speak much more quietly than you think

A normal speaking voice is more than enough. A quiet, low voice at conversational pace usually transcribes just as well, which is the point: you can dictate in a shared office without narrating your work to the room.

When whispering works

True whispering, with no voice behind it, is harder. Whispered speech loses the pitch information that makes vowels distinct, so accuracy drops for everyone, not just for Vox. What tends to work:
  • Quiet voiced speech, sometimes called stage whisper. Keep the vocal cords engaged and just turn the volume down. This is nearly as accurate as normal speech.
  • Slowing down slightly. Whispered consonants are the first thing to get lost.
  • Short sentences. Less context for a misheard word to cascade through.
What tends not to work:
  • Pure unvoiced whispering for long passages.
  • Whispering technical terms, names, and acronyms. Say those at normal volume.

In a noisy room

Vox handles this well, because it is much closer to your mouth than the noise is. Speak normally. Do not raise your voice, which changes how you articulate and usually hurts more than the noise does.
TODO: does noise suppression handle this? There is a DeepFilterNet option, so say whether it is on by default and whether it helps here.
TODO
TODO: cover whether the speaker gate is shipping and what it does.

Things that help regardless

  • Pause before you press. Starting to talk before the button registers clips your first word.
  • Punctuate out loud if you want punctuation. Say “comma”, “period”, “new line”.
  • Add names to your vocabulary. Settings has a vocabulary list. Proper nouns, product names, and jargon are worth adding once and never fighting again.
  • Do not trail off. The last word of a sentence is the most commonly dropped one.

Noise suppression

Vox can run neural noise suppression on the audio before transcribing it. TODO: say where the setting is, what the default is, what it costs in latency, and in which situations it is worth turning on. Note that it is currently Mac only.