Dictation: how to speak so it works
Voice and multimodal AI · Lesson 2 / 20
Two different modes
They're often confused, and the difference matters. Dictation turns your speech into the text of a request; the answer comes back as text. Voice conversation means you speak, the assistant replies aloud, and you get a dialogue.
Dictation is almost always better when you'll use the answer afterwards: copy it, edit it, send it. Voice conversation suits the case where your hands and eyes are busy and you need the answer once, right now.
How to speak so it's recognised
- Normal pace, normal words. Slowed-down "dictating" speech is recognised worse — these systems were trained on ordinary speech.
- Say the punctuation if structure matters: "full stop", "new line", "colon".
- Spell out difficult words. Names, companies, technical terms and addresses are where errors concentrate.
- Don't fear stumbling. Say "no, not that" and repeat — the model understands the correction from context.
The technique that saves the most
Don't try to dictate a perfectly phrased request. Say what you're actually thinking, with all the qualifications and repetitions, and add at the end: "work out what I mean and ask me if anything's unclear." The speech comes out messy and the result is better than a polished written request, because you put in more context.
Where dictation loses
Precise text with terms, figures and names. Anything containing addresses, reference numbers, codes. Multi-level lists and tables. Situations where it's noisy or other people can hear you.
Check what was recognised
Reread the text before sending, especially figures and names. A recognition error inside a number looks perfectly normal — "fifteen" became "fifty", and you get a confident answer to a question you never asked.
Cheat sheet
- Dictation if you'll use the answer afterwards.
- Voice conversation if your hands are busy.
- Normal pace; spell out difficult words.
- Always reread numbers and names.
Dictate the same real work or personal request two ways: (1) messily, aloud, with all your qualifications and a closing request to work out what you meant and ask questions; (2) as a pre-phrased short text. Compare the answers. Separately, dictate a text containing at least 3 numbers and 2 names, and check the recognition — what got distorted. Paste both requests, both answers and the recognition check.