Multimodality: what it changes in practice
Voice and multimodal AI · Lesson 1 / 20
An AI that hears and sees
Until recently you talked to an assistant only in text. Now most of them accept voice, photographs, screenshots, documents and spreadsheets. That sounds like a technical detail, and it changes the important thing — when you turn to AI at all.
A text request means sitting down, opening something, phrasing it and typing. Because of that, a great many small questions never get asked: it's easier to shrug. Voice and camera remove that barrier, and the assistant starts being used in situations where it was previously useless — in the kitchen, in a shop, on the road, standing in front of something broken.
Four channels and what each gives you
- Voice. Speed. You speak roughly three times faster than you type, and you can do it with your hands full.
- Photo. Context without description. Showing a broken tap is faster and more precise than describing it.
- Screenshot. Accuracy. You don't retell the error, you show it — along with everything around it.
- Files. Volume. You can't paraphrase a forty-page contract into a request, but you can upload it.
The rule for choosing a channel
Choose not by convenience but by what is cheapest to convey. If explaining in words takes three sentences and still isn't precise — show it. If the question is simple but your hands are busy — say it. If you need exact text with specific terms — type, because dictation stumbles on names and terminology.
What multimodality doesn't change
The channel has no effect on the quality of the model's thinking. A badly framed question stays badly framed when spoken. Errors and inventions don't go anywhere — and people trust an answer given about a photograph more readily than one given about text, with no grounds for doing so.
Cheat sheet
- Voice for speed, photo for context, screenshot for accuracy, file for volume.
- Choose the channel by what's cheapest to convey.
- The barrier drops — you ask more.
- The model's thinking doesn't depend on the channel.
Recall 6 situations from the past week where you needed a quick answer but didn't ask AI — because it was awkward, slow or hard to explain. For each, decide which channel would have suited it better (voice, photo, screenshot, file) and why. Then pick two and actually do them through the chosen channel. Describe what happened and whether it matched your expectation.