AI that hears and sees
People used to talk to AI only in text. Now many assistants are multimodal: they understand voice, photos, screenshots, documents and tables. This changes how convenient they are — often it is faster to say or show than to type.
What it gives you in practice
- Voice: dictate a question or email on the go.
- Photo: snap a problem and ask what to do.
- Screenshot: show an error or interface and ask for help.
- Files: upload a PDF or table and ask about the content.
(photo of a plant) What plant is this and how do I care for it?
It looks like a peace lily (Spathiphyllum). It likes shade and humidity, water when the top soil dries, and mist the leaves. Want a short care sheet?
How to turn it on
Assistant apps have a microphone icon and a paperclip or camera icon — that is how you add voice, photos and files. Features depend on the service and plan; we will cover each in this course.
Rule: pick the format for the situation. Sometimes one photo explains the task better than three paragraphs of text.
🧠 What does "multimodal" AI mean?