SSkilvy

AI that hears and sees

People used to talk to AI only in text. Now many assistants are multimodal: they understand voice, photos, screenshots, documents and tables. This changes how convenient they are — often it is faster to say or show than to type.

What it gives you in practice

  • Voice: dictate a question or email on the go.
  • Photo: snap a problem and ask what to do.
  • Screenshot: show an error or interface and ask for help.
  • Files: upload a PDF or table and ask about the content.
(photo of a plant) What plant is this and how do I care for it?
It looks like a peace lily (Spathiphyllum). It likes shade and humidity, water when the top soil dries, and mist the leaves. Want a short care sheet?

How to turn it on

Assistant apps have a microphone icon and a paperclip or camera icon — that is how you add voice, photos and files. Features depend on the service and plan; we will cover each in this course.

Rule: pick the format for the situation. Sometimes one photo explains the task better than three paragraphs of text.
VoicePhotoFile AI understands all Answer
Modern AI understands not only text but voice, images and documents.

🧠 What does "multimodal" AI mean?