What an AI avatar actually does
AI avatars and virtual presenters · Lesson 1 / 20
Two independent machines under one video
An AI avatar is a synthesized video of a person saying words they never said. Technically there are two independent systems inside. The first turns text into sound: speech synthesis, which may use a clone of your voice or a stock one. The second turns sound into picture: a model predicts how lips, jaw, brows and head should move for each phoneme and redraws them on the source face.
This split is not a technical detail — it is the main working tool of this course. Almost every complaint about a finished clip lives in exactly one of the two layers. "Sounds like a robot reading a manual" is the speech layer; the picture has nothing to do with it. "The mouth lags behind the sound" or "the face goes rubbery on long vowels" is the video layer, and rewriting the script will not help. People who do not separate the two spend months editing scripts hoping to fix lip sync.
What avatars do well
- Repeatability. Twenty clips of identical quality in an identical frame — a task that exhausts a live presenter by take five.
- Edits without reshoots. A price changed in the offer: you change one line of text, not reassemble a studio.
- Languages. One script becomes eight localized versions over a lunch break.
What avatars do badly
- Emotion that is not in the text. The model voices what is written, not what is felt. An apology to customers from an avatar reads as mockery.
- Spontaneity. Conversation, reaction, real laughter get imitated — and the imitation shows.
- Anything that must prove presence. Congratulations, condolences, a founder's personal address after a failure.
Cheat sheet
- Avatar = speech synthesis + facial synthesis; separate layers, separate problems.
- Strengths: repeatability, cheap edits, languages.
- Weaknesses: real emotion, spontaneity, proof of presence.
- The comparison point is not a shoot, it is the absence of a video.