What the model does with your text

AI image generation · Lesson 1 / 20

It draws what's written, not what's intended

Generators are trained on millions of image-and-caption pairs. You supply text, the model assembles the most probable picture for that description. From which the essential point follows: it draws what you wrote, not what you pictured.

The gap between those two things is the source of nearly every disappointment. In your head there's a specific scene with details you never articulated, because they seem obvious. For the model they don't exist.

What doesn't exist until it's stated

  • The angle. Unstated, you get a random one — usually a frontal mid-shot.
  • The light. By default it'll be flat studio lighting, rarely what you imagined.
  • What else is in frame. The model fills the background with something plausible, usually superfluous.
  • The mood. Unspecified, you get pleasantly neutral — the most averaged option available.

Why negative phrasing works badly

"No text in the background", "not cartoonish" are unreliable, because the model works with probabilities rather than prohibitions: a mentioned word raises the likelihood of related elements appearing. It's more reliable to say what should be there: instead of "not cartoonish", say "photorealistic, shot on 50mm".

The practical conclusion

Before writing a prompt, describe the picture aloud as though explaining it over the phone to someone who has to draw it. Everything you say goes into the prompt. Everything that "goes without saying" goes in too, because it doesn't go without saying for the model.

Insight. The gap between what's intended and what's written causes nearly every failure. Describe the picture as if explaining it by phone to someone who must draw it.
Common mistake. Using negations like "no text" and "not cartoonish". The model works with probabilities, not prohibitions, and a mentioned word is more likely to summon an element than remove it.
Pro tip. Everything that feels "obvious" — angle, light, what else is in frame — doesn't exist for the model. Those omissions are exactly what spoils the result.

Cheat sheet

  • The model draws what's written, not what's intended.
  • Angle, light, surroundings and mood are stated explicitly.
  • Say what should be there, not what shouldn't.
  • Test: would you explain it this way by phone?
1. What's the main cause of failed generations?
2. Why does "without something" phrasing work badly?
3. What happens if you don't specify the light?
Task — checked by AI

Take a picture you want to produce and write its description two ways: first as you'd normally write a prompt, then as you'd explain it by phone to someone who has to draw it, with angle, light, surroundings and mood. Generate from both. Paste both descriptions, describe the difference in results, and list which details in the second version proved decisive.

🔒 Answer the question correctly to move on to the next lesson.

What the model does with your text — AI image generation — Skilvy