No magic β just description
Image generators (Midjourney, DALLΒ·E, Stable Diffusion and others) are trained on millions of "image + caption" pairs. You give text β the model assembles the most fitting image.
Key idea: the model draws what you described in words, not what you pictured in your head. The more precise the description, the closer the result.
What to grasp right away
- The result is probabilistic. The same prompt yields different images β that's normal.
- Details decide. "A cat" vs "a ginger cat by a window in warm sunset light" β worlds apart.
- Iteration is inevitable. The first image is a draft; then you refine.
Three beginner expectations
- "It'll read my mind." No β describe it explicitly, or you'll get an average version.
- "The very first image will be perfect." Usually not: budget 3β5 tries.
- "Naming the object is enough." The object is just one of four layers (next lesson).
Exercise
Generate an image from the word "coffee", then from "a cup of espresso on a wooden table, morning light from a window, close-up". Compare β you'll instantly feel what details decide.
In this course you'll learn to describe an image so you get what you want in a few tries, not by luck.
π§ Why does one prompt give different images?