Dream
Text-to-image with a fixed seed, so you can change one word in the prompt and see precisely what that word was doing.
Generated images collect here for this session only. Nothing is uploaded or stored on a server — refresh the page and they are gone.
This is the only clean way to read a prompt. With the seed fixed, a single changed word produces a single traceable change in the image. Change the seed as well and you are comparing two unrelated pictures.
It begins with a field of pure random noise. Guided by your prompt, the model repeatedly predicts which part of what it sees is noise and subtracts it. After a few dozen steps an image is left. The seed decides the starting noise, which is why the same seed and prompt reproduce the same picture.
Trained on a large scrape of web images, with the aesthetic and demographic skew that implies. Counting, spelling and precise spatial relationships remain weak.
- Word order and emphasis matter more than word count. Terms near the start of the prompt carry more weight.
- Same seed, one word changed, is the only clean way to see what a word contributes. Everything else changes too much at once.
- It has no model of physics or anatomy — it has a model of what pictures of those things tend to look like. Hands and text are the usual tells.
This page says “the model predicts”, not “the model knows”. That is deliberate. None of these systems understand the images or sentences you give them; they map inputs to outputs using patterns fixed at training time. The difference matters most exactly when the output is impressive.