Image & Video
Generative visual models create new images from learned visual patterns and a prompt.
They are not simply searching for an existing photograph. They generate a new visual arrangement that matches the prompt.
Your prompt
A vibrant Holi festival at Connaught Place, golden hour.
The prompt describes the scene and details you want.
Model
Learned visual patterns
The model learned relationships between language and visual structure from training data.
Drag from noise to image
Many modern image generators use diffusion-style processes that iteratively refine noise toward a visual consistent with the prompt. This interaction is a simplified illustration.
Pure noise
NoiseColourStructureRefineFinal
Why video is harder
An image must look coherent once. A video has to keep identity, objects, motion and lighting coherent across many frames.
Same character across three frames.
Key insight: image generation creates spatial coherence. Video generation must also preserve coherence over time.