Image & Video

Generative visual models create new images from learned visual patterns and a prompt.

They are not simply searching for an existing photograph. They generate a new visual arrangement that matches the prompt.

Your prompt

A vibrant Holi festival at Connaught Place, golden hour.

The prompt describes the scene and details you want.

Model

Learned visual patterns

The model learned relationships between language and visual structure from training data.

Drag from noise to image

Many modern image generators use diffusion-style processes that iteratively refine noise toward a visual consistent with the prompt. This interaction is a simplified illustration.

Pure noise
NoiseColourStructureRefineFinal

Why video is harder

An image must look coherent once. A video has to keep identity, objects, motion and lighting coherent across many frames.

Same character across three frames.

Key insight: image generation creates spatial coherence. Video generation must also preserve coherence over time.
Next foundation: Vibe Coding. Continue →