AI VideoAI Video · Lesson 01
AI Video: The Landscape
What generative video actually does, what it can't do, and how a modern pipeline fits together.
Video tutorialTrack course · Youri van Hofwegen
Speed
Transcript & captions
10/10
AI video generation turns text, images or existing footage into new moving images using diffusion and transformer models. Instead of filming, you describe — then direct the model through iteration. The skill is no longer camera operation; it is specification, selection and assembly.
- Text-to-video (T2V): a prompt becomes a clip, usually 4-10 seconds.
- Image-to-video (I2V): a still frame is animated — the most controllable route.
- Video-to-video (V2V): restyle or re-time existing footage.
- Lip-sync / avatar: a face plus audio becomes a talking presenter.
| Stage | Tool type | Output |
|---|---|---|
| Concept | Chat model | Logline, beat sheet, shot list |
| Stills | Image model | Keyframes, character sheets |
| Motion | Video model | 4-10s clips per shot |
| Audio | TTS + music model | VO, score, SFX |
| Assembly | Editor | Final cut, colour, export |
Nearly every good AI video is many short clips edited together — not one long generation. Plan in shots, not scenes.
Knowledge check
0/2 answeredWhich workflow gives you the most control over composition?
Typical single-generation clip length is about...