AI VideoAI Video · Lesson 01

AI Video: The Landscape

What generative video actually does, what it can't do, and how a modern pipeline fits together.

Video tutorialTrack course · Youri van Hofwegen
Speed
Next lesson
Watch 00:00 · 2 checkpoints Watch on YouTube More on this topic
Transcript & captions
10/10

AI video generation turns text, images or existing footage into new moving images using diffusion and transformer models. Instead of filming, you describe — then direct the model through iteration. The skill is no longer camera operation; it is specification, selection and assembly.

  • Text-to-video (T2V): a prompt becomes a clip, usually 4-10 seconds.
  • Image-to-video (I2V): a still frame is animated — the most controllable route.
  • Video-to-video (V2V): restyle or re-time existing footage.
  • Lip-sync / avatar: a face plus audio becomes a talking presenter.
StageTool typeOutput
ConceptChat modelLogline, beat sheet, shot list
StillsImage modelKeyframes, character sheets
MotionVideo model4-10s clips per shot
AudioTTS + music modelVO, score, SFX
AssemblyEditorFinal cut, colour, export

Nearly every good AI video is many short clips edited together — not one long generation. Plan in shots, not scenes.

Knowledge check

0/2 answered

Which workflow gives you the most control over composition?

Typical single-generation clip length is about...