AI Video & Storyboard Glossary
Plain-English definitions of the terms behind AI video and storyboard generation — from storyboard and shot list to the generation modes (t2v, i2v, r2v) you will meet in any modern tool.
Storyboarding & pre-production
- Storyboard
- A sequenced set of shots — each shown as a frame or description with framing and action — that lets you see a film before you shoot it. In AI workflows, each panel often carries the prompt that will generate it.
- Shot list
- A structured, text breakdown of every shot in a scene: shot number, type (wide / medium / close-up), camera angle, movement, subject, and duration. The shot list is the spreadsheet version of a storyboard.
- Animatic
- A timed slideshow of storyboard frames, often with scratch audio, used to test pacing before any footage exists. An animatic is a presentation of stills — not generated video.
- Logline
- A single sentence that captures a story: protagonist, goal, obstacle, and stakes. It is the brief that every later creative decision is judged against.
- Treatment
- A short prose summary of the whole story — usually one to two pages, no dialogue — written before the screenplay so structural problems can be fixed cheaply.
- Screenplay
- The formatted script of a film: scene headings, action, and dialogue. In an AI storyboard workflow, the screenplay is locked first, then broken into shots.
- Establishing shot
- A wide shot, usually opening a scene, that sets the location and spatial geography so the audience knows where they are before the action tightens in.
- Cut
- The transition from one shot to the next. Planning the cut between shots is what makes a storyboard read as continuous motion rather than a gallery of separate images.
AI generation modes
- Text-to-video (t2v)
- Generating a video clip directly from a text prompt, with no input image.
- Image-to-video (i2v)
- Animating a still image: you supply a starting frame and the model generates motion from it.
- Reference-to-video (r2v)
- Generating video from reference inputs that define a character, style, or subject — without designating a strict first frame. On multimodal models the references can include not just images but reference videos and audio (see video-to-video).
- First-last-frame (firstlast)
- Supplying both a starting and an ending image; the model generates the in-between motion that connects them.
- Text-to-image (t2i)
- Generating a still image from a text prompt.
- Image-to-image (i2i)
- Editing or restyling an existing image with a prompt while preserving its core content — used for outfit swaps, expression changes, and reference cleanup.
AI video concepts
- Character consistency
- Keeping a character looking like the same person — face, outfit, build — across multiple shots and generations, by binding identity reference images and a persona to a reusable character rather than re-describing it in each prompt.
- Video prompt
- The structured description sent to a video model for one shot — typically a first frame, the motion through the shot, and where it lands — plus any reference images and style.
- Multi-model platform
- A platform that gives access to many AI models (e.g. Sora 2, Kling, Seedance, VEO) from one account, so you can route each shot to the model that suits it instead of being locked to one engine.
- Frame chaining
- Using the last frame of one shot as the first frame of the next, so a continuous action stays coherent across a cut.
- Upscaling (super-resolution)
- Increasing the pixel resolution of an image or video — e.g. 1080p to 4K — while an AI model reconstructs plausible fine detail (edges, texture) instead of just stretching pixels. Commonly used to finish AI-generated footage generated cheaply at a lower resolution, then upscaled. FlyAIgh’s Upscaler runs on the Topaz Labs API.
Video-to-video & references
- Video-to-video (v2v)
- Generating a new video guided by an existing one, rather than from text alone. The reference clip can drive the style, motion, or subject of a fresh generation. On modern multimodal models this is done by passing the source as a reference video alongside a prompt — not by repainting the original frame by frame.
- Style transfer
- Carrying the visual aesthetic of one piece of media — its color, era, or animation style — onto new content. In AI video, style transfer is one of the most common video-to-video jobs (for example, turning live footage into anime).
- Reference video
- A video clip supplied as an input to guide a generation — its look, camera move, or motion is carried into the new output. Unlike a first frame, a reference video is a directing signal, not the literal starting image.
- Reference audio
- An audio track supplied as an input so the generated video moves to it or incorporates it — used for music-led or audio-synced generation. Multimodal models like Seedance 2.0 accept reference audio alongside reference images and video.
Editing & continuity
- Match cut
- A cut where a shape, motion, or composition in one shot carries into the next, so the transition feels deliberate and continuous. The opposite of a jarring, unrelated cut.
- J-cut and L-cut
- Edits where the audio and picture change at different moments. In a J-cut the next shot’s sound starts before its picture; in an L-cut the previous shot’s sound lingers into the next. Both smooth the seam between shots.
- Jump cut
- A cut between two shots of the same subject that are only slightly different, creating a deliberate jump in time or position. Used for energy or to compress time; avoided when invisible continuity is the goal.
- 180-degree rule
- A continuity guideline: keep the camera on one side of an imaginary line between two subjects so their left/right positions stay consistent across cuts. Crossing the line disorients the viewer.
Camera language (shots & movement)
- Wide shot (long shot)
- A shot that frames the subject within its full surroundings, showing the whole scene and where everyone stands. In AI video it is the establishing beat — generate it first to lock the geography, then cut to tighter shots.
- Medium shot
- A shot framing a person roughly from the waist up — close enough to read expression, wide enough to show gesture. The workhorse shot for dialogue.
- Close-up
- A shot that fills the frame with a single subject — usually a face — to carry emotion or detail. In AI video, close-ups are where character consistency matters most, because the face is large.
- Pan
- A camera move that pivots horizontally from a fixed point, sweeping left or right across a scene. In an AI video prompt, describe it as "the camera pans left/right" rather than moving the subject.
- Tilt
- A camera move that pivots vertically from a fixed point, tipping up or down — used to reveal height or follow a subject as they rise. The vertical counterpart of a pan.
- Dolly shot
- A shot where the whole camera physically moves toward or away from the subject on a track, changing the perspective (unlike a zoom, which only magnifies). "Dolly in" builds intensity; "dolly out" reveals context.
- Tracking shot
- A shot where the camera travels alongside or behind a moving subject, holding it in frame as it moves. Also called a traveling or follow shot — the backbone of dynamic AI video motion prompts.
- Zoom
- Changing the lens focal length to magnify or widen the view without moving the camera. A zoom flattens space; a dolly changes it — they look different even when the framing matches.
- Low-angle shot
- A shot taken from below the subject, looking up, which makes the subject feel powerful, large, or dominant. The opposite of a high-angle shot.
- High-angle shot
- A shot taken from above the subject, looking down, which tends to make the subject feel small, vulnerable, or observed.
- Dutch angle
- A shot where the camera is tilted so the horizon runs diagonally, creating unease, tension, or disorientation. Also called a canted or oblique angle.
- Over-the-shoulder shot
- A shot framed from behind one person’s shoulder onto another, anchoring a conversation in space. The staple two-character dialogue setup, and where the 180-degree rule keeps eyelines consistent.
Ready to put these into practice?