Stills First: The 5-Stage AI Video Pipeline That Stops Credit Burn
An AI video pipeline in five stages - plus the stage-4 rule that stops credit burn: never prompt video directly, generate stills first. With seven upgrades I would make to it.
Wizy produces AI video in five stages: idea, script, scene breakdown, image-then-video generation, and audio finishing. The single rule that saves the most money sits in stage 4 — never prompt a video model directly. Generate still images first, approve them, and only then animate the approved frames. Video generation costs far more per second than image generation, and a clip you reject costs exactly as much as one you keep.
Below is the pipeline as it runs today, followed by seven upgrades I would make to it.
The five stages at a glance
| Stage | The work | Tools | Output |
|---|---|---|---|
| 1. Idea | Define topic, objective, audience, key message | Gemini | Video concept |
| 2. Script | Hook → Content → Climax → CTA | NotebookLM, Gemini | Finished script |
| 3. Scene breakdown | Split into scenes; define characters, setting, action | Whisk AI, CapCut Pro | Storyboard / scene list |
| 4. Image → video | Detailed prompt per scene; stills first, then animate | Veo, Runway, Gemini, CapCut Pro | Per-scene clips |
| 5. Audio & effects | Voice sync, music, ambience, colour, subtitles | CapCut, ElevenLabs | Final video |
Stage 1: Idea — decide the message before you touch a tool
Every video starts by naming four things: the topic, the objective, the audience, and the one key message. The source is usually something that already exists — a blog post, a grand opening, a promotion. Gemini turns that raw material into a first script direction.
The discipline here is subtraction. A video that carries two messages carries none.
Stage 2: Script — Hook → Content → Climax → CTA
The script follows a fixed four-beat shape. The hook earns the first three seconds, the content delivers the substance, the climax lands the emotional or logical payoff, and the CTA tells the viewer exactly what to do next.
NotebookLM does heavy lifting at this stage: given source material — article links, existing videos, images — it can draft a complete video outline rather than just prose. Existing written content becomes a rough video concept without starting from a blank page.
Stage 3: Scene breakdown — the cheapest place to kill a bad video
The script is cut into individual scenes, each with its characters, setting and action defined. When there is already rough footage, the team selects the promising sections and keeps the original audio as a timing reference for rebuilding the scene later.
This stage produces the storyboard. It is also the last point where changing your mind is free.
Stage 4: Stills first, then motion — the rule that protects the budget
This is the most valuable habit in the whole pipeline, and it is worth stating plainly:
Write a detailed prompt per scene. Generate AI images first. Assemble the approved images into video. Do not prompt video directly.
Two reasons it works:
- Cost. Video generation burns credits at a far higher rate than image generation. Iterating on stills means you fail cheaply and only spend video credits on frames you have already approved.
- Control. A still can be judged, corrected and re-rolled in seconds. Once motion is baked in, fixing a detail means regenerating the whole clip.
The corollary is that prompts must be specific enough to succeed on the first generation. Vague prompts are not free — they are just a slower way to spend credits.
Stage 5: Audio and effects — where a good video becomes a finished one
The final stage synchronises AI-generated voice-over with the visuals, then layers background music, sound effects and ambience. Colour, lighting and contrast are graded, AI artefacts are hunted down, and subtitles are added and proofread.
Subtitles are not an afterthought. Most short-form video is watched muted.
Seven upgrades I would make to this pipeline
The process above is sound. These are the gaps I would close.
1. Add a Stage 0: a one-line brief with a success metric
The pipeline starts at "idea" but never defines done. Before anything else, write one line: audience, single message, platform and duration, and the metric that decides success — watch-through rate, clicks, replies. Without it, stage 5 has no acceptance test and "finished" becomes a matter of taste.
2. Lock characters and style before stage 4, not during it
The most common failure in AI video is identity drift: the same person changes face between scene 2 and scene 5. Build a small style bible first — locked character reference frames, a palette, a lens and lighting convention — and pass those reference images into every scene generation. Fixing drift after the fact means regenerating whole scenes.
3. Put a human approval gate between stage 3 and stage 4
Storyboard sign-off should be explicit and mandatory. Every hour spent generating footage for a scene that gets cut is money spent proving something the storyboard already knew.
4. Adopt a naming convention and keep a shot log
scene_04_v03 beats final_final_2. Log which prompt produced which accepted shot. Without this, a re-render three days later cannot be reproduced, and the team relights the same scene twice.
5. Build a prompt library
Prompts that produced accepted shots are a compounding asset — the most undervalued output of the whole pipeline. Save the winners with a note on what they solved. After twenty videos, this library is worth more than any single tool subscription.
6. Lock the voice-over before generating visuals
For short-form especially, record or generate the VO first and cut the shot list to the actual read. Generating visuals first and then discovering the narration runs four seconds long forces re-renders that stills-first was supposed to prevent.
7. Add a Stage 6: distribution
The pipeline currently ends at "final video", which is where the value has not yet been collected. One master should produce vertical cuts, a thumbnail, platform-specific captions, and a publishing schedule. For Vietnamese content, add a mandatory diacritics proofread — auto-captioning is reliably wrong about Vietnamese tone marks.
And one number to track above all: cost per finished minute, including rejected generations. It is the only figure that tells you whether the pipeline is actually improving or just getting busier.
FAQ
Why generate images before video instead of prompting video directly? Because video credits cost far more per second than image credits, and a rejected clip costs the same as an accepted one. Iterating on stills lets you fail cheaply, then spend video credits only on approved frames.
What is the biggest quality risk in an AI video pipeline? Identity and style drift between scenes. Lock character reference images and a visual style before generating any scene, and pass those references into every generation.
Where should the human approval step go? At the storyboard, between scene breakdown and generation. That is the last point where changing direction costs nothing.
Do these tools handle Vietnamese? Voice-over and subtitles work, but auto-generated Vietnamese captions get diacritics wrong often enough that a manual proofread must be a required step, not an optional one.
What single metric should a video team track? Cost per finished minute, counting rejected generations. It exposes waste that a per-video budget hides.
✍️ The Author: Do Ngoc Hoan Founder of CookConnects.ca & Wizy.ca. Bridging the gap between advanced algorithms and business execution. I write for technical founders looking to scale their impact with AI and robust engineering.