Making product videos
One idea to a finished cut on every platform. Demos, ads, Shorts, Reels.
Three layers
Every element in a finished video sits in one of three layers, and which layer it sits in decides how you make it.
Captured is real. Your screen recording, a phone clip, a voice track. Exact, free to copy, and expensive only when you have to shoot it again.
Composed is drawn at render time. Captions, titles, callouts, the zoom path, the end card, the grade. Exact, free, editable forever.
Generated is a sample from a model. Plausible rather than exact, dollars a clip, and — the part that catches people out — asking twice returns two different answers. There is no earlier version to go back to.
Almost every failure in an AI-made product video is an element sitting in the wrong layer. Your UI in the generated layer comes back with invented buttons. Your price in the generated layer comes back as $4,900, and fixing the typo means re-rolling the shot, which returns a different shot. Your hook in the composed layer can be swapped ten times before lunch; baked into footage, it can be swapped once.
So the question for any element isn’t whether a model can do it. It’s which layer it belongs in. Text, logos, prices, UI, and anything a lawyer would read: composed. Your product: captured. What’s left — an opening image, an environment, b-roll, a presenter — is the only thing worth generating, and it’s a wrapper around the real thing.
One more constraint shapes the schedule rather than the layers. A video model returns a few seconds and remembers nothing, so continuity between shots has to live outside it, in files you pass back in on every call.
Four ways a video fails
They’re independent. Each has its own signal, and only one of them is fixable in the edit.
| Fails as | What happened | Signal | Fixable in the edit |
|---|---|---|---|
| Nobody watches | the hook missed | drop inside two seconds | no — write more hooks |
| They watch, don’t follow | crop too wide, cuts too fast | drop through the middle | partly — recrop, slow down |
| They follow, don’t care | the claim isn’t one they want | watch-through, no clicks | no — that’s positioning |
| It looks wrong | drift, garbled text, an uncanny shot | your own eyes | yes |
The last row is the only one that responds to working harder on the video, which is why it absorbs the attention. Afternoons go into re-rolling a b-roll clip nobody notices, while the row that decides the outcome got written once, in five minutes, and never tested.
Decide what ships
Who’s watching, on what platform, and what they do next. Distribution sets the cut, so it’s settled first.
One claim per video. If you have two, make two videos.
Master vertical at 1080×1920 and treat everything else as a derived crop. Starting 16:9 and cropping down throws away the shot; starting 9:16 and adding bars for YouTube doesn’t.
Settle the CTA now. It changes the last three seconds and the end card, and retrofitting it means recomposing.
Capture, then index
Record long and messy, and narrate while you do it. The footage, the copy, and the look all come out of that one capture.
Record at twice the output resolution so a crop still lands sharp. Do each action twice and slowly, because cursor moves that look fine live read as jitter once slowed. Seed real-looking data first — empty states and test test kill more demos than bad editing. Kill notifications. And narrate even badly: whatever you say is the only grounded description of what the thing does, and the script comes out of it later.
Then turn the capture into a time-coded map before cutting anything:
- Transcript, time-coded.
- Frame delta — the difference between consecutive frames over time. Typing is small and local. Navigation is a medium abrupt jump. Loading is near zero. A payoff, where the result appears or the chart fills, is a large delta right after a near-zero stretch.
- OCR on sampled frames, so the index knows which screen is which.
The delta curve says when to cut. The region that changed says where to crop when you reframe 16:9 into 9:16. One signal, both jobs.
Cutting then falls out of the index mechanically. Keep the payoffs — most of a product recording is navigation and a small fraction is the moment something appears, and that fraction is the video. Drop near-zero delta with no cursor motion. Speed typing 4–8×. Punch into the changed region instead of showing the whole screen, because full-screen UI at 9:16 is unreadable on a phone.
Two things here read as amateur if you get them wrong. Speech and screen don’t align — people say “here’s the result” two seconds before it renders — so treat the transcript as claims and the delta map as events, then align each claim to the nearest event. And raw captures leak: a customer name, a live inbox, a key in a settings pane, a notification sliding in. Redaction is a gate before anything leaves the machine, and you’re the least likely person to catch it, because you look at that data every day.
Write the hooks
Written off the transcript, against one body. There’s no arc at twenty seconds — hook, tension, payoff, CTA — and the hook is the only part worth iterating.
How many is set by traffic, not ambition. A variant needs enough impressions for its curve to separate from the others, so ten variants across five hundred impressions tells you nothing you couldn’t have guessed. Below that threshold the loop is theatre: make one video and show it to five people who’ve never seen the product.
Ground every claim in something the transcript actually said, because a hook the product can’t pay off loses the back half of the curve. Write for mute. And vary the kind of hook rather than the wording — the problem stated, the result first, the number, the contradiction, the question. Ten rewordings of one hook is one test.
The words the models know
Models are trained on captioned footage and on scripts, so they respond to the vocabulary those captions use and guess at everything else. “Nice light from the side” is a guess. “Hard key at 45 degrees camera left, 4:1 key-to-fill” is an instruction.
Pick one word per axis and hold it across the whole project:
- Shot size — extreme wide, wide, medium wide, medium, medium close-up, close-up, extreme close-up, over-the-shoulder, insert
- Angle — eye level, low angle, high angle, overhead, Dutch
- Movement — locked-off, pan, tilt, dolly in, dolly out, truck, pedestal, crane, orbit, handheld, push in. Name one, and name the speed
- Lens — focal length in mm and a T-stop. “35mm, T2.8” beats “cinematic”
- Light — key, fill, rim, practical; hard or soft; direction in degrees or clock position; key-to-fill ratio; color temperature in kelvin. Named patterns land well: Rembrandt, butterfly, split, high-key, low-key
- Motion blur — shutter angle. 180° is the film default; 90° gives the crisp staccato look
- Grade — the palette in named colors, plus the stock or camera whose response curve you want
Three words are overloaded enough to be worth pinning down:
- Keyframe means a defined pose in animation and an I-frame in a codec. Here it’s the conditioning frame, the still you pass as
first_frame. - Reference means a mood board to a person and
inputReferencesto the gateway. Say which one. - Style frame is the single rendered still that fixes the look. The look block is the text you paste into every prompt. Different artifacts, easy to conflate.
Lock the look
Pull the palette and type off your own sampled frames. It should look like the product, not like a stock template.
Then write the look once, as a block you paste into every generated shot and never reword, so the only thing varying between two shots is the action clause. Keep it under 90 words — long style preambles crowd out the fifteen words of action that matter.
Look block template
[ACTION — 15-30 words, one continuous motion]
Clean product-film look. 50mm, T2.8, eye level, locked-off unless
stated. Soft large key 45 degrees camera left, 3:1 key-to-fill, no
hard shadow. Palette: [two brand colors], neutral ground, no
competing color. Fine grain, no lens flare, no bloom.
9:16, 1080x1920, 24fps, 180° shutter.
No text, no captions, no watermark, no UI, no zoom.
Render it as a style frame on the hardest shot, then freeze it. Changing it later means re-rendering everything before it.
Turnarounds, when a subject recurs
Build these only when a character, prop, or location appears in more than one shot. One-off b-roll doesn’t need them.
The artifact is an orthographic character turnaround sheet, and it has to be a single image. Reference conditioning takes an image, not a folder — and a sheet drawn in one pass reconciles its own angles, where five separately generated images drift apart from each other.
Specify it in those terms:
- Orthographic projection, no perspective, no foreshortening. Perspective baked into a reference distorts every shot that uses it
- Uniform scale and a shared eye line across all views, aligned on one baseline, so the model reads them as one subject
- Views: front, three-quarter left, left profile, rear three-quarter left, rear. Mirror to the right only if the subject is asymmetric
- A-pose, neutral expression
- Flat even lighting, no cast shadows, plain mid-grey background. The turnaround carries identity; the look block carries mood. A sheet lit hard from the left fights every shot you light from the right
- Generate it from one hero image, so drift has a single origin
An object turnaround is the same sheet plus a scale reference and an in-hand view, since objects fail on material and scale rather than shape. A location needs an establishing wide and one variant per lighting state.
Review it the way a casting director reads headshots — whole sheet at once, is this one subject. If the rear view has shorter hair than the front, every rear shot fights you and you’ll blame the shot.
Turnaround prompt
Orthographic character turnaround sheet. Single image, five views in
one horizontal row: front, three-quarter left, left profile, rear
three-quarter left, rear. Orthographic projection, no perspective, no
foreshortening. Identical scale and eye line across all views, aligned
on a shared baseline. A-pose, neutral expression. Flat even studio
lighting, no cast shadows, plain mid-grey background. Full figure head
to feet in every view. No text, no labels, no watermark.
Subject: [hero image as reference, or description]
Generating the wrapper
One continuous motion per shot, sayable in one clause. If you need a comma, it’s two shots.
Approve it as a still first — render the conditioning frame, look at it, then animate. Images cost cents and clips cost dollars. Keep the camera locked-off or on a simple linear move for anything you’ll composite onto, because tracking generated footage is worse than tracking real footage.
Coverage is a budget line, and two things set it: your hit rate on that kind of shot, and how much of the video rests on it. Roughly one take in four is usable, so four to six for an ordinary shot and ten for the one carrying the video — but measure your own rate after the first project and use that instead. If eight takes fail, the shot is wrong rather than the sampling. Re-block it, or cut it in two.
Fix the seed while tuning the prompt and vary it for coverage, so you’re never changing two things at once. Judge in motion, at speed and at quarter speed. Stills lie: gravity gets lazy, hands grow a finger, the walk slides.
Compose and ship
Everything above the footage — captions, titles, callouts, the zoom path, the end card. Remocn ships this layer as 297 components, MIT, code in your repo, driven by an agent with live preview. Remotion underneath.
Captions burned in, always; muted autoplay is the default viewing condition. Keep anything critical inside the middle 80% vertically, since platform UI covers the bottom fifth and the right edge. The zoom path comes from the index rather than by hand. Sound does more than another twenty takes — room tone, three foley hits, a bed. And one grade over everything, so generated shots and screen capture read as one video.
Then one body, ten hooks, rendered programmatically. That’s the thing composing in code buys that an NLE can’t. Derive 1:1 and 16:9 by re-cropping through the same index rather than re-editing, and re-check safe areas per crop, because an end card that clears the UI at 9:16 can collide at 1:1. Check duration and file limits at upload time; they move.
Read retention, not views. The first-two-second drop says which hook won, and nothing else in the analytics matters until that’s settled.
What gets in the way
The method above is mechanical. The reasons it doesn’t get followed aren’t.
- You can’t see your own product any more. You skip the step that confuses everyone because you stopped noticing it years ago. The narration in your capture is the best defence — a stranger reading the transcript will find the gap you can’t.
- Generating is the fun part. Re-rolling a shot feels like craft. Writing the eleventh hook feels like admin. The second one decides the outcome.
- Polish reads as progress. A grade pass takes an afternoon and moves no number. Shipping one more variant takes ten minutes and moves the only number.
- Nobody wants to hear the claim is wrong. Row three of the failure table has no edit that fixes it, so it gets rediagnosed as row one, and the team makes five more hooks for a video nobody wanted.
- The redaction gate gets skipped by the only person who could have caught it. Real customer data looks normal to you.
And the honest limit: all of this assumes you can measure retention. If you’re shipping to a few hundred people, the variant loop tells you nothing, and you’re better off making one careful video and watching five strangers try to explain it back to you.
Models
August 2026, on the Vercel AI Gateway. These turn over every few months; the slots don’t.
- Direct and script —
anthropic/claude-opus-5. Bulk prompt expansion ongoogle/gemini-3-flash. - Watch the recording —
google/gemini-2.5-protakes video input directly. - Conditioning frames and turnarounds —
google/gemini-3.1-flash-imagefor volume,bfl/flux-2-profor heroes. - Prototype motion —
minimax/minimax-h3-max(480p and 768p; the naming is inverted, Max is the small fast one),bytedance/seedance-2.0-fast, orspacexai/grok-imagine-videoat 480p. - Final —
bytedance/seedance-2.5oralibaba/wan-v3.0-video. Both do image-to-video, first-and-last frame, references, and audio.google/veo-3.1-generate-001runs around 23× the cost per clip of the cheap tier, so keep it for hero shots and useveo-3.1-fast-generate-001at roughly a fifth of that elsewhere.
Prototype and finish inside one family, or the prototype won’t predict the final.
What the models can’t do
- hands in fast motion: crop them out or shorten the shot
- text: never generate it
- counts: five apples come back as four
- two people touching: cut between them
- hitting a music beat: generate long, let the composition time it
- crowds: keep them soft or out of focus
- screens and UI: capture, never generate
- physics past about two seconds
- dialogue past a sentence: separate VO, cut to reactions
Shorten the shot, cut around it, or move it up a layer.