Every ad in here is generated, not filmed. No shoot, no crew, no talent, no licensing. That changes which formats are cheap and which are expensive, and not in the order you would expect.
These are the twelve we build for direct-to-consumer brands on Meta. For each one: what it is made of, where it wins, and the specific place it breaks. Most brands run two or three of these, then wonder why scaling stopped.
Take the whole thing as a file and run it through your own AI
The entire catalog as plain text, structured so Claude or ChatGPT can use it. Paste it in and it will pick formats for your brand, audit the last twenty ads you ran, or write a full batch plan. No signup, no email, nothing gated.
format-catalog.md · 24 KB · four ready-made prompts inside
Motion analysed $1.3 billion in ad spend across 550,000+ ads from 6,000+ advertisers on Facebook and Instagram, September 2025 through early January 2026. Four findings from it are the whole argument for this document.
~5%
of ads launched become winners. A winner spends at least 10x the account's median ad.
~55%
of all spend flows to that 5%. Roughly 17% goes to outright losers.
2–3x
more creative launched by the top quartile in every budget tier, versus same-budget peers.
2x
the winners, on an identical budget, for brands that launch more creative.
Read the last two together and it gets uncomfortable. More money does not buy more winners. More swings does. And a swing only counts if it is genuinely different, because twenty versions of one talking head is one swing with twenty price tags on it. Generating the ads instead of filming them is what makes twelve formats a realistic month rather than a wish. There is no shoot to schedule, so the constraint stops being budget and starts being whether anyone decided what to make.
02
What an AI ad is actually made of
Three ingredients. The mix is what makes one format cheap and another expensive.
Images
The frames come first
Every clip starts from a generated still. Some formats need a matched pair, a start state and an end state, which doubles the image count and makes that format far less forgiving.
Clips
Each one is a dice roll
Clips are the video generations. This number is the best single predictor of how long a format takes and what it costs, better than run time is.
Voice
Two lanes, and they are not equal
Either generated inside the video model along with the picture, or recorded separately and laid over silent footage. This choice decides more than people expect.
The voice split is the most useful line in this document
Voice generated in-model
You get a character who genuinely speaks on camera.
Lip sync degrades past roughly ten seconds in one clip.
The same character drifts into sounding like a different person across a sequence.
Fix: lock the voice description word for word, reuse it on every clip.
Voice recorded separately
Neither problem exists. Not reduced, gone.
Casts of five cost nothing extra in drift.
Fastest turnaround of anything in the catalog.
Cost: you cannot have a talking face.
Half the failure notes in this catalog are downstream of that one choice. If a format keeps breaking on you, check which lane it is in before you rewrite the script.
03
The whole catalog, on one screen
Images and clips are a typical build, what a normal version of each format actually takes. A typical build is more useful than a range when you are planning a month.
Format
Run band
Images
Clips
Voice
Characters
Look
Talking Head UGC
15–35s
7
7
In-model
1
Phone-shot realism
Longform Testimonial
60–120s
18
18
In-model
1
Phone-shot realism
Podcast Two-Hander
30–90s
12
12
In-model
2
Studio conversation
Judged Panel
60–90s
10
24
In-model
3+
Produced television
Blind Sense Test
30–60s
9
9
Separate
1
Phone-shot realism
Clinical Explainer
20–50s
16
8
Separate
0
Evidence
Object Talk
8–40s
2
3
In-model
0
Stylized animation
Claymation
30–60s
14
9
Separate
0–3
Stylized animation
Founder Saga
60–120s
18
18
Separate
3–5
Narrated story
VO-Driven B-Roll
30–60s
10
10
Separate
0
Tactile lifestyle
Product Hero
24–48s
12
7
Separate
0
Commercial polish
Native Feed Ad
8–44s
6
6
Either
0–1
Anti-polish organic
Two rows are worth staring at. Judged Panel builds twenty-four clips off ten images, because the character designs get reused. That is the only reason a five-character produced-TV ad is affordable. Clinical Explainer does the opposite, sixteen images for eight clips, because every clip needs a matched pair. It is the least forgiving format here.
04
The twelve
Talking Head UGC
01
A generated person talks to camera like they are recording on their phone. The workhorse, and the format most batches are built on.
Typical build
Run band
15–35s
Characters
1
Frames
Single
Where it wins
Cold traffic, where a peer voice beats an authority voice
Problem and solution stories that need a face to carry them
Volume. The cheapest format to run many versions of
Where it breaks
The voice drifts. Past roughly five clips the same character starts sounding like someone else, unless the voice description is locked word for word and reused on every clip
Lip sync degrades past about ten seconds in one clip. Write to that rather than fixing it after
One room for forty seconds is boring. Move the character or cut away
Non-obvious moveRebuild your best performing script with a different age of character in a different room. Same words. A separate swing, not a variation, and it costs one build.
Longform Testimonial
02
A generated person tells a transformation story, with a cutaway on almost every line so each claim gets a picture.
Typical build
Run band
60–120s
Characters
1
Frames
Single
Where it wins
High intent audiences who will genuinely watch two minutes
Before and after proof, and products that need explaining time
A time-anchored hook that maps to the viewer's own calendar
Where it breaks
Clip count. Eighteen generations makes this a hero piece, not a batch filler
Keeping one generated person recognisably the same across all those cutaways is the hard part, and it is invisible until it is wrong
Cutaways planned at script time work. Cutaways added later never match the character
Non-obvious moveWrite the script, then mark every sentence that makes a claim. Each mark is one cutaway you owe. Four marks means this is the wrong format.
Podcast Two-Hander
03
Two generated characters in conversation, one pitching the other. Each is generated separately and cut together. They never share a frame.
Typical build
Run band
30–90s
Characters
2
Frames
Single
Where it wins
Reveal hooks. "Wait, what is that?" is an entire opening
Authority and status positioning, where being overheard beats being told
Industry claims that land better as gossip than as a pitch
Where it breaks
Twice the everything. Two characters is double the generations and double the work keeping each face consistent
Wrong pairing kills the read instantly. Cinematic music under a raw conversation is fake in one second
The escape hatchIf the second character never has to be seen, do not generate them. Put them off camera with their own voice. One body, two people in the scene, roughly half the build. State in the prompt that the visible character is not speaking, or the model syncs their mouth to the wrong voice.
Judged Panel
04
An authority rules on competing products and yours wins. Every character is generated alone; the panel exists only through eyelines, seating geometry and the edit.
Typical build
Run band
60–90s
Characters
3+
Frames
Reused
Where it wins
A verdict story. Products compete, an expert rules, yours wins on a stated reason
Brands that want a produced television feel instead of phone footage
Three or more characters who never share a frame, which is exactly what AI is good at and a real shoot is not
Where it breaks
The run time explodes and nobody notices. Beats times rivals times clip length is your run time. Three rivals with a full arc quietly becomes a four minute ad
The fix is to cut rivals, not to trim clips at random
Ten images carrying twenty-four clips only works if the character designs are locked first. Skip that and the reuse collapses
Non-obvious moveDo the arithmetic before writing. Beats x rivals x seconds. Over ninety, drop a rival. Two rivals plus your hero lands in a normal band every time.
Blind Sense Test
05
A generated person blind-tests products against name brands and the hidden one wins. The idea does the work; the build is simple.
Typical build
Run band
30–60s
Characters
1
Frames
Single
Where it wins
Anything testable by smell, taste, sound, feel or texture
The curiosity gap is free. "Which one wins" holds attention with no effort
Reactions are non-verbal, so it sidesteps lip sync entirely while still having a person on camera. Very few formats can say that
Where it breaks
A weak idea takes the whole ad with it. There is no craft layer to hide behind
A reaction that looks performed kills it, and models overshoot expressions by default. Ask for restraint explicitly
Competitor products in frame is generally fine. Disparaging them in the voiceover is a legal question, not a creative one
Non-obvious moveName the three biggest brands in your category out loud. If you would genuinely bet on your product in a blind test against them, this format is nearly free.
Clinical Explainer
06
An isolated subject on a plain background morphing from one state to the next, with a calm narrator over the top. Evidence, not advertising.
Typical build
Run band
20–50s
Characters
None
Frames
Paired
Where it wins
Teaching a mechanism nobody can see, which is most supplements and most skincare
It is modular. A few beats drop into any other ad as the "how it works" section
Internal biology that photoreal models often refuse to generate will usually pass in this stylized register. It is the practical route to showing what happens inside the body
Where it breaks
It has no warmth and it is not supposed to. Use it to explain, never to make someone feel something
A flat narrator sinks it. The voice is doing more work than the picture
Two images for every clip. The least forgiving format here, because a sloppy start and end pair does not make a weak clip, it makes garbage
Non-obvious moveBuild six beats once as a standalone ad, then reuse the same six as the middle of three other ads. The only genuinely modular format here.
Object Talk
07
An object with a face talks to camera in first person. An ingredient, a body process, or the cheap thing the customer tried before you.
Typical build
Run band
8–40s
Characters
None human
Frames
Reused
Where it wins
The strongest scroll stop in the catalog. Nothing else in the feed looks like it
It turns a lecture into a character monologue, so people sit through the science
The cheapest build here by a wide margin. One character design carries the whole ad
Where it breaks
A cartoon register does not fit every brand. If you are locked to photoreal, skip it
Getting a genuinely expressive face onto an abstract object takes four to six image tries. Budget for it
Animation timed to specific words is unreliable across a long clip. Keep beats short
The angle that convertsMake the object the customer's old failed solution, confessing. "I'm your drugstore moisturizer. I'm full of cheap oils that keep my price down and clog your pores." The old product indicts itself and yours becomes the upgrade, with no hard sell and no competitor named.
Claymation
08
A clay, Pixar or otherwise made-not-filmed register. Products become characters and worlds get built.
Typical build
Run band
30–60s
Characters
0–3
Frames
Mixed
Where it wins
Visual relief in a batch that is otherwise wall to wall talking heads
It carries claims that sound forced from a human mouth. Small workers building your skin layer by layer is fine in clay and absurd in live action
No lip sync, no face consistency, no uncanny valley. Three of the hardest things in AI video simply do not apply
Where it breaks
Stylized and photoreal cannot be mixed. A different rule set, not a filter you apply
Not every brand survives being cute. If the category is serious, this can read as unserious
The finish decides it. A rough clay ad reads as a cheap cartoon; a scored and graded one reads as a television spot
Non-obvious moveTake the one claim your compliance reviewer keeps softening. Animation is usually where an over-literal claim becomes an obviously figurative one, which is a different conversation.
Founder Saga
09
A multi-character origin story told over narration and fast cuts. We went there, we found this, we brought it back.
Typical build
Run band
60–120s
Characters
3–5
Frames
Single
Where it wins
Heritage and sourced brands where the journey genuinely is the pitch
Narration means nothing lip syncs, so a five character cast costs nothing in voice drift. This is the format that makes a large cast affordable
Long form absorbs a volume of footage a thirty second spot cannot hold
Where it breaks
Highest build of anything here. A hero creative, not something to run five of
Keeping three to five generated characters recognisably themselves across eighteen clips is the hardest continuity job in the catalog
If there is no real story, this format exposes that faster than any other
Non-obvious moveWrite the story as six sentences first. If sentence four is not surprising to someone outside your company, you have an About page, not a saga.
VO-Driven B-Roll
10
No talking head at all. Unboxing, texture, hands, product in use, with a voice over the top.
Typical build
Run band
30–60s
Characters
None
Frames
Single
Where it wins
Products where opening the box is the experience. Bedding, beauty, candles, drinks
Buyers who care how something feels more than what an expert says about it
No lip sync, no voice drift, no face to keep consistent. Fastest turnaround in the catalog
Where it breaks
Nobody vouches for the product on camera. No face to trust means no personal credibility
It cannot hold an authority claim or a heavy mechanism. Pair it with something that can
Hands are still the least reliable thing to generate. Frame tighter than you think and keep the motion simple
Non-obvious moveRecord the voiceover before generating anything. The script decides the clip count. The other order is how brands end up with pretty footage and nothing to say over it.
Product Hero
11
No people. Product shots, pours, opens and compositions, with a narrator laid over.
Typical build
Run band
24–48s
Characters
None
Frames
Paired
Where it wins
Hook economics. One set of visuals plus four voiceovers and four text hooks is four ads at close to the cost of one. Nothing else here is that cheap to multiply
Batches where building and keeping a character consistent is not worth the lift
Categories sold on how the thing looks and feels in the hand
Where it breaks
A narrator sounds like a third party, so hard closes land cold. First person does not work here
No person means no emotional journey. Problem and solution stories fall flat
Your packaging has to survive being generated. Legible labels and exact logos are the most common failure in this format. Composite the real pack where it matters
Non-obvious moveThe cheapest way to test four hooks properly. Build visuals once, write four completely different opening lines, ship four ads. Whatever wins tells you what to say everywhere else.
Native Feed Ad
12
Built to not look like an ad. Lower polish on purpose, sitting flush with organic content.
Typical build
Run band
8–44s
Characters
0–1
Frames
Single
Where it wins
The cheapest and fastest lane. The split-screen version is two images and no clips at all
Faceless versions avoid the two hardest problems in AI video, faces and lip sync, entirely
Tired categories where anything polished instantly reads as an ad and gets scrolled
Where it breaks
Deliberately rough is one inch from actually bad. Models default to over-polishing, so roughness has to be asked for explicitly
Trend audio does not work here. A licensed track and a real micro-expression cannot be reproduced
Non-obvious moveScreenshot the first five organic posts in your own feed. Match that polish level exactly, not the level of the ads around them. That is the whole brief.
05
Run the whole thing through your own AI
Four prompts that do the actual work once the file is in the conversation.
Paste the file, then paste one of these
Replace anything in brackets.
Pick formats for your brand
My brand is [BRAND], we sell [PRODUCT] to [WHO]. We spend about [SPEND] a month on Meta. Pick the five formats that fit us best. For each, tell me why it fits, what our version would show on screen, and the specific way it would fail for a product like ours.
Audit what you already run
Here are the last twenty ads we launched: [ONE LINE EACH]. Tag each with a format from the catalog. Tell me how many distinct formats we run, which of the twelve we have never touched, and which untouched one you would bet on first.
Write a batch plan
Build me a twenty-ad batch for [BRAND] using the allocation model in the file. For each ad give me the format, the awareness stage, a one-line concept and the hook line. Do not repeat a concept shape twice.
Brief one ad properly
Using the [FORMAT] entry, write a full brief for [BRAND] selling [PRODUCT]. Give me the image list, the clip list with a duration for each, what is said over each clip, and flag anything on that format's "where it breaks" list my brief is at risk of.
Operator reads from the accounts we build for, not a published study. We are calling them directional on purpose, because a format ranking drawn from a handful of accounts is a starting bet, not a benchmark. Point your first three swings here, then let your own data overrule us.
01
Supplements
Leads
Clinical explainer and object talk. The mechanism is invisible, so a person describing it is the weakest tool available. Stylized registers also carry claims that sound forced from a human mouth, and they sidestep the realism problem entirely.
Lags
Pure product hero. Nobody buys a capsule because the bottle looked good.
The trap
Every competitor is running an ingredient lecture. The differentiator is the format the lecture arrives in, not the ingredients.
02
Skincare and beauty
Leads
Education-led structures. Judged panel and clinical explainer both carry them. The strongest account in our own book runs an education-first spine, not a testimonial-first one.
Lags
Generic before-and-after talking heads. Saturated, and they attract the most claim scrutiny of anything here.
The trap
The actives work under the surface, so a generated camera can only ever show a face. Diagram or animate the layer that cannot be filmed at all. That is the one real advantage of building this way.
03
GLP-1, telehealth and regulated
Leads
Native feed and clinical explainer. Faceless and diagrammatic avoids the body-morph problem entirely, and both are cheap enough to run at the volume a regulated account needs.
Lags
Dramatic transformation footage. Most likely to be rejected, and body transformation is one of the least reliable things to generate convincingly.
The trap
Use transformation proxies instead of transformations. A simultaneous split screen, a loose waistband, a visualisation of the invisible. Same story, nothing morphs.
04
Home, comfort and consumable CPG
Leads
On-camera demonstration and two-person dialogue. Both consistently outrun a plain talking head. The mechanism is physical, so show it moving in the first five seconds.
Lags
Plain talking head with no demo. It drops off unless an unusually strong hook is doing all the work alone.
The trap
If your product can be judged by feel, taste or smell, the blind sense test is nearly free and almost nobody in the space runs one.
One pattern held across all four and it had nothing to do with format. When we audited a competing script set against ours in one account, their thirty-one scripts mentioned the brand's headline guarantee zero times. Ours carried the same four elements every time: the guarantee, the mechanism, a reason to move now, and an actual ask. Before changing format, check all four are in the script you already have. That is a free fix and it outranks everything else on this page.
07
How to spread twenty ads
Motion's finding was that more creative wins, not more of the same creative. Here is how we allocate a twenty-ad month so the swings are actually different. Adjust the counts, keep the shape.
8
The engine
Talking head and native feed. High volume, low build, fast to iterate. This is where you find hooks, not where you find formats.
40%
5
The teachers
Clinical explainer, object talk, product hero. No characters to keep consistent, and they carry the mechanism. If your product needs explaining, this block is not optional.
25%
4
The pattern breaks
Claymation, blind sense test, podcast. These are what stop a batch looking like one batch. Cut this block and the account fatigues faster.
20%
2
The heroes
Longform testimonial, founder saga, judged panel. Expensive, slow, and worth it once or twice a month. Never more.
10%
1
The one you have never run
Whatever is left on the list. It is one ad. The cost of being wrong is one ad, and the cost of never trying is the format you would have found.
5%
The second axis, and the one most brands miss. Cut the same twenty by how much the viewer already knows: 30 to 40% for people who do not yet know they have the problem, 40 to 50% for people comparing solutions, 10 to 20% for people ready to buy. Almost every account we open runs that last group at close to 100%, which is exactly why it stops scaling. You cannot buy new customers with ads written for existing ones.
08
Picking one, in three questions
Question one
Who has to be believed?
A peer, an authority, or nobody. Peer sends you to talking head or testimonial. Authority sends you to judged panel or clinical explainer. Nobody sends you to product hero, b-roll or object talk.
Question two
Does anyone need to speak on camera?
The AI-specific question, and it decides more than people expect. Yes puts you in the in-model voice lane with its lip sync limits and voice drift. No makes every one of those problems disappear and your build gets faster immediately.
Question three
What is already in the batch?
Run the character counts and the voice column against what you shipped last month. If every ad is one generated person talking for thirty seconds, the gap is structural and no amount of copywriting closes it.
If you want these built instead of just listed
We are an AI creative studio for direct-to-consumer brands. We build every format on this page as performance ads, at volume, for brands spending real money on Meta.
18%
lower CPA for a DTC supplement brand at $40K a month in spend
3.1x
ROAS on winning creatives
8–12
winning creatives produced per month
12+
brand partners
Prove It is $997. Three performance videos and five statics, and it credits toward any of our packages.
And the honest version: if you already have a studio you like, the file above works fine without us. The picking is the part that pays.