Whim Creative  /  Creative Reference No. 05 Free. Nothing gated.
The AI Ad Format Catalog

Twelve AI ad formats
that are working right now

Every ad in here is generated, not filmed. No shoot, no crew, no talent, no licensing. That changes which formats are cheap and which are expensive, and not in the order you would expect.

These are the twelve we build for direct-to-consumer brands on Meta. For each one: what it is made of, where it wins, and the specific place it breaks. Most brands run two or three of these, then wonder why scaling stopped.

Take the whole thing as a file and run it through your own AI

The entire catalog as plain text, structured so Claude or ChatGPT can use it. Paste it in and it will pick formats for your brand, audit the last twenty ads you ran, or write a full batch plan. No signup, no email, nothing gated.

format-catalog.md  ·  24 KB  ·  four ready-made prompts inside
Download the file
01

Why this beats raising your budget

Motion analysed $1.3 billion in ad spend across 550,000+ ads from 6,000+ advertisers on Facebook and Instagram, September 2025 through early January 2026. Four findings from it are the whole argument for this document.

~5%
of ads launched become winners. A winner spends at least 10x the account's median ad.
~55%
of all spend flows to that 5%. Roughly 17% goes to outright losers.
2–3x
more creative launched by the top quartile in every budget tier, versus same-budget peers.
2x
the winners, on an identical budget, for brands that launch more creative.

Source: Motion, Creative Benchmarks 2026. motionapp.com/thumbstop-pulse/creative-benchmarks-2026

Read the last two together and it gets uncomfortable. More money does not buy more winners. More swings does. And a swing only counts if it is genuinely different, because twenty versions of one talking head is one swing with twenty price tags on it. Generating the ads instead of filming them is what makes twelve formats a realistic month rather than a wish. There is no shoot to schedule, so the constraint stops being budget and starts being whether anyone decided what to make.

02

What an AI ad is actually made of

Three ingredients. The mix is what makes one format cheap and another expensive.

Images

The frames come first

Every clip starts from a generated still. Some formats need a matched pair, a start state and an end state, which doubles the image count and makes that format far less forgiving.

Clips

Each one is a dice roll

Clips are the video generations. This number is the best single predictor of how long a format takes and what it costs, better than run time is.

Voice

Two lanes, and they are not equal

Either generated inside the video model along with the picture, or recorded separately and laid over silent footage. This choice decides more than people expect.

The voice split is the most useful line in this document

Voice generated in-model
  • You get a character who genuinely speaks on camera.
  • Lip sync degrades past roughly ten seconds in one clip.
  • The same character drifts into sounding like a different person across a sequence.
  • Fix: lock the voice description word for word, reuse it on every clip.
Voice recorded separately
  • Neither problem exists. Not reduced, gone.
  • Casts of five cost nothing extra in drift.
  • Fastest turnaround of anything in the catalog.
  • Cost: you cannot have a talking face.

Half the failure notes in this catalog are downstream of that one choice. If a format keeps breaking on you, check which lane it is in before you rewrite the script.

03

The whole catalog, on one screen

Images and clips are a typical build, what a normal version of each format actually takes. A typical build is more useful than a range when you are planning a month.

FormatRun bandImagesClipsVoiceCharactersLook
Talking Head UGC15–35s77In-model1Phone-shot realism
Longform Testimonial60–120s1818In-model1Phone-shot realism
Podcast Two-Hander30–90s1212In-model2Studio conversation
Judged Panel60–90s1024In-model3+Produced television
Blind Sense Test30–60s99Separate1Phone-shot realism
Clinical Explainer20–50s168Separate0Evidence
Object Talk8–40s23In-model0Stylized animation
Claymation30–60s149Separate0–3Stylized animation
Founder Saga60–120s1818Separate3–5Narrated story
VO-Driven B-Roll30–60s1010Separate0Tactile lifestyle
Product Hero24–48s127Separate0Commercial polish
Native Feed Ad8–44s66Either0–1Anti-polish organic

Two rows are worth staring at. Judged Panel builds twenty-four clips off ten images, because the character designs get reused. That is the only reason a five-character produced-TV ad is affordable. Clinical Explainer does the opposite, sixteen images for eight clips, because every clip needs a matched pair. It is the least forgiving format here.

04

The twelve

Talking Head UGC

01
A frame in the Talking Head UGC format

A generated person talks to camera like they are recording on their phone. The workhorse, and the format most batches are built on.

Typical build
Run band
15–35s
Characters
1
Frames
Single

Where it wins

  • Cold traffic, where a peer voice beats an authority voice
  • Problem and solution stories that need a face to carry them
  • Volume. The cheapest format to run many versions of

Where it breaks

  • The voice drifts. Past roughly five clips the same character starts sounding like someone else, unless the voice description is locked word for word and reused on every clip
  • Lip sync degrades past about ten seconds in one clip. Write to that rather than fixing it after
  • One room for forty seconds is boring. Move the character or cut away
Non-obvious moveRebuild your best performing script with a different age of character in a different room. Same words. A separate swing, not a variation, and it costs one build.

Longform Testimonial

02
A frame in the Longform Testimonial format

A generated person tells a transformation story, with a cutaway on almost every line so each claim gets a picture.

Typical build
Run band
60–120s
Characters
1
Frames
Single

Where it wins

  • High intent audiences who will genuinely watch two minutes
  • Before and after proof, and products that need explaining time
  • A time-anchored hook that maps to the viewer's own calendar

Where it breaks

  • Clip count. Eighteen generations makes this a hero piece, not a batch filler
  • Keeping one generated person recognisably the same across all those cutaways is the hard part, and it is invisible until it is wrong
  • Cutaways planned at script time work. Cutaways added later never match the character
Non-obvious moveWrite the script, then mark every sentence that makes a claim. Each mark is one cutaway you owe. Four marks means this is the wrong format.

Podcast Two-Hander

03
A frame in the Podcast Two-Hander format

Two generated characters in conversation, one pitching the other. Each is generated separately and cut together. They never share a frame.

Typical build
Run band
30–90s
Characters
2
Frames
Single

Where it wins

  • Reveal hooks. "Wait, what is that?" is an entire opening
  • Authority and status positioning, where being overheard beats being told
  • Industry claims that land better as gossip than as a pitch

Where it breaks

  • Twice the everything. Two characters is double the generations and double the work keeping each face consistent
  • Wrong pairing kills the read instantly. Cinematic music under a raw conversation is fake in one second
The escape hatchIf the second character never has to be seen, do not generate them. Put them off camera with their own voice. One body, two people in the scene, roughly half the build. State in the prompt that the visible character is not speaking, or the model syncs their mouth to the wrong voice.

Judged Panel

04
A frame in the Judged Panel format

An authority rules on competing products and yours wins. Every character is generated alone; the panel exists only through eyelines, seating geometry and the edit.

Typical build
Run band
60–90s
Characters
3+
Frames
Reused

Where it wins

  • A verdict story. Products compete, an expert rules, yours wins on a stated reason
  • Brands that want a produced television feel instead of phone footage
  • Three or more characters who never share a frame, which is exactly what AI is good at and a real shoot is not

Where it breaks

  • The run time explodes and nobody notices. Beats times rivals times clip length is your run time. Three rivals with a full arc quietly becomes a four minute ad
  • The fix is to cut rivals, not to trim clips at random
  • Ten images carrying twenty-four clips only works if the character designs are locked first. Skip that and the reuse collapses
Non-obvious moveDo the arithmetic before writing. Beats x rivals x seconds. Over ninety, drop a rival. Two rivals plus your hero lands in a normal band every time.

Blind Sense Test

05
A frame in the Blind Sense Test format

A generated person blind-tests products against name brands and the hidden one wins. The idea does the work; the build is simple.

Typical build
Run band
30–60s
Characters
1
Frames
Single

Where it wins

  • Anything testable by smell, taste, sound, feel or texture
  • The curiosity gap is free. "Which one wins" holds attention with no effort
  • Reactions are non-verbal, so it sidesteps lip sync entirely while still having a person on camera. Very few formats can say that

Where it breaks

  • A weak idea takes the whole ad with it. There is no craft layer to hide behind
  • A reaction that looks performed kills it, and models overshoot expressions by default. Ask for restraint explicitly
  • Competitor products in frame is generally fine. Disparaging them in the voiceover is a legal question, not a creative one
Non-obvious moveName the three biggest brands in your category out loud. If you would genuinely bet on your product in a blind test against them, this format is nearly free.

Clinical Explainer

06
A frame in the Clinical Explainer format

An isolated subject on a plain background morphing from one state to the next, with a calm narrator over the top. Evidence, not advertising.

Typical build
Run band
20–50s
Characters
None
Frames
Paired

Where it wins

  • Teaching a mechanism nobody can see, which is most supplements and most skincare
  • It is modular. A few beats drop into any other ad as the "how it works" section
  • Internal biology that photoreal models often refuse to generate will usually pass in this stylized register. It is the practical route to showing what happens inside the body

Where it breaks

  • It has no warmth and it is not supposed to. Use it to explain, never to make someone feel something
  • A flat narrator sinks it. The voice is doing more work than the picture
  • Two images for every clip. The least forgiving format here, because a sloppy start and end pair does not make a weak clip, it makes garbage
Non-obvious moveBuild six beats once as a standalone ad, then reuse the same six as the middle of three other ads. The only genuinely modular format here.

Object Talk

07
A frame in the Object Talk format

An object with a face talks to camera in first person. An ingredient, a body process, or the cheap thing the customer tried before you.

Typical build
Run band
8–40s
Characters
None human
Frames
Reused

Where it wins

  • The strongest scroll stop in the catalog. Nothing else in the feed looks like it
  • It turns a lecture into a character monologue, so people sit through the science
  • The cheapest build here by a wide margin. One character design carries the whole ad

Where it breaks

  • A cartoon register does not fit every brand. If you are locked to photoreal, skip it
  • Getting a genuinely expressive face onto an abstract object takes four to six image tries. Budget for it
  • Animation timed to specific words is unreliable across a long clip. Keep beats short
The angle that convertsMake the object the customer's old failed solution, confessing. "I'm your drugstore moisturizer. I'm full of cheap oils that keep my price down and clog your pores." The old product indicts itself and yours becomes the upgrade, with no hard sell and no competitor named.

Claymation

08
A frame in the Claymation format

A clay, Pixar or otherwise made-not-filmed register. Products become characters and worlds get built.

Typical build
Run band
30–60s
Characters
0–3
Frames
Mixed

Where it wins

  • Visual relief in a batch that is otherwise wall to wall talking heads
  • It carries claims that sound forced from a human mouth. Small workers building your skin layer by layer is fine in clay and absurd in live action
  • No lip sync, no face consistency, no uncanny valley. Three of the hardest things in AI video simply do not apply

Where it breaks

  • Stylized and photoreal cannot be mixed. A different rule set, not a filter you apply
  • Not every brand survives being cute. If the category is serious, this can read as unserious
  • The finish decides it. A rough clay ad reads as a cheap cartoon; a scored and graded one reads as a television spot
Non-obvious moveTake the one claim your compliance reviewer keeps softening. Animation is usually where an over-literal claim becomes an obviously figurative one, which is a different conversation.

Founder Saga

09
A frame in the Founder Saga format

A multi-character origin story told over narration and fast cuts. We went there, we found this, we brought it back.

Typical build
Run band
60–120s
Characters
3–5
Frames
Single

Where it wins

  • Heritage and sourced brands where the journey genuinely is the pitch
  • Narration means nothing lip syncs, so a five character cast costs nothing in voice drift. This is the format that makes a large cast affordable
  • Long form absorbs a volume of footage a thirty second spot cannot hold

Where it breaks

  • Highest build of anything here. A hero creative, not something to run five of
  • Keeping three to five generated characters recognisably themselves across eighteen clips is the hardest continuity job in the catalog
  • If there is no real story, this format exposes that faster than any other
Non-obvious moveWrite the story as six sentences first. If sentence four is not surprising to someone outside your company, you have an About page, not a saga.

VO-Driven B-Roll

10
A frame in the VO-Driven B-Roll format

No talking head at all. Unboxing, texture, hands, product in use, with a voice over the top.

Typical build
Run band
30–60s
Characters
None
Frames
Single

Where it wins

  • Products where opening the box is the experience. Bedding, beauty, candles, drinks
  • Buyers who care how something feels more than what an expert says about it
  • No lip sync, no voice drift, no face to keep consistent. Fastest turnaround in the catalog

Where it breaks

  • Nobody vouches for the product on camera. No face to trust means no personal credibility
  • It cannot hold an authority claim or a heavy mechanism. Pair it with something that can
  • Hands are still the least reliable thing to generate. Frame tighter than you think and keep the motion simple
Non-obvious moveRecord the voiceover before generating anything. The script decides the clip count. The other order is how brands end up with pretty footage and nothing to say over it.

Product Hero

11
A frame in the Product Hero format

No people. Product shots, pours, opens and compositions, with a narrator laid over.

Typical build
Run band
24–48s
Characters
None
Frames
Paired

Where it wins

  • Hook economics. One set of visuals plus four voiceovers and four text hooks is four ads at close to the cost of one. Nothing else here is that cheap to multiply
  • Batches where building and keeping a character consistent is not worth the lift
  • Categories sold on how the thing looks and feels in the hand

Where it breaks

  • A narrator sounds like a third party, so hard closes land cold. First person does not work here
  • No person means no emotional journey. Problem and solution stories fall flat
  • Your packaging has to survive being generated. Legible labels and exact logos are the most common failure in this format. Composite the real pack where it matters
Non-obvious moveThe cheapest way to test four hooks properly. Build visuals once, write four completely different opening lines, ship four ads. Whatever wins tells you what to say everywhere else.

Native Feed Ad

12
A frame in the Native Feed Ad format

Built to not look like an ad. Lower polish on purpose, sitting flush with organic content.

Typical build
Run band
8–44s
Characters
0–1
Frames
Single

Where it wins

  • The cheapest and fastest lane. The split-screen version is two images and no clips at all
  • Faceless versions avoid the two hardest problems in AI video, faces and lip sync, entirely
  • Tired categories where anything polished instantly reads as an ad and gets scrolled

Where it breaks

  • Deliberately rough is one inch from actually bad. Models default to over-polishing, so roughness has to be asked for explicitly
  • Trend audio does not work here. A licensed track and a real micro-expression cannot be reproduced
Non-obvious moveScreenshot the first five organic posts in your own feed. Match that polish level exactly, not the level of the ads around them. That is the whole brief.
05

Run the whole thing through your own AI

Four prompts that do the actual work once the file is in the conversation.

Paste the file, then paste one of these

Replace anything in brackets.

Pick formats for your brand
My brand is [BRAND], we sell [PRODUCT] to [WHO]. We spend about [SPEND] a month on Meta. Pick the five formats that fit us best. For each, tell me why it fits, what our version would show on screen, and the specific way it would fail for a product like ours.
Audit what you already run
Here are the last twenty ads we launched: [ONE LINE EACH]. Tag each with a format from the catalog. Tell me how many distinct formats we run, which of the twelve we have never touched, and which untouched one you would bet on first.
Write a batch plan
Build me a twenty-ad batch for [BRAND] using the allocation model in the file. For each ad give me the format, the awareness stage, a one-line concept and the hook line. Do not repeat a concept shape twice.
Brief one ad properly
Using the [FORMAT] entry, write a full brief for [BRAND] selling [PRODUCT]. Give me the image list, the clip list with a duration for each, what is said over each clip, and flag anything on that format's "where it breaks" list my brief is at risk of.
format-catalog.md
Plain text  ·  ~24 KB
Download the file
  • All twelve formats, structured
  • Typical build for each
  • The four category reads
  • The allocation model
  • The scoring instrument
  • Four ready-made prompts
06

What is working, by category

Directional

Operator reads from the accounts we build for, not a published study. We are calling them directional on purpose, because a format ranking drawn from a handful of accounts is a starting bet, not a benchmark. Point your first three swings here, then let your own data overrule us.

01

Supplements

Leads

Clinical explainer and object talk. The mechanism is invisible, so a person describing it is the weakest tool available. Stylized registers also carry claims that sound forced from a human mouth, and they sidestep the realism problem entirely.

Lags

Pure product hero. Nobody buys a capsule because the bottle looked good.

The trap

Every competitor is running an ingredient lecture. The differentiator is the format the lecture arrives in, not the ingredients.

02

Skincare and beauty

Leads

Education-led structures. Judged panel and clinical explainer both carry them. The strongest account in our own book runs an education-first spine, not a testimonial-first one.

Lags

Generic before-and-after talking heads. Saturated, and they attract the most claim scrutiny of anything here.

The trap

The actives work under the surface, so a generated camera can only ever show a face. Diagram or animate the layer that cannot be filmed at all. That is the one real advantage of building this way.

03

GLP-1, telehealth and regulated

Leads

Native feed and clinical explainer. Faceless and diagrammatic avoids the body-morph problem entirely, and both are cheap enough to run at the volume a regulated account needs.

Lags

Dramatic transformation footage. Most likely to be rejected, and body transformation is one of the least reliable things to generate convincingly.

The trap

Use transformation proxies instead of transformations. A simultaneous split screen, a loose waistband, a visualisation of the invisible. Same story, nothing morphs.

04

Home, comfort and consumable CPG

Leads

On-camera demonstration and two-person dialogue. Both consistently outrun a plain talking head. The mechanism is physical, so show it moving in the first five seconds.

Lags

Plain talking head with no demo. It drops off unless an unusually strong hook is doing all the work alone.

The trap

If your product can be judged by feel, taste or smell, the blind sense test is nearly free and almost nobody in the space runs one.

One pattern held across all four and it had nothing to do with format. When we audited a competing script set against ours in one account, their thirty-one scripts mentioned the brand's headline guarantee zero times. Ours carried the same four elements every time: the guarantee, the mechanism, a reason to move now, and an actual ask. Before changing format, check all four are in the script you already have. That is a free fix and it outranks everything else on this page.

07

How to spread twenty ads

Motion's finding was that more creative wins, not more of the same creative. Here is how we allocate a twenty-ad month so the swings are actually different. Adjust the counts, keep the shape.

8

The engine

Talking head and native feed. High volume, low build, fast to iterate. This is where you find hooks, not where you find formats.

40%
5

The teachers

Clinical explainer, object talk, product hero. No characters to keep consistent, and they carry the mechanism. If your product needs explaining, this block is not optional.

25%
4

The pattern breaks

Claymation, blind sense test, podcast. These are what stop a batch looking like one batch. Cut this block and the account fatigues faster.

20%
2

The heroes

Longform testimonial, founder saga, judged panel. Expensive, slow, and worth it once or twice a month. Never more.

10%
1

The one you have never run

Whatever is left on the list. It is one ad. The cost of being wrong is one ad, and the cost of never trying is the format you would have found.

5%

The second axis, and the one most brands miss. Cut the same twenty by how much the viewer already knows: 30 to 40% for people who do not yet know they have the problem, 40 to 50% for people comparing solutions, 10 to 20% for people ready to buy. Almost every account we open runs that last group at close to 100%, which is exactly why it stops scaling. You cannot buy new customers with ads written for existing ones.

08

Picking one, in three questions

Question one

Who has to be believed?

A peer, an authority, or nobody. Peer sends you to talking head or testimonial. Authority sends you to judged panel or clinical explainer. Nobody sends you to product hero, b-roll or object talk.

Question two

Does anyone need to speak on camera?

The AI-specific question, and it decides more than people expect. Yes puts you in the in-model voice lane with its lip sync limits and voice drift. No makes every one of those problems disappear and your build gets faster immediately.

Question three

What is already in the batch?

Run the character counts and the voice column against what you shipped last month. If every ad is one generated person talking for thirty seconds, the gap is structural and no amount of copywriting closes it.

If you want these built instead of just listed

We are an AI creative studio for direct-to-consumer brands. We build every format on this page as performance ads, at volume, for brands spending real money on Meta.

18%
lower CPA for a DTC supplement brand at $40K a month in spend
3.1x
ROAS on winning creatives
8–12
winning creatives produced per month
12+
brand partners

Prove It is $997. Three performance videos and five statics, and it credits toward any of our packages.

And the honest version: if you already have a studio you like, the file above works fine without us. The picking is the part that pays.

Book a call