# The AI Ad Format Catalog

Twelve AI video ad formats that are working right now, structured for a language model.
From Whim Creative. Free to use, copy, or hand to anyone.

Every format in this file is generated, not filmed. No shoot, no crew, no talent, no licensing.
That changes which formats are cheap and which are expensive, and it is not the order you would
expect from traditional production.

---

## HOW TO USE THIS FILE

Paste this entire file into Claude or ChatGPT, then add one of the prompts below.

**To pick formats for your brand**

> Above is a catalog of twelve AI video ad formats. My brand is [BRAND], we sell [PRODUCT] to
> [WHO]. We spend about [SPEND] a month on Meta.
> Pick the five formats that fit us best. For each one, tell me why it fits, what our version
> would actually show on screen, and the specific way it would fail for a product like ours.

**To audit what you already run**

> Above is a catalog of twelve AI video ad formats. Here are the last twenty ads we launched:
> [PASTE A LIST, one line each, describing what happens in the ad].
> Tag each ad with a format from the catalog. Tell me how many distinct formats we run,
> which of the twelve we have never touched, and which untouched one you would bet on first.

**To write a batch plan**

> Above is a catalog of twelve AI video ad formats. Build me a twenty-ad batch for [BRAND].
> Use the allocation model in the file. For each ad give me the format, the awareness stage,
> a one-line concept, and the hook line. Do not repeat a concept shape twice.

**To brief one ad properly**

> Using the [FORMAT NAME] entry above, write a full brief for [BRAND] selling [PRODUCT].
> Give me the image list, the clip list with a duration for each, what is said over each clip,
> and flag anything on that format's "where it breaks" list that my brief is at risk of.

---

## HOW TO READ EACH ENTRY

An AI ad is built from three things, and the mix is what makes one format cheap and another
expensive.

- **Images**: the still frames generated first. Every clip starts from one. Some formats need a
  matched pair, a start state and an end state, which roughly doubles the image count for that
  format and makes it much less forgiving.
- **Clips**: the video generations. Each one is a separate roll of the dice, so this number is the
  best single predictor of how long a format takes and how much it costs.
- **Voice**: either generated inside the video model along with the picture, or recorded separately
  and laid over silent footage.

**The voice split is the most useful line in this file.** In-model voice gives you a character who
genuinely speaks on camera, and it brings two problems that only exist in that lane: lip sync
degrades on longer clips, and the same character drifts into sounding like a different person
across a sequence. Recorded-separately voice has neither problem, ever. It just cannot give you a
talking face.

Half the "where it breaks" notes below are downstream of that one choice.

- **Run band**: the length range the format works inside. Outside it, the format stops behaving.
- **Characters**: how many generated people appear on screen.
- **Look**: the visual register. These do not blend. Photoreal and stylized are different rule sets.

Typical builds below are what a normal version of each format actually takes. Ranges exist, but a
typical build is more useful than a range when you are planning a month.

---

# THE TWELVE

## 1. Talking Head UGC

A generated person talks to camera like they are recording on their phone. The workhorse, and the
format most batches are built on.

- Typical build: 7 images, 7 clips
- Voice: generated in-model
- Frames: single start frame per clip
- Run band: 15 to 35 seconds
- Characters: 1
- Look: phone-shot realism

**Where it wins**
- Cold traffic, where a peer voice beats an authority voice.
- Problem and solution stories that need a face to carry them.
- Volume. It is the cheapest format to run many versions of.

**Where it breaks**
- The voice drifts. Past roughly five clips the same character starts sounding like someone else,
  unless the voice description is locked word for word and reused on every single clip.
- Lip sync degrades past about ten seconds in one clip. Write to that rather than fixing it after.
- One room for forty seconds is boring. Move the character or cut away.

**Non-obvious move**
Rebuild your best performing script with a different age of character in a different room. Same
words. That is a genuinely separate swing, not a variation, and it costs one build.

---

## 2. Longform Testimonial

A generated person tells a transformation story, with a cutaway on almost every line so each claim
gets a picture.

- Typical build: 18 images, 18 clips
- Voice: in-model on the talking clips, silent cutaways
- Frames: single start frame per clip
- Run band: 60 to 120 seconds
- Characters: 1
- Look: phone-shot realism

**Where it wins**
- High intent audiences who will genuinely watch two minutes.
- Before and after proof, and products that need explaining time.
- A time-anchored hook that maps to the viewer's own calendar.

**Where it breaks**
- Clip count. Eighteen generations makes this a hero piece, not a batch filler.
- Keeping one generated person recognisably the same across all those cutaways is the hard part,
  and it is invisible until it is wrong.
- Cutaways planned at script time work. Cutaways added later never match the character.

**Non-obvious move**
Write the script first, then mark every sentence that makes a claim. Each mark is one cutaway you
owe. If you have four marks, this is the wrong format.

---

## 3. Podcast Two-Hander

Two generated characters in conversation, one pitching the other. Each is generated separately and
cut together. They never share a frame.

- Typical build: 12 images, 12 clips
- Voice: generated in-model
- Frames: single start frame per clip
- Run band: 30 to 90 seconds
- Characters: 2
- Look: studio conversation

**Where it wins**
- Reveal hooks. "Wait, what is that?" is an entire opening.
- Authority and status positioning, where being overheard beats being told.
- Industry claims that land better as gossip than as a pitch.

**Where it breaks**
- Twice the everything. Two characters is double the generations and double the work keeping each
  face consistent across a sequence.
- Wrong pairing kills the read instantly. Cinematic music under a raw conversation is fake in one
  second.

**Non-obvious move**
If the second character never has to be seen, do not generate them. Put them off camera with their
own voice. One body on screen, two people in the scene, roughly half the build. State explicitly in
the prompt that the visible character is not the one speaking, or the model will sync their mouth
to the wrong voice.

---

## 4. Judged Panel

An authority rules on competing products and yours wins. Every character is generated alone; the
panel exists only through eyelines, seating geometry and the edit.

- Typical build: 10 images, 24 clips
- Voice: generated in-model
- Frames: single start frame per clip, character designs reused across many clips
- Run band: 60 to 90 seconds
- Characters: 3 or more
- Look: produced television

**Where it wins**
- A verdict story. Products compete, an expert rules, yours wins on a stated reason.
- Brands that want a produced television feel instead of phone footage.
- Three or more characters who never need to share a frame, which is exactly what AI is good at and
  a real shoot is not.

**Where it breaks**
- The run time explodes and nobody notices. Beats times rivals times clip length is your run time.
  Three rivals with a full arc quietly becomes a four minute ad.
- The fix is to cut rivals, not to trim clips at random.
- Note the ratio above: ten images carrying twenty-four clips. That reuse is the only reason this
  format is affordable, and it only works if the character designs are locked first.

**Non-obvious move**
Do the arithmetic before writing. Beats x rivals x seconds. If the answer is over ninety, drop a
rival. Two rivals plus your hero lands in a normal band every time.

---

## 5. Blind Sense Test

A generated person blind-tests products against name brands and the hidden one wins. The idea does
the work; the build is simple.

- Typical build: 9 images, 9 clips
- Voice: recorded separately, reactions carry the story
- Frames: single start frame per clip
- Run band: 30 to 60 seconds
- Characters: 1
- Look: phone-shot realism

**Where it wins**
- Anything testable by smell, taste, sound, feel or texture.
- The curiosity gap is free. "Which one wins" holds attention with no effort.
- Reactions are non-verbal, so this format sidesteps lip sync entirely while still having a person
  on camera. Very few formats can say that.

**Where it breaks**
- A weak idea takes the whole ad with it. There is no craft layer to hide behind.
- A reaction that looks performed kills it instantly, and models overshoot expressions by default.
  Ask for restraint explicitly.
- Competitor products in frame is generally fine. Disparaging them in the voiceover is a legal
  question, not a creative one. Check it.

**Non-obvious move**
Name the three biggest brands in your category out loud. If you would genuinely bet on your product
in a blind test against them, this format is nearly free. If you would not, that is worth knowing.

---

## 6. Clinical Explainer

An isolated subject on a plain background morphing from one state to the next, with a calm narrator
over the top. Evidence, not advertising.

- Typical build: 16 images, 8 clips
- Voice: recorded separately
- Frames: matched pairs, a start state and an end state for every clip
- Run band: 20 to 50 seconds
- Characters: none
- Look: evidence and documentary

**Where it wins**
- Teaching a mechanism nobody can see, which is most supplements and most skincare.
- Buyers who trust authority more than a stranger's testimonial.
- It is modular. A few of these beats drop into any other ad as the "how it works" section.
- Internal biology that photoreal models often refuse to generate will usually pass in this
  stylized register. It is the practical route to showing what happens inside the body.

**Where it breaks**
- It has no warmth and it is not supposed to. Use it to explain, never to make someone feel
  something.
- A flat narrator sinks it. The voice is doing more work than the picture.
- Look at the image-to-clip ratio: two images for every clip. This is the least forgiving format in
  the catalog, because a sloppy start and end pair does not produce a weak clip, it produces
  garbage.

**Non-obvious move**
Build six beats once as a standalone ad, then reuse the same six as the middle section of three
other ads. It is the only genuinely modular format here.

---

## 7. Object Talk

An object with a face talks to camera in first person. An ingredient, a body process, or the cheap
thing the customer tried before you.

- Typical build: 2 images, 3 clips
- Voice: generated in-model
- Frames: one character design, reused across every clip
- Run band: 8 to 40 seconds
- Characters: none human
- Look: stylized animation

**Where it wins**
- The strongest scroll stop in the catalog. Nothing else in the feed looks like it.
- It turns a lecture into a character monologue, so people sit through the science.
- The cheapest build here by a wide margin. One character design carries the whole ad, so
  consistency costs almost nothing once the design is right.

**Where it breaks**
- A cartoon register does not fit every brand. If you are locked to photoreal, skip it.
- Getting a genuinely expressive face onto an abstract object takes several attempts. Budget four
  to six image tries before you get the character.
- Animation timed to specific words is unreliable across a long clip. Keep beats short.

**Non-obvious move**
Make the talking object the customer's old failed solution, confessing. "I'm your drugstore
moisturizer. I'm full of cheap oils that keep my price down and clog your pores." The old product
indicts itself and yours becomes the obvious upgrade, with no hard sell and no competitor named.

---

## 8. Claymation and Stylized Animation

A clay, Pixar or otherwise made-not-filmed register. Products become characters and worlds get
built.

- Typical build: 14 images, 9 clips
- Voice: recorded separately, usually with a custom score
- Frames: mixed, some matched pairs for transitions
- Run band: 30 to 60 seconds
- Characters: 0 to 3 animated
- Look: stylized animation

**Where it wins**
- Visual relief in a batch that is otherwise wall to wall talking heads.
- It carries claims that sound forced from a human mouth. Small workers building your skin layer by
  layer is fine in clay and absurd in live action.
- No lip sync problem, no face consistency problem, no realism uncanny valley. Three of the hardest
  things in AI video simply do not apply.

**Where it breaks**
- Stylized and photoreal cannot be mixed. It is a different rule set, not a filter you apply.
- Not every brand survives being cute. If the category is serious, this can read as unserious.
- The finish is where it is won or lost. A rough clay ad reads as a cheap cartoon; a scored and
  graded one reads as a television spot.

**Non-obvious move**
Take the one claim your compliance reviewer keeps softening. Animation is usually where an
over-literal claim becomes an obviously figurative one, which is a different conversation.

---

## 9. Founder Saga

A multi-character origin story told over narration and fast cuts. We went there, we found this, we
brought it back.

- Typical build: 18 images, 18 clips
- Voice: recorded separately
- Frames: single start frame per clip
- Run band: 60 to 120 seconds
- Characters: 3 to 5
- Look: narrated story

**Where it wins**
- Heritage and sourced brands where the journey genuinely is the pitch.
- Narration means nothing lip syncs, so a five character cast costs nothing in voice drift. This is
  the format that makes a large cast affordable.
- Long form absorbs a volume of footage a thirty second spot cannot hold.

**Where it breaks**
- Highest build of anything here. This is a hero creative, not something to run five of.
- Keeping three to five generated characters recognisably themselves across eighteen clips is the
  hardest continuity job in the catalog.
- If there is no real story, this format exposes that faster than any other.

**Non-obvious move**
Write the story as six sentences first. If sentence four is not surprising to someone outside your
company, you do not have a saga, you have an About page.

---

## 10. VO-Driven B-Roll

No talking head at all. Unboxing, texture, hands, product in use, with a voice over the top.

- Typical build: 10 images, 10 clips
- Voice: recorded separately
- Frames: single start frame per clip
- Run band: 30 to 60 seconds
- Characters: none
- Look: tactile lifestyle

**Where it wins**
- Products where opening the box is the experience. Bedding, beauty, candles, drinks.
- Buyers who care how something feels more than what an expert says about it.
- No lip sync, no voice drift, no face to keep consistent. Fastest turnaround in the catalog.

**Where it breaks**
- Nobody vouches for the product on camera. No face to trust means no personal credibility.
- It cannot hold an authority claim or a heavy mechanism. Pair it with something that can.
- Hands are still the least reliable thing to generate. Frame tighter than you think and keep the
  motion simple.

**Non-obvious move**
Record the voiceover before generating anything. The script decides the clip count. The other order
is how brands end up with forty seconds of pretty footage and nothing to say over it.

---

## 11. Product Hero

No people. Product shots, pours, opens and compositions, with a narrator laid over.

- Typical build: 12 images, 7 clips
- Voice: recorded separately
- Frames: matched pairs on the motion clips
- Run band: 24 to 48 seconds
- Characters: none
- Look: commercial polish

**Where it wins**
- Hook economics. One set of visuals plus four voiceovers and four text hooks is four ads at close
  to the cost of one. Nothing else here is that cheap to multiply.
- Batches where building and keeping a character consistent is not worth the lift.
- Categories sold on how the thing looks and feels in the hand.

**Where it breaks**
- A narrator sounds like a third party, so hard closes land cold. First person does not work here.
- No person means no emotional journey. Problem and solution stories fall flat.
- Your actual packaging has to survive being generated. Legible labels and exact logos are the
  single most common failure in this format, so plan to composite the real pack where it matters.

**Non-obvious move**
This is the cheapest way to test four hooks properly. Build the visuals once, write four completely
different opening lines, ship four ads. Whatever wins tells you what to say in every other format.

---

## 12. Native Feed Ad

Built to not look like an ad. Lower polish on purpose, sitting flush with organic content.

- Typical build: 6 images, 6 clips. The split-screen version is 2 images and no clips at all
- Voice: either
- Frames: single start frame per clip
- Run band: 8 to 44 seconds
- Characters: 0 to 1
- Look: anti-polish organic

**Where it wins**
- The cheapest and fastest lane. The split-screen version generates no video whatsoever.
- Faceless versions avoid the two hardest problems in AI video, faces and lip sync, entirely.
- Tired categories where anything polished instantly reads as an ad and gets scrolled.

**Where it breaks**
- Deliberately rough is one inch from actually bad. Lo-fi is a craft, not the absence of one, and
  models default to over-polishing, so roughness has to be asked for explicitly.
- Trend audio does not work here. A licensed track and a real micro-expression cannot be
  reproduced.

**Non-obvious move**
Screenshot the first five organic posts in your own feed. Match that polish level exactly, not the
level of the ads around them. That is the whole brief.

---

# WHAT IS WORKING, BY CATEGORY

Evidence grade: DIRECTIONAL. These are operator reads from the accounts we build for, not a
published study. Treat them as where to point your first three swings, then let your own data
overrule us.

## Supplements
- **Leads**: clinical explainer and object talk. The mechanism is invisible, so a person describing
  it is the weakest tool available. Stylized registers also carry claims that sound forced from a
  human mouth, and they sidestep the realism problem entirely.
- **Lags**: pure product hero. Nobody buys a capsule because the bottle looked good.
- **The trap**: every competitor is running an ingredient lecture. The differentiator is the format
  the lecture arrives in, not the ingredients.

## Skincare and beauty
- **Leads**: education-led structures. Judged panel and clinical explainer both carry them.
- **Lags**: generic before-and-after talking heads. Saturated, and they attract the most claim
  scrutiny of anything here.
- **The trap**: the actives work under the surface, so a generated camera can only ever show a
  face. Diagram or animate the layer that cannot be filmed at all, which is the one real advantage
  of building this way.

## GLP-1, telehealth and regulated
- **Leads**: native feed and clinical explainer. Faceless and diagrammatic avoids the body-morph
  problem entirely, and both are cheap enough to run at the volume a regulated account needs.
- **Lags**: dramatic transformation footage. Most likely to be rejected, and body transformation is
  one of the least reliable things to generate convincingly.
- **The trap**: use transformation proxies instead of transformations. A simultaneous split screen,
  a loose waistband, a visualisation of the invisible. Same story, nothing morphs.

## Home, comfort and consumable CPG
- **Leads**: on-camera demonstration and two-person dialogue. Both consistently outrun a plain
  talking head. The mechanism is physical, so show it moving in the first five seconds.
- **Lags**: plain talking head with no demo. It drops off unless an unusually strong hook carries
  it alone.
- **The trap**: if your product can be judged by feel, taste or smell, the blind sense test is
  nearly free and almost nobody in the space runs one.

## The pattern that held across all four
When we audited a competing script set against ours in one account, their thirty-one scripts
mentioned the brand's headline guarantee zero times. Ours carried the same four elements every
time: the guarantee, the mechanism, a reason to move now, and an actual ask.

**Before changing format, check that all four are in the script you already have.** That is a free
fix and it outranks everything else in this file.

---

# THE ALLOCATION MODEL

For a twenty-ad month. Adjust the counts, keep the shape.

| Block | Count | Formats | Job |
|---|---|---|---|
| The engine | 8 | Talking head, native feed | Find hooks. High volume, low build, fast to iterate |
| The teachers | 5 | Clinical explainer, object talk, product hero | Carry the mechanism. No characters to keep consistent |
| The pattern breaks | 4 | Claymation, blind sense test, podcast | Stop the batch looking like one batch |
| The heroes | 2 | Longform testimonial, founder saga, judged panel | Expensive, slow, worth it twice a month. Never more |
| The unknown | 1 | Whatever you have never run | The cost of being wrong is one ad |

**The second axis, and the one most brands miss.** Cut the same twenty by how much the viewer
already knows:

- 30 to 40 percent for people who do not yet know they have the problem
- 40 to 50 percent for people comparing solutions
- 10 to 20 percent for people ready to buy

Almost every account we open runs that last group at close to 100 percent, which is exactly why it
stops scaling. You cannot buy new customers with ads written for existing ones.

---

# THE SCORING INSTRUMENT

Ten minutes on your last twenty ads. No tool required.

1. List the last twenty creatives you launched.
2. Tag each with a format from this catalog. If it does not fit, write "other."
3. Count how many **distinct** formats appear.

| Distinct formats | Read |
|---|---|
| 1 to 2 | You have one swing |
| 3 to 4 | Where most brands land |
| 5 to 7 | Healthy |
| 8 or more | Top quartile behaviour |

If your number is under five, your next winner is probably a different shape, not a better script.

---

# PICKING ONE, IN THREE QUESTIONS

**1. Who has to be believed?**
A peer, an authority, or nobody. Peer sends you to talking head or testimonial. Authority sends you
to judged panel or clinical explainer. Nobody sends you to product hero, b-roll or object talk.

**2. Does anyone need to speak on camera?**
This is the AI-specific question and it decides more than people expect. If yes, you are in the
in-model voice lane and you inherit lip sync limits and voice drift. If no, every one of those
problems disappears and your build gets faster and cheaper immediately.

**3. What is already in the batch?**
Run the character counts and voice column against what you shipped last month. If every ad is one
generated person talking for thirty seconds, the gap is structural and no amount of copywriting
closes it.

---

# WHY ANY OF THIS MATTERS

Motion analysed 1.3 billion dollars in ad spend across more than 550,000 ads from over 6,000
advertisers on Facebook and Instagram, September 2025 through early January 2026.

- About 5 percent of ads launched become winners. A winner spends at least ten times the account's
  median ad.
- About 55 percent of all spend flows to that 5 percent. Roughly 17 percent goes to outright
  losers.
- The top quartile in every budget tier launches two to three times more creative than same-budget
  peers.
- Brands that launch more creative get about twice the winners on an identical budget.

Source: Motion, Creative Benchmarks 2026.
https://motionapp.com/thumbstop-pulse/creative-benchmarks-2026

More money does not buy more winners. More swings does. And a swing only counts if it is genuinely
different, because twenty versions of one talking head is one swing with twenty price tags on it.

Generating the ads instead of filming them is what makes twelve formats a realistic month rather
than a wish. There is no shoot to schedule, so the constraint stops being budget and starts being
whether anyone decided what to make.

---

# WHO MADE THIS

Whim Creative. We build AI video ads for direct-to-consumer brands on Meta.

Prove It is $997: three performance videos and five statics, credited toward any package.

Book a call: https://calendly.com/leo-whim/discovery

If you already have a studio you like, this file works fine without us. The picking is the part
that pays, not the file.
