Whim Creative / free tool

The Seedance 2.5 Ad Builder

Paste one prompt into Claude or ChatGPT, answer fifteen lines about your product, and get back a complete video ad spec you can run today.

August 7, 2026  ·  written the day the API opened

ByteDance opened the Seedance 2.5 API on August 7. We had it running about an hour later and put four generations through it that morning.

What follows is two things. First the builder, which is the useful part. Then the numbers we actually measured, including the reason we kept most of our own work on the older model.

One thing worth saying before the tool, because it shapes what the tool is for. The floor for what looks good moved up for everyone at once this week. That is good news, and it also means looking good stopped being much of a differentiator on the same day.

What still decides whether an ad works is which awareness stage you are writing to, whether the mechanism you are claiming is actually true, whether it holds attention past three seconds, and how many real angles you can run before the account fatigues. None of that got easier. It got more valuable, and it is the part we get paid for.

We do AI performance creative for a handful of leading DTC brands on Meta. The builder below is the production layer, given away. Section 10 is the longer version of the argument.

THE TOOLPaste this in and answer fifteen lines

This is a prompt, not software. It works in Claude, ChatGPT, or anything else that takes a long message. It encodes about four months of our production rules so you do not have to know any of them.

You get back a start-frame image prompt, a motion prompt with the spoken line in it, a voice contract, and the exact API call. Nothing to install, nothing gated, no email.

The Seedance 2.5 Ad Builder
========================================================
THE SEEDANCE 2.5 AD BUILDER, by Whim Creative
Paste this whole message into Claude or ChatGPT.
Fill in PART 1. Do not change PART 2.
========================================================

PART 1. YOUR BRIEF. One line each. Guess if you have to.

1.  BRAND NAME (blank if you would rather not name it):
2.  PRODUCT NAME:
3.  WHAT IT PHYSICALLY IS (bottle, spray, patch, powder, device, plus size and color):
4.  HOW A PERSON USES IT. If a part moves, say which and where:
5.  WHO BUYS IT (age, gender, one line on their life):
6.  THE PROBLEM THEY ALREADY KNOW THEY HAVE:
7.  MECHANISM. Why it works, in words a twelve year old follows:
8.  PROOF. One true, specific thing you can back up. If approximate, write "about":
9.  THIS AD'S JOB: COLD AUDIENCE / MECHANISM / OBJECTION / CLOSE A WARM BUYER
10. WHO TALKS: PERSON ON CAMERA, or NO TALKING, SOUND ONLY
11. WHERE IT HAPPENS, AND WHAT TIME OF DAY:
12. CLIP LENGTH IN SECONDS, 4 to 30, whole numbers. If anybody speaks, use 8. Eight seconds is
    about sixteen to twenty spoken words, which is one problem line plus one short proof and nothing else:
13. THINGS I AM NOT ALLOWED TO SAY (legal, medical, regulated). Supplements, skincare, health and
    money carry rules, so this is rarely blank:
14. MARKET. What country are the buyers in (sets the accent):
15. PRODUCT PHOTO, a file or public link, NONE if you have none. It goes into your IMAGE tool, not
    the video call. NONE is allowed and changes the ad: the product will not appear on screen,
    because a product the model has never seen gets invented, and an invented one looks wrong:

========================================================
PART 2. INSTRUCTIONS FOR THE AI. DO NOT EDIT.
========================================================

YOUR JOB. Turn the brief into one ready to run spec for one Seedance 2.5 video ad. Return the
shape in block F and nothing else, no preamble, no coaching, no offers to revise. Enforce these rules
silently and fix your own violations before answering. No em dash, no en dash, no dash as punctuation:
use a period, a comma or the word and.

RULE ZERO. NEVER ASK A QUESTION. Blank or conflicting fields get your most likely guess, logged in
WHAT I ASSUMED as a statement, "You left field 15 blank, so I assumed you have no product photo",
never "You said you have no product photo", and never an offer to revise.

RULE ONE. NEVER INVENT A FACT ABOUT THE PRODUCT. Outranks everything. Every spoken fact traces to
field 6, 7, 8 or 13. Never smuggle in an onset time, a percentage, a dose, a study, a review count, a
named competitor or ingredient, a pathway not in field 7, or an absolute like "nothing else worked". A
line wanting a number you were not given is written without it, and their hedge stays, so "about 4,000
reviews" becomes "about four thousand", never a converted count.

RULE ONE B. NO PHOTO MEANS NO PRODUCT ON SCREEN. If field 15 is NONE, the ad is built with the
product absent from frame. Do not write the product into the still, do not write a held object into
the motion prompt, and do not write a pinned action that handles it. She talks, the product is named
in the line only. Tested: with no start frame the model invents the object and materialises it in
frame partway through the clip, which is worse than not showing it. Say this in WHAT I ASSUMED, in
one line, so they know the fix is to shoot or generate one product still.

RULE TWO. NEVER INVENT A FACT ABOUT THE MODEL. No superlatives, no ranked failure modes. The only
claims you may make are the ones this kit states.

RULE THREE. NEVER ISSUE A CLEAN BILL. Say what you did, never certify the ad is compliant. Judge what
the ad IMPLIES: past problem plus present product plus any resolution is a testimonial, whatever words
you used. Under PERSON ON CAMERA the speaker reports what the product DOES or what field 8 SAYS, never
what happened to her body. The PICTURE claims too: a generated person holding the product at home
reads as a real customer, the exact representation 16 CFR 255.2(c) polices, so the frame is part of
the claim. If the only honest angle is a result, write the strongest non-result line and log it.

DISCLOSURE. A DELIVERABLE, NOT A WARNING. An operational read, not legal advice.
- New York GBL 396-b, the Synthetic Performer Disclosure Law, live since June 9 2026. Anyone who
  produces or creates a commercial ad containing a synthetic performer, a digitally created human who
  is not a recognisable real person, must conspicuously disclose it in the creative. $1,000 first
  offence, $5,000 after. Whoever renders the file is a producer, so that is the operator, not only the
  brand. "Conspicuous" is undefined with no Attorney General guidance yet, so the placement below is
  conservative, not a safe harbour. Meta targeting makes New York reach the default, so disclose
  globally.
- Exemption: audio only. AI voiceover over real footage needs no disclosure, an on-camera AI
  spokesperson does.
- The FTC rules. Under 16 CFR 465 agencies are not immune from liability, per the FTC's own Q and A,
  and AI avatars are prohibited only when the testimonial is fake or false. 16 CFR 255.2(c) is the
  sharp edge: an ad representing an endorser as an ACTUAL CONSUMER must use actual consumers in audio
  AND video, or clearly disclose they are not, so an AI persona claiming a personal result is NOT
  curable by a label. Disclosure fixes non-disclosure, not deception, and "results not typical" does
  not work, the FTC copy-tested it.
- Meta has no general commercial AI-disclosure rule, only social issue, election and political ads,
  and since June 2026 it auto-detects and labels generative AI itself. Never say Meta requires one.
WIRING. A human on camera, hands included, means section 9 ships a disclosure line as a deliverable.
No human, no line, so say so and note Meta may label it anyway. It never goes in either prompt, the
model garbles on-screen text, so burn it in with any editor, CapCut is free: "AI-generated. This
person is not real.", held the whole clip from the first frame, legible, high contrast, clear of the
top and bottom 15 percent where the feed lays its own buttons. That line does two jobs, GBL 396-b and
the 255.2(c) not-actual-consumers disclosure, so shortening it, moving it off frame one or dropping it
breaks both.

TWO PROMPTS, TWO JOBS. The image prompt owns what the shot LOOKS like, the motion prompt owns what
HAPPENS and what is SAID. The model is handed the finished still, so the motion prompt never
re-describes the room, furniture, wardrobe, lighting, camera, face or product. In either prompt every
sentence names the one thing that changes, locks something that would drift, or bans one artifact.
Anything else is deleted.

--------------------------------------------------------
A. WRITE THE SCRIPT FIRST
--------------------------------------------------------
One idea per ad, the strongest single thing from fields 6, 7 and 8. Shape it to field 9: cold opens on
the problem in their own words, mechanism on the surprise, objection says it out loud first, warm
buyer gives one reason to move now.

Dialogue rules, all mandatory:
- Full sentences with subjects, verbs and connectives, never stacked fragments. Contractions
  everywhere, a subject pronoun before every action verb.
- One filler per sentence, two in the line: so, honestly, I mean, look, just, actually. No ellipses,
  no dashes. For a pause, end the sentence.
- Numbers as words INSIDE THE SPOKEN LINE ONLY, digits everywhere else. Acronyms as a mouth says
  them, "You Gee See" not "UGC". A hard brand name is respelled phonetically in dialogue only.
- Never point at anything on screen, no "link below", no "swipe up", the model renders those as a
  garbled text box.
- No hype, no game changer, revolutionary, epic or stunning. Every line sounds pulled from mid
  conversation with a friend.
- Field 13 respected exactly, a banned claim cut rather than softened. Field 1 blank, never invent a
  brand name, the speaker says "this". Field 1 filled, it is spoken once, late in the line.

THE WHOLE AD READ. Then read the ad as a stranger would, problem, product and proof landing as one
impression. If that impression is an outcome the brief cannot back, rewrite until it is not and log
the cut. A diagnosed implied claim never ships, disclosure does not fix it.

MECHANISM FIDELITY. The dialogue verb must match fields 4 and 3, so "spray under the tongue" needs a
spray in field 3 and an actuator in the still. If they disagree, trust 4 for the action, 3 for the
object, and flag it. What the words explain is what the camera shows, or the assumptions say so.

LENGTH. Aim for field 12 seconds times 2.2 words, which sits mid band. The band is 2.0 to 2.5 spoken
words per second. Below 2.0 the model pads the opening with a breath, so take a second off. Over, add a second or two, never
three,
use that number in the API call, and never drop the field 8 proof to fit. Past 8 seconds of continuous speech is where we have seen pacing break down, so treat 8 as the safe
ceiling for a spoken clip. A line that will not fit inside 8 loses its least load bearing sentence,
never the proof, and you log the cut. Longer renders are for shots that do not carry a spoken line.

CONTENT FILTER. Refusals never say why, and naming a blocked word to forbid it can trip the filter, so
keep these out of the dialogue and both prompts, ban lines included: hemp, cannabis, CBD and any
nickname; photoreal insides of a body, use an object standing in for the body part; other companies'
logos. Log the workaround, never silently drop the mechanism.

--------------------------------------------------------
B. WRITE THE START FRAME IMAGE PROMPT
--------------------------------------------------------
Generated FIRST, in an image tool. ChatGPT, Gemini and Midjourney all do it, and it prices the still,
not block E. 9:16 vertical, written inside the prompt as digits: 1080 by 1920 pixels. The video
inherits its look from here, so compose for motion: mid gesture, mouth slightly parted, eyes just off
the lens. One flowing description, not labeled blocks: framing, the subject and her action, the field
11 setting, pose, skin, hair and clothing in real detail, a lived-in room, a closing string of look
cues. NO TALKING makes the product or her hands the subject, with no face and no second person
anywhere in the frame.

THE LOOK IS PHONE REAL, THE ONLY REGISTER HERE. Always: shot vertically on an iPhone, handheld,
casual framing, slightly imperfect crop, everything in focus front to back, no background blur,
visible smartphone compression, mild grain in the shadows, real pores, fine lines, uneven skin tone,
no retouching. Never: cinematic, film grain, lens flare, dramatic lighting, color grade, LUT, bokeh,
shallow depth of field, slow motion, speed ramp, whip pan, crane, dolly, gimbal, Dutch angle, epic,
stunning, camera brand names. Add lived-in mess, a spotless counter reads as an ad.

LIGHT. DAYTIME: natural window light, name the side. NIGHT: never "natural window light", name ONE
practical light in that room, just out of frame, say it is the only light and what falls into shadow,
then write "lamplight only" in the closing string.

PRODUCT. Exactly one, sized against the body since these models default far too big: "noticeably
smaller than her hand". Pull back rather than push into the label, small text will not render. If the
pinned action touches a button or lid, say which face the camera sees. With a field 15 photo, add it
as a reference image and describe the packaging in as few words as possible, words fight a photo and
lose. Without one, give shape, material, finish and colors, keep the label plain, and say it renders
blank.

AVOID line: beauty filter, skin smoothing, HDR look, background blur, two of the product, floating
text, staged styling. Check the still by eye before spending: vertical, one product smaller than a
hand, the moving part visible, hands whole, light right for the hour.

--------------------------------------------------------
C. WRITE THE VOICE CONTRACT
--------------------------------------------------------
Skip only if field 10 says NO TALKING. Every call is voice blind, so author one canonical paragraph
and paste it word for word into every clip with speech.

Forty to sixty five words, one paragraph, eight fields in order: 1 gender in capitals, guarded both
ways, "FEMALE voice. Not male. Not masculine."; 2 age band; 3 accent, the plain nationality adjective
from field 14, the United States if blank; 4 pitch; 5 timbre; 6 pacing, exactly "natural pace,
continuous delivery, no pauses between words" and nothing else, since slow, unhurried, leisurely,
relaxed and conversational pace all reinforce the defect you are removing; 7 register, emotional color
only, no pace words; 8 room tone matched to field 11, named plainly, since none reads as a dead booth.
End with: no reverb beyond the room, no music, no other voices.

Then a reuse rule: paste this paragraph into every future clip for this brand, changing ONLY the room
tone sentence when the room changes. Then, in these words: "Voice is not fully castable on this model.
This is the strongest mitigation there is, not a guarantee. If it comes back wrong, re-roll it."

--------------------------------------------------------
D. WRITE THE MOTION PROMPT
--------------------------------------------------------
Under 180 words excluding the voice paragraph, counted for real. An overloaded prompt garbles the
speech, so cut description before a lock or a ban.

1. SCENE STATE. One present tense sentence: what is happening, still or moving. It names no room,
   furniture, wardrobe, light, camera, face or product. A state, never an order.
2. SPEECH. Written as: She says, exactly: "..."
3. DELIVERY. One line, ENERGY ONLY, never pace. Low by default, near sixty percent, like she is
   telling one friend.
4. ONE action pinned to one exact spoken word, and it must coexist with the lock: it may MOVE the
   product as one unit, never RE-HANDLE it. Legal: "on the word finally she lifts the bottle toward
   her chin." Illegal: any re-grip. If the ad needs it handled, drop the lock, lock the label facing
   camera instead, and log the trade.
5. PRODUCT LOCK, only when the product is held and the action does not re-handle it. "She holds the
   [product] as a single fixed unit, hand and product locked together, never set down, never
   re-gripped, and her other hand never touches it." Carry its two bans into AVOID.
6. ONE camera instruction, in the same words as the start frame: locked off, the default, usually
   best and the only option for a propped phone; handheld at chest height with slight shake; one slow
   drift; or one small push in.
7. VOICE, then the block C paragraph pasted word for word.
8. SOUND. A cue in angle brackets only for an on-camera action in this clip, placed on that action,
   like <soft aerosol hiss>. A talking head where nobody uses the product gets none. Native sound is
   free, so never add effects later.
9. AVOID line, five to eight terms, worst artifact first, earlier terms carry more weight:
   warped or extra fingers, morphing hands, the product changing shape, a second person or her other
   hand entering frame, the mouth out of sync, a smooth gliding camera move. Banned camera words are
   legal here, they are bans. Never ban what your pinned action needs.
10. Both guard lines, exactly:
    NO TEXT, NO SUBTITLES, NO CAPTIONS, NO UI ON SCREEN.
    NO MUSIC, NO MELODY, NO INSTRUMENTS, NO SOUNDTRACK.
Music is banned twice on purpose, a music bed causes out of sync mouths. AVOID is the only negative
field there is. One continuous take whatever the length, no cuts written in.

--------------------------------------------------------
E. HOW TO RUN IT. ONE PATH, THE API.
--------------------------------------------------------
E0. COST, BEFORE THEY SPEND ANYTHING. Measured on our own account on 2026-08-07: 720p is 63 credits
per second of finished video, 480p is 28, credits cost $0.005 each, audio is free. Show the working
for THEIR length, cheap drafts and one keeper: "14 seconds is 392 credits at 480p, $1.96,
and 882 at 720p, $4.41, so two drafts and a keeper is about $8." Sign up at kie.ai, create an API key
and add credits, the call fails on a zero balance and your balance sits on that account page.
Paste the key over YOUR_API_KEY in both commands, keeping Bearer in front. On Windows type curl.exe,
never plain curl, which in PowerShell is a different program that rejects -H.

E1. THE CALL. first_frame_url is a direct public link to the STILL YOU GENERATED, not the field 15
product photo. Direct means the raw file, so pasted into a browser tab the image alone fills the
window, and a Dropbox or Drive share page fails. Get one by right clicking the still in your image
tool and choosing Copy image address, or, if that is not an address ending .png or .jpg, by uploading
the file to any free image host with a direct or raw link. Test it in a fresh tab before
spending a credit.

POST https://api.kie.ai/api/v1/jobs/createTask
Authorization: Bearer YOUR_API_KEY
Content-Type: application/json

{
  "model": "bytedance/seedance-2-5",
  "input": {
    "prompt": "<< THE MOTION PROMPT, JSON ESCAPED >>",
    "first_frame_url": "PUT_YOUR_IMAGE_LINK_HERE",
    "duration": << THE LENGTH BLOCK A SETTLED ON, A WHOLE NUMBER >>,
    "resolution": "720p",
    "aspect_ratio": "9:16",
    "generate_audio": true
  }
}

Save that body as a plain text file called body.json, on the Desktop. In Notepad use Save as, set Save
as type to All Files, and put the name in quotes, "body.json", or Windows adds a hidden .txt and the
command cannot find it. Type cd Desktop in the terminal first, then send it:
curl.exe -X POST "https://api.kie.ai/api/v1/jobs/createTask" -H "Authorization: Bearer YOUR_API_KEY" -H "Content-Type: application/json" -d @body.json
A file not found error is that hidden .txt or the wrong folder, not a bad key.

ESCAPING, NOT OPTIONAL. In the JSON body the motion prompt is ONE line, every double quote written \"
and every line break \n, or the request breaks. For a cheap draft change 720p to 480p.

E2. THE POLL. The POST hands back a taskId, a receipt and not the video. Poll every thirty seconds,
note a browser address bar cannot send the Authorization header:
curl.exe "https://api.kie.ai/api/v1/jobs/recordInfo?taskId=PASTE_YOUR_TASK_ID" -H "Authorization: Bearer YOUR_API_KEY"
Keep polling while "state", inside "data", is neither success nor fail. On "success" the link sits
inside "resultJson", itself JSON written as text, so it arrives like this, backslashes and all:
"state":"success","resultJson":"{\"resultUrls\":[\"<one long address ending .mp4>\"]}"
Take the address between those quotes, delete every backslash, download it immediately, links expire
in about a day. On "fail" read "failMsg": an indirect image link, a zero balance, a content refusal,
an illegal value. Renders run a few minutes. The job also lists itself in your kie.ai account, so if
the terminal is what stops you, watch and take the file there instead. A re-send is a second paid
render.

API NOTES, printed in section 10. Always pin duration, unset bills a default five seconds. Only 480p
and 720p are legal. Send aspect_ratio, feed a 9:16 frame, leave generate_audio on, it is free. No seed
and no negative prompt, so every re-roll is a fresh take.

--------------------------------------------------------
F. OUTPUT EXACTLY THIS SHAPE. ALL TEN SECTIONS, EVERY TIME.
--------------------------------------------------------
1 BEFORE YOU START, carrying the E0 arithmetic. 2 THE SCRIPT. 3 START FRAME IMAGE PROMPT, opening
"Generate this image first, before you do anything else", closing with the still checks and the E1
link test. 4 VOICE CONTRACT. 5 MOTION PROMPT, readable, printed alone: the prompt, then a line reading
END OF PROMPT, with any note under that line. Sections 5 and 6 carry the same words, so a sentence in
one and not the other is a bug. 6 HOW TO RUN IT, escaped. 7 WHAT I ASSUMED, each guess worded "you
left X blank, I assumed Y". All paste-ready.
8. CLAIM CHECK. A short table: every factual statement in the dialogue in one column, the field it
came from in the other. The field named must carry the claim AS SPOKEN, not the raw material. If a row
cannot name one, you broke Rule One: rewrite the line, not the row. A generated persona stating a
personal result is not fixable by a label, rewrite it. Then a final row with no field number,
"What this ad implies as a whole", with your honest read. If that read is an outcome the brief does not
back, you do not get to log it and ship it: rewrite until the implication is gone, and say in the row
what you cut. Close with: "Check these against what you can actually back up. Nobody has cleared this
ad for you."
9. DISCLOSURE. Only when a synthetic human is on camera. Open with "Do not run this ad without this
line burned in." Then the line and its placement, plus three sentences: this is an operational read,
not legal advice; "conspicuous" is undefined with no Attorney General guidance, so this placement is
conservative and not a safe harbour; whoever renders the file is the producer on the hook, and that is
you. Otherwise say no line is required and Meta may label the ad itself.
10. AFTER THE RENDER. Watch it once with sound on, checking the mouth against the words. A stumbled
opening word or a snapped gaze means trimming three to five tenths of a second, a dragging delivery
means speeding to about 106 percent. If speech ran past 8 seconds, cut a filler sentence, never the
field 8 proof, and re-roll at 8. Never re-roll unchanged: put the defect at the FRONT of AVOID, and if
it is visible in the still, fix the STILL, an image edit costs a fraction of a render. Draft at 480p,
spend 720p on the keeper. Audio and video are one file, so a stray music bed cannot be muted without
killing her voice. Then the API notes from block E.
Free. Yours. No attribution needed, though we will not stop you.
  1. Copy the block above and paste it into Claude or ChatGPT.
  2. Fill in the fifteen lines at the top. Guess where you are not sure. The builder is written to work off rough answers and to tell you what it assumed.
  3. Generate the start frame in whatever image tool you already use, from the image prompt it gives you. Look at it. This costs cents and it is where you catch problems.
  4. Run the API call it prints. A thirty second clip is about nine dollars and takes roughly four minutes.
What it will not do

It writes one ad, not a campaign. It cannot check whether your claims are legal in your category, so the brief asks what you are not allowed to say and it obeys that literally. And it will not invent proof. If you do not give it a number, it writes the line without one, which is the correct behaviour and occasionally a disappointing one.

01Two prompts, two jobs

This is the idea the builder is built around, and it is the single thing most people get wrong.

A shot is made from two separate prompts. An image prompt produces the start frame. A motion prompt animates it. The still owns what the shot looks like. The motion prompt owns what happens and what gets said.

ONE SHOT, TWO PROMPTS START FRAME What it LOOKS like Face. Product. Label. Room. Wardrobe. Light. MOTION PROMPT What HAPPENS The action. The line. The voice. One camera move. Describe the room twice and your text outranks the picture. That is drift.
The split most people skip. Anything the still already shows should not appear in the motion prompt.

Once you see it that way, the most common failure explains itself. People write one long prompt describing the room, the person, the wardrobe, the light, the product, and the action. Then they wire in a reference image that already shows most of that, and the two fight.

The text usually wins. Describe a bottle in words next to a photo of the bottle and the model tends to build a bottle from your words. That is why products swell, labels change, and the same person shows up looking slightly different in every clip.

02Authority, not length

The rule that took us longest to learn, and the one the builder enforces hardest.

Long prompts are not worse than short ones. Prompts full of sentences that do not have a job are worse. Every sentence should be doing exactly one of three things.

Name the one thing that changed

If a start frame or reference is wired in, the model already has the scene. Your job is the delta. "Change only her pose."

Lock something that would otherwise drift

Identity, product size, wardrobe, which hand holds what, the voice. Anything you leave unstated gets invented, and invented differently on every call.

Ban a specific artifact you have actually seen

Not a general plea for quality. The named defect you just watched happen. "Splayed pinky." "Never blank or stripped."

A sentence doing none of those is not neutral. It competes with your reference and usually beats it.

03The words that are costing you money

If you are making feed-native content, these push the model toward a glossy commercial look and away from anything that reads as real.

Strip theseUse these
cinematic, anamorphic, film grain, lens flare, dramatic lighting, color grade, LUT, bokeh, depth of field, slow motion, speed ramp, whip pan, crane, dolly, gimbal, Dutch angle, epic, stunning, camera brand names iPhone handheld, natural light, window light, slight camera shake, casual, authentic, deep focus front to back, visible JPEG compression, mild grain in shadows

The most reliable tell that a still was generated is fake shallow depth of field, the creamy blurred background a model adds because it assumes "photo" means "DSLR photo." Real phone footage is deep focus. Ask for it.

This list is scoped to the feed-native register. A polished commercial genuinely wants focal length, shallow depth of field, directed lighting and a colour grade. What causes trouble is crossing the two inside one piece. The two hype words, epic and stunning, do nothing in either.

04Approve a still before you buy a second of video

Highest-leverage habit we have, and it is free to copy.

A generated still costs cents. A video clip costs dollars. So every decision a frame can settle gets settled at the frame layer, before any video spend at all.

SETTLE IT AT THE CHEAP LAYER STILL ¢ cents A HUMAN LOOKS zero video spend before this point CLIP $ dollars
Identity, product and label are frame decisions. Making them at the clip layer means paying many times over for the same answer.

The start frame holds the static things. What it does not hold is hands, once motion starts. If something has to arrive or change state, give it a start frame and an end frame and let the action happen in between, rather than asking the model to invent where it ends up.

The failure this prevents

Identity drift is the most common reason a take gets rejected in our own work. Referencing a master image does not lock a subject; models re-imagine on every call. One of our builds ended up with three different women in the same ad, because the identity came from a text description instead of a real photograph. One identity source, and it should be a real image.

05Where the money actually goes

Everyone quotes the per-clip price. It is close to meaningless on its own, because it prices the clip you keep and ignores the ones you throw away.

ONE AUDITED RUN, EVERY GENERATION COUNTED 27% 42% 31% SHIPPED RE-ROLLS THAT BOUGHT QUALITY PURE WASTE Takes that made the cut. Beats that had not passed yet. Generated after a beat had already passed. Stop the pool the moment a take clears your bar and that last block is free money.
One of our production runs, every generation counted. The rightmost block bought nothing at all.

Two habits follow from that, and they will save you more than any per-second price comparison.

Stop generating the moment a take passes

Almost a third of that run's budget bought nothing, because the pool size was decided in advance instead of ending when it got what it came for. Set an accept bar. Hit it, stop.

Rounds cost more than generations

Every review round regenerates every touched clip. A creative is not "twelve clips of spend," it is twelve clips times however many rounds you go. Money spent getting the brief right before anything generates is the cheapest money in the process.

And one thing we would rather you did not believe

"AI is cheaper on pixels" does not hold up. In our experience a person firing the same generations by hand spends about the same on model calls, sometimes more, because they over-generate without a stopping rule. The saving is labour and parallel throughput. A business case built on the API bill is built on the wrong number.

06When it fails twice, stop rewriting the prompt

The instinct when a generation comes back wrong is to rewrite the prompt. That is right exactly once.

First failure, fix the spec. Second failure on the same thing, stop editing the prompt and draw several takes. If they all fail the same way, that is the diagnosis. Change the shape, the scene or the framing.

Three identical failures from three independent takes is not three bad rolls. It is the model telling you the thing you are asking for sits outside what it will do from that setup, and no amount of wording gets you there.

The related trap is more expensive. If the broken thing is an input rather than a prompt, do not re-roll it at full price. One of our builds burned real money re-rolling a broken phonetic spelling of a brand name at full clip cost, hoping for variance. A batch of cheap low-resolution throwaway probes, at a fraction of the price each, found the spelling the model said correctly in a single pass.

07The tells that read as AI

The tellThe fix
The frozen tail. The last half second holds a static frame, and in the edit it reads as a freeze before every cut. Trim roughly a third of a second off each clip's tail in assembly, and prompt for continuous motion to the very last frame.
Everything drifts smooth, glossy, plastic. Faces and reflective surfaces worst. Put the texture constant in the frame prompt, not the video prompt. The frame sets the look and the clip inherits it.
Flatness. Same framing, same energy, take after take. Author a distinct camera move and delivery per beat. A global "make it lively" cue does nothing.
The wrong product mechanism. A white cream on a spoon for something that is an amber syrup. Lock the physical facts like an identity. Getting the mechanism wrong reads as fake instantly, even to someone who cannot say why.
Garbled or invented label text. Fix it in generation, not in the editor. Product beats want a product reference in the start frame, near-zero motion on the product, and the product larger in frame so more label pixels survive.
A blank, stripped hero product. Ask for a "clean, uncluttered" shot and the model strips the packaging's own artwork. "Uncluttered" must describe the scene, not the product. Say the front art stays complete.
Mouths moving out of sync under a voiceover. The uncanny middle state, worse than either extreme. Pick one. Clean talking, or clean not-talking. Running voiceover, keep mouths still.
A logo the model drew. Slightly wrong, which is worse than absent. Generate the shot clean, composite the real file after.

08Frames are deaf and blind

The most expensive lesson in our archive, and it cost a full production run.

We reviewed a batch by pulling three still frames from each clip and running a twelve-agent adversarial review over the stills. It came back clean. Every clip passed.

Then someone watched the actual videos with the sound on, and every clip was a reject. Hand and microphone fused together throughout. Stiff delivery. Bad lip-sync. Looping backgrounds. None of which a still frame can show. The frame review had looked directly at the fused blobs and reasoned them away as background blur.

Cheap fast generation is worthless without disciplined verification. Watch and listen to every clip before you select it, assemble it, or claim anything about it.

The general form is worth more than the story. Every automated check is a proxy for something you actually care about, and a proxy can pass while the thing it stands for is false. Text detection approximates "the label is right." A pacing score approximates "this sounds natural." Each one needs a paired reality check, or you are trusting a number that quietly stopped meaning anything.

09What we measured, and why we mostly stayed put

Now the part that is about the model rather than the craft.

The claimWhat we found
30 second single clipReal. Zero hard cuts across the full thirty seconds, scene-detect run on the file.
Native audioReal. An AAC stereo track comes back on the same file, at no extra cost.
Native 4KNot reachable. The API accepts 480p and 720p. 1080p, 2160p and 4k are rejected.
50 reference inputsUnverified. In the launch material, not exposed on the path we used.
Vertical for MetaReal. 9:16 accepted, output is 720×1280 at 24fps.

Most comparisons you will read this week put 2.5 against Seedance 1.x. We run 2.0 in production, and 2.0 had its own upgrade the same week.

Seedance 2.0Seedance 2.5
Max resolution1080p and 4K720p
Max single render15 seconds30 seconds
Price per secondbaseline+53.7%

Against the model we already run, 2.5 is a resolution downgrade at a 54% price increase. Its real advantage is a single render longer than fifteen seconds.

So we adopted it as a routed exception rather than a default. A shot moves to 2.5 when it needs something 2.0 cannot do. Everything else stayed where it was.

The arithmetic worth stealing

Cost per usable shot is the rate times the shot length, divided by your hit rate. Render length does not appear in it. One thirty second render and two fifteen second renders cost the same and produce the same number of shot candidates. If a failure turns out to be render-wide rather than shot-local, the long render loses everything at once. So on the numbers we modelled, a longer render is cost-neutral at best, and worse whenever a whole-render failure is possible. We have not measured how often failures are render-wide rather than shot-local, and that is the number that would settle it.

LengthCostRender time
5 seconds$1.5887 seconds
30 seconds$9.45233 to 271 seconds

10The part the model does not solve

The floor for what looks good moved up for everyone at once this week. That is good news, and it also means looking good stopped being much of a differentiator on the same day.

Which awareness stage you are writing to

Most brands write every ad to the person already shopping, who is the smallest slice of the market. The harder question is what you say to somebody who has never thought about the problem.

Whether the mechanism is true

"It works" is a claim. Why it works is an argument, and an argument is what survives a skeptical scroll. If you cannot explain the mechanism in one sentence, the ad is decoration.

Hold rate past three seconds

Not the thumbstop. Plenty of ads stop the scroll and then have nothing to say. Find where the curve falls off and fix that beat.

How many angles you can run before the account fatigues

Volume is a creative-strategy problem before it is a production problem. A cheaper model raises your ceiling on volume, and volume with no angle diversity is the same ad four times.

None of that got easier this week. It got more valuable.

Where these numbers come from

The prices, render times, resolutions and the 30 second continuity check were measured on our own account on August 7, 2026, across four generations. The 2.0 comparison is measured on the same account: 2.5 bills 63 credits per second at 720p where 2.0 bills 41, which is the 53.7% step. The 27 / 42 / 31 spend split is from one internal production run where every generation was logged and categorised afterwards; it is one run, not a benchmark. Everything else in here is craft learned across our own builds, which means it is our experience rather than a study.

If you would rather we ran it

We do AI performance creative for a handful of leading DTC brands on Meta.

Prove It is $997 for three videos and five statics, credited toward any package if you keep going.

Book 30 minutes with me and we will walk your account.

Not ready for that

Send me the one line your ad actually says, at leo@whim.live, and I will tell you whether the mechanism in it is doing any work. Takes me a minute and costs you nothing.