The shot list, the prompt templates and the mistakes — written so you can run this yourself today. If you get to the end and decide you'd rather not spend the hours, that's what we do.
Written by Aman Rai · ₹50Cr+ / US$6M+ ad spend managed · 280+ brands since 2020 · Last updated 26 August 2026
The model's rarely the problem. Kling, Higgsfield, Runway, Google Veo and Sora all produce footage good enough for the feed now. What kills the ad is that people write one prompt describing a beautiful person holding a product in soft light, and the feed treats that exactly like what it is — an advert.
Real creator video doesn't look like that. It's shot at arm's length, in a kitchen with the wrong colour temperature, and it starts mid-sentence because the person already hit record. Meta's own ad creative guidance and the TikTok Creative Center both keep pointing at the same thing: native-looking, sound-on, front-loaded video beats polished film in an in-feed placement.
You're not prompting for a beautiful video. You're prompting for a video that doesn't look like it was prompted.
Write the shot list before you open any tool. This is the part people skip, and it's the part that decides whether the ad works.
| # | Shot | What it has to do | Length |
|---|---|---|---|
| 1 | The hook | Earn the next second. A face mid-sentence, a mess, a surprising object — never a logo, never a slow pan. | 0–2s |
| 2 | The problem | Name the specific irritation your buyer already feels. Specific beats dramatic. | 2–6s |
| 3 | Product enters | The product appears in an ordinary way — pulled out of a bag, already on the counter. Not presented. | 6–10s |
| 4 | The demonstration | The one thing it does, shown, not described. Hands in frame. | 10–20s |
| 5 | Proof or reaction | A result, a face reacting, a before-and-after that stays within the platform's rules. | 20–26s |
| 6 | The ask | One instruction. Spoken, not just captioned. | 26–30s |
Six shots, one prompt each. If a shot can't say what it's for in one sentence, cut it.
Replace anything in [brackets]. These are written for image-to-video and text-to-video models like Kling, Higgsfield, Runway and Veo — the structure matters more than which one you use.
Text-to-video will get your bottle shape roughly right and your label completely wrong, and a wrong label is worse than showing no product at all. The fix isn't a better prompt.
| What you see | Why it happens | Fix |
|---|---|---|
| Reads as an advert, not a person | Prompt asked for studio lighting and a beautiful subject | Ask for the room, the hour and the imperfection explicitly |
| Hands go wrong on the product | Too much hand movement in one generation | Shorter clips, one hand action per shot, image-to-video |
| Label is not your label | Text-to-video invents packaging | Reference frame of the real product |
| Voice does not match the face | Generic text-to-speech over a specific face | Cast the voice to the face; keep lines short and conversational |
| Everything blends together | All six shots generated with the same prompt skeleton | Change room, distance and energy between shots so cuts feel real |
Here is the honest arithmetic, from doing this every week.
One finished thirty-second ad is six shots. Each shot takes two to five generations before the framing, the hands and the label all survive at once. That is twenty to forty generations per ad, before you have touched the audio, the captions, the aspect-ratio cutdowns or the file naming that keeps your reporting readable a month later.
And one ad is not the format. The economics of AI UGC only work at volume — twenty or thirty genuinely different hooks in the auction, letting the platform find the two that work. Testing three variants and concluding the format does not work is a sample-size problem, not a creative one. It's the single most common mistake we see.
So the real question isn't whether you can do this. You can — everything above is the actual method, not a teaser. The question is whether twenty to forty generations an ad, thirty ads a month, is where your hours should go.
You have someone whose actual job is creative, you are testing fewer than ten variants a month, and you enjoy the iteration. The method above is complete. Nothing's held back.
You need twenty to forty ad-ready variants a month, you are already spending on Meta or TikTok, and the hours are currently coming out of your own week. That is the point where this stops being a skill and starts being a production line.
If it's the second one, we run it inside your ad account, on your naming convention, with the hooks rewritten each month from what actually won. Scoped on a call — and if we're not the right fit we'll say so.
More than people expect. A single finished thirty-second ad is usually six to eight shots, and each shot takes two to five prompt attempts before the framing, the hands and the product label all survive at once. So one ad is realistically twenty to forty generations. That is the number nobody puts in the tutorials, and it is the main reason brands start with AI UGC and quietly stop after three weeks.
Almost always because the shot list is wrong, not because the render is bad. Real creator footage is handheld, badly lit in a believable way, and cuts mid-sentence. A prompt that asks for a beautiful person in soft studio light gives you a beautiful person in soft studio light — which reads as an advert, and the feed scrolls past adverts. Ask for the imperfection explicitly: slight camera shake, a kitchen at 7am, the product already opened and used.
Text-to-video alone will invent your packaging. The reliable route is image-to-video: generate or shoot a clean frame that contains your real product, then animate that frame. Tools like Kling and Higgsfield accept a reference image for exactly this. Any prompt-only workflow will get the bottle shape close and the label wrong, and a wrong label is worse than no product shot at all.
Not for being AI-generated in itself. Restrictions come from the usual places — before-and-after claims, health and body promises, implied endorsements, and text overlays that make claims the landing page does not support. Write the script to the same rules a human creator would follow and the render being synthetic is not the issue.
Do not judge on one. The whole point of this format is volume — the economics only work when you can put twenty or thirty genuinely different hooks into the auction and let it pick. Testing three variants and concluding AI UGC does not work is the most common mistake we see, and it is a sample-size problem rather than a creative one.
You can do all of it. The honest question is whether you want to spend the hours. The work is the shot list, the twenty to forty generations per ad, the audio pass, the captioning, the naming convention so the reporting stays readable, and then reading the results and rewriting next month's hooks from what won. If that sounds like your best use of a month, keep it in-house — we'd tell you the same on a call.
Send me what you sell and roughly what you are spending. I will tell you whether AI UGC is the right format for it — including when it is not.
Message Me on WhatsApp Prefer not to message? Send your details instead →These come straight to me. I reply myself, usually the same working day — WhatsApp is faster if you are in a hurry.