AI UGC · Prompts · Storyboards · Scripts

AI UGC PROMPTS AND STORYBOARDS THAT ACTUALLY CONVERT

The shot list, the prompt templates and the mistakes — written so you can run this yourself today. If you get to the end and decide you'd rather not spend the hours, that's what we do.

Written by Aman Rai · ₹50Cr+ / US$6M+ ad spend managed · 280+ brands since 2020 · Last updated 26 August 2026

The short version. An AI UGC ad that holds attention is a six-shot storyboard, not one long prompt. Shot one is the hook in the first second, shot two is the problem, shot three is the product entering the frame, shot four is the demonstration, shot five is the proof or reaction, shot six is the ask. Each shot gets its own prompt, and each prompt has to specify the imperfection — handheld camera, ordinary room, product already opened — because clean studio lighting is what makes synthetic footage read as an advert and get scrolled past. Expect two to five attempts per shot before the framing, the hands and the product label all survive together. One finished thirty-second ad is realistically twenty to forty generations. That number's the whole story of why most brands try this and then quietly stop.

What you will get from this page

  • The six-shot storyboard, with what each shot has to accomplish.
  • Copy-ready prompt templates for every shot — swap the bracketed parts for your product.
  • The image-to-video step that stops the model inventing your packaging.
  • Five failure patterns that make AI UGC look fake, and the fix for each.
  • An honest count of the hours, so you can decide whether to run it in-house.

WHY MOST AI UGC GETS SCROLLED PAST

The model's rarely the problem. Kling, Higgsfield, Runway, Google Veo and Sora all produce footage good enough for the feed now. What kills the ad is that people write one prompt describing a beautiful person holding a product in soft light, and the feed treats that exactly like what it is — an advert.

Real creator video doesn't look like that. It's shot at arm's length, in a kitchen with the wrong colour temperature, and it starts mid-sentence because the person already hit record. Meta's own ad creative guidance and the TikTok Creative Center both keep pointing at the same thing: native-looking, sound-on, front-loaded video beats polished film in an in-feed placement.

You're not prompting for a beautiful video. You're prompting for a video that doesn't look like it was prompted.

THE SIX-SHOT STORYBOARD EVERY UGC AD NEEDS

Write the shot list before you open any tool. This is the part people skip, and it's the part that decides whether the ad works.

#ShotWhat it has to doLength
1The hookEarn the next second. A face mid-sentence, a mess, a surprising object — never a logo, never a slow pan.0–2s
2The problemName the specific irritation your buyer already feels. Specific beats dramatic.2–6s
3Product entersThe product appears in an ordinary way — pulled out of a bag, already on the counter. Not presented.6–10s
4The demonstrationThe one thing it does, shown, not described. Hands in frame.10–20s
5Proof or reactionA result, a face reacting, a before-and-after that stays within the platform's rules.20–26s
6The askOne instruction. Spoken, not just captioned.26–30s

Six shots, one prompt each. If a shot can't say what it's for in one sentence, cut it.

PROMPTS YOU CAN COPY AND USE

Replace anything in [brackets]. These are written for image-to-video and text-to-video models like Kling, Higgsfield, Runway and Veo — the structure matters more than which one you use.

Shot 1 — the hookVertical 9:16 selfie video, [woman in her late twenties] holding the phone at arm's length, already mid-sentence, slightly off-centre framing. [Ordinary kitchen] at [7am], overhead light only, mild colour cast. Natural skin texture, no retouching. Slight handheld shake. She looks straight into the lens with an irritated expression. Shallow depth, phone-camera look, not cinematic.
Shot 3 — product enters (use image-to-video)Animate this frame: the hand picks up [the product] from the counter and turns it once toward the lens, label facing camera for about a second, then lowers it out of frame. Keep the packaging exactly as shown in the reference image. Handheld, minor wobble, same lighting as the frame.
Shot 4 — the demonstrationClose handheld shot of hands [applying / opening / assembling] [the product] on [a bathroom counter]. Camera dips slightly as she adjusts her grip. Real fingernails, visible skin texture, one small imperfection in the background such as [a folded towel out of place]. No music-video polish, no slow motion, no lens flare.
Shot 6 — the askSame [woman] back at arm's length, softer expression now, mid-sentence again. She says one short line and shrugs slightly at the end. Frame slightly tighter than shot 1 so the cut reads as a jump cut, not a new scene. Same room, same light.
The line that fixes most of it. Add “phone-camera look, natural skin texture, slight handheld shake, no cinematic lighting” to every prompt. On its own that single instruction moves more footage from “advert” to “creator” than any other change we make.

STOP THE MODEL INVENTING YOUR PACKAGING

Text-to-video will get your bottle shape roughly right and your label completely wrong, and a wrong label is worse than showing no product at all. The fix isn't a better prompt.

  1. Get one clean, well-lit frame of the real product — a phone photo on a plain surface is enough.
  2. Feed that frame in as the reference image and animate it. Kling and Higgsfield both take a reference frame for exactly this.
  3. Keep the product shot short. Two seconds of accurate packaging beats eight seconds of a hallucinated version.
  4. For the talking shots, the product does not need to be in frame at all — that's where text-to-video is safe.

FIVE THINGS THAT MAKE IT LOOK FAKE

What you seeWhy it happensFix
Reads as an advert, not a personPrompt asked for studio lighting and a beautiful subjectAsk for the room, the hour and the imperfection explicitly
Hands go wrong on the productToo much hand movement in one generationShorter clips, one hand action per shot, image-to-video
Label is not your labelText-to-video invents packagingReference frame of the real product
Voice does not match the faceGeneric text-to-speech over a specific faceCast the voice to the face; keep lines short and conversational
Everything blends togetherAll six shots generated with the same prompt skeletonChange room, distance and energy between shots so cuts feel real

THE PART THE TUTORIALS LEAVE OUT

Here is the honest arithmetic, from doing this every week.

One finished thirty-second ad is six shots. Each shot takes two to five generations before the framing, the hands and the label all survive at once. That is twenty to forty generations per ad, before you have touched the audio, the captions, the aspect-ratio cutdowns or the file naming that keeps your reporting readable a month later.

And one ad is not the format. The economics of AI UGC only work at volume — twenty or thirty genuinely different hooks in the auction, letting the platform find the two that work. Testing three variants and concluding the format does not work is a sample-size problem, not a creative one. It's the single most common mistake we see.

So the real question isn't whether you can do this. You can — everything above is the actual method, not a teaser. The question is whether twenty to forty generations an ad, thirty ads a month, is where your hours should go.

RUN IT YOURSELF, OR HAND IT OVER

Keep it in-house if

You have someone whose actual job is creative, you are testing fewer than ten variants a month, and you enjoy the iteration. The method above is complete. Nothing's held back.

Hand it over if

You need twenty to forty ad-ready variants a month, you are already spending on Meta or TikTok, and the hours are currently coming out of your own week. That is the point where this stops being a skill and starts being a production line.

If it's the second one, we run it inside your ad account, on your naming convention, with the hooks rewritten each month from what actually won. Scoped on a call — and if we're not the right fit we'll say so.

QUESTIONS PEOPLE ASK ABOUT AI UGC PROMPTS

How many prompts does one AI UGC ad actually need?

More than people expect. A single finished thirty-second ad is usually six to eight shots, and each shot takes two to five prompt attempts before the framing, the hands and the product label all survive at once. So one ad is realistically twenty to forty generations. That is the number nobody puts in the tutorials, and it is the main reason brands start with AI UGC and quietly stop after three weeks.

Why does my AI UGC look fake even when the model looks real?

Almost always because the shot list is wrong, not because the render is bad. Real creator footage is handheld, badly lit in a believable way, and cuts mid-sentence. A prompt that asks for a beautiful person in soft studio light gives you a beautiful person in soft studio light — which reads as an advert, and the feed scrolls past adverts. Ask for the imperfection explicitly: slight camera shake, a kitchen at 7am, the product already opened and used.

Can AI UGC show my actual product accurately?

Text-to-video alone will invent your packaging. The reliable route is image-to-video: generate or shoot a clean frame that contains your real product, then animate that frame. Tools like Kling and Higgsfield accept a reference image for exactly this. Any prompt-only workflow will get the bottle shape close and the label wrong, and a wrong label is worse than no product shot at all.

Do AI UGC ads get restricted by Meta or TikTok?

Not for being AI-generated in itself. Restrictions come from the usual places — before-and-after claims, health and body promises, implied endorsements, and text overlays that make claims the landing page does not support. Write the script to the same rules a human creator would follow and the render being synthetic is not the issue.

How many variants should I test before judging AI UGC?

Do not judge on one. The whole point of this format is volume — the economics only work when you can put twenty or thirty genuinely different hooks into the auction and let it pick. Testing three variants and concluding AI UGC does not work is the most common mistake we see, and it is a sample-size problem rather than a creative one.

What do you actually do that I cannot do myself?

You can do all of it. The honest question is whether you want to spend the hours. The work is the shot list, the twenty to forty generations per ad, the audio pass, the captioning, the naming convention so the reporting stays readable, and then reading the results and rewriting next month's hooks from what won. If that sounds like your best use of a month, keep it in-house — we'd tell you the same on a call.

WANT THIS RUN FOR YOU?

Send me what you sell and roughly what you are spending. I will tell you whether AI UGC is the right format for it — including when it is not.

Message Me on WhatsApp Prefer not to message? Send your details instead →

SEND YOUR DETAILS

These come straight to me. I reply myself, usually the same working day — WhatsApp is faster if you are in a hurry.