TL;DR: UGC-style AI video works by looking unproduced, so the prompt has to ask for the opposite of a cinematic one: handheld instead of stabilized, available light instead of a lighting setup, a phone lens instead of a cinema lens, a real cluttered room instead of a styled set. Below: the inversion principle, the honest limit on fabricated testimonials, and 27 copy-paste prompts across 8 UGC formats plus B-roll.
What makes a UGC-style AI video prompt different from a cinematic one?
A UGC-style prompt asks for the opposite of everything a cinematic prompt asks for. Cinematic prompting spends its whole vocabulary on control: a named camera move, a matched lens, a lighting setup, a graded look. UGC-style prompting spends the same vocabulary undoing that control on purpose, because the entire reason a UGC clip reads as trustworthy is that it looks like nobody was steering it.
That's the inversion this whole post is built on. Every prompt below names the unproduced version of a cinematic default, deliberately:
| Cinematic element | A cinematic prompt asks for | A UGC-style prompt asks for instead |
|---|---|---|
| Camera | A locked-off tripod shot or a smooth gimbal move | Handheld sway, a small correction mid-shot, a phone propped on a shelf |
| Lens | A cinema lens with shallow depth of field | A phone-style wide lens, mostly in focus, slight distortion at the frame edge |
| Lighting | A three-point setup: key, fill, rim | One available light source: window light, an overhead kitchen light, a ring light left slightly off-axis |
| Framing | Rule-of-thirds, centered subject, clean headroom | Off-center, a forehead or chin cropped by the frame, the phone held a little too close |
| Environment | A styled set with a clean background | A real room: an unmade bed, a charger cable on the floor, a half-full mug |
| Delivery | A scripted line, one clean take | A mid-sentence restart, a short laugh, a trailed-off word |
| Sound | A mixed track with music and foley | Room tone, an appliance hum, one ambient noise, phone-mic compression |
| Color | A graded, filmic look | Flat, slightly warm or blue-tinted phone-camera white balance |
Every published prompting guide for these models, ours included, teaches the left column. Nobody's vendor documentation teaches the right column, because the right column isn't a model capability, it's a choice you make about what to ask for. That's the whole craft skill this post is trying to hand you.
Why does asking for "cinematic" work against you here?
Because the word does real work in the model, just work you don't want. Current video models were trained on enormous amounts of captioned, professionally shot footage, so a word like cinematic, or an unmodified request for a "product ad," pulls in a dense cluster of associations: stabilized movement, graded color, deliberate lighting. That's a feature when you want a commercial. It's the exact thing to fight when you want a clip that reads as somebody's unscripted phone video.
The fix isn't deleting the word cinematic and hoping the model lands somewhere rawer on its own. It's naming the opposite explicitly, in every field a cinematic prompt would normally control: say handheld instead of leaving camera behavior unstated, name one available light source instead of leaving lighting to the model's default, ask for a phone-style lens instead of leaving focal length open. An underspecified prompt doesn't default to native. Left alone, most of today's models default toward the more produced look, because that's what the bulk of their training footage looked like.
Sound behaves the same way, and it differs by vendor in a way worth knowing before you build a UGC clip around it. Veo 3.1's audio is on by default, generating synchronized sound and dialogue whether you describe it or not, so a raw UGC feel there means actively asking for restrained room tone instead of a full mix. Other current models ship audio off by default, so getting any ambient sound at all means turning audio on and describing it yourself. Don't assume one vendor's audio behavior transfers to another; check the model you're actually generating on.
Is it OK to generate an AI "customer" saying your product worked?
No, and this is the part of this post that actually matters. Everything above is a craft technique. This is an ethics point, and it doesn't get softer because the video looks convincing.
A UGC ad's whole persuasive power comes from the viewer believing a real customer made it. That belief is the product. Generated UGC borrows that trust without earning it, because the viewer is reacting to what they think is a stranger's honest experience, not to a piece of creative direction. Inventing a customer who says your product worked is a fabricated endorsement, and that's a different act from making a stylized ad, even when the two look similar on screen.
This isn't a hypothetical concern. In the US, the FTC finalized a rule in 2024 addressing exactly this, and its own language names the technology directly: the rule addresses reviews and testimonials that "misrepresent that they are by someone who does not exist, such as AI-generated fake reviews, or who did not have actual experience with the business or its products or services". A generated speaker who never used your product is exactly that: someone with no actual experience with the business or its products, dressed up as if they did. The FTC's separate endorsement guides state the underlying principle plainly: "endorsements must be honest and not misleading", and that "An endorsement must reflect the honest opinion of the endorser". A generated speaker with no real experience of your product fails that test by definition, regardless of how the clip was made.
Platform policy adds a second, separate layer on top of the regulatory one, and it's worth checking directly rather than assuming a rule. YouTube's own policy on disclosing altered or synthetic content requires creators to disclose content that:
- "Makes a real person appear to say or do something they didn’t do."
- "Alters footage of a real event or place."
- "Generates a realistic scene that didn’t actually occur."
Its own list of examples requiring disclosure includes, verbatim, "Making it appear as if someone gave advice that they did not actually give". A fabricated customer testimonial is that example, just with a product name attached. TikTok's newsroom describes its own policy as one that "requires people to label AI-generated content that contains realistic images, audio or video". Meta applies broader "AI info" labels across its apps and, separately, its own account of the policy notes that "since January, advertisers have to disclose when they digitally create or alter a political or social issue ad in certain cases." That's a narrower rule, scoped to political and social-issue advertising rather than every ad. None of these three platform policies are identical, and they keep changing, so read the current version for wherever you're publishing rather than trusting a blog post, this one included, to have the latest wording.
How do you turn the inversion into a reusable prompt?
Build every prompt on the same seven-part shape a cinematic prompt would use (shot, subject, action, setting, light, look, sound, covered in more depth in our breakdown of video prompt structure), but fill each slot with its unproduced answer instead of its produced one. Shot: handheld or propped-phone, not a named camera move. Light: one real source, not a lighting design. Setting: a specific, slightly cluttered real room, not a described set. Sound: one ambient detail, not a mix.
Every prompt below follows that shape and uses square-bracket placeholders you swap for your own details: [PRODUCT], [ROOM], [COMPLAINT], [FEATURE], [TIME OF DAY]. Replacing every placeholder is where the value is; a prompt that still says [PRODUCT] produces a generic result the same way a cinematic prompt with an unfilled subject does.
Talking-head testimonial prompts (3)
These are for pre-vis and internal review only, never for a claim you publish as a customer's word. Write the person generically (a role or an age range, not a name or a fabricated backstory) and keep the language a monologue, not sworn testimony.
1. Kitchen counter monologue
Handheld selfie-style shot, phone held at arm's length, slight tremor and one small
reframe mid-take. A person in their late 20s to 30s stands in a real kitchen, [ROOM
DETAIL: dish rack, half-empty mug, sticky note on the fridge] visible behind them.
One available light source: overhead kitchen light, slightly yellow-toned, no fill.
They talk directly to the lens about [PRODUCT] and [FEATURE], pause mid-sentence to
find the next word, and let out a short laugh partway through. Frame is off-center,
their forehead brushing the top edge. Audio: their voice close and slightly clipped,
a fridge hum in the background, no music.
2. Bathroom mirror monologue
Phone propped against a bathroom mirror, slightly crooked angle, static but not
locked-off. A person mid-morning routine talks to their reflection about [PRODUCT]
while [ACTION: towel-drying hair, applying moisturizer]. Available light only:
bathroom vanity bulbs, a little harsh, faint reflection glare on the mirror edge.
Delivery is casual and unscripted, trailing off once, restarting the sentence.
Framing crops the top of their head. Audio: bathroom fan running low, water
dripping once, voice slightly echoey off the tile.
3. Parked car monologue
Phone mounted on a dashboard clip, slight handheld jitter from engine idle. A
person sits in the driver's seat, car parked, and talks about [PRODUCT] and why
they'd recommend it for [USE CASE], glancing at the road once out of habit. Light
is whatever's available: overcast daylight through the windshield, uneven across
their face. Framing is tight and a little too close, one ear cut off by the frame
edge. Audio: engine idle hum, a turn signal ticking once, no score.
Unboxing prompts (3)
1. Coffee table unboxing
Overhead-ish handheld angle, phone held above a coffee table cluttered with
[CLUTTER: a remote, a coaster ring, a half-drunk glass]. Hands open a [PRODUCT]
box, slightly fumbling the tape, natural pace, no narration script beyond a quiet
"okay, let's see." Available light: a window to one side, casting a soft uneven
falloff across the table. Framing is imperfect, one corner of the box outside
frame. Audio: cardboard tearing, a couch creak, ambient room tone, no music bed.
2. Doorstep unboxing
Handheld vertical shot, phone held at chest height just after bringing a package
inside. Person sets [PRODUCT] box on a hallway floor or entry table still wearing
outdoor shoes, opens it standing up, slightly unsteady framing as they crouch.
Available light: whatever daylight comes through the open door, mixed with dim
interior light, visibly uneven color temperature. Voice has a slight breathlessness
from just walking in. Audio: door latch, keys set down, street sound briefly audible
before the door closes.
3. Bedroom floor unboxing
Phone held low, handheld, person sitting cross-legged on a bedroom floor with
[PRODUCT] box between their knees, a laundry basket and phone charger visible in
frame. Single available light: a bedside lamp, warm and dim, leaving one side of
their face underlit. Delivery is quiet, half to themselves, half to the camera,
with one genuine pause of surprise at what's inside. Audio: fabric rustle, a
notification buzz once in the background, no added sound.
Problem/solution demo prompts (3)
1. Kitchen frustration to fix
Handheld, slightly shaky start: person visibly frustrated with [PROBLEM: a tangled
cord, a stuck jar lid, a cluttered drawer] in a real kitchen. Cut only implied by a
head turn, not a clean edit. They reach for [PRODUCT], demonstrate [FEATURE] solving
the problem in one continuous unscripted take, react with visible relief rather
than scripted enthusiasm. Available light: overhead kitchen fixture only. Framing
loose and off-center throughout. Audio: the actual sound of the problem (rattling,
scraping) followed by the actual sound of the fix, ambient kitchen noise underneath.
2. Desk-clutter fix
Phone propped on a stack of books at desk height, static but slightly tilted.
Person sits at a real desk with [CLUTTER: tangled cables, sticky notes, a cold cup
of coffee], gestures at the mess with visible annoyance, then uses [PRODUCT] to fix
[SPECIFIC PROBLEM]. Delivery includes one filler word and one unscripted aside.
Available light: laptop screen glow plus one desk lamp, mismatched color
temperature. Audio: keyboard clatter faintly in the background, a chair creak,
no music.
3. Skincare complaint to result
Handheld, phone held close to the face at bathroom-mirror distance. Person points
out [SKIN COMPLAINT] under harsh available bathroom light, no diffusion, no
retouched look. They apply [PRODUCT], narrate what it feels like in plain,
unpolished language, and note they'll show the result later rather than claiming
an instant fix on camera. Framing crops slightly at the chin. Audio: bathroom fan,
a bottle cap click, voice close to the mic with slight breathiness.
Before-and-after prompts (3)
1. Same-room split take
Two handheld clips implied as one continuous scene: the same real room shown in
[BEFORE STATE], then again in [AFTER STATE] after using [PRODUCT], same static
framing both times so the room itself is the comparison rather than a produced
side-by-side. Available light held consistent between the two: whatever natural
light the room actually gets at that hour. Voice-over is plain, not read from a
script: "okay, here's before... and here's after." Audio: ambient room tone
throughout, no transition sound effect.
2. Continuous-take transformation
One unbroken handheld take, no visible cut, following [SUBJECT: a countertop, a
closet, a face] from a messy or [BEFORE STATE] condition to a tidied or [AFTER
STATE] condition using [PRODUCT] in real time, not sped up. Camera drifts and
resettles at least once mid-take. Available light: single overhead or window
source, unchanged throughout. Audio: the real sounds of the process (wiping,
folding, application), ambient room tone, no music swell at the reveal.
3. Held-up-to-camera comparison
Handheld, phone held at arm's length in a bathroom or bedroom. Person holds up
[BEFORE OBJECT/RESULT] to the lens, describes it plainly, then holds up [AFTER
OBJECT/RESULT] next to it for a direct side-by-side, hand slightly unsteady.
Available light only, no ring light diffusion implied. Framing is close and a
little off-level. Audio: their voice close to the mic, a slight rustle of fabric
or paper, no added sound design.
Get-ready-with-me (GRWM) style prompts (3)
1. Bathroom mirror GRWM
Phone propped against the mirror, slightly crooked, static. Person moves through
a real morning routine ([ACTION: brushing hair, applying skincare, getting
dressed]), talking to the camera in fragments between steps rather than one
continuous script, mentioning [PRODUCT] once mid-routine rather than as the whole
video's subject. Available light: bathroom vanity bulbs only. Framing crops at
the shoulders, sometimes losing the top of the head as they lean in. Audio: water
running briefly, a hairdryer clicking on and off, voice trailing during tasks.
2. Closet GRWM
Handheld, phone held at chest height, person moving between a closet and a mirror
in a real bedroom, clothes visible on a chair in the background. They narrate
outfit choices casually, folding in a mention of [PRODUCT] as one small part of
the routine, not the center of it. Available light: bedroom window plus one lamp,
mixed color temperature. Delivery includes one "wait, actually" self-correction.
Audio: hangers clicking, fabric rustle, ambient room tone.
3. Car GRWM before leaving
Phone mounted on a dashboard clip or held at arm's length, parked car, engine off
or idling. Person finishes a last step of getting ready ([ACTION: applying
lip balm, fixing hair in the rearview mirror]) using [PRODUCT], talking to the
camera in short bursts between glancing at themselves. Available light: daylight
through the windshield, uneven and directional. Audio: keys jingling, a car door
sound once, no music.
Hands-only product demo prompts (3)
1. Kitchen counter hands-only
Overhead-ish handheld shot, phone held above a real kitchen counter with
[CLUTTER: crumbs, a dish towel, a coffee ring] visible. Only hands and [PRODUCT]
in frame, no face. Hands demonstrate [FEATURE] at a natural, slightly imperfect
pace, one small fumble not edited out. Available light: overhead kitchen fixture,
uneven shadow from the hands themselves. Audio: the actual sound of the product
in use, a countertop tap, ambient kitchen noise, no voice-over.
2. Desk hands-only
Phone propped at a low angle on a real desk, static. Hands only, unboxing or
demonstrating [PRODUCT] on [FEATURE] next to a keyboard and a half-empty water
glass. Available light: laptop glow plus a desk lamp, visibly mismatched
temperature. Pace is unhurried and includes a brief pause to reposition the
object. Audio: a faint keyboard click from off-frame, the product's actual sound,
no music.
3. Bathroom counter hands-only
Handheld, phone held close over a real bathroom counter with [CLUTTER: a toothbrush
cup, a stray bobby pin] visible at the frame edge. Hands demonstrate [PRODUCT] and
[FEATURE] up close, natural unsteady framing, no face in shot. Available light:
vanity bulbs only, slightly warm. Audio: water running briefly, a cap click,
ambient bathroom tone, no voice-over.
Reaction prompts (3)
1. Unboxing reaction
Handheld, phone held at arm's length or propped nearby. Person opens [PRODUCT]
box in a real living room, [CLUTTER: a couch cushion out of place, a remote on
the floor] visible. Reaction is a genuine-seeming pause before speaking, not an
exaggerated gasp, followed by an unscripted, slightly awkward first sentence.
Available light: whatever the room actually has at that hour, possibly a lamp
plus daylight. Framing loose, off-center. Audio: cardboard sound, a quiet "oh
wow," ambient room tone, no music sting.
2. First-time-trying reaction
Phone propped on a shelf or held by a second, unseen hand. Person tries [PRODUCT]
for the first time on camera, starting mildly skeptical ("I don't know if this is
going to work") and shifting to genuine surprise partway through, without a
scripted turn. Available light: single real source only. Delivery includes at
least one pause where they're clearly thinking, not reciting. Audio: the actual
sound of the product working, their unscripted reaction, ambient room tone.
3. Two-person reaction
Handheld, phone held by one person filming a second person trying [PRODUCT] for
the first time in a real room ([ROOM DETAIL]). Slight camera drift as the filmer
reacts too. Dialogue is casual and overlapping in places, not alternating cleanly
like a script. Available light: whatever the room has, uneven between the two
people. Framing occasionally clips one person's shoulder. Audio: both voices,
ambient room tone, one laugh, no music.
Day-in-the-life prompts (3)
1. Morning routine montage
A sequence of short handheld clips, each in a real room of the same home (kitchen,
bathroom, entryway), following a morning routine where [PRODUCT] appears once,
naturally, not as the focus of every clip. Each clip uses only the available
light of that room at that hour, so lighting shifts clip to clip rather than
staying matched. Framing is loose and slightly different each time, like separate
grabbed moments rather than a planned sequence. Audio: each room's own ambient
sound, no continuous music bed across clips.
2. Work-from-home day
Handheld or propped-phone clips through a work-from-home day: desk in the morning,
kitchen at lunch, couch in the evening, with [PRODUCT] used once at [POINT IN DAY].
Each setting keeps its own real clutter (cables, dishes, a blanket) and its own
available light rather than a matched look across the day. Delivery, where there
is any, is a quick aside to camera, not a narrated recap. Audio: the ambient sound
specific to each setting, a keyboard, a microwave, a TV low in the background.
3. Errand-day sequence
Short handheld clips across a day running errands: getting in the car, inside a
store, back home, with [PRODUCT] appearing once at [POINT: in the car, at home
afterward]. Available light shifts naturally from daylight to indoor light across
the clips, unmatched on purpose. Framing is quick and a little rough, like clips
grabbed between tasks. Audio: car engine, store ambience, front door, each kept
distinct rather than smoothed into one continuous soundtrack.
B-roll and storyboard prompts (no performer claim) (3)
These make no claim at all, no face implying a customer, no voice implying an opinion. Use them to previsualize a scene or as literal cutaway footage.
1. Product-only counter shot
Static-handheld shot (slight natural drift, no locked tripod stillness), [PRODUCT]
sitting on a real counter with [CLUTTER] softly out of focus behind it. Available
light only, from one side, casting a real shadow rather than a lit product-shot
falloff. No hands, no face, no voice. Audio: ambient room tone only, no music,
no narration.
2. Hands-placing-object B-roll
Handheld, hands only, placing or arranging [PRODUCT] in [SETTING: a bag, a shelf,
a bathroom cabinet] as part of a normal routine, not a display. Available light
matching the real setting, uneven. No face, no dialogue implied anywhere in the
prompt. Audio: the actual sound of the object being placed, ambient tone, nothing
added.
3. Room establishing shot
Handheld, slow unsteady pan across a real room where [PRODUCT] is actually used
([ROOM: kitchen, bathroom, home office]), showing genuine clutter and use-wear,
no styling pass. Available light only, whatever the room has at [TIME OF DAY].
No subject, no voice, no music, just the space itself, for use as a storyboard
reference or a literal cutaway between other clips.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.
Create An AccountWhich Prompt Architects plan includes video prompt generation?
Video Prompt Generation sits on the Advanced and Team plans, not on Pro and not on the free plan, per the feature comparison table on our pricing page. Image Prompt Generation is available starting on Pro, so if all you need is stills, you don't need to move up a tier for it; video prompting specifically is gated one level higher. If you're building a library of these formats, saving structured versions to reuse across a product line lives in the same Video Prompt Library use case.
None of this replaces judgment on the ethics section above. A tool that helps you draft the prompt faster doesn't change what you're allowed to publish once the clip exists. The 27 templates above cover the format; the two callouts above them cover what you can honestly do with the result.
For the durations and 4K handling on one specific model rather than the general craft, see our breakdown of Kling 3.0's duration and resolution limits, and for why so much AI video reads as generic in the first place, see why AI video looks fake. For the camera-movement vocabulary this post deliberately avoids, it's covered in full in our camera movement vocabulary reference, and for a wider view of the current model landscape, see the best AI video prompt tools in 2026.