TL;DR: Google's published Veo prompt formula is five parts in this order: cinematography, subject, action, context, style and ambiance. Add an explicit audio block, because Veo generates a synchronized soundtrack from the same prompt text. One thing changed since this guide first ran: the Veo 3 API models were shut down on June 30, 2026, and Veo 3.1 is what you are actually prompting.
What is the best Veo 3 prompt structure?
Google's own five-part formula is [Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance]. Camera work goes first. That ordering comes straight from Google Cloud's Veo 3.1 prompting guide, which calls cinematography "the most powerful tool for conveying tone and emotion" (Google Cloud, ultimate prompting guide for Veo 3.1, published October 16, 2025, accessed August 26, 2026).
Here is Google's own example prompt, unedited:
Medium shot, a tired corporate worker, rubbing his temples in exhaustion, in
front of a bulky 1980s computer in a cluttered office late at night. The scene
is lit by the harsh fluorescent overhead lights and the green glow of the
monochrome monitor. Retro aesthetic, shot as if on 1980s color film, slightly
grainy.
Read it back against the formula and every slot is filled. "Medium shot" is cinematography. "A tired corporate worker" is subject. "Rubbing his temples" is action. The 1980s office is context. The final sentence is style and ambiance. Nothing is left for the model to guess.
That formula is the skeleton. The rest of this guide is what you hang on it: an audio block, a camera vocabulary, a character-consistency method, 25 shot patterns, and the hard limits that stop a prompt from working.
Which Veo model are you actually prompting?
Three: Veo 3.1, Veo 3.1 Fast and Veo 3.1 Lite. All three carry -preview in their Gemini API model IDs as of August 26, 2026, while the Gemini Enterprise Agent Platform serves veo-3.1-generate-001 at GA with a stated retirement date of November 17, 2026 or later.
| Surface | What you get | Model IDs |
|---|---|---|
| Gemini API | Veo 3.1, Fast, Lite, all Preview | veo-3.1-generate-preview, veo-3.1-fast-generate-preview, veo-3.1-lite-generate-preview |
| Gemini Enterprise Agent Platform | Veo 3.1, Fast, Lite at GA | veo-3.1-generate-001 and siblings |
| Google Flow | Veo 3.1 Lite, Fast, Quality, plus Gemini Omni Flash | Model picker, no IDs |
| Retired June 30, 2026 | Veo 3, Veo 3 Fast, Veo 2 | veo-3.0-generate-001, veo-3.0-fast-generate-001, veo-2.0-generate-001 |
Sources: Google Gemini API Veo documentation and deprecations page, plus Google Flow's model help page, all accessed August 26, 2026.
One nuance worth carrying: Google's Gemini API video overview page now tells developers to "Use Gemini Omni Flash as your default model for video generation," and positions Veo 3.1 for scene extension, last-frame control, or integration with legacy pipelines (ai.google.dev/gemini-api/docs/video, accessed August 26, 2026). That recommendation is scoped to that page. DeepMind's Veo page still headlines Veo 3.1 as its video model. If you are choosing a model rather than learning to prompt one, read Veo 3 vs Sora vs Kling before you commit.
Why does prompt structure matter so much in Veo?
Because three of Veo's hard constraints are all paid for in words, and structure is how you budget them.
- A 1,024-token text ceiling. Every Veo 3.1 model lists "Text input: 1,024 tokens" in its Gemini API model card. You cannot solve a vague prompt by piling on adjectives forever. Spend tokens on the layers that change the frame.
- Native audio you are paying for either way. Audio on Veo 3.1 is listed as "Always on." There is no switch. A prompt with no sound cues does not produce a silent clip; it produces a soundtrack you did not choose.
- Per-second billing with no free tier. Google's pricing table shows Veo 3.1 Standard at $0.40 per second for 720p and 1080p and $0.60 per second for 4K, Veo 3.1 Fast at $0.10 / $0.12 / $0.30, and Veo 3.1 Lite at $0.05 / $0.08 with no 4K. Free tier: "Not available" on every row (ai.google.dev/gemini-api/docs/pricing, page last updated August 13, 2026, accessed August 26, 2026).
At eight seconds a clip, a sloppy prompt that needs five re-rolls on the Standard tier costs real money. Structure is the cheapest quality lever you have, and it is free.
What are the parts of a Veo 3 prompt?
Google's five parts are the contract. In practice I write seven blocks, because splitting Google's "context" and "style and ambiance" into scene, lighting and style makes it obvious when a block is empty. The extra rows are an authoring convenience, not extra syntax the model parses.
| Block | Maps to Google's formula | Example |
|---|---|---|
| 1. Camera | Cinematography | "medium close-up tracking shot from her right, 35mm lens, slight handheld feel" |
| 2. Subject | Subject | "a 30-year-old woman with curly red hair, wearing a long wool coat, holding a leather portfolio" |
| 3. Action | Action | "walking briskly across a wet cobblestone street, glancing back once over her shoulder" |
| 4. Scene | Context | "Paris at dusk in late autumn, light rain falling, Notre Dame faintly visible, lamp posts lit" |
| 5. Lighting | Style & ambiance | "golden-hour warm light from the west mixing with cool blue streetlamp light, reflections on wet stone" |
| 6. Style | Style & ambiance | "cinematic, shot on 35mm film, slight grain, muted teal-and-orange grade" |
| 7. Audio | Not in the five-part formula, documented separately | "footsteps on wet stone, distant traffic hum, faint church bells, sparse melancholic piano" |
You do not always need every row. A product turntable does not need dialogue. An abstract mood piece may not need a named subject. But keep the order stable, and never ship a prompt with an empty audio row.
How does each block change the output?
- Camera is the block most beginners under-use and the one Google ranks first. It sets tone before the viewer knows what they are looking at.
- Subject is the anchor. Vague subjects produce averaged, faceless results. Concrete ones with age, hair, wardrobe and a prop are repeatable.
- Action drives motion and physics. "Sprinting," "ambling" and "stumbling" all read differently, so pick the exact verb.
- Scene sets the world. Skip it and the model invents a location, usually a bland studio-grey nowhere.
- Lighting carries mood. Golden hour, overcast diffused and neon underlit give three different emotional registers from one subject.
- Style tells the model what kind of image this is: claymation, film noir on 35mm, anime, photoreal documentary.
- Audio is the multiplier, and it gets its own section below.
What does a complete Veo prompt look like?
Here is the seven-block template assembled into one copy-pasteable prompt, camera first, in Google's order:
Camera: Medium close-up tracking shot from her right side, 35mm lens, slight
handheld feel. Camera moves at her walking speed.
Subject: A 30-year-old woman with curly red hair, light freckles, wearing a
long charcoal wool coat, holding a leather portfolio.
Action: Walking briskly across a wet cobblestone street, glancing back over
her shoulder once, breath faintly visible in the cold air.
Scene: Paris at dusk in late autumn, light rain falling, Notre Dame just
visible in the background, soft fog, lamp posts lit.
Lighting: Golden-hour warm light from the west mixing with cool blue from
the streetlamps. Reflections on the wet cobblestones.
Style: Cinematic, shot on 35mm film, fine grain, muted teal-and-orange grade.
Audio: Leather shoes on wet stone, distant traffic hum, faint church bells,
a sparse melancholic piano score. She murmurs to herself: "Not yet."
Generate that and you get eight seconds of cinematic video with a synchronized soundtrack: footsteps timed to her stride, a murmured line matched to her mouth, ambient city sound underneath. The labeled-block format is not required syntax. It is an audit tool. If a block is empty, you know exactly which lever you left on the table.
You do not have to write blocks by hand every time. Prompt Architects expands a one-line idea into a structured prompt in under two seconds and saves it to a reusable library, so your best shot templates are one click away. We generate the prompt, not the video, and the free plan covers 5 enhancements a day forever (see our FAQ for the current limits). Most of the quality difference between amateur and professional Veo output is just consistency of structure, and a saved template enforces it.
How do you write audio prompts for Veo 3.1?
Google documents three audio cue types, and they map cleanly onto three layers you should write separately. The Gemini API Veo page puts it plainly: "You can provide Veo with cues for sound effects, ambient noise, and dialogue. The model captures the nuance of these cues to generate a synchronized soundtrack" (ai.google.dev/gemini-api/docs/veo, page last updated July 30, 2026, accessed August 26, 2026).
Layer 1: Dialogue
Put spoken lines in quotation marks. That is Google's documented convention on both the Gemini API page and the Cloud prompting guide.
A woman says: "We have to leave now."
A grizzled fisherman mutters, "Storm's coming. Tie her down."
Google's own worked example goes further and attaches delivery notes to the speaker: Man: (Hand on his hunting knife) "That's no ordinary bear." You can specify whispered, shouted, trembling or deadpan the same way. Keep lines short. Eight seconds is not much runway, and a long line gets rushed. Dialogue prompting in Veo 3.1 goes deeper on timing and lip sync.
Layer 2: Sound effects and ambience
Google separates these two, and so should you. Its documented example formats are SFX: thunder cracks in the distance for discrete effects and Ambient noise: the quiet hum of a starship bridge for the background soundscape.
SFX: waves crashing, seagulls calling, a distant beach voice.
Ambient noise: low office hum, a printer in the next room, occasional keyboard clicks.
Without an ambience layer, scenes feel acoustically dead. The picture moves and the soundstage is empty. One sentence usually fills it.
Layer 3: Score
Specify mood, instrumentation and tempo for any musical bed.
Soft melancholic piano with sparse strings, slow tempo, building gently.
Skip the score in dialogue-heavy scenes where music fights the voice. Add it when atmosphere is the point: abstract mood pieces, product hero shots, montages.
| Audio layer | When to use it | Prompt snippet |
|---|---|---|
| Dialogue | Character is speaking on camera | A woman says: "We have to leave now." |
| SFX | Any discrete, timed sound | SFX: thunder cracks in the distance |
| Ambient noise | Almost every shot, to fill the soundstage | Ambient noise: city traffic, distant sirens, light rain |
| Score | Mood-driven shots, montages, commercials | warm acoustic guitar, mid-tempo, hopeful |
How do you direct multiple shots in one generation?
With timestamp prompting. Google documents assigning actions to timed segments inside a single prompt so one generation produces a multi-shot sequence with controlled pacing. This is the technique most Veo guides still do not mention, and it is the biggest structural addition since this article first ran.
Here is Google's documented example, reproduced from the Cloud prompting guide:
[00:00-00:02] Medium shot from behind a young female explorer with a leather
satchel and messy brown hair in a ponytail, as she pushes aside a large jungle
vine to reveal a hidden path.
[00:02-00:04] Reverse shot of the explorer's freckled face, her expression
filled with awe as she gazes upon ancient, moss-covered ruins in the
background. SFX: The rustle of dense leaves, distant exotic bird calls.
[00:04-00:06] Tracking shot following the explorer as she steps into the
clearing and runs her hand over the intricate carvings on a crumbling stone
wall. Emotion: Wonder and reverence.
[00:06-00:08] Wide, high-angle crane shot, revealing the lone explorer
standing small in the center of the vast, forgotten temple complex,
half-swallowed by the jungle. SFX: A swelling, gentle orchestral score begins
to play.
Note what each segment still carries: a camera instruction, a subject, an action, and an audio cue. Timestamp prompting does not replace the five-part formula. It runs the formula four times inside one eight-second budget, which is why it eats tokens fast. Two-second beats are about the shortest that read as a shot rather than a flicker.
Which camera modifiers work in Veo?
The ones Google names in its own guide, and standard film vocabulary generally. You do not need to translate "dolly in" into plain English.
| Category | Modifiers Google documents or demonstrates |
|---|---|
| Camera movement | dolly shot, tracking shot, crane shot, aerial view, slow pan, POV shot |
| Composition | wide shot, close-up, extreme close-up, low angle, two-shot |
| Lens and focus | shallow depth of field, wide-angle lens, soft focus, macro lens, deep focus |
| Also parsed reliably | medium shot, medium close-up, over-the-shoulder, eye-level, high angle, top-down, Dutch tilt, handheld, gimbal-smooth, push-in |
The pattern that works: one composition modifier, one movement modifier, one lens note. "Medium close-up, slow dolly in, shallow depth of field" is a clear, non-contradictory instruction. Five stacked modifiers is the fastest way to get mush. For a fuller vocabulary, see 30 cinematic camera prompts.
What about lighting modifiers?
Lighting carries the emotional weight. Specify source, direction and mood, and let the scene do some of the work.
| Category | Modifiers |
|---|---|
| Source | natural daylight, golden hour, blue hour, overcast diffused, studio softbox, neon, candlelight, firelight, practicals |
| Direction | front-lit, side-lit, backlit, top-lit, underlit, rim light |
| Mood | warm, cool, high-contrast, low-contrast, moody, ethereal, gritty, cinematic, dreamy |
If you already wrote "Paris at dusk, light rain" into the scene block, you have implied a cool, low-contrast, reflective condition. The lighting block refines it rather than starting over. Scene and lighting should reinforce each other, never contradict.
How do you keep a character consistent across multiple shots?
Reference images beat every text method, and Veo 3.1 accepts up to three. Google's documentation is direct about what they do: "Provide up to three asset images of a single person, character, or product. Veo preserves the subject's appearance in the output video."
You have three approaches, weakest to strongest.
Approach 1: Prompt repetition
Lock the subject description verbatim at the top of every shot prompt. Same name, age, hair, wardrobe, distinguishing features, word for word. Lowest effort, acceptable for two or three shots, drifts over longer sequences.
Approach 2: JSON character object
Define the character once in a structured object and paste it into every shot. Identical tokens each time leaves the model less room to drift:
{
"character": {
"name": "Sarah",
"age": 30,
"appearance": "curly red hair, shoulder-length, green eyes, light freckles",
"wardrobe": "long charcoal wool coat, black leather boots, leather portfolio"
},
"world": {
"location": "Paris, autumn dusk, light rain",
"palette": "warm golden plus cool blue contrast",
"style": "cinematic 35mm, fine grain"
}
}
Paste character and world at the top of each shot prompt, then add only the per-shot camera and action below. Because the appearance tokens are byte-for-byte identical, drift drops sharply. Note that Veo's documented request body takes a plain text prompt string, so this is a discipline for structuring your text, not a separate API field. Our JSON video prompt templates for Veo has ready-made objects.
Approach 3: Reference images
The strongest method. Supply up to three images through the referenceImages parameter. One catch worth planning around: reference-image generations are locked to eight seconds, and Google lists them alongside 1080p and 4K as cases where durationSeconds must be "8".
| Feature | Reference images | JSON character object | Prompt repetition |
|---|---|---|---|
| Setup effort | Medium | Medium | Low |
| Holds a face across 10 shots | Partially | ||
| Works text-only, no assets | |||
| Forces 8-second duration | |||
| Available on Veo 3.1 Lite | |||
| Best for | Faces, branded products | 5 to 10 shot sequences | Quick tests |
For anything where a real face or a branded product has to stay identical, reach for reference images. Text-only methods are for speed and prototyping. If your character is melting between frames rather than merely drifting, that is a different failure: see why your AI video morphs and warps.
What are 25 tested Veo prompt examples?
Twenty-five patterns in five categories. Each is a starting point. Adapt the subject, scene and audio, but keep the skeleton intact.
Cinematic narrative (5)
- Emotional close-up. Eye-level medium close-up, shallow depth of field, a solo character in a quiet moment, soft window light, sparse piano score, one murmured line of dialogue.
- Two-character dialogue. Over-the-shoulder framing at 50mm, two named characters with distinct wardrobes, dialogue in quotes for both, room-tone ambience.
- City tracking shot. Third-person follow through a crowded street, gimbal-smooth at walking pace, golden hour, layered street ambience.
- Slow push-in on an object. Static-to-dolly-in on a key prop such as a letter, a ring or a phone. Macro detail, dramatic side light, a slow building string note.
- Wide establishing reveal. Crane-up wide shot revealing a skyline at blue hour, slow rise, wind and distant city ambience, a swelling orchestral cue.
Product and commercial (5)
- Hero turntable. Product centered on a slow-rotating turntable, studio softbox lighting, sweep backdrop, 85mm lens, no dialogue, clean ambient pad with a subtle whoosh on the reveal.
- Lifestyle placement. Product in a natural domestic setting such as a kitchen counter or desk, warm practical lighting, a hand entering frame to use it, cozy room ambience.
- Liquid pour. Slow-motion liquid pouring into a glass, side lighting to catch translucency, crisp pour and fizz SFX.
- Top-down reach. Top-down framing, a hand reaching for the product on a styled flat-lay surface, even soft light, tactile foley as fingers make contact.
- Gradient sweep reveal. Product reveal as a light gradient sweeps left to right, dark-to-bright, minimalist studio, a single rising synth tone.
Abstract and mood (5)
- Fabric in wind. Slow-motion silk billowing against a colored gradient, backlit, dreamy soft focus, an airy ambient drone.
- Particles in light. Dust drifting through colored volumetric beams, dark background, slow real-time motion, soft pad score.
- Macro surface tension. Extreme macro of a water droplet hitting a surface, high-speed slow motion, ring light, crisp impact SFX with a low resonant tail.
- Cloud time-lapse. Time-lapse of clouds rolling over a horizon, shifting golden-to-blue light, no subject, wind and faint atmospheric score.
- Morphing geometry. Geometric shapes morphing on an unbroken neutral background, clean studio light, minimalist, a rhythmic ambient pulse.
Documentary and interview (5)
- Talking head. Subject seated, eye-level, soft natural window light, slight rack focus to background, room ambience, dialogue in quotes.
- Working hands. Close-up of skilled hands at a craft such as pottery or woodworking, shallow focus, warm workshop light, rich tool and material foley.
- Walk-and-talk. Handheld follow shot of a subject walking and speaking to camera, natural daylight, dialogue in quotes, footsteps and street ambience.
- B-roll insert. Detail shots of an environment, textures and signage and objects, no people, gentle camera drift, location-specific ambience.
- Environmental portrait. Subject framed within their setting, a baker in a bakery, wide-to-medium, practical light, ambient workplace sound, optional single line.
Action and kinetic (5)
- Skateboard trick. Low-angle shot of a skateboarder landing a trick, slow-motion look, harsh midday sun, wheels-on-concrete and board-clack SFX.
- Sunrise runner. Side-tracking shot of a runner at sunrise, long lens compression, backlit rim light, breathing and footfall ambience, driving score.
- Cooking sequence. A timestamped four-beat sequence of chopping, sizzling and plating, top-down and macro mix, warm kitchen light, layered cooking SFX.
- Marketplace movement. High-angle crowd movement in a busy market, smooth slow pan, vibrant natural light, dense market ambience.
- Vehicle drive-by. Tracking a car passing camera with motion blur, low angle, overcast diffused light, engine doppler and tire SFX.
Want the expanded prompt text for each of these? We keep a maintained set in our free Veo 3 prompt generator, formatted for one-click copy.
What are the most common Veo prompt mistakes?
The failure patterns are predictable once you have burned enough credits.
- No audio cues. Still the number-one mistake, and now an expensive one because audio is always on and always billed. Fix: at least one ambient noise line in every prompt.
- Generic subjects. "A woman walks" gives you faceless stock footage. Fix: add age, hair, wardrobe and a prop.
- Contradictory framing. "Wide shot close-up" or "static handheld tracking shot" produces mush. Fix: one composition, one movement, one lens.
- Skipped context. No location, no weather, no time of day means bland defaults. Fix: always say where, when and what the weather is.
- Asking for a square video. Veo supports
16:9and9:16only. There is no 1:1 option on any Veo 3.1 model. Fix: shoot 9:16 for vertical feeds and crop in post if you need square. An older version of this guide told you to write "1:1 square" into the prompt. It never worked. - Setting aspect ratio in prose. Aspect ratio is the
aspectRatioAPI parameter, not a phrase in your prompt text. Writing "9:16 vertical" in the prompt body is not what switches the output. - Negative prompts written as negatives. Google's guidance is to describe what you want excluded positively: "a desolate landscape with no buildings or roads" rather than "no man-made structures."
- One mega-prompt for a whole sequence. Trying to describe five shots freeform overruns the token budget. Fix: use timestamp prompting, or one prompt per shot with a shared character object.
- Over-stuffing. Past roughly 300 words the model starts dropping constraints, and 1,024 tokens is a hard wall. Cut redundant adjectives.
If you generate across ChatGPT, Claude, Gemini, Midjourney and Veo regularly, these compound. A consistent enhancement workflow catches most of them before you spend a generation credit.
What are the technical specs and limits of Veo 3.1?
Knowing the hard constraints stops you writing prompts the model cannot honor. Everything below is from Google's Gemini API Veo documentation, page last updated July 30, 2026, accessed August 26, 2026.
| Spec | Veo 3.1 and 3.1 Fast | Veo 3.1 Lite |
|---|---|---|
| Resolution | 720p (default), 1080p, 4K | 720p (default), 1080p |
| Clip duration | 4, 6 or 8 seconds. Must be 8 for 1080p, 4K, reference images or extension | 4, 6 or 8 seconds. Must be 8 for 1080p or reference images |
| Frame rate | 24fps | 24fps |
| Aspect ratios | 16:9 (default), 9:16 | 16:9 (default), 9:16 |
| Text input limit | 1,024 tokens | 1,024 tokens |
| Audio | Native, always on | Native, always on |
| Reference images | Up to 3 | Not supported |
| Video extension | Supported | Not supported |
| Videos per request | 1 | 1 |
| Status on Gemini API | Preview | Preview |
A few practical notes that catch people out:
- Scene extension has real numbers now. Veo 3.1 extends a Veo-generated video by 7 seconds at a time, up to 20 times, producing a combined output of up to 148 seconds. Input videos must be 720p, 16:9 or 9:16, and 141 seconds or shorter. Google's launch blog described this loosely as "even lasting for a minute or more"; the documentation is more precise.
- Extension runs at 720p only. If you need a 4K finish, generate the full sequence first and upscale outside Veo.
- Videos are deleted after 2 days. Generated videos are stored server-side for 2 days, then removed. Referencing one for extension resets its timer. Download anything you care about.
- Everything is watermarked. All Veo output carries a SynthID watermark, verifiable through Google's SynthID platform. Plan for that in any client deliverable.
- Latency runs 11 seconds to 6 minutes. Google documents a minimum of 11 seconds and a maximum of 6 minutes at peak. Build your review loop around the slow case.
What changed between Veo 3 and Veo 3.1?
Veo 3.1 launched on October 15, 2025 in paid preview on the Gemini API, alongside Veo 3.1 Fast (Google Developers Blog, accessed August 26, 2026). Google described it as richer native audio, greater narrative control, and enhanced image-to-video with better character consistency across multiple scenes, plus three new capabilities: reference images, scene extension, and first-and-last-frame transitions.
Two later dates matter as much:
- January 13, 2026 — Google added 4K output and improved 1080p for Veo, plus native vertical generation for Ingredients to Video rather than cropping from landscape.
- March 31, 2026 — Veo 3.1 Lite launched as the cost-efficient tier.
The prompt grammar did not change across any of it. Every technique in this guide works on Veo 3.1, Fast and Lite. That is precisely why a Veo 3 structure guide is still useful after the Veo 3 endpoint went away.
How do you build a repeatable Veo workflow?
One great clip is luck. A repeatable look is a workflow.
- Pick a pattern from the 25 above that matches your shot.
- Fill the blocks in Google's order: camera, subject, action, context, style, then audio. Do not leave a block empty unless it genuinely does not apply.
- Add all three audio layers that fit: dialogue, SFX and ambient noise, plus score if atmosphere is the point.
- Prototype on Veo 3.1 Lite or Fast at 720p. At $0.05 to $0.10 a second you can afford to be wrong. Composition and pacing read fine at 720p.
- Change one variable at a time. Only the camera, or only the lighting, between attempts. Change three things and you learn nothing.
- Save what works. When a modifier combination reliably produces your look, store it. Your library of proven shot templates is the actual asset.
- Finish on Veo 3.1 Standard at 1080p or 4K, eight seconds, once the structure is locked.
Step 6 is where most people leave value on the table. The difference between a hobbyist and someone shipping consistent video is not talent. It is a library of battle-tested prompt structures they can reuse and remix, with global variables so swapping the subject across ten saved templates is one edit instead of ten. If you also work across image and text models, the same discipline applies: How to direct AI video like a filmmaker covers the cross-model view.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.
Create An AccountBy Nafiul Hasan, founder of Prompt Architects, building AI prompt-enhancement tools used across ChatGPT, Claude, Gemini, Midjourney and Veo. All Veo specifications in this article were verified against Google-owned documentation on August 26, 2026. Model availability changes fast; check Google's Veo documentation before you build against a spec here.