TL;DR: A video prompt has seven parts: shot, subject, action, setting, light, look and sound. That grouping is mine. No vendor enforces one, and Google alone publishes four different decompositions across four of its own pages. Below, each part is defined, built cumulatively, and rendered for Veo 3.1, Kling 3.0, Sora 2 and Wan 2.7.
What is video prompt structure?
Video prompt structure is the order and grouping of information you hand a text-to-video model. It is not a syntax. Every current model takes one free-text field, and none of them parse your prompt into slots. If you write your seven parts in a random order, nothing errors.
Structure earns its keep somewhere else. When a clip comes back wrong and your prompt is one 90-word run-on sentence, you cannot tell whether the problem was the framing, the light or the action, so you rewrite the whole thing and get a differently wrong clip. When the same prompt is built from named parts, you change one part, regenerate, and learn something. That is the entire argument for structure, and it is a craft argument rather than a technical one.
Does any vendor publish an official video prompt structure?
Several do. None of them agree, and Google disagrees with itself four ways.
| Google page | What it publishes | Count |
|---|---|---|
ai.google.dev/gemini-api/docs/veo → "Veo prompt guide" | Subject, Action, Style, then Camera positioning and motion, Composition, Focus and lens effects, Ambiance (last four each marked [Optional]) | 7 |
cloud.google.com/blog/…/ultimate-prompting-guide-for-veo-3-1 | "Consider this five-part formula for optimal control": [Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance] | 5 |
| Gemini Enterprise Agent Platform, "Video generation prompt guide" → "Anatomy of a prompt" | Subject · Action · Scene or context · Camera angles · Camera movements · Lens and optical effects · Visual style & aesthetics · Temporal elements · Audio · Cinematic terms · Negative prompts | 11 |
deepmind.google/models/veo/prompt-guide/ | "What to think about while writing prompts": Shot framing and motion · Style · Lighting · Character descriptions · Location · Action · Dialogue | 7 |
All four pages were live on August 27, 2026. Note that the two sevens are not the same seven: the Gemini API list has Composition and Ambiance and no Lighting or Dialogue, and the DeepMind list has Lighting and Dialogue and no Composition. DeepMind also frames its list as something to draw from rather than fill in, in its own words asking you to "use a mix of the elements below". Only the Cloud blog states an order, and only the Cloud blog calls its list a formula.
Other vendors publish less, and what they publish is shorter:
- ByteDance / Seedance, on the BytePlus ModelArk video generation tutorial: "Prompt = subject + motion, background + motion, camera + motion ...", plus the instruction to "place important content at the front".
- Luma, in the Ray 3.2 generation reference for the
promptfield: "A text description of the video, 1–6,000 characters. Be specific about subject, motion, camera movement, lighting, and pacing." - OpenAI, in the Sora 2 prompting guide's "Prompt anatomy that works": "State the camera framing, note depth of field, describe the action in beats, and set the lighting and palette."
- Kling publishes no formula at all. Its VIDEO 3.0 user guide teaches by worked example, and its API spec documents a shot-list grammar instead, covered below.
So the honest summary is: five categories recur across every vendor (subject, action, setting, camera, style), lighting and sound recur across most, and nobody publishes the same list twice. What follows is my synthesis of those, tested against each vendor's documented syntax.
The seven parts, and what each one controls
Shot, Subject, Action, Setting, Light, Look, Sound. Seven because that is where the categories stop overlapping, not because seven is a magic number. Each section below gives you what the part controls, what happens when you drop it, the vendor-specific syntax where one exists, and the prompt so far. The example builds one part at a time, so you can paste any intermediate version and see what changed.
Part 1 — Shot: what does the camera do?
The shot controls framing, height, angle and movement. It is the only part that describes the viewer rather than the world.
Omit it and the model picks. On a text-to-video generation with no framing cue, you will get a different distance and a different move on every seed, which is the single biggest reason a set of clips will not cut together.
Google's Cloud blog puts [Cinematography] first in its formula and calls it "the most powerful tool for conveying tone and emotion". The Gemini Enterprise guide splits it into three sections: camera angles, camera movements, and lens and optical effects. Sora's guide asks for camera framing and depth of field as separate lines.
Medium shot, eye level, slow dolly in, shallow depth of field.
Part 2 — Subject: who or what is in the frame?
The subject is the thing the shot is about, described with two or three fixed identifying details. Fixed matters more than vivid: the details you repeat across shots are what holds identity together.
Omit it and you get a competent, generic clip of a place with nobody in it, or an anonymous figure who changes between generations.
Google's Gemini API Veo guide lists Subject as one of three non-optional elements. DeepMind's list splits the same job into "Character descriptions", and gives the sharpest advice on the page: "A woman in her twenties with wavy brown hair and light freckles" will produce more specific results than "a brown-haired woman."
Medium shot, eye level, slow dolly in, shallow depth of field. A woman in her
sixties in a canvas apron, grey hair tied back, forearms bare.
Part 3 — Action: what happens, in beats?
Action is the verb, timed. One clear subject action per shot, described as counted beats rather than a general state.
Omit it and the model animates ambience: hair moves, light shifts, nothing happens. Over-specify it, with four unrelated actions in an eight-second clip, and you get morphing as the model tries to fit them all in.
OpenAI's guide is the most concrete here. Its weak-versus-strong pair is "Actor walks across the room" against "Actor takes four steps to the window, pauses, and pulls the curtain in the final second." That is a rule you can apply to any model.
Medium shot, eye level, slow dolly in, shallow depth of field. A woman in her
sixties in a canvas apron, grey hair tied back, forearms bare. She unhooks the
awning strap, folds the canvas in three, and sets it on the counter in the
final second.
Part 4 — Setting: where and when is this?
Setting is place, time of day, weather and period. Google's Gemini Enterprise guide calls it "the 'where' and the 'when'"; the Cloud blog formula calls it [Context].
Omit it and the background becomes a soft, unplaceable blur that reads as stock footage. Setting is also what carries continuity across a sequence, so it is worth writing once and pasting verbatim into every shot.
Medium shot, eye level, slow dolly in, shallow depth of field. A woman in her
sixties in a canvas apron, grey hair tied back, forearms bare. She unhooks the
awning strap, folds the canvas in three, and sets it on the counter in the
final second. A covered market stall at closing time, empty aisles behind her,
wet floor, late October.
Part 5 — Light: where is the key, and what colour is everything?
Light is the source, its direction, its quality, and a short list of named palette colours. It is the part most people skip and the part that most reliably breaks an edit.
Omit it and each generation invents its own key direction. Two clips of the same subject, one lit from camera left and one from camera right, will not cut together no matter how good either one is.
Sora's guide gives the tightest template here, including the idea of palette anchors: "Naming three to five colors helps keep the palette stable across shots." Google's Gemini Enterprise guide puts Lighting under "Visual style & aesthetics" alongside tone, artistic style and ambiance.
… A covered market stall at closing time, empty aisles behind her, wet floor,
late October. Low sun through the market's end doors from camera left, one
sodium lamp overhead behind her. Palette: rust, cream, wet slate.
Part 6 — Look: what medium is this pretending to be?
Look is the format claim: film stock, sensor, era, grain, grade, animation style. One line, usually at the end, and it colours everything above it.
Omit it and you get the model's house style, which is currently a clean, slightly plastic digital render. That is a real aesthetic choice made on your behalf.
OpenAI's guide argues for putting this early rather than late: "Establish this style early so the model can carry it through consistently." Google's Gemini API list has Style as one of its three non-optional elements. Both are worth more than the specific vocabulary.
… Palette: rust, cream, wet slate. Shot on 35mm, fine grain, gentle halation
on the overhead lamp.
Part 7 — Sound: what do you hear, including silence?
Sound is dialogue, sound effects and ambient bed. On every current audio-capable model it lives in the prompt text, not in a parameter.
Omit it and an audio-capable model invents a soundtrack. That is usually a generic music cue you then have to strip in the edit, which is why "no music" is a line worth writing.
The syntax is close to standard across vendors. Google's Gemini API Veo guide asks for quotes around dialogue, explicit description for SFX, and a described soundscape for ambience. The Cloud blog shows the same three as labelled fragments: A woman says, "We have to leave now.", SFX: thunder cracks in the distance, Ambient noise: the quiet hum of a starship bridge. Kling 3.0 gates audio behind settings.audio: "native" but still takes the content from the prompt.
Here is the complete seven-part prompt:
Medium shot, eye level, slow dolly in, shallow depth of field. A woman in her
sixties in a canvas apron, grey hair tied back, forearms bare. She unhooks the
awning strap, folds the canvas in three, and sets it on the counter in the
final second. A covered market stall at closing time, empty aisles behind her,
wet floor, late October. Low sun through the market's end doors from camera
left, one sodium lamp overhead behind her. Palette: rust, cream, wet slate.
Shot on 35mm, fine grain, gentle halation on the overhead lamp.
Ambient noise: distant roller shutters and a single trolley wheel. SFX: a
canvas snap on the fold. No music.
Full detail on audio-only phrasing is in Veo audio prompts.
What about duration, resolution and aspect ratio?
None of them are parts of the prompt. They are the container, and they live in API fields.
OpenAI states this more plainly than anyone: "These parameters are the video's container: resolution, duration, and character references will not change based on prose like 'make it longer.' Set them explicitly in the API call; your prompt controls everything else."
The one live exception is Seedance on BytePlus ModelArk, which still honours a legacy inline flag syntax appended to the prompt text, with what its own docs call "loose validation":
<your seven-part prompt> --rs 720p --rt 16:9 --dur 5 --seed 11 --cf false --wm true
BytePlus recommends the JSON body instead, because that path uses strict validation and returns an error on a bad value rather than ignoring it. Treat the flags as a compatibility surface, not a feature.
Where does each part actually land, per model?
| Feature | Veo 3.1 | Kling 3.0 Omni | Sora 2 | Wan 2.7 | Luma Ray 3.2 |
|---|---|---|---|---|---|
| Shot / camera | Named element, listed first in the Cloud formula | Per-shot in the multi-shot grammar | Own labelled line: Camera shot | Prose, inside a timestamped shot header | Named in the prompt field docs |
| Subject | Non-optional element | Prose, plus bindable Elements | Prose, anchor with fixed details | Prose | Named in the prompt field docs |
| Action | Non-optional element | Prose | Prose, described in beats | Prose | Named as "motion" |
| Setting | Called Context in the formula | Prose | Prose | Prose | Not named separately |
| Light | Inside Style & Ambiance | Prose | Own line, plus palette anchors | Prose | Named in the prompt field docs |
| Look / style | Non-optional element | Prose | Advised to go early | Prose, or via prompt_extend | Not named separately |
| Sound | Prose, separate sentences advised | Prose, gated by settings.audio | Dialogue block plus ambience line | Prose, or an uploaded audio_url | Not documented on the generation page |
| Negative prompt field | negative_prompt, 500 chars | ||||
| Frame rate published | 24fps | Not published | Not published | Not published | 24fps grid implied by keyframes |
| Prompt length cap published | 1,024 tokens | 3,072 chars | Not published | 5,000 chars | 6,000 chars |
Two rows deserve a note. Kling publishes no frame rate, anywhere. There is no occurrence of "fps", "frame rate" or "frames per second" in the full Kling 3.0 and 3.0 Omni text-to-video API specification, and Kling's own llms.txt gives contradictory ranges across models, so it cannot be used to settle it. Any specific Kling frame rate you find online is not coming from Kling. And negative prompt support is a per-model question rather than a per-vendor one, which the negative prompt support matrix works through properly.
What order should the seven parts go in?
Vendors disagree, and the disagreement is small enough that it is not worth agonising over.
Google's Cloud formula runs cinematography first. Sora's guide advises establishing style early. BytePlus tells you to "place important content at the front". Nobody publishes evidence that a different order fails.
My default is the order above, for a non-technical reason: it reads like a shot list, so a human collaborator can review it. Order the parts however you like, but keep the order fixed across a project. Consistent order is what makes a set of prompts diffable, and diffable prompts are what let you attribute a change to a cause.
The same prompt, rendered for four models
Same seven parts. Four surface forms.
Veo 3.1 takes one continuous paragraph, with audio in separate sentences, under 1,024 tokens:
Medium shot, eye level, slow dolly in, shallow depth of field. A woman in her sixties in a canvas apron, grey hair tied back, forearms bare. She unhooks the awning strap, folds the canvas in three, and sets it on the counter in the final second. A covered market stall at closing time, empty aisles behind her, wet floor, late October. Low sun through the end doors from camera left, one sodium lamp overhead behind her; palette of rust, cream and wet slate; shot on 35mm with fine grain and gentle halation.
Ambient noise: distant roller shutters and a single trolley wheel. SFX: a canvas snap on the fold. No music.
For a timed sequence in a single Veo generation, the Cloud blog's timestamp workflow assigns each beat a range:
[00:00-00:03] Medium shot, eye level, slow dolly in: she unhooks the awning strap at a covered market stall at closing time, wet floor, empty aisles behind. SFX: distant roller shutters.
[00:03-00:06] Close-up of her hands folding the canvas in three, same low side light, shallow depth of field. SFX: a canvas snap on the fold.
[00:06-00:08] Medium shot, static: she sets the folded canvas on the counter and looks off left. Emotion: quiet, unhurried.
Kling 3.0 Omni documents an explicit multi-shot grammar in its API spec: "shot n, m, words; shot n, m, words;", where n is the shot number, m is that shot's duration in seconds, and words is that shot's prompt at up to 512 characters. Shots run from 1 to 6, each is at least one second, and the durations must sum to the total:
{
"prompt": "shot 1, 3, Medium shot, eye level, slow dolly in. A woman in her sixties in a canvas apron, grey hair tied back, unhooks the awning strap at a covered market stall at closing time, wet floor, empty aisles behind her. Low sun through the end doors from camera left, one sodium lamp behind her, rust and wet slate palette, 35mm fine grain; shot 2, 3, Close-up on her hands folding the canvas in three, same low side light, shallow depth of field; shot 3, 4, Medium shot, static. She sets the folded canvas on the counter and looks off left. Ambient roller shutters and one trolley wheel, a canvas snap on the fold, no music.",
"settings": {
"multi_shot": true,
"audio": "native",
"resolution": "1080p",
"aspect_ratio": "16:9",
"duration": 10
}
}
Kling's web interface uses a looser version of the same idea. Its own Custom Multi-Shot example drops the durations and just numbers the shots:
Shot 1, medium shot eye level, slow dolly in, a woman in a canvas apron unhooks an awning strap at a market stall at closing time, cinematic handheld.
Shot 2, close-up of her hands folding the canvas in three, same low side light.
Shot 3, medium static shot, she sets the folded canvas down and looks off left.
Note that settings.multi_shot defaults to true on Kling 3.0, so a single-shot prompt can still come back cut. Set it to false when you want one unbroken take. The finer points of Kling's shot and camera vocabulary are in Kling AI prompt format.
Sora 2 responds well to labelled lines, which is the shape its own guide's strong examples take:
Camera shot: medium shot, eye level, slow dolly in
Depth of field: shallow, sharp on subject, market aisles soft behind
Subject: a woman in her sixties in a canvas apron, grey hair tied back, forearms bare
Action: she unhooks the awning strap, folds the canvas in three, sets it on the counter in the final second
Setting: a covered market stall at closing time, wet floor, empty aisles, late October
Lighting + palette: low sun through the end doors from camera left, one sodium lamp behind her
Palette anchors: rust, cream, wet slate
Look: 35mm film, fine grain, gentle halation on the overhead lamp
Sound: distant roller shutters, one trolley wheel, a canvas snap on the fold, no music
Wan 2.7 documents timestamped shot headers in prose, and is explicit that its shot_type parameter "has no effect". It is also the one model here with a real negative prompt field:
{
"model": "wan2.7-t2v",
"input": {
"prompt": "Generate a multi-shot video. Shot 1 [0-3 seconds] medium shot, eye level, slow dolly in: a woman in her sixties in a canvas apron unhooks an awning strap at a covered market stall at closing time, wet floor, empty aisles behind her, low sun from camera left, one sodium lamp behind, 35mm fine grain. Shot 2 [3-6 seconds] close-up: her hands fold the canvas in three under the same low side light. Shot 3 [6-10 seconds] medium shot, static: she sets the folded canvas on the counter and looks off left. Ambient roller shutters and a trolley wheel, a canvas snap on the fold, no music.",
"negative_prompt": "modern signage, brand logos, on-screen text, extra fingers, low resolution"
},
"parameters": {
"resolution": "1080P",
"ratio": "16:9",
"duration": 10,
"prompt_extend": false
}
}
prompt_extend defaults to true on Wan, which means an LLM rewrites your prompt before generation. That is helpful for a one-line idea and actively unhelpful once you have spent an hour tuning seven parts, so turn it off when you are iterating.
How do you debug a prompt that is not working?
Subtract, do not add. The instinct when a clip comes back wrong is to write more, and more prose usually makes it worse.
- Strip to Shot plus Subject plus Action. Three parts, one sentence. If that is wrong, nothing you add later fixes it.
- Add Setting. Regenerate. This is where morphing and warping usually appear, because a busy environment gives the model more to lose track of.
- Add Light and Look together. They interact; splitting them wastes a generation.
- Add Sound last. It cannot break the picture, so it belongs at the end of the ladder.
Here is step one, which is worth keeping around as a control prompt:
Medium shot, eye level, slow dolly in. A woman in her sixties in a canvas apron
unhooks an awning strap and folds the canvas in three.
OpenAI's guide gives the same advice from the other direction: "If a shot keeps misfiring, strip it back: freeze the camera, simplify the action, clear the background. Once it works, layer additional complexity step by step." It also notes that shorter clips follow instructions more reliably, and suggests stitching two four-second clips rather than generating one eight-second clip. If your renders keep dissolving mid-shot, that is usually step two on the ladder: too much environment, or two actions competing for the same second.
Where to keep the seven-part template
Seven parts is easy to hold in your head for one prompt and impossible to hold across a project. The practical failure is not forgetting the structure, it is retyping it: you write a good seven-part prompt for a client, adapt it three times in a chat window, and the best version is the one you cannot find next week.
A prompt template only compounds when the structure is stored once and the changing parts are marked as variables. That is the layer Prompt Architects builds. You save the skeleton below, mark the brackets as variables, and fill them per shot instead of retyping the grammar:
[SHOT]. [SUBJECT]. [ACTION, in beats]. [SETTING, time of day, weather].
[LIGHT: source, direction, quality]. Palette: [THREE TO FIVE COLOURS].
[LOOK: format, stock, grain, grade].
Ambient noise: [BED]. SFX: [ONE OR TWO SPECIFIC SOUNDS]. [MUSIC OR "no music"].
Video prompt generation sits on the Advanced and Team plans rather than Pro, and the free plan includes 5 prompt enhancements per day, forever, per our FAQ.
To be clear about what this page is: we build the prompt layer. We do not generate video, we do not host any of these models, and we are not affiliated with Google, Kuaishou, OpenAI, Alibaba or Luma. Running anything above needs an account and credits with that vendor. For the model-specific versions rather than the cross-model structure, Veo 3 prompt structure and Kling AI prompt format go deeper.
Every specification on this page was read from the vendor's own documentation on August 27, 2026. This category ships fast enough that the vendor's page is always the final authority.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An Account