TL;DR: Vidu prompts are two things at once: a scene written as prose, and a JSON body whose fields differ per endpoint. These 22 templates are the prose half, grouped by job and left as fill-in-the-blank skeletons, with the settings each one needs printed above it. All checked against platform.vidu.com on 27 August 2026.
This is a reference sheet, not a tutorial. If you want the mechanics behind the fields, our guide to prompting Vidu Q3 reads the API schema line by line. This page assumes you already know which endpoint you are calling and just want something to fill in.
What goes in a Vidu prompt template?
Five slots, in this order, because Q3 gives you nowhere else to put them.
Vidu disabled the two fields you would normally reach for. movement_amplitude carries the note "This parameter does not take effect when using the q2 & q3 model" on text2video, and the style field says "This parameter is not effective when using q2 or q3 models". Both still accept a value and echo it back in the response. Neither changes the output. So motion and register are prose problems, and a template that leaves them out will drift.
| Slot | What it fixes | Why it cannot live elsewhere |
|---|---|---|
| Subject and action | Who or what moves, and what they do | The only place it exists |
| Camera | Framing, lens feel, and the move | movement_amplitude is inert on Q3 |
| Light | Key direction, quality, time of day | No lighting field on any endpoint |
| Sound | Speech, effects, and whether there is music | Q3 generates audio from the prompt text |
| Cut plan | One take, or N shots with their lengths | Q3 does its own scene cuts unless told not to |
The sound slot is the one people leave off. Q3 replaced the old bgm boolean, which text2video now marks with "BGM does not available in q3 models", and generates speech, effects and score from the prompt instead. Write silence if you want silence. Our sound design prompts for AI video has the vocabulary for that clause.
Here is one template filled in, so the bracket convention below is unambiguous:
SKELETON
[SUBJECT] moves [HOW] across [SPACE]. The camera [MOVE] and settles.
[LIGHT]. Sound: [SOUND]. One continuous take, no cuts.
FILLED
A grey whippet trots stiff-legged across a wet supermarket car park. The camera
tracks left at knee height and settles. Flat overcast light, faint sodium spill
from the shopfront. Sound: paws on wet tarmac, distant trolley rattle, no music.
One continuous take, no cuts.
How do you spend a 16-second Vidu Q3 budget?
Carefully, because you cannot top it up afterwards.
Sixteen seconds is the documented ceiling for viduq3-pro and viduq3-turbo on text2video, img2video and start-end2video, with a default of 5. On reference2video the same two models are documented at 3 to 16, and viduq3-mix alone at 1 to 16. That is the whole budget for one call.
The two endpoints that would extend it are both closed to Q3. Vidu's Video Extension endpoint, which adds 1 to 7 seconds to an existing clip, publishes a model enum containing only viduq2-pro and viduq2-turbo. The Intelligent Multi-Frame endpoint, which strings keyframes together and states "Each task supports up to 9 frames, Each task supports at least 2 frames.", publishes the same two-model enum. Neither lists a Q3 string.
Practical shot math for a single call. Every Q3 row on the Model Map runs at 24fps, so a beat you can actually read needs about two seconds. Sixteen seconds is therefore three shots of five with a second of air, or two shots of eight, or one take. Four or more cuts inside 16 seconds tends to arrive as a montage of half-formed ideas. Duration is billed per second on Vidu's pricing page, so a 16-second draft also costs roughly three times a 5-second one before you have seen whether the idea works.
Which endpoint and model string does each job need?
Route the job first, then pick the template. The model enums differ per endpoint and copying a string sideways is the fastest way to a rejected request.
| Job | Endpoint | Q3 strings in that enum | Duration | Frame shape from |
|---|---|---|---|---|
| Build a scene from nothing | text2video | viduq3-pro, viduq3-turbo | 1–16s, default 5 | aspect_ratio |
| Animate one still | img2video | viduq3-pro-fast, viduq3-turbo, viduq3-pro | 1–16s, default 5 | the input image |
| Move between two stills | start-end2video | viduq3-turbo, viduq3-pro | 1–16s, default 5 | the input images |
| Hold subjects consistent | reference2video, images form | viduq3-mix, viduq3-turbo, viduq3 | 3–16s, viduq3-mix 1–16 | aspect_ratio |
| Name subjects with handles | reference2video, subjects form | viduq3-turbo, viduq3 | 3–16s, default 5 | aspect_ratio |
Two notes that decide half the routing. viduq3-pro is absent from the reference2video model enum in both request forms, even though that page's own duration table names it. And the subjects form carries a flat exclusion, "Note: viduq3-mix does not support the use of entities for the time being", so if your template uses at-sign handles, viduq3-turbo is the Q3 model to reach for.
Which Vidu prompts work for product demos and e-commerce?
Four skeletons. Product work lives or dies on what you forbid, so each one ends with an exclusion line.
1 — HERO ORBIT
viduq3-pro · text2video · 8s · 16:9
[PRODUCT] stands on [SURFACE] against [BACKDROP]. The camera orbits [CLOCKWISE
OR ANTICLOCKWISE] through ninety degrees at one constant speed, product held
dead centre and in focus throughout. [KEY LIGHT DIRECTION AND QUALITY], one
fill from the opposite side. Sound: [ROOM TONE OR SILENCE].
No hands, no people, no on-screen text, no reflections of a studio.
2 — UNBOXING, HANDS ONLY
viduq3-turbo · img2video · 10s · frame from the source image
Two hands enter from the bottom of the frame, lift the lid of [PACKAGE] straight
up and set it aside off-camera left, then withdraw. [PRODUCT] is revealed still
seated in [INSERT MATERIAL]. The camera holds absolutely still for the first six
seconds, then pushes in two hand-widths. [LIGHT]. Sound: [CARDBOARD OR FOAM
SOUND], no music, no voice.
Never show a face, a wrist watch or a sleeve logo.
3 — PRODUCT IN SITU FROM REFERENCES
viduq3-mix · reference2video, images form · 12s · 16:9
Place the product from the reference images into [SCENE FROM THE LAST REFERENCE
IMAGE], at [POSITION IN FRAME], at the scale it would really be. Nothing about
the product changes: keep [COLOUR], [FINISH] and [LABEL ORIENTATION] exactly as
referenced. The camera makes one slow lateral move [DIRECTION] and stops.
[TIME OF DAY AND LIGHT]. Sound: [AMBIENCE], no music.
No second copy of the product anywhere in frame.
4 — THREE-BEAT FEATURE CALLOUT
viduq3-pro · text2video · 16s · 9:16
Three shots, hard cuts, no dissolves. Shot one, five seconds: a tight macro of
[FEATURE ONE] with the rest of the product out of focus behind it. Shot two,
five seconds: a mid shot of [FEATURE TWO] being used by a hand entering from
[SIDE]. Shot three, six seconds: a static wide of the whole product on
[SURFACE], nothing moving. [CONSISTENT LIGHT ACROSS ALL THREE].
Sound: [ONE SOUND PER SHOT], no voice-over, no music.
Keep the product identical in colour and finish across all three shots.
Talking-head, UGC and dialogue templates
Q3 makes speech from the prompt text, so write the line you want said, in quotes, with the delivery around it. Note that voice_id on img2video is labelled "Voice ID, The Q3 series model is not effective", so voice casting on Q3 is done in words, not by id.
5 — SINGLE LINE TO CAMERA
viduq3-turbo · img2video · 6s · frame from the source image
The person in the image looks straight down the lens and says, in a [AGE]
[ACCENT] voice, [PACE] and [MOOD]: "[ONE SENTENCE, UNDER FIFTEEN WORDS]."
One blink, a small natural head movement, hands stay out of frame. The camera
creeps in almost imperceptibly and never cuts. [AMBIENCE] behind, no music.
No gestures, no lean, no smile at the end.
6 — TWO-HANDER WITH NAMED SUBJECTS
viduq3-turbo · reference2video, subjects form · 12s · 16:9
@[NAME A] is [POSITION AND ACTION] when @[NAME B] enters from [SIDE].
@[NAME A], [DELIVERY]: "[LINE ONE]."
@[NAME B], [DELIVERY]: "[LINE TWO]."
Two-shot at chest height for the first line, then a single on @[NAME B] for the
second. [LIGHT SOURCE, PRACTICAL IF POSSIBLE]. Sound: their voices, [AMBIENCE],
no score.
Both faces stay matched to their reference images throughout.
7 — VERTICAL TESTIMONIAL WITH A BEAT CHANGE
viduq3-pro · text2video · 14s · 9:16
A [DESCRIPTION] person sits in [SETTING], framed waist up, slightly off centre
to the [SIDE]. For the first eight seconds they speak evenly to camera:
"[SETUP SENTENCE]." Then they pause, look away, and finish quieter:
"[PAYOFF SENTENCE]." No cut. Soft [KEY DIRECTION] key, practical lamp behind.
Sound: voice and room tone only, no music, no swell on the pause.
No captions, no lower third, no hand gestures.
8 — VOICE OVER ACTION, NO LIP SYNC
viduq3-pro · text2video · 10s · 16:9
[SCENE AND ACTION, WRITTEN AS IF SILENT]. No character's mouth is visible at any
point. Over the top, an unseen narrator says once, [DELIVERY]:
"[ONE SENTENCE]." The camera [MOVE]. [LIGHT].
Sound: the narrator, [TWO DIEGETIC SOUNDS], no music.
Nobody in frame speaks or reacts to the narration.
What do b-roll, texture and food templates need?
Short, silent, repeatable. These are the shots you will need twelve of, which is exactly why they should be templates.
9 — MACRO TEXTURE PASS
viduq3-turbo · text2video · 5s · 1:1
Extreme close on [MATERIAL], filling the frame edge to edge. [WHAT MOVES: DUST,
FIBRES, LIQUID, HEAT SHIMMER] moves slowly across it. The camera drifts
[DIRECTION] by a few centimetres and no more. Shallow focus, one plane sharp.
[LIGHT AT A RAKING ANGLE]. Sound: [ONE TEXTURAL SOUND] only.
No object enters or leaves the frame.
10 — HANDS-ONLY PROCESS SHOT
viduq3-pro · text2video · 9s · 4:3
Overhead, slightly off axis, on [WORK SURFACE]. Two hands complete [TASK] in
three unhurried movements: [MOVE ONE], [MOVE TWO], [MOVE THREE]. Never show a
face, a forearm above the elbow, or a clock. The camera does not move at all.
[LIGHT]. Sound: [THE SOUNDS THE TASK MAKES], no music, no voice.
Nothing enters the frame that was not already on the surface.
11 — POUR AND STEAM
viduq3-pro · text2video · 6s · 9:16
[LIQUID] pours from [VESSEL] into [RECEIVER] in one steady stream, filling it to
[LEVEL]. Steam rises and drifts [DIRECTION] out of the top of the frame. The
camera holds still at [HEIGHT], product centred. [BACKLIGHT SO THE STEAM READS].
Sound: the pour, then settling, no music.
No hand tremor, no splash outside the receiver, no text.
12 — PLATING REVEAL
viduq3-turbo · img2video · 8s · frame from the source image
The dish in the image is finished in place: [GARNISH] is placed with [TOOL] at
[POSITION], then [FINAL TOUCH]. The plate never moves. The camera lifts about
ten degrees to a flatter angle across the whole eight seconds. [LIGHT].
Sound: [PLATING SOUND], room tone, no music.
No steam, no hands after the fifth second, no second plate.
13 — AMBIENT ENVIRONMENT HOLD
viduq3-turbo · text2video · 7s · 16:9
A locked-off wide of [PLACE] with nobody in it. The only movement is [ONE
AMBIENT MOVEMENT: CURTAIN, LEAVES, SIGNAGE, WATER, TRAFFIC IN THE FAR
DISTANCE]. [TIME OF DAY, WEATHER, LIGHT]. The camera does not move, zoom or
rack focus. Sound: [LAYERED AMBIENCE], no music, no voice.
No person or animal appears at any point.
Transition, establishing and architecture templates
The first two are for start-end2video, where the two frames must be close in shape: the documented requirement is a ratio between start and end frame in the 0.8 to 1.25 window. Crop both before you upload.
14 — MATCH CUT ACROSS TWO FRAMES
viduq3-pro · start-end2video · 5s · frames from the input images
Move from the first frame to the second by keeping [THE MATCHING ELEMENT] in the
same position and scale the whole way through, and letting everything around it
change. No dissolve, no fade to black, no morph of faces. The camera stays
locked. Sound: [SOUND THAT BRIDGES BOTH SPACES], no music.
The matching element never leaves the centre of the frame.
15 — SCALE TRANSITION, MACRO TO WIDE
viduq3-pro · start-end2video · 8s · frames from the input images
Pull back continuously from the macro detail of the first frame to the wide of
the second, revealing [WHAT THE DETAIL BELONGED TO] at about the halfway point.
One unbroken move, constant speed, no cut. [LIGHT CONSISTENT ACROSS BOTH].
Sound: [ONE SOUND THAT GROWS AS THE FRAME WIDENS], no music.
Nothing new enters the frame during the pull back.
16 — ESTABLISHING REVEAL
viduq3-pro · text2video · 12s · 16:9
Open tight on [FOREGROUND DETAIL], hold for three seconds, then rise and pull
back to reveal [THE WIDER PLACE] behind it. End on a static wide. One continuous
move. [TIME OF DAY, WEATHER, HAZE, LIGHT DIRECTION]. Sound: [ONE FOREGROUND
SOUND] that gives way to [WIDER AMBIENCE], no music.
No cut, no people in the wide, no aerial view above roof height.
17 — INTERIOR WALK-THROUGH
viduq3-pro · text2video · 16s · 16:9
One continuous take at eye height moving forward through [BUILDING TYPE]: begin
in [ROOM ONE], pass through [THRESHOLD], continue into [ROOM TWO], stop facing
[FOCAL POINT]. Constant walking pace, no cuts, no doubling back. [NATURAL LIGHT
SOURCE AND DIRECTION], [MATERIAL PALETTE]. Sound: footsteps on [FLOOR
MATERIAL], [BUILDING AMBIENCE], no music.
No furniture changes between rooms, no people, no text on walls.
18 — FACADE LIGHT STUDY
viduq3-turbo · text2video · 10s · 4:3
A static frontal view of [BUILDING FACADE], filling the frame with a small
margin of sky. Over ten seconds the light moves from [START CONDITION] to [END
CONDITION] and shadows shift across [SPECIFIC FEATURE]. The camera never moves.
No time-lapse stutter; the change is smooth. Sound: [DISTANT AMBIENCE], no music.
No people, no vehicles, no birds crossing the frame.
Sports, motion and motion-graphic templates
Vidu does not document legible on-screen text as a capability on any video endpoint, so treat type as a post step and use these to generate the plate underneath it.
19 — SPEED CHANGE WITHOUT A SPEED FIELD
viduq3-pro · text2video · 8s · 16:9
[SUBJECT] performs [ACTION]. The first two seconds play at ordinary speed. At
the moment of [SPECIFIC INSTANT], the action slows to roughly a quarter speed
for three seconds, then resumes ordinary speed to the end. The camera [MOVE]
continuously and does not change speed with the subject. [LIGHT].
Sound: [SOUND] at ordinary pitch, then muffled during the slow section.
One take, no cuts, no ramping of the camera move.
20 — SPORTS FOLLOW SHOT
viduq3-pro · text2video · 12s · 9:16
Track [ATHLETE] in profile from [START POINT] to [END POINT] across [SURFACE],
holding them at the same size in frame the entire way. The camera moves parallel
to them at matching speed. Background passes behind at speed and stays out of
focus. [WEATHER AND LIGHT]. Sound: [BREATH OR IMPACT SOUNDS], crowd at low
level, no music.
No cut, no slow motion, no overtake by a second athlete.
21 — KINETIC PLATE FOR TYPE
viduq3-turbo · text2video · 6s · 1:1
Abstract motion field on a [COLOUR] background: [SHAPES OR MATERIAL] move
[DIRECTION] at an even pace, never fully covering the central third of the
frame. No literal objects, no logos, no letters or numbers anywhere. The camera
does not move. Even, flat lighting with no hot spot. Sound: silence.
Keep the central third clean enough to place text over later.
22 — LOOPING AMBIENT ABSTRACT
viduq3-turbo · text2video · 5s · 16:9
[MATERIAL: INK, SMOKE, SAND, FOIL, LIQUID] moves continuously in the centre of
the frame with no beginning or end to the movement. The first and last frames
should look nearly identical so the clip can be cut back to back with itself.
[SINGLE LIGHT SOURCE AND COLOUR]. The camera is locked. Sound: silence.
No hands, no vessel, no visible edge to the material.
What do you send alongside the prompt?
Two wrappers. The first shows every field worth setting on a text-driven Q3 call, including the two most template libraries ignore.
{
"model": "viduq3-pro",
"prompt": "PASTE TEMPLATE 17 HERE, FULLY FILLED IN",
"duration": 16,
"aspect_ratio": "16:9",
"resolution": "1080p",
"audio": true,
"seed": 41207,
"off_peak": true,
"payload": "tpl:interior-walkthrough v4 | client:northgate | take:2"
}
seed and payload are the two that make a template library work rather than just exist. Seed is documented as defaulting to a random number, with manually set values overriding it, so pinning one lets you change a single slot and see only that change. payload is documented on every video endpoint as a transparent transmission parameter that receives no processing, is limited to 1048576 characters, and comes back on the response. Nothing else in a Vidu response records which template produced the clip, so if you generate at any volume, put the template name and version in there.
The second wrapper is the subjects form, which is the only place Vidu's at-sign handles work. The name field is documented as "Usable in prompts via @subjectname", and Vidu's own worked example uses positional ids instead, with the prompt "@1 and @2 are cooking together, and both say they love hot pot."
{
"model": "viduq3-turbo",
"subjects": [
{
"name": "rafa",
"images": [
"https://example.com/rafa-front.jpg",
"https://example.com/rafa-profile.jpg"
],
"voice_id": ""
},
{
"name": "ledger",
"images": ["https://example.com/ledger-closed.jpg"]
}
],
"prompt": "PASTE TEMPLATE 6 HERE, WITH @rafa AND @ledger IN THE HANDLE SLOTS",
"duration": 12,
"aspect_ratio": "16:9",
"resolution": "1080p",
"audio": true,
"payload": "tpl:two-hander v2"
}
One honest gap. Each subject may carry its own voice_id, and unlike the img2video version of that field, the subjects form does not repeat the note that it is ineffective on Q3. Whether subject-level voice casting works on a Q3 model is not published either way, so the example leaves it empty rather than implying it does.
Two more caps worth knowing before you batch. Each subject supports up to three images, and the total across the array is capped, phrased as "The maximum number of images or textual content should not exceed 7". And Vidu's usage page states that "Each organization can use up to 5 concurrent tasks.", per organization rather than per key, so a twenty-variation sweep of one template is four waves, not one.
What do you change when a template underperforms?
In this order, one change at a time, with the seed pinned.
- Cut the cut plan. If the output arrives as a montage you did not ask for, the fix is the sentence "one continuous take, no cuts" rather than a shorter duration. Q3's own scene-cutting is on by default.
- Move the exclusion to the end. The last line of a Vidu prompt is where the forbidden list belongs. Buried mid-paragraph it competes with the action.
- Name the sound explicitly, including silence. A missing sound clause is not a request for silence on Q3; it is a request for whatever the model thinks fits.
- Shorten the dialogue. A line that cannot be said comfortably in the duration you set will be rushed or truncated. Roughly two and a half words per second is a safe ceiling.
- Check the endpoint, not the prompt. A template that produces a bizarre crop is usually a template written with an
aspect_ratioline running against img2video or start-end2video, neither of which has that field. - Drop the resolution before you drop the idea. Billing is per second and per resolution tier, so iterate at 540p on
viduq3-turboand only spend 1080p onviduq3-proonce the shot is right.
If the problem is morphing rather than framing, that is a different failure with different fixes, and the same templating logic applied to other models sits in our JSON video prompt templates for Veo 3 and the Grok Imagine Video 1.5 templates. The camera clause in every skeleton above draws on our camera movement vocabulary.
Where a prompt library fits in this
Prompt Architects does not generate video. It stores and versions the prompt template, which on Vidu is the part that keeps breaking.
A Vidu template is not one string. It is a prose skeleton, a set of slots, an endpoint, a model string, a duration and an aspect_ratio that may not exist on the endpoint you moved it to. Twenty-two of those, edited by two people across four client folders, is a versioning problem long before it is a creative one. The specific failure it prevents is small and expensive: shot nine of a sequence quietly loses the "no on-screen text" line that shots one through eight carried, and nobody notices until the edit. Video prompt generation sits on the Advanced and Team plans; the library itself is on every tier including the free one, which the FAQ page describes as 5 prompt enhancements per day, forever.
Everything above was read off platform.vidu.com/docs on 27 August 2026. Vidu's Update Notice records three separate Q3 launches between January and April 2026, and its own FAQ still answers the duration question with 4 and 8 seconds. Date your templates, and re-read the endpoint page before you scale a batch.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An Account