TL;DR: Vidu publishes no prose prompting guide for Q3. It publishes a request schema, and that schema is the guide. Pick a Q3 model string, put duration, resolution and aspect ratio in the body, and write motion, camera and dialogue as prose, because movement_amplitude, style and bgm are all inert on Q3.
Vidu prompting is unusual in one specific way. On most video models the prompt is where you fight for control. On Vidu Q3 half the controls you would reach for are switched off at the model level, and the rest live in the JSON around the prompt rather than inside it. Everything below comes from platform.vidu.com/docs, read on 27 August 2026.
Which Vidu Q3 model are you actually prompting?
There is no single Q3. Vidu's Model Map lists five members of the ViduQ3 Series, and a sixth string turns up in an endpoint enum without appearing there at all.
| Model string | Where it appears | Documented duration | Resolutions |
|---|---|---|---|
viduq3-pro | Model Map, text2video, img2video, start-end2video, pricing | 1–16s | 540p, 720p, 1080p |
viduq3-mix | Model Map, reference2video, pricing | 1–16s (pricing says 3–16s) | 720p, 1080p |
viduq3-turbo | Model Map, all four video endpoints, pricing | 1–16s; 3–16s on reference2video | 540p, 720p, 1080p |
viduq3-drama | Model Map only | 3–15s | 720p, 1080p |
viduq3-ad | Model Map only | 3–15s | 720p, 1080p |
viduq3-pro-fast | img2video enum and pricing, not the Model Map | 1–16s | 720p, 1080p |
viduq3 | reference2video enum and pricing | 3–16s, per the pricing page | 540p, 720p, 1080p |
Two are worth pausing on. viduq3-drama is described as a "Comic/drama industry model" with the "Best quality in Q3 series; strong dialogue, character positioning, motion effects, and cinematic storytelling". viduq3-ad is an "Advertising industry model" that is "suited for 5–8s short-form ads". Both sound like exactly what you want. Neither appears in any endpoint's accepted-values list, nor on the pricing page. How to call them is not published, so treat them as announced rather than available.
| Feature | viduq3-pro | viduq3-mix | viduq3-turbo |
|---|---|---|---|
| Listed in text2video model enum | |||
| Listed in img2video model enum | |||
| Listed in start-end2video model enum | |||
| Listed in reference2video model enum | As a bare viduq3 | ||
| Documented resolutions | 540p / 720p / 1080p | 720p / 1080p | 540p / 720p / 1080p |
| Named subjects (subjects array) | Not listed | Explicitly unsupported | |
| Off-peak mode on reference2video | n/a |
The rollout explains the mess. Vidu's Update Notice dates viduq3-pro to 27 January 2026 for text and image inputs, viduq3-turbo to 11 February 2026 across three endpoints, and the reference2video trio to 13 April 2026: "Support the addition of viduq3-mix、viduq3-turbo、viduq3 model for reference2video". Three launches, three enums, one Model Map trying to summarise all of it.
Does Vidu publish a Vidu Q3 prompting guide?
No, and it is worth saying so plainly rather than inventing a framework and attributing it to them. There is no prompt-writing tutorial anywhere in the developer docs. What exists instead is machinery for not writing the prompt yourself.
Vidu ships a Prompt Recommendation endpoint that takes one image and returns between 1 and 10 suggested prompts, tagged either img2video or template. Both img2video and start-end2video also carry an is_rec boolean; set it and, in Vidu's words, "the system will automatically generate and apply a recommended prompt to create the video", at a documented cost of an additional 10 credits per task. Note the trade: the same field description warns that if is_rec is used, "the model will ignore the manually entered prompt".
That is the closest Vidu comes to house style. Its FAQ is blunter still, answering whether an optimised prompt can be returned with "Currently, returning an optimized result for the prompt is not supported." So the documented structure worth teaching is the request body, and the discipline is knowing which fields the model ignores.
What belongs in the prompt and what belongs in the request body?
Shape goes in the body. Content goes in the prompt. Vidu Q3 makes the split sharper than most models, because it disabled the two fields that would otherwise straddle the line.
The movement_amplitude field takes auto, small, medium or large, and on three separate endpoint pages it carries the same note: "This parameter does not take effect when using the q2 & q3 model". The style field on text2video, with its general and anime options, says "This parameter is not effective when using q2 or q3 models". Neither is deprecated. Both accept your value, return it in the response body, and change nothing.
So on Q3 there is no motion dial and no style toggle. Pacing, camera behaviour and visual register are prose problems now. If you want a slow push-in, you write a slow push-in. Our camera movement vocabulary for AI video is the working list for that, and the seven-part video prompt anatomy is the shape most of these prompts end up taking.
Prompt length is generous. Every video endpoint gives the same limit: "A textual description for video generation, with a maximum length of 5000 characters". The FAQ page disagrees, and we will come back to that.
How long can a Vidu Q3 clip be?
One to sixteen seconds, defaulting to five, on three of the four endpoints. On text2video, img2video and start-end2video the duration line reads "viduq3-pro , viduq3-turbo : default 5s, available: 1 - 16". The reference2video endpoint is stricter: "viduq3-turbo、viduq3-pro: Default is 5 seconds, available option: 3-16", with viduq3-mix alone documented as "Default is 5 seconds, available option: 1-16".
Sixteen seconds is a real number, not a stitched one, and it is why Q3 is worth the trouble. Vidu Q2 tops out at 10 seconds on text and reference inputs and 8 on first-and-last-frame; Vidu Q1 accepts exactly one value, 5. For that comparison across the field, we keep a video duration parameter table by model.
Duration is billed per second, so a 16-second draft at 1080p is not a cheap way to test an idea. Vidu documents an off_peak boolean for this: it "consumes lower points", and "Tasks submitted in off peak mode will be generated within 48 hours". Unfinished off-peak tasks are cancelled and refunded automatically. One caveat the pricing page adds and the endpoint pages do not: off-peak is listed as Not Supported for viduq3-mix on reference2video.
How do you get sound out of Vidu Q3?
With the audio boolean, and only with it. This is the biggest behavioural change from Q2, and the docs are unusually direct about it.
The old bgm flag is dead on this series. Its description on text2video ends with "BGM does not take effect when the duration of the q2 model is 9 or 10 seconds;BGM does not available in q3 models". The reference2video version says "q3 model does not support this parameter". If you have a Q2 wrapper that sets bgm: true, that flag stops doing anything the moment you point it at Q3.
What replaces it is broader. audio is documented as "Whether to use direct audio-video generation capability", and its note reads "Only the q3 models supports this parameter". Set to true it "Outputs video with generated speech and background music based on the prompt". Speech, effects and score, all from the prompt text, all in one pass.
Two related fields do not help you on Q3. audio_type, which splits vocals from sound effects, notes that it "only supports audio splitting for q2, q1, and 2.0 series models". And voice_id, which pins a specific voice, is labelled "Voice ID, The Q3 series model is not effective" on img2video. So on Q3 you cannot choose the voice by id and you cannot ask for effects without dialogue. You describe the voice in the prompt instead: age, accent, delivery, volume.
How do you keep one subject consistent across a Vidu Q3 clip?
This is the part of Vidu with no close equivalent elsewhere, and it lives on reference2video, which publishes two request shapes at one URL.
The plain form takes an images array. The doc states that "viduq3-mix viduq3-turbo viduq3 viduq2 viduq1 vidu2.0 accepts 1 to 7 images", each at least 128x128, under 50MB, with an aspect ratio between 1:4 and 4:1. Seven angles of one jacket, or three characters and four props, held across the shot.
The named form is the interesting one. Instead of images you pass a subjects array, where each entry has a name and up to three images. Vidu's note on that name field is one line long and does all the work: it is "Usable in prompts via @subjectname". The prompt then addresses cast members by handle rather than by description. Vidu's own worked example uses positional ids instead: "@1 and @2 are cooking together, and both say they love hot pot."
Three constraints on that form, all documented. The subjects cap is the same 7, phrased as "The maximum number of images or textual content should not exceed 7". Each individual subject supports up to 3 images. And there is a hard exclusion: "Note: viduq3-mix does not support the use of entities for the time being". The named-subject enum is viduq3-turbo, viduq3, viduq2, viduq1 and vidu2.0, so if you want handles, viduq3-turbo is the Q3 model to reach for.
That covers consistency inside one generation. Holding a face steady across a whole sequence of separate clips is a different problem with different tactics, and we wrote it up separately in keeping a character consistent across video clips.
What resolution and aspect ratio can Vidu Q3 produce?
1080p at 24fps, at most. Every Q3 row on the Model Map reads 24fps, and no Q3 row lists anything above 1080p. If you need more, that is a second call: the Upscale-pro endpoint takes 1080p, 2K, 4K or 8K as upscale_resolution, requires the target to be higher than the source, and accepts videos up to 300 seconds and under 60fps.
Aspect ratio is where the endpoints diverge again:
text2videoaccepts16:9,9:16,3:4,4:3and1:1, with the note "3:4 & 4:3 only support q2 & q3 model". Q3 gets the full set.reference2videodocuments16:9,9:16and1:1in one section and adds "q2、q3 model supports any aspect ratio", then in the other section lists all five with "3:4&4:3only support q2 model".img2videoandstart-end2videohave noaspect_ratiofield at all. The frame comes from your input image. Crop before you upload, not after.
start-end2video adds one more rule that has nothing to do with the prompt: the two frames must be close in shape, because the "ratio between start/end frame must be in 0.8~1.25". What happens if you break it is not published, so match your crops.
Copy-paste Vidu Q3 prompts
Full request bodies first. Swap the image URLs and the key, and these run as written.
1. Text to video, 16 seconds, with sound
{
"model": "viduq3-pro",
"prompt": "A rain-slicked Tokyo side street at 2am, neon reflected in standing water. A courier in a yellow shell jacket wheels a bike into frame from the right, stops, and looks up at a shuttered ramen counter. She says, quietly and in Japanese, 'Closed again.' Slow push-in from a low angle over eleven seconds, then a hard cut to a wide shot of the empty street from the far end. Ambient rain, distant traffic, no music until the cut.",
"duration": 16,
"aspect_ratio": "16:9",
"resolution": "1080p",
"audio": true,
"seed": 20260827,
"off_peak": false
}
2. Image to video, cheap draft pass
{
"model": "viduq3-turbo",
"images": ["https://example.com/frame-01.jpg"],
"prompt": "The subject exhales, shoulders dropping, and turns her head slowly to camera left. The handheld camera drifts three degrees right and settles. Nothing else in the frame moves. Room tone only, no speech.",
"duration": 6,
"resolution": "540p",
"audio": false,
"off_peak": true
}
Note what is absent: no aspect_ratio, because img2video has none, and no movement_amplitude, because it does nothing here.
3. Reference to video with named subjects
{
"model": "viduq3-turbo",
"subjects": [
{
"name": "mara",
"images": [
"https://example.com/mara-front.jpg",
"https://example.com/mara-three-quarter.jpg",
"https://example.com/mara-profile.jpg"
]
},
{
"name": "kettle",
"images": ["https://example.com/copper-kettle.jpg"]
}
],
"prompt": "@mara sets @kettle down on the hob, wipes her hands on her apron and says in a low, tired English accent, 'Give it five minutes.' She stays in frame the whole time. Static camera at chest height, shallow depth of field, late afternoon window light from the left. Kettle ticks as it heats. No music.",
"duration": 8,
"aspect_ratio": "16:9",
"resolution": "1080p",
"audio": true
}
4. Reference to video from loose images, best-balance model
{
"model": "viduq3-mix",
"images": [
"https://example.com/product-front.png",
"https://example.com/product-angle.png",
"https://example.com/product-detail.png",
"https://example.com/kitchen-counter.png"
],
"prompt": "The grinder from the reference images sits on the marble counter from the fourth image. A hand enters from the right, twists the hopper a quarter turn, and withdraws. Cut to a macro shot of beans dropping into the burr. Cut back to a wide static shot. Warm morning light, no on-screen text, no people in shot beyond the hand.",
"duration": 10,
"aspect_ratio": "16:9",
"resolution": "720p"
}
5. First and last frame
{
"model": "viduq3-pro",
"images": [
"https://example.com/open-door.jpg",
"https://example.com/empty-room.jpg"
],
"prompt": "The door swings shut over four seconds as the camera holds still. Light narrows across the floorboards until the room is lit only by the window. Hinge creak, then a soft latch click, then silence.",
"duration": 6,
"resolution": "1080p",
"audio": true
}
Thirteen more prompt bodies, each written for a named model, endpoint and duration. Drop them into the prompt field above.
6. viduq3-pro / text2video / 12s / 9:16
A single unbroken handheld shot following a chef's hands through a plating sequence: tweezers place three leaves, a spoon drags sauce in an arc, a thumb wipes the rim. Never show a face. Overhead, slightly off axis. Kitchen clatter and one voice calling an order in the background.
7. viduq3-pro / text2video / 16s / 16:9
Three shots, cut hard. Shot one, five seconds: a wide of a wheat field at golden hour, wind moving through it left to right. Shot two, five seconds: a mid shot of an old man in a linen shirt, watching, saying nothing. Shot three, six seconds: extreme close on his hand closing around a fistful of grain. Wind and cicadas throughout. No score.
8. viduq3-turbo / text2video / 5s / 1:1
A cast-iron pan on a gas flame, seen from directly above. Butter hits the surface and foams instantly. Steam rises into the top of the frame. Locked-off camera, no movement at all. Sizzle only.
9. viduq3-turbo / img2video / 8s / from source image
The person in the image begins speaking to camera, warm and unhurried, in a mid-Atlantic accent: 'We rebuilt the whole thing in six weeks.' Small natural head movement, one blink, hands stay out of frame. The camera creeps in almost imperceptibly. Quiet office ambience behind.
10. viduq3-pro / img2video / 14s / from source image
Hold on the still image for two seconds, then let the scene come alive: dust in the light shaft, a curtain lifting, a cat crossing the background from left to right and out. The camera arcs slowly clockwise around the room without ever leaving the space. No dialogue, no music, only house sounds.
11. viduq3-turbo / reference2video (subjects) / 10s / 9:16
@host holds @bottle at chest height, turns it label-out to camera, and says brightly, 'It's the same formula, half the plastic.' Then she sets it down and steps back. Vertical framing, waist up, soft key from the front, plain seamless backdrop. Upbeat room tone, no backing track.
12. viduq3-turbo / reference2video (subjects) / 16s / 16:9
@detective enters the room, crosses to the desk, and picks up @folder. @rookie stays by the door. Detective, gravelly and low: 'You read this yet?' Rookie, defensive: 'Twice.' Two-shot, then push in on the folder. Practical lamp light only, heavy shadow, faint rain outside the window.
13. viduq3-mix / reference2video / 12s / 16:9
Assemble the subjects from the reference images into a single continuous scene: the bicycle leans against the wall from image three, the jacket from image two hangs on the handlebars, the street from image four is empty behind it. Slow lateral dolly right. Late autumn light, leaves moving. Street ambience, no music.
14. viduq3-mix / reference2video / 6s / 1:1
Keep the character identical to the reference images across a two-shot cut: first a tight close on the eyes, then a wide of the same character standing in a doorway. Match the lighting between both. No dialogue. One low sustained tone under the cut.
15. viduq3-pro / start-end2video / 4s / from frames
Transform the first frame into the second by dissolving the ink on the page upward into a flock of birds. Camera drifts up with them and settles on the empty desk. Paper rustle, then wingbeats, then nothing.
16. viduq3-pro / start-end2video / 10s / from frames
Move from the packed stadium of the first frame to the empty one of the second by way of a time-lapse: crowd thins in waves, floodlights cut out bank by bank, shadows swing across the pitch. Locked camera. Crowd noise decaying to wind.
17. viduq3-turbo / text2video / 16s / 9:16
Six seconds of a phone screen being scrolled by a thumb, shot over the shoulder. Cut. Six seconds of the same person outdoors, phone in pocket, walking fast. Cut. Four seconds of a static wide of an empty desk. No speech at all. Sound design only: notification chimes, then footsteps, then room tone.
18. viduq3-pro / text2video / 16s / 16:9
An anamorphic-style single take through a night market: start tight on a wok, pull back and track right past three stalls, end on a wide of the street with the vendor calling out to the camera in Cantonese. Practical lantern light, shallow focus, visible flare. Layered market noise, one voice on top.
Where do Vidu's own docs contradict each other?
In four places we could verify, all on platform.vidu.com, read on 27 August 2026. Trust the endpoint reference over the general pages when they clash: the endpoint tables are model-scoped and the general pages are not.
Prompt length. Every video endpoint says "with a maximum length of 5000 characters". The FAQ says "The maximum supported input length is approximately 1,500 characters."
Duration and resolution. That same FAQ answers "What video generation durations does Vidu support?" with "Vidu supports two video generation durations: 4 seconds and 8 seconds." and lists three resolutions, "360p, 720p, and 1080p". Those are Vidu 2.0 numbers. Q3 does 1 to 16 seconds and offers 540p, which the FAQ never mentions.
viduq3-mix duration. The reference2video page gives "viduq3-mix: Default is 5 seconds, available option: 1-16". The pricing page's Q3 table lists Vidu Q3-mix on reference2video as 3-16S.
The audio field on reference2video. The named-subjects request body documents audio and audio_type. The plain images request body on the same page documents neither, yet its response body returns audio. Send it and it may well work; the doc does not promise it.
One more oddity that is not a contradiction but will waste your afternoon: the Prompt Recommendation page prints its endpoint as https://api.vidu.cn/ent/v2/img2video-prompt-recommendation, on the .cn domain, while every other endpoint page prints api.vidu.com.
What breaks Vidu Q3 prompts most often?
Setting bgm and expecting music. It is a no-op on Q3. Use audio.
Assuming a field that returns in the response was honoured. movement_amplitude and style both echo back in the response body on Q3 and both do nothing. The response is a record of what you sent, not proof it mattered.
Reusing a model string across endpoints. viduq3-mix is a reference2video model. viduq3-pro-fast is an img2video model. Neither travels.
Reaching for a negative prompt. There is no negative-prompt field on any Vidu video endpoint. Exclusions have to be phrased positively inside the prompt, and the general reliability of that approach is what we mapped in the negative prompt support matrix.
Asking for 16 seconds of continuous action. Q3's headline feature is smart scene cuts, so a long single prompt will often be interpreted as several shots. If you want one take, say one continuous take. If you want cuts, write the cuts. Morphing between them is a separate failure, covered in why AI video morphs and warps.
Mismatched start and end frames. The 0.8 to 1.25 ratio window on start-end2video is stated as a requirement, not a hint.
Queueing more than five jobs and wondering why they crawl. Vidu's usage limits page states that "Each organization can use up to 5 concurrent tasks." Extras queue in submission order. The FAQ also claims support for input "in more than 200 languages worldwide", so language is rarely the bottleneck; concurrency usually is.
Where a prompt tool fits, and where it does not
Prompt Architects does not generate video. It generates the prompt template, and then keeps it. That distinction matters more on Vidu than on most models, because a Q3 prompt is not really one string. It is a shot list, a cast block of @handles, a lighting clause and a sound clause, wrapped in a JSON body whose fields differ per endpoint. That is a structure with variables in it, not a sentence you retype.
The failure mode this fixes is not creative. It is that clip nine of a series quietly loses the "no on-screen text" clause that clips one through eight carried, and nobody notices until the edit. Video prompt generation sits on the Advanced and Team plans, while the prompt library itself is on every tier, including the free one, which the FAQ page describes as 5 prompt enhancements per day, forever.
Vidu's own docs will keep moving. Three Q3 launches in three months, two models on the map with no way to call them, and an FAQ still describing 2026 as a four-and-eight-second world. Save the request shape, date it, and re-read the endpoint page before you scale a batch.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An Account