Back to blog
Video18 min read

AI Video Sound Design: Prompts for Veo, Seedance, Kling and Vidu

What each AI video model documents about audio: Veo's cue labels, Seedance's brackets, Kling's off-by-default switch, Vidu Q3's audio boolean, plus 20 copy-paste sound prompts.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: AI video sound design is now a prompting problem, not a post-production one, and every model handles it differently. Veo 3.1 has audio always on with three documented cue labels. Seedance uses brackets. Kling defaults to silent. Vidu Q3 replaced bgm with a boolean. The craft transfers; the syntax does not.

Sound is the half of AI video that most prompts leave to chance. There are a hundred guides on lens choice and camera movement for these models, and almost nothing on what to write when you want a door to close convincingly. Which is odd, because audio is one of the few parts of a video prompt where the vendors have written the rules down. Everything below comes from vendor-owned documentation read on 27 August 2026.

Which models generate audio by default, and can you turn it off?

The single most useful fact in this whole area is that the default state is not consistent. Four vendors, four different contracts, and one of them is the opposite of what everyone assumes.

Google Gemini API Veo page, BytePlus ModelArk create-task reference and Seedance 2.5 tutorial, Kling 3.0 Omni API reference, Vidu text2video and img2video references. All read 27 August 2026.
FeatureVeo 3.1Seedance 2.5 / 2.0Kling 3.0 & OmniVidu Q3
Audio field in the request bodyNonegenerate_audiosettings.audioaudio
Default stateAlways ontrueofftrue (varies by endpoint)
Can you switch audio off?
Documented cue labels for the promptDialogue, SFX, Ambient NoiseBrackets per sound typeNot documentedNot documented
Separate endpoint to add audio later
Output channel count publishedMono

Google's model-features table gives Veo 3.1, Veo 3.1 Fast and Veo 3.1 Lite the row "Audio: Natively generates audio with video." marked "Always on", against Veo 2's "Silent only". There is no audio parameter in the Veo request body at all. The only knob is your prose.

Kling is the surprise. Its 3.0 text-to-video reference lists settings.audio as a string with a default of off and an enum of native or off, described simply as "Whether to generate audio for the video". A default Kling 3.0 request returns a silent clip. On the Omni endpoint the enum gains a third value, original, which means "The generated video retains the original sound of the reference video." Plenty of people conclude Kling has no audio because they never set the field.

Vidu's default depends on which endpoint you call. On text2video and start-end2video the audio boolean reads "Default: true" with the note "Only the q3 models supports this parameter". On img2video the same field reads "Default: false", with a separate note that "When model is q3, the default value for this parameter is true". On reference2video it says "By default, audio is set to true for viduq3 and viduq3-turbo, and false for other models". Three phrasings of roughly one rule. Set it explicitly.

What audio syntax does each vendor actually document?

This is the part that does not transfer. Copy a Seedance bracket into a Veo prompt and you get brackets read as scene description.

Veo 3.1: three cue labels and separate sentences

Google's Gemini API Veo page says "You can provide Veo with cues for sound effects, ambient noise, and dialogue" and that "The model captures the nuance of these cues to generate a synchronized soundtrack." It names exactly three: Dialogue in quotation marks, "Sound Effects (SFX): Explicitly describe sounds." and "Ambient Noise: Describe the environment's soundscape."

The placement rule lives on a different Google page. The Gemini Enterprise video prompt guide says "Clearly specify if you want audio. We recommend that you use separate sentences in your prompt to describe the audio." Google's Cloud blog for Veo 3.1 then adds literal inline labels, which the Gemini API page does not use, and shows them inside timestamped shot blocks.

A woodworker planes a board in a small shop, late afternoon light through
sawdust. Ambient noise: a radio playing faintly two rooms away.
SFX: the long shearing hiss of a hand plane, a shaving curling away.
A rain-soaked fire escape at 3am, viewed from the alley below. A woman in a
wet coat climbs down, one rung at a time, favouring her left hand.

Ambient noise: steady rain on metal, a distant siren two blocks over,
water running into a drain.
SFX: the flat clang of a boot on wet steel, once per step.
[00:00-00:03] Wide shot of an empty station platform at dawn, camera static.
Ambient noise: low platform hum, wind through the tunnel mouth.
[00:03-00:06] Medium shot as a train brakes into frame.
SFX: rail screech rising, then air brakes releasing.
[00:06-00:08] Close on the doors. SFX: a two-tone chime, then doors sliding open.
A cramped record shop, afternoon light through dusty glass. Two collectors
face each other across a crate.
Older man: "You're holding the only clean copy in the city."
Younger man: "Then you already know what I'm asking for it."
Ambient noise: the faint crackle of a record playing in the back room.

For what you can and cannot control inside Veo's mix specifically, see our Veo audio prompts guide.

Seedance: brackets per sound type

ByteDance publishes the most explicit sound syntax of any vendor here. Its Seedance 2.5 tutorial, under the heading "Use special characters to distinguish sounds", assigns one bracket type per sound: () for music, <> for sound effects, {} for dialogue, and 【】 for subtitles. It adds that "For non-Chinese dialogue, it is recommended to specify the language before the dialogue."

A night market stall, steam rising off a wok. The cook plates a dish and
slides it across the counter.
(warm lo-fi guzheng loop, low in the mix)
<oil hissing, metal spatula on steel, a cleaver hitting the board twice>
{English: "Careful. That plate is hotter than it looks."}
【Careful. That plate is hotter than it looks.】

The 2.0 series prompt guide organises prompts by shot instead, and lists audio as a per-shot element: for each shot you give camera movement, subject action, spatial position and audio, where audio means "describe the sound effects, voices, background music, etc. corresponding to the shot."

Shot 1: Fixed camera on a workshop bench. Hands lift a cracked ceramic bowl
into frame. <a soft ceramic clink as the pieces meet> (no music)
Shot 2: Slow push-in on the join as gold lacquer is drawn along the crack.
<the fine scrape of a brush, breathing just audible>
Shot 3: The bowl is turned to catch the light; the camera holds.
(a single low string note enters and sustains)

Negative audio control is documented too. The 2.5 prompt guide says "Supports negative audio control for finer dimensions, including sound effects, background music (BGM), and dialogue", with its own examples reading "No BGM; generate only environmental sounds and action sounds." and "No audio."

Render Video 1. No BGM; generate only environmental sounds and action sounds.
Keep it subtitle-free.

Version matters here. ByteDance states that "Seedance 2.0 does not respond to timestamps and only responds to shot numbers, while Seedance 2.5 supports integer-second timestamps", and its 2.5 guide gives the bracket form below while warning you to "Pay attention to timeline continuity" and leave no gaps.

[0s-4s] A courier freewheels down a wet hill at dusk.
<tyres on wet asphalt, a freewheel ticking>
[4s-8s] She brakes hard at a junction and puts a foot down.
<a rim brake squeal, one shoe scuffing grit>
(no music)

The generate_audio field itself carries one prompt-level instruction most people miss: "It is recommended to put dialogue content in double quotes to optimize the audio generation effect." Our Seedance prompting guide covers the rest of the request shape.

Kling: a switch, then two separate audio prompts

Kling documents no in-prompt sound vocabulary at all. What it documents instead is architecture. Turn settings.audio to native and you get "native audio matching the visuals" generated from the same prose. Or leave the video silent and score it afterwards through a dedicated endpoint.

{
  "prompt": "shot 1, 4, a beekeeper lifts a frame from a hive, bees drifting around the veil; shot 2, 4, close on gloved hands scraping wax from the comb",
  "settings": { "audio": "native", "duration": 8, "resolution": "1080p" }
}

That video-to-audio endpoint forces the separation this craft actually needs. It takes sound_effect_prompt and bgm_prompt as two different fields, each capped at 200 characters, plus an asmr_mode boolean that "enhances detailed sound effects and is suitable for highly immersive content scenarios". Kling's capability map says it "Supports adding audio to all videos generated by Kling models and user-uploaded videos".

{
  "video_url": "https://example.com/beekeeper.mp4",
  "sound_effect_prompt": "close, dry bee hum; the woody knock of a frame lifted from a hive; gloves brushing canvas",
  "bgm_prompt": "sparse acoustic guitar, no percussion, very low in the mix",
  "asmr_mode": false
}
{
  "prompt": "a heavy oak door closing in a stone corridor, then the latch dropping",
  "duration": 3.5
}

Two caveats from the same docs. On the Omni endpoint, when you supply a feature reference video, "Native audio generation is not supported and the audio parameter can only be off in this case." And a 200-character cap on each audio prompt is not a limitation to work around. It is a hint.

Vidu Q3: bgm is gone

Vidu removed a parameter rather than adding one. The bgm boolean still appears in the request tables, annotated across endpoints with "BGM does not available in q3 models" and, on reference2video, "q3 model does not support this parameter". In its place sits audio, described as "Whether to use direct audio-video generation capability", where true "outputs video with sound (including dialogue and sound effects)".

There is no prompt vocabulary for sound. Write it as prose, in the same paragraph as everything else.

{
  "model": "viduq3-turbo",
  "prompt": "A locksmith works a stubborn lock in a quiet stairwell. Metal picks tick against the pins, a neighbour's television murmurs through a wall, and the lock finally gives with a dull thunk. No music.",
  "duration": 8,
  "audio": true
}
{
  "model": "viduq3-pro",
  "prompt": "Waves break against a harbour wall in slow motion, spray catching low sun. Only water and wind. No score, no voices.",
  "duration": 8,
  "audio": false
}
{
  "model": "viduq3-turbo",
  "prompt": "A watchmaker seats a balance wheel under a loupe. Tweezers tick against brass, a case-back is set down on felt, and a clock in the corner keeps time. She says, \"Almost.\"",
  "duration": 8,
  "audio": true
}

The Vidu Q3 prompting guide has the full variant map. One extra note from the schema: "The voice_id parameter takes effect only when this parameter is true", so voice selection is gated behind the audio switch.

Sora 2: two forms, both from OpenAI

OpenAI's Sora 2 prompting guide contradicts itself in the space of one screen. The prose tells you to place dialogue in an XML-style dialogue block "so the model clearly distinguishes visual description from spoken lines". The worked example immediately below uses a plain Dialogue: label with a bulleted speaker list, and the reusable template later on the same page uses Dialogue: and Background Sound: as plain labels too. Both are OpenAI's own. The label form is the one the examples actually use.

A cramped, windowless room with walls the colour of old ash. A bare bulb
pools light onto a scarred metal table. Two people face each other across it.

Dialogue:
- Detective: "You already told me where you were."
- Suspect: "And you already decided you didn't believe me."

Background Sound:
The hum of a failing fluorescent tube, a chair leg scraping once, rain on
a high window. No score.
Sound
Diegetic only: faint rail screech, train brakes hiss, distant announcement
muffled, low ambient hum. Footsteps and paper rustle; no score or added foley.
Mix: prioritize train and ambient detail over footstep transients.

One flag before you build on it: OpenAI's deprecations page removes the Videos API, sora-2 and sora-2-pro on 24 September 2026, recommended-replacement column blank. See our Sora migration guide.

Why does naming the source and distance beat naming the mood?

Because a model can render a source. It cannot render an adjective. "Tense atmosphere" gives the model nothing to synthesise; "a chair leg scraping once, then nothing" gives it an object, an action and a duration.

The pattern is unmissable in the vendors' own examples. Google's are "tires screeching loudly, engine roaring" and "A faint, eerie hum resonates in the background". OpenAI's production-grade example reads "faint rail screech, train brakes hiss, distant announcement muffled". Each names a physical source, and most also encode distance: faint, distant, muffled, in the background. Distance is the cheapest mix control you have, and it works on every model here because it is just adjectives attached to a real object.

The rewrite discipline is mechanical. Take each mood word and ask what is making that noise, how far away it is, and how often it happens.

Before: eerie, unsettling atmosphere, ominous soundscape
After:  a refrigerator compressor cycling on and off in the next room;
        one floorboard settling; no traffic, no voices, no music

How do you time a sound to an action?

By putting the sound inside the same time block as the action, using whatever segmentation the model accepts. This is where the models genuinely differ.

Google's Cloud blog demonstrates [00:00-00:02] timestamp blocks with SFX: lines dropped inside them. Seedance splits by version, so the same prompt shape is right on 2.5 and wrong on 2.0. Kling takes a third route with a documented multi-shot string, "shot n, m, words; shot n, m, words;", where m is the shot duration in seconds and the durations must sum to the total.

shot 1, 3, a hand hovers over a piano key, the room silent apart from a
clock ticking; shot 2, 2, the key is struck and the note rings out;
shot 3, 3, the note decays as the camera pulls back through the doorway

The general rule underneath all three syntaxes: one clearly named sound per beat, anchored to a visible action in the same clause. A sound with no visible cause is where sync failures come from.

When is silence the right instruction?

More often than people expect, and it is an instruction rather than the absence of one. OpenAI puts it best: "If your shot is silent, you can still suggest pacing with one small sound", and "Think of it as a rhythm cue rather than a full soundtrack."

Three routes, by model. Seedance takes negative audio control by category, so you can kill the score and keep the room. Kling and Vidu take a hard switch. Veo 3.1 has no off switch at all, so silence there means describing a genuinely quiet space and muting the track in an editor if it matters.

An anechoic chamber, one figure standing in the centre. Only their own
breathing, and the faint rustle of a sleeve. (no music) <no room tone>

Why does stacking cues muddy the mix?

Because you are writing a brief, not a session, and there is no fader. Every cue you add competes for the same generated track, and nothing in any of these APIs lets you set a level. The closest anyone comes is a prose priority line, which OpenAI models in its own example with "Mix: prioritize train and ambient detail over footstep transients."

Kling's two 200-character caps are the most honest thing any vendor publishes about this. That is room for roughly three named sounds and a placement note, and if you cannot say it in that space you are asking for a mix rather than a sound. Three cues per shot is a good working ceiling on every model here: one hero sound tied to the action, one ambience bed, and at most one musical instruction.

Why are dialogue and lip sync a different problem from SFX?

Because effects only have to be plausible, while speech has to land on a mouth. Every vendor treats it separately, and their docs are unusually candid about the limits.

OpenAI is specific about budget: "a 4-second shot will usually accommodate one or two short exchanges, while an 8-second clip can support a few more", and warns that "Long, complex speeches are unlikely to sync well and may break pacing." Kling does not try to solve it inside the video prompt at all: it ships Lip Sync and Voice Management as separate endpoints. Vidu gates voice_id behind the audio switch. ByteDance's Seedance troubleshooting page documents mispronunciation of homophones and a reference voice that "differs significantly from the reference voice" as known failure modes, and its 2.0 guide says the language of dialogue must be consistent and that mixing languages should be avoided.

Practical consequence: budget one short line per four seconds, keep it in one language, put it in quotation marks or the model's dialogue container, and do not write jokes that depend on precise timing.

When should you replace generated audio in post?

For anything finished, assume you will. Generated audio is a scratch track that proves the edit works, and the vendors say so themselves if you read the troubleshooting pages rather than the launch posts.

ByteDance documents that "When a video contains narration, abrupt clicking sounds and cut-off noise are likely to appear at the end of the video", and its recommended fix is to "use editing tools such as CapCut to apply audio fade-out processing to the ending audio track through the volume envelope". That is a vendor telling you to finish the audio in an editor. Its own Seedance 2.0 launch post concedes the model "still needs to address issues like occasional audio distortion". Google's Veo limitations note that the model "will sometimes block a video from generating because of safety filters or other processing issues with the audio".

Keep the generated track when the clip is social-native, short and effects-led, or when you are testing timing before committing to a mix. Replace it when dialogue has to be understood, when music has to hit a cut, when the deliverable will be graded and mastered, or when one sound has to match across shots generated in separate calls. Nothing here guarantees the same door sounds the same twice.

A reusable audio block for any model

Paste this under your visual description, then convert the labels to whatever the target model documents: separate sentences plus SFX: and Ambient noise: for Veo, brackets for Seedance, sound_effect_prompt and bgm_prompt for Kling's audio endpoint, plain prose for Vidu and Sora.

AMBIENCE: [one line naming the space and its constant background sound]
HERO SOUND: [one named source, tied to the action that causes it, with distance]
SECOND SOUND: [optional, quieter, different frequency range from the hero]
MUSIC: [either one instrument and a placement, or the word none]
DIALOGUE: ["one short line" per four seconds, one language, in quotes]
EXCLUDE: [what must not appear: score, foley, crowd, subtitles]

If you want a starting structure for the visual half that this slots into, the seven-part video prompt anatomy is the shape most of these end up taking. Sound is the seventh part, and it is the one people delete when the prompt gets long. It should be the last thing you cut.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account