Back to blog
Video17 min read

Veo Audio Prompts: Sound, Music and Ambience in Veo 3.1

How to direct sound effects, ambience, score and silence in Veo 3.1, what Google actually documents, and why Gemini Omni Flash is now the default video model.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: Veo 3.1 generates audio natively in the same request as the video, and Google documents exactly three cue types: dialogue in quotes, sound effects, and ambient noise. Music has no documented category. Write audio as separate sentences, keep the whole prompt under 1,024 tokens, and check which model you are on first.

What is a Veo audio prompt, and which Google video model should you use?

A Veo audio prompt is not a separate field. It is the part of your ordinary text prompt that describes what the shot sounds like, and Veo generates that soundtrack in the same pass as the picture.

Before you write one, check which model you are pointed at, because Google changed the recommendation. The Gemini API video documentation now says to use Gemini Omni Flash as your default model for video generation, describing it as a fast, multimodal model for video generation and conversational video editing. Veo 3.1 is positioned for scene extension, last-frame control, or integration with legacy pipelines (Google, ai.google.dev/gemini-api/docs/video, accessed August 26, 2026).

That matters for audio specifically, because both models generate it. Omni Flash's own prompt guide says the model will try to generate an appropriate audio track by default, and it is the model whose documentation explicitly tells you how to ask for music. Veo 3.1 is the model with the more developed sound effect and ambient noise vocabulary, plus the editing controls that carry audio through an extended scene.

So: start on Omni Flash for a new project. Reach for Veo 3.1 when you need scene extension, first-and-last-frame transitions, or the specific audio cue grammar below. Everything from here down is about Veo, because that is what most people searching for audio prompts are already using.

Does Veo 3.1 generate sound automatically?

Yes, and you cannot turn it off. Google's Gemini API documentation lists audio support for Veo 3.1 and Veo 3.1 Lite as native generation, always on. There is no generateAudio flag in the documented Veo 3.1 request body, which lists prompt, image, lastFrame, referenceImages, video, aspectRatio, durationSeconds, personGeneration and resolution.

This is the single most important thing to internalise. Silence is not the default state that your prompt adds sound to. A soundtrack is the default state, and your prompt is the only thing standing between you and whatever the model thinks a rain-slicked alley sounds like.

The predecessor made the contrast obvious. Google's own Veo documentation describes Veo 2 as "Silent only" and Veo 3.1 as natively generating audio with video. Everything between those two lines is the reason this article exists.

That ceiling is your real constraint. 1,024 tokens covers the entire prompt: cinematography, subject, action, context, style, and every audio line. A florid three-paragraph camera description leaves very little room for a soundscape, which is a good argument for the terse, labelled audio syntax in the next section.

Diegetic sound, score and ambience: what is the difference?

Film sound splits into three layers, and Veo's documented vocabulary maps onto two of them cleanly and one of them badly. Getting the distinction right is what stops your prompt from producing mud.

LayerWhat it isVeo's documented handle
Diegetic soundSounds with a source inside the frame: a door, an engine, a footstep. The character could hear it.Sound effects. Google's instruction is to describe sounds explicitly, for example "tires screeching loudly, engine roaring".
AmbienceThe continuous background bed that makes a place feel real: room tone, traffic, wind, hum.Ambient noise. Google's instruction is to describe the environment's soundscape, for example "A faint, eerie hum resonates in the background".
ScoreNon-diegetic music laid over the scene. Nobody in the frame can hear it.Not a documented category. See the music section below.

Dialogue is the fourth layer and gets its own treatment: Google's guidance is to use quotes for specific speech, as in "This must be the key," he murmured. Dialogue has enough failure modes to deserve its own article, and we wrote one, so this post stays out of its way. See Dialogue Prompting in Veo 3.1 for speech, timing and the lip sync question.

The reason to keep these three straight in your head is that they behave differently under pressure. Ambience wants to be continuous and unremarkable. Diegetic effects want to be timed to a visible action. Score wants to be independent of both. When you dump all three into one run-on sentence, Veo tends to average them.

How do you write sound effects into a Veo prompt?

Write them as their own sentences, and describe the sound rather than naming the object. Google's Vertex prompt guide is explicit about the first half of that: it recommends using separate sentences in your prompt to describe the audio.

The Google Cloud prompting guide for Veo 3.1 goes further and uses a label prefix, SFX:, in its worked examples. That prefix is not a parsed API syntax, but it is Google's own house convention in published example prompts, and it does a useful job of fencing the audio off from the cinematography.

Medium shot, a mechanic slams the hood of a rusted pickup and wipes her hands
on a rag, in a corrugated-iron workshop at dusk, warm tungsten light, grainy
16mm aesthetic.

SFX: the heavy metallic clang of the hood, the squeak of an unoiled hinge just
before it lands, a wrench clattering onto concrete two seconds later.

Ambient noise: a distant highway, an oscillating fan, cicadas outside the open
roller door.

Three habits separate effects that land from effects that smear:

Name the material, not just the object. "A door closes" gives the model nothing. "The soft thud of a padded studio door sealing shut" gives it a texture, a weight and a decay.

Tie the effect to a visible action. Veo synchronises sound to what it renders, so an effect with no on-screen cause has nothing to lock to. If you want a wrench to clatter, put the wrench in the action line.

Order effects roughly in time. Google's own examples read as sequences: "A rough bark, snapping twigs, footsteps on the damp earth. A lone bird chirps." That is four events in rough chronological order, not a bag of nouns.

How do you prompt ambience without burying the scene?

Ambience is the layer people over-write. A bed is supposed to sit under the scene, so describe one or two continuous sources and stop, rather than cataloguing everything that could theoretically be audible.

Google's published ambience examples are notably restrained: "the sounds of city traffic and distant sirens", "waves crashing on the shore", "the quiet hum of an office". Two elements, one clause. Compare that with the temptation to write a paragraph about the acoustics of a room, which eats your token budget and gives the model four things to blend badly.

There is a second, less obvious use for the ambient line: it establishes scale. A cathedral and a broom cupboard can hold the same subject and the same lighting, and the only thing telling the model which one it is may be whether you wrote "a long reverberant tail on every footstep" or "close, dead room tone with no echo".

Ambient noise: close, dead room tone. No reverb. The faint electrical hum of a
fridge compressor in the next room.
Ambient noise: a long, cold reverberant tail on every sound. Distant water
dripping into stone. The high whistle of wind through a broken window.

Same subject, same camera, two entirely different spaces, and the only variable is the ambient line. If you build reusable prompt blocks, ambience is the highest-leverage thing to keep in a variable, which is the pattern we walk through in JSON Video Prompt Templates for Veo 3.

Can you prompt music in Veo 3.1?

You can write music into the prompt, but Google does not document music as a Veo cue type the way it documents sound effects and ambient noise. That gap is real, and pretending otherwise is how people end up confused when a score request gets ignored.

Here is the evidence, dated. Google's Veo prompt guidance names three audio cues: dialogue, sound effects, ambient noise. Music is absent from the list. Yet Google's own Veo example prompts include music inline, such as "upbeat electronic music with a rhythmical beat is playing, high energy professional video". And in Google Cloud's timestamped prompting example, an orchestral score is smuggled in through the effects channel: SFX: A swelling, gentle orchestral score begins to play.

So music works in practice, is present in Google's examples, and is documented nowhere as a controllable element. Treat it accordingly: describe genre, energy and instrumentation in plain language, do not expect tempo, key or arrangement control, and do not expect the score to duck under dialogue.

How do you layer and time three audio channels in one 8-second shot?

Use timestamp prompting, and attach the audio line to the timestamp block where it belongs. Google Cloud's Veo 3.1 prompting guide demonstrates this pattern directly, assigning actions to timed segments inside a single generation.

Veo 3.1 supports 4, 6 and 8 second durations at 24fps, so you are budgeting between roughly 96 and 192 frames. That is enough for three or four beats, and no more.

[00:00-00:02] Close-up on a violinist's hands, rosin dust catching the light,
bow poised above the strings. Ambient noise: a large hall, cold and empty,
audible room tone.

[00:02-00:04] The bow drops. Medium shot, her shoulders release.
SFX: a single sustained note, the faint scrape of horsehair on gut.

[00:04-00:06] Slow crane rise revealing 900 empty red velvet seats.
SFX: the note swells and gains a low string section beneath it.

[00:06-00:08] Wide, high-angle. She is a small figure in a vast room.
Ambient noise: the hall reverb lengthens until the note decays into silence.

Two things are doing the work there. First, each block owns its own audio line, so the model is not asked to distribute one soundscape across four different shots. Second, the ambience persists as a described quality while the effects change, which is how a real mix behaves.

This is also the point where a prompt template stops being a nicety. Four timestamp blocks, each with a camera line, an action line and an audio line, is twelve components to keep consistent across a dozen shots. Nobody retypes that reliably at 11pm.

How do you prompt silence in Veo?

Google does not document a way to request silence. There is no mute parameter, no level control, and no documented negative prompt route on the audio track.

The negative prompt situation is worth stating precisely, because it is the first thing people reach for. Google's Vertex prompt guide does document negative prompts for video, with the instruction to describe what you do not want to see rather than using instructive words like "no" or "don't". But the Gemini API's documented Veo 3.1 parameter list does not include a negativePrompt field, and Gemini Omni Flash's limitations state outright that negative prompts are not supported. Three Google surfaces, three different answers, and none of them promises anything about audio.

What actually works is describing a place that is quiet, rather than asking for the absence of sound.

Ambient noise: near-total silence. The dead, padded quiet of a recording booth.
The only audible sound is the subject's own breathing, close and slightly
uneven.
Ambient noise: the muffled hush of heavy snowfall. Every sound is absorbed.
No wind. A single distant crow, three seconds in, then nothing.

Note the trick in both: they give the model one small permitted sound. A prompt that asks for pure nothing tends to get filled in. A prompt that says "the only thing you may generate is breathing" gives the model a target it can actually hit.

If silence is a hard requirement rather than a stylistic preference, strip the audio track in an editor. That is not a workaround to be embarrassed about. It is the only guaranteed route, because the audio is always on.

What can you not control in Veo's audio?

More than the tutorials admit. Here is the honest list, with the source for each, as of August 26, 2026.

Lip sync. Not documented. There is no lip sync feature, parameter, or quality guarantee for Veo 3.1 in Google's documentation. Google DeepMind's Veo page states that creating videos with natural and consistent spoken audio, particularly for shorter speech segments, remains an area of active development, and that the team is continuously working to refine audio synchronization and eliminate instances of incoherent speech.

Specific voices. Not documented for Veo. You can describe a voice in prose, as in Google's own example of "a voiceover with a polished British accent speaks in a serious, urgent tone", but there is no voice ID, no reference audio upload, and no consistency guarantee between runs. Gemini Omni Flash's limitations say plainly that voice editing is not supported and that uploading audio references is unsupported in the current version of the API.

Levels, mixing and ducking. Not documented on any Google surface I could find. You cannot set a music bed 12dB under dialogue. You describe intent and accept the mix.

Language. Google's Veo documentation states that English is fully supported, but other languages have not been evaluated.

Audio on some editing operations. Google Cloud's Veo 3.1 announcement notes that the add and remove object feature currently utilizes the Veo 2 model and does not generate audio. Not every Veo-branded operation is an audio operation.

Reliability. Google's Veo limitations note that the model will sometimes block a video from generating because of safety filters or other processing issues with the audio. Budget retries.

Is native audio still Veo's advantage over Kling and Runway?

Against Google's own back catalogue, the gap is documented and large. Against the competition, the honest answer in August 2026 is that I could not verify it from the vendors' own pages, and I am not going to infer it.

Google video models, audio surface only. Source: Google Gemini API and Google Cloud documentation, accessed August 26, 2026.
FeatureVeo 3.1Gemini Omni FlashVeo 2
Native audio with video
Audio toggle in requestNone, always onNone, always onN/A
Documented sound effect cueDescribed generally
Documented ambient noise cueDescribed generally
Documented music guidance
Negative promptsIn Vertex guide, not in Gemini API paramsExplicitly not supportedIn Vertex guide
Text input limit1,024 tokensNot documentedNot documented

On the rivals: Runway's developer documentation at docs.dev.runwayml.com lists Gen-4.5, Aleph 2.0 and Seedance 2.5 and documents no audio output parameter for its video generation endpoints, but Runway's help centre returns 403 to automated requests and its product pages did not return readable text through a reader proxy, so I could not confirm the current state of its audio features from Runway's own writing. Kling's API reference at app.klingai.com returns HTTP 446 to direct fetches and only a navigation stub through a proxy, so Kling's audio capability is likewise unverified here. Third-party coverage of both is contradictory on dates and features, which is exactly why I am not repeating it.

What I can say without hedging: Google documents audio as part of the generation, in the same request, driven by the same prompt text, with a named cue vocabulary. That is a specific, checkable claim. "Nobody else has audio" is not, and it may well be stale. If you are choosing between models on capability rather than on audio alone, Veo 3 vs Sora vs Kling covers the wider comparison, and How to Direct AI Video Like a Filmmaker covers the visual side of the same prompt.

A reusable audio block you can paste into any Veo prompt

Keep the audio section structurally identical across every shot you generate. Consistency in the prompt is what makes results comparable when you change one variable.

SFX: [1-3 sounds, each tied to a visible action, in rough time order,
described by material and weight rather than by object name]

Ambient noise: [1-2 continuous sources, plus one explicit statement about
the acoustic space: dead / close / reverberant / open]

Music: [genre + energy + one instrument, or omit this line entirely if the
scene should stay diegetic]

Filled in for a single shot:

Wide shot, a lighthouse keeper climbs an iron spiral staircase, handheld,
cold blue pre-dawn light through salt-crusted glass.

SFX: the dull ring of boot soles on wet iron, one step at a time, and the
brief metallic groan of the handrail taking his weight.

Ambient noise: a long reverberant tail inside the stone tower. Storm surf
muffled through thick walls. No wind inside.

Music: none.

That last line is doing real work. "Music: none" is not a documented instruction and will not reliably suppress a score, but it is a clear statement of intent inside the prompt text, and it costs you four tokens.

If you write these often, the structure is the asset, not any individual prompt. That is the whole argument behind Veo 3 Prompt Structure: a consistent skeleton you fill in, rather than a fresh paragraph every time you have an idea.

Prompt Architects does not generate video. It generates the prompt, then keeps it: the video prompt builder produces the structured block, the Prompt Library stores it, and Global Variables swap the ambience or the subject without you retyping the other eleven components. The rendering is still Google's job, and the audio limits above are still Google's limits.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account