TL;DR: An AI video music prompt only works where the vendor documents music. Seedance and LTX name it as a prompt element. Veo generates a soundtrack from prose but documents no music cue label. Kling directs music on a separate call. Runway Gen-4.5 and Luma Ray 3.2 publish no audio at all. Never name a real song.
Music is the fastest way to change what a shot means, and it is the part of a video prompt most people either skip or fill with the words "epic cinematic". Both are mistakes, but only one is fixable by writing better. The other is a boundary problem: several of these models do not let you direct music at all, and a prompt that assumes otherwise is just wasted characters in the field.
So start with the boundary. Everything below was read on the vendors' own documentation on 29 August 2026.
Which AI video models let you direct music by prompt?
Three different answers. Some models generate music and let you describe it. Some generate audio with dialogue and effects but no music layer you can address. Some generate nothing at all.
| Model | Generates audio? | Music specifically? | Directable by prompt? |
|---|---|---|---|
| Veo 3.1 (Gemini API) | Yes, natively | Yes, in practice | Prose only. Documented cue labels are dialogue, SFX, ambient noise |
| Seedance 2.5 / 2.0 | Yes | Yes, with a bracket for it | Yes, plus negative control |
| LTX-2.5 | Yes, generate_audio | Yes, named prompt element | Yes, including continuity across cuts |
| Kling 3.0 / 3.0 Omni | settings.audio, default off | Not at generation time | No. Use the video-to-audio endpoint |
| Vidu Q3 | Yes, audio boolean | bgm is gone on Q3 | No |
| Grok Imagine Video 1.5 | Yes, on by default | Not documented | No. The audio docs cover voices |
| Runway Gen-4.5 | No audio field | No | No |
| Luma Ray 3.2 | Nothing published | Nothing published | No |
Four of those rows deserve their exact wording, because a documented absence and a control you have not found are very different things.
Veo 3.1 generates 8-second videos "with natively generated audio", and its prompting section says you "can provide Veo with cues for sound effects, ambient noise, and dialogue" so the model can "generate a synchronized soundtrack". Three cue labels: Dialogue, Sound Effects (SFX), Ambient Noise. Music is not one of them. Google's own worked example on the same page nonetheless ends with "upbeat electronic music with a rhythmical beat is playing", so music written as plain prose clearly lands. It just is not a documented category, which means you get no guarantees about it. (ai.google.dev/gemini-api/docs/veo.)
Seedance is the opposite: ByteDance publishes a symbol for music. The 2.5 tutorial says to "Use () for music, <> for sound effects, {} for dialogue, and 【】 for subtitles". The 2.0 series prompt guide gives the same table with the music symbol as full-width () and an example reading (fast-paced rock music is playing in the background). ByteDance also documents negative control: the 2.5 guide "Supports negative audio control for finer dimensions, including sound effects, background music (BGM), and dialogue", with the example "No BGM; generate only environmental sounds and action sounds." (docs.byteplus.com, ModelArk articles 2607688, 2607689 and 2222480.)
LTX-2.5 treats music as one of six key prompt elements: "Clearly describe ambient sound, music, speech, or singing." Its multi-shot guidance goes further than any other vendor here, telling you that at every cut you should "say whether music / dialogue / ambience continues or changes", with worked phrasing of "the piano score continues across the cut" or "the dialogue drops; only wind remains." That is real score direction, written down. (docs.ltx.io.)
Kling ships silent. The 3.0 Omni text-to-video spec sets settings.audio to a default of off, with native producing "native audio matching the visuals" and no music field of any kind. Music is directable on Kling, but on a different endpoint, which is the next section.
The four no-music rows are documented absences rather than gaps in our reading. Vidu publishes a bgm boolean whose true value means "the system will automatically add a suitable BGM", but the same entry adds "BGM does not available in q3 models" (sic), and the audio boolean that replaced it yields sound described as dialogue and effects. Grok Imagine Video 1.5 says "Generated videos include an audio track by default", then spends its whole audio section on preset voices; the word music does not appear on the page. Runway Gen-4.5 carries no audio field on its own endpoints, although the Veo and Seedance models Runway hosts do. Luma Ray 3.2's video generation guide mentions audio, sound and music zero times.
Should the music even come from the video model?
Usually not, and this is the part most template posts leave out.
For real work, the score is not generated with the video. It is licensed or written separately and laid under the cut in an editor, because that is the only way to get a track that runs the length of the piece, hits your actual edit points, survives a client revision, and comes with a licence. An eight-second generation gives you none of that.
Prompted music direction has a narrower job: mood coherence inside a single generated clip. You tell the model what emotional world the shot lives in so the motion, the performance and the sound agree. That is worth doing. It is not scoring.
| Feature | Prompt it into the clip | Score it separately |
|---|---|---|
| Runs the full length of the edit | ||
| Hits your real cut points | Only inside one clip | |
| Survives a client revision | ||
| Comes with a licence you can show | Check vendor terms | |
| Keeps motion and sound agreeing | You have to match it | |
| Costs a second tool |
There is a middle path, better documented than most people realise. Three endpoints take a finished video and score it from a text brief:
- Kling video-to-audio (
POST /v1/audio/video-to-audio) takes abgm_promptand a separatesound_effect_prompt, each capped at 200 characters, on an input video between 3 and 20 seconds. - Pika Soundtrack takes a video and an
instructiondescribing "sound effects, ambience, dialogue, music, or overall style". Its spec notes the picture and duration are preserved and only the soundtrack is replaced. - Sonilo v1.1 Music, in video-to-scored-video mode, takes a
promptdocumented simply as "Guide the music.", plus apreserve_speechflag that will "Isolate and keep the original speech under the score."
That last flag is the ambient-bed-under-dialogue problem solved as a parameter, which tells you the problem is real.
Twenty-eight music prompts, organised by what the music has to do
Labelled by target model. Nothing here is cross-pasted between vendors. For the audio syntax underneath these, see our AI video sound design guide; for Veo specifically, Veo audio prompts. Camera and shot vocabulary lives in camera movement vocabulary for AI video.
Set an emotional register
# 1 — Veo 3.1 — plain prose, music written as a clause
A woman sits alone in a parked car at night, rain running down the
windscreen, streetlight cutting across her face. She does not move.
A single sustained cello note holds under the whole shot, quiet,
slightly out of tune, no percussion. Rain on the roof stays audible
above it.
# 2 — LTX-2.5 — music as one of the six prompt elements
A handheld medium shot follows a boy carrying a cardboard box down an
empty school corridor at the end of term. Late afternoon light through
high windows, dust in the air. His face is set and blank. Warm, simple
piano, two hands, mid-tempo, slightly behind the beat, with room tone
and distant traffic under it. No strings, no swell.
# 3 — Seedance 2.5 — bracketed music cue
A market stall at dawn. The vendor lifts the shutter and steam rises
off the first pot.
(gentle nylon-string guitar, single instrument, unhurried, no drums)
<the shutter rattling up, water coming to the boil, one moped passing>
Match cut rhythm and pace
# 4 — LTX-2.5 — tempo tied to the edit, continuity stated
A hard cut sequence of three shots in a busy commercial kitchen: hands
scoring a fish, a pan going up in flame, a plate landing on the pass.
Fast, clipped percussion with a wood-block pulse at roughly 120 beats
per minute drives the whole sequence; the pulse continues across every
cut without pause. Kitchen noise sits under it, never above it.
# 5 — Seedance 2.5 — music pace named against shot rhythm
Shot 1: quick pan across a cyclist clipping in. Shot 2: close on the
chain catching. Shot 3: wide, she pulls away down the hill.
(driving four-on-the-floor electronic pulse, steady tempo, no fills,
no build)
<the click of a cleat, chain tension, wind rising>
# 6 — Pika Soundtrack — instruction field on a finished clip
Score this cut with a steady mid-tempo motorik drum pattern that never
changes speed, so every visual cut lands against an unbroken pulse.
Keep the existing footsteps and door sounds audible under the music.
Build tension
# 7 — Veo 3.1 — a rising shape described, not a genre named
A corridor of identical office doors, seen from a slow tracking shot
that never quite reaches the end. Fluorescent tubes flicker. A low
string drone begins almost inaudibly and rises in pitch and volume
across the whole shot without ever resolving. No drums. The hum of the
lights stays constant underneath.
# 8 — LTX-2.5 — tension with a stated non-resolution
A static wide of a farmhouse kitchen at night; a woman watches the
door. The camera does not move. A single detuned piano note repeats
every two seconds, getting slightly louder each time, and stops
abruptly before the shot ends, leaving only the fridge hum.
# 9 — Sonilo v1.1 Music — prompt field, video-to-scored-video
Slow build from near silence: sub-bass pulse first, then a bowed
metallic scrape entering halfway, then a low choir tone in the last
quarter. Never reach a full climax. Keep dialogue audible throughout.
Put a hit or a sting on a specific beat
# 10 — Seedance 2.5 — integer-second timestamps
[0s-3s] A man reads a letter at a kitchen table, calm.
(sparse piano, single notes, quiet)
[3s-4s] He stops reading. His hand tightens on the paper.
(one low struck piano chord on the moment his hand tightens, then
silence)
[4s-8s] He sets the letter down and looks up.
(room tone only, no music)
# 11 — LTX-2.5 — hit tied to a described action, not a timecode
A close-up of a hand hovering over a light switch in a dark hallway.
The music is a quiet sustained organ chord. At the exact moment the
switch clicks down, a single hard bass drum hit lands and the organ
cuts out. Only the buzz of the bulb remains for the rest of the shot.
# 12 — Kling video-to-audio — bgm_prompt, 200 characters maximum
Sparse low piano throughout. One hard low brass stab at the moment the
door slams, then silence for the rest. No melody, no build.
Put a bed under dialogue and keep it out of the way
# 13 — Veo 3.1 — music described as subordinate to speech
Two colleagues talk across a desk in a small office, late afternoon.
"You already knew," she says. He does not answer. Very quiet
mid-register strings sit far under the dialogue, no melody, barely
present, and do not change while either of them is speaking. Room tone
and a distant printer are audible in the gaps.
# 14 — Sonilo v1.1 Music — preserve_speech is the point
Understated ambient pad, low in the mix, no percussion, no melodic
hook, nothing that competes with a speaking voice. It should be
noticeable only when nobody is talking.
# 15 — Seedance 2.5 — dialogue bracket plus a quiet music bracket
A father and daughter sit on the back step at dusk.
{Are you going to tell her, or am I}
(soft sustained synth pad, very low in the mix, no rhythm)
<crickets, a screen door creaking on its hinge>
Direct genre and instrumentation
# 16 — LTX-2.5 — instrumentation, not genre labels
An establishing wide of a fishing harbour before sunrise, boats
knocking against the dock. The music is solo accordion and a single
upright bass, played slowly and slightly loosely, with no drums and no
reverb, as if recorded in a small room.
# 17 — Seedance 2.5 — era and instrument stack
A teenager rides a skateboard through an empty multi-storey car park.
(distorted electric guitar, live drums, no synths, recorded rough and
loud, mid-tempo)
<wheels over expansion joints, an echo off concrete>
# 18 — Kling video-to-audio — bgm_prompt, instrumentation only
Solo harpsichord, dry, no reverb, moderate tempo, repeating four-bar
figure, no bass, no percussion, no swell.
Direct an era or a period feel
# 19 — Veo 3.1 — period conveyed by recording character
A newsreel-style shot of a crowded factory floor, black and white,
slightly overexposed. The music sounds like a small brass band
recorded onto optical film: narrow frequency range, audible surface
noise, no low end at all, played briskly.
# 20 — Seedance 2.5 — period without naming an artist
A couple dances in an empty ballroom, one lamp lit.
(1940s-style small-group swing: muted trumpet, brushed snare, upright
bass, played quietly as if from another room)
<their shoes on parquet, a chair scraping somewhere off-screen>
# 21 — Pika Soundtrack — instruction field
Score as if recorded in the early 1980s on analogue tape: monophonic
synth bass, gated reverb on a simple drum machine, one held pad chord,
tape hiss audible between phrases.
Use silence on purpose
On Kling, silence is a parameter rather than a sentence: leave settings.audio at its default of off and the render is silent, so you can lay your own track under it in an editor. Prompt 23 is for the other case, where you want the room but not a score.
# 22 — Seedance 2.5 — documented negative audio control
A surgeon scrubs in, alone, before an operation. No BGM; generate only
environmental sounds and action sounds.
<running water, the brush on skin, the extractor fan>
# 23 — Kling 3.0 Omni — sound but no music, with settings.audio: native
A locksmith works on a cylinder at a workbench under one lamp. The
scene has no music at all. The only sound is the room: the pick moving
in the barrel, a spring clicking, his breathing, and the hum of a
strip light overhead.
# 24 — LTX-2.5 — music stopping as a dramatic event
A wide shot of a stadium tunnel; a player walks toward the light with
a full crowd roaring and a brass fanfare playing. The instant she
steps onto the pitch, the fanfare stops dead and the crowd drops to a
distant murmur. Only her breathing remains close.
Write diegetic music, coming from inside the scene
# 25 — Veo 3.1 — source named and placed in the room
A house party seen from the kitchen doorway. Dance music plays from a
speaker in the next room, muffled by the wall, bass-heavy and
indistinct, with the top end mostly gone. Conversation and glassware
are louder than the music.
# 26 — LTX-2.5 — diegetic source that changes with the camera
A tracking shot follows a busker's guitar case along a subway platform
as a train arrives. The guitar and voice are close and clear at the
start of the move, then progressively swallowed by the train, until
only the rhythm of the strumming is left under the brakes.
# 27 — Seedance 2.5 — a car radio as the only music source
A man drives at night, one hand on the wheel.
(tinny country music through car speakers, thin, no bass, with a
faint station hiss)
<indicator ticking, tyres on wet road, wipers>
# 28 — Pika Soundtrack — instruction field, diegetic only
All music must sound like it comes from the small radio visible on the
workbench: narrow band, mono, slightly detuned, with the room's
reverb on it. No score, no music that the characters could not hear.
What actually makes music direction work?
The prompts above share five decisions. Those are the transferable part.
Tempo reads against motion, not against mood. A slow shot with fast music feels urgent; the same music over fast motion feels ordinary. Specify tempo in relation to what is moving. "Steady pulse at roughly 120 beats per minute over a slow push-in" tells the model something. "Upbeat" does not.
Instrumentation is the fastest emotional shorthand you have. Solo cello, solo accordion and solo synth pad over the same shot produce three different films. One instrument plus its playing style beats three adjectives, because an instrument is a physical thing a model has heard thousands of examples of, and "melancholy" is not.
Decide whether you are scoring the picture or the subtext. Scoring the picture means the music matches what is visible: bright music over a happy scene. Scoring the subtext means it matches what the scene is actually about, which is frequently the opposite. Cheerful music over a funeral is a decision, not a mistake, and it is one you have to state explicitly or the model will default to matching the picture.
"Epic cinematic" is not a music instruction. It names a marketing category and contains no acoustic information at all. Everything that makes a cue feel epic is missing from those two words: which instruments, in which register, at what tempo, and with what shape over the clip's length. Replace it with "low brass swells under a slow string ostinato, quiet at the start and full by the end" and there is something to render.
Music sets meaning the image cannot supply. A person walking down a street is neutral footage. What tells a viewer whether they are late, hunted, or free is the sound. This is the strongest argument for prompting music at all, even for eight seconds: it is the cheapest way to make the model commit to an interpretation, and it keeps performance and motion honest to it. Continuity across shots is the same problem, which our transitions in AI video prompts guide covers from the edit side.
Why should you never name an artist, band or song?
Two reasons, and the second one is the one people ignore.
It is a rights problem. OpenAI's video generation guide lists content restrictions and states plainly that "Copyrighted characters and copyrighted music will be rejected." Vendors filter for this, and a rejected generation is the good outcome. The bad one is a clip you ship with something recognisable in it.
It also does not work. The model has no access to the recording you are thinking of. A name buys you a wobbly average of everything associated with it in the training data, which is less specific than the description you could have written yourself.
The substitute is mechanical and it is better: describe instrumentation, tempo, texture and era. Instead of naming a band, write "distorted electric guitar, live drums, no synths, recorded rough and loud, mid-tempo". That is four decisions the model can act on, and none belongs to anybody.
Who owns the music the model made?
Nobody can answer that for you in a blog post.
Generated audio arrives with no licence document attached. Whether you can put it in a client deliverable depends on the vendor's terms, the specific model, and often the plan you are on, and the audio terms are not always the same document as the video terms. Read them before the clip goes into paid work: OpenAI publishes usage policies at openai.com/policies/usage-policies and Google publishes generative-AI terms at policies.google.com/terms/generative-ai; Kling, Runway, Lightricks and ByteDance each publish their own. If your client needs a licence they can point at in an audit, a licensed library is still the answer.
If what you actually want is a soundtrack rather than eight seconds of coherent mood, use a music model. Suno, ElevenLabs Music, MiniMax Music and Google's Lyria all exist for this, produce a track of usable length, and let you iterate on it independently of the picture. Our Suno prompt templates cover the prompting side. For avatar and presenter video, where music almost always goes on underneath in the edit rather than into the generation, see our HeyGen prompt templates.
A reusable music block for any model that supports it
Five slots, in this order. The block works on every model above that generates music. Fill the slots in prose for Veo and LTX, inside Seedance's music bracket (() per the 2.5 tutorial, () on the 2.0 guide's table), and inside bgm_prompt for Kling's video-to-audio call.
1. Instrument(s) and how many — "solo cello", "muted trumpet and brushed snare"
2. Playing style — "slow, slightly behind the beat, no vibrato"
3. Tempo, stated against motion — "steady pulse under a slow push-in"
4. Dynamic shape over the clip — "quiet at the start, full by the end", "never resolves"
5. What it must NOT do — "no drums", "no swell", "never above the dialogue"
Slot five earns its place more often than the others. Most disappointing generated music is not wrong in kind, it is wrong in amount: too busy, too loud, or building when it should hold. Saying what the music must not do is the cheapest correction available.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An AccountMusic direction is a small discipline with an unusually clear boundary. Establish what your model supports, keep the score in an editor, and use the prompt for what it is good at: making eight seconds of picture and sound agree about what they mean.