Back to blog
Video19 min read

Dialogue Prompting in Veo 3.1 (Speech, Timing and Lip Sync)

How to write dialogue into a Veo 3.1 prompt: exact speech syntax, delivery direction, timing inside an 8-second clip, what lip sync actually does, and 20 copy-paste prompts.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: Put the spoken line in quotation marks attached to an attribution clause — A woman says, "We have to leave now." — which is the syntax Google's own Veo prompt guide uses. Direct delivery in that same sentence. Veo 3.1 clips run 4, 6 or 8 seconds, so time your line with a stopwatch before you generate.

How do you write dialogue into a Veo prompt?

You write it as prose, inside the prompt, in quotation marks, attached to whoever says it. There is no separate dialogue field.

Google's Veo documentation gives one sentence of guidance and one example. Under its audio section it says to use quotes for specific speech, and the example it ships is:

"This must be the key," he murmured.

The Veo 3.1 prompting guide on the Google Cloud blog (published October 16, 2025) uses the other word order, with the attribution first:

A woman says, "We have to leave now."

Both work. The second is easier to control because the attribution clause is where you put your delivery direction, and putting it before the line means the model reads the direction before it reads the words.

That is the entire documented syntax. Everything else in this post is craft built on top of it, and I will label it as such.

Which Veo version are you actually prompting?

Veo 3.1 is the current Veo model as of August 26, 2026, and it is still in preview. The Gemini API model IDs are veo-3.1-generate-preview and veo-3.1-fast-generate-preview, with a Veo 3.1 Lite variant also listed.

The documented specs that matter for dialogue:

SpecVeo 3.1 (verified Aug 26, 2026)
Clip duration4s, 6s, 8s — 8s required for reference images, interpolation, or higher resolutions
Resolution720p, 1080p (8s only), 4k (8s only)
Aspect ratios16:9, 9:16
AudioNatively generated with video, always on
Prompt text limit1,024 tokens
Language"English (EN) is fully supported, but other languages have not been evaluated"

That last row is Google's own wording, and it is the most under-reported fact in this category. If you are writing dialogue in Spanish, Hindi or Bengali, you are outside what Google says it has evaluated. It may work. It is not documented to work.

One more thing changed recently, and it will date this post faster than anything else here. Google's Gemini API video generation page now recommends Gemini Omni Flash (gemini-omni-flash-preview) as the default video model, and positions Veo 3.1 for specific capabilities like scene extension, last-frame control, and existing pipelines. Omni Flash is a different model with different behaviour. Its own docs note that uploading audio references is unsupported and that voice editing is not supported. Check ai.google.dev/gemini-api/docs/video before you build anything on either.

If you want the full non-dialogue prompt anatomy first, our Veo 3 prompt structure guide covers subject, action, camera and style in depth. This post assumes you have that and are adding speech.

How do you direct delivery — tone, pace, accent, volume?

You direct it inside the attribution clause, in plain adjectives and adverbs, before the quoted line.

Google's own Veo 3.1 guide demonstrates this with he says in a weary voice, "Of all the offices in this town, you had to walk into mine." The delivery instruction is not a parameter. It is an adverbial phrase, and it sits between the speaker and the speech.

The direction levers worth using, roughly in order of how reliably they land in my experience with the documented syntax:

  • Volume: whispered, hushed, barely audible, raised, shouting over the noise
  • Emotional state: weary, delighted, defensive, resigned, quietly furious
  • Pace: clipped, unhurried, rushing the words together, trailing off
  • Register: formal, casual, clinical, conspiratorial
  • Physical condition: out of breath, mouth full, over a bad phone line
  • Accent: stated plainly, e.g. a soft Scottish accent

Pace and pauses are the two you will fight most. A comma inside the quoted line reads as a small beat. An explicit stage direction between two quoted fragments reads as a longer one, and that is more reliable than punctuation alone:

A tired detective in a rumpled coat leans back in his chair.
He says, quietly, "I know exactly who did it."
He pauses, exhales, then adds, "I just can't prove it."

How much dialogue fits in one clip?

Google does not publish a word limit for spoken lines. What it publishes is the clip length: 4, 6 or 8 seconds. So the constraint is arithmetic, not policy.

The method that works is unglamorous. Read your line out loud, at the pace you actually want it delivered, and time it. If it does not finish inside your chosen duration with at least a beat of silence to spare at each end, it is too long. Cut words, not delivery.

Why the beat matters: a clip that starts on the first syllable and ends on the last one feels wrong even when the sync is fine. You want a moment of the character existing before they speak and a moment after they stop.

For an 8-second clip, in practice, that leaves you room for one substantial line, or two short ones, or one line plus a reaction. It does not leave you room for a conversation. Anyone promising you a four-line exchange in 8 seconds is showing you a clip where the speech is rushed.

There is a second, harder ceiling: the Gemini API Veo 3.1 model page documents a 1,024-token text input limit. Long cinematographic prompts with several quoted lines plus a page of style description burn tokens fast and will hit it. When they do, something gets dropped, and dialogue buried at the end of a long prompt is a good candidate.

Timestamp prompting for pacing

Google's Veo 3.1 prompting guide includes a timestamped workflow that carves the 8 seconds into explicit beats. Its published example uses this shape:

[00:00-00:02] Medium shot from behind a young female explorer as she pushes aside a large jungle vine.
[00:02-00:04] Reverse shot of her face, filled with awe. SFX: rustle of dense leaves, distant bird calls.
[00:04-00:06] Tracking shot as she runs her hand over carvings. Emotion: wonder and reverence.
[00:06-00:08] Wide, high-angle crane shot revealing the temple complex.

That is Google's own example, and it is for visuals and SFX rather than dialogue. Adapting it to speech is my extension of the pattern, not documented behaviour. It is also the single most useful thing you can do to stop a line being rushed, because it tells the model which two seconds the line has to fit inside.

Does Veo 3.1 actually do lip sync?

It generates mouth movement and speech in the same pass, which is not the same thing as a lip-sync feature, and the difference matters.

Google does not document lip sync as a named, specified capability of Veo 3.1 anywhere I could find on August 26, 2026: not on the Gemini API Veo page, not in the Veo 3.1 launch post, not in the Cloud prompting guide. What Google DeepMind's Veo model page does say, in its own words, is that creating videos with natural and consistent spoken audio, particularly for shorter speech segments, remains an area of active development.

That is the vendor telling you where the seam is. Short lines are the hard case, and short lines are exactly what fits in an 8-second clip.

What this means in practice. Wide and medium shots hold up well, because the mouth is small in frame and small errors are invisible. Tight close-ups are where you inspect frame by frame. If your shot is a talking head at close range and the sync is not landing, you have two options: reframe wider, or take the speech out of Veo entirely.

For that second route, you generate the performance silently and add speech with a dedicated tool afterwards. Runway documents Generative Audio as a way to add speech to your images and videos, and its API changelog lists dedicated speech models including eleven_v3, which supports audio tags like [laughs] and [whispers]. Kling ships a separate Lipsync feature on its own site. These are separate steps in a pipeline, with all the version drift and cost that implies. What they buy you is a retake on the voice without regenerating the video.

How do you handle two characters talking?

By making attribution unambiguous, because that is the failure mode. In a single-character clip the model cannot get the speaker wrong. In a two-hander it can, and it will.

The rule: attach every line to a visual description, not to a name. The model has never met "Sarah." It has met "the woman in the red coat."

Two people at a diner counter. On the left, a woman in a red coat with wet hair.
On the right, an older man in a grey work jacket holding a coffee cup.
The woman in the red coat turns to him and says, "You still haven't answered me."
The man in the grey jacket sets down the cup and replies, flatly, "I know."

Keep the turn-taking explicit (turns to him, replies) so the model has an ordering signal, not just two quoted strings. And keep both lines short, because you are now fitting two lines plus two reaction beats into 8 seconds.

This is one place a competitor is genuinely ahead on documentation. Kling's own VIDEO 3.0 Omni model guide (published February 6, 2026) shows element-tag attribution, where characters are bound to @Name references and dialogue is written as Close-up, @Alan says, "He just likes cookies more than me." Google publishes no equivalent tagging syntax for Veo. If multi-character attribution is the core of your project, that is worth knowing before you pick. Our Veo 3 vs Sora vs Kling comparison covers the wider trade-offs.

20 copy-paste Veo dialogue prompts

Escalating from one line to a two-hander. Each is written for an 8-second clip unless noted. Swap the descriptions; keep the shape.

1. Bare minimum, one line

Medium shot of a barista behind a counter. She looks up and says, "We're out of oat milk."

2. One line with delivery direction

Medium shot of a barista behind a counter. She looks up and says, apologetically, "We're out of oat milk."

3. Emotional state plus physical action

Close-up of a man in a hospital waiting room, hands clasped. He exhales, looks at the floor, and says quietly, "She's going to be fine."

4. Whisper

Tight two-shot in a dark stairwell. A young woman leans toward her friend and whispers, barely audible, "Don't move." Ambient: distant dripping water, faint hum of a fluorescent light.

5. Raised voice over ambient noise

Wide shot on a windy dock. A fisherman in yellow oilskins shouts over the wind, "Tie it off, now!" SFX: howling wind, rigging clattering against metal.

6. Accent and register

Medium shot of an older woman in a tweed jacket in a village post office. She says in a soft Scottish accent, warm and unhurried, "You'll be wanting the second-class, then."

7. Pause mid-line

Medium close-up of a lawyer at a desk. He says, "I read the contract." He pauses, taps the page once, then adds, "Twice."

8. Line landing on a sound effect

Close-up of a mechanic wiping her hands on a rag. She says, dryly, "Try it now." SFX: an engine turning over and catching on the second attempt.

9. Direct to camera, presenter register

Medium shot, eye level, a woman in a plain navy shirt against a soft grey backdrop. She looks directly into the lens and says, confidently and unhurried, "Most of what you've been told about this is wrong."

10. Off-screen narration over action

Wide tracking shot following a cyclist through wet city streets at dawn. A calm male voice narrates over the scene, unhurried: "Nobody starts here because it's easy." Ambient: tyres on wet asphalt, distant traffic.

11. Breathless delivery

Handheld medium shot, a runner stopping at a trailhead, hands on knees. She says between breaths, "I made it. Barely."

12. Timestamped single speaker, 8 seconds

[00:00-00:02] Close-up of a chef's hands plating a dish under warm kitchen light.
[00:02-00:05] Cut to a medium shot of the chef looking up. He says, evenly, "Taste it before you salt it."
[00:05-00:08] Slow push in on the finished plate. SFX: low kitchen hum, a knife set down on steel.

13. Reaction shot, one speaker, one listener

Two-shot at a kitchen table. A man in a grey sweater says, carefully, "I quit this morning." The woman opposite him does not speak; her expression shifts from confusion to a slow smile.

14. Two-hander, one line each

Medium two-shot in a stairwell. On the left, a woman in a red coat. On the right, a man in a grey work jacket. The woman in the red coat says, "You're late." The man in the grey jacket replies, without looking up, "I'm here."

15. Two-hander with delivery contrast

Two-shot in a parked car at night, rain on the windscreen. The driver, a woman with short dark hair, says flatly, "Say it." The passenger, a teenage boy in a school blazer, answers, barely above a whisper, "I lied."

16. Timestamped two-hander

[00:00-00:02] Medium shot of a woman in a red coat entering a diner, shaking rain off her sleeve.
[00:02-00:04] Reverse shot, an older man in a grey jacket at the counter. He says, "You came."
[00:04-00:06] Reverse to the woman in the red coat. She replies, tightly, "I always do."
[00:06-00:08] Wide shot, both in frame, neither speaking. Ambient: coffee machine hiss, low radio.

17. One side of a phone call

Medium shot of a man pacing a small apartment, phone to his ear. He says, increasingly frustrated, "No — no, I already sent it. Check the second folder." SFX: faint unintelligible voice through the phone speaker.

18. Crowd chatter that must stay unintelligible

Wide shot of a crowded market. In the foreground, a vendor turns to the camera and says clearly, "Fresh this morning." Background: dense crowd chatter, indistinct and unintelligible, no discernible words.

19. Rising argument, 6-second clip

Two-shot in a cramped office. A woman in a blazer says, controlled, "That's not what we agreed." The man opposite answers louder, cutting across her, "It's what we're doing."

20. Spokesperson with a product line

Medium shot, a man in his thirties in a plain workshop, soft window light from camera left. He holds up a small notebook, looks into the lens, and says, plainly and without hype, "I write everything down. That's the whole system."

Veo 3.1 vs Kling 3.0 Omni vs Runway Gen-4.5 on audio

Every cell below was checked at the vendor's own documentation on August 26, 2026. Where a vendor does not publish a spec, the cell says so.

Veo 3.1Kling VIDEO 3.0 OmniRunway Gen-4.5
Max clip length8s (also 4s, 6s)Up to 15s2–10s
Resolution720p, 1080p, 4k1080p, 720pNot published in the API changelog entry
Native audio with videoYes, always onYes, native audio-visual outputRunway's Gen-4.5 announcement describes native audio generation and editing; the API changelog entry for gen4_5 does not mention audio
Documented dialogue syntaxQuoted speech with attribution clauseQuoted speech with @Element speaker tagsNot published
Multi-speaker attributionNo dedicated syntax published@Name element referencesNot published
Lip syncNot documented as a named feature"More precise lip-syncing" claimed when voice recordings are bound to character elementsNot published
Speech languages"English (EN) is fully supported, but other languages have not been evaluated"Not stated in the Omni model guideNot published
Model IDveo-3.1-generate-previewNot stated in the model guidegen4_5

Two honest caveats. Runway's Gen-4.5 research page returned server errors on every attempt on August 26, 2026, so the audio claim in that row comes from the announcement text as surfaced in search rather than a page I could read end to end; verify it yourself before relying on it. And Kling's own site rejects direct fetches; the Kling figures come from its VIDEO 3.0 Omni model guide read through a reader proxy.

For a fuller model-by-model comparison beyond audio, see our Veo 3 vs Sora vs Kling breakdown. Note that Sora is no longer a live option: OpenAI's deprecation page states it notified developers on March 24, 2026 that the Videos API and the sora-2 and sora-2-pro aliases would be removed from the API on September 24, 2026, and OpenAI discontinued the Sora consumer app on April 26, 2026.

Troubleshooting: four things that go wrong

The dialogue is ignored entirely. Almost always position or budget. Move the quoted line into the first third of the prompt, attach it to an explicit attribution clause, and cut style description until you are comfortably under the documented 1,024-token input limit. A line at the end of a 400-word cinematic prompt reads as decoration.

The voice is wrong: wrong age, wrong gender, wrong accent. There is no voice selection parameter in the documented Veo 3.1 interface. Your only lever is description, so describe the speaker's age, build and manner in the visual description before the line, and put the accent in the attribution clause. If you need the same voice across multiple clips, this is where Veo will disappoint you: Google publishes no voice-persistence mechanism.

The speech is garbled or rushed. Overrun. Time the line aloud, cut it to fit with a beat spare at each end, and consider using the timestamp format to give the line its own two-second window. Do not fix it by asking for slower delivery in a line that was always too long.

The mouth is out of sync. Reframe wider first. It is the cheapest fix and it works more often than re-rolling. If the shot must be a close-up, generate the performance and replace the audio with a dedicated speech tool afterwards, accepting the seam. Google DeepMind's own page flags short spoken segments as an area of active development, so this is a known limit, not something you will prompt your way out of.

Where to check all of this yourself

Version-specific claims about a third-party model go stale in months. These are the pages I verified against on August 26, 2026, so you can re-check them rather than trust this post in six months:

  • Gemini API Veo docs, for syntax, durations, resolutions and language support: ai.google.dev/gemini-api/docs/veo
  • Gemini API video generation overview, for which model Google currently recommends: ai.google.dev/gemini-api/docs/video
  • Veo 3.1 model page, for the 1,024-token input limit and preview status: ai.google.dev/gemini-api/docs/models/veo-3.1-generate-preview
  • Google DeepMind Veo page, for the spoken-audio limitation quote: deepmind.google/models/veo
  • Ultimate prompting guide for Veo 3.1, Google Cloud blog, October 16, 2025, for prompt structure and timestamp prompting: cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1
  • Kling VIDEO 3.0 Omni model guide, February 6, 2026: kling.ai/quickstart/klingai-video-3-omni-model-user-guide
  • Runway API changelog, for gen4_5 specs and audio model IDs: docs.dev.runwayml.com/api-details/api_changelog
  • OpenAI deprecations, for the Sora 2 API removal date: developers.openai.com/api/docs/deprecations

If you want templates rather than syntax, our free Veo 3 prompt generator post has cinematic starting points you can bolt a dialogue clause onto, and the JSON video prompt templates are worth a look if you are managing dozens of these and need them structured.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account