Back to blog
Video20 min read

AI Video Lip Sync and Dialogue: What Each Model Documents

AI video lip sync is three separate jobs with three prompt shapes. What Veo, Kling, Sora, Seedance, LTX, Runway, Vidu, Wan and Grok each document, plus 27 copy-paste prompts.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: AI video lip sync is three different jobs wearing one name. Some models generate speech straight from your prompt. Some take an audio file and animate a mouth to match it. A few replace dialogue over footage you already shot. Pick the route first, because the prompt shape is different for each one.

What does AI video lip sync actually mean right now?

It means whichever of three jobs the vendor happens to be selling, and they rarely say which.

Route one: native speech. The model writes the audio and the mouth movement in the same generation. You supply words in the prompt. Veo 3.1, Sora 2, Seedance, LTX and Grok Imagine all work this way.

Route two: audio-driven. You supply an audio file and a face, and a dedicated endpoint animates the mouth to match. Kling, Vidu and Alibaba's wan2.2-s2v all ship this as a separate API call, not as a prompt.

Route three: replacement. You already have footage of someone speaking, and you want different words out of the same face. Lightricks calls this Dub-It. ByteDance handles it as a video edit instruction.

The routes are not interchangeable, and the biggest single wasted effort in this whole area is writing a beautiful dialogue prompt for a model that was never going to speak. Sort the route first. Everything below is organised by the job, not the tool. If you want the underlying shot grammar rather than the speech layer, the seven parts of a video prompt is the companion piece.

Which models generate speech, and which just generate sound?

Here is what each vendor publishes on its own pages, checked on August 29, 2026. "Not published" means exactly that: the vendor does not document it, which is not the same as the model being unable to do it.

ModelAudio in the boxHow dialogue gets inLip sync documented
Veo 3.1 (Gemini API)Yes, model table says always onInline, quoted speechNot published, zero mentions
Veo 2No, marked silent onlyn/an/a
Kling 3.0 / 3.0 OmniOff by default (settings.audio)Shot grammar published, no dialogue syntaxSeparate endpoint
Kling Lip Sync APIAudio you supplyAudio file or TTS idYes, /v1/videos/advanced-lip-sync
Sora 2YesInline, guide and example disagree on formNot published
Seedance 2.5Yes{} braces, inside timestamped shotsVia video edit instruction
Seedance 2.0YesCurly braces per symbol tableNot published
Grok Imagine video 1.5Yes, generate_audio to disablePreset voices tagged in the promptNot published
Runway Gen-4.5No audio field on that branchn/aAct-Two is a separate endpoint
Runway Act-TwoDriven by a reference videoPerformance transfer, not textYes, facial performance
LTX-2.5Yes, generate_audio: false to disableQuotation marks, plus language and accentNot as a parameter
LTX-2.3 Dub-It (beta)Regenerates the speechPrompt template, single speakerYes, explicitly
Vidu Lip SyncAudio you supply, or TTS from textaudio_url or textYes, separate endpoint
Luma Ray 3.2Not published anywhere in its docsn/aNot published
Pika 2.5No audio field in the specn/aKling's, via Pika's catalog
Wan 2.7Yes, or supply audio_urlInline in the vendor's own exampleNot published for 2.7
wan2.2-s2vAudio you supplyAudio file plus one imageYes, digital human
MiniMax H3Speech in the vendor's own example; no audio parameter publishedCharacter speaks: plus a reference audio for timbreNot published

Two rows deserve a second look. Kling is the only audio-capable model here that ships silent by default: its 3.0 text-to-video spec sets settings.audio to off, with native as the opt-in, described as "native: The generated video includes native audio matching the visuals." And Luma publishes nothing at all about audio: the words audio, voice, speech and dialogue return zero hits across its Models page, its video generation guide and its API reference index.

Duration matters more than it looks here, because a line that will not fit is the most common cause of a mouth still moving at the cut. Duration parameters by model has the full matrix.

Where is dialogue prompted inline, and in what syntax?

Five vendors publish a form. They do not agree, and the differences are not cosmetic.

Google. The Veo page says: "Dialogue: Use quotes for specific speech." Its own worked example goes further and labels speakers with parenthetical direction, in the shape Man: (Hand on his hunting knife) "That's no ordinary bear."

OpenAI. The Sora 2 prompting guide contradicts itself on one page. The prose says "Dialogue must be described directly in your prompt." Then it tells you to place that dialogue in a <dialogue> block below the prose. The worked example immediately underneath uses a plain Dialogue: label with a bulleted speaker list. Both are OpenAI's. Pick the example, since it is the one they actually ran.

ByteDance. The Seedance 2.0 prompt guide publishes a symbol table: music in full-width parentheses, sound effects in angle brackets, dialogue in {}, subtitles in 【】. The 2.5 tutorial keeps all four markers, including {} for dialogue — it does not replace them — with one quiet difference: the 2.0 page prints the music brackets full-width and the 2.5 page prints them ASCII. What 2.5 adds is time. ByteDance says it outright: "Seedance 2.0 does not respond to timestamps and only responds to shot numbers, while Seedance 2.5 supports integer-second timestamps". Its own worked example opens shots as Shot 1 (0-3s):.

Lightricks. LTX asks you to place spoken dialogue in quotation marks and to "Specify language and accent if needed".

Kling. Kling's API does publish a prompt grammar, and it is a shot grammar, not a dialogue one: "shot n, m, words; shot n, m, words;" with one to six shots, each at least one second, the durations summing to the total, and 512 characters per shot. There is no documented dialogue syntax in the 3.0 Omni spec. Kling's speech story is the separate Lip Sync and Avatar endpoints.

For Veo specifically, dialogue prompting in Veo 3.1 goes considerably deeper than this page can.

How do you prompt one character speaking to camera?

The workhorse. Keep the line short, name the framing, and say who is speaking before you say what they say.

Veo 3.1 (Gemini API), native audio

Medium close-up, [SUBJECT: 30s woman, dark curly hair, olive linen shirt] seated at a
[LOCATION: sunlit kitchen table], addressing the camera directly. Soft window light from
frame left, shallow depth of field. She speaks one line, calm and unhurried:
"[LINE, 8-12 WORDS MAX]"
Ambient: faint refrigerator hum, no music.

Kling Avatar API, audio-driven

POST /v1/videos/avatar/image2video
image: [PORTRAIT URL]
sound_file: [AUDIO URL, mp3/wav/m4a/aac, 2-300s, <=5MB]
prompt: "[SUBJECT] speaks to camera, [EMOTION: warm, steady], slight head movement,
blinks naturally, hands out of frame. Camera static, medium close-up."
mode: pro

LTX-2.5, single continuous take

A single continuous take. [SUBJECT] stands in [LOCATION], facing the lens. The camera
holds still. Present tense throughout. She says, in [LANGUAGE/ACCENT]: "[LINE]".
Ambient: [ROOM TONE]. No music. No cuts.

Wan 2.7, inline dialogue in Alibaba's own house style

[SHOT SIZE] of [SUBJECT] in [LOCATION], [LIGHTING]. Its mouth clearly moving as it speaks
in a [VOICE QUALITY] voice: "[LINE]". [CAMERA MOVE]. Background: [AMBIENCE].

Lightricks makes the case for the single take explicitly: use one when you want "dialogue that must stay lip-synced in one framing." A cut mid-sentence is where sync most often falls apart.

How do you write a two-hander conversation?

Two speakers doubles every problem: identity, turn-taking and timing. Label speakers consistently, alternate turns, and cap the exchange.

Veo 3.1, speaker labels with parenthetical direction

Wide shot, [LOCATION]. [CHARACTER A: description] and [CHARACTER B: description] face each
other. [CAMERA]. 
A: ([PHYSICAL BEAT]) "[LINE 1, SHORT]"
B: ([PHYSICAL BEAT]) "[LINE 2, SHORTER]"
Ambient: [SOUNDSCAPE].

Sora 2, following the guide's worked example rather than its prose

[PROSE SCENE DESCRIPTION: room, light, two figures, mood. 2-3 sentences.]

Dialogue:
- [NAME A]: "[LINE 1]"
- [NAME B]: "[LINE 2]"
- [NAME A]: "[LINE 3]"

Grok Imagine video 1.5, reference-to-video with two preset voices

{
  "model": "grok-imagine-video-1.5",
  "prompt": "The person from <IMAGE_1> sits across from a second speaker in [LOCATION], speaking with the voice from <AUDIO_0>. A second speaker with the voice from <AUDIO_1> replies.",
  "reference_images": [{"url": "[IMAGE_URL_1]"}],
  "reference_audios": [{"voice_id": "[VOICE_A]"}, {"voice_id": "[VOICE_B]"}],
  "duration": 8,
  "aspect_ratio": "9:16",
  "resolution": "720p"
}

Seedance 2.5, timestamped shots

Shot 1 (0-4s): [SHOT SIZE] of [CHARACTER A] in [LOCATION]. [CHARACTER A] says
{[LINE 1]}.
Shot 2 (4-8s): Reverse angle on [CHARACTER B]. [CHARACTER B] says {[LINE 2]}.
Shot 3 (8-12s): Two-shot, both in frame. [CHARACTER A] says {[LINE 3]}.
No subtitles. No BGM; generate only environmental sounds and action sounds.

Two ceilings are published, and they are not the same kind of thing. xAI's is a hard request limit: reference-to-video takes at most three preset voices, and an unrecognised voice_id returns a 400 with the list of valid ones. OpenAI's is softer and narrower — it is about reusable character assets, not about how many people may speak. The video guide says "A single video can include up to two characters", and the Sora 2 cookbook words the same number as advice: "We recommend no more than 2 characters per generation." Neither sentence is a documented cap on speakers.

How do you do voiceover over action without lip sync?

This is the distinction that saves the most time on this whole page. If nobody is on camera speaking, you do not need lip sync at all, and asking for it invites the exact artefacts you were trying to avoid. Narration over b-roll is an audio problem, not a mouth problem.

Veo 3.1, narration with no speaker in frame

[SHOT: hands only, product on a workbench, no face visible]. [CAMERA MOVE]. 
Voiceover, off-screen, [TONE]: "[LINE]"
Ambient: [SOUNDSCAPE]. No on-screen speaker.

Kling 3.0, audio explicitly on

settings.audio: native
prompt: "shot 1, 3, [SCENE A, no visible speaker]; shot 2, 4, [SCENE B]; shot 3, 3,
[SCENE C]. Ambient throughout: [SOUNDSCAPE]. No on-camera dialogue."
settings.duration: 10

Seedance 2.0, symbol table for the audio bed

[SCENE DESCRIPTION]. ([MUSIC: slow piano, low in the mix]) <[SFX: rain on a tin roof]>
No subtitles. No visible speaker.

Runway Gen-4.5, silent by design

promptText: "[SHOT SIZE] of [SUBJECT] [ACTION] in [LOCATION], [LIGHTING], [CAMERA MOVE].
No people speaking on camera."
duration: 10
ratio: 1280:720

That last one is a fact worth internalising: Runway's own gen4.5 branches carry no audio field at all, so the voiceover is a separate job you do elsewhere. For everything about the non-speech layer, sound design prompts for AI video covers it properly.

How do you prompt a reaction with no line?

Silence is an instruction, and most models need to be told, because a face in close-up plus an audio-on model is a strong invitation to invent mumbling.

Veo 3.1

Extreme close-up on [SUBJECT]'s face, [LIGHTING]. They do not speak. [MICRO-EXPRESSION:
eyes widen, jaw tightens, one slow blink]. Mouth closed throughout.
Ambient only: [SOUND]. No dialogue, no voiceover, no music.

Kling 3.0, silent by default is the feature here

settings.audio: off
prompt: "[SHOT SIZE] on [SUBJECT], [EXPRESSION BEAT 1] then [EXPRESSION BEAT 2].
Lips closed and still. [CAMERA MOVE]."

LTX-2.5, one rhythm cue instead of a soundtrack

A single continuous take. [SUBJECT] listens, saying nothing. [PHYSICAL CUE: swallows,
looks away, exhales]. Lips remain closed. One sound only: [DISTANT SOUND].

How do you direct a line with a specific emotion?

Physical cues beat adjectives on every model that documents its preference. ByteDance's own guidance for expressions is to use descriptive sentences and reduce the use of idioms. Google's example carries the direction in parentheses before the line rather than as a mood label.

Veo 3.1, direction before the quote

[SHOT]. [SUBJECT] [PHYSICAL STATE: gripping the doorframe, shoulders drawn up].
Speaking, (voice tight, barely above a whisper): "[LINE]"

Sora 2, emotion in the prose, line in the block

[PROSE: describe posture, breath, hands, eyeline. The emotion lives here.]

Dialogue:
- [NAME]: "[LINE]"

LTX-2.5, delivery vocabulary the vendor lists

[SCENE]. [SUBJECT] speaks with a [DIALOGUE STYLE: resonant voice with gravitas /
childlike curiosity / robotic monotone] at [VOLUME: whisper / mutter / shout]:
"[LINE]".

Seedance 2.5, emotion as a beat inside the shot

Shot 2 (5-9s): Facial close-up. [SUBJECT] [PHYSICAL BEAT: pupils contract, tears fall].
[SUBJECT] says {[LINE]}.

How do you get an accent or another language?

Two vendors are unusually explicit here, and one is unusually honest about the limits.

Google publishes the caveat that matters most: for Veo, "English (EN) is fully supported, but other languages have not been evaluated, so they may work but results can vary." That is a documented uncertainty, not a rumour, and it should change what you promise a client.

ByteDance's rule for Seedance 2.0 is stricter and more mechanical: "The language of dialogue must be consistent, and mixing Chinese and English should be avoided (except for proper nouns)", and an uncommon language must be named alongside the braced line.

LTX Dub-It, the vendor's own template

[SPEAKER] is speaking [LANGUAGE/ACCENT], saying: "[DIALOGUE IN NATIVE SCRIPT]"

Seedance 2.0, language marked explicitly

[SCENE]. [SUBJECT] says in [LANGUAGE] {[LINE IN NATIVE SCRIPT]}. No subtitles.

Veo 3.1, hedged deliberately

[SHOT] of [SUBJECT] in [LOCATION]. They speak one short line in [LANGUAGE] with a
[REGION] accent: "[LINE]". Keep the line under [N] words.

Runway voice dubbing, a separate endpoint entirely

{
  "model": "eleven_voice_dubbing",
  "audioUri": "[AUDIO URL]",
  "targetLang": "[es|fr|de|ja|ko|hi|pt|zh|ar|...]",
  "numSpeakers": [N],
  "dropBackgroundAudio": false
}

Two honest constraints. Lightricks lists Dub-It's validated languages as English, French, Spanish, German and Russian, and states that "it does not translate automatically", so you supply the translated line yourself, in the target script. And Runway's dubbing endpoint is not a video model at all: it dubs the audio, and you still need a face that matches.

How do you replace dialogue over existing footage?

This is ADR, and it is the route with the most documented capability and the tightest constraints.

LTX-2.3 Dub-It, open weights, ComfyUI or a Python script

python -m ltx_pipelines.dub_it \
  --reference-video ./[SOURCE].mp4 \
  --prompt "[SPEAKER] is speaking [LANGUAGE/ACCENT], saying: \"[NEW LINE]\"" \
  --height 720 --width 1280 --num-frames 161 --seed 42

Seedance 2.5, as a video edit instruction

Translate the spoken dialogue in the video into [LANGUAGE], with no subtitles. Precisely
adjust the lip movements to match the translated speech, while keeping everything else
unchanged.

Kling advanced lip sync, face selection first

{
  "session_id": "[FROM FACE RECOGNITION API]",
  "face_choose": [{
    "face_id": "[ID]",
    "sound_file": "[AUDIO URL, 2-60s, <=5MB]",
    "sound_start_time": 0,
    "sound_end_time": 3000,
    "sound_insert_time": 1000,
    "sound_volume": 1,
    "original_audio_volume": 0
  }]
}

Vidu Lip Sync, text-driven

POST https://api.vidu.com/ent/v2/lip-sync
{
  "video_url": "[SOURCE VIDEO URL]",
  "text": "[NEW LINE]",
  "voice_id": "[VOICE ID FROM VIDU'S VOICE LIST]",
  "ref_photo_url": "[FRONTAL FACE IMAGE OF THE TARGET SPEAKER]"
}

Lightricks describes Dub-It as a tool that "re-generates speech in video, producing lip-synced output with new dialogue while preserving the speaker's visual appearance and vocal identity", and lists among its behaviours "Preserves the full video except the lip region". Kling's endpoint states flatly that it "Currently only supports one person lip-sync." Vidu's is the same shape: when the input video contains several faces, "the lip-sync API can only select one face as the target for lip synchronization", which is why the reference photo exists. Alibaba's wan2.2-s2v takes the audio-first route instead, driving a still image so that lip movements, facial expressions and actions follow the audio.

Which lip sync failures can you prompt around?

Sort the symptom before you rewrite anything. About half of what goes wrong here is not a wording problem, and rewriting the sentence against a capability limit is the most expensive habit in this field.

Prompt-fixable.

  1. The line does not fit the clip. Count syllables against seconds. OpenAI's guidance is that "a 4-second shot will usually accommodate one or two short exchanges, while an 8-second clip can support a few more", and that "Long, complex speeches are unlikely to sync well and may break pacing." Cut words or extend the shot.
  2. Wrong speaker moves. Label speakers consistently and alternate turns, which is what OpenAI's guide asks for explicitly.
  3. Unwanted mumbling in a silent shot. Say lips closed, no dialogue, no voiceover. Or use Kling with settings.audio left at off.
  4. Flat delivery. Replace mood adjectives with physical cues, in parentheses before the line.
  5. The wrong language or accent. Name it. LTX asks for it directly; Seedance requires it for uncommon languages.
  6. Sync breaking at a cut. Keep dialogue inside one continuous take, per Lightricks' own advice.
  7. Unwanted subtitles burned in. Add the negative instruction, and lower your expectations: ByteDance says "Currently, it is not possible to directly avoid generating subtitles 100%."

Capability limits, where no rewrite helps.

  1. Phoneme-level accuracy. No vendor in this list publishes a phoneme fidelity claim, a target, or a control for it. There is nothing to prompt.
  2. Teeth and tongue artefacts. Same. Not a documented parameter anywhere.
  3. Multiple speakers on a single-speaker endpoint. Kling's lip sync takes one face. Vidu's takes one face. Lightricks says Dub-It "does not distinguish between multiple speakers." That is architecture, not phrasing.
  4. Identity drift mid-line. ByteDance documents the causes as reference-image problems rather than prompt problems, and its fix is a separate headshot with the face large in frame, not a better sentence. Keeping a character consistent across clips is the deeper treatment.
  5. Audio and video of different lengths. Alibaba spells out the mechanical outcome for Wan: "If the audio is shorter than the video, the remaining video is silent." Trim in an editor.
  6. Lip sync where the vendor documents none. Luma, Pika and Runway's Gen-4.5 branch publish no speech or lip sync capability. Asking harder does not create one.

What are you not allowed to make?

Two constraints. Both are short, both are real, and neither is a footnote.

Do not generate a real, identifiable person saying words they did not say. This is the defining harm of the technology and the platforms are explicit. OpenAI's video guide lists among its restrictions that "Real people—including public figures—cannot be generated" and that "Character uploads that depict human likeness are blocked by default." Runway's usage policy prohibits "Use of the service to impersonate an individual or entity, or to misrepresent your affiliation with an individual or entity". Read the policy of whichever platform you are on, because they differ. And note the part no policy covers: likeness, right-of-publicity and defamation law apply to you regardless of what a tool permits. A model that lets a generation through is not a legal opinion.

Label it where the platform requires labelling. Requirements vary by platform, so check the one you publish to rather than assuming a single rule. YouTube's own policy says it requires "creators to disclose when they use AI to meaningfully alter or generate photorealistic content", and lists among the cases needing disclosure content that “Makes a real person appear to say or do something they didn’t do.” Its non-disclosure list is instructive too: cloning your own voice for voiceovers or dubs sits there. Some models also mark their output automatically. Google states that videos created by Veo are watermarked using SynthID.

Where to check all of this yourself

Every claim above is on a vendor page you can read today, and all of it will move.

All read August 29, 2026.

Write the prompt once, keep it parameterised, and reuse it across the three routes rather than rebuilding it every time the model list changes. That is the whole reason a prompt library beats a folder of screenshots.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account