TL;DR: AI video lip sync is three different jobs wearing one name. Some models generate speech straight from your prompt. Some take an audio file and animate a mouth to match it. A few replace dialogue over footage you already shot. Pick the route first, because the prompt shape is different for each one.
What does AI video lip sync actually mean right now?
It means whichever of three jobs the vendor happens to be selling, and they rarely say which.
Route one: native speech. The model writes the audio and the mouth movement in the same generation. You supply words in the prompt. Veo 3.1, Sora 2, Seedance, LTX and Grok Imagine all work this way.
Route two: audio-driven. You supply an audio file and a face, and a dedicated endpoint animates the mouth to match. Kling, Vidu and Alibaba's wan2.2-s2v all ship this as a separate API call, not as a prompt.
Route three: replacement. You already have footage of someone speaking, and you want different words out of the same face. Lightricks calls this Dub-It. ByteDance handles it as a video edit instruction.
The routes are not interchangeable, and the biggest single wasted effort in this whole area is writing a beautiful dialogue prompt for a model that was never going to speak. Sort the route first. Everything below is organised by the job, not the tool. If you want the underlying shot grammar rather than the speech layer, the seven parts of a video prompt is the companion piece.
Which models generate speech, and which just generate sound?
Here is what each vendor publishes on its own pages, checked on August 29, 2026. "Not published" means exactly that: the vendor does not document it, which is not the same as the model being unable to do it.
| Model | Audio in the box | How dialogue gets in | Lip sync documented |
|---|---|---|---|
| Veo 3.1 (Gemini API) | Yes, model table says always on | Inline, quoted speech | Not published, zero mentions |
| Veo 2 | No, marked silent only | n/a | n/a |
| Kling 3.0 / 3.0 Omni | Off by default (settings.audio) | Shot grammar published, no dialogue syntax | Separate endpoint |
| Kling Lip Sync API | Audio you supply | Audio file or TTS id | Yes, /v1/videos/advanced-lip-sync |
| Sora 2 | Yes | Inline, guide and example disagree on form | Not published |
| Seedance 2.5 | Yes | {} braces, inside timestamped shots | Via video edit instruction |
| Seedance 2.0 | Yes | Curly braces per symbol table | Not published |
| Grok Imagine video 1.5 | Yes, generate_audio to disable | Preset voices tagged in the prompt | Not published |
| Runway Gen-4.5 | No audio field on that branch | n/a | Act-Two is a separate endpoint |
| Runway Act-Two | Driven by a reference video | Performance transfer, not text | Yes, facial performance |
| LTX-2.5 | Yes, generate_audio: false to disable | Quotation marks, plus language and accent | Not as a parameter |
| LTX-2.3 Dub-It (beta) | Regenerates the speech | Prompt template, single speaker | Yes, explicitly |
| Vidu Lip Sync | Audio you supply, or TTS from text | audio_url or text | Yes, separate endpoint |
| Luma Ray 3.2 | Not published anywhere in its docs | n/a | Not published |
| Pika 2.5 | No audio field in the spec | n/a | Kling's, via Pika's catalog |
| Wan 2.7 | Yes, or supply audio_url | Inline in the vendor's own example | Not published for 2.7 |
| wan2.2-s2v | Audio you supply | Audio file plus one image | Yes, digital human |
| MiniMax H3 | Speech in the vendor's own example; no audio parameter published | Character speaks: plus a reference audio for timbre | Not published |
Two rows deserve a second look. Kling is the only audio-capable model here that ships silent by default: its 3.0 text-to-video spec sets settings.audio to off, with native as the opt-in, described as "native: The generated video includes native audio matching the visuals." And Luma publishes nothing at all about audio: the words audio, voice, speech and dialogue return zero hits across its Models page, its video generation guide and its API reference index.
Duration matters more than it looks here, because a line that will not fit is the most common cause of a mouth still moving at the cut. Duration parameters by model has the full matrix.
Where is dialogue prompted inline, and in what syntax?
Five vendors publish a form. They do not agree, and the differences are not cosmetic.
Google. The Veo page says: "Dialogue: Use quotes for specific speech." Its own worked example goes further and labels speakers with parenthetical direction, in the shape Man: (Hand on his hunting knife) "That's no ordinary bear."
OpenAI. The Sora 2 prompting guide contradicts itself on one page. The prose says "Dialogue must be described directly in your prompt." Then it tells you to place that dialogue in a <dialogue> block below the prose. The worked example immediately underneath uses a plain Dialogue: label with a bulleted speaker list. Both are OpenAI's. Pick the example, since it is the one they actually ran.
ByteDance. The Seedance 2.0 prompt guide publishes a symbol table: music in full-width parentheses, sound effects in angle brackets, dialogue in {}, subtitles in 【】. The 2.5 tutorial keeps all four markers, including {} for dialogue — it does not replace them — with one quiet difference: the 2.0 page prints the music brackets full-width and the 2.5 page prints them ASCII. What 2.5 adds is time. ByteDance says it outright: "Seedance 2.0 does not respond to timestamps and only responds to shot numbers, while Seedance 2.5 supports integer-second timestamps". Its own worked example opens shots as Shot 1 (0-3s):.
Lightricks. LTX asks you to place spoken dialogue in quotation marks and to "Specify language and accent if needed".
Kling. Kling's API does publish a prompt grammar, and it is a shot grammar, not a dialogue one: "shot n, m, words; shot n, m, words;" with one to six shots, each at least one second, the durations summing to the total, and 512 characters per shot. There is no documented dialogue syntax in the 3.0 Omni spec. Kling's speech story is the separate Lip Sync and Avatar endpoints.
For Veo specifically, dialogue prompting in Veo 3.1 goes considerably deeper than this page can.
How do you prompt one character speaking to camera?
The workhorse. Keep the line short, name the framing, and say who is speaking before you say what they say.
Veo 3.1 (Gemini API), native audio
Medium close-up, [SUBJECT: 30s woman, dark curly hair, olive linen shirt] seated at a
[LOCATION: sunlit kitchen table], addressing the camera directly. Soft window light from
frame left, shallow depth of field. She speaks one line, calm and unhurried:
"[LINE, 8-12 WORDS MAX]"
Ambient: faint refrigerator hum, no music.
Kling Avatar API, audio-driven
POST /v1/videos/avatar/image2video
image: [PORTRAIT URL]
sound_file: [AUDIO URL, mp3/wav/m4a/aac, 2-300s, <=5MB]
prompt: "[SUBJECT] speaks to camera, [EMOTION: warm, steady], slight head movement,
blinks naturally, hands out of frame. Camera static, medium close-up."
mode: pro
LTX-2.5, single continuous take
A single continuous take. [SUBJECT] stands in [LOCATION], facing the lens. The camera
holds still. Present tense throughout. She says, in [LANGUAGE/ACCENT]: "[LINE]".
Ambient: [ROOM TONE]. No music. No cuts.
Wan 2.7, inline dialogue in Alibaba's own house style
[SHOT SIZE] of [SUBJECT] in [LOCATION], [LIGHTING]. Its mouth clearly moving as it speaks
in a [VOICE QUALITY] voice: "[LINE]". [CAMERA MOVE]. Background: [AMBIENCE].
Lightricks makes the case for the single take explicitly: use one when you want "dialogue that must stay lip-synced in one framing." A cut mid-sentence is where sync most often falls apart.
How do you write a two-hander conversation?
Two speakers doubles every problem: identity, turn-taking and timing. Label speakers consistently, alternate turns, and cap the exchange.
Veo 3.1, speaker labels with parenthetical direction
Wide shot, [LOCATION]. [CHARACTER A: description] and [CHARACTER B: description] face each
other. [CAMERA].
A: ([PHYSICAL BEAT]) "[LINE 1, SHORT]"
B: ([PHYSICAL BEAT]) "[LINE 2, SHORTER]"
Ambient: [SOUNDSCAPE].
Sora 2, following the guide's worked example rather than its prose
[PROSE SCENE DESCRIPTION: room, light, two figures, mood. 2-3 sentences.]
Dialogue:
- [NAME A]: "[LINE 1]"
- [NAME B]: "[LINE 2]"
- [NAME A]: "[LINE 3]"
Grok Imagine video 1.5, reference-to-video with two preset voices
{
"model": "grok-imagine-video-1.5",
"prompt": "The person from <IMAGE_1> sits across from a second speaker in [LOCATION], speaking with the voice from <AUDIO_0>. A second speaker with the voice from <AUDIO_1> replies.",
"reference_images": [{"url": "[IMAGE_URL_1]"}],
"reference_audios": [{"voice_id": "[VOICE_A]"}, {"voice_id": "[VOICE_B]"}],
"duration": 8,
"aspect_ratio": "9:16",
"resolution": "720p"
}
Seedance 2.5, timestamped shots
Shot 1 (0-4s): [SHOT SIZE] of [CHARACTER A] in [LOCATION]. [CHARACTER A] says
{[LINE 1]}.
Shot 2 (4-8s): Reverse angle on [CHARACTER B]. [CHARACTER B] says {[LINE 2]}.
Shot 3 (8-12s): Two-shot, both in frame. [CHARACTER A] says {[LINE 3]}.
No subtitles. No BGM; generate only environmental sounds and action sounds.
Two ceilings are published, and they are not the same kind of thing. xAI's is a hard request limit: reference-to-video takes at most three preset voices, and an unrecognised voice_id returns a 400 with the list of valid ones. OpenAI's is softer and narrower — it is about reusable character assets, not about how many people may speak. The video guide says "A single video can include up to two characters", and the Sora 2 cookbook words the same number as advice: "We recommend no more than 2 characters per generation." Neither sentence is a documented cap on speakers.
How do you do voiceover over action without lip sync?
This is the distinction that saves the most time on this whole page. If nobody is on camera speaking, you do not need lip sync at all, and asking for it invites the exact artefacts you were trying to avoid. Narration over b-roll is an audio problem, not a mouth problem.
Veo 3.1, narration with no speaker in frame
[SHOT: hands only, product on a workbench, no face visible]. [CAMERA MOVE].
Voiceover, off-screen, [TONE]: "[LINE]"
Ambient: [SOUNDSCAPE]. No on-screen speaker.
Kling 3.0, audio explicitly on
settings.audio: native
prompt: "shot 1, 3, [SCENE A, no visible speaker]; shot 2, 4, [SCENE B]; shot 3, 3,
[SCENE C]. Ambient throughout: [SOUNDSCAPE]. No on-camera dialogue."
settings.duration: 10
Seedance 2.0, symbol table for the audio bed
[SCENE DESCRIPTION]. ([MUSIC: slow piano, low in the mix]) <[SFX: rain on a tin roof]>
No subtitles. No visible speaker.
Runway Gen-4.5, silent by design
promptText: "[SHOT SIZE] of [SUBJECT] [ACTION] in [LOCATION], [LIGHTING], [CAMERA MOVE].
No people speaking on camera."
duration: 10
ratio: 1280:720
That last one is a fact worth internalising: Runway's own gen4.5 branches carry no audio field at all, so the voiceover is a separate job you do elsewhere. For everything about the non-speech layer, sound design prompts for AI video covers it properly.
How do you prompt a reaction with no line?
Silence is an instruction, and most models need to be told, because a face in close-up plus an audio-on model is a strong invitation to invent mumbling.
Veo 3.1
Extreme close-up on [SUBJECT]'s face, [LIGHTING]. They do not speak. [MICRO-EXPRESSION:
eyes widen, jaw tightens, one slow blink]. Mouth closed throughout.
Ambient only: [SOUND]. No dialogue, no voiceover, no music.
Kling 3.0, silent by default is the feature here
settings.audio: off
prompt: "[SHOT SIZE] on [SUBJECT], [EXPRESSION BEAT 1] then [EXPRESSION BEAT 2].
Lips closed and still. [CAMERA MOVE]."
LTX-2.5, one rhythm cue instead of a soundtrack
A single continuous take. [SUBJECT] listens, saying nothing. [PHYSICAL CUE: swallows,
looks away, exhales]. Lips remain closed. One sound only: [DISTANT SOUND].
How do you direct a line with a specific emotion?
Physical cues beat adjectives on every model that documents its preference. ByteDance's own guidance for expressions is to use descriptive sentences and reduce the use of idioms. Google's example carries the direction in parentheses before the line rather than as a mood label.
Veo 3.1, direction before the quote
[SHOT]. [SUBJECT] [PHYSICAL STATE: gripping the doorframe, shoulders drawn up].
Speaking, (voice tight, barely above a whisper): "[LINE]"
Sora 2, emotion in the prose, line in the block
[PROSE: describe posture, breath, hands, eyeline. The emotion lives here.]
Dialogue:
- [NAME]: "[LINE]"
LTX-2.5, delivery vocabulary the vendor lists
[SCENE]. [SUBJECT] speaks with a [DIALOGUE STYLE: resonant voice with gravitas /
childlike curiosity / robotic monotone] at [VOLUME: whisper / mutter / shout]:
"[LINE]".
Seedance 2.5, emotion as a beat inside the shot
Shot 2 (5-9s): Facial close-up. [SUBJECT] [PHYSICAL BEAT: pupils contract, tears fall].
[SUBJECT] says {[LINE]}.
How do you get an accent or another language?
Two vendors are unusually explicit here, and one is unusually honest about the limits.
Google publishes the caveat that matters most: for Veo, "English (EN) is fully supported, but other languages have not been evaluated, so they may work but results can vary." That is a documented uncertainty, not a rumour, and it should change what you promise a client.
ByteDance's rule for Seedance 2.0 is stricter and more mechanical: "The language of dialogue must be consistent, and mixing Chinese and English should be avoided (except for proper nouns)", and an uncommon language must be named alongside the braced line.
LTX Dub-It, the vendor's own template
[SPEAKER] is speaking [LANGUAGE/ACCENT], saying: "[DIALOGUE IN NATIVE SCRIPT]"
Seedance 2.0, language marked explicitly
[SCENE]. [SUBJECT] says in [LANGUAGE] {[LINE IN NATIVE SCRIPT]}. No subtitles.
Veo 3.1, hedged deliberately
[SHOT] of [SUBJECT] in [LOCATION]. They speak one short line in [LANGUAGE] with a
[REGION] accent: "[LINE]". Keep the line under [N] words.
Runway voice dubbing, a separate endpoint entirely
{
"model": "eleven_voice_dubbing",
"audioUri": "[AUDIO URL]",
"targetLang": "[es|fr|de|ja|ko|hi|pt|zh|ar|...]",
"numSpeakers": [N],
"dropBackgroundAudio": false
}
Two honest constraints. Lightricks lists Dub-It's validated languages as English, French, Spanish, German and Russian, and states that "it does not translate automatically", so you supply the translated line yourself, in the target script. And Runway's dubbing endpoint is not a video model at all: it dubs the audio, and you still need a face that matches.
How do you replace dialogue over existing footage?
This is ADR, and it is the route with the most documented capability and the tightest constraints.
LTX-2.3 Dub-It, open weights, ComfyUI or a Python script
python -m ltx_pipelines.dub_it \
--reference-video ./[SOURCE].mp4 \
--prompt "[SPEAKER] is speaking [LANGUAGE/ACCENT], saying: \"[NEW LINE]\"" \
--height 720 --width 1280 --num-frames 161 --seed 42
Seedance 2.5, as a video edit instruction
Translate the spoken dialogue in the video into [LANGUAGE], with no subtitles. Precisely
adjust the lip movements to match the translated speech, while keeping everything else
unchanged.
Kling advanced lip sync, face selection first
{
"session_id": "[FROM FACE RECOGNITION API]",
"face_choose": [{
"face_id": "[ID]",
"sound_file": "[AUDIO URL, 2-60s, <=5MB]",
"sound_start_time": 0,
"sound_end_time": 3000,
"sound_insert_time": 1000,
"sound_volume": 1,
"original_audio_volume": 0
}]
}
Vidu Lip Sync, text-driven
POST https://api.vidu.com/ent/v2/lip-sync
{
"video_url": "[SOURCE VIDEO URL]",
"text": "[NEW LINE]",
"voice_id": "[VOICE ID FROM VIDU'S VOICE LIST]",
"ref_photo_url": "[FRONTAL FACE IMAGE OF THE TARGET SPEAKER]"
}
Lightricks describes Dub-It as a tool that "re-generates speech in video, producing lip-synced output with new dialogue while preserving the speaker's visual appearance and vocal identity", and lists among its behaviours "Preserves the full video except the lip region". Kling's endpoint states flatly that it "Currently only supports one person lip-sync." Vidu's is the same shape: when the input video contains several faces, "the lip-sync API can only select one face as the target for lip synchronization", which is why the reference photo exists. Alibaba's wan2.2-s2v takes the audio-first route instead, driving a still image so that lip movements, facial expressions and actions follow the audio.
Which lip sync failures can you prompt around?
Sort the symptom before you rewrite anything. About half of what goes wrong here is not a wording problem, and rewriting the sentence against a capability limit is the most expensive habit in this field.
Prompt-fixable.
- The line does not fit the clip. Count syllables against seconds. OpenAI's guidance is that "a 4-second shot will usually accommodate one or two short exchanges, while an 8-second clip can support a few more", and that "Long, complex speeches are unlikely to sync well and may break pacing." Cut words or extend the shot.
- Wrong speaker moves. Label speakers consistently and alternate turns, which is what OpenAI's guide asks for explicitly.
- Unwanted mumbling in a silent shot. Say lips closed, no dialogue, no voiceover. Or use Kling with
settings.audioleft atoff. - Flat delivery. Replace mood adjectives with physical cues, in parentheses before the line.
- The wrong language or accent. Name it. LTX asks for it directly; Seedance requires it for uncommon languages.
- Sync breaking at a cut. Keep dialogue inside one continuous take, per Lightricks' own advice.
- Unwanted subtitles burned in. Add the negative instruction, and lower your expectations: ByteDance says "Currently, it is not possible to directly avoid generating subtitles 100%."
Capability limits, where no rewrite helps.
- Phoneme-level accuracy. No vendor in this list publishes a phoneme fidelity claim, a target, or a control for it. There is nothing to prompt.
- Teeth and tongue artefacts. Same. Not a documented parameter anywhere.
- Multiple speakers on a single-speaker endpoint. Kling's lip sync takes one face. Vidu's takes one face. Lightricks says Dub-It "does not distinguish between multiple speakers." That is architecture, not phrasing.
- Identity drift mid-line. ByteDance documents the causes as reference-image problems rather than prompt problems, and its fix is a separate headshot with the face large in frame, not a better sentence. Keeping a character consistent across clips is the deeper treatment.
- Audio and video of different lengths. Alibaba spells out the mechanical outcome for Wan: "If the audio is shorter than the video, the remaining video is silent." Trim in an editor.
- Lip sync where the vendor documents none. Luma, Pika and Runway's Gen-4.5 branch publish no speech or lip sync capability. Asking harder does not create one.
What are you not allowed to make?
Two constraints. Both are short, both are real, and neither is a footnote.
Do not generate a real, identifiable person saying words they did not say. This is the defining harm of the technology and the platforms are explicit. OpenAI's video guide lists among its restrictions that "Real people—including public figures—cannot be generated" and that "Character uploads that depict human likeness are blocked by default." Runway's usage policy prohibits "Use of the service to impersonate an individual or entity, or to misrepresent your affiliation with an individual or entity". Read the policy of whichever platform you are on, because they differ. And note the part no policy covers: likeness, right-of-publicity and defamation law apply to you regardless of what a tool permits. A model that lets a generation through is not a legal opinion.
Label it where the platform requires labelling. Requirements vary by platform, so check the one you publish to rather than assuming a single rule. YouTube's own policy says it requires "creators to disclose when they use AI to meaningfully alter or generate photorealistic content", and lists among the cases needing disclosure content that “Makes a real person appear to say or do something they didn’t do.” Its non-disclosure list is instructive too: cloning your own voice for voiceovers or dubs sits there. Some models also mark their output automatically. Google states that videos created by Veo are watermarked using SynthID.
Where to check all of this yourself
Every claim above is on a vendor page you can read today, and all of it will move.
- Google. The Veo model-features table, the dialogue-in-quotes line and the SynthID note: ai.google.dev/gemini-api/docs/veo.
- Kling. kling.ai/document-api serves clean markdown for every page — the 3.0 Omni text-to-video spec for
settings.audioand the shot grammar, lip-sync and avatar for the two speech endpoints. - OpenAI. The video generation guide carries the guardrails and the character limit; the Sora 2 prompting guide carries the dialogue block and the timing advice. The Sora 2 API models are listed for removal on September 24, 2026 on the deprecations page.
- Lightricks. Dub-It (beta) and the prompting guide.
- ByteDance. Three separate pages, and they differ: the Seedance 2.0 prompt guide for the symbol table and the subtitle caveat, the 2.5 tutorial for the same markers restated, and the 2.5 prompt guide for timestamps and the edit instruction.
- xAI. Video generation and reference-to-video.
- Runway. docs.dev.runwayml.com/api.md for the Gen-4.5, Act-Two and voice-dubbing request shapes; the usage policy for the impersonation rule; and the help centre's Act-Two dialogue article, which is where Runway ties Act-Two to lip sync — the API reference describes it only as facial expressions and body control.
- Vidu. The Lip Sync endpoint.
- Alibaba. Wan2.7 text-to-video for
audio_urland the inline-dialogue example, and wan2.2-s2v for the digital-human route. - Also checked and found empty: Luma (no audio, voice, speech or dialogue on the Models page, the video generation guide or the API reference index), Pika 2.5 (no audio field in the spec), and MiniMax H3 (a
Character speaks:example, but no published audio parameter and no lip sync).
All read August 29, 2026.
Write the prompt once, keep it parameterised, and reuse it across the three routes rather than rebuilding it every time the model list changes. That is the whole reason a prompt library beats a folder of screenshots.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An Account