TL;DR: Video character consistency has no complete solution in August 2026. Ranked by what actually holds: one long take beats everything, then vendor character references such as Veo 3.1's three images or Kling Elements, then first-and-last-frame chaining, then extending a clip, then repeating a text description. Shot choice covers the rest.
Everyone hits this in week one. You generate a great eight-second shot of your character, you write the next shot, and a stranger walks into frame wearing the same jacket. The face is close. It is not the same face.
The advice circulating for this problem is mostly wrong, mostly because it was written for a different model than the one you are using. So this post ranks the real strategies by how well they actually work, and states every per-model limit with the vendor page it came from. All capability claims below were checked against the vendors' own documentation on 27 August 2026 and are dated for that reason: this is the fastest-moving corner of the field.
One warning belongs at the top, because it deletes the obvious workflow on one major model.
Why does the same character change between clips?
Because nothing carries over. Two calls to the same model are two independent samples, and the only state passed between them is whatever you attached: text, images, a video, an asset ID. Text underdetermines a face by an enormous margin. "A woman in her thirties with short dark hair" describes millions of faces, and the model picks a fresh one each time.
Reference features help, but they reconstruct rather than copy. The face is regenerated at every frame, which is why identity error tends to arrive gradually rather than all at once, and why it gets worse during fast motion, occlusion and long shots. That mechanism is the subject of a separate piece on why AI video morphs and warps, and it is worth reading before you blame your prompt.
Which strategy actually works? The honest ranking
- One long take instead of several clips. Most reliable, most constrained.
- Vendor character references. Real features with documented caps, and very different per model.
- First-and-last-frame chaining. Good for continuity of framing, weaker for identity.
- Extending an existing clip. Excellent continuity, limited reach.
- A detailed text description repeated verbatim. Weakest. Degrades predictably.
- Directing around the problem. Not a fix, but the only technique available on every model.
Everything below expands one of those, in order.
1. Can you just generate one long take?
If the story fits inside the ceiling, yes, and you should. Inside a single generation the model holds identity itself and there are no seams for drift to appear in. The problem is that the ceilings are low and each one carries conditions.
Veo 3.1 generates eight-second videos, and eight seconds is compulsory rather than optional whenever you use reference images, 1080p or 4K. Kling 3.0 generates "up to 15 seconds of continuous video, with a flexible duration ranging from 3 to 15 seconds", and its Multi-Shot toggle lets one generation contain several planned camera setups. Seedance 2.5 raised its ceiling to 30 seconds specifically to present "a complete story without multi-segment stitching". LTX-2.5 offers a 20-second option, but only on the fast variant at 720p or 1080p at 24 or 25fps; every other combination caps at 10 seconds, and durations are discrete even integers.
LTX's angle here is the strongest claim any vendor makes on this topic. Lightricks describes LTX-2.5's native multi-shot as a single generation producing "multiple connected shots, holding character, scene, lighting, visual style, and voice consistent across cuts". A cut inside one generation is not the same risk as a cut between two generations.
A one-take prompt needs to be written as a timeline rather than as a paragraph, and here the version you are on matters. ByteDance is explicit: "Seedance 2.0 does not respond to timestamps and only responds to shot numbers, while Seedance 2.5 supports integer-second timestamps." Writing 0-6s on a 2.0 model is writing to nobody.
Single 30-second take, one continuous character, timestamp blocks.
0-6s: Medium shot. MAYA (see character block) walks out of a
tiled subway stairwell into rain. She pulls her collar up.
6-14s: The camera tracks left with her along shopfronts. She
checks a paper ticket, keeps walking, does not look at camera.
14-22s: She stops under a red awning. Push in to a medium
close-up as she reads the ticket properly.
22-30s: She looks up off-camera left, folds the ticket into her
coat pocket, and steps out of frame right.
Hold constant across the whole take: same face, same short
dark hair, same olive canvas coat, same paper ticket, same
overcast rainy daylight, same handheld 35mm look.
Kling, Multi-Shot on, one generation.
Shot 1 (4s): Wide. @Maya crosses an empty car park at dusk.
Shot 2 (5s): Reverse, medium. @Maya reaches the car door and
stops, keys in hand.
Shot 3 (6s): Close-up over her shoulder on the windscreen note.
Same character, same wardrobe and same lighting in all three
shots. No change of location.
2. Do character reference images actually work?
They are the best per-clip tool available, and they differ so much between vendors that generic advice is useless. Here is what each vendor publishes.
| Feature | Veo 3.1 | Kling 3.0 Omni | Seedance 2.5 | LTX-2.5 / 2.3 |
|---|---|---|---|---|
| Longest single take | 8s | 15s | 30s | 20s on fast at 720p/1080p, 24-25fps; otherwise 10s |
| Character reference in the video model | Up to 3 asset images | Elements: 2-4 images, or one 3-8s clip | Up to 30 images, but no real human faces | Not published |
| First and last frame | Yes, Veo 3.1 only | Yes | Yes, ratio must be adaptive | last_frame_uri on image-to-video |
| Extend an existing clip | +7s, up to 20 times, 720p only | Not published | Yes, 4-30s per round | Only on ltx-2-3-pro, which caps at 10s |
| Reference the asset by name in the prompt | No | Yes, @ElementName | Yes, @Image 1 / @Video 1 | No |
Each cell above traces to one of these pages, all read directly on 27 August 2026:
| Model | Claim source |
|---|---|
| Veo 3.1 | ai.google.dev/gemini-api/docs/veo and Google Cloud, "Guide video generation using asset and style images" |
| Kling 3.0 Omni | kling.ai/document-api/api/video/3-0-omni/elements, omni video generation, Element Library user guide |
| Seedance 2.5 | docs.byteplus.com/en/docs/ModelArk/2607688 and the 2.5 prompt guide |
| LTX-2.5 / 2.3 | docs.ltx.io/models/ltx-2-5 and docs.ltx.io/models/ltx-2-3 |
Veo 3.1 takes up to three reference images. Google's wording: "Provide up to three asset images of a single person, character, or product. Veo preserves the subject's appearance in the output video." Two constraints matter in practice. Reference images force durationSeconds to 8, and the text prompt that accompanies them is capped at 1,024 tokens on the Gemini API. Google also contradicts Google on what those three images can be: the Gemini API parameter table calls referenceImages "up to three images to be used as style and content references", while the Cloud documentation lists the separate style-image reference as supported on Veo 2.0 models only, not on any 3.1 variant. Both pages were live on 27 August 2026. Treat subject reference as the reliable half.
Kling goes further than anyone with Elements: a saved, reusable character asset. A multi-image Element needs at least one frontal image plus one to three additional views, so two to four images total. A video character Element is built from a single 3-8 second clip of one character, and Kling notes that only realistic humanoid figures can be customised from video. Once created, you call the character by name in the prompt with @, and Kling's own guide claims video generation supports up to seven reference characters. Bind an Element to a start frame and the cap tightens to three, with the note that "the elements must appear in the reference frames".
Seedance 2.5 has the largest reference budget of any model here, 50 assets in one request, made up of 30 images, 10 videos and 10 audio clips, against 15 assets on the 2.0 series. The prompt addresses them positionally as @Image 1, @Video 1, and ByteDance's advice is to state what each asset provides rather than assuming: "Specify what each asset provides, such as appearance, action, or timbre, and what should not be referenced." All of which is undercut for live-action work by the face restriction in the warning above. For animated, stylised or product characters, it is the most flexible reference system available.
Veo 3.1, three asset images of the same character.
Reference images: three views of the same woman, olive canvas
coat, short dark hair.
Prompt: The woman from the reference images stands at a rain-
streaked bus shelter at night, reading a paper ticket. She
looks up as headlights sweep across her face. Handheld 35mm,
shallow depth of field, sodium streetlight, light rain. Keep
her face, hair and coat exactly as in the references.
Kling Element, created once and reused.
element_name: Maya
element_description: Woman, early thirties, short dark hair,
olive canvas coat, small scar above left eyebrow, calm posture.
Images: 1 frontal, 1 three-quarter left, 1 profile, 1 close-up
of the coat collar and scar.
Kling prompt calling the saved Element.
Medium shot. @Maya sits at a diner counter at 3am, turning a
coffee cup without drinking. Neon sign through the window
behind her. She glances at the door, then back down. Keep
@Maya's face, hair and coat identical to the Element.
Seedance 2.5, omni reference with explicit asset roles.
@Image 1 provides the character's appearance and wardrobe.
@Image 2 provides the prop only, not the lighting.
@Video 1 provides the camera movement and pacing only. Do not
reference its subject, colour grade or location.
Shot: the character from @Image 1 walks the length of a
harbour wall at sunrise, carrying the bag from @Image 2. Keep
the character's face and clothing consistent for the whole
clip and avoid face changes.
3. Does first-and-last-frame chaining hold a face?
Partly. Every model above supports supplying a starting frame, and most support an ending frame too, which gives you exact control over where a shot begins and ends. Google calls it interpolation and describes the benefit precisely: it "gives you precise control over your shot's composition by letting you define the starting and ending frame".
The trick that makes this a continuity technique is chaining. Generate shot one, take its final frame, use that frame as the first frame of shot two. Seedance ships a parameter for exactly this. return_last_frame returns the final frame as a watermark-free PNG at the video's own dimensions, and ByteDance's tip spells out the intended use: "Use the last frame of the previously generated video as the first frame of the next video task to quickly generate a sequence of videos."
Two honest caveats. First, chaining preserves the pose, framing and wardrobe at the seam, and preserves identity only as well as the model reconstructs it from that one frame, so error compounds across a long chain. Second, a last frame is often the worst frame in the clip: motion blur, a half-blink, a mouth mid-word. Pick the cleanest frame you can rather than the literal last one when the tooling lets you.
On Seedance the chained frame has a second benefit worth knowing. Face-containing outputs from Seedance 2.5 and 2.0, and their returned last frames, are trusted as inputs to the same models under the same account for 30 days, which is how a face-consistent sequence is legitimately assembled on a platform that will not accept a photograph of a face.
Chaining, shot 2 of 4.
First frame: final frame of shot 1 (character mid-stride,
three-quarter back view, harbour wall, dawn light).
Last frame: character stopped at the end of the wall, facing
the water, same wardrobe and light.
Prompt: Continue the walk without a cut. She slows over the
last few steps and stops at the end of the wall. Same person,
same coat, same dawn light, no change of location.
Seedance, request the last frame for the next link.
Add to the task body: "return_last_frame": true
Then pass the returned PNG as the first_frame of the next
task, and repeat the character block verbatim.
4. Should you extend a clip instead of regenerating it?
When it is available, extending beats chaining, because the model gets the moving footage rather than a single frame. Google's implementation is the clearest documented: Veo 3.1 extends a Veo-generated video "by 7 seconds and up to 20 times", the output combines input and extension "for up to 148 seconds of video", and the mechanism is stated openly, that extend "finalizes the final second or 24 frames of your video and continues the action".
The conditions are strict. Extension is 720p only, the input must be a Veo-generated video of 141 seconds or less, videos are stored for two days, and Veo 3.1 Lite is excluded. Google also warns that voice does not extend effectively if it is not present in the last second of video.
Seedance extends too, choosing which supplied video to continue from the prompt's intent, with output duration from 4 to 30 seconds or automatic. LTX has a proper extend endpoint that adds between 2 and 20 seconds at either the end or, unusually, the start of a clip, with a context parameter controlling how many seconds of the original the model reads.
That last one carries a trap worth stating plainly, and it is the reason a 20-second LTX take is an all-or-nothing bet: retake, extend and reframe require ltx-2-3-pro, and every duration option on the Pro variants tops out at 10 seconds. The long take and the repair tools are mutually exclusive. If you generate 20 seconds on a fast variant and second 14 goes wrong, there is no extend, no retake, and no reframe. There is only a full re-roll. That trade sits underneath the longer piece on structuring a 20-second generation.
Kling is the exception here. No video-extension endpoint appears in Kling's API documentation index as of 27 August 2026, so on Kling the continuity tools are Elements, frames and Multi-Shot rather than extension.
Veo 3.1 extension, shot 1 into shot 2.
Input: the previously generated 8-second clip, 720p.
Prompt: Extend the shot. She keeps walking away down the
harbour wall as the camera holds still. Same coat, same light,
same pace. No cut, no new character entering frame.
LTX extend, 6 seconds onto the end.
video_uri: the finished clip
duration: 6
mode: end
prompt: Continue the same movement at the same speed. Do not
change the character, the wardrobe, the lens or the light.
5. Does repeating a detailed character description work?
It works a little, it is free, and you should do it. It also degrades, and it is important to know how.
A verbatim character block, pasted identically into every prompt in a sequence, does two useful things. It stops you from accidentally re-describing your character differently in shot four, which is a surprisingly common cause of drift, and it gives the model a stable set of distinctive features to reconstruct from. ByteDance's own guidance points the same way for unreferenced subjects: "For subjects without additional image or video references, describe the subject's appearance and key features in detail."
Three failure modes are predictable. First, prompt length is capped, and the cap is often tight: Veo 3.1 on the Gemini API carries a 1,024-token text input limit, and Runway's video endpoints cap promptText at 1,000 characters, so a 200-word character bible eats the budget you need for the shot itself. Second, a longer description does not converge on one face. It narrows a distribution, and past a point the extra clauses start competing with each other for the model's attention. Third, wardrobe and props hold far better than faces do, which is why a consistent silhouette can carry a sequence whose face is quietly wrong.
Seeds are the other half of this myth. Google's note on Veo is refreshingly blunt: the seed parameter "doesn't guarantee determinism, but slightly improves it". Runway documents seed on Gen-4.5 as producing "similar results" for an identical request, not identical ones. Kling's 3.0 Omni text-to-video, image-to-video and omni references contain no seed field at all. Seedance lists seed as supported on the 1.x pro models rather than the 2.x series, with the same warning attached: for the same request and the same seed, "similar results will be generated, but complete consistency is not guaranteed". A seed re-rolls a sampling path. It does not store an identity.
Write the block once and store it as a reusable template, then paste it unchanged. That is the whole technique.
Character block, pasted verbatim into every shot prompt.
MAYA: woman, early thirties, 5'6", light olive skin, short
dark brown hair tucked behind the left ear, thin scar above
the left eyebrow, dark brown eyes, slightly asymmetric smile.
Wardrobe: olive canvas coat, black crew-neck, dark jeans,
scuffed brown boots. Always the same coat, always buttoned to
the second button. Posture: upright, hands usually in pockets.
Continuity clause to append to every shot prompt.
Continuity: same person, same wardrobe, same hair, same time of
day and same weather as the previous shot. Do not restyle the
character. Do not change the coat colour. Do not age the face.
Turn a shot list into per-shot prompts, in ChatGPT or Claude.
Here is a character block and a five-shot list. For each shot,
write a single video prompt that: repeats the character block
verbatim at the top, describes only that shot's action, camera
and light, and ends with the continuity clause. Do not
paraphrase the character block in any of the five. Return the
five prompts, numbered, with no commentary.
CHARACTER BLOCK: [paste]
SHOT LIST: [paste]
6. Where is identity drift simply unsolved, and what do you do instead?
Some shots will not hold, on any model, today. A clean frontal close-up of a human face held for eight seconds is the hardest case in AI video, because it puts maximum resolution on the exact features the model is least stable about. Dialogue makes it worse, because the mouth is moving through shapes the reference never showed. Fast motion, heavy occlusion and shot-reverse-shot on two characters are all in the same category.
The answer is old and slightly deflating: direct around it. Editors have hidden continuity errors for a century, and the same grammar works here.
- Choose shot sizes that carry less identity. Wide and medium shots ask less of the face than a close-up does.
- Break the eyeline. Profile, three-quarter back and over-the-shoulder framings show less of what drifts.
- Use light against the face. Backlight, silhouette, hard side light and rain all reduce how much face is legible without reducing how much story is.
- Cut away. Hands, feet, an object, a door, a road. A cutaway between two shots of your character also resets the viewer's memory of the face they just saw.
- Keep the coat. Silhouette, colour and props read as identity to an audience faster than bone structure does, and they survive generation far better.
- Shorten the shot. Drift is a function of time on screen. Two three-second shots hide what one eight-second shot exposes.
None of that is a workaround for a broken tool. It is what the shot list should have looked like anyway.
Cutaway that carries the story without a face.
Close-up, 3 seconds. Two hands in olive canvas sleeves fold a
damp paper ticket in half, then in half again. Rain on the
back of the hands. Shallow focus, sodium streetlight, handheld.
No face in frame.
Face-light shot for a hard continuity moment.
Wide, 5 seconds. The character stands in a doorway backlit by
a corridor light, seen mostly in silhouette, coat outline
clear. She turns her head slightly toward the camera but stays
in shadow. Rain visible outside. No facial detail required.
Two shorter shots instead of one long one.
Shot A (3s): Medium, character steps off a kerb into the rain,
head down.
Shot B (3s): Cut to a low angle on her boots hitting a puddle,
then tilt up to her back as she walks away.
One last thing worth checking before you build a pipeline: OpenAI's deprecations page lists sora-2 and sora-2-pro for removal from the API on 24 September 2026, with the recommended-replacement column left blank. Do not architect a character-continuity workflow on it this month.
Prompt Architects does not generate video. It stores, versions and enhances the prompts you feed to the models that do, which is exactly the layer where a character block, a continuity clause and a per-model reference syntax want to live. The rest of the craft is covered elsewhere on this blog: Seedance's prompt syntax, and character references in Midjourney for building the reference sheet in the first place.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An Account