TL;DR: Video transition prompts do two unrelated jobs. Inside a single clip you can prompt camera moves and, on some models, a described cut. Between two finished clips there is almost nothing to prompt, because that is an editing operation, and exactly one vendor documents an exception. Below: what each vendor actually documents, 28 copy-paste prompts, and the craft that makes two clips cut together.
Before anything else: Prompt Architects writes the prompt. It does not render video and it is not an editor. Everything below is words to paste into whichever model renders.
What Are Video Transition Prompts, Exactly?
Two completely different things share the name, and almost every page on this keyword blurs them.
Type one is a transition inside a single generated clip. A push-in, a whip pan, a rack focus, a crash zoom, or a described cut from one framing to another, all rendered by the model as part of one output file. This is a real prompting job with real vocabulary.
Type two is a transition between two clips. Clip A ends, clip B begins, and something happens at the join: a hard cut, a dissolve, a wipe. That is normally an editing operation, and you do it in an editor. Almost no prompt reaches across two separate MP4 files, because by the time file two is generated, file one has been written and closed. Exactly one vendor in this survey publishes a feature that does, and it generates the join rather than cutting on it.
Most people searching this phrase want type two and are trying to prompt their way to it. With one documented exception, covered below, it does not work, and the failure mode is specific: ask a model for "then it cuts to a close-up" and many will render a soft, boneless morph instead, because they are trained to produce continuous motion and a cut is a discontinuity.
Which Transitions Can You Actually Prompt Inside One Clip?
Camera moves, reliably. Described cuts, on two models that document them and a handful that tolerate them.
Camera-move vocabulary is settled and documented elsewhere: the terms and what each physically does are in the camera movement vocabulary reference, and where movement sits alongside subject, setting and light is in the seven-part anatomy of a video prompt. This section covers what those posts do not.
LTX-2.5 is the vendor in this set that names specific edits most explicitly. Lightricks' prompting guide says LTX-2.5 generates multi-shot scenes: "several distinct shots joined by explicit cuts inside one prompt". Its comparison table tells you what to write at each join, in these words: "Name the edit: hard cut, match cut, dissolve, etc." It also names what has to travel across the cut — "Re-identify subjects when they reappear; say what carries across the cut", and "At every cut, say whether music / dialogue / ambience continues or changes". The model page states the intent plainly: "Native multi-shot means a single generation can produce multiple connected shots, holding character, scene, lighting, visual style, and voice consistent across cuts." (docs.ltx.io, accessed August 29, 2026.)
Kling 3.0 documents shots but not transitions. Its API spec gives a multi-shot prompt format, shot n, m, words; shot n, m, words;, with n as the "shot sequence number (1–6 shots supported)", m as duration in seconds, and each shot prompt capped at 512 characters. It tells you where the cuts land. It publishes no vocabulary for what happens at them — Kling's user guide says the model "will automatically plan scene transitions, shot framing, and camera angle changes based on the prompts", which is the opposite trade: it takes the decision off you rather than handing you words for it (kling.ai/document-api and kling.ai/quickstart, accessed August 29, 2026).
Worth knowing before you paste anything into Kling: on the 3.0 and 3.0 Omni text-to-video spec, settings.multi_shot defaults to true. Want one unbroken take? Set it to false: "When set to false, multi-shot prompts will not produce multi-shot output".
Seedance 2.5 is the second vendor that documents transitions, and it is the only one that reaches across two finished clips. ByteDance's own prompt guide instructs: "For transition shots, clearly specify both the trigger point and the transition method", and asks you to "include both the transition timing and method" — its worked example names a left wipe combined with a natural dissolve at the five-second mark. Its capability table then lists a two-video transition entry, under the #video-transition anchor, that "Takes two input videos and generates the missing in-between segment" (docs.byteplus.com/en/docs/ModelArk/2607689, accessed August 29, 2026). Read that carefully: it generates the missing middle, which is not the same thing as placing a cut. It is the single documented case in this survey of a prompt operating on two separate files.
Everywhere else, a described cut is undocumented behaviour. It sometimes works. Label it as an experiment, not a feature.
In-clip prompts (copy-paste)
# 1 — LTX-2.5 — in-clip — push in
A single continuous take. Medium shot of a woman at a workbench, warm tungsten
key from frame left. The camera pushes in slowly and steadily toward her hands
as she sets down a soldering iron. Ambient room tone and the faint hum of a
bench fan throughout.
# 2 — Veo 3.1 — in-clip — pull out reveal
A close-up of a single brass key on a wooden table, shallow depth of field.
The camera pulls back steadily to reveal the key is one of forty hanging on a
hotel key board, cold fluorescent overhead light, empty lobby behind.
# 3 — LTX-2.5 — in-clip — whip pan
A single continuous take. A man stands at a bus stop reading a paper timetable
in flat overcast light. A fast whip pan to the right blurs the street into
horizontal streaks and settles on the arriving bus, motion blur resolving as
the camera stops. Traffic noise rises through the pan.
# 4 — LTX-2.5 — in-clip — match cut on shape
A close-up of a chrome ceiling fan turning slowly in a still bedroom, dawn
light through slatted blinds. A match cut connects the spinning fan to a
helicopter rotor at the same size and position in frame, now in hard daylight
over open water. The rotor holds the same rotation direction and speed.
# 5 — LTX-2.5 — in-clip — match cut on motion
A wide shot of a child throwing a red ball upward in a park, camera tilting up
to follow it against a pale sky. A match cut on the upward motion connects to
a rocket rising through the same patch of sky, the tilt continuing unbroken.
The park ambience drops out; only engine rumble remains.
# 6 — LTX-2.5 — in-clip — dissolve for elapsed time
A wide shot of an empty restaurant dining room at 5pm, chairs still stacked,
low golden light through the front window. The image dissolves into the same
room at 9pm, every table full, warm interior lamps on, chairs down. Camera
position and framing are identical across the dissolve. The room tone builds
from silence to conversation across the transition.
# 7 — LTX-2.5 — in-clip — hard cut with audio continuity
A medium shot of a cellist mid-phrase in a small rehearsal room, single window
light. A hard cut transitions to a low-angle close-up of her bow hand at the
strings; the cello line continues across the cut without interruption and the
lighting stays identical.
# 8 — Runway Gen-4.5 — in-clip — rack focus
A static locked-off medium shot. A paper receipt in sharp focus in the
foreground, a woman blurred at the far end of the room behind it. Focus racks
from the receipt to her face over about two seconds; the foreground goes soft
as she comes sharp. No camera movement.
# 9 — Seedance 2.5 — in-clip — crash zoom
A wide shot of a suburban street at midday, one front door visible at the far
end. A fast crash zoom in to a tight shot of the door handle, the zoom
completing in under a second and holding steady. Handheld micro-shake on
arrival. Bright flat daylight throughout.
# 10 — Kling 3.0 — in-clip — orbit
A single continuous take. A ceramic vase on a matte grey plinth in a white
studio. The camera orbits slowly clockwise around the vase at plinth height,
completing about a quarter turn. Soft top light, no visible fixtures, shadow
travelling across the plinth as the camera moves.
# 11 — Veo 3.1 — in-clip — reveal from behind a foreground object
The camera tracks slowly right, passing behind a thick concrete pillar that
fills the frame and blacks it out for about half a second, then continues to
reveal a man waiting on the platform beyond. Cold underground station light,
tiled walls, distant train rumble.
# 12 — Luma Ray 3.2 — in-clip — time-of-day shift in one take
A locked-off wide shot of a mountain ridge. The light travels from blue
pre-dawn through low golden sunrise to flat mid-morning, clouds moving right
to left across the frame. Framing does not change; only light, shadow length
and cloud position do.
# 13 — Seedance 2.5 — in-clip — scale shift, macro to wide
Extreme macro on a single water droplet on a leaf, surface tension visible.
The camera pulls back continuously and smoothly through the leaf, the branch,
the whole tree, ending on a wide shot of a single tree in an open field. Soft
overcast light held constant through the whole move.
# 14 — Kling 3.0 Omni — in-clip — three shots, documented multi-shot syntax
shot 1, 4, A wide shot of a rain-slick harbour at dusk, one figure walking
toward the moored boats, sodium lamps reflecting on the water;
shot 2, 3, A medium shot of the same figure untying a mooring rope, hands wet,
lamp light from frame right;
shot 3, 3, A low-angle close-up of the rope slipping free of the cleat, water
moving in the background.
What Does Each Vendor Actually Document?
Frame conditioning and clip extension, mostly. Transition vocabulary, almost never.
The single clearest cross-vendor evidence sits in Runway's own developer API, which fans one image_to_video endpoint out into thirteen model-specific request shapes and publishes a separate schema for each. Runway's own Gen-4.5 accepts position values of first only, and so do the grok_imagine_1_5 and gemini_omni_flash branches. The veo3.1 and seedance2 branches on the same endpoint accept first and last. Same request shape, same vendor, different capability, because the capability belongs to the model, not the API (docs.dev.runwayml.com/api.md, accessed August 29, 2026).
| Model (surface) | Pin first frame | Pin last frame | Extend an existing clip | Transition vocabulary published |
|---|---|---|---|---|
| Veo 3.1 (Gemini API) | Yes, image | Yes, lastFrame | Yes, +7s, up to 20 times | No |
| Kling 3.0 / 3.0 Omni | Yes, first_frame | Yes, last_frame | No, extension is 1.0/1.5/1.6 only | No, but multi-shot syntax is |
| Sora 2 (OpenAI) | Yes, input_reference | Not documented | Yes, up to 20s each, 6 times | No |
| Seedance 2.5 (BytePlus) | Yes | Yes | Yes, plus return_last_frame | Yes, method plus timing |
| Wan 2.7 (Alibaba Cloud) | Yes | Yes | Yes, video continuation | No |
| LTX-2.5 | Yes | Yes, last_frame_uri | No, extend is 2.3-pro only | Yes, named edits |
| Luma Ray 3.2 | Yes, start_frame | Yes, end_frame | Yes, forward and backward | No |
| Vidu Q3 | Yes | Yes, start-end2video | No, extend is Q2 only | No |
| Runway Gen-4.5 | Yes | Not supported | No | No |
| Pika 2.5 Pikaframes | Yes | Yes, 2 to 5 keyframes | Not documented | Partially, transition_duration_s |
| Grok Imagine 1.5 | Yes | Not documented | Not on 1.5 — /v1/videos/extensions exists, but every documented example passes grok-imagine-video | No |
Verified August 29, 2026 at each vendor's own documentation. Where a cell says "not documented", that is what the vendor publishes, not a claim that the behaviour is impossible.
One dated caveat on that Sora row: OpenAI's deprecations page records the "removal from the API on September 24, 2026" of the Videos API and the Sora 2 aliases. That extension route has weeks left on it, not years — read what to use instead before you build on it.
| Feature | Veo 3.1 | Kling 3.0 Omni | Sora 2 | LTX-2.5 |
|---|---|---|---|---|
| Pin the last frame | Yes, lastFrame | Yes, last_frame | Not documented | Yes, last_frame_uri |
| Continue a finished clip | Yes, +7s per extension | No | Yes, up to 120s total | No, 2.3-pro only |
| Cuts inside one generation | Not documented | Yes, shot syntax | Not documented | Yes, named edits |
| Names specific transitions | No | No | No | Yes: hard cut, match cut, dissolve |
How Do You Hand Off From One Clip to the Next?
By deciding what the last frame of clip A and the first frame of clip B are, before you generate either.
Three mechanisms exist, and they are not interchangeable. A fourth, Seedance 2.5 generating the missing middle between two supplied videos, is the one covered above, and it is the only route here that starts from two finished clips rather than a still.
Frame conditioning. You supply the start image, the end image, or both, and the model fills the middle. Google documents this as interpolation: lastFrame is "The final image for an interpolation video to transition. Must be used in combination with the image parameter", and the feature "gives you precise control over your shot's composition" by letting you "define the starting and ending frame". Luma generalises it furthest: start_frame and end_frame pin the ends, while multi-keyframe image-to-video lets you pin up to 64 images at chosen frame positions, "so you choreograph the motion beat by beat instead of hoping a single prompt lands". If you are starting from a still, the still-to-video handoff guide covers the input side in more depth.
Extension. The vendor continues a clip it already generated. Google: "Use Veo 3.1 to extend videos that you previously generated with Veo by 7 seconds and up to 20 times", and "Extend finalizes the final second or 24 frames of your video and continues the action". OpenAI: "Use extensions when you want to preserve motion, camera direction, and scene continuity." xAI puts it on its own endpoint, /v1/videos/extensions, returning "a single video that picks up seamlessly from the last frame of the input and continues with the generated content" — its duration there sets the extended portion only, not the finished file. The catch elsewhere is ownership: the "Gemini API only supports video extensions for Veo-generated videos", so extension is not a way to continue footage from somewhere else.
Last-frame export. BytePlus is the quiet standout. Its capability table lists "Return the last frame of the generated video" against every Seedance model, exposed as return_last_frame and returned as last_frame_url — the whole hand-off in one parameter: generate clip A, take the returned frame, feed it in as clip B's first frame. Alibaba documents the same intent differently: "Wan 2.7 - image-to-video supports first-frame-to-video, first-and-last-frame-to-video, and video continuation", while "The image-to-video (based on first frame) feature for Wan 2.6 and earlier models supports only first-frame-to-video".
Hand-off and between-clip prompt pairs (copy-paste)
# 15 — Veo 3.1 — between-clips — clip A, matched motion direction
Medium shot, camera static. A woman in a navy coat walks from frame left to
frame right along a canal path in flat morning light, passing the camera and
exiting frame right. She does not look at camera. Water and distant traffic in
the background.
# 16 — Veo 3.1 — between-clips — clip B, matched motion direction
Medium-wide shot, camera static, same flat morning light and same colour
temperature as the previous shot. The same woman in a navy coat enters from
frame left already walking at the same pace and crosses to frame centre, where
she stops at a green metal gate.
# 17 — Kling 3.0 — between-clips — clip A, cut on action
Medium shot of a man in a workshop reaching for a hanging coat with his right
hand. The shot ends with his hand closed on the collar and the coat starting
to lift off the hook. Warm side light from a single window frame right.
# 18 — Kling 3.0 — between-clips — clip B, cut on action
Close-up on the same man's shoulder, same warm side light from frame right.
The shot begins with the coat already mid-lift and continues into him swinging
it onto his shoulder in one motion. Same wardrobe, same workshop background
detail behind him.
# 19 — Seedance 2.5 — between-clips — clip A, matched lighting, request last frame
A wide shot of a chef plating a dish at a stainless pass, hard overhead
kitchen light, cool white balance. Camera static. The shot ends with the plate
finished and both hands withdrawing from frame.
(API: set return_last_frame so the final frame comes back for clip B.)
# 20 — Seedance 2.5 — between-clips — clip B, first frame is clip A's last frame
Continue from the supplied first frame. The same plate under the same hard
overhead kitchen light and cool white balance. A server's hands enter from
frame right, lift the plate, and carry it out of frame right. Camera static,
identical framing.
# 21 — Luma Ray 3.2 — between-clips — forward extend
Forward extend the supplied generation. The cyclist continues along the ridge
road at the same speed and in the same screen direction, camera holding the
same distance and height. Light continues to warm as the sun drops. No cut, no
change of framing.
# 22 — Luma Ray 3.2 — between-clips — designed end frame for the next clip
End the clip on a clean, still, well-lit frame: the cyclist stopped at the
roadside, bike upright, full body in frame centre, sky occupying the top
third. Hold that composition for the final beat with no motion blur, so the
final frame can be reused as the opening frame of the next shot.
# 23 — Pika 2.5 Pikaframes — between-clips — keyframe chain
(Supply 3 ordered keyframes: empty studio, half-built set, finished set.)
Prompt applied to every transition: the camera holds the same locked-off wide
framing and the same soft north light while the set assembles; crew members
move quickly through frame but the walls, floor and camera position stay put.
# 24 — Vidu Q3 — between-clips — start and end frame hand-off
(Supply two images with closely matched aspect ratios: start frame is the last
frame of the previous clip; end frame is the opening composition of the next.)
The camera drifts slowly right and the light warms from cool daylight to low
sun over the course of the shot. No subject enters or leaves frame.
# 25 — Runway Gen-4.5 — between-clips — clip A, respecting the 180-degree line
Two people at a cafe table, camera on the near side of the table looking
across it. The woman sits at frame left facing right, the man at frame right
facing left. Medium two-shot, soft window light from frame left. She begins
speaking.
# 26 — Runway Gen-4.5 — between-clips — clip B, same side of the line
Single close-up of the man from the same side of the table as the previous
shot, so he still faces frame left. Soft window light still from frame left,
same colour temperature. He listens, then begins to answer. Background
matches the cafe interior behind his shoulder.
# 27 — Sora 2 — between-clips — extension prompt
Continue the scene. The camera keeps rising at the same rate and in the same
direction, clearing the rooftop line and revealing the sunrise beyond. The
same wind noise continues; no new subject enters frame.
# 28 — Veo 3.1 — between-clips — extension prompt
Extend this video. The paraglider continues its descent at the same speed and
the same screen direction, the camera holding its distance. Light and colour
grade stay identical to the input clip. Do not change framing or introduce a
cut.
What Craft Makes Two Clips Actually Cut Together?
Four things, all decided before you generate, none of which any vendor documents because they are film craft rather than API behaviour.
Match motion direction. If a subject exits frame right in clip A, they should enter frame left in clip B. Reverse it and the audience reads the subject as having turned around — the most common reason AI sequences feel wrong without viewers being able to say why.
Cut on action. A cut placed mid-movement is close to invisible; a cut between two static moments announces itself. End clip A partway through a gesture and start clip B partway through the same gesture, from a different framing. Prompts 17 and 18 are built on exactly this.
Respect the 180-degree line. Pick a side of the imaginary line running through your two subjects and keep the camera on it. Cross it and the two people appear to face the same direction. Models have no concept of the line, so describing camera position consistently across both prompts is entirely on you.
Change the framing enough. Two shots at nearly the same size and angle produce a jump cut, which reads as an error unless you meant it. Change shot scale by a clear step, or the angle by more than about thirty degrees.
Two more matter specifically for generated footage: hold your light logic constant, because a colour-temperature shift across a cut is far more visible than a framing one, and keep your subject description byte-identical between the two prompts. On that second point, keeping a character consistent across clips goes deeper than this post can.
And a dissolve is not a texture, it is a meaning. A hard cut says the next thing happens immediately; a dissolve says time passed or the place changed. Use it when time really has moved. Audio carries a comparable load at every join, which is why LTX asks you to state whether music continues or changes; music direction in video prompts covers that side.
What Should You Not Expect From Any of This?
Frame-accurate control, editing grammar, or consistency between vendors.
No vendor here publishes frame-accurate timing for a described cut. Luma comes closest with explicit keyframe indices on a 24fps grid, but that pins where a supplied image lands, not where a cut falls. Kling's shot syntax takes whole seconds. Nothing in these APIs gives you a cut on a named frame.
No model "understands" editing grammar either. LTX documents multi-shot behaviour and gives you vocabulary for it, which means the training deliberately covered it — a very different claim from the model reasoning about continuity. It is pattern completion that happens to have been taught the pattern.
One observation of my own, not documented behaviour: prompts that describe the state of the frame at the end of a shot produce more reusable last frames than prompts describing only the action. "Hold that composition for the final beat with no motion blur" is in nobody's docs. It has consistently given me cleaner frames to hand to the next clip, and a blurred final frame is a bad first frame for anything.
None of this is an editing workflow. If your sequence involves presenters or avatar footage, avatar video prompt templates covers that path. The cut itself still happens in your editor.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An Account