TL;DR: Every useful LTX prompt is really a prompt plus a configuration. Twenty seconds exists only on a Fast variant at 720p or 1080p and 24 or 25 fps, while retake, extend and reframe exist only on ltx-2-3-pro. The 22 templates below each carry the model, resolution, frame rate and duration they need.
One honesty note first. Prompt Architects writes the prompt, not the video. Nothing on this page renders a clip. These are words to paste into LTX, whether you go through the API, the playground, or your own GPU.
Which LTX Model Does Each Prompt Need?
Whichever one supports the duration, resolution and endpoint your shot actually requires, which is a narrower set than most LTX guides admit. Four model IDs are live: ltx-2-5-fast, ltx-2-5-pro, ltx-2-3-fast and ltx-2-3-pro. Plain LTX-2 was removed on August 16, 2026, per Lightricks' own changelog, so ltx-2-fast and ltx-2-pro now return an error.
Here is the routing that decides which template you can use, compressed to the choices that actually bind:
| If your shot needs | Use | Because |
|---|---|---|
| 12 to 20 seconds, text or image driven | ltx-2-5-fast or ltx-2-3-fast, 720p/1080p, 24 or 25 fps | The only lane where durations above 10 exist |
| 1440p or 4K output | ltx-2-5-fast or either 2.3 variant | ltx-2-5-pro stops at 1080p |
| 48 or 50 fps | Any variant except ltx-2-5-pro at 48 | Caps the clip at 10 seconds regardless |
| Retake, extend or reframe | ltx-2-3-pro only | No 2.5 variant serves those endpoints |
| More than one shot in one generation | LTX-2.5 either variant | Native multi-shot is a 2.5 feature |
| Audio-driven output up to 20 seconds | ltx-2-5-fast or ltx-2-3-pro, 720p/1080p | Length follows the audio, not duration |
Source: the LTX-2.5 and LTX-2.3 model pages and the endpoint compatibility table, accessed August 27, 2026.
Two constraints inside that table are easy to miss. Durations are discrete even integers, 6 through 20, so there is no 7-second and no 15-second request. And 20 seconds at 4K does not exist on any model at any frame rate. If you want the cross-model version of that arithmetic, the video duration parameters reference covers what every major model actually accepts.
Cost is the other reason 2.3 has not gone away. On Lightricks' published pricing page as of August 27, 2026, text-to-video at 1080p runs $0.13 per second on ltx-2-5-fast and $0.06 per second on ltx-2-3-fast. Rates change; check the page before you budget. The shape of the gap is the point: the legacy tier is the draft tier.
Can You Get 20 Seconds and Still Repair the Clip?
Yes, but only through audio-to-video, and this is the single most useful routing fact on this page. It also overturns a rule that looks airtight if you only read the duration matrix.
The apparent trap goes like this. A 20-second generation requires a Fast variant. Retake, extend and reframe require ltx-2-3-pro, which caps at 10 seconds. So a long take should be all-or-nothing, with a full reroll as the only recovery. That is exactly right for text-to-video and image-to-video, and our companion piece on structuring a 20-second LTX take treats it as the governing constraint.
The audio path is the exception. Lightricks' model pages state that duration "applies to text-to-video and image-to-video only", and that "For audio-to-video, output length follows your input audio". Then the August 18, 2026 changelog gives audio-to-video full resolution parity and says ltx-2-3-pro "accepts up to 20 seconds at 720p/1080p and up to 10 seconds at 1440p/4K". The API reference publishes the same table on the audio_uri field.
Note the asymmetry that catches people out: ltx-2-3-fast does not serve audio-to-video at all, and ltx-2-5-pro caps audio at 10 seconds even at 1080p. So among the four variants, exactly two reach 20 seconds of audio-driven video.
How Long Should an LTX Prompt Be?
Lightricks publishes two answers, and they do not agree. Reconciling them is worth doing explicitly rather than picking one and hoping.
The LTX-2 GitHub README instructs you to write "all in a single flowing paragraph", to "Start directly with the action, and keep descriptions literal and precise", and to "Keep within 200 words." The API prompting guide says something different under its own Length heading: "Match length to complexity. A simple single shot is often 4–8 sentences; a longer screenplay-style scene can run longer, provided every sentence adds concrete visual or audio detail." Its first sample prompt, a screenplay-formatted news broadcast, runs to 237 words by a direct count of the published text.
The same split shows up in shape. The prompting guide tells you to write a single continuous take as a "single flowing paragraph", and warns that for multi-shot prompts you should not use a shot list, numbered beats or screenplay sluglines "unless you also describe the cut in prose". Yet Lightricks' own long-shot blog post formats every 20-second example as a screenplay, with scene headers and character cues.
The guide resolves that one itself, if you read past the single-shot section: "When a scene involves dialogue, multiple beats, or precise timing, write it in a screenplay style, with scene headers, character cues, and quoted dialogue". A 20-second take is precisely that case. The slugline warning is scoped to multi-shot prompts, where an undescribed cut is genuinely ambiguous.
So the working rule the templates below follow:
- Six to ten seconds, no dialogue: one flowing paragraph, four to eight sentences, under 200 words.
- Twelve to twenty seconds, or any dialogue: screenplay formatting is fine, and 200 words is a soft floor rather than a ceiling.
- Multi-shot on LTX-2.5: chronological prose, every cut named in words, never a bare numbered list.
The hard limit is separate from all of that. The published OpenAPI specification sets maxLength: 5000 on the prompt field across text-to-video, image-to-video, audio-to-video, retake and extend. You will hit incoherence long before you hit 5,000 characters.
Long Single Takes: Four Templates for the 20-Second Lane
These four all require a Fast variant at 720p or 1080p and 24 or 25 fps. Nothing else reaches 20 seconds through duration.
[01 · ltx-2-5-fast · 1920x1080 · 24 fps · 20 s · text-to-video]
EXT. QUARRY ROAD - LATE AFTERNOON
Low amber sun, long shadows, chalk dust hanging in still air.
The shot opens on an extreme close-up of a woman in her fifties, deep
tan, grey hair cropped short, a faded orange hi-vis vest over a work
shirt. She squints past the lens, jaw set. She lifts a bottle, drinks,
and lowers it. The camera begins a slow steady pull back, revealing her
standing at the lip of a chalk quarry, machinery small and still on the
floor below. She turns and walks along the rim, boots printing the dust.
The camera keeps easing back until the whole quarry wall fills the
frame, the woman in the orange vest a single moving point along its top
edge. She stops at a marker post, rests one hand on it, and looks down
into the pit. The camera slows to a stop and holds as dust drifts across
the light.
Audio: dry wind, distant diesel engine idling, boots on loose chalk, no music.
[02 · ltx-2-5-fast · 1280x720 · 24 fps · 20 s · text-to-video · dialogue]
INT. NIGHT BUS - LATE
Single overhead strip light, green-blue cast, rain streaking the windows.
A man in his twenties, shaved head, black puffer jacket, sits alone near
the back with a guitar case upright between his knees. He watches the
window. A woman in a red rain jacket, hair soaked flat, sits down across
the aisle and sets a wet paper bag on her lap.
WOMAN (not looking over): "You played the corner by the station."
He looks across, surprised, then back at the window.
MAN (quietly): "Nobody stopped."
The bus turns; the light swings across both faces. She opens the bag,
takes out a pastry, tears it in half and holds one piece out across the
aisle without a word. He takes it. The camera pushes in slightly and
settles on the two of them eating in silence as the rain runs down the
glass behind.
Audio: diesel engine drone, rain on metal roof, wet tyres, no music.
[03 · ltx-2-3-fast · 1280x720 · 24 fps · 20 s · text-to-video · cheap draft pass]
A locked-off wide shot of a launderette at night, six machines along the
back wall, two ceiling tubes flickering slightly out of sync. A man in
his sixties in a brown cardigan sits on a plastic chair reading a folded
newspaper, one machine spinning in front of him. The spin cycle slows
and stops. He lowers the paper, looks at the machine, and does not move.
A woman in a grey tracksuit enters through the glass door, crosses to a
machine at the far end, and begins loading it. Neither speaks. The man
folds the newspaper, stands, opens his machine, and pulls out a single
white sheet. He holds it, looks at it, and pushes it back in. He sits
down again and reopens the paper.
Audio: fluorescent hum, a machine spinning down, a door hinge, coins in a slot.
[04 · ltx-2-5-fast · 1080x1920 · 24 fps · 20 s · text-to-video · vertical]
A vertical handheld shot follows a pair of hands from behind as they
work a lump of clay on a spinning wheel, framed tight from just above
the wrists to the wheel head. Warm side light from a window at frame
left. The clay is off-centre and wobbling. The hands close around it,
thumbs pressing down into the middle, and the wobble evens out. A wall
begins to rise between the fingers, thin and slightly uneven. One hand
lifts away, wipes on an apron, and returns wet. The wall rises higher
and straightens. The wheel slows. A wire passes underneath the base. The
hands lift the finished cylinder clear and set it on a board at frame
right, and the wheel comes to a stop.
Audio: wheel motor hum, wet clay, water in a bowl, a distant radio, no music.
Which Templates Keep Retake and Extend Available?
Every template in this section, because they all run on ltx-2-3-pro. Two of them reach 20 seconds by supplying the audio; the rest are 10-second Pro shots built to be patched rather than rerolled.
[05 · ltx-2-3-pro · 1920x1080 · 24 fps · length = input audio, up to 20 s · audio-to-video]
A single continuous take in a small radio studio at night. A woman in
her forties, dark hair tied back, headphones half on with one ear free,
sits at a foam-screened microphone under a single warm desk lamp. The
rest of the room is dark. She leans in slightly as she speaks, one hand
flat on the desk, the other turning a pen end over end. The camera holds
a medium close-up and drifts in almost imperceptibly across the take.
Her expression follows the meaning of the words: steady at first, then
softer, then still. Behind her, a small red ON AIR panel glows. She
finishes, sits back, and takes the headphones off.
[06 · ltx-2-3-pro · 1080x1920 · 24 fps · length = input audio, up to 20 s · audio-to-video · vertical]
A vertical shot of a street drummer on an upturned plastic crate at the
mouth of an underpass, lit by a hard sodium light above and behind him.
He is in his twenties, sleeveless, a towel around his neck, playing on
paint tins and a cracked cymbal. The camera starts at a low three-quarter
angle and arcs slowly around him as he plays, keeping his hands in frame.
Passers-by cross behind in soft focus. His movements land on the beat of
the supplied track; hits fall with the accents, and he holds still on the
rests. The arc completes at a frontal medium shot and stops there.
[07 · ltx-2-3-pro · 3840x2160 · 25 fps · 10 s · text-to-video · product hero]
A slow overhead shot of a matte black wristwatch on a slab of wet slate,
lit by one soft source from the upper left and nothing else. Water beads
sit on the crystal. The camera descends steadily toward the dial, the
beads growing and separating as it approaches. The second hand sweeps.
The descent stops with the dial filling the centre of the frame, the
bezel edge just inside the border, and holds there without drifting.
Audio: a faint room tone and the mechanical tick of the movement, no music.
[08 · ltx-2-3-pro · 1920x1080 · 24 fps · 10 s · text-to-video · testimonial]
A medium close-up of a man in his thirties, close-cropped beard, navy
crew-neck sweater, seated on a stool against a bare brick wall. One
large soft source off frame left is the only light, and it falls away
across the brick behind him. He looks slightly off-lens and speaks: "We
were rebuilding the same thing every quarter. Nobody had written down
why." He pauses, glances down, then back up. The camera pushes in four
inches and stops. He nods once and settles.
Audio: quiet room tone, no music, no ambience beyond the room.
[09 · ltx-2-3-pro · 1920x1080 · 48 fps · 10 s · text-to-video · high frame rate]
A tracking shot alongside a cyclocross rider carrying her bike up a
muddy bank, camera moving at her pace and at her shoulder height. She
wears a mud-spattered white and green kit, number 47 pinned at the hip.
Cold flat overcast light, no shadows. Mud sprays from her cleats with
each step. She reaches the top of the bank, drops the bike down, swings
a leg over, and clips in. The camera holds level as she accelerates away
and the frame empties to churned mud and tape.
Audio: heavy breathing, mud underfoot, chain and freewheel, distant cowbells.
How Do You Prompt Image-to-Video in LTX?
Describe the motion, not the picture. The still already carries the composition, wardrobe and light, so a prompt that re-describes them competes with the frame it was given. Lightricks positions the workflow this way in their long-shot guide: "Image-to-video gives you control over style, framing, and character consistency." If you want the general craft of that handoff, going from a still image to AI video covers it across models.
[10 · ltx-2-5-pro · 1920x1080 · 24 fps · 8 s · image-to-video]
Animate the supplied frame as one continuous take. The subject stays
where she is. She shifts her weight slightly, turns her head to look off
frame right, and holds there. The steam from the cup on the table
continues rising and bends toward the window. The curtain at frame left
moves once in a draught and settles. The camera pushes in slowly, two or
three inches, and stops before her shoulders leave the frame. No cut, no
change of location, no new light source.
Audio: room tone, a radiator ticking, faint street noise through glass.
[11 · ltx-2-3-pro · 1920x1080 · 24 fps · 10 s · image-to-video with last_frame_uri]
A single continuous take that begins exactly on the supplied first frame
and ends exactly on the supplied last frame. The camera makes one
uninterrupted move between them, easing out of the opening framing and
settling into the closing one without overshooting or reversing. The
subject moves once, early, then holds. Light and wardrobe stay identical
throughout. Nothing enters or leaves the frame that is not present in
one of the two supplied images.
Audio: one continuous ambience with no change in level.
[12 · ltx-2-5-fast · 1080x1920 · 24 fps · 16 s · image-to-video · vertical]
Animate the supplied vertical frame as one long continuous take with no
cut. The subject in the foreground begins the action already implied by
his posture and carries it through to a stop. The camera holds his
position in the lower third and drifts upward slowly across the take,
revealing more of the space above him, and comes to rest when the top of
the structure behind him is fully in frame. He completes the action,
lowers his hands, and looks up. The camera stops moving and holds for
the final two seconds.
Audio: one continuous outdoor ambience, no music, no dialogue.
When Should One Generation Contain More Than One Shot?
When the beats belong to different framings and the cut itself carries meaning. Multi-shot is an LTX-2.5 capability only: the model page describes it as "a single generation can produce multiple connected shots, holding character, scene, lighting, visual style, and voice consistent across cuts." LTX-2.3 has no equivalent, which is a real reason to stay on 2.5 for anything edited.
Lightricks' guidance is specific and worth following exactly. Prefer two to four shots, because "more cuts usually need clearer, shorter beats per shot". Name every transition in natural language. Re-establish scale, angle, who is in frame and lighting after each cut. Reuse the same visual identifiers for people who reappear. And state audio continuity at each cut, one way or the other.
[13 · ltx-2-5-fast · 1920x1080 · 24 fps · 12 s · text-to-video · two shots]
A wide shot frames a snow-covered petrol station forecourt at dusk, one
pump lit, the shop window yellow behind it. A woman in a long grey coat
and red woollen hat stands at the pump with the nozzle in the tank,
breath clouding, watching the numbers climb. Wind and the tick of the
pump fill the air. A hard cut transitions to a medium close-up of the
same woman in the red woollen hat seen from inside the shop through the
glass, the numbers reflected across it; the wind drops away and shop
radio takes over. She looks up, meets the eye of someone off frame
inside, and raises a hand in a small greeting.
[14 · ltx-2-5-fast · 1920x1080 · 24 fps · 20 s · text-to-video · three shots]
A wide establishing shot of a boatyard at first light, hulls up on
cradles, frost on the tarpaulins, a single figure crossing the yard with
a toolbox. Gulls and halyards ring in the wind. The view cuts to a
close-up of his hands opening the toolbox on an upturned crate, the same
frost visible on the metal; the gulls continue, the wind quieter behind
the hull. He lifts out a scraper and turns it over. A hard cut
transitions to a medium shot of the man, mid-fifties, oilskin jacket and
a grey watch cap, working the scraper along the hull in long strokes,
flakes of old paint falling past the lens; the wind returns to full
level and the halyards ring again. He stops, steps back, and looks up at
the length of the hull.
[15 · ltx-2-5-pro · 1920x1080 · 24 fps · 10 s · text-to-video · audio stated at the cut]
A close-up of a violin case being opened on a bed, the instrument dull
with dust, a folded programme tucked beside the neck. Quiet room tone
only, no music. A match cut connects the shape of the violin to the
shape of a hospital chart hanging at the foot of a bed in a bright ward;
the room tone gives way to a low ventilator rhythm and distant footsteps.
A woman in her thirties, dark plait over one shoulder, stands at the end
of the bed holding the folded programme from the case. She looks at it,
then folds it away into her coat pocket, and the ventilator continues
underneath.
What Do the Repair Prompts Look Like?
Short, scoped and about the section, not the film. Retake, extend and reframe run only on ltx-2-3-pro, and each carries constraints the prompt has to respect.
Retake needs an input video of at least 73 frames, around three seconds at 24 fps, and a section of at least 2 seconds defined by start_time and duration. Its mode chooses what gets replaced: replace_audio, replace_video or replace_audio_and_video, which is the default. Output is 1080p landscape or portrait.
[16 · ltx-2-3-pro · 1920x1080 · retake · mode replace_video · section ≥ 2 s]
In this section only, the man in the grey overcoat keeps both hands in
his pockets and does not gesture. His head turns once to the left,
slowly, and stops. Everything else in the frame is unchanged: same
lighting from the window at frame left, same wardrobe, same background,
same camera position. No new objects enter the frame.
[17 · ltx-2-3-pro · 1920x1080 · retake · mode replace_audio · section ≥ 2 s]
Over this section, replace the audio with quiet interior room tone, a
radiator ticking twice, and one distant car passing outside. No music,
no dialogue, no footsteps. The level stays constant and matches the
ambience either side of the section.
Extend generates 2 to 20 seconds onto the start or end of an existing clip, and preserves the input resolution. The advanced context parameter sets how many seconds of the source the model reads first, and the specification states the sum of context plus extension "cannot exceed 505 frames (~21 seconds at 24fps)".
[18 · ltx-2-3-pro · preserves input resolution · extend · mode end · 6 s]
Continue directly from the final frame with no cut. The woman stays in
the same position and finishes the movement she has begun, then comes to
rest. The camera completes its move, slows, and stops. Light, wardrobe
and background remain exactly as they are at the end of the input. The
final two seconds are still, with only the ambient motion already
present in the scene.
[19 · ltx-2-3-pro · preserves input resolution · extend · mode start · 4 s]
Generate the moments immediately before the first frame, ending exactly
on it. The scene is already established and unchanged: same location,
same light, same wardrobe. The subject arrives into the framing the
input begins with, settles, and is still by the time the input starts.
The camera is already in its opening position and does not move.
Reframe takes no prompt at all. It is a pure request, and its resolution enum is the useful part: 1:1 at 720x720 or 1080x1080, 4:5 at 720x900 or 1080x1350, 5:4, 9:16 and 16:9. That is your social repurposing path, and it exists on exactly one model.
{
"video_uri": "https://example.com/your-1920x1080-master.mp4",
"model": "ltx-2-3-pro",
"resolution": "1080x1920"
}
Which 20-Second Route Should You Take?
There are three, and they trade different things away. Most LTX writing treats the first as the only one, which is why the choice usually gets made by accident.
| Feature | Fast, text-driven | Pro, audio-driven | Stitched Pro clips |
|---|---|---|---|
| Model required | ltx-2-5-fast or ltx-2-3-fast | ltx-2-3-pro | ltx-2-3-pro |
| Reaches 20 seconds | |||
| Length set by the duration field | |||
| Needs an audio track you supply | |||
| Retake a bad section afterwards | |||
| Extend the head or tail | |||
| Reframe to 9:16, 4:5 or 1:1 | |||
| Maximum resolution | 1080p | 1080p | 4K |
| One unbroken take | |||
| Cost of four bad seconds | Full reroll | Retake the section | Reroll one clip |
The text-driven route is the one to reach for when the shot is exploratory and you have no soundtrack yet. It is the cheapest way to find out whether a 20-second idea holds at all, and rerolling it is a normal part of the workflow rather than a failure.
The audio-driven route suits anything where the sound already exists: a scripted voiceover, a music cut, a recorded line of dialogue. You lose the ability to dial a length, because the track decides, and you gain every repair endpoint. For narration-led work that is close to a straight upgrade.
Stitching wins whenever you need 4K, 48 or 50 fps, or fine control over rhythm, none of which any single 20-second generation offers. It costs you the unbroken camera move, which is often the only reason someone wanted 20 seconds in the first place.
Why Do the Templates Describe Camera Moves in Words?
Because the API cannot express a move that starts, travels and lands. The camera_motion field on text-to-video, image-to-video and, since the August 19, 2026 changelog, audio-to-video is a single enum with eight values: dolly_in, dolly_out, dolly_left, dolly_right, jib_up, jib_down, static and focus_shift. One value applies to the whole generation.
That is fine for a locked frame or a simple push. It cannot describe a pull-back that reveals a location and then stops on a specific composition, which is the shape almost every template above uses. So the move lives in the prompt text, and the parameter stays unset or set to static when you want the model to hold still.
The prompting guide gives the reason this matters more than it sounds: describing how subjects appear after the movement "helps the model complete the motion accurately". A move without a stated destination either finishes early and leaves the rest of the clip unassigned, or drifts for the whole duration. Naming what is in frame when the camera stops is the instruction that ends it.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An AccountWhat Should You Reuse Across Every LTX Prompt?
Three reusable blocks and one request body, because the parameter half of an LTX prompt is where most of the avoidable failures live.
[20 · any variant · reusable header — paste above any single take]
Single continuous take, no cuts. Present tense throughout. One named
light source, consistent for the whole clip. The camera move states
where it starts, what it passes, and where it stops. The final two
seconds are a settling action, not a new beat. No on-screen text.
[21 · any variant · reusable identity block — repeat verbatim at each beat]
IDENTITY: age and build, hair, one garment named with colour and
material, one distinguishing feature. Repeat this description in full at
every beat rather than referring back to it.
LIGHT: one source, named, with its direction. Describe its effect again
after each camera move.
AUDIO: one continuous soundscape - ambient bed, one intermittent
element, and either named music or explicitly none.
[22 · open-source pipelines · Dub-It IC-LoRA · speech replacement]
A woman speaking in Spanish, saying: "Grabamos esto tres veces antes de
que saliera bien, y la tercera fue la unica que nadie miro."
That last one has rules of its own worth stating, because it fails silently otherwise. Lightricks' guide says to provide the full dialogue text, since "the model follows the content of the prompt" and it will not translate dialogue for you, to write in the native script of the target language, and to use a single speaker only. On timing, "keep your prompt at roughly the same timing and syllable length as the original dialogue" — too long and words get skipped, too short and the delivery drags. Validated languages listed are English, French, Spanish, German and Russian.
And the request body most people get wrong, because duration is required even when it is null:
{
"model": "ltx-2-5-fast",
"prompt": "<a beat schedule, not a description>",
"duration": null,
"resolution": "1280x720",
"fps": 24,
"generate_audio": true
}
That is automatic duration, where the model picks the length from your prompt. It exists on both LTX-2.5 variants and neither 2.3 one, and Lightricks states "The result never exceeds the longest duration available for your resolution and frame rate in the table above." Two caveats from the docs. It cannot be combined with last_frame_uri on image-to-video, since a fixed last frame requires a known length. And on a prepaid account, credits are held against the maximum until the job finishes, so your balance has to cover 20 seconds even for a clip that comes back at 8.
Keep the versions that worked. The API exposes no seed and no negative prompt field, so a prompt you cannot reproduce is a result you cannot reproduce, and the prompt text is the only durable artefact you own. Treating each of these as a versioned prompt template with the variable parts marked is the whole difference between a library and a scratchpad. The same discipline applies across models: our JSON video prompt templates for Veo 3 and Grok Imagine Video 1.5 templates are built the same way.
The Short Version
An LTX prompt is never just text. It is text plus a variant, a resolution, a frame rate and a duration, and three of those four silently constrain the others. Twenty seconds through duration means a Fast variant at 720p or 1080p and 24 or 25 fps, with no repair endpoints available afterwards. Twenty seconds through audio means ltx-2-3-pro, a track you supply, and retake, extend and reframe all still on the table. Ten seconds on a Pro variant buys 4K, 48 fps and repairability, and costs you the long unbroken take. Pick the constraint you cannot lose, then pick the template that matches it.