Back to blog
Video25 min read

Structuring a 20-Second Generation: LTX 2.5 and LTX 2.3 Prompting

LTX 2.3 prompting for long takes: where the real 20-second ceiling is, how to pace beats across it, and 18 copy-paste prompts built to hold for the full duration.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: LTX 2.3 prompting for a 20-second take only works in one lane: a Fast variant, 720p or 1080p, at 24 or 25 fps. Everything else caps at 10 seconds. A take that long needs beat pacing, one light logic, repeated identity anchors, and a camera move that ends.

One honesty note first. Prompt Architects writes the prompt, not the video. Nothing on this page renders a clip. These are words to paste into LTX, whether you run it through the API, the playground, or your own GPU.

Which LTX Model Actually Does 20 Seconds, and at What Resolution?

Twenty seconds is real, but it lives in a much narrower box than the headline suggests. Lightricks publishes a per-variant support matrix, and it is the single most useful table in their docs.

ModelResolutionFPSDurations available (seconds)
ltx-2-5-fast720p24, 256, 8, 10, 12, 14, 16, 18, 20
ltx-2-5-fast720p48, 506, 8, 10
ltx-2-5-fast1080p24, 256, 8, 10, 12, 14, 16, 18, 20
ltx-2-5-fast1080p48, 506, 8, 10
ltx-2-5-fast1440p, 4K24, 25, 48, 506, 8, 10
ltx-2-5-pro720p, 1080p24, 25, 506, 8, 10
ltx-2-3-fast720p, 1080p24, 256, 8, 10, 12, 14, 16, 18, 20
ltx-2-3-pro720p to 4K24, 25, 48, 506, 8, 10

Source: LTX-2.5 and LTX-2.3 model pages on docs.ltx.io, accessed August 27, 2026.

Read the shape of that, not just the numbers. The long-take lane is Fast-only, cinematic-frame-rate-only, and 1080p-or-below. Raise the resolution to 1440p, raise the frame rate to 48, or switch to Pro for the extra fidelity, and your ceiling halves to 10 seconds. There is also a floor: 6 seconds is the shortest value on offer, and the steps are even integers, so there is no 7-second and no 15-second option.

Is LTX-2.3 still the model to prompt for?

Mostly not, and this matters for anyone working from a guide written earlier in 2026. LTX-2.5 arrived on the API on August 11, 2026, per Lightricks' own API changelog, and the LTX-2 GitHub repository files LTX-2.3 under a heading that reads "Legacy: LTX-2.3". The predecessor, plain LTX-2, was removed from the API entirely on August 16, 2026.

LTX-2.3 is not dead, though, and one detail makes it permanently relevant to long-form work. Three endpoints still require ltx-2-3-pro and are unavailable on every 2.5 model: retake, extend, and reframe. Which produces a genuinely awkward trade, and it is the single most important structural fact on this page:

Weights are genuinely open, with one condition worth knowing before you build a business on them. LTX-2.5 is published on Hugging Face under the LTX-2.x Community License Agreement dated August 11, 2026: royalty-free use, modification, distribution and derivatives, subject to the use restrictions in Attachment A, except that entities with annual revenue of at least $10,000,000 must obtain a paid Commercial Use Agreement from Lightricks.

Why Is a 20-Second Take a Different Prompting Problem?

Because a short clip only has to be plausible, and a long one has to be governed. In six seconds a model can carry a single gesture on momentum alone. In twenty it has roughly three times as many frames to invent, with nothing anchoring the last third except whatever it decided in the first.

Lightricks names the failure mode directly in their long-shot guide. Order the prompt by actions and make sure the amount of movement fits comfortably within your selected duration, they write, because "Too much detail, and the motion will feel rushed. Too little, and you'll see unwanted pauses or drifting behaviour" (Prompting Long Shots with LTX-2, November 7, 2025).

That single sentence contains both long-take failure modes. Overload the prompt and the model compresses your beats, sprinting through in fast-forward. Underload it and the model runs out of instructions, which is when it starts looping a gesture, drifting the camera, or quietly morphing a face. If you have watched that happen and wondered why, why AI video morphs and warps covers the mechanism.

The practical consequence: your prompt is not a description any more. It is a schedule.

How Do You Pace Beats Across 20 Seconds?

Budget one distinct beat every four to five seconds, and write them in the order they happen. Four or five beats fills 20 seconds without crowding it. Seven or eight will be compressed into a rush.

A beat is a change the viewer can name: a head turns, a door opens, the camera lands, the light shifts, a second subject enters. Ambient continuation is not a beat. "Rain falls" is texture; "the rain stops" is a beat.

The structure Lightricks recommends for long shots is a four-part scaffold: a scene header for place and time, a short description for tone and atmosphere, blocking for how subjects move in sequence with the camera, and dialogue with bracketed performance cues. That maps cleanly onto a beat schedule.

[LTX long-take scaffold — paste and fill]
EXT. <PLACE> – <TIME OF DAY>
<One sentence of tone, palette, and atmosphere.>

Beat 1 (0–4s): <opening framing + the first action>
Beat 2 (4–9s): <the change that follows from beat 1>
Beat 3 (9–14s): <camera repositions; state what is now in frame>
Beat 4 (14–18s): <the reveal, reaction, or arrival>
Beat 5 (18–20s): <closing action that lets the shot settle>

Audio throughout: <one continuous soundscape, named once.>

Two things to know before you use that scaffold verbatim. First, Lightricks' current prompting guide asks you to write a single continuous take as "a single flowing paragraph" rather than a numbered list, and to reserve screenplay formatting for scenes with dialogue, multiple beats or precise timing. A 20-second take is exactly that case, so headers and cues are fine here. Second, do not leave bare numbered beats in a multi-shot prompt; that guide is explicit that shot lists and sluglines only work if you also describe each cut in prose.

Here is the scaffold filled in, as a real prompt.

EXT. HARBOUR WALL – FIRST LIGHT
Cold blue-grey dawn, flat overcast light, wet stone, muted palette.

The shot opens on an extreme close-up of a woman in her sixties, short
silver hair pushed flat by wind, a navy waxed jacket zipped to the chin.
She looks out past the lens, eyes steady, breath visible. She holds the
gaze for a moment. Then she turns her head to the left, and the camera
begins a slow pull back, revealing her standing alone at the end of a
harbour wall, the sea heaving grey behind her. She takes three unhurried
steps toward the water and stops at the edge. The camera continues easing
back until the whole wall is in frame, the town small and dark behind it.
She raises one hand against the wind, holds it there, then lowers it and
looks down at the water. The camera slows to a stop and holds on her
silhouette as gulls cross the frame.

Audio: constant low wind, water slapping stone, distant gulls, no music.

Notice what carries the last two seconds: "the camera slows to a stop and holds." Lightricks recommends exactly this. If your dialogue or performance cues end too early, add a soft closing action such as "character takes a beat and looks up" or "camera drifts upward as the wind passes" to carry the motion to the finish. Without one, the final seconds are unassigned, and unassigned time is where drift lives.

[Closing-action fragments — append one to any long take]
…the camera slows to a stop and holds on the empty frame.
…she takes a beat, exhales, and looks up.
…the camera drifts upward as the wind passes through the trees.
…he lowers his hands into his lap and the room settles into stillness.
…the last of the steam rises out of frame and the surface goes still.
…the camera eases back a few inches and stops, holding the composition.

How Do You Keep Motion Going Without Looping or Drifting?

Give the model motion that has somewhere to go. A looping artefact is usually a gesture with no terminal state, so the model repeats it because nothing in the prompt says it ends.

Two motion types survive 20 seconds reliably. Continuous processes with natural persistence, such as water, smoke, traffic, weather or crowd flow, because they are supposed to keep going. And a single long action with a clear finish, such as walking a defined distance, pouring until the glass is full, or turning a piece of work over and setting it down. What does not survive is a short repeatable gesture with no endpoint: waving, nodding, stirring, typing. Ask for twenty seconds of stirring and you get a loop.

[Continuous-process long take — rain]
A static wide shot of a narrow city street at night, seen from under an
awning. Heavy rain falls steadily through the cone of a single sodium
streetlamp, water sheeting off the awning edge in an unbroken curtain in
the foreground. Puddles fill and overflow across the camber of the road.
A car passes left to right, its headlights raking the wet asphalt, and
its tail lights recede into the dark. Rain continues. A second car passes
the other way, slower. The gutter stream widens and begins carrying
leaves toward a drain at frame right, where they gather and turn. The
rain eases, the sheet off the awning thinning to separate drops, and the
street goes quiet.
Audio: dense rain on metal and asphalt, two passing cars, a distant siren.
[Single long action with a terminal state — pour]
A tight overhead shot of a scratched steel bar top under warm tungsten
light. A bartender's hands enter frame holding a chilled bottle. He sets
a tall glass down at centre, adjusts it a quarter turn, then begins
pouring slowly, the liquid climbing the glass and a pale head building at
the rim. He tilts the bottle upright as the head reaches the top, sets
the bottle down at frame left, and waits while the foam settles and
compacts. He wipes a single drip from the base with his thumb, slides the
glass six inches toward the lens, and lifts his hands out of frame. The
shot holds on the settled glass.
Audio: pouring liquid, glass on steel, low room murmur, no music.
[Travelling motion — nothing to loop]
A locked-off side view from a moving train, camera parallel to the
window. Flat farmland runs past in the middle distance under a high grey
sky, hedgerows flickering close to the glass. A line of pylons crosses
the frame one by one, evenly spaced. The land begins to rise, and the
hedgerows give way to a cutting of dark wet rock that fills the frame and
darkens the light. The cutting ends and the view opens onto a wide
estuary, silver and still, boats grounded on mud. The train slows and the
landscape settles to a walking pace before stopping.
Audio: rhythmic rail noise slowing gradually, wind against glass.

How Do You Hold a Face, a Wardrobe, and a Light Source for 20 Seconds?

Repeat the identifiers. Once is not enough over 20 seconds, and repeating them is not redundancy; it is the only continuity mechanism you have in a text prompt.

Lightricks' long-shot guide adds a specific craft note here that is easy to miss: start in a close-up, then move out, because it grounds the scene and helps the model retain facial and material detail. Wider shots can soften likeness across all AI models, they add, so if you must go wide, have your character turn away or stay at a consistent distance from the lens.

Lighting is the other drifter. The current prompting guide asks you to keep one coherent light logic per shot, on the grounds that mixed light sources confuse the result. Over 20 seconds that means naming the source once, early, and naming its consequence again later. If you want the vocabulary for that, 40 lighting terms for AI prompts is the reference.

[Identity-anchor template — repeat the anchor at every beat]
<Shot opens on> a man in his forties, shaved head, grey wool overcoat,
tortoiseshell glasses, a thin white scar through his left eyebrow.
He <action one>. The camera <move>. The man in the grey overcoat
<action two>, glasses catching the light. The camera <move>. He
<action three>, one hand pushing the tortoiseshell frames back up his
nose. The camera <settles>, holding on the man in the grey overcoat as
<closing action>.
[One-light-logic template]
The only light source is <named source: a single window at frame left /
a bare overhead bulb / a fire in the grate>. Everything in frame is lit
by it and falls off away from it. <Subject> stands <position relative to
that source>. As the camera <moves>, the light on <subject>'s face
<changes in a way consistent with that one source>. Shadows stay <hard /
soft> throughout. No other light enters the scene.
[Wardrobe and prop lock — 20-second interview setup]
A medium close-up of a woman in her thirties, dark curly hair tied back,
a mustard corduroy shirt with the sleeves rolled twice, a plain steel
watch on her left wrist. She sits on a wooden stool against a bare
plaster wall, lit only by a large window off frame left. She looks
slightly off-lens and speaks: "I kept the first one. It's terrible. I
look at it whenever I think I've got worse." She glances down at the
steel watch, turns it half a turn on her wrist, and looks back up. The
camera pushes in six inches and stops. The woman in the mustard corduroy
shirt smiles briefly, shakes her head once, and settles.
Audio: quiet room tone, faint traffic through glass, no music.

How Do You Write a Camera Move That Has a Beginning and an End?

State where the move starts, what it passes, where it stops, and what is in frame when it stops. A move with no stated destination will either finish in four seconds and leave sixteen unassigned, or drift for the whole clip.

Lightricks' guide makes the useful version of this point: describing how subjects appear after the movement helps the model complete the motion accurately. So the terminal framing is not decoration, it is the instruction that tells the move when it is done.

There is also an API-level trap. The camera_motion parameter on text-to-video and image-to-video is a single enum with eight values: dolly_in, dolly_out, dolly_left, dolly_right, jib_up, jib_down, static, and focus_shift. One value applies to the whole generation. A 20-second move that starts, lands and holds cannot be expressed in that field, so it has to live in the prompt text. If you want the full move vocabulary, 30 camera-movement terms covers what each one actually means.

[Push-in that lands and holds]
The shot opens wide on a school gymnasium at night, empty bleachers,
half the overhead lights dead. A boy of about twelve in a red tracksuit
stands alone at the free-throw line, ball on his hip. The camera begins
a slow, steady push in along the centre line. He bounces the ball twice.
The camera keeps pushing as he sets his feet and raises the ball. It
arrives at a medium shot, chest to head, and stops there. He shoots.
The camera holds, framing locked, as he watches the ball out of frame
and his shoulders drop.
Audio: two ball bounces echoing in a large empty room, a distant hum of
failing fluorescent lights.
[Jib up that stops on a horizon]
The shot begins low, a few inches above wet sand, looking along the beach
at dawn as a shallow wave runs in and retreats past the lens. The camera
rises slowly and steadily, the sand falling away beneath it, the wave
line shrinking. It passes the height of a standing person, revealing a
long empty curve of coast and a single figure walking far off at frame
right. The camera continues rising until the horizon sits a third of the
way up the frame, then stops and holds. The figure keeps walking, small
and unhurried, toward the far headland.
Audio: waves in a slow regular cycle, wind, one gull.
[Pan that ends on a second subject]
A static medium shot of an old woman at a kitchen table, morning light
from a window at frame left, a mug of tea steaming in front of her. She
looks toward the doorway off frame right and says, "It's gone cold
again." The camera begins a slow pan right across a wall of framed
photographs, each one passing in and out of focus. The pan continues past
a doorframe and settles on a man in his seventies asleep upright in an
armchair, a newspaper collapsed across his lap, lit by the same window
light spilling through the doorway. The camera stops. He does not wake.
Audio: a wall clock, faint birdsong, one page settling.

Should You Use Multi-Shot or a Single Continuous Take?

LTX-2.5 can produce several connected shots inside one generation. Lightricks calls it native multi-shot and describes it as holding "character, scene, lighting, visual style, and voice consistent across cuts." That is not available in LTX-2.3, and it changes the arithmetic of a 20-second slot: you can spend it as one take, or as two to four shots.

Their guidance is specific. Prefer two to four shots. Name each transition in natural language, such as "a hard cut transitions to" or "the image dissolves into". Re-establish the new shot after every cut with scale, angle, who is in frame, and lighting if it changed. Reuse the same visual identifiers for recurring people. And state audio continuity explicitly at every cut, for example that the score continues or that the dialogue drops and only wind remains.

Stay single-shot, they say, when you want unbroken camera motion, intimate performance, or dialogue that has to stay lip-synced in one framing. For image-to-video from a first frame, prefer a single continuous take unless you deliberately describe a cut away from that opening image.

[Two-shot cut, 20 seconds, LTX-2.5 multi-shot]
A wide shot frames a rain-slick loading bay behind a theatre at night,
a single work light above a steel door. A young man in a soaked green
bomber jacket paces, checking a phone whose screen lights his face. Low
rain and the hum of an extractor fan fill the air. A hard cut transitions
to a medium close-up of the steel door from inside, warm and dry, a woman
in a black headset leaning against it with her eyes shut; the extractor
hum continues across the cut, the rain now muffled behind the door. She
opens her eyes, says quietly, "Two minutes," and pushes the bar down. The
door swings out into the rain and the work light floods across her.
[Three shots — establish, detail, reaction]
A wide establishing shot of a workshop at dusk, sawdust hanging in the
last light through high windows, a half-built chair on a bench. An older
woman in a leather apron and steel-rimmed glasses runs her palm along the
back rail. Warm ambient tone, a radio playing faintly. The view cuts to
an extreme close-up of the joint under her thumb, grain lit hard from the
side, a hairline gap visible where two pieces meet; the radio continues,
quieter. A match cut connects that gap to the same gap seen from above as
she sets a chisel against it. The image cuts back to a medium shot of the
woman in the leather apron, glasses pushed up on her head now, looking at
the joint with her jaw set. She exhales through her nose. The radio plays
on.
[Match cut with stated audio continuity]
A close-up of a red kettle on a gas hob, flame ring blue beneath it,
steam beginning to rise into cold morning light. Kitchen ambience, a
clock, no music. A match cut connects the rising steam to the exhaust of
a bus pulling away from a stop on a grey street; the clock is gone, city
traffic takes over. A woman in a red duffel coat, the same red as the
kettle, stands where the bus was, holding a paper cup in both hands. The
view cuts to a medium close-up of her face over the rim of the cup, steam
crossing the frame; traffic drops to a low background wash. She looks off
frame left, and something changes in her expression.

Where Do Long LTX Generations Characteristically Fail?

Five failure modes account for almost everything, and four of them are prompt-side.

Beat compression. You wrote seven beats into 20 seconds and the model played all seven, fast. Cut to four or five and let each one breathe.

Unassigned tail. The last three or four seconds have no instruction, so the model invents. Always end with a closing action.

Identity softening on the wide. The face was right in the close-up and generic by the pull-back. Anchor with repeated identifiers, and keep the subject turned or at a consistent distance when you go wide.

Light logic breaking. Two implied sources fight over the run and the key wanders. Name one source and describe its consequence at each beat.

On-screen text and chaotic physics. These are model-side, and Lightricks says so plainly: LTX-2.5 improves short-text accuracy but "exact spelling and consistency across frames are not guaranteed," so keep text short, verify it across the clip, and add critical titles or logos in post. Highly chaotic motion can still introduce artefacts; simpler, plausible motion is more reliable.

Single 20-Second Take or Stitched Short Clips: Which Wins?

Both, on different jobs. The honest comparison:

Verified against the LTX support matrix and endpoint compatibility table at docs.ltx.io, accessed 27 August 2026.
FeatureOne 20-second takeStitched short clips
Maximum resolution1080p4K
Maximum frame rate25 fps50 fps
Pro-variant fidelity available
Retake or extend a bad section
Continuity across the whole runtimeYou maintain it
Unbroken camera motion
Lip-synced dialogue in one framingPer clip only
Cost of one bad secondFull rerollOne clip reroll
Editing control over rhythmPrompt onlyFull

The single take wins when the unbroken run is the point: a continuous camera move, a performance that must not cut, dialogue that has to stay lip-synced in one framing, or a slow reveal whose whole effect is that it never cuts away.

Stitching wins on almost everything else, and for one blunt reason that the endpoint table makes unavoidable: repair. A 20-second generation is all-or-nothing, because retake and extend require ltx-2-3-pro, which will not produce 20 seconds in the first place. Build your 20 seconds from three ltx-2-3-pro clips instead and you can retake a two-second section, extend the tail, or reframe the whole thing to another aspect ratio. Those endpoints also carry their own limits worth knowing: retake needs a section of at least 2 seconds, extend runs from 2 to 20 seconds per call with the sum of context plus extension capped at 505 frames, roughly 21 seconds at 24 fps, and both need an input of at least 73 frames.

Stitching also unlocks 4K and 48 fps, neither of which exists at 20 seconds in any configuration.

The decision, then, is not about ambition. It is about whether continuity or repairability is the thing you cannot lose.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account

What Should You Reuse Across Every Long LTX Prompt?

Three fragments, and a habit.

The habit is versioning. Long-take prompts get long, and the API accepts up to 5,000 characters on the prompt field, so there is room to bloat. The version that worked is worth keeping as a prompt template with the variable parts marked, rather than as a paragraph you rewrite from memory each time. That is doubly true here, because the published API schema exposes no seed field: you cannot pin a generation and rerun it, so the prompt text is the only reproducible artefact you own. The open-source pipelines do accept a seed flag on the command line, which is one real reason to run locally.

[Reusable header — paste at the top of any 20-second prompt]
Single continuous take, no cuts. 20 seconds. One light source, named
below, consistent throughout. Present tense. Camera move states where it
starts, what it passes, and where it stops. Final two seconds are a
settling action, not a new beat.
[Reusable identity block]
IDENTITY (repeat verbatim at each beat): <age> <build>, <hair>,
<one garment with colour and material>, <one distinguishing feature>.
[Reusable audio block]
AUDIO: one continuous soundscape for the full duration —
<ambient bed>, <one intermittent element>, <music or explicitly none>.
Any dialogue in quotation marks, with language and accent named.

And two request bodies, so the parameter side is not guesswork. Note that duration is required even when it is null.

{
  "model": "ltx-2-5-fast",
  "prompt": "<your 20-second single take>",
  "duration": 20,
  "resolution": "1920x1080",
  "fps": 24,
  "generate_audio": true
}
{
  "model": "ltx-2-5-fast",
  "prompt": "<let the model decide the length from the beats>",
  "duration": null,
  "resolution": "1280x720",
  "fps": 24
}

That second one is automatic duration, added to ltx-2-5-fast on August 10, 2026 and to both 2.5 variants the following day. The model reads your prompt and picks the length, and Lightricks states the result never exceeds the longest duration your resolution and frame rate allow. It is a good fit for the beat-schedule approach in this article, because the beats are what the duration head is reading. Two caveats from the docs: it cannot be combined with last_frame_uri on image-to-video, since a fixed last frame requires a known length, and on a prepaid account credits are held against the maximum until the job finishes, so your balance has to cover 20 seconds even for a clip that comes back at 8.

The Short Version

The 20-second lane is narrow and specific: Fast variant, 720p or 1080p, 24 or 25 fps, and no way to patch the result afterwards. Inside that lane, a take that holds is a scheduled take. Four or five beats, one light source named and re-named, identity anchors repeated at every beat, a camera move with a stated stopping point, and a closing action that carries the final seconds. Everything else, and there is a lot of it, is better built from short clips you can actually fix. For the wider craft of directing a shot rather than describing one, how to direct AI video like a filmmaker picks up where this leaves off.

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account