TL;DR: A Kling 15 second video is a plain parameter, not a special mode. settings.duration is an integer enum of every value from 3 to 15, default 5, and 4K does not cost you length. The hard part is filling 15 seconds with something that sustains.
One honesty note before anything else. Prompt Architects writes the prompt, not the video. Nothing on this page renders a clip. This is a planning guide for the request you send Kling, and a record of what Kling's own documentation says and does not say.
How long can a Kling 3.0 video actually be?
Fifteen seconds, in one generation, and the specification is unusually clean about it. On POST /text-to-video/kling-3.0, the field settings.duration is typed int, defaults to 5, and its enum lists every integer from 3 through 15. Not 5 or 10. Every value.
Kling's VIDEO 3.0 user guide, dated February 6, 2026, says the same thing in prose: "The new model generates up to 15 seconds of continuous video, with a flexible duration ranging from 3 to 15 seconds." It frames the point as narrative rather than spec, closing with "Say goodbye to fragmented assembly and embrace a story with real progression and flow."
That window is a property of the model family, not of any feature you have to unlock. Here is how it lands across the 3.0 endpoints, taken from the request-body tables of each specification page.
| Feature | Kling 3.0 | Kling 3.0 Turbo | Kling 3.0 Omni |
|---|---|---|---|
| Duration enum | 3–15 | 3–15 | 3–15 |
| Duration default | 5 | 5 | 5 |
| Resolution enum | 720p, 1080p, 4k | 720p, 1080p | 720p, 1080p, 4k |
| settings.multi_shot field | |||
| multi_shot default | true | n/a | true |
| Multi-shot prompt grammar documented | |||
| settings.audio enum | native, off | not in request body | native, original, off |
| Prompt cap (chars) | 3072 | 3072 | 3072 |
Two things in that table are worth pausing on. Turbo carries the full 3 to 15 range but has no 4K option. And the multi-shot toggle defaults to true on 3.0 and 3.0 Omni, which means a single-take prompt on those endpoints is a prompt that has to actively avoid triggering shot changes.
Does asking for 4K cost you duration?
No. This is the genuinely notable fact about planning a long Kling shot, and it runs opposite to how most video models behave.
Kling's own request example on the 3.0 text-to-video page sends "resolution": "4k" and "duration": 15 in the same body. Not a footnote, not a caveat page. The vendor's canonical example is the maximum of both axes at once.
Compare that to the pattern documented across the field, which we mapped in AI video duration parameters by model: on most surfaces the ceiling on length drops as resolution rises, so the two settings are effectively fighting for the same budget. Kling 3.0 is the exception. You choose the length you need and the resolution you need, separately.
Native 4K arrived after the model itself. The API changelog dates it to April 23, 2026, and Kling's blog describes the launch as the "native 4K video generation feature for the Kling 3.0 video model series", with the pixel figure appearing only inside a partner evaluation quoted in that post, as "native 3840×2160 resolution content during the generation stage itself". Treat 3840×2160 as reported by an integration partner rather than as a line in Kling's spec, because that is exactly where it sits.
What does a 15 second generation cost?
Enough that you want the shot planned before you press generate. Kling's API pricing page publishes per-second rates, and multiplying by 15 is the whole calculation.
| Kling 3.0, no native audio | Per second | A 15 second shot |
|---|---|---|
| 720p | $0.084 | $1.26 |
| 1080p | $0.112 | $1.68 |
| 4K | $0.42 | $6.30 |
Those per-second figures are Kling's, from kling.ai/document-api/pricing/base/video, read August 27, 2026. The 15 second column is my arithmetic, not a published number. The ratio is worth internalising: 4K is exactly five times the 720p rate, so a fifteen second 4K generation costs about as much as twenty-five seconds of 720p. Iterating a shot list at 4K is an expensive way to find out the shot list does not work. Block at 720p, finish at 4K.
What does Kling not publish about a 15 second shot?
Two absences matter, and both get invented on the internet rather than looked up.
The second absence is more practical. You cannot extend a Kling 3.0 clip. The video capability map lists Video Extension as Supported only on Kling 1.6, 1.5 and 1.0, and marks it Not Supported for 3.0 Turbo, 3.0, 3.0 Omni, O1, 2.6 and 2.5 Turbo. Fifteen seconds is not a default you can nudge. It is the ceiling for a single 3.0 asset, and anything longer is an edit of multiple generations in a timeline. That changes how you plan: the shot has to resolve inside the window, because there is no continuation button.
How do you structure 15 seconds of screen time?
Three beats minimum, four if the shot has a turn in it. A five second clip carries one idea. Fifteen seconds carries a small arc, and a prompt that describes only a state will produce twelve seconds of a model looking for something to do.
The useful evidence here is Kling's own worked example. Its 15-second long-shot demonstration in the VIDEO 3.0 guide does not describe a scene. It marks time. The prompt says "At the 4th second, the camera accelerates forward with her", then "At the 8th second, the camera gradually zooms in to a medium shot", then "At the 12th second, the music and movement reach a climax", and closes with what happens "In the final 3 seconds". That is a beat sheet written as prose, at roughly four second intervals, by the vendor demonstrating its own feature.
The second example on that page opens by declaring its own form: "This is a 15-second cinematic long take, a single unbroken shot with no edited transitions." Telling the model what kind of object it is making, before describing any content, is cheap and appears to help.
Here is the template we use, derived from the shape of those examples rather than from any published Kling rule.
[Form declaration] This is a 15-second continuous shot, one unbroken take, no cuts.
[0–3s] ESTABLISH. Frame, subject, light, lens. Everything the viewer needs to
read the scene. No new information is introduced after this window.
[3–7s] FIRST CHANGE. One motion, one direction. Camera or subject, not both
equally. The shot commits to a vector.
[7–11s] THE TURN. Something enters, reverses, or is revealed. This is the beat
that stops the clip from being a loop.
[11–15s] RESOLVE. Camera settles or commits to a final framing. End on a
composition you would be happy to freeze.
Three craft rules follow from the fact that 15 seconds is three times the length most of these models were tuned on:
One camera vector, not a sequence of moves. A dolly that becomes a crane that becomes an orbit reads as three shots badly joined. If you want three moves you want Multi-Shot.
Give the shot a clock. Something in frame that changes visibly and does not repeat: liquid falling, tracks accumulating in snow, a door opening, light shifting. A walking cycle is not a clock. Repeating actions are where a long generation reveals its seams, and a non-repeating change gives the model an unambiguous arrow of time to follow.
Put subject motion across camera motion. If the camera pushes in while the subject walks toward it, you have spent 15 seconds on a shot that is over in four. Perpendicular motion sustains.
The seven components a video prompt should carry regardless of length are covered separately in the anatomy of a video prompt in seven parts; this section is only about how they get distributed across time.
When should you split into Multi-Shot instead?
The decision is not about ambition, it is about whether there is a cut in your idea.
Stay in one take when the action is physically continuous, when the camera move is the point, when two people are talking in one space, or when the subject is a single object being revealed. Kling's own multi-character dialogue examples run as single takes with four lines of dialogue inside them.
Switch to Multi-Shot when the idea contains more than one location, more than one moment in time, or needs coverage of a single beat from more than one angle. A cut you would make in an edit is a cut Kling should make in the generation.
Then check the arithmetic, because it is unforgiving. Six shots is the documented maximum and 15 seconds is the maximum total, which averages 2.5 seconds a shot. Under about two seconds a shot cannot establish framing before it is gone. Three to five shots inside 15 seconds is the range that actually reads.
What is the Multi-Shot syntax, exactly?
This is the part almost nobody publishes, because it lives in the API reference rather than the marketing guide. Kling documents a formal grammar for per-shot timing, and it is short.
The format is given as "shot n, m, words; shot n, m, words;", separated by semicolons. The 3.0 Omni page calls them "standard semicolons"; the 3.0 Turbo page calls them "half-width semicolons", which is the same instruction phrased for a different audience. The three positions are documented as:
- n — the shot sequence number. "n: shot sequence number (1–6 shots supported)".
- m — that shot's duration in seconds, with the constraint stated as "each shot ≥ 1s; sum of all shot durations must equal the total video duration".
- words — that shot's prompt, capped at 512 characters.
The whole prompt field, meanwhile, has "Maximum length: 3072 characters (recommended: ≤ 2500)", and is documented as accepting both directions of instruction at once: "The prompt can include positive and negative descriptions". There is no separate negative_prompt field on this endpoint. Six shots at 512 characters each is 3,072, which is presumably not a coincidence.
There is a second, structured way to do this, and it is going away. The legacy endpoint POST /v1/videos/text2video, whose model_name enum now includes kling-v3, exposes a multi_prompt array of objects with index, prompt and duration. Its field notes repeat the same constraints in different words: "Supports up to 6 storyboards, with a minimum of 1 storyboard" and "The sum of the durations of all storyboards equals the total duration of the current task". It also inverts the design of the new endpoint, warning that when multi_shot is true, "the prompt parameter is invalid, and the first/end frame generation is not supported".
That page carries a retirement notice: "This model or capability will be retired on September 15, 2026." If you are building today, build against /text-to-video/kling-3.0 and the string grammar, not the array.
Two more porting traps between the endpoints, both easy to miss: duration is an int on the new endpoint and a string on the legacy one, and multi_shot defaults to true on the new endpoint and false on the legacy one.
Where do long Kling generations characteristically fail?
Honestly labelled: what follows is our observation from generating, plus an inference from which features Kling chose to ship. Kling publishes no failure-mode documentation, and I am not going to invent a vendor citation for craft advice.
The inference is worth making explicit, because it is unusually legible. Kling's 3.0 release notes lead on element binding, describing it as locking a subject so that "Regardless of camera movements and scene development, the key subjects remain stable and consistent throughout". They also lead on native text, which "can automatically identify text content in uploaded images (such as signs, captions, or logos) and maintain text consistency, avoiding issues such as text displacement or blurring". Those are the two things a vendor fixes when a longer window has made them worse. Identity drift and text degradation are the failures the 15 second ceiling made expensive enough to engineer around.
In practice, five things go wrong late in a long clip rather than early:
- Identity drift after the midpoint. Faces, clothing details and product markings are stable for the first several seconds and then negotiate. Bind an element rather than describing the subject harder.
- Text and logos degrading in the back half. Same fix, different feature: supply the text in a start frame instead of asking for it in prose.
- Motion decay. The shot decelerates toward a static frame because the prompt described a state, not a trajectory. This is the clock rule above, restated as a symptom.
- Loop tells. Repeating actions resample and the seam becomes visible somewhere around the second cycle.
- Warping through the turn. The most demanding beat structurally is the one where something reverses, and it is where geometry is most likely to smear. We wrote up the general case in why AI video morphs and warps.
The debugging move for all five is the same and it is cheap: regenerate the failing beat alone as a 4 or 5 second clip at 720p. If it holds at 5 seconds and fails at 15, the problem is that the model is being asked to sustain, and the answer is Multi-Shot. If it fails at 5 seconds too, the problem is the description, and the length was never the issue.
Camera and motion control are their own subject, with four separate systems that are easy to confuse; they are covered in Kling camera control and motion control rather than repeated here. And if you are still choosing a model for a long shot, Veo 3 vs Sora vs Kling compares the three on the axes that actually differ.
18 copy-paste prompts for 15-second Kling shots
Every Multi-Shot example below has been checked so the per-shot seconds sum to the stated total. Swap the nouns, keep the timing skeleton. Each fence contains only the prompt, so it copies clean.
Single-take, beat-marked
1. Product hero, static-to-push.
This is a 15-second continuous shot, one unbroken take, no cuts. A ceramic coffee
cup on a wet concrete counter, low morning side light from a window on the left,
35mm, shallow depth of field. The camera begins in a slow push from one metre
out. At the 4th second, steam begins to rise and the camera continues forward
without accelerating. At the 8th second, a hand enters from the right and rotates
the cup a quarter turn, revealing a glazed maker's mark. At the 12th second, the
camera lifts fractionally and settles. In the final 2 seconds, the steam thins and
the frame holds on the mark, centred.
2. Character walk, lateral tracking.
This is a 15-second continuous shot, one unbroken take, no cuts. A woman in a grey
wool coat walking left to right along a rain-slick harbour wall at dusk, sodium
lamps behind her, handheld with slight weight, 50mm. The camera tracks laterally
with her, matching her pace. At the 5th second, she pulls her collar up against
the wind and the camera holds its distance. At the 9th second, she stops and turns
to look out over the water, the camera drifting past her by half a metre before
settling. In the final 4 seconds, she exhales, visible in the cold, and the camera
holds on her profile against the lamps.
3. Top-down food assembly, fixed camera.
This is a 15-second continuous shot, one unbroken take, no cuts. An overhead
top-down view of a chef's hands assembling a plate on dark slate, hard key light
from the upper right, macro. The camera holds still. In the first 3 seconds a
smear of green sauce is drawn across the slate. At the 4th second, three scallops
are placed in a diagonal. At the 8th second, a hand dusts crushed pistachio from
height and the particles fall slowly. At the 12th second, the camera begins a slow
rise. In the final 2 seconds it holds on the completed plate, symmetrical in frame.
4. Single constant-rate orbit, the one-vector rule.
This is a 15-second continuous shot, one unbroken take, no cuts. A vintage
motorcycle parked in an empty concrete underpass, cold blue ambient light, one
warm practical bulb overhead, anamorphic, 40mm. The camera begins in a low wide
and orbits the bike clockwise at a constant rate for the entire duration, never
changing speed. At the 5th second the overhead bulb flickers once. At the 10th
second the orbit reaches the tank and the reflection of the bulb crosses the paint.
In the final 3 seconds the camera completes the arc and stops square to the front
wheel.
5. Falling object, descending camera.
This is a 15-second continuous shot, one unbroken take, no cuts. A single sheet of
paper falling through a shaft of dusty light in an empty warehouse, tracked from
the side, 85mm, high contrast. The camera descends with the paper at the same
rate. At the 4th second the paper catches an air current and turns over. At the
8th second it passes out of the light shaft into shadow and the exposure adjusts.
At the 12th second it re-enters light near the floor. In the final 2 seconds it
settles on the concrete and the camera stops above it.
6. Two-person dialogue in one take.
This is a 15-second continuous shot, one unbroken take, no cuts. Two friends at a
small kitchen table, late evening, one warm overhead lamp, 35mm, static camera at
eye level, medium two-shot. The woman on the left says, "You actually did it,
then." The man opposite looks down at his hands and says, "I handed in the notice
this morning." She waits, then says, "And how does that feel?" He looks up and
says, "Terrifying. Correct, though." Natural pauses between lines, no music, room
tone only.
7. Landscape with a clock element.
This is a 15-second continuous shot, one unbroken take, no cuts. A wide landscape
of a snowfield at first light, mountains in the far background, a single figure
walking away from camera into the frame's centre. The camera begins static. At the
4th second it begins a very slow push forward. At the 9th second the figure crests
a low rise and their tracks become visible behind them, unbroken and lengthening.
At the 13th second the camera stops. In the final 2 seconds only the tracks and
drifting snow move.
8. Macro with no camera move at all.
This is a 15-second continuous shot, one unbroken take, no cuts. A glass of amber
liquid on a bar, a single ice cube, backlit by a window with venetian blinds,
100mm macro, very shallow focus. Camera fixed. In the first 4 seconds condensation
forms and one bead begins to run. At the 6th second the bead reaches the base. At
the 9th second the ice shifts and cracks audibly. At the 12th second the surface
settles. In the final 3 seconds the light through the blinds shifts slightly warmer
as if a cloud has passed.
Multi-Shot, using the documented prompt-string grammar
9. Three shots, 5 + 5 + 5 = 15.
shot 1, 5, wide establishing shot of a modern glass office lobby at dawn, empty,
cool light, camera slowly pushing in from the entrance; shot 2, 5, medium shot of a
woman in a navy suit stepping out of a lift, adjusting her cuff, camera tracking
with her; shot 3, 5, close-up of her hand placing a keycard on a reader, the light
turning green, shallow depth of field
10. Four shots, 4 + 4 + 4 + 3 = 15.
shot 1, 4, wide shot of a coastal road at sunset, a red convertible entering frame
from the left, low golden light; shot 2, 4, interior shot from the passenger side,
the driver's hand on the wheel, wind in her hair, handheld; shot 3, 4, low-angle
shot of the front wheel throwing up dust on the shoulder; shot 4, 3, extreme wide
drone shot of the car cresting the final rise, silhouetted against the sun
11. Five shots, 3 + 3 + 3 + 3 + 3 = 15.
shot 1, 3, close-up of an espresso machine portafilter locking into place, steam
and brass, warm cafe light; shot 2, 3, medium shot of the barista watching the
pour, focused, shallow depth of field; shot 3, 3, macro shot of crema forming in
the cup, top-down; shot 4, 3, medium shot of the cup sliding across the counter
toward a waiting customer; shot 5, 3, wide shot of the cafe interior, morning
light, the customer sitting down by the window
12. Six shots, 2 + 3 + 3 + 3 + 3 + 1 = 15. This mirrors the shape of the six-shot example in Kling's own legacy API docs, which uses those exact durations.
shot 1, 2, low-angle rear wide shot of a cyclist on a mountain switchback,
tracking behind; shot 2, 3, side close-up of the front wheel and the loose gravel
it displaces; shot 3, 3, first-person POV over the handlebars, the road falling
away ahead; shot 4, 3, frontal medium shot tracking backward, the rider's face set
with effort; shot 5, 3, side-on eye-level tracking with slight lateral drift, pine
forest passing; shot 6, 1, high-angle wide shot as the rider disappears around the
bend
13. Two shots, 8 + 7 = 15. Use two shots when the idea has exactly one cut in it.
shot 1, 8, a woman sitting alone in a hospital waiting room at night, fluorescent
overhead light, static camera, medium shot, she checks her phone twice and puts it
face down; shot 2, 7, the same woman outside on the steps at dawn, warm low sun,
handheld medium shot, she looks up as someone off-camera calls her name and begins
to smile
14. Six shots, 2 + 2 + 2 + 2 + 2 + 2 = 12. Not every Multi-Shot job has to run the full window.
shot 1, 2, macro shot of flour falling onto a wooden board; shot 2, 2, hands
cracking an egg into a well of flour; shot 3, 2, hands kneading dough, dusted with
flour; shot 4, 2, the dough resting under a cloth, warm kitchen light; shot 5, 2,
hands rolling the dough thin with a wooden pin; shot 6, 2, the finished pasta hung
to dry on a rack, backlit
15. Three shots, 3 + 3 + 3 = 9. A short, tight product cut.
shot 1, 3, extreme close-up of a matte black watch case rotating on a turntable,
single hard key light, black background; shot 2, 3, macro shot of the crown being
turned by a fingertip, shallow focus; shot 3, 3, medium shot of the watch on a
wrist, cuff pulled back, natural window light
Request bodies
16. A 15-second single-take 4K request, matching the pairing of 4k with 15 in Kling's own documented example.
{
"prompt": "This is a 15-second continuous shot, one unbroken take, no cuts. A ceramic coffee cup on a wet concrete counter, low morning side light, 35mm, shallow depth of field. The camera pushes in slowly and continuously. At the 8th second a hand enters from the right and rotates the cup a quarter turn. In the final 3 seconds the camera settles and holds on the maker's mark.",
"settings": {
"resolution": "4k",
"aspect_ratio": "16:9",
"duration": 15,
"audio": "off",
"multi_shot": false
}
}
17. The same 15 seconds as five shots of 3. Note that multi_shot stays true, which is also the documented default on this endpoint.
{
"prompt": "shot 1, 3, wide establishing shot of a rain-soaked city street at night, neon reflections; shot 2, 3, medium shot of a courier dismounting a bike, helmet still on; shot 3, 3, close-up of a package changing hands in a doorway; shot 4, 3, medium shot of the door closing, warm light narrowing to a line; shot 5, 3, wide shot of the empty street, the bike gone, rain continuing",
"settings": {
"resolution": "1080p",
"aspect_ratio": "16:9",
"duration": 15,
"audio": "native",
"multi_shot": true
}
}
18. The blocking pass. Fill this in at 720p, then re-run the winner at 4K with one setting changed.
Beat sheet, 15 seconds, four beats:
0-3s ESTABLISH: ____________________ (frame, subject, light, lens)
3-7s FIRST CHANGE: _________________ (one motion, one direction)
7-11s THE TURN: _____________________ (enters, reverses, or is revealed)
11-15s RESOLVE: ______________________ (final framing worth freezing)
Clock element (a visible change that does not repeat): ____________________
Camera vector (one only): _________________________________________________
Check: does the shot still work if the subject stops moving?
If no, the camera is doing nothing.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An AccountWhere to keep the shot grammar
The beat sheet above is worth more the tenth time you use it than the first, which is the argument for keeping it as a prompt template with the subject, light and clock element as variables, rather than rewriting it per shot. Fifteen seconds of 4K is a real cost per attempt, and the fastest way to reduce attempts is to stop re-deriving the structure every time.
Everything factual on this page comes from kling.ai, read on August 27, 2026: the request-body tables at /document-api/api/video/3-0-omni/text-to-video, /3-0-omni/image-to-video, /3-0-omni/video-omni and /3-0-turbo/text-to-video; the legacy endpoint at /document-api/api/video/2-1-master/text-to-video; the capability map at /document-api/guides/capability-map/video; the pricing table at /document-api/pricing/base/video; the changelog at /document-api/updates/api; the VIDEO 3.0 user guide at /quickstart/klingai-video-3-model-user-guide; and the 4K announcement at /blog/kling-ai-introduces-native-4k-video-model. Anything not on those pages is labelled above as inference or as our own craft observation. Vendor specs move; check the enum before you ship a batch.