TL;DR: Pika's current model is Pika 2.5, and its API publishes eleven Pika-owned operations. Text-to-video accepts a prompt and is locked to five seconds. Pikaffects, the feature Pika is famous for, accepts no prompt at all: you pick one of 62 named effects. Templates below, checked August 27, 2026.
Everything numeric on this page was read on August 27, 2026 from Pika's own properties: the developer documentation at dev.pika.art, its public model catalog API, the pricing page and the consumer FAQ on pika.art. Where two Pika pages disagree, and they do, both readings are shown. No third-party aggregator was used for any figure, because for this product they are all a model version or two behind.
What do Pika prompt templates actually control?
Less than you would guess from the search results, and in a way that is worth knowing before you write anything.
Pika's public catalog lists 11 Pika-owned operations inside a wider catalog of "133 callable operations" spanning seventeen vendors. Seven of the eleven are video, four are audio. The video seven split cleanly into three groups by how they treat free text:
| Feature | Pika 2.5 text-to-video | Pika 2.5 image-to-video | Pikaframes | Pikaffects |
|---|---|---|---|---|
| Accepts a prompt | ||||
| Prompt is required | ||||
| negative_prompt field | ||||
| seed field | ||||
| Resolution you choose | 720p or 1080p | 720p or 1080p | 720p or 1080p | Not a parameter |
| Length you choose | 5s only | 5s or 10s | 1-10s per transition | Not a parameter |
| Aspect ratio parameter | ||||
| What decides the look | Your words | The image, then your words | The keyframes | A named effect |
Three rows in that table decide how you write.
Nothing here has an aspect ratio parameter. Not one of the seven. On image-to-video, Pikaframes, Pikaswaps and Pikadditions the frame shape comes from the file you upload. On text-to-video the only lever you have is the sentence itself, which is exactly why Pika's own worked example prompt in the documentation ends by naming a square composition in words.
Duration is a setting, not a prompt clause. Writing "a ten second shot" into a Pika 2.5 text prompt does nothing, because the schema pins duration_s to a constant of 5 and describes it as "Video length in seconds (5s only for text-to-video)." Length belongs in the request, not the sentence. That is true across most current video models, and our video duration reference has the cross-vendor version.
A negative prompt is real here. Five of the seven publish negative_prompt, described as "Text describing what to avoid in the video." That is genuinely unusual right now: our negative prompt support matrix records Kling's newer 3.0 endpoint, Veo 3.1 on the Gemini API, Seedance and Luma Ray 3.2 as publishing no such parameter at all. If you have been trained by those models to stuff exclusions into the positive prompt, stop doing it on Pika and use the field.
Which Pika model are you prompting in 2026?
Pika 2.5. That is the only Pika video model in the catalog, it is what the generator on the homepage offers, and it is what every plan on the pricing page grants access to.
Getting this right matters more than usual, because Pika's own consumer FAQ has not kept up. That page still introduces the model family with a paragraph saying Pika 2.2 "introduces Pikaframes and generations up to 10 seconds" and tells you to "Choose your generation length from 1-10 seconds for ultimate control." It also carries a heading, "Different Pika models have different outputs:", followed by a resolution list running from Pika 2.2 down to Pika 1.5, which it calls the "proud parent of Pikaffects." None of those version strings exist in the API today.
The same FAQ answers the question "Do you have a publicly accessible API?" with "We’re flattered you asked! We’re taking on select partners." The link it offers now redirects to dev.pika.art, which is an open, self-serve developer site with a public catalog endpoint that needs no key at all. The answer and its own link target disagree.
Two more names in that pricing line are worth flagging: Pikascenes and Pikatwists appear on the consumer product and appear zero times anywhere in the developer documentation. They are not in the catalog, not in llms.txt, and not in any operation spec. If you are writing prompts against the app, they exist. If you are writing prompts against the API, they do not. Scope your templates to the surface you are actually using.
Why do Pikaffects accept no prompt at all?
Because they are not a prompt feature. They are a menu.
This is the single most useful thing to know about prompting Pika, and it is invisible from the outside. The Pikaffects image endpoint takes exactly two required fields, an effect name and a source image URL, plus an optional seed. Its input schema sets additionalProperties to false, which means a prompt key is not ignored, it is rejected. The effect name is an enum described as "Name of the effect to apply to the image", and the operation itself is summarised as "Playful physics effects applied to a still image."
There are 62 named image effects and 9 named video effects. Here is the full list, which is the part of this page worth bookmarking:
PIKAFFECTS — image → video (62 effects, pick exactly one)
90s Me Epic Me Melt
Action Me Everythings Bonsai Mona Me
Arcade Winner Explode Mrs Me
Baby Me Eye-pop Museum Me
Bald Me Eyes Zoom in Peel
Balloonify it Fairytale Me Pick Me
CCTV Me Goth Dream Poke
Cake-ify Hazmat Fit Princess Me
Captain Me Hearts Bouquet Proposal
Choco Me Hero Me Puppy Me
Classy Me Human Pet Rose
Clown Fit Inflate Royal Me
Crazy in Love Jungle Me Squish
Crumble Labubu Style Ta-da
Crush Leprechaun Me Tarot Transform
Cupid Strike Levitate Tear
Decapitate Lo-Fi Me VIP Me
Deflate Looong Hair Warrior Me
Dissolve Love Bomb Yarn Charm
Doom Stroll Magic Polaroid Zen Me
Eat a Rat Make it Real
PIKAFFECTS — video → video (9 effects, pick exactly one)
Anime Cat Happy Asteroid Pink Hair
Cute Shroom Its Alive Swan Head
Duplicate it Its Computer Wizard Cat
Read as prompt engineering, that list is a hard boundary. If the transformation you want is in it, your entire creative decision moves upstream into the source image, because the image is the only variable you control. If it is not in the list, no wording gets you there and you should be generating the effect as a normal text-to-video shot instead.
That reframes the work usefully. For a Pikaffects run, the thing to write carefully is the prompt for the still you feed it, and the discipline is the one in our image to video handoff guide: make the subject unambiguous, isolate it against a background that will survive being deformed, and leave physical headroom in frame for whatever the effect is about to do.
PIKAFFECTS SOURCE-IMAGE BRIEF — write this for your image model, not for Pika
SUBJECT: [one object or one person, named plainly, centred]
SEPARATION: [what puts it clearly in front of the background]
HEADROOM: [empty space on the side the effect will expand or collapse into]
SURFACE: [the material that has to read as real while it deforms]
FRAME: [shot size and what is deliberately excluded]
EFFECT: [the Pikaffect you will pick afterwards, so you brief the still for it]
How do you write a Pika 2.5 text-to-video prompt?
Dense, descriptive, single-paragraph, with the framing written into the words. Pika's own worked example in the specification is a good calibration point: it is roughly 110 words, runs subject and wardrobe and setting and colour grade in one continuous block, and closes by naming both the composition and a short list of exclusions.
Since there is no aspect ratio field and no separate style field, everything lives in the sentence. Six slots is enough:
PIKA 2.5 TEXT-TO-VIDEO TEMPLATE — one paragraph, no line breaks in the real prompt
SUBJECT: [who or what, plus two details a stranger could not guess]
ACTION: [one continuous motion that can finish inside five seconds]
SETTING: [where, and what is behind the subject]
CAMERA: [shot size + one move, or "static" if you want none]
LIGHT/GRADE: [source, direction, and the film or video look]
FRAMING: [aspect described in words, e.g. "vertical 9:16 composition"]
SETTINGS: resolution [720p|1080p] · duration_s 5 · seed [optional]
NEGATIVE: [send exclusions in negative_prompt, not in the sentence]
Six ready prompts. Each one is a single paragraph in the request body, written to finish its action inside five seconds:
1. PRODUCT, TABLETOP
A matte ceramic pour-over dripper on a walnut counter, steam rising in a thin
ribbon as water spirals into the bed of grounds. Morning light from a window at
camera left rakes across the glaze and throws a soft shadow to the right. Slow
push in on a locked tripod, no handheld drift. Warm daylight grade, gentle film
grain, shallow depth so the kettle behind is a soft shape. Vertical 9:16
composition, product centred with headroom above.
2. STREET, DOCUMENTARY
A courier in a rain-darkened jacket wheels a bicycle across wet cobbles at dusk,
tyres throwing a fine spray. Sodium street lamps behind her put an orange rim on
her shoulders while the sky holds the last cold blue. Camera tracks laterally at
walking pace, eye level. Handheld with a small amount of weight, 35mm look,
lifted blacks, visible grain. Wide 16:9 composition.
3. INTERIOR, CHARACTER
An older man in a knitted cardigan sets a chess piece down on a worn board and
leaves his fingers resting on it, thinking. Lamplight from the right, the rest of
the room falling into shadow. Static camera, medium close-up from slightly below
eye level. Tungsten warmth, soft contrast, quiet. Square 1:1 composition.
4. NATURE, WIDE
A single birch on a ridge line bends under a gust as low cloud tears past behind
it, grass laying flat in a moving wave across the slope. Overcast light, no sun,
desaturated greens and greys. Camera static on a long lens, subject small in
frame with sky occupying the upper two thirds. Wide 16:9 composition.
5. FOOD, MACRO
Butter melts into a hot cast-iron pan and slides, foaming at the edges, as heat
distortion rises off the metal. Hard key light from above and behind creates
specular highlights on the moving fat. Slow overhead descent, no rotation. Rich
contrast, clean colour, no colour cast. Vertical 9:16 composition.
6. ABSTRACT, TEXTURE
Thick indigo ink drops into clear water and unfurls in slow filaments, edges
feathering as the pigment loses density. Backlit against a plain white field so
the ink reads almost black at the core. Static macro camera, no move at all.
High key, clinical, no grain. Square 1:1 composition.
Send exclusions separately. The field exists on all five text-accepting video operations, so use it:
NEGATIVE PROMPT — paste into negative_prompt, not into prompt
General: text, watermark, logo, caption, subtitles, distorted hands,
extra fingers, extra limbs, warped faces, morphing geometry
Motion: camera shake, whip pans, sudden zoom, speed ramp, jump cut
Look: oversaturated colour, plastic skin, heavy vignette, HDR halo
How do you describe motion when the image already carries the subject?
By describing only the change. On Pika 2.5 image-to-video the prompt is optional, and the schema tells you what it is for in one sentence: "Text prompt guiding the motion. The source image carries the subject."
That is a precise instruction and most people ignore it. Re-describing the subject on an image-to-video call spends your prompt competing with a picture the model can already see. What earns its place is the verb, the camera, and the thing that should stay still.
7. The subject stays where she is; only her breath moves the scarf. Camera
holds static. Nothing else in frame changes.
8. Slow dolly in toward the doorway, ending on a medium shot. Subject does not
move. Dust in the light beam drifts downward.
9. The liquid begins to pour from the left edge of the glass and fills toward
the base. The glass itself does not move or rotate.
10. Wind enters from camera right and travels through the field left to right in
a single wave. The horizon line stays level. No camera move.
11. Handheld with slight weight, camera arcs a quarter turn to the left around
the subject at walking pace. Subject holds their pose and their eyeline
follows the lens.
Two things to keep off that list. Do not ask for a cut, because a five or ten second single generation has no edit point. And do not ask for the shot size you already have, because you set that when you chose the still. The full checklist for what belongs in a video prompt at all is in our anatomy of a video prompt.
What changes for Pikaframes, Pikaswaps and Pikadditions?
The prompt narrows each time, and the schema tells you how.
Pikaframes is described as "Animate between 2-5 ordered keyframes with Pika 2.5." The images field takes "2-5 ordered keyframe image URLs. Consecutive pairs become transitions." The prompt is optional and is "Text prompt guiding the motion. Applied to every transition." That last clause is the trap: one prompt covers all of them, so write something true of every hop, not of one. Timing is transition_duration_s, and the docs note "Seconds per transition. Over 5 seconds needs exactly 2 keyframes."
12. A continuous slow dolly forward through every transition, constant speed,
no cuts, no rotation. Lighting stays consistent between frames.
13. Each change happens as a soft dissolve rather than a physical movement.
Camera is locked off throughout.
14. Morph the subject smoothly between states, preserving silhouette and scale.
Background remains fixed and unlit by the change.
Pikaswaps is "Swap subjects or objects in a video while keeping the scene." Here the prompt is required, and there are three ways to point at the region: modify_region_roi, described as "Free-text description of the region to replace. Provide this or a mask", modify_region_mask, "URL of a mask image marking the region to replace. Provide this or an ROI", and an optional image, "URL of a reference image for the replacement subject." Name what changes, then name what must not.
15. Replace the plain white mug in the subject's right hand with a dark green
enamel camp mug with a chipped rim. Keep her hand, the table, the lighting
and every shadow exactly as they are.
modify_region_roi: the mug held in the right hand
16. Replace the view through the window with a night city skyline in rain,
lights blurred by the glass. The room, the curtains, the person and the
interior lighting stay unchanged.
modify_region_roi: the rectangle of the window pane
Pikadditions is "Insert any subject into an existing video." The prompt is required and the optional image is "URL of a reference image of the subject to add." Say what to add, where it sits, and how it behaves relative to the motion already in the shot.
17. Add a tabby cat asleep on the left arm of the sofa, rising and falling
slightly with its own breathing. It stays asleep and does not react to the
people talking. Everything already in the shot is unchanged.
18. Add three paper lanterns drifting upward in the top right of frame, small
and far away, moving slower than the foreground. They pass behind the tree
line rather than in front of it.
Does Pika publish a prompt guide?
No. This is the part of the reporting that surprised us, so it is worth stating plainly rather than papering over.
There is no prompting page on Pika's developer documentation. Its llms.txt index, which is a 36 KB routing file covering authentication, uploads, polling, idempotency, error codes and billing in real detail, contains no prompt-writing section and no link to one. The only sentence in Pika's consumer FAQ that resembles craft advice is the answer to a question about camera moves, which reads: "Include sexy keywords in your prompt, like bullet time, vertigo, timelapse or dolly down." That is the whole published guidance.
What Pika does publish instead is a worked example prompt inside every operation specification, and those are considerably more instructive than the FAQ. They are long, they are prose rather than tag lists, they name lighting and grade explicitly, and they end with exclusions. Every template on this page was calibrated against them.
There is a second consequence of that design. Pika describes its API as "agent-native", and the documentation is written to be read by a coding agent as much as by a person: a routing index, per-operation specs, a public catalog endpoint that returns the live JSON schema without a key. If you are building prompts programmatically, fetch the schema and generate against it rather than hardcoding a field list from an article, including this one.
How does Pika compare with the models we have already documented?
On raw length, badly. On what the prompt controls, differently rather than worse.
| Model | Text-to-video length the API accepts | Where that comes from |
|---|---|---|
| Pika 2.5 | Exactly 5s | Pika catalog schema, read August 27, 2026 |
| Veo 3.1 (Gemini API) | 4, 6 or 8s | Our duration reference, checked August 27, 2026 |
| Kling 3.0 | Any integer from 3 to 15s | Our duration reference, checked August 27, 2026 |
| Vidu Q3 Pro and Turbo | 1 to 16s | Our duration reference, checked August 27, 2026 |
The three non-Pika rows come from our video duration parameters by model, which was checked against each vendor's own API documentation. Pika sits at the bottom of that column and it is not close.
But length is the wrong axis for what Pika is for. The models above are shot generators, and a Seedance or Kling prompt is doing cinematography. Pika's distinguishing surfaces are editing operations on footage that already exists: swap a region, insert a subject, apply a named transformation, animate between stills you supply. Prompting those well means writing about change and preservation, not about composition. "Replace X, keep everything else" is a genuinely different sentence from "a wide shot of X at golden hour", and the second habit actively hurts you on the first job.
Pika also has an audio family, four operations, that the video specialists mostly do not. Pika Soundtrack takes a video and an optional instruction field, and the schema note on it is unusually candid: it "Reaches the model verbatim", and you can "Omit it, or leave it blank, for natural sound design derived from the video alone." Pika SFX is a plain text-to-audio operation described as "Generate any sound effect from a text prompt." Pika Music adds a rewrite_prompt boolean, documented as "Expand the prompt with a model before generating. Leave off when the prompt already specifies the arrangement; turn it on for a short or vague one", which is a rare case of a vendor telling you when not to use its own helper.
19. PIKA SFX
A heavy oak door in a stone corridor: the iron latch lifts with a dull clack,
the hinges give one long dry groan as the door swings, and the boards settle
with a low knock as it comes to rest. Close perspective, no reverb tail beyond
the corridor's own.
20. PIKA SOUNDTRACK — instruction field
Sound design only, no music and no dialogue. Follow the visible action exactly:
footfalls timed to each step, fabric movement when the arms swing, and a quiet
room tone underneath. Nothing dramatic, nothing that anticipates the cut.
If you want the cross-model version of that, our sound design prompts for AI video covers how Veo, Seedance, Kling and Vidu each handle it.
When is Pika the wrong choice?
Often enough that saying so is the useful part of this page.
When you need more than five seconds from text. Text-to-video is pinned at 5. Kling, Vidu and LTX all accept substantially longer single generations, and stitching five-second Pika clips is not a substitute for one continuous take.
When aspect ratio has to be exact. There is no parameter for it. A prompt clause is a request, not a setting, and if your deliverable is a hard 9:16 master you are better off with a model that takes the ratio as an input.
When you want an effect that is not on the list. The 62 image effects and 9 video effects are the complete set. There is no free-text route into that pipeline, and a prompt asking for a transformation that is not in the enum has nowhere to go. Generate it as a normal shot instead, or use a model built around motion control.
When you are on the free tier and need text-to-video. Pika's pricing page describes Basic as image-to-video only, at 480p, and marks commercial use as not included. Its FAQ also states that "If you share a video directly from Pika, videos have a watermark for users of all tiers."
When you are pricing a project off a blog post. Pika publishes per-operation pricing on each model page and a live catalog endpoint that returns it as structured data. Read those, not a summary, and read them on the day you quote.
None of that makes Pika a bad tool. It makes it a specific one. The honest positioning is that Pika is an editing and effects platform with a competent five-second generator attached, and prompts written for it should reflect that shape.
Keeping the ones that work
The templates on this page are deliberately parameterised, because the value of a Pika prompt is almost never in a single run. It is in the fourth client who needs the same product shot with a different bottle, or the twelfth social cut that has to match the eleven before it. That is a library problem, not a writing problem, and it is the same discipline we apply to text prompts.
Prompt Architects sits on the prompt side of this line and stays there. We do not generate video and we are not going to. What we do is turn a rough shot idea into a structured brief with the slots filled, store the version that actually rendered well, and let you swap the bracketed variables per client or per campaign with structured output you can paste straight into Pika or into a request body. The frames are Pika's job.
One last piece of hygiene, and it is the reason this page has a date stamp on every claim. Pika moved from Pikaffects and a Discord bot to an eleven-operation API and a 133-model aggregation catalog inside about eighteen months, and its own FAQ did not keep up. Anything you write against a version number should be re-checked against the catalog before you rely on it, including the numbers here.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An Account