TL;DR: Grok Imagine prompts land best as one flowing paragraph that names the camera move first, sequences action with the word "then", and stacks quality descriptors at the end. Below are 29 copy-paste templates across product, UGC, cinematic, explainer, social and reference-driven video, with the API settings and per-clip cost for each.
What are the best Grok Imagine prompts?
The best ones look almost nothing like the bulleted shot lists people write for other video models. They read as a single unbroken paragraph in lowercase, they open by stating the camera, they use "then" as a beat marker between actions, and they close with a short run of quality words.
That shape is not a preference. It is copied from the only substantial prompt xAI has published for this model.
Here is the interesting thing about grok-imagine-video-1.5: xAI does not publish a prompting guide for it. They publish one for speech-to-speech, at /developers/model-capabilities/audio/speech-to-speech/prompting-guide. Nothing equivalent exists for Imagine video or Imagine images. Every prompting-guide URL you would guess returns a 404, and the documentation index at docs.x.ai/llms.txt confirms no such page is listed. Checked August 26, 2026.
So the entire official corpus of Grok Imagine video prompt craft is one runway-fashion example buried in the reference-to-video page, plus a handful of one-line demo prompts. That example is worth reading closely, because it is the closest thing to a house style anyone has:
slow zoom in on the white fashion runway stage. then, the model from <IMAGE_1> walks in from the back of the shot from the white opening, and gracefully walk out onto the front of the white stage platform. they wear the shirt from <IMAGE_2> and black flared jeans. they look dramatically at the camera. high quality slow motion shot. fun, playful. skin pores. highly detailed faces. perfect shot. they reach the end of the runway and look at the camera as the camera slowly zooms. subtle smile.
Four things stand out. It is all lowercase. It states the camera move in the first four words. It marks the second beat with "then". And it restates the camera move at the very end, after the quality tokens, which is unusual enough to be deliberate. Note also that it passes three reference images but only ever names two of them in the prompt text.
Every template below is built to that shape. If you want the mechanics underneath it, the modes, the limits, the prompt upsampler and the parameters that do not exist, post 156 covers how to prompt Grok Imagine Video 1.5 in full. This page is the copy-paste pack, so I will not repeat it.
What is the skeleton every template follows?
Six slots, in this order, written as continuous prose rather than labelled fields.
[camera move] on [subject] in [location].
[first action]. then, [second action].
[lighting direction and quality]. [lens or depth-of-field note].
[mood words]. [material or texture detail]. photoreal.
[restate the camera move resolving].
Two slots do most of the work. Naming the camera move first gives the model a stable frame before it has to decide what moves inside that frame. Restating how the move resolves at the end tends to stop the shot drifting into a second, unrequested camera move halfway through.
Keep it to one paragraph. Do not use JSON here. Grok Imagine's video endpoint takes a plain string, and unlike Veo there is no documented schema to fill, so a JSON prompt buys you nothing on this model. If you work in JSON for other video models, the Veo 3 JSON template pack is the right page for that; it does not transfer here.
Product and commerce prompts
Five shots that cover most of a product launch: the hero, the unboxing, the pack shot, the texture macro, and the demonstration. Run these as text-to-video at 1080p, or as image-to-video if you already have a rendered still and want the product's exact geometry preserved.
1. Hero product orbit
slow orbit around a matte black espresso machine on a polished concrete plinth. the camera circles left to right at a constant speed. then, steam begins to rise gently from the cup below the spout. soft key light from the upper left, single rim light picking out the right edge. shallow depth of field, background falling to soft grey. crisp brushed metal texture. photoreal. commercial product film. the orbit slows and settles with the machine centred in frame.
2. Unboxing, hands only
locked-off top-down shot of two hands lifting the lid from a cream gift box on a walnut table. the lid comes away. then, the hands fold back the tissue paper to reveal a folded navy scarf. natural window light from the left, soft directional shadows. no faces, hands only. warm, tactile, unhurried. visible paper grain and fabric weave. photoreal. the camera holds completely still throughout.
3. Pack shot with moving light
static wide shot of three skincare bottles arranged on a curved white riser against a pale sand backdrop. a shaft of hard sunlight sweeps slowly across the set from the right. then, a palm frond shadow tracks over the bottles and off the left edge. the bottles do not move. glossy glass, condensation beading on the tallest bottle. clean, editorial, high-end beauty advertising. photoreal. camera locked off for the entire shot.
4. Texture macro
extreme macro push in on the surface of a dark chocolate bar as a knife tip presses into the corner. the crack spreads outward. then, a shard breaks away and settles flat. tight top light with deep falloff to black. shallow focus holding the crack line sharp. slow, deliberate. visible cocoa bloom and sugar crystals. photoreal food cinematography. the push in stops just as the shard lands.
5. Before and after in one take
single locked-off shot of a scuffed leather boot on a workbench. a cloth enters frame from the right and wipes across the toe in one pass. then, the cloth leaves frame and the wiped section sits clean and glossy against the dull leather beside it. warm workshop light from a window behind. the contrast between the two halves of the boot is clearly visible. photoreal. the camera does not move at any point.
UGC and talking-head prompts
These use reference-to-video, which is where Grok Imagine gets genuinely interesting. You pass a face through reference_images and a voice through reference_audios, then tag both inside the prompt text. Reference-to-video is capped at 720p, which is fine for social and wrong for a broadcast spot.
Voices come from xAI's built-in preset roster, the same identifiers as their text-to-speech API, passed as {"voice_id": "eve"} and matched case-insensitively. You can pass at most three. Supplying your own voice clips is restricted to trusted partners on request, so for everyone else these are preset voices only.
6. To-camera testimonial
the person from <IMAGE_1> sits on a beige sofa in a bright living room and speaks directly to camera with the voice from <AUDIO_0>. they gesture once with an open hand. then, they settle back and keep talking. handheld framing with a very slight sway. soft daylight from a window on the left. natural skin texture, visible pores. casual, warm, unpolished. shot on a phone front camera. photoreal. the handheld drift continues gently to the end.
7. Problem then solution
the person from <IMAGE_1> stands in a small kitchen holding a phone and speaks to camera with the voice from <AUDIO_0>. they start with a frustrated shrug. then, they hold the phone up towards the lens and their expression opens into a smile. handheld, slight reframe as the phone rises. overhead kitchen light, warm. natural, conversational, not acted. photoreal. the camera settles as they lower the phone.
8. Hands-only demo with voiceover
close top-down shot of two hands assembling a small wooden desk organiser on a linen mat, narrated by the voice from <AUDIO_0>. the hands slot the divider into place. then, they press the base flat until it clicks. no face in frame at any point. soft overhead light with no hard shadows. calm, instructional, tactile. photoreal. the camera drifts a few centimetres closer as the piece comes together.
9. Street vox pop
handheld medium shot of the person from <IMAGE_1> standing on a busy pavement, answering an off-camera interviewer with the voice from <AUDIO_0>. they glance away as they think. then, they look back and continue. pedestrians pass behind them out of focus. overcast daylight, flat and even. documentary feel, unscripted energy. photoreal. the frame bobs naturally and reframes once towards the end.
10. Founder desk note
medium close-up of the person from <IMAGE_1> at a cluttered desk, speaking to camera with the voice from <AUDIO_0>. they lean in slightly as they begin. then, they settle back into the chair. a monitor glows off-frame to the right and lights one side of their face. late evening, lamp light, warm and low. sincere, low-energy, direct. photoreal. the camera holds still throughout.
Cinematic and narrative prompts
Six shots you can cut together into a sequence. All text-to-video, all 1080p, all better at 10 seconds than 8. If you want a deeper vocabulary for the camera itself, 30 cinematic camera prompts has the full move list, and the terms in it drop straight into the first slot of the skeleton above.
11. Establishing aerial
aerial shot flying forward over a ridge at dawn. the fog-filled valley below opens up as the camera crests the ridge. then, the camera rises slightly and levels off, revealing a single road cutting through the fog. low golden sun behind the ridge, long shadows, haze in the air. wide anamorphic feel. epic, still, unhurried. photoreal cinematography. the forward drift slows as the valley fills the frame.
12. Slow push-in
slow dolly push in on a woman standing at a rain-streaked window with her back to camera. rain runs down the glass in front of her. then, she turns her head towards the window as the camera closes. cool blue light from outside, one warm practical lamp behind her. shallow focus, background bokeh from street lights. quiet, tense, restrained. photoreal, subtle anamorphic flare on the lamp. the push in stops just short of her shoulder.
13. Handheld chase
handheld camera running behind a figure sprinting down a narrow alley at night. the frame shakes and swings with the pace. then, it whips left as the figure turns a corner and disappears from view. wet cobbles reflecting sodium street lamps. motion blur across the edges of frame. urgent, kinetic, breathless. photoreal, high shutter. the camera keeps running as the corner clears.
14. Golden-hour portrait
medium shot of a man leaning against a fence in a wheat field at golden hour. wind moves the wheat and his shirt as he looks off to the left. then, he slowly turns his head towards the camera. backlit by low sun, rim light through his hair, warm haze and floating dust in the air. still, warm, contemplative. photoreal. the camera drifts slowly to the right and holds as he meets the lens.
15. Night neon tracking shot
tracking shot moving alongside a woman walking down a wet neon-lit street at night. she steps through a puddle carrying the reflection of a sign overhead. then, she passes under a second sign and the colour on her face shifts. magenta and cyan practical light, deep shadows between the signs. she does not look at the camera. moody, saturated, cinematic. photoreal. the camera keeps pace at a constant distance to the end.
16. Match on action through a doorway
a hand reaches for a door handle in a dim hallway and pulls it open. then, the camera follows through the doorway into a bright sunlit courtyard as the figure steps out. hard contrast between the dark interior and the blown-out exterior, the exposure settling as the shot moves outside. one continuous move, no cut. photoreal, single take. the camera comes to rest in the courtyard behind the figure.
Explainer and B2B prompts
The hardest category for any video model, because the request is usually abstract. The fix is to give it a physical metaphor with real materials rather than an idea. Five second clips are usually enough here, which also makes this the cheapest group to iterate on.
17. Concept metaphor
a single tangled ball of red thread sits on a white surface. one end lifts and pulls away from the tangle. then, the whole knot unravels into a straight taut line running out of frame. even soft studio light, no visible shadows. clean, minimal, conceptual. photoreal macro. the camera pushes in slightly as the line straightens and holds.
18. Diagram build
top-down shot of a hand drawing three connected boxes on white paper with a black marker. the hand draws the first box. then, it draws a connecting arrow and the second box beside it. the paper fills most of the frame. even overhead light, camera static. clean, instructional, no clutter. photoreal, visible ink bleed into the paper fibres. the hand withdraws and the frame holds on the finished drawing.
19. Data in motion
macro tracking shot along a row of frosted acrylic bar-chart blocks on a dark reflective surface. the camera moves left to right past the blocks. then, each block lights from within in sequence as the camera reaches it. cool white internal glow, dark reflections beneath. shallow depth of field. precise, modern, corporate. photoreal. the track slows as the tallest block lights.
20. Office ambient B-roll
slow lateral dolly past a glass meeting room where three people are talking around a table. the glass carries a soft reflection of the office behind the camera. then, one person stands and moves to the whiteboard. large windows, flat daylight, muted palette. observational, corporate, unstaged. photoreal. the dolly continues past at a steady walking pace and does not stop.
21. Process line
tracking shot moving alongside a conveyor belt carrying identical sealed cartons through a bright factory. the camera matches the belt speed so the cartons sit still in frame. then, the machinery behind them slides past and a robotic arm swings into view. cool industrial lighting, clean stainless surfaces. steady, mechanical, precise. photoreal. the camera holds its match to the belt right to the end.
Social and short-form prompts
All vertical, all short, all built to survive a caption overlay. Set aspect_ratio to 9:16 explicitly, because the documented default is 16:9 and it will not infer vertical from the content of your prompt.
22. Vertical food hook
vertical shot of a bowl of noodles on a dark counter with steam rising. a pair of chopsticks enters from the top of frame. then, they lift a single strand high and hold it there. warm overhead light, deep shadows around the bowl. the top third of the frame stays clear of the bowl. appetising, high contrast. photoreal. the camera pushes in slightly as the strand rises and stops.
23. Satisfying loop
vertical macro shot looking directly down at thick pink paint pouring onto a slowly rotating white disc. the paint spreads outward across the surface. then, the disc completes its turn and the surface sits perfectly flat and even. soft even light, no hard shadow. hypnotic, clean, minimal. photoreal. the first and last frames of the shot look identical.
24. Trend-format lifestyle walk
vertical handheld shot following a person from behind as they walk into a sunlit bakery. the door swings shut behind them. then, they stop in front of the counter and look up at the menu board. camera at chest height with natural sway. warm morning light through the shopfront. casual, aspirational, phone-shot feel. photoreal. the camera settles behind their shoulder and holds.
25. Text-safe empty frame
vertical static shot of a plain sage green wall with a single potted olive tree in the lower left corner. soft daylight sits across the wall. then, a cloud passes and the light dims slowly across the whole surface. nothing enters or leaves frame. the upper two thirds of the frame stay completely empty. calm, minimal, muted. photoreal. the camera does not move at any point.
Reference-driven prompts
Reference-to-video is the mode that separates Grok Imagine from a plain text-to-video model. xAI describes it as incorporating specific people, objects or clothing without locking the first frame, which is what makes it different from image-to-video. The subject appears in the clip, but the model still composes the shot.
Four use cases xAI names directly: virtual try-on, product placement, character-consistent storytelling, and voice identity.
26. Virtual try-on
the model from <IMAGE_1> walks slowly towards camera down a plain white studio corridor wearing the jacket from <IMAGE_2> over dark jeans. they reach the middle of the corridor and stop. then, they turn a quarter to the left and turn back to camera. even soft studio light from above. the jacket's cut, colour and hardware stay consistent throughout. fabric moves naturally as they walk. high quality fashion film. photoreal, visible weave. the camera holds wide as they settle.
27. Product placement in a scene
a woman sits at a cafe table and reaches for the bottle from <IMAGE_1>. she lifts it towards her. then, she sets it back down in front of her with the label facing camera. handheld medium shot with a slight reframe as she reaches. warm afternoon light from a window on the right, soft shadows across the table. the bottle's shape, label and colour stay consistent throughout. natural, unstaged. photoreal. the frame settles once the bottle lands.
28. Character-consistent series beat
the character from <IMAGE_1> stands at the edge of a rooftop at dusk looking out over the city. they hold the position for a moment. then, they turn and walk back towards the stairwell door. their face, hair and clothing match the reference exactly. low warm sun behind the skyline, long shadows across the roof surface. quiet, cinematic, restrained. photoreal. the camera holds wide and does not move at any point.
29. Voice-matched presenter
the person from <IMAGE_1> stands in front of a plain grey backdrop and delivers a short address to camera with the voice from <AUDIO_0>. they hold eye contact with the lens as they begin. then, they gesture once and return their hands to their sides. even soft frontal light with no visible shadow on the backdrop. corporate, calm, authoritative. natural blinking and micro-expressions. photoreal. the camera stays locked off throughout.
Which settings should you pair with each template?
Match the group to the row. All prices are the flat published rate multiplied by duration.
| Template group | Mode | duration | aspect_ratio | resolution | Cost per clip |
|---|---|---|---|---|---|
| Product hero (1, 3, 4) | text-to-video | 8 | 16:9 | 1080p | $0.64 |
| Product in-hand (2, 5) | image-to-video | 6 | 1:1 | 1080p | $0.48 |
| UGC talking head (6–10) | reference-to-video | 8 | 9:16 | 720p | $0.64 |
| Cinematic (11–16) | text-to-video | 10 | 16:9 | 1080p | $0.80 |
| Explainer (17–21) | text-to-video | 5 | 16:9 | 720p | $0.40 |
| Social (22–25) | text-to-video | 4 | 9:16 | 720p | $0.32 |
| Reference-driven (26–29) | reference-to-video | 10 | 9:16 or 16:9 | 720p | $0.80 |
The documented ranges behind that table, all read from xAI's REST reference and video generation guide on August 26, 2026:
| Field | Documented values | Default |
|---|---|---|
duration | integer, 1 to 15 seconds | 8 |
aspect_ratio | 1:1 16:9 9:16 4:3 3:4 3:2 2:3 | 16:9 |
resolution | 480p 720p 1080p | 480p |
reference_audios | up to 3 entries, each a preset voice_id | none |
| Audio track | included by default | on |
Two of the five request modes ignore that table entirely, and both catch people out because the parameters are accepted rather than rejected.
| Mode | Endpoint | What it does with your settings |
|---|---|---|
| Video editing | /v1/videos/edits | Ignores duration, aspect_ratio and resolution. The output inherits all three from the source video, and the source is capped at 8.7 seconds. A 1080p input is downsized to 720p on the way out. |
| Video extension | /v1/videos/extensions | duration here means the length of the new segment only, not the finished file. It accepts 2 to 10 seconds and defaults to 6, so a 10-second source extended by 5 returns a 15-second video. |
That second row is the one worth remembering, because it is also a workaround. The generation cap is 15 seconds, but extension adds its segment on top of the original length, so a longer piece is a chain of generations rather than one long request.
Because the price is flat per second, running all 29 templates once at 5 seconds each costs $11.60. Resolution buys you render time, not billing. That changes how you should iterate: draft at short durations rather than at low resolution, since duration is the only lever that actually moves the invoice.
What do xAI's own docs disagree about?
Three things, as of August 26, 2026. None of them are catastrophic, and all three will silently waste your time if you assume the docs are internally consistent.
Reference tag indexing. One line of prose on the reference-to-video page says to tag voices as <AUDIO_0>, <AUDIO_1>, <AUDIO_2>, "with <IMAGE_0>… when you also pass images". Every code example on that same page and its sibling does something different. Across the six worked prompts in the documentation, the first reference image is always <IMAGE_1> and the second is always <IMAGE_2>. <IMAGE_0> appears exactly once in the entire documentation set, in that one line of prose. Audio, by contrast, is consistently zero-indexed everywhere. Follow the examples.
Batch API support. The model page for grok-imagine-video-1.5 lists "Batch API: Supported". The Batch API documentation says the opposite in plain language: image and video batch requests support grok-imagine-image and grok-imagine-video, and other Imagine models "including grok-imagine-image-2.0 and grok-imagine-video-1.5" are rejected with "not supported for batch processing". If you were planning to batch a 29-template sweep, plan on concurrent individual requests instead.
Whether resolution affects price. The Imagine overview page states that video generation "uses per-second pricing where both duration and resolution affect the total cost". No resolution-tiered rate card exists anywhere in xAI's documentation. The pricing page lists one figure for the model, $0.080 per second, and the model page repeats the same single figure under Output. I suspect this is where the tiered pricing tables floating around third-party blogs came from. Until xAI publishes tiers, the only defensible thing to quote is the flat rate, dated.
What should you do when a template does not land?
There is no seed, so you cannot re-roll deterministically, and there is no negative prompt, so you cannot subtract. That leaves four moves, in order of how often they work.
Cut the prompt down. A long prompt does not give the model more control, it gives it more chances to average competing instructions together. If a shot comes back muddled, delete the third and fourth sentences before you add anything.
Move the camera to the front. If the shot drifts or invents a move you did not ask for, check that the camera instruction is the first thing in the prompt and restated at the end. That bracketing is the one structural trick visible in xAI's official example.
Switch modes rather than rewrite. If text-to-video keeps missing the subject, generate a still you like and drive it through image-to-video, which locks the first frame. If you need the subject present but the composition free, use reference-to-video instead.
Only then add descriptors. Most generic output is under-specified in materials and light rather than over-specified in adjectives. Why AI videos look generic breaks down which details actually change the render and which are decoration.
There is a fifth move that people forget exists. If a clip is 80 percent right and one element is wrong, you do not have to regenerate it. Send it to the edit endpoint with a prompt naming only the change you want, in the style of xAI's own examples: "give the woman a silver necklace", "change the colour of the outfit to red". Editing preserves the rest of the scene, which is a very different failure profile from rolling the dice on a fresh generation and losing a shot you already liked. The cost is that the edit output comes back at 720p or lower, so keep the original file.
Where does prompt tooling fit?
Templates like these are only worth having if you can find them again in six weeks. Twenty-nine prompts across six use cases, each with its own settings, is exactly the kind of thing that starts in a note file and ends nowhere.
That is the problem Prompt Architects exists for. Save each template once, turn the variable parts into reusable placeholders, and pull the right one into whatever tool you are working in. Video prompts are a first-class type in the library alongside text and image prompts.
Two things to be straight about. We do not generate video. We generate the prompt, and you take it to xAI's API or to Grok Imagine yourself. And we do not have an API, so if you want this pipeline fully automated, you will be scripting the xAI call directly and using us for the authoring and storage layer only.
There is a free plan, capped at 5 prompt enhancements per day, forever. Paid plans start at $4.99/month at the time of writing, and current pricing is on the pricing page. If you would rather start from a broader base, 100+ prompt templates covers the non-video categories.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An Account