TL;DR: Grok Imagine Video 1.5 is xAI's current video model, priced at $0.080 per second, with clips of 1 to 15 seconds and native 1080p on text-to-video and image-to-video. It has no seed and no negative prompt. Write one paragraph of plain prose, because an upsampler rewrites it before generation.
How do you prompt Grok Imagine Video 1.5?
Write a single paragraph that describes a still frame first, then the motion through it, then the sound. That order matches how the model works. Everything else in this post follows from that one sentence.
The reason the order matters is buried in xAI's July release notes rather than in any prompting guide: "Text-to-video on this model runs as text-to-image then image-to-video under the hood." The generation docs repeat it. The model builds a first frame from your prompt, then animates that frame, and you never see the intermediate image. If your prompt is all mood and no frame, the model invents the frame. If your prompt is all frame and no motion, you get a photograph with a slow drift on it.
There is a second reason, and it is the one most guides miss entirely. Your prompt is not sent to the renderer as written. xAI's video API documents a prompt-rewriting model in its usage accounting, so what you type is a brief for a rewriter, not an instruction to a renderer. More on that below.
One more thing worth saying up front, because it saves people money: xAI publishes a prompting guide for its voice model. It does not publish one for Imagine video. Every "official Grok video prompt structure" you find online is somebody's reverse-engineering, including this one. The difference is that everything here is tied to a documented parameter, a quoted sentence, or a sample xAI itself published.
What can Grok Imagine Video 1.5 actually do?
It generates video from text, from an image, or from reference images and voices. It does not, on the evidence of its own model card, take video as an input.
Here is what xAI documents, all checked at docs.x.ai on August 26, 2026.
| Setting | Parameter | Documented value |
|---|---|---|
| Clip length | duration | 1 to 15 seconds, default 8 |
| Aspect ratio | aspect_ratio | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, default 16:9 |
| Resolution | resolution | 480p, 720p, 1080p, default 480p |
| Native 1080p | n/a | Text-to-video and image-to-video only; reference-to-video capped at 720p |
| Audio | n/a | Generated by default |
| Preset voices | reference_audios | Up to 3 per request, by voice_id |
| Reference images | reference_images | URL, base64 data URI, or Files API file_id; JPEG, PNG or WebP |
| Seed | n/a | Not documented |
| Negative prompt | n/a | Not documented |
| Maximum prompt length | n/a | Not documented |
| Price | n/a | $0.080 per second of output |
| Rate limit | n/a | 10 requests per second |
| Regions | n/a | us-east-1, us-west-2 |
At that rate an 8-second clip costs $0.64 and the 15-second maximum costs $1.20. The older grok-imagine-video is listed at $0.050 per second and is not marked deprecated, unlike grok-voice-think-fast-1.0 which is. That older model's card lists its modalities as text, image, video to video, which is why it shows up in the editing and extension examples.
Why does text-to-video behave differently from image-to-video?
Because on text-to-video the model has to invent your first frame, and on image-to-video you hand it one.
That single difference should change how you write. For text-to-video, spend the first third of your prompt on things a photographer would care about: lens, framing, subject, wardrobe, set, light. You are commissioning a still. Only then describe what moves. For image-to-video, skip all of it. The frame already exists. Describe motion, camera, and sound, and nothing else, because re-describing the image just gives the rewriter room to argue with it.
There is a trap in image-to-video that xAI states plainly and almost nobody quotes. The output defaults to the input image's aspect ratio. If you pass aspect_ratio as well, the docs say it "will override this and stretch the image to the desired aspect ratio." Stretch is xAI's word, not mine. If you have a 16:9 still and you want a 9:16 vertical, crop the still before you send it. Do not ask the API to reframe it.
If you want the fuller theory of what a shot description should contain, our guide on directing AI video like a filmmaker covers lens, lighting and blocking in a model-agnostic way. What follows is the Grok-specific version.
What does a good Grok Imagine video prompt look like?
The best available specimen is xAI's own. Its reference-to-video documentation ships a sample prompt for a fashion runway shot, and it is worth reading as evidence of house style rather than as a demo:
"slow zoom in on the white fashion runway stage. then, the model from
<IMAGE_1>walks in from the back of the shot from the white opening, and gracefully walk out onto the front of the white stage platform. they wear the shirt from<IMAGE_2>and black flared jeans. they look dramatically at the camera. high quality slow motion shot. fun, playful. skin pores. highly detailed faces. perfect shot. they reach the end of the runway and look at the camera as the camera slowly zooms. subtle smile."
Five things are true of that prompt, and all five are copyable.
It is lowercase plain prose, not a schema. It opens with the camera move, so the frame is established before anything happens in it. It sequences beats with "then," and full stops rather than nesting clauses. It binds wardrobe and identity to reference tags instead of describing them. And it closes by restating the final beat, which reads like insurance against the clip drifting off in its last two seconds.
That last habit is the one I would steal first. A 10-second generation has time to lose the plot. Telling it how the shot ends, twice, costs you nine words.
Here is that structure as a prompt template you can paste and fill:
[camera move and framing]. [subject, wardrobe, set, light, described as a still].
then, [beat one]. [beat two]. [beat three].
[audio: dialogue, ambience, score].
[quality and format tags].
[restate the final beat and where the camera ends].
Filled in for a product shot:
slow push in on a matte black espresso machine on a concrete counter, morning
side light from a window on the left, shallow depth of field, 50mm.
then, a hand enters from the right and sets a white cup under the spout.
then, espresso pours in a thin steady stream and crema builds.
then, steam rises and catches the light.
ambient kitchen room tone, the low hum of the pump, no music.
high quality macro shot. highly detailed. shallow focus. perfect shot.
the pour finishes and the camera holds on the full cup, still pushing in slowly.
Filled in for a talking-head clip, which is where the audio track earns its keep:
medium close-up, static camera, subject centred, soft key from front left,
grey studio backdrop, 85mm, shallow depth of field.
the person from <IMAGE_1> looks directly at camera and speaks with the voice
from <AUDIO_0>, saying: "we shipped it on a tuesday and nobody noticed."
then, they pause and half-smile.
clean room tone, no music, no reverb.
highly detailed face. natural skin texture. accurate lip sync. perfect shot.
they hold the look at camera as the line lands.
Filled in for b-roll, where the whole job is texture:
handheld tracking shot moving right to left past a row of market stalls at dusk,
practical string lights overhead, wet cobblestones, 35mm, slight lens flare.
then, a vendor lifts a crate of oranges into frame.
then, the camera drifts past and the stall falls out of focus.
crowd murmur, distant traffic, no music.
cinematic. film grain. highly detailed. perfect shot.
the shot ends on the empty wet street ahead, still moving.
How do reference images and voices work?
Reference-to-video is the mode most likely to be underused, because it does something image-to-video cannot: it carries a person, an object, or a garment into a shot without locking the first frame to a picture.
You pass images in reference_images and refer to them inside the prompt with angle-bracket tags. You pass preset voices in reference_audios, each as an object like {"voice_id": "eve"}, drawn from the same voice roster as xAI's text-to-speech API, and refer to them as <AUDIO_0>. Identifiers are case-insensitive, and an unknown one returns a 400 with the list of valid voices, which is a genuinely well-designed error.
curl -X POST https://api.x.ai/v1/videos/generations \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-d '{
"model": "grok-imagine-video-1.5",
"prompt": "medium close-up, static camera, soft key light. the person from <IMAGE_1> speaks to camera with the voice from <AUDIO_0>. then, they glance down and back up. clean room tone, no music. highly detailed face. accurate lip sync. they hold the look at camera.",
"reference_images": [{"url": "<IMAGE_URL_1>"}],
"reference_audios": [{"voice_id": "eve"}],
"duration": 8,
"aspect_ratio": "9:16",
"resolution": "720p"
}'
Four constraints to hold in your head:
- Three voices maximum, tagged
<AUDIO_0>,<AUDIO_1>,<AUDIO_2>. Voice references built from your own audio files are limited to trusted partners on request, and caller-supplied clips must be 15 seconds or shorter. - Reference-to-video is capped at 720p. Native 1080p is text-to-video and image-to-video only. If you need both references and 1080p, you cannot have them in one call today.
imageandreference_imagesare mutually exclusive. Sending both returns a 400. Pick the mode.- The image tag index is inconsistent in xAI's own docs. The runway sample uses
<IMAGE_1>and<IMAGE_2>; the reference-audio section says to use<IMAGE_0>alongside<AUDIO_0>. xAI does not resolve this anywhere I could find. Test both against your own account before you build a pipeline on one of them.
What is the prompt upsampler, and how should it change your writing?
Every video response comes back with a usage object, and the field descriptions in it are the most useful undocumented documentation xAI has published. Input text_tokens are described as "Prompt text tokens consumed by the prompt-rewriting (upsampler) LLM." Output reasoning_tokens are "Reasoning (thinking) tokens generated by the prompt-rewriting (upsampler) LLM." Output text_tokens are "Rewritten-prompt text tokens."
So there is a reasoning model between you and the renderer, it thinks about your prompt, it rewrites it, and you are billed for all of that on top of the per-second output price.
Three practical consequences.
Short prompts are not cheap prompts. A six-word prompt gets expanded into something long by a model guessing at your intent. You pay for the expansion either way. You may as well write the expansion yourself and get the version you wanted.
Specify the things you care about and stay silent on the rest. Anything you leave out, the rewriter fills in. That is fine for the colour of a background wall and terrible for the brand of a product or the age of a person. Silence is a request for invention.
Do not expect determinism. There is no seed on this endpoint, and now there is a reasoning model in the path as well. Two identical requests can diverge twice over. If you need a repeatable look across a set of clips, pin the first frame with image-to-video and vary only the motion text. That is the closest thing to a seed the API offers.
This is also why a structured, reusable brief beats freestyling. If you already keep your video prompts as fielded templates, the pattern in our JSON video prompt templates piece transfers cleanly: keep the fields, flatten them to prose at send time, because Grok's endpoint takes one prompt string and nothing else.
How does it compare with Veo 3.1 and Gemini Omni Flash?
Only on what the vendors publish. Anything a vendor does not publish is marked as such rather than guessed at.
| Feature | Grok Imagine Video 1.5 | Veo 3.1 | Gemini Omni Flash |
|---|---|---|---|
| Model ID | grok-imagine-video-1.5 | veo-3.1-generate-preview | gemini-omni-flash-preview |
| Clip length | 1 to 15s, default 8 | 4, 6 or 8s | not published |
| Max resolution | 1080p | 4K, 8s only | not published |
| Aspect ratios | 7 options, default 16:9 | 16:9 and 9:16 | 16:9 and 9:16 |
| Native audio | |||
| Published price per second | $0.080 | $0.40 at 720p and 1080p | about $0.10 at 720p |
| Vendor status label | GA on the Imagine API | preview | preview |
Two notes on that table. Google's own model list labels Gemini Omni Flash as being in preview and gives Veo 3.1 a -preview model ID, while its video documentation now points developers to Omni Flash as the default choice for video. And the price gap is real but not like-for-like: Grok's $0.080 buys a second of output at whichever resolution you asked for, since xAI publishes no resolution tiers, whereas Veo's rate card prices 720p, 1080p and 4K separately.
If you want the model-versus-model view rather than the prompting view, our Veo 3 vs Sora vs Kling comparison covers the field. This post is deliberately about one model's parameters.
What does xAI not document?
This matters more than usual here, because Grok Imagine is newer than Veo or Kling and the gap between what xAI publishes and what circulates online is wide.
Consumer app limits. Grok Imagine also runs on grok.com/imagine and in the iOS and Android apps. Per-tier generation caps for those surfaces are not documented at docs.x.ai. Numbers do circulate on aggregator blogs. None of them trace back to xAI, so none of them are in this post.
Maximum prompt length. The API's error table lists "a prompt that is too long" as a cause of invalid_argument. No character or token limit is published.
Maximum reference images via the API. The REST reference states a limit of three for reference_audios and no limit for reference_images. xAI's July 31, 2026 consumer announcement mentions up to seven references per generation on grok.com/imagine and iOS. Whether that number applies to the API is not stated.
Whether 1.5 accepts video input. Covered above: the model card says text and image to video, and every editing and extension sample uses grok-imagine-video. xAI never says "1.5 rejects video."
How the upsampler rewrites. No published rules, no published system prompt, no documented way to turn it off.
Where does prompt tooling fit?
Honestly: not in the generation. Prompt Architects does not generate video. There is no version of this post where we are the answer to "which AI video model should I use," because we are not one, and if you want the model, xAI's API is right there.
What we do is the part before the API call. Our video prompt feature turns a rough idea into the structured shot description this model wants, with the frame, the beats, the audio and the closing note in the order described above, and the Prompt Library keeps the ones that worked so you are not rewriting a runway shot from memory next month. Global Variables let you swap the subject, the product or the voice across a set of prompts without editing each one. Given that there is no seed on this API, a saved prompt is the only version control you get.
The wider failure mode, which is not Grok-specific at all, is covered in why your AI videos look generic. Most of it comes down to under-specifying the frame and letting the model choose. On a model with an upsampler in the path, that tendency is amplified.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An Account