TL;DR: Hailuo 2.3 prompts work best as one subject, one action beat, one bracketed camera command and a light description, all inside MiniMax's documented 2,000-character prompt limit. The model produces 6 or 10 seconds at 768P, or 6 seconds at 1080P, at 24fps. Audio output is not documented in MiniMax's video API reference.
What are the best MiniMax Hailuo 2.3 prompts?
The best Hailuo prompts name one subject, give it one physical action, and attach one camera command in square brackets. Everything else is set dressing.
That sounds reductive until you look at what MiniMax says the model is for. The Hailuo 2.3 launch post claims improvements in "the portrayal of physical actions, stylization, and character micro-expressions, while further optimizing its response to motion commands." Body motion and camera obedience are the two things this model was tuned on. A prompt that spends 400 words on colour grading and 6 words on what the person is doing is asking the model to compete on the axis it was not built for.
So the template in this post is deliberately blunt. Subject, action, camera, environment, light, hold. Six slots. Fill them and stop.
Everything below is checked against MiniMax's own developer documentation rather than a reseller's spec sheet, because the aggregator pages disagree with each other and several of them describe Hailuo 02 while calling it 2.3.
What does Hailuo 2.3 actually support?
Six or ten seconds at 768P, six seconds at 1080P, 24 frames per second, and a prompt field capped at 2,000 characters. There is no 512P option and no ten-second 1080P option.
Here are the numbers as MiniMax publishes them, read on 26 August 2026:
| Spec | MiniMax-Hailuo-2.3 | Source |
|---|---|---|
| Endpoint | POST https://api.minimax.io/v1/video_generation | API reference |
| Duration | 6 or 10 at 768P; 6 at 1080P; default 6 | Text-to-Video reference |
| Resolution | 768P (default) or 1080P at 6s; 768P at 10s | Text-to-Video reference |
| Frame rate | 24 fps | Models overview |
| Prompt limit | 2,000 characters | Text-to-Video reference |
| Camera commands | 15 bracketed commands | Text-to-Video reference |
| Audio | Not documented | Not published |
| Aspect ratio parameter | Not documented on the v1 endpoint | Not published |
| Negative prompt parameter | Not documented on the v1 endpoint | Not published |
Two of those rows matter more than they look. There is no documented negative_prompt field on this endpoint, so every exclusion you want has to be phrased positively inside the prompt itself. And on image-to-video, MiniMax's own note says video resolution follows the first frame image, so your input still governs the frame.
What is the six-part Hailuo prompt template?
Use this shape for every text-to-video shot. It fits comfortably inside 2,000 characters and it puts the two things Hailuo 2.3 is good at, body motion and camera obedience, in the positions the model reads first.
[SUBJECT] — one person or object, with two identity anchors (age/build/wardrobe).
[ACTION] — one physical beat, verb first, that can complete in the clip length.
[CAMERA] — one bracketed command from MiniMax's list of 15.
[SPACE] — where this happens, in under ten words.
[LIGHT] — direction and quality of light, in under ten words.
[HOLD] — what stays still or unchanged for the whole shot.
Filled in, it looks like this:
A woman in her thirties, cropped denim jacket, short dark hair, pushes a heavy
steel door open with her shoulder and steps through [Tracking shot]. Narrow
service corridor, painted concrete. Hard overhead strip light from directly
above, cool white. Her jacket, hair length and the corridor walls stay
unchanged for the whole shot.
Note what is missing. No lens, no film stock, no artist name, no adjective pile. Those help image models more than they help this one. Note also that the action is a single beat that finishes inside six seconds. "Pushes a door and steps through" completes. "Walks down the corridor, opens a door, greets a colleague, sits down" does not, and Hailuo will compress it into mush or drop half of it.
If you want a second beat, use the sequential form MiniMax documents rather than a comma splice:
A man picks up a book [Pedestal up], then reads it at the window [Static shot].
That is MiniMax's own example prompt from the text-to-video reference, and it is the pattern to copy: beat, camera, "then", beat, camera.
Which camera commands does Hailuo 2.3 support?
Fifteen, in nine families. MiniMax lists them explicitly in the prompt field description for the text-to-video and image-to-video endpoints.
| Family | Commands |
|---|---|
| Truck | [Truck left] · [Truck right] |
| Pan | [Pan left] · [Pan right] |
| Push | [Push in] · [Pull out] |
| Pedestal | [Pedestal up] · [Pedestal down] |
| Tilt | [Tilt up] · [Tilt down] |
| Zoom | [Zoom in] · [Zoom out] |
| Shake | [Shake] |
| Follow | [Tracking shot] |
| Static | [Static shot] |
Three usage rules come straight from the docs. Commands inside a single bracket take effect simultaneously, written comma-separated, with a recommended maximum of three. Commands placed in sequence through the prompt execute in that order. And free-form natural language works, but MiniMax says explicit commands yield more accurate results.
That last line is the useful one. If you have been writing "the camera slowly drifts left while rising above the subject" and getting inconsistent results, the documented equivalent is one bracket:
[Truck left,Pedestal up]
Here are combinations worth keeping. Each is two or three commands in one bracket, which is inside MiniMax's recommended ceiling:
Reveal a room: [Truck right,Pan left]
Rise into a wide: [Pedestal up,Pull out]
Press into a face: [Push in,Tilt down]
Handheld chase feel: [Tracking shot,Shake]
Lock off for dialogue: [Static shot]
Drop to a low angle: [Pedestal down,Tilt up]
One caution the docs do not spell out: [Zoom in] and [Push in] are not the same move. Zoom changes focal length, push moves the camera body. If your subject's background compresses when you wanted it to stay put, you asked for the wrong one. Our cinematic camera prompt collection covers the same distinction for other models.
What are good Hailuo prompts for character motion?
Character motion is where Hailuo 2.3 was pitched, so it is worth spending your prompt budget there. These four templates cover the beats that break most often: weight shifts, hand-offs, turns, and sitting down.
WEIGHT SHIFT
A [AGE] [PERSON] in [WARDROBE] lifts a [HEAVY OBJECT] from the floor to
waist height, bracing through the legs [Static shot]. [SPACE]. [LIGHT].
The object's size and the person's grip stay consistent throughout.
TURN AND FACE
A [AGE] [PERSON] in [WARDROBE] turns from three-quarter profile to face
the camera, chin leading, shoulders following [Push in]. [SPACE].
[LIGHT]. Wardrobe, hair and background remain unchanged.
HAND-OFF
A [PERSON A] passes a [SMALL OBJECT] to [PERSON B]; the object leaves one
hand and is taken by the other in a single continuous motion [Static shot].
[SPACE]. [LIGHT]. Both faces stay visible for the whole exchange.
SIT DOWN
A [AGE] [PERSON] in [WARDROBE] lowers into a [CHAIR TYPE], one hand on the
armrest, weight settling [Pedestal down]. [SPACE]. [LIGHT]. The chair
position and the room behind stay fixed.
Two things make these work. The HOLD clause at the end is doing quiet work against drift: naming what should not change gives the model something to anchor on across frames, which is the same trick that keeps subjects stable in image prompting. And every action is described mechanically rather than emotionally. "Braces through the legs" and "chin leading, shoulders following" describe physics. "Looks determined" describes a feeling, and the model has to guess at the body language that produces it.
How do you write Hailuo image-to-video prompts?
Add a first frame image and keep the prompt to what changes. The image already carries subject, wardrobe, framing and light, so repeating them wastes characters and invites the model to redraw what you already fixed.
MiniMax's image requirements for first_frame_image: JPG, JPEG, PNG or WebP, under 20MB, short edge greater than 300px, aspect ratio between 2:5 and 5:2. Public URLs and base64 data URLs both work.
{
"model": "MiniMax-Hailuo-2.3",
"first_frame_image": "https://example.com/frame.jpg",
"prompt": "She turns her head to look off-frame left, then back to camera [Static shot]. Everything else in the frame is unchanged.",
"prompt_optimizer": false,
"duration": 6,
"resolution": "1080P"
}
A working request end to end, using the two-step async pattern MiniMax documents:
import os, time, requests
KEY = os.environ["MINIMAX_API_KEY"]
H = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}
task = requests.post(
"https://api.minimax.io/v1/video_generation",
headers=H,
json={
"model": "MiniMax-Hailuo-2.3",
"prompt": "A man picks up a book [Pedestal up], then reads [Static shot].",
"prompt_optimizer": False,
"duration": 6,
"resolution": "1080P",
},
).json()
task_id = task["task_id"]
while True:
time.sleep(10)
r = requests.get(
"https://api.minimax.io/v1/query/video_generation",
headers=H, params={"task_id": task_id},
).json()
if r["status"] == "Success":
print(r["file_id"], r["video_width"], r["video_height"])
break
One version-specific trap. Hailuo 2.3-Fast is listed in the model list for the image-to-video endpoint but not in the text-to-video endpoint's list, which means the Fast variant needs a first frame image. If you want the cheaper variant, you need to supply a starting frame, and MiniMax notes that video resolution follows that image.
Generation is asynchronous in both directions. The create call returns a task_id and nothing else. You then either poll the query endpoint until status comes back as Success, or register a callback_url and let MiniMax push updates to you. The callback path has one gotcha worth knowing before you wire it: MiniMax first sends a POST containing a challenge field, and your server has to echo that value back within three seconds or the URL never gets validated. Callback statuses are processing, success and failed.
Should you render at 768P or 1080P?
Render at 768P while you are still iterating on the prompt, then switch to 1080P for the take you want to keep. The reason is arithmetic rather than taste.
At MiniMax's published list prices, a 768P six-second Hailuo 2.3 clip costs $0.28 and the 1080P version of the same six seconds costs $0.49. That is roughly a 75% premium for a resolution step you cannot judge until the motion is already right. Twenty exploratory renders at 768P run about $5.60. The same twenty at 1080P run about $9.80, and nineteen of them were going in the bin regardless.
The Fast variant makes the exploration phase cheaper still, at $0.19 for 768P six seconds, though it needs a first frame image, so it only helps on image-to-video work.
There is a second reason to stay at 768P early. Ten-second renders are 768P-only, so if you are still deciding whether an action needs six seconds or ten, you have to be at 768P to compare them at all. Lock the timing first, then take the resolution step once.
One thing worth flagging honestly: MiniMax does not publish a bitrate or codec for either resolution, and the query response returns only file_id, video_width and video_height. If your delivery spec has a bitrate floor, you will have to measure the output rather than read it off a page.
Should you leave prompt_optimizer on?
Turn it off as soon as your prompt is good. MiniMax states that prompt_optimizer defaults to true and that setting it to false gives "more precise control", which is a polite way of saying the default rewrites your text before it reaches the model.
That is genuinely helpful when you type one line and want it fleshed out. It is actively harmful once you have tuned camera commands and a HOLD clause, because you are no longer sending the prompt you wrote. If your renders drift between runs of an identical prompt, this flag is the first thing to check.
There is a third setting most people miss. fast_pretreatment defaults to false and reduces optimisation time when the optimizer is on. It applies only to Hailuo 2.3, Hailuo 2.3-Fast and Hailuo 02. It is a latency knob, not a quality knob.
The practical loop: draft with prompt_optimizer on, look at what comes back, then rewrite the prompt yourself, set the flag to false, and iterate from there. That is the same discipline behind treating prompts as versioned artefacts rather than one-off typing, which we cover in prompt versioning for developers.
Hailuo 2.3 vs 2.3-Fast vs MiniMax H3: what changed?
| Feature | MiniMax-Hailuo-2.3 | MiniMax-Hailuo-2.3-Fast | MiniMax-H3 |
|---|---|---|---|
| Text-to-video | Yes | Not in the model list | Yes |
| Image-to-video | Yes | Yes (first frame required) | Yes |
| First + last frame | Not supported | Not supported | Yes |
| Longest clip | 10s at 768P | 10s at 768P | 15s |
| Top resolution | 1080P (6s only) | 1080P (6s only) | 2K |
| Frame rate | 24 fps | 24 fps | 24 fps |
| Prompt limit | 2,000 characters | 2,000 characters | 7,000 characters |
| Native audio | Not documented | Not documented | Native stereo sound |
| Endpoint | /v1/video_generation | /v1/video_generation | /v2/video_generation |
| Listed by MiniMax as | Legacy | Legacy | Current |
| 768P 6s list price | $0.28 | $0.19 | $0.08 per second |
The honest read: if your work is six-second character beats and you already have prompts that land, Hailuo 2.3 is cheap, predictable and still supported. If you need audio in the same generation, longer clips, 2K, or first-and-last-frame control, MiniMax has moved those to H3 and a different endpoint, and your Hailuo prompts will not port across unchanged. H3 takes multimodal content arrays rather than a flat prompt string.
How does Hailuo 2.3 compare to Veo 3.1 and Kling 3.0?
Only on what each vendor actually publishes. Everything unverified below says so.
| Hailuo 2.3 | Veo 3.1 (preview) | Kling 3.0 | |
|---|---|---|---|
| Durations | 6s or 10s at 768P; 6s at 1080P | 4, 6 or 8 seconds | Not published |
| Top resolution | 1080P | 4K (8s only) | 4K, per Kling's pricing tiers |
| Frame rate | 24 fps | 24 fps | Not published |
| Native audio | Not documented | Always on | Not published |
| Prompt limit | 2,000 characters | 1,024 tokens | Not published |
| Status | Listed as legacy | Preview | Not published |
Three notes on that table. Note first that the units differ: Hailuo counts characters, Veo counts tokens, so the two limits are not directly comparable. Google's Gemini API video page carries the line "Use Gemini Omni Flash as your default model for video generation", so Veo 3.1 is not the default even inside Google's own API docs. Both Veo 3.1 and Omni Flash carry -preview in their model IDs. And Kling's specification pages are client-rendered with no server-side spec text, so the widely repeated "15 seconds at 4K and 60fps" figure could not be confirmed at any Kling-owned page. The 4K claim is corroborated by Kling's pricing tiers; the duration and frame rate are not, so they read as not published here rather than as facts.
Separately, if you were weighing Sora: OpenAI's deprecations table lists the Videos API, sora-2 and sora-2-pro for removal on 24 September 2026, with no named replacement. Do not build a 2026 pipeline on it. For the wider model-by-model breakdown, see Veo 3 vs Sora vs Kling.
Why do Hailuo prompts fail?
Five failure modes cover most of it, and four of the five are prompt-side.
Too many beats. Six seconds holds one action. Two if the second is a small one. A prompt with four verbs gets compressed, and compression is what "sped-up AI video" looks like.
Camera commands buried in prose. MiniMax says free-form descriptions work but explicit commands are more accurate. If your camera move is a sentence rather than a bracket, you are on the less reliable path.
Contradictory camera commands. [Push in] with [Pull out] in the same bracket is not a smooth in-and-out; it is two conflicting instructions arriving simultaneously. Use the sequential form with "then" instead.
The optimizer rewriting your work. Covered above. If identical prompts give different results, check the flag before you blame the model.
Sensitive content rejection. MiniMax returns status code 1026 for "sensitive content detected in prompt" and 2013 for invalid input parameters. A 1026 is a prompt rewrite, not a retry. A 2013 usually means a duration and resolution pair that the model does not support, which for Hailuo 2.3 almost always means someone asked for 10 seconds at 1080P.
If your output looks technically fine but generically bad, that is a different problem, and why your AI videos look generic is the better read.
How do you keep Hailuo prompts reusable?
Turn the six-part template into a stored prompt template with variables for subject, action, camera, space and light, then fill the variables per shot instead of retyping the structure. A shot list of twelve becomes twelve variable sets against one template.
Being direct about what our own product does here: Prompt Architects does not generate video. We build the prompt. You still render in Hailuo or through MiniMax's API, and you still pay MiniMax per clip. What we handle is the layer above that, which is the part that gets messy once you have more than about twenty prompts: video prompt generation, a personal prompt library that syncs across devices, global variables so one wardrobe change updates every shot, and JSON output when you want to hand a structured object to a script rather than a paragraph to a text box.
The free plan gives you 5 prompt enhancements per day, forever, per our FAQ, which is enough to test whether the template shape above holds up on your own footage before you pay anything. Current paid pricing sits on the pricing page, and it is discounted at the time of writing, so check it there rather than trusting a number in a blog post.
If you would rather build the library by hand, how to build a personal AI prompt library walks through the folder and tagging structure, and JSON video prompt templates covers the structured-output version of the same idea for other models.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.
Create An Account