Back to blog
Video21 min read

Wan 2.7 Prompt Templates: 20 Copy-Paste Prompts (2026)

Wan 2.7 prompt templates for text-to-video, image-to-video and reference-to-video, with every parameter checked against Alibaba Cloud Model Studio's own API docs (August 2026).

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: Wan 2.7 is the newest Wan video model in Alibaba Cloud Model Studio, served as wan2.7-t2v, wan2.7-i2v, wan2.7-r2v and wan2.7-videoedit. Prompts run up to 5,000 characters, output is 2 to 15 seconds at 720P or 1080P, 30fps, with a generated audio track. The weights are not published for local use.

I went looking for Wan 2.7 weights to run on a home GPU. They are not there. That single fact reshapes how you should write prompts for this model, so it goes first, before a single template.

What are Wan 2.7 prompts, and what does the model actually accept?

A Wan 2.7 prompt is a plain-text description sent to Alibaba Cloud Model Studio's video-synthesis endpoint, and the interesting part is how much of the shot it is expected to carry. Wan 2.7 removed the structured shot_type field that earlier versions used, so shot count, cuts, timing, dialogue and sound now all live inside the prompt string.

Here is the full parameter surface, taken from the Wan2.7 text-to-video API reference on Alibaba Cloud Model Studio, last updated July 1, 2026 and read on August 26, 2026.

ParameterAccepted valuesDefault
promptChinese or English, up to 5,000 characters, silently truncated beyond thatrequired
negative_promptup to 500 characters, truncated beyond thatnone
resolution720P or 1080P1080P
ratio16:9, 9:16, 1:1, 4:3, 3:416:9
durationinteger, 2 to 15 seconds5
prompt_extendtrue or falsetrue
watermarktrue or false (adds "AI Generated", lower right)false
seedinteger, 0 to 2147483647random
audio_urlWAV or MP3, 2 to 30s, up to 15 MBnone

The tier plus ratio combination resolves to fixed pixel dimensions: 1920x1080 at 1080P and 16:9, 1080x1920 at 9:16, 1440x1440 at 1:1, 1648x1248 at 4:3. The 720P equivalents are 1280x720, 720x1280, 960x960 and 1104x832. Every output is MP4 with H.264 encoding at 30fps, per Model Studio's text-to-video guide.

Two operational details matter more than they look. First, calls are asynchronous only: you create a task, then poll for a task_id, and Alibaba recommends a 15-second polling interval because generation takes one to five minutes. Second, the returned video URL expires after 24 hours and is then purged, so a pipeline that does not download and re-store immediately is a pipeline that loses its own output.

Is Wan 2.7 open weights you can run locally?

No, and any guide telling you to download it is wrong. I checked both channels Wan publishes through, on August 26, 2026. The Wan-Video GitHub organisation lists five repositories: Wan2.1, Wan2.2, Wan-Animate-2, Wan-Dancer and Wan-skills. There is no Wan2.5, Wan2.6 or Wan2.7 repository. A HuggingFace model search for "Wan2.7" returns zero results, and the Wan-AI organisation's newest entries are Wan2.2-Animate-2-14B and Wan-Dancer-14B, both from July 2026.

So Wan 2.7 is reachable as a hosted API and not as a checkpoint. Alibaba's own pricing page lists wan2.7-t2v, wan2.7-i2v, wan2.7-r2v, wan2.7-videoedit and the wan2.7-image pair, with dated snapshots including wan2.7-t2v-2026-06-12, wan2.7-t2v-2026-04-25, wan2.7-i2v-2026-04-25 and wan2.7-r2v-2026-06-12. Those snapshot dates are the closest thing to a public release timeline that Alibaba's documentation gives you.

If local generation is the requirement, the open branch stops at Wan 2.2. That family is licensed Apache 2.0, stated in the repository's own License Agreement section and in the model cards. Wan2.2-TI2V-5B runs 720P at 24fps and, per the README, works on a GPU with at least 24GB of VRAM such as an RTX 4090, producing a 5-second 720P clip in under nine minutes without further optimisation. The larger T2V-A14B and I2V-A14B mixture-of-experts models support 480P and 720P and the README's single-GPU command asks for at least 80GB of VRAM.

What is the Wan 2.7 prompt formula?

Alibaba publishes six of them, in the Model Studio prompt guide last updated June 3, 2026. They are worth quoting because they tell you the order the model was trained to expect, and order matters more in video prompting than most people assume.

Basic:            Entity + Scene + Motion
Advanced:         Entity + Scene + Motion + Aesthetic control + Stylization
Image-to-video:   Motion + Camera movement
Sound:            Entity + Scene + Motion + Sound description
Multi-shot:       Overall description + Shot number + Timestamp + Shot content
Reference:        Reference identifier + Action + Scene + Lines + Background music

Aesthetic control is where most weak prompts fail. Alibaba defines it as light source, lighting environment, shot size, camera angle, lens and camera movement, and its documented vocabulary includes terms like daylight, firelight, overcast light, rim light, side light, soft light, close-up, close shot, medium full shot, extreme full shot, long-focus lens, ultra-wide-angle fisheye, tilt-shift, over-the-shoulder shot, high-angle shot, aerial shot, clean single shot, two-shot, group shot, centre composition, left-heavy composition and balanced composition. Use those words rather than invented synonyms.

The guide also attaches intent to camera moves, which is unusually useful: a push-in creates intimacy or tension, a pull-out reveals scale, a tracking shot places the viewer alongside the subject, an orbit signals that the subject is central, and a fixed camera signals stillness. It advises keeping an orbit arc under 45 degrees, because wider arcs risk spatial distortion.

One caveat on the sample renders in that guide. The sound-section examples are labelled as generated with Wan 2.5 Preview, and the cinematic dictionary and stylization examples are labelled Wan 2.2. The vocabulary is current; the reference videos beside it were not made with 2.7.

Wan 2.7 text-to-video prompt templates

Fill the bracketed slots and delete what you do not need. Each of these follows the advanced formula, front-loading aesthetic control the way Alibaba's own examples do.

Template 1: cinematic character close-up

Generate single shot. Soft light, side light, warm tones, low saturation, medium close-up,
eye-level shot, long-focus lens, centre composition.
[AGE + APPEARANCE] wearing [WARDROBE] sits in [SPECIFIC INTERIOR].
[LIGHT SOURCE] falls from [DIRECTION], leaving [SHADOW BEHAVIOUR] across their face.
They [SMALL PHYSICAL ACTION], then slowly [SECOND ACTION], eyes [EXPRESSION].
The background is blurred but readable as [PLACE].
The camera holds steady. No dialogue.

Template 2: product hero on a locked-off camera

Generate single shot. Studio lighting, top light plus soft fill, high contrast, cool tones,
close shot, eye-level, medium-focus lens, centre composition, fixed camera.
A [PRODUCT] in [MATERIAL AND FINISH] stands on a [SURFACE] against a [BACKDROP].
A slow highlight travels across [SPECIFIC EDGE OR SURFACE] as the [PRODUCT] rotates
a quarter turn. Dust motes drift through the key light.
Sound effects: a single low resonant tone as the rotation settles. No dialogue.

Template 3: establishing landscape with a drone move

Generate single shot. Clear sky light, hard light, warm tones, extreme full shot, aerial shot,
wide-angle, establishing shot, balanced composition.
[TERRAIN] stretches to the horizon under [WEATHER]. [SECONDARY ELEMENT] moves through
the lower third of the frame.
The camera flies forward and slightly upward, revealing [WHAT IS HIDDEN AT FIRST].
Sound effects: wind across open ground, distant [AMBIENT SOURCE].
Background music: a slow ascending string score.

Template 4: macro texture and liquid

Generate single shot. Backlight, soft light, low contrast, close-up, macro, shallow depth
of field, centre composition.
[SUBSTANCE] [ACTION VERB: pours, cracks, blooms, separates] across [SURFACE] in slow motion.
Individual [PARTICLES OR DROPLETS] catch the backlight and scatter.
The camera pushes in slowly, holding focus on [EXACT DETAIL].
Sound effects: the close, wet [SPECIFIC SOUND] of [SUBSTANCE] meeting [SURFACE].
No dialogue. No background music.

Template 5: handheld street documentary

Generate single shot. Overcast light, soft light, low saturation, cool tones, medium shot,
eye-level, medium-focus lens, left-heavy composition, documentary photography style.
[PERSON] walks through [STREET DESCRIPTION] carrying [OBJECT].
Pedestrians cross frame in the foreground and briefly obscure the subject.
The camera follows from a parallel side view, handheld, with slight natural instability.
Sound effects: traffic, footsteps on wet pavement, fragments of unintelligible conversation.

Template 6: negative prompt block

low resolution, error, worst quality, low quality, deformed, extra fingers,
bad proportions, watermark, on-screen text, subtitles, jump cut, duplicated limbs,
warped face, flicker

That negative-prompt phrasing is adapted from Alibaba's own example in the API reference. Remember the 500-character cap: anything past it is truncated without warning.

Wan 2.7 camera-movement prompt templates

Camera language is the fastest quality lever in any video model, and it is the same skill whichever one you use. If you want the wider vocabulary, our 30 cinematic camera prompts for Veo 3 and Kling covers the same grammar in a different dialect.

Template 7: push-in for tension

Generate single shot. [LIGHTING BLOCK], medium shot narrowing to medium close-up,
centre composition. The camera pushes in slowly on [SUBJECT] as they [ACTION].
The push begins wide enough to show [CONTEXT] and ends tight enough to exclude it.
The subject does not move toward camera. Movement is camera only.

Template 8: pull-out reveal

Generate single shot. [LIGHTING BLOCK], starting close shot, ending extreme full shot.
The shot opens on [SMALL DETAIL]. The camera pulls out steadily to reveal that
[DETAIL] is part of [MUCH LARGER THING], and continues pulling until [FINAL SCALE]
fills the frame. One continuous move, no cut.

Template 9: orbit under 45 degrees

Generate single shot. Backlight, sunset, soft light, silhouette, medium shot,
centre composition, orbiting camera movement.
The camera arcs from behind [SUBJECT] to a three-quarter front angle, no more than
40 degrees of travel, keeping [SUBJECT] centred throughout.
[SUBJECT] stays planted and [SMALL ACTION].

Template 10: compound move

Generate single shot. Drone shot, fast fly-through, [LIGHTING BLOCK].
The camera flies forward through [ENCLOSED SPACE], then rises sharply and exits into
[OPEN SPACE], then slows and tilts down toward [SUBJECT].
Three linked movements, no cuts between them.

Wan 2.7 multi-shot prompt templates

This is the section that most rewards switching from Wan 2.6 to 2.7. On 2.6 you set shot_type to multi and enabled prompt rewriting. On 2.7 that field is gone and the structure is written out longhand. Alibaba's FAQ is explicit: if the prompt contains no shot instruction, the model reads the semantics and decides for itself whether to cut, which is a coin flip you do not want in a client deliverable.

The pattern is: one overall description, then numbered shots with timestamp ranges, then shot content. Keep the timestamps inside your duration value.

Template 11: five-shot brand story at 15 seconds

Generate a multi-shot video. [ONE-SENTENCE THEME AND EMOTIONAL ARC].
Shot 1 [0-3s] Extreme full shot: [ESTABLISHING LOCATION], [LIGHT], [SUBJECT ENTERS].
Shot 2 [3-6s] Medium shot: [SUBJECT] [ACTION], camera tracks parallel.
Shot 3 [6-9s] Close-up: [SPECIFIC DETAIL OR EXPRESSION], fixed camera.
Shot 4 [9-12s] Medium shot: [COMPLICATION OR TURN], camera pushes in.
Shot 5 [12-15s] Wide shot: [RESOLUTION], camera pulls out and holds.
Keep [SUBJECT] visually identical across all five shots.

Template 12: three-shot product explainer at 10 seconds

Generate a multi-shot video. A calm, precise product demonstration in a bright studio.
Shot 1 [0-3s] Close shot, top light, centre composition: hands lift [PRODUCT] into frame.
Shot 2 [3-7s] Medium close-up, side light: hands operate [SPECIFIC MECHANISM],
the action clearly readable.
Shot 3 [7-10s] Close shot, soft light, shallow depth of field: [PRODUCT] rests on
[SURFACE], camera pushes in fractionally and stops.
Consistent hands, consistent product finish, consistent lighting temperature throughout.

Template 13: hard lock to a single shot

Generate single shot. One continuous take, no cuts, no transitions, no scene changes.
[FULL ADVANCED-FORMULA PROMPT GOES HERE]

Template 13 looks trivial and is the one I would keep closest to hand. Video models cut when they run out of ideas, and an explicit single-shot instruction is the cheapest fix for a clip that mysteriously teleports halfway through. The same failure mode shows up across every model in this category, which we unpacked in why your AI videos look generic.

Wan 2.7 audio and dialogue prompt templates

Wan 2.7 writes its own soundtrack. If you do not pass audio_url, the docs say the model generates background music or sound effects matched to the video. If you do pass a track, it is used as-is, truncated to the video length, with silence for any remainder.

Alibaba's sound formula splits into three components, each with its own sub-formula: voice is lines plus emotion plus tone plus speed plus timbre plus accent; a sound effect is source material plus action plus ambient sound; background music is score plus style.

Template 14: single speaker to camera

Generate single shot. Single speaker. [LIGHTING BLOCK], medium close-up, eye-level,
centre composition.
[SPEAKER DESCRIPTION] stands in [LOCATION], [POSTURE AND HANDS].
They look directly at the camera and say, "[EXACT LINE]", in a [EMOTION] tone,
at a [SPEED] pace, with a [TIMBRE] voice, in [ACCENT] English.
Ambient audio: [ROOM TONE]. No background music.

Template 15: two-hander dialogue

Generate single shot. Group conversation. [LIGHTING BLOCK], medium shot,
over-the-shoulder angle, warm tones.
[Character A: BLACK-COATED ENGINEER] and [Character B: SEATED ANALYST] face each
other across [SURFACE].
[Character A] sets down [OBJECT]. [Character A, low steady voice]: "[LINE ONE]"
Immediately, [Character B] looks up. [Character B, quieter and faster]: "[LINE TWO]"
Hands stay on the table. No other dialogue.

Template 16: Foley-led scene, no speech

Generate single shot. [LIGHTING BLOCK], [SHOT SIZE], [COMPOSITION].
[SUBJECT] [ACTION] in [ENVIRONMENT].
Sound effects: [OBJECT] striking [SURFACE] with a [ONOMATOPOEIA] in a [ROOM
CHARACTERISTIC] space; [SECOND SOUND]; [AMBIENT BED].
No dialogue. No background music.

Template 17: score-led montage

Generate a multi-shot video. [THEME].
Shot 1 [0-4s] [CONTENT]. Shot 2 [4-8s] [CONTENT]. Shot 3 [8-12s] [CONTENT].
Background music: [INSTRUMENT] score in a [GENRE] style, [TEMPO], building through
Shot 2 and dropping to ambience for the last two seconds.
No dialogue.

The suppression phrasings are worth memorising: Alibaba's guide names "No dialogue." and "No background music." as the literal English strings that turn each off.

Wan 2.7 image-to-video prompt templates

The image-to-video endpoint takes a media array with typed entries, and only certain combinations are valid: first_frame alone; first_frame plus driving_audio; first_frame plus last_frame; that pair plus driving_audio; first_clip alone; or first_clip plus last_frame. Each type may appear at most once. Input images run 240 to 8000 pixels per side, up to 20 MB; input clips run 2 to 10 seconds, up to 100 MB.

There is no ratio parameter here. The output aspect ratio follows your input material.

Template 18: first frame only

[SUBJECT IN THE IMAGE] [MOTION VERB] [SPEED ADVERB], while [SECONDARY ELEMENT]
[SECONDARY MOTION].
The camera [pushes in slowly / moves left / holds fixed].
Everything not described stays exactly as it appears in the image.

Keep it to motion and camera. Alibaba's image-to-video formula is deliberately short because the image already supplies entity, scene and style, and re-describing them invites the model to redraw what you gave it.

Template 19: first frame to last frame

The scene transforms continuously from the first frame to the last frame.
[SUBJECT] [PATH OF MOTION BETWEEN THE TWO STATES].
The camera angle gradually [MOVEMENT], finishing at the framing shown in the last frame.
One smooth transition, no cut, no dissolve.

Template 20: video continuation

Continuing from the supplied clip: [SUBJECT] [NEXT ACTION], then [FOLLOWING ACTION].
Match the lighting, wardrobe, framing and camera behaviour of the input clip exactly.
No cut at the join.

Continuation billing catches people out. Alibaba states that if duration is 15 and your input clip is 3 seconds, the model generates a 12-second continuation and the final 15-second output is billed for all 15 seconds.

Here is the full request shape for a first-frame plus audio job, adapted from the API reference:

{
  "model": "wan2.7-i2v-2026-04-25",
  "input": {
    "prompt": "YOUR TEMPLATE 18 TEXT HERE",
    "media": [
      { "type": "first_frame",   "url": "https://your-host/frame.png" },
      { "type": "driving_audio", "url": "https://your-host/voice.mp3" }
    ]
  },
  "parameters": {
    "resolution": "720P",
    "duration": 10,
    "prompt_extend": false,
    "watermark": false
  }
}

How do you write a Wan 2.7 reference-to-video prompt?

Reference-to-video is the mode with the most unusual prompt grammar, and the one most likely to be described incorrectly elsewhere. You pass a media array of reference_image and reference_video entries, optionally with a reference_voice on each, and then you refer to them positionally inside the prompt as "Image 1", "Image 2", "Video 1" and so on. Images and videos are numbered separately. In English, capitalise the word and put a space before the number.

The documented limits: at least one reference image or video, reference images plus reference videos totalling five or fewer, and at most one first_frame. When a reference asset stands in for a character, it must contain only one character. Duration is 2 to 15 seconds when references are images only, and 2 to 10 seconds when any reference video is included.

Bonus template: multi-subject reference scene

Video 1 sits in [LOCATION] holding Image 2, and says, "[LINE ONE]".
Image 1 walks in from frame left carrying Image 3, places it on [SURFACE from Image 4],
and replies, "[LINE TWO]".
The scene is [SCENE DESCRIPTION]. [LIGHTING BLOCK], medium shot, centre composition.
Background music: [STYLE], low in the mix.

Bonus template: storyboard panel reference

Based on the reference image, in the style of [STYLE].
Keep the characters and the [SETTING] consistent with the panels. Do not add text.
Atmosphere: [THREE ADJECTIVES].
Characters: [ONE-LINE DESCRIPTION EACH].
Follow the panel order as shot order.

That second one mirrors an official example in which a single nine-panel image drives story, composition and character design at once. It is the closest thing Wan 2.7 has to a storyboard input, and it is far cheaper than describing nine shots in prose.

What should you never put in a Wan 2.7 prompt?

Alibaba publishes a short list of patterns that fail, which is more candid than most vendors manage. Reproduced with their stated reasons:

AvoidWhy it fails
Named real peopleRejected by most platforms, or rendered inconsistently
Rapid scene changes inside one clipCuts belong between clips; one clip is one continuous shot
Exact text legibilityText renders approximately; specific words will not be precise
Long, complex choreographyA 30-second sequence will not hold; keep actions short
Lip sync to exact wordsDoes not work reliably

Add one more from the API contract rather than the guide: do not rely on seed for reproducibility. Alibaba states plainly that because generation is probabilistic, the same seed does not guarantee identical results. A fixed seed improves reproducibility; it does not deliver it.

How do Wan 2.7, Wan 2.6 and Wan 2.2 differ for prompting?

Prompt-surface differences. Wan 2.7 and 2.6 fields from Alibaba Cloud Model Studio API references; Wan 2.2 from the Wan-Video/Wan2.2 repository README. All read August 26, 2026.
FeatureWan 2.7 (API)Wan 2.6 (API)Wan 2.2 (open weights)
Weights published for local use
LicenceNot published as weightsNot published as weightsApache 2.0
Multi-shot controlNatural language in promptshot_type parameterNot documented
Generates its own audio track
Reference-to-videoUp to 5 image or video refsUp to 3 character refs
Resolution control fieldsresolution + ratiosize--size CLI flag
Max duration15s15sNot stated as a fixed cap

The migration note in Alibaba's own FAQ is two lines long: size becomes resolution plus ratio, and shot_type becomes prose. If you are porting a Wan 2.6 template library, those are the only two mechanical changes, though the prose rewrite is real work.

One documentation inconsistency worth flagging before you build. The Wan2.7 text-to-video API reference states that the DashScope SDK supports Wan 2.6 and earlier and that Wan 2.7 is not supported, pointing you at raw HTTP. The text-to-video user guide, updated the same day, ships a Python SDK sample calling wan2.7-t2v-2026-06-12 and asks for DashScope Python 1.25.16 or later. Both pages were live on August 26, 2026. Check your SDK version against the guide, and keep an HTTP fallback.

What does a Wan 2.7 generation cost?

From Alibaba Cloud Model Studio's model pricing page, International deployment scope, read August 26, 2026: wan2.7-t2v, wan2.7-i2v and wan2.7-r2v are all listed at $0.10 per second at 720P and $0.15 per second at 1080P, with a free quota of 50 seconds valid for 90 days after you activate Model Studio. Only output is billed, and Alibaba states that failed calls do not incur charges or consume free quota. Prices are region-specific and change, so check the pricing page before you budget.

A 10-second 1080P clip is therefore about $1.50 of output. That reframes prompt discipline as a cost question rather than an aesthetic one. Every regenerated clip is another dollar-fifty, and the default prompt_extend: true means some of those regenerations are you fighting a rewrite you never asked for.

Where to keep these templates so you actually reuse them

Twenty templates in a browser tab is a library for about four days. The failure mode is predictable: you tweak Template 11 for a client, the tweak is better, and it lives in that one chat window until it is gone.

A prompt template only compounds if it is stored with variables rather than pasted and edited. That is what Prompt Architects does. You save the template once, mark [SUBJECT], [LIGHTING BLOCK] and [STYLE] as global variables, and fill them per project instead of retyping the shot grammar. The same library covers your text prompts, your JSON prompts and your video prompts in one place, and the free plan gives you 5 enhancements per day, forever, per our FAQ.

If you work in JSON rather than prose, the structural argument carries across models, and our JSON video prompt templates for Veo 3 show the same shot grammar in object form. For the underlying craft of shot design, how to direct AI video like a filmmaker is the companion piece to this one, and if you want a broader starting set, 100+ ChatGPT prompt templates covers the non-video half of the workflow.

One last honest note. We build the prompt layer. We do not generate the video, we do not host Wan, and we are not affiliated with Alibaba. For generation you need a Model Studio account with an API key in the same region as the endpoint you call, or a local Wan 2.2 setup if Apache 2.0 weights are what you actually needed.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account