Back to blog
Video10 min read

Crowd and Background Character Prompts

Crowd scene AI video prompts: why OpenAI documents a character-consistency limit but Google doesn't, per-model figure caps, and the composition workarounds that actually help.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: Crowds and background characters break AI video and image models harder than any single subject does, because per-generation consistency is documented as a limitation for a handful of named figures at most. OpenAI admits this directly; Google's own limitations page does not. The workaround isn't a parameter. It's describing crowds as texture and movement, keeping them off-center and out of focus, and accepting that any legible background face is a re-roll.

One honesty note first. Prompt Architects writes and structures the prompt text; it does not render video or images. Nothing here generates a crowd scene. This is a guide to prompting crowds and background characters for whichever model, Veo, Kling, Seedance, or a Gemini image model, actually produces the clip.

Why Do Crowds and Background Characters Break AI Video Models?

Because consistency across generations is already a documented weak point for a single character, and a crowd asks the model to hold that same weak point steady across dozens of faces at once, with no reference image anchoring any single one of them.

OpenAI states this outright. Its current GPT Image guide, under its own "Limitations" heading, says: "While capable of producing consistent imagery, the model may occasionally struggle to maintain visual consistency for recurring characters or brand elements across multiple generations." That's a vendor admitting the failure mode in its own documentation, for its own current model family.

Google's equivalent page does not carry the same admission. Its current Gemini image-generation limitations section lists language coverage, the model not always following an exact requested image count, per-model reference-image caps, and a tip about generating text before the image it sits on. Nothing in that list states that character consistency degrades across generations the way OpenAI's page does. That's a genuine difference in what the two vendors choose to document, not a gap in this research: don't write "Google admits the same limitation" because, on the page checked, it doesn't.

Either way, a background crowd is the worst-case version of a problem both vendors' models have to some degree. One named character already risks drifting between generations. Ten unnamed background figures drift independently of each other and of the foreground subject, and nothing in any of the vendor docs checked for this post claims otherwise.

How Many People Can a Model Actually Keep Visually Consistent at Once?

A small, specific, and surprisingly low number, once you look at what the vendors actually publish rather than assume.

On the image side, Google's Gemini API documentation states that gemini-3.1-flash-image "supports character resemblance of up to 4 characters and the fidelity of up to 10 objects in a single workflow." That's a hard, vendor-stated ceiling on named-character consistency, for the current flagship Nano Banana 2 model, not an inferred one.

On the video side, Kling's Elements system, the feature built specifically for keeping a named character or object consistent across a generation, caps how many distinct elements you can reference at once. Kling's own onboarding guide states that when a video reference is supplied, up to four images or elements can be uploaded; without a video reference, that rises to seven.

Four to seven is the honest range for named, individually consistent figures across the vendors checked here. A wedding-hall scene with forty guests, a stadium crowd, a busy trading floor, none of that is in the same category as directing four named Elements. Past that single-digit ceiling, the model isn't tracking individuals anymore, whatever the prompt asks for.

The Composition Workarounds That Actually Help

None of these are parameters. They're how you write the prompt once you accept the consistency ceiling above.

Describe crowds as texture and movement, not as people. "A dense crowd shifting and murmuring near the platform edge" gives the model a pattern to render. "Twelve distinct commuters, each with their own face and outfit" gives it a request it cannot reliably fulfil, and the failure won't be a refusal, it'll be a handful of warped, half-formed faces in the background.

Keep background figures out of focus and away from the center of frame. Shallow depth of field and off-center placement aren't just cinematographic choices here; they're the two levers that keep a viewer's eye off exactly the part of the image least likely to hold up. A background crowd rendered as a soft blur reads as intentional. The same crowd rendered in sharp focus reads as broken.

Accept that a legible background face is a re-roll, not a request. If a shot genuinely needs one background extra to be recognisable, closer to camera, in focus, doing something specific, treat that figure as a named subject with its own reference image and its own generation, not as part of the crowd. Asking the crowd-level prompt to also produce one sharp, consistent face inside it is asking for two different things from the same instruction.

Use camera_fixed where it exists, and read its own honesty about itself. Seedance's camera_fixed boolean (default false) is one of the only genuine camera-behavior parameters documented in this space, and BytePlus's own docs describe exactly what it does and does not guarantee: setting it to true appends a fixed-camera instruction to your prompt, but, in the vendor's own words, "the actual result is not guaranteed." A locked-off camera at least removes one variable, camera drift, from a crowd shot that already has enough going on.

A native square frame changes how much crowd you're committing to. Kling 3.0 Omni's settings.aspect_ratio accepts 16:9, 9:16, or 1:1, a native square Veo doesn't offer at any resolution tier; Veo 3.1's aspectRatio is 16:9 or 9:16 only, across every current variant checked. A square frame crops out the wide peripheral space where background figures usually drift furthest from the prompt, which is a real, if incidental, way to reduce how much crowd behavior you're asking the model to hold together at once.

Model / fieldWhat it actually controlsDefaultRelevant to crowd shots because
Seedance camera_fixedWhether the camera is locked offfalseRemoves camera drift as a variable; result still "not guaranteed" per BytePlus's own docs
Kling settings.aspect_ratioFrame shape: 16:9, 9:16, 1:116:9Native square isn't available on Veo at any tier; changes how much peripheral crowd is in frame
Kling settings.audioWhether the clip generates soundoffA silent crowd scene is the default on Kling; Veo always generates audio
Kling settings.multi_shotWhether the prompt's shot grammar produces multiple shotstrueLets a crowd stay background in a wide shot, then cut to one named foreground figure in a second shot

How Do You Prompt for Multiple People Without Losing the Composition?

The instinct is to count. "Twelve distinct commuters" or "a group of six friends, each with a different outfit" reads like a precise, well-specified prompt, and it's exactly the kind of instruction that produces the worst results, because it asks the model to solve individual identity for every person named, at the same consistency ceiling that struggles with four to seven Elements on the vendors checked above. A multiple people video prompt that actually holds together stops counting heads and starts describing density and arrangement instead.

Think in layers rather than a headcount. A foreground layer holds the one or two figures the shot is actually about, described with real detail: clothing, pose, expression. A midground layer holds loosely described secondary figures, "two or three people at the next table, mid-conversation," where some variation between generations is expected and fine. A background layer holds everything past that as pure density and motion, no count, no individual description at all. Most crowd and background actors AI video failures come from writing all three layers with the same level of specific, individual detail, which is the one thing none of the models checked for this post can actually deliver past a handful of named figures.

Does Kling Generate Crowd Noise or Ambient Sound by Default?

No, and this catches people who assume every current video model works like Veo. Kling 3.0 Omni's settings.audio field defaults to off, with native and original as the other documented values, so a filled café or a busy street renders silent unless you set it explicitly. Veo 3.1 is the opposite: it always generates audio alongside video on every variant, so the same crowd-ambience expectation, background chatter, footsteps, traffic, produces a real result on one model and dead silence on the other unless you know to ask.

A Worked Example: Prompting a Filled Café Background Without Naming Anyone

The goal here is one named presenter in focus, with a believable but unresolved crowd behind them, using Kling's multi-shot grammar and its settings.audio field explicitly:

settings: { multi_shot: true, duration: 8, audio: "native", aspect_ratio: "16:9" }
prompt:
"shot 1, 4, medium shot on @Presenter speaking directly to camera at a
café counter, sharp focus, shallow depth of field; behind them, a dense,
softly blurred crowd of café patrons shifts and murmurs, no individual
face resolved, warm ambient chatter and espresso-machine hiss; shot 2, 4,
@Presenter gestures toward the counter, background crowd remains a soft,
out-of-focus mass of movement and color, camera holds steady;"

Only one figure is named or referenced as an Element. Everything behind them is described as density, blur, and sound, not as a roster of individual people, which is the actual difference between a crowd shot that holds up and one that produces a background of half-formed faces.

Where Do Presenters and Named Characters Fit Into a Crowd Scene?

Separately from the crowd itself, and worth planning that way from the start. If your shot needs one presenter or protagonist to stay recognisable across multiple generations, that's the problem covered in our guide to video character consistency: reference images, saved character assets, and the other levers that exist specifically because a single named figure is a tractable problem in a way a crowd isn't. If that same shot is part of a longer sequence built shot by shot, our guide to the anatomy of a video prompt and to Kling's camera and motion control cover the surrounding fields this post didn't need to repeat.

This split, one named figure handled deliberately, everyone else handled as texture, is also exactly what an explainer video with a background office or classroom scene needs: a presenter you can actually keep consistent, and a background crowd you were never going to hold together anyway. Fighting the model on the second half is how a background of warped faces ends up in an otherwise clean clip, and it's a large part of why AI video still looks generic even when the foreground subject is well directed. Plan the split before you write the first shot, not after a crowd shot comes back wrong, and the rest of the prompt gets noticeably easier to write.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account