TL;DR: AI character consistency in September 2026 comes from three real tools, not one trick: attach-a-reference-image parameters (Midjourney's Edit Model, Nano Banana, GPT Image 2), single-photo consistency models (Ideogram Character), and trained identity models you build once (Higgsfield Soul ID). Seeds do not do it. Every option still drifts sometimes; none is pixel-perfect.
Every image model you have used has done this to you: you generate a character you like, you write the next prompt to put them somewhere else, and a stranger comes back wearing the same jacket. The face is close. It is not the same face. This is the single most requested fix from anyone building a comic, a brand mascot, a course's recurring host, or a set of product photos with one model in them, and it is the reason AI character consistency now gets searched thousands of times a month with almost no page actually naming what each vendor supports.
The honest answer is that no image model, as of this writing, guarantees an identical face across separate generations. What has changed is which levers each vendor gives you to reduce the drift, and those levers were rebuilt substantially in the last two months. This post names them, with the vendor's own words and the date each claim was checked.
This post is about stills. If your character needs to hold up across a moving clip instead of a set of images, the constraints are different enough to need a separate answer: see keeping a character consistent across video clips. And if what you actually need is one consistent art style rather than one consistent face, that is a different problem too, covered in keeping one visual style across a whole series.
Why Does the Same AI Character Never Quite Match Between Images?
Because nothing carries over between two generations except what you hand the model directly. A diffusion or autoregressive image model does not remember the character it drew for you five minutes ago. Each call samples a fresh image from scratch, guided by whatever text and reference pixels you provide in that call. Without a reference image or a trained identity in the request, the only thing tying two generations of "a woman in a red coat" together is the words, and words underdetermine a face: jaw width, eye spacing, the exact shade of hair, all of it gets re-guessed every time.
This is also why some fixes that sound plausible do not work. A random number does not encode a face, a short text description cannot specify sixty facial landmarks, and a style preset controls the rendering, not the identity underneath it. The tools that actually move the needle all do one of two things: they hand the model literal reference pixels to copy from at generation time, or they train a small model on a specific person's photos once, so that model can reproduce that face without a fresh reference every time.
What Are Your Three Real Options for Character Consistency?
Grouped by mechanism, not by vendor, here is what is actually available in September 2026:
- Attach reference images at generation time. You supply one or more photos of the character in the same request as your prompt, and the model conditions on them. This covers Midjourney's Edit Model, Google's Nano Banana family, and OpenAI's GPT Image 2.
- A single-photo consistency model. You upload one reference photo once per character and the tool handles identity extraction automatically, no dataset and no training step. Ideogram Character is the clearest example.
- A trained identity model. You feed the tool a batch of photos of one person, it trains a small dedicated model on that identity, and you reuse the trained character indefinitely without reattaching a photo. Higgsfield's Soul ID is the named example here.
| Feature | Midjourney Edit Model | Nano Banana 2 / Pro | GPT Image 2 | Ideogram Character | Trained model (Soul ID) |
|---|---|---|---|---|---|
| Reference images per generation | Up to 4 | 4 (Nano Banana 2) or 5 (Pro) | No published cap | 1 | 0 — reuses the trained identity |
| Requires training before first use | |||||
| Vendor's own docs name a consistency limit | |||||
| Reattach a photo on every generation |
No row in that table is a guarantee. It is a map of which lever each vendor gives you, so you stop reaching for a parameter that was quietly retired on the version you are running.
Does Midjourney's --cref or --oref Still Work in 2026?
Not on the version most people are running. Midjourney's parameter history here is a genuine trap if you learned it a year ago. Character Reference (--cref, plus --cw for weight) was scoped to Midjourney version 6. Omni Reference (--oref) replaced it as a V7-only feature. Midjourney's Version compatibility chart marks both features unsupported on the merged V8.1 & V8.2 column, and the Character Reference parameter has since been folded into a "Legacy Features" reference page rather than kept as its own current article.
The current path is a different feature entirely, not a renamed parameter. Midjourney's Omni Reference article tells you directly: "When using V8.X, use the Edit Model instead—our improved approach to editing and creating with reference images." The Edit Model article itself, updated September 2, 2026, describes what it does in three parts: make changes to an existing image from written instructions, generate new images using "up to 4 reference images (replacing Omni Reference and Character Reference)", and inpaint or outpaint inside the Editor. You attach it in Discord with the --edit parameter followed by an image URL, or from the website by dragging up to four images into the Imagine bar.
Midjourney's own documentation carries a practical tip worth repeating here: if a scene mixes up features between two characters, combine those characters into a single reference image first rather than feeding the model two separate photos and hoping it keeps them apart. That single instruction fixes more multi-character scenes than any prompt wording does. For the full version-by-version parameter history on this one vendor, see character consistency in Midjourney.
How Many Reference Images Can Nano Banana and GPT Image 2 Actually Take?
Google documents this with real numbers, and they differ by model within the same family, which is easy to miss if you learned "Nano Banana" as one thing (our full Nano Banana Pro prompting guide covers the rest of its parameters). "Nano Banana" is Google's marketing umbrella for four distinct models: Nano Banana 2 Lite (gemini-3.1-flash-lite-image), Nano Banana 2 (gemini-3.1-flash-image), Nano Banana Pro (gemini-3-pro-image), and the older Nano Banana (gemini-2.5-flash-image). Only two of them are built for this task. Google's image-generation guide is blunt about Lite: "Not optimized for multiple reference inputs or multi-turn sequential editing." Nano Banana 2 supports up to 4 images of characters specifically for consistency, and Nano Banana Pro raises that to 5, alongside 3 dedicated style-reference images that Google's table lists as Pro-only. All three sit inside a shared ceiling of 14 total reference images per request across objects, characters, and style.
OpenAI takes a different approach with GPT Image 2 and does not publish an equivalent numeric cap on reference images for an edit call, though its documentation confirms you can supply more than one input image and that a mask, if you use one, applies to the first image in the set. What OpenAI does publish is a direct admission and a direct contradiction of that admission, on two of its own pages. The image-generation guide's limitations section states the model "may occasionally struggle to maintain visual consistency for recurring characters or brand elements across multiple generations." OpenAI's own cookbook, on developers.openai.com, lists among GPT Image 2's key capabilities: "Robust facial and identity preservation for edits, character consistency, and multi-step workflows". Both are live OpenAI pages as of this writing; they are not describing the same model the same way.
There is a second, smaller contradiction worth knowing if you call the API directly. OpenAI's guide says to omit the input_fidelity parameter for GPT Image 2, stating flatly that it "does not work for this model because output is already high fidelity by default". Its own cookbook's code examples do the opposite: every edit call shown sets input_fidelity="high" explicitly, including on an identity-preservation example that asks the model to remove an object from a photo without changing the person in it. If you are building against the API rather than the consumer app, try setting it anyway; the code that OpenAI itself ships disagrees with the prose OpenAI itself wrote next to it.
Should You Use a Single Reference Photo or Train a Dedicated Character Model?
This is the real fork in the road, and it depends on how many images you are about to generate. If you need a character in a handful of scenes, a reference-image tool is faster to start: no dataset, no wait, no separate account setup. Ideogram's Character feature is built entirely around this case. Ideogram's own page states it plainly: "Ideogram Character builds a consistent character from a single reference photo. You don't need to train a custom model, prepare a 10-photo dataset, or fine-tune a LoRA. Upload one image, write a prompt, and generate." It identifies the character through automated facial and hair detection rather than manual tagging, and it is available free on ideogram.ai and the iOS app, with the same model reachable through Ideogram's developer API.
If you are generating a character across dozens or hundreds of images, over weeks, a trained identity model earns its setup cost back quickly, because you stop reattaching a photo on every call. Higgsfield's Soul ID is the clearest named example: you train it once on 20 or more photos of one person, up to 80, and training itself takes a few minutes. After that, the trained character shows up as a selectable option, reused across future generations without a fresh reference image. Higgsfield's own help documentation sets the expectation honestly rather than oversells it: expect "clearly the same person" rather than a pixel-identical face across every generation, and a Soul ID holds one person at a time, with a separate feature, Elements, for scenes needing two or more consistent characters together. It is also worth knowing where the trained identity lives: Higgsfield's docs state "it's used within the Soul models and connected tools and is not exportable as a standalone file", so a trained Soul ID does not travel to Midjourney, Ideogram, or any tool outside Higgsfield's own ecosystem.
Neither path is objectively better. A single-photo tool costs you nothing upfront and works today; a trained identity costs you a short setup and a handful of good source photos, and pays that back once you are past a few dozen generations of the same face.
A Character Sheet Prompt You Can Reuse Across Tools
Whichever mechanism you pick, the prompt text you write around it does real work too. A "character sheet" prompt fixes the details a reference image cannot fully carry on its own, so the model has fewer decisions left to improvise when the pose or scene changes. Here is a template structure worth keeping on hand and adapting per character, written to be paired with whichever reference-image or trained-identity feature you are using:
CHARACTER: [name or short identifier]
IDENTITY ANCHORS (do not vary): [face shape, eye color, hair color and style,
skin tone, any distinguishing marks — scars, freckles, glasses]
BUILD: [height impression, build, posture]
WARDROBE ANCHOR: [one signature item that should appear in most shots —
a jacket, a specific color, an accessory]
STYLE: [photographic / illustrated / 3D render, and the specific look]
SCENE: [what changes this generation — pose, setting, lighting, action]
Keep IDENTITY ANCHORS and WARDROBE ANCHOR fixed. Only SCENE should change
between generations of this character.
Feeding a prompt built this way alongside a reference image gives the model two independent signals pointing at the same identity, which is exactly why Midjourney's own tip about combining multiple character photos into a single reference works: it reduces the number of competing signals the model has to reconcile in one generation. This is describing what a well-structured prompt produces when paired with a model's reference feature, not a guarantee of an exact result; every vendor covered above still documents drift as a real possibility. If you are keeping several of these character sheets for a recurring cast, Prompt Architects' library lets you save the identity block once as a reusable prompt and pull it back out for the next scene instead of retyping it, which matters more than it sounds once you are past your third character.
What Should You Do When the Face Still Drifts Anyway?
Direct around it rather than fighting it. A few things consistently help across every tool above. First, do not use a seed as a consistency mechanism: Midjourney's own Seeds article is explicit that seeds "can’t capture or bookmark a specific style, character, or appearance across different prompts", and recommends style references, the Edit Model's reference images, and Personalization instead, because a seed only fixes the initial noise pattern, not the subject. Second, when a scene needs more than one of your characters, combine their reference photos into a single composite image before generating rather than supplying two separate photos and hoping the model keeps them apart, per Midjourney's own advice on the Edit Model.
Third, treat difficult angles as a craft problem rather than a prompting problem. Extreme close-ups and unusual angles are exactly where a trained or referenced identity is most likely to show visible drift, according to Higgsfield's own limitations note; the same logic applies across every model here, because more of the face is exposed to error at close range. Shooting your character at a slight distance, in partial profile, or with consistent lighting removes a lot of the drift a straight-on beauty shot would expose. None of this makes any model perfect. It does make the difference between a character your audience recognizes across ten images and one that quietly turns into someone else by image four. If your next step is turning these stills into a moving sequence, the style has to hold up shot to shot too, which is its own separate problem covered in style consistency across a whole video sequence.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.
Create An Account