TL;DR: A portrait prompt has five real levers: light, lens, framing, expression, and wardrobe, plus a choice about how much retouching language to add on top. Naming one classic lighting setup, one focal length, and exactly one emotion beats stacking adjectives every time. This is the technique itself, not another list of finished prompts, plus the one question about whose face you're allowed to put in one.
What do good portrait prompts actually have in common?
Portrait prompts AI models render well share a discipline that general image prompts don't need: restraint. A landscape prompt can stack five mood words and let them blend into an atmosphere. A face has one nose, two eyes, and one mouth to carry all of it, so five lighting adjectives don't blend. They fight over the same eight square inches, and the model has to pick a loser.
Our own list of 35 tested headshot prompts gives you finished text across professional, casual, and creative looks, and our headshot prompt generator turns five fields into a platform-ready prompt for whichever tool you're using. Neither one stops to explain why one lighting setup reads as real and another reads as a video-game cutscene, or why "confident smile" plus "serious expression" produces a face that looks like neither. That's what this page is for: the five variables that separate a usable portrait from a rejected one, applied specifically to a face, plus the one question about consent that a prompt-writing guide has no business skipping.
How does lighting direction and quality actually read on a face?
Pick one light source, one direction, and one quality (hard or soft) before you touch anything else. This is the single highest-leverage decision in a portrait prompt, because a face reveals lighting mistakes a landscape hides: a shadow falling wrong under the nose or across one eye socket is instantly legible as wrong to anyone who has ever looked at another human face, which is everyone.
Three classic setups cover almost every portrait job. Rembrandt lighting, a single key roughly 45 degrees to the side and slightly above eye level, reads as trustworthy and corporate, the look behind most executive and finance headshots. A soft, large light near the lens axis (a beauty dish, or a big window) flattens shadows and reads as editorial or beauty work. A single window light from the side, left mostly untouched, reads as candid and natural, which is what most lifestyle and personal-brand portraits actually want. Naming the setup by name does more work than naming an emotion about it: "soft, moody, atmospheric" tells a model almost nothing; "Rembrandt lighting, soft window light camera-left" tells it exactly where the shadow triangle falls.
For the other 37 lighting terms (rim, split, butterfly, golden hour, and the rest), our full lighting vocabulary reference defines each one with a copy-paste fragment. What matters for a portrait specifically is picking exactly one classic combination and stopping, not working through the list.
Does naming a specific lens actually change an AI portrait?
Yes, and it's worth understanding why before you use it. OpenAI's own prompting guidance for its image models is direct about this: to get believable photorealism, "Use photography language (lens, lighting, framing) and explicitly ask for real texture (pores, wrinkles, fabric wear, imperfections). Avoid words that imply studio polish or staging" (developers.openai.com/cookbook/examples/multimodal/image-gen-models-prompting-guide, fetched September 4, 2026). The model isn't calculating a focal length or simulating glass. It learned what millions of real photos captioned with a lens name typically look like, and reproduces that signature.
For a portrait specifically, that signature has a sweet spot. A lens named in the 50–135mm equivalent range compresses the background and flatters facial proportions, which is why it's the default for LinkedIn and executive work. A wide lens named for a close-up shot pulls in real wide-angle distortion along with it: the same "big nose" effect a real 24mm lens produces two feet from a face, because that's what the training data captioned that way actually looks like. Our camera and lens terms reference covers the full range, the aperture and depth-of-field vocabulary that pairs with it, and exactly where the optics analogy breaks.
Framing and crop, matched to where the portrait is going
Aspect ratio gets most of the attention, but crop distance and eye placement do more of the actual work. A tight beauty crop that fills the frame with a face reads as intimate or editorial; pull back to head-and-shoulders and it reads as a standard professional headshot; pull back further to include hands or torso and it reads as environmental, which needs a reason for the extra space to exist (a desk, a window, a tool of the trade) or it just looks unfinished.
| Crop type | Distance language to use | Reads as |
|---|---|---|
| Tight beauty crop | "fills the frame," "close crop above the shoulders" | Editorial, intimate, high-fashion |
| Head-and-shoulders | "medium close-up," "shoulders-up" | Standard professional headshot |
| Half-body environmental | "three-quarter length," "waist-up with context" | Founder, editorial profile, "about" page |
| Full-body lifestyle | "full-length," "environmental portrait" | Lifestyle, dating profile, personal brand |
Eye placement matters more than most prompts bother with. Naming "eyes on the upper third of the frame" keeps a crop from feeling either cramped at the top or floating in dead space at the bottom, and it's the one framing instruction that survives being translated across Midjourney's --ar, a plain pixel size, or an aspect-ratio field, because it's describing composition in words rather than asking a parameter to do it.
Why does naming two emotions in one prompt ruin the expression?
Because the model doesn't choose between them: it averages. "Confident smile" and "serious, composed expression" in the same sentence doesn't produce a face that's confident and serious in sequence; it produces one flat, uncertain middle-ground stare that reads as neither, which is the single most common reason a technically correct portrait prompt still comes back looking wrong.
The fix is one emotion, described once, plus exactly one asymmetric detail to break the geometric perfection a model defaults to. A real smile is never perfectly symmetrical; a real gaze is rarely dead-center on the lens. "Confident, closed-mouth smile, gaze slightly off-camera" reads as a held expression. "Confident, warm, friendly, approachable, engaging smile" reads as five adjectives arguing, and the model resolves the argument by giving you nothing in particular. Where you point the eyes matters as much as the mouth: direct eye contact with the lens reads as authority; a gaze just past the camera reads as candid and approachable. Pick the one your subject actually needs, name it once, and stop.
Wardrobe language that survives the render
Vague clothing renders a vague garment. "A blazer" gives you a generic blazer in a generic color; "a charcoal wool blazer, notch lapel, open collar, no tie" gives you the one you were picturing, because you've fixed the fabric, the cut, and how much skin shows at the neckline in one pass.
Three details usually cover it: the garment and cut, the color relationship to the background you've already described, and one texture or fit cue. Color matters more than most prompts credit: a pure-white shirt against a bright, blown-out background reads as clipped rather than clean, and a dark suit against a dark backdrop loses its silhouette entirely. Name the wardrobe color with the background already in mind rather than as an afterthought. Past three or four wardrobe details, you're not adding information anymore, you're adding words the model has to reconcile against the lighting and background you already specified, and reconciliation is where contradictions creep in.
How do you separate a subject from the background without losing realism?
Name what's behind the subject, then name how out-of-focus it is, in that order. "Blurred background" alone gives the model permission to blur anything; "a softly blurred bookshelf, camera-left window light spilling across the top shelf" gives it a real scene to blur, which reads as a photograph of a place rather than a face floating in front of a gradient.
This is where lighting and lens choices from earlier compound each other. A longer lens name (85mm-equivalent) plus a named light source produces the shallow-depth-of-field falloff that separates a subject from its background convincingly; a wide lens name plus flat, even lighting collapses that separation and the subject reads as pasted onto the background rather than standing in front of it. Swapping the background entirely (an office for a studio backdrop, or matching a subject to a scene shot separately) is a distinct skill with its own failure modes; our background replacement prompting guide covers the swap itself rather than the separation technique here.
How much retouching language is too much?
Less than you'd think, and the direction matters more than the amount. The instinct is to ask for "flawless skin" or "perfect complexion," and that instinct is exactly backwards: real skin has texture, and asking a model to remove all of it produces the plastic, waxy look that reads as synthetic faster than almost any other single mistake in a portrait prompt.
The practical version: describe the skin the way a documentary photographer would, not the way a beauty retoucher would. "Real skin texture, visible pores, natural under-eye shadow" reads as a photograph. "Smooth, flawless, airbrushed" reads as an illustration wearing a photograph's clothes.
Putting the five variables together
None of the five decisions above works in isolation. A portrait prompt is the five of them landing on the same face at once, in a fairly stable order: who the subject is, what they're wearing, one emotion, one light source, one lens and framing choice, what's behind them, and how much retouching restraint to name. Here's that order as a skeleton, not a finished prompt to copy:
[subject: age band, build, defining feature]
[expression: exactly one emotion + one asymmetric detail]
[wardrobe: garment, cut, color relative to background]
[light: one classic setup, direction, quality]
[lens + framing: focal-length feel, crop distance, eye placement]
[background: a named scene, then its blur quality]
[restraint clause: real texture, no heavy retouching]
[platform settings: whatever the destination model needs]
Fill each bracket with one decision, not a list of options, and you get a prompt built the same way a photographer plans a shoot: subject and wardrobe first, light and lens second, everything else in service of those two. For finished, ready-to-paste versions of that same structure, our 35-prompt list and headshot generator do the assembly for you; this is the reasoning behind why their fields are in that order.
Whose face are you actually allowed to put in a prompt?
Your own, or a face you have real permission to use. Nothing else. The reference-image technique that anchors a portrait to your actual likeness works exactly as well on any other face you feed it, and the technique being easy doesn't grant you consent you don't otherwise have. A stranger's headshot pulled from a search result, an old team photo of someone who's since left, or a public figure's press photo are all the same mistake wearing a different excuse.
The workable rule is the same one that applies to using someone's photo for anything else: if you'd need to ask before publishing their picture, you need to ask before generating a version of their face too. Generate from your own reference, from a hired model who has agreed to the use, or from a genuine, specific yes, and disclose that a portrait is AI-generated when the context calls for honesty, which is most professional contexts. Our headshot prompt list already notes that LinkedIn permits AI-generated photos as long as they reflect a real likeness. The same floor, real likeness, real consent, applies everywhere else, written policy or not.
Stop rewriting prompts. Start shipping.
Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.
Create An Account