You make AI images of the same person consistently by conditioning every generation on a reference image of that person, not by describing them in words — words alone cannot hold an identity steady across generations.
The reference image does the identity work; your prompt handles pose, lighting, framing, and style. Midjourney V7, for example, ships two features built for exactly this: Omni Reference, which locks a character's look to a supplied image, and Personalization v2, which tunes output toward your own aesthetic preferences. According to our AI tool database, Midjourney's V7 release includes Draft Mode, Omni Reference, Personalization v2, Niji 7 for anime, and a V1 Video Model, and the tool carries an editorial rating of 4.8 out of 5 as the benchmark for AI art quality.
The practical consequence is simple: identity consistency is a reference-image problem first and a prompt problem second.
Start with a clean reference image
The mechanism behind identity drift is worth understanding, because it explains why your first attempt usually fails. An image model does not store a memory of "Sarah." Each generation is a fresh sampling from a probability distribution shaped by your inputs.
If your only input is the phrase "a woman with brown hair and glasses," the model samples a new face every time, because nothing in the input pins down which face. A reference image changes what the model conditions on: instead of sampling any plausible face, it samples faces near the one you supplied.
That is why Omni Reference exists — it gives the model a concrete anchor rather than a verbal sketch. The anchor is only as good as the image you feed it. A sharp, front-facing, evenly lit photo with the face clearly visible will hold far better than a blurry group shot where your subject is three people deep.
Shoot or choose one reference image per person and reuse it across every generation. If you want a second angle, add a second reference rather than describing the angle in words.
Structure the prompt around everything except the face
Once the reference is doing the identity work, your prompt should spend its words on the variables you actually want to change. A workable structure is: subject reference, then action, then setting, then lighting, then style, then aspect ratio. A concrete example: with a reference photo of your subject attached, the prompt "[reference] standing in a rain-soaked Tokyo alley at night, neon reflections on wet pavement, shallow depth of field, cinematic lighting, 16:9" will keep the face anchored while changing everything else.
If instead you write "a man in his thirties with a short beard and tired eyes standing in a rain-soaked Tokyo alley," you have handed the model a fresh face description and it will happily invent one. The rule of thumb: if a word in your prompt describes what the person looks like, delete it and let the reference carry it.
Keep the same reference image, change only the scene words, and generate several variations at once so you can compare identity retention side by side.
Iterate by changing one variable at a time
Identity drift compounds when you change too much at once. If you switch the reference image, the pose, the lighting, and the aspect ratio in a single step, you cannot tell which change broke the likeness. Change one thing per round.
If the face drifts when you move from a portrait to a full-body shot, the cause is usually that the face occupies fewer pixels, so the model has less identity signal to work with. Fix that by generating a tighter crop first, then widening the frame in a later pass, or by adding a second reference image that shows the same face at a similar scale.
Draft Mode in Midjourney V7 is useful here: it produces faster, lower-cost iterations, so you can test a batch of scene variations cheaply before committing to final renders. Treat the first few generations as calibration, not as deliverables.
Know where this breaks down
Reference-image conditioning is not a guarantee. It holds well for moderate changes in pose, clothing, and setting, and it struggles when the face is small, turned sharply, heavily shadowed, or stylised into a medium the model has less training signal for. It also does not give you a permanent, reusable "character" the way a game engine would; each session starts from the reference image again.
Cost matters too. Midjourney's plans, per our database snapshot, run Basic at $10 per month, Standard at $30, and Pro at $60, with a Mega tier at $120 for the V7 entry — and the faster you iterate, the more of your monthly generation allowance you burn. Pricing and plan details change, so check the vendor's own page before you commit.
Finally, if you need a specific real person's likeness, check the platform's rules on depicting real people; most have restrictions, and consent is your responsibility, not the tool's. The honest summary: reference conditioning gets you recognisable consistency, not forensic identity, and the gap between those two is where most disappointment lives.