AI image and video models have no memory. Every generation starts from zero and rebuilds your character out of the words in your prompt — and "blonde woman, mid-thirties" describes about four million different faces.
The fix isn't a better tool or a longer prompt. It's a reference pack: a small set of locked images and a fixed written spec that you feed into every single generation from that point on. Build it once, never rebuild it.
You spend an evening generating your avatar. You get one you love. You save it, you post it, you feel like you've finally started.
Then you go to make the second piece of content and she's gone. The nose is wrong. The jaw is heavier. Her hair sits differently. She's recognisably a similar sort of woman and completely not the same person, and now you're sat there generating variation after variation trying to find your way back to a face you already had.
Nearly every woman I've watched start a faceless AI brand loses her first two weeks to this. Not to strategy, not to content, not to anything that matters. To chasing a face.
It isn't a skill problem and it isn't your prompt writing. It's that you've been asked to do something the tools were never set up to do.
Why the drift happens
Image models are stateless. There's no file somewhere called "your avatar" that the model opens each time. Every generation is the first generation as far as the model is concerned. It reads your words, builds a person who satisfies those words, renders her, and forgets her completely.
So the question stops being "how do I get her back" and becomes: how much of her did my words actually pin down?
Usually, almost none of it. Take a typical avatar prompt:
"Beautiful blonde woman in her early thirties, blue eyes, natural makeup, wearing a cream blazer, studio lighting."
Everything in there is a category, not a person. Blonde is a range. Early thirties is a range. Blue eyes covers everything from pale grey to deep navy. Nothing in that sentence constrains the distance between her eyes, the width of her nose, the shape of her jaw, the height of her cheekbones — and those are precisely the things your audience uses to decide whether two images are the same woman.
The four things that actually drift
When people say their avatar is inconsistent, they usually can't say what changed. It's worth being precise, because each one has a different fix.
1. Face geometry
The proportions: eye spacing, nose width and bridge shape, jaw and chin, cheekbone height, brow position. This is the one that makes her read as a different human being, and it's the one words are worst at controlling. No amount of adjectives will hold it.
2. Colouring
Skin tone, undertone, hair colour and hair level, eye colour. Drifts more subtly, but it's why your avatar can look tanned in one image and porcelain in the next. Often caused by lighting changes rather than the model itself.
3. Build and age read
Height, frame, how old she reads. AI has a strong pull towards a default: young, slim, symmetrical, mid-twenties. If your avatar is meant to be forty-five with a real body, the model will quietly drag her back towards that default every time you give it room to.
4. Her world
Wardrobe, setting, colour palette, the light she lives in. Not identity exactly, but it's a big part of why an account feels like one person. An avatar in a different room, different clothes and different light every post feels like a stock photo library, not a brand.
The first three are fixed with reference images. The fourth is fixed with a written spec you reuse. You need both.
The fix: build a reference pack once
A reference pack is the small set of assets you generate once, at the start, and then never regenerate. Everything you make afterwards is anchored to it.
It contains two halves.
The locked images
Three to five images of your character, generated in a single session, that you commit to permanently:
- A front-facing portrait. Neutral expression, even lighting, plain background, shoulders up. This is your primary anchor and the one you'll use most.
- A three-quarter turn. Same lighting, same day, same session. Gives the model information about the sides of her face it can't get from the front.
- A full-body shot. Locks her build and proportions so she doesn't quietly get taller and thinner over three months.
- One or two in-world shots. Her in the setting she'll actually appear in, so you have a tonal reference as well as an identity one.
The mistake almost everyone makes here is generating a gorgeous reference. Dramatic side lighting, strong expression, moody background. Don't. Every bit of drama in your reference gets inherited by every image you make from it, and harsh shadow across a face actively hides the geometry the model needs to read.
Keep the reference boring. Flat light, plain background, neutral face, highest resolution your tool will give you. The beautiful images come later, generated from the boring one.
The written spec
A short block of text you paste into every prompt, word for word, forever. Not a description you rewrite each time — the same characters in the same order, treated like a name rather than a sentence.
It should pin down what images alone can't hold: her age read, her build, her hair at a specific level and length, her wardrobe register, her palette, her setting. Written once, saved somewhere you can copy it in two seconds.
The reason it must be identical every time is the same reason the drift happened in the first place. Rewording is regenerating. If you write "warm blonde" one week and "honey blonde" the next, you've asked for two different women.
Build her reference pack in about ten minutes
The Avatar Quiz walks you through twelve questions — her face, her colouring, her build, her world — and hands you a complete prompt pack at the end: her reference images plus her first podcast and yap stills, written in the exact format that holds. It's the same method taught inside Human AI, and it's free.
How to actually use the pack
The pack only works if you change one habit: you never generate your character again. You generate scenes, and you attach your character to them.
In practice that means every prompt from now on has three parts, in this order:
- The reference image, attached as an image input — character reference, face reference, or whatever your tool calls it. Not described. Attached.
- Your written spec, pasted in unchanged.
- The scene — and this is the only part you write fresh. Where she is, what she's doing, the light, the framing, the mood.
Most tools now have some version of a character reference or identity feature, and they've got much better over the last year. The naming is different everywhere, the underlying idea is the same: give it a face to hold rather than a description to interpret.
If your tool has an identity feature where you train a character once from several photos and then call it by name — use it. That's the strongest version of this, and it's what your reference pack is for.
If you want to see what the scene half looks like written properly, the UND Bot prompt collection has nine of mine you can copy and paste straight into Higgsfield, OpenArt or anything similar. Swap in your own reference and spec, keep the scene structure.
Test it before you build on it
Before you make forty pieces of content, make four, deliberately different: one close-up, one wide, one in different lighting, one from a different angle. Put them side by side.
If a stranger would say those are four photos of the same woman, you're locked and you can start building. If they wouldn't, go back and fix the reference now — because a drifting avatar doesn't get better with volume, it just gets more expensive to fix.
What will still drift, and what to let go
Full consistency doesn't exist yet and anyone selling it to you is overselling. Some things are still genuinely hard:
- Profile views. The further you get from the angle of your reference, the more the model is inventing. Side-on is where most characters break.
- Extreme close-ups. More facial detail visible means more chances to be wrong.
- Hands, teeth, and jewellery. Still unreliable. Frame around them where you can.
- Big lighting changes. Harsh coloured light rewrites skin tone and can shift the whole read of a face.
Here's the thing worth holding onto, though. Your audience is not doing forensic comparison. They're scrolling. They recognise people the way we all do — hair, colouring, silhouette, the general shape of a face, the world around her.
You don't need pixel-identical. You need unmistakably her at scroll speed. That's a much lower bar, and a reference pack clears it comfortably.
The short version
- Models are stateless. Words describe categories, not people. That's the whole cause.
- Build a reference pack in one session: front portrait, three-quarter, full body, one or two in-world. Keep the lighting flat and the expression neutral.
- Write a spec once and paste it in unchanged forever. Rewording is regenerating.
- Every future prompt: reference image + unchanged spec + new scene. Only the scene changes.
- Test with four deliberately different shots before you build a content library on top of it.
- Aim for unmistakably her at scroll speed, not pixel-perfect. Perfect isn't available and isn't needed.
Get this right in your first week and the rest of it — the content, the voice, the actual business — stops being blocked by a face you can't find twice.