← Blog

Your AI avatar looks like a different person in every video. Here's why.

Emma Jackson · 8 min read

Short answer

AI image and video models have no memory. Every generation starts from zero and rebuilds your character out of the words in your prompt — and "blonde woman, mid-thirties" describes about four million different faces.

The fix isn't a better tool or a longer prompt. It's a reference pack: a small set of locked images and a fixed written spec that you feed into every single generation from that point on. Build it once, never rebuild it.

You spend an evening generating your avatar. You get one you love. You save it, you post it, you feel like you've finally started.

Then you go to make the second piece of content and she's gone. The nose is wrong. The jaw is heavier. Her hair sits differently. She's recognisably a similar sort of woman and completely not the same person, and now you're sat there generating variation after variation trying to find your way back to a face you already had.

Nearly every woman I've watched start a faceless AI brand loses her first two weeks to this. Not to strategy, not to content, not to anything that matters. To chasing a face.

It isn't a skill problem and it isn't your prompt writing. It's that you've been asked to do something the tools were never set up to do.

Why the drift happens

Image models are stateless. There's no file somewhere called "your avatar" that the model opens each time. Every generation is the first generation as far as the model is concerned. It reads your words, builds a person who satisfies those words, renders her, and forgets her completely.

So the question stops being "how do I get her back" and becomes: how much of her did my words actually pin down?

Usually, almost none of it. Take a typical avatar prompt:

"Beautiful blonde woman in her early thirties, blue eyes, natural makeup, wearing a cream blazer, studio lighting."

Everything in there is a category, not a person. Blonde is a range. Early thirties is a range. Blue eyes covers everything from pale grey to deep navy. Nothing in that sentence constrains the distance between her eyes, the width of her nose, the shape of her jaw, the height of her cheekbones — and those are precisely the things your audience uses to decide whether two images are the same woman.

Your prompt was never describing a person. It was describing a category, and getting a random member of it.

The four things that actually drift

When people say their avatar is inconsistent, they usually can't say what changed. It's worth being precise, because each one has a different fix.

1. Face geometry

The proportions: eye spacing, nose width and bridge shape, jaw and chin, cheekbone height, brow position. This is the one that makes her read as a different human being, and it's the one words are worst at controlling. No amount of adjectives will hold it.

2. Colouring

Skin tone, undertone, hair colour and hair level, eye colour. Drifts more subtly, but it's why your avatar can look tanned in one image and porcelain in the next. Often caused by lighting changes rather than the model itself.

3. Build and age read

Height, frame, how old she reads. AI has a strong pull towards a default: young, slim, symmetrical, mid-twenties. If your avatar is meant to be forty-five with a real body, the model will quietly drag her back towards that default every time you give it room to.

4. Her world

Wardrobe, setting, colour palette, the light she lives in. Not identity exactly, but it's a big part of why an account feels like one person. An avatar in a different room, different clothes and different light every post feels like a stock photo library, not a brand.

The first three are fixed with reference images. The fourth is fixed with a written spec you reuse. You need both.

The fix: build a reference pack once

A reference pack is the small set of assets you generate once, at the start, and then never regenerate. Everything you make afterwards is anchored to it.

It contains two halves.

The locked images

Three to five images of your character, generated in a single session, that you commit to permanently:

The mistake almost everyone makes here is generating a gorgeous reference. Dramatic side lighting, strong expression, moody background. Don't. Every bit of drama in your reference gets inherited by every image you make from it, and harsh shadow across a face actively hides the geometry the model needs to read.

Keep the reference boring. Flat light, plain background, neutral face, highest resolution your tool will give you. The beautiful images come later, generated from the boring one.

The written spec

A short block of text you paste into every prompt, word for word, forever. Not a description you rewrite each time — the same characters in the same order, treated like a name rather than a sentence.

It should pin down what images alone can't hold: her age read, her build, her hair at a specific level and length, her wardrobe register, her palette, her setting. Written once, saved somewhere you can copy it in two seconds.

The reason it must be identical every time is the same reason the drift happened in the first place. Rewording is regenerating. If you write "warm blonde" one week and "honey blonde" the next, you've asked for two different women.

Free tool

Build her reference pack in about ten minutes

The Avatar Quiz walks you through twelve questions — her face, her colouring, her build, her world — and hands you a complete prompt pack at the end: her reference images plus her first podcast and yap stills, written in the exact format that holds. It's the same method taught inside Human AI, and it's free.

How to actually use the pack

The pack only works if you change one habit: you never generate your character again. You generate scenes, and you attach your character to them.

In practice that means every prompt from now on has three parts, in this order:

  1. The reference image, attached as an image input — character reference, face reference, or whatever your tool calls it. Not described. Attached.
  2. Your written spec, pasted in unchanged.
  3. The scene — and this is the only part you write fresh. Where she is, what she's doing, the light, the framing, the mood.

Most tools now have some version of a character reference or identity feature, and they've got much better over the last year. The naming is different everywhere, the underlying idea is the same: give it a face to hold rather than a description to interpret.

If your tool has an identity feature where you train a character once from several photos and then call it by name — use it. That's the strongest version of this, and it's what your reference pack is for.

If you want to see what the scene half looks like written properly, the UND Bot prompt collection has nine of mine you can copy and paste straight into Higgsfield, OpenArt or anything similar. Swap in your own reference and spec, keep the scene structure.

Test it before you build on it

Before you make forty pieces of content, make four, deliberately different: one close-up, one wide, one in different lighting, one from a different angle. Put them side by side.

If a stranger would say those are four photos of the same woman, you're locked and you can start building. If they wouldn't, go back and fix the reference now — because a drifting avatar doesn't get better with volume, it just gets more expensive to fix.

What will still drift, and what to let go

Full consistency doesn't exist yet and anyone selling it to you is overselling. Some things are still genuinely hard:

Here's the thing worth holding onto, though. Your audience is not doing forensic comparison. They're scrolling. They recognise people the way we all do — hair, colouring, silhouette, the general shape of a face, the world around her.

You don't need pixel-identical. You need unmistakably her at scroll speed. That's a much lower bar, and a reference pack clears it comfortably.

Consistency isn't a tool you buy. It's a decision you make once and then stop renegotiating.

The short version

Get this right in your first week and the rest of it — the content, the voice, the actual business — stops being blocked by a face you can't find twice.

Now open

humanAI

Talk-to-camera content — reels, podcasts, story-driven video — without ever filming yourself.
Course and community in one place. The free training opens the moment you join.

Join Free

Free to join. Start today.

Keep reading

Why AI content feels fake — and the part no tool will fix

The technical tells are mostly solved. What's still fake is the message underneath.

How to start a faceless Instagram account in 2026, after the repost rule

Theme pages didn't die — reposting did. The two routes that still work.