avatar-lab

← Blog

We trained a character into a platform. It kept 0% of its markers.

If you ask anyone — including us, four days before the run — how to get the strongest character lock on a generation platform, the answer is the same: train the character in. Fine-tune it, make it a native object, pay the training cost and collect the consistency. That is what the platforms recommend, and it is what intuition says.

So we measured it. One stylised character with declared, checkable identity markers — hair colour, a face mark, a specific collar palette. Two conditions on the same platform: the trained-in version, and plain reference images with the written specification. Same scenes, same judge.

The result

The trained version kept 0% of the declared markers. Hair colour wrong, face mark gone, collar palette replaced. The plain-reference condition kept 100%. We assumed a mistake and re-ran it on a second platform: same direction, same conclusion.

The mechanism, as far as the runs let us see it: training compresses a character into the model’s own style space. For a stylised character, the nearest point in that space is a different character — smoothed, prettified, regressed toward what the model already likes to draw. Reference images at generation time do not get compressed; they sit next to the prompt as evidence the model has to reconcile on every call.

Why a similarity score hides this

The trained condition scored well on whole-image similarity. It reproduced composition, lighting and rendering style beautifully. It just wasn’t the same character — and a single similarity number cannot tell those two things apart. This is why every result we publish is reported on axes that can disagree: face identity, whole-subject similarity, prompt adherence, and how rigid the output set became. The platform with the highest single score in our main run had the lowest identity — it got its score by repeating the reference framing and ignoring three of the five scenes.

What changed because of this

Every adapter pack this registry issues is built on references-plus-specification, not on training, and the per-platform wording is written from measured behaviour rather than platform documentation. Where training is the right call — it sometimes is, for photoreal humans on specific platforms — the pack says so explicitly, with the run that justifies it.

Two of the four headline findings in our benchmark contradicted what we ourselves believed when we designed it. That is the strongest argument we know for measuring instead of assuming — and for publishing the misses along with the hits.

Read the full measurement note →