A recurring presenter is the hardest thing a brand can ask a generative pipeline for, because it is the case where the audience has the strongest prior. People are extraordinarily good at faces. A product that is five per cent different between shots is unnoticed; a face that is five per cent different is a different person.
The character brief
Everything starts with a paragraph you will paste unchanged for months. It has to specify things that do not change with mood or lighting.
- Structure: face shape, bone, proportion, the asymmetry. Measurable things.
- Two or three fixed marks: a scar, a mole, a gap, a crooked tooth. These carry more identity than any amount of description.
- Hair: length, texture, parting, and how it behaves when disturbed.
- Wardrobe: exact, including fastenings and wear. A missing button is an identity anchor.
- What is deliberately unspecified: expression and pose, which have to vary.
What must not be in it: personality adjectives. "Warm, approachable, confident" describes nobody, which is why every generation from it returns a different person who happens to be warm, approachable and confident.
The reference set
One frame is not enough, and twenty is worse than five. What you want is a small fixed set covering the angles the campaign actually needs — front, three-quarter, profile, and one at the shot size you will use most — approved once and never quietly extended.
The discipline that matters is that the reference is used for identity only. If it is also carrying the pose, the light and the grade, every output will be a near-copy of the reference frame, and you will have one shot rendered fifty times.
The drift check
Two checks, both quick, both non-negotiable at volume.
The first is the overlay: take the new frame and the original approved reference, align them on the eyes, and flick between the two at full size. Structural differences that are invisible side by side are obvious in a flick, because the eye is comparing positions rather than impressions.
The second is the strip: every approved frame of that presenter, in a row, at thumbnail size. Drift is a gradient and gradients are only visible over distance. A face that has moved two per cent per asset over thirty assets is a different person at the end, and nobody spots it one asset at a time.
When to train instead
| REFERENCE CONDITIONING | TRAINED IDENTITY | |
|---|---|---|
| Setup cost | Almost none | Significant: material, time, iteration |
| Consistency ceiling | Good, degrades with unusual angles | High, holds across poses |
| Flexibility | Full | Constrained to what it was trained on |
| Portability | Survives a model change | Does not |
| Worth it above | — | Roughly twenty to thirty assets, or a presenter recurring for over a quarter |
| Rights burden | Consent for the source material | Consent plus an explicit derivative-training clause |
The threshold is not a rule, it is a break-even. Below it, training costs more than it saves. Above it, the consistency is worth the setup, and the additional benefit is that a trained identity can hold a pose that references struggle with.
The rights half, which is not optional
If the presenter derives from a real person, the release has to grant the right to train on the supplied material explicitly. Photography-era releases almost never do, because the concept did not exist when they were drafted. Scope, term, territory, withdrawal and end-of-term disposal all belong in it.
If the presenter is wholly synthetic, the consent question falls away and the disclosure question does not. A synthetic person presented as a real one engages transparency obligations regardless of whether anybody’s likeness was used.
A likeness and voice release drafted for generative production, including the derivative-training clause photography releases do not contain.
THE CONSENT TEMPLATE →