A campaign in eleven markets used to mean eleven voice sessions, eleven timing passes and a lot of scheduling. It now means one process that produces eleven tracks in an afternoon and a much harder question about which of them are good enough to run.
The four separable problems
| PROBLEM | HOW WELL IT AUTOMATES | WHERE IT BREAKS |
|---|---|---|
| Translation of meaning | Well | Idiom, humour, claims that are regulated differently by market |
| Timing to picture | Moderately | Languages expand and contract by up to a third; the cut does not |
| Voice and performance | Poorly | Emphasis lands in the wrong place, and emotion reads as flat or as pantomime |
| Mouth shapes | Well within a narrow envelope | Angle, distance, motion, facial hair, low light |
Treating these as one problem is the standard mistake. They automate at completely different rates, and a workflow that outputs all four together gives you no way to accept two and redo two.
Where visual lip sync works
The envelope is narrower than the demos suggest. It works when the face is close to camera, roughly front-on, evenly lit, not moving much, and unobstructed. Outside that, in rough order of how quickly it degrades: profile angles, distance from camera, head movement, beards and moustaches, hard side light, and anything crossing the mouth.
The practical consequence is that lip sync is a shot design decision. If a campaign is going to be localised, the speaking shots should be composed for it — closer, flatter, steadier — and that decision has to be taken before anything is produced.
What has to be re-recorded
- Anything where the emphasis carries the meaning. Converted performance places stress by rule and the rule is wrong often enough to matter.
- Humour. Timing is the entire mechanism and it does not survive conversion.
- Regulated claims, where the exact wording is legally load-bearing and a translation is a new claim requiring its own substantiation.
- Anything with a named person’s voice, which is a consent question before it is a quality one.
- The market that matters most. If eighty per cent of spend is in one country, that version is worth a human.
A workable review process
Automated output needs a review pass by somebody who speaks the language, and the useful version of that pass is structured rather than "does this sound alright".
- Read the translated script alone, without audio. Catch meaning errors before performance distracts from them.
- Listen without picture. Catch emphasis and pace problems on their own.
- Watch with picture at full speed. Catch sync.
- Watch the mouth at 50 per cent speed. Catch the sync errors that full speed hides.
- Check every claim against the local regulator’s position, because a translation is a new claim.
The disclosure position
Two things are happening and they are treated differently. Translating and re-voicing a real performer’s words is, in most markets, ordinary post-production. Synthesising a performer’s own voice in a language they do not speak is a synthetic performance, and it engages both consent and transparency obligations.
Since August 2026, EU transparency obligations under Article 50 apply to synthetic audio and video content that qualifies as a deepfake, which captures a voice clone of an identifiable person. Platform policies frequently reach further. The safe operating position is to treat a cloned voice as disclosable and to get the consent explicitly, per language, in writing.
The UK, EU and platform positions on one page, decided once per campaign rather than argued about at delivery.
THE DISCLOSURE CHECKLIST →