TALECRAFTERS
← All posts
PRODUCTIONPOSTCOMPLIANCE

AI Dubbing and Lip Sync for Multi-Market Campaigns

Konstantinos Chatzimichail3 min read
PRODUCTION

Automated dubbing handles the words reliably and the performance poorly. Visual lip sync works well on close, front-on, well-lit faces and degrades quickly outside that. Plan the shot for localisation before you shoot or generate it, because almost nothing about this is fixable afterwards.

A note on what this is. A working summary written by a production studio, current at the date above, not legal advice. Regulation in this area is moving. Check the primary sources linked at the foot of the piece and take advice before relying on any of it commercially.

A campaign in eleven markets used to mean eleven voice sessions, eleven timing passes and a lot of scheduling. It now means one process that produces eleven tracks in an afternoon and a much harder question about which of them are good enough to run.

The four separable problems

PROBLEMHOW WELL IT AUTOMATESWHERE IT BREAKS
Translation of meaningWellIdiom, humour, claims that are regulated differently by market
Timing to pictureModeratelyLanguages expand and contract by up to a third; the cut does not
Voice and performancePoorlyEmphasis lands in the wrong place, and emotion reads as flat or as pantomime
Mouth shapesWell within a narrow envelopeAngle, distance, motion, facial hair, low light
What localisation actually consists of, and how well each part automates

Treating these as one problem is the standard mistake. They automate at completely different rates, and a workflow that outputs all four together gives you no way to accept two and redo two.

Where visual lip sync works

The envelope is narrower than the demos suggest. It works when the face is close to camera, roughly front-on, evenly lit, not moving much, and unobstructed. Outside that, in rough order of how quickly it degrades: profile angles, distance from camera, head movement, beards and moustaches, hard side light, and anything crossing the mouth.

The practical consequence is that lip sync is a shot design decision. If a campaign is going to be localised, the speaking shots should be composed for it — closer, flatter, steadier — and that decision has to be taken before anything is produced.

What has to be re-recorded

  • Anything where the emphasis carries the meaning. Converted performance places stress by rule and the rule is wrong often enough to matter.
  • Humour. Timing is the entire mechanism and it does not survive conversion.
  • Regulated claims, where the exact wording is legally load-bearing and a translation is a new claim requiring its own substantiation.
  • Anything with a named person’s voice, which is a consent question before it is a quality one.
  • The market that matters most. If eighty per cent of spend is in one country, that version is worth a human.

A workable review process

Automated output needs a review pass by somebody who speaks the language, and the useful version of that pass is structured rather than "does this sound alright".

  1. Read the translated script alone, without audio. Catch meaning errors before performance distracts from them.
  2. Listen without picture. Catch emphasis and pace problems on their own.
  3. Watch with picture at full speed. Catch sync.
  4. Watch the mouth at 50 per cent speed. Catch the sync errors that full speed hides.
  5. Check every claim against the local regulator’s position, because a translation is a new claim.

The disclosure position

Two things are happening and they are treated differently. Translating and re-voicing a real performer’s words is, in most markets, ordinary post-production. Synthesising a performer’s own voice in a language they do not speak is a synthetic performance, and it engages both consent and transparency obligations.

Since August 2026, EU transparency obligations under Article 50 apply to synthetic audio and video content that qualifies as a deepfake, which captures a voice clone of an identifiable person. Platform policies frequently reach further. The safe operating position is to treat a cloned voice as disclosable and to get the consent explicitly, per language, in writing.

The UK, EU and platform positions on one page, decided once per campaign rather than argued about at delivery.

THE DISCLOSURE CHECKLIST

Questions people actually ask

How good is AI dubbing in 2026?

Reliable for translating meaning, moderate for timing, and poor for performance. Emphasis lands by rule rather than by intention, which is wrong often enough to matter in anything where the stress carries the meaning.

When does AI lip sync work well?

On faces that are close to camera, roughly front-on, evenly lit, not moving much and unobstructed. It degrades with profile angles, distance, head movement, facial hair, hard side light and anything crossing the mouth — so it is a shot design decision, made before production.

What has to be re-recorded rather than converted?

Anything where emphasis carries meaning, humour, regulated claims where the wording is legally load-bearing, anything using a named person’s voice, and the single market carrying most of the spend.

Why does language length break a localised edit?

Because the same sentence can be up to a third longer in another language while the cut stays where it is. Leave handles on every speaking shot or plan for a different edit per market, and decide which at the edit stage rather than at delivery.

Does AI dubbing have to be disclosed?

Re-voicing a performer with another actor is ordinary post. Synthesising a performer’s own voice in a language they do not speak is a synthetic performance, which engages consent and, since August 2026, EU transparency obligations under Article 50 where it qualifies as a deepfake. Treat a cloned voice as disclosable and get consent per language in writing.

WRITTEN BY

Konstantinos Chatzimichail
FOUNDER AND CREATIVE DIRECTOR, TALECRAFTERS

Founder of TaleCrafters. Writes the pipelines the studio works to, directs the films that come out of them, and publishes both.

More from Konstantinos

TERMS USED HERE

TAKE THE TOOL WITH YOU

READ NEXT