Audiences forgive a great deal in a generated face and forgive almost nothing in the mouth, because lip reading is partly involuntary. Sync that is slightly wrong reads as dubbing, and dubbing reads as untrustworthy.
The technical failure modes are consistent: plosives arriving late, the jaw moving without the lips shaping, and a mouth that keeps moving fractionally after the audio stops. All three are visible at normal speed once you know to look.
The consent position is stricter here than anywhere else in synthetic production. Putting words a person did not say into a recognisable mouth is the definition of the thing regulation is aimed at, and it needs an explicit release naming synthetic dialogue.
Why does generated lip sync look like dubbing?
Usually timing: plosives landing late, or the mouth continuing fractionally after the audio ends. Lip reading is partly involuntary, so small errors register as wrongness rather than as detail.
What consent does synthetic dialogue need?
An explicit release covering synthetic dialogue specifically, separate from any release signed for the original footage. Words a person did not say in a recognisable mouth is exactly what transparency regulation targets.
