Voice is the modality that arrived first and got least attention, probably because nobody makes a showreel out of a voiceover. It is also the one where the quality argument is effectively over for a large category of work, which moves the interesting questions elsewhere.
Where it passes and where it does not
| COPY TYPE | PASSES? | WHY |
|---|---|---|
| Neutral informational VO | Yes | Even pace, no stress decisions, no subtext to carry |
| Corporate narration | Yes | The register is already flat, which is a low bar to clear |
| E-learning and instruction | Yes | Clarity is the whole requirement |
| Emotional read | Rarely | Emotion is produced by decisions mid-sentence, which is exactly the gap |
| Comedy | No | Timing is the mechanism, and it does not survive synthesis |
| Character performance | No | Listening, hesitation and self-interruption are the craft |
| Anything with a name on it | Depends entirely on consent | A rights question before a quality one |
The six controls that make it sound human
Most bad synthetic reads are undirected rather than incapable. The controls that matter, in order of effect:
- Breath placement. Insert breaths where a person would take them, which is before a clause they are about to emphasise, not at the ends of sentences. This single change does more than any other.
- Pace variation. A human read speeds up through the familiar and slows through the important. A uniform pace is the loudest tell there is.
- Emphasis by rewriting, not by markup. Move the word you want stressed to the end of the clause. Prosody follows structure more reliably than it follows tags.
- Sentence length variation in the script. Synthetic reads expose monotonous sentence rhythm far more than human ones, because a human unconsciously varies against it.
- One imperfection. A slightly early breath, a very small stumble, one word taken at a different pace. One, not three.
- Room. A completely clean voice in a completely silent mix is not a recording of anything. Put it in a space.
The consent structure
A cloned voice needs a release that a standard performance contract does not contain, because the standard contract licenses a recording and this licenses a capability.
- Explicit grant to create a model from the supplied recordings, named as such.
- Scope: which brands, which product categories, which media. A voice licensed for internal training that appears in an advert is a breach nobody documented.
- Term, with an end date rather than "in perpetuity", and a defined disposal obligation for the model at the end of it.
- Territory, which matters because personality and likeness rights differ substantially between jurisdictions.
- Exclusions: categories the performer will not be used for. Political, gambling, alcohol, health claims — whatever they choose.
- Withdrawal: a mechanism, a notice period, and what happens to assets already in market.
- Rate structure for use, not only for the session. A session fee for something that runs forever is the arrangement performers are right to resist.
Disclosure
A wholly synthetic voice that is not presented as a specific person is, in most markets, not a disclosure issue in itself, though platform policies vary and are frequently stricter than law.
A cloned voice of an identifiable person is a different matter. Since August 2026, EU transparency obligations under Article 50 apply to synthetic audio that qualifies as a deepfake, and an identifiable person’s cloned voice sits squarely in that. The working position is to disclose, to hold the consent in writing, and to decide both at brief stage.
What we will not do
Clone a voice without a signed release from the person, including for a test, including internally, including when the recordings are publicly available. The availability of material has never been the same thing as permission, and the fact that it is now technically trivial is an argument for the rule rather than against it.
The definition, the consent structure it requires, and the related terms.
VOICE CLONING, DEFINED →