TALECRAFTERS
← All posts
PRODUCTIONETHICSCRAFT

Voice Cloning for Brand Narration: What It Costs, What It Owes

Konstantinos Chatzimichail3 min read
PRODUCTION

Synthetic narration is now indistinguishable from a human read on neutral, informational copy. It still fails on emphasis that carries meaning, on humour, and on anything requiring a decision mid-sentence. Where it succeeds, the constraint is no longer quality — it is consent, disclosure and direction.

A note on what this is. A working summary written by a production studio, current at the date above, not legal advice. Regulation in this area is moving. Check the primary sources linked at the foot of the piece and take advice before relying on any of it commercially.

Voice is the modality that arrived first and got least attention, probably because nobody makes a showreel out of a voiceover. It is also the one where the quality argument is effectively over for a large category of work, which moves the interesting questions elsewhere.

Where it passes and where it does not

COPY TYPEPASSES?WHY
Neutral informational VOYesEven pace, no stress decisions, no subtext to carry
Corporate narrationYesThe register is already flat, which is a low bar to clear
E-learning and instructionYesClarity is the whole requirement
Emotional readRarelyEmotion is produced by decisions mid-sentence, which is exactly the gap
ComedyNoTiming is the mechanism, and it does not survive synthesis
Character performanceNoListening, hesitation and self-interruption are the craft
Anything with a name on itDepends entirely on consentA rights question before a quality one
Synthetic narration by copy type

The six controls that make it sound human

Most bad synthetic reads are undirected rather than incapable. The controls that matter, in order of effect:

  1. Breath placement. Insert breaths where a person would take them, which is before a clause they are about to emphasise, not at the ends of sentences. This single change does more than any other.
  2. Pace variation. A human read speeds up through the familiar and slows through the important. A uniform pace is the loudest tell there is.
  3. Emphasis by rewriting, not by markup. Move the word you want stressed to the end of the clause. Prosody follows structure more reliably than it follows tags.
  4. Sentence length variation in the script. Synthetic reads expose monotonous sentence rhythm far more than human ones, because a human unconsciously varies against it.
  5. One imperfection. A slightly early breath, a very small stumble, one word taken at a different pace. One, not three.
  6. Room. A completely clean voice in a completely silent mix is not a recording of anything. Put it in a space.

The consent structure

A cloned voice needs a release that a standard performance contract does not contain, because the standard contract licenses a recording and this licenses a capability.

  • Explicit grant to create a model from the supplied recordings, named as such.
  • Scope: which brands, which product categories, which media. A voice licensed for internal training that appears in an advert is a breach nobody documented.
  • Term, with an end date rather than "in perpetuity", and a defined disposal obligation for the model at the end of it.
  • Territory, which matters because personality and likeness rights differ substantially between jurisdictions.
  • Exclusions: categories the performer will not be used for. Political, gambling, alcohol, health claims — whatever they choose.
  • Withdrawal: a mechanism, a notice period, and what happens to assets already in market.
  • Rate structure for use, not only for the session. A session fee for something that runs forever is the arrangement performers are right to resist.

Disclosure

A wholly synthetic voice that is not presented as a specific person is, in most markets, not a disclosure issue in itself, though platform policies vary and are frequently stricter than law.

A cloned voice of an identifiable person is a different matter. Since August 2026, EU transparency obligations under Article 50 apply to synthetic audio that qualifies as a deepfake, and an identifiable person’s cloned voice sits squarely in that. The working position is to disclose, to hold the consent in writing, and to decide both at brief stage.

What we will not do

Clone a voice without a signed release from the person, including for a test, including internally, including when the recordings are publicly available. The availability of material has never been the same thing as permission, and the fact that it is now technically trivial is an argument for the rule rather than against it.

The definition, the consent structure it requires, and the related terms.

VOICE CLONING, DEFINED

Questions people actually ask

Is AI voiceover good enough for brand work?

For neutral informational narration, corporate voice and e-learning, yes — the quality argument is effectively over. It still fails on emotional reads, comedy and character performance, because all three depend on decisions made mid-sentence.

How do you make a synthetic voice sound human?

Place breaths before clauses about to be emphasised rather than at sentence ends, vary the pace, create emphasis by moving the stressed word to the end of a clause rather than by markup, vary sentence length in the script, add exactly one imperfection, and put the voice in a room rather than in silence.

What consent does voice cloning require?

An explicit grant to build a model from the recordings — not just to use them — plus scope by brand and media, a term with an end date and disposal obligation, territory, category exclusions, a withdrawal mechanism, and a rate structure for use rather than only for the session.

Does a cloned voice have to be disclosed?

A cloned voice of an identifiable person generally yes: since August 2026 EU transparency obligations under Article 50 apply to synthetic audio qualifying as a deepfake. A wholly synthetic voice not presented as a specific person is usually a platform-policy question rather than a legal one, and platform policy is often stricter.

Why do synthetic reads sound flat even when the voice is good?

Usually because the copy was written for the eye. A uniform pace is the loudest tell there is, and synthetic reads expose monotonous sentence rhythm far more than human ones. Read the script aloud, mark where you naturally breathe and slow, and rewrite so those fall on clause boundaries.

WRITTEN BY

Konstantinos Chatzimichail
FOUNDER AND CREATIVE DIRECTOR, TALECRAFTERS

Founder of TaleCrafters. Writes the pipelines the studio works to, directs the films that come out of them, and publishes both.

More from Konstantinos

TERMS USED HERE

TAKE THE TOOL WITH YOU

READ NEXT