TALECRAFTERS
← All posts
SYSTEMSPERFORMANCESTRATEGY

Creative Testing at Volume: How Many Variants Is Enough?

Konstantinos Chatzimichail3 min read
SYSTEMS

Vary one axis at a time and you learn something reusable. Vary four and you get a winner you cannot explain and cannot repeat. The right volume is however many variants it takes to isolate one axis, which is usually between five and eight, not ninety.

The pitch for generative creative is volume: test a hundred concepts a week instead of four. The pitch is broadly true and it hides a problem, which is that a hundred variants differing on everything at once produce a ranked list and no explanation. Next month you start again from nothing, because you learned which asset won rather than why.

The difference between a test and a tournament

A tournament ranks options. It is useful and it is not knowledge: the winner tells you what to run this month and nothing about what to make next month. A test isolates a variable and produces a claim you can carry forward, which compounds.

Both have a place. The mistake is running a tournament while believing you ran a test, which is what happens whenever a matrix varies more than one thing.

Designing a matrix that produces knowledge

ROUNDWHAT VARIESWHAT IS HELDWHAT YOU LEARN
1Hook (6 variants, one per mechanism)Everything else identicalWhich mechanism this audience responds to
2Opening visual (4)Winning hook, everything else identicalWhether the cliff is visual or verbal
3Body structure (3)Winning hook and visualWhether the middle slope is a structure problem
4Call to action (3)Everything aboveWhether the ask or the argument is the limit
5Presenter or register (4)Everything aboveHow much of performance is who is delivering it
One axis at a time

Twenty assets across five rounds, and at the end you hold four reusable claims about your audience rather than one winning file. The next campaign starts from those claims, which is the only mechanism by which testing gets cheaper over time.

Where the returns stop

  • Statistical: past a point, additional variants divide the same budget and nothing reaches significance. More variants means slower learning, not faster.
  • Practical: someone has to watch all of them. Ninety assets is over an hour of review before anything is watched twice.
  • Creative: variants generated to fill a matrix rather than to test an idea are noise, and they dilute the average performance of the batch.
  • Platform: several delivery systems concentrate spend on early leaders, so a large matrix gets pruned by the algorithm before your test finishes. You measured the platform, not the creative.

The last one catches sophisticated teams. If the platform allocates on early signal, a fifty-variant test is a five-variant test with forty-five assets that never got a chance, and the five were chosen by delivery rather than by design.

What to hold constant, and how

A test is only valid if everything except the tested axis is genuinely identical, which is harder in generative production than it sounds, because two generations from the same prompt are not the same asset.

  1. Generate the invariant portion once and reuse the file. Do not regenerate it per variant.
  2. For hook tests, change only the audio and the first two seconds of picture. Everything after is one master.
  3. Keep the same presenter, same wardrobe, same light, same grade across a round. Vary those in their own round.
  4. Name files so the axis is in the name. Reading a results table where the filenames do not encode the variable is how findings get lost.
  5. Log which round each asset belongs to, so a winner from round one can be traced when it stops winning in month four.

The metric to test on

For hook rounds, the first-three-seconds retention, not the conversion. A hook that wins on conversion may have won for reasons downstream of the hook, and you will have attributed it wrongly.

For structure rounds, the middle slope of the retention curve. For call-to-action rounds, conversion. Matching the metric to the axis is the step that makes the round interpretable, and it is skipped more often than any other.

Which feature of the graph corresponds to which part of the piece, so a test round measures what it thinks it is measuring.

READING A RETENTION CURVE

Questions people actually ask

How many ad variants should you test at once?

Enough variants of a single axis to distinguish them, which for most paid social is five to eight. Beyond that, additional variants divide the same budget, nothing reaches significance, and learning gets slower rather than faster.

Why is testing ninety variants a bad idea?

Because if they differ on several axes at once you get a winner you cannot explain and cannot repeat. You also hit four ceilings: statistical significance, review time, creative dilution, and platforms that concentrate spend on early leaders and prune your test before it finishes.

How do you design a creative test that produces reusable knowledge?

Vary one axis per round and hold everything else identical: hooks first, then opening visual, then body structure, then call to action, then presenter or register. Twenty assets across five rounds yields four reusable claims about your audience rather than one winning file.

What metric should each test round use?

Match the metric to the axis. Hook rounds are judged on first-three-seconds retention, structure rounds on the middle slope of the retention curve, and call-to-action rounds on conversion. A hook judged on conversion may have won for reasons downstream of the hook.

How do you hold variables constant in generative testing?

Generate the invariant portion once and reuse the file rather than regenerating it per variant — two generations from the same prompt are not the same asset. For hook tests, change only the audio and the first two seconds of picture over one master.

WRITTEN BY

Konstantinos Chatzimichail
FOUNDER AND CREATIVE DIRECTOR, TALECRAFTERS

Founder of TaleCrafters. Writes the pipelines the studio works to, directs the films that come out of them, and publishes both.

More from Konstantinos

TERMS USED HERE

TAKE THE TOOL WITH YOU

READ NEXT