Every fortnight a new model takes the top of a benchmark and every fortnight somebody asks whether we are switching. The honest answer is that we mostly are not, and the reason is not loyalty. It is that the benchmark measures the thing that matters least to a production schedule.
Rankings measure output quality on prompts written to show output quality. What decides whether a model can carry a campaign is a different list, and almost none of it appears in a comparison table.
The six things that actually decide it
| CRITERION | THE QUESTION | WHY IT OUTRANKS RAW QUALITY |
|---|---|---|
| Reference conditioning | Can it take an image and hold identity from it? | Without it, every shot after the first is a new person. This alone rules out otherwise excellent models. |
| Clip length and extension | How long before drift, and can it chain? | Sets your shot-length discipline and therefore your shot count and therefore your budget. |
| First and last frame control | Can you pin both ends? | The difference between directing a shot and receiving one. |
| Licence and provenance | Is commercial use unambiguous, and what does the output carry? | A client legal team will ask. "The terms of service imply it" is not an answer. |
| Throughput and queue | How many concurrent jobs, and what happens at 6pm on a Thursday? | A model that is twenty per cent better and four times slower loses on every deadline. |
| Failure shape | How does it go wrong? | Predictable failure is cheaper than occasionally spectacular output, because you can gate for it. |
Notice that quality is not on the list. That is not because it does not matter; it is because above a threshold every serious model clears it, and below that threshold none of the other criteria save you. Quality is a gate, not a ranking.
The failure shape argument
This is the one people find counter-intuitive. Given two models, one that produces a brilliant frame seven times in ten and something unusable three times in ten, and one that produces a good frame nine times in ten and a slightly soft frame once, the second is cheaper to produce on even if the first has the higher ceiling.
The reason is that a pipeline is a set of gates, and gates catch predictable failures cheaply. A model that fails the same way every time can have a check written for it. A model that fails differently each time needs a human looking at everything, and human attention is the most expensive thing in the building.
The number that converts a per-second price into a production budget, and the one most studios do not track.
ACCEPTANCE RATE, DEFINED →The thing that makes switching expensive
It is never the prompts. Prompt vocabulary transfers reasonably well between models because they were all trained on film and respond to film language. What does not transfer is everything around the prompt.
- Conditioning method. A pipeline built on image references does not move to a model that only takes text without rebuilding the identity strategy from scratch.
- Aspect and duration assumptions. Shot lists are written against a usable clip length. Change the length and the shot count changes and the edit changes.
- Your negatives. Negative prompts are written from artefacts you have personally seen, so they are model-specific by construction. A new model means a new empty list and a fortnight of rebuilding it.
- Acceptance-rate history. The moment you switch, every budget you have quoted from is describing a different machine.
- The lock file. Palette and light behaviour tuned to one model’s response are a starting point, not a spec, on the next one.
What to do about deprecation
Models get withdrawn, and 2026 has already provided the case study: a widely adopted model was deprecated and its consumer product closed, which stranded workflows that had been built around its specific behaviour. Anyone who had treated it as infrastructure discovered it was a product.
The defence is not to predict which model survives. It is to keep the things that would have to be rebuilt outside the model:
- Keep the lock file model-agnostic. Write palette, light direction and material behaviour in plain production language, not in phrasing tuned to one model’s quirks.
- Keep identity in an asset, not in a checkpoint. A set of reference frames survives a model change. A fine-tune does not.
- Keep the shot list in beats, not in generations. A beat can be produced by anything; a generation ID cannot be reproduced anywhere else.
- Log acceptance rate per model, not per campaign, so that when you do move you can quote from the new machine rather than the old one.
- Never let a client deliverable depend on a model still existing. Deliver the frames, not the recipe.
The one benchmark worth running yourself
Take the three hardest shots from a brief you actually ran. Not showreel shots: the ones with legible packaging type, a recurring face, and a hand doing something. Run twenty attempts of each on any model you are considering, and count how many you would have sent.
That number is your acceptance rate on that model for that class of work, and it is worth more than every comparison article published this year, including this one. It takes an afternoon and it is the only figure that describes your work rather than somebody else’s prompt.
The scaffolds we test a new model against, including the shot-list lock block that has to survive a switch.
THE PROMPTING LIBRARY →