TrendingThe thresholds you restructure your account around were never published.
PPC

Most creative tests end long before they could have concluded.

Two things break creative testing at once: the platform decides who sees what, and almost nobody runs the numbers on how much data a real answer needs. The arithmetic is not close. Detecting a 20% difference takes roughly 400 conversions per variant, and most tests are called at a tenth of that.

MSMikołaj Salecki, portrait
Editor-in-chief
Aug 14, 2026·7 min read
Two nearly identical plaster busts on a balance beam that is tipping decisively to one side, the fulcrum resting on a single small pebble, hairline measurement rules running behind them and one brand-blue disk above
The beam moved. That is not the same thing as one side being heavier.Illustration: Mediovsky · generated with AI
TL;DR
  • Detecting a 20% difference in conversion rate takes roughly 400 conversions per variant. A 10% difference takes roughly 1,600.
  • That holds at almost any baseline rate, which is why a low conversion rate is a traffic problem, not a testing excuse.
  • Sample size scales with the inverse square of the effect. Halving the difference you chase quadruples the cost.
  • Across 663 Facebook experiments, non-experimental estimates inflated true lifts of 29%, 18%, and 5% into 83%, 58%, and 24%. [2]
  • An earlier study across 15 experiments and 1.6 billion impressions found the same direction of failure. [1]
  • Google publishes ceilings, not recommendations: 3 responsive search ads and 50 active text ads per ad group. [4]
  • When you cannot afford to test, select instead: fewer, more different concepts, judged over a longer window.

A creative test finishes on a Thursday. Variant B has 61 conversions, variant A has 48. That is a quarter more, which is a big-looking gap, and somebody screenshots it. Variant B becomes the new control, the losing concept is retired, and the next test starts from a base that may never have been better in the first place.

Nothing in that story is unusual, and almost nothing in it is sound. Two separate problems arrived at once. The comparison was not randomized, and even if it had been, there was nowhere near enough data in it to see a gap that size through the noise.

Your comparison is not an experiment

The most useful research here is old enough to be settled and large enough to be uncomfortable. Gordon and colleagues, working with 15 US advertising experiments at Facebook covering 500 million user-experiment observations and 1.6 billion ad impressions, found that observational methods “often fail to produce the same effects as the randomized experiments, even after conditioning on extensive demographic and behavioral variables.” [1]

A later study pushed the sample to 663 large-scale Facebook experiments and tested the modern tooling people assume has solved this: double machine learning and stratified propensity score matching. The result is worth reading twice. Median true experimental effects at three funnel stages were 29%, 18%, and 5%. Double machine learning estimated them at 83%, 58%, and 24%. Propensity score matching was worse still. The authors concluded that despite large-scale experiments and rich user-level data, they were “unable to reliably estimate an ad campaign’s causal effect” without running the experiment. [2]

Funnel stage True effect, from the experiment Estimated by double machine learning
Upper 29% 83%
Middle 18% 58%
Lower 5% 24%

Read the direction, not just the size. These methods do not scatter randomly around the truth. They inflate it, consistently, in the direction that makes advertising look better. Two ads inside one ad set are a milder version of exactly this problem, because the system decides who sees which ad, and it makes that decision using its prediction of who will convert.

The number nobody runs

Set the bias aside and grant yourself a perfectly randomized test. You still need enough events, and the requirement is larger than intuition suggests.

The standard sample-size calculation for comparing two proportions takes the baseline rate, the difference you want to detect, a confidence level, and a power level. [3] Run it at 95% confidence and 80% power, and something clean falls out of the arithmetic [3].

Baseline conversion rate To detect a 10% relative difference To detect a 20% relative difference
1% 163,100 sessions per variant 42,700 sessions per variant
2% 80,700 21,100
3% 53,200 13,900
5% 31,200 8,200
Conversions per variant, any of the above about 1,600 about 400

These are computed rather than quoted, with the standard two-proportion power calculation at 95% confidence and 80% power [3], so you can reproduce every cell. The last row is the point. Multiply any traffic figure by its conversion rate and you land on the same place: about 400 conversions per variant to see a 20% difference, about 1,600 to see a 10% one.

The number to remember

About 400 conversions per variant buys you the right to notice a 20% difference. Not to be certain of it.

Computed: two-proportion power calculation, 95% confidence, 80% power

Two consequences follow immediately. First, a low conversion rate does not excuse you from the requirement, it just means the same 400 conversions cost more traffic. Second, halving the difference you want to detect roughly quadruples the data, because sample size scales with the inverse square of the effect. That single fact should reorganize your testing calendar: small differences are not slightly more expensive to resolve, they are catastrophically more expensive, and most of them are not worth resolving at all.

A steeply rising stack of thin paper sheets forming a curve that climbs off the top of the frame, a small plaster hand resting near the low flat end, hairline coordinate marks behind and one brand-blue sheet at the point the curve turns upward
Halve the difference you want to see, and the cost of seeing it roughly quadruples.Illustration: Mediovsky · generated with AI

What the platform does to the middle of your test

Even a test that will eventually reach 400 conversions per variant rarely gets to spend them evenly. Delivery systems allocate impressions toward whatever they currently predict will perform, which means the ad that looks better in its first few hundred impressions receives more of the next few thousand. The gap you observe at the end is partly a real difference and partly the compounding of an early accident.

This is not a defect to be complained about. It is the system doing the job it was built for, which is to spend your money on the best-performing thing it can currently identify. It just means that the output of that process is an allocation decision, not a measurement.

Two parallel channels of falling sand, the left one wide and lit brand blue, the right narrowed to a thread, both feeding plaster vessels of unequal fill on a dotted grid
The system pours more into whichever side looked better first. The final gap includes that decision.Illustration: Mediovsky · generated with AI

Three designs that survive

Pick by the question, not by the tool

Randomized experiment tools: the platform splits deliberately rather than letting delivery do it. Google lets you choose what share of the original campaign’s budget goes to the experiment and recommends 50% “to provide the best comparison between the original and experiment campaigns.” [5] Meta’s split-testing API “automates audience division, ensures no overlap between groups.” [6] Use these instead of duplicating an ad set by hand.

Geo or audience holdout: withhold advertising from randomly assigned regions and compare outcomes. The right tool for whether spend is producing incremental sales, and the wrong tool for choosing between two thumbnails. It shares its logic with incrementality testing.

Sequential with a clean baseline: run one concept, establish a stable base over a full purchase cycle, then switch. Weak against seasonality, but the only option when volume cannot support two arms at once, and honest if you say so.

When you cannot afford a test, stop testing

Most accounts cannot reach 400 conversions per variant on any reasonable timescale. That is a real constraint, and the useful response is to change the activity rather than to run underpowered tests and pretend.

The alternative is selection. Produce concepts that differ enough to matter, put them in, and let the system allocate. You lose the ability to say why one worked, which is a genuine loss. You keep the ability to end up with the better one, which is what the budget is for. The escape route hiding in the arithmetic is that big differences are cheap to detect: a difference of half again as good resolves in a small fraction of the data a tenth-better difference needs, so the way to make testing affordable is to test things that are actually different.

Volume has its own arithmetic, and it is the one that decides how many ads you should be running. Google publishes ceilings of 3 enabled responsive search ads and 50 active text or non-image ads per ad group. [4] Those are limits, not advice. The number that matters is your own: take the conversions the ad set produces in the window you are willing to wait, divide by the number of creatives, and if the answer is not a number you would act on, you are not running a test, you are running a lottery with extra steps.

  • Write the decision rule before launch: the metric, the smallest difference worth acting on, the conversions that implies, and the date you look.
  • Compute the sample size first. If the budget cannot buy it in a reasonable window, do not start the test.
  • Test concepts, not variations. Reserve the expensive machinery for differences large enough to be worth resolving.
  • Use the platform’s randomized experiment tool rather than duplicating ad sets, and never change anything else while it runs.
  • Keep one audience per test. Overlapping audiences across concurrent campaigns contaminate both.
  • Report the conversion count next to every result, permanently. A lift with no denominator is a rumor.
  • When volume will not support testing, say so out loud and switch to selection rather than quietly lowering the bar.

None of this makes creative less important. It makes creative more important, because creative is one of the few remaining inputs that genuinely moves outcomes while structure and settings converge across every advertiser in your category. What it makes less important is the ceremony around choosing creative, most of which produces confident conclusions from data that could not support them, in a direction that flatters whatever you already did.

Sources

  1. Marketing Science · A comparison of approaches to advertising measurement: evidence from big field experiments at FacebookGordon, Zettelmeyer, Bhargava, and Chapsky: 15 US experiments, 500 million user-experiment observations, 1.6 billion impressions, and the finding that observational methods often fail to reproduce randomized results
  2. arXiv · Close enough? A large-scale exploration of non-experimental approaches to advertising measurement663 large-scale Facebook experiments: median true effects of 29%, 18%, and 5% estimated as 83%, 58%, and 24% by double machine learning
  3. NIST/SEMATECH · Engineering Statistics Handbook: sample sizes required for proportionsthe standard method, with the confidence and power terms, behind the computed table in this piece
  4. Google Ads Help · About Google Ads account limits3 enabled responsive search ads, 50 active text and non-image ads, and 300 image or gallery ads per ad group
  5. Google Ads Help · Set up a custom experimentchoosing the share of the original campaign’s budget given to the experiment, and the recommendation to use 50%
  6. Meta for Developers · Marketing API: split testingautomated audience division with no overlap between test groups

Frequently asked questions

How many conversions do I need to call a creative winner?

Roughly 400 per variant to detect a 20% relative difference in conversion rate at 95% confidence and 80% power, and roughly 1,600 per variant to detect a 10% difference. The striking part is that this holds almost regardless of your baseline conversion rate, because a lower rate needs proportionally more traffic to produce the same number of conversions.

Why does halving the effect I want to detect cost so much more data?

Because sample size scales with the inverse square of the effect size. Halving the difference you want to detect multiplies the data you need by about four. That is why chasing small creative differences is usually not worth the budget it would take to resolve them.

Is a platform split test a real randomized experiment?

The platform’s own experiment tools do randomize, and they are far better than duplicating an ad set and splitting budget by hand. But the moment you compare ads inside one ad set, delivery is deciding who sees what, and that comparison is observational rather than randomized.

How badly do non-randomized comparisons mislead?

Badly, and in a consistent direction. Across 663 large-scale Facebook experiments, researchers found that even double machine learning turned median true effects of 29%, 18%, and 5% at different funnel stages into estimates of 83%, 58%, and 24%. Propensity score matching was worse. The authors concluded they could not reliably recover a campaign’s causal effect without the experiment.

How many creatives should I run per ad set or ad group?

Enough that each one can accumulate a judgeable number of conversions in the time you are willing to wait, which is arithmetic rather than doctrine. If an ad set produces 200 conversions a month and you run 10 ads, no ad will ever have enough behind it to judge. Google publishes hard ceilings, 3 enabled responsive search ads and 50 active text ads per ad group, but a ceiling is not a recommendation.

What should I do when I cannot reach the sample size?

Stop trying to test and start trying to select. Produce genuinely different concepts rather than variations, let the system allocate, and judge at the concept level over a longer window. You give up the ability to say why something won, and you keep the ability to keep winning.

Do bigger differences need less data?

Much less, and this is the practical escape route, because sample size scales with the inverse square of the effect. A concept-level difference of half again as good resolves in a fraction of the data a tenth-better difference needs. Testing two versions of a headline is usually unaffordable. Testing two genuinely different propositions is often quite cheap.

What is a holdout and when should I use one?

A holdout withholds advertising from a randomly chosen group or geography and compares outcomes against the exposed group. It is the right tool for measuring whether a channel or a strategy produces incremental sales, and the wrong tool for choosing between two thumbnails, because it answers a question about spend rather than about creative.

When should I write the decision rule?

Before launch, always. Fix the metric, the minimum difference worth acting on, the sample size that implies, and the date you will look. A rule written afterward is not a rule, it is a description of what you already decided to believe.

Found this useful?
MSMikołaj Salecki, portrait
Editor-in-chief

Mikołaj Salecki

Writes about media, tech, and AI business for people who actually run digital. Former agency lead. Skeptic of frameworks that read better than they perform.

More articles →