- Incrementality builds a counterfactual. A treated group sees your ads, a control group does not, and the gap is causal, which is exactly what last-click and multi-touch cannot produce. [4]
- In practitioner geo holdouts, multi-touch attribution overstates lower-funnel channels like brand search by 30% or more, crediting demand it only harvested. [5]
- Use experiments as the ground truth and calibrate a marketing mix model (MMM) to them, not the reverse. If a geo holdout shows near-zero lift, the test overrules the model. [6]
- To read a clean signal you must sacrifice real spend: a 15-30% or full on/off swing, held 2-4 weeks plus carryover. [5][6]
- Write the decision rule before the test, hold the line on the holdout, and pre-define the shocks that void it. [4]
Last-click loves your brand search. It is faithful, cheap per conversion, and almost entirely a passenger. The people who typed your name were coming anyway. Turn the campaign off in half your markets and watch what actually changes. That is incrementality testing, and in 2026 it is the only measurement that answers the question every budget meeting is really asking. Not which channel got the credit, but which spend caused the sale.
Why a test beats the dashboard and the model
Attribution follows paths. Incrementality builds a counterfactual. A test splits the world into a treated group that sees your ads and a control group that does not, matched on everything else, so any difference in conversions is caused by the ads and nothing else. Last-click and multi-touch attribution cannot make that split. They watch observed journeys and hand credit to whoever appeared last, which systematically flatters demand-harvesting channels and starves the demand-creating ones behind a healthy split of brand versus performance. [5]
A mix model gets closer, because it works at the portfolio level. But it is still a regression on historical spend, exposed to correlated budgets, seasonality, and its own priors. The 2026 consensus is not to pick one. Run experiments as the causal ground truth and use their lift to calibrate your mix model, not the other way around. When MMM says a channel returns 3x and a geo holdout shows near-zero incremental lift at current frequency, the test wins the budget argument. [6][7]
There are three designs, and they suit different questions.
| Design | How it randomizes | Best for |
|---|---|---|
| Geo experiment | by region, treatment markets vs control | TV, out-of-home (OOH), retail media, broad or brand search [7] |
| Audience holdout | withhold ads from a random slice of users | retargeting, customer relationship management (CRM), prospecting segments [3] |
| Conversion lift/ghost ads | platform-run control group inside a walled garden | always-on high-spend campaigns [6] |
Size it before you run it
Every design costs the same thing: a deliberate sacrifice. To read a clean lift you have to withhold real spend from real markets, and the first planning question is how much you can afford to hold dark. Cuts under a 15-30% swing rarely clear the noise, so a 5% nudge over 10 days is monitoring, not a test. [5]
Gross spend → media you hold dark
Enter the channel's monthly spend, how hard you cut it in the test markets, what share of the budget those markets carry, and how long the test runs. The output is the media you deliberately withhold to buy a clean read.
A rough planning model, not a forecast of lost sales. It assumes even pacing, 4.345 weeks a month, and that the withheld spend, not the incremental revenue it would have earned, is the cost you control. Size against what you can afford to hold dark, then confirm statistical power with a proper geo tool.
Running a geo holdout, end to end
Design
Pick one channel and one budget question. Write the decision rule first. If the full 95% confidence interval for incremental return on ad spend (iROAS) clears 1.0x, scale. Below it, cut. Straddling it, retest. [4]
Split
Gather at least 6 months of clean region-level history, then pair or cluster comparable markets into treatment and control. A powered setup often looks like 12 matched geo pairs, [5] though a large program like Wayfair's optimizes across all its geos rather than just a subset. [2][1]
Hold out
Cut spend hard in the treatment markets, 15-30% or all the way to zero, and keep control markets at business as usual. No back-door targeting, no safety campaigns. [5]
Run
Leave it alone for 2-4 weeks, plus about 2 weeks of carryover where purchase cycles are long. Haus TikTok geo experiments average 21 days to detect lift. [6][7]
Read lift
Estimate lift with confidence intervals using a causal tool (GeoLift, CausalImpact, synthetic controls), compute incremental cost per acquisition (CPA) and iROAS, then apply the rule you already wrote. [5]
The pitfalls that void a read
- Spillover. Ads bleed from test markets into control, especially TV and OOH. Use non-adjacent regions, tighter geo targeting, and drop border areas from the analysis. [1][5]
- A timid swing. Under a 15-30% change, a real effect can hide in the noise. Commit to a meaningful on/off or a 30%+ shift. [5]
- Too short. Ad effects lag, so a window under 2-4 weeks with no carryover buffer misses delayed conversions. [6]
- External shocks. Your own promo, a competitor discount, a holiday, or a news spike can invalidate the read. Pre-define the invalidating conditions and rerun. [5]
- Overlapping experiments. Two live tests on the same audience interfere. Keep a test calendar and one experiment per audience. [4][6]
- Moving the goalposts. Changing the hypothesis after seeing results is self-deception. The pre-registered rule is binding. [4]
- Trusting platform lift alone. Ghost-ad methods are a black box and can flatter the platform's own objective. Triangulate with geo tests and MMM. [7]
Incrementality is not another dashboard. It is a discipline that costs money to run and only pays off if the result is allowed to change the budget. Write the rule, hold the line on the holdout, and let the test overrule the click. The teams that keep losing are the ones still grading spend on the credit it claimed, not the sales it caused.
Sources
- Lifesight · Geo-Based Incrementality Testing: Marketer's Guide 20266 months of clean history and 80% power before a geo test
- Wayfair · How Wayfair Uses Geo Experiments to Measure Incrementality
- Sellforte · Best incrementality testing toolsaudience holdout designs and randomized slices
- StellaHeyStella · Incrementality Testing in 2026: Geo Experiments, Holdouts & iROASthe pre-registered decision rule and the 1.0x iROAS threshold
- Davies Meyer · Incrementality Testing 2026: Geo-Holdouts, Conversion Lift, and AI2-6 week duration, 15-30% spend cut, 12 geo pairs, carryover, 30% multi-touch attribution (MTA) overstatement
- LiftLab · Incrementality Testing and MMM Calibrationgeo holdout designs, four-to-six-week reads, and calibrating a marketing mix model to experiment results
- eMarketer · FAQ on incrementality: how to prove your ads actually work in 2026Haus TikTok 21-day average, 3-4 week minimums, which design suits which channel
Frequently asked questions
What is incrementality testing, and how is it different from attribution?
Incrementality testing is a controlled experiment that splits your audience into a treated group that sees your ads and a control group that does not, identical in every other way, so any gap in conversions is caused by the ads. Attribution does the opposite: it watches observed journeys and hands credit to whoever appeared last or along the path. That flatters demand-harvesting channels like brand search and cannot tell a 'would have converted anyway' user from one the ad actually moved.
Why run experiments if you already have a marketing mix model?
MMM works at the portfolio level, but it is a regression on historical spend and is vulnerable to correlated budgets, seasonality, and its own priors. The 2026 consensus is to run incrementality tests as the causal ground truth and feed their lift estimates back to calibrate the mix model, not the other way around. When MMM says a channel returns 3x and a geo holdout shows near-zero incremental lift at current frequency, the test wins the budget argument.
How big a spend change and how long a test do I actually need?
Small tweaks lie. Cuts under 15-30% rarely move a signal above noise, so a 5% nudge over 10 days is monitoring, not a test. Practitioner guidance points to a meaningful on/off or 30%+ swing held for 2-4 weeks, plus roughly 2 weeks of carryover observation where purchase cycles are long. Haus TikTok geo experiments average 21 days to detect lift.
What is a ghost-ad or conversion-lift test?
It is a platform-run experiment inside a walled garden. Meta, Google, or TikTok randomly assign users to a treatment group eligible to see your ad and a control group shown a placebo or nothing (the ghost ad), then report the incremental conversions between them. The granularity is good, but the method is a black box and can flatter the platform's own objective, so treat it as continuous measurement to be validated periodically with geo tests, not as ground truth.
Which test design should I use for my channel?
Geo experiments randomize by region and suit TV, out-of-home, retail media, and broad or brand search. Audience holdouts withhold ads from a random slice of users and fit retargeting, CRM, and prospecting segments. Platform-run conversion lift, or ghost ads, suits always-on high-spend campaigns, though its black-box method should be validated periodically with geo tests.
How much historical data do I need before running a geo holdout?
Plan on at least 6 months of clean region-level history so you can pair or cluster comparable markets into treatment and control. A powered setup often looks like 12 matched geo pairs, though larger programs may instead optimize across all their geos rather than a subset. Confirm statistical power with a proper geo tool before you commit spend.
What can invalidate an incrementality test?
Spillover from test markets into control is a common one, especially with TV and out-of-home, along with external shocks like your own promo, a competitor discount, a holiday, or a news spike. Overlapping experiments on the same audience interfere, and changing the hypothesis after seeing results is self-deception. Pre-define the conditions that void the read and rerun when they hit.
How do I decide whether to scale a channel from the result?
Write the decision rule before the test, not after. A common rule is that if the full 95% confidence interval for incremental ROAS clears 1.0x you scale, if it sits below you cut, and if it straddles you retest. The pre-registered rule is binding, which is what stops you from moving the goalposts once the numbers land.
Get the next issue by email.
One letter, once a week. Sharp coverage of media, tech, and AI business. No filler.




