
From the field
The holdout test
The incrementality method that measures what advertising actually causes by withholding it from a comparable region or audience and counting the conversions that happen anyway, because platform-reported conversions include sales that would have occurred regardless.
Platform attribution answers the question "who saw an ad and later converted?", which is not the question the invoice asks. A retargeting campaign that shows ads to people already on the way to buying will claim those purchases and report heroic returns. The only clean answer to "what did the money cause?" is to stop spending somewhere comparable and watch what happens.
The holdout test
Split
Pre-period
Run
Count
Gap against noise
Use it when
- Brand search campaigns claim conversions the brand would get for free
- A retargeting campaign reports suspiciously beautiful ROAS
- Deciding whether a channel deserves its budget at renewal time
Do not use it when
- Conversion volumes are too small to detect a difference. A rule of thumb, not a law: under roughly 100 conversions per group over the test, chance alone moves the gap between the groups by about 28 either way, because noise grows with the square root of the count, so only very large lifts show up
- The business cannot tolerate a temporary dip in one region
The steps
- Pick the split. Two comparable city districts, or a random 20% of the audience excluded from the campaign. Comparable matters more than clever.
- Check the pre-period. Count both groups' conversions for the same number of weeks before the test, with ads running in both. If district A normally books 5 more than district B, the lift is whatever the test adds on top of those 5, not the whole gap.
- Run long enough to cover a buying cycle. Four to eight weeks for most local services; shorter tests measure noise.
- Count conversions per group from the business's records, not the ad platform's: bookings, calls, invoices.
- Compute the lift, then check it against the noise. Incremental conversions = exposed group minus holdout, less any pre-period gap. Divided by the conversions the platform claimed, that is the share of the platform's claim that is real. Divided by the exposed group's conversions, it is the share of that group's sales the advertising caused. Then the noise check: counts like these vary by about their square root through chance alone, so a gap smaller than about 2 × √(exposed + holdout) cannot show a lift. It can still show that a large claim is impossible.
Worked example
A dental clinic's brand-keyword search campaign, tested by pausing it in one of two comparable districts for six weeks. Illustrative numbers:
| Group | Ad spend, test weeks | Bookings, 6 weeks before (ads on in both) | Bookings, 6 test weeks | Platform-claimed |
|---|---|---|---|---|
| District A (ads on) | €900 | 61 | 64 | 41 |
| District B (ads off) | €0 | 61 | 58 | n/a |
| Difference | 0 | 6 |
Claimed lift against what the test allows
The numbers
| Group | patients |
|---|---|
| Measured gap, district A minus B | 6 patients |
| Highest lift the noise allows6 + 2 × √(64 + 58), rounded | 28 patients |
| Conversions the platform claimed | 41 patients |
The districts ran level before the test, so the whole gap of 6 counts as lift. Against the platform's claim that is 6 ÷ 41, about 15% of the patients the campaign took credit for. Against district A's bookings it is 6 ÷ 64: about 9% of them were caused by the ads.
Those percentages are not precise, and the example should say so. At about 60 bookings a group, chance moves each count by roughly 8 and the gap between them by roughly 11, so a gap of 6 sits well inside the noise. The true lift could be zero, and it could be twenty. This clinic is under the volume in the "Do not use it when" list above, and the test still earned its six weeks, because it answered the question that mattered: the platform's 41 would need a gap of 41, and nothing the test allows comes close. The likeliest reading is that people searching the clinic's name find its organic listing directly below the ad either way. The brand budget moved to generic keywords the following quarter, where there is no organic listing to fall back on.
Why it works
Holdouts inherit their logic from controlled experiments. In a random split the advertising is the only systematic difference between the groups, so the difference in outcomes is the advertising's effect plus chance. A district split gets close to that once the pre-period shows the districts ran level, and the noise check puts a size on the chance.
The answer varies by account, which is the reason to test your own. Google's analysis of more than 400 pause studies on search accounts found that on average 89% of paid clicks were not replaced by organic clicks when the ads stopped, across all keyword types. A field experiment published in Econometrica in 2015, run at a large online marketplace, found its brand-keyword ads had no measurable short-term benefit. Both results are true of the accounts they measured, and neither is true of yours by default.
Ad platforms do run lift studies of their own: Google Ads offers Conversion Lift, with holdouts by user or by region, and where it is available it is worth using. A holdout counted from the business's own records is the check that does not depend on the platform grading its own work, and it is the one framework on this site built to show that a running campaign should be switched off.
AARRR shows which stage of the funnel is leaking; the holdout test shows whether the money spent on a stage caused anything. A campaign that survives it still has to clear the spend floor to keep learning.
Sources
- Blake, T., Nosko, C., & Tadelis, S. (2015). Consumer heterogeneity and paid search effectiveness: a large-scale field experiment. Econometrica, 83(1), 155-174.
- Chan, D., Yuan, Y., Koehler, J., & Kumar, D. (2011). Incremental clicks: the impact of search advertising. Journal of Advertising Research, 51(4), 643-647.
- Google Ads Help. About Conversion Lift.
Who this is for
The same method, read three ways.
01
Running it
Count bookings per district from the clinic's own diary, for six weeks before and six weeks during. Divide the gap by what the platform claimed, and check it against 2 × √(A + B) before believing a single percentage.
02
Teaching it
A controlled experiment in its smallest useful form, and the clearest demonstration that attribution and causation answer different questions. The noise check is the lesson most versions skip: a small test can refute a big claim long before it can confirm a small one.
03
Renewing a budget
Before renewing a channel, ask what happened in a district where it was switched off. If nobody has checked, the platform's conversion count is the platform grading its own work.