Incrementality Testing: The Method, Step by Step
Incrementality testing measures whether ad spend caused a sale that wouldn't have happened anyway, by comparing an exposed group against a genuinely unexposed or matched control group. The method has five steps: define the hypothesis, choose exposed and control groups, hold the design constant for a set period, measure the gap, and decide what the gap means before acting on it.
What this looks like in a real account
Step one: write the hypothesis down before you touch a bid
A usable incrementality test starts with a specific, falsifiable claim — not "is DSP working" but "does prospecting audience X produce purchases that wouldn't have happened without it, over a 30-day window, at a rate that justifies its cost per acquisition." The specificity matters because it determines every other design choice: which group gets held out, how long the test runs, and what result would actually change the budget decision. A vague hypothesis produces a test that's hard to design and even harder to act on once it's done.
Step two: choose your exposed and control groups
Two common designs: a holdout, where a randomly selected slice of your eligible audience is deliberately excluded from the campaign while the rest sees it normally, and a geo-lift, where spend is paused in a matched set of geographic markets while running normally elsewhere. A holdout answers the audience-level question directly but requires the platform to support true randomised exclusion, not just an exclusion list you build manually. A geo-lift is easier to execute cleanly on Amazon DSP and sponsored ads together, since it doesn't require per-shopper randomisation — you're comparing markets, not individuals.
Step three: hold the design constant, and control for known confounders
Once the test starts, don't touch bids, budgets or creative in either the exposed or control condition — any mid-test change reintroduces the confounding the design was built to remove. Also check the calendar: our own data shows Prime week alone roughly doubles DSP click costs across the whole platform, agency-wide, a 116% increase we've measured directly. Launching a test that straddles a tentpole event without accounting for it means the test period's most expensive days dominate the read, in a way that has nothing to do with whether the campaign itself is incremental. Either exclude tentpole weeks from the measurement window, or extend the test long enough that they're a small share of it.
Step four: the arithmetic
Take a geo-lift test across 10 matched markets, 5 exposed and 5 held out, running for 30 days. The 5 exposed markets generate $180,000 in sales; the 5 control markets, matched for size and historical performance, generate $150,000. The lift is (180,000 − 150,000) / 150,000 = 20%. If the campaign spent $15,000 across the exposed markets during the test, the incremental revenue is $30,000 — the $180,000 actual minus the $150,000 the matched control implies would have happened anyway — giving an incremental ROAS of $30,000 / $15,000 = 2.0x. That's a meaningfully different, and more honest, number than whatever last-touch attribution reported for the same campaign, because it's measuring causation rather than mere sequence.
Step five: decide what the result means, including when it's bad news
If the lift comes back small or statistically inconclusive — a common outcome, and not a wasted test — the honest response isn't to discard the design and try again until a favourable result appears. It's to check test power first: was the sample size, in markets or audience, actually large enough to detect a lift of the size you'd expect, given normal variance? Our own DSP book shows individual advertiser ROAS ranging from 0.85x to 18.03x across 27 accounts in one 31-day window — that much natural variance means an underpowered test can easily produce a null result even when a real, smaller lift exists. If power was adequate and the result is still genuinely flat, that's real information: the spend, at that budget level and audience definition, isn't producing incremental sales, and the honest move is to reallocate rather than keep funding it on the strength of attributed numbers the test just contradicted.
The common mistake, including ours
The mistake is running a holdout or geo-lift designed and computed entirely by the same party being measured on its results, without writing the test design into the scope of work in advance. "Is this incremental, or would they have bought anyway" is the single most common objection we hear in DSP conversations, and the honest response to it is naming who designs and computes the test — a holdout run by the party being paid on the outcome, without independent oversight of the design, isn't a holdout a skeptical stakeholder should fully trust, us included. We now write the test design, the holdout definition and the reporting cadence into the SOW before spend starts, specifically so that objection has a real answer rather than a reassurance.
| Design element | What it controls for | Common failure |
|---|---|---|
| Randomised or matched control group | Selection bias | Using an exclusion list instead of true randomisation |
| Fixed test period, no mid-test changes | Confounding from unrelated optimisation | Adjusting bids or creative during the test |
| Excluding or accounting for tentpole weeks | Seasonal cost and demand spikes | Launching a test that straddles Prime week unadjusted |
| Adequate sample size / test power | False negatives from natural variance | Concluding 'no lift' from an underpowered test |
Which one you should actually pick
Any brand with enough scale to reach test power can design and run this themselves, particularly a geo-lift, which needs no special platform access beyond standard reporting and a willingness to hold spend constant for the test window. Where reMKTR's practice differs is writing the test design into the SOW before spend starts and running it on a fixed cadence rather than as a one-off — a discipline that's easy to state and genuinely harder to keep to under deadline pressure, as part of the same practice behind Full Circle's $500M+ in managed Amazon spend across 100+ brands.
Shortlist on the job, not the feature grid. Pull your search-term report for the last 90 days and total the spend against terms that produced no orders — 33.6% on the account above. Then ask each vendor on your list what they would do about it in week one, and see who answers with a process rather than a screenshot.
Common questions
How long should an incrementality test run?
Long enough to reach adequate statistical power for the lift size you expect, and ideally excluding or spanning past any tentpole event that would distort the read — often 4-8 weeks as a practical range, though it depends on baseline conversion volume.
What's the difference between a holdout and a geo-lift test?
A holdout randomly excludes a slice of an audience from exposure; a geo-lift pauses spend in matched geographic markets instead. Geo-lift is often easier to execute cleanly on Amazon ad products since it doesn't require per-shopper randomisation.
What does it mean if my incrementality test shows no lift?
First check whether the test had adequate power to detect a realistic lift size — an underpowered test can show no lift even when a real, smaller one exists. If power was adequate, a null result is real information worth acting on.
Who should design an incrementality test — the agency running the media, or someone independent?
The honest answer is that the design, the holdout definition and the reporting cadence should be agreed and written into the scope of work before spend starts, regardless of who ultimately runs it — that's a fair question to ask of any vendor, including us.
We show the method before the number.
Claim the free auditRead next
- Tinuiti Pricing: How the Quote Gets BuiltPricing · tinuiti pricing
- Kenshoo Alternative: Start By Noticing It Is Skai NowAlternative · kenshoo alternative
- Pacvue Pricing: No Public Number — What to AskPricing · pacvue pricing
- Skai vs Pacvue: Contracts, Not Feature GridsHead to head · skai vs pacvue