HomeLearnHoldout Testing on Amazon DSP: A Practical Guide
Comparison

Holdout Testing on Amazon DSP

Updated 2026-08-21 · 1257 words · Written against what currently ranked for “Holdout testing on Amazon DSP”
The short answer

A holdout test on Amazon DSP withholds a randomly selected audience segment from exposure to a campaign while the rest of the eligible audience sees it normally, then compares purchase rates between the two groups over a fixed window. The design is built and measured inside Amazon Marketing Cloud, and the most common mistake is ending the test before the audience models have finished learning.

What this looks like in a real account

$89,885
of ad spend — 33.6% of everything the account spent — went to search terms that produced zero orders
Walkize · Amazon account data, Dec 2025–Aug 2026
89,045
individual search terms took money over the same period and returned nothing at all
Walkize · Amazon account data, Dec 2025–Aug 2026
75.5%
of all sales came from the top 1% of search terms. The other 99% is where the decisions actually are
Walkize · Amazon account data, Dec 2025–Aug 2026
2.25x
$267,131 of spend against $601,614 of sales — a 44.4% ACoS, with all of the waste above still sitting inside it
Walkize · Amazon account data, Dec 2025–Aug 2026

What makes a DSP holdout different from a general incrementality test

DSP's audience-based targeting makes true randomised holdouts more achievable than on many other channels — because the platform is already segmenting eligible shoppers into an audience, carving out a randomly selected control slice from that same audience pool is a design DSP's own targeting infrastructure supports directly, rather than requiring a separate geographic proxy. That's the main practical advantage of a DSP holdout over a geo-lift: it tests the actual audience being targeted, not a market-level stand-in for it.

Setting up the test

Define the audience you want to test — a prospecting segment, a retargeting pool, a specific supply source — and split it randomly, holding out a slice (commonly 10-20% of the eligible audience, sized to reach adequate statistical power for your expected lift and baseline conversion rate). Run the campaign normally against the exposed majority while the holdout slice sees no exposure to that specific campaign. Both groups need to be tracked through to purchase inside AMC, using the pseudonymized join that lets you measure conversion in the control group even though, by definition, they received no ad exposure to attribute a sale to.

The three-week mistake

The single most common design failure we see is ending the test too early. In our own experience, three weeks is the point at which most brands ask to pause a DSP test, and it's close to the worst possible moment to do it — the audience models are typically still resolving at that stage, and pausing discards the learning along with the spend. We've watched one supplement brand ask to stop after three weeks and roughly $6,000 spent, at exactly the point its return had quadrupled inside a Prime week that had raised CPCs across the whole platform — pulling the test then would have measured the platform's most expensive, least-settled period as if it were representative. The useful commitment to agree in advance isn't a contract length; it's a learning window, stated and agreed before the test starts, that both sides commit to honouring even if early numbers look unimpressive.

A worked example

Take a prospecting audience of 200,000 eligible shoppers. Hold out 15% — 30,000 — leaving 170,000 exposed. Over a 6-week test window, the exposed group converts at 0.42% (714 purchases) and the holdout group converts at 0.31% (93 purchases, scaled to the same population size for comparison — 0.31% of 30,000 is 93). Scaling the holdout rate up to the exposed group's population size for a fair comparison: 0.31% of 170,000 would be roughly 527 purchases if the exposed group had converted at the holdout's rate. The difference — 714 actual versus 527 expected at the control rate — is 187 incremental purchases, or roughly 26% lift over what would have happened without the campaign.

The common mistake, including ours

Beyond ending tests too early, the mistake we've made ourselves is under-sizing the holdout slice on a lower-volume account, producing a control group too small to detect anything but a very large lift reliably. A 30,000-person holdout works cleanly on a large prospecting pool; carve the same 15% out of a 10,000-person audience and the resulting 1,500-person control group may simply not have enough statistical power to distinguish a real, moderate lift from noise. We now check expected sample size against a power calculation before committing to a holdout percentage, rather than defaulting to the same ratio on every account regardless of scale.

What to do with the result once you trust it

A confirmed positive lift is the point to size the finding, not just celebrate it — recompute the incremental cost per acquisition using the incremental purchase count rather than the attributed one, since that's the number that actually reflects what the campaign cost you per real new customer. In the worked example above, 187 incremental purchases against whatever the campaign spent over the six weeks gives an incremental CPA that's a genuinely more honest efficiency number than the platform's own attributed CPA, and it's the number worth carrying into a scaling decision rather than the attributed figure the dashboard already shows you.

When the holdout result is bad news

If a well-powered, full-duration holdout test shows minimal or no lift, the honest response is to treat that as real information about that specific audience and campaign, not as a reason to distrust holdout testing generally. Before reallocating the budget entirely, check whether the exposed and control groups were genuinely comparable at the start — a randomisation error is more common than it should be — and whether the test window included a distorting event like a tentpole sale that wasn't accounted for. If both check out clean and the lift is still genuinely absent, that's the signal to move the budget to a different audience or channel, not to keep funding the tested segment on the strength of its attributed numbers alone.

Side by side — Holdout testing on Amazon DSP
Design choiceTypical rangeWhy it matters
Holdout size10-20% of eligible audienceToo small undermines statistical power; too large wastes reach
Test duration6+ weeks, learning-window agreed in advanceAudience models are still resolving around the 3-week mark
Tentpole handlingExclude or extend past major sale eventsPrime week alone has driven a measured 116% DSP click-cost increase agency-wide
Population comparison methodScale holdout rate to exposed group sizeA raw percentage comparison without scaling misreads the lift

Which one you should actually pick

Any advertiser with AMC access and a large enough audience can design and run a DSP holdout test themselves, following the method above. reMKTR builds the learning-window commitment into every test upfront specifically because we've seen a promising test pulled too early before, as part of the same discipline behind Full Circle's $500M+ in managed Amazon spend across 100+ brands.

What to do with this

Shortlist on the job, not the feature grid. Pull your search-term report for the last 90 days and total the spend against terms that produced no orders — 33.6% on the account above. Then ask each vendor on your list what they would do about it in week one, and see who answers with a process rather than a screenshot.

Common questions

How big should a DSP holdout group be?

Commonly 10-20% of the eligible audience, but the right size depends on a power calculation against your expected lift and baseline conversion rate — smaller accounts need a careful check that the holdout isn't too small to detect a real result.

How long should a DSP holdout test run?

At least 6 weeks in most cases, and specifically past the roughly 3-week mark where audience models are typically still resolving — ending early is the most common design mistake we see.

What if a tentpole event like Prime Day falls during my test?

Either exclude that window from the measurement period or extend the test long enough that it's a small share of the total — an unaccounted-for tentpole event will distort the read in ways that have nothing to do with whether the campaign is incremental.

Can I run a DSP holdout without Amazon Marketing Cloud?

Not reliably — measuring purchase behaviour in a genuinely unexposed control group requires the pseudonymized join AMC provides. Native DSP or sponsored ads reporting alone can't measure a group that, by design, received no attributable ad exposure.

We show the method before the number.

Claim the free audit
Written against what currently ranked for “Holdout testing on Amazon DSP”, checked 2026-08-21: advertising.amazon.com. Vendor prices change without notice — check the vendor's own page before you budget. Our own figures are labelled with the account and period they came from.