What Is Incrementality Testing?
Incrementality testing is a controlled experiment — a holdout, geo test, or matched-market test — that measures how many conversions your ads actually caused, versus what would have happened without them. It replaces attribution guesswork with a test group and a control group, then measures the gap.
What this looks like across the book we manage
What It Actually Measures
Attribution tells you which touchpoint a customer saw last before they bought. Incrementality testing asks a different question: would that sale have happened anyway? Those are not the same question, and the gap between them is where most wasted ad spend hides.
Last-click attribution can't answer the causal question because it has no way to see the customer who would have converted with zero ads. It just counts who touched what. Incrementality testing fixes that by building a control group — a slice of the audience that sees no ad, or a placebo ad — and comparing its conversion rate to the group that was exposed. The difference is the real, incremental contribution of the media.
This matters most for channels where the platform grades its own homework: Google, Meta, and any walled garden reporting self-attributed conversions. It's also the only honest way to size the contribution of channels that generate impressions without clicks, like display and video, where last-click credit is close to meaningless.
The Formula and a Worked Example
The standard calculation: (Test group conversion rate − Control group conversion rate) ÷ Test group conversion rate = Incrementality %.
Say a test group of 20,000 people sees the campaign and converts at 3.0%. A matched control group of 20,000 sees nothing and converts at 1.8%. (3.0% − 1.8%) ÷ 3.0% = 40% incrementality. That means 60% of the conversions in the test group would have happened anyway — the campaign only caused four in ten of them.
Here's why the platform number and the incremental number can diverge sharply. Across 30 of our advertisers in July 2026, the book attributed 57,137 purchases at a blended $5.49 cost per acquisition, with 20.1% of those from shoppers new to the brand. That figure is real and it's useful — but it's an attributed number, not a tested one. It tells you what the platform counted, not what the platform caused. Reconciling that number against a holdout, inside Amazon Marketing Cloud where DSP and sponsored ads stop double-counting each other, is the only way to know how much of it was incremental.
Test Types, Side by Side
Not every channel can run the same test. The table below covers the four structures you'll actually encounter.
When the Number Comes Back Bad
A low or negative incrementality result isn't a failed test. It's the test working. Three things to check before you act on it:
- Holdout too small. If your control group is under roughly 10% of total reach, noise can swamp the signal and a real lift can look flat.
- Seasonality contamination. If demand for your category swings independently of ad spend — a launch, a competitor promo, a holiday — a test run across that window can blame or credit the ads for something else entirely.
- Sample ratio mismatch. If the test and control groups aren't actually comparable in size, geography, or prior purchase behavior, the comparison is void before you've even read the result.
We've made the small-holdout mistake ourselves, early on, sizing a control group to protect media efficiency rather than statistical power — the test came back inconclusive and had to be rerun for another cycle. The fix isn't to distrust every bad number. It's to check the design before you trust any number, good or bad.
Google, Meta, and the Specialist Vendors
Google runs its own conversion lift studies inside Google Ads, and Meta runs conversion lift studies inside Ads Manager — both are native, both are free to run, and both are graded by the platform whose spend is being tested. That's a real limitation, not a fatal one: they're a reasonable first check, especially for single-channel questions.
Independent vendors go further. Measured runs continuous, always-on experimentation across hundreds of data sources and uses the results to calibrate a media mix model — a genuinely strong fit for brands running many channels at once who want one reconciled view. Haus specializes in geo-experiments and incrementality-as-a-service, and suits teams that want a dedicated testing partner without building the infrastructure themselves. Both are legitimate answers to this question and outperform doing nothing.
What none of them do natively is see inside Amazon's own walled garden the way a clean room built for it can. Sponsored ads and DSP report separately by default, which means a brand running both can double-count the same shopper twice before any incrementality test even starts.
Where Amazon DSP Fits
Dr. DSP is Amazon DSP — Demand-Side Platform, Amazon's programmatic display, video, and audio buying — run as a managed product, not the Delivery Service Partner courier franchise. It's built on the same holdout-and-control logic described above, reconciled inside Amazon Marketing Cloud so DSP and sponsored ads stop claiming the same sale twice. Every change we make carries the evidence behind it, a measurement plan, and a rollback trigger before it goes live. Full Circle has managed more than $500M in revenue across 100+ brands; there's no published price, a demo and the first 30 days are free, and pricing is set on the call against actual budget and scope. Whether or not that's the right fit for you, the underlying test — holdout versus exposed, causal versus correlated — is the same one worth demanding from anyone measuring your media.
| Test type | How it works | Best fit | Watch out for |
|---|---|---|---|
| Holdout / PSA-placebo | Randomly withholds the ad from a slice of the audience; shows nothing or a placebo instead | Single-channel tests, Meta and Google lift studies | Holdout too small to detect real lift |
| Geo holdout | Turns spend off in some markets, keeps it running in matched markets | Channels with weak user-level tracking: CTV, audio, linear TV | Regional demand differences unrelated to the ad |
| Matched-market test | Pairs similar stores or regions rather than randomizing individuals | Retail, offline sales, DTC with a regional footprint | Picking 'similar' markets that aren't actually comparable |
| Ghost bids / counterfactual | Programmatic platform logs what would have shown if the bid had lost the auction | Programmatic display and video, DSP-specific testing | Requires platform-side data you can't get from a walled garden's own dashboard |
Which one you should actually pick
Use Google's or Meta's native lift studies for a fast, single-channel check — they're free and built in. Use Measured or Haus if you're running many channels and want continuous testing tied to an MMM. Use Amazon DSP's clean-room approach specifically when DSP and sponsored ads need to stop double-counting each other.
Shortlist on the job, not the feature grid. Pull your search-term report for the last 90 days and total the spend against terms that produced no orders — 48.5% across the 47 brands above. Then ask each vendor on your list what they would do about it in week one, and see who answers with a process rather than a screenshot.
Common questions
Is incrementality testing the same as A/B testing?
They're built the same way — a test group and a control group — but they answer different questions. A/B testing usually compares two versions of the same ad to see which performs better. Incrementality testing compares exposed versus unexposed to see whether the ad caused anything at all.
How big does my holdout group need to be?
Most practitioners treat roughly 10% of total reach as a working minimum for the control group, though the real number depends on your baseline conversion rate and how small a lift you need to detect. Too small and a real effect gets lost in noise; too large and you're needlessly withholding revenue.
What's the difference between incrementality testing and a media mix model (MMM)?
Incrementality testing is an experiment — it directly measures cause and effect for one channel or tactic over a defined window. An MMM is a statistical model fit to historical spend and outcomes across all channels at once. The two work best together: incrementality results are commonly used to calibrate and correct the MMM's assumptions.
Does Google or Amazon have their own incrementality testing tool?
Google offers conversion lift studies inside Google Ads. Amazon's version runs through Amazon Marketing Cloud, which can reconcile DSP and sponsored ads so they're not both claiming credit for the same sale — useful specifically because Amazon's own reporting otherwise keeps those two ad products separate.
What does it mean if my incrementality result comes back near zero?
It usually means the channel is reaching people who were going to convert anyway — not that the channel is worthless, but that at current spend levels it isn't adding new demand. Before cutting the budget, check the test design for a too-small holdout, seasonality overlap, or a mismatched control group; a near-zero result built on a flawed test isn't a real answer.
Dr. DSP is Amazon DSP — the Demand-Side Platform, not the delivery franchise — run daily by Fable 5 with operators from a $500M+ Amazon team supervising. You pick the approval level, we reconcile in Amazon Marketing Cloud, and Orbit is included. First 30 days free, priced on the call.
Book a Dr. DSP demoRead next
- Acorn Cost: What Acorn-i Charges, and What to AskPricing · acorn cost
- ChannelAdvisor Alternative: Two Jobs, One Hard ExitAlternative · channeladvisor alternative
- Pacvue Pricing: What Quote-Only Really Means for BuyersPricing · pacvue pricing
- Intentwise Pricing: Quote-Only, and What It BuysPricing · intentwise pricing
- Perpetua Pricing: What the Page Shows, and What It Doesn'tPricing · perpetua pricing
- Quartile Pricing: What the Terms Commit You ToPricing · quartile pricing