The question incrementality testing actually answers

Platform-reported ACOS and ROAS tell you what sales are attributed to ads, not what sales advertising actually caused — see Attribution in Marketplace Advertising for why those aren't the same thing. Incrementality testing is a more rigorous (and more effortful) attempt to answer the harder, more useful question directly: how many additional sales did this advertising cause that wouldn't have happened otherwise?

The basic logic: a holdout test

The cleanest version of an incrementality test compares two similar groups — one that receives advertising (the "exposed" group) and one that doesn't (the "holdout" or "control" group) — over the same period, then compares the resulting sales difference. If the exposed group's sales are meaningfully higher than the holdout group's, after accounting for any pre-existing baseline difference between them, that gap is a reasonable estimate of the advertising's true incremental effect.

Why this is harder in a marketplace context than it sounds

Unlike a controlled lab experiment, a seller can't perfectly randomize which shoppers see ads and which don't — marketplace ad platforms don't typically offer a formal, built-in holdout/control tool the way some larger advertising platforms do for major brand advertisers. Sellers with meaningful scale approximate this in a few practical ways instead:

Product-pair testing. If you have two genuinely comparable products (similar price, similar demand, similar review count) in the same category, run ads on one and not the other for a defined period, then compare their sales growth relative to their own recent baseline. This isn't a perfect randomized test, but it's far more informative than assuming any single product's ad-driven sales are fully incremental.

Time-based on/off testing. Turn a campaign off entirely for a defined period (a week or two, ideally during a period without unusual seasonality or promotions) after establishing a stable baseline, and compare total (not just ad-attributed) revenue during the off period against a comparable prior period with ads running. A meaningful drop in total revenue during the off period is evidence the advertising was driving real incremental sales, not just capturing sales that would have happened anyway; little to no drop suggests more cannibalization than the ACOS alone would have suggested.

Geo or marketplace-based testing. For sellers operating in multiple regions or on multiple marketplaces, running ads in one geography/marketplace and holding off in a comparable one over the same period offers another rough approximation of a holdout test.

Worked example: a time-based on/off test

A seller with a stable-selling product running Sponsored Products at a steady $50/day for months wants to know how incremental that spend really is. They pause the campaign entirely for two weeks (choosing a period with no other planned pricing or promotional changes), having tracked total weekly revenue for the prior eight weeks as a baseline (averaging $3,000/week). During the two-week pause, weekly revenue falls to an average of $2,600/week — a drop of about 13%, smaller than the roughly 20% of revenue that had been ad-attributed while campaigns were running. This suggests some, but not all, of the ad-attributed revenue was genuinely incremental — a meaningful portion likely reflects sales that would have happened organically anyway, information the standalone ACOS number couldn't have revealed.

What to watch out for that can bias the result

Seasonality and external events. Running a test during a period with unrelated demand swings (a holiday, a competitor stockout, a viral moment) can produce a misleading result — choose as stable and "normal" a period as possible, and be cautious interpreting a test that happened to coincide with an unusual event.

Too short a test window. A one- or two-day test is unlikely to produce a reliable signal given normal day-to-day sales variance — a longer window (at least a week, often two) gives a more trustworthy read.

Organic ranking effects during the pause. If turning ads off causes a temporary dip in organic ranking (because ad-driven sales velocity had been propping up ranking), the test may understate baseline organic performance in a way that's specific to having recently run ads, rather than reflecting how the product would perform if it had never been advertised at all.

When incrementality testing is (and isn't) worth the effort

This is a meaningful undertaking — pausing a working campaign deliberately costs some sales in the short term to gain a clearer long-term signal. It's most worth doing for larger-spend campaigns where getting the true incrementality answer materially changes a budget decision, and least worth it for small campaigns where even a wrong assumption about incrementality doesn't change what you'd do differently.

Mistakes to avoid

  • Running a test during an atypical period (holiday, promotion, external demand shock) and treating the result as representative.
  • Testing for too short a window to distinguish a real effect from ordinary day-to-day variance.
  • Concluding "0% incremental" or "100% incremental" from one test — real incrementality is almost always somewhere in between, and a single test is a data point, not a permanent verdict.

FAQs

Do I need a data science background to run this? No — a basic before/after or product-pair comparison, with attention to avoiding an atypical test period, is accessible without formal statistical training, though a longer test window and more careful baseline comparison improve confidence in the result.

How often should I re-run an incrementality test? Not continuously — it's a periodic check (perhaps annually, or when a major strategy or budget decision is on the table) rather than an ongoing operational metric like ACOS.

Is there a simpler alternative to a full test? Tracking TACOS over time (see ROAS vs. ACOS vs. TACOS) is a lower-effort, ongoing directional signal, even though it's less precise than a deliberate holdout test.