Skip to content

4 min read

The Unfavorable Economics of Knowing If Your Ads Work

Twenty-five large randomised advertising experiments, $2.8M of spend, and confidence intervals over 100 points wide. Why ad measurement is hard for reasons no dashboard can fix.

  • Advertising
  • Measurement
  • Marketing Science

The reason you cannot tell whether your advertising works is not that your tracking is set up badly. It is that the signal is small relative to the noise in sales, and no amount of instrumentation changes that ratio.

This is the least comfortable post in this series, because its conclusion is that a service you might want to buy — reliable measurement of advertising return — is frequently unpurchasable at your scale, from anyone, at any price.

The core result

Randall Lewis and Justin Rao published in the Quarterly Journal of Economics the definitive statistical critique of advertising measurement Lewis & Rao 2015. Their evidence base is unusually strong: 25 large-scale digital advertising randomised controlled trials with well-known U.S. firms — 19 retailers and 6 financial services companies — collectively $2.8 million in ad spend, each experiment covering more than 500,000 unique users and most over a million.

Their finding is not that advertising does not work. It is that even these experiments could not tell.

The median standard error on ROI was 26.1% for the retail campaigns and 115% for the brokerages. Median confidence intervals on ROI exceeded 100 percentage points in width.

Why the noise is so large

The mechanism is simple once stated. Advertising campaigns are cheap per person and sales are variable per person.

Lewis and Rao note that campaign cost ranged from $0.02 to $0.35 per exposed user, most near $0.10. Meanwhile individual purchase behaviour has enormous variance — most people buy nothing, a few buy a lot.

Write the detectable effect from the sample-size relation:

Δ    σn\Delta \;\propto\; \frac{\sigma}{\sqrt{n}}

The problem is that σ\sigma for individual sales is very large relative to the plausible per-person effect of a $0.10 ad exposure. You are trying to detect a small shift in the mean of a heavy-tailed distribution. The required nn is enormous, and it grows with the square of how small the effect is.

Figure 1Why ad measurement is structurally hard: the effect being estimated is small relative to the natural spread of per-customer spend.Diagram by the author; illustrative.

The observational estimates are not merely noisy — they are biased

If randomised estimates are imprecise, the natural fallback is observational modelling: regress conversions on exposures, control for what you can. Blake, Nosko and Tadelis showed at eBay exactly how badly that fails Blake et al. 2015.

Two experiments. In the first, eBay stopped paid search on all brand queries on Yahoo! and MSN while continuing on Google as a control. The result: 99.5% of clicks were retained. Almost all the traffic that paid search was being credited with would have arrived anyway through organic results.

In the second, covering 30% of U.S. traffic over 60 days across more than 100 million managed keywords, the estimates diverged spectacularly by method:

MethodEstimated ROI
OLS> 4000%
With fixed effects> 1500%
Experimental / IV−63%

26.1%

Median standard error on ROI, retail RCTs

Lewis & Rao, QJE 2015

99.5%

Brand-search clicks retained when ads were switched off

Blake et al., Econometrica 2015

−63%

Experimental ROI where OLS estimated >4000%

Blake et al., Econometrica 2015

The bias runs in one direction — toward flattering the channel — because people who were already going to buy are disproportionately the ones who see and click your ads. This is selection, not a tracking gap, and better attribution software does not touch it.

A note on the counter-evidence

You will encounter the claim that roughly 89% of paid search clicks are incremental. It originates with Chan, Yuan, Koehler and Kumar, based on over 400 "search ads pause studies" Chan et al. 2011.

What to do instead

Three practices survive this literature.

Prefer big interventions you can switch off. A holdout with a large expected effect is measurable at small samples. Pausing a channel entirely for four weeks tells you more than a year of attribution modelling.

Measure closer to the decision. Qualified conversations are less noisy than revenue, arrive sooner, and are what your acquisition system actually produces. Revenue variance is the enemy; do not put it in the denominator of your learning loop if you do not have to.

Accept "we do not know" as a finding. The most useful thing measurement can tell a small business is that a given question cannot be answered at their scale — which converts an expensive attribution project into a decision made on judgement, deliberately, with the uncertainty acknowledged rather than laundered through a dashboard.

That is what "measure honestly," the third of Ariadne's working principles, actually commits me to. Sometimes the honest measurement is a wide interval and a shrug.

References

  1. Lewis, R. A., & Rao, J. M. (2015). The unfavorable economics of measuring the returns to advertising. The Quarterly Journal of Economics, 130(4), 1941–1973. https://doi.org/10.1093/qje/qjv023Authors at Google and Microsoft Research at publication; experiments run at Yahoo!, their former employer.
  2. Blake, T., Nosko, C., & Tadelis, S. (2015). Consumer heterogeneity and paid search effectiveness: A large-scale field experiment. Econometrica, 83(1), 155–174. https://doi.org/10.3982/ECTA12423
  3. Chan, D. X., Yuan, Y., Koehler, J., & Kumar, D. (2011). Incremental clicks: The impact of search advertising. Journal of Advertising Research, 51(4), 643–647. https://doi.org/10.2501/JAR-51-4-643-647Peer-reviewed, but all four authors were Google employees and the result favours the publisher's product.
  4. Kohavi, R., Longbotham, R., Sommerfield, D., & Henne, R. M. (2009). Controlled experiments on the web: Survey and practical guide. Data Mining and Knowledge Discovery, 18(1), 140–181. https://doi.org/10.1007/s10618-008-0114-1

Next: stage 08, and the law that caps every automation project before it starts.

Sourena Khanzadeh

Founder & Growth Engineer, Ariadne Growth Systems

Toronto, Canada

Ariadne Growth SystemsGrowth System Auditsupport@ariadne.fyi