4 min read
The Experiment You Cannot Run
What A/B testing actually costs in sample size, why most small businesses can never afford it, and the four things to do instead.
- Experimentation
- Statistics
- Decision Making
The honest answer to "should we A/B test the new page" is usually no — not because testing is wrong, but because the sample size required exceeds the traffic you will receive before the question stops mattering.
I get asked to set up A/B testing more often than almost anything else, and I turn it down more often than almost anything else. Here is the arithmetic behind that.
The sample size formula
The canonical source is Kohavi, Longbotham, Sommerfield and Henne, whose Equation 2 gives, for 95% confidence and 80% power Kohavi et al. 2009:
where is the variance of the metric and the absolute change you want to detect. For 90% power, replace 16 with 21.
Zhou, Lu and Shallah derive this from the general two-sample power expression and confirm the constant Zhou et al. 2023:
Checking: . The rule of thumb is the formula, not a simplification of it.
What that means for a firm with 40 inquiries a month
Take a realistic case. You want to know whether a new proposal template raises your close rate from 25% to 30%. Conversion is Bernoulli, so , and .
Twelve hundred deals per variant. Two thousand four hundred in total. At 40 inquiries a month that is five years, assuming nothing about your business, market, or offer changes in the meantime.
Kohavi and colleagues state the scaling explicitly: "to increase the experiment sensitivity by a factor of 10, say from 5% delta to 0.5%, you need = 100 times more users" Kohavi et al. 2013. Their general guidance is that teams need at least thousands of active users for this machinery to be worth deploying.
The failure mode that makes it worse
Faced with a slow test, the natural move is to check on it. This is the single most damaging thing you can do.
Johari, Koomen, Pekelis and Walsh quantified it: under continuous monitoring with no correction, Type I error "can easily increase fivefold" at 10,000 samples — a nominal 5% false-positive rate becomes roughly 25% Johari et al. 2022. One test in four that you declare a winner is noise.
Their fix is always-valid inference via a mixture sequential probability ratio test, which permits continuous monitoring by construction Johari et al. 2017. If you do run tests and you will peek — and you will — use sequential methods rather than pretending you did not look.
1,200
Deals per variant to detect a 25% → 30% close-rate change
n = 16σ²/Δ²
~5×
Type I error inflation from unconstrained peeking
Johari et al., Operations Research 2022
~50%
Variance reduction achievable with CUPED, halving required traffic
Deng et al., WSDM 2013
Four things that do work at low volume
Reduce variance instead of adding traffic. CUPED uses pre-experiment data as a control variate. The adjusted estimator has variance , where is the correlation between the metric and its pre-period value. Deng and colleagues report roughly 50% variance reduction in practice — "effectively achieving the same statistical power with only half of the users" Deng et al. 2013. Halving your required traffic is the single largest lever available, and it costs no traffic at all.
Move the metric earlier in the funnel. Kohavi's own worked example makes the point: an e-commerce test needed over 409,000 users per variant to detect a 5% revenue change, but under 122,000 when the metric was switched from revenue to conversion, because Bernoulli variance is so much smaller than revenue variance Kohavi et al. 2009. Pick the earliest metric that still means something.
Use bandits when you want the outcome, not the answer. If your goal is more customers rather than publishable knowledge, randomised probability matching allocates traffic to each arm in proportion to its probability of being best. Scott's simulation found regret from equal allocation "more than an order of magnitude greater than under probability matching" Scott 2010. You give up a clean p-value and get more conversions along the way — usually the right trade for a business.
Run holdouts rather than page tests. Switching a channel off entirely for a defined window is a real experiment with a large effect size, which is exactly what small samples can detect. Big effects need small samples; that is the same formula read in the other direction.
And keep Twyman's law in view, which Kohavi and Longbotham put at the head of their paper on unexpected results: any figure that looks interesting or different is usually wrong Kohavi & Longbotham 2010. At small samples, the striking result is almost always the artifact.
References
- Kohavi, R., Longbotham, R., Sommerfield, D., & Henne, R. M. (2009). Controlled experiments on the web: Survey and practical guide. Data Mining and Knowledge Discovery, 18(1), 140–181. https://doi.org/10.1007/s10618-008-0114-1
- Kohavi, R., Deng, A., Frasca, B., Walker, T., Xu, Y., & Pohlmann, N. (2013). Online controlled experiments at large scale. Proceedings of KDD 2013, 1168–1176. https://doi.org/10.1145/2487575.2488217
- Zhou, J., Lu, J., & Shallah, A. (2023). All about sample-size calculations for A/B testing: Novel extensions and practical guide. Proceedings of CIKM 2023. https://doi.org/10.1145/3583780.3614779
- Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2022). Always valid inference: Continuous monitoring of A/B tests. Operations Research, 70(3), 1806–1821. https://doi.org/10.1287/opre.2021.2135
- Johari, R., Koomen, P., Pekelis, L., & Walsh, D. (2017). Peeking at A/B tests: Why it matters, and what to do about it. Proceedings of KDD 2017, 1517–1525. https://doi.org/10.1145/3097983.3097992
- Deng, A., Xu, Y., Kohavi, R., & Walker, T. (2013). Improving the sensitivity of online controlled experiments by utilizing pre-experiment data. Proceedings of WSDM 2013, 123–132. https://doi.org/10.1145/2433396.2433413
- Scott, S. L. (2010). A modern Bayesian look at the multi-armed bandit. Applied Stochastic Models in Business and Industry, 26(6), 639–658. https://doi.org/10.1002/asmb.874
- Kohavi, R., & Longbotham, R. (2010). Unexpected results in online controlled experiments. ACM SIGKDD Explorations Newsletter, 12(2), 31–35. https://doi.org/10.1145/1964897.1964905
Next: the deeper reason advertising measurement is hard — and it is not a tooling problem.
Sourena Khanzadeh
Founder & Growth Engineer, Ariadne Growth Systems
Toronto, Canada
Ariadne Growth SystemsGrowth System Auditsupport@ariadne.fyi