5 min read
Measure Honestly: Goodhart's Law Inside Your Own Dashboard
Why the agency model produces optimised parts and an unimproved whole — traced to a formal result in contract theory, not to bad intentions.
- Measurement
- Incentives
- Growth Engineering
Ariadne's pitch against the agency model is that each specialist is measured on their own piece and none of them is measured on whether you got more customers. That is not a marketing line. It is the multitask principal-agent problem, and it has a theorem.
Every founder who has hired an SEO agency, a paid-media agency and a CRM consultant has lived this: each reports success, and the business does not grow. The instinct is to suspect the vendors. The economics says you would get the same outcome from honest, competent people.
The theorem
Holmström and Milgrom's multitask paper is the foundational formal result Holmström & Milgrom 1991. From page 25:
In general, when there are multiple tasks, incentive pay serves not only to allocate risks and to motivate hard work, it also serves to direct the allocation of the agents' attention among their various duties.
The consequence is the part that matters. Let an agent choose effort on a measurable task and on an unmeasurable one, paid at rate on the first, against a cost of effort . When the two tasks compete for the same attention — that is, when they are substitutes in the cost function — the comparative static runs the wrong way:
Paying harder for the thing you can count reduces the thing you cannot. The optimal contract may therefore involve weak or no incentives at all on the measurable task — because a strong incentive would distort the allocation of attention more than it would motivate effort.
George Baker supplies the complementary result: what happens when the metric you pay on is not the thing you actually want Baker 1992. Contracts based on such performance measures "will not in general provide first-best incentives," even with a risk-neutral agent — because the optimal piece rate must be shrunk below the true value of output to limit distortion. The measure being noisy is not the problem. The measure being different from the objective is.
Getting the attributions right
Three related laws are routinely conflated. Since this post is about honesty in measurement, the citations should be honest too.
Goodhart's law originates in two 1975 papers presented by Charles Goodhart at a Reserve Bank of Australia conference: "any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." The 1975 volume is not digitised, and I have not read it at source — the wording is confirmed only through secondary sources, which disagree slightly on the exact phrasing and on which of the two papers it appeared in Rodamar 2018.
Campbell's law is the sharper sibling, and it is verifiable at source. Donald Campbell, writing on the corrupting effect of quantitative indicators Campbell 1979, states that the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort the processes it was intended to monitor.
The sentence everyone quotes — "when a measure becomes a target, it ceases to be a good measure" — is Marilyn Strathern's, from a 1997 article in European Review about audit in British universities, page 308 Strathern 1997. She was paraphrasing Goodhart. If you have attributed that sentence to Goodhart, you have attributed Strathern's compression of his idea to him.
And the oldest management statement of the same problem is Steven Kerr's 1975 Academy of Management Journal paper, whose title says it: "On the Folly of Rewarding A, While Hoping for B" Kerr 1975.
Naming which failure you have
Manheim and Garrabrant turn the slogan into four mechanically distinct failure modes, which is far more useful diagnostically Manheim & Garrabrant 2018:
| Variant | Mechanism | Marketing example |
|---|---|---|
| Regressional | The proxy correlates imperfectly; selecting hard on it selects noise | Best-converting keyword last month regresses to the mean |
| Extremal | The relationship holds in the normal range and breaks at the extreme | Cheap leads stay decent until you push volume, then collapse |
| Causal | You act on a correlation that is not causal | Boosting a metric that merely accompanied good outcomes |
| Adversarial | Someone optimises against your metric | A vendor generating cheap unqualified inquiries to hit CPL |
4
Distinct Goodhart failure modes worth diagnosing separately
Manheim & Garrabrant 2018
1997
Year Strathern wrote the sentence usually credited to Goodhart
European Review 5(3), p. 308
~70%
Marketing leads never worked — the classic unowned arrow
Sabnis et al., J. Marketing 2013
What follows for how I contract
The theory has three uncomfortable, concrete implications, and I try to live by them.
One person accountable for the whole path. Not because a generalist beats specialists at every stage — they do not — but because Holmström and Milgrom show that splitting measurable duties across separately-incentivised agents predictably starves the unmeasured seams between them. The arrows are where the money leaks.
Weak incentives on proxies, deliberately. The multitask result says the optimal contract on a distorted measure often has a low powered incentive. In practice: fixed scope and fixed fee, with the proxy metrics reported for diagnosis rather than compensation.
Report the metric that can embarrass me. Qualified conversations and closed customers, not impressions and rankings. Those numbers move slowly and sometimes move the wrong way, which is precisely what makes them worth reporting. A metric that only ever goes up is not measuring anything.
The name for the alternative is comfortable measurement: a dashboard of green numbers on a business that is not growing. Every one of those numbers is real. None of them is the objective.
References
- Holmström, B., & Milgrom, P. (1991). Multitask principal–agent analyses: Incentive contracts, asset ownership, and job design. Journal of Law, Economics, & Organization, 7(Special Issue), 24–52. https://doi.org/10.1093/jleo/7.special_issue.24
- Baker, G. P. (1992). Incentive contracts and performance measurement. Journal of Political Economy, 100(3), 598–614. https://doi.org/10.1086/261831
- Kerr, S. (1975). On the folly of rewarding A, while hoping for B. Academy of Management Journal, 18(4), 769–783. https://doi.org/10.5465/255378
- Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67–90. https://doi.org/10.1016/0149-7189(79)90048-X
- Strathern, M. (1997). 'Improving ratings': Audit in the British University system. European Review, 5(3), 305–321 (quote at p. 308). https://doi.org/10.1002/(SICI)1234-981X(199707)5:3<305::AID-EURO184>3.0.CO;2-4
- Rodamar, J. (2018). There ought to be a law! Campbell versus Goodhart. Significance, 15(6), 9. https://doi.org/10.1111/j.1740-9713.2018.01205.xBest short source for getting the Goodhart/Campbell attributions right.
- Manheim, D., & Garrabrant, S. (2018). Categorizing variants of Goodhart's law. arXiv preprint arXiv:1803.04585. https://arxiv.org/abs/1803.04585Preprint; not peer-reviewed.
- Sabnis, G., Chatterjee, S. C., Grewal, R., & Lilien, G. L. (2013). The sales lead black hole: On sales reps' follow-up of marketing leads. Journal of Marketing, 77(1), 52–67. https://doi.org/10.1509/jm.10.0047
Next: the fourth principle — leave the keys — and the industrial-organisation literature on what lock-in is worth to the person doing the locking.
Sourena Khanzadeh
Founder & Growth Engineer, Ariadne Growth Systems
Toronto, Canada
Ariadne Growth SystemsGrowth System Auditsupport@ariadne.fyi