Top Performers Are Not the Most Impressive When Extreme Performance Indicates Unreliability

Paper · 2012

Denrell and Liu's 2012 PNAS paper showing that where luck moves outcomes a lot, the very best result in a field is weak evidence of the best ability, because an extreme number says more about how noisy the setting is than about the person who posted it. Their conclusion is that the top performers should not automatically be the ones you imitate.

Published
2012

The question

If someone posted the best number in the field, are they the most able person in the field?

The method

Two mathematical models rather than a field study. “Model 1 added heterogeneity to a model of self-reinforcing performance, and model 2 added heterogeneity to the standard model of performance as a” linear combination of skill and noise [2]. The models’ prescription was then tested against people: the authors ran experiments in which participants inferred skill from observed performance [3].

The findings

The headline result is counterintuitive and narrow in a useful way: “the highest performers may not have the highest expected ability and should not be imitated or praised” [1]. The mechanism is that “an extreme performance may be more informative about the level of noise and the strength of rich-get-richer dynamics than about skill” [1]. A very large number tells you the setting can produce very large numbers, which is a fact about the setting.

Whether it applies is a question about the domain, not about the person: “whether higher performance indicates higher ability depends on whether extreme performance could be achieved by skill or requires luck” [1]. Where luck cannot manufacture an extreme result, the ordinary intuition holds and then some — “performance is a good indicator of skill and, as our second model shows, extreme performance may be especially informative” [2].

People do not make this correction on their own. In the experiments, “despite clear feedback and incentives to be accurate, 69 out of 119 participants never predicted higher performers to be less skilled than those with moderately high performance” [3].

The limits

The result is conditional and the authors say so twice. It needs a domain “in which performance can be substantially influenced by chance events and even relatively unskilled agents can achieve high performance” [2], and it needs an evaluator who cannot see how noisy the setting is: “Our results would not hold for evaluators who have detailed information about the setting” [2]. This is a modelling paper with supporting experiments, not a study of any particular labour market.

  • Choosing who to call for a Calibration Call. Without this paper you reach for the most celebrated name you can get to, on the reasonable theory that the best result came from the best judgment. Denrell and Liu say that in a luck-heavy field — which describes startups — the most extreme outcome is the least clean signal in the distribution. Two or three people currently doing the work well beat one legend.
  • Deciding whose playbook to copy. Without this paper, “it worked for them” is enough. The paper’s question is prior to that one: could an unskilled operator have produced this outcome in this domain? If yes, the outcome is not the endorsement it looks like.
  • Reading a candidate’s biggest number. A record quarter, a hypergrowth stretch, a fund’s best year: the paper’s argument applies to a résumé line as much as to a mutual fund. Ask what the variance in that role looks like before treating the peak as a measurement.

1
Jerker Denrell and Chengwei Liu, “Top Performers Are Not the Most Impressive When Extreme Performance Indicates Unreliability,” Proceedings of the National Academy of Sciences 109, no. 24 (2012): 9331-9336, abstract and introduction,
https://doi.org/10.1073/pnas.1116048109
2
Denrell and Liu, “Top Performers Are Not the Most Impressive,” § “Summary.”
3
Denrell and Liu, “Top Performers Are Not the Most Impressive,” § “Implications.”