- Over twenty years of research, Tetlock scored nearly 30,000 probability forecasts from 284 political and economic experts against reality using Brier scores — the same rule used to grade weather forecasters [1][2][4].
- In aggregate, experts barely beat a dart-tossing chimp, and failed to beat both “dilettantes” (informed non-specialists reading outside their field) and simple mechanical extrapolation algorithms that just projected the recent past forward — only Berkeley undergraduates did worse than chance [1].
- Experts were measurably overconfident: forecasts they called 100% certain came true about 80% of the time, and 80%-confidence calls came true about 65% of the time [1].
- Accuracy tracked how an expert thought more than what they thought or who they were: “foxes,” who drew on many small models, outperformed “hedgehogs,” who reasoned from one big theory — and the gap was widest on long-range forecasts inside a hedgehog’s own specialty, where confidence in the theory let them embellish furthest from what actually happened [1].
- Media visibility ran the opposite direction from accuracy: the only consistent predictor of forecasting performance was a forecaster’s fame by Google-count, and the better-known a forecaster was, the worse calibrated they tended to be [3].
- Tetlock is explicit that the “experts know nothing” sound bite the book is often reduced to overstates what the data show — the tournaments ran on political and economic questions with a resolvable ground truth (elections, wars, growth rates) through the 1980s and 1990s, and a subset of forecasters beat chance consistently on near-term predictions [1].
Expert Political Judgment: How Good Is It? How Can We Know?
Tetlock's twenty-year forecasting study of expert political judgment, including the fox-hedgehog distinction, the inverse relationship between an expert's media fame and forecasting accuracy, and the finding that many experts underperformed simple benchmarks.
- Author
- Published
- 2005
- Purchase
- Amazon
- When you’re running a Calibration Call and one advisor is far more confident and well-known than the others: don’t let that decide it. In this data, fame and confidence ran inversely to accuracy [3] — ask for their track record, not their certainty.
- When one advisor has a single organizing theory of what makes someone great for the role (“it’s all pipeline,” “only ex-Stripe people work”): that’s a hedgehog, and hedgehogs did worse, not better, the further out and more in-domain the forecast [1] — weight the advisor who synthesizes several angles over the one with the cleanest story.
- When you’re forecasting whether a hire, a market, or a bet will still look good in two or three years: resist a confident expert’s long-range extrapolation from their pet theory. In aggregate, experts didn’t reliably beat mechanical algorithms that just projected the recent past forward [1] — a base-rate estimate deserves real weight next to a compelling story.
1