Intuitive Evaluation of Likelihood Judgment Producers: Evidence for a Confidence Heuristic

Paper · 2004

Price and Stone's experiments on how evaluators judge advisors: given identical track records, people consistently prefer the more confident advisor — including the overconfident one — a bias Management Craft flags as a trap for anyone judging candidates or peers on stated conviction rather than proof.

Author
Paul C. PriceEric R. Stone
Published
2004

The question

When two advisors have made the same underlying calls, does the one who states them with more confidence look more competent?

The method

In three experiments, college students studied two fictional financial advisors’ stock-direction judgments paired with the actual outcomes, then said which advisor they’d hire — one advisor (“moderate”) was reasonably well-calibrated, the other (“extreme”) was overconfident but built to have identical discrimination and identical categorical correctness (both were right 75% of the time in Experiment 3) [1].

The findings

Participants preferred the overconfident advisor in all three experiments — 71% in Experiment 1 [1], 64% in Experiment 2 [1], 63% in Experiment 3 [1] — even though the two advisors’ actual accuracy was held equal by design. Experiment 2 showed this tracked a belief about knowledge: participants who preferred the extreme advisor mostly rated him more knowledgeable, with no matching pattern for who was “more honest” [1]. Experiment 3 pinned down the mechanism directly: participants overestimated the confident advisor’s percentage of correct calls and underestimated the moderate advisor’s, even though both were correct exactly 75% of the time [1] — the authors call this the confidence heuristic, and present a quantitative model in which each point of perceived confidence gap buys the more confident advisor roughly 0.21 points of assumed correctness advantage [1]. A secondary result: participants high in both need for cognition and right-wing authoritarianism were most likely to prefer the extreme advisor (86% vs. 54% for everyone else) [1].

The limits

This is a lab paradigm with fictional advisors and no personal stakes beyond course credit and a small bonus; the paper doesn’t test whether the effect survives repeated real-world feedback over many more trials, and the authors themselves note 48 trials may be too few for genuine calibration-learning to set in [1].

  • When you’re comparing two people’s confident-sounding takes on a candidate or a call, and one is simply more assertive: that assertiveness is not evidence of being right. This paper’s design held the two advisors’ actual accuracy identical and people still preferred the more confident one, in all three experiments [1].
  • When a reference call or Calibration Call advisor states an opinion with total certainty: don’t let the certainty substitute for their track record. Participants in this study didn’t just prefer the confident advisor — they misremembered his hit rate as higher than it actually was [1]. Ask what they’ve been right and wrong about before, not just how sure they sound now.

1
Paul C. Price and Eric R. Stone, "Intuitive Evaluation of Likelihood Judgment Producers: Evidence for a Confidence Heuristic," Journal of Behavioral Decision Making 17 (2004): 39–57,
https://doi.org/10.1002/bdm.460