Using advice from multiple sources to revise and improve judgments

Paper · 2007

Yaniv and Milyavsky's studies of judges revising estimates with advice from multiple advisors, showing accuracy gains that flatten quickly as the number of opinions grows and that judges egocentrically discount others' advice.

Author
Ilan YanivMaxim Milyavsky
Published
2007

The question

When someone revises an initial estimate using two to eight other people’s opinions at once, how do they combine them — and how much do they actually gain?

The method

Undergraduate students gave an initial estimate for 24 historical-date questions [1], then saw two, four, or eight other estimates drawn from a real pool of prior respondents’ answers rather than artificial advice [1], and gave a final estimate, with a real cash bonus for accuracy [1].

The findings

Using advice helped — mean error fell by roughly 27% with two opinions, 28% with four [1], and 33% with eight, gains that grew with more opinions but at a fast-diminishing marginal rate [1]. The best-fitting description of what people actually did wasn’t averaging everything: it was “egocentric trimming” [1] — dropping the one or two opinions furthest from their own initial guess, then averaging what was left — rather than trimming based on distance from the group’s consensus. And people left real accuracy on the table: their actual final estimates were roughly as accurate as the crudest mechanical rule tested (the midrange, which uses only the two extremes) [1] and significantly less accurate than a simple median [1] or a rule that discarded outliers relative to the group [1]. A second experiment (artificial near/far advice rather than a real pool) found the same self-centered pattern from a different angle: participants weighted their own initial estimate at 0.71 on average [1], when equal weighting of self plus advice would call for about 0.33 [1].

The limits

The task was estimating historical dates from numeric advice under a real but modest cash incentive — not the qualitative, higher-stakes advice a hiring manager collects on a reference call. The paper doesn’t test whether experienced professional judgment (rather than undergraduates guessing dates) shows the same egocentric-trimming pattern.

  • When you’re deciding how many Calibration Calls are “enough”: the data says stop worrying about six or eight. Gains from extra opinions flattened fast in this study — 27% → 28% → 33% error reduction going from two to four to eight advisors [1] — two or three well-chosen calls capture most of the available signal. See The Calibration Call.
  • When one advisor’s take is wildly different from the rest and your instinct is to just drop it: that instinct — egocentric trimming — is exactly the heuristic this paper’s participants used, and it was measurably worse than an objective rule like taking the median or trimming against the group’s consensus rather than your own gut [1]. Write down a mechanical rule for combining opinions before you hear them, not after.
  • When you’ve already formed a strong initial read before gathering other opinions: this paper’s weighting analysis found people lean on their own prior view about twice as heavily as an equal-weighting rule would justify (0.71 vs. 0.33) [1] — advice you solicit after you’ve made up your mind is advice you’re structurally likely to discount, not use.

1
Ilan Yaniv and Maxim Milyavsky, "Using advice from multiple sources to revise and improve judgments," Organizational Behavior and Human Decision Processes 103 (2007): 104–120.
https://doi.org/10.1016/j.obhdp.2006.05.006