The question
When you have two estimates from two people, is averaging them better than picking the better person, and do people know?
The method
Two experiments in which participants predicted how well averaging would perform against the individual judges, across materials whose bracketing rates the authors varied [1].
The findings
The mathematics is not in dispute and is the reason the paper matters: “if the estimates of two imperfect judges ever fall on either side of the truth, which we term bracketing, averaging must outperform the average judge” [1]. That makes averaging “an effective way of improving accuracy when combining expert judgments, integrating group members’ judgments, or using advice to modify personal judgments” [1].
People do not believe it. The common intuition is the one the authors set out to test, “falsely believing that the average of two judges’ estimates would be no more accurate than the average judge” [1] — which, as they point out, “is the worst possible performance of averaging, and occurs only when judges never bracket the truth” [2]. “Because the judges in our stimuli bracketed the truth between 24% and 90% of the time, averaging was always superior to the average judge (and, in these cases, superior to both judges)” [2].
The misconception is fixable when the structure is visible: “Experiment 1 confirmed that this misconception was common across a range of bracketing rates. Experiment 2 demonstrated that the effectiveness of averaging can be recognized when bracketing was made transparent” [1]. Their explanation for why nobody learns this by living is the most quotable line in the paper: “life abounds with data on individual accuracy and error perhaps more so in an individualistic culture — but rarely reveals interpersonal patterns of error” [2].
The limits
The estimates being combined are numerical quantities with a knowable truth. A calibration call does not produce numbers, and “average two accounts of what great looks like” is a metaphor for what this paper proves, not the thing it proves. Take the argument for why several imperfect sources beat one good one; do not take a formula.