The question
Schmidt and Hunter’s (1998) meta-analytic rankings of selection-procedure validity anchored the field for over two decades. Paul R. Sackett, Charlene Zhang, Christopher M. Berry, and Filip Lievens ask whether the range-restriction correction behind those numbers (the statistical adjustment for the fact that a validity study can only measure the people who were actually hired, a narrower and higher-scoring slice than the full applicant pool) was ever built the way the field assumed, or whether decades of meta-analyses had quietly overcorrected.
The method
A predictor-by-predictor revisit of Schmidt and Hunter’s 1998 summary of 19 selection procedures, scrutinizing the range-restriction correction each underlying meta-analysis used and substituting a corrected or uncorrected estimate wherever the original correction wasn’t credible [1].
The findings
The overcorrection has a specific mechanical cause. A range-restriction correction assumes a predictive study design (test applicants, then hire and validate later), where restriction on the predictor is direct. But 74 to 98 percent of the underlying studies behind most procedures were concurrent designs that tested people already on the job [1], hired on the basis of something other than the predictor now being scored. In a concurrent study, restriction can only be indirect, and indirect restriction is mathematically small: unless the original hiring method happens to correlate very highly with the predictor under study, which it rarely does, the correction factor runs from zero to about 10 percent [1]. Applying a predictive-study-sized correction to a population that was mostly concurrent is where the overcorrection came from.
Run that fix through Schmidt and Hunter’s 19 procedures and cognitive ability’s validity falls from .51 to .31, work samples from .54 to .33 (the largest point drop of any predictor), and structured interviews from .51 to .42, the smallest drop of the three [1]. That smaller drop reorders the field: structured interviews now rank as the single strongest predictor of job performance, ahead of cognitive ability [1]. The revised top five keeps nearly the same cast as Schmidt and Hunter’s, but its mean validity is .37 against their .49: the ordering barely moved, and the confidence behind the numbers did.
The paper is careful to distinguish “these tools don’t work” from “these tools don’t work as well as claimed”: in its own words, “our predictors are useful; but the predictive relationships are considerably weaker than previously thought” [1]. Where no trustworthy correction was possible, the authors applied none at all rather than risk another overcorrection, and even under the most generous plausible range-restriction assumptions, cognitive ability’s estimate rises only to .36, still far below Schmidt and Hunter’s .51 [1].
The limits
The authors flag two limits on their own findings. Their reanalysis still assumes predictive studies (testing applicants) and concurrent studies (testing incumbents) measure the same underlying validity, an assumption they note is itself unresolved: if faking or effort differs between a real hiring decision and a research setting, both the old and new numbers could be biased in ways the correction doesn’t touch. And the authors are explicit that their estimates are the best correction the available data supports today, not a final word: “we do not put a stake in the ground regarding the meta-analytic estimates we put forward here” [1].