Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range

Paper · 2022

Sackett, Zhang, Berry, and Lievens' 2022 meta-analysis correcting a systematic statistical overcorrection in decades of personnel-selection research, re-ranking structured interviews (r = .42) above cognitive-ability tests (r = .31) as predictors of job performance. The revised numbers behind why MC treats structured interviews as the strongest lever in a hiring process, not cognitive testing.

Published
2022

The question

Schmidt and Hunter’s (1998) meta-analytic rankings of selection-procedure validity anchored the field for over two decades. Paul R. Sackett, Charlene Zhang, Christopher M. Berry, and Filip Lievens ask whether the range-restriction correction behind those numbers (the statistical adjustment for the fact that a validity study can only measure the people who were actually hired, a narrower and higher-scoring slice than the full applicant pool) was ever built the way the field assumed, or whether decades of meta-analyses had quietly overcorrected.

The method

A predictor-by-predictor revisit of Schmidt and Hunter’s 1998 summary of 19 selection procedures, scrutinizing the range-restriction correction each underlying meta-analysis used and substituting a corrected or uncorrected estimate wherever the original correction wasn’t credible [1].

The findings

The overcorrection has a specific mechanical cause. A range-restriction correction assumes a predictive study design (test applicants, then hire and validate later), where restriction on the predictor is direct. But 74 to 98 percent of the underlying studies behind most procedures were concurrent designs that tested people already on the job [1], hired on the basis of something other than the predictor now being scored. In a concurrent study, restriction can only be indirect, and indirect restriction is mathematically small: unless the original hiring method happens to correlate very highly with the predictor under study, which it rarely does, the correction factor runs from zero to about 10 percent [1]. Applying a predictive-study-sized correction to a population that was mostly concurrent is where the overcorrection came from.

Run that fix through Schmidt and Hunter’s 19 procedures and cognitive ability’s validity falls from .51 to .31, work samples from .54 to .33 (the largest point drop of any predictor), and structured interviews from .51 to .42, the smallest drop of the three [1]. That smaller drop reorders the field: structured interviews now rank as the single strongest predictor of job performance, ahead of cognitive ability [1]. The revised top five keeps nearly the same cast as Schmidt and Hunter’s, but its mean validity is .37 against their .49: the ordering barely moved, and the confidence behind the numbers did.

The paper is careful to distinguish “these tools don’t work” from “these tools don’t work as well as claimed”: in its own words, “our predictors are useful; but the predictive relationships are considerably weaker than previously thought” [1]. Where no trustworthy correction was possible, the authors applied none at all rather than risk another overcorrection, and even under the most generous plausible range-restriction assumptions, cognitive ability’s estimate rises only to .36, still far below Schmidt and Hunter’s .51 [1].

The limits

The authors flag two limits on their own findings. Their reanalysis still assumes predictive studies (testing applicants) and concurrent studies (testing incumbents) measure the same underlying validity, an assumption they note is itself unresolved: if faking or effort differs between a real hiring decision and a research setting, both the old and new numbers could be biased in ways the correction doesn’t touch. And the authors are explicit that their estimates are the best correction the available data supports today, not a final word: “we do not put a stake in the ground regarding the meta-analytic estimates we put forward here” [1].

  • When you’ve read a popular retelling of “structured interviews plus GMA testing, validity around .6” and are sizing a hiring process around it: that figure is Schmidt and Hunter’s 1998 number, not this paper’s. The revised structured-interview estimate is .42 and cognitive ability’s is .31: still useful predictors, not the near-certain signal the older number implied.
  • When you’re deciding how much weight a cognitive-ability test should carry relative to a well-run Recruiting & hiring process built on structured interviews: the revised numbers rank structured interviews above cognitive ability on point estimate alone (.42 vs .31), the reverse of Schmidt and Hunter’s ordering.
  • When someone argues a validated test is airtight because it “corrects for range restriction”: ask whether the correction was built from predictive studies (applicants) or concurrent studies (existing employees) and applied uniformly across both. That mismatch is the specific statistical move this paper argues inflated decades of numbers.
  • When you’re weighing test batteries purely on reported validity coefficients without checking whether the underlying meta-analysis is pre- or post-2022: a number lifted from Schmidt and Hunter’s 1998 summary without this paper’s correction is very likely an overestimate for the reasons above.

1
Paul R. Sackett, Charlene Zhang, Christopher M. Berry, and Filip Lievens, "Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range," Journal of Applied Psychology 107, no. 11 (2022): 2040–2068.
https://doi.org/10.1037/apl0000994