The Validity and Utility of Selection Methods in Personnel Psychology: Practical and Theoretical Implications of 85 Years of Research Findings

Paper · 1998

Schmidt and Hunter's meta-analysis of 85 years of personnel-selection research, ranking predictors of job performance by validity — general mental ability at the top for decades. The estimates Sackett et al. (2022) later revised, demoting cognitive-ability tests below structured interviews.

Published
1998

The question

By 1998, eighty-five years of studies had reported wildly different validity numbers for the same selection method on what looked like the same job. Frank L. Schmidt and John E. Hunter ask which of 19 personnel-selection procedures actually predict future job performance, and how much validity a second measure adds once an employer is already using the single best one available.

The method

A meta-analysis of decades of prior validity studies [1], correcting each one for two statistical artifacts: measurement error in the job-performance ratings used as the criterion, and range restriction, the fact that people who get hired are a narrower slice of ability than the full applicant pool. Schmidt and Hunter call the resulting figures “operational” or “true” validity [1], a correction that raises the numbers relative to raw, uncorrected study results, and one later study would challenge directly.

The findings

General mental ability (GMA) is the single best predictor of job performance for hiring inexperienced workers, with a corrected validity of .51 for the medium-complexity jobs that make up 62% of the U.S. workforce [1]. Structured interviews (.51) clear unstructured ones (.38) by a wide margin [1]. No single addition to a GMA test beats an integrity test: the pairing reaches a composite validity of .65, with GMA plus a structured interview close behind at .63 [1]. Job experience only predicts performance for roughly its first five years on a job, after which more tenure adds almost nothing [1]. At the bottom of the list, graphology has essentially no validity at all, indistinguishable from hiring at random [1].

The limits

Schmidt and Hunter name their own limitation: they tested only two-predictor combinations built around GMA, not three-predictor stacks or combinations that exclude GMA entirely, and they explicitly set differential validity and predictive fairness across gender and racial subgroups outside the paper’s scope. A larger limit came later, from outside the paper. Sackett, Zhang, Berry, and Lievens’ 2022 reanalysis found that the range-restriction correction at the center of Schmidt and Hunter’s method had been systematically overcorrecting, inflating validity estimates for GMA and several other predictors well above what the underlying data support. Treat this paper as the method that shaped three decades of hiring practice, not as today’s number.

  • When you’re building or auditing a hiring loop and deciding which parts of it to trust: this paper’s ranking says structured interviews and cognitive tests carry far more signal than unstructured chats or credential checklists, so running a Calibration Call built around a fixed rubric is buying real predictive power, not just the appearance of fairness.
  • When you’re tempted to add a “gut feel” conversation after a candidate has already cleared a validated test: the incremental-validity math here shows a second predictor’s payoff depends on how much it overlaps with the first, not just its own validity, so an unstructured chat bolted onto a strong test can add almost nothing.
  • When a candidate’s long tenure elsewhere is the deciding factor between two finalists: past roughly five years, additional experience stops predicting future performance in this data, so favoring 12 years over 6 is rewarding a number that quit predicting years ago.
  • When you’re citing “GMA predicts job performance at .51” as settled fact: that headline number is exactly the range-restriction-corrected estimate Sackett et al.’s 2022 reanalysis found overcorrected, so treat it as historically important, not current.

1
Frank L. Schmidt and John E. Hunter, "The Validity and Utility of Selection Methods in Personnel Psychology: Practical and Theoretical Implications of 85 Years of Research Findings," Psychological Bulletin 124, no. 2 (1998): 262–274.
https://doi.org/10.1037/0033-2909.124.2.262