The question
By 1998, eighty-five years of studies had reported wildly different validity numbers for the same selection method on what looked like the same job. Frank L. Schmidt and John E. Hunter ask which of 19 personnel-selection procedures actually predict future job performance, and how much validity a second measure adds once an employer is already using the single best one available.
The method
A meta-analysis of decades of prior validity studies [1], correcting each one for two statistical artifacts: measurement error in the job-performance ratings used as the criterion, and range restriction, the fact that people who get hired are a narrower slice of ability than the full applicant pool. Schmidt and Hunter call the resulting figures “operational” or “true” validity [1], a correction that raises the numbers relative to raw, uncorrected study results, and one later study would challenge directly.
The findings
General mental ability (GMA) is the single best predictor of job performance for hiring inexperienced workers, with a corrected validity of .51 for the medium-complexity jobs that make up 62% of the U.S. workforce [1]. Structured interviews (.51) clear unstructured ones (.38) by a wide margin [1]. No single addition to a GMA test beats an integrity test: the pairing reaches a composite validity of .65, with GMA plus a structured interview close behind at .63 [1]. Job experience only predicts performance for roughly its first five years on a job, after which more tenure adds almost nothing [1]. At the bottom of the list, graphology has essentially no validity at all, indistinguishable from hiring at random [1].
The limits
Schmidt and Hunter name their own limitation: they tested only two-predictor combinations built around GMA, not three-predictor stacks or combinations that exclude GMA entirely, and they explicitly set differential validity and predictive fairness across gender and racial subgroups outside the paper’s scope. A larger limit came later, from outside the paper. Sackett, Zhang, Berry, and Lievens’ 2022 reanalysis found that the range-restriction correction at the center of Schmidt and Hunter’s method had been systematically overcorrecting, inflating validity estimates for GMA and several other predictors well above what the underlying data support. Treat this paper as the method that shaped three decades of hiring practice, not as today’s number.