The question
Do intelligence and aptitude tests actually predict how well someone will do the job, and if they don’t, what should replace them?
The method
A skeptical review of the validity evidence the testing movement rested on rather than a new study — Thorndike and Hagen’s correlations, Ghiselli’s fifty-year survey, Terman’s gifted-children follow-ups — followed by six proposals for a different kind of test [2].
The findings
The evidence was thinner than the field believed. Thorndike and Hagen “obtained 12,000 correlations between aptitude test scores and various measures of later occupational success on over 10,000 respondents and concluded that the number of significant correlations did not exceed what would be expected by chance” [1]. Ghiselli’s review, the result most often cited the other way, reported that “general intelligence tests correlate .42 with trainability and .23 with proficiency across all types of jobs” [1] — but never said how proficiency had been measured, which McClelland argues is the whole problem: the measures “depend heavily on the credentials the man brings to the job” [1], so “the correlation between intelligence test scores and job success often may be an artifact” [1] of both being tied to social class.
His replacement is criterion sampling: test the behaviour you care about instead of a proxy for it. “If you want to know how well a person can drive a car (the criterion), sample his ability to do so by giving him a driver’s test” [2]. Applied to hiring, that means leaving the office: “If you want to test who will be a good policeman, go find out what a policeman does. Follow him around, make a list of his activities, and sample from that list in screening applicants” [2].
And the warning that matters most to anyone gathering a standard from other people: “do not rely on supervisors’ judgments of who are the better policemen because that is not, strictly speaking, job analysis but analysis of what people think involves better performance” [2].
The limits
McClelland is explicit that the second half of the paper is not evidence. “My goal is to brainstorm a bit on how things might be different, not to present hard evidence that my proposals are better than what has been done to date” [2]. The critique of the validity literature is documented; the criterion-sampling programme is a proposal, and the field spent the following decades testing it.