
If you have ever sat through an assessment vendor’s pitch, you have seen the table: a ranked list of selection methods by how well they predict job performance. At the top, general cognitive ability at r = 0.51. Work samples at 0.54. Structured interviews at 0.51. And personality, further down, with modest numbers.
That table comes from Schmidt and Hunter’s (1998) meta-analysis in Psychological Bulletin. It is the most cited paper in work psychology. It shows up in HR tech decks, in HR textbooks, in master’s theses and — let’s be honest — it showed up in our own decks too.
The problem is that in 2022 a team led by Paul Sackett published in the Journal of Applied Psychology that those numbers are systematically inflated. Not slightly: by 0.10 to 0.20 correlation points across most predictors. The ranking changes. Cognitive ability falls off the podium. And the practical conclusion about what to measure and how to combine it flips.
This post is about that: how well personality actually predicts, compared against what, and with what caveats. It is not an explainer of the model itself. If you want to know what each dimension measures, that is in our OCEAN model explained. Here we talk evidence.
The 1998 table, as everyone cites it
Schmidt and Hunter reviewed 85 years of research in personnel selection and published a hierarchy of 19 methods. These are the original values (operational validity for overall job performance, corrected for criterion unreliability and for range restriction):
| Selection method | Schmidt & Hunter (1998) |
|---|---|
| Work sample tests | 0.54 |
| General cognitive ability | 0.51 |
| Structured interviews | 0.51 |
| Job knowledge tests | 0.48 |
| Integrity tests | 0.41 |
| Unstructured interviews | 0.38 |
| Assessment centers | 0.37 |
| Biodata | 0.35 |
| Conscientiousness | 0.31 |
| Years of job experience | 0.18 |
| Years of education | 0.10 |
| Graphology | 0.02 |
| Age | −0.01 |
Two readings of this table dominated for 25 years, and both need revising.
The first: “cognitive ability is the single best predictor, full stop.” The second: “personality predicts, but it’s the poor cousin of the table.”
Neither survives intact.
What went wrong: the range restriction correction
Here comes the technical part, and it is worth understanding because it is literally the heart of the problem.
When you validate a selection test, you almost never get to do it on the applicant population. You work with incumbents, people who already passed a filter. And that group has less variance than the original pool: the low scorers never got in. That compression flattens the observed correlation between test and performance.
The classic fix is to correct: estimate how much it flattened and restore the correlation to the size it “would have had” in the full population. That is legitimate. The question is how much you correct.
Sackett, Zhang, Berry and Lievens (2022) examined the five approaches historically used to estimate that correction and found they all share the same flaw: they assume far more range restriction than actually exists. The classic meta-analyses applied corrections based on artifact distributions that would only be plausible if the original hiring criterion correlated almost perfectly with the test being validated. In the real world, that does not happen. The highest correlation the authors could find between predictors used in selection is 0.50, and most are well below that.
Translated without the jargon: for decades the field inflated its numbers by correcting for a bias much smaller than assumed.
The revised numbers (2022)
This is the table almost nobody in HR tech is showing. Revised operational validity, same predictors, same criterion:
| Selection method | Schmidt & Hunter (1998) | Sackett et al. (2022) |
|---|---|---|
| Structured interviews | 0.51 | 0.42 |
| Job knowledge tests | 0.48 | 0.40 |
| Biodata (empirically keyed) | 0.35 | 0.38 |
| Work sample tests | 0.54 | 0.33 |
| General cognitive ability | 0.51 | 0.31 |
| Integrity tests | 0.41 | 0.31 |
| Assessment centers | 0.37 | 0.29 |
| Conscientiousness — contextualized | — | 0.25 |
| Emotional stability — contextualized | — | 0.23 |
| Extraversion — contextualized | — | 0.21 |
| Conscientiousness — general | 0.31 | 0.19 |
| Unstructured interviews | 0.38 | 0.19 |
| Years of job experience | 0.18 | 0.07 |
Three things worth flagging.
One: the structured interview came out on top. Not cognitive ability. The best-ranked predictor today is a well-designed interview, with fixed questions and scoring criteria defined in advance. Which reinforces something we have argued for a while in interviews vs assessments: the interview was never the problem. The unstructured interview was — and it collapses from 0.38 to 0.19.
Two: cognitive ability dropped from 0.51 to 0.31. It is still a useful predictor. It is no longer the undisputed king. And that shift matters, because for 25 years the large adverse impact of cognitive tests was justified with “but it predicts the most.” At 0.31, that argument loses a lot of force.
Three: personality dropped too, but dropped less. And — this is the key part — how much it dropped depends on how you measure it.
The finding that changes everything for personality: contextualize it
Look at the two conscientiousness rows in the table:
- Conscientiousness measured in general (“I am an organized person”): 0.19
- Conscientiousness contextualized to work (“at work, I am an organized person”): 0.25
Same trait, different frame of reference, roughly 30% better prediction. And the pattern repeats across every dimension: emotional stability goes from 0.09 to 0.23, extraversion from 0.10 to 0.21, agreeableness from 0.10 to 0.19, openness from 0.05 to 0.12.
This is not a methodological footnote. It is the difference between a personality test that belongs in a hiring process and one that does not. A clinical or self-discovery inventory asks how you are in life. A selection instrument has to ask how you are at work, because people behave differently in different contexts and the criterion you want to predict is a work criterion.
If a vendor offers you a personality assessment and cannot tell you whether its items are contextualized to the workplace, you already know which of the two rows they are standing on.
What Barrick and Mount saw first
Before all of this debate there was the foundational meta-analysis: Barrick and Mount (1991), in Personnel Psychology. They crossed the five Big Five dimensions against three performance criteria (job proficiency, training proficiency and personnel data) across five occupational families: professionals, police, managers, sales and skilled/semi-skilled.
The result that made the paper a classic:
Conscientiousness predicts performance across all five occupational families and all three criteria, with estimated true score correlations between 0.20 and 0.23. It is the only dimension that does. No other trait is universal.
The rest are conditional:
- Extraversion predicts in the two socially loaded occupations: managers and sales.
- Openness to experience predicts training proficiency (ρ = 0.25) — who learns faster.
- The remaining dimensions contribute in specific contexts, at magnitudes of 0.10 or less.
This is the exact opposite of the magical thinking much of the market sells. There is no single “winning” personality profile. There is one trait that helps almost everywhere — conscientiousness — and several that help depending on the role. Which is why an ideal role profile can’t be copied from a generic template.
Does personality add anything on top of cognitive ability?
This is the question that actually matters for designing a process, and it has a clear answer: yes, for a simple reason.
Personality and cognitive ability are nearly independent. Average correlations between ability and personality measures hover around 0.10 (Meriac et al., 2008, as cited in Sackett et al., 2022), and Schmidt and Hunter already reported that integrity and conscientiousness tests correlate approximately zero with general cognitive ability. Two predictors that do not overlap carry different information about the same person.
In the 1998 table, adding conscientiousness to a cognitive test lifted the combined validity from 0.51 to 0.60: a gain of 0.09, an 18% improvement. Adding an integrity test — which is largely a conscientiousness measure — lifted it to 0.65, the strongest combination in the entire paper.
Now, intellectual honesty: those incremental validity figures are computed on the inflated 1998 validities. Sackett’s team did not republish an equivalent increments table, so there is no direct replacement. What they did publish in the follow-up paper (Sackett et al., 2023, in Industrial and Organizational Psychology) is even more interesting for our purposes.
They compared every possible combination of 1, 2, 3, 4 and 5 predictors. The mean validity of those composites is 0.51 using the old values and 0.47 using the revised values. In other words: individual predictors deflated a lot, but composites barely moved. Well-designed selection practice does not collapse under the revision.
And the most provocative data point in the paper: if you assign zero weight to cognitive ability in a six-predictor composite, the old estimates said you lost 0.20 of validity (from 0.66 to 0.46). Under the revised estimates, you lose just 0.05 (from 0.61 to 0.56). A composite without a cognitive test can be nearly as valid, with far less adverse impact.
What r = 0.25 actually means
Time to cool the enthusiasm, which is the part we care most about.
A correlation of 0.25 between conscientiousness and performance does not mean you can predict who will perform. It means that if you hire a lot of people using that criterion, on average you will do better than hiring at random. In any single decision, there is a great deal of variance the trait does not explain.
Put as bluntly as possible: personality does not tell you who your best employee will be. It moves the odds in your favor. Nothing more, and nothing less.
Anyone selling you a personality assessment as a crystal ball is overselling. And anyone telling you personality “doesn’t work” because 0.25 is low is ignoring that the best-ranked predictor in the entire field — the structured interview — reaches 0.42, and that no selection tool, none, goes higher.
That is the honest conclusion of the current state of the art: no single predictor is enough. The right question is not “which test is best,” it is “which signals do I combine, and how.”
What to do with this in your process
Five concrete decisions that follow from the evidence:
One. Structure your interviews. It is the best-ranked predictor (0.42) and you are probably doing it wrong. Fixed questions, same order, scoring criteria defined before you meet the first candidate. The open-ended chat is worth half as much (0.19).
Two. Demand personality measures contextualized to work. 0.25 versus 0.19 is the difference between an instrument that contributes and one that barely moves the needle. Ask your vendor about the frame of reference of the items.
Three. Combine signals that do not overlap. Personality, values, job knowledge evidence and a structured interview measure different things. That is where the real gain lives: in the composite, not in the star test.
Four. Drop the predictors that do not predict. Years of experience collapsed from 0.18 to 0.07. If your first CV filter is still “minimum five years in the industry,” you are screening on one of the weakest predictors available.
Five. Demand to know how the signals are combined. If you get a single score, you have a right to know what each component weighs. That is the basis of our insistence on algorithmic transparency in HR tech: a number you cannot audit is not science, it is faith.
Where this leaves us
We are a company that sells personality assessments. Publishing a post that says “personality predicts moderately” is unusual. We do it because the alternative — inflating the number, hiding the 2022 revision, citing Schmidt and Hunter as if nothing happened — is exactly what the rest of the category does, and it is unsustainable.
Our instrument measures OCEAN+ with items contextualized to the workplace, crosses it with organizational values and produces a Fit Score. That Fit Score is a composite, not a single test, precisely because the evidence says no individual predictor is enough. And the same engine evaluates candidates and the team you already have, which is the logic we lay out in one engine for the whole talent lifecycle.
What we will not tell you is that we predict the future. We move the odds in your favor with the best available evidence, and we show you where every number comes from.
If you want the full conceptual frame for why personality and values belong together, it is in the complete culture fit guide. If you want to understand why MBTI does not enter this conversation at all — spoiler: its predictive validity does not even reach the floor of this table — that is in why MBTI does not work for hiring.
And if you already measure candidates with science but evaluate your current team by gut feel, Talento Index runs on exactly the same engine.
Sources
- Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1-26.
- Meriac, J. P., Hoffman, B. J., Woehr, D. J., & Fleisher, M. S. (2008). Further evidence for the validity of assessment center dimensions: A meta-analysis of the incremental criterion-related validity of dimension ratings. Journal of Applied Psychology, 93(5), 1042-1052. (Cited in Sackett et al., 2022, for the ability-personality correlations of around 0.10.)
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040-2068.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology, 16(3), 283-300.
- Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262-274.
Related OCEAN+ profiles
Discover which personality dimensions to look for in each role.
Related Articles
Culture and Retention: What the Data Really Says
Values fit predicts staying better than self-reported satisfaction. What the person-organization fit meta-analyses actually show, with nothing exaggerated.
Self-Awareness: The Predictor of Whether Coaching Works
The gap between how someone sees themselves and how their team sees them predicts whether feedback lands or bounces. What the science says, and how to use it.
MBTI for Hiring: Why It Doesn't Work
50% of people get a different MBTI type when they retake the test. Here's what psychology says about its validity — and why it matters for your hiring process.