Science 10 min read

Pre-Employment Assessments: Which Ones Actually Predict Performance

The market sells dozens of tests and almost none publish their validity. What each family of instruments measures and which ones predict performance.

Clara Bellini

Clara Bellini

Head of People Science

Share:
Pre-Employment Assessments: Which Ones Actually Predict Performance
pre-employment assessment psychometric testing hiring predictive validity selection

A company shopping for pre-employment assessments will find twenty vendors in one afternoon. Almost all of them talk about accuracy, colors, and profiles. Very few publish a validity coefficient.

That asymmetry is the whole problem. The useful question is not which test is best, but what evidence each family of instruments has for the only thing that matters, which is predicting performance on the job.

What makes a questionnaire a psychometric instrument

Three properties, and you can ask for all three in writing.

Reliability. The same person gets a similar result on a retake. An instrument that reassigns half your people to a different category in five weeks is not measuring a stable trait.

Validity. The result relates to something outside the test itself. Criterion validity is the one that matters in hiring, and it means correlation with independently measured performance.

Norms. The population the score is compared against. A percentile with no declared reference population is not a percentile.

If a vendor cannot give you all three, what you have is a shared vocabulary for talking about people. That can be useful in a team workshop. It is not a hiring criterion.

What the updated evidence says

For twenty years the standard reference was Schmidt and Hunter’s 1998 review of 85 years of selection research. In 2022, Sackett, Zhang, Berry and Lievens revisited it in the Journal of Applied Psychology and corrected a methodological problem. The range-restriction corrections in use had been inflating the coefficients.

The result pushed estimates down by 0.10 to 0.20 points across most of the board. The ranking moved less than expected, but first place changed hands.

The mean operational validities that paper reports, in order:

  • Structured interview · 0.42. It came out first. Also with wide spread around that mean, so implementation quality carries weight.
  • Job knowledge tests · 0.40.
  • Empirically keyed biodata · 0.38.
  • Work samples · 0.33.
  • General cognitive ability · 0.31. No longer the top predictor, which had been the consensus since 1998.
  • Interests · 0.24.

On personality, the relevant finding is one of nuance. General inventories perform worse than contextualized ones, meaning those that ask about behavior at work rather than in life generally. It is the difference between “I am organized” and “I am organized when working on a team under pressure.”

The five families, and what each one is for

Trait personality. The Big Five is the model with the most accumulated support. It does not predict as much as a structured interview, and it predicts different things — how the person will work, not how much they know. It is the layer that explains fit with the role and with the team. The breakdown is in Big Five predictive validity.

Cognitive ability. A good predictor, especially in complex roles. Also the one carrying the most adverse-impact risk in several jurisdictions, which forces you to document its relationship to the job.

Work samples and knowledge tests. They measure whether the person can do the task. Expensive to build and hard to scale, which is why they are rare. Where they exist, they work.

Situational judgment. They present a work situation and ask you to pick a response. They measure judgment applied to a context, and they depend heavily on that context resembling the real one.

Behavioral typologies. DISC, MBTI, and their derivatives. They are built to give people a shared vocabulary, not to rank candidates. The detail is in DISC vs OCEAN and in MBTI for hiring.

The six questions to ask before buying

  1. What construct does it measure, and under which model? If the answer is the product name, there is no answer.
  2. What is the test-retest reliability, and at what interval? A number without an interval says nothing.
  3. Against what criterion was it validated? Supervisor-rated performance, tenure, production indicators. And with what sample size.
  4. What is the normative population? A Spanish norm applied to a Mexican sample shifts the percentiles. Percentiles are also locked to the instrument they were built on, so they do not transfer between tests.
  5. Can an individual result be explained? If a candidate asks why they ranked below someone else, somebody has to be able to answer. That argument is developed in algorithmic transparency in HR tech.
  6. Is adverse impact monitored? No instrument erases bias. What a serious process does is measure it, show it, and correct when it appears. How it gets measured is in hiring bias.

The mistake of buying a single instrument

The most common reading of the validity table is wrong. People look for the highest number and buy that.

Predictors do not compete, they add up. A structured interview and a personality inventory explain different parts of performance, and together they predict more than either alone. The design that pays off has two layers. The instrument measures stable dispositions before the interview, and the structured interview explores specific competencies through lived examples. The guide is in competency-based interviews.

What does not work is stacking three behavioral typologies and calling it triangulation.

Where Talen.to sits

The engine measures on the Big Five and adds two layers a generic test does not have. Fit against the role profile, with expected ranges per dimension, and fit against the declared values of the company doing the hiring. That cross is what a catalog instrument cannot do, because it knows nothing about your company.

The archetype reading that accompanies the report is qualitative. It is there to have a conversation about a profile, and it is not used to rank candidates.

If you want to see what the instrument produces before deciding, the Talent Diagnostic takes about ten minutes and requires no signup. And the model breakdown is in the OCEAN model explained.

About the author

Clara Bellini

Clara Bellini

Marketing Director

Marketing Director @ Talen.to. Former agency, now product. Believer in data > intuition and culture > everything.

Related OCEAN+ profiles

Discover which personality dimensions to look for in each role.

Related Articles

Talen.to

Ready to improve your hiring?

Start assessing culture fit with science.