The competency-based interview has a reputation problem. Everyone says they use it and very few people use it as designed.
Here is what usually happens. Someone drafts a list of competencies, each interviewer picks whichever questions feel right, and at the end the team gets together to say what they thought of each candidate. That is not a structured interview. That is a conversation with a list nearby.
The difference matters, because structure is exactly the part with evidence behind it.
What the evidence says
In Schmidt and Hunter’s review of eighty-five years of selection research, the structured interview lands among the strongest predictors of job performance, well above the unstructured interview. The gap between the two is not the interviewer or the questions. It is that one compares and the other does not.
An unstructured interview produces impressions that are not comparable to each other. If you asked one candidate about a team conflict and another about their proudest achievement, you do not have two data points. You have two anecdotes.
The four elements that make it work
1. Same questions, same order, every candidate. This is the hardest element to hold and the one that contributes most. The moment one interviewer improvises a follow-up they did not ask another candidate, comparability breaks.
2. Questions about past behavior, not hypotheticals. “Tell me about a time you had to ship without full information” works better than “what would you do if…”. The second one measures the ability to imagine a correct answer.
3. A scale defined before you interview. For each competency, written in advance, what counts as a 1, a 3, and a 5. Without this, every interviewer carries their own scale and the team average means nothing.
4. Individual scores before the debrief. Everyone scores alone, and only then do you discuss. Discuss first and the most senior opinion pulls the rest along. The result is one opinion with four signatures.
The question format
STAR is still the most practical scaffold: Situation, Task, Action, Result. The part everyone skips is the A.
When someone tells you about an achievement, they tend to narrate the context and the result. What you need to know is what that person specifically did, as distinct from what the team did. The useful follow-up is blunt. “Which part did you do?” And then: “What would have happened if you had not been there?”
It pays to prepare two fixed follow-ups per competency, identical for every candidate. That way you go deeper without breaking the structure.
Where it breaks in practice
Too many competencies. With eight competencies and one hour, you get seven minutes each. Four competencies explored properly beat eight mentioned.
Competency gets confused with trait. “Conscientiousness” is a personality trait and it gets measured with an instrument, not a question. “Prioritizing with incomplete information” is a competency and it gets explored through a lived case. Asking someone whether they are conscientious measures how well they know the expected answer.
Nobody calibrates the interviewers. Two people working from the same written scale will score differently the first time. One calibration session, on a recorded or invented case, aligns the scales before the real interviews start.
What the interview cannot give you
Even done well, the structured interview measures what the person did and can narrate. It does not measure stable dispositions, and it does not measure fit with your culture.
That is why the design that works best has two layers. The assessment measures traits with a model that has published validity, and compares them against the role profile and against your company’s actual values. The structured interview explores the competencies specific to the job, through lived cases.
Each layer contributes what the other cannot see. The interview alone leaves you with well-organized impressions. The assessment alone leaves you unsure whether this person can do this particular job.
Multi-evaluator competency assessment covers how the two layers combine when the role justifies it.
If you want to see what the first layer measures before building the second, the Talent Diagnostic takes about ten minutes and needs no signup.
Related OCEAN+ profiles
Discover which personality dimensions to look for in each role.
Related Articles
The Job Profile: How to Define It With Data, Not a Template
The job description you publish and the profile you decide against are two different documents. Most companies have the first and none of the second.
Talent Review and the 9-Box: How to Fill It With Data
The grid crosses performance with potential. You already measure performance. Potential is usually the manager's opinion, and that is where the exercise breaks.
Performance Reviews: The Methods and How to Choose
Rating scales, 360s, OKRs, forced ranking. Each answers a different question, and picking wrong is the most common reason the whole process changes nothing.