Vetting & Technical Assessment · 3 min read
Structured Interview Scorecards That Actually Work
How to build an interview scorecard engineers will use, why independent scoring matters more than the questions, and how to spot a rubric that needs rewriting.
A useful interview scorecard defines what weak, adequate and strong answers look like before anyone is interviewed, and requires each interviewer to score independently before discussion. The rubric matters less than the discipline of scoring blind, which is what removes anchoring on the loudest voice.
Why unstructured interviews feel informative and are not
A conversation with a candidate produces a strong impression, and that impression feels like evidence. It is mostly noise. Without a rubric agreed in advance, an interviewer's judgement drifts toward candidates who communicate in a familiar register, who share a background, or who happened to be interviewed on a good day. This is one of the better-established findings in selection research and it survives every attempt to argue that a particular team is exempt.
The failure is not that interviewers are careless. It is that human judgement in the absence of structure reliably substitutes an easy question for a hard one. 'Would this person do the job well' is hard; 'do I like talking to this person' is easy, and the second gets answered while the first appears to have been.
Structure does not remove judgement. It constrains where judgement is applied, forcing it onto specific observable behaviours rather than onto a global impression formed in the first four minutes.
What belongs on the scorecard
Each criterion needs three things: a name, a definition of what you are assessing, and concrete descriptions of weak, adequate and strong responses. The concrete descriptions are the part teams skip and the part that does the work, because 'strong communication' means something different to every interviewer while 'explains a trade-off without being asked to' does not.
Keep the number of criteria small. Four or five assessed properly beats twelve assessed superficially, and long scorecards are quietly abandoned within a quarter. The criteria should also map to the actual failure modes of the role: if engineers on your team most often struggle with ambiguous requirements, that belongs on the card whatever the job description says.
Include a space for evidence, not just a score. Requiring an interviewer to write down what the candidate actually said in support of a score is the cheapest available check against unconscious substitution, and it makes the later discussion about evidence rather than impressions.
Independent scoring is the part that matters
If you change only one thing, change the ordering. Every interviewer records their scores and evidence before any discussion happens. This sounds procedural and it is the single highest-return intervention in technical hiring, because discussion-first reliably produces convergence on whoever speaks with the most authority rather than on the strongest evidence.
The value shows up in the disagreements. When two competent interviewers score the same candidate very differently on the same criterion, that is information. Sometimes it means the candidate performed differently in two sessions, which is worth knowing. More often it means the criterion is ambiguous and two reasonable people read it differently, which means the rubric needs rewriting.
That second case is how an assessment process improves rather than merely persisting. Teams that never surface rubric ambiguity keep applying it inconsistently for years while believing they have a structured process.
Signs your rubric needs rewriting
- Two experienced interviewers consistently score the same answers differently
- Everyone scores everything in the middle band, which usually means the anchors are vague
- Interviewers fill in scores after the debrief rather than before it
- A criterion has never once distinguished between candidates you hired and those you did not
- The evidence boxes are consistently empty or contain restatements of the score
- Nobody can remember what one of the criteria was actually intended to capture
- Scores cluster by interviewer rather than by candidate
Calibration, and how to run it cheaply
A rubric drifts even when it is followed, because interviewers recalibrate against whoever they saw most recently. The correction is a periodic calibration session: take three past submissions with the scores removed, have everyone score them independently, then compare. It takes an hour a quarter and it catches drift that would otherwise be invisible.
This is also the cheapest way to onboard a new interviewer. Handing someone five past assessments with the recorded scores hidden, then comparing their judgement against the actual outcomes, teaches the bar far faster than shadowing does and produces a documented indication of whether they are ready to assess independently.
Keep the scored submissions. Over a year they become a calibration set that answers a question most teams cannot answer at all: has our bar moved, and in which direction.
Part of the Vetting & Technical Assessment cluster · Read the pillar page
More in Vetting & Technical Assessment
Vetting & Technical Assessment
How to Vet a Software Developer: A Practical Framework
An evidence-based framework for assessing engineers: work samples over puzzles, structured interviews over conversations, and calibration to stop drift.
3 min read
Vetting & Technical Assessment
Work Sample Tests for Engineers: Design and Scoring
How to design a work sample that predicts real performance, keep it under three hours, and score it consistently across reviewers without arguing.
3 min read
Vetting & Technical Assessment
Assessing Technical Debt in an Inherited Codebase
How to evaluate a codebase you did not write, which signals predict future pain, and how to brief an engineer joining a system nobody fully understands.
4 min read
Frequently asked questions
What should be on an engineering interview scorecard?
Four or five criteria, each with a name, a definition, and concrete descriptions of weak, adequate and strong responses, plus space to record the evidence behind each score rather than only the score itself.
Why score independently before discussing?
Because discussion-first produces convergence on whoever speaks with most authority rather than on the strongest evidence. Independent scoring preserves genuine disagreement, which is informative about both the candidate and the rubric.
How many interview criteria should there be?
Four or five assessed properly. Longer scorecards are assessed superficially and quietly abandoned within a quarter, which leaves you with the appearance of structure and none of the benefit.
How often should we calibrate?
About once a quarter, using past submissions with the scores removed. It takes an hour, catches drift that is otherwise invisible, and doubles as the fastest way to onboard a new interviewer to your actual bar.
Should the scorecard change for senior candidates?
Yes, and it should change in kind rather than only in degree. Weight judgement about what not to build and the quality of their questions above implementation fluency, since that is what distinguishes senior capability.