Hiring: scorecard first, and defending the bar
What it is
Hiring as a lead's discipline has three parts, and the first one determines the other two:
1. WRITE THE SCORECARD BEFORE THE LOOP.
What this role must be able to do, which signals prove it,
and which interview produces each signal. Written down before
anyone is interviewed.
2. INTERVIEW TO IT, CONSISTENTLY.
Same questions, same rubric, same evidence standard across
candidates. Otherwise you are measuring rapport and calling
it judgment.
3. DEFEND THE BAR, WHICH MEANS SAYING NO TO A "FINE" CANDIDATE
AND ARTICULATING WHY.
A hire you are unsure about is a no. Bar defense is a lead
responsibility, and it is unpopular exactly when it matters
most.
And a fourth that is usually treated as someone else's job and is not: onboarding. A named buddy, a shipped change in week one, and a 30/60/90 with explicit success criteria.
What this is confused with: hiring as filtering. A filter asks who to eliminate. A scorecard asks what evidence would convince you, which is a different question and produces different interviews: an interviewer with a filter looks for reasons to reject, and an interviewer with a scorecard looks for specific signals and notices when they are absent.
Also confused: a high bar and a slow process. They are independent, and the two failures are symmetric: hiring someone you were unsure about, and losing a strong candidate to a three-week gap between rounds.
The problem it solves
Without a scorecard written first, the loop measures rapport.
Five interviewers, no shared scorecard. The debrief:
"Strong. Really enjoyed talking to them."
"Solid. Nothing concerning."
"I liked them. Good culture fit."
"Fine on the coding, bit quiet."
"I'd hire them."
Four hires, one weak yes, zero evidence, and one phrase
("culture fit") that in practice usually means similarity to
the interviewer.
Six months later the same team cannot say why the hire is
struggling, because nobody ever wrote down what the role
required.
And the asymmetry that makes bar defense hard is real, not imagined:
The cost of a wrong yes:
- the person struggles, which is bad for them
- the team absorbs the gap, usually the strongest people
- 6 to 12 months to a performance process, which consumes an
enormous amount of a lead's time
- the team's own bar quietly resets, because the standard is
now visible
The cost of a wrong no:
- you interview more candidates
These are not comparable, and every process pressure pushes the
other way: the requisition is open, the team is stretched, the
recruiter has a metric, and you have just spent six hours on
someone who is "fine".
Mechanics
The scorecard, written first
For a senior backend engineer on a payments team:
MUST HAVE Signal from Bar
--------------------------------------------------------------
designs a system with system design can name
explicit failure modes failure modes
unprompted and
say what
happens in each
writes correct code under coding works, handles
ambiguity, and asks about the edge case
the ambiguity they identified
debugs production, reasons debugging / forms a
from evidence incident round hypothesis and
says what would
disprove it
operates in a domain with domain deep-dive asks about
correctness constraints idempotency,
(money) reconciliation
or audit
without prompting
disagrees productively behavioural a real example
where they
changed their
position
NICE TO HAVE
payments domain experience
Kotlin specifically (we can teach this)
NOT REQUIRED, and say so explicitly so nobody screens on it
a degree
experience at a company of our size
familiarity with our exact stack
Two properties do the work:
EACH SIGNAL IS OWNED BY EXACTLY ONE INTERVIEW. Otherwise three
people assess the same thing and nobody assesses the fourth,
which is the most common loop design failure. Assign it in
writing, in the loop invite.
THE BAR IS WRITTEN AS OBSERVABLE EVIDENCE, not as a level. "Can
name failure modes unprompted" is checkable. "Strong system
design" is a feeling.
The "not required" list is worth writing explicitly, because unwritten preferences operate anyway, and a screening step that quietly filters on a stack or a company tier is where most of a pipeline's diversity disappears before anyone is interviewed.
Structure, and why it beats rapport
SAME QUESTIONS across candidates for a given role, so you have a
comparison rather than five separate impressions.
A RUBRIC PER QUESTION with what a weak, adequate and strong
answer contains, written in advance.
EVIDENCE IN THE WRITE-UP: what the candidate said, not how you
felt. "Asked what happens if the callback is delivered twice,
then designed for it" is evidence. "Great instincts" is not.
INDEPENDENT WRITE-UPS BEFORE THE DEBRIEF, submitted, so nobody
anchors on the loudest voice. This is the single cheapest
quality improvement in a hiring loop.
WRITTEN FEEDBACK WITHIN 24 HOURS, because memory degrades fast
and a write-up produced three days later is a reconstruction.
The research finding behind all of this is consistent: unstructured interviews are weak predictors of job performance and strongly reflect interviewer similarity, and structure is what converts an interview from a social interaction into a measurement.
Defending the bar
THE RULE: a "fine" candidate is a no. If you are talking
yourself into it, that is the answer.
THE ARTICULATION, which is the part that makes it defensible:
not "I just didn't feel it"
but "the scorecard says they must be able to reason from
evidence in a production debugging scenario. In the
debugging round they guessed three times without checking
anything, and when I asked what would disprove their
hypothesis they didn't have an answer. That is the
signal we said we needed and it was absent."
That is arguable, checkable and specific, and it is what makes
a no survive pressure from a recruiter with a metric and a hiring
manager with an open requisition.
The pressures are predictable, so prepare the responses:
"We've been looking for four months."
-> "And a wrong hire costs us a year. What changed about the
role that would make this candidate right?"
"They're better than nobody."
-> Not true. The team absorbs the gap, usually the strongest
people, and the standard resets visibly.
"We can coach them up."
-> Sometimes, and only for teachable things you have named in
advance and have capacity to teach. "We can teach Kotlin"
is credible; "we can teach them to reason from evidence"
usually is not.
"The other four said yes."
-> Then say your specific evidence and let the room weigh it.
A no with evidence is a contribution; a no without it is
an obstruction, which is why the articulation matters.
And the inverse discipline, which is less discussed: do not block on a preference. A no that is really "they did not solve it the way I would have" is the same failure in the other direction, and it is why the scorecard is written before you meet anyone.
Calibrating interviewers
NEW INTERVIEWERS shadow twice, then are shadowed twice, before
running a round alone.
DISAGREEMENT IS THE TRAINING SIGNAL. When two interviewers reach
opposite conclusions from the same round, that conversation is
worth more than any training material.
TRACK OUTCOMES where you can: which interviewers' strong-hire
signals correlate with people who do well at 12 months. Small
samples, so treat it as a conversation starter rather than a
score.
ROTATE, so that the loop is not four people who think alike, and
so the load does not concentrate on the same three seniors.
Onboarding, which is part of hiring
A NAMED BUDDY. One person, named before day one, whose explicit
job is to be interruptible for six weeks. Not the lead.
A SHIPPED CHANGE IN WEEK ONE. Small, real, in production.
It proves the pipeline works for them, it forces every access
and tooling problem to surface immediately, and it is the
single strongest early signal to the new person that they are
going to be effective here.
Keep a standing list of small, safe, genuinely useful changes
for exactly this.
A 30/60/90 WITH EXPLICIT SUCCESS CRITERIA, written and shared:
30 environment working, shipped 2-3 small changes, met the
people they will work with, can describe what the team
owns
60 owns a small feature end to end, on-call shadow completed,
contributing in design reviews
90 fully in the rotation, owns a meaningful piece of work,
has raised at least one thing they think we do badly
The last one at 90 days is deliberate: a new person's outside
view has a short shelf life, and asking for it explicitly is the
only way most people will offer it.
Onboarding failures show up as hiring failures, and the distinction matters because the fix is different. A person who is struggling at four months when nobody named a buddy, they shipped nothing for three weeks, and no success criteria existed, is not a hiring mistake yet.
A worked example: a loop that was measuring rapport
A team hiring two senior engineers. Twelve months of history: 34 onsite loops, 6 offers, 5 hires, of whom 2 were struggling at the twelve-month mark and one had left.
The audit of the existing loop:
No scorecard existed. The loop was:
- two coding rounds (different interviewers, same kind of
problem)
- one system design
- one "culture / values"
- hiring manager conversation
Signal coverage, reconstructed from write-ups:
coding covered 3 times (both coding rounds
and half the design round)
system design covered once, inconsistently
debugging / production NEVER ASSESSED
domain correctness NEVER ASSESSED
disagreement assessed by 3 different people with
3 different standards
Write-up quality:
contained specific evidence 31%
contained only impressions 69%
submitted before the debrief 22%
Of the 34 loops, the correlation between "which interviewer
liked them most" and the eventual decision was the strongest
pattern in the data.
Debugging and production reasoning were never assessed, and both struggling hires were struggling on exactly that, which is not a coincidence but a direct consequence: the loop could not have detected it.
The rebuild:
1. SCORECARD, written by the lead with two senior engineers, in
90 minutes. Five must-haves, each owned by exactly one round.
2. LOOP REDESIGNED to match:
coding (1 round, was 2)
system design (1)
production debugging (1, NEW: a real incident from our own
history, with the logs and dashboards)
domain deep-dive (1, folded into the hiring manager
conversation)
behavioural, structured on disagreement and ownership (1)
3. RUBRICS per question, one page each, with weak/adequate/strong
descriptions written before the first candidate.
4. INDEPENDENT WRITE-UPS submitted before the debrief, enforced
by the scheduling tool. This one change took write-ups
containing specific evidence from 31% to 88%, because a
write-up you cannot revise after hearing others' opinions has
to stand on its own.
5. 24-HOUR FEEDBACK SLA, tracked.
The independent-write-up change was free and produced the largest single improvement in evidence quality, which is consistent with the anchoring research and was the easiest thing to sell internally because it removed no one's autonomy.
The bar defense, tested twice in the first quarter:
CANDIDATE 1: four yeses, one no (the debugging round).
The no, articulated: "the scorecard says reasons from evidence
in production. Given real logs, they proposed three causes in
four minutes without looking at anything, and when I asked
what would rule out the first one they said 'we'd have to try
it'. That is the signal we said we needed."
The recruiter noted the requisition had been open five months.
The hiring manager asked whether it was coachable. The
articulation was specific enough that the answer was "not in
the timeframe we'd need", and it was a no.
CANDIDATE 2: three yeses, one no, one weak yes.
The no was "they didn't solve the design the way I would
have", which under questioning did not map to any scorecard
line. That no was overridden, correctly, and the interviewer
was re-calibrated.
Both outcomes are the process working, and the second is the
one people forget: the bar is defended against soft yeses AND
against preference-based nos.
Results over the following twelve months:
before after
onsite loops 34 29
offers 6 7
hires 5 6
struggling at 12 months 2 0
regretted attrition 1 0
write-ups with specific
evidence 31% 88%
write-ups before debrief 22% 100%
median time from onsite to
decision 6 days 1.5 days
candidate-declined offers 1 0
Offer rate went up while the bar went up, which surprised the team and has a straightforward explanation: a structured loop with a 24-hour feedback SLA and a 1.5-day decision is a better candidate experience, and the one previously declined offer had cited the slow process.
Onboarding, changed at the same time:
Before: no buddy, first commit at a median of 19 days, no
written success criteria. Two of the five previous hires had
said in retrospect that they "didn't feel useful for two
months".
After: named buddy before day one, a standing list of small
safe changes, a written 30/60/90.
median days to first production change: 19 -> 3
the 90-day "what do we do badly" question produced, from the
first three hires: an undocumented deploy step, a misleading
runbook, and an onboarding doc that had been wrong for a year.
All three were fixed, which is a return on a question that
costs nothing.
Asking a new hire at 90 days what the team does badly is the cheapest audit available, and the answers expire: within six months they will have normalised everything they noticed.
Production evidence
Structured interviewing's predictive advantage over unstructured interviewing is one of the most consistently replicated findings in industrial and organisational psychology (the Schmidt and Hunter meta-analyses being the widely cited synthesis), and it is the empirical basis for same-questions, same-rubric, evidence-in-the-write-up.
Google's re:Work materials document their move to structured interviewing with defined rubrics and independent written feedback before the debrief, along with their published finding that unstructured interviews correlate weakly with performance while adding a strong similarity bias.
Amazon's Bar Raiser program is the clearest institutionalised form of bar defense: a trained interviewer outside the hiring team, empowered to block, existing precisely because a hiring manager under pressure will lower the bar. The asymmetry it encodes, that a wrong yes is more costly and less reversible than a wrong no, is the argument this page makes.
The scorecard-first practice is codified in Geoff Smart and Randy Street's Who as the "scorecard" step, with the same rationale: define the outcomes and competencies before meeting anyone, or you will evaluate against an impression formed in the first minutes.
Anchoring effects in group evaluation are well documented, and the practical countermeasure, collecting independent judgments before discussion, is standard in structured hiring at Google, Amazon and others.
Onboarding research consistently associates early role clarity and early meaningful contribution with time-to-productivity and first-year retention, which is the basis for the shipped-change-in-week-one practice.
The debate
Is a high bar worth a longer search? Yes, and the arithmetic is not close: a wrong hire costs six to twelve months of a lead's attention plus the team absorbing the gap, and a wrong no costs more interviews. The legitimate counter-argument is that an unfilled role also has a cost, and the honest resolution is to fix the pipeline rather than to lower the bar, because lowering it is irreversible in a way that a slower search is not.
Does "culture fit" have any legitimate use? As commonly used, no: in practice it means similarity to the interviewer and it is where bias enters an otherwise structured loop. The legitimate version is values alignment against written, behaviourally defined values ("gives and receives direct feedback", "makes decisions with incomplete information"), assessed with structured questions. If you cannot say which written value the concern maps to, it is not a values concern.
Should the hiring manager be in the loop? Yes, and they should not be the only strong voice. Amazon's Bar Raiser exists precisely because the hiring manager has an interest in filling the requisition, and the structural answer is someone in the room whose incentive is the bar rather than the role.
Is a take-home better than a live coding round? It measures something closer to the job and it excludes people with caregiving responsibilities and second jobs unless it is genuinely time-boxed and respected. The position: offer a choice where you can, cap the take-home at about two hours, and never let a "two-hour" take-home be one where the best candidates spend eight.
Should you hire for potential? For junior roles, largely yes. For senior and staff roles, potential is not the bar; demonstrated scope is, and hiring someone into a level they have not operated at is setting them up in front of an audience. The honest version is to hire them at the level they have demonstrated and say what would move them.
Is fast feedback worth the process cost? Yes, and it is close to free. A 24-hour written feedback SLA improves evidence quality because memory degrades, and a decision within two days of the onsite measurably improves offer acceptance, which the worked example saw directly.
Follow-up Q&A
"Why write the scorecard before the loop?"
Because otherwise the loop measures rapport and calls it judgment. A scorecard names what the role must be able to do, which observable signal proves each capability, and which single interview owns each signal. In one audit the loop assessed coding three times, system design once and inconsistently, and production debugging never, and both hires struggling at twelve months were struggling on production debugging. That is not bad luck; the loop was structurally incapable of detecting it. Writing the "not required" list explicitly matters too, because unwritten preferences operate anyway and a screen that quietly filters on a stack or a company tier removes people before anyone is interviewed.
"What makes a no defensible?"
Specific absent evidence tied to a named scorecard line. Not "I didn't feel it" but "the scorecard says they must reason from evidence in a production scenario; given real logs they proposed three causes in four minutes without checking anything, and when asked what would rule out the first they said we would have to try it." That is checkable and arguable, and it survives a recruiter with a metric and a five-month-old requisition. The inverse discipline matters equally: a no that is really "they did not solve it the way I would have" maps to no scorecard line and should be overridden.
"What is the single cheapest improvement to a hiring loop?"
Independent written feedback submitted before the debrief. It costs nothing, removes nobody's autonomy, and it stops the loudest voice from anchoring the room. In one case it took the share of write-ups containing specific evidence from 31 percent to 88 percent, because a write-up you cannot revise after hearing others' opinions has to stand on its own. Pair it with a 24-hour feedback deadline, since a write-up produced three days later is a reconstruction rather than a record.
"How do you handle the pressure to lower the bar?"
Prepare the responses, because the pressures are predictable. To "we've been looking four months": a wrong hire costs us a year, and what changed about the role that would make this candidate right. To "they're better than nobody": not true, because the team absorbs the gap, usually the strongest people, and the standard resets visibly. To "we can coach them up": credible only for things you named as teachable in advance and have capacity to teach, so "we can teach Kotlin" is fine and "we can teach them to reason from evidence" usually is not. The asymmetry is the whole argument: a wrong yes costs six to twelve months and a wrong no costs more interviews.
"What does good onboarding look like, and why is it a hiring topic?"
A named buddy before day one whose explicit job is to be interruptible for six weeks, a shipped production change in week one, and a written 30/60/90 with success criteria. It is a hiring topic because onboarding failures present as hiring failures, and the fix is different: someone struggling at four months who had no buddy, shipped nothing for three weeks and had no written criteria is not a hiring mistake yet. Keep a standing list of small safe changes so the week-one ship is always possible, because it forces every access and tooling problem to surface immediately.
"What is the most under-used thing in onboarding?"
Asking at 90 days what the team does badly, as an explicit written expectation rather than an offhand question. A new person's outside view has a short shelf life; within six months they will have normalised everything they noticed. From the first three hires in one team it produced an undocumented deploy step, a misleading runbook, and an onboarding document that had been wrong for a year, all of which were fixed. It is the cheapest audit available and it also tells the new person their perspective is wanted.
Common misconceptions
"You know a good candidate when you see one." Unstructured interviews are weak predictors of performance and strongly reflect similarity to the interviewer, which is what "I just liked them" usually measures.
"A high bar means a slow process." They are independent. In one rebuild the bar rose, evidence quality rose, and time from onsite to decision went from six days to a day and a half, which improved offer acceptance.
"Better than nobody." The team absorbs the gap, usually the strongest people, and the visible standard resets. The comparison is not to an empty seat, it is to the next candidate plus the cost of being wrong.
"Culture fit is a real signal." As commonly used it means similarity to the interviewer. The legitimate version is written, behaviourally defined values, and if a concern maps to none of them it is not a values concern.
"The hiring manager should decide." They have an interest in filling the requisition, which is exactly why a bar-raising voice with a different incentive exists in mature processes.
"Hire for potential." For junior roles, largely. At senior and staff, potential is not the bar and hiring someone into a level they have not operated at sets them up to struggle publicly.
Interview delivery note
Say this verbatim: "I write the scorecard before the loop: what the role must be able to do, which observable signal proves it, and which single round owns that signal. Otherwise you assess coding three times, never assess production debugging, and then cannot explain why the hire is struggling on production debugging a year later." It names the practice and the exact failure it prevents, with a consequence that is concrete.
The senior-versus-staff separator is articulating a no in scorecard terms under pressure. A senior engineer says a candidate did not feel right. A staff engineer says which named signal was required, what the candidate actually did and said that showed its absence, and whether it is teachable in the time available, which is what makes a no survive a recruiter with a metric and a five-month-old requisition. And the same person overrides a no that turns out to be "they did not solve it the way I would have," because the discipline runs in both directions.
The second signal is treating onboarding as part of hiring. Naming a buddy before day one, keeping a standing list of small safe changes so a new hire ships to production in week one, and writing a 30/60/90 with success criteria is what makes "this hire is struggling" a diagnosable statement rather than a conclusion about the person.
Further reading
- Google's re:Work guides on structured interviewing, rubrics and independent written feedback.
- Amazon's published description of the Bar Raiser role, for institutionalised bar defense and the wrong-yes asymmetry.
- Geoff Smart and Randy Street, Who, for the scorecard-first method.
- Schmidt and Hunter's meta-analytic work on selection methods, for the predictive validity of structured versus unstructured interviews.
- The promotions and the calibration room page, for the same evidence-versus-adjectives discipline applied internally.