Promotions won two quarters early, and the calibration room

What it is

A promotion is decided in a room you are not in, by people who mostly have not seen the work, using evidence that was produced months earlier. So the promotion is won two quarters before the cycle, by deliberately manufacturing the evidence and the witnesses, not by writing a good packet in the last two weeks.

The calibration room, mechanically:

  - 6 to 12 managers and senior ICs, plus a chair (often the
    director or an HR partner)
  - each candidate is presented by their manager, in 5 to 15
    minutes
  - the room compares candidates AGAINST THE RUBRIC and against
    each other, to keep the bar consistent across teams
  - one confident objection is far more powerful than three
    supportive nods, because "not yet" is the safe answer
  - the people in the room mostly have not seen your work, so
    they are evaluating your MANAGER'S ACCOUNT of it, and the
    corroboration of anyone in the room who has seen it

What this is confused with: promotion as a reward for a good year. Promotion is a statement that someone is already operating at the next level, which is a claim about scope and evidence rather than about effort or output. "They worked incredibly hard and delivered a lot" is the packet that loses, because it argues the current level well.

Also confused: the packet and the case. The packet is a document written in a two-week window. The case is the set of artifacts, witnesses and scope decisions accumulated over the preceding two to four quarters. The packet can only describe the case; it cannot create one.

The problem it solves

The predictable failure is a strong engineer who has done next-level work invisibly.

Q1  builds the migration tooling three teams later depend on.
    Nobody outside the team knows.
Q2  quietly prevents two incidents by catching design flaws in
    review. No artifact exists.
Q3  mentors two juniors to independence. Verbal only.
Q4  manager writes the packet. It says: "highly impactful,
    trusted by peers, drives quality."

In the room:
  chair:  "What is the evidence of scope beyond their team?"
  manager: "Three teams use the migration tooling."
  senior IC from another org: "I've not heard of it."
  chair:  "Let's revisit next cycle."

Nothing in that sequence was unfair. The work happened, and no
durable, attributable, corroborated evidence of it existed.

And the second failure: the person who has done the work but not at the required scope.

Rubric line, Staff: "influences technical direction beyond their
own team."

Candidate's evidence: excellent design work, all within the
team's own services, plus one cross-team consultation.

That is a genuine senior engineer performing very well. It is not
the next level, and a manager who argues otherwise loses
credibility for the next candidate they bring.

The remedy for both is the same and it is temporal: decide what the evidence will be, then assign the work that produces it, at least two quarters out.

Mechanics

Working backwards from the rubric

Take the ACTUAL rubric, line by line, not your summary of it.

  Rubric line              Current evidence      Gap
  ------------------------------------------------------------
  "solves ambiguous        the ingest redesign,  none
   problems"               Q2
  "influences technical    none outside the      THE GAP
   direction beyond        team
   their own team"
  "grows other             mentored one intern   partial: needs
   engineers"              informally            a sustained,
                                                 attributable
                                                 example
  "operational             on-call competent,    partial: no
   ownership"              no incident command   incident command

Two gaps. Both take a quarter of deliberate assignment to close,
and neither can be closed by writing better prose in the packet.

Then convert each gap into a named artifact with a date:

Gap: influence beyond the team

  Assignment: own the schema-evolution RFC that the three
  consuming teams need.
  Artifact:   the RFC document, the review meeting notes with
              the three teams, and the recorded adoption
              decision.
  Witnesses:  the two staff engineers who will be in the
              calibration room need to have read it and,
              ideally, commented on it.
  Date:       RFC circulated by end of Q1, adopted by mid-Q2,
              two quarters before the Q3 cycle.

The witness line is the one leads omit, and it is the difference between "the manager says they did this" and "someone in this room read it."

What counts as evidence

STRONG                              WEAK
------------------------------------------------------------
a written design doc with           "they're a great designer"
  named reviewers
an RFC other teams adopted          "they influence people"
an incident they commanded,         "they're strong
  with the write-up                  operationally"
a person they grew, who says so     "they mentor a lot"
  in writing
a measured outcome with a           "improved performance
  before/after number                significantly"
a system they own that other        "trusted with critical
  teams depend on                     work"

The test: could a stranger in the calibration room verify this in two minutes without asking you? Anything failing that test is an adjective.

And peer feedback is evidence, which means it must be solicited deliberately:

Ask, two quarters early, and be specific about what you need:

  "In the peer feedback round, could you write about the
   schema-evolution RFC specifically? What was useful about it
   for your team, and what would have happened without it. A
   short, concrete paragraph is worth more than a general
   endorsement."

Unsolicited peer feedback is uniformly positive and uniformly
vague, which makes it useless in the room. Solicited, specific
feedback is a citation.

The scope question, which is the real bar

Most promotion denials at senior-to-staff are about SCOPE, not
quality. The three dimensions:

  BLAST RADIUS   team -> several teams -> org -> company
  AMBIGUITY      given a task -> given a problem -> finds the
                 problem
  TIME HORIZON   this sprint -> this quarter -> this year ->
                 multi-year

A candidate can be excellent on quality and unarguably not at
the next level on all three. The lead's job in the two quarters
before is to CREATE the scope, not to argue about it later.

Creating scope is an assignment decision, and it usually means giving something up:

The cross-team migration coordination is the scope. Someone has
it now. Giving it to the candidate means the current owner does
something else, and that is a real cost with a real conversation
attached.

A lead who wants the promotion but will not reassign the scope
is asking the calibration room to promote someone for work they
were not allowed to do.

Preparing the room, not just the packet

THE CHAIR: knows the rubric best and asks the hardest question.
  Anticipate it. If the weakest line is operational ownership,
  the packet should address it before it is asked, with the
  specific evidence and, where the evidence is thin, an honest
  acknowledgment plus what has changed.

THE SKEPTIC: there is usually one, and their objection is
  usually specific. Find it in advance by pre-socialising: show
  the case to one or two people who will be in the room, a
  month before, and ask "what would you push back on?"
  That conversation is worth more than a week of packet
  editing.

THE CORROBORATOR: someone in the room who has seen the work
  first-hand. If nobody in the room has, the case rests entirely
  on your account, and one confident objection outweighs it.
  Manufacturing a corroborator is a two-quarter project: it
  means the candidate's work has to touch that person's
  world.

One confident detractor beats several supporters, because "not yet" is the reversible decision and the room is optimising against promoting someone who then struggles. So the work is not accumulating support, it is removing the specific objection.

The packet itself

Structure that survives scrutiny:

  1. THE CLAIM, in one sentence.
     "X is operating at Staff and has been for two quarters."
  2. RUBRIC LINE BY RUBRIC LINE. For each, the specific artifact
     and who can corroborate it.
  3. THE STRONGEST ITEM FIRST, and it should be a scope item,
     not an output item.
  4. THE HONEST GAP. Name the weakest line yourself, say what
     the evidence is, and what has changed. A packet with no
     weaknesses is either dishonest or describing someone who
     was promoted late.
  5. NO ADJECTIVES that are not immediately followed by an
     artifact.

Naming the weakest line yourself is counter-intuitive and it works, because the room will find it anyway, and finding it themselves reads as the manager either not knowing or concealing it.

The conversation with the candidate

TWO QUARTERS OUT, explicitly:
  "Here is the rubric. Here is where I think you are on each
   line. These two lines are the gap. Here is the work I am
   going to give you to close them, and here is the artifact I
   need to exist by June. This is not a guarantee, and I will
   tell you if my read changes."

NO SURPRISES. If the case is not going forward, say so before
the cycle, not after. The most damaging version is a candidate
who believes they are being put up and discovers in the
announcement that they were not.

AND SAY WHAT IS OUT OF YOUR CONTROL. Calibration compares across
teams and is subject to budget and headcount reality in some
organisations. Pretending it is purely meritocratic sets up a
worse conversation later.

A worked example: two candidates, one cycle

A team of nine. Two senior engineers, both strong, both wanting Staff. The lead started planning in Q1 for the Q3 cycle.

The rubric mapping, done in Q1:

Rubric line                     Candidate A       Candidate B
------------------------------------------------------------------
solves ambiguous problems       STRONG            STRONG
  (evidence)                    (ingest redesign) (search rewrite)
technical direction beyond      NONE              partial: one
  own team                                        cross-team design
                                                  review
grows other engineers           informal only     STRONG (two
                                                  people promoted
                                                  to senior)
operational ownership           STRONG (commanded partial: on-call
                                two incidents)    competent, no
                                                  incident command
written communication           weak: no durable  STRONG (three
                                artifacts         published design
                                                  docs)

The read: both had one large gap and one partial gap, and they were different gaps, which meant different assignments rather than the same advice.

Candidate A's two quarters:

GAP: influence beyond the team, and no durable written artifacts.

Assignment: own the event-schema standard that four teams needed
  and nobody owned. Explicitly a scope stretch in a familiar
  technology.

  Q1  RFC drafted. Circulated to four teams. Two rounds of
      revision from genuinely hostile feedback from one team,
      which A handled by scheduling a call rather than replying
      in the doc. That behaviour was itself later cited.
  Q2  adopted by three of four teams. The fourth declined with a
      documented reason, which A wrote up rather than hiding.
      A ran a 30-minute session at the engineering all-hands.

Artifacts by end of Q2:
  - the RFC, with 31 comments from 9 people across 4 teams
  - the adoption decision record, including the declining team's
    rationale
  - the all-hands recording
Witnesses in the room: the staff engineer from the platform team
  who had pushed back hardest and then supported it, and the
  director who chaired the adoption decision.

The hostile reviewer becoming a supporter was worth more than an easy adoption would have been, because it produced a corroborator with credibility precisely from having disagreed.

Candidate B's two quarters:

GAP: operational ownership, no incident command.

Assignment: incident commander training, then commander on the
  rotation, plus ownership of the quarterly reliability review.

  Q1  shadowed two incidents as scribe, took the IC training.
  Q2  commanded three incidents including one Sev1 lasting
      4 hours. Wrote all three postmortems. Ran the reliability
      review with the director present.

Artifacts:
  - three postmortems, one of which changed a platform-wide
    practice (the canary scope change)
  - the reliability review deck and the resulting funded work

A complication, handled honestly: the Sev1 went badly for the
first 40 minutes. B misdiagnosed and pursued the wrong path.
The postmortem said so, in B's own words, and described what
they changed.

The lead's decision was to include it in the packet rather than
omit it, on the reasoning that the room would otherwise hear
about it from the platform director who was on the call, and
hearing it from the manager first with the learning attached is
strictly better.

Including the failure was the right call and it was argued about internally. In the room, the platform director's comment was that B's postmortem was the most honest one they had read that year, which is not a sentence that would have existed had the packet omitted it.

Pre-socialisation, one month before the cycle:

The lead showed both cases to two people who would be in the
room and asked: "what would you push back on?"

On A: "The RFC is good. Is there evidence they can operate when
       there is no obvious right answer, or was the schema
       standard a case where the answer was clear and the hard
       part was coordination?"
  -> a real gap in the framing. The lead added the ingest
     redesign as the ambiguity evidence and reframed the RFC as
     the scope evidence, rather than trying to make one artifact
     carry both.

On B: "Three incidents in one quarter, and one of them went
       badly. Is that enough of a track record?"
  -> anticipated. The packet added the two shadowed incidents,
     the training, and a statement from the on-call lead about
     B's subsequent shifts.

Both objections were raised in the actual room, in almost the
same words, and both were already addressed in the packet.

Finding the objection a month early and answering it in the packet is the single highest-return activity in the whole process, and it costs two 20-minute conversations.

Outcome:

A: promoted. The discussion took 6 minutes. The platform staff
   engineer's corroboration was the decisive contribution: "I
   disagreed with the first draft and they came to me rather
   than arguing in comments. Three of my team's services depend
   on the outcome."

B: promoted. The discussion took 19 minutes, almost all of it
   about the Sev1. The chair's summary: "the postmortem is
   next-level work regardless of how the incident went."

The lead's read afterwards, recorded for the next cycle: A's
case was won in Q1 when the schema standard was reassigned away
from the person who had it. Everything after that was
documentation.

"The case was won when the scope was reassigned" is the compressed version of this whole page.

And one candidate who was not put up, handled explicitly:

A third senior engineer had asked about Staff in Q1. The lead's
honest read was two gaps and no realistic path within two
quarters, because the scope needed did not exist on the team.

That was said in Q1, in those words, along with what would need
to be true and roughly when. It was an uncomfortable
conversation and it was better than the alternative, which was
an ambiguous "let's see how the year goes" followed by an
announcement in which they were not included.

They were promoted the following year.

Production evidence

Calibration as a formal process is documented practice at Google, Meta, Amazon, Microsoft and most large technology companies, with the stated purpose of maintaining a consistent bar across managers and teams. The structure, a manager presenting a case to a panel that mostly has not observed the work, is what makes corroboration and artifacts decisive.

Published engineering career ladders (Rent the Runway's, CircleCI's, Dropbox's, Square's, and the collection at progression.fyi) all express levels in terms of scope and observable evidence rather than tenure or output volume, which is the basis for working backwards from rubric lines.

Amazon's Bar Raiser program, though a hiring mechanism, is the clearest documented instance of the asymmetry that also governs promotion panels: a designated participant empowered to block, on the principle that a wrong "yes" is more costly and less reversible than a wrong "no."

Research on structured evaluation and bias underpins why written artifacts and specific solicited feedback outperform general impressions: unstructured evaluation correlates more strongly with similarity to the evaluator, which is one reason "could a stranger verify this in two minutes" is a useful test for evidence.

Google's guidance on promotion packets and comparable published internal materials from other companies consistently emphasise concrete impact statements with measurable outcomes over descriptive praise, which is the same distinction as artifacts against adjectives.

The debate

Is it cynical to plan a promotion two quarters out? The alternative is a system where the outcome depends on whether the year's work happened to produce visible evidence, which favours people already positioned to do visible work. Deliberate planning makes the criteria explicit and available to everyone, and the lead who does it for one person and not another is the one behaving unfairly.

Should you include a failure in a packet? When the room will hear about it anyway, yes, with the learning attached. The counter-argument, that it hands the skeptic ammunition, is real, and the resolution is that an unmentioned failure surfacing from another participant is far worse: it reads as the manager not knowing or concealing. When the failure is genuinely unknown outside the team, the call is closer and the deciding question is whether the learning from it is itself next-level evidence.

Is scope creation fair to the person who currently has the scope? It is a real cost and it must be an explicit conversation, not a quiet reassignment. A lead who will not have that conversation is asking the room to promote someone for work they were not allowed to do, which is the version that actually harms the candidate.

Does pre-socialising the case game the process? It surfaces objections early so they can be addressed with evidence rather than with argument in the room. That improves the decision's information quality, and a case that cannot survive a friendly skeptic a month early was not going to survive a hostile one on the day.

Should you tell someone they are not going up? Always, early, with the specific gaps and what would close them. The counter-argument, that it demotivates, has it backwards: the demotivating version is an ambiguous "let's see," followed by an announcement that does not include them, which also destroys trust in every subsequent conversation.

Is the calibration room meritocratic? Partly, and pretending it is entirely so sets up a worse conversation later. Budget, headcount, cross-team comparison and the relative persuasiveness of different managers all affect outcomes, and saying that plainly to a candidate, alongside the parts that are within their control, is more respectful than the alternative.

Follow-up Q&A

"When is a promotion actually decided?"

Two quarters before the cycle, when the scope is assigned. The calibration room evaluates artifacts and corroboration, and both take a quarter to produce and a quarter to be noticed. The packet can only describe the case; it cannot create one. In one instance the compressed postmortem of a successful case was that it was won in Q1 when the cross-team schema standard was reassigned to the candidate from the person who held it, and everything after that was documentation.

"How do you decide what work to assign?"

Work backwards from the actual rubric, line by line, not from a summary of it. For each line, write what evidence exists today. The lines with no evidence are the gaps, and each gap converts into a named artifact with a date and a set of witnesses: the RFC document, the review notes from the three consuming teams, the recorded adoption decision, circulated by end of Q1 so that the two staff engineers who will be in the calibration room have read it. The witness line is the one leads omit, and it is what turns "the manager says they did this" into "someone in this room saw it."

"What counts as evidence in a calibration room?"

Anything a stranger could verify in two minutes without asking you. A design doc with named reviewers, an RFC other teams adopted, an incident someone commanded with the write-up, a person they grew who says so in writing, a measured outcome with a before and after number. Everything else is an adjective: "great designer", "influences people", "strong operationally", "improved performance significantly." A packet made of adjectives argues the current level well.

"Why is one objection more powerful than several supporters?"

Because "not yet" is the reversible decision and the room is optimising against promoting someone who then struggles. So the work is not accumulating support, it is removing the specific objection. That is why pre-socialising matters: show the case to one or two people who will be in the room a month early and ask what they would push back on. In one cycle both objections found that way were raised in the actual room in nearly the same words, and both were already answered in the packet. Two twenty-minute conversations, and it is the highest-return activity in the process.

"Would you include a failure in a promotion packet?"

If the room will hear about it anyway, yes, with the learning attached, because hearing it from the manager first is strictly better than hearing it from a participant who was on the call. In one case a candidate misdiagnosed the first forty minutes of a Sev1, said so in their own postmortem, and the platform director in the room called it the most honest postmortem they had read that year. That sentence would not have existed had the packet omitted the failure. When the failure is genuinely unknown outside the team the call is closer, and the deciding question is whether the learning is itself next-level evidence.

"What do you tell someone who is not going up this cycle?"

Tell them in the quarter before, not after the announcement, with the specific rubric lines, what evidence is missing, what would close it, and roughly when that is realistic. If the required scope does not exist on the team, say that too, because it is the truth and it changes what they should do. The uncomfortable early conversation is better than an ambiguous "let's see how the year goes" followed by an announcement that does not include them, which damages trust in every subsequent conversation you have with them.

Common misconceptions

"Promotion rewards a great year." It is a statement that someone is already operating at the next level, which is a claim about scope and evidence, not about effort or output.

"Write a strong packet." The packet describes a case that already exists. Two weeks of writing cannot create two quarters of artifacts and witnesses.

"Impact is enough." At senior-to-staff, most denials are about scope: blast radius, ambiguity and time horizon. Excellent work entirely inside one team is a strong senior engineer.

"Never mention a failure." An unmentioned failure surfacing from someone else in the room reads as the manager not knowing or concealing it, which costs more than the failure did.

"Support in the room is what matters." One confident objection outweighs several supporters, because deferral is the safe decision. Remove the objection rather than accumulating nods.

"The process is purely meritocratic." Budget, cross-team comparison and manager persuasiveness all affect outcomes, and saying so alongside what is controllable is more respectful than pretending otherwise.

Interview delivery note

Say this verbatim: "A promotion is decided two quarters before the cycle, when the scope is assigned. The room evaluates artifacts and corroboration, and both take a quarter to produce and a quarter to be noticed, so the packet can only describe a case it cannot create." It states the mechanism and rules out the thing most people optimise.

The senior-versus-staff separator is manufacturing a corroborator deliberately. A senior lead writes a good packet. A staff lead identifies that nobody in the calibration room has seen the candidate's work first-hand, treats that as a two-quarter problem, and assigns work that touches a specific person's world so that a credible participant can say "three of my team's services depend on this." In one case the decisive contribution came from the engineer who had pushed back hardest on the first draft, which is worth more than easy support precisely because they had disagreed.

The second signal is pre-socialising to find the objection. Showing the case to a future participant a month early and asking what they would push back on costs twenty minutes, surfaces the specific objection that would otherwise land unanswered, and converts the room from a debate into a confirmation. It is the same instinct as taking a recommendation upward rather than a question.

Further reading

  • Published engineering career ladders (Rent the Runway, CircleCI, Dropbox, Square) and the progression.fyi collection, for rubric lines expressed as scope and evidence.
  • Amazon's published description of the Bar Raiser role, for the asymmetry between a wrong yes and a wrong no in a panel decision.
  • Google's re:Work materials on structured evaluation, for why specific written evidence outperforms general impressions.
  • The growing people page, for the assignment mechanics that produce the evidence.
  • The underperformance sequence page, for the same no-surprises principle applied in the other direction.