Drills 37 to 42: leadership
Six role-plays, ninety seconds each, out loud. These are harder to rehearse than technical drills because the failure mode is not being wrong, it is being vague.
Every one of these answers uses the same skeleton, and you should say it out loud before answering: first move, information I would gather, line I would not cross. Announcing the structure buys you the benefit of the doubt for the next ninety seconds, and it stops you rambling.
The second discipline: name your own contribution. In four of these six, part of the cause is a management failure. Candidates who run these purely as conversations about the other person are scored as having missed it.
Drill 37. Your best engineer's PR comments are demoralising juniors. First move?
My first move isn't the conversation, it's reading the actual comments. I'd pull five to ten from the last two weeks, because "people feel bad" isn't actionable and "here are four comments and what each one costs the author" is. I'd sort them into correct-and-well-delivered, correct-and-badly-delivered, and not-actually-correct, and that third pile is the most damaging, because the author can't tell it from the real findings so they have to treat everything as blocking.
Then a private conversation using SBI: the specific PR, the specific wording, and what it cost. And a concrete ask, not "be kinder", which nobody can act on. Say what's blocking and what isn't. State the finding, not a judgement of the author.
Then I'd change the system, framed for the whole team so nobody's named: a comment taxonomy with blocking, suggestion, nit, question and praise prefixes, because most of the harm is ambiguity rather than tone; automation of everything mechanical so humans never comment on style; and reviewer rotation so nobody is a single gate.
The line I wouldn't cross is lowering the bar. Their standard is why they're valuable. It's the delivery I'm changing.
Depth signal: the system change alongside the conversation, and naming how you would know it worked (review latency, queue depth, whether the juniors' PR rate recovers).
Full treatment: The toxic code reviewer.
Drill 38. Review queue depth doubled after the AI tooling rollout. What do you do?
First, I'd say that this is expected rather than surprising, because when generation speeds up the bottleneck moves from writing to reviewing. Teams hit this in month two of adoption almost universally.
The measurement first: review queue depth, time to first review, and merge time, split by whether the PR was AI-assisted. Without the split I'm guessing.
Then four counters. Require authors to be able to explain generated code as their own, which is both a quality gate and a learning one. Label AI-assisted PRs so reviewers calibrate their attention. Raise test requirements on generated code, because tests are the check that scales when volume rises and human review doesn't. And cap PR size, because review effectiveness collapses past roughly 400 lines and generated PRs are often large.
The thing I'd raise unprompted with leadership is the two-sided data: AI adoption correlates with higher throughput and also with higher change failure rate. So I'd pair every speed metric with a quality guardrail rather than reporting deployment frequency alone, because otherwise we'll celebrate a number that's getting worse underneath.
Depth signal: naming the throughput-versus-stability tradeoff as the thing to instrument, rather than treating the queue as a staffing problem.
Drill 39. Your director wants a date you cannot commit to.
My first move isn't to answer, it's to ask what the date is anchored to. A contractual deadline, a customer commitment already made, a conference, and a stretch target need completely different responses, and quite often the real constraint has a better answer than either of us started with.
Then options with costs, never a yes or no. Full scope at 60 percent confidence on the 29th and 90 percent on the 12th. Or the 15th with the bulk import and admin UI cut, which is the core flow working end to end. Or the 15th at full scope with two borrowed engineers, at about 70 percent, and I'd be honest that it slows the team I'm borrowing from. And I'd recommend one, because handing over three options with no opinion is abdication rather than collaboration.
The confidence numbers come from the actual cycle time of the last eighteen comparable items, not from story points, because points measure imagined effort and cycle time measures what happened.
If they push anyway, I'd commit and make the risk explicit: "we hit this about one time in four, so let's agree now what we drop and what we tell the customer if it slips." Then execute properly, because a recorded objection followed by half-hearted delivery is the worst of both.
The line I wouldn't cross is giving a date I don't believe, because they'd plan on it and the cost lands on people downstream.
Depth signal: percentile forecasting from historical cycle time, and asking what the date is for before answering.
Full treatment: The impossible date.
Drill 40. Make the case for 25 percent reliability investment to a product VP.
I wouldn't open by asking for capacity. I'd open by showing we already spend it: 18 percent of the team's time went to unplanned work last quarter, up from 11 two quarters before, and 60 percent of it traces to deploy failures we catch by hand after users notice.
Then three specific items rather than a budget line: automated canary analysis with rollback, load-test gates in CI, and the two dependency timeouts that caused four incidents. The expected return stated up front: unplanned work under 8 percent, so a net gain of about 10 percent of capacity. Bounded to one quarter with a review.
And the sentence that makes it credible: if the number hasn't moved by the review, we should stop rather than keep spending. That converts a request into an experiment.
If they cut me to 10 percent, I take it, scope it honestly, and put the consequence on the record: "with 10 I'd do the canary work, which should take us from 18 to about 12; the capacity incidents continue and I'd want to revisit after the launch." Ten percent with a measured result beats 25 percent with an argument.
Separately I'd push for an error budget policy, so this stops being a quarterly negotiation. The catch is that leadership has to sign it before the budget runs out, not during the incident.
Depth signal: the reframe (already spending it, invisibly, at a worse exchange rate), and accepting the smaller number gracefully while pricing what is lost.
Full treatment: The reliability investment case.
Drill 41. Two teams are building the same service. You have no authority over either.
First I'd verify the duplication is real, because "the same thing" is often two teams solving different problems that look alike from outside. Two hours reading both. If they're genuinely different, saying so publicly is the most valuable thing I can do, and forcing a merge would destroy value.
Assuming it's real, I don't argue that duplication is bad, because nobody disputes that and nobody acts on it. I quantify: three engineer-years a year of duplicated maintenance, seven downstream consumers with four integrating against both, and a live correctness divergence on partial refunds that nobody owns. Then I translate that into whatever the decision-maker already said they wanted capacity for, so consolidating becomes the route to their goal rather than a tidiness project.
Then I find the lowest common manager and give them a written decision to make: the situation, the cost, four real options including do-nothing, a recommendation, and one specific ask, which is a decision by a date plus an announcement that it's decided. DACI if they want the vocabulary; the value is a single named approver.
Two things I'd do that people skip. Talk to both tech leads before writing anything, so neither is ambushed and the document is accurate. And make sure the team whose service is retired owns the migration and gets its distinctive features ported, because "your year of work is deleted" is why these get agreed and then quietly not done.
The line I wouldn't cross is trying to decide it myself. I have no authority, so my job is to make the decision easy for someone who has it, and then support it whichever way it goes.
Depth signal: checking the premise, and giving the losing team something real.
Full treatment: Two teams building the same service.
Drill 42. An engineer wants promotion; they are one level of scope short.
Before the conversation I'd go through the next level's rubric line by line and ask the question that decides whose problem this is: have they had the opportunity to demonstrate what's missing? If the gap is cross-team influence and every project I've given them was inside the team, the gap is mine, because at this level the evidence comes from the work someone is assigned.
Then I'd lead with the answer so they're not spending the conversation guessing. Specific against the rubric, including what's met: depth is there, quality is there, mentoring is there, scope isn't. Name my own contribution honestly. And convert the gap into named work rather than an instruction to be more strategic: lead the auth migration across four teams, drive the API standards RFC to adoption, present at architecture review.
Then a review date, and a very careful promise. I promise the packet and my advocacy; I never promise the outcome, because I don't control the calibration room. And I commit to telling them in January if it isn't tracking, so they don't find out in March.
The line I wouldn't cross is being vague to be kind. "You're really close, keep going" is heard as a yes, it costs them a year of the wrong work, and the next conversation is far worse.
Depth signal: "have they had the opportunity", and promising the packet rather than the promotion.
Full treatment: Promotion when they are one level short.
How to practise these
Not by reading. These need to be said to a person, because the failure mode is specific to speech: under mild pressure people become vague, and vagueness is exactly what is being scored.
The drill: read only the question, set ninety seconds, and answer out loud to a colleague who is instructed to interrupt once with "can you be more specific?" That single interruption is the whole exercise, because the answer to it is where the grade is.
Three tests to apply to your own answer afterwards:
- Did you name a specific first action, or did you describe a category of action? "I'd have a conversation" fails; "I'd pull five of their recent comments" passes.
- Did you name your own contribution? In four of these six there is one, and volunteering it is the strongest single signal available.
- Did you say how you would know it worked? Almost nobody does. Adding one sentence about the metric you would watch converts a plausible answer into an operator's answer.
And the calibration rule that governs all of them: every story and every role-play answer should contain one number. Review latency, interrupt rate, unplanned work percentage, confidence level. Leadership answers without numbers sound like opinions, because that is what they are.