Building a story portfolio, not answers to questions

What it is

A set of 12 to 16 rehearsed stories from your own work, each with numbers and a stated decision, that you map onto whatever question you are asked, rather than a list of answers to anticipated questions.

THE WRONG MODEL
  "Tell me about a conflict"       -> the conflict answer
  "Tell me about a failure"        -> the failure answer
  "Tell me about influence"        -> the influence answer

  Requires anticipating the questions, which is impossible,
  and produces obviously-rehearsed answers to the ones you
  guessed and nothing for the ones you did not.

THE PORTFOLIO MODEL
  16 stories, each tagged with 3 to 5 themes it can serve.

  "Tell me about a conflict"    -> the auth migration story,
                                   told through the
                                   disagreement with the
                                   platform lead
  "Tell me about influence"     -> the SAME story, told
                                   through how the three EMs
                                   were brought along
  "Tell me about a failure"     -> the SAME story, told
                                   through the six weeks lost
                                   to the wrong sequencing

One story serves several questions because the same events contain conflict, influence, failure and technical judgement. The skill is knowing which thread to pull.

Commonly confused with having good answers. The portfolio is inventory; the skill is retrieval and framing. A candidate with twenty excellent stories and no index will search for one under pressure and tell it badly.

The problem it solves

Behavioural interviews sample from a large question space and you cannot cover it by anticipation.

Questions a staff loop might ask:
  disagreement with a peer / with your manager / with a
  senior leader / that you lost
  a failure / a mistake you caused / one you caught
  influence without authority / across teams / upward
  a technical decision you regret / one you were right
  about / one you changed your mind on
  mentoring / underperformance / hiring / firing
  ambiguity / conflicting priorities / a bad deadline
  the hardest thing you have built / the thing you are
  proudest of / the thing you would redo

That is thirty-plus questions. Nobody has thirty stories.

With sixteen well-chosen stories, each serving three to five themes, the space is covered, and the alternative is being caught without material for something you did not predict, which is where candidates freeze.

The second problem it solves is repetition. A candidate who tells the same story to four interviewers looks like they have one experience. The portfolio makes it deliberate: you know which stories you have used and can choose differently.

Mechanics

Choosing the sixteen

Not your sixteen favourite projects. Sixteen that cover the themes and demonstrate the level.

COVERAGE REQUIREMENTS (aim for at least one strong story in
each; several stories will cover several)

  TECHNICAL DEPTH        a decision only you could have made
  TECHNICAL BREADTH      a system you designed end to end
  SCALE                  something large, with numbers
  AMBIGUITY              a problem with no clear right answer
  FAILURE                one that was your fault
  RECOVERY               an incident you led
  CONFLICT (peer)        a disagreement you resolved
  CONFLICT (upward)      a disagreement with a manager or
                         senior leader
  INFLUENCE              cross-team, without authority
  MENTORING              someone who grew, with the mechanism
  UNDERPERFORMANCE       a hard people conversation
  HIRING                 a decision you made and why
  PRIORITISATION         something you chose not to do
  CHANGED YOUR MIND      with the evidence that changed it
  LONG-HORIZON           something that took quarters
  ORGANISATIONAL         a change to how the team worked

The three that candidates are most often missing:

Underperformance. Most senior engineers have never had the conversation, and staff loops ask about it because staff engineers influence people they do not manage. If you genuinely have no story, the closest legitimate substitute is a peer whose work you had to address, and being honest that you have not managed someone is better than inventing it.

Changed your mind. It is asked constantly and it needs a specific piece of evidence that changed your view. "I became more open to X over time" is not a story; "I ran the benchmark and it was 3x slower than I claimed, so I withdrew the proposal in the design review" is.

Something you chose not to do. Prioritisation stories are almost always told as "we shipped a lot", and the interesting version is what you killed and how you defended the decision.

The format: prepare in SCOR, deliver in STAR

PREPARE IN SCOR         because it forces the two things that
                        make a story land:
  Situation             context, compressed
  Complication          what made it hard
  Options               what you considered and rejected,
                        with costs
  Resolution            what you chose and what happened

DELIVER IN STAR         because that is the scoring rubric
  Situation + Task      including one explicit sentence about
                        what was YOURS to decide
  Action                the options compressed to one sentence
                        each, then what you did, in first
                        person singular
  Result                the number, and the lesson

The Options section is the staff signal and it is the one that disappears in a STAR telling. Compress each alternative to one sentence with its cost and keep all of them: "I evaluated scaling the cluster, which would have masked a cause I didn't understand; caching, which has near-zero hit rate for personalised results; and another day of root-causing with the regression visible." Twelve seconds, and it is the highest-density evidence of judgement in the whole story.

See SCOR, STAR and the scar-tissue story for the conversion mechanics.

The index: what makes it a portfolio rather than a pile

STORY                     THEMES                        NUMBERS
------------------------------------------------------------------
Auth migration across     influence, conflict-peer,     4 teams,
4 teams                   long-horizon, ambiguity       2 quarters,
                                                        3 EMs

Search p99 regression     technical depth, recovery,    4k QPS,
after personalisation     changed-my-mind               180->1400
                                                        ->210 ms

Killed the graph          prioritisation, conflict-     6 months
database project          upward, changed-my-mind       saved, 1
                                                        VP annoyed

Priya to senior           mentoring, organisational     18 months,
                                                        2 promoted

The reindex outage        failure, recovery,            40 min,
                          organisational                97% cache
                                                        miss
...

Building this table is the actual preparation, and it takes an evening. Once it exists, the interview becomes a lookup: hear the question, identify the theme, pick a story you have not used, and tell the version that emphasises that theme.

Track which stories you have used with which interviewer, because a loop is four to six conversations and interviewers compare notes.

Rehearsing the numbers, not the sentences

For each story, memorise 4 to 5 figures:
  the SCALE          "4,000 queries per second across nine
                     locales"
  the BEFORE         "p99 was 180 ms"
  the AFTER          "210 ms, better than before the
                     regression"
  one DETAIL only a  "six shards for 80,000 documents"
  participant knows
  the DURATION       "two days to find it"

The prose should vary between tellings and SHOULD. The
numbers must not, because inconsistency across a loop is
noticed and it is the fastest way to lose credibility.

The detail only a participant would know is the highest-value one. "Six shards for 80,000 documents" is not a fact anyone would recall unless they were there, and it does more for credibility than any adjective.

Timing, and the three lengths

Every story needs three versions and rehearsing only one is a common mistake.

30 SECONDS   the scar-tissue version, dropped inside a
             technical answer. Three sentences: what we did,
             what went wrong with a number, what we do now.

90 SECONDS   the standard behavioural answer. Full STAR,
             one number per section, stop.

3 MINUTES    the deep-dive version, when an interviewer says
             "tell me more about that". Adds the options in
             detail, the technical mechanics, and the
             follow-on consequences.

The failure mode is telling the three-minute version when ninety seconds was asked for, and it is extremely common. An interviewer who wants more will ask; one who is waiting for you to stop will not interrupt, and the impression is that you cannot calibrate.

What makes a story fail

NO NUMBER              nothing is verifiable, and it sounds
                       like a description of a category of
                       event rather than a memory.

NO "I"                 STAR scores individual contribution.
                       Say what the team did, then what you
                       did. This is rubric compliance, not
                       credit-taking.

NO DECISION            a narration of events you were present
                       for. The Options section is what makes
                       it a decision.

NO RETROSPECT          "what I'd do differently" is the
                       cheapest credibility available and
                       most candidates omit it.

TOO OLD                a story from six years ago suggests
                       nothing has happened since. Prefer the
                       last two to three years.

TOO SMALL              a story whose scope is one sprint does
                       not demonstrate staff level, however
                       well told.

A worked example: one story, four questions

THE EVENTS
  A search personalisation launch caused a p99 regression
  from 180 ms to 1.4 s, but only in small locales. Took two
  days to find because the team looked at aggregate latency.
  Cause: one query per shard per user segment, against
  locales over-sharded at six shards for 80,000 documents.
  Fixed by resharding and batching. p99 went to 210 ms.
  Added per-locale alerting, which caught an unrelated
  regression six weeks later.

"TELL ME ABOUT A TECHNICAL PROBLEM YOU SOLVED"
  Lead with the inversion: the smallest indices were the
  slowest, which ruled out capacity. Emphasise the diagnosis
  and the options considered.

"TELL ME ABOUT A TIME YOU CHANGED YOUR MIND"
  Lead with: I was convinced it was a capacity problem and
  argued for scaling the cluster. Segmenting by locale
  proved me wrong within an hour, and I withdrew the
  proposal in front of the team that had been about to
  approve the spend.

"TELL ME ABOUT PRESSURE FROM LEADERSHIP"
  Lead with: the business asked daily whether to roll
  personalisation back, and I asked for one more day with
  the regression visible. Emphasise the negotiation, the
  fallback plan I offered, and that I'd have rolled back if
  the day produced nothing.

"TELL ME ABOUT A FAILURE"
  Lead with: it took us two days to find, and it shouldn't
  have. We were alerting on aggregate p99, which hid an
  inversion that would have been obvious per locale on day
  one. That's the failure, and the alerting change is what
  came out of it.

Same events, four openings, four emphases, and the numbers are identical in all four. That is what a portfolio buys, and it is why sixteen stories cover thirty questions.

Production evidence

Amazon's Leadership Principles loop trains interviewers to collect STAR-structured evidence and to probe specifically for individual contribution, which is why the "we" to "I" conversion is mechanical rather than stylistic, and Amazon's own candidate guidance names STAR explicitly.

Google's published interview guidance tells candidates to use STAR and states that the interviewer is assessing what you did, which is the same rubric field.

Structured behavioural interviewing research consistently finds that structured questions predict job performance better than unstructured ones, which is why interviewers hold to the format: the structure is what makes candidates comparable, and a story that will not fit it is genuinely harder to score.

Barbara Minto's The Pyramid Principle is the source of the situation-complication-resolution structure, developed for consulting communication, and the reason leading with the complication works is that it makes the audience want the answer.

Will Larson's Staff Engineer contains interview accounts from staff engineers at many companies, and the recurring theme is that the stories that landed were about influence and judgement rather than about technical difficulty.

The debate

The case for a prepared portfolio: the question space is too large to anticipate, and under pressure people become vague. Preparation is what prevents that, and it is the single highest-return activity in behavioural preparation.

The case against over-preparing: rehearsed answers sound rehearsed, interviewers notice, and it reads as inauthentic. Some interviewers deliberately ask unusual questions to get past prepared material.

The case for preparing themes rather than stories: know what you want to convey and let the specifics come naturally. More flexible, and it fails under pressure precisely because specifics are what disappears when you are nervous.

My position: prepare sixteen stories in SCOR, index them by theme, rehearse the numbers rather than the sentences, and vary the prose deliberately between tellings.

The distinction between rehearsing numbers and rehearsing sentences is what resolves the over-preparation objection. A story told with identical wording sounds recited; a story told with identical figures and different wording sounds like a memory. So I would fix the four or five numbers per story and let everything else vary, which also means the story adapts to the question rather than being delivered regardless of it.

The three coverage gaps I would specifically hunt for are underperformance, changed my mind, and something you chose not to do, because they are asked constantly and most candidates have no material. For "changed my mind" the requirement is a specific piece of evidence, and "I became more open to X" is not a story. For underperformance, if you genuinely have not managed anyone, saying so and offering the closest real thing is better than constructing something, because an interviewer probing a fabricated people story finds the bottom of it in two questions.

The mechanical thing I would not skip is the index table, because it converts the interview from recall to lookup. Hearing a question and searching your memory for a relevant story is where candidates freeze; hearing a question, identifying the theme, and picking from a pre-tagged set is not. That table is an evening of work and it is worth more than rehearsing any individual answer.

And three lengths per story, because telling the three-minute version when ninety seconds was asked for is extremely common and reads as an inability to calibrate. An interviewer who wants more will ask.

Where I would push back on the framing: the portfolio is inventory, and the skill is retrieval. Twenty excellent stories with no index is worse than twelve indexed ones, because the failure under pressure is not lacking material, it is not finding it.

Follow-up Q&A

"Why a portfolio rather than answers to expected questions?" Because the question space is thirty-plus questions and nobody has thirty stories. Sixteen stories, each tagged with three to five themes it can serve, covers it, because the same events contain conflict, influence, failure and technical judgement, and the skill is knowing which thread to pull. Preparing answers requires anticipating the questions, which produces obviously-rehearsed answers to the ones you guessed and nothing for the ones you did not.

"How do you choose the sixteen?" By coverage, not by favourite. At least one strong story for each of: technical depth, breadth, scale, ambiguity, failure that was your fault, incident recovery, peer conflict, upward conflict, influence without authority, mentoring, underperformance, hiring, prioritisation, changing your mind, something long-horizon, and an organisational change. Several stories will cover several themes, which is the point.

"Which coverage gaps do candidates usually have?" Three. Underperformance, because most senior engineers have never had that conversation and staff loops ask because staff engineers influence people they do not manage. Changed your mind, which needs a specific piece of evidence, since "I became more open to X" is not a story. And something you chose not to do, because prioritisation stories are almost always told as "we shipped a lot" and the interesting version is what you killed.

"Doesn't this sound rehearsed?" Only if you rehearse the sentences. I fix four or five numbers per story, the scale, the before, the after, one detail only a participant would know, and the duration, and let the prose vary between tellings. Identical wording sounds recited; identical figures with different wording sounds like a memory. And it means the story adapts to the question rather than being delivered regardless of it.

"Show me how one story serves several questions." Take a search latency regression: p99 went from 180 milliseconds to 1.4 seconds, only in small locales, took two days to find. For "a technical problem you solved", lead with the inversion, that the smallest indices were slowest, which ruled out capacity. For "a time you changed your mind", lead with having argued for scaling the cluster and being proven wrong within an hour. For "pressure from leadership", lead with the business asking daily whether to roll back and me asking for one more day. For "a failure", lead with the two days, which was our alerting hiding an inversion that would have been obvious per locale. Same events, four openings, identical numbers.

"How long should a story be?" Three lengths, all rehearsed. Thirty seconds for the scar-tissue version dropped inside a technical answer. Ninety seconds for the standard behavioural answer. Three minutes for when someone says "tell me more". The common failure is giving the three-minute version when ninety seconds was asked for, which reads as an inability to calibrate, and an interviewer who wants more will ask.

"What makes a story fail?" Four things in order. No number, so nothing is verifiable and it sounds like a description of a category of event rather than a memory. No "I", because the rubric has a field for individual contribution and saying what the team did without saying what you did leaves it empty. No decision, so it is a narration of events you were present for, which is why the Options section matters. And no retrospect, which is the cheapest credibility available and most people omit it.

"What's the actual preparation work?" The index table: story, the three to five themes it serves, and its numbers. An evening. Once it exists the interview is a lookup rather than a recall problem, and recall is what fails under pressure. I would also track which stories I have used with which interviewer, because a loop is four to six conversations and interviewers compare notes, so telling the same story twice makes it look like you have one experience.

"What if you genuinely lack a story for a theme?" Say so and offer the closest real thing. For underperformance, if you have never managed anyone, "I haven't managed, but here's a peer whose work I had to address and how I handled it" is a good answer. Constructing something is much worse, because an interviewer probing a fabricated people story finds the bottom of it in two questions, and at that point the rest of the interview is spent recovering credibility.

Common misconceptions

"Prepare an answer for each likely question." The space is too large. Prepare stories and index them by theme.

"Rehearsing makes you sound fake." Rehearsing sentences does. Rehearsing numbers and varying the prose does the opposite.

"The best stories are the most technically impressive." At staff level the stories that land are about influence, judgement and decisions, and technical difficulty is the setting rather than the point.

"One great story is enough." A loop is four to six conversations and interviewers compare notes. Repeating a story makes it look like you have one experience.

"More detail is better." Beyond about a hundred seconds the interviewer stops tracking and starts waiting. Compress the Situation ruthlessly; it is the part everyone over-tells.

Interview delivery note

This topic is usually assessed indirectly, through whether your stories are good, but it does come up as "how do you prepare" and as advice-giving in a leadership conversation.

Frame it as inventory plus retrieval: "I'd build a portfolio rather than answers, because the question space is thirty-plus questions and nobody has thirty stories. Sixteen stories, each tagged with three to five themes, covers it, because the same events contain conflict, influence, failure and technical judgement. The skill is knowing which thread to pull."

Name the mechanism that makes it work under pressure: "And the actual artifact is an index table: story, themes it serves, and its numbers. That converts the interview from a recall problem to a lookup, and recall is what fails when you're nervous."

Give the over-preparation answer, because it is the obvious objection: "I'd rehearse the numbers, not the sentences. Four or five figures per story, fixed, and the prose varies every telling. Identical wording sounds recited; identical numbers with different wording sounds like a memory. And inconsistent numbers across a loop are noticed immediately."

The coverage point worth volunteering: "and I'd hunt specifically for the three gaps most people have: underperformance, changed-my-mind with a specific piece of evidence, and something you chose not to do. Those are asked constantly and most candidates have nothing."

Further reading

  • Barbara Minto, The Pyramid Principle, for the situation-complication-resolution structure.
  • Amazon's and Google's published interview preparation guidance, for how the STAR rubric is actually applied.
  • Will Larson, Staff Engineer, for the interview accounts and what distinguished the stories that landed.
  • SCOR, STAR and the scar-tissue story, for the format conversion mechanics.