Building a story portfolio, not answers to questions
What it is
A set of 12 to 16 rehearsed stories from your own work, each with numbers and a stated decision, that you map onto whatever question you are asked, rather than a list of answers to anticipated questions.
THE WRONG MODEL
"Tell me about a conflict" -> the conflict answer
"Tell me about a failure" -> the failure answer
"Tell me about influence" -> the influence answer
Requires anticipating the questions, which is impossible,
and produces obviously-rehearsed answers to the ones you
guessed and nothing for the ones you did not.
THE PORTFOLIO MODEL
16 stories, each tagged with 3 to 5 themes it can serve.
"Tell me about a conflict" -> the auth migration story,
told through the
disagreement with the
platform lead
"Tell me about influence" -> the SAME story, told
through how the three EMs
were brought along
"Tell me about a failure" -> the SAME story, told
through the six weeks lost
to the wrong sequencing
One story serves several questions because the same events contain conflict, influence, failure and technical judgement. The skill is knowing which thread to pull.
Commonly confused with having good answers. The portfolio is inventory; the skill is retrieval and framing. A candidate with twenty excellent stories and no index will search for one under pressure and tell it badly.
The problem it solves
Behavioural interviews sample from a large question space and you cannot cover it by anticipation.
Questions a staff loop might ask:
disagreement with a peer / with your manager / with a
senior leader / that you lost
a failure / a mistake you caused / one you caught
influence without authority / across teams / upward
a technical decision you regret / one you were right
about / one you changed your mind on
mentoring / underperformance / hiring / firing
ambiguity / conflicting priorities / a bad deadline
the hardest thing you have built / the thing you are
proudest of / the thing you would redo
That is thirty-plus questions. Nobody has thirty stories.
With sixteen well-chosen stories, each serving three to five themes, the space is covered, and the alternative is being caught without material for something you did not predict, which is where candidates freeze.
The second problem it solves is repetition. A candidate who tells the same story to four interviewers looks like they have one experience. The portfolio makes it deliberate: you know which stories you have used and can choose differently.
Mechanics
Choosing the sixteen
Not your sixteen favourite projects. Sixteen that cover the themes and demonstrate the level.
COVERAGE REQUIREMENTS (aim for at least one strong story in
each; several stories will cover several)
TECHNICAL DEPTH a decision only you could have made
TECHNICAL BREADTH a system you designed end to end
SCALE something large, with numbers
AMBIGUITY a problem with no clear right answer
FAILURE one that was your fault
RECOVERY an incident you led
CONFLICT (peer) a disagreement you resolved
CONFLICT (upward) a disagreement with a manager or
senior leader
INFLUENCE cross-team, without authority
MENTORING someone who grew, with the mechanism
UNDERPERFORMANCE a hard people conversation
HIRING a decision you made and why
PRIORITISATION something you chose not to do
CHANGED YOUR MIND with the evidence that changed it
LONG-HORIZON something that took quarters
ORGANISATIONAL a change to how the team worked
The three that candidates are most often missing:
Underperformance. Most senior engineers have never had the conversation, and staff loops ask about it because staff engineers influence people they do not manage. If you genuinely have no story, the closest legitimate substitute is a peer whose work you had to address, and being honest that you have not managed someone is better than inventing it.
Changed your mind. It is asked constantly and it needs a specific piece of evidence that changed your view. "I became more open to X over time" is not a story; "I ran the benchmark and it was 3x slower than I claimed, so I withdrew the proposal in the design review" is.
Something you chose not to do. Prioritisation stories are almost always told as "we shipped a lot", and the interesting version is what you killed and how you defended the decision.
The format: prepare in SCOR, deliver in STAR
PREPARE IN SCOR because it forces the two things that
make a story land:
Situation context, compressed
Complication what made it hard
Options what you considered and rejected,
with costs
Resolution what you chose and what happened
DELIVER IN STAR because that is the scoring rubric
Situation + Task including one explicit sentence about
what was YOURS to decide
Action the options compressed to one sentence
each, then what you did, in first
person singular
Result the number, and the lesson
The Options section is the staff signal and it is the one that disappears in a STAR telling. Compress each alternative to one sentence with its cost and keep all of them: "I evaluated scaling the cluster, which would have masked a cause I didn't understand; caching, which has near-zero hit rate for personalised results; and another day of root-causing with the regression visible." Twelve seconds, and it is the highest-density evidence of judgement in the whole story.
See SCOR, STAR and the scar-tissue story for the conversion mechanics.
The index: what makes it a portfolio rather than a pile
STORY THEMES NUMBERS
------------------------------------------------------------------
Auth migration across influence, conflict-peer, 4 teams,
4 teams long-horizon, ambiguity 2 quarters,
3 EMs
Search p99 regression technical depth, recovery, 4k QPS,
after personalisation changed-my-mind 180->1400
->210 ms
Killed the graph prioritisation, conflict- 6 months
database project upward, changed-my-mind saved, 1
VP annoyed
Priya to senior mentoring, organisational 18 months,
2 promoted
The reindex outage failure, recovery, 40 min,
organisational 97% cache
miss
...
Building this table is the actual preparation, and it takes an evening. Once it exists, the interview becomes a lookup: hear the question, identify the theme, pick a story you have not used, and tell the version that emphasises that theme.
Track which stories you have used with which interviewer, because a loop is four to six conversations and interviewers compare notes.
Rehearsing the numbers, not the sentences
For each story, memorise 4 to 5 figures:
the SCALE "4,000 queries per second across nine
locales"
the BEFORE "p99 was 180 ms"
the AFTER "210 ms, better than before the
regression"
one DETAIL only a "six shards for 80,000 documents"
participant knows
the DURATION "two days to find it"
The prose should vary between tellings and SHOULD. The
numbers must not, because inconsistency across a loop is
noticed and it is the fastest way to lose credibility.
The detail only a participant would know is the highest-value one. "Six shards for 80,000 documents" is not a fact anyone would recall unless they were there, and it does more for credibility than any adjective.
Timing, and the three lengths
Every story needs three versions and rehearsing only one is a common mistake.
30 SECONDS the scar-tissue version, dropped inside a
technical answer. Three sentences: what we did,
what went wrong with a number, what we do now.
90 SECONDS the standard behavioural answer. Full STAR,
one number per section, stop.
3 MINUTES the deep-dive version, when an interviewer says
"tell me more about that". Adds the options in
detail, the technical mechanics, and the
follow-on consequences.
The failure mode is telling the three-minute version when ninety seconds was asked for, and it is extremely common. An interviewer who wants more will ask; one who is waiting for you to stop will not interrupt, and the impression is that you cannot calibrate.
What makes a story fail
NO NUMBER nothing is verifiable, and it sounds
like a description of a category of
event rather than a memory.
NO "I" STAR scores individual contribution.
Say what the team did, then what you
did. This is rubric compliance, not
credit-taking.
NO DECISION a narration of events you were present
for. The Options section is what makes
it a decision.
NO RETROSPECT "what I'd do differently" is the
cheapest credibility available and
most candidates omit it.
TOO OLD a story from six years ago suggests
nothing has happened since. Prefer the
last two to three years.
TOO SMALL a story whose scope is one sprint does
not demonstrate staff level, however
well told.
A worked example: one story, four questions
THE EVENTS
A search personalisation launch caused a p99 regression
from 180 ms to 1.4 s, but only in small locales. Took two
days to find because the team looked at aggregate latency.
Cause: one query per shard per user segment, against
locales over-sharded at six shards for 80,000 documents.
Fixed by resharding and batching. p99 went to 210 ms.
Added per-locale alerting, which caught an unrelated
regression six weeks later.
"TELL ME ABOUT A TECHNICAL PROBLEM YOU SOLVED"
Lead with the inversion: the smallest indices were the
slowest, which ruled out capacity. Emphasise the diagnosis
and the options considered.
"TELL ME ABOUT A TIME YOU CHANGED YOUR MIND"
Lead with: I was convinced it was a capacity problem and
argued for scaling the cluster. Segmenting by locale
proved me wrong within an hour, and I withdrew the
proposal in front of the team that had been about to
approve the spend.
"TELL ME ABOUT PRESSURE FROM LEADERSHIP"
Lead with: the business asked daily whether to roll
personalisation back, and I asked for one more day with
the regression visible. Emphasise the negotiation, the
fallback plan I offered, and that I'd have rolled back if
the day produced nothing.
"TELL ME ABOUT A FAILURE"
Lead with: it took us two days to find, and it shouldn't
have. We were alerting on aggregate p99, which hid an
inversion that would have been obvious per locale on day
one. That's the failure, and the alerting change is what
came out of it.
Same events, four openings, four emphases, and the numbers are identical in all four. That is what a portfolio buys, and it is why sixteen stories cover thirty questions.
Production evidence
Amazon's Leadership Principles loop trains interviewers to collect STAR-structured evidence and to probe specifically for individual contribution, which is why the "we" to "I" conversion is mechanical rather than stylistic, and Amazon's own candidate guidance names STAR explicitly.
Google's published interview guidance tells candidates to use STAR and states that the interviewer is assessing what you did, which is the same rubric field.
Structured behavioural interviewing research consistently finds that structured questions predict job performance better than unstructured ones, which is why interviewers hold to the format: the structure is what makes candidates comparable, and a story that will not fit it is genuinely harder to score.
Barbara Minto's The Pyramid Principle is the source of the situation-complication-resolution structure, developed for consulting communication, and the reason leading with the complication works is that it makes the audience want the answer.
Will Larson's Staff Engineer contains interview accounts from staff engineers at many companies, and the recurring theme is that the stories that landed were about influence and judgement rather than about technical difficulty.
The debate
The case for a prepared portfolio: the question space is too large to anticipate, and under pressure people become vague. Preparation is what prevents that, and it is the single highest-return activity in behavioural preparation.
The case against over-preparing: rehearsed answers sound rehearsed, interviewers notice, and it reads as inauthentic. Some interviewers deliberately ask unusual questions to get past prepared material.
The case for preparing themes rather than stories: know what you want to convey and let the specifics come naturally. More flexible, and it fails under pressure precisely because specifics are what disappears when you are nervous.
My position: prepare sixteen stories in SCOR, index them by theme, rehearse the numbers rather than the sentences, and vary the prose deliberately between tellings.
The distinction between rehearsing numbers and rehearsing sentences is what resolves the over-preparation objection. A story told with identical wording sounds recited; a story told with identical figures and different wording sounds like a memory. So I would fix the four or five numbers per story and let everything else vary, which also means the story adapts to the question rather than being delivered regardless of it.
The three coverage gaps I would specifically hunt for are underperformance, changed my mind, and something you chose not to do, because they are asked constantly and most candidates have no material. For "changed my mind" the requirement is a specific piece of evidence, and "I became more open to X" is not a story. For underperformance, if you genuinely have not managed anyone, saying so and offering the closest real thing is better than constructing something, because an interviewer probing a fabricated people story finds the bottom of it in two questions.
The mechanical thing I would not skip is the index table, because it converts the interview from recall to lookup. Hearing a question and searching your memory for a relevant story is where candidates freeze; hearing a question, identifying the theme, and picking from a pre-tagged set is not. That table is an evening of work and it is worth more than rehearsing any individual answer.
And three lengths per story, because telling the three-minute version when ninety seconds was asked for is extremely common and reads as an inability to calibrate. An interviewer who wants more will ask.
Where I would push back on the framing: the portfolio is inventory, and the skill is retrieval. Twenty excellent stories with no index is worse than twelve indexed ones, because the failure under pressure is not lacking material, it is not finding it.
Follow-up Q&A
"Why a portfolio rather than answers to expected questions?" Because the question space is thirty-plus questions and nobody has thirty stories. Sixteen stories, each tagged with three to five themes it can serve, covers it, because the same events contain conflict, influence, failure and technical judgement, and the skill is knowing which thread to pull. Preparing answers requires anticipating the questions, which produces obviously-rehearsed answers to the ones you guessed and nothing for the ones you did not.
"How do you choose the sixteen?" By coverage, not by favourite. At least one strong story for each of: technical depth, breadth, scale, ambiguity, failure that was your fault, incident recovery, peer conflict, upward conflict, influence without authority, mentoring, underperformance, hiring, prioritisation, changing your mind, something long-horizon, and an organisational change. Several stories will cover several themes, which is the point.
"Which coverage gaps do candidates usually have?" Three. Underperformance, because most senior engineers have never had that conversation and staff loops ask because staff engineers influence people they do not manage. Changed your mind, which needs a specific piece of evidence, since "I became more open to X" is not a story. And something you chose not to do, because prioritisation stories are almost always told as "we shipped a lot" and the interesting version is what you killed.
"Doesn't this sound rehearsed?" Only if you rehearse the sentences. I fix four or five numbers per story, the scale, the before, the after, one detail only a participant would know, and the duration, and let the prose vary between tellings. Identical wording sounds recited; identical figures with different wording sounds like a memory. And it means the story adapts to the question rather than being delivered regardless of it.
"Show me how one story serves several questions." Take a search latency regression: p99 went from 180 milliseconds to 1.4 seconds, only in small locales, took two days to find. For "a technical problem you solved", lead with the inversion, that the smallest indices were slowest, which ruled out capacity. For "a time you changed your mind", lead with having argued for scaling the cluster and being proven wrong within an hour. For "pressure from leadership", lead with the business asking daily whether to roll back and me asking for one more day. For "a failure", lead with the two days, which was our alerting hiding an inversion that would have been obvious per locale. Same events, four openings, identical numbers.
"How long should a story be?" Three lengths, all rehearsed. Thirty seconds for the scar-tissue version dropped inside a technical answer. Ninety seconds for the standard behavioural answer. Three minutes for when someone says "tell me more". The common failure is giving the three-minute version when ninety seconds was asked for, which reads as an inability to calibrate, and an interviewer who wants more will ask.
"What makes a story fail?" Four things in order. No number, so nothing is verifiable and it sounds like a description of a category of event rather than a memory. No "I", because the rubric has a field for individual contribution and saying what the team did without saying what you did leaves it empty. No decision, so it is a narration of events you were present for, which is why the Options section matters. And no retrospect, which is the cheapest credibility available and most people omit it.
"What's the actual preparation work?" The index table: story, the three to five themes it serves, and its numbers. An evening. Once it exists the interview is a lookup rather than a recall problem, and recall is what fails under pressure. I would also track which stories I have used with which interviewer, because a loop is four to six conversations and interviewers compare notes, so telling the same story twice makes it look like you have one experience.
"What if you genuinely lack a story for a theme?" Say so and offer the closest real thing. For underperformance, if you have never managed anyone, "I haven't managed, but here's a peer whose work I had to address and how I handled it" is a good answer. Constructing something is much worse, because an interviewer probing a fabricated people story finds the bottom of it in two questions, and at that point the rest of the interview is spent recovering credibility.
Common misconceptions
"Prepare an answer for each likely question." The space is too large. Prepare stories and index them by theme.
"Rehearsing makes you sound fake." Rehearsing sentences does. Rehearsing numbers and varying the prose does the opposite.
"The best stories are the most technically impressive." At staff level the stories that land are about influence, judgement and decisions, and technical difficulty is the setting rather than the point.
"One great story is enough." A loop is four to six conversations and interviewers compare notes. Repeating a story makes it look like you have one experience.
"More detail is better." Beyond about a hundred seconds the interviewer stops tracking and starts waiting. Compress the Situation ruthlessly; it is the part everyone over-tells.
Interview delivery note
This topic is usually assessed indirectly, through whether your stories are good, but it does come up as "how do you prepare" and as advice-giving in a leadership conversation.
Frame it as inventory plus retrieval: "I'd build a portfolio rather than answers, because the question space is thirty-plus questions and nobody has thirty stories. Sixteen stories, each tagged with three to five themes, covers it, because the same events contain conflict, influence, failure and technical judgement. The skill is knowing which thread to pull."
Name the mechanism that makes it work under pressure: "And the actual artifact is an index table: story, themes it serves, and its numbers. That converts the interview from a recall problem to a lookup, and recall is what fails when you're nervous."
Give the over-preparation answer, because it is the obvious objection: "I'd rehearse the numbers, not the sentences. Four or five figures per story, fixed, and the prose varies every telling. Identical wording sounds recited; identical numbers with different wording sounds like a memory. And inconsistent numbers across a loop are noticed immediately."
The coverage point worth volunteering: "and I'd hunt specifically for the three gaps most people have: underperformance, changed-my-mind with a specific piece of evidence, and something you chose not to do. Those are asked constantly and most candidates have nothing."
Further reading
- Barbara Minto, The Pyramid Principle, for the situation-complication-resolution structure.
- Amazon's and Google's published interview preparation guidance, for how the STAR rubric is actually applied.
- Will Larson, Staff Engineer, for the interview accounts and what distinguished the stories that landed.
- SCOR, STAR and the scar-tissue story, for the format conversion mechanics.