Why we interview like Amazon hires: the method behind a readiness score you can trust

Laptop showing a group video call during an AI readiness assessment interview with several leaders on screen.
Content

The interviewer asks about data governance, and the answer arrives instantly, polished by a dozen board meetings. “We take governance extremely seriously. It is a top-three priority this year. The interviewer nods, writes nothing, and asks a smaller question. “Tell me about the last time a data problem stopped an AI project. What happened next?”

The room slows down. The prepared answer has nowhere to go.

If you have sat in either chair, you already know the difference between those two answers. The first is a message. The second, when it finally comes, is a memory. A score built on messages measures the messaging, not the readiness. This post covers how interview design keeps messages out of the score, and why that design is the reason a score deserves an investor’s trust.

Key points

  • A readiness score can only contain what its questions extract. A question that can be answered with a claim collects claims. A question that demands a specific, dated example collects evidence. The score is decided at question design, long before anyone answers.
  • Hiring science has tested this question hard. When the field corrected its own inflated estimates in 2022, structured interviews came out as the strongest single predictor of job performance, at roughly twice the validity of unstructured conversation. Our interviews run on the same discipline: fixed questions, written rubrics, behaviour-based prompts.
  • Stories beat claims. Our questions ask for specific, implemented examples with measurable outcomes. “We are planning to” does not score, because plans are messages and outcomes are evidence.
  • One chair cannot see a company. Questions are tailored by role, multiple leaders are interviewed per organisation, and two interviewers attend each session so evidence is gathered consistently.
  • The method is tested before it is trusted. The question set went through more than five iterations, and we ran the interviews on ourselves before any client sat the process.

A score inherits its interviews

Start with the failure mode. It is almost never lying. Ask an executive whether the company has a data governance framework and the answer, yes, will usually be true. There is a framework. It is a document. It was approved eighteen months ago, and no project has touched it since. The answer was honest, the score built on it is wrong, and nothing downstream can catch the error, because nothing false was ever said.

A vague question does exactly this. It lets a true answer carry a false picture. So the discipline of the interview was never about catching liars. It is about making honesty informative, and that cannot be done in the room. The informativeness has to be built into the question itself. When a prompt asks for the last specific instance, with dates and outcomes, a truthful answer has no room left to be empty. The picture arrives with the honesty, by construction.

That work is finished before anyone sits down. Which is why, for anyone allocating money against a readiness signal, the first question is not what the company scored. It is what the company was asked.

Borrowed from the science of hiring

Interviewing happens to be one of the most-studied measurement problems in organisational research, because hiring forced the issue long ago. And the field has been unusually willing to audit itself. In 2022, a team led by Paul Sackett re-examined the statistical corrections underlying decades of selection research and found they had systematically inflated the results. Most validity estimates came down. Structured interviews, where every candidate faces the same questions scored against a fixed rubric, landed at .42, ahead of cognitive ability tests, and roughly twice the .19 recorded for unstructured conversation. The numbers got smaller. The ranking of structure did not.

Amazon’s famous hiring process is what that structure looks like in practice. Candidates are asked for specific past events, not opinions, and answers are scored against defined criteria, the company’s written leadership principles. When we designed our interview set, Amazon’s methodology was the explicit model, pointed at organisations instead of candidates.

A caveat, because this series is fussy about evidence. The hiring research proves nothing about readiness assessments. It studied hiring. What it does show is that structure is what keeps an interview honest whenever the person answering has every reason to present well. A job candidate has rehearsed for a week. A leadership team describing its own company has rehearsed for years, to boards, to investors, to itself. So our method makes one bet. The discipline that keeps a week of rehearsal honest will hold against years of it, because structure does not care how polished the story is. It asks for the last specific example either way.

Stories are harder to polish than claims

The core rule of our interviews is that answers must be based on real examples, not future plans. “Yes, we have a governance framework” does not score. The answer that scores sounds different. “We implemented the framework in the second quarter. Data quality has improved since, time-to-access for AI projects has dropped from weeks to days, and the dashboard tracking both is reviewed monthly.” The gap between those two answers is the gap between a message and a memory, and you can usually hear it within seconds. A message arrives fluent and general, and would sound the same in any company. A memory arrives with texture, a specific quarter, a dashboard someone checks, a project that had a name.

Stories resist polish for a simple reason. A claim rests entirely on the person making it, and there is nothing else to check. A story is tangled up with reality. Other people were in the room, the dates sit in calendars, the consequences left traces in budgets and dashboards, and the next interview, two doors down, will describe the same events from a different chair. A polished story has to survive cross-examination it cannot see coming. That is why yes/no questions were removed from the set entirely. A yes has one author. A story has an audit trail.

A plan is the limiting case. However sincere, “we are planning to” has an author and nothing else, because nothing has happened yet. No witnesses, no dates, no traces. That is why it earns no score. The first post in this series called the wider pattern the optimism gap, organisations grading themselves on trajectory rather than position. The interview is where that gap gets closed, one requested example at a time.

Seeing a company takes more than one pair of eyes

That interview two doors down is not a spare part. It is the other half of the design, because a CEO, a CIO, and an operations lead do not see the same organisation. Strategy looks aligned from the top and contradictory from the middle. So the questions are tailored by role, and no assessment rests on a single interview, because one perspective, however senior, is a sample size of one.

Each session also seats two interviewers. One asks, one tracks evidence, clarifies acronyms, and catches what the flow of conversation would otherwise lose. The retro from our first assessments was blunt about why this matters: consistent evidence gathering is what makes scores comparable across interviews, and comparability is what makes a portfolio readable.

Triangulation is not a courtesy. It is how the method protects the score from the most charming person in the company.

What this means for executives and investors

Every readiness report that reaches you began as a set of conversations you did not attend. The questions were chosen, the answers were heard, and whether the trust you are now asked to extend is deserved was settled in rooms you will never see. You cannot sit in on those interviews. You can decide what you require of them. Three practices follow:

Ask what the company was asked. Before weighing any readiness score, request the methodology. Fixed questions, written rubrics, and evidence requirements are the difference between measurement and impressions. A real instrument is easy to show. A provider who cannot produce one is not offering measurement. They are offering impressions of a conversation, with a verdict attached.

Distrust single-voice verdicts. A readiness picture assembled from one executive interview is a self-portrait. Ask how many roles were interviewed and whether their accounts were tested against each other. Disagreement between chairs is not noise. It is usually the finding.

Listen for stories in your own reviews. The message-versus-memory test works in any boardroom, no methodology required. When a team reports readiness, ask for the last specific instance: the pilot that shipped, the issue that stopped one, the number that moved. What comes back, and how fast, will tell you most of what an assessment would.

Three practices, one instinct. Trust scores the way you trust people: by what they can show, not what they can say.

The question is the quality control

Go back to the room where this post started. The question that slowed everything down was small, specific, and almost gentle. And that pause was not the interview failing. It was the interview starting. The moment a prepared answer runs out is the moment information begins to flow, and every discipline in this method exists to reach that moment faster and more often.

That is what a trustworthy readiness signal is. Not a cleverer scoring model, but a conversation designed so that honesty and information arrive together. The status at the end means something because the questions made it mean something.

There is one more discipline behind the method, and it is the humblest one. Before any client sat these interviews, we sat them ourselves.

This is the third post in a series on evidence-based AI readiness. The next post covers what happened when we ran our own assessment on ourselves, and what dogfooding taught us about sitting on the other side of the rubric. Continue to the final post: We found our own optimism gap: What a readiness assessment showed about us

Recommended articles