A score is a starting point: why readiness is measured as a trend, not an event

Executives around a boardroom table during a results meeting, viewed through a glass wall.
Content

The leadership team commissioned the AI readiness assessment eight weeks ago. Today is the readout. Twenty minutes in, the operating-model pillar comes up amber, and the COO wants it noted that the reorganisation lands next quarter. Two seats down, someone is already typing an email disputing the findings. Nobody has yet asked what amber means.

The assessment has not failed. It has been received as a verdict.

If you have sat through one of these readouts, you know the moment. The next ten minutes decide whether eight weeks of work produce a plan or an argument. And hidden inside that choice is a larger one, because a score received as a verdict gets filed, and a filed score decides nothing at all. This post covers how to design the reception of a score, and why the score only becomes valuable when it repeats.

Key points

  • People react emotionally to being scored. Kluger and DeNisi’s landmark meta-analysis found feedback improves performance on average, yet more than a third of interventions made performance worse, typically when attention shifts to the self instead of the task.

  • Reception can be designed. Colour bands in the room with precise scores held for internal tracking, and published definitions of what good looks like at every level.

  • A single score is a photograph. Its value appears at the second measurement, when a trajectory becomes visible, and readiness evidence is perishable enough that the trajectory is the only current thing you own.

  • Every reading ends with a roadmap to the next maturity level, and the next reading opens by testing it. The company answers for the effort. We answer for the advice.

  • For executives and investors, three practices follow: commission a cadence before the first score exists, watch how teams receive their results, and read deltas before levels.

The verdict reflex

That mid-meeting email is one of the most reliable findings in performance science. Kluger and DeNisi’s meta-analysis of feedback interventions, spanning 607 effect sizes and 23,663 observations, found that feedback improves performance on average, and that over a third of interventions actively made performance worse. The deciding variable was attention. Feedback that keeps people on the task helps. Feedback that pulls attention toward the self, toward “what does this number say about me”, backfires. A readiness score handed to a leadership team is aimed squarely at the self.

None of this is peculiar to AI readiness. Many executives have already lived it in a different instrument. An employee engagement survey lands, and the objections start before the discussion does. The sample was small. The timing was bad. Question nine was badly worded. The team was mid-restructure. Some of those objections are fair, and none of them change what the survey is for. Run it once and it is a complaint. Run it every year with the same questions and it becomes the only honest account an organisation has of whether it is getting better at being a place to work.

So the reflex will show up. The only question is whether the assessment is designed to absorb it or to feed it.

Reception can be designed

We learned this by watching it go sideways in a real room. In our early assessments, precise numeric scores produced confusion and contested findings instead of conversations about gaps. The redesign that followed changed how results are delivered, not the standard behind them.

External reports now lead with colour-band statuses, red, amber, and green, paired with named maturity levels rather than decimals. The precise numeric scores still exist, but they serve the assessment team’s internal tracking, where interviewers and graders weigh evidence in half-step increments. That resolution is real and useful behind the scenes. It is also more precision than a headline can honestly carry, and more than a leadership conversation can use. A colour invites a conversation about direction. A decimal invites a negotiation.

The statuses do the heavier lifting. Every maturity level carries a written definition of the behaviours that earn it, so when a pillar reads amber at Emerging, the room can see exactly what Developing requires. The difference between a 3.5 and a 4.0 is an argument. The difference between Emerging and Developing is a roadmap. And interviewees hear upfront what the assessment is, what it is not, and how the result should be read.

None of this softened the evidence bar, the implemented-examples standard that the first post in this series makes the case for. The redesign moved attention off the self and back onto the task.

A baseline becomes valuable at the second measurement

Here is the reframe that connects reception to value. A single score is a photograph, and arguing with a photograph is a strange use of an afternoon. The value compounds when the measurement repeats, because two honest photographs make a trajectory, and trajectory is what boards and investors actually allocate against.

Repetition is not optional housekeeping. Readiness pillars sound durable, but the evidence behind them is perishable. Sponsors move on, pilots get cancelled, key people leave, and six months later the report describes a company that no longer exists. This is also why arguing the first score upward is such an expensive win. An inflated baseline is a debt collected at every measurement after it. A modest, accurate one lets the second reading show real movement.

What happens between two measurements

Every reading in our framework ends with a roadmap. Each pillar closes with specific recommendations written against the published definition of the level above. The second measurement then arrives carrying an agenda: which recommendations were acted on, which moved the pillar, and which were tried but moved nothing. When a company does the work and the pillar does not move, the advice is what failed, and we say so. The second measurement tests both parties. The company answers for the effort. We answer for the advice.

None of this works if the loop can be gamed. So the instrument holds still, same pillars, same rubric, same level definitions, and effort alone moves nothing. A trend that goodwill could bend is not a trend an investor can steer by.

What this means for executives and investors

Somewhere in your portfolio, a company is drifting and does not know it, because its last snapshot said ready and nothing has asked the question since. Three practices follow:

Commission a cadence, not an event. Write the second measurement into the first engagement, before anyone has seen a score. Teams report differently to a measurement they know will return, and a roadmap with a test date is a commitment rather than advice.

Watch the reception, not just the result. A team that responds to amber with questions is a better bet than a team that responds to green with a press release. The posture toward the score predicts what the organisation will do about it. Fund the gaps you asked to see, or every future score becomes a negotiation.

Read deltas before levels. A level answers where a company stands. A delta answers where it is heading, and an allocation decision is a bet on where it is heading. Ask for last reading’s roadmap alongside this reading’s result. Levels flatter; deltas inform.

Three practices, one habit. Stop asking where a company stands. Start asking which way it is moving.

The meeting after the score

The organisations that get the most from a readiness assessment are rarely the ones with the best numbers. They are the ones where the meeting after the score is about sequencing work rather than contesting findings, and where that meeting happens again next reading, with last time’s roadmap open on the table. A score can only ever tell you where a company was standing when the shutter clicked. A cadence of honest scores tells you what it is becoming.

The score is the starting point. The trend is the distance travelled.

This is the second post in a series on evidence-based AI readiness. The next post covers the interview methodology behind the scores. Continue to the next post: Why we interview like Amazon hires: the method behind a readiness score you can trust

Recommended articles