The question was about incentives. How does this organisation reward people for adopting AI?
The answer was honest, thoughtful, and entirely in the future tense. We are still figuring that out.
Under the rule, an answer in the future tense does not score. So it did not score. The finding went into the report, and a recommendation was written against it, the same way one would be for any client.
Everyone in the room knew the rule. The rule was ours.
That report has never been shown to a client. It runs on the same seven pillars, the same rubric and the same evidence standard as every report we issue. The organisation on its cover is ours.
By the time we sat those interviews, our own transformation was well underway. Projects were staffed across functions with named accountability owners. A directed AI literacy programme was running, with courses assigned by function. The people problem our own report recorded was that more employees wanted onto the projects than there were places for them. Nobody in the building needed persuading.
The reading still came back with gaps, because commitment is not evidence. That distance, between how ready an organisation believes it is and what its record supports, is the optimism gap. It is the subject this series opened with. We were not exempt from it.
Key points
-
Our transformation was already running when we assessed ourselves. The reading did not start it. It told us where we actually were.
-
Enthusiasm was our strength, not our gap. Commitment, staffing and training were in place. What was missing was evidence.
-
The gap that cost us most was not technical. It sat in how AI adoption is recognised and rewarded, and its recommendation pointed upstream at something that had to be working before the AI part could be addressed at all.
-
The internal run also changed the instrument. How we interview, how evidence is handled and how results are reported all moved. Some of it is in force, some is booked against the next reading.
-
For anyone assessing a portfolio: confidence and evidence are two separate measurements. Expect the gap between them to be widest in the teams that are trying hardest.
Commitment is not evidence
The optimism gap is easy to misread as a story about hype, or about executives who have not done the reading. Our own case suggests something less comfortable. The gap shows up most clearly in organisations that are genuinely committed, because commitment produces confidence faster than it produces evidence.
Every signal we would have pointed to was real. Teams volunteering. Training being consumed. Projects running with owners attached. None of that was theatre, and none of it was enough, because the questions do not ask what an organisation is doing. They ask what it can show, with an implemented example and an outcome attached.
That is a narrow standard and it is deliberately narrow. Under it, an initiative in flight scores lower than a smaller initiative already finished and measured. Plans do not count. Intentions do not count. An answer in the future tense, however sincere, does not count.
We wrote that rule believing it was fair. Sitting on the other side of it, we can report that it is fair and unwelcome at the same time.
The gap that cost us most was not technical
The incentives finding was the one that stung, and it is worth being precise about why. Nothing about it was an AI problem.
The recommendation written against that finding did not ask us to adopt a tool or build a capability. It pointed upstream, at a system of working practice that had to be genuinely in use before AI contribution could be written into how people are reviewed and progressed at all.
That is an uncomfortable finding to receive and a clarifying one. The gap could not be closed by an AI initiative, because the AI part was never the blocker. It also meant the work was slower and less visible than anything we could have announced. There was no launch to point at.
What we can point to is the sequence. The finding was kept rather than softened. The prerequisite was named rather than skipped. The work sits against the next reading, where the same question gets asked again, under the same rule, with the same consequence for an answer in the future tense.
We are not going to claim this one is closed. Under our own evidence standard, an assessor writing that a gap is closed would need to show the implemented example with outcomes, and we would fail our own request. The honest status is that the prerequisite is being worked and the next reading is the test.
The assessment changed the assessment
The internal run also changed the instrument, more than any design review had.
Some of that came from the practice interviews, before any client was involved. Whether the sessions worked better from slides, from open questions, or from pre-reads was argued in the abstract and settled in use. The question set ran through at least five iterations, and the standing rule was that each version had to be validated through tests and real interviews rather than read-throughs. Testing on ourselves meant every revision cost an afternoon rather than a client relationship.
More of it came from the completed assessment. In force now: a preparation guide, so anyone sitting the interview knows what evidence will be asked for before the first question. Written definitions of what each maturity level looks like for each pillar, so the same answer receives the same score regardless of who is listening, and so reports explain a level rather than asserting it. Transcript handling that is cleaned before anything is scored, with a validation step that checks whether the evidence being credited actually exists in the record rather than trusting a summary of it.
Booked against the next reading: a two-interviewer model, so answers are cross-checked and followed up in the room rather than reconstructed afterwards, and a move away from the current document format toward reporting a leadership team can navigate rather than read end to end.
None of that is glamorous. All of it is the kind of change that only surfaces when the people who wrote the questions have to answer them.
There is a quieter effect worth naming. Seven pillars, separate leaders, separate conversations, and then one document. Each account was reasonable on its own. Assembled, the distance between them was visible in a way it had not been before, which is most of what a readiness reading does for an organisation before it does anything else.
What this means for executives and investors
If readiness assessments inform your decisions, our own results suggest three habits.
Measure confidence and evidence separately. They are different quantities and they move at different speeds. A team that reports high readiness and cannot produce implemented examples is not being dishonest. It is describing its commitment accurately and its record optimistically.
Expect the gap to be widest where effort is highest. The organisations most likely to overstate readiness are the ones with the most activity in flight, because activity feels like progress and does not yet evidence like it. Treat a confident answer from a busy team as a prompt to ask for the example, not as a reason to relax.
Ask what the last reading changed. Not whether the methodology is sound. What moved, and what is still open with a date attached. Gap lists that only ever shrink are being managed for presentation. A live list with items on it is the sign of a process that is running.
The optimism gap closes the same way for everyone
This series has argued that readiness claims need evidence, that a score is a starting point rather than a judgment, that a trend tells you more than a snapshot, and that trust in a score is built in the design of the questions behind it.
We built the instrument that makes those arguments, and it found the same thing in us that it finds elsewhere. Real commitment, running further ahead of the record than we would have guessed. The assessment is the short part. It is a few weeks of interviews and a report. The long part is what happens next, when the gap list either changes how the organisation works or quietly stops being mentioned.
Ours is not a confession and not a credential. It is a working document with a test date, which is exactly what we tell every client a score should be. Some of it is closed. Some of it is not, and it will be asked about again.
An assessor with a perfect self-score has something to sell. An assessor with a maintained gap list has something to prove, on a cadence, with the next reading to answer to.
The gap closes the same way for everyone, one implemented example at a time. Including for the people holding the rubric.
This is the final post in a series on evidence-based AI readiness. The series began with the optimism gap, the distance between self-assessed and evidenced readiness, and ends here, with the same instrument pointed at its makers. Start the series from the beginning: The AI readiness optimism gap: why organisations score themselves higher than reality


