The AI Maturity Assessment
Playbook
Why the Group is doing this, how it will run, and how to complete the assessment well.
The Group is José de Mello and its companies.
The working group is the people each company names to run its own assessment, join the session where learnings are shared, and take part in shaping the next edition of the baseline.
The gap between activity and value
If you are looking for certainty about AI, there is only one on offer: it is accelerating change — in business, in the economy, in society — at a speed without precedent, and nobody knows exactly where it leads.
What an organisation can know is where it stands. And right now, the market has split into two very unequal groups.
Nearly nine in ten organisations are deploying AI. Roughly six in a hundred can point to value on the P&L. The difference between the two groups is not the tools they bought. The high performers know where they stand, measure what matters, and redesign the work, not just the software.
A familiar management maxim, often attributed to Peter Drucker, says that what gets measured gets managed. A shared baseline makes movement visible: it helps the Group distinguish activity from progress, see where capability is strengthening, and decide where attention is most likely to create value.
And the market is not waiting. In the sectors closest to this portfolio, AI agents are past the pilot stage and operating at scale, with published, measured results:
made by autonomous agents across US health systems, rated ~9/10 by the patients themselves.
run with AI agents at Celanese, grown from a single process unit at a single plant.
These are results, not projections, and disruption rarely announces itself from inside the sector. Which makes the question for every company in this Group a practical one, not a philosophical one:
Which side of the gap are we on, and how would we know?
That question is what this initiative begins to answer. This AI Maturity Assessment is qualitative and contextual by design. Its value is the shared language and common starting point it gives very different companies. That shared language will underpin the separate AI Value Measurement that follows, which is not included in this tool. In that phase, each company can test what AI is actually changing and what value it is creating. The next chapter explains the thinking behind the first phase.
Measurement first: a baseline built to evolve
The usual way to start with AI is to pick the tools, launch the pilots, and wonder afterwards whether any of it worked. The leaders, the organisations that actually capture value from AI, choose the opposite order: measurement first. Before deciding what to build next, they agree on how they will know where they stand, and whether they are moving.
That order shows up clearly in the data:
The most valuable practice is also one of the least practised. High performers set outcome-based objectives tied to business KPIs and rigorously measure adoption, quality and results. The CFO finding sharpens the point: AI remains a cross-functional responsibility, but its value needs financial accountability and numbers credible enough for finance to stand behind. Measurement is not overhead. It is the discipline that separates the two groups in chapter 1.
So what does an assessment actually contribute? An honest answer first: every maturity assessment in the world is self-scored, which means every one of them is, in the end, a structured opinion. The number is not the product. The product is a shared language, and the value is the conversation it makes possible. For five entities that share almost nothing operationally, a common way to name things (what counts as an AI initiative, what "governed" means, what "value" means) is the asset everything else builds on.
What actually distinguishes organisations that scale AI with solidity are durable fundamentals: leadership that owns the agenda, work redesigned rather than decorated, people brought along, independence preserved, governance that enables, decisions informed by data, value that gets checked.
An assessment is, above all, a choice about what to pay attention to, and that choice was made thoughtfully. Each dimension earns its place because it is where value is being won or lost right now. Independence is a clear example: with AI capability concentrated in a handful of global providers and increasingly entangled in geopolitics, knowing what you own and what you merely rent has become a board-level question.
Those fundamentals are what this assessment measures. The baseline was distilled from the best of the best: the strongest current thinking across consulting firms, the technology firms deploying AI inside enterprises, AI natives, AI labs, and the European regulatory landscape. Chapter 5 shows exactly where it comes from.
And it is a baseline built to evolve. The entire purpose of this first exercise is to test it against the Group's reality: keep what fits, challenge what does not, and decide together what the next edition adopts, including what the Group chooses to pay attention to next. Chapter 6 explains how that works in practice.
The case was made on measurement, not on tooling. That is still the case we are making.
From baseline to proof, and back again
The two phases are connected, but they answer different questions. This AI Maturity Assessment asks where each company stands today. The AI Value Measurement asks what AI is changing and what that change is worth.
This tool creates a shared language. It gives very different companies a common way to describe their starting point, recognise strengths and gaps, and decide which questions deserve closer attention. Repeated over time, it shows how capabilities are evolving and where progress is visible.
The AI Value Measurement is separate and is not included in this tool. It will have its own tool and playbook. At a high level, its purpose is to give each operating company a fit-for-purpose way to demonstrate AI value while preserving portfolio autonomy.
Over time, the two tools can inform each other. What companies learn through AI Value Measurement can sharpen what the next maturity assessment pays attention to. Each new reading strengthens the next.
Each tool stands on its own and can be repeated when useful. Together, they create a long-lived measurement cycle: this AI Maturity Assessment establishes a shared baseline, the AI Value Measurement brings the value question into each company's reality, and the learning returns to the next assessment.
First we agree where we are, in one language. Then we prove what AI is worth.
Shared language, company choice
José de Mello committed to promote an AI Maturity Assessment within the Group as a first step towards building shared language and understanding.
The assessment is a self-assessment. Each company records where its practices stand on the same simple grid, at the depth it chooses. It may use the tool internally to go deeper with its own teams, or simply share the baseline practice results with the holding. Each company also decides what it brings into working group discussions, in service of shared language and learning.
Nothing you type leaves your device, there is no central database, and a completed assessment is uploaded into a private José de Mello environment.
The holding's role is equally specific. It convenes the working group to align language and discuss any learnings the companies choose to share.
One more property completes the contract. The assessment is periodic: it will be repeated, and what matters over time is direction.
Shared language comes first. It gives José de Mello a baseline for measuring progress over time and gives the portfolio a foundation for alignment and learning.
That principle is carried into the working agreement each entity records inside its own assessment, in its own words, before answering a single practice. The next chapter shows where the instrument itself comes from.
Where the dimensions come from
A fair question to ask of any assessment: who says these are the right things to measure? The answer here is: no single opinion. The instrument was built from an industry scan across five families of sources, chosen because each one sees a part of the picture the others miss.
At the centre of the scan sits business agility: the capacity to sense change, decide well, and redesign the work faster than the environment shifts. Every source family reads that capacity from a different angle, and where they agree is where the assessment looks.
What they agree on became eight dimensions. Dimensions describe organisational conditions, not departments; nobody owns a dimension. Each carries a conviction about how AI adoption actually works:
Underneath the eight dimensions sit 34 practices, each answerable with a simple state and, where useful, grounded in a real example or piece of evidence. Chapter 6 covers how that works when you sit down to complete it.
Nothing in these dimensions requires believing any vendor's roadmap. They are deliberately bread-and-butter: the conditions that keep paying off regardless of which tools win the season. The next chapter shows how to complete the assessment in practice.
Fundamentals that endure, not tool fashions.
Completing the assessment
The practical part is deliberately light. You open the assessment link in a browser. There is no account, no sign-in and nothing to install; your work is saved on your device as you type, and the file you export at the end is yours.
The first step is scope, and it is not a formality: you record the entity and the scope you are answering for, because the same answer means something different inside a single business unit than it does across a whole company. Then come the 34 practices. Each one is answered with a state: where that practice stands today, on a four-step scale.
There is a clear, shared view of how AI may affect the business, its people, and its customers or beneficiaries.
Bring the answer to life: Which concrete choice changed because of this view? What was deliberately left out?
The practice statement tells you what to assess. The follow-up question helps you distinguish a general statement about AI from a view that is genuinely steering choices.
The state is the answer. A short justification alongside it helps the reading; everything else you could write is optional and means it. When none of the four states fits, Chapter 7 explains the three alternative responses available.
How deep you go is a choice you make per practice, not a commitment you make as a company:
Depth is proportional to the decision, and you cannot get it wrong by going too shallow. A deeper answer raises confidence in the reading; it never raises the state. That rule is what keeps a portfolio-wide instrument from turning into a compliance exercise.
When an example names something real, you can record it as a row on the initiatives register: one row per thing actually going on. Software you have adopted goes to the separate tool register; buying a tool and running an initiative are different facts about an organisation. And the deep dive is for the three or four initiatives the conversation will actually turn on. You are not meant to deep-dive everything.
Nor is this a solo exercise. The best readings come from a small group of contributors who can answer across strategy, work, people, technology, governance, data and market, with one person responsible for validating the final reading. If the work has to pass between people, it travels as a file, one person at a time; the User Guide covers the mechanics of exporting, resuming and handing over.
The assessment closes with four questions that help each company decide what it may want to bring into working group discussions: where you would share how, where you would welcome support, where you disagree with a practice as written, and what you would want the portfolio to discuss. The distribution is the input to that conversation, not its conclusion.
A richer assessment makes a richer conversation. What the portfolio can learn starts with what each company chooses to record.
Where your philosophy differs, it is recorded and discussed
The baseline was built from the best of the best. But best of the best is not the same as right for this Group. Each company arrives with its own method, its own philosophy of how change happens, and its own regulatory reality. The assessment is designed to hear that, not to flatten it.
Alongside the scale, every practice accepts three answers that are not scores:
The third one is the invitation. If your philosophy differs from the conviction behind a practice, you do not need to force it into the scale or resolve the disagreement before completing the assessment. You record it, in the file, as an answer. Where several entities mark the same practice, that is a signal about the instrument, and it is treated as one.
The disagreements companies choose to bring into working group discussions help the portfolio decide what a future edition adopts: a practice reworded, a conviction challenged, a dimension rebalanced. Testing and evolving the baseline is the stated purpose of the first exercise.
That is also why the current edition stays fixed while everyone completes it. Every exported file records the exact wording it was answered against; change the wording mid-flight and the readings stop being comparable. Edition one is how edition two gets earned.
An instrument with nowhere to say "I don't know" produces confident fiction.
Use these answers wherever they are true. They preserve the context behind the reading and give the instrument a way to learn.
The opportunity
The companies in this Group operate in different sectors, with different customers, different regulators and different clocks. You do not compete with each other. In the corporate world that is a rare configuration, and it changes what can be shared.
Market peers show each other polished surfaces. Inside this Group, a company can show the parts that actually teach: what a redesign really took, what failed before something worked, what a capability costs to keep. The assessment gives that sharing a structure. Where one entity is further along in a dimension, the whole Group gains a shortcut. Where several entities carry the same gap, it can be solved once instead of five times.
And the absence of a ranking does not remove the pull of seeing what a peer has made work. It makes that pull useful. Nobody is defending a position in a league table; everyone can afford to be curious.
This is also where this exercise pays off. The assessment is periodic, and its baseline evolves with the Group, so every edition sharpens the shared language. The value measurement phase then turns that language into demonstrated value, each company in its own operational reality. Over time the Group builds the one thing none of its companies could build alone: a portfolio that learns as a whole.
Alone, each company learns at its own pace. Together, the Group learns at the pace of its furthest member.
That is the opportunity on the table. Complete the assessment, bring what you record, and take part in the conversation. This playbook stays with you; return to it whenever you need it.
before completing the assessment
Entity & Scope
Introduce your organisation and let us know if you are answering on behalf of the entire organisation or part of it.
The eight dimensions are conditions that run across the organisation. They are not departments, and no single team owns one. “Data for decision” is not the data team’s dimension, and “Proportionate governance” is not the legal team’s. Answer each one for the area you declared at the start, not for one team.
Distribution
Eight dimensions, read side by side. No total, no average, no ranking.
the reading, and what it is for
Profile & Reading
There is no total score, no entity average and no comparison with anyone else — the distribution is the input to a conversation, not its conclusion.
Summary
Distribution of practice state, by dimension
Answer depth and evidence, by dimension
The conversation this reading is for
Four questions to take into the conversation with the working group.