If you drop the ROI curve, what do you actually put in front of a board?
The percentages usually cited to justify an architecture maturity programme do not survive a source check. They trace back to a paywall, an untraceable book claim, or a vendor measuring its own product. The full audit is in the companion piece.
This is the argument I use instead. It needs no statistics, and it is stronger than the number it replaces, because it is about your organisation rather than someone else's.
Disclosure: I am building an architecture maturity instrument, and I sell advisory work and products in the field it covers. The method below is the core of it. Whether the instrument itself ends up published openly or sold is not settled — either way my interest in it is commercial.
Argue the mechanism, not the curve
Maturity buys specific named failure modes avoided, and those are documentable in a way a percentage never is. Ranked by how much scrutiny each survives — which is itself the discipline this article is arguing for.
Best evidenced — avoided failure, not gained upside.
- The ivory tower. Svyatoslav Kotusev's empirical work — more than 1,700 EA publications reviewed, and a study of artifacts in 27 organisations — finds that maturity shows up in whether artifacts get used in real decisions, not in how many exist. His artifacts framework, revised to v3.0 in April 2026, identifies 34 artifact types found in practice and reports that mature practices use 12 to 18 of them — about half. A practice with 20 artifacts on the IT side alone is not more mature than one running 15 across both the business and IT sides. It is the failure mode Kotusev calls the so-called ivory tower syndrome, and its cost is rework and stalled decisions that no one attributes to architecture.
- Governance appointed but never built. In 2006 the US GAO applied its own EA maturity framework across 27 departments and agencies (GAO-06-831) and found the steering-committee and oversight requirement the weakest of the governance elements it singles out. Agencies had on average fully satisfied 77% of governance-related elements, and 57% satisfied the oversight-committee element fully, against 96% who had appointed a chief architect and 93% with a programme office. Naming a chief architect is nearly universal. Building the structure that reviews results and holds people to them is where four in ten implementations fall short. GAO's own conclusion: the key is sustained executive leadership.
- Reactive, incident-driven AI governance instead of a working intake gate. Two independent 2025 papers point the same way without citing each other. Kooy, Piest & Bemthuis (EDOC 2025) report that organisations with formalised AI policies, ethical oversight and traceability mechanisms are "more likely" to integrate generative AI into architecture workflows effectively. Alexander Ettinger's dissertation argues that architecture management, treated as a dynamic capability, "can enhance" GenAI adoption, partly by improving governance frameworks.
Moderately evidenced — better decisions, shown in cases, not proven statistically. Say so explicitly when you use these. Törmer & Henningsson (2017), both at Copenhagen Business School, found across three acquisition cases that pre-existing EA maturity enables distinct post-merger integration strategies — it shapes which approach is available to a deal team, not merely how well it executes one. Directly relevant if your technology decisions sit near transaction activity — but a three-case theory-building study, not a statistical finding.
Weakest evidenced. Strategic optionality from a modular architecture is the most intuitively appealing claim in the whole literature and the one I could least verify. Name it as a widely held view. Do not present it as demonstrated.
Not supportable at all: any specific percentage — profitability, cost, defect rate — tied to a maturity score.
How to build a case a CFO cannot dismantle
The mechanism argument only works if you can show the mechanism operating in your own organisation. That is a measurement problem, and it is solvable in a quarter. The instrument is a decision log, not a maturity score.
Instrument decisions, not artifacts. Pick five decisions your organisation keeps re-litigating from scratch — the ones that come back every 18 months with new faces and the same arguments. Write them down with dates. That list is simultaneously your baseline, your first Considerations artifact, and the only before/after evidence that will ever be genuinely yours.
For each decision, record four fields: what was decided, which artifact was consulted (or none), how long it took from raising to resolution, and whether it was revisited. 90 days of that beats every external citation in the companion article, because it is about your organisation and the CFO cannot argue the sample is unrepresentative.
Three concrete measures worth capturing, all of which are events rather than opinions:
| Measure | How you capture it | What it demonstrates |
|---|---|---|
| Decisions reopened within 12 months | Decision log, date raised vs date revisited | Whether artifacts are actually consulted |
| AI initiatives that reached production without passing an intake gate | Compare your gate's register against a list of what is live | Whether governance is real or nominal |
| Standard overrides, logged with reason | An exceptions register the architecture board reviews | Whether the standard is a standard or a suggestion |
Failure mode to expect: the programme reports maturity-level progress instead of these three. Levels rise because people learn the instrument. That is scoring drift, not improvement, and a CFO who has seen one maturity programme will recognise it.
CMMI publishes its appraisal register openly, and the shape of it is the argument. Filter it by achieved maturity level and you get 11,338 appraisals at level 3 and 3,461 at level 5 — against 135 at level 2 and 35 at level 4. Two of the five levels anyone actually aims for are all but empty.
Those are counts, not shares. I read them on 20 August 2026, when the register held 14,955 appraisals in total. Two caveats, because this article is in no position to skip them. The level filters overlap: the five of them plus the capability-level option sum to 14,985 against a register of 14,955, so they do not partition it and I am not going to present them as a distribution. And PARS posts a form rather than serving a linkable result, so you will have to run the filter yourself instead of following a link — which is a small irony in a piece about reaching the underlying data, and worth the thirty seconds.
Level 1 holds exactly one appraisal, which is not the anomaly it looks like: in CMMI level 1 is the un-appraised baseline, so nobody sets out to be appraised at it. Levels 2 and 4 have no such excuse.
A genuine spread of organisational capability does not have holes in it. A spread of what people want to be able to claim does.
That is the strongest argument I know for measuring on evidence rather than on self-assessment, and for not turning the result into a comparison. The moment a level becomes a claim, the claim is what improves. If your reporting cannot distinguish "we got better" from "we got better at answering the questionnaire," you have built the thing the companion article warns about.
What this changes for AI governance maturity
The reason this stopped being an academic argument in the last two years is that AI made the gate urgent.
Both 2025 papers cited above describe governance capability as an enabler of AI adoption. I would put it more strongly from practice: an organisation that cannot get a standard override logged is not going to catch a business unit wiring a model into a customer-facing process, because the same missing structure explains both. That is my inference, not their finding, and by this article's own test that makes it an argument rather than evidence. I use it because the mechanism is visible in any organisation that has tried.
What is not an inference is the deadline. The EU AI Act's Article 26 obligations on deployers of high-risk systems — human oversight, monitoring, keeping logs — apply from 2 December 2027 for stand-alone Annex III systems, on the schedule set out in the AI governance operating model piece. Whoever is supposed to exercise that oversight needs to see the system before it goes live, which is the same gate this whole argument is about.
So the question to put in front of a board is not "what does level 4 buy us." It is: when a business unit puts an AI capability into production next quarter, which named person sees it before it goes live, and what artifact do they consult? If there is no answer, you have a mechanism argument that needs no statistics, and a failure mode with a statutory date attached rather than a hypothetical cost. Who owns that gate is usually the first thing to settle.
That is also why the parts of the architect's job that AI is changing sharpen this rather than soften it. Producing artifacts is getting cheaper and faster. Deciding which ones matter, and making sure they are consulted, is not.
Your first 30 days
- Week 1 — write down the five decisions your organisation keeps re-litigating. No tooling, no template. A document with dates.
- Week 2 — open a decision log with the four fields above, and start with the next real decision instead of writing up old ones.
- Week 3 — list every AI initiative currently live, and compare it against whatever intake register exists. The gap is your finding, and it will be larger than expected.
- Week 4 — take that gap to whoever chairs your architecture governance. Not a maturity score. One number, on one page: how many things went live without anyone looking.
If you do nothing else from this article, do week 3. It takes an afternoon and it produces the only evidence in the whole exercise that no one will argue with.
If the decision is already in front of you and you would rather not build the case alone, that is advisory work — schedule a briefing.
Common questions
What do I say when a stakeholder asks for the ROI number? Say the number does not survive a source check, and offer the mechanism instead: here are three failure modes we can currently demonstrate, here is what each costs us when it happens, here is what changes if we fix them. Then show the decision log. Stakeholders are more tolerant of "we don't have that figure and here's why" than most architects expect, and considerably less tolerant of a figure that falls apart under one question.
Where does AI Act or NIS2 compliance fit into this? Compliance is a property of a working governance mechanism, not a separate programme. If AI initiatives reach production without passing a gate, that is simultaneously your maturity finding, your AI Act Article 26 exposure, and — for entities in scope — a gap in the NIS2 Article 21 risk-management measures that Article 20 makes the management body answerable for. The intake register is the same artifact in both conversations.
Is 90 days really enough to show anything? It is enough to show whether the mechanism exists, which is the question a board is actually asking. It is not enough to show a trend. Say which of the two you are presenting — conflating them is how maturity programmes lose credibility in the second year.
Sources & further reading
Checked on 2026-08-19. The full source audit, including the claims that failed, is in the companion article.
- The Practice of Enterprise Architecture: A Modern Approach to Business and IT Alignment — Svyatoslav Kotusev, 2021 (more than 1,700 EA publications reviewed, case studies across 35 organisations)
- Enterprise Architecture: Leadership Remains Key to Establishing and Leveraging Architectures for Organizational Transformation (GAO-06-831) — US Government Accountability Office, 14 August 2006 (the 57% oversight-committee finding)
- Impact and Implications of Generative AI for Enterprise Architects in Agile Environments: A Systematic Literature Review — Kooy, Piest & Bemthuis, University of Twente, arXiv, 2025 (to appear, EDOC 2025 Workshops)
- Enterprise Architecture as a Dynamic Capability for Scalable and Sustainable Generative AI Adoption — Ettinger, arXiv, 2025 (MBA dissertation, Warwick Business School; the author discloses employment at SAP LeanIX, and the sample skews toward that vendor's customer base — weigh accordingly)
- An Affordance Perspective on Generative AI in Enterprise Architecture — Banh, Tran & Strobel, IEEE Access 13, pp. 192711–192730, 2025 (an affordance study of what GenAI offers architects; it does not address governance maturity, and is listed here because it is often grouped with the two above)
- How Enterprise Architecture Maturity Enables Post-Merger IT Integration — Törmer & Henningsson, Copenhagen Business School, Business Informatics Research (BIR) 2017, Springer LNBIP vol. 295, pp. 16–30
- Regulation (EU) 2024/1689 (AI Act), Article 26 — obligations of deployers of high-risk AI systems
- Regulation (EU) 2026/1744 (Digital Omnibus on AI) — the amendment that moved Annex III high-risk obligations to 2 December 2027; the unamended 2024 text still carries the old date
- Directive (EU) 2022/2555 (NIS2), Articles 20 and 21 — management-body accountability, and the risk-management measures it answers for
- CMMI Published Appraisal Results System (PARS) — ISACA (the open appraisal register; figures in this article were read from it in August 2026 and change as appraisals are added and expire)