Somewhere in the deck you are about to present, there is a slide with a number on it. Maturity goes up, cost goes down, or profitability goes up 20%. It came from a respected source, and you have seen it in three other decks this year.
Try to find the data underneath it.
I did, for every number that turns up in these conversations, and what I found is worse than a handful of loose citations. Enterprise architecture cannot currently prove its own worth. The field's own research names the claim that EA creates value as one of its persistent unsubstantiated myths, and the figures most often used to argue otherwise trace back to a paywall, an untraceable book claim, or a vendor measuring its own product.
That is not an academic embarrassment. It is why architecture practices lose budget arguments they should win, why an incoming CIO can dissolve a function that was doing good work and meet no resistance, and why a competent architect ends up defending a case that falls apart under one good question from a CTO who has read the footnote.
I do not think the conclusion is that architecture delivers nothing. I think the profession has spent thirty years trying to prove the wrong thing. There is a version of this argument that holds — it is narrower, it is harder to make, and it does not fit on a slide — and this article is the first of a short series that gets to it. What the evidence actually supports. Why the burden of proof has been in the wrong place all along. How you would measure the thing that matters. And what falls out of it.
This first piece clears the ground. It is the demolition, and it is not meant to be comfortable.
Disclosure, because this article would otherwise be exactly the thing it criticises: I am building an architecture maturity instrument, and I sell advisory work and products in the field it covers, though I have no commercial relationship with any platform named below. Whether the instrument itself ends up published openly or sold is not settled — either way my interest in it is commercial. Everything below applies to it too, and the closing section says how.
And a word on method, since this article spends its length demanding one of others. What follows is a source audit, not a systematic review. I took the claims that turn up in almost every maturity business case — the four-stage model, the 20% figure, the CMMI statistics, the framework verdicts — and followed each back until I reached data you can open, or ran out of trail. Where I ran out, I say so rather than rounding it up to a citation.
That method has a known blind spot, and it is worth stating before you rely on any of this: it finds what the field repeats. A well-evidenced claim that nobody cites would not show up here at all.
What the EA maturity literature actually says about business value
The most rigorously verifiable source on this question is a systematic literature review from TU Delft: Yiwei Gong and Marijn Janssen, "The value of and myths about enterprise architecture" (2019). Its finding, in the authors' own words:
"Only half of the articles provide empirical evidence supporting the EA value claims. Frequently, values are assumed to be the result of EA efforts, but many alternative explanations are possible."
The paper's first myth is stated plainly — "EA creates value" — and it concludes that enterprise architecture is "an instrument enabling the creation of value", not a value generator in its own right. That distinction is the whole article in one line. An instrument can be used well or badly. Its existence guarantees nothing.
Note what Gong and Janssen do not do: they never study maturity. The word appears once in the paper, in an appendix table. So the maturity-to-value claim is a further step out again, resting on a foundation their review already finds unproven. That is a harder position than most maturity business cases assume they are arguing from.
The review's search funnel is worth knowing, because you will need the precision when challenged. Searching Web of Science for "Enterprise Architecture" or "IT Architecture" across 2006–2016 returned 254 journal articles. After excluding inaccessible and non-English work, 199 remained. Of those 199, only 47 mention EA value at all. Within that 47: 11 offer no supporting material, 25 rely on citation-based support, and 18 carry empirical evidence, with some overlap between the last two groups.
A statistic phrased as "18 out of 199 articles provided empirical evidence" circulates online. The 18 genuinely is a subset of the 199, so it is a real ratio — but the denominator is doing all the work. The abstract says half. The funnel gives 18 of the 47 articles that discuss value at all, which is 38%, and 18 of the 199 analysed, which is 9%. Those are three different statements, and the paper does not reconcile them. Name your denominator, every time. That the field's most-cited review of EA value cannot be pinned to one number is not a footnote about sloppy citation. It is the finding.
Why the most-cited numbers cannot be checked
Three sources carry most of the weight in EA maturity business cases. Here is what happened when I tried to reach the data behind each.
Ross, Weill & Robertson, Enterprise Architecture as Strategy (2006). The four-stage model — Business Silos, Standardized Technology, Optimized Core, Business Modularity — is genuinely influential and the underlying 103-company survey is real. The stage-to-outcome numbers are another matter. I went looking through the book's public excerpts, MIT CISR's published material and secondary summaries, and did not find figures tying stages to profitability, cost or time-to-market that a reader could reconstruct. I cannot prove they are absent from the book. I can say I could not reach them, which for a business case amounts to the same problem.
Weill & Ross, IT Savvy (2009), the "IT-savvy companies are 20% more profitable" claim. This one is everywhere. It appears as an assertion in slide decks, vendor pages and consultancy reports. A sample size is published — roughly 1,800 organisations across 60-plus countries — but I could not trace how "savvy" or "profitability" were operationalised. A number whose method you cannot state is a number you cannot defend.
CMMI Institute's performance results — 9,500+ appraisals, double-digit improvements in defects and cost variance. Two problems, and the second is fatal. It is self-published by the body that sells CMMI appraisals, which is non-independence by construction. And it measures software process improvement, not enterprise architecture. A more credible relative exists in SEI's CMU/SEI-2006-TR-004, but SEI describes it in its own words as "brief case descriptions" created with representatives from ten organisations "that have achieved notable quantitative performance results" — a hand-picked showcase, not a controlled study.
There is a sharper irony in the CMMI record. SEI told acquirers, in writing, not to use levels the way the market went on to use them. Its own implementation guide states that "maturity level should never be used as a screening criterion at this stage" — and, in SEI's narrative rather than its sample wording, that "the SCE is not pass/fail — that is, being less than Defined maturity level will not exclude an offeror from the competitive range." 14 years later SEI acknowledged what had happened anyway, naming "misusing appraisal ratings as entry criteria" and observing that "in the rush to achieve a maturity level, the focus on improving organizational performance is lost." SEI itself put the phrase in quotation marks: "level mania". Demand for the rating never came from evidence that the rating predicted anything; it came from buyers treating it as an entry ticket.
There is a structural point hiding in all this. A large share of the EA maturity market — Gartner's IT Score, Forrester's equivalents, vendor tooling — makes value claims that are proprietary and paywalled by design. That is not an access inconvenience. It means those claims are structurally unverifiable by anyone outside the organisation selling them, which is a reason for scepticism rather than a reason to assume the data is fine behind the wall.
So there is one question to apply before you cite anything, and it is fast:
Can I get to the underlying data, or only to someone asserting the conclusion?
If the trail ends at a paywall, a book with no published dataset, or a report from the organisation selling the assessment, you have found an assertion. Assertions are fine in conversation and fatal in a business case, because the first person who checks becomes the most credible person in the room and it will not be you.
Apply it to this article too. Everything above links to something you can reach, and where it does not, I have said so.
What AI changed here, and what it did not
Something changed while I was doing this work, and it cuts against the instinct that better tools make source-checking obsolete.
A general-purpose model reproduces what the record most often says. The 20% figure is in slide decks, vendor pages and consultancy reports; the source check is in almost none of them. So when you ask for help justifying a maturity programme, what comes back is the field's most-repeated claims — fluently, in the register of something settled, and often with the hedging that the original authors did put in quietly removed. That is not the model inventing a number. It is the model being accurate about a literature that is itself unreliable, which is a harder problem.
The operational consequence is an asymmetry you will feel directly. The cost of producing a confident business case has fallen to almost nothing. The cost of checking one has not moved at all. Someone in the room can now generate a persuasive, well-structured, fully-referenced argument in a minute, and verifying its four load-bearing citations still takes an afternoon.
Which means the scarce skill is no longer assembling the case. It is knowing which claim to pull on first. That is a judgement, it does not scale, and it is the part of this job that has just become more valuable rather than less.
Which frameworks to use, and what each is actually good for
- GAO EAMMF (GAO-03-584G, 2003; expanded in GAO-10-846G, 2010). Public, primary, free, and the most directly on-point EA-specific source for how assessment becomes accountable governance. Use it for governance design. Do not use it as your maturity spine — it is built for US federal agencies.
- Van Steenbergen's DyAMM (Utrecht University, Sogeti-linked). Dutch, focus-area structured, and empirically validated against a dataset of 56 reliable assessments — one of very few in this field that can say that. Use it, particularly its checkpoint-question method for self-assessment. Note the wording: 56 assessments, not 56 organisations.
- TOGAF. Its Architecture Content Framework enumerates 73 distinct artifacts — 20 catalogs, 14 matrices, 33 diagrams, 4 maps and 2 tables. I counted them myself, in the live standard on 20 August 2026, because no total is published anywhere in it. That absence is its own small comment on the question. Kotusev's current framework, revised April 2026, identifies 34 artifact types actually found in practice — and reports that mature practices use 12 to 18 of them. So TOGAF prescribes four to six times what a working practice uses. TOGAF is right that tailoring is required, and says so itself, but gives almost no evidence-based guidance on which artifacts earn their keep. Its own Series Guide on maturity models (G203) states: "The benefits of using CMMs are well documented. Future versions of the TOGAF standard may include a maturity model to measure adoption of the TOGAF standard itself." The guide reproduces the US Department of Commerce's ACMM and CMMI instead, which implies it has no native maturity model of its own. Use TOGAF for vocabulary, not for maturity.
- US DoC ACMM (v1.2, December 2007). TOGAF's worked example, no longer maintained, hosting domain no longer serving it. Historically important, not something to run today.
- CMMI. The conceptual ancestor of the whole staged-maturity style, and genuinely not EA-specific — no EA domain, ever. Do not use as a spine.
- Gartner IT Score for Enterprise Architecture. Real, current, actively maintained, and paywalled. Fine to use if you already have the subscription; do not cite it externally, because your reader cannot check it and you have just failed your own source test.
And a note on the tools, because it is the question I get asked most. I went through the current documentation of the eight platforms that show up on most shortlists — SAP LeanIX, Ardoq, Bizzdesign (now also holding MEGA's Hopex and the former Software AG Alfabet), Orbus, Avolution, BlueDolphin. None ships a maturity model of the architecture practice inside the product. Several publish one on their website — LeanIX and Orbus both put a five-level model online — but what sits in the tool you actually work in is maturity scoring of the landscape: Alfabet scores business capabilities, Orbus the operating model, Avolution capability maturity over time. Useful, and not the same thing. Nor does any of them benchmark your practice against anyone else's — that sits with the analysts, and it is the one part of this market Gartner genuinely owns. Every one of them now markets AI over the repository — querying it, populating it, spotting inconsistencies — and none of that touches the maturity question either. Producing architecture content is getting cheaper. Establishing whether any of it changed a decision is not.
One qualification worth holding onto. Several platforms ship something that looks like architecture QA — LeanIX Quality Seal, Alfabet data quality rules, Orbus repository governance reports — but every one checks whether a record is complete and approved, not whether the architecture is any good. That is where most conversations about EA tooling quietly go wrong.
The same standard, applied to my own model
I am finishing an architecture maturity model built on the sources above. It faces every criticism in this article, and more worth naming. Pöppelbuß & Röglinger (ECIS 2011) and Mettler (2011) both argue that maturity models generally — not just EA ones — rest on weak theoretical foundations. Mettler names the failure modes directly: unnecessary bureaucracy, poor theoretical foundation, and "the impression of a falsified certainty to achieve success." Mine is not exempt from any of them.
So it makes no maturity-to-value claim, ranks every business-value argument by how much scrutiny it survives, and treats an organisation's own before-and-after evidence as stronger than any citation it contains. What that looks like in practice — the mechanism argument, and the three measures that carry it — is the companion piece to this one.
There is a gap in this field that no vendor is going to fill, because the honest answer is not sellable as a percentage: there is no good open data on what architecture maturity actually correlates with in practice. I would rather help produce some than keep citing around the hole. If you run or work in an architecture practice and would take part in a properly designed study — published instrument, stated sample size, disclosed conflict of interest, raw data released, findings published whether or not they support the thesis — that is what the next piece I am writing is about.
And if you are not waiting for research because the decision is already in front of you, that is advisory work — schedule a briefing. I would rather name that here than let it sit implied.
Including if the finding is that the link is not there. Especially then.
Common questions
Does this mean EA maturity models are worthless? No. It means they are instruments, not causes. Gong & Janssen's own conclusion is that EA is "an instrument enabling the creation of value." A maturity model tells you where you are and what to do next. It does not, by existing, produce a return — and the research does not support claiming that it does.
Is there any EA maturity model with real empirical validation behind it? Van Steenbergen's DyAMM is the strongest case — validated against a repository of 56 assessments, with a published methodology. Note the distinction: that validates the instrument, not a maturity-to-business-value relationship. Those are different claims, and the second one is the one that keeps failing.
What do I say when a stakeholder asks for the ROI number? Say the number does not survive a source check, and offer the mechanism instead. That is a long enough answer to need its own piece: the mechanism argument.
Sources & further reading
Every link below was checked on 2026-08-19. All resolve to something you can read without a subscription, except the three noted.
- The value of and myths about enterprise architecture — Gong & Janssen, International Journal of Information Management 46, pp. 1–9, Delft University of Technology, 2019 (the systematic literature review this article rests on)
- Information Technology: A Framework for Assessing and Improving Enterprise Architecture Management, Version 1.1 (GAO-03-584G) — US Government Accountability Office, 2003 (full text, govinfo)
- Organizational Transformation: A Framework for Assessing and Improving Enterprise Architecture Management, Version 2.0 (GAO-10-846G) — US Government Accountability Office, 2010
- The TOGAF® Standard, 10th Edition — Architecture Content, §3.6 Architectural Artifacts by ADM Phase — The Open Group (requires a free Open Group account; this is the section the 73 artifacts were counted from, in the build dated 2025-05-15)
- TOGAF® Series Guide: Architecture Maturity Models (G203) — The Open Group, February 2020 (requires a free Open Group account; the document is also mirrored openly here)
- The Practice of Enterprise Architecture: A Modern Approach to Business and IT Alignment — Svyatoslav Kotusev, 2021 (more than 1,700 EA publications reviewed, case studies across 35 organisations)
- The Dynamic Architecture Maturity Matrix: Instrument Analysis and Refinement — van Steenbergen, Schipper, Bos & Brinkkemper, ICSOC/ServiceWave 2009 Workshops, Springer LNCS, 2010 (closed access; the 56-assessment figure is stated in the dissertation below)
- Maturity and effectiveness of enterprise architecture — Marlies van Steenbergen, doctoral dissertation, Utrecht University, 2011 (the full DyAMM account)
- What Makes a Useful Maturity Model? — Pöppelbuß & Röglinger, ECIS 2011 Proceedings, AIS Electronic Library
- Maturity assessment models: a design science research approach — Mettler, International Journal of Society Systems Science, Vol. 3, Nos. 1/2, pp. 81–98, 2011 (open-access copy, University of St. Gallen)
- Performance Results of CMMI-Based Process Improvement (CMU/SEI-2006-TR-004) — Software Engineering Institute, Carnegie Mellon University, 2006 (SEI's own description: brief case descriptions from ten organisations that achieved notable results)
- Software Capability Evaluation Version 2.0 Implementation Guide (CMU/SEI-94-TR-5) — Software Engineering Institute, 1994 (the instruction not to use maturity level as a screening criterion)
- CMMI or Agile: Why Not Embrace Both! (CMU/SEI-2008-TN-003) — Software Engineering Institute, 2008 (SEI's own account of level mania and appraisal misuse)
- Gartner IT Score for Enterprise Architecture — Gartner (subscription required; listed because it is discussed, not as a citable source)
The survey of EA platform capability was done against each vendor's own current product and documentation pages in August 2026. Feature sets in this market change quickly; check the vendor before relying on it.
Deliberately not linked, because they are the point: the stage-to-outcome figures in Enterprise Architecture as Strategy (Ross, Weill & Robertson, 2006), the "20% more profitable" claim in IT Savvy (Weill & Ross, 2009), and CMMI Institute's aggregate appraisal statistics. Each is real and each is cited constantly. None could be traced to data a reader can check.