← Writing

AI · MBA

What Gartner's 28% actually says

A manager stands up in a board meeting and says only 28% of AI projects succeed. He has made three mistakes before he finishes the sentence. The number is real. What he thinks it measures is not what it measures, and the gap between the two is where budgets get funded, or killed, for the wrong reasons.

This essay defends nothing about AI spending, and it offers no fresh failure statistic to trade against the old ones. It takes apart a single number. The claim is narrow and, I think, defensible: 28% is not an "AI failure rate". It is the share of use cases inside one corner of IT that the people running them rated as meeting their own expectations. Before that figure earns a slide, you need to know what was counted, on whom, and by what threshold.

How a survey becomes a headline

Start with the source. As reported, Gartner surveyed 782 infrastructure and operations leaders across November and December 2025, and published the result on 7 April 2026. The headline line: 28% of AI use cases fully succeed and meet ROI expectations, while 20% fail outright.

One honesty note before we build on that. Gartner's newsroom blocks automated access, so the wording here comes from outlets that quoted the release, The Register and CIO among them. I flag it because the exact phrasing matters, and I have not read it on Gartner's own page. Everything below treats the number as reported, not as verified against the primary text.

Now take the sentence apart, one word at a time. The noun first. The 28% counts use cases, not projects and not companies. A use case is a single application of AI inside infrastructure and operations, such as auto-remediation, self-healing infrastructure, or an agent-led workflow. It is a slippery unit, because one company can run twenty of them, and one "project" can contain several. Count use cases and you count the small, numerous thing. That choice shapes the result before any respondent answers.

The same survey reports a second number that rarely travels with the first. As reported, 77% of leaders have at least one successful use case. Read the two together. At the level of the use case, the picture looks bleak. At the level of the organisation, most respondents already have something that works. One survey, two denominators, two moods. The headline picks the bleaker denominator, because a low number reads as a better story. Neither figure is wrong. They answer different questions, and the question is hidden in the choice of what to divide by.

A percentage without its denominator is not a fact. It is half of one.

Make it concrete. Imagine a hundred firms, each running four AI use cases in I&O. Suppose most firms get one of their four to work and the other three stall. Count by use case and roughly a quarter succeed, a grim-looking 25%. Count by firm and nearly all have a win, a cheerful figure north of 90%. Nothing changed except the unit of the count. The same underlying reality supports the pessimistic headline and the optimistic one, and a presenter can pick whichever fits the argument he walked in with. That is not fraud, only the ordinary freedom a choice of denominator hands to whoever holds the microphone.

Now the verb. "Meet ROI expectations" is a judgement, and the judge is the manager who owns the project. Nothing in the reporting says the return was measured independently. There is no audit trail, no profit-and-loss reconciliation, no third party checking the claim. "Success" here means an I&O leader told a surveyor that the thing worked. That is a legitimate instrument, and self-reported surveys are a normal way to read an industry. But self-report has a known direction of error. The release, as reported, carries no caveat about that bias and no published sampling or weighting frame. When the methodology sits behind a paywall, treat the output as a signal from a group, not a measurement of the world.

Then the adjectives, which are the softest part. "Success", "ROI", and "failure" are elastic words, and each survey stretches them differently. Gartner's success is "fully succeeds and meets ROI expectations". That bar is high and soft at once: high because it demands full success, soft because the expectation is set by the respondent. A leader who expected little and got a little counts the same as one who transformed an operation. Move the definition an inch and 28 becomes 40, or 20.

The number is downstream of a definition you were never shown.

There is a second thing a binary count hides, and it is where the money lives. A use case that shaved two percent off a cloud bill and one that automated an entire tier both register as a single "success". Two failed pilots that cost a thousand dollars each count the same as one that burned a million. The board cares about the size of the wins and the size of the losses, and the percentage is silent on both. A portfolio can be 28% successful and wildly profitable, or 72% successful and underwater, depending entirely on where the value and the cost happened to concentrate. Rate says nothing about weight.

None of this is an accusation. It is the ordinary anatomy of a survey statistic, and it applies to every figure in this debate, not only Gartner's. The point of walking through it is that the anatomy is invisible in the headline. By the time "only 28%" reaches a slide, the noun, the verb, and the adjectives have all been stripped, and what remains is a bare number wearing a certainty it never had.

Why the numbers disagree

If 28% were an AI failure rate, you would expect other studies to land near it. They do not. They land at 5%, at more than 80%, at 46%, at 30%. That spread is not measurement noise, and it is not a sign that one study is honest and the rest are junk. Each number counts a different thing, on a different population, against a different threshold. Line them up and the disagreement mostly dissolves, because they were never measuring the same quantity.

Take MIT first. The NANDA initiative published "The GenAI Divide: State of AI in Business 2025". Its instrument was 52 interviews, 153 executive surveys, and more than 300 public deployments. Its threshold was measurable impact on profit and loss. Its unit was a GenAI pilot or deployment across the whole company, not one corner of IT. The finding travels as "95% fail".

It is widely misread. The 95% refers to pilots without measurable P&L impact, not to companies. As Sify documented, "95% of companies failing" is a misreading of the report. The same report notes that roughly 90% of employees already use private LLMs at work, the shadow AI that no formal census captures. So the 5% "real impact" figure and the 90% "quiet usage" figure sit in the same document, describing different layers of the same firms. A study can find that almost no pilots clear a P&L bar and that almost everyone is using the tools. Both are true. They are not in tension once you see that they count different things.

Now RAND. It is quoted everywhere at "more than 80% of AI projects fail". Read the source. The report, RRA2680-1, phrases it as "by some estimates" — it is citing a figure, not producing one. RAND's own study is qualitative, built on interviews with data scientists and engineers. Read honestly, it says a large majority of projects struggle. It does not certify a precise 80%, and it is not a meta-analysis of a counted set of projects. The companion line, that AI projects fail "twice as often" as non-AI IT, is framing rather than a measured ratio.

A number inside quotation marks is not a measurement. It is a quotation.

S&P Global's Voice of the Enterprise surveyed 1,006 IT and line-of-business professionals across North America and Europe in 2025. Two numbers come out of it, and they are easy to confuse. On average, 46% of proof-of-concept projects are scrapped before production. Separately, the share of firms abandoning most of their initiatives rose from 17% to 42% year over year. 46% is not 42%, and neither one is a return figure. Both measure abandonment, which is a different construct from ROI. A PoC can be killed for a good reason, a bad reason, or because the team learned what it needed and moved on. Abandonment counts the exit, not the value.

Gartner has a second number in circulation, and it is the one most often quoted out of tense. At least 30% of generative AI projects will be abandoned after proof-of-concept by the end of 2025. That is a prediction, made in July 2024, attached to an estimated cost of five to twenty million dollars per deployment. Quoting a 2024 forecast as a 2026 measured result is a category error with a date stamp on it.

Now widen the lens, because the genre is older than AI. "X% of technology projects fail" has been a headline since the 1990s, and its most-cited source is the Standish Group's CHAOS report. Standish defines success as on-time and on-budget and on-scope, measured against the original estimates. Anything that overran cost, time, or scope is "challenged"; anything cancelled is "failed".

That definition has a problem, and Eveleens and Verhoef named it in IEEE Software in 2010. Working from 5,457 forecasts across 1,211 projects, they set out four faults. A success defined purely by the accuracy of an estimate is misleading. The measure is one-sided, so it understates success. Managing to that definition corrupts good estimation, because it rewards padding your forecasts so reality can beat them. And averaging figures of unknown bias produces numbers that mean little. The sampling frame, on top of all that, is closed and US-centric.

If a decades-old failure statistic does not survive scrutiny, a two-year-old one deserves the same knife.

So here is the shape of the disagreement. Five studies, five denominators: use case, pilot, project, firm, forecast. Five thresholds: self-rated ROI, measured P&L, "fails", abandonment, estimate accuracy. Five populations, from I&O leaders to interviewed engineers to line-of-business managers. Stacking 28 against 5 against 80 as "the AI failure rate" dresses a category error as a trend. Comparing them directly is like averaging a temperature, a distance, and a weight because all three came back as the number forty.

FigureWho / sampleUnit of analysisThreshold (what it counts)What it structurally omits
28% succeed (Gartner, 2026)782 I&O leaders, Nov–Dec 2025use case in I&Ofull success and ROI expectations — self-ratedbusiness/finance/end users, product AI, shadow AI, non-ROI value, ROI past the horizon
5% real impact / 95% not (MIT NANDA, 2025)52 interviews + 153 surveys + 300+ deploymentsGenAI pilot/deployment (whole firm)measurable P&L impactshadow AI (though the report describes it), non-tech sectors, public sector
>80% fail (RAND, 2024)qualitative interviews with practitionersAI project"fails" — a cited estimate, not a RAND countprecision of the figure (RAND itself says "large majority"), representativeness
46% PoCs / 42% firms (S&P, 2025)1,006 IT + LOB pros, N. America + EuropePoC / firmabandonment before production (not ROI)business vs technical cause, value of learning
≥30% abandoned (Gartner, 2024)— (a model/forecast)GenAI projectprediction of abandonment by end-2025it is not an ex-post measurement
~16–31% succeed (Standish, background)closed Standish databaseIT projecton-time + on-budget + on-scope vs estimatesvalue delivered despite overruns; sampling bias

How a board should read the number

None of this makes the number useless. It makes the number something to interrogate before it earns a slide. The interrogation is a short list of questions, and each one is designed to put a stripped word back.

  • What exactly was counted? Use case, pilot, project, firm, or forecast. The unit decides everything downstream, and it is the first thing the headline hides.
  • On whom? I&O leaders, the C-suite, interviewed practitioners. Then the harder question: does that population resemble us? A survey of infrastructure teams says little about a product company shipping a model to customers.
  • Who defined success, and was it measured or declared? Self-rating is not an audit. If the respondent set the bar and then reported clearing it, you have an opinion with a percentage on it.
  • Over what horizon? A project younger than the survey window counts as a failure today, even if its return is designed to arrive in eighteen to twenty-four months. The clock can manufacture the failure.
  • Measurement or forecast? Gartner's 30% is a prediction. Do not quote it as a result, and notice when someone else does.
  • Sourced or pasted? "More than 80% by some estimates" is not a measurement by whoever repeats it next. Trace the number to the body that produced it, or mark it as hearsay.
  • Cui bono? An advisory analyst, a governance vendor, a transformation consultant — each has a story to sell. Ask what the "failure" narrative is for before you accept its arithmetic.
  • Does the sample see the value you care about? Shadow AI, capability built, options bought, lessons paid for by a failed pilot. If the instrument cannot observe non-ROI value, its silence on that value is not evidence of absence.

The horizon question deserves more than a bullet, because it is the one that quietly manufactures failure. Enterprise technology returns rarely arrive on a straight line. Cost lands first, in licences, integration, and retraining, and the return, if it comes, arrives on a lag. Measure a portfolio while most of it sits in that early trough, and you read cost without return, which the instrument files as failure. A survey is a photograph, not a film. It cannot tell a project that is failing from one that is merely early, and in the month the shutter opens the two look identical.

With those in hand, read the Gartner set on its own terms, as a set rather than a headline. 28% of use cases fully succeed. 20% fail completely. That leaves a 52% middle the headline erases, use cases that are neither triumphs nor write-offs, most of them presumably still maturing. "Only 28% succeed" quietly implies that 72% failed. The survey's own second number says 20% did. The other 52% are doing what most work does, which is neither winning nor dying but grinding forward.

Then the most useful figure in the release, and the one least likely to reach a board. Of the leaders reporting a failure, 57% attributed it to "expected too much, too fast". Read that slowly. The single largest named cause of failure is not the technology. It is the expectation set around the technology. Gartner's own framing points the same way: as reported, return depended on integration with existing workflows, on governance, and on fit to a real operational need, not on the sophistication of the model. Roughly 38% pointed to weak data quality and availability, and integration shows up as a top barrier.

The survey is less a verdict on AI than a mirror held up to procurement.

There is one more discipline, and it is the one most often skipped. Ask what a survey of I&O leaders structurally cannot observe. It cannot see the business units, the CFO, or the end user, because it did not ask them. It cannot see product AI, the model inside a shipped feature that drives revenue, because that lives outside infrastructure and operations. It cannot see shadow AI, the private tools employees already lean on, because no one is surveying that. And it cannot price non-ROI value: the option you bought, the capability you built, the thing your team learned by failing a pilot cheaply. A number this narrow is a keyhole. Do not mistake it for the room.

Picture the board scene done well. Someone puts up "72% of AI projects fail". The right response is not another statistic. Ask a question instead: of what population, measured how, over what window, and blind to what? If the answer is "use cases inside I&O, self-rated, in a two-month window, at firms that are not ours", the slide has not been refuted. It has been priced. You now know exactly how much weight it can bear, which is some, and not the weight of a kill decision on its own.

Where the skepticism itself fails

A teardown of a statistic is satisfying, and it has its own failure mode. The methodological skeptic can talk himself into ignoring a number that, read honestly, still carries a signal. That move is its own kind of motivated reasoning, the reasoning of someone who wants to keep spending and has found a way to dismiss the inconvenient data. "It's just methodology" is the budget-holder's version of denial, and it is at least as dangerous as the naïve reading it corrects.

So state the honest reading plainly. The claim here is not that 28% is wrong, or that AI is working fine, but that it does not mean what the board slide says it means. Those are different claims, and only the last one is being defended here. The number is not invalidated by its limits. It is bounded by them.

Notice what survives the scrutiny. The instruments disagree on the figure, but they agree on the direction. MIT, RAND, S&P, and Gartner used different units, thresholds, and populations, and every one of them found that most AI initiatives do not yet clear their own return bar. You cannot average those numbers, and the category error still stands if you try. But convergence from independent, badly-aligned instruments is itself a form of evidence. When four crude thermometers all read "hot", you do not know the temperature, and you still should not put your hand on the stove.

Notice, too, what defends these surveys best. The "expected too much, too fast" finding barely depends on methodology at all. It is a statement about human expectation, and it would survive almost any reasonable definition of failure you substituted in. Gartner's structural claim, that return tracks integration and governance rather than model quality, is a genuine and useful result, independent of whether the headline reads 28 or 38. The person who defends this research well does not defend the headline percentage. He defends the findings sitting underneath it, which are sturdier than the number on top.

There is a subtler defence, one that cuts against my own thesis. Imprecise numbers coordinate. A board, a market, and a procurement team need a shared reference point more than they need a perfectly specified one. A figure that is directionally true and widely cited can do more work than a precise figure nobody has heard of. "28% of use cases" is becoming that kind of Schelling point for the maturity of enterprise AI, and it earns some of its influence honestly, by being roughly right about the direction. The teardown in this essay lowers the weight you should put on the number. It does not lower that weight to zero.

The disciplined posture sits between two errors. One error reads "72% fail" and concludes that AI is hype to be waited out. The other reads "just a survey" and concludes the data can be ignored while the spending continues. Both are motivated reasoning, aimed in opposite directions, and both feel like rigour to the person doing them. The number is neither a verdict nor noise. It is a partial, self-reported, narrow-population signal, and the job is to use it as exactly that, no heavier and no lighter.

There is a cost to over-correction worth naming directly. A founder or operator who learns to dismantle every uncomfortable statistic acquires a tool for never updating. Every inconvenient number has a denominator you can quibble with, a self-report you can distrust, a horizon you can call unfair. Deployed selectively, that skill becomes a machine for justifying whatever you already wanted to do. The scalpel that opens the Gartner number opens the case for your own project just as cleanly, and intellectual honesty means turning it on both.

The question to make it answer

So here is the discipline, compressed. Before a percentage goes on a slide, make it answer one question: of what? Of use cases or of companies. Measured or declared. Over what horizon, and blind to what. A statistic that cannot survive that question is not evidence. It is a mood with a decimal point.

The 28% is not the answer to "is AI working?". It is the start of a better question — working for whom, at what, measured how, and compared against what we already tolerate from every other technology project we have ever funded. Ask it that way, and the number stops being a headline and becomes what it always was: one narrow reading, honestly worth having, and dangerous only when it travels alone.

Sources

  1. primaryGartner, AI Projects in I&O Stall Ahead of Meaningful ROI Returns (press release, 2026) — (newsroom returned 403 to automated access; figures verified via the secondary outlets below, not against the primary page)
  2. primaryGartner, 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025 (press release, 2024).
  3. primaryMIT NANDA, The GenAI Divide: State of AI in Business 2025 (report, 2025).
  4. primaryRAND — Ryseff, De Bruhl, Newberry et al., The Root Causes of Failure for Artificial Intelligence Projects (RRA2680-1, 2024).
  5. primaryS&P Global Market Intelligence, Voice of the Enterprise: AI & Machine Learning, Use Cases 2025.
  6. primaryEveleens & Verhoef, The Rise and Fall of the Chaos Report Figures, IEEE Software 27(1):30–36 (2010).
  7. secondaryThe Register, Only 28% of AI infrastructure projects fully pay off (2026).
  8. secondaryCIO, AI often doesn't deliver ROI for IT departments either.
  9. secondarySify, 95% Companies Failing with AI? An MIT NANDA Report Misread by All.
  10. secondaryCIO Dive, AI project failure rates are on the rise: report.
  11. secondaryHenrico Dolfing, Project Failure Is Largely Misunderstood (critique of the Chaos report).
  12. secondarymetamorphOS, Why 72% of AI projects fail, and what Gartner's research isn't telling you(opinion/marketing; used here as framing, not as data)