AI · MBA
What Gartner's 28% actually says
A manager stands up in a board meeting and says only 28% of AI projects succeed. He has made three mistakes before he finishes the sentence. The number is real. What he thinks it measures is not what it measures, and the gap between the two is where budgets get funded, or killed, for the wrong reasons.
This essay defends nothing about AI spending, and it offers no fresh failure statistic to trade against the old ones. It takes apart a single number. The claim is narrow and, I think, defensible: 28% is not an "AI failure rate". It is the share of use cases inside one corner of IT that the people running them rated as meeting their own expectations. Before that figure earns a slide, you need to know what was counted, on whom, and by what threshold.
How a survey becomes a headline
Start with the source. As reported, Gartner surveyed 782 infrastructure and operations leaders across November and December 2025, and published the result on 7 April 2026. The headline line: 28% of AI use cases fully succeed and meet ROI expectations, while 20% fail outright.
One honesty note before we build on that. Gartner's newsroom blocks automated access, so the wording here comes from outlets that quoted the release, The Register and CIO among them. I flag it because the exact phrasing matters, and I have not read it on Gartner's own page. Everything below treats the number as reported, not as verified against the primary text.
Now take the sentence apart, one word at a time. The noun first. The 28% counts use cases, not projects and not companies. A use case is a single application of AI inside infrastructure and operations, such as auto-remediation, self-healing infrastructure, or an agent-led workflow. It is a slippery unit, because one company can run twenty of them, and one "project" can contain several. Count use cases and you count the small, numerous thing. That choice shapes the result before any respondent answers.
The same survey reports a second number that rarely travels with the first. As reported, 77% of leaders have at least one successful use case. Read the two together. At the level of the use case, the picture looks bleak. At the level of the organisation, most respondents already have something that works. One survey, two denominators, two moods. The headline picks the bleaker denominator, because a low number reads as a better story. Neither figure is wrong. They answer different questions, and the question is hidden in the choice of what to divide by.
A percentage without its denominator is not a fact. It is half of one.
Make it concrete. Imagine a hundred firms, each running four AI use cases in I&O. Suppose most firms get one of their four to work and the other three stall. Count by use case and roughly a quarter succeed, a grim-looking 25%. Count by firm and nearly all have a win, a cheerful figure north of 90%. Nothing changed except the unit of the count. The same underlying reality supports the pessimistic headline and the optimistic one, and a presenter can pick whichever fits the argument he walked in with. That is not fraud, only the ordinary freedom a choice of denominator hands to whoever holds the microphone.
Now the verb. "Meet ROI expectations" is a judgement, and the judge is the manager who owns the project. Nothing in the reporting says the return was measured independently. There is no audit trail, no profit-and-loss reconciliation, no third party checking the claim. "Success" here means an I&O leader told a surveyor that the thing worked. That is a legitimate instrument, and self-reported surveys are a normal way to read an industry. But self-report has a known direction of error. The release, as reported, carries no caveat about that bias and no published sampling or weighting frame. When the methodology sits behind a paywall, treat the output as a signal from a group, not a measurement of the world.
Then the adjectives, which are the softest part. "Success", "ROI", and "failure" are elastic words, and each survey stretches them differently. Gartner's success is "fully succeeds and meets ROI expectations". That bar is high and soft at once: high because it demands full success, soft because the expectation is set by the respondent. A leader who expected little and got a little counts the same as one who transformed an operation. Move the definition an inch and 28 becomes 40, or 20.
The number is downstream of a definition you were never shown.
There is a second thing a binary count hides, and it is where the money lives. A use case that shaved two percent off a cloud bill and one that automated an entire tier both register as a single "success". Two failed pilots that cost a thousand dollars each count the same as one that burned a million. The board cares about the size of the wins and the size of the losses, and the percentage is silent on both. A portfolio can be 28% successful and wildly profitable, or 72% successful and underwater, depending entirely on where the value and the cost happened to concentrate. Rate says nothing about weight.
None of this is an accusation. It is the ordinary anatomy of a survey statistic, and it applies to every figure in this debate, not only Gartner's. The point of walking through it is that the anatomy is invisible in the headline. By the time "only 28%" reaches a slide, the noun, the verb, and the adjectives have all been stripped, and what remains is a bare number wearing a certainty it never had.
Why the numbers disagree
If 28% were an AI failure rate, you would expect other studies to land near it. They do not. They land at 5%, at more than 80%, at 46%, at 30%. That spread is not measurement noise, and it is not a sign that one study is honest and the rest are junk. Each number counts a different thing, on a different population, against a different threshold. Line them up and the disagreement mostly dissolves, because they were never measuring the same quantity.
Take MIT first. The NANDA initiative published "The GenAI Divide: State of AI in Business 2025". Its instrument was 52 interviews, 153 executive surveys, and more than 300 public deployments. Its threshold was measurable impact on profit and loss. Its unit was a GenAI pilot or deployment across the whole company, not one corner of IT. The finding travels as "95% fail".
It is widely misread. The 95% refers to pilots without measurable P&L impact, not to companies. As Sify documented, "95% of companies failing" is a misreading of the report. The same report notes that roughly 90% of employees already use private LLMs at work, the shadow AI that no formal census captures. So the 5% "real impact" figure and the 90% "quiet usage" figure sit in the same document, describing different layers of the same firms. A study can find that almost no pilots clear a P&L bar and that almost everyone is using the tools. Both are true. They are not in tension once you see that they count different things.
Now RAND. It is quoted everywhere at "more than 80% of AI projects fail". Read the source. The report, RRA2680-1, phrases it as "by some estimates" — it is citing a figure, not producing one. RAND's own study is qualitative, built on interviews with data scientists and engineers. Read honestly, it says a large majority of projects struggle. It does not certify a precise 80%, and it is not a meta-analysis of a counted set of projects. The companion line, that AI projects fail "twice as often" as non-AI IT, is framing rather than a measured ratio.
A number inside quotation marks is not a measurement. It is a quotation.
S&P Global's Voice of the Enterprise surveyed 1,006 IT and line-of-business professionals across North America and Europe in 2025. Two numbers come out of it, and they are easy to confuse. On average, 46% of proof-of-concept projects are scrapped before production. Separately, the share of firms abandoning most of their initiatives rose from 17% to 42% year over year. 46% is not 42%, and neither one is a return figure. Both measure abandonment, which is a different construct from ROI. A PoC can be killed for a good reason, a bad reason, or because the team learned what it needed and moved on. Abandonment counts the exit, not the value.
Gartner has a second number in circulation, and it is the one most often quoted out of tense. At least 30% of generative AI projects will be abandoned after proof-of-concept by the end of 2025. That is a prediction, made in July 2024, attached to an estimated cost of five to twenty million dollars per deployment. Quoting a 2024 forecast as a 2026 measured result is a category error with a date stamp on it.
Now widen the lens, because the genre is older than AI. "X% of technology projects fail" has been a headline since the 1990s, and its most-cited source is the Standish Group's CHAOS report. Standish defines success as on-time and on-budget and on-scope, measured against the original estimates. Anything that overran cost, time, or scope is "challenged"; anything cancelled is "failed".
That definition has a problem, and Eveleens and Verhoef named it in IEEE Software in 2010. Working from 5,457 forecasts across 1,211 projects, they set out four faults. A success defined purely by the accuracy of an estimate is misleading. The measure is one-sided, so it understates success. Managing to that definition corrupts good estimation, because it rewards padding your forecasts so reality can beat them. And averaging figures of unknown bias produces numbers that mean little. The sampling frame, on top of all that, is closed and US-centric.
If a decades-old failure statistic does not survive scrutiny, a two-year-old one deserves the same knife.
So here is the shape of the disagreement. Five studies, five denominators: use case, pilot, project, firm, forecast. Five thresholds: self-rated ROI, measured P&L, "fails", abandonment, estimate accuracy. Five populations, from I&O leaders to interviewed engineers to line-of-business managers. Stacking 28 against 5 against 80 as "the AI failure rate" dresses a category error as a trend. Comparing them directly is like averaging a temperature, a distance, and a weight because all three came back as the number forty.
| Figure | Who / sample | Unit of analysis | Threshold (what it counts) | What it structurally omits |
|---|---|---|---|---|
| 28% succeed (Gartner, 2026) | 782 I&O leaders, Nov–Dec 2025 | use case in I&O | full success and ROI expectations — self-rated | business/finance/end users, product AI, shadow AI, non-ROI value, ROI past the horizon |
| 5% real impact / 95% not (MIT NANDA, 2025) | 52 interviews + 153 surveys + 300+ deployments | GenAI pilot/deployment (whole firm) | measurable P&L impact | shadow AI (though the report describes it), non-tech sectors, public sector |
| >80% fail (RAND, 2024) | qualitative interviews with practitioners | AI project | "fails" — a cited estimate, not a RAND count | precision of the figure (RAND itself says "large majority"), representativeness |
| 46% PoCs / 42% firms (S&P, 2025) | 1,006 IT + LOB pros, N. America + Europe | PoC / firm | abandonment before production (not ROI) | business vs technical cause, value of learning |
| ≥30% abandoned (Gartner, 2024) | — (a model/forecast) | GenAI project | prediction of abandonment by end-2025 | it is not an ex-post measurement |
| ~16–31% succeed (Standish, background) | closed Standish database | IT project | on-time + on-budget + on-scope vs estimates | value delivered despite overruns; sampling bias |
How a board should read the number
None of this makes the number useless. It makes the number something to interrogate before it earns a slide. The interrogation is a short list of questions, and each one is designed to put a stripped word back.
- What exactly was counted? Use case, pilot, project, firm, or forecast. The unit decides everything downstream, and it is the first thing the headline hides.
- On whom? I&O leaders, the C-suite, interviewed practitioners. Then the harder question: does that population resemble us? A survey of infrastructure teams says little about a product company shipping a model to customers.
- Who defined success, and was it measured or declared? Self-rating is not an audit. If the respondent set the bar and then reported clearing it, you have an opinion with a percentage on it.
- Over what horizon? A project younger than the survey window counts as a failure today, even if its return is designed to arrive in eighteen to twenty-four months. The clock can manufacture the failure.
- Measurement or forecast? Gartner's 30% is a prediction. Do not quote it as a result, and notice when someone else does.
- Sourced or pasted? "More than 80% by some estimates" is not a measurement by whoever repeats it next. Trace the number to the body that produced it, or mark it as hearsay.
- Cui bono? An advisory analyst, a governance vendor, a transformation consultant — each has a story to sell. Ask what the "failure" narrative is for before you accept its arithmetic.
- Does the sample see the value you care about? Shadow AI, capability built, options bought, lessons paid for by a failed pilot. If the instrument cannot observe non-ROI value, its silence on that value is not evidence of absence.
The horizon question deserves more than a bullet, because it is the one that quietly manufactures failure. Enterprise technology returns rarely arrive on a straight line. Cost lands first, in licences, integration, and retraining, and the return, if it comes, arrives on a lag. Measure a portfolio while most of it sits in that early trough, and you read cost without return, which the instrument files as failure. A survey is a photograph, not a film. It cannot tell a project that is failing from one that is merely early, and in the month the shutter opens the two look identical.
With those in hand, read the Gartner set on its own terms, as a set rather than a headline. 28% of use cases fully succeed. 20% fail completely. That leaves a 52% middle the headline erases, use cases that are neither triumphs nor write-offs, most of them presumably still maturing. "Only 28% succeed" quietly implies that 72% failed. The survey's own second number says 20% did. The other 52% are doing what most work does, which is neither winning nor dying but grinding forward.
Then the most useful figure in the release, and the one least likely to reach a board. Of the leaders reporting a failure, 57% attributed it to "expected too much, too fast". Read that slowly. The single largest named cause of failure is not the technology. It is the expectation set around the technology. Gartner's own framing points the same way: as reported, return depended on integration with existing workflows, on governance, and on fit to a real operational need, not on the sophistication of the model. Roughly 38% pointed to weak data quality and availability, and integration shows up as a top barrier.
The survey is less a verdict on AI than a mirror held up to procurement.
There is one more discipline, and it is the one most often skipped. Ask what a survey of I&O leaders structurally cannot observe. It cannot see the business units, the CFO, or the end user, because it did not ask them. It cannot see product AI, the model inside a shipped feature that drives revenue, because that lives outside infrastructure and operations. It cannot see shadow AI, the private tools employees already lean on, because no one is surveying that. And it cannot price non-ROI value: the option you bought, the capability you built, the thing your team learned by failing a pilot cheaply. A number this narrow is a keyhole. Do not mistake it for the room.
Picture the board scene done well. Someone puts up "72% of AI projects fail". The right response is not another statistic. Ask a question instead: of what population, measured how, over what window, and blind to what? If the answer is "use cases inside I&O, self-rated, in a two-month window, at firms that are not ours", the slide has not been refuted. It has been priced. You now know exactly how much weight it can bear, which is some, and not the weight of a kill decision on its own.
Where the skepticism itself fails
A teardown of a statistic is satisfying, and it has its own failure mode. The methodological skeptic can talk himself into ignoring a number that, read honestly, still carries a signal. That move is its own kind of motivated reasoning, the reasoning of someone who wants to keep spending and has found a way to dismiss the inconvenient data. "It's just methodology" is the budget-holder's version of denial, and it is at least as dangerous as the naïve reading it corrects.
So state the honest reading plainly. The claim here is not that 28% is wrong, or that AI is working fine, but that it does not mean what the board slide says it means. Those are different claims, and only the last one is being defended here. The number is not invalidated by its limits. It is bounded by them.
Notice what survives the scrutiny. The instruments disagree on the figure, but they agree on the direction. MIT, RAND, S&P, and Gartner used different units, thresholds, and populations, and every one of them found that most AI initiatives do not yet clear their own return bar. You cannot average those numbers, and the category error still stands if you try. But convergence from independent, badly-aligned instruments is itself a form of evidence. When four crude thermometers all read "hot", you do not know the temperature, and you still should not put your hand on the stove.
Notice, too, what defends these surveys best. The "expected too much, too fast" finding barely depends on methodology at all. It is a statement about human expectation, and it would survive almost any reasonable definition of failure you substituted in. Gartner's structural claim, that return tracks integration and governance rather than model quality, is a genuine and useful result, independent of whether the headline reads 28 or 38. The person who defends this research well does not defend the headline percentage. He defends the findings sitting underneath it, which are sturdier than the number on top.
There is a subtler defence, one that cuts against my own thesis. Imprecise numbers coordinate. A board, a market, and a procurement team need a shared reference point more than they need a perfectly specified one. A figure that is directionally true and widely cited can do more work than a precise figure nobody has heard of. "28% of use cases" is becoming that kind of Schelling point for the maturity of enterprise AI, and it earns some of its influence honestly, by being roughly right about the direction. The teardown in this essay lowers the weight you should put on the number. It does not lower that weight to zero.
The disciplined posture sits between two errors. One error reads "72% fail" and concludes that AI is hype to be waited out. The other reads "just a survey" and concludes the data can be ignored while the spending continues. Both are motivated reasoning, aimed in opposite directions, and both feel like rigour to the person doing them. The number is neither a verdict nor noise. It is a partial, self-reported, narrow-population signal, and the job is to use it as exactly that, no heavier and no lighter.
There is a cost to over-correction worth naming directly. A founder or operator who learns to dismantle every uncomfortable statistic acquires a tool for never updating. Every inconvenient number has a denominator you can quibble with, a self-report you can distrust, a horizon you can call unfair. Deployed selectively, that skill becomes a machine for justifying whatever you already wanted to do. The scalpel that opens the Gartner number opens the case for your own project just as cleanly, and intellectual honesty means turning it on both.
The question to make it answer
So here is the discipline, compressed. Before a percentage goes on a slide, make it answer one question: of what? Of use cases or of companies. Measured or declared. Over what horizon, and blind to what. A statistic that cannot survive that question is not evidence. It is a mood with a decimal point.
The 28% is not the answer to "is AI working?". It is the start of a better question — working for whom, at what, measured how, and compared against what we already tolerate from every other technology project we have ever funded. Ask it that way, and the number stops being a headline and becomes what it always was: one narrow reading, honestly worth having, and dangerous only when it travels alone.
Sources
- primaryGartner, AI Projects in I&O Stall Ahead of Meaningful ROI Returns (press release, 2026) — (newsroom returned 403 to automated access; figures verified via the secondary outlets below, not against the primary page)
- primaryGartner, 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025 (press release, 2024).
- primaryMIT NANDA, The GenAI Divide: State of AI in Business 2025 (report, 2025).
- primaryRAND — Ryseff, De Bruhl, Newberry et al., The Root Causes of Failure for Artificial Intelligence Projects (RRA2680-1, 2024).
- primaryS&P Global Market Intelligence, Voice of the Enterprise: AI & Machine Learning, Use Cases 2025.
- primaryEveleens & Verhoef, The Rise and Fall of the Chaos Report Figures, IEEE Software 27(1):30–36 (2010).
- secondaryThe Register, Only 28% of AI infrastructure projects fully pay off (2026).
- secondaryCIO, AI often doesn't deliver ROI for IT departments either.
- secondarySify, 95% Companies Failing with AI? An MIT NANDA Report Misread by All.
- secondaryCIO Dive, AI project failure rates are on the rise: report.
- secondaryHenrico Dolfing, Project Failure Is Largely Misunderstood (critique of the Chaos report).
- secondarymetamorphOS, Why 72% of AI projects fail, and what Gartner's research isn't telling you — (opinion/marketing; used here as framing, not as data)
Co naprawdę mówi 28% Gartnera
Menedżer wstaje na posiedzeniu zarządu i mówi, że udaje się tylko 28% projektów AI. Popełnił trzy błędy, zanim skończył zdanie. Liczba jest prawdziwa. Tyle że mierzy co innego, niż on sądzi — a w tej różnicy zapadają decyzje o finansowaniu i ucinaniu budżetów, z niewłaściwych powodów.
Ten esej nie broni niczego w sprawie wydatków na AI i nie podrzuca nowej statystyki porażek w miejsce starej. Rozbiera jedną liczbę. Teza jest wąska i — jak sądzę — do obronienia: 28% mierzy nie „porażki AI", tylko odsetek use case'ów w jednym zakątku IT, które ludzie nimi zarządzający ocenili jako spełniające ich własne oczekiwania. Zanim ta liczba zasłuży na slajd, musisz wiedzieć, co policzono, na kim i przy jakim progu.
Jak z ankiety powstaje nagłówek
Zacznijmy od źródła. Gartner przepytał 782 liderów infrastruktury i operacji (I&O) w listopadzie i grudniu 2025, a wynik opublikował 7 kwietnia 2026. Nagłówek: 28% use case'ów AI odnosi pełny sukces i spełnia oczekiwania ROI, a 20% zawodzi całkowicie.
Jedna uwaga o źródle, zanim pójdziemy dalej. Liczby i cytaty pochodzą wprost z komunikatu Gartnera, opublikowanego 7 kwietnia 2026 w formie pytań i odpowiedzi z udziałem Melanie Freeze z zespołu badawczego. To wciąż komunikat prasowy, nie pełny raport: metodologia ankiety wraz z doborem i ważeniem próby siedzi za paywallem — i to ograniczenie będzie w tym eseju wracać.
Teraz rozłóżmy to zdanie na słowa. Najpierw rzeczownik. 28% liczy use case'y, nie projekty i nie firmy. Use case to pojedyncze zastosowanie AI w infrastrukturze i operacjach — auto-remediacja, samonaprawialna infrastruktura, proces prowadzony przez agenta. To jednostka śliska, bo jedna firma może mieć ich dwadzieścia, a jeden „projekt" — kilka. Kiedy liczysz use case'y, liczysz coś małego i licznego. Ten wybór kształtuje wynik, zanim odpowie pierwszy respondent.
Ta sama ankieta podaje drugą liczbę, która rzadko wędruje z pierwszą. 77% liderów ma co najmniej jeden udany use case. Przeczytajmy obie razem. Na poziomie use case'a obraz jest ponury. Na poziomie organizacji większość respondentów ma już coś, co działa. Jedna ankieta, dwa mianowniki, dwa nastroje. Nagłówek wybiera mianownik ciemniejszy, bo niska liczba to lepszy materiał na historię. Żadna z liczb nie jest błędna. Odpowiadają na różne pytania, a pytanie kryje się w wyborze tego, przez co dzielimy.
Procent bez mianownika nie jest faktem. Jest jego połową.
Ujmijmy to konkretnie. Wyobraźmy sobie sto firm, każda z czterema use case'ami AI w I&O. Załóżmy, że w większości firm jeden z czterech use case'ów działa, a pozostałe trzy stoją w miejscu. Policzmy według use case'a — udaje się mniej więcej jedna czwarta, ponuro wyglądające 25%. Policzmy według firmy — prawie każda ma trafienie, radosna liczba powyżej 90%. Nie zmieniło się nic poza jednostką liczenia. Ta sama rzeczywistość podpiera nagłówek pesymistyczny i optymistyczny, a prowadzący wybierze ten, który pasuje do tezy, z jaką wszedł na salę. Nie ma w tym oszustwa, jest zwykła swoboda, którą wybór mianownika daje temu, kto trzyma mikrofon.
Teraz czasownik. „Spełnia oczekiwania ROI" to ocena, a sędzią jest menedżer, który prowadzi projekt. Nic w komunikacie nie mówi, że zwrot zmierzono niezależnie. Nie ma śladu audytu, nie ma uzgodnienia z rachunkiem zysków i strat (P&L), nie ma trzeciej strony sprawdzającej deklarację. „Sukces" znaczy tu tyle, że lider I&O powiedział ankieterowi, że rzecz zadziałała. To uprawnione narzędzie — ankieta z samooceną (self-report) to normalny sposób czytania branży. Ale samoocena ma znany kierunek błędu. W komunikacie nie ma ani zastrzeżenia o tym obciążeniu, ani opublikowanego schematu doboru i ważenia próby. Gdy metodologia siedzi za paywallem, traktuj wynik jako sygnał od grupy, nie pomiar świata.
Potem przymiotniki — najmiększa część. „Sukces", „ROI" i „porażka" to słowa elastyczne, a każda ankieta rozciąga je inaczej. Sukces u Gartnera to „pełny sukces i spełnienie oczekiwań ROI". Ten próg jest naraz wysoki i miękki: wysoki, bo żąda pełnego sukcesu; miękki, bo oczekiwanie ustala sam respondent. Lider, który spodziewał się mało i dostał mało, liczy się tak samo jak ten, który przeobraził całą operację. Przesuń definicję odrobinę i 28 robi się 40 albo 20.
Liczba jest pochodną definicji, której ci nigdy nie pokazano.
Jest druga rzecz, którą liczenie zero-jedynkowe ukrywa — i to tam mieszkają pieniądze. Use case, który ściął dwa procent z rachunku za chmurę, i taki, który zautomatyzował całą warstwę, zapisują się oba jako jeden „sukces". Dwa nieudane piloty po tysiąc dolarów liczą się tak samo jak jeden, który spalił milion. Zarząd interesuje wielkość wygranych i wielkość strat, a procent milczy o obu. Portfel może być udany w 28% i szalenie rentowny albo udany w 72% i pod kreską — zależnie wyłącznie od tego, gdzie skupiły się wartość i koszt. Wskaźnik nie mówi nic o wadze.
Nic z tego nie jest oskarżeniem. To zwykła anatomia statystyki z ankiety i dotyczy każdej liczby w tym sporze, nie tylko tej od Gartnera. Sens przejścia przez nią jest taki, że w nagłówku ta anatomia jest niewidoczna. Zanim „tylko 28%" dotrze na slajd, rzeczownik, czasownik i przymiotniki zostają zdarte, a zostaje naga liczba ubrana w pewność, której nigdy nie miała.
Dlaczego te liczby się nie zgadzają
Gdyby 28% było wskaźnikiem porażek AI, inne badania lądowałyby blisko. Nie lądują. Lądują na 5%, na ponad 80%, na 46%, na 30%. Ten rozrzut to nie szum pomiaru ani znak, że jedno badanie jest uczciwe, a reszta to śmieć. Każda liczba liczy coś innego, na innej populacji, wobec innego progu. Ustaw je w rzędzie, a niezgoda w większości znika — bo nigdy nie mierzyły tej samej wielkości.
Najpierw MIT. Inicjatywa NANDA (Networked Agents and Decentralized AI) opublikowała raport The GenAI Divide: State of AI in Business 2025. Jej narzędziem były 52 wywiady, 153 ankiety kadry kierowniczej i ponad 300 publicznych wdrożeń. Progiem był mierzalny wpływ na rachunek zysków i strat. Jednostką był pilot lub wdrożenie GenAI w skali całej firmy, nie jeden zakątek IT. Wniosek wędruje jako „95% zawodzi".
Czyta się go powszechnie źle. 95% odnosi się do pilotów bez mierzalnego wpływu na P&L, nie do firm. Jak udokumentował Sify, „95% firm zawodzi" to błędne odczytanie raportu. Ten sam raport zauważa, że około 90% pracowników już używa prywatnych modeli językowych w pracy — to shadow AI, którego żaden formalny spis nie łapie. Liczba 5% „realnego wpływu" i liczba 90% „cichego użycia" sąsiadują więc w jednym dokumencie i opisują różne warstwy tych samych firm. Badanie może stwierdzić, że niemal żaden pilot nie przechodzi progu P&L, a zarazem że niemal każdy używa narzędzi. Jedno i drugie jest prawdą. Nie ma między nimi napięcia, gdy zobaczysz, że liczą co innego.
Teraz RAND. Cytowany wszędzie jako „ponad 80% projektów AI zawodzi". Przeczytajmy źródło. Raport RRA2680-1 ujmuje to jako „według niektórych szacunków" — cytuje liczbę, nie wytwarza jej. Własne badanie RAND jest jakościowe, oparte na wywiadach z data scientistami i inżynierami. Czytane uczciwie, mówi, że znaczna większość projektów grzęźnie. Nie poświadcza precyzyjnych 80% i nie jest meta-analizą policzonego zbioru projektów. Towarzysząca teza, że projekty AI zawodzą „dwa razy częściej" niż IT bez AI, to rama, a nie zmierzony stosunek.
Liczba w cudzysłowie nie jest pomiarem. Jest cytatem.
S&P Global w badaniu Voice of the Enterprise przepytał w 2025 roku 1006 specjalistów IT i biznesu (line-of-business) w Ameryce Północnej i Europie. Wychodzą z niego dwie liczby, łatwe do pomylenia. Średnio 46% projektów proof-of-concept (PoC) zostaje ubitych przed produkcją. Osobno: odsetek firm porzucających większość swoich inicjatyw wzrósł rok do roku z 17% do 42%. 46% to nie 42%, i żadna z tych liczb nie mówi o zwrocie. Obie mierzą porzucenie, a to inna kategoria niż ROI. PoC można ubić z dobrego powodu, ze złego albo dlatego, że zespół nauczył się, czego potrzebował, i ruszył dalej. Porzucenie liczy wyjście, nie wartość.
Obok krąży druga liczba Gartnera, cytowana niemal zawsze w złym czasie gramatycznym, jakby coś, co dopiero miało się wydarzyć, już się wydarzyło. W lipcu 2024 roku firma przewidziała, że do końca 2025 porzuconych zostanie co najmniej 30% projektów generatywnej AI po fazie proof-of-concept, z doczepionym szacunkiem kosztu wdrożenia: od pięciu do dwudziestu milionów dolarów. Kiedy ta prognoza wraca w 2026 roku podana jako zmierzony wynik, powstaje błąd kategorii z własnym datownikiem.
Teraz poszerzmy kadr, bo gatunek jest starszy niż AI. „X% projektów technologicznych zawodzi" to nagłówek od lat 90., a jego najczęściej cytowanym źródłem jest raport CHAOS grupy Standish. Standish definiuje sukces jako projekt w terminie, w budżecie i w zakresie, mierzony wobec pierwotnych szacunków. Cokolwiek przekroczyło koszt, czas albo zakres, jest „zagrożone"; cokolwiek anulowane — „nieudane".
Ta definicja ma problem, a Eveleens i Verhoef nazwali go w IEEE Software w 2010 roku. Pracując na 5457 prognozach z 1211 projektów, wyłożyli cztery wady. Sukces zdefiniowany wyłącznie przez trafność szacunku wprowadza w błąd. Miara jest jednostronna, więc zaniża sukces. Zarządzanie pod tę definicję psuje dobre szacowanie, bo nagradza pompowanie prognoz, żeby rzeczywistość mogła je pobić. A uśrednianie liczb o nieznanym obciążeniu daje liczby, które znaczą niewiele. Do tego schemat doboru próby jest zamknięty i skupiony na USA.
Jeśli statystyka porażek sprzed dekad nie przetrzymuje kontroli, ta sprzed dwóch lat zasługuje na ten sam nóż.
Oto kształt niezgody. Pięć badań, pięć mianowników: use case, pilot, projekt, firma, prognoza. Pięć progów: samoocena ROI, mierzony P&L, „porażka", porzucenie, trafność szacunku. Pięć populacji, od liderów I&O przez przepytanych inżynierów po menedżerów biznesu. Zestawianie 28 z 5 i z 80 jako „wskaźnika porażek AI" przebiera błąd kategorii za trend. Porównywać je wprost to jak uśredniać temperaturę, odległość i wagę, bo wszystkie trzy pomiary wyszły na czterdzieści.
| Liczba | Kto / próba | Jednostka analizy | Próg (co liczy) | Co strukturalnie pomija |
|---|---|---|---|---|
| 28% sukcesu (Gartner, 2026) | 782 liderów I&O, XI–XII 2025 | use case w I&O | pełny sukces i oczekiwania ROI — samoocena | biznes/finanse/użytkownicy końcowi, produktowe AI, shadow AI, wartość spoza ROI, ROI za horyzontem |
| 5% realnego wpływu / 95% bez (MIT NANDA, 2025) | 52 wywiady + 153 ankiety + 300+ wdrożeń | pilot/wdrożenie GenAI (cała firma) | mierzalny wpływ na P&L | shadow AI (choć raport go opisuje), sektory nietechnologiczne, sektor publiczny |
| >80% porażek (RAND, 2024) | jakościowe wywiady z praktykami | projekt AI | „porażka" — cytowany szacunek, nie policzenie RAND | precyzja liczby (sam RAND mówi „znaczna większość"), reprezentatywność |
| 46% PoC / 42% firm (S&P, 2025) | 1006 specjalistów IT + LOB, Ameryka Płn. + Europa | PoC / firma | porzucenie przed produkcją (nie ROI) | przyczyna biznesowa vs techniczna, wartość nauki |
| ≥30% porzuconych (Gartner, 2024) | — (model/prognoza) | projekt GenAI | prognoza porzucenia do końca 2025 | to nie jest pomiar ex post |
| ~16–31% sukcesu (Standish, tło) | zamknięta baza Standish | projekt IT | w terminie + w budżecie + w zakresie vs szacunki | wartość dostarczona mimo przekroczeń; obciążenie próby |
Jak zarząd powinien czytać tę liczbę
Nic z tego nie czyni liczby bezużyteczną. Czyni ją czymś, co trzeba przepytać, zanim zasłuży na slajd. Przesłuchanie to krótka lista pytań, a każde ma wstawić z powrotem jedno zdarte słowo.
- Co dokładnie policzono? Use case, pilot, projekt, firmę czy prognozę. Jednostka decyduje o wszystkim, co dalej, i jest pierwszą rzeczą, którą nagłówek ukrywa.
- Na kim? Liderach I&O, zarządzie, przepytanych praktykach. Potem pytanie trudniejsze: czy ta populacja przypomina nas? Ankieta wśród zespołów infrastruktury mówi mało o firmie produktowej, która dostarcza model klientom.
- Kto zdefiniował sukces i czy go zmierzono, czy zadeklarowano? Samoocena to nie audyt. Jeśli respondent sam ustawił próg, a potem zgłosił, że go przekroczył, masz opinię z doczepionym procentem.
- W jakim horyzoncie? Projekt młodszy niż okno ankiety liczy się dziś jako porażka, nawet jeśli zwrot ma przyjść za osiemnaście do dwudziestu czterech miesięcy. Zegar potrafi wyprodukować porażkę.
- Pomiar czy prognoza? 30% Gartnera to przewidywanie. Nie cytuj go jako wyniku i wyłap, gdy robi to ktoś inny.
- Ze źródła czy z kopiuj-wklej? „Ponad 80% według niektórych szacunków" nie staje się pomiarem w ustach kolejnego, kto to powtórzy. Prześledź liczbę wstecz, aż do instytucji, która ją wytworzyła, albo oznacz ją jako pogłoskę.
- Cui bono? Analityk doradczy, dostawca narzędzi governance, konsultant od transformacji — każdy ma historię na sprzedaż. Zapytaj, czemu służy narracja o „porażce", zanim przyjmiesz jej arytmetykę.
- Czy próba widzi wartość, na której ci zależy? Shadow AI, zbudowaną zdolność, kupione opcje, lekcje opłacone nieudanym pilotem. Jeśli narzędzie nie umie zaobserwować wartości spoza ROI, jego milczenie o niej nie jest dowodem jej braku.
Pytanie o horyzont zasługuje na więcej niż punkt listy, bo to ono po cichu produkuje porażkę. Zwroty z technologii w firmie rzadko przychodzą po linii prostej. Najpierw ląduje koszt licencji, integracji i przeszkolenia, a zwrot, jeśli w ogóle przyjdzie, przychodzi z opóźnieniem. Zmierz portfel, gdy większość siedzi w tym wczesnym dołku, a odczytasz koszt bez zwrotu, który narzędzie zaksięguje jako porażkę. Ankieta to zdjęcie, nie film. Nie odróżni projektu, który zawodzi, od takiego, na którego wynik jest po prostu za wcześnie — a w miesiącu, gdy otwiera się migawka, oba wyglądają identycznie.
Uzbrojeni w te pytania wróćmy do liczb Gartnera i przeczytajmy je tak, jak zostały podane — jako zestaw, nie nagłówek. 28% use case'ów odnosi pełny sukces. 20% zawodzi całkowicie. Zostaje 52% środka, które nagłówek wymazuje: use case'y ani triumfalne, ani spisane na straty, w większości pewnie wciąż dojrzewające. „Tylko 28% się udaje" po cichu sugeruje, że 72% zawiodło. Druga liczba tej samej ankiety mówi, że zawiodło 20%. Pozostałe 52% robi to, co większość pracy: ani nie wygrywa, ani nie umiera, tylko mozolnie sunie naprzód.
Potem najużyteczniejsza liczba w komunikacie — i ta, która najrzadziej dotrze do zarządu. 57% liderów zgłosiło co najmniej jedną porażkę, a jako jej powód wielu z nich — Gartner nie podaje odsetka — wskazało, że „oczekiwano zbyt wiele, zbyt szybko". Przeczytajmy to powoli. Powodem, który komunikat wymienia przed wszystkimi innymi, nie jest technologia, tylko to, jakie oczekiwania wokół niej ustawiono. Własna rama Gartnera wskazuje w tę samą stronę: zwrot zależał od integracji z istniejącymi procesami, od governance i od dopasowania do realnej potrzeby operacyjnej, nie od wyrafinowania modelu. 38% wskazało słabą jakość lub ograniczoną dostępność danych jako bezpośrednią przyczynę porażki i dokładnie tyle samo — utrzymujące się luki kompetencyjne.
Ta ankieta nie tyle wydaje wyrok na AI, ile podstawia lustro działom zakupów.
Jest jeszcze jedna dyscyplina, najczęściej pomijana. Zapytaj, czego ankieta wśród liderów I&O strukturalnie nie może zobaczyć. Nie widzi jednostek biznesowych, dyrektora finansowego ani użytkownika końcowego, bo ich nie zapytała. Nie widzi produktowego AI — modelu w dostarczonej funkcji, który napędza przychód — bo to żyje poza infrastrukturą i operacjami. Nie widzi shadow AI, prywatnych narzędzi, na których pracownicy już się opierają, bo nikt tego nie bada. I nie umie wycenić wartości spoza ROI: kupionej opcji, zbudowanej zdolności, tego, czego zespół nauczył się, tanio oblewając pilota. Liczba tak wąska to dziurka od klucza. Nie pomyl jej z pokojem.
Wyobraźmy sobie scenę na zarządzie zagraną dobrze. Ktoś wyświetla „72% projektów AI zawodzi". Właściwą odpowiedzią nie jest kolejna statystyka. Zadaj pytanie: na jakiej populacji, mierzone jak, w jakim oknie i ślepe na co? Jeśli odpowiedź brzmi „use case'y w I&O, samoocena, w dwumiesięcznym oknie, w firmach, którymi nie jesteśmy", slajd nie został obalony. Został wyceniony. Wiesz teraz dokładnie, ile ciężaru uniesie — trochę uniesie, ale nie ciężar decyzji o ubiciu projektu.
Gdzie zawodzi sam sceptycyzm
Rozbieranie statystyki daje satysfakcję i ma własny scenariusz porażki. Sceptyk metodologiczny potrafi wmówić sobie, że można zignorować liczbę, która czytana uczciwie wciąż niesie sygnał. Ten ruch to osobna odmiana umotywowanego rozumowania (motivated reasoning): rozumowanie kogoś, kto chce wydawać dalej i znalazł sposób, by odrzucić niewygodne dane. „To tylko metodologia" to wersja zaprzeczenia w ustach dysponenta budżetu i jest co najmniej tak groźna jak naiwny odczyt, który poprawia.
Wyłóżmy więc uczciwy odczyt wprost. Teza nie brzmi, że 28% jest błędne ani że AI działa świetnie — brzmi, że liczba nie znaczy tego, co mówi slajd zarządu. To różne tezy, a broniona jest tu tylko ostatnia. Granice nie unieważniają liczby. One ją ograniczają.
Zauważmy, co przetrzymuje kontrolę. Narzędzia nie zgadzają się co do liczby, ale zgadzają się co do kierunku. MIT, RAND, S&P i Gartner użyły różnych jednostek, progów i populacji, a każde z nich stwierdziło, że większość inicjatyw AI wciąż nie przekracza własnego progu zwrotu. Tych liczb nie da się uśrednić i błąd kategorii nie znika, gdy próbujesz. Ale zbieżność niezależnych, źle zestrojonych narzędzi jest sama w sobie formą dowodu. Gdy cztery prymitywne termometry pokazują „gorąco", nie znasz temperatury, a ręki na piecu i tak nie kładziesz.
Zauważmy też, co broni tych ankiet najlepiej. Wniosek „oczekiwano zbyt wiele, zbyt szybko" niemal wcale nie zależy od metodologii. To zdanie o ludzkim oczekiwaniu i przetrwałoby prawie każdą rozsądną definicję porażki, jaką byś podstawił. Strukturalna teza Gartnera — że zwrot idzie za integracją i governance, a nie za jakością modelu — to wynik prawdziwy i użyteczny, niezależny od tego, czy nagłówek mówi 28, czy 38. Kto broni tych badań dobrze, nie broni procentu z nagłówka. Broni wniosków leżących pod nim, a te są solidniejsze niż liczba na wierzchu.
Jest obrona subtelniejsza, cięta przeciw mojej własnej tezie. Liczby nieprecyzyjne koordynują. Zarządowi bardziej potrzebny jest wspólny punkt odniesienia niż idealnie doprecyzowany. Rynkowi i zespołowi zakupów tak samo. Liczba prawdziwa co do kierunku i szeroko cytowana zrobi więcej roboty niż precyzyjna, o której nikt nie słyszał. „28% use case'ów" staje się takim punktem Schellinga dla dojrzałości firmowego AI i część swojego wpływu zdobywa uczciwie, bo co do kierunku ma z grubsza rację. Rozbiórka w tym eseju obniża wagę, jaką powinieneś nadać liczbie. Nie obniża tej wagi do zera.
Zdyscyplinowana postawa leży między dwoma błędami. Jeden czyta „72% zawodzi" i wnioskuje, że AI to hype, który należy przeczekać. Drugi czyta „tylko ankieta" i wnioskuje, że dane można zignorować, a wydawać dalej. Oba to umotywowane rozumowanie, wycelowane w przeciwne strony, a temu, kto je uprawia, oba wydają się rygorem. Liczba nie jest ani wyrokiem, ani szumem. To sygnał cząstkowy, z samooceny, z wąskiej populacji, a zadanie polega na tym, żeby użyć go w tej właśnie roli, nie czyniąc go ani cięższym, ani lżejszym.
Nadkorekta ma koszt wart nazwania wprost. Founder albo operator, który uczy się rozkładać każdą niewygodną statystykę, zyskuje narzędzie, dzięki któremu nigdy nie musi aktualizować poglądów. Każda nieprzyjemna liczba ma mianownik, o który można się kłócić, samoocenę, której można nie ufać, horyzont, który można nazwać niesprawiedliwym. Stosowana wybiórczo ta umiejętność staje się maszyną do usprawiedliwiania tego, co i tak chciałeś zrobić. Skalpel, który otwiera liczbę Gartnera, otwiera równie czysto sprawę twojego własnego projektu — a uczciwość intelektualna każe ciąć nim w obie strony.
Pytanie, na które liczba ma odpowiedzieć
Oto dyscyplina w skrócie. Zanim procent trafi na slajd, każ mu odpowiedzieć na jedno pytanie: z czego? Z use case'ów czy z firm. Zmierzony czy zadeklarowany. W jakim horyzoncie i ślepy na co. Statystyka, która nie przetrzyma tego pytania, nie jest dowodem. Jest nastrojem z przecinkiem dziesiętnym.
28% nie jest odpowiedzią na pytanie „czy AI działa?". To początek lepszego pytania: dla kogo działa, w czym, jak to zmierzono i jak wypada na tle tego, co tolerujemy przy każdym innym projekcie technologicznym, jaki kiedykolwiek sfinansowaliśmy. Zadaj je tak, a liczba przestaje być nagłówkiem i staje się tym, czym zawsze była: jednym wąskim odczytem, który naprawdę warto mieć, groźnym tylko wtedy, gdy wędruje sam.
Bibliografia
- primaryGartner, AI Projects in I&O Stall Ahead of Meaningful ROI Returns (komunikat prasowy w formie pytań i odpowiedzi z udziałem Melanie Freeze, 7 kwietnia 2026) — (odczytany bezpośrednio z newsroomu Gartnera; pełny raport z metodologią pozostaje za paywallem)
- primaryGartner, 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025 (komunikat prasowy, 2024).
- primaryMIT NANDA, The GenAI Divide: State of AI in Business 2025 (raport, 2025).
- primaryRAND — Ryseff, De Bruhl, Newberry i in., The Root Causes of Failure for Artificial Intelligence Projects (RRA2680-1, 2024).
- primaryS&P Global Market Intelligence, Voice of the Enterprise: AI & Machine Learning, Use Cases 2025.
- primaryEveleens & Verhoef, The Rise and Fall of the Chaos Report Figures, IEEE Software 27(1):30–36 (2010).
- secondaryThe Register, Only 28% of AI infrastructure projects fully pay off (2026).
- secondaryCIO, AI often doesn't deliver ROI for IT departments either.
- secondarySify, 95% Companies Failing with AI? An MIT NANDA Report Misread by All.
- secondaryCIO Dive, AI project failure rates are on the rise: report.
- secondaryHenrico Dolfing, Project Failure Is Largely Misunderstood (krytyka raportu Chaos).
- secondarymetamorphOS, Why 72% of AI projects fail, and what Gartner's research isn't telling you — (opinia/marketing; użyte tu jako rama, nie jako dane)