AI · MBA
95, 80, 28: the war of AI failure numbers
Three numbers run most board conversations about AI this year. MIT's NANDA report says 95% of GenAI pilots show no measurable P&L impact. RAND is quoted for over 80% of AI projects failing, and Gartner reports that 28% of AI use cases fully succeed. They do not contradict one another, because they never counted the same object. The choice of what to count was made before any data was collected, and that choice is the story.
Failure statistics about AI are a publishing genre with its own economy. The methodology behind each number sits downstream of the incentives of whoever produced it. Read that as an analysis of incentives; nobody had to lie for the numbers to land where they did.
A companion piece takes one of these apart in detail: what the Gartner 28% actually says. This one is about the genre.
Who makes the numbers, and what they sell
Start with the balance sheet, which is public and explains the shape of the output.
Gartner closed FY2025 with revenue of $6,497m, up from $6,267m in 2024 — growth of about 3.7%. The Insights segment, formerly Research, is the largest contributor. Around 83% of revenue comes from recurring research subscriptions, with contract value in the region of $5.0–5.2bn during 2025, and consulting adding roughly 9%. These figures come from an aggregator of the 10-K, since SEC EDGAR returns a 403 to automated access, so treat them as reported. The business model is a subscription to decision-grade research, plus conferences, plus advisory.
The quotable number in the press release is the free sample; the product sits behind it. That is a description of a funnel. It is not the pay-to-play claim, which was litigated in NetScout v. Gartner and rejected in 2014. Nobody is buying a result. The point is subtler and harder to defend against: the number that reaches you has already been selected for being quotable.
The selection is visible in one producer's own output. Inside thirteen months Gartner published three figures. In June 2025 it forecast that more than 40% of agentic AI projects would be cancelled by the end of 2027. Days later it published a survey result: 45% of organisations with high AI maturity keep AI projects operational for at least three years. In April 2026 came the 28%. The first and third travelled everywhere. The 45% has almost no media life at all.
I have the titles and figures from a search index, because Gartner's newsroom also returns a 403, so I am not quoting the releases themselves. The pattern still holds. The pessimistic figures were not pushed harder than the optimistic one; they were picked up harder. Selection happens on the receiving end too.
The university lab has a product as well. The 95% figure came out of MIT Project NANDA — Networked Agents and Decentralized AI, based at the MIT Media Lab, with Ramesh Raskar as principal investigator. NANDA builds infrastructure for AI agents: a decentralised registry described as a DNS for agents, agent-to-agent protocols, identity and reputation verification. Its mission statement draws the analogy directly, comparing the missing agent layer to DNS and HTTP in the early web.
Now read the report's diagnosis next to that. Enterprise deployments fail, the report argues, because the systems lack memory and cannot learn from feedback. That is the same gap the NANDA infrastructure exists to fill. The report was not peer reviewed.
The consultant's version is the oldest one in the genre. BCG's October 2024 study, "Where's the Value in AI?", found that 74% of companies struggle to achieve and scale value from AI. Only 26% had the capabilities to move beyond proof of concept, and 4% were described as cutting-edge. The sample was 1,000 CxOs and senior executives across 59 countries and more than 20 sectors. Read the instrument carefully. It is a self-assessment of capability maturity, and it never claimed to measure project outcomes. It nonetheless circulates as a 74% failure rate, and BCG sells AI transformation.
The last exhibit is what happens when one of these numbers stops being a talking point. On 19 and 20 August 2025 AI-linked equities sold off: the Nasdaq fell more than 1.2%, Nvidia about 3.5%, Palantir close to 10%. Market commentary at the time named the MIT report alongside Sam Altman's bubble remarks as the proximate triggers. Both were in the same week, and the sell-off should not be pinned on the report alone. Still, a non-peer-reviewed working paper from a lab building agent infrastructure moved public markets, and its methodology did not travel with the headline number.
The number arrived. The denominator stayed home.
The five settings that decide the answer
Before any respondent is contacted, a study of failure has already chosen five things. Each choice moves the result by tens of points, and none of them is visible in a headline.
Unit of analysis. Use case, pilot, proof of concept, project, firm, or forecast. Six different denominators are in active circulation, and they are not convertible into one another.
Success threshold. Measurable P&L impact, reaching production, staying in production for three years, meeting a manager's ROI expectations, or matching the original estimate. Each threshold has its own literature and its own number.
Population. Infrastructure and operations leaders, C-suite respondents, interviewed practitioners, a fintech's customer base, or a nationally representative government sample.
Horizon. A project younger than the survey window counts as a failure by construction. An eighteen-month payback period registers as nothing at month nine.
Self-report versus measurement. A manager's judgement of their own ROI is not an audit, and the two can point in opposite directions. More on that below.
Set those five knobs and the answer is largely determined. The data collection that follows is expensive and careful, and mostly ornamental to the headline.
| Number | Producer and sample | Unit counted | Threshold | Structurally invisible to it |
|---|---|---|---|---|
| 28% fully succeed | Gartner, 782 I&O leaders, Nov–Dec 2025 | AI use case inside I&O | Full success plus ROI expectations, self-assessed | Business, finance, end users, product-side AI. The same survey reports 77% of leaders with at least one successful use case |
| 95% of pilots | MIT NANDA, 2025; not peer reviewed | GenAI pilot or deployment | Measurable P&L impact | Shadow AI, non-P&L value, anything below the measurement threshold |
| over 80% fail | RAND, quoted estimate inside a qualitative study | AI project | "Fails" — an estimate cited, not measured by RAND | Precision itself; RAND's own reading is "the large majority" |
| 46% of PoCs killed | S&P Global, 1,006 IT and line-of-business professionals | Proof of concept | Abandonment before production, not ROI | Whether abandonment was a business decision or a technical one. Distinct from the 42% of firms abandoning most initiatives, up from 17% |
| at least 30% abandoned | Gartner, July 2024 | GenAI project | Forecast to end of 2025 | It is not a measurement of anything that happened |
| 48% reach production | Gartner, May 2024, as reported | AI project | Deployment throughput; eight months from prototype to production | ROI, retention, abandonment — a sixth construct again |
Every row has its own denominator and its own threshold. Stacking them into a single "AI failure rate" is a category error, and the detailed dissection of the first row lives in the companion essay.
The genre is older than AI
None of this started in 2023. The template — an alarming percentage of technology projects fail — has been running since the mid-1990s, and it has already been taken apart twice by people with data.
The Standish Group's CHAOS report defines success as on time, on budget and on scope against the original estimate. Everything that overruns becomes "challenged"; everything cancelled becomes "failed".
Two independent critiques matter, and most citations of CHAOS acknowledge neither.
Moløkken-Østvold and Jørgensen, writing in Information and Software Technology 48(4) in 2006, examined the flagship 1994 figure: an average cost overrun of 189%. They found it far higher than comparable estimation studies of the same period, and traced the gap to a sampling method skewed toward failing projects. Their conclusion was that 189% is very likely much too high for typical 1990s projects. Continued use of the figure as a benchmark, they argued, produces bad decisions and holds back estimation practice.
Eveleens and Verhoef, in IEEE Software 27(1) in 2010, worked from 5,457 forecasts across 1,211 projects and raised four objections. A definition of success built purely on estimate accuracy is misleading. The measure is one-sided, so successes are systematically undercounted. Managing to the measure actively damages estimation practice, because the cheapest route to "success" is to inflate the original forecast. And averaging data of unknown bias produces figures with no interpretation.
That third objection has a name. Goodhart's law, in Marilyn Strathern's 1997 formulation, holds that "when a measure becomes a target, it ceases to be a good measure". Apply it forward. If your board starts reporting the percentage of successful AI use cases, your teams will start filing only the safe ones. The number will improve while nothing else does.
The same knife cuts 95, 80 and 28. It was sharpened on a report from 1994.
How a number mutates in transit
The genre has a transmission mechanism, and it is easiest to see on a single page.
Folio3 AI, a firm selling custom AI development and consulting, maintains a statistics page on AI project failure rates. Three of its entries repay a close look, purely as artefacts.
First: "80% of AI projects fail to deliver their intended business value", attributed to RAND's analysis of more than 2,400 enterprise AI initiatives. RAND analysed no such thing. The RAND report is a qualitative study built on interviews with practitioners, and its 80% appears as an estimate cited from elsewhere, hedged with "by some estimates". The figure acquired a sample it never had.
Second: "42% of companies abandoned at least one AI initiative in 2025", attributed to S&P. In the S&P data, 42% is the share of firms abandoning most of their initiatives, up from 17% the year before. Swapping the quantifier changes the construct entirely, and makes the number sound both smaller and more universal.
Third: "85% of AI project failures trace back to poor data quality (Gartner, 2025)". The origin is a Gartner forecast from 2018, which predicted that through 2022, 85% of AI projects would deliver erroneous outcomes because of bias in data, algorithms or the teams managing them. A forecast about wrong answers became a measured cause of failure, and 2018 became 2025. Three mutations on one page, all pointing the same way.
The canonical number of the genre never had a study behind it at all. "87% of data science projects never make it into production" comes from a July 2019 VentureBeat article which was sponsored content reporting a panel discussion at that publication's own conference. There was no sample and no source paper. The 78% figure that travels beside it came from a Dimensional Research survey commissioned by Alegion, a data-labelling vendor. The original article now returns a 403 to automated access, so the provenance above is as reported through secondary accounts. Dozens of MLOps vendors still cite it as the evidence base for buying their tooling.
The mechanism has been named since 1979. Beverly Houghton coined the Woozle effect for evidence by citation: a claim gains credibility through the frequency with which it is repeated rather than through any evidence for it. Her own description was reification by accretion. It remains in current methodological use: a 2024 paper in Electronic Commerce Research applies it directly to belief persistence in research literatures.
Repetition is not replication. It only sounds like it.
One counterexample from inside the same commercial genre is instructive. Pertama Partners, an AI training and advisory firm, runs a page headlined with an 80% failure rate for 2026. The body then states that RAND's study is qualitative and built on interviews. It is best read as "the large majority fail" rather than as a precise percentage, the page says, and it links the RAND report. The fine print is correct; the headline is alarmist; the page sells readiness audits and workshops. The incentive lives in the title and the call to action, not in the analysis.
Why the extreme number is the one that travels
It is tempting to explain the selection with editorial malice. The evidence points at something duller and harder to fix.
Soroka, Fournier and Nir, in PNAS 116(38) in 2019, ran a psychophysiological experiment across 17 countries on six continents using real news material. The average person shows stronger arousal responses to negative content than to positive content. It is the broadest cross-national demonstration of negativity bias available, and the authors are careful to note substantial individual variation.
Trussler and Soroka, in the International Journal of Press/Politics 19(3) in 2014, showed the behavioural half. In a controlled setting, readers select negative and cynical stories, regardless of what they say they prefer. Stated preference and revealed preference come apart.
No editor needs to conspire. An outlet optimising for clicks and a conference optimising for attendance will converge on the most extreme available figure by ordinary means. The optimistic 45% loses to the pessimistic 40% forecast for the same reason a quiet quarter loses to a scandal.
Then there is the one control that would settle any of this, and that the genre almost never uses. The MIT NANDA report did not go through peer review. Kevin Werbach of Wharton, reading it repeatedly, reported that he could not reconstruct where the 95% came from, and set a clean condition: publish the full supporting data, or retract the report. The report mentions 52 interviews and hundreds of data points without disclosing sample demographics or collection method. As of the research date for this essay, neither the data release nor the retraction has happened, and I have no evidence of either. Werbach's wording here is as reported, since the original outlet also returns a 403. Absence of review proves nothing about the number either way. It only removes the mechanism that would have caught an error.
What measurement gives you instead
Replace the survey with a measurement and the numbers do not converge. They fragment differently, and the fragments are more useful.
METR ran a randomised controlled trial published in July 2025: 16 experienced open-source developers, 246 real tasks in their own mature repositories, using early-2025 tooling. Developers forecast that AI would make them 24% faster. After finishing, they estimated they had been 20% faster. Measured, they were 19% slower. Self-assessment and measurement pointed in opposite directions, on the same tasks, for the same people.
That is the strongest available argument against reading any self-report survey as a measurement of return. It is also a small study, in a narrow context, with tools from the first half of 2025, and METR itself revised the experimental design in February 2026. It does not show that AI slows developers down. It shows that asking people and measuring them give different answers.
Measurement also runs strongly positive, in the same period, on different work. Brynjolfsson, Li and Raymond, in the Quarterly Journal of Economics 140(2) in 2025, used a staggered rollout of a GenAI assistant across 5,172 customer support agents. Productivity, measured as issues resolved per hour, rose 15% on average — 34% for novices and close to nothing for experienced agents. Customer sentiment improved and employee retention rose.
Side by side, the two studies look contradictory. They agree on the part that matters. For experienced practitioners, the measured effect is near zero or negative. For novices, it is large. Measurement yields no failure rate at all, only effects that depend on task, tooling and tenure.
Even adoption, which sounds like the simplest thing to count, splits by roughly thirty points depending on who counts it. The US Census Bureau's Business Trends and Outlook Survey is a nationally representative government instrument. It puts AI use at 17–20% of firms between mid-December 2025 and early May 2026. That rises to 37% among firms with 250 or more employees and falls below 20% for firms under 20. One supplement reports 18% of firms, but 32% when weighted by employment. The question itself changed in November 2025, from AI used in producing goods and services to AI used in any business function, which moved the count independently of any behaviour.
Ramp's AI Index measures something else entirely: corporate card and bill-pay transaction data across more than 70,000 businesses. It recorded 47.6% of firms paying for AI tools in February 2026, crossing 50% in March, against roughly 35% a year earlier. That measures paid commitment rather than use, on a sample of one fintech's customers, skewed toward technology firms. Ramp does not publish a sampling frame or representativeness caveat on the index page.
The gap between 17–20% and 50% is not a factual dispute. Two populations and two definitions of adoption produce two correct answers to two different questions. The mechanism is identical to 95 versus 80 versus 28.
Application: build your own reference class
The productive move is to stop importing anyone's number as a forecast for your own organisation.
The method has an author and a date. Bent Flyvbjerg set out reference class forecasting in "From Nobel Prize to Project Management: Getting Risks Right", Project Management Journal 37(3), August 2006. He developed the practice in "Curbing Optimism Bias and Strategic Misrepresentation in Planning", European Planning Studies 16(1), 2008. Three canonical steps:
- Identify a reference class of past, similar undertakings.
- Establish a probability distribution for that class on the parameter you are forecasting.
- Compare your specific undertaking with that distribution to establish the most likely outcome.
The parent theory is Kahneman and Tversky's distinction between the inside view and the outside view, developed further with Lovallo in 1993. The inside view uses the details of your project. The outside view uses the distribution of outcomes for similar projects and deliberately ignores your details. Flyvbjerg's contribution is that this corrects two different errors at once: optimism bias, which is unconscious, and strategic misrepresentation, which is the deliberate understatement of cost to get a project approved.
That distinction transfers cleanly to the producers of failure numbers. Some of the skew is cognitive. Some of it is an editorial choice with a business behind it. The remedies differ.
The method was first applied in practice to the Edinburgh Tram Line 2 and later to Crossrail for the UK Department for Transport. Its advantage is largest for non-routine undertakings, which is what enterprise AI is for an organisation attempting it the first time.
Hold two base rates while you build your own.
Flyvbjerg's database, assembled over decades at Oxford's Saïd Business School, covers more than 16,000 large projects. As reported in How Big Things Get Done (2023), 8.5% come in on budget and on time, and 0.5% come in on budget, on time and deliver the promised benefits. That figure comes from a trade book rather than a peer-reviewed paper, so treat it as reported. The implication is uncomfortable for the whole genre. "Most initiatives fail to deliver what was promised" is the base rate for projects in general, not a peculiarity of AI.
Flyvbjerg and Budzier, in Harvard Business Review 89(9), September 2011, studied 1,471 IT projects. Average cost overrun: 27%. But one project in six was a black swan, with an average cost overrun of 200% and a schedule overrun near 70%. The mean describes almost nobody. The tail decides whether the company survives the project.
The entire war over 95 versus 80 versus 28 is being fought over a parameter that was the wrong one to ask for.
Here is the procedure I would put in front of a board, in order.
- Define the unit before looking at any number. Use case, pilot, project or firm. Without this, nothing you read is comparable to anything you run.
- Define the success threshold ex ante and in writing. Measurable P&L impact? Reaching production? Still running after three years? Each threshold has a different number attached to it in the literature, and picking one afterwards is how projects get graded on a curve.
- Collect your own reference class. Ten to thirty closed initiatives from the last three to five years inside your own organisation — not only AI, but any change project of comparable complexity and novelty.
- Build a distribution, not an average. Cost, time and benefit against the original forecast. Record the median and the P80 or P90.
- Locate the new project inside that distribution and write down the resulting uplift on cost and on time. This is step three of reference class forecasting, and it is the step people skip.
- Separate optimism bias from strategic misrepresentation. The first is treated with data. The second is treated by changing incentives: check whether the person promoted for getting a project approved is the person accountable for delivering it.
- Price the tail before the mean. If one IT project in six overruns by 200%, the first question is what you do if this is that one, not what the average looks like.
- Use external numbers only as an external reference class, with the denominator stated out loud, and never as a forecast for yourself.
Steps three and four are the whole exercise. The rest is bookkeeping.
Most organisations cannot execute step three, and that is the real finding. Suppose you cannot list thirty closed initiatives with their original forecasts and their actual outcomes. Then you have no institutional memory of your own delivery, which is why someone else's percentage looks attractive. It fills a gap that should not exist.
Limits, and the counterargument that matters
Four objections have force, and the third is the one that should change how you read everything above.
One: incentive analysis is not falsification. A number produced by an interested party can still be correct, and pointing at a business model does not refute a finding. Gartner sells subscriptions and may still be right about 28%. NANDA sells a protocol thesis and may still be right that memory is the binding constraint. This essay has a producer too, with an interest in appearing more careful than the sources it examines. Incentive analysis tells you where to check. It does not tell you what you will find.
Two: what we have here are disclosure gaps, not proven errors. For Gartner, BCG, McKinsey, Wharton and Ramp we have no published sampling frame, no weighting scheme and no question wording. That is undisclosed methodology, and it is a reason to hold the numbers loosely. It is not evidence that any of them is wrong. I have also read several of these sources through secondary outlets, because the originals block automated access, and that is a limitation of this essay rather than of the sources.
Three: the direction converges, and the convergence is real. This is the objection that has to win half the argument.
Take the numbers pointing the other way first, and apply the same scepticism to them. Wharton's Human-AI Research initiative, working with the consultancy GBK Collective, surveyed more than 800 senior leaders at firms with over 1,000 employees and $50m in revenue in October 2025. It found 74% already seeing positive ROI from GenAI, 72% tracking formalised ROI metrics, and 82% of leaders using GenAI at least weekly. That is the same class of evidence as the 95%: executive self-assessment, a sample selected by an interested party, and no independent audit of return. Scepticism that only runs in the pessimistic direction is not scepticism. It is a mood.
Now hold the convergence. McKinsey's State of AI published a wave in November 2025, with 1,993 respondents across 105 countries. It reports 88% of organisations using AI in at least one function, but only 39% attributing any EBIT impact to it. Most of those put the impact below 5%, and roughly 6% qualify as high performers. The earlier March 2025 wave had more than 80% of respondents reporting no measurable enterprise-level EBIT impact. Those figures are as reported through secondary outlets, since McKinsey's page and PDF time out to automated access.
Then leave the commercial producers entirely. Daron Acemoglu, in "The Simple Macroeconomics of AI", Economic Policy 40(121), 2025, applies a task-based model and Hulten's theorem. The upper bound he derives is about +0.66% total factor productivity over ten years. An economist with nothing to sell, using a method with no survey in it, also arrives at a small number.
A consultant, an analyst house, a university lab, a survey of 1,993 executives and an independent macroeconomist measure five different things by five different methods and point the same way. Most AI initiatives do not deliver measurable return quickly. The dispersion of the values justifies rejecting the specific figures. It does not justify rejecting the conclusion. And rejecting every number is cheaper than building your own, which is why that move deserves suspicion.
Four: "pilot did not scale" is not the same as "pilot failed", and nobody has good data on the difference. Arnon Shimoni argued in August 2025 that MIT counts as a failure every pilot that did not reach full production. He offered his own breakdown, in which only a quarter of pilots are genuine failures. The rest are learning exercises, vendor evaluations, and discoveries that a problem was not worth solving. Those are private observations with no published methodology, from someone working in the industry, and they are not data. As an argument about construct validity they still stand. A portfolio of cheap, reversible experiments should have a high termination rate by design, and any metric that scores termination as failure will punish exactly the behaviour you want.
Closing
The useful question for your next board meeting is how many of your own closed initiatives you have ever counted, and against which forecast. Until that list exists, every percentage you quote is somebody else's organisation, measured on somebody else's terms, for somebody else's purpose.
Sources
- primaryBent Flyvbjerg, "From Nobel Prize to Project Management: Getting Risks Right," Project Management Journal 37(3), August 2006, pp. 5–15 (DOI 10.1177/875697280603700302); "Curbing Optimism Bias and Strategic Misrepresentation in Planning: Reference Class Forecasting in Practice," European Planning Studies 16(1), 2008, pp. 3–21 (DOI 10.1080/09654310701747936); and, with Alexander Budzier, "Why Your IT Project May Be Riskier Than You Think," Harvard Business Review 89(9), September 2011, pp. 23–25 — the three steps of reference class forecasting, the optimism-bias versus strategic-misrepresentation distinction, and the 1,471-project fat tail. The 16,000-project database figures (8.5% and 0.5%) come from the trade book How Big Things Get Done (2023) and are quoted as reported.
- primaryKjetil Moløkken-Østvold & Magne Jørgensen, "How large are software cost overruns? A review of the 1994 CHAOS report," Information and Software Technology 48(4), 2006, pp. 297–301 (the 189% figure as an outlier; a sampling frame skewed toward failing projects), and J. Laurenz Eveleens & Chris Verhoef, "The Rise and Fall of the Chaos Report Figures," IEEE Software 27(1), 2010, pp. 30–36 (5,457 forecasts across 1,211 projects; the four objections). Marilyn Strathern's 1997 formulation of Goodhart's law names the third objection.
- primaryMETR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (10 July 2025; arXiv:2507.09089) — 16 developers, 246 tasks; forecast −24%, post-hoc estimate −20%, measured +19%; design revised February 2026. With Erik Brynjolfsson, Danielle Li & Lindsey Raymond, "Generative AI at Work," Quarterly Journal of Economics 140(2), 2025, pp. 889–942 — 5,172 support agents; +15% average, +34% for novices, near zero for the experienced.
- primaryDaron Acemoglu, "The Simple Macroeconomics of AI," Economic Policy 40(121), 2025, pp. 13–58 (NBER WP 32487, May 2024) — task-based model plus Hulten's theorem; no more than +0.66% TFP over ten years.
- primaryStuart Soroka, Patrick Fournier & Lilach Nir, "Cross-national evidence of a negativity bias in psychophysiological reactions to news," PNAS 116(38), 2019, pp. 18888–18892 (17 countries, six continents), and Marc Trussler & Stuart Soroka, "Consumer Demand for Cynical and Negative News Frames," International Journal of Press/Politics 19(3), 2014, pp. 360–379 (revealed preference diverges from stated preference).
- primaryUS Census Bureau, "Large Firms With at Least 20 Employees Biggest AI Users" (BTOS, 26 May 2026) — 17–20% of firms; 37% at 250+ employees; 18% unweighted against 32% weighted by employment; question rewritten in November 2025.
- primaryThe sources behind the headline numbers. MIT Media Lab, "NANDA — Project Overview" (the agent registry, agent-to-agent protocols and the DNS-and-HTTP analogy) plus MIT NANDA, "The GenAI Divide: State of AI in Business 2025" (the 95%, not peer reviewed); RAND, Ryseff, De Bruhl & Newberry, "The Root Causes of Failure for AI Projects" (RRA2680-1, 2024) — qualitative interviews, with "more than 80%" appearing as an estimate cited from elsewhere; S&P Global Market Intelligence, "Voice of the Enterprise: AI & Machine Learning, Use Cases 2025" — 1,006 respondents, 46% of PoCs killed, 42% of firms abandoning most initiatives.
- primaryBCG, "AI Adoption in 2024: 74% of Companies Struggle to Achieve and Scale Value" (24 October 2024) — 1,000 executives, 59 countries; a self-assessment of capability maturity. Wharton Human-AI Research with GBK Collective, "Accountable Acceleration: Gen AI Fast-Tracks Into the Enterprise" (October 2025) — 800+ senior leaders; 74% reporting positive ROI, on the same class of evidence.
- secondaryGartner press-release titles and figures, retrieved via search index because the newsroom returns a 403 to automated access: "Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027" (25 June 2025, a forecast); "45% of Organizations With High AI Maturity Keep AI Projects Operational for at Least Three Years" (30 June 2025, a survey result); "Generative AI is Now the Most Frequently Deployed AI Solution" (7 May 2024, 48% reaching production over eight months); the 2018 forecast that 85% of AI projects would deliver erroneous outcomes through 2022. Gartner FY2024/FY2025 revenue and contract value via stockanalysis.com (SEC EDGAR returns 403). NetScout v. Gartner (2014) rejected the pay-to-play claim.
- secondaryFortune, "An MIT report finding 95% of AI pilots fail spooked investors…" (21 August 2025) — the 19–20 August sell-off, with Altman's bubble remarks in the same week. Futuriom, "Why We Don't Believe MIT NANDA's Weird AI Study" (August 2025) — the source for Kevin Werbach's data-or-retraction condition, quoted as reported since the page returns a 403. Arnon Shimoni, "MIT's 95% AI failure rate is wrong" (27 August 2025) — a practitioner's decomposition, an argument rather than data.
- secondaryMcKinsey QuantumBlack, "The State of AI" (November 2025 wave; 1,993 respondents, 105 countries) — 88% using AI, 39% attributing any EBIT impact, ~6% high performers; March 2025 wave, over 80% with no measurable EBIT impact. Figures as reported through secondary outlets, since the page and PDF time out. Ramp, "AI Index" and "AI Index March 2026" — card and bill-pay data across 70,000+ businesses; 47.6% in February 2026, above 50% in March, ~35% a year earlier; no published sampling frame.
- secondaryProvenance and mutation. Beverly D. Houghton's Woozle effect (1979), described by the author as reification by accretion, with current methodological use in Electronic Commerce Research (Springer, 2024). VentureBeat, "Why do 87% of data science projects never make it into production?" (July 2019) — sponsored content reporting a conference panel; provenance confirmed through secondary accounts as the page returns a 403. Exhibits, cited as artefacts and not as data: Folio3 AI, "AI Project Failure Rate in 2026" (the RAND sample of 2,400+ initiatives that does not exist, the S&P quantifier swap, the 2018 Gartner forecast redated to 2025) and Pertama Partners, "AI Project Failure Rate 2026: 80% Fail" (correct qualitative caveat in the body, alarmist headline above it).
95, 80, 28: wojna liczb o porażki AI
Trzy liczby prowadzą w tym roku większość rozmów o AI na zarządach. Raport MIT NANDA mówi, że 95% pilotaży GenAI nie ma mierzalnego wpływu na rachunek zysków i strat (P&L). Liczbę ponad 80% nieudanych projektów AI cytuje się z powołaniem na RAND, a Gartner podaje, że 28% use case'ów AI odnosi pełny sukces. Nie przeczą sobie, bo nigdy nie liczyły tego samego. Wybór przedmiotu liczenia zapadł, zanim zebrano jakiekolwiek dane — i właśnie o tym wyborze jest ten esej.
Statystyki porażek AI to gatunek wydawniczy z własną ekonomią. Metodologia każdej liczby jest pochodną bodźców tego, kto ją wyprodukował. Czytaj to jako analizę bodźców, nie jako oskarżenie: nikt nie musiał kłamać, żeby liczby wylądowały tam, gdzie wylądowały.
Jedną z nich rozbieram osobno, po kawałku: co naprawdę mówi 28% Gartnera. Ten esej jest o gatunku.
Kto produkuje te liczby i co przy okazji sprzedaje
Zacznij od bilansu. Jest publiczny i tłumaczy kształt produktu.
Gartner zamknął rok obrotowy 2025 przychodem 6497 mln USD wobec 6267 mln w 2024 — wzrost o około 3,7%. Największy udział ma segment Insights, dawniej Research. Około 83% przychodu pochodzi z odnawialnych subskrypcji badawczych, przy wartości kontraktów rzędu 5,0–5,2 mld USD w 2025 roku; doradztwo dokłada mniej więcej 9%. Te dane biorę z agregatora raportu 10-K, bo SEC EDGAR zwraca 403 przy dostępie automatycznym — traktuj je jako podane. Model biznesowy to subskrypcja badań pod decyzje, plus konferencje, plus doradztwo.
Cytowalna liczba w komunikacie prasowym to darmowa próbka; właściwy produkt jest dopiero za nią. To opis lejka sprzedażowego, nie zarzut kupowania wyników — ten przerabiał sąd w sprawie NetScout przeciw Gartnerowi i odrzucił go w 2014 roku. Nikt nie kupuje rezultatu. Rzecz jest subtelniejsza i trudniej się przed nią bronić: liczba, która do ciebie dociera, przeszła już selekcję pod kątem cytowalności.
Selekcję widać w dorobku jednego producenta. W ciągu trzynastu miesięcy Gartner opublikował trzy liczby. W czerwcu 2025 prognozował, że do końca 2027 roku skasowanych zostanie ponad 40% projektów agentowych. Kilka dni później podał wynik ankiety: 45% organizacji o wysokiej dojrzałości AI utrzymuje projekty AI w ruchu przez co najmniej trzy lata. W kwietniu 2026 przyszło 28%. Pierwsza i trzecia obiegły wszystko. Czterdzieści pięć procent nie ma w mediach niemal żadnego życia.
Tytuły i liczby mam z indeksu wyszukiwarki, bo newsroom Gartnera też zwraca 403 — nie cytuję samych komunikatów. Wzorzec i tak się trzyma. Pesymistycznych liczb nie pchano mocniej niż optymistycznej. Mocniej je podchwycono. Selekcja dzieje się także po stronie odbiorcy.
Laboratorium uniwersyteckie również ma produkt. Liczba 95% wyszła z projektu MIT NANDA — Networked Agents and Decentralized AI, prowadzonego w MIT Media Lab, z Rameshem Raskarem jako głównym badaczem. NANDA buduje infrastrukturę dla agentów AI: zdecentralizowany rejestr opisywany jako DNS dla agentów, protokoły agent–agent, weryfikację tożsamości i reputacji. Deklaracja misji stawia analogię wprost i porównuje brakującą warstwę agentową do DNS i HTTP z wczesnej sieci.
Teraz przeczytaj obok tego diagnozę raportu. Wdrożenia korporacyjne zawodzą, twierdzi raport, bo systemom brakuje pamięci i nie uczą się z informacji zwrotnej. To ta sama luka, którą infrastruktura NANDA ma wypełnić. Raport nie przeszedł recenzji naukowej.
Wersja konsultanta jest w tym gatunku najstarsza. Badanie BCG z października 2024, Where's the Value in AI?, ustaliło, że 74% firm ma kłopot z osiągnięciem i wyskalowaniem wartości z AI. Tylko 26% miało zdolności, by wyjść poza proof of concept (PoC), a 4% opisano jako czołówkę. Próba: 1000 członków zarządów i wyższej kadry z 59 krajów i ponad 20 branż. Przyjrzyj się narzędziu. To samoocena dojrzałości zdolności; badanie nigdy nie twierdziło, że mierzy wyniki projektów. Mimo to krąży jako 74% porażek, a BCG sprzedaje transformację AI.
Ostatni dowód rzeczowy pokazuje, co się dzieje, gdy taka liczba przestaje być tylko tematem rozmów. 19 i 20 sierpnia 2025 akcje spółek związanych z AI poszły w dół: Nasdaq spadł o ponad 1,2%, Nvidia o około 3,5%, Palantir o blisko 10%. Ówczesne komentarze rynkowe wskazywały raport MIT obok wypowiedzi Sama Altmana o bańce jako bezpośrednie zapalniki. Oba wydarzenia przypadły na ten sam tydzień, więc wyprzedaży nie da się przypisać samemu raportowi. Mimo to nierecenzowany working paper z laboratorium budującego infrastrukturę agentową poruszył rynki publiczne, a metodologia nie pojechała razem z liczbą z nagłówka.
Liczba dojechała. Mianownik został w domu.
Pięć ustawień, które przesądzają odpowiedź
Zanim ktokolwiek skontaktuje się z pierwszym respondentem, badanie porażek wybrało już pięć rzeczy. Każdy z tych wyborów przesuwa wynik o dziesiątki punktów i żadnego nie widać w nagłówku.
Jednostka analizy. Use case, pilotaż, PoC, projekt, firma albo prognoza. W obiegu jest sześć różnych mianowników i nie da się ich na siebie przeliczyć.
Próg sukcesu. Mierzalny wpływ na P&L, dojście do produkcji, utrzymanie się w produkcji przez trzy lata, spełnienie oczekiwań ROI menedżera albo trafienie w pierwotny szacunek. Każdy próg ma własną literaturę i własną liczbę.
Populacja. Liderzy infrastruktury i operacji (I&O), członkowie zarządów, przepytani praktycy, baza klientów jednego fintechu albo reprezentatywna próba rządowa.
Horyzont. Projekt młodszy niż okno ankiety liczy się jako porażka z samej konstrukcji pomiaru. Osiemnastomiesięczny okres zwrotu w dziewiątym miesiącu nie pokazuje niczego.
Samoocena kontra pomiar. Ocena własnego ROI przez menedżera to nie audyt; jedno i drugie potrafi wskazywać w przeciwne strony. Wracam do tego niżej.
Ustaw te pięć pokręteł, a odpowiedź jest w dużej mierze przesądzona. Zbieranie danych, które potem następuje, jest kosztowne i staranne — i wobec nagłówka głównie ozdobne.
| Liczba | Kto i próba | Jednostka liczenia | Próg | Czego strukturalnie nie widzi |
|---|---|---|---|---|
| 28% pełnego sukcesu | Gartner, 782 liderów I&O, XI–XII 2025 | use case AI wewnątrz I&O | pełny sukces plus oczekiwania ROI, samoocena | biznes, finanse, użytkownicy końcowi, AI po stronie produktu. Ta sama ankieta podaje 77% liderów z co najmniej jednym udanym use case'em |
| 95% pilotaży | MIT NANDA, 2025; bez recenzji naukowej | pilotaż lub wdrożenie GenAI | mierzalny wpływ na P&L | shadow AI, wartość spoza P&L, wszystko poniżej progu pomiaru |
| ponad 80% porażek | RAND, szacunek cytowany wewnątrz badania jakościowego | projekt AI | „zawodzi" — szacunek zacytowany, nie zmierzony przez RAND | sama precyzja; własny odczyt RAND brzmi „znaczna większość" |
| 46% ubitych PoC | S&P Global, 1006 specjalistów IT i biznesu | proof of concept | porzucenie przed produkcją, nie ROI | czy porzucenie było decyzją biznesową, czy techniczną. To nie to samo co 42% firm porzucających większość inicjatyw, wobec 17% wcześniej |
| co najmniej 30% porzuconych | Gartner, lipiec 2024 | projekt GenAI | prognoza do końca 2025 | to nie pomiar czegokolwiek, co się wydarzyło |
| 48% dochodzi do produkcji | Gartner, maj 2024, jak podano | projekt AI | przepustowość wdrożeń; osiem miesięcy od prototypu do produkcji | ROI, utrzymanie, porzucenie — znowu szósty konstrukt |
Każdy wiersz ma własny mianownik i własny próg. Sklejanie ich w jeden „wskaźnik porażek AI" to błąd kategorii, a szczegółową rozbiórkę pierwszego wiersza znajdziesz w eseju towarzyszącym.
Gatunek jest starszy niż AI
Nic z tego nie zaczęło się w 2023 roku. Szablon — alarmujący odsetek projektów technologicznych zawodzi — chodzi od połowy lat 90. i był już dwa razy rozebrany przez ludzi z danymi.
Raport CHAOS grupy Standish definiuje sukces jako projekt w terminie, w budżecie i w zakresie, mierzony wobec pierwotnego szacunku. Co przekroczy, staje się „zagrożone"; co anulowane — „nieudane".
Liczą się dwie niezależne krytyki, a większość cytowań CHAOS nie wspomina o żadnej.
Moløkken-Østvold i Jørgensen w Information and Software Technology 48(4) z 2006 roku wzięli pod lupę sztandarową liczbę z 1994: średnie przekroczenie kosztu o 189%. Wyszła im dużo wyższa niż w porównywalnych badaniach szacowania z tego samego okresu, a różnicę wyprowadzili z doboru próby przechylonego ku projektom nieudanym. Ich wniosek: 189% jest bardzo prawdopodobnie stanowczo za wysokie dla typowych projektów lat 90. Dalsze używanie tej liczby jako punktu odniesienia, argumentowali, produkuje złe decyzje i hamuje praktykę szacowania.
Eveleens i Verhoef w IEEE Software 27(1) z 2010 roku pracowali na 5457 prognozach z 1211 projektów i podnieśli cztery zarzuty. Sukces zdefiniowany wyłącznie przez trafność szacunku wprowadza w błąd. Miara jest jednostronna, więc sukcesy są systematycznie zaniżane. Zarządzanie pod miarę psuje praktykę szacowania, bo najtańsza droga do „sukcesu" prowadzi przez napompowanie pierwotnej prognozy. A uśrednianie danych o nieznanym obciążeniu daje liczby bez interpretacji.
Trzeci zarzut ma nazwę. Prawo Goodharta w sformułowaniu Marilyn Strathern z 1997 roku mówi: „gdy miara staje się celem, przestaje być dobrą miarą". Przyłóż to do dziś. Jeśli twój zarząd zacznie raportować odsetek udanych use case'ów AI, zespoły zaczną zgłaszać tylko te bezpieczne. Liczba się poprawi, a poza nią nic.
Ten sam nóż tnie 95, 80 i 28. Naostrzono go na raporcie z 1994 roku.
Jak liczba mutuje w drodze
Gatunek ma mechanizm transmisji i najłatwiej zobaczyć go na jednej stronie.
Folio3 AI sprzedaje budowę AI na zamówienie i doradztwo, a przy okazji prowadzi stronę ze statystykami porażek projektów AI. Trzy wpisy warto obejrzeć z bliska — wyłącznie jako okazy.
Pierwszy: „80% projektów AI nie dowozi zamierzonej wartości biznesowej", przypisane analizie RAND obejmującej ponad 2400 korporacyjnych inicjatyw AI. RAND nie analizował niczego takiego. Raport RAND to badanie jakościowe oparte na wywiadach z praktykami, a jego 80% występuje jako szacunek zacytowany skądinąd, obwarowany frazą „według niektórych szacunków". Liczba dorobiła się próby, której nigdy nie miała.
Drugi: „42% firm porzuciło w 2025 roku co najmniej jedną inicjatywę AI", przypisane S&P. W danych S&P 42% to odsetek firm porzucających większość swoich inicjatyw, wobec 17% rok wcześniej. Podmiana kwantyfikatora zmienia konstrukt w całości i sprawia, że liczba brzmi zarazem łagodniej i powszechniej.
Trzeci: „85% porażek projektów AI ma źródło w złej jakości danych (Gartner, 2025)". Źródłem jest prognoza Gartnera z 2018 roku, według której do 2022 roku 85% projektów AI miało dawać błędne wyniki z powodu obciążeń w danych, w algorytmach albo w zespołach, które nimi zarządzają. Prognoza o złych odpowiedziach stała się zmierzoną przyczyną porażki, a z 2018 zrobił się 2025. Trzy mutacje na jednej stronie, wszystkie w tę samą stronę.
Kanoniczna liczba gatunku nigdy nie miała za sobą żadnego badania. „87% projektów data science nigdy nie trafia na produkcję" pochodzi z artykułu VentureBeat z lipca 2019 roku, który był treścią sponsorowaną relacjonującą panel na konferencji tego samego serwisu. Nie było próby ani pracy źródłowej. Towarzysząca jej liczba 78% wzięła się z ankiety Dimensional Research zamówionej przez Alegion, dostawcę etykietowania danych. Oryginalny artykuł zwraca dziś 403 przy dostępie automatycznym, więc powyższe pochodzenie podaję za relacjami wtórnymi. Dziesiątki dostawców MLOps wciąż cytują tę liczbę jako podstawę dowodową zakupu swoich narzędzi.
Mechanizm ma nazwę od 1979 roku. Beverly Houghton ukuła termin „efekt Woozle" na dowód przez cytowanie: teza zyskuje wiarygodność przez częstość powtarzania, a nie przez jakikolwiek dowód na jej rzecz. Sama opisała to jako urzeczowienie przez narastanie. Termin wciąż jest w użyciu metodologicznym — praca z 2024 roku w Electronic Commerce Research stosuje go wprost do trwałości przekonań w literaturze naukowej.
Powtórzenie to nie replikacja. Brzmi tylko podobnie.
Pouczający jest jeden kontrprzykład z wnętrza tego samego gatunku komercyjnego. Pertama Partners, firma szkoleniowo-doradcza od AI, prowadzi stronę z nagłówkiem o 80% porażek w 2026 roku. W treści pisze potem, że badanie RAND jest jakościowe i oparte na wywiadach. Najlepiej czytać je jako „zawodzi znaczna większość", a nie jako precyzyjny procent — tak stwierdza sama strona i linkuje raport RAND. Drobny druk jest poprawny, nagłówek alarmistyczny, a strona sprzedaje audyty gotowości i warsztaty. Bodziec siedzi w tytule i w wezwaniu do działania, nie w analizie.
Dlaczego podróżuje akurat liczba skrajna
Kusi, żeby wytłumaczyć tę selekcję złą wolą redakcji. Dowody wskazują na coś nudniejszego i trudniejszego do naprawienia.
Stuart Soroka z Patrickiem Fournierem i Lilach Nir przeprowadzili eksperyment psychofizjologiczny w 17 krajach na sześciu kontynentach, na prawdziwym materiale informacyjnym; wyniki opublikowali w PNAS 116(38) z 2019 roku. Przeciętny człowiek reaguje silniejszym pobudzeniem na treść negatywną niż na pozytywną. To najszersze dostępne międzynarodowe potwierdzenie skrzywienia ku negatywności, a autorzy uczciwie odnotowują dużą zmienność indywidualną.
Trussler i Soroka w International Journal of Press/Politics 19(3) z 2014 roku pokazali połowę behawioralną. W warunkach kontrolowanych czytelnicy wybierają teksty negatywne i cyniczne, niezależnie od tego, co deklarują. Preferencja deklarowana rozjeżdża się z ujawnioną.
Żaden redaktor nie musi spiskować. Serwis optymalizujący kliknięcia i konferencja optymalizująca frekwencję dojdą do najskrajniejszej dostępnej liczby zwykłą drogą. Optymistyczne 45% przegrywa z pesymistyczną prognozą 40% z tego samego powodu, dla którego spokojny kwartał przegrywa ze skandalem.
Jest jedna kontrola, która rozstrzygnęłaby cokolwiek z tego — i gatunek prawie nigdy jej nie używa. Raport MIT NANDA nie przeszedł recenzji naukowej. Kevin Werbach z Wharton, po wielokrotnej lekturze, podał, że nie umie odtworzyć, skąd wzięło się 95%, i postawił jasny warunek: opublikować pełne dane źródłowe albo wycofać raport. Raport wspomina o 52 wywiadach i setkach punktów danych, ale nie ujawnia struktury próby ani metody zbierania. Na dzień zbierania materiałów do tego eseju nie nastąpiło ani udostępnienie danych, ani wycofanie, i nie mam dowodu na żadne z nich. Sformułowanie Werbacha podaję za relacją, bo źródłowy serwis też zwraca 403. Brak recenzji nie dowodzi niczego o samej liczbie, w żadną stronę. Usuwa tylko mechanizm, który wyłapałby błąd.
Co daje w zamian pomiar
Zamień ankietę na pomiar, a liczby się nie zbiegną. Rozpadną się inaczej, a te odłamki są bardziej użyteczne.
METR przeprowadził badanie z randomizacją, opublikowane w lipcu 2025: 16 doświadczonych programistów open source, 246 realnych zadań w ich własnych dojrzałych repozytoriach, narzędzia z początku 2025 roku. Programiści prognozowali, że AI przyspieszy ich o 24%. Po skończeniu szacowali, że byli szybsi o 20%. Pomiar pokazał, że byli o 19% wolniejsi. Samoocena i pomiar wskazały w przeciwne strony, na tych samych zadaniach, u tych samych ludzi.
To najmocniejszy dostępny argument przeciw czytaniu jakiejkolwiek ankiety samooceny jako pomiaru zwrotu. To zarazem małe badanie, w wąskim kontekście, na narzędziach z pierwszej połowy 2025 roku, a sam METR zmienił projekt eksperymentu w lutym 2026. Nie pokazuje, że AI spowalnia programistów. Pokazuje, że pytanie ludzi i mierzenie ich dają różne odpowiedzi.
Pomiar bywa też mocno pozytywny — w tym samym okresie, przy innej pracy. Erik Brynjolfsson z Danielle Li i Lindsey Raymond w Quarterly Journal of Economics 140(2) z 2025 roku wykorzystali stopniowe wdrożenie asystenta GenAI u 5172 konsultantów obsługi klienta. Produktywność mierzona liczbą zamkniętych zgłoszeń na godzinę wzrosła średnio o 15% — o 34% u początkujących i o niemal nic u doświadczonych. Poprawiły się nastroje klientów, wzrosło też utrzymanie pracowników.
Postawione obok siebie oba badania wyglądają na sprzeczne. Zgadzają się w tym, co ważne. U doświadczonych praktyków zmierzony efekt jest bliski zeru albo ujemny. U początkujących jest duży. Pomiar nie daje żadnego wskaźnika porażek — daje efekty, które zależą od zadania, od narzędzi i od stażu.
Nawet adopcja, która brzmi jak najprostsza rzecz do policzenia, rozjeżdża się o mniej więcej trzydzieści punktów zależnie od tego, kto liczy. Business Trends and Outlook Survey amerykańskiego Census Bureau to reprezentatywne narzędzie rządowe. Daje użycie AI na poziomie 17–20% firm między połową grudnia 2025 a początkiem maja 2026. U firm z 250 pracownikami i więcej rośnie do 37%, a poniżej 20 pracowników spada pod 20%. Jeden z aneksów podaje 18% firm, ale 32% po zważeniu zatrudnieniem. Samo pytanie zmieniono w listopadzie 2025 — z AI używanej przy wytwarzaniu towarów i usług na AI używaną w dowolnej funkcji biznesowej — co przesunęło wynik niezależnie od jakiegokolwiek zachowania.
AI Index firmy Ramp mierzy coś zupełnie innego: dane transakcyjne z kart firmowych i płatności faktur w ponad 70 000 przedsiębiorstw. W lutym 2026 zanotował 47,6% firm płacących za narzędzia AI, w marcu przekroczył 50%, wobec około 35% rok wcześniej. To pomiar płatnego zobowiązania, nie użycia, na próbie klientów jednego fintechu, przechylonej ku firmom technologicznym. Ramp nie publikuje na stronie indeksu ani operatu losowania, ani zastrzeżenia o reprezentatywności.
Przepaść między 17–20% a 50% nie jest sporem o fakty. Dwie populacje i dwie definicje adopcji dają dwie poprawne odpowiedzi na dwa różne pytania. Mechanizm jest identyczny jak przy 95 kontra 80 kontra 28.
Zastosowanie: zbuduj własną klasę referencyjną
Ruch, który się opłaca, to przestać importować czyjąkolwiek liczbę jako prognozę dla własnej organizacji.
Metoda ma autora i datę. Bent Flyvbjerg wyłożył prognozowanie z klasy referencyjnej w tekście „From Nobel Prize to Project Management: Getting Risks Right", Project Management Journal 37(3), sierpień 2006. Praktykę rozwinął w „Curbing Optimism Bias and Strategic Misrepresentation in Planning", European Planning Studies 16(1), 2008. Trzy kanoniczne kroki:
- Wskaż klasę referencyjną minionych, podobnych przedsięwzięć.
- Ustal dla tej klasy rozkład prawdopodobieństwa na parametrze, który prognozujesz.
- Porównaj swoje konkretne przedsięwzięcie z tym rozkładem i wyznacz najbardziej prawdopodobny wynik.
Teorią nadrzędną jest rozróżnienie Kahnemana i Tversky'ego na widok od wewnątrz i widok z zewnątrz, rozwinięte z Lovallo w 1993 roku. Widok od wewnątrz opiera się na szczegółach twojego projektu. Widok z zewnątrz opiera się na rozkładzie wyników podobnych projektów i celowo ignoruje twoje szczegóły. Wkład Flyvbjerga polega na tym, że koryguje to naraz dwa różne błędy: nieświadome skrzywienie optymizmu oraz strategiczne zafałszowanie, czyli celowe zaniżanie kosztu, żeby uzyskać zgodę na projekt.
To rozróżnienie przenosi się czysto na producentów liczb o porażkach. Część skrzywienia jest poznawcza. Część to wybór redakcyjny z biznesem w tle. Lekarstwa są inne.
W praktyce metodę zastosowano najpierw przy drugiej linii tramwajowej w Edynburgu, a później przy Crossrail dla brytyjskiego Ministerstwa Transportu. Przewaga metody jest największa przy przedsięwzięciach nierutynowych — a takim jest korporacyjne AI dla organizacji, która robi to pierwszy raz.
Póki budujesz własną wartość bazową, trzymaj w ręku dwie cudze.
Baza Flyvbjerga, budowana przez dekady w Saïd Business School w Oksfordzie, obejmuje ponad 16 000 dużych projektów. Jak podano w How Big Things Get Done (2023), 8,5% mieści się w budżecie i w terminie, a 0,5% mieści się w budżecie, w terminie i dowozi obiecane korzyści. Ta liczba pochodzi z książki popularnej, nie z pracy recenzowanej, więc traktuj ją jako podaną. Wniosek jest niewygodny dla całego gatunku. „Większość inicjatyw nie dowozi tego, co obiecano" to wartość bazowa dla projektów w ogóle, nie osobliwość AI.
Flyvbjerg i Budzier w Harvard Business Review 89(9) z września 2011 zbadali 1471 projektów IT. Średnie przekroczenie kosztu: 27%. Ale co szósty projekt był czarnym łabędziem, ze średnim przekroczeniem kosztu o 200% i harmonogramu o blisko 70%. Średnia nie opisuje prawie nikogo. To ogon rozstrzyga, czy firma przeżyje projekt.
Cała wojna o 95 kontra 80 kontra 28 toczy się o parametr, o który nie należało pytać.
Oto procedura, którą położyłbym przed zarządem, po kolei.
- Zdefiniuj jednostkę, zanim spojrzysz na jakąkolwiek liczbę. Use case, pilotaż, projekt albo firma. Bez tego nic, co czytasz, nie jest porównywalne z niczym, co prowadzisz.
- Ustal próg sukcesu z góry i na piśmie. Mierzalny wpływ na P&L? Dojście do produkcji? Działanie po trzech latach? Każdy próg ma w literaturze doczepioną inną liczbę, a wybieranie progu po fakcie to prosta droga do oceniania projektów z taryfą ulgową.
- Zbierz własną klasę referencyjną. Od dziesięciu do trzydziestu zamkniętych inicjatyw z ostatnich trzech–pięciu lat we własnej organizacji — nie tylko AI, ale każdy projekt zmiany o porównywalnej złożoności i nowości.
- Zbuduj rozkład, nie średnią. Porównaj koszt i czas z pierwotną prognozą; to samo zrób z korzyścią. Zapisz medianę oraz P80 albo P90.
- Umieść nowy projekt w tym rozkładzie i zapisz wynikający z tego narzut na koszt i na czas. To trzeci krok prognozowania z klasy referencyjnej i to jego ludzie pomijają.
- Oddziel skrzywienie optymizmu od strategicznego zafałszowania. Pierwsze leczy się danymi. Drugie leczy się zmianą bodźców: sprawdź, czy osoba awansowana za zdobycie zgody na projekt odpowiada potem za jego dowiezienie.
- Wyceń ogon przed średnią. Jeśli co szósty projekt IT przekracza budżet o 200%, pierwsze pytanie brzmi, co zrobisz, jeśli to jest właśnie ten, a nie jak wygląda średnia.
- Używaj cudzych liczb wyłącznie jako zewnętrznej klasy referencyjnej, z mianownikiem powiedzianym na głos, i nigdy jako prognozy dla siebie.
Kroki trzeci i czwarty to całe ćwiczenie. Reszta to księgowość.
Większość organizacji nie umie wykonać kroku trzeciego i to jest prawdziwe odkrycie. Załóżmy, że nie potrafisz wymienić trzydziestu zamkniętych inicjatyw z ich pierwotnymi prognozami i faktycznymi wynikami. Wtedy nie masz instytucjonalnej pamięci własnego dowożenia — i dlatego cudzy procent wygląda atrakcyjnie. Zapełnia lukę, której nie powinno być.
Granice i kontrargument, który ma wagę
Cztery zarzuty mają siłę, a trzeci powinien zmienić sposób, w jaki czytasz wszystko powyżej.
Pierwszy: analiza bodźców to nie falsyfikacja. Liczba wyprodukowana przez stronę zainteresowaną wciąż może być poprawna, a wskazanie modelu biznesowego nie obala wyniku. Gartner sprzedaje subskrypcje i może mieć rację co do 28%. NANDA sprzedaje tezę o protokole i może mieć rację, że pamięć jest wiążącym ograniczeniem. Ten esej też ma producenta, z interesem w tym, żeby uchodzić za staranniejszy od źródeł, które bada. Analiza bodźców mówi ci, gdzie sprawdzić. Nie mówi, co znajdziesz.
Drugi: mamy tu luki w ujawnianiu, a nie udowodnione błędy. Przy Gartnerze, BCG, McKinseyu, Wharton i Rampie nie ma opublikowanego operatu losowania, schematu ważenia ani brzmienia pytań. To metodologia nieujawniona i powód, żeby traktować te liczby z rezerwą. To nie dowód, że któraś z nich jest błędna. Część tych źródeł czytałem przez serwisy wtórne, bo oryginały blokują dostęp automatyczny — a to ograniczenie tego eseju, nie źródeł.
Trzeci: wyniki zbiegają się co do kierunku, a ta zbieżność jest prawdziwa. To zarzut, który musi wygrać połowę sporu.
Weź najpierw liczby wskazujące w drugą stronę i przyłóż do nich ten sam sceptycyzm. Inicjatywa Human-AI Research w Wharton, wspólnie z firmą doradczą GBK Collective, przepytała w październiku 2025 ponad 800 osób z wyższej kadry w firmach powyżej 1000 pracowników i 50 mln USD przychodu. Wyszło 74% widzących już dodatnie ROI z GenAI i 72% mierzących sformalizowane wskaźniki ROI; 82% liderów sięga po GenAI co najmniej raz w tygodniu. To ta sama klasa dowodu co 95%: samoocena kadry, próba dobrana przez stronę zainteresowaną, brak niezależnego audytu zwrotu. Sceptycyzm działający wyłącznie w stronę pesymistyczną nie jest sceptycyzmem. Jest nastrojem.
Teraz spójrz na samą zbieżność. McKinsey opublikował falę badania State of AI w listopadzie 2025, na 1993 respondentach ze 105 krajów. Podaje 88% organizacji używających AI w co najmniej jednej funkcji, ale tylko 39% przypisujących temu jakikolwiek wpływ na EBIT. Większość z nich wskazuje wpływ poniżej 5%, a mniej więcej 6% kwalifikuje się jako najlepsi. We wcześniejszej fali z marca 2025 ponad 80% respondentów zgłaszało brak mierzalnego wpływu na EBIT na poziomie firmy. Te liczby podaję za serwisami wtórnymi, bo strona i PDF McKinseya nie odpowiadają na dostęp automatyczny.
Potem wyjdź całkiem poza producentów komercyjnych. Daron Acemoglu w tekście „The Simple Macroeconomics of AI", Economic Policy 40(121), 2025, stosuje model zadaniowy i twierdzenie Hultena. Górne ograniczenie, które wyprowadza, to około +0,66% łącznej produktywności czynników produkcji przez dziesięć lat. Ekonomista, który nie ma nic do sprzedania, metodą bez ankiety w środku, też dochodzi do małej liczby.
Konsultant, dom analityczny, laboratorium uniwersyteckie, ankieta wśród 1993 menedżerów i niezależny makroekonomista mierzą pięć różnych rzeczy pięcioma różnymi metodami — i wskazują w tę samą stronę. Większość inicjatyw AI nie dowozi mierzalnego zwrotu szybko. Rozrzut wartości uzasadnia odrzucenie konkretnych liczb. Nie uzasadnia odrzucenia wniosku. A odrzucenie każdej liczby kosztuje mniej niż zbudowanie własnej — i dlatego ten ruch zasługuje na podejrzliwość.
Czwarty: „pilotaż się nie wyskalował" znaczy co innego niż „pilotaż zawiódł", a nikt nie ma dobrych danych o tej różnicy. Arnon Shimoni argumentował w sierpniu 2025, że MIT liczy jako porażkę każdy pilotaż, który nie doszedł do pełnej produkcji. Zaproponował własny podział, w którym prawdziwe porażki to tylko jedna czwarta pilotaży. Reszta to ćwiczenia z uczenia się, oceny dostawców i odkrycia, że problemu nie warto było rozwiązywać. To prywatne obserwacje bez opublikowanej metodologii, od człowieka pracującego w branży, i nie są danymi. Jako argument o trafności konstruktu wciąż się bronią. Portfel tanich, odwracalnych eksperymentów powinien mieć z założenia wysoki wskaźnik ubijania, a każda metryka licząca ubicie jako porażkę ukarze dokładnie to zachowanie, na którym ci zależy.
Zamknięcie
Użyteczne pytanie na najbliższy zarząd brzmi: ile własnych zamkniętych inicjatyw kiedykolwiek policzyliście i wobec której prognozy. Póki ta lista nie istnieje, każdy cytowany przez ciebie procent opisuje cudzą organizację, zmierzoną na cudzych warunkach, w cudzym celu.
Bibliografia
- primaryBent Flyvbjerg, „From Nobel Prize to Project Management: Getting Risks Right", Project Management Journal 37(3), sierpień 2006, s. 5–15 (DOI 10.1177/875697280603700302); „Curbing Optimism Bias and Strategic Misrepresentation in Planning: Reference Class Forecasting in Practice", European Planning Studies 16(1), 2008, s. 3–21 (DOI 10.1080/09654310701747936); oraz, z Alexandrem Budzierem, „Why Your IT Project May Be Riskier Than You Think", Harvard Business Review 89(9), wrzesień 2011, s. 23–25 — trzy kroki prognozowania z klasy referencyjnej, rozróżnienie skrzywienia optymizmu od strategicznego zafałszowania oraz gruby ogon w 1471 projektach. Liczby z bazy 16 000 projektów (8,5% i 0,5%) pochodzą z książki popularnej How Big Things Get Done (2023) i podane są za nią.
- primaryKjetil Moløkken-Østvold i Magne Jørgensen, „How large are software cost overruns? A review of the 1994 CHAOS report", Information and Software Technology 48(4), 2006, s. 297–301 (189% jako wartość odstająca; dobór próby przechylony ku projektom nieudanym), oraz J. Laurenz Eveleens i Chris Verhoef, „The Rise and Fall of the Chaos Report Figures", IEEE Software 27(1), 2010, s. 30–36 (5457 prognoz z 1211 projektów; cztery zarzuty). Sformułowanie prawa Goodharta przez Marilyn Strathern z 1997 roku nazywa zarzut trzeci.
- primaryMETR, „Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (10 lipca 2025; arXiv:2507.09089) — 16 programistów, 246 zadań; prognoza −24%, szacunek po fakcie −20%, pomiar +19%; projekt badania zmieniony w lutym 2026. Oraz Erik Brynjolfsson, Danielle Li i Lindsey Raymond, „Generative AI at Work", Quarterly Journal of Economics 140(2), 2025, s. 889–942 — 5172 konsultantów obsługi klienta; +15% średnio, +34% u początkujących, blisko zera u doświadczonych.
- primaryDaron Acemoglu, „The Simple Macroeconomics of AI", Economic Policy 40(121), 2025, s. 13–58 (NBER WP 32487, maj 2024) — model zadaniowy plus twierdzenie Hultena; nie więcej niż +0,66% TFP przez dziesięć lat.
- primaryStuart Soroka, Patrick Fournier i Lilach Nir, „Cross-national evidence of a negativity bias in psychophysiological reactions to news", PNAS 116(38), 2019, s. 18888–18892 (17 krajów, sześć kontynentów), oraz Marc Trussler i Stuart Soroka, „Consumer Demand for Cynical and Negative News Frames", International Journal of Press/Politics 19(3), 2014, s. 360–379 (preferencja ujawniona rozjeżdża się z deklarowaną).
- primaryUS Census Bureau, „Large Firms With at Least 20 Employees Biggest AI Users" (BTOS, 26 maja 2026) — 17–20% firm; 37% przy 250 pracownikach i więcej; 18% bez ważenia wobec 32% po zważeniu zatrudnieniem; pytanie przeredagowane w listopadzie 2025.
- primaryŹródła liczb z nagłówków. MIT Media Lab, „NANDA — Project Overview" (rejestr agentów, protokoły agent–agent i analogia do DNS oraz HTTP) plus MIT NANDA, „The GenAI Divide: State of AI in Business 2025" (95%, bez recenzji naukowej); RAND, Ryseff, De Bruhl i Newberry, „The Root Causes of Failure for AI Projects" (RRA2680-1, 2024) — wywiady jakościowe, w których „ponad 80%" występuje jako szacunek zacytowany skądinąd; S&P Global Market Intelligence, „Voice of the Enterprise: AI & Machine Learning, Use Cases 2025" — 1006 respondentów, 46% ubitych PoC, 42% firm porzucających większość inicjatyw.
- primaryBCG, „AI Adoption in 2024: 74% of Companies Struggle to Achieve and Scale Value" (24 października 2024) — 1000 menedżerów, 59 krajów; samoocena dojrzałości zdolności. Wharton Human-AI Research z GBK Collective, „Accountable Acceleration: Gen AI Fast-Tracks Into the Enterprise" (październik 2025) — ponad 800 osób z wyższej kadry; 74% zgłaszających dodatnie ROI, na tej samej klasie dowodu.
- secondaryTytuły i liczby z komunikatów prasowych Gartnera, pobrane przez indeks wyszukiwarki, bo newsroom zwraca 403 przy dostępie automatycznym: „Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027" (25 czerwca 2025, prognoza); „45% of Organizations With High AI Maturity Keep AI Projects Operational for at Least Three Years" (30 czerwca 2025, wynik ankiety); „Generative AI is Now the Most Frequently Deployed AI Solution" (7 maja 2024, 48% dochodzących do produkcji w ciągu ośmiu miesięcy); prognoza z 2018 roku, że do 2022 roku 85% projektów AI da błędne wyniki. Przychód i wartość kontraktów Gartnera za FY2024/FY2025 przez stockanalysis.com (SEC EDGAR zwraca 403). Sprawa NetScout przeciw Gartnerowi (2014) odrzuciła zarzut kupowania wyników.
- secondaryFortune, „An MIT report finding 95% of AI pilots fail spooked investors…" (21 sierpnia 2025) — wyprzedaż z 19–20 sierpnia, z wypowiedziami Altmana o bańce w tym samym tygodniu. Futuriom, „Why We Don't Believe MIT NANDA's Weird AI Study" (sierpień 2025) — źródło warunku Kevina Werbacha „dane albo wycofanie", podane za relacją, bo strona zwraca 403. Arnon Shimoni, „MIT's 95% AI failure rate is wrong" (27 sierpnia 2025) — rozbiór praktyka, argument, a nie dane.
- secondaryMcKinsey QuantumBlack, „The State of AI" (fala z listopada 2025; 1993 respondentów, 105 krajów) — 88% używających AI, 39% przypisujących jakikolwiek wpływ na EBIT, ~6% najlepszych; fala z marca 2025, ponad 80% bez mierzalnego wpływu na EBIT. Liczby podane za serwisami wtórnymi, bo strona i PDF nie odpowiadają. Ramp, „AI Index" oraz „AI Index March 2026" — dane z kart i płatności faktur w ponad 70 000 firm; 47,6% w lutym 2026, powyżej 50% w marcu, ~35% rok wcześniej; brak opublikowanego operatu losowania.
- secondaryPochodzenie i mutacje. Efekt Woozle Beverly D. Houghton (1979), opisany przez autorkę jako urzeczowienie przez narastanie, wciąż w użyciu metodologicznym w Electronic Commerce Research (Springer, 2024). VentureBeat, „Why do 87% of data science projects never make it into production?" (lipiec 2019) — treść sponsorowana relacjonująca panel konferencyjny; pochodzenie potwierdzone przez relacje wtórne, bo strona zwraca 403. Okazy, cytowane jako artefakty, nie jako dane: Folio3 AI, „AI Project Failure Rate in 2026" (nieistniejąca próba RAND na ponad 2400 inicjatywach, podmiana kwantyfikatora u S&P, prognoza Gartnera z 2018 przedatowana na 2025) oraz Pertama Partners, „AI Project Failure Rate 2026: 80% Fail" (poprawne zastrzeżenie jakościowe w treści, alarmistyczny nagłówek nad nim).