AI · MBA
Too much, too fast: the mechanics of expectation inflation
Most AI post-mortems reach for the technology. The model was not good enough, the data was dirty, the tooling was immature. Gartner's 2026 survey points somewhere less comfortable: among the leaders who failed, many gave the same self-diagnosis, and it was not the model. They said they had expected too much, too fast. If that is right, the failure was priced before the project began — in the expectations, which is a layer you can write down and negotiate while the sponsor is still neutral.
This essay is about that layer: why the gap opens so reliably between what a sponsor expects and what a team can deliver. And what you can put in writing to close it before it costs you a budget. You inherit the technology. You author the expectation.
The one figure worth taking seriously
The survey is easy to quote and easy to quote wrongly, so start with what it says. Gartner questioned about 782 infrastructure and operations leaders across November and December 2025, and published on 7 April 2026. One honesty note first. Gartner's newsroom blocks automated access with a 403, so the wording here comes from The Register and CIO, which quoted the release. The two even disagree on the sample: 782 in one, 783 in the other. I use 782, and I have not read the figure on Gartner's own page.
Now the number itself. Of those leaders, 57% reported at least one failure in applying AI to their area. That is the whole of what 57% measures — the share of leaders who had a failure. It is not the share of projects that failed, and it is not the share who overreached. Among that 57%, many attributed the failure to one thing: they had expected too much, too fast. The phrase belongs to Melanie Freeze, the Gartner analyst quoted in the coverage. It is her gloss on the pattern, not a variable the survey measured.
That distinction matters more than it looks. "Many of the leaders with a failure named expectations" is a modest claim. "57% failed because of expectations" is a different, larger claim the data does not support, and you will see the second one everywhere. Hold the modest version.
What did "too much, too fast" mean concretely? Freeze unpacks it: leaders assumed AI would immediately automate complex tasks, cut costs, or fix long-standing operational problems. They expected speed and full autonomy. When the expectation was not set realistically and the result did not arrive quickly, trust fell and the project stalled. That is a description of an expectation collapsing, not of a model underperforming.
The failures also cluster in a telling place. Gartner reports them concentrated in auto-remediation, self-healing infrastructure, and agent-led workflows — the tasks where leaders expect AI to outrun what the tools actually deliver in messy, unpredictable operations. Inflation rises with the ambition of the task, not with hype in the abstract.
And the survey's own framing puts the cause outside the model. As reported, return depended on integration with existing workflows, on governance, and on fit to a real operational need — not on the sophistication of the model. That single sentence is the whole argument for this essay. If the cause sits in expectations and organisation rather than in the model, then it sits in a layer you can specify before anyone writes code.
The same data set names the barriers that travel with expectation. About 38% of leaders pointed to persistent skill gaps, and the same share cited poor data quality or limited availability as a direct cause of failure. Reported success was highest in IT service management and cloud operations, at 53%. Per-business-unit funding showed up as a structural weakness — each unit paying for its own pilot, nobody owning the integration. None of these is a claim about model quality. They are claims about the organisation around the model, which is the layer the contract works on.
One caution. "Expected too much, too fast" is a self-assessment of cause, not an independent audit. The respondent is explaining their own failure, and that explanation is convenient — it lifts the blame off the choice of use case and off the data. So "it was expectations, not technology" is a thesis supported by self-report, not a root-cause analysis. Treat it as a strong, self-serving signal, and build on it with your eyes open.
In a companion piece I took the same survey's 28% headline apart, because that number is routinely misread as an AI failure rate. This essay does the opposite with the 57%. It takes the finding seriously — and then asks the question the finding invites. If expectation is the failure point, what makes expectation inflate, and can you contract against it?
Where the inflation comes from
Call the phenomenon expectation inflation: the steady drift upward of what a sponsor believes a project will deliver, and how fast, unmoored from what the work can produce. It has a narrative shape, a material channel, and a missing artifact.
The narrative shape has a name Gartner itself coined. Jackie Fenn built the hype cycle at Gartner in 1995, with its "Peak of Inflated Expectations" followed by a "Trough of Disillusionment". As a story about how a market talks itself into overreach, it is useful. As a law, it is not. The most careful review, Steinert and Leifer's 2010 PICMET paper, found no solid empirical basis for the curve and noted that few technologies actually traverse the full trajectory. So borrow the phrase for the mechanism and drop the promise. The peak names how expectations balloon. It does not guarantee that a plateau of productivity waits politely on the far side.
The material channel is the demo-to-production gap. A sponsor buys a vision from a curated demo — a keynote, a vendor's best case, a Friday-afternoon prototype that worked once. The team then meets production, where the edge cases live, the data is partial, and the integration is the actual job. The demo is engineered to succeed; production is engineered by nobody. MIT's NANDA report put a hard number near the gap: roughly 95% of GenAI pilots showed no measurable impact on profit and loss. The unit there is a pilot, not a company, and the figure carries its own caveats — I walk through them in the 28% piece. But the direction is plain. The distance from a demo that impresses to a system that pays is where most of the disappointment lives.
A human detail sits inside the channel. The sponsor rarely buys the vision from the team that will build it. They buy it from a keynote, a peer at a dinner, a competitor's press release — a source with no responsibility for delivery and every reason to compress the timeline. The vision arrives pre-inflated and socially validated, which makes it harder to discount than a number the team itself proposed. By the time it reaches the people who have to ship it, the expectation has already been set by someone who will never be measured against it.
Then the missing artifact: a defined "done". Ask a sponsor and a team, separately, what success looks like, and you will often get two different answers, neither written down. When the target is unstated, it floats, and it floats upward, because every party fills the blank with their own best case. A number nobody committed to on paper is a number that can always have been higher.
Three forces, one direction. Now the engine underneath them.
The mechanism, from first principles
Inflation is not one mistake repeated. It is three independent forces that happen to push the same way, and it is worth deriving each rather than taking any on authority.
The first is information asymmetry. George Akerlof's 1970 paper on the market for lemons showed what happens when one side of a deal knows the quality and the other does not. The informed party's signal dominates, and the market bends around it. Akerlof studied used cars, not AI projects, so this is a lens, not a measurement. But the lens fits, and it points in two directions at once. The vendor and the demo know the best case and sell that signal; the sponsor buys it. And the team, which can see the production truth, has every incentive to keep the bad news quiet — reputation and the next tranche of funding both reward optimism. So the sponsor hears an inflated signal from below and an edited one from within. The two errors compound instead of cancelling.
The second force is escalation of commitment. Barry Staw's 1976 study, "Knee-deep in the Big Muddy", ran 240 students through a resource-allocation decision. It found something durable: people put more money into a losing course when they feel personally responsible for having chosen it. Sunk cost is the folk name; personal responsibility is the actual driver. A sponsor who bought the vision publicly is not a neutral judge of whether to continue. Mark Keil carried the finding out of the lab and into technology, showing in "Pulling the Plug" (1995) and later work that runaway IT projects are escalation of commitment to a failing course. The moment to write a stopping rule is before the sponsor has staked their name, because after that they are motivated to add, not to end.
The third force is the planning fallacy. Kahneman and Tversky named it in 1979 and Kahneman returned to it in Thinking, Fast and Slow (2011). We forecast a project from the inside view: we look at the specific plan in front of us, imagine it going roughly as intended, and read off a best-case estimate. What we skip is the outside view — the distribution of how projects like this one actually went. Estimates inflate not from dishonesty but from method. They are built from within the project, where the plan looks clean, rather than from the record of comparable projects, where the plan usually did not.
You can watch the three hand off to each other on a real timeline. At the pitch, asymmetry sets the price: the sponsor sees the demo's best case and funds against it. At kickoff, commitment locks it: the sponsor has told their peers the number, so the number is now theirs to defend. Through delivery, the planning fallacy keeps the story intact: every status update is written from the inside, where the plan still looks recoverable. By the time production truth arrives, all three are pulling the same rope, and the honest signal — this will not clear the bar — is the one thing nobody is paid to send.
Stack the three and the pattern is not mysterious. The sponsor is sold an inflated signal, is then bound to it by public commitment, and forecasts the rest from inside a story that was optimistic to begin with. None of this needs a bad model to produce a failed project. It only needs an unmanaged expectation.
The expectation contract
If the failure lives in expectations, the intervention is an artifact, negotiated with the sponsor before the project starts, while everyone is still neutral. Call it the expectation contract. It has four parts, and each one disarms a specific force from the section above.
A success criterion, drawn from a reference class. The single most important move is to set the bar from outside the project. Bent Flyvbjerg calls this reference class forecasting: anchor the estimate in the distribution of comparable, completed projects rather than in the plan you are looking at. Kahneman called it "the single most important piece of advice regarding how to increase accuracy in forecasting". In practice, it means the success bar is a measurable outcome taken from how similar deployments actually landed — not a figure lifted from the demo.
A horizon, taken from the same class. When is the criterion judged? Not when the sponsor's patience runs out, but at a point drawn from how long comparable projects took to show a return. The horizon is where "too fast" gets defused, because it converts an impatient instinct into a date agreed in advance.
An explicit statement of what this is not. The floating target is fixed by naming the non-goals. This project deflects tier-one support tickets; it does not replace the support team, it does not touch billing disputes, it does not aim at tier-three. Non-goals stop the scope from inflating every time the sponsor imagines a new best case.
Kill criteria, written before launch. The conditions under which you stop, agreed while the sponsor is still a neutral judge. This is the direct antidote to escalation of commitment. If the sponsor sets the stopping rule before staking their reputation, the rule survives the moment when sunk cost would otherwise take over.
The negotiation itself is a single move repeated: replace the sponsor's inside-view story with an outside-view reference class. When the sponsor says "this should deflect 60% of tickets in a quarter", you do not argue the number. You ask what comparable agent deployments actually delivered, and in what time, and you set the bar and the horizon there. You are not lowering ambition. You are pricing it against evidence.
| Inflation force | How it shows up | Contract clause that disarms it |
|---|---|---|
| Hype peak / demo-to-production gap | success bar set by a curated demo | success criterion drawn from a reference class of similar deployments |
| Planning fallacy (inside view) | horizon read off the plan, best-case | horizon taken from how comparable projects actually ran |
| Escalation of commitment | sponsor keeps funding a stalled project | kill criteria written before launch, while the sponsor is neutral |
| Information asymmetry / floating target | "done" is never defined, scope drifts up | explicit non-goals: a written statement of what this is not |
Two practical questions decide whether this works. First, where does the reference class come from? Three sources, in falling order of trust. Your own past pilots rank first, because they share your data and your constraints. Then public write-ups in the same task family, discounted because people mostly publish wins. Then vendor references, discounted harder, because the vendor chose them. You will rarely get a clean distribution. You are looking for an order of magnitude and a rough time-to-value, which is enough to move the bar off the demo. Second, how is this different from a statement of work? A SOW lists deliverables. The expectation contract governs the belief around them — how good, how soon, judged how, and stopped when. You can hit every deliverable in a SOW and still have a sponsor who feels the project failed, because the deliverables cleared and the expectation did not.
Run it once, concretely. A founder greenlights an AI agent for tier-one support. The inside-view expectation arrives fast and confident: it will deflect most tickets within a quarter and pay for itself by summer. Now build the contract. The reference class here is agent-led workflows — exactly where Gartner's failures concentrate. It sets a sober bar. Suppose comparable deployments deflect somewhere in a modest band, with quality held, and take two to three quarters to get there. Write the criterion as measurable deflection with a quality floor, judged at that horizon. Write the non-goals: not a replacement for the team, not billing, not anything a wrong answer makes expensive. Write the kill line: if deflection has not cleared the floor with quality intact by a named month, you stop, and you agreed that in month zero. The whole artifact costs an afternoon and no capital, and it moves the argument from "did it disappoint" to "did it clear a bar we set together, from evidence, before we started".
The sponsor will push back, and the pushback is where the contract earns its place. "Two to three quarters is too slow" is the honest objection, and it deserves an honest answer, not a lecture. The answer is a question: which comparable deployment hit the bar faster, and what did it have that we do not? Sometimes there is one, and the horizon moves, on evidence. Usually there is not, and the silence is the negotiation. You have not talked the sponsor down. You have shown them the class they were exempting themselves from, and let the exemption fail on its own.
That is the value. The contract does not make the model better. It makes the disappointment impossible to manufacture out of an expectation nobody ever wrote down.
Where the contract fails
A tool sold as universal is sold dishonestly, and the expectation contract has real limits. Three of them matter enough that ignoring any one turns the artifact into theatre.
First: sometimes the high expectation is the point. Ambition is not the enemy here. A sponsor with no vision funds nothing, and a bar set only at the safe, evidenced level can leave most of the available value on the table. Sitkin and colleagues examined this directly in "The Paradox of Stretch Goals" (2011). Stretch goals — targets beyond what current capability suggests — can produce real gains, but only under two conditions: slack resources to absorb the failures, and recent wins that build the confidence to try. They also raise the variance of outcomes, which is the honest cost of reaching. So the contract is not a plea for timid goals; it is the discipline that keeps ambition from becoming a blind gamble. When you have slack and a track record, a stretch bar is warranted, and the contract's job shifts to bounding the downside of the reach, not to capping the reach itself.
Second: expectation management is a fine alibi for sandbagging. The same tool that disarms inflation can be turned to lowball a target when there is real slack to do more. "Managing expectations" then becomes cover for under-ambition, and the reference class becomes a shield for a team that wants an easy win. Sitkin's finding cuts both ways. An organisation with slack and recent successes that sets a timid bar is misusing the outside view exactly as badly as the sponsor who ignores it. The contract has to be honest in both directions. The reference class sets the bar — not the sponsor's fear of looking foolish, and not the team's preference for a number it can beat in its sleep. The tell is simple. If the class supports a higher bar and the team argues it down without new evidence, that is sandbagging wearing the language of prudence.
Third: some inflation is not error, and the contract cannot reach it. Flyvbjerg draws the line that matters. Alongside the planning fallacy, which is an honest cognitive mistake, sits strategic misrepresentation: the deliberate understatement of cost and risk to get a decision approved. A contract negotiated in good faith does nothing against a bad-faith game, because the other party never intended to be bound by evidence. Its cousin is the uniqueness bias — "our project is different, so your reference class does not apply" — which lets a sponsor wave away the outside view entirely. If they will not accept the class, the contract has no anchor to set. And two structural failures finish the list. If the sponsor lacks the power to enforce the kill criteria, because someone above them owns the narrative, the written rule is decoration. And if the field moves fast enough, the reference class itself decays. Comparable deployments from twelve months ago may not be comparable now, and the drift runs both ways. A task that was impossible last year becomes routine; a demo that dazzled collapses at production scale. The outside view assumes a stable class, and AI does not always supply one.
So the contract's reach is precise. It disarms inflation that is honest error, made by parties with the authority to be bound. It does little against inflation that is strategy, and less against a field that keeps rewriting its own reference class. Name that limit out loud, because a founder who believes the artifact covers the political case has swapped one overconfidence for another.
Closing
The comfortable reading of the 57% is that AI disappointed its buyers. The useful reading is that expectation was the cheapest thing in the project to fix, and the one nobody wrote down. You cannot pre-negotiate the model; it will be as good or as ordinary as it turns out to be. You can pre-negotiate the expectation — the bar, the horizon, the non-goals, the kill line — and you can do it while the sponsor is still a neutral judge. That window closes the moment the demo lands. After the demo the sponsor has seen the best case and staked a little of their name on it, and so, quietly, have you. Write the contract before then, or spend next year explaining why the model was fine and the project still failed.
Sources
- primaryGartner, AI Projects in I&O Stall Ahead of Meaningful ROI Returns (press release, 2026) — (newsroom returned 403 to automated access; the 57% figure and Melanie Freeze's "too much, too fast" gloss are verified via the secondary outlets below, not against the primary page)
- primaryBarry M. Staw, "Knee-deep in the Big Muddy: A Study of Escalating Commitment to a Chosen Course of Action," Organizational Behavior and Human Performance 16:27–44 (1976).
- primaryMark Keil, "Pulling the Plug: Software Project Management and the Problem of Project Escalation," MIS Quarterly 19(4):421–447 (1995); extended in Keil et al., MIS Quarterly 24(4) (2000) — escalation carried from the lab into IT projects.
- primaryDaniel Kahneman & Amos Tversky, "Intuitive Prediction: Biases and Corrective Procedures," TIMS Studies in Management Science 12:313–327 (1979) — the planning fallacy and the inside/outside view.
- primaryDaniel Kahneman, Thinking, Fast and Slow (2011), ch. 23 "The Outside View" — popularisation of the inside/outside view and reference class forecasting.
- primarySim B. Sitkin, Kelly E. See, C. Chet Miller, Michael W. Lawless & Andrew M. Carton, "The Paradox of Stretch Goals: Organizations in Pursuit of the Seemingly Impossible," Academy of Management Review 36(3):544–566 (2011).
- primaryBent Flyvbjerg, "From Nobel Prize to Project Management: Getting Risks Right," Project Management Journal 37(3):5–15 (2006); see also "Curbing Optimism Bias and Strategic Misrepresentation in Planning" (2008) — reference class forecasting, strategic misrepresentation, uniqueness bias.
- primaryGeorge A. Akerlof, "The Market for 'Lemons': Quality Uncertainty and the Market Mechanism," Quarterly Journal of Economics 84(3):488–500 (1970) — information asymmetry as an explanatory lens, not a measurement of AI projects.
- primaryMartin Steinert & Larry Leifer, "Scrutinizing Gartner's Hype Cycle Approach," PICMET 2010 Proceedings — the hype cycle lacks a solid empirical basis; few technologies traverse the full curve.
- secondaryThe Register, Only 28% of AI infrastructure projects fully pay off (2026) — (source for the 782 sample and the 57% / 38% figures; the exact phrase "too much, too fast" does not appear in its body — it paraphrases the analyst)
- secondaryCIO, AI often doesn't deliver ROI for IT departments either — (source for the 783 sample and the quoted "many said… expected too much, too fast")
- secondaryWikipedia, Gartner hype cycle — (orientation on Fenn's 1995 origin and the phases; for source facts go to Fenn and to Steinert & Leifer)
Za dużo, za szybko: mechanika inflacji oczekiwań
Większość rozliczeń z porażek AI sięga po technologię. Model był za słaby, dane brudne, narzędzia niedojrzałe. Ankieta Gartnera z 2026 wskazuje miejsce mniej wygodne: wśród liderów, którym nie wyszło, wielu postawiło tę samą samodiagnozę — i nie był to model. Powiedzieli, że oczekiwali za dużo, za szybko. Jeśli to prawda, porażka była wliczona w cenę, zanim projekt ruszył — tkwiła w oczekiwaniach, a to warstwa, którą można spisać i wynegocjować, póki sponsor jest jeszcze bezstronny.
Ten esej jest o tej warstwie: o tym, dlaczego luka między oczekiwaniem sponsora a tym, co zespół dowozi, otwiera się tak niezawodnie. I o tym, co wpisać do umowy, żeby ją zamknąć, zanim kosztuje budżet. Technologię dziedziczysz. Oczekiwanie piszesz sam.
Jedna liczba warta potraktowania serio
Tę ankietę łatwo cytować i łatwo zacytować źle, więc zacznij od tego, co mówi. Gartner przepytał około 782 liderów infrastruktury i operacji w listopadzie i grudniu 2025, a wynik opublikował 7 kwietnia 2026. Najpierw jedna uwaga o uczciwości. Newsroom Gartnera blokuje dostęp automatyczny błędem 403, więc sformułowania biorę z The Register i CIO, które cytowały komunikat. Te dwa źródła różnią się nawet co do próby: 782 w jednym, 783 w drugim. Używam 782 i nie czytałem tej liczby na własnej stronie Gartnera.
Teraz sama liczba. Spośród tych liderów 57% zgłosiło co najmniej jedną porażkę we wdrażaniu AI w swoim obszarze. Te 57% nie mierzy niczego więcej — to odsetek liderów, którym coś się nie udało. To nie odsetek projektów, które padły, ani odsetek tych, którzy przeszarżowali. Wśród tych 57% wielu przypisało porażkę jednej rzeczy: oczekiwali za dużo, za szybko. Fraza należy do Melanie Freeze, analityczki Gartnera cytowanej w doniesieniach. To jej interpretacja wzorca, nie zmienna, którą ankieta zmierzyła.
To rozróżnienie waży więcej, niż wygląda. „Wielu liderów z porażką wskazało oczekiwania" to twierdzenie skromne. „57% padło przez oczekiwania" to inne, większe twierdzenie, na które w danych nie ma pokrycia — i to jego zobaczysz wszędzie. Trzymaj się wersji skromnej.
Co „za dużo, za szybko" znaczyło konkretnie? Freeze to rozwija: liderzy zakładali, że AI od razu zautomatyzuje złożone zadania, zetnie koszty albo naprawi zastarzałe problemy operacyjne. Oczekiwali tempa i pełnej autonomii. Gdy oczekiwania nie ustawiono realistycznie, a wynik nie przyszedł szybko, zaufanie spadło i projekt utknął. To opis oczekiwania, które się załamuje, nie modelu, który dowozi za mało.
Porażki skupiają się też w miejscu wymownym. Gartner podaje, że koncentrują się w auto-remediacji, samonaprawialnej infrastrukturze i procesach prowadzonych przez agenta — tam, gdzie liderzy oczekują, że AI wyprzedzi to, co narzędzia realnie dowożą w bałaganie nieprzewidywalnych operacji. Inflacja rośnie z ambicją zadania, nie z hype'em w oderwaniu.
A sama ankieta stawia przyczynę poza modelem. Jak podano, zwrot zależał od integracji z istniejącymi procesami, od governance i od dopasowania do realnej potrzeby operacyjnej — nie od wyrafinowania modelu. To jedno zdanie jest całym argumentem tego eseju. Jeśli przyczyna siedzi w oczekiwaniach i organizacji, a nie w modelu, to siedzi w warstwie, którą można określić, zanim ktokolwiek napisze kod.
Ten sam zbiór danych nazywa bariery, które chodzą w parze z oczekiwaniem. Około 38% liderów wskazało utrzymujące się luki kompetencyjne, a tyle samo — słabą jakość lub ograniczoną dostępność danych jako bezpośrednią przyczynę porażki. Najwyższą skuteczność zgłoszono w zarządzaniu usługami IT i operacjach chmurowych: 53%. Osobną słabością strukturalną było finansowanie per jednostka biznesowa — każda płaci za własny pilotaż, nikt nie odpowiada za integrację. Żadne z tego nie mówi o jakości modelu. To twierdzenia o organizacji wokół modelu, czyli o warstwie, na której pracuje kontrakt.
Jedno zastrzeżenie. „Oczekiwali za dużo, za szybko" to samoocena przyczyny, nie niezależny audyt. Respondent tłumaczy własną porażkę, a to tłumaczenie jest wygodne — zdejmuje winę z wyboru use case'u i z danych. Więc „to oczekiwania, nie technologia" to teza wsparta samoopisem, nie analizą przyczyny źródłowej. Traktuj ją jako sygnał mocny, lecz interesowny, i buduj na niej z otwartymi oczami.
W tekście towarzyszącym rozłożyłem nagłówkowe 28% z tej samej ankiety, bo tę liczbę rutynowo czyta się jak wskaźnik porażek AI. Ten esej robi z 57% coś odwrotnego. Bierze wynik serio — i stawia pytanie, które ten wynik sam podsuwa. Jeśli punktem porażki jest oczekiwanie, co je nadmuchuje i czy da się przeciw temu zawrzeć umowę?
Skąd bierze się inflacja
Nazwijmy to zjawisko inflacją oczekiwań: powolne dryfowanie w górę tego, co sponsor sądzi, że projekt dowiezie i jak szybko, oderwane od tego, co praca może wyprodukować. Ma kształt narracyjny, kanał materialny i brakujący artefakt.
Kształt narracyjny ma nazwę, którą ukuł sam Gartner. Jackie Fenn zbudowała hype cycle w Gartnerze w 1995. Krzywa ma swój „szczyt napompowanych oczekiwań”, a zaraz po nim „dolinę rozczarowania”. Jako opowieść o tym, jak rynek wmawia sobie przesadę, jest użyteczna. Jako prawo — nie. Najstaranniejszy przegląd zrobili Steinert i Leifer w pracy z konferencji PICMET 2010: nie znaleźli solidnej podstawy empirycznej dla tej krzywej i zauważyli, że mało która technologia przechodzi całą trajektorię. Pożycz więc samą frazę do opisu mechanizmu, a obietnicę odrzuć. Szczyt opisuje, jak oczekiwania puchną. Nie gwarantuje, że po drugiej stronie grzecznie czeka plateau produktywności.
Kanał materialny to przepaść między demem a produkcją. Sponsor kupuje wizję z wyreżyserowanego dema — keynote'u, najlepszego przypadku dostawcy, piątkowego prototypu, który zadziałał raz. Potem zespół spotyka produkcję, gdzie mieszkają przypadki brzegowe, dane są cząstkowe, a integracja jest właściwą robotą. Demo zaprojektowano tak, by się udało; produkcji nie projektuje nikt. Raport MIT NANDA przyłożył do tej przepaści twardą liczbę: około 95% pilotaży generatywnej AI nie pokazało mierzalnego wpływu na wynik finansowy. Jednostką jest tam pilotaż, nie firma, a liczba ma swoje zastrzeżenia — omawiam je w tekście o 28%. Ale kierunek jest jasny. Dystans od dema, które robi wrażenie, do systemu, który się spłaca, to miejsce, w którym rodzi się większość rozczarowań.
W kanale siedzi ludzki szczegół. Sponsor rzadko kupuje wizję od zespołu, który ją zbuduje. Kupuje ją z keynote'u, od znajomego przy kolacji, z komunikatu konkurenta — od źródła, które nie odpowiada za dowóz i ma każdy powód, by ścisnąć harmonogram. Wizja przychodzi wstępnie napompowana i społecznie potwierdzona, przez co trudniej ją zdyskontować niż liczbę zaproponowaną przez sam zespół. Zanim dotrze do ludzi, którzy mają ją dowieźć, oczekiwanie ustawił już ktoś, kogo nikt nigdy z niego nie rozliczy.
Potem brakujący artefakt: zdefiniowane „gotowe". Zapytaj sponsora i zespół osobno, jak wygląda sukces, a często dostaniesz dwie różne odpowiedzi, żadnej spisanej. Gdy cel jest niewypowiedziany, płynie — i płynie w górę, bo każda strona wypełnia lukę własnym najlepszym przypadkiem. Liczba, do której nikt nie zobowiązał się na papierze, to liczba, która zawsze mogła być wyższa.
Trzy siły, jeden kierunek. Teraz silnik pod nimi.
Mechanizm od pierwszych zasad
Inflacja nie jest jednym błędem powtarzanym w kółko. Napędzają ją trzy niezależne siły, które akurat pchają w tę samą stronę — i każdą warto wyprowadzić, a nie brać na autorytet.
Pierwsza to asymetria informacji. Praca George'a Akerlofa z 1970 o rynku cytryn pokazała, co się dzieje, gdy jedna strona transakcji zna jakość, a druga nie. Sygnał strony poinformowanej dominuje, a rynek się pod niego wygina. Akerlof badał używane auta, nie projekty AI, więc to soczewka, nie pomiar. Ale soczewka pasuje i celuje w dwie strony naraz. Dostawca i demo znają najlepszy przypadek i sprzedają ten sygnał; sponsor go kupuje. A zespół, który widzi prawdę produkcji, ma wszelkie powody, by trzymać złe wieści przy sobie — reputacja i kolejna transza finansowania nagradzają optymizm. Sponsor słyszy więc dwa sygnały: napompowany z dołu i wygładzony od środka. Dwa błędy się kumulują, zamiast znosić.
Druga siła to eskalacja zaangażowania. Badanie Barry'ego Stawa z 1976, „Knee-deep in the Big Muddy", postawiło 240 studentów przed decyzją o alokacji zasobów. Znalazło coś trwałego: ludzie dokładają więcej pieniędzy do przegrywającego przedsięwzięcia, gdy czują się osobiście odpowiedzialni za to, że je wybrali. Koszt utopiony to nazwa potoczna; prawdziwym motorem jest osobista odpowiedzialność. Sponsor, który publicznie kupił wizję, nie jest bezstronnym sędzią tego, czy kontynuować. Mark Keil wyniósł to odkrycie z laboratorium do technologii: w „Pulling the Plug” (1995) i późniejszych pracach pokazał, że rozpędzone projekty IT to eskalacja zaangażowania w przegrywające przedsięwzięcie. Moment na spisanie reguły stopu jest, zanim sponsor postawi swoje nazwisko — bo potem ma motyw, by dokładać, nie kończyć.
Trzecia siła to błąd planowania. Kahneman i Tversky nazwali go w 1979, a Kahneman wrócił do niego w Thinking, Fast and Slow (2011). Projekt prognozujemy z widoku od wewnątrz: patrzymy na konkretny plan przed nami, wyobrażamy sobie, że idzie mniej więcej zgodnie z zamysłem, i odczytujemy najlepszy przypadek. Pomijamy widok z zewnątrz — rozkład tego, jak naprawdę poszły projekty podobne do tego. Szacunki puchną nie z nieuczciwości, lecz z metody. Budujemy je od wewnątrz projektu, gdzie plan wygląda czysto, a nie z zapisu porównywalnych projektów, gdzie plan zwykle czysty nie był.
Widać, jak cała trójka podaje sobie pałeczkę na realnej osi czasu. Przy pitchu asymetria ustala cenę: sponsor widzi najlepszy przypadek dema i finansuje pod niego. Przy starcie zobowiązanie ją blokuje: sponsor podał kolegom liczbę, więc teraz to on musi jej bronić. Przez cały dowóz błąd planowania trzyma opowieść w całości: każdy status pisze się od wewnątrz, gdzie plan wciąż wygląda, jakby dało się go uratować. Zanim przyjdzie prawda produkcji, wszystkie trzy ciągną za tę samą linę, a uczciwy sygnał — poprzeczki nie przeskoczymy — jest jedyną rzeczą, za której wysłanie nikomu nie płacą.
Ułóż te trzy w stos, a wzorzec przestaje być tajemnicą. Sponsorowi sprzedano napompowany sygnał, potem związano go publicznym zobowiązaniem, a resztę prognozuje z wnętrza opowieści, która od początku była optymistyczna. Nic z tego nie potrzebuje złego modelu, żeby dać nieudany projekt. Potrzebuje tylko niezarządzanego oczekiwania.
Kontrakt na oczekiwania
Jeśli porażka rodzi się w oczekiwaniach, odpowiedzią jest artefakt — wynegocjowany ze sponsorem, zanim projekt ruszy, póki wszyscy są jeszcze bezstronni. Nazwijmy go kontraktem na oczekiwania. Ma cztery części, a każda rozbraja konkretną siłę z poprzedniej sekcji.
Kryterium sukcesu wzięte z klasy referencyjnej. Najważniejszy pojedynczy ruch to ustawić poprzeczkę spoza projektu. Bent Flyvbjerg nazywa to prognozowaniem z klasy referencyjnej: zakotwicz szacunek w rozkładzie porównywalnych, ukończonych projektów, a nie w planie, na który patrzysz. Kahneman nazwał to „najważniejszą pojedynczą radą, jak zwiększyć trafność prognoz". W praktyce znaczy to, że poprzeczka sukcesu jest mierzalnym wynikiem wziętym z tego, jak naprawdę wylądowały podobne wdrożenia — nie liczbą wyjętą z dema.
Horyzont wzięty z tej samej klasy. Kiedy oceniamy kryterium? Nie gdy sponsorowi kończy się cierpliwość, lecz w punkcie wziętym z tego, ile porównywalne projekty potrzebowały, by pokazać zwrot. Horyzont jest miejscem, gdzie rozbraja się „za szybko", bo zamienia niecierpliwy odruch w datę uzgodnioną z góry.
Jawne określenie, czym to nie jest. Płynący cel przybijasz, nazywając nie-cele. Ten projekt odbija zgłoszenia wsparcia pierwszej linii; nie zastępuje zespołu wsparcia, nie dotyka sporów o płatności, nie celuje w trzecią linię. Nie-cele powstrzymują zakres przed puchnięciem za każdym razem, gdy sponsor wyobrazi sobie nowy najlepszy przypadek.
Kryteria stopu spisane przed startem. Warunki, w których przerywasz, uzgodnione, póki sponsor jest jeszcze bezstronnym sędzią. To bezpośrednia odtrutka na eskalację zaangażowania. Jeśli sponsor ustawi regułę stopu, zanim postawi reputację, reguła przeżyje moment, w którym koszt utopiony inaczej wziąłby górę.
Sama negocjacja to jeden powtarzany ruch: zamień opowieść sponsora z widoku od wewnątrz na klasę referencyjną z widoku z zewnątrz. Gdy sponsor mówi „to powinno odbić 60% zgłoszeń w kwartał", nie spierasz się o liczbę. Pytasz, co realnie dowiozły porównywalne wdrożenia agentów i w jakim czasie, i tam ustawiasz poprzeczkę oraz horyzont. Nie obniżasz ambicji. Wyceniasz ją wobec dowodów.
| Siła inflacji | Jak się objawia | Klauzula kontraktu, która ją rozbraja |
|---|---|---|
| Szczyt hype'u / przepaść demo–produkcja | poprzeczka ustawiona przez wyreżyserowane demo | kryterium sukcesu z klasy referencyjnej podobnych wdrożeń |
| Błąd planowania (widok od wewnątrz) | horyzont odczytany z planu, najlepszy przypadek | horyzont wzięty z tego, jak naprawdę szły porównywalne projekty |
| Eskalacja zaangażowania | sponsor wciąż finansuje utknięty projekt | kryteria stopu spisane przed startem, póki sponsor jest bezstronny |
| Asymetria informacji / płynący cel | „gotowe" nigdy nie zdefiniowane, zakres dryfuje w górę | jawne nie-cele: spisane określenie, czym to nie jest |
O tym, czy to działa, rozstrzygają dwa praktyczne pytania. Pierwsze: skąd bierze się klasa referencyjna? Trzy źródła, od najbardziej do najmniej wiarygodnego. Twoje własne dawne pilotaże są pierwsze, bo powstały na twoich danych i w twoich ograniczeniach. Potem publiczne opisy z tej samej rodziny zadań, z dyskontem, bo ludzie publikują głównie zwycięstwa. Potem referencje dostawcy, z dyskontem większym, bo to dostawca je wybrał. Rzadko dostaniesz czysty rozkład. Szukasz rzędu wielkości i zgrubnego czasu do zwrotu, a to wystarcza, by zdjąć poprzeczkę z dema. Drugie pytanie: czym to różni się od specyfikacji zakresu prac? SOW wylicza produkty do dostarczenia. Kontrakt na oczekiwania rządzi przekonaniem wokół nich — jak dobrze, jak szybko, ocenione jak i zatrzymane kiedy. Możesz trafić w każdy produkt z SOW i wciąż mieć sponsora, który czuje, że projekt zawiódł, bo produkty się dowiozły, a oczekiwanie nie.
Przećwicz to raz, na konkrecie. Founder daje zielone światło agentowi AI do wsparcia pierwszej linii. Oczekiwanie z widoku od wewnątrz przychodzi szybko i pewnie: odbije większość zgłoszeń w kwartał i zwróci się do lata. Teraz zbuduj kontrakt. Klasa referencyjna to tutaj procesy prowadzone przez agenta — dokładnie tam, gdzie skupiają się porażki Gartnera. Ustawia trzeźwą poprzeczkę. Załóżmy, że porównywalne wdrożenia odbijają zgłoszenia w skromnym paśmie, przy utrzymanej jakości, i potrzebują na to dwóch–trzech kwartałów. Zapisz kryterium jako mierzalne odbicie z podłogą jakości, ocenione w tym horyzoncie. Zapisz nie-cele: nie zastępstwo zespołu, nie płatności, nie nic, co przy błędnej odpowiedzi robi się kosztowne. Zapisz linię stopu: jeśli odbicie nie przejdzie podłogi przy zachowanej jakości do nazwanego miesiąca, przerywasz — a zgodziłeś się na to w miesiącu zero. Cały artefakt kosztuje jedno popołudnie i zero kapitału, a przenosi spór z „czy rozczarowało" na „czy przeszło poprzeczkę, którą ustawiliśmy razem, z dowodów, zanim zaczęliśmy".
Sponsor się postawi, a ten opór jest miejscem, w którym kontrakt zarabia na swoje istnienie. „Dwa–trzy kwartały to za wolno" to uczciwy zarzut i zasługuje na uczciwą odpowiedź, nie na wykład. Odpowiedzią jest pytanie: które porównywalne wdrożenie trafiło w poprzeczkę szybciej i co miało, czego my nie mamy? Czasem takie jest i horyzont się przesuwa — na dowodach. Zwykle takiego nie ma, a cisza jest negocjacją. Nie zagadałeś sponsora. Pokazałeś mu klasę, z której sam siebie wyłączał, i pozwoliłeś, by to wyłączenie runęło pod własnym ciężarem.
To jest ta wartość. Kontrakt nie czyni modelu lepszym. Czyni rozczarowanie niemożliwym do wyprodukowania z oczekiwania, którego nikt nigdy nie spisał.
Gdzie kontrakt zawodzi
Narzędzie sprzedawane jako uniwersalne jest sprzedawane nieuczciwie, a kontrakt na oczekiwania ma realne granice. Trzy z nich ważą tyle, że zignorowanie którejkolwiek zmienia artefakt w teatr.
Pierwsza: czasem o wysokie oczekiwanie właśnie chodzi. Ambicja nie jest tu wrogiem. Sponsor bez wizji nie finansuje niczego, a poprzeczka ustawiona tylko na bezpiecznym, udowodnionym poziomie zostawia większość dostępnej wartości na stole. Sitkin ze współpracownikami zbadał to wprost w „The Paradox of Stretch Goals" (2011). Cele rozciągnięte, czyli stretch goals ustawione ponad to, na co wskazują obecne możliwości, potrafią dać realny zysk, ale tylko przy dwóch warunkach: zapasie zasobów, który wchłonie porażki, i świeżych zwycięstwach, które budują pewność, by próbować. Podnoszą też wariancję wyników, co jest uczciwym kosztem sięgania. Kontrakt nie jest więc prośbą o cele nieśmiałe; jest dyscypliną, która nie pozwala ambicji zamienić się w ślepy zakład. Gdy masz zapas i historię wygranych, poprzeczka ambitna jest uzasadniona, a rola kontraktu przesuwa się na ograniczanie downside'u sięgnięcia, nie na przycinanie samego zasięgu.
Druga: zarządzanie oczekiwaniami to świetne alibi dla sandbaggingu. To samo narzędzie, które rozbraja inflację, można obrócić w zaniżanie celu, gdy realnie jest zapas, by zrobić więcej. „Zarządzanie oczekiwaniami" staje się wtedy przykrywką dla braku ambicji, a klasa referencyjna — tarczą dla zespołu, który chce łatwej wygranej. Odkrycie Sitkina tnie w obie strony. Organizacja z zapasem i świeżymi sukcesami, która stawia nieśmiałą poprzeczkę, nadużywa widoku z zewnątrz dokładnie tak samo źle jak sponsor, który go ignoruje. Kontrakt musi być uczciwy w obie strony. Poprzeczkę ustawia klasa referencyjna — nie lęk sponsora przed ośmieszeniem i nie preferencja zespołu do liczby, którą pobije z zamkniętymi oczami. Sygnał jest prosty. Jeśli klasa uzasadnia wyższą poprzeczkę, a zespół zbija ją bez nowych dowodów, to sandbagging w języku roztropności.
Trzecia: część inflacji to nie błąd i kontrakt jej nie sięga. Flyvbjerg rysuje granicę, która się liczy. Obok błędu planowania, który jest uczciwą pomyłką poznawczą, stoi strategiczne zafałszowanie: celowe zaniżanie kosztu i ryzyka, żeby przepchnąć decyzję do zatwierdzenia. Kontrakt zawarty w dobrej wierze nic nie zrobi przeciw grze w złej wierze, bo druga strona nigdy nie zamierzała dać się związać dowodom. Jego kuzynem jest uniqueness bias, czyli przeświadczenie o własnej inności: „nasz projekt jest inny, więc twoja klasa referencyjna nie ma zastosowania”. Ono pozwala sponsorowi machnąć ręką na cały widok z zewnątrz. Jeśli nie przyjmie klasy, kontrakt nie ma kotwicy do ustawienia. A listę domykają dwie porażki strukturalne. Jeśli sponsor nie ma władzy, by wyegzekwować kryteria stopu, bo narracją włada ktoś nad nim, spisana reguła jest dekoracją. A jeśli dziedzina zmienia się dość szybko, sama klasa referencyjna się psuje. Porównywalne wdrożenia sprzed dwunastu miesięcy mogą już nie być porównywalne, a dryf idzie w obie strony. Zadanie niemożliwe rok temu staje się rutyną; demo, które olśniewało, wali się w skali produkcji. Widok z zewnątrz zakłada stabilną klasę, a AI nie zawsze taką dostarcza.
Zasięg kontraktu jest zatem ściśle wytyczony. Rozbraja inflację, która jest uczciwym błędem popełnionym przez strony na tyle umocowane, że umowa naprawdę je wiąże. Niewiele wskóra przeciw inflacji, która jest strategią, i jeszcze mniej przeciw dziedzinie, która wciąż przepisuje własną klasę referencyjną. Nazwij tę granicę wprost, bo founder, który wierzy, że artefakt obejmuje też przypadek polityczny, zamienił jedną nadmierną pewność na inną.
Zamknięcie
Wygodne odczytanie 57% brzmi: AI zawiodło kupujących. Użyteczne odczytanie brzmi: oczekiwanie było najtańszą rzeczą w projekcie do naprawy — i jedyną, której nikt nie spisał. Modelu nie wynegocjujesz z góry; będzie tak dobry albo tak zwyczajny, jak wyjdzie. Oczekiwanie wynegocjujesz z góry — poprzeczkę, horyzont, nie-cele, linię stopu — i zrobisz to, póki sponsor jest jeszcze bezstronnym sędzią. To okno zamyka się w chwili, gdy ląduje demo. Po demie sponsor zobaczył najlepszy przypadek i postawił na niego trochę swojego nazwiska; po cichu ty także. Spisz kontrakt wcześniej albo spędź przyszły rok, tłumacząc, dlaczego model był w porządku, a projekt i tak padł.
Bibliografia
- primaryGartner, AI Projects in I&O Stall Ahead of Meaningful ROI Returns (komunikat prasowy, 2026) — (newsroom zwrócił 403 przy dostępie automatycznym; liczba 57% oraz interpretacja „too much, too fast" Melanie Freeze zweryfikowane przez źródła wtórne poniżej, nie wobec strony pierwotnej)
- primaryBarry M. Staw, "Knee-deep in the Big Muddy: A Study of Escalating Commitment to a Chosen Course of Action," Organizational Behavior and Human Performance 16:27–44 (1976).
- primaryMark Keil, "Pulling the Plug: Software Project Management and the Problem of Project Escalation," MIS Quarterly 19(4):421–447 (1995); rozszerzone w Keil i in., MIS Quarterly 24(4) (2000) — eskalacja przeniesiona z laboratorium do projektów IT.
- primaryDaniel Kahneman & Amos Tversky, "Intuitive Prediction: Biases and Corrective Procedures," TIMS Studies in Management Science 12:313–327 (1979) — błąd planowania oraz widok od wewnątrz i z zewnątrz.
- primaryDaniel Kahneman, Thinking, Fast and Slow (2011), rozdz. 23 "The Outside View" — popularyzacja widoku od wewnątrz i z zewnątrz oraz prognozowania z klasy referencyjnej.
- primarySim B. Sitkin, Kelly E. See, C. Chet Miller, Michael W. Lawless & Andrew M. Carton, "The Paradox of Stretch Goals: Organizations in Pursuit of the Seemingly Impossible," Academy of Management Review 36(3):544–566 (2011).
- primaryBent Flyvbjerg, "From Nobel Prize to Project Management: Getting Risks Right," Project Management Journal 37(3):5–15 (2006); zob. też "Curbing Optimism Bias and Strategic Misrepresentation in Planning" (2008) — prognozowanie z klasy referencyjnej, strategiczne zafałszowanie, błąd wyjątkowości.
- primaryGeorge A. Akerlof, "The Market for 'Lemons': Quality Uncertainty and the Market Mechanism," Quarterly Journal of Economics 84(3):488–500 (1970) — asymetria informacji jako soczewka wyjaśniająca, nie pomiar projektów AI.
- primaryMartin Steinert & Larry Leifer, "Scrutinizing Gartner's Hype Cycle Approach," PICMET 2010 Proceedings — hype cycle nie ma solidnej podstawy empirycznej; mało która technologia przechodzi całą krzywą.
- secondaryThe Register, Only 28% of AI infrastructure projects fully pay off (2026) — (źródło próby 782 oraz liczb 57% / 38%; dokładna fraza „too much, too fast" nie pojawia się w treści — serwis parafrazuje analityczkę)
- secondaryCIO, AI often doesn't deliver ROI for IT departments either — (źródło próby 783 oraz cytatu „many said… expected too much, too fast")
- secondaryWikipedia, Gartner hype cycle — (orientacja co do genezy z 1995 (Fenn) i faz; po fakty źródłowe idź do Fenn oraz Steinerta i Leifera)