← Writing

AI · MBA

Too much, too fast: the mechanics of expectation inflation

Most AI post-mortems reach for the technology. The model was not good enough, the data was dirty, the tooling was immature. Gartner's 2026 survey points somewhere less comfortable: among the leaders who failed, many gave the same self-diagnosis, and it was not the model. They said they had expected too much, too fast. If that is right, the failure was priced before the project began — in the expectations, which is a layer you can write down and negotiate while the sponsor is still neutral.

This essay is about that layer: why the gap opens so reliably between what a sponsor expects and what a team can deliver. And what you can put in writing to close it before it costs you a budget. You inherit the technology. You author the expectation.

The one figure worth taking seriously

The survey is easy to quote and easy to quote wrongly, so start with what it says. Gartner questioned about 782 infrastructure and operations leaders across November and December 2025, and published on 7 April 2026. One honesty note first. Gartner's newsroom blocks automated access with a 403, so the wording here comes from The Register and CIO, which quoted the release. The two even disagree on the sample: 782 in one, 783 in the other. I use 782, and I have not read the figure on Gartner's own page.

Now the number itself. Of those leaders, 57% reported at least one failure in applying AI to their area. That is the whole of what 57% measures — the share of leaders who had a failure. It is not the share of projects that failed, and it is not the share who overreached. Among that 57%, many attributed the failure to one thing: they had expected too much, too fast. The phrase belongs to Melanie Freeze, the Gartner analyst quoted in the coverage. It is her gloss on the pattern, not a variable the survey measured.

That distinction matters more than it looks. "Many of the leaders with a failure named expectations" is a modest claim. "57% failed because of expectations" is a different, larger claim the data does not support, and you will see the second one everywhere. Hold the modest version.

What did "too much, too fast" mean concretely? Freeze unpacks it: leaders assumed AI would immediately automate complex tasks, cut costs, or fix long-standing operational problems. They expected speed and full autonomy. When the expectation was not set realistically and the result did not arrive quickly, trust fell and the project stalled. That is a description of an expectation collapsing, not of a model underperforming.

The failures also cluster in a telling place. Gartner reports them concentrated in auto-remediation, self-healing infrastructure, and agent-led workflows — the tasks where leaders expect AI to outrun what the tools actually deliver in messy, unpredictable operations. Inflation rises with the ambition of the task, not with hype in the abstract.

And the survey's own framing puts the cause outside the model. As reported, return depended on integration with existing workflows, on governance, and on fit to a real operational need — not on the sophistication of the model. That single sentence is the whole argument for this essay. If the cause sits in expectations and organisation rather than in the model, then it sits in a layer you can specify before anyone writes code.

The same data set names the barriers that travel with expectation. About 38% of leaders pointed to persistent skill gaps, and the same share cited poor data quality or limited availability as a direct cause of failure. Reported success was highest in IT service management and cloud operations, at 53%. Per-business-unit funding showed up as a structural weakness — each unit paying for its own pilot, nobody owning the integration. None of these is a claim about model quality. They are claims about the organisation around the model, which is the layer the contract works on.

One caution. "Expected too much, too fast" is a self-assessment of cause, not an independent audit. The respondent is explaining their own failure, and that explanation is convenient — it lifts the blame off the choice of use case and off the data. So "it was expectations, not technology" is a thesis supported by self-report, not a root-cause analysis. Treat it as a strong, self-serving signal, and build on it with your eyes open.

In a companion piece I took the same survey's 28% headline apart, because that number is routinely misread as an AI failure rate. This essay does the opposite with the 57%. It takes the finding seriously — and then asks the question the finding invites. If expectation is the failure point, what makes expectation inflate, and can you contract against it?

Where the inflation comes from

Call the phenomenon expectation inflation: the steady drift upward of what a sponsor believes a project will deliver, and how fast, unmoored from what the work can produce. It has a narrative shape, a material channel, and a missing artifact.

The narrative shape has a name Gartner itself coined. Jackie Fenn built the hype cycle at Gartner in 1995, with its "Peak of Inflated Expectations" followed by a "Trough of Disillusionment". As a story about how a market talks itself into overreach, it is useful. As a law, it is not. The most careful review, Steinert and Leifer's 2010 PICMET paper, found no solid empirical basis for the curve and noted that few technologies actually traverse the full trajectory. So borrow the phrase for the mechanism and drop the promise. The peak names how expectations balloon. It does not guarantee that a plateau of productivity waits politely on the far side.

The material channel is the demo-to-production gap. A sponsor buys a vision from a curated demo — a keynote, a vendor's best case, a Friday-afternoon prototype that worked once. The team then meets production, where the edge cases live, the data is partial, and the integration is the actual job. The demo is engineered to succeed; production is engineered by nobody. MIT's NANDA report put a hard number near the gap: roughly 95% of GenAI pilots showed no measurable impact on profit and loss. The unit there is a pilot, not a company, and the figure carries its own caveats — I walk through them in the 28% piece. But the direction is plain. The distance from a demo that impresses to a system that pays is where most of the disappointment lives.

A human detail sits inside the channel. The sponsor rarely buys the vision from the team that will build it. They buy it from a keynote, a peer at a dinner, a competitor's press release — a source with no responsibility for delivery and every reason to compress the timeline. The vision arrives pre-inflated and socially validated, which makes it harder to discount than a number the team itself proposed. By the time it reaches the people who have to ship it, the expectation has already been set by someone who will never be measured against it.

Then the missing artifact: a defined "done". Ask a sponsor and a team, separately, what success looks like, and you will often get two different answers, neither written down. When the target is unstated, it floats, and it floats upward, because every party fills the blank with their own best case. A number nobody committed to on paper is a number that can always have been higher.

Three forces, one direction. Now the engine underneath them.

The mechanism, from first principles

Inflation is not one mistake repeated. It is three independent forces that happen to push the same way, and it is worth deriving each rather than taking any on authority.

The first is information asymmetry. George Akerlof's 1970 paper on the market for lemons showed what happens when one side of a deal knows the quality and the other does not. The informed party's signal dominates, and the market bends around it. Akerlof studied used cars, not AI projects, so this is a lens, not a measurement. But the lens fits, and it points in two directions at once. The vendor and the demo know the best case and sell that signal; the sponsor buys it. And the team, which can see the production truth, has every incentive to keep the bad news quiet — reputation and the next tranche of funding both reward optimism. So the sponsor hears an inflated signal from below and an edited one from within. The two errors compound instead of cancelling.

The second force is escalation of commitment. Barry Staw's 1976 study, "Knee-deep in the Big Muddy", ran 240 students through a resource-allocation decision. It found something durable: people put more money into a losing course when they feel personally responsible for having chosen it. Sunk cost is the folk name; personal responsibility is the actual driver. A sponsor who bought the vision publicly is not a neutral judge of whether to continue. Mark Keil carried the finding out of the lab and into technology, showing in "Pulling the Plug" (1995) and later work that runaway IT projects are escalation of commitment to a failing course. The moment to write a stopping rule is before the sponsor has staked their name, because after that they are motivated to add, not to end.

The third force is the planning fallacy. Kahneman and Tversky named it in 1979 and Kahneman returned to it in Thinking, Fast and Slow (2011). We forecast a project from the inside view: we look at the specific plan in front of us, imagine it going roughly as intended, and read off a best-case estimate. What we skip is the outside view — the distribution of how projects like this one actually went. Estimates inflate not from dishonesty but from method. They are built from within the project, where the plan looks clean, rather than from the record of comparable projects, where the plan usually did not.

You can watch the three hand off to each other on a real timeline. At the pitch, asymmetry sets the price: the sponsor sees the demo's best case and funds against it. At kickoff, commitment locks it: the sponsor has told their peers the number, so the number is now theirs to defend. Through delivery, the planning fallacy keeps the story intact: every status update is written from the inside, where the plan still looks recoverable. By the time production truth arrives, all three are pulling the same rope, and the honest signal — this will not clear the bar — is the one thing nobody is paid to send.

Stack the three and the pattern is not mysterious. The sponsor is sold an inflated signal, is then bound to it by public commitment, and forecasts the rest from inside a story that was optimistic to begin with. None of this needs a bad model to produce a failed project. It only needs an unmanaged expectation.

The expectation contract

If the failure lives in expectations, the intervention is an artifact, negotiated with the sponsor before the project starts, while everyone is still neutral. Call it the expectation contract. It has four parts, and each one disarms a specific force from the section above.

A success criterion, drawn from a reference class. The single most important move is to set the bar from outside the project. Bent Flyvbjerg calls this reference class forecasting: anchor the estimate in the distribution of comparable, completed projects rather than in the plan you are looking at. Kahneman called it "the single most important piece of advice regarding how to increase accuracy in forecasting". In practice, it means the success bar is a measurable outcome taken from how similar deployments actually landed — not a figure lifted from the demo.

A horizon, taken from the same class. When is the criterion judged? Not when the sponsor's patience runs out, but at a point drawn from how long comparable projects took to show a return. The horizon is where "too fast" gets defused, because it converts an impatient instinct into a date agreed in advance.

An explicit statement of what this is not. The floating target is fixed by naming the non-goals. This project deflects tier-one support tickets; it does not replace the support team, it does not touch billing disputes, it does not aim at tier-three. Non-goals stop the scope from inflating every time the sponsor imagines a new best case.

Kill criteria, written before launch. The conditions under which you stop, agreed while the sponsor is still a neutral judge. This is the direct antidote to escalation of commitment. If the sponsor sets the stopping rule before staking their reputation, the rule survives the moment when sunk cost would otherwise take over.

The negotiation itself is a single move repeated: replace the sponsor's inside-view story with an outside-view reference class. When the sponsor says "this should deflect 60% of tickets in a quarter", you do not argue the number. You ask what comparable agent deployments actually delivered, and in what time, and you set the bar and the horizon there. You are not lowering ambition. You are pricing it against evidence.

Inflation forceHow it shows upContract clause that disarms it
Hype peak / demo-to-production gapsuccess bar set by a curated demosuccess criterion drawn from a reference class of similar deployments
Planning fallacy (inside view)horizon read off the plan, best-casehorizon taken from how comparable projects actually ran
Escalation of commitmentsponsor keeps funding a stalled projectkill criteria written before launch, while the sponsor is neutral
Information asymmetry / floating target"done" is never defined, scope drifts upexplicit non-goals: a written statement of what this is not

Two practical questions decide whether this works. First, where does the reference class come from? Three sources, in falling order of trust. Your own past pilots rank first, because they share your data and your constraints. Then public write-ups in the same task family, discounted because people mostly publish wins. Then vendor references, discounted harder, because the vendor chose them. You will rarely get a clean distribution. You are looking for an order of magnitude and a rough time-to-value, which is enough to move the bar off the demo. Second, how is this different from a statement of work? A SOW lists deliverables. The expectation contract governs the belief around them — how good, how soon, judged how, and stopped when. You can hit every deliverable in a SOW and still have a sponsor who feels the project failed, because the deliverables cleared and the expectation did not.

Run it once, concretely. A founder greenlights an AI agent for tier-one support. The inside-view expectation arrives fast and confident: it will deflect most tickets within a quarter and pay for itself by summer. Now build the contract. The reference class here is agent-led workflows — exactly where Gartner's failures concentrate. It sets a sober bar. Suppose comparable deployments deflect somewhere in a modest band, with quality held, and take two to three quarters to get there. Write the criterion as measurable deflection with a quality floor, judged at that horizon. Write the non-goals: not a replacement for the team, not billing, not anything a wrong answer makes expensive. Write the kill line: if deflection has not cleared the floor with quality intact by a named month, you stop, and you agreed that in month zero. The whole artifact costs an afternoon and no capital, and it moves the argument from "did it disappoint" to "did it clear a bar we set together, from evidence, before we started".

The sponsor will push back, and the pushback is where the contract earns its place. "Two to three quarters is too slow" is the honest objection, and it deserves an honest answer, not a lecture. The answer is a question: which comparable deployment hit the bar faster, and what did it have that we do not? Sometimes there is one, and the horizon moves, on evidence. Usually there is not, and the silence is the negotiation. You have not talked the sponsor down. You have shown them the class they were exempting themselves from, and let the exemption fail on its own.

That is the value. The contract does not make the model better. It makes the disappointment impossible to manufacture out of an expectation nobody ever wrote down.

Where the contract fails

A tool sold as universal is sold dishonestly, and the expectation contract has real limits. Three of them matter enough that ignoring any one turns the artifact into theatre.

First: sometimes the high expectation is the point. Ambition is not the enemy here. A sponsor with no vision funds nothing, and a bar set only at the safe, evidenced level can leave most of the available value on the table. Sitkin and colleagues examined this directly in "The Paradox of Stretch Goals" (2011). Stretch goals — targets beyond what current capability suggests — can produce real gains, but only under two conditions: slack resources to absorb the failures, and recent wins that build the confidence to try. They also raise the variance of outcomes, which is the honest cost of reaching. So the contract is not a plea for timid goals; it is the discipline that keeps ambition from becoming a blind gamble. When you have slack and a track record, a stretch bar is warranted, and the contract's job shifts to bounding the downside of the reach, not to capping the reach itself.

Second: expectation management is a fine alibi for sandbagging. The same tool that disarms inflation can be turned to lowball a target when there is real slack to do more. "Managing expectations" then becomes cover for under-ambition, and the reference class becomes a shield for a team that wants an easy win. Sitkin's finding cuts both ways. An organisation with slack and recent successes that sets a timid bar is misusing the outside view exactly as badly as the sponsor who ignores it. The contract has to be honest in both directions. The reference class sets the bar — not the sponsor's fear of looking foolish, and not the team's preference for a number it can beat in its sleep. The tell is simple. If the class supports a higher bar and the team argues it down without new evidence, that is sandbagging wearing the language of prudence.

Third: some inflation is not error, and the contract cannot reach it. Flyvbjerg draws the line that matters. Alongside the planning fallacy, which is an honest cognitive mistake, sits strategic misrepresentation: the deliberate understatement of cost and risk to get a decision approved. A contract negotiated in good faith does nothing against a bad-faith game, because the other party never intended to be bound by evidence. Its cousin is the uniqueness bias — "our project is different, so your reference class does not apply" — which lets a sponsor wave away the outside view entirely. If they will not accept the class, the contract has no anchor to set. And two structural failures finish the list. If the sponsor lacks the power to enforce the kill criteria, because someone above them owns the narrative, the written rule is decoration. And if the field moves fast enough, the reference class itself decays. Comparable deployments from twelve months ago may not be comparable now, and the drift runs both ways. A task that was impossible last year becomes routine; a demo that dazzled collapses at production scale. The outside view assumes a stable class, and AI does not always supply one.

So the contract's reach is precise. It disarms inflation that is honest error, made by parties with the authority to be bound. It does little against inflation that is strategy, and less against a field that keeps rewriting its own reference class. Name that limit out loud, because a founder who believes the artifact covers the political case has swapped one overconfidence for another.

Closing

The comfortable reading of the 57% is that AI disappointed its buyers. The useful reading is that expectation was the cheapest thing in the project to fix, and the one nobody wrote down. You cannot pre-negotiate the model; it will be as good or as ordinary as it turns out to be. You can pre-negotiate the expectation — the bar, the horizon, the non-goals, the kill line — and you can do it while the sponsor is still a neutral judge. That window closes the moment the demo lands. After the demo the sponsor has seen the best case and staked a little of their name on it, and so, quietly, have you. Write the contract before then, or spend next year explaining why the model was fine and the project still failed.

Sources

  1. primaryGartner, AI Projects in I&O Stall Ahead of Meaningful ROI Returns (press release, 2026) — (newsroom returned 403 to automated access; the 57% figure and Melanie Freeze's "too much, too fast" gloss are verified via the secondary outlets below, not against the primary page)
  2. primaryBarry M. Staw, "Knee-deep in the Big Muddy: A Study of Escalating Commitment to a Chosen Course of Action," Organizational Behavior and Human Performance 16:27–44 (1976).
  3. primaryMark Keil, "Pulling the Plug: Software Project Management and the Problem of Project Escalation," MIS Quarterly 19(4):421–447 (1995); extended in Keil et al., MIS Quarterly 24(4) (2000) — escalation carried from the lab into IT projects.
  4. primaryDaniel Kahneman & Amos Tversky, "Intuitive Prediction: Biases and Corrective Procedures," TIMS Studies in Management Science 12:313–327 (1979) — the planning fallacy and the inside/outside view.
  5. primaryDaniel Kahneman, Thinking, Fast and Slow (2011), ch. 23 "The Outside View" — popularisation of the inside/outside view and reference class forecasting.
  6. primarySim B. Sitkin, Kelly E. See, C. Chet Miller, Michael W. Lawless & Andrew M. Carton, "The Paradox of Stretch Goals: Organizations in Pursuit of the Seemingly Impossible," Academy of Management Review 36(3):544–566 (2011).
  7. primaryBent Flyvbjerg, "From Nobel Prize to Project Management: Getting Risks Right," Project Management Journal 37(3):5–15 (2006); see also "Curbing Optimism Bias and Strategic Misrepresentation in Planning" (2008) — reference class forecasting, strategic misrepresentation, uniqueness bias.
  8. primaryGeorge A. Akerlof, "The Market for 'Lemons': Quality Uncertainty and the Market Mechanism," Quarterly Journal of Economics 84(3):488–500 (1970) — information asymmetry as an explanatory lens, not a measurement of AI projects.
  9. primaryMartin Steinert & Larry Leifer, "Scrutinizing Gartner's Hype Cycle Approach," PICMET 2010 Proceedings — the hype cycle lacks a solid empirical basis; few technologies traverse the full curve.
  10. secondaryThe Register, Only 28% of AI infrastructure projects fully pay off (2026) — (source for the 782 sample and the 57% / 38% figures; the exact phrase "too much, too fast" does not appear in its body — it paraphrases the analyst)
  11. secondaryCIO, AI often doesn't deliver ROI for IT departments either(source for the 783 sample and the quoted "many said… expected too much, too fast")
  12. secondaryWikipedia, Gartner hype cycle(orientation on Fenn's 1995 origin and the phases; for source facts go to Fenn and to Steinert & Leifer)