← Writing

AI · MBA

ROI doesn't live in the model

A stalled AI project has an obvious-looking fix: swap the model for a better one. It rarely works. Gartner's own research points somewhere else entirely — return tracks how well the system is integrated, governed, and fitted to the work, not the sophistication of the model. If that is right, most of the money spent chasing a smarter model is being spent on the wrong link in the chain.

This essay makes one argument, and it is narrow. The value from an AI system is a property of the whole chain around the model, not of the model in isolation. A chain's throughput is set by its weakest link, and the weakest link is almost never the model. Drop a better model into an unchanged process and you have upgraded a part that was not the bottleneck.

The number does not move, and people are surprised that it does not.

The rule from the factory floor

The cleanest way to see this is a model built for a physical factory, decades before anyone shipped a language model into production. Eliyahu Goldratt and Jeff Cox laid it out in The Goal (North River Press, 1984): the Theory of Constraints. A system's throughput is set by a single constraint — the bottleneck — and improving anything that is not the bottleneck does nothing for output.

Goldratt's own line is worth keeping literal. "An hour lost at a bottleneck is an hour lost for the entire system. An hour saved at a non-bottleneck is a mirage." The mirage is the trap. Speed up a station that was never the limiting one and the local metric improves, the team feels productive, and the plant ships exactly as much as before.

Picture the plant he was describing. Machines feed one another in sequence, and one of them is slower than the rest. Speed up any machine except the slow one, and inventory just piles up in front of the bottleneck, which still releases work at its own unchanged pace. The plant's output is the bottleneck's output, full stop. Every improvement not aimed at it is motion without gain, busywork that shows on a local dashboard and nowhere on the shipping dock.

Now map it onto an AI system. The chain runs roughly: source data, access and permissions, integration into the systems where work happens, the model, its output, a path to action, and the people who have to adopt it. The model is one station on that line. If the constraint sits at integration or data or adoption, a stronger model is an hour saved at a non-bottleneck. It is real work, and it is a mirage.

There is a first-principles version of the same idea, and it comes from computing. Gene Amdahl gave it in 1967: the maximum speedup from improving one component is capped by the fraction of the system you did not improve. Even an infinite improvement to one part leaves the untouched fraction as a hard ceiling. The frame is from performance engineering, not from AI, so treat it as an analogy rather than a measurement.

The analogy is exact where it counts. Suppose the model accounts for a third of the value an AI system delivers. The figure is illustrative, not measured, and nobody has a credible number for it. Then a flawless model, perfect and free, lifts you by at most that third. The other two-thirds — integration, data, adoption — set the ceiling, and you cannot buy your way past them with weights.

The chain is worse than additive. Treat each stage as a yield. Say integration passes only 60% of the model's good outputs into a usable action, and people then act on only half of those. A perfect model now delivers 0.6 × 0.5 — thirty cents on the dollar of its own quality. The fractions are illustrative, and the exact numbers do not matter. The shape matters: non-model stages multiply, so two mediocre links downstream can swallow almost all of a great model's value. Upgrading the model raises the first term in a product whose later terms sit near zero. You sharpened the one number that was never the problem.

A third frame widens the lens from the chain to the whole system. Donella Meadows, in Thinking in Systems: A Primer (Chelsea Green, 2008), argues that a system's behaviour comes from its structure and its relationships, not from any single element. Swapping one component is, in her language, a low-leverage intervention. The high-leverage places to intervene are the flows of information, the rules, and the goals — the structure the model sits inside.

Meadows is careful, and honesty requires the caveat she insisted on. She presents her leverage points as a heuristic, and warns explicitly against treating them as a formula or a validated ranking. So take the direction, not a ladder to climb. The direction is enough: the model is a component, and components are where the leverage is lowest.

Why a better model doesn't move the number

Set the models aside and reason it out. Where, concretely, does an AI system create value? Not at the moment the model produces a good answer. At the moment that answer reaches a system, triggers an action, and someone trusts it enough to let it run. Everything before that is a demo.

That is the last-mile problem, borrowed from logistics and adapted to machine learning: the distance between a working prototype and a system that delivers value in production. The last mile is the most expensive stretch and the one where most projects break. MLOps exists as a discipline precisely to close it: integration, monitoring, retraining, and the path from output to action. The value is manufactured in that mile, which sits entirely downstream of the model.

So the chain has a structural feature that decides everything. Value accrues at the end, and the constraint usually sits before the model gets its say. ROI is a function of the whole pipeline, and a function dominated by its weakest term. The model is rarely that term.

This is not only a deduction. The survey evidence points the same way, from more than one direction.

Start with the anchor. Gartner surveyed 782 infrastructure and operations leaders across November and December 2025 and published on 7 April 2026. Melanie Freeze, a Director of Research at Gartner, framed the headline result directly: "ROI from AI is not driven by the sophistication of the model, but by how well the technology is integrated, governed, and aligned with real operational needs." One honesty note. Gartner's newsroom blocks automated access, so this wording comes from outlets that quoted the release — CIO and The Register among them — not from the primary page.

Read the failure side of that same survey and the pattern holds. Of the leaders reporting a failure, 57% blamed "expected too much, too fast" — an expectation problem, not a capability one. Roughly 38% named weak data quality and availability. On adoption, integration difficulty (48%) and lack of budget (50%) were the top barriers. Not one leading cause is "the model was not good enough."

The success side agrees by omission. Among leaders with a working use case, the credited factors were wiring AI into existing workflows and systems, full backing from business leadership, and cross-functional collaboration. No top success factor concerns the class, size, or vintage of the model. The people closest to the work locate their wins and losses in the process, and the model is not on either list.

A second, independent population reaches the same verdict. MIT's NANDA initiative, in The GenAI Divide: State of AI in Business 2025, names the main barrier a "learning gap" — the inability of firms to fold models into their own workflows, structures, and culture. As reported, the report is blunt: workflow integration, not model quality, is the real barrier. Different unit, different sample, same finding. (The often-quoted "95% of pilots" line from that report is a separate argument about denominators, dismantled in the 28% teardown.)

The model is the part everyone can see. That is exactly why it gets the blame it has not earned.

There is an incentive layer under the reflex, too. The model is the thing with a vendor, a price list, and a sales team, so it is the thing with someone paid to sell you the upgrade. Integration, data hygiene, and change management have no equivalent salesperson. Nobody cold-calls you about your permissions model or your onboarding flow. So the market pushes attention toward the one link in the chain that has a storefront, and away from the links that usually bind. The reflex is not only cognitive. Someone is paid to sustain it.

There is an economic foundation under all of this, and it predates the AI debate. Erik Brynjolfsson and Lorin Hitt spent the 2000s on the same point about information technology. Its value comes mostly from complementary organisational capital: redesigned processes, changed work practices, retraining. That capital is often worth far more than the technology itself, and much harder to measure. Their later work with Chad Syverson named the consequence: a productivity J-curve, where a general-purpose technology first depresses measured productivity while the complementary investments are built, and only later raises it.

Two cautions keep that honest. The data is about computers and IT from the 1990s and 2000s, so its transfer to language models is an analogy, strong but still an analogy. And "general-purpose technology" here is an economic category, not a product — do not read GPT as the model. With those flagged, the lesson is durable. The return on a general-purpose tool is a return on the reorganisation around it, and "no ROI yet" is sometimes the bottom of the J rather than a failure. Swapping the model does not dig you out of that trough.

Reading the constraint before you spend

None of this tells you where your constraint is. It tells you where to look first, and it warns you off the reflex. The reflex — buy a better model — is a bet that the model is the bottleneck, made without checking. Goldratt's whole method is the checking: find the constraint before you spend a cent improving anything.

Here is a diagnostic to find it. It is my own synthesis, not a Gartner finding or a Goldratt theorem, and it should be read as a framework rather than a measurement. Each question separates a model problem from a process problem before the budget moves.

SymptomLikely constraintRight lever (not a model upgrade)
The model clears a hand-picked gold set offline, yet the live deployment underdeliversDownstream: integration, data, or adoptionWire the output into a system with a real action path
The failure repeats after you swap in a stronger model on the same use caseNot the modelStop upgrading; instrument and fix the process
The blocker is input data, access, latency, or format — not answer qualityData and integrationFix the pipeline, not the weights
The output lands in a report nobody executesLast mile and adoptionBuild the action path and the change management
No public evidence that any model of this class does the task at allThe modelThen, and only then, a better model is the right lever

The first four questions point away from the model far more often than toward it. That is the whole claim, restated as a procedure. If the model already performs on a clean, curated set and the system still fails, the model is not your problem, and a better one will fail the same way at the same station.

The order of investment follows from the same logic. Spend first on integration and the path to action. Then on data quality and availability. Then on adoption and the change management that makes people trust the output. Then on the governance of the budget itself — funding that reaches across business units instead of fragmenting inside them. The next model upgrade comes after all of that, unless the diagnostic's last row is the one that lit up.

Governance earns its place on that list for a reason that is easy to miss. Freeze points at how AI actually gets funded: "Today, many AI initiatives are still funded by individual business units. However, as AI infrastructure spending continues to rise, CEOs and CFOs need to play a more active role." Per-unit funding fragments the very things the whole chain shares: common data, common integration, common standards. The weak link there is an org chart, not an architecture, and no model upgrade touches it.

That ordering is not only my inference. Gartner published a second study on 16 April 2026, surveying 353 data-and-analytics and AI leaders. That is a different sample and a different construct from the 782 above, so the two should not be summed or conflated. Its finding: organisations with successful AI initiatives invest up to four times more, as a share of revenue, in data quality, governance, AI-ready people, and change management than those that fail. "Up to four times" is a ceiling, not an average. And only 39% of technology leaders were confident their current AI spend would improve the financials at all.

Read those two Gartner surveys together and the shape is clear. The winners are not the ones who bought the smartest model. They are the ones who spent, heavily, on everything the smart model needs to be useful.

Spending on the model is visible and fast. Spending on the process is slow and looks like overhead. The reflex favours the visible one, which is exactly why it is usually wrong.

A stalled project, walked through the chain

Abstract levers are cheap, so here is the diagnostic run once, concretely. A mid-size company builds an assistant to triage inbound support tickets. It reads each ticket, classifies it, drafts a reply, and routes it to the right queue. In the demo it is excellent. Six months later, deflection is flat and the team is quietly back to triaging by hand. The instinct in the room is to blame the model and ask for the next one.

Run the diagnostic instead of the purchase order. First question: does the model clear a curated gold set offline? Someone pulls two hundred historical tickets and runs them through, and the classifications come back 94% correct. The model can do the task. So the constraint is downstream, and a better model would be an hour saved at a non-bottleneck.

Now find the actual station. The tickets live in a helpdesk tool the assistant cannot write back into, so a human copies every draft by hand — and stops bothering by week three. That is an integration constraint, not a model one. The output has no action path, so it lands in a queue nobody works. Adoption dies for a reason that has nothing to do with answer quality.

The fix costs an engineer a fortnight, not a model licence. Wire the assistant's output into the helpdesk API so routing happens automatically. Log every action for the audit trail governance will ask for. Let the humans approve rather than retype. Deflection moves the following month, and the model that was "not good enough" is the same model, untouched.

Swap the constraint and the lesson holds. Suppose the assistant instead classified badly in production despite a clean demo. You check the input and find the live tickets arrive with attachments stripped and customer history missing — the model sees less than it saw in testing. That is a data-availability constraint. A stronger model reasons better over the same impoverished input and still misses, because you cannot reason your way to information that was never in the prompt.

The pattern generalises past this one case. When a working demo fails in production, the break is almost always at the seam between the model and the system, not inside the model. The demo runs in a clean room. Production runs in your real stack, with its permissions, its latency, its formats, and its people — and that is where throughput is actually decided.

When the model really is the constraint

A rule sold as universal is being sold dishonestly. Three limits matter enough that ignoring any one of them turns this argument against you.

One: sometimes the model genuinely is the bottleneck. If a task exceeds what the current generation can do, the model is the binding constraint. Goldratt's own logic then says to strengthen that link. Improving the actual constraint is the one move that raises throughput. Capability jumps are real. Reasoning models have unlocked tasks that were previously unsolvable; on ARC-AGI-2, a benchmark built to resist surface pattern-matching, scores have moved sharply upward in a single generation. Treat that as evidence a threshold exists, not as a conversion rate from model quality to ROI — capability benchmarks are sometimes marketing. The thesis was never "the model doesn't matter." It is "identify the real constraint," and occasionally the real constraint is the model.

How do you know you are in that row rather than fooling yourself? There is a clean test. Look for public evidence that some model of the current class already does your task — a paper, a benchmark, a shipped product. If it exists, the ceiling is not the model, and your work is downstream. If nothing of the kind exists anywhere, you may genuinely be early, and a stronger model is the honest lever.

Two: you can over-invest in process on a foundation that cannot hold it. Building integrations, data pipelines, and change programmes around a model that simply cannot do the task is the same error in reverse. You polish a non-bottleneck until the process is thicker and the throughput is still zero. It is the symmetric mistake to the endless model upgrade, and it wastes more money because process work is slower to unwind. The stop condition is a single test: before you build the scaffolding, prove that some model of this class can do the task on your gold set. If none can, you are not looking at a process problem yet.

Three: "it's the process, not the model" is a comfortable alibi. The same sentence that correctly redirects a team away from a pointless upgrade also excuses a department that does not want to change its process. If it is "not a model question," then it is "not our fault," and one can skip both the model evaluation and the workflow redesign. The thesis can justify inaction in either direction. You can watch it happen. One team whose assistant fails in production says "Gartner is right, it's a process problem," and uses the line to avoid ever testing whether a newer model would clear the task. Another team, handed the same sentence, keeps the broken workflow and blames the model forever. Same words, opposite excuse, identical result: nothing changes. The antidote is the diagnostic above, which refuses the general answer and forces you to name a specific constraint — this data feed, that missing action path, this team's refusal to adopt.

Hold the three limits together and a discipline falls out. The model is the constraint when no model of its class can do the task. The process is the constraint the rest of the time. And the way to tell them apart is a test on a gold set, not a preference for whichever answer lets you avoid work. The thesis here is a default, not a law. Defaults are for the common case, and the common case is a process problem — but the whole job is to check, because the expensive mistakes live in the exceptions.

One more limit belongs to the evidence itself, and it is the same discipline applied to the 28%. The Gartner anchor is a self-assessment survey of one corner of IT. The unit is the use case, success is declared rather than audited, and the population is I&O leaders, not the whole enterprise. Freeze's "not driven by the sophistication of the model" is the data author's interpretation of that survey, reported at second hand. It is a considered position from Gartner, not a controlled experiment that pitted model against process and measured the winner. It converges with MIT's independent finding and with the economics, which is worth something. It is not proof, and it should not be quoted as one.

Closing

A better model is the right answer to exactly one question: is the model the constraint? Ask it honestly, on a gold set, before the purchase order — and most of the time the answer is no. That "no" is not a dead end. It is a map to where the return has been sitting the whole time, one station downstream, unglamorous and unbought.

Sources

  1. primaryEliyahu M. Goldratt & Jeff Cox, The Goal (North River Press, 1984) — the Theory of Constraints; throughput is set by the bottleneck, and improving a non-bottleneck is a mirage.
  2. primaryGene M. Amdahl, "Validity of the Single Processor Approach to Achieving Large-Scale Computing Capabilities," AFIPS Conference Proceedings (1967) — the un-improved fraction caps the achievable speedup.
  3. primaryDonella H. Meadows, Thinking in Systems: A Primer (Chelsea Green, 2008), and "Leverage Points: Places to Intervene in a System" (1999) — behaviour follows structure; leverage points offered as heuristic, not formula.
  4. primaryErik Brynjolfsson & Lorin Hitt, "Beyond Computation: Information Technology, Organizational Transformation and Business Performance," Journal of Economic Perspectives 14(4):23–48 (2000) — IT value comes largely from complementary organisational capital.
  5. primaryErik Brynjolfsson, Daniel Rock & Chad Syverson, "The Productivity J-Curve: How Intangibles Complement General Purpose Technologies," NBER Working Paper 25148 (2018) — measured productivity dips before it rises while complementary intangibles are built.
  6. primaryGartner, AI Projects in I&O Stall Ahead of Meaningful ROI Returns (press release, 7 April 2026) — 782 I&O leaders; 28% of use cases fully succeed; the Freeze framing on integration, governance, and fit. (newsroom returned 403 to automated access; quotes and figures verified via the secondary outlets below)
  7. primaryGartner, Organizations with Successful AI Initiatives Invest Up to Four Times More in Data and Analytics Foundations (press release, 16 April 2026) — 353 D&A/AI leaders; "up to four times" more in foundations; 39% confident of financial upside.
  8. primaryMIT NANDA, The GenAI Divide: State of AI in Business 2025 — the "learning gap"; workflow integration, not model quality, as the real barrier.
  9. secondaryCIO, AI often doesn't deliver ROI for IT departments either — verbatim Freeze quotes on model sophistication and per-business-unit funding.
  10. secondaryThe Register, Only 28% of AI infrastructure projects fully pay off (7 April 2026) — sample, analyst, and headline figures.
  11. secondaryBuilt In, Artificial Intelligence Has a "Last Mile" Problem — framing of MLOps and the last-mile gap between prototype and production.
  12. secondaryRCR Wireless, 7 takeaways from the GenAI Divide (MIT NANDA) — secondary framing of the "learning gap" and workflow-integration finding.