AI · MBA
Mainframe and the boundary of GenAI competence
Gartner forecasts that more than 70% of mainframe exit projects started in 2026 will fail to produce their intended benefits. The blame it assigns is an overestimation of what generative AI tooling can do. Read as a headline, that is one more failure statistic in a year full of them. Read as an experiment, it is more useful: a clean test of where a model's competence ends and someone else's work begins. Legacy code is close to the worst case you could design for this tool, and the reasons are specific enough to check against the pitch on your desk.
Language models can program. The limit is that the distribution they were trained on and the distribution they are asked to work on have almost nothing in common. And the expensive half of a migration was never the half a model touches.
What the forecast actually says
On 18 June 2026 Gartner published a press release with a sentence that travelled fast. More than 70% of mainframe exit projects initiated in 2026 will fail to produce the intended benefits, due to an overestimation of generative AI tooling capabilities. The named analyst is Alessandro Galimberti, VP Analyst. Gartner's newsroom returns a 403 to automated access, so I am working from three trade outlets that quote the release verbatim: IT-Online, DIGIT and TechEdgeAI. Treat the wording below as reported rather than read at source.
Three qualifiers sit inside that sentence. All three get dropped in retelling.
It is a forecast, not a measurement. There is no published sample, no method, no sampling frame. It is an analyst position built on client contact. That is a legitimate genre, and it is a different genre from a study.
The denominator is narrow. Projects initiated in 2026 — not migrations in general, not the installed base of past attempts, not a rate observed after the fact.
The threshold is soft. "Fail to produce the intended benefits" is not "collapse" and not "get cancelled". A project can land, run, and still miss its business case by a distance nobody wants to write down.
A second number from the same release gets welded onto the first, and should not be. By 2030, Gartner forecasts, 75% of vendors operating in the mainframe exit market will pivot their business models or cease operations. That is a prediction about a supply-side shakeout. Different denominator, different phenomenon. Adding it to the 70% produces a sentence with no referent.
Gartner's own recommendation is narrower and more interesting than "GenAI does not work on legacy". As reported, Galimberti describes a widening gap between the marketing promise of GenAI and its real-world ability to transform and migrate complex legacy code. The recommendation that follows is the half that gets cut: for many mainframe customers, GenAI can be used more effectively to enable modernisation in place than to accelerate migration off the platform. Same tool, different job. Reporting only the first half misquotes the source.
That is a claim about job selection, not about the strength of the instrument.
A circle drawn around a tool
Buffett's frame comes from the 1996 Berkshire Hathaway letter to shareholders: "You only have to be able to evaluate companies within your circle of competence. The size of that circle is not very important; knowing its boundaries, however, is vital." Munger developed it alongside him for decades. It was written about investing, and about the investor.
Move it one step. Apply it to the instrument, not the person.
The useful question about a language model is where its edge sits. That edge is set by the distribution of its training data measured against the distribution of the work. A model trained overwhelmingly on modern, open, well-tested code is inside its circle when it writes modern, open, well-tested code. A mainframe estate sits outside that circle on nearly every axis at once: language, runtime, idiom, era, and the absence of anything to check the answer against.
Three more frames name the specific parts of the problem. Each has an author.
Chesterton's fence. G.K. Chesterton, The Thing (1929), in the chapter "The Drift from Domesticity". The reformer who finds a fence across a road and cannot see why it is there has not earned the right to remove it. He must first go and find out. Every odd line in forty-year-old code is that fence: a workaround for a driver bug, a regulatory requirement from 1987, an ordering that a downstream batch job silently depends on. The fence is still standing. Nobody remembers the road.
Hyrum's Law. Observed by Hyrum Wright at Google around 2011, then named and published in Winters, Manshreck and Wright, Software Engineering at Google (O'Reilly, 2020). With a sufficient number of users of an API, it does not matter what you promise in the contract: all observable behaviours of your system will be depended on by somebody. The consequence for migration is severe and rarely priced. A rewrite that matches the written specification perfectly can still break users. The de facto contract is the observed behaviour: date format, record ordering, response latency, the exact text of an error message that some downstream script parses.
Legacy code is code without tests. Michael Feathers, Working Effectively with Legacy Code (Prentice Hall, 2004). The definition is deliberately blunt, and his method follows from it. Before you change anything, wrap it in characterisation tests that pin down actual behaviour rather than intended behaviour. That names the ground-truth problem exactly. With no tests, there is nothing to compare the model's output against.
Underneath all of it sits a hard ceiling. Rice's theorem (H.G. Rice, Transactions of the AMS 74(2), 1953) says that every non-trivial semantic property of a Turing-recognisable language is undecidable. Functional equivalence of two programs is such a property. In plain terms: "prove that the new system does the same thing as the old one" is not, in the general case, a task you can close with a proof. What remains is differential testing on real traffic: a parallel run. That is calendar time and operating cost, not a model licence.
Put the four together and the shape of the problem appears. The specification is the code. The contract is the observed behaviour. The oracle does not exist. And equivalence cannot be settled by argument.
Why the syntax is the small part
A migration looks like a translation problem. COBOL in, Java out. That framing is precisely what makes GenAI look like the obvious accelerator, because translation is the thing these models visibly do well.
Decompose the work instead. A running mainframe application is at least four things stacked on each other.
- Text. Syntax and control flow. Written down, machine-readable, complete.
- Behaviour nobody wrote down. Timing windows, batch ordering, retry semantics, the error path that evolved incident by incident, the reason a field is padded a particular way.
- Environmental coupling. CICS, JCL, VSAM, DB2, the scheduler, the operator runbook, the job that must finish before the other one starts.
- A behavioural contract with everything downstream, in Hyrum's sense, including consumers nobody has a list of.
Only layer one is in the repository. Layers two, three and four exist as consequences, not as artefacts. A model reads layer one flawlessly and can do nothing but infer the rest.
After forty years of amendments, the specification is the code. That is not a metaphor. Whatever document once described intent has drifted from what runs, because every production incident added a branch and every regulator added a rule. The system's real behaviour is the accumulated residue of those changes, and it lives in exactly one place.
Two structural reasons make this the hard case for a model, and both are checkable.
The first is distribution. COBOL is, in machine-learning terms, a low-resource language with distinct logic patterns. That is not my framing; it is stated directly in the 2026 SEDCoT paper, which names the property as the reason general-purpose models show suboptimal correctness on COBOL translation. Competence degrades systematically off-distribution. Here the work is off-distribution in every direction simultaneously.
The second is asymmetry. Generating code that looks right is cheap. Proving that it behaves identically is expensive, and Rice's theorem says the general case has no shortcut. So the part of the work that AI compresses is not the part that costs. You can buy a translation in an afternoon. You cannot buy an oracle.
This is the same shape as the argument in ROI doesn't live in the model. The model is one station on a chain, and the chain here runs: establish ground truth, extract rules, generate code, build tests, run in parallel, cut over, support afterwards. GenAI can eat one station entirely and leave the total cost roughly where it was, because the constraint sits at the stations that never appear in a demo.
What the benchmarks measure, and what they don't
The evidence on this boundary is unusually good for a question this young, and it points in a consistent direction. It also gets quoted badly, so the unit of analysis matters as much as the number.
Start with the easy case. Pan and colleagues published "Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code" at ICSE 2024, out of IBM Research and UIUC. They tested 1,700 code samples across five high-resource languages, using three benchmarks, two real projects and over 10,000 tests. The share of correct translations ranged from 2.1% to 47.3% depending on the model. The team hand-labelled 1,748 bugs into 15 categories of translation error. This is C and Java and Python, with abundant training data and code that already has tests. It is the friendliest possible version of the task.
Now the unfriendly one. COBOL-Coder (Dau and colleagues, arXiv:2604.03986, 5 April 2026) evaluated models on COBOLEval, a COBOL adaptation of HumanEval. GPT-4o reached 41.8% compilability and 16.4 pass@1. Open-source baselines (CodeGemma, CodeLlama, StarCoder2) mostly failed to produce a program that compiled at all. A domain-adapted model reached 73.95% compilability and 49.33 pass@1, and 34.93 pass@1 on Java-to-COBOL where general models scored near zero. The methodological caveat is the important part. These are standalone HumanEval-style functions, not production modules wired into CICS, JCL, VSAM and DB2. If the number is 16.4% on toy functions, it is a lower bound on difficulty for a real estate, not an upper one.
The sharpest result is AgentModernize (Ahmed and Galib, arXiv:2605.17535, 17 May 2026). Naive baselines (single-prompt translation and chain-of-thought) preserved system behaviour in 0.0% of cases, on every scenario and every base model tested. A multi-agent framework built around an explicit intermediate artefact, a Behavioral Specification Graph, extracted 91.2% of gold-standard rules with a best behavioural error of 8.1%. The authors' own conclusion is the one worth carrying: the bottleneck is code generation, not knowledge extraction. Reading is cheap. Proving equivalence is not.
And the direction of improvement is specific. SEDCoT (Entin and colleagues, arXiv:2607.04092, 5 July 2026) reports at least a 12% gain over the prior state of the art. It came from symbolic execution, generated test suites and delta debugging to minimise counterexamples. The thing that moved the boundary was machinery for proving, not machinery for writing.
Meanwhile, the tasks that consist of reading rather than rewriting come out well. Diggs and colleagues (LLM4Code 2025 at ICSE, arXiv:2411.14971) found LLM-generated documentation for MUMPS and IBM Assembly Language Code to be generally hallucination-free, complete, readable and useful against ground truth. Assembly was the harder of the two. Their own caveat deserves equal billing: no automatic metric correlated strongly with comment quality, so there is no cheap way to measure whether the documentation is good other than a human reading it. COBRAIN (EASE 2025) extracted business rules from COBOL at precision 1.0 and recall 0.746 against the rule-based COBREX tool, with F1 0.73 against ground truth versus 0.59. In a comprehension study with 28 participants, over 80% preferred the LLM output.
One more result belongs here, from a different setting. METR ran a randomised controlled trial with 16 experienced open-source developers across 246 tasks on their own mature repositories (arXiv:2507.09089, July 2025). With early-2025 AI tools they were 19% slower, while estimating afterwards that the tools had made them 20% faster. This is not a law about AI and productivity, and METR does not present it as one. It is a snapshot of one context, with a small sample and tools that have since changed. It is the closest available analogue to deep context that lives outside the code, and it carries one transferable lesson. Self-reported acceleration is not evidence.
| Study | Unit of analysis | Result | What it does not show |
|---|---|---|---|
| Pan et al., ICSE 2024 | Code sample, high-resource languages | 2.1–47.3% correct translations; 1,748 bugs in 15 categories | Anything about COBOL, or about code without tests |
| COBOL-Coder, 2026 | Standalone HumanEval-style COBOL function | GPT-4o 41.8% compilable, 16.4 pass@1; domain-adapted 73.95% / 49.33 | Production modules with CICS, JCL, VSAM, DB2 |
| AgentModernize, 2026 | Behaviour preservation across scenarios | Naive baselines 0.0%; multi-agent 91.2% of gold rules, 8.1% best behavioural error | That the extracted rules survive contact with a real cutover |
| SEDCoT, 2026 | COBOL translation correctness | ≥12% over prior state of the art, from symbolic execution and delta debugging | That scale alone produces the same gain |
| Diggs et al., 2025 | Generated documentation for MUMPS / assembly | Generally hallucination-free, complete, useful | A cheap automatic quality metric — none correlated |
| COBRAIN, EASE 2025 | Business rules extracted from COBOL | Precision 1.0, recall 0.746; F1 0.73 vs 0.59; 28 participants, >80% preferred | That extraction equals migration |
| METR, 2025 | Task completion time, 16 devs, 246 tasks | 19% slower with AI; self-estimated 20% faster | A general law; sample is small and tools have moved |
The table splits in two. Where the task is to read, describe or extract, the results are good. Where the task is to rewrite and be trusted, the results are poor. They improve when someone bolts on a verification apparatus.
The estate you are holding
Before evaluating a pitch, it helps to know what is being pitched against. The volume figures everyone quotes are weak, and the weakness has to be stated every time they are used.
Two numbers circulate. Reuters, in 2017, reported roughly 220 billion lines of COBOL in use, 43% of banking systems built on it, and 95% of ATM card swipes touching it. That is a journalistic compilation of industry figures with no published counting method. The other is "over 800 billion lines in daily use", which comes from a 2022 Micro Focus survey of 1,104 respondents across 49 countries. That is a vendor measuring perception, not an audit of repositories. Both figures are orders of magnitude, not measurements, and should be labelled as such whenever they appear on a slide.
The audited numbers are better and worse. The US Government Accountability Office reported in July 2025 (GAO-25-107795) that CFO Act agencies identified 69 legacy systems. GAO singled out the 11 most critical: between 23 and 60 years old, costing around $754 million a year to maintain. Eight of the eleven run on outdated languages; Treasury's runs on COBOL and assembler. Seven operate with known security vulnerabilities. Four run on unsupported hardware or software. Only three of the eleven had a complete modernisation plan, and two had none at all.
The IRS Individual Master File is the canonical case. Over 60 years old, written in assembler and COBOL, designed for the IBM System/360. Its replacement programme, CADE, was abandoned in 2009. CADE-2 has still not replaced it and is not expected to before 2030. In March 2025 the IRS told GAO it had paused modernisation programmes while priorities were being revised. That is roughly twenty years of attempts, all of them before GenAI existed.
One commercial reference point for doing it properly. Commonwealth Bank of Australia replaced its core banking platform over five years, at a cost above one billion Australian dollars (around 750 million US dollars at 2017 exchange rates). Hold that number next to any proposal promising a migration in two quarters because the tooling has improved.
One narrative needs correcting, because it is the one most often used to create urgency. "All the COBOL programmers are retiring" is weaker evidence than its confidence suggests, and the best available data points the other way. The BMC Mainframe Survey 2025 is the twentieth edition, with over 1,100 respondents. It found the 18–49 age group rising from 53% in 2018 to roughly 80% in 2025, and Gen Z from 1% to 15%. In the same survey 97% held a positive view of the platform, 93% planned further investment, and only 3% were considering alternatives, down from 10%. Around 75% treated GenAI as a strategic initiative, but only about 30% would allow AI to execute tasks autonomously. That survey polls a mainframe vendor's own customers, so it has an obvious interest in the conclusion "the mainframe is alive". Keep both biases in frame.
The accurate sentence is narrower and harder. What thins out is the number of people who remember why a particular line is there. That is a different problem from a shortage of COBOL programmers, and you cannot hire your way out of it.
There is a figure I would like to put here and cannot. No independent public number exists for the share of mainframe estates carrying meaningful regression test coverage. That is the single most decision-relevant statistic in the whole argument, and nobody has counted it.
Reading a modernisation pitch
The frames above collapse into a short interrogation. None of these questions requires technical depth to ask, and the quality of the answers separates a plan from a brochure.
- Where will the ground truth come from? A parallel run on production traffic? Recorded inputs and outputs? Or a specification nobody has seen?
- What is the coverage of the tests that already exist — not the ones you intend to generate? If it is zero, who signs the statement that outputs are equivalent?
- How will you detect differences that are not functional bugs? Timing, batch ordering, throughput, message formats. This is the Hyrum's Law question, and it is the one that gets answered worst.
- Who on our side knows this system? How many people, how many years, and when do they leave?
- What is the unit in your benchmarks? A HumanEval-style function, or a production module with CICS, JCL and VSAM? The gap between those two is where most of the marketing lives.
- What is the measure of success in the contract? Lines translated, or transactions passing a parallel run without divergence? Only one of those is a business outcome.
- Big bang or incremental? With a running original and a path back, or a single cutover weekend?
- Which step of the chain does AI actually shorten, and what share of today's budget does that step represent?
Question one carries most of the weight, and there is a cautionary case for what happens when it goes unanswered. Michigan's MiDAS system auto-adjudicated 22,427 unemployment fraud cases between October 2013 and August 2015. The Michigan Office of the Auditor General reported in February 2016 a 93% error rate for appeals filed between October 2013 and June 2015, and 20,965 decisions were subsequently reversed. Tens of thousands of people were assessed quadruple penalties, with wages and tax refunds garnished. MiDAS was rule-based, not machine learning, and calling it an AI failure would be wrong. The pattern is what transfers: automation deployed at scale where the ground truth had never been verified.
The other half of the answer is where GenAI genuinely earns its place in a legacy programme. Gartner's own recommendation, modernisation in place rather than migration acceleration, points at the same set of tasks the research supports.
| Task | Is there an oracle? | Cost of a wrong answer | Verdict |
|---|---|---|---|
| Documenting undocumented modules | Human review only; no automatic metric correlates | Low — a bad comment gets discarded | Use it |
| Extracting business rules into a reviewable artefact | Comparable against rule-based tools | Low, if the artefact is reviewed before use | Use it, then review it |
| Generating characterisation tests against the running system | The running system is the oracle | Low — a wrong test fails visibly | Use it; highest value per hour |
| Impact analysis and dependency mapping | Partially checkable against the code | Medium | Use it, verify the edges |
| Wholesale translation of production modules | None, unless a parallel run exists | High — silent behavioural drift | Do not buy this as the plan |
| Migrating batch orchestration and timing | Nothing in the code to check against | High, and discovered late | People, plus a parallel run |
The pattern in the right-hand column is the same one Feathers described in 2004. The tool is strong wherever the running system can act as its own judge, and weak wherever a human has to certify equivalence from nothing.
Limits, and the counterargument that matters
Three objections have real force, and a reader with a mainframe estate should hold all three.
One: the boundary moves, and it is moving now. Between April and July 2026 at least three published papers pushed it: COBOL-Coder, AgentModernize, SEDCoT. The Gartner forecast is about projects initiated in 2026, not about a permanent property of the technology. Reading it as a verdict on GenAI is the misreading worth avoiding. There is a more useful version of the claim, and it is falsifiable. The boundary appears to be moving through verification machinery (symbolic execution, generated test suites, delta debugging, explicit intermediate artefacts), not through model scale. If the next large gain on COBOL translation comes from a bigger general-purpose model with no verification apparatus attached, the mechanism argued here is wrong.
Two: most of these failures will have ordinary causes. Flyvbjerg and Budzier, writing in Harvard Business Review in September 2011, studied 1,471 IT projects. The average cost overrun was 27%. But one in six was a black swan, with an average cost overrun of 200% and a schedule overrun near 70%. The distribution is fat-tailed, so the mean tells you nothing about the risk. The TSB migration of April 2018 makes the point concrete. Around five million customers were moved off Lloyds systems in a single weekend. Branch, telephone, online and mobile banking became unavailable for much of a 5.2 million customer base, and fraud attacks ran at 70 times normal levels at the peak. The regulatory fines came to £48.65m (FCA £29.75m plus PRA £18.9m, after a 30% settlement discount; £69.5m without it), alongside £32.7m in redress and a total reported cost of roughly £366m. The independent review by Slaughter and May, published in November 2019, pointed at the scale and pace of the migration and at the failure to assess the main IT supplier's capability. Scope, pace, supplier governance, and a big-bang cutover. No AI involved anywhere. A meaningful share of the projects Gartner is forecasting about would miss their benefits with or without a model in the room. Attributing everything to GenAI hides the causes that were already there.
Three: a competence boundary makes an excellent alibi. "AI cannot do this" does not imply "so we leave it alone", and the slide from the first to the second is the most expensive move available. GAO's eleven systems cost around $754 million a year to run, seven of them with known vulnerabilities and four on unsupported hardware or software. The IRS has been trying since before 2009 and paused again in March 2025. The cost of not moving is measurable, compounding, and larger than the cost of moving carefully. Modernisation in place is a real option. Incremental replacement with a parallel run and a path back is a real option. Another decade of paused programmes, justified by a correct observation about model limits, is not.
There is a fourth limit, and it applies to the anchor itself. The Gartner forecast has no published sample, no method and no sampling frame, and I read it through three secondary outlets because the source returns a 403. It converges with the research above, which is worth something. It is not a measurement, and it should not be quoted as one — including by essays that agree with it.
Closing
Put two questions to a modernisation pitch. Which step of the work does the tool remove? And who signs the statement that the new system does the same thing as the old one? If nobody in the room can name that person, you are not buying a migration. You are buying a translation, and hoping.
Sources
- primaryGartner, "Gartner Predicts More Than 70% of Mainframe Exit Projects Will Fail Due to Overestimation of Generative AI's Capabilities" (press release, 18 June 2026) — the 70% forecast, the 75% vendor-shakeout forecast, and the "modernisation in place" recommendation; analyst Alessandro Galimberti. (Newsroom returned 403 to automated access; wording verified through the secondary outlets listed below.)
- primaryWarren Buffett, letter to Berkshire Hathaway shareholders (1996) — circle of competence; "knowing its boundaries, however, is vital".
- primaryG.K. Chesterton, The Thing (1929), chapter "The Drift from Domesticity" — the fence across the road.
- primaryTitus Winters, Tom Manshreck & Hyrum Wright, Software Engineering at Google (O'Reilly, 2020) — Hyrum's Law, observed by Hyrum Wright: all observable behaviours of a system will be depended on by somebody.
- primaryMichael Feathers, Working Effectively with Legacy Code (Prentice Hall, 2004) — legacy code as code without tests; characterisation tests pin actual, not intended, behaviour.
- primaryH.G. Rice, "Classes of Recursively Enumerable Sets and Their Decision Problems," Transactions of the American Mathematical Society 74(2) (1953) — non-trivial semantic properties, including program equivalence, are undecidable.
- primaryRangeet Pan et al., "Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code," ICSE 2024 — 1,700 samples, five languages, 2.1–47.3% correct translations, 1,748 labelled bugs. (arXiv:2308.03109 · DOI 10.1145/3597503.3639226)
- primaryA.T.V. Dau et al., "COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation" (arXiv:2604.03986, 5 April 2026), and P. Entin et al., "SEDCoT: Enhancing LLM-Based COBOL Code Translation via Symbolic Execution and Delta Debugging" (arXiv:2607.04092, 5 July 2026) — COBOLEval results, COBOL as a low-resource language, and gains from verification machinery.
- primaryS.N. Ahmed & M. Galib, "AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs" (arXiv:2605.17535, 17 May 2026) — 0.0% behaviour preservation for naive baselines; the bottleneck is code generation, not knowledge extraction.
- primaryC. Diggs et al., "Leveraging LLMs for Legacy Code Modernization: Challenges and Opportunities for LLM-Generated Documentation," LLM4Code 2025 at ICSE (arXiv:2411.14971), and "LLM vs Rule-Based — The COBRAIN Tool and An Empirical Study on Extracting Business Rules from COBOL," EASE 2025 (DOI 10.1145/3756681.3756982) — documentation and rule extraction results, with the authors' caveat that no automatic metric correlates with quality.
- primaryMETR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (10 July 2025) — RCT, 16 developers, 246 tasks, 19% slower against a self-estimate of 20% faster. (arXiv:2507.09089)
- primaryU.S. GAO, "Information Technology: Agencies Need to Plan for Modernizing Critical Decades-Old Legacy Systems," GAO-25-107795 (17 July 2025), and "Information Technology: IRS Is Developing a New Modernization Framework," GAO-25-107611 — the 11 critical systems, $754m annual maintenance, and the IRS Individual Master File timeline.
- primaryTSB Bank / Slaughter and May, "Independent Review of TSB's 2018 migration to a new IT platform" (19 November 2019), and FCA, "TSB fined £48.65m for operational resilience failings" (20 December 2022) — scale, pace, supplier assessment, fines and redress.
- primaryBent Flyvbjerg & Alexander Budzier, "Why Your IT Project May Be Riskier Than You Think," Harvard Business Review (September 2011) — 1,471 projects, 27% average cost overrun, one in six with a 200% overrun.
- secondaryIT-Online, "Overestimating GenAI will see mainframe exit projects fail" (23 June 2026); DIGIT, "Orgs are overestimating GenAI when it comes to mainframe exits" (22 June 2026); TechEdgeAI, "Mainframe Exit Projects Face GenAI Reality Check" (18 June 2026) — three outlets quoting the Gartner release verbatim; the basis for every Gartner quotation above.
- secondaryEstate-scale figures, all of them weak, and used above only as orders of magnitude: Reuters via CNBC, "Banks scramble to fix old systems as IT 'cowboys' ride into sunset" (11 April 2017) — the 220bn-lines / 43% / 95% figures and the Commonwealth Bank of Australia cost, a journalistic compilation with no published counting method; The Stack, "There's over 800 billion lines of COBOL in daily use" — the Micro Focus 2022 survey of 1,104 respondents across 49 countries, a vendor measuring perception; TechChannel, "Mainframers Are Getting Younger: BMC Survey" — BMC Mainframe Survey 2025, over 1,100 respondents, vendor-sponsored with a self-selecting sample of that vendor's own customers.
- secondaryGovTech, "Michigan Integrated Data Automated System Experiences 93 Percent Error Rate" (on the Michigan Office of the Auditor General report, February 2016), and Ford School STPP, "Cahoo v. SAS — MiDAS explainer" (2024) — 22,427 auto-adjudicated cases, 20,965 reversals; a rule-based system, not machine learning.
Mainframe i granica kompetencji GenAI
Gartner prognozuje, że ponad 70% projektów wyjścia z mainframe'a rozpoczętych w 2026 nie dowiezie zakładanych korzyści. Winą obarcza przecenianie tego, co potrafią narzędzia generatywnej AI. Czytana jak nagłówek, prognoza jest kolejną statystyką porażek w roku, który ich nie skąpi. Czytana jak eksperyment, daje więcej: czysty test na to, gdzie kończy się kompetencja modelu, a zaczyna cudza robota. Kod legacy jest bliski najgorszemu przypadkowi, jaki dałoby się dla tego narzędzia zaprojektować, a powody są dość konkretne, żeby sprawdzić je na ofercie leżącej na twoim biurku.
Modele językowe potrafią programować. Ograniczenie leży gdzie indziej: rozkład, na którym je wytrenowano, i rozkład pracy, którą mają wykonać, nie mają ze sobą prawie nic wspólnego. A droga połowa migracji nigdy nie była tą połową, której model dotyka.
Co dokładnie mówi prognoza
18 czerwca 2026 Gartner opublikował komunikat prasowy ze zdaniem, które poszło w świat szybko. Ponad 70% projektów wyjścia z mainframe'a rozpoczętych w 2026 nie dowiezie zakładanych korzyści, z powodu przeceniania możliwości narzędzi generatywnej AI. Wskazany analityk to Alessandro Galimberti, VP Analyst. Newsroom Gartnera zwraca 403 przy dostępie automatycznym, więc pracuję na trzech serwisach branżowych, które cytują komunikat dosłownie: IT-Online, DIGIT i TechEdgeAI. Traktuj poniższe brzmienie jako relacjonowane, nie czytane u źródła.
W tym zdaniu siedzą trzy zastrzeżenia. Wszystkie trzy giną przy opowiadaniu dalej.
To prognoza, nie pomiar. Nie ma opublikowanej próby, nie ma metody, nie ma operatu losowania. To stanowisko analityka zbudowane na kontakcie z klientami. Taki gatunek jest uprawniony, tylko nie należy go czytać jak badania.
Mianownik jest wąski. Projekty rozpoczęte w 2026 — nie migracje w ogóle, nie dorobek wcześniejszych podejść, nie wskaźnik zaobserwowany po fakcie.
Próg jest miękki. „Nie dowiezie zakładanych korzyści" to nie „zawali się" i nie „zostanie skasowany". Projekt może wylądować, działać i nadal rozminąć się z własnym uzasadnieniem biznesowym o dystans, którego nikt nie chce zapisać.
Do pierwszej liczby dospawuje się drugą z tego samego komunikatu i nie powinno się tego robić. Do 2030 roku, prognozuje Gartner, 75% dostawców działających na rynku wyjścia z mainframe'a zmieni model biznesowy albo zakończy działalność. To przewidywanie o czystce po stronie podaży. Inny mianownik, inne zjawisko. Dodanie go do 70% daje zdanie bez desygnatu.
Własna rekomendacja Gartnera jest węższa i ciekawsza niż „GenAI nie działa na legacy". Jak podano, Galimberti opisuje rosnącą lukę między marketingową obietnicą GenAI a jej realną zdolnością do przekształcania i migrowania złożonego kodu legacy. Rekomendacja, która z tego wynika, to ta połowa, którą się ucina: dla wielu klientów mainframe'owych GenAI da się użyć skuteczniej do modernizacji na miejscu niż do przyśpieszenia zejścia z platformy. To samo narzędzie, inne zadanie. Kto relacjonuje tylko pierwszą połowę, przekręca źródło.
Ta teza dotyczy doboru zadania, a nie siły instrumentu.
Krąg zakreślony wokół narzędzia
Rama Buffetta pochodzi z listu do akcjonariuszy Berkshire Hathaway z 1996 roku: „Musisz umieć ocenić spółki jedynie w obrębie swojego kręgu kompetencji. Rozmiar tego kręgu nie ma wielkiego znaczenia; znajomość jego granic jest natomiast niezbędna". Munger rozwijał ją razem z nim przez dekady. Napisano ją o inwestowaniu i o inwestorze.
Przesuń ją o krok. Przyłóż do instrumentu, nie do człowieka.
Użyteczne pytanie o model językowy brzmi: gdzie leży jego krawędź. Wyznacza ją rozkład danych treningowych zestawiony z rozkładem pracy. Model wytrenowany w przeważającej części na współczesnym, otwartym, otestowanym kodzie jest wewnątrz swojego kręgu, kiedy pisze współczesny, otwarty, otestowany kod. Park mainframe'owy leży poza tym kręgiem na niemal każdej osi naraz: język, środowisko uruchomieniowe, idiom, epoka i brak czegokolwiek, o co dałoby się oprzeć sprawdzenie odpowiedzi.
Trzy kolejne ramy nazywają konkretne części tego problemu. Każda ma autora.
Płot Chestertona. G.K. Chesterton, The Thing (1929), rozdział „The Drift from Domesticity". Reformator, który zastaje płot w poprzek drogi i nie widzi, po co tam stoi, nie zasłużył sobie na prawo do usunięcia go. Najpierw musi pójść i się dowiedzieć. Każda dziwna linijka w czterdziestoletnim kodzie jest takim płotem: obejście błędu sterownika, wymóg regulacyjny z 1987 roku, kolejność, na której po cichu opiera się jakieś zadanie wsadowe niżej w łańcuchu. Płot wciąż stoi. Drogi nikt już nie pamięta.
Prawo Hyruma. Zaobserwował je Hyrum Wright w Google około 2011 roku; nazwane i opublikowane zostało w książce Software Engineering at Google (O'Reilly, 2020), napisanej wspólnie z Titusem Wintersem i Tomem Manshreckiem. Przy dostatecznej liczbie użytkowników API nie ma znaczenia, co obiecujesz w kontrakcie: na każdym obserwowalnym zachowaniu twojego systemu ktoś się oprze. Konsekwencja dla migracji jest dotkliwa i rzadko wyceniana. Przepisanie, które idealnie zgadza się ze spisaną specyfikacją, wciąż może zepsuć robotę użytkownikom. Kontraktem de facto jest zachowanie obserwowane: format daty, kolejność rekordów, czas odpowiedzi, dokładna treść komunikatu o błędzie, który parsuje jakiś skrypt niżej w łańcuchu.
Kod legacy to kod bez testów. Michael Feathers, Working Effectively with Legacy Code (Prentice Hall, 2004). Definicja jest celowo brutalna, a metoda z niej wynika. Zanim cokolwiek zmienisz, obłóż to testami charakteryzującymi, które utrwalają zachowanie faktyczne, a nie zamierzone. Ta definicja nazywa problem prawdy odniesienia po imieniu. Bez testów nie ma do czego porównać tego, co model wypluje.
Pod tym wszystkim leży twardy sufit. Twierdzenie Rice'a (H.G. Rice, Transactions of the AMS 74(2), 1953) mówi, że każda nietrywialna własność semantyczna języka rozpoznawalnego przez maszynę Turinga jest nierozstrzygalna. Równoważność funkcjonalna dwóch programów jest taką własnością. Po ludzku: „udowodnij, że nowy system robi to samo co stary" nie jest, w przypadku ogólnym, zadaniem, które zamkniesz dowodem. Zostaje testowanie różnicowe na prawdziwym ruchu, czyli praca równoległa. To czas kalendarzowy i koszt operacyjny, nie licencja na model.
Złóż te cztery rzeczy razem, a wyłania się kształt problemu. Specyfikacją jest kod. Kontraktem jest zachowanie obserwowane. Wyroczni nie ma. A równoważności nie rozstrzygnie się argumentem.
Dlaczego składnia to mała część
Migracja wygląda na problem tłumaczenia. COBOL na wejściu, Java na wyjściu. Właśnie to ujęcie sprawia, że GenAI wygląda na oczywisty przyśpieszacz — tłumaczenie to rzecz, którą te modele robią dobrze i którą łatwo pokazać.
Rozłóż tę pracę na części. Działająca aplikacja mainframe'owa to co najmniej cztery rzeczy ułożone jedna na drugiej.
- Tekst. Składnia i przepływ sterowania. Spisany, czytelny maszynowo, kompletny.
- Zachowanie, którego nikt nie spisał. Okna czasowe, kolejność wsadów, semantyka ponowień, ścieżka błędu narosła incydent po incydencie, powód, dla którego jedno pole jest dopełniane akurat tak.
- Sprzężenie z otoczeniem. CICS, JCL, VSAM, DB2, harmonogram, runbook operatora, zadanie, które musi się skończyć, zanim ruszy następne.
- Kontrakt behawioralny ze wszystkim, co niżej w łańcuchu, w sensie Hyruma, razem z odbiorcami, których listy nikt nie ma.
W repozytorium leży tylko warstwa pierwsza. Pozostałe trzy istnieją jako konsekwencje, nie jako artefakty. Model czyta warstwę pierwszą bezbłędnie, a resztę może wyłącznie wywnioskować.
Po czterdziestu latach poprawek specyfikacją jest kod. To nie metafora. Dokument, który kiedyś opisywał zamiar, odjechał od tego, co działa, bo każdy incydent produkcyjny dokładał gałąź, a każdy regulator regułę. Prawdziwe zachowanie systemu to zbiorczy osad tych zmian i mieszka dokładnie w jednym miejscu.
Dwa strukturalne powody czynią z tego trudny przypadek dla modelu i oba da się sprawdzić.
Pierwszy to rozkład. COBOL jest, w języku uczenia maszynowego, językiem niskozasobowym o odrębnych wzorcach logiki. To nie moje ujęcie; mówi to wprost praca SEDCoT z 2026 roku, która wskazuje tę własność jako powód, dla którego modele ogólnego przeznaczenia wypadają na tłumaczeniu COBOL-a poniżej optimum. Kompetencja degraduje się systematycznie poza rozkładem. Tutaj praca leży poza rozkładem we wszystkich kierunkach naraz.
Drugi to asymetria. Wygenerowanie kodu, który wygląda poprawnie, jest tanie. Udowodnienie, że zachowuje się identycznie, jest drogie, a twierdzenie Rice'a mówi, że w przypadku ogólnym nie ma skrótu. Ta część pracy, którą AI skraca, nie jest więc tą, która kosztuje. Tłumaczenie kupisz w jedno popołudnie. Wyroczni nie kupisz.
To ten sam kształt co argument w ROI nie mieszka w modelu. Model jest jednym stanowiskiem w łańcuchu, a łańcuch biegnie tak: ustal prawdę odniesienia, wyciągnij reguły, wygeneruj kod, zbuduj testy, puść pracę równoległą, przełącz, potem utrzymuj. GenAI potrafi zjeść jedno stanowisko w całości i zostawić koszt sumaryczny mniej więcej tam, gdzie był, bo ograniczenie siedzi na stanowiskach, które nigdy nie pokazują się w demie.
Co mierzą benchmarki, a czego nie
Dowody na tę granicę są nietypowo dobre jak na pytanie postawione tak niedawno i wskazują spójny kierunek. Cytuje się je też źle, więc jednostka analizy waży tyle samo co liczba.
Zacznij od łatwego przypadku. Pracę „Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code" opublikowali na ICSE 2024 Pan i współpracownicy z IBM Research oraz UIUC. Przetestowali 1700 próbek kodu w pięciu językach wysokozasobowych, na trzech benchmarkach, dwóch prawdziwych projektach i ponad 10 000 testów. Udział tłumaczeń poprawnych wyniósł od 2,1% do 47,3%, zależnie od modelu. Zespół ręcznie zaklasyfikował 1748 błędów do 15 kategorii pomyłek tłumaczeniowych. To i C, i Java, i Python — obfite dane treningowe i kod, który już ma testy. Najżyczliwsza możliwa wersja tego zadania.
Teraz nieżyczliwa. COBOL-Coder (Dau i współpracownicy, arXiv:2604.03986, 5 kwietnia 2026) ocenił modele na COBOLEval, czyli adaptacji HumanEval do COBOL-a. GPT-4o osiągnął 41,8% kompilowalności przy 16,4 pass@1. Bazowe modele otwarte (CodeGemma, CodeLlama, StarCoder2) w większości nie wyprodukowały programu, który w ogóle się kompiluje. Model zaadaptowany do dziedziny doszedł do 73,95% kompilowalności przy 49,33 pass@1, a na tłumaczeniu z Javy na COBOL do 34,93 pass@1 tam, gdzie modele ogólne miały wynik bliski zeru. Najwięcej waży tu zastrzeżenie metodologiczne. To samodzielne funkcje w stylu HumanEval, nie moduły produkcyjne wpięte w CICS, JCL, VSAM i DB2. Skoro na funkcjach zabawkowych wychodzi 16,4%, dla prawdziwego parku jest to dolne ograniczenie trudności, nie górne.
Najostrzejszy wynik daje AgentModernize (Ahmed i Galib, arXiv:2605.17535, 17 maja 2026). Naiwne podejścia bazowe — tłumaczenie jednym promptem i chain-of-thought — zachowały działanie systemu w 0,0% przypadków, na każdym scenariuszu i każdym testowanym modelu bazowym. Wieloagentowy układ zbudowany wokół jawnego artefaktu pośredniego, grafu specyfikacji behawioralnej, wyciągnął 91,2% reguł wzorcowych przy najlepszym błędzie behawioralnym 8,1%. Wniosek samych autorów jest tym, co trzeba stąd wynieść: wąskim gardłem jest generowanie kodu, nie wydobywanie wiedzy. Czytanie jest tanie. Dowodzenie równoważności nie.
Kierunek poprawy też jest konkretny. SEDCoT (Entin i współpracownicy, arXiv:2607.04092, 5 lipca 2026) raportuje co najmniej 12% zysku wobec najlepszego wcześniejszego wyniku. Wziął się z wykonania symbolicznego i z generowanych zestawów testów; do tego delta debugging minimalizował kontrprzykłady. Granicę przesunęła maszyneria do dowodzenia, nie maszyneria do pisania.
Tymczasem zadania polegające na czytaniu, a nie na przepisywaniu, wypadają dobrze. Diggs i współpracownicy (LLM4Code 2025 przy ICSE, arXiv:2411.14971) uznali dokumentację wygenerowaną przez LLM dla MUMPS-a i IBM Assembly Language Code za w zasadzie wolną od halucynacji, kompletną, czytelną i użyteczną wobec prawdy odniesienia. Asembler był trudniejszy z tej dwójki. Ich własne zastrzeżenie zasługuje na równe miejsce: żadna metryka automatyczna nie korelowała mocno z jakością komentarza, więc nie ma taniego sposobu na zmierzenie, czy dokumentacja jest dobra — poza człowiekiem, który ją przeczyta. COBRAIN (EASE 2025) wyciągnął z COBOL-a reguły biznesowe z precyzją 1,0 przy czułości 0,746 wobec regułowego narzędzia COBREX, a wobec prawdy odniesienia osiągnął F1 0,73 kontra 0,59. W badaniu zrozumiałości z 28 uczestnikami ponad 80% wolało wynik LLM-a.
Należy tu jeszcze jeden wynik, z zupełnie innego ustawienia. METR przeprowadził badanie z randomizacją na 16 doświadczonych programistach open source, na 246 zadaniach w ich własnych dojrzałych repozytoriach (arXiv:2507.09089, lipiec 2025). Z narzędziami AI z początku 2025 roku byli o 19% wolniejsi, a po fakcie szacowali, że narzędzia przyśpieszyły ich o 20%. METR nie przedstawia tego wyniku jako prawa o AI i produktywności, bo prawem nie jest: to migawka z jednego kontekstu, zrobiona na małej próbie i na narzędziach, które od tamtej pory się zmieniły. To najbliższy dostępny odpowiednik sytuacji, w której głęboki kontekst mieszka poza kodem, i zostawia jedną lekcję, która przenosi się dalej. Przyśpieszenie znane wyłącznie z samooceny nie jest dowodem.
| Badanie | Jednostka analizy | Wynik | Czego nie pokazuje |
|---|---|---|---|
| Pan i in., ICSE 2024 | Próbka kodu, języki wysokozasobowe | 2,1–47,3% poprawnych tłumaczeń; 1748 błędów w 15 kategoriach | Niczego o COBOL-u ani o kodzie bez testów |
| COBOL-Coder, 2026 | Samodzielna funkcja COBOL w stylu HumanEval | GPT-4o 41,8% kompilowalności, 16,4 pass@1; model zaadaptowany 73,95% / 49,33 | Modułów produkcyjnych z CICS, JCL, VSAM, DB2 |
| AgentModernize, 2026 | Zachowanie działania w różnych scenariuszach | Naiwne podejścia bazowe 0,0%; wieloagentowe 91,2% reguł wzorcowych, najlepszy błąd behawioralny 8,1% | Że wyciągnięte reguły przetrwają zderzenie z prawdziwym przełączeniem |
| SEDCoT, 2026 | Poprawność tłumaczenia COBOL-a | ≥12% powyżej wcześniejszego stanu sztuki, z wykonania symbolicznego i delta debuggingu | Że sama skala daje ten sam zysk |
| Diggs i in., 2025 | Wygenerowana dokumentacja dla MUMPS-a i asemblera | W zasadzie bez halucynacji, kompletna, użyteczna | Taniej automatycznej metryki jakości — żadna nie korelowała |
| COBRAIN, EASE 2025 | Reguły biznesowe wyciągnięte z COBOL-a | Precyzja 1,0, czułość 0,746; F1 0,73 kontra 0,59; 28 uczestników, ponad 80% wolało | Że wydobycie równa się migracji |
| METR, 2025 | Czas wykonania zadania, 16 programistów, 246 zadań | 19% wolniej z AI; samoocena 20% szybciej | Prawa ogólnego; próba mała, narzędzia poszły dalej |
Tabela dzieli się na dwoje. Gdzie zadaniem jest czytać, opisać albo wydobyć, wyniki są dobre. Gdzie zadaniem jest przepisać i zasłużyć na zaufanie, wyniki są słabe. Poprawiają się, kiedy ktoś dokręci do tego aparaturę weryfikacyjną.
Co masz na stanie
Zanim ocenisz ofertę, dobrze wiedzieć, z czym ma się ona zmierzyć. Liczby o skali, które wszyscy cytują, są słabe, a tę słabość trzeba powtarzać za każdym razem, gdy się ich używa.
W obiegu są dwie. Reuters podał w 2017 roku około 220 miliardów linii COBOL-a w użyciu, 43% systemów bankowych zbudowanych na nim i 95% transakcji kartą w bankomacie, które go dotykają. To dziennikarska kompilacja danych branżowych bez opublikowanej metody liczenia. Druga to „ponad 800 miliardów linii w codziennym użyciu" i pochodzi z ankiety Micro Focus z 2022 roku, na 1104 respondentach z 49 krajów. To dostawca mierzący percepcję, nie audyt repozytoriów. Obie są rzędami wielkości, nie pomiarami, i tak trzeba je podpisywać, ilekroć trafiają na slajd.
Liczby audytowane są lepsze i gorsze zarazem. Amerykańskie GAO (Government Accountability Office) podało w lipcu 2025 roku (GAO-25-107795), że agencje objęte ustawą CFO Act wskazały 69 systemów legacy. GAO wyodrębniło 11 najbardziej krytycznych: od 23 do 60 lat, około 754 milionów dolarów rocznie na utrzymanie. Osiem z jedenastu działa na przestarzałych językach; system Departamentu Skarbu na COBOL-u i asemblerze. Siedem pracuje ze znanymi podatnościami. Cztery chodzą na sprzęcie albo oprogramowaniu bez wsparcia. Kompletny plan modernizacji miały tylko trzy z jedenastu, a dwa nie miały żadnego.
Kanonicznym przypadkiem jest Individual Master File w amerykańskim IRS. Ma ponad 60 lat i jest napisany w asemblerze i COBOL-u; projektowano go jeszcze na IBM System/360. Program jego wymiany, CADE, porzucono w 2009 roku. CADE-2 wciąż go nie zastąpił i nie zapowiada się na to przed 2030. W marcu 2025 IRS poinformował GAO, że wstrzymał programy modernizacyjne na czas rewizji priorytetów. To mniej więcej dwadzieścia lat podejść, wszystkie sprzed istnienia GenAI.
Jeden komercyjny punkt odniesienia, ile kosztuje zrobienie tego porządnie. Commonwealth Bank of Australia wymieniał swoją centralną platformę bankową przez pięć lat, kosztem powyżej miliarda dolarów australijskich (około 750 milionów dolarów amerykańskich po kursie z 2017 roku). Trzymaj tę liczbę obok każdej propozycji obiecującej migrację w dwa kwartały, bo narzędzia się poprawiły.
Jedna narracja wymaga sprostowania, bo to jej najczęściej używa się do budowania presji czasu. „Wszyscy programiści COBOL-a przechodzą na emeryturę" jest dowodem słabszym, niż sugeruje pewność siebie tego zdania, a najlepsze dostępne dane wskazują w drugą stronę. BMC Mainframe Survey 2025 to dwudziesta edycja, na ponad 1100 respondentach. Grupa wiekowa 18–49 lat urosła w niej z 53% w 2018 roku do około 80% w 2025, a pokolenie Z z 1% do 15%. W tej samej ankiecie 97% oceniało platformę pozytywnie, 93% planowało dalsze inwestycje, a alternatywy rozważało tylko 3%, wobec 10% wcześniej. Około 75% traktowało GenAI jako inicjatywę strategiczną, ale tylko jakieś 30% pozwoliłoby AI wykonywać zadania samodzielnie. Ta ankieta pyta własnych klientów dostawcy mainframe'ów, więc ma oczywisty interes we wniosku „mainframe żyje". Miej przed oczami obie te stronniczości naraz.
Zdanie trafne jest węższe i trudniejsze. Kurczy się liczba ludzi, którzy pamiętają, dlaczego dana linijka tam jest. To inny problem niż brak programistów COBOL-a i rekrutacją się z niego nie wyjdzie.
Jedną liczbę chciałbym tu wstawić i nie mogę. Nie ma niezależnego publicznego źródła, które podawałoby, jaki odsetek parków mainframe'owych ma sensowne pokrycie testami regresji. To ta jedna statystyka, która w całym tym wywodzie waży przy decyzji najwięcej, i nikt jej nie policzył.
Jak czytać ofertę modernizacji
Ramy powyżej sprowadzają się do krótkiego przesłuchania. Żadne z tych pytań nie wymaga technicznej głębi, żeby je zadać, a jakość odpowiedzi oddziela plan od broszury.
- Skąd weźmiecie prawdę odniesienia? Praca równoległa na ruchu produkcyjnym? Nagrane wejścia i wyjścia? Czy specyfikacja, której nikt nie widział?
- Jakie pokrycie mają testy, które już istnieją — nie te, które zamierzacie wygenerować? Jeśli zerowe, kto podpisuje oświadczenie, że wyjścia są równoważne?
- Jak wykryjecie różnice, które nie są błędami funkcjonalnymi? Czasy, kolejność wsadów, przepustowość, formaty komunikatów. To pytanie z prawa Hyruma i na nie odpowiada się najgorzej.
- Kto po naszej stronie zna ten system? Ile osób, ile lat stażu i kiedy odchodzą?
- Jaka jest jednostka w waszych benchmarkach? Funkcja w stylu HumanEval czy moduł produkcyjny z CICS i JCL, z danymi w VSAM? W przepaści między nimi mieszka większość marketingu.
- Co jest miarą sukcesu w umowie? Przetłumaczone linie czy transakcje przechodzące pracę równoległą bez rozjazdu? Tylko jedno z tego jest wynikiem biznesowym.
- Big bang czy przyrostowo? Z działającym oryginałem i drogą powrotną, czy jeden weekend na przełączenie?
- Który krok łańcucha AI realnie skraca i jaki udział w obecnym budżecie ma ten krok?
Największy ciężar niesie pytanie pierwsze, a na to, co się dzieje, kiedy zostaje bez odpowiedzi, jest przykład ostrzegawczy. System MiDAS w stanie Michigan rozstrzygnął automatycznie 22 427 spraw o wyłudzenie zasiłku dla bezrobotnych między październikiem 2013 a sierpniem 2015. Stanowy Office of the Auditor General podał w lutym 2016 wskaźnik błędu 93% dla odwołań złożonych między październikiem 2013 a czerwcem 2015, a 20 965 decyzji następnie uchylono. Dziesiątkom tysięcy ludzi naliczono poczwórne kary; ściągano je z wynagrodzeń i ze zwrotów podatku. MiDAS był regułowy, nie uczył się maszynowo, i nazwanie go porażką AI byłoby błędem. Przenosi się wzorzec: automatyzacja wdrożona na skalę tam, gdzie prawdy odniesienia nigdy nie zweryfikowano.
Druga połowa odpowiedzi to miejsca, w których GenAI w programie legacy naprawdę się sprawdza. Własna rekomendacja Gartnera — modernizacja na miejscu zamiast przyśpieszania migracji — wskazuje ten sam zestaw zadań, który wspierają badania.
| Zadanie | Czy jest wyrocznia? | Koszt złej odpowiedzi | Werdykt |
|---|---|---|---|
| Dokumentowanie nieudokumentowanych modułów | Tylko przegląd człowieka; żadna metryka automatyczna nie koreluje | Niski — zły komentarz idzie do kosza | Używaj |
| Wyciąganie reguł biznesowych do artefaktu do przeglądu | Porównywalne z narzędziami regułowymi | Niski, jeśli artefakt przejrzysz przed użyciem | Używaj, potem przejrzyj |
| Generowanie testów charakteryzujących wobec działającego systemu | Wyrocznią jest działający system | Niski — zły test widocznie nie przechodzi | Używaj; najwyższa wartość na godzinę |
| Analiza wpływu i mapowanie zależności | Częściowo sprawdzalne wobec kodu | Średni | Używaj, weryfikuj krawędzie |
| Hurtowe tłumaczenie modułów produkcyjnych | Żadnej, o ile nie ma pracy równoległej | Wysoki — cichy rozjazd zachowań | Nie kupuj tego jako planu |
| Migracja orkiestracji wsadów i okien czasowych | Nie ma w kodzie do czego przyłożyć | Wysoki, i odkryty późno | Ludzie plus praca równoległa |
Wzorzec w prawej kolumnie jest ten sam, który Feathers opisał w 2004 roku. Narzędzie jest mocne wszędzie tam, gdzie działający system może być własnym sędzią, i słabe wszędzie tam, gdzie człowiek musi poświadczyć równoważność z niczego.
Granice i kontrargument, który waży
Trzy zarzuty mają realną siłę i czytelnik z parkiem mainframe'owym powinien mieć w głowie wszystkie trzy.
Pierwszy: granica się przesuwa i przesuwa się właśnie teraz. Między kwietniem a lipcem 2026 pchnęły ją co najmniej trzy opublikowane prace: COBOL-Coder, AgentModernize, SEDCoT. Prognoza Gartnera dotyczy projektów rozpoczętych w 2026 roku, nie trwałej własności technologii. Czytanie jej jako wyroku na GenAI to pomyłka, której lepiej uniknąć. Jest użyteczniejsza wersja tej tezy i da się ją obalić. Granica przesuwa się, jak widać, przez maszynerię weryfikacyjną — wykonanie symboliczne, generowane zestawy testów, delta debugging, jawne artefakty pośrednie — a nie przez skalę modelu. Jeśli kolejny duży zysk na tłumaczeniu COBOL-a przyniesie większy model ogólnego przeznaczenia bez doczepionej aparatury weryfikacyjnej, mechanizm opisany tutaj jest błędny.
Drugi: większość tych porażek będzie miała zwyczajne przyczyny. Flyvbjerg i Budzier pisali w Harvard Business Review we wrześniu 2011 roku o 1471 przebadanych projektach IT. Średnie przekroczenie kosztu wyniosło 27%. Ale jeden na sześć był czarnym łabędziem, ze średnim przekroczeniem kosztu 200% i harmonogramu blisko 70%. Rozkład ma gruby ogon, więc średnia nic nie mówi o ryzyku. Migracja TSB z kwietnia 2018 czyni tę myśl namacalną. Około pięciu milionów klientów przeniesiono z systemów Lloyds w jeden weekend. Bankowość oddziałowa, telefoniczna, internetowa i mobilna stały się niedostępne dla dużej części bazy 5,2 miliona klientów, a ataki oszustów szły w szczycie na poziomie 70 razy wyższym niż normalnie. Kary regulacyjne wyniosły 48,65 mln funtów: FCA 29,75 mln plus PRA 18,9 mln, po 30-procentowym rabacie ugodowym, a bez niego 69,5 mln. Do tego doszły rekompensaty na 32,7 mln funtów, a całkowity raportowany koszt sięgnął około 366 mln funtów. Niezależny przegląd Slaughter and May, opublikowany w listopadzie 2019 roku, wskazał na skalę i tempo migracji oraz na brak oceny zdolności głównego dostawcy IT. Zakres, tempo, nadzór nad dostawcą i przełączenie na raz. Nigdzie żadnej AI. Sporo projektów, których dotyczy prognoza Gartnera, rozminęłoby się z korzyściami z modelem w pokoju albo bez niego. Przypisanie wszystkiego GenAI zasłania przyczyny, które już tam były.
Trzeci: granica kompetencji jest znakomitym alibi. „AI tego nie potrafi" nie znaczy „więc zostawiamy to w spokoju", a ześlizg z pierwszego zdania w drugie to najdroższy dostępny ruch. Jedenaście systemów GAO kosztuje w utrzymaniu około 754 milionów dolarów rocznie, siedem ze znanymi podatnościami, cztery na sprzęcie albo oprogramowaniu bez wsparcia. IRS próbuje od czasów sprzed 2009 roku i w marcu 2025 wstrzymał się ponownie. Koszt nieruszania się jest mierzalny, kumuluje się i jest większy niż koszt ostrożnego ruszenia. Modernizacja na miejscu to realna opcja. Przyrostowa wymiana z pracą równoległą i drogą powrotną to realna opcja. Kolejna dekada wstrzymanych programów, uzasadniona trafną obserwacją o granicach modelu, opcją nie jest.
Jest czwarta granica i dotyczy samej kotwicy. Prognoza Gartnera nie ma opublikowanej próby, nie ma metody, nie ma operatu losowania, a ja czytałem ją przez trzy źródła wtórne, bo źródło zwraca 403. Zbiega się z badaniami powyżej i to ma swoją wagę. Pomiarem nie jest i nie należy jej cytować jako pomiaru — również w esejach, które się z nią zgadzają.
Zamknięcie
Zadaj ofercie modernizacji dwa pytania. Który krok pracy narzędzie usuwa? I kto podpisuje oświadczenie, że nowy system robi to samo co stary? Jeśli nikt w pokoju nie umie wskazać tej osoby, nie kupujesz migracji. Kupujesz tłumaczenie i nadzieję, że wyjdzie.
Bibliografia
- primaryGartner, „Gartner Predicts More Than 70% of Mainframe Exit Projects Will Fail Due to Overestimation of Generative AI's Capabilities" (komunikat prasowy, 18 czerwca 2026) — prognoza 70%, prognoza 75% o czystce wśród dostawców i rekomendacja „modernizacja na miejscu"; analityk Alessandro Galimberti. (Newsroom zwrócił 403 przy dostępie automatycznym; brzmienie zweryfikowane przez źródła wtórne poniżej.)
- primaryWarren Buffett, list do akcjonariuszy Berkshire Hathaway (1996) — krąg kompetencji; „znajomość jego granic jest natomiast niezbędna".
- primaryG.K. Chesterton, The Thing (1929), rozdział „The Drift from Domesticity" — płot w poprzek drogi.
- primaryTitus Winters, Tom Manshreck i Hyrum Wright, Software Engineering at Google (O'Reilly, 2020) — prawo Hyruma, zaobserwowane przez Hyruma Wrighta: na każdym obserwowalnym zachowaniu systemu ktoś się oprze.
- primaryMichael Feathers, Working Effectively with Legacy Code (Prentice Hall, 2004) — kod legacy jako kod bez testów; testy charakteryzujące utrwalają zachowanie faktyczne, a nie zamierzone.
- primaryH.G. Rice, „Classes of Recursively Enumerable Sets and Their Decision Problems", Transactions of the American Mathematical Society 74(2) (1953) — nietrywialne własności semantyczne, w tym równoważność programów, są nierozstrzygalne.
- primaryRangeet Pan i in., „Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code", ICSE 2024 — 1700 próbek, pięć języków, 2,1–47,3% poprawnych tłumaczeń, 1748 zaklasyfikowanych błędów. (arXiv:2308.03109 · DOI 10.1145/3597503.3639226)
- primaryA.T.V. Dau i in., „COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation" (arXiv:2604.03986, 5 kwietnia 2026) oraz P. Entin i in., „SEDCoT: Enhancing LLM-Based COBOL Code Translation via Symbolic Execution and Delta Debugging" (arXiv:2607.04092, 5 lipca 2026) — wyniki na COBOLEval, COBOL jako język niskozasobowy i zyski z maszynerii weryfikacyjnej.
- primaryS.N. Ahmed i M. Galib, „AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs" (arXiv:2605.17535, 17 maja 2026) — 0,0% zachowania działania dla naiwnych podejść bazowych; wąskim gardłem jest generowanie kodu, nie wydobywanie wiedzy.
- primaryC. Diggs i in., „Leveraging LLMs for Legacy Code Modernization: Challenges and Opportunities for LLM-Generated Documentation", LLM4Code 2025 przy ICSE (arXiv:2411.14971), oraz „LLM vs Rule-Based — The COBRAIN Tool and An Empirical Study on Extracting Business Rules from COBOL", EASE 2025 (DOI 10.1145/3756681.3756982) — wyniki dla dokumentacji i wydobywania reguł, z zastrzeżeniem autorów, że żadna metryka automatyczna nie koreluje z jakością.
- primaryMETR, „Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (10 lipca 2025) — badanie z randomizacją, 16 programistów, 246 zadań, 19% wolniej wobec samooceny 20% szybciej. (arXiv:2507.09089)
- primaryU.S. GAO, „Information Technology: Agencies Need to Plan for Modernizing Critical Decades-Old Legacy Systems", GAO-25-107795 (17 lipca 2025), oraz „Information Technology: IRS Is Developing a New Modernization Framework", GAO-25-107611 — 11 systemów krytycznych, 754 mln dolarów rocznego utrzymania i kalendarium Individual Master File w IRS.
- primaryTSB Bank / Slaughter and May, „Independent Review of TSB's 2018 migration to a new IT platform" (19 listopada 2019), oraz FCA, „TSB fined £48.65m for operational resilience failings" (20 grudnia 2022) — skala, tempo, ocena dostawcy, kary i rekompensaty.
- primaryBent Flyvbjerg i Alexander Budzier, „Why Your IT Project May Be Riskier Than You Think", Harvard Business Review (wrzesień 2011) — 1471 projektów, średnie przekroczenie kosztu 27%, jeden na sześć z przekroczeniem 200%.
- secondaryIT-Online, „Overestimating GenAI will see mainframe exit projects fail" (23 czerwca 2026); DIGIT, „Orgs are overestimating GenAI when it comes to mainframe exits" (22 czerwca 2026); TechEdgeAI, „Mainframe Exit Projects Face GenAI Reality Check" (18 czerwca 2026) — trzy serwisy cytujące komunikat Gartnera dosłownie; podstawa każdego cytatu z Gartnera powyżej.
- secondaryLiczby o skali parku, wszystkie słabe, użyte wyżej wyłącznie jako rzędy wielkości: Reuters przez CNBC, „Banks scramble to fix old systems as IT 'cowboys' ride into sunset" (11 kwietnia 2017) — liczby 220 mld linii / 43% / 95% oraz koszt Commonwealth Bank of Australia, dziennikarska kompilacja bez opublikowanej metody liczenia; The Stack, „There's over 800 billion lines of COBOL in daily use" — ankieta Micro Focus z 2022 roku, 1104 respondentów z 49 krajów, dostawca mierzący percepcję; TechChannel, „Mainframers Are Getting Younger: BMC Survey" — BMC Mainframe Survey 2025, ponad 1100 respondentów, sponsorowana przez dostawcę, na samodobierającej się próbie jego własnych klientów.
- secondaryGovTech, „Michigan Integrated Data Automated System Experiences 93 Percent Error Rate" (o raporcie Michigan Office of the Auditor General, luty 2016), oraz Ford School STPP, „Cahoo v. SAS — MiDAS explainer" (2024) — 22 427 spraw rozstrzygniętych automatycznie, 20 965 uchyleń; system regułowy, nie uczenie maszynowe.