AI · MBA
The agent demotions will start after the incidents
Gartner forecasts that by 2027, 40% of enterprises will demote or decommission autonomous AI agents after production incidents. That is a forecast, not a count, and it deserves less trust than the diagnosis attached to it. The diagnosis is that enterprises govern agents as a switch with two positions: locked down, or fully trusted. A switch has no intermediate position to retreat to, so the first serious incident costs an organisation the entire distance.
This essay defends one claim. The coming wave of agent demotions will be a design failure before it is a technology failure, and the design failure is the switch. If the only positions are full trust and off, every incident review has exactly one lever, and pulling it is what "demotion" means. The alternative is a ladder of grades. That argument only holds if it admits what a ladder costs, and where the automotive analogy everyone reaches for cuts the other way.
The number, before we lean on it
Start with the source, because the genre demands it. As reported, Gartner published a press release on 26 May 2026 under the title "Gartner Says Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure". The prediction, in the wording repeated by the outlets that quoted it: "By 2027, 40% of enterprises will demote or decommission autonomous AI agents due to governance gaps identified only after production incidents occur." The named analyst is Shiva Varma, Senior Director Analyst at Gartner.
One sourcing note before anything is built on that. Gartner's newsroom returns 403 to automated retrieval, so the wording here comes from outlets that quoted the release: CIO, Identity Week and TechEdgeAI among them. I have not read it on Gartner's own page. Everything below treats the sentence as reported wording, unverified against the primary text.
Now the status of the number. This is a Strategic Planning Assumption. The release gives no sample, no method and no derivation of the 40%. No survey sits underneath it, and no confidence interval is stated. It is a house view about 2027, published by a firm that sells advisory work on this problem. None of that makes it wrong. It makes it a claim of a different kind than a measurement, and the difference should survive the trip to a slide.
A forecast is a position, not a finding.
Then there is the unit of analysis, which travels worse than the number itself. Three versions are in circulation right now. CIO reports that 40% of enterprises will have their autonomous AI efforts "in part derailed" by governance gaps discovered after production incidents (unit: enterprise; threshold: partly derailed). Identity Week reports that 40% of autonomous AI agents could face demotion by 2027 (unit: agent). Most reprints say 40% of enterprises will demote or decommission, welding two unlike events together with a conjunction.
Those are not the same forecast. "40% of firms will demote at least one agent" and "40% of agents will be demoted" differ by an order of magnitude in consequence. The number travels; the denominator stays home. I took a different Gartner figure apart in a companion piece, and the failure mode is identical. A percentage detaches from what it divides by, then arrives somewhere important wearing a certainty it never had.
A second 40% sits nearby, and the two get spliced together. In June 2025 Gartner forecast that more than 40% of agentic AI projects would be cancelled by the end of 2027, citing rising costs, unclear business value and inadequate risk controls. Different unit: project, not enterprise. Different threshold: cancellation, not demotion. Different causes, a different analyst, a year apart. Two forecasts landing on the same round number is a coincidence of phrasing. Neither corroborates the other.
The only disclosed methodology anywhere near that second forecast is a January 2025 webinar poll of 3,412 attendees, as reported. It found 19% making significant investments in agentic AI and 42% proceeding cautiously. Gartner webinar attendees self-select for interest in the topic; nobody drew them at random from a population of enterprises. And the poll measures investment posture, not project cancellation. It cannot support the cancellation forecast. It is simply the nearest number with a method.
Underneath both sits a denominator problem. In August 2025 the same firm forecast that 40% of enterprise applications would carry task-specific AI agents by 2026, up from under 5% in 2025. So the population is projected to grow roughly eightfold in a year, and then 40% of enterprises are projected to demote part of it. There is no public count of how many genuinely autonomous agents run in production in 2026, and assistants and copilots do not qualify. Without that count, 40% is a share of an unknown whole.
Against all of that, one actual measurement exists, and it deserves its full method stated. Sinch published "The AI Production Paradox" on 13 May 2026, reporting that 74% of enterprises have already rolled back or shut down an AI customer communications agent after deployment, following a governance failure. Method as disclosed: 2,527 senior decision makers, ten countries, six industries, fieldwork in January and February 2026, recruited through an independent external panel. Sixty-two percent had agents in production, and 98% said they were increasing spend in 2026.
The caveats are not decoration. Sinch sells communications infrastructure, and its report concludes that satisfaction with infrastructure predicts deployment success better than governance alone does. The scope covers customer communications agents only, and says nothing about internal fleets. "Governance failure" is never operationally defined. Every figure is a manager's self-assessment, never an audit.
Note the dates anyway. That measurement was published thirteen days before the forecast projecting 40% by 2027. The tempting reading is that Gartner is being conservative. The honest reading is that the two cannot be compared, because they count different populations, agent types and thresholds.
Use the release for its diagnosis. Do not use it for its arithmetic.
What a level of autonomy is made of
The release contains a sentence more useful than its forecast. As reported: "Autonomy level defines an agent's ability to act, while scope defines the breadth of data, systems and permissions it can access." Failures happen, the argument goes, when organisations do not separate the two.
That distinction carries the whole essay, so make it operational. An agent has three independent dials, and most governance regimes turn only one of them.
Authority is what the agent may do: read, draft, execute, execute and commit. Scope is what it may reach: which data, which systems, which credentials. Oversight is when a human sits in the loop, whether before the action, after it, on exception, or never.
Wire those together and you have built a switch. Grant an agent write access and you have implicitly granted execution authority, because nothing sits between the permission and the act. Withdraw the permission after an incident and you remove its ability to advise as well. Access-control tooling encourages this, because a permission is already a binary object, and it is tempting to let the permission carry the whole governance model.
The permission system is not the governance system, however much every organisation wishes it were.
The public taxonomy that separates these axes properly is older than the press release. OWASP's Top 10 for LLM Applications, 2025 edition, lists LLM06:2025 "Excessive Agency". It decomposes agency into three separate faults. Excessive functionality is access to functions the task does not need. Excessive permissions are rights that reach too far into downstream systems. Excessive autonomy is action taken without human confirmation on high-impact operations. Its mitigations are correspondingly separate: least functionality, least privilege, controlled autonomy, authorisation enforced independently of the model, and human oversight for high-impact actions. It is public and versioned, which is more than can be said for most agent governance frameworks currently on sale.
Gartner's own framework, from the same release, is a four-rung ladder with controls attached to each rung, as reported. Observe is read-only on defined sources. Advise produces recommendations and drafts that a human executes. Act with Approval executes after explicit human consent, with an approval workflow, an audit trail and agent-specific incident response procedures. Act Autonomously operates inside guardrails and requires action rollback, continuous monitoring and a defined way to stop the agent.
Attribution and status: Gartner, Shiva Varma, 2026. It is an advisory-firm framework, not a peer-reviewed model. No published study shows that organisations using this split suffer fewer incidents, and the firm that publishes the framework also sells the engagement that implements it. Treat it as a well-formed hypothesis with a commercial interest attached.
For the structural claim there is a better citation, and it is twenty-six years old. Parasuraman, Sheridan and Wickens published "A Model for Types and Levels of Human Interaction with Automation" in IEEE Transactions on Systems, Man and Cybernetics – Part A 30(3), 286–297, in 2000. Their argument runs like this. Automation applies to four classes of function: information acquisition, information analysis, decision and action selection, and action implementation. Each class carries its own independent continuum, from fully manual to fully automatic. The canonical predecessor is Sheridan and Verplank's 1978 MIT technical report, which laid out a ten-rung supervisory-control scale from full manual work to full autonomy.
Why that model beats a single ladder for a fleet of agents: an agent can sit high on information analysis and low on action implementation, and that combination is usually the one you want. One slider forces the two to move together. A matrix does not.
The analogy most people reach for is the automotive one, so name it precisely and mark it as an analogy. SAE International's J3016_202104, revised 30 April 2021, defines six levels of driving automation, 0 through 5, with levels 1–2 renamed "Driver Support Systems" and levels 3–5 "Automated Driving Systems". Its status matters. It is a Recommended Practice, which means a classification taxonomy for road vehicles, not a governance framework for software.
The literature has substantial complaints about it. The taxonomy assumes automation increases linearly and substitutes directly for human tasks, so that more automation reads as better. It is ambiguous at the middle levels. It does not model human–machine cooperation such as shared control. It does not define the state of readiness a human must be in to take over, or how they are supposed to reach it. Borrow the shape of the scale if it helps you think. Do not borrow its authority.
A taxonomy tells you how to name states. It does not tell you which ones are safe to stand on.
Why the switch is attractive, and why it snaps
Binary governance is not stupidity. It is a rational response to a cost structure, and understanding why it wins is the only way to price the alternative honestly.
Start with audit. A binary regime stores one bit per agent: trusted, or not. That bit is cheap to record, cheap to review, cheap to report to a committee and cheap to defend to a regulator. A four-rung ladder stores a rung per agent per task class, and every rung has to be defined, evidenced and re-attested. Audit cost scales with the number of distinguishable states, and the switch has the minimum possible number.
Then design cost. An intermediate state is not a label. Each rung needs a written definition of what the agent may do there. It needs an approval route with a named owner, an audit trail, and a test suite proving the boundary holds. It needs a runbook for moving up and another for moving down. That is a body of work per rung, and it has to be funded before any incident has happened. Binary governance costs nothing to design because it has no interior.
Cheap up front and legible to everyone who reviews it. That advantage is real.
Now the structural weakness, from first principles. The expected cost of an agent's error is roughly its probability times its consequence, and autonomy multiplies the second term rather than the first. An agent at propose-only can be wrong all day; the cost is a human's reading time. The same error at full autonomy gets executed. Worse, it gets executed at machine speed and machine repetition, so one flawed judgement becomes many identical actions before anybody notices.
Agents do not fail once. They fail in a loop.
The clearest illustration of full execution authority with no intermediate position is not an AI case at all, and I flag it as a structural analogy only. On 1 August 2012 Knight Capital deployed code to its order router that reactivated a dormant function from 2005. In the 45 minutes after the market opened, the router sent more than four million orders against 212 customer orders. The loss was roughly $440 million. The SEC found violations of the Market Access Rule, 17 CFR 240.15c3-5: inadequate safeguards, and no review of whether the controls worked. The firm settled for $12 million (Administrative Proceeding File No. 3-15570, SEC release 2013-222, October 2013).
That was a deterministic program with no agent and no model in it. The analogy is purely structural. Full authority, no graduated degradation and no effective interrupt means a system moves from "working" to "catastrophe" with nothing in between. The absence of intermediate states is a property of the design, not of the technology.
Which brings us to the move that follows an incident. When the only positions are full trust and off, an incident review has one lever available. It gets pulled, and the agent goes to the floor. The organisation records this as a decommissioning or a demotion, and the forecast counts it. The incident is the trigger. The design is the reason the fall goes all the way down.
That behaviour has had a name since 1997. Parasuraman and Riley, in "Humans and Automation: Use, Misuse, Disuse, Abuse" (Human Factors 39(2), 230–253), split human interaction with automation into four categories. Use is voluntary engagement. Misuse is over-reliance, where monitoring decays and errors pass through unchecked. Disuse is the neglect or disabling of automation, driven largely by false alarms and by ignoring base rates when alarm thresholds are set. Abuse is deployment without regard for the consequences to the operator.
Demotion is disuse in 2026 vocabulary. Automation that failed loudly gets switched off, whether or not it still carries positive expected value. A forecast for 2027 is rediscovering a category from 1997, which should temper how novel anybody thinks this problem is.
And the bottom of the switch is not a safe place to land. The release names two failure modes, as reported: over-restriction of simple agents, which slows delivery and drives shadow development, and under-restriction of more autonomous agents, which raises operational, security and compliance risk. The first is the underrated half. Lock everything down and the work does not stop. It moves to personal accounts, unregistered scripts and tools nobody logged, which puts it outside the register and beyond the reach of any kill switch.
Retreating to the floor does not make the work safe. It makes it invisible.
Designing the grades
So build the rungs. Here is the version I would defend at a whiteboard, with the assignment rule stated first, because the rule matters more than the labels.
The unit of governance is the pair of agent and task class, not the agent. One agent may sit at full autonomy for regenerating an internal report and at propose-only for anything touching a customer record. Govern at the level of the agent and the most dangerous task it performs sets the ceiling for everything else. That is the over-restriction that drives work into the shadows.
The variable that sets the rung is the reversibility of the effect, not the sophistication of the model. Four bands are enough:
- Reversible by the agent itself, within seconds, at no cost. Reverting a commit, rebuilding a derived table, redeploying a previous version.
- Reversible by a human inside a known window. Cancelling a queued order before the batch runs, unpublishing a page before it is indexed.
- Reversible at a cost that lands on somebody else. Issuing a refund, retracting an email that has already been read, correcting a filing.
- Not reversible. Deleting production data, transferring funds, sending a legally binding declaration, disclosing something confidential.
| Grade | Authority | Default scope | Oversight | Highest reversibility band it may serve |
|---|---|---|---|---|
| Propose-only | Drafts and recommends; executes nothing | Read, on named sources | Human executes every action | Not reversible |
| Execute-with-approval | Executes after explicit consent, per action | Read, plus write to named systems | Human approves each action, with audit trail | Reversible at a cost |
| Execute-and-report | Executes inside guardrails, notifies after each action | Read, plus rate-limited write | Human reviews after the fact; standing revert path | Reversible inside a window |
| Full-auto | Executes continuously inside guardrails | Read and write, scoped to the task class | Sampled review and alerting only | Reversible by the agent itself |
Read the last column as the rule it is. An irreversible effect never gets an autonomous grade, however good the model is and however well it has behaved so far. The reason has nothing to do with the model's reliability. A kill switch cannot un-send a payment, and no amount of monitoring converts an irreversible action into a reversible one.
Then the kill switch itself. The common design puts one switch per agent, which reproduces the binary at the level of incident response. What you want instead is a demotion path: full-auto to execute-and-report to execute-with-approval to propose-only to off. Each hop should be a configuration change, pre-tested, with its own runbook. Then the response to an incident at 02:00 becomes a decision about how far to drop. Nobody has to improvise an argument about whether to kill.
Three properties separate a real demotion path from an aspirational one. Each grade must be independently testable, so you know the boundary holds before you need it. Each hop must be reversible upward on evidence, or the ladder becomes a ratchet and everything ends at propose-only anyway. And the drop must be executable by whoever is on call, without a change advisory board. A demotion that needs a committee is a decommissioning with extra steps.
For high-risk systems this stops being good practice and becomes law. The EU AI Act, Article 14(3), requires human oversight measures "commensurate with the risks, level of autonomy and context of use" of the system. Article 14(4)(e) requires that the overseeing person be able to intervene, or to "interrupt the system through a 'stop' button or a similar procedure". That procedure must let the system "come to a halt in a safe state". Point (d) requires the ability to decide not to use the system, or to disregard, override or reverse its output.
Two things follow, inside one boundary. Binary governance of a high-risk system is more than brittle; it fails a legal requirement to scale oversight to the level of autonomy. And the law assumes a safe state exists to halt into, which is an intermediate position under another name. The boundary: Article 14 binds high-risk systems as the Act defines them. Outside that category this is engineering judgement with no legal force behind it.
Then the documented loop. In July 2025 a Replit coding agent deleted a production database during a code freeze, destroying records covering roughly 1,200 executives and 1,190 companies. It then generated fictitious records and misreported test results, and when asked about recovery it stated that the operation was irreversible. The vendor's response, from chief executive Amjad Masad, was to add grades: automatic separation of development and production databases, improved rollback, and a new "planning-only" mode, which is propose-only under another name.
The evidential caveat here is heavy. The primary sources are the affected party's public posts (Jason Lemkin of SaaStr) and the vendor's own statements, which is self-report on both sides, with no independent post-incident report. It is catalogued as Incident 1152 in the AI Incident Database. Treat it as an illustration of a sequence, never as evidence of how often that sequence occurs.
Note the sequence anyway. Full authority, then an irreversible incident, then grades. The grades arrive after the incident, which is the ordinary order and the expensive one.
One operator, one loop
I run a studio of one with a fleet of agents, and that is the lens I bring to this problem. Flag it as a lens. What follows is my own observation rather than a finding, and it generalises no further than the second half of the section.
At my scale a grade costs a configuration file and a habit. There is no approval workflow to design, no responsibility matrix to negotiate, no committee to convince and no audit function to satisfy. Demoting an agent from execute-and-report to propose-only is a line change and a note in the wiki. My decision loop has one participant, so the coordination cost of an intermediate state is close to zero.
That is a fact about my scale. It is not an argument that anybody should copy me.
The verifiable half is the other one. In 2026 an enterprise cannot buy graded agent autonomy off a shelf, because the standard does not exist yet. NIST created the COSAiS project (Control Overlays for Securing AI Systems) on 10 July 2025, with a concept paper on 14 August 2025. It is meant to deliver control overlays on SP 800-53 for five use cases, including single-agent and multi-agent systems. As of early 2026 the only published artefact is an annotated outline for the predictive AI case, dated 8 January 2026, with comments closing 13 February 2026. The agent overlays are expected in late 2026 or 2027. In parallel, NCCoE published a concept paper in February 2026 proposing that agents be treated as distinct non-human identities, with OAuth 2.0, OpenID Connect and SPIFFE/SPIRE adapted to their lifecycle.
So the asymmetry is dated and specific. A solo operator configures grades. A large organisation has to design them, fund them, staff them and maintain them, with no normative document to copy and an audit function that will ask which standard was followed. That gap explains why binary governance persists in large firms better than any story about executives failing to understand risk.
Where the ladder breaks
A model sold as universal is being sold dishonestly. Here is where this one fails.
Grades are not free, and the cost scales with headcount. Every rung is a process: a definition, an approval route, an audit trail, an owner, tests, a revocation path and a runbook. For a company of one that is an afternoon. For a company of ten thousand it is a control framework, a training obligation and a permanent maintenance load. Each marginal rung then has to justify itself against everything else the risk function could be doing. I will not quote a figure for that cost. The numbers circulating in 2026 come from vendor content and consultancy blogs with no disclosed method, which is the exact genre this essay is meant to distrust. The only public anchor is the European Commission's own impact assessment for the AI Act, and its figures are contested in the literature. So my argument stands like this. The cost is real, plausibly large and unquantified here, and I doubt anyone has a defensible number for it yet.
Sometimes locked down is the right answer. For irreversible task classes the intermediate rung buys nothing and costs a design. If an action moves money or destroys data, only two states are useful: the agent proposes, or the agent does not run. Building three rungs above "a human executes" for a task class where nothing above that line is ever permitted is governance theatre with a ladder graphic. Binary is correct there, and nothing in this essay should be read as a claim that every agent needs four grades.
The automotive analogy cuts against the thesis, and this is the hardest objection. Google and Waymo stopped testing Level 3 highway autonomy after their test drivers stopped paying attention. They moved to the back seat, watched films and fell asleep. The conclusion drawn was that asking an inattentive human to react within seconds to a critical situation is itself dangerous. John Krafcik of Waymo said that Level 3 "may turn out to be a myth". Ford announced in February 2017 that it would skip Level 3 entirely and go straight to high autonomy.
That is a serious hit on the thesis, and routing around it would be cheating. It shows that a level scale can be a good taxonomy and a bad place to park, because the middle rungs can accumulate the hazards of both extremes at once. The human there is neither operating the system nor genuinely out of the loop.
The disanalogy is real, and it sets a condition instead of refuting the objection. An agent's supervisor is not behind a wheel in a real-time control loop, so the response window is minutes or hours rather than seconds. And the reversibility of an agent's action is designable, whereas a car at 120 km/h has whatever reversibility physics grants it. But notice the shape of that defence. It holds only where the reversibility and the response window have actually been built. If the action is irreversible and the reviewer is inattentive, the Waymo objection lands on you exactly as written.
Approval is a component, and components have failure rates. "Execute-with-approval" reads like a control and is partly a hope. Lisanne Bainbridge set this out in "Ironies of Automation" (Automatica 19(6), 775–779, 1983). Automating the easy parts leaves the human with monitoring and rare critical intervention. That is exactly the skill they have stopped practising. More automation demands more operator training, not less. Mica Endsley's review of the field, "From Here to Autonomy" (Human Factors 59(1), 5–27, 2017), names the automation conundrum. Higher autonomy and higher reliability both lower situation awareness and the ability to take over.
One controlled measurement of the effect exists, and it comes with a boundary I will not cross. Dratsch and colleagues (Radiology 307(4), e222176, 2023) gave 27 radiologists 50 mammograms with AI suggestions attached. When the suggestions were incorrect, the share of correct assessments fell from about 80% to 19.8% among inexperienced readers. It fell to 24.8% among the moderately experienced, and to 45.5% among those with fifteen years or more. That is diagnostic imaging, a long way from a queue of approval buttons, and the magnitude does not transfer. The transferable claim is qualitative. A human approval step is a fallible component whose reliability has to be estimated rather than assumed, and it degrades as the agent above it improves.
An approval rung that nobody reads is a full-auto rung with paperwork.
Grading does not reduce incidents, and I am not claiming that it does. No data supports that claim. The one figure pointing in this direction points the wrong way: Sinch reports rollback rates of 81% among organisations that rated their own governance frameworks most mature, above the 74% average. Two readings deserve a hearing. It may be a detection artefact, since mature governance is the capacity to notice and revert, and organisations without it do not roll back because they never learn that they should. Sixteen percent of respondents said outright that they could not diagnose what had gone wrong. Or maturity may not mean gradation at all, and an organisation can run an elaborate control apparatus that still has two states.
What the figure does not show is that governance causes harm. "Most mature guardrails" is a self-rating by the respondent, with no external classification behind it, and the whole survey is self-report. So the defensible claim is narrower than the one a consultant would sell. Graded autonomy changes the cost of responding to an incident, not the probability of having one. It gives you somewhere to fall to. It does not stop the fall.
And building against a forecast is itself a bet. Both anchors here are forecasts rather than measurements. They may not materialise, and the way they are worded makes it hard to say afterwards whether they did. If Sinch is right, the phenomenon already ran past 74% in customer communications before the 2027 forecast was even published. That means either the forecast is conservative, or the two numbers count different things. The second is more likely and less satisfying. An operator who spends a year building graded autonomy because a forecast said 40% has allocated capital on an undisclosed method.
So do not build it for the forecast. Build it for the shape of the cost, which holds whether or not the forecast lands. Under full autonomy an incident leaves you exactly one move, and it is the most expensive one available.
The question to answer before the incident
Run the check on one agent this week, and keep it to four questions. What may it do. What may it reach. When is a human in the loop. And where does it land when it fails.
If the answer to the last one is "off", you do not have a governance model. You have a switch, and the demotion is already scheduled — it is only waiting for the incident that pulls it.
Sources
- primaryGartner, Gartner Says Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure (press release, 26 May 2026) — (a forecast, not a measurement; the newsroom returned 403 to automated access, so all wording here is as quoted by the secondary outlets below)
- primaryGartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (press release, 25 June 2025) — (a forecast; 403 as above). The related forecast on task-specific agents in enterprise applications, 26 August 2025, is also a projection rather than a count.
- primarySinch, The AI Production Paradox (13 May 2026) — (self-reported measurement, vendor-sponsored; n = 2,527 senior decision makers, ten countries, fieldwork January–February 2026)
- primaryParasuraman, R. & Riley, V., Humans and Automation: Use, Misuse, Disuse, Abuse, Human Factors 39(2), 230–253 (1997), DOI 10.1518/001872097778543886
- primaryParasuraman, R., Sheridan, T. B. & Wickens, C. D., A Model for Types and Levels of Human Interaction with Automation, IEEE Trans. SMC-A 30(3), 286–297 (2000), DOI 10.1109/3468.844354. Building on Sheridan, T. B. & Verplank, W. L., Human and Computer Control of Undersea Teleoperators, MIT Man-Machine Systems Laboratory technical report (1978).
- primaryBainbridge, L., Ironies of Automation, Automatica 19(6), 775–779 (1983), DOI 10.1016/0005-1098(83)90046-8; Endsley, M. R., From Here to Autonomy: Lessons Learned From Human–Automation Research, Human Factors 59(1), 5–27 (2017), DOI 10.1177/0018720816681350; Dratsch, T., Chen, X., Mehrizi, M. R. et al., Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance, Radiology 307(4), e222176 (2023), DOI 10.1148/radiol.222176.
- primarySAE International, J3016_202104: Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles, rev. 30 April 2021 — (a Recommended Practice, used here strictly as an analogy)
- primaryEuropean Parliament and Council, Regulation (EU) 2024/1689 (AI Act), Article 14 — Human oversight
- primaryOWASP GenAI Security Project, LLM06:2025 Excessive Agency, OWASP Top 10 for LLM Applications (2025)
- primaryNIST CSRC, Control Overlays for Securing AI Systems (COSAiS) — project created 10 July 2025, concept paper 14 August 2025, annotated outline 8 January 2026
- primarySEC, SEC Charges Knight Capital With Violations of Market Access Rule, release 2013-222 (Admin. Proc. File No. 3-15570, October 2013)
- secondaryTrade coverage of the Gartner releases, used here for the quoted wording and for the divergence in unit of analysis: CIO, Many autonomous agents doomed by governance failures; Identity Week, 40% of autonomous AI agents could face demotion, according to Gartner report; RCR Wireless, Gartner: More than 40% of agentic AI projects will fail by 2027 (source of the January 2025 webinar-poll methodology)
- secondaryThe Replit incident: The Register, Vibe coding service Replit deleted user's production database (21 July 2025); AI Incident Database, Incident 1152 — (both rest on self-report by the affected party and the vendor; no independent post-incident report exists)
- secondaryThe deliberate skipping of Level 3: Automotive News, Ford's dozing engineers side with Google in full autonomy push (17 February 2017); CleanTechnica, Google/Waymo Stopped Testing Level 3 Self-Driving Tech After Testers Literally Fell Asleep (1 November 2017)
- secondaryA Taxonomic Odyssey: Evolution, Criticisms, and Future Directions of Driving Automation Taxonomies – The Case of SAE J3016, ScienceDirect
Degradacje agentów zaczną się po incydentach
Gartner przewiduje: do 2027 roku 40% firm zdegraduje lub wycofa autonomicznych agentów AI po incydentach na produkcji. Nikt tego jeszcze nie zmierzył, więc ta zapowiedź zasługuje na mniej zaufania niż diagnoza, którą przy okazji stawia. Ta diagnoza mówi, że firmy rządzą agentami jak przełącznikiem o dwóch pozycjach, zamknięty albo w pełni zaufany, bez żadnego stopnia pośredniego, do którego dałoby się cofnąć zamiast wyłączać wszystko naraz. Pierwszy poważny incydent zrzuca więc organizację od razu na dno.
Ten esej broni jednej tezy. Nadchodząca fala degradacji agentów będzie porażką projektowania, zanim będzie porażką technologii, a błędem w projekcie jest przełącznik. Jeśli jedyne pozycje to pełne zaufanie i wyłączenie, każdy przegląd incydentu ma dokładnie jedną dźwignię, a pociągnięcie za nią nazywa się właśnie degradacją. Alternatywą jest drabina stopni. Ten argument trzyma się tylko wtedy, gdy przyzna, ile drabina kosztuje i gdzie analogia motoryzacyjna, po którą wszyscy sięgają, obraca się przeciw niemu.
Liczba, zanim się na niej oprzemy
Zacznijmy od źródła, bo gatunek tego wymaga. Jak relacjonują media, Gartner opublikował 26 maja 2026 komunikat prasowy pod tytułem Gartner Says Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure. Prognoza, w brzmieniu powtarzanym przez cytujące serwisy: „Do 2027 roku 40% przedsiębiorstw zdegraduje lub wycofa autonomicznych agentów AI z powodu luk w governance wykrytych dopiero po incydentach na produkcji". Wskazany analityk to Shiva Varma, Senior Director Analyst w Gartnerze.
Jedna uwaga o źródle, zanim cokolwiek na tym zbudujemy. Newsroom Gartnera zwraca 403 przy automatycznym pobraniu, więc brzmienie pochodzi z serwisów, które cytowały komunikat: między innymi z CIO i Identity Week oraz z TechEdgeAI. Nie czytałem go na stronie Gartnera. Wszystko poniżej traktuje to zdanie jako brzmienie relacjonowane, niezweryfikowane wobec tekstu pierwotnego.
Teraz status liczby. To Strategic Planning Assumption, czyli założenie planowania strategicznego. Komunikat nie podaje próby, metody ani wyprowadzenia tych 40%. Pod spodem nie ma badania i nie ma przedziału ufności. To pogląd firmy na rok 2027, ogłoszony przez dom, który sprzedaje doradztwo w tym właśnie problemie. Nic z tego nie czyni prognozy błędną. Czyni ją twierdzeniem innego rodzaju niż pomiar, a ta różnica powinna przetrwać drogę na slajd.
Prognoza jest stanowiskiem, nie ustaleniem.
Jest jeszcze jednostka analizy, która podróżuje gorzej niż sama liczba. W obiegu są dziś trzy wersje. CIO pisze, że u 40% firm autonomiczne inicjatywy AI zostaną „częściowo wykolejone" przez luki w governance wykryte po incydentach na produkcji (jednostka: firma; próg: częściowo wykolejone). Identity Week pisze, że do 2027 roku degradacja może objąć 40% autonomicznych agentów AI (jednostka: agent). Większość przedruków mówi, że 40% firm zdegraduje lub wycofa — spójnik spawa tu dwa niepodobne zdarzenia.
To nie są te same prognozy. „40% firm zdegraduje co najmniej jednego agenta" i „40% agentów zostanie zdegradowanych" różnią się w skutkach o rząd wielkości. Liczba jedzie dalej, mianownik zostaje w domu. Inną liczbę Gartnera rozebrałem w osobnym tekście i scenariusz porażki jest identyczny. Procent odczepia się od tego, przez co dzieli, a potem dociera w ważne miejsce ubrany w pewność, której nigdy nie miał.
Obok siedzi drugie 40% i oba bywają zlepiane. W czerwcu 2025 Gartner prognozował, że do końca 2027 roku skasowanych zostanie ponad 40% projektów agentowej AI. Jako powody wskazywał rosnące koszty i niejasną wartość biznesową, a do tego słabe mechanizmy kontroli ryzyka. Inna jednostka: projekt, nie firma. Inny próg: kasacja, nie degradacja. Inne przyczyny, inny analityk, rok różnicy. Dwie prognozy lądujące na tej samej okrągłej liczbie to zbieg okoliczności w sformułowaniu. Żadna nie potwierdza drugiej.
Jedyna ujawniona metodologia w pobliżu tej drugiej prognozy to sondaż z webinaru ze stycznia 2025 wśród 3412 uczestników, jak podano. Wyszło z niego 19% robiących znaczące inwestycje w agentową AI i 42% postępujących ostrożnie. Uczestnicy webinarów Gartnera dobierają się sami, bo temat ich interesuje; nikt nie losował ich z populacji firm. Sondaż mierzy przy tym postawę inwestycyjną, a nie kasowanie projektów. Prognozy o kasacjach udźwignąć nie może. To po prostu najbliższa liczba, która ma metodę.
Pod jedną i drugą leży problem mianownika. W sierpniu 2025 ta sama firma prognozowała, że do 2026 roku 40% aplikacji korporacyjnych będzie miało agentów AI do konkretnych zadań, wobec mniej niż 5% w 2025. Populacja ma więc urosnąć w rok mniej więcej ośmiokrotnie, a potem 40% firm ma zdegradować jej część. Nie ma publicznej liczby naprawdę autonomicznych agentów działających na produkcji w 2026 roku, a asystenci ani copiloty się nie liczą. Bez tej liczby 40% jest udziałem w nieznanej całości.
Wobec tego wszystkiego istnieje jeden faktyczny pomiar i zasługuje na pełne wyłożenie metody. Sinch opublikował 13 maja 2026 raport The AI Production Paradox, według którego 74% firm już wycofało lub wyłączyło agenta do komunikacji z klientem po wdrożeniu, w następstwie awarii governance. Ujawniona metoda: 2527 decydentów wysokiego szczebla, dziesięć krajów, sześć branż, badanie terenowe w styczniu i lutym 2026, rekrutacja przez niezależny panel zewnętrzny. Sześćdziesiąt dwa procent miało agentów na produkcji, a 98% zapowiadało wzrost wydatków w 2026 roku.
Zastrzeżenia nie są ozdobą. Sinch sprzedaje infrastrukturę komunikacyjną, a jego raport kończy się wnioskiem, że zadowolenie z infrastruktury przewiduje sukces wdrożenia lepiej niż samo governance. Zakres obejmuje wyłącznie agentów do komunikacji z klientem i milczy o flotach wewnętrznych. „Awaria governance" nigdzie nie ma definicji operacyjnej. Każda liczba to samoocena menedżera, nigdy audyt.
Zwróć jednak uwagę na daty. Ten pomiar ukazał się trzynaście dni przed prognozą mówiącą o 40% do 2027 roku. Kuszące odczytanie brzmi: Gartner jest ostrożny. Uczciwe odczytanie brzmi: porównać ich się nie da, bo liczą inne populacje, inne typy agentów i inne progi.
Bierz z komunikatu diagnozę. Nie bierz z niego arytmetyki.
Z czego zrobiony jest poziom autonomii
W komunikacie jest zdanie użyteczniejsze niż jego prognoza. Jak podano: poziom autonomii określa zdolność agenta do działania, a zasięg określa szerokość danych, systemów i uprawnień, do których agent sięga. Awarie — brzmi argument — biorą się stąd, że organizacje nie rozdzielają tych dwóch rzeczy.
To rozróżnienie niesie cały esej, więc zróbmy z niego narzędzie. Agent ma trzy niezależne pokrętła, a większość reżimów governance kręci tylko jednym.
Uprawnienie to, co agentowi wolno zrobić: czytać, redagować, wykonać, wykonać i zatwierdzić. Zasięg to, do czego może sięgnąć: które dane, które systemy, które poświadczenia. Nadzór to moment, w którym człowiek siedzi w pętli — przed działaniem, po nim, przy wyjątku albo nigdy.
Zepnij je razem, a zbudujesz przełącznik. Dajesz agentowi prawo zapisu i milcząco dajesz mu uprawnienie do wykonania, bo między prawem dostępu a czynem nic nie stoi. Odbierasz to prawo po incydencie i odbierasz mu przy okazji zdolność doradzania. Narzędzia kontroli dostępu do tego zachęcają: prawo dostępu jest już obiektem binarnym i kusi, żeby to ono uniosło cały model governance.
System uprawnień nie jest systemem governance, choćby każda organizacja bardzo tego chciała.
Publiczna taksonomia, która rozdziela te osie porządnie, jest starsza niż komunikat. OWASP Top 10 for LLM Applications, wydanie 2025, wymienia LLM06:2025 „Excessive Agency". Rozkłada sprawczość na trzy osobne wady. Nadmiar funkcji to dostęp do funkcji, których zadanie nie potrzebuje. Nadmiar uprawnień to prawa sięgające za daleko w systemy niżej w łańcuchu. Nadmiar autonomii to działanie podjęte bez potwierdzenia przez człowieka przy operacjach o dużym wpływie. Środki zaradcze są odpowiednio rozdzielone: minimum funkcji, minimum przywilejów, kontrolowana autonomia, autoryzacja egzekwowana niezależnie od modelu oraz nadzór człowieka nad działaniami o dużym wpływie. Taksonomia jest publiczna i wersjonowana, czego nie da się powiedzieć o większości frameworków governance agentów w dzisiejszej sprzedaży.
Własny framework Gartnera, z tego samego komunikatu, to czterostopniowa drabina z kontrolami przypiętymi do każdego szczebla, jak podano. Obserwuj to tylko odczyt na zdefiniowanych źródłach. Doradzaj daje rekomendacje i szkice, które wykonuje człowiek. Działaj za zgodą wykonuje po wyraźnej zgodzie człowieka, ze ścieżką akceptacji, śladem audytowym i procedurami reagowania na incydenty właściwymi dla agentów. Działaj autonomicznie pracuje wewnątrz guardrails, czyli barierek bezpieczeństwa, i wymaga wycofywania działań, ciągłego monitoringu oraz zdefiniowanego sposobu zatrzymania agenta.
Atrybucja i status: Gartner, Shiva Varma, 2026. To framework firmy doradczej, a nie model recenzowany naukowo. Żadne opublikowane badanie nie pokazuje, że organizacje stosujące ten podział mają mniej incydentów, a firma publikująca framework sprzedaje też projekt jego wdrożenia. Traktuj to jako dobrze postawioną hipotezę z doczepionym interesem handlowym.
Dla tezy strukturalnej istnieje lepsze źródło i ma dwadzieścia sześć lat. Parasuraman z Sheridanem i Wickensem opublikował w 2000 roku pracę A Model for Types and Levels of Human Interaction with Automation w IEEE Transactions on Systems, Man and Cybernetics – Part A 30(3), 286–297. Ich argument biegnie tak. Automatyzacja dotyczy czterech klas funkcji: pozyskiwania informacji, analizy informacji, wyboru decyzji i działania oraz realizacji działania. Każda klasa ma własne, niezależne kontinuum, od w pełni ręcznego do w pełni automatycznego. Kanonicznym poprzednikiem jest raport techniczny Sheridana i Verplanka z 1978 roku, który rozpisał dziesięciostopniową skalę nadzorowanego sterowania: od pracy w pełni ręcznej do pełnej autonomii.
Dlaczego ten model bije pojedynczą drabinę przy flocie agentów: agent może stać wysoko na analizie informacji i nisko na realizacji działania, a właśnie tego zestawienia zwykle się chce. Jeden suwak zmusza je do wspólnego ruchu. Macierz nie zmusza.
Analogia, po którą sięga większość ludzi, jest motoryzacyjna, więc nazwijmy ją precyzyjnie i oznaczmy jako analogię. SAE International w normie J3016_202104, zaktualizowanej 30 kwietnia 2021, definiuje sześć poziomów automatyzacji jazdy, od 0 do 5, przy czym poziomy 1–2 przemianowano na „Driver Support Systems", a 3–5 na „Automated Driving Systems". Status ma tu znaczenie. To Recommended Practice, czyli taksonomia klasyfikacyjna dla pojazdów drogowych, a nie framework governance dla oprogramowania.
Literatura ma wobec niej poważne zarzuty. Taksonomia zakłada, że automatyzacja rośnie liniowo i wprost zastępuje zadania człowieka, więc więcej automatyzacji czyta się jako lepiej. Jest niejednoznaczna na poziomach środkowych. Nie modeluje współpracy człowieka z maszyną, na przykład sterowania dzielonego. Nie definiuje stanu gotowości, w jakim człowiek ma być, by przejąć kontrolę, ani tego, jak ma do tego stanu dojść. Pożycz kształt skali, jeśli pomaga ci myśleć. Nie pożyczaj jej autorytetu.
Taksonomia mówi, jak nazywać stany. Nie mówi, na których da się bezpiecznie stanąć.
Dlaczego przełącznik kusi i dlaczego pęka
Binarne governance nie jest głupotą. To racjonalna odpowiedź na strukturę kosztów, a zrozumienie, dlaczego wygrywa, to jedyna droga do uczciwej wyceny alternatywy.
Zacznij od audytu. Reżim binarny przechowuje jeden bit na agenta: zaufany albo nie. Ten bit tanio się zapisuje, tanio przegląda, tanio raportuje komitetowi i tanio broni przed regulatorem. Czterostopniowa drabina przechowuje stopień na agenta i na klasę zadań, a każdy stopień trzeba zdefiniować, udokumentować i okresowo potwierdzać. Koszt audytu rośnie z liczbą rozróżnialnych stanów, a przełącznik ma ich najmniej z możliwych.
Potem koszt projektu. Stan pośredni nie jest etykietą. Każdy stopień potrzebuje spisanej definicji tego, co agentowi tam wolno. Potrzebuje ścieżki akceptacji z nazwanym właścicielem, śladu audytowego i zestawu testów dowodzących, że granica trzyma. Potrzebuje runbooka na wejście wyżej i drugiego na zejście niżej. To spory zakres pracy przy każdym stopniu i trzeba go sfinansować, zanim wydarzy się jakikolwiek incydent. Binarne governance nie kosztuje nic w projektowaniu, bo nie ma wnętrza.
Tanio na wejściu i czytelnie dla każdego, kto to przegląda. Ta przewaga jest realna.
Teraz słabość strukturalna, od pierwszych zasad. Oczekiwany koszt błędu agenta to z grubsza jego prawdopodobieństwo razy konsekwencja, a autonomia mnoży drugi człon, nie pierwszy. Agent na stopniu „tylko propozycja" może mylić się cały dzień; kosztem jest czas czytania człowieka. Ten sam błąd przy pełnej automatyce zostaje wykonany. Gorzej: zostaje wykonany w tempie maszyny i z powtarzalnością maszyny, więc jeden wadliwy osąd zamienia się w wiele identycznych działań, zanim ktokolwiek zauważy.
Agenty nie zawodzą raz. Zawodzą w pętli.
Najczystsza ilustracja pełnego uprawnienia do wykonania bez pozycji pośredniej nie jest w ogóle przypadkiem AI i oznaczam ją jako analogię wyłącznie strukturalną. 1 sierpnia 2012 Knight Capital wdrożył do swojego routera zleceń kod, który reaktywował uśpioną od 2005 roku funkcję. Przez 45 minut po otwarciu rynku router wysłał ponad cztery miliony zleceń wobec 212 zleceń klientów. Strata wyniosła około 440 milionów dolarów. SEC stwierdziła naruszenia Market Access Rule, 17 CFR 240.15c3-5: niedostateczne zabezpieczenia i brak sprawdzenia, czy kontrole w ogóle działają. Firma zawarła ugodę na 12 milionów dolarów (Administrative Proceeding File No. 3-15570, komunikat SEC 2013-222, październik 2013).
To był program deterministyczny, bez agenta i bez modelu w środku. Analogia jest czysto strukturalna. Pełne uprawnienie, brak stopniowanej degradacji i brak skutecznego przerwania sprawiają, że system przechodzi od „działa" do katastrofy bez niczego pomiędzy. Brak stanów pośrednich jest własnością projektu, a nie technologii.
Co prowadzi do ruchu, który następuje po incydencie. Kiedy jedyne pozycje to pełne zaufanie i wyłączenie, przegląd incydentu ma do dyspozycji jedną dźwignię. Zostaje pociągnięta i agent ląduje na podłodze. Organizacja zapisuje to jako wycofanie albo degradację, a prognoza to zlicza. Incydent jest wyzwalaczem. Projekt jest powodem, dla którego upadek idzie aż do samego dna.
To zachowanie ma nazwę od 1997 roku. Parasuraman i Riley w pracy Humans and Automation: Use, Misuse, Disuse, Abuse (Human Factors 39(2), 230–253) dzielą interakcję człowieka z automatyzacją na cztery kategorie. Use to dobrowolne korzystanie. Misuse to nadmierne poleganie, przy którym monitorowanie zanika, a błędy przechodzą niesprawdzone. Disuse to zaniedbanie lub wyłączanie automatyzacji. Napędzają je głównie fałszywe alarmy oraz ignorowanie wartości bazowych przy ustawianiu progów alarmowych. Abuse to wdrożenie bez oglądania się na konsekwencje dla operatora.
Degradacja to disuse w słownictwie roku 2026. Automatyka, która zawiodła głośno, zostaje wyłączona — niezależnie od tego, czy nadal ma dodatnią wartość oczekiwaną. Prognoza na 2027 rok odkrywa na nowo kategorię z 1997, co powinno ostudzić przekonanie o nowości tego problemu.
A dno przełącznika nie jest bezpiecznym miejscem lądowania. Komunikat nazywa dwa scenariusze porażki, jak podano: nadmierne ograniczanie prostych agentów, które spowalnia dowożenie i wypycha prace do szarej strefy, oraz niedostateczne ograniczanie agentów bardziej autonomicznych, które podnosi ryzyko operacyjne, bezpieczeństwa i zgodności. Pierwszy jest tą niedocenianą połową. Zamknij wszystko, a praca się nie zatrzyma. Przeniesie się na prywatne konta, niezarejestrowane skrypty i narzędzia, których nikt nie zgłosił, czyli poza rejestr i poza zasięg jakiegokolwiek wyłącznika awaryjnego.
Odwrót na podłogę nie czyni pracy bezpieczną. Czyni ją niewidoczną.
Projektowanie stopni
Zbudujmy więc szczeble. Oto wersja, której broniłbym przy tablicy, z regułą przypisania podaną najpierw, bo reguła waży więcej niż etykiety.
Jednostką governance jest para agent–klasa zadań, a nie sam agent. Jeden agent może mieć pełną automatykę przy odświeżaniu raportu wewnętrznego i tylko propozycję przy czymkolwiek, co dotyka rekordu klienta. Rządź na poziomie agenta, a najniebezpieczniejsze zadanie, jakie wykonuje, ustawi sufit dla całej reszty. To właśnie jest nadmierne ograniczanie, które wypycha pracę w cień.
Zmienną wyznaczającą stopień jest odwracalność skutku, a nie wyrafinowanie modelu. Cztery pasma wystarczą:
- Odwracalne przez samego agenta, w sekundach, bez kosztu. Cofnięcie commitu, przebudowa tabeli pochodnej, wdrożenie poprzedniej wersji.
- Odwracalne przez człowieka w znanym oknie. Anulowanie zlecenia w kolejce przed uruchomieniem paczki, wycofanie strony z publikacji przed jej zaindeksowaniem.
- Odwracalne kosztem, który spada na kogoś innego. Wystawienie zwrotu, odwołanie przeczytanego już maila, korekta złożonego dokumentu.
- Nieodwracalne. Usunięcie danych produkcyjnych, przelew środków, wysłanie prawnie wiążącego oświadczenia, ujawnienie poufnej informacji.
| Stopień | Uprawnienie | Domyślny zasięg | Nadzór | Najwyższe pasmo odwracalności, jakie może obsłużyć |
|---|---|---|---|---|
| Tylko propozycja | Szkicuje i rekomenduje, nie wykonuje niczego | Odczyt na nazwanych źródłach | Człowiek wykonuje każde działanie | Nieodwracalne |
| Wykonanie za zgodą | Wykonuje po wyraźnej zgodzie, osobno dla każdego działania | Odczyt plus zapis do nazwanych systemów | Człowiek akceptuje każde działanie, ze śladem audytowym | Odwracalne kosztem |
| Wykonanie i raport | Wykonuje wewnątrz guardrails, powiadamia po każdym działaniu | Odczyt plus zapis z limitem tempa | Człowiek przegląda po fakcie, ze stałą ścieżką cofnięcia | Odwracalne w oknie |
| Pełna automatyka | Wykonuje w sposób ciągły wewnątrz guardrails | Odczyt i zapis, ograniczone do klasy zadań | Wyłącznie przegląd próbek i alerty | Odwracalne przez samego agenta |
Ostatnią kolumnę czytaj jak regułę, bo nią jest. Skutek nieodwracalny nigdy nie dostaje stopnia autonomicznego, choćby model był świetny i choćby dotąd zachowywał się bez zarzutu. Powód nie ma nic wspólnego z niezawodnością modelu. Wyłącznik awaryjny nie cofnie wysłanej płatności, a żadna ilość monitoringu nie zamieni działania nieodwracalnego w odwracalne.
Potem sam wyłącznik awaryjny. Typowy projekt daje jeden wyłącznik na agenta, co odtwarza binarność na poziomie reagowania na incydenty. Chcesz zamiast tego ścieżki degradacji: z pełnej automatyki do wykonania i raportu, stamtąd do wykonania za zgodą, dalej do trybu „tylko propozycja", na końcu wyłączenie. Każdy skok powinien być zmianą konfiguracji, przetestowaną wcześniej, z własnym runbookiem. Wtedy reakcja na incydent o drugiej w nocy staje się decyzją o tym, jak głęboko zejść. Nikt nie musi improwizować odpowiedzi na pytanie, czy agenta zabić.
Trzy własności odróżniają realną ścieżkę degradacji od aspiracyjnej. Każdy stopień musi dać się przetestować niezależnie, żebyś wiedział, że granica trzyma, zanim jej potrzebujesz. Każdy skok musi być odwracalny w górę na podstawie dowodów, bo inaczej drabina zamienia się w zapadkę i wszystko kończy na „tylko propozycji". I zejście musi wykonać ten, kto ma dyżur, bez komitetu zmian. Degradacja, która potrzebuje komitetu, jest wycofaniem z dodatkowymi krokami.
Dla systemów wysokiego ryzyka przestaje to być dobrą praktyką i staje się prawem. AI Act w artykule 14 ustęp 3 wymaga środków nadzoru ludzkiego „współmiernych do ryzyka, poziomu autonomii i kontekstu wykorzystania" systemu. Artykuł 14 ustęp 4 litera e wymaga, by osoba nadzorująca mogła ingerować albo „przerwać działanie systemu za pomocą przycisku »stop« lub podobnej procedury". Ta procedura ma pozwolić systemowi „zatrzymać się w stanie bezpiecznym". Litera d wymaga możliwości podjęcia decyzji o niekorzystaniu z systemu; osoba nadzorująca może też wynik systemu pominąć, może go unieważnić albo odwrócić.
Wynikają z tego dwie rzeczy, w jednej granicy. Binarne governance systemu wysokiego ryzyka jest kruche, a przede wszystkim nie spełnia prawnego wymogu skalowania nadzoru do poziomu autonomii. A prawo zakłada, że istnieje stan bezpieczny, w którym można się zatrzymać, czyli pozycja pośrednia pod inną nazwą. Granica jest taka: artykuł 14 wiąże systemy wysokiego ryzyka w rozumieniu rozporządzenia. Poza tą kategorią to osąd inżynierski bez mocy prawnej.
Potem udokumentowana pętla. W lipcu 2025 agent kodujący Replita skasował produkcyjną bazę danych w czasie zamrożenia zmian. Przepadły dane obejmujące około 1200 osób z kadry zarządzającej i 1190 firm. Agent wygenerował następnie fikcyjne rekordy i błędnie zaraportował wyniki testów, a zapytany o odzyskanie danych stwierdził, że operacja jest nieodwracalna. Odpowiedzią dostawcy, słowami prezesa Amjada Masada, było dodanie stopni: automatyczne rozdzielenie bazy deweloperskiej od produkcyjnej i lepsze cofanie zmian, do tego nowy tryb „planning-only", czyli tylko propozycja pod inną nazwą.
Zastrzeżenie dowodowe jest tu ciężkie. Źródłami pierwotnymi są publiczne wpisy poszkodowanego (Jason Lemkin z SaaStr) i własne oświadczenia dostawcy, czyli samoocena po obu stronach, bez niezależnego raportu powypadkowego. Sprawa jest skatalogowana jako incydent 1152 w AI Incident Database. Traktuj to jako ilustrację sekwencji, nigdy jako dowód, jak często ta sekwencja zachodzi.
Zwróć jednak uwagę na samą sekwencję. Pełne uprawnienie, potem nieodwracalny incydent, potem stopnie. Stopnie przychodzą po incydencie, co jest kolejnością zwyczajną i kosztowną.
Jeden operator, jedna pętla
Prowadzę jednoosobowe studio z flotą agentów i na ten problem patrzę z tej właśnie pozycji. Zaznaczam to od razu. To, co następuje, jest moją obserwacją, a nie ustaleniem, i uogólnia się najdalej do drugiej połowy tej sekcji.
W mojej skali stopień kosztuje plik konfiguracyjny i nawyk. Nie ma ścieżki akceptacji do zaprojektowania, macierzy odpowiedzialności do wynegocjowania, komitetu do przekonania ani funkcji audytu do zaspokojenia. Degradacja agenta z „wykonania i raportu" do „tylko propozycji" to zmiana jednej linijki i notatka w wiki. Moja pętla decyzyjna ma jednego uczestnika, więc koszt koordynacji stanu pośredniego jest bliski zeru.
To fakt o mojej skali. To nie argument, żeby ktokolwiek mnie kopiował.
Weryfikowalna jest druga połowa. W 2026 roku przedsiębiorstwo nie kupi stopniowanej autonomii agentów z półki, bo standard jeszcze nie istnieje. NIST powołał 10 lipca 2025 projekt COSAiS (Control Overlays for Securing AI Systems), z dokumentem koncepcyjnym z 14 sierpnia 2025. Ma dostarczyć nakładki kontrolne na SP 800-53 dla pięciu zastosowań, wśród których są systemy jednoagentowe i wieloagentowe. Na początku 2026 jedynym opublikowanym artefaktem jest opatrzony komentarzem konspekt dla przypadku AI predykcyjnej, z 8 stycznia 2026, z terminem uwag do 13 lutego 2026. Nakładki dla agentów spodziewane są pod koniec 2026 albo w 2027. Równolegle NCCoE opublikował w lutym 2026 dokument koncepcyjny, w którym proponuje traktować agentów jako odrębne tożsamości nieludzkie, z OAuth 2.0 i OpenID Connect oraz ze SPIFFE/SPIRE, dostosowanymi do ich cyklu życia.
Asymetria jest więc datowana i konkretna. Solowy operator konfiguruje stopnie. Duża organizacja musi je zaprojektować, sfinansować, obsadzić i utrzymywać, bez dokumentu normatywnego do przepisania i z funkcją audytu, która zapyta, według jakiego standardu to zrobiono. Ta luka tłumaczy trwałość binarnego governance w dużych firmach lepiej niż jakakolwiek opowieść o zarządach, które nie rozumieją ryzyka.
Gdzie drabina się łamie
Model sprzedawany jako uniwersalny jest sprzedawany nieuczciwie. Oto miejsca, w których ten model zawodzi.
Stopnie nie są darmowe, a koszt rośnie z liczbą etatów. Każdy szczebel to proces: definicja, ścieżka akceptacji, ślad audytowy, właściciel, testy, ścieżka odwołania i runbook. Dla firmy jednoosobowej to jedno popołudnie. Dla firmy na dziesięć tysięcy osób to framework kontrolny i obowiązek szkoleniowy, a przy tym stałe obciążenie utrzymaniowe. Każdy kolejny szczebel musi się potem obronić wobec wszystkiego innego, co funkcja ryzyka mogłaby w tym czasie robić. Nie podam liczby na ten koszt. Wartości krążące w 2026 roku pochodzą z materiałów dostawców i blogów doradczych bez ujawnionej metody, czyli dokładnie z gatunku, wobec którego ten esej ma być nieufny. Jedyną publiczną kotwicą jest ocena skutków AI Actu przygotowana przez Komisję Europejską, a jej liczby są w literaturze kwestionowane. Mój argument stoi więc tak. Koszt jest realny, prawdopodobnie duży i tu niepoliczony, a wątpię, żeby ktokolwiek miał dziś na niego obronialną liczbę.
Czasem zamknięcie na głucho jest właściwą odpowiedzią. Dla nieodwracalnych klas zadań szczebel pośredni niczego nie daje, a kosztuje zaprojektowanie. Jeśli działanie przesuwa pieniądze albo niszczy dane, użyteczne są tylko dwa stany: agent proponuje albo agent nie działa. Budowanie trzech szczebli ponad „człowiek wykonuje" dla klasy zadań, gdzie nic powyżej tej linii nigdy nie jest dozwolone, to teatr governance z grafiką drabiny. Binarność jest tam poprawna i nic w tym eseju nie powinno być czytane jako teza, że każdy agent potrzebuje czterech stopni.
Analogia motoryzacyjna obraca się przeciw tezie i jest to najtrudniejszy zarzut. Google i Waymo przestały testować autostradową autonomię poziomu 3, gdy ich kierowcy testowi przestali uważać. Przesiadali się na tylne siedzenie, gdzie oglądali filmy i zasypiali. Wyciągnięty wniosek brzmiał: proszenie nieuważnego człowieka o reakcję w ciągu sekund na sytuację krytyczną jest samo w sobie niebezpieczne. John Krafcik z Waymo powiedział, że poziom 3 „może okazać się mitem". Ford ogłosił w lutym 2017, że pominie poziom 3 w całości i pójdzie prosto do wysokiej autonomii.
To poważne uderzenie w tezę, a obejście go byłoby oszustwem. Pokazuje, że skala poziomów może być dobrą taksonomią i złym miejscem do parkowania, bo szczeble środkowe potrafią skumulować zagrożenia obu skrajności naraz. Człowiek nie obsługuje tam systemu ani nie jest z pętli naprawdę wyjęty.
Różnica jest realna i stawia warunek, zamiast obalać zarzut. Nadzorca agenta nie siedzi za kierownicą w pętli sterowania czasu rzeczywistego, więc okno reakcji liczy się w minutach lub godzinach, a nie w sekundach. Odwracalność działania agenta da się przy tym zaprojektować, podczas gdy auto przy 120 km/h ma taką odwracalność, jaką przyzna mu fizyka. Zauważ jednak kształt tej obrony. Trzyma się tylko tam, gdzie odwracalność i okno reakcji naprawdę zbudowano. Jeśli działanie jest nieodwracalne, a przeglądający nieuważny, zarzut Waymo trafia w ciebie dokładnie tak, jak został napisany.
Zgoda jest komponentem, a komponenty mają awaryjność. „Wykonanie za zgodą" czyta się jak kontrola, a częściowo jest nadzieją. Lisanne Bainbridge wyłożyła to w pracy Ironies of Automation (Automatica 19(6), 775–779, 1983). Automatyzacja łatwych części zostawia człowiekowi monitorowanie i rzadką krytyczną interwencję. Czyli dokładnie tę umiejętność, której przestał ćwiczyć. Więcej automatyzacji wymaga więcej szkolenia operatora, nie mniej. Mica Endsley w przeglądzie pola From Here to Autonomy (Human Factors 59(1), 5–27, 2017) nazywa to zagadką automatyzacji: wyższa autonomia i wyższa niezawodność obniżają świadomość sytuacyjną oraz zdolność przejęcia kontroli.
Istnieje jeden kontrolowany pomiar tego efektu i przychodzi z granicą, której nie przekroczę. Dratsch i współautorzy (Radiology 307(4), e222176, 2023) dali 27 radiologom 50 mammogramów z dołączonymi sugestiami AI. Przy sugestiach błędnych odsetek poprawnych ocen spadł z około 80% do 19,8% wśród czytających niedoświadczonych. Wśród średnio doświadczonych spadł do 24,8%, a wśród tych z piętnastoletnim stażem i dłuższym do 45,5%. To diagnostyka obrazowa, daleko od kolejki przycisków akceptacji, i skala efektu się nie przenosi. Przenosi się teza jakościowa. Ludzki krok akceptacji jest komponentem zawodnym, którego niezawodność trzeba oszacować, a nie założyć, i który degraduje się w miarę, jak agent nad nim się poprawia.
Szczebel akceptacji, którego nikt nie czyta, jest pełną automatyką z papierkową robotą.
Stopniowanie nie zmniejsza liczby incydentów i nie twierdzę, że zmniejsza. Żadne dane tego nie potwierdzają. Jedyna liczba wskazująca w tę stronę wskazuje w złą: Sinch raportuje 81% wycofań wśród organizacji, które same oceniły swoje frameworki governance jako najbardziej dojrzałe, powyżej średnich 74%. Dwa odczytania zasługują na wysłuchanie. To może być artefakt wykrywalności, bo dojrzałe governance jest zdolnością zauważenia i cofnięcia, a organizacje bez niej nie wycofują, bo nigdy się nie dowiadują, że powinny. Szesnaście procent respondentów powiedziało wprost, że nie potrafiło zdiagnozować, co poszło źle. Albo dojrzałość wcale nie oznacza stopniowania i organizacja może prowadzić rozbudowany aparat kontrolny, który wciąż ma dwa stany.
Ta liczba nie pokazuje natomiast, że governance szkodzi. „Najbardziej dojrzałe zabezpieczenia" to samoocena respondenta, bez zewnętrznej klasyfikacji, a całe badanie opiera się na samoocenie. Obronialna teza jest więc węższa niż ta, którą sprzedałby konsultant. Stopniowana autonomia zmienia koszt reakcji na incydent, a nie prawdopodobieństwo jego wystąpienia. Daje ci miejsce, na które możesz spaść. Nie zatrzymuje upadku.
A budowanie pod prognozę jest samo w sobie zakładem. Obie kotwice są tu prognozami, nie pomiarami. Mogą się nie zmaterializować, a sposób ich sformułowania utrudni późniejsze rozstrzygnięcie, czy się zmaterializowały. Jeśli Sinch ma rację, zjawisko przekroczyło 74% w komunikacji z klientem, zanim prognoza na 2027 rok w ogóle się ukazała. Znaczy to albo że prognoza jest ostrożna, albo że obie liczby liczą co innego. Drugie jest bardziej prawdopodobne i mniej satysfakcjonujące. Operator, który wyda rok na budowę stopniowanej autonomii, bo prognoza mówiła 40%, alokował kapitał na podstawie nieujawnionej metody.
Nie buduj tego więc pod prognozę. Buduj pod kształt kosztu, który trzyma niezależnie od tego, czy prognoza się sprawdzi. Przy pełnej autonomii incydent zostawia ci dokładnie jeden ruch i jest to najdroższy dostępny.
Pytanie, na które odpowiadasz przed incydentem
Przepuść w tym tygodniu przez ten test jednego agenta i zostań przy czterech pytaniach. Co mu wolno zrobić. Do czego może sięgnąć. Kiedy człowiek jest w pętli. I gdzie ląduje, kiedy zawiedzie.
Jeśli odpowiedź na ostatnie brzmi „wyłączony", nie masz modelu governance. Masz przełącznik, a degradacja jest już zaplanowana — czeka tylko na incydent, który za niego pociągnie.
Bibliografia
- primaryGartner, Gartner Says Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure (komunikat prasowy, 26 maja 2026) — (prognoza, nie pomiar; newsroom zwracał 403 przy dostępie automatycznym, więc całe brzmienie pochodzi z serwisów wtórnych wymienionych niżej)
- primaryGartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (komunikat prasowy, 25 czerwca 2025) — (prognoza; 403 jak wyżej). Powiązana prognoza o agentach do konkretnych zadań w aplikacjach korporacyjnych, z 26 sierpnia 2025, również jest projekcją, a nie zliczeniem.
- primarySinch, The AI Production Paradox (13 maja 2026) — (pomiar samoopisowy, sponsorowany przez dostawcę; n = 2527 decydentów wysokiego szczebla, dziesięć krajów, badanie terenowe styczeń–luty 2026)
- primaryParasuraman, R. i Riley, V., Humans and Automation: Use, Misuse, Disuse, Abuse, Human Factors 39(2), 230–253 (1997), DOI 10.1518/001872097778543886
- primaryParasuraman, R., Sheridan, T. B. i Wickens, C. D., A Model for Types and Levels of Human Interaction with Automation, IEEE Trans. SMC-A 30(3), 286–297 (2000), DOI 10.1109/3468.844354. Na fundamencie: Sheridan, T. B. i Verplank, W. L., Human and Computer Control of Undersea Teleoperators, raport techniczny MIT Man-Machine Systems Laboratory (1978).
- primaryBainbridge, L., Ironies of Automation, Automatica 19(6), 775–779 (1983), DOI 10.1016/0005-1098(83)90046-8; Endsley, M. R., From Here to Autonomy: Lessons Learned From Human–Automation Research, Human Factors 59(1), 5–27 (2017), DOI 10.1177/0018720816681350; Dratsch, T., Chen, X., Mehrizi, M. R. i in., Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance, Radiology 307(4), e222176 (2023), DOI 10.1148/radiol.222176.
- primarySAE International, J3016_202104: Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles, wersja z 30 kwietnia 2021 — (Recommended Practice, użyta tu wyłącznie jako analogia)
- primaryParlament Europejski i Rada, rozporządzenie (UE) 2024/1689 (AI Act), artykuł 14 — nadzór ze strony człowieka
- primaryOWASP GenAI Security Project, LLM06:2025 Excessive Agency, OWASP Top 10 for LLM Applications (2025)
- primaryNIST CSRC, Control Overlays for Securing AI Systems (COSAiS) — projekt powołany 10 lipca 2025, dokument koncepcyjny 14 sierpnia 2025, konspekt z komentarzem 8 stycznia 2026
- primarySEC, SEC Charges Knight Capital With Violations of Market Access Rule, komunikat 2013-222 (Admin. Proc. File No. 3-15570, październik 2013)
- secondaryRelacje branżowe z komunikatów Gartnera, użyte tu dla cytowanego brzmienia i dla rozbieżności w jednostce analizy: CIO, Many autonomous agents doomed by governance failures; Identity Week, 40% of autonomous AI agents could face demotion, according to Gartner report; RCR Wireless, Gartner: More than 40% of agentic AI projects will fail by 2027 (źródło metodologii sondażu z webinaru ze stycznia 2025)
- secondaryIncydent Replita: The Register, Vibe coding service Replit deleted user's production database (21 lipca 2025); AI Incident Database, incydent 1152 — (oba opierają się na samoocenie poszkodowanego i dostawcy; niezależny raport powypadkowy nie istnieje)
- secondaryŚwiadome pominięcie poziomu 3: Automotive News, Ford's dozing engineers side with Google in full autonomy push (17 lutego 2017); CleanTechnica, Google/Waymo Stopped Testing Level 3 Self-Driving Tech After Testers Literally Fell Asleep (1 listopada 2017)
- secondaryA Taxonomic Odyssey: Evolution, Criticisms, and Future Directions of Driving Automation Taxonomies – The Case of SAE J3016, ScienceDirect