AI · MBA
AI-ready data is a power problem
Every AI programme eventually hits the same wall, and the wall is always described in engineering language. The data is not ready. In February 2025 Gartner forecast that through 2026, organisations would abandon 60% of AI projects unsupported by AI-ready data — a prediction, not a post-mortem. Read that as an engineering problem and you buy pipelines, catalogues and quality dashboards. Read it as a problem of ownership and you get a different budget and a different set of arguments.
This essay makes one argument, and it is narrow. Part of data readiness really is technical. But the part that kills projects is usually a question of property: who owns the data, whose budget pays for its quality, and whose position weakens when it flows to someone else. Silos are not a failure of discipline. They are a reasonable response by people measured on something else.
Your data quality problem has a job title and a bonus scheme.
Read the numbers before you quote them
The 60% figure travels badly. Gartner's press release of 26 February 2025 says that "through 2026, organizations will abandon 60% of AI projects unsupported by AI-ready data." Three things fall off in transit. It is a forecast issued in early 2025 about a window ending in 2026, and nobody has published a check against outcomes. The denominator is projects unsupported by AI-ready data, not all AI projects. The popular version, "60% of AI projects will fail," is a different and false claim. And the release does not publish the model behind the prediction. One honesty note on sourcing: Gartner's newsroom blocks automated access, so I am working from outlets that quoted the release verbatim rather than from the primary page.
The companion number is more interesting, and more often mangled. Gartner reports that 63% of organisations "either do not have or are unsure if they have the right data management practices for AI." That comes from a survey of 248 data management leaders in the third quarter of 2024. So: a self-assessment, from a narrow professional subpopulation, about their own organisations. Not an audit of practice, and not a view from the board.
And it welds together two states that are not the same thing. "We do not have the practices" is a resourcing answer. "We are not sure whether we have them" is an answer about who is accountable for knowing. Gartner does not split the two, and no percentage breakdown exists in the source. That matters, because the second half of the sentence is where this essay lives. Building an argument on it is my interpretation, not a Gartner finding. An organisation that cannot say whether it has data management practices has told you something precise about ownership.
One definitional point sits underneath all of this, and it is mine. "Ready data" in general does not exist. Data is ready or not ready for a particular use case, against the requirements of a particular technique, and only until the use case changes. Readiness is a relation, not a property. That has an organisational consequence people skip: you cannot centralise accountability for a moving relation into one horizontal programme, because the programme has no stable object to manage. Whoever owns "data quality" in the abstract owns a noun that changes shape whenever someone else's project changes shape.
Which is roughly how a programme ends up owning a catalogue instead of an outcome.
Information was always the currency
The mechanism is not new, and the best description of it predates the current wave by more than three decades. Thomas Davenport, Robert Eccles and Laurence Prusak published "Information Politics" in MIT Sloan Management Review on 15 October 1992, based on two years of fieldwork across more than 25 companies. Their finding runs against the era's optimism. Information technology was supposed to free the flow of information and flatten the hierarchy. It did close to the opposite. Information became the currency of organisational life, too valuable for most managers to give away.
Their frame is a taxonomy of five information-political models: technocratic utopianism, anarchy, feudalism, monarchy and federalism. Feudalism is the familiar one, where business units run their own information estates with their own vocabularies. The authors describe the firm as a set of city-states, each with its own culture, leadership and language. Federalism, meaning negotiated common terms across units that retain real autonomy, is their realistic goal, not monarchy.
The model worth naming today is the first. Technocratic utopianism is the belief that a sufficiently good technical programme will dissolve an organisational problem. Build the warehouse, model the enterprise, standardise the definitions, and the politics will follow. In 1992 that meant enterprise data modelling. Today it means "AI-ready data" as a platform initiative with a delivery date.
The age of the source is an argument, not a weakness. The same behaviour survived the data warehouse, master data management, the data lake, the lakehouse and data mesh. Five technical answers, one unchanged organisational pattern. A constraint that outlives every solution aimed at it is probably not in the layer those solutions address.
If the tool keeps changing and the outcome does not, the tool was not the variable.
Why the silo is rational
Assume, for the rest of this section, that everyone involved is competent and acting in good faith. The silo still forms. That is the interesting case, and in my experience it is also the common one.
Start with power, defined operationally rather than morally. Hickson, Hinings, Lee, Schneck and Pennings set out a strategic contingencies theory of intraorganisational power in Administrative Science Quarterly in 1971. A subunit's power rises with three things: how well it copes with uncertainty for others, how central it is to the workflow, and how hard it is to substitute. Uncertainty here has a specific meaning: a lack of information about future events.
Now put data into that model. Data is precisely the resource that reduces someone else's uncertainty. A unit holding data that others need is, by the theory, powerful in proportion to that need. Hand the data over cleanly, in a form others can use without asking, and you have reduced their dependence on you. You have also made yourself easier to substitute.
Inside this model, sharing data is unilateral disarmament.
Then the cost side, which is where the incentive problem becomes exact. Jensen and Meckling's 1976 paper in the Journal of Financial Economics defined the agency relationship. When one party acts on another's behalf and interests diverge, you get monitoring costs, bonding costs and a residual loss. Their subject was shareholders and managers, so applying it to departments is an extension of the model, not their claim.
The extension holds because the structure is the same. A data owner is an agent of the firm and a principal of their own budget. Bringing a source system to the quality another team's model needs costs their engineers and their roadmap slot. The benefit accrues to a project owned by a different director and measured on a metric in someone else's part of the P&L. Nothing in that arrangement is irrational on either side. The refusal is the predicted output.
Mancur Olson explained why nobody volunteers to pay. In The Logic of Collective Action (Harvard University Press, 1965) he set out the arithmetic of concentrated costs and diffuse benefits. His formulation: "the larger a group is, the farther it will fall short of obtaining an optimal supply of any collective good." Data quality is a collective good for every downstream consumer, funded by one upstream producer. Olson's prediction is chronic undersupply, worsening with organisational size. That is what data quality budgets look like in practice. His remedy is a selective incentive: a benefit available only to the party bearing the cost.
Now the shape of the failure. The usual metaphor gets it backwards. People reach for Garrett Hardin's tragedy of the commons (Science, 1968) to describe data silos. Hardin describes overuse of a shared resource that has no owner. A silo is the opposite: underuse of a resource that has too many owners. The right model is Michael Heller's, from "The Tragedy of the Anticommons" in the Harvard Law Review (1998). When many parties hold rights of exclusion and nobody holds an effective privilege of use, the resource is systematically underused.
Heller's opening observation is worth carrying. After communism fell, storefronts in Eastern European cities sat empty while street kiosks multiplied in front of them. Demand existed. What did not exist was a path through the several agencies and entities that each held a veto over the lease. That is the architecture of a corporate data estate: a surplus of veto rights, a deficit of use rights.
Getting this right changes the prescription. For a commons, you add ownership. For an anticommons, you must bundle or waive rights of exclusion. So the standard governance response (convene a council, give every domain a seat, require sign-off) adds veto rights to a system already dying of them. It is the correct remedy applied to the wrong problem, and it makes the problem worse.
Elinor Ostrom's Governing the Commons (Cambridge University Press, 1990) supplies the constructive half, and it also kills the lazy conclusion. Her case work showed that shared resources are often governed successfully without privatisation and without a central authority, following a set of design principles. The second of those principles is the operational one here: benefits should be proportional to costs. Anyone paying for a shared resource must receive a share of the value tied to what they paid. Read that as an instruction and it produces a budgeting rule, which I will come back to.
The silo is the organisation chart working.
Governance without decision rights is theatre
The definition of data governance sharp enough to detect theatre comes from the academic literature, not from a vendor. Vijay Khatri and Carol Brown, writing in Communications of the ACM in 2010, define governance as the allocation of decision rights and accountabilities for data-related processes. They separate governance (which decisions must be made, and who makes them) from management, which is who executes.
That gives a test you can apply in a single meeting. Does the programme move any decision right from one person to another? Does it change any line in anyone's annual objectives? If both answers are no, the programme is management wearing governance as a costume. Catalogues and committees without transferred decision rights meet that description exactly.
It also explains the specific way these programmes decay. A catalogue is documentation of an asset someone else owns. Keeping an entry current is a private cost paid by the source team for a benefit spread across consumers they will never meet. That is Olson's structure again, one level down, at the granularity of a single table description. The undersupply is predicted by the same model that predicts the underfunding. One gap deserves flagging. I could not find a credible public study measuring how many enterprise catalogues fall out of maintenance; the claims that circulate come from firms selling catalogues. Treat the mechanism as reasoned, not measured.
The same test explains the data steward with no authority. The role is usually designed as accountability without decision rights: responsible for quality, unable to set a source system's roadmap. That is a structural impossibility handed to a person.
One proxy for how that goes needs handling with care. Bean and Davenport reported in Harvard Business Review in 2021, as summarised by MIT Sloan the following year, that the average chief data officer tenure runs about two and a half years. The comparison figures are roughly seven years for a CEO and four and a half for a CFO or CIO. More recent industry surveys land in the same region, with over half of respondents reporting under three years. The reasons given centre on an unclear mandate and a role carved out of the CIO's territory, with transformation expected inside about eighteen months. Three cautions. These are self-report surveys of data leaders, not a registry of appointments. MIT Sloan does not name a lack of budget or formal authority as the principal cause, and I am not going to put that in their mouths. And tenure is a proxy for a hard job, not proof of a mechanism.
Gartner's own forecast on governance is the sharpest thing they have published on this, and it is rarely read closely. In February 2024 they predicted that by 2027, 80% of data and analytics governance initiatives will fail "due to a lack of a real or manufactured crisis." Saul Judah, a VP Analyst, is quoted as saying that a governance programme which does not enable prioritised business outcomes fails. The recommendation is to move away from a centre-out, command-and-control posture toward tangible business outcomes. Again: a forecast, not a measurement, and again quoted at second hand.
Sit with the word manufactured. An advisory firm is openly recommending that you produce a crisis, because governance without one does not take. That is an admission about what these programmes actually run on: attention and priority. Both are political currencies, allocated by people with the standing to allocate them. Gartner does not describe this as a power problem. That reading is mine, and I think it is the only honest one available.
Which returns us to the 63%, and to the half of it that says "we are not sure." An organisation gives that answer when nobody owns the question. Owning a question means being the person whose review includes the answer. Where no such person exists, the honest response to an auditor is a shrug, and the shrug gets recorded as a data management finding. It is really a finding about accountability.
A political diagnostic, run before the project
None of this is useful until it changes what you do in week one. So here is the procedure I use. It is my own synthesis, not a Gartner method or a theorem, and it should be read as a framework rather than a measurement.
Five questions, asked before anyone writes an integration ticket.
Who owns each data set the project needs? Not which system holds it — which named person controls the roadmap of the system that produces it. If you cannot name a human with a budget, you have found the first problem.
What does the handover cost them this year? In their engineering hours and their quarterly commitments. Get an estimate from them, not from your architect. The number is usually larger than the requesting team assumes, and that gap is where projects stall.
Who is worse off if this project succeeds? In scope, headcount, budget line, or in owning a metric they currently get to explain. Someone almost always is. Resistance from that quarter is information about the distribution of power, not evidence of stupidity.
Which decision right would have to move, and who signs that? Name the right: approving schema changes, or granting access without case-by-case review. If no right moves, expect no flow.
Is the legal and technical path already open? If permission exists, the interface exists, and the data still does not move, you are looking at a political constraint. You have just diagnosed it without a workshop.
The table below is how I keep the two readings apart in practice.
| Symptom | Engineering reading | Ownership reading | The test that separates them |
|---|---|---|---|
| "The data quality isn't good enough" | Cleansing and validation backlog | Nobody's budget funds quality that benefits someone else | Ask whose P&L pays for the fix and whose gets the benefit |
| "Access takes months" | Missing interfaces, security tooling | Rights of exclusion held by several parties, no privilege of use | Count the parties who can say no; count those who can say yes |
| "We don't know what data we have" | Catalogue coverage is incomplete | Nobody is accountable for knowing | Ask who is evaluated on the answer |
| "The source system can't export that" | Legacy without an API | True constraint until proven otherwise | Ask whether a funded remediation plan exists — and survived last year |
| "Governance is in place" | Policies and a committee | No decision right has moved | Name one right that changed hands, in writing |
| "There's no budget for it" | Genuine cost ceiling | Budget exists and is reallocated annually | Check whether the line was approved and then spent elsewhere |
Then three moves. They are deliberately small, because large ones require authority a project lead usually does not have.
Put the data quality budget inside the AI project budget, not beside it. This is Ostrom's proportionality principle applied directly. A separate data quality programme asks a source team to supply a collective good for free. A line item inside the project budget converts it into a purchased service with a customer, a price and a delivery date. The same work, moved from charity to trade. It also puts an honest number on the project, which some sponsors will not enjoy.
Put the data owner inside the project team, with a share of the outcome. Not on a steering committee, which is a place where accountability goes to be diluted. Skin in the game means their objectives contain a piece of the project's result. This is Olson's selective incentive in its smallest usable form: a benefit reaching only the party bearing the cost.
Move exactly one decision right, and write it down. One is enough to test whether the organisation is willing. If nothing can move, you have learned something important early. The programme will produce documentation and no flow, and you can decide whether to fund it on that basis.
The nearest thing to supporting evidence I can offer comes with heavy caveats. Gartner reported in April 2026 that organisations with self-reported AI success invest up to four times more, as a percentage of revenue, in foundations. Their list of foundations covers data quality, governance, AI-ready people and change management. The survey ran across 353 data and analytics and AI leaders in November and December 2025. In the same release, Rita Sallam is quoted saying only 39% of technology leaders are confident current AI investment will improve financial results. Now the caveats, which matter more than the headline. Success is self-assessed. Causality is unresolved: success can fund foundations as easily as the reverse. And "up to four times" is an upper bound, not an average. What I take from it is the composition of that list of foundations, which puts people and change management alongside data quality. Those are organisational line items, and they are sitting in the same sentence as the technical one.
The argument connects to a broader point about where AI value hides. The return from an AI system rarely sits in the model; it lives in the links around it. Data ownership is the least visible of those links, because it appears in no architecture diagram.
Where the thesis is wrong
A thesis that explains every failure explains none, and "it's a power problem" can be stretched to cover anything. That makes it comfortable, and comfortable arguments deserve the hardest test I can build.
Sometimes the constraint really is that the data does not exist. RAND's 2024 report on the root causes of failure for AI projects lists insufficient data among its five. The organisation simply does not hold the data needed to train an effective model. Two rigour notes. The study is qualitative, based on 65 interviews with data scientists and engineers with at least five years of experience, so 65 is a count of people, not of projects. And the report's ">80% of AI projects fail" line is quoted from other estimates, not measured by RAND.
That root cause is a hard boundary for everything above. Data never captured at the required granularity will not appear because you fixed the incentives. Better incentives can start the capture today. That does nothing for a project needing twenty-four months of history, and no amount of stakeholder mapping shortens that wait.
Sometimes the constraint is tedious engineering that nobody wants to do. Sculley and colleagues made the case in "Hidden Technical Debt in Machine Learning Systems" at NIPS 2015, and it has aged well. The model code is a small fraction of a real production system; the rest is collection, feature extraction, serving and monitoring. Data dependencies cost more to maintain than code dependencies and receive less engineering rigour. Their CACE principle (changing anything changes everything) describes why these systems resist clean decomposition. That work is expensive and unglamorous. It is also exactly the work that gets relabelled "an organisational problem" by people who would rather run a workshop than a migration.
I want to be blunt about that failure mode, because this essay supplies the vocabulary for it. A political diagnosis is an excellent alibi. It converts a hard engineering backlog into someone else's character flaw, and it lets a team spend a quarter on stakeholder maps while the pipeline stays broken. If reading this makes you feel better about not doing the pipeline work, you have used it wrong.
Sometimes the money genuinely is not there. The UK National Audit Office reported in January 2025 on the state as of March 2024. Government departments ran around 228 significant legacy systems, of which 63 (28%) were red-rated for high likelihood and high impact of operational or security risk. For 120 of those 228, more than half, no fully funded remediation plan existed. Separately, the DSIT and Government Digital Service review of digital government, also published in January 2025, found that 28% of the central government technology estate is legacy. That is up from 26% in 2023, with a range of 10% to 60% across organisations, and 22% of legacy systems red-rated. Note the different denominators in the two red-rated figures. They are not the same measurement and should not be compared.
No vendor lock, missing interface or absent supplier support is fixed by an incentive change. Those are real and countable, and they will still be there after the workshop.
So where is the line? The same DSIT review contains the best answer I have found, and it is the most useful sentence in this essay's research. Half of survey respondents indicated that when a budget exists for legacy system remediation, it frequently gets reallocated to other initiatives. Read that carefully. The problem is technical. The money was approved. It vanished into someone else's priority, repeatedly and by decision.
The same review adds two findings that point the same way. Only 27% of respondents believe their data infrastructure gives a full operational or transactional view, and 70% describe the data landscape as uncoordinated, lacking interoperability and without a single source of truth. And departments treat data as their own, which blocks aggregation even where the legal basis for sharing already exists under the Digital Economy Act 2017.
That gives three tests, and they are empirical rather than rhetorical. First: does the data physically exist, at the granularity the use case needs? If not, this is not politics, and no diagnostic of mine will help you. Second: are legal permission and technical capability already in place while nothing flows? Then it is politics, and the evidence is the gap between what is allowed and what happens. Third: does a budget exist and get reallocated year after year? Then it is politics too, and the reallocation decision is where you should be looking.
Forget "technical versus political." The line runs between what cannot be done, what cannot be funded two years running, and what is permitted but not consented to by the person who holds it. Only the first is an engineering problem.
Closing
Before you approve another catalogue, find the person whose budget pays for the quality and whose bonus does not depend on it. If you cannot name them, you do not have a data problem yet — you have an ownership vacancy, and the catalogue will document it beautifully.
Sources
- primaryGartner, Lack of AI-Ready Data Puts AI Projects at Risk (press release, 26 February 2025) — the "through 2026 … abandon 60% of AI projects unsupported by AI-ready data" forecast, and the 63% figure from a Q3 2024 survey of 248 data management leaders. (newsroom returned 403 to automated access; wording verified via Freevacy's report of the release, 26 February 2025)
- primaryGartner, Gartner Predicts 80% of D&A Governance Initiatives Will Fail by 2027, Due to a Lack of a Real or Manufactured Crisis (press release, 28 February 2024) — the forecast and the Saul Judah quote on prioritised business outcomes. (403 to automated access; quotes as reported by BizTechReports, 29 February 2024)
- primaryGartner, Organizations with Successful AI Initiatives Invest Up to Four Times More in Data and Analytics Foundations (press release, 16 April 2026) — 353 D&A and AI leaders, November–December 2025; the "up to four times" upper bound; Rita Sallam on 39% confidence. (403 to automated access; figures as reported by TechEdge AI, 20 April 2026)
- primaryThomas H. Davenport, Robert G. Eccles & Laurence Prusak, "Information Politics," MIT Sloan Management Review (15 October 1992) — two years of fieldwork across more than 25 firms; information as organisational currency; the five models from technocratic utopianism to federalism.
- primaryVijay Khatri & Carol V. Brown, "Designing Data Governance," Communications of the ACM 53(1):148–152 (2010) — governance as the allocation of decision rights and accountabilities, distinct from management.
- primaryDavid J. Hickson, C. R. Hinings, C. A. Lee, R. E. Schneck & J. M. Pennings, "A Strategic Contingencies' Theory of Intraorganizational Power," Administrative Science Quarterly 16(2):216–229 (1971) — subunit power as coping with uncertainty, workflow centrality and non-substitutability.
- primaryMichael C. Jensen & William H. Meckling, "Theory of the Firm: Managerial Behavior, Agency Costs and Ownership Structure," Journal of Financial Economics 3(4):305–360 (1976) — monitoring costs, bonding costs and residual loss when ownership and control separate. Applied here to departments, which extends beyond the paper's subject.
- primaryMichael A. Heller, "The Tragedy of the Anticommons: Property in the Transition from Marx to Markets," Harvard Law Review 111(3):621–688 (1998) — too many rights of exclusion and no effective privilege of use produce underuse; the empty-storefront observation.
- primaryGarrett Hardin, "The Tragedy of the Commons," Science 162(3859):1243–1248 (1968) — used here as the counterpoint: overuse without an owner, the opposite sign to the silo problem.
- primaryElinor Ostrom, Governing the Commons: The Evolution of Institutions for Collective Action (Cambridge University Press, 1990) — self-governance of shared resources without privatisation or central authority; design principle two, proportionality between benefits and costs.
- primaryMancur Olson, The Logic of Collective Action: Public Goods and the Theory of Groups (Harvard University Press, 1965) — concentrated costs and diffuse benefits produce chronic undersupply; selective incentives as the remedy.
- primaryD. Sculley et al., "Hidden Technical Debt in Machine Learning Systems," NIPS 2015 — ML code as a small fraction of the system; data dependencies costlier than code dependencies; CACE.
- primaryRAND Corporation (James Ryseff, Brandon De Bruhl, Sydne Newberry), The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed, RRA2680-1 (2024) — qualitative study, 65 interviews; insufficient data as one of five root causes; the ">80% fail" line is a quoted estimate, not a RAND measurement.
- primaryNational Audit Office, Government cyber resilience (29 January 2025) — 228 significant legacy systems as of March 2024; 63 (28%) red-rated; 120 of 228 without fully funded remediation plans.
- primaryDepartment for Science, Innovation and Technology / Government Digital Service, State of digital government review (21 January 2025) — 28% of the central government estate is legacy, range 10–60%, 22% red-rated; remediation budgets frequently reallocated; 27% report a full operational view; 70% describe an uncoordinated landscape; data treated as departmental property despite the Digital Economy Act 2017.
- secondaryRandy Bean & Thomas H. Davenport, "Why Do Chief Data Officers Have Such Short Tenures?", Harvard Business Review (August 2021), as summarised by Brian Eastwood, "Chief data officers don't stay in their roles long. Here's why," MIT Sloan (1 September 2022) — average CDO tenure around 2.5 years against roughly 7 for CEOs and 4.5 for CFOs and CIOs; unclear mandate as a stated cause. Self-reported survey data, not a registry.
Dane gotowe na AI to problem władzy
Każdy program AI uderza w końcu w tę samą ścianę, a opisuje się ją zawsze językiem inżynierii. Dane nie są gotowe. W lutym 2025 Gartner prognozował, że do 2026 organizacje porzucą 60% projektów AI niewspartych danymi gotowymi na AI. Okno tej prognozy dopiero się domyka. Odczytaj ją jako problem inżynierski, a kupisz pipeline'y, katalogi i dashboardy jakości. Odczytaj jako problem własności, a dostaniesz inny budżet i inny zestaw argumentów.
Ten esej stawia jedną tezę i jest ona wąska. Część gotowości danych naprawdę jest techniczna. Ale ta część, która zabija projekty, jest zwykle pytaniem o własność: kto jest właścicielem danych, czyj budżet płaci za ich jakość i czyja pozycja słabnie, gdy dane trafiają do kogoś innego. Nie jest tak, że komuś zabrakło dyscypliny: silos jest rozsądną reakcją ludzi rozliczanych z czegoś innego.
Twój problem z jakością danych ma stanowisko i system premiowy.
Przeczytaj liczby, zanim je zacytujesz
Liczba 60% źle znosi podróż. Komunikat prasowy Gartnera z 26 lutego 2025 mówi, że „do 2026 organizacje porzucą 60% projektów AI niewspartych danymi gotowymi na AI". Po drodze gubią się trzy rzeczy. To prognoza wydana na początku 2025 o oknie kończącym się w 2026, a nikt dotąd nie zestawił jej publicznie z wynikami. Mianownikiem są projekty niewsparte danymi gotowymi na AI, nie wszystkie projekty AI. Popularna wersja, „60% projektów AI upadnie", to twierdzenie inne i fałszywe. Komunikat nie podaje też modelu stojącego za prognozą. Jedna uwaga o uczciwości źródłowej: newsroom Gartnera blokuje dostęp automatyczny, więc pracuję na serwisach, które zacytowały komunikat dosłownie, a nie na stronie pierwotnej.
Towarzysząca liczba jest ciekawsza i częściej przekręcana. Gartner podaje, że 63% organizacji „nie ma właściwych praktyk zarządzania danymi dla AI albo nie jest pewnych, czy je ma". To wynik ankiety wśród 248 liderów zarządzania danymi z trzeciego kwartału 2024. Czyli: samoocena, z wąskiej podpopulacji zawodowej, o własnych organizacjach. Nikt tu nie audytował praktyki ani nie patrzył z poziomu zarządu.
To zdanie zlepia też dwa stany, które nie są tym samym. „Nie mamy praktyk" to odpowiedź o zasobach. „Nie jesteśmy pewni, czy je mamy" to odpowiedź o tym, kto odpowiada za wiedzę. Gartner tych dwóch nie rozdziela, a w źródle nie ma rozbicia procentowego. To ma znaczenie, bo cały ten esej wyrasta z drugiej połowy tamtego zdania. Budowanie na niej argumentu to moja interpretacja, nie ustalenie Gartnera. Organizacja, która nie umie powiedzieć, czy ma praktyki zarządzania danymi, powiedziała ci coś precyzyjnego o własności.
Pod tym wszystkim leży jedno rozstrzygnięcie definicyjne i jest ono moje. „Gotowe dane" w ogólności nie istnieją. Dane są gotowe albo niegotowe dla konkretnego zastosowania, wobec wymagań konkretnej techniki, i tylko dopóki to zastosowanie się nie zmieni. Gotowość jest relacją: wiąże dane z zastosowaniem, a w samych danych jej nie znajdziesz. Ma to konsekwencję organizacyjną, którą się pomija: odpowiedzialności za ruchomą relację nie da się scentralizować w jednym programie poprzecznym, bo program nie ma stabilnego przedmiotu do zarządzania. Kto ma na głowie „jakość danych" w abstrakcji, ten ma rzeczownik zmieniający kształt za każdym razem, gdy zmienia kształt cudzy projekt.
Tak mniej więcej program zostaje z katalogiem zamiast z wynikiem.
Informacja zawsze była walutą
Mechanizm nie jest nowy, a najlepszy jego opis wyprzedza obecną falę o ponad trzy dekady. Thomas Davenport, Robert Eccles i Laurence Prusak ogłosili „Information Politics" w MIT Sloan Management Review 15 października 1992, na podstawie dwóch lat badań terenowych w ponad 25 firmach. Ich ustalenie idzie pod prąd optymizmu epoki. Technologia informatyczna miała uwolnić przepływ informacji i spłaszczyć hierarchię. Zrobiła niemal odwrotnie. Informacja stała się walutą życia organizacji, zbyt cenną, by większość menedżerów chciała ją oddawać.
Ich rama to taksonomia pięciu modeli polityki informacyjnej: technokratyczna utopia, anarchia, feudalizm, monarchia i federalizm. Feudalizm jest tym znajomym modelem, w którym jednostki biznesowe prowadzą własne włości informacyjne z własnym słownictwem. Autorzy opisują firmę jako zbiór miast-państw, każde z własną kulturą, przywództwem i językiem. Realistycznym celem nie jest u nich monarchia, tylko federalizm: wynegocjowane wspólne pojęcia między jednostkami, które zachowują realną autonomię.
Model wart nazwania dzisiaj to pierwszy. Technokratyczna utopia to wiara, że dostatecznie dobry program techniczny rozpuści problem organizacyjny. Zbuduj hurtownię, zamodeluj przedsiębiorstwo, ujednolić definicje, a polityka pójdzie za tym. W 1992 znaczyło to korporacyjne modelowanie danych. Dziś znaczy „dane gotowe na AI" jako inicjatywę platformową z datą wdrożenia.
Wiek tego źródła działa tu na korzyść tezy. Od 1992 to samo zachowanie przetrwało hurtownię danych, master data management, data lake, lakehouse i data mesh. Pięć odpowiedzi technicznych, jeden niezmieniony wzorzec organizacyjny. Ograniczenie, które przeżywa każde skierowane w nie rozwiązanie, prawdopodobnie nie leży w warstwie, którą te rozwiązania ruszają.
Skoro narzędzie zmieniało się raz za razem, a wynik nie, to zmienną nie było narzędzie.
Dlaczego silos jest racjonalny
Załóż na czas tej sekcji, że wszyscy zaangażowani są kompetentni i działają w dobrej wierze. Silos i tak się tworzy. To jest ciekawy przypadek, a z mojego doświadczenia także częsty.
Zacznij od władzy, zdefiniowanej operacyjnie, a nie moralnie. Hickson, Hinings, Lee, Schneck i Pennings przedstawili teorię strategicznych uwarunkowań władzy wewnątrzorganizacyjnej w Administrative Science Quarterly w 1971. Trzy rzeczy podnoszą władzę podjednostki: to, jak dobrze radzi sobie z niepewnością za innych, jak centralna jest dla przepływu pracy i jak trudno ją zastąpić. Niepewność ma tu znaczenie ścisłe: brak informacji o przyszłych zdarzeniach.
Teraz wstaw dane do tego modelu. Dane są właśnie tym zasobem, który redukuje cudzą niepewność. Jednostka trzymająca dane potrzebne innym jest, według tej teorii, tym silniejsza, im większa jest ta potrzeba. Oddaj dane w czystej formie, takiej, której inni użyją bez pytania, a zmniejszysz ich zależność od siebie. Zrobisz się przy tym łatwiejszy do zastąpienia.
W tym modelu dzielenie się danymi to jednostronne rozbrojenie.
Teraz strona kosztów, bo tam problem bodźców staje się ścisły. Praca Jensena i Mecklinga z 1976 w Journal of Financial Economics zdefiniowała relację agencji. Gdy jedna strona działa w imieniu drugiej, a interesy się rozjeżdżają, powstają koszty monitorowania, koszty zabezpieczenia i strata rezydualna. Ich przedmiotem byli akcjonariusze i menedżerowie, więc przeniesienie tego na działy to rozszerzenie modelu, nie ich twierdzenie.
To rozszerzenie się broni, bo struktura jest ta sama. Właściciel danych jest agentem firmy i pryncypałem własnego budżetu. Doprowadzenie systemu źródłowego do jakości, jakiej potrzebuje model innego zespołu, pochłania godziny jego inżynierów i miejsce w jego roadmapie. Korzyść spływa na projekt, który ma innego dyrektora i mierzy się metryką z cudzej części P&L. Nic w tym układzie nie jest nieracjonalne po żadnej ze stron. Odmowa to przewidziany wynik.
Mancur Olson wyjaśnił, dlaczego nikt nie zgłasza się do płacenia. W książce The Logic of Collective Action z 1965 roku (Harvard University Press) wyłożył arytmetykę skoncentrowanych kosztów i rozproszonych korzyści. Jego sformułowanie: „im większa jest grupa, tym bardziej będzie odbiegać od optymalnej podaży jakiegokolwiek dobra zbiorowego". Jakość danych jest dobrem zbiorowym dla każdego odbiorcy niżej w łańcuchu, a finansuje ją jeden producent wyżej. Olson przewiduje chroniczne niedostarczanie, tym gorsze, im większa organizacja. Tak właśnie wyglądają w praktyce budżety na jakość danych. Jego lekarstwem jest bodziec selektywny: korzyść dostępna wyłącznie dla strony ponoszącej koszt.
Teraz kształt tej porażki. Zwyczajowa metafora myli kierunek. Do opisania silosów danych ludzie sięgają po tragedię wspólnego pastwiska Garretta Hardina (Science, 1968). Hardin opisuje nadużywanie wspólnego zasobu, który nie ma właściciela. Silos jest odwrotnością: zasób ma zbyt wielu właścicieli i przez to leży odłogiem. Właściwy model jest Michaela Hellera, z „The Tragedy of the Anticommons" w Harvard Law Review (1998). Gdy wiele stron trzyma prawa wykluczania, a nikt nie ma skutecznego prawa użycia, zasób zostaje wykorzystany daleko poniżej możliwości.
Heller zaczyna od obserwacji, która zostaje w głowie. Po upadku komunizmu witryny sklepowe w miastach Europy Wschodniej stały puste, a przed nimi mnożyły się uliczne kioski. Popyt istniał. Nie istniała droga przez kilka urzędów i podmiotów, z których każdy miał weto wobec najmu. Tak wygląda architektura korporacyjnego majątku danych: nadmiar praw weta, niedobór praw użycia.
Trafne rozpoznanie zmienia zalecenie. Wspólnemu pastwisku dodaje się właściciela. Dla anticommons trzeba scalić albo znieść prawa wykluczania. Standardowa odpowiedź governance (zwołaj radę, daj każdej domenie miejsce, wymagaj akceptacji) dokłada prawa weta do systemu, który już na nie umiera. To właściwe lekarstwo podane na niewłaściwą chorobę, więc pogarsza sprawę.
Governing the Commons Elinor Ostrom z 1990 roku (Cambridge University Press) dokłada konstruktywną połowę i zarazem zabija leniwy wniosek. Jej studia przypadków pokazały, że wspólne zasoby bywają skutecznie zarządzane bez prywatyzacji i bez centralnej władzy, według zestawu zasad projektowych. Druga z nich jest tu operacyjna: korzyści mają być proporcjonalne do kosztów. Kto płaci za wspólny zasób, musi dostać udział w wartości powiązany z tym, co zapłacił. Przeczytaj to jako polecenie, a wyjdzie reguła budżetowa, do której jeszcze wrócę.
Silos to działający schemat organizacyjny.
Governance bez praw decyzyjnych to teatr
Definicja data governance na tyle ostra, by wykryć teatr, pochodzi z literatury naukowej, nie od dostawcy. Vijay Khatri i Carol Brown w Communications of the ACM z 2010 definiują governance jako przydział praw decyzyjnych i odpowiedzialności za procesy związane z danymi. Oddzielają governance (które decyzje trzeba podjąć i kto je podejmuje) od zarządzania, czyli od tego, kto wykonuje.
To daje test, który przeprowadzisz w trakcie jednego spotkania. Czy program przenosi choć jedno prawo decyzyjne z jednej osoby na drugą? Czy zmienia choć jedną linijkę w czyichś celach rocznych? Jeśli obie odpowiedzi brzmią „nie", to zarządzanie przebrane za governance. Katalogi i komitety bez przeniesionych praw decyzyjnych pasują do tego opisu co do joty.
To wyjaśnia też, w jaki konkretnie sposób takie programy się rozkładają. Katalog jest dokumentacją zasobu, którego właścicielem jest ktoś inny. Utrzymanie wpisu w aktualności to koszt prywatny, płacony przez zespół źródłowy dla korzyści rozsianej wśród odbiorców, których nigdy nie pozna. To znowu struktura Olsona, poziom niżej, na ziarnistości opisu jednej tabeli. Niedostarczanie przewiduje ten sam model, który przewiduje niedofinansowanie. Jedną lukę trzeba tu oznaczyć. Nie znalazłem wiarygodnego publicznego badania mierzącego, ile katalogów korporacyjnych wypada z utrzymania; krążące twierdzenia pochodzą od firm sprzedających katalogi. Traktuj ten mechanizm jako wyrozumowany, nie zmierzony.
Ten sam test tłumaczy data stewarda bez uprawnień. Steward odpowiada za jakość danych, ale nie może ustawić roadmapy systemu źródłowego. Tak zwykle projektuje się tę rolę: odpowiedzialność bez praw decyzyjnych. To strukturalna niemożliwość wręczona człowiekowi.
Jak to wygląda w praktyce, pokazuje jedna liczba. Podaję ją ostrożnie. Bean i Davenport podali w Harvard Business Review w 2021, że przeciętna kadencja chief data officera trwa około dwóch i pół roku; MIT Sloan streścił ich tekst rok później. Dla porównania: CEO wytrzymuje mniej więcej siedem lat. CFO albo CIO wytrzymuje cztery i pół roku. Nowsze ankiety branżowe pokazują podobny obraz: ponad połowa respondentów wskazuje mniej niż trzy lata. Powody, jakie podają sami badani, to niejasny mandat i rola wykrojona z terytorium CIO, przy transformacji oczekiwanej w mniej więcej osiemnaście miesięcy. Podkreślam „sami": liderzy danych odpowiadali w ankiecie o własnej sytuacji, nikt nie sprawdzał rejestru powołań. MIT Sloan nie wskazuje braku budżetu ani formalnej władzy jako głównej przyczyny i nie zamierzam wkładać im tego w usta. Krótka kadencja mówi tylko tyle, że robota jest trudna, i nie wskazuje, skąd ta trudność się bierze.
Własna prognoza Gartnera o governance jest najostrzejszą rzeczą, jaką w tej sprawie opublikowali, i rzadko czyta się ją uważnie. W lutym 2024 przewidzieli, że do 2027 upadnie 80% inicjatyw governance w danych i analityce „z powodu braku realnego lub wykreowanego kryzysu". Cytowany w komunikacie Saul Judah, VP Analyst, mówi, że program governance, który nie umożliwia priorytetyzowanych wyników biznesowych, upada. Zalecenie brzmi: odejść od postawy centralno-nakazowej w stronę namacalnych wyników biznesowych. Znowu: prognoza, nie pomiar, i znowu cytowana z drugiej ręki.
Zatrzymaj się przy słowie wykreowany. Firma doradcza otwarcie zaleca, żebyś wyprodukował kryzys, bo governance bez niego się nie przyjmuje. To przyznanie, na czym te programy naprawdę jadą: na uwadze i na priorytecie. Jedno i drugie to waluty polityczne, rozdzielane przez ludzi mających pozycję, żeby je rozdzielać. Gartner nie nazywa tego problemem władzy. To odczytanie jest moje i uważam je za jedyne uczciwe z dostępnych.
Wracamy tym do 63% i do tej połowy, która mówi „nie jesteśmy pewni". Organizacja daje taką odpowiedź, kiedy nikt nie odpowiada za to pytanie. Odpowiadać za pytanie znaczy być osobą, której ocena roczna zawiera odpowiedź. Tam, gdzie takiej osoby nie ma, uczciwą reakcją na audytora jest wzruszenie ramion, a wzruszenie ramion zostaje zapisane jako ustalenie o zarządzaniu danymi. Naprawdę jest ustaleniem o odpowiedzialności.
Diagnostyka polityczna, przeprowadzona przed projektem
Nic z tego nie jest użyteczne, dopóki nie zmieni tego, co robisz w pierwszym tygodniu. Więc oto procedura, której używam. To moja własna synteza, nie metoda Gartnera ani twierdzenie, i należy ją czytać jako ramę, nie jako pomiar.
Pytań jest pięć i zadajesz je, zanim ktokolwiek napisze pierwsze zadanie integracyjne.
Kto jest właścicielem każdego zbioru danych, którego potrzebuje projekt? Nie który system go trzyma — która osoba z imienia i nazwiska kontroluje roadmapę systemu, który go produkuje. Jeśli nie umiesz wskazać człowieka z budżetem, znalazłeś pierwszy problem.
Ile to przekazanie będzie ich kosztować w tym roku? W ich godzinach inżynierskich i ich zobowiązaniach kwartalnych. Oszacowanie weź od nich, nie od swojego architekta. Liczba wychodzi zwykle większa, niż zakłada zespół proszący, a w tej różnicy grzęzną projekty.
Kto wychodzi na tym gorzej, jeśli projekt się uda? W zakresie, etatach, linii budżetowej albo we władaniu metryką, którą dziś ma prawo tłumaczyć. Prawie zawsze ktoś taki jest. Opór z tej strony to informacja o rozkładzie władzy, nie dowód głupoty.
Które prawo decyzyjne musiałoby się przenieść i kto to podpisuje? Nazwij prawo: zatwierdzanie zmian schematu albo przyznawanie dostępu bez rozpatrywania każdej sprawy z osobna. Jeśli żadne prawo się nie przenosi, nie licz na przepływ.
Czy droga prawna i techniczna jest już otwarta? Jeśli zgoda istnieje, interfejs istnieje, a dane nadal nie płyną, patrzysz na ograniczenie polityczne. Właśnie postawiłeś diagnozę bez warsztatów.
Tabelą poniżej rozdzielam te dwa odczyty w praktyce.
| Objaw | Odczyt inżynierski | Odczyt własnościowy | Test, który je rozdziela |
|---|---|---|---|
| „Jakość danych jest za słaba" | Zaległości w czyszczeniu i walidacji | Żaden budżet nie finansuje jakości, z której korzysta ktoś inny | Zapytaj, czyj P&L płaci za naprawę, a czyj dostaje korzyść |
| „Dostęp trwa miesiącami" | Brak interfejsów i narzędzi bezpieczeństwa | Prawa wykluczania w rękach kilku stron, brak prawa użycia | Policz strony, które mogą powiedzieć „nie"; policz te, które mogą powiedzieć „tak" |
| „Nie wiemy, jakie mamy dane" | Katalog jest niepełny | Nikt nie odpowiada za wiedzę | Zapytaj, kogo się z tej odpowiedzi rozlicza |
| „System źródłowy tego nie wyeksportuje" | Legacy bez API | Ograniczenie prawdziwe, dopóki nie dowiedziesz inaczej | Zapytaj, czy istnieje sfinansowany plan naprawy — i czy przetrwał zeszły rok |
| „Governance mamy wdrożone" | Polityki i komitet | Żadne prawo decyzyjne się nie przeniosło | Wskaż jedno prawo, które zmieniło właściciela, na piśmie |
| „Nie ma na to budżetu" | Realny sufit kosztowy | Budżet jest i co roku ląduje gdzie indziej | Sprawdź, czy linia została zatwierdzona, a potem wydana na coś innego |
Potem trzy ruchy. Są celowo małe, bo duże wymagają władzy, której lider projektu zwykle nie ma.
Wsadź budżet na jakość danych do środka budżetu projektu AI, nie obok niego. To wprost zastosowana zasada proporcjonalności Ostrom. Osobny program jakości danych prosi zespół źródłowy o dostarczenie dobra zbiorowego za darmo. Pozycja w budżecie projektu zamienia to w kupioną usługę, która ma klienta, a do tego cenę i termin dostawy. Ta sama praca, przeniesiona z jałmużny do handlu. Stawia to też przy projekcie uczciwą liczbę, co niektórym sponsorom się nie spodoba.
Wsadź właściciela danych do zespołu projektowego i daj mu udział w wyniku. Nie do komitetu sterującego, bo w komitetach odpowiedzialność się rozcieńcza. Skin in the game znaczy, że jego cele zawierają kawałek rezultatu projektu. To bodziec selektywny Olsona w najmniejszej użytecznej postaci: korzyść docierająca wyłącznie do strony, która ponosi koszt.
Przenieś dokładnie jedno prawo decyzyjne i zapisz to. Jedno wystarczy, żeby sprawdzić, czy organizacja jest chętna. Jeśli nic nie da się przenieść, dowiedziałeś się wcześnie czegoś ważnego. Program wyprodukuje dokumentację i zero przepływu, a ty możesz na tej podstawie zdecydować, czy go finansować.
Mam jeszcze jedną rzecz najbliższą dowodowi i poważne zastrzeżenia do niej. Gartner podał w kwietniu 2026, że organizacje, które same zgłaszają sukces w AI, inwestują w fundamenty do czterech razy więcej jako procent przychodu. Na tej liście fundamentów są jakość danych, governance, ludzie gotowi na AI i zarządzanie zmianą. Ankieta objęła 353 liderów obszaru danych i analityki oraz AI; Gartner przeprowadził ją w listopadzie i grudniu 2025. Cytowana w tym samym komunikacie Rita Sallam przywołuje inną liczbę: tylko 39% liderów technologicznych wierzy, że obecne inwestycje w AI poprawią wyniki finansowe. Nagłówek znaczy tu mniej niż to, co go podważa: sukces oceniły sobie same badane organizacje, a przyczynowość zostaje nierozstrzygnięta, bo sukces może finansować fundamenty równie dobrze jak fundamenty sukces. Zostaje jeszcze „do czterech razy": to pułap, pod którym mieści się cały rozrzut, więc czytanie go jako przeciętnej zawyża wynik. Biorę z tego jedno: skład tej listy fundamentów, na której ludzie i zarządzanie zmianą stoją obok jakości danych. To pozycje organizacyjne, a Gartner wymienia je jednym tchem z tą techniczną.
Argument łączy się z szerszą tezą o tym, gdzie chowa się wartość z AI. Zwrot z systemu AI rzadko siedzi w modelu; mieszka w ogniwach wokół niego. Własność danych jest najmniej widocznym z tych ogniw, bo nie pojawia się na żadnym diagramie architektury.
Gdzie ta teza jest błędna
Teza, która tłumaczy każdą porażkę, nie tłumaczy żadnej, a „to problem władzy" da się naciągnąć na wszystko. To czyni ją wygodną, a wygodne argumenty zasługują na najtwardszy test, jaki umiem zbudować.
Czasem ograniczeniem naprawdę jest to, że danych nie ma. Raport RAND z 2024 o źródłowych przyczynach porażek projektów AI wymienia niedostateczne dane jako jedną z pięciu. Organizacja po prostu nie ma danych potrzebnych, żeby wytrenować skuteczny model. Ważne jest, jak autorzy do tego doszli. Badanie jest jakościowe i opiera się na 65 wywiadach z data scientistami i inżynierami z co najmniej pięcioletnim stażem. Liczba 65 mówi więc, z iloma ludźmi RAND rozmawiał, a nie ile projektów zbadał. Zdanie z tego samego raportu o „ponad 80% porażek projektów AI" RAND zacytował z cudzych szacunków i sam go nie zmierzył.
Ta przyczyna źródłowa wyznacza twardą granicę wszystkiemu, co napisałem wyżej. Dane, których nigdy nie zbierano w wymaganej ziarnistości, nie pojawią się dlatego, że naprawiłeś bodźce. Lepsze bodźce mogą uruchomić zbieranie od dziś. Nic to nie daje projektowi, który potrzebuje dwudziestu czterech miesięcy historii, a żadna mapa interesariuszy tego czekania nie skróci.
Czasem ograniczeniem jest żmudna inżynieria, której nikt nie chce robić. Sculley i współautorzy postawili sprawę w „Hidden Technical Debt in Machine Learning Systems" na NIPS 2015 i tekst dobrze się zestarzał. Kod modelu to niewielki ułamek prawdziwego systemu produkcyjnego; reszta to zbieranie danych, ekstrakcja cech, serwowanie i monitoring. Zależności danych kosztują w utrzymaniu więcej niż zależności kodu, a dostają mniej rygoru inżynierskiego. Ich zasada CACE (zmiana czegokolwiek zmienia wszystko) opisuje, dlaczego te systemy opierają się czystej dekompozycji. Ta praca jest droga i nieefektowna. Jest też akurat tą pracą, którą przemianowuje na „problem organizacyjny" ktoś, kto woli poprowadzić warsztat niż migrację.
Chcę być w tej sprawie ostry, bo to ten esej dostarcza jej słownika. Diagnoza polityczna jest znakomitym alibi. Zamienia trudny dług inżynierski w cudzą wadę charakteru i pozwala zespołowi przepalić kwartał na mapy interesariuszy, podczas gdy pipeline dalej jest zepsuty. Jeśli po tej lekturze masz spokojniejsze sumienie, choć roboty przy pipelinie nie tykasz, użyłeś jej źle.
Czasem pieniędzy naprawdę nie ma. Brytyjski National Audit Office raportował w styczniu 2025 stan na marzec 2024. Resorty rządowe utrzymywały około 228 istotnych systemów legacy, z czego 63 (28%) miało czerwoną ocenę za wysokie prawdopodobieństwo i wysoki wpływ ryzyka operacyjnego lub bezpieczeństwa. Dla 120 z tych 228, czyli ponad połowy, nie istniał w pełni sfinansowany plan naprawy. Osobno przegląd cyfrowego państwa autorstwa DSIT i Government Digital Service, także ze stycznia 2025, ustalił, że 28% majątku technologicznego administracji centralnej to legacy. To wzrost z 26% w 2023, z rozrzutem od 10% do 60% między organizacjami, przy 22% systemów legacy z oceną czerwoną. Zwróć uwagę na różne mianowniki w obu liczbach o czerwonej ocenie. To nie jest ten sam pomiar i nie należy ich zestawiać.
Żadnego uzależnienia od dostawcy, brakującego interfejsu ani nieobecnego wsparcia producenta nie naprawia zmiana bodźców. To rzeczy realne i policzalne. Po warsztacie nigdzie nie znikną.
Gdzie więc biegnie granica? Ten sam przegląd DSIT zawiera najlepszą odpowiedź, jaką znalazłem, i jest to najużyteczniejsze zdanie z całego researchu do tego eseju. Połowa badanych wskazała, że gdy budżet na naprawę systemów legacy istnieje, często zostaje przesunięty na inne inicjatywy. Przeczytaj to uważnie. Problem jest techniczny. Pieniądze były zatwierdzone. Znikały w cudzym priorytecie wielokrotnie, za każdym razem na mocy czyjejś decyzji.
Ten sam przegląd dokłada dwa ustalenia, które wskazują w tę samą stronę. Tylko 27% badanych uważa, że ich infrastruktura danych daje pełny obraz operacyjny lub transakcyjny, a 70% opisuje swoje środowisko danych jako nieskoordynowane, bez interoperacyjności i bez jednego źródła prawdy. Resorty traktują też dane jak swoje, co blokuje agregację nawet tam, gdzie podstawa prawna do dzielenia się już istnieje na mocy Digital Economy Act 2017.
To daje trzy testy, empiryczne, nie retoryczne. Pierwszy: czy dane fizycznie istnieją, w ziarnistości, której potrzebuje to zastosowanie? Jeśli nie, to nie jest polityka i żadna moja diagnostyka ci nie pomoże. Drugi: czy zgoda prawna i możliwość techniczna są na miejscu, a mimo to nic nie płynie? Wtedy to polityka, a dowodem jest przepaść między tym, co wolno, a tym, co się dzieje. Trzeci: czy budżet istnieje i rok po roku ląduje gdzie indziej? Wtedy to też polityka, a przyglądać się powinieneś właśnie decyzji o przesunięciu.
Zapomnij o podziale „techniczne kontra polityczne". Granica biegnie między tym, czego nie da się zrobić, tym, czego nie da się sfinansować dwa lata z rzędu, i tym, na co jest zgoda formalna, ale nie ma zgody człowieka, który to trzyma. Tylko pierwsze jest problemem inżynierskim.
Zamknięcie
Zanim zatwierdzisz kolejny katalog, znajdź osobę, której budżet płaci za jakość, a której premia od tej jakości nie zależy. Jeśli nie umiesz jej nazwać, nie masz jeszcze problemu z danymi — masz wakat na właściciela, a katalog udokumentuje go pięknie.
Bibliografia
- primaryGartner, Lack of AI-Ready Data Puts AI Projects at Risk (komunikat prasowy, 26 lutego 2025) — prognoza „do 2026 … porzucą 60% projektów AI niewspartych danymi gotowymi na AI" oraz liczba 63% z ankiety wśród 248 liderów zarządzania danymi z III kwartału 2024. (newsroom zwrócił 403 przy dostępie automatycznym; brzmienie zweryfikowane przez relację Freevacy z komunikatu, 26 lutego 2025)
- primaryGartner, Gartner Predicts 80% of D&A Governance Initiatives Will Fail by 2027, Due to a Lack of a Real or Manufactured Crisis (komunikat prasowy, 28 lutego 2024) — prognoza i cytat Saula Judaha o priorytetyzowanych wynikach biznesowych. (403 przy dostępie automatycznym; cytaty za relacją BizTechReports z 29 lutego 2024)
- primaryGartner, Organizations with Successful AI Initiatives Invest Up to Four Times More in Data and Analytics Foundations (komunikat prasowy, 16 kwietnia 2026) — 353 liderów D&A i AI, listopad–grudzień 2025; górna granica „do czterech razy"; Rita Sallam o 39% przekonanych. (403 przy dostępie automatycznym; liczby za relacją TechEdge AI z 20 kwietnia 2026)
- primaryThomas H. Davenport, Robert G. Eccles i Laurence Prusak, „Information Politics", MIT Sloan Management Review (15 października 1992) — dwa lata badań terenowych w ponad 25 firmach; informacja jako waluta organizacji; pięć modeli od technokratycznej utopii do federalizmu.
- primaryVijay Khatri i Carol V. Brown, „Designing Data Governance", Communications of the ACM 53(1):148–152 (2010) — governance jako przydział praw decyzyjnych i odpowiedzialności, odrębny od zarządzania.
- primaryDavid J. Hickson, C. R. Hinings, C. A. Lee, R. E. Schneck i J. M. Pennings, „A Strategic Contingencies' Theory of Intraorganizational Power", Administrative Science Quarterly 16(2):216–229 (1971) — władza podjednostki jako radzenie sobie z niepewnością, centralność w przepływie pracy i niezastępowalność.
- primaryMichael C. Jensen i William H. Meckling, „Theory of the Firm: Managerial Behavior, Agency Costs and Ownership Structure", Journal of Financial Economics 3(4):305–360 (1976) — koszty monitorowania, koszty zabezpieczenia i strata rezydualna przy rozdzieleniu własności i kontroli. Zastosowane tu do działów, co wykracza poza przedmiot tej pracy.
- primaryMichael A. Heller, „The Tragedy of the Anticommons: Property in the Transition from Marx to Markets", Harvard Law Review 111(3):621–688 (1998) — zbyt wiele praw wykluczania i brak skutecznego prawa użycia dają wykorzystanie zasobu poniżej możliwości; obserwacja o pustych witrynach sklepowych.
- primaryGarrett Hardin, „The Tragedy of the Commons", Science 162(3859):1243–1248 (1968) — użyty tu jako kontrapunkt: nadużywanie bez właściciela, znak przeciwny do problemu silosów.
- primaryElinor Ostrom, Governing the Commons: The Evolution of Institutions for Collective Action (Cambridge University Press, 1990) — samorządność wspólnych zasobów bez prywatyzacji i bez centralnej władzy; druga zasada projektowa, proporcjonalność korzyści do kosztów.
- primaryMancur Olson, The Logic of Collective Action: Public Goods and the Theory of Groups (Harvard University Press, 1965) — skoncentrowane koszty i rozproszone korzyści dają chroniczne niedostarczanie; bodźce selektywne jako lekarstwo.
- primaryD. Sculley i in., „Hidden Technical Debt in Machine Learning Systems", NIPS 2015 — kod ML jako niewielki ułamek systemu; zależności danych droższe niż zależności kodu; CACE.
- primaryRAND Corporation (James Ryseff, Brandon De Bruhl, Sydne Newberry), The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed, RRA2680-1 (2024) — badanie jakościowe, 65 wywiadów; niedostateczne dane jako jedna z pięciu przyczyn źródłowych; zdanie o „ponad 80% porażek" to cytowany szacunek, nie pomiar RAND.
- primaryNational Audit Office, Government cyber resilience (29 stycznia 2025) — 228 istotnych systemów legacy na marzec 2024; 63 (28%) z oceną czerwoną; 120 z 228 bez w pełni sfinansowanych planów naprawy.
- primaryDepartment for Science, Innovation and Technology / Government Digital Service, State of digital government review (21 stycznia 2025) — 28% majątku administracji centralnej to legacy, rozrzut 10–60%, 22% z oceną czerwoną; budżety na naprawę często przesuwane; 27% zgłasza pełny obraz operacyjny; 70% opisuje środowisko jako nieskoordynowane; dane traktowane jak własność resortu mimo Digital Economy Act 2017.
- secondaryRandy Bean i Thomas H. Davenport, „Why Do Chief Data Officers Have Such Short Tenures?", Harvard Business Review (sierpień 2021), w streszczeniu Briana Eastwooda, „Chief data officers don't stay in their roles long. Here's why", MIT Sloan (1 września 2022) — przeciętna kadencja CDO około 2,5 roku wobec mniej więcej 7 dla CEO i 4,5 dla CFO i CIO; niejasny mandat jako podana przyczyna. Dane z ankiet samoopisowych, nie z rejestru.