← Writing

AI · MBA

The company brain and the knowledge nobody wrote down

One vendor page states the promise cleanly enough to test. "Company Brain is the memory that never resigns: everything your organization knows, structured, governed, and answerable in seconds." Take the middle clause literally. Everything your organisation knows. The claim I want to defend is narrower and harder to argue away. A company brain can only be assembled from what the company recorded, and the knowledge that decides expensive questions was mostly never recorded.

That boundary is no argument against the technology. It sets the price worth paying, and it names the half of the problem that stays yours. The interesting part is that the limit does not move when models improve. It is set by how much an organisation writes down, and that is a property of the organisation.

What the category actually promises

The pitch settled into a recognisable shape during 2026. Ingest the working substrate of the company: chat, mail, tickets, documents, the CRM, meeting transcripts, the code host. Extract entities and relations into a knowledge graph. Define the vocabulary once, in a semantic layer. Retrieve with a hybrid of vectors, keywords and graph traversal, and cite sources. Keep episodic memory across sessions. Put a governance plane on top for permissions and masking.

That is competent engineering. Each component does something real, and I have no quarrel with the parts.

The same page carries three numbers, and they are worth reading as a specimen. Over 35,000 operational documents at "a global aviation operator". Sixty times faster resolution for dispatch teams. Six hours down to seven seconds at "a professional-services firm". No customer is named. No baseline is defined. No method is published. Those illustrate a capability. They do not measure an effect, and a buyer should file them accordingly.

Strip the language and one sentence remains: we will make your company queryable. Everything below is an argument about how much of a company can be turned into a query at all.

A knowledge problem moved indoors

Friedrich Hayek published "The Use of Knowledge in Society" in the American Economic Review 35(4), September 1945. It runs from page 519 to 530. He wrote it against Oskar Lange and the case for a planned economy. The argument he makes is one about information.

Hayek's opening move restates what the economic problem is. Society does not face a problem of allocating resources whose properties a planner already knows. It faces "a problem of the utilization of knowledge which is not given to anyone in its totality" (§I). In section III he names the kind of knowledge he means: "the knowledge of the particular circumstances of time and place".

Section IV states the constraint that matters here. The knowledge he is concerned with "is knowledge of the kind which by its nature cannot enter into statistics and therefore cannot be conveyed to any central authority in statistical form". His conclusion follows in section V. "The ultimate decisions must be left to the people who are familiar with these circumstances, who know directly of the relevant changes and of the resources immediately available to meet them."

Hayek wrote about an entire economy, and his coordinating device was the price system. Carrying the argument inside a single firm is my extension, not his. It cannot be done mechanically either. Ronald Coase explained why in "The Nature of the Firm" (Economica, New Series, 4(16), November 1937, 386–405). Firms exist because organising some activity internally costs less than transacting for it in a market. If Hayek's objection held unchanged inside organisations, firms of any size would be impossible. They are not, so centralising knowledge within a company must sometimes pay.

What transfers is the filter. The verdict does not. Hayek's central authority could receive only what fitted into statistics. A retrieval system can receive only what fitted into a record. Both centres are limited at the same step. Something has to be written down before it can travel, and the writing loses most exactly where the knowledge was tied to a situation.

Where an organisation keeps what it knows

James Walsh and Gerardo Ungson drew the best available map of this in "Organizational Memory". It appeared in the Academy of Management Review 16(1), 1991, pages 57 to 91. They define organisational memory as stored information from an organisation's history that can be brought to bear on present decisions. They then locate it in five retention facilities plus external archives.

Hold the ingestion pipeline against that list.

  1. External archives. Documents, mail, tickets, records held inside the firm and outside it. Ingested well. This is where the category delivers.
  2. Transformations. The procedures that turn inputs into outputs. Ingested to the degree somebody documented them. That varies enormously by function, and it is worst where the work leans on judgement.
  3. Structures. Roles and the links between people. An org chart is ingestible. Who actually decides, and whose objection stops a project, is not on the chart.
  4. Individuals. Memory, skill, the pattern library of people who have seen the failure before. Not ingestible.
  5. Culture. Shared frames and language that settle which interpretations get taken seriously. Not ingestible.
  6. Ecology. The physical arrangement of the workplace, which encodes status and shapes who talks to whom. Not ingestible, and mostly not even noticed.

The honest description of a company brain follows. A strong index over one retention facility, a partial index over two, silence on three. Model quality does not change that mapping. Neither does context length, nor the care taken over the ontology.

The word "brain" is doing heavy lifting, and it deserves resistance. Walsh and Ungson spend their opening pages on the anthropomorphism problem for good reason. Once you accept that an organisation has a memory the way a person does, you stop asking where that memory physically sits.

What the record systematically loses

An archive is not a random sample of what happened. It is a biased one, and the bias runs in a direction that hurts.

Chris Argyris and Donald Schön separated two things in Theory in Practice: Increasing Professional Effectiveness (Jossey-Bass, 1974). The espoused theory is what a person says they would do, and why. The theory-in-use is what actually governs the behaviour, and it is often unknown even to the actor. The gap between them is not dishonesty. It is the normal condition of skilled work.

A corpus of company documents is a corpus of espoused theory. The memo says the supplier was chosen on total cost of ownership. The theory-in-use was that the head of operations had been burned by the other vendor in 2023 and would not say so in writing. Both are true. Only one is retrievable.

The second bias is selection. Things get written down when somebody has to approve them, defend them, or bill for them. So documentation is densest where a regulator, an auditor or a client demanded it, and thinnest where the work is pure judgement. The archive over-represents the parts of the business that were already legible, and a brain built on it inherits that shape.

The third bias is decay. Artefacts survive; reasons do not. The contract is still in the drive five years later, and the argument that produced clause 7 lives in the head of somebody who has since moved to a competitor. So the graph ends up rich in objects and poor in causes. The cause was what you needed.

The barrier was never access

If the missing bins were merely inconvenient, the pitch would still hold. You could argue that indexing the archive removes the main obstacle to reusing what the firm knows. There is a measurement on exactly that question, and it is thirty years old.

Gabriel Szulanski studied 122 transfers of best practice inside eight companies, 271 observations in total. He published the results as "Exploring Internal Stickiness: Impediments to the Transfer of Best Practice Within the Firm" in the Strategic Management Journal 17, Winter Special Issue 1996, 27–43. The conventional view at the time blamed motivation. People hoard, departments compete, nobody shares.

The data pointed elsewhere. The largest barriers were the recipient's lack of absorptive capacity, causal ambiguity, and an arduous relationship between source and recipient.

Read the pitch against that finding. A company brain is an excellent answer to the problem of finding and asking. Causal ambiguity means nobody can state precisely why the practice works, so there is no correct text to retrieve. Absorptive capacity means the recipient lacks the prior knowledge to use the answer once it arrives. An arduous relationship is a problem of trust between two people, and a search box is not a party to it.

Michael Polanyi named the underlying limit in The Tacit Dimension (1966), on page 4: "I shall reconsider human knowledge by starting from the fact that we can know more than we can tell."

The standard reply is that tacit knowledge can be converted. Ikujiro Nonaka and Hirotaka Takeuchi built the SECI model around that claim in The Knowledge-Creating Company (Oxford University Press, 1995), with externalisation as the step where tacit becomes explicit. It is the licence under which knowledge management has operated for thirty years. It is also contested. Stephen Gourlay argued in "Conceptualizing Knowledge Creation: A Critique of Nonaka's Theory" that the conversion modes lack evidence which cannot be explained more simply. The paper ran in the Journal of Management Studies 43(7), 2006, 1415–1436, and it holds that the framework leaves out knowledge that is inherently tacit.

Gourlay may be wrong. The point for a buyer is different. "We will externalise the tacit knowledge into the graph" is a contested research programme sold as an implementation step, and the plan rarely says which of the two it is.

Permissions are part of the architecture

Every serious product in this category enforces access control, and it should. A brain that leaked the compensation review to the whole company would last one afternoon.

Follow that requirement to its consequence. The system answers per viewer. Two people asking the same question get different answers, and neither of them can tell which parts were withheld. The union of what the platform holds is available to nobody, including the person who signed the contract.

So "everything your organisation knows" is false even for the codified part. What any individual can reach is bounded by clearance, and clearance follows the org chart. The answers reproduce the hierarchy that produced the documents.

None of this is a flaw in the products. It is the correct behaviour, and it quietly refutes the sentence used to sell them. A single source of truth that renders differently for every reader is a permissioned index. The word for that already exists.

What the measurements say

Until recently there was no benchmark for this task on company-internal data. That is an odd gap given how much is being sold into it. EnterpriseRAG-Bench closes part of it. The authors are Yuhong Sun, Joachim Rahmfeld, Chris Weaver, Weijia Chen, Roshan Desai, Wenxi Huang and Mark H. Butler, and the paper is arXiv:2605.05253, posted 5 May 2026 and revised on 19 May. The corpus holds roughly 500,000 documents across nine enterprise sources: Slack, Gmail, Linear, Google Drive, HubSpot, Fireflies, GitHub, Jira and Confluence. There are 500 questions in ten categories.

Three retrieval approaches were run against it. BM25, a keyword ranking function from the 1990s, scored highest on correctness at 68.8%, with 68.4% document recall. Vector search over OpenAI's text-embedding-3-large reached 51.4%. An agent exploring the corpus as a filesystem landed at 60.6% correctness, with the best completeness at 61.1%.

Vector search also lost the semantic category, at 32.8%. That category was built to favour embeddings. Whatever else the result means, the default architecture of most company-brain products was beaten by lexical matching on the kind of data those products are aimed at.

The category breakdown is more instructive than the totals. Recognising that the answer is absent scored 100% across all three systems. Reconciling conflicting sources reached 80 to 90%, and reasoning inside a single document 75 to 85%. Then it falls away. Project-level questions ran from 37.5 to 60%. Completeness sat at 35 to 40% for BM25 and vector search.

Read down that list and a pattern appears. The system performs when the question is phrased in the vocabulary of the record and the answer sits in one place. It degrades when the answer must be assembled, and when the asker uses their own words instead of the archive's. That is absorptive capacity restated in engineering terms. You get a good answer when you already know enough to ask in the right language.

The authors' caveats deserve as much weight as their numbers. The corpus is synthetic, produced by a generation pipeline and then degraded with deliberate noise: 5% of documents relocated at random, 3% misfiled by a model, near-duplicates carrying conflicting facts. Documents are flattened into JSON key-value pairs, which the authors say does not capture the nested structure of real enterprise documents. Gold answers are described as revisable hypotheses rather than fixed ground truth, because exhaustive annotation at half a million documents is not feasible.

Every one of those caveats makes the benchmark friendlier than production. A clean generated corpus with known answers is the easy version. Roughly two answers in three is the ceiling observed on it.

One more result belongs here, because it kills the obvious workaround. If retrieval is the weak step, put the whole corpus in the context window. Nelson Liu and colleagues measured what happens then, in "Lost in the Middle: How Language Models Use Long Contexts" (Transactions of the ACL, 2024). Accuracy traces a U-shaped curve. It is highest when the relevant passage sits at the start or the end of the context, and it degrades in the middle, including in models built for long contexts. A larger window relocates the retrieval problem.

SourceUnit of analysisResultWhat it does not show
Szulanski, SMJ 1996Transfer of a best practice between units122 transfers, 8 firms; top barriers were absorptive capacity, causal ambiguity, arduous relationshipAnything about software; the study predates all of it
EnterpriseRAG-Bench, 2026Question over a synthetic 500k-document corpusBM25 68.8% correct, vector 51.4%, agent 60.6%; vector 32.8% on semantic questionsPerformance on real corpora; the gold labels are revisable
Liu et al., TACL 2024Position of the relevant passage in contextU-shaped accuracy; degradation in the middle, long-context models includedThat the effect survives unchanged in current models
Buechsenschuss et al., 2026Employee network position and self-reported outputCollaboration centrality +7.77 against +1.12 in control (p < .001)Whether the answers people got were correct
METR, 2025Task completion time, 16 developers, 246 tasks19% slower with AI tools; participants estimated 20% fasterA general law; small sample, tools have moved since

The audit, worked through

Abstract arguments about codification are easy to nod along to and hard to act on. So here is the exercise I would run before signing anything, worked through on an invented firm. The numbers below are illustrative. They are not a case study, and nobody measured them.

Take a services company of 140 people. List the last ten decisions that moved real money. For each one, write a single sentence giving the actual reason it went that way. Then check whether that sentence exists in any system the platform would ingest.

A legal settlement: the reason is in the board minutes, because the board had to approve it. Recorded. A contract renegotiated at a 12% discount: the CRM records the approval and the approver, and says nothing about why the discount was granted. Not recorded. A senior hire: the scorecard is in the applicant system, and the decisive factor was a reference call that nobody wrote up. Not recorded.

A supplier switch: the memo cites total cost of ownership, and the operations lead had a bad experience with the incumbent in 2023. Half recorded, and the recorded half is the espoused one. A module rewritten from scratch: the design document argues from maintainability, while the immediate cause was one engineer refusing to keep patching it. Not recorded. Five remain: an office move, a paused product line, a retention deal, a walked-away tender, a price rise on the smallest tier. Three of them have a written rationale. Two have a chat thread that names the decision and skips the cause.

Count comes out at four out of ten with a recorded reason. That is the ceiling on what any brain can know about how this company decides. It is not a bad score for a firm of that size, and it is nowhere near "everything your organisation knows".

Now notice what a demo would do with the same ten. Asked what was decided, it answers ten out of ten, with citations. The gap between those two scores is the entire subject of this essay, and it is invisible in every product evaluation that only asks questions with documented answers.

The strongest evidence against this essay

If a company brain answers the easy questions and cannot hold the hard ones, the natural prediction is atrophy. People stop asking colleagues. The human network thins. The organisation leans on the archive it can query, and less on the people who understand it. That prediction has now been tested, and it failed.

Ralf Buechsenschuss, Irmela Koch-Bayram, Torsten Biemann and Phanish Puranam ran a randomised field experiment at a technology services firm in Central Europe. They reported it as "GenAI Adoption Increases the Density of Knowledge and Collaboration Networks", INSEAD Working Paper 2026/22/STR, dated 1 April 2026 and posted as SSRN 6028034. Three hundred and sixteen employees in 42 teams took part. Teams were randomly assigned either an assistant grounded on the firm's own knowledge bases and communication data through retrieval-augmented generation, or work as usual. Outcomes were measured over two weeks before and three months after.

Collaboration-network degree centrality rose by 7.77 in the treated group against 1.12 in the control, a difference of 6.65 at p < .001. Knowledge-network centrality rose 5.21 against 0.84, a difference of 4.37 at the same threshold. Specialists gained more in knowledge centrality. Generalists gained more in output. Satisfaction with knowledge access improved.

People talked to each other more. The authors read the assistant as lowering the cost of coordination and raising the value of individuals as sources of knowledge, and their reading fits the data they collected.

Two limits sit on the result, and I would want both on the table before treating it as a purchase justification. It is a working paper under review, covering one firm over three months. More importantly, the primary outcomes are self-reported by employees: interactions, projects handled, satisfaction. That is the exact measure which failed in METR's randomised trial of experienced open-source developers (arXiv:2507.09089, July 2025). Participants were 19% slower with AI tools across 246 tasks, while estimating afterwards that they had been 20% faster.

What survives contact between the two studies is a distinction. The INSEAD experiment measures how the assistant changed the shape of the network, and how people felt about access. It does not measure whether the answers were right, and neither does any figure a vendor has shown me. Both things can hold at once. A grounded assistant makes an organisation more connected, and it still cannot contain what nobody recorded.

The second objection is stronger, and this one I accept. Enormous amounts of enterprise knowledge are codified already, or codify cheaply: contracts, tickets, policies, invoices, code, specifications, the history of a customer account. For those domains a good retrieval layer is worth its price. The EnterpriseRAG-Bench figures look weak in the abstract, and they describe work that used to take a person an hour of reading. The boundary here is about job selection, the same shape as the argument in Mainframe and the boundary of GenAI competence. Same instrument, different task, wildly different value.

A test you can run before you sign

The decision is not whether to buy. It is how much of your problem you are buying a solution to, and the difference is measurable in about a week.

  1. Audit ten decisions. Run the exercise from the worked example above on your own last ten. Count the ones whose reason exists in writing. That fraction is your ceiling.
  2. Log the questions. For one week, have your team record the questions they ask each other directly. Sort them into three piles: answerable from a document, answerable only by a specific person, answerable by nobody today. Pile one is what you are actually buying.
  3. Pilot against pile two. Vendor demos are built from pile one, which is why they impress and why they predict nothing. Run the trial on questions that currently need a person, and have that person mark the answers for correctness.
  4. Price pile three as an organisational problem. Questions nobody can answer are not a retrieval failure. They are a record-keeping failure, and the fix is a habit: decision records, incident write-ups, a written reason for every rejected option.

That last habit has the longest half-life, and nobody sells it, because it cannot be installed. The cost of a company brain is a line in the budget. The cost of writing down reasons is paid by senior people in fifteen-minute increments, forever. It is also the only thing on the list that raises the ceiling in step one.

Which lets you read the vendor's sentence more usefully. The archive is indeed the memory that never resigns. Whoever knows why the archive says what it says can still hand in their notice on Friday.

What the boundary is worth

A company brain is a good product with a false name. It indexes the residue of work. That residue has real value, and it is a different object from what the organisation knows.

Buying it as an index is a sound decision at a defensible price. Buying it as a mind is how you end up with a system that answers every question except the ones worth asking.

Before the next demo, run step one. Count how many of your last ten expensive decisions have their reason written down anywhere. Whatever that number turns out to be, no vendor can sell you a way past it.

Sources

  1. primaryF.A. Hayek, "The Use of Knowledge in Society," American Economic Review 35(4), September 1945, 519–530 — knowledge "not given to anyone in its totality" (§I); "the knowledge of the particular circumstances of time and place" (§III); knowledge that "by its nature cannot enter into statistics" (§IV); decisions left to those familiar with the circumstances (§V). Written against Oskar Lange's case for a planned economy. Applying the argument inside a single firm is this essay's extension, not Hayek's.
  2. primaryR.H. Coase, "The Nature of the Firm," Economica, New Series, 4(16), November 1937, 386–405 — firms exist because internal organisation is sometimes cheaper than transacting in a market; used here as the limit on transferring Hayek's conclusion indoors.
  3. primaryJames P. Walsh & Gerardo R. Ungson, "Organizational Memory," Academy of Management Review 16(1), 1991, 57–91 — organisational memory as stored information from an organisation's history brought to bear on present decisions; five retention facilities (individuals, culture, transformations, structures, ecology) plus external archives; the anthropomorphism warning.
  4. primaryChris Argyris & Donald A. Schön, Theory in Practice: Increasing Professional Effectiveness (Jossey-Bass, 1974) — espoused theory as what a person says governs their action, theory-in-use as what actually governs it, and the gap between the two.
  5. primaryGabriel Szulanski, "Exploring Internal Stickiness: Impediments to the Transfer of Best Practice Within the Firm," Strategic Management Journal 17, Winter Special Issue 1996, 27–43 — 271 observations of 122 transfers in eight firms; absorptive capacity, causal ambiguity and an arduous relationship outrank motivational barriers.
  6. primaryMichael Polanyi, The Tacit Dimension (Doubleday, 1966), p. 4 — "we can know more than we can tell".
  7. primaryIkujiro Nonaka & Hirotaka Takeuchi, The Knowledge-Creating Company (Oxford University Press, 1995) — the SECI model and externalisation as the conversion of tacit knowledge into explicit knowledge.
  8. primaryStephen Gourlay, "Conceptualizing Knowledge Creation: A Critique of Nonaka's Theory," Journal of Management Studies 43(7), 2006, 1415–1436 — the conversion modes lack evidence that cannot be explained more simply; inherently tacit knowledge is omitted from the framework.
  9. primaryYuhong Sun, Joachim Rahmfeld, Chris Weaver, Weijia Chen, Roshan Desai, Wenxi Huang & Mark H. Butler, "EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge" (arXiv:2605.05253, 5 May 2026; revised 19 May 2026) — ~500,000 synthetic documents across nine enterprise sources, 500 questions in ten categories; BM25 68.8% correctness and 68.4% document recall, vector search 51.4%, filesystem agent 60.6% with 61.1% completeness, vector 32.8% on the semantic category; category spread from 100% on absent answers down to 35–40% on completeness; the authors' own limitations on synthetic generation, JSON flattening and revisable gold labels.
  10. primaryNelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni & Percy Liang, "Lost in the Middle: How Language Models Use Long Contexts," Transactions of the Association for Computational Linguistics 12 (2024) — U-shaped accuracy by position of the relevant passage, long-context models included.
  11. primaryRalf Buechsenschuss, Irmela Koch-Bayram, Torsten Biemann & Phanish Puranam, "GenAI Adoption Increases the Density of Knowledge and Collaboration Networks: Evidence from a Field Experiment," INSEAD Working Paper 2026/22/STR (1 April 2026, SSRN abstract 6028034) — 316 employees in 42 teams at a Central European technology services firm, RAG-grounded assistant, two-week baseline and three-month follow-up; collaboration centrality +7.77 against +1.12 (p < .001), knowledge centrality +5.21 against +0.84 (p < .001). (Working paper under review; primary outcomes self-reported by employees.)
  12. primaryMETR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (arXiv:2507.09089, July 2025) — 16 developers, 246 tasks, 19% slower with AI tools against a self-estimated 20% speed-up; cited here for the divergence between self-report and measurement.
  13. secondarySphereIQ, "Company Brain — Enterprise Knowledge Platform" (product page, accessed 21 August 2026) — the quoted headline, the component list (knowledge graph and ontology, semantic layer, hybrid retrieval with citations, episodic memory, governance plane) and the unattributed customer figures: 35,000+ operational documents at "a global aviation operator", 60× faster resolution, six hours to seven seconds at "a professional-services firm". Vendor marketing, cited as evidence that the pitch exists, not that it works.