← Writing

MBA · ∩

Stop fighting symptoms: the structure is the problem

A founder's day is mostly firefighting: this outage, that angry customer, the deploy that broke again. The trap is that every fire feels like the problem, when the fires are outputs of a structure that keeps producing them. Systems thinking makes an uncomfortable claim — recurring problems are not bad luck, they are your system working as built — and it points at the far cheaper intervention: change the structure, not the symptom.

Behaviour comes from structure

Donella Meadows put the core idea plainly: the behaviour of a system arises from its structure — its stocks, its flows, and above all its feedback loops. If a problem keeps coming back after you fix it, you did not fix it; you relieved a symptom while leaving intact the structure that generates it. Meadows called the tendency of systems to defeat well-meaning symptomatic fixes "policy resistance," and it is why so much effort produces so little durable change.

Peter Senge catalogued the recurring shapes of this failure as systems archetypes. The most common in a company is "shifting the burden": a symptomatic fix relieves the pressure so effectively that the underlying problem never gets addressed, and worse, the organisation's capacity to address it atrophies because the quick fix is always there. You discount to stop churn, the discount works, and you never fix the onboarding that caused the churn — and now you are dependent on the discount.

The move: find the structure that keeps producing the fire

The operational discipline is to treat a recurring symptom as a question about structure. Constant on-call fires are not an on-call problem; they are a fragile-deploy problem. Repeated bad hires are not a run of bad luck; they are a hiring-system problem. Churn is rarely a pricing problem; it is usually a value-delivery or onboarding problem wearing a pricing costume. In each case the symptomatic fix — more heroics, another interview round, a discount — treats the output, and the durable fix changes the thing generating the output.

The original System 1.0 problem has not been solved, and a new problem — dependence — has been added.Donella Meadows, on shifting the burden

This is the whole logic of resilience work, which is my day job: you do not make a system reliable by responding faster to failures, you make it reliable by engineering the structure so those failures stop recurring. Meadows' leverage-points insight is that the highest-leverage interventions are not at the level of events (the fire) but at the level of the rules, incentives, and goals that shape the flows — and those are exactly the levels founders are most tempted to skip, because they are slow and invisible while the fire is fast and loud.

A worked fix: the deploy that keeps breaking

Take the most ordinary recurring fire: every few weeks a deploy takes down production, someone scrambles, it gets patched at midnight, and everyone moves on relieved. The symptomatic response is to get better at the scramble — faster alerts, a slicker rollback, a hero who knows the incantation. Each of those treats the output, the outage, and each makes the team a little more dependent on the heroics that keep it alive. Meadows would call this shifting the burden: the quick fix works well enough that the structure generating the outages never has to be examined, and the organisation's capacity to fix it properly quietly atrophies.

The structural question is different: what about the system reliably produces this outage? Usually the answer is a fragile deploy pipeline — no staging that mirrors production, no automated tests on the path that keeps breaking, a manual step that a tired human performs differently each time. Those are structure, not events, and they are where the leverage is. Fix the pipeline and the class of outage stops recurring; keep fixing outages and you are signing up to fix them forever, faster each time.

But here is where systems purism has to yield to triage, because the honest sequence matters. When the site is down at midnight, you do not open a discussion about pipeline architecture; you grab the bucket, restore service, and keep the system alive. The structural fix is tomorrow's work, scheduled deliberately once the fire is out, not philosophised about while it burns. The discipline is to actually schedule it — to not let the relief of a working patch become the reason the structure is never touched. Stabilise the symptom now; change the structure on purpose after, so the same fire does not book the same midnight next month.

The move generalises: constant on-call is a fragility problem, repeated bad hires are a hiring-system problem, recurring churn is usually an onboarding problem. In each case the recurring symptom is a question about structure, and the answer is worth more than any amount of getting faster at the response.

Where "fix the root cause" becomes a trap of its own

Systems purism has failure modes as real as firefighting's. Three.

One: sometimes the symptom is the emergency. When the building is burning, "let us address the root cause" is negligence. Triage comes first: stabilise the symptom, keep the system alive, and only then fix the structure. A founder who philosophises about deploy architecture while the site is down has confused the essay with the job.

Two: "the root cause" is often a myth. Complex systems rarely have a single root; they have interacting causes and feedback loops with no clean origin. The confident hunt for the one true cause can be as reductive as symptom-chasing, just slower. Meadows' own point is that you look for leverage, not for a single villain.

Three: restructuring is expensive, and optionality argues against premature systems. For a young or small company, building the elaborate system to prevent a problem that has happened once is over-engineering — the cheap symptomatic patch is correct until the pattern proves it recurs. Fix the structure when the fire is a pattern, not the first time it burns.

So when a problem comes back, stop reaching for a faster bucket and ask what structure keeps lighting the fire. Then fix that — unless the building is actually on fire right now, in which case grab the bucket first and be honest that you will owe the structural fix tomorrow.

Sources

  1. primaryDonella H. Meadows, Thinking in Systems: A Primer (2008); "Leverage Points: Places to Intervene in a System" (1999).
  2. primaryPeter M. Senge, The Fifth Discipline (1990) — systems archetypes.
  3. secondaryTaiichi Ohno, Toyota Production System (1978) — the "5 Whys" root-cause method.
  4. secondaryW. Edwards Deming, on common-cause vs. special-cause variation.