Ideas

When Correction Becomes Part of the Failure

A system does not become fragile only when it starts failing. A more important transition may happen earlier: when its attempts to correct problems stop restoring the system and start helping those problems persist.

I called that threshold the Drift Point.

A simple way to think about it is:

Problem → correction → recovery

becomes:

Problem → correction → accommodation

and eventually:

Problem → correction → reinforcement

At that point, the organization may still recognize problems and devote substantial effort to fixing them. What it is losing is the ability to return to a healthier state.

The phrase that first helped me see it was simpler:

Failures start talking to each other.

The Drift Point

My working definition is:

The Drift Point is the transition where a system’s corrective mechanisms stop reliably reducing deviation and begin preserving, transferring, or reinforcing it.

I first arrived at the idea while writing science fiction. I needed a way to describe a system that had not simply broken, but had reorganized itself around accumulated failures. Then I started seeing versions of the same pattern in organizations, software, product architecture, governance, and AI.

The interesting question stopped being when does a system fail? It became: when does a system stop being able to correct itself?

Failure is usually late

We tend to recognize failure by its visible outcome. A company misses its targets. A platform becomes unreliable. Customers leave. A product collapses under accumulated complexity.

Those events matter, but they are often late.

A healthy system can be wrong. It detects a deviation, responds, learns, and moves closer to an acceptable state. A drifting system behaves differently.

A problem appears. The organization responds. The response introduces a workaround. The workaround creates a dependency. The dependency creates another problem. Eventually, the intervention itself becomes part of the structure sustaining the original condition.

That transition matters because a system can continue performing reasonably well after its ability to recover has begun deteriorating.

A green dashboard does not necessarily mean a healthy system.

Drift is not the same as decline

Organizations drift constantly. Products accumulate exceptions. Teams work around limitations. Processes change. Strategies become misaligned with reality.

None of that automatically means a system is approaching failure. The more useful distinction is recoverability.

Imagine two organizations facing the same bad decision. The first recognizes it, changes priorities, removes the failed initiative, and recovers. It drifted, but its corrective capacity remained intact.

The second recognizes the same problem and launches a corrective program. The program creates a team, processes, technology, metrics, incentives, and budget. Careers become attached to it. Other systems begin depending on it.

Now reversing the original direction means unwinding much more than the original decision. The organization may know that something is wrong while becoming progressively less capable of acting on what it knows.

The system has begun reorganizing itself around the deviation.e may tell us more about future resilience than current performance does.

Follow the intervention

There is substantial prior work around neighboring ideas. Jens Rasmussen modeled safety as a dynamic control problem in which local adaptations and organizational pressures can push complex systems toward unsafe boundaries. Diane Vaughan’s analysis of the Challenger launch decision showed how repeated acceptance of anomalies could normalize behavior that later appeared obviously dangerous. Chris Argyris distinguished between correcting errors within an existing set of assumptions and learning that questions the assumptions themselves.

Research on critical transitions offers another useful parallel: some systems recover more slowly from disturbances as they approach a tipping point.

Drift Point is not meant to replace those ideas. What interests me is using the sequence of correction itself as the unit of analysis.

Don’t just study what went wrong. Study what happened after the system noticed it, and then what happened after that.

Correction can become accommodation

Consider a product organization with a complicated checkout. A new requirement adds a field. Compliance adds a step. Growth adds an upsell. Support adds explanatory copy.

Each change is locally reasonable.

Customers begin struggling, so UX simplifies the presentation. Progressive disclosure helps. Better hierarchy helps. Copy improves. Then another requirement arrives, and another design pattern contains it.

The local design work can be excellent while the underlying system becomes increasingly difficult to simplify. The corrective mechanism has quietly shifted from:

remove complexity

to:

accommodate complexity

That is why Drift Point is not another name for bad design. Good local work can participate in system-level drift.

Competent people can create systemic drift

Drift does not require stupidity, negligence, or bad intentions. It can emerge from competent people solving local problems.

Operations creates a workaround. Engineering abstracts it. UX makes it easier to use. Management adds oversight. Each intervention may be entirely reasonable from where that team sits.

Collectively, they can make the original condition increasingly expensive to reverse. That exposes an important distinction:

Individual competence and systemic corrective capacity are not the same thing.

Smart people paying attention does not guarantee that the larger system remains capable of correcting itself.

When failures start talking

Early in a system’s decline, problems may still be relatively independent. An unreliable application can be fixed. A broken process can be redesigned. A misleading metric can be replaced.

But as drift progresses, those problems begin interacting.

An unreliable application creates manual work. The manual work changes staffing. Staffing creates process rules. Those rules distort measurement. Distorted measurement shifts incentives. The new incentives place more pressure on the application.

Now changing one component affects several others. The organization no longer has a collection of problems.

It has a problem system.

That is what I mean when I say:

Failures start talking to each other.

Increasing coupling helps explain why correction becomes less effective. Every intervention must now unwind more relationships than the one before it.

Watch what corrective effort produces

One way to recognize a system approaching its Drift Point is to pay attention to the changing relationship between corrective effort and durable improvement.

Early:

small corrective effort → meaningful improvement

Later:

larger effort → smaller improvement

Near the Drift Point:

large effort → little durable change

Past it:

large effort → more complexity, dependency, or lock-in

This is not meant as a literal mathematical curve. The useful signal is simpler.

If every corrective program requires more coordination, governance, tooling, people, exceptions, and explanation while producing less underlying change, then adding another intervention may not be solving the problem. Declining response to corrective effort may itself be evidence that the system has changed.

A green dashboard can miss this

Most organizational dashboards measure current output: revenue, conversion, throughput, availability, incidents, cost. Those measures matter, but they tell us much less about whether the system can still change when credible evidence contradicts its assumptions.

A profitable business can become harder to redirect. A reliable platform can accumulate dependencies faster than it removes them. Current performance can remain green while recoverability deteriorates.

The questions become different:

Can bad news still change priorities?

Can temporary exceptions still be removed?

Does corrective work reduce the underlying complexity, or merely make it easier to tolerate?

These are questions about corrective capacity, not simply performance.

AI can stabilize the wrong state

AI makes this distinction more important. Imagine an organization automating a dysfunctional workflow.

Friction falls. Throughput improves. Fewer people have to deal with the painful parts manually. More processes begin depending on the workflow because it is now cheaper and faster.

Operationally, the intervention succeeds. Structurally, the original workflow may have become harder to remove.

AI did not necessarily correct the system. It may have stabilized its current state.

That leads to a distinction I think matters:

An intelligent organization is not one that knows more. It is one that can still change because of what it knows.

More analysis, documentation, forecasts, recommendations, and automation do not automatically create corrective capacity. Information can move faster while the organization becomes harder to change.

A practical Drift Point test

I don’t think this needs to begin as another maturity model or scorecard. Start with one question:

Is this intervention restoring the system, or helping it tolerate the deviation?

If you want to go further, ask:

State
What is the system trying to preserve?

Signal
What evidence says it is moving away from that state?

Correction
What happens when the system receives that evidence?

Propagation
Does the intervention change the underlying condition, or mainly accommodate it?

Coupling
Are previously independent problems becoming dependent on one another?

Recoverability
Could the system still return to a healthier state without disproportionate disruption?

The important move is temporal. Don’t take a snapshot. Follow the sequence.

What did the organization know? What did it do? What happened next?

Over time, the pattern may move from:

Problem → correction → recovery

to:

Problem → correction → accommodation

and eventually:

Problem → correction → reinforcement

That transition is the thing I’m interested in.

Maybe the point is not a point

The name Drift Point sounds precise. Real systems rarely are.

There may be no clean mathematical boundary. It may be better understood as a threshold region where behavior changes: on one side, corrective intervention tends to restore capability; on the other, correction increasingly preserves the changed state.

The exact boundary may only become obvious in retrospect. The distinction can still be useful if the system behaves differently on either side.

Drift becomes direction

Most organizations eventually make bad decisions. That isn’t exceptional.

The more interesting question is why some systems recover while others gradually build themselves around those decisions. That is what makes the Drift Point useful to me. It shifts attention away from the original mistake and toward something more fundamental: the system’s relationship with correction.

A resilient system can be wrong. A drifting system can know that it is wrong. A system approaching its Drift Point can expend enormous effort trying to correct itself and still move further in the same direction.

The catastrophe comes later.

The more important transition may be when drift becomes direction.