Process Decay: Why Every Process Erodes Itself From the Inside
Updated: Sep 14
Processes don't fail because people ignore them. They fail because people follow them for a while, and then a little less, and then only when it's convenient.
This is process decay, and it isn't a discipline problem. It's a predictable, mechanical outcome of how habits and procedures interact.
And it's 100% preventable.
The stabilizing force that only stabilizes some of the process
Habits are what make a process survivable in the real world. Once a step becomes automatic, a person does it without second thought: no willpower or memory required. That's a good thing. When we run by habit, humanity is able to make decisions at extremely high speed and with very few errors.
It's also the reason process decay is so selective.
A habit only forms around what gets repeated identically every time. The steps in your process that run on every single cycle, the ones with no variation, get locked in by repetition. They become part of the muscle memory of the job. They're protected.
When habits are allowed to run without interruption, they reliably repeat.
However, one of the results of high stress scenarios (rushing, repeat interruptions, fatigue, etc...) is that they interrupt habits which forces our reasoning brain into the fight. This is the part of the brain that decides to change processes.
You see, not every step in a written process is indispensable every time.
Verification steps that only catch a problem once in a while and setup steps that only matter on certain configurations could be skipped most times and nobody would know. Even cleaning steps that might not be checked every time could be skipped when someone is in a rush and it's likely nobody would notice... this time.
But what if we keep rushing and keep skipping steps? What happens to the habit?
What happens if that verification step is there to prevent catastrophe?
Why nothing stops the cut
Here's the part that makes this worse than a simple discipline lapse: when a non-critical step gets skipped, nothing happens. That's what "non-critical on this run" means. No alarm, no defect, no injury, nobody even notices. The step wasn't needed this time.
The person who skipped it gets direct, immediate proof that skipping it was fine.
And that invites us to make it normal.
Sociologist Diane Vaughan documented exactly this pattern in her study of the Challenger disaster, coining the term normalization of deviance: a deviation from procedure that produces no bad outcome gets repeated, and each repetition without consequence makes the next one easier... until the deviation stops registering as a deviation at all. It becomes "how we do it here."
Scott Snook, analyzing the 1994 friendly-fire shootdown of two U.S. Army Black Hawks, described the organizational version of the same mechanism as practical drift: the slow, steady uncoupling of what people actually do from what the written procedure says. Rules are written to avoid all the terrible things that can happen. Day-to-day work seems very far from all the potential disasters, and it's easy to imagine those things can't happen here. Operators quietly optimize around the gap, and because each individual optimization looks harmless in isolation, no one intervenes.
Both researchers are describing natural process decay. It isn't sabotage.
It's optimization, which through stress and interruptions, can be injected into a pre-existing habit almost without rational recognition of the change.
What's left when the decay finishes
Run this forward long enough without something actively holding a process in place, and you get a predictable result: every step that isn't needed on every single run eventually gets truncated. What survives is the critical-path minimum: the steps that are always necessary, because those are the only ones habit ever protected.
Everything else: the verification steps, the setup checks that only matter sometimes, the tasks that only have to be done every week or two, the redundancies built in for the exception... those erode away. The process doesn't fail all at once. It quietly loses its margin, one skipped step at a time, until the day the exception shows up and the step that was supposed to catch it isn't there anymore.
This is why post-incident investigations so often turn up a procedure that "everyone knew" wasn't being followed. It wasn't a secret. It wasn't negligence in the way people mean when they say the word. It was process decay... slow, silent, and locally rational at every single step along the way.
The fix isn't more discipline. It's more stability and process reinforcement.
You can't out-willpower this. Telling people to "just follow the procedure" doesn't work, because the whole mechanism runs beneath conscious commitment.
There are 2 pieces to the puzzle of preventing process decay.
The first is to run a calm operation with a minimum of interruptions and urgency for operators. If you're always hammering them for speed, you will keep disrupting their habits and they will cut steps where they can. That's just how that works. If you can keep them operating with habit more often, they won't turn on the decision cycle that skips steps.
The second fix is that process stabilization has to be a deliberate system, not an assumption. Any task that could be skipped without immediate consequence needs to be verified. If the operator forgets to start a machine... you'll notice. You probably don't need to be checking that. But if the operator doesn't verify something that could cause a huge recall... you need a system to find that.
The moment a non-critical step becomes reliably un-checked, it's already on its way out.
And if you don't have a system that's designed to prevent process decay... you are likely in an advanced state of process decay already.
Processes don't decay because people stop caring. They decay because nothing was built to notice the parts that were never going to protect themselves.






Comments