Maintenance Detective: Gremlins in the Drives
First time it happened, no big deal. We just power cycled the machine and it went away.
One drive in a bank of them dropped out mid-shift: the fault said something about "F111 Safety Hardware". Machine just shut down and wouldn't run. No E-stop pressed. No gate opened. Safety relay had all green lights, no issue there. After power cycling, the machine fired up and ran the rest of the shift with zero complaints.
Then nothing for three weeks.
Then it came back. Once a week for a few weeks. Mostly the same drive but always the same bank. It always reset cleanly on a power cycle. Techs swapped a drive just to rule it out... didn't change a thing.
Then it stopped being "once a week." It was twice a day. Then four times. Random drives, random times. We just kept power cycling.
Then the whole bank dropped at once and the line was down for a shift before we found the issue.
No E-stop activation in the logs. No gate ever opened. We just kept getting those faults.
What was actually going on here? And why does something that "resets clean" every single time get more frequent instead of just... failing?
This is a 2 part puzzle:
What likely electrical failure am I seeing that would fail exponentially more frequently as it progresses?
In this example, where is the issue?

Answer:
The STO (Safe Torque Off) enable on that drive bank runs as a daisy-chained safety loop (common configuration)... one signal on each line that must match, wired through every drive's dual-channel safety input in series. That matters, because it means a fault anywhere in either channel doesn't just take out one drive. It potentially takes out everything downstream of wherever the break is, which is exactly why the "faulting" drive kept moving around the bank.
Swapping drives never fixed it because no drive was ever the problem.
Now, one drive will typically fault first when voltage sags on the STO channel, so it's not uncommon for the same drive to always trip first.
Because one STO line was experiencing a slow voltage drop... and the machine stopped when the first drive faulted... you might always see the symptom on the same drive.
Here's the tell that should point you at wiring instead of an actual safety event: STO runs dual-channel for a reason: two independent legs that have to agree. When a real E-stop or gate opens, both legs drop together and you get a message that says "Safety Open" (on a Powerflex 525, it's different on every drive).
When one leg blips open for a few milliseconds while the other stays high, the drive doesn't just see "STO commanded off", it declares a channel discrepancy fault. That's a different fault code, and it's the fingerprint of a wiring problem, not a safety device doing its job.
Why it got worse instead of just failing once:
A bad connection will generate heat and fretting corrosion, which accelerates the worsening of the connection. The exponential increase in frequency is the signature of wiring issues.
When you see an accelerating electrical condition like this it is often a loose wire and that's something that should be checked when you see this signature.
Taking a screwdriver and tightening every lug in a cabinet is often easier than troubleshooting (and a very smart thing to do on an annual PM).
In power circuits, the contact area gets hot and burns as the connection gets loose. In lower power circuits, fretting corrosion oxidizes surfaces slowly. The result is the same. You get an accelerating drop in voltage as the contact point builds resistance. Once it gets enough voltage-drop to create a real problem, you get to deal with a condition that'll start intermittently and accelerate.
How it was actually found:
Logically, if you look up F111 on the Powerflex 525 and realize it means one of the STO inputs is low... that would take you to check the voltage across the STO daisy-chains. One would probably be low and variable because of a bad connection. You would then chase that back to a safety relay output (or it's power source)... and you would find a bad connection somewhere along the way.
The point:
A loose connection frequently doesn't fail once and announce itself. It fails on a curve: rare, then occasional, then constant, then dead... and the corrosion/burn from each event is what shortens the distance to the next one.
The description of how this plays out on an STO control line just seemed like a logical way to illustrate it.
Resetting a recurring issue doesn't solve it. It only buys time until it happens again.
180 OPEX helps manufacturers identify and address the conditions behind recurring issues before they become accepted as part of the operation.






Comments