The clearest sign that a problem was fixed at the wrong level is that it comes back. Not identically, which would be too easy to spot, but in a family resemblance: the same kind of failure, in a nearby place, with a different last step. Each time it is investigated, a cause is found, a fix is applied, and everybody moves on, and the accumulation of those fixes is visible in any long-running operation as a layer of workarounds on top of something nobody has been back to.
Section 5 of the PMBOK® Guide Eighth Edition includes the analytical techniques a team uses for this work, and the technique matters considerably less than the discipline around it. Most failed analyses fail for the same two reasons, and both are about stopping.
The symptom is what was noticed. The report was late, the feed dropped, the batch failed, the invoice went out wrong. It is the thing that raised the alarm and it carries almost no information about why, which is why a fix aimed at it produces a workaround: a manual check, a second approval, somebody arriving early on Fridays.
The sequence is what happened, in order. Establishing it is genuinely useful work and it is not the same as finding a cause, because the last event before a failure is rarely the reason for it. Sequences are where investigations produce their most confident wrong answers, since the final step is always the most visible and usually belongs to whoever was on shift.
The cause is the condition that made the sequence possible, and the test is counterfactual: if this condition had been different, would the failure still have occurred? A cause you can act on is one where the answer is no, and where changing it is within somebody's authority. That second half matters, because an analysis that terminates at a cause nobody can change has produced an explanation and not a fix.
At human error. When the answer is that somebody made a mistake, the analysis has stopped one step short, because people make mistakes continuously and most of them are caught by something. The useful question is what made this mistake easy to make and what allowed it to reach the outcome undetected, and those two questions almost always point at a design, a procedure or a check that does not exist.
At the first plausible cause, under pressure to resume. Investigations happen while something is broken and somebody senior is asking when it will be working again. The first explanation that fits gets adopted because it permits the restart, and the review that would have tested it never happens, because by then the thing is running. Building a short, scheduled second look into the process, a week later, with the pressure gone, catches a surprising proportion of these.
A broadcaster's regional playout had a recurring fault: intermittent loss of audio on one channel at the top of the hour, perhaps once a fortnight, never reproducible on demand.
The first investigation found a faulty patch on a router and replaced it, which was a real fault and not the cause; the loss returned three weeks later. The second found an operator error in a changeover procedure and issued a clarified instruction, together with a strip of tape across two faders to stop them being moved, which is how the desk acquired its first workaround. The third found a timing tolerance in an automation profile and adjusted it, and a cable tie went on a switch to hold it in position while the adjustment was tested. The tie stayed for four months.
Each fix addressed the last step in a sequence. What none of them addressed was the condition underneath: the automation system and the audio router held separate copies of the channel configuration, and a change made in one did not propagate to the other. Every failure had followed a configuration change somewhere in the previous fortnight, which nobody had noticed because the changes were made by a different team, logged in a different system, and none of the three investigations had looked outside the gallery.
The fix, once the condition was named, was unglamorous: a single source for the channel configuration, with the router reading from it, and a check at handover that the two agreed. It took about three weeks of work and it ended the fault permanently. The tape and the cable tie stayed on the desk for another six months, because nobody could remember what they were for and nobody wanted to be the person who removed them.
Change a condition, not a behaviour. A fix that depends on people remembering to do something new has added a demand to an already busy job, and it degrades. A fix that removes the possibility, or that surfaces the discrepancy automatically, holds. Where only a procedural fix is available, it needs a check attached to it, because otherwise nobody will know when it has stopped being followed.
Verify by the symptom. The measure of a successful analysis is that the original problem stops occurring, and that has to be looked at deliberately some months later. Teams close investigations when the fix is implemented, which is the point at which the useful evidence begins rather than ends.
Decide the level to fix at. A cause found in one system is often present in four, and the decision about whether to fix the instance or the class belongs to somebody with a view of all of them. Project managers tend to fix the instance, because that is what they are accountable for, and the honest move is to fix the instance and pass the class upwards with the evidence attached.
For a PMP® candidate, what helps is seeing that a recurring problem indicates an analysis that terminated early, so a situation where the same issue reopens is asking what condition was never changed. A response that applies a tighter control at the point of failure is adding a workaround with a formal name. A structured PMP exam preparation course is full of situations where every previous fix was correct and the problem persists.
This exercise is worth doing while nothing is on fire. Open the issue log and look for items that have been closed and reopened, or closed twice in different words. Each of those is an analysis that stopped somewhere, and the second look is considerably cheaper now than it was during the incident, because nothing is currently broken and nobody is waiting on the answer.
Knowing when an investigation has stopped too early, and holding it open while somebody senior wants the service back, is a judgement with real pressure attached. Omega's PMP® Exam Preparation works through problem analysis where the obvious answer is available and wrong.
The analytical techniques behind this work are collected in the PMBOK® Guide Eighth Edition.
Ad · Amazon affiliate link.
A177: Five Whys vs Fishbone Diagram
A178: SWOT Analysis in Project Management
A183: Cost-Benefit Analysis: Why the Cheapest Option Is Not Always Best
A180: Facilitation Techniques for Difficult Project Conversations
A184: Trend Analysis vs Variance Analysis
PMP and PMBOK are registered marks of the Project Management Institute, Inc.