How to Ask Better Questions in a Root Cause Analysis

Most advice about root cause analysis is advice about method: pick a tool, go down a few levels, draw the diagram. I spend most of my time a layer below that, working with the people who will run the process to shape the criteria that tell them what to ask, and then seeing the results again from the other side when I audit systems I had nothing to do with building. The difference between an analysis that goes somewhere and one that stops early usually sits in those criteria, and it is not reliably the difference between a careful analyst and a lazy one. The safety examples below are illustrations rather than accounts of any particular engagement.

Why generic criteria produce shallow analysis

A criterion that anyone can answer without knowing much about the work is unlikely to produce depth.

"Are there any safety risks here?" is a real criterion and it appears on real forms. Everybody in the room can answer it, they can answer it without having been to the site, and the record comes back complete either way. Now put a different set next to it. Have the contractors on this site completed the orientation. Has the hazard analysis been done on the specific processes this project runs. What chemicals are on this site, and what risks come with them here rather than in general. Those are hard to answer without having gone and looked, which is most of what the specificity is for.

The generic version is not always wrong. An organization without much safety specificity to speak of may be served perfectly well by a broad question, and adding four narrow ones there buys nothing but administration. The failure is more often generality sitting on top of an operation that has real specificity nobody wrote down.

Where the questions come from

The questions worth asking come from two lists: what the organization already asks at that point in the work, and what it should be asking.

Those are the two I open with, and the second one is doing most of the work. What do you already ask usually turns out to mean what you think you should be asking and are maybe actually asking some of the time, which is a useful thing to discover out loud. What should you be asking is where the room starts arguing, and the arguing is the point; the stakeholders and the people who know the work shape the criteria, and my job is holding the purpose and the risk in front of them while they do it and asking the perspective and challenge questions that keep the expansion honest.

There is an obvious objection here: asking an organization to write down the questions it should be asking sounds circular, because if they knew, they would already be asking. The gap between the two lists is what breaks that. Very few groups are asking everything they already believe they should, and the distance between the stated practice and the actual one is visible to the people in the room in a way it is not visible from a checklist somebody else wrote.

How purpose turns into a question

Purpose is what the activity is for, and it comes in three kinds: controlling or responding to risk, satisfying a requirement or an expectation, and pursuing an opportunity.

That has to get concrete before it generates much. Nobody gets hurt is a purpose most organizations would sign, and on its own it will not tell you whether to ask about contractor orientation or about chemicals or about the third thing nobody has raised. What produces a question is working down from purpose through the sources of harm, the scenarios that present them, and what the likelihood and severity look like if the harm actually occurs. That context is what shapes a control that works rather than one that merely exists.

Where something is unlikely and would be catastrophic, you have to understand what an acceptable outcome even looks like before you can design toward it, and the criteria come out of that understanding rather than out of a template. Where something is likely and minor, the control gets normalized: wear the gloves so you do not take a paper cut on the line. The instruction has the same shape in both cases and the rigor behind it does not, and the rigor tracks the criticality of the impact rather than the size of the control.

None of this is specific to safety, which is where people tend to hear it first. Quality carries cost and regulatory implications, energy use and waste can turn into legal liabilities, and a supplier approved without a real evaluation reaches into commercial and central parts of the operation. Safety is just the domain where consequence is easiest to see, which makes it a clean illustration and a misleading boundary.

A note on proportion, because this reads as a program otherwise. My position is that every activity should have the right questions asked of it, and those questions should push at the fringes rather than sitting in the middle of what everybody already knows. How you came to know the risk can vary a lot: it can be methodical, with evaluative criteria and scoring behind it, or it can be generally known and socialized through the organization without ever having been written down. Both are legitimate inputs. What has to hold is that the questions stay proportionate and relevant to the risk and the operational context, which is a different requirement from running a formal analysis forty times.

What too much specificity costs

Criteria that are too specific tend to get answered without much thought, and past a point they cost more than the depth they add.

The failure here is not that the form is long. Criteria that are too specific get skipped, skimmed, pencil-whipped, or just answered with less rigor than the question deserved, and the trivial ones do more harm than leaving them out: they produce fatigue and cost trust in the process and in the leadership that requires it, because somebody answering a question they can see is pointless has learned something about how seriously the rest of it was meant.

The subtler cost lands on the evaluator. Blinders go on, and the person answering gets tunnel-visioned on the specific question in front of them rather than the intent behind it, which erodes their ability to discern whether the question is even relevant in this particular risk context. That is close to the opposite of what the specificity was supposed to buy. The narrow question was meant to send somebody to go and look; past a certain density it does the looking for them, badly.

So both ends fail, and they fail differently: a generic criterion produces an answer carrying little information, and an over-specified set produces a great many answers, confidently, to questions that may have stopped mattering. Neither is purely a discipline problem in the way it usually gets written up, and treating it as one will cost you the credibility you need for the actual fix.

How to tell whether a criterion still earns its place

A criterion earns its place when there is a purpose for answering it.

The answers should have implications somewhere. They inform the control parameters for the process, or they go out to a customer, a regulator, or another stakeholder, or they give management the visibility it is supposed to have. Any of those is a purpose, including the ones that terminate in a contract and inform nothing further, because satisfying a requirement is one of the three things an activity does. What you are looking for is the criterion with none of them, and where a question seems to inform nothing at all, that is worth asking about.

Criteria also go stale quietly, because a question that no longer fits the operation still gets answered. Where a process changes, change control should trigger an evaluation of what else has to change with it, and the criteria are part of that. Where nothing triggers, tracing the answers forward is what catches it. That is not a clean test and I would not present it as one; somebody who wants the trace to come back full can usually make it come back full, so it works where the person running it is willing to hear the answer.

The last piece is what happens to the reasoning after the room breaks up, and this is the least settled part of the method. The thinking can be institutionalized in the criteria definition itself, so the why travels with the what, or it can go into training material for the management teams who will maintain it. Both of those work. What I cannot tell you is how reliably either survives turnover, because the criteria tend to look fine from the outside and the reasoning is usually the part that leaves first. The analysis this all feeds is also the front half of a 

corrective action process, and a plan built on questions nobody could answer without going and looking is a different object from a plan built on questions anybody could answer from their desk.

If you want to test your own set, run two passes. Pull your investigation criteria and trace three answers forward to what they inform: a control parameter, a report going to somebody outside the process, a decision an actual person makes. Then find one criterion a person could answer without going and looking. Those two passes will tell you more about the depth of your analyses than reading the analyses will.

I would rather have your version than your agreement. When the criteria are technically complete and you still cannot tell whether the thinking happened, what do you do instead?

I design these systems with the people who will run them, which mostly means facilitating the room where the criteria get written rather than writing them and handing them over.

Frequently asked questions

What makes a root cause investigation criterion specific enough?

It should be a question that is hard to answer without going and looking at the actual work. "Are there any risks here" fails that; asking whether the hazard analysis was done on the specific processes this project runs does not. The test is whether the criterion is specific to this organization and its industry rather than to the topic in general.

Can investigation criteria be too specific?

Yes, and the failure is quieter than the generic one. Over-specified criteria tend to get skipped, skimmed, or answered thinly, trivial ones cost trust in the process, and the evaluator starts tracking the question in front of them rather than its intent, which erodes their ability to judge whether it applies in this risk context.

How do you know whether a criterion is still worth asking?

Trace the answer forward to what it informs: a control parameter, a report to a customer or regulator, a decision somebody makes, or a requirement it satisfies. Where a process changes, change control should trigger a review of the criteria with it. A question whose answer informs nothing is worth asking about.

Does every activity need a formal risk assessment before you can write criteria?

Not necessarily. What an activity needs is questions proportionate and relevant to its risk and operational context, and the risk can be known methodically with scoring behind it or generally known and socialized without being written down. Both are legitimate inputs. What is harder to defend is criteria set without reference to either.

Next
Next

How to Tell If a Root Cause Analysis Was Actually Thorough