A Finished Evaluation Doesn't Show You What Nobody Asked
A completed evaluation record shows you the answer. It does not show you how wide the question was. Every rating is filled in, every cause is identified, every action has an owner and a date. Nothing is missing and nothing is wrong. What no record carries is the set of things nobody thought to consider before writing it down — and that is where a management review program most often loses its reach.
What does a complete-looking evaluation actually hide?
Take a risk register with an impact rating on every line. Nothing is blank. The register reads as finished, and as a record it is finished.
Impact is not one thing. A given failure or change lands on product quality, on the customer commitment, on health and safety, on the environment, on the delivery schedule. A single rating collapses several distinct evaluations into one number. What the record cannot tell you is which of those areas the person actually held in mind before writing it.
This is not a claim that the rating is wrong. It usually is not. It is a claim about what a record can carry, and a record carries answers rather than the shape of the question that produced them.
Why does the scope of a question disappear?
Because the answer is the artifact and the question is not.
When someone asks what the impact of a change is, they are asking a question with a boundary around it — a set of areas they are willing to think about. That boundary becomes invisible the moment the answer is written down. An evaluation that considered one area and an evaluation that considered five produce records that look the same. Nobody reading the record afterward can tell them apart, and neither, a quarter later, can the person who wrote it.
Which is why widening an evaluation is usually not a matter of collecting more information. The information was available. The question was narrow.
Where does this show up in a management review?
In at least three separate places, which is what makes it a property of evaluation rather than a quirk of any one process.
The impact rating. One number standing in for several judgments, described above.
The cause. Root cause analysis tends to terminate at the point where the nonconformity appeared. But the nonconformity was produced by a chain of processes, and the cause can sit several steps upstream of where the failure surfaced. Did the training process contribute to this? Did the way the work was scheduled? Questions like those take seconds to dismiss when the answer is no. They do not get asked at all when nobody has written them down.
Worth being precise about where the analysis should stop. Five whys does not stop where the evidence gets murky. It stops at rock — the thing actually standing in the way, the one you can push on. Stopping short of rock and stopping at rock produce records that look the same.
The action. An action gets assigned against a cause. Implementing a change and sustaining it require more than the change itself: the people who now have to work differently, and whatever keeps the change in place once attention moves elsewhere. An action can be complete on the page and narrow in scope at the same time.
What does it cost when the aperture closes?
Here is the case that made this concrete for me, from the audit side.
Back-tracing a realized safety event that reached a consumer, the trail ran to a change control record. The change impact assessment on that record had no predefined criteria — the field was empty, waiting for someone to write in what they believed the change would affect. What got written included no safety hazards and no safety evaluation, because the change was assumed not to introduce one.
The record was complete. The assessment was signed. No step in the process was skipped. The scope of the question was one person’s sense of what mattered, formed at the moment they filled in the field, and nothing in the record indicates that any narrowing took place.
That is the whole mechanism sitting in one artifact. Not a documentation failure, and not negligence. A question with a boundary nobody could see.
Is this a paperwork problem or a thinking problem?
A thinking problem, and the distinction matters because the obvious fix is the wrong one.
The obvious fix is more columns — four impact fields instead of one. That reliably produces the same number entered four times at the end of a Friday. Four columns carrying one repeated value is narrower thinking with more evidence of it.
The move that works is a single consolidated field with the areas named in the header: product quality, customer, health and safety, delivery. The operator considers the full set and records one value. The design lands in the thought pattern rather than in additional column data collected for show.
How far to take this is a gradient, and it should be. The rigor appropriate to a device manufacturer whose product carries patient risk is not the rigor appropriate to a staffing company. The scope of a criteria set follows criticality and sector, and getting that wrong in the direction of over-prescription carries its own cost.
The cost objection deserves a straight answer. Adding criteria is design work performed once, up front. The marginal cost per cycle is the time to consider areas that turn out not to apply, and dismissing an area that does not apply takes seconds. Where an area does apply, the time spent is the point. An argument that the wider question costs too much is usually quoting the cost of finding something as a reason not to look.
Why can’t the people running the system generate these questions themselves?
They can. What they cannot reliably do is generate them from memory, in the moment, while doing something else.
These questions are conditionally relevant. They bear on some cases and not others, they hit rarely, and an aperture question often turns up nothing because often nothing was there. A question with a low hit rate does not survive in a working manager’s recall, and it probably should not — retention tuned to hit rate is a reasonable way for a mind to run an operation. That is a limit on recall, not a comment on competence.
Which is why the criteria belong in the form rather than in a person’s head, and why they get chosen at design time, cold, by people who are not under pressure to remember everything at once. Filling a blank field is a recall task performed mid-task by one person. Choosing criteria is a design task performed once.
It should also not be one person doing the choosing. The functions that own each area should have input, because certain activities are relevant to certain areas under certain conditions and to a certain depth, and the people who own those areas know that better than whoever administers the form. Safety does not get left off a criteria list that safety helped write.
I should be direct about my position, since I sell this kind of work. What I hand a client is a question, and a question is checkable against their own record, by them, this week, without me. If a criterion never turns anything up across several cycles, it should be iterated out. That is the test I would want applied to my own recommendations, and it is the reason this does not reduce to a standing engagement.
How do you widen your own evaluation criteria?
Start with an inventory rather than a verdict. List the evaluation activities running in the organization — risk assessment, change impact, cause analysis, supplier assessment, action planning. For each one, find the defined criteria. Where the field is blank, that is the one to work on first.
Then ask the function owners what their area needs considered, and under what conditions: regulatory exposure, customer commitments, quality, health and safety, financial. This is process design work, not a documentation exercise, and the cost of it lands up front rather than every cycle.
One caution on reading the results. When a recommendation to widen something goes nowhere, there are three explanations and only two of them are failures: the question did not land, the question landed on someone without the standing to move it, or the business considered it and decided the change was not worth making. The third is the review working as intended. A post that treated every unmoved recommendation as decay would be wrong.
Telling those apart is easier than it looks, and it does not depend on a written rationale. Whether reasoning gets documented is a scoped question of its own, and a written record is one signal among many — sometimes close to neutral. But what an auditor is trained to evaluate is not confined to documents. An auditor can hear a real decision even where nothing was written: the person explains the reasoning, and other stakeholders describe it the same way. A decision that was actually made leaves a distributed trace.
The reader’s move here is not to go find an outsider. It is to stay open to evaluation criteria you did not generate, run each one against your own record instead of reading it for agreement, adopt the ones that prove relevant into standard management work, and drop the ones that stop earning their place. A criterion that only sometimes hits is still worth carrying. Turning something up every cycle was never the goal.
Frequently asked questions
Does adding evaluation criteria just create more paperwork?
It can, if the criteria arrive as separate fields to fill in. The version that works is a single consolidated field with the relevant areas named in the header, so the operator considers the set and records one value. The aim is a wider question, not a wider form.
Who should define the criteria?
The functions that own each area, contributing to a list assembled at design time. A criteria list drafted by one person carries that person’s blind spots into every evaluation that follows, which is the failure the list exists to prevent.
Do we need an outside consultant for this?
No. The work is an inventory of your evaluation activities and a conversation with your function owners. An outside look can supply candidate criteria faster, and independent internal audit work is a reasonable route to that, but the criteria themselves need to come from the people who own the areas.
What if a criterion never turns anything up?
Then it has stopped earning its place and it should come out. A low hit rate is not by itself a disqualifier, since many of the useful criteria apply rarely. But a list that never gets pruned turns into the same dead form it replaced, which is what structured audit and compliance programs tend to surface first.