How Many Levels Should a Competence Rating Scale Have?

How many levels should a competence rating scale have?

As many as there are meaningfully different levels of performance that change a decision, and no more. The number is not a property of the competence. It is a property of the decisions the scale has to serve and the risk carried by getting those decisions wrong.

This is why one organization can correctly run two different scales inside the same program. A contract manufacturer supplying aerospace and medical device customers ran a binary skills and competence matrix at the program level, with written performance criteria defining what entry into the competent group required. Inside a handful of manufacturing departments, the same organization ran graduated matrices, because the work in those departments needed distinctions the binary could not hold.

Neither scale was a compromise. Each was matched to what the decisions in that area actually required.

When is a binary competent or not competent scale the right choice?

Binary is the right choice in two distinct situations, and they are not the same situation.

Validated collapse. An external body has already resolved the distinctions. Published acceptance criteria define what adequate work looks like. Qualification is granted through demonstrated capability against those criteria rather than through attendance, and the person granting it holds a qualification that someone else granted them, traceable back through the scheme, with an expiry date attached. When the term lapses, the authority to perform the qualified function lapses with it. The gradient exists — it just exists outside your organization, built by people with more data than you have. Collapsing to binary there is not a shortcut. It is declining to reinvent a distinction that has already been resolved more rigorously than you will resolve it.

Low criticality. The distinctions do not matter to any decision. Plenty of administrative and support processes are genuinely yes-or-no: the person can do the task or cannot, and planning needs nothing further. There is no hidden gradient here. There is simply nothing worth resolving.

What unites both cases is that the binary was chosen. The failure mode is binary arrived at by default — a checkbox that nobody examined, carrying no criteria and encoding no judgment. That failure looks identical on paper to the validated kind, which is why the distinction has to be tested rather than asserted. The same principle governs why training requirements are satisfied by demonstrated competence rather than by records of attendance.

How do you tell a designed binary scale from a default one?

Two tells, one objective and one conversational.

The artifact test. Ask for the criteria. A designed binary gate has written conditions that had to be satisfied for the answer to be yes — qualifiers, thresholds, what had to be demonstrated and observed by whom. If nobody can produce that definition, the gate was never a gate. It was a guess wearing a checkbox.

The compression test. Take everyone rated yes on a single competence and ask leadership directly whether there are material differences between those people in that competence. This is more useful than it sounds, because it is falsifiable from inside the population you already have. No external benchmark required.

The test needs a governor, though, or it proves too much. There are almost always some differences between people, because skill is continuous. If any difference counts, every binary fails and you have argued yourself into resolving everything, which is its own failure.

The governor is the decision. Not "are there differences," but "are there differences that would change how these people get scheduled, assigned, developed, or evaluated?" One operator may work ten percent faster and generate more rework, or need a second set of eyes. That is a real difference and it may net to nothing that changes a deployment decision. If two people differ in polish but you would put either on the same job tomorrow, the difference is real and irrelevant. If the difference means one can run the hard job unsupervised and the other cannot, the binary is costing you exactly the information planning needs.

When does a graduated scale earn its complexity?

When no external body has resolved the distinctions and the distinctions carry decision-relevant risk. Then you have to construct them yourself, tailored to the operating context of that organization, department, or function.

A graduated scale earns its administrative cost when it serves specific decisions:

  • Distinguishing resource capability so leadership can see what the operation can actually take on

  • Allocating utilization — putting the right capability against the right work

  • Supporting performance evaluation with something other than impression

  • Identifying the people who can move between functions as flexible capacity when volume fluctuates

That last one is the operational payoff that most often justifies the overhead in a manufacturing environment. Knowing who can float across functions is how a shop absorbs demand variability without either carrying excess headcount or missing delivery. A binary matrix will tell you who can run a given line. A graduated one tells you who can run it well enough that moving them there during a surge does not create a quality problem downstream.

One construction note that matters more than it appears. Sort the population into similarly-performing groups first, then write the criteria from what actually separates those groups. But set the minimum and maximum against best and worst case, not against the people currently in the building. A scale fitted to the current roster stops being a measurement instrument and becomes a description of the status quo, and it will need rebuilding the moment the roster changes.

Does meeting a prescribed scale mean you have the right scale?

No. A prescribed scale is a floor, not a sufficiency test.

When a standard or regulation prescribes levels and criteria, what it encodes is a risk judgment: based on the performance observed and the evidence collected across an industry, these are the distinctions that control the failures that matter. That is all a standard is. Someone identified what can go wrong and published what has to be true to prevent it.

The catch is that this judgment is generalized deliberately. A standard reads the same for the five-person shop and the five-hundred-person shop, because it has to be portable across everyone who will ever apply it. The universality is the feature and the limitation in the same breath. Your operating complexity, volume, and consequence profile are not the abstract case the standard was calibrated against.

So there are two obligations, and organizations routinely collapse them into one. Observe the prescribed scale, which is not negotiable. Then contextualize on top of it, because the prescribed scale addresses generalized risk and yours is specific. Compliance with a prescribed scale is evidence that someone resolved the general case. It is not evidence that you have resolved yours. Regulated environments make this especially visible, where personnel must be qualified for their assigned responsibilities rather than merely trained, and the qualification has to hold up against the specific work being performed.

What happens when scale complexity is mismatched?

The scale stops producing decision-useful information. That is the entire cost, and it shows up in two symmetrical ways.

Binary where a gradient was needed produces a flattened view. Real capability differences become invisible because the instrument cannot hold them. Planning proceeds on the assumption that everyone in the competent group is interchangeable, and the operation discovers otherwise at the worst possible moment — usually during a surge, a shift change, or an absence.

A gradient forced where binary would do manufactures distinctions. Levels get created that nothing depends on. The result is clutter and administrative drag: time spent arguing about whether someone is truly at level three or level four, evidence gathered to support a distinction that changes no decision, and a matrix that people maintain out of obligation rather than use.

Both are the same failure wearing opposite costumes. Under-resolution hides differences that matter. Over-resolution invents differences that do not. In each case the scale has stopped doing the only job it has.

Why do well-designed competence scales still fail?

Because the design was never the load-bearing part. A competence scale is a risk mitigation instrument, and it is only as sound as the risk judgment it encodes and the organizational sanction behind that judgment.

The common diagnosis is that these programs die from rating quality — managers cluster everyone in the middle, self-assessments inflate, ratings vary more by who is doing the rating than by who is being rated. That is real, but it is largely self-limiting under a risk-matched design. Resolution only exists where risk justified it, which means there are far fewer places for rating drift to hide. A shop that over-resolves everything hands it a hundred hiding places. And where evaluation quality does become the constraint, the honest read is that the ability to discern performance is itself a competence in the system — one that has to be defined and verified like any other, or the whole matrix inherits an unrated dependency at its root. That is the same logic that makes the gap between a certificate of attendance and demonstrated capability worth taking seriously in evaluation and audit roles as well.

The more common failure is quieter. Risk assessment is not an event that happens in a room with a template. It happens at every altitude, including informally. A team leader who thinks "I need to manage my team’s competence" has performed a risk assessment. It is real even though nothing was logged.

What goes wrong is that the judgment never travels. The leader identified a genuine risk, designed a genuine mitigation, and it stayed local. No consensus with other leaders, no shared view of scope, no sanction from above, and therefore no mandate to resource the evaluation process the scale depends on. The scale is not wrong. It is unsanctioned — one person’s private risk judgment wearing the institutional costume of a program.

That is a failure of risk culture rather than of scale design, and it is the reason level count is the last question rather than the first. Before asking how many levels, ask what risk the scale controls, who agreed that risk was worth controlling, and whether the effort to sustain it was ever actually priced. Organizations working through this usually find it is a process design problem before it is a documentation problem.

Frequently asked questions

Is a three-level or five-level competence scale better?

Neither is better in the abstract. Five levels is the most common default in published guidance, but the number should follow from how many distinctions change a decision in that specific area. If three levels capture every distinction that affects scheduling, assignment, and development, adding two more creates argument without adding information.

Should the same rating scale be used across the whole organization?

Usually not. Consistency of format is worth having so the matrix stays readable, but resolution should vary by area, because risk varies by area. The same skill can justify a binary gate in one department and a graduated scale in another when the consequence profile differs. Forcing uniform resolution optimizes for tidiness at the expense of usefulness.

How do you keep competence ratings honest?

Define the criteria in behavioral terms, require evidence with a date attached, and specify the conditions under which evaluation happens — including when, by whom, and with what preparation. Evaluation performed at the end of a difficult week produces different results than evaluation performed as scheduled work, and the process should account for that rather than hope. Where honest evaluation is not happening, treat the evaluator’s capability as a competence requiring verification, not as a character problem. This is ordinarily addressed as part of building operational discipline into how the work runs.

What is the first step in designing a competence rating scale?

Name the decisions the scale has to support and the risk carried by getting them wrong. Everything else — level count, criteria, evidence requirements, review cadence — follows from that. Starting with the scale format and working backward toward justification is how organizations end up with matrices that are maintained but never consulted. For a broader view of how competence fits into system design, see the Insights and Articles library, or how these requirements are handled inside a structured compliance program.

Previous
Previous

Why Performance Dips When a Manager Changes — and What the Dip Actually Reveals

Next
Next

Why Management Reviews Become Shell Meetings — and How to Map One That Is Not