How to Build a Competence Management System That Gets Used
Most competence systems get built in the wrong order. The rating scale is designed first, roles are fitted into it, and the risk question that should have governed every decision after it never gets asked directly. This is the architecture in the order it actually runs.
What is a competence architecture?
A competence architecture is a chain of dependent decisions, not a list of elements. Role risk determines the shape of the scale. The shape of the scale determines how evidence has to be designed. The evidence design determines what a gap is allowed to do. Each decision constrains the next one.
This matters because the failure is silent. A wrong call on scale shape does not produce an error in the matrix — the matrix looks complete. It produces work performed under a verification regime that was never calibrated for the person performing it, and the first visible symptom is defective output or a missed commitment.
The risk evaluation is not a separate exercise bolted onto the front. It is the same discipline organizations already apply through organizational risk assessment — criticality, and the impact a role has on commitments, requirements, objectives, and the measures leadership actually watches. Competence is a control, and controls are sized to risk.
Where do you start if nothing is documented yet?
There are two entry points, and the second one is faster in operations that already run well.
The forward build starts from assumed roles: define the role, identify its general and contextual competence requirements, and build the scale from the practical risk inherent in that role’s performance variability. It works when roles are stable and documented.
The backward pass starts from the people. Establish which external standards apply. Evaluate the risk carried by each position. Then sort the workforce that actually exists into groups with genuinely distinguishable capability — and the boundaries between those groups are the scale. You do not invent levels and fit people into them. You find the distinctions already present in the operation and name them.
Run backward and it validates a forward build: if the population will not sort into the levels you designed, the levels were theoretical. Run it first and it is the way in when nothing exists on paper. Either way the people who already hold this knowledge are the source — ask them what separates their strongest performer on a task from their weakest, because that difference is the scale boundary. If extracting it takes months rather than hours, the knowledge was not there. Extraction cost is a diagnostic.
How do you decide whether a role needs a graduated scale or a binary one?
The discriminator is performance variability in the work, not seniority. Bounded output means binary is the honest scale. Variable complexity and context means a graduated scale is required.
A machine technician operates the machine correctly or does not — there is no meaningful middle, so a binary rating is not a simplification, it is accurate. Data entry into a CRM behaves the same way. A final-touch operator at the end of that same line handles finish work across a range of complexities the role definition cannot enumerate in advance, and a graduated scale describes where in that range the person can currently operate. A software sales engineer sits on the same side for a different reason.
Two graduated scales can look identical on paper and measure different things. Where the risk sits in the decision, the scale is anchored to the oversight and guidance the person requires. Where it sits in the output, it is anchored to the complexity of work the person can take on unsupervised. The wrong anchor produces a scale that rates people accurately against the wrong variable.
How many competencies should an organization track?
As many as the risk justifies, and no more. Competence is rated per context, not once per role — the same person sits at different levels across different products, processes, and equipment, and one number per role is an average that hides exactly what the system was built to see.
The obvious objection is that this explodes into a cell count nobody maintains, which is the failure mode the whole exercise is meant to avoid. The control is risk-based prioritization against organizational direction. A genuinely complex, high-consequence operation may legitimately carry a large number of competence maturities — and if the risk is high enough, tracking and optimizing them is worth the cost. Where it is not, the competency should not be in the matrix at all.
This is the same logic that governs any proportionate control set, and organizations that already run structured operational risk management will recognize it. Breadth is not an administrative preference. It is an output of the risk evaluation.
What counts as evidence of competence?
A record is evidence when it contains objective evidence that defined success criteria were satisfied. That is the whole test, and it is why completion records fail it — not because records are the wrong artifact, but because a completion record carries no success criteria, so there is nothing in it to verify against.
Evidence design follows scale design directly. If a rating claims someone can work at a given complexity, the assessment has to have been performed at that complexity and verified by someone qualified to judge it — an evaluation or test process built at varying difficulty, or production of specific outputs inspected against defined criteria. What it cannot be is an assessment built at one difficulty used to substantiate a rating at another.
Criteria have to be defined in advance and in context — general expectations do not survive contact with an auditor or a planner trying to use the data. Organizations building this alongside training programs will find the competence and awareness requirements in ISO standards point the same way: what is evaluated is whether competence is defined, achieved, and verified.
What should happen when someone is rated below what the role requires?
The consequence mirrors the scale shape. Binary role, binary consequence: the work is blocked, either technologically in operations mature enough to enforce it, or through management evaluation and decision. The person does not enter that role until the gap closes.
Graduated role, graduated consequence: the work proceeds under increased oversight and verification of outputs — the same variable the oversight-anchored scale was measuring in the first place. The oversight level is the gap. Not a workaround for an inconvenient rating — the rating, expressed as an operating condition.
The development side gets designed at the same time, not improvised when a gap appears. If the criteria for each level are defined, the development actions common to reaching that level can be defined alongside them and paired to the criteria. Leadership owns delivering and sustaining that development — a management responsibility, not a records exercise, and where a competence system connects to any real continuous improvement framework.
Why do competence matrices stop being used?
Because the matrix gets built and nothing consumes it. That is the most common failure, and it happens at both ends.
Downstream, the data never reaches the decisions it should be constraining. Production schedules get committed and sales makes promises without anyone asking whether the capability exists, at the levels the work requires, in the quantity the schedule assumes. Tracked availability of competent individuals is a direct read on production capability and constraint, and it usually sits unused two clicks from the people making the commitment.
Upstream, the operating mechanics are never standardized — how assessment happens, on what frequency, on what triggers, who owns it. Ratings get assigned once and quietly decay. And without history there is no trend, because every reading is a snapshot with nothing to compare it against. You cannot tell whether capability is improving, holding, or eroding.
Both failures have the same root. The matrix is not the program. The program is oversight, review frequency, ownership, defined criteria, and process — the matrix is one artifact inside it. A high-touch leader with hands on the work produces the same information without any of it, and that holds exactly as long as that person is present and their attention is available. It does not transfer.
Frequently asked questions
Who validates the competence of your most experienced people?
Peer consensus, or external validation through a recognized certification program where one exists for that discipline. Where no external normalization is available, validation through reputation and organizational preference is legitimate — but declare it as the method rather than defaulting to it invisibly. The top of the scale still needs validating. It does not get to be the one level that is assumed.
Can a competence rating go down?
In practice it functions as a ratchet and downward movement is rare. What moves is performance — a different instrument with a different purpose. Keeping the two separate prevents a competence system from being read as a performance verdict.
How do you handle soft skills and the things a rating cannot capture?
A numeric scale rarely captures everything that makes a specific person the right choice for a specific job. Run a qualitative assessment alongside the numeric one, holding intuition and closeness against a consistent framework rather than in place of it. Where a soft attribute genuinely drives outcomes — if projects are stressful enough that being easy to work with is a real constraint — track it properly and feed it into planning, typically from performance evaluation. Sensitive attributes do not belong in a broadly shared document.
Is a simple role, requirement, evidence, gap matrix ever sufficient?
Yes. If the risk has been evaluated and a four-column structure with defined evidence covers it, that is not a shortcut — it is the control being proportionate to the risk, which is the objective. The failure mode is the reverse: treating the audit as the only constraint and producing documentation retrospectively, which costs more in administrative work and pre-audit cleanup than building it properly, and leaves nothing usable behind. Organizations running several frameworks see this when they consolidate through integrated management system consulting and find how much of that work was duplicated.
Where this usually starts
The useful first move is not building a matrix. It is naming the three roles where performance variability is highest, then asking what the current control is and whether anyone could produce evidence for it today. That is answerable in an afternoon, and it tells you whether the problem is documentation or design. In aerospace and defense, where flowdown and competence evidence are inspected directly, that answer arrives on its own. The same discipline applies through structured risk assessment wherever the consequence of variable work is high.