Agents get marked down for things they could not control whenever a scorecard criterion depends on something outside their hands: a system that failed, a customer who left, or a team that never replied. The fix is not a faster appeals process, it is scorecard design. Every control-dependent criterion should be excluded, marked not applicable for that conversation, or rerouted to the team that owns the failure, so the score keeps measuring the agent while the finding still gets acted on.
In short
- A criterion the agent cannot influence is not a quality measure, it is a tax on bad luck, and agents work out the difference within about a month.
- Three categories cover nearly all of it: system and tooling failures, customer behaviour, and upstream teams.
- Each has a different fix. Exclude the criterion, mark it not applicable for that conversation, or reroute the finding to the team that owns it.
- Not applicable has to exist on every criterion and it has to shrink the denominator. If it does not, marking it is arithmetically identical to a fail and reviewers stop using it.
- You can see these in the data before anyone complains. A criterion that fails far more often in one queue, one hour or after one release date is a system problem wearing an agent’s name.
- A finding that traces to a broken process should update the process, not the person’s score.
Why one uncontrollable criterion discredits the whole score
Four situations come up again and again when support agents describe a score they think was unfair, and none of them is really an argument about judgement.
- A system failure meant a step the agent completed was never logged, so the reviewer could not see it and scored it as missing.
- The customer disconnected before the closing sequence, and every closing criterion returned an automatic zero.
- An escalation went to a supervisor who never answered, and the agent lost the resolution and follow-up points.
- A caller was hard to hear, an email address came back mistyped, and a data-accuracy criterion recorded it as a privacy failure.
Take the arithmetic of the second one seriously, because it is the clearest. If your closing section holds four criteria and a disconnect zeroes three of them, then customer behaviour, not agent behaviour, decides which band that agent lands in. The agents who take the angriest callers get hung up on most, so the people handling your hardest work score worst.
Two things follow, in this order. First, the score stops carrying information. If a reviewer cannot tell you whether a 72 means poor work or a bad afternoon on the tooling, nobody downstream can use the number for a decision. Second, and more expensive, the score loses its authority with the people it is supposed to help. Once an agent believes the number is partly arbitrary, every piece of coaching attached to it arrives pre-discounted. You are not coaching, you are negotiating, and you have lost the conditions for running a QA programme agents trust. Deming put the general case more bluntly than any QA guide does: a bad system will beat a good person every time, and measuring the person harder does not change which of the two is losing.
This is not a failure of the QA team. Reviewers work inside the same form as everyone else, and most forms give them two options on a criterion that could not possibly have been met: fail it, or quietly pass it and hope nobody audits. Both are wrong, and there is no third choice to reach for. That is a design problem, fixable in the rubric rather than in the argument about the rubric.
How to find the control-dependent criteria on your scorecard
Take your current QA scorecard and walk it one line at a time. For each criterion, ask three questions in this order.
- Could the agent have passed this with the tools, permissions and information they had at that moment? Not with the tools they should have had. The ones actually in front of them.
- Does a pass require anybody else to act first? Another team, a customer, a batch job, an approval.
- Could a strong agent pass this every single time by choosing to? If the honest answer is no, the criterion is measuring circumstances as well as skill.
Sort every criterion into three buckets as you go. Fully controllable, where most of your soft-skill and process criteria should sit. Conditionally controllable, meaning the agent controls it only when a precondition holds. Not controllable, meaning a pass depends on something the agent has no influence over at all.
The second bucket is where the work is. For every conditionally controllable criterion, write the precondition down in one plain sentence: the customer stayed on the line, the refund tool was up, the second-line team replied inside its service level, the recording captured the full call. Those sentences become the not-applicable triggers you will need later, and writing them out is what turns a vague sense of unfairness into a rule a reviewer can apply the same way twice.
Run it with agents in the room
Do this pass with two reviewers and two senior agents together, not as a QA-only exercise. The agents will find the control-dependent criteria in a fraction of the time, because they have been losing points to them for months and can name the exact ticket. It also changes what the exercise signals: you are asking the floor to help fix the form rather than announcing a correction to it. Budget two hours for a fifteen-criterion scorecard and expect to find between two and five problem lines.
The three categories of uncontrollable failure
Almost everything you find will sort into one of three categories, and the category tells you what to do about it. The distinction that matters is who could have prevented the failure, because that is also who should receive the finding.
| Category | What it looks like | Criteria usually affected | Default treatment |
|---|---|---|---|
| System and tooling | An outage or a slow tool, a completed step that was never written to the record, a recording or transcript that cuts out, a field that did not save | Process adherence, documentation, verification steps, hold and handle time | Not applicable for that conversation, plus a defect raised against the system |
| Customer behaviour | The customer disconnects before the close, refuses verification, will not stay for the recap, talks over the agent throughout, arrives already furious | Closing sequence, identity verification, satisfaction-linked criteria, tone and interruption criteria | Exclude the criterion, or make it conditional on the customer having stayed and engaged |
| Upstream and adjacent teams | An escalation nobody answers, a second-line backlog, a policy that forbids the resolution the customer needs, a knowledge article that is wrong or missing, a macro with outdated wording | Resolution, first contact resolution, accuracy, promised follow-through | Score the agent only on what they did with what they had, and reroute the finding to the owning team |
Exclude, mark not applicable, or reroute
Three fixes, and picking the wrong one is what makes these corrections fail. The test is not how serious the failure was, it is how often the criterion is impossible and who owns the cause.
Exclude the criterion when it fails the control test most of the time. If a criterion depends on an unreliable tool, or on the customer behaving a particular way, it is not a standard, it is a coin toss with a points value attached. A useful rule of thumb: if you would never coach someone on it, do not score them on it either.
Mark it not applicable when the criterion is normally controllable but was genuinely impossible in that one conversation. This is the common case, and it needs two things: a named trigger from the preconditions you wrote down earlier, and evidence in the record that the trigger fired. A disconnect timestamp before the closing sequence is evidence. A reviewer’s impression is not.
Reroute the finding when the failure is real, matters to the customer, and belongs to somebody. The customer did not get their refund, which is a genuine quality failure worth knowing about, but the cause was a policy the agent is not allowed to override. The finding leaves QA as a process defect with an owner and a date. The agent’s score does not move.
There is a fourth option teams reach for that you should refuse: fail the criterion and add a comment saying it was not the agent’s fault. The comment does not change the number, and the number is what reaches the monthly review, the ranking and the performance conversation. Sympathy in a free-text box is not a control, and it leaves the agent with nothing to work with except disputing the score.
Give reviewers the authority, then audit it
Reviewers should be able to apply not applicable without asking permission. Teams that make it an escalation are effectively removing the option, because a reviewer with forty reviews to finish will not open a ticket to skip one line. Control it on the back end instead: report not-applicable usage by reviewer and by criterion, and put the outliers on the agenda of your next calibration session. That is a far better conversation than the one you get from forcing false fails.
Why not applicable has to be a first-class option on every criterion
Not applicable only works if the arithmetic works. The score has to be points earned divided by points applicable, not divided by the full form. If eleven of fifteen criteria applied and the agent passed eight, the score is eight out of eleven. If your form leaves all fifteen in the denominator, marking a line not applicable costs exactly what failing it costs, reviewers notice within a week, and the option goes unused. Plenty of teams already score out of the applicable questions only. Others have a not-applicable checkbox that does nothing, which is worse than not having one, because it looks like a fix.
Consider what happens on a form with no working not-applicable option. The strict reviewers fail the impossible criterion, the generous ones pass it, and both are guessing at a rule that was never written. You have manufactured reviewer disagreement that no amount of calibration will resolve, because it is structural rather than perceptual: two reviewers can agree completely about what happened and still produce different scores. Agents then experience their score as a function of who reviewed them, which is the fastest way to lose a QA programme’s credibility.
The data damage lasts longer. A criterion passed when it was never tested inflates your pass rate. A criterion failed when it was impossible manufactures a defect that never happened. Either way that criterion’s statistics become unusable, and those statistics are exactly what you need to weight the scorecard sensibly and to spot the system problems in the next section.
Four requirements, and they are all cheap:
- Available on every criterion, not on a chosen few. The one you did not enable is the one that will break.
- A reason code from a short list, five or six options, so the usage is countable. Free text is not countable.
- Visible to the agent on the review, so they can see the reviewer noticed. Half the resentment in this whole subject is agents assuming nobody did.
- Reported monthly by criterion and by reason.
One threshold worth adopting. If a criterion comes back not applicable in more than roughly a third of conversations, it does not belong on that scorecard. It belongs on a channel-specific or queue-specific one, because you are asking a single form to grade two different kinds of work.
How to spot the problem in your data before an agent tells you
Complaints are a lagging indicator, and they arrive from your most confident agents rather than your most affected ones. The data gets there first, and the method runs in a spreadsheet.
For each criterion, take its failure rate over the last quarter, then break that rate down by hour of day, queue or skill, channel, agent tenure band, and against the dates of releases and policy changes. Agent skill is spread fairly evenly across those buckets. System problems are not, and that asymmetry is the whole diagnostic. A criterion that fails 6% of the time overall but 35% of the time in one queue between two and five in the afternoon is not telling you about coaching.
Four patterns and what each one usually means:
- Concentrated in a time window.Understaffing, queue pressure, or a batch job that locks a system every afternoon. Check what else happens at that hour before you write a coaching note.
- Concentrated in one queue, skill or channel. A tool only that group uses, a routing rule, or a scorecard that does not fit that group’s work.
- A step change on a date. A release, a policy change, or a change to the scorecard itself. A criterion that was fine in March and fails constantly from April did not become harder because your agents got worse.
- Spread evenly across everyone, including your strongest agents. The criterion is ambiguous or impossible. If your top decile fails it about as often as your bottom decile, it is not measuring skill.
There is a sampling problem hiding in this method, and it is worth naming. At the 3% sample most programmes run, you cannot do it at all. Slice one criterion by queue and by hour and most of the cells are empty, so the pattern that would have exonerated the agent is invisible and the anecdote wins the argument instead. This is the practical case for scoring everything: 100% coverage revealing trends that 3% sampling never could. The point is not that more gets seen, it is that the selection bias disappears, so a concentrated failure can be recognised as concentrated.
Kaizo’s Auto QA scores every conversation against your own scorecard and keeps the evidence attached criterion by criterion, so you can filter one criterion’s failures by queue and by hour and see whether they cluster before anyone is coached on them. That traceability matters more than the coverage does: the useful question is not how many conversations were scored, it is whether you can show why a specific criterion failed on a specific conversation. More on the coverage side of this in what 100% QA coverage actually means.
Send the finding to the owner, not to the agent’s score
Everything above points at one principle. A finding that traces to a broken process should update the process, not the person’s score. QA reads more of what actually happens to customers than any other function, which makes it the best defect-detection system most support organisations own. Spending that signal on individual point deductions wastes it.
The mechanics are ordinary and that is the point:
- Tag the owner at review time. Product, workforce management, second line, knowledge, policy. Doing it during the review takes seconds. Reconstructing it in a monthly retro takes an afternoon and gets skipped.
- Give the finding the fields a defect gets. What failed, how often, which queues, what it costs in reopens or handle time. A process finding with a number attached gets fixed. A complaint does not. This is standard root cause analysis, applied to the scorecard itself rather than to the ticket.
- Suspend the criterion while the defect is open, and say so out loud on the floor. This is the step that buys back credibility, because it proves the scorecard responds to evidence.
- Report the fixes next to the scores. Criteria suspended, defects raised, defects closed. It changes what QA is understood to be for, and it gives your team a second set of numbers that is not the average score.
What is left after this work is the part the agent actually controls, which is the only part worth a one-to-one. Coaching stops opening with a dispute about whether the score was fair, and turning QA data into coaching gets easier because the findings are all things a person can change. Kaizo’s AI coaching builds its coaching cards from scored criteria, so the criteria you exclude or gate stop generating coaching noise as well as stopping the point deductions.
Two housekeeping rules to close on. Version the scorecard whenever you remove or gate a criterion, and never recalculate old scores under new rules, because a March score was produced under March’s form. And announce every change before it takes effect, with the reason. A scorecard that visibly corrects itself is treated very differently from one that quietly appears in a new version each quarter.
Frequently asked questions
Should agents be scored on things they cannot control?
No. A criterion an agent cannot influence measures circumstances rather than performance, and it makes the total score unusable for coaching or for ranking. The failure may still be worth recording, but it should be recorded against the system, the customer situation or the team that caused it. Sort each criterion into fully controllable, conditionally controllable and not controllable, then exclude, mark not applicable or reroute accordingly.
What should you do when an agent fails QA for a reason that was not their fault?
Do not correct it with a note in the comments, because the comment does not change the number that reaches their review. If the criterion was impossible in that specific conversation, mark it not applicable so it leaves the denominator. If the criterion is regularly impossible, remove it from the scorecard. If the failure was real but caused by another team, keep the finding, attach an owner to it and leave the agent’s score untouched.
How do you handle a call where the customer hung up before the closing script?
Trigger a not-applicable status on every closing criterion, evidenced by the disconnect timestamp, rather than letting them return automatic zeros. Score the parts of the conversation that happened. A disconnect zeroing three or four closing criteria can move an agent an entire performance band on customer behaviour alone, and it hits hardest on the agents handling the most difficult callers.
Should a QA scorecard have a not applicable option?
Yes, on every criterion, and it has to remove that criterion from the denominator so the score is points earned out of points applicable. A not-applicable option that leaves the total intact is arithmetically identical to a fail and reviewers stop using it. Require a reason code from a short list, show the status to the agent on the review, and report usage by reviewer and criterion so calibration can audit it.
How do you tell a system problem from an agent problem in QA data?
Break each criterion’s failure rate down by hour of day, queue, channel, tenure and release date. Skill is distributed fairly evenly across those buckets, so a criterion that fails far more often in one queue, in one time window, or from one release date onward is almost always a system or process problem. A criterion failed at similar rates by your strongest and weakest agents is either ambiguous or impossible.
How do you use root cause analysis in quality assurance?
Ask what would have had to be true for the agent to pass, then follow that chain until you reach something someone owns. If the chain ends at a tool that did not save a field, a knowledge article that was wrong or an escalation nobody answered, the finding belongs to that owner as a defect with a frequency and a cost, and the scorecard criterion should be suspended until it is fixed. A finding that traces to a broken process should update the process, not the person’s score.
Related terms
Find out which of your criteria your agents cannot actually pass
Bring the scorecard you use today and a quarter of scored conversations. We will show you which criteria fail in patterns that follow your queues and your release dates rather than your agents.