An agent error is a failed QA criterion the agent could have passed by doing their job differently. A process error is a failed criterion the agent could not have passed, because a macro, a knowledge article, a policy, an upstream system or the queue made the correct behaviour unavailable. Telling them apart at review time is what turns a quality score from a verdict on a person into a diagnosis of an operation, and it is the difference between coaching someone harder and fixing the thing that made them fail.
In short
- A scorecard has one name on it, so every failed criterion is attributed to that person by default. Nobody decides this. The form just has nowhere else to put the answer.
- Seven causes cover almost everything: agent skill, agent judgement, macro or template, missing or wrong knowledge article, unfollowable policy, upstream system or handoff, and staffing.
- Only the first two are coachable. If your reviews have never been tagged as anything else, that is a measurement gap, not a clean operation.
- Tag the failed criterion, not the conversation, keep the list to seven, and require evidence for any tag that files work against another team.
- If 40% of your failed criteria trace to process rather than people, coaching harder will not move the number.
- Process causes spread thin across many agents, so a small sample can prove an agent problem and structurally cannot prove a macro problem.
Why a scorecard blames the person by default
Look at the shape of a QA review. There is one name at the top, a list of criteria lifted from your QA scorecard, and a column for pass or fail. There is nowhere to record why a criterion failed. So the answer defaults to the person whose name is on the form, not because anyone chose that, but because the form has no other field.
Three things follow, and they compound.
Coaching absorbs work that belongs to operations. A team lead spends a one to one explaining a behaviour to an agent who already knew the behaviour and could not perform it, because the saved reply they are required to use contains the wrong wording. The conversation is polite, the agent nods, and nothing changes, because nothing that caused the failure was in the room. Repeat that for a quarter and you have a coaching programme with no mechanism of action.
The score stops being diagnostic. Once environmental failures are folded into individual scores, the number measures how much broken process a person had to absorb that week. Two agents on 78 can mean completely different things, and the score cannot tell you which. Anything you build on top of that number, ranking, bonuses, promotion decisions, inherits the confusion.
QA turns adversarial. This is the part practitioners actually talk about. Agents describe being marked down for a system that timed out mid transaction, for a customer who hung up before the closing script could be delivered, for an escalation that no supervisor picked up. Every one of those is a genuinely failed criterion. Not one of them is a coaching problem. A programme that cannot say so out loud will be treated as a hunt for mistakes, and it will be right to be. Naming the cause honestly is the precondition for running a QA programme agents trust.
Note what this is not. Most support teams already run root cause analysis, and it answers a different question: why did the customer contact us at all. That is contact driver work, and it is worth doing. This is the other question, the one almost nobody runs: why did this criterion fail, and who owns the fix. Your QA rubric records what happened. It does not record what caused it, and until it does, the answer will keep defaulting to the agent.
A working taxonomy: seven causes of a failed criterion
You need a list short enough that a reviewer picks from it without thinking, and complete enough that nothing gets forced into the wrong box. Seven is about right. Anything more and reviewers pick inconsistently, which produces a dataset you cannot aggregate.
| Cause | What it looks like in the conversation | Who owns the fix | Does coaching help? |
|---|---|---|---|
| Agent skill | They knew what to do and did it poorly. Clumsy wording, wrong sequence, incomplete explanation | Team lead | Yes, directly |
| Agent judgement | They chose the wrong option among several defensible ones. Escalated too early, refunded when they should have troubleshot | Team lead with QA | Yes, using calibrated examples |
| Macro or template | The required wording came from a saved reply that is out of date, or is wrong for this case, and the agent used it as instructed | Whoever owns the macro library | No |
| Knowledge article missing or wrong | They searched, found nothing or found something stale, and improvised. The improvisation is what failed | Knowledge owner | No |
| Policy that cannot be followed | The rule contradicts another rule, or cannot be completed in the time the same scorecard demands | Policy owner | No |
| Upstream system or handoff | A tool timed out, the data shown was wrong, or another team returned the ticket unresolved | Ops or engineering | No |
| Staffing and workload | Queue depth forced the exact shortcut the criterion penalises | Workforce planning | No |
Why skill, knowledge and will is not enough
Plenty of teams, especially in outsourced operations, already split QA findings three ways: skill, knowledge, and will or motivation. It is a good instinct and it is better than nothing. It has one flaw that makes it unusable for the job here.
All three branches end inside the agent. A taxonomy where every path terminates at the person can only ever produce coaching. It cannot tell you the refund macro is broken, because it has no branch for the refund macro. Run a year of reviews through it and you will have a year of evidence that your people need more training, and no evidence at all about your operation, whatever the truth happens to be. Safety research settled this argument a long time ago. James Reason’s BMJ paper contrasting the person approach and the system approach to human error treats errors as consequences of upstream conditions rather than as causes in their own right, which is exactly the move a three branch model cannot make.
Keep skill and knowledge, fold will into judgement, and add the four environmental causes. The model immediately starts producing work for teams other than yours, which is the point and also the reason it will meet resistance.
Two of the four deserve naming because they are the ones most often mistaken for agent problems. A knowledge base that is technically complete but unfindable under time pressure produces failures that look exactly like ignorance. And understaffing produces failures that look exactly like carelessness, because an agent working a queue two hours deep will skip the verification step the scorecard rewards. In both cases the agent is the last link in a chain and the only one being measured.
One diagnostic before you go further. Pull last month’s reviews and count how many failures were attributed to anything other than the person. If the answer is zero, that is not a clean operation. It is a form with one field.
One failed criterion, five different causes
The taxonomy only becomes real when you run a single criterion through it. Take a common one: confirmed the resolution and set a clear next step before closing. It failed. Here is the same failure, five times.
- Agent skill. They said “anything else I can help with?” and closed. They know the standard, they have hit it before, they were rushing. Coach it.
- Macro. The closing template for this ticket type ends at the apology and contains no next step field. Every agent using it fails the same criterion. The macro owner rewrites it once and the failure disappears across the whole team.
- Knowledge gap. The next step depends on a refund timeline that is documented in three places with three different numbers. The agent gave a vague answer because a specific one carried more risk than a vague one. Fix the article.
- Policy. Agents are not permitted to commit to a date without supervisor approval, and the same scorecard penalises hold time. The two rules cannot both be satisfied. Someone senior has to pick one.
- Upstream system. The order status field showed stale data, so the agent could not state a next step they trusted. Route it to the team that owns the integration.
Five identical entries in your quality report. Five completely different pieces of work, four of which no amount of coaching will touch. This is also why criterion level tagging beats conversation level tagging: the same conversation might pass on tone, fail on next step because of the macro, and fail on verification because the agent skipped it. One tag would flatten all of that.
Notice something else. Four of the five fixes are permanent and apply to everyone. Coaching an individual is the only branch that has to be repeated for every new hire, which is exactly why programmes that only produce coaching feel like they are running to stand still.
How to tag root cause at review time without doubling the work
The objection is always the same, and it is fair: reviews already take too long. Six rules keep the cost to seconds rather than minutes.
- Tag the failed criterion, not the conversation. One conversation can fail three criteria for three different reasons. A single conversation level tag averages them into something meaningless.
- Only failures get tagged. A passed criterion is not a diagnostic event. A review with one failure costs one extra click, and most reviews have very few.
- Seven options, single choice, required. Free text feels flexible and will not aggregate. If reviewers can type, you will end up with forty spellings of “macro” and no report at the end of the quarter.
- Default to agent skill and require an affirmative change. This costs nothing when the agent genuinely was the cause, and it stops the tag becoming a sympathy vote handed out to people the reviewer likes. The burden of proof should sit on the claim that creates work for another team.
- Any tag pointing at another team needs evidence in the review. The macro name, the article the agent actually opened, the timestamp of the timeout. A process tag with no evidence is an opinion, and the team you send it to will treat it as one, correctly.
- Calibrate the tag, not just the score. Two reviewers who agree the score is 85 and disagree about why are producing a dataset you cannot use. Add three or four tagging cases to every calibration session and check agreement on the cause separately from agreement on the number.
The two questions that settle most cases
Reviewers stall on the hard ones. These two resolve most of them in under a minute, and both are answerable by someone who was not there.
Could a strong performer, following the documented process exactly, have passed this criterion on this conversation? If the honest answer is no, it is not an agent cause, whatever else it is. This one question does most of the work, and it deliberately sets the bar at a strong performer rather than a perfect one.
Has the same criterion failed for several unrelated agents on the same type of conversation this month? If yes, treat it as process until someone proves otherwise. Individual weakness does not cluster by ticket type. Broken process does, every time.
Expect the first month of tags to be poor. Some reviewers over tag process because it feels kinder than marking a colleague down. Others under tag it because filing work against another team is politically expensive. Both settle after one properly run calibration round, and neither is a reason to skip the exercise.
What the tag distribution tells you when you aggregate it
A tag on a single review is trivia. The distribution across a few thousand of them is the actual product, and there are four ways to read it.
By share. What proportion of all failed criteria are environmental rather than personal? This is the headline number and it should be on the same slide as your quality score. If 40% of your failed criteria trace to process rather than people, coaching harder will not move the number, and a coaching plan built on that data is a plan to help people work around a defect faster.
By criterion. When one criterion carries most of the process tags, the usual cause is the criterion itself. It is asking for something your operation does not support. That is a rubric fix, not an ops fix, and it is much cheaper.
By agent. An agent whose failures are overwhelmingly environmental is usually on the hardest queue, not the weakest. Ranking that person against a team handling simpler work is the fastest way to lose the floor’s belief in the programme, and it is the complaint practitioners raise most often.
By trend. A process tag that keeps appearing after the fix shipped means the fix did not land, or nobody told the agents it had. That is worth catching in weeks rather than quarters.
The report this makes possible
What you can now produce is a ranked list of process defects ordered by how many failed criteria each one caused, with a named owner against every line. That is a document a head of support can walk into a product or operations review with, and it changes what the QA function is for. It is also the only version of this argument that survives contact with a team outside support, because it converts a quality complaint into a countable defect with a cost attached. Feeding that back into your wider support insights is where the QA programme stops being an internal audit and starts being an input to how the operation is run.
Why a sample can prove an agent problem and not a process problem
Here is the reason most teams have never seen this distribution, and it is arithmetic rather than effort.
Agent causes concentrate. Process causes spread thin. A weak agent produces the same failure repeatedly across their own conversations, so a handful of reviews will surface it. A broken refund macro touches forty agents once each. At the three percent sample most QA sampling programmes settle on, you might catch one or two instances, in different weeks, reviewed by different people, tagged inconsistently if they are tagged at all. Nothing aggregates. The macro stays broken and forty agents each collect a coaching note about wording.
That asymmetry, and not a general preference for more data, is the honest argument for scoring beyond a sample. It is the difference between 100% coverage revealing trends that 3% sampling never could and a review programme that can only ever see individuals. Process defects are population level objects. You cannot find them one interview at a time.
Two things have to hold for the population to be worth having.
- The tag has to be consistent. Volume built on inconsistent tagging just gives you a confident wrong answer, which is worse than no answer. This is why the calibration step above is not optional.
- Every tagged failure has to be traceable. The owner of a macro will not rewrite it on the strength of a count. They will rewrite it when they can open five of the conversations and read the exchange that failed. “The refund macro caused 137 failed criteria, here are the transcripts” gets fixed. “Quality is down in refunds” does not.
Kaizo’s Auto QA scores every conversation against your own scorecard inside Zendesk and Salesforce Service Cloud and keeps the reasoning attached to each criterion, so a failure you tag as a macro problem carries the specific exchange that proves it rather than a reviewer’s recollection of one. The tagging judgement stays with your reviewers. What changes is that the tag sits on a full population instead of a thin slice of one.
Routing each cause, and closing the loop
A tag that does not create work is decoration. Every cause needs a destination, and the destinations are different.
- Agent skill and judgement. Straight to coaching, with the conversation attached rather than described from memory. This is the only branch where a one to one is the correct response, and turning QA findings into coaching is a discipline in itself.
- Macro, knowledge article, policy. A named owner and a date, filed in whatever backlog that team already works from. Not a QA spreadsheet that nobody outside QA opens. If the fix does not enter someone else’s queue, it does not exist.
- Upstream system and handoff. The engineering or operations intake, with the volume attached. Volume is what gets a ticket prioritised; a well argued anecdote is not.
- Staffing and workload. Workforce planning, and it is a forecasting conversation rather than a quality one. Sending it to a team lead guarantees nothing happens.
Then close the loop, which is the step almost every programme skips. After a fix ships, re-measure the same criterion on the next month’s conversations. If the failure rate does not move, the diagnosis was wrong, and knowing that is worth more than the fix was. A programme that measures its own corrections stops being an audit function and becomes something other teams bring problems to voluntarily.
One last thing worth saying plainly to the floor. Separating agent error from process error is not a way to lower the bar. Agent causes still get coached, and a strong programme finds more of them, not fewer, because reviewers stop hedging on genuine performance issues once they have somewhere honest to put everything else.
Frequently asked questions
What is the root cause if an agent did not pass their QA target?
It is one of seven things, and only two of them are the agent. Agent skill and agent judgement are coachable. The other five are environmental: a macro or template with the wrong content, a knowledge article that is missing or out of date, a policy that cannot be followed as written, an upstream system or handoff failure, and staffing levels that force the shortcut the scorecard penalises. Before writing a coaching plan, ask whether a strong performer following the documented process could have passed that criterion on that conversation. If not, the fix belongs to another team.
What are the root cause categories in a BPO QA review?
Most outsourced operations use skill, knowledge and will or motivation. Those three cover the agent well and they share one blind spot: every branch ends at the person, so the model can only ever produce coaching. Keep skill and knowledge, fold will into judgement, and add four environmental categories for macros and templates, knowledge gaps, unfollowable policy, and upstream system or handoff failures. Staffing is worth a seventh. The moment those exist, findings start routing to the teams that can actually fix them.
How do you use root cause analysis for quality assurance?
Attach a single required cause tag to each failed criterion at review time, not to the conversation as a whole. Keep the list to seven options, default it to agent skill so it has to be actively moved, and require supporting evidence for any tag that points at another team. Then read the distribution rather than the individual tags: by share, by criterion, by agent and by trend over time. The output is a ranked list of process defects with an owner against each, which is a different and more useful artefact than a list of underperforming agents.
Should an agent be scored down for a system error?
The criterion should still be recorded as failed, because the customer experience was genuinely worse and you want that visible. What should not happen is the failure counting against the agent’s score or feeding their coaching plan. Record it, tag the cause as an upstream system failure, route it to the team that owns the system, and exclude it from the personal number. Marking someone down for a timeout they could not prevent is the single fastest way to make a QA programme resented, and it also corrupts the score as a measure of anything.
How do you tag root cause without slowing reviews down?
Tag only failed criteria, which in most reviews is a small number, and make it a one click single choice from a fixed list rather than free text. That adds seconds per failure, not minutes. The real cost is not the click, it is the calibration: reviewers who agree on the score but disagree on the cause produce data you cannot aggregate. Add a few tagging cases to each calibration session and measure agreement on the cause separately from agreement on the number.
Related terms
Find out how much of your failure rate is process, not people
Bring your current scorecard and a month of reviews you have already completed. We will walk through where your failed criteria actually come from, and which of them a coaching conversation was never going to fix.