Agents distrust QA when the process feels arbitrary: a tiny hand-picked sample of their conversations gets reviewed, the rubric is never published, two reviewers score the same ticket differently, and the resulting number affects their bonus with no way to challenge it. That is not agents being difficult, it is a reasonable response to a system that cannot show its working. You fix it by making the standard visible before it is applied, scoring a representative or complete set of conversations so nothing was picked, calibrating reviewers so the score does not depend on who opened the ticket, attaching the evidence to every score so an agent can check the reasoning line by line, and giving agents a genuine right of reply. Two of those matter more than the rest, and they matter equally: reviewing everything without traceable evidence is just a bigger audit, and traceable evidence on a hand-picked sample is still a hand-picked sample. Do that and QA stops being something done to agents and becomes something they use.
In short
- Most agent resentment of QA is a fairness problem, not an attitude problem, and it usually traces back to how conversations are selected for review.
- An agent’s real protection is a score they can trace, check and dispute, not a promise that everything was reviewed. Reviewing more without showing the reasoning is just a bigger audit.
- Cherry-picked sampling is the other half of the problem: when 2% of conversations are reviewed, agents rightly suspect the sample was chosen to prove a point.
- Publish the rubric and its weightings before you score against it, so nobody is graded on a standard they have never seen.
- Calibrate reviewers regularly, because a score that changes depending on who reviewed is not a measurement, it is an opinion.
- Every score needs the evidence attached: the exact line in the transcript that produced it, so it can be checked or challenged.
- A real dispute path with a named owner and a deadline turns a grievance into a process, and it improves the rubric as a side effect.
- Give agents their own data and coach on patterns rather than incidents, and QA becomes the thing that gets them promoted rather than the thing that gets them in trouble.
Why agents distrust QA, and why they are usually right
Ask a support agent what they think of quality assurance and you will rarely get a neutral answer. The complaint is remarkably consistent across teams, industries and channels: QA feels like something that happens to you, decided by someone you cannot see, using a standard nobody explained, on conversations you did not choose.
It is tempting for a QA lead to read that as resistance to accountability. It almost never is. Agents are not objecting to being measured. They object to being measured badly, and then having that measurement attached to their pay, their ranking and their next conversation with their manager.
Pull the complaints apart and they reduce to five specific, fixable failures.
- Opaque scoring: the rubric is either unpublished, or published but so vague that agents cannot predict how any given conversation will land.
- Selection bias, real or perceived: a handful of conversations get reviewed each month, and agents believe the reviewer chose the ones that make a point.
- Inconsistent reviewers: the same conversation would score differently depending on who picked it up, and everyone knows it.
- Consequences without recourse: the score moves a bonus or a rating, but there is no credible way to challenge it.
- QA as punishment: the only time an agent hears from QA is when something went wrong, so QA becomes synonymous with being in trouble.
None of these are complaints about high standards. They are complaints about an unaccountable process. Fix the process and the standard can go up, not down.
There is a cost to leaving this unaddressed, and it is not just morale. A QA program agents do not believe in stops producing behavior change, which is the only reason it exists. Agents learn to satisfy the criteria they think get checked rather than the ones that matter, feedback sessions become defensive, and the scores drift toward a number everyone quietly agrees to ignore. Meanwhile the QA team spends most of its week grading tickets and defending the grades rather than improving anything. The experience on the agent side and the return on the program fail together, for the same reason.
Signs your agents do not trust QA
Distrust is rarely stated outright, because saying “I think QA is unfair” sounds like an excuse. It shows up sideways instead. If you recognize the phrases in the left-hand column, you have a fairness problem rather than a performance problem, and the fix is structural.
| What agents say | What it actually signals | The fix |
|---|---|---|
| “You only ever pick my worst tickets” | Sampling is small enough that selection bias is plausible, so the score is not believed | Score every conversation, or make the sample random and provably so |
| “It depends who reviews you” | Reviewers are not calibrated and score the same work differently | Run regular calibration sessions and publish the agreed answers |
| “I did not know that was a fail” | The rubric is unpublished, vague, or changed without notice | Publish the rubric and weightings up front, version any change |
| “There is no point arguing” | No dispute path exists, or disputes disappear without a decision | A named owner, a deadline, and a written outcome for every appeal |
| “QA only shows up when I mess up” | QA is being used as a disciplinary tool, not a coaching input | Surface strong conversations too, and route findings into coaching |
| Silence in feedback sessions | Agents have concluded the score is not negotiable, so engagement stops | Give agents their own dashboard and let them self-review first |
Fix the sampling problem first
Start here, then read the next three sections as one piece of work. Almost every fairness argument in QA collapses back to how conversations were selected, and no amount of rubric polishing will fix a sample agents do not believe in. But fixing selection on its own is not the finish line: it removes the suspicion that someone went looking, and it is what makes the evidence trail worth building.
The arithmetic is unforgiving. A typical manual program reviews something in the order of two to five conversations per agent per month. For an agent handling several hundred, that is a single-digit percentage of their work standing in for all of it. When the sample is that small, two things are true at once: the score is statistically fragile, and the agent has no way to rule out that the reviewer went looking for a problem. Even a scrupulously fair reviewer cannot prove they were fair, because the selection itself is invisible.
Random beats hand-picked, and complete beats random
The minimum viable fix is genuinely random selection, documented and explained, so that nobody chose anything. That removes the accusation, though it does not remove the fragility: a random handful of tickets is still a handful.
The complete fix is scoring 100% of conversations. When every conversation is evaluated against the same rubric, the argument disappears entirely, because nothing was picked. An agent’s score stops being a verdict on six tickets and becomes a description of their actual month, including the difficult customers, the quiet Tuesday and the queue at end of shift. Agents who were skeptical of a sample are frequently the strongest advocates of full coverage, because it is the first time their good work has been counted alongside their bad.
One caveat, and it decides how this lands with your team. Coverage without a visible standard and a traceable score is just a wider net, and agents will read it exactly that way. Scoring everything and explaining nothing is the version of this that feels like being watched. Coverage earns its place because it means every agent, not only the sampled few, can hold a score up against the transcript and argue with it. That is why the next three sections matter more than this one.
This is how Kaizo is built: auto-scoring every conversation against your own rubric, so there is no sample and no selection to argue about, with the evidence attached to each score so an agent can see what was scored and why. At UiPath, Kaizo automated 100% of QA, delivering 200% ROI and an 8% lift in quality score.
Publish the rubric before you score against it
Nobody should be graded on a standard they have not read. That sounds obvious, and yet in many teams the scorecard lives in a spreadsheet the QA team owns, is revised quietly, and reaches agents only as a number.
Publishing the rubric means three things, and all three matter.
- The criteria are visible and specific. “Showed empathy” is not a criterion, it is an invitation to disagree. “Acknowledged the customer’s stated frustration before moving to troubleshooting” is something two people can look at a transcript and agree on.
- The weightings are visible. Agents need to know what actually moves the score. If a compliance step is an automatic fail and tone is worth three points, say so, because agents will otherwise optimise for the wrong thing and feel ambushed when it does not pay off.
- Changes are announced and dated. A rubric that shifts silently is worse than a strict one. Version it, announce the change, and do not apply new criteria to conversations that happened before the change.
The fastest way to pressure-test a QA rubric is to hand it to three agents with the same conversation and ask them to score it. Anywhere they disagree with each other is a criterion that is too vague to be fair, and you have found it before it cost you any trust.
It is worth going one step further and involving agents in writing the criteria in the first place. This is not a morale exercise. Agents know which parts of a scorecard reward behavior that is bad for customers, which criteria are impossible to satisfy on certain contact types, and which requirements quietly conflict with each other. A rubric that has survived that scrutiny is both fairer and more accurate, and agents argue with it far less, because they can no longer describe it as something invented by people who do not take the calls.
Calibrate, so the score does not depend on the reviewer
“It depends who reviews you” is the most corrosive thing agents can say about a QA program, because if it is true then the score measures the reviewer as much as the agent. And it is usually at least partly true. Human reviewers drift over time, disagree with each other on borderline cases, and score differently on a busy Friday than a quiet Monday.
What a calibration session actually looks like
QA calibration is straightforward and does not need to be a big production. Everyone who scores conversations, including team leads who only do it occasionally, scores the same two or three conversations independently. You then compare, and spend the time on the disagreements rather than the agreements. Every disagreement resolves into one of two outcomes: a reviewer was applying the rubric wrongly, or the rubric itself is ambiguous and needs rewriting. Both are useful.
Two practices make calibration credible to agents rather than invisible to them. First, run it on a fixed cadence, monthly is usually enough, so it is a routine and not a reaction to a complaint. Second, publish the agreed answers. If reviewers argued about whether a particular response counted as a policy breach and landed on a decision, agents should be able to read that decision. It teaches the standard far better than the rubric text alone, and it demonstrates that the standard is being maintained rather than improvised.
Automated scoring helps here for a specific reason: it does not drift. It applies the same rubric to the first conversation of the month and the ten-thousandth. Humans still set and calibrate the standard, but they are no longer the source of the inconsistency agents complain about.
Attach evidence, and give agents a real right of reply
A score with no evidence behind it cannot be discussed, only accepted or resented. A score with the exact line of the transcript attached becomes a conversation about the work.
So make evidence non-negotiable. Every criterion that was marked down should point at the specific moment that caused it. This changes the feedback session completely: instead of “you scored 72 on empathy”, it is “here is where the customer said they had already called twice, and here is the next reply, which went straight to the reset steps”. The first is a judgement about the agent. The second is an observation about a conversation, and it is coachable.
Building a dispute path agents will actually use
A right of reply that exists on paper but never overturns anything is worse than none, because it proves the point agents were making. A credible appeal process needs four things:
- A named owner. Someone specific reviews disputes, and it is not the person whose score is being disputed.
- A deadline. Five working days, say. Disputes that expire quietly are the fastest way to lose the room.
- A written outcome. Upheld or overturned, with a reason that references the rubric.
- A feedback loop into the rubric. If three agents dispute the same criterion, the criterion is the problem. Fix it and say publicly that it was fixed because agents raised it.
Expect a spike in disputes when you first open the channel, and treat it as a good sign rather than a crisis. It is backlogged frustration finding a legitimate outlet. Volumes settle quickly once agents can see that appeals are read and that some of them succeed.
The same rules apply when the reviewer is software
If part of your scoring is automated, none of the above gets easier or softer. It gets more important, because an agent cannot walk over to a model and ask what it was thinking. Two commitments keep it fair. First, the score has to point at the line that produced it, so an agent can check the reasoning themselves instead of being told the system is accurate. Second, the grader should be open to being tested: run it across conversations your reviewers have already scored, show the team where it agrees with them and where it does not, and publish that. An agent who has seen the grader checked will accept a score from it. An agent who has only been told it is reliable will not, and that is a rational position.
It is also worth knowing who built the grader and what else they sell. Most QA and CX platforms now sell their own AI support agents, which means the same vendor supplies the work and the mark. Kaizo does not sell AI agents, so when human-handled and AI-handled conversations are scored against the same rubric, the grader has nothing of its own to protect. Every score links back to the evidence in the transcript, so it can be checked, coached on or challenged rather than taken on faith.
Make QA a coaching input, not a disciplinary one
The last piece is framing, and it is the one that decides whether the first five stick. If the only time QA appears in an agent’s week is when something went wrong, no amount of procedural fairness will make it feel like anything other than surveillance.
Coach on patterns, not incidents
A single low-scoring conversation is mostly noise. A pattern across a month is a skill gap, and skill gaps are coachable. With full coverage you can say something genuinely useful: this agent handles billing queries at or above target but loses points on complex technical escalations, consistently, across dozens of conversations. That is a training plan. A single bad ticket is just a bad day. Turning QA data into coaching is what converts the program from an audit into agent development.
Surface the good, not only the bad
If your QA system only ever flags failures, it is a failure-detection system, not a quality system. Full coverage means you can also surface the best conversations of the week, use them as calibration examples, and let agents see their own strong work recognized in the same system that flags the weak. This costs nothing and changes the emotional register of the whole program.
Be careful about tying scores to pay
This is the question every QA lead eventually faces. Scores can inform pay, but the higher the stakes, the higher the bar for the process: full or provably random coverage, a published rubric, calibrated reviewers, evidence on every score, a working appeal path. If you do not have all five, linking scores to compensation will damage trust faster than the incentive improves behavior. Many teams get better results using QA for development and progression, and reserving formal consequences for clear, evidence-backed compliance failures.
Give agents their own data
Finally, let agents see their own scores, their own trend, and how the rubric was applied, without waiting for a monthly meeting. Agents who can self-review before their one-to-one arrive with their own analysis instead of a defence. That is the point at which QA has stopped being something done to them. It is also, not coincidentally, when the quality program starts to move the numbers it was built to move.
Frequently asked questions
How do I get agent buy-in for a QA program?
Buy-in follows fairness, so fix the process before you try to sell it. Publish the rubric before you score against it, remove selection bias by reviewing everything or sampling randomly, calibrate your reviewers, attach evidence to every score and give agents a working appeal path. Then involve agents directly by asking them to help rewrite the criteria they find unfair, because a rubric agents helped shape is one they argue with far less.
Should QA scores affect agent pay?
They can, but only if the process can withstand the scrutiny that comes with it. That means full or provably random coverage, a published rubric, calibrated reviewers, evidence on every score and a real dispute path. Without all five, linking scores to compensation will cost you more in trust than it gains you in behavior, and many teams get better results using QA for coaching and progression while reserving formal consequences for clear compliance failures.
How many conversations should I review per agent per month?
Manual programs typically manage two to five per agent per month, which is too few to be statistically meaningful and small enough that agents reasonably suspect the sample was chosen. If you are reviewing manually, make the selection genuinely random and be transparent about how it works. The better answer is to score 100% of conversations automatically, which removes the sampling argument entirely because nothing was picked, provided each of those scores links back to the evidence so an agent can check it.
How do I make an automated QA score feel fair to agents?
Show your working and let the grader be tested. Every automated score should point at the exact line in the conversation that produced it, so an agent can check the reasoning rather than being told the system is accurate. Before you roll it out, run it across conversations your reviewers have already graded and share where it agreed with them and where it did not, then keep the same dispute path you would use for a human reviewer. It is also worth asking who built the grader and what else they sell, because a vendor that supplies both the AI handling your conversations and the score on those conversations is grading its own work.
Should agents be able to see their own QA scores?
Yes, and they should be able to see them as they happen rather than in a monthly summary. Hidden scores guarantee distrust, because an agent who cannot see the standard being applied has no way to act on it. Agents with access to their own data and the evidence behind each score tend to arrive at coaching sessions with their own analysis instead of a defence.
How should I handle a disputed QA score?
Give disputes a named owner who was not the original reviewer, a deadline of about five working days, and a written outcome that references the rubric whether the score is upheld or overturned. Track the criteria that get disputed most, because repeated disputes on the same criterion usually mean the criterion is ambiguous rather than that agents are wrong. Fixing it and saying publicly that agents prompted the fix does more for trust than any amount of communication about QA.
Why do agents feel QA is unfair even when reviewers are trying to be fair?
Because good intentions are invisible and selection is not. When only a few conversations per agent are reviewed, a scrupulously fair reviewer has no way of proving they did not go looking for problems, so the suspicion survives regardless of how carefully the review was done. Fairness has to be demonstrable, which is why coverage, published criteria, calibration and evidence matter more than reviewer intent.
Related terms
Build a QA program your agents will actually trust
See how Kaizo links every score to the exact evidence in the transcript, so agents can check it and dispute it, scores every conversation against your own rubric so nothing was picked, and turns the results into coaching rather than a monthly verdict.