A QA dispute process is the documented route an agent uses to challenge a score they believe is wrong, and it becomes more important, not less, once scoring is automated. When a human graded you, the appeal route existed by default: you walked over and argued with the person who wrote the number. Software does not have a desk to walk over to, so unless you build the route deliberately it simply disappears. A working process needs five things: the agent can see the score, the exact criterion wording it was judged against, and the evidence behind it; a defined window to raise a dispute; a reviewer who is not the person or system being challenged; a turnaround time the agent can rely on; and a written outcome that references the rubric. Without the first of those, nothing else matters, because a score that does not link to the transcript line behind it cannot be disputed at all, only resented.
In short
- Automation deletes an appeal route that manual QA had for free. Nobody decides to remove it, which is exactly why it goes missing.
- A score that cannot be challenged is a verdict, and people do not improve against verdicts. They learn to work around them.
- Evidence is the precondition. If a score does not link to the specific transcript line that produced it, there is nothing to dispute and the process is theater.
- The reviewer of a dispute must never be the person or the system being disputed. Everything else in the design is negotiable, this is not.
- There are four outcomes, and the most valuable one is not the agent being right about the score. It is the agent being right about the criterion, because that fixes every future score.
- Dispute rate is a management metric. Near zero is usually not success, it means agents have given up or do not know the route exists.
- Read disputes by criterion rather than by agent. A rising rate on one criterion is a rubric defect, not a discipline problem.
- Disputed conversations are the best possible sample for a calibration session, because they are by definition the ambiguous ones.
Why the dispute path is not an optional nicety
Every manual QA program has an appeal process whether it is written down or not. An agent who thinks a score is unfair finds the reviewer, or their team lead, and argues. Sometimes the score changes, sometimes it does not, but the route exists and everybody knows it does.
Automate the scoring and that route quietly disappears. Not because anyone decided to remove it, which is the point. There is no reviewer to find, no inbox that owns the decision, and no obvious first step. The agent is left with a number, a feeling that it is wrong, and nowhere to take it. That is the moment trust in the whole program dies, and it usually happens in the first month of a rollout without anybody noticing.
An unchallengeable score is a verdict
The distinction matters more than it sounds. A score you can question is a claim: it says here is what happened, here is the standard, here is how the two line up. You can engage with a claim. You can accept it, argue with it, or learn from it.
A score you cannot question is a verdict. It tells you the outcome and nothing about the reasoning. People do not improve against verdicts, they adapt to them. They optimize for whatever they believe gets checked, they stop raising the ambiguous cases, and they treat the monthly number as weather rather than feedback. The program still produces scores. It just stops producing behavior change, which is the only reason it exists.
There is a second cost that lands on the QA team directly. Every dispute an agent cannot raise formally still gets raised, in one-to-ones, in team chat, and in the way agents talk about QA to new joiners. A dispute process does not create the disagreement. It gives you a version of it you can see, count and act on.
What has to be true before a dispute is even possible
Before you design any workflow, check that a dispute is technically possible in your setup. Three things have to be visible to the agent, and if any one of them is missing, the process you build on top will be theater.
- The agent can see the score, without asking. Not a monthly summary and not a number their team lead reads out. The individual score on the individual conversation, available when it happens. An agent who only learns their scores in a monthly review can only dispute things they can no longer remember.
- The agent can see the exact criterion wording. Not the category name, the actual sentence they were judged against. “Empathy: 2 out of 5” is not disputable, because there is nothing specific to disagree with. A written criterion is a promise about what will be measured, and a dispute is a claim that the promise was not kept.
- The agent can see the evidence. Which lines of the transcript produced the mark. This is the one people skip, and it is the one that decides everything.
No evidence, no dispute
If a score does not link back to the moment in the conversation that caused it, then a dispute reduces to the agent saying it feels wrong and the reviewer saying it looks right. There is nothing to examine, so the outcome depends on who is more senior. Agents work this out fast, and once they have, they stop bothering. That is where most near-zero dispute rates come from.
Evidence changes the conversation from a judgment about the agent into an observation about a transcript. Instead of “you scored low on ownership”, it is “this criterion asks whether you confirmed the next step before closing, and here is the closing message”. Now the agent can point at the line above it where they did confirm, or accept that they did not. Either way the disagreement is about something both people can read.
This is why Kaizo ties every score to the specific evidence in the transcript that produced it. It is not a transparency feature bolted on for comfort. It is the thing that makes a score arguable, and a score that cannot be argued with cannot be trusted either. If you are choosing a system, treat traceable evidence as a requirement rather than a preference, because validating automated scoring and disputing an individual score are the same capability at different scales.
Designing the process: who, when, and who decides
Once the preconditions hold, the process itself is short. Resist the urge to make it elaborate. Every extra step is a place where an agent decides it is not worth it.
Who can raise one, and against what
Any agent, on any individual score, on their own conversations. Do not restrict disputes to auto-fails or to scores below a threshold. A four out of five that should have been a five matters to the person it belongs to, and restricting the route sends the message that small unfairness is acceptable. Let team leads raise disputes on behalf of their team too, because some agents will never raise one for themselves.
The window
Set a deadline for raising a dispute and make it generous enough to survive a week of annual leave. Ten working days from the score being visible to the agent is a reasonable recommendation: long enough that a normal shift pattern or a holiday does not silently remove the right, short enough that the conversation is still fresh in the agent’s memory when it is examined. The reasoning matters more than the number. Pick a window that fits how your team actually works and publish it.
Who reviews it
This is the one rule with no flexibility: the reviewer must not be the person or the system being disputed. If a human scored the conversation, a different human reviews the dispute. If software scored it, a human reviews it, and ideally not the person who configured that criterion. A process where the original decision maker also judges the challenge is not a dispute process, it is a request for reconsideration, and agents can tell the difference immediately.
Turnaround, and what happens while it is open
Commit to a turnaround and publish it. Five working days is a reasonable recommendation because it is short enough that the dispute still feels connected to the conversation, and long enough that it can wait for a calibration discussion if the case is genuinely ambiguous. Disputes that expire quietly do more damage than a score that is upheld.
While a dispute is open, the score should be visibly marked as under review wherever it appears, and it should not be used in a formal performance conversation, a ranking, or anything tied to pay until it is resolved. Leave the original score in place rather than deleting it, and record the outcome against it. The audit trail is the point: an agent should be able to see that they raised something, that it was read, and what was decided.
The four outcomes, and what each one teaches you
Every dispute resolves into one of four outcomes. Teams tend to treat only the first as a real result, which wastes three quarters of the value. The second outcome is the most useful one you can get.
| Outcome | What happened | What it teaches | What to do next |
|---|---|---|---|
| Score changes | The agent was right about this conversation. The evidence did not support the mark | One score was wrong. Useful, but local | Correct the score, tell the agent plainly that they were right, and check whether the same misread affected similar conversations |
| Criterion changes | The agent was right about the standard. The criterion was ambiguous, outdated, or impossible to meet on this contact type | The most valuable outcome. Fixing it corrects every future score, not just this one | Rewrite the criterion, version it, announce the change, and say publicly that an agent prompted it |
| Score stands, reasoning was never explained | The mark was correct but nobody had shown the agent why until they challenged it | This is a communication failure in the program, not an agent problem | Uphold with a written explanation, then fix the gap that made the challenge necessary |
| Score stands, agent was wrong | The evidence supports the mark and the criterion is sound | The dispute surfaced a genuine skill or knowledge gap the agent cared enough to raise | Treat it as the coaching conversation it is. An agent arguing their case is already engaged with the standard |
What changes when a machine did the scoring
Most of the design above applies to any QA program. Two things are genuinely different when the score came from software, and both cut in the program’s favor if you handle them properly.
There is no intent to argue about
A large share of manual QA disputes are really disputes about the reviewer: they were having a bad week, they do not understand this queue, they have it in for me. Those arguments are unwinnable because nobody can prove a state of mind, and they poison the relationship between agents and the QA team.
Automated scoring removes that argument entirely, which narrows every dispute to two candidate causes: either the evidence does not support the mark, or the criterion is wrong. That is a much better conversation, and it is one that ends in something being fixed. Say this to agents openly when you launch the process. The point is not that the software is always right. It is that when it is wrong, you can find out exactly why.
One upheld dispute is never a one-off
This is the part teams get wrong. A human reviewer’s mistake affects one conversation. A systematic error in automated scoring affects every conversation with the same shape, all at once, and every agent who handled one.
So change the reflex. When a dispute is upheld on the evidence, do not just correct that score. Immediately check every other score on the same criterion with the same pattern, over whatever period it applies to, and correct them proactively. Then tell the affected agents, including the ones who never disputed anything. An agent who receives an unprompted correction learns more about whether the program is fair than any amount of communication about fairness. This is also the strongest argument for scoring every conversation against the same rubric rather than a hand-picked sample: consistent scoring means a systematic error is findable and fixable in one sweep, instead of hiding in whichever conversations happened to be reviewed.
It is worth knowing who built the grader and what else they sell. Most QA and CX platforms now sell their own AI support agents, which means the same vendor supplies some of the work and the mark on that work. Kaizo does not sell AI agents, so it has nothing of its own to defend in the conversations it scores, and a dispute about an automated score stays a question about the evidence and the criterion rather than a question about whose product is being protected.
Dispute rate is a management metric, not a nuisance
Once the route exists, the volume flowing through it becomes one of the most honest signals you have about the health of the program. Most teams read it exactly backwards.
Near zero is not success
A dispute rate at or near zero is almost never evidence that the scoring is perfect. In practice it means one of four things, and three of them are bad: agents do not know the route exists, the route is buried behind too many steps, agents have concluded that disputes never succeed, or agents believe raising one will be held against them. Only the fourth possibility is good news, and it is the least likely.
So treat a silent process as a broken one until proven otherwise. Ask a handful of agents directly whether they know how to dispute a score and whether they think it would work. Their answer is more reliable than the number.
Expect a spike when you open the channel
The first weeks after launching a dispute process usually produce a rush of them. This is backlogged frustration finding a legitimate outlet, not a sudden collapse in scoring quality. Volumes settle quickly once agents can see that disputes are read and that some of them succeed. If you panic and tighten the process during the spike, you have confirmed the suspicion agents already had.
Read it by criterion, not by agent
This is the discipline that turns dispute rate into something useful. If you slice disputes by agent, you get a list of people who complain, and you will start managing the complaining. Slice by criterion instead and you get a map of where your rubric is failing.
- Disputes concentrated on one criterion: the criterion is ambiguous, or it is being applied to contact types it was never written for. This is a rubric defect. Rewrite it. If several agents independently read the same sentence differently from you, they are not all wrong.
- Disputes concentrated on one queue or contact type: the standard probably does not fit that work. A criterion written for billing queries may be unsatisfiable on a complex technical escalation.
- Disputes concentrated in one team: more often a communication gap than a quality gap. The standard may not have been explained the same way there.
- Disputes concentrated on one agent: this is the only pattern that is genuinely about an individual, and even then it is usually a knowledge gap worth coaching rather than a discipline issue.
- A sudden rise across the board: something changed. A rubric revision, a policy update, a new contact type, or a scoring configuration change. Go and find it.
One number is worth watching alongside volume: the proportion of disputes that are upheld. If almost none succeed, agents will stop using the process and they will be right to. If almost all succeed, your scoring or your rubric needs work well beyond the dispute queue.
Feed disputes back into calibration
Here is the part that makes the whole process pay for itself. Disputed conversations are the single best sample you will ever have for a calibration session, because they are, by definition, the ambiguous ones.
Most calibration sessions struggle with sample selection. Someone picks a few conversations, and if they happen to be clear-cut, everyone agrees, the session finishes early and nothing is learned. Agreement on easy cases tells you nothing. The cases where reasonable people disagree are the ones that reveal where your standard is actually unclear, and a dispute queue is a self-populating list of exactly those cases, identified by the people closest to the work.
How to run it
Take the disputed conversations from the period, strip the outcome, and have every reviewer score them independently before the session. Then compare. Where reviewers disagree with each other, you have found a criterion that needs rewriting, not an agent who needs correcting. Where reviewers agree with each other but disagree with the original score, you have found a scoring problem worth investigating properly.
Publish what you decide. If reviewers argued about whether a particular response counted as a policy breach and landed on an answer, agents should be able to read that answer. It teaches the standard better than the rubric text ever will, and it demonstrates that disputes go somewhere. That loop, dispute to calibration to a rewritten criterion to a published decision, is what converts an appeal mechanism from a complaints desk into the main engine that improves your scoring criteria over time.
Common mistakes when designing a dispute process
- Building the workflow before the evidence exists. If scores do not link to the transcript, a dispute form just collects frustration. Fix traceability first.
- Routing disputes to the person who wrote the score. Or, in an automated program, to the person who configured the criterion being challenged. Independence is the whole point.
- No deadline on your side. Agents accept an upheld score far more readily than a dispute that was never answered.
- Correcting one score and stopping there. When automation is wrong it is usually wrong in a pattern. Sweep every similar score rather than fixing the one that got noticed.
- Making the agent prove the score wrong. The burden belongs with the score. If the evidence for a mark cannot be shown, the mark should not stand.
- Reading dispute volume as a discipline problem. A rising rate on one criterion is information about your rubric. Managing the people who raise disputes is how you get to a silent process that everyone distrusts.
- Burying the route. If disputing a score takes more than a click or two from the score itself, most agents will not start. The effort required should match the effort of leaving a comment, not filing a grievance.
- Keeping outcomes private. When a criterion changes because an agent challenged it, say so publicly. That single act does more for trust in the program than any communication campaign.
- Letting a disputed score count while it is open. Using a contested number in a ranking or a pay conversation before it is resolved tells agents that the process is decorative. See the wider set of auto QA mistakes for how this pattern repeats elsewhere.
Frequently asked questions
Should agents be able to challenge an AI QA score?
Yes, and more so than with a human reviewer, not less. When a person graded you, the appeal route existed by default because you could go and argue with them. Software has no desk to walk over to, so unless the route is built deliberately it disappears, and a score nobody can challenge stops being feedback and becomes a verdict. The practical requirement is that every automated score links to the evidence in the transcript that produced it, so the challenge can be about something specific rather than about whether the system is trustworthy in general.
What is a normal QA dispute rate?
There is no credible benchmark, and you should be wary of anyone who quotes one. Dispute rate depends almost entirely on things that vary between teams: how easy the route is to find, whether agents believe it works, how strict the rubric is, how many criteria are subjective, and whether scores affect pay. A team with a hidden process and a distrusted one will report a low number for opposite reasons, so the figure is not comparable across companies. Read your own rate as a trend instead. Watch which criteria disputes cluster on, watch the proportion that are upheld, and treat a rate near zero as a sign the route is unknown or not believed rather than as proof the scoring is perfect.
Who should review a disputed QA score?
Someone who was not involved in producing it. If a human scored the conversation, a different reviewer takes the dispute. If it was scored automatically, a human reviews it, and preferably not the person who wrote or configured the criterion being challenged. This is the one rule worth protecting above all the others, because a process where the original decision maker judges the challenge is not a dispute process and agents recognize that immediately.
How long should a QA dispute take to resolve?
Publish a turnaround and hold to it. Five working days is a reasonable recommendation: short enough that the decision still feels connected to the conversation, long enough that a genuinely ambiguous case can wait for a calibration discussion. Pair it with a window for raising disputes, around ten working days from the score becoming visible, so that a normal shift pattern or a week of leave does not quietly remove the right. Both numbers are recommendations rather than standards, so adjust them to how your team works and then be predictable about them, because an unanswered dispute damages trust more than an upheld one.
What happens to the score while a dispute is open?
Mark it as under review wherever it appears, and keep it out of anything with consequences until it is resolved. That means no ranking, no formal performance conversation and nothing tied to pay while the question is open. Leave the original score in place rather than deleting it and record the outcome against it, so there is an audit trail an agent can see. Using a contested number before it is resolved is the fastest way to teach agents that the process is decorative.
Will opening a dispute process mean agents dispute everything?
Expect a spike in the first few weeks and then a settling. The initial rush is backlogged frustration finding a legitimate outlet, not a sudden drop in scoring quality, and volumes fall once agents can see that disputes are read and that some of them succeed. Tightening the process during that spike is the one response that guarantees the problem persists, because it confirms exactly what agents suspected. If high volume does not settle, the signal is about your rubric rather than your agents: look at which criteria the disputes cluster on.
Related terms
Give every score something an agent can argue with
See how Kaizo ties every automated score to the exact evidence in the transcript, so an agent can check the reasoning line by line, raise a dispute against something specific, and get an answer that references the criterion. Kaizo does not sell AI agents, so it has nothing of its own to defend in the conversations it scores.