Most support agents do not like being scored by AI, and their reasons are consistent: deductions for phrases they did say, scores that fell the week the system arrived, and no route to dispute anything. A smaller group prefers it, and their reasons are just as specific: the same standard applied every time, no reviewer carrying a personal history, criteria written down in advance, and a score that moves when they change what they do. The split has very little to do with AI. It tracks whether a score can be traced to evidence, explained, and overturned by a person.
In short
- The backlash is real and it is the larger side. Agents report being marked down for phrases they did say, for saying them in the wrong order, and for calls the customer ended early.
- The loudest complaint is not accuracy. It is that no dispute route exists and nobody, including supervisors, can explain how the score was produced.
- Almost every complaint aimed at AI QA was first aimed at human QA. Nitpicking, anonymous graders and retaliation after an appeal predate automated scoring by decades.
- The minority who prefer it name four things: consistency, no personal grudge, criteria published in advance, and a supervisor who can overturn the result.
- The reported failure mode is keyword matching. A grader that checks whether a phrase appeared produces exactly the complaints agents are making.
- If you cannot show an agent the moment in the conversation that caused a deduction, you do not have a QA programme. You have a number.
What agents say when AI starts scoring their calls
Search Google for AI quality assurance in a call center and, at the time of writing, the first result is not a vendor. It is an agent rant thread. That is worth sitting with before reading any further, because the people being scored are currently outranking the people selling the scoring.
Read enough of those threads and the complaints stop being vague. They are specific, repeated across different employers, and mostly checkable. Paraphrasing the recurring ones:
- Marked down for things they demonstrably do. One agent is coached weekly for never asking for a phone number or email, and asks on every call.
- Marked down for not reading a script verbatim, having read it verbatim. Others report saying the right words in the wrong order and losing the point for it, because the grader checks sequence and the conversation did not cooperate.
- Punished for the customer’s audio. Background noise on the caller’s end means the required phrase is not captured, and the whole section scores zero.
- Punished for the customer’s behaviour. If the caller hangs up before the closing, several closing items auto-zero and an otherwise good call drops below expectation.
- Judged on unequal samples. One agent’s month is built on twenty calls, a colleague’s on thirty-four, decided by the system.
- Sentiment read as fault. An angry caller who was handled well still flags the conversation negative.
- No way to dispute. The most striking version: a team can contest its four human evaluations a month but has no process at all for the automated ones, and asking repeatedly produced no process.
- Nobody can explain the number. Agents describe being coached every week against a score their own supervisors cannot break down.
- The switch was not announced. In one account the employer moved from human to automated grading without telling anyone and changed some requirements at the same time.
Two things make this worse than a normal grumble. These are frequently people who were scoring well before and can tell you their old numbers. And the behaviour it produces is visible to customers: agents describe repeating a caller’s name four times so it registers, and restarting a pitch from the top after an interruption. They are talking to the grader, not to the person on the line.
Most of these complaints are older than AI
Before treating this as an argument about automation, it is worth being honest about the baseline. Practitioner forums have been full of the same grievances since long before a model graded anything.
Agents describe one reviewer who consistently pulls their weakest call of the month while every previous reviewer scored them twenty to thirty points higher. QA leads describe two reviewers scoring the same ticket sixty and eighty-five because showed empathy meant something different to each of them. Agents describe never meeting the people who grade them and never learning which conversations were pulled.
The appeal stories are worse than the scoring stories. Some programmes cap appeals at three per month and forbid them on compliance items entirely. One agent reports their review count spiking in the weeks after they challenged one. Another won a dispute and got the evaluation back with a lower score and eight newly discovered errors attached. A trainer, introducing a scoring platform to a new cohort, told them they could challenge a score but would probably lose because the team is rarely wrong.
So the grievance was never a machine graded me. It was someone graded me, I cannot see why, and I cannot do anything about it. Automated scoring did not invent that. In plenty of places it inherited it and then applied it to more conversations.
Good human QA shows up in the same threads, and it is instructive. One analyst describes the goal as consistency, fairness and transparency, plus a challenge route that is actually used. That is the same list the pro-AI agents give. Nobody in either camp wants to be measured less. They want to be measured the same way twice, which is what a quality monitoring form is supposed to guarantee and frequently does not.
The agents who prefer it, and the reasons they give
The minority is smaller, quieter and easy to miss under the volume of the backlash. Their accounts are worth reading closely because they are unusually concrete about why, which makes them useful rather than merely reassuring. It is also less surprising than it first looks: Logg, Minson and Moore’s work on algorithm appreciation found that people often prefer algorithmic judgement to human judgement, rather than instinctively rejecting it.
No personal history
The most repeated reason is the absence of personal judgement. The grader was not in the meeting where you disagreed with your team lead, does not remember that you appealed last month, and has no view about you. One agent whose team switched away from human reviewers summarised it as the AI being more forgiving, precisely because there is no personal judgement in it. Set that against the agent above whose reviewer reliably picks their worst call, and the appeal of a grader with no memory of you becomes obvious.
The criteria exist as a document
One agent describes having a written specification of exactly what the programme looks for and where in the call it needs to appear. That single artifact removes most of the argument, because there is nothing left to interpret differently. It is the same job a well-built QA rubric does, except that here it was handed to the people being scored rather than kept by the people scoring.
The score responds to what you change
The same agent’s lowest score came from missing a specific phrase inside the window it was expected in. She moved her closing earlier in the call and went straight back to full marks. That is the whole mechanism: a deduction she could locate, a change she could make, and a result she could see. Compare it with the agent who delivers a required pitch on roughly nine and a half calls out of ten and watches that criterion sit at seventy percent.
A human is still the last word
In the accounts that go well, the automated score is not final. The system escalates anything it considers seriously wrong to a supervisor, who checks it. There is no dedicated review team pulling calls at random, so the thing agents used to dread, the hand-picked sample, stopped existing. The same agent notes the second-order effect: supervisors spend their time coaching instead of everyone bracing for the monthly verdict. Whether that holds depends entirely on whether the coaching loop was built at the same time.
A running number instead of a monthly event
Several accounts mention simply seeing their score across time frames rather than a monthly judgement on a handful of calls.
Two honest caveats, because this is preference and not endorsement. One agent notes scores went up and expects management to raise the targets in response. Another says the grader only registers what you say, so the team learned which words to use. Neither is an argument that automated scoring is good. They are arguments that it is legible, which is a lower bar and, on this evidence, the bar that matters.
What actually separates the two experiences
Line the two sets of accounts up and the dividing line is not the technology. Both groups are describing automated scoring, in different implementations.
| What the agent is reacting to | The version agents reject | The version agents accept |
|---|---|---|
| Criteria | Discovered by losing points, and nobody can explain how the total was produced | Published before the first score, including where in the conversation each item is expected |
| Evidence | A number on its own, or a sentiment reading management cannot break down further | Every deduction cites the moment in the conversation that caused it |
| Appeal | No process, or a process that exists on paper and has never changed a score | A named route with an owner and a deadline, and results that visibly do get overturned |
| Consistency | Depends which reviewer you drew, or which words the transcript happened to catch | The same rule applied the same way every time, with drift checked on purpose |
| Selection | Different agents judged on different numbers of calls, chosen by someone or by luck | Selection stops being a step, so nobody is judged on a hand-picked few |
| Responsiveness | The behaviour changes and the score does not | The score moves when the behaviour moves, and the agent can tell which change did it |
| Authority | The automated score is final and feeds ranking, pay or discipline directly | A human can overturn it, and sometimes does |
The failure mode is keyword matching, not AI
Look again at the complaints in the first section. Almost all share one shape: the agent did the thing and the grader failed to detect it. That is a detection failure, or in scoring terms a false positive, and false positives have a cause.
Buyers say it out loud in the same forums. A large share of what is sold as AI QA is rules automation with a model wrapped around the outside, and checking whether a string appeared is still the core of it. Practitioners who have run these projects describe it as if it contains X then fail in a better interface, and name the maintenance burden as the reason they abandoned it.
Every one of the agent complaints falls out of that design. Exact wording is required, so a paraphrase fails. Order is checked, so a natural conversation fails. An interruption breaks the sequence. Background noise breaks the capture. A hang-up removes the closing. None of these are the agent’s performance and all of them are the agent’s score.
Scoring a conversation as a narrative rather than a word list changes the shape of the error. The question becomes whether the agent acknowledged the problem, asked the relevant follow-ups, offered the right resolution and confirmed it, regardless of exact phrasing. That is what using a language model as a judge is for, and it is the difference between auto-scoring that reads a conversation and auto-scoring that searches it. It does not remove errors. It changes them from errors the agent can disprove to errors the agent can argue about, which is a real improvement and not a solved problem.
One consequence deserves its own rule. Any deduction that depends on a transcript depends on audio. If the caller’s line was poor, the score is a measurement of the line. Programmes that take this seriously suppress a criterion when transcript confidence is low rather than scoring it zero, because a zero is a claim and low confidence is an absence of one.
How to run automated scoring agents will accept
None of what follows requires software you do not already have, and none of it is specific to automation: it is the same fairness contract behind running a QA programme agents trust. Build it by hand first, because buying a tool on top of an unwritten standard reproduces the unwritten standard at scale.
- Publish the criteria before the first score. Each item, what counts as meeting it, and where in the conversation it is expected. Give it to agents, not just reviewers. On this evidence that document does more for acceptance than any accuracy claim you can make.
- Announce the change and name what is changing about the standard. The worst account in the whole corpus is a team whose employer switched graders silently and altered requirements in the same week.
- Require an evidence citation on every deduction. If a point comes off, the record shows the sentence or the timestamp. Deductions that cannot be shown do not get applied. This is the single rule that most reliably converts the first group of agents into the second.
- Score conversations you have already graded, both ways, before going live. Take a hundred reviewed conversations, run them through the new grader, and look only at the disagreements, which is the core of validating AI QA scoring. Most will be criteria that were always ambiguous and were previously resolved by whichever reviewer picked up the ticket. Fix the criterion, not the agent. This is ordinary QA calibration pointed at the grader instead of at each other.
- Keep a dispute route with a name and a deadline, and publish the upheld rate. The rate being visibly above zero is the point. A programme where nothing is ever overturned tells agents the route is decorative, and they stop using it.
- Do not let an automated score reach pay, ranking or discipline unread. A human signs off, and the human can override.
- Recalibrate on a schedule, and treat a drift in agreement as a defect in the rubric rather than a change in the team.
That standard is what Kaizo is built to. Conversations are scored against the scorecard you wrote rather than a fixed template, each deduction links back to the moment in the conversation that produced it, and a reviewer’s correction replaces the automated score rather than sitting next to it. If you are evaluating anything in this category, including us, those three properties are the ones to test on your own conversations before the demo data.
Five numbers that tell you whether agents accept it
Acceptance is measurable, and measurable early. Watch these from the first week rather than waiting for an engagement survey to tell you what the forums already would have.
- Dispute rate, and the share upheld. A dispute rate of zero is almost never agreement. It is usually no route, or no belief that the route works.
- Reviewer agreement. Have humans blind re-score a sample of automatically graded conversations, and track how often the two land in the same band. This is the number that belongs on the wall, because it answers what agents are actually asking, and it is the only honest way to establish how accurate AI QA is on your own work.
- Score distribution against the previous baseline. A cliff in the week the system arrived is a rubric problem, not a collapse in performance. Several of the angriest accounts are exactly this, diagnosed as a people problem.
- Movement after coaching. Take agents who changed a specific behaviour and check whether the relevant criterion moved. If it did not, the grader is not measuring the behaviour, and no amount of coaching will fix that.
- The language agents use. Listen for whether people describe the standard in the words of your criteria document or in folklore, as codewords and tricks. Folklore means the criteria never reached them.
Reviewer agreement and post-coaching movement are the two Kaizo reports on directly, and they are worth insisting on from anything you buy. A vendor that cannot show you agreement on your own conversations is asking you to take accuracy on trust, which is the same request agents are refusing to grant.
The pattern underneath all of this is unglamorous. Agents do not divide over whether a machine should grade them. They divide over whether the grade can be traced, explained and reversed. Build for that and the technology question mostly stops being interesting, which is the point at which automated scoring starts being useful to the people it is scoring.
Frequently asked questions
Do call center agents like AI QA scoring?
Most do not, and the negative accounts outnumber the positive ones by a wide margin. The recurring complaints are deductions for phrases the agent did say, automatic zeros caused by background noise or a caller hanging up, and no route to dispute the result. A smaller group prefers automated scoring to the human reviewers they had before, and they consistently cite the same four reasons: consistency, no personal grudge, criteria published in advance, and a supervisor who can overturn the score. Preference tracks the implementation, not the technology.
Why did my QA score drop when AI started scoring calls?
Usually because the standard changed at the same time as the grader, and nobody said so. A grader that checks for exact phrases in an expected order will penalise natural conversation, paraphrasing, interruptions and calls the customer ends early, none of which a human reviewer would have marked. A sharp, team-wide drop in the week a system goes live is evidence of a rubric problem rather than a performance problem, and the fix is recalibrating the criteria against conversations that were already graded by hand.
Can you dispute an AI QA score?
You can if your employer built a route, and many have not. The most common complaint in agent forums is teams that can contest their human evaluations but have no process at all for the automated ones. A defensible programme names an owner, sets a deadline, requires the original deduction to cite the moment in the conversation that caused it, and publishes how many disputes were upheld. If nothing is ever overturned, the route is decorative and agents will treat it that way.
Is AI QA scoring fairer than a human reviewer?
It is more consistent, which is not the same thing. A machine applies the same rule to every conversation and carries no history with the person it is grading, which removes reviewer-to-reviewer variance and the single-reviewer problem agents complain about most. It also applies a bad rule just as consistently. Fairness comes from the rubric, the evidence attached to each deduction and the appeal path. Consistency is what automation adds to whichever of those you already had.
How do you grade customer service calls consistently?
Write the criteria down before you grade anything, including where in the conversation each item is expected. Require every deduction to point at the sentence or timestamp that caused it. Have multiple reviewers grade the same conversations regularly and treat disagreement as a defect in the criterion rather than a difference of opinion. Give agents a real way to challenge a score. Those four habits produce consistency whether the grading is done by people, by software, or by both.
Should an AI QA score affect pay or discipline?
Not on its own. Agents in these threads describe bonuses and performance ratings driven by a number their supervisors cannot explain, which is the fastest way to lose a QA programme. Keep a human sign-off between the automated score and any consequence, make sure that person can see the evidence behind each deduction, and make sure they can override it. If the override never happens, the sign-off is not real.
Related terms
Show your agents the evidence behind every score
Bring a month of conversations your team has already reviewed by hand. We will score them against your own scorecard, show you where the two graders disagree, and show you the moment in the conversation behind every point that came off.