Shadow Scoring: Run Auto QA in Parallel Before You Trust It

Run automated QA scoring alongside your manual reviews, hidden from agents, and compare. How long to shadow, what to measure, and the exit test to go live.
Playbook · Automated QA verification

Shadow scoring means running automated QA scoring on your live conversations at the same time as your existing manual reviews, with the automated result hidden from agents and attached to no consequence, so you can compare the two before you switch anything on. You run it long enough for agreement to settle criterion by criterion, usually four to eight weeks, and you go live one tier of the scorecard at a time rather than flipping the whole thing at once. It is the field trial that follows a frozen-set validation test, and what it produces is a written exit test rather than an opinion.

In short

  • Shadow mode is a live parallel run on production traffic, not a lab test on a frozen set. Do the lab test first, then shadow.
  • Measure agreement one criterion at a time. Two systems can produce an 84 and an 86 on the same conversation and disagree on every line that made them.
  • Size the run per criterion, not per run. A rare compliance criterion can appear zero times in a month while the totals look healthy.
  • Healthy disagreement is concentrated in a few criteria, consistent in direction, and narrowing. Broken disagreement is scattered, two-directional, and moves week to week for no reason.
  • Procedural and factual criteria converge in weeks. Tone and empathy may never, often because your own reviewers do not agree on them either. Exit per criterion or you will shadow forever.
  • Expect cases where the automated score was right and the reviewer was wrong. Decide how you will handle that before it happens, because mishandling it is what ends shadow runs.
  • Shadow mode costs real reviewer time on top of the reviews they already do, and most teams that quit, quit in week three. Run the narrow version rather than none.

What shadow scoring is, and what it is not

Shadow scoring is running automated QA scoring over your live conversations while your existing manual review process carries on unchanged. Both produce a result for the same conversation. The agent sees only the human one. Nothing the automated system produces touches a scorecard, a coaching conversation, a bonus or a performance record for as long as the run lasts.

The term is borrowed from software deployment, where a new service is run against real production traffic in parallel with the old one and its output is logged rather than served. The logic transfers exactly: you collect evidence under real conditions without exposing anyone to the consequences of a system you have not yet tested. Google’s Rules of Machine Learning makes the same move in Rule 24, which tells you to run both models over the same sample and measure the difference between them before any user sees the new one.

Three things it is not

  • Not a pilot. A pilot replaces the existing process for part of the team. Shadow mode adds to it. That difference is the entire cost problem, and the last section is honest about it.
  • Not the validation protocol. A frozen reference set, scored blind by two reviewers and reconciled, is a lab test: run once, off to the side, before you commit. That is the right thing to do first, and the method is in how to validate AI QA scoring. Shadow scoring is what comes after it, on live traffic, over weeks, with your real mix and your reviewers on an ordinary week.
  • Not a vendor demo. If the conversations were selected by the party being measured, you have not learned anything about your operation.

The reason to do both is that they fail differently. A frozen set tells you whether the grader can read your rubric. A shadow run tells you whether it holds up on the Monday after a product launch, on the queue nobody has documented, in the fourteen-message thread that switched language halfway through. If you are still deciding whether to automate at all, that is a different question and it is answered in manual QA versus auto QA. Shadow mode assumes the decision is made and the open question is whether to trust the specific system in front of you.

How long to shadow, and how many conversations you need

Two numbers decide the run and only one of them is the number people ask about. The question is always how many conversations. The answer that matters is how many conversations per criterion.

Count per criterion, not per run

A criterion that applies to every conversation gets exercised hundreds of times in a month. A criterion that only applies to refund requests might appear ten times. A compliance criterion that fires once in five hundred conversations will appear never. Those three do not finish at the same moment, and a run sized on the total will quietly leave your most consequential criteria untested while the headline looks fine. Build the table before you start: every criterion, how often it applies in your real mix, and how many applicable conversations you will have accumulated by the date you planned to stop.

As a working target, we would not read much into a criterion’s agreement rate below about thirty applicable conversations, and below fifteen it is an anecdote. That is a recommendation with reasoning rather than a derived threshold: you need the criterion exercised often enough that one unusual week cannot move the number on its own. Rare criteria will not get there in a normal run. Either accept that they stay under human review and say so, or assemble a separate oversampled set of known cases and test those off to the side, which is the same technique the frozen-set protocol uses for auto-fails.

How long in calendar time

Four to eight weeks. Shorter than four and you cannot tell a real agreement rate from a good fortnight, and you miss the weekly shape of your own operation, because Monday conversations do not resemble Friday conversations and month-end does not resemble mid-month. Longer than eight and two things go wrong at once: reviewer willingness collapses, and the environment moves underneath you, so the thing you spent ten weeks measuring is no longer the thing you are about to switch on.

Set the end date before the first conversation is scored, and put it in writing. An open-ended shadow run does not end in a decision. It ends in attrition.

What to measure: agreement per criterion, never the overall score

The most common way to waste a shadow run is to compare totals. Two systems can return an 84 and an 86 on the same conversation and disagree on every line that produced them, because a harsh call on one criterion cancels a soft call on another. Aggregate agreement is the one number almost guaranteed to look reassuring, and the one number that cannot tell you what to do next.

Measure one criterion at a time, using the same arithmetic you would use to compare two human reviewers, including the correction for chance. On a criterion that ninety five percent of conversations pass, a grader that simply always says pass will agree with your reviewers almost every time and will have demonstrated nothing at all. The arithmetic is the same one you already use to check agreement between two people in a QA calibration session, and the reason to use the identical arithmetic is that it makes the machine’s number directly comparable to your team’s.

What to log, per conversation

What to record Why it earns its place
Human result per criterion, plus the evidence the reviewer pointed at Without evidence a disagreement is two opinions, and it cannot be diagnosed weeks later
Automated result per criterion, plus the lines of transcript that produced it Same reason, and it is how you separate a reasoning error from the system reading the wrong message
Whether the criterion applied at all Agreement is only meaningful over applicable conversations, and applicability mismatches are their own failure mode
Queue, channel, conversation length, agent tenure Agreement that is fine overall and poor on one queue is a bias you want to find before go-live, not after
Which reviewer scored it, and how far into their session Late-session human drift is real and shows up as machine error if you do not track it
Date The trend across weeks is the actual output of a shadow run. A single snapshot is what the lab test already gave you

What a healthy disagreement pattern looks like

The shape of the disagreement tells you more than its size. Two runs can report the same headline agreement and mean opposite things.

Signal Healthy Broken
Where disagreement sits Concentrated in two or three named criteria, usually the subjective ones Spread thinly across every criterion, including the objective ones
Direction Consistently one way on a given criterion: reliably harsher, or reliably softer Scattered both ways on the same criterion with no pattern
Movement across weeks Flat or narrowing, with any step change traceable to something you changed Moves week to week with nothing having changed
Explainability You can read the evidence from both sides and see why each answered as it did The evidence does not account for the result, or there is no evidence to read
Spread across the operation Roughly even across queues, channels, languages and agent tenure Materially worse on one queue, one channel, one language, or on new agents
Versus your own reviewers Machine-to-human disagreement is similar in size to human-to-human disagreement on that criterion The machine disagrees far more than your reviewers disagree with each other

Reading the pattern, including the failure that looks like success

A one-directional disagreement is good news badly presented. If the grader marks a criterion harshly and does it consistently, you have a threshold or a wording problem, and problems like that get fixed in an afternoon. Scatter in both directions on the same criterion is the harder finding, because there is no adjustment that fixes noise. Sort your criteria by business weight before you read any of this, because the answer you need is about the three criteria that carry consequences, not the twelve that do not.

Then watch for the pattern that looks like progress. If agreement climbs steadily through the run for no reason you can name, suspect contamination. The usual cause is that a reviewer has gained sight of the automated result somewhere and is anchoring to it, consciously or not, at which point the run is measuring how well the machine agrees with itself. Keep the automated output out of the reviewing interface for the whole run, check that it has not appeared in an export or a dashboard someone opened, and treat any improvement you cannot attribute to a specific change as a defect in the run rather than a result.

All of this rests on the automated results being traceable in the first place. A system that returns a criterion result with no pointer to the conversation cannot be shadowed, only believed. Kaizo’s Auto QA attaches the reasoning and the specific exchange to every criterion result, which is what turns a disagreement into something you can diagnose rather than something you argue about from memory.

The criteria that will never converge

Do not expect a single convergence date. Criteria settle at very different rates, and some will not settle at all.

Fast, usually two to three weeks

Procedural and factual criteria. Was identity verified, was the required disclosure given, was the correct policy quoted, was the ticket tagged, were the documented steps followed. These have a checkable answer sitting in the transcript, both sides are reading the same evidence, and the disagreements that survive are usually applicability mismatches rather than judgment gaps.

Slow, or never

Tone, empathy, ownership, professionalism, and every variant of whether the agent made the customer feel heard. These do not converge on a schedule, and the reason is usually not the model. It is that your own reviewers do not agree with each other on them either. Before concluding that the model acting as judge is bad at empathy, look at the human-to-human agreement rate for that criterion from your last calibration. If two experienced reviewers agree seventy percent of the time on tone, then a grader agreeing with them seventy percent of the time has matched your team, and holding it to a standard your own people do not meet is not a test, it is a way of never finishing.

This is why the exit has to be per criterion. Waiting for tone to converge before switching anything on means shadowing forever while the twelve criteria that were ready in week three go unused. Split the scorecard into tiers, take the fast ones live, and leave the slow ones with a human reviewing every flag. There is nothing embarrassing about that arrangement. For where automated scoring is strong and weak by nature, see how accurate AI QA is.

When the AI is right and the human was wrong

This happens more often than people expect, and most rollout advice leaves it out because it is uncomfortable. Somewhere around week three a reviewer will read the evidence attached to a disagreement and concede. Then it happens again, and a pattern appears: one criterion, or one reviewer, or the last hour of a long scoring session.

Handled badly, this ends the run. The reviewers doing the extra work to evaluate the system discover that the system is being used to evaluate them, and cooperation disappears within a week. Set the rules before the first disagreement, not after.

  • Never revise the human record after seeing the machine disagree. The reconciled human answer is the benchmark. Log the concession as a diagnosed disagreement and leave the record alone. A benchmark corrected in hindsight measures nothing.
  • Report at criterion level, never at reviewer level. Nobody should receive a per-reviewer accuracy ranking out of a shadow run, and no such view should reach anyone above them. The unit of analysis is the criterion.
  • Say plainly that the run is not a reviewer assessment, then behave that way when it is inconvenient. One screenshot of a reviewer league table circulating on the team is the end of the exercise.
  • Treat a cluster as a calibration signal. Repeated human misses on one criterion mean the criterion is ambiguous or the team has drifted, and the fix is calibration rather than a personnel conversation.
  • Fix the conditions, not the person. Concessions bunched at the end of long sessions are a workload finding. Shorter review blocks solve more of those than feedback does.

There is a real prize behind the discomfort. A shadow run that surfaces genuine human misses is the strongest internal case you will ever have for automating that criterion, and it is far more persuasive than any vendor accuracy figure, because it happened on your conversations in front of your own team. Just do not spend that credibility on blaming anyone. Attaching consequences before the evidence is in is a reliable way to lose the floor, covered in the auto-QA mistakes that kill agent trust.

The exit test: when to stop shadowing and go live

Write the exit test before you start. If you set the criteria after seeing the results, you will set them to match the results, and everyone in the room will know it.

Per criterion, all four have to hold

  1. Enough evidence. The criterion was exercised on enough applicable conversations to be readable, roughly thirty as a working floor.
  2. Agreement at your ceiling. Machine-to-human agreement on that criterion is at or near the human-to-human agreement rate for the same criterion. Your reviewers set the bar, not a number from a vendor deck.
  3. Stability. The last three weekly readings sit inside a narrow band, with no unexplained step changes.
  4. Diagnosed disagreement. Every remaining disagreement has a named cause, and the direction of the error is one you have decided you can live with. Favour catching more on compliance criteria, and favour fewer false positives on coaching criteria.

Then four conditions that are about the operation, not the criterion

  • Agents are told before the switch rather than after, and told which criteria are automated and which are not.
  • A dispute route exists and a named person owns it. See how to dispute an AI QA score.
  • The first live period carries no consequence. Results publish, and nothing attaches to a bonus, a ranking or a performance record for an agreed number of weeks.
  • One tier goes live at a time. Objective criteria first. Subjective criteria stay in assist mode with a human reviewing every flag.

Then write down the failure exit, which almost nobody does. If a criterion has not met the test by the end date it does not go live, and the run does not get extended a third time. If most of the scorecard has not met it, the answer for this system is no. A documented no after eight weeks is a far better outcome than a one-year contract defended by sunk cost.

Once the objective tier is live, the coverage argument finally earns its keep. The same rubric applies to every conversation instead of the handful a reviewer can reach, which is what gives you 100% coverage revealing trends that 3% sampling never could. The shadow run is what earns you the right to believe those trends.

The honest cost, and the minimum viable version

Shadow mode is not free and it is not cheap. For the length of the run your reviewers do the reviews they already do, and then a second body of work on top: recording evidence in a form that can be compared, sitting in a weekly session working through disagreements, and doing all of it while nothing visibly improves, because by design nothing changes until the run ends. Teams that abandon shadow mode mostly abandon it in week three, once the novelty has gone and the disagreement log has turned into an admin task with no audience.

Three things prevent that, and all three have to exist on day one. A named owner whose job includes the run rather than a volunteer who fits it around everything else. A short fixed weekly session, thirty minutes, on the calendar, cancelled by nobody. And a visible end date, so people can see they are paying into something that finishes.

The version to run when you cannot afford the full one

If a full parallel run across the whole team is not realistic, do not skip shadowing. Shrink it. A narrow run that finishes beats a comprehensive one that dies in week three, and it beats going live on faith by a wide margin.

  • One queue, or five agents. Pick the queue that carries the most consequence, not the easiest one to review.
  • Three to five criteria. Only the ones that will drive a decision or attach to a consequence. Ignore the rest of the scorecard for now.
  • Thirty conversations a week for four weeks. Sampled randomly inside that queue rather than chosen. That is roughly a hundred and twenty conversations, enough for a handful of criteria and nowhere near enough for twenty.
  • One hour a week. Reviewers score their normal sample as usual, and the owner spends the hour on the comparison and the disagreement log.
  • The same exit test, narrowed. It applies only to the criteria you tested. Everything untested stays under human review and gets named as untested in whatever you present.

Be precise about what the small version cannot tell you. It does not cover rare compliance criteria, it does not surface bias in queues you left out, and four weeks is short enough that a quiet month can flatter the result. Say all of that out loud when you present it. A narrow finding presented honestly is worth more than a broad one presented as more than it is, and it hands you something no vendor accuracy claim ever will: a number measured on your conversations, by your reviewers, that you can defend to anyone who asks.

Frequently asked questions

What is shadow scoring in QA?

Shadow scoring is running automated QA scoring on your live conversations at the same time as your existing manual reviews, with the automated result hidden from agents and attached to no consequence. Both systems score the same conversations, you compare them criterion by criterion over several weeks, and you only switch a criterion over once agreement is good enough and stable enough. The term is borrowed from software deployment, where a new service runs against real traffic in parallel and its output is logged rather than served.

How long should you run automated QA scoring in shadow mode?

Four to eight weeks for most teams. Under four weeks you cannot tell a real agreement rate from one good fortnight, and you miss the weekly and monthly shape of your own volume. Over eight weeks reviewer willingness collapses and the environment moves underneath you, so what you measured is no longer what you are switching on. Set the end date before you start, because an open-ended run does not end in a decision, it ends in attrition.

How many conversations does a shadow run need?

Size it per criterion rather than per run. A criterion that applies to every conversation is exercised hundreds of times in a month, while a compliance criterion that fires once in five hundred conversations may not appear at all, and a run sized on the total leaves your most consequential criteria untested. As a working floor, do not read much into a criterion below about thirty applicable conversations. Rare criteria either stay under human review or get tested separately on an oversampled set of known cases.

What should you measure during a shadow run?

Agreement per criterion, corrected for chance, never the overall score. Two systems can return an 84 and an 86 on the same conversation and disagree on every line that produced them, because a harsh call on one criterion cancels a soft call on another. Use the same arithmetic you would use to compare two human reviewers, so the machine’s number is directly comparable to your team’s, and log the evidence behind both results so disagreements can be diagnosed later rather than argued about.

What does a healthy disagreement pattern look like?

Disagreement concentrated in two or three named criteria rather than spread across all of them, consistent in direction on any given criterion, narrowing or flat across weeks, explainable from the evidence both sides pointed at, and roughly even across queues, channels and agent tenure. Broken looks like the opposite: scattered, two-directional, moving week to week with nothing changed, and worse on one queue or on new agents. Agreement that climbs for no nameable reason usually means a reviewer has seen the automated result and is anchoring to it.

What do you do when the automated score was right and the reviewer was wrong?

Expect it, and decide the rules before it happens. Never revise the human benchmark after seeing the machine disagree, log the concession as a diagnosed disagreement instead. Report at criterion level and never produce a per-reviewer accuracy ranking, because the moment reviewers realise the run is assessing them, cooperation ends. Treat a cluster of misses on one criterion as a calibration signal or an ambiguous criterion, and treat concessions bunched at the end of long sessions as a workload problem.

When is it safe to stop shadowing and go live?

Per criterion, when four things hold: the criterion has enough applicable conversations to be readable, machine-to-human agreement is at or near the human-to-human agreement rate for that same criterion, the last three weekly readings are stable, and every remaining disagreement has a named cause with an error direction you can accept. Then check the operational conditions: agents told in advance, a dispute route with a named owner, a first live period with no consequences attached, and one tier going live at a time.

Shadow Kaizo against your own reviewers

Run Kaizo in parallel with your existing manual reviews on your own queue, with the automated result hidden from agents, and compare criterion by criterion with every result traced to the exact exchange that produced it. You should not have to trust a scoring system before you have watched it work next to your team.

Book a demoExplore Agentic Auto QA

Choose your help desk

Not using either? We’ll let you know as soon as we can support your help desk solution.

Kaizo
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.