To validate AI QA scoring, you build a frozen reference set of real conversations, have two or more reviewers score it blind and reconcile their answers, run the automated scoring over the identical set, and then measure agreement criterion by criterion rather than as one blended number. For pass or fail and auto-fail criteria you measure precision and recall separately, because a system can look broadly accurate while missing most of what it was meant to catch. Every disagreement is then diagnosed by cause, and most turn out to be ambiguous criteria rather than model errors. The whole protocol fits inside a working week, and it is the only way to replace a vendor accuracy claim with a number that describes your own operation.
In short
- An accuracy claim with no published method is a marketing number. This protocol replaces it with a measurement taken on your conversations, your rubric and your policies.
- The reference set must be frozen and representative of your real mix, including the boring and the ambiguous conversations, or you are measuring the interesting cases someone chose.
- Humans score blind first. If your reviewers do not agree with each other on a criterion, you have no benchmark for that criterion yet, and that is a finding worth having.
- Measure agreement criterion by criterion. A blended score can look strong while the system is unusable on the three criteria your business actually cares about.
- For pass or fail and auto-fail criteria, precision and recall answer two different questions and they trade off against each other. Compliance and coaching pull in opposite directions.
- Diagnose disagreements by cause, not by blame. Most trace back to a criterion that a human and a machine can read two different ways, which is a rubric fix rather than a model fix.
- Re-test on a schedule. Accuracy is not a property you certify once; it drifts as products, policies, channels and models change.
What this protocol proves, and what it does not
Automated QA vendors advertise accuracy figures. Very few publish the method behind them: which conversations the number was measured on, which criteria it covers, who defined the correct answer, or whether the sample was chosen by the vendor. An accuracy figure without a method is a claim, not a measurement.
This protocol replaces the claim with your own number. It is deliberately vendor neutral: you can run it against any automated scoring system, including one you already pay for, and it works the same way whether the grader is a large language model, a rules engine, or a mix of both. What it produces is a per criterion agreement profile between the automated scores and a human standard you agreed on in advance, plus a diagnosis of every place the two disagreed.
What it does not produce is proof that the system is objectively right. There is no universal truth about whether a conversation was good, only your rubric consistently applied. For the background on what accuracy means and where automated scoring is strong and weak, read how accurate AI QA is first. This article assumes you want the measurement.
One structural note before you start. Most QA and CX platforms now sell their own AI agents, which means the system scoring your AI conversations was often built by the company that also built the thing being scored. That is not a reason to skip the test. It is a reason to run it, and to ask who has a commercial interest in the result. Kaizo is one of the few QA platforms that does not sell AI agents, and we would rather you ran this protocol against us than accepted a number on trust.
What you need in place before day one
The protocol takes about five working days of elapsed time and far less than five days of effort, mostly reviewer time in short blocks. Get four things ready first.
A rubric that is not moving
You cannot validate scoring against a rubric that is being edited mid-test. Freeze the current version of your scorecard for the week, including criterion wording, weights and auto-fail definitions. If you already know the rubric needs work, validate the version you run today so you have a baseline to improve against, then see how to write a rubric an AI can score for the rewrite.
Two or three reviewers, and a tie-breaker
We recommend at least two independent reviewers plus a third person who can resolve deadlocks. A single reviewer gives you an opinion rather than a standard, and no way to tell a machine error from a human one. The third voice keeps reconciliation from becoming a negotiation between two people who both want to be right.
Export access and a plain spreadsheet
You need full transcripts, including the internal notes reviewers would normally see, and somewhere to record scores per conversation and per criterion. A spreadsheet is sufficient. Do not run the human scoring inside the QA tool you are testing, because reviewers will see the automated score and anchor to it.
An agreement that nobody tunes anything mid-test
No prompt changes, no criterion rewrites, no model switches between the human pass and the machine pass. If something changes, the run is void and the machine pass starts again on the same frozen set. This is the rule most validation attempts break.
Day 1: build and freeze the reference set
The reference set is the foundation of the test. If it is not representative, every number you produce afterwards describes a population you do not have.
Sample the mix you actually run
Pull a stratified random sample rather than a convenience sample. Stratify on the dimensions that genuinely change how a conversation is scored: channel, queue or topic, length, outcome, and whether it was handled by a tenured agent, a new agent, or an AI agent. If forty percent of your volume is one repetitive billing question, roughly forty percent of your reference set should be that question. It will feel like a waste of reviewer time. It is the only way the result generalizes.
Include the boring and the ambiguous on purpose
Teams instinctively build reference sets out of interesting conversations, because interesting conversations are what QA meetings are about. That produces a set that is unrepresentatively hard, and it hides a grader that is quietly wrong on routine work, which is where most volume lives. Deliberately include short, uneventful, obviously fine conversations, and a handful of genuinely ambiguous ones where you expect reviewers to argue. The boring cases test consistency. The ambiguous cases test the rubric.
How big it should be
Our recommendation is sixty to one hundred conversations, and the reasoning matters more than the number. It needs to be large enough that each criterion is exercised many times, because a criterion that appears in four conversations tells you nothing, and small enough that two reviewers can each score every conversation inside a week. If many of your criteria are scenario-specific, size up rather than down, or accept that those criteria are untested this round and say so.
Handle rare auto-fail events separately
Serious compliance failures are usually rare. If a required disclosure is missed in one conversation in five hundred, a hundred conversation sample will contain roughly none of them, and the system will score perfectly by never flagging anything. Build a second, deliberately oversampled set of known failures for those criteria and analyze it separately.
Freeze it
Record the conversation identifiers in a fixed list, take an immutable copy of the transcripts, and never add, drop or swap a conversation once scoring begins. Every run, including re-tests months later, has to happen on identical data.
Days 2 and 3: score it blind, by humans, and reach agreement
Now you build the benchmark. Each reviewer scores the full frozen set independently, criterion by criterion, with no sight of the other reviewer’s answers and none of any automated score. Blindness is not a formality: once a reviewer has seen a machine score, their independent judgment is gone for that conversation.
Record scores at criterion level
Capture a value for every criterion on every conversation, a not applicable option, and a short note pointing at the evidence that justified the answer. That note is what makes reconciliation fast, and it is the same discipline you will demand of the automated system later.
Reconcile, and treat human disagreement as a result
Work through every criterion where the reviewers differed. The goal is a single agreed answer per criterion per conversation, decided by evidence rather than seniority, with the third reviewer breaking deadlocks. That agreed set is your benchmark.
Pay attention to which criteria produced the disagreements. If your reviewers disagree substantially on a criterion, you have no benchmark for it and no automated system can be validated against it. That is a genuine finding, often the most valuable output of the week, and it means the criterion is ambiguous rather than your people careless. Fix the wording, then either re-score that criterion or exclude it and mark it untested. If disagreement is widespread rather than isolated, stop and run a calibration session first.
Lock the benchmark
Write the reconciled answers to a file and do not touch them again. In particular, never revise a human answer after seeing the machine disagree with it. That habit invalidates more validation exercises than any technical mistake.
Day 3: run the automated scoring over the frozen set
The machine pass is the easy part. Run the automated scoring over exactly the same conversations, with the same rubric version, and export the results at criterion level.
Four things to insist on. First, you choose the conversations, not the vendor: if the sample is filtered by the party being measured, the result is not a measurement. Second, no configuration changes and no re-runs after seeing the comparison, because tuning and re-running fits the system to the reference set and your next number will be optimistically wrong. Third, export the evidence alongside the score. Every result should point at the specific lines in the transcript that produced it, because on day five you need to know whether a disagreement came from bad reasoning or from the system reading the wrong thing. A score with no traceable evidence cannot be diagnosed, only argued about. Fourth, if you are comparing vendors, run all of them over the identical frozen set. Two accuracy numbers measured on two different samples are not comparable, which is exactly why published vendor figures cannot be compared to each other.
Day 4: measure agreement criterion by criterion, never as one number
Now compare the machine output to the locked human benchmark. The temptation is to compute a single overall agreement percentage and treat it as the answer. Resist it. A blended number is the most misleading artifact in this whole process.
Imagine a rubric of ten criteria: seven simple and objective, such as whether identity was verified, and three that are the judgment calls your business genuinely cares about, such as whether the resolution was correct. A system that is near perfect on the seven and close to useless on the three still produces a strong looking overall number. You would adopt it, and it would fail you on exactly the things you bought it for.
What to compute for each criterion
- Agreement rate: the share of conversations where the machine result matched the reconciled human answer, counted only over conversations where the criterion applied.
- Direction of disagreement: when it differs, is the machine consistently harsher or softer than your reviewers? A bias in one direction is often a threshold or wording problem you can fix in an afternoon. Disagreement scattered both ways is a deeper problem.
- Size of disagreement: on scaled criteria, a one point gap and a full scale gap are not the same failure. Track the distribution, not just the count.
- Business weight: sort the results by how much each criterion matters to your operation, not alphabetically. Your decision depends on the top of that list.
Present the output as a table of criteria with those columns. It is far more useful to a buying or go-live decision than any headline percentage, and it tells your team where to work.
Day 4: for pass or fail criteria, measure precision and recall separately
Agreement rate is the wrong tool for pass or fail criteria and auto-fails, because those events are rare. If a serious failure occurs in two percent of conversations, a system that flags nothing agrees with your reviewers ninety eight percent of the time while catching nothing. That is why compliance criteria need two numbers instead of one.
Precision, in plain language
Of everything the system flagged, how much was real? Low precision means false alarms: reviewers spend their time overturning flags, agents lose trust because they keep being marked down for things they did not do, and the queue fills with noise. Precision protects your team’s time and the credibility of the program.
Recall, in plain language
Of everything that was genuinely a failure, how much did the system catch? Low recall means silent misses. Nothing looks wrong, nobody is complaining, and the failures are simply not surfacing. Recall protects you from risk you cannot see.
Why they trade off, and which one to favor
The two move against each other. Loosen the system so it flags aggressively and recall goes up while precision falls, because more of what it flags is innocent. Tighten it to clear cases only and precision goes up while recall falls, because borderline real failures slip through. No setting maximizes both, so decide which error you would rather make, criterion by criterion.
For regulatory and safety criteria, favor recall. A missed disclosure is not a risk worth trading for a tidier review queue, and the cost of a false alarm is a human confirming a flag. For coaching and tone criteria, favor precision. A wrong coaching flag does real damage to an agent’s trust, while a missed coaching moment costs little because there will be another next week. Write the choice down per criterion: it is a policy decision, not a technical one. Report both numbers on the oversampled rare event set from day one, and say clearly which set each came from.
Day 5: diagnose every disagreement by cause, not by blame
The counting is finished. The most valuable hour of the week is this one: going through the disagreements one at a time and asking what caused each. Teams usually skip it and jump to a verdict on the vendor, throwing away the finding that would have improved their program most.
Take every conversation and criterion where the machine and the reconciled human answer differed, read the evidence both sides pointed at, and assign a cause. In our experience the largest bucket is almost never model error. It is criteria a careful human and a careful machine can both read reasonably and still answer differently, which means the rubric is doing the failing.
| Cause of disagreement | What it looks like | How to fix it |
|---|---|---|
| Ambiguous criterion | Your two human reviewers also disagreed here, or the machine’s evidence is reasonable but answers a slightly different question than the one you meant | Rewrite the criterion so there is only one reading, define the edge cases explicitly, then re-score that criterion |
| Missing context | The judgment depended on something not in the transcript: a phone call, a policy exception, a known account history | Either feed the system that context, or mark the criterion as not machine scorable and keep it with humans |
| Genuine model error | The evidence is present and unambiguous in the transcript, the criterion is clear, and the system still got it wrong | Raise it with the vendor with the transcript attached, and hold that criterion below the automation threshold until it is fixed |
| Human error or drift | On review, your reviewers concede the machine was right, often on a long or repetitive conversation late in a scoring session | Note it, keep the benchmark as it was reconciled, and treat a cluster of these as a calibration signal for the team |
| Criterion does not apply | The criterion was scored by one side and skipped by the other because it was not relevant to that conversation type | Fix the applicability rules so both sides skip the same conversations, and exclude these from the agreement math |
Set your thresholds and decide what happens below them
A validation result is only useful if it changes what you do. Before you look at the final numbers, decide what each level of agreement will mean, so the thresholds are not reverse engineered from the result you got.
We deliberately will not tell you that a specific agreement percentage is good. Any published figure would be invented, because the right threshold depends on how consequential the criterion is, how much review capacity you have, and how much your own reviewers agree on that criterion. Your reviewer agreement rate from day three is the practical ceiling, and holding a machine to a tighter standard than your own team meets is not a fair test.
Three tiers, decided per criterion
- Automate: agreement is at or near your human ceiling and the failure mode is tolerable. The score stands on its own, with a dispute path so any individual result can be challenged and overturned.
- Assist: agreement is decent but not at the ceiling, or precision and recall are unbalanced. The system scores everything, humans review the flags and the borderline cases. Most subjective criteria belong here, and there is nothing embarrassing about that.
- Hold: agreement is poor, or the criterion failed the ambiguity diagnosis. Do not automate it, do not attach it to any consequence, and rewrite it before the next round.
Then write down the consequence rules: which criteria can auto-fail a conversation without a human confirming it, which can appear in an agent’s scorecard, and which can affect anything an agent cares about, after how long a trust building period. Attaching consequences to automated scores before running this test is one of the fastest ways to lose the floor, covered in more depth in the auto-QA mistakes that kill agent trust.
The payoff for getting the thresholds right is that automation can then run across everything rather than a sample. At UiPath, Kaizo automated 100% of QA, delivering 200% ROI and an 8% lift in quality score. Full coverage is also what makes this protocol possible, because it removes the selection bias from sampling. It is a precondition for a trustworthy measurement, not the achievement itself.
Re-test on a schedule, and how to read your result
Accuracy is not a certificate you earn once. Products and policies change, new channels appear, agents adopt new macros, and the underlying models are updated by their providers without asking you. Any of those can move scoring behavior on criteria that were solid six months ago.
A cadence we recommend, and why
Quarterly, plus an immediate re-test whenever any of four things change: the rubric, the scoring model or configuration, a policy or product a criterion depends on, or the addition of a new channel or language. Quarterly is a judgment call rather than a derived number: frequent enough to catch drift before a quarter of scoring decisions are built on it, infrequent enough that reviewers do not resent it. Keep the original frozen set as a regression set so results stay comparable, and refresh a portion each cycle so it does not stop resembling current traffic. Log every run: date, rubric version, set version, per criterion results.
Reading the result honestly
A good outcome does not look like one impressive number. It looks like this: high agreement on the objective, evidence checkable criteria that make up most of your rubric; middling agreement on the subjective ones, sitting in the assist tier with humans reviewing flags; one or two criteria pulled out because the week revealed they were ambiguously worded; and a documented reason for every threshold. That is a program you can defend to an auditor, to a skeptical support floor, and to yourself.
A bad outcome is a strong overall percentage with no per criterion breakdown, no precision and recall on the auto-fails, and no diagnosis of the disagreements. That is the same artifact the vendors publish, and it is what this protocol exists to replace. The measure of a scoring system is not the number it claims. It is whether it hands you the evidence to check that number yourself, on conversations you chose.
Frequently asked questions
How do you validate AI QA scoring?
Build a frozen reference set of real conversations, have two or more reviewers score it blind and reconcile into a single agreed answer, run the automated scoring over the identical set with no configuration changes in between, then measure agreement criterion by criterion. For pass or fail and auto-fail criteria, measure precision and recall separately. Finish by diagnosing every disagreement by cause, because most trace back to an ambiguous criterion rather than a model error.
How many conversations do you need in a reference set?
Sixty to one hundred is our recommendation, and the reasoning matters more than the number. It needs to be big enough that every criterion is exercised many times, since a criterion that appears in only a handful of conversations tells you nothing, and small enough that two reviewers can each score the whole set within a week without rushing. Size up if your rubric has many scenario-specific criteria, and build a separate oversampled set for rare compliance failures.
What is a good agreement rate between AI and human QA scores?
There is no universal figure, and any vendor quoting one without publishing its method is not giving you a usable number. The practical benchmark is your own reviewers: the rate at which two calibrated humans agree with each other on a criterion is the realistic ceiling for that criterion, and holding a machine to a tighter standard than your own team meets is not a fair test. Set thresholds per criterion against that ceiling, weighted by how consequential the criterion is.
What is the difference between precision and recall in QA auto-fails?
Precision asks: of everything the system flagged, how much was real. Recall asks: of everything that was genuinely a failure, how much did it catch. Low precision means false alarms that waste reviewer time and erode agent trust. Low recall means silent misses that look like everything is fine. They trade off against each other, so favor recall on regulatory and safety criteria and precision on coaching and tone criteria.
Why measure accuracy per criterion instead of overall?
Because a blended number hides exactly the failures that matter. A system that is near perfect on the many simple, objective criteria and close to useless on the few judgment calls your business cares about will still produce a strong looking overall figure. Breaking the result out by criterion, sorted by business weight, tells you which parts you can automate, which need a human reviewing the flags, and which criteria need rewriting.
How often should you re-validate AI QA scoring?
We recommend quarterly, plus an immediate re-test whenever the rubric, the scoring model or configuration, a dependent policy or product, or the channel and language mix changes. Accuracy drifts because the environment around it moves, including model updates you did not request. Keep the original frozen set as a regression set so runs stay comparable, refresh part of it each cycle so it keeps resembling current traffic, and log every result to build a trend line.
Related terms
Run this protocol against Kaizo
Bring your own frozen reference set, scored and reconciled by your reviewers, and we will run Kaizo over the identical conversations and hand back the comparison criterion by criterion, with every score traced to the exact lines in the transcript that produced it. Kaizo does not sell AI agents, so it has nothing to defend in the conversations it grades. You should not have to take an accuracy claim on faith, from us or from anyone else.