Why a 3% QA Sample Cannot Support Agent-Level Decisions

The math behind QA sampling is unforgiving: a few conversations per agent carries a margin of error so wide that most agent-level comparisons are noise. Here is the statistics, in plain terms, and why coverage is the only fix that survives it.
Guide · Automated QA

When a QA program samples a small number of conversations per agent, the resulting score carries a margin of error large enough that most conclusions drawn from it are unreliable. This is not an opinion about sampling; it is a property of small samples. A handful of scored conversations estimates an agent’s true quality with a confidence interval so wide that two agents who genuinely differ can look identical, and two who are identical can look different, purely from which conversations were drawn. Understanding the arithmetic is what turns the move to full coverage from a preference into a requirement for any decision made about an individual.

In short

  • A QA score from a small sample is an estimate with a margin of error, not a fact, and the margin is wide when the sample is small.
  • Margin of error shrinks with the square root of the sample size, so small samples are punishingly imprecise and adding a few conversations barely helps.
  • With only a handful of conversations per agent, the confidence interval around a score can span many points, wide enough to swamp real differences.
  • That makes ranking or comparing agents on small samples mostly an exercise in reading noise.
  • The fix is not a slightly bigger sample, which the square-root law makes expensive and slow; it is scoring every conversation, which removes sampling error for the agent-level question entirely.
  • The numbers used here are illustrative of the mechanism, not benchmarks; the point is the shape of the math, which holds for any small sample.

A sampled score is an estimate, not a measurement

The first thing to be clear about is what a sampled QA score actually is. When you score a few of an agent’s conversations and report an average, you are not measuring their quality. You are estimating it from a sample, exactly like a poll estimates an election from a subset of voters. And like a poll, the estimate comes with a margin of error that says how far the true value could plausibly be from the number you got.

QA reporting almost never states that margin, which is the root of the problem. A score of 84 gets treated as the fact that the agent is at 84, when what the sample actually supports is something like the agent is somewhere in a range around 84, and how wide that range is depends entirely on how many conversations you scored. Ignore the range and you will read differences that are not there. This is the statistical spine under the definitional point in what QA sampling is.

The square-root law, and why small samples hurt

The margin of error on an estimate shrinks with the square root of the sample size, not with the sample size itself. That single fact is why small QA samples are so weak. To halve your margin of error you have to quadruple the number of conversations you score, so the first few conversations buy you a very imprecise estimate and getting to a precise one by sampling is enormously expensive.

An illustrative sense of scale, offered to show the mechanism rather than as a benchmark: scoring a handful of conversations per agent leaves a margin of error wide enough to span many percentage points on either side of the reported score. So an agent reported at 84 and one reported at 78 may have identical true quality, their scores separated only by which conversations happened to be pulled. The gap you are about to coach on, or rank on, is inside the noise. Adding a few more conversations narrows that band only slightly, because of the square root, which is why you cannot sample your way out of the problem at a reasonable cost.

What this does to common QA decisions

The imprecision is not academic. It directly undermines the decisions QA scores are used for.

Decision What the small sample does to it Consequence
Ranking agents Rankings shuffle with each sampling draw You reward and penalize based on noise
Comparing two agents Differences smaller than the margin of error are unreadable Real gaps are missed, false gaps are acted on
Tracking an agent over time Month-to-month moves are mostly sampling variation You chase swings that are not real changes
Tying scores to pay or review The decision rests on an estimate too wide to defend A fairness and trust problem, not just a data one

Why coverage is the only fix that survives the math

There are, in principle, two ways to shrink the margin of error: score more conversations, or stop sampling. The square-root law makes the first one slow and expensive, and it never fully closes, because any sample retains some error. The second one closes it completely for the question that matters.

When you score every conversation, there is no sampling error in the agent-level number, because you are not estimating the agent’s quality from a subset, you are measuring all of it. The confidence interval collapses, and the ranking, comparison, and trend questions become answerable rather than noisy. This is the statistical reason the shift from sampling to coverage is not a nice-to-have: it is the only move that makes agent-level QA decisions defensible. It also removes the argument covered in QA call selection, because with everything scored there is no sample to be biased.

Full coverage introduces one new obligation in place of the old one: instead of trusting a sample to be representative, you have to trust the automated scorer to be accurate, which is why every score should trace back to the evidence and be checkable against conversations your reviewers agreed on. That is the subject of validating AI QA scoring. You trade a statistical problem you cannot solve for a verification problem you can.

Frequently asked questions

Why can’t a small QA sample support agent-level decisions?

Because a sampled score is an estimate with a margin of error, and a small sample produces a wide one. With only a handful of conversations per agent, the confidence interval around the score can span many points, wide enough that two agents who genuinely differ look identical and two who are identical look different. Ranking or comparing agents on that is mostly reading noise.

How much does adding more sampled conversations help?

Less than you would hope. Margin of error shrinks with the square root of sample size, so halving the error requires quadrupling the conversations scored. The first few conversations give a very imprecise estimate, and reaching a precise one by sampling is slow and expensive, which is why you cannot practically sample your way to a defensible agent-level score.

Does full coverage remove sampling error?

For the agent-level question, yes. When every conversation is scored, you are measuring an agent’s quality rather than estimating it from a subset, so there is no sampling error in the number. The confidence interval collapses and comparisons, rankings, and trends become answerable. It replaces the statistical problem of sampling with the verification problem of trusting the scorer, which is solvable.

Are the numbers in this article benchmarks?

No. Any specific figures are illustrative of the mechanism, chosen to show how the margin of error behaves. The point is the shape of the math, which holds for any small sample regardless of the exact numbers: small samples carry wide margins of error, and that error shrinks only with the square root of sample size.

See what your sampled scores can and cannot tell you

Tell us how many conversations you sample per agent. We will show you, on your own numbers, roughly how wide the margin of error around each score is, which of your current comparisons sit inside the noise, and what changes when every conversation is scored instead. Every score traces back to the evidence in the transcript.

Book a demoExplore Agentic Auto QA

Choose your help desk

Not using either? We’ll let you know as soon as we can support your help desk solution.

Kaizo
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.