Whether a 95% QA target is achievable is an arithmetic question, not an opinion. It depends on how many scored criteria your scorecard has, how large each deduction is, and how many conversations an agent is measured on, and for a lot of scorecards the honest answer is no: one ordinary miss already puts a conversation well below 95. Work out the ceiling your own scorecard allows before you commit to a number, then set the target from your current score distribution rather than from a figure that sounded demanding in a meeting.
In short
- A target does not say be good. It says how many misses a conversation is allowed, and deduction size decides that, not the number on the slide.
- Scorecards have a resolution. If your criteria are worth sixteen points each, no score exists between 84 and 100, so a 95 target is a 100 target with friendlier packaging.
- Targets and scorecards are almost always written by different people at different times, and nobody multiplies them together. That is how impossible targets get published.
- One auto-fail in a month of ten reviews caps the average at 90. No amount of later perfection recovers it.
- If your scores cluster at the top and the bottom with nothing in between, the scorecard has stopped measuring and started sorting, and there is nothing left to coach toward.
- An unreachable target inverts the incentive, because the agents taking the hardest conversations have the most surface area to lose points on.
- Set the target at a percentile of your current distribution, publish it in misses rather than points, and move it on a schedule.
What a 95% QA target is actually asking for
Quality targets get chosen the way round numbers always get chosen. Ninety-five sounds demanding without sounding cruel, it fits on a slide, and nobody in the room has to defend it. It is then handed to a scorecard that was designed by different people on a different day, and the two are never multiplied together.
Do the multiplication and the target stops being a motivational number. It becomes a specific instruction about how many mistakes a conversation is permitted to contain, and that permission is set entirely by the size of your deductions.
Your scorecard has a resolution, and it is coarser than you think
A percentage implies a hundred possible values. Almost no scorecard has a hundred. Most have between six and fifteen criteria, each marked pass or fail with no partial credit, which means the scores an agent can actually receive land on a small grid of reachable values with gaps between them. The gaps are the size of your deductions.
That grid is what makes a target meaningful or decorative. If your smallest realistic deduction is four points, a 95 target permits exactly one small miss. If your smallest deduction is sixteen points, there is no reachable score anywhere between 84 and 100, so a 95 target does not permit one small miss. It permits nothing at all. It is a 100 target wearing a friendlier number, and the agents being measured worked that out long before the person who set it did.
This is not an argument against demanding standards. It is an argument against standards nobody has checked against the instrument that measures them. Before you defend or attack a target, read the scorecard as arithmetic rather than as a list of good behaviours. Manufacturing formalised this decades ago. NIST’s engineering statistics handbook defines process capability as the comparison between what an in-control process actually produces and the specification it is being held to, and a specification the process cannot reach is a specification problem rather than an effort problem.
Work out the ceiling for your own scorecard
Here is the method, run against an example scorecard. The scorecard below is invented for illustration and is not a benchmark or a recommendation. Use the method, not the numbers.
It has two auto-fail gates that void the interaction, six substantive criteria worth sixteen points each, and one administrative criterion worth four. That totals a hundred.
| Criterion (example scorecard, illustrative only) | Type | Cost of failing |
|---|---|---|
| Identity verified before account details discussed | Auto-fail gate | Score becomes 0 |
| Customer data handled per policy | Auto-fail gate | Score becomes 0 |
| Resolution was factually correct | Scored | 16 points |
| Policy applied correctly | Scored | 16 points |
| Answered the question that was actually asked | Scored | 16 points |
| Next steps and ownership made explicit | Scored | 16 points |
| Tone appropriate to the situation | Scored | 16 points |
| Effort acknowledged before the fix was offered | Scored | 16 points |
| Ticket tagged and categorised correctly | Scored | 4 points |
Four steps from a scorecard to a defensible number
Step one. Separate the gates from the pool. Auto-fails do not deduct, they void. Pull them out and treat them separately, because mixing them into the arithmetic is what produces most of the confusion. If you are not sure which of your criteria are genuinely gates, that distinction is worth settling first, and what an auto-fail is and when to use one covers it.
Step two. Enumerate the reachable scores from the top down. For the example scorecard that is 100, then 96 if only the tagging item fails, then 84 if one substantive criterion fails, then 80 for both, then 68. Write out the top five. It takes two minutes and it is the single most useful thing in this article.
Step three. Find the one-miss score. That is the highest reachable total below a perfect score once something real has gone wrong. Here it is 84. The 96 does not count, because failing to tag a ticket correctly is not the kind of miss a target is trying to prevent.
Step four. Compare the one-miss score with the target. If the one-miss score is below the target, the target is asking for perfection on every reviewed conversation. Say that out loud and see whether the room still wants it. In the example, 84 is a long way below 95, so a per-conversation 95 target means flawless, every time.
Now run it again where people are actually measured
Almost nobody is held to a target on a single conversation. The target usually applies to a monthly average across a handful of reviews, and that changes the arithmetic in ways worth working through.
Say each agent gets ten conversations reviewed a month, scored on the example card. To average 95 you need seven of the ten to be perfect. Six perfect and four with a single substantive miss averages 93.6, which fails. Seven and three averages 95.2, which passes. So the target, stated honestly, is this: you may have one real miss in three conversations, and nothing worse.
Then add a gate. One auto-fail in the month means nine perfect conversations and one zero, which averages 90. The target is unreachable for that month from the moment the gate trips, and no amount of later perfection recovers it, because the month has a fixed denominator. At five reviews a month instead of ten, one auto-fail averages 80 before anything else has gone wrong.
Two things follow. The first is that a target and a review volume are a single decision, not two, because halving the sample doubles the damage a single zero does. That makes your QA monitoring cadence part of the target, whether or not anyone wrote it down that way. The second is that when the review count is small, the score is a noisy estimate as well as a coarse one, which is a separate problem covered in what QA sampling is and where a thin sample breaks. If the point values themselves look arbitrary once you write them out, that is a weighting problem rather than a target problem, and it starts with the rubric the criteria came from.
Why impossible targets get published in the first place
Nobody sets out to publish a target their own scorecard cannot produce. It happens because the two artefacts have different authors and different birthdays.
Targets usually arrive from outside the QA function. They get inherited from someone’s previous employer, copied out of a client service agreement, agreed in a commercial negotiation before a scorecard existed, or set as last year’s number plus a bit. In each of those routes the number is chosen without reference to any criterion list, because at the moment of choosing there often is not one.
Scorecards, meanwhile, drift in one direction. An incident happens, so a criterion gets added. A regulator asks a question, so a gate gets added. A team lead notices an annoyance, so another line appears. Every addition is individually reasonable and every addition lowers the achievable ceiling, because there is more to get wrong and the same hundred points to spread across it. The target does not move, because nobody owns the pairing.
Ask when your current target was set, then ask when the scorecard was last changed. If the second date is later than the first, and it usually is, the target has gone unchecked for that entire gap.
One governance rule fixes most of this: a scorecard change is a target change. Version them as one artefact, re-derive the ceiling and the one-miss score every time a criterion is added or a gate is introduced, and publish both numbers together. If the ceiling moved and the target did not, you have quietly raised the standard without telling anyone, which is the version of this that does the most damage to trust.
The bimodal distribution problem: when a scorecard sorts instead of measuring
Coarse deductions and a long gate list produce a specific failure state, easy to diagnose and fatal to a QA programme’s credibility. Scores stop spreading out. They pile up at the top, pile up at the bottom, and leave the middle empty.
Agents describe it the same way every time. Every item is either an instant zero or costs sixteen points, so the only outcomes are a perfect score or a disaster. That is not a complaint about strictness. It is an accurate description of a distribution with two humps and nothing between them.
Three checks will tell you whether you have it, and all three run off scores you already hold.
- Resolution check. Count the distinct score values that actually occurred last quarter. Compare that with the number of values your scorecard could theoretically produce. If a hundred-point scale produced six observed values, you do not have a percentage. You have a six-point scale that is dressed as one, and everyone should stop discussing changes of two or three points as though they mean something.
- Middle-band check. What share of scores landed between the one-miss score and the target? In the example scorecard that band is 84 to 95, and if it is close to empty then the target has no approach ramp. Nobody is near it. They are either past it or nowhere near it.
- Variance source check. Recompute the spread with all auto-fail zeroes removed. If the spread collapses, your quality metric is functionally an auto-fail counter with decoration. That may be a legitimate thing to measure, but it should be named as such, and it should not be reported as an internal quality score.
When a distribution is bimodal, the scorecard has stopped measuring and started sorting. Sorting has its uses, but it leaves nothing to coach with, because there is no gradient. Two agents on 84 got there by failing entirely different criteria, and the number cannot tell you which.
The distribution, not the average, is the artefact worth putting in front of leadership, because a single mean hides all of this. Kaizo’s scorecard reporting in Insights shows the shape rather than only the mean, and because every conversation is scored rather than a sample of them, that shape is the whole population rather than an estimate of it.
How an unreachable target inverts the incentive
Here is the consequence that turns an arithmetic problem into a retention problem, and the reason an unreachable target is incompatible with running a QA programme agents trust. Difficulty is not distributed evenly across conversations, and neither is the opportunity to lose points.
A short password reset touches three criteria and has almost no way to go wrong. A billing dispute from an angry customer with a policy exception touches every criterion on the card, involves an identity check, requires a judgement call on tone under pressure, and depends on a handoff the agent does not control. The second conversation has several times the surface area of the first. Score both against the same target and the arithmetic quietly punishes the agent who took the hard one.
The behaviour that follows is rational, not dishonest. People cherry-pick easy tickets when the queue lets them, stop volunteering for escalations, argue about ownership at handoff, and ask whether a particular ticket can be excluded from review. And your strongest agents, the ones you deliberately route difficult work to, sit at the bottom of a team scorecard that leadership reads as a ranking of ability.
Reviewers usually notice and start compensating, marking a hard conversation more generously because it was hard. That is a humane instinct and it destroys comparability, because the score now encodes a private difficulty adjustment that no two reviewers make the same way.
The fix is to compare like with like rather than to compensate informally. Segment scores by conversation type or queue, set a target per segment where the distributions genuinely differ, and never rank across segments. This depends on how conversations reach the review queue in the first place, which is its own problem and is covered in how conversations get selected for monitoring. Segmenting is also where scoring every conversation earns its keep rather than being a coverage boast: when Auto QA scores the whole population, each segment still has enough conversations in it to have a real distribution, whereas a manual sample splits into slices too small to read.
Set the target from your distribution, not from a round number
The alternative to picking a number is deriving one. It takes an afternoon and it produces a target you can defend to both the board and the floor, which no round number can.
- Pull last quarter’s scores at the level the target applies to. If the target is a monthly average per agent, pull agent-months. If it is per conversation, pull conversations. Do not average anything yet, because averaging is what hid the problem.
- Compute the ceiling and the one-miss score using the four steps above, and write both at the top of the page.
- Compute the median, the 75th and the 90th percentile of what actually happened. These three numbers describe your programme far better than any single average does.
- Set the target at a percentile. The 75th is the usual defensible choice: a quarter of the team already clears it, so it is demonstrably achievable by people doing this job on this scorecard, and the other three quarters have a reachable place to go. If you want to be gentler, use the median as a floor and the 75th as the stretch.
- Check the target sits on the grid. The target must be at or below a reachable score, and at least one reachable score must sit between the median and the target. If there is no rung on the ladder between where people are and where you are asking them to be, the target cannot be approached, only jumped at.
- Publish it in misses, not in points. Say what it permits in plain language: one non-critical miss in three reviewed conversations and no gate failures. Agents can act on that sentence. They cannot act on 95.
- Set a review date and move it deliberately. A target is meant to move. Raise it when the distribution has shifted under it, not on a calendar reflex, and announce the move as a change of standard rather than slipping it in.
Two cautions. A percentile target moves as the distribution moves, so re-derive it rather than freezing it, or you end up with the same stale number you were trying to escape. And the whole exercise assumes reviewers are calibrated with each other: if two reviewers score the same conversation differently, your percentiles are partly a picture of reviewer variation rather than agent quality. Run a calibration session before you derive anything.
The point is modest and worth stating plainly. A target claims that a standard is both worth reaching and possible to reach. The first half is a judgement call and yours to make. The second half is arithmetic, and arithmetic does not care how the number sounded in the meeting.
Frequently asked questions
Is a 95% QA target realistic?
It depends entirely on your scorecard, and you can settle it in about ten minutes. Work out the highest score an agent can reach after one genuine mistake. If that number is below 95, then a per-conversation 95 target is asking for a flawless conversation every time, whatever the intention was. Then check it at the level people are measured: if an agent gets ten reviews a month, a single auto-fail caps the monthly average at 90 no matter how good everything else is. A 95 target is realistic on a scorecard with small, granular deductions and few gates, and unrealistic on one built from large deductions and a long gate list.
What is a good QA score?
There is no score that is good in general, because a score only means something relative to the scorecard that produced it. An 88 on a scorecard with two gates and fifteen small criteria describes a very different conversation from an 88 on a scorecard with six gates and six large ones. The useful question is not what number is good but where a number sits in your own distribution: the median, the 75th percentile and the reachable grid of your scorecard tell you far more than any figure quoted from elsewhere.
What is the average QA score?
No trustworthy cross-industry average exists, and you should be sceptical of any figure presented as one. Averages are not comparable between organisations because scorecards are not comparable: criterion counts, deduction sizes, gate policies, review volumes and reviewer strictness all differ, and each of them moves the average independently of actual quality. Compute your own average, but treat it as a description of your instrument as much as of your team, and look at the distribution rather than the mean.
How do I calculate the achievable ceiling for my QA scorecard?
Separate the auto-fail gates out, because they void rather than deduct. Then list the scored criteria with their point values and write out the highest few totals an agent could actually receive, from a perfect score downward. The highest total below perfect that follows a genuine mistake is your one-miss score. If your target sits above the one-miss score, the target requires perfection. Repeat the calculation across the number of conversations an agent is reviewed on per period, because that is where the target usually applies.
Why do our QA scores cluster at the top and the bottom with nothing in between?
Because the deductions are large relative to the scale, the gate list is long, or both. When every criterion is worth a big slice and several failures void the score outright, there are very few reachable values in the middle, so scores pile up at the extremes. Count the distinct score values you actually observed last quarter to confirm it. If a hundred-point scale produced a handful of values, the scorecard has stopped measuring degrees of quality and started sorting conversations into acceptable and unacceptable, and there is no gradient left to coach along.
Should every agent have the same QA target?
Only if they handle comparable work. Complex conversations touch more criteria, involve more judgement and carry more gate exposure than simple ones, so the same target is materially harder to hit in an escalations queue than in a password-reset queue. The result is that the people taking your hardest conversations score worst, which inverts the incentive you were trying to create. Segment scores by conversation type or queue, set targets per segment where the distributions genuinely differ, and avoid ranking across segments at all.
Related terms
Set a target your scorecard can actually produce
Bring the scorecard you use today and a quarter of scored conversations. We will plot the distribution, work out the ceiling your criteria allow, and show you which targets are reachable and which are arithmetic.