When your internal quality score and CSAT point in different directions, the first thing to check is whether they were computed on the same conversations, because usually they were not. CSAT comes from the minority of customers who chose to respond and your quality score comes from a small sample a reviewer chose, so the two sets often barely overlap and the disagreement can be arithmetic rather than real. Once you have ruled that out, a persistent gap means your scorecard is measuring compliance with a process instead of the outcome the customer cares about.
In short
- Before diagnosing anything, check the overlap. If the surveyed conversations and the reviewed conversations are different sets, you are comparing two populations, not two verdicts.
- High QA with low CSAT is the common case and it usually means the scorecard rewards process compliance over outcome.
- Low QA with high CSAT is more interesting: often the agent broke process to help the customer, which is a policy finding rather than a coaching one.
- The two metrics measure different objects by design. QA measures the process you control; CSAT measures the outcome plus everything you do not control.
- Neither number wins by default. Escalations, repeat contacts and refunds are the tiebreaker, because they are behaviour rather than opinion.
- A quality score that never disagrees with CSAT is not well calibrated, it is redundant.
Are the two numbers even describing the same conversations?
Check this before you diagnose anything, because it is the most common explanation and almost nobody checks it.
Consider how each number is actually produced. CSAT comes from customers who received a survey and chose to answer it, typically a small fraction of contacts. Your quality score comes from conversations a reviewer selected, typically around 3% of volume. These are two independent selection processes running over the same month, and the intersection between them is often close to empty.
So when the two disagree, the honest first hypothesis is not that one is wrong. It is that you have measured different conversations and are treating the difference as a contradiction.
The five-minute test. Pull the conversations that received a CSAT response last month. Pull the conversations that were reviewed. Count the overlap. If it is in single figures, you do not currently have a comparison, and no amount of analysis on top of those two numbers will produce one. The fix is to deliberately review a set of conversations that have CSAT responses attached, which makes the comparison possible for the first time.
There is a second selection problem stacked on the first. Survey response is not random: people who respond differ systematically from people who do not, which is why non-response bias is treated as a core methodological problem rather than a rounding issue. Pew Research’s work on what low response rates mean for surveys is the clearest available treatment, and the implication for CSAT is direct. A 20% response rate is not a 20% sample of your customers’ experience, it is a complete sample of the customers who felt strongly enough to reply.
If fewer than about ten conversations last month have both a CSAT response and a QA review, your two metrics are not in disagreement. They are describing separate populations, and the gap you are looking at may contain no information at all.
What does each number actually measure?
They are not two attempts at the same measurement, which is why expecting them to match is the underlying error. They are measuring different objects, and a well-designed pair of metrics should sometimes disagree.
- Your quality score measures the process you control. Did the agent do the things your organisation decided constitute good handling? It is internal, it is comparable across agents, and it is only as good as the judgement encoded in the rubric.
- CSAT measures the outcome, plus everything else. The answer the customer got, the policy behind it, the wait, the product, the price, and their mood. Very little of that is the agent.
Once framed that way, the disagreement becomes informative rather than alarming. A customer can be unhappy about a correctly delivered no. An agent can be scored down for skipping a step while the customer got exactly what they needed. Both are true statements about different things, and an internal quality score exists precisely because the customer’s rating cannot tell you what to change.
A quality score that always agrees with CSAT is not measuring anything additional. If your two numbers move in lockstep, you have either built a scorecard that predicts satisfaction rather than describing handling, or somebody is scoring with one eye on the rating. Perfect agreement is a warning sign and not a success condition.
The four quadrants, and what each one means
Assuming you have ruled out the population problem, plot the two against each other. Each quadrant has a specific cause and a specific action, and they are not interchangeable.
| Pattern | Most likely cause | How to confirm it | What to do |
|---|---|---|---|
| High QA, low CSAT | The scorecard rewards process compliance rather than resolution, or the customer is unhappy about policy rather than handling | Read ten high-scoring conversations with low ratings. Ask whether the customer’s problem was actually solved | Reweight toward resolution, or reclassify the complaint as a policy issue and route it out of QA |
| Low QA, high CSAT | The agent broke process to help, or the scorecard is penalising things customers do not notice | Check what they were marked down for. If it is tone, formatting or macro use, look at whether the criterion earns its weight | This is a policy or rubric finding, not a coaching one. Do not coach an agent for a satisfied customer |
| Both low | Genuine performance or capacity problem, or a process that is failing everyone | Check whether it is concentrated in specific agents or spread across all of them | Spread means process. Concentrated means coaching. Diagnose before acting |
| Both high, and you do not believe it | Reviewer leniency, a scorecard everyone passes, or survey selection favouring resolved contacts | Look at the score distribution. If almost nothing fails, the scorecard has stopped discriminating | Recalibrate reviewers and check whether any criterion has passed 95% of the time for two quarters |
Why is high QA with low CSAT the most common pattern?
Because scorecards drift toward what is easy to observe, and what is easy to observe is rarely what matters. Greeting, sign-off, tagging and tone can all be graded quickly and consistently. Whether the customer’s problem is actually gone requires reading the whole conversation and knowing the policy.
So the criteria that are cheap to score accumulate, the criterion that matters most gets one line, and you end up with a scorecard where a conversation can score 95% while leaving the customer with an unresolved issue and a polite sign-off. Reweighting toward outcome is the fix, and weighting criteria from evidence is the method: compare each criterion’s failure rate against a downstream outcome such as repeat contact, and let the criteria that predict nothing lose their weight.
There is a second cause worth separating out, because the action is completely different. Sometimes the handling was genuinely good and the customer is unhappy with the answer. A correctly delivered refusal, an accurate but unwelcome policy, a price. That is not a quality failure and it should not be coached as one. It is a finding for whoever owns the policy, and the most useful thing QA can do with it is count it, route it, then stop. Treating it as customer dissatisfaction to be reduced by better handling will not work, because the handling was not the problem.
Tag low ratings by whether the complaint is about the handling or about the answer. Two months of that single distinction tells you how much of your dissatisfaction is a support problem at all, and it is usually less than leadership assumes.
Which number wins when you have to choose?
Neither. Use behaviour as the tiebreaker. Both numbers are opinions, one held by a reviewer and one by a customer. What customers did afterwards is not an opinion, and it is already in your helpdesk.
Three tiebreakers, in order of usefulness:
- Repeat contact within seven days. The cleanest signal available. A customer coming back is a statement that the first contact did not resolve it, regardless of what either number said. If your high-scoring conversations generate repeat contacts, the scorecard is wrong.
- Escalations and complaints. Costly, deliberate customer actions. A team with excellent ratings and rising escalations has a measurement problem somewhere, and escalation handling quality is where to look first.
- Refunds and goodwill given to settle service failures. Money leaving the business to fix something is the least deniable evidence there is.
Then run the check the other way round. Take the conversations customers rated badly and have reviewers score them blind, without seeing the rating. If reviewers consistently score them well, your rubric and your customers disagree about what good means, and the rubric is the thing you can change. This is the same mechanism as a calibration session, pointed at the scorecard rather than at the reviewers.
The scale problem is worth naming, because it is what makes this hard rather than merely tedious. Any of these checks needs enough conversations with both a rating and a review to be readable, and at a 3% sample that intersection is too small to segment. Kaizo’s Auto QA scores the whole population on your own rubric, which means every surveyed conversation also has a quality score attached, and the comparison stops depending on two samples happening to overlap. That is what 100% coverage revealing trends that 3% sampling never could buys you here: not a better score, but the ability to ask the question at all.
What should you change first?
Work in this order. Each step is cheap and each one removes a possible explanation, which is what stops you reweighting a scorecard to fix a sampling artefact.
- Measure the overlap. Count conversations with both a rating and a review. If it is tiny, fix that first by reviewing surveyed conversations deliberately, and stop analysing until you have.
- Check reviewer agreement. If two reviewers disagree with each other, the quality score is not stable enough to be compared with anything. Reviewer variance masquerades as a CSAT-versus-QA problem more often than it gets diagnosed as itself.
- Separate handling complaints from answer complaints. Two months of tagging tells you how much of the gap is even a support problem.
- Test each criterion against repeat contact. Criteria that fail without moving any downstream outcome are carrying weight they have not earned.
- Only then reweight. And when you do, version the scorecard so the historic scores remain readable against the rubric that produced them.
What not to do is set a target that forces the two numbers to agree. A quality score managed toward CSAT stops being an independent measurement and becomes a second, more expensive way of reporting the same thing. The reason to run both is that they see different failures, and what CSAT can and cannot measure is the wider version of that argument. Keeping them independent is the point, and occasional disagreement is the evidence that you have.
Frequently asked questions
Why are our QA scores high but CSAT low?
Most often because the scorecard rewards process compliance rather than resolution, so a conversation can score highly while leaving the customer’s problem unsolved. The second common cause is that the customer is unhappy with the answer rather than the handling, which is a policy finding and should not be coached. Before either, check whether the surveyed conversations and the reviewed conversations are actually the same set, because usually they are not.
Which is more reliable, CSAT or an internal quality score?
Neither, because they measure different objects. A quality score measures the process you control and is comparable across agents. CSAT measures the outcome plus the policy, the wait, the product and the customer’s mood, very little of which is the agent. Use behaviour as the tiebreaker instead: repeat contact, escalations and refunds are actions rather than opinions.
Should QA scores and CSAT match?
No, and a pair that always agrees is a warning sign. It means either the scorecard was built to predict satisfaction rather than to describe handling, or reviewers are scoring with one eye on the rating. The reason to run both metrics is that they detect different failures, so occasional disagreement is evidence they are working independently.
What does it mean when an agent has low QA scores and happy customers?
Usually that they broke a process to help the customer, or that the scorecard penalises things customers do not notice such as tone, formatting or macro use. Check what they lost points for before coaching anyone. If the criterion does not predict any downstream outcome like repeat contact, this is a rubric finding rather than a performance one.
Is a 20% survey response rate a 20% sample?
No. It is a complete sample of the customers who felt strongly enough to respond, which is a systematically different group from those who did not. Non-response bias is a well-documented methodological problem rather than a rounding issue, and it means CSAT describes a self-selected population. That is the main reason it needs an internal measure alongside it rather than instead of it.
Related terms
Find out whether your two numbers describe the same conversations
Bring a month of surveyed conversations and your current scorecard. We will score them on your own rubric so you can see where the rating and the review genuinely disagree, and where they never met.