CSAT is the only measure of support quality that comes from the customer, and nothing internal replaces it, but it can only carry a narrow load. It reports an outcome rather than a process, it arrives from a self-selected minority, and it cannot be attributed cleanly once a conversation has changed hands. The answer is not a better survey. It is a second, internal measure of how conversations were handled, so that when the two disagree you can read the disagreement as information instead of treating it as a broken number.
In short
- CSAT is not worthless and you should not switch it off. It is the only number in your reporting that comes from the person you actually served.
- A 20% response rate is not a 20% sample. The bias does not come from the rate, it comes from the fact that whatever makes someone answer is usually related to what you are trying to measure.
- CSAT measures the outcome, not the process, so it can tell you a conversation went badly and never what to change.
- On a transferred, reopened or bot-assisted conversation the rating lands on whoever touched it last, which makes agent-level CSAT the least reliable version of the metric.
- An internal quality score and CSAT answer different questions. A high score on a conversation the customer hated is not a contradiction, it is a finding.
- Most teams treat the disagreement between the two as a measurement fault and adjust one until they match. That removes the only independent check either of them had.
Why does CSAT look fine while escalations keep coming?
Because your satisfaction score and your escalation queue are measuring two different things, and only one of them is a measure of how the work was done. CSAT records how a customer felt at the moment you asked them. Escalations record what the conversation actually cost, in second touches, supervisor time, credits and rework. Those two can move in opposite directions for months without either of them being wrong.
The situation has a shape. Response rates sit in the teens, the team hits its target most weeks, and yet the same contact reasons keep producing escalations and nobody can point at the score and say what to fix on Monday. Eventually someone asks the question that matters, which is not how to raise the number but what it measures.
This page answers that and maps where to go next. It is not an argument that CSAT is a bad metric. It is an argument about load: what it can carry alone, what it cannot, and what has to be built next to it.
What is CSAT genuinely good at?
It is the only number in your reporting that comes from outside your own organisation. Your quality score, your handle time, your resolution rate and your first-contact resolution are all your company grading its own work against its own definition of good. Giving that outside signal up to rely entirely on an internal one is a downgrade, and any version of this argument that ends with “so turn the surveys off” has gone wrong somewhere.
It is also cheap, immediate and comparable across time. A support leader can put a CSAT trend in front of a board without explaining the methodology first, which is not true of anything you build internally.
Swapping the question does not fix what follows, either. CSAT, NPS and CES differ in what they ask and barely at all in who answers, so a loyalty question inherits the same self-selection. The same holds for most other ways to measure customer satisfaction: post-hoc reports of a feeling, from whoever felt like replying.
The useful question is not whether CSAT is good or bad. It is which questions it can actually answer.
| What you want to know | Can CSAT answer it | Why |
|---|---|---|
| Did this specific interaction land badly for this specific customer? | Yes, when they answer | It is a direct report from the customer rather than an inference drawn from operational data |
| Is dissatisfaction rising or falling across a large population over time? | Yes, in aggregate | Systematic bias stays roughly constant, so the direction of travel survives even when the level is unreliable |
| Which contact reasons produce the worst experiences? | Partly | Only where volume is high enough that the responding minority is still a usable number of people |
| What did the agent do, or fail to do? | No | The customer is rating an outcome, and frequently rating the policy, the product or the wait rather than the handling |
| Is this agent performing better or worse than that one? | No | Per-agent samples are small, self-selected, and the mix of contacts each agent handles is not comparable |
| Was the conversation handled correctly? | No | A customer can be satisfied by a wrong answer and dissatisfied by a correct one |
Is a 20% response rate a 20% sample?
No, and the reason matters more than the arithmetic. It would be a 20% sample only if the customers who answered were indistinguishable from those who did not, on the exact dimension you are measuring. In support they never are, because whatever makes someone bother to answer is usually related to how satisfied they are.
Be precise here, because the sloppy version of this argument gets knocked down constantly. Pew Research Center, which has tracked non-response for decades, found the response rate on its own is an unreliable indicator of bias, and that estimates from surveys answered by roughly 9% of those contacted were still accurate across a wide range of measures. A low response rate is not automatically broken data.
What it means is that your error is governed by whether the propensity to respond correlates with what you are measuring, and you cannot detect that correlation from inside the results. In a support queue it usually does.
- Customers who had a bad experience answer at higher rates than customers who had an unremarkable one. Effort predicts response.
- So do customers who had an unusually good one, so responses skew to both ends and the average between them describes almost nobody. It is the same reason the dissatisfaction half of the distribution is often more informative.
- Customers who gave up and left do not answer at all. The worst outcome removes itself from your sample.
- High-value accounts often route through named contacts or account managers, so their conversations never enter the survey population, and response rates differ by channel and locale anyway, so your sample composition shifts every time you change routing.
None of this is repaired by asking a better question. It is a property of who chooses to answer, which is why a rating from a fifth of your customers is not a small copy of all of them.
Why can’t the score tell you what to change?
Because it measures an outcome, and outcomes are produced by several things at once. A 2 out of 5 tells you the customer was unhappy. It does not tell you whether the agent misread the question, applied a policy that is indefensible said out loud, waited three days on a broken handoff, or gave a technically perfect answer the customer did not want to hear.
This is the difference between a result measure and a process measure, and it explains why setting a target on CSAT so reliably produces movement without improvement. When the score is the goal, the cheapest ways to move it are to influence who gets asked and how they feel when asked. Neither touches how conversations are handled.
The research on the underlying weakness is older than most CX teams. Dixon, Freeman and Toman’s study of more than 75,000 service interactions, published in Harvard Business Review, found that exceeding customer expectations produced very little additional loyalty over simply meeting them, and that reducing customer effort predicted loyalty considerably better than delight did.
What has to sit next to it is a measure of what happened inside the conversation. That is the distinction between an agent error and a process error, the only one that tells you whether to coach a person or change a rule. It is also why working the dissatisfied conversations beats raising an average: they are individually readable, and reading them is the work.
Who does the score belong to when the conversation changed hands?
Nobody, and any system that assigns it to one person is inventing an answer. This is the limitation least often written about and the one that does the most damage, because it is where a fuzzy team metric becomes an unfair individual one.
Five failures, all configuration defaults rather than deliberate choices:
- Last touch wins. Most survey configurations attribute the response to whoever closed the ticket. The agent who spent forty minutes on a hard diagnosis gets nothing, the colleague who sent the closing note gets the rating.
- The transfer penalty. On a conversation that crossed two teams, the rating is a verdict on both agents, the handoff and the queue time between them, compressed into one digit on one person’s record.
- The reopen. A customer satisfied on Tuesday who reopened on Friday leaves two ratings, one obsolete, both attributed to somebody.
- The bot handoff. Where an AI agent handled the first half, the human who picked up the rest inherits whatever the bot did.
- The delay. A customer answering four days later is rating a memory that has been coloured by everything since.
The consequence is specific: agent-level CSAT is the least reliable version of this metric and the one most often used in performance conversations. That turns a measurement debate into a trust problem, because agents can see the attribution is wrong long before anyone reports it. A per-agent signal has to come from what the agent actually did, which means measuring conversation quality at the level of the messages they sent.
What about the conversations nobody ever rates?
They are the majority, and not a random majority. Between conversations that never trigger a survey and surveys nobody answers, most of what your team does in a month produces no satisfaction data at all. That would be tolerable if the missing set were arbitrary. It is not. The conversations that systematically disappear tend to be these:
- Channels with no survey attached: phone callbacks, social, community forums, anything outside the helpdesk.
- Tickets closed by automation, merged, or resolved after the customer stopped replying, all of which usually suppress the survey.
- Conversations handled end to end by an AI agent, where the survey is frequently switched off.
- Proactive outreach, where a rating measures your timing rather than your handling.
- Internal escalations resolved without the customer ever being told there was a problem.
Notice what the list has in common: it is disproportionately the automated, the transferred, the abandoned and the complicated. The rated set skews toward straightforward conversations one person handled and closed.
This is where the argument for measuring conversations internally comes from, and it is worth being exact about what it buys you. Not vigilance. A review population you selected on purpose behaves very differently from one your customers selected for you. Kaizo’s Insights layer reports against the full conversation set rather than the rated subset, which is the practical version of scoring every conversation: the point is removing a selection step, not adding scrutiny.
What does an internal quality score measure that CSAT does not?
A different question, which is the entire point of having both. CSAT asks the customer how the outcome felt. An internal quality score asks a trained reviewer, or a grader working to the same rubric, whether the conversation was handled the way your team said it should be. These are not two attempts at one measurement, and expecting them to agree is the mistake that starts most of the arguments here.
| CSAT | Internal quality score | |
|---|---|---|
| Who produces it | The customer, if they choose to respond | A reviewer or an automated grader, on whichever conversations you select |
| What it measures | How the outcome felt | What was done, against a written standard |
| Coverage | A self-selected minority of a subset of channels | Whatever you decide, up to all of it |
| Attribution | Last touch, in most configurations | Per message and per agent, traceable to the evidence that produced it |
| Can it be wrong | Not really. It is an accurate report of a feeling | Yes, which is exactly why it has to be calibrated and auditable |
| What it tells you to do | Nothing directly | The specific criterion that failed, on the specific conversation |
| What it cannot tell you | Whether the process is sound | Whether the customer was actually happy |
What should you do when your QA score and CSAT disagree?
Read the disagreement. Do not resolve it. The most expensive mistake here is treating a gap between the internal score and the customer score as a measurement fault, then adjusting one until they line up. Teams reweight the scorecard toward what customers mention, or add sentiment into the rubric as a criterion. Within two quarters they have two metrics that say the same thing and no independent check on either.
Correlation is not the objective. If your quality score predicted CSAT perfectly you would need only one of them, and the one you would keep is the customer’s. The value of the internal measure is precisely that it can be high when CSAT is low, and low when CSAT is high, because those two situations have different causes and different fixes.
Note the honest cost, because it is what keeps the argument straight. A conversation can be fully compliant, accurately documented, policy-correct and still be the worst thing that happened to that customer all week. A well-handled no is a 5 on your scorecard and a 1 in their inbox. That is not a defect in either measure.
Here is how to read each combination. For the mechanics, see how teams run an internal score day to day.
Before reading anything into the gap, check whether the two numbers were computed on the same conversations at all, which is usually where this resolves and is the first step in what to do when your QA score and CSAT disagree.
| Quality score | CSAT | What it usually means | What to do first |
|---|---|---|---|
| High | High | The standard and the customer agree. This is the boring quadrant and most of your volume should sit in it | Nothing, except sample it occasionally to check your reviewers are not simply being generous |
| High | Low | The conversation followed the process and the customer still lost. The policy, the product or the process is the problem, not the person | Read ten of them end to end. You are looking for a rule that is defensible internally and indefensible to a customer |
| Low | High | The agent was liked and the standard was not met. Usually an over-promise, a skipped verification step, or a workaround that will surface again later | Check specifically for policy, verification and compliance failures. This is the quadrant where the expensive mistakes hide, and it is the one nobody investigates |
| Low | Low | Straightforward. The handling was poor and the customer noticed | Coach it, then check whether the same criterion is failing across the whole team rather than one person |
| Disagreement concentrated in one team, channel or contact reason | Any | Not a measurement problem. Something structural is different there | Compare the two populations before you compare the two scores. The rated set for that group is probably not comparable to anyone else’s |
How do you build the internal half without abandoning CSAT?
Keep CSAT as it is, stop asking it to do jobs it cannot do, and build the second measure alongside it. The order matters, because most teams start at step four and then spend a year arguing about the rubric.
- Write down what CSAT is now responsible for. Usually a trend line and an early warning, both legitimate. Everything else, meaning agent performance, coaching topics and root cause, is being asked of a measure that cannot supply it.
- Stop using it at agent level. One decision, a week to communicate, and it removes most of the attribution arguments above.
- Define the standard first. A scorecard is a written definition of good handling on your product, in your policy environment. If you cannot write it down, every reviewer scores their own version instead.
- Score a stratified sample, not a convenient one. Cover contact reasons and channels in proportion to volume, and over-sample what CSAT never reaches: transfers, reopens and anything an AI agent touched.
- Calibrate before you report. Two reviewers, the same twenty conversations, compare. Where they disagree the rubric is ambiguous rather than the reviewers careless, and the rubric is what you fix.
- Only then put the two numbers side by side, at team and contact-reason level rather than agent level, and read the quadrants above.
- Review the disagreement monthly, as a standing item. It is the cheapest source of process findings you have.
None of that requires software. It requires a written standard, reviewer time and the discipline to sample properly, and building the QA programme is the substance of the work. What automation changes is the sampling constraint: when scoring is not rationed by reviewer hours, human time moves onto the disagreements. Procede Software, a Kaizo customer, reports a 13% annual increase in CSAT, which is the direction the causation runs in. You do not raise a satisfaction score by measuring satisfaction harder.
Two conditions hold this together. The internal score has to be traceable back to the evidence, so that when it disagrees with the customer you can open the conversation and see which criterion produced the number. And reviewers have to agree with each other, which makes calibration a prerequisite.
Frequently asked questions
What are the limitations of CSAT?
Four matter. It comes from a self-selected minority, so the sample is not representative in a way you can correct for. It measures the outcome rather than the process, so it cannot tell you what to change. It cannot be attributed cleanly once a conversation has been transferred, reopened or partly handled by a bot. And it is collected on only a fraction of conversations.
Is CSAT a good measure of support quality?
It is a good measure of customer satisfaction and a poor measure of support quality, and those are different things. It is the only signal that comes directly from the customer, which makes it irreplaceable, but it reports how an outcome felt rather than how the work was done. Use it for direction of travel at team level, and an internal quality score to answer what to change.
What are common CSAT mistakes?
Setting a target on it and managing to the target, which changes behaviour rather than handling. Using it at individual agent level, where the sample is smallest and attribution weakest. Reading the average instead of the distribution. And, most damaging, adjusting the internal quality rubric until the two scores agree.
Why is our CSAT high when customers still complain?
Usually because the customers who complain and the customers who answer surveys are not the same people. Complaints concentrate in transfers, reopens and long-running threads, exactly the conversations least likely to produce a clean rating. A high average alongside a persistent escalation problem is a sign that the rated set does not represent the difficult set.
What is the difference between a CSAT score and a QA score?
CSAT is produced by the customer and measures how the outcome felt. A QA or internal quality score is produced by a reviewer or an automated grader and measures whether the conversation met a written standard. CSAT covers a self-selected minority and cannot be decomposed. A quality score traces back to the criterion that failed.
How do you measure support quality without relying on CSAT?
Define what good handling looks like in writing, turn it into a scorecard, score a stratified sample against it, and calibrate reviewers until they agree. That gives you a process measure you can act on. Keep CSAT alongside it as the outcome measure, and treat disagreement between the two as a finding rather than an error.
Related terms
- What DSAT is, and why the dissatisfied half is more actionable
- How to turn QA data into coaching people accept
- How to QA conversations an AI agent handled
- Building a QA scorecard that measures handling, not outcome
- Customer service metrics and KPIs, and which ones matter
- When your QA score and CSAT disagree, who is right?
Find out what your unrated conversations look like
Bring a month of conversations and your current satisfaction results. We will score the whole set against your own standard and show you where the two measures disagree, because that gap is where the process findings are.