AI customer service metrics are the measures used to judge whether an AI agent is handling conversations well. The two most commonly reported, containment rate (the share of conversations the AI handled without a human) and post-AI CSAT (satisfaction surveys sent after an AI conversation), are also the two most misleading, because containment rewards the bot for stonewalling the customer and CSAT is answered by a small self-selected group. A measurement stack that holds up looks at resolution quality, factual accuracy, escalation quality, repeat contact rate, silent abandon rate, policy adherence, and cost per genuinely resolved contact. None of those can be read off a survey or a 2% sample, so they require scoring every conversation against a rubric. The decisive question is who holds the rubric: most QA and CX platforms now sell AI agents of their own, which makes their AI performance reporting a self-assessment. Insist on an evaluator that is independent from the AI being graded and on scores you can verify, each one traced to the evidence in the transcript.
In short
- Containment rate measures whether a human got involved, not whether the customer got helped, so a conversation someone abandoned in frustration counts as a win.
- Post-AI CSAT is answered by a self-selected minority and systematically misses the silent quitters, who are the customers your AI failed hardest.
- The replacement stack is resolution quality, factual accuracy, escalation quality, repeat contact rate, silent abandon rate, policy adherence, and cost per genuinely resolved contact.
- Resolution quality is the anchor metric: it asks whether the customer’s issue was actually solved, which is a judgment about the transcript, not about the routing.
- Repeat contact rate and silent abandon rate are the two cheapest lie detectors you have, because they catch failures that containment scores as successes.
- Directional reading beats target setting here: a repeat contact rate that rises after AI handling is a red flag no matter what its absolute value is.
- The grader’s independence decides whether any of these numbers are usable: most QA and CX platforms now sell their own AI agents, so their AI performance reporting is a self-assessment.
- Every number here should be verifiable, meaning each score traces to the exact evidence in the transcript and the grader itself can be validated against conversations your own reviewers scored.
- None of these can be measured from a 2% sample or a survey response rate, so scoring every conversation is the precondition that makes an unbiased measurement possible in the first place.
Why the two default AI metrics mislead
Every support leader deploying AI is being asked the same question by their board: prove it works. The two numbers that get reached for first are containment rate and post-AI CSAT, because both are easy to pull and both go up and to the right if you squint. They are also the two numbers least likely to describe what actually happened to your customers.
The problem is structural, not a matter of tuning. Containment measures a routing outcome: did a human touch this conversation. Post-AI CSAT measures an opinion, but only the opinions of people who chose to give one. Neither of them looks at the substance of the conversation, which is the only place the answer lives. That is a very different problem from the one solved by the standard method for measuring AI agent performance, and a very different problem from choosing which of the general customer service metrics belong on your dashboard. This is about which numbers are quietly lying to you.
Here is the short version of the argument, metric by metric, before we go into each in detail.
| Metric | What it claims to show | How it misleads | What to measure instead |
|---|---|---|---|
| Containment / deflection rate | The AI handled the conversation successfully | Counts abandonment and stonewalling as success, because no human got involved | Resolution quality plus silent abandon rate |
| Post-AI CSAT | Customers are happy with the AI | Answered by a self-selected minority, and the angriest customers just leave | Repeat contact rate and resolution quality on 100% of conversations |
| Average handle time on AI conversations | The AI is fast | A fast wrong answer is faster than a slow correct one | Factual accuracy and cost per genuinely resolved contact |
| Escalation rate | Fewer escalations means the AI is coping | Suppressing escalations is easy and harmful; the timing and handover matter more | Escalation quality: right moment, full context |
| Volume handled by AI | Scale and ROI | Says nothing about what happened inside that volume | Policy adherence scored against your rubric |
Containment rate: the metric that rewards stonewalling
Containment rate, sometimes called deflection rate, is the share of conversations that ended without a human agent getting involved. It became the default AI metric because it maps directly onto the business case: every contained conversation is a contact you did not pay a human to handle.
The flaw is in the definition. Containment does not ask whether the customer got what they came for. It asks whether anyone else got pulled in. That means several very different outcomes all get counted identically as a win:
- The genuine resolution: the customer asked, the AI answered correctly, the customer left satisfied. This is the outcome the metric was designed to capture.
- The abandonment: the customer asked three times, got a variation of the same unhelpful answer, and closed the window. No human was involved, so it counts as contained.
- The stonewall: the AI could not or would not hand over, kept offering help center articles, and the customer gave up and went to social media or the app store review section instead.
- The displaced contact: the customer closed the chat and called the phone line, or emailed, or opened a new ticket the next day. Contained in one channel, still costing you in another.
Three of those four are failures, and the stonewall is actively worse than doing nothing. Yet a containment number cannot distinguish between them, because the distinguishing information is inside the conversation and containment never looks there.
This is not an argument for abandoning containment. It is a useful capacity and cost input. It is an argument against reporting it as a quality metric, because optimizing containment directly means making it harder for customers to reach a human, which is exactly the behavior nobody intends and everybody recognizes as a customer.
Post-AI CSAT: the metric that never hears from the people who left
CSAT is a good metric in its home context: a human conversation, a reasonable response rate, a stable population answering it. Bolted onto an AI conversation it develops a specific and serious blind spot.
Self-selection at the exact wrong moment
Survey responses come from a minority of customers, and that minority is never random. People who felt strongly enough to respond skew toward the two ends. But the AI failure mode that matters most, the customer who quietly concludes this is not going anywhere and disengages, is precisely the customer least likely to fill in a survey afterward. They did not stay to be surveyed. They are the silent quitters, and post-AI CSAT is structurally deaf to them.
The survey asks about the wrong thing
Even among people who do respond, a satisfaction score conflates whether the answer was correct, whether the tone was pleasant, and whether the outcome was one the customer wanted to hear. An AI that is charming and confidently wrong can score well. An AI that correctly delivers a policy answer the customer dislikes can score badly. Neither result tells you whether the system is working.
The comparison is unfair in both directions
Teams often compare AI CSAT against human agent CSAT and draw conclusions. Those two populations are not comparable: the AI usually takes the simple, high-volume, easily satisfied contacts first, and escalates the hard ones. A flattering AI CSAT can simply mean the AI was given the easy queue. A poor one can mean it was handed the queue nobody could win.
Keep CSAT. Just stop treating it as evidence about your AI. Treat it as one signal among several, weighted by the knowledge that it only hears from the people who stayed.
Resolution quality: the anchor metric
If you replace containment with one thing, replace it with resolution quality: did this conversation actually solve the customer’s problem. It is the anchor of the whole stack because every other metric here is either an input to it or a check on it.
What it is. A judgment, made against the transcript, about whether the customer’s underlying issue was addressed. Not whether a ticket was closed, not whether a response was sent, not whether the intent was matched. Whether the thing the customer needed to happen, happened.
How to measure it. As a scored criterion on your QA rubric, applied to AI-handled conversations exactly as it is applied to human-handled ones. The criterion needs to be written so that a reviewer and an automated evaluator would read it the same way: state what evidence in the transcript counts as resolution for each major contact type. Refund requested and refund confirmed. Address changed and change verified back to the customer. Question asked and correct answer given with the customer acknowledging it.
What a bad number looks like. The signal to watch is not an absolute threshold, it is the gap between resolution quality and containment. If a large share of conversations are contained but a much smaller share are scored as resolved, that gap is the size of your problem, expressed in conversations. It is also the number worth putting in front of a board, because it is the honest version of the one they were shown.
Note that this is deliberately stricter than first contact resolution, which infers resolution from the absence of a follow up. FCR is a useful proxy. Resolution quality reads the transcript instead of inferring from silence, and silence is exactly what an abandoning customer produces.
Factual accuracy and hallucination rate
An AI agent that invents a refund window, misstates a delivery timeline, or confidently describes a feature you do not have is not a quality problem, it is a liability problem. Factual accuracy is the metric that catches it, and it is the one most teams have no measurement for at all.
What it is. The share of AI responses containing a factual claim that is checkable against your knowledge base, policy documents, or order data, and the share of those claims that are correct. Hallucination rate is the inverse: claims that are confidently stated and unsupported by any source.
How to measure it. Every factual assertion in a conversation gets checked against the grounded source of truth. This is only practical automatically, at volume, which is one of the strongest arguments for scoring every conversation rather than sampling: a hallucination in the 98% you did not read costs exactly as much as one in the 2% you did. Track it separately by contact type, because accuracy is rarely uniform. The billing questions and the shipping questions usually behave very differently.
What a bad number looks like. Any nonzero rate on claims that carry legal, financial, or safety weight is a bad number, regardless of how small it is. On lower-stakes claims, the direction is what matters: a hallucination rate that climbs after a knowledge base change or a model update is telling you something specific and urgent about that change.
One caution: the AI that generated the answer is not a reliable judge of whether the answer was true. Accuracy has to be assessed by something outside the system that produced it, which is the same principle behind LLM as a judge being kept separate from the model under evaluation.
Escalation quality: not how often, but how well
Escalation rate is reported as though lower is better. It is not. Escalation is a feature, and an AI that never escalates is not a strong AI, it is a trapped customer. What matters is escalation quality: whether the handover happened at the right moment and arrived with the right context.
Timing
The two failure modes sit on either side of correct. Escalating too early wastes the automation entirely and irritates a customer who had a simple question. Escalating too late, after three or four failed loops, means the human inherits a customer who is already angry and a conversation that has to be restarted. Score the moment: was there an identifiable turn where a competent handover should have happened, and did it happen there?
Context transfer
The most expensive escalation failure is the one where the human agent opens the conversation and has to ask the customer to explain everything again. That single behavior undoes the goodwill of the entire AI interaction. Score whether the handover carried the customer’s stated issue, what was already attempted, and any account or order context the AI had already retrieved.
Correctness
Was escalation the right call at all? An AI that escalates every ambiguous message is generating human workload for nothing, which shows up in cost per resolved contact rather than in quality scores. Both directions need scoring, or you will optimize one into the other.
What a bad number looks like. A falling escalation rate alongside a falling resolution quality score is the clearest bad signal in this entire article. It means the AI is holding onto conversations it cannot finish.
Repeat contact rate and silent abandon rate
These two are grouped together because they are your cheapest lie detectors. Both catch failures that containment scores as successes, and neither requires a customer to fill anything in.
Repeat contact rate after an AI conversation
What it is. The share of customers who come back within a defined window, in any channel, about the same underlying issue after an AI-handled conversation.
How to measure it. Link contacts by customer and by issue, not by ticket, and specifically look across channels. A customer who abandoned a chat and then phoned is the single most important case to catch, and a per-channel view will miss them completely. Set the window to match your product: same day for a delivery issue, a week or more for a billing dispute.
What a bad number looks like. A repeat contact rate that rises after AI handling compared with the equivalent human-handled contacts is a red flag, whatever its absolute value. It means containment moved work rather than removing it.
Silent abandon rate
What it is. The share of AI conversations the customer exits without resolution and without escalating. No survey, no complaint, no follow up. These conversations look identical to successes in every routing-based metric.
How to measure it. Identify conversations that ended on a customer turn, or ended with the AI’s last message unanswered, and score whether the issue had been resolved at that point. Splitting satisfied departures from silent quits is a transcript judgment, which is exactly why this metric does not exist on most dashboards: it cannot be counted, it has to be read.
What a bad number looks like. Any silent abandon rate that is a meaningful fraction of your containment rate means your containment number is inflated by exactly that much. Reporting containment without it is reporting a number you know is wrong.
Policy adherence and cost per genuinely resolved contact
The last two close the loop between quality and the business case.
Policy adherence
What it is. Whether the AI followed the rules it is required to follow: verifying identity before disclosing account details, giving mandatory disclosures, refusing out-of-policy discounts, handling a vulnerable customer or a complaint disclosure correctly, staying inside regulatory language where it applies.
How to measure it. These are binary, evidence-checkable criteria, which makes them the easiest thing on this list to score automatically and the most damaging thing to sample. Run them as pass or fail checks on every conversation, and treat each fail as an incident to review rather than a percentage to average away. This is the same rubric logic your human team is already scored against, which is the point: your internal quality score should cover the AI too, or you have two standards and no comparison.
What a bad number looks like. On identity verification and regulated disclosures, one fail is a bad number. Do not average these.
Cost per genuinely resolved contact
What it is. Total cost of the AI conversations divided by the number of conversations that were actually resolved, not the number contained.
How to measure it. Take your resolution quality score, apply it to your contained volume, and use the resulting count as the denominator. Include the downstream cost the AI created: repeat contacts, the extra handle time on escalations that arrived without context, and the human time spent cleaning up.
What a bad number looks like. If cost per genuinely resolved contact is not meaningfully better than your human-handled equivalent, the deployment is not paying for itself no matter how good the containment chart looks. This is the number that turns a quality argument into a finance argument, which is usually the argument that gets budget.
Why none of this works from a sample or a survey
Read back through the seven metrics and one thing is true of every single one: it requires reading the conversation. Resolution quality is a judgment about the transcript. Hallucination rate is a check of claims against sources. Escalation quality is an assessment of a specific turn. Silent abandon is the absence of a signal, which means it can only be found by looking at conversations where nothing happened.
That has two consequences most measurement programs have not absorbed yet.
Sampling does not work here. Traditional QA reviews a small percentage of conversations, and for a human team of a known size that is a defensible statistical compromise. It is not defensible for AI, for two reasons. First, AI failures are systematic rather than distributed: a bad knowledge article or a policy edge case produces the same failure every time it is hit, so a sample either catches the whole cluster or misses it entirely. Second, AI handles far more volume, so 2% is a smaller window on a bigger operation. Scoring every conversation is the precondition rather than the point: it removes the selection bias that makes the other seven numbers meaningless. It is not a matter of watching anyone, it is a matter of not letting a sampling rule decide what your metrics say.
The grader has to be independent, and this is now the harder problem. Full coverage has quietly become ordinary. Independence has gone the other way. Most major QA and CX platforms now sell AI agents of their own or are building them, which means the system reporting on your AI performance is assessing a product its own company built. That is not a comment on anyone’s integrity, it is a structural conflict: a vendor’s evaluation criteria will tend to reflect what its system does well, and it chooses what counts as a resolution. Every metric in this article is only as trustworthy as the independence of whoever produced it. Ask directly whether the evaluator sells, or plans to sell, the agents it is grading.
And you should not have to take independence on trust either. Neutrality that cannot be checked is just a nicer claim. Insist that every score links back to the exact evidence in the transcript that produced it, that any score can be challenged and adjudicated by a human, and that the grader can be run against a set of conversations your own reviewers scored so you can see the agreement rate criterion by criterion before you report anything upward. A score you can trace is a score you can challenge. A score you cannot trace is a marketing claim with a number attached.
This is the reason Kaizo does not sell AI agents, and one of the very few QA platforms that can still say so. It has nothing in the conversation to defend, so it scores AI agents and your human team against one rubric you define, with every score traced back to the transcript and testable against your own reference set. It applies that rubric across all your conversations rather than a chosen slice. At UiPath, that approach automated 100% of QA with 200% ROI and an 8% lift in quality score. Practically, it means QA for AI agents and QA for humans stop being two separate programs producing two incomparable sets of numbers.
Frequently asked questions
What are the most important AI customer service metrics?
Resolution quality, factual accuracy, escalation quality, repeat contact rate, silent abandon rate, policy adherence, and cost per genuinely resolved contact. These seven describe what actually happened inside the conversation. Containment rate and post-AI CSAT, the two most commonly reported, describe routing and self-selected opinion instead, which is why they mislead.
What is containment rate and why is it misleading?
Containment rate, also called deflection rate, is the share of conversations an AI handled without a human getting involved. It is misleading because it measures whether anyone else got pulled in, not whether the customer got helped. A customer who asked three times, got nowhere, and closed the window in frustration counts as contained, and so does a customer who gave up and phoned instead.
Is CSAT a good way to measure AI agents?
Only as one signal among several. Post-AI CSAT is answered by a small self-selected group, and the customers your AI failed hardest are the ones least likely to stay and answer a survey. It also conflates correctness with tone, so a confidently wrong AI can score well. Compare it against transcript-based measures rather than treating it as evidence on its own.
How do you measure whether an AI agent actually resolved an issue?
Score resolution as a criterion on your QA rubric, applied to the transcript rather than inferred from ticket status. Define, per contact type, what evidence in the conversation counts as resolution: the refund confirmed, the change verified back to the customer, the correct answer acknowledged. Then compare that resolved share against your containment rate. The gap between the two is the size of the problem.
What is silent abandon rate?
It is the share of AI conversations a customer exits without resolution and without escalating: no complaint, no survey, no follow up. These conversations are invisible to every routing-based metric because no human was involved, so they inflate containment. Finding them means reading conversations that ended on a customer turn and judging whether the issue had actually been addressed.
Why does the evaluator need to be independent from the AI it is scoring?
Because a platform that supplies your AI agents and also reports on their performance is producing a self-assessment, and its evaluation criteria will tend to reflect what its system does well. Most QA and CX platforms now sell AI agents of their own, so this is the normal case rather than the exception. Independence is what makes the number usable outside your own team, and verification is what makes independence real: ask whether the evaluator sells the agents it is grading, whether every score links back to the specific evidence in the transcript, and whether it will show you its agreement rate against conversations your reviewers already scored.
Related terms
Measure what your AI agents actually did
See your containment rate next to the resolution quality score behind it, graded against your own rubric by an evaluator that does not sell you AI agents and so has nothing to defend in the result. Every score traces to the evidence in the transcript, and you can check it against conversations your reviewers have already scored.