Call center metrics fall into three groups that behave very differently: efficiency metrics like average handle time and speed of answer, which measure how fast work moves; quality metrics like your QA score, which measure how well it was done; and outcome metrics like CSAT, first contact resolution and repeat contact rate, which measure what the customer actually got. Only outcome metrics describe satisfaction directly, and only quality metrics can predict it in advance, because they are the only ones you can read before the customer replies. The catch is that a quality score built on a 2% manual sample is too small and too late to be a leading indicator, while a score built on 100% of conversations is early enough to act on. Efficiency metrics predict CSAT only in the negative: push them hard enough and resolution falls, repeat contacts rise, and satisfaction follows them down.
In short
- Group your metrics into efficiency, quality and outcome. Most dashboards are almost entirely efficiency, which is why they explain speed but not satisfaction.
- Efficiency metrics do not predict CSAT on their own. They predict it in reverse: squeeze handle time far enough and first contact resolution drops, repeat contacts climb, and CSAT follows.
- Outcome metrics tell you the truth but arrive late, and CSAT only ever arrives from the minority of customers who answer the survey.
- Quality scores are the bridge: they are the only measure you can read before the customer responds, on every conversation, including the silent ones.
- A QA score from a 2% sample cannot be a leading indicator. It is a small, lagging, opinion-weighted number with too little coverage to spot a trend early.
- Repeat contact rate is one of the most underrated numbers in a contact center, because it catches the failures a resolved ticket and a good handle time both hide.
- The right target for any of these metrics depends on your channels, product and customer base. Copying someone else’s benchmark is how teams end up optimizing the wrong number.
Why most call center metrics lists do not help you move CSAT
Open almost any article on call center metrics and you get the same thing: twenty-five definitions in a row, alphabetical or arbitrary, each with a formula and a suggested target. Average handle time, average speed of answer, occupancy, adherence, abandonment rate, service level. Then, somewhere near the end, customer satisfaction, treated as one more item on the list rather than the thing the other twenty-four are supposed to influence.
That format hides the only question a support leader actually has. Not what does this metric mean, but which of these numbers, if it moves this week, tells me something about what my customers will feel next week?
The answer is that most of them do not tell you anything of the sort, and a few of them actively mislead. Speed metrics measure how quickly work leaves the queue. They say nothing about whether the customer’s problem went away. Quality metrics could say exactly that, but in most contact centers they are built on a sample so small that by the time the number moves, the quarter is over.
So the useful way to read a metrics list is not one metric at a time. It is in three groups, with an honest account of how those groups fight each other. That is what this roundup does. Each metric gets a short paragraph and a link to the full definition, because the point here is not to redefine average handle time for the hundredth time. The point is to show you what it does to everything else.
The three groups: efficiency, quality and outcome
Every number on a contact center dashboard belongs to one of three families, and the family it belongs to tells you what it can and cannot be used for.
Efficiency metrics measure the movement of work
Handle time, speed of answer, occupancy, transfer rate, cost per contact. These are operational metrics. They exist to answer capacity and staffing questions: can we absorb the volume, are people waiting, is the team over or under loaded. They are genuinely important, and they are cheap to collect because your helpdesk generates them automatically. They are also the easiest to game, which matters more than most teams admit.
Quality metrics measure how the work was done
Your QA score, internal quality score, compliance pass rate, adherence to the required process. These are the only metrics that describe the content of a conversation rather than its timing or its aftermath. They are also the only ones you control directly, because you write the rubric that produces them.
Outcome metrics measure what the customer got
Satisfaction, effort, loyalty, resolution, repeat contacts. These are the closest thing to truth on the dashboard, and they are the reason the other two families exist. Their weakness is timing and coverage: they arrive after the fact, and the survey-based ones arrive only from the customers who chose to reply.
Almost every dysfunctional dashboard has the same shape. Eighty percent efficiency, a thin slice of quality, and CSAT sitting on its own at the end with nothing connecting it to anything else.
Efficiency metrics: what they are for, and where they mislead
Efficiency metrics are not the enemy. They become one the moment they are treated as performance targets for individual agents rather than as capacity signals for the operation.
Average handle time
The average time an agent spends on a contact, including talk, hold and after-call work. It is the single most useful staffing input you have and the single most abused agent target. Push it down and agents rush, skip verification steps, and close before the issue is fully resolved. Read the full breakdown of what average handle time is and how to calculate it, then treat it as a planning number, not a scorecard line.
First response time
How long the customer waits for a first human reply. It matters, and unlike handle time it is measured from the customer’s side of the glass, which makes it a fairer thing to hold the team to. It still says nothing about whether that first reply was any good. See first response time explained.
Average speed of answer and service level
Queue metrics: how long callers wait before someone picks up, and the share answered within your threshold. Both are staffing diagnostics. Neither tells you anything about the conversation that followed.
Occupancy and adherence
How much of a logged-in agent’s time is spent on contacts, and how closely the team follows the schedule. Useful for workforce planning, dangerous as an individual target, because the fastest way to raise occupancy is to reduce the time agents spend thinking, checking and documenting.
Transfer rate and cost per contact
Transfer rate is the most quality-adjacent efficiency metric you have: a high one usually means routing or enablement is broken, and it tends to move with repeat contacts. Cost per contact is a board-level number that should never be optimized without a resolution metric next to it.
Outcome metrics: the truth, arriving late
Outcome metrics describe the customer’s experience directly. That makes them the most honest numbers on the dashboard and the least useful for steering the week, because they land after the conversation is over.
CSAT
Satisfaction with a specific interaction, usually from a one-question post-contact survey. It is the metric everyone answers to, and it has a structural problem: response rates are low and skewed, so you hear disproportionately from customers who were delighted or furious. See what CSAT is and how it is calculated. Everything in this article is ultimately about predicting this number before it arrives.
First contact resolution
The share of contacts fully resolved without the customer coming back. FCR is the hinge between efficiency and outcome, and it is the metric most directly traded away when handle time is squeezed. Of every number on this page, it is the one that tends to track satisfaction most closely, for the obvious reason that solving the problem the first time is what customers came for. See first contact resolution explained.
Repeat contact rate
The mirror image of FCR, and worth tracking separately because it is harder to fool. A ticket can be marked resolved and still generate a second contact two days later. Repeat contacts catch exactly the failure that a good handle time and a closed ticket conspire to hide.
CES and NPS
Customer effort score asks how hard the customer had to work, which in support contexts often correlates with loyalty more tightly than satisfaction does. Net promoter score measures relationship-level advocacy and moves too slowly to be a contact center operating metric. See CES explained, NPS explained, and the side-by-side on CSAT vs NPS vs CES if you are deciding which to run.
How the three groups trade off against each other
The reason a flat list of metrics is useless is that these numbers are not independent. Move one and you move others, often in the opposite direction to the one you wanted.
The best-documented trade-off is handle time against resolution. When average handle time becomes an individual target, agents optimize for it, because people optimize for what they are measured on. The available savings come from the parts of a conversation that take time and do not shorten the call: verifying details, checking the knowledge base, confirming the customer actually understood. Cut those and handle time falls immediately. First contact resolution falls a little later. Repeat contacts rise after that, which quietly puts the volume back into the queue you were trying to protect, and CSAT drifts down last, by which point the causal chain is a quarter old and hard to argue about.
Occupancy behaves the same way. Raising it looks like efficiency until it turns into fatigue, and tired agents produce worse conversations, more errors and more escalations. Transfer rate is often the first visible symptom.
The trade-off runs in the useful direction too. Raise resolution quality and repeat contacts fall, which reduces total volume, which relieves the queue pressure that made you want to cut handle time in the first place. Teams that get this right often see handle time rise slightly per contact while total handling effort per customer issue falls, which is the outcome you actually wanted and the one a handle time target would have blocked.
Resist the urge to copy a benchmark here. The right handle time for a two-minute password reset and a forty-minute billing dispute are not comparable, and the right target for your operation depends on your channel mix, product complexity and customer base. Set targets from your own baseline and watch what happens to the metrics on either side of them.
Which metrics actually move with satisfaction
If you are choosing what to put in front of the leadership team, the honest ranking looks roughly like this, in descending order of how tightly the metric relates to what customers report.
- Resolution. Whether the issue was actually solved, and whether the customer had to come back, is the closest operational proxy for satisfaction there is. This is the pair to watch: FCR up, repeat contacts down.
- Effort. How many steps, channels, repetitions and transfers the customer went through. Transfer rate and reopen rate are the operational fingerprints of high effort, and they move in the same direction as dissatisfaction.
- Quality score, when it has real coverage. A well-built scorecard measures the behaviors that cause resolution and low effort: correct information, clear explanation, right process followed, expectations set. On full coverage it becomes a genuine leading indicator.
- Waiting. Response and answer times matter, but with a ceiling. Fast and wrong does not beat slow and right, and beyond a reasonable threshold, further speed buys very little satisfaction.
- Handle time and occupancy. No direct relationship with satisfaction at all in the positive direction. Only a negative one, when they are pushed hard enough to damage resolution.
Notice that the top three are all things a quality program measures or influences, and the bottom two are the ones that dominate most dashboards. That inversion is the whole problem.
The eight-metric call center dashboard worth keeping
You do not need twenty-five metrics. You need enough coverage across the three groups to see a trade-off happening while you can still do something about it. Eight is usually enough, and every one of them has a trap, so the trap is listed next to it.
| Metric | Group | What it tells you | The trap |
|---|---|---|---|
| First contact resolution | Outcome | Whether the issue was actually solved on the first try | Easy to overstate if it is measured from ticket status rather than customer behavior |
| Repeat contact rate | Outcome | The failures that a closed ticket hides | Needs a sensible time window, or you will count unrelated new issues |
| CSAT | Outcome | What the customers who replied felt about the interaction | Low and skewed response rates. Never read a small sample as a trend |
| Customer effort | Outcome | How hard the customer had to work to get resolved | Survey fatigue if you run it alongside CSAT on every contact |
| QA score | Quality | Whether the behaviors that cause resolution were present | Meaningless as a leading indicator on a 2% sample. Needs full coverage |
| Critical error or compliance rate | Quality | The failures that matter regardless of the overall score | Burying them inside an average score, where a 92% hides a legal problem |
| Average handle time | Efficiency | Capacity, staffing and where effort is concentrated | The moment it becomes an individual target, resolution starts paying for it |
| First response time | Efficiency | How long customers wait before a human engages | Answering fast with a holding message that resolves nothing |
How QA scores connect efficiency to outcome
The gap in every dashboard described so far is causal. Efficiency metrics tell you how fast. Outcome metrics tell you how it went. Neither tells you why, and neither arrives in time to change anything.
Quality scoring is the only measure that sits in the middle. A conversation that got a low score for missing information, an unclear explanation or a skipped verification step is a conversation that is likely to generate a repeat contact and a poor survey response, and you can see that on the day it happens rather than a week later. That is the definition of a leading indicator: it moves first, and the outcome follows.
For that to work, the scorecard has to measure causes rather than cosmetics. A rubric full of tone and formatting criteria will produce a stable score that never predicts anything. A rubric built on resolution behaviors, accuracy, process adherence and expectation setting will move before CSAT does. If you are building or rebuilding one, start with the QA scorecard framework and the wider call center quality assurance guide, and read what the internal quality score actually tells you before you decide how to aggregate it.
The other requirement is that the score has to be comparable across agents and across weeks. If two reviewers grade the same conversation differently, the metric is measuring reviewer opinion as much as agent behavior, and it will not correlate with anything. Calibration is not a nice-to-have here. It is what makes the number a metric rather than an anecdote.
Why a 2% QA sample cannot be a leading indicator
Here is the structural problem with manual quality assurance as a metric. A reviewer scoring a handful of conversations per agent per month is producing a number from a sample that is tiny, non-random and slow.
Tiny means the number is dominated by noise. A single unusually bad conversation can swing an agent’s monthly score by several points, so you cannot tell a real decline from a bad Tuesday. Non-random means the sample is usually whatever was convenient or whatever got escalated, which skews it in ways nobody documents. Slow means the score arrives at the end of the review cycle, weeks after the conversations it describes, by which time the customers involved have already answered the survey, already contacted you again, or already left.
None of that is a criticism of the reviewers. It is arithmetic. You cannot build an early warning system out of 2% of the evidence.
A quality score built on 100% coverage changes the arithmetic rather than the concept. Every conversation is scored against the same rubric, so a week-over-week move is a real signal, not sampling noise. Every agent, queue, channel and issue type has enough volume to trend independently, so you can see that the decline is confined to billing contacts on chat before it shows up in the aggregate. And the score exists within hours of the conversation, which is what makes it early enough to act on. This is what automated QA is actually for: not saving reviewer time, though it does that, but producing a quality metric with enough coverage and speed to behave like a leading indicator.
A leading indicator is only worth acting on if you can check it. Kaizo links each score back to the evidence in the transcript, so a trend can be traced to the specific behaviors that caused it, and so the grading itself can be tested against conversations your own reviewers have already scored. Kaizo scores against the scorecard your team defines, natively inside Zendesk and Salesforce, and it does not sell AI support agents, so AI-handled conversations are measured on the same standard by a grader with no product of its own in the result. Full coverage is the precondition for both: it is what takes the selection bias out of the trend and gives you the whole population to check the grading against. At UiPath, that meant 100% of QA automated, 200% ROI, and an 8% lift in quality score.
Frequently asked questions
What are the most important call center metrics?
The eight worth keeping are first contact resolution, repeat contact rate, CSAT, customer effort, QA score, critical error rate, average handle time and first response time. Together they cover all three metric groups: efficiency, quality and outcome. Any dashboard made mostly of efficiency metrics will tell you how fast the team is working and nothing about whether customers are being helped.
Which call center metrics predict CSAT?
Resolution metrics come closest, because solving the issue on the first contact is what customers came for. Effort signals like transfers and reopens move with dissatisfaction. A quality score built on full conversation coverage is the only genuine leading indicator, since it can be read before the customer replies. Efficiency metrics like handle time do not predict CSAT positively at all, only negatively when they are pushed hard enough to damage resolution.
What is the difference between call center metrics and KPIs?
A metric is anything you measure. A KPI is the small subset you have decided the operation will be judged on. The practical difference is behavioral: people optimize for KPIs, so making average handle time a KPI causes agents to shorten conversations whether or not that helps the customer. Choose KPIs from the outcome and quality groups, and keep efficiency metrics as diagnostics you monitor rather than targets you enforce.
Does lowering average handle time hurt customer satisfaction?
Lowering it by removing genuine waste, such as bad tooling or slow systems, does not. Lowering it by targeting agents does, because the time comes out of verification, knowledge checks and confirming understanding. The usual sequence is handle time down, first contact resolution down, repeat contacts up, and CSAT down last. Watch resolution and repeat contacts alongside any handle time initiative so you can see which of the two is happening.
How many metrics should a contact center track?
Around eight on the working dashboard, spread across efficiency, quality and outcome, with more available for diagnosis when something moves. Long metric lists do not improve decisions, they dilute attention, and they make it easy to miss the trade-offs happening between numbers that sit on different tabs. The test of a dashboard is not how much it measures but whether it shows you a problem early enough to fix.
What is a good target for these call center metrics?
There is no universal answer, and any article giving you one is guessing. The right handle time, resolution rate and quality score depend on your channel mix, product complexity, customer base and how you define a resolved contact. Set targets from your own historical baseline, then move them deliberately and watch what happens to the metrics on either side. A borrowed benchmark is how teams end up optimizing a number that was never their constraint.
Related terms
See which of your metrics is about to move
Kaizo scores 100% of your conversations against your own scorecard inside Zendesk and Salesforce, so the metric arrives early enough to act on instead of confirming a CSAT drop you already knew about, and links every score to the evidence in the transcript so you can prove the number is right before you report it.