Escalation handling quality is how well a conversation is managed after it leaves the person who first received it: the handoff, the recovery of a customer who is already unhappy, the accuracy of what gets promised, and whether the escalation needed to happen at all. It cannot be judged with a standard QA scorecard, because handle time, template adherence and raw satisfaction scores all mislead on a conversation somebody inherited already broken. Score escalations on their own rubric, and attribute the failure to the point where it was created rather than to the person who cleaned it up.
In short
- An escalated conversation is a different unit of work. Running your standard scorecard over it penalises the person who inherited the problem.
- Three criteria are actively wrong on an escalation: handle time, template or script adherence, and unadjusted satisfaction scores.
- The handoff is its own reviewable event with two owners. Most of the damage is done in the transfer, not in the resolution.
- Attribute the escalation to where it was created. Charging the escalation handler for an upstream failure is the most common unfairness in QA.
- The finding that saves money is not how well the escalation was handled. It is whether it should have existed, and that only shows up when you review the originating contact too.
- One escalation is a coaching conversation. A hundred escalations is a process map, and escalations are too rare for a small sample to show you the map.
What counts as an escalation, and why your definition decides your data
Before you can score escalations you have to agree what one is, and most support teams quietly run three incompatible definitions at once. The customer asks for a manager. An agent reassigns a ticket to a specialist queue. A rule fires because a service level was about to be breached. All three get called escalations, they have nothing in common operationally, and averaging them produces a number nobody can act on.
Pick a definition, write it down, and make sure it is recorded on the ticket rather than inferred later from a tag somebody added from memory. The useful test is ownership: an escalation is any conversation whose ownership changes because the current owner cannot resolve it within their authority, information or skill. That excludes a routine reassignment for capacity and includes the case where nobody moved the ticket but a supervisor had to authorise the outcome.
The type matters because it decides who the quality question belongs to.
| Type | What triggers it | Who owns the quality | What a spike in it tells you |
|---|---|---|---|
| Hierarchical | The customer or the agent pulls in a supervisor or manager | Split: the originating agent and the supervisor | Agents lack authority, or lack confidence in the authority they have |
| Functional | The issue needs a specialist team, billing, engineering, trust and safety | The receiving team, plus whoever wrote the handoff | Frontline scope is too narrow, or routing rules are wrong |
| Priority or service level | A rule fires on age, breach risk or account tier | The queue and the staffing model, not an individual | A capacity or backlog problem wearing an escalation costume |
| Customer-demanded | The customer explicitly asks for someone else | The originating contact, almost always | Trust broke in the first exchange. Review that exchange, not the escalation |
| External | It leaves the company: a regulator, a public review, a legal threat | The programme, and it should trigger a formal review | Something got through several internal checks first |
Why your normal scorecard is wrong on an escalated ticket
Most teams score escalations with the scorecard they already have, because it is the scorecard they already have. It produces numbers, so nobody notices that several of the criteria have stopped measuring anything.
The problem is not that escalations are harder. It is that the standard criteria were designed around assumptions that an escalation breaks: that the agent is meeting the customer for the first time, that the customer is neutral, that the fastest resolution is the best one, and that a good answer exists inside the standard playbook. On an escalated ticket, none of those hold. Before you change anything, it is worth being clear about what your QA scorecard is actually claiming to measure.
Criterion by criterion, here is what breaks.
| Criterion | What it measures on a routine ticket | Why it breaks on an escalation | Use instead |
|---|---|---|---|
| Handle time or resolution time | Efficiency | The work is deliberately slower. The escalation exists because the fast path failed, and rewarding speed here rewards fobbing the customer off | Time to first meaningful update, and whether promised timings were met |
| Template, macro or script adherence | Consistency | The standard response is what the customer already rejected. Reusing it reads as a wall | Whether the response addressed the specific objection the customer raised |
| Opening and greeting | Professionalism | The customer has already told the story once. A fresh greeting that ignores that history is the complaint, not the courtesy | Context acknowledgement: did the reply prove the history had been read |
| Satisfaction score on the ticket | Customer outcome | The customer arrived unhappy. The score partly grades the previous contact, so the agent inherits somebody else’s rating | Sentiment movement across the escalated conversation, plus repeat contact within 14 days |
| First contact resolution | Effectiveness | It is definitionally not a first contact, so the criterion is either always failed or quietly excluded | Durable resolution: did the same issue come back |
| Knowledge or policy accuracy | Correctness | This one survives intact and matters more here than anywhere else | Keep it, and weight it higher than you do on routine work |
What to evaluate on an escalated conversation instead
Six criteria carry almost all the signal on an escalated ticket. Each one needs evidence a second reviewer could point at in the transcript, otherwise you have written adjectives rather than a rubric.
- Context absorption. Did the handler read the thread before replying? The evidence is specific: they referenced something the customer said earlier without being told it again. The failure evidence is even easier to spot, and it is the next item.
- The re-ask. Did the handler ask for information that is already in the thread? Order number, account email, what went wrong. This is binary, it takes a reviewer ten seconds, and agents never dispute it because the transcript settles it.
- Acknowledgement without blame-shifting. The customer needs the failure named. What fails this criterion is not a missing apology, it is deflection: blaming the previous agent, the system, the policy or the customer. Naming a colleague as the cause in front of a customer should be an auto-fail.
- Commitment accuracy. Did the handler promise something the company will actually do, by a date it will actually happen? This is the criterion that produces the next escalation when it fails, so it earns heavy weight.
- Use of the authority they were escalated for. An escalation that ends in the same answer, delivered by someone more senior, has cost you two agents and bought the customer nothing. Either the handler used discretion the first agent did not have, or the escalation path is decorative.
- Loop closure. Was the customer told the outcome, and was the originating agent told what the resolution was? The second half is skipped almost everywhere, and skipping it guarantees the same agent creates the same escalation next week.
Calibrate before you roll this out
Escalations are exactly the conversations reviewers disagree about, because they are emotive and because the right answer often depends on context the transcript only half contains. Run one calibration session on three escalated tickets before the rubric goes live. If your reviewers cannot agree on whether the handler used real discretion, agents certainly will not accept the score. The format is the same as any other calibration session you run, with escalations as the only sample.
Score the handoff, not just the resolution
Ask a support team where escalations go wrong and they will describe the resolution. Read a hundred escalated threads and you will find that the damage was usually done in the ninety seconds of the transfer. The Harvard Business Review case for reducing customer effort rather than trying to delight applies with unusual force here, because an escalation is already a second attempt and a sloppy handoff turns it into a third.
The handoff is a reviewable event in its own right, and it has two owners. Score it as its own block, with the sending side attributed to the originating agent and the receiving side to the handler. Keeping them separate is the whole point, because it is the only way the score lands on the right person.
Sending side
- The reason for the escalation is recorded in the ticket, in words, not just as a tag.
- What has already been tried is summarised, so the receiver does not repeat it.
- The customer was told the handoff was happening, why, and roughly when to expect a reply. A silent transfer is a failure even if the resolution is perfect.
- Nothing was promised on behalf of the receiving team that the receiving team has not agreed to.
Receiving side
- The thread was read before the first reply went out. The re-ask test settles this.
- The wait was acknowledged. The customer has been sitting in a queue twice now.
- The customer was not asked to restate the problem as a way of opening the conversation.
- If the ticket was sent back or moved on again, the reason was recorded with the same discipline as the first handoff.
Two numbers fall straight out of scoring the handoff this way, and both are more useful than escalation rate on its own. Re-ask rate is the share of escalated conversations where the receiving handler asked for information already in the thread. Bounce rate is the share of escalations that get sent back or forward again before resolution, which is usually a routing defect rather than a person’s failing. Neither requires a survey, and both are countable from the conversation record you already hold. If you are building a broader measurement set, this sits alongside the other ways to measure conversation quality.
The attribution problem: who actually caused this escalation?
This is the fairness question that decides whether you can run a QA programme agents believe in, and almost no escalation guidance addresses it.
Most escalations are created upstream. By the time a conversation reaches a specialist or a supervisor, the decision that caused it was taken earlier: by a first contact that missed something, by a policy nobody on the frontline can bend, by a product or system failure, or by an expectation somebody else set. Scoring the escalation handler for the existence of the escalation is the single most common unfairness in support QA, and agents spot it immediately.
The fix is a rule and a habit.
The rule. The escalation handler is scored on the escalation only. Whether the escalation should have happened is a separate finding, attributed to a separate owner, recorded on a separate record. Never let one score carry both judgements, because the moment it does, the people handling your hardest work score worst.
The habit. Every escalation review is two reviews. Open the escalated conversation, and open the contact that created it. Then tag the origin against one of four buckets:
- Handling. The first agent had the information and the authority and did not use them. This one belongs to an individual and it is coachable.
- Policy. The agent did exactly what the policy said and the policy produced an angry customer. This belongs to whoever owns the policy, and coaching the agent for it is worse than useless.
- System or product. Something was broken, slow or wrong. This belongs on an engineering or operations backlog, not in a one to one.
- Expectation. A promise made elsewhere, by pricing, marketing, a previous ticket or an outage notice, did not survive contact with reality.
Two guardrails keep the tagging honest. The origin tag must not be set by the escalation handler, who has an obvious interest in it, and it must not be set by the originating agent either. It is a reviewer’s call. And when the origin bucket is policy, system or expectation, no agent-level score changes at all, which is the part you should say out loud in the launch. Once the tags accumulate, they are the input to real root cause analysis rather than a monthly guess.
Should this escalation have happened at all?
Handling quality is worth improving. Escalation volume is worth more, because an escalation you prevent costs nothing to handle well, and because every avoidable one is a piece of customer dissatisfaction you manufactured yourself.
So attach a verdict to every escalation review, separate from the score. Three options, and the wording matters because the whole point is to make the middle bucket visible.
- Avoidable. The first agent had the authority, the information and the tools to resolve it. Nothing structural stopped them. This is a coaching finding, and it should be a small share of the total in a healthy programme.
- Structural. Nobody at the first tier could have resolved it, because of a permission they do not hold, a system they cannot reach, or a policy they cannot bend. This is the important bucket. It is not an agent problem, it is a design problem, and every ticket in it is a candidate for pushing authority down a level.
- Legitimate. The escalation is the process working. Specialist knowledge was genuinely required. Nothing to fix, and it should not be counted as a defect in anyone’s numbers.
The structural bucket is where the money is, and it is invisible to every escalation programme that only scores the escalation itself. When forty of last quarter’s escalations all needed the same refund approval that sits one level above the frontline, you are not looking at a training gap. You are looking at an approval threshold set in the wrong place, and you can price the fix by counting the tickets.
A note on cost, deliberately without a number. It is tempting to multiply escalations by an average handling cost and produce a headline. Do not, unless you can defend the inputs: a real escalation costs at least two people’s time, usually a supervisor’s attention, the elapsed wait the customer experiences and the increased chance of churn that follows a bad recovery, and none of those are constant across issue types. Count the tickets in the structural bucket and describe the specific fix each cluster needs. That argument survives scrutiny. An invented cost per escalation does not.
What escalation patterns tell you in aggregate
One escalation is a coaching conversation. A hundred escalations, read together, is a map of where your service design fails, and it says things no individual review can.
Five reads worth running monthly
- Concentration by topic. Group escalations by the thing the customer wanted, not by the tag the agent applied. A handful of topics almost always produce most of the volume, and they are usually policy edges rather than complex problems.
- Origin mix over time. Track the four attribution buckets as shares. If the handling share is falling while the structural share is flat, your coaching is working and your process is not.
- Repeat escalations. Same customer, same underlying issue, escalated twice. This is the strongest single indicator that the first resolution was cosmetic, and it is worth reviewing every instance by hand.
- Bounce and re-ask rates by receiving team. These expose the teams that are receiving work they are not set up to take, which is a routing conversation rather than a quality one.
- Time of day and staffing. Escalations that cluster at shift boundaries or coverage gaps are a rota finding, and no amount of coaching will move them.
The denominator problem
Comparing escalation rate between agents is the fastest way to lose the room, because the agent on the hardest queue escalates most. Escalation rate belongs with the rest of your customer service metrics, and it is only comparable within the same queue, the same channel and roughly the same issue mix, and even then it belongs in a conversation rather than on a leaderboard. A high escalation rate on complex work can be correct behaviour. A low one can mean an agent is refusing to escalate things they should, which is the more expensive failure and it will show up as repeat contacts instead.
Why sampling struggles here
Escalations are rare by design, which is exactly what makes them hard to review well. A small random sample of a month’s conversations, drawn the way most QA sampling works, will contain almost no escalations, and the ones it does contain will not be representative of the topics that drive them. Most teams respond by hand-picking escalations to review, which introduces a different bias: you find what you went looking for. Reviewing every escalation instead of a sample is what turns this from anecdote into a pattern, and it is the same argument behind 100% coverage revealing trends that 3% sampling never could.
That is where automated scoring earns its place, once the rubric above exists and reviewers agree on it. Kaizo’s Auto QA applies your own escalation rubric to every escalated conversation rather than to the handful that get sampled, so the topic concentration and the origin mix are counted rather than estimated. The part that matters more than the coverage is that each deduction stays attached to the moment in the conversation that caused it, so when an agent disputes a score on a ticket they inherited, the reviewer can look at the exchange instead of arguing from memory. What comes out of the aggregate then feeds the coaching conversations that are worth having, and quietly retires the ones that are not.
Frequently asked questions
What is escalation handling?
Escalation handling is the work of resolving a conversation after ownership has moved from the person who first received it, usually because the issue needed more authority, more specialist knowledge or a decision the first agent could not make. It covers the transfer itself, the recovery of a customer who is already frustrated, the accuracy of what gets committed, and closing the loop with both the customer and the originating agent. In quality assurance it is treated as its own category of work because the criteria that judge a routine ticket do not apply cleanly to an inherited one.
How do you handle an escalation well?
Read the full thread before you reply, so the customer never has to repeat themselves. Name what went wrong without blaming a colleague, a policy or the customer. Use whatever discretion you were escalated for, because an escalation that produces the same answer from someone more senior has helped nobody. Commit only to things that will actually happen, on dates that are real. Then close the loop twice: tell the customer the outcome, and tell the agent who escalated it what the resolution was, so the same escalation does not happen again next week.
What is an example of a customer escalation?
A customer is refused a refund by an agent following policy, replies that this is unacceptable and asks for a manager. The ticket moves to a supervisor who has authority to make an exception. That is a hierarchical, customer-demanded escalation, and it is the most common shape. A different example: a billing discrepancy that the frontline cannot see in their tools gets routed to a finance specialist. That one is functional, nobody did anything wrong, and it should not count against the first agent.
Should escalated tickets be scored with the same QA scorecard?
No. At least four standard criteria mislead on an escalation. Handle time punishes the slower work the escalation exists to allow, template adherence rewards reusing the response the customer already rejected, greeting and opening criteria ignore that the story has already been told once, and satisfaction scores partly grade the previous contact rather than the current one. Keep accuracy and policy compliance, drop or replace the rest, and add criteria for the handoff, commitment accuracy and use of authority.
How do you measure escalation handling quality?
Use four measures together and none of them alone. A rubric score on the escalated conversation covering context absorption, acknowledgement, commitment accuracy, use of authority and loop closure. Re-ask rate, the share of escalations where the receiving handler asked for something already in the thread. Bounce rate, the share sent back or moved on again before resolution. And repeat escalation rate, the share where the same customer escalates the same issue twice, which is the clearest sign the first resolution was cosmetic.
What is a good escalation rate?
There is no benchmark worth quoting, because the number depends entirely on how you define an escalation, what your product is, and how much authority sits at the first tier. A team that counts every specialist routing as an escalation will report several times the rate of a team that counts only supervisor involvement. Compare your rate against your own trend and against comparable queues internally, and pay more attention to the mix than the level. A stable rate with a falling share of avoidable escalations is a programme that is working.
Who should be scored when a ticket is escalated?
Both people, on different things. The originating agent is scored on the handoff they sent and on whether their handling caused the escalation. The receiving handler is scored on the escalation itself and never on the fact that it exists. Whether the escalation should have happened is recorded as a separate finding against one of four origins: handling, policy, system, or an expectation set elsewhere. When the origin is policy, system or expectation, no agent-level score should move at all.
Related terms
Find out how many of your escalations were avoidable
Bring a month of escalated tickets and the scorecard you run on them today. We will show you where those escalations were created, how many sit in the structural bucket you can design away, and what changes when the rubric is applied to all of them rather than the few that get sampled.