A bad conversation costs you across four lines: the rework of handling it again, the money given away to settle it, the escalation labour it consumes, and the share of revenue that leaves. The first three are calculable from data already in your helpdesk, and together they are usually enough to fund a quality programme without touching the fourth. The churn line is the largest and the least defensible, so quote it separately or not at all, because an inflated total is easier for a finance team to dismiss than a small one is.
In short
- Aggregate industry figures about the cost of poor service cannot survive a budget meeting, because nobody accepts that a share of them is theirs.
- Four cost lines: rework, goodwill and refunds, escalation labour, and lost revenue. Only the first three are defensible from your own data.
- Rework is the strongest line and the easiest to compute. Repeat contact rate multiplied by fully loaded cost per contact.
- Do not add rework and churn for the same conversation without saying so. Double counting is the fastest way to lose the room.
- Effort predicts disloyalty better than satisfaction does, which is why a resolved-but-exhausting conversation still carries cost.
- Present a range with the method attached, not a single number. A defensible small figure beats an impressive one you cannot source.
Why does the trillion-dollar statistic not help you?
Search for the cost of poor customer service and you will find very large aggregate numbers, usually in the trillions, usually attributed to a survey of consumers across an entire economy. They are probably directionally true. They are also useless to you, for three reasons worth being explicit about.
Nobody accepts that a share of an economy-wide figure is theirs. The moment you put a global number on a slide, the question becomes what portion applies here, and you do not have an answer. The conversation moves to the credibility of the statistic instead of the decision you came to get.
The unit is wrong. A figure denominated in economies cannot be divided down to a conversation, which is the unit your quality programme actually operates on. You need a cost per bad conversation, because that is the thing you can change the count of.
It has no method attached. Anyone can dismiss a number they cannot reproduce. A smaller figure with the arithmetic shown beside it survives scrutiny that a bigger one does not, and surviving scrutiny is the entire job here.
So build your own. It takes an afternoon, most of the inputs are already in your helpdesk, and the result belongs to you, which is what makes it usable in the reporting you take to executives.
A defensible small number beats an impressive unsourced one. You are not trying to win an argument about the industry, you are trying to price a decision about your own operation.
What counts as a bad conversation, before you price anything?
Define this first or the arithmetic is meaningless. A bad conversation is one that produced the wrong outcome, or produced the right outcome at unreasonable cost to the customer. Those are two different populations and they cost different amounts.
The second half of that definition is the one teams miss. A conversation can end correctly, with the customer satisfied enough to leave a decent rating, having taken four contacts and two weeks. That conversation cost you three unnecessary handles and it did real damage. The Harvard Business Review argument in Stop Trying to Delight Your Customers is the useful frame here: the effort a customer has to expend predicts disloyalty more reliably than their satisfaction predicts loyalty. High-effort resolutions carry cost that a satisfaction score will not show you.
Practically, take these four as your population, because each one is detectable without a survey:
- Wrong or incomplete resolution. Detected by repeat contact on the same issue.
- Avoidable escalation. Went to a senior or a specialist for something that should not have needed one.
- High effort. More contacts, channels or elapsed time than the issue warranted.
- Avoidable goodwill. Money given away to settle something that should not have needed settling.
Note what is not on the list. A low satisfaction score is not on it, because dissatisfaction arrives from a self-selected minority and misses most of the population you are trying to count. Use it as a cross-check, not as the definition. This is the same reason a quality programme needs an internal quality score alongside the customer one rather than instead of it.
How do you calculate the four cost lines?
Work through them in this order, which is also the order of how defensible they are. Stop wherever your data stops, and say where you stopped.
| Cost line | How to calculate it | Where the data is | How defensible |
|---|---|---|---|
| Rework | Extra contacts on the same issue, multiplied by fully loaded cost per contact | Helpdesk: repeat contacts within seven days, linked or same-requester same-topic | Strong. Both inputs are yours and neither is an estimate |
| Goodwill and refunds | Value of credits, refunds and discounts issued to settle a service failure, minus what policy would have given anyway | Billing or refund log, filtered to service-recovery reasons | Strong, though separating policy from goodwill takes judgement |
| Escalation labour | Escalated volume multiplied by the loaded cost of the senior time it consumes, which is typically several times a first-line contact | Helpdesk escalation flags plus your own salary bands | Moderate. The multiple needs stating as an assumption |
| Lost revenue | Churn or reduced spend attributable to the service failure | Requires joining support data to revenue data, which most teams cannot do cleanly | Weak. Real, large, and the least provable line you have |
What does the arithmetic look like in practice?
Here is the shape of the calculation. Substitute your own figures; the point is the structure, not the illustration.
Step 1. Get a fully loaded cost per contact
Take total support cost for a period, including salaries, employer costs, tooling, management overhead and premises, and divide by contacts handled in that period. Use the loaded figure, not the hourly wage. Finance will use the loaded one and you do not want to be corrected on your own arithmetic in the meeting.
Step 2. Count the rework
Pull repeat contacts within seven days on the same issue for a full month. Multiply the count by your cost per contact. This is your rework line, and for most teams it is the single largest defensible number available, often several times what anybody in the room expected.
Step 3. Add goodwill and escalation
Sum service-recovery credits for the same month. Then count escalations and multiply by your best estimate of the senior time each consumes. State the multiple you used out loud, because an assumption you volunteer is treated as rigour and the same assumption discovered later is treated as a flaw. Worth separating the avoidable escalations from the correct ones while you are in there, which is what escalation handling quality measures; only the avoidable ones belong in a cost of failure.
Step 4. Divide, then sanity-check
Divide the total by the number of bad conversations you identified to get a cost per bad conversation. Then check it against something you already know. If the figure implies your worst 5% of conversations cost more than your entire support budget, something is double counted. Go and find it before somebody else does.
The most common error is counting the same conversation in two lines: rework as extra handles, then again as churn risk. If you present both, say explicitly that they are alternative framings of the same event rather than two costs that add. One discovered double count discredits every other number on the slide.
Should you put a number on churn at all?
Quote it separately, clearly labelled as an estimate, or leave it out. Never let it sit inside a total that also contains your defensible lines, because the weakest line in a sum sets the credibility of the whole sum.
The problem is attribution rather than arithmetic. A customer who left after a bad conversation may have left because of it, or because of price, or because they no longer needed the product, and the bad conversation is simply the last thing you have a record of. Assigning the full lifetime value of that customer to one interaction is a claim you cannot support, and any finance team will spot it.
Two honest ways to handle it:
- Present it as a sensitivity, not a figure. If one in twenty of these customers leaves because of this, the cost is X. That framing hands the assumption to the audience, which is both more honest and more persuasive than asserting it yourself.
- Use retention as a direction rather than a value. Show that customers with a bad conversation in the period renewed at a lower rate than those without, and let the correlation stand as a correlation. It is a weaker claim and a much harder one to dismiss.
Where a customer-facing figure does exist, use the one that has a name attached. Kaizo’s own published customer result is a 13% annual increase in CSAT at Procede Software, which is attributable and checkable, and one attributable number is worth more in this conversation than a modelled total. The general case for the link between quality work and satisfaction is covered in how to measure conversation quality.
What do you actually do with the number?
The cost figure is not the deliverable. It is the denominator for a decision, and there are only three decisions it supports.
Fund the programme. Compare the annual rework line to the cost of the review capacity or tooling that would reduce it. This is the strongest version of the argument because both sides are cost, so nobody has to believe a revenue projection. The Auto QA ROI case is built on exactly this comparison.
Prioritise what to fix. Cost per bad conversation varies enormously by contact reason, and the expensive ones are rarely the frequent ones. Segment the rework line by issue type and you will usually find a small number of reasons generating most of it, which is a process problem rather than an agent problem almost every time.
Set a target somebody will believe. A reduction in repeat contact rate is a better target than a rise in a quality score, because it is denominated in a unit the business already budgets. Quality score movement is the mechanism; repeat contact reduction is the result.
One practical constraint. All of this depends on being able to identify the bad conversations, and a 3% review sample cannot find them reliably at any useful granularity: segment a small sample by issue type and you are left with single-figure counts per segment. Kaizo’s Auto QA scores the full population against your own rubric, which is what makes the segmentation above possible at all, and it keeps the reasoning attached to each conversation so a cost line can be traced back to the specific conversations behind it rather than asserted. That is the practical meaning of 100% coverage revealing trends that 3% sampling never could: not more scores, but a number you can take into a budget meeting and defend line by line.
If the honest answer this quarter is that quality is flat, that is still a finding worth presenting, and the pillar on proving the value of support quality covers how to present it without either overclaiming or apologising.
Frequently asked questions
How do you calculate the cost of poor customer service?
Build it from four lines rather than borrowing an industry statistic. Rework is repeat contacts multiplied by your fully loaded cost per contact. Goodwill is service-recovery credits and refunds. Escalation labour is escalated volume multiplied by the senior time it consumes. Lost revenue is real but weakly attributable, so quote it separately as a sensitivity or leave it out. The first three come from data already in your helpdesk.
What is a fully loaded cost per contact?
Total support cost for a period divided by contacts handled in that period, where total cost includes salaries, employer costs, tooling, management overhead and premises rather than just hourly wages. Use the loaded figure because that is the one a finance team will use, and being corrected on your own arithmetic mid-meeting costs you more than the difference between the two numbers.
Should you include customer churn in the cost of a bad conversation?
Only as a clearly labelled estimate, and never inside a total that contains your defensible lines, because the weakest line in a sum sets the credibility of the whole sum. The problem is attribution: a customer who left after a bad conversation may have left for price or need, with the conversation simply being the last recorded event. Present it as a sensitivity, or show renewal rates with and without a bad conversation and let the correlation stand.
What is the most defensible cost of a service failure?
Rework, meaning the extra contacts generated by a conversation that did not resolve the issue. Both inputs are yours, neither is an estimate, and the resulting figure is usually larger than anyone in the room expects. It is also the line most directly reduced by a quality programme, which makes it the natural basis for a funding argument where both sides of the comparison are cost.
How does a quality programme reduce these costs?
By identifying which contact reasons generate the rework, which is almost always a small number of process or knowledge problems rather than a broad agent-performance issue. The constraint is granularity: a 3% review sample cannot be segmented by issue type without collapsing to single-figure counts per segment, so the diagnosis needs far wider coverage than a routine sample provides.
Related terms
Put a defensible number on the conversations that cost you most
Bring a month of conversations and your repeat contact data. We will show you which contact reasons are generating the rework, and what the quality findings behind them look like line by line.