The ROI of auto QA is the sum of four value pools measured against the cost of running it: reviewer time reclaimed, the value of finding failures in the conversations nobody was reviewing, quality improvements that flow through to customer outcomes, and risk avoided when a compliance or misinformation problem is caught early. Most business cases only count the first pool, which is the smallest and the easiest to argue with. A credible case builds each pool from your own operational numbers, states the confidence level of each one honestly, and subtracts the real costs of implementation, calibration, and change management. This page gives you the arithmetic for all four so you can build a model your finance team can audit rather than a vendor percentage you have to defend.
In short
- Build the model from your own numbers. A vendor ROI percentage is not evidence, and a finance team will treat it as marketing.
- There are four value pools: reviewer time, coverage, quality outcomes, and risk avoided. Counting only reviewer time understates the case badly.
- Reviewer time reclaimed is the easiest pool to calculate and usually the smallest. Do not build the whole case on it.
- Coverage value is the largest pool in practice and the hardest to quantify. Measure it by sampling the conversations you were never reading and counting what you find.
- Quality improvements are only credible with a before-and-after baseline, so capture your baseline before anything changes.
- Subtract the honest costs: implementation, calibration effort, and change management. A case with no costs in it reads as a sales pitch.
- Do not promise headcount savings. Reclaimed reviewer time is almost always redeployed into coaching, not removed from the budget.
- The model only holds if the findings are verifiable. Every score should trace to the evidence in the transcript so a finance reviewer can audit a finding instead of believing it.
- Check who owns the grader before you build on its numbers. Most QA and CX platforms now sell their own AI agents, so their scoring engine is assessing a product their own company built.
Why most auto QA business cases are too small to win
The typical business case for auto QA is one line of arithmetic: reviewers spend hours grading conversations, automation removes most of those hours, multiply by an hourly cost and you have a saving. It is easy to calculate and easy for a CFO to shoot down. The saving looks modest, it is obviously soft cost rather than cash, and the first question back is whether you are actually going to remove anyone from the payroll.
The arithmetic is not wrong. It just accounts for the least interesting fraction of the value. Automating quality assurance does not primarily make an existing process cheaper. It changes what the process can see. A team reviewing a small sample manually and a team scoring every conversation automatically are not doing the same job at different prices. They are doing different jobs. A case that stops at reviewer hours argues for a cheaper version of the current process. One that includes coverage, outcomes, and risk argues for a capability the organization does not have.
Build it, do not borrow it
Every figure should come from your own operation or be visibly labeled as an assumption you are asking finance to accept. Industry averages and vendor ROI percentages weaken the case, because the first thing a skeptical reviewer questions is where the number came from. If the answer is “a vendor published it,” the whole model inherits that doubt.
Pool 1: reviewer time reclaimed
Start here because it is the pool everyone expects, but keep it short. Take the evaluations your team completes in a typical month. Multiply by the average minutes a reviewer spends per evaluation, including finding and opening the conversation, not just scoring it. Divide by sixty for hours. Multiply by the fully loaded hourly cost of a reviewer, which finance can give you and which is higher than salary divided by hours because it includes employer costs, benefits, and overhead.
- Evaluations per month multiplied by average minutes per evaluation, divided by 60, multiplied by loaded hourly reviewer cost.
- Then apply the share of that grading work automation actually absorbs. It is not all of it. Human reviewers still calibrate, handle disputes, and review what the system flags as ambiguous.
Be conservative on that share, and present this pool as a soft saving rather than a budget line, for reasons the counterweights section returns to.
If the manual versus automated cost comparison is the specific question you are answering, the detailed side by side is covered in manual QA versus auto QA. For a business case, treat this pool as the opening line rather than the argument.
Pool 2: coverage value, or what you find in the conversations nobody read
This is the pool that makes the case, and almost nobody quantifies it. A manual program reviews a sample, and the remaining conversations have never been looked at by anyone. Not lightly reviewed, not spot checked. Never read. Reviewing 2 percent means 98 percent of your customer interactions are unexamined, and the failures in them are not distributed politely in proportion to the sample. Compliance breaches, incorrect information given confidently, customers signaling they are about to leave, escalations that should have happened: these live in the unreviewed pile, and their cost is real whether or not anyone measures it. Full coverage is the difference between knowing what happens in your operation and estimating it.
How to put a number on it honestly
You do not need the automation in place to measure this pool. You need one deliberate exercise you can run in a week.
- Draw a genuinely random sample from the conversations you never review. Not the escalated ones, not the low CSAT ones, because a biased sample inflates your finding and the inflation will be discovered.
- Have an experienced reviewer read all of them against your existing scorecard, recording material problems separately: policy breaches, incorrect information, missed escalations, unrecovered churn signals.
- Count findings by category, not just the score. A business case needs “this many instances of this specific problem,” not an average.
- Extrapolate carefully and state the method. The rate found in the sample, applied to the unreviewed population, gives an estimated annual volume per problem type. Show the arithmetic so a reviewer can check it.
- Attach cost only where you can defend it. Use your own figures for categories your business already tracks, such as a repeat contact or a churned account, and leave the rest as counted findings with no value attached.
Be disciplined about that last step. A model that prices only what it can defend beats one that prices everything and gets challenged on its weakest line. The unpriced findings still sit alongside the number, as the reason the estimate is conservative.
This is also the pool where one real customer result beats a benchmark. At UiPath, Kaizo automated 100 percent of QA, which the customer associated with a 200 percent return on investment and an 8 percent lift in quality score. That is one organization’s outcome under their own conditions, not a figure for your model, but it shows the shape of the case: the value appeared when coverage went from a sample to everything.
Pool 3: quality improvement that flows to outcomes
The third pool is the one your CFO cares about most and the hardest to prove: quality scores improving, and those improvements reaching metrics that have money attached. Faster coaching cycles because a manager sees a problem days after it happens instead of weeks. Fewer repeat contacts because the root cause was found and fixed. CSAT moving because the behaviors that drive it are measured on every conversation instead of a handful.
Attribution is the whole problem
Support metrics move for many reasons at once: seasonality, a product release, a pricing change, a hiring wave. Claim a CSAT movement as the return on a QA investment without controlling for any of that and it will not survive its first serious review, which damages the credible parts of your case too. Three practices make the pool defensible:
- Capture the baseline before anything changes. Record your quality score distribution, repeat contact rate, average handle time, CSAT, and coaching cadence, at team level, before implementation begins. A baseline captured afterward is not a baseline.
- Prefer the metrics closest to the intervention. Coaching cycle time and repeat contact rate sit close to what QA changes, so the causal story is short. CSAT and retention sit further away and have many other inputs. Claim the near metrics confidently and the far ones cautiously.
- Stage the rollout if you can. If one group of teams goes first, the teams that have not yet moved are a natural comparison group, and a difference between them is far more convincing than a before-and-after line on its own.
Then be conservative on purpose. Model the pool at a fraction of the improvement you expect and say so in the document. Coming in ahead of a modest projection earns the right to make a bigger ask next year. Missing an ambitious one makes every future number you present suspect. The mechanics of converting findings into coaching that changes behavior are covered in turning QA data into coaching.
Pool 4: risk avoided, framed as insurance
The fourth pool needs careful handling. In a regulated environment, a single compliance failure that reaches a regulator, or a systematic pattern of agents giving incorrect information about a product, can cost more than the entire QA program many times over. Sampling finds these late or not at all, because a rare event in a small sample is usually an absent event.
The temptation is to price the worst case and put it in the total. Do not. A number built from a hypothetical fine is the weakest line in any model and it will be attacked first, which puts your solid pools on the defensive. Present risk the way insurance is presented: as coverage against a class of event, with the cost drawn from your own history where you have one.
What makes this credible is specificity rather than size. Name the failure modes that actually apply to you: the disclosure your agents are required to give and do not always give, the product claim that has drifted into something not quite true, the identity verification step that gets skipped when the queue is long. Then state plainly what the current program can and cannot see. With sampling you detect a pattern like this only if it happens to land in the reviewed fraction. With full coverage you detect it the first time it recurs.
For most finance reviewers, one clearly described risk the current process demonstrably cannot see is worth more than a large speculative number. It reframes the spend from an efficiency purchase to a control, and controls are judged on whether they work, not on their payback period.
The four value pools side by side
Present all four together so the reader can see the shape of the case at a glance, including how confident you are in each line. Showing your confidence level is not a weakness in the document. It is the thing that makes the confident lines believable.
| Value pool | How to calculate it | How confident you can be | Typical size relative to the others |
|---|---|---|---|
| Reviewer time reclaimed | Evaluations per month x minutes each / 60 x loaded hourly cost, then the share automation absorbs | High. Every input comes from your own review logs and payroll | Smallest. Real but modest, and soft cost rather than cash |
| Coverage value | Sample the conversations you never review, count findings by category, extrapolate to the unreviewed population | Medium. The count is solid, the extrapolation and pricing are estimates you must show working for | Largest in practice, and the pool most cases omit entirely |
| Quality flowing to outcomes | Baseline before implementation, then track coaching cycle time, repeat contacts and CSAT against it | Medium to low without a control group. Near metrics confidently, far metrics cautiously | Substantial over time, near zero in the first quarter |
| Risk avoided | Name the failure modes sampling cannot see, price only what your own history supports | Low as a number, high as an argument. Insurance, not a line item | Unbounded and rare. Decisive in regulated operations |
The honest counterweights: what to subtract
A business case with no costs in it is not a business case. The counterweights make the document more credible, and they stop objections arriving as surprises in the meeting where you ask for the money.
Implementation and calibration
Someone has to connect the system to your helpdesk, translate your scorecard into criteria a machine can apply, and get the first results into a shape people trust. That is internal time from your QA lead and probably an admin, over weeks rather than days. Scoring then has to be calibrated against a human standard and re-checked as policies change, so budget recurring reviewer time for calibration and disputes as well. This is the work a QA analyst should be doing rather than grading by hand, but it still occupies hours you must show, costed at the loaded rate from pool one.
Change management
Moving from a small sample to every conversation being scored is a significant change for agents, and it lands badly if it feels like surveillance. Team leads need training on how to use the output, agents need a clear dispute path, and someone has to explain the change credibly. This cost is mostly management attention, but it determines whether the other three pools ever materialize, so name it.
The headcount trap
This is the most important paragraph in the model. Reclaimed reviewer time is almost never removed from the budget. It gets redeployed into coaching, root cause work, calibration, and acting on the findings the coverage pool surfaced, which is the correct outcome, because those findings are worthless if nobody has time to act on them.
So do not present pool one as headcount savings unless you genuinely intend to reduce headcount and have agreement to do so. If you promise a reduction you will be held to it, and when it does not appear the whole case is retrospectively judged to have failed, including the pools that delivered. Present reclaimed time as capacity redirected to higher value work, state where it goes, and let the reader see that the redirection is what makes pools two and three possible.
A worked model you fill in with your own numbers
Below is the model with placeholders in square brackets. Every bracketed item is a number you replace with one from your own operation. There are deliberately no example values, because a plausible-looking figure has a habit of surviving into the final document and being defended as though it were real.
Inputs to gather first
- [EVALUATIONS_PER_MONTH] from your review log, and [MINUTES_PER_EVALUATION] measured over a week of real reviews rather than estimated.
- [LOADED_HOURLY_REVIEWER_COST] from finance, including employer costs and overhead.
- [SHARE_OF_GRADING_AUTOMATED] as a conservative percentage after allowing for calibration and disputes.
- [UNREVIEWED_CONVERSATIONS_PER_MONTH], and [FINDINGS_PER_HUNDRED_SAMPLED] per problem category from the pool two exercise.
- [COST_PER_REPEAT_CONTACT] and any other unit costs your business already tracks and stands behind.
- [ANNUAL_PROGRAM_COST], covering licensing plus the implementation, calibration, and change management time costed above.
The model
- Pool 1: [EVALUATIONS_PER_MONTH] x [MINUTES_PER_EVALUATION] / 60 x [LOADED_HOURLY_REVIEWER_COST] x [SHARE_OF_GRADING_AUTOMATED] x 12, labeled clearly as capacity redirected rather than cash removed.
- Pool 2: per problem category, [FINDINGS_PER_HUNDRED_SAMPLED] / 100 x [UNREVIEWED_CONVERSATIONS_PER_MONTH] x 12 gives the estimated annual volume now visible. Multiply by your own unit cost only where you have one, and list the rest as counted findings.
- Pool 3: a deliberately conservative fraction of your expected improvement in the metric closest to the intervention, valued at your own unit cost, with the baseline date and comparison method noted.
- Pool 4: no number. A short list of named failure modes your sampling cannot detect, with the note that any single occurrence would exceed [ANNUAL_PROGRAM_COST].
- Net: pools 1 to 3 minus [ANNUAL_PROGRAM_COST], with pool 4 alongside as the risk argument.
- Payback: [ANNUAL_PROGRAM_COST] divided by the monthly total of pools 1 to 3, in months, excluding the first quarter because pool 3 will not have arrived yet.
Then show the model as a range: a conservative case, a base case, and an upside, differing only in the assumptions you flagged as estimates. A single confident number invites a debate about that number. A range invites a debate about the assumptions, which is the conversation you want.
How to present it so it survives the meeting
The document that wins is rarely the one with the biggest number. It is the one whose author has clearly already asked the hard questions.
- Lead with the coverage gap, not the savings. Open with the share of conversations nobody reads and what you found in a sample of them. That is the problem statement. Everything else responds to it.
- Separate cash, capacity, and risk. Do not add a soft time saving to a hard cost avoided. Finance will separate them anyway.
- Show the costs and your confidence per line. Implementation, calibration, and change management all visible, and a high, medium or low against each value pool. Flagging the weak lines buys credibility for the strong ones.
- Commit to a measurement plan. Name the baseline, the review date, and the metrics you will be judged on. An offer to be measured signals that you believe the model.
- Bring one real example. A single conversation from the unreviewed pile containing a genuine failure does more in the room than any slide of arithmetic.
- Say who graded it, and whether they had a stake in the answer. A finance reviewer will ask, and the answer should be short.
Two properties that decide whether the model holds
The first is verifiability. Your whole case rests on findings being checkable. If a score does not link back to the exact evidence that produced it, you cannot audit pool two, you cannot defend pool three, and the model is asking finance to take a vendor’s word for its own inputs. Insist that every score points at the transcript line behind it, that any score can be challenged and reviewed by a human, and that the grader can be run against conversations your own reviewers already scored so you can show an agreement rate rather than assert one. That single test converts your model from an estimate into something a skeptical reviewer can reproduce.
The second is neutrality, and it gets more important every quarter as more of your volume is handled by AI. Most QA and CX platforms now sell AI agents of their own, which means the system producing your quality numbers is assessing a product its own company built. Any pool you calculate on top of those numbers inherits that conflict, and it is a fair question for a CFO to ask. Kaizo is one of the only QA platforms that does not sell AI agents, so it has nothing in the conversation to defend. It scores against the scorecard your business already uses, native to Zendesk and Salesforce, and traces each score to the transcript line behind it, so a finance reviewer can check a finding instead of believing it. Whichever system you evaluate, insist on both properties.
Frequently asked questions
How do you calculate the ROI of QA automation?
Build it from four value pools rather than one. Calculate reviewer time reclaimed from your own evaluation volume and loaded hourly cost, estimate coverage value by sampling the conversations you never review and counting what you find, project quality improvements against a baseline you captured before implementation, and describe risk avoided qualitatively. Then subtract implementation, calibration, and change management costs. Present the pools separately with a confidence level on each.
What is the biggest source of value in an auto QA business case?
Coverage, in almost every case. Reviewer time reclaimed is the easiest pool to calculate and the smallest. The larger value comes from the failures sitting in the conversations nobody was reviewing: compliance breaches, incorrect information, missed escalations, and churn signals that a small manual sample was statistically unlikely to ever surface.
Should I promise headcount savings in the business case?
No, unless you genuinely intend to reduce headcount and have agreement to do so. Reclaimed reviewer time is almost always redeployed into coaching, calibration, and acting on the new findings rather than removed from the budget. If you promise a reduction you will be held to it, and failing to deliver it will discredit the parts of the model that did deliver.
How do I prove that a quality score improvement was caused by auto QA?
You cannot prove it outright, so build for credibility instead. Capture a baseline before anything changes, prefer metrics close to the intervention such as coaching cycle time and repeat contact rate over distant ones such as retention, and stage the rollout so the teams that have not moved yet act as a comparison group. Then model the pool conservatively and say that you have.
What costs should be subtracted from the auto QA business case?
Licensing plus three internal costs that teams routinely forget: implementation and configuration time to connect the system and translate your scorecard, ongoing calibration and dispute handling time, and the management attention required for change management with agents and team leads. Cost the internal time at the same loaded hourly rate you used for the reviewer time pool.
How long does auto QA take to pay back?
It depends entirely on your conversation volume, your current review sample size, and your program cost, which is why the honest answer is to calculate it rather than quote a benchmark. Divide your annual program cost by the monthly value of the pools you can quantify, and exclude the first quarter from the calculation, because implementation and calibration happen before the outcome improvements arrive.
Related terms
Put real numbers in your business case
Bring a slice of the conversations nobody is reviewing today and we will score them with you, so the coverage pool in your model comes from your own data instead of an assumption. Every finding traces back to the transcript line behind it, so your finance reviewer can audit it. Kaizo does not sell AI agents, so it has no stake in the numbers it hands you.