Skip to content

Best practice

BPO Quality Assurance: Why a 2% Sample Is Not Enough

What a vendor-scored sample leaves unread, why coverage alone does not fix it, and how one shared scorecard replaces the monthly QA spreadsheet.

· 13 min read

On this page

A 2% sample fails because the vendor reads 2% to 3% of its own tickets and grades its own work, leaving 97% to 98% unread. Governing a BPO takes evaluation of every conversation and a scorecard both client and vendor can audit.

Part 1 of 4 in the outsourced governance series. Written for support, CX and vendor operations leaders.

In short

  • Small samples miss the rare failures that cost accounts.
  • Full coverage is now standard. Traceable, challengeable scores are what differ.
  • Several vendors need one shared scorecard to compare or enforce an SLA.
  • A shared scorecard helps the BPO prove its quality too.

Jump to

  1. What it means today
  2. How much a sample misses
  3. Coverage is only the start
  4. Multiple vendors, one baseline
  5. Why BPOs benefit too
  6. How to make the move
  7. The question to ask

What BPO quality assurance actually means today

When a scaling software company hands frontline support to a Business Process Outsourcing partner, the unit economics are compelling. Coverage extends to 24 hours a day. Headcount scales without a hiring pipeline. Cost per contact falls. None of that is in dispute, and the companies doing it are not making a mistake.

What often does not scale with the headcount is the governance. The support conversations move offshore, but the measurement system that tells you whether those conversations are any good usually stays exactly as it was when the team sat in your building: a sample of tickets, read by a person, scored on a form.

The difference is that the person reading them now works for the vendor being measured.

Outsourcing the headcount is a sound commercial decision. Outsourcing the scorecard is a different decision, and most companies make it by accident.

In a typical arrangement the BPO employs its own quality analysts. Those analysts pull a sample of the week’s tickets, read them against a rubric, assign scores, and roll the result into a monthly report for the client. The report arrives as a spreadsheet. It usually says quality is high.

It is not that the number is dishonest. It is that nobody on the client side can check it, because the tickets behind it were selected by the party being graded and the reasoning behind each score was never written down anywhere the client can see.

How much does a 2% sample actually miss?

Take a mid-market support operation running 15,000 conversations a month through an outsourcing partner. At a 3% sample, the quality programme reads 450 of them. The remaining 14,550 are never evaluated by anyone. 450conversations read per month at a 3% sample of 15,000 14,550conversations nobody evaluates, out of the same volume 1 in 50failure rate that a 50-ticket sample will usually miss entirely

The problem is not only the size of the gap. It is which failures live inside it. Sampling is good at detecting things that happen constantly and bad at detecting things that happen occasionally, and the failures that cost you real money are almost always occasional.

A tone problem that shows up in one conversation out of three will appear in any sample you take. A migration instruction that is wrong once a fortnight, given to whichever customer happens to ask, will not. The statistics of small samples are unforgiving here, and no amount of analyst diligence fixes them.

The second sampling problem nobody mentions

A sample is only unbiased if it is drawn at random. In practice, ticket selection is often left to the analyst or shaped by convenience: recent tickets, short tickets, tickets from queues the analyst knows well. Every one of those habits makes the resulting score less representative than the sample size alone would suggest. How conversations get selected for review quietly determines what the score means.

Sampled vendor QA compared with evaluation across every conversation | | Sampled QA, scored by the vendor | Full-coverage QA on a shared scorecard | |---|---|---| | Conversations evaluated | 2% to 3%, selected by the party being measured | Every conversation, selected by nobody | | Who owns the rubric | Each vendor maintains its own | One rubric, applied identically to every partner | | Reporting lag | Monthly or bi-weekly, after the fact | Continuous, visible to both sides at once | | Rare failure modes | Statistically invisible | Surfaced as a pattern once they recur | | Disputing a score | Argue about a spreadsheet cell | Open the conversation and read the evidence | | Comparing two vendors | Not possible, different rubrics | Direct, same model and same criteria |

Why full coverage is the starting line, not the answer

It would be convenient to end the argument here: score every conversation instead of 3% of them and the governance problem disappears. That was a defensible claim a few years ago. It is not one now.

Scoring 100% of conversations has become ordinary. Most serious automated QA platforms do it, and a buyer comparing vendors will hear the same number from all of them. Coverage is the precondition for governing an outsourced operation. It is not the thing that makes the governance trustworthy.

Once every vendor claims to read every conversation, the only question left is whether you can check the reading.

Two things decide that.

Can you trace the score back to the evidence?

A quality score is an assertion until someone can open it and see the specific moment in the conversation that produced it. If a BPO agent is marked down for a compliance miss, the account manager should be able to click the criterion and land on the sentence that triggered it. Without that, the client has swapped one unauditable number for another, faster one. Score traceability is what makes a governance conversation possible instead of a standoff.

Is the grader independent of the work being graded?

This is the structural point, and it applies at two levels. The vendor should not be the sole author of its own quality score, for the same reason auditors do not work for the company they audit. And as outsourced operations start blending human agents with AI agents, the platform doing the grading should not also be selling the AI agents it grades. Kaizo is one of the few QA platforms that does not sell its own AI agents, which is precisely why it can score AI-handled conversations without marking its own homework.

A note on how this lands with agents

Full coverage is easy to present badly. Told that every conversation is now read, outsourced agents reasonably hear surveillance, and QA programmes that arrive sounding like monitoring meet resistance that no dashboard fixes. The honest framing is the accurate one: coverage removes selection bias. Under a 3% sample, most good work was invisible and a single unlucky ticket could define a quarter. Programmes agents trust are the ones that make that case first.

What one shared scorecard changes across multiple vendors

Few companies stay on one partner. A mature operation might run an overnight vendor in one region, a bilingual partner in another, and an internal team handling complex enterprise escalations. Each arrives with its own quality system, its own rubric, and its own analysts.

The result is three quality scores that cannot be compared. Vendor A reports 94%, Vendor B reports 91%, and there is no way to know whether that gap is real or an artefact of two different rubrics weighted differently. Meanwhile the internal team is measured on a fourth system, so nobody can answer the only question that matters commercially: which of these is actually better, and at what cost?

Putting every partner on one scorecard changes four things at once.

  • Comparison becomes possible. The same criteria, weighted the same way, applied by the same model to every conversation from every vendor. How you weight the criteria becomes a decision you make once, deliberately, rather than one each vendor makes for you.
  • SLA enforcement gets an evidentiary basis. A contractual quality target is only enforceable if both parties accept the measurement. A traceable score on full coverage is defensible in a commercial review in a way that a sampled self-report is not.
  • Coaching gets specific. Instead of a generic instruction to improve resolution time, you can see that a particular group is failing a particular criterion on a particular ticket type, which is the level of detail coaching actually needs.
  • Routing decisions get evidence. Which partner should absorb the next volume increase stops being a procurement negotiation and becomes a data question. See your own vendor mix on one scorecard Bring last month’s conversations from each partner. We will score them against a single rubric and show you where the vendors genuinely differ.

Book a demo

Why this is good news for the BPO as well

It is easy to write about outsourced quality as though the vendor were the problem. That framing is popular, it is adversarial, and it is mostly wrong. BPOs do not want to deliver poor conversations. They are running a business whose entire commercial case rests on delivering service the client would struggle to deliver as cheaply, and a client who cannot see quality is a client who renegotiates on price alone.

A vendor with no way to prove its quality can only compete on rate. Evidence is what lets a good BPO argue for something better than the cheapest bid.

The monthly self-graded spreadsheet is a bad deal for the outsourcing partner too. It generates suspicion it cannot answer, it makes every quality conversation a negotiation about methodology, and it gives the vendor no mechanism to demonstrate that its team is better than a cheaper competitor. A shared, traceable scorecard replaces all of that with a record both sides read the same way.

This is not theoretical. EverHelp, an outsourcing provider, put its own operation on automated QA and grew QA ratings by 270% while holding its internal quality score steady, and cut the time its leads spent preparing coaching sessions by 75%. That is a vendor using evaluation as a commercial asset, not submitting to an audit. The full EverHelp story is worth reading if you sit on the vendor side of this relationship.

What good looks like from both sides

  • The client can see quality continuously instead of monthly, and can explain any score it questions.
  • The vendor can prove its performance with evidence rather than assertion, and can point to specific coaching it ran in response.
  • Both parties argue about the same conversations rather than about whose rubric is fairer.
  • Calibration between client reviewers and vendor reviewers becomes a scheduled exercise rather than a source of friction. Running calibration sessions is how the two sides stay aligned on what the criteria mean.

How to move an outsourced programme onto one scorecard

This does not have to start as a contract renegotiation. In practice the sequence that works is deliberately unthreatening.

  1. Agree the rubric before you measure anything. Write the criteria with the vendor in the room. A rubric imposed on a partner produces disputes; a rubric co-authored produces a baseline. What belongs in a rubric is a harder question than it looks.
  2. Run in shadow first. Score every conversation without attaching consequences to the result, and compare the output against the vendor’s existing sampled scores. Shadow scoring is how both sides build confidence in the grader before anything depends on it.
  3. Validate the grader on your own conversations. Take a set of tickets your team has already scored by hand, run them through, and check where the two disagree and why. Validating AI scoring against a known set is the step most programmes skip and later regret.
  4. Publish the dispute route on day one. Agents and team leads need a defined way to challenge a score and see it reviewed. A working dispute process is what separates a governance tool from a surveillance tool.
  5. Only then attach it to the commercial review. Once both sides trust the measurement, it can carry SLA weight.

Is your BPO governance auditable? Nine questions

Tick every statement you can support with evidence today, not with a vendor assurance.

  • I know what percentage of my outsourced conversations were evaluated last month.
  • I know who selected the conversations that were evaluated.
  • Every vendor and internal team is scored against the same rubric, weighted the same way.
  • I can open any individual score and see the evidence in the conversation that produced it.
  • I could challenge a specific score in a vendor review and resolve it from the record.
  • I know which failure category costs me the most, across every vendor.
  • Agents at the vendor have a defined route to dispute a score.
  • My quality measurement is independent of the party being measured.
  • I could say today which of my partners delivers the best quality per unit of cost.

Six or fewer ticks means your quality reporting is an assertion rather than a record. That is the normal starting position, and it is fixable without changing vendors.

The question worth asking at the next vendor review

Of the customers who left last quarter, how many had conversations handled by an outsourcing partner, and what happened in those conversations?

Most companies cannot answer that, because the conversations in question were part of the 97% nobody read. The answer is not to bring support back in house. It is to stop accepting a sampled self-report as governance when the volume, the contract value and the customer relationships involved all deserve a record.

UiPath now automates 100% of its QA with Kaizo, measures a 200% return against its own baseline, and has lifted its quality score by 8 points while continuing to improve it quarter on quarter. That is the scale at which quality measurement stops being an operations task and starts being an input to how the business is run.

Outsourcing the work was the right call. Keeping the standard is the part you do not delegate.

Frequently asked questions What is BPO quality assurance?

BPO quality assurance is the process of evaluating the quality of customer conversations handled by an outsourcing partner against a defined set of criteria. Traditionally it is performed by quality analysts employed by the BPO, who manually review a sample of 2% to 3% of tickets and report scores to the client each month. Modern programmes evaluate every conversation automatically against a single rubric shared by the client and the vendor, so that both parties read the same record. What percentage of tickets does a typical BPO quality programme review?

Between 2% and 3% of total ticket volume is the common range for manual sampling, which leaves 97% to 98% of conversations unevaluated. On a volume of 15,000 conversations a month, a 3% sample means 450 conversations are read and 14,550 are not. The size of the sample matters less than the fact that failures which occur occasionally rather than constantly are unlikely to appear in it at all. Why is sampled quality assurance a problem when support is outsourced?

Two reasons compound. First, a small sample cannot reliably detect rare failure modes, and rare failures are usually the expensive ones. Second, when the vendor selects the sample and scores it, the client has no independent way to verify the result. The score may well be accurate, but it is an assertion rather than a record, which makes it difficult to act on commercially. Is scoring 100% of conversations enough to fix BPO governance?

No. Full coverage is now standard across serious automated QA platforms, so it no longer distinguishes one approach from another. What matters is whether each score can be traced back to the specific evidence in the conversation that produced it, and whether the party doing the grading is independent of the work being graded. Coverage is the precondition; traceability and independence are what make the measurement trustworthy. How do you compare two BPO vendors fairly?

Put both on the same rubric, weighted identically, applied by the same evaluation model to every conversation from each vendor. As long as each partner maintains its own quality system, reported scores are not comparable, because the difference between 94% and 91% may be entirely explained by how each rubric is constructed rather than by any difference in the conversations. Does this approach work for the BPO as well as the client?

Yes, and it is usually a better deal for a capable vendor. Without shared evidence, a BPO can only compete on rate, because it has no way to demonstrate that its service is better than a cheaper bid. A traceable scorecard lets a vendor prove quality, target coaching precisely, and defend its position in a commercial review. EverHelp, an outsourcing provider, grew QA ratings by 270% on automated QA while cutting coaching preparation time by 75%. Do we have to change BPO vendors to fix this?

Generally no. The measurement layer sits over the existing operation through the helpdesk rather than replacing the vendor’s tooling, so the same partners keep doing the same work while both sides gain a shared record. Most programmes start in shadow mode, scoring conversations without commercial consequence, and only attach the result to SLA reviews once both parties trust the grader. Which helpdesks does Kaizo work with for outsourced operations?

Kaizo integrates natively with Zendesk and Salesforce Service Cloud. Because the integration is native rather than a generic connector, the evaluation layer sits directly over the conversations your vendors already handle, which is why onboarding is measured in days rather than months.

Keep reading

Put every vendor on the same scorecard

Bring one month of conversations from each of your outsourcing partners. We will score them all against a single rubric and show you, with the evidence attached, where the vendors genuinely differ.

Book a demoKaizo for BPOs

EU AI Act ready . SOC 2 and ISO 27001 certified . Native Zendesk and Salesforce Service Cloud integrations

In Kaizo Kaizo for BPOs BPOs live or die on demonstrable quality across clients who each define it differently. Kaizo scores every conversation against every client’s own criteria. See Kaizo for BPOs

On this page

See this on your own conversations

We will score a sample of your real tickets against your standards, so the example is yours.

Trusted by global support teams

  • Foot Locker
  • SteelSeries
  • Canva
  • GetYourGuide
  • Instacart