How to Build a Customer Service QA Program From Scratch: a 7 Step Playbook

A zero to one playbook for building a customer service QA program: define quality, build a small scorecard, set coverage, calibrate reviewers, pilot, coach, and evolve. With a 12 week rollout plan.
Playbook · Quality Assurance

Building a customer service QA program from scratch means working through seven steps in order: define what quality means for your business and tie it to an outcome you already track, build a deliberately small scorecard, decide how much of your conversation volume you will review, choose and calibrate your reviewers, pilot on one small team with the rubric published in advance, turn scores into coaching on a fixed cadence, and then review the rubric on a schedule. The order matters more than the tooling. Most programs that fail did not fail because the software was wrong, they failed because scoring started before anyone agreed on what a good conversation looks like. A realistic timeline from nothing to a working, trusted program is about twelve weeks.

In short

  • Do the steps in order. Defining quality before building a scorecard, and calibrating reviewers before scoring anything real, are the two sequencing decisions that decide whether the program survives.
  • Tie your definition of quality to a business outcome you already report on, otherwise nobody outside the QA team will care about the score.
  • Start with a small scorecard and resist adding to it. Rubric bloat is the single most common way a new program collapses under its own weight.
  • Publish the rubric to agents before the first real score is given. A program agents did not see coming is a program they will not trust.
  • A score with no coaching attached is an audit, not a program. Book the coaching cadence at the same time you launch the scorecard.
  • A spreadsheet is a legitimate way to start and forces you to define the rubric properly. It stops scaling at the point where you want coverage beyond a small sample.
  • Give the program a named owner. Programs without an owner do not get killed, they quietly stop being run.

Step 1: Define what quality means for your business, and tie it to an outcome you already care about

Give this a week, and do it before you look at any tooling. The output of step one is a single written paragraph that says what a good conversation looks like in your business, and one sentence naming the business outcome that paragraph is supposed to move.

That second half is the part most new programs skip. A QA score that only exists to describe itself has no constituency: the support leadership team will nod at it, and the rest of the business will ignore it. A QA score that is explicitly built to reduce repeat contacts, or to cut escalations, or to protect a compliance obligation, has a reason to exist that survives the first budget conversation.

How to actually produce it

Pull twenty to thirty recent conversations, deliberately mixed: some that ended well, some that generated a complaint, some that were reopened. Read them with two or three people who know the operation. Write down what separates the good ones from the bad ones in plain language, not in QA vocabulary. You will usually find three or four themes, something like: the agent understood the actual problem, the information given was correct, the customer knew what would happen next, and the tone fit the situation.

Those themes become the spine of your scorecard in step two. If you want the wider conceptual grounding before you start, the complete guide to customer service quality assurance covers what QA is for and how it fits alongside your other quality signals.

Step 2: Build the scorecard small, then resist adding to it

Take one week. Turn the themes from step one into criteria, and then keep cutting until it hurts a little.

Our opinionated recommendation is to launch with somewhere around eight to twelve criteria, and no more. This is a judgment call rather than an industry standard, and the reasoning is practical: a reviewer holding twelve criteria in their head can score a conversation in a few minutes and stay consistent across a shift. At thirty criteria they cannot. They start pattern-matching, the scores drift, and the resulting number stops meaning anything. A smaller scorecard scored consistently is worth far more than a thorough one scored erratically.

Make each criterion something two people would score the same way

The test for every criterion is: if two reviewers read the same conversation, would they give the same answer? Prefer criteria that can be checked against evidence in the transcript, such as whether identity was verified or whether the next step was stated, over criteria that ask for a general impression. Keep one or two subjective criteria if tone genuinely matters to you, but write them so the intent is unmistakable.

Weight the criteria rather than treating them all as equal, and decide up front which failures are automatic fails, usually the compliance and data-handling ones. For the mechanics of scales, weighting and pass thresholds, the QA scorecard guide and the explainer on what a QA rubric is go deeper than this playbook should. If you would rather start from something that already exists and cut it down, the scorecard templates are a faster starting point than a blank page, and the customer service QA checklist is a useful sanity check on what you may have left out.

Step 3: Decide coverage and sampling, and be honest that sampling is a compromise

Take a few days. You need one number: how many conversations per agent, per period, will actually get reviewed.

If you are reviewing manually, our recommendation for a starting point is two to four evaluations per agent per month. Again, this is a considered recommendation rather than a benchmark. The reasoning is that fewer than two gives you a sample so thin that a single unusual conversation swings an agent’s score, and more than four is more reviewing time than most teams can sustain once the novelty wears off. Pick the number you can genuinely staff every month, not the number that sounds rigorous.

Say out loud what sampling costs you

Here is the honest part, and it is worth telling your stakeholders in week three rather than month six. A manual sample of a few conversations per agent tells you something about those conversations and almost nothing about the rest. The failures that generate complaints and churn are, by definition, rare, so they are unlikely to land in a small random sample. Sampling is good enough to coach individuals on visible habits. It is not good enough to find the systematic problem hiding in the conversations nobody read.

That does not make sampling wrong. It makes it a compromise you are choosing knowingly. Automated scoring is what removes the compromise, because the constraint stops being reviewer hours: at UiPath, Kaizo automated 100% of QA, delivering 200% ROI and an 8% lift in quality score. If you want to understand what changes when the sample becomes the whole population, see 100% QA coverage. For your first twelve weeks, though, a sample you actually complete beats full coverage you have not bought yet.

Step 4: Pick who reviews, and calibrate them before you score anything real

Take a week, and do not skip the second half of it.

Who reviews depends on your size. Below roughly forty agents, team leads reviewing their own people is usually the pragmatic answer, with the caveat that leads score their own team generously and you should expect it. Above that, a dedicated reviewer or two gives you far more consistency, because they are scoring across teams and can see the difference. The QA analyst role explains what that job actually involves if you are considering hiring for it.

Calibrate first, score second

This is the step new programs most often reverse, and it is expensive to reverse. Before a single score goes on an agent’s record, get every reviewer to independently score the same three to five conversations, then sit together and compare. You will be surprised how far apart they are on their first attempt. That gap is not a people problem, it is a rubric problem: wherever reviewers disagree, a criterion is ambiguous and needs rewriting.

Rerun the exercise until the reviewers land close to each other, then rerun it monthly forever, because reviewers drift. QA calibration covers why this works, and how to run calibration sessions gives you the format for the meeting itself. If you launch a program on an uncalibrated rubric, the first agent to compare their score with a teammate’s will find the inconsistency, and you will spend your credibility defending it.

Step 5: Run a pilot on a small group, and publish the rubric first

Give this two to four weeks. Pick one team, ideally one where the lead is on board, and run the full loop on them only. Our suggested pilot size is a single team of roughly eight to fifteen agents: small enough that you can fix problems quickly, large enough that the scores are not dominated by one person’s week.

Publishing the rubric is not optional

Send agents the scorecard, the weightings, the automatic fails, and worked examples of a passing and a failing conversation, before the first review happens. Then say plainly what the scores will and will not be used for. Almost every story of a QA program that agents resent starts the same way: scores appeared, nobody knew the criteria, and the first time an agent saw the rubric was when they were being marked down against it.

Treat the pilot as a test of the rubric rather than a test of the agents. Expect to rewrite two or three criteria in the first fortnight, and tell the pilot team that in advance, because inviting them to break your rubric turns them from subjects into collaborators. Ask them directly: which of these scores felt unfair, and why? Nearly every unfair-feeling score traces back to a criterion that reads differently to the person doing the work than to the person doing the scoring. Running a QA program agents trust goes further into the fairness mechanics, including the dispute path you should have in place before you scale beyond the pilot.

Step 6: Turn scores into coaching, on a cadence you have already booked

This is the step that separates a program from an audit, and it is where most of the value is. A score that nobody discusses with the agent changes nothing at all. Worse, it teaches agents that QA is surveillance, which is very hard to undo later.

Book the cadence at the same moment you launch the scorecard, before the first scores exist. Our recommendation is a short one-to-one every month per agent, roughly twenty to thirty minutes, focused on exactly one thing to improve. The reasoning behind picking one thing: a coaching session that lists six weaknesses produces no change at all, whereas a session that names one behavior, shows the transcript where it happened, and agrees on what to do differently produces change you can see in next month’s scores.

Coach on patterns, not on individual bad conversations

The useful unit is a repeated behavior across several conversations, not a single bad day. Bring two or three examples of the same pattern and the agent cannot reasonably dispute it, which changes the conversation from defending one score to fixing one habit. Turning QA data into coaching covers the session structure in detail. And watch for the pattern that is not an individual issue at all: when six agents fail the same criterion, that is a process, knowledge base or training problem, and coaching individuals on it is the wrong fix.

Step 7: Review and evolve the rubric on a schedule

Set the date now, in your first month, for a review roughly ninety days out. Then repeat it quarterly. Putting it in the calendar before you need it is the whole trick, because a rubric that is only revisited when someone complains will be revisited under pressure, defensively, and badly.

At each review, look at three things. First, criteria where nearly everyone scores full marks: they no longer discriminate between good and poor work, so they are costing reviewer time and buying no information. Retire them or raise the bar. Second, criteria where reviewers still disagree after calibration: rewrite them or drop them. Third, what has changed in the business, because new products, policies and channels create quality expectations your rubric has never heard of.

Change slowly and announce it

Every rubric change breaks comparability with your historical scores, so change on a schedule, in batches, and tell agents what changed and why before the new version goes live. A rubric that shifts quietly under agents is a rubric they will stop trusting, and once they see the score as arbitrary you have lost the coaching loop that step six depends on. Keep the scorecard roughly the same size it started at: for every criterion you add, ask what you are removing.

Your first 12 weeks, and how to prove it is working

What to measure to prove the program is working

You will be asked to justify the time this takes, usually around month four. Decide now what you will show, and start capturing it in week one so you have a baseline to compare against.

  • Movement in the outcome you named in step one: repeat contact rate, escalations, reopens, whatever you tied quality to. This is the only number that makes the program matter to people outside support.
  • Reviewer agreement over time: how close your reviewers are on the same conversation. Rising agreement means the rubric is getting clearer and the scores are getting more trustworthy.
  • Score movement for coached agents: specifically on the one criterion each agent was coached on. This shows the coaching loop works, rather than showing that averages wobble.
  • Coaching sessions actually held versus scheduled: the least glamorous metric here and the best early warning. When this slips, the program is dying and the score chart will not tell you for months.
  • Agent perception: ask two questions in your engagement survey, whether agents understand how they are scored and whether they think it is fair. If those fall, nothing else you measure is reliable.

Do not commit to a target score. A QA average is an artifact of your rubric’s difficulty, so chasing it just pressures reviewers into leniency. If you want a composite view of quality alongside your other signals, the internal quality score is a more honest place to build a headline number than a raw scorecard average.

The 12 week rollout

Here is the sequence compressed into a schedule. It assumes a QA owner spending part of their week on this, not a full-time team, which is the situation most people building a first program are actually in.

Phase Weeks What you do What done looks like
Define Week 1 Read 20 to 30 real conversations, write your definition of quality, name the business outcome it should move One page anyone in the business can read and agree with
Design Weeks 2 to 4 Draft 8 to 12 weighted criteria, mark the automatic fails, set your coverage number, pick and calibrate reviewers Reviewers score the same conversation within a narrow range
Pilot Weeks 5 to 8 Publish the rubric to one team, score them, hold the first coaching sessions, rewrite the criteria that felt unfair A stable rubric and a pilot team that can explain how they are scored
Scale Weeks 9 to 12 Roll out to the rest of the team, lock the monthly coaching cadence and calibration session, book the 90 day rubric review Every agent scored and coached on a predictable schedule, with an owner named

Spreadsheet or software: an honest look at build versus buy

You do not need to buy anything to start, and there is a real argument for not buying immediately. Building your first rubric in a spreadsheet forces you to define every criterion yourself rather than accepting a vendor’s defaults, and that definition work is the part that actually determines whether the program succeeds. Teams that start in a spreadsheet often understand their own quality standard better than teams that started with a tool.

A spreadsheet will carry you honestly through the first twelve weeks and, for a small team with a stable operation, considerably longer. If that is where you land, you are not doing QA wrong: most of the value in this playbook sits in steps one, two, four and six, and none of those require software.

Where the spreadsheet stops

It stops at coverage. Everything a spreadsheet does well depends on a human reading each conversation, so your review volume is capped by reviewer hours, and the gap between what you review and what you handle only widens as you grow. The symptoms are recognizable: reviews slipping later each month, scores that arrive too late to coach on, no way to see whether a problem is one agent or the whole team, and constant nagging about whether the sample is representative. If you are asking those questions, the constraint has moved from your rubric to your capacity.

That is the point where automated scoring earns its cost, because it applies the rubric you already wrote to every conversation instead of a sample of them. Two questions should decide which tool you buy. The first is whether you can verify it: every score should link back to the evidence in the transcript, and you should be able to run conversations your reviewers have already scored through it and measure the agreement yourself. The second is whether the grader is neutral. Most QA and CX platforms now sell their own AI agents, so the system scoring your automation is often the same system supplying it. Kaizo is built for that transition: it scores against your own scorecard, links every score back to the evidence in the transcript so agents and analysts can check it, does not sell AI support agents of its own, and is native to Zendesk and Salesforce, so the QA work happens where your team already works. Full coverage is what makes the checking possible, because you are no longer limited to whichever conversations a reviewer had time for. Bring the rubric you built in the spreadsheet, because the definition work in steps one and two is the part that carries over.

The five reasons new QA programs die

Almost every program that quietly stops running was killed by one of these. All five are cheaper to prevent in week two than to fix in month six.

  • Rubric bloat. Every stakeholder wants their criterion added, the scorecard grows from ten items to forty, scoring takes too long, reviewers start pattern-matching, and the scores stop meaning anything. Prevention: a hard cap agreed in advance, and a rule that adding a criterion requires removing one.
  • No coaching loop. Scores are produced and filed. Nothing changes, so nobody can point to a reason the program exists, and it is the first thing cut when the team gets busy. Prevention: book the coaching cadence before the first score exists, and track sessions held.
  • Agents are not bought in. The rubric was never published, scores appeared without explanation, and there was no way to challenge one. Agents conclude QA is something done to them and disengage. Prevention: publish the rubric first, run the pilot as a test of the rubric, and give people a real dispute path.
  • No owner. QA is everyone’s responsibility, which means it is nobody’s. Reviews slip a week, then a month, then stop, and nobody notices because it was never on anyone’s objectives. Prevention: one named person, with the time for it written down.
  • Scores used punitively. The moment QA scores feed performance management or compensation, reviewers soften their scoring to protect their people and agents optimize for the rubric rather than the customer. You lose the honest signal permanently, and it is very hard to win back. Prevention: state in writing, before launch, that scores are for coaching, and hold that line.

Frequently asked questions

How long does it take to build a customer service QA program?

Plan for about twelve weeks from nothing to a working program: one week to define quality, three weeks to design and calibrate the scorecard, four weeks piloting on one team, and four weeks to roll out and lock the coaching cadence. You can score conversations sooner than that, but scoring before reviewers are calibrated and before agents have seen the rubric tends to cost you more trust than it buys you data.

What is the first step in setting up a QA program?

Define what quality means for your business, in writing, before you touch a scorecard or a tool. Read twenty to thirty real conversations with two or three people who know the operation and write down in plain language what separates the good ones from the bad ones. Then name the business outcome that definition is meant to move, such as repeat contacts or escalations, so the program has a reason to exist that survives a budget review.

How many criteria should a QA scorecard have?

Our recommendation is to launch with roughly eight to twelve weighted criteria rather than a comprehensive list. This is an opinionated starting point, not an industry standard, and the reasoning is that a reviewer can hold about a dozen criteria in their head and stay consistent across a shift. A smaller scorecard scored consistently produces a more useful number than a thorough one scored erratically.

How many conversations should we review per agent?

If you are reviewing manually, two to four evaluations per agent per month is a reasonable place to start, chosen because fewer than two lets one unusual conversation swing an agent’s score and more than four is rarely sustainable. Pick the number you can genuinely staff every month. Be explicit with stakeholders that any small sample tells you about the conversations you read and very little about the ones you did not.

Can you run a QA program in a spreadsheet?

Yes, and it is a legitimate way to start. A spreadsheet forces you to define every criterion yourself instead of accepting a tool’s defaults, which is the work that actually decides whether the program succeeds. It stops scaling when you want coverage beyond a small sample, because review volume is capped by reviewer hours, and that gap widens as the team grows.

Why do new QA programs fail?

Five causes account for most of them: rubric bloat that makes scoring too slow to stay consistent, no coaching loop so scores change nothing, agents who never saw the rubric and therefore do not trust it, no named owner so reviews quietly stop, and scores used for performance management, which destroys the honest signal. Each one is far cheaper to prevent in the design phase than to repair once the program is running.

Built the rubric? See it run on every conversation

Bring the scorecard you designed in steps one and two, along with a few conversations your reviewers have already scored, and we will show you the same rubric applied to your conversations with every score traced back to the transcript, so you can check the grading against your own judgment before you trust it.

Book a demoExplore Agentic Auto QA

Choose your help desk

Not using either? We’ll let you know as soon as we can support your help desk solution.

Kaizo
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.