Rolling out a QA tool is a change management project with a software installation inside it. The configuration takes days. Getting a support team to accept a new instrument grading their work takes weeks, and that is where almost every rollout actually fails. A realistic plan runs about twelve weeks: configure and connect, announce the change before any score lands, calibrate the tool against your own reviewers, pilot on a deliberately mixed team with no consequences attached, then widen and retire the old process once leadership has seen both sets of numbers side by side.
In short
- Configuration is the easy part. Budget for the fortnight of conversations with your team, not the afternoon of setup.
- Tell agents the instrument is changing before the first score lands. Say nothing and the rumour arrives first, and you spend the rollout arguing with it instead of running it.
- Do not pilot on your best team. They pass everything, surface no configuration gaps and file no disputes, so you learn nothing about the path that actually breaks.
- Scores usually fall in month one because scrutiny went up, not because quality fell. Tell leadership that before it happens, not after they have already read it as the tool failing.
- Running the old and the new process together roughly doubles the QA function’s monthly hours for six to ten weeks. Budget those hours or you will abandon the comparison halfway through and lose the evidence you were collecting.
- Never retro-apply a new scorecard to old conversations. Freeze the historical scores, label the version that produced them, and start the new trend line from a clean date.
What a QA tool rollout actually is, and what it is not
A QA tool rollout is not an integration project. The integration is the smallest, most predictable part of it. What you are actually doing is changing the instrument that measures people’s work, on people who are already being measured, and who did not ask for the change.
That distinction decides how you spend your time. Teams that treat the rollout as a technical task book two days with an admin, switch it on, and then spend the next quarter managing the fallout. Teams that treat it as a change project book the same two days for configuration and then eight weeks for everything else. Kotter’s study of a hundred companies in Leading Change: Why Transformation Efforts Fail is the reference text for the difference, and none of the failure patterns he found were technical.
Two jobs sit next to this one and are commonly confused with it.
- Designing the programme. If you have not yet decided what quality means, built a scorecard or agreed who reviews, that comes first and it is a separate piece of work. Start with how to build a QA program from scratch and come back here.
- Choosing the tool. If you are still comparing vendors, how to pick call center QA software covers the selection criteria. This guide starts the day the contract is signed.
So the assumption here is: you have a scorecard you already use, reviewers who already score against it, and a piece of customer service QA software arriving that will do some or all of that work differently. Everything below is about getting from that state to a working one without the team deciding the tool is something being done to them.
What to tell agents before the first score lands
The most expensive mistake in a QA rollout costs nothing to avoid. It is saying nothing until the first automated score appears in someone’s queue.
Here is what happens when you stay quiet. Somebody in the pilot mentions it in a channel. By the end of the week the version circulating on the floor is that management has bought software to monitor everyone and the layoffs are coming. That version is now the thing you have to argue against, and you will be arguing against it for months. You cannot get ahead of a rumour you let start.
The five things the announcement has to contain
One message, from the support lead rather than from the QA team, delivered in team meetings and then written down somewhere agents can go back to. It needs to answer, in this order:
- What is changing, precisely. Not “we are improving quality”. The scoring is moving from spreadsheets to a system, and here is what the reviewer used to do that the system will now do.
- What is not changing. Usually the scorecard itself, the pass mark, and whether the score affects pay. Say this explicitly. In the absence of a statement, everyone assumes all three are changing.
- When the first score with consequences attached will happen. Give a date, and make it later than you think you need. “Scores will be visible from the 3rd, and nothing counts towards anything until the 1st of next month” is a sentence that buys enormous goodwill.
- How to challenge a score. Who to go to, what happens next, how long it takes. A dispute path that exists on day one is worth more than any amount of reassurance.
- Why now. The honest reason, which is usually that the current sample is too small to be fair to anyone.
Then show them the tool. An actual scored conversation on a screen, with the reasoning visible, defuses more anxiety than a paragraph of policy. The ongoing version of this argument, the one you keep having after launch, is covered in how to run a QA program agents do not hate. This section is only about the one conversation you have before anything starts.
Why your best team is the wrong pilot cohort
Almost every rollout pilots on the strongest team. It is the obvious choice: they are cooperative, they are good at their jobs, and the numbers will look fine. That is exactly why it is the wrong choice.
A high performing team produces a misleading result in four specific ways. They pass nearly everything, so the scoring distribution is flat and you cannot tell whether the tool discriminates. They handle the cleanest queue, so the messy ticket types that break your configuration never appear. They rarely disagree with a score, so the dispute path is never tested and you go live with an untested dispute path. And when you widen the rollout, the wider team can see that the pilot ran on the group least likely to be hurt by it, which costs you exactly the credibility the pilot was supposed to buy.
Build the cohort deliberately instead
Eight to fifteen agents, and pick them for coverage rather than for goodwill:
- Mixed tenure. At least one person in their first three months and one with years on the queue. New agents surface where the rubric is unclear; long-tenured agents surface where it has drifted from reality.
- Your messiest ticket type. Refunds, escalations, whatever the queue is where conversations run long and policy gets bent. This is where configuration breaks.
- One credible sceptic. Someone respected who will say what everyone else is thinking. A pilot with no complaints in it has not tested anything, and a sceptic who is won over is worth more than ten volunteers.
- The reviewer who is losing work. Whoever currently does the manual scoring for that queue. If they are not in the pilot, they will quietly keep their spreadsheet and you will not find out for a quarter.
Which conversations get scored during the pilot matters as much as who is in it, and the selection rules are the same ones that govern QA generally. If the pilot scores a hand-picked set, agents will read the result as hand-picked. How conversations get selected for monitoring covers why.
What running the old and the new process at once actually costs
You will run both processes side by side for a while. That is correct, and every vendor recommends it. Almost nobody tells you what it costs, which is why so many teams abandon it in week four and lose the comparison they were building.
Take a team of forty agents doing around twelve thousand conversations a month, with four manual reviews per agent. Here is where the QA function’s hours actually go.
| Activity | Before rollout, hours per month | During the parallel run | After go-live |
|---|---|---|---|
| Manual reviews and feedback | 40 | 40, unchanged, this is the point | 8 |
| Checking the tool’s scores against your reviewers’ | 0 | 10 | 2 |
| Configuration, rubric mapping and retuning | 0 | 24 | 3 |
| Handling disputes and questions | 1 | 6 | 4 |
| Reporting, and explaining the two sets of numbers | 2 | 6 | 2 |
| Total | 43 | 86 | 19 |
Who pays for the parallel period, and how long it lasts
Roughly double, for six to ten weeks. That is the honest number, and it is the line item that gets left out of every auto QA business case. Two consequences follow.
Somebody has to stop doing something. The hours do not appear. In practice this means pausing a coaching cycle, cutting the manual sample for the queues not in the pilot, or borrowing a team lead for two days a week. Decide which before you start, because if you do not, the thing that gets dropped is the comparison itself, and then you go live with no evidence and no answer when someone asks whether the tool agrees with your reviewers.
Your reviewers are doing double work for the tool that changes their job. That is a real grievance and it deserves a real answer. Name the end date, show them the after column, and make sure they are the ones who get to say when agreement is good enough to stop. A reviewer who signed off on the switch will defend it. A reviewer it happened to will not.
How long the parallel run needs to last is a measurement question rather than a calendar question: you stop when per-criterion agreement between the tool and your reviewers has stopped moving, not when the eight weeks are up. The method for that, including what to measure and what a healthy disagreement pattern looks like, is set out in how to validate AI QA scoring, and how accurate AI QA is covers what agreement levels are realistic to expect. Run the comparison with the automated result hidden from agents and attached to no consequence until it converges.
Scores drop in the first month, and it is usually not the tool
Expect the average score to fall when you go live. It happens to almost everyone, and it is the single most dangerous moment in a rollout, because it looks exactly like the tool failing.
Three things change at once, and none of them is quality.
- The denominator moved. You used to score a small sample. Now you are scoring everything, or close to it. The hard conversations that a reviewer with limited time never got to are now in the average. The old number was not wrong, it was measuring a different, easier population.
- Selection stopped being merciful. Human reviewers unconsciously skip the eleven-message escalation and score the clean three-message ticket. Nobody does this deliberately. It happens because the reviewer has an hour and a queue.
- Consistency arrived. A criterion two reviewers used to enforce differently is now enforced the same way every time. For any criterion that was being applied leniently, that reads as a drop.
The fix is entirely in the sequencing. Brief your leadership on this before the first report, not after. Put it in writing, with a number: tell them you expect the average to fall by some amount in month one and that you will explain the delta rather than defend it. Then report the first two months as two series, the old sample-based score and the new one, so nobody compares across the break. And freeze consequences. No bonus, no performance conversation and no league table runs off the new score until you have a full month of it, plus a reviewer spot-check confirming the drops are real findings.
Underneath all of this sits a statistical point worth understanding before you have to argue it: a small sample was never precise enough to support agent-level decisions in the first place, which what QA sampling is and where it breaks makes uncomfortably clear. The new number is not worse. It is the first one that was ever defensible.
What happens to historical scores under a new scorecard
Nobody writes about this, and it comes up in week one of every rollout. You have two or three years of scores in spreadsheets. The new tool has a scorecard that is similar but not identical. Somebody asks whether the history can be brought across.
Do not retro-apply the new scorecard to old conversations. Not to prove the tool works, not to build a trend line, not because it is technically possible. An agent’s March score was produced under March’s rules by March’s reviewers. Recalculating it in September destroys the one property a QA number depends on, which is that it meant something at the time it was given. It also produces a trend line that measures your scorecard changing rather than your team improving, and someone will eventually present that chart to leadership as evidence of something.
Four rules that keep the record honest:
- Freeze and keep the old scores readable. Export them somewhere durable before the old system is switched off. You will need them for a dispute, an audit or an appraisal.
- Label the scorecard version that produced each score, with the date range it was in force. If you cannot say which rubric produced a number, the number is not evidence.
- Start the new trend line from a clean date and say so on every chart that crosses it. A vertical line and a one-line note is enough.
- Backfill only for calibration, never for the record. Rescoring old conversations to test the tool is useful. Publishing those results as an agent’s history is not.
The general version of this rule, and why it applies whenever the criteria or their weights change rather than only at a tool switch, belongs with the QA scorecard itself. What a tool has to give you here is version history: which rubric produced a score, when it changed, and what the reasoning was on the specific conversation. Kaizo’s Auto QA keeps the criterion, the version and the exchange that triggered the deduction attached to each scored conversation, so a score from six months ago can still be explained rather than just retrieved.
A realistic rollout timeline
Twelve weeks, assuming a named owner and a scorecard that already exists and is not being argued about. Add four to six weeks if either of those is untrue. Anyone promising two weeks is describing the configuration, not the rollout.
The effort column is what the QA function actually spends. It does not include the agents’ time, which is small but real, or the admin work on the helpdesk side.
| Phase | Weeks | Who is doing the work | Rough effort | Do not move on until |
|---|---|---|---|---|
| Configure and connect | 1 to 2 | QA lead plus a helpdesk admin | 15 to 20 hours total | The scorecard in the system matches the one you use today, criterion for criterion, and you have checked it on ten real conversations |
| Announce to the team | 2 | Support lead, with the QA lead in the room | 3 hours | Every agent has heard it from their own manager rather than from a channel |
| Calibrate the tool against your reviewers | 3 to 6 | QA lead plus two reviewers | 10 hours a week | Per-criterion agreement has stopped moving, not when the weeks are up |
| Pilot cohort live, scores visible, nothing counts | 5 to 8 | Pilot team plus QA lead | 8 hours a week | Disputes have actually been filed and resolved. Zero disputes means the path is untested |
| Widen to the remaining queues | 9 to 11 | QA lead | 6 hours a week | Two consecutive weeks with no new configuration defects |
| Retire the old process | 12 onward | QA lead | 4 hours a week, falling | Leadership has seen both series side by side and understands the break |
Six ways a rollout fails, and the signal that shows up first
Each of these is recoverable if you catch it in the first month and expensive if you do not. The signal matters more than the description, because the signal is what you can actually check on a Friday afternoon.
- The shadow spreadsheet. A team lead keeps scoring in the old sheet because it is faster and they trust it. Signal: review volume in the tool is lower than the number of coaching conversations happening. Somebody is scoring somewhere else.
- The pilot that never ends. The pilot goes well, and then nothing happens for two months because nobody owns the decision to widen. Signal: the widening date has moved twice without a new date being set.
- Configuration debt. Criteria were mapped roughly during setup with an intention to fix them later. Later never arrives, and by month three the scores are wrong in a way everybody has stopped mentioning. Signal: reviewers routinely override the same criterion, or a criterion passes on essentially every conversation.
- Calibration quietly stops. The tool is consistent, so the reviewer sessions get dropped as redundant. Consistency is not correctness, and without calibration sessions nobody notices when the tool and the humans have drifted apart. Calibration is not made redundant by automation, it is what checks it. Signal: no calibration has been run since go-live.
- Scores nobody can explain. An agent asks why they lost points, their manager cannot answer, and the manager stops using the score rather than admit that. Signal: QA scores disappear from one-to-ones while the reviews keep being generated.
- The owner leaves. The person who ran the rollout changes role and it turns out the whole thing lived in their head. Signal: nobody else can say when the scorecard was last changed or why.
One check catches four of the six. Once a month, pick five scored conversations at random and ask the agent’s manager to explain the score without opening the vendor’s documentation. If they cannot, the rollout has not landed yet, whatever the adoption dashboard says.
Frequently asked questions
How long does it take to implement a QA tool?
About twelve weeks from signed contract to retiring the old process, if you already have a scorecard and a named owner. The configuration itself takes one to two weeks. The remaining ten are calibration, the pilot, the parallel run and widening. Add four to six weeks if the scorecard is still being designed or if nobody owns the project. Two-week implementations are describing the technical setup only.
Should we tell agents before rolling out a QA tool?
Yes, and before the first score lands rather than alongside it. If you say nothing, the first version of the story that reaches the floor is the worst one, and you will spend the rollout arguing with it. The announcement should come from the support lead, state what is changing and what is not, give a date when scores start counting, explain how to challenge a score, and be followed by showing the team an actual scored conversation on a screen.
Which team should pilot a new QA tool?
Not your best one. A high-performing team passes almost everything, handles the cleanest queue and files no disputes, so you learn nothing about where configuration breaks or whether the dispute path works. Build a cohort of eight to fifteen agents deliberately: mixed tenure, your messiest ticket type, one credible sceptic, and the reviewer whose manual scoring the tool is replacing.
Why did our QA scores drop after switching tools?
Usually because scrutiny increased, not because quality fell. Three things change at once: you are scoring far more conversations so the hard ones a reviewer never reached are now in the average, selection stopped skipping the long escalations, and a criterion two reviewers enforced differently is now enforced the same way every time. Report the first two months as two separate series and freeze consequences until you have a full month of the new score plus a reviewer spot-check.
Should we recalculate old QA scores under the new scorecard?
No. A score produced under last year’s rubric by last year’s reviewers meant something at the time it was given, and recalculating it destroys that. You also end up with a trend line that measures your scorecard changing rather than your team improving. Freeze and export the historical scores, label which scorecard version produced them, start the new trend line from a clean date, and rescore old conversations only to calibrate the tool, never to restate someone’s record.
How long should we run manual QA and automated QA in parallel?
Six to ten weeks in most cases, but the stopping condition is measurement rather than calendar: stop when per-criterion agreement between the tool and your reviewers has stopped moving. Budget for it properly, because the parallel period roughly doubles the QA function’s monthly hours. Decide in advance what gets paused to fund it, otherwise the comparison is what gets dropped and you go live with no evidence.
Related terms
See where your scores land before a single agent does
Bring the scorecard you use today and one month of conversations your team has already reviewed. We will run them through the same criteria so you can see the size of the month-one shift, and brief your leadership on it, before anything goes live.