When a support team scales, quality rarely falls because the people got worse. It falls because the review method stayed the same while the volume behind it multiplied, so the sample stops representing the work, reviewers quietly drift apart, and the quality number becomes impossible to argue with or against. The repair is sequenced rather than simultaneous: at roughly 10 agents you write the standard down, at 30 you fix sampling and calibration, and at 100 you make the score segmented and traceable.
In short
- Scaling breaks QA, not size. A 40-person team that has been 40 people for two years is usually fine. A 20-person team that was 8 people in March is not.
- The first thing to go is not the score, it is the shared definition of good. Once that lives in several heads and nowhere on paper, every later fix is built on sand.
- A fixed percentage sample is not a fixed level of confidence. Reviewing 3% of 200 conversations a week and 3% of 20,000 answer completely different questions.
- Reviewer drift is silent, and it is the most common reason a quality trend turns out to be wrong. Two people scoring the same conversation differently is not a rounding error.
- The tell that you have outgrown the spreadsheet is not the row count. It is that nobody can reconstruct why a given conversation got the score it got.
- Change one thing per threshold. Teams that rebuild the rubric, the sample, the tooling and the coaching cadence in one quarter cannot tell afterwards which change worked.
Our support team doubled and quality is slipping. What actually broke?
Almost nothing broke. A method that was quietly working stopped scaling, and the reporting carried on as though it had not. That is the honest diagnosis for most teams in this situation, and it matters because it points at a different set of fixes than the ones people reach for first.
Here is the usual shape. When the team was small the lead read a lot of tickets without calling it review, and standards passed between people sitting next to each other. Quality was genuinely known, even though it was never measured.
Then headcount grew. The lead now has more reports, more escalations and more meetings, so the informal reading shrinks to a handful of conversations a week, while a quality percentage still ships into a deck every month. It looks stable. It is stable because it is produced from a sample too small and too casually chosen to move.
The tell is that the number no longer surprises anyone. A quality metric that has not told you something you did not already know in six months is not measuring quality, it is confirming a prior.
Separate two problems before fixing either. Capacity problems are about handling more volume with the people you have, and there is plenty of decent advice on them, including our own on strategies for scaling a support team. Measurement problems are about whether you still know what is happening in the conversations. Growth creates both, and teams reliably fix the first and assume the second came free.
What breaks first, at around 10 agents
At this size the failure is definitional, not statistical. Volume is still small enough that a lead could read a meaningful share of it. What has gone missing is agreement on what a good conversation looks like now that more than two or three people are having them. The symptoms are specific.
- Two people give the same customer opposite answers and both believe they were right.
- Feedback is delivered as personal preference, because there is nothing else to point at.
- New starters take longer to become useful than the last cohort did, and nobody can say why.
- The lead is the only person who knows whether a given ticket was handled well.
Change exactly one thing here: write the standard down. Not a programme, not a tool, not a target. A short written definition of what you are judging and what each level looks like, which is what a QA rubric is for. Four or five criteria is plenty. Teams that launch a full QA programme at ten agents usually produce a process nobody has time to run.
Do it now rather than later because everything downstream depends on it. Sampling, calibration, coaching and reporting are all operations performed on a definition. If the definition arrives after the tooling, you spend the next year discovering that your historical data measured something you have since stopped believing in, and that ramping new joiners against it produced a second cohort worse than the first.
What breaks at around 30 agents, and why the sample is the culprit
At thirty agents the definition usually exists and the measurement stops being trustworthy. This is where most teams first notice the problem, and where they most often misdiagnose it as a people problem. Two mechanisms are at work and they compound.
The first is that the sample stopped being representative. A fixed percentage feels like a constant, but it is not a constant level of confidence. Three percent of a hundred and fifty conversations a week is four or five, which tells you almost nothing about an agent and nothing at all about a channel. Worse, the sample is usually not random: reviewers pull what is visible, meaning escalations, complaints and the tickets already on someone’s mind. Useful to read, terrible to average. If sampling design is new to you, start with what QA sampling actually is, then work out how many conversations your own team needs; the NIST handbook section on sample size for a proportion gives the arithmetic in one page.
The second is reviewer drift. One reviewer is internally consistent by definition. Three reviewers are three standards wearing the same rubric, and the divergence hides inside the reported average because it cancels out. Grader agreement is a measured thing, not an assumed one, and the standard benchmarks for how much agreement is enough are older than the support industry. The practical version of this is running calibration sessions, which at thirty agents moves from nice-to-have to load-bearing.
Two supporting fixes belong here rather than earlier. Revisit how your scorecard criteria are weighted, because an unweighted scorecard lets a greeting failure cancel a compliance failure. And check whether your target is doing any work: a 95% quality target everyone clears every month is a ceiling, not a standard.
This is also the first threshold where automated scoring earns its place, and the reason is not effort. Scoring all of the conversations rather than a slice removes the sampling argument from the room, which is what automated QA is for. It does not remove the calibration argument, and you should be suspicious of anyone who says it does.
What breaks at around 100 agents
Past a hundred agents the score has to survive contact with people who were not in the room. The problems from thirty do not go away, they acquire an audience. Three new failures show up.
The aggregate hides everything worth knowing. One company-wide quality percentage across four queues, three languages and two shift patterns is an average of populations that have nothing to do with each other. It moves when the mix moves, which reads as a quality change and is not one. Segmentation becomes the only way the number stays honest.
Coaching becomes a distributed system. At thirty agents one person holds the whole picture. At a hundred it is held by eight or ten team leads who each see a tenth of it, and who were promoted for being good at the job rather than for reading quality data. What they do with a score is now the highest-leverage thing in the programme, which is the argument in enabling team leads to use QA. A lead who reads a score and cannot turn it into a specific conversation is a reporting line, not a coach.
The number gets challenged. Somebody outside support asks why it moved, or an agent disputes a score in a performance review, and the programme either produces the evidence or loses the argument permanently. This is why score traceability matters more than coverage at this size: showing which conversations produced a number, and why each was graded the way it was, is what survives scrutiny. Coverage is table stakes and every vendor claims it. Proving a grader was right is not.
Reporting is the other half. A support lead, a VP and a finance audience need three different views of the same data, which is the whole of reporting support quality to executives. Teams that get budget at this stage are the ones whose quality data is already denominated in something the reader is accountable for.
What to change at each threshold, and what to leave alone
The most common way this goes wrong is doing all of it at once. A team notices quality slipping and in one quarter rewrites the rubric, changes the sample, buys a tool and changes the coaching cadence. Quality moves. Nobody can say which change moved it, so nothing is learned and the next slip is handled by guessing again.
Find your row, do that row, and deliberately do not do the rows below it yet. The last column is the important one.
| Roughly | What breaks | The tell | Change this | Not yet |
|---|---|---|---|---|
| Under 10 agents | Nothing structural. Informal reading still covers the volume | You can name what every agent is currently bad at, without looking it up | Nothing. Keep reading tickets and keep notes | A programme, a tool, a target |
| Around 10 agents | The shared definition of good stops being shared | Two agents give opposite answers and both are confident | Write the rubric. Four or five criteria, with what each level looks like | Sampling design, targets, tooling |
| Around 30 agents | The sample stops representing the work, and reviewers drift apart | The reported number has not surprised anyone in six months | Fix how conversations are selected, then calibrate the reviewers against each other | Executive dashboards, per-agent league tables |
| Around 100 agents | The aggregate hides the segments, and the score gets challenged | Someone asks why it moved and the answer takes two days to assemble | Segment the reporting, and make every score traceable to its evidence | Another reporting layer before traceability exists |
Should we hire someone to run QA, or just hire another agent?
The question is not whether reviewing is worth someone’s time, it is whether it is worth a whole person’s time before reviewing is repeatable. Hiring a reviewer into an undefined standard produces one person’s opinion at higher cost. A workable sequence, in order of what it costs you.
- Give it to a lead as named time, not as goodwill. Four hours a week, in the calendar, with the rubric written. If it does not survive a busy month, that is data.
- Split it across leads with calibration. Cheaper than a hire, and it builds the muscle where coaching actually happens, but only if the leads are measured against each other rather than trusted individually.
- Hire the dedicated reviewer. Justified when review is a real backlog with a real queue, when there are enough segments that nobody has the whole picture, or when an external obligation puts a floor under how much must be reviewed.
One caution about the second option. Leads reviewing their own reports quietly biases the number, because the reviewer is grading the outcome of their own coaching. Cross-review between teams costs nothing extra and removes the conflict.
Whatever you choose, what decides whether the programme sticks is not who does the scoring. It is whether the people being scored believe the scores, which is a separate and harder problem covered in running a QA programme agents trust.
When should we stop running QA in a spreadsheet?
Not when the file gets big. When nobody can reconstruct why a conversation got the score it got. Row count is a bad trigger, because spreadsheets stay usable long past the point where they stop being defensible. Four tells, any one of which is sufficient.
- Somebody disputes a score and answering means opening three tabs and remembering a conversation from six weeks ago.
- You cannot tell whether a movement is real without manually recounting the sample.
- The rubric changed and the historical rows now silently mean something different.
- Two reviewers have started keeping private tabs, because the shared one does not fit how they work.
Marley Spoon is the clearest version of this in our own customer base: as the support team grew tenfold across three time zones, the spreadsheet went, not because of its size but because a distributed team could not keep one standard alive inside it. Procede Software reported a 13% annual increase in CSAT after moving review off manual sampling.
If you do move, treat it as a change-management exercise rather than a purchase. The failure mode is not picking the wrong tool, it is rolling it out so that agents read it as surveillance, which is what a QA tool rollout plan exists to prevent. And the data it produces has to be readable outside the QA team, which is the job of your support analytics layer rather than of the scoring itself.
How do you know the quality number is real and not noise?
Ask what would have to be true for the number to be wrong, then check whether it is. Three checks catch almost everything, and none needs a statistician.
Check the denominator before the movement. A two-point swing on forty reviewed conversations sits well inside the range you would get from a different forty. Work out how many you would need for the movement to be distinguishable from chance, and if that is more than you review, say so in the report rather than after someone else notices.
Check the mix. If queue composition changed, a new market went live, or a bot started absorbing the easy contacts, the conversations left for humans got harder and the score should have fallen with no change in performance. Segmented reporting makes this visible; a single aggregate does not.
Check the graders. Take one conversation, have every reviewer score it blind, and look at the spread. If they disagree by more than a grade on a criterion, that criterion is not measuring anything yet and volume will not fix it.
Only then does the trend deserve a causal story, which is what measuring conversation quality at scale depends on: a number is only as good as your ability to say where it came from.
Frequently asked questions
How do you know if you need to scale your support QA?
Two signals, and you need only one. The reported quality number has not surprised anyone in six months, which usually means the sample is too small or too selectively chosen to move. Or you cannot say, without going and looking, which part of the team is currently weakest and why. If a lead can still answer that from memory, informal review is working and you should leave it alone.
Our support team doubled and quality is slipping. Where do we start?
Work out which threshold you have crossed rather than fixing everything. If the team is now around ten people, the standard was probably never written down, so write the rubric and stop there. If it is around thirty, the standard exists and the measurement has stopped being trustworthy, so fix how conversations are selected for review and then calibrate the reviewers against each other. Doing both in one quarter means you will not know which worked.
How many conversations should we review as the team grows?
Stop thinking in percentages. A fixed percentage gives a shrinking level of confidence per agent as the team grows, because the sample spreads across more people. Decide instead what you want to be able to say, for example that you can talk to an individual agent about a pattern with a straight face, then work backwards to conversations per agent per period. For most teams that lands between eight and twenty per agent per month, a far larger absolute number than three percent of the queue produces once the queue is big.
At what team size do you need a dedicated QA person?
Later than most people assume, and the trigger is workload rather than headcount. Named time in a lead’s calendar, with a written rubric, covers most teams up to roughly thirty agents. A dedicated hire is justified when review has become a real backlog with a queue of its own, when there are enough segments or languages that nobody holds the whole picture, or when an external obligation puts a floor under how much must be reviewed. Hiring before the standard is written just buys one person’s opinion at a higher price.
Does automating QA fix the problem of scaling?
It fixes one half. Scoring all of the conversations instead of a slice removes the sampling argument, which is the mechanism that breaks around thirty agents. It does not remove the need for an agreed definition of good, and it does not remove calibration, because an automated grader still has to be checked against human judgement and shown to be right. Coverage is the precondition now, not the differentiator. Being able to show why a conversation received the score it did is the part that holds up under challenge.
What does it mean to scale customer support without losing quality?
Keeping three things intact while volume and headcount grow: a written definition of what good looks like, a review sample that represents the work rather than the loudest tickets, and reviewers who agree with each other. Most advice on scaling support is about capacity, which is a different problem. You can double throughput and lose all three, and the reporting will not tell you until something visible goes wrong.
Related terms
- How to build a QA programme from scratch
- What QA sampling is and how to choose a sample
- How to run QA calibration sessions
- How to weight QA scorecard criteria
- Reporting support quality to executives
- Running a QA programme agents trust
- Turning QA data into coaching
- Onboarding and the quality ramp
- Keeping quality honest through peak
- When to stop running your QA programme in a spreadsheet
Find out which threshold your team has actually crossed
Bring one month of your own conversations. We will show you what your current sample can and cannot support a claim about, where your reviewers disagree, and the one change worth making at the size you are now.