What a Team Lead Actually Does With QA Scores Each Week

Most team leads coach the score instead of the cause. A weekly method for reading QA results, choosing who to coach, and knowing when to coach nobody at all.
Playbook · Agent Coaching

A team lead’s job with QA scores is not to review them, it is to convert them into one or two behaviour changes a week. That means reading the week’s results for causes rather than totals, deciding who genuinely needs coaching instead of coaching everyone equally, and being willing to say that a week produced nothing worth coaching. The QA reviewer owns whether a score is correct. The lead owns what happens next.

In short

  • A score tells you what happened. It does not tell you why, and the why is the only thing you can coach.
  • Read the week in a fixed order: the whole team first, then movement, then individuals. Most leads do it backwards and coach a rubric problem to eleven people one at a time.
  • Most of the improvement available in any given week sits with two or three agents and one broken process. Spreading coaching evenly spreads it thin.
  • Some weeks there is genuinely nothing to coach. Saying so out loud protects the programme. Manufacturing a development point to fill the slot destroys it faster than silence ever would.
  • The reviewer decides whether a score is right. The lead decides what to do about it. Leads who blur that line spend the week re-litigating scores instead of using them.
  • When scoring is automated the bottleneck moves. Finding problems stops being hard and choosing which two are worth the week becomes the whole skill.

Why a score is not yet something you can coach

An agent comes back at 78 on resolution accuracy, three weeks running. The obvious move is to open the one to one with the number: your accuracy is 78, let us get it to 85.

That is coaching the score, and it almost never works, because 78 is a symptom with at least four different diseases behind it. The agent may not know the current policy. The policy may have changed without reaching the floor. The right answer may be buried three clicks deep in a tool nobody has time to search mid-conversation. Or the agent may be reading the criterion differently to the reviewer, in which case the score is measuring a disagreement about the rubric rather than a performance.

Each of those has a different fix, and only one of them is yours.

What the score says What you may find underneath Who actually fixes it
One agent’s resolution accuracy is drifting down They do not know the current policy You, through coaching. This is the one that is genuinely yours
One agent’s resolution accuracy is drifting down The policy changed and the floor was never told Enablement or the policy owner. Coaching the agent here is unfair and it will not hold
Tone is marked down repeatedly, same agent They are getting hit at a specific hour with no recovery time Scheduling first, then you. Coaching someone out of queue pressure does not work
The same criterion is failing across most of the team The criterion means one thing to the reviewer and another to everyone else QA, through calibration and a rubric fix. Never coach a rubric problem person by person
One agent, one bad week, no pattern A genuinely off week, or three brutal tickets routed their way Nobody. Note it and look again next week

The question that turns a score into a coachable insight

There is one question that does most of the sorting, and it takes about ten seconds per case.

What would have to be true for this to be the agent’s to fix?

If the honest answer involves another team, a policy, a tool, a staffing decision or an ambiguous rubric line, it is not a coaching point. It is an escalation with a coaching point stapled to it, and if you coach it anyway you will burn credibility on something the agent cannot change.

Two habits make this question answerable rather than theoretical.

  • Open the conversation before you form a view. Not all of them. One or two, the ones the score points at. A criterion failure looks completely different once you have read the exchange that produced it, and about a third of the time the read changes your conclusion.
  • Say the cause out loud before you say the fix. If you cannot state the cause in one sentence without using the word score, you do not have one yet. “Your accuracy is low” is not a cause. “You are quoting the pre-June refund window” is.

Once you have a cause, the handoff into an actual conversation is a separate discipline with its own steps, covered in how to turn QA data into coaching. This page is about getting to the point where you have something worth handing off. The underlying number itself, and what it is and is not capable of telling you, is covered in what an internal quality score measures.

What to look at first when the week’s results land

The order matters more than the tooling. Most leads open the lowest-scoring individual first, because that is where the discomfort is. That is the wrong end. Work outside in, in four passes, in about forty minutes.

Pass 1: the whole team, one criterion at a time

Before you look at a single person, look at the criteria on the scorecard. You are hunting for one where the failure rate is high across many agents rather than concentrated in a few. That is a systemic finding and it is the most valuable thing on the page, because fixing it once fixes it for everyone, and coaching it individually would have cost you eleven conversations and taught nobody anything. If the pattern is about why customers are contacting you rather than how agents are handling it, that belongs in root cause analysis rather than in anyone’s one to one.

Pass 2: movement, not level

Now sort by change rather than by score. A steady 82 tells you very little. An 88 that fell to 79 tells you something happened, and it happened recently enough that the agent still remembers the week. Movement is coachable in a way that a stable low number rarely is, because a stable low number usually means a training gap or a role mismatch, and neither of those is a Tuesday conversation.

Pass 3: the bottom of the table, with the ticket mix in view

Only now look at who is lowest, and do not look at the score on its own. Check what they were handling. The agent who takes the escalations, the billing disputes and the angry reopens will score below the agent who takes password resets, and coaching them for it is how you teach a team to avoid hard work.

Pass 4: the top, for one thing worth copying

Pick one specific behaviour from a high scorer that you can name and describe. Not “be more like Sam”. Something like “Sam confirms the refund amount and the date in the same sentence, every time, and it kills the follow-up ticket”. That is a reusable asset and it costs you five minutes.

Forty minutes, once a week, and you now know the one systemic issue, the two or three people worth your time, and one thing worth spreading. Everything after this is execution.

Coaching everyone equally is the most expensive habit in the job

It looks like fairness. Twelve agents, twelve slots, everybody gets their turn and nobody can say they were singled out. It is also the single most reliable way to convert a QA programme into a calendar obligation. The stakes are higher than the calendar suggests, because Gallup estimates that managers account for at least 70% of the variance in employee engagement across business units, so how a lead spends this time is one of the larger levers in the operation.

The arithmetic is unforgiving. A team lead with twelve agents has perhaps eight hours a week of genuine coaching capacity once queues, escalations, staffing and their own meetings are paid for. Split evenly that is forty minutes per person, most of which goes on recapping numbers the agent has already seen. Concentrated, it is two real conversations that change something, one systemic issue fixed at the source, and one behaviour publicly named and spread.

Sort the week into four buckets instead.

Bucket What puts someone here Typical size What you actually do
Coach A repeated, specific behaviour that is genuinely within their control 1 or 2 people One behaviour, one agreed action, evidence open on the screen
Watch Something moved, once. Could easily be noise or ticket mix 2 or 3 people Nothing this week. Write the name down and look again next week
Escalate, do not coach The cause sits with a policy, a tool, staffing or the rubric Whatever the week produced Write it up, send it to the owner, tell the agents you did
Reinforce Something specific and good, worth other people hearing 1 person Say it in public, name the behaviour, not the person’s general excellence

Two rules that stop the buckets going wrong

The buckets are weekly, not permanent. They describe what this week’s evidence supports, not who is good and who is not. An agent who lands in Watch four weeks running has stopped being noise and has become a Coach, and an agent who has been in Coach for two months is not a coaching case any more, they are a training or a role conversation.

Rotate the Reinforce bucket deliberately. If the same two names appear in Coach every week and never anywhere else, those two agents are not receiving coaching, they are receiving surveillance, and they will describe it that way to everyone else on the team. This is the fastest way to lose a floor’s cooperation with QA, and it is entirely avoidable.

When you get to the conversation itself, the structure that keeps it from drifting into a numbers recap is a separate artifact, and there is a working agenda in what a coaching framework is. The buckets decide who is in the room and why. The template decides what happens once they are.

What to do in a week when there is genuinely nothing to coach

This is common and almost every programme handles it badly.

On a stable team with a settled scorecard, a real proportion of weeks produce no individual coaching point worth an agent’s time. Nothing moved. Nothing repeated. The one thing that did look odd turned out to be a hard ticket handled reasonably. That is not a failure of the programme, it is the programme working.

What happens instead is that the calendar says the one to one is at eleven and the template has a field labelled development area, so the lead fills it. Somebody gets told their greetings could be warmer. This is manufacturing, and it costs three specific things.

  • Feedback stops being a signal. Once an agent has worked out that a development point arrives every single week regardless of what they did, they stop treating any of them as information. You have trained them to discount you.
  • The real ones get buried. When week nine contains an actual problem, it arrives in the same slot, in the same tone, with the same weight as the eight invented ones before it.
  • The scorecard starts to look arbitrary. If the agent scored 94 and you still found something, the obvious inference is that the number was never the point and you were going to find something anyway.

Four things that are all better than inventing one

  • Say it plainly. “Nothing came up this week that is worth your time. Here is the thing that looked good.” Ninety seconds. This is the single most trust-building sentence available to a team lead and it costs nothing.
  • Give the slot back. Cancelling a one to one because there is nothing to say is not a failure of management. Holding one to fill the calendar is.
  • Turn it around. Ask what is getting in their way, what the tooling makes harder than it should be, which ticket type they dread. You will learn more in that fifteen minutes than in the previous four sessions.
  • Spend the time on the systemic issue. There was one on your team-level pass. Nobody is coaching it because it is not a person. Go and fix it.

One caution, because this can fail in the other direction. “Nothing to coach” is a conclusion you reach after doing the weekly pass, not a reason to skip it. Skipping the pass and calling the result nothing to coach is the same neglect wearing better clothes. Do the forty minutes, then be willing to conclude that the answer is no.

Where the reviewer’s job ends and the lead’s begins

These two roles get confused constantly, and the symptom is unmistakable: the lead spends the week arguing about whether scores were right instead of doing anything with them. The programme grinds, QA gets defensive, and no agent changes any behaviour.

Split it by question rather than by job title.

The question Whose call it is Where it gets settled
Was this conversation scored correctly? The QA reviewer The review itself, and the dispute route
Do two reviewers score this the same way? QA A calibration session, not a one to one
Is the rubric measuring the right thing? QA, with the lead as a loud input Scorecard review, quarterly
Why does this agent keep failing this criterion? The team lead Reading the conversations, then coaching
What is the one thing this agent changes next week? The lead, agreed with the agent The one to one
Is this score reliable enough to act on at all? Both, once A dispute route, used sparingly and on the record

The two ways leads get this boundary wrong

The second reviewer. This lead re-scores everything before they will use it. They have effectively doubled the cost of the programme, they are doing a job someone else is already paid for, and they have no time left to coach. If you genuinely cannot trust the scores, that is a calibration problem and it is fixed by running a calibration session, not by quietly running a shadow QA function out of your own week.

The rule worth holding: dispute a score once, through the route, and then act on the corrected picture. Re-opening scores inside a one to one teaches the agent that the number is negotiable and puts you in the strange position of defending someone else’s work instead of doing your own.

The postman. The opposite failure. This lead forwards the score with “please review and improve” and adds nothing. The agent receives a number and an instruction, has no cause and no example, and changes nothing. The interpretation step is the entire value the lead adds, and it is the step that gets dropped first when the week is busy.

If the split still feels blurry, it helps to look at what the reviewing role is actually accountable for, which is set out in what a QA analyst does. Everything not on that list is yours.

What changes when every conversation is scored

Under sampled review the lead’s scarce resource was evidence. A few percent of conversations reviewed, a handful per agent per month, and half a day spent hunting for an example concrete enough to make a point stick. The common complaint was that the sample missed the thing you knew was happening.

Automated scoring inverts that, and the change is bigger than it first looks. The evidence problem goes away and an attention problem takes its place. Three consequences worth planning for.

  • Selection becomes the skill. The job stops being “find a problem” and becomes “choose which two of the eleven visible problems are worth this week”. That is a harder judgement and nobody is trained for it, which is why leads who get full coverage often report feeling less in control rather than more for the first month.
  • One-offs get less interesting and patterns get much more so. With 100% coverage revealing trends that 3% sampling never could, a single poor conversation stops being worth a meeting. Three of the same failure in a month is worth a meeting, and now you can actually see the difference between the two.
  • Every deduction has to be traceable or you cannot use it. An agent asked to change something will ask which conversation. If you cannot open it and point at the exchange, the coaching does not land, and it should not. Coverage on its own is now table stakes across the category. Being able to show why a grader marked something down, and to have that survive a challenge, is the part that makes a score usable by a team lead at all.

Kaizo’s AI coaching is built around that last point: the pattern is surfaced with the underlying conversations still attached to it, which is where EverHelp reported a 90% reduction in coaching prep time. Note what that does and does not remove. The preparation is the part that automates. Choosing which two people are worth the week, and having the conversation, are still yours, and there is no version of this where they are not.

If your programme is moving from a sample to full scoring, it is worth reading how to score 100% of conversations alongside this, because the operational shift lands on the team lead before it lands anywhere else.

Frequently asked questions

How to be a good supervisor in a call center?

Convert evidence into change rather than reporting it. Concretely: run one weekly pass over the QA results in a fixed order, team-level patterns before individuals, decide which two or three people are genuinely worth coaching that week, escalate anything whose cause sits with a policy, tool or staffing decision, and be willing to tell an agent there is nothing to coach when there is nothing to coach. The supervisors agents rate highly are not the ones who give the most feedback, they are the ones whose feedback turns out to be worth listening to.

How do you coach an agent on a low QA score?

Do not coach the score. Find the cause first by opening one or two of the conversations behind it and asking what would have to be true for this to be the agent’s to fix. If the answer involves a policy, another team, a tool or an ambiguous rubric line, it is an escalation rather than a coaching point. If it genuinely is theirs, name one specific behaviour, show the exchange that demonstrates it, and agree one action. One behaviour per session, with evidence on screen.

What should a team lead do if there is nothing to coach that week?

Say so. Tell the agent that nothing came up worth their time, name one thing that looked good, and give the rest of the slot back. Inventing a development point to fill the calendar is the fastest way to teach a team that your feedback is a ritual rather than a signal, and it means the real coaching point in week nine arrives with no weight behind it. The important caveat is that nothing to coach must be a conclusion you reached after doing the weekly review, not a substitute for doing it.

What is the difference between a team lead and a QA analyst?

The QA analyst owns whether a score is correct, whether reviewers agree with each other, and whether the rubric measures the right thing. The team lead owns what happens next: why a pattern exists, which behaviour to change, and what the agent commits to. Leads who cross that line end up re-scoring conversations, which doubles the cost of the programme and leaves no time for coaching. Disagree with a score once, through the dispute route, then act on the corrected picture.

How often should a team lead review QA scores?

Weekly, in one sitting of about forty minutes, is the cadence most teams can actually sustain. Daily is noise, because a single conversation rarely justifies a conversation. Monthly is too late, because the agent no longer remembers the week and the systemic issues have had four weeks to spread. Fix the slot in the calendar and protect it, since the pass is what makes the rest of the week’s coaching decisions cheap.

How many agents can one team lead coach properly?

Ten to fifteen is where most support teams land, and the constraint is coaching hours rather than span of control on paper. A lead with twelve agents has roughly eight hours of real coaching capacity a week once queues, escalations and their own meetings are paid for. Spread evenly that produces twelve shallow conversations. Concentrated on the two or three cases the week’s evidence actually supports, it produces change. If a lead has more than about twenty agents, the honest answer is that they are monitoring, not coaching.

See which two coaching points your week was actually worth

Bring a month of your own scored conversations. We will run the same weekly pass against your data and show you the one systemic issue and the two people it says to spend the week on.

Book a demoExplore AI Coaching

Choose your help desk

Not using either? We’ll let you know as soon as we can support your help desk solution.

Kaizo
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.