A QA calibration session, called a call calibration session when the sample is voice, is a recurring meeting where everyone who scores conversations grades the same sample independently, then compares results and agrees on one interpretation of the scorecard. You run it by picking two or three conversations, having each reviewer score them blind before the meeting, opening with the criteria where scores disagreed most, and ending by rewriting whatever rubric wording caused the disagreement. Sessions work best every two weeks, at 45 to 60 minutes, with every active reviewer plus the team leads who use the scores for coaching. The output is not a winning score, it is a clearer rubric and a documented decision that future reviews follow. You measure success by tracking how far apart reviewers are on the same conversation over time, not by how pleasant the meeting felt.
In short
- Calibration exists to make QA defensible to agents. If two reviewers would score the same conversation differently, no agent has a reason to accept either score.
- Score blind before the meeting. If reviewers see each other’s scores first, the most senior voice in the room quietly becomes the standard.
- Pick borderline and disputed conversations, not clean examples. Clean conversations produce agreement that teaches you nothing.
- Spend the session on the criteria with the widest spread, not on every criterion in the scorecard.
- The output of a good session is an edited rubric. If nothing in the scorecard changed, you held a discussion, not a calibration.
- Measure agreement with a real number: the percentage of criteria where reviewers scored identically, and the average spread on a criterion. Run the same measurement against any automated scoring you use, treating it as one more rater.
- Persistent disagreement is almost always a rubric problem, not a reviewer problem. Fix the wording before you retrain the person.
Why calibration is the fix for every QA doing their own thing
The most common complaint about quality assurance, from agents and reviewers alike, is that scoring feels arbitrary. One reviewer marks a conversation down for tone, another passes the identical conversation. One takes a hard line on process steps, another treats them as guidance. Agents notice this quickly, and once they do, every score becomes negotiable and every coaching conversation starts with an argument about the number instead of the behavior.
Calibration is the mechanism that fixes this. Put simply, it is a session where everyone who scores conversations grades the same sample independently and then reconciles the differences (the full definition lives in what is QA calibration). Its real purpose is not to make reviewers agree for the sake of harmony. It is to make the score defensible: an agent should be able to ask any reviewer why they lost a point and get the same answer.
That reframing changes how you run the session. You are not there to decide who scored a conversation correctly. You are there to find the places where your scorecard is ambiguous enough that two competent people read it differently, and to close those gaps in writing. Every session should leave a trail: a criterion reworded, an example added, a rule documented.
Who should attend, and how often to run it
Keep the room small and the cadence steady. Calibration fails far more often from irregularity than from bad facilitation.
Who attends
- Every person who scores conversations. This is not optional. A reviewer who skips calibration is by definition uncalibrated, and their scores carry the same weight as everyone else’s.
- The team leads who coach on the scores. They are the ones who have to defend a number to an agent, so they need to have heard the reasoning first hand.
- One owner of the scorecard. Usually the QA lead. This person facilitates, records decisions, and is accountable for actually editing the rubric afterwards.
- An agent representative, occasionally. Rotating a senior agent into one session a quarter is the single fastest way to build trust in QA. They see that the standard is argued over seriously rather than handed down.
How often
Every two weeks is the sweet spot for most teams. Weekly is worth it during the first two months of a new scorecard, or after a significant policy change, because that is when ambiguity surfaces fastest. Monthly is the floor: below that, reviewers drift far enough apart between sessions that you spend the whole meeting relitigating basics. Run sessions per channel or per queue if your rubric differs meaningfully between them, since calibrating a chat conversation teaches you very little about a complex voice escalation.
Time-box to 45 to 60 minutes. A session that regularly overruns is a signal that you are trying to cover too many conversations, not that you need a longer meeting.
How to choose the conversations you calibrate on
Sample selection is where most calibration sessions are lost before they begin. The instinct is to pick a clear example of good work and a clear example of bad work. Both are useless. Everyone will agree, agreement will feel like success, and nothing about the rubric will improve.
Calibrate on the conversations that sit in the grey zone. Practical sources:
- Disputed scores. Any conversation an agent challenged in the last two weeks is a gift. The dispute is direct evidence that the rubric was unclear to someone.
- Outlier scores. Conversations that scored unusually high or unusually low against the norm for that queue, where the reason is not immediately obvious.
- Criteria you already suspect. If empathy, tone or resolution have been contentious, deliberately pick conversations that stress those criteria rather than sampling at random.
- Edge cases from new policy. After any change to process or policy, pick the first conversations where the new rule was tested in the wild.
- Where automated and human scores diverge. If you score every conversation automatically, the gaps between the auto-score and a manual review are a precise, high-volume shortlist of the exact cases where the standard is unclear.
Two or three conversations is the right number for a 60 minute session. Teams that try to cover six end up skimming all of them and resolving none. Send the conversations out at least 48 hours ahead so nobody scores them in a rush ten minutes before the call.
One practical warning: do not let the sample become a collection of disasters. If every calibration conversation is a failure, reviewers start reading the sample as a warning about how harsh they should be, and scores across the whole team creep downward for two weeks afterwards. Mix in at least one conversation that is genuinely borderline in the other direction, where the question is whether it deserves full marks rather than whether it should have failed.
Calibrating on calls rather than tickets
Call calibration follows the same rules with one practical difference: length. A written ticket can be read in the meeting, a 15 minute call cannot. Send the recordings out with the blind scoring request, and in the session replay only the 60 to 90 seconds around each disputed criterion rather than the whole call. Agree in advance whether reviewers are scoring the transcript, the audio, or both, because tone is exactly where reviewers diverge most and where a rubric written for text quietly stops working.
The session agenda, minute by minute
The order matters more than anything else in this article. Reviewers must score blind, individually, and submit before the meeting starts. The moment scores are visible in advance, the most senior or most confident person in the room becomes the standard by default, and you have replaced calibration with deference.
Here is a working agenda for a 60 minute session covering three conversations. Adapt the timings, keep the sequence.
| Stage | Time | What happens | Output |
|---|---|---|---|
| Blind scoring (before the session) | 48h ahead | Each reviewer scores all sample conversations independently, with no visibility of anyone else’s scores | One completed scorecard per reviewer per conversation |
| Spread review | 5 min | Facilitator shares the range of scores per criterion without naming who gave what | A ranked list of the criteria with the widest disagreement |
| Deep dive on the widest gap | 15 min | Reviewers at each end of the range explain their reasoning, pointing to specific evidence in the transcript | The competing interpretations stated out loud |
| Second widest gap | 15 min | Same format, second criterion | The competing interpretations stated out loud |
| Third gap or second conversation | 10 min | Same format, timeboxed hard | The competing interpretations stated out loud |
| Agree the standard | 10 min | Facilitator proposes one interpretation per contested criterion, room confirms or amends | A written decision per criterion, with the example conversation attached |
| Rubric actions | 5 min | Assign the specific wording changes, named owner and date | Edited scorecard shipped within a week |
How to discuss variance without it becoming a debate
The discussion phase is where calibration sessions either produce a standard or turn into a status contest. A few facilitation rules keep it productive.
Anonymise the scores, name the reasoning
Show the spread on a criterion (for example, four reviewers scored between 2 and 5) without attributing scores to people. Then ask for the reasoning at each end. You want the arguments in the open and the egos out of it. In practice this makes junior reviewers far more willing to defend a position that differs from their manager’s.
Demand evidence, not impressions
Every position has to point at something in the transcript. If a reviewer says the tone was dismissive, the next sentence should be the line that was dismissive. Impressions that cannot be located in the conversation are exactly the impressions that agents cannot act on, and they are the reason QA feels unfair.
Separate the score from the rubric
The question is never who was right. It is what the rubric actually asks for, and whether it asks for it clearly enough that both readings were reasonable. If both readings were reasonable, the reviewers are not the problem.
Close every criterion before you move on
Do not let the session drift to the next conversation with the previous criterion unresolved. A calibration session with three open questions at the end has produced nothing. It is better to cover one criterion properly than five superficially.
Write the decision down while everyone is in the room
Record the agreed interpretation and link it to the example conversation. That pairing, a rule plus a real case, is what makes the standard usable for a reviewer who was not in the session and for onboarding the next hire.
Close the loop: change the rubric, not the reviewer
A calibration session that ends with everyone nodding and nothing changing has failed. The measurable output is an edited scorecard. When reviewers disagree, the cause is almost always one of four things, and all four are fixable in the rubric rather than in the person.
- Vague wording. Criteria like showed empathy or handled the issue well invite interpretation. Replace them with observable behavior: acknowledged the customer’s frustration before moving to the solution.
- Multiple things in one criterion. If a criterion bundles accuracy, tone and process, reviewers will weight the parts differently. Split it.
- No defined failure threshold. Reviewers need to know what a 3 looks like versus a 4. Anchor each point on the scale to a description, and attach a real conversation as the reference example.
- Missing edge case rules. What happens when the customer was abusive, when the system was down, when the policy did not cover the situation? Decide once, write it down, and stop deciding it again every two weeks.
Maintain a short living document of these decisions alongside the scorecard itself. Every entry should be one line: the criterion, the agreed interpretation, and a link to the conversation that settled it. New reviewers should read it before their first review, which shortens onboarding from months of osmosis to an afternoon.
One more loop to close: tell the agents. If a criterion changed because reviewers could not agree on it, that is worth saying out loud in the team channel. It shows that the standard is maintained rather than imposed, and it makes the feedback that follows much easier to accept.
How to measure whether calibration is working
Calibration is one of the few QA activities you can put a hard number on, and you should. Without a metric, sessions become a ritual that everyone attends and nobody can justify.
The two numbers to track
Agreement rate is the percentage of criteria on which all reviewers gave an identical score for the same conversation. It is blunt but honest, and it is the number to report upward. A team starting out is often under 50 percent. Above 80 percent on a mature scorecard is a reasonable target, and 100 percent should make you suspicious that your sample is too easy.
Average variance is how far apart reviewers are when they do disagree, measured per criterion. This is the more useful diagnostic, because it tells you where the ambiguity lives. A criterion with consistently high variance is a criterion to rewrite, regardless of how the overall agreement rate looks.
Track both per session and plot them over time. The trend matters more than any single reading: agreement should climb after a rubric change and hold, and a sudden drop usually means a policy shift or a new reviewer who has not yet been calibrated. It is also worth watching your internal quality score for sudden movement that has no operational cause, since that is often reviewer drift rather than a real change in quality.
What to do when reviewers keep disagreeing
If a criterion stays contentious after two or three sessions, stop trying to talk it into alignment. Take one of these routes instead:
- Split it into an objective part that can be checked against evidence and a subjective part that is coached on but not scored.
- Reduce the scale. A five point scale on a subjective criterion invites disagreement that a pass or fail does not. Precision you cannot achieve is not precision.
- Drop it from the score and keep it as a coaching note. Not everything worth discussing needs to affect an agent’s number.
- Check for a genuine business disagreement. Sometimes reviewers disagree because the organization itself has not decided whether speed or thoroughness wins. That is a leadership decision, not a QA one, and calibration is how you surface it.
Where automation changes the maths
Human reviewers drift: they get tired, standards move as the team turns over, and they only ever read a small sample. Automated scoring applies one identical rubric to 100 percent of conversations, so it cannot drift the way a group of raters does. That does not remove the need for calibration, it changes what calibration is for. Instead of policing each other, the session becomes the place where humans agree the standard, and then the place where you prove the automated scoring is applying that standard the way you meant it.
Run the same measurement you already run on your reviewers, with the automated score treated as one more rater. Take the reference set the group scored blind, run it through the system, and calculate agreement rate and average variance per criterion exactly as above. A criterion where the system disagrees with a calibrated group is either a criterion the system is reading wrongly or, more often, a criterion your rubric never defined clearly enough. Repeat the check on a schedule, because a rubric edit or a policy change can move agreement without anyone noticing. Nobody should be asked to accept a vendor’s accuracy claim in place of this.
That check is only possible when scores are explainable. Because Kaizo traces every score back to the exact evidence in the transcript, a calibration session can inspect why a score was given rather than argue about impressions, and a disagreement turns into a rubric edit instead of a standoff. It also matters who is doing the grading: most QA and CX platforms now sell their own AI agents, so when they score an AI-handled conversation they are grading their own product. Kaizo does not sell AI agents, so the standard your group calibrated is applied to your people and your automation on the same terms. At UiPath, that combination let the team automate 100% of QA with 200% ROI and an 8% lift in quality score.
Frequently asked questions
How often should you run QA calibration sessions?
Every two weeks suits most teams. Move to weekly for the first two months of a new scorecard or after a significant policy change, when ambiguity surfaces fastest. Monthly is the practical floor, because below that reviewers drift far enough apart between sessions that you spend the meeting relitigating basics rather than improving the rubric.
What is a call calibration session?
A call calibration session is the same exercise applied to voice: every reviewer scores the same two or three recorded calls independently, then the group compares scores criterion by criterion and agrees on one reading of the scorecard. Voice adds two wrinkles worth planning for. Listening happens in real time, so a 12 minute call costs every attendee 12 minutes, which is why you score before the meeting rather than during it. And tone carries meaning that a transcript loses, so agree explicitly whether reviewers are scoring what was said, how it was said, or both.
Who should attend a calibration session?
Everyone who scores conversations, plus the team leads who coach on those scores and the person who owns the scorecard. Keep the group small enough that everyone speaks. Rotating a senior agent in once a quarter is one of the fastest ways to build trust, because they see the standard being argued over seriously rather than handed down.
How many conversations should you calibrate on in one session?
Two or three in a 60 minute session. Teams that try to cover six end up skimming all of them and resolving none. Pick borderline, disputed or outlier conversations rather than clean examples, since clean conversations produce easy agreement that teaches you nothing about where your rubric is ambiguous.
How do you measure whether calibration is working?
Track two numbers. Agreement rate is the percentage of criteria where all reviewers gave an identical score on the same conversation. Average variance is how far apart they are when they disagree, measured per criterion. Agreement rate is the headline; variance per criterion is the diagnostic that tells you which rubric wording to fix next.
What do you do when reviewers keep disagreeing on the same criterion?
Stop trying to talk it into alignment and change the criterion. Split it into an objective part and a subjective part, reduce a five point scale to pass or fail, or move it out of the score and keep it as a coaching note. Persistent disagreement is a rubric problem, not a reviewer problem, unless the underlying cause is that the business has never decided whether speed or thoroughness wins.
Does automated QA scoring remove the need for calibration?
No, but it changes the purpose. Automated scoring applies one identical rubric to every conversation, so it does not drift the way a group of human raters does. Calibration stops being about policing each other and becomes the process of agreeing the standard the automation then applies everywhere. You still need humans to define what good looks like, and you should calibrate the system the way you calibrate people: run your reference set through it, measure agreement criterion by criterion, and fix the rubric wording wherever it disagrees.
Related terms
Calibrate the standard, then apply it to every conversation
Bring the reference set your team calibrated on and we will run it through Kaizo so you can compare criterion by criterion. Every score links to the evidence in the transcript, which is what lets calibration sessions inspect the reasoning instead of arguing about impressions.