A QA scorecard template is a fixed list of criteria you score every conversation against, each with a plain-English definition and a weight that reflects how much it matters. A good one covers five things: whether the issue was resolved, whether the information was accurate, whether policy and process were followed, how the customer was treated, and how efficiently it was handled. Which template you should use depends on the channel, because what counts as good in a two-minute live chat is not what counts as good in a formal complaint response. The five templates below are complete and ready to copy: pick the one that matches your channel, delete the criteria that do not apply to your business, and start scoring.
In short
- Five complete templates below, for general support, live chat, email, voice and complaints, with criteria, definitions and suggested weightings you can copy today.
- Six to eight criteria is the sweet spot. Fewer than five and the score says nothing useful, more than ten and reviewers start rushing.
- Weight by business impact, not by how easy something is to score. Resolution and accuracy should always carry the most weight.
- Keep auto-fail criteria separate from the weighted score, and keep the list short: compliance breaches, data handling and serious customer harm only.
- Every criterion needs a one-line definition, or two reviewers will read it differently and your scores will not be comparable.
- Channel matters. Chat rewards pace and brevity, email rewards completeness, voice rewards call control, complaints reward ownership.
- A template only creates value if every score can be traced back to the line in the conversation that produced it, and if it is applied consistently, which is the part manual sampling almost always fails at.
Before you copy: how to use these templates
A QA scorecard is the fixed set of criteria your team scores conversations against, so that quality is measured the same way by everyone. If you want the full framework behind it, the QA scorecard guide covers how to design one from scratch, and the QA rubric explainer covers how to write scoring levels that reviewers actually agree on. This page is the shortcut: five finished scorecards you can lift straight into your QA tool, your quality monitoring form or a spreadsheet.
Three rules before you start
- Delete before you add: every template below is deliberately slightly over-specified. Cut anything your business does not genuinely care about rather than bolting on more criteria.
- Keep the definitions: the one-line definition next to each criterion is the part that makes scores comparable between reviewers. Copy those across too.
- Score on a 1 to 5 scale, or pass and fail: pick one and use it everywhere. Mixing scales across criteria makes the weighted total meaningless.
Each template uses percentage weights that add up to 100, so your final score lands as a clean percentage. Adjust the weights to your priorities, but keep the total at 100.
Template 1: General customer support scorecard
The default, channel-agnostic scorecard. Use this if you run a mixed inbox and want one standard across everything, or as the base you fork the other templates from.
- Issue resolved (25%): the customer’s actual problem was solved, not just answered. If a follow-up contact on the same issue was inevitable, this is not a full score.
- Accuracy of information (20%): everything the agent stated was factually correct against the knowledge base, product behavior and current pricing or policy.
- Process and policy adherence (15%): required steps were completed in order, including verification, logging, tagging and any mandatory disclosures.
- Tone and empathy (15%): the agent matched the customer’s situation, acknowledged frustration where it existed, and stayed warm without being robotic or over-familiar.
- Clarity (10%): the response was easy to understand first time, free of internal jargon, and structured so the customer knew what to do next.
- Ownership and proactivity (10%): the agent took the issue on, anticipated the obvious next question, and did not push the customer to another team without cause.
- Efficiency (5%): the issue was handled without unnecessary back-and-forth, holds or repeated information requests.
Suggested pass mark: 85%. Anything below 70% should trigger a coaching conversation rather than a note in a report.
Template 2: Live chat scorecard
Chat is short, fast and concurrent, so completeness matters less than pace, brevity and not leaving the customer staring at an empty window. Pair this with your live chat metrics so the quality score and the speed numbers tell one story.
- Opening and acknowledgement (10%): the agent greeted the customer, confirmed they had read the issue, and did not open with a generic script that ignored what was written.
- Response pace and silence management (15%): replies came within your target gap, and any longer pause was flagged with a short holding message rather than silence.
- Issue resolved (25%): the chat ended with the problem solved or a clearly owned next step, not with the customer giving up or being told to email.
- Accuracy of information (15%): facts, links and instructions given in the chat were correct and current.
- Tone in short form (15%): the agent stayed human in a compressed format. Brevity did not become bluntness, and there was no unnecessary formality.
- Quality of canned content (10%): macros and saved replies were edited to fit the customer’s situation rather than pasted in whole, and no irrelevant boilerplate was dumped into the chat.
- Closing and confirmation (10%): the agent confirmed the issue was resolved, set out any follow-up, and closed properly instead of letting the chat time out.
Suggested pass mark: 85%. Score whole chats, not individual messages, or the pace criteria cannot be judged fairly.
Template 3: Email and ticket scorecard
Email is asynchronous, so every avoidable round trip costs the customer hours or days. This template weights completeness and structure far more heavily than speed.
- Every question answered (25%): the reply addressed all of the questions the customer asked, including the ones buried at the bottom of a long message.
- Accuracy of information (20%): facts, steps, links and figures were correct, and any account-specific detail was checked rather than assumed.
- Structure and scannability (15%): the answer was organised so the customer could find it, using short paragraphs, numbered steps or bold labels where the content warranted it.
- Tone and personalization (15%): the email read as though a person had actually read the customer’s message, referenced their specific situation, and matched the seriousness of the issue.
- Grammar and formatting (10%): no spelling or grammar errors, correct name and details, no broken merge fields, no leftover template placeholders.
- Next steps and expectations (10%): the customer was told what happens next, who does it and by when, so they do not have to chase.
- Round trips avoided (5%): the agent gathered or looked up what they needed instead of sending a one-line reply asking for information already available.
Suggested pass mark: 88%. Email is the channel where a written record exists, so hold it to a slightly higher bar.
Template 4: Voice and phone scorecard
Voice adds two things no other channel has: real-time call control and a much higher compliance risk. This template covers both. If you are building out call center QA more broadly, the call center QA framework guide sits above this template.
- Verification and compliance (15%): identity was verified to the required standard and all mandatory statements, disclosures or consent wording were given before the relevant part of the call.
- Issue resolved (20%): the caller’s problem was solved on the call, or a specific, owned commitment was made with a date attached.
- Accuracy of information (15%): everything said on the call was correct, including anything quoted from memory rather than looked up.
- Active listening (15%): the agent let the caller finish, did not make them repeat information already given, and reflected the issue back accurately before solving it.
- Voice quality and pace (10%): clear speech, appropriate pace, no talking over the customer, and language pitched to the caller rather than to a colleague.
- Hold and transfer handling (10%): holds were requested, explained, kept short and thanked for. Transfers were warm, with context passed to the receiving agent.
- Call control and structure (10%): the agent guided the call, kept it on track without rushing the customer, and avoided dead air while working.
- Close and wrap-up (5%): the outcome was summarised back to the caller, further questions were invited, and the notes left behind matched what actually happened.
Suggested pass mark: 85%, with verification and compliance also carried as an auto-fail (see below).
Template 5: Complaints and escalations scorecard
Complaints are scored differently because the outcome is only half the job. How the customer was handled on the way to the outcome determines whether they stay. Ownership and de-escalation carry real weight here.
- Acknowledgement and de-escalation (20%): the complaint was acknowledged in the customer’s own terms early in the conversation, with a genuine apology where one was warranted and no defensiveness.
- Remedy or resolution offered (20%): a concrete outcome was put forward, within policy, proportionate to the harm caused, and explained clearly rather than hinted at.
- Root cause established (15%): the agent worked out what actually went wrong rather than treating the symptom, and captured it well enough for the business to act on.
- Ownership, no bouncing (15%): one person owned the case end to end. The customer was not asked to re-explain the complaint to a second or third person.
- Policy and regulatory handling (15%): required complaint wording, timelines, escalation rights and record-keeping obligations were met exactly.
- Expectations and follow-through (10%): the customer was told what happens next and by when, and any commitment made was specific enough to be checked later.
- Handover documentation (5%): the case notes are good enough that the next person picking it up needs nothing else to continue.
Suggested pass mark: 90%. Complaints are your highest-risk conversations, so the bar should be higher than your general scorecard, not lower.
How to weight criteria and set a pass mark
Weighting is where most scorecards go wrong. The temptation is to weight what is easy to score, which is usually formatting and greetings, because those are quick and objective. The result is a scorecard that rewards tidy conversations that did not solve anything. Weight by business impact instead, and use the grouping below as a sanity check on whatever you build.
| Criterion group | What it measures | Suggested weight | Notes |
|---|---|---|---|
| Outcome | Was the customer’s issue actually resolved | 20-30% | Should always be the single heaviest criterion |
| Accuracy | Was the information given correct and current | 15-25% | Objective and evidence-checkable, so easy to score consistently |
| Compliance and process | Were required steps, disclosures and policy followed | 10-20% | Higher in regulated industries, and often an auto-fail as well |
| Communication and tone | How the customer was treated and understood | 15-25% | The most subjective group, so keep the definitions tight |
| Structure and clarity | Was the response organised and easy to act on | 10-15% | Weight higher on email, lower on voice |
| Efficiency | Was it handled without avoidable friction | 5-10% | Keep it small. Over-weighting this drives agents to rush |
Auto-fails, and how to avoid the tick-box trap
Two things separate a scorecard people trust from one they quietly game. The first is a short, well-drawn auto-fail list. The second is a habit of checking that the score still means something.
When to use an auto-fail
An auto-fail sets the whole conversation to zero regardless of the weighted score. It exists for the failures that cannot be averaged away by good behavior elsewhere. Keep the list to four or five items, for example:
- Identity or verification not completed where your policy required it.
- Customer data mishandled, shared with the wrong party or logged where it should not be.
- Deliberately incorrect or misleading information given to close the contact.
- Abusive, discriminatory or dismissive conduct toward the customer.
- A regulatory obligation missed, such as a required disclosure, complaint right or consent.
Everything else belongs in the weighted score. If your auto-fail list is longer than five items, some of them are really just heavily weighted criteria, and treating them as auto-fails will make your quality trend unreadable.
How to avoid the tick-box trap
A scorecard becomes a tick-box exercise when reviewers can score it without thinking and agents can pass it without helping anyone. Three habits prevent that:
- Require evidence for every score: a score with no quoted line from the conversation cannot be verified, coached on or challenged. This applies to automated scoring too: if a tool cannot show you the line it scored against, you cannot check whether it read the conversation correctly.
- Calibrate regularly: have reviewers score the same conversations and compare. QA calibration is what stops the same rubric drifting into two different standards.
- Review the scorecard itself twice a year: if a criterion has scored above 95% for two quarters, it is no longer telling you anything and the weight is better spent elsewhere. If you want a list of the failure patterns to design out, the DO NOT checklist for QA scorecards is a useful companion, and the 10 unique QA templates e-book covers a wider set of niche scenarios than the five above.
Making the template actually stick
The hard part is not choosing criteria. It is applying them consistently once the novelty wears off. Most teams start strong, then quietly slip to a handful of reviews per agent per month, at which point the scorecard describes a sample so small that a single bad week can swing an agent’s average. The scorecard is fine. The coverage is the problem, and no amount of redesigning the criteria will fix it.
Kaizo applies your scorecard, including these templates, with every score linked back to the exact line in the conversation that produced it, so any score can be checked, coached on or overturned rather than taken on faith. Full coverage is what makes that possible: when every conversation is scored rather than a 2% sample, the evidence exists for all of them instead of only the ones someone picked. And because Kaizo does not sell the AI agents that increasingly handle a share of your conversations, the same scorecard can be applied to human-handled and AI-handled work by a grader with nothing of its own to protect. At UiPath, Kaizo automated 100% of QA, delivering 200% ROI and an 8% lift in quality score. Kaizo reads conversations natively from Zendesk and Salesforce.
Whether you automate it or not, the sequence is the same: copy a template, define each criterion in one line, agree the weights with the people who will be scored on them, calibrate on a handful of real conversations, then use the results to drive coaching rather than to file a report. Publish the scorecard to the team before you score anyone against it, so agents know the standard they are being held to. A scorecard that changes what an agent does next week is worth ten that only produce a number.
Frequently asked questions
How many criteria should a QA scorecard have?
Six to eight is the practical sweet spot, and ten should be your hard ceiling. Below five criteria the score is too blunt to coach from, and above ten reviewers start pattern-matching instead of reading. If you need more detail, add scoring levels within a criterion rather than adding another criterion.
How should I weight QA scorecard criteria?
Weight by business impact, not by how easy something is to score. Resolution should be the heaviest single criterion at 20 to 30%, accuracy and compliance next, communication and tone in the middle, and efficiency smallest at 5 to 10%. Keep the weights adding up to 100 so your final score reads as a clean percentage.
What is a good pass mark for a QA scorecard?
85% works for most general support and chat scorecards, 88% for email where there is a written record, and 90% for complaints and other high-risk conversations. Set the threshold where roughly the top half of your team clears it comfortably, then hold it steady. A pass mark that moves every quarter cannot show a trend.
When should a QA criterion be an auto-fail instead of a weighted score?
Make it an auto-fail only when the failure cannot be offset by doing everything else well: a missed verification, mishandled customer data, deliberately misleading information, abusive conduct, or a missed regulatory obligation. Keep the list to four or five items. Longer lists turn ordinary mistakes into zeroes and make your quality trend unreadable.
Do I need a different scorecard for each channel?
You need a shared core and channel-specific extras. Resolution, accuracy and compliance should be scored the same way everywhere so results stay comparable. Around that core, chat needs pace and silence management, email needs completeness and structure, and voice needs hold, transfer and call control criteria.
How often should I change my QA scorecard?
Review it every six months, and change it only when the data or the business demands it. If a criterion scores above 95% for two consecutive quarters it has stopped being informative and the weight is better spent elsewhere. Avoid changing weights mid-quarter, because it breaks the comparison with everything scored before the change.
Related terms
See your scorecard applied, and check the grader on your own conversations
Bring the template you just copied and a set of conversations your reviewers have already graded. We will score them against your rubric, show you the exact evidence behind every score, and let you check the grader before you trust it.