Quality assurance on Salesforce Service Cloud means evaluating the support work recorded on a case: the replies an agent actually wrote, the channel the customer arrived on, and the state the case was left in. The difficult part is not the scorecard, it is deciding what counts as the conversation, because a case is a container for related records rather than a single thread. Settle that definition first and sampling, attribution and reporting all follow from it.
In short
- A Service Cloud case is a container, not a conversation. Decide which records inside it you are scoring before you write a single criterion.
- The case owner is often not the person who wrote the reply. On a case that changed hands, scoring the owner attributes someone else’s work to them.
- Every scorecard criterion has to map to something that exists on the case in your org, or reviewers end up scoring from memory.
- Sampling starts with a report that defines the population, not with whichever cases are easiest to open.
- One case can hold email, chat and voice. Decide whether that is one score or several before rollout, not during the first dispute.
- Almost nothing about a case model is standard between orgs, so a QA scorecard has to be mapped to your fields rather than copied from a template.
What quality assurance on Salesforce Service Cloud actually means
Most QA guidance assumes a ticket: one customer, one thread, one assignee, read top to bottom. Service Cloud does not work that way, and that single structural difference is why generic QA advice tends to fall apart in the first week.
A case is a record. It carries fields such as status, origin, reason, priority and owner, though the values in those picklists are configured by your own admin rather than fixed by the platform. Around that record sit the things a customer and an agent actually did: emails, chat or messaging transcripts, internal posts, tasks, and, where the org has voice provisioned, call records. The case feed shows a chronology of that activity, but a chronology is not the same thing as a conversation. It mixes agent replies with system entries, field changes and posts from people who never spoke to the customer.
So the first question in Service Cloud QA is not “what should we score”. It is what are we calling the conversation. Three defensible answers exist, and you need to pick one deliberately:
- The customer-facing exchange only. Every message the customer could see, in order, ignoring internal posts and field history. This is the closest analogue to a ticket thread and the easiest to explain to an agent.
- The exchange plus the record state. The same messages, judged alongside how the case was categorised and closed. Right for teams whose data quality feeds reporting.
- The whole case including internal work. Messages, internal notes, handoffs and tasks. The most complete view, and the hardest to score consistently, because reviewers disagree about how much internal work is enough.
Write the answer down before you write criteria. If you do not, each reviewer quietly picks their own, and you get the disagreement pattern every QA lead recognises: two people score the same case sixty and eighty-five, and neither can say why. The scorecard mechanics that sit on top of this are the same ones in any support QA programme, and if you have not built one yet, start with how to build a QA scorecard and come back.
What there is to evaluate on a case, and what there is not
Before writing criteria, take one recently closed case and inventory what is genuinely on it. Do this in your own org rather than from a template, because what is present varies enormously with how the org was configured.
Usually available
- Agent-authored messages. The email replies, chat or messaging turns, and any public posts. This is the core evidence for anything about communication.
- Timestamps. When the customer wrote, when the agent replied, when ownership changed, when the case closed. Enough to reason about responsiveness, with caveats below.
- Field values at close. Status, reason, product or category fields, and whatever custom fields your org requires. This is what makes a case reportable, so it is legitimately part of quality.
- Internal notes and posts. The handoff quality, the investigation, the escalation reasoning.
- Ownership history. Who held the case, and when it moved.
Frequently missing, and worth checking before you promise a criterion
- Voice. Whether a call is attached at all, and whether there is a transcript rather than only a recording or a note, depends on how voice is provisioned in your org. Confirm it before writing a criterion that depends on call content.
- What the agent could see. Field-level security and sharing rules mean the reviewer’s view of a case is not necessarily the agent’s view. If a criterion assumes the agent had information in hand, verify that they did.
- Knowledge state at the time. If an article changed after the case closed, the agent may have been right against the version they had.
- Anything that happened off the record. A phone callback logged as a one-line note, a conversation in a chat tool, a warm handoff to another team. These are real work and they are invisible to QA.
The rule that follows is short. If a criterion cannot be evidenced from something on the case, it is not a QA criterion, it is an opinion. Either find the field or the message that evidences it, or move it out of the scorecard and into coaching. A general-purpose customer service QA checklist is a useful starting inventory, but every line of it has to survive this test against your own case layout.
How case data maps to scorecard criteria
This is the mapping exercise that makes or breaks the programme. For each criterion, name the evidence, name the check, and name the failure mode you expect. The last column is the one teams skip, and it is the one that predicts every argument you are going to have.
| Criterion type | Evidence on the case | How a reviewer checks it | Where it goes wrong |
|---|---|---|---|
| Resolution accuracy | The final agent message, read against the fields the case was closed with | Does the answer given match the outcome recorded, and was it correct at the time | Status gets set by an automation or a macro, so the record says resolved without any evidence that it was |
| Process compliance | Required fields, internal notes, and any entitlement or service-level records the org has configured | Compare what the case records against your documented process for that case type | The requirement lives in a wiki rather than a field, so two reviewers apply two different versions of it |
| Communication quality | Agent-authored messages only, never the whole feed | Score the text the customer could actually read | Reviewers absorb tone from internal posts and system entries and attribute it to the agent |
| Responsiveness and effort | Timestamps on messages and on ownership changes | Measure gaps between a customer message and the next agent reply | Queue time, business hours and waiting on another team all get charged to whoever happens to own the case |
| Categorisation and data quality | Origin, reason, record type and your own reporting fields | Do the values describe what actually happened in the conversation | Agents pick whatever value clears the validation rule fastest, and nobody scored it until now |
| Routing | Ownership history, queue membership, record type | Did the case reach the right team first, and if not, why not | A misroute caused by a routing rule is scored as an agent failure |
How scoring a case differs from scoring a Zendesk ticket
Plenty of support ops leads have run QA on Zendesk and are now being asked to run it on Service Cloud, or to run both at once. The scorecard philosophy transfers. The mechanics do not, and these are the six differences that cost the most time. If you are running the Zendesk side as well, the platform-specific mechanics for that one are covered separately in quality assurance for Zendesk.
| Dimension | Zendesk ticket | Salesforce Service Cloud case |
|---|---|---|
| The unit of review | A ticket is a thread. Requester, assignee and comments read as one linear conversation | A case is a record. The conversation is assembled from related messages, transcripts and feed items |
| Who did the work | Assignee plus the visible comment authors | The owner field says who holds the case now. Authorship has to be read off each individual message |
| Channel | Usually a property of the ticket itself | Channels arrive as different related records, and one case can carry more than one of them |
| Defining a sample population | Views and search filters | A report or list view built on your org’s own case fields |
| How much is standard | The field set is broadly consistent between accounts | Record types, custom fields and picklist values are configured per org, so no scorecard transfers unmodified |
| Closing state | A short, familiar set of statuses | Status values, and what each one means operationally, are defined by your admin |
How to pull a sample you can defend
The most common accusation levelled at a QA programme is that the reviewer picked the worst cases. On Service Cloud that accusation is easy to make, because there is no natural review queue and it is genuinely tempting to work from whatever list view is already open.
Define the population first, in a report, and save the report. A workable filter is closed date inside the review period, record type or queue that identifies the work you are scoring, channel if you score channels separately, and an exclusion for cases with no agent-authored message at all. Duplicates, auto-closed cases and cases that never reached a human are noise, and leaving them in the population quietly inflates or deflates every average you publish afterwards.
Then sample from that population rather than from the list. Two rules matter more than the sample size:
- Randomise inside strata, not across the whole set. Stratify by agent and by case type, then draw randomly within each stratum. Pure random sampling across a whole month gives your busiest agents twenty cases and your newest agent one, which makes individual scores incomparable.
- Fix the sample before anyone reads a case. Draw the list, save it, and review what you drew. A sample chosen after the reviewer has skimmed the population is not a sample.
Be honest with yourself about what a manual sample buys you. Reviewing thirty cases per agent per month is a serious amount of work, and it still tells you very little about the specific case types that occur rarely and hurt most. That limitation is arithmetic, not effort, and it is the same constraint NIST’s engineering handbook sets out when it asks whether a sampled proportion of defectives meets a requirement: the sample has to be large enough for the test to be valid before the answer means anything. Which is why a sampled score and an internal quality score reported to leadership need to carry their confidence interval alongside the number, and why it is worth knowing where QA sampling stops being able to answer the question.
Who wrote the reply: attribution on a case that changed hands
Here is the problem no template prepares you for, and the one that does the most damage to agent trust in a Service Cloud QA programme.
Cases move. They get routed, escalated, reassigned, picked out of a queue, handed to a specialist team and handed back. The owner field records who holds the case at the moment you look at it. It does not record who wrote the reply that annoyed the customer, or who did the investigation that solved it. If your QA process scores a case and posts the result against the owner, then on every transferred case you have attributed one person’s work to another.
The practical consequence is predictable and specific. The agents who handle escalations inherit cases that are already going badly, and they collect the resulting scores. Those are usually your most capable people. Within a quarter they learn that taking a hard case costs them, and the behaviour you most wanted to encourage becomes the behaviour your scoreboard punishes.
Four rules that fix it
- Score the message, not the record. Every deduction should attach to a specific agent-authored message with a timestamp and an author, not to the case as a whole.
- Attribute by author. Read authorship off each message rather than off the owner field. If your reporting cannot do that, this is the first thing to fix, before you tune a single weight.
- Handle multi-agent cases explicitly. Either score each agent’s own contribution separately, or exclude the case from individual scoring and use it for process review. Both are defensible. Silently scoring the last owner is not.
- Keep the ownership history next to the score. When a score is disputed, the first question is always who did what. If the answer takes twenty minutes to reconstruct, the dispute is already lost.
Transferred cases are also the richest source you have for process work rather than individual coaching. A case that moved four times before it was solved is telling you something about routing, entitlements or team boundaries, and that belongs in root cause analysis rather than on somebody’s scorecard.
Scoring every case instead of a sample
Once the manual method is settled, the obvious question is whether you can stop sampling. You can, and the argument for it on Service Cloud is stronger than on a simpler ticket model, because the case types that matter most are frequently the rare ones that a sample never reaches. There is a real difference between 100% coverage revealing trends that 3% sampling never could, and a monthly review of thirty cases per agent.
Coverage on its own is not the interesting part, though, and it is no longer a differentiator either. Every vendor in this category now scores everything. The question that decides whether an automated score survives contact with your team is whether you can show the grader was right. On Service Cloud specifically, that means four things:
- Traceability into the case. A deduction has to name the criterion and point at the exact message inside the case that triggered it. “Communication: minus ten” on a case with forty feed items is unusable.
- The same evidence the grader used. A reviewer checking a score should see the message set the score was computed from, not a summary of it.
- A dispute route. An agent has to be able to challenge a score and have somebody inspect the reasoning. If they cannot, the score is a verdict.
- Measured agreement before it counts. Run the automated score alongside your own reviewers on the same cases for a few weeks and validate the scoring before the number is attached to anything with consequences.
Kaizo’s Auto QA applies your own scorecard criteria to cases and keeps the reasoning attached to the conversation, so a disputed deduction can be traced back to the exchange that caused it rather than argued from memory. It runs on Service Cloud through the Salesforce integration, described in more detail in Kaizo on Salesforce, alongside the same integration for Zendesk, which matters if you are running QA across both while a migration is in progress.
A rollout that survives the first dispute
Service Cloud QA programmes fail at the same point, which is the first time an agent seriously challenges a score and the programme cannot answer. Sequence the rollout so that the answer exists before the challenge does.
| Stage | What you do | Done when |
|---|---|---|
| 1. Define the unit | Pick one of the three definitions of the conversation and write it down. Inventory what is actually present on a closed case in your org | A reviewer can state, without hedging, which records are in scope |
| 2. Map criteria to fields | For each criterion, name the evidence, the check and the expected failure mode. Drop anything you cannot evidence | Every criterion points at a field or a message type that exists |
| 3. Build the population report | One saved report that defines the reviewable set, with exclusions for auto-closed and no-agent cases | Two people running the report get the same case count |
| 4. Shadow score and calibrate | Three reviewers score the same fifteen cases independently, including at least three that changed owner. Compare and argue | You know your reviewers’ disagreement rate, and it is falling |
| 5. Publish with an appeals route | Announce the scorecard, the sample method and the appeal process together. An appeal with no route is a grievance | Agents can name the process without looking it up |
| 6. Report, then re-check the mapping | Report scores next to operational outcomes such as reopens and escalations. Revisit the mapping each quarter | Score movement explains something a manager already suspected |
Reporting a case score so somebody acts on it
A QA score that lives only in a QA tool gets read by the QA team. To be worth the effort, the score has to sit next to the operational numbers the support leadership already looks at, and it has to be broken down in a way that names an action.
Three cuts do most of the work on Service Cloud data:
- By case type or record type, not just by agent. Quality varies far more by the kind of work than by the person doing it. A low average on one record type is a process, knowledge or routing problem, and it is fixable centrally.
- By criterion, across the whole team. One criterion failing broadly is almost never a coaching problem. It is an unclear policy, a missing knowledge article, or a criterion that is badly written.
- Score against reopen and escalation rate. This is the check on the scorecard itself. If high-scoring cases reopen at the same rate as low-scoring ones, the scorecard is measuring something other than quality and needs reweighting.
That third cut is the one worth protecting. It is what stops a QA programme drifting into a compliance ritual that everybody performs and nobody uses. Pairing quality scores with the operational picture is exactly what a support analytics view is for, and it is also the fastest way to demonstrate to a sceptical leadership team that the QA programme is measuring something real.
Frequently asked questions
Does Salesforce Service Cloud include QA scoring out of the box?
Service Cloud gives you the raw material: the case record, the messages and transcripts related to it, the field history and a reporting layer over all of it. A weighted scorecard, a reviewer queue, calibration between reviewers and an auditable appeals trail are not part of the standard case object. Teams either build a lightweight version using custom fields and reports, which works until reviewer disagreement becomes the bottleneck, or they add a QA tool that reads cases. Check what your own org already has configured before assuming either way, because customisation varies enormously.
What should you evaluate on a Salesforce case?
Only what the case evidences. In practice that is the agent-authored messages, the timestamps on those messages and on ownership changes, the field values the case was closed with, and the internal notes if your definition of the conversation includes them. Anything you cannot point at on the record, such as what the agent knew at the time or a callback logged as a one-line note, belongs in coaching rather than in a score.
How is QA on Service Cloud different from QA on Zendesk?
The scorecard philosophy is the same, the mechanics are not. A Zendesk ticket is a linear thread with a clear assignee, so the unit of review is obvious. A Salesforce case is a record with related messages, transcripts and feed items, so you have to define the unit yourself. Authorship also has to be read off individual messages rather than taken from the owner field, and because record types and custom fields are configured per org, no scorecard transfers between two Salesforce orgs unmodified.
How do you sample Salesforce cases for QA review?
Build a saved report that defines the population, filtering on closed date, record type or queue, channel where relevant, and excluding cases with no agent-authored message. Then draw randomly within strata, by agent and by case type, rather than randomly across the whole month, so every agent contributes a comparable number of cases. Fix the sample before anyone opens a case. Reviewing a list you have already skimmed is not sampling.
Who gets the QA score when a case changes owner?
Not automatically the current owner, which is the default most programmes fall into and the one that does the most damage. Attribute each scored message to whoever wrote it. On a case handled by several people, either score each agent’s own contribution separately or take the case out of individual scoring and use it for process review. Scoring the last owner for work they inherited teaches your strongest agents to avoid escalations.
Can you score email, chat and voice on the same case?
You can, but decide the approach before rollout rather than during the first dispute. A single blended score across a case containing an email thread and a chat transcript is hard to coach on, because the standards for the two are genuinely different. Scoring each channel segment separately and reporting them separately is usually clearer. If voice is involved, confirm first whether your org actually has transcripts attached rather than only recordings, since a criterion that depends on call content is unenforceable without them.
Related terms
See your own Service Cloud cases scored end to end
Bring one month of closed cases and the scorecard you use today. We will show you what your criteria actually pick up on real case data, and where a score cannot yet be traced back to the message that caused it.