Scoring soft skills in QA means evaluating things like empathy, tone, and de-escalation without reducing them to a reviewer’s gut feeling. The usual failure is a criterion like was the agent empathetic, which two reviewers answer differently because they are scoring their own impression rather than anything in the conversation. The fix is to decompose each soft skill into the observable behaviors that produce it, acknowledging the customer’s situation, matching the response to their emotional state, avoiding blame, and score those behaviors against the evidence in the transcript. Done that way, empathy stops being an opinion and becomes something you can measure, coach, and even have an AI score consistently.
In short
- Soft skills fail to score not because they do not matter, but because they are usually written as impressions rather than behaviors.
- The fix is to decompose each skill into observable behaviors: what did the agent actually do that demonstrates empathy or de-escalation.
- Behavior-based criteria are answerable from the transcript, so two reviewers, and an AI grader, can reach the same score.
- Empathy is a Kaizo-distinctive strength: scoring the empathy axis well is where many QA programs are weakest.
- Calibration matters most on soft-skill criteria, because that is where reviewer drift is largest.
- Scored well and at full coverage, soft skills become coachable with the exact moment in the conversation attached.
Why soft skills usually get scored badly
Most scorecards handle soft skills in one of two bad ways. They skip them, scoring only the objective mechanics because those are easy, which quietly tells agents that how they made the customer feel does not count. Or they include them as impressions, with criteria like was the agent empathetic or was the tone appropriate, which sound reasonable and are nearly unscorable, because each reviewer answers from their own sense of what empathy looks like.
The result is the worst of both. The criteria that arguably matter most to the customer experience are the ones scored least consistently, and agents learn that their soft-skill score is really a measure of who reviewed them. This is the same consistency problem that calibration exists to solve, concentrated on the criteria where it bites hardest.
Decompose the skill into observable behaviors
The move that makes a soft skill scorable is to stop scoring the abstraction and start scoring the behaviors that produce it. Empathy is not a feeling you detect; it is a set of things an agent did or did not do, each visible in the transcript.
| Soft skill | Unscorable version | Observable behaviors to score instead |
|---|---|---|
| Empathy | Was the agent empathetic? | Acknowledged the specific problem before solving; named the impact on the customer; avoided minimizing language |
| Tone | Was the tone appropriate? | Matched formality to the customer; avoided dismissive or robotic phrasing; adjusted when the customer was upset |
| De-escalation | Did the agent de-escalate? | Acknowledged the frustration explicitly; took ownership rather than deflecting; offered a concrete next step early |
| Active listening | Did the agent listen? | Referenced details the customer gave; did not ask for information already provided; confirmed understanding before acting |
| Reassurance | Was the customer reassured? | Set a concrete expectation; explained what happens next; avoided vague promises |
Write the behaviors so the evidence decides
Once you have the behaviors, phrase each as something the transcript answers rather than something the reviewer judges. Acknowledged the customer’s specific problem before offering a solution is answerable: you can point to the line, or its absence. Was warm is not. This is the same discipline as writing a rubric an AI can score, and it has a useful side effect: criteria clear enough for an AI to score are criteria clear enough that your human reviewers finally agree too.
Keep two cautions in mind. First, do not over-specify into a script. The behavior is acknowledged the customer’s situation, not said the exact words I understand this is frustrating, because scoring the script produces the robotic tone covered in positive scripting. Score whether the agent did the thing, in their own words. Second, allow for context: the empathy a routine password reset needs is different from a billing error that cost the customer money, so the strongest scorecards judge whether the response was proportionate to the situation.
Calibrate hardest here, then score everything
Soft-skill criteria are where reviewers drift most, so they are where calibration pays off most. Run your empathy and tone criteria through calibration specifically: have reviewers score the same emotionally charged conversations, surface where they diverge, and sharpen the behavior definitions until they converge. If reviewers cannot agree, the criterion is still an impression, not a behavior, and needs rewriting.
Two things then make soft-skill scoring genuinely useful. Coverage: scoring every conversation rather than a sample is what lets you see whether an agent’s empathy holds up across their full range, not just in the few calls that got pulled. And traceability: because each behavior score points at the exact line, coaching on something as delicate as empathy can start from a real moment in a real conversation rather than from a manager’s general impression, which is the difference between feedback an agent accepts and feedback they resent. Scoring the empathy axis well is, in practice, where many programs are weakest and where a QA approach built for it stands out.
Frequently asked questions
How do you score soft skills like empathy in QA?
By decomposing each soft skill into the observable behaviors that produce it, then scoring those against the transcript. Instead of asking was the agent empathetic, you score whether they acknowledged the customer’s specific problem, named its impact, and avoided minimizing language. Behaviors are answerable from the evidence, so reviewers, and an AI grader, can score them consistently.
Why are soft skills scored so inconsistently?
Because they are usually written as impressions rather than behaviors. A criterion like was the tone appropriate asks each reviewer to answer from their own sense of appropriate, so scores drift with who reviewed the conversation. Rewriting the criterion as specific observable behaviors removes most of that drift.
Can an AI score empathy and tone?
Yes, once the criteria are behavior-based rather than impressionistic. An AI can check whether an agent acknowledged the customer’s situation, referenced details they gave, and avoided dismissive phrasing, because those are observable in the transcript. It cannot reliably score was the agent empathetic phrased as an abstraction, which is exactly the version human reviewers score inconsistently too.
How do you score empathy without making agents sound scripted?
Score the behavior, not the exact words. The criterion is that the agent acknowledged the customer’s situation, however they phrased it, not that they said a specific sentence. Scoring a required script produces the robotic tone that undermines the point. Judge whether the agent did the thing, and allow the response to be proportionate to the situation.
Related terms
Make empathy and tone something you can actually score
Bring your current soft-skill criteria and a set of emotionally charged conversations. We will show you how to turn empathy and tone into behaviors your reviewers and an AI can score the same way, across every conversation, with each score traced back to the exact line so coaching starts from a real moment.