QA call selection is how you decide which conversations get reviewed. It sounds like an administrative detail, but it is one of the most consequential choices in a QA program, because the conversations you pick determine both the scores you produce and whether agents believe those scores are fair. The most common grievance on the support floor, that QA only ever looks at an agent’s worst calls, is usually not paranoia. It is an accurate observation about selection bias, and it is the fastest way to lose the trust a QA program depends on. Fixing it means either making selection genuinely random and representative, or removing selection from the equation by reviewing everything.
In short
- Call selection decides which conversations get scored, which shapes both the results and whether agents trust them.
- Manual selection is almost never neutral: reviewers gravitate to short calls, flagged calls, or calls near a complaint.
- The complaint that QA only picks an agent’s worst calls is usually a correct read on that bias, not a defensive excuse.
- Biased selection produces scores that do not represent an agent’s real work, so coaching built on them lands as unfair.
- The two honest fixes are genuinely random, representative sampling, or full coverage that removes selection entirely.
Why selection is where trust is won or lost
Every QA program that reviews a subset of conversations has to answer one question first: which ones? It is easy to treat that as logistics, but the answer quietly determines everything downstream. The conversations you select become the entire evidence base for an agent’s score, so if selection is skewed, the score is skewed, no matter how fair the scoring itself is.
This is also where agents form their opinion of the whole program. An agent does not see your scoring rubric or your calibration process. They see which of their calls got reviewed, and they notice the pattern. If the reviewed calls are consistently the hard ones, the program reads as a search for mistakes, and that perception, once formed, is very hard to reverse. A QA program’s credibility is decided at selection, before a single score is given. The wider trust question is covered in how to run a QA program agents trust.
Why manual selection is quietly biased
Reviewers rarely set out to pick unfairly. The bias creeps in through entirely reasonable habits, which is what makes it so persistent.
| Selection habit | Why it seems reasonable | The bias it creates |
|---|---|---|
| Reviewing flagged or escalated calls | Those calls seem most worth attention | Over-samples problems, so scores skew negative |
| Picking calls near a complaint or low CSAT | Wanting to understand what went wrong | Reviews an agent at their worst moments |
| Choosing shorter calls | They are faster to get through | Systematically excludes complex work |
| Reviewing recent calls only | They are top of mind | Misses patterns and rewards recency |
| Letting managers pick | They know their team | Selection reflects existing opinions of each agent |
The agent complaint is usually correct
When an agent says QA only picks their worst calls, the instinct is to treat it as deflection. It is worth taking literally instead, because more often than not it is an accurate description of the selection habits above. If reviewers gravitate to flagged calls, complaint-adjacent calls, and escalations, then an agent’s reviewed sample really is weighted toward their difficult moments, and their score really does understate their typical work.
That has two costs. The obvious one is morale: coaching built on an unrepresentative sample feels like an ambush, and agents disengage from a process they experience as unfair. The less obvious one is that the data is genuinely wrong. You are making decisions about people based on a slice of their work that was selected precisely because it was atypical. The complaint is not just a feelings problem to manage. It is a measurement problem to fix.
The two honest fixes
There are only two ways to remove selection bias, and they sit at different points on the same line.
Make selection genuinely random and representative
If you must sample, sample properly: pull conversations at random across each agent’s full range of call types, lengths, and outcomes, rather than letting reviewers or managers choose. This is harder than it sounds, because true randomness has to be built into the process rather than left to good intentions, and even done well a small random sample carries the statistical limits covered in why a small sample cannot support agent-level decisions.
Remove selection entirely with full coverage
The cleaner fix is to stop selecting. If every conversation is scored, there is no sample to bias, and the argument about which calls got picked simply disappears. An agent’s score reflects all of their work, so the objection that QA cherry-picks the bad ones has no purchase. Scoring 100% of conversations is what makes that possible, and it changes the conversation on the floor from was this fair to what does the work actually show. At UiPath, Kaizo automated 100% of QA with 200% ROI and an 8% lift in quality score. Because every score traces back to the specific evidence, a coaching conversation starts from what happened rather than from a dispute about the sample.
Frequently asked questions
What is QA call selection?
It is how a QA program decides which conversations get reviewed. Because the selected conversations become the entire evidence base for an agent’s score, selection shapes both the results and whether agents believe those results are fair. It is one of the most consequential and most overlooked choices in a QA program.
Why do agents say QA only picks their worst calls?
Usually because it is true. Reviewers naturally gravitate to flagged calls, escalations, and calls near a complaint, which weights an agent’s reviewed sample toward their hardest moments. The result is a score that understates their typical work, so the complaint is an accurate read on selection bias rather than a defensive excuse.
How do you make QA call selection fair?
Two ways. Either make selection genuinely random and representative, pulling conversations across each agent’s full range of call types and outcomes rather than letting people choose, or remove selection entirely by scoring every conversation. Full coverage is the cleaner fix, because with no sample there is no selection to bias.
Does full coverage remove selection bias?
Yes, by removing selection. If every conversation is scored, there is no subset to skew, so an agent’s score reflects all of their work and the objection that QA cherry-picks bad calls no longer applies. It also shifts coaching from a dispute about which calls were chosen to a discussion of what the work actually shows.
Related terms
Take the argument about call selection off the table
Tell us how you choose which conversations to review today. We will show you what scoring every conversation changes, how it removes the selection bias agents complain about, and how each score traces back to the evidence so coaching starts from what happened rather than a dispute about the sample.