Call monitoring is the practice of reviewing customer calls against a defined quality standard so you can measure how service is actually delivered and coach agents on it. It is done in three ways: listening live while a call is in progress, reviewing recordings after the fact, and scoring calls automatically against a scorecard. Manual monitoring typically covers a small sample of calls, often two to five per agent per month, while automated scoring can cover every call. The output is a score, the evidence behind it, and a coaching conversation. Done well, call monitoring is a feedback system, not a paperwork exercise.
In short
- Call monitoring has three modes: live monitoring, recorded review, and automated scoring of every call. Most teams need a mix.
- Before you monitor anything, confirm your recording, consent, retention, and payment-data obligations with your own legal and compliance advisors.
- A good call QA form has 8 to 12 weighted criteria, each one written so two reviewers would score it the same way.
- Sampling a handful of calls per agent tells you about those calls, not about your operation. The failures usually sit in the calls nobody reviewed.
- Feedback is where monitoring pays off: one strength, one fix, the exact moment in the call, and an agreed next step.
- Automated scoring changes the economics because coverage stops being a budget decision and reviewer time moves from grading to coaching.
- When evaluating call monitoring software, judge it on evidence-linked scoring you can verify on your own calls, scorecard flexibility, coaching workflow, and whether the grader is independent of any AI agents you deploy.
What call monitoring is, and the three ways teams do it
Call monitoring means reviewing customer calls against a defined standard so you can see how service is really delivered, then acting on what you find. It sits inside the broader discipline of quality monitoring, which covers every channel. This guide is about the calls specifically: how to review them, what to score, how often, and what to do with the result.
There are three practical modes, and they answer different questions. Teams that treat them as competing options usually pick the wrong one. Teams that stack them get the coverage of automation, the depth of recorded review, and the immediacy of live monitoring where it genuinely helps.
| Mode | How it works | Best for | Main limitation |
|---|---|---|---|
| Live monitoring | A supervisor listens while the call is in progress, sometimes with whisper or barge-in | New hires, escalations, live-floor coaching | Consumes a supervisor in real time, so it never scales |
| Recorded review | A reviewer scores a recording against a form after the call ends | Depth, calibration, disputes, sensitive cases | Slow and expensive, so coverage stays tiny |
| Automated scoring | Every call is transcribed and scored against your scorecard automatically | Coverage, trends, finding the calls worth a human look | Objective criteria score cleanly, subjective nuance still needs human calibration |
Get the consent and compliance basics settled first
Recording and monitoring calls involves personal data and, in many places, specific rules about notifying the people on the call. The rules vary widely by country, by state or region, and by industry, and they change. Nothing here is legal advice. Treat the points below as the questions to take to your own legal and compliance advisors, and confirm what applies in every jurisdiction you take calls from before you record or monitor anything.
The questions worth answering before you start
- Notification and consent: some jurisdictions require only that one party to the call knows it is being recorded, and others require that everyone on the call is informed or agrees. Because you often cannot control where an inbound caller is located, many teams default to the stricter approach: announce the recording clearly at the start of every call and make the announcement part of the QA standard itself.
- Agent awareness: employees who are monitored should know that monitoring happens, what is scored, and how the results are used. Transparency is both an obligation in many places and the difference between QA that agents accept and QA they resent.
- Retention: decide how long recordings and transcripts are kept, and delete them on schedule. Keeping everything forever is a liability, not a safety net.
- Payment and sensitive data: if card details, health information, or other sensitive data can be spoken aloud, you need a defined way to keep it out of stored recordings and transcripts, such as pausing recording during payment capture or automatic redaction.
- Access control: limit who can play back a recording, and log who listened to what. A recording library that anyone can browse is a problem waiting to happen.
Write the answers down as a short policy and reference it in your QA documentation. When a monitoring program gets challenged, the policy is what you are defending.
What to listen for on a call
The single biggest failure in call monitoring is reviewing calls without a fixed idea of what matters. A reviewer who listens for “was this a good call” produces an opinion. A reviewer who listens against defined criteria produces a measurement you can compare, trend, and coach on.
The four things every call review should cover
Compliance and process. Did the required things happen? The recording announcement, identity verification, mandatory disclosures, the correct steps for that call type. These are binary. Either it happened or it did not, and the transcript shows which.
Accuracy. Was the information the agent gave actually correct, and was the action they took the right one? An agent can be warm, fast, and completely wrong. Accuracy failures are the ones that generate repeat contacts and complaints, so they usually deserve the heaviest weight on the form.
Communication and tone. Did the agent listen without interrupting, acknowledge the customer’s situation, avoid jargon, and set clear expectations about what happens next? This is where calls differ most from written channels: pace, interruptions, dead air, and hold handling only exist on a call.
Resolution and next steps. Was the customer’s actual issue solved, and did they leave the call knowing what happens next and when? A call that ends politely with an unresolved issue is a failed call that scores well on tone.
Two things are worth listening for that rarely appear on forms. First, dead air and hold discipline: long unexplained silences are a common driver of poor call experiences. Second, the reason the customer called at all. Every monitored call is also a data point about why contacts happen, and that pattern is often worth more than the individual score.
A free call QA form you can copy
Here is a starting call QA form. Keep it to 8 to 12 criteria, weight it so the total is 100, and write every criterion so two different reviewers would score it the same way. If a criterion needs a debate to score, it is written badly. Adapt the weights to your business: a regulated financial support line will weight compliance far higher than a consumer retail line.
Section 1: Opening and compliance, 20 points
- Recording and identification announcement, 5 points: the required opening statement was given clearly and completely.
- Identity verification, 10 points: the customer was verified using the approved method before any account information was discussed.
- Required disclosures, 5 points: any disclosures required for this call type were given in full.
Section 2: Accuracy and resolution, 40 points
- Correct information, 15 points: everything the agent stated was accurate and consistent with current policy.
- Correct action taken, 15 points: the account changes, refunds, escalations, or orders made were the right ones and were completed.
- Issue resolved or correctly routed, 10 points: the customer’s actual issue was solved on the call, or escalated to the right place with full context.
Section 3: Communication, 25 points
- Active listening, 8 points: the agent let the customer finish, did not repeat questions already answered, and reflected the issue back accurately.
- Clarity and language, 7 points: plain language, no unexplained jargon, appropriate pace.
- Acknowledgment of the customer’s situation, 5 points: the impact on the customer was recognized before moving to the fix.
- Hold and dead-air handling, 5 points: holds were requested, explained, time-boxed, and thanked for, with no long unexplained silences.
Section 4: Close and documentation, 15 points
- Clear next steps, 8 points: the customer was told what happens next, by whom, and by when.
- Notes and disposition, 7 points: the call was logged accurately enough that the next agent could pick it up cold.
Auto-fail criteria, scored separately
- Missing or incomplete verification before account information was shared.
- Sensitive data mishandled, such as card details captured outside the approved process.
- Materially incorrect information that would cause the customer financial or account harm.
- Rudeness or dismissiveness toward the customer.
Auto-fails should be rare, unambiguous, and separate from the percentage score. Mixing them into the weighting hides them. Keeping them separate means one serious failure is visible even on a call that otherwise scored 90.
Two rules make this form work in practice. Every criterion needs a written one-line definition of what a pass looks like, and every score needs to point at the moment in the call that produced it. If you want the underlying anatomy of a scoring form, the guide to what a quality monitoring form is covers structure and weighting in more depth, and the QA scorecard templates give you variants for other channels.
How many calls should you monitor, and why sampling fails
The traditional answer is two to five calls per agent per month. That number did not come from statistics. It came from how many calls one reviewer can grade in a working week. It is a budget constraint that got repeated until it sounded like a standard.
Do the arithmetic on your own team. An agent taking 40 calls a day handles roughly 800 a month, so reviewing four of them covers half of one percent of their work. If a serious failure happens on one call in fifty, a four-call sample misses it most months, and when it finally appears nobody can tell whether it is an isolated slip or a quarter-long pattern.
The three ways a small sample misleads you
It is not random in practice. Reviewers pick calls that are convenient: recent, average length, from agents they are already reviewing. Calls that are very short, very long, or transferred are exactly where failures cluster, and they are exactly what convenience sampling skips.
It cannot support the decisions you make with it. A four-call sample gives a score with an error range wide enough to swallow the difference between your best and worst agent. Ranking a team, gating a bonus, or building a coaching plan on that number is guessing with a spreadsheet.
It flatters the average and hides the tail. Most teams do not have a general quality problem. They have a specific one: a call type, a policy, a shift, a queue. Averages from a tiny sample smooth exactly that signal away.
A practical cadence
If you are monitoring manually, target coverage where it changes decisions rather than a flat quota. New hires and agents in a coaching plan get frequent review. Complaints, escalations, and repeat contacts get reviewed every time, because those calls already told you something went wrong. Everyone else gets a genuinely random sample, chosen by a rule rather than by the reviewer.
If you can score automatically, the question changes entirely. You score 100% of calls and spend human review time only on the calls the scoring flagged, plus a calibration set to keep the standard honest. Coverage stops being the thing you ration. See what 100% QA coverage means for the fuller argument.
How to turn a monitored call into feedback that lands
A score with no conversation attached changes nothing. It is also the fastest way to make agents hate QA. The monitoring is the cheap part. The feedback is where the return is.
A structure that works
- Start with the agent’s own read. Ask how they thought the call went before you give your view. Agents identify their own issue more often than managers expect, and a self-identified issue is one they will actually work on.
- Name one strength, specifically. Not “good tone” but the exact moment: the way they reframed the policy at four minutes in so the customer stopped arguing with it.
- Pick one thing to fix. One. A form with four amber criteria produces a coaching session about one of them. The rest wait for the next session or resolve themselves once the main issue moves.
- Play the moment. Do not describe the failure, listen to it together. Thirty seconds of the actual call ends the debate about whether it happened and moves the conversation to what to do instead.
- Agree what different sounds like. Practice the alternative phrasing out loud in the session. “Be more empathetic” is not coachable. A sentence the agent has said once already is.
- Set the check-in. Name when you will look at this again and what you will look at. Feedback with no follow-up is an opinion, not coaching.
Two things to avoid. Do not save up a month of monitored calls for one long session, because feedback delivered weeks late has lost the context that made it useful. And give agents a route to challenge a score: a dispute path costs a few reversed scores and buys a team that treats the numbers as real. Our guide to giving quality feedback that works has scripts and examples for the harder conversations.
What to look for in call monitoring software
Call monitoring tools are easy to demo and hard to evaluate, because every product looks capable in a scripted walkthrough. These are the criteria that separate tools that get used from tools that get bought and abandoned. Use them as your own checklist regardless of which vendors you shortlist.
- Transcription quality on your actual calls. Everything downstream depends on the transcript. Test it on your worst audio, your accents, your product names, and your industry vocabulary, not on a clean demo recording.
- Does every score link to its evidence? A score you cannot trace back to a specific moment in the call cannot be coached on, disputed, or trusted. This is the single most useful question to ask in a demo.
- Can you build your own scorecard? A fixed vendor rubric will not match your compliance requirements or your definition of resolution. You need to define criteria, weights, auto-fails, and different forms for different call types.
- Coverage and cost model. Ask specifically what percentage of calls gets scored and what happens to the price when volume doubles. Coverage that is technically possible but priced out of reach is not coverage.
- Calibration support. Can several reviewers score the same call and compare, so you can see and close disagreement? Without this your standard drifts silently. QA calibration is what keeps scores comparable over time.
- Does it connect to coaching? If findings have to be copied by hand into another system to become a coaching session, the workflow breaks and the tool becomes a reporting exercise.
- Where it sits relative to your helpdesk or CRM. A tool your reviewers and team leads have to leave their main system to use gets used less. Check how deep the integration actually goes rather than whether an integration exists.
- Data handling and admin controls. Retention settings, redaction, access logs, and role permissions should be configurable by you and easy to evidence when compliance asks.
- Independence of the grader. Most QA and CX platforms now sell their own AI agents. If you are also deploying AI to handle calls, the system scoring those calls should not be the same system handling them, because that vendor ends up grading its own product.
Run a trial on real calls you have already scored by hand, and compare. A vendor’s accuracy claim on their own data is marketing. Agreement with your own reviewers on your own calls is evidence.
How automated scoring changes the economics of call monitoring
Every constraint described above traces back to the same root cause: a human has to listen to a call in real time to score it, so coverage is capped by reviewer headcount. Automated scoring removes that cap, and several things change at once.
Coverage stops being a rationing decision. When every call is transcribed and scored against your form, you stop debating whether four calls per agent is enough and start looking at the whole population. Outliers become visible because you can see the distribution rather than a handful of points from it.
Reviewer time moves from grading to coaching. The expensive, valuable thing a QA analyst does is not assigning points. It is spotting why a pattern exists and fixing it. When the grading is done, that time goes to the calls that need judgment, to calibration, and to the coaching conversations that actually move scores.
Consistency improves. The same rubric is applied to every call at 9am on Monday and 6pm on Friday. Human reviewers drift, disagree with each other, and score differently when tired. Automated scoring does not, which is why measuring agreement against a calibrated human standard, and re-checking it periodically, is the honest way to judge whether it is working.
The trend becomes reliable. A monthly quality score built from a 0.5% sample moves with sampling noise. Built from every call, it moves when quality moves.
Kaizo links every score back to the evidence in the transcript, so you can take a score apart on your own calls and see exactly what produced it instead of trusting a claimed accuracy figure. Kaizo also does not sell AI support agents, which is now unusual, so when it scores an AI-handled call it has no product of its own to defend. Full coverage is what makes both of those useful: scoring the whole population rather than a sample is what lets you verify the grading on the calls you pick, and what stops a reviewer’s selection deciding how a team looks. Kaizo scores against the scorecard your team defines, reads conversations natively from Zendesk and Salesforce, and routes what it finds into coaching rather than into a report nobody opens. At UiPath, Kaizo automated 100% of QA with 200% ROI and an 8% lift in quality score.
Automation does not remove humans from call monitoring. It removes them from the part of it that never needed a human: pressing play, ticking boxes, and hoping the four calls you picked were representative. For the wider program view, including how monitoring fits with scorecards, calibration and coaching, see the call center quality assurance guide and the guide to contact center quality assurance.
Frequently asked questions
What is call monitoring?
Call monitoring is the practice of reviewing customer calls against a defined quality standard to measure how service is delivered and to coach agents on it. It is done live while a call is in progress, by reviewing recordings afterward, or by scoring calls automatically against a scorecard. The output is a score, the evidence behind it, and a coaching conversation.
Is it legal to record and monitor customer calls?
Recording and monitoring calls is common practice, but the rules differ by country, by region, and by industry, and they cover notification, consent, retention, and how sensitive data is handled. This is not legal advice. Confirm the requirements that apply to every jurisdiction you take calls from with your own legal and compliance advisors before you record or monitor, and write the resulting policy down.
How many calls should you monitor per agent?
The common benchmark of two to five calls per agent per month is a budget constraint, not a statistical one, and it typically covers under one percent of an agent’s calls. If you are monitoring manually, prioritize new hires, complaints, escalations, and a genuinely random sample rather than a flat quota. If you can score automatically, review 100% and spend human time only on what gets flagged.
What should be on a call QA form?
A workable call QA form has 8 to 12 weighted criteria totaling 100 points, grouped into opening and compliance, accuracy and resolution, communication, and close and documentation, plus a small set of separate auto-fail criteria. Every criterion needs a one-line definition of what a pass looks like so two reviewers score it the same way. Weight accuracy and compliance most heavily, since those failures cause the most downstream damage.
What is the difference between live call monitoring and recorded call review?
Live monitoring means a supervisor listens while the call is happening, which allows immediate intervention and is useful for new hires and escalations, but it occupies a supervisor in real time so it never scales. Recorded review means scoring the call afterward against a form, which allows more careful and consistent scoring but is slow, so coverage stays small. Most teams need both, plus automated scoring for coverage.
How do you give feedback after monitoring a call?
Ask the agent how they thought the call went first, name one specific strength, pick one thing to fix, and play the actual moment from the call rather than describing it. Agree out loud what the alternative sounds like, then set a specific check-in. Deliver it within days of the call, and give agents a route to challenge a score they disagree with.
Related terms
See whether automated scoring agrees with your own reviewers
Bring a set of calls your team has already scored by hand, and we will show you how Kaizo scores every one of them against your own form, with each score linked back to the exact moment in the call so you can compare criterion by criterion. Kaizo does not sell AI agents, so it has nothing of its own to protect in the score.