When peak volume arrives you have three things you can give up: how much you review, how often you review it, and how strictly you grade. Give up the first two, because coverage and cadence are recoverable in January and a loosened standard is not. Criteria that protect identity, data, money, safety or a vulnerable customer never move at all, no matter what the queue looks like, because the consequence of failing them does not scale down just because you are busy.
In short
- There are three levers, not one: coverage, cadence and strictness. Almost every team pulls the wrong one.
- Cutting coverage is reversible and honest. Loosening the standard silently rewrites what your historic scores mean.
- Five criterion classes never relax: identity verification, data protection, financial accuracy, safety, and vulnerable-customer handling.
- Tone, formatting, structure and process-hygiene criteria are the legitimate candidates, and only for a declared period.
- If you do change the scorecard, version it with an effective date and annotate the peak window, or your January trend is unreadable.
- Relaxing a criterion for temporary staff and not permanent staff is a scoring decision that needs saying out loud, not a quiet exception.
Why is loosening the standard the wrong lever to pull?
You have three levers, and they are not equivalent. When the queue doubles and your reviewers get redeployed onto it, the instinct is to make the scorecard more forgiving so the numbers hold up. That is the one option of the three that does lasting damage.
- Coverage. How many conversations get reviewed. Cut this and you lose precision for a period, which is recoverable and obvious.
- Cadence. How often reviews and coaching happen. Cut this and you delay feedback, which is a real cost and also recoverable.
- Strictness. Where the bar sits. Cut this and every score before and after the change becomes incomparable, permanently.
The asymmetry is the whole argument. A month at 1% sample instead of 3% leaves you with a smaller, weaker but still honest dataset. A month where the bar moved leaves you with a dataset that reads as an improvement in quality when what improved was the marking. In January somebody will look at that line and draw a conclusion from it.
There is a second-order cost too. Standards that move under pressure teach the team that the standard is negotiable, and it does not become non-negotiable again in February just because you say so. That is a harder thing to rebuild than coverage.
Review fewer conversations. Do not review them more generously. The first is a resourcing decision you can explain and reverse; the second silently changes what every score in your history means.
Which criteria can actually be relaxed?
If you are going to move the bar anywhere, move it here, and only for a declared period. The test for a relaxable criterion is simple: does failing it cost the customer anything they will still care about next week? If not, it is a candidate.
| Criterion class | Example | At peak | Why |
|---|---|---|---|
| Presentation and structure | Greeting used, sign-off present, formatting tidy | Relax freely | Costs the customer nothing. This is the first and largest saving available |
| Efficiency and process hygiene | Correct macro used, tags applied, notes complete | Relax, with one exception | Internal cost only. Keep tagging if you need the peak data to diagnose anything afterwards |
| Tone and empathy depth | Personalisation, acknowledgement of feeling | Lower the bar, do not remove | Genuinely matters, but expecting peak-week warmth from an exhausted team is a target you set to fail |
| Proactive and value-add | Anticipated the next question, offered the relevant extra | Suspend entirely | This is what capacity buys, and at peak you do not have it |
| Resolution correctness | The answer was right and complete | Hold | The whole point. Relaxing this means accepting wrong answers as a policy |
| Identity and data protection | Verified before disclosing, no over-sharing | Never | Regulatory exposure does not scale down with your staffing |
| Financial accuracy | Correct price, refund, entitlement, charge | Never | Directly costly, and it generates the January contacts that make recovery worse |
| Safety and vulnerability | Recognised risk, followed the escalation route | Never | The consequence is not commercial |
Why do some criteria never move, whatever the queue looks like?
Because their consequence is not proportional to your workload. A missing sign-off in December costs the same as a missing sign-off in June, which is roughly nothing. A verification step skipped in December costs exactly what it costs in June, and possibly more, because peak is when impersonation attempts rise and when the person handling the contact is least likely to be experienced.
Data protection is the clearest case. Regulators do not recognise seasonal pressure as a mitigating circumstance, and accountability under data protection law is an ongoing obligation to be able to demonstrate compliance rather than a best-efforts commitment, as the UK regulator’s accountability and governance guidance sets out. A temporary relaxation of an identity check is not a quality decision, it is a compliance decision, and it is almost certainly not yours to take alone. The same applies to anything covered by customer data protection in your scorecard.
Put the never-relax set in the auto-fail gate before peak, not in the weighted pool. This is the structural version of the argument. A criterion sitting in the weighted pool can be survived: an agent fails it, loses points, and still finishes acceptably. A criterion in the auto-fail gate voids the interaction regardless of everything else, which is what you actually mean when you say it cannot be relaxed. Doing this in November means the December pressure has nowhere to push, and it removes the judgement call from a tired reviewer at eleven at night.
Keep the gate small while you do it. Three to five items, each answerable yes or no by two reviewers who have never spoken. If your list of things that can never be relaxed runs to fifteen, you have not prioritised, and the practical effect at peak will be that all fifteen get quietly ignored.
Relaxing an identity-verification or data-handling criterion is not a QA decision. If somebody proposes it because the queue is bad, that decision belongs with whoever owns compliance, in writing. A quality programme should not be the place this gets waived informally.
How do you relax a criterion without corrupting the data?
If you are changing anything, change it the way you would change an instrument you intend to trust again afterwards. Four steps, and the fourth is the one everybody skips.
- Version the scorecard with an effective date. Not an edit in place. The peak scorecard is a new version that starts on a date and ends on a date, and the previous version stays readable. This is what keeps a score traceable to the rules that produced it.
- Announce it to the team as a change, with the reason and the end date. Agents notice a scorecard change immediately whether or not you mention it. Announcing it costs nothing and not announcing it costs you the assumption of good faith that a programme agents trust runs on.
- Do not change the target at the same time. Change the instrument or change the number, never both in one week, or you will not be able to attribute the movement afterwards. If the criteria got easier, a stable score means quality fell.
- Annotate the window in the reporting itself. A shaded band on the chart with the dates and one sentence on what changed. Do this in December, because in February nobody will remember the exact dates and the line will get read as though it were continuous.
Restore on the date you announced, not when it feels calm. A relaxation with no end date is not a relaxation, it is a new standard that arrived without a decision. Put the restore date in a calendar with an owner before you start, and treat it as the first cadence commitment of the new year. Then run one calibration session on the restored rubric before it produces numbers anyone acts on, because reviewers drift during a period of deliberate leniency and they do not snap back automatically.
What about seasonal and temporary staff?
The common approach is to score temporary staff more leniently, informally, without recording that this is happening. That is the worst of the available options, because it corrupts the blended number and gives the temporary cohort no usable feedback either.
Three defensible options, and the choice matters less than making it explicitly:
- Same scorecard, separate reporting. Usually right. One standard, but the cohorts are reported separately so the blended figure is not a mixture of two different populations at two different points on their quality ramp.
- A reduced scorecard for the first weeks. Score only the criteria a two-week hire can reasonably be held to, and say which ones are excluded. This is honest and it is more useful for coaching, because the feedback is about things they can act on.
- Same scorecard, same reporting, different expectation. Defensible only if the expectation is written down. In practice this collapses into informal leniency, which is where you started.
Whichever you pick, the criteria to protect for a new cohort are the ones where they are most likely to fail and most likely to be blamed for a process rather than a behaviour. A temporary hire who gives wrong information because the knowledge base is stale has been scored for something outside their control, and at peak that is the most common single cause of a bad score. Separating agent error from process error is worth more during peak than at any other time of year, because the process is under more strain than the people.
Do not set a numeric quality target for a cohort that has been in the job under a month. Report their scores as a ramp against the previous cohort’s ramp instead. A 95% target applied to a two-week hire is a number designed to be missed.
What do you keep reviewing when you can barely review anything?
If coverage is the lever you are pulling, pull it deliberately rather than letting it decay. A reduced sample chosen well is more useful than the larger random one it replaces, because at peak you are no longer measuring average quality, you are hunting for the specific failures that generate January work.
Three populations earn their review time when almost nothing else does:
- Every conversation that touched the never-relax set. Refunds, verification, anything flagged as vulnerable. Not a sample of them, all of them. The volume is small and the consequence is the highest on the list.
- Repeat contacts. A customer coming back within a few days is the cheapest available signal that something failed, and it needs no survey response to detect. This is also the population that turns into your January backlog if it is not caught in December.
- Escalations, read for cause rather than for scoring.Escalation handling at peak tells you where the process is breaking, which is more actionable than any average, because a process fix helps every remaining conversation and a coaching note helps one agent.
What you drop, explicitly, is the routine random sample and individual coaching reviews for agents who are not in difficulty. Say that is what you are doing and when it comes back. If you would rather not lose the routine layer at all, that is the honest argument for automated scoring at peak: Kaizo’s Auto QA keeps the full population scored on your own rubric while your reviewers spend their reduced hours on the three lists above, which is 100% coverage revealing trends that 3% sampling never could, applied in the month when a sample is least representative. Either way, record the coverage decision. A percentage that fell because nobody had time is indistinguishable in the data from one that fell because you chose it, and only one of them is defensible when somebody asks in February.
Frequently asked questions
Should you lower QA standards during peak season?
No. Lower coverage and cadence instead, because both are recoverable and a loosened standard is not. A month at a smaller sample leaves you with a weaker but honest dataset, whereas a month where the bar moved makes every score before and after it incomparable, and the movement reads as a quality improvement when what improved was the marking.
Which QA criteria can be safely relaxed at peak?
Presentation and structure criteria such as greetings, sign-offs and formatting, most process-hygiene criteria, the depth expected on tone and empathy, and anything proactive or value-add. The test is whether failing the criterion costs the customer something they will still care about next week. Resolution correctness, identity verification, data protection, financial accuracy and vulnerable-customer handling never relax.
How do you score temporary or seasonal staff?
Use the same scorecard but report the cohort separately, or use a reduced scorecard covering only the criteria a short-tenure hire can reasonably be held to. What to avoid is informal leniency, which corrupts the blended number and gives the cohort no usable feedback. Do not set a numeric target for anyone under a month in the job; report their ramp against the previous cohort’s ramp instead.
How do you change a QA scorecard without breaking your reporting?
Version it with an effective date rather than editing in place, keep the previous version readable, and annotate the changed window directly on the chart with the dates and one sentence on what changed. Never change the scorecard and the target in the same week, because you will not be able to attribute the movement afterwards. Set the restore date before you start.
What should you still review when reviewer time is cut?
Every conversation touching refunds, identity verification or a vulnerable customer, all repeat contacts, and all escalations read for cause rather than for scoring. Drop the routine random sample and individual coaching reviews for agents who are not in difficulty, and record that you dropped them, because a coverage figure that fell by accident looks identical to one you chose.
Related terms
Go into peak knowing which standards are actually protected
Bring your current scorecard and we will show you which criteria are sitting in the weighted pool that should be gating, and what your coverage looks like when reviewer hours are cut in half.