Peak Season Customer Service: How to Keep QA Honest

QA is the first thing cut when volume spikes, so visibility drops when risk peaks. How to resize sampling, pick what to relax, and score seasonal hires.
Playbook · Peak Season QA

Peak season quality management means protecting your QA programme while conversation volume rises, instead of quietly suspending it until the surge passes. It matters because reviewers are the first people redeployed onto the queue, so measurement stops at the exact moment your workforce, your contact reasons and your customers are all least familiar. The work is decided in advance: a review budget you can actually staff, a re-stratified sample, a written list of criteria you will relax and criteria you never will, and a post-peak review booked before peak starts.

In short

  • QA is the first thing cut during peak, because reviewers are the easiest people to redeploy onto the queue. Visibility disappears exactly when risk is highest.
  • Quality and volume genuinely trade against each other. The question is not whether you trade, it is what you trade and whether you decided it or drifted into it.
  • A sample sized for normal volume is not representative of peak. The population changes, so the stratification has to change with it.
  • Write the relax and never-relax list before peak. Making that call under pressure is how a standard quietly becomes a suggestion.
  • Seasonal agents need a shorter scorecard, not a softer one. A rubric built for tenured agents mostly measures tenure.
  • Book the post-peak review before peak starts. It is the only chance to find out what the relaxations actually cost you.

Why QA is the first thing cut when volume spikes

Every peak season plan has a staffing section, a scheduling section and a self-service section. Almost none of them has a quality section. That is not an oversight in the writing, it reflects what actually happens on the floor.

Here is the mechanism. Your reviewers are usually senior agents or team leads, which means they are queue-capable. When volume doubles, the fastest lever a manager has is to put every queue-capable person back on the queue. QA review is the only support activity with no same-day customer consequence, so it is the cheapest thing to stop. Nobody announces the decision. The weekly review count just falls, then falls again, and by week three the programme has stopped without anyone choosing to stop it.

Three consequences follow, and they compound.

  • Measurement stops when the inputs change most. You are running the least familiar workforce you will run all year, on contact reasons that did not exist last quarter, for customers who buy from you once a year. That is the period you understand least well and the period you are now measuring least.
  • The record gets a hole in it. Peak is the season you will most want to study in January, when you plan next year. It will also be the season with the thinnest quality data, because you were not reviewing.
  • The programme restarts cold. Reviewers who have not scored anything for six weeks have drifted apart from each other, and the first scores of the new year are the least trustworthy ones you will publish.

Be honest about the tradeoff rather than pretending it away. An hour spent reviewing is an hour not spent answering, and you cannot run a full-fidelity manual QA programme and clear a doubled queue with the same people. Anyone who tells you quality and volume are compatible by default has not staffed a December. The useful question is what you give up, and whether you chose it in September or discovered it in December. This is the same constraint that makes understaffing expensive in ways that never show up on the staffing spreadsheet, and the same one that makes scaling a support team harder than adding headcount.

One practical note on timing. If you are a retailer, the window you are planning against is the November and December period the National Retail Federation tracks as the winter holiday season, which means the quality decisions in this guide need to be made in late summer, not in November when there is no room left to make them.

What peak does to your sample, and why the old sample size lies

Most QA programmes review a small random percentage of conversations. That works when the population is stable. Peak is the one time of year it is not.

Four things change at once: who is answering, what customers are asking, who the customers are, and when the work happens. A random draw across the whole month will be dominated by your highest-volume contact reason handled by your highest-volume tenured agents, because that is what a proportional sample does. The sample will look healthy and will tell you nothing about any of the four things that changed.

The fix is stratification, not more reviews. Fix a review budget you can actually staff, then allocate it deliberately instead of drawing it at random.

A worked example

Say a normal month is 15,000 conversations and you review 3% of them, so 450 reviews. December is 40,000. Holding 3% means 1,200 reviews, which you cannot staff, and which is why the programme collapses instead of shrinking.

So go the other way. Decide the number of reviews you can genuinely complete, say 300, and spend them:

  • 120 on agents under 90 days. They may be 20% of headcount and less than 20% of volume, but they carry most of the variance.
  • 90 on peak-specific contact reasons. Delivery exceptions, gift orders, returns, promotion disputes. Reasons that barely existed in October.
  • 60 on business as usual, so you can still tell whether your baseline moved.
  • 30 on reopens and escalations, which are the cheapest early warning you have.

Three hundred well-chosen reviews beat 700 random ones, and 300 is a number a single reviewer can actually deliver alongside queue time. The point of QA sampling at peak is not statistical purity, it is making sure the parts of the operation you know least about are the parts you look at most.

What changes at peak Why a normal-season sample misses it What to do instead
Who is answering Seasonal and cross-trained agents are a small share of total volume, so a proportional sample barely touches them Set a minimum number of reviews per week for anyone under 90 days, independent of their share of volume
What customers ask about Peak-specific reasons are new, and a random draw over-samples the familiar top reason you already understand Stratify by contact reason and reserve a fixed share of the budget for reasons that did not exist last quarter
Who the customers are Once-a-year buyers do not know your policies or your product, and they escalate differently, but they are invisible in an agent-first sample Add a stratum for first-contact and first-purchase customers
When the work happens Extended hours and weekend cover fall outside the shifts reviewers work, so late and weekend conversations are systematically under-reviewed Sample by time block as well as by agent, and check coverage by hour in week one rather than in January
How long a conversation is Backlog makes threads longer and multi-agent, so one score lands on whoever happened to touch it last Score the handoff explicitly, or pull multi-agent threads out of individual scoring and review them as a process problem

Which criteria to relax, and which you never relax

Under a doubled queue, something gives. If you do not choose what, your agents will choose for you, and they will choose based on what gets measured and what gets shouted about in stand-up, which is almost always handle time.

There is one test for whether a criterion can be relaxed. Does failing it cost the customer something they will still notice after the conversation ends, or cost the business something it cannot undo? If neither, it is a candidate.

Criterion Peak status Why
Greeting, sign-off, brand voice, formatting Relax These are consistency criteria. A correct answer in four minutes beats a warm one in four hours, and peak customers say so
Proactive personalisation, upsell prompts, satisfaction-probing questions Suspend Discretionary value-add that assumes spare cognitive room. At peak there is none, so scoring it only manufactures failures
Documentation and tagging Relax the depth, hold the minimum Full notes are expensive. A ticket with no contact reason recorded is a ticket you cannot analyse in January, which is exactly when you will want to
Accuracy of the answer given Never relax A wrong answer at peak creates a second contact at peak, so it costs you twice at your most expensive moment
Identity verification, data handling, consent, regulated disclosures Never relax These are gates, not weighted criteria. A gate that bends under pressure was never a gate
Resolution, ownership, correct escalation and handoff Never relax Reopens and transfers are the mechanism by which December backlog becomes February backlog

Put the peak scorecard in writing before peak starts

The difference between a plan and a panic is that the plan is written down and dated. A verbal understanding that the team is “being pragmatic about tone right now” is not a standard, it is deniability, and agents can tell the difference immediately.

Produce one page, before peak, and publish it to the whole team including the seasonal cohort on their first day. It needs five things:

  1. A start date and an end date. The peak scorecard expires. Say when, or it becomes the permanent scorecard by accident.
  2. The list of relaxed and suspended criteria, named individually, with one sentence of reasoning each. Reasoning is what stops it reading as management giving up.
  3. The gates, restated, so nobody reads “we are relaxing the scorecard” as “we are relaxing everything”. Anything that should void an interaction belongs in your auto-fail list, not in the weighted pool where it can be traded off against speed.
  4. The review budget and how it is allocated, so anyone can see why a seasonal agent is being reviewed four times a week and a tenured agent once a fortnight.
  5. An explicit note that peak scores are a separate version. Scores produced under peak rules do not average into the annual number as though they were the same measurement.

That last point is the one teams skip and regret. If your QA scorecard changes and the score keeps the same name, you have quietly broken the only thing a quality number is good for, which is comparison over time. Version it, label the period, and keep the two comparable to themselves rather than to each other.

How to score seasonal and temporary agents fairly

A rubric built for tenured agents mostly measures tenure. A three-week hire will lose points on product depth, tool fluency, policy edge cases and judgement calls, all of which are functions of time on the job rather than of effort or care. Score them against it and you get a number that predicts nothing, demoralises a workforce you cannot afford to lose, and drags the team average down in a way that misleads you about your tenured agents too.

Three adjustments fix most of it.

1. Cut the scorecard down, do not soften it

Reduce to the criteria a competent adult can meet in week one: correct answer, correct process followed, correct escalation, respectful tone, verification completed. Five to seven items. This is closer to a QA checklist than to a nuanced rubric, and that is appropriate. A short scorecard applied strictly is fairer and more useful than a long one applied with silent allowances.

2. Report the cohort separately, always

Never mix a seasonal cohort into the team average. Publish two numbers. Mixing them tells you the team got worse when what actually happened is that the team got bigger and newer, and it hides whether your tenured agents held up under load.

3. Compare against a ramp, not a target

Define what acceptable looks like in week one and what it looks like in week three, then compare an agent to their own cohort at the same tenure. A seasonal agent at 72% in week one who reaches 84% by week three is a success. The same agent measured against a 90% team target is a failure on paper and will behave like one.

Two things about what you do with the score. Bring it forward: the first review of a seasonal agent should land inside their first 48 hours of live work, because that is when feedback still changes habits rather than correcting them. And it should arrive as a conversation, not a number, which is the difference between turning QA data into coaching and simply reporting it. A score delivered during peak that produces no coaching is administrative overhead you cannot afford this month.

Calibrate before peak, and once during it

Reviewer agreement drifts fastest under time pressure, and peak applies three pressures at once. Reviewers are rushing. They are scoring contact reasons they have never scored before. And they are applying a modified scorecard that nobody has used yet.

Two sessions are enough.

  • Pre-peak, on the peak scorecard. There is no point calibrating the normal rubric in October if you plan to use a different one in November. Run the session on last peak’s conversations if you kept them, using this peak’s rules, and settle the arguments before they cost you anything.
  • Mid-peak, short. Thirty minutes, three conversations, week two or three. The goal is catching drift while it is still correctable, not being thorough. A rushed calibration that happens beats a proper one that gets cancelled.

One detail that gets missed. If you have kept the programme alive by deputising team leads to review, those people are the least calibrated reviewers you have and the least likely to be invited to a session designed for the regular QA team. Invite them first. The mechanics are the same as any other QA calibration session, the difference is only who is in the room and which rubric is on the screen.

The post-peak review almost nobody runs

Book it before peak begins, for the second week of the quiet period. Not later. Memories fade fast, and the seasonal cohort will have left, taking with them the only people who can tell you what the onboarding actually missed.

It should establish three things, and it is worth being strict that it establishes all three rather than turning into a general debrief.

1. What the relaxations actually cost

Compare reopen rate, escalation rate and customer dissatisfaction across the relaxed period against the equivalent normal-season window. This is the single most valuable output and almost nobody produces it. It also cuts both ways, which is what makes it interesting: if suspending personalisation for six weeks cost you nothing measurable, that criterion may not deserve the weight it carries for the other forty-six weeks either. Peak is an unplanned experiment on your own scorecard. Read the results.

2. Which failures were process, not people

Peak concentrates process defects until they are impossible to miss. If forty agents all failed the same criterion in the same week, that is a policy, a macro, a knowledge gap or a broken system, and coaching individuals for it wastes the coaching budget twice over. Sort failures by criterion and by contact reason before you ever sort them by agent, which is the same discipline as any other root cause analysis.

3. Who to rehire

You will run peak again. Keep a per-agent quality record for the seasonal cohort with a rehire recommendation attached, written while you still remember. Most teams lose this entirely and re-recruit from zero every year, then wonder why the ramp is the same length every time.

Then archive the peak scorecard, note the version change on the timeline, and put the standard rubric back with a date. The programme should restart on a named day, not drift back.

What changes when review coverage does not depend on reviewer headcount

Every decision above exists because of one constraint: manual review capacity comes out of the same pool of people as queue capacity, so the two compete directly. Sampling, budgets and allocation are all ways of rationing a fixed number of reviewer hours. Change that constraint and some of the decisions disappear while others get sharper.

What goes away is the rationing. When conversations are scored automatically, stratification stops being a budget exercise and becomes a reporting filter. Week-one agents and peak-specific contact reasons become cohorts you can look at whenever you want, rather than allocations you had to argue about in September. This is the argument behind the line Kaizo puts on its own service pages, that 100% coverage reveals trends 3% sampling never could, and peak is where that difference is most visible, because peak is when the population changes fastest.

What does not go away is the judgement. What to relax, what stays a gate, how to compare a seasonal cohort to a tenured one, whether a failure is a person or a process. No system decides those, and any vendor implying otherwise is selling you out of a decision you still have to make.

And one thing gets harder. Consistency cuts both ways. A miscalibrated criterion applied to a 3% sample distorts a handful of conversations and a good reviewer quietly compensates. Applied to every conversation it distorts all of them, in the same direction, during your highest-volume weeks. So the pre-peak calibration matters more when scoring is automated, not less.

Which means the thing worth checking before peak is not coverage, since most platforms now offer it, but whether you can show a score was right when someone disputes it. A seasonal agent challenging a deduction three days before Christmas deserves the specific exchange that triggered it, not a total. Kaizo’s Auto QA applies your own scorecard to conversations inside Zendesk and Salesforce Service Cloud and leaves the reasoning attached to the conversation, so a peak-period score can be traced back and argued from evidence rather than from memory. Reporting through Kaizo Insights then lets you hold the seasonal cohort and the tenured cohort apart without rebuilding the report every week.

Frequently asked questions

How do you manage high support ticket volume while maintaining quality support?

You decide in advance which parts of quality you will trade and which you will not, then you write it down. Relax consistency criteria like greeting, formatting and proactive personalisation. Never relax accuracy, verification and correct escalation, because errors in those three create a second contact at your most expensive moment. Shrink the review count to a number you can actually staff and spend it on new agents and new contact reasons rather than drawing it at random.

Should you pause QA during peak season?

No, but you should resize it. Pausing removes visibility at the exact point your workforce, your contact reasons and your customers are least familiar, and it leaves a hole in the record for the season you will most want to study afterwards. Cut the review count to what one reviewer can genuinely deliver alongside queue time, cut the scorecard to the criteria that matter under load, and keep going.

How many conversations should you review during peak season?

Stop thinking in percentages and start with a number you can staff. A 3% sample of a normal month becomes an impossible workload when volume nearly triples, which is why programmes collapse rather than shrink. Fix an absolute review budget, then allocate it: roughly 40% to agents under 90 days, 30% to contact reasons that are new this season, 20% to business as usual so you can still see your baseline, and 10% to reopens and escalations.

How do you score seasonal customer service agents fairly?

Give them a shorter scorecard, not a softer one. Cut to five to seven criteria a competent person can meet in week one: correct answer, correct process, correct escalation, respectful tone, verification completed. Report the seasonal cohort separately from tenured agents so neither number misleads you, and compare each agent to their own cohort at the same tenure rather than to a team target built for people with two years on the job.

What is backlog in customer service?

Backlog is the set of conversations that have arrived and not yet been resolved, carried forward from one day to the next. It matters for quality because a backlogged thread tends to be longer, touched by more than one agent, and reopened more often, which makes it hard to attribute a score to any single person. Score the handoff explicitly, or pull multi-agent threads out of individual scoring and review them as a process problem instead.

When should you run the post-peak QA review?

Book it before peak starts, scheduled for the second week of the quiet period. Any later and memories have faded and the seasonal cohort has already left. It should establish three things: what the relaxed criteria actually cost you in reopens, escalations and dissatisfaction, which failures were process rather than people, and which seasonal agents you would rehire.

Go into peak knowing what your quality actually looks like

Bring last peak’s conversations and the scorecard you used. We will show you what your sample missed, where the seasonal cohort really sat against your tenured agents, and which of the criteria you relaxed turned out to matter.

Book a demoExplore Kaizo Insights

Choose your help desk

Not using either? We’ll let you know as soon as we can support your help desk solution.

Kaizo
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.