Nesting is the supervised phase between classroom training and full independence, where a new hire handles real conversations with a trainer or senior agent close at hand. A nesting program works when those agents are evaluated against expectations built for the onboarding period rather than compared to tenured agents, and when every evaluation turns into a coaching conversation within days rather than weeks. Two structures hold up in practice: a dedicated nesting scorecard, or a single shared scorecard whose criteria are weighted by tenure.
In short
- Nesting fails as a measurement exercise when new hires are graded against the tenured team’s bar. Build the bar for the phase instead.
- Two models work. A dedicated nesting scorecard, or one shared scorecard where the weight of each criterion changes with tenure.
- Hold nesting agents to a higher quality target on a narrower set of contact types, then ramp volume. Quality first, throughput second.
- Read adherence per criterion across the whole cohort. One criterion sitting at 30% is a training defect, not five coaching problems.
- Coaching is the point of the scoring. If an evaluation does not produce a specific behaviour to change, it was an audit, not a program.
- Judge the program on whether adherence moved after coaching, and on what happens to volume and CSAT once nesting ends.
What is a nesting program, and how is it different from shadowing?
Nesting is the phase where a new hire handles live conversations themselves, under close supervision, before joining the general pool. Some teams call it incubation, guided practice or on-the-job training. It sits after classroom training and after shadowing, and it is the first time real quality data exists for that person.
The distinction that matters operationally is who is holding the conversation. In shadowing, the new hire watches someone experienced work. In nesting, the new hire does the work and someone experienced is available when it goes wrong. Everything about how you measure the phase follows from that difference, because only one of the two produces a transcript with the hire’s name on it.
Nesting length varies far more by contact complexity than by team size. Two to four weeks is typical for a narrow product with good documentation. Six to eight is normal where agents carry policy judgement, handle regulated work or cover many product lines.
Most published guidance stops at that definition. The harder question, and the one this page is about, is what you do with the conversations nesting produces: what standard to hold them to, whose scorecard to use, and how quickly the feedback has to come back to be worth collecting at all.
Why nesting agents need their own quality bar
The default is to run new hires through the same evaluation everyone else gets and compare the numbers. It feels fair and it is the wrong instrument, for three reasons that all push in the same direction.
The work is not the same work. Nesting agents get routed the simpler contact types deliberately. Scoring them on a scorecard built for the full contact mix means grading them partly on criteria their queue never gives them a chance to satisfy.
A comparison to tenured agents produces no action. Knowing a hire scores below the team average tells a manager something they already assumed. It does not say which behaviour to change on Monday, which is the only output that matters during onboarding.
The purpose of the score is different. For a tenured agent, a quality score is a performance signal. For someone in week two it is a diagnostic: it exists to find the gap while the gap is still cheap to close. Those two jobs need different targets, different weighting and different consequences.
The stakes are higher than a scorecard argument suggests. Gallup’s research on why the onboarding experience drives retention found only 12% of employees strongly agree their organisation onboards new people well, and the nesting weeks are where that judgement gets formed.
So the first design decision in a nesting program is not which tool to use. It is whether nesting agents get a separate scorecard, or the same one read differently. Both work. They fail in different ways.
Whichever you pick, the criteria underneath have to be specific enough that two reviewers grade the same conversation the same way. If the rubric you are scoring against is loose, a nesting trend line is measuring reviewer disagreement rather than learning, and you will coach the wrong thing with total confidence.
Model 1: build a dedicated nesting scorecard
In this model the nesting agents stay part of the same support team, but they are assessed against a scorecard written for the onboarding period. Criteria, targets and expectations are all set for where the hire actually is.
The target is higher than the team’s, not lower. One Kaizo customer running this model holds nesting agents to a 90% internal quality score while the organisation’s benchmark sits at 80%, and ramps the number of tickets solved over the phase. That inversion surprises people, so it is worth being explicit about why it works: the contact types are narrower and the supervision is closer, so a high bar is achievable. Holding it there teaches the standard before speed pressure arrives, and it means volume is the thing that ramps rather than quality being the thing that catches up later.
What the scorecard is actually for
Scoring on its own changes nothing. In this customer’s program each evaluation feeds a coaching conversation, built from that agent’s own strengths and weak spots rather than from a generic new-hire checklist. Kaizo’s AI-generated Coaching Cards assemble that from the scored conversations, which matters less as a feature than as a cadence: the constraint on coaching a nesting agent daily is almost never willingness, it is the hours a manager spends gathering evidence before the conversation.
Then the loop is closed at the level of the criterion rather than the person. Adherence Analysis shows, per criterion, what share of evaluated conversations met it across the whole nesting group, and which agents are pulling the number down.
In week 28 this customer found adherence to “Acknowledge Customer Concern” sitting at 30% across their nesting agents, with two agents below 50% on that criterion specifically. Targeted coaching took it to 75% over the following weeks. The number to watch is not the composite score, it is the criterion nobody is meeting.
That week-28 read is the argument for this model in one example. A composite score of 82% would have looked acceptable and hidden the fact that a specific, teachable behaviour was being missed in seven conversations out of ten. Criterion-level reading is what makes a nesting program a training instrument instead of a report.
Model 2: weight the same scorecard by tenure
The second model refuses the separate scorecard entirely. Every agent, new hire or veteran, is evaluated against one quality framework. What changes is how much each criterion moves the final score, weighted by the agent’s tenure group.
Newer agents are coached against the same standards, and the point deductions are simply more forgiving. A missed verification step is a miss for everybody, and it is discussed in coaching for everybody. It just costs a first-month agent fewer points than it costs someone in their second year.
The benefit is continuity. Because the criteria never change, the trajectory stays readable all the way from nesting into tenure. Nothing has to be re-explained at the handover, there is no moment where an agent discovers the real scorecard was different, and the manager is not maintaining two frameworks that drift apart. Only the weighting moves as expectations rise. If you already weight scorecard criteria deliberately, this is a small extension of machinery you have rather than a second system.
What the customer measured
Running this model, one Kaizo customer tracked agents moving from their first tenure group into the second and saw:
- 14% increase in ticket volume solved
- 21% improvement in customer sentiment
- 33% increase in CSAT
Read those together rather than separately. Volume and quality usually trade against each other during a ramp, and the pattern support leaders fear is throughput rising while the customer experience quietly degrades. Here all three moved the same way, which is the outcome a graduated framework is designed to produce: development is paced, but the definition of good never moves.
Tenure weighting only works if the tenure groups and their weights are written down and shared before anyone is scored against them. Discovering after the fact that your mistakes counted for more than a colleague’s does more damage to trust than a harsh score ever does.
Which model should you choose?
Neither is more rigorous than the other. They suit different teams, and the honest deciding factor is usually how many intakes you run and how much scorecard maintenance you can carry.
| Dedicated nesting scorecard | Tenure-weighted single scorecard | |
|---|---|---|
| Best when | Nesting is a distinct phase with its own restricted queue, and intakes come in defined cohorts | Hiring is continuous, tenure is a gradient rather than a phase, and one framework has to serve everyone |
| What the new hire sees | A clearly separate standard for a clearly separate period, with an obvious end point | The real standard from day one, with the stakes rising as they gain experience |
| Main advantage | Criteria and targets fit the narrow queue exactly, so every scored item is one they could actually have satisfied | Perfect continuity. The trajectory is readable from week one to year two with no break in the series |
| Main risk | The handover. Moving to the tenured scorecard can feel like a cliff if the transition is not planned | Weighting is invisible unless you publish it, and an unexplained weighting reads as favouritism |
| Maintenance cost | Two scorecards to keep calibrated and in sync as the product changes | One scorecard, plus a weighting table per tenure group |
| Where the coaching signal comes from | Criterion-level adherence across the cohort | Criterion-level fail rates compared against the next tenure group up |
How to run the coaching loop during nesting
Both models depend on the same loop, and the loop is where nesting programs usually break. Scoring is the easy half and the half that gets bought; coaching is the half that gets scheduled and then cancelled.
Coach the criterion, not the score
Telling a nesting agent they scored 78 is not coaching. The usable version names the criterion, points at the specific exchange that triggered the deduction, and says what to do instead next time. That is also the only form a hire can dispute, which matters more in week three than at any other point in a career, because an agent who concludes early that quality scores are arbitrary will hold that view for years. The mechanics of delivering quality feedback apply here with the volume turned up.
Shorten the gap between the conversation and the feedback
A habit formed in week one and reinforced a few hundred times costs far more to unlearn in week eight. Same-day or next-day is the target during nesting, against a weekly or fortnightly cadence for tenured agents. Nothing else on this list buys as much.
Run one calibration session on nesting conversations specifically before you make any decision off nesting data. Reviewer agreement is reliably worse on new-hire work, because there are more errors per conversation and reviewers differ on how much slack a beginner earns. Skip it and you will read reviewer variance as a learning curve.
Read the cohort before you coach the individual
When one hire misses a criterion, coach the hire. When most of an intake misses the same criterion, coaching them one by one is a misdiagnosis: the training module is the defect, and you will reproduce it next intake. This is exactly what the week-28 example above surfaced, and it is invisible on individual scorecards no matter how carefully you read them. Turning that group pattern into a plan is the same discipline as turning QA data into coaching, applied to an intake instead of a person.
All of this is runnable by hand and considerably harder when the data is thin. Under conventional QA sampling, a nesting agent’s first week is two reviewed conversations, which cannot separate learning from luck and will never show a criterion-level pattern across six people. With 100% coverage the cohort read becomes readable in week one rather than week six. Kaizo’s Auto QA scores against your own scorecard and keeps the reasoning attached to each conversation, so a nesting-week deduction can be traced back to the exchange that caused it rather than argued from memory, and coaching support is what compresses evidence-gathering enough to make a daily cadence survive contact with a real week.
How do you know the nesting program worked?
Most teams evaluate nesting by whether it finished on schedule. That measures the calendar. Four checks measure the program.
- Did adherence move after coaching? Pick the criteria you coached and compare adherence before and after, at the criterion level. This is the tightest available test of whether coaching changed behaviour rather than just occurring.
- Did quality hold when volume ramped? The whole design intent is quality first, throughput second. If the score falls as tickets rise, the ramp is too fast or the standard never really took.
- Did it hold after nesting ended? Check the first month on the tenured scorecard. A sharp drop is a handover problem, not a competence problem, and it usually means the nesting bar was set on criteria the general queue does not exercise.
- Did the next intake do better on the criterion you fixed? This is the one nobody runs. Training changes are made on instinct and almost never verified, which is why the same module leaks for years.
Where nesting programs go wrong
- Reading the composite score only. An acceptable average hides a criterion at 30% adherence. The average is the last place a training defect shows up.
- Leaving it ambiguous whether nesting scores count. Score during nesting, use the same reviewers, tag it as onboarding and keep it out of anything attached to pay or ranking. Then say so on day one. It is the ambiguity that damages trust, not the scoring. There is more on reading a new hire’s quality data week by week if you are setting the exit test.
- Changing the scorecard mid-cohort. It makes the trajectory unreadable, because improvement and a scoring change look identical. Version it and start new versions at intake boundaries.
- Benchmarking the cohort against your best agent rather than a realistic percentile of the tenured team, which sets a bar most of the existing team would fail.
- Treating the last day of nesting as the decision date. If the standard is met early, sign them off early. If it was never met, another week rarely fixes it.
None of this is exotic, and the reason it goes wrong so consistently is that nesting is the phase where everyone is busiest. The programs that hold up are the ones where the criterion-level read is a standing weekly habit rather than something a manager does when there is time, which there never is. If you are still assembling the scorecard itself, the do-not checklist for QA scorecards covers the design traps that make criterion-level reading impossible later.
Frequently asked questions
What does nesting mean in a call center?
Nesting is the phase between classroom training and full independence, where a new hire handles real customer conversations themselves while a trainer or senior agent stays close at hand. They usually work a restricted set of simpler contact types with a lower volume target. It is also called incubation, guided practice or on-the-job training, and it is the first period in which genuine quality data exists for that person.
What is the difference between nesting and shadowing?
In shadowing, the new hire observes an experienced agent working and handles nothing themselves. In nesting, the new hire handles the conversation and the experienced person is available when it goes wrong. Shadowing usually comes first and lasts days; nesting follows and lasts weeks. The practical consequence is that only nesting produces scoreable conversations with the new hire’s name on them, which is why the evaluation design matters from the first day of nesting and not before.
How long is nesting in a call center?
Two to four weeks is typical for a narrow product with strong documentation, and six to eight weeks is normal where agents carry real policy judgement, cover many product lines or work in a regulated environment. The length is a property of the work rather than of the hire. A better approach than a fixed duration is a written exit standard, so an agent who meets it in week three moves on in week three and an agent who has not met it by the scheduled end date does not get signed off by the calendar.
Should nesting QA scores count toward performance reviews?
No. Score during nesting using the same rubric and the same reviewers, but tag those scores as onboarding and keep them out of the ranked record and out of anything attached to pay or promotion. The data is genuinely valuable for finding where training leaked and unreliable as a measure of the person, because the sample is small and the queue is deliberately easier. Tell the hire on day one which scores count and which do not, because the ambiguity is what damages trust rather than the scoring itself.
What quality score should a nesting agent be held to?
Higher than the tenured team’s benchmark, not lower, provided their queue is genuinely restricted. One approach in practice is a 90% internal quality score target for nesting agents against an 80% organisation benchmark, with ticket volume ramping across the phase rather than quality catching up afterwards. The logic is that a narrow contact mix with close supervision makes a high bar achievable, and it establishes the standard before speed pressure arrives.
What does nesting refer to in contact center training?
It refers to the supervised live-work stage of onboarding, named for the idea of a bird leaving the nest in stages rather than all at once. In contact center training curricula it is normally the fourth stage, after orientation, classroom or virtual training, and shadowing. The stage itself is well covered in most training plans. What is usually missing is the measurement laid over the top of it: which scorecard applies, what target the agent is held to, and how quickly evaluations turn into coaching.
Related terms
See the ramp instead of guessing at it
Score every conversation a nesting agent handles from their first day, with the criterion-level view that shows which behaviour the whole intake is missing. See what your next cohort actually looks like.