When to Stop Running Your QA Programme in a Spreadsheet

A spreadsheet is the right place to start a QA programme and the wrong place to keep it. The five failure signatures, and what to carry across when you move.
Guide · QA Programme Design

A spreadsheet stops being the right home for QA at the point where you need to prove something about the scores rather than simply record them. Recording verdicts is what a spreadsheet is good at; establishing that two reviewers agree, that the sample represents the work, and that a specific score can be traced back to the conversation and the rubric version that produced it are the three things it cannot do at all. Team size is a weak signal for this. Review volume, reviewer count and how often a score gets disputed are the ones that actually decide it.

In short

  • Starting in a spreadsheet is correct. A team that cannot describe its rubric in a spreadsheet is not ready to buy anything.
  • The trigger is not headcount. It is the first time somebody disputes a score and you cannot reconstruct how it was produced.
  • Separate what a spreadsheet makes painful from what it makes impossible. Only the second list is an argument for moving.
  • The three impossibilities are reviewer agreement, sampling audit, and traceability from a score back to a versioned rubric.
  • A spreadsheet has no memory of its own structure, so last quarter’s scores were produced by a rubric you can no longer read.
  • Most migrations quietly abandon the history and restart the programme at zero. Decide what you are carrying across before you move.

Is a spreadsheet actually the wrong tool for QA?

No, and starting anywhere else is usually a mistake. A spreadsheet forces you to write down what you are measuring, and a team that cannot express its QA rubric as a set of columns does not yet have a rubric. It has opinions.

The spreadsheet is doing four jobs well at the start, and it is worth being clear about them because they are the jobs any replacement has to keep doing:

  • It costs nothing and needs no approval, so the programme can exist before anyone has agreed it should.
  • It changes in seconds. Early on, the rubric is wrong and you need to be able to correct it during the meeting where you discover that.
  • Everybody can read it. No permissions conversation, no training, no seat cost for the team lead who needs to look at one thing once.
  • It makes the design visible. The weights are right there in the formula bar, which is more transparency than many tools offer.

So this guide is not an argument that spreadsheets are bad. It is an argument that they have a specific ceiling, that the ceiling is lower than most teams realise, and that hitting it looks like your own disorganisation rather than a tooling limit. That confusion is what keeps teams in a spreadsheet about a year longer than they should be.

Information

If you have not yet run a calibration session or written down how conversations get selected for review, tooling is not your constraint. Build the programme in the spreadsheet first. A tool applied to an undefined rubric produces the same confusion faster.

What are the five signs your spreadsheet has stopped working?

These arrive in a fairly reliable order, and each one is a symptom of a different underlying limit. Two or more at once means you are past the ceiling rather than having a bad month.

  1. Reviewers have started keeping their own copies. Someone made a personal tab because the shared one was awkward, and now two versions of the truth exist. This is the earliest signal and the most ignored.
  2. Nobody can answer what the rubric was three months ago. Columns got added and renamed in place, so historic scores were produced by a structure that no longer exists anywhere.
  3. A disputed score cannot be reconstructed. An agent asks why they lost points, and the honest answer is that the reviewer remembers deciding it. See how a score dispute should actually work for what the answer needs to contain.
  4. Reporting has become a monthly manual build. Somebody spends a day a month copying, pivoting and pasting into a deck. That day is the real licence fee you are already paying.
  5. Coverage is quietly falling and nobody decided that. Volume grew, review capacity did not, so the percentage reviewed drifts down without a decision being made or recorded.

Notice that only the fourth and fifth are about effort. The first three are about defensibility, and they are the ones that matter, because a quality programme whose outputs cannot be defended stops being used for anything consequential.

What does a spreadsheet make impossible, rather than painful?

This is the distinction that should drive the decision, and it is the one vendor material blurs. Painful things are a cost you can choose to keep paying. Impossible things are capability you do not have at any price.

What you need to do In a spreadsheet Verdict
Record a verdict against a rubric Straightforward, and arguably clearer than most tools Painful: no, it is genuinely fine
Report a monthly average by team A manual build, an hour or a day depending on discipline Painful. A real cost, not a blocker
Pull the conversations that were reviewed Manual linking, and links rot as tickets are merged Painful, degrading over time
Measure whether two reviewers agree Requires double-scoring the same conversations blind, which a shared sheet cannot enforce Impossible in practice
Prove the sample represented the work Selection happened outside the sheet and was never recorded, so it cannot be audited after the fact Impossible
Trace a score to the rubric version that produced it Columns were edited in place, so the old structure is gone Impossible
Review every conversation rather than a sample Not a tooling problem, a labour one Impossible at any realistic headcount

Why does the missing history matter more than the missing features?

A spreadsheet keeps your data and forgets its own structure. That sentence is the whole problem, and it is worth sitting with, because it is the failure that arrives last and costs the most.

When you add a criterion, rename a column, or change a weight, the change applies silently to everything the sheet can still calculate. There is no version, no effective date, and no record that the old structure ever existed. So March’s scores sit in the same grid as July’s while having been produced by two different instruments, and nothing in the file says so.

Three consequences follow, and they compound:

  • Trends become uninterpretable. Did quality improve in Q2, or did somebody reweight the scorecard criteria in April? If you cannot rule out the second, the first is not a finding.
  • Targets stop meaning anything. A 95% target set against one rubric is a different target against a stricter one, and moving the instrument while holding the number constant is how teams accidentally raise the bar without announcing it.
  • Disputes become unwinnable in both directions. You cannot show the agent the rules they were graded against, and they cannot show you they met them. Score traceability is the thing being lost here, and it is the foundation everything else in a quality programme rests on.

The underlying principle is not specific to support. The UK government’s AQuA Book, which sets out how analysis used for decisions should be assured, makes proportionality its central idea: the assurance around an analysis should scale with the consequence of that analysis being wrong. QA scores that feed coaching, bonuses or headcount decisions have crossed the threshold where a tool with no audit trail is a reasonable choice.

Attention

If your QA scores influence pay, promotion or performance management, an instrument with no version history is a genuine exposure and not merely untidy. The first serious dispute is the moment you find out, and by then the evidence you needed was overwritten months ago.

How do you decide? Use volume and stakes, not headcount

Team size is the signal everybody reaches for and it is close to useless, because a 12-agent team handling financial complaints has more assurance need than a 40-agent team answering order status. Three questions decide it.

How many reviews per week, and by how many people? One reviewer is a spreadsheet problem. Three or more reviewers is an agreement problem, and agreement cannot be measured in a shared sheet because you cannot stop reviewers seeing each other’s scores. Once you have multiple reviewers, the biggest source of error in your data is no longer agent behaviour, it is reviewer variance.

What decisions do the scores feed? If the answer is coaching conversations, a spreadsheet is defensible for a long time. If the answer includes pay, ranking or reporting to executives, the bar is higher, because those decisions get challenged and the challenge is evidential.

What percentage of conversations are you reviewing, and did you choose that number? Most teams land near a 3% sample by accident, as the residue of how much time reviewers happen to have. If your coverage is an outcome rather than a decision, you have already lost control of the thing the programme is supposed to measure. That is also the point where manual and automated review stop being interchangeable options: reviewing everything is not a harder version of sampling, it is a different activity, and full coverage answers questions a sample cannot be asked.

Tip

Run one diagnostic before deciding anything. Have two reviewers score the same ten conversations without seeing each other’s work, then compare. If they disagree on more than about two, your problem is calibration rather than tooling, and no purchase fixes it.

What must you carry across when you move?

The standard failure of a QA tool migration is not a bad tool. It is that the history stays behind and the programme silently restarts at zero, which costs you the one asset that took eighteen months to build. Decide these five before you move anything.

  1. The rubric, with its reasoning. Not just the criteria, but why each weight is what it is. If that only exists in one person’s head, write it down first, because migration is when it gets lost.
  2. A frozen copy of the old sheet, read-only, with a date. This is your only evidence for any pre-migration score. Never migrate by editing the original.
  3. The historic averages, labelled as a different instrument. Carry the numbers across, and carry the caveat with them. A chart that runs through the migration date without a marker is a chart that will mislead somebody in a board meeting.
  4. The calibration record. Whatever you have on how reviewers were aligned, since this is what you rebuild agreement from rather than starting over.
  5. The open disputes. Anything unresolved needs to be settled under the old rubric before the new one starts, or it becomes unsettleable.

Then set the new baseline deliberately. Run both in parallel for two to four weeks rather than cutting over, which is the same logic as shadow scoring: you want to know how the new instrument reads your work before it starts producing numbers people act on. Kaizo’s Auto QA is designed to take the scorecard you already use rather than replacing it with a vendor default, which is what makes a parallel run readable, since a difference between the two is then a difference in scoring and not a difference in rubric. The wider sequencing is in the QA tool rollout plan.

How do you move without losing the team’s trust?

This is the part that gets skipped, and it is the part that determines whether the new system is believed. A tooling change is read by agents as a change in how closely they are being watched, whether or not that is what it is.

Say what is changing and what is not, in that order. If the rubric is unchanged and only the recording moved, lead with that, because it is the question everyone is actually asking. If the rubric did change, say so plainly and separately, and do not let a rubric change hide inside a tooling change. Those two announcements should never be the same announcement.

Show one scored conversation end to end before go-live. Not a slide about the system: an actual conversation, with the criterion, the moment that triggered it, and the route to challenge it. That single demonstration does more for adoption than the training session, and it is the core of running a programme agents trust.

Do not backfill. Rescoring old conversations under the new system is tempting and it is corrosive. An agent’s March score was produced under March’s rules, and recalculating it in July destroys the property a quality programme depends on, which is that the number meant something at the time it was given. Start clean, keep the old record intact, and say out loud that you are doing both.

Coverage is worth being careful about in the same conversation. Moving from a sample to reviewing everything changes an agent’s experience more than any feature does, and 100% coverage revealing trends that 3% sampling never could is genuinely useful to the team as well as to you, but only if the first thing it is used for is finding process problems rather than people. Get that sequence wrong once and the tool is understood as surveillance for the rest of its life.

Frequently asked questions

Can you run a QA programme in Excel or Google Sheets?

Yes, and for a new programme with one or two reviewers it is the sensible choice. A spreadsheet forces you to define the rubric and costs nothing to change while that rubric is still wrong. It stops being sufficient when you need to prove something about the scores rather than just record them, specifically that reviewers agree, that the sample was representative, and that a given score can be traced to the rubric version behind it.

How many agents before you need a QA tool?

Headcount is the wrong test. A 12-agent team handling financial complaints needs more assurance than a 40-agent team answering order status. Use three questions instead: how many people review, what decisions the scores feed, and whether your coverage percentage was chosen or just happened. Three or more reviewers is usually the real threshold, because reviewer agreement cannot be measured in a shared sheet.

What is the biggest risk of keeping QA in a spreadsheet?

That it has no memory of its own structure. Columns get renamed and weights get edited in place, so historic scores were produced by a rubric that no longer exists anywhere. This makes trends uninterpretable, since you cannot rule out that the instrument changed rather than the quality, and it makes disputes unwinnable because you cannot show an agent the rules they were graded against.

Should we import our historic QA scores into a new tool?

Import them, but label them as having come from a different instrument and keep a frozen read-only copy of the original sheet. Never let a chart run through the migration date without a marker. And do not rescore old conversations under the new rubric: a score was produced under the rules in force at the time, and recalculating it later removes the reason anyone should trust the number.

How do we tell the team we are moving QA into a tool?

Separate the two announcements. Say what is changing about the recording and what is not changing about the rubric, in that order, and never let a rubric change travel inside a tooling change. Then show one real scored conversation end to end, including how to challenge it. If coverage is also increasing, use the first month of findings on process problems rather than on individuals.

See what your current scorecard looks like with the history kept

Bring the sheet you use today and a month of conversations you have already reviewed. We will run your own rubric across them and show you what changes when every score keeps its reasoning.

Book a demoExplore Agentic Auto QA

Choose your help desk

Not using either? We’ll let you know as soon as we can support your help desk solution.

Kaizo
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.