An auto-fail, also called an automatic failure or a critical error, is a QA criterion that sets an entire evaluation to zero when it is breached, no matter how well the rest of the conversation went. It exists because some failures are not averageable: a compliance breach is not offset by good rapport and a fast resolution. A criterion earns that power only if it is objective, checkable against evidence in the transcript, unambiguous enough that two reviewers always agree, genuinely severe, and within the agent’s control. Anything failing those tests is a heavily weighted criterion, not an auto-fail.
In short
- An auto-fail zeroes a whole evaluation on one criterion, which makes it the highest-stakes scoring decision on your scorecard and the one most in need of a defensible bar.
- The bar is five tests: objective, evidence-checkable, unambiguous, severe, and within the agent’s control. Failing any one means the criterion belongs in the weighted section instead.
- If your auto-fail list is long, you do not have auto-fails, you have a weighting problem. Our view is that most teams should be able to justify five or fewer.
- Under automated scoring, tune auto-fail criteria for precision over recall: one wrong auto-fail costs more trust than several missed ones, because it makes every other score look arguable.
- Verify each auto-fail criterion against a set of conversations your reviewers have already scored before it goes live, and require every fatal score to trace back to the evidence in the transcript.
- Every auto-fail needs a named appeal route with a deadline, and the appeal rate on a criterion is itself a signal about how well it is written.
- After an auto-fail, coach the specific failure rather than the zero, and separate a one-off from a pattern before concluding anything about the agent.
What an auto-fail is, and why some failures are not averageable
Most QA scoring is an averaging exercise: a scorecard splits a conversation into criteria, each earns points, and the total says how good the interaction was. That works because most quality signals trade off against each other. A slow resolution is balanced by an unusually clear explanation.
An auto-fail is the exception. It says this one thing is not part of the average: if it is breached, the evaluation is zero regardless of what else happened. Teams also call it an automatic failure, a critical error, or a fatal error, and it usually sits in its own section at the top of the scorecard rather than as a weighted line item.
The logic behind zeroing a score
Some failures are categorically different from a low score. If an agent makes an account change without verifying who they are speaking to, that is not a mediocre interaction. It is one that should not have happened, and averaging it away says a warm tone compensates for a control that protects the customer.
So the test is not “is this important?” Nearly everything on a scorecard is important. The test is whether the failure is non-compensable. If an otherwise excellent conversation would still leave a reasonable person calling the interaction unacceptable, you have an auto-fail. If not, you have a weighted criterion.
The bar a criterion has to clear before it can zero a score
Because an auto-fail overrides everything else, it should be the hardest thing on your scorecard to qualify for. Here is the test we would apply, which a candidate has to pass in full.
1. Objective
It describes something that either happened or did not. No scale, no “to what extent,” no “sufficiently.”
2. Evidence-checkable in the transcript
You can point at the lines that prove the breach, or at the absence of lines that should have been there. If proving it needs context from outside the conversation, it cannot zero a score automatically.
3. Unambiguous enough that two reviewers always agree
The strictest test, and it eliminates most candidates. Not usually agree. Always agree. Run the criterion through a calibration session, and if two reviewers land differently even once, rewrite it or demote it.
4. Genuinely severe
One breach carries real consequences: regulatory exposure, a security or privacy risk, a commitment the business has to honor, or direct harm to the customer. Severity is not “management cares a lot about this.” It is that one event is disproportionate to any normal deduction.
5. Within the agent’s control
The agent could have done otherwise. If a step was impossible because a system was down or a policy was ambiguous, the failure belongs to the process, and auto-failing someone for the unavoidable is the fastest way to make the scorecard look unfair.
These are stricter versions of the tests in writing a QA rubric an AI can score, and this bar applies only to criteria you want to make fatal.
Worked examples: which criteria clear the bar and which do not
What matters below is the reason, not the verdict: the test a criterion fails tells you what would have to change for it to move.
| Criterion | Auto-fail or weighted | Why it falls on that side |
|---|---|---|
| Identity not verified before an account change | Auto-fail | Objective, visible in the transcript, severe, controllable. Reviewers always agree on whether it happened. |
| Required disclosure or consent language omitted | Auto-fail | The words were either said or not, and the omission can invalidate the interaction on its own. |
| Sensitive data mishandled, such as payment details taken in chat | Auto-fail | Binary and evidence-backed, and the exposure is not offset by anything else in the call. |
| A prohibited claim made, such as a refund the policy does not allow | Auto-fail | Checkable against written policy, and it commits the business to something it has to honor. |
| Rude or dismissive tone | Heavily weighted | Serious, but reviewers draw the line differently on the same transcript. Fails the always-agree test. |
| Poor empathy | Weighted, low | Not a single observable event, and not tight enough for an agent to argue against. |
| Missed opportunity to upsell | Weighted, low | A commercial preference, not a severe failure. |
| A required step skipped because the system was down | Not scored | Fails the control test. A process defect, routed to whoever owns the process. |
The five-item rule: a long auto-fail list is a weighting problem
An opinion, offered as an opinion rather than a benchmark: if your auto-fail criteria do not fit on one hand, you do not have auto-fails. You have a weighting problem.
The reasoning
Auto-fails get added one at a time, and each feels justified. An incident happens, someone asks how QA missed it, a new fatal criterion goes on the scorecard. Nobody ever removes one, because that feels like saying it does not matter.
The problem is what that does to the score. Once many criteria can zero an evaluation, the rest of the scorecard stops carrying information and improvement becomes invisible, because two agents with very different quality look identical once both have hit a fatal criterion. When a zero is rare it gets investigated. When zeros are common they get explained away.
How to shorten the list honestly
Name the specific consequence of one breach of each criterion. Not the category of risk, the consequence. If you cannot describe an outcome disproportionate to a scoring deduction, move it into the weighted section with a high weight, where it still hurts and still shows up in coaching.
Five or fewer is our recommended target, and a team sitting at fifteen should treat that as a finding rather than a configuration. Compliance-heavy operations may sit higher, and auto QA for compliance covers those criteria in more depth.
Auto-fails under automated scoring: tune for precision, not recall
All of that applies whether a human or a machine is scoring. What automation changes is the cost of being wrong. A grader can be wrong two ways on an auto-fail: it fires when it should not have, or it misses a real breach. These are not symmetrical.
Why a false-positive auto-fail costs more
A missed auto-fail is the same gap you already had with manual sampling, where most conversations were never read. Quiet, and fixable by tuning. A false positive is loud. An agent who did nothing wrong is handed a zero by a machine, and once the floor has seen the system get a fatal call wrong, every other score becomes arguable. “It got that auto-fail wrong last month” becomes the standing answer to any coaching conversation. You do not lose one evaluation, you lose the program’s credibility.
That asymmetry has a clear design consequence: auto-fail criteria should be tuned for precision even at the cost of recall. Write them narrowly, and where a case is borderline, route it to a human instead of letting the zero land automatically. A criterion that catches fewer breaches but is never wrong beats one that catches more and is sometimes wrong.
Verify before you switch it on
Freeze a set of conversations your reviewers have already scored, run the automated scoring over it, and measure precision and recall separately per criterion rather than as one accuracy number. Then decide what precision a criterion must hit before it may zero a score unconfirmed. Our validation protocol walks through that test.
One structural point belongs here. If the platform doing the scoring also sells the AI agents being graded, it is marking its own homework, and the fatal criteria are the most exposed. Kaizo is one of the few QA platforms that does not sell its own AI agents, so it has nothing to protect in the conversations it grades, and every score, including an auto-fail, traces back to the evidence that produced it. That is what makes a fatal score checkable rather than arguable.
Design the appeal path before the first auto-fail lands
Every auto-fail needs a route to challenge it, and it has to exist before the first zero is issued. An appeal process invented in response to the first angry agent reads as damage control.
Four things to decide in advance
- Who reviews: a named human, not “the QA team.” Not the agent’s direct manager, who has an interest in the team’s numbers.
- How fast: publish a deadline you commit to. We would suggest a couple of business days, because an unresolved zero shapes how someone works meanwhile.
- What happens to the score: decide whether an upheld appeal restores the underlying score or voids the evaluation. Either is defensible. Leaving the zero visible with a note attached is not.
- What happens to the criterion: a successful appeal is evidence about the criterion itself, so feed it back into the definition.
Appeal rate is a quality signal, not a nuisance
The most useful output is the pattern. Track appeals per criterion and the share upheld:
- Many appeals, mostly upheld: the criterion is firing wrongly. Take it out of automatic operation.
- Many appeals, few upheld: the criterion works, but agents do not understand it. A communication gap, not a scoring one.
- Almost no appeals: not automatically good news. Either the criterion is well written, or nobody thinks appealing is worth it. Ask the floor which.
Visible recourse is what a QA program agents trust is built on.
What to do with the score after an auto-fail
The zero is the start of the process, not the end. Handled badly it is a punishment with a number attached. Handled well it is the clearest coaching signal you get.
Coach the failure, not the zero
The score is a summary, and summaries are not coachable. The behavior is. Talk about the exact moment: what the transcript shows, what should have happened, and what got in the way. Not knowing the requirement is a training gap, skipping it under pressure is a workload issue, and a judgment call is worth understanding before correcting. The zero tells you none of that. The evidence does. Same principle as turning QA data into coaching, at higher stakes.
Do not average it away, and do not let it swamp the trend
Do not quietly drop auto-fails from the monthly average to keep the team number looking normal, which defeats the point of having them. Do not bury them in a blended score either, where one zero hides whether quality improved. Report them separately: quality on the weighted criteria, auto-fail count on its own line.
Separate a one-off from a pattern
One auto-fail from a strong agent is an incident. Three from the same agent on the same criterion is a pattern. The same criterion firing across many agents is not an agent problem at all, it is a process or training problem surfacing as individual failures. Reading auto-fails across 100% of conversations rather than a sample is what makes that visible: a 2% sample shows two incidents and nothing about whether they are a trend. At UiPath, Kaizo automated 100% of QA with 200% ROI and an 8% lift in quality score. Coverage is the precondition for reading auto-fails correctly, not a claim that nothing gets past you.
Common mistakes with auto-fail criteria
- Auto-failing on subjective criteria: rudeness and poor empathy feel severe enough to justify it, but reviewers will not agree on where the line sits. A fatal score that depends on who reviewed it is a lottery, not a standard.
- Letting the list grow after every incident: criteria get added after bad news and never removed. Review the list on a schedule and make each one re-justify itself.
- Auto-failing things outside the agent’s control: outages, ambiguous policies, and missing tooling produce failures that belong to the process.
- Turning auto-fails on for everyone at once: run them in observation mode first, where the criterion fires and gets logged but zeroes nothing.
- Issuing a fatal score with no evidence attached: if an agent cannot see the lines that produced the zero, they can neither learn from it nor challenge it. One of several auto-QA mistakes that kill agent trust.
- Tying auto-fails to pay or discipline on day one: a new criterion needs time to prove it fires correctly. Attach consequences once you have the precision data, and re-test it as policies change.
Frequently asked questions
What is an auto-fail in QA?
An auto-fail, also called an automatic failure or critical error, is a QA criterion that sets an entire evaluation to zero when it is breached, regardless of how the rest of the conversation went. It exists for failures that are not compensable: a compliance breach is not offset by good rapport or a fast resolution.
What should count as an auto-fail criterion?
Only a criterion that passes five tests: it is objective, it is checkable against evidence in the transcript, it is unambiguous enough that two experienced reviewers always agree, a single breach is genuinely severe, and it was within the agent’s control. Anything failing even one test belongs in the weighted section instead.
How many auto-fail criteria should a QA scorecard have?
Our view, offered as a recommendation rather than an industry figure, is five or fewer for most teams. Once many criteria can zero an evaluation, scores collapse toward full marks or zero, the rest of the scorecard stops carrying information, and zeros get normalized instead of investigated. A long auto-fail list is usually a weighting problem rather than a set of critical errors.
Can AI automatically apply auto-fail criteria?
Yes, and objective criteria such as whether identity was verified or a required disclosure was given are exactly what an automated system scores most reliably. The condition is that each criterion is verified against conversations your reviewers have already scored, so you know its precision before it can zero a score without a human confirming.
What is the risk of automating auto-fails?
False positives. A missed auto-fail is the same gap you already had with manual sampling, but an auto-fail issued wrongly hands a zero to someone who did nothing wrong, and once the floor has seen the system get a fatal score wrong, every other score becomes arguable. That is why auto-fail criteria should be tuned for precision over recall and routed to a human when the evidence is borderline.
Should an agent be able to appeal an auto-fail?
Always, and the route should exist before the first zero is issued. Decide in advance who reviews it, how fast they respond, and what happens to the score if the appeal succeeds. The appeal rate per criterion is a signal in its own right: many upheld appeals mean the criterion is firing wrongly, many rejected ones mean agents do not understand it, and almost none can mean either that it is well written or that nobody thinks appealing is worth it.
Related terms
See what your auto-fail criteria would actually catch
Bring your current critical error list and a set of conversations your reviewers have already scored. We will show you which criteria fire cleanly, which are too ambiguous to carry an automatic zero, and what each would have caught across every conversation rather than a sample. Every score traces back to the exact evidence in the transcript, so an auto-fail can be checked and challenged rather than taken on faith.