Humans Catch 9-26% of AI Errors: Your Human in the Loop Is a Receipt, Not a Control

The most repeated piece of AI safety advice in business is also one of the least effective. I have watched it fail in client after client, and the research explains exactly why.

Adding a human approval step to an automated workflow does not make it safer. In most implementations it makes the organisation feel safer while changing almost nothing about what actually happens — and it quietly relocates the blame.

I have had the same meeting perhaps sixty times in the last two years.

Someone senior asks the reasonable question: what happens when the AI gets it wrong? And someone else — usually the person who built the workflow, sometimes me in my first year doing this — says the magic words. There is a human in the loop.

Everyone relaxes. The meeting moves on. The risk item gets marked green. And in almost every case I have followed up on, nobody ever measures whether that human is doing anything at all.

I want to make an argument I have become increasingly confident about: for most business automations, the human approval step is not a control. It is a receipt. It produces a record that a person looked, without producing the outcome that a person looking is supposed to produce.

Start With the Number

The research on this is not new, not thin, and not ambiguous. It is just inconvenient enough that the automation industry has quietly declined to read it.

The finding that should end the conversation: when a genuine problem was surfaced directly in front of a human reviewer, intervention success held at only 9 to 26% across every oversight strategy tested. The human approved the bad action anyway, roughly three times out of four.

Read that again, because the framing matters. This is not the failure rate for problems that slipped past unnoticed. This is the failure rate for problems that were successfully put in front of a person whose entire job was to catch them.

The researchers named the mechanism a recognition bottleneck. The constraint is not attention — people saw the action. The constraint is recognition: seeing something and classifying it as a problem rather than rationalising it into the system's framing.

The gate exists. The audit log shows an approval. The bad action went through anyway, most of the time. That is not oversight. That is documentation of a decision nobody actually made.

Four Reasons It Fails, and None of Them Are Laziness

I want to be careful not to make this about bad reviewers. It is not. The people doing this work in my clients' businesses are conscientious and often extremely good at their actual jobs. The failure is structural, and it has four distinct mechanisms that compound.

1. Automation Bias

The tendency to over-trust an automated system: accepting its suggestions without sufficient scrutiny, and eventually ceasing to monitor it at all. It has been documented for three decades across aviation, radiology, criminal justice and hiring, long before anyone was deploying language models.

It is not a character flaw. It is what happens to any human supervising a system that is usually right. Being usually right is precisely what makes a system dangerous to supervise.

2. Approval Fatigue

Distinct from automation bias, and caused by volume rather than perceived authority. A reviewer processing hundreds of decisions per shift anchors on the first few, applies steadily lower scrutiny as the queue grows, and eventually treats approval as the path of least resistance.

Google DeepMind researchers identified this explicitly as a human-in-the-loop trap, documenting how high-volume approval requests degrade oversight quality over time. The trap is that the volume which makes automation worth doing is the same volume that destroys the review step you attached to it.

3. The Explainability Paradox

This is the one that surprised me most, because it contradicts advice I used to give.

Conventional wisdom says that showing a human the AI's reasoning will improve scrutiny. The evidence suggests the opposite: exposing the reasoning frequently increases deference rather than reducing it. A plausible explanation makes the conclusion feel more supported, not less.

I spent a year telling clients to surface model reasoning in their approval interfaces. I now think that advice was, at best, neutral and quite possibly counterproductive.

4. Institutional Design

Reviewers face large queues, throughput incentives, and organisational confidence in the system's accuracy. Under those conditions, reviewers converge on approving most outputs regardless of individual diligence.

If your approval step is measured on how fast the queue clears, you have not built a control. You have built a conveyor with a button on it.

The Study That Should Worry You Most

If the above only shows that human review adds little, this next finding shows it can subtract.

A study involving 450 clinicians gave them assistance from deliberately biased AI tools during diagnosis. Their accuracy fell from 73% to 61.7% — not because they lacked the knowledge, but because they deferred to the system.

These were trained experts, in a high-stakes domain, on decisions they were entirely capable of making correctly alone. Oversight existed on paper. It had been hollowed out by interaction design: confidence displays, suggestion formatting, and institutional culture all nudging toward agreement.

The uncomfortable implication is that a human placed in the loop of a subtly wrong system can perform worse than that human working with no system at all. "We have a human reviewing it" is not automatically a risk reduction. Under some conditions it is a risk transfer with a worse outcome.

The 5% Test

Here is the diagnostic I now run with every client, and it takes one query against your workflow logs.

What percentage of items does your human reviewer override, reject, or send back?

The threshold from the oversight literature is clear: sustained override rates below 5% in non-trivial domains typically indicate rubber-stamping rather than effective review.

Below 5%, you are not looking at a control. You are looking at a logging mechanism with a person attached. It may still be worth having — for accountability, for training data, for regulatory paperwork — but you should stop describing it to your board as a safety measure, and you should stop letting it substitute for the controls that would actually work.

Now the Part Nobody Wants to Hear

Set the 5% test beside the Automation Ratio — the percentage of AI-assisted outputs that ship without human correction — and something awkward falls out.

Suppose you tell me your Automation Ratio is 95%. Excellent. That means your reviewer is correcting 5% of what they see, which sits exactly at the rubber-stamping threshold.

Suppose you tell me it is 98%. Better still — and now your override rate is 2%, which is well inside the range the research says indicates a reviewer who has stopped genuinely reviewing.

You cannot simultaneously claim a high Automation Ratio and claim human review as your primary safety control. Those two claims describe the same number, and the number cannot be good news for both at once.

Yet I meet businesses making both claims in the same conversation, usually within about ninety seconds of each other. The pitch deck says the automation is 97% accurate. The risk register says a human reviews every output. Both cannot be doing the work being claimed.

There are only three honest resolutions. Your Automation Ratio is inflated and the reviewer is genuinely catching things. Or the ratio is accurate and the review is theatre. Or — and this is the good outcome — your reviewer is not reviewing everything, but is reviewing the small subset the system flagged as uncertain, at a volume where attention is sustainable.

Only the third one is a system. The first two are stories.

What the Approval Step Is Actually For

I want to name the thing that is usually happening, because I think most people know it and nobody says it.

In a large share of the deployments I see, the human approval step exists to answer a question about liability rather than a question about quality. If something goes wrong, someone's name is on the approval. The organisation has a person to point at.

Researchers have a term for this: the moral crumple zone. The human absorbs the impact of a system failure they had no realistic capacity to prevent. Analyses of documented cases — from Robodebt to the Mata v. Avianca filing — keep finding the same shape: the paperwork records a human choice that the human was not, in practice, positioned to make.

Even regulators have half-noticed. The EU AI Act mandates human oversight for high-risk systems, and then, in the same article, obliges providers to make overseers aware of their own tendency toward automation bias. It is a requirement that quietly concedes the requirement may not work.

And a body of research on government algorithms goes further, arguing that human discretion does not reliably improve outcomes even when reviewers have genuine agency rather than a rubber stamp.

If you are deploying a control principally so that responsibility has somewhere to land, that is a defensible business decision. Just do not confuse it with a control that reduces the frequency of bad outcomes, and do not let it crowd out the ones that would.

What To Build Instead

This is where I part company with most critiques of human oversight, which tend to stop at the diagnosis. The answer is not to remove humans. It is to stop putting them where they are cheapest to put and start putting them where they are effective.

Constrain Authority Rather Than Review Output

This is the whole argument in one line: a human approving what a system did is structurally weaker than a system that could not have done it.

Review is post-hoc and depends on human recognition under load — the exact thing the research says fails 74 to 91% of the time. Constraint is structural and depends on nothing.

Concretely, for any automated step:

  • Scope the credentials to the job. If the workflow only reads, it should not hold write access. Most integration tokens in the wild are far broader than the task requires, because scoping them properly took an afternoon nobody had.
  • Enumerate the permitted actions. An allowlist of things the system may do beats a reviewer checking what it did. The former cannot be fatigued.
  • Cap the blast radius. Rate limits, spend limits, record-count limits. A system that can touch ten records before it must stop and ask is safer than one reviewed by a tired person after it touched ten thousand.
  • Prefer reversible actions. Draft rather than send. Stage rather than publish. Flag rather than delete. Reversibility converts a category of failure from incident to inconvenience.
  • Make the system declare uncertainty. A workflow that escalates when confidence is low gives your reviewer twenty meaningful decisions instead of two thousand meaningless ones.

Put the Human Before, Not After

The highest-value human contribution to an automated process is specification — deciding what good output looks like, what the edge cases are, and what the system must never do. That work happens once, with full attention, and it compounds across every subsequent run.

Compare that to reviewing the four-hundredth output of the day, where attention is a depleting resource and the marginal value of each review approaches zero. Same person, same salary, wildly different return.

This is exactly the shape of the Bending Spoons model: humans doing specification and judgment, machines doing production. Not humans checking machines.

Audit the Distribution, Not the Item

Sampling beats reviewing. Pull twenty random outputs a week and examine them properly, with time and without a queue. You will learn more about your system's actual failure modes than from a reviewer approving everything with three seconds of attention each.

Better still, track the shape of the output over time. Sudden shifts in length, tone, refusal rate, tool-call frequency or retry rate tell you something changed upstream — long before any individual output looks wrong enough to reject.

Reduce What Needs Judgment at All

Most of what people route through an approval queue does not require judgment. It requires a rule that nobody wrote down. The four-category classification exists for exactly this: if a step can be expressed as a rule, expressing it as a rule removes the need for both the agent and the reviewer.

Every approval gate you can delete by making the underlying logic deterministic is a gate that cannot be rubber-stamped, cannot fatigue, and cannot fail silently at 4pm on a Friday.

Where I Think I Might Be Wrong

I hold this position strongly, so I should state the strongest version of the opposing case rather than a convenient one.

Human review does work under specific conditions, and they are identifiable. Low volume, where fatigue never sets in. High stakes per item, where the reviewer's attention is naturally engaged. Trained reviewers with genuine domain expertise and real authority to reject. Adversarial framing, where the reviewer's job is explicitly to find fault rather than to confirm. And accountability that is real rather than nominal.

Radiology second-reads, senior code review on critical systems, legal sign-off on material contracts — these can be genuine controls. The common thread is that the reviewer is not processing a queue.

There is also a legitimate argument that even a weak control is better than none, particularly during the first months of a deployment when you do not yet know the failure modes. I accept that. My objection is not to the existence of the step. It is to the step being counted as sufficient, and to the constraints that would actually work never getting built because the box is already ticked.

And in regulated domains you may simply be required to have a human decision-maker regardless of its efficacy. Fine. Have one. Just build the structural controls as well, and be honest internally about which of the two is doing the work.

What This Explains

I think this framework accounts for a set of statistics that have never fitted together neatly.

Why do 74% of AI agent deployments get rolled back? Partly because the oversight attached to them was theatre, so the first serious failure arrived with no warning and no containment — and the organisation concluded the technology was unreliable rather than that the controls were.

Why did 88.4% of organisations experience agent-related security incidents in a year when nearly all of them had approval workflows in place? Because approval workflows govern output, and security incidents come from authority.

Why do 80% of executives report no measurable AI ROI? Partly because a workflow with a human bottleneck on every output has not removed the labour, only relocated it — while adding a licence fee. You cannot get the economics of automation while paying a person to look at every unit of production.

That last one is the commercial argument, and for most readers it will matter more than the safety one. The rubber stamp is not merely ineffective. It is the reason your automation has not paid for itself.

What To Do This Week

One number, then one decision.

  • Measure your override rate. For each workflow with a human approval step, what percentage gets rejected or amended? If you cannot query this, that is itself the finding.
  • If it is above 15%, your reviewer is genuinely working — and your underlying system is weaker than you think. Fix the system.
  • If it is between 5 and 15%, you probably have real review. Protect it by keeping the volume sustainable, and resist the urge to route more through it.
  • If it is below 5%, stop calling it a control. Then pick one structural constraint from the list above and implement it this week.
  • Whatever the number, move one human upstream. Take the person currently approving outputs and give them two hours on specification instead. Compare the results after a month.

None of this is difficult. It is just unglamorous, and it requires admitting that a control everyone signed off on is not doing what its name implies. In my experience that admission is the hard part, not the engineering.

Boring disciplined moves compound. A reviewer who overrides nothing is not a safety net — it is a person being paid to agree with software, and a story the organisation tells itself instead of building the thing that would work.

If your first instinct on reading this was that your reviewer is different, I would gently suggest that is the most common response and also the least tested one. Run the query. The number will tell you more in five minutes than I can in four thousand words.

This is, in the end, the same argument I keep making in different clothes: most businesses do not have an AI problem, they have a systems problem. The approval queue is what a systems problem looks like when you try to solve it with a person.

Frequently Asked Questions

Are you saying we should remove human oversight entirely?

No, and I want to be precise about this. I am saying that reviewing every output is the least effective place to spend human attention, and that it frequently displaces controls that would work better. Move the human upstream to specification, and downstream to sampling and genuine exceptions. Keep them out of the middle of a high-volume queue, because that is where the evidence says their contribution approaches zero.

What if regulation requires a human decision-maker?

Then have one — that is not negotiable, and I would not advise otherwise. But build the structural constraints alongside it and be clear internally about which is doing the work. A compliance requirement satisfied is not the same as a risk reduced, and treating them as identical is how organisations end up surprised.

How do I know if my override rate is meaningful or just noise?

Look at the trend rather than a single reading, and look at what gets overridden. A healthy review function rejects a varied set of things for varied reasons. A rubber stamp rejects an occasional obvious formatting error and nothing else. If every override in the last quarter was cosmetic, the substantive review is not happening.

Doesn't showing the AI's reasoning help reviewers catch errors?

The evidence suggests it often does the opposite, which surprised me enough that I changed my own advice. A plausible-sounding explanation increases confidence in the conclusion rather than prompting scrutiny of it. If you keep reasoning displays, treat them as a debugging aid for your team rather than as a control that improves review quality.

Our reviewer is a senior expert, not a queue processor. Does this still apply?

Less so, and that is the genuine exception. Expertise, low volume and real authority to reject are exactly the conditions under which review works. The question to ask is whether that person is still operating under those conditions today, or whether volume has crept up quarter by quarter until they are processing a queue without anyone deciding they should be.

What is the single highest-leverage change here?

Scoping credentials to the task. It takes an afternoon, it cannot be fatigued, it degrades gracefully, and it addresses the failure class that approval queues are structurally incapable of catching. Everything else on the list is worth doing. That one I would do first.

Related Reading

The Automation Ratio Is the Only Metric That Predicts Survival — the number that, set beside your override rate, exposes whether your review step is real.

You Don't Need an Agent, You Need a Rule — the four-category classification, and how every rule you write removes an approval gate.

You Don't Have an AI Problem, You Have a Systems Problem — the parent argument this piece is a specific case of.

The 'AI Business' Advice Is Wrong: Tool Delivery vs Process Ownership — why owning an outcome forces you to build controls that work, and selling tools does not.

80% of Executives Report No Measurable AI ROI — the measurement gap, and how a human bottleneck on every output quietly guarantees it.

74% of AI Agent Deployments Get Rolled Back — the failure rate that looks different once you assume the oversight was nominal.

The Five Eyes AI Agent Security Guide — governance-first architecture, and why authority scoping beats output review on the security axis too.

About the Author

Hamza Baig is the founder of Hexona Systems, an AI automation agency serving clients across six continents, and the AI Automation Institute, a community of more than 40,000 entrepreneurs building with AI.

He has been featured in the GHL Top 50, Yahoo Finance and Brainz Magazine, and writes regularly on automation architecture, agent governance and the operational realities of AI deployment.

Read more analysis on the Hamza Automates blog, including the no-code automation workflow guide, or get in touch to discuss an automation build.

Follow @hamza_automates on Instagram for daily automation breakdowns.


About

Hamza Baig is the founder of Hexona Systems—an automation agency and softwareplatform that helps thousands of entrepreneurs and business owners implement AI-powered workflows at scale.

Share

Related Posts