All Four Big Four Firms Have Now Been Caught Publishing AI Slop

PwC published reports about agentic AI that were themselves generated by AI, complete with a governance framework that does not exist and governments that never deployed it.

PwC published reports about agentic AI that were themselves generated by AI, complete with a governance framework that does not exist and governments that never deployed it. Deloitte, EY and KPMG got there first. The lesson is not that AI writing is unreliable.

Four firms that sell AI adoption advice at premium rates have now each been caught shipping unverified AI output under their own logo. The technology worked exactly as designed. What failed was every process wrapped around it.

This is the most-shared business story of the week, and most of the sharing is schadenfreude. That is understandable and I am not above enjoying it. It is also a waste of a genuinely instructive failure.

Because if you strip out the brand names, what happened at PwC is the single most common way AI deployments fail in ordinary businesses — and the reason it made the Financial Times is not that it was unusual. It is that somebody checked.

What the FT and GPTZero Actually Found

The AI-detection firm GPTZero examined four PwC Middle East "thought leadership" reports published between 2024 and 2026, covering agentic AI, government transformation and electric vehicles. The Financial Times independently verified the findings — a conclusion echoed across the accounting trade press.

That independent verification matters, and I want to flag it early. GPTZero sells AI detection, so it has a commercial interest in finding AI. The FT does not. When a vendor's finding is confirmed by an outlet with no stake in it, the finding survives the obvious objection.

The Framework That Does Not Exist

The centrepiece is a 2025 report titled Transforming Governance. GPTZero scored it at an 84% likelihood of being entirely AI-generated — rising to 100% once the reference section was excluded.

The report contains an entire section on a framework called Citizen Pulse, described as a PwC methodology deployed by governments including Denmark, Saudi Arabia, the United States and Australia.

The framework does not exist. It was hallucinated in full, along with the citations supporting it. No public evidence backs any of the claimed government deployments.

A consulting firm published a document promoting its own product, and the product was invented by a language model. That is not a citation error. That is a marketing claim about a service the firm does not offer, made to governments the firm wanted as clients.

The Tells

The forensic detail is where this becomes genuinely useful, because these are the same tells you should be looking for in your own output:

GPTZero policy analyst Paul Esau described the pattern as chaotic signposting, and offered the sharpest diagnostic line in the whole investigation: no human is going to cite the same fact three times in two pages using three different sources.

That is the tell that generalises. Individual hallucinations are hard to spot. Structural incoherence — citations that do not correspond to anything, numbering that does not run in order, the same claim sourced three ways — is visible to anyone who reads the document once with attention.

Which tells you, with some precision, that nobody did.

PwC's Response

In fairness, the firm responded. PwC Middle East said it takes the accuracy of its published research seriously and is updating a limited number of supporting citations.

"A limited number of supporting citations" is doing an enormous amount of work in that sentence, given that one of the reports promoted a framework that does not exist. But the statement is on the record and it belongs here.

All Four. Every Single One.

PwC is not an outlier. It completes the set.

Four firms. Four independent detections. Same failure, same year.

When four separate organisations with different leadership, different processes and different geographies produce the identical failure, you are not looking at four lapses in judgment. You are looking at a structural incentive that all four share.

The Big Four pumped out hundreds of AI thought-leadership pieces to win clients while pushing their own staff to use AI to work faster. Volume was the objective. Nobody's job was to be the person who slowed it down.

The Same Week, PwC's Own CEO Explained How This Happens

Here is the detail that elevates this from embarrassing to genuinely instructive.

In the same week the FT investigation ran, PwC's US chief executive Paul Griggs was giving an interview about where companies go wrong with AI. Among the traps he named: bolting it onto broken processes.

He is right. That is exactly the diagnosis, and it is the argument I have been making in this newsletter since I started writing it. Most organisations do not have an AI problem, they have a systems problem, and adding a capable model to a process that was already producing unchecked output simply produces unchecked output faster.

PwC's thought-leadership pipeline was a process for generating documents to win mandates. Its quality control was presumably some combination of partner review and editorial sign-off. Introduce a tool that can produce a plausible forty-page report in an afternoon, and the bottleneck moves from writing to reviewing — except nobody resourced the reviewing, because reviewing was never the constraint before.

The chief executive of the firm articulated the failure mode in public during the same week his firm demonstrated it. I do not say that to score a point. I say it because it is the clearest possible evidence that knowing the diagnosis is not the same as having the controls.

This Is an Automation Ratio Failure

If you have read me before, you know the metric I keep returning to: the Automation Ratio — the percentage of AI-assisted outputs that ship without human correction.

PwC's Automation Ratio on these reports was, functionally, 100%. The output shipped without correction. Published, distributed to prospective government clients, used to solicit mandates.

And that is the trap in the metric, which I want to state plainly because I have been guilty of underselling it. A high Automation Ratio is only good news if you are measuring correctness. If you are measuring throughput, a 100% ratio means nobody is checking — which is indistinguishable from excellence right up until the Financial Times calls.

The number has to be paired with a verification standard or it measures the wrong thing. Ship-without-correction is only meaningful if somebody was genuinely in a position to correct.

The Contrast Worth Studying

Set PwC beside Bending Spoons: roughly 90% of pull requests AI-written, $2.57 million of revenue per employee, and no scandal.

Same technology. Same level of dependence on machine-generated output. Opposite outcome.

The difference is that software has a verification layer that prose does not. Code compiles or it does not. Tests pass or they fail. A hallucinated function name is caught in seconds by a machine, for free, before any human sees it.

A hallucinated citation is caught by a person who reads the source. There is no compiler for a footnote.

That is the actual lesson, and it is more useful than anything about consultancies. Domains with cheap automated verification can absorb enormous AI leverage safely. Domains without it need the verification built deliberately, because it will not happen by default — and the higher your output volume, the less likely it becomes.

The Review Process Was a Receipt

Somebody approved these reports. A firm like PwC does not publish externally without sign-off. There was a process, and the process produced an approval.

It just did not produce a reader.

I wrote earlier this week that a human approval step in most workflows is a receipt rather than a control — it generates a record that a person looked without producing the outcome that a person looking is supposed to produce. Research on oversight puts intervention success at somewhere between 9% and 26% even when a genuine problem is placed directly in front of a reviewer.

Four Big Four firms have now supplied the field evidence. These are organisations with more review infrastructure than almost any business reading this, and the reviews caught nothing — not the invented framework, not the footnote ending in chatgpt.com, not the same statistic cited three ways in two pages.

If their approval steps were receipts, yours probably are too. That is not an insult. It is the base rate.

The Uncomfortable Part

I would be a fraud if I wrote this article without addressing the obvious.

I publish thought leadership. I use AI to help produce it. So does much of my audience, and the entire premise of my business is helping people produce more output with AI assistance. The story I have just spent two thousand words on is a story about people exactly like me getting it catastrophically wrong.

So let me state the standard I hold rather than pretend the risk does not apply.

  • Every statistic in anything I publish is traced to a primary source before it ships. Not an aggregator summarising a primary source — the original. Aggregators get numbers wrong constantly, and I have caught it happening more than once this month.
  • Every link is opened. If I cite a report, I have read enough of that report to know the cited claim is in it. A URL that has not been opened is not a citation, it is a decoration.
  • Vendor statistics are labelled as vendor statistics. If a number comes from a company with a commercial interest in that number being alarming, that gets said in the sentence, not in a footnote.
  • Where a figure cannot be verified, it is either cut or explicitly flagged as unverified. Never smoothed over with confident phrasing.
  • Nothing goes out that I have not read end to end as a reader. Not skimmed for errors — read.

That standard costs time and it is the reason I publish less than I could. It is also the entire product. Thought leadership whose claims cannot be checked is not thought leadership; it is content marketing with a lower conversion rate and a much higher tail risk.

You are welcome to hold me to this. If you find a claim in anything I have published that does not trace to its source, I would genuinely like to know.

The reputational asset the Big Four spent a century building was the assumption that someone competent had checked. That assumption is now a question — and questions like that are far cheaper to answer before they are asked.

What To Do About It

1. Separate Generation From Verification

The person or system that produced a draft cannot be the thing that verifies it. That is true for humans and it is more true for models, which will confidently confirm their own fabrications when asked.

Practically: verification should be a separate pass, ideally by a different person, with an explicit checklist rather than a general instruction to review.

2. Make Citations Mechanically Checkable

This is the closest thing to a compiler that prose has. Every factual claim gets a source. Every source gets opened. A simple script that fetches every URL in a document and flags dead links, redirects and anything with a tracking parameter would have caught PwC's most embarrassing tell in under a second.

Nobody at four global firms wrote that script. It is perhaps forty lines.

3. Track Corrections, Not Output

If your team's AI metric is documents produced, you have built PwC's incentive. Track the correction rate alongside it — what percentage of AI-assisted output needed substantive fixing before shipping?

If that number is near zero, do not celebrate. Investigate. Near zero means either your process is exceptional or nobody is looking, and those two states are indistinguishable from a dashboard.

4. Use the Right Tier for the Job

A model asked to write a forty-page authoritative report from a thin brief will fill the gaps, because filling gaps is what it does. A model asked to restructure material you supplied, or to draft a section from sources you provided, has far less room to invent.

This is task-model matching applied to scope rather than to model choice, and it is also why some tasks want a rule rather than a model — a template with slots cannot hallucinate a framework.

5. Decide What You Are Actually Selling

PwC was selling AI-adoption advice while demonstrating AI-adoption failure. That gap is fatal precisely because the product was credibility.

If you sell process ownership rather than tool delivery, your published work is evidence of the process. It has to survive being checked, because being checked is the point.

The Broader Read

It would be easy to read this as a story about AI being unreliable. I think that is exactly backwards, and it is the reading I would push against hardest.

The models did what they do. Asked to produce an authoritative report with citations, they produced an authoritative-looking report with citation-shaped objects. That behaviour has been documented since 2023 and is not news to anybody at a Big Four firm.

What is news is that four organisations with effectively unlimited resources for quality control shipped the output anyway, repeatedly, over two years, in documents intended to win government business.

That is a governance failure with a technology component, not a technology failure. It is the same shape as 74% of AI agent deployments getting rolled back and the same shape as 80% of executives reporting no measurable AI ROI. In every case the tool worked and the system around it did not.

Gartner still projects $206 billion in AI agent spending this year. A meaningful share of that will be spent by organisations with exactly PwC's control environment, and some of them will find out the same way.

The firms that come out of this decade well will not be the ones that used AI most aggressively. They will be the ones whose published claims held up when somebody went and checked — which was always the business, before any of this.

Frequently Asked Questions

Does this mean AI-written content is unreliable?

It means unverified AI-written content is unreliable, which is a meaningfully different claim. The failure here was not that a model produced a plausible-sounding citation — that is expected behaviour. It was that four global firms published without anyone checking whether the citations pointed anywhere real.

How can I tell if something I am reading was AI-generated?

Structural incoherence is the strongest signal, more reliable than any writing-style intuition. Citations that do not correspond to claims, numbering that does not run in order, the same fact sourced three different ways, and links that redirect or die. Detection tools exist but produce false positives; opening the citations is slower and much more conclusive.

Should I stop using AI to write reports for clients?

No, but you should separate drafting from verification and treat every factual claim as unverified until someone opens the source. The organisations doing this successfully are not using less AI. They are checking more of what comes out of it.

Is GPTZero's detection reliable enough to accuse someone?

On its own, treat detection scores as a prompt to investigate rather than as proof — these tools produce false positives, and the firm sells detection. What makes this case solid is that the Financial Times independently verified the substantive findings: a framework that does not exist and citations that do not support their claims are checkable facts, not probability scores.

What is the single cheapest safeguard here?

Open every link before publishing. It is unglamorous, it takes minutes, and it would have caught the most damaging tell in the entire PwC investigation — a footnote URL with a ChatGPT tracking parameter still attached.

Why did this happen to all four firms at once?

Because they share an incentive rather than a flaw. All four scaled thought-leadership output to win mandates while encouraging staff to use AI for speed. When volume is the target and verification is nobody's specific job, the failure is a matter of time rather than of judgment.

Related Reading

The Automation Ratio Is the Only Metric That Predicts Survival — and why a ratio near 100% can mean excellence or mean nobody is checking.

You Don't Have an AI Problem, You Have a Systems Problem — the diagnosis PwC's own chief executive gave the same week his firm demonstrated it.

Bending Spoons: $2.57M Revenue Per Employee — the same reliance on AI output with the opposite result, and the verification layer that explains the difference.

The 'AI Business' Advice Is Wrong: Tool Delivery vs Process Ownership — why selling credibility means your published work has to survive being checked.

80% of Executives Report No Measurable AI ROI — the measurement gap that lets throughput masquerade as performance.

74% of AI Agent Deployments Get Rolled Back — the same governance failure in a different department.

The No-Code Automation Workflow Guide — building workflows with verification steps designed in rather than bolted on.

About the Author

Hamza Baig is the founder of Hexona Systems, an AI automation agency serving clients across six continents, and the AI Automation Institute, a community of more than 40,000 entrepreneurs building with AI.

He has been featured in the GHL Top 50, Yahoo Finance and Brainz Magazine, and writes regularly on automation architecture, agent governance and the operational realities of AI deployment.

Read more analysis on the Hamza Automates blog, or get in touch to discuss an automation build.

Follow @hamza_automates on Instagram for daily automation breakdowns.

Note: PwC Middle East's statement on the investigation is quoted above. Findings described here are those of GPTZero as independently verified by the Financial Times; readers are encouraged to consult both primary sources directly.


About

Hamza Baig is the founder of Hexona Systems—an automation agency and softwareplatform that helps thousands of entrepreneurs and business owners implement AI-powered workflows at scale.

Share

Related Posts