Eighty percent of executives report no measurable AI ROI. Everyone reads that as a failure rate. It is not. It is a measurement rate. Most of those companies never wrote down what the process cost before they automated it — which means they cannot tell a win from a loss, and they are killing and scaling projects essentially at random.
I have had the same meeting on four continents now. A client wants to know whether their AI investment is working. We open the deck. There are usage charts, token counts, adoption curves, a slide with a robot on it. Somebody has counted how many times the tool was opened last month.
Then I ask one question, and the room goes quiet.
How long did this take before?
Nobody knows. Not approximately, not directionally, not on a napkin. The process ran for years and produced revenue and employed people, and no one ever wrote down what it cost per unit, how long it took, or how often it had to be redone. So there is no before. And without a before, there is no after — there is only an after, floating unattached, meaning whatever the person presenting it needs it to mean.
Here is my position, and it is less comforting than the one everybody else is selling: your AI project probably did not fail. It probably did not succeed either. It produced a result that nobody can evaluate, in a company that will now make an expensive decision about it based on vibes.
Start with the numbers, because they are consistent across every research house that has looked. Gartner's 2026 research found that only 28 percent of AI use cases in operations fully succeed and meet ROI expectations, and it identified a consistent failure pattern: organisations deploy without baselines and then cannot prove value.
The same picture appears from the executive seat. Per IBM's 2026 research, only 29 percent of executives can confidently measure AI ROI today — and the primary reason given is an absent or imprecise pre-deployment baseline. Not a bad model. Not a failed integration. A missing measurement, taken before anything was built.
Read those two findings together and the standard narrative collapses. The problem is not that AI does not work. The problem is that most organisations built no mechanism capable of detecting whether it worked. Those are radically different diagnoses with radically different treatments, and the industry has spent two years prescribing for the wrong one.
The consequences are already priced in. Deloitte's 2026 State of AI research found 42 percent of companies abandoned at least one AI initiative in 2025, with average sunk cost reaching $7.2 million. That is a great deal of money spent on projects that were cancelled by people who, in most cases, could not have told you whether they were working.
Eighty percent report no measurable ROI. Note the word. Not 'no ROI'. No measurable ROI. The sentence has been telling us the truth the whole time and we keep reading past the adjective.
Most people hear my argument and relax slightly. If the projects did not really fail, that is good news, no? No. It is considerably worse, and this is the part I want to be unambiguous about.
A failure you can measure is cheap. You detect it, you stop, you keep the lesson, you move the budget. Failure with evidence is just expensive tuition.
An unmeasurable outcome is different. It does not stop anything and it does not teach anything. What it does is corrupt every subsequent decision you make, because you now have a portfolio of projects you cannot rank. And you will still have to rank them — the budget cycle does not wait for your epistemics. So you will rank them on other things: which sponsor is most senior, which demo was most impressive, which team tells the best story, which vendor sent the best deck.
Without a before number, you do not stop making decisions. You just start making them on charisma. The projects that get scaled are the ones with the best storytellers, and the ones that get killed are the ones with the most honest ones.
That is the real cost, and it compounds. Kill a working automation because its owner undersold it, and you lose the return plus the institutional belief that automation returns anything. Scale a broken one because it demoed well, and you multiply the loss across the business while congratulating yourself. Do both in the same quarter — which is common — and your AI programme is now actively selecting against the disciplined teams.
It gets structurally worse than bad decisions. One 2026 analysis for IT leaders found that 44 percent of enterprises fund AI investments from prior automation savings without validating whether those savings were actually delivered. Unvalidated past returns fund present investments, which then claim future returns against an equally unvalidated baseline.
Sit with the shape of that. Money is moving through the business on the strength of savings nobody confirmed, to buy more savings nobody will confirm, justified by the first set. It is a closed loop with no contact with reality at any point. The same analysis has a name for the resulting behaviour — faith-based AI investment, where programmes continue because they feel strategically important rather than because they demonstrate measured value.
And when the auditors eventually arrive, the failure is not what people expect. Per the same research, audit failure most often results from missing pre-programme baselines and undocumented attribution — not from inaccurate data. The numbers you have are usually fine. There is simply nothing to compare them to.
Here is the second half of the argument, and it is the part that makes people defensive, so let me be careful with it.
When you evaluate an AI process, you evaluate it against perfection. Every error is visible, logged, screenshotted, forwarded. Somebody says: look, it got this wrong. And that is treated as evidence about AI.
When you evaluate the human process it replaced, you evaluate it against memory. And memory says the humans were fine. The humans were fine because nobody counted the invoices that were keyed wrong, the emails that sat for three days, the report that got rebuilt twice because the first version used last quarter's numbers, the handoff that failed every time a particular person was on leave. None of that was logged, so none of it exists, so the incumbent process has an error rate of zero in the only place it is ever compared: somebody's recollection.
Your AI has a measured error rate. Your humans have a remembered one. Those are not comparable numbers, and the comparison is happening in every review meeting anyway.
I am not arguing that AI is better than your team. I have no idea whether it is, and neither do you — that is precisely the point. I am arguing that you are running a competition in which one contestant is scored on video review and the other is scored on nostalgia, and then drawing conclusions about the technology from the result.
The honest version of this is uncomfortable in both directions. Measure properly and some of your automations will turn out to be worse than what they replaced, and you will have to switch them off. Others will turn out to have been quietly saving real money that you were about to cancel. You do not get to know which is which in advance. That is what makes it measurement rather than opinion.
So here is the framework I now insist on before I will build anything for a client, and the one I would like you to steal. I call it the Before Number, because the name is the instruction.
A Before Number is a structured, pre-deployment capture of how a specific process actually performs today, taken before a single line of automation is written. The research consensus on what to capture is refreshingly boring: cycle time, error rate, labour hours per unit, and cost per transaction, measured across a 30-to-90-day window before go-live.
One warning from the same research, because it is the most common way this goes wrong: the most frequent baseline failure is measuring the wrong thing. Before capturing any metric, define the process boundary — the start point, the end point, and the handoffs in between. If the automation touches invoice processing, the boundary runs from invoice receipt to payment authorisation. Measure that, not "the finance team."
Four numbers are not enough on their own. Three conditions turn them into evidence.
First, an owner. One assessment of why enterprise AI spend fails to demonstrate return ranks the absence of a baseline first — describing it as fatal and irreversible, because if you did not measure the before state, no amount of after measurement can create a credible return. It ranks the absence of an accountable owner second, on the grounds that value does not capture itself: a pilot can produce useful output and still fail financially if nobody is responsible for banking or redeploying the value.
Second, kill criteria, written down before launch. The specific metric thresholds that end the programme. The same 2026 guidance for IT leaders calls executive ownership of kill criteria the single most effective governance control available to a CFO. Deciding what failure looks like after you have seen the results is not a decision. It is a negotiation with your own sunk costs.
Third, attribution. If your sales cycle shortens after you deploy an AI tool, how much of that is the tool and how much is a new hire, a seasonal pattern, or a pricing change? The answer requires controlled comparison — tracking work handled with AI against matched work handled without it. Teams that skip this overstate returns early, then cannot defend the claim when the CFO asks for the methodology.
Four numbers, one owner, kill criteria in advance, and a control group. That is the entire framework. It requires no new technology, no vendor, and no budget — which is exactly why it is the thing that gets skipped.
This is not theoretical. One documented case involved a logistics operator that defined specific, countable KPIs — freight diversions, empty miles, operational efficiency points — and captured baseline measurements for each before deployment, then routed every decision through an attribution framework so financial impact was captured automatically. The outcome was not just an improvement; it was a traceable one, where management could tie each efficiency point to a specific dollar figure.
Notice what made that work. Not a better model. Not a bigger budget. Somebody decided what to count, counted it before starting, and built the counting into the system rather than bolting a report on afterwards. The measurement was part of the architecture.
Almost nobody does. The rigorous version of this asks for several years of cost, volume and quality data plus a modelled counterfactual of what the trend would have produced without intervention. That is the standard for an audit committee.
You are not defending a bond prospectus. You need enough to tell a real improvement from noise. Thirty days of honest measurement on one process beats three years of reconstructed guesses across twelve. Start narrow and start now — the data you do not collect this month is the data you will not have next quarter, and the window closes permanently the moment you go live.
It will, by roughly a month. Weigh that against a 42 percent abandonment rate at an average sunk cost of $7.2 million. One assessment of the pattern puts it plainly: enterprise AI ROI fails when pilots are measured after launch instead of designed for measurement before launch.
And be honest about where the pressure comes from. The reason baseline capture gets cut is not that it is slow. It is that it is unexciting, it produces no demo, and it occasionally reveals that the process you were about to automate is not worth automating. That last one is not a bug. That is the framework doing its most valuable work, on the cheapest possible day.
Some genuinely are not, and I am not going to pretend otherwise. Morale, optionality, the value of your team learning to work with these tools — real, and hard to put a number on.
But notice how that argument gets used. It almost never appears at the start of a project, when it would be an honest scoping decision. It appears at the end, when the hard numbers came back ambiguous, as a retreat position. Serious frameworks handle this deliberately by pairing financial metrics with operational ones and knowing which to prioritise at which stage — not by reaching for the unquantifiable only after the quantifiable disappoints.
The rule I use: if a benefit is unquantifiable, say so in writing before you start, and say what would count as evidence anyway. An unquantifiable benefit named in advance is a legitimate strategic bet. An unquantifiable benefit discovered in the results meeting is an alibi.
Then measuring will be quick and you will win the argument permanently, with numbers, in front of the person holding the budget. If you are right, this costs you a month and buys you every subsequent approval. Reluctance to measure something you are certain about is not confidence. It is a preference for the certainty over the answer.
Regular readers know I argue that the Automation Ratio — the percentage of AI-assisted outputs that ship without human correction — is the number that predicts whether AI investment returns anything. The Before Number is not a competing idea. It is the prerequisite that makes the ratio mean something.
An Automation Ratio of 70 percent is meaningless in isolation. If the human process shipped clean 95 percent of the time, you have made things worse and dressed it as progress. If the human process shipped clean 40 percent of the time, you have transformed the operation. Same ratio, opposite conclusions, and the only thing that distinguishes them is a number you had to capture before you started.
The Automation Ratio tells you where you are. The Before Number tells you which direction you travelled to get there. One without the other is a coordinate with no map.
It also closes the loop on the argument I have made from the beginning: you do not have an AI problem, you have a systems problem. A company that cannot state what its own core process costs per unit does not have a measurement gap in its AI programme. It has a measurement gap in its business, which the AI programme merely made visible. The automation did not create the blindness. It walked into it.
And it explains why 74 percent of agent deployments get rolled back with such consistency. Some of those rollbacks are correct. But a rollback decision requires knowing the thing was worse than the alternative, and most organisations rolling back cannot demonstrate that. They are reverting to a process whose performance they also never measured, on the basis of a comparison they cannot make.
Concrete, in order, and none of it requires a purchase.
That last one is the hardest and the most valuable. Half the AI opinions inside most companies rest on projects nobody could evaluate. Retiring those claims — in both directions, the triumphant and the dismissive — clears the ground for actual decisions. And if you are building the process from scratch, the workflow discipline comes before the tooling, not after.
The industry has spent two years arguing about whether AI delivers returns. It has been the wrong argument, conducted by people who could not have settled it either way, using evidence that does not exist.
The finding is consistent everywhere anyone has looked: organisations deploy without baselines and then cannot prove value, and the primary reason executives cannot measure AI ROI is an absent pre-deployment baseline. Not model quality. Not integration difficulty. Not the vendor. A number nobody wrote down, on a Tuesday, before the project started, because writing it down felt like a delay.
That is oddly good news, because it is the cheapest problem in this entire field to fix. It costs one person, thirty days, and one spreadsheet. It requires no frontier model, no new platform, no migration, and no budget approval. And it is skipped almost universally, which means doing it is a genuine competitive advantage available to anyone reading this by the end of the month.
Everyone is arguing about which model to use. Almost nobody can tell you what their process cost before they used any of them. The second question is worth more than the first, and it has been sitting there unanswered the whole time.
The endpoint of doing this properly is not a better model. It looks like Bending Spoons — $2.57 million of revenue per employee, humans doing specification and judgment, machines doing execution. You do not get to that by picking well. You get there by knowing, precisely, what every process costs and what changed when you touched it. That is not glamorous work. It is just the work.
So before you evaluate another model, ship another agent, or sit through another deck with a robot on the slide, answer the question that stopped the room in every one of those meetings. How long did this take before? How often was it wrong? What did it cost?
If you cannot answer, you do not have an AI result to interpret. You have an anecdote with a budget attached. Go get the before number. Everything else you do afterwards will finally mean something.
A structured, pre-deployment capture of how a process performs today, taken before any automation is built. The research consensus is four measurements across a 30-to-90-day window: cycle time, error and rework rate, labour hours per unit, and fully-loaded cost per transaction — scoped to a clearly defined process boundary with a stated start point, end point and handoffs. Without it, post-deployment numbers are assertions rather than evidence.
Only partially, and you should be honest about the limits. You can sometimes recover cost and volume data from finance systems, and error rates from support tickets or rework logs if anyone kept them. What you cannot recover is the counterfactual — what the process would have done anyway. The practical move is to stop citing that project as evidence in either direction, capture a proper baseline for the next one, and where possible run a control group now so you have something to attribute against going forward.
Weigh it against the alternative. Deloitte found 42 percent of companies abandoned at least one AI initiative in 2025 at an average sunk cost of $7.2 million, and the pattern is that pilots get measured after launch instead of designed for measurement before it. A month of baseline capture is cheap insurance against a seven-figure write-off you cannot even diagnose.
Kill criteria are the specific metric thresholds, defined before launch, that end a programme. One 2026 guide for IT leaders calls executive ownership of kill criteria the single most effective governance control available to a CFO. Setting them in advance matters because deciding what failure looks like after seeing the results is not a decision — it is a negotiation with your own sunk costs, and sunk costs always win that negotiation.
Controlled comparison. The standard approach is tracking cohorts of work handled with AI against matched cohorts handled without it, so improvements can be attributed rather than assumed. If your cycle time drops after deployment, a new hire, a seasonal pattern or a pricing change could each explain it. Teams that skip attribution overstate returns early and cannot defend the methodology when finance asks.
The Before Number is the prerequisite. The Automation Ratio — the share of AI outputs shipping without human correction — tells you where you are, but a ratio of 70 percent is a triumph against a 40 percent human baseline and a disaster against a 95 percent one. Same number, opposite conclusions. Capture the baseline and the ratio becomes decision-grade; skip it and the ratio is a statistic you can argue about forever.
Pick one process you have not automated yet and start counting: cycle time, rework rate, hours per unit, cost per unit. Thirty days, one spreadsheet, one named owner. It requires no budget, no vendor and no technology, and it is the only work in your AI programme whose window closes permanently — the moment that process goes live, the chance to measure its before state is gone for good. Everything else can be done later. This cannot.
Continue the argument:
1. The Automation Ratio: The Metric That Predicts Survival — the number the Before Number makes interpretable.
2. 80% of Executives Report No Measurable AI ROI — the statistic this article re-reads from the other direction.
3. You Don't Have an AI Problem, You Have a Systems Problem — why a missing baseline is a business gap, not an AI gap.
4. The 'AI Business' Advice Is Wrong: Tool Delivery vs Process Ownership — owning a process means knowing what it costs.
5. 74% of AI Agent Deployments Get Rolled Back — rollback decisions made without the comparison that would justify them.
6. Satya Nadella's Learning Loop Warning — what happens when you stop being able to see your own operation.
7. Gartner's $206 Billion AI Agent Spending Forecast for 2026 — the scale of spending this discipline is protecting.
Hamza Baig is the founder of Hexona Systems, an AI automation agency operating across six continents, and the AI Automation Institute, where he has trained more than 40,000 entrepreneurs in practical AI systems design.
He has been featured in the GHL Top 50, Yahoo Finance and Brainz Magazine. His work focuses on the gap between what AI can do in a demo and what it reliably ships in production — the Automation Ratio, task-model matching, and governance-first architecture.
Read more analysis on the Hamza Automates blog, get in touch about your automation stack, or follow @hamza_automates on Instagram for daily breakdowns.
Research referenced: Gartner 2026 operations AI success rates; IBM 2026 executive ROI measurement research; Deloitte 2026 State of AI in the Enterprise; 2026 enterprise AI ROI measurement guidance for IT leaders; enterprise AI ROI failure-pattern assessments; and published baseline-capture frameworks. Figures are as reported by the cited secondary sources. Opinions are the author's own.
Further reading on measurement practice: pre-deployment baseline frameworks, common ROI measurement pitfalls, and the model-layer discipline that makes measurement portable across vendors.
Hamza Baig is the founder of Hexona Systems—an automation agency and softwareplatform that helps thousands of entrepreneurs and business owners implement AI-powered workflows at scale.