Skip to content
AI.info

Future Horizons

The AI ROI Reality Gap: Why the Promised Productivity Revolution Has Not Arrived

McKinsey's 2026 survey found 37% of companies report any EBIT impact from AI and Gartner found 22% have scaled it. A year of heavier spending bought more scaling and no more earnings.

The AI ROI Reality Gap: Why the Promised Productivity Revolution Has Not Arrived

Gabriele Masetti ·

The bill has come due

Enterprises had poured somewhere between $30 billion and $40 billion into generative AI initiatives by the time the MIT NANDA project published its "State of AI in Business 2025" report in August 2025. The return on that spending, measured where it actually matters — the profit-and-loss statement — is, for the overwhelming majority of companies, close to zero.

Based on 52 structured interviews with executives, a survey of 153 leaders, and an analysis of more than 300 public AI deployments, the NANDA researchers concluded that 95% of enterprise generative-AI pilots are delivering no measurable P&L impact. Only about 5% of organizations are seeing what the report calls rapid, transformative returns.

That is not a claim that the technology doesn't work. Ask a software engineer using an AI coding assistant, or a customer-support team running a well-scoped chatbot, and you will hear that the tools are frequently excellent at the narrow task in front of them. The MIT finding is something more specific and more uncomfortable for corporate boards that approved eight- and nine-figure AI budgets on the strength of vendor demos: excellent task-level performance is not converting into enterprise-level financial results.

That is the ROI reality gap, and it is visible in report after report from Gartner, S&P Global Market Intelligence, McKinsey, and Boston Consulting Group, all pointing at the same seam — not whether the models are capable, but whether companies know how to turn capability into cash. A year on, the 2026 survey rounds have not closed it.

The gap is a distinct problem from the macroeconomic productivity puzzle that has occupied economists since the 1980s — the observation that transformative technologies show up everywhere except the GDP statistics, and the argument that AI, like electricity before it, needs a slow rewiring of business processes before its effects appear in aggregate output. That debate is real and worth having on its own terms.

But it is a story about national accounts and long time lags. The ROI gap is about what is happening right now, inside individual companies, between the AI budget line and the earnings call — and the news there is not that the payoff is merely delayed. It is that most current deployments are structurally unlikely to pay off at all, for reasons that have far more to do with corporate execution than with model capability.

What "95% failure" actually means

The NANDA report's headline number has been widely, and somewhat sloppily, summarized as "95% of AI projects fail." That is not quite what it says. The pilots studied were mostly succeeding at the task they were built for; what failed to materialize was P&L-level impact after the pilot supposedly moved into production. The report's more diagnostic finding is about how companies got AI into the building.

More than 80% of firms had experimented with general-purpose tools like ChatGPT or Microsoft Copilot, and those experiences were generally judged positive. But only about 5% of custom, internally built enterprise AI pilots ever reached production. Tools purchased from a specialized vendor and integrated through a partnership succeeded at roughly triple the rate of tools built in-house — about two-thirds of the time, versus one-third for internal builds.

Deployment approach Reported success rate
Vendor-purchased, partnered tools About two-thirds
Internally built tools About one-third
Custom internal pilots reaching production ~5%

The report's own explanation is not a talent shortage, a compute shortage, or a regulatory obstacle. It is what the authors call a "learning gap": most deployed systems don't retain feedback, adapt to context, or improve with use the way a competent new hire would over a few months. They function more like a static tool bolted onto a workflow than a colleague absorbed into one.

Successful deployments, by contrast, shared three traits: they targeted a narrow, well-defined workflow rather than a broad ambition; ownership sat with the frontline managers who actually ran that workflow, not with a centralized AI lab; and the system was wired into the tools and data the team already used, rather than living as a separate destination employees had to remember to visit.

That last point connects directly to another of the report's more startling findings: a "shadow AI economy" thriving in parallel with the failing official one. NANDA found that roughly 90% of employees are already using personal AI tools for work on a regular basis, even though only about 40% of their employers have an official AI subscription.

Millions of workers are, in effect, running their own private pilot programs on ChatGPT, Claude, or Gemini accounts that were never sanctioned, funded, or measured by the enterprise — and getting more day-to-day value out of them than out of the AI systems their employer actually purchased. It is a damning contrast: the tools employees weren't given a budget for are outperforming, in perceived usefulness, the tools the company spent millions building.

The integration bill nobody budgeted for

Ask any CIO who has tried to move a generative-AI pilot from a demo to a department-wide rollout, and the same list of costs surfaces: cleaning and structuring the underlying data so a model can actually reason over it, rebuilding workflows and approval chains around the tool rather than around the old process, retraining staff, building monitoring and evaluation harnesses to catch when the system quietly degrades, and managing the change-management friction of a workforce that did not ask for the tool in the first place.

How many companies have actually scaled depends on the definition, which is itself part of the problem. Gartner, surveying 1,303 organisations with at least $50 million in revenue between January and April 2026, found 22% had scaled AI across multiple business units or adopted an AI-first approach. McKinsey's 2026 round, asking a looser question about enterprise-wide use, reports a considerably higher share. Both agree the binding constraint has moved from model access to whether a company's data is governed, connected and reusable enough for an AI system to draw on reliably. None of that shows up in a vendor's per-seat licensing quote, and all of it is expensive, slow, and unglamorous compared with the initial pilot.

The financial-industry data bears this out at scale. S&P Global Market Intelligence's 2025 research found that the share of companies expecting to abandon most of their AI initiatives had jumped to 42%, up from just 17% the year before — a signal that the failure rate was not a one-time hangover from early hype but was getting worse as more organizations tried to move past the pilot stage. No 2026 refresh of the series has appeared, so those remain the most recent numbers in it.

The same research, drawn from a survey of over 1,000 respondents across North America and Europe, found that companies are, on average, discarding 46% of AI proof-of-concepts before they ever reach implementation, citing cost, data privacy exposure, and security risk as the recurring reasons.

Metric Value
Firms expecting to abandon most AI initiatives (2024) 17%
Firms expecting to abandon most AI initiatives (2025) 42%
AI proof-of-concepts discarded before implementation (2025) 46%
Organisations scaled across multiple business units (Gartner, 2026) 22%
Respondents reporting any enterprise-level EBIT impact (McKinsey, 2026) 37%

Boston Consulting Group's AI Radar 2025 puts a number on the executive-level disappointment behind those abandonment rates: even as one in three companies planned to spend more than $25 million on AI in 2025, only about 25% of executives reported generating substantial value from their AI investments. BCG's diagnosis is that most companies are aiming too low and too shallow — chasing incremental productivity tweaks at the edges of the business rather than committing the 80%-plus share of AI investment that BCG's small cohort of leaders devote to redesigning core functions and building new offerings around the technology.

McKinsey's 2026 global survey, published on 25 August 2026 as "The state of AI in 2026: On the road to ROI," lands on the same split a year later. Only 37% of respondents attribute any enterprise-level EBIT impact to AI at all, a share that held firm while adoption and spending climbed. The cohort McKinsey calls AI high performers stayed flat too, at roughly 6% of respondents.

Gartner's 2026 numbers describe the same two-speed field from the other side. Among the organisations it rates as high performers, 81% of AI initiatives produced a positive return; among low performers, nearly a third of initiatives had a return nobody in the company could name. Meanwhile 85% of functional leaders planned to increase AI spending in 2026, and around 11% of organisations could not say what their function had spent on AI the year before.

The gap between the median company and this small cohort is, in miniature, the entire ROI reality gap: the technology clearly can produce outsized returns for organizations that rebuild their operating model around it, and clearly is not doing so for the much larger group that is bolting it onto business-as-usual.

Agentic AI is repeating the pattern faster

If generative-AI chatbots took two or three years to expose this gap, the next wave — autonomous "agentic" AI systems that plan and execute multi-step tasks with limited human oversight — is on pace to expose it faster. Gartner predicted in mid-2025 that more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls as the primary drivers.

Gartner's research also flags a market distortion behind the hype: of the thousands of vendors marketing themselves as agentic-AI providers, the firm estimates only around 130 offer genuine agentic capability, with the rest engaged in what Gartner calls "agent washing" — rebranding existing chatbots and robotic-process-automation tools with agentic language rather than agentic substance.

A January 2025 Gartner poll of over 3,400 webinar attendees found only 19% of organizations describing their agentic-AI investment as significant, with 42% calling it conservative and nearly a third still watching and waiting. Enterprises are, in short, about to relearn the generative-AI lesson — that a capable demo and a production-grade, financially accountable deployment are two very different projects — on a faster and more expensive cycle.

Where the ROI is real

None of this means enterprise AI return is a mirage everywhere. It means the return is concentrated in a narrower set of use cases than the initial hype implied, and those cases share a profile: a bounded task, a clear before-and-after metric, and a workflow the company was actually willing to redesign.

Coding assistants are the clearest individual-task success story with hard experimental backing. A randomized controlled trial run with researchers from Microsoft, GitHub, and MIT found that professional developers given access to GitHub Copilot completed a benchmark coding task 55.8% faster than a control group working without it. A separate pooled analysis of field experiments across three companies, covering nearly 4,867 developers, found Copilot use associated with roughly a 26% increase in throughput as measured by pull requests merged.

That is about as close to an unambiguous, reproducible productivity win as generative AI has produced anywhere in the enterprise — and it is precisely the kind of case the NANDA framework predicts should work: a narrow, well-defined task, embedded directly inside the tool developers already use, with an obvious and immediately visible unit of output.

Customer support offers a real but messier example, and Klarna's much-cited case captures both the promise and the overreach. In February 2024, the Swedish fintech said its OpenAI-powered assistant was handling the equivalent workload of 700 full-time agents, resolving inquiries in under two minutes versus eleven for a human, and projected to add $40 million to that year's profit.

Klarna's own reported figures show the cost of each customer-service interaction falling from $0.32 in the first quarter of 2023 to $0.19 by the first quarter of 2025 — a genuine, measurable, sustained unit-cost improvement, not a one-quarter press-release number.

But the full story is more instructive than the headline: by May 2025, CEO Sebastian Siemiatkowski was telling Bloomberg that the company had cut human staff too aggressively in the rush to automate, and Klarna began rehiring for a hybrid, human-inclusive support model even as it kept crediting AI with doing the bulk of routine volume.

The lesson enterprises should take from Klarna is not "AI support doesn't work" — the unit economics say it plainly does, for a large share of routine volume — but that ROI and headcount strategy are two separate decisions, and companies that collapse them into one press release tend to have to walk part of it back in public eighteen months later.

The gap between a task and a P&L line

The common thread connecting the winners to the losers is the same gap the NANDA researchers named directly: individual-task productivity studies are not the same measurement as organizational P&L impact, and enterprises have been treating them as interchangeable. A 55.8% speed-up on a coding benchmark, or a sharp drop in repeat customer inquiries, is real and measurable — but it only shows up on an income statement if the enterprise has redesigned staffing, pricing, or output targets around that new capacity.

Most companies haven't. They have bought or built a tool, observed that it works in isolated hands, and then left the surrounding organization — headcount plans, workflow ownership, incentive structures, data pipelines — almost entirely unchanged. The tool performs; the org chart doesn't move; the P&L doesn't notice.

That is the actual shape of the AI ROI reality gap: not a broken technology, and not simply a delayed macroeconomic payoff waiting on the slow diffusion that history's other general-purpose technologies required. It is a mismatch between where AI capability actually lives — narrow, well-integrated, frontline-owned deployments — and where most enterprise AI budgets were actually spent, which is on broad, centrally run, loosely integrated initiatives optimized for an impressive pilot demo rather than for a number a CFO can defend.

McKinsey titled its 2026 round "On the road to ROI," and the road is doing a lot of work in that phrase: more companies have scaled, spending is still rising, and the share reporting any enterprise-level EBIT impact is 37%, essentially where it stood a year earlier. Scaling moved. The earnings line did not.

Until more companies close that specific, addressable gap — narrowing scope, buying rather than over-building, wiring the tool into existing workflow ownership, and measuring P&L rather than pilot adoption — the productivity revolution enterprises were promised will keep showing up in the 5% and staying stubbornly absent from the other 95%.

Explore

More articles