It's tempting to read a stat like "only 25% of AI initiatives deliver expected ROI" as evidence that the technology is overhyped. It's the wrong conclusion. The same CEOs reporting disappointing returns are, in most cases, running pilots built on the same underlying models that are quietly delivering real value elsewhere in their own organizations. The gap isn't in what the models can do — it's in how the projects around them get scoped, deployed, and supported after launch.

25%
of AI initiatives deliver their expected ROI, per IBM's CEO study
16%
of AI projects ever reach enterprise-wide scale
4x
more successful orgs invest in data quality & governance, per Gartner

Where the 75% actually go wrong

It rarely fails at the model. It fails at the handoff points around it: a pilot proves the concept in a controlled demo, then stalls trying to connect to real production data; or it launches successfully in one department and never gets adapted for the next one because nobody planned for that step. Gartner's April 2026 research puts a number on the difference — organizations with successful AI outcomes invest up to four times more, as a share of revenue, in data quality, governance, and change management than the ones that stall. The technology spend looks similar across both groups. The supporting investment doesn't.

Why "enterprise-wide scale" is the real bar, not "it works"

A pilot that works for one team but can't be extended to a second is, from a return-on-investment standpoint, indistinguishable from a pilot that never worked at all — it cost the same to build and it's generating value for a fraction of the organization it was pitched to justify. That's the gap between the 25% seeing ROI and the 16% reaching real scale: most of the value-generating minority still hasn't cleared the harder bar of making that value available company-wide.

A successful pilot proves the model works. Enterprise-wide scale proves the deployment does — and that's the harder problem almost nobody budgets enough time for.

What the 25% are actually doing

Three patterns show up consistently in the initiatives that clear the ROI bar: they start with a workflow narrow enough to measure precisely, rather than an open-ended "AI for X" mandate; they build the data plumbing and access governance once, upfront, instead of treating it as a phase-two problem after the demo impresses leadership; and they plan the second and third deployment before the first one ships, so scaling isn't a separate project competing for a new budget approval. That last point connects directly to the agent-sprawl problem — the fastest path past a stalled pilot is often not a better model, it's not having to rebuild the integration work from scratch for the next team.

The takeaway for anyone scoping their next AI project

Measure the project against enterprise-wide scale from day one, not against "did the pilot work." If a proposal can't describe how a second department adopts the same deployment without a parallel six-month integration effort, that's the ROI risk hiding in plain sight — long before the first invoice for compute ever shows up.