Why most AI pilots stall, and how to avoid it
Most AI pilots stall because the data behind them was never organized for the job, and because no one defined what “working” would look like before starting. The research on this is now blunt: MIT’s Project NANDA reported that “95% of organizations are getting zero return” from generative AI, and Gartner has predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025. Neither report blames the underlying models.
This is for owners, managing partners and COOs who’ve either tried an AI pilot that fizzled, or want to avoid running one that will.
What the research actually found
MIT’s Project NANDA report, based on 52 executive interviews, a survey of 153 leaders, and analysis of roughly 300 public AI deployments, reported: “Despite $30–40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return.” Its explanation for why: “most AI tools don’t learn and don’t integrate well into workflows.” (MIT NANDA, “The GenAI Divide: State of AI in Business 2025,” July 2025)
Separately, Gartner predicted at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, unclear business value, inadequate risk controls and escalating costs (Gartner press release, July 29, 2024).
The pattern behind both numbers
Read together, three things show up repeatedly:
- The data wasn’t ready. A pilot pointed at scattered, uncleaned information produces answers no one trusts, so it never leaves the pilot stage.
- No one defined success first. Without a specific outcome agreed in advance, a pilot can run for months without anyone being able to say whether it worked.
- It stayed generic. A tool that doesn’t know your firm’s procedures, clients and standards produces work that looks plausible but isn’t usable without heavy rework — which erases the time savings that justified the pilot.
MIT’s report also found that “external partnerships with learning-capable, customized tools reached deployment ~67% of the time, compared to ~33% for internally built tools.” The authors note these percentages come from their interview sample of 52 organizations and may not represent the broader market.
What a firm can do differently
- Fix the data first. A pipeline that connects and cleans your firm’s and clients’ information is the difference between a demo and something your team will actually rely on.
- Name the outcome before you start. Decide what you’re measuring (hours saved on a specific task, turnaround time, error rate) and how you’ll know in a defined window.
- Scope it narrow. One process, not a department-wide rollout, so a stall is cheap and a success is easy to point to.
- Keep a person in the loop. A result nobody has to trust blindly is a result people will actually use.
Where Precision AI OS fits
Precision AI OS is built around the failure mode the research describes: we build the data pipeline first, so agents work from your firm’s real information instead of guessing, and every output goes to a person on your team to approve before it reaches a client. See how we protect your data or book a discovery call.
Frequently asked questions
What percentage of AI pilots actually fail?
MIT's Project NANDA reported that 95% of organizations in its research were getting zero return from generative AI (July 2025). Gartner separately predicted at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025.
Is the failure rate about the AI models themselves?
Usually not. MIT's report points to tools that "don't learn and don't integrate well into workflows" rather than model quality, which is why scoping, data and workflow fit matter more than which model you pick.
What's the single biggest fix?
Get the underlying data organized before you pilot anything. A pilot run on scattered, unverified information produces answers no one trusts, which is a common reason pilots quietly stop.
Does bringing in outside help actually change the odds?
In MIT's interview sample, external partnerships reached deployment about 67% of the time versus about 33% for internal builds. It's one study of 52 organizations, not a guarantee.
Related articles
Put AI to work in your firm
Start with the business outcome, the data it depends on, and the people who will approve the work.