WHY AI PILOTS STALL
95% of AI pilots stall. 74% of leaders report positive ROI. Both are true.
MIT says 95% of companies get zero return from generative AI. Wharton says three out of four enterprise leaders already see positive ROI. The two studies measure different points on the same path, and the distance between those points is where a mid-market company either wins or wastes a year.
This page explains what each study actually measured, why they agree more than the headlines suggest, and the three things the companies at the far end of the path have in common.
THE 95%
What MIT measured.
The number comes from The GenAI Divide: State of AI in Business 2025, published by MIT’s Project NANDA in July 2025. The team reviewed more than 300 publicly disclosed AI initiatives, interviewed people at 52 organizations, and surveyed 153 senior leaders between January and June 2025.
The headline sentence is blunt: “95% of organizations are getting zero return.” The same paragraph continues: “Just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact.”
MIT set a high bar. A pilot counted as successful only if it was deployed beyond the pilot phase with measurable KPIs, and ROI was measured six months after the pilot ended. In the authors’ words, success meant “a marked and sustained productivity and/or P&L impact.” Anecdotes didn’t clear it.
The authors also say, in a research-limitations note beside the chart, that the figures are “directionally accurate based on individual interviews rather than official company reporting.” Take the 95% as a serious finding from a serious team, and treat the reasons behind it as more useful than the decimal.
Four findings from the report that matter to a company your size
-
01
Mid-market companies move faster.
Top-performing mid-market companies (MIT draws the enterprise line at $100 million in revenue) reported average timelines of about 90 days from pilot to full implementation. Enterprises took nine months or longer.
-
02
Your people are already using AI.
Only 40% of companies had bought an official AI subscription. Workers at more than 90% of the companies surveyed were using personal AI tools for work anyway. MIT calls it the shadow AI economy.
-
03
Buying beat building.
Pilots run with an outside partner reached deployment about 67% of the time. Internal builds reached it about 33% of the time. The authors flag that this is correlation rather than proof of cause, and that the gap was consistent across interviews.
-
04
Budgets go where the demo is.
Asked to allocate a hypothetical AI budget, executives put roughly half of it into sales and marketing, while the clearest returns showed up in back-office work: document processing, finance, operations.
The report’s conclusion on organizational design: “Success depends less on resources and more on decentralizing authority with clear ownership.”
THE 74%
What Wharton measured.
Accountable Acceleration: Gen AI Fast-Tracks Into the Enterprise came out in October 2025 from Wharton Human-AI Research and GBK Collective. It’s the third year of the same survey, fielded June 26 to July 11, 2025, to 801 senior decision-makers at US companies with at least 1,000 employees and more than $50 million in revenue.
The numbers: 74% of those leaders report positive ROI on generative AI, and four in five expect positive returns within two to three years. Some 82% use the tools at least weekly and 46% daily. And 72% say their company formally tracks ROI, on metrics like profitability, throughput and workforce productivity.
The tools are being used for ordinary knowledge work: data analysis (73% of respondents), summarizing documents and meetings (70%), and editing and writing documents (68%). Wharton notes the most-used cases are also the highest rated for performance. These leaders got faster at work they already do.
Two details for a mid-market reader. First, the companies under $2 billion, and the smallest tier at $50 million to $250 million in particular, reported quicker ROI than the largest enterprises; 34% of the $2 billion-plus companies reported neutral or too-early outcomes. Second, the report’s own summary of what holds companies back: “people and processes are the new constraint.”
Chief AI Officer roles now exist at 60% of the enterprises surveyed, 64% have adopted data security policies, and 61% are running training programs. Wharton’s read is that the companies pulling ahead are the ones aligning talent, training and trust with the spend. Put plainly: someone in charge, and the rules written down.
THE PATH
Why both numbers are true.
Every company’s AI work sits somewhere on the same five-step path. The two studies measure different steps on it.
-
1
Experimentation
People try ChatGPT. Nothing is owned, measured, or connected to a process.
MIT’s 95% sit here -
2
Individual assistance
Faster emails, summaries, first drafts. Real, and invisible in the P&L.
MIT’s 95% sit here -
3
Workflow adoption
AI sits inside one recurring process, with a person accountable for it.
Wharton’s 74% report from here -
4
Measured improvement
Hours, cycle time or error rate move against a written baseline.
Wharton’s 74% report from here -
5
Sustained P&L impact
The gain holds for six months and shows up in the numbers.
MIT’s 5% cleared this
Wharton asked leaders whether they see positive returns, and most of their answers describe steps two through four: productivity, throughput, incremental profit. MIT asked whether a pilot became a sustained system, built into the workflow, with a measurable financial result six months later, which is step five. The 74% and the 5% are measuring different steps.
For a company between $10 million and $200 million, the path is shorter than it is for the enterprises in either sample. Fewer approvals, fewer systems, one owner who can decide on a Monday. That’s the 90 days in MIT’s data. The disadvantage is that nobody hands you a Chief AI Officer or a governance committee. You have to decide to own it.
A note on scale: both studies skew larger than most companies reading this. MIT’s sample runs from small businesses to enterprises; Wharton’s smallest tier is $50 million in revenue with a thousand employees. The mechanisms transfer. The exact percentages may not.
WHERE PILOTS STALL
What stalling sounds like from the inside.
I spend most weeks in rooms with CEOs of construction, manufacturing, professional services and PE-backed companies. Almost none of them describe a failed AI project. They describe not knowing who is using what. Here is how it sounds, in their words, with names removed.
“I couldn’t tell you how many people are signed up for ChatGPT. I just don’t know. We don’t have an AI policy. People are still wandering around.”
CEO, construction firm, New York
“Somebody’s using Gemini, somebody’s using NotebookLM, I’m using Claude and ChatGPT. It’s just craziness. We’re all paying a different fee.”
VP of people and operations, 115-person fleet services company
“Maybe there’s three or four people using it on the side. Now, I’m not naive. There’s 100 people in the company and I think 100 people are using it.”
Executive, commercial real estate
“They’re siloed, because we don’t have a business AI with a shared conscience.”
COO, electrical and mechanical contractor
“Do we need an AI designer, an AI implementer, an AI specialist? We don’t know.”
The same VP of operations
MIT heard the same thing. A mid-market manufacturing COO told the researchers: “The hype on LinkedIn says everything has changed, but in our operations, nothing fundamental has shifted. We’re processing some contracts faster, but that’s all that has changed.”
The five ways a pilot stalls
-
01
Nobody owns it.
The pilot has a sponsor and a champion but no one whose job it is on Monday. When the champion changes roles, the pilot goes with them. MIT’s top-ranked barrier was “unwillingness to adopt new tools,” and unwillingness is what an unowned tool gets.
-
02
There’s no baseline.
Nobody wrote down how many hours the process took before, so nobody can prove it takes fewer now. The pilot ends in a debate instead of a decision.
-
03
The AI sits beside the work instead of inside it.
It drafts, someone copies the draft into the real system, and within a month people skip the middle step. Nobody decided to stop using it. They just did.
-
04
The tool has no company context.
A general chatbot doesn’t know your approval rules, your naming conventions or your exceptions. MIT calls this the learning gap, and it is the report’s central explanation for the 95%. One of MIT’s interviewees, on tools that don’t learn: “It’s useful the first week, but then it just repeats the same mistakes. Why would I use that?”
-
05
The company built when it should have bought.
Internal builds reached deployment half as often as partnered ones. MIT is careful to call that a correlation. My own read, from watching it happen: the team underestimates integration, evaluation and upkeep, and the build becomes a second job nobody was hired for.
I’ve watched a sixth one up close. We spent a full day with a client’s team. They said, “We’ve got it from here.” We checked in 30 days later and everybody had gone back to doing things the old way. Training without an owner and a measured workflow is a nice day out.
THE FAR END
What the companies that get results have in common.
Both reports land on the same three things. Wharton’s thesis is that returns follow leadership, governance and skills. MIT’s 5% had clear ownership, workflow integration and a measurement bar. In a company your size: an owner, a policy, one measured workflow.
1. An owner
One person whose name is on AI at your company. They approve tools, answer questions, decide exceptions and report the number. Someone with a P&L instinct and the authority to say yes and no. Usually an existing leader, rarely IT.
Monday question: who approved the AI tools we’re using right now? If the answer is a shrug, you’ve found the first job.
2. A policy
One page that says what can go into AI tools, what can’t, which accounts to use, and who to call when something goes wrong. MIT found workers at over 90% of companies using personal AI accounts for work while only 40% of companies had bought a subscription. That gap is company data sitting on personal logins the company can’t see or shut off. The policy closes it in an hour.
The one-hour version: leadership team in a room, a draft, an argument about section four (what stays out of the tools), signed before lunch. One Vistage member did it in six minutes with the template and ChatGPT open side by side. The template is further down this page.
3. One measured workflow
One recurring process, a baseline taken before you start, AI built into the actual steps, and a number checked at the end. The four-filter test for picking it: it recurs weekly, it runs on rules, it’s measured in hours, and it’s nobody’s favorite job. Anything passing all four is your pilot.
Weekly recurrence means the hours add up. Rules mean today’s tools can handle it. Hours give you something to measure. And nobody defends a job nobody likes.
THE 90 DAYS
What a measured quarter looks like.
MIT’s top mid-market performers went from pilot to full implementation in about 90 days. Here is the shape of a quarter that earns that number.
-
Weeks 1 to 2
Owner, policy, baseline.
Name the owner. Run the one-hour policy meeting. Move everyone onto company accounts. Pick the workflow with the four filters and write down the baseline: hours per week, who does it, error rate if there is one.
-
Weeks 3 to 6
Build it into the work.
Put the AI inside the actual steps, with the handoffs and the exceptions. Buy or partner if the build is more than a few days. Train the two or three people who run the process on the new version of their job, not on prompting in general.
-
Weeks 7 to 12
Measure and decide.
Compare against the baseline every two weeks. In week twelve the owner makes one of three calls: scale it, redesign it, or stop it. The decision rule gets written in week one so nobody negotiates it later.
-
Day 91
The second workflow.
The first one paid for the second. Now the math compounds.
ONE WORKFLOW
What it looks like with the numbers in.
A construction firm we worked with had one person spending 20 hours a week processing certificates of insurance. Every subcontractor, every project, every week. It passed all four filters.
Baseline: 20 hours a week. We built the workflow and deployed it in about 24 hours. Result: 20 hours a week became 2. Call it 1,000 hours a year returned to the business, roughly $600,000 in enterprise value by our estimate, from one process nobody enjoyed doing.
Run your own version. Take your loaded hourly rate and multiply it by the hours you personally spend each week on work AI should already be doing. The rewriting, the summarizing, the first drafts, the report you rebuild every Monday. Four hours a week at $500 an hour is $8,000 a month, for one person. Do it for the two people who report to you and the business case is two lines of arithmetic.
THE SCORECARD
Which side of the divide are you on?
Six yes-or-no questions. Tick the ones you can answer yes to today.
-
0 to 2 yes
You’re in MIT’s 95%, and you have plenty of company. The good news is that the first two items take an afternoon.
-
3 to 4 yes
You’re where most of Wharton’s 74% are: real, self-reported productivity gains, not yet measured the way MIT measured. One measured workflow moves you.
-
5 to 6 yes
You’re on the far side of the divide. Pick the second workflow.
QUESTIONS
Asked about the two studies, and about starting.
Why do AI pilots fail?
Most stall for organizational reasons rather than technical ones. MIT NANDA’s 2025 research ranked unwillingness to adopt new tools as the top barrier, with model output quality and poor user experience next. The report’s explanation for all three is that most tools don’t learn a company’s context or fit its workflows. In practice a pilot fails when nobody owns it, there is no baseline to measure against, and the AI sits beside the work instead of inside it.
What percentage of AI pilots fail?
MIT NANDA’s July 2025 report found 95% of organizations saw no measurable return from generative AI, and only about 5% of integrated pilots produced sustained value. The bar was strict: deployment beyond the pilot phase with measurable KPIs, assessed six months later. Softer measures of success show much higher rates; Wharton’s October 2025 survey found 74% of enterprise leaders reporting positive ROI.
Do the MIT and Wharton AI studies contradict each other?
No. Wharton’s survey of 801 enterprise leaders found 74% reporting positive ROI, mostly from productivity, throughput and incremental profit. MIT measured whether pilots became sustained systems, built into the workflow, with P&L impact six months on. Wharton describes the middle of the value path and MIT describes the far end. Both conclude that people, ownership and process decide the outcome more than model quality does.
Who should own AI at a mid-size company?
One named leader with a P&L instinct and the authority to approve tools, set the use policy and report the results. Wharton found Chief AI Officer roles at 60% of the large enterprises it surveyed. A mid-market company rarely needs a new hire, but it does need one accountable owner. MIT’s finding was that success depends less on resources and more on decentralizing authority with clear ownership.
How do you measure an AI pilot?
Write down the baseline before you start: hours per week, cycle time, or error rate for the specific workflow. Build the AI into the actual steps of the work, then compare against the baseline at fixed intervals with a written decision rule: scale, redesign or stop. Track a lagging number the CFO cares about, such as hours returned or cost per unit, rather than usage alone.
How long should an AI pilot run?
About 90 days for a mid-market company. MIT found top mid-market performers moved from pilot to full implementation in roughly 90 days, against nine months or longer for large enterprises. A workable shape is two weeks to name an owner, adopt a policy and take a baseline, four weeks to build the AI into the workflow, and six weeks to measure and decide.
Do we need an AI use policy before starting a pilot?
Yes, and it takes about an hour. MIT found workers at more than 90% of companies using personal AI tools for work while only 40% of companies had bought an official subscription, which means company data sitting on personal logins. A one-page policy that says what can go into AI tools, what stays out, which accounts to use and who owns the call closes that gap. The template is below.
Why do mid-size companies deploy AI faster than large ones?
Fewer approvals, fewer systems, and an owner who can decide the same week. MIT reported roughly 90-day pilot-to-deployment timelines for top mid-market performers against nine months or more for enterprises. Wharton found companies under $2 billion, the $50 million to $250 million tier in particular, reporting quicker ROI than firms above $2 billion. The trade-off is that nobody hands a mid-market CEO a governance structure. Someone has to decide to own it.
Sources
- MIT NANDA, The GenAI Divide: State of AI in Business 2025 (Challapally, Pease, Raskar, Chari), July 2025. Figures cited: executive summary, sections 3.1 to 3.4, 4.1, 5.1, 6.1, 6.3, and the methodology appendix.
- Wharton Human-AI Research and GBK Collective, Accountable Acceleration: Gen AI Fast-Tracks Into the Enterprise, October 2025. Figures cited: methodology (PDF p. 6), executive summary (PDF pp. 7 to 14), ROI findings (PDF pp. 38 and 47). Summary at Knowledge at Wharton.
- Executive quotations under “What stalling sounds like” are verbatim from recorded advisory calls in 2026, used with identifying details removed. The construction case is a Chief AI Officer client engagement.
THE AI USE POLICY
Start with the policy. It takes an hour.
Every room I speak in leaves with a draft AI use policy. This is that template: what your people can put into AI tools, what stays out of them, which accounts to use, and who owns the call. One Vistage member turned it into his company’s policy in six minutes.
On its way.
Check your inbox in the next few minutes. If it’s not there, look in spam once, then email me.
NEXT
If you want the hour run for you.
The Private Practice is AI counsel for one executive at a time. Two working sessions a month on your actual calendar, inbox and decisions, building the owner, the policy and the first measured workflow with you. Ten seats, application only. Or bring the argument on this page to your leadership team as a working session.
The Private Practice Speaking & Workshops
Or say hello: doc@chrisdaigle.com / (504) 606-7519.