Skip to content

Strategy

Why most GenAI pilots return nothing measurable

MIT NANDA found ~95% of GenAI pilots show no P&L return. What that study actually measured, what it cannot show, and the six things a kill criterion contains.

Published
Reading
8 min
Based on
Published research — MIT Project NANDA 2025, Gartner, Eurostat — plus our own portfolio gating method

The number has been on every slide since August 2025: roughly 95% of GenAI pilots produce no measurable return. It is quoted as a verdict on the technology. It is not. It is a verdict on how the pilots were set up to be judged — and the correction is one paragraph of writing that almost nobody produces before the build starts.

The study says less than the headline, and something more useful

The source is MIT Project NANDA’s The GenAI Divide: State of AI in Business 2025: 52 structured interviews, a survey of 153 leaders, and a review of 300-plus public deployments. That is a purposive sample, not a random one, and “no measurable P&L impact” is largely self-reported by organisations that, in most cases, never established a baseline to measure against.

Read strictly, the finding is that around 95% of organisations could not produce a number. That is a different claim from 95% produced no value, and the study cannot separate the two. It is also cross-sectional: it photographs a population whose pilots are mostly less than a year old, at an adoption stage where the honest answer to “what did it return?” is often “too early”.

Hold that distinction and the more interesting results are elsewhere in the same data. The enterprise funnel for task-specific GenAI tools runs about 60% investigated, about 20% piloted, about 5% successfully implemented. But generic chatbot tools converted from pilot to implementation at roughly 83%. The collapse is concentrated in custom and vendor enterprise systems, not in language models as such. Whatever is failing is failing at the point where a model meets a system of record and a process owner.

Two further numbers survive the sampling objections because they are comparative — both arms were measured the same way. Pilots delivered through external partnerships reached deployment about 67% of the time, against about 33% for internal builds, with employee usage rates nearly double for the externally built tools. And roughly half of GenAI budget went to sales and marketing, where attribution is easy, while the largest documented savings in the sample were back-office: $2–10M annually from BPO elimination, a 30% reduction in external agency spend, about $1M annually from outsourced risk-management checks.

Budget followed measurability rather than payback. That, not the 95%, is the finding worth acting on.

Pilots rarely fail. They stop being mentioned.

Gartner’s June 2025 forecast is that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. Cancellation is what happens when a budget cycle ends. It arrives late, it arrives at the sponsor’s convenience, and it teaches the organisation nothing, because by then nobody can reconstruct what the pilot was supposed to beat.

A written kill criterion differs in one respect that carries all the weight: it has a date. It converts “is this working?” from a judgement call, which the person who commissioned the work will always answer optimistically, into a lookup.

A kill criterion is six specific things, and most written ones have two

The ones we see in practice contain a metric and a hope. A criterion that can actually fire contains:

  1. A metric at the unit of work — cost per successful task, per resolved ticket, per processed invoice, per drafted document. Cost per token is an input, not a KPI.
  2. A baseline measured before the build, with the measurement method frozen in writing. Process-mining extract or a two-week time study; never a manager’s estimate, and never re-measured by a different instrument afterwards.
  3. A threshold with a direction and a tolerance — the number at which the pilot proceeds, iterates, or stops.
  4. A measurement window and a calendar date, fixed before the first result is seen.
  5. A population, including a holdout that does not get the tool. Without a counterfactual you cannot tell the model from seasonality.
  6. A named signatory. Not a committee.

Where defaults exist, use them. Adoption below 30% of licensed seats is a failed rollout regardless of model quality; healthy sustained adoption runs above 60%. For a RAG system, retrieval recall@5 below 80% means generation-quality work is wasted effort, because most of what gets reported as hallucination is retrieval failure. Groundedness gates on a golden set typically sit between 90% and 97% depending on risk tier. Those are release gates. The kill criterion sits above them, on the business metric — a system can pass every release gate and still be worth switching off.

We got this wrong ourselves in the way that is easiest to get wrong. Our first written criterion was defined on accuracy against a golden dataset: precise, falsifiable, and ours. It measured a thing we controlled, and it would have passed comfortably while the tool sat unopened. A criterion on a metric the supplier owns cannot fire.

The signature is the mechanism, not the ceremony

The signatory is the process owner whose queue, headcount or cycle time actually changes — not the executive sponsor who approved the budget, not IT, and not the supplier.

The sponsor’s incentive is continuation; a sponsor who kills a project is reporting their own misallocation upward. The process owner’s incentive runs the other way, and more practically, only the process owner can execute the revert, because the revert is a change to how their team works on a Monday morning.

If nobody in the organisation will put their name on the threshold, that refusal is itself the output of the gate. It means the value case was never believed at the level where the work happens.

The day it triggers, you are not starting from zero

“Off” has to be tested before it is needed. Measure mean time to revert to the prior manual process while the pilot is still healthy — a rollback plan that has never been executed is a paragraph, not a capability.

Keep the logs. Article 26(6) of the AI Act requires deployers to retain automatically generated logs for at least six months, and a decommission that tears down the environment on the same afternoon destroys evidence you are obliged to hold. Non-enforcement is not non-liability: Bulgaria had no designated market surveillance authority and no national sanctions regime as of January 2026, and the obligation attaches anyway.

Then write down which of the six conditions failed, and take the inventory of what survives. The golden dataset labelled by your own subject-matter experts survives. The exception taxonomy — several hundred real items coded by reason — survives, and it is usually the most valuable artefact in the whole pilot, because in document-heavy back office a large share of exceptions turn out to be upstream data problems (wrong PO reference, missing master data) that no model was ever going to fix. The integration and permissions work survives. The measured baseline survives, and it is what makes the next attempt cheaper.

A portfolio with a 0% kill rate has decorative gates. A healthy one retires 30–50% of candidates before build. Ask any advisor — us included — for the last three use cases they recommended against. An advisor who cannot name one is not gating anything.

Where this stops applying

A kill criterion is an instrument for efficiency cases with a countable unit of work. It fits badly on capability bets, on compliance-driven work with no discretionary alternative, and on anything where the counterfactual genuinely cannot be isolated. Forcing a false threshold onto those is worse than admitting they are being funded on judgement.

Windows have to match your own tempo, too. NANDA’s top mid-market performers went from pilot to full implementation in around 90 days; large enterprises took nine months or more. A 60-day criterion inside a nine-month organisation does not enforce discipline — it kills things before the integration debt clears, and the correct response is to fix the cycle time, not the threshold.

And none of this substitutes for the unglamorous part. Eurostat’s 2025 figures put Bulgarian enterprise AI use at 8.55% against a 19.95% EU average. That gap will not close because someone wrote a better gate. It closes when a first use case is measured against something, and the measurement was agreed before anyone saw the result.

Abstract warm light on a dark field

Is this the problem you are living with?

If this article describes your situation, the fastest next step is a call with the person who wrote it.

30 minutes, no obligation, and you keep whatever we work out.