This page does two things. It explains why pilots stall, using the honest version of the statistic everyone quotes, and it hands you an ungated business case template mapped to the UK government’s own standard. Sources were read on 21 September 2026 and are linked in Sources.
At a glance
- 5 per cent of task-specific AI tools reached production in the most widely cited study. 20 per cent got as far as a pilot.
- The famous 95 per cent figure rests on 52 interviews and 153 survey responses, and its authors state its limits. We repeat them here.
- Four failures account for most stalls, and all four are visible before anything is built.
- A business case template is published in full on this page. No form, no email, no gate.
- One system we built runs daily and has priced a real issued job in public: 63 measured lines, £98,214.72.
What does the 95 per cent figure actually measure?
It comes from MIT’s Project NANDA report, The GenAI Divide: State of AI in Business 2025, published in July 2025. The headline is that despite very large enterprise investment, 95 per cent of organisations are getting zero return, and that just 5 per cent of integrated AI pilots are extracting significant value.
The more useful number is inside the report rather than on its cover. For task-specific AI tools, the funnel runs at roughly 60 per cent of organisations evaluating, 20 per cent reaching pilot stage, and 5 per cent reaching production. The report attributes the drop to brittle workflows, a lack of contextual learning, and misalignment with day-to-day operations.
Now the part almost nobody repeats. The report states its own limitations in writing: the figures are “directionally accurate based on individual interviews rather than official company reporting”, sample sizes vary by category, and success definitions differ between organisations. The evidence base is 52 structured interviews, survey responses from 153 senior leaders collected at four industry conferences, and a systematic review of more than 300 publicly disclosed initiatives, over a research period of January to June 2025. Success is defined as deployment beyond pilot with measurable indicators.
That is a substantial piece of work and it is not a measured failure rate for UK businesses of £1m to £50m turnover. Nobody has measured that. Use 95 per cent as evidence that a pattern is real and widespread. Do not use it as a probability for your project, and be sceptical of any supplier who does, including us.
Failure 1: did anyone own it?
Name the person who will be asked about this system in six months. If the answer is a department, a steering group or “IT”, the pilot has already failed and the failure has simply not surfaced yet.
The pattern is consistent. A pilot is sponsored by someone senior, run by someone technical, and used by someone in operations who was told about it late. When it produces an awkward result, nobody has both the authority and the motivation to change the surrounding work, so the awkward result stands and the team quietly reverts. The system keeps running and nobody uses it.
The fix costs nothing and takes ten minutes: one named owner, who owns the process rather than the technology, agreed in writing before the pilot starts.
Failure 2: was the data fine only in the demo?
Demonstrations run on prepared data. Production runs on whatever arrived this morning, including the supplier who changed their invoice layout, the record with three addresses in it, and the field somebody has been using for notes since 2019.
The symptom is a pilot that scored well and a live system whose accuracy fell off a cliff in week two. It is almost never a model problem. It is a distribution problem: the sample was cleaner than reality, usually because the person preparing it unconsciously chose the tidy examples.
The check is to insist that the pilot runs on an unselected slice of live data, picked by someone who is not building the thing, including the last three cases that went wrong.
Failure 3: was the process ever redesigned?
A pilot that bolts a tool onto an unchanged process produces a team doing the old process plus the new tool. The work does not get smaller. It gets an extra step, and the extra step is the thing people drop first when they are busy.
This is the failure that makes adoption look like resistance. Staff are not refusing the technology. They are declining to do two versions of the same job. If the pilot does not remove a step, a form, a spreadsheet or a hand-off, nothing has been automated and the return exists only on the slide.
The check is to write down, before the build, exactly which existing step disappears on the day this goes live. If no step disappears, stop and go back to the process.
Failure 4: was the payback horizon ever agreed?
A pilot without an agreed decision date does not fail. It fades. It stays in “evaluation” for three quarters, the sponsor changes role, and the budget is absorbed with nobody ever writing the word “no”.
Agreeing the horizon in advance means three things are fixed before the start: the date on which a decision will be taken, the number that will be looked at on that date, and who takes the decision. If the number is not agreed in advance, the review meeting becomes an argument about which number to use, and that argument is always won by whoever is most invested.
Be realistic about the horizon itself. Published UK automation work commonly describes payback across several months rather than weeks, and an honest pilot plan says which months.
Which four checks would have caught it?
All four failures are visible before anything is built, which is why we run these before a build is quoted rather than after it disappoints.
- Owner check. One named person, owning the process, agreed in writing. If nobody can be named, there is no project.
- Data check. Run the idea against an unselected slice of live data, chosen by someone outside the build, including recent failures.
- Step-removal check. Name the existing step that disappears on go-live. If none does, the return is theoretical.
- Decision check. Fix the date, the metric and the decision-maker before the work starts.
None of these requires a supplier. They require an afternoon, and they are the cheapest thing in this entire subject.
What do UK rules add to this?
Every page we read on this question was written for a United States audience and none of them mentions UK data protection, the ICO, or a pound sign. That is a real gap for a British reader, because the UK rules change the sequence of work rather than just adding paperwork at the end.
If the pilot processes personal data, UK GDPR applies while it is a pilot, not from the day it goes into production. The ICO publishes guidance on AI and data protection covering how the data protection principles apply to AI systems, and separate guidance on explaining decisions made with AI. The practical consequences for a mid-sized business are usually three: you need a lawful basis before the data is used rather than after the results look good; you need to be able to explain an individual decision in terms a person affected by it would accept; and you need to know where the data physically goes if a third-party model is involved.
Doing this at the end is how a successful pilot fails to reach production. The system works, legal asks a question nobody prepared for, and the project waits.
The public sector has a further layer worth knowing about even if you are not in it: the government publishes an AI Playbook setting out how it expects AI to be used safely and effectively, and business cases follow HM Treasury’s Five Case Model. Both are free to read and both are reasonable templates for a private business that wants a defensible process.
What does it cost to check properly, and what does it cost not to?
The checks above cost an afternoon of internal time. A structured version, run across the whole operation with the arithmetic written down, is our fourteen-day AI Constraint Audit at a fixed £2,000, with three sixty-minute working sessions and four documents you own. It is described in full, including what it excludes, on the AI Constraint Audit page. The fee is not credited against a build, which is the mechanism that allows the answer to be no.
The cost of not checking is not the pilot budget. It is the second year: the team has concluded AI does not work here, the next proposal is harder to fund, and the process that was actually costing the money is still costing it. That is the expensive part, and it does not appear on any invoice.
When should you kill a pilot instead of rescuing it?
Kill it when the data check failed and the fix is a data project nobody has budgeted. You are no longer testing the idea, you are funding a different project through the back door.
Kill it when no step has been removed after two attempts. Two attempts is generous. If the work has not got smaller by then, the tool is not the problem.
Kill it when the owner has changed twice. Ownership churn is the clearest predictor in this list, because each new owner restarts the argument about what success means.
And kill it when the honest answer to “what would we do differently at scale” is “spend more”. Scaling something that does not work at small size is how a stalled pilot becomes an expensive stalled programme.
A business case template you can copy
Here is the template, in full, with no form and no email gate. It maps to the HM Treasury Five Case Model, which is the standard the UK government publishes for its own business cases: Strategic, Economic, Commercial, Financial and Management. Using it does not make you a public body. It makes your paper recognisable to a finance director who has seen one before.
# AI business case: [process name]
## 1. Strategic case: why this, and why now
- The process, in one sentence, as it runs today.
- What it is holding up: capacity, cash, risk, or a customer promise.
- What happens over the next twelve months if nothing changes.
- Why this process and not the other candidates. (Name at least two you rejected.)
## 2. Economic case: the options, including doing nothing
| Option | First-year cost | Annual value released | Assumption that matters most |
|---|---|---|---|
| Do nothing | £0 | £0 | [the cost that continues] |
| Change the process only | £ | £ | |
| Buy an off-the-shelf tool | £ | £ | |
| Commission a build | £ | £ | |
- Loaded hourly cost used: £[x] per hour. (Salary plus employer costs, divided by working hours.)
- Measured, not estimated: [which numbers were counted, and by whom]
- Estimated: [which numbers are judgement, and how wrong they could be]
## 3. Commercial case: how it would be bought
- Supplier, and how they were chosen.
- Pricing model: fixed price, day rate, or licence. Who carries the overrun.
- What is explicitly not included in the price.
- Ownership of the code and the data at the end.
- Exit: what we would have to do to stop using this.
## 4. Financial case: the money, over three years
| | Year 1 | Year 2 | Year 3 |
|---|---|---|---|
| Build or licence | £ | £ | £ |
| Run rate (hosting, usage, support) | £ | £ | £ |
| Internal time | £ | £ | £ |
| Total cost | £ | £ | £ |
| Value released | £ | £ | £ |
- Budget it comes from, and whether it is capital or operating spend.
- Payback month, stated as a date.
- The single number that, if wrong by 30 per cent, changes the decision.
## 5. Management case: who does it and how we will know
- Named owner (a person, not a department).
- Decision date, and the metric that will be read on that date.
- What "working" looks like, in one sentence containing a number.
- The existing step that disappears on go-live.
- Data protection: lawful basis, whether a DPIA is required, where the data goes.
- Kill criteria: the three conditions under which we stop.
## Appendix: what we cannot show yet
- Claims in this paper that are not yet evidenced, listed plainly.
Two notes on using it. First, the rejected options in section 2 are the part that makes it persuasive. A paper with one option in it is a request, not a case. Second, the appendix is not optional. Naming what you cannot yet prove is what makes the rest of the paper believable, and it is the section that most people delete before the meeting.
A system that did reach production
Every page on this subject explains why pilots fail. Very few link to anything that did not.
QSQuoter is a system we built and run. On 14 July 2026 it priced a first-floor build-over extension in Heaton Moor, Stockport, issued by Cheadle Construction under reference JAY-AA-20260713: sixty-three measured lines totalling £98,214.72 including VAT, published in full so that every line traces to a dimension somebody can check. Preliminaries and site set-up at £17,374. Bathrooms and en-suites at £15,962. Eight more sections, all readable.
That is a bounded claim and we keep it bounded. It shows we can build software that prices a construction job to an auditable standard and put it into daily use. It does not prove hours saved, a return on investment, or anything at all about a different industry. What we cannot show you yet is listed on our evidence page, which is also where the job can be opened.
What this does not apply to
This is written for UK businesses of roughly £1m to £50m turnover running their own operations: an operations director at a 40-person distributor, a contracts manager at a construction firm, a finance lead at a clinic group.
It does not apply to research pilots, where a negative result is the output and production was never the goal. It does not cover regulated financial or clinical decisioning, which carries approval requirements far beyond this page. It is not legal advice on UK GDPR; the ICO’s own guidance is linked below and should be read directly. And the statistics quoted are drawn from a study with a mostly enterprise and United States sample, so they describe a pattern rather than your odds.
When we are the wrong choice
If you want a proof of concept for a board slide, we are the wrong supplier. We build things that run, and that is a slower and less impressive first meeting.
If the pilot exists to justify a purchase already agreed, the four checks will be unwelcome.
If nobody can be named as the owner, no supplier can fix that, and we will say so on the free fit call rather than sell you an audit.
If the process changes every few months, a pilot will be out of date before it finishes. Wait.
Sources
Read on 21 September 2026.
- MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 (PDF, July 2025). The 95 per cent claim, the 60 / 20 / 5 per cent funnel, the named causes, the research period of January to June 2025, the methodology of 52 interviews, 153 survey responses and 300 public initiatives, and the stated limitations quoted above.
- ICO, artificial intelligence guidance. Guidance on AI and data protection, and on explaining decisions made with AI.
- GOV.UK, AI Playbook for the UK Government. Published 10 February 2025.
- GOV.UK AI, write a business case. The Five Case Model structure used in the template above: Strategic, Economic, Commercial, Financial, Management.
- GOV.UK, the Green Book. The appraisal standard the Five Case Model sits inside.
- QSQuoter. Reference JAY-AA-20260713, 14 July 2026, 63 measured lines, £98,214.72.
Related reading
The short answers
Each answer is the opening of the section it links to, so nothing here is written for a crawler that a reader cannot also see.
What does the 95 per cent AI failure figure actually measure?
It comes from MIT's Project NANDA report, The GenAI Divide: State of AI in Business 2025, published in July 2025. The headline is that despite very large enterprise investment, 95 per cent of organisations are getting zero return, and that just 5 per cent of integrated AI pilots are extracting significant value.
Which four checks stop an AI pilot stalling before production?
All four failures are visible before anything is built, which is why we run these before a build is quoted rather than after it disappoints.
What do UK data protection rules add to an AI pilot?
Every page we read on this question was written for a United States audience and none of them mentions UK data protection, the ICO, or a pound sign. That is a real gap for a British reader, because the UK rules change the sequence of work rather than just adding paperwork at the end.
What does it cost to check an AI pilot properly?
The checks above cost an afternoon of internal time. A structured version, run across the whole operation with the arithmetic written down, is our fourteen-day AI Constraint Audit at a fixed £2,000, with three sixty-minute working sessions and four documents you own. It is described in full, including what it excludes, on the AI Constraint Audit page. The fee is not credited against a build, which is the mechanism that allows the answer to be no.
When should you kill an AI pilot instead of rescuing it?
Kill it when the data check failed and the fix is a data project nobody has budgeted. You are no longer testing the idea, you are funding a different project through the back door.
What goes into an AI business case a UK board will sign off?
Here is the template, in full, with no form and no email gate. It maps to the HM Treasury Five Case Model, which is the standard the UK government publishes for its own business cases: Strategic, Economic, Commercial, Financial and Management. Using it does not make you a public body. It makes your paper recognisable to a finance director who has seen one before.
What percentage of AI pilots fail?
The most cited figure is 95 per cent of organisations getting zero return, from MIT's Project NANDA report published in July 2025. The same report's funnel for task-specific tools is more precise: about 60 per cent evaluated, 20 per cent piloted, 5 per cent reached production. The authors describe these as directionally accurate rather than measured.
Why do AI pilots stall between proof of concept and production?
Four reasons recur. Nobody owned the change, the data behaved differently outside the demo set, the surrounding process was never redesigned, and no payback horizon was agreed so no decision date ever arrived. None of them is about model quality.
How long should an AI pilot run before you decide?
Set the decision date before the pilot starts, and make it weeks rather than quarters. A pilot without a decision date does not fail, it fades, which is worse because nobody learns anything and the budget is spent either way.
What is the difference between a pilot and a proof of concept?
A proof of concept answers whether the thing can work, usually on prepared data. A pilot answers whether it works here, on live data, with the people who will use it. Most stalled projects are proofs of concept that were presented as pilots.
Do UK GDPR and the ICO apply to an internal AI pilot?
If the pilot processes personal data, yes, including when it is internal and small. The ICO publishes guidance on AI and data protection covering accountability, fairness and explaining decisions made with AI. Deciding this after the pilot succeeds is the expensive order to do it in.
Who should own an AI pilot?
The person who owns the process being changed, not the person who owns the technology. A pilot owned by IT tends to be judged on whether the system worked. A pilot owned by operations is judged on whether the work got easier, which is the question that matters.
What should a pilot cost?
Less than the annual cost of the problem it addresses, by a comfortable margin. Published UK entry points run from about £500 for a single automated workflow to £5,000 for a supported build. If you cannot say what the problem costs today, the pilot has no budget, it has a guess.
Is the business case template really free?
Yes. It is on this page in full, there is no form, no email address and no download gate. Copy it, rename the headings, delete the sections you do not need. If it is useful you will remember where it came from.