AI & automation Our take

Why most AI pilots fail, and how to run one that does not

Pilots impress in the demo and stall before production. Six reasons, two gates nobody defines, and how to run one that ends in a decision.

4 minread 808words last updated

The short answer

Most AI pilots impress in the demo, run for a few weeks with a small group, and then neither ship nor stop. They fade. The reasons are rarely technical: no baseline was measured, so nobody can say whether it helped; no success criteria were agreed, so nobody can say whether it passed; the demo data was cleaner than the real data; nobody was assigned to own it in production; no monthly budget existed for running it; and nobody planned for the people whose work would change. Two gates prevent the fade, and both are easy to leave unwritten: after the proof of concept, is the real data good enough; after the pilot, who owns it and what does it cost per month.

Diagram: three stages, proof of concept, pilot and production, each with what it proves and does not prove and its typical duration, with two gates between them: gate one asks whether the real data is good enough, gate two, marked as where projects stall, asks who owns it and what it costs per month.
Each stage proves one thing. Each gate asks one question.

The six reasons

ReasonWhat it looks likePrevention
No baseline”It feels faster”Measure the current process for two weeks before building
No success criteriaThe pilot ends and nobody knows if it passedAgree numbers and a date before it starts
Demo dataWorked on ten clean examples; real inputs are messyEvaluate on real cases from day one
No production owner”IT will take it from here”; IT was not askedName the owner before the pilot
No running budgetNobody priced the monthly linePut usage, hosting, monitoring and review in the business case
No people planThe team keeps working the old way alongsideInvolve them early; change the procedure with them

Running a pilot that ends in a decision

  1. Baseline the current process: time, volume, errors, speed.
  2. Write the criteria and the date: what the numbers must be, what the constraints are, when the decision is made.
  3. Name the production owner and the monthly budget now, as a condition of starting.
  4. Build with an evaluation set from real cases, labelled by the people who do the work.
  5. Run with a person at the gate and a log of every decision.
  6. At the date, decide from the numbers: go to production with the owner and budget, or stop with a written lesson.
  7. If go, treat production as a project: integration, monitoring, procedures, training, not a flag flipped.

What the stages actually prove

A proof of concept proves the idea can work, on sample data, in days, and should be thrown away. A pilot proves it works for real people on real data at small scale, against a baseline, in weeks. Production proves it keeps working: monitored, owned, budgeted, supported, with a way to switch it off. Each stage answers one question. Skipping a stage means answering its question in the next one, at greater cost and in front of more people.

What this means for you

Before the next pilot, insist on five things in writing: a measured baseline, success criteria with a date, a named production owner, a monthly running budget and a plan for the people whose work changes. Evaluate on real cases from the start. At the date, decide. The pilot that ends in a decision, either way, is the one that moves the business forward. The one that fades moves nothing except the calendar.

Written by the CivSec S.M.A.R.T team

We build and run websites, software and AI systems for businesses. We write about what we see in that work, in plain language, and we update articles when things change.

Last checked . Spotted something outdated? Tell us.

Frequently asked questions

Our pilot worked well. Why is it still not in production a year later?

Because the pilot answered a different question from the one production asks. The pilot showed it can work for a few people on selected data. Production needs an owner, monitoring, a monthly budget, integration with the systems everyone uses, and changes to how people work. Nobody assigned those, so the pilot stayed a pilot until interest moved on.

How do we set success criteria for a pilot?

From the baseline. Take the metrics of the current process, time, errors, speed, volume, and state what the pilot must achieve on each to be worth taking to production, plus the constraints: cost per case, review load, no unacceptable failures. Write it down with a date. At the date, the numbers decide.

Is it bad to stop a pilot?

It is the second-best outcome and far better than the most common one, which is the fade. A pilot stopped on evidence teaches the business what its data and processes can support and costs a few weeks. A pilot that neither stops nor ships consumes attention for a year and teaches nothing.