Skip to content
AVEXIAINNOVATIONS PRIVATE LIMITED

Technology · 2 June 2026 · 8 min

Most AI pilots should fail. Design them so failure is cheap

The problem with enterprise AI is not that pilots fail. It is that they are structured so failure cannot be admitted.

Most AI pilots should fail. Design them so failure is cheap

A well-designed portfolio of AI pilots should have a failure rate somewhere around half. If yours is at ninety percent success, you are either extraordinarily lucky or you are choosing problems too small to matter.

The difficulty is rarely the technology. It is that most pilots are funded as commitments rather than as experiments, so admitting the result costs someone their credibility. The pilot then quietly becomes a production system nobody wanted.

Set the threshold before you build

Every pilot we run carries a written threshold agreed at kickoff: the specific cost or revenue line it must move, by how much, measured how, by when. If it does not clear the threshold, the recommendation is to stop, and that recommendation is not a failure of the team.

  • Name the line in the P&L the pilot is supposed to move.
  • State the minimum movement that justifies phase two, in currency, not in percentages of a percentage.
  • Fix the measurement method in advance so it cannot be reinterpreted after the fact.
  • Cap the pilot at six weeks. Anything longer becomes politically expensive to cancel.

Evaluation is the deliverable

For anything built on a language model, the evaluation harness matters more than the prompt. Model behaviour drifts, your data changes, and an unmeasured system degrades silently. We treat the eval suite as the primary artefact and the application as the thing attached to it.

A pilot without a pre-agreed kill threshold is not an experiment. It is a commitment wearing an experiment's clothes.

Design the human path first

Where the cost of a wrong answer is real — a credit decision, a clinical note, a legal summary — the review path is part of the design, not a mitigation added after the first incident. That means deciding what a reviewer sees, how long they have, and what happens to the model output when they disagree.

In practice

Of the AI pilots we ran for clients last year, a little under half did not clear their threshold. Every one of those was stopped at the end of six weeks, and that is the reason clients keep funding the next one.

RI

R. Iyer

Partner, Technology & Digital Innovation

INNOVATE · BUILD · SCALE · TOGETHER

Turn the reading into a decision.

If any of this describes your situation, the first call is with the partner who would run the work.

Or write to hello@avexia.co · +91 22 4890 1200