Why pilots drift instead of ending
Most AI pilots do not end with a decision. They end with a demo, an enthusiastic team, a sponsor who has moved on to the next priority, and a quiet extension “while we gather more data”. Six months later the pilot is still running, costing time and licences, and nobody can say whether it worked.
The cause is usually that nobody agreed beforehand what evidence would justify expanding it. The scorecard fixes that by making the criteria, weights and decision rules explicit.
The criteria and how they are weighted
| Criterion | Weight | What a 5 looks like |
|---|---|---|
| Value against baseline | 20 | Clear, measured improvement on the baseline |
| Output quality and error rate | 20 | Quality at or above the pre-pilot level |
| Adoption by intended users | 15 | Most intended users use it repeatedly, unprompted |
| Risk and compliance | 15 | Reviewed and signed off by the right people |
| Operating cost and effort | 10 | Running cost and effort known and acceptable |
| Owner readiness | 10 | A named owner with time and budget to run it |
| Scalability | 10 | Repeatable by other teams with documented steps |
Each score from 1 to 5 is converted to a 0–100 scale and weighted. The decision rules are:
- Stop if the weighted score is below 45, or if value and adoption (or value and quality) are both very weak.
- Go only if the score is at least 70, no criterion scored 1, value, quality and owner readiness are all 3 or higher, and no critical issue is flagged.
- Revise in every other case, with the reasons that blocked a go listed.
A ticked critical issue on risk always prevents a go, however high the score.
Stop is a valid result
A stop recommendation is not a failure of the pilot. It means the pilot did its job: it tested an idea cheaply and found that this workflow, tool or timing does not justify more effort now. The scorecard lists what would need to change to revisit it, so the learning is kept, and the people and budget can move to something with better odds.
Illustrative example
Illustrative example, not a client result. A support team piloted AI-drafted first responses for eight weeks. The sponsor scores value 4 (handling time fell against a measured baseline), quality 3, adoption 4, operating cost 3, risk 2 with the critical-issue box ticked (some drafts quoted customer data from other tickets), owner readiness 4 and scalability 3.
The weighted score is 58. The recommendation is revise: the score is below the go threshold, the critical issue blocks expansion in any case, and risk is the weakness listed to fix before re-scoring. The team fixes the retrieval filter that let other customers’ data into drafts, has the data protection lead sign it off, and re-scores four weeks later.
Assumptions and limitations
The scorecard structures your judgement; it does not replace it. Scores are only as good as the evidence behind them, so use measured figures where you can. The thresholds and weights are a reasonable default rather than a standard, and they are shown so you can disagree with them. The scorecard does not provide risk, legal or compliance sign-off; those stay with the accountable people in your organisation.
What to do next
Gather the evidence for value, quality and adoption with the adoption measurement worksheet. If you are scoring a pilot that has not yet started, run the readiness self-assessment first.
If the result is revise and you want help running the next iteration, with a baseline, coaching and a named internal owner, see AI enablement. If several pilots have stalled and the organisation has lost confidence in the programme, AI adoption rescue is the more direct route.