Free scorecard · Enablement

Decide whether to expand, revise or stop an AI pilot on evidence

When a pilot looks promising but nobody can say whether to expand it, score it. Rate seven weighted criteria from 1 to 5: value against baseline, output quality, adoption by intended users, operating cost, risk and compliance, owner readiness and scalability. You get a weighted score and a go, revise or stop recommendation with reasons. A critical risk issue blocks a go, and stop is a valid outcome.

Free tool · runs in your browser · nothing is sent unless you choose to send a brief

Score the pilot on seven criteria from 1 (poor) to 5 (strong), using the evidence you have rather than impressions. Weights are shown next to each criterion. Tick the critical-issue box if any risk or compliance issue would block wider use; that alone prevents a “Go”.

Value against baseline · weight 20

1 = No measured change, or no baseline. 5 = Clear, measured improvement on the baseline.

Output quality and error rate · weight 20

1 = Errors frequent or unacceptable to reviewers. 5 = Quality at or above the pre-pilot level.

Adoption by intended users · weight 15

1 = Few of the intended users use it after the first weeks. 5 = Most intended users use it repeatedly, unprompted.

Operating cost and effort · weight 10

1 = Running it costs more effort or money than it returns. 5 = Running cost and effort are known and acceptable.

Risk and compliance · weight 15

1 = Open issues with data, security or compliance. 5 = Reviewed and signed off by the right people.

Owner readiness · weight 10

1 = Nobody will own it after the pilot team moves on. 5 = A named owner with time and budget to run it.

Scalability to more teams or volume · weight 10

1 = Works only with the pilot team’s special effort. 5 = Repeatable by other teams with documented steps.

0 of 7 scored

This is a good fit if…

  • You sponsor an AI pilot that is ending, and the team wants to expand it but the evidence is mixed.
  • You run several pilots and need a consistent way to decide which ones continue.
  • You want to agree the decision rules before a pilot starts, so the result is not argued after the fact.

Look elsewhere if…

  • You have not started a pilot yet and need to know whether the basics are in place. Use the readiness self-assessment.
  • The rollout has already stalled across the organisation and you need someone to find out why. Use AI adoption rescue.

What you get

Decision worksheet covering task quality, usage, risk, operating effort and ownership

  • A weighted score out of 100 across seven criteria.
  • A go, revise or stop recommendation with the specific reasons, including what blocked a go.
  • A list of weaknesses to fix before re-scoring, or what would need to change to revisit a stopped pilot.
  • A copyable record of the decision and the scores behind it.

How it runs

  1. 01

    Agree the criteria up front

    Ideally before the pilot starts, agree with the sponsor that this is how the decision will be made.

  2. 02

    Score from evidence

    Use measured figures where you have them, such as the adoption measurement worksheet, and the reviewers’ judgement on quality. Tick the critical-issue box if any risk would block wider use.

  3. 03

    Act on the recommendation

    Expand in a controlled step, fix the named weaknesses and re-score on a date, or stop and record what was learned.

Not included

  • No risk or compliance sign-off. The scorecard records your judgement; accountable owners still approve.

Why pilots drift instead of ending

Most AI pilots do not end with a decision. They end with a demo, an enthusiastic team, a sponsor who has moved on to the next priority, and a quiet extension “while we gather more data”. Six months later the pilot is still running, costing time and licences, and nobody can say whether it worked.

The cause is usually that nobody agreed beforehand what evidence would justify expanding it. The scorecard fixes that by making the criteria, weights and decision rules explicit.

The criteria and how they are weighted

CriterionWeightWhat a 5 looks like
Value against baseline20Clear, measured improvement on the baseline
Output quality and error rate20Quality at or above the pre-pilot level
Adoption by intended users15Most intended users use it repeatedly, unprompted
Risk and compliance15Reviewed and signed off by the right people
Operating cost and effort10Running cost and effort known and acceptable
Owner readiness10A named owner with time and budget to run it
Scalability10Repeatable by other teams with documented steps

Each score from 1 to 5 is converted to a 0–100 scale and weighted. The decision rules are:

  • Stop if the weighted score is below 45, or if value and adoption (or value and quality) are both very weak.
  • Go only if the score is at least 70, no criterion scored 1, value, quality and owner readiness are all 3 or higher, and no critical issue is flagged.
  • Revise in every other case, with the reasons that blocked a go listed.

A ticked critical issue on risk always prevents a go, however high the score.

Stop is a valid result

A stop recommendation is not a failure of the pilot. It means the pilot did its job: it tested an idea cheaply and found that this workflow, tool or timing does not justify more effort now. The scorecard lists what would need to change to revisit it, so the learning is kept, and the people and budget can move to something with better odds.

Illustrative example

Illustrative example, not a client result. A support team piloted AI-drafted first responses for eight weeks. The sponsor scores value 4 (handling time fell against a measured baseline), quality 3, adoption 4, operating cost 3, risk 2 with the critical-issue box ticked (some drafts quoted customer data from other tickets), owner readiness 4 and scalability 3.

The weighted score is 58. The recommendation is revise: the score is below the go threshold, the critical issue blocks expansion in any case, and risk is the weakness listed to fix before re-scoring. The team fixes the retrieval filter that let other customers’ data into drafts, has the data protection lead sign it off, and re-scores four weeks later.

Assumptions and limitations

The scorecard structures your judgement; it does not replace it. Scores are only as good as the evidence behind them, so use measured figures where you can. The thresholds and weights are a reasonable default rather than a standard, and they are shown so you can disagree with them. The scorecard does not provide risk, legal or compliance sign-off; those stay with the accountable people in your organisation.

What to do next

Gather the evidence for value, quality and adoption with the adoption measurement worksheet. If you are scoring a pilot that has not yet started, run the readiness self-assessment first.

If the result is revise and you want help running the next iteration, with a baseline, coaching and a named internal owner, see AI enablement. If several pilots have stalled and the organisation has lost confidence in the programme, AI adoption rescue is the more direct route.

Questions buyers ask

Why is stop presented as a good outcome?

Because a pilot exists to find out whether something is worth doing. A clear stop, with the reasons recorded, saves the budget and attention that a half-alive pilot would otherwise consume for months. It also makes it easier to run the next pilot, because people trust that the decision will be made on evidence.

Why can a single risk issue block a go?

Because risk does not average out. A pilot that saves time but exposes personal data, or draws a regulatory objection, cannot be expanded until that issue is resolved, however well it scores elsewhere. The scorecard turns such a pilot into revise or stop and says why.

Can we change the weights?

Not in the tool, but you can read the table: value and quality carry the most weight, followed by adoption and risk. If your situation differs, for example a regulated workflow where risk should weigh more, agree that with your sponsor and note it alongside the result.

What if we have no baseline to score value against?

Score value low, and the recommendation will reflect it. A pilot without a baseline can still produce useful learning, but it cannot show value. The usual next step is to revise: record a baseline on a comparable group and re-score after a few weeks.

Use the worksheet or discuss your requirement

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics