Free worksheet · Enablement

Measure whether AI is changing the work, not just whether licences are used

Licence dashboards show who logged in, not whether the work changed. Enter your own figures for one workflow: people, licences, weekly active users, task volume, minutes per task before and with AI, and rework and quality pass rates. The worksheet returns licence utilisation, workflow adoption, gross and rework-adjusted hours saved, quality change, and flags such as high usage with no time saved.

Free tool · runs in your browser · nothing is sent unless you choose to send a brief

Measure one workflow at a time. Enter your own figures; the grey numbers are an example only, and blank fields are treated as missing, not zero. The worksheet does the arithmetic and flags patterns worth questioning. It cannot check that your figures are right.

Reach

Who does this workflow.

Seats paid for this group.

From the tool’s admin report, for this group.

Volume and time

Across the whole group.

Per week.

Measured, not recalled, if possible.

min

Include checking and editing.

min
Quality

Share of tasks sent back or redone.

%

Same definition as before.

%

Same reviewer or checklist.

%

Same reviewer or checklist.

%

This is a good fit if…

  • You own an AI rollout and the only number you can report is how many licences are active.
  • A finance sponsor has asked what the tools are worth, and you do not want to claim every saved minute as cash.
  • You are running a pilot on one workflow and need a consistent before-and-after measure.

Look elsewhere if…

  • You want someone to design and run measurement across a programme with several teams. That is part of an enablement pilot or rollout rescue, not a worksheet.
  • You need to decide whether to expand a finished pilot. Use the pilot scorecard, with these figures as its evidence.

What you get

Editable measurement specification with baseline, comparison and capacity-versus-cash distinctions

  • Licence utilisation, reach and workflow adoption, kept separate so broad but shallow use is visible.
  • Gross hours saved per week and a rework-adjusted figure that counts the cost of fixing outputs.
  • Quality change in percentage points, using the same check before and after.
  • Flags that point to the next question: no baseline, high use with no saving, or time saved at the cost of quality.

How it runs

  1. 01

    Choose one workflow

    Measure a single, recurring task with a clear output. Averages across unrelated work hide everything useful.

  2. 02

    Record a baseline

    Ideally two weeks of time per task and a quality measure before AI is introduced, on the same group or a comparable one.

  3. 03

    Measure again on the same definitions

    After a few weeks of use, enter the with-AI figures. Same task definition, same reviewer or checklist, same counting rules.

Not included

  • No financial valuation of hours saved. The worksheet reports capacity, not cash.
  • No check that your figures are accurate or comparable.

Why licence usage is the wrong headline

A rollout report that says “80% of licences active” answers the question the vendor cares about. It does not tell you whether any piece of work is faster, better or cheaper. People can open a tool every week and use it for nothing that matters, or use it for a task where checking the output takes longer than doing the work.

The worksheet separates three questions that usually get blended together:

  • Are people using it? Licence utilisation (weekly active users ÷ licences) and reach (weekly active users ÷ people in scope).
  • Is it used for this work? Workflow adoption: the share of tasks of this type done with AI assistance.
  • Did the work change? Time per task, rework and quality, before and after, on the same definitions.

How the figures are calculated

  • Gross hours saved per week = AI-assisted tasks per week × (baseline minutes − minutes with AI) ÷ 60. “Minutes with AI” must include checking and editing.
  • Rework-adjusted hours saved = AI-assisted tasks × (baseline minutes × (1 + rework rate before) − minutes with AI × (1 + rework rate after)) ÷ 60. This assumes a reworked task costs roughly its own time again.
  • Quality change = quality check pass rate with AI − pass rate before, in percentage points.
  • Capacity is shown as working days per week at 7.5 hours a day, with the reminder that capacity is not cash.

Blank fields are treated as missing, not zero, so a missing baseline produces a flag rather than a misleading figure.

Capacity is not cash

Hours saved become a financial result only when something changes: more work handled by the same team, less overtime or contractor spend, or a role not backfilled. Until then they are capacity, and the honest report says what that capacity was used for. This distinction is what lets a finance sponsor trust the rest of the numbers.

Illustrative example

Illustrative example using the worksheet’s example figures, not client data. A support team of 24 people has 24 licences and 15 weekly active users. Of 120 complaint replies a week, 70 are drafted with AI. Baseline time was 40 minutes per reply; with AI, including review, it is 28. Rework rose from 8% to 12%, and the quality check pass rate fell from 90% to 88%.

The worksheet reports licence utilisation and reach of 63%, workflow adoption of 58%, 14.0 gross hours saved a week and 13.8 after rework, which is about 1.8 working days of capacity. It raises one flag: time saved but quality fell. The right next step is not to celebrate the hours. It is to ask the reviewer whether a two-point fall in pass rate is acceptable for complaint replies, and what change to the prompt, template or review step would recover it.

Assumptions and limitations

  • Every figure is yours. The worksheet does arithmetic and pattern checks; it cannot tell whether the figures are accurate or whether the before and after groups are comparable.
  • It measures one workflow at a time. Adding unrelated tasks together hides the patterns the flags look for.
  • The rework adjustment is a simplification. Where rework is a quick correction rather than a full redo, the adjusted figure understates the saving.
  • It says nothing about what other organisations achieve. There are no benchmarks in it, deliberately.

What to do next

Use these figures as the evidence for the pilot scorecard when deciding whether to expand, revise or stop. If you have not yet chosen a workflow or recorded a baseline, start with the readiness self-assessment.

If the flags show a rollout where usage is high and the work has not changed, that is the situation AI adoption rescue is for. If you want help setting up measurement as part of a bounded pilot, with a named internal owner at the end, see AI enablement.

Questions buyers ask

Why not just report hours saved as money?

Because saved time only becomes money if something changes: the team takes on more work with the same people, overtime or contractor spend falls, or a vacancy is not backfilled. Otherwise it is capacity. Reporting it as cash invites a finance team to look for a saving that does not exist, which damages trust in the whole programme.

Where do I get weekly active users and task counts?

Weekly active users usually come from the AI tool’s admin console, filtered to the group in scope. Task counts come from the system where the work happens: tickets closed, documents drafted, invoices processed. If neither exists, a two-week tally by the team is better than an estimate from memory.

What if we have no baseline?

The worksheet flags it. You can still compare against a group that does not use AI yet, or pause AI use on a sample of tasks for a week. Without either, treat any time saving as an estimate and say so when you report it.

Why is rework part of the time saving?

Because a draft produced in half the time that has to be redone more often may save nothing. The rework-adjusted figure assumes a reworked task takes roughly its own time again. If rework in your workflow is lighter or heavier than that, adjust your interpretation accordingly.

How often should we measure?

Weekly for the first month of a pilot, then monthly. Look at trends over several weeks rather than a single week: early enthusiasm and holiday periods both distort short samples.

Use the worksheet or discuss your requirement

A short, non-confidential description is enough to start. I read every brief personally and reply within two business days, including when the answer is that I am not the right fit.

Step 1 of 2 · The basics