A practical 30-day scorecard for UK SMEs testing AI: compare complete tasks, count review and rework, price the true cost, and decide whether to renew or stop.

An AI pilot is worth renewing when it improves a real task after you count the whole job, including checking, corrections and running costs. After 30 days, compare a defined set of cases with the old process, check that quality has held up, and decide whether to continue, change the design or stop. A busy AI dashboard is not the answer.

I would make that decision on one page. If you cannot describe the task, the cases tested, the time spent by people and the mistakes found, you have a demonstration rather than evidence of an operating improvement.

Why make the decision after a month?

A month is long enough for the tool to meet some ordinary interruptions: missing information, awkward requests, holidays and a colleague who was not in the original demo. It is short enough to avoid paying for a year of licences while hoping that value will appear. The right period depends on the task. A monthly finance close may need more than one cycle; a high-volume enquiry process may give useful evidence sooner.

Recent research explains why the question matters, but it does not answer it for your firm. In September, BGF reported that concerns about return on investment appeared among the growth businesses it interviewed across the UK and Ireland. A SAS report drawing on IDC research found that its UK impact index lagged its trustworthiness index. Those are different samples and measures. Neither tells you what your invoice, enquiry or document workflow achieved last Tuesday.

The UK government’s AI evaluation guidance is written for public interventions, but two of its practical points apply here: define the outcome early and document what the normal process does before changing it. I would scale that advice down to a small, repeatable test.

What exactly are you measuring?

Choose one job with a clear start and finish. “Help the team with AI” is too broad. “Turn an incoming supplier invoice into an approved, correctly coded entry” is measurable. The finish is the approved entry, rather than the moment an AI tool produces a draft.

Write down the cases that qualify. For an invoice pilot, that might mean ordinary domestic supplier invoices arriving by email, with exceptions such as credit notes and disputed amounts kept in a separate lane. Record the exclusions too. If a tool handles only the easy half of the queue, reporting its performance as the result for all invoices will mislead you.

Name the person who owns the measure. That person should be able to trace a sample from arrival to a usable outcome and ask the people doing the work where time actually went. Tool usage reports can help explain adoption; they do not establish business value.

How do you build a fair baseline?

Before switching on the AI step, sample cases from the current process. Record the date, case type, who did the work, elapsed time, active working time, review time, rework and whether the result was accepted first time. Keep a note of unusual cases. Do the same during the pilot.

Compare like with like. A month with a seasonal rush or a newly trained colleague is not directly comparable with a quiet fortnight run by your most experienced person. If the volume or case mix shifts, show that in the decision record. Where possible, keep a small set of similar cases on the original route during the pilot. That gives you a contemporary comparison as well as a historical baseline.

The government’s guidance calls this the “business-as-usual” comparison. You do not need to copy its full evaluation framework for a small internal trial. You do need to say what was compared, when, and where the comparison is weak. If there was no baseline, start one now and postpone any claim that the pilot saved money. A retrospective guess is useful for forming a hypothesis, not proving a result.

Which numbers belong on the scorecard?

I would keep the sheet short enough that someone will update it. For each period, show the same denominators and the same case rules.

MeasureRecord it this wayWhy it matters
Eligible casesNumber entering the agreed workflow, plus excluded casesStops easy-case selection from flattering the result
Completed outcomesCases accepted and usable at the defined finishSeparates drafts from finished work
Human minutesPreparation, prompting, checking, correction and hand-off per completed caseCaptures work moved from one step to another
End-to-end timeArrival to usable result, with waiting time noted separatelyShows what the customer or next team actually experiences
QualityFirst-time acceptance and serious errors, using a written definitionProtects against a fast but unreliable process
AdoptionEligible cases where the agreed AI step was actually usedExplains whether a weak result reflects the tool or non-use
CostSupplier fees, setup, integration, training and ongoing supportKeeps the subscription price from posing as total cost

Use a count and a rate for quality. “Two errors” means little without the number of cases inspected. More seriously, do not average away a privacy breach, incorrect price or unsafe recommendation because many easy cases passed. Some failures should trigger an immediate pause regardless of the mean result.

For the first month, check every output if the consequence of an error is material. Where the volume makes that impractical, use a documented sample and review every flagged exception. Your reviewer needs the original input and the approved source, not only the AI’s answer. Our Friday afternoon workflow test explains why ordinary messy cases should be in that sample.

How do you calculate the value without fooling yourself?

Start with minutes per completed case, then multiply by the volume that genuinely followed the new route. Subtract the time spent checking, repairing and administering the tool. Use the same costing rate on both routes and state what it includes. Add licence, setup and support costs separately.

Here is illustrative arithmetic, not an Ajairu client result. Suppose 40 comparable cases took 20 human minutes each before the pilot. During the pilot, 40 completed cases took 8 minutes each for preparation and checking, plus 80 minutes of extra rework in total. The apparent time released is 40 × (20 − 8) − 80 = 400 minutes, or six hours and 40 minutes. If the month also required training or configuration, count that time before claiming a net benefit.

Do not automatically multiply those minutes by an hourly rate and call the result cash saved. Capacity released is useful when the team can spend it on a valuable queue, meet a deadline or avoid extra paid hours. Cash savings require an actual spending change, such as fewer paid overtime hours or a retired supplier fee. Revenue needs its own evidence: a completed order or paid invoice, with a sensible comparison, not a lead that the AI marked “qualified”.

I would report three lines: capacity released, actual spending changed and service quality. If the first improves while the other two stay steady, that can still be a good result. It is simply a different kind of result. Our guide to what AI adoption means in practice makes the same distinction between using a tool and changing a business outcome.

What can make the result look better than it is?

The easiest trap is counting the AI’s first draft and ignoring the rest of the job. A second is giving the pilot only clean inputs while the original process handles the difficult ones. A third is watching a motivated pilot group for four weeks, then assuming everyone else will use the tool in the same way.

Watch for work that moves out of sight. If customer service closes tickets faster but finance spends longer fixing the resulting credits, the business has not saved the full amount shown on the service team’s dashboard. If a manager must read every output in detail, include those minutes. If an integration breaks and a colleague manually catches up, record the incident and recovery time.

The UK Business Data Survey 2026 found mixed AI governance and awareness among surveyed businesses. Its adoption figures vary with the definition and sample. That is a reminder to describe your own measure precisely. “We used AI” is not a scorecard. The team should know which data the pilot can access, who can pause it and what to do when an output looks wrong. The AI system owner article covers those operating responsibilities in more detail.

How do you decide: renew, repair or stop?

Agree the decision rule before the trial. I would write a short threshold for benefit, a quality floor and a list of failures that require a pause. Thresholds must fit the job; a missed appointment and an incorrect clinical instruction do not carry the same consequence.

At day 30, make one of three calls:

  1. Continue within the tested scope when the complete process improved, quality stayed within the agreed limit, the operating cost is understood and somebody owns the service. Review it again after a longer period.
  2. Change and retest when the tool helps but a specific hand-off, input rule or review step consumes the gain. Name the change and set another decision date. Do not roll out a new version on the strength of the old test.
  3. Stop or pause when the benefit disappears after checking and rework, the risk is unacceptable, or there is no reliable owner for errors and supplier changes. Stopping a weak pilot is a useful decision, not a failure to “do AI”.

Keep a one-page decision record: task and scope, dates and case counts, baseline and comparison, the seven measures above, incidents, total cost, decision, owner and next review date. Attach the raw case log so another person can challenge the conclusion. If the numbers are too thin, say “inconclusive” and collect better evidence rather than manufacturing a percentage.

The AI implementation service starts with a bounded workflow and a measurable outcome. If you have a pilot nearing renewal and want a second pair of eyes on its scorecard, contact me with the task, your baseline and the awkward cases. That is enough to have a useful first conversation.

AI Pilots

Recommended Reads