← All insights

Measure AI value at the level of work

Count the improvement in a completed business outcome, including the review, correction, and downstream work required to achieve it.

AI value should be measured where work reaches an acceptable outcome. A system can generate more content, answer more questions, or prepare recommendations faster while the surrounding process becomes harder to manage. If people spend the saved time checking, correcting, or explaining the output, the organisation needs to account for that effort.

Usage is useful evidence of adoption. It tells leaders whether people are trying the service and where demand is growing. It does not establish that the work is better. The business case needs a connection between use, process performance, and the strategic outcome that justified the investment.

Define acceptable completion

Consider an illustrative assistant that drafts responses to supplier exceptions. Producing a draft is an intermediate event. A useful outcome is a response that is accurate, authorised, understood by the recipient, and capable of moving the issue toward resolution. An attractive draft that requires extensive correction has not yet delivered the intended benefit.

Before evaluation, the team should agree what acceptable completion means and how it will judge it. The criteria might include factual correctness, use of current information, compliance with the decision owner’s mandate, and a clear next action. The criteria should reflect the actual consequences of the work rather than a generic score for writing quality.

The baseline should capture the existing process with the same boundary. If the current measure ends when an employee starts drafting but the new measure ends at final approval, the comparison is misleading. Both need to include the steps that materially affect the outcome, including waiting for information and resolving an exception after the initial response.

Follow effort across the handoff

I would examine elapsed time and human effort separately. A response can require fewer staff minutes yet spend longer waiting in an approval queue. Conversely, a team may deliberately invest more review effort to reduce the risk of a consequential mistake. The right interpretation depends on the purpose and constraints of the service.

Downstream effects also deserve attention. A planning assistant that releases more recommendations could overload the people authorised to implement them. A drafting tool could increase the review backlog. The evaluation should ask whether the next team receives usable work and whether the overall decision cycle improves. Local speed is only one part of that picture.

The service should be examined across different cases. Routine requests and ambiguous exceptions may have very different results. A single average can conceal a useful narrow application or a serious weakness in a consequential part of the workload. Separating the cases helps leaders decide where to expand, where to redesign, and where to keep a manual process.

Turn the evidence into an operating decision

Each evaluation needs a business owner who can decide what the results mean. The owner should compare quality, time, effort, cost, and failure consequences with the original objective. Some benefits may be qualitative, such as clearer explanations or better access to knowledge, but they should still be supported by observable evidence and user feedback.

The review should also identify confounding changes. Better training, cleaner source data, or a revised approval process may contribute to improvement alongside AI. Recognising those contributions gives the organisation a more accurate account of what should be sustained. It also prevents technology from receiving credit for every improvement in a redesigned workflow.

The useful question is whether the organisation can now complete important work more dependably within acceptable cost and risk. That question keeps measurement connected to the purpose of the investment and gives leaders a defensible basis for the next decision.

Developed from my EA 878 capstone on the Cognitive Enterprise Architecture Framework (CEAF); EA 878 leadership paper on governance and adoption. These recommendations extend the coursework; examples are illustrative and do not report an employer assessment or measured results.

Bring this topic to your team or event. Connect with David.