Automation in a factory has an easy relationship with proof. There is a unit, the unit is countable, and the count either goes up or it does not.

Knowledge work has no unit. A sales team does not produce widgets, it produces attention allocated to opportunities, some of which convert for reasons nobody fully controls. A back office does not produce output, it produces the absence of problems. When an AI system lands in the middle of that, the honest answer to did it work? is frequently everyone feels like it did, which is not an answer a finance function can sign.

This is the most common reason we see AI programmes quietly defunded. Not because they failed. Because nobody could demonstrate that they had not.

The fix is not a cleverer measurement technique after the fact. It is deciding, before anything is built, what would count as evidence.

Define the unit during Diagnose, not after Deploy

Our second stage exists partly to produce this. Before a system is designed, we agree on the unit of work it is meant to affect and on how that unit is currently counted.

A unit has to be three things. Countable, so there is no argument about the number. Attributable, so a change in it can plausibly be traced to the system rather than to the season. Owned, so a specific person is accountable for it moving.

Calls scored per reviewer per day is a unit. Enquiries answered within the firm's own response standard is a unit. Documents cleared without a second pass is a unit. "Productivity" is not a unit, and neither is "efficiency."

Where a business cannot name the unit, that is the finding. It usually means the process has never been managed numerically, and the first value we deliver is the measurement itself, before any automation at all.

Three measures that hold up

Across engagements, three measures survive scrutiny from both operators and finance. We try to land on at least two.

Throughput per person. How many units one person clears in a period. This is the measure most directly attributable to a system, because headcount is known and the period is fixed. Its weakness is that it invites the assumption that the gain converts to reduced headcount. Usually it does not and should not; it converts to the same team covering more, which is a different and often better argument.

Cycle time. How long a unit takes from arrival to resolution, measured end to end including the waiting. This is the measure clients underestimate and customers notice. In sales operations it is frequently the whole ballgame, because the commercial value of a response decays sharply with delay and the decay is measurable in the business's own historical data.

Escape rate. How many units get through with an error, or fall out of the process entirely and are never resolved. This is the measure that converts an operational story into a revenue story, because each escape can be priced. It is also the one that most often justifies the project on its own.

Revenue is the measure everyone asks for first and the one we are most cautious about, for reasons below.

Accuracy is a floor, not a result

A number like accuracy is necessary and routinely misread.

Our call auditing work holds at 93% agreement with the client's own reviewers scoring the same conversations. That figure does one job: it establishes that the system's output can be trusted enough to act on without a human re-checking every case. It is a licence to operate.

It is not a business result. Nothing about 93% tells you whether the business is better off. The business result is what happens downstream: that every call gets reviewed instead of the four percent a manager had time for, that coaching is directed at the specific behaviours costing deals, that a problem surfaces this week rather than in a quarterly sample.

Two disciplines follow from this. Measure accuracy against the client's own experts on the client's own cases, not against a public benchmark. And never report accuracy as the outcome. The outcome is the operational change it permitted.

Baseline before you deploy, or you never get one

This is the single most expensive mistake available, and it is entirely avoidable.

Once a system is live, the pre-system number is gone. You cannot reconstruct it from memory, and the estimates people offer afterwards are shaped by how they feel about the system. Every argument about value from that point forward is an argument about anecdotes.

So we capture the baseline during Diagnose, from the business's own records, and we get it agreed in writing by the person who will be asked about the value later. Where records do not exist, we measure for a period before building. That period is never wasted: it almost always reveals something about the process that changes the design.

The baseline is one of the named artefacts of our method, and Measure is a stage of its own rather than a closing paragraph in a report.

The honest problem: attribution

One caution, because the alternative is to oversell.

Knowledge work sits inside a business where other things are also changing. A new hire joins. A competitor exits. Pricing moves. When throughput improves by a third, some portion of that is the system and some portion is everything else, and no method available to us separates them perfectly.

What can be done is to narrow the claim until it is defensible. Compare the same team against itself rather than against a different team. Compare the same period last year to account for seasonality. Hold out a segment where the system is not deployed where the business can tolerate it. Report the range rather than the flattering end of it.

A client who is told the system contributed somewhere between a fifth and a third of an improvement, with the reasoning shown, will trust the next number we bring them. A client told it contributed all of it will eventually discover otherwise, and then nothing we measure will be believed again.

What this looks like in practice

Measure produces a document with four things in it: the unit, the baseline, the current figure, and the honest attribution range. It goes to the person who signed off on the business case at the start, and it is compared against what that business case promised.

Sometimes the comparison is uncomfortable. That is the point of writing the promise down. A stage that can only confirm success is not a measurement stage, it is marketing, and it is worth nothing the second time.

Book your free AI transformation audit.

Sixty minutes. We map one function of your business, show you where the time and the money are leaking, and hand you the business case — whether or not you build it with us.

What you walk away with

  • A map of how one function really runs today
  • The hours and rupees it is currently costing you
  • A build sequence, priced and sequenced
  • A clear answer on whether AI is even the right fix

Book a free AI transformation audit