Measuring AI success: metrics that are not vanity

Usage counts make AI look successful; business metrics tell you whether it is. The four layers of measurement and the vanity traps.

3 minread 616words last updated

The short answer

AI systems come with dashboards full of numbers that go up: conversations, messages, documents processed, users, satisfaction scores. Those are usage and sentiment, and they make any system look successful. Whether it is successful is a different question, answered by business metrics measured against a baseline: did revenue, cost, speed to the customer, error rate or capacity change, and by how much. Four layers of measurement keep this honest: business outcomes at the top, process metrics beneath, system quality beneath that, and adoption at the base. Vanity lives in the bottom two layers and is presented as if it were the top. A quarterly scorecard with one line per system across the four layers is enough for a small business.

The four layers

LayerExamplesVanity version
Business outcomesConversion rate, cost per enquiry, days sales outstanding, revenue per employee, complaints”Customers love it”
Process metricsTime per case, cases per person per day, response time, error and rework rate, backlog”We processed ten thousand documents”
System qualityMonthly sample accuracy, confident-wrong rate, escalation rate, guardrail firings, cost per case”Ninety-five percent accurate” with no baseline or sample method
AdoptionShare of eligible work going through the system, workarounds found, team sentiment”Two thousand active users”

Building the scorecard

  1. Per system, pick one or two business outcomes it should move, and take their baseline.
  2. Pick the process metrics that connect the system to those outcomes: time, volume, errors, speed.
  3. Set the quality measures: monthly sample score, escalation rate, cost per case.
  4. Set the adoption measures: share of eligible work through the system, workarounds.
  5. Agree the review dates before launch: three months and six months.
  6. One line per system, four columns, on the quarterly governance agenda.

Reading the scorecard honestly

A system with high adoption and quality but no movement in business outcomes is a well-built tool doing the wrong job, or one whose freed time was never redeployed. A system with strong outcomes and falling quality is living on borrowed time. A system with low adoption and everything else fine has a change-management problem, not a technology one. The four layers together tell you which conversation to have; any one alone tells you a story someone chose.

What this means for you

Measure AI the way you measure anything else the business spends on: by what changed in the numbers that matter, against a baseline, at dates agreed in advance. Keep usage and satisfaction in their place as means, not ends. A one-line-per-system scorecard across the four layers, read every quarter, is enough to know what is working, what needs fixing and what should be switched off.

Written by the CivSec S.M.A.R.T team

We build and run websites, software and AI systems for businesses. We write about what we see in that work, in plain language, and we update articles when things change.

Last checked . Spotted something outdated? Tell us.

Frequently asked questions

Our assistant has thousands of conversations a month. Is that success?

It is usage, which is necessary and not sufficient. Success is what those conversations changed: fewer tickets reaching staff, faster resolution, higher conversion, lower cost per enquiry, measured against the months before. Thousands of conversations that resolve nothing, or that people have because they cannot find a human, are a cost.

What is a fair timescale to judge an AI system?

Three months for a first reading and six for a judgement, against a baseline taken before launch. Earlier readings measure the ramp, not the steady state. Later readings without a baseline measure nothing. Agree the dates and the metrics before launch; it is the only way the judgement is fair to everyone.

How do we measure quality without a data science team?

Sample. Each month, someone who knows the work scores a random set of the system's outputs as good, acceptable or wrong, the same way the evaluation set was scored before launch. Thirty cases a month gives a trend. It takes an hour and it is more informative than any dashboard the vendor provides.