Many AI projects are judged by impressions: a good demo, a confident estimate of hours saved, an enthusiastic team. Those are easy to produce and easy to inflate. A better approach is to measure the automation against a baseline, with quality checks and the full cost, and to agree in advance what result would make you continue, revise, or stop.
Start with success criteria and a baseline
Anthropic's prompt engineering guide assumes you begin with a clear definition of success criteria and some way to test against them empirically. That guide is about prompts, but the same discipline applies to an automation. Anthropic's agents guide also describes customer support as a case where success can be clearly measured through user-defined resolutions. In our view, you should write down, before you build, what a good result looks like and how the work is done today.
Record the baseline for a representative period: how many items, how long each takes, how many are wrong or reworked, and how much it costs in people's time.
Measure five things
- Throughput and time. How many items are processed and how long each takes from start to finish, including waiting for review.
- Quality. How often is the result correct or acceptable? Check against a defined set of examples and a sample of live items, not impressions. OpenAI's guidance recommends tests and evaluation suites so you can monitor performance.
- Rework and exceptions. How many items need human correction, and how long does that take? An automation that is fast but needs constant fixing may not save effort.
- Cost. Count the whole cost: tool and model usage, review time, setup, and maintenance, not only the time saved on the main task.
- Adoption. Do the people who are supposed to use it actually use it, and what do they say? A technically good automation that nobody trusts delivers little.
Compare like with like
Use the same period type, the same kinds of items, and the same definition of "done" for the baseline and the automated version. Include difficult cases, not just easy ones, and note seasonal effects or changes in volume that could distort the comparison.
Be careful with these traps
- Estimated hours saved. Hours saved on one step are not hours saved overall if the work moves elsewhere. Measure end-to-end time and effort.
- Averages that hide failures. A high average accuracy can hide a category of cases that fails badly. Look at the failures separately.
- Ignoring review and maintenance costs. They are real and ongoing.
- Measuring too early or on too few items. Small samples give noisy results; say how many items the result is based on.
- Changing the definition of success after the fact.
Decide with a rule agreed in advance
Before you start, write the decision rule. For example: continue if quality meets the agreed threshold on a defined sample and total effort is lower than the baseline; revise if quality is close but specific failure modes are fixable; stop if neither holds. Then report the results, including what did not work, and make the decision.
Report honestly
A good report states the baseline, the method, the sample size, the quality results, the full cost, the failure modes, and the unresolved questions. It avoids guarantees and multipliers, and it makes clear what the numbers do and do not show.
This is the structure behind our AI Iteration Lab: one bounded workflow, a baseline comparison, an evaluation, and an executive decision at the end of the month. To choose a workflow worth measuring, read how to choose your first workflow to automate with AI.