Skip to article

Measure the outcome, not just the model calls

Arun MalikSeriesAI AgentsAutomation

Series date follows the editorial schedule. First published ; updated .

A workflow can make fewer model calls and still cost more to operate.

Perhaps the rule rejects more inputs, so reviewers spend longer resolving them. Perhaps the new implementation needs frequent updates. Perhaps it looks fast because its timer stops before a report reaches the person who needs it. An inference bill cannot answer those questions.

The measurement task starts with a definition of finished work. For our fictional lab, that might be an authorized, evidence-supported status report, or an appropriate stop with a usable reason. It is not “the function returned without throwing.”

Keep the population visible

Suppose a proposed router sends exact enum responses to the deterministic checker and ambiguous notes to the hybrid workflow. The checker will probably receive different work. Comparing their average times as though the inputs were interchangeable would answer a poorly defined question.

Instead, define the comparison's eligible population first. If the question is whether a rule is a useful replacement for interpreting the agreed enum, compare routes on those same eligible records. Keep ambiguous notes as a separate category, and report how much of the intended task remains outside the rule's scope.

A narrow rule can still be valuable. The problem is letting its narrowness disappear from the denominator.

Fixed question and eligible cases

For the lab's agreed structured format, can a direct rule produce acceptable status reports with less total work than the proposed interpretation route?

Same task, authority, evidence, and outcome criteria

Route under consideration

Record collection, reasoning, verification, approvals, retries, and handoffs.

Alternative route

Record the same boundaries, including work moved to people or maintenance.

Retain failures, exclusions, unknowns, and timeouts.

Reviewed comparison

Assess outcome quality and effort together. State mismatches, uncertainty, and what was not measured.

A proposed study design, not a completed experiment. The equal-sized boxes encode comparable scope, not equal results.

A metric dictionary before a dashboard

These are candidate definitions. They contain no measured values and do not imply that every team should collect every metric.

Runtime model calls per attempted eligible run
Count every call inside the declared boundary, including retries and assisted recovery. Record model, prompt, and workflow versions. A no-runtime-AI component is not a no-runtime-AI end-to-end process if another step calls a model.
Time to verified outcome or terminal stop
Measure from the agreed start event through queues, collection, approval, interpretation, and verification. Preserve timeout cases rather than removing them from the distribution. Report slow cases as well as the middle.
Total effort and cost
Separate runtime and tool charges, model usage, storage, review, authoring, maintenance, and recovery. State the observation window and how shared work is allocated. Keep one-time implementation effort distinct from recurring effort.
Outcome quality and error
Define expected outcomes independently of the implementation being evaluated. Distinguish unsupported reports, missed reports, wrong diagnoses, appropriate abstentions, and unavailable evidence. Severity matters as well as frequency.
Reviewed handoff
Record why the route stopped, what the reviewer received, how much work remained, and whether the handoff resolved the request. Fewer escalations can mean better coverage or fewer safeguards; the count alone cannot distinguish them.
Coverage by pattern
Count eligible, excluded, attempted, reported, and reviewed cases separately for each input pattern. Do not quietly expand or shrink eligibility between comparison windows.
Contributor effort
Include domain experts' authoring time, reviewers' corrections, operator follow-up, and engineering maintenance. A playbook with no runtime model may still depend on substantial human work.
Recurrence, only when relevant
If preventing repeat failures is an objective, define the event, exposure, and recurrence window. A status-reporting improvement need not prevent failures. Do not use recurrence as the universal outcome for every playbook.

For an economic summary, one candidate is total attributable cost over verified outcomes in a declared window. Put failed attempts and recovery in the numerator. If there are no verified outcomes, the ratio is undefined; do not display a reassuring zero. And show the underlying outcome counts, because a single ratio can conceal a drop in useful coverage.

Google's automation chapter discusses both the value of consistency and the difficulty of calculating time savings. It also describes automation falling out of step with changing systems. Those cautions are good reasons to include maintenance and recovery instead of treating inference cost as the whole bill.

A comparison protocol for the lab

I would record the following before collecting any comparative result.

  1. Task and eligibility. Compare reporting for the supported structured format. Keep ambiguous notes, unknown schemas, and denied requests visible as separate categories rather than silently treating them as successful exclusions.
  2. Alternatives and versions. Freeze the rule, interpretation contract, model configuration if used, and verification policy. A changed model or a more permissive reviewer changes the experiment.
  3. Expected outcomes. Have someone qualified to judge the scenario review the labels without using the evaluated system's answer as truth. Record disagreements and unresolved cases.
  4. Observation boundary. Include authoring and recurring work under explicit accounting rules. State whether collection, approval, delivery, and maintenance are included or excluded.
  5. Failure handling. Preserve timeouts and unsupported outputs. Define whether an appropriate review outcome satisfies the task or remains unfinished work.
  6. Analysis and limitations. Separate patterns, disclose workload changes, and avoid treating repeated copies of one fixture as independent evidence about a broad population.

For a real authorized evaluation, a randomized comparison may help separate the route's effect from changing traffic, but it is only appropriate when the alternatives and assignment are themselves permitted. An offline replay can compare parts of a workflow without performing external actions, yet cannot reproduce every queue, human decision, or side effect. A matched observational comparison still has possible confounders.

The article proposes a protocol; it reports none of those studies. The passing fixtures in Part 6 provide unit-level contract evidence, not comparative accuracy, latency, or savings.

Put each intended claim through two gates

The first gate is permission to disclose. The second is whether the evidence supports the sentence. Passing either one does not pass the other.

An internal result does not become publishable because someone rounds it, turns it into a percentage, or draws it relative to an arbitrary baseline. A qualitative sentence such as “it saved substantial effort” can still disclose an internal outcome. Public availability of a paper does not establish employer approval for every reuse of its claims.

For a claims ledger, record the proposed sentence, its source and version, the applicable disclosure status, the population and window, and its limitations. A pending authorization means omit the claim. An independently constructed synthetic example should be labeled as such and should not be made by perturbing private results.

The evidence gate then asks whether the comparison was fair, whether failures were retained, and whether the design supports causal language. “We observed fewer calls in this window” and “this design reduced total operating cost” are different claims.

If the evaluation eventually shows fewer calls but more reviewer work, report both. That is a decision-relevant result, not an inconvenience to remove from the chart.

Sources and scope

  1. Niall Murphy with John Looney and Michael Kacirek, The Evolution of Automation at Google, Chapter 7: consistency, time-saving tradeoffs, and changing systems.
  2. Alex Perry and Max Luebbe, Testing for Reliability, Chapter 17: different test scopes and limits of evidence from passing tests.

The metric dictionary and protocol are this series' proposed methodology. There are no operational measurements, estimated gains, internal aggregates, or transformed private results in this article. No disclosure approval or empirical validation is asserted.