Measure the outcome, not just the model calls
Series date follows the editorial schedule. First published ; updated .
Progressive Crystallization · Part 8 of 10
A workflow can make fewer model calls and still cost more to operate.
Perhaps the rule rejects more inputs, so reviewers spend longer resolving them. Perhaps the new implementation needs frequent updates. Perhaps it looks fast because its timer stops before a report reaches the person who needs it. An inference bill cannot answer those questions.
The measurement task starts with a definition of finished work. For our fictional lab, that might be an authorized, evidence-supported status report, or an appropriate stop with a usable reason. It is not “the function returned without throwing.”
Keep the population visible
Suppose a proposed router sends exact enum responses to the deterministic checker and ambiguous notes to the hybrid workflow. The checker will probably receive different work. Comparing their average times as though the inputs were interchangeable would answer a poorly defined question.
Instead, define the comparison's eligible population first. If the question is whether a rule is a useful replacement for interpreting the agreed enum, compare routes on those same eligible records. Keep ambiguous notes as a separate category, and report how much of the intended task remains outside the rule's scope.
A narrow rule can still be valuable. The problem is letting its narrowness disappear from the denominator.
For the lab's agreed structured format, can a direct rule produce acceptable status reports with less total work than the proposed interpretation route?
Same task, authority, evidence, and outcome criteria
Record collection, reasoning, verification, approvals, retries, and handoffs.
Record the same boundaries, including work moved to people or maintenance.
Retain failures, exclusions, unknowns, and timeouts.
Assess outcome quality and effort together. State mismatches, uncertainty, and what was not measured.
A metric dictionary before a dashboard
These are candidate definitions. They contain no measured values and do not imply that every team should collect every metric.
- Runtime model calls per attempted eligible run
- Count every call inside the declared boundary, including retries and assisted recovery. Record model, prompt, and workflow versions. A no-runtime-AI component is not a no-runtime-AI end-to-end process if another step calls a model.
- Time to verified outcome or terminal stop
- Measure from the agreed start event through queues, collection, approval, interpretation, and verification. Preserve timeout cases rather than removing them from the distribution. Report slow cases as well as the middle.
- Total effort and cost
- Separate runtime and tool charges, model usage, storage, review, authoring, maintenance, and recovery. State the observation window and how shared work is allocated. Keep one-time implementation effort distinct from recurring effort.
- Outcome quality and error
- Define expected outcomes independently of the implementation being evaluated. Distinguish unsupported reports, missed reports, wrong diagnoses, appropriate abstentions, and unavailable evidence. Severity matters as well as frequency.
- Reviewed handoff
- Record why the route stopped, what the reviewer received, how much work remained, and whether the handoff resolved the request. Fewer escalations can mean better coverage or fewer safeguards; the count alone cannot distinguish them.
- Coverage by pattern
- Count eligible, excluded, attempted, reported, and reviewed cases separately for each input pattern. Do not quietly expand or shrink eligibility between comparison windows.
- Contributor effort
- Include domain experts' authoring time, reviewers' corrections, operator follow-up, and engineering maintenance. A playbook with no runtime model may still depend on substantial human work.
- Recurrence, only when relevant
- If preventing repeat failures is an objective, define the event, exposure, and recurrence window. A status-reporting improvement need not prevent failures. Do not use recurrence as the universal outcome for every playbook.
For an economic summary, one candidate is total attributable cost over verified outcomes in a declared window. Put failed attempts and recovery in the numerator. If there are no verified outcomes, the ratio is undefined; do not display a reassuring zero. And show the underlying outcome counts, because a single ratio can conceal a drop in useful coverage.
Google's automation chapter discusses both the value of consistency and the difficulty of calculating time savings. It also describes automation falling out of step with changing systems. Those cautions are good reasons to include maintenance and recovery instead of treating inference cost as the whole bill.
A comparison protocol for the lab
I would record the following before collecting any comparative result.
- Task and eligibility. Compare reporting for the supported structured format. Keep ambiguous notes, unknown schemas, and denied requests visible as separate categories rather than silently treating them as successful exclusions.
- Alternatives and versions. Freeze the rule, interpretation contract, model configuration if used, and verification policy. A changed model or a more permissive reviewer changes the experiment.
- Expected outcomes. Have someone qualified to judge the scenario review the labels without using the evaluated system's answer as truth. Record disagreements and unresolved cases.
- Observation boundary. Include authoring and recurring work under explicit accounting rules. State whether collection, approval, delivery, and maintenance are included or excluded.
- Failure handling. Preserve timeouts and unsupported outputs. Define whether an appropriate review outcome satisfies the task or remains unfinished work.
- Analysis and limitations. Separate patterns, disclose workload changes, and avoid treating repeated copies of one fixture as independent evidence about a broad population.
For a real authorized evaluation, a randomized comparison may help separate the route's effect from changing traffic, but it is only appropriate when the alternatives and assignment are themselves permitted. An offline replay can compare parts of a workflow without performing external actions, yet cannot reproduce every queue, human decision, or side effect. A matched observational comparison still has possible confounders.
The article proposes a protocol; it reports none of those studies. The passing fixtures in Part 6 provide unit-level contract evidence, not comparative accuracy, latency, or savings.
Put each intended claim through two gates
The first gate is permission to disclose. The second is whether the evidence supports the sentence. Passing either one does not pass the other.
An internal result does not become publishable because someone rounds it, turns it into a percentage, or draws it relative to an arbitrary baseline. A qualitative sentence such as “it saved substantial effort” can still disclose an internal outcome. Public availability of a paper does not establish employer approval for every reuse of its claims.
For a claims ledger, record the proposed sentence, its source and version, the applicable disclosure status, the population and window, and its limitations. A pending authorization means omit the claim. An independently constructed synthetic example should be labeled as such and should not be made by perturbing private results.
The evidence gate then asks whether the comparison was fair, whether failures were retained, and whether the design supports causal language. “We observed fewer calls in this window” and “this design reduced total operating cost” are different claims.
If the evaluation eventually shows fewer calls but more reviewer work, report both. That is a decision-relevant result, not an inconvenience to remove from the chart.
Sources and scope
- Niall Murphy with John Looney and Michael Kacirek, The Evolution of Automation at Google, Chapter 7: consistency, time-saving tradeoffs, and changing systems.
- Alex Perry and Max Luebbe, Testing for Reliability, Chapter 17: different test scopes and limits of evidence from passing tests.
The metric dictionary and protocol are this series' proposed methodology. There are no operational measurements, estimated gains, internal aggregates, or transformed private results in this article. No disclosure approval or empirical validation is asserted.