Skip to article

When a rule has earned its place

Arun MalikSeriesAI AgentsAutomation

Series date follows the editorial schedule. First published ; updated .

A rule earns its place by matching a stated contract within a stated boundary. Repeating an agent's answer is not enough.

That sounds obvious until the demonstration looks convincing. An assistant interprets several responses correctly. We extract the pattern, write a conditional, and run the same examples again. Everything passes. What we have shown is that the conditional reproduces those examples. We still need to ask whether the intended behavior is right, where it applies, and what happens outside it.

This article includes a small executable example. It has no network access, no model calls, and no connection to real equipment. Its fixtures and clock values are independently invented for teaching. It exercises a contract; it does not measure operational performance.

Make the eligible domain smaller than the problem

The problem is “help someone understand an interface.” The rule's domain is much narrower: classify one locally supplied record in the fictional lab-status/1 format when its target, timestamp, and status field meet the contract.

The function accepts only the exact fields schema, target, observedAtMs, evidenceId, and state. The status must be exactly up or down. It returns “appears up” or “appears down,” with an evidence reference. It does not interpret free text, infer a cause, or change anything.

That makes the rule an alternative to interpretation only for the structured variant. It is not a replacement for the ambiguous-note workflow in Part 5. Keeping that workflow hybrid, or requiring human interpretation, remains defensible.

Inside the boundary

Supported schema, exact target, usable timestamp, valid evidence ID, and an agreed status enum. Existing permission covers reporting.

Outside the boundary

Unknown format or state, extra fields, stale or future evidence, missing data, or a different target. Return a review reason.

Authority refused

Denied read, denied report, missing context, or an operation other than status reporting. Return a denial, not an interpretation.

The rule has distinct report, review, and denied outcomes. An out-of-scope case is not silently routed to a model.

Run the demonstration

The source files are lab-status.cjs and lab-status.test.cjs. Save both in the same directory. With Node.js installed, run:

node --test lab-status.test.cjs

No package installation is required. The test file uses Node's built-in test runner and assertions. The implementation is a pure function over supplied JavaScript values. Production input decoding, request-size limits, source authentication, real authorization, and external delivery are intentionally absent.

The harness supplies readAllowed and reportAllowed flags. These are test doubles for permission decisions, not a security mechanism. A real caller must not be allowed to grant itself access by setting a Boolean. The demonstration checks that the rule respects the supplied decision; it does not prove that an upstream authorization service made the right one.

The fixture clock uses milliseconds and a deliberately chosen freshness boundary. The test accepts an observation exactly at the limit and rejects one just beyond it. These values are teaching inputs, not a recommended operational timeout. A real freshness policy must account for its producer, clocks, and intended use.

Write expected behavior independently of the answer

The named cases encode expectations directly from the contract. They are not labeled by asking a model whether its own answer looks correct. That is a useful separation, but it is not independent external adjudication: the example and its tests were prepared together for this article.

Accepted observations

Known up and down states. Evidence exactly at the declared age boundary. Expected result: the corresponding report.

Rejected evidence

Missing fields, unexpected fields, wrong target, unsupported schema, invalid time, stale evidence, or unknown state. Expected result: a specific review reason.

Denied authority

Read or report denied, a truthy string instead of authorization, or a reset request. Expected result: denial before interpreting the observation.

Repeat evaluation

Frozen inputs, evaluated twice. Expected result: equal output and unchanged input values. No external side effect exists in this function.

This is a fixture coverage map, not an accuracy chart. Passing these categories says nothing about how frequently they occur in a real workload.

Some cases are deliberately inconvenient. Uppercase UP is rejected because the fictional contract defines a case-sensitive enum. An added instruction field is rejected because this version permits no extra fields. That strictness is a design choice, not a universal parser rule. If forward compatibility is required, specify how unknown fields are handled and test that different contract.

The implementation also denies a read before examining malformed evidence. This makes denial behavior independent of whether the supplied record happens to look healthy. It does not replace checks at the actual collection boundary, because this local function receives an observation that has already been supplied.

A serious replacement decision would add cases withheld from rule development, independent domain review of expected outputs, and tests at integration boundaries. Do not call a new fixture a holdout after using its failure to tune the rule. Once it influences development, it belongs to the development evidence.

What was actually verified

On September 27, 2026, the files linked above were run on Windows with Node.js v24.12.0. The built-in runner reported 29 passing tests and no failures. These are named synthetic contract tests, not a sample-based estimate of accuracy, a comparison with an agent, or a benchmark of cost or latency.

The exact SHA-256 hashes for that run are:

lab-status.cjs
72016c55b5f85cc0ec870217382f6f13ddd66039845c956b3e312b4c7bf60171

lab-status.test.cjs
ce8c029622001627ac54f907447fe4acd3c4c88e334ac1ac53411db220c51f22

The rule assumes the producer's field meaning is correct and the supplied context is trustworthy. It cannot detect a producer that confidently reports the wrong state. It does not exercise a real collection timeout, report delivery, a model, or approval expiry. It does not establish general equivalence with an interpretation system.

The Google SRE testing chapter explicitly cautions that passing tests does not necessarily prove reliability. It also distinguishes unit, integration, and system tests. Our pure-function cases cover one unit. Calling them end-to-end validation would conceal precisely the boundaries we still need to examine.

The acceptance decision is narrower than “replace AI”

For this toy contract, acceptance means the documented fixture expectations pass, there is no runtime inference in the function, and unsupported evidence does not produce a status guess. The useful conclusion is that this explicit enum mapping can be implemented as ordinary code.

For a real replacement, the reviewer must also decide whether the eligible domain is useful and whether false reports, rejected cases, and human follow-up are acceptable for the actual task. Those criteria come from consequences and workload, not a universal passing percentage.

If a new schema breaks the contract, disable that route and return a review outcome while its owner investigates. Reverting to a previous implementation is an option only if that version still fits the input and remains authorized. Calling a more flexible model automatically would change the boundary.

The rule has earned one small job. It has not earned the rest of the workflow.

Sources and scope

  1. Alex Perry and Max Luebbe, Testing for Reliability, Google SRE, Chapter 17: test levels, uncertainty under change, and limitations of passing tests.
  2. The two linked source files and their named assertions are the primary evidence for the local demonstration. The recorded environment and hashes identify what was run.

The example is original, synthetic, and nonproduction. It reports no employer implementation, internal measurement, production accuracy, or comparative savings.