Skip to article

One scenario, several defensible designs

Arun MalikSeriesAI AgentsAutomation

Series date follows the editorial schedule. First published ; updated .

The lab operator asks a small question: “What does this observation tell us about the interface?” There is no single execution form that is right for every version of that question.

If the observation follows an agreed enum, a rule may be enough. If it is an ambiguous note, interpretation remains. If the next useful observation is not known in advance, an authorized investigation may need agent-led orchestration. All three can produce reusable work.

This final article follows those choices through the same fictional scenario. The traces are explanatory sketches, not recorded agent runs. The only executed implementation is the local, side-effect-free fixture rule linked in Part 6.

Fix the task before comparing designs

The invented lab has an approved target called fixture-link. The intended output is a supported status report with an evidence reference, or an explicit reason that reporting cannot proceed. No design has permission to reset the interface, investigate another target, or assert a root cause from a status field alone.

The structured observation uses lab-status/1 and reports state: down. An alternate input is a human-written note that may require interpretation. Those inputs belong to different eligibility categories; they should not be mixed into an apparent performance comparison.

A lab expert can supply the known format and its meaning directly. An investigation can instead reveal that the response is ambiguous or that the current procedure made an unsupported causal inference. The source changes what evidence the reviewer has. It does not dictate the execution form.

Knowledge already available

Expert-authored meaning, scenario boundaries, and exceptions.

Knowledge learned through investigation

Observations, corrections, competing explanations, and unresolved questions.

Review either source against the same task and authority.

Next step depends on findings

Consider an agent-led investigation within a declared scope.

Procedure fixed, interpretation unresolved

Consider a hybrid flow or explicit human interpretation.

Eligible behavior fully specified

Consider a deterministic rule from the start.

Any selected design still needs its own evidence, permissions, and owner.

No supported or authorized route?

Stop and hand off for review. Do not use a more capable executor to bypass the missing condition.

A selection guide, not an automatic router. The criteria are design questions requiring review, not a claim that the system can classify its own uncertainty perfectly.

Three sketches of the work

Agent-led sketch. The operator authorizes a bounded investigation. The agent inspects the permitted observation, sees “down,” and asks whether the supplied format definition distinguishes operational state from administrative intent. If an approved source answers that question, it can use the answer; otherwise it hands the unresolved question to the DRI. It presents the evidence without inventing a cause. The agent chose the next useful step, even if the investigation followed a reusable playbook.

This design may be useful when the evidence or path is unfamiliar. It has more discretion over permitted investigation, not more authority over the target. The sketch assumes a supported tool boundary; it is not an implemented agent.

Hybrid sketch. A fixed procedure validates access and evidence, then supplies the alternate operator note to a bounded interpretation step. The model proposes “appears down” or “uncertain” with evidence references. The validator checks the output's structure and reference membership. Where semantic support is not mechanically established, an authorized reviewer decides whether to release the status interpretation. The fixed branch either reports the supported result or produces a review packet.

This design retains judgment where the wording calls for it. It may remain hybrid indefinitely. Its cost includes the reviewer when review is required, and the sketch does not claim that constraining the output makes the interpretation accurate.

Deterministic sketch. The rule receives the supported structured observation and trusted fixture context. It checks the requested operation, supplied permission decisions, exact fields, target, timestamp, and enum. For the eligible down value, it returns:

{
  "outcome": "report",
  "status": "appears_down",
  "evidenceId": "observation-a",
  "contractVersion": "lab-status-rule/1"
}

That output is an object returned locally, not a report delivered to a person. The “known down” fixture checks this behavior. The function makes no runtime model calls and performs no collection or external delivery. Its permission flags are harness inputs, not real authorization.

The expert could have authored this rule before any agent ran. The investigation route is not a prerequisite.

Now change the input

Change the structured response to lab-status/2. The implemented rule returns a review outcome for an unsupported schema. It does not ask a model to guess the new meaning. The owner can investigate under separately granted authority and decide whether a revised contract is justified.

Make the observation stale, and the rule refuses to treat it as current. Deny read permission, and the rule returns a denial before interpreting the supplied observation. Ask for a reset, and the operation is refused because it is outside the contract. These are exercised fixture cases, not production access-control tests.

The hypothetical agent-led and hybrid designs need corresponding failure behavior at their own boundaries. A model's ability to understand unfamiliar text does not establish permission to consume it or act on it. A natural-language explanation of a stale observation does not make the observation fresh.

If a future workflow adds report delivery, it also adds uncertain side effects and retry obligations. The pure function's repeatability does not cover them. The recovery runbook explains why a lost acknowledgement needs reconciliation rather than a blind restart.

Choose where the behavior belongs

For the published demonstration, keep the rule local. There is no deployed service to integrate with and no authorized production workload. That is the concrete decision in the fictional architecture record.

For an actual authorized system, engineers could keep the understood behavior in a maintained playbook or incorporate it into service code. They would compare ownership, change cadence, caller needs, and recovery responsibilities. A service implementation might make reporting more consistent without preventing the interface from going down.

A mixed design is also coherent: a deterministic status checker inside a hybrid reporting flow, itself called during an agent-led investigation. Name each boundary honestly. Do not advertise the whole process as no-runtime-AI because one useful component no longer needs inference.

Choose your next question

The series is organized in reading order, but the reusable pieces are easier to find by question:

Discovering or supplying knowledge, capturing a contract, isolating reasoning, validating a rule, and choosing a location are useful development activities. They can repeat, overlap, or be skipped when unnecessary. They are not five additional execution types.

What this series establishes

The public sources provide engineering distinctions and cautions. Anthropic's agent and workflow guidance supports asking who directs execution and starting with a simpler design when it fits. The Google SRE testing chapter explains why tests establish evidence at particular boundaries rather than universal reliability.

The series contributes a proposed way to organize those choices, original fictional contracts, and a reproducible toy rule. It does not establish comparative model accuracy, production safety, cost savings, automatic promotion, or a universally optimal lifecycle. No employer deployment or internal result is offered as validation.

For your own scenario, the first useful artifact may be a tested rule. It may be a better question for the DRI. Choose the form that fits what is known, and leave what is not known visible.