Skip to article

Keep the playbook, or move behavior into the service?

Arun MalikSeriesAI AgentsAutomation

Series date follows the editorial schedule. First published ; updated .

Once the status check has an explicit contract, a new question appears: should it stay in the playbook or move into the service?

Removing the runtime model does not answer that question. A deterministic workflow can remain outside the service. Engineers can also implement the understood behavior in service code. Those choices change ownership, deployment, and failure boundaries, even when the decision rule stays the same.

And neither choice necessarily fixes why the interface went down. A clearer status report is a useful outcome in its own right.

Keep the contract stable while comparing locations

In our fictional lab, the agreed rule reports an eligible status observation and refuses unknown formats. We can place it in two different designs without changing that promise.

Shared behavioral contract

Validate scope and evidence. Interpret only supported status values. Report with provenance, or stop with a reason. Do not diagnose cause or authorize mitigation.

Two optional locations for understood behavior

Maintained runtime playbook

A separate executor collects permitted evidence and applies the rule. The playbook owner manages its release, dependencies, and recovery.

Service implementation

Engineers place the understood behavior behind a service interface. The service owner manages its deployment, compatibility, and operational load.

Both retain permission checks, observability, tests, and an owner.

This is a location fork, not a final stage of crystallization. Service integration can improve handling while leaving an underlying fault unchanged.

The playbook version might obtain the raw response through a permitted adapter and call the local rule. The service version might expose a status-report operation that applies the same contract near the producer. In the second design, the service still needs to establish who may ask about which target and what it may disclose.

The function from Part 6 could inform either implementation. Moving its file into a service repository would not, by itself, supply authentication, rate limits, a trustworthy clock, or delivery semantics. Integration is engineering work, not a change of folder.

Who can keep the meaning current?

Suppose the lab's status producer changes its schema. A separately maintained playbook has to learn about that change, test it, and coordinate its release. A service implementation maintained with the producer may make that coordination easier. But only if the owner actually treats the status contract as part of the service's responsibilities.

The Google SRE automation chapter discusses external automation falling out of step with the system it manages. It argues for handling some use cases directly in applications rather than depending on external glue. That is a useful reason to consider integration. It is not a requirement that every useful playbook disappear into service code.

A playbook may need to coordinate several independently owned systems. Putting that procedure inside one service can give that service responsibility for other teams' release schedules and access policies. A separate workflow with a clear owner may be easier to maintain.

Change cadence also matters. A diagnostic procedure used by authorized operators may need revisions on a different schedule from a core service. Conversely, a status contract relied on by many callers may belong with the service that defines its semantics. Count the coordination work on both sides.

Deployment ownership is part of correctness here. A perfectly tested rule in a component nobody knows how to update will eventually become a problem someone has to diagnose.

A fictional architecture decision record

Here is a filled decision record for the lab example. Its assumptions are invented, not inferred from an employer's system.

Context
The current teaching exercise has one local fixture producer and one consumer. There is no deployed service or authorized external collection. The goal is to demonstrate the bounded status contract, not operate a network.
Decision
Keep the demonstration as a side-effect-free rule called by the local test harness. Do not add a service merely to give the development story a destination.
Alternative: maintained playbook
For a future authorized lab, a separate playbook could own collection, contract selection, reporting, and reviewed handoff. It would need an operational owner and tests for the external boundaries absent from the fixture demonstration.
Alternative: service integration
If the producer's maintainers decide that consistent status reporting belongs in their service, they could implement the contract there and expose a versioned interface. That requires a deployment and compatibility decision, not just approval of the rule.
Consequences
The present example stays easy to inspect but demonstrates no service behavior. A future playbook adds orchestration responsibility. A future endpoint adds service maintenance and exposure. Neither has been measured as cheaper.
Reconsider when
There is an actual authorized consumer, a stable owner, a demonstrated coordination problem, or a producer change that makes the current placement unsuitable. Reopen the decision when those assumptions change.

The most important entry is the decision not to build the service yet. The absence of a service is a limitation of the exercise, not unfinished progress toward a prescribed endpoint.

Moving the behavior changes the tests

Suppose a future review approves the endpoint. Preserve the narrow contract tests, then add tests for the new boundary: caller authorization, target binding, producer identity, response-size limits, clock handling, and version compatibility. Unit tests passing in both repositories would not establish that the deployed endpoint uses the expected configuration.

Define how clients handle an unavailable endpoint. A client should not interpret an error page as “down,” call an older incompatible parser, or fall back to an agent with more access. A reviewed fallback is a separate supported path.

During an authorized rollout, compare permitted outputs without duplicating consequential actions. If the design merely reports local fixture status, comparison is straightforward. If it later submits external reports or performs changes, running both implementations can itself create duplicate effects. The AWS idempotency discussion is relevant to those receiver-side semantics, not a guarantee that parallel rollout is harmless.

Rollback needs a compatibility condition. Restoring old code helps only when the old contract still fits the producer and callers. If the schema changed irreversibly, disabling the affected path and handing off may be the appropriate recovery instead.

Finally, decide what to retire. An integrated status rule might let you remove a duplicated parser while retaining the playbook that gathers other evidence or coordinates review. Retirement can be partial. The location decision should eliminate unnecessary duplication without pretending that all diagnostic work has been absorbed.

Put the behavior where someone can keep its contract true. That may be a service. It may be a playbook that remains useful for years.

Sources and scope

  1. Niall Murphy with John Looney and Michael Kacirek, The Evolution of Automation at Google, Chapter 7, “The Use Cases for Automation”: external automation, changing systems, and handling use cases inside applications.
  2. Amazon Builders' Library, Making retries safe with idempotent APIs: explicit intent and receiver-side coordination for repeated effects.

The architecture decision record is fictional. No service endpoint, migration, production rollout, or root-cause repair was implemented or measured for this article.