Two ways to get a useful playbook
Progressive Crystallization · Part 3 of 10
“We learned this during an incident” and “the operator already knew this” can lead to the same useful playbook. They leave different questions for its reviewer.
An investigation gives you observations, decisions, and sometimes a successful outcome. An expert-authored procedure gives you a proposed explanation of how the task should work. Neither arrives with its applicability already proved.
Progressive crystallization needs both routes. Otherwise we make people run an agent to discover knowledge they already have, or mistake a plausible execution trace for a specification.
Two notes from an invented lab
Return to the fictional interface-status check from Part 2. Everything in the notes below is newly constructed teaching material. There is no real device, incident, internal troubleshooting guide, or operational result behind it.
The permitted response said state: down. The assistant proposed that the interface was unavailable. The operator pointed out that the interface might have been deliberately disabled. The final report preserved the observation without asserting a cause.
For the lab's fictional lab-status/1 format, report the state field when it is exactly up or down. Do not infer cause. Unsupported values, missing evidence, or a mismatched target require review.
Both sources enter the same review, but contribute different evidence.
Report an eligible observation with its provenance. Specify format, target, freshness, authorization, and rejection behavior before choosing an execution form.
The investigation note is useful because it records a correction. If we kept only the assistant's successful final answer, we would lose the reason to prohibit a causal claim. The operator's intervention belongs in the capture record.
The expert note is useful because it can make exploration unnecessary. If its contract is confirmed, the status mapping can start as a deterministic parser. But the reviewer still needs to ask what “state” means, which producer owns that meaning, and whether the field says anything about the user's experience. A procedure can be confidently authored and wrong.
Capture the disagreement, not just the winning path
A transcript tends to preserve the sequence that happened. A reusable procedure must say what should happen in cases the transcript never saw. Those are different artifacts.
For an investigation-derived candidate, keep the observation separate from the interpretation. In the invented note, the observation is a field value. “The interface appears down” is an interpretation within the fictional schema. “The cable is faulty” would be an unsupported diagnosis. Resetting the interface would be a proposed action, requiring a different task and authority.
This separation makes rejected explanations valuable. The expert's correction tells the next author not to turn an operational status into a causal rule. A timeout might likewise tell us that a missing response cannot be treated as “down.” Absence of evidence is a distinct outcome.
OpenTelemetry's tracing documentation describes a trace as the path of a request through an application, built from spans representing units of work. Timestamps, parent relationships, and events help reconstruct execution. That is useful infrastructure for capture. It does not decide whether an inference was justified or whether a branch should be permitted next time.
Do not turn “record the evidence” into “log everything.” For an actual system, the capture design must decide what may be retained, who can read it, and when it expires. An evidence reference can point to a controlled record. It need not copy sensitive input into a broadly readable playbook.
A worksheet that works for either source
I would put the following record beside a candidate procedure. It is a proposed authoring aid, not an implementation requirement for every organization. The filled entries describe only our fictional lab.
- Scenario and intended outcome
- Report whether one approved lab interface appears up or down from an eligible status response. Do not diagnose cause or change the interface.
- Knowledge source and provenance
- Either the invented investigation note above, including the operator's correction, or the invented operator-authored format description. Keep the source identifiable instead of blending them into “the system learned.”
- Applicability
- The agreed response format, requested target, and evidence freshness must match the contract. A response from another target is not supporting evidence.
- Assumptions
- The producer and its schema are trusted within the lab exercise. Real deployment would need to establish producer identity and the reliability of its observations separately.
- Exceptions
- Missing, stale, malformed, contradictory, or unsupported evidence produces a review outcome, not a guessed status.
- Owner and review
- A named lab maintainer owns the candidate contract. A reviewer checks its interpretation and tests. Authorship alone does not grant collection or execution privileges.
- Unresolved questions
- How does the producer represent a deliberately disabled interface? What clock assumptions govern freshness? Can two valid observations disagree? These questions block broader claims even if the narrow status report remains useful.
- Revision trigger
- A schema change, an observed contradiction, or a change in the intended outcome reopens review. Record the affected contract version.
The blank version is just these headings without the lab answers. Keeping the unresolved-questions field is important. Otherwise the polished procedure hides uncertainty that was visible in the original conversation.
The source does not choose the form
A domain expert might author an agent-led playbook: investigate within this scope, consult these permitted sources, stop before these actions, and present competing explanations to the DRI. The procedure captures expertise without predicting every step.
An investigation might instead reveal an already documented, stable enum. After reviewing the contract, an engineer can write a direct parser. There is no obligation to build an intermediate hybrid just to preserve a story about stages.
And both sources may leave a genuinely ambiguous interpretation. A hybrid workflow can fix collection and reporting while preserving that decision for a model or human. Keeping the ambiguity visible is more useful than claiming that capture finished it.
The Google SRE chapter on testing describes tests as a way to demonstrate specific areas of equivalence when systems change, and cautions that passing tests does not necessarily prove reliability. Applied here, the next step is to turn the candidate contract into checkable expectations. Neither the transcript nor the expert's name substitutes for that work.
Before asking how to automate a captured procedure, ask what was actually captured: an observation, an interpretation, an accepted requirement, or an unresolved question. A useful playbook keeps those distinctions intact.
Sources and scope
- OpenTelemetry, Traces: request paths, spans, timestamps, events, and parent relationships. A tracing data model, not a specification-extraction method or a logging-permission policy.
- Alex Perry and Max Luebbe, Testing for Reliability, Google SRE, Chapter 17: testing changes, uncertainty, and limits of passing tests.
The notes and worksheet are original fictional examples. No investigation was performed, no employee account is being retold, and no empirical advantage is claimed for either knowledge source.