Skip to article

Permissions, recovery, and changed inputs

Arun MalikSeriesAI AgentsAutomation

Series date follows the editorial schedule. First published ; updated .

A changed input should not silently grant a workflow a more powerful executor.

Imagine the fictional lab producer starts returning a new status format. The deterministic checker refuses it. An agent might be able to interpret the response, but that does not establish permission to send the response to a model, query additional sources, or change the interface. The failed assumption and the proposed recovery are separate decisions.

This is where a useful automation design needs a runbook of its own. It must say what stops, what evidence survives, who owns the next decision, and how a permitted restart differs from continuing a failed run.

Stop the affected path, then establish what happened

The local fixture checker makes no external calls and has no side effects. Its recovery is simple: return the denial or review reason. A real workflow can already have done work before it discovers a problem. That changes the questions.

Suppose an extended, still fictional design collects a status observation, builds a report, and submits it to an approved lab mailbox. Collection succeeds. The delivery acknowledgement times out. We now know that an observation was collected and a delivery was attempted. We do not know whether the report arrived.

Re-running the whole playbook hides that distinction. It may collect a different observation and deliver a second report. “Try again” is not a recovery specification.

Trigger: denied, stale, changed, or incomplete

Stop the affected route. Do not infer a status from missing data, and do not infer failure of a side effect from a missing acknowledgement.

Preserve the permitted record of completed and uncertain work.

Reviewed handoff

The owner checks authority, evidence, contract version, and whether any effect may already have occurred.

Choose an authorized response, not an automatic escalation.

Close without restart

The request is denied, no longer needed, or cannot be supported by the evidence.

Reconcile or retry

Confirm prior effects, use supported idempotency semantics, and recheck current permission.

Revise or investigate

Approve a new contract or a separately scoped investigation before execution resumes.

The authority boundary remains in force for agent-led, hybrid, and deterministic recovery. A different form does not inherit broader privileges.

Give each fault its own meaning

A practical runbook should distinguish the following cases. These are proposed responses for our fictional lab, not reports of a deployed system.

Read denied
Do not collect the observation. Return a policy-safe reason to the requester. Do not switch identities or ask an agent to retrieve the same data by another route.
Approval expired
Stop before the operation that depended on it. Ask for a fresh decision covering the current target, action, and conditions. A previous approval should not become a reusable credential.
Stale or future-dated evidence
Do not classify it as current state. If a new observation is permitted, create a new collection attempt and retain the distinction. Investigate clock assumptions rather than automatically widening the freshness limit.
Changed format or meaning
Withdraw the affected parser route. Preserve a permitted example for contract review. A familiar field name does not guarantee that its meaning stayed the same.
Collection timeout
Record observation unavailable. A bounded retry may be allowed by the source's contract and current policy. Timeout does not mean the interface is down.
Delivery timeout
Record outcome unknown. Query the delivery system's supported status mechanism if authorized, or hand off. Retry only under a contract that handles a possibly completed prior delivery.
Duplicate work or partial completion
Identify which logical request and steps were completed. Reconcile before repeating effects. Do not assume every action has an inverse, or that an inverse would restore the original situation.

Calling every row “failed” throws away information the operator needs. Calling every row “escalate to an agent” changes the execution form without answering the recovery question.

The receiving system owns part of the retry contract

The Amazon Builders' Library discussion of idempotent APIs explains how caller-provided request identifiers express retry intent. It also explains that recording the identifier and performing the mutation must be coordinated atomically. Otherwise the system can record a request without completing it, or complete the effect without recording the request.

For the fictional mailbox extension, the playbook cannot create those semantics on its own. The receiver must define whether a repeated logical submission returns the prior result, how long it remembers requests, and how callers can resolve uncertainty. If that contract is unavailable, the recovery path may require human reconciliation rather than an automatic retry.

Even an idempotent request can require a fresh permission check. Idempotency concerns repeated effects; authorization concerns whether this identity may make this request now. The concepts solve different problems.

A pure local function returning the same object twice, as in Part 6, tests neither distributed delivery nor persistence across a crash. It is useful evidence at a smaller boundary.

Record enough to support the handoff

The owner should be able to distinguish attempted collection, accepted evidence, interpretation, review, and delivery. Attach the contract and implementation versions, relevant policy decision references, and terminal reason. Keep the observation's timestamp separate from the time a model interpreted it or a report was delivered.

OpenTelemetry traces provide spans and parent relationships for reconstructing a request's path. That can help correlate these steps. A completed span, however, is not proof that the reported status is correct or that a recipient received the intended report. Outcome verification needs the corresponding contract evidence.

Retention has boundaries too. Record authorized references rather than copying every input into a log. Protect evidence and approval records according to their sensitivity. A reviewer who can see a failure summary need not automatically gain access to all source material.

Reopening the design is part of operating it

Assign an owner before enabling the route. That owner needs a way to disable it, inspect permitted evidence, revise the contract, and decide whether to keep the current form. A drift alert without a responsible recipient is only a message.

The review may choose a new parser, a retained hybrid interpretation, or an agent-led investigation with explicitly authorized inputs and tools. It may choose no restart. None of those outcomes is a failure to progress.

OWASP's authorization guidance recommends checking permissions on every request and periodically reviewing permissions for unwanted expansion. That principle applies just as much to recovery as to the happy path.

The refusal to proceed is doing useful work when it tells the owner which promise no longer holds. Preserve that signal instead of hiding it behind a more flexible executor.

Sources and scope

  1. Amazon Builders' Library, Making retries safe with idempotent APIs: explicit retry intent and atomic handling of effects.
  2. OpenTelemetry, Traces: spans, context, events, and request-path reconstruction.
  3. OWASP, Authorization Cheat Sheet: least privilege, per-request checks, and permission review.

The mailbox, fault cases, and runbook are hypothetical. No distributed recovery system was implemented or tested for this article. The runnable example remains the side-effect-free fixture function from Part 6.