When an AI workflow should become ordinary automation
Progressive Crystallization · Part 1 of 10
A working AI workflow is a good place to start asking where you can stop using AI.
If a model keeps choosing the same operation for a well-defined input, there may be no useful decision left for it to make. Parsing a known file format, checking a required field, or sorting records by a declared key can be ordinary software. The difficult part is knowing whether you have found a stable rule or merely seen a few similar cases.
I use progressive crystallization here for a proposed engineering practice: move a bounded part of an AI-assisted workflow into explicit code when you can state its contract and test it. Leave unresolved judgment visible. This is an argument about how to design work, not a claim of a new algorithm or a report of measured savings.
The knowledge need not start with an agent. Operators, domain experts, and engineers can author a playbook from a procedure they already understand, or learn it through an investigation. Either source can support an agent-led, hybrid, or deterministic implementation. A fully agent-orchestrated playbook is still a playbook. Reuse does not require removing AI.
My public paper's execution taxonomy is the conceptual starting point. This series makes the choices explicit rather than prescribing a ladder: skip an unnecessary stage, keep a mixed workflow, or start with rules. Service integration is an optional location decision. It may improve handling without fixing the underlying cause. The paper is cited for terminology, not as evidence of savings, automatic promotion, or safety guarantees.
Investigation findings or proactively supplied expertise. Both need review.
Agent-led, hybrid, or deterministic. Choose for a stated scope, not a maturity score.
A maintained playbook, or understood behavior incorporated into service code.
Start with a reading list
Imagine a local experiment that turns synthetic book records into a reading-list page. There are no customer records, production services, or outgoing messages. An assistant reads two fixture files, works out their column names, combines the records, and suggests subject groupings. A person reviews the generated page.
At first, asking the assistant to inspect the files is convenient. Suppose the experiment then adopts two fixed input formats, each with an explicit version field. The mappings are now agreed: one format's book_title and the other's title both become title. Records have stable identifiers. The output order is specified.
Why ask a model to rediscover those mappings on every run?
A small program can accept the two known versions, require a nonempty title and a unique identifier, escape text for HTML, and write a draft page in one designated output directory. It can reject everything else. Those are proposed requirements for this example, not results from a deployed system.
Subject grouping is a different question. A book about the history of computing might reasonably belong under technology or history. If the reader wants help exploring such choices, a model can suggest groups for review. If the reader instead supplies an explicit category for every record, that part can become a lookup too.
The useful boundary runs through the workflow. It need not put the whole thing on one side.
A run tells you what happened, not what should happen
A successful run might show the assistant merging two records with the same title. That does not establish a deduplication rule. Different editions can share a title. A trace also cannot tell you whether a missing author is acceptable, whether input order matters, or whether the reviewer quietly corrected an error.
You have to decide those things. For this experiment, duplicate identifiers might be an error while duplicate titles remain separate records. A missing author might be displayed as unknown. The specification should say so before the implementation is judged against it.
Then test the contract, including cases that the attractive demonstration never encountered: an unsupported version, a duplicate identifier, an empty title, and text that would become executable markup if inserted unescaped. Keep additional fixtures aside when developing the converter so the final check is not just a replay of its development examples. Compare the rendered output with independently prepared expectations, not only with what the assistant produced.
Passing those tests gives evidence about those cases. It does not prove correctness for every input. A program can consistently apply a mistaken requirement, and a test suite can faithfully encode the same mistake.
This is familiar software engineering. The AI-assisted beginning does not exempt the result from it.
Keep the decisions that still need reasoning
Anthropic's Building effective agents draws a useful distinction: workflows orchestrate LLMs and tools through predefined code paths, while agents let LLMs direct their processes and tool use. It recommends starting with simpler solutions. That is architectural guidance from a vendor, not experimental proof that one design will outperform another on your task.
A predefined workflow can still contain model calls. Removing model-directed orchestration is different from removing inference. Our reading-list converter might be entirely conventional code, while its optional grouping assistant still uses a model. Both can sit behind the same review step.
Nor is the problem that all agents forget everything. Systems can retain context, retrieve prior work, and use memory. But remembering a previous solution and executing an explicit rule are different ways to reuse work. If the desired behaviour is fully specified, first compare the model-assisted route with a direct implementation. If the inputs remain ambiguous or the desired output depends on interpretation, the model may still be useful.
Sometimes the rule is simple enough to write before running an agent at all. You do not need an AI discovery phase to justify a parser.
The model call can go. The obligations stay.
For the local experiment, the converter should have access only to the fixture inputs and its draft-output directory. A malformed record should stop generation with a useful explanation, rather than trigger a guess. A person should still check the draft before copying it anywhere public. Replacing reasoning with code does not grant permission to publish.
Retries need their own design. Repeating a run should not append the same books again or destroy an unrelated file. For this example, generating a new draft and replacing only the designated prior draft is a more sensible contract than accumulating output. Even then, write failures and interrupted runs need tests.
The Amazon Builders' Library article on idempotent APIs explains the general issue: retrying a request should not create additional side effects. It describes caller-provided request identifiers and the need to coordinate recording the identifier with mutations atomically. An identifier alone is not protection. That API design discussion does not certify our file converter, but it makes clear why “run the same steps again” is an incomplete recovery plan.
And the world can change. A new export version might introduce different meanings for familiar columns. The converter should reject the unknown version until its contract is revised. Silently handing rejected inputs to a more capable agent would change both the behaviour and the risk. Any assisted fallback needs its own input limits, permissions, and approval conditions. A stop with a human handoff is a valid outcome.
Google's SRE chapter The Evolution of Automation at Google warns that automation can create problems as well as solve them, and discusses automation falling out of step with the systems it manages. These are engineering cautions, not evidence of a guaranteed safety improvement from removing a model. Ordinary software can repeat the wrong operation very efficiently.
What would make the change worthwhile?
The conventional converter needs no model inference for its own work. That says nothing about its total cost.
Count model tokens separately from compute and storage. Then account for implementing and maintaining the converter, verifying its output, obtaining approvals, monitoring failures, and revising it when formats or expectations change. The optional grouping assistant still has inference costs. A human reviewer still has work to do.
Compare both routes on the same permitted task and representative inputs. Measure correct output and rejected cases alongside elapsed time and human effort. Include recovery work. If the program looks cheaper only because the comparison quietly removed review or ignored maintenance, it has answered a different question.
There is no break-even result in this hypothetical example. A frequently used, stable conversion may repay the effort; a rarely used format that keeps changing may not. The decision needs evidence from the actual task, not a promised percentage reduction in tokens.
I would start with the fixed column mapping, not the ambiguous subject groupings. Write down what it accepts, what it produces, and when it stops. Then test whether moving that one piece into code improves the work. Leave the next decision alone until you can explain it just as clearly.
Sources and scope
The reading-list experiment is fictional and nonproduction. No deployment results or employer-specific designs are reported. The proposal above is my synthesis; these sources support the stated distinctions and cautions, not a validated method or a claim of novelty.
- Anthropic, Building effective agents (originally December 2024; living page). See “What are agents?” and “When (and when not) to use agents.” Definitions and design guidance, not a comparative trial.
- Amazon Builders' Library, Making retries safe with idempotent APIs. See “Reducing client complexity with idempotent API design” for request identifiers and atomic handling of side effects.
- Niall Murphy with John Looney and Michael Kacirek, The Evolution of Automation at Google, Site Reliability Engineering, Chapter 7 (2016). See the opening caution and “The Use Cases for Automation” discussion of changing systems and fragile automation.
- Arun Malik, Progressive Crystallization, arXiv v1, Section III. Conceptual taxonomy only. This series does not reproduce or independently validate the paper's operational results.
Public sources checked September 27, 2026.