The Amplification Playbook
Series date follows the editorial schedule. First published ; updated .
Pick one workflow. Describe what would count as better work, who would benefit, and what evidence would make you stop. That is a more useful start than an organisation-wide promise that AI will amplify everyone.
The essays in this series do not prove that human-plus-AI is always best. They show why the answer depends on the comparison, the task and the people doing it. The playbook below is my proposed way to turn those distinctions into a trial. It is not a report of a deployment or a validated organisational recipe.
Start with an outcome somebody wants
Consider a hypothetical adult-education centre answering questions about its courses. Staff want to send accurate, useful replies without spending their day reconstructing information scattered across course pages. Learners want to know whether a course fits their needs. Neither group is asking for more generated text.
Write the outcome in those terms: correct answers that help learners decide, within a useful response time and an affordable workload. Ask staff which questions are repetitive, which require conversation, and which are hard because the records disagree. Ask learners what an unhelpful answer costs them.
Set the initial scope narrowly. The assistant may draft answers from approved course information. It may not change enrolments, promise exceptions, approve refunds or invent accessibility provision. Those are proposed limits for this example, not universal rules for education services. Avoid putting personal or sensitive information into a tool unless the organisation has established that the use is permitted.
Name a person responsible for the trial and a route for resolving policy questions. If nobody can answer whether a proposed use is acceptable, a more fluent draft will not fix the missing decision.
Keep a simpler alternative in the comparison
The current process is one baseline. A cleaned-up reference page or a deterministic answer lookup may be another. Compare those with an AI-assisted draft-and-review workflow. Do not let the trial become a contest between an improved AI process and an intentionally neglected manual one.
Where an autonomous route is lawful and appropriate, evaluate it separately. Where it is not, say why it is excluded. The amplification thesis requires a clear baseline, not an unsafe experiment carried out to fill every column.
For the centre, a useful first evaluation could use a de-identified, independently reviewed set of representative questions. Include conflicting information and questions that should be deferred. Separate the cases used to improve the tool from those used to assess it. An assistant that looks good only on the examples used to tune it has not passed a useful test.
Decide where a person can change the result
In the assisted route, the draft should show which source supports the answer and flag missing information. A reviewer needs time, relevant knowledge and the ability to correct or withhold the reply. Adding a person who cannot inspect the evidence does not create useful oversight.
For a simple published start time, an exact lookup may be sufficient. For conflicting course descriptions, the assistant can identify the conflict and send it to the information owner. If neither the person nor the tool has the answer, the workflow should acknowledge that uncertainty rather than reward whichever produces a confident sentence first.
The routing essay explains why this allocation needs testing. A model's self-reported confidence is not a measurement of its advantage over a particular reviewer. Route design should follow checked outcomes and permissions, not a borrowed numerical threshold.
Measure the whole bill
Choose the main outcome and acceptance conditions before looking at results. For this example, check correctness against approved information, unsupported promises, appropriate deferrals and useful response time. Count staff effort spent checking sources, editing drafts, handling escalations and repairing mistakes. Include setup and ongoing maintenance in the cost discussion.
A hypothetical saving that disappears at review
Suppose the current process takes 60 minutes for a fixed batch of questions. The assisted route uses 20 minutes to prepare drafts, 25 minutes to check them and 20 minutes to resolve exceptions: 65 minutes of staff effort.
Drafting may feel much faster, but this batch has not saved labour. All four numbers are invented solely for this arithmetic example. They are not a forecast or a result from an organisation.
Better answers could still justify the extra effort. Faster elapsed delivery might also matter. Record those benefits separately rather than calling the result a labour saving.
A live trial needs an appropriate comparison design. Random assignment of eligible cases can help when feasible, but account for staff learning, case difficulty and shared queues. Record exclusions, unfinished cases and departures from the assigned process. A small trial with no observed serious errors does not establish that serious errors cannot occur.
The coding evidence shows why perceived speed, task selection and measured time can diverge. The creativity evidence adds another warning: if the workflow generates options, assess their range separately from the quality of the strongest one.
Make room for learning and for workers
Ask whether new staff are becoming more capable or simply completing supported tasks. The two can move differently. Use safe practice cases to check later independent performance, with a clear explanation of how results will improve training. Do not insert surprise exams into live work or use a learning trial as covert performance surveillance.
The support-agent evidence gives reasons to examine who benefits rather than reporting only an average. The classroom evidence shows why assisted output is not a learning measure. Neither supplies a universal onboarding timetable for the centre.
Discuss what happens if the trial succeeds. Do staff get time for harder enquiries, shorter queues, training, or a different workload? Who maintains the course information and who receives credit for that work? A productivity result does not settle those decisions. Job-posting trends cannot promise a promotion to the people affected.
Give stopping the same status as expanding
Agree on what pauses the trial: for example, an unauthorised disclosure, repeated unsupported commitments, or a reviewer queue that cannot meet the service's obligations. The responsible owner should be able to withdraw access and return to the previous process. Keep a way for staff and learners to report problems without having to prove that AI caused them.
Expansion should name the scope justified by the evidence. Passing on ordinary timetable questions does not justify answering every policy exception. Reassess when the model, sources, permissions or working population changes. If the simple reference guide performs as well with less burden, choosing it is a successful evaluation result.
The NIST AI RMF Playbook is useful background for organising this work around Govern, Map, Measure and Manage. NIST describes it as voluntary suggested actions, not a checklist to follow in full. AI RMF 1.0 is under revision; the linked guidance does not certify this proposed trial or establish its return on investment.
At the end, write a short decision record: which workflow was tested, the evidence, the remaining uncertainty, the authorised next scope and the conditions for reversing it. If the record cannot explain why the people doing the work are better supported, revise the design before calling it amplification.
Sources and status
This concludes the overview and eight essays. The study-specific evidence is linked in the relevant articles above. NIST AI RMF 1.0 and its companion Playbook provide voluntary risk-management guidance, not empirical validation of this framework. The centre and cost calculation are hypothetical. Earlier draft claims of a proven organisational playbook and private operating results have been removed.