The Amplification Thesis

Arun MalikSeriesEssayAI Strategy

Series date follows the editorial schedule. First published ; updated .

Making someone better at a task does not mean you have built the best way to do that task. An AI assistant can improve a person's work while the same model, working alone, still does better. That distinction changes the case for amplification.

I want AI to make people more capable. But "human plus AI wins" is too easy a slogan, and the evidence does not support it as a general rule. Amplification is a design objective, not a guaranteed result. Give the person a useful contribution to make, then test the whole workflow against the alternatives.

Augmentation Human + AI versus human alone The stronger test Human + AI versus the better standalone option Passing the first test does not mean passing the second.
Two questions, different baselines. A conceptual diagram, not an indexed score chart.
"Where can we cut?" a head is a cost to remove
A cost question is only a starting point
automate stay manual two options, and that's the whole menu
A collaboration is a third option to test
Different baselines, different pooled results
Creative performance: 21 observations, human baseline
human context, judgment AI generate, search test different knowledge, different strengths
Different strengths are a possibility to test
1960Licklidersymbiosis 1962Engelbartaugment intellect todaytest theworkflow a design tradition, not a result
Not a new idea
full AI full human ask about the task preferences vary across occupations
Conceptual spectrum, not a survey score
Designer: define constraints AI: propose variants Designer: test with users
Hypothetical example, not production evidence
Compare feasible workflowssame task, same quality bar Count the whole jobincluding review and rework Inspect the errorsan average can hide costly failures
A proposed evaluation, not a claimed result
Explain the choice? Spot a flawed suggestion? Work without the tool? output quality is not a skills test
Check whether capability lasts
valuebetter work costthe whole workflow keep the design only if it earns its place
No automatic moat
route automationwhen it meets the bar collaborationwhen the role helps or keep the human-only workflow
Not everything should be amplified
What can AI amplify? Does this workflow help?
Make the objective answerable
The starting point

"Where can we cut headcount?"

That is a legitimate budget question. It is a poor specification for a product. Removing a task from someone's day might save money, or it might free them to do work that was previously out of reach. Those are different outcomes, and a headcount target cannot tell you which one happened.

Start with the work you want done better. A faster first draft is useful only if getting to an acceptable final result also gets easier.

The false choice

Automate, or stay manual

There is a third option: change how the work is done. A person might frame the problem, use AI to explore possibilities, and return to information the model cannot supply. That is a candidate for amplification. It still has to earn its place.

Researchers distinguish human augmentation, beating the human-only baseline, from beating the stronger of the human-only and AI-only baselines. The latter is the harder test. Passing one tells you nothing certain about passing the other.

The broader evidence

Better than a person. Worse than the best option.

Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone analyzed 106 experiments reported in 74 papers, with 370 effect sizes. On average, human-AI combinations improved on humans alone. But they performed worse than the better standalone option. Both results came from the same review.

The plotted values are standardized effects, not percentage improvements. The papers covered January 2020 through June 2023, so this is evidence about those systems and experiments, not a verdict on every current model.

Vaccaro et al. (2024), pooled task performance. Positive values favor the combination.
ComparisonHedges' g95% CI
Versus human alone0.640.53 to 0.74
Versus better standalone-0.23-0.39 to -0.07

Decision tasks showed losses against the stronger baseline. Content-creation tasks were more promising, but their pooled advantage over that baseline was not statistically distinguishable from zero. "Creative work always wins" would overstate this result too.

The evidence

A narrower case for creative work

The May 2025 version of Niklas Holzner, Sebastian Maier, and Stefan Feuerriegel's meta-analysis included 28 studies and 8,214 participants overall. Its human-with-AI versus human-without-AI comparison used 21 observations and found a small positive effect on creative performance.

Holzner et al. (May 2025, section 3.3). Creative performance with AI versus without AI, not versus AI alone.
ObservationsHedges' g95% CI
210.2730.018 to 0.528

Results varied substantially; removing some individual studies made the confidence interval cross zero. A separate comparison found no statistically significant difference between AI alone and humans alone. That does not establish equivalence, and stitching the comparisons together does not establish that the combination beats both.

The review also found lower idea diversity with AI assistance, based on six observations from four studies. Better-rated individual outputs and a narrower pool of ideas can coexist. If you want exploration, measure variety as well as the quality of the selected answer.

The mechanism

Give the person something useful to contribute

Complementarity is a design hypothesis: the person brings information or a capability that improves what the model can do, or vice versa. A designer may know why customers abandoned a feature. A model may produce variants the designer had not considered.

Neither contribution is guaranteed. People can miss errors; a fluent model can invent context. Asking someone to approve an answer does not, by itself, add expertise to it. The useful question is specific: what changes because this person is involved?

The long arc

An old ambition, not an old proof

In 1960, J.C.R. Licklider proposed a partnership in which people set goals and evaluate results while computers do routinizable work. Douglas Engelbart's 1962 framework treated improving human problem-solving as a system-design problem, including methods and tools.

That is a useful tradition to work in. It is not six decades of experimental proof that any human-AI pairing will win.

The preference

Ask the people doing the work

Stanford's WORKBank study collected responses from 1,500 workers across 104 occupations. It found varying preferences for human involvement, alongside interest in automating low-value, repetitive tasks.

Preferences are design input, not evidence that a preferred arrangement performs better. Ask which parts of the work people want help with, and which knowledge they need to keep using. Then test the proposed division of work with them.

A hypothetical example

A designer working on a sign-up form

Imagine a designer trying to improve a public library's sign-up form. This is a hypothetical example, not a production case or a disguised set of operational results. The designer brings observations from user interviews and sets constraints: accessibility, plain language, and the information the library actually needs.

AI proposes alternative wording and layouts. The designer rejects options that misunderstand the service, tests the remaining options with users, and revises the form. Generating variants is only one part of the job. Discovering that a question should not be on the form at all may matter more.

The evaluation

Count the whole job

Compare the designer's existing workflow with the assisted version on comparable briefs. Where safe and feasible, include an AI-only output given the same documented constraints. Decide the quality bar before looking at the results. Assess usability and accessibility, and count time spent prompting, checking, and reworking.

Use more than a favorite demo. Rotate task assignments to reduce practice effects, have outputs evaluated without revealing how they were made where possible, and inspect costly errors separately from average quality. This is a proposed evaluation, not a claim that the hypothetical workflow succeeds.

The longer term

Better output is not proof of learning

A good form does not tell you whether the designer learned anything. If skill growth is part of the investment case, test that separately later. Can the person explain the design choices, recognize a flawed suggestion, and handle a comparable brief without assistance?

Those are checks I would build into the rollout, not findings established by the short-task studies above. Keep opportunities for independent practice if that capability matters. Do not book "compounding expertise" as a return before measuring it.

The strategy

The business case has to survive the costs

Domain knowledge and a well-tested workflow can be valuable. Calling them an uncopyable moat goes too far. Collaboration costs time, and automation can also improve as systems and processes improve. Neither strategy comes with a predetermined growth curve.

A useful investment case says what got better, what it cost, and which alternative it beat. If the benefit disappears once review and rework are included, redesign the workflow or stop using it for that task.

The nuance

Some things should be automated

If automation meets the required quality and risk constraints at lower cost, use it. In the library example, checking whether required fields are present may need ordinary code rather than a generative model or a reviewer.

And if an AI-only option scores well but is not legally or operationally feasible, say so. The best benchmark score and the best deployable workflow are different questions. Human involvement may be required for accountability even when it does not improve the measured score.

The thesis

Design for amplification. Make it prove itself.

I would start by looking for work people could do better with AI, rather than assuming the only return is a smaller team. That opens possibilities a replacement-only plan misses. It does not excuse a weak comparison.

Keep the person involved for a reason you can name. Keep the workflow because the results justify it.

The task what should improve? The baseline better than what? The contribution what does each add? The evaluation does the design help? Measure the whole workflow, including review and rework.
A design objective becomes useful when you can test it.

Where this series goes next

The next essay, The Centaur Advantage, asks when a division of work actually beats the stronger standalone option. That claim needs evidence for a particular task and workflow.

Charts show standardized mean differences (Hedges' g), not percentages or absolute scores. Narrow bars show 95% confidence intervals for pooled effects, not the range of outcomes a new team should expect. The library example is hypothetical; no private operational data is used.