The Great Equalizer
Series date follows the editorial schedule. First published ; updated .
The same assistant can be valuable to a beginner and unhelpful to an expert. That sounds obvious. It gets lost surprisingly quickly when a team reports one average productivity number and uses it to decide everyone's tooling, training, and workload.
The strongest reason to take the equalizer idea seriously is a specific workplace study, not a general promise that AI turns juniors into seniors. It found large differences in who benefited. Those differences are more useful than the slogan.
What the support-agent study measured
Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the staggered introduction of a conversational assistant to customer-support workers. The tool suggested responses during chats. Agents remained responsible for the conversation and could ignore the suggestions.
I use the November 2023 version of their NBER working paper, Generative AI at Work, for the figures here. It covers 5,179 agents and reports a 14% average increase in productivity, measured as customer issues resolved per hour. Its headline estimate for novice and low-skilled workers is 34%, with minimal productivity gains for experienced and highly skilled workers. These are that version's figures, not a blend of numbers from the later journal publication and earlier summaries.
The outcome matters. Resolved issues per hour combines handling time, the number of chats handled, and the share resolved. It is not a measure of general intelligence, total professional expertise, or how much somebody learned. Nor does a 34% increase in that rate mean each conversation took 34% less time.
This was a workplace rollout analysed using differences in adoption timing, not an individually randomised trial. The authors use a difference-in-differences design and examine pre-adoption trends. Interpreting the estimates causally depends on the design's assumptions, including whether the comparison workers provide a credible account of what would have happened without access.
Within that setting, the result is encouraging: assistance improved the measured performance of the people who had more room to improve. It does not establish that beginners in every profession benefit more, or that experts have nothing to gain from a differently designed assistant.
Borrowing a good answer versus learning to produce one
The authors offer evidence consistent with useful practices spreading from stronger agents to others. Lower-skill agents' language became more similar to higher-skill agents' language. Workers also retained some productivity gains during software outages when suggestions were unavailable.
That is more informative than assisted performance alone. But the paper calls the mechanism evidence suggestive, and outages are not a randomised long-term test of independent expertise. We should not turn the finding into the claim that AI is the fastest knowledge-transfer mechanism ever built.
There is a practical distinction here. An assistant might help a new worker locate the relevant policy, phrase a clear response, and avoid a known mistake. The worker could still struggle when the policy is missing, two records conflict, or a customer describes an unfamiliar problem. Excellent performance in the supported task does not answer those questions.
Read the percentage without changing its meaning
Suppose, purely to explain the arithmetic, a baseline were 10 resolved issues per hour. A 34% increase in that rate would be 13.4 issues per hour. Neither number is a reported baseline from the study.
You could not conclude that a trainee became 34% more knowledgeable, that their salary should change by 34%, or that staffing could safely fall by 34%. Those are different quantities with different constraints.
Even the time-per-issue conversion would require assumptions about concurrency, case mix and quality. A rate is a good result to measure. It is a poor substitute for every other result.
A better onboarding experiment
Consider a hypothetical organisation training staff to answer membership questions. Rather than announcing that AI will replace months of training, give the experiment a narrower question: does access help new staff resolve representative questions correctly, and do they retain enough understanding to handle new cases?
Keep experienced staff in the evaluation. Record the time they spend correcting suggestions and helping trainees. If the assistant saves the trainee time but doubles the expert's interruptions, the trainee's dashboard is not the whole result.
Use an independently reviewed set of cases, including exceptions and questions that should be deferred. Compare an ordinary reference guide with AI-assisted access to the same material. Where practical, assign access randomly and keep outcome review blind to the condition. Separate tenure from baseline performance: someone new to the organisation may bring substantial experience from elsewhere.
Then test a fresh set of cases without the assistant, after a suitable interval. Explain to participants that this checks the training design, not whether they can be blamed for using an approved tool. The delay and sample should fit the work. There is no universal number of days or cases that makes an onboarding trial trustworthy.
A useful result might be faster assisted work with unchanged independent performance. Another might be improved learning but extra review time. Both deserve more precise names than “equalization.” If less-supported workers fall behind, investigate access, language, training and case assignment before describing the difference as ability.
Who owns the improvement?
The support-agent paper explicitly does not determine aggregate employment or wage effects. A firm could respond to improved novice productivity in several ways, including changing hiring or redesigning roles. The study cannot choose that response for it.
That matters for experts too. If their examples help train or maintain a tool, writing and checking those examples is work. Calling it knowledge sharing does not make the time free. A fair implementation would make that contribution visible and discuss how the gains affect workload and progression.
I would therefore ask a manager for two results: the distribution of assisted outcomes, and what workers can do after the assistance is removed. A narrower performance gap is worth having. It is not yet proof of a narrower opportunity gap.
Source and scope
Brynjolfsson, Li and Raymond, Generative AI at Work, NBER working paper 31161, April 2023, revised November 2023. The linked PDF supplies the sample and estimates used here, along with the staggered-rollout design and learning analysis. Earlier draft claims about engineering onboarding, invented skill-quartile estimates and universal human-only capabilities are not retained. The membership workflow is hypothetical.