Skip to content
AI.info

Future Horizons

AI and the Future of Work: The 300-Million-Job Question

Goldman's 300 million, the WEF's net 78 million and Acemoglu's 1.1% answer narrower questions than people ask of them. What the randomized trials and the 2026 payroll data actually measured.

AI and the Future of Work: The 300-Million-Job Question

Gabriele Masetti ·

The Number That Launched a Thousand Headlines

In March 2023, Goldman Sachs economists Joseph Briggs and Devesh Kodnani published a report, "The Potentially Large Effects of Artificial Intelligence on Economic Growth," that gave the AI-and-jobs debate its defining statistic: generative AI could expose the equivalent of 300 million full-time jobs worldwide to some degree of automation. The number spread fast, and in the retelling it usually hardened into something the report never actually said — that 300 million people would lose their jobs.

The underlying analysis was narrower and more careful than the headlines suggested. Briggs and Kodnani estimated that about 18% of work globally could be computerized by current generative AI capabilities, with roughly two-thirds of U.S. occupations exposed to some degree of automation and, among those exposed occupations, somewhere between a quarter and half of workload potentially replaceable.

Exposure was heavily concentrated by occupation: office and administrative support came in at 46% exposed, legal work at 44%, architecture and engineering at 37%, while construction, installation and maintenance, and building-and-grounds work sat in the single digits. Critically, "exposed" meant tasks within a job could be automated or augmented, not that the job itself would disappear.

The same report projected that widespread adoption could lift global GDP by about 7%, or nearly $7 trillion, over a ten-year horizon — a productivity story sitting right next to the displacement story, in the same document, largely unread by the people who cited only the 300 million figure.

The geography of exposure was equally uneven, though this detail rarely made the headlines either. Goldman's model put Hong Kong, Israel, Japan, Sweden, and the United States among the countries most exposed to AI-driven automation, reflecting economies with a high concentration of white-collar, task-routine occupations. Workers in mainland China, Nigeria, Vietnam, Kenya, and India sat at the low end of exposure, largely because their labor markets skew toward occupations — agriculture, informal services, physical trades — that generative AI, as distinct from robotics, does not yet touch.

The 300 million figure is therefore not a forecast of where automation will hit hardest in human terms, but of where task composition overlaps most with what language models are currently good at.

What the World Economic Forum Actually Says for 2025

The other number that circulates constantly is from the World Economic Forum's Future of Jobs Report. It is easy to get this one wrong, because the WEF has published overlapping estimates for different target years, and reporting on AI's labor effects has sometimes conflated them. The 2020 edition, looking ahead to 2025, projected 85 million jobs displaced against 97 million newly created. That pairing is six years old and describes a different forecast horizon entirely.

The report that still sets the terms is the Future of Jobs Report 2025, published by the WEF in January 2025 and — the series being biennial — the standing edition eighteen months later. It is built on a survey of over 1,000 employers representing more than 14 million workers across 55 economies. Its horizon is 2030, not 2025, and its figures are 170 million new roles created against 92 million roles displaced, for a net gain of 78 million jobs — job disruption affecting the equivalent of 22% of today's total employment.

The fastest-growing roles cluster around technology, data, and AI specialists, but also include less obviously "AI-adjacent" work such as delivery drivers, care roles, educators, and farmworkers, reflecting demographic and green-transition trends layered on top of automation. The report also flags a skills problem distinct from headcount: employers expect 39% of workers' current skill sets to be transformed or rendered outdated within five years, which is arguably the more consequential number in the whole document, since it applies to people who keep their jobs, not just those who lose them.

The World Economic Forum's Future of Jobs Report 2025 projects both large-scale displacement and larger-scale job creation through 2030.

The Skeptic's Rebuttal: Acemoglu and the Macroeconomics of AI

Against these output-and-employment forecasts stands a pointed academic counterweight: MIT economist and Nobel laureate Daron Acemoglu's paper "The Simple Macroeconomics of AI," circulated by NBER and MIT's economics department in 2024 and later published in Economic Policy. Acemoglu starts from task-level exposure estimates similar in spirit to Goldman's, but applies a more conservative lens on which exposed tasks are actually profitable to automate given current accuracy, integration costs, and human oversight requirements.

His conclusion: only around 5% of tasks are likely to be automated cost-effectively within the coming decade, translating into a "nontrivial, but modest" boost to U.S. GDP of roughly 1.1% to 1.6% cumulative over ten years, alongside total factor productivity gains he estimates at no more than about 0.7% over the same period — well below the multi-percent annual gains implied by more bullish industry projections.

Acemoglu goes further, arguing that even these modest estimates may be optimistic, because the early productivity evidence comes from easy-to-learn, well-defined tasks, while much of the remaining exposed work involves context-dependent judgment that current models handle far less reliably. Goldman Sachs published a direct rebuttal defending its more optimistic framework, and the exchange is less a settled dispute than a live fault line in how economists model AI diffusion: how fast will accuracy improve, how quickly will firms restructure workflows around a new tool, and how much of "exposure" ever becomes "adoption."

Acemoglu's framework does not deny that AI will automate a great deal eventually; it argues that "eventually" is doing a lot of work in headline-grabbing forecasts, and that the gap between what a model can do in a demo and what a firm will trust it to do unsupervised, at scale, on the full messiness of real business tasks, closes far more slowly than exposure percentages imply. That gap is exactly what the field experiments described next were designed to measure directly, rather than infer from occupational classifications.

What the Randomized Trials Actually Show

While macroeconomists argue over aggregate projections, field experiments have measured what AI tools actually do to output when real workers use them on real tasks. The set is no longer small, and it no longer points one way: the strongest single result in it is negative.

The most cited is "Navigating the Jagged Technological Frontier," a randomized controlled trial run by Harvard Business School, MIT, Warwick, and Boston Consulting Group researchers (Dell'Acqua and coauthors) with 758 BCG consultants, published in September 2023. Consultants given GPT-4 access completed 12.2% more tasks, 25.1% faster, and produced work judged more than 40% higher in quality than a control group — but only for tasks that fell inside what the authors called the model's "jagged frontier."

For a task explicitly designed to sit outside GPT-4's reliable capability, AI-assisted consultants were 19 percentage points less likely to produce a correct answer, and the effect was often invisible to the consultants themselves, who could not tell which side of the frontier they were on. The productivity gains were also uneven by skill level: below-average performers improved roughly 43%, while top performers gained about 17% — AI narrowing the gap between strong and weak performers, not just amplifying everyone equally.

Study Setting Productivity result
BCG consultants (Dell'Acqua et al., 2023) Consulting, GPT-4 +12.2% tasks, 25.1% faster, >40% higher quality
GitHub Copilot RCT (Peng et al., 2023) Software engineering, single task ~55.8% faster
Multi-company RCT (Microsoft/Accenture/Fortune 100) Software engineering, real workflows ~26% increase in output
Customer support (Brynjolfsson, Li, Raymond) 5,179 agents, Fortune 500 firm +14% avg; +34% for novice agents
METR RCT (Becker, Rush, Barnes, Rein, 2025) 16 experienced open-source developers, 246 tasks 19% SLOWER with AI, while believing they were 20% faster

A parallel pattern shows up in software engineering. A 2023 controlled experiment on GitHub Copilot (Peng, Kalliamvakou, Cihon, and Demirer, arXiv:2302.06590) found developers with Copilot access completed a defined coding task about 55% faster than a control group, with a wide confidence interval reflecting task-specific variance.

A separate, larger set of three randomized trials across Microsoft, Accenture, and an anonymous Fortune 100 manufacturer, covering thousands of developers and led by researchers from Microsoft, MIT, Princeton, and Wharton, measured real workflow metrics — pull requests, commits, and builds — rather than a single benchmark task, and found a smaller but still substantial roughly 26% increase in output, with less-experienced and newer developers benefiting disproportionately.

Then the software evidence stopped pointing one way. In July 2025 METR, a research nonprofit, published a randomized controlled trial in which 16 experienced open-source developers completed 246 tasks in mature repositories they had worked on for an average of five years, each task randomly assigned to allow or forbid the AI tools available at the February-to-June 2025 frontier. Before starting, the developers forecast that AI would cut their completion time by 24%. Afterwards, they estimated it had cut it by 20%. Measured, allowing AI increased completion time by 19%.

The sample is 16 people, and the authors say plainly they are not claiming AI fails to speed up most developers. Taken narrowly, the gap inside the result is still the most useful number in the literature: a 39-point spread between the speedup developers believed they were getting and the slowdown a stopwatch recorded. The people using these tools are not reliable instruments for measuring them, which is worth remembering whenever an adoption survey reports self-assessed gains.

The customer-support evidence tells the same distributional story most cleanly. Erik Brynjolfsson, Danielle Li, and Lindsey Raymond studied the staggered rollout of a generative AI assistant to 5,179 customer support agents at a Fortune 500 software firm (NBER Working Paper 31161, later published in the Quarterly Journal of Economics).

Average productivity, measured as issues resolved per hour, rose 14% — but that average masked a 34% gain for novice and lower-skilled agents against essentially no measurable gain for the most experienced, highest-skilled agents, whose conversation quality even dipped slightly with AI assistance. The AI tool appeared to encode and distribute the conversational patterns of the company's best agents, letting newcomers with about two months of tenure perform on par with untreated agents who had six months or more of experience.

The Pattern Underneath the Numbers

Read together, these experiments describe something more specific and more useful than either "AI takes your job" or "AI barely matters." They describe a leveling effect: generative AI tools compress the performance gap between novice and expert workers within tasks the AI can reliably handle, while contributing little or even something negative once a task drifts outside that reliable zone. That is a very different distributional story from the occupation-level exposure percentages in the Goldman Sachs and WEF reports, which say which jobs contain automatable tasks but nothing about who within those jobs gains or loses.

That reading also explains why the Goldman Sachs 300 million figure and Acemoglu's modest 1.1-1.6% GDP estimate are not actually as contradictory as they look side by side. Exposure is a ceiling on what could eventually be automated under ideal conditions; realized productivity depends on integration costs, error tolerance, and how much of a given task genuinely sits inside the "jagged frontier" today.

Both can be true: a very large share of tasks touched by generative AI, and a comparatively modest near-term macroeconomic effect, because task-level exposure converts into economy-wide output only as fast as firms redesign workflows, retrain staff, and trust the tool's outputs enough to remove the human check.

The Entry-Level Evidence That Arrived After the Trials

Every randomized trial above measures output per worker. None of them says anything about how many workers a firm hires, and by 2026 that second question had data of its own.

Erik Brynjolfsson, Bharat Chandar and Ruyu Chen of Stanford's Digital Economy Lab used ADP payroll records covering millions of American workers through June 2026 to look for AI's footprint in employment counts, not productivity. Their paper, revised in August 2026 and called "Canaries in the Coal Mine," finds employment of 22-to-25-year-olds in the most AI-exposed occupations running about 19% below where the comparison with less-exposed peers puts it, while experienced workers in the same occupations show no equivalent gap.

Three features make that number hard to wave away. The gap has widened steadily since August 2025, not as a one-off; it operates through reduced hiring rather than increased separations, which is why it does not surface as a layoff story; and it concentrates in roles where AI substitutes for human work, while employment holds flat or rises where AI complements it.

The authors call these early descriptive indicators, canaries rather than causal estimates, and the same data shows no widespread economy-wide displacement. Both halves matter: aggregate employment is not collapsing, and the entry door is narrower. That reframes the levelling effect the trials found. A tool that lets a two-month newcomer match a six-month veteran is a training accelerator when a worker holds it and an argument for hiring fewer newcomers when an employer does.

Why Reskilling Is the Actual Bottleneck

That integration lag is where the WEF's 39% figure earns its weight. If nearly two in five current job-relevant skills are expected to be transformed or made obsolete by 2030, the binding constraint is not technological. It is how fast training systems, employer investment and individual workers can move people from the skills AI is displacing into the ones the same disruption is creating.

The customer-support and consulting studies both suggest that AI assistance itself can function as an accelerated, embedded training mechanism, compressing a newcomer's learning curve to match a veteran's. It is also the mechanism for which employers, universities, and governments have the least developed playbook.

If 92 million existing roles are expected to be displaced by 2030 even as 170 million new ones appear, the workers most at risk are not simply those in the highest-exposure occupations. They are the ones whose employers lack the training infrastructure to move them onto the AI-augmented side of the ledger before their old tasks are automated out from under them.

But an informal tutor is not a substitute for deliberate investment, and it only reaches workers who got hired. The randomized trials measured augmentation, not displacement. The payroll evidence that has since filled part of that gap counts jobs that were never opened, not workers who were let go. Neither follows an individual through a transition, which remains the thing nobody is measuring well and the thing the 300-million and 78-million-net figures both paper over.

The honest reading of the evidence, then, is not that the 300-million-job number was wrong or that Acemoglu's skepticism refutes it — it is that both numbers answer questions narrower than the ones people ask of them. Task exposure is not job loss. A modest macro GDP estimate is not evidence that nothing is changing inside individual workplaces.

The randomized trials, measuring what happens when specific workers use these tools on specific tasks, point toward redistribution of who performs well at a job rather than replacement of who holds it. The payroll data adds the part the trials could not see: at one end of the labour market, the first job after university, that redistribution has already turned into a hiring decision. Four years into the generative AI era, the sharper policy question is not whether to slow adoption to protect jobs in the aggregate, which is not where the evidence puts the damage, but what replaces the on-the-job training an entry-level role used to provide once the entry door narrows.

Explore

More articles