How many AI agents could run on the chips shipped through 2027? Epoch says up to 171 million

Epoch AI counts memory, not GPUs, to estimate how many frontier agents the 2025 to 2027 chip supply could run at once. The answer is tens of millions, and the author warns demand may not keep up.

By Yash Malviya

Published

Close-up of a vintage computer circuit board highlighting chips and connectors
Photo: Nicolas Foster / Pexels

How many AI agents could run on the chips already shipped or on order through 2027? Epoch AI's answer, published October 2, 2026 by Jason Li, is somewhere between 33 million and 171 million running at the same time, if every chip were put to work on frontier-model agents. The more interesting line comes at the end of the report: the buildout may run ahead of the demand to fill it.

The number matters because it turns an abstract compute race into a labor-market comparison. If you believe the phrase "country of geniuses in a datacenter," which Epoch cites from Dario Amodei's essay "Machines of Loving Grace," then someone should be able to count the geniuses. Epoch's count is the best public attempt we have found this run, and it comes with assumptions worth reading.

Why Epoch counts memory instead of chips

The report's key move is to size supply by high-bandwidth memory, the stacked DRAM that sits next to AI accelerators, rather than by GPU units alone. Epoch does not spell out its reasoning in the part of the report we read beyond the method itself, so take the choice as its modeling assumption. Epoch counts HBM3E and HBM4 or HBM4E shipped from 2025 through 2027, converts it to GB300-equivalent units, and adjusts for HBM4 being faster, with a central assumption of 2x and a tested range of 1x to 4x.

Shipment volumes are reconstructed from TrendForce releases and assumed generation shares. TrendForce's August 4, 2026 release projects HBM bit shipments to grow 50 to 60% year on year in 2027, and says that still will not keep pace with demand, with DRAM supply constrained through 2027. That context is why memory is the bottleneck worth counting. Our own coverage of the AI memory shortage and DRAM prices shows the same constraint from the buyer's side.

From memory to agents

For closed models, Epoch infers how many agents a chip can serve from what an agent-hour costs. It uses $30 an hour as an API-equivalent price, a $5 per GB300-hour rental rate and a 5 to 10 times ratio of revenue to cost. Spending data comes from TraceLab agent traces, and the pooled hourly rates were about $18.19 for GPT-5.5, $15.50 for GPT-5.6 Sol, $24.34 for Opus 4.8 and $50.16 for Fable 5.

“These estimates suggest a risk that the compute buildout could run ahead of inference demand.”

Jason Li, Epoch AI, October 2, 2026

For open models, Epoch uses SemiAnalysis's AgentX serving benchmark at 50 and 100 tokens per second per user. AgentX replays recorded Claude Code sessions as a dependency graph of main-agent and subagent requests, and counts live agent clients rather than HTTP requests. SemiAnalysis states that AgentX does not evaluate answer quality, only serving throughput.

The results:

  • Closed frontier agents: a central 20 to 40 million concurrent on shipments through 2026, and 50 to 101 million through 2027. The full range through 2027 is 33 to 171 million.
  • Labor equivalent: running 168 hours a week, about 140 to 720 million full-time employees' working hours.
  • Open models: applying DeepSeek V4 Pro benchmarks gives about 1.9 billion concurrent agents, because a cheaper model fits far more sessions per chip.

The open-model figure shows how soft these numbers are. The same silicon supports roughly ten times more agents if the agents are smaller and cheaper. Which model does the work moves the answer more than any chip shipment does.

Detailed view of an Intel i486 DX2 CPU installed on a vintage motherboard with chips and circuits
Memory modules on a circuit board: high-bandwidth memory is the constraint Epoch counts. Photo: Nicolas Foster / Pexels

The part that should worry the builders

Epoch then asks what it would take to fill the capacity. At 20% effective use, made up of a 40% allocation to agents and 50% utilization, the capacity implies $2.6 to $5.3 trillion a year in API-equivalent spending through 2027. Epoch's own projection of developer revenue is roughly $1 trillion by the end of 2027, and that already assumes fivefold annual growth.

Jason Li puts it plainly: "Even modest use of this capacity would require a massive increase in global demand for AI." Epoch is careful to say these are scenarios, not forecasts. They assume full deployment, and the spending figures are not break-even revenue requirements, since provider margins are assumed, not observed.

The demand side is moving, though. Epoch's October 5 analysis found that the median OpenAI researcher's coding-agent usage rose from under $1 a day in January 2026 to $601 by mid-August, with growth near 1.8 times a month, and Epoch called such growth "probably unsustainable." That is one population of unusually heavy users, and the pattern is real. We covered it in our piece on OpenAI researchers' coding-agent spending. Whether ordinary companies adopt agents at anything like that rate is the unanswered question.

What the number does not say

An agent here means a unit of serving capacity, not a worker who can do a job. The comparison with 140 to 720 million employees counts hours of compute, and says nothing about whether the output is as good as a person's. As Epoch's InnovationEval report showed days later, agents can still fail at hard end-to-end tasks.

Our take

Treat 171 million agents as an upper bound on hardware, not a claim about work. The useful finding is the mismatch: a supply chain building for trillions in spending, against a market that Epoch's own projection puts near $1 trillion. If the demand does not appear, the first signs would be falling prices and idle memory, and the labs with the biggest commitments would feel it first. If it does, memory is the constraint to watch. Either way, the race is now measured less in parameters than in how many agents a lab can afford to keep running.

Frequently asked questions

How many AI agents could run on the chips shipped through 2027?

Epoch AI estimates about 33 to 171 million concurrent frontier-model agents on high-bandwidth memory shipped from 2025 through 2027, with central estimates of 50 to 101 million. That assumes full deployment and allocation, so it is a capacity scenario, not a forecast.

Why does Epoch count HBM instead of GPUs?

Long-running agents are limited by the context they hold, which lives in memory. Epoch counts HBM3E and HBM4 shipped from 2025 through 2027 and converts it to GB300-equivalent units, adjusting for HBM4 being faster.

How many human workers is that equivalent to?

Running 168 hours a week, Epoch says those agents would supply the working hours of roughly 140 to 720 million full-time employees. It counts hours of compute, not the quality of the work done.

Is there enough demand to use all that capacity?

Epoch warns there may not be. Using 20% of capacity implies $2.6 to $5.3 trillion a year in API-equivalent spending, against projected developer revenue of about $1 trillion by end-2027.

What did Epoch assume for open models?

Using SemiAnalysis's AgentX benchmarks with DeepSeek V4 Pro, Epoch gets about 1.9 billion concurrent agents. Smaller, cheaper models fit many more sessions on the same chips.

How fast is HBM supply growing?

TrendForce projected on August 4, 2026 that HBM bit shipments would grow 50 to 60% year on year in 2027, and said that would still lag demand, with DRAM supply constrained through 2027.

Sources

What each one is, and whose it is.

  1. DatasetIndependent of the vendor
  2. Press reportIndependent of the vendor
  3. 3

    AgentX methodology, SemiAnalysis InferenceX (June 21, 2026)

    DocumentationIndependent of the vendor
  4. DatasetIndependent of the vendor