AI compute explained: what a frontier training run costs in chips, power and money
One frontier training run costs hundreds of millions of dollars by outside estimates, inside data centers that cost tens of billions. Here is what FLOP, GPUs and gigawatts mean, and how to read the numbers.
Published

AI compute is the processing capacity used to train and run models, and the honest answer to "what does a frontier training run cost" is: a few hundred million dollars in chips, power and staff for one final run by the best public estimate, inside a build-out that costs tens of billions. Those are different numbers, and mixing them up is the most common error in coverage. This explainer separates them, dates them and names who estimated each.
The units, in plain terms
FLOP means floating-point operations, the basic arithmetic a chip performs. Total training compute is the number of those operations spent on one run, which is why a model's size is quoted in FLOP rather than hours. Epoch AI, the research group that keeps the main public database of this, tracks over 3,600 models and grades its own estimates: "confident" ones are good to within 3x, "likely" within 10x, and "speculative" within 30x. Read any single FLOP figure with that spread in mind. Our piece on why 10^25 and 10^26 FLOP thresholds disagree covers how regulators use these numbers.
GPUs are the accelerator chips doing the arithmetic, and the usual accounting unit is the GPU-hour: one chip running for one hour. A run on 100,000 GPUs for 100 days is about 240 million GPU-hours. Stanford's 2026 AI Index, published April 13, 2026, counts global AI compute capacity at 17.1 million H100-equivalents, growing 3.3x a year since 2022, with Nvidia supplying over 60% of it.
Gigawatts (GW) measure power draw at any instant, while gigawatt-hours measure energy over time. The same Index puts AI data center power capacity at 29.6 GW, which it compares to New York state at peak demand.
What one run costs: the Epoch estimate
The best-known worked example is Epoch's estimate for xAI's Grok 4, published September 12, 2025. Epoch put the run at about 246 million H100-hours, 310 GWh of electricity and a median cost of $490 million. It reached the cost two ways: renting H100s at roughly $1.90 to $2.20 per hour, and depreciating owned hardware plus electricity at $0.08 to $0.20 per kWh.
“a stylized model, not an estimate for any specific facility”
Notice what that figure is. It is a cost to produce one final training run, with hardware spread over its useful life. It is not the cash xAI spent to build the cluster, and it leaves out failed runs, experiments and the models that came before.
Epoch's older study of training costs, by Ben Cottier and colleagues, published June 3, 2024, found the amortized hardware and energy cost of final runs for frontier models growing 2.4x a year since 2016 (95% interval 2.0x to 3.1x). It projected that the largest runs would pass $1 billion by 2027. In its breakdown, hardware was 47 to 67% of the cost, research staff 29 to 49%, and energy only 2 to 6%. Power is a small line in a run's bill and a large one in the site's planning.

Run cost versus capex
Capex is the money spent up front on chips, buildings and networks. Run cost is the slice of that capital, plus electricity and people, attributable to one training job. A lab can spend $38 billion on a site and still have a training run that "costs" half a billion on paper, because the site trains many models and also serves customers.
Epoch's May 14, 2026 data insight models a stylized one-gigawatt data center built on Nvidia GB200 NVL72 systems: $38 billion up front, $0.9 billion a year to operate, and about $8.5 billion a year all-in once the capital is spread over 5-year server and 14-year facility lives. Its own range runs from $7 billion to $12 billion depending on server lifespan. Epoch calls it "a stylized model, not an estimate for any specific facility."
At the corporate level, Epoch's tally of SEC filings shows combined capex at Microsoft, Amazon, Alphabet, Meta and Oracle reaching $140.6 billion in the fourth quarter of 2025, up from about $36.8 billion in the second quarter of 2023. Epoch notes the companies do not disclose how much of it is AI-specific. The run is the visible tip; the filings show the iceberg.
How power sets the pace
Epoch's August 11, 2025 analysis found the largest runs already above 100 MW and projected that single runs in 2030 could draw 4 to 16 GW, with power demand growing about 2.2x a year historically. For perspective, the Grok 4 energy estimate of 310 GWh is about what a 100 MW facility would use in roughly four months at full load, which fits that scale. On the grid side, see why interconnection is the bottleneck.
Reading these numbers skeptically
- The labs do not publish their costs. Every per-run figure here is an outside estimate, from Epoch, not an audited cost from xAI or any lab. Epoch itself flags "significant uncertainty around our point estimate" for Grok 4 because xAI's public statements about GPU-hours were vague.
- Rental prices are a proxy. A lab that owns its chips, or buys at a discount, pays something else.
- Final-run cost understates total R&D. Failed runs, ablations and data work sit outside it.
- Projections are extrapolations. "More than $1 billion by 2027" was a 2024 trend line, and the 2026 data insights show how much depends on whether capex keeps outrunning cash flow.
- Stanford is secondary here. The AI Index compiles estimates, including Epoch's, and itself notes that training compute can only be estimated independently now that labs have stopped disclosing parameters.
Our take
Treat the $490 million figure as a defensible order of magnitude, not a price tag, and treat any headline that sets it beside a $38 billion site as comparing a slice with the whole. The number worth watching is not the cost of one run but whether the capex behind it earns revenue. Until labs publish their own run costs, every figure in this piece is someone's model, and the useful habit is to ask which estimator, which date and which cost definition.
Frequently asked questions
What does AI compute cost for a frontier training run?
Epoch AI estimated xAI's Grok 4 run at about $490 million median (September 2025), using 246 million H100-hours and 310 GWh. That is an amortized cost for one final run, not the cash spent building the cluster, and xAI has not published its own figure.
What is a FLOP in AI training?
A FLOP is one floating-point operation, the basic arithmetic a chip performs. Total training compute is the count of those operations used in a run. Epoch grades its estimates as confident (within 3x), likely (within 10x) or speculative (within 30x).
How much power does a frontier AI training run use?
Epoch's August 2025 analysis found the largest runs already exceeding 100 MW and projected 4 to 16 GW for single runs by 2030. Its Grok 4 estimate was 310 GWh of electricity. Stanford's 2026 AI Index puts total AI data center capacity at 29.6 GW.
What is the difference between capex and the cost of a training run?
Capex is the up-front spend on chips, buildings and networks. Run cost is the share of that capital, plus electricity and staff, attributed to one training job. Epoch models a one-gigawatt data center at $38 billion up front, about $8.5 billion a year all-in.
Is the cost of training AI models rising?
Epoch's 2024 study found amortized hardware and energy cost of final runs growing 2.4x a year since 2016, projecting the largest runs above $1 billion by 2027. That is a trend extrapolation, and labs do not publish their own costs to check it.
Who estimates AI training costs, and can they be trusted?
Epoch AI is the main public estimator, with Stanford's AI Index compiling its data. Both are independent of the labs, but the estimates rest on rental prices and public statements, so treat them as orders of magnitude.
Sources
What each one is, and whose it is.
- 1
How much does it cost to train frontier AI models?, Epoch AI (June 2, 2024)
PaperIndependent of the vendorNot peer reviewed, preprint - 2
Grok 4 training resources, Epoch AI (September 11, 2025)
DatasetIndependent of the vendor - 3
Total cost of ownership of a one-gigawatt AI data center, Epoch AI (May 13, 2026)
DatasetIndependent of the vendor - 4
Hyperscaler capex has quadrupled since GPT-4's release, Epoch AI (February 25, 2026)
DatasetIndependent of the vendor - 5
Power demands of frontier AI training, Epoch AI (August 10, 2025)
PaperIndependent of the vendorNot peer reviewed, preprint - 6
2026 AI Index Report: Research and development, Stanford HAI (April 12, 2026)
OtherIndependent of the vendor