Models & capabilityIndependently verified

Can AI automate real jobs? Epoch tested six frontier models on its own work

Epoch AI gave six agents 11 real tasks from its own research shop and published the results on October 8, 2026. Reliable on well-defined work, weak on judgment, and a warning about what benchmark scores hide.

By Himanshu Sakre

Published

Bearded man with eyeglasses working on a laptop in a minimalist office setting
Photo: https://kaboompics.com/ / Pexels

Most AI job claims rest on benchmarks. Epoch AI tried something plainer: it gave six frontier agents real tasks from its own work and checked what came back. The answer, published October 8, 2026 by Kelly Hong and Greg Burnham, is that current models cannot yet replace workers, "at least not at Epoch." The details are more useful than the headline.

How the test worked

Epoch paired six models with their own coding harnesses and top reasoning settings: GPT-6 Astra in Codex, Claude Fable 5.1 in Claude Code, Grok 4.6 in Grok Build, Gemini 3.8 Flash in Antigravity, Kimi K3 in Kimi Code, and Qwen 3.8 Max in Qwen Code.

The suite had 11 tasks in five categories: graphic design (3 tasks), data insight generation (3), data explorer generation (2), AI data center research (2) and research design (1). Each model got full permissions in its own workspace and Google account, plus tools such as Figma and reference material, "as much context as we would to a new hire."

Rubrics were built with Epoch staff and mixed objective checks, such as word limits, with subjective ones, such as clear titles. One human grader reviewed each output, and each model ran once per task. The authors say they put more weight on qualitative observations than on the scores.

What the models did well

Fable 5.1 and GPT-6 Astra were broadly tied in the lead. Their strength was reliability on well-defined work such as coding and computational analysis.

“When an AI agent fails to improve with practice, has it failed to collect useful evidence, or failed to understand evidence it already has?”

Research question proposed by GPT-6 Astra, quoted in Epoch AI, Can AI automate Epoch?, October 8, 2026

Computer use improved noticeably. Earlier models failed to port an article to Substack. The newer ones succeeded. And design is already moving: Epoch removed a restyling task from the suite because Fable 5.1 produced output at Epoch quality, and the team now largely automates that step. That is a concrete, in-house example of a task leaving the human column.

An office setting with a person analyzing charts and graphs at a desk with a laptop and smartphone
A research analyst reviewing charts, the kind of well-defined task agents handled best. Photo: Yan Krukau / Pexels

Where they fell short

The failures cluster around judgment, and they are specific.

  • Implicit conventions. For a model performance chart, Fable 5.1 produced a far more information-heavy graphic than the simple chart an Epoch designer made, and used "A" and "B" labels as one-off markers instead of the consistent system in the reference example.
  • Flawed experiments treated as findings. GPT-6 Astra's pilot gave models a 4096-token output limit, which cut off 61 of 280 responses before they named an input. After raising the budget and seeing improvement, it called the result sensitivity to the acquisition budget. The authors call that misleading, since the truncated cases produced no answers at all.
  • Vague designs. Astra proposed a promising research question: "When an AI agent fails to improve with practice, has it failed to collect useful evidence, or failed to understand evidence it already has?" Its experiment never explained what "matched externally selected experience" meant or how its reference bounds would be used.
  • Verbosity. Fable 5.1 wrote dense executive summaries and wordy captions, against Epoch's principle of clear, minimal communication.
  • Convergence. Fable 5.1, Gemini 3.8 Flash and Grok 4.6 all chose the same polling topic for a data insight, and Kimi K3 and Qwen 3.8 Max wrote similar self-evaluation proposals.

The open-weight gap, and what it says about benchmarks

Open-weight models lagged further behind, even on well-defined tasks. Kimi K3's insight about arXiv rested on a filtering error: it analyzed all arXiv papers rather than only physics papers and never checked its filter. Its diagram of wealth redistribution did not follow the described flow of returns from capital assets.

Yet Kimi K3 scores 158 on the Epoch Capabilities Index, roughly tied with Grok 4.6. The authors argue that this shows existing benchmarks miss real weaknesses. A score that ties two models can hide a large difference in whether you would trust the output unsupervised. For the broader problem, see our explainer on what AI coding benchmarks actually measure.

Three October readings on AI doing AI work

This test lands alongside two other recent pieces of evidence, and they point the same direction. The State of AI Report 2026 cites Anthropic's own internal index, saying Claude led about 26 percent of its measured model R&D work in August, up from under 1 percent in February, with researchers still setting tasks and supervising. And Epoch's InnovationEval asked agents to devise a real post-training improvement and found, in its own words, early evidence that the answer is no. We covered it in Epoch's InnovationEval.

Put side by side: models can do a growing share of measured, well-specified work, and they still struggle to originate good ideas or to catch their own mistakes. Both are consistent with the Epoch conclusion that closed-weight frontier models are reliable on defined tasks and consistently fall short on judgment.

The limits of the test

Epoch lists its own caveats, and they are real. The sample is small, with few tasks and one run per model per task, so scores are noisy. Grading is subjective, and the authors treat the numbers as less informative than the qualitative findings. The tasks reflect Epoch's work and standards, not all work. Some gaps may reflect the agent harness rather than the model. The suite differs from GDPval, AutomationBench, the Remote Labor Index and CRUX by requiring models to gather context with tools in messy, open-ended tasks.

Our take

The finding to carry away is not that AI cannot do jobs. It is that the part of a job AI does best is the part that is easiest to specify, and the part it does worst is knowing when its own output is wrong. We would use these agents on bounded tasks with a human checking conclusions, not on open-ended research. Watch Epoch's reruns: if the judgment gap closes with the next model generation, this article will need correcting.

Frequently asked questions

Can AI automate real jobs?

Partly. Epoch's October 8, 2026 test found frontier agents reliable on well-defined tasks like coding, data analysis and computer use, but short on judgment, experiment design and implicit conventions. The authors conclude models cannot yet replace workers, at least not at Epoch.

Which models did Epoch test?

GPT-6 Astra, Claude Fable 5.1, Grok 4.6, Gemini 3.8 Flash, Kimi K3 and Qwen 3.8 Max, each in its own coding harness at a top reasoning setting.

Which model did best in Epoch's test?

Claude Fable 5.1 and GPT-6 Astra were broadly tied in the lead, according to Epoch. Open-weight models lagged further behind, including on well-defined tasks.

What did Epoch automate after the test?

Epoch removed a graphic design restyling task from its suite because Fable 5.1 produced Epoch-quality output, and says the team now largely automates that step.

How reliable is Epoch's test?

It is a small sample: 11 tasks, one run per model per task, graded by one person. The authors say scores are noisy and qualitative findings are more informative.

Why does Kimi K3's score matter?

Kimi K3 scores 158 on the Epoch Capabilities Index, roughly tied with Grok 4.6, yet handled basic tasks less reliably. Epoch argues this shows existing benchmarks miss real weaknesses.

Sources

What each one is, and whose it is.

  1. 1

    Can AI automate Epoch?, Epoch AI (October 8, 2026)

    BenchmarkIndependent of the vendor
  2. 2

    Can AI automate AI R&D yet?, Epoch AI (October 7, 2026)

    BenchmarkIndependent of the vendor
  3. 3

    State of AI Report 2026, Air Street Capital (October 8, 2026)

    OtherIndependent of the vendor
  4. 4

    Epoch AI latest publications, Epoch AI (October 8, 2026)

    DatasetIndependent of the vendor