GPT-6 Astra aced Epoch's board-game benchmark with one card, so Epoch banned the card
Epoch AI's EBR-bench update, published October 7, 2026, shows GPT-6 Astra scoring 21 of 21 by leaning on a single rule-bending card. With the card banned it still averages 16, about 50 percent above the next best model.
Published

A benchmark score is only as good as the rules it was built on, and Epoch AI just found a hole in one. In an update published October 7, 2026, the research group reported that GPT-6 Astra hit a perfect 21 out of 21 on EBR-bench, and that every one of its top scores relied on a single card that sidesteps the game's central pressure. Epoch banned the card. Astra's average fell from 19.8 to 16, and it is still far ahead of everything else.
Epoch's update carries no on-record quotes from named people beyond the write-up itself, so this piece relies on the published numbers.
What EBR-bench tests
EBR-bench has models play Earthborne Rangers, a cooperative board game. One playthrough takes human players roughly two to four hours, which makes the benchmark a test of long-horizon, multi-step work. The point is learning from experience: whether a model gets better across repeated playthroughs, explores different decks and improves its tactical play.
Epoch's earlier report, on July 1, 2026, found that models struggled with all of that. They did not explore deck options well, did not optimize tactics over time and did not improve their scores much with experience. That is why a model that does improve is news.
The score, and the card
On the original benchmark, GPT-6 Astra reached a perfect 21 of 21, and did so in more than half of its attempts. Its average was 19.8, against 10.5 for Claude Opus 5, the second-highest model.
Then Epoch looked at how. All of Astra's top scores relied on one card that bypasses the game's default time constraints. With the card, Astra finished in 88 turns. Without it, 42. A top human baseline participant also used the card and reached a maximum score, taking 183 turns.
Two details keep this from being a simple cheating story.
- The card was not necessary. The second-ranked human baseline participant scored 100 percent without it, in 76 turns.
- Epoch found no direct evidence of memorization or cheating. It also notes that Earthborne Rangers players have known about the card since near the game's 2023 release, and that the designers did not fully anticipate how leaving it unbanned would bypass the central fatigue mechanic.
In other words, the benchmark had a loophole that humans and a model both found. The fault lies in the test design as much as in the model.

What the ban changed
Epoch banned the card in the default experiment, the one that produces headline scores and feeds the Epoch Capabilities Index. It will still run internal experiments with the card allowed, to track whether other models catch up.
For most models, the ban did not matter much. It produced no statistically significant change for Claude Fable 5.1, Claude Opus 5 or GPT-5.6 Sol, and the direction of any effect was mixed. Epoch will report post-ban scores only for those three models and Astra, plus future models, and will not rerun older results.
For Astra, the ban cost about four points. Under the ban it averaged 16, with a best score of 20 of 21. Two human baseliners each scored 21. That is still roughly a 50 percent jump over the strongest previous models.
The learning curve is the more interesting part. Astra improved from 11 of 21 on its first playthrough to 21 of 21 on its second. The top human went from 1 of 21 on the first try to 21 of 21 on the sixth. A model that learns faster than a person on a board game is a fair headline, if the caveats come with it. Epoch adds that Astra's fatigue management, a proxy for tactical decision-making, remains mediocre and shows no improvement across playthroughs.
Multi-agent setups did not help much
Epoch also tested whether handing the game to a team helps. Using the Inspect Deep Agent harness, developed by the UK AI Security Institute, it allowed up to four subagents under the same ten-playthrough limit. Prototyping with up to eight subagents showed no significant difference.
Multi-agent setups increased deck exploration for most of the four models tested, but the effect was statistically significant only for Claude Opus 5. Topline scores were not significantly affected, and the effect was not consistently positive or negative. More agents is not automatically more intelligence, a point worth remembering when companies pitch swarms of them.
Why this matters beyond a board game
Benchmarks shape which models look best, which gets priced into funding rounds, and which gets cited in policy. When a single rules exploit can swing a headline result, that chain is fragile. We made a similar point in our explainer on what AI leaderboards rank and what a top score hides.
Epoch is also candid about the limits. It expects EBR-bench to be saturated within a few months, is retiring some run configurations, reducing sample sizes from ten to five, and says these are the final planned changes. It plans to test models on more complex games.
Our take
This is how benchmark maintenance should look: find the loophole, say who else used it, ban it, and publish the before and after. The result we would trust is the card-free one. Astra is a real step up at learning from experience on this task, at an average of 16 of 21 against 10.5 for the next best model before the ban, and still a long way from reliable tactical judgment. Watch the harder games Epoch has promised, because a benchmark that saturates in months tells you less each week.
Frequently asked questions
What is EBR-bench?
EBR-bench is Epoch AI's benchmark in which models play the cooperative board game Earthborne Rangers. A playthrough takes humans two to four hours, so it tests whether models learn from experience over long, multi-step tasks.
What did GPT-6 Astra score on EBR-bench?
A perfect 21 of 21 on the original benchmark, averaging 19.8. Claude Opus 5, the second-highest model, averaged 10.5. With the card banned, Astra averaged 16 with a best of 20.
Why did Epoch ban a card?
All of Astra's top scores relied on one card that bypasses the game's default time constraints and its central fatigue mechanic. A top human also used it. Epoch banned it from the default experiment.
Did GPT-6 Astra cheat or memorize the game?
Epoch says it found no direct evidence of memorization or cheating. Players have known about the card since near the game's 2023 release, and the designers did not fully anticipate its effect.
Did the card ban change other models' scores?
Not significantly for Claude Fable 5.1, Claude Opus 5 or GPT-5.6 Sol, and the direction of the effect was mixed. Epoch will report post-ban scores only for those three models, Astra and future models.
Do multiple agents improve EBR-bench scores?
Not reliably. Up to four subagents increased deck exploration for most models, but the effect was statistically significant only for Claude Opus 5, and topline scores were not significantly affected.
Sources
What each one is, and whose it is.
- 1
EBR-bench update, Epoch AI (October 7, 2026)
BenchmarkIndependent of the vendor - 2
Epoch AI latest publications, Epoch AI (October 8, 2026)
DatasetIndependent of the vendor - 3
Can AI automate Epoch?, Epoch AI (October 8, 2026)
BenchmarkIndependent of the vendor - 4
State of AI Report 2026, Air Street Capital (October 8, 2026)
OtherIndependent of the vendor