Beyond One-Shot Testing
For years, artificial intelligence evaluation has suffered from a stubborn blind spot: it treats intelligence like a multiple-choice quiz. We feed models static problems, measure their one-shot answers, and rank them on leaderboards. But real-world problem solving isn't a single prompt and response. It requires adaptation, trial and error, and the ability to learn from past mistakes over multiple iterations.
That is precisely what Epoch AI’s EBR-bench sets out to measure. Instead of testing how well a model can solve a puzzle on its first try, this benchmark asks a sharper question: Can an AI system improve its score across repeated playthroughs of a complex, relatively obscure cooperative card and strategy game called Earthborne Rangers?
By forcing models to navigate the same game world ten times in a row, with nothing but their own episodic memory and accumulated context to guide them, EBR-bench offers a rare window into actual in-context learning capabilities. And as the results show, getting models to learn from their own virtual experiences is a lot messier than simple scaling laws might suggest.
The Setup: Ten Runs and a Text Interface
Earthborne Rangers is not your standard benchmark fare like Chess or Go. It is an open-world, narrative-driven tabletop game where players act as rangers protecting a futuristic wilderness. Translating that to an AI benchmark required boiling the physical components down to a text-based interface built on the game's core engine.
In a standard EBR-bench evaluation, an AI agent is given a set of rules, a basic scenario, and a specific objective: survive, explore, and rack up a high topline score across a sequence of playthroughs. Under the v2 methodology adopted by Epoch AI, a default run consists of 10 playthroughs. The model is told only its final score after each run, leaving it to figure out on its own what went wrong, which strategies failed, and how to optimize its deck and actions next time around.
To keep things rigorous, Epoch updated their scoring metric from the mean of a sample's final 20% of playthroughs to the best of its final two playthroughs. This change mirrors their ongoing human baseline studies and ensures that anomalous bad luck doesn't completely sink an otherwise competent agent's final grade.
What the Leaderboard Reveals
When you look at how frontier models stack up on EBR-bench, the picture defies easy categorization. Epoch has run evaluations on models including Claude Opus 4.8, Claude Opus 5, Claude Fable 5, GPT-5.5, and GPT-5.6 Sol.
The differences in learning trajectories are fascinating. Some models manage to steadily climb their score curves as they accumulate context across the ten playthroughs, while others plateau early or struggle to maintain coherence as their conversation history balloons. Crucially, the benchmark also tests models with a "strategy guide"—effectively an answer sheet reflecting expert human gameplay. Comparing unguided runs against guided runs reveals whether a model is genuinely figuring out the mechanics or simply relying on memorized heuristics.
The results highlight a recurring frustration in modern AI evaluation: raw capability doesn't automatically translate to effective adaptation. A model might be brilliant at writing code or summarizing dense texts, but put it in a persistent environment where it has to remember why it failed turn four, and cracks start to show. This mirrors the finding from enterprise agent benchmarks like GTM Bench, where realistic multi-step workflows expose capability gaps that static exams miss.
Context Compaction and Hidden Pitfalls
One of the most revealing aspects of Epoch's methodology is how sensitive these agentic benchmarks are to technical scaffolding. During the evolution of EBR-bench, researchers experimented heavily with context compaction thresholds.
Initially set at a conservative 250,000 tokens, Epoch later raised the compaction threshold to approximately 90% of each model’s full context window—usually capped around 880,000 tokens for GPT models to prevent compaction endpoint issues. The impact was striking: some models scored significantly better at lower compaction settings, while others saw their performance drop. For instance, Claude Fable 5 scored 11.9 percentage points worse at the lower compaction setting, whereas GPT-5.5 dropped 13.3 percentage points when restricted.
These swings demonstrate that when we test "AI learning," we are often testing the delicate plumbing of context management, memory summarization, and prompt engineering just as much as core reasoning. If an agent forgets a crucial rule from playthrough two because its memory was prematurely compacted, its "learning curve" flattens out entirely. It is a sharp illustration of why engineering teams increasingly treat context-window resilience as a first-class pipeline design concern, not an afterthought.
Measuring Against Human Baselines
To ground these machine scores in reality, Epoch ran a human baseline experiment between July and August 2026. They recruited 15 participants with varying familiarity with the game, paying them $30 an hour with performance bonuses for beating AI benchmarks or hitting expert levels.
The humans played through the exact same text-based interface, giving researchers a direct point of comparison. While humans brought their own cognitive quirks and fatigue limits to the table, their learning trajectory provided a vital sanity check against the endless optimism of synthetic scaling. Seeing how humans adapt to Earthborne Rangers highlights just how much nuance is required to solve an open-ended survival game—nuance that LLMs are only beginning to approximate.
Where Benchmarking Goes From Here
EBR-bench is a reminder that the frontier of AI evaluation has shifted. Static datasets and multiple-choice benchmarks are rapidly saturating, pushing researchers toward interactive, multi-turn, persistent environments that score whole cohorts of runs rather than single traces, capturing systemic behavior like learning curves and regressions.
By testing whether models can actually learn from experience rather than just parroting training data, Epoch AI has built a tough, highly informative diagnostic tool. It strips away the illusion of competence and exposes the real limitations of in-context adaptation. Until models can reliably learn from their mistakes without choking on their own memory logs, true autonomous agency remains just out of reach.