Research

Epoch AI hides game identities to stop benchmark cheating

Epoch AI has launched Mystery Game Puzzles, a benchmark that hides the identity of its underlying game to prevent AI labs from artificially inflating their models' reasoning scores.

AlphaSignal4 days agoResearch
Image: AlphaSignal

Research organization Epoch AI has introduced a new benchmark called Mystery Game Puzzles to combat data contamination and benchmark-specific tuning. The evaluation tests artificial intelligence models on 100 positions from a popular but undisclosed game, requiring them to determine the single best next move. To prevent labs from intentionally or accidentally optimizing their systems for this specific test, Epoch AI is keeping the game's identity, prompts, example positions, and model transcripts entirely secret.

Currently, the closed-source model Opus 5 holds the top score on the benchmark at 59 percent, followed closely by GPT-5.5, which previously achieved 56 percent. In the open-weight category, Qwen3.8-Max leads with a score of 38 percent, narrowly outperforming GPT-5.4 and Opus 4.8. Although performance on the benchmark surged from 25 percent to 56 percent within a two-month window early in the year, progress has plateaued, and scores have remained largely flat since April.

The benchmark's design reveals that raw computing power is not the limiting factor for these systems. While models are allocated a generous budget of 1 million output tokens, the leading models utilize fewer than 400,000 tokens. This indicates that errors stem from fundamental reasoning failures rather than resource constraints. For practitioners, this benchmark shifts the focus of evaluation toward genuine, zero-shot planning and spatial reasoning. By eliminating the possibility of post-training or fine-tuning on curated examples of the target game, the benchmark provides a more honest assessment of how a model generalizes to unfamiliar logical environments.

This is our own summary of reporting by AlphaSignal

More in Research