Epoch AI launches new game puzzles benchmark for AI reasoning
Summary
A new "game puzzles" benchmark has been launched to assess AI systems on reasoning-heavy tasks, expanding beyond the previously established Chess Puzzles benchmark. The current record-holder for this benchmark is Opus 5, which achieved a score of 59%. Earlier this year, rapid advancements were made, with scores increasing from 25% in February to 56% in April. However, top scores have plateaued since then, indicating that while AI has made strides in general-purpose reasoning, there remains limited room for further improvement. Notably, open-weight models, like Qwen3.8-Max, are also showing competitive results, scoring 38% and suggesting that these systems are improving in unfamiliar reasoning contexts.
Analysis
Opus 5: Opus 5 is a leading AI model focused on advanced reasoning capabilities. In the context of this news, it has achieved the top score on the newly launched game puzzles benchmark, highlighting its strength in handling unfamiliar domains. The model demonstrates efficient token usage even on challenging tasks where it ultimately errs. GPT-5.4: GPT-5.4 is a GPT-series AI model designed for complex reasoning and problem-solving. It appears in the benchmark results as a strong closed model that Qwen3.8-Max has edged out, providing a comparison point for open-weight advancements in unfamiliar task domains. GPT-5.5: GPT-5.5 is an advanced AI model from the GPT series emphasizing general-purpose reasoning. It previously set a record on the game puzzles benchmark earlier in the year, illustrating the rapid early progress in AI performance on novel puzzle-based evaluations before recent stabilization. Opus 4.6: Opus 4.6 is an earlier version in the Opus AI model line with strong reasoning foundations. It established an initial benchmark record on the game puzzles evaluation in February, marking the starting point of rapid AI improvement tracked in this news. Opus 4.8: Opus 4.8 is a mid-series Opus AI model optimized for iterative reasoning. It features in recent benchmark comparisons as a model surpassed by newer open-weight entries like Qwen3.8-Max, underscoring evolving performance across model types on the puzzles test. Qwen3.8-Max: Qwen3.8-Max is an open-weight AI model developed for high-performance language tasks. It is relevant here as the highest-scoring open-weight entry on the game puzzles benchmark, outperforming several closed models and signaling progress in accessible AI systems for out-of-distribution reasoning. Open-Weight Progress: Open-weight models are achieving competitive results on out-of-distribution reasoning challenges, offering evidence of broader advancements beyond proprietary systems. Benchmark Development: A new game puzzles benchmark has been introduced to evaluate AI reasoning in domains where models lack specific post-training, complementing existing evaluations like chess puzzles. Model Improvement Trends: AI systems have demonstrated gains in general-purpose reasoning on unfamiliar tasks, though progress on top benchmark scores has stabilized after initial rapid advances.
Categories
aicryptotechai_agentsmachine_learning