GPT-6 Astra leads Epoch Capabilities Index with ECI of 166
Summary
GPT-6 Astra has achieved the top position in the Epoch Capabilities Index (ECI) with a score of 169 at launch, surpassing competitors like Claude Fable 5.1. However, its performance in software engineering has regressed to an ECI of 164, following the addition of new benchmark results, still allowing it to be competitive with Fable 5.1’s score of 167. The ECI reflects a model's abilities across multiple benchmarks, with domain-specific scores, such as SWE-ECI for software engineering, providing a focused evaluation of a model's performance in specialized areas.
Analysis
Kimi K3: Kimi K3 is a large language model featured in Epoch AI's ongoing ECI assessments. It provides context for capability comparisons among current frontier systems. The model is displayed in the latest Domain-specific ECI Explorer results. GPT-5.6 Sol: GPT-5.6 Sol is an advanced large language model tracked in Epoch AI's ECI rankings. It serves as a reference point in recent comparisons of frontier model capabilities. The model appears in the Domain-specific ECI Explorer alongside other top performers. GPT-6 Astra: GPT-6 Astra is a frontier large language model evaluated on a wide range of benchmarks. It currently leads the overall Epoch Capabilities Index with strong performance across multiple domains. In the latest ECI update, it also achieved the highest Math-ECI score while maintaining competitive results in software engineering. Claude Fable 5.1: Claude Fable 5.1 is a frontier large language model from the Claude family known for high performance on specialized tasks. It remains state-of-the-art on software engineering benchmarks according to the latest domain-specific ECI evaluation. The model is included in Epoch AI's comparative analysis of leading AI systems. AI Benchmarking: Epoch AI maintains the Epoch Capabilities Index as a composite metric that aggregates performance across dozens of benchmarks to compare general model capabilities. Domain Specialization: Domain-specific ECI scores isolate model performance within areas like mathematics or software engineering by refitting on relevant benchmarks while preserving overall difficulty scaling.
Categories
aitechai_agents