CyberGym-E2E-AA benchmark reveals cost-effective AI models for cybersecurity tasks
Summary
The CyberGym-E2E-AA benchmark, which assesses cybersecurity defense capabilities, has revealed that a majority of frontier intelligence models are restricted from effectively responding to over 85% of the tasks due to safety concerns tied to identifying security vulnerabilities. This benchmark, adapted from a public design to evaluate how AI agents handle vulnerabilities in widely used open-source software, highlights the difficulty of balancing offensive and defensive cybersecurity strategies. Despite the limitations faced by many models, more capable options like GPT-6 Luna and MiMo-V2.6-Pro have emerged as cost-effective solutions, with the ability to conduct around 100 bug hunts on extensive codebases for approximately $20 each, significantly cheaper than alternatives like Grok 4.7.