CyberGym-E2E-AA benchmark reveals cost-effective AI models for cybersecurity tasks

Summary

The CyberGym-E2E-AA benchmark, which assesses cybersecurity defense capabilities, has revealed that a majority of frontier intelligence models are restricted from effectively responding to over 85% of the tasks due to safety concerns tied to identifying security vulnerabilities. This benchmark, adapted from a public design to evaluate how AI agents handle vulnerabilities in widely used open-source software, highlights the difficulty of balancing offensive and defensive cybersecurity strategies. Despite the limitations faced by many models, more capable options like GPT-6 Luna and MiMo-V2.6-Pro have emerged as cost-effective solutions, with the ability to conduct around 100 bug hunts on extensive codebases for approximately $20 each, significantly cheaper than alternatives like Grok 4.7.

Analysis

Berkeley RDI: Berkeley RDI created the original CyberGym-E2E benchmark designed to test AI agents on full-cycle defensive cybersecurity from vulnerability discovery through patching. The news describes Artificial Analysis' adaptation of this benchmark into CyberGym-E2E-AA for its own public leaderboard and evaluations. CyberGym-E2E-AA: CyberGym-E2E-AA is Artificial Analysis' implementation of the CyberGym-E2E benchmark, which evaluates AI agents on end-to-end defensive cybersecurity tasks focused on memory-safety vulnerabilities in real C/C++ open-source projects. It requires agents to discover a vulnerability, reproduce a crash with a proof-of-concept input, and apply a patch that resolves the issue while ensuring the project's tests continue to pass. The benchmark is featured in the news to assess and compare frontier AI models on these capabilities. Artificial Analysis: Artificial Analysis develops and runs independent evaluations and leaderboards that measure AI model performance across coding, reasoning, agentic workflows, and specialized domains such as cybersecurity. In the news, it created and hosts the CyberGym-E2E-AA benchmark and leaderboard to provide standardized comparisons of model effectiveness and efficiency on complex defensive tasks. Benchmark Implementation: CyberGym-E2E-AA adapts a public defensive cybersecurity benchmark originally developed for evaluating AI agents on realistic vulnerability handling in widely used open-source software. Cost-Effectiveness Trends: Among the highest-performing models on cybersecurity benchmarks, several stand out for delivering strong results at significantly lower computational cost compared to alternatives. Model Safety Considerations: Frontier AI models encounter safety restrictions when presented with tasks that involve identifying or addressing security vulnerabilities, limiting their ability to respond on a large share of evaluations.

Categories

aiai_agentstechmachine_learning
View Original Tweet