Artificial Analysis Coding Agent Index v1.5 adds safety refusal reporting

Summary

The latest update to the Artificial Analysis Coding Agent Index, version 1.5, introduces safety refusal reporting, which tracks how often agents decline to start or continue tasks on safety grounds. This new feature sheds light on agent behavior and scoring discrepancies; notably, Claude Fable 5.1 recorded the highest fallback rates, with fallback attempts contributing 8.8% of the Index's weight in Claude Code and 7.1% in Devin Fusion. Safety refusals can lead to either a fallback to another model, allowing the task to continue, or a blocked attempt, which scores zero. This update enhances the benchmarking process for coding agents, evaluating their performance across various software engineering tasks, including implementation, terminal use, and technical questions.

Analysis

Claude Fable: Claude Fable 5.1 is an AI model variant assessed within coding agent evaluations on the Artificial Analysis platform. It recorded the highest fallback rates among tested models when encountering safety-related refusals during tasks in both Claude Code and Devin Fusion setups. These results incorporate performance from fallback models to maintain continuity in scoring. Artificial Analysis: Artificial Analysis operates a benchmarking platform focused on evaluating AI coding agents through real-world software engineering tasks across multiple benchmarks. It aggregates results into composite indices while tracking additional metrics such as token usage, cost, and execution time. The platform introduced safety refusal reporting in Coding Agent Index v1.5 to better contextualize model behavior and performance differences. Coding Agent Index v1.5: The Coding Agent Index v1.5 serves as a composite metric averaging scores from DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA benchmarks covering repository understanding, implementation, and terminal workflows. Version 1.5 adds safety refusal reporting to explain instances where models decline tasks on safety grounds and to account for fallback or blocked attempts. The index supports comparisons across agent variants, models, and execution configurations. Benchmark Scope: The index aggregates binary pass/fail outcomes across implementation, terminal, and repository Q&A tasks to provide a balanced view of coding agent capabilities. Safety Handling: Coding agents respond to safety refusals either by falling back to another model to complete the task or by stopping the attempt with a block that results in a zero score.

Categories

aiai_agentsmachine_learningtechvirtuals
View Original Tweet