Artificial Analysis Coding Agent Index v1.5 adds safety refusal reporting
Summary
The latest update to the Artificial Analysis Coding Agent Index, version 1.5, introduces safety refusal reporting, which tracks how often agents decline to start or continue tasks on safety grounds. This new feature sheds light on agent behavior and scoring discrepancies; notably, Claude Fable 5.1 recorded the highest fallback rates, with fallback attempts contributing 8.8% of the Index's weight in Claude Code and 7.1% in Devin Fusion. Safety refusals can lead to either a fallback to another model, allowing the task to continue, or a blocked attempt, which scores zero. This update enhances the benchmarking process for coding agents, evaluating their performance across various software engineering tasks, including implementation, terminal use, and technical questions.