Agent Arena analyzes 20,840 coding traces to reveal model feedback patterns

Summary

In a recent analysis of user feedback on coding agents from Agent Arena, researchers examined 20,840 traces across 29 models, revealing that newer frontier models generally receive more positive user evaluations compared to their predecessors. Notably, broken code remains the leading cause of user complaints, accounting for 69.1%, followed by issues such as incomplete outputs and usability. Interestingly, while models exhibit common weaknesses, they also vary in specific complaint frequencies; for instance, Fable 5.1 has notably fewer complaints about sloppiness and flaky behavior than Astra.

Analysis

Astra: Astra refers to GPT-6 Astra, OpenAI's frontier model released in early September 2026, optimized for computer use, automation, and certain agentic workflows. Arena traces indicate it draws more complaints around slop, design, and flakiness than Fable 5.1. The news uses it as a comparative example in feedback balance among leading models. Fable: Fable 5.1 is Anthropic's frontier coding and agentic model released in early September 2026, noted for strong performance on intelligence benchmarks and agentic tasks. It shows distinct complaint profiles compared to peers, with fewer issues around sloppiness and flakiness in Arena data. The news highlights its relative strengths in the analyzed traces versus other models. Agent Arena: Agent Arena is a platform that evaluates AI agents through real-world user interactions, collecting traces from tasks like coding and analyzing feedback signals such as praise, complaints, success rates, and error recovery. It powers leaderboards and Pareto frontiers across dozens of models using causal tracing on live sessions. In this news, it supplied the 20,840 coding traces that reveal shifting feedback patterns among frontier models. Dawid Galarowicz: Dawid Galarowicz authors detailed analyses of AI agent performance based on platform data from Agent Arena. His work focuses on qualitative insights from user feedback to track model strengths and weaknesses. The news directs readers to his full article for deeper exploration of the coding agent study. Model Differentiation: Frontier models exhibit shared failure modes but vary noticeably in the frequency of specific complaints such as sloppy output or inconsistent behavior. Model Feedback Trends: Newer frontier models from multiple labs are receiving a more positive balance of user feedback in real-world agent evaluations. Coding Agent Weaknesses: Broken or non-functional code remains the dominant source of user complaints across coding agents, followed by incomplete outputs and usability issues.

Categories

machine_learningaiai_agentsvirtuals

Related sources

View Original Tweet