AgentBug-Smith builds 200-bug benchmark for coding agents
Summary
Researchers have developed AgentBug-Smith, an advanced system designed to automate the identification and reproduction of real-world harness bugs in AI coding agents. Traditional coding agents have struggled to fix these types of bugs, achieving only about 9% success compared to 40% for regular software bugs. AgentBug-Smith constructs a benchmark of 200 evolving harness bugs, derived from real GitHub reports, which helps enhance the capabilities of software agents. The research highlights the unique challenges presented by an agent's code, which includes tool calls and prompts that depend on live model interactions, making it difficult to reproduce and evaluate bugs effectively. This system aims to provide a scalable foundation for improving the self-repair capabilities of coding agents.