AgentBug-Smith builds 200-bug benchmark for coding agents

Summary

Researchers have developed AgentBug-Smith, an advanced system designed to automate the identification and reproduction of real-world harness bugs in AI coding agents. Traditional coding agents have struggled to fix these types of bugs, achieving only about 9% success compared to 40% for regular software bugs. AgentBug-Smith constructs a benchmark of 200 evolving harness bugs, derived from real GitHub reports, which helps enhance the capabilities of software agents. The research highlights the unique challenges presented by an agent's code, which includes tool calls and prompts that depend on live model interactions, making it difficult to reproduce and evaluate bugs effectively. This system aims to provide a scalable foundation for improving the self-repair capabilities of coding agents.

Analysis

AgentBug-Smith: AgentBug-Smith is an automated harness bug reproduction approach designed to continuously discover and reproduce real-world harness bugs from open-source agentic systems. It turns GitHub bug reports into runnable tests, outperforming prior techniques for general software bugs across different backbone LLMs. The tool directly supports the news by enabling scalable creation of benchmarks and evaluation of coding agents on fixing issues in their own code harnesses, such as tool calls, memory, and prompts. Live-Harness-Bench: Live-Harness-Bench is a live and extensible benchmark of reproducible harness bugs built by applying AgentBug-Smith to open-source agentic systems in the wild. It provides both an evaluation set for testing software agents on real harness bug repair and a knowledge base for distilling reusable repair skills from past fixes. This benchmark underpins the paper's findings on agent performance limitations and potential improvements for self-improving AI systems. Agent Code Challenges: An agent's own code encompasses everything around the model, including tool calls, memory, and prompts, with bugs that depend on live model calls and are therefore difficult to recreate and test. Self-Improvement Path: Before trusting a coding agent with an agent's code, testing it on bugs already fixed provides a practical way to assess readiness for autonomous fixes.

Categories

aimachine_learningai_agentstechvirtuals
View Original Tweet