MIT and Sakana AI's SIFT framework reduces coding agent evaluation costs with 35% accuracy

Summary

MIT and Sakana AI have introduced the SIFT framework, which utilizes a language model to significantly reduce the costs associated with evaluating coding agents, achieving an accuracy of 35.1% on the Polyglot benchmark while using fewer CPU hours and less financial resources. Traditionally, evaluating candidate changes in coding agents can be costly and time-consuming, requiring thousands of CPU hours. SIFT addresses this issue by allowing a separate language model to conduct preliminary assessments of modifications, enabling multiple evaluations to occur in parallel and streamlining the search for effective improvements. This innovative approach is part of a broader trend in AI research, where academic institutions and specialized AI labs work together to enhance the practical applications of self-improving agents.

Analysis

MIT: MIT is a prominent research university with strong programs in computer science and artificial intelligence. Researchers from MIT collaborated with Sakana AI to develop the SIFT framework for more efficient evaluation of coding agent modifications. SIFT: Recursive Self-Improvement via Fast Tree Search (SIFT) is a framework that incorporates language model pairwise judgments to prioritize candidate agents before full benchmark runs. Developed by MIT and Sakana AI, it enables asynchronous exploration of agent modifications while building on prior self-improvement methods. o3-mini: o3-mini is a closed language model optimized for reasoning and code-related tasks. SIFT achieved strong results when using o3-mini as the underlying agent model during benchmark testing. Polyglot: Polyglot is a multi-language coding benchmark for assessing agent performance across diverse tasks. SIFT was evaluated on Polyglot to demonstrate more efficient discovery of high-performing coding agent variants. SWE-bench: SWE-bench is a software engineering benchmark with verified task subsets for evaluating coding agents. SIFT experiments on its verified subset highlighted advantages of judge-guided candidate selection. Sakana AI: Sakana AI is an AI research organization focused on novel approaches to agentic systems and self-improving models. It partnered with MIT on the SIFT project to address high costs in recursive self-improvement loops for coding agents. TerminalBench: TerminalBench is a benchmark suite targeting coding agents in command-line and terminal settings. SIFT used it to show how LLM judges can identify promising agents even when small test sets produce misleading signals. Qwen3-Coder-30B: Qwen3-Coder-30B is an open-weight language model designed for coding and programming tasks. It served as one of the base models in SIFT experiments comparing the framework against earlier self-improvement approaches. Darwin Gödel Machine: Darwin Gödel Machine (DGM) is a recursive self-improvement system that maintains an archive of agents and evaluates candidates on expanding task sets. SIFT extends DGM by integrating LLM-based ranking to improve search efficiency. Huxley-Gödel Machine: Huxley-Gödel Machine (HGM) is a self-improvement method that factors in performance across agent lineages when selecting search directions. SIFT was directly compared to HGM in experiments on coding benchmarks. Benchmark Diversity: Researchers are applying self-improvement frameworks across multiple coding and software engineering benchmarks to validate broader applicability. Evaluation Efficiency: Pairwise comparisons by language models are emerging as a complementary signal to full benchmark runs when tuning coding agents. AI Research Collaboration: Academic institutions and specialized AI labs are jointly developing methods to make self-improving agents more practical for real-world use.

Categories

techaimachine_learningai_agents
View Original Tweet