MIT and Sakana AI's SIFT framework reduces coding agent evaluation costs with 35% accuracy
Summary
MIT and Sakana AI have introduced the SIFT framework, which utilizes a language model to significantly reduce the costs associated with evaluating coding agents, achieving an accuracy of 35.1% on the Polyglot benchmark while using fewer CPU hours and less financial resources. Traditionally, evaluating candidate changes in coding agents can be costly and time-consuming, requiring thousands of CPU hours. SIFT addresses this issue by allowing a separate language model to conduct preliminary assessments of modifications, enabling multiple evaluations to occur in parallel and streamlining the search for effective improvements. This innovative approach is part of a broader trend in AI research, where academic institutions and specialized AI labs work together to enhance the practical applications of self-improving agents.