Furniture Assembly Benchmark scores improve from 28% to 80% in 10 months
Summary
A new benchmark called the Furniture Assembly Benchmark (FAB) has been introduced to evaluate AI models' ability to identify mistakes in IKEA furniture assembly, improving the top score from 28% to 80% in just 10 months. This benchmark tasks AI with assessing builds based on a manual and photos of half-completed furniture, highlighting the distinct error-handling behaviors of different models. For instance, while Gemini 3.1 Pro tends to miss correct builds, GPT-5.4 struggles with spotting mistakes. Notably, GPT-6 Astra excels, being both the fastest and most accurate model tested. This development aligns with trends in AI evaluation, emphasizing the importance of combining visual analysis with instructional understanding for real-world tasks.