Perplexity AI releases pplx-embed-v2-late models for multimodal search
Summary
Perplexity AI has released two late-interaction embedding models, known as pplx-embed-v2-late, which enhance multimodal retrieval by allowing searches over text, images, and pages within a shared embedding space. These models maintain a separate vector for each token, employing MaxSim scoring for matching, thus preserving detail when querying longer documents and visual content. By embedding images and rendered pages directly, they facilitate effective searches through various formats like PDFs and slides, while maintaining the original layout and elements that traditional text extraction may overlook. Both models are publicly available on Hugging Face and have demonstrated improved performance on key benchmarks, significantly enhancing retrieval capabilities in the process.