Perplexity AI releases pplx-embed-v2-late models for multimodal search

Summary

Perplexity AI has released two late-interaction embedding models, known as pplx-embed-v2-late, which enhance multimodal retrieval by allowing searches over text, images, and pages within a shared embedding space. These models maintain a separate vector for each token, employing MaxSim scoring for matching, thus preserving detail when querying longer documents and visual content. By embedding images and rendered pages directly, they facilitate effective searches through various formats like PDFs and slides, while maintaining the original layout and elements that traditional text extraction may overlook. Both models are publicly available on Hugging Face and have demonstrated improved performance on key benchmarks, significantly enhancing retrieval capabilities in the process.

Analysis

Hugging Face: Hugging Face operates a platform for hosting, sharing, and deploying machine learning models. It provides the public repository where Perplexity AI has made the pplx-embed-v2-late models available for download and use. Perplexity AI: Perplexity AI develops AI-powered search and retrieval systems. The company has introduced pplx-embed-v2-late, a pair of late-interaction embedding models designed for multimodal document search. The models are released publicly via Hugging Face. pplx-embed-v2-late: pplx-embed-v2-late comprises two late-interaction embedding models from Perplexity AI that support retrieval across text, images, and full pages. They use per-token vectors with MaxSim scoring to retain fine-grained details in lengthy or visual content. Both sizes are distilled from a shared teacher model for consistent embedding spaces. Compatibility: Models of different sizes distilled from the same teacher share one embedding space, supporting queries from a smaller model against a corpus indexed by a larger one. Model Architecture: Late-interaction designs keep separate vectors per token and apply MaxSim scoring to match query tokens against the most relevant document tokens. Multimodal Retrieval: Direct embedding of images and rendered pages allows searches over PDFs, slides, and scans while preserving tables, figures, and layout.

Categories

machine_learningaiai_agentstech
View Original Tweet