Perplexity AI's pplx-embed-v2-context-9b-preview sets new benchmark for contextual embedding models

Summary

Perplexity AI has announced the release of its new contextual embedding model, pplx-embed-v2-context-9b-preview, which sets a new standard for performance in the context-bench evaluation framework. This model innovatively encodes document chunks with the complete context of the document, addressing common limitations in retrieval systems that separate long texts into isolated chunks. By using a query-aware context compression approach, the model distills relevance signals from all tokens in a document, allowing it to train without the constraints of traditional gold-chunk supervision. It was submitted for blind evaluation, demonstrating leading performance in answer and evidence retrieval tasks.

Analysis

turbopuffer: turbopuffer develops vector search infrastructure and created the privately held context-bench benchmark. The benchmark draws its design and capabilities from direct customer feedback on retrieval needs. It serves as an evaluation platform for new models including pplx-embed-v2-context-9b-preview. context-bench: context-bench is a privately held benchmark for context-aware retrieval developed by turbopuffer. Its queries and documents draw from conversations with the company's customers and support evaluation of models on answer, evidence, and document retrieval tasks. The benchmark enables blind submissions for standardized comparison of contextual embedding approaches. pplx-embed-v2-context-9b-preview: pplx-embed-v2-context-9b-preview is a contextual embedding model from Perplexity AI designed for RAG systems that processes a document as a list of chunks and encodes them jointly. Each resulting chunk embedding incorporates surrounding document context while producing one vector per chunk. It is released as a preview on Hugging Face with separate query and document encoding methods and is currently leading on the ConTEB benchmark. Benchmark Focus: context-bench evaluates retrieval performance specifically on answer and evidence tasks at multiple cutoffs in addition to standard document retrieval metrics. Training Approach: The model overcomes reliance on limited gold-chunk supervision by distilling relevance signals from a query-aware context compression model that scores document tokens against queries. Contextual Embeddings: Contextual embedding models overcome information loss from splitting documents into isolated chunks by encoding each segment while retaining visibility of the full surrounding document.

Categories

machine_learningaitech
View Original Tweet