Models

Perplexity Releases Efficient pplx-embed-v2 Model

Perplexity released preview weights for its pplx-embed-v2-context-9b-preview model, delivering high-accuracy contextual retrieval with an eightfold reduction in vector storage.

AlphaSignal1 day agoModels
Image: AlphaSignal

Perplexity has launched preview weights for pplx-embed-v2-context-9b-preview, a 9-billion-parameter contextual embedding model. Unlike traditional models that rely on single gold-chunk labels, this model is trained by distilling chunk relevance from a query-aware context compression teacher. It was initialized from an in-house 9B-parameter ColBERT retrieval model, using Matryoshka training to support 1024-dimensional and 2048-dimensional representations, alongside quantization-aware training for native int8 output. The training mixture spanned roughly 430 public and internal datasets across more than 50 languages.

In evaluations, the model achieved the highest average nDCG@10 on the public ConTEB benchmark among tested contextual embedding models, though pplx-context-v1-4B scored higher on NarrativeQA and Nemotron-3-8B led on COVID-QA. Its chunk-size sensitivity was modest, with average nDCG@10 declining from 81.0% with 64-token chunks to 79.9% with 512-token chunks. On context-bench—a private benchmark by turbopuffer containing 2,099 queries across 38,894 long documents with a median length of 6,100 tokens in 21 domains—the new model outperformed voyage-context-4 by 14.4 points on Answer and Evidence recall at K=10. Crucially, it matched the quality of voyage-context-4 while using just 1 KB per vector instead of 8 KB, representing an eightfold reduction in raw storage. However, the older pplx-context-v1-4B maintained higher Document@3 and Document@5 scores.

For practitioners, adopting this model requires several operational adjustments. Because it uses late chunking to process entire documents in a single forward pass, retrieval pipelines must encode complete documents together rather than embedding chunks independently. Any update to a document section requires re-embedding the entire document, as changes alter contextual representations elsewhere. Additionally, vector databases must natively support 1,024-dimensional int8 vectors to realize the storage benefits. Teams must also measure indexing costs, as whole-document encoding alters GPU memory usage, batching, and throughput compared to standard chunk embedding.

This is our own summary of reporting by AlphaSignal

More in Models