← Back to explorer

pplx-embed-v2-context-9b-preview: contextual embedding beyond the gold passage

Type
model
Venue
Perplexity (2026-09-30; preview)
Year
2026
Source
blog
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

New method to train contextual embedding models: instead of embedding each document chunk on its own, the 9B model encodes the whole document once and pools chunk vectors afterward, so every chunk embedding sees the full document. The new training signal: instead of one labeled "gold" chunk per query, relevance is distilled from Perplexity's own context compression model, which scores every document token against the query; this teaches the model to retrieve both the answer and the context that supports it. Claimed results: pplx-embed-v2-context-9b-preview sets a new SOTA on ConTEB and on turbopuffer's new private context-bench (evaluated blind): +14.4 points answer recall@10 over voyage-context-4; at 1024 dims in int8 (1KB per vector) it still beats voyage-context-4 at 8KB per vector.

Keywords

embeddings · retrieval · RAG · context compression

Topics

embeddings, retrieval, RAG, context compression

Research notes

  • Discovery: Denis Yarats (@denisyarats, verified, cofounder & CTO @perplexity_ai) 2026-09-30 post: https://x.com/denisyarats/status/2105380502203195835?s=20 (quotes a @perplexity_ai announcement)
  • Blog: https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage
  • Model on Hugging Face: https://huggingface.co/perplexity-ai/pplx-embed-v2-context-9b-preview
  • License not stated in the post.
  • Connects to the collection's embedding and retrieval entries (Cohere Embed 5).