pplx-embed-v2-context-9b-preview: contextual embedding beyond the gold passage
- Type
- model
- Venue
- Perplexity (2026-09-30; preview)
- Year
- 2026
- Source
- blog
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
New method to train contextual embedding models: instead of embedding each document chunk on its own, the 9B model encodes the whole document once and pools chunk vectors afterward, so every chunk embedding sees the full document. The new training signal: instead of one labeled "gold" chunk per query, relevance is distilled from Perplexity's own context compression model, which scores every document token against the query; this teaches the model to retrieve both the answer and the context that supports it. Claimed results: pplx-embed-v2-context-9b-preview sets a new SOTA on ConTEB and on turbopuffer's new private context-bench (evaluated blind): +14.4 points answer recall@10 over voyage-context-4; at 1024 dims in int8 (1KB per vector) it still beats voyage-context-4 at 8KB per vector.
Keywords
embeddings · retrieval · RAG · context compression
Topics
embeddings, retrieval, RAG, context compression
Research notes
- Discovery: Denis Yarats (@denisyarats, verified, cofounder & CTO @perplexity_ai) 2026-09-30 post: https://x.com/denisyarats/status/2105380502203195835?s=20 (quotes a @perplexity_ai announcement)
- Blog: https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage
- Model on Hugging Face: https://huggingface.co/perplexity-ai/pplx-embed-v2-context-9b-preview
- License not stated in the post.
- Connects to the collection's embedding and retrieval entries (Cohere Embed 5).