← Back to explorer

Cohere Embed 5 (Pro and Fast embeddings family)

Type
model
Venue
Cohere, generally available 2026-09-30
Year
2026
Source
blog
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

Cohere's new state-of-the-art embeddings family in two tiers: Embed 5 Pro for frontier capabilities and Embed 5 Fast for low-latency performance. Snapshot specs (identical for both unless noted): 128K token context; text, image, and fused text+image inputs (page images embedded directly or fused with metadata into a single vector); 100+ languages with cross-lingual retrieval; output dims 2048/1536/1024/768/512/256 with Matryoshka representations and float/int8/binary formats (1024-dim int8 recommended sweet spot; 256-dim binary = 32 bytes/vector, 256x reduction; up to 96% compression without quality loss); Pro and Fast vectors share an embedding space, so teams can index with Pro and query with Fast without re-indexing (cross-model combos lose only 1.6-2.7% mean retrieval quality). Both self-hostable with vLLM for private VPC/on-prem deployment. Benchmarks: ViDoRe V3 (RCP-nDCG@10, 8 enterprise document domains) Pro 85.8 (+8.8 over Embed 4), ahead of Voyage 4 Large (83.7), Gemini Embedding 2 (83.2), OpenAI text-embedding-3-large (75.5); Fast 84.5, ahead of Gemini Embedding 2 and Voyage 4 Large. Finance: Pro leads on FinanceBench (80.1), FinQA (90.0), ViDoRe V3 Finance (85.0), +3.3 avg over the next non-Cohere competitor. Fused text-image: Pro 82.3 vs Gemini Embedding 2 61.3. Page-image retrieval: Pro 77.0. Multilingual: Pro highest average across DE/FR/ES/IT/RU (77) and strong on JA/CN/KO/AR/FA/HI/BN/TE/ID/TH. Fast: ~2.4x higher document throughput than Pro, beats Voyage 4 Nano by ~7 points on ViDoRe V3, outperforms Qwen3-VL-Embedding-2B by ~20 points. Pricing: Pro $0.12/1M tokens, Fast $0.08/1M tokens. Access: Cohere API, Model Vault, Microsoft Foundry, Amazon SageMaker, North; batch embedding; integrations with LangChain, Haystack, Weaviate, Qdrant, Pinecone, Elasticsearch, MongoDB, Redis, Milvus, OpenSearch.

Keywords

embeddings · retrieval · multimodal · RAG · multilingual · compression

Topics

embeddings, retrieval, multimodal, RAG

Research notes

  • Discovery: Cohere (@cohere, verified) 2026-09-30: https://x.com/cohere/status/2105285142394896435
  • Product page: https://cohere.com/embed
  • Technical evaluation annotations/code: https://github.com/cohere-ai/rcp-ndcg
  • API model IDs: embed-v5.0-pro / embed-v5.0-fast.
  • Connects to the collection's embeddings, retrieval, and RAG entries (alongside topk-embed-v1).