topk-embed-v1 (TopK multimodal late-interaction retriever)
- Type
- model
- Venue
- TopK
- Year
- 2026
- Source
- huggingface
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
A family of open-source and hosted text/image embedding models from TopK, announced as frontier retrieval quality at $0.05/1M tokens with highly compressible representations for serving at scale with storage comparable to dense baselines. The 2B variant (topk-embed-v1-small) is a multimodal late-interaction retriever finetuned from Qwen3.5-2B: instead of one vector per input it stores multiple embeddings per text token or image patch and scores with MaxSim. Text queries search both text documents and images (scanned pages, reports, slides); 1024 tokens per query, 8192 tokens per text document; 2048-dim embeddings with Matryoshka (MRL) prefix support (e.g. 256) and Ward token pooling (2x/4x/8x). Evaluated on the eight public ViDoRe v3 test datasets: image-native nDCG@10 65.22 (Recall 69.26), image-crosslingual 63.17/67.46, markdown-native 62.80/67.00, markdown-crosslingual 60.49/65.08. Requires a CUDA GPU with bfloat16 (Ampere or newer) and custom trust_remote_code code.
Keywords
embeddings · retrieval · multimodal · late interaction · RAG · ViDoRe
Topics
embeddings, retrieval, multimodal, late interaction
Research notes
- Discovery: topk.io (@topk_io, verified) 2026-09-29: https://x.com/topk_io/status/2104963920469307671
- Variants: https://huggingface.co/topk-io/topk-embed-v1-small (2B, current) and https://huggingface.co/topk-io/topk-embed-v1-xsmall (0.9B).
- Note: the originally announced repo path topk-io/topk-embed-v1 now 404s (renamed to -small).
- No paper, blog, or code-repo link published on the model page as of the read; company site https://topk.io.
- Connects to the collection's embeddings, retrieval, and RAG entries.