← Back to explorer

topk-embed-v1 (TopK multimodal late-interaction retriever)

Type
model
Venue
TopK
Year
2026
Source
huggingface
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

A family of open-source and hosted text/image embedding models from TopK, announced as frontier retrieval quality at $0.05/1M tokens with highly compressible representations for serving at scale with storage comparable to dense baselines. The 2B variant (topk-embed-v1-small) is a multimodal late-interaction retriever finetuned from Qwen3.5-2B: instead of one vector per input it stores multiple embeddings per text token or image patch and scores with MaxSim. Text queries search both text documents and images (scanned pages, reports, slides); 1024 tokens per query, 8192 tokens per text document; 2048-dim embeddings with Matryoshka (MRL) prefix support (e.g. 256) and Ward token pooling (2x/4x/8x). Evaluated on the eight public ViDoRe v3 test datasets: image-native nDCG@10 65.22 (Recall 69.26), image-crosslingual 63.17/67.46, markdown-native 62.80/67.00, markdown-crosslingual 60.49/65.08. Requires a CUDA GPU with bfloat16 (Ampere or newer) and custom trust_remote_code code.

Keywords

embeddings · retrieval · multimodal · late interaction · RAG · ViDoRe

Topics

embeddings, retrieval, multimodal, late interaction

Research notes

  • Discovery: topk.io (@topk_io, verified) 2026-09-29: https://x.com/topk_io/status/2104963920469307671
  • Variants: https://huggingface.co/topk-io/topk-embed-v1-small (2B, current) and https://huggingface.co/topk-io/topk-embed-v1-xsmall (0.9B).
  • Note: the originally announced repo path topk-io/topk-embed-v1 now 404s (renamed to -small).
  • No paper, blog, or code-repo link published on the model page as of the read; company site https://topk.io.
  • Connects to the collection's embeddings, retrieval, and RAG entries.