Cohere Embed 5 Uses One Vector Space for Pro and Fast
marktechpost.com

Cohere Embed 5 Uses One Vector Space for Pro and Fast

Tech News
3 min read

Published by AINave Editorial

TL;DRCohere’s Embed 5 lets teams index with Pro and serve queries with Fast without rebuilding the index, as long as both use the same output dimension. The practical appeal is clearest in that workflow; its headline benchmark comparisons are Cohere-reported and measure reranking more than first-stage retrieval.

Cohere’s Embed 5 embedding model makes its Pro and Fast tiers more than separate quality and speed options: they share an embedding space, so a system can index documents with Pro and embed live queries with Fast without re-indexing. That compatibility depends on matching output dimensions, and the performance evidence comes from Cohere’s own evaluations.supported details

One index, two query paths

Pro is designed for maximum retrieval quality; Fast targets lower latency and cost on the query path. In Cohere’s development evaluation across 40 datasets, a Pro-index/Fast-query setup scored 98.4 when Pro/Pro was normalized to 100. Fast/Fast scored 96.6. Cohere also measured 377.3 documents per second for Fast, versus 159.7 for Pro.evaluation figures

That makes the shared space the most useful product distinction: teams can keep the corpus vectors produced by Pro while using Fast for repeated queries. Cohere’s reported figures suggest this can preserve much of the Pro/Pro score while increasing document throughput, though the evaluation is not independent evidence of performance in every retrieval system.

Both tiers accept text, images, or fused text-plus-image inputs, support more than 100 languages, and handle inputs up to 128K tokens. Embedding a page image directly, or combining it with metadata, is relevant for documents such as scanned pages, slides, schematics, and charts, where text alone may leave visual information out.model capabilities

Strong reported scores, with a metric caveat

On ViDoRe V3, Cohere reports averages of 85.8 for Pro and 84.5 for Fast, compared with 83.7 for Voyage 4 Large, 83.2 for Gemini Embedding 2, and 75.5 for OpenAI text-embedding-3-large. These are vendor-reported comparisons, not independently replicated rankings.ViDoRe V3 results

Most of the reported results use Cohere’s RCP-nDCG@10 metric, which reorders a fixed candidate set. It therefore says more about reranking quality than first-stage retrieval recall. That distinction matters when applying the scores to a search system whose candidate generation is itself a major part of the problem.metric description and replication status

The results also vary by language: Pro leads the European-language average, while Gemini Embedding 2 beats it on nine of ten additional languages in Cohere’s table. A single overall score does not settle which model fits a multilingual corpus.language comparison

Price and vector size are separate trade-offs

The reported text-input prices are $0.12 per million tokens for Pro and $0.08 for Fast; image inputs are listed at $0.40 per million tokens for either tier. These rates make Fast cheaper per text token, but they do not by themselves establish total workload cost.reported pricing

Embed 5 offers six output dimensions, from 256 to 2048, and float, int8, or binary formats. Cohere recommends 1024-dimensional int8 as a storage-quality compromise. At the extreme, a 256-dimensional binary vector takes 32 bytes, compared with 8 KB for a 2048-dimensional float32 vector; the smaller binary representation can trade away accuracy, so storage savings are not a free quality gain.vector options and storage examples

The key implementation choice is therefore not simply Pro versus Fast. It is whether the shared-space query path and vector format fit the workload: matching dimensions enables tier mixing, while compression changes the size and potentially the quality of the vectors being searched.

FAQs

It is an embedding-model family for enterprise search, RAG, and agentic retrieval, with text, image, and fused text-plus-image inputs.

Sources

Latest Tech News