Build or buy now applies to retrieval too. What Cohere's Compass Cloud and Embed 5 change
Cohere is turning its retrieval stack into a managed service with API, MCP and permission enforcement, plus a new embedding family. How to decide what to build and what to buy.
Listen to this article · 5 min
AI-generated narration of the full article.

Ask an enterprise AI team where its time goes, and retrieval usually tops the list. Connecting SharePoint and Google Drive, parsing scanned PDFs and slide decks, choosing an embedding model, running a vector index, adding keyword search and a reranker, and, hardest of all, making sure nobody retrieves a document they are not allowed to open.
Two Cohere announcements in late September show that layer becoming a product you can buy as a service. That does not make the decision easy, but it changes the options.
What Cohere announced
Compass Cloud, announced on September 25, is a managed version of Cohere’s retrieval platform, now in private beta with a limited number of enterprise partners. According to Cohere, it covers the whole pipeline:
- connect to sources such as SharePoint, OneDrive and Google Drive;
- parse documents, including multimodal content;
- embed with dense and sparse representations, and index them separately from source files;
- retrieve with hybrid semantic, sparse and keyword search, then rerank;
- govern with multi-tenant access control and document-level permissions enforced during retrieval.
It is exposed three ways: an API for custom search and RAG applications, a dedicated and independently versioned MCP server for agents, and a Python SDK. The announcement does not specify deployment regions.
Embed 5, announced on September 30, is a new embedding family in two tiers, Pro and Fast, generally available on Cohere’s API, Microsoft Foundry and Amazon SageMaker. Cohere lists a 128,000-token context, more than 100 languages, text and image inputs, output sizes from 2,048 down to 256 dimensions and compressed formats. Published prices are US$0.12 per million text tokens for Pro and US$0.08 for Fast. Cohere’s benchmark claims, including top scores on visual document retrieval and financial retrieval benchmarks, are its own.
- Connect sources
- Parse documents
- Embed and index
- Hybrid retrieval and rerank
- Enforce document permissions
- Define relevance and evaluate
- Managed in Compass Cloud, per Cohere (private beta): Connect sources · Parse documents · Embed and index · Hybrid retrieval and rerank · Enforce document permissions
- Still yours: Define relevance and evaluate
Can now be bought as a serviceCannot be outsourced
How to think about build or buy
Permissions are the deciding factor. Most internal retrieval projects underestimate permission enforcement. Mirroring access controls from SharePoint or Drive into a vector index, keeping them in sync and enforcing them at query time is hard to get right. A service that does it as part of retrieval removes a major risk. It also means you must verify it: test with users who should and should not see specific documents before go-live.
Embeddings create a quiet lock-in. Changing embedding models means re-embedding and re-indexing every document. With a managed service, that is the provider’s job, on the provider’s schedule. With your own stack, you control the timing and pay the cost. Either way, plan model upgrades as migrations, not as configuration changes.
Dimensions are a cost lever. Smaller vectors and compressed formats reduce storage and search cost, usually with some loss in quality. Test the trade-off on your own documents rather than assuming the largest option is needed.
MCP changes who consumes retrieval. A retrieval layer exposed through MCP is designed to be called by agents, not only by a chat interface. That makes it shared infrastructure, with its own owner, service levels and access policy.
Evaluation stays with you. No managed service knows what a relevant result is for your claims team or your lawyers. The discipline we described in our article on measuring retrieval applies whatever you buy.
For US enterprises
- List your sources and their permission models before talking to vendors. That list decides more than any benchmark.
- Run a permission test suite: real users, real documents, expected allow and deny results.
- Benchmark embeddings on your own corpus, including scanned documents and tables if you have them.
- Ask about re-indexing: who triggers it, how long it takes and what it costs when the model changes.
- Treat a preview as a preview. Compass Cloud is in private beta; plan production timelines accordingly.
The bottom line
Retrieval is modularizing like the rest of the AI stack: parsing, embedding, indexing and permission-aware search can now be bought as one service, or assembled from parts. The right answer depends less on model quality than on your sources, your permission models and your appetite to operate infrastructure. Either way, the definition of a good result remains yours. Our 2026 enterprise AI stack guide places this layer in context.