Most enterprise RAG failures start in retrieval. Measure it as its own system
Cohere's RCP-nDCG@10 uses rubric-based, calibrated AI judgments to score retrieval when labeled test sets fall short. Why enterprise teams should evaluate retrieval separately from the final answer.
Listen to this article · 5 min
AI-generated narration of the full article.

When an enterprise assistant gives a bad answer, the first suspect is usually the model. Often the model did exactly what it was asked: it summarized the documents it was given. The documents were the problem. The search step returned an outdated policy, the wrong contract version or nothing relevant at all.
Most teams still evaluate only the final answer. That makes retrieval failures look like model failures, and leads to the wrong fix. Cohere’s new evaluation method for retrieval is a good occasion to separate the two.
What Cohere published
On September 30, Cohere introduced RCP-nDCG@10 (Rubric-Calibrated Preferences nDCG@10), a way to measure how relevant a search system’s top ten results are.
The problem it targets is real. Standard retrieval metrics depend on pre-labeled relevance judgments for each query. When a newer system surfaces relevant documents that were never labeled, the metric counts them as misses. Cohere reports that, among documents its benchmarks labeled irrelevant, 28% were judged useful by humans.
RCP-nDCG replaces fixed labels with a calibrated AI judge that assesses every retrieved document using two signals:
- a relevance rubric of five yes or no questions, applied the same way to every query;
- pairwise comparisons between documents to estimate their relative order.
A calibration step, based on item response theory, combines them. In Cohere’s words, “the comparisons set the order, and the rubric places it on a scale shared by every query.”
Cohere reports that RCP-nDCG picked the system human reviewers preferred 77% of the time, versus 52% for conventional nDCG, and that at the widest margins reviewers agreed with it 97% of the time. These are Cohere’s own results. The paper, code and data are published, which allows independent checking.
- Question
- Retrieval: top 10 documents
- Generation: answer from those documents
- Answer to the user
Measure relevance of what was foundMeasure faithfulness and correctness of the answer
Why this matters for enterprise teams
Retrieval deserves its own scorecard. A RAG system is two systems: one that finds and one that writes. They fail differently and are fixed differently. Better chunking, metadata filters, hybrid search or a reranker fix retrieval. A better prompt or model fixes generation. Without separate metrics, teams change the wrong thing.
Labeled test sets go stale. Enterprise content changes every week: new policies, new products, new contracts. A test set labeled six months ago penalizes a system for finding documents that did not exist then. Rubric-based judging is one way to keep evaluation current without relabeling everything by hand.
The rubric is a business artifact. “Relevant” means different things to a claims analyst and to a lawyer. Writing the five questions that define relevance for a use case is work for the people who own that use case, not only for engineers.
AI judges need calibration too. Cohere’s method is explicit about calibration because an uncalibrated judge simply moves the problem. Any LLM judge should be checked against a sample of human judgments, and checked again when the judge model changes. We made the same point about production evaluation in our analysis of agent observability.
For US enterprises
- Log retrieval results separately from answers, for every query in testing and a sample in production.
- Write a relevance rubric with the business owner of each use case: what makes a document useful for this question.
- Build a small human-labeled set and use it to calibrate any AI judge before trusting its scores.
- Track a ranking metric over time, such as nDCG@10 or a calibrated variant, and rerun it when content, embeddings or chunking change.
- Fix retrieval before swapping models. When answers degrade, check what was retrieved first.
The bottom line
RCP-nDCG@10 is a research contribution, not a product you need to buy. Its value for enterprise teams is the discipline it encodes: treat retrieval as a production system with its own definition of quality, its own tests and its own owner. As we argued in our analysis of Palantir’s AIP Evolve, evaluation is becoming the control surface for AI systems. Retrieval is where many of those controls should start.