← All insights

News analysisLens: United States3 min read

Most enterprise RAG failures start in retrieval. Measure it as its own system

Cohere's RCP-nDCG@10 uses rubric-based, calibrated AI judgments to score retrieval when labeled test sets fall short. Why enterprise teams should evaluate retrieval separately from the final answer.

Listen to this article · 5 min

AI-generated narration of the full article.

Curved library shelves full of books seen from below, in La Madre duotone, beside the words Measure retrieval
Photo: Patrik Goethe (StockSnap, CC0)

When an enterprise assistant gives a bad answer, the first suspect is usually the model. Often the model did exactly what it was asked: it summarized the documents it was given. The documents were the problem. The search step returned an outdated policy, the wrong contract version or nothing relevant at all.

Most teams still evaluate only the final answer. That makes retrieval failures look like model failures, and leads to the wrong fix. Cohere’s new evaluation method for retrieval is a good occasion to separate the two.

What Cohere published

On September 30, Cohere introduced RCP-nDCG@10 (Rubric-Calibrated Preferences nDCG@10), a way to measure how relevant a search system’s top ten results are.

The problem it targets is real. Standard retrieval metrics depend on pre-labeled relevance judgments for each query. When a newer system surfaces relevant documents that were never labeled, the metric counts them as misses. Cohere reports that, among documents its benchmarks labeled irrelevant, 28% were judged useful by humans.

RCP-nDCG replaces fixed labels with a calibrated AI judge that assesses every retrieved document using two signals:

  • a relevance rubric of five yes or no questions, applied the same way to every query;
  • pairwise comparisons between documents to estimate their relative order.

A calibration step, based on item response theory, combines them. In Cohere’s words, “the comparisons set the order, and the rubric places it on a scale shared by every query.”

Cohere reports that RCP-nDCG picked the system human reviewers preferred 77% of the time, versus 52% for conventional nDCG, and that at the widest margins reviewers agreed with it 97% of the time. These are Cohere’s own results. The paper, code and data are published, which allows independent checking.

Two places to measure an enterprise RAG system01Question02Retrieval: top 10documents03Generation: answerfrom those documents04Answer to the userMeasure relevance of what was foundMeasure faithfulness and correctness of the answer
  1. Question
  2. Retrieval: top 10 documents
  3. Generation: answer from those documents
  4. Answer to the user

Measure relevance of what was foundMeasure faithfulness and correctness of the answer

If you only measure the last box, you cannot tell a search failure from a model failure.

Why this matters for enterprise teams

Retrieval deserves its own scorecard. A RAG system is two systems: one that finds and one that writes. They fail differently and are fixed differently. Better chunking, metadata filters, hybrid search or a reranker fix retrieval. A better prompt or model fixes generation. Without separate metrics, teams change the wrong thing.

Labeled test sets go stale. Enterprise content changes every week: new policies, new products, new contracts. A test set labeled six months ago penalizes a system for finding documents that did not exist then. Rubric-based judging is one way to keep evaluation current without relabeling everything by hand.

The rubric is a business artifact. “Relevant” means different things to a claims analyst and to a lawyer. Writing the five questions that define relevance for a use case is work for the people who own that use case, not only for engineers.

AI judges need calibration too. Cohere’s method is explicit about calibration because an uncalibrated judge simply moves the problem. Any LLM judge should be checked against a sample of human judgments, and checked again when the judge model changes. We made the same point about production evaluation in our analysis of agent observability.

For US enterprises

  1. Log retrieval results separately from answers, for every query in testing and a sample in production.
  2. Write a relevance rubric with the business owner of each use case: what makes a document useful for this question.
  3. Build a small human-labeled set and use it to calibrate any AI judge before trusting its scores.
  4. Track a ranking metric over time, such as nDCG@10 or a calibrated variant, and rerun it when content, embeddings or chunking change.
  5. Fix retrieval before swapping models. When answers degrade, check what was retrieved first.

The bottom line

RCP-nDCG@10 is a research contribution, not a product you need to buy. Its value for enterprise teams is the discipline it encodes: treat retrieval as a production system with its own definition of quality, its own tests and its own owner. As we argued in our analysis of Palantir’s AIP Evolve, evaluation is becoming the control surface for AI systems. Retrieval is where many of those controls should start.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case