Cohere announced a new evaluation metric called Rubric‑Calibrated Preferences nDCG@10 (RCP‑nDCG@10). It uses a calibrated AI judge to score every retrieved document against a relevance rubric, instead of relying only on pre‑labeled answer keys. In a blind human study, the metric chose the system reviewers preferred 77% of the time, versus 52% for conventional nDCG. Cohere says it will use RCP‑nDCG@10 to optimise its next‑generation Embed and Rerank models and will add it to the MTEB benchmark.
The company is also looking for private‑beta partners for its Compass Cloud search platform.
Why it matters
The metric gives a clearer picture of search quality, helping developers build systems that return truly relevant results.