From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Our new standard for measuring enterprise retrieval quality, validated against human judgment.
How well does nDCG capture the quality of modern retrieval systems?
nDCG is the established yardstick for measuring how well search systems rank results, and is widely used across benchmarks such as MTEB and BEIR .
It is preferred because it reflects both relevance and position, giving more weight to useful results that appear near the top.
However, it relies on pre-existing relevance labels, which, if incomplete, mean other relevant results can be missed or scored incorrectly.
Here, we introduce Rubric-Calibrated Preferences nDCG@10 (RCP-nDCG@10), an evaluation methodology that grades each retrieved document against the same explicit relevance criteria for every query, using a calibrated AI judge.
Designed to better reflect real-world enterprise retrieval quality, RCP-nDCG@10 is the metric against which we have optimised our next generation search models.
Here we explain how it works, what's different from the status quo, how we validated it against human judgment. nDCG is only as good as its labels …