mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-30 00:18:32 +02:00
measuring-retrieval-relevance: accessibility pass on ranking metrics
Expands MRR with Mean Reciprocal Rank and a Wikipedia link where the metric becomes operational, matching the inline expansion treatment already given to NDCG. Drops the IR acronym in favor of the plainer "ranking metrics" phrasing. Strips the parenthetical subtitle from the Pitfalls heading and normalizes the intro to use "golden sets". Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.7
parent
4d5a3b8f60
commit
d615b8fc30
+4
-4
@@ -58,7 +58,7 @@ Document:
|
||||
|
||||
## Using the Golden Set
|
||||
|
||||
<a href="https://amenra.github.io/ranx/" target="_blank">ranx</a> is a Python library for ranking-metric evaluation. It covers the standard IR metrics (`recall@k`, `MRR`, `NDCG@k`, `Precision@k`, MAP, and others) through one consistent interface, so you don't hand-roll each metric or juggle different libraries as needs grow.
|
||||
<a href="https://amenra.github.io/ranx/" target="_blank">ranx</a> is a Python library for ranking-metric evaluation. It covers the standard ranking metrics (`recall@k`, `MRR`, `NDCG@k`, `Precision@k`, MAP, and others) through one consistent interface, so you don't hand-roll each metric or juggle different libraries as needs grow.
|
||||
|
||||
The evaluation runs in three steps: load the labeled queries into the shape ranx expects, run each through Qdrant, then compute metrics.
|
||||
|
||||
@@ -144,15 +144,15 @@ Higher is better on all three. Which metric matters most depends on what your pi
|
||||
| Single-answer retrieval (FAQ or Q&A) | `MRR` or `Hits@1` | The first result is what the user acts on; lower ranks matter little |
|
||||
| Re-ranking or recommendation feeds | `NDCG@k` | Order within the result list matters; a highly relevant doc at rank 5 is worse than at rank 1 |
|
||||
|
||||
[NDCG (Normalized Discounted Cumulative Gain)](https://en.wikipedia.org/wiki/Discounted_cumulative_gain) needs graded labels (for example, 0/1/2 scores per query-document pair). For binary labels, stick with `recall@k` and `MRR`. For the full metric list (Precision@k, MAP, ERR, and others), see the <a href="https://amenra.github.io/ranx/" target="_blank">ranx docs</a>.
|
||||
[NDCG (Normalized Discounted Cumulative Gain)](https://en.wikipedia.org/wiki/Discounted_cumulative_gain) needs graded labels (for example, 0/1/2 scores per query-document pair). For binary labels, stick with `recall@k` and [`MRR` (Mean Reciprocal Rank)](https://en.wikipedia.org/wiki/Mean_reciprocal_rank). For the full metric list (Precision@k, MAP, ERR, and others), see the <a href="https://amenra.github.io/ranx/" target="_blank">ranx docs</a>.
|
||||
|
||||
On choosing `k`: set it to match actual usage. If the application shows 5 results to the user, measure `@5`. If a RAG pipeline passes 10 chunks to the LLM, measure `@10`. Reporting `@100` for a UI that surfaces 5 results makes the metric look artificially good.
|
||||
|
||||
Re-run whenever the retrieval stack changes: new embedding model (which also requires re-embedding queries and re-indexing), new index config, or new reranker. In CI, compute `recall@10` against a fixed golden set and fail the job when the score drops below your target threshold.
|
||||
|
||||
## Pitfalls to Watch For (Data Leakage and Friends)
|
||||
## Pitfalls to Watch For
|
||||
|
||||
In golden query sets, **data leakage** means any setup that makes offline metrics look better than production reality. Unlike classic train/test leakage, the issue is often evaluation design. Keep source documents in the index (they are the expected relevant answers). Focus on these risks:
|
||||
In golden sets, **data leakage** means any setup that makes offline metrics look better than production reality. Unlike classic train/test leakage, the issue is often evaluation design. Keep source documents in the index (they are the expected relevant answers). Focus on these risks:
|
||||
|
||||
**Synthetic-query unrealism.** LLMs often mirror source wording, creating easier queries than real user input. This inflates offline scores. Mitigate it by instructing the LLM to generate queries as a user who hasn't seen the source document, then compare synthetic and real-query distributions (length and specificity).
|
||||
|
||||
|
||||
Reference in New Issue
Block a user