retrieval-quality: rename to "Measuring ANN Precision"

- Update title, H1, and Time in retrieval-quality.md
    (30 min -> 15 min reflects the pivot to Web UI)
  - Rename references in tutorials-lp-overview.md and the headless
    tutorial index; swap the pill from Python to Web UI to reflect
    the new primary flow
  - Replace "ANN recall" with "ANN precision" in Fundamentals
    (4 places: intro, comparison note, ladder table, cross-link)
    and in the golden-set tutorial's layer-1 cross-reference
  - Filename kept as retrieval-quality.md so existing URLs and
    aliases still work
This commit is contained in:
Dylan Couzon
2026-04-22 17:13:27 -04:00
parent 175f92ddbe
commit 92af150b00
5 changed files with 10 additions and 10 deletions
@@ -7,7 +7,7 @@
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/) | Understand evaluation levels and choose the right metric. | <span class="pill">Python</span> | 20m | <span class="text-yellow">Intermediate</span> |
| [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Generate ground truth data at scale and avoid data leakage. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Retrieval Quality Evaluation](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall and tune HNSW parameters. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN precision with the Web UI and tune HNSW parameters. | <span class="pill">Web UI</span> | 15m | <span class="text-yellow">Intermediate</span> |
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Multivectors and Late Interaction](/documentation/tutorials-search-engineering/using-multivector-representations/) | Effective use of multivector representations. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
@@ -70,7 +70,7 @@ partition: qdrant
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/) | Understand evaluation levels and choose the right metric. | <span class="pill">Python</span> | 20m | <span class="text-yellow">Intermediate</span> |
| [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Generate ground truth data at scale and avoid data leakage. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Retrieval Quality Evaluation](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall and tune HNSW parameters. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN precision with the Web UI and tune HNSW parameters. | <span class="pill">Web UI</span> | 15m | <span class="text-yellow">Intermediate</span> |
| [Reranking for Better Search](/documentation/search-precision/reranking-semantic-search/) | Use multivector representations for better ranking. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
@@ -12,9 +12,9 @@ aliases:
Before measuring retrieval quality, it's worth understanding what you're measuring. Retrieval quality operates at three distinct levels, and it's easy to optimize for the wrong one.
The first level is **ANN recall**: does the approximate search return the same results as an exact nearest-neighbor search? This is a purely algorithmic question about how faithfully HNSW approximates exhaustive search. It has nothing to do with whether those results are useful to a human.
The first level is **ANN precision**: does the approximate search return the same results as an exact nearest-neighbor search? This is a purely algorithmic question about how faithfully HNSW approximates exhaustive search. It has nothing to do with whether those results are useful to a human.
The second level is **retrieval relevance**: of the results returned, how many are relevant to the query intent? This requires a labeled ground-truth dataset or human judgment. A pipeline can achieve near-perfect ANN recall and still surface irrelevant documents if the embeddings are a poor fit for the task.
The second level is **retrieval relevance**: of the results returned, how many are relevant to the query intent? This requires a labeled ground-truth dataset or human judgment. A pipeline can achieve near-perfect ANN precision and still surface irrelevant documents if the embeddings are a poor fit for the task.
The third level is **business impact**: does better retrieval lead to better outcomes like lower hallucination rates in downstream LLMs, higher task-completion rates, or improved user satisfaction scores? This is what stakeholders care about, but it's the hardest to measure directly. The causal chain from a vector match to a user outcome is long and easily dominated by generator behavior, UI, and other confounders, so no single offline metric is a reliable proxy for a KPI. The next section describes how teams bridge this gap in practice.
@@ -24,7 +24,7 @@ The three levels aren't measured in isolation. Teams that successfully connect r
| # | Layer | What it measures | Cadence | Cost |
|---|---|---|---|---|
| 1 | ANN recall | `Recall@k` vs exact kNN on a sampled query set | On index or embedding changes | Low |
| 1 | ANN precision | `Recall@k` vs exact kNN on a sampled query set | On index or embedding changes | Low |
| 2 | Retrieval relevance | `Recall@k` / `NDCG@k` vs a labeled golden set | Weekly, or on retrieval-stack changes | Low per run; **golden set is the real cost** (see next tutorial) |
| 3 | End-to-end answer quality | LLM-as-judge or human rating on the golden set | Weekly, or on retrieval or generator changes | Moderate (LLM-judge cost × eval size) |
| 4 | Business impact | Online A/B behind a flag | Per release, once offline layers pass | High (traffic, experimentation infra) |
@@ -71,4 +71,4 @@ On choosing `k`: set it to match actual usage. If the application shows 5 result
## Next Steps
To build a labeled dataset for relevance evaluation, see [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/). To measure and tune the ANN recall of a Qdrant collection in practice, see [Retrieval Quality Evaluation](/documentation/tutorials-search-engineering/retrieval-quality/).
To build a labeled dataset for relevance evaluation, see [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/). To measure and tune the ANN precision of a Qdrant collection in practice, see [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/).
@@ -10,7 +10,7 @@ aliases:
| Time: 40 min | Level: Intermediate | | |
|--------------|---------------------|--|----|
This tutorial covers **layer 2** of the <a href="/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/#connecting-the-levels-in-practice" target="_blank">evaluation ladder</a>: **retrieval relevance**. Measuring how well retrieved results match real user intent requires a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). For layer 1 (ANN recall against exact kNN), which needs no relevance labels, use the **Search Quality** tab in the <a href="/documentation/web-ui/" target="_blank">Qdrant Web UI</a>.
This tutorial covers **layer 2** of the <a href="/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/#connecting-the-levels-in-practice" target="_blank">evaluation ladder</a>: **retrieval relevance**. Measuring how well retrieved results match real user intent requires a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). For layer 1 (ANN precision against exact kNN), which needs no relevance labels, use the **Search Quality** tab in the <a href="/documentation/web-ui/" target="_blank">Qdrant Web UI</a>.
## Generating Queries
@@ -1,14 +1,14 @@
---
title: Retrieval Quality Evaluation
title: Measuring ANN Precision
aliases:
- /documentation/tutorials/retrieval-quality/
- /documentation/beginner-tutorials/retrieval-quality/
weight: 6
---
# Evaluate Retrieval Quality
# Measuring ANN Precision
| Time: 30 min | Level: Intermediate | | |
| Time: 15 min | Level: Intermediate | | |
|--------------|---------------------|--|----|
This tutorial measures **layer 1** of the <a href="/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/#connecting-the-levels-in-practice" target="_blank">evaluation ladder</a>, **ANN precision**: the share of Qdrant's approximate nearest-neighbor top-k that appears in the exact kNN top-k. For retrieval relevance (layer 2), see the <a href="/documentation/tutorials-search-engineering/retrieval-quality-golden-set/" target="_blank">Building a Golden Query Set</a> tutorial.