From 92af150b00cddf5d8e1af776090601cdd68ceac6 Mon Sep 17 00:00:00 2001 From: Dylan Couzon Date: Wed, 22 Apr 2026 17:13:27 -0400 Subject: [PATCH] retrieval-quality: rename to "Measuring ANN Precision" - Update title, H1, and Time in retrieval-quality.md (30 min -> 15 min reflects the pivot to Web UI) - Rename references in tutorials-lp-overview.md and the headless tutorial index; swap the pill from Python to Web UI to reflect the new primary flow - Replace "ANN recall" with "ANN precision" in Fundamentals (4 places: intro, comparison note, ladder table, cross-link) and in the golden-set tutorial's layer-1 cross-reference - Filename kept as retrieval-quality.md so existing URLs and aliases still work --- .../headless/content/tutorials/search-engineering.md | 2 +- .../content/documentation/tutorials-lp-overview.md | 2 +- .../retrieval-quality-fundamentals.md | 8 ++++---- .../retrieval-quality-golden-set.md | 2 +- .../tutorials-search-engineering/retrieval-quality.md | 6 +++--- 5 files changed, 10 insertions(+), 10 deletions(-) diff --git a/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md b/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md index 5041da568..3ff505852 100644 --- a/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md +++ b/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md @@ -7,7 +7,7 @@ | [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | Python | 30m | Intermediate | | [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/) | Understand evaluation levels and choose the right metric. | Python | 20m | Intermediate | | [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Generate ground truth data at scale and avoid data leakage. | Python | 40m | Intermediate | -| [Retrieval Quality Evaluation](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall and tune HNSW parameters. | Python | 30m | Intermediate | +| [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN precision with the Web UI and tune HNSW parameters. | Web UI | 15m | Intermediate | | [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | Python | 40m | Intermediate | | [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | Python | 45m | Intermediate | | [Multivectors and Late Interaction](/documentation/tutorials-search-engineering/using-multivector-representations/) | Effective use of multivector representations. | Python | 30m | Intermediate | diff --git a/qdrant-landing/content/documentation/tutorials-lp-overview.md b/qdrant-landing/content/documentation/tutorials-lp-overview.md index aae6e5e3f..103e6060d 100644 --- a/qdrant-landing/content/documentation/tutorials-lp-overview.md +++ b/qdrant-landing/content/documentation/tutorials-lp-overview.md @@ -70,7 +70,7 @@ partition: qdrant | [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | Python | 30m | Intermediate | | [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/) | Understand evaluation levels and choose the right metric. | Python | 20m | Intermediate | | [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Generate ground truth data at scale and avoid data leakage. | Python | 40m | Intermediate | -| [Retrieval Quality Evaluation](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall and tune HNSW parameters. | Python | 30m | Intermediate | +| [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN precision with the Web UI and tune HNSW parameters. | Web UI | 15m | Intermediate | | [Reranking for Better Search](/documentation/search-precision/reranking-semantic-search/) | Use multivector representations for better ranking. | Python | 30m | Intermediate | | [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | Python | 40m | Intermediate | | [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | Python | 45m | Intermediate | diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md index 30848008b..9a5fa8cb6 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md @@ -12,9 +12,9 @@ aliases: Before measuring retrieval quality, it's worth understanding what you're measuring. Retrieval quality operates at three distinct levels, and it's easy to optimize for the wrong one. -The first level is **ANN recall**: does the approximate search return the same results as an exact nearest-neighbor search? This is a purely algorithmic question about how faithfully HNSW approximates exhaustive search. It has nothing to do with whether those results are useful to a human. +The first level is **ANN precision**: does the approximate search return the same results as an exact nearest-neighbor search? This is a purely algorithmic question about how faithfully HNSW approximates exhaustive search. It has nothing to do with whether those results are useful to a human. -The second level is **retrieval relevance**: of the results returned, how many are relevant to the query intent? This requires a labeled ground-truth dataset or human judgment. A pipeline can achieve near-perfect ANN recall and still surface irrelevant documents if the embeddings are a poor fit for the task. +The second level is **retrieval relevance**: of the results returned, how many are relevant to the query intent? This requires a labeled ground-truth dataset or human judgment. A pipeline can achieve near-perfect ANN precision and still surface irrelevant documents if the embeddings are a poor fit for the task. The third level is **business impact**: does better retrieval lead to better outcomes like lower hallucination rates in downstream LLMs, higher task-completion rates, or improved user satisfaction scores? This is what stakeholders care about, but it's the hardest to measure directly. The causal chain from a vector match to a user outcome is long and easily dominated by generator behavior, UI, and other confounders, so no single offline metric is a reliable proxy for a KPI. The next section describes how teams bridge this gap in practice. @@ -24,7 +24,7 @@ The three levels aren't measured in isolation. Teams that successfully connect r | # | Layer | What it measures | Cadence | Cost | |---|---|---|---|---| -| 1 | ANN recall | `Recall@k` vs exact kNN on a sampled query set | On index or embedding changes | Low | +| 1 | ANN precision | `Recall@k` vs exact kNN on a sampled query set | On index or embedding changes | Low | | 2 | Retrieval relevance | `Recall@k` / `NDCG@k` vs a labeled golden set | Weekly, or on retrieval-stack changes | Low per run; **golden set is the real cost** (see next tutorial) | | 3 | End-to-end answer quality | LLM-as-judge or human rating on the golden set | Weekly, or on retrieval or generator changes | Moderate (LLM-judge cost × eval size) | | 4 | Business impact | Online A/B behind a flag | Per release, once offline layers pass | High (traffic, experimentation infra) | @@ -71,4 +71,4 @@ On choosing `k`: set it to match actual usage. If the application shows 5 result ## Next Steps -To build a labeled dataset for relevance evaluation, see [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/). To measure and tune the ANN recall of a Qdrant collection in practice, see [Retrieval Quality Evaluation](/documentation/tutorials-search-engineering/retrieval-quality/). +To build a labeled dataset for relevance evaluation, see [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/). To measure and tune the ANN precision of a Qdrant collection in practice, see [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/). diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md index f164ca4b6..f2dd19977 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md @@ -10,7 +10,7 @@ aliases: | Time: 40 min | Level: Intermediate | | | |--------------|---------------------|--|----| -This tutorial covers **layer 2** of the evaluation ladder: **retrieval relevance**. Measuring how well retrieved results match real user intent requires a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). For layer 1 (ANN recall against exact kNN), which needs no relevance labels, use the **Search Quality** tab in the Qdrant Web UI. +This tutorial covers **layer 2** of the evaluation ladder: **retrieval relevance**. Measuring how well retrieved results match real user intent requires a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). For layer 1 (ANN precision against exact kNN), which needs no relevance labels, use the **Search Quality** tab in the Qdrant Web UI. ## Generating Queries diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md index 5073953e9..d50b561b6 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md @@ -1,14 +1,14 @@ --- -title: Retrieval Quality Evaluation +title: Measuring ANN Precision aliases: - /documentation/tutorials/retrieval-quality/ - /documentation/beginner-tutorials/retrieval-quality/ weight: 6 --- -# Evaluate Retrieval Quality +# Measuring ANN Precision -| Time: 30 min | Level: Intermediate | | | +| Time: 15 min | Level: Intermediate | | | |--------------|---------------------|--|----| This tutorial measures **layer 1** of the evaluation ladder, **ANN precision**: the share of Qdrant's approximate nearest-neighbor top-k that appears in the exact kNN top-k. For retrieval relevance (layer 2), see the Building a Golden Query Set tutorial.