Rename retrieval-quality slugs to match tutorial scope

Tutorial slugs now match their content:
- retrieval-quality/ -> ann-recall/ (alias preserves the old published path)
- retrieval-quality-golden-set/ -> retrieval-relevance/
- retrieval-quality-pipeline-output/ -> pipeline-output-quality/

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Dylan Couzon
2026-04-30 16:32:46 -04:00
co-authored by Claude Opus 4.7
parent 33bd663d80
commit 9fbcd1adad
7 changed files with 14 additions and 13 deletions
@@ -5,7 +5,7 @@
| [Relevance Feedback](/documentation/tutorials-search-engineering/using-relevance-feedback/) | Relevance Feedback Retrieval in Qdrant | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall with the Web UI and tune HNSW parameters. | <span class="pill">Web UI</span> | 15m | <span class="text-green">Beginner</span> |
| [Measuring ANN Recall](/documentation/tutorials-search-engineering/ann-recall/) | Measure ANN recall with the Web UI and tune HNSW parameters. | <span class="pill">Web UI</span> | 15m | <span class="text-green">Beginner</span> |
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Multivectors and Late Interaction](/documentation/tutorials-search-engineering/using-multivector-representations/) | Effective use of multivector representations. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
@@ -10,5 +10,5 @@ partition: ecosystem
| Tutorial | Objective | Stack | Time | Level |
| :--- | :--- | :--- | :--- | :--- |
| [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) | Build a labeled golden set and score retrieval relevance with ranx. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/) | Build a labeled golden set and score retrieval relevance with ranx. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Evaluating Pipeline Output Quality](/documentation/improve-search/pipeline-output-quality/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
@@ -14,9 +14,9 @@ partition: ecosystem
This tutorial focuses on **pipeline output quality**: whether the full retrieval pipeline produces the right output once retrieved results reach a consumer, most often an LLM generator in a RAG system.
To measure pipeline output quality, you run your golden set through the full pipeline, capture each `(question, retrieved_context, answer)` triple, and score the triples against judgment metrics like faithfulness, answer relevancy, and context precision.
Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) (does the approximate index match exact kNN?) and [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) (do the top-k results match query intent?).
Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/ann-recall/) (does the approximate index match exact kNN?) and [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/) (do the top-k results match query intent?).
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/)), LLM access for generation and judging, and Python with `ragas` installed.
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/)), LLM access for generation and judging, and Python with `ragas` installed.
## Wiring the RAG Pipeline
@@ -162,7 +162,7 @@ Ragas isn't the only tool in this space: <a href="https://docs.confident-ai.com/
## Isolating Retrieval vs Generation
If you're also running [retrieval evaluation](/documentation/improve-search/retrieval-quality-golden-set/) against the same golden set, pairing the two scores on every run gives a diagnostic 2x2 for attributing score changes. When a metric drops after a change (new embedding model, new prompt, or new chunking strategy), the pair tells you which half of the pipeline to investigate.
If you're also running [retrieval evaluation](/documentation/improve-search/retrieval-relevance/) against the same golden set, pairing the two scores on every run gives a diagnostic 2x2 for attributing score changes. When a metric drops after a change (new embedding model, new prompt, or new chunking strategy), the pair tells you which half of the pipeline to investigate.
Pair `recall@10` from the retrieval evaluation with `faithfulness` from the pipeline-output evaluation. In the table, High and Low are relative to the target thresholds you set per metric.
@@ -14,7 +14,7 @@ partition: ecosystem
This tutorial focuses on **retrieval relevance**: how well retrieved results match real user intent.
To measure retrieval relevance, you need a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). This tutorial covers both building that dataset and running it through Qdrant to compute relevance metrics.
Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) (does the approximate index match exact kNN?) and [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/) (does the end-to-end pipeline produce the right output?).
Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/ann-recall/) (does the approximate index match exact kNN?) and [Evaluating Pipeline Output Quality](/documentation/improve-search/pipeline-output-quality/) (does the end-to-end pipeline produce the right output?).
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + optional payload), an embedding model available to encode queries at evaluation time, and Python with `ranx` installed.
@@ -178,4 +178,4 @@ In golden sets, **data leakage** means any setup that makes offline metrics look
## Next Steps
Once retrieval relevance is on target, the next layer is pipeline output quality: whether the full pipeline produces the right output when retrieval feeds into a consumer (LLM generator, ranker, or UI). See [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/).
Once retrieval relevance is on target, the next layer is pipeline output quality: whether the full pipeline produces the right output when retrieval feeds into a consumer (LLM generator, ranker, or UI). See [Evaluating Pipeline Output Quality](/documentation/improve-search/pipeline-output-quality/).
@@ -68,7 +68,7 @@ partition: develop
| [Semantic Search Basics](/documentation/tutorials-search-engineering/neural-search/) | Deploy a search service for company descriptions. | <span class="pill">FastAPI</span> | 30m | <span class="text-green">Beginner</span> |
| [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall with the Web UI and tune HNSW parameters. | <span class="pill">Web UI</span> | 15m | <span class="text-green">Beginner</span> |
| [Measuring ANN Recall](/documentation/tutorials-search-engineering/ann-recall/) | Measure ANN recall with the Web UI and tune HNSW parameters. | <span class="pill">Web UI</span> | 15m | <span class="text-green">Beginner</span> |
| [Reranking for Better Search](/documentation/search-precision/reranking-semantic-search/) | Use multivector representations for better ranking. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
@@ -3,6 +3,7 @@ title: Measuring ANN Recall
aliases:
- /documentation/tutorials/retrieval-quality/
- /documentation/beginner-tutorials/retrieval-quality/
- /documentation/tutorials-search-engineering/retrieval-quality/
weight: 5
---
@@ -20,8 +21,8 @@ This tutorial focuses on **ANN recall**: how closely approximate nearest-neighbo
ANN recall measures how closely approximate search matches exact kNN. It's the first of four evaluation layers; each higher layer measures a different property of the retrieval system, with different tools.
- **ANN recall** (this tutorial). Is the approximate index close to exact kNN?
- **Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/)). Do the top-k results match query intent?
- **Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/)). Does the end-to-end pipeline (retrieval + generator, ranker, or UI) produce the right output?
- **Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/)). Do the top-k results match query intent?
- **Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/improve-search/pipeline-output-quality/)). Does the end-to-end pipeline (retrieval + generator, ranker, or UI) produce the right output?
- **Business impact**. Do the KPIs the business cares about move? Application-specific, out of scope for these tutorials.
A high score on a higher layer requires acceptable scores on the layers below. Embedding quality (separately measured by benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard)) sets the ceiling on every downstream metric.
@@ -86,4 +87,4 @@ Wire it into CI and fail the job when recall falls below your target threshold.
## Next Steps
Once ANN recall is on target, continue with [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) to check how well those results match user intent.
Once ANN recall is on target, continue with [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/) to check how well those results match user intent.
@@ -135,7 +135,7 @@ ranking quality of search results, with higher scores indicating better performa
Binary Quantization definitely speeds up the retrieval, and make it cheaper, but also seems not to affect the quality of
the retrieval much in some cases. **However, that's something you should carefully verify on your own data**. If you are
a Qdrant user, then you can just enable quantization on an existing collection and [measure the impact on the retrieval
quality](/documentation/tutorials-search-engineering/retrieval-quality/).
quality](/documentation/tutorials-search-engineering/ann-recall/).
All the tests we did were performed using [`beir-qdrant`](https://github.com/kacperlukawski/beir-qdrant), and might be
reproduced by running [the script available on the project