rename Building a Golden Query Set to Measuring Retrieval Relevance

Aligns the layer-2 tutorial title with the parallel "Measuring X" /
"Evaluating X" pattern used by the other two and maps directly to the
four-layer framework. Slug stays the same to preserve URLs and the
golden-set artifact identity in the path. Also updates the nav descriptions
to reflect the tutorial's full scope (build + score, not just build).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Dylan Couzon
2026-04-23 23:29:44 -04:00
co-authored by Claude Opus 4.7
parent e02bfb0bf1
commit 6c10996b66
5 changed files with 7 additions and 7 deletions
@@ -6,7 +6,7 @@
| [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN precision with the Web UI and tune HNSW parameters. | <span class="pill">Web UI</span> | 15m | <span class="text-green">Beginner</span> |
| [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Generate ground truth data at scale and avoid data leakage. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Build a labeled golden set and score retrieval relevance with ranx. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
@@ -69,7 +69,7 @@ partition: develop
| [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN precision with the Web UI and tune HNSW parameters. | <span class="pill">Web UI</span> | 15m | <span class="text-green">Beginner</span> |
| [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Generate ground truth data at scale and avoid data leakage. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Build a labeled golden set and score retrieval relevance with ranx. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Reranking for Better Search](/documentation/search-precision/reranking-semantic-search/) | Use multivector representations for better ranking. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
@@ -1,11 +1,11 @@
---
title: Building a Golden Query Set
title: Measuring Retrieval Relevance
weight: 6
aliases:
- /documentation/tutorials/retrieval-quality-golden-set/
---
# Building a Golden Query Set
# Measuring Retrieval Relevance
| Time: 40 min | Level: Intermediate | | |
|--------------|---------------------|--|----|
@@ -15,7 +15,7 @@ To measure pipeline output quality, you run your golden query set through the fu
For orientation on the four layers of retrieval evaluation and where this tutorial fits, see [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/#the-four-layers-of-retrieval-evaluation).
**Prerequisites.** A Qdrant collection with your corpus indexed (chunk text in a `text` payload field), a labeled golden set (see [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)), LLM access for generation and judging, and Python with `ragas` installed. The Wiring section shows the exact entry shape this tutorial expects.
**Prerequisites.** A Qdrant collection with your corpus indexed (chunk text in a `text` payload field), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)), LLM access for generation and judging, and Python with `ragas` installed. The Wiring section shows the exact entry shape this tutorial expects.
## Wiring the RAG Pipeline
@@ -19,7 +19,7 @@ To measure ANN precision, you compare Qdrant's approximate top-k against the exa
Retrieval quality operates at four layers. Each catches different failure modes at a different cadence and cost. This tutorial covers layer 1.
- **Layer 1: ANN precision** (this tutorial). How closely approximate nearest-neighbor search matches exact kNN. Run on every index or embedding change.
- **Layer 2: Retrieval relevance** ([Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)). How well the results match query intent against a labeled dataset. Run weekly, or on retrieval-stack changes.
- **Layer 2: Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)). How well the results match query intent against a labeled dataset. Run weekly, or on retrieval-stack changes.
- **Layer 3: Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/)). Whether the full pipeline (retrieval plus an LLM generator, a ranker, or a UI) produces the right output. Run weekly, or on retrieval or generator changes.
- **Layer 4: Business impact**. Whether better retrieval moves the KPIs the business cares about. Measured per release once the offline layers pass.
@@ -93,4 +93,4 @@ Wire it into CI and fail the job when precision falls below your target threshol
Measuring ANN precision keeps HNSW tuning honest. The Search Quality tab gives you a quick interactive read; the Python helper plugs into CI to catch regressions after embedding model changes or index config updates.
Once ANN precision is on target, the next layer is whether the retrieved results are relevant to users. See [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/).
Once ANN precision is on target, the next layer is whether the retrieved results are relevant to users. See [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/).