From 9fbcd1adad76798b286e38492b6a9b89a03daf82 Mon Sep 17 00:00:00 2001 From: Dylan Couzon Date: Thu, 30 Apr 2026 16:32:46 -0400 Subject: [PATCH] Rename retrieval-quality slugs to match tutorial scope Tutorial slugs now match their content: - retrieval-quality/ -> ann-recall/ (alias preserves the old published path) - retrieval-quality-golden-set/ -> retrieval-relevance/ - retrieval-quality-pipeline-output/ -> pipeline-output-quality/ Co-Authored-By: Claude Opus 4.7 (1M context) --- .../headless/content/tutorials/search-engineering.md | 2 +- .../content/documentation/improve-search/_index.md | 4 ++-- ...ality-pipeline-output.md => pipeline-output-quality.md} | 6 +++--- ...rieval-quality-golden-set.md => retrieval-relevance.md} | 4 ++-- .../content/documentation/tutorials-lp-overview.md | 2 +- .../{retrieval-quality.md => ann-recall.md} | 7 ++++--- .../tutorials-search-engineering/static-embeddings.md | 2 +- 7 files changed, 14 insertions(+), 13 deletions(-) rename qdrant-landing/content/documentation/improve-search/{retrieval-quality-pipeline-output.md => pipeline-output-quality.md} (94%) rename qdrant-landing/content/documentation/improve-search/{retrieval-quality-golden-set.md => retrieval-relevance.md} (97%) rename qdrant-landing/content/documentation/tutorials-search-engineering/{retrieval-quality.md => ann-recall.md} (91%) diff --git a/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md b/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md index 40afe228d..ee8d15d3a 100644 --- a/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md +++ b/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md @@ -5,7 +5,7 @@ | [Relevance Feedback](/documentation/tutorials-search-engineering/using-relevance-feedback/) | Relevance Feedback Retrieval in Qdrant | Python | 30m | Intermediate | | [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | Python | 45m | Intermediate | | [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | Python | 30m | Intermediate | -| [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall with the Web UI and tune HNSW parameters. | Web UI | 15m | Beginner | +| [Measuring ANN Recall](/documentation/tutorials-search-engineering/ann-recall/) | Measure ANN recall with the Web UI and tune HNSW parameters. | Web UI | 15m | Beginner | | [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | Python | 40m | Intermediate | | [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | Python | 45m | Intermediate | | [Multivectors and Late Interaction](/documentation/tutorials-search-engineering/using-multivector-representations/) | Effective use of multivector representations. | Python | 30m | Intermediate | diff --git a/qdrant-landing/content/documentation/improve-search/_index.md b/qdrant-landing/content/documentation/improve-search/_index.md index bf7c6ecf1..76bd911dd 100644 --- a/qdrant-landing/content/documentation/improve-search/_index.md +++ b/qdrant-landing/content/documentation/improve-search/_index.md @@ -10,5 +10,5 @@ partition: ecosystem | Tutorial | Objective | Stack | Time | Level | | :--- | :--- | :--- | :--- | :--- | -| [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) | Build a labeled golden set and score retrieval relevance with ranx. | Python | 40m | Intermediate | -| [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | Python | 45m | Intermediate | +| [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/) | Build a labeled golden set and score retrieval relevance with ranx. | Python | 40m | Intermediate | +| [Evaluating Pipeline Output Quality](/documentation/improve-search/pipeline-output-quality/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | Python | 45m | Intermediate | diff --git a/qdrant-landing/content/documentation/improve-search/retrieval-quality-pipeline-output.md b/qdrant-landing/content/documentation/improve-search/pipeline-output-quality.md similarity index 94% rename from qdrant-landing/content/documentation/improve-search/retrieval-quality-pipeline-output.md rename to qdrant-landing/content/documentation/improve-search/pipeline-output-quality.md index 29582fb36..3aac26446 100644 --- a/qdrant-landing/content/documentation/improve-search/retrieval-quality-pipeline-output.md +++ b/qdrant-landing/content/documentation/improve-search/pipeline-output-quality.md @@ -14,9 +14,9 @@ partition: ecosystem This tutorial focuses on **pipeline output quality**: whether the full retrieval pipeline produces the right output once retrieved results reach a consumer, most often an LLM generator in a RAG system. To measure pipeline output quality, you run your golden set through the full pipeline, capture each `(question, retrieved_context, answer)` triple, and score the triples against judgment metrics like faithfulness, answer relevancy, and context precision. -Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) (does the approximate index match exact kNN?) and [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) (do the top-k results match query intent?). +Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/ann-recall/) (does the approximate index match exact kNN?) and [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/) (do the top-k results match query intent?). -**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/)), LLM access for generation and judging, and Python with `ragas` installed. +**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/)), LLM access for generation and judging, and Python with `ragas` installed. ## Wiring the RAG Pipeline @@ -162,7 +162,7 @@ Ragas isn't the only tool in this space: FastAPI | 30m | Beginner | | [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | Python | 45m | Intermediate | | [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | Python | 30m | Intermediate | -| [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall with the Web UI and tune HNSW parameters. | Web UI | 15m | Beginner | +| [Measuring ANN Recall](/documentation/tutorials-search-engineering/ann-recall/) | Measure ANN recall with the Web UI and tune HNSW parameters. | Web UI | 15m | Beginner | | [Reranking for Better Search](/documentation/search-precision/reranking-semantic-search/) | Use multivector representations for better ranking. | Python | 30m | Intermediate | | [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | Python | 40m | Intermediate | | [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | Python | 45m | Intermediate | diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md b/qdrant-landing/content/documentation/tutorials-search-engineering/ann-recall.md similarity index 91% rename from qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md rename to qdrant-landing/content/documentation/tutorials-search-engineering/ann-recall.md index 29601d7f0..031193488 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/ann-recall.md @@ -3,6 +3,7 @@ title: Measuring ANN Recall aliases: - /documentation/tutorials/retrieval-quality/ - /documentation/beginner-tutorials/retrieval-quality/ + - /documentation/tutorials-search-engineering/retrieval-quality/ weight: 5 --- @@ -20,8 +21,8 @@ This tutorial focuses on **ANN recall**: how closely approximate nearest-neighbo ANN recall measures how closely approximate search matches exact kNN. It's the first of four evaluation layers; each higher layer measures a different property of the retrieval system, with different tools. - **ANN recall** (this tutorial). Is the approximate index close to exact kNN? -- **Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/)). Do the top-k results match query intent? -- **Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/)). Does the end-to-end pipeline (retrieval + generator, ranker, or UI) produce the right output? +- **Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/)). Do the top-k results match query intent? +- **Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/improve-search/pipeline-output-quality/)). Does the end-to-end pipeline (retrieval + generator, ranker, or UI) produce the right output? - **Business impact**. Do the KPIs the business cares about move? Application-specific, out of scope for these tutorials. A high score on a higher layer requires acceptable scores on the layers below. Embedding quality (separately measured by benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard)) sets the ceiling on every downstream metric. @@ -86,4 +87,4 @@ Wire it into CI and fail the job when recall falls below your target threshold. ## Next Steps -Once ANN recall is on target, continue with [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) to check how well those results match user intent. \ No newline at end of file +Once ANN recall is on target, continue with [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/) to check how well those results match user intent. \ No newline at end of file diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/static-embeddings.md b/qdrant-landing/content/documentation/tutorials-search-engineering/static-embeddings.md index a984f96ac..264c97a43 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/static-embeddings.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/static-embeddings.md @@ -135,7 +135,7 @@ ranking quality of search results, with higher scores indicating better performa Binary Quantization definitely speeds up the retrieval, and make it cheaper, but also seems not to affect the quality of the retrieval much in some cases. **However, that's something you should carefully verify on your own data**. If you are a Qdrant user, then you can just enable quantization on an existing collection and [measure the impact on the retrieval -quality](/documentation/tutorials-search-engineering/retrieval-quality/). +quality](/documentation/tutorials-search-engineering/ann-recall/). All the tests we did were performed using [`beir-qdrant`](https://github.com/kacperlukawski/beir-qdrant), and might be reproduced by running [the script available on the project