diff --git a/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md b/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md
index 44a4ad68c..40afe228d 100644
--- a/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md
+++ b/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md
@@ -6,8 +6,6 @@
| [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | Python | 45m | Intermediate |
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | Python | 30m | Intermediate |
| [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall with the Web UI and tune HNSW parameters. | Web UI | 15m | Beginner |
-| [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Build a labeled golden set and score retrieval relevance with ranx. | Python | 40m | Intermediate |
-| [Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | Python | 45m | Intermediate |
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | Python | 40m | Intermediate |
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | Python | 45m | Intermediate |
| [Multivectors and Late Interaction](/documentation/tutorials-search-engineering/using-multivector-representations/) | Effective use of multivector representations. | Python | 30m | Intermediate |
diff --git a/qdrant-landing/content/documentation/improve-search/_index.md b/qdrant-landing/content/documentation/improve-search/_index.md
new file mode 100644
index 000000000..bf7c6ecf1
--- /dev/null
+++ b/qdrant-landing/content/documentation/improve-search/_index.md
@@ -0,0 +1,14 @@
+---
+title: Improve Search
+weight: 1450
+partition: ecosystem
+---
+
+# Improve Search
+
+*Embedding choice, chunking strategies, and retrieval evaluation using Python ecosystem tools.*
+
+| Tutorial | Objective | Stack | Time | Level |
+| :--- | :--- | :--- | :--- | :--- |
+| [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) | Build a labeled golden set and score retrieval relevance with ranx. | Python | 40m | Intermediate |
+| [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | Python | 45m | Intermediate |
diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md b/qdrant-landing/content/documentation/improve-search/retrieval-quality-golden-set.md
similarity index 95%
rename from qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md
rename to qdrant-landing/content/documentation/improve-search/retrieval-quality-golden-set.md
index 99be593b7..b504a249d 100644
--- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md
+++ b/qdrant-landing/content/documentation/improve-search/retrieval-quality-golden-set.md
@@ -3,6 +3,7 @@ title: Measuring Retrieval Relevance
weight: 6
aliases:
- /documentation/tutorials/retrieval-quality-golden-set/
+partition: ecosystem
---
# Measuring Retrieval Relevance
@@ -13,7 +14,7 @@ aliases:
This tutorial focuses on **retrieval relevance**: how well retrieved results match real user intent.
To measure retrieval relevance, you need a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). This tutorial covers both building that dataset and running it through Qdrant to compute relevance metrics.
-This tutorial is part of a four-layer retrieval evaluation framework; see [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/#the-four-layers-of-retrieval-evaluation) for the full overview.
+Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) (does the approximate index match exact kNN?) and [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/) (does the end-to-end pipeline produce the right output?).
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + optional payload), an embedding model available to encode queries at evaluation time, and Python with `ranx` installed.
@@ -177,4 +178,4 @@ In golden sets, **data leakage** means any setup that makes offline metrics look
## Next Steps
-Once retrieval relevance is on target, the next layer is pipeline output quality: whether the full pipeline produces the right output when retrieval feeds into a consumer (LLM generator, ranker, or UI). See [Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/).
+Once retrieval relevance is on target, the next layer is pipeline output quality: whether the full pipeline produces the right output when retrieval feeds into a consumer (LLM generator, ranker, or UI). See [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/).
diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output.md b/qdrant-landing/content/documentation/improve-search/retrieval-quality-pipeline-output.md
similarity index 90%
rename from qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output.md
rename to qdrant-landing/content/documentation/improve-search/retrieval-quality-pipeline-output.md
index 41bc61f0a..29582fb36 100644
--- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output.md
+++ b/qdrant-landing/content/documentation/improve-search/retrieval-quality-pipeline-output.md
@@ -3,6 +3,7 @@ title: Evaluating Pipeline Output Quality
weight: 7
aliases:
- /documentation/tutorials/retrieval-quality-pipeline-output/
+partition: ecosystem
---
# Evaluating Pipeline Output Quality
@@ -13,9 +14,9 @@ aliases:
This tutorial focuses on **pipeline output quality**: whether the full retrieval pipeline produces the right output once retrieved results reach a consumer, most often an LLM generator in a RAG system.
To measure pipeline output quality, you run your golden set through the full pipeline, capture each `(question, retrieved_context, answer)` triple, and score the triples against judgment metrics like faithfulness, answer relevancy, and context precision.
-This tutorial is part of a four-layer retrieval evaluation framework; see [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/#the-four-layers-of-retrieval-evaluation) for the full overview.
+Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) (does the approximate index match exact kNN?) and [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) (do the top-k results match query intent?).
-**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)), LLM access for generation and judging, and Python with `ragas` installed.
+**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/)), LLM access for generation and judging, and Python with `ragas` installed.
## Wiring the RAG Pipeline
@@ -161,7 +162,7 @@ Ragas isn't the only tool in this space: Python | 45m | Intermediate |
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | Python | 30m | Intermediate |
| [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall with the Web UI and tune HNSW parameters. | Web UI | 15m | Beginner |
-| [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Build a labeled golden set and score retrieval relevance with ranx. | Python | 40m | Intermediate |
-| [Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | Python | 45m | Intermediate |
| [Reranking for Better Search](/documentation/search-precision/reranking-semantic-search/) | Use multivector representations for better ranking. | Python | 30m | Intermediate |
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | Python | 40m | Intermediate |
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | Python | 45m | Intermediate |
diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md
index 4ef7f706b..29601d7f0 100644
--- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md
+++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md
@@ -15,16 +15,16 @@ This tutorial focuses on **ANN recall**: how closely approximate nearest-neighbo
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + optional payload).
-## The Four Layers of Retrieval Evaluation
+## The Retrieval Evaluation Stack
-This tutorial is part of a four-layer retrieval evaluation framework.
+ANN recall measures how closely approximate search matches exact kNN. It's the first of four evaluation layers; each higher layer measures a different property of the retrieval system, with different tools.
-- **Layer 1: ANN recall** (this tutorial). How closely approximate nearest-neighbor search matches exact kNN.
-- **Layer 2: Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)). How well the results match query intent against a labeled dataset.
-- **Layer 3: Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/)). Whether the full pipeline (retrieval plus an LLM generator, a ranker, or a UI) produces the right output.
-- **Layer 4: Business impact**. Whether better retrieval moves the KPIs the business cares about. This layer is application-specific and out of scope for these tutorials.
+- **ANN recall** (this tutorial). Is the approximate index close to exact kNN?
+- **Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/)). Do the top-k results match query intent?
+- **Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/)). Does the end-to-end pipeline (retrieval + generator, ranker, or UI) produce the right output?
+- **Business impact**. Do the KPIs the business cares about move? Application-specific, out of scope for these tutorials.
-Retrieval quality sits on top of embedding quality, measured separately by benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard), which sets the ceiling on every downstream metric.
+A high score on a higher layer requires acceptable scores on the layers below. Embedding quality (separately measured by benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard)) sets the ceiling on every downstream metric.
## Measure ANN Recall with the Web UI
@@ -86,4 +86,4 @@ Wire it into CI and fail the job when recall falls below your target threshold.
## Next Steps
-Once ANN recall is on target, continue with [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) to check how well those results match user intent.
\ No newline at end of file
+Once ANN recall is on target, continue with [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) to check how well those results match user intent.
\ No newline at end of file