diff --git a/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md b/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md index 44a4ad68c..40afe228d 100644 --- a/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md +++ b/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md @@ -6,8 +6,6 @@ | [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | Python | 45m | Intermediate | | [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | Python | 30m | Intermediate | | [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall with the Web UI and tune HNSW parameters. | Web UI | 15m | Beginner | -| [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Build a labeled golden set and score retrieval relevance with ranx. | Python | 40m | Intermediate | -| [Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | Python | 45m | Intermediate | | [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | Python | 40m | Intermediate | | [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | Python | 45m | Intermediate | | [Multivectors and Late Interaction](/documentation/tutorials-search-engineering/using-multivector-representations/) | Effective use of multivector representations. | Python | 30m | Intermediate | diff --git a/qdrant-landing/content/documentation/improve-search/_index.md b/qdrant-landing/content/documentation/improve-search/_index.md new file mode 100644 index 000000000..bf7c6ecf1 --- /dev/null +++ b/qdrant-landing/content/documentation/improve-search/_index.md @@ -0,0 +1,14 @@ +--- +title: Improve Search +weight: 1450 +partition: ecosystem +--- + +# Improve Search + +*Embedding choice, chunking strategies, and retrieval evaluation using Python ecosystem tools.* + +| Tutorial | Objective | Stack | Time | Level | +| :--- | :--- | :--- | :--- | :--- | +| [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) | Build a labeled golden set and score retrieval relevance with ranx. | Python | 40m | Intermediate | +| [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | Python | 45m | Intermediate | diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md b/qdrant-landing/content/documentation/improve-search/retrieval-quality-golden-set.md similarity index 95% rename from qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md rename to qdrant-landing/content/documentation/improve-search/retrieval-quality-golden-set.md index 99be593b7..b504a249d 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md +++ b/qdrant-landing/content/documentation/improve-search/retrieval-quality-golden-set.md @@ -3,6 +3,7 @@ title: Measuring Retrieval Relevance weight: 6 aliases: - /documentation/tutorials/retrieval-quality-golden-set/ +partition: ecosystem --- # Measuring Retrieval Relevance @@ -13,7 +14,7 @@ aliases: This tutorial focuses on **retrieval relevance**: how well retrieved results match real user intent. To measure retrieval relevance, you need a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). This tutorial covers both building that dataset and running it through Qdrant to compute relevance metrics. -This tutorial is part of a four-layer retrieval evaluation framework; see [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/#the-four-layers-of-retrieval-evaluation) for the full overview. +Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) (does the approximate index match exact kNN?) and [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/) (does the end-to-end pipeline produce the right output?). **Prerequisites.** A Qdrant collection populated with your documents as points (vectors + optional payload), an embedding model available to encode queries at evaluation time, and Python with `ranx` installed. @@ -177,4 +178,4 @@ In golden sets, **data leakage** means any setup that makes offline metrics look ## Next Steps -Once retrieval relevance is on target, the next layer is pipeline output quality: whether the full pipeline produces the right output when retrieval feeds into a consumer (LLM generator, ranker, or UI). See [Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/). +Once retrieval relevance is on target, the next layer is pipeline output quality: whether the full pipeline produces the right output when retrieval feeds into a consumer (LLM generator, ranker, or UI). See [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/). diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output.md b/qdrant-landing/content/documentation/improve-search/retrieval-quality-pipeline-output.md similarity index 90% rename from qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output.md rename to qdrant-landing/content/documentation/improve-search/retrieval-quality-pipeline-output.md index 41bc61f0a..29582fb36 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output.md +++ b/qdrant-landing/content/documentation/improve-search/retrieval-quality-pipeline-output.md @@ -3,6 +3,7 @@ title: Evaluating Pipeline Output Quality weight: 7 aliases: - /documentation/tutorials/retrieval-quality-pipeline-output/ +partition: ecosystem --- # Evaluating Pipeline Output Quality @@ -13,9 +14,9 @@ aliases: This tutorial focuses on **pipeline output quality**: whether the full retrieval pipeline produces the right output once retrieved results reach a consumer, most often an LLM generator in a RAG system. To measure pipeline output quality, you run your golden set through the full pipeline, capture each `(question, retrieved_context, answer)` triple, and score the triples against judgment metrics like faithfulness, answer relevancy, and context precision. -This tutorial is part of a four-layer retrieval evaluation framework; see [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/#the-four-layers-of-retrieval-evaluation) for the full overview. +Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) (does the approximate index match exact kNN?) and [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) (do the top-k results match query intent?). -**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)), LLM access for generation and judging, and Python with `ragas` installed. +**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/)), LLM access for generation and judging, and Python with `ragas` installed. ## Wiring the RAG Pipeline @@ -161,7 +162,7 @@ Ragas isn't the only tool in this space: Python | 45m | Intermediate | | [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | Python | 30m | Intermediate | | [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN recall with the Web UI and tune HNSW parameters. | Web UI | 15m | Beginner | -| [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Build a labeled golden set and score retrieval relevance with ranx. | Python | 40m | Intermediate | -| [Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | Python | 45m | Intermediate | | [Reranking for Better Search](/documentation/search-precision/reranking-semantic-search/) | Use multivector representations for better ranking. | Python | 30m | Intermediate | | [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | Python | 40m | Intermediate | | [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | Python | 45m | Intermediate | diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md index 4ef7f706b..29601d7f0 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md @@ -15,16 +15,16 @@ This tutorial focuses on **ANN recall**: how closely approximate nearest-neighbo **Prerequisites.** A Qdrant collection populated with your documents as points (vectors + optional payload). -## The Four Layers of Retrieval Evaluation +## The Retrieval Evaluation Stack -This tutorial is part of a four-layer retrieval evaluation framework. +ANN recall measures how closely approximate search matches exact kNN. It's the first of four evaluation layers; each higher layer measures a different property of the retrieval system, with different tools. -- **Layer 1: ANN recall** (this tutorial). How closely approximate nearest-neighbor search matches exact kNN. -- **Layer 2: Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)). How well the results match query intent against a labeled dataset. -- **Layer 3: Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/)). Whether the full pipeline (retrieval plus an LLM generator, a ranker, or a UI) produces the right output. -- **Layer 4: Business impact**. Whether better retrieval moves the KPIs the business cares about. This layer is application-specific and out of scope for these tutorials. +- **ANN recall** (this tutorial). Is the approximate index close to exact kNN? +- **Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/)). Do the top-k results match query intent? +- **Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/)). Does the end-to-end pipeline (retrieval + generator, ranker, or UI) produce the right output? +- **Business impact**. Do the KPIs the business cares about move? Application-specific, out of scope for these tutorials. -Retrieval quality sits on top of embedding quality, measured separately by benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard), which sets the ceiling on every downstream metric. +A high score on a higher layer requires acceptable scores on the layers below. Embedding quality (separately measured by benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard)) sets the ceiling on every downstream metric. ## Measure ANN Recall with the Web UI @@ -86,4 +86,4 @@ Wire it into CI and fail the job when recall falls below your target threshold. ## Next Steps -Once ANN recall is on target, continue with [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) to check how well those results match user intent. \ No newline at end of file +Once ANN recall is on target, continue with [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) to check how well those results match user intent. \ No newline at end of file