diff --git a/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md b/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md index 3ff505852..d6daa777f 100644 --- a/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md +++ b/qdrant-landing/content/documentation/headless/content/tutorials/search-engineering.md @@ -6,8 +6,8 @@ | [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | Python | 45m | Intermediate | | [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | Python | 30m | Intermediate | | [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/) | Understand evaluation levels and choose the right metric. | Python | 20m | Intermediate | -| [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Generate ground truth data at scale and avoid data leakage. | Python | 40m | Intermediate | | [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN precision with the Web UI and tune HNSW parameters. | Web UI | 15m | Intermediate | +| [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Generate ground truth data at scale and avoid data leakage. | Python | 40m | Intermediate | | [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | Python | 40m | Intermediate | | [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | Python | 45m | Intermediate | | [Multivectors and Late Interaction](/documentation/tutorials-search-engineering/using-multivector-representations/) | Effective use of multivector representations. | Python | 30m | Intermediate | diff --git a/qdrant-landing/content/documentation/tutorials-lp-overview.md b/qdrant-landing/content/documentation/tutorials-lp-overview.md index 103e6060d..ac6eda5c2 100644 --- a/qdrant-landing/content/documentation/tutorials-lp-overview.md +++ b/qdrant-landing/content/documentation/tutorials-lp-overview.md @@ -69,8 +69,8 @@ partition: qdrant | [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | Python | 45m | Intermediate | | [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | Python | 30m | Intermediate | | [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/) | Understand evaluation levels and choose the right metric. | Python | 20m | Intermediate | -| [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Generate ground truth data at scale and avoid data leakage. | Python | 40m | Intermediate | | [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure ANN precision with the Web UI and tune HNSW parameters. | Web UI | 15m | Intermediate | +| [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) | Generate ground truth data at scale and avoid data leakage. | Python | 40m | Intermediate | | [Reranking for Better Search](/documentation/search-precision/reranking-semantic-search/) | Use multivector representations for better ranking. | Python | 30m | Intermediate | | [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | Python | 40m | Intermediate | | [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | Python | 45m | Intermediate | diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md index 43a7eb52f..2ae50c5d4 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md @@ -29,7 +29,7 @@ The four layers aren't measured in isolation. Teams that successfully connect re | # | Layer | What it measures | Cadence | Cost | |---|---|---|---|---| | 1 | ANN precision | `Recall@k` vs exact kNN on a sampled query set | On index or embedding changes | Low | -| 2 | Retrieval relevance | `Recall@k` / `NDCG@k` vs a labeled golden set | Weekly, or on retrieval-stack changes | Low per run; **golden set is the real cost** (see next tutorial) | +| 2 | Retrieval relevance | `Recall@k` / `NDCG@k` vs a labeled golden set | Weekly, or on retrieval-stack changes | Low per run; **golden set is the real cost**| | 3 | End-to-end answer quality | LLM-as-judge or human rating on the golden set | Weekly, or on retrieval or generator changes | Moderate (LLM-judge cost × eval size) | | 4 | Business impact | Online A/B behind a flag | Per release, once offline layers pass | High (traffic, experimentation infra) | diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md index 3c3a388e0..f655433fe 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md @@ -1,6 +1,6 @@ --- title: Building a Golden Query Set -weight: 5 +weight: 6 aliases: - /documentation/tutorials/retrieval-quality-golden-set/ --- @@ -10,23 +10,28 @@ aliases: | Time: 40 min | Level: Intermediate | | | |--------------|---------------------|--|----| -This tutorial covers **layer 2** of the evaluation ladder: **retrieval relevance**. Measuring how well retrieved results match real user intent requires a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). For layer 1 (ANN precision against exact kNN), which needs no relevance labels, use the **Search Quality** tab in the Qdrant Web UI. +This tutorial focuses on **retrieval relevance**: how well retrieved results match real user intent. +To measure retrieval relevance, you need a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). + +To evaluate other layers of your retrieval pipeline, see the evaluation ladder. ## Generating Queries -There are three practical approaches to building a golden set. They trade quality against cost and scale, so most teams use a mix. Pick the ones that match your resources and quality bar. +There are three practical approaches to building a golden set. Each one trades quality against cost and scale. -### 1. Human Annotation (Highest Quality, Highest Cost) +### 1. Human Annotation -Domain experts assign relevance scores on a binary (relevant / not relevant) or graded (0/1/2 or 1–5) scale. This is the cleanest approach for high-stakes applications and produces the graded labels NDCG needs. Expert time is the bottleneck, so reserve it for a small, high-value subset — the hardest queries or the ones that matter most commercially — and use the other two approaches for coverage. +Domain experts assign relevance scores on a binary (relevant / not relevant) or graded (0/1/2 or 1–5) scale. Human-labeled data produces the highest-fidelity signal and is the only practical source for graded labels, which ranking metrics like [Normalized Discounted Cumulative Gain (NDCG)](https://en.wikipedia.org/wiki/Discounted_cumulative_gain) use to reward relevant results appearing at higher positions. Expert time is the bottleneck, which typically limits this approach to a small set of high-value queries. -### 2. Real User Queries from Logs (High Realism, Requires Production Traffic) +### 2. Real User Queries from Logs -If your app records queries with click or explicit-feedback signals, sample query-document pairs directly. This captures real user intent and vocabulary, and should be your first choice once production traffic exists. Stratify sampling so rare-but-important cases aren't drowned out. For search-style traffic, that usually means query type or topic cluster. For RAG or agentic retrieval, it often means conversation turn or intent class. A few hundred labeled pairs can detect large metric differences; per-slice analysis or small ranking deltas need substantially more. Treat any number as a starting point and widen confidence intervals if the signal is noisy. +If your app records queries with click or explicit-feedback signals, sample query-document pairs directly. Log-based pairs reflect real user intent and vocabulary that synthetic queries cannot replicate, though the approach requires production traffic and a signal that maps to relevance. Frequent queries dominate uniform samples, so stratifying by query type, topic cluster, conversation turn, or intent class keeps rare-but-important cases represented. A few hundred labeled pairs typically detects large metric differences; per-slice analysis or small ranking deltas require substantially more. -### 3. LLM-Based Synthetic Generation (Scales Cheaply, Lowest Fidelity) +### 3. LLM-Based Synthetic Generation -When logs and reviewers aren't available, prompt an LLM to generate plausible queries for each document. This scales to thousands of pairs, but synthetic queries are easier to retrieve than what real users type. Frameworks such as Ragas provide ready-made testset generators if you want a maintained tool; the example below is a minimal prompt shape you can adapt to any model. +An LLM can generate plausible queries for each document. This scales to thousands of pairs cheaply, but synthetic queries are typically easier to retrieve than real user queries, which inflates offline scores relative to production behavior. Frameworks such as Ragas provide ready-made testset generators if you want a maintained tool. + +The example below prompts an LLM to produce short, realistic queries for each document, with the source document serving as the labeled relevant answer. ```python import os @@ -39,6 +44,8 @@ client = anthropic.Anthropic( ) def generate_queries_for_doc(doc_text: str, n: int = 3) -> list[str]: + # doc_text is one document from your corpus; iterate over the corpus + # to build the full golden set, with each source doc as the relevance label. response = client.messages.create( model=os.environ.get("ANTHROPIC_MODEL", "claude-sonnet-4-6"), max_tokens=256, @@ -65,9 +72,10 @@ Construct the labels as a `Qrels` object (a dict mapping query IDs to `{doc_id: from qdrant_client import QdrantClient from ranx import Qrels, Run, evaluate -client = QdrantClient("http://localhost:6333") +client = QdrantClient("http://localhost:6333") # or QdrantClient(url="https://.cloud.qdrant.io", api_key="...") for Qdrant Cloud def retrieval_run(golden_set: list, collection: str, k: int = 10) -> Run: + # Each entry: {"query_id": str, "query_vector": list[float], "labels": {doc_id: score}} run = {} for entry in golden_set: results = client.query_points( diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md index 61163b835..c55e80119 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md @@ -3,7 +3,7 @@ title: Measuring ANN Precision aliases: - /documentation/tutorials/retrieval-quality/ - /documentation/beginner-tutorials/retrieval-quality/ -weight: 6 +weight: 5 --- # Measuring ANN Precision @@ -11,9 +11,10 @@ weight: 6 | Time: 15 min | Level: Intermediate | | | |--------------|---------------------|--|----| -This tutorial measures **layer 1** of the evaluation ladder, **ANN precision**: the share of Qdrant's approximate nearest-neighbor top-k that appears in the exact kNN top-k. For retrieval relevance (layer 2), see the Building a Golden Query Set tutorial. +This tutorial focuses on **ANN precision**: how closely approximate nearest-neighbor (ANN) search matches exact kNN search. +To measure ANN precision, you compare Qdrant's approximate top-k against the exact kNN top-k, then tune HNSW parameters to control the precision/latency trade-off. -We'll measure Qdrant's ANN precision with `precision@k` and tune HNSW parameters to control the precision/latency trade-off. The ANN algorithm is one of several levers that shape retrieval quality in a production pipeline, alongside the embedding model, retrieval strategy (dense, sparse, hybrid, and multi-vector), filtering, and reranking. +To evaluate other layers of your retrieval pipeline, see the evaluation ladder. ## ANN Precision