diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md new file mode 100644 index 000000000..ff3881ac6 --- /dev/null +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md @@ -0,0 +1,70 @@ +--- +title: Retrieval Quality Fundamentals +weight: 4 +aliases: + - /documentation/tutorials/retrieval-quality-fundamentals/ +--- + +# Retrieval Quality Fundamentals + +| Time: 15 min | Level: Intermediate | | | +|--------------|---------------------|--|----| + +Before measuring retrieval quality, it's worth understanding what you're measuring. Retrieval quality operates at three distinct levels, and it's easy to optimize for the wrong one. + +The first level is **ANN recall**: does the approximate search return the same results as an exact nearest-neighbor search? This is a purely algorithmic question about how faithfully HNSW approximates exhaustive search. It has nothing to do with whether those results are useful to a human. + +The second level is **retrieval relevance**: of the results returned, how many are relevant to the query intent? This requires a labeled ground-truth dataset or human judgment. A pipeline can achieve near-perfect ANN recall and still surface irrelevant documents if the embeddings are a poor fit for the task. + +The third level is **business impact**: does better retrieval lead to better outcomes like lower hallucination rates in downstream LLMs, higher task-completion rates, or improved user satisfaction scores? This is what stakeholders care about, but it's the hardest to measure directly. The causal chain from a vector match to a user outcome is long and easily dominated by generator behavior, UI, and other confounders, so no single offline metric is a reliable proxy for a KPI. The next section describes how teams bridge this gap in practice. + +## Connecting the Levels in Practice + +The three levels aren't measured in isolation. Teams that successfully connect retrieval work to business outcomes tend to build an **evaluation ladder** that runs each layer at a different cadence and cost, and uses the result of each layer to decide whether to invest effort at the next: + +| Layer | Question it answers | Cadence | Cost | +|---|---|---|---| +| `Recall@k` vs exact kNN on a sampled query set | Is the ANN plumbing sound? | On index-config or embedding changes | Low, once a sample query set exists | +| `Recall@k` / `NDCG@k` vs labeled golden set | Are the right documents surfacing? | Weekly, or on retrieval-stack changes | Low per run; **building the golden set is the real cost** (see the next page) | +| End-to-end answer quality on golden set, scored by LLM-as-judge or human rating | Does the user get a correct answer? | Weekly, or on retrieval- or generator-stack changes | Moderate (LLM-judge cost per query × eval size) | +| Online A/B behind a flag | Does the business KPI move? | Per release, once offline layers pass | High (traffic allocation, experimentation infra) | + +Each layer is necessary but not sufficient. A win at layer 2 that doesn't carry through to layer 3 usually means the generator or the prompt is the bottleneck, not retrieval. This is the most useful diagnostic the ladder provides, and the reason teams shouldn't collapse layers 2 and 3 into a single score. + +**Isolate the component under test.** When end-to-end quality moves, hold one side fixed: evaluate retrieval with the generator frozen, and evaluate the generator with retrieval frozen. Without this, attribution collapses into guesswork and the ladder stops being diagnostic. + +**Proxy KPIs for teams without A/B infrastructure.** Most teams don't have the traffic or tooling to run a proper A/B. Cheap production signals that correlate with business value (click position on surfaced results, answer copy or share rate, session-level task completion, thumbs-up/down) can be instrumented long before a formal experimentation platform exists, and sit usefully between the golden-set layer and the full A/B. + +**Pre-register the decision rule.** Before running the A/B, write down what constitutes a win and what constitutes a no-ship, in terms of both the retrieval metric and the KPI. This is the highest-leverage discipline for avoiding "recall improved but the KPI didn't, the KPI is noisy, let's ship anyway" rationalization. + +**Tooling.** Qdrant owns layers 1 and 2 directly. For layer 3, the ecosystem has mature tooling like [Ragas](https://docs.ragas.io/), [Phoenix](https://phoenix.arize.com/), and [DeepEval](https://docs.confident-ai.com/) that handles LLM-as-judge scoring and offline answer-quality eval. Use those rather than rebuilding the scoring harness in-house. + +## Quality Metrics + +There are various ways to quantify the quality of semantic search. Some of them, such as [Precision@k](https://en.wikipedia.org/wiki/Evaluation_measures_(information_retrieval)#Precision_at_k), +are based on the number of relevant documents in the top-k search results. Others, such as [Mean Reciprocal Rank (MRR)](https://en.wikipedia.org/wiki/Mean_reciprocal_rank), +take into account the position of the first relevant document in the search results. [DCG and NDCG](https://en.wikipedia.org/wiki/Discounted_cumulative_gain) +metrics are, in turn, based on the relevance score of the documents. + +If we treat the search pipeline as a whole, we could use any of them. For the ANN algorithm itself, however, the natural question is much narrower: did the approximate search recover the same set of items that an exact kNN search would have returned? What approximation loses is *items* (some true nearest neighbors are missed and replaced by further ones), so the most informative metric is set-overlap against exact kNN. This is **`recall@k`**: of the true `k` nearest neighbors returned by exact search, how many does the approximate search recover? It's calculated as `|ANN results ∩ exact results| / k`. When both ANN and exact search return exactly `k` items, `recall@k` and `precision@k` are numerically identical. The community uses "recall" to stay aligned with the ANN-benchmarks convention and to make it explicit that the ground truth is the exact top-k set. + +### Choosing the Right Metric + +The right choice of metric depends on what the search pipeline does with its results, and on what ground truth is available. The table below is a starting point, not a prescription: pick the metric that matches your ground truth and your user-visible behavior. + +| Scenario | Recommended metric | Ground truth | Why | +|---|---|---|---| +| Tuning HNSW parameters | `Recall@k` | Exact kNN search | Approximation manifests as missed items from the true top-k set, so set-overlap against exact search is the quantity that changes with index parameters | +| RAG pipeline (LLM reads top-k chunks) | `Recall@k` | Labeled relevant chunks | The LLM can recover if a relevant doc is at position 3 vs 1; missing it entirely hurts more | +| Single-answer retrieval (FAQ, Q&A) | `MRR` or `Hits@1` | Labeled correct answer | The first result is what the user acts on; lower ranks matter little | +| Re-ranking or recommendation feeds | `NDCG@k` | Graded relevance labels (e.g. 0/1/2) | Order within the result list matters; a highly relevant doc at rank 5 is worse than at rank 1 | + +Note that `Recall@k` appears twice with different ground truths: against exact kNN when tuning the index, against labeled data when evaluating the pipeline end-to-end. They share a formula but answer different questions. + +On choosing `k`: set it to match actual usage. If the application shows 5 results to the user, measure `@5`. If a RAG pipeline passes 10 chunks to the LLM, measure `@10`. Reporting `@100` for a UI that surfaces 5 results makes the metric look artificially good. + +`NDCG` is worth the added complexity only when you have **graded relevance labels** (for example 0/1/2 scores per query-document pair rather than binary relevant/not-relevant) and when the downstream system benefits from fine-grained ranking. Without multi-grade annotations, the simpler metrics give a cleaner signal with less labeling overhead. + +## Next Steps + +To build a labeled dataset for relevance evaluation, see [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/). To measure and tune the ANN recall of a Qdrant collection in practice, see [Retrieval Quality Evaluation](/documentation/tutorials-search-engineering/retrieval-quality/). diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md new file mode 100644 index 000000000..68f81be9e --- /dev/null +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md @@ -0,0 +1,57 @@ +--- +title: Building a Golden Query Set +weight: 5 +aliases: + - /documentation/tutorials/retrieval-quality-golden-set/ +--- + +# Building a Golden Query Set + +| Time: 20 min | Level: Intermediate | | | +|--------------|---------------------|--|----| + +Evaluating retrieval relevance requires a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). The [Retrieval Quality Evaluation](/documentation/tutorials-search-engineering/retrieval-quality/) tutorial measures **ANN recall** against exact kNN, which needs no relevance labels. This page covers the separate task of building labeled data to measure **retrieval relevance** against real user intent. + +## Generating Queries + +The highest-fidelity source is **real user queries mined from logs**. If the application records clicks or explicit feedback, sample query-document pairs and use them directly. Reach for this first before generating synthetic data. When sampling, stratify by query type or topic cluster rather than sampling uniformly at random: a random sample over-represents frequent queries and leaves rare-but-important cases uncovered. As a rough heuristic, a few hundred labeled pairs is enough to detect large metric differences; detecting small ranking differences or slicing by query type requires substantially more. The right number depends on your effect size and query-level variance, so treat any specific number as a starting point and widen confidence intervals if the signal is noisy. + +When logs aren't available, **LLM-based synthetic generation** is a practical alternative. For each document in the corpus, prompt a capable LLM to generate a handful of queries that a user would plausibly ask if they were looking for that document. This scales cheaply to thousands of query-document pairs. + +```python +import os + +import anthropic + +client = anthropic.Anthropic() + +def generate_queries_for_doc(doc_text: str, n: int = 3) -> list[str]: + response = client.messages.create( + model=os.environ.get("ANTHROPIC_MODEL", "claude-sonnet-4-6"), + max_tokens=256, + messages=[{ + "role": "user", + "content": ( + f"Generate {n} short, realistic search queries that would lead a user to the " + f"following document. Return only the queries, one per line.\n\n{doc_text}" + ), + }], + ) + return response.content[0].text.strip().splitlines() +``` + +For high-stakes applications, **human annotation** is the cleanest approach. Domain experts rate query-document pairs on a 3 to 5 point relevance scale, which is expensive but produces the signal needed for NDCG-based evaluation. + +## Pitfalls to Watch For (Data Leakage and Friends) + +The term **data leakage** is often used loosely to cover any situation where evaluation scores come out higher than real-world performance would justify. Unlike classical ML train/test leakage, the failure modes for golden query sets are more subtle. The document that a query was generated from **must** be indexed (it's the relevant answer the evaluation expects to find), so "hold out the labeled docs from the index" is *not* the right fix. The real risks are the following. + +**Synthetic-query unrealism.** LLMs tend to paraphrase the source document's wording. The resulting queries are much easier to retrieve than what real users type, which are typically shorter, vaguer, and use different vocabulary. Offline scores on synthetic queries therefore overstate production quality. Two practical mitigations: (1) prompt the LLM to write queries "as a user who has not seen this document", and (2) anchor a sample of synthetic queries against any real queries you do have, and verify the distributions of length and specificity are comparable. + +**Embedding-model contamination.** If the embedding model was fine-tuned on (query, document) pairs that overlap with the golden set, the evaluation measures in-distribution performance and overstates real-world behavior. When using an off-the-shelf model, check its training mixture if published; when fine-tuning in-house, keep a strict split between fine-tuning data and evaluation data. + +**Near-duplicate documents.** If two documents are nearly identical, a query generated from one will retrieve the other, but the other won't be in the labels, so **precision is under-reported** because the labeled relevant set is incomplete. Deduplicate the corpus before generating labels (e.g. drop documents with cosine similarity > 0.95 to an earlier document), or label near-duplicate clusters jointly. + +**Temporal drift.** If the corpus evolves over time, evaluating with queries generated from documents that post-date the index version under test is unfair: those documents aren't there to retrieve. Pin the corpus snapshot, generate queries from that snapshot, and re-generate the golden set when the corpus changes materially. + +**Reviewer reproducibility.** Whatever approach you pick, pin the data: the corpus snapshot, the query-generation prompt, the LLM model version, and any deduplication threshold. Without this, a later "the score got worse" investigation can't tell whether the retrieval regressed or the evaluation set changed underneath it.