Move retrieval relevance and pipeline-output tutorials to Improve Search

Splits the three retrieval-evaluation tutorials across two sections.
ANN Recall (Web UI, language-agnostic) stays under Search Engineering.
The two Python ecosystem tutorials (ranx, ragas) move to a new top-level
Improve Search section under the Ecosystem partition.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Dylan Couzon
2026-04-30 12:30:08 -04:00
co-authored by Claude Opus 4.7
parent 8ca87dc442
commit 33bd663d80
6 changed files with 30 additions and 18 deletions
@@ -0,0 +1,14 @@
---
title: Improve Search
weight: 1450
partition: ecosystem
---
# Improve Search
*Embedding choice, chunking strategies, and retrieval evaluation using Python ecosystem tools.*
| Tutorial | Objective | Stack | Time | Level |
| :--- | :--- | :--- | :--- | :--- |
| [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) | Build a labeled golden set and score retrieval relevance with ranx. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
@@ -0,0 +1,181 @@
---
title: Measuring Retrieval Relevance
weight: 6
aliases:
- /documentation/tutorials/retrieval-quality-golden-set/
partition: ecosystem
---
# Measuring Retrieval Relevance
| Time: 40 min | Level: Intermediate | | |
|--------------|---------------------|--|----|
This tutorial focuses on **retrieval relevance**: how well retrieved results match real user intent.
To measure retrieval relevance, you need a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). This tutorial covers both building that dataset and running it through Qdrant to compute relevance metrics.
Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) (does the approximate index match exact kNN?) and [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/) (does the end-to-end pipeline produce the right output?).
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + optional payload), an embedding model available to encode queries at evaluation time, and Python with `ranx` installed.
## Generating Queries
There are three practical approaches to building a golden set. Each one trades quality against cost and scale.
### 1. Human Annotation
Domain experts assign relevance scores on a binary (relevant / not relevant) or graded (0/1/2 or 1–5) scale. Human-labeled data produces the highest-fidelity signal and is the primary source for graded labels. Expert time is the bottleneck, which typically limits this approach to a small set of high-value queries.
### 2. Real User Queries from Logs
Sample query-document pairs from your production logs, using clicks or explicit feedback (thumbs up/down, ratings) as the relevance signal. Real user queries capture intent and vocabulary that synthetic generation can't match, but you need enough traffic and a signal that maps to relevance.
Balance the sample so frequent queries don't crowd out rare ones: group by query type, topic, or intent class. Start with a few hundred labeled pairs to detect large metric differences; per-slice analysis or small ranking deltas need substantially more.
### 3. LLM-Based Synthetic Generation
An LLM can generate plausible queries for documents sampled from your corpus. This scales cheaply to thousands of pairs, but synthetic queries are typically easier to retrieve than real user queries, which inflates offline scores. For very large corpora, log-based sampling is often more practical.
The document you feed the LLM (the **source document**) becomes the relevance label for every query it generates:
```text
You are helping build an evaluation dataset for a search system.
Generate 3 realistic search queries for the document below.
Each query should be what a real user would type to find it.
Phrase queries naturally, not as paraphrases of the document.
Return exactly 3 lines, one query per line. No numbering, no bullets, no preamble. Example:
how does X work
best way to configure Y
what is Z used for
Document:
{document_text}
```
**Tune the prompt to your corpus:**
- **Query style.** Questions for FAQ/RAG, keyword phrases for e-commerce, intent phrases for code search, or technical terms for specialist domains.
- **Count per document.** `3` is a default; tune to document length and golden-set size.
- **Persona.** A generic "user" works broadly; specialist corpora (medical, legal, technical) benefit from targeted personas.
- **Language.** Default English; state multilingual explicitly.
## Using the Golden Set
<a href="https://amenra.github.io/ranx/" target="_blank">ranx</a> is a Python library for ranking-metric evaluation. It covers the standard ranking metrics (`recall@k`, `MRR`, `NDCG@k`, `Precision@k`, MAP, and others) through one consistent interface, so you don't hand-roll each metric or juggle different libraries as needs grow.
The evaluation runs in three steps: load the labeled queries into the shape ranx expects, run each through Qdrant, then compute metrics.
**1. Load and assemble.** For each labeled query, build an entry with `query_id`, `query_text`, `query_vector` (embedded with the same model your Qdrant collection uses), and `labels`:
```python
{
"query_id": "q1",
"query_text": "how does X work",
"query_vector": [0.12, -0.48, 0.33, ...], # embedding of query_text
"labels": {"doc_42": 1}, # source doc for synthetic queries, relevant docs otherwise
}
```
Build the full `golden_set` by normalizing whatever your generation pipeline produced, then looping through it:
```python
from your_embedding_model import embed
# Normalize whatever your generation pipeline produced into this shape:
# - Synthetic: one item per generated query, labels = {source_doc_id: 1}
# - Logs: one item per query-click pair, labels = {clicked_doc_id: 1}
# - Human: one item per annotated query, labels = {doc_id: score, ...}
labeled_data = [
{"query_text": "how does X work", "labels": {"doc_42": 1}},
{"query_text": "what is Y used for", "labels": {"doc_55": 1, "doc_88": 1}},
# ...one entry per labeled query
]
golden_set = []
for i, item in enumerate(labeled_data):
golden_set.append({
"query_id": f"q{i}",
"query_text": item["query_text"],
"query_vector": embed(item["query_text"]),
"labels": item["labels"],
})
```
**2. Build `Qrels` and `Run`.** ranx compares two inputs, both shaped as `{query_id: {doc_id: score}}`:
- **`Qrels`** (query relevance judgments). The labeled ground truth. Use `1` for binary labels or the raw `0/1/2` for graded labels.
- **`Run`** (retrieval output). What Qdrant returned for each query, with similarity scores.
```python
from qdrant_client import QdrantClient
from ranx import Qrels, Run, evaluate
client = QdrantClient("http://localhost:6333") # or QdrantClient(url="https://<id>.cloud.qdrant.io", api_key="...") for Qdrant Cloud
def retrieval_run(golden_set: list, collection: str, k: int = 10) -> Run:
run = {}
for entry in golden_set:
results = client.query_points(
collection_name=collection,
query=entry["query_vector"],
limit=k,
).points
# p.id type must match the doc_id type in labels (ranx matches by equality).
run[entry["query_id"]] = {p.id: p.score for p in results}
return Run(run)
qrels = Qrels({entry["query_id"]: entry["labels"] for entry in golden_set})
run = retrieval_run(golden_set, collection="my_collection", k=10)
```
**3. Compute metrics.** `evaluate(qrels, run, [...])` compares the two and returns a dict of metric names to floats.
```python
metrics = evaluate(qrels, run, ["recall@10", "mrr", "ndcg@10"])
```
`evaluate()` returns:
```python
{"recall@10": 0.82, "mrr": 0.71, "ndcg@10": 0.76}
```
Higher is better on all three.
### Choosing the Right Metric
Which metric matters most depends on what your pipeline does with results:
| Scenario | Recommended Metric | Why |
|---|---|---|
| RAG pipeline (LLM reads top-k chunks) | `Recall@k` | The LLM can recover if a relevant doc is at position 3 vs 1; missing it entirely hurts more |
| Single-answer retrieval (FAQ or Q&A) | `MRR` or `Hits@1` | The first result is what the user acts on; lower ranks matter little |
| Re-ranking or recommendation feeds | `NDCG@k` | Order within the result list matters; a highly relevant doc at rank 5 is worse than at rank 1 |
[NDCG (Normalized Discounted Cumulative Gain)](https://en.wikipedia.org/wiki/Discounted_cumulative_gain) needs graded labels (for example, 0/1/2 scores per query-document pair). For binary labels, stick with `recall@k` and [`MRR` (Mean Reciprocal Rank)](https://en.wikipedia.org/wiki/Mean_reciprocal_rank). For the full metric list (Precision@k, MAP, ERR, and others), see the <a href="https://amenra.github.io/ranx/" target="_blank">ranx docs</a>.
On choosing `k`: set it to match actual usage. If the application shows 5 results to the user, measure `@5`. If a RAG pipeline passes 10 chunks to the LLM, measure `@10`. Reporting `@100` for a UI that surfaces 5 results makes the metric look artificially good.
### Re-running in CI
Re-run whenever the retrieval stack changes: new embedding model (which also requires re-embedding queries and re-indexing), new index config, or new reranker. In CI, compute `recall@10` against a fixed golden set and fail the job when the score drops below your target threshold.
## Pitfalls to Watch For
In golden sets, **data leakage** means any setup that makes offline metrics look better than production reality. Unlike classic train/test leakage, the issue is often evaluation design. Keep source documents in the index (they are the expected relevant answers). Focus on these risks:
**Synthetic-query unrealism.** LLMs often mirror source wording, creating easier queries than real user input. This inflates offline scores. Mitigate it by instructing the LLM to generate queries as a user who hasn't seen the source document, then compare synthetic and real-query distributions (length and specificity).
**Embedding-model contamination.** If your embedding model was trained on pairs overlapping with the golden set, results will look better than true generalization. For hosted models, review published training data when possible. For in-house fine-tuning, keep strict train/eval separation.
**Near-duplicate documents.** Your retrieval may return a near-duplicate of a labeled document that isn't in the label set. That makes **metrics look worse** because labels are incomplete, not because retrieval is failing. A score dip here is a signal to audit your labels before tuning retrieval. Deduplicate before labeling (for example, cosine similarity > 0.95), or label duplicate clusters together.
**Temporal drift.** If the corpus changes after labeling, labels go stale: referenced docs may be removed or superseded by newer versions. Pin a corpus snapshot for each run and regenerate the golden set after material corpus changes.
**Setup reproducibility.** Version the full evaluation setup: corpus snapshot, how labels were produced, and any preprocessing thresholds. Otherwise you can't tell whether a later score drop is model/index regression or dataset drift.
## Next Steps
Once retrieval relevance is on target, the next layer is pipeline output quality: whether the full pipeline produces the right output when retrieval feeds into a consumer (LLM generator, ranker, or UI). See [Evaluating Pipeline Output Quality](/documentation/improve-search/retrieval-quality-pipeline-output/).
@@ -0,0 +1,192 @@
---
title: Evaluating Pipeline Output Quality
weight: 7
aliases:
- /documentation/tutorials/retrieval-quality-pipeline-output/
partition: ecosystem
---
# Evaluating Pipeline Output Quality
| Time: 45 min | Level: Intermediate | | |
|--------------|---------------------|--|----|
This tutorial focuses on **pipeline output quality**: whether the full retrieval pipeline produces the right output once retrieved results reach a consumer, most often an LLM generator in a RAG system.
To measure pipeline output quality, you run your golden set through the full pipeline, capture each `(question, retrieved_context, answer)` triple, and score the triples against judgment metrics like faithfulness, answer relevancy, and context precision.
Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/retrieval-quality/) (does the approximate index match exact kNN?) and [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/) (do the top-k results match query intent?).
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-quality-golden-set/)), LLM access for generation and judging, and Python with `ragas` installed.
## Wiring the RAG Pipeline
<a href="https://docs.ragas.io/" target="_blank">Ragas</a> is a Python library that uses an LLM as a judge to score RAG outputs (rating each answer against criteria like faithfulness and relevancy). It expects samples shaped as `(question, retrieved_context, answer)` triples, so you build a fresh evaluation set from your labeled data. Three steps: prepare the evaluation data, define a grounding prompt, and run the retrieve-generate-record loop.
**1. Prepare the evaluation data.** Each entry needs a `query_id`, a `query_text` (for prompting the generator), a `query_vector` (for retrieval), and `labels`. For `context_precision` only, also include a `ground_truth` reference answer.
If your queries came from synthetic generation, they don't carry ground-truth answers natively. A simple workaround: make one more LLM pass per query, constrained to the source document, asking for a one-to-two-sentence reference answer. Skip this step if you're only scoring `faithfulness` and `answer_relevancy` (both are reference-free).
```python
# Example of an evaluation-ready entry.
{
"query_id": "q1",
"query_text": "how does X work",
"query_vector": [0.12, -0.48, 0.33, ...],
"labels": {"doc_42": 1},
"ground_truth": "...", # optional; required for context_precision only
}
```
Different golden-set sources (human annotation, log sampling, or LLM synthesis) produce different raw shapes. Normalize to this structure before running the loop.
**2. Define the grounding prompt.** The prompt is the seam between retrieval and generation. Keep it in a versioned string so you can swap models without touching the evaluation code:
```python
PROMPT_TEMPLATE = """You are answering questions using retrieved source material.
Answer the question below using only the provided context.
If the context does not contain the answer, say so explicitly.
Do not rely on outside knowledge.
Context:
{retrieved_context}
Question:
{query_text}
"""
```
The prompt above is a starting point; tune it for your domain: answer style, refusal behavior, whether outside knowledge is allowed, and output format.
**3. Run retrieval and generation.** For each entry, retrieve the top-k chunks, pass them through the generator, and record a `SingleTurnSample` (Ragas's data class for one evaluation record: question, retrieved context, generated answer, and optional reference).
```python
import os
import anthropic
from qdrant_client import QdrantClient
from ragas import SingleTurnSample
client = QdrantClient("http://localhost:6333") # or QdrantClient(url="https://<id>.cloud.qdrant.io", api_key="...") for Qdrant Cloud
# The example uses Anthropic, but any LLM provider works.
anthropic_client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
def generate_answer(query_text: str, contexts: list) -> str:
"""Fill the prompt template with context + question, then call the LLM."""
prompt = PROMPT_TEMPLATE.format(
retrieved_context="\n\n".join(contexts),
query_text=query_text,
)
response = anthropic_client.messages.create(
model=os.environ.get("ANTHROPIC_MODEL", "claude-sonnet-4-6"),
max_tokens=512,
messages=[{"role": "user", "content": prompt}],
)
return response.content[0].text
def build_eval_set(golden_set: list, collection: str, k: int = 10) -> list:
"""For each labeled query: retrieve from Qdrant, generate an answer, package as a Ragas sample."""
samples = []
for entry in golden_set:
# Retrieve top-k chunks from Qdrant.
results = client.query_points(
collection_name=collection,
query=entry["query_vector"],
limit=k,
).points
contexts = [p.payload["text"] for p in results] # adjust the payload key to match your schema
# Generate an answer grounded in those chunks.
answer = generate_answer(entry["query_text"], contexts)
# Package into a Ragas sample: question, context, answer, optional reference.
samples.append(SingleTurnSample(
user_input=entry["query_text"],
retrieved_contexts=contexts,
response=answer,
reference=entry.get("ground_truth", ""),
))
return samples
```
If an entry's `ground_truth` is empty, Ragas silently skips that sample for metrics that need a reference (like `context_precision`). Populate it only when you'll actually score those metrics.
## Scoring with Ragas
Three Ragas metrics cover the common failure modes for pipeline output quality:
- **`faithfulness`** checks whether the answer only makes claims supported by the retrieved context. It drops when the generator hallucinates or uses its training knowledge instead of the retrieved context.
- **`answer_relevancy`** checks whether the answer addresses the question. It drops when the generator pads, dodges, or drifts off-topic.
- **`context_precision`** checks whether the retrieved chunks are relevant to the ground-truth answer and ranked highly. It drops when retrieval surfaces noise that crowds out the useful chunks. `context_precision` compares against the `reference` field, so it only scores queries that carry a ground-truth answer.
Pass the eval samples into `evaluate()` with those three metrics:
```python
from ragas import EvaluationDataset, evaluate
from ragas.metrics.collections import faithfulness, answer_relevancy, context_precision
dataset = EvaluationDataset(samples=samples)
scores = evaluate(
dataset,
metrics=[faithfulness, answer_relevancy, context_precision],
# For a non-OpenAI judge, pass llm= and embeddings= (see Ragas docs).
)
```
`evaluate()` returns an `EvaluationResult` object. Its aggregate scores print like this:
```python
{"faithfulness": 0.88, "answer_relevancy": 0.81, "context_precision": 0.74}
```
Higher is better on all three. Aggregates hide the distribution that tells you what's breaking, so drop into the per-query view to find the worst-scoring samples:
```python
per_query = scores.to_pandas() # row-per-query scores
worst = per_query.nsmallest(10, "faithfulness")
```
### Running in CI
If you ship retrieval changes regularly, this evaluation earns its place in CI. Running it on every change against a fixed golden set catches generator regressions from prompt edits, model swaps, or chunking changes before they reach production. The usual pattern: set a target threshold per metric and fail the job when any score drops below.
### Alternatives
**Without a golden set.** `faithfulness` and `answer_relevancy` are reference-free; swap `context_precision` for `LLMContextPrecisionWithoutReference`. You can then score synthetic queries offline or sampled production traffic live, at the cost of no fixed baseline for regression gating.
Ragas isn't the only tool in this space: <a href="https://docs.confident-ai.com/" target="_blank">DeepEval</a> has a pytest-native API that fits the CI story more directly.
## Isolating Retrieval vs Generation
If you're also running [retrieval evaluation](/documentation/improve-search/retrieval-quality-golden-set/) against the same golden set, pairing the two scores on every run gives a diagnostic 2x2 for attributing score changes. When a metric drops after a change (new embedding model, new prompt, or new chunking strategy), the pair tells you which half of the pipeline to investigate.
Pair `recall@10` from the retrieval evaluation with `faithfulness` from the pipeline-output evaluation. In the table, High and Low are relative to the target thresholds you set per metric.
| Recall@10 | Faithfulness | Diagnosis |
|---|---|---|
| High | High | Ready to ship. |
| High | Low | Generator or prompt problem. Retrieval is surfacing the right context; something downstream (prompt, model, or temperature) is misusing it. |
| Low | Low | Fix retrieval first. The generator can't be faithful to context it never saw. |
| Low | High | Rare. Usually means either the golden-set labels are incomplete (retrieval found useful docs the label set doesn't cover) or the generator punted with a non-committal answer that has no claims to fail on. Read a sample of per-query outputs before acting. |
This split is the reason to keep retrieval and pipeline-output evaluation separate. Collapsing them into one end-to-end score tells you the pipeline moved, but not which half moved, so the next iteration becomes guesswork.
## Non-RAG Use Cases
Ragas's metrics assume the consumer is an LLM generator. If retrieval feeds something else (a ranker, a recommendation surface, an agent, a search UI), swap the metrics to match: CTR or dwell time for a UI, graded rubrics for a ranker, task-completion rate for an agent. The method stays the same: freeze the consumer, run the golden set through the full pipeline, score the end-to-end output. Only the metric changes.
## Pitfalls to Watch For
**Judge bias.** LLM judges reward verbose, confident, or well-formatted answers even when the underlying claim is weaker. Calibrate by running a sample of outputs through human raters and comparing; if judge and human scores disagree often, adjust the rubric or swap the judge model.
**Self-judging contamination.** Using the same model to generate and to judge inflates scores because the judge recognizes and rewards its own output style. Pick a different model family for the judge than for the generator, and record both versions in every run so score shifts can't be blamed on a silent upgrade.
**Cost scaling.** LLM-as-judge cost grows with queries times metrics times judge calls per metric, and Ragas makes multiple judge calls per sample. A 500-query golden set with three metrics runs into the thousands of judge-model calls per run. Sample aggressively during iteration and reserve the full sweep for release candidates.
## Wrapping Up
You now have a Ragas-based scoring loop for the full RAG pipeline, a 2x2 to attribute regressions to retrieval or generation, and a CI pattern to gate releases on `faithfulness`, `answer_relevancy`, and `context_precision`.