Semantics, improve clarity

This commit is contained in:
Dylan Couzon
2026-04-24 12:02:41 -04:00
parent 9f7c1f3fe7
commit cea3fd6092
3 changed files with 36 additions and 27 deletions
@@ -13,7 +13,7 @@ aliases:
This tutorial focuses on **retrieval relevance**: how well retrieved results match real user intent. This tutorial focuses on **retrieval relevance**: how well retrieved results match real user intent.
To measure retrieval relevance, you need a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). This tutorial covers both building that dataset and running it through Qdrant to compute relevance metrics. To measure retrieval relevance, you need a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). This tutorial covers both building that dataset and running it through Qdrant to compute relevance metrics.
For orientation on the four layers of retrieval evaluation and where this tutorial fits, see [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/#the-four-layers-of-retrieval-evaluation). This tutorial is part of a four-layer retrieval evaluation framework; see [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/#the-four-layers-of-retrieval-evaluation) for the full overview.
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + optional payload), an embedding model available to encode queries at evaluation time, and Python with `ranx` installed. **Prerequisites.** A Qdrant collection populated with your documents as points (vectors + optional payload), an embedding model available to encode queries at evaluation time, and Python with `ranx` installed.
@@ -141,7 +141,11 @@ metrics = evaluate(qrels, run, ["recall@10", "mrr", "ndcg@10"])
{"recall@10": 0.82, "mrr": 0.71, "ndcg@10": 0.76} {"recall@10": 0.82, "mrr": 0.71, "ndcg@10": 0.76}
``` ```
Higher is better on all three. Which metric matters most depends on what your pipeline does with results: Higher is better on all three.
### Choosing the Right Metric
Which metric matters most depends on what your pipeline does with results:
| Scenario | Recommended Metric | Why | | Scenario | Recommended Metric | Why |
|---|---|---| |---|---|---|
@@ -153,6 +157,8 @@ Higher is better on all three. Which metric matters most depends on what your pi
On choosing `k`: set it to match actual usage. If the application shows 5 results to the user, measure `@5`. If a RAG pipeline passes 10 chunks to the LLM, measure `@10`. Reporting `@100` for a UI that surfaces 5 results makes the metric look artificially good. On choosing `k`: set it to match actual usage. If the application shows 5 results to the user, measure `@5`. If a RAG pipeline passes 10 chunks to the LLM, measure `@10`. Reporting `@100` for a UI that surfaces 5 results makes the metric look artificially good.
### Re-running in CI
Re-run whenever the retrieval stack changes: new embedding model (which also requires re-embedding queries and re-indexing), new index config, or new reranker. In CI, compute `recall@10` against a fixed golden set and fail the job when the score drops below your target threshold. Re-run whenever the retrieval stack changes: new embedding model (which also requires re-embedding queries and re-indexing), new index config, or new reranker. In CI, compute `recall@10` against a fixed golden set and fail the job when the score drops below your target threshold.
## Pitfalls to Watch For ## Pitfalls to Watch For
@@ -13,9 +13,9 @@ aliases:
This tutorial focuses on **pipeline output quality**: whether the full retrieval pipeline produces the right output once retrieved results reach a consumer, most often an LLM generator in a RAG system. This tutorial focuses on **pipeline output quality**: whether the full retrieval pipeline produces the right output once retrieved results reach a consumer, most often an LLM generator in a RAG system.
To measure pipeline output quality, you run your golden set through the full pipeline, capture each `(question, retrieved_context, answer)` triple, and score the triples against judgment metrics like faithfulness, answer relevancy, and context precision. To measure pipeline output quality, you run your golden set through the full pipeline, capture each `(question, retrieved_context, answer)` triple, and score the triples against judgment metrics like faithfulness, answer relevancy, and context precision.
For orientation on the four layers of retrieval evaluation and where this tutorial fits, see [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/#the-four-layers-of-retrieval-evaluation). This tutorial is part of a four-layer retrieval evaluation framework; see [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/#the-four-layers-of-retrieval-evaluation) for the full overview.
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)), LLM access for generation and judging, and Python with `ragas` installed. The Wiring section shows the exact entry shape this tutorial expects. **Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)), LLM access for generation and judging, and Python with `ragas` installed.
## Wiring the RAG Pipeline ## Wiring the RAG Pipeline
@@ -57,7 +57,8 @@ Question:
The prompt above is a starting point; tune it for your domain: answer style, refusal behavior, whether outside knowledge is allowed, and output format. The prompt above is a starting point; tune it for your domain: answer style, refusal behavior, whether outside knowledge is allowed, and output format.
**3. Run retrieval and generation.** For each entry, retrieve the top-k chunks, pass them through the generator, and record a `SingleTurnSample`. `SingleTurnSample` is Ragas's data class for one evaluation record: question, retrieved context, generated answer, and optional reference. The example uses Anthropic, but any LLM provider works (OpenAI, Cohere, a local model). Only the `generate_answer` body changes: **3. Run retrieval and generation.** For each entry, retrieve the top-k chunks, pass them through the generator, and record a `SingleTurnSample` (Ragas's data class for one evaluation record: question, retrieved context, generated answer, and optional reference).
```python ```python
import os import os
@@ -67,6 +68,8 @@ from qdrant_client import QdrantClient
from ragas import SingleTurnSample from ragas import SingleTurnSample
client = QdrantClient("http://localhost:6333") # or QdrantClient(url="https://<id>.cloud.qdrant.io", api_key="...") for Qdrant Cloud client = QdrantClient("http://localhost:6333") # or QdrantClient(url="https://<id>.cloud.qdrant.io", api_key="...") for Qdrant Cloud
# The example uses Anthropic, but any LLM provider works.
anthropic_client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY")) anthropic_client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
@@ -109,7 +112,7 @@ def build_eval_set(golden_set: list, collection: str, k: int = 10) -> list:
return samples return samples
``` ```
The result is a list of `SingleTurnSample` objects, one per query. Each sample carries the question, retrieved contexts, generated answer, and optional reference. The list feeds directly into Ragas's `evaluate()` in the next section. If an entry's `ground_truth` is empty, Ragas silently skips that sample for metrics that need a reference (like `context_precision`). Populate it only when you'll actually score those metrics.
## Scoring with Ragas ## Scoring with Ragas
@@ -146,11 +149,15 @@ per_query = scores.to_pandas() # row-per-query scores
worst = per_query.nsmallest(10, "faithfulness") worst = per_query.nsmallest(10, "faithfulness")
``` ```
### Running in CI
If you ship retrieval changes regularly, this evaluation earns its place in CI. Running it on every change against a fixed golden set catches generator regressions from prompt edits, model swaps, or chunking changes before they reach production. The usual pattern: set a target threshold per metric and fail the job when any score drops below. If you ship retrieval changes regularly, this evaluation earns its place in CI. Running it on every change against a fixed golden set catches generator regressions from prompt edits, model swaps, or chunking changes before they reach production. The usual pattern: set a target threshold per metric and fail the job when any score drops below.
### Alternatives
**Without a golden set.** `faithfulness` and `answer_relevancy` are reference-free; swap `context_precision` for `LLMContextPrecisionWithoutReference`. You can then score synthetic queries offline or sampled production traffic live, at the cost of no fixed baseline for regression gating. **Without a golden set.** `faithfulness` and `answer_relevancy` are reference-free; swap `context_precision` for `LLMContextPrecisionWithoutReference`. You can then score synthetic queries offline or sampled production traffic live, at the cost of no fixed baseline for regression gating.
Ragas isn't the only tool in this space: <a href="https://docs.confident-ai.com/" target="_blank">DeepEval</a> has a pytest-native API that fits the CI story more directly, and teams that want full rubric control often build a small set of custom LLM-as-judge prompts instead. Ragas isn't the only tool in this space: <a href="https://docs.confident-ai.com/" target="_blank">DeepEval</a> has a pytest-native API that fits the CI story more directly.
## Isolating Retrieval vs Generation ## Isolating Retrieval vs Generation
@@ -167,6 +174,10 @@ Pair `recall@10` from the retrieval evaluation with `faithfulness` from the pipe
This split is the reason to keep retrieval and pipeline-output evaluation separate. Collapsing them into one end-to-end score tells you the pipeline moved, but not which half moved, so the next iteration becomes guesswork. This split is the reason to keep retrieval and pipeline-output evaluation separate. Collapsing them into one end-to-end score tells you the pipeline moved, but not which half moved, so the next iteration becomes guesswork.
## Non-RAG Use Cases
Ragas's metrics assume the consumer is an LLM generator. If retrieval feeds something else (a ranker, a recommendation surface, an agent, a search UI), swap the metrics to match: CTR or dwell time for a UI, graded rubrics for a ranker, task-completion rate for an agent. The method stays the same: freeze the consumer, run the golden set through the full pipeline, score the end-to-end output. Only the metric changes.
## Pitfalls to Watch For ## Pitfalls to Watch For
**Judge bias.** LLM judges reward verbose, confident, or well-formatted answers even when the underlying claim is weaker. Calibrate by running a sample of outputs through human raters and comparing; if judge and human scores disagree often, adjust the rubric or swap the judge model. **Judge bias.** LLM judges reward verbose, confident, or well-formatted answers even when the underlying claim is weaker. Calibrate by running a sample of outputs through human raters and comparing; if judge and human scores disagree often, adjust the rubric or swap the judge model.
@@ -175,11 +186,9 @@ This split is the reason to keep retrieval and pipeline-output evaluation separa
**Cost scaling.** LLM-as-judge cost grows with queries times metrics times judge calls per metric, and Ragas makes multiple judge calls per sample. A 500-query golden set with three metrics runs into the thousands of judge-model calls per run. Sample aggressively during iteration and reserve the full sweep for release candidates. **Cost scaling.** LLM-as-judge cost grows with queries times metrics times judge calls per metric, and Ragas makes multiple judge calls per sample. A 500-query golden set with three metrics runs into the thousands of judge-model calls per run. Sample aggressively during iteration and reserve the full sweep for release candidates.
**Non-RAG consumers.** Ragas metrics assume a generator output. If retrieval feeds a ranker, a recommendation surface, or a UI, swap Ragas for metrics that match the consumer: CTR and dwell time for a UI, graded rubrics for a ranker, and task-completion scores for an agent. The method (freeze the consumer, score the end-to-end output against the golden set) stays the same; only the metric changes.
## Connecting to Business Impact ## Connecting to Business Impact
Once pipeline output quality is on target, the remaining question is whether those offline wins move the KPIs the business responds to. It's the hardest measurement to get right, but a few disciplines make it manageable without a full experimentation platform. Once pipeline output quality is on target, the remaining question is whether those offline wins move the KPIs the business responds to. The standard answer is an A/B test: ship the change to a subset of users, compare their KPIs against a control group receiving the old behavior, and isolate the change's effect from unrelated drift. It's the hardest measurement to get right, but a few disciplines make it manageable without a full experimentation platform.
### Pick KPIs That Match the Product Shape ### Pick KPIs That Match the Product Shape
@@ -192,18 +201,14 @@ No single KPI captures "retrieval is working." The right one depends on what ret
| Agentic | Step count to completion, tool-selection accuracy | Cost per completed task | | Agentic | Step count to completion, tool-selection accuracy | Cost per completed task |
| Recommendations | Engagement rate, diversity-adjusted engagement | Retention, long-term value | | Recommendations | Engagement rate, diversity-adjusted engagement | Retention, long-term value |
### Pre-Register the Decision Rule ### Measurement Best Practices
Before running an A/B, write down what counts as a win, what counts as a no-ship, and which guardrails (latency, cost, slice performance, safety) can't regress. This is the highest-leverage habit for avoiding "the KPI is noisy, let's ship anyway" rationalization after results land. Three practices make A/B results more credible:
### Calibrate the Offline-Online Transfer Function - **Pre-register the decision rule.** Before the test, write down win criteria, no-ship criteria, and which guardrails (latency, cost, slice performance, safety) must not regress.
- **Calibrate offline against online.** Track how offline metric changes map to KPI changes over several launches. Three or four paired measurements usually reveal the slope (how much offline translates), the threshold (below which online noise dominates), and the lag (how long the KPI takes to move).
Pair offline metric changes with online KPI changes over a few launches. After three or four paired measurements, patterns emerge: the slope (how much offline translates to online), the threshold (below what offline delta online noise dominates), and the lag (how long after launch the KPI moves). Until calibrated, offline wins are unfalsifiable claims. - **Instrument proxy signals.** When a formal A/B isn't feasible, proxies like thumbs up/down, regeneration rate, source-click rate, copy/share actions, and session abandonment can catch large regressions and wins. Add them to production telemetry before they're needed.
### Proxy Signals When You Can't A/B
Most teams don't have a proper experimentation platform, or don't have the traffic to power a test quickly. Cheap-to-instrument signals correlate enough with value to catch big regressions and big wins: thumbs up or down on answers, regeneration rate, source-click rate, copy or share actions, and session abandonment. Build these into production telemetry before you need them.
## Wrapping Up ## Wrapping Up
That completes the four-layer retrieval evaluation stack: ANN precision, retrieval relevance, pipeline output quality, and business impact. Each layer catches a different class of regression, and running them together means retrieval changes can ship with defensible evidence at every stage. That completes the four-layer retrieval evaluation stack: ANN precision, retrieval relevance, pipeline output quality, and business impact. Each layer catches a different class of regression.
@@ -18,14 +18,14 @@ To measure ANN precision, you compare Qdrant's approximate top-k against the exa
## The Four Layers of Retrieval Evaluation ## The Four Layers of Retrieval Evaluation
Retrieval quality operates at four layers. Each catches different failure modes at a different cadence and cost. This tutorial covers layer 1. This tutorial is part of a four-layer retrieval evaluation framework.
- **Layer 1: ANN precision** (this tutorial). How closely approximate nearest-neighbor search matches exact kNN. Run on every index or embedding change. - **Layer 1: ANN precision** (this tutorial). How closely approximate nearest-neighbor search matches exact kNN. Run on every index or embedding change.
- **Layer 2: Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)). How well the results match query intent against a labeled dataset. Run weekly, or on retrieval-stack changes. - **Layer 2: Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/)). How well the results match query intent against a labeled dataset. Run weekly, or on retrieval-stack changes.
- **Layer 3: Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/)). Whether the full pipeline (retrieval plus an LLM generator, a ranker, or a UI) produces the right output. Run weekly, or on retrieval or generator changes. - **Layer 3: Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output/)). Whether the full pipeline (retrieval plus an LLM generator, a ranker, or a UI) produces the right output. Run weekly, or on retrieval or generator changes.
- **Layer 4: Business impact**. Whether better retrieval moves the KPIs the business cares about. Measured per release once the offline layers pass. - **Layer 4: Business impact**. Whether better retrieval moves the KPIs the business cares about. Measured per release once the offline layers pass.
Retrieval quality sits on top of embedding quality. Embedding quality is measured separately by benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard) and sets the ceiling on every downstream metric. Retrieval quality sits on top of embedding quality, measured separately by benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard), which sets the ceiling on every downstream metric.
## Measure ANN Precision with the Web UI ## Measure ANN Precision with the Web UI
@@ -41,7 +41,7 @@ HNSW is a hierarchical graph where each node has a set of links to other nodes.
For the full list of HNSW parameters, including on-disk storage and precision/memory trade-offs, see [Optimize Performance](/documentation/ops-optimization/optimize/). For the full list of HNSW parameters, including on-disk storage and precision/memory trade-offs, see [Optimize Performance](/documentation/ops-optimization/optimize/).
Toggle **advanced mode** in the Search Quality tab to tune these parameters inline. Raise `m` to 32 and `ef_construct` to 200, then run the evaluation again. Toggle **advanced mode** in the Search Quality tab to tune these parameters inline. To see the effect, try raising `m` and `ef_construct` (for example, to 32 and 200), then run the evaluation again.
![Search Quality advanced mode with HNSW parameters](/documentation/tutorials/retrieval-quality/search-quality-advanced.png) ![Search Quality advanced mode with HNSW parameters](/documentation/tutorials/retrieval-quality/search-quality-advanced.png)
@@ -49,7 +49,7 @@ Precision should increase at the cost of higher build time and memory.
![Search Quality results after HNSW tuning](/documentation/tutorials/retrieval-quality/search-quality-after-tuning.png) ![Search Quality results after HNSW tuning](/documentation/tutorials/retrieval-quality/search-quality-after-tuning.png)
Tune until you hit the point that matches your quality and cost targets. Tune until the balance between precision and cost matches your targets.
## Automate in CI with Python ## Automate in CI with Python
@@ -93,6 +93,4 @@ Wire it into CI and fail the job when precision falls below your target threshol
## Next Steps ## Next Steps
Measuring ANN precision keeps HNSW tuning honest. The Search Quality tab gives you a quick interactive read; the Python helper plugs into CI to catch regressions after embedding model changes or index config updates. Once ANN precision is on target, continue with [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/) to check how well those results match user intent.
Once ANN precision is on target, the next layer is whether the retrieved results are relevant to users. See [Measuring Retrieval Relevance](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/).