mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-02 17:38:31 +02:00
Merge pull request #2305 from qdrant/retrival-quality-guide-improvement
docs(tutorials): split retrieval-quality into three evaluation layers
This commit is contained in:
+1
-1
@@ -5,7 +5,7 @@
|
|||||||
| [Relevance Feedback](/documentation/tutorials-search-engineering/using-relevance-feedback/) | Relevance Feedback Retrieval in Qdrant | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
| [Relevance Feedback](/documentation/tutorials-search-engineering/using-relevance-feedback/) | Relevance Feedback Retrieval in Qdrant | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
||||||
| [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
|
| [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
|
||||||
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
||||||
| [Retrieval Quality Evaluation](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure quality and tune HNSW parameters. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
| [Measuring ANN Recall](/documentation/tutorials-search-engineering/ann-recall/) | Measure ANN recall with the Web UI and tune HNSW parameters. | <span class="pill">Web UI</span> | 15m | <span class="text-green">Beginner</span> |
|
||||||
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
|
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
|
||||||
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
|
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
|
||||||
| [Multivectors and Late Interaction](/documentation/tutorials-search-engineering/using-multivector-representations/) | Effective use of multivector representations. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
| [Multivectors and Late Interaction](/documentation/tutorials-search-engineering/using-multivector-representations/) | Effective use of multivector representations. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
||||||
|
|||||||
@@ -0,0 +1,14 @@
|
|||||||
|
---
|
||||||
|
title: Improve Search
|
||||||
|
weight: 1450
|
||||||
|
partition: ecosystem
|
||||||
|
---
|
||||||
|
|
||||||
|
# Improve Search
|
||||||
|
|
||||||
|
*Embedding choice, chunking strategies, and retrieval evaluation using Python ecosystem tools.*
|
||||||
|
|
||||||
|
| Tutorial | Objective | Stack | Time | Level |
|
||||||
|
| :--- | :--- | :--- | :--- | :--- |
|
||||||
|
| [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/) | Build a labeled golden set and score retrieval relevance with ranx. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
|
||||||
|
| [Evaluating Pipeline Output Quality](/documentation/improve-search/pipeline-output-quality/) | Score a RAG pipeline with Ragas and isolate retrieval vs generation failures. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
|
||||||
@@ -0,0 +1,206 @@
|
|||||||
|
---
|
||||||
|
title: Evaluating Pipeline Output Quality
|
||||||
|
weight: 7
|
||||||
|
aliases:
|
||||||
|
- /documentation/tutorials/retrieval-quality-pipeline-output/
|
||||||
|
partition: ecosystem
|
||||||
|
---
|
||||||
|
|
||||||
|
# Evaluating Pipeline Output Quality
|
||||||
|
|
||||||
|
| Time: 45 min | Level: Intermediate | | |
|
||||||
|
|--------------|---------------------|--|----|
|
||||||
|
|
||||||
|
This tutorial focuses on **pipeline output quality**: whether the full retrieval pipeline produces the right output once retrieved results reach a consumer, most often an LLM generator in a RAG system.
|
||||||
|
To measure pipeline output quality, you run your golden set through the full pipeline, capture each `(question, retrieved_context, answer)` triple, and score the triples against judgment metrics like faithfulness, answer relevancy, and context precision.
|
||||||
|
|
||||||
|
Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/ann-recall/) (does the approximate index match exact kNN?) and [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/) (do the top-k results match query intent?).
|
||||||
|
|
||||||
|
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + a `text` payload field for the chunk content), a labeled golden set (see [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/)), LLM access for generation and judging, and Python with `ragas` installed.
|
||||||
|
|
||||||
|
## Wiring the RAG Pipeline
|
||||||
|
|
||||||
|
Several frameworks score RAG outputs with an LLM judge, including Ragas, <a href="https://docs.confident-ai.com/" target="_blank">DeepEval</a>, and others. We use Ragas here because it's the lightest setup for the three metrics this tutorial covers. If your team has standardized on a different framework or prefers to call the judge LLM directly, the same workflow applies.
|
||||||
|
|
||||||
|
<a href="https://docs.ragas.io/" target="_blank">Ragas</a> is a Python library that uses an LLM as a judge to score RAG outputs (rating each answer against criteria like faithfulness and relevancy). It expects samples shaped as `(question, retrieved_context, answer)` triples, so you build a fresh evaluation set from your labeled data. Three steps: prepare the evaluation data, define a grounding prompt, and run the retrieve-generate-record loop.
|
||||||
|
|
||||||
|
**1. Prepare the evaluation data.** Each entry needs a `query_id`, a `query_text` (used for both prompting the generator and embedding for retrieval), and `labels`. For `context_precision` only, also include a `ground_truth` reference answer.
|
||||||
|
|
||||||
|
Synthetic queries don't ship with ground-truth answers. If you're only scoring `faithfulness` and `answer_relevancy`, skip this step since both are reference-free. Otherwise, generate references by running each query through an LLM scoped to its source document. Have the model return `NO_ANSWER` when the source can't answer, and drop those rows before scoring, or `context_precision` ends up judging retrieval against a reference the source doc doesn't support.
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Example of an evaluation-ready entry.
|
||||||
|
{
|
||||||
|
"query_id": "q1",
|
||||||
|
"query_text": "how does X work",
|
||||||
|
"labels": {"doc_42": 1},
|
||||||
|
"ground_truth": "...", # optional; required for context_precision only
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Different golden-set sources (human annotation, log sampling, or LLM synthesis) produce different raw shapes. Normalize to this structure before running the loop.
|
||||||
|
|
||||||
|
**2. Define the grounding prompt.** The prompt is the seam between retrieval and generation. Keep it in a versioned string so you can swap models without touching the evaluation code:
|
||||||
|
|
||||||
|
```python
|
||||||
|
PROMPT_TEMPLATE = """You are answering questions using retrieved source material.
|
||||||
|
|
||||||
|
Answer the question below using only the provided context.
|
||||||
|
If the context does not contain the answer, say so explicitly.
|
||||||
|
Do not rely on outside knowledge.
|
||||||
|
|
||||||
|
Context:
|
||||||
|
{retrieved_context}
|
||||||
|
|
||||||
|
Question:
|
||||||
|
{query_text}
|
||||||
|
"""
|
||||||
|
```
|
||||||
|
|
||||||
|
The prompt above is a starting point; tune it for your domain: answer style, refusal behavior, whether outside knowledge is allowed, and output format.
|
||||||
|
|
||||||
|
**3. Run retrieval and generation.** For each entry, retrieve the top-k chunks, pass them through the generator, and record a `SingleTurnSample` (Ragas's data class for one evaluation record: question, retrieved context, generated answer, and optional reference).
|
||||||
|
|
||||||
|
|
||||||
|
```python
|
||||||
|
import os
|
||||||
|
|
||||||
|
import anthropic
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
from ragas import SingleTurnSample
|
||||||
|
|
||||||
|
from your_embedding_model import embed # must match the model your Qdrant collection uses
|
||||||
|
|
||||||
|
client = QdrantClient("http://localhost:6333") # or QdrantClient(url="https://<id>.cloud.qdrant.io", api_key="...") for Qdrant Cloud
|
||||||
|
|
||||||
|
# The example uses Anthropic, but any LLM provider works.
|
||||||
|
anthropic_client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
|
||||||
|
|
||||||
|
|
||||||
|
def generate_answer(query_text: str, contexts: list) -> str:
|
||||||
|
"""Fill the prompt template with context + question, then call the LLM."""
|
||||||
|
prompt = PROMPT_TEMPLATE.format(
|
||||||
|
retrieved_context="\n\n".join(contexts),
|
||||||
|
query_text=query_text,
|
||||||
|
)
|
||||||
|
response = anthropic_client.messages.create(
|
||||||
|
model=os.environ.get("ANTHROPIC_MODEL", "claude-sonnet-4-6"),
|
||||||
|
max_tokens=512,
|
||||||
|
messages=[{"role": "user", "content": prompt}],
|
||||||
|
)
|
||||||
|
return response.content[0].text
|
||||||
|
|
||||||
|
|
||||||
|
def build_eval_set(golden_set: list, collection: str, k: int = 10) -> list:
|
||||||
|
"""For each labeled query: retrieve from Qdrant, generate an answer, package as a Ragas sample."""
|
||||||
|
samples = []
|
||||||
|
for entry in golden_set:
|
||||||
|
# Retrieve top-k chunks from Qdrant.
|
||||||
|
results = client.query_points(
|
||||||
|
collection_name=collection,
|
||||||
|
query=embed(entry["query_text"]),
|
||||||
|
limit=k,
|
||||||
|
).points
|
||||||
|
contexts = [p.payload["text"] for p in results] # adjust the payload key to match your schema
|
||||||
|
|
||||||
|
# Generate an answer grounded in those chunks.
|
||||||
|
answer = generate_answer(entry["query_text"], contexts)
|
||||||
|
|
||||||
|
# Package into a Ragas sample: question, context, answer, optional reference.
|
||||||
|
samples.append(SingleTurnSample(
|
||||||
|
user_input=entry["query_text"],
|
||||||
|
retrieved_contexts=contexts,
|
||||||
|
response=answer,
|
||||||
|
reference=entry.get("ground_truth", ""),
|
||||||
|
))
|
||||||
|
return samples
|
||||||
|
```
|
||||||
|
|
||||||
|
If an entry's `ground_truth` is empty, Ragas silently skips that sample for metrics that need a reference (like `context_precision`). Populate it only when you'll actually score those metrics.
|
||||||
|
|
||||||
|
## Scoring with Ragas
|
||||||
|
|
||||||
|
Three Ragas metrics cover the common failure modes for pipeline output quality:
|
||||||
|
|
||||||
|
- **`faithfulness`** checks whether the answer only makes claims supported by the retrieved context. It drops when the generator hallucinates or uses its training knowledge instead of the retrieved context.
|
||||||
|
- **`answer_relevancy`** checks whether the answer addresses the question. It drops when the generator pads, dodges, or drifts off-topic.
|
||||||
|
- **`context_precision`** checks whether the retrieved chunks are relevant to the ground-truth answer and ranked highly. It drops when retrieval surfaces noise that crowds out the useful chunks. `context_precision` compares against the `reference` field, so it only scores queries that carry a ground-truth answer.
|
||||||
|
|
||||||
|
Pass the eval samples into `evaluate()` with those three metrics:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from anthropic import Anthropic
|
||||||
|
from openai import OpenAI
|
||||||
|
from ragas import EvaluationDataset, evaluate
|
||||||
|
from ragas.embeddings.base import embedding_factory
|
||||||
|
from ragas.llms import llm_factory
|
||||||
|
from ragas.metrics.collections import AnswerRelevancy, ContextPrecision, Faithfulness
|
||||||
|
|
||||||
|
# Judge LLM, Use a different LLM family as the generator to avoid self-evaluation bias
|
||||||
|
judge_client = OpenAI() # reads OPENAI_API_KEY from the environment
|
||||||
|
judge_llm = llm_factory("gpt-5.4", client=judge_client)
|
||||||
|
|
||||||
|
# This is the judge's question-similarity check; it does not need to match the retrieval embedder.
|
||||||
|
judge_embeddings = embedding_factory("openai", model="text-embedding-3-large", client=judge_client)
|
||||||
|
|
||||||
|
metrics = [
|
||||||
|
Faithfulness(llm=judge_llm),
|
||||||
|
AnswerRelevancy(llm=judge_llm, embeddings=judge_embeddings),
|
||||||
|
ContextPrecision(llm=judge_llm),
|
||||||
|
]
|
||||||
|
|
||||||
|
dataset = EvaluationDataset(samples=samples)
|
||||||
|
scores = evaluate(dataset, metrics=metrics)
|
||||||
|
```
|
||||||
|
|
||||||
|
`evaluate()` returns an `EvaluationResult` object. Its aggregate scores print like this:
|
||||||
|
|
||||||
|
```python
|
||||||
|
{"faithfulness": 0.88, "answer_relevancy": 0.81, "context_precision": 0.74}
|
||||||
|
```
|
||||||
|
|
||||||
|
Higher is better on all three. Aggregates hide the distribution that tells you what's breaking, so drop into the per-query view to find the worst-scoring samples:
|
||||||
|
|
||||||
|
```python
|
||||||
|
per_query = scores.to_pandas() # row-per-query scores
|
||||||
|
worst = per_query.nsmallest(10, "faithfulness")
|
||||||
|
```
|
||||||
|
|
||||||
|
### Running in CI
|
||||||
|
|
||||||
|
If you ship retrieval changes regularly, this evaluation earns its place in CI. Running it on every change against a fixed golden set catches generator regressions from prompt edits, model swaps, or chunking changes before they reach production. The usual pattern: set a target threshold per metric and fail the job when any score drops below.
|
||||||
|
|
||||||
|
### Alternatives
|
||||||
|
|
||||||
|
**Without a golden set.** `faithfulness` and `answer_relevancy` are reference-free; swap `context_precision` for `LLMContextPrecisionWithoutReference`. You can then score synthetic queries offline or sampled production traffic live, at the cost of no fixed baseline for regression gating.
|
||||||
|
|
||||||
|
## Isolating Retrieval vs Generation
|
||||||
|
|
||||||
|
If you're also running [retrieval evaluation](/documentation/improve-search/retrieval-relevance/) against the same golden set, pairing the two scores on every run gives a diagnostic 2x2 for attributing score changes. When a metric drops after a change (new embedding model, new prompt, or new chunking strategy), the pair tells you which half of the pipeline to investigate.
|
||||||
|
|
||||||
|
Pair `recall@10` from the retrieval evaluation with `faithfulness` from the pipeline-output evaluation. In the table, High and Low are relative to the target thresholds you set per metric.
|
||||||
|
|
||||||
|
| Recall@10 | Faithfulness | Diagnosis |
|
||||||
|
|---|---|---|
|
||||||
|
| High | High | Ready to ship. |
|
||||||
|
| High | Low | Generator or prompt problem. Retrieval is surfacing the right context; something downstream (prompt, model, or temperature) is misusing it. |
|
||||||
|
| Low | Low | Fix retrieval first. The generator can't be faithful to context it never saw. |
|
||||||
|
| Low | High | Rare. Usually means either the golden-set labels are incomplete (retrieval found useful docs the label set doesn't cover) or the generator punted with a non-committal answer that has no claims to fail on. Read a sample of per-query outputs before acting. |
|
||||||
|
|
||||||
|
This split is the reason to keep retrieval and pipeline-output evaluation separate. Collapsing them into one end-to-end score tells you the pipeline moved, but not which half moved, so the next iteration becomes guesswork.
|
||||||
|
|
||||||
|
## Non-RAG Use Cases
|
||||||
|
|
||||||
|
Ragas's metrics assume the consumer is an LLM generator. If retrieval feeds something else (a ranker, a recommendation surface, an agent, a search UI), swap the metrics to match: CTR or dwell time for a UI, graded rubrics for a ranker, task-completion rate for an agent. The method stays the same: freeze the consumer, run the golden set through the full pipeline, score the end-to-end output. Only the metric changes.
|
||||||
|
|
||||||
|
## Pitfalls to Watch For
|
||||||
|
|
||||||
|
**Judge bias.** LLM judges reward verbose, confident, or well-formatted answers even when the underlying claim is weaker. Calibrate by running a sample of outputs through human raters and comparing; if judge and human scores disagree often, adjust the rubric or swap the judge model.
|
||||||
|
|
||||||
|
**Self-judging contamination.** Using the same model to generate and to judge inflates scores because the judge recognizes and rewards its own output style. Pick a different model family for the judge than for the generator, and record both versions in every run so score shifts can't be blamed on a silent upgrade.
|
||||||
|
|
||||||
|
**Cost scaling.** LLM-as-judge cost grows with queries times metrics times judge calls per metric, and Ragas makes multiple judge calls per sample. A 500-query golden set with three metrics runs into the thousands of judge-model calls per run. Sample 50 to 100 queries with a cheap judge (`claude-haiku-4-5` or `gpt-4o-mini`) during iteration; reserve the full sweep with the strong judge for release candidates.
|
||||||
|
|
||||||
|
## Wrapping Up
|
||||||
|
|
||||||
|
You now have a Ragas-based scoring loop for the full RAG pipeline, a 2x2 to attribute regressions to retrieval or generation, and a CI pattern to gate releases on `faithfulness`, `answer_relevancy`, and `context_precision`.
|
||||||
@@ -0,0 +1,179 @@
|
|||||||
|
---
|
||||||
|
title: Measuring Retrieval Relevance
|
||||||
|
weight: 6
|
||||||
|
aliases:
|
||||||
|
- /documentation/tutorials/retrieval-quality-golden-set/
|
||||||
|
partition: ecosystem
|
||||||
|
---
|
||||||
|
|
||||||
|
# Measuring Retrieval Relevance
|
||||||
|
|
||||||
|
| Time: 40 min | Level: Intermediate | | |
|
||||||
|
|--------------|---------------------|--|----|
|
||||||
|
|
||||||
|
This tutorial focuses on **retrieval relevance**: how well retrieved results match real user intent.
|
||||||
|
To measure retrieval relevance, you need a labeled dataset of queries paired with their expected relevant documents (commonly called a *golden query set* or *ground truth*). This tutorial covers both building that dataset and running it through Qdrant to compute relevance metrics.
|
||||||
|
|
||||||
|
Two related tutorials cover the other retrieval-evaluation concerns: [Measuring ANN Recall](/documentation/tutorials-search-engineering/ann-recall/) (does the approximate index match exact kNN?) and [Evaluating Pipeline Output Quality](/documentation/improve-search/pipeline-output-quality/) (does the end-to-end pipeline produce the right output?).
|
||||||
|
|
||||||
|
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + optional payload), an embedding model available to encode queries at evaluation time, and Python with `ranx` installed.
|
||||||
|
|
||||||
|
## Generating Queries
|
||||||
|
|
||||||
|
There are three practical approaches to building a golden set. Each one trades quality against cost and scale.
|
||||||
|
|
||||||
|
### 1. Human Annotation
|
||||||
|
|
||||||
|
Domain experts assign relevance scores on a binary (relevant / not relevant) or graded (0/1/2 or 1–5) scale. Human-labeled data produces the highest-fidelity signal and is the primary source for graded labels. Expert time is the bottleneck, which typically limits this approach to a small set of high-value queries.
|
||||||
|
|
||||||
|
### 2. Real User Queries from Logs
|
||||||
|
|
||||||
|
Sample query-document pairs from your production logs, using clicks or explicit feedback (thumbs up/down, ratings) as the relevance signal. Real user queries capture intent and vocabulary that synthetic generation can't match, but you need enough traffic and a signal that maps to relevance.
|
||||||
|
|
||||||
|
Balance the sample so frequent queries don't crowd out rare ones: group by query type, topic, or intent class. Start with a few hundred labeled pairs to detect large metric differences; per-slice analysis or small ranking deltas need substantially more.
|
||||||
|
|
||||||
|
### 3. LLM-Based Synthetic Generation
|
||||||
|
|
||||||
|
An LLM can generate plausible queries for documents sampled from your corpus. This scales cheaply to thousands of pairs, but synthetic queries are typically easier to retrieve than real user queries, which inflates offline scores. For very large corpora, log-based sampling is often more practical.
|
||||||
|
|
||||||
|
The document you feed the LLM (the **source document**) becomes the relevance label for every query it generates:
|
||||||
|
|
||||||
|
```text
|
||||||
|
You are helping build an evaluation dataset for a search system.
|
||||||
|
|
||||||
|
Generate 3 realistic search queries for the document below.
|
||||||
|
Each query should be what a real user would type to find it.
|
||||||
|
Phrase queries naturally, not as paraphrases of the document.
|
||||||
|
|
||||||
|
Return exactly 3 lines, one query per line. No numbering, no bullets, no preamble. Example:
|
||||||
|
how does X work
|
||||||
|
best way to configure Y
|
||||||
|
what is Z used for
|
||||||
|
|
||||||
|
Document:
|
||||||
|
{document_text}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Tune the prompt to your corpus:**
|
||||||
|
|
||||||
|
- **Query style.** Questions for FAQ/RAG, keyword phrases for e-commerce, intent phrases for code search, or technical terms for specialist domains.
|
||||||
|
- **Count per document.** `3` is a default; tune to document length and golden-set size.
|
||||||
|
- **Persona.** A generic "user" works broadly; specialist corpora (medical, legal, technical) benefit from targeted personas.
|
||||||
|
- **Language.** Default English; state multilingual explicitly.
|
||||||
|
|
||||||
|
## Using the Golden Set
|
||||||
|
|
||||||
|
<a href="https://amenra.github.io/ranx/" target="_blank">ranx</a> is a Python library for ranking-metric evaluation. It covers the standard ranking metrics (`recall@k`, `MRR`, `NDCG@k`, `Precision@k`, MAP, and others) through one consistent interface, so you don't hand-roll each metric or juggle different libraries as needs grow.
|
||||||
|
|
||||||
|
The evaluation runs in three steps: load the labeled queries into the shape ranx expects, run each through Qdrant, then compute metrics.
|
||||||
|
|
||||||
|
**1. Load and assemble.** For each labeled query, build an entry with `query_id`, `query_text`, and `labels`:
|
||||||
|
|
||||||
|
```python
|
||||||
|
{
|
||||||
|
"query_id": "q1",
|
||||||
|
"query_text": "how does X work",
|
||||||
|
"labels": {"doc_42": 1}, # source doc for synthetic queries, relevant docs otherwise
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Build the full `golden_set` by normalizing whatever your generation pipeline produced, then looping through it:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Normalize whatever your generation pipeline produced into this shape:
|
||||||
|
# - Synthetic: one item per generated query, labels = {source_doc_id: 1}
|
||||||
|
# - Logs: one item per query-click pair, labels = {clicked_doc_id: 1}
|
||||||
|
# - Human: one item per annotated query, labels = {doc_id: score, ...}
|
||||||
|
labeled_data = [
|
||||||
|
{"query_text": "how does X work", "labels": {"doc_42": 1}},
|
||||||
|
{"query_text": "what is Y used for", "labels": {"doc_55": 1, "doc_88": 1}},
|
||||||
|
# ...one entry per labeled query
|
||||||
|
]
|
||||||
|
|
||||||
|
golden_set = []
|
||||||
|
for i, item in enumerate(labeled_data):
|
||||||
|
golden_set.append({
|
||||||
|
"query_id": f"q{i}",
|
||||||
|
"query_text": item["query_text"],
|
||||||
|
"labels": item["labels"],
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
**2. Build `Qrels` and `Run`.** ranx compares two inputs, both shaped as `{query_id: {doc_id: score}}`:
|
||||||
|
|
||||||
|
- **`Qrels`** (query relevance judgments). The labeled ground truth. Use `1` for binary labels or the raw `0/1/2` for graded labels.
|
||||||
|
- **`Run`** (retrieval output). What Qdrant returned for each query, with similarity scores.
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
from ranx import Qrels, Run, evaluate
|
||||||
|
|
||||||
|
from your_embedding_model import embed # must match the model your Qdrant collection uses
|
||||||
|
|
||||||
|
client = QdrantClient("http://localhost:6333") # or QdrantClient(url="https://<id>.cloud.qdrant.io", api_key="...") for Qdrant Cloud
|
||||||
|
|
||||||
|
def retrieval_run(golden_set: list, collection: str, k: int = 10) -> Run:
|
||||||
|
run = {}
|
||||||
|
for entry in golden_set:
|
||||||
|
results = client.query_points(
|
||||||
|
collection_name=collection,
|
||||||
|
query=embed(entry["query_text"]),
|
||||||
|
limit=k,
|
||||||
|
).points
|
||||||
|
# p.id type must match the doc_id type in labels (ranx matches by equality).
|
||||||
|
run[entry["query_id"]] = {p.id: p.score for p in results}
|
||||||
|
return Run(run)
|
||||||
|
|
||||||
|
qrels = Qrels({entry["query_id"]: entry["labels"] for entry in golden_set})
|
||||||
|
run = retrieval_run(golden_set, collection="my_collection", k=10)
|
||||||
|
```
|
||||||
|
|
||||||
|
**3. Compute metrics.** `evaluate(qrels, run, [...])` compares the two and returns a dict of metric names to floats.
|
||||||
|
|
||||||
|
```python
|
||||||
|
metrics = evaluate(qrels, run, ["recall@10", "mrr", "ndcg@10"])
|
||||||
|
```
|
||||||
|
|
||||||
|
`evaluate()` returns:
|
||||||
|
|
||||||
|
```python
|
||||||
|
{"recall@10": 0.82, "mrr": 0.71, "ndcg@10": 0.76}
|
||||||
|
```
|
||||||
|
|
||||||
|
Higher is better on all three.
|
||||||
|
|
||||||
|
### Choosing the Right Metric
|
||||||
|
|
||||||
|
Which metric matters most depends on what your pipeline does with results:
|
||||||
|
|
||||||
|
| Scenario | Recommended Metric | Why |
|
||||||
|
|---|---|---|
|
||||||
|
| RAG pipeline (LLM reads top-k chunks) | `Recall@k` | The LLM can recover if a relevant doc is at position 3 vs 1; missing it entirely hurts more |
|
||||||
|
| Single-answer retrieval (FAQ or Q&A) | `MRR` or `Hits@1` | The first result is what the user acts on; lower ranks matter little |
|
||||||
|
| Re-ranking or recommendation feeds | `NDCG@k` | Order within the result list matters; a highly relevant doc at rank 5 is worse than at rank 1 |
|
||||||
|
|
||||||
|
[NDCG (Normalized Discounted Cumulative Gain)](https://en.wikipedia.org/wiki/Discounted_cumulative_gain) needs graded labels (for example, 0/1/2 scores per query-document pair). For binary labels, stick with `recall@k` and [`MRR` (Mean Reciprocal Rank)](https://en.wikipedia.org/wiki/Mean_reciprocal_rank). For the full metric list (Precision@k, MAP, ERR, and others), see the <a href="https://amenra.github.io/ranx/" target="_blank">ranx docs</a>.
|
||||||
|
|
||||||
|
On choosing `k`: set it to match actual usage. If the application shows 5 results to the user, measure `@5`. If a RAG pipeline passes 10 chunks to the LLM, measure `@10`. Reporting `@100` for a UI that surfaces 5 results makes the metric look artificially good.
|
||||||
|
|
||||||
|
### Re-running in CI
|
||||||
|
|
||||||
|
Re-run whenever the retrieval stack changes: new embedding model (which also requires re-embedding queries and re-indexing), new index config, or new reranker. In CI, compute `recall@10` against a fixed golden set and fail the job when the score drops below your target threshold.
|
||||||
|
|
||||||
|
## Pitfalls to Watch For
|
||||||
|
|
||||||
|
In golden sets, **data leakage** means any setup that makes offline metrics look better than production reality. Unlike classic train/test leakage, the issue is often evaluation design. Keep source documents in the index (they are the expected relevant answers). Focus on these risks:
|
||||||
|
|
||||||
|
**Synthetic-query unrealism.** LLMs often mirror source wording, creating easier queries than real user input. This inflates offline scores. Mitigate it by instructing the LLM to generate queries as a user who hasn't seen the source document, then compare synthetic and real-query distributions (length and specificity).
|
||||||
|
|
||||||
|
**Embedding-model contamination.** If your embedding model was trained on pairs overlapping with the golden set, results will look better than true generalization. For hosted models, review published training data when possible. For in-house fine-tuning, keep strict train/eval separation.
|
||||||
|
|
||||||
|
**Near-duplicate documents.** Your retrieval may return a near-duplicate of a labeled document that isn't in the label set. That makes **metrics look worse** because labels are incomplete, not because retrieval is failing. A score dip here is a signal to audit your labels before tuning retrieval. Deduplicate before labeling (for example, cosine similarity > 0.95), or label duplicate clusters together.
|
||||||
|
|
||||||
|
**Temporal drift.** If the corpus changes after labeling, labels go stale: referenced docs may be removed or superseded by newer versions. Pin a corpus snapshot for each run and regenerate the golden set after material corpus changes.
|
||||||
|
|
||||||
|
**Setup reproducibility.** Version the full evaluation setup: corpus snapshot, how labels were produced, and any preprocessing thresholds. Otherwise you can't tell whether a later score drop is model/index regression or dataset drift.
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
Once retrieval relevance is on target, the next layer is pipeline output quality: whether the full pipeline produces the right output when retrieval feeds into a consumer (LLM generator, ranker, or UI). See [Evaluating Pipeline Output Quality](/documentation/improve-search/pipeline-output-quality/).
|
||||||
@@ -68,7 +68,7 @@ partition: develop
|
|||||||
| [Semantic Search Basics](/documentation/tutorials-search-engineering/neural-search/) | Deploy a search service for company descriptions. | <span class="pill">FastAPI</span> | 30m | <span class="text-green">Beginner</span> |
|
| [Semantic Search Basics](/documentation/tutorials-search-engineering/neural-search/) | Deploy a search service for company descriptions. | <span class="pill">FastAPI</span> | 30m | <span class="text-green">Beginner</span> |
|
||||||
| [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
|
| [Collaborative Filtering](/documentation/tutorials-search-engineering/collaborative-filtering/) | Collaborative filtering using sparse embeddings. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
|
||||||
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
| [Multivector Document Retrieval](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) | PDF RAG using ColPali and embedding pooling. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
||||||
| [Retrieval Quality Evaluation](/documentation/tutorials-search-engineering/retrieval-quality/) | Measure quality and tune HNSW parameters. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
| [Measuring ANN Recall](/documentation/tutorials-search-engineering/ann-recall/) | Measure ANN recall with the Web UI and tune HNSW parameters. | <span class="pill">Web UI</span> | 15m | <span class="text-green">Beginner</span> |
|
||||||
| [Reranking for Better Search](/documentation/search-precision/reranking-semantic-search/) | Use multivector representations for better ranking. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
| [Reranking for Better Search](/documentation/search-precision/reranking-semantic-search/) | Use multivector representations for better ranking. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
||||||
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
|
| [Hybrid Search with Reranking](/documentation/tutorials-search-engineering/reranking-hybrid-search/) | Implement late interaction and sparse reranking. | <span class="pill">Python</span> | 40m | <span class="text-yellow">Intermediate</span> |
|
||||||
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
|
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
|
||||||
|
|||||||
@@ -0,0 +1,90 @@
|
|||||||
|
---
|
||||||
|
title: Measuring ANN Recall
|
||||||
|
aliases:
|
||||||
|
- /documentation/tutorials/retrieval-quality/
|
||||||
|
- /documentation/beginner-tutorials/retrieval-quality/
|
||||||
|
- /documentation/tutorials-search-engineering/retrieval-quality/
|
||||||
|
weight: 5
|
||||||
|
---
|
||||||
|
|
||||||
|
# Measuring ANN Recall
|
||||||
|
|
||||||
|
| Time: 15 min | Level: Beginner | | |
|
||||||
|
|--------------|---------------------|--|----|
|
||||||
|
|
||||||
|
This tutorial focuses on **ANN recall**: how closely approximate nearest-neighbor (ANN) search matches exact kNN search, measured with `recall@k`.
|
||||||
|
|
||||||
|
**Prerequisites.** A Qdrant collection populated with your documents as points (vectors + optional payload).
|
||||||
|
|
||||||
|
## The Retrieval Evaluation Stack
|
||||||
|
|
||||||
|
ANN recall measures how closely approximate search matches exact kNN. It's the first of four evaluation layers; each higher layer measures a different property of the retrieval system, with different tools.
|
||||||
|
|
||||||
|
- **ANN recall** (this tutorial). Is the approximate index close to exact kNN?
|
||||||
|
- **Retrieval relevance** ([Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/)). Do the top-k results match query intent?
|
||||||
|
- **Pipeline output quality** ([Evaluating Pipeline Output Quality](/documentation/improve-search/pipeline-output-quality/)). Does the end-to-end pipeline (retrieval + generator, ranker, or UI) produce the right output?
|
||||||
|
- **Business impact**. Do the KPIs the business cares about move? Application-specific, out of scope for these tutorials.
|
||||||
|
|
||||||
|
A high score on a higher layer requires acceptable scores on the layers below. Embedding quality (separately measured by benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard)) sets the ceiling on every downstream metric.
|
||||||
|
|
||||||
|
## Measure ANN Recall with the Web UI
|
||||||
|
|
||||||
|
Qdrant's Web UI includes an ANN Recall tab that measures the gap between approximate and exact search without writing evaluation code. Open the dashboard at `http://localhost:6333/dashboard` (or your cluster's dashboard on Qdrant Cloud), navigate to your collection, open the ANN Recall tab, and click **Check Index Quality** to run the comparison.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
The tab reports average **recall@k** (1.0 = perfect overlap; 0.95+ is typical for well-tuned HNSW).
|
||||||
|
|
||||||
|
## Tuning Search Recall
|
||||||
|
|
||||||
|
Toggle **advanced mode** in the ANN Recall tab to tune search-time parameters inline. The main one is `hnsw_ef`: the number of candidates evaluated during a search. Raising it explores more of the graph, improving recall at the cost of higher query latency. To see the effect, raise `hnsw_ef` (for example, to 256) and run the evaluation again.
|
||||||
|
|
||||||
|
Recall should increase at the cost of higher query latency.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
If `hnsw_ef` alone does not get you to your recall target, the build-time parameters `m` and `ef_construct` set the ceiling on the recall approximate search can achieve. Changing them requires rebuilding the HNSW index. For the trade-offs and how to choose values, see [HNSW Indexing Fundamentals](/course/essentials/day-2/what-is-hnsw/) in the Qdrant Essentials course.
|
||||||
|
|
||||||
|
## Automate in CI with Python
|
||||||
|
|
||||||
|
The Web UI is the fastest way to check recall interactively. For continuous integration or scripted regression tests, the Qdrant client exposes the same exact-search mode via `search_params=models.SearchParams(exact=True)`. Compare the ANN and exact top-k sets yourself and compute recall@k.
|
||||||
|
|
||||||
|
This helper takes a list of query vectors and returns the average recall@k. Use a representative sample of query vectors from your workload (typically 20–50, embedded with the same model your collection uses) as your test set.
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
|
|
||||||
|
def avg_recall_at_k(
|
||||||
|
client: QdrantClient,
|
||||||
|
collection_name: str,
|
||||||
|
test_vectors: list,
|
||||||
|
k: int,
|
||||||
|
) -> float:
|
||||||
|
recalls = []
|
||||||
|
for vector in test_vectors:
|
||||||
|
ann_ids = {
|
||||||
|
p.id for p in client.query_points(
|
||||||
|
collection_name=collection_name,
|
||||||
|
query=vector,
|
||||||
|
limit=k,
|
||||||
|
).points
|
||||||
|
}
|
||||||
|
knn_ids = {
|
||||||
|
p.id for p in client.query_points(
|
||||||
|
collection_name=collection_name,
|
||||||
|
query=vector,
|
||||||
|
limit=k,
|
||||||
|
search_params=models.SearchParams(exact=True),
|
||||||
|
).points
|
||||||
|
}
|
||||||
|
recalls.append(len(ann_ids & knn_ids) / k)
|
||||||
|
|
||||||
|
return sum(recalls) / len(recalls)
|
||||||
|
```
|
||||||
|
|
||||||
|
Wire it into CI and fail the job when recall falls below your target threshold. This catches regressions from embedding model swaps or index config changes before they reach production.
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
Once ANN recall is on target, continue with [Measuring Retrieval Relevance](/documentation/improve-search/retrieval-relevance/) to check how well those results match user intent.
|
||||||
-230
@@ -1,230 +0,0 @@
|
|||||||
---
|
|
||||||
title: Retrieval Quality Evaluation
|
|
||||||
aliases:
|
|
||||||
- /documentation/tutorials/retrieval-quality/
|
|
||||||
- /documentation/beginner-tutorials/retrieval-quality/
|
|
||||||
weight: 4
|
|
||||||
---
|
|
||||||
|
|
||||||
# Evaluate Retrieval Quality with Qdrant
|
|
||||||
|
|
||||||
| Time: 30 min | Level: Intermediate | | |
|
|
||||||
|--------------|---------------------|--|----|
|
|
||||||
|
|
||||||
Semantic search pipelines are as good as the embeddings they use. If your model cannot properly represent input data, similar objects might
|
|
||||||
be far away from each other in the vector space. No surprise, that the search results will be poor in this case. There is, however, another
|
|
||||||
component of the process which can also degrade the quality of the search results. It is the ANN algorithm itself.
|
|
||||||
|
|
||||||
In this tutorial, we will show how to measure the quality of the semantic retrieval and how to tune the parameters of the HNSW, the ANN
|
|
||||||
algorithm used in Qdrant, to obtain the best results.
|
|
||||||
|
|
||||||
## Embeddings quality
|
|
||||||
|
|
||||||
The quality of the embeddings is a topic for a separate tutorial. In a nutshell, it is usually measured and compared by benchmarks, such as
|
|
||||||
[Massive Text Embedding Benchmark (MTEB)](https://huggingface.co/spaces/mteb/leaderboard). The evaluation process itself is pretty
|
|
||||||
straightforward and is based on a ground truth dataset built by humans. We have a set of queries and a set of the documents we would expect
|
|
||||||
to receive for each of them. In the [evaluation process](https://qdrant.tech/rag/rag-evaluation-guide/), we take a query, find the most similar documents in the vector space and compare
|
|
||||||
them with the ground truth. In that setup, **finding the most similar documents is implemented as full kNN search, without any approximation**.
|
|
||||||
As a result, we can measure the quality of the embeddings themselves, without the influence of the ANN algorithm.
|
|
||||||
|
|
||||||
## Retrieval quality
|
|
||||||
|
|
||||||
Embeddings quality is indeed the most important factor in the semantic search quality. However, vector search engines, such as Qdrant, do not
|
|
||||||
perform pure kNN search. Instead, they use **Approximate Nearest Neighbors** (ANN) algorithms, which are much faster than the exact search,
|
|
||||||
but can return suboptimal results. We can also **measure the retrieval quality of that approximation** which also contributes to the overall
|
|
||||||
search quality.
|
|
||||||
|
|
||||||
### Quality metrics
|
|
||||||
|
|
||||||
There are various ways of how quantify the quality of semantic search. Some of them, such as [Precision@k](https://en.wikipedia.org/wiki/Evaluation_measures_(information_retrieval)#Precision_at_k),
|
|
||||||
are based on the number of relevant documents in the top-k search results. Others, such as [Mean Reciprocal Rank (MRR)](https://en.wikipedia.org/wiki/Mean_reciprocal_rank),
|
|
||||||
take into account the position of the first relevant document in the search results. [DCG and NDCG](https://en.wikipedia.org/wiki/Discounted_cumulative_gain)
|
|
||||||
metrics are, in turn, based on the relevance score of the documents.
|
|
||||||
|
|
||||||
If we treat the search pipeline as a whole, we could use them all. The same is true for the embeddings quality evaluation. However, for the
|
|
||||||
ANN algorithm itself, anything based on the relevance score or ranking is not applicable. Ranking in vector search relies on the distance
|
|
||||||
between the query and the document in the vector space, however distance is not going to change due to approximation, as the function is
|
|
||||||
still the same.
|
|
||||||
|
|
||||||
Therefore, it only makes sense to measure the quality of the ANN algorithm by the number of relevant documents in the top-k search results,
|
|
||||||
such as `precision@k`. It is calculated as the number of relevant documents in the top-k search results divided by `k`. In case of testing
|
|
||||||
just the ANN algorithm, we can use the exact kNN search as a ground truth, with `k` being fixed. It will be a measure on **how well the ANN
|
|
||||||
algorithm approximates the exact search**.
|
|
||||||
|
|
||||||
## Measure the quality of the search results
|
|
||||||
|
|
||||||
Let's build a quality [evaluation](https://qdrant.tech/rag/rag-evaluation-guide/) of the ANN algorithm in Qdrant. We will, first, call the search endpoint in a standard way to obtain
|
|
||||||
the approximate search results. Then, we will call the exact search endpoint to obtain the exact matches, and finally compare both results
|
|
||||||
in terms of precision.
|
|
||||||
|
|
||||||
Before we start, let's create a collection, fill it with some data and then start our evaluation. We will use the same dataset as in the
|
|
||||||
[Loading a dataset from Hugging Face hub](/documentation/tutorials-basics/huggingface-datasets/) tutorial, `Qdrant/arxiv-titles-instructorxl-embeddings`
|
|
||||||
from the [Hugging Face hub](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings). Let's download it in a streaming
|
|
||||||
mode, as we are only going to use part of it.
|
|
||||||
|
|
||||||
```python
|
|
||||||
from datasets import load_dataset
|
|
||||||
|
|
||||||
dataset = load_dataset(
|
|
||||||
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
We need some data to be indexed and another set for the testing purposes. Let's get the first 50000 items for the training and the next 1000
|
|
||||||
for the testing.
|
|
||||||
|
|
||||||
```python
|
|
||||||
dataset_iterator = iter(dataset)
|
|
||||||
train_dataset = [next(dataset_iterator) for _ in range(60000)]
|
|
||||||
test_dataset = [next(dataset_iterator) for _ in range(1000)]
|
|
||||||
```
|
|
||||||
|
|
||||||
Now, let's create a collection and index the training data. This collection will be created with the default configuration. Please be aware that
|
|
||||||
it might be different from your collection settings, and it's always important to test exactly the same configuration you are going to use later
|
|
||||||
in production.
|
|
||||||
|
|
||||||
<aside role="status">
|
|
||||||
Distance function is another parameter that may impact the retrieval quality. If the embedding model was not trained to minimize cosine
|
|
||||||
distance, you can get suboptimal search results by using it. Please test different distance functions to find the best one for your embeddings,
|
|
||||||
if you don't know the specifics of the model training.
|
|
||||||
</aside>
|
|
||||||
|
|
||||||
```python
|
|
||||||
from qdrant_client import QdrantClient, models
|
|
||||||
|
|
||||||
client = QdrantClient("http://localhost:6333")
|
|
||||||
client.create_collection(
|
|
||||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
|
||||||
vectors_config=models.VectorParams(
|
|
||||||
size=768, # Size of the embeddings generated by InstructorXL model
|
|
||||||
distance=models.Distance.COSINE,
|
|
||||||
),
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
We are now ready to index the training data. Uploading the records is going to trigger the indexing process, which will build the HNSW graph.
|
|
||||||
The indexing process may take some time, depending on the size of the dataset, but your data is going to be available for search immediately
|
|
||||||
after receiving the response from the `upsert` endpoint. **As long as the indexing is not finished, and HNSW not built, Qdrant will perform
|
|
||||||
the exact search**. We have to wait until the indexing is finished to be sure that the approximate search is performed.
|
|
||||||
|
|
||||||
```python
|
|
||||||
client.upload_points( # upload_points is available as of qdrant-client v1.7.1
|
|
||||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
|
||||||
points=[
|
|
||||||
models.PointStruct(
|
|
||||||
id=item["id"],
|
|
||||||
vector=item["vector"],
|
|
||||||
payload=item,
|
|
||||||
)
|
|
||||||
for item in train_dataset
|
|
||||||
]
|
|
||||||
)
|
|
||||||
|
|
||||||
while True:
|
|
||||||
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
|
|
||||||
if collection_info.status == models.CollectionStatus.GREEN:
|
|
||||||
# Collection status is green, which means the indexing is finished
|
|
||||||
break
|
|
||||||
```
|
|
||||||
|
|
||||||
## Standard mode vs exact search
|
|
||||||
|
|
||||||
Qdrant has a built-in exact search mode, which can be used to measure the quality of the search results. In this mode, Qdrant performs a
|
|
||||||
full kNN search for each query, without any approximation. It is not suitable for production use with high load, but it is perfect for the
|
|
||||||
evaluation of the ANN algorithm and its parameters. It might be triggered by setting the `exact` parameter to `True` in the search request.
|
|
||||||
We are simply going to use all the examples from the test dataset as queries and compare the results of the approximate search with the
|
|
||||||
results of the exact search. Let's create a helper function with `k` being a parameter, so we can calculate the `precision@k` for different
|
|
||||||
values of `k`.
|
|
||||||
|
|
||||||
```python
|
|
||||||
def avg_precision_at_k(k: int):
|
|
||||||
precisions = []
|
|
||||||
for item in test_dataset:
|
|
||||||
ann_result = client.query_points(
|
|
||||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
|
||||||
query=item["vector"],
|
|
||||||
limit=k,
|
|
||||||
).points
|
|
||||||
|
|
||||||
knn_result = client.query_points(
|
|
||||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
|
||||||
query=item["vector"],
|
|
||||||
limit=k,
|
|
||||||
search_params=models.SearchParams(
|
|
||||||
exact=True, # Turns on the exact search mode
|
|
||||||
),
|
|
||||||
).points
|
|
||||||
|
|
||||||
# We can calculate the precision@k by comparing the ids of the search results
|
|
||||||
ann_ids = set(item.id for item in ann_result)
|
|
||||||
knn_ids = set(item.id for item in knn_result)
|
|
||||||
precision = len(ann_ids.intersection(knn_ids)) / k
|
|
||||||
precisions.append(precision)
|
|
||||||
|
|
||||||
return sum(precisions) / len(precisions)
|
|
||||||
```
|
|
||||||
|
|
||||||
Calculating the `precision@5` is as simple as calling the function with the corresponding parameter:
|
|
||||||
|
|
||||||
```python
|
|
||||||
print(f"avg(precision@5) = {avg_precision_at_k(k=5)}")
|
|
||||||
```
|
|
||||||
|
|
||||||
Response:
|
|
||||||
|
|
||||||
```text
|
|
||||||
avg(precision@5) = 0.9935999999999995
|
|
||||||
```
|
|
||||||
|
|
||||||
As we can see, the precision of the approximate search vs exact search is pretty high. There are, however, some scenarios when we
|
|
||||||
need higher precision and can accept higher latency. HNSW is pretty tunable, and we can increase the precision by changing its parameters.
|
|
||||||
|
|
||||||
## Tweaking the HNSW parameters
|
|
||||||
|
|
||||||
HNSW is a hierarchical graph, where each node has a set of links to other nodes. The number of edges per node is called the `m` parameter.
|
|
||||||
The larger the value of it, the higher the precision of the search, but more space required. The `ef_construct` parameter is the number of
|
|
||||||
neighbours to consider during the index building. Again, the larger the value, the higher the precision, but the longer the indexing time.
|
|
||||||
The default values of these parameters are `m=16` and `ef_construct=100`. Let's try to increase them to `m=32` and `ef_construct=200` and
|
|
||||||
see how it affects the precision. Of course, we need to wait until the indexing is finished before we can perform the search.
|
|
||||||
|
|
||||||
```python
|
|
||||||
client.update_collection(
|
|
||||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
|
||||||
hnsw_config=models.HnswConfigDiff(
|
|
||||||
m=32, # Increase the number of edges per node from the default 16 to 32
|
|
||||||
ef_construct=200, # Increase the number of neighbours from the default 100 to 200
|
|
||||||
)
|
|
||||||
)
|
|
||||||
|
|
||||||
while True:
|
|
||||||
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
|
|
||||||
if collection_info.status == models.CollectionStatus.GREEN:
|
|
||||||
# Collection status is green, which means the indexing is finished
|
|
||||||
break
|
|
||||||
```
|
|
||||||
|
|
||||||
The same function can be used to calculate the average `precision@5`:
|
|
||||||
|
|
||||||
```python
|
|
||||||
print(f"avg(precision@5) = {avg_precision_at_k(k=5)}")
|
|
||||||
```
|
|
||||||
|
|
||||||
Response:
|
|
||||||
|
|
||||||
```text
|
|
||||||
avg(precision@5) = 0.9969999999999998
|
|
||||||
```
|
|
||||||
|
|
||||||
The precision has obviously increased, and we know how to control it. However, there is a trade-off between the precision and the search
|
|
||||||
latency and memory requirements. In some specific cases, we may want to increase the precision as much as possible, so now we know how
|
|
||||||
to do it.
|
|
||||||
|
|
||||||
## Wrapping up
|
|
||||||
|
|
||||||
Assessing the quality of retrieval is a critical aspect of [evaluating](https://qdrant.tech/rag/rag-evaluation-guide/) semantic search performance. It is imperative to measure retrieval quality when aiming for optimal quality of.
|
|
||||||
your search results. Qdrant provides a built-in exact search mode, which can be used to measure the quality of the ANN algorithm itself,
|
|
||||||
even in an automated way, as part of your CI/CD pipeline.
|
|
||||||
|
|
||||||
Again, **the quality of the embeddings is the most important factor**. HNSW does a pretty good job in terms of precision, and it is
|
|
||||||
parameterizable and tunable, when required. There are some other ANN algorithms available out there, such as [IVF*](https://github.com/facebookresearch/faiss/wiki/Faiss-indexes#cell-probe-methods-indexivf-indexes),
|
|
||||||
but they usually [perform worse than HNSW in terms of quality and performance](https://nirantk.com/writing/pgvector-vs-qdrant/#correctness).
|
|
||||||
+1
-1
@@ -135,7 +135,7 @@ ranking quality of search results, with higher scores indicating better performa
|
|||||||
Binary Quantization definitely speeds up the retrieval, and make it cheaper, but also seems not to affect the quality of
|
Binary Quantization definitely speeds up the retrieval, and make it cheaper, but also seems not to affect the quality of
|
||||||
the retrieval much in some cases. **However, that's something you should carefully verify on your own data**. If you are
|
the retrieval much in some cases. **However, that's something you should carefully verify on your own data**. If you are
|
||||||
a Qdrant user, then you can just enable quantization on an existing collection and [measure the impact on the retrieval
|
a Qdrant user, then you can just enable quantization on an existing collection and [measure the impact on the retrieval
|
||||||
quality](/documentation/tutorials-search-engineering/retrieval-quality/).
|
quality](/documentation/tutorials-search-engineering/ann-recall/).
|
||||||
|
|
||||||
All the tests we did were performed using [`beir-qdrant`](https://github.com/kacperlukawski/beir-qdrant), and might be
|
All the tests we did were performed using [`beir-qdrant`](https://github.com/kacperlukawski/beir-qdrant), and might be
|
||||||
reproduced by running [the script available on the project
|
reproduced by running [the script available on the project
|
||||||
|
|||||||
BIN
Binary file not shown.
|
After Width: | Height: | Size: 196 KiB |
BIN
Binary file not shown.
|
After Width: | Height: | Size: 240 KiB |
Reference in New Issue
Block a user