- Add a layer-1 anchor sentence under the Time/Level table pointing at the evaluation ladder in Fundamentals, mirroring the golden-set tutorial's anchor. Addresses abdonpijpelink's ask for a levels-table link and a "this tutorial focuses on level 1" framing. - Title-case the six H2 headers for consistency across the tutorials-search-engineering set. Deferred: streaming the 60K training items into upload_points instead of materializing as a list (abdonpijpelink line 69) — pending manager confirmation before proceeding. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
11 KiB
title, aliases, weight
| title | aliases | weight | ||
|---|---|---|---|---|
| Retrieval Quality Evaluation |
|
6 |
Evaluate Retrieval Quality
| Time: 30 min | Level: Intermediate |
|---|
This tutorial measures layer 1 of the evaluation ladder, ANN recall: the share of exact kNN results that Qdrant's approximate nearest-neighbor search recovers. For retrieval relevance (layer 2), see the Building a Golden Query Set tutorial.
Semantic search pipelines are as good as the embeddings they use. If your model cannot properly represent input data, similar objects might be far away from each other in the vector space. No surprise, that the search results will be poor in this case. There is, however, another component of the process which can also degrade the quality of the search results. It is the ANN algorithm itself.
In this tutorial, we will show how to measure the quality of the semantic retrieval and how to tune the parameters of the HNSW, the ANN algorithm used in Qdrant, to obtain the best results.
Embeddings Quality
The quality of the embeddings is a topic for a separate tutorial. In a nutshell, it is usually measured and compared by benchmarks, such as Massive Text Embedding Benchmark (MTEB). The evaluation process itself is pretty straightforward and is based on a ground truth dataset built by humans. We have a set of queries and a set of the documents we would expect to receive for each of them. In the evaluation process, we take a query, find the most similar documents in the vector space and compare them with the ground truth. In that setup, finding the most similar documents is implemented as full kNN search, without any approximation. As a result, we can measure the quality of the embeddings themselves, without the influence of the ANN algorithm.
Retrieval Quality
Embeddings quality is indeed the most important factor in the semantic search quality. However, vector search engines, such as Qdrant, do not perform pure kNN search. Instead, they use Approximate Nearest Neighbors (ANN) algorithms, which are much faster than the exact search, but can return suboptimal results. We can also measure the retrieval quality of that approximation which also contributes to the overall search quality.
For a broader discussion of what to measure and when (ANN recall vs retrieval relevance vs business impact, and which metric fits which scenario),
see Retrieval Quality Fundamentals. This tutorial focuses on the
ANN-algorithm layer and measures it with recall@k: the fraction of the true top-k items returned by exact search that the approximate search
recovers. When both ANN and exact search return exactly k items, recall@k and precision@k are numerically identical; we use "recall" to
match the ANN-benchmarks convention.
Measure the Quality of the Search Results
Let's build a quality evaluation of the ANN algorithm in Qdrant. We will, first, call the search endpoint in a standard way to obtain the approximate search results. Then, we will call the exact search endpoint to obtain the exact matches, and finally compare both results in terms of recall.
Before we start, let's create a collection, fill it with some data and then start our evaluation. We will use the same dataset as in the
Loading a dataset from Hugging Face hub tutorial, Qdrant/arxiv-titles-instructorxl-embeddings
from the Hugging Face hub. Let's download it in a streaming
mode, as we are only going to use part of it.
from datasets import load_dataset
dataset = load_dataset(
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
)
We need some data to be indexed and another set for the testing purposes. Let's get the first 60000 items for the training and the next 1000 for the testing.
dataset_iterator = iter(dataset)
train_dataset = [next(dataset_iterator) for _ in range(60000)]
test_dataset = [next(dataset_iterator) for _ in range(1000)]
Now, let's create a collection and index the training data. This collection will be created with the default configuration. Please be aware that it might be different from your collection settings, and it's always important to test exactly the same configuration you are going to use later in production.
from qdrant_client import QdrantClient, models
client = QdrantClient("http://localhost:6333")
client.create_collection(
collection_name="arxiv-titles-instructorxl-embeddings",
vectors_config=models.VectorParams(
size=768, # Size of the embeddings generated by InstructorXL model
distance=models.Distance.COSINE,
),
)
We are now ready to index the training data. Uploading the records is going to trigger the indexing process, which will build the HNSW graph.
The indexing process may take some time, depending on the size of the dataset, but your data is going to be available for search immediately
after receiving the response from the upsert endpoint. As long as the indexing is not finished, and HNSW not built, Qdrant will perform
the exact search. We have to wait until the indexing is finished to be sure that the approximate search is performed.
client.upload_points( # upload_points is available as of qdrant-client v1.7.1
collection_name="arxiv-titles-instructorxl-embeddings",
points=[
models.PointStruct(
id=item["id"],
vector=item["vector"],
payload=item,
)
for item in train_dataset
]
)
while True:
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
if collection_info.status == models.CollectionStatus.GREEN:
# Collection status is green, which means the indexing is finished
break
Standard Mode vs Exact Search
Qdrant has a built-in exact search mode, which can be used to measure the quality of the search results. In this mode, Qdrant performs a
full kNN search for each query, without any approximation. It is not suitable for production use with high load, but it is perfect for the
evaluation of the ANN algorithm and its parameters. It might be triggered by setting the exact parameter to True in the search request.
We are simply going to use all the examples from the test dataset as queries and compare the results of the approximate search with the
results of the exact search. Let's create a helper function with k being a parameter, so we can calculate the recall@k for different
values of k.
def avg_recall_at_k(k: int):
recalls = []
for item in test_dataset:
ann_result = client.query_points(
collection_name="arxiv-titles-instructorxl-embeddings",
query=item["vector"],
limit=k,
).points
knn_result = client.query_points(
collection_name="arxiv-titles-instructorxl-embeddings",
query=item["vector"],
limit=k,
search_params=models.SearchParams(
exact=True, # Turns on the exact search mode
),
).points
# We can calculate the recall@k by comparing the ids of the search results
ann_ids = set(item.id for item in ann_result)
knn_ids = set(item.id for item in knn_result)
recall = len(ann_ids.intersection(knn_ids)) / k
recalls.append(recall)
return sum(recalls) / len(recalls)
Calculating the recall@5 is as simple as calling the function with the corresponding parameter:
print(f"avg(recall@5) = {avg_recall_at_k(k=5)}")
Response:
avg(recall@5) = 0.9935999999999995
As we can see, the recall of the approximate search vs exact search is pretty high. There are, however, some scenarios when we need higher recall and can accept higher latency. HNSW is pretty tunable, and we can increase the recall by changing its parameters.
Tweaking the HNSW Parameters
HNSW is a hierarchical graph, where each node has a set of links to other nodes. The number of edges per node is called the m parameter.
The larger the value of it, the higher the recall of the search, but more space required. The ef_construct parameter is the number of
neighbours to consider during the index building. Again, the larger the value, the higher the recall, but the longer the indexing time.
The default values of these parameters are m=16 and ef_construct=100. Let's try to increase them to m=32 and ef_construct=200 and
see how it affects the recall. Of course, we need to wait until the indexing is finished before we can perform the search.
client.update_collection(
collection_name="arxiv-titles-instructorxl-embeddings",
hnsw_config=models.HnswConfigDiff(
m=32, # Increase the number of edges per node from the default 16 to 32
ef_construct=200, # Increase the number of neighbours from the default 100 to 200
)
)
while True:
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
if collection_info.status == models.CollectionStatus.GREEN:
# Collection status is green, which means the indexing is finished
break
The same function can be used to calculate the average recall@5:
print(f"avg(recall@5) = {avg_recall_at_k(k=5)}")
Response:
avg(recall@5) = 0.9969999999999998
The recall has obviously increased, and we know how to control it. However, there is a trade-off between the recall and the search latency and memory requirements. In some specific cases, we may want to increase the recall as much as possible, so now we know how to do it.
Wrapping Up
Assessing the quality of retrieval is a critical aspect of evaluating semantic search performance. It is imperative to measure retrieval quality when aiming for optimal quality of. your search results. Qdrant provides a built-in exact search mode, which can be used to measure the quality of the ANN algorithm itself, even in an automated way, as part of your CI/CD pipeline.
Again, the quality of the embeddings is the most important factor. HNSW does a pretty good job in terms of recall, and it is parameterizable and tunable, when required. There are some other ANN algorithms available out there, such as IVF*, but they usually perform worse than HNSW in terms of quality and performance.