diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md index 2038e297c..b484ef475 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md @@ -15,18 +15,9 @@ This tutorial measures **layer 1** of the +![Search Quality tab with default evaluation results](/documentation/tutorials/retrieval-quality/search-quality-tab.png) -```python -from datasets import load_dataset +The tab reports average **precision@k**. The score is typically high but not always perfect. When you need higher recall and can accept higher latency or more memory, HNSW is tunable. -dataset = load_dataset( - "Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True -) -``` +## Tweaking the HNSW Parameters -We need some data to be indexed and another set for the testing purposes. Let's get the first 60000 items for the training and the next 1000 -for the testing. +HNSW is a hierarchical graph where each node has a set of links to other nodes. The `m` parameter controls the number of edges per node: higher `m` means higher recall at the cost of more memory. The `ef_construct` parameter controls how many neighbours are considered during index building: higher `ef_construct` means higher recall at the cost of longer indexing time. Defaults are `m=16` and `ef_construct=100`. -```python -dataset_iterator = iter(dataset) -train_dataset = [next(dataset_iterator) for _ in range(60000)] -test_dataset = [next(dataset_iterator) for _ in range(1000)] -``` +For the full list of HNSW parameters, including on-disk storage and precision/memory trade-offs, see [Optimize Performance](/documentation/operations/optimize/). -Now, let's create a collection and index the training data. This collection will be created with the default configuration. Please be aware that -it might be different from your collection settings, and it's always important to test exactly the same configuration you are going to use later -in production. +Toggle **advanced mode** in the Search Quality tab to tune these parameters inline. Raise `m` to 32 and `ef_construct` to 200, then run the evaluation again. - + +![Search Quality advanced mode with HNSW parameters](/documentation/tutorials/retrieval-quality/search-quality-advanced.png) + +Precision should increase at the cost of higher build time and memory. + + +![Search Quality results after HNSW tuning](/documentation/tutorials/retrieval-quality/search-quality-after-tuning.png) + +Tune until you hit the point that matches your quality and cost targets. + +## Automate in CI with Python + +The Web UI is the fastest way to check recall interactively. For continuous integration or scripted regression tests, the Qdrant client exposes the same exact-search mode via `search_params=models.SearchParams(exact=True)`. Compare the ANN and exact top-k sets yourself and compute recall. + +The helper below takes a list of query vectors and returns the average recall@k. Supply your own test set: a representative sample of query vectors from your workload, held out from training. ```python from qdrant_client import QdrantClient, models -client = QdrantClient("http://localhost:6333") -client.create_collection( - collection_name="arxiv-titles-instructorxl-embeddings", - vectors_config=models.VectorParams( - size=768, # Size of the embeddings generated by InstructorXL model - distance=models.Distance.COSINE, - ), -) -``` -We are now ready to index the training data. Uploading the records is going to trigger the indexing process, which will build the HNSW graph. -The indexing process may take some time, depending on the size of the dataset, but your data is going to be available for search immediately -after receiving the response from the `upsert` endpoint. **As long as the indexing is not finished, and HNSW not built, Qdrant will perform -the exact search**. We have to wait until the indexing is finished to be sure that the approximate search is performed. - -```python -client.upload_points( # upload_points is available as of qdrant-client v1.7.1 - collection_name="arxiv-titles-instructorxl-embeddings", - points=[ - models.PointStruct( - id=item["id"], - vector=item["vector"], - payload=item, - ) - for item in train_dataset - ] -) - -while True: - collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings") - if collection_info.status == models.CollectionStatus.GREEN: - # Collection status is green, which means the indexing is finished - break -``` - -## Standard Mode vs Exact Search - -Qdrant has a built-in exact search mode, which can be used to measure the quality of the search results. In this mode, Qdrant performs a -full kNN search for each query, without any approximation. It is not suitable for production use with high load, but it is perfect for the -evaluation of the ANN algorithm and its parameters. It might be triggered by setting the `exact` parameter to `True` in the search request. -We are simply going to use all the examples from the test dataset as queries and compare the results of the approximate search with the -results of the exact search. Let's create a helper function with `k` being a parameter, so we can calculate the `recall@k` for different -values of `k`. - -```python -def avg_recall_at_k(k: int): +def avg_recall_at_k( + client: QdrantClient, + collection_name: str, + test_vectors: list, + k: int, +) -> float: recalls = [] - for item in test_dataset: - ann_result = client.query_points( - collection_name="arxiv-titles-instructorxl-embeddings", - query=item["vector"], - limit=k, - ).points - - knn_result = client.query_points( - collection_name="arxiv-titles-instructorxl-embeddings", - query=item["vector"], - limit=k, - search_params=models.SearchParams( - exact=True, # Turns on the exact search mode - ), - ).points + for vector in test_vectors: + ann_ids = { + p.id for p in client.query_points( + collection_name=collection_name, + query=vector, + limit=k, + ).points + } + knn_ids = { + p.id for p in client.query_points( + collection_name=collection_name, + query=vector, + limit=k, + search_params=models.SearchParams(exact=True), + ).points + } + recalls.append(len(ann_ids & knn_ids) / k) - # We can calculate the recall@k by comparing the ids of the search results - ann_ids = set(item.id for item in ann_result) - knn_ids = set(item.id for item in knn_result) - recall = len(ann_ids.intersection(knn_ids)) / k - recalls.append(recall) - return sum(recalls) / len(recalls) ``` -Calculating the `recall@5` is as simple as calling the function with the corresponding parameter: - -```python -print(f"avg(recall@5) = {avg_recall_at_k(k=5)}") -``` - -Response: - -```text -avg(recall@5) = 0.9935999999999995 -``` - -As we can see, the recall of the approximate search vs exact search is pretty high. There are, however, some scenarios when we -need higher recall and can accept higher latency. HNSW is pretty tunable, and we can increase the recall by changing its parameters. - -## Tweaking the HNSW Parameters - -HNSW is a hierarchical graph, where each node has a set of links to other nodes. The number of edges per node is called the `m` parameter. -The larger the value of it, the higher the recall of the search, but more space required. The `ef_construct` parameter is the number of -neighbours to consider during the index building. Again, the larger the value, the higher the recall, but the longer the indexing time. -The default values of these parameters are `m=16` and `ef_construct=100`. Let's try to increase them to `m=32` and `ef_construct=200` and -see how it affects the recall. Of course, we need to wait until the indexing is finished before we can perform the search. - -```python -client.update_collection( - collection_name="arxiv-titles-instructorxl-embeddings", - hnsw_config=models.HnswConfigDiff( - m=32, # Increase the number of edges per node from the default 16 to 32 - ef_construct=200, # Increase the number of neighbours from the default 100 to 200 - ) -) - -while True: - collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings") - if collection_info.status == models.CollectionStatus.GREEN: - # Collection status is green, which means the indexing is finished - break -``` - -The same function can be used to calculate the average `recall@5`: - -```python -print(f"avg(recall@5) = {avg_recall_at_k(k=5)}") -``` - -Response: - -```text -avg(recall@5) = 0.9969999999999998 -``` - -The recall has obviously increased, and we know how to control it. However, there is a trade-off between the recall and the search -latency and memory requirements. In some specific cases, we may want to increase the recall as much as possible, so now we know how -to do it. +Drop it into your CI pipeline and fail the job if recall drops below a threshold after an embedding model change or index config update. ## Wrapping Up -Assessing the quality of retrieval is a critical aspect of evaluating semantic search performance. It is imperative to measure retrieval quality when aiming for optimal quality of. -your search results. Qdrant provides a built-in exact search mode, which can be used to measure the quality of the ANN algorithm itself, -even in an automated way, as part of your CI/CD pipeline. +Measuring ANN recall keeps HNSW tuning honest. The Search Quality tab gives you a quick interactive read; the Python helper above plugs into CI to catch regressions after embedding model changes or index config updates. -Again, **the quality of the embeddings is the most important factor**. HNSW does a pretty good job in terms of recall, and it is -parameterizable and tunable, when required. There are some other ANN algorithms available out there, such as [IVF*](https://github.com/facebookresearch/faiss/wiki/Faiss-indexes#cell-probe-methods-indexivf-indexes), -but they usually [perform worse than HNSW in terms of quality and performance](https://nirantk.com/writing/pgvector-vs-qdrant/#correctness). +HNSW covers most workloads well and is tunable when you need more recall. Other ANN algorithms exist, such as [IVF*](https://github.com/facebookresearch/faiss/wiki/Faiss-indexes#cell-probe-methods-indexivf-indexes), but they generally [perform worse than HNSW on quality and performance](https://nirantk.com/writing/pgvector-vs-qdrant/#correctness).