retrieval-quality: pivot to Web UI, add CI skeleton, trim setup

- Drop the dataset-setup walkthrough (HF loading, collection
    create, upload, wait-for-green). Readers at this phase already
    have a collection.
  - Replace the Python evaluation block with a "Measure ANN Recall
    with the Web UI" section built around the Search Quality tab.
    Default run is one-click (sample size 10); HNSW tuning uses the
    tab's advanced mode instead of update_collection. Three
    screenshot placeholders at
    /documentation/tutorials/retrieval-quality/*.png.
  - Reflect that the tab reports precision@k; note the recall@k
    equivalence already spelled out in the ANN Recall section.
  - Keep Python but move it to an "Automate in CI" section with a
    reusable skeleton function.
  - Collapse the standalone "Embeddings Quality" section into a
    one-sentence MTEB pointer inside ANN Recall.
  - Rewrite Wrapping Up to match the new scope.
  - Link the HNSW tuning section to Optimize Performance for the
    full parameter reference.
This commit is contained in:
Dylan Couzon
2026-04-22 16:34:16 -04:00
parent 80bae7403e
commit 921b645e70
@@ -15,18 +15,9 @@ This tutorial measures **layer 1** of the <a href="/documentation/tutorials-sear
We'll measure Qdrant's ANN recall with `recall@k` and tune HNSW parameters to control the recall/latency trade-off. The ANN algorithm is one of several levers that shape retrieval quality in a production pipeline, alongside the embedding model, retrieval strategy (dense, sparse, hybrid, and multi-vector), filtering, and reranking. We'll measure Qdrant's ANN recall with `recall@k` and tune HNSW parameters to control the recall/latency trade-off. The ANN algorithm is one of several levers that shape retrieval quality in a production pipeline, alongside the embedding model, retrieval strategy (dense, sparse, hybrid, and multi-vector), filtering, and reranking.
## Embeddings Quality
The quality of the embeddings is a topic for a separate tutorial. In a nutshell, it is usually measured and compared by benchmarks, such as
[Massive Text Embedding Benchmark (MTEB)](https://huggingface.co/spaces/mteb/leaderboard). The evaluation process itself is pretty
straightforward and is based on a ground truth dataset built by humans. We have a set of queries and a set of the documents we would expect
to receive for each of them. In the evaluation process, we take a query, find the most similar documents in the vector space and compare
them with the ground truth. In that setup, **finding the most similar documents is implemented as full kNN search, without any approximation**.
As a result, we can measure the quality of the embeddings themselves, without the influence of the ANN algorithm.
## ANN Recall ## ANN Recall
The embedding model sets a baseline for search quality, but the retrieval pipeline can still underperform it. Vector search engines such as Qdrant don't run pure kNN at query time; they use **approximate nearest-neighbor** (ANN) algorithms for speed. ANN is faster than exact search but can return suboptimal results. **ANN recall** measures that gap. Embedding quality sets the ceiling on search quality and is measured separately via benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard). The retrieval pipeline can still underperform that ceiling: vector search engines such as Qdrant don't run pure kNN at query time but use **approximate nearest-neighbor** (ANN) algorithms for speed. ANN is faster than exact search but can return suboptimal results. **ANN recall** measures that gap.
For a broader discussion of what to measure and when (ANN recall vs retrieval relevance vs business impact, and which metric fits which scenario), For a broader discussion of what to measure and when (ANN recall vs retrieval relevance vs business impact, and which metric fits which scenario),
see [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/). This tutorial focuses on the see [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/). This tutorial focuses on the
@@ -34,180 +25,75 @@ ANN-algorithm layer and measures it with `recall@k`: the fraction of the true to
recovers. When both ANN and exact search return exactly `k` items, `recall@k` and `precision@k` are numerically identical; we use "recall" to recovers. When both ANN and exact search return exactly `k` items, `recall@k` and `precision@k` are numerically identical; we use "recall" to
match the ANN-benchmarks convention. match the ANN-benchmarks convention.
## Measure the Quality of the Search Results ## Measure ANN Recall with the Web UI
Let's build a quality evaluation of the ANN algorithm in Qdrant. We will, first, call the search endpoint in a standard way to obtain Qdrant's Web UI has a Search Quality tab that measures the gap between approximate and exact search without requiring evaluation code. Open the dashboard at `http://localhost:6333/dashboard` (or your cluster's dashboard on Qdrant Cloud), navigate to your collection, and click the Search Quality tab. A run launches automatically with a default sample size of 10 queries, comparing ANN against exact kNN.
the approximate search results. Then, we will call the exact search endpoint to obtain the exact matches, and finally compare both results
in terms of recall.
Before we start, let's create a collection, fill it with some data and then start our evaluation. We will use the same dataset as in the <!-- SCREENSHOT 1: Search Quality tab with the default run results visible (sample size 10, ANN vs kNN). Filename: search-quality-tab.png -->
[Loading a dataset from Hugging Face hub](/documentation/tutorials-basics/huggingface-datasets/) tutorial, `Qdrant/arxiv-titles-instructorxl-embeddings` ![Search Quality tab with default evaluation results](/documentation/tutorials/retrieval-quality/search-quality-tab.png)
from the [Hugging Face hub](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings). Let's download it in a streaming
mode, as we are only going to use part of it.
```python The tab reports average **precision@k**. The score is typically high but not always perfect. When you need higher recall and can accept higher latency or more memory, HNSW is tunable.
from datasets import load_dataset
dataset = load_dataset( ## Tweaking the HNSW Parameters
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
)
```
We need some data to be indexed and another set for the testing purposes. Let's get the first 60000 items for the training and the next 1000 HNSW is a hierarchical graph where each node has a set of links to other nodes. The `m` parameter controls the number of edges per node: higher `m` means higher recall at the cost of more memory. The `ef_construct` parameter controls how many neighbours are considered during index building: higher `ef_construct` means higher recall at the cost of longer indexing time. Defaults are `m=16` and `ef_construct=100`.
for the testing.
```python For the full list of HNSW parameters, including on-disk storage and precision/memory trade-offs, see [Optimize Performance](/documentation/operations/optimize/).
dataset_iterator = iter(dataset)
train_dataset = [next(dataset_iterator) for _ in range(60000)]
test_dataset = [next(dataset_iterator) for _ in range(1000)]
```
Now, let's create a collection and index the training data. This collection will be created with the default configuration. Please be aware that Toggle **advanced mode** in the Search Quality tab to tune these parameters inline. Raise `m` to 32 and `ef_construct` to 200, then run the evaluation again.
it might be different from your collection settings, and it's always important to test exactly the same configuration you are going to use later
in production.
<aside role="status"> <!-- SCREENSHOT 2: Advanced mode panel with HNSW parameter inputs (m, ef_construct) visible. Filename: search-quality-advanced.png -->
Distance function is another parameter that may impact the retrieval quality. If the embedding model was not trained to minimize cosine ![Search Quality advanced mode with HNSW parameters](/documentation/tutorials/retrieval-quality/search-quality-advanced.png)
distance, you can get suboptimal search results by using it. Please test different distance functions to find the best one for your embeddings,
if you don't know the specifics of the model training. Precision should increase at the cost of higher build time and memory.
</aside>
<!-- SCREENSHOT 3: Results view after m=32 / ef_construct=200, showing higher precision than the first run. Filename: search-quality-after-tuning.png -->
![Search Quality results after HNSW tuning](/documentation/tutorials/retrieval-quality/search-quality-after-tuning.png)
Tune until you hit the point that matches your quality and cost targets.
## Automate in CI with Python
The Web UI is the fastest way to check recall interactively. For continuous integration or scripted regression tests, the Qdrant client exposes the same exact-search mode via `search_params=models.SearchParams(exact=True)`. Compare the ANN and exact top-k sets yourself and compute recall.
The helper below takes a list of query vectors and returns the average recall@k. Supply your own test set: a representative sample of query vectors from your workload, held out from training.
```python ```python
from qdrant_client import QdrantClient, models from qdrant_client import QdrantClient, models
client = QdrantClient("http://localhost:6333")
client.create_collection(
collection_name="arxiv-titles-instructorxl-embeddings",
vectors_config=models.VectorParams(
size=768, # Size of the embeddings generated by InstructorXL model
distance=models.Distance.COSINE,
),
)
```
We are now ready to index the training data. Uploading the records is going to trigger the indexing process, which will build the HNSW graph. def avg_recall_at_k(
The indexing process may take some time, depending on the size of the dataset, but your data is going to be available for search immediately client: QdrantClient,
after receiving the response from the `upsert` endpoint. **As long as the indexing is not finished, and HNSW not built, Qdrant will perform collection_name: str,
the exact search**. We have to wait until the indexing is finished to be sure that the approximate search is performed. test_vectors: list,
k: int,
```python ) -> float:
client.upload_points( # upload_points is available as of qdrant-client v1.7.1
collection_name="arxiv-titles-instructorxl-embeddings",
points=[
models.PointStruct(
id=item["id"],
vector=item["vector"],
payload=item,
)
for item in train_dataset
]
)
while True:
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
if collection_info.status == models.CollectionStatus.GREEN:
# Collection status is green, which means the indexing is finished
break
```
## Standard Mode vs Exact Search
Qdrant has a built-in exact search mode, which can be used to measure the quality of the search results. In this mode, Qdrant performs a
full kNN search for each query, without any approximation. It is not suitable for production use with high load, but it is perfect for the
evaluation of the ANN algorithm and its parameters. It might be triggered by setting the `exact` parameter to `True` in the search request.
We are simply going to use all the examples from the test dataset as queries and compare the results of the approximate search with the
results of the exact search. Let's create a helper function with `k` being a parameter, so we can calculate the `recall@k` for different
values of `k`.
```python
def avg_recall_at_k(k: int):
recalls = [] recalls = []
for item in test_dataset: for vector in test_vectors:
ann_result = client.query_points( ann_ids = {
collection_name="arxiv-titles-instructorxl-embeddings", p.id for p in client.query_points(
query=item["vector"], collection_name=collection_name,
limit=k, query=vector,
).points limit=k,
).points
knn_result = client.query_points( }
collection_name="arxiv-titles-instructorxl-embeddings", knn_ids = {
query=item["vector"], p.id for p in client.query_points(
limit=k, collection_name=collection_name,
search_params=models.SearchParams( query=vector,
exact=True, # Turns on the exact search mode limit=k,
), search_params=models.SearchParams(exact=True),
).points ).points
}
# We can calculate the recall@k by comparing the ids of the search results recalls.append(len(ann_ids & knn_ids) / k)
ann_ids = set(item.id for item in ann_result)
knn_ids = set(item.id for item in knn_result)
recall = len(ann_ids.intersection(knn_ids)) / k
recalls.append(recall)
return sum(recalls) / len(recalls) return sum(recalls) / len(recalls)
``` ```
Calculating the `recall@5` is as simple as calling the function with the corresponding parameter: Drop it into your CI pipeline and fail the job if recall drops below a threshold after an embedding model change or index config update.
```python
print(f"avg(recall@5) = {avg_recall_at_k(k=5)}")
```
Response:
```text
avg(recall@5) = 0.9935999999999995
```
As we can see, the recall of the approximate search vs exact search is pretty high. There are, however, some scenarios when we
need higher recall and can accept higher latency. HNSW is pretty tunable, and we can increase the recall by changing its parameters.
## Tweaking the HNSW Parameters
HNSW is a hierarchical graph, where each node has a set of links to other nodes. The number of edges per node is called the `m` parameter.
The larger the value of it, the higher the recall of the search, but more space required. The `ef_construct` parameter is the number of
neighbours to consider during the index building. Again, the larger the value, the higher the recall, but the longer the indexing time.
The default values of these parameters are `m=16` and `ef_construct=100`. Let's try to increase them to `m=32` and `ef_construct=200` and
see how it affects the recall. Of course, we need to wait until the indexing is finished before we can perform the search.
```python
client.update_collection(
collection_name="arxiv-titles-instructorxl-embeddings",
hnsw_config=models.HnswConfigDiff(
m=32, # Increase the number of edges per node from the default 16 to 32
ef_construct=200, # Increase the number of neighbours from the default 100 to 200
)
)
while True:
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
if collection_info.status == models.CollectionStatus.GREEN:
# Collection status is green, which means the indexing is finished
break
```
The same function can be used to calculate the average `recall@5`:
```python
print(f"avg(recall@5) = {avg_recall_at_k(k=5)}")
```
Response:
```text
avg(recall@5) = 0.9969999999999998
```
The recall has obviously increased, and we know how to control it. However, there is a trade-off between the recall and the search
latency and memory requirements. In some specific cases, we may want to increase the recall as much as possible, so now we know how
to do it.
## Wrapping Up ## Wrapping Up
Assessing the quality of retrieval is a critical aspect of evaluating semantic search performance. It is imperative to measure retrieval quality when aiming for optimal quality of. Measuring ANN recall keeps HNSW tuning honest. The Search Quality tab gives you a quick interactive read; the Python helper above plugs into CI to catch regressions after embedding model changes or index config updates.
your search results. Qdrant provides a built-in exact search mode, which can be used to measure the quality of the ANN algorithm itself,
even in an automated way, as part of your CI/CD pipeline.
Again, **the quality of the embeddings is the most important factor**. HNSW does a pretty good job in terms of recall, and it is HNSW covers most workloads well and is tunable when you need more recall. Other ANN algorithms exist, such as [IVF*](https://github.com/facebookresearch/faiss/wiki/Faiss-indexes#cell-probe-methods-indexivf-indexes), but they generally [perform worse than HNSW on quality and performance](https://nirantk.com/writing/pgvector-vs-qdrant/#correctness).
parameterizable and tunable, when required. There are some other ANN algorithms available out there, such as [IVF*](https://github.com/facebookresearch/faiss/wiki/Faiss-indexes#cell-probe-methods-indexivf-indexes),
but they usually [perform worse than HNSW in terms of quality and performance](https://nirantk.com/writing/pgvector-vs-qdrant/#correctness).