mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-11 13:58:30 +02:00
retrieval-quality: pivot to Web UI, add CI skeleton, trim setup
- Drop the dataset-setup walkthrough (HF loading, collection
create, upload, wait-for-green). Readers at this phase already
have a collection.
- Replace the Python evaluation block with a "Measure ANN Recall
with the Web UI" section built around the Search Quality tab.
Default run is one-click (sample size 10); HNSW tuning uses the
tab's advanced mode instead of update_collection. Three
screenshot placeholders at
/documentation/tutorials/retrieval-quality/*.png.
- Reflect that the tab reports precision@k; note the recall@k
equivalence already spelled out in the ANN Recall section.
- Keep Python but move it to an "Automate in CI" section with a
reusable skeleton function.
- Collapse the standalone "Embeddings Quality" section into a
one-sentence MTEB pointer inside ANN Recall.
- Rewrite Wrapping Up to match the new scope.
- Link the HNSW tuning section to Optimize Performance for the
full parameter reference.
This commit is contained in:
+47
-161
@@ -15,18 +15,9 @@ This tutorial measures **layer 1** of the <a href="/documentation/tutorials-sear
|
||||
|
||||
We'll measure Qdrant's ANN recall with `recall@k` and tune HNSW parameters to control the recall/latency trade-off. The ANN algorithm is one of several levers that shape retrieval quality in a production pipeline, alongside the embedding model, retrieval strategy (dense, sparse, hybrid, and multi-vector), filtering, and reranking.
|
||||
|
||||
## Embeddings Quality
|
||||
|
||||
The quality of the embeddings is a topic for a separate tutorial. In a nutshell, it is usually measured and compared by benchmarks, such as
|
||||
[Massive Text Embedding Benchmark (MTEB)](https://huggingface.co/spaces/mteb/leaderboard). The evaluation process itself is pretty
|
||||
straightforward and is based on a ground truth dataset built by humans. We have a set of queries and a set of the documents we would expect
|
||||
to receive for each of them. In the evaluation process, we take a query, find the most similar documents in the vector space and compare
|
||||
them with the ground truth. In that setup, **finding the most similar documents is implemented as full kNN search, without any approximation**.
|
||||
As a result, we can measure the quality of the embeddings themselves, without the influence of the ANN algorithm.
|
||||
|
||||
## ANN Recall
|
||||
|
||||
The embedding model sets a baseline for search quality, but the retrieval pipeline can still underperform it. Vector search engines such as Qdrant don't run pure kNN at query time; they use **approximate nearest-neighbor** (ANN) algorithms for speed. ANN is faster than exact search but can return suboptimal results. **ANN recall** measures that gap.
|
||||
Embedding quality sets the ceiling on search quality and is measured separately via benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard). The retrieval pipeline can still underperform that ceiling: vector search engines such as Qdrant don't run pure kNN at query time but use **approximate nearest-neighbor** (ANN) algorithms for speed. ANN is faster than exact search but can return suboptimal results. **ANN recall** measures that gap.
|
||||
|
||||
For a broader discussion of what to measure and when (ANN recall vs retrieval relevance vs business impact, and which metric fits which scenario),
|
||||
see [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/). This tutorial focuses on the
|
||||
@@ -34,180 +25,75 @@ ANN-algorithm layer and measures it with `recall@k`: the fraction of the true to
|
||||
recovers. When both ANN and exact search return exactly `k` items, `recall@k` and `precision@k` are numerically identical; we use "recall" to
|
||||
match the ANN-benchmarks convention.
|
||||
|
||||
## Measure the Quality of the Search Results
|
||||
## Measure ANN Recall with the Web UI
|
||||
|
||||
Let's build a quality evaluation of the ANN algorithm in Qdrant. We will, first, call the search endpoint in a standard way to obtain
|
||||
the approximate search results. Then, we will call the exact search endpoint to obtain the exact matches, and finally compare both results
|
||||
in terms of recall.
|
||||
Qdrant's Web UI has a Search Quality tab that measures the gap between approximate and exact search without requiring evaluation code. Open the dashboard at `http://localhost:6333/dashboard` (or your cluster's dashboard on Qdrant Cloud), navigate to your collection, and click the Search Quality tab. A run launches automatically with a default sample size of 10 queries, comparing ANN against exact kNN.
|
||||
|
||||
Before we start, let's create a collection, fill it with some data and then start our evaluation. We will use the same dataset as in the
|
||||
[Loading a dataset from Hugging Face hub](/documentation/tutorials-basics/huggingface-datasets/) tutorial, `Qdrant/arxiv-titles-instructorxl-embeddings`
|
||||
from the [Hugging Face hub](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings). Let's download it in a streaming
|
||||
mode, as we are only going to use part of it.
|
||||
<!-- SCREENSHOT 1: Search Quality tab with the default run results visible (sample size 10, ANN vs kNN). Filename: search-quality-tab.png -->
|
||||

|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
The tab reports average **precision@k**. The score is typically high but not always perfect. When you need higher recall and can accept higher latency or more memory, HNSW is tunable.
|
||||
|
||||
dataset = load_dataset(
|
||||
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
|
||||
)
|
||||
```
|
||||
## Tweaking the HNSW Parameters
|
||||
|
||||
We need some data to be indexed and another set for the testing purposes. Let's get the first 60000 items for the training and the next 1000
|
||||
for the testing.
|
||||
HNSW is a hierarchical graph where each node has a set of links to other nodes. The `m` parameter controls the number of edges per node: higher `m` means higher recall at the cost of more memory. The `ef_construct` parameter controls how many neighbours are considered during index building: higher `ef_construct` means higher recall at the cost of longer indexing time. Defaults are `m=16` and `ef_construct=100`.
|
||||
|
||||
```python
|
||||
dataset_iterator = iter(dataset)
|
||||
train_dataset = [next(dataset_iterator) for _ in range(60000)]
|
||||
test_dataset = [next(dataset_iterator) for _ in range(1000)]
|
||||
```
|
||||
For the full list of HNSW parameters, including on-disk storage and precision/memory trade-offs, see [Optimize Performance](/documentation/operations/optimize/).
|
||||
|
||||
Now, let's create a collection and index the training data. This collection will be created with the default configuration. Please be aware that
|
||||
it might be different from your collection settings, and it's always important to test exactly the same configuration you are going to use later
|
||||
in production.
|
||||
Toggle **advanced mode** in the Search Quality tab to tune these parameters inline. Raise `m` to 32 and `ef_construct` to 200, then run the evaluation again.
|
||||
|
||||
<aside role="status">
|
||||
Distance function is another parameter that may impact the retrieval quality. If the embedding model was not trained to minimize cosine
|
||||
distance, you can get suboptimal search results by using it. Please test different distance functions to find the best one for your embeddings,
|
||||
if you don't know the specifics of the model training.
|
||||
</aside>
|
||||
<!-- SCREENSHOT 2: Advanced mode panel with HNSW parameter inputs (m, ef_construct) visible. Filename: search-quality-advanced.png -->
|
||||

|
||||
|
||||
Precision should increase at the cost of higher build time and memory.
|
||||
|
||||
<!-- SCREENSHOT 3: Results view after m=32 / ef_construct=200, showing higher precision than the first run. Filename: search-quality-after-tuning.png -->
|
||||

|
||||
|
||||
Tune until you hit the point that matches your quality and cost targets.
|
||||
|
||||
## Automate in CI with Python
|
||||
|
||||
The Web UI is the fastest way to check recall interactively. For continuous integration or scripted regression tests, the Qdrant client exposes the same exact-search mode via `search_params=models.SearchParams(exact=True)`. Compare the ANN and exact top-k sets yourself and compute recall.
|
||||
|
||||
The helper below takes a list of query vectors and returns the average recall@k. Supply your own test set: a representative sample of query vectors from your workload, held out from training.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("http://localhost:6333")
|
||||
client.create_collection(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
vectors_config=models.VectorParams(
|
||||
size=768, # Size of the embeddings generated by InstructorXL model
|
||||
distance=models.Distance.COSINE,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
We are now ready to index the training data. Uploading the records is going to trigger the indexing process, which will build the HNSW graph.
|
||||
The indexing process may take some time, depending on the size of the dataset, but your data is going to be available for search immediately
|
||||
after receiving the response from the `upsert` endpoint. **As long as the indexing is not finished, and HNSW not built, Qdrant will perform
|
||||
the exact search**. We have to wait until the indexing is finished to be sure that the approximate search is performed.
|
||||
|
||||
```python
|
||||
client.upload_points( # upload_points is available as of qdrant-client v1.7.1
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=item["id"],
|
||||
vector=item["vector"],
|
||||
payload=item,
|
||||
)
|
||||
for item in train_dataset
|
||||
]
|
||||
)
|
||||
|
||||
while True:
|
||||
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
|
||||
if collection_info.status == models.CollectionStatus.GREEN:
|
||||
# Collection status is green, which means the indexing is finished
|
||||
break
|
||||
```
|
||||
|
||||
## Standard Mode vs Exact Search
|
||||
|
||||
Qdrant has a built-in exact search mode, which can be used to measure the quality of the search results. In this mode, Qdrant performs a
|
||||
full kNN search for each query, without any approximation. It is not suitable for production use with high load, but it is perfect for the
|
||||
evaluation of the ANN algorithm and its parameters. It might be triggered by setting the `exact` parameter to `True` in the search request.
|
||||
We are simply going to use all the examples from the test dataset as queries and compare the results of the approximate search with the
|
||||
results of the exact search. Let's create a helper function with `k` being a parameter, so we can calculate the `recall@k` for different
|
||||
values of `k`.
|
||||
|
||||
```python
|
||||
def avg_recall_at_k(k: int):
|
||||
def avg_recall_at_k(
|
||||
client: QdrantClient,
|
||||
collection_name: str,
|
||||
test_vectors: list,
|
||||
k: int,
|
||||
) -> float:
|
||||
recalls = []
|
||||
for item in test_dataset:
|
||||
ann_result = client.query_points(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
query=item["vector"],
|
||||
for vector in test_vectors:
|
||||
ann_ids = {
|
||||
p.id for p in client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=vector,
|
||||
limit=k,
|
||||
).points
|
||||
|
||||
knn_result = client.query_points(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
query=item["vector"],
|
||||
}
|
||||
knn_ids = {
|
||||
p.id for p in client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=vector,
|
||||
limit=k,
|
||||
search_params=models.SearchParams(
|
||||
exact=True, # Turns on the exact search mode
|
||||
),
|
||||
search_params=models.SearchParams(exact=True),
|
||||
).points
|
||||
|
||||
# We can calculate the recall@k by comparing the ids of the search results
|
||||
ann_ids = set(item.id for item in ann_result)
|
||||
knn_ids = set(item.id for item in knn_result)
|
||||
recall = len(ann_ids.intersection(knn_ids)) / k
|
||||
recalls.append(recall)
|
||||
}
|
||||
recalls.append(len(ann_ids & knn_ids) / k)
|
||||
|
||||
return sum(recalls) / len(recalls)
|
||||
```
|
||||
|
||||
Calculating the `recall@5` is as simple as calling the function with the corresponding parameter:
|
||||
|
||||
```python
|
||||
print(f"avg(recall@5) = {avg_recall_at_k(k=5)}")
|
||||
```
|
||||
|
||||
Response:
|
||||
|
||||
```text
|
||||
avg(recall@5) = 0.9935999999999995
|
||||
```
|
||||
|
||||
As we can see, the recall of the approximate search vs exact search is pretty high. There are, however, some scenarios when we
|
||||
need higher recall and can accept higher latency. HNSW is pretty tunable, and we can increase the recall by changing its parameters.
|
||||
|
||||
## Tweaking the HNSW Parameters
|
||||
|
||||
HNSW is a hierarchical graph, where each node has a set of links to other nodes. The number of edges per node is called the `m` parameter.
|
||||
The larger the value of it, the higher the recall of the search, but more space required. The `ef_construct` parameter is the number of
|
||||
neighbours to consider during the index building. Again, the larger the value, the higher the recall, but the longer the indexing time.
|
||||
The default values of these parameters are `m=16` and `ef_construct=100`. Let's try to increase them to `m=32` and `ef_construct=200` and
|
||||
see how it affects the recall. Of course, we need to wait until the indexing is finished before we can perform the search.
|
||||
|
||||
```python
|
||||
client.update_collection(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=32, # Increase the number of edges per node from the default 16 to 32
|
||||
ef_construct=200, # Increase the number of neighbours from the default 100 to 200
|
||||
)
|
||||
)
|
||||
|
||||
while True:
|
||||
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
|
||||
if collection_info.status == models.CollectionStatus.GREEN:
|
||||
# Collection status is green, which means the indexing is finished
|
||||
break
|
||||
```
|
||||
|
||||
The same function can be used to calculate the average `recall@5`:
|
||||
|
||||
```python
|
||||
print(f"avg(recall@5) = {avg_recall_at_k(k=5)}")
|
||||
```
|
||||
|
||||
Response:
|
||||
|
||||
```text
|
||||
avg(recall@5) = 0.9969999999999998
|
||||
```
|
||||
|
||||
The recall has obviously increased, and we know how to control it. However, there is a trade-off between the recall and the search
|
||||
latency and memory requirements. In some specific cases, we may want to increase the recall as much as possible, so now we know how
|
||||
to do it.
|
||||
Drop it into your CI pipeline and fail the job if recall drops below a threshold after an embedding model change or index config update.
|
||||
|
||||
## Wrapping Up
|
||||
|
||||
Assessing the quality of retrieval is a critical aspect of evaluating semantic search performance. It is imperative to measure retrieval quality when aiming for optimal quality of.
|
||||
your search results. Qdrant provides a built-in exact search mode, which can be used to measure the quality of the ANN algorithm itself,
|
||||
even in an automated way, as part of your CI/CD pipeline.
|
||||
Measuring ANN recall keeps HNSW tuning honest. The Search Quality tab gives you a quick interactive read; the Python helper above plugs into CI to catch regressions after embedding model changes or index config updates.
|
||||
|
||||
Again, **the quality of the embeddings is the most important factor**. HNSW does a pretty good job in terms of recall, and it is
|
||||
parameterizable and tunable, when required. There are some other ANN algorithms available out there, such as [IVF*](https://github.com/facebookresearch/faiss/wiki/Faiss-indexes#cell-probe-methods-indexivf-indexes),
|
||||
but they usually [perform worse than HNSW in terms of quality and performance](https://nirantk.com/writing/pgvector-vs-qdrant/#correctness).
|
||||
HNSW covers most workloads well and is tunable when you need more recall. Other ANN algorithms exist, such as [IVF*](https://github.com/facebookresearch/faiss/wiki/Faiss-indexes#cell-probe-methods-indexivf-indexes), but they generally [perform worse than HNSW on quality and performance](https://nirantk.com/writing/pgvector-vs-qdrant/#correctness).
|
||||
|
||||
Reference in New Issue
Block a user