mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-11 13:58:30 +02:00
retrieval-quality: pivot to Web UI, add CI skeleton, trim setup
- Drop the dataset-setup walkthrough (HF loading, collection
create, upload, wait-for-green). Readers at this phase already
have a collection.
- Replace the Python evaluation block with a "Measure ANN Recall
with the Web UI" section built around the Search Quality tab.
Default run is one-click (sample size 10); HNSW tuning uses the
tab's advanced mode instead of update_collection. Three
screenshot placeholders at
/documentation/tutorials/retrieval-quality/*.png.
- Reflect that the tab reports precision@k; note the recall@k
equivalence already spelled out in the ANN Recall section.
- Keep Python but move it to an "Automate in CI" section with a
reusable skeleton function.
- Collapse the standalone "Embeddings Quality" section into a
one-sentence MTEB pointer inside ANN Recall.
- Rewrite Wrapping Up to match the new scope.
- Link the HNSW tuning section to Optimize Performance for the
full parameter reference.
This commit is contained in:
+51
-165
@@ -15,18 +15,9 @@ This tutorial measures **layer 1** of the <a href="/documentation/tutorials-sear
|
|||||||
|
|
||||||
We'll measure Qdrant's ANN recall with `recall@k` and tune HNSW parameters to control the recall/latency trade-off. The ANN algorithm is one of several levers that shape retrieval quality in a production pipeline, alongside the embedding model, retrieval strategy (dense, sparse, hybrid, and multi-vector), filtering, and reranking.
|
We'll measure Qdrant's ANN recall with `recall@k` and tune HNSW parameters to control the recall/latency trade-off. The ANN algorithm is one of several levers that shape retrieval quality in a production pipeline, alongside the embedding model, retrieval strategy (dense, sparse, hybrid, and multi-vector), filtering, and reranking.
|
||||||
|
|
||||||
## Embeddings Quality
|
|
||||||
|
|
||||||
The quality of the embeddings is a topic for a separate tutorial. In a nutshell, it is usually measured and compared by benchmarks, such as
|
|
||||||
[Massive Text Embedding Benchmark (MTEB)](https://huggingface.co/spaces/mteb/leaderboard). The evaluation process itself is pretty
|
|
||||||
straightforward and is based on a ground truth dataset built by humans. We have a set of queries and a set of the documents we would expect
|
|
||||||
to receive for each of them. In the evaluation process, we take a query, find the most similar documents in the vector space and compare
|
|
||||||
them with the ground truth. In that setup, **finding the most similar documents is implemented as full kNN search, without any approximation**.
|
|
||||||
As a result, we can measure the quality of the embeddings themselves, without the influence of the ANN algorithm.
|
|
||||||
|
|
||||||
## ANN Recall
|
## ANN Recall
|
||||||
|
|
||||||
The embedding model sets a baseline for search quality, but the retrieval pipeline can still underperform it. Vector search engines such as Qdrant don't run pure kNN at query time; they use **approximate nearest-neighbor** (ANN) algorithms for speed. ANN is faster than exact search but can return suboptimal results. **ANN recall** measures that gap.
|
Embedding quality sets the ceiling on search quality and is measured separately via benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard). The retrieval pipeline can still underperform that ceiling: vector search engines such as Qdrant don't run pure kNN at query time but use **approximate nearest-neighbor** (ANN) algorithms for speed. ANN is faster than exact search but can return suboptimal results. **ANN recall** measures that gap.
|
||||||
|
|
||||||
For a broader discussion of what to measure and when (ANN recall vs retrieval relevance vs business impact, and which metric fits which scenario),
|
For a broader discussion of what to measure and when (ANN recall vs retrieval relevance vs business impact, and which metric fits which scenario),
|
||||||
see [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/). This tutorial focuses on the
|
see [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/). This tutorial focuses on the
|
||||||
@@ -34,180 +25,75 @@ ANN-algorithm layer and measures it with `recall@k`: the fraction of the true to
|
|||||||
recovers. When both ANN and exact search return exactly `k` items, `recall@k` and `precision@k` are numerically identical; we use "recall" to
|
recovers. When both ANN and exact search return exactly `k` items, `recall@k` and `precision@k` are numerically identical; we use "recall" to
|
||||||
match the ANN-benchmarks convention.
|
match the ANN-benchmarks convention.
|
||||||
|
|
||||||
## Measure the Quality of the Search Results
|
## Measure ANN Recall with the Web UI
|
||||||
|
|
||||||
Let's build a quality evaluation of the ANN algorithm in Qdrant. We will, first, call the search endpoint in a standard way to obtain
|
Qdrant's Web UI has a Search Quality tab that measures the gap between approximate and exact search without requiring evaluation code. Open the dashboard at `http://localhost:6333/dashboard` (or your cluster's dashboard on Qdrant Cloud), navigate to your collection, and click the Search Quality tab. A run launches automatically with a default sample size of 10 queries, comparing ANN against exact kNN.
|
||||||
the approximate search results. Then, we will call the exact search endpoint to obtain the exact matches, and finally compare both results
|
|
||||||
in terms of recall.
|
|
||||||
|
|
||||||
Before we start, let's create a collection, fill it with some data and then start our evaluation. We will use the same dataset as in the
|
<!-- SCREENSHOT 1: Search Quality tab with the default run results visible (sample size 10, ANN vs kNN). Filename: search-quality-tab.png -->
|
||||||
[Loading a dataset from Hugging Face hub](/documentation/tutorials-basics/huggingface-datasets/) tutorial, `Qdrant/arxiv-titles-instructorxl-embeddings`
|

|
||||||
from the [Hugging Face hub](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings). Let's download it in a streaming
|
|
||||||
mode, as we are only going to use part of it.
|
|
||||||
|
|
||||||
```python
|
The tab reports average **precision@k**. The score is typically high but not always perfect. When you need higher recall and can accept higher latency or more memory, HNSW is tunable.
|
||||||
from datasets import load_dataset
|
|
||||||
|
|
||||||
dataset = load_dataset(
|
## Tweaking the HNSW Parameters
|
||||||
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
We need some data to be indexed and another set for the testing purposes. Let's get the first 60000 items for the training and the next 1000
|
HNSW is a hierarchical graph where each node has a set of links to other nodes. The `m` parameter controls the number of edges per node: higher `m` means higher recall at the cost of more memory. The `ef_construct` parameter controls how many neighbours are considered during index building: higher `ef_construct` means higher recall at the cost of longer indexing time. Defaults are `m=16` and `ef_construct=100`.
|
||||||
for the testing.
|
|
||||||
|
|
||||||
```python
|
For the full list of HNSW parameters, including on-disk storage and precision/memory trade-offs, see [Optimize Performance](/documentation/operations/optimize/).
|
||||||
dataset_iterator = iter(dataset)
|
|
||||||
train_dataset = [next(dataset_iterator) for _ in range(60000)]
|
|
||||||
test_dataset = [next(dataset_iterator) for _ in range(1000)]
|
|
||||||
```
|
|
||||||
|
|
||||||
Now, let's create a collection and index the training data. This collection will be created with the default configuration. Please be aware that
|
Toggle **advanced mode** in the Search Quality tab to tune these parameters inline. Raise `m` to 32 and `ef_construct` to 200, then run the evaluation again.
|
||||||
it might be different from your collection settings, and it's always important to test exactly the same configuration you are going to use later
|
|
||||||
in production.
|
|
||||||
|
|
||||||
<aside role="status">
|
<!-- SCREENSHOT 2: Advanced mode panel with HNSW parameter inputs (m, ef_construct) visible. Filename: search-quality-advanced.png -->
|
||||||
Distance function is another parameter that may impact the retrieval quality. If the embedding model was not trained to minimize cosine
|

|
||||||
distance, you can get suboptimal search results by using it. Please test different distance functions to find the best one for your embeddings,
|
|
||||||
if you don't know the specifics of the model training.
|
Precision should increase at the cost of higher build time and memory.
|
||||||
</aside>
|
|
||||||
|
<!-- SCREENSHOT 3: Results view after m=32 / ef_construct=200, showing higher precision than the first run. Filename: search-quality-after-tuning.png -->
|
||||||
|

|
||||||
|
|
||||||
|
Tune until you hit the point that matches your quality and cost targets.
|
||||||
|
|
||||||
|
## Automate in CI with Python
|
||||||
|
|
||||||
|
The Web UI is the fastest way to check recall interactively. For continuous integration or scripted regression tests, the Qdrant client exposes the same exact-search mode via `search_params=models.SearchParams(exact=True)`. Compare the ANN and exact top-k sets yourself and compute recall.
|
||||||
|
|
||||||
|
The helper below takes a list of query vectors and returns the average recall@k. Supply your own test set: a representative sample of query vectors from your workload, held out from training.
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from qdrant_client import QdrantClient, models
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
client = QdrantClient("http://localhost:6333")
|
|
||||||
client.create_collection(
|
|
||||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
|
||||||
vectors_config=models.VectorParams(
|
|
||||||
size=768, # Size of the embeddings generated by InstructorXL model
|
|
||||||
distance=models.Distance.COSINE,
|
|
||||||
),
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
We are now ready to index the training data. Uploading the records is going to trigger the indexing process, which will build the HNSW graph.
|
def avg_recall_at_k(
|
||||||
The indexing process may take some time, depending on the size of the dataset, but your data is going to be available for search immediately
|
client: QdrantClient,
|
||||||
after receiving the response from the `upsert` endpoint. **As long as the indexing is not finished, and HNSW not built, Qdrant will perform
|
collection_name: str,
|
||||||
the exact search**. We have to wait until the indexing is finished to be sure that the approximate search is performed.
|
test_vectors: list,
|
||||||
|
k: int,
|
||||||
```python
|
) -> float:
|
||||||
client.upload_points( # upload_points is available as of qdrant-client v1.7.1
|
|
||||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
|
||||||
points=[
|
|
||||||
models.PointStruct(
|
|
||||||
id=item["id"],
|
|
||||||
vector=item["vector"],
|
|
||||||
payload=item,
|
|
||||||
)
|
|
||||||
for item in train_dataset
|
|
||||||
]
|
|
||||||
)
|
|
||||||
|
|
||||||
while True:
|
|
||||||
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
|
|
||||||
if collection_info.status == models.CollectionStatus.GREEN:
|
|
||||||
# Collection status is green, which means the indexing is finished
|
|
||||||
break
|
|
||||||
```
|
|
||||||
|
|
||||||
## Standard Mode vs Exact Search
|
|
||||||
|
|
||||||
Qdrant has a built-in exact search mode, which can be used to measure the quality of the search results. In this mode, Qdrant performs a
|
|
||||||
full kNN search for each query, without any approximation. It is not suitable for production use with high load, but it is perfect for the
|
|
||||||
evaluation of the ANN algorithm and its parameters. It might be triggered by setting the `exact` parameter to `True` in the search request.
|
|
||||||
We are simply going to use all the examples from the test dataset as queries and compare the results of the approximate search with the
|
|
||||||
results of the exact search. Let's create a helper function with `k` being a parameter, so we can calculate the `recall@k` for different
|
|
||||||
values of `k`.
|
|
||||||
|
|
||||||
```python
|
|
||||||
def avg_recall_at_k(k: int):
|
|
||||||
recalls = []
|
recalls = []
|
||||||
for item in test_dataset:
|
for vector in test_vectors:
|
||||||
ann_result = client.query_points(
|
ann_ids = {
|
||||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
p.id for p in client.query_points(
|
||||||
query=item["vector"],
|
collection_name=collection_name,
|
||||||
limit=k,
|
query=vector,
|
||||||
).points
|
limit=k,
|
||||||
|
).points
|
||||||
knn_result = client.query_points(
|
}
|
||||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
knn_ids = {
|
||||||
query=item["vector"],
|
p.id for p in client.query_points(
|
||||||
limit=k,
|
collection_name=collection_name,
|
||||||
search_params=models.SearchParams(
|
query=vector,
|
||||||
exact=True, # Turns on the exact search mode
|
limit=k,
|
||||||
),
|
search_params=models.SearchParams(exact=True),
|
||||||
).points
|
).points
|
||||||
|
}
|
||||||
|
recalls.append(len(ann_ids & knn_ids) / k)
|
||||||
|
|
||||||
# We can calculate the recall@k by comparing the ids of the search results
|
|
||||||
ann_ids = set(item.id for item in ann_result)
|
|
||||||
knn_ids = set(item.id for item in knn_result)
|
|
||||||
recall = len(ann_ids.intersection(knn_ids)) / k
|
|
||||||
recalls.append(recall)
|
|
||||||
|
|
||||||
return sum(recalls) / len(recalls)
|
return sum(recalls) / len(recalls)
|
||||||
```
|
```
|
||||||
|
|
||||||
Calculating the `recall@5` is as simple as calling the function with the corresponding parameter:
|
Drop it into your CI pipeline and fail the job if recall drops below a threshold after an embedding model change or index config update.
|
||||||
|
|
||||||
```python
|
|
||||||
print(f"avg(recall@5) = {avg_recall_at_k(k=5)}")
|
|
||||||
```
|
|
||||||
|
|
||||||
Response:
|
|
||||||
|
|
||||||
```text
|
|
||||||
avg(recall@5) = 0.9935999999999995
|
|
||||||
```
|
|
||||||
|
|
||||||
As we can see, the recall of the approximate search vs exact search is pretty high. There are, however, some scenarios when we
|
|
||||||
need higher recall and can accept higher latency. HNSW is pretty tunable, and we can increase the recall by changing its parameters.
|
|
||||||
|
|
||||||
## Tweaking the HNSW Parameters
|
|
||||||
|
|
||||||
HNSW is a hierarchical graph, where each node has a set of links to other nodes. The number of edges per node is called the `m` parameter.
|
|
||||||
The larger the value of it, the higher the recall of the search, but more space required. The `ef_construct` parameter is the number of
|
|
||||||
neighbours to consider during the index building. Again, the larger the value, the higher the recall, but the longer the indexing time.
|
|
||||||
The default values of these parameters are `m=16` and `ef_construct=100`. Let's try to increase them to `m=32` and `ef_construct=200` and
|
|
||||||
see how it affects the recall. Of course, we need to wait until the indexing is finished before we can perform the search.
|
|
||||||
|
|
||||||
```python
|
|
||||||
client.update_collection(
|
|
||||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
|
||||||
hnsw_config=models.HnswConfigDiff(
|
|
||||||
m=32, # Increase the number of edges per node from the default 16 to 32
|
|
||||||
ef_construct=200, # Increase the number of neighbours from the default 100 to 200
|
|
||||||
)
|
|
||||||
)
|
|
||||||
|
|
||||||
while True:
|
|
||||||
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
|
|
||||||
if collection_info.status == models.CollectionStatus.GREEN:
|
|
||||||
# Collection status is green, which means the indexing is finished
|
|
||||||
break
|
|
||||||
```
|
|
||||||
|
|
||||||
The same function can be used to calculate the average `recall@5`:
|
|
||||||
|
|
||||||
```python
|
|
||||||
print(f"avg(recall@5) = {avg_recall_at_k(k=5)}")
|
|
||||||
```
|
|
||||||
|
|
||||||
Response:
|
|
||||||
|
|
||||||
```text
|
|
||||||
avg(recall@5) = 0.9969999999999998
|
|
||||||
```
|
|
||||||
|
|
||||||
The recall has obviously increased, and we know how to control it. However, there is a trade-off between the recall and the search
|
|
||||||
latency and memory requirements. In some specific cases, we may want to increase the recall as much as possible, so now we know how
|
|
||||||
to do it.
|
|
||||||
|
|
||||||
## Wrapping Up
|
## Wrapping Up
|
||||||
|
|
||||||
Assessing the quality of retrieval is a critical aspect of evaluating semantic search performance. It is imperative to measure retrieval quality when aiming for optimal quality of.
|
Measuring ANN recall keeps HNSW tuning honest. The Search Quality tab gives you a quick interactive read; the Python helper above plugs into CI to catch regressions after embedding model changes or index config updates.
|
||||||
your search results. Qdrant provides a built-in exact search mode, which can be used to measure the quality of the ANN algorithm itself,
|
|
||||||
even in an automated way, as part of your CI/CD pipeline.
|
|
||||||
|
|
||||||
Again, **the quality of the embeddings is the most important factor**. HNSW does a pretty good job in terms of recall, and it is
|
HNSW covers most workloads well and is tunable when you need more recall. Other ANN algorithms exist, such as [IVF*](https://github.com/facebookresearch/faiss/wiki/Faiss-indexes#cell-probe-methods-indexivf-indexes), but they generally [perform worse than HNSW on quality and performance](https://nirantk.com/writing/pgvector-vs-qdrant/#correctness).
|
||||||
parameterizable and tunable, when required. There are some other ANN algorithms available out there, such as [IVF*](https://github.com/facebookresearch/faiss/wiki/Faiss-indexes#cell-probe-methods-indexivf-indexes),
|
|
||||||
but they usually [perform worse than HNSW in terms of quality and performance](https://nirantk.com/writing/pgvector-vs-qdrant/#correctness).
|
|
||||||
|
|||||||
Reference in New Issue
Block a user