diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md index c55e80119..219e7fa6c 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md @@ -8,43 +8,34 @@ weight: 5 # Measuring ANN Precision -| Time: 15 min | Level: Intermediate | | | +| Time: 15 min | Level: Beginner | | | |--------------|---------------------|--|----| This tutorial focuses on **ANN precision**: how closely approximate nearest-neighbor (ANN) search matches exact kNN search. -To measure ANN precision, you compare Qdrant's approximate top-k against the exact kNN top-k, then tune HNSW parameters to control the precision/latency trade-off. +To measure ANN precision, you compare Qdrant's approximate top-k against the exact kNN top-k using `precision@k`, then tune HNSW parameters to trade memory and build time for higher precision. -To evaluate other layers of your retrieval pipeline, see the evaluation ladder. - -## ANN Precision - -Embedding quality sets the ceiling on search quality and is measured separately via benchmarks like [MTEB](https://huggingface.co/spaces/mteb/leaderboard). The retrieval pipeline can still underperform that ceiling: vector search engines such as Qdrant don't run pure kNN at query time but use **approximate nearest-neighbor** (ANN) algorithms for speed. ANN is faster than exact search but can return suboptimal results. **ANN precision** measures that gap. - -For a broader discussion of the full evaluation ladder (ANN quality, retrieval relevance, business impact), see [Retrieval Quality Fundamentals](/documentation/tutorials-search-engineering/retrieval-quality-fundamentals/). This tutorial measures the ANN layer with `precision@k`: the fraction of ANN's top-k results that appear in the exact kNN top-k. In ANN-benchmarks terminology this is equivalent to `recall@k`, since both searches return exactly `k` items. +To learn more about retrieval quality evaluation, see the evaluation ladder. ## Measure ANN Precision with the Web UI -Qdrant's Web UI has a Search Quality tab that measures the gap between approximate and exact search without requiring evaluation code. Open the dashboard at `http://localhost:6333/dashboard` (or your cluster's dashboard on Qdrant Cloud), navigate to your collection, and click the Search Quality tab. A run launches automatically with a default sample size of 10 queries, comparing ANN against exact kNN. +Qdrant's Web UI has a Search Quality tab that measures the gap between approximate and exact search without requiring evaluation code. Open the dashboard at `http://localhost:6333/dashboard` (or your cluster's dashboard on Qdrant Cloud), navigate to your collection, open the Search Quality tab, and click **Check Index Quality** to run the comparison. - ![Search Quality tab with default evaluation results](/documentation/tutorials/retrieval-quality/search-quality-tab.png) -The tab reports average **precision@k**. The score is typically high but not always perfect. When you need higher precision and can accept higher latency or more memory, HNSW is tunable. +The tab reports average **precision@k** (1.0 = perfect overlap; 0.95+ is typical for well-tuned HNSW). HNSW has tunable parameters that trade memory and index build time for higher precision. -## Tweaking the HNSW Parameters +## Tuning the HNSW Parameters -HNSW is a hierarchical graph where each node has a set of links to other nodes. The `m` parameter controls the number of edges per node: higher `m` means higher precision at the cost of more memory. The `ef_construct` parameter controls how many neighbours are considered during index building: higher `ef_construct` means higher precision at the cost of longer indexing time. Defaults are `m=16` and `ef_construct=100`. +HNSW is a hierarchical graph where each node has a set of links to other nodes. The `m` parameter controls the number of edges per node: higher `m` means higher precision at the cost of more memory. The `ef_construct` parameter controls how many neighbors are considered during index building: higher `ef_construct` means higher precision at the cost of longer indexing time. Defaults are `m=16` and `ef_construct=100`. For the full list of HNSW parameters, including on-disk storage and precision/memory trade-offs, see [Optimize Performance](/documentation/ops-optimization/optimize/). Toggle **advanced mode** in the Search Quality tab to tune these parameters inline. Raise `m` to 32 and `ef_construct` to 200, then run the evaluation again. - ![Search Quality advanced mode with HNSW parameters](/documentation/tutorials/retrieval-quality/search-quality-advanced.png) Precision should increase at the cost of higher build time and memory. - ![Search Quality results after HNSW tuning](/documentation/tutorials/retrieval-quality/search-quality-after-tuning.png) Tune until you hit the point that matches your quality and cost targets. @@ -53,7 +44,7 @@ Tune until you hit the point that matches your quality and cost targets. The Web UI is the fastest way to check precision interactively. For continuous integration or scripted regression tests, the Qdrant client exposes the same exact-search mode via `search_params=models.SearchParams(exact=True)`. Compare the ANN and exact top-k sets yourself and compute precision@k. -The helper below takes a list of query vectors and returns the average precision@k. Supply your own test set: a representative sample of query vectors from your workload, held out from training. +This helper takes a list of query vectors and returns the average precision@k. Use a representative sample of query vectors from your workload as your test set. ```python from qdrant_client import QdrantClient, models @@ -87,10 +78,10 @@ def avg_precision_at_k( return sum(precisions) / len(precisions) ``` -Drop it into your CI pipeline and fail the job if precision drops below a threshold after an embedding model change or index config update. +Wire it into CI and fail the job when precision falls below your target threshold. This catches regressions from embedding model swaps or index config changes before they reach production. ## Wrapping Up -Measuring ANN precision keeps HNSW tuning honest. The Search Quality tab gives you a quick interactive read; the Python helper above plugs into CI to catch regressions after embedding model changes or index config updates. +Measuring ANN precision keeps HNSW tuning honest. The Search Quality tab gives you a quick interactive read; the Python helper plugs into CI to catch regressions after embedding model changes or index config updates. -HNSW covers most workloads well and is tunable when you need higher precision. Other ANN algorithms exist, such as [IVF*](https://github.com/facebookresearch/faiss/wiki/Faiss-indexes#cell-probe-methods-indexivf-indexes), but they generally [perform worse than HNSW on quality and performance](https://nirantk.com/writing/pgvector-vs-qdrant/#correctness). +Once ANN precision is on target, the next layer is whether the retrieved results are relevant to users. See [Building a Golden Query Set](/documentation/tutorials-search-engineering/retrieval-quality-golden-set/). \ No newline at end of file