diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md index c13aaa69e..43a7eb52f 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-fundamentals.md @@ -39,7 +39,7 @@ Each layer is necessary but not sufficient. A win at layer 2 that doesn't carry **Proxy KPIs for teams without A/B infrastructure.** A/B design for RAG and agentic systems is still evolving, and most teams don't have the traffic or tooling for a proper A/B regardless. Cheap production signals that correlate with business value (click position on surfaced results, answer copy or share rate, session-level task completion, thumbs-up/down) can be instrumented long before a formal experimentation platform exists, and sit usefully between the golden-set layer and the full A/B. -**Pre-register the decision rule.** Before running the A/B, write down what constitutes a win and what constitutes a no-ship, in terms of both the retrieval metric and the KPI. This is the highest-leverage discipline for avoiding "recall improved but the KPI didn't, the KPI is noisy, let's ship anyway" rationalization. See Evan Miller's [How Not to Run an A/B Test](https://www.evanmiller.org/how-not-to-run-an-ab-test.html) on why post-hoc decisions wreck A/B integrity. +**Pre-register the decision rule.** Before running the A/B, write down what constitutes a win and what constitutes a no-ship, in terms of both the retrieval metric and the KPI. This is the highest-leverage discipline for avoiding "recall improved but the KPI didn't, the KPI is noisy, let's ship anyway" rationalization. **Tooling.** For layer 1, the Qdrant Web UI ships with a Search Quality tab that measures ANN vs exact kNN without code (see [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/)). For layer 2, [ranx](https://amenra.github.io/ranx/) is the standard Python library for ranking metrics (recall@k, MRR, NDCG@k, and others). For layer 3, the ecosystem has mature tooling like [Ragas](https://docs.ragas.io/), [Arize Phoenix](https://phoenix.arize.com/), and [DeepEval](https://docs.confident-ai.com/) for LLM-as-judge scoring and offline answer-quality eval.