retrieval-quality-fundamentals: drop Evan Miller citation

The 2010 piece is too aged to serve as the credibility anchor
this paragraph needs, and the other prescriptive bullets in this
section don't cite external sources either. Keeping the advice
prose-only for consistency.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Dylan Couzon
2026-04-22 22:52:20 -04:00
co-authored by Claude Opus 4.7
parent ebe01452d6
commit ecddc5c65a
@@ -39,7 +39,7 @@ Each layer is necessary but not sufficient. A win at layer 2 that doesn't carry
**Proxy KPIs for teams without A/B infrastructure.** A/B design for RAG and agentic systems is still evolving, and most teams don't have the traffic or tooling for a proper A/B regardless. Cheap production signals that correlate with business value (click position on surfaced results, answer copy or share rate, session-level task completion, thumbs-up/down) can be instrumented long before a formal experimentation platform exists, and sit usefully between the golden-set layer and the full A/B.
**Pre-register the decision rule.** Before running the A/B, write down what constitutes a win and what constitutes a no-ship, in terms of both the retrieval metric and the KPI. This is the highest-leverage discipline for avoiding "recall improved but the KPI didn't, the KPI is noisy, let's ship anyway" rationalization. See Evan Miller's [How Not to Run an A/B Test](https://www.evanmiller.org/how-not-to-run-an-ab-test.html) on why post-hoc decisions wreck A/B integrity.
**Pre-register the decision rule.** Before running the A/B, write down what constitutes a win and what constitutes a no-ship, in terms of both the retrieval metric and the KPI. This is the highest-leverage discipline for avoiding "recall improved but the KPI didn't, the KPI is noisy, let's ship anyway" rationalization.
**Tooling.** For layer 1, the Qdrant Web UI ships with a Search Quality tab that measures ANN vs exact kNN without code (see [Measuring ANN Precision](/documentation/tutorials-search-engineering/retrieval-quality/)). For layer 2, [ranx](https://amenra.github.io/ranx/) is the standard Python library for ranking metrics (recall@k, MRR, NDCG@k, and others). For layer 3, the ecosystem has mature tooling like [Ragas](https://docs.ragas.io/), [Arize Phoenix](https://phoenix.arize.com/), and [DeepEval](https://docs.confident-ai.com/) for LLM-as-judge scoring and offline answer-quality eval.