- Add "why measure retrieval quality" lead-in paragraph
- Expand ANN on first use; flag sparse vectors as out of scope
for layer 1 (they use exact matching)
- Add LLM-as-judge to layer-2 ground-truth options
- Break the Tooling bullet into per-layer recommendations: Web UI
for L1, ranx for L2, Ragas/Phoenix/DeepEval for L3
- Reorder Quality Metrics so layer 1 (ANN recall formula + exact
kNN equivalence) comes before the generic layer-2 relevance
metrics; trim a redundant sentence
- Add end-to-end answer quality as a distinct third layer in the
prose intro so it matches the ladder table's four rows
- Standardize vocabulary on "layer" (was mixing "level" in the
intro with "layer" everywhere else); update the section anchor
to #connecting-the-layers-in-practice in this file and the two
cross-linking tutorials
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Update title, H1, and Time in retrieval-quality.md
(30 min -> 15 min reflects the pivot to Web UI)
- Rename references in tutorials-lp-overview.md and the headless
tutorial index; swap the pill from Python to Web UI to reflect
the new primary flow
- Replace "ANN recall" with "ANN precision" in Fundamentals
(4 places: intro, comparison note, ladder table, cross-link)
and in the golden-set tutorial's layer-1 cross-reference
- Filename kept as retrieval-quality.md so existing URLs and
aliases still work
- Switch the metrics section from three hand-rolled functions
(recall@k, MRR, NDCG@k) and a bespoke evaluate() loop to a
single ranx-based example. Handles binary and graded labels
in one call (mrscoopers, line 59).
- Reframe the Anthropic synthetic-generation snippet as a
minimal prompt shape, point at Ragas for readers who want a
maintained testset generator, and fix a bug (Anthropic() was
called without importing the class; switched to
anthropic.Anthropic()). Addresses abdonpijpelink line 53 and
mrscoopers line 31.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Open the tutorial by anchoring it at layer 2 of the evaluation
ladder and linking to the levels table in retrieval-quality-
fundamentals (abdonpijpelink, line 14).
- Point layer-1 readers (ANN recall vs exact kNN) at the Search
Quality tab in the Qdrant Web UI instead of the retrieval-quality
tutorial (mrscoopers, line 13).
- Trim the duplicate layer-1 pointer at the end of "Using the
Golden Set" (mrscoopers, line 113).
- Open cross-tutorial links in a new tab to match the repo
convention.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Introduces two new conceptual tutorials under tutorials-search-engineering:
- Retrieval Quality Fundamentals covers the three-level evaluation
framework (ANN recall, retrieval relevance, business impact), the
evaluation ladder that connects them in practice, and a which-metric-
when decision table keyed by scenario and available ground truth.
- Building a Golden Query Set covers query generation at scale (logs,
LLM synthesis, human annotation) and the failure modes commonly
lumped together as data leakage: synthetic-query unrealism,
embedding-model contamination, near-duplicate documents, temporal
drift, and reviewer reproducibility.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>