- Soften "how teams bridge this gap" to "common patterns for
bridging this gap" so we don't imply we harvested real client
pipelines for this writeup
- Reframe the layer-2/3 diagnostic and the "Isolate the component
under test" bullet so they name multiple downstream consumer
types (LLM generator, ranker, UI) rather than assuming RAG
- Add a one-line caveat that A/B design for RAG and agentic
systems is still evolving to the Proxy KPIs paragraph
- Simplify the recall@k / precision@k equivalence note and link
ann-benchmarks.com as the citation for community convention
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add "why measure retrieval quality" lead-in paragraph
- Expand ANN on first use; flag sparse vectors as out of scope
for layer 1 (they use exact matching)
- Add LLM-as-judge to layer-2 ground-truth options
- Break the Tooling bullet into per-layer recommendations: Web UI
for L1, ranx for L2, Ragas/Phoenix/DeepEval for L3
- Reorder Quality Metrics so layer 1 (ANN recall formula + exact
kNN equivalence) comes before the generic layer-2 relevance
metrics; trim a redundant sentence
- Add end-to-end answer quality as a distinct third layer in the
prose intro so it matches the ladder table's four rows
- Standardize vocabulary on "layer" (was mixing "level" in the
intro with "layer" everywhere else); update the section anchor
to #connecting-the-layers-in-practice in this file and the two
cross-linking tutorials
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Update title, H1, and Time in retrieval-quality.md
(30 min -> 15 min reflects the pivot to Web UI)
- Rename references in tutorials-lp-overview.md and the headless
tutorial index; swap the pill from Python to Web UI to reflect
the new primary flow
- Replace "ANN recall" with "ANN precision" in Fundamentals
(4 places: intro, comparison note, ladder table, cross-link)
and in the golden-set tutorial's layer-1 cross-reference
- Filename kept as retrieval-quality.md so existing URLs and
aliases still work
Introduces two new conceptual tutorials under tutorials-search-engineering:
- Retrieval Quality Fundamentals covers the three-level evaluation
framework (ANN recall, retrieval relevance, business impact), the
evaluation ladder that connects them in practice, and a which-metric-
when decision table keyed by scenario and available ground truth.
- Building a Golden Query Set covers query generation at scale (logs,
LLM synthesis, human annotation) and the failure modes commonly
lumped together as data leakage: synthetic-query unrealism,
embedding-model contamination, near-duplicate documents, temporal
drift, and reviewer reproducibility.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>