The Search Quality tab uses the Query Points API, so it can only
tune search-time SearchParams. Replace the m/ef_construct walkthrough
with hnsw_ef and link to the Essentials course for the build-time
trade-offs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Keep a brief Layer 4 mention in the ANN guide and remove the dedicated section from the pipeline-output guide.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds an inline gloss for SingleTurnSample where readers first encounter it,
names EvaluationResult precisely as the return type of evaluate(), and
normalizes "golden query set" to "golden set" in the intro to match the
rest of the tutorial. Ends the tutorial with a Wrapping Up section that
closes the series.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Expands MRR with Mean Reciprocal Rank and a Wikipedia link where the
metric becomes operational, matching the inline expansion treatment
already given to NDCG. Drops the IR acronym in favor of the plainer
"ranking metrics" phrasing. Strips the parenthetical subtitle from the
Pitfalls heading and normalizes the intro to use "golden sets".
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds an explicit Prerequisites line matching the pattern in Measuring
Retrieval Relevance and Evaluating Pipeline Output Quality. Renames
"Wrapping Up" to "Next Steps" for naming consistency across the series.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Aligns the layer-2 tutorial title with the parallel "Measuring X" /
"Evaluating X" pattern used by the other two and maps directly to the
four-layer framework. Slug stays the same to preserve URLs and the
golden-set artifact identity in the path. Also updates the nav descriptions
to reflect the tutorial's full scope (build + score, not just build).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The four-layer framework, metric-selection table, and business-impact
guidance now live inside the three execution tutorials. The fundamentals
page is no longer needed as a shared reference and readers don't have to
leave the tutorial flow to get context.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Folds the business-impact framework (KPI selection, pre-registered decision
rules, offline-online calibration, proxy signals) into the tutorial so it
stands alone without the external fundamentals page. The ladder pointer in
the intro now references the four-layer section inside ANN Precision.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replaces the external pointer to retrieval-quality-fundamentals with an
inline scenario-to-metric table plus guidance on picking k. The ladder
pointer now references the four-layer section inside ANN Precision rather
than the fundamentals page.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replaces the external pointer to retrieval-quality-fundamentals with a
self-contained section introducing the four evaluation layers. The tutorial
no longer depends on fundamentals for orientation.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Soften "how teams bridge this gap" to "common patterns for
bridging this gap" so we don't imply we harvested real client
pipelines for this writeup
- Reframe the layer-2/3 diagnostic and the "Isolate the component
under test" bullet so they name multiple downstream consumer
types (LLM generator, ranker, UI) rather than assuming RAG
- Add a one-line caveat that A/B design for RAG and agentic
systems is still evolving to the Proxy KPIs paragraph
- Simplify the recall@k / precision@k equivalence note and link
ann-benchmarks.com as the citation for community convention
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add "why measure retrieval quality" lead-in paragraph
- Expand ANN on first use; flag sparse vectors as out of scope
for layer 1 (they use exact matching)
- Add LLM-as-judge to layer-2 ground-truth options
- Break the Tooling bullet into per-layer recommendations: Web UI
for L1, ranx for L2, Ragas/Phoenix/DeepEval for L3
- Reorder Quality Metrics so layer 1 (ANN recall formula + exact
kNN equivalence) comes before the generic layer-2 relevance
metrics; trim a redundant sentence
- Add end-to-end answer quality as a distinct third layer in the
prose intro so it matches the ladder table's four rows
- Standardize vocabulary on "layer" (was mixing "level" in the
intro with "layer" everywhere else); update the section anchor
to #connecting-the-layers-in-practice in this file and the two
cross-linking tutorials
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Update title, H1, and Time in retrieval-quality.md
(30 min -> 15 min reflects the pivot to Web UI)
- Rename references in tutorials-lp-overview.md and the headless
tutorial index; swap the pill from Python to Web UI to reflect
the new primary flow
- Replace "ANN recall" with "ANN precision" in Fundamentals
(4 places: intro, comparison note, ladder table, cross-link)
and in the golden-set tutorial's layer-1 cross-reference
- Filename kept as retrieval-quality.md so existing URLs and
aliases still work
The Search Quality tab in the Web UI reports precision@k, not
recall@k, so the whole tutorial now uses precision@k as the metric
label: "ANN recall" -> "ANN precision" in the anchor, section
headers, Python helper (avg_precision_at_k), and prose. Kept a
one-line bridge note that ANN-benchmarks terminology calls this
recall@k, since both searches return exactly k items.
- Drop the dataset-setup walkthrough (HF loading, collection
create, upload, wait-for-green). Readers at this phase already
have a collection.
- Replace the Python evaluation block with a "Measure ANN Recall
with the Web UI" section built around the Search Quality tab.
Default run is one-click (sample size 10); HNSW tuning uses the
tab's advanced mode instead of update_collection. Three
screenshot placeholders at
/documentation/tutorials/retrieval-quality/*.png.
- Reflect that the tab reports precision@k; note the recall@k
equivalence already spelled out in the ANN Recall section.
- Keep Python but move it to an "Automate in CI" section with a
reusable skeleton function.
- Collapse the standalone "Embeddings Quality" section into a
one-sentence MTEB pointer inside ANN Recall.
- Rewrite Wrapping Up to match the new scope.
- Link the HNSW tuning section to Optimize Performance for the
full parameter reference.
- Rewrite the intro so the ANN algorithm reads as one of several
levers shaping retrieval quality (alongside the embedding model,
retrieval strategy, filtering, reranking) rather than the only
factor beyond embeddings. Addresses mrscoopers on the reductive
"embeddings + ANN" framing.
- Rename the "Retrieval Quality" section to "ANN Recall" and
rewrite its opening paragraph to match; ANN approximation quality
isn't the same as retrieval quality broadly.
- Drop the RAG-evaluation-guide link from the three places it
appeared in this file. This tutorial isn't RAG-specific.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add a layer-1 anchor sentence under the Time/Level table
pointing at the evaluation ladder in Fundamentals, mirroring
the golden-set tutorial's anchor. Addresses abdonpijpelink's
ask for a levels-table link and a "this tutorial focuses on
level 1" framing.
- Title-case the six H2 headers for consistency across the
tutorials-search-engineering set.
Deferred: streaming the 60K training items into upload_points
instead of materializing as a list (abdonpijpelink line 69) —
pending manager confirmation before proceeding.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Switch the metrics section from three hand-rolled functions
(recall@k, MRR, NDCG@k) and a bespoke evaluate() loop to a
single ranx-based example. Handles binary and graded labels
in one call (mrscoopers, line 59).
- Reframe the Anthropic synthetic-generation snippet as a
minimal prompt shape, point at Ragas for readers who want a
maintained testset generator, and fix a bug (Anthropic() was
called without importing the class; switched to
anthropic.Anthropic()). Addresses abdonpijpelink line 53 and
mrscoopers line 31.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Open the tutorial by anchoring it at layer 2 of the evaluation
ladder and linking to the levels table in retrieval-quality-
fundamentals (abdonpijpelink, line 14).
- Point layer-1 readers (ANN recall vs exact kNN) at the Search
Quality tab in the Qdrant Web UI instead of the retrieval-quality
tutorial (mrscoopers, line 13).
- Trim the duplicate layer-1 pointer at the end of "Using the
Golden Set" (mrscoopers, line 113).
- Open cross-tutorial links in a new tab to match the repo
convention.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Clicking a language tab now switches all other code-snippet widgets on
the page to the same language, and the choice is saved to localStorage
so it is auto-applied on future page loads.
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>