- Rewrite the intro so the ANN algorithm reads as one of several
levers shaping retrieval quality (alongside the embedding model,
retrieval strategy, filtering, reranking) rather than the only
factor beyond embeddings. Addresses mrscoopers on the reductive
"embeddings + ANN" framing.
- Rename the "Retrieval Quality" section to "ANN Recall" and
rewrite its opening paragraph to match; ANN approximation quality
isn't the same as retrieval quality broadly.
- Drop the RAG-evaluation-guide link from the three places it
appeared in this file. This tutorial isn't RAG-specific.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add a layer-1 anchor sentence under the Time/Level table
pointing at the evaluation ladder in Fundamentals, mirroring
the golden-set tutorial's anchor. Addresses abdonpijpelink's
ask for a levels-table link and a "this tutorial focuses on
level 1" framing.
- Title-case the six H2 headers for consistency across the
tutorials-search-engineering set.
Deferred: streaming the 60K training items into upload_points
instead of materializing as a list (abdonpijpelink line 69) —
pending manager confirmation before proceeding.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Switch the metrics section from three hand-rolled functions
(recall@k, MRR, NDCG@k) and a bespoke evaluate() loop to a
single ranx-based example. Handles binary and graded labels
in one call (mrscoopers, line 59).
- Reframe the Anthropic synthetic-generation snippet as a
minimal prompt shape, point at Ragas for readers who want a
maintained testset generator, and fix a bug (Anthropic() was
called without importing the class; switched to
anthropic.Anthropic()). Addresses abdonpijpelink line 53 and
mrscoopers line 31.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Open the tutorial by anchoring it at layer 2 of the evaluation
ladder and linking to the levels table in retrieval-quality-
fundamentals (abdonpijpelink, line 14).
- Point layer-1 readers (ANN recall vs exact kNN) at the Search
Quality tab in the Qdrant Web UI instead of the retrieval-quality
tutorial (mrscoopers, line 13).
- Trim the duplicate layer-1 pointer at the end of "Using the
Golden Set" (mrscoopers, line 113).
- Open cross-tutorial links in a new tab to match the repo
convention.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Updates both search engineering index files (the headless partial and
the tutorials-lp overview) to list Retrieval Quality Fundamentals and
Building a Golden Query Set alongside the existing Retrieval Quality
Evaluation row. The Evaluation row is also retitled from "Measure
quality and tune HNSW parameters" to "Measure ANN recall and tune
HNSW parameters" to match the refactored page.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Renames precision to recall throughout the ANN-evaluation tutorial so
the page aligns with the ANN-benchmarks convention and with the new
Retrieval Quality Fundamentals page. The numerical formula is
unchanged: when ANN and exact search both return exactly k items,
recall@k and precision@k are numerically identical.
Other changes:
- Remove the Quality metrics subsection, now covered by the
Fundamentals page, and replace it with a short link across.
- Bump weight from 4 to 6 so the three retrieval-quality pages
order as Fundamentals, Golden Query Set, Evaluation.
- Fix a pre-existing prose/code mismatch: the prose said "first
50000 items" while the code uses range(60000).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Introduces two new conceptual tutorials under tutorials-search-engineering:
- Retrieval Quality Fundamentals covers the three-level evaluation
framework (ANN recall, retrieval relevance, business impact), the
evaluation ladder that connects them in practice, and a which-metric-
when decision table keyed by scenario and available ground truth.
- Building a Golden Query Set covers query generation at scale (logs,
LLM synthesis, human annotation) and the failure modes commonly
lumped together as data leakage: synthetic-query unrealism,
embedding-model contamination, near-duplicate documents, temporal
drift, and reviewer reproducibility.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds a documentation page for Superlinked (SIE) as a Qdrant embedding
provider. The sie-qdrant package provides SIEVectorizer for dense
embeddings and SIENamedVectorizer for multi-type (dense, sparse, and
multivector/ColBERT) embeddings, enabling hybrid search via Qdrant's
Reciprocal Rank Fusion and native MaxSim retrieval via MultiVectorConfig.
Python-only.
* add thanks to sparse embeddings series
* move acknowledgements to part 5, trim copy
---------
Co-authored-by: thierrypdamiba <thierrypdamiba@users.noreply.github.com>