Commit Graph
15 Commits
Author SHA1 Message Date
Dylan Couzon 0e6e3e3c5c golden-set final update 2026-04-23 21:41:16 -04:00
Dylan Couzon c21062782b clean up golden set 2026-04-23 20:22:29 -04:00
Dylan Couzon 340c6528ca Clean up, tone 2026-04-23 11:06:18 -04:00
Dylan CouzonandClaude Opus 4.7 e2df842fb5 retrieval-quality-fundamentals: structure & scope pass
- Add "why measure retrieval quality" lead-in paragraph
- Expand ANN on first use; flag sparse vectors as out of scope
  for layer 1 (they use exact matching)
- Add LLM-as-judge to layer-2 ground-truth options
- Break the Tooling bullet into per-layer recommendations: Web UI
  for L1, ranx for L2, Ragas/Phoenix/DeepEval for L3
- Reorder Quality Metrics so layer 1 (ANN recall formula + exact
  kNN equivalence) comes before the generic layer-2 relevance
  metrics; trim a redundant sentence
- Add end-to-end answer quality as a distinct third layer in the
  prose intro so it matches the ladder table's four rows
- Standardize vocabulary on "layer" (was mixing "level" in the
  intro with "layer" everywhere else); update the section anchor
  to #connecting-the-layers-in-practice in this file and the two
  cross-linking tutorials

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 22:31:35 -04:00
Dylan Couzon 92af150b00 retrieval-quality: rename to "Measuring ANN Precision"
- Update title, H1, and Time in retrieval-quality.md
    (30 min -> 15 min reflects the pivot to Web UI)
  - Rename references in tutorials-lp-overview.md and the headless
    tutorial index; swap the pill from Python to Web UI to reflect
    the new primary flow
  - Replace "ANN recall" with "ANN precision" in Fundamentals
    (4 places: intro, comparison note, ladder table, cross-link)
    and in the golden-set tutorial's layer-1 cross-reference
  - Filename kept as retrieval-quality.md so existing URLs and
    aliases still work
2026-04-22 17:13:27 -04:00
Dylan Couzon 81883cbecc golden set: framing 2026-04-22 14:31:10 -04:00
Dylan CouzonandClaude Opus 4.7 a47c57ebdc golden set: use ranx for metrics, reframe Anthropic example
- Switch the metrics section from three hand-rolled functions
  (recall@k, MRR, NDCG@k) and a bespoke evaluate() loop to a
  single ranx-based example. Handles binary and graded labels
  in one call (mrscoopers, line 59).
- Reframe the Anthropic synthetic-generation snippet as a
  minimal prompt shape, point at Ragas for readers who want a
  maintained testset generator, and fix a bug (Anthropic() was
  called without importing the class; switched to
  anthropic.Anthropic()). Addresses abdonpijpelink line 53 and
  mrscoopers line 31.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 13:52:23 -04:00
Dylan CouzonandClaude Opus 4.7 acc85e65ca golden set: anchor intro at layer 2, link layer 1 to Web UI
- Open the tutorial by anchoring it at layer 2 of the evaluation
  ladder and linking to the levels table in retrieval-quality-
  fundamentals (abdonpijpelink, line 14).
- Point layer-1 readers (ANN recall vs exact kNN) at the Search
  Quality tab in the Qdrant Web UI instead of the retrieval-quality
  tutorial (mrscoopers, line 13).
- Trim the duplicate layer-1 pointer at the end of "Using the
  Golden Set" (mrscoopers, line 113).
- Open cross-tutorial links in a new tab to match the repo
  convention.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 10:56:34 -04:00
Dylan Couzon 6eba7e2ceb docs(golden-set): anchor intro at layer 2 2026-04-22 10:30:25 -04:00
Dylan CouzonandAbdon Pijpelink 4d99153c90 make Anthropic API key explicit
Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com>
2026-04-22 09:17:40 -04:00
Dylan Couzon 1a9821cc95 golden set: reduce verbosity 2026-04-21 12:06:48 -04:00
Dylan Couzon bb8cac5a6e update completion times 2026-04-21 12:00:10 -04:00
Dylan Couzon 8cdc9acd14 golden dataset - improve tutorial quality 2026-04-20 12:42:53 -04:00
Dylan Couzon 02710a1f9b Add using golden set instructions 2026-04-20 12:41:21 -04:00
Dylan CouzonandClaude Sonnet 4.6 5db9d106aa Add Retrieval Quality Fundamentals and Golden Query Set tutorials
Introduces two new conceptual tutorials under tutorials-search-engineering:

- Retrieval Quality Fundamentals covers the three-level evaluation
  framework (ANN recall, retrieval relevance, business impact), the
  evaluation ladder that connects them in practice, and a which-metric-
  when decision table keyed by scenario and available ground truth.
- Building a Golden Query Set covers query generation at scale (logs,
  LLM synthesis, human annotation) and the failure modes commonly
  lumped together as data leakage: synthetic-query unrealism,
  embedding-model contamination, near-duplicate documents, temporal
  drift, and reviewer reproducibility.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 10:42:05 -04:00