Commit Graph
23 Commits
Author SHA1 Message Date
Dylan Couzon cea3fd6092 Semantics, improve clarity 2026-04-24 12:02:41 -04:00
Dylan Couzon 9f7c1f3fe7 stuff 2026-04-24 00:48:34 -04:00
Dylan Couzon 3e378a3647 Adding more details to the instructions 2026-04-24 00:38:27 -04:00
Dylan Couzon d14bbd3fc1 add inline comments to code snippets 2026-04-24 00:27:04 -04:00
Dylan CouzonandClaude Opus 4.7 d615b8fc30 measuring-retrieval-relevance: accessibility pass on ranking metrics
Expands MRR with Mean Reciprocal Rank and a Wikipedia link where the
metric becomes operational, matching the inline expansion treatment
already given to NDCG. Drops the IR acronym in favor of the plainer
"ranking metrics" phrasing. Strips the parenthetical subtitle from the
Pitfalls heading and normalizes the intro to use "golden sets".

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 23:40:40 -04:00
Dylan CouzonandClaude Opus 4.7 6c10996b66 rename Building a Golden Query Set to Measuring Retrieval Relevance
Aligns the layer-2 tutorial title with the parallel "Measuring X" /
"Evaluating X" pattern used by the other two and maps directly to the
four-layer framework. Slug stays the same to preserve URLs and the
golden-set artifact identity in the path. Also updates the nav descriptions
to reflect the tutorial's full scope (build + score, not just build).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 23:29:44 -04:00
Dylan CouzonandClaude Opus 4.7 1ab6ff384e golden-set: inline metric-selection table and choosing-k guidance
Replaces the external pointer to retrieval-quality-fundamentals with an
inline scenario-to-metric table plus guidance on picking k. The ladder
pointer now references the four-layer section inside ANN Precision rather
than the fundamentals page.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 23:16:37 -04:00
Dylan Couzon 5ca06869c2 Improve dataset generation for next guide. 2026-04-23 23:07:32 -04:00
Dylan Couzon 0e6e3e3c5c golden-set final update 2026-04-23 21:41:16 -04:00
Dylan Couzon c21062782b clean up golden set 2026-04-23 20:22:29 -04:00
Dylan Couzon 340c6528ca Clean up, tone 2026-04-23 11:06:18 -04:00
Dylan CouzonandClaude Opus 4.7 e2df842fb5 retrieval-quality-fundamentals: structure & scope pass
- Add "why measure retrieval quality" lead-in paragraph
- Expand ANN on first use; flag sparse vectors as out of scope
  for layer 1 (they use exact matching)
- Add LLM-as-judge to layer-2 ground-truth options
- Break the Tooling bullet into per-layer recommendations: Web UI
  for L1, ranx for L2, Ragas/Phoenix/DeepEval for L3
- Reorder Quality Metrics so layer 1 (ANN recall formula + exact
  kNN equivalence) comes before the generic layer-2 relevance
  metrics; trim a redundant sentence
- Add end-to-end answer quality as a distinct third layer in the
  prose intro so it matches the ladder table's four rows
- Standardize vocabulary on "layer" (was mixing "level" in the
  intro with "layer" everywhere else); update the section anchor
  to #connecting-the-layers-in-practice in this file and the two
  cross-linking tutorials

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 22:31:35 -04:00
Dylan Couzon 92af150b00 retrieval-quality: rename to "Measuring ANN Precision"
- Update title, H1, and Time in retrieval-quality.md
    (30 min -> 15 min reflects the pivot to Web UI)
  - Rename references in tutorials-lp-overview.md and the headless
    tutorial index; swap the pill from Python to Web UI to reflect
    the new primary flow
  - Replace "ANN recall" with "ANN precision" in Fundamentals
    (4 places: intro, comparison note, ladder table, cross-link)
    and in the golden-set tutorial's layer-1 cross-reference
  - Filename kept as retrieval-quality.md so existing URLs and
    aliases still work
2026-04-22 17:13:27 -04:00
Dylan Couzon 81883cbecc golden set: framing 2026-04-22 14:31:10 -04:00
Dylan CouzonandClaude Opus 4.7 a47c57ebdc golden set: use ranx for metrics, reframe Anthropic example
- Switch the metrics section from three hand-rolled functions
  (recall@k, MRR, NDCG@k) and a bespoke evaluate() loop to a
  single ranx-based example. Handles binary and graded labels
  in one call (mrscoopers, line 59).
- Reframe the Anthropic synthetic-generation snippet as a
  minimal prompt shape, point at Ragas for readers who want a
  maintained testset generator, and fix a bug (Anthropic() was
  called without importing the class; switched to
  anthropic.Anthropic()). Addresses abdonpijpelink line 53 and
  mrscoopers line 31.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 13:52:23 -04:00
Dylan CouzonandClaude Opus 4.7 acc85e65ca golden set: anchor intro at layer 2, link layer 1 to Web UI
- Open the tutorial by anchoring it at layer 2 of the evaluation
  ladder and linking to the levels table in retrieval-quality-
  fundamentals (abdonpijpelink, line 14).
- Point layer-1 readers (ANN recall vs exact kNN) at the Search
  Quality tab in the Qdrant Web UI instead of the retrieval-quality
  tutorial (mrscoopers, line 13).
- Trim the duplicate layer-1 pointer at the end of "Using the
  Golden Set" (mrscoopers, line 113).
- Open cross-tutorial links in a new tab to match the repo
  convention.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 10:56:34 -04:00
Dylan Couzon 6eba7e2ceb docs(golden-set): anchor intro at layer 2 2026-04-22 10:30:25 -04:00
Dylan CouzonandAbdon Pijpelink 4d99153c90 make Anthropic API key explicit
Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com>
2026-04-22 09:17:40 -04:00
Dylan Couzon 1a9821cc95 golden set: reduce verbosity 2026-04-21 12:06:48 -04:00
Dylan Couzon bb8cac5a6e update completion times 2026-04-21 12:00:10 -04:00
Dylan Couzon 8cdc9acd14 golden dataset - improve tutorial quality 2026-04-20 12:42:53 -04:00
Dylan Couzon 02710a1f9b Add using golden set instructions 2026-04-20 12:41:21 -04:00
Dylan CouzonandClaude Sonnet 4.6 5db9d106aa Add Retrieval Quality Fundamentals and Golden Query Set tutorials
Introduces two new conceptual tutorials under tutorials-search-engineering:

- Retrieval Quality Fundamentals covers the three-level evaluation
  framework (ANN recall, retrieval relevance, business impact), the
  evaluation ladder that connects them in practice, and a which-metric-
  when decision table keyed by scenario and available ground truth.
- Building a Golden Query Set covers query generation at scale (logs,
  LLM synthesis, human annotation) and the failure modes commonly
  lumped together as data leakage: synthetic-query unrealism,
  embedding-model contamination, near-duplicate documents, temporal
  drift, and reviewer reproducibility.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 10:42:05 -04:00