FormulaQuery is the Python class name; TS uses formula. Switch prose,
heading, link anchor, and SEO meta to the neutral term so the docs read
correctly regardless of SDK.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Move FormulaQuery from Hybrid Search into Multi-Stage Queries; soften
weighted-RRF tuning prose; tighten the RRF-k intro.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Address findings from a code-bot adversarial review of the hybrid-search
materials:
- `hybrid-formula-decay/` (all 7 language sources + http.md +
`_description.md`): wrap the decay term in `MultExpression(mult=[0.1, ...])`
in every language. Previously the snippets summed `$score` with
`ExpDecayExpression` directly, modeling the failure mode the docs
explicitly warn against (un-weighted decay crowds out small RRF scores).
Also document the `defaults` requirement and the recommended datetime
payload index in `_description.md`. Build validated across all 6 SDKs.
- `hybrid-rrf/go.go`: add `Limit: qdrant.PtrOf(uint64(20))` to both
prefetches so the Go snippet matches the other language tabs.
- `hybrid-queries.md`:
- Reframe the weighted-RRF intro to drop "semantic search model
understands meaning better than a simple keyword matcher". On
SciFact (the corpus in the companion notebook) BM25 actually beats
dense, so the universal claim was contradicted by our own data.
- Clarify that the notebook provides a tuning helper to adapt to a
train/val split, not that it demonstrates the split itself.
- Add a one-line note that Qdrant uses zero-based rank positions so
readers can verify the RRF formula against actual scores.
- Apply brand-voice fixes: Title Case on "Multi-Stage Queries" and
"Re-Scoring Examples", replace "all the above techniques" with
"all of these techniques".
`generated/*.md` regenerated via `./docker.sh ./generate-md.py`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replaces the older `FusionQuery(fusion=Fusion.RRF)` enum form with
`RrfQuery(rrf=Rrf())` across all language tabs (Python, TypeScript,
Rust, Go, Java, C#) plus the REST body in `http.md`, for both the
`hybrid-rrf/` snippet and the inner RRF prefetch inside
`hybrid-formula-decay/`. The newer dedicated `Rrf` message is the
recommended path going forward; the old enum stays supported for
backward compatibility. Server-side both forms converge to the same
`FusionInternal::Rrf { k: 2, weights: None }`, verified against the
qdrant/qdrant source.
`generated/*.md` files in both directories regenerated via
`./docker.sh ./generate-md.py`. `./docker.sh ./check.py build` passes
across all six SDKs.
Also softens the DBSF prose in hybrid-queries.md to drop the
"weighted RRF tends to win" framing. Neither method dominates the
other in general.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds the missing keyword index on document_id so grouping works under
strict mode (Cloud default), and tunes the upload_points call to
batch_size=256, parallel=2 for faster ingestion. Mirrors the notebook
in qdrant/examples#103.
Switches the two body references to the accompanying notebook from
GitHub URLs to githubtocolab so readers can run it without cloning.
The header table's GitHub link stays for readers who want the source
view.
Tightens awkward and stale phrasing across the intro and Dataset
section (per a full review pass), repositions FormulaQuery in Wrapping
Up as an alternative rather than part of the default pipeline, drops
the redundant 'document-side equivalent' closing line, and adds a
one-line pointer to the notebook right before the Setup section so
readers can pivot to runnable code.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds dense_abstract as a fourth prefetch in retrieve() since we
ingest it already. Removes the unactionable dense_abstract hedge
paragraph, the When to Group section (the prefetch-limit gotcha
folds into a code comment), the duplicate schema-section vector
bullets, and the awkward transition sentence between the code and
the design subsections. Replaces 'summary' with 'abstract as a
whole' in descriptions for consistency with the schema rename done
earlier.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds a one-line comment above the query_points_groups call so readers
see what the function does without waiting for the When to Group
subsection (which now sits at the end of the design-decisions list).
Removes the entire 'Where This Pattern Doesn't Fit' section: the
short/homogeneous paragraph read as obvious (a reader who's deep into
this tutorial wouldn't try to multi-rep a tweet), and the
inconsistent-metadata paragraph was too vague to be actionable.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Renames sparse_keywords to sparse_title (the vector now indexes only
the title text), adds a keyword index on the tags payload field, and
gives retrieve() an optional tags parameter that builds a query_filter
when set. Pre-filtering on the tags payload is faster and more precise
than mixing categories into BM25 lexical matching.
Updates the schema description bullets, the prefetch justification,
the intro failure-mode list, and the dataset framing to match the
new design. Drops avg_len from 15 to 10 to reflect title-only word
counts.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Moves the standalone chunking-strategy sentence into the Dataset
paragraph that already explains why we chunk, and drops the dedicated
Open Ends section. The chunking-method link list (POMA-AI VST, Jina,
Chonkie) goes away with the section, and the BM25F note is also
removed since the workaround it describes is what the tutorial
already demonstrates. Cleans up two stale step-number references in
Wrapping Up.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Renames "How to Fuse" to "Which Fusion to Use" and expands it to
cover RRF (default), Weighted RRF, DBSF, and custom formulas as four
named options, with the FormulaQuery code block moved in from what
was the boosting section. Softens RRF framing from "stick with RRF
unless..." to "reasonable starting point; variants often do better
once you have an eval set." Adds the distribution-alignment
explanation in the custom-formula paragraph and links the in-repo
Decay Functions and Score Boosting references along with the RRF vs
DBSF FAQ entry.
The "When to Boost, When to Rerank" section now only covers true
boosting (recency, authority, decay) and reranking, and is reordered
to sit immediately after fusion so the ranking decisions stay
together. The "When to Group, When Not To" section moves to the end
as a presentation concern.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Calibrates BM25 length normalization for the short title+categories
sparse field with a comment on why. Removes redundant k1/b/avg_len
prose from the tutorial Open Ends section and the cross-link paragraph
in text-search.md, since the BM25 Parameters subsection above already
documents calibration with a working example.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The two lookup_from mentions were misleading: the feature is for
querying by ID across collections, not for splitting representation
storage. The line-107 paragraph now points readers to the documented
with_lookup pattern for the payload-split case and stays silent on
vector splits, which are a separate design problem (multiple queries
plus client-side fusion) that doesn't fit this tutorial's scope.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The arxiv data has abstracts, not summaries. Renaming the named
vector and prose throughout removes the ambiguity flagged on the PR.
Adds a short paragraph to the Dataset section explaining that
abstracts fit any embedding model's context window, so chunking is
included to mirror the pipeline shape you'd use on full bodies.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Drops FastEmbed in favor of server-side embedding via Cloud Inference
for dense vectors and core BM25 (in Qdrant since 1.15) for sparse.
Simplifies ingestion and query code; adds an aside covering the
self-host path. Also clears two em dashes from the tutorial prose.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replaces the FastEmbed tutorial link with the hybrid search section
of the core Text Search guide, since FastEmbed is a satellite library.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The FAQ linked to /documentation/tutorials-search-engineering/retrieval-quality/,
which only exists as a Hugo alias on the ANN recall tutorial. Aliases emit
HTML redirects but not Markdown ones, so the link checker hits the .md
output and 404s. The surrounding prose (golden query set, NDCG@10) is
relevance-tutorial territory anyway.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- hybrid-queries: see-also link at end of grouping section
- vectors: clarify MaxSim returns one combined score and point to named vectors + tutorial
- text-search: BM25 short-field calibration note plus BM25F workaround pointer
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Tutorial slugs now match their content:
- retrieval-quality/ -> ann-recall/ (alias preserves the old published path)
- retrieval-quality-golden-set/ -> retrieval-relevance/
- retrieval-quality-pipeline-output/ -> pipeline-output-quality/
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Splits the three retrieval-evaluation tutorials across two sections.
ANN Recall (Web UI, language-agnostic) stays under Search Engineering.
The two Python ecosystem tutorials (ranx, ragas) move to a new top-level
Improve Search section under the Ecosystem partition.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Web UI is renaming the tab, card title, and column header to
"ANN Recall" alongside the precision-to-recall switch. Update the
prose and screenshot alt text to match. Button label
(Check Index Quality) stays as is.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Recall@k is the standard metric for ANN benchmarks. Rename the
tutorial title, headings, helper function, and prose; update the
recent inline edits and merge sentence; and update cross-links from
the relevance and pipeline-output tutorials and the tutorials index.
Generic Precision@k mentions in the ranx metric list are unrelated
and left alone.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The Search Quality tab uses the Query Points API, so it can only
tune search-time SearchParams. Replace the m/ef_construct walkthrough
with hnsw_ef and link to the Essentials course for the build-time
trade-offs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Keep a brief Layer 4 mention in the ANN guide and remove the dedicated section from the pipeline-output guide.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds an inline gloss for SingleTurnSample where readers first encounter it,
names EvaluationResult precisely as the return type of evaluate(), and
normalizes "golden query set" to "golden set" in the intro to match the
rest of the tutorial. Ends the tutorial with a Wrapping Up section that
closes the series.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Expands MRR with Mean Reciprocal Rank and a Wikipedia link where the
metric becomes operational, matching the inline expansion treatment
already given to NDCG. Drops the IR acronym in favor of the plainer
"ranking metrics" phrasing. Strips the parenthetical subtitle from the
Pitfalls heading and normalizes the intro to use "golden sets".
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds an explicit Prerequisites line matching the pattern in Measuring
Retrieval Relevance and Evaluating Pipeline Output Quality. Renames
"Wrapping Up" to "Next Steps" for naming consistency across the series.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Aligns the layer-2 tutorial title with the parallel "Measuring X" /
"Evaluating X" pattern used by the other two and maps directly to the
four-layer framework. Slug stays the same to preserve URLs and the
golden-set artifact identity in the path. Also updates the nav descriptions
to reflect the tutorial's full scope (build + score, not just build).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The four-layer framework, metric-selection table, and business-impact
guidance now live inside the three execution tutorials. The fundamentals
page is no longer needed as a shared reference and readers don't have to
leave the tutorial flow to get context.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Folds the business-impact framework (KPI selection, pre-registered decision
rules, offline-online calibration, proxy signals) into the tutorial so it
stands alone without the external fundamentals page. The ladder pointer in
the intro now references the four-layer section inside ANN Precision.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replaces the external pointer to retrieval-quality-fundamentals with an
inline scenario-to-metric table plus guidance on picking k. The ladder
pointer now references the four-layer section inside ANN Precision rather
than the fundamentals page.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replaces the external pointer to retrieval-quality-fundamentals with a
self-contained section introducing the four evaluation layers. The tutorial
no longer depends on fundamentals for orientation.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The 2010 piece is too aged to serve as the credibility anchor
this paragraph needs, and the other prescriptive bullets in this
section don't cite external sources either. Keeping the advice
prose-only for consistency.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Back the "pre-register the decision rule" advice with a link to
Evan Miller's "How Not to Run an A/B Test," which is the canonical
reference for why post-hoc decisions and peeking wreck A/B
validity. Addresses mrscoopers' ask to support the advice with
external resources.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Soften "how teams bridge this gap" to "common patterns for
bridging this gap" so we don't imply we harvested real client
pipelines for this writeup
- Reframe the layer-2/3 diagnostic and the "Isolate the component
under test" bullet so they name multiple downstream consumer
types (LLM generator, ranker, UI) rather than assuming RAG
- Add a one-line caveat that A/B design for RAG and agentic
systems is still evolving to the Proxy KPIs paragraph
- Simplify the recall@k / precision@k equivalence note and link
ann-benchmarks.com as the citation for community convention
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add "why measure retrieval quality" lead-in paragraph
- Expand ANN on first use; flag sparse vectors as out of scope
for layer 1 (they use exact matching)
- Add LLM-as-judge to layer-2 ground-truth options
- Break the Tooling bullet into per-layer recommendations: Web UI
for L1, ranx for L2, Ragas/Phoenix/DeepEval for L3
- Reorder Quality Metrics so layer 1 (ANN recall formula + exact
kNN equivalence) comes before the generic layer-2 relevance
metrics; trim a redundant sentence
- Add end-to-end answer quality as a distinct third layer in the
prose intro so it matches the ladder table's four rows
- Standardize vocabulary on "layer" (was mixing "level" in the
intro with "layer" everywhere else); update the section anchor
to #connecting-the-layers-in-practice in this file and the two
cross-linking tutorials
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Update title, H1, and Time in retrieval-quality.md
(30 min -> 15 min reflects the pivot to Web UI)
- Rename references in tutorials-lp-overview.md and the headless
tutorial index; swap the pill from Python to Web UI to reflect
the new primary flow
- Replace "ANN recall" with "ANN precision" in Fundamentals
(4 places: intro, comparison note, ladder table, cross-link)
and in the golden-set tutorial's layer-1 cross-reference
- Filename kept as retrieval-quality.md so existing URLs and
aliases still work
The Search Quality tab in the Web UI reports precision@k, not
recall@k, so the whole tutorial now uses precision@k as the metric
label: "ANN recall" -> "ANN precision" in the anchor, section
headers, Python helper (avg_precision_at_k), and prose. Kept a
one-line bridge note that ANN-benchmarks terminology calls this
recall@k, since both searches return exactly k items.
- Drop the dataset-setup walkthrough (HF loading, collection
create, upload, wait-for-green). Readers at this phase already
have a collection.
- Replace the Python evaluation block with a "Measure ANN Recall
with the Web UI" section built around the Search Quality tab.
Default run is one-click (sample size 10); HNSW tuning uses the
tab's advanced mode instead of update_collection. Three
screenshot placeholders at
/documentation/tutorials/retrieval-quality/*.png.
- Reflect that the tab reports precision@k; note the recall@k
equivalence already spelled out in the ANN Recall section.
- Keep Python but move it to an "Automate in CI" section with a
reusable skeleton function.
- Collapse the standalone "Embeddings Quality" section into a
one-sentence MTEB pointer inside ANN Recall.
- Rewrite Wrapping Up to match the new scope.
- Link the HNSW tuning section to Optimize Performance for the
full parameter reference.
- Rewrite the intro so the ANN algorithm reads as one of several
levers shaping retrieval quality (alongside the embedding model,
retrieval strategy, filtering, reranking) rather than the only
factor beyond embeddings. Addresses mrscoopers on the reductive
"embeddings + ANN" framing.
- Rename the "Retrieval Quality" section to "ANN Recall" and
rewrite its opening paragraph to match; ANN approximation quality
isn't the same as retrieval quality broadly.
- Drop the RAG-evaluation-guide link from the three places it
appeared in this file. This tutorial isn't RAG-specific.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add a layer-1 anchor sentence under the Time/Level table
pointing at the evaluation ladder in Fundamentals, mirroring
the golden-set tutorial's anchor. Addresses abdonpijpelink's
ask for a levels-table link and a "this tutorial focuses on
level 1" framing.
- Title-case the six H2 headers for consistency across the
tutorials-search-engineering set.
Deferred: streaming the 60K training items into upload_points
instead of materializing as a list (abdonpijpelink line 69) —
pending manager confirmation before proceeding.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>