* fix(docs): correct C# MatchExcept snippet to use MatchExcept instead of Match
The C# code example in the match-except filter condition snippet was
incorrectly using Match() instead of MatchExcept(), making it identical
to the match-any example. This is misleading for developers trying to
implement the MatchExcept (NOT IN) filter condition in C#.
Fixes#1656
* Add Markdownified code snippet
---------
Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com>
Adds the missing keyword index on document_id so grouping works under
strict mode (Cloud default), and tunes the upload_points call to
batch_size=256, parallel=2 for faster ingestion. Mirrors the notebook
in qdrant/examples#103.
Switches the two body references to the accompanying notebook from
GitHub URLs to githubtocolab so readers can run it without cloning.
The header table's GitHub link stays for readers who want the source
view.
Tightens awkward and stale phrasing across the intro and Dataset
section (per a full review pass), repositions FormulaQuery in Wrapping
Up as an alternative rather than part of the default pipeline, drops
the redundant 'document-side equivalent' closing line, and adds a
one-line pointer to the notebook right before the Setup section so
readers can pivot to runnable code.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds dense_abstract as a fourth prefetch in retrieve() since we
ingest it already. Removes the unactionable dense_abstract hedge
paragraph, the When to Group section (the prefetch-limit gotcha
folds into a code comment), the duplicate schema-section vector
bullets, and the awkward transition sentence between the code and
the design subsections. Replaces 'summary' with 'abstract as a
whole' in descriptions for consistency with the schema rename done
earlier.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds a one-line comment above the query_points_groups call so readers
see what the function does without waiting for the When to Group
subsection (which now sits at the end of the design-decisions list).
Removes the entire 'Where This Pattern Doesn't Fit' section: the
short/homogeneous paragraph read as obvious (a reader who's deep into
this tutorial wouldn't try to multi-rep a tweet), and the
inconsistent-metadata paragraph was too vague to be actionable.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Renames sparse_keywords to sparse_title (the vector now indexes only
the title text), adds a keyword index on the tags payload field, and
gives retrieve() an optional tags parameter that builds a query_filter
when set. Pre-filtering on the tags payload is faster and more precise
than mixing categories into BM25 lexical matching.
Updates the schema description bullets, the prefetch justification,
the intro failure-mode list, and the dataset framing to match the
new design. Drops avg_len from 15 to 10 to reflect title-only word
counts.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Moves the standalone chunking-strategy sentence into the Dataset
paragraph that already explains why we chunk, and drops the dedicated
Open Ends section. The chunking-method link list (POMA-AI VST, Jina,
Chonkie) goes away with the section, and the BM25F note is also
removed since the workaround it describes is what the tutorial
already demonstrates. Cleans up two stale step-number references in
Wrapping Up.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Renames "How to Fuse" to "Which Fusion to Use" and expands it to
cover RRF (default), Weighted RRF, DBSF, and custom formulas as four
named options, with the FormulaQuery code block moved in from what
was the boosting section. Softens RRF framing from "stick with RRF
unless..." to "reasonable starting point; variants often do better
once you have an eval set." Adds the distribution-alignment
explanation in the custom-formula paragraph and links the in-repo
Decay Functions and Score Boosting references along with the RRF vs
DBSF FAQ entry.
The "When to Boost, When to Rerank" section now only covers true
boosting (recency, authority, decay) and reranking, and is reordered
to sit immediately after fusion so the ranking decisions stay
together. The "When to Group, When Not To" section moves to the end
as a presentation concern.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Calibrates BM25 length normalization for the short title+categories
sparse field with a comment on why. Removes redundant k1/b/avg_len
prose from the tutorial Open Ends section and the cross-link paragraph
in text-search.md, since the BM25 Parameters subsection above already
documents calibration with a working example.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The two lookup_from mentions were misleading: the feature is for
querying by ID across collections, not for splitting representation
storage. The line-107 paragraph now points readers to the documented
with_lookup pattern for the payload-split case and stays silent on
vector splits, which are a separate design problem (multiple queries
plus client-side fusion) that doesn't fit this tutorial's scope.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>