Commit Graph
6296 Commits
Author SHA1 Message Date
Dylan CouzonandClaude Opus 4.7 3de7b1ddfb tighten intro, dataset, and wrap-up; add notebook pointer before setup
Tightens awkward and stale phrasing across the intro and Dataset
section (per a full review pass), repositions FormulaQuery in Wrapping
Up as an alternative rather than part of the default pipeline, drops
the redundant 'document-side equivalent' closing line, and adds a
one-line pointer to the notebook right before the Setup section so
readers can pivot to runnable code.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 20:23:42 -04:00
Dylan CouzonandClaude Opus 4.7 b30140820f tighten retrieval section and remove vector-description duplication
Adds dense_abstract as a fourth prefetch in retrieve() since we
ingest it already. Removes the unactionable dense_abstract hedge
paragraph, the When to Group section (the prefetch-limit gotcha
folds into a code comment), the duplicate schema-section vector
bullets, and the awkward transition sentence between the code and
the design subsections. Replaces 'summary' with 'abstract as a
whole' in descriptions for consistency with the schema rename done
earlier.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 19:53:39 -04:00
Dylan CouzonandClaude Opus 4.7 479d305aca explain query_points_groups inline and remove pattern-doesnt-fit section
Adds a one-line comment above the query_points_groups call so readers
see what the function does without waiting for the When to Group
subsection (which now sits at the end of the design-decisions list).

Removes the entire 'Where This Pattern Doesn't Fit' section: the
short/homogeneous paragraph read as obvious (a reader who's deep into
this tutorial wouldn't try to multi-rep a tweet), and the
inconsistent-metadata paragraph was too vague to be actionable.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 19:08:37 -04:00
Dylan CouzonandClaude Opus 4.7 8ea4653ee8 move categories from BM25 input to filterable payload
Renames sparse_keywords to sparse_title (the vector now indexes only
the title text), adds a keyword index on the tags payload field, and
gives retrieve() an optional tags parameter that builds a query_filter
when set. Pre-filtering on the tags payload is faster and more precise
than mixing categories into BM25 lexical matching.

Updates the schema description bullets, the prefetch justification,
the intro failure-mode list, and the dataset framing to match the
new design. Drops avg_len from 15 to 10 to reflect title-only word
counts.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 18:44:37 -04:00
Dylan CouzonandClaude Opus 4.7 8abae12a8c consolidate chunking note and remove open ends section
Moves the standalone chunking-strategy sentence into the Dataset
paragraph that already explains why we chunk, and drops the dedicated
Open Ends section. The chunking-method link list (POMA-AI VST, Jina,
Chonkie) goes away with the section, and the BM25F note is also
removed since the workaround it describes is what the tutorial
already demonstrates. Cleans up two stale step-number references in
Wrapping Up.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 18:18:51 -04:00
Dylan CouzonandClaude Opus 4.7 a9610d2a30 restructure fusion section to cover four named strategies
Renames "How to Fuse" to "Which Fusion to Use" and expands it to
cover RRF (default), Weighted RRF, DBSF, and custom formulas as four
named options, with the FormulaQuery code block moved in from what
was the boosting section. Softens RRF framing from "stick with RRF
unless..." to "reasonable starting point; variants often do better
once you have an eval set." Adds the distribution-alignment
explanation in the custom-formula paragraph and links the in-repo
Decay Functions and Score Boosting references along with the RRF vs
DBSF FAQ entry.

The "When to Boost, When to Rerank" section now only covers true
boosting (recency, authority, decay) and reranking, and is reordered
to sit immediately after fusion so the ranking decisions stay
together. The "When to Group, When Not To" section moves to the end
as a presentation concern.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 17:33:32 -04:00
daniel-azoulai 1f2197b7c0 Update case-study-sapu.md 2026-05-11 13:57:08 -07:00
daniel-azoulai 484061ff80 Update case-study-sapu.md 2026-05-11 13:49:12 -07:00
Ivan Pleshkov d91e94a365 review remark 2026-05-11 22:29:43 +02:00
Dylan CouzonandClaude Opus 4.7 68cec300f2 set avg_len for BM25 ingestion and trim parameter notes
Calibrates BM25 length normalization for the short title+categories
sparse field with a comment on why. Removes redundant k1/b/avg_len
prose from the tutorial Open Ends section and the cross-link paragraph
in text-search.md, since the BM25 Parameters subsection above already
documents calibration with a working example.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 16:06:03 -04:00
Dylan CouzonandClaude Opus 4.7 376202a113 remove lookup_from references, point to with_lookup for payload splits
The two lookup_from mentions were misleading: the feature is for
querying by ID across collections, not for splitting representation
storage. The line-107 paragraph now points readers to the documented
with_lookup pattern for the payload-split case and stays silent on
vector splits, which are a separate design problem (multiple queries
plus client-side fusion) that doesn't fit this tutorial's scope.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 15:24:08 -04:00
Dylan CouzonandClaude Opus 4.7 433cae62a5 rename dense_summary to dense_abstract and explain chunking choice
The arxiv data has abstracts, not summaries. Renaming the named
vector and prose throughout removes the ambiguity flagged on the PR.
Adds a short paragraph to the Dataset section explaining that
abstracts fit any embedding model's context window, so chunking is
included to mirror the pipeline shape you'd use on full bodies.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 15:03:52 -04:00
Dylan CouzonandClaude Opus 4.7 972568940d switch tutorial to qdrant cloud inference and core bm25
Drops FastEmbed in favor of server-side embedding via Cloud Inference
for dense vectors and core BM25 (in Qdrant since 1.15) for sparse.
Simplifies ingestion and query code; adds an aside covering the
self-host path. Also clears two em dashes from the tutorial prose.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 14:48:21 -04:00
Dylan CouzonandClaude Opus 4.7 300397b467 point hybrid search prerequisite at text search guide
Replaces the FastEmbed tutorial link with the hybrid search section
of the core Text Search guide, since FastEmbed is a satellite library.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 14:32:47 -04:00
m-qdrant aca47a9ee7 Merge pull request #2346 from qdrant/maddie-qdrant-patch-6
Update top-banner.md
2026-05-11 13:39:58 -04:00
m-qdrant 86fbed28a0 Update top-banner.md
http://qdrant.tech/blog/qdrant-1.18.x/
2026-05-11 13:32:52 -04:00
Dylan Couzon 6e073b853b Merge pull request #2344 from qdrant/fix-faq-retrieval-quality-link
Fix broken FAQ link to retrieval-quality tutorial
2026-05-11 12:40:47 -04:00
Ivan Pleshkov 1cf25ad0d9 dataset links 2026-05-11 18:18:59 +02:00
Dylan CouzonandClaude Opus 4.7 a5e4ef0532 Fix FAQ link to point at the relevance tutorial directly
The FAQ linked to /documentation/tutorials-search-engineering/retrieval-quality/,
which only exists as a Hugo alias on the ANN recall tutorial. Aliases emit
HTML redirects but not Markdown ones, so the link checker hits the .md
output and 404s. The surrounding prose (golden query set, NDCG@10) is
relevance-tutorial territory anyway.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 12:08:02 -04:00
Dylan Couzon d21b5a5435 Merge pull request #2305 from qdrant/retrival-quality-guide-improvement
docs(tutorials): split retrieval-quality into three evaluation layers
2026-05-11 11:55:53 -04:00
Ivan Pleshkov b2cc944a3e images and links 2026-05-11 17:42:30 +02:00
Ivan Pleshkov c95d156197 images 2026-05-11 17:21:06 +02:00
Tim Visée c37393a17e Merge pull request #2298 from qdrant/version-1.18
Publish docs for version 1.18
2026-05-11 17:02:50 +02:00
jojii e1090927b0 Wording + Grammar 2026-05-11 16:54:10 +02:00
Ivan Pleshkov a73cac3336 review remarks 2026-05-11 16:27:04 +02:00
Abdon Pijpelink 7fd26f0e63 Test code snippets against 1.18 release 2026-05-11 16:14:29 +02:00
jojii 42443a6e7e SIMD + other review 2026-05-11 16:05:35 +02:00
Ivan Pleshkov eb7e866208 better calibration and simd parts 2026-05-11 15:14:25 +02:00
Ivan Pleshkov bb835f71ec simplify simd 2026-05-11 14:46:42 +02:00
Abdon Pijpelink 4e319c65bc Update release blog date 2026-05-11 13:32:11 +02:00
Abdon Pijpelink 728ea9caa7 Test python against dev branch 2026-05-11 13:04:36 +02:00
Ivan Pleshkov 5106a76bbd review remarks 2026-05-11 11:58:47 +02:00
Abdon Pijpelink cba308dc31 Fix tracing ID Python snippet import 2026-05-11 11:32:10 +02:00
jojii 144d288e2d Grammar and wording improvements 2026-05-11 10:31:05 +02:00
Mohamed Arbi c3a3c1b0fd fix typos in hybrid-cloud and private cloud docs (#2343) 2026-05-11 08:46:00 +02:00
Abdon Pijpelink dca1d290f5 Update release blog date 2026-05-11 08:22:41 +02:00
46e81449a8 Add TurboQuant quantization documentation (v1.18.0) (#2313)
* Add TurboQuant quantization documentation (v1.18.0)

Adds a new TurboQuant section to the quantization guide covering the
four encoding options (bits1/bits1_5/bits2/bits4), automatic asymmetric
quantization, distance metric support, and the automatic TQ+ precision
enhancement for sealed segments. Updates the comparison table and
method-selection guidance to recommend TurboQuant over Binary and
Scalar Quantization for new collections. Adds HTTP-only snippets for
basic setup and explicit bit-depth selection.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* Consistently use title case for headers

* Restructure doc to lead with TurboQuant

* Clarify rescoring

* Edits

* Soften TQ advice

* Updates

* Fix

* Apply suggestions from code review

Co-authored-by: Jojii <15957865+JojiiOfficial@users.noreply.github.com>

* Review feedback

* Add Go/Java/C# code snippets

* Add list of 4 quantization methods to introduction

* Review feedback

* Stronger advice for TQ4

* Add Python snippets

* Add Rust snippets

* Add TS snippets

* Update recommendation table

* Update production checklist

* Remove link to article

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Jojii <15957865+JojiiOfficial@users.noreply.github.com>
2026-05-11 08:13:03 +02:00
Ivan Pleshkov 73d63fb5ca preview 2026-05-11 00:39:28 +02:00
Ivan Pleshkov 356489d51d review remarks 2026-05-11 00:33:50 +02:00
generall c830a15c40 Update GitHub Stars 2026-05-10 17:17:34 +00:00
daniel-azoulai f3533c9662 Update 2026-05-09 10:56:44 -07:00
daniel-azoulai fb28e57495 Update case-study-sapu.md 2026-05-08 14:14:23 -07:00
daniel-azoulai cad42938d5 Update case-study-sapu.md 2026-05-08 14:11:00 -07:00
daniel-azoulai 94dd626fef Update case-study-sapu.md 2026-05-08 14:10:33 -07:00
daniel-azoulai 84a73f1759 Update 2026-05-08 14:07:29 -07:00
daniel-azoulai 5d31169b1b sapu 2026-05-08 14:03:02 -07:00
daniel-azoulai 3c6def3336 case-study-sapu 2026-05-08 13:46:23 -07:00
Tim ViséeandJojii 2746db8762 Apply suggestions from code review
Co-authored-by: Jojii <15957865+JojiiOfficial@users.noreply.github.com>
2026-05-08 17:03:33 +02:00
Ivan Pleshkov 6b16f68996 review remarks 2026-05-08 12:04:17 +02:00
Abdon Pijpelink 7048d0cd54 Remove links to article 2026-05-08 08:25:09 +02:00