diff --git a/qdrant-landing/content/documentation/search/text-search.md b/qdrant-landing/content/documentation/search/text-search.md index 694dd1fe7..4207bfd14 100644 --- a/qdrant-landing/content/documentation/search/text-search.md +++ b/qdrant-landing/content/documentation/search/text-search.md @@ -208,7 +208,7 @@ For instance, book titles are generally shorter than 256 words. To achieve more {{< code-snippet path="/documentation/headless/snippets/text-search/ingest-bm25-avglen/" >}} -When designing a multi-representation collection (combining short fields like titles and tags with longer body text), the practical default is BM25 on the shorter, structured fields with dense vectors carrying the longer ones. Default `k` and `b` values are calibrated for document-length text and may need recalibration when applied to titles or short tags. BM25F is the principled extension for multi-field text of varying length; Qdrant doesn't support it natively today, but the [Multi-Representation Search](/documentation/tutorials-search-engineering/multi-representation-search/) tutorial shows the workaround: separate sparse vectors per field, fused via the Query API. +When designing a multi-representation collection (combining short fields like titles and tags with longer body text), the practical default is BM25 on the shorter, structured fields with dense vectors carrying the longer ones. BM25F is the principled extension for multi-field text of varying length; Qdrant doesn't support it natively today, but the [Multi-Representation Search](/documentation/tutorials-search-engineering/multi-representation-search/) tutorial shows the workaround: separate sparse vectors per field, fused via the Query API. #### Language-specific Settings diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/multi-representation-search.md b/qdrant-landing/content/documentation/tutorials-search-engineering/multi-representation-search.md index 895e583ea..6b19ef484 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/multi-representation-search.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/multi-representation-search.md @@ -126,9 +126,12 @@ for paper in papers: # Cloud Inference embeds each Document on the server, so you don't need a client-side embedding library. title_doc = models.Document(text=paper["title"], model=DENSE_MODEL) abstract_doc = models.Document(text=paper["abstract"], model=DENSE_MODEL) + # avg_len is the average word count of the indexed text. + # Default is 256 (document-length); setting it to the actual field length (~15 here) improves BM25 scoring accuracy. sparse_doc = models.Document( text=paper["title"] + " " + " ".join(paper["categories"]), model=BM25_MODEL, + options={"avg_len": 15.0}, ) for i, chunk in enumerate(chunks): @@ -246,7 +249,7 @@ The pattern also assumes you can identify the representations cleanly. If your c **Chunking strategy.** This tutorial uses a fixed-length sentence chunker for clarity. The chunking choice has a measurable effect on retrieval quality, and the right strategy depends on document structure. Worth considering for your corpus: hierarchical chunking ([POMA-AI VST](https://arxiv.org/abs/2406.04590)), late chunking ([Jina](https://jina.ai/news/late-chunking-in-long-context-embedding-models/)), and semantic chunking ([Chonkie](https://github.com/chonkie-ai/chonkie)). -**BM25F.** The technically correct extension of BM25 to multi-field text of varying length is BM25F, which weights term statistics per field. Qdrant doesn't support BM25F natively today. The workaround used in step 3, separate sparse vectors per field fused via the Query API, gives you most of the practical benefit. Default BM25 parameters (k1, b) are calibrated for document-length text and may need recalibration for short fields like titles or tags; treat them as tunable hyperparameters when working with mixed-length sparse representations. +**BM25F.** The technically correct extension of BM25 to multi-field text of varying length is BM25F, which weights term statistics per field. Qdrant doesn't support BM25F natively today. The workaround used in step 3 (separate sparse vectors per field fused via the Query API) gives you most of the practical benefit. ## Wrapping Up