diff --git a/qdrant-landing/content/documentation/manage-data/vectors.md b/qdrant-landing/content/documentation/manage-data/vectors.md index faf761f03..c9425786a 100644 --- a/qdrant-landing/content/documentation/manage-data/vectors.md +++ b/qdrant-landing/content/documentation/manage-data/vectors.md @@ -121,6 +121,8 @@ There are two scenarios where multivectors are useful: * **Late interaction embeddings** - Some text embedding models can output multiple vectors for a single text. For example, a family of models such as ColBERT output a relatively small vector for each token in the text. +MaxSim returns a single combined score per point, not per subvector. For per-representation control across title, summary, and chunk embeddings, see [Named Vectors](#named-vectors) and the [Multi-Representation Search tutorial](/documentation/tutorials-search-engineering/multi-representation-search/). The [multivectors course](/course/multi-vector-search/) covers limitations at scale. + In order to use multivectors, we need to specify a function that will be used to compare between matrices of vectors Currently, Qdrant supports `max_sim` function, which is defined as a sum of maximum similarities between each pair of vectors in the matrices. diff --git a/qdrant-landing/content/documentation/search/hybrid-queries.md b/qdrant-landing/content/documentation/search/hybrid-queries.md index c68ebd7db..e4806a039 100644 --- a/qdrant-landing/content/documentation/search/hybrid-queries.md +++ b/qdrant-landing/content/documentation/search/hybrid-queries.md @@ -133,3 +133,5 @@ REST API ([Schema](https://api.qdrant.tech/master/api-reference/search/query-poi {{< code-snippet path="/documentation/headless/snippets/query-groups/basic/" >}} For more information on the `grouping` capabilities refer to the reference documentation for search with [grouping](/documentation/search/search/#search-groups) and [lookup](/documentation/search/search/#lookup-in-groups). + +**See also:** the [Multi-Representation Search](/documentation/tutorials-search-engineering/multi-representation-search/) tutorial for a worked end-to-end example of grouping in a hybrid retrieval pipeline. diff --git a/qdrant-landing/content/documentation/search/text-search.md b/qdrant-landing/content/documentation/search/text-search.md index 166fdbb19..694dd1fe7 100644 --- a/qdrant-landing/content/documentation/search/text-search.md +++ b/qdrant-landing/content/documentation/search/text-search.md @@ -208,6 +208,8 @@ For instance, book titles are generally shorter than 256 words. To achieve more {{< code-snippet path="/documentation/headless/snippets/text-search/ingest-bm25-avglen/" >}} +When designing a multi-representation collection (combining short fields like titles and tags with longer body text), the practical default is BM25 on the shorter, structured fields with dense vectors carrying the longer ones. Default `k` and `b` values are calibrated for document-length text and may need recalibration when applied to titles or short tags. BM25F is the principled extension for multi-field text of varying length; Qdrant doesn't support it natively today, but the [Multi-Representation Search](/documentation/tutorials-search-engineering/multi-representation-search/) tutorial shows the workaround: separate sparse vectors per field, fused via the Query API. + #### Language-specific Settings By default, BM25 uses English-specific settings for tokenization, stemming, and stopword removal. Words are reduced to their English root form, and common English stopwords are removed. If your data is not in English, this leads to suboptimal search results. To achieve optimal results for other languages, configure language-specific BM25 settings.