add cross-links from search and vectors docs to multi-rep tutorial

- hybrid-queries: see-also link at end of grouping section
- vectors: clarify MaxSim returns one combined score and point to named vectors + tutorial
- text-search: BM25 short-field calibration note plus BM25F workaround pointer

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Dylan Couzon
2026-05-05 19:10:24 -04:00
co-authored by Claude Opus 4.7
parent 120895c75d
commit bb877b23f7
3 changed files with 6 additions and 0 deletions
@@ -121,6 +121,8 @@ There are two scenarios where multivectors are useful:
* **Late interaction embeddings** - Some text embedding models can output multiple vectors for a single text.
For example, a family of models such as ColBERT output a relatively small vector for each token in the text.
MaxSim returns a single combined score per point, not per subvector. For per-representation control across title, summary, and chunk embeddings, see [Named Vectors](#named-vectors) and the [Multi-Representation Search tutorial](/documentation/tutorials-search-engineering/multi-representation-search/). The [multivectors course](/course/multi-vector-search/) covers limitations at scale.
In order to use multivectors, we need to specify a function that will be used to compare between matrices of vectors
Currently, Qdrant supports `max_sim` function, which is defined as a sum of maximum similarities between each pair of vectors in the matrices.
@@ -133,3 +133,5 @@ REST API ([Schema](https://api.qdrant.tech/master/api-reference/search/query-poi
{{< code-snippet path="/documentation/headless/snippets/query-groups/basic/" >}}
For more information on the `grouping` capabilities refer to the reference documentation for search with [grouping](/documentation/search/search/#search-groups) and [lookup](/documentation/search/search/#lookup-in-groups).
**See also:** the [Multi-Representation Search](/documentation/tutorials-search-engineering/multi-representation-search/) tutorial for a worked end-to-end example of grouping in a hybrid retrieval pipeline.
@@ -208,6 +208,8 @@ For instance, book titles are generally shorter than 256 words. To achieve more
{{< code-snippet path="/documentation/headless/snippets/text-search/ingest-bm25-avglen/" >}}
When designing a multi-representation collection (combining short fields like titles and tags with longer body text), the practical default is BM25 on the shorter, structured fields with dense vectors carrying the longer ones. Default `k` and `b` values are calibrated for document-length text and may need recalibration when applied to titles or short tags. BM25F is the principled extension for multi-field text of varying length; Qdrant doesn't support it natively today, but the [Multi-Representation Search](/documentation/tutorials-search-engineering/multi-representation-search/) tutorial shows the workaround: separate sparse vectors per field, fused via the Query API.
#### Language-specific Settings
By default, BM25 uses English-specific settings for tokenization, stemming, and stopword removal. Words are reduced to their English root form, and common English stopwords are removed. If your data is not in English, this leads to suboptimal search results. To achieve optimal results for other languages, configure language-specific BM25 settings.