mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-30 00:18:32 +02:00
add cross-links from search and vectors docs to multi-rep tutorial
- hybrid-queries: see-also link at end of grouping section - vectors: clarify MaxSim returns one combined score and point to named vectors + tutorial - text-search: BM25 short-field calibration note plus BM25F workaround pointer Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.7
parent
120895c75d
commit
bb877b23f7
@@ -121,6 +121,8 @@ There are two scenarios where multivectors are useful:
|
||||
* **Late interaction embeddings** - Some text embedding models can output multiple vectors for a single text.
|
||||
For example, a family of models such as ColBERT output a relatively small vector for each token in the text.
|
||||
|
||||
MaxSim returns a single combined score per point, not per subvector. For per-representation control across title, summary, and chunk embeddings, see [Named Vectors](#named-vectors) and the [Multi-Representation Search tutorial](/documentation/tutorials-search-engineering/multi-representation-search/). The [multivectors course](/course/multi-vector-search/) covers limitations at scale.
|
||||
|
||||
In order to use multivectors, we need to specify a function that will be used to compare between matrices of vectors
|
||||
|
||||
Currently, Qdrant supports `max_sim` function, which is defined as a sum of maximum similarities between each pair of vectors in the matrices.
|
||||
|
||||
@@ -133,3 +133,5 @@ REST API ([Schema](https://api.qdrant.tech/master/api-reference/search/query-poi
|
||||
{{< code-snippet path="/documentation/headless/snippets/query-groups/basic/" >}}
|
||||
|
||||
For more information on the `grouping` capabilities refer to the reference documentation for search with [grouping](/documentation/search/search/#search-groups) and [lookup](/documentation/search/search/#lookup-in-groups).
|
||||
|
||||
**See also:** the [Multi-Representation Search](/documentation/tutorials-search-engineering/multi-representation-search/) tutorial for a worked end-to-end example of grouping in a hybrid retrieval pipeline.
|
||||
|
||||
@@ -208,6 +208,8 @@ For instance, book titles are generally shorter than 256 words. To achieve more
|
||||
|
||||
{{< code-snippet path="/documentation/headless/snippets/text-search/ingest-bm25-avglen/" >}}
|
||||
|
||||
When designing a multi-representation collection (combining short fields like titles and tags with longer body text), the practical default is BM25 on the shorter, structured fields with dense vectors carrying the longer ones. Default `k` and `b` values are calibrated for document-length text and may need recalibration when applied to titles or short tags. BM25F is the principled extension for multi-field text of varying length; Qdrant doesn't support it natively today, but the [Multi-Representation Search](/documentation/tutorials-search-engineering/multi-representation-search/) tutorial shows the workaround: separate sparse vectors per field, fused via the Query API.
|
||||
|
||||
#### Language-specific Settings
|
||||
|
||||
By default, BM25 uses English-specific settings for tokenization, stemming, and stopword removal. Words are reduced to their English root form, and common English stopwords are removed. If your data is not in English, this leads to suboptimal search results. To achieve optimal results for other languages, configure language-specific BM25 settings.
|
||||
|
||||
Reference in New Issue
Block a user