mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-06 19:38:30 +02:00
Expand BM25 section
This commit is contained in:
@@ -404,16 +404,18 @@ Full-text search in Qdrant is powered by [sparse vectors](/articles/sparse-vecto
|
||||
### BM25
|
||||
|
||||
BM25 (Best Matching 25) is a popular ranking algorithm that takes a probabilistic approach to score calculation. For each search term, BM25 considers several statistics about the term and the document to calculate a relevance score:
|
||||
- Term frequency: the more often a term appears in a document, the more relevant that document is likely to be.
|
||||
- Inverse document frequency: the rarer a term is across all documents, the higher the weight of that term.
|
||||
- Term frequency (TF): the more often a term appears in a document, the more relevant that document is likely to be.
|
||||
- Inverse document frequency (IDF): the rarer a term is across all documents, the higher the weight of that term.
|
||||
- Document length: a term appearing in a shorter document is more relevant than the same term appearing in a longer document.
|
||||
|
||||
Qdrant offers native support for BM25 in the form of an [inference model](/documentation/concepts/inference/#server-side-inference-bm25) that generates sparse vectors, or you can generate vectors on the client side using the [FastEmbed](/documentation/fastembed/) library. The BM25 model supports the same [text processing steps](#text-processing) as text indices, including tokenization, lowercasing, ASCII folding, stemming, and stopword removal.
|
||||
Qdrant provides native support for BM25 through an [inference model](/documentation/concepts/inference/#server-side-inference-bm25) that generates sparse vectors, or you can generate vectors on the client side using the [FastEmbed](/documentation/fastembed/) library.
|
||||
|
||||
The BM25 model supports the same [text processing](#text-processing) options as text indices, including tokenization, lowercasing, ASCII folding, stemming, and stopword removal. A notable difference with text indices is that BM25 defaults to English stemming and stopword removal. If you are using a language other than English, ensure that you [configure](#language-specific-settings) the model accordingly.
|
||||
|
||||
To use BM25, configure a sparse vector when creating a collection:
|
||||
|
||||
```json
|
||||
PUT /collections/books
|
||||
PUT /collections/books?wait=true
|
||||
{
|
||||
"sparse_vectors": {
|
||||
"title-bm25": {
|
||||
@@ -425,8 +427,187 @@ PUT /collections/books
|
||||
|
||||
Note the [IDF modifier](/documentation/concepts/indexing/#idf-modifier), which configures the sparse vector for queries that use the inverse document frequency (IDF).
|
||||
|
||||
Now you can ingest data. The following example ingests a book with its title represented as a sparse vector generated by the BM25 model:
|
||||
|
||||
```json
|
||||
PUT /collections/books/points?wait=true
|
||||
{
|
||||
"points": [
|
||||
{
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"title-bm25": {
|
||||
"text": "The Time Machine",
|
||||
"model": "qdrant/bm25"
|
||||
}
|
||||
},
|
||||
"payload": {
|
||||
"title": "The Time Machine",
|
||||
"author": "H.G. Wells",
|
||||
"isbn": "9780553213515"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
After ingesting data, you can query the sparse vector. The following example searches for books with "time travel" in the title using the BM25 model:
|
||||
|
||||
```json
|
||||
POST /collections/books/points/query
|
||||
{
|
||||
"query": {
|
||||
"text": "time travel",
|
||||
"model": "qdrant/bm25"
|
||||
},
|
||||
"using": "title-bm25",
|
||||
"limit": 10,
|
||||
"with_payload": true
|
||||
}
|
||||
```
|
||||
|
||||
#### Language-specific Settings
|
||||
|
||||
By default, BM25 uses English-specific settings for tokenization, stemming, and stopword removal. Words are reduced to their English root form, and common English stopwords are removed. If your data is not in English, this leads to suboptimal search results. To achieve optimal results for other languages, configure language-specific BM25 settings.
|
||||
|
||||
<aside role="status">
|
||||
If you set any of the options discussed in this section, ensure that you apply the same settings at ingest and query time.
|
||||
</aside>
|
||||
|
||||
**Stemming and Stopwords**
|
||||
|
||||
To configure stemming and stopword removal, use the following options:
|
||||
|
||||
- `language`: sets the language for stemming and stopword removal. Defaults to `english`. To disable stemming and stopword removal, set `language` to `none`.
|
||||
- `stemmer`: defaults to the [Snowball stemmer](https://github.com/qdrant/rust-stemmers) for `language` (if set), but can be configured independently.
|
||||
- `stopwords`: defaults to a set of stopwords for `language` (if set) but can be configured independently. You can configure a specific `language` and/or configure an explicit set of stopwords that will be merged with the stopword set of the configured language.
|
||||
|
||||
For example, to use Spanish stemming and stopwords during data ingestion, use:
|
||||
|
||||
```json
|
||||
PUT /collections/books/points?wait=true
|
||||
{
|
||||
"points": [
|
||||
{
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"title-bm25": {
|
||||
"text": "La Máquina del Tiempo",
|
||||
"model": "qdrant/bm25",
|
||||
"options": {
|
||||
"language": "spanish"
|
||||
}
|
||||
}
|
||||
},
|
||||
"payload": {
|
||||
"title": "La Máquina del Tiempo",
|
||||
"author": "H.G. Wells",
|
||||
"isbn": "9788411486880"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
At query time, use the exact same parameters to ensure consistent text processing:
|
||||
|
||||
```json
|
||||
POST /collections/books/points/query
|
||||
{
|
||||
"query": {
|
||||
"text": "tiempo",
|
||||
"model": "qdrant/bm25",
|
||||
"options": {
|
||||
"language": "spanish"
|
||||
}
|
||||
},
|
||||
"using": "title-bm25",
|
||||
"limit": 10,
|
||||
"with_payload": true
|
||||
}
|
||||
```
|
||||
|
||||
To configure only a stemmer or a stopword set, rather than both, set `language` to `none` and specify the configuration for the desired stemmer or stopwords.
|
||||
|
||||
**ASCII Folding**
|
||||
|
||||
ASCII folding is the process of removing diacritics (accents) from characters. By removing diacritics, ASCII folding enables you to ignore accents when searching. For instance, with ASCII folding enabled, searching for "cafe" matches both "cafe" and "café".
|
||||
|
||||
To enable ASCII folding, set the `ascii_folding` option to `true` at both ingest and query time:
|
||||
|
||||
```json
|
||||
POST /collections/books/points/query
|
||||
{
|
||||
"query": {
|
||||
"text": "Mieville",
|
||||
"model": "qdrant/bm25",
|
||||
"options": {
|
||||
"ascii_folding": true
|
||||
}
|
||||
},
|
||||
"using": "author-bm25",
|
||||
"limit": 10,
|
||||
"with_payload": true
|
||||
}
|
||||
```
|
||||
|
||||
**Tokenizer**
|
||||
|
||||
The tokenizer breaks down text into individual tokens (words). By default, the BM25 model uses the `word` tokenizer, which splits text based on word boundaries like whitespace and punctuation. This method is effective for Latin-based languages but may not work well for languages with non-Latin alphabets or languages that do not use spaces to separate words. For those languages, use the `multilingual` tokenizer. This tokenizer supports multiple languages, including those with non-Latin alphabets and non-space delimiters.
|
||||
|
||||
```json
|
||||
POST /collections/books/points/query
|
||||
{
|
||||
"query": {
|
||||
"text": "村上春樹",
|
||||
"model": "qdrant/bm25",
|
||||
"options": {
|
||||
"tokenizer": "multilingual"
|
||||
}
|
||||
},
|
||||
"using": "author-bm25",
|
||||
"limit": 10,
|
||||
"with_payload": true
|
||||
}
|
||||
```
|
||||
|
||||
**Language-neutral Text Processing**
|
||||
|
||||
In some situations, you may want to disable language-specific processing altogether. For example, when searching for author names, that don't necessarily conform to the rules of a specific language.
|
||||
|
||||
To disable language-specific processing, set the following options:
|
||||
- `language`: set to `none` to disable language-specific stemming and stopword removal.
|
||||
- `tokenizer`: set to `multilingual` for multilingual tokenization and lemmatization.
|
||||
- Optionally, set `ascii_folding` to `true` to enable ASCII folding and ignore diacritics.
|
||||
|
||||
```json
|
||||
POST /collections/books/points/query
|
||||
{
|
||||
"query": {
|
||||
"text": "Mieville",
|
||||
"model": "qdrant/bm25",
|
||||
"options": {
|
||||
"language": "none",
|
||||
"tokenizer": "multilingual",
|
||||
"ascii_folding": true
|
||||
}
|
||||
},
|
||||
"using": "author-bm25",
|
||||
"limit": 10,
|
||||
"with_payload": true
|
||||
}
|
||||
```
|
||||
|
||||
#### Configuring BM25 Parameters
|
||||
|
||||
The BM25 [ranking function](https://en.wikipedia.org/wiki/Okapi_BM25#The_ranking_function) includes three adjustable parameters that you can set to optimize search results for your specific use case:
|
||||
|
||||
- `k`. Controls term frequency saturation. Higher values increase the influence of term frequency. Defaults to 1.2.
|
||||
- `b`. Controls document length normalization. Ranges from 0 (no normalization) to 1 (full normalization). A higher value means longer documents have less impact. Defaults to 0.75.
|
||||
- `avg_len`. Average number of words in the field being queried. Defaults to 256.
|
||||
|
||||
For instance, book titles are generally shorter than 256 words. To achieve more accurate scoring when searching for book titles, you could calculate or estimate the average title length and set the `avg_len` parameter accordingly:
|
||||
|
||||
```json
|
||||
POST /collections/books/points/query
|
||||
{
|
||||
@@ -434,7 +615,7 @@ POST /collections/books/points/query
|
||||
"text": "time travel",
|
||||
"model": "qdrant/bm25",
|
||||
"options": {
|
||||
"ascii_folding": true
|
||||
"avg_len": 5.0
|
||||
}
|
||||
},
|
||||
"using": "title-bm25",
|
||||
@@ -464,6 +645,8 @@ POST /collections/books/points/query
|
||||
}
|
||||
```
|
||||
|
||||
For a tutorial on using SPLADE++ with FastEmbed, refer to [How to Generate Sparse Vectors with SPLADE](/documentation/fastembed/fastembed-splade/).
|
||||
|
||||
### miniCOIL
|
||||
|
||||
[miniCOIL](/articles/minicoil/) strikes a balance between the flexibility of BM25 and the performance of SPLADE++. Like SPLADE++, miniCOIL is a transformer-based model that generates sparse vectors for text. However, unlike SPLADE++, miniCOIL does not use a fixed vocabulary, making it an effective model for lexical search that ranks results based on the contextual meaning of keywords.
|
||||
|
||||
Reference in New Issue
Block a user