Release 1 10 aggregated (#1005)
* Qdrant 1.10: describe IDF computation (#996) * describe IDF computation * add python snippet, fix typo * docs: Java, Csharp * add ts and rust examples * docs: Formatting Java, Csharp indexing.md * fix language --------- Co-authored-by: Luis Cossío <luis.cossio@outlook.com> Co-authored-by: Anush <anushshetty90@gmail.com> Co-authored-by: davidmyriel <davidmyriel@gmail.com> * Qdrant 1.10: add vectors page with description of vector types and datatypes (#995) * add vectors page with description of vector types and datatypes * docs: csharp vectors.md * docs: Java vectors.md * docs: Java, Csharp formatting * wip: rust and typescript snippets * Update qdrant-landing/content/documentation/concepts/vectors.md Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com> * Update qdrant-landing/content/documentation/concepts/vectors.md Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com> * add rest of snippets * add proofread and edit text --------- Co-authored-by: Anush <anushshetty90@gmail.com> Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com> Co-authored-by: davidmyriel <davidmyriel@gmail.com> * Qdrant 1.10.0: document S3 snapshots, reorganize storage text (#994) * Move snapshots storage section down, add local fs and S3 sub sections * Remove yaml suffix from S3 secrets * Update S3 example configuration, it didn't match latest release * Update GitHub Stars (#993) Co-authored-by: generall <generall@users.noreply.github.com> * Add documentation for the Query API [v1.10] (#990) * Add documentation for the Query API * review suggestions * upd url * Add rrf image, fix typos * easier intro image to fusion * add python snippets * docs: java, csharp * docs: csharp hybrid-queries.md * docs: Format hybrid-queries.md * Add Rust client search examples using Query API * docs: Java hybrid-queries.md * doc: Helper comment * Add Rust examples in hybrid queries * v1.10.0 typescript snippets (#999) * v1.10.0 typescript snippets * missed using * missed comments * missed docs * Address @agourlay's review * fix proofread & edit --------- Co-authored-by: generall <andrey@vasnetsov.com> Co-authored-by: Anush <anushshetty90@gmail.com> Co-authored-by: timvisee <tim@visee.me> Co-authored-by: Ivan Pleshkov <pleshkov.ivan@gmail.com> Co-authored-by: davidmyriel <davidmyriel@gmail.com> * [WIP] Version 1.10. Release Article (#992) * add release article draft * add wip links * fix article linnk * add more information * fix link * add parameters * add benchmark info * better issues window screenshot * fix image and descriptions * add last bits of info * update release header * fix title * docs: Java, Csharp snippets * docs: Fixed typos * Removed first Java, C# snippets for a simpler intro * Update S3 snapshot storage, link to configuration section * Update S3 configuration example * Add Rust examples * Update new Rust client text * Mention ColBERT Rust snippet as good client example * Update qdrant-landing/content/blog/qdrant-1.10.x.md Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> * Update qdrant-landing/content/blog/qdrant-1.10.x.md Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> * Update qdrant-landing/content/blog/qdrant-1.10.x.md Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> * Update qdrant-landing/content/blog/qdrant-1.10.x.md Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> * Update qdrant-landing/content/blog/qdrant-1.10.x.md Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> * Update qdrant-landing/content/blog/qdrant-1.10.x.md Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> * Update qdrant-landing/content/blog/qdrant-1.10.x.md Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> * Update qdrant-landing/content/blog/qdrant-1.10.x.md Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> * Update qdrant-landing/content/blog/qdrant-1.10.x.md Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> * add last few changes * Fix internal link * Update Rust documentation links to point to specific version --------- Co-authored-by: Luis Cossío <luis.cossio@outlook.com> Co-authored-by: Anush <anushshetty90@gmail.com> Co-authored-by: timvisee <tim@visee.me> Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> * upd links in release blog post * dont use links with domain * [draft] bm 42 article (#969) * bm 42 article draft * upd the article * BPE -> wordpiece * azeret mono as fallback font * testing other fallback for monospace font * self hosted fonts * only proofread & edit * Update qdrant-landing/content/articles/bm42.md * last fixes --------- Co-authored-by: trean <trean.mi@gmail.com> Co-authored-by: davidmyriel <davidmyriel@gmail.com> * fix link --------- Co-authored-by: Luis Cossío <luis.cossio@outlook.com> Co-authored-by: Anush <anushshetty90@gmail.com> Co-authored-by: davidmyriel <davidmyriel@gmail.com> Co-authored-by: Arnaud Gourlay <arnaud.gourlay@gmail.com> Co-authored-by: Tim Visée <tim@visee.me> Co-authored-by: generall <generall@users.noreply.github.com> Co-authored-by: Luis Cossío <luis.cossio@qdrant.com> Co-authored-by: Ivan Pleshkov <pleshkov.ivan@gmail.com> Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> Co-authored-by: trean <trean.mi@gmail.com>
@@ -0,0 +1,362 @@
|
||||
---
|
||||
title: "BM42: New Baseline for Hybrid Search"
|
||||
short_description: "Introducing next evolutionary step in lexical search."
|
||||
description: "Introducing BM42 - a new sparse embedding approach, which combines the benefits of exact keyword search with the intelligence of transformers."
|
||||
social_preview_image: /articles_data/bm42/social-preview.jpg
|
||||
preview_dir: /articles_data/bm42/preview
|
||||
weight: -140
|
||||
author: Andrey Vasnetsov
|
||||
date: 2024-07-01T12:00:00+03:00
|
||||
draft: false
|
||||
keywords:
|
||||
- hybrid search
|
||||
- sparse embeddings
|
||||
- bm25
|
||||
---
|
||||
|
||||
For the last 40 years, BM25 has served as the standard for search engines.
|
||||
It is a simple yet powerful algorithm that has been used by many search engines, including Google, Bing, and Yahoo.
|
||||
|
||||
Though it seemed that the advent of vector search would diminish its influence, it did so only partially.
|
||||
The current state-of-the-art approach to retrieval nowadays tries to incorporate BM25 along with embeddings into a hybrid search system.
|
||||
|
||||
However, the use case of text retrieval has significantly shifted since the introduction of RAG.
|
||||
Many assumptions upon which BM25 was built are no longer valid.
|
||||
|
||||
For example, the typical length of documents and queries vary significantly between traditional web search and modern RAG systems.
|
||||
|
||||
In this article, we will recap what made BM25 relevant for so long and why alternatives have struggled to replace it. Finally, we will discuss BM42, as the next step in the evolution of lexical search.
|
||||
|
||||
## Why has BM25 stayed relevant for so long?
|
||||
|
||||
To understand why, we need to analyze its components.
|
||||
|
||||
The famous BM25 formula is defined as:
|
||||
|
||||
$$
|
||||
\text{score}(D,Q) = \sum_{i=1}^{N} \text{IDF}(q_i) \times \frac{f(q_i, D) \cdot (k_1 + 1)}{f(q_i, D) + k_1 \cdot \left(1 - b + b \cdot \frac{|D|}{\text{avgdl}}\right)}
|
||||
$$
|
||||
|
||||
Let's simplify this to gain a better understanding.
|
||||
|
||||
- The $score(D, Q)$ - means that we compute the score for each pair of document $D$ and query $Q$.
|
||||
|
||||
- The $\sum_{i=1}^{N}$ - means that each of $N$ terms in the query contribute to the final score as a part of the sum.
|
||||
|
||||
- The $\text{IDF}(q_i)$ - is the inverse document frequency. The more rare the term $q_i$ is, the more it contributes to the score. A simplified formula for this is:
|
||||
|
||||
$$
|
||||
\text{IDF}(q_i) = \frac{\text{Number of documents}}{\text{Number of documents with } q_i}
|
||||
$$
|
||||
|
||||
It is fair to say that the `IDF` is the most important part of the BM25 formula.
|
||||
`IDF` selects the most important terms in the query relative to the specific document collection.
|
||||
So intuitively, we can interpret the `IDF` as **term importance within the corpora**.
|
||||
|
||||
That explains why BM25 is so good at handling queries, which dense embeddings consider out-of-domain.
|
||||
|
||||
The last component of the formula can be intuitively interpreted as **term importance within the document**.
|
||||
This might look a bit complicated, so let's break it down.
|
||||
|
||||
$$
|
||||
\text{Term importance in document }(q_i) = \color{red}\frac{f(q_i, D)\color{black} \cdot \color{blue}(k_1 + 1) \color{black} }{\color{red}f(q_i, D)\color{black} + \color{blue}k_1\color{black} \cdot \left(1 - \color{blue}b\color{black} + \color{blue}b\color{black} \cdot \frac{|D|}{\text{avgdl}}\right)}
|
||||
$$
|
||||
|
||||
- The $\color{red}f(q_i, D)\color{black}$ - is the frequency of the term $q_i$ in the document $D$. Or in other words, the number of times the term $q_i$ appears in the document $D$.
|
||||
- The $\color{blue}k_1\color{black}$ and $\color{blue}b\color{black}$ are the hyperparameters of the BM25 formula. In most implementations, they are constants set to $k_1=1.5$ and $b=0.75$. Those constants define relative implications of the term frequency and the document length in the formula.
|
||||
- The $\frac{|D|}{\text{avgdl}}$ - is the relative length of the document $D$ compared to the average document length in the corpora. The intuition befind this part is following: if the token is found in the smaller document, it is more likely that this token is important for this document.
|
||||
|
||||
#### Will BM25 term importance in the document work for RAG?
|
||||
|
||||
As we can see, the *term importance in the document* heavily depends on the statistics within the document. Moreover, statistics works well if the document is long enough.
|
||||
Therefore, it is suitable for searching webpages, books, articles, etc.
|
||||
|
||||
However, would it work as well for modern search applications, such as RAG? Let's see.
|
||||
|
||||
The typical length of a document in RAG is much shorter than that of web search. In fact, even if we are working with webpages and articles, we would prefer to split them into chunks so that
|
||||
a) Dense models can handle them and
|
||||
b) We can pinpoint the exact part of the document which is relevant to the query
|
||||
|
||||
As a result, the document size in RAG is small and fixed.
|
||||
|
||||
That effectively renders the term importance in the document part of the BM25 formula useless.
|
||||
The term frequency in the document is always 0 or 1, and the relative length of the document is always 1.
|
||||
|
||||
So, the only part of the BM25 formula that is still relevant for RAG is `IDF`. Let's see how we can leverage it.
|
||||
|
||||
## Why SPLADE is not always the answer
|
||||
|
||||
Before discussing our new approach, let's examine the current state-of-the-art alternative to BM25 - SPLADE.
|
||||
|
||||
The idea behind SPLADE is interesting—what if we let a smart, end-to-end trained model generate a bag-of-words representation of the text for us?
|
||||
It will assign all the weights to the tokens, so we won't need to bother with statistics and hyperparameters.
|
||||
The documents are then represented as a sparse embedding, where each token is represented as an element of the sparse vector.
|
||||
|
||||
And it works in academic benchmarks. Many papers report that SPLADE outperforms BM25 in terms of retrieval quality.
|
||||
This performance, however, comes at a cost.
|
||||
|
||||
* **Inappropriate Tokenizer**: To incorporate transformers for this task, SPLADE models require using a standard transformer tokenizer. These tokenizers are not designed for retrieval tasks. For example, if the word is not in the (quite limited) vocabulary, it will be either split into subwords or replaced with a `[UNK]` token. This behavior works well for language modeling but is completely destructive for retrieval tasks.
|
||||
|
||||
* **Expensive Token Expansion**: In order to compensate the tokenization issues, SPLADE uses *token expansion* technique. This means that we generate a set of similar tokens for each token in the query. There are a few problems with this approach:
|
||||
- It is computationally and memory expensive. We need to generate more values for each token in the document, which increases both the storage size and retrieval time.
|
||||
- It is not always clear where to stop with the token expansion. The more tokens we generate, the more likely we are to get the relevant one. But simultaneously, the more tokens we generate, the more likely we are to get irrelevant results.
|
||||
- Token expansion dilutes the interpretability of the search. We can't say which tokens were used in the document and which were generated by the token expansion.
|
||||
|
||||
* **Domain and Language Dependency**: SPLADE models are trained on specific corpora. This means that they are not always generalizable to new or rare domains. As they don't use any statistics from the corpora, they cannot adapt to the new domain without fine-tuning.
|
||||
|
||||
* **Inference Time**: Additionally, currently available SPLADE models are quite big and slow. They usually require a GPU to make the inference in a reasonable time.
|
||||
|
||||
At Qdrant, we acknowledge the aforementioned problems and are looking for a solution.
|
||||
Our idea was to combine the best of both worlds - the simplicity and interpretability of BM25 and the intelligence of transformers while avoiding the pitfalls of SPLADE.
|
||||
|
||||
And here is what we came up with.
|
||||
|
||||
## The best of both worlds
|
||||
|
||||
As previously mentioned, `IDF` is the most important part of the BM25 formula. In fact it is so important, that we decided to build its calculation into the Qdrant engine itself.
|
||||
Check out our latest [release notes](https://github.com/qdrant/qdrant/releases/tag/v1.10.0). This type of separation allows streaming updates of the sparse embeddings while keeping the `IDF` calculation up-to-date.
|
||||
|
||||
As for the second part of the formula, *the term importance within the document* needs to be rethought.
|
||||
|
||||
Since we can't rely on the statistics within the document, we can try to use the semantics of the document instead.
|
||||
And semantics is what transformers are good at. Therefore, we only need to solve two problems:
|
||||
|
||||
- How does one extract the importance information from the transformer?
|
||||
- How can tokenization issues be avoided?
|
||||
|
||||
|
||||
### Attention is all you need
|
||||
|
||||
Transformer models, even those used to generate embeddings, generate a bunch of different outputs.
|
||||
Some of those outputs are used to generate embeddings.
|
||||
|
||||
Others are used to solve other kinds of tasks, such as classification, text generation, etc.
|
||||
|
||||
The one particularly interesting output for us is the attention matrix.
|
||||
|
||||
{{< figure src="/articles_data/bm42/attention-matrix.png" alt="Attention matrix" caption="Attention matrix" width="60%" >}}
|
||||
|
||||
The attention matrix is a square matrix, where each row and column corresponds to the token in the input sequence.
|
||||
It represents the importance of each token in the input sequence for each other.
|
||||
|
||||
The classical transformer models are trained to predict masked tokens in the context, so the attention weights define which context tokens influence the masked token most.
|
||||
|
||||
Apart from regular text tokens, the transformer model also has a special token called `[CLS]`. This token represents the whole sequence in the classification tasks, which is exactly what we need.
|
||||
|
||||
By looking at the attention row for the `[CLS]` token, we can get the importance of each token in the document for the whole document.
|
||||
|
||||
|
||||
```python
|
||||
sentences = "Hello, World - is the starting point in most programming languages"
|
||||
|
||||
features = transformer.tokenize(sentences)
|
||||
|
||||
# ...
|
||||
|
||||
attentions = transformer.auto_model(**features, output_attentions=True).attentions
|
||||
|
||||
weights = torch.mean(attentions[-1][0,:,0], axis=0)
|
||||
# ▲ ▲ ▲ ▲
|
||||
# │ │ │ └─── [CLS] token is the first one
|
||||
# │ │ └─────── First item of the batch
|
||||
# │ └────────── Last transformer layer
|
||||
# └────────────────────────── Averate all 6 attention heads
|
||||
|
||||
for weight, token in zip(weights, tokens):
|
||||
print(f"{token}: {weight}")
|
||||
|
||||
# [CLS] : 0.434 // Filter out the [CLS] token
|
||||
# hello : 0.039
|
||||
# , : 0.039
|
||||
# world : 0.107 // <-- The most important token
|
||||
# - : 0.033
|
||||
# is : 0.024
|
||||
# the : 0.031
|
||||
# starting : 0.054
|
||||
# point : 0.028
|
||||
# in : 0.018
|
||||
# most : 0.016
|
||||
# programming : 0.060 // <-- The third most important token
|
||||
# languages : 0.062 // <-- The second most important token
|
||||
# [SEP] : 0.047 // Filter out the [SEP] token
|
||||
|
||||
```
|
||||
|
||||
|
||||
The resulting formula for the BM42 score would look like this:
|
||||
|
||||
$$
|
||||
\text{score}(D,Q) = \sum_{i=1}^{N} \text{IDF}(q_i) \times \text{Attention}(\text{CLS}, q_i)
|
||||
$$
|
||||
|
||||
|
||||
Note that classical transformers have multiple attention heads, so we can get multiple importance vectors for the same document. The simplest way to combine them is to simply average them.
|
||||
|
||||
These averaged attention vectors make up the importance information we were looking for.
|
||||
The best part is, one can get them from any transformer model, without any additional training.
|
||||
Therefore, BM42 can support any natural language as long as there is a transformer model for it.
|
||||
|
||||
In our implementation, we use the `sentence-transformers/all-MiniLM-L6-v2` model, which gives a huge boost in the inference speed compared to the SPLADE models. In practice, any transformer model can be used.
|
||||
It doesn't require any additional training, and can be easily adapted to work as BM42 backend.
|
||||
|
||||
|
||||
### WordPiece retokenization
|
||||
|
||||
The final piece of the puzzle we need to solve is the tokenization issue. In order to get attention vectors, we need to use native transformer tokenization.
|
||||
But this tokenization is not suitable for the retrieval tasks. What can we do about it?
|
||||
|
||||
Actually, the solution we came up with is quite simple. We reverse the tokenization process after we get the attention vectors.
|
||||
|
||||
Transformers use [WordPiece](https://huggingface.co/learn/nlp-course/en/chapter6/6) tokenization.
|
||||
In case it sees the word, which is not in the vocabulary, it splits it into subwords.
|
||||
|
||||
Here is how that looks:
|
||||
|
||||
```text
|
||||
"unbelievable" -> ["un", "##believ", "##able"]
|
||||
```
|
||||
|
||||
What can merge the subwords back into the words. Luckily, the subwords are marked with the `##` prefix, so we can easily detect them.
|
||||
Since the attention weights are normalized, we can simply sum the attention weights of the subwords to get the attention weight of the word.
|
||||
|
||||
After that, we can apply the same traditional NLP techniques, as
|
||||
|
||||
- Removing of the stop-words
|
||||
- Removing of the punctuation
|
||||
- Lemmatization
|
||||
|
||||
In this way, we can significantly reduce the number of tokens, and therefore minimize the memory footprint of the sparse embeddings. We won't simultaneously compromise the ability to match (almost) exact tokens.
|
||||
|
||||
## Practical examples
|
||||
|
||||
|
||||
| Trait | BM25 | SPLADE | BM42 |
|
||||
|-------------------------|--------------|--------------|--------------|
|
||||
| Interpretability | High ✅ | Ok 🆗 | High ✅ |
|
||||
| Document Inference speed| Very high ✅ | Slow 🐌 | High ✅ |
|
||||
| Query Inference speed | Very high ✅ | Slow 🐌 | Very high ✅ |
|
||||
| Memory footprint | Low ✅ | High ❌ | Low ✅ |
|
||||
| In-domain accuracy | Ok 🆗 | High ✅ | High ✅ |
|
||||
| Out-of-domain accuracy | Ok 🆗 | Low ❌ | Ok 🆗 |
|
||||
| Small documents accuracy| Low ❌ | High ✅ | High ✅ |
|
||||
| Large documents accuracy| High ✅ | Low ❌ | Ok 🆗 |
|
||||
| Unknown tokens handling | Yes ✅ | Bad ❌ | Yes ✅ |
|
||||
| Multi-lingual support | Yes ✅ | No ❌ | Yes ✅ |
|
||||
| Best Match | Yes ✅ | No ❌ | Yes ✅ |
|
||||
|
||||
|
||||
Starting from Qdrant v1.10.0, BM42 can be used in Qdrant via FastEmbed inference.
|
||||
|
||||
Let's see how you can setup a collection for hybrid search with BM42 and [jina.ai](https://jina.ai/embeddings/) dense embeddings.
|
||||
|
||||
```http
|
||||
PUT collections/my-hybrid-collection
|
||||
{
|
||||
"vectors": {
|
||||
"jina": {
|
||||
"size": 768,
|
||||
"distance": "Cosine"
|
||||
}
|
||||
},
|
||||
"sparse_vectors": {
|
||||
"bm42": {
|
||||
"modifier": "idf" // <--- This parameter enables the IDF calculation
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient()
|
||||
|
||||
client.create_collection(
|
||||
collection_name="my-hybrid-collection",
|
||||
vectors_config={
|
||||
"jina": models.VectorParams(
|
||||
size=768,
|
||||
distance=models.Distance.COSINE,
|
||||
)
|
||||
},
|
||||
sparse_vectors_config={
|
||||
"bm42": models.SparseVectorParams(
|
||||
modifier=models.Modifier.IDF,
|
||||
)
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
The search query will retrieve the documents with both dense and sparse embeddings and combine the scores
|
||||
using the Reciprocal Rank Fusion (RRF) algorithm.
|
||||
|
||||
```python
|
||||
from fastembed import SparseTextEmbedding, TextEmbedding
|
||||
|
||||
query_text = "best programming language for beginners?"
|
||||
|
||||
model_bm42 = SparseTextEmbedding(model_name="Qdrant/bm42-all-minilm-l6-v2-attentions")
|
||||
model_jina = TextEmbedding(model_name="jinaai/jina-embeddings-v2-base-en")
|
||||
|
||||
sparse_embedding = list(embedding_model.embed_query(query_text))[0]
|
||||
dense_embedding = list(model_jina.embed_query(query_text))[0]
|
||||
|
||||
client.query_points(
|
||||
collection_name="my-hybrid-collection",
|
||||
prefetch=[
|
||||
models.Prefetch(query=sparse_embedding, using="bm42", limit=10),
|
||||
models.Prefetch(query=dense_embedding, using="jina", limit=10),
|
||||
],
|
||||
query=models.FusionQuery(fusion=models.Fusion.RRF), # <--- Combine the scores
|
||||
limit=10
|
||||
)
|
||||
|
||||
```
|
||||
|
||||
### Benchmarks
|
||||
|
||||
To prove the point further we have conducted some benchmarks to highlight the cases where BM42 outperforms BM25.
|
||||
Please note, that we didn't intend to make an exhaustive evaluation, as we are presenting a new approach, not a new model.
|
||||
|
||||
For out experiments we choose [quora](https://huggingface.co/datasets/BeIR/quora) dataset, as a good representative of the Question-Answering task.
|
||||
The typical example of the dataset is the following:
|
||||
|
||||
```text
|
||||
{"_id": "109", "text": "How GST affects the CAs and tax officers?"}
|
||||
{"_id": "110", "text": "Why can't I do my homework?"}
|
||||
{"_id": "111", "text": "How difficult is it get into RSI?"}
|
||||
```
|
||||
|
||||
As you can see, it has pretty short texts, there are not much of the statistics to rely on.
|
||||
|
||||
After encoding with BM42, the average vector size is only **5.6 elements per document**.
|
||||
|
||||
With `datatype: uint8` available in Qdrant, the total size of the sparse vector index is about **13Mb** for ~530k documents.
|
||||
|
||||
|
||||
| | BM25 | BM42 |
|
||||
|---------------|------|----------|
|
||||
|Precision @ 10 | 0.45 | **0.49** |
|
||||
|
||||
|
||||
Please note, that both BM25 and BM42 won't work well on their own in a production environment.
|
||||
Best results are achieved with a combination of sparse and dense embeddings in a hybrid approach.
|
||||
In this scenario, the two models are complementary to each other.
|
||||
The sparse model is responsible for exact token matching, while the dense model is responsible for semantic matching.
|
||||
|
||||
Some more advanced models might outperform default `sentence-transformers/all-MiniLM-L6-v2` model we were using.
|
||||
We encourage developers involved in training embedding models to include a way to extract attention weights and contribute to the BM42 backend.
|
||||
|
||||
## Fostering curiosity and experimentation
|
||||
|
||||
Despite all of its advantages, BM42 is not always a silver bullet.
|
||||
For large documents without chunks, BM25 might still be a better choice.
|
||||
|
||||
There might be a smarter way to extract the importance information from the transformer. There could be a better method to weigh IDF against attention scores.
|
||||
|
||||
Qdrant does not specialize in model training. Our core project is the search engine itself. However, we understand that we are not operating in a vacuum. By introducing BM42, we are stepping up to empower our community with novel tools for experimentation.
|
||||
|
||||
We truly believe that the sparse vectors method is at exact level of abstraction to yield both powerful and flexible results.
|
||||
|
||||
Many of you are sharing your recent Qdrant projects in our [Discord channel](https://discord.com/invite/qdrant). Feel free to try out BM42 and let us know what you come up with.
|
||||
|
||||
@@ -0,0 +1,574 @@
|
||||
---
|
||||
title: "Qdrant 1.10 - Universal Query, Built-in IDF & ColBERT Support"
|
||||
draft: false
|
||||
short_description: "Single search API. Server-side IDF. Native multivector support."
|
||||
description: "Consolidated search API, built-in IDF, and native multivector support. "
|
||||
preview_image: /blog/qdrant-1.10.x/social_preview.png
|
||||
social_preview_image: /blog/qdrant-1.10.x/social_preview.png
|
||||
date: 2024-06-25T00:00:00-08:00
|
||||
author: David Myriel
|
||||
featured: false
|
||||
tags:
|
||||
- vector search
|
||||
- ColBERT late interaction
|
||||
- BM25 algorithm
|
||||
- search API
|
||||
- new features
|
||||
---
|
||||
|
||||
[Qdrant 1.10.0 is out!](https://github.com/qdrant/qdrant/releases/tag/v1.10.0) This version introduces some major changes, so let's dive right in:
|
||||
|
||||
**Universal Query API:** All search APIs, including Hybrid Search, are now in one Query endpoint.</br>
|
||||
**Built-in IDF:** We added the IDF mechanism to Qdrant's core search and indexing processes.</br>
|
||||
**Multivector Support:** Native support for late interaction ColBERT is accessible via Query API.
|
||||
|
||||
## One Endpoint for All Queries
|
||||
**Query API** will consolidate all search APIs into a single request. Previously, you had to work outside of the API to combine different search requests. Now these approaches are reduced to parameters of a single request, so you can avoid merging individual results.
|
||||
|
||||
You can now configure the Query API request with the following parameters:
|
||||
|
||||
|Parameter|Description|
|
||||
|-|-|
|
||||
|no parameter|Returns points by `id`|
|
||||
|`nearest`|Queries nearest neighbors ([Search](/documentation/concepts/search/))|
|
||||
|`fusion`|Fuses sparse/dense prefetch queries ([Hybrid Search](/documentation/concepts/hybrid-queries/#hybrid-search))|
|
||||
|`discover`|Queries `target` with added `context` ([Discovery](/documentation/concepts/explore/#discovery-api))|
|
||||
|`context` |No target with `context` only ([Context](/documentation/concepts/explore/#context-search))|
|
||||
|`recommend`|Queries against `positive`/`negative` examples. ([Recommendation](/documentation/concepts/explore/#recommendation-api))|
|
||||
|`order_by`|Orders results by [payload field](/documentation/concepts/hybrid-queries/#re-ranking-with-payload-values)|
|
||||
|
||||
For example, you can configure Query API to run [Discovery search](/documentation/concepts/explore/#discovery-api). Let's see how that looks:
|
||||
|
||||
```http
|
||||
POST collections/{collection_name}/points/query
|
||||
{
|
||||
"query": {
|
||||
"discover": {
|
||||
"target": <vector_input>,
|
||||
"context": [
|
||||
{
|
||||
"positive": <vector_input>,
|
||||
"negative": <vector_input>
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
We will be publishing code samples in [docs](/documentation/concepts/hybrid-queries/) and our new [API specification](http://api.qdrant.tech).</br> *If you need additional support with this new method, our [Discord](https://qdrant.to/discord) on-call engineers can help you.*
|
||||
|
||||
### Native Hybrid Search Support
|
||||
Query API now also natively supports **sparse/dense fusion**. Up to this point, you had to combine the results of sparse and dense searches on your own. This is now sorted on the back-end, and you only have to configure them as basic parameters for Query API.
|
||||
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"prefetch": [
|
||||
{
|
||||
"query": {
|
||||
"indices": [1, 42], // <┐
|
||||
"values": [0.22, 0.8] // <┴─sparse vector
|
||||
},
|
||||
"using": "sparse",
|
||||
"limit": 20,
|
||||
},
|
||||
{
|
||||
"query": [0.01, 0.45, 0.67, ...], // <-- dense vector
|
||||
"using": "dense",
|
||||
"limit": 20,
|
||||
}
|
||||
],
|
||||
"query": { "fusion": "rrf" }, // <--- reciprocal rank fusion
|
||||
"limit": 10
|
||||
}
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{Fusion, PrefetchQueryBuilder, Query, QueryPointsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.add_prefetch(PrefetchQueryBuilder::default()
|
||||
.query(Query::new_nearest([(1, 0.22), (42, 0.8)].as_slice()))
|
||||
.using("sparse")
|
||||
.limit(20u64)
|
||||
)
|
||||
.add_prefetch(PrefetchQueryBuilder::default()
|
||||
.query(Query::new_nearest(vec![0.01, 0.45, 0.67]))
|
||||
.using("dense")
|
||||
.limit(20u64)
|
||||
)
|
||||
.query(Query::new_fusion(Fusion::Rrf))
|
||||
).await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
import static io.qdrant.client.QueryFactory.fusion;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.Fusion;
|
||||
import io.qdrant.client.grpc.Points.PrefetchQuery;
|
||||
import io.qdrant.client.grpc.Points.QueryPoints;
|
||||
|
||||
QdrantClient client = new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client.queryAsync(
|
||||
QueryPoints.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.addPrefetch(PrefetchQuery.newBuilder()
|
||||
.setQuery(nearest(List.of(0.22f, 0.8f), List.of(1, 42)))
|
||||
.setUsing("sparse")
|
||||
.setLimit(20)
|
||||
.build())
|
||||
.addPrefetch(PrefetchQuery.newBuilder()
|
||||
.setQuery(nearest(List.of(0.01f, 0.45f, 0.67f)))
|
||||
.setUsing("dense")
|
||||
.setLimit(20)
|
||||
.build())
|
||||
.setQuery(fusion(Fusion.RRF))
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
prefetch: new List < PrefetchQuery > {
|
||||
new() {
|
||||
Query = new(float, uint)[] {
|
||||
(0.22f, 1), (0.8f, 42),
|
||||
},
|
||||
Using = "sparse",
|
||||
Limit = 20
|
||||
},
|
||||
new() {
|
||||
Query = new float[] {
|
||||
0.01f, 0.45f, 0.67f
|
||||
},
|
||||
Using = "dense",
|
||||
Limit = 20
|
||||
}
|
||||
},
|
||||
query: Fusion.Rrf
|
||||
);
|
||||
```
|
||||
|
||||
Query API can now pre-fetch vectors for requests, which means you can run queries sequentially within the same API call. There are a lot of options here, so you will need to define a strategy to merge these requests using new parameters. For example, you can now include **rescoring within Hybrid Search**, which can open the door to strategies like iterative refinement via matryoshka embeddings.
|
||||
|
||||
*To learn more about this, read the [Query API documentation](/documentation/concepts/search/#query-api).*
|
||||
|
||||
## Inverse Document Frequency [IDF]
|
||||
|
||||
IDF is a critical component of the **TF-IDF (Term Frequency-Inverse Document Frequency)** weighting scheme used to evaluate the importance of a word in a document relative to a collection of documents (corpus).
|
||||
There are various ways in which IDF might be calculated, but the most commonly used formula is:
|
||||
|
||||
$$
|
||||
\text{IDF}(q_i) = \ln \left(\frac{N - n(q_i) + 0.5}{n(q_i) + 0.5}+1\right)
|
||||
$$
|
||||
|
||||
Where:</br>
|
||||
`N` is the total number of documents in the collection. </br>
|
||||
`n` is the number of documents containing non-zero values for the given vector.
|
||||
|
||||
This variant is also used in BM25, whose support was heavily requested by our users. We decided to move the IDF calculation into the Qdrant engine itself. This type of separation allows streaming updates of the sparse embeddings while keeping the IDF calculation up-to-date.
|
||||
|
||||
The values of IDF previously had to be calculated using all the documents on the client side. However, now that Qdrant does it out of the box, you won't need to implement it anywhere else and recompute the value if some documents are removed or newly added.
|
||||
|
||||
You can enable the IDF modifier in the collection configuration:
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"sparse_vectors": {
|
||||
"text": {
|
||||
"modifier": "idf"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
sparse_vectors={
|
||||
"text": models.SparseVectorParams(
|
||||
modifier=models.Modifier.IDF,
|
||||
),
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{CreateCollectionBuilder, sparse_vectors_config::SparseVectorsConfigBuilder, Modifier, SparseVectorParamsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
let mut config = SparseVectorsConfigBuilder::default();
|
||||
config.add_named_vector_params(
|
||||
"text",
|
||||
SparseVectorParamsBuilder::default().modifier(Modifier::Idf),
|
||||
);
|
||||
|
||||
client
|
||||
.create_collection(
|
||||
CreateCollectionBuilder::new("{collection_name}")
|
||||
.sparse_vectors_config(config),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||
import io.qdrant.client.grpc.Collections.Modifier;
|
||||
import io.qdrant.client.grpc.Collections.SparseVectorConfig;
|
||||
import io.qdrant.client.grpc.Collections.SparseVectorParams;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.createCollectionAsync(
|
||||
CreateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setSparseVectorsConfig(
|
||||
SparseVectorConfig.newBuilder()
|
||||
.putMap("text", SparseVectorParams.newBuilder().setModifier(Modifier.Idf).build()))
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
sparseVectorsConfig: ("text", new SparseVectorParams {
|
||||
Modifier = Modifier.Idf,
|
||||
})
|
||||
);
|
||||
```
|
||||
|
||||
### IDF as Part of BM42
|
||||
|
||||
This quarter, Qdrant also introduced BM42, a novel algorithm that combines the IDF element of BM25 with transformer-based attention matrices to improve text retrieval. It utilizes attention matrices from your embedding model to determine the importance of each token in the document based on the attention value it receives.
|
||||
|
||||
We've prepared the standard `all-MiniLM-L6-v2` Sentence Transformer so [it outputs the attention values](https://huggingface.co/Qdrant/all_miniLM_L6_v2_with_attentions). Still, you can use virtually any model of your choice, as long as you have access to its parameters. This is just another reason to stick with open source technologies over proprietary systems.
|
||||
|
||||
In practical terms, the BM42 method addresses the tokenization issues and computational costs associated with SPLADE. The model is both efficient and effective across different document types and lengths, offering enhanced search performance by leveraging the strengths of both BM25 and modern transformer techniques.
|
||||
|
||||
> To learn more about IDF and BM42, read our [dedicated technical article](/articles/bm42/).
|
||||
|
||||
**You can expect BM42 to excel in scalable RAG-based scenarios where short texts are more common.** Document inference speed is much higher with BM42, which is critical for large-scale applications such as search engines, recommendation systems, and real-time decision-making systems.
|
||||
|
||||
## Multivector Support
|
||||
We are adding native support for multivector search that is compatible, e.g., with the late-interaction [ColBERT](https://github.com/stanford-futuredata/ColBERT) model. If you are working with high-dimensional similarity searches, **ColBERT is highly recommended as a reranking step in the Universal Query search.** You will experience better quality vector retrieval since ColBERT’s approach allows for deeper semantic understanding.
|
||||
|
||||
This model retains contextual information during query-document interaction, leading to better relevance scoring. In terms of efficiency and scalability benefits, documents and queries will be encoded separately, which gives an opportunity for pre-computation and storage of document embeddings for faster retrieval.
|
||||
|
||||
**Note:** *This feature supports all the original quantization compression methods, just the same as the regular search method.*
|
||||
|
||||
**Run a query with ColBERT vectors:**
|
||||
|
||||
Query API can handle exceedingly complex requests. The following example prefetches 1000 entries most similar to the given query using the `mrl_byte` named vector, then reranks them to get the best 100 matches with `full` named vector and eventually reranks them again to extract the top 10 results with the named vector called `colbert`. A single API call can now implement complex reranking schemes.
|
||||
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"prefetch": {
|
||||
"prefetch": {
|
||||
"query": [1, 23, 45, 67], // <------ small byte vector
|
||||
"using": "mrl_byte"
|
||||
"limit": 1000,
|
||||
},
|
||||
"query": [0.01, 0.45, 0.67, ...], // <-- full dense vector
|
||||
"using": "full"
|
||||
"limit": 100,
|
||||
},
|
||||
"query": [ // <─┐
|
||||
[0.1, 0.2, ...], // < │
|
||||
[0.2, 0.1, ...], // < ├─ multi-vector
|
||||
[0.8, 0.9, ...] // < │
|
||||
], // <─┘
|
||||
"using": "colbert",
|
||||
"limit": 10
|
||||
}
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{PrefetchQueryBuilder, Query, QueryPointsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.add_prefetch(PrefetchQueryBuilder::default()
|
||||
.add_prefetch(PrefetchQueryBuilder::default()
|
||||
.query(Query::new_nearest(vec![1.0, 23.0, 45.0, 67.0]))
|
||||
.using("mlr_byte")
|
||||
.limit(1000u64)
|
||||
)
|
||||
.query(Query::new_nearest(vec![0.01, 0.45, 0.67]))
|
||||
.using("full")
|
||||
.limit(100u64)
|
||||
)
|
||||
.query(Query::new_nearest(vec![
|
||||
vec![0.1, 0.2],
|
||||
vec![0.2, 0.1],
|
||||
vec![0.8, 0.9],
|
||||
]))
|
||||
.using("colbert")
|
||||
.limit(10u64)
|
||||
).await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.PrefetchQuery;
|
||||
import io.qdrant.client.grpc.Points.QueryPoints;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.queryAsync(
|
||||
QueryPoints.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.addPrefetch(
|
||||
PrefetchQuery.newBuilder()
|
||||
.addPrefetch(
|
||||
PrefetchQuery.newBuilder()
|
||||
.setQuery(nearest(1, 23, 45, 67)) // <------------- small byte vector
|
||||
.setUsing("mrl_byte")
|
||||
.setLimit(1000)
|
||||
.build())
|
||||
.setQuery(nearest(0.01f, 0.45f, 0.67f)) // <-- dense vector
|
||||
.setUsing("full")
|
||||
.setLimit(100)
|
||||
.build())
|
||||
.setQuery(
|
||||
nearest(
|
||||
new float[][] {
|
||||
{0.1f, 0.2f}, // <─┐
|
||||
{0.2f, 0.1f}, // < ├─ multi-vector
|
||||
{0.8f, 0.9f} // < ┘
|
||||
}))
|
||||
.setUsing("colbert")
|
||||
.setLimit(10)
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
prefetch: new List <PrefetchQuery> {
|
||||
new() {
|
||||
Prefetch = {
|
||||
new List <PrefetchQuery> {
|
||||
new() {
|
||||
Query = new float[] { 1, 23, 45, 67 }, // <------------- small byte vector
|
||||
Using = "mrl_byte",
|
||||
Limit = 1000
|
||||
},
|
||||
}
|
||||
},
|
||||
Query = new float[] {0.01f, 0.45f, 0.67f}, // <-- dense vector
|
||||
Using = "full",
|
||||
Limit = 100
|
||||
}
|
||||
},
|
||||
query: new float[][] {
|
||||
[0.1f, 0.2f], // <─┐
|
||||
[0.2f, 0.1f], // < ├─ multi-vector
|
||||
[0.8f, 0.9f] // < ┘
|
||||
},
|
||||
usingVector: "colbert",
|
||||
limit: 10
|
||||
);
|
||||
```
|
||||
|
||||
**Note:** *The multivector feature is not only useful for ColBERT; it can also be used in other ways.*</br>
|
||||
For instance, in e-commerce, you can use multi-vector to store multiple images of the same item. This serves as an alternative to the [group-by](/documentation/concepts/search/#grouping-api) method.
|
||||
|
||||
## Sparse Vectors Compression
|
||||
|
||||
In version 1.9, we introduced the `uint8` [vector datatype](/documentation/concepts/vectors/#datatypes) for sparse vectors, in order to support pre-quantized embeddings from companies like JinaAI and Cohere.
|
||||
This time, we are introducing a new datatype **for both sparse and dense vectors**, as well as a different way of **storing** these vectors.
|
||||
|
||||
**Datatype:** Sparse and dense vectors were previously represented in larger `float32` values, but now they can be turned to the `float16`. `float16` vectors have a lower precision compared to `float32`, which means that there is less numerical accuracy in the vector values - but this is negligible for practical use cases.
|
||||
|
||||
These vectors will use half the memory of regular vectors, which can significantly reduce the footprint of large vector datasets. Operations can be faster due to reduced memory bandwidth requirements and better cache utilization. This can lead to faster vector search operations, especially in memory-bound scenarios.
|
||||
|
||||
When creating a collection, you need to specify the `datatype` upfront:
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"vectors": {
|
||||
"size": 1024,
|
||||
"distance": "Cosine",
|
||||
"datatype": "float16"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```java
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||
import io.qdrant.client.grpc.Collections.Datatype;
|
||||
import io.qdrant.client.grpc.Collections.Distance;
|
||||
import io.qdrant.client.grpc.Collections.VectorParams;
|
||||
import io.qdrant.client.grpc.Collections.VectorsConfig;
|
||||
|
||||
QdrantClient client = new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.createCollectionAsync(
|
||||
CreateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setVectorsConfig(VectorsConfig.newBuilder()
|
||||
.setParams(VectorParams.newBuilder()
|
||||
.setSize(1024)
|
||||
.setDistance(Distance.Cosine)
|
||||
.setDatatype(Datatype.Float16)
|
||||
.build())
|
||||
.build())
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{CreateCollectionBuilder, Datatype, Distance, VectorParamsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client
|
||||
.create_collection(
|
||||
CreateCollectionBuilder::new("{collection_name}").vectors_config(
|
||||
VectorParamsBuilder::new(1024, Distance::Cosine).datatype(Datatype::Float16),
|
||||
),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
vectorsConfig: new VectorParams {
|
||||
Size = 1024,
|
||||
Distance = Distance.Cosine,
|
||||
Datatype = Datatype.Float16
|
||||
}
|
||||
);
|
||||
```
|
||||
|
||||
**Storage:** On the backend, we implemented bit packing to minimize the bits needed to store data, crucial for handling sparse vectors in applications like machine learning and data compression. For sparse vectors with mostly zeros, this focuses on storing only the indices and values of non-zero elements.
|
||||
|
||||
You will benefit from a more compact storage and higher processing efficiency. This can also lead to reduced dataset sizes for faster processing and lower storage costs in data compression.
|
||||
|
||||
## New Rust Client
|
||||
|
||||
Qdrant’s Rust client has been fully reshaped. It is now more accessible and
|
||||
easier to use. We have focused on putting together a minimalistic API interface.
|
||||
All operations and their types now use the builder pattern, providing an easy
|
||||
and extensible interface, preventing breakage with future updates. See the Rust
|
||||
[ColBERT query](#multivector-support) as great example. Additionally,
|
||||
Rust supports safe concurrent execution, which is crucial for handling multiple
|
||||
simultaneous requests efficiently.
|
||||
|
||||
Documentation got a significant improvement as well. It is much better organized
|
||||
and provides usage examples across the board. Everything links back to our main
|
||||
documentation, making it easier to navigate and find the information you need.
|
||||
|
||||
<p align="center">
|
||||
Visit our
|
||||
<a href="https://docs.rs/qdrant-client/1.10/qdrant_client/">client</a> and
|
||||
<a href="https://docs.rs/qdrant-client/1.10/qdrant_client/struct.Qdrant.html">operations</a> documentation
|
||||
</p>
|
||||
|
||||
## S3 Snapshot Storage
|
||||
Qdrant **Collections**, **Shards** and **Storage** can be backed up with [Snapshots](/documentation/concepts/snapshots/) and saved in case of data loss or other data transfer purposes. These snapshots can be quite large and the resources required to maintain them can result in higher costs. AWS S3 and other S3-compatible implementations like [min.io](https://min.io/) is a great low-cost alternative that can hold snapshots without incurring high costs. It is globally reliable, scalable and resistant to data loss.
|
||||
|
||||
You can configure S3 storage settings in the [config.yaml](https://github.com/qdrant/qdrant/blob/master/config/config.yaml), specifically with `snapshots_storage`.
|
||||
|
||||
For example, to use AWS S3:
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
snapshots_config:
|
||||
# Use 's3' to store snapshots on S3
|
||||
snapshots_storage: s3
|
||||
|
||||
s3_config:
|
||||
# Bucket name
|
||||
bucket: your_bucket_here
|
||||
|
||||
# Bucket region (e.g. eu-central-1)
|
||||
region: your_bucket_region_here
|
||||
|
||||
# Storage access key
|
||||
# Can be specified either here or in the `AWS_ACCESS_KEY_ID` environment variable.
|
||||
access_key: your_access_key_here
|
||||
|
||||
# Storage secret key
|
||||
# Can be specified either here or in the `AWS_SECRET_ACCESS_KEY` environment variable.
|
||||
secret_key: your_secret_key_here
|
||||
```
|
||||
|
||||
*Read more about [S3 snapshot storage](/documentation/concepts/snapshots/#s3) and [configuration](/documentation/guides/configuration/).*
|
||||
|
||||
This integration allows for a more convenient distribution of snapshots. Users of **any S3-compatible object storage** can now benefit from other platform services, such as automated workflows and disaster recovery options. S3's encryption and access control ensure secure storage and regulatory compliance. Additionally, S3 supports performance optimization through various storage classes and efficient data transfer methods, enabling quick and effective snapshot retrieval and management.
|
||||
|
||||
## Issues API
|
||||
Issues API notifies you about potential performance issues and misconfigurations. This powerful new feature allows users (such as database admins) to efficiently manage and track issues directly within the system, ensuring smoother operations and quicker resolutions.
|
||||
|
||||
You can find the Issues button in the top right. When you click the bell icon, a sidebar will open to show ongoing issues.
|
||||
|
||||

|
||||
|
||||
## Minor Improvements
|
||||
|
||||
- Pre-configure collection parameters; quantization, vector storage & replication factor - [#4299](https://github.com/qdrant/qdrant/pull/4299)
|
||||
|
||||
- Overwrite global optimizer configuration for collections. Lets you separate roles for indexing and searching within the single qdrant cluster - [#4317](https://github.com/qdrant/qdrant/pull/4317)
|
||||
|
||||
- Delta encoding and bitpacking compression for sparse vectors reduces memory consumption for sparse vectors by up to 75% - [#4253](https://github.com/qdrant/qdrant/pull/4253), [#4350](https://github.com/qdrant/qdrant/pull/4350)
|
||||
|
||||
@@ -30,6 +30,10 @@ A [Payload](/documentation/concepts/payload/) describes information that you can
|
||||
|
||||
[Explore](/documentation/concepts/explore/) includes several APIs for exploring data in your collections.
|
||||
|
||||
## Hybrid Queries
|
||||
|
||||
[Hybrid Queries](/documentation/concepts/hybrid-queries/) combines multiple queries or performs them in more than one stage.
|
||||
|
||||
## Filtering
|
||||
|
||||
[Filtering](/documentation/concepts/filtering/) defines various database-style clauses, conditions, and more.
|
||||
|
||||
@@ -604,6 +604,136 @@ Unlike a dense vector index, a sparse vector index does not require a pre-define
|
||||
|
||||
**Note:** A sparse vector index only supports dot-product similarity searches. It does not support other distance metrics.
|
||||
|
||||
### IDF Modifier
|
||||
|
||||
*Available as of v1.10.0*
|
||||
|
||||
For many search algorithms, it is important to consider how often an item occurs in a collection.
|
||||
Intuitively speaking, the less frequently an item appears in a collection, the more important it is in a search.
|
||||
|
||||
This is also known as the Inverse Document Frequency (IDF). It is used in text search engines to rank search results based on the rarity of a word in a collection.
|
||||
|
||||
IDF depends on the currently stored documents and therefore can't be pre-computed in the sparse vectors in streaming inference mode.
|
||||
In order to support IDF in the sparse vector index, Qdrant provides an option to modify the sparse vector query with the IDF statistics automatically.
|
||||
|
||||
The only requirement is to enable the IDF modifier in the collection configuration:
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"sparse_vectors": {
|
||||
"text": {
|
||||
"modifier": "idf"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
sparse_vectors={
|
||||
"text": models.SparseVectorParams(
|
||||
modifier=models.Modifier.IDF,
|
||||
),
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient, Schemas } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.createCollection("{collection_name}", {
|
||||
sparse_vectors: {
|
||||
"text": {
|
||||
modifier: "idf"
|
||||
}
|
||||
}
|
||||
});
|
||||
```
|
||||
|
||||
|
||||
```rust
|
||||
use qdrant_client::qdrant::{
|
||||
CreateCollectionBuilder, SparseVectorParamsBuilder,
|
||||
Modifier, sparse_vectors_config::SparseVectorsConfigBuilder
|
||||
};
|
||||
use qdrant_client::qdrant::;
|
||||
use qdrant_client::Qdrant;
|
||||
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
let mut sparse_vectors_config = SparseVectorsConfigBuilder::default();
|
||||
|
||||
sparse_vectors_config.add_named_vector_params(
|
||||
"text",
|
||||
SparseVectorParamsBuilder::default()
|
||||
.modifier(Modifier::Idf),
|
||||
);
|
||||
|
||||
client
|
||||
.create_collection(
|
||||
CreateCollectionBuilder::new("{collection_name}")
|
||||
.sparse_vectors_config(sparse_vectors_config)
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
|
||||
```java
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||
import io.qdrant.client.grpc.Collections.Modifier;
|
||||
import io.qdrant.client.grpc.Collections.SparseVectorConfig;
|
||||
import io.qdrant.client.grpc.Collections.SparseVectorParams;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.createCollectionAsync(
|
||||
CreateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setSparseVectorsConfig(
|
||||
SparseVectorConfig.newBuilder()
|
||||
.putMap("text", SparseVectorParams.newBuilder().setModifier(Modifier.Idf).build()))
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
sparseVectorsConfig: ("text", new SparseVectorParams {
|
||||
Modifier = Modifier.Idf,
|
||||
})
|
||||
);
|
||||
```
|
||||
|
||||
Qdrant uses the following formula to calculate the IDF modifier:
|
||||
|
||||
$$
|
||||
\text{IDF}(q_i) = \ln \left(\frac{N - n(q_i) + 0.5}{n(q_i) + 0.5}+1\right)
|
||||
$$
|
||||
|
||||
Where:
|
||||
|
||||
- `N` is the total number of documents in the collection.
|
||||
- `n` is the number of documents containing non-zero values for the given vector element.
|
||||
|
||||
## Filtrable Index
|
||||
|
||||
Separately, a payload index and a vector index cannot solve the problem of search using the filter completely.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: Payload
|
||||
weight: 40
|
||||
weight: 45
|
||||
aliases:
|
||||
- ../payload
|
||||
---
|
||||
|
||||
@@ -8,7 +8,18 @@ aliases:
|
||||
# Points
|
||||
|
||||
The points are the central entity that Qdrant operates with.
|
||||
A point is a record consisting of a vector and an optional [payload](../payload/).
|
||||
A point is a record consisting of a [vector](../vectors/) and an optional [payload](../payload/).
|
||||
|
||||
It looks like this:
|
||||
|
||||
```json
|
||||
// This is a simple point
|
||||
{
|
||||
"id": 129,
|
||||
"vector": [0.1, 0.2, 0.3, 0.4],
|
||||
"payload": {"color": "red"},
|
||||
}
|
||||
```
|
||||
|
||||
You can search among the points grouped in one [collection](../collections/) based on vector similarity.
|
||||
This procedure is described in more detail in the [search](../search/) and [filtering](../filtering/) sections.
|
||||
@@ -20,40 +31,6 @@ At the first stage, the operation is written to the Write-ahead-log.
|
||||
|
||||
After this moment, the service will not lose the data, even if the machine loses power supply.
|
||||
|
||||
## Awaiting result
|
||||
|
||||
If the API is called with the `&wait=false` parameter, or if it is not explicitly specified, the client will receive an acknowledgment of receiving data:
|
||||
|
||||
```json
|
||||
{
|
||||
"result": {
|
||||
"operation_id": 123,
|
||||
"status": "acknowledged"
|
||||
},
|
||||
"status": "ok",
|
||||
"time": 0.000206061
|
||||
}
|
||||
```
|
||||
|
||||
This response does not mean that the data is available for retrieval yet. This
|
||||
uses a form of eventual consistency. It may take a short amount of time before it
|
||||
is actually processed as updating the collection happens in the background. In
|
||||
fact, it is possible that such request eventually fails.
|
||||
If inserting a lot of vectors, we also recommend using asynchronous requests to take advantage of pipelining.
|
||||
|
||||
If the logic of your application requires a guarantee that the vector will be available for searching immediately after the API responds, then use the flag `?wait=true`.
|
||||
In this case, the API will return the result only after the operation is finished:
|
||||
|
||||
```json
|
||||
{
|
||||
"result": {
|
||||
"operation_id": 0,
|
||||
"status": "completed"
|
||||
},
|
||||
"status": "ok",
|
||||
"time": 0.000206061
|
||||
}
|
||||
```
|
||||
|
||||
## Point IDs
|
||||
|
||||
@@ -306,6 +283,26 @@ await client.UpsertAsync(
|
||||
|
||||
are both possible.
|
||||
|
||||
## Vectors
|
||||
|
||||
Each point in qdrant may have one or more vectors.
|
||||
Vectors are the central component of the Qdrant architecture,
|
||||
qdrant relies on different types of vectors to provide different types of data exploration and search.
|
||||
|
||||
Here is a list of supported vector types:
|
||||
|
||||
|||
|
||||
|-|-|
|
||||
| Dense Vectors | A regular vectors, generated by majority of the embedding models. |
|
||||
| Sparse Vectors | Vectors with no fixed length, but only a few non-zero elements. <br> Useful for exact token match and collaborative filtering recommendations. |
|
||||
| MultiVectors | Matrices of numbers with fixed length but variable height. <br> Usually obtained from late interraction models like ColBERT. |
|
||||
|
||||
It is possible to attach more than one type of vector to a single point.
|
||||
In Qdrant we call it Named Vectors.
|
||||
|
||||
Read more about vector types, how they are stored and optimized in the [vectors](../vectors/) section.
|
||||
|
||||
|
||||
## Upload points
|
||||
|
||||
To optimize performance, Qdrant supports batch loading of points. I.e., you can load several points into the service in one API call.
|
||||
@@ -2306,3 +2303,39 @@ client
|
||||
|
||||
To batch many points with a single operation type, please use batching
|
||||
functionality in that operation directly.
|
||||
|
||||
|
||||
## Awaiting result
|
||||
|
||||
If the API is called with the `&wait=false` parameter, or if it is not explicitly specified, the client will receive an acknowledgment of receiving data:
|
||||
|
||||
```json
|
||||
{
|
||||
"result": {
|
||||
"operation_id": 123,
|
||||
"status": "acknowledged"
|
||||
},
|
||||
"status": "ok",
|
||||
"time": 0.000206061
|
||||
}
|
||||
```
|
||||
|
||||
This response does not mean that the data is available for retrieval yet. This
|
||||
uses a form of eventual consistency. It may take a short amount of time before it
|
||||
is actually processed as updating the collection happens in the background. In
|
||||
fact, it is possible that such request eventually fails.
|
||||
If inserting a lot of vectors, we also recommend using asynchronous requests to take advantage of pipelining.
|
||||
|
||||
If the logic of your application requires a guarantee that the vector will be available for searching immediately after the API responds, then use the flag `?wait=true`.
|
||||
In this case, the API will return the result only after the operation is finished:
|
||||
|
||||
```json
|
||||
{
|
||||
"result": {
|
||||
"operation_id": 0,
|
||||
"status": "completed"
|
||||
},
|
||||
"status": "ok",
|
||||
"time": 0.000206061
|
||||
}
|
||||
```
|
||||
@@ -11,13 +11,171 @@ Searching for the nearest vectors is at the core of many representational learni
|
||||
Modern neural networks are trained to transform objects into vectors so that objects close in the real world appear close in vector space.
|
||||
It could be, for example, texts with similar meanings, visually similar pictures, or songs of the same genre.
|
||||
|
||||

|
||||
|
||||
{{< figure src="/docs/encoders.png" caption="This is how vector similarity works" width="70%" >}}
|
||||
|
||||
## Query API
|
||||
|
||||
*Available as of v1.10.0*
|
||||
|
||||
Qdrant provides a single interface for all kinds of search and exploration requests - the `Query API`.
|
||||
Here is a reference list of what kind of queries you can perform with the `Query API` in Qdrant:
|
||||
|
||||
Depending on the `query` parameter, Qdrant might prefer different strategies for the search.
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Nearest Neighbors Search | Vector Similarity Search, also known as k-NN |
|
||||
| Search By Id | Search by an already stored vector - skip embedding model inference |
|
||||
| [Recommendations](../explore/#recommendation-api) | Provide positive and negative examples |
|
||||
| [Discovery Search](../explore/#discovery-api) | Guide the search using context as a one-shot training set |
|
||||
| [Scroll](../points/#scroll-points) | Get all points with optional filtering |
|
||||
| [Order By](../hybrid-queries/#re-ranking-with-stored-values) | Order points by payload key |
|
||||
| [Hybrid Search](../hybrid-queries/#hybrid-search) | Combine multiple queries to get better results |
|
||||
| [Multi-Stage Search](../hybrid-queries/#multi-stage-queries) | Optimize performance for large embeddings |
|
||||
|
||||
|
||||
**Nearest Neighbors Search**
|
||||
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"query": [0.2, 0.1, 0.9, 0.7] // <--- Dense vector
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query=[0.2, 0.1, 0.9, 0.7], # <--- Dense vector
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.query("{collection_name}", {
|
||||
query: [0.2, 0.1, 0.9, 0.7], // <--- Dense vector
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{Condition, Filter, Query, QueryPointsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client
|
||||
.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.query(Query::new_nearest(vec![0.2, 0.1, 0.9, 0.7]))
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import java.util.List;
|
||||
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.QueryPoints;
|
||||
|
||||
QdrantClient client = new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client.queryAsync(QueryPoints.newBuilder()
|
||||
.setCollectionName("{collectionName}")
|
||||
.setQuery(nearest(List.of(0.2f, 0.1f, 0.9f, 0.7f)))
|
||||
.build()).get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
query: new float[] { 0.2f, 0.1f, 0.9f, 0.7f }
|
||||
);
|
||||
```
|
||||
|
||||
**Search By Id**
|
||||
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"query": "43cf51e2-8777-4f52-bc74-c2cbde0c8b04" // <--- point id
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query="43cf51e2-8777-4f52-bc74-c2cbde0c8b04", # <--- point id
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.query("{collection_name}", {
|
||||
query: '43cf51e2-8777-4f52-bc74-c2cbde0c8b04', // <--- point id
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{Condition, Filter, PointId, Query, QueryPointsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client
|
||||
.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.query(Query::new_nearest(PointId::new("43cf51e2-8777-4f52-bc74-c2cbde0c8b04")))
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import java.util.UUID;
|
||||
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.QueryPoints;
|
||||
|
||||
QdrantClient client = new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client.queryAsync(QueryPoints.newBuilder()
|
||||
.setCollectionName("{collectionName}")
|
||||
.setQuery(nearest(UUID.fromString("43cf51e2-8777-4f52-bc74-c2cbde0c8b04")))
|
||||
.build()).get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
query: Guid.Parse("43cf51e2-8777-4f52-bc74-c2cbde0c8b04")
|
||||
);
|
||||
```
|
||||
|
||||
## Metrics
|
||||
|
||||
There are many ways to estimate the similarity of vectors with each other.
|
||||
In Qdrant terms, these ways are called metrics.
|
||||
The choice of metric depends on vectors obtaining and, in particular, on the method of neural network encoder training.
|
||||
The choice of metric depends on the vectors obtained and, in particular, on the neural network encoder training method.
|
||||
|
||||
Qdrant supports these most popular types of metrics:
|
||||
|
||||
@@ -37,22 +195,8 @@ It happens only once for each vector.
|
||||
The second step is the comparison of vectors.
|
||||
In this case, it becomes equivalent to dot production - a very fast operation due to SIMD.
|
||||
|
||||
## Query planning
|
||||
|
||||
Depending on the filter used in the search - there are several possible scenarios for query execution.
|
||||
Qdrant chooses one of the query execution options depending on the available indexes, the complexity of the conditions and the cardinality of the filtering result.
|
||||
This process is called query planning.
|
||||
|
||||
The strategy selection process relies heavily on heuristics and can vary from release to release.
|
||||
However, the general principles are:
|
||||
|
||||
* planning is performed for each segment independently (see [storage](../storage/) for more information about segments)
|
||||
* prefer a full scan if the amount of points is below a threshold
|
||||
* estimate the cardinality of a filtered result before selecting a strategy
|
||||
* retrieve points using payload index (see [indexing](../indexing/)) if cardinality is below threshold
|
||||
* use filterable vector index if the cardinality is above a threshold
|
||||
|
||||
You can adjust the threshold using a [configuration file](https://github.com/qdrant/qdrant/blob/master/config/config.yaml), as well as independently for each collection.
|
||||
Depending on the query configuration, Qdrant might prefer different strategies for the search.
|
||||
Read more about it in the [query planning](#query-planning) section.
|
||||
|
||||
## Search API
|
||||
|
||||
@@ -1486,3 +1630,21 @@ The looked up result will show up under `lookup` in each group.
|
||||
```
|
||||
|
||||
Since the lookup is done by matching directly with the point id, any group id that is not an existing (and valid) point id in the lookup collection will be ignored, and the `lookup` field will be empty.
|
||||
|
||||
|
||||
## Query planning
|
||||
|
||||
Depending on the filter used in the search - there are several possible scenarios for query execution.
|
||||
Qdrant chooses one of the query execution options depending on the available indexes, the complexity of the conditions and the cardinality of the filtering result.
|
||||
This process is called query planning.
|
||||
|
||||
The strategy selection process relies heavily on heuristics and can vary from release to release.
|
||||
However, the general principles are:
|
||||
|
||||
* planning is performed for each segment independently (see [storage](../storage/) for more information about segments)
|
||||
* prefer a full scan if the amount of points is below a threshold
|
||||
* estimate the cardinality of a filtered result before selecting a strategy
|
||||
* retrieve points using payload index (see [indexing](../indexing/)) if cardinality is below threshold
|
||||
* use filterable vector index if the cardinality is above a threshold
|
||||
|
||||
You can adjust the threshold using a [configuration file](https://github.com/qdrant/qdrant/blob/master/config/config.yaml), as well as independently for each collection.
|
||||
|
||||
@@ -15,29 +15,6 @@ This feature can be used to archive data or easily replicate an existing deploym
|
||||
|
||||
For a step-by-step guide on how to use snapshots, see our [tutorial](/documentation/tutorials/create-snapshot/).
|
||||
|
||||
## Store snapshots
|
||||
|
||||
The target directory used to store generated snapshots is controlled through the [configuration](../../guides/configuration/) or using the ENV variable: `QDRANT__STORAGE__SNAPSHOTS_PATH=./snapshots`.
|
||||
|
||||
You can set the snapshots storage directory from the [config.yaml](https://github.com/qdrant/qdrant/blob/master/config/config.yaml) file. If no value is given, default is `./snapshots`.
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
# Specify where you want to store snapshots.
|
||||
snapshots_path: ./snapshots
|
||||
```
|
||||
|
||||
*Available as of v1.3.0*
|
||||
|
||||
While a snapshot is being created, temporary files are by default placed in the configured storage directory.
|
||||
This location may have limited capacity or be on a slow network-attached disk. You may specify a separate location for temporary files:
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
# Where to store temporary files
|
||||
temp_path: /tmp
|
||||
```
|
||||
|
||||
## Create snapshot
|
||||
|
||||
<aside role="status">If you work with a distributed deployment, you have to create snapshots for each node separately. A single snapshot will contain only the data stored on the node on which the snapshot was created.</aside>
|
||||
@@ -527,3 +504,68 @@ For example:
|
||||
```bash
|
||||
./qdrant --storage-snapshot /snapshots/full-snapshot-2022-07-18-11-20-51.snapshot
|
||||
```
|
||||
|
||||
## Storage
|
||||
|
||||
Created, uploaded and recovered snapshots are stored as `.snapshot` files. By
|
||||
default, they're stored on the [local file system](#local-file-system). You may
|
||||
also configure to use an [S3 storage](#s3) service for them.
|
||||
|
||||
### Local file system
|
||||
|
||||
By default, snapshots are stored at `./snapshots` or at `/qdrant/snapshots` when
|
||||
using our Docker image.
|
||||
|
||||
The target directory can be controlled through the [configuration](../../guides/configuration/):
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
# Specify where you want to store snapshots.
|
||||
snapshots_path: ./snapshots
|
||||
```
|
||||
|
||||
Alternatively you may use the environment variable `QDRANT__STORAGE__SNAPSHOTS_PATH=./snapshots`.
|
||||
|
||||
*Available as of v1.3.0*
|
||||
|
||||
While a snapshot is being created, temporary files are placed in the configured
|
||||
storage directory by default. In case of limited capacity or a slow
|
||||
network attached disk, you can specify a separate location for temporary files:
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
# Where to store temporary files
|
||||
temp_path: /tmp
|
||||
```
|
||||
|
||||
### S3
|
||||
|
||||
*Available as of v1.10.0*
|
||||
|
||||
Rather than storing snapshots on the local file system, you may also configure
|
||||
to store snapshots in an S3-compatible storage service. To enable this, you must
|
||||
configure it in the [configuration](../../guides/configuration/) file.
|
||||
|
||||
For example, to configure for AWS S3:
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
snapshots_config:
|
||||
# Use 's3' to store snapshots on S3
|
||||
snapshots_storage: s3
|
||||
|
||||
s3_config:
|
||||
# Bucket name
|
||||
bucket: your_bucket_here
|
||||
|
||||
# Bucket region (e.g. eu-central-1)
|
||||
region: your_bucket_region_here
|
||||
|
||||
# Storage access key
|
||||
# Can be specified either here or in the `AWS_ACCESS_KEY_ID` environment variable.
|
||||
access_key: your_access_key_here
|
||||
|
||||
# Storage secret key
|
||||
# Can be specified either here or in the `AWS_SECRET_ACCESS_KEY` environment variable.
|
||||
secret_key: your_secret_key_here
|
||||
```
|
||||
|
||||
@@ -98,6 +98,15 @@ storage:
|
||||
# Where to store snapshots
|
||||
snapshots_path: ./snapshots
|
||||
|
||||
snapshots_config:
|
||||
# "local" or "s3" - where to store snapshots
|
||||
snapshots_storage: local
|
||||
# s3_config:
|
||||
# bucket: ""
|
||||
# region: ""
|
||||
# access_key: ""
|
||||
# secret_key: ""
|
||||
|
||||
# Where to store temporary files
|
||||
# If null, temporary snapshot are stored in: storage/snapshots_temp/
|
||||
temp_path: null
|
||||
@@ -130,10 +139,17 @@ storage:
|
||||
performance:
|
||||
# Number of parallel threads used for search operations. If 0 - auto selection.
|
||||
max_search_threads: 0
|
||||
# Max total number of threads, which can be used for running optimization processes across all collections.
|
||||
# Note: Each optimization thread will also use `max_indexing_threads` for index building.
|
||||
# So total number of threads used for optimization will be `max_optimization_threads * max_indexing_threads`
|
||||
max_optimization_threads: 1
|
||||
|
||||
# Max number of threads (jobs) for running optimizations across all collections, each thread runs one job.
|
||||
# If 0 - have no limit and choose dynamically to saturate CPU.
|
||||
# Note: each optimization job will also use `max_indexing_threads` threads by itself for index building.
|
||||
max_optimization_threads: 0
|
||||
|
||||
# CPU budget, how many CPUs (threads) to allocate for an optimization job.
|
||||
# If 0 - auto selection, keep 1 or more CPUs unallocated depending on CPU size
|
||||
# If negative - subtract this number of CPUs from the available CPUs.
|
||||
# If positive - use this exact number of CPUs.
|
||||
optimizer_cpu_budget: 0
|
||||
|
||||
# Prevent DDoS of too many concurrent updates in distributed mode.
|
||||
# One external update usually triggers multiple internal updates, which breaks internal
|
||||
@@ -141,6 +157,18 @@ storage:
|
||||
# If null - auto selection.
|
||||
update_rate_limit: null
|
||||
|
||||
# Limit for number of incoming automatic shard transfers per collection on this node, does not affect user-requested transfers.
|
||||
# The same value should be used on all nodes in a cluster.
|
||||
# Default is to allow 1 transfer.
|
||||
# If null - allow unlimited transfers.
|
||||
#incoming_shard_transfers_limit: 1
|
||||
|
||||
# Limit for number of outgoing automatic shard transfers per collection on this node, does not affect user-requested transfers.
|
||||
# The same value should be used on all nodes in a cluster.
|
||||
# Default is to allow 1 transfer.
|
||||
# If null - allow unlimited transfers.
|
||||
#outgoing_shard_transfers_limit: 1
|
||||
|
||||
optimizers:
|
||||
# The minimal fraction of deleted vectors in a segment, required to perform segment optimization
|
||||
deleted_threshold: 0.2
|
||||
@@ -186,33 +214,76 @@ storage:
|
||||
# Interval between forced flushes.
|
||||
flush_interval_sec: 5
|
||||
|
||||
# Max number of threads, which can be used for optimization per collection.
|
||||
# Note: Each optimization thread will also use `max_indexing_threads` for index building.
|
||||
# So total number of threads used for optimization will be `max_optimization_threads * max_indexing_threads`
|
||||
# If `max_optimization_threads = 0`, optimization will be disabled.
|
||||
max_optimization_threads: 1
|
||||
# Max number of threads (jobs) for running optimizations per shard.
|
||||
# Note: each optimization job will also use `max_indexing_threads` threads by itself for index building.
|
||||
# If null - have no limit and choose dynamically to saturate CPU.
|
||||
# If 0 - no optimization threads, optimizations will be disabled.
|
||||
max_optimization_threads: null
|
||||
|
||||
# This section has the same options as 'optimizers' above. All values specified here will overwrite the collections
|
||||
# optimizers configs regardless of the config above and the options specified at collection creation.
|
||||
#optimizers_overwrite:
|
||||
# deleted_threshold: 0.2
|
||||
# vacuum_min_vector_number: 1000
|
||||
# default_segment_number: 0
|
||||
# max_segment_size_kb: null
|
||||
# memmap_threshold_kb: null
|
||||
# indexing_threshold_kb: 20000
|
||||
# flush_interval_sec: 5
|
||||
# max_optimization_threads: null
|
||||
|
||||
# Default parameters of HNSW Index. Could be overridden for each collection or named vector individually
|
||||
hnsw_index:
|
||||
# Number of edges per node in the index graph. Larger the value - more accurate the search, more space required.
|
||||
m: 16
|
||||
|
||||
# Number of neighbours to consider during the index building. Larger the value - more accurate the search, more time required to build index.
|
||||
ef_construct: 100
|
||||
|
||||
# Minimal size (in KiloBytes) of vectors for additional payload-based indexing.
|
||||
# If payload chunk is smaller than `full_scan_threshold_kb` additional indexing won't be used -
|
||||
# in this case full-scan search should be preferred by query planner and additional indexing is not required.
|
||||
# Note: 1Kb = 1 vector of size 256
|
||||
full_scan_threshold_kb: 10000
|
||||
# Number of parallel threads used for background index building. If 0 - auto selection.
|
||||
|
||||
# Number of parallel threads used for background index building.
|
||||
# If 0 - automatically select.
|
||||
# Best to keep between 8 and 16 to prevent likelihood of building broken/inefficient HNSW graphs.
|
||||
# On small CPUs, less threads are used.
|
||||
max_indexing_threads: 0
|
||||
|
||||
# Store HNSW index on disk. If set to false, index will be stored in RAM. Default: false
|
||||
on_disk: false
|
||||
|
||||
# Custom M param for hnsw graph built for payload index. If not set, default M will be used.
|
||||
payload_m: null
|
||||
|
||||
# Default shard transfer method to use if none is defined.
|
||||
# If null - don't have a shard transfer preference, choose automatically.
|
||||
# If stream_records, snapshot or wal_delta - prefer this specific method.
|
||||
# More info: https://qdrant.tech/documentation/guides/distributed_deployment/#shard-transfer-method
|
||||
shard_transfer_method: null
|
||||
|
||||
# Default parameters for collections
|
||||
collection:
|
||||
# Number of replicas of each shard that network tries to maintain
|
||||
replication_factor: 1
|
||||
|
||||
# How many replicas should apply the operation for us to consider it successful
|
||||
write_consistency_factor: 1
|
||||
|
||||
# Default parameters for vectors.
|
||||
vectors:
|
||||
# Whether vectors should be stored in memory or on disk.
|
||||
on_disk: null
|
||||
|
||||
# shard_number_per_node: 1
|
||||
|
||||
# Default quantization configuration.
|
||||
# More info: https://qdrant.tech/documentation/guides/quantization
|
||||
quantization: null
|
||||
|
||||
service:
|
||||
|
||||
# Maximum size of POST data in a single request in megabytes
|
||||
max_request_size_mb: 32
|
||||
|
||||
@@ -265,6 +336,12 @@ service:
|
||||
# Uncomment to enable.
|
||||
# read_only_api_key: your_secret_read_only_api_key_here
|
||||
|
||||
# Uncomment to enable JWT Role Based Access Control (RBAC).
|
||||
# If enabled, you can generate JWT tokens with fine-grained rules for access control.
|
||||
# Use generated token instead of API key.
|
||||
#
|
||||
# jwt_rbac: true
|
||||
|
||||
cluster:
|
||||
# Use `enabled: true` to run Qdrant in distributed deployment mode
|
||||
enabled: false
|
||||
|
||||
|
After Width: | Height: | Size: 27 KiB |
|
After Width: | Height: | Size: 1.9 MiB |
|
After Width: | Height: | Size: 30 KiB |
|
After Width: | Height: | Size: 24 KiB |
|
After Width: | Height: | Size: 213 KiB |
|
After Width: | Height: | Size: 77 KiB |
|
After Width: | Height: | Size: 60 KiB |
|
After Width: | Height: | Size: 529 KiB |
|
After Width: | Height: | Size: 337 KiB |
|
After Width: | Height: | Size: 1.0 MiB |
|
After Width: | Height: | Size: 42 KiB |
|
After Width: | Height: | Size: 80 KiB |
|
After Width: | Height: | Size: 35 KiB |
@@ -107,11 +107,10 @@ $table-striped-bg: $neutral-94;
|
||||
|
||||
$font-family-sans-serif: 'Satoshi', 'Helvetica Neue', 'Noto Sans', 'Liberation Sans', Arial, sans-serif,
|
||||
'Apple Color Emoji', 'Segoe UI Emoji', 'Segoe UI Symbol', 'Noto Color Emoji';
|
||||
$font-family-monospace: 'Source Code Pro', SFMono-Regular, Menlo, Monaco, Consolas, 'Liberation Mono', 'Courier New',
|
||||
monospace;
|
||||
$font-family-monospace: 'Source Code Pro', 'Overpass Mono', SFMono-Regular, 'Droid Sans Mono', Menlo, Monaco, Consolas, 'Lucida Console', 'Monaco', monospace;
|
||||
$font-family-base: 'Satoshi', 'Helvetica Neue', 'Noto Sans', 'Liberation Sans', Arial, sans-serif, 'Apple Color Emoji',
|
||||
'Segoe UI Emoji', 'Segoe UI Symbol', 'Noto Color Emoji';
|
||||
$font-family-code: 'Source Code Pro', SFMono-Regular, Menlo, Monaco, Consolas, 'Liberation Mono', 'Courier New', monospace;
|
||||
$font-family-code: 'Source Code Pro', 'Overpass Mono', SFMono-Regular, 'Droid Sans Mono', Menlo, Monaco, Consolas, 'Lucida Console', 'Monaco', monospace;
|
||||
|
||||
$font-size-xs: rem(12);
|
||||
$font-size-s: rem(14);
|
||||
|
||||
@@ -1,6 +1,8 @@
|
||||
@use 'sass:math';
|
||||
@use 'helpers/functions' as *;
|
||||
@import 'satoshi-font';
|
||||
@import 'fonts/satoshi-font';
|
||||
@import 'fonts/source-code-pro';
|
||||
@import 'fonts/overpass-mono';
|
||||
@import 'bootstrap-custom';
|
||||
|
||||
html {
|
||||
|
||||
@@ -0,0 +1,15 @@
|
||||
@font-face {
|
||||
font-family: 'Overpass Mono';
|
||||
font-style: normal;
|
||||
font-weight: 300 700;
|
||||
font-display: swap;
|
||||
src: url('/fonts/overpass_mono/overpass-mono-variable-font_wght.ttf') format('truetype');
|
||||
}
|
||||
|
||||
@font-face {
|
||||
font-family: 'Overpass Mono';
|
||||
font-style: italic;
|
||||
font-weight: 300 700;
|
||||
font-display: swap;
|
||||
src: url('/fonts/overpass_mono/overpass-mono-variable-font_wght.ttf') format('truetype');
|
||||
}
|
||||
@@ -0,0 +1,15 @@
|
||||
@font-face {
|
||||
font-family: 'Source Code Pro';
|
||||
font-style: normal;
|
||||
font-weight: 200 900;
|
||||
font-display: swap;
|
||||
src: url('/fonts/source_code_pro/source-code-pro-variablefont_wght.ttf') format('truetype');
|
||||
}
|
||||
|
||||
@font-face {
|
||||
font-family: 'Source Code Pro';
|
||||
font-style: italic;
|
||||
font-weight: 200 900;
|
||||
font-display: swap;
|
||||
src: url('/fonts/source_code_pro/source-code-pro-italic-variablefont_wght.ttf') format('truetype');
|
||||
}
|
||||
@@ -99,6 +99,7 @@ article code:not([class]) {
|
||||
.chroma .line {
|
||||
color: $neutral-94;
|
||||
font-family: $font-family-monospace;
|
||||
font-optical-sizing: auto;
|
||||
font-size: 16px;
|
||||
font-style: normal;
|
||||
font-weight: 400;
|
||||
@@ -338,7 +339,6 @@ article code:not([class]) {
|
||||
/* CommentSingle */
|
||||
.chroma .c1 {
|
||||
color: $neutral-80;
|
||||
font-style: italic;
|
||||
}
|
||||
/* CommentSpecial */
|
||||
.chroma .cs {
|
||||
|
||||
@@ -32,10 +32,4 @@
|
||||
|
||||
{{ partial "seo_schema" . }}
|
||||
|
||||
|
||||
<link rel="preconnect" href="https://fonts.googleapis.com">
|
||||
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
|
||||
<link href="https://fonts.googleapis.com/css2?family=Source+Code+Pro:ital,wght@0,300..600;1,300..600&display=swap" rel="stylesheet">
|
||||
{{ partial "js-head" . }}
|
||||
|
||||
</head>
|
||||
|
||||
@@ -0,0 +1,93 @@
|
||||
Copyright 2021 The Overpass Project Authors (https://github.com/RedHatOfficial/Overpass)
|
||||
|
||||
This Font Software is licensed under the SIL Open Font License, Version 1.1.
|
||||
This license is copied below, and is also available with a FAQ at:
|
||||
https://openfontlicense.org
|
||||
|
||||
|
||||
-----------------------------------------------------------
|
||||
SIL OPEN FONT LICENSE Version 1.1 - 26 February 2007
|
||||
-----------------------------------------------------------
|
||||
|
||||
PREAMBLE
|
||||
The goals of the Open Font License (OFL) are to stimulate worldwide
|
||||
development of collaborative font projects, to support the font creation
|
||||
efforts of academic and linguistic communities, and to provide a free and
|
||||
open framework in which fonts may be shared and improved in partnership
|
||||
with others.
|
||||
|
||||
The OFL allows the licensed fonts to be used, studied, modified and
|
||||
redistributed freely as long as they are not sold by themselves. The
|
||||
fonts, including any derivative works, can be bundled, embedded,
|
||||
redistributed and/or sold with any software provided that any reserved
|
||||
names are not used by derivative works. The fonts and derivatives,
|
||||
however, cannot be released under any other type of license. The
|
||||
requirement for fonts to remain under this license does not apply
|
||||
to any document created using the fonts or their derivatives.
|
||||
|
||||
DEFINITIONS
|
||||
"Font Software" refers to the set of files released by the Copyright
|
||||
Holder(s) under this license and clearly marked as such. This may
|
||||
include source files, build scripts and documentation.
|
||||
|
||||
"Reserved Font Name" refers to any names specified as such after the
|
||||
copyright statement(s).
|
||||
|
||||
"Original Version" refers to the collection of Font Software components as
|
||||
distributed by the Copyright Holder(s).
|
||||
|
||||
"Modified Version" refers to any derivative made by adding to, deleting,
|
||||
or substituting -- in part or in whole -- any of the components of the
|
||||
Original Version, by changing formats or by porting the Font Software to a
|
||||
new environment.
|
||||
|
||||
"Author" refers to any designer, engineer, programmer, technical
|
||||
writer or other person who contributed to the Font Software.
|
||||
|
||||
PERMISSION & CONDITIONS
|
||||
Permission is hereby granted, free of charge, to any person obtaining
|
||||
a copy of the Font Software, to use, study, copy, merge, embed, modify,
|
||||
redistribute, and sell modified and unmodified copies of the Font
|
||||
Software, subject to the following conditions:
|
||||
|
||||
1) Neither the Font Software nor any of its individual components,
|
||||
in Original or Modified Versions, may be sold by itself.
|
||||
|
||||
2) Original or Modified Versions of the Font Software may be bundled,
|
||||
redistributed and/or sold with any software, provided that each copy
|
||||
contains the above copyright notice and this license. These can be
|
||||
included either as stand-alone text files, human-readable headers or
|
||||
in the appropriate machine-readable metadata fields within text or
|
||||
binary files as long as those fields can be easily viewed by the user.
|
||||
|
||||
3) No Modified Version of the Font Software may use the Reserved Font
|
||||
Name(s) unless explicit written permission is granted by the corresponding
|
||||
Copyright Holder. This restriction only applies to the primary font name as
|
||||
presented to the users.
|
||||
|
||||
4) The name(s) of the Copyright Holder(s) or the Author(s) of the Font
|
||||
Software shall not be used to promote, endorse or advertise any
|
||||
Modified Version, except to acknowledge the contribution(s) of the
|
||||
Copyright Holder(s) and the Author(s) or with their explicit written
|
||||
permission.
|
||||
|
||||
5) The Font Software, modified or unmodified, in part or in whole,
|
||||
must be distributed entirely under this license, and must not be
|
||||
distributed under any other license. The requirement for fonts to
|
||||
remain under this license does not apply to any document created
|
||||
using the Font Software.
|
||||
|
||||
TERMINATION
|
||||
This license becomes null and void if any of the above conditions are
|
||||
not met.
|
||||
|
||||
DISCLAIMER
|
||||
THE FONT SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
|
||||
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO ANY WARRANTIES OF
|
||||
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT
|
||||
OF COPYRIGHT, PATENT, TRADEMARK, OR OTHER RIGHT. IN NO EVENT SHALL THE
|
||||
COPYRIGHT HOLDER BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY,
|
||||
INCLUDING ANY GENERAL, SPECIAL, INDIRECT, INCIDENTAL, OR CONSEQUENTIAL
|
||||
DAMAGES, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING
|
||||
FROM, OUT OF THE USE OR INABILITY TO USE THE FONT SOFTWARE OR FROM
|
||||
OTHER DEALINGS IN THE FONT SOFTWARE.
|
||||
@@ -0,0 +1,93 @@
|
||||
Copyright 2010, 2012 Adobe Systems Incorporated (http://www.adobe.com/), with Reserved Font Name 'Source'. All Rights Reserved. Source is a trademark of Adobe Systems Incorporated in the United States and/or other countries.
|
||||
|
||||
This Font Software is licensed under the SIL Open Font License, Version 1.1.
|
||||
This license is copied below, and is also available with a FAQ at:
|
||||
https://openfontlicense.org
|
||||
|
||||
|
||||
-----------------------------------------------------------
|
||||
SIL OPEN FONT LICENSE Version 1.1 - 26 February 2007
|
||||
-----------------------------------------------------------
|
||||
|
||||
PREAMBLE
|
||||
The goals of the Open Font License (OFL) are to stimulate worldwide
|
||||
development of collaborative font projects, to support the font creation
|
||||
efforts of academic and linguistic communities, and to provide a free and
|
||||
open framework in which fonts may be shared and improved in partnership
|
||||
with others.
|
||||
|
||||
The OFL allows the licensed fonts to be used, studied, modified and
|
||||
redistributed freely as long as they are not sold by themselves. The
|
||||
fonts, including any derivative works, can be bundled, embedded,
|
||||
redistributed and/or sold with any software provided that any reserved
|
||||
names are not used by derivative works. The fonts and derivatives,
|
||||
however, cannot be released under any other type of license. The
|
||||
requirement for fonts to remain under this license does not apply
|
||||
to any document created using the fonts or their derivatives.
|
||||
|
||||
DEFINITIONS
|
||||
"Font Software" refers to the set of files released by the Copyright
|
||||
Holder(s) under this license and clearly marked as such. This may
|
||||
include source files, build scripts and documentation.
|
||||
|
||||
"Reserved Font Name" refers to any names specified as such after the
|
||||
copyright statement(s).
|
||||
|
||||
"Original Version" refers to the collection of Font Software components as
|
||||
distributed by the Copyright Holder(s).
|
||||
|
||||
"Modified Version" refers to any derivative made by adding to, deleting,
|
||||
or substituting -- in part or in whole -- any of the components of the
|
||||
Original Version, by changing formats or by porting the Font Software to a
|
||||
new environment.
|
||||
|
||||
"Author" refers to any designer, engineer, programmer, technical
|
||||
writer or other person who contributed to the Font Software.
|
||||
|
||||
PERMISSION & CONDITIONS
|
||||
Permission is hereby granted, free of charge, to any person obtaining
|
||||
a copy of the Font Software, to use, study, copy, merge, embed, modify,
|
||||
redistribute, and sell modified and unmodified copies of the Font
|
||||
Software, subject to the following conditions:
|
||||
|
||||
1) Neither the Font Software nor any of its individual components,
|
||||
in Original or Modified Versions, may be sold by itself.
|
||||
|
||||
2) Original or Modified Versions of the Font Software may be bundled,
|
||||
redistributed and/or sold with any software, provided that each copy
|
||||
contains the above copyright notice and this license. These can be
|
||||
included either as stand-alone text files, human-readable headers or
|
||||
in the appropriate machine-readable metadata fields within text or
|
||||
binary files as long as those fields can be easily viewed by the user.
|
||||
|
||||
3) No Modified Version of the Font Software may use the Reserved Font
|
||||
Name(s) unless explicit written permission is granted by the corresponding
|
||||
Copyright Holder. This restriction only applies to the primary font name as
|
||||
presented to the users.
|
||||
|
||||
4) The name(s) of the Copyright Holder(s) or the Author(s) of the Font
|
||||
Software shall not be used to promote, endorse or advertise any
|
||||
Modified Version, except to acknowledge the contribution(s) of the
|
||||
Copyright Holder(s) and the Author(s) or with their explicit written
|
||||
permission.
|
||||
|
||||
5) The Font Software, modified or unmodified, in part or in whole,
|
||||
must be distributed entirely under this license, and must not be
|
||||
distributed under any other license. The requirement for fonts to
|
||||
remain under this license does not apply to any document created
|
||||
using the Font Software.
|
||||
|
||||
TERMINATION
|
||||
This license becomes null and void if any of the above conditions are
|
||||
not met.
|
||||
|
||||
DISCLAIMER
|
||||
THE FONT SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
|
||||
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO ANY WARRANTIES OF
|
||||
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT
|
||||
OF COPYRIGHT, PATENT, TRADEMARK, OR OTHER RIGHT. IN NO EVENT SHALL THE
|
||||
COPYRIGHT HOLDER BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY,
|
||||
INCLUDING ANY GENERAL, SPECIAL, INDIRECT, INCIDENTAL, OR CONSEQUENTIAL
|
||||
DAMAGES, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING
|
||||
FROM, OUT OF THE USE OR INABILITY TO USE THE FONT SOFTWARE OR FROM
|
||||
OTHER DEALINGS IN THE FONT SOFTWARE.
|
||||