mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-28 23:48:31 +02:00
Merge branch 'master' into pincone-vs-q
This commit is contained in:
@@ -0,0 +1,362 @@
|
||||
---
|
||||
title: "BM42: New Baseline for Hybrid Search"
|
||||
short_description: "Introducing next evolutionary step in lexical search."
|
||||
description: "Introducing BM42 - a new sparse embedding approach, which combines the benefits of exact keyword search with the intelligence of transformers."
|
||||
social_preview_image: /articles_data/bm42/social-preview.jpg
|
||||
preview_dir: /articles_data/bm42/preview
|
||||
weight: -140
|
||||
author: Andrey Vasnetsov
|
||||
date: 2024-07-01T12:00:00+03:00
|
||||
draft: false
|
||||
keywords:
|
||||
- hybrid search
|
||||
- sparse embeddings
|
||||
- bm25
|
||||
---
|
||||
|
||||
For the last 40 years, BM25 has served as the standard for search engines.
|
||||
It is a simple yet powerful algorithm that has been used by many search engines, including Google, Bing, and Yahoo.
|
||||
|
||||
Though it seemed that the advent of vector search would diminish its influence, it did so only partially.
|
||||
The current state-of-the-art approach to retrieval nowadays tries to incorporate BM25 along with embeddings into a hybrid search system.
|
||||
|
||||
However, the use case of text retrieval has significantly shifted since the introduction of RAG.
|
||||
Many assumptions upon which BM25 was built are no longer valid.
|
||||
|
||||
For example, the typical length of documents and queries vary significantly between traditional web search and modern RAG systems.
|
||||
|
||||
In this article, we will recap what made BM25 relevant for so long and why alternatives have struggled to replace it. Finally, we will discuss BM42, as the next step in the evolution of lexical search.
|
||||
|
||||
## Why has BM25 stayed relevant for so long?
|
||||
|
||||
To understand why, we need to analyze its components.
|
||||
|
||||
The famous BM25 formula is defined as:
|
||||
|
||||
$$
|
||||
\text{score}(D,Q) = \sum_{i=1}^{N} \text{IDF}(q_i) \times \frac{f(q_i, D) \cdot (k_1 + 1)}{f(q_i, D) + k_1 \cdot \left(1 - b + b \cdot \frac{|D|}{\text{avgdl}}\right)}
|
||||
$$
|
||||
|
||||
Let's simplify this to gain a better understanding.
|
||||
|
||||
- The $score(D, Q)$ - means that we compute the score for each pair of document $D$ and query $Q$.
|
||||
|
||||
- The $\sum_{i=1}^{N}$ - means that each of $N$ terms in the query contribute to the final score as a part of the sum.
|
||||
|
||||
- The $\text{IDF}(q_i)$ - is the inverse document frequency. The more rare the term $q_i$ is, the more it contributes to the score. A simplified formula for this is:
|
||||
|
||||
$$
|
||||
\text{IDF}(q_i) = \frac{\text{Number of documents}}{\text{Number of documents with } q_i}
|
||||
$$
|
||||
|
||||
It is fair to say that the `IDF` is the most important part of the BM25 formula.
|
||||
`IDF` selects the most important terms in the query relative to the specific document collection.
|
||||
So intuitively, we can interpret the `IDF` as **term importance within the corpora**.
|
||||
|
||||
That explains why BM25 is so good at handling queries, which dense embeddings consider out-of-domain.
|
||||
|
||||
The last component of the formula can be intuitively interpreted as **term importance within the document**.
|
||||
This might look a bit complicated, so let's break it down.
|
||||
|
||||
$$
|
||||
\text{Term importance in document }(q_i) = \color{red}\frac{f(q_i, D)\color{black} \cdot \color{blue}(k_1 + 1) \color{black} }{\color{red}f(q_i, D)\color{black} + \color{blue}k_1\color{black} \cdot \left(1 - \color{blue}b\color{black} + \color{blue}b\color{black} \cdot \frac{|D|}{\text{avgdl}}\right)}
|
||||
$$
|
||||
|
||||
- The $\color{red}f(q_i, D)\color{black}$ - is the frequency of the term $q_i$ in the document $D$. Or in other words, the number of times the term $q_i$ appears in the document $D$.
|
||||
- The $\color{blue}k_1\color{black}$ and $\color{blue}b\color{black}$ are the hyperparameters of the BM25 formula. In most implementations, they are constants set to $k_1=1.5$ and $b=0.75$. Those constants define relative implications of the term frequency and the document length in the formula.
|
||||
- The $\frac{|D|}{\text{avgdl}}$ - is the relative length of the document $D$ compared to the average document length in the corpora. The intuition befind this part is following: if the token is found in the smaller document, it is more likely that this token is important for this document.
|
||||
|
||||
#### Will BM25 term importance in the document work for RAG?
|
||||
|
||||
As we can see, the *term importance in the document* heavily depends on the statistics within the document. Moreover, statistics works well if the document is long enough.
|
||||
Therefore, it is suitable for searching webpages, books, articles, etc.
|
||||
|
||||
However, would it work as well for modern search applications, such as RAG? Let's see.
|
||||
|
||||
The typical length of a document in RAG is much shorter than that of web search. In fact, even if we are working with webpages and articles, we would prefer to split them into chunks so that
|
||||
a) Dense models can handle them and
|
||||
b) We can pinpoint the exact part of the document which is relevant to the query
|
||||
|
||||
As a result, the document size in RAG is small and fixed.
|
||||
|
||||
That effectively renders the term importance in the document part of the BM25 formula useless.
|
||||
The term frequency in the document is always 0 or 1, and the relative length of the document is always 1.
|
||||
|
||||
So, the only part of the BM25 formula that is still relevant for RAG is `IDF`. Let's see how we can leverage it.
|
||||
|
||||
## Why SPLADE is not always the answer
|
||||
|
||||
Before discussing our new approach, let's examine the current state-of-the-art alternative to BM25 - SPLADE.
|
||||
|
||||
The idea behind SPLADE is interesting—what if we let a smart, end-to-end trained model generate a bag-of-words representation of the text for us?
|
||||
It will assign all the weights to the tokens, so we won't need to bother with statistics and hyperparameters.
|
||||
The documents are then represented as a sparse embedding, where each token is represented as an element of the sparse vector.
|
||||
|
||||
And it works in academic benchmarks. Many papers report that SPLADE outperforms BM25 in terms of retrieval quality.
|
||||
This performance, however, comes at a cost.
|
||||
|
||||
* **Inappropriate Tokenizer**: To incorporate transformers for this task, SPLADE models require using a standard transformer tokenizer. These tokenizers are not designed for retrieval tasks. For example, if the word is not in the (quite limited) vocabulary, it will be either split into subwords or replaced with a `[UNK]` token. This behavior works well for language modeling but is completely destructive for retrieval tasks.
|
||||
|
||||
* **Expensive Token Expansion**: In order to compensate the tokenization issues, SPLADE uses *token expansion* technique. This means that we generate a set of similar tokens for each token in the query. There are a few problems with this approach:
|
||||
- It is computationally and memory expensive. We need to generate more values for each token in the document, which increases both the storage size and retrieval time.
|
||||
- It is not always clear where to stop with the token expansion. The more tokens we generate, the more likely we are to get the relevant one. But simultaneously, the more tokens we generate, the more likely we are to get irrelevant results.
|
||||
- Token expansion dilutes the interpretability of the search. We can't say which tokens were used in the document and which were generated by the token expansion.
|
||||
|
||||
* **Domain and Language Dependency**: SPLADE models are trained on specific corpora. This means that they are not always generalizable to new or rare domains. As they don't use any statistics from the corpora, they cannot adapt to the new domain without fine-tuning.
|
||||
|
||||
* **Inference Time**: Additionally, currently available SPLADE models are quite big and slow. They usually require a GPU to make the inference in a reasonable time.
|
||||
|
||||
At Qdrant, we acknowledge the aforementioned problems and are looking for a solution.
|
||||
Our idea was to combine the best of both worlds - the simplicity and interpretability of BM25 and the intelligence of transformers while avoiding the pitfalls of SPLADE.
|
||||
|
||||
And here is what we came up with.
|
||||
|
||||
## The best of both worlds
|
||||
|
||||
As previously mentioned, `IDF` is the most important part of the BM25 formula. In fact it is so important, that we decided to build its calculation into the Qdrant engine itself.
|
||||
Check out our latest [release notes](https://github.com/qdrant/qdrant/releases/tag/v1.10.0). This type of separation allows streaming updates of the sparse embeddings while keeping the `IDF` calculation up-to-date.
|
||||
|
||||
As for the second part of the formula, *the term importance within the document* needs to be rethought.
|
||||
|
||||
Since we can't rely on the statistics within the document, we can try to use the semantics of the document instead.
|
||||
And semantics is what transformers are good at. Therefore, we only need to solve two problems:
|
||||
|
||||
- How does one extract the importance information from the transformer?
|
||||
- How can tokenization issues be avoided?
|
||||
|
||||
|
||||
### Attention is all you need
|
||||
|
||||
Transformer models, even those used to generate embeddings, generate a bunch of different outputs.
|
||||
Some of those outputs are used to generate embeddings.
|
||||
|
||||
Others are used to solve other kinds of tasks, such as classification, text generation, etc.
|
||||
|
||||
The one particularly interesting output for us is the attention matrix.
|
||||
|
||||
{{< figure src="/articles_data/bm42/attention-matrix.png" alt="Attention matrix" caption="Attention matrix" width="60%" >}}
|
||||
|
||||
The attention matrix is a square matrix, where each row and column corresponds to the token in the input sequence.
|
||||
It represents the importance of each token in the input sequence for each other.
|
||||
|
||||
The classical transformer models are trained to predict masked tokens in the context, so the attention weights define which context tokens influence the masked token most.
|
||||
|
||||
Apart from regular text tokens, the transformer model also has a special token called `[CLS]`. This token represents the whole sequence in the classification tasks, which is exactly what we need.
|
||||
|
||||
By looking at the attention row for the `[CLS]` token, we can get the importance of each token in the document for the whole document.
|
||||
|
||||
|
||||
```python
|
||||
sentences = "Hello, World - is the starting point in most programming languages"
|
||||
|
||||
features = transformer.tokenize(sentences)
|
||||
|
||||
# ...
|
||||
|
||||
attentions = transformer.auto_model(**features, output_attentions=True).attentions
|
||||
|
||||
weights = torch.mean(attentions[-1][0,:,0], axis=0)
|
||||
# ▲ ▲ ▲ ▲
|
||||
# │ │ │ └─── [CLS] token is the first one
|
||||
# │ │ └─────── First item of the batch
|
||||
# │ └────────── Last transformer layer
|
||||
# └────────────────────────── Averate all 6 attention heads
|
||||
|
||||
for weight, token in zip(weights, tokens):
|
||||
print(f"{token}: {weight}")
|
||||
|
||||
# [CLS] : 0.434 // Filter out the [CLS] token
|
||||
# hello : 0.039
|
||||
# , : 0.039
|
||||
# world : 0.107 // <-- The most important token
|
||||
# - : 0.033
|
||||
# is : 0.024
|
||||
# the : 0.031
|
||||
# starting : 0.054
|
||||
# point : 0.028
|
||||
# in : 0.018
|
||||
# most : 0.016
|
||||
# programming : 0.060 // <-- The third most important token
|
||||
# languages : 0.062 // <-- The second most important token
|
||||
# [SEP] : 0.047 // Filter out the [SEP] token
|
||||
|
||||
```
|
||||
|
||||
|
||||
The resulting formula for the BM42 score would look like this:
|
||||
|
||||
$$
|
||||
\text{score}(D,Q) = \sum_{i=1}^{N} \text{IDF}(q_i) \times \text{Attention}(\text{CLS}, q_i)
|
||||
$$
|
||||
|
||||
|
||||
Note that classical transformers have multiple attention heads, so we can get multiple importance vectors for the same document. The simplest way to combine them is to simply average them.
|
||||
|
||||
These averaged attention vectors make up the importance information we were looking for.
|
||||
The best part is, one can get them from any transformer model, without any additional training.
|
||||
Therefore, BM42 can support any natural language as long as there is a transformer model for it.
|
||||
|
||||
In our implementation, we use the `sentence-transformers/all-MiniLM-L6-v2` model, which gives a huge boost in the inference speed compared to the SPLADE models. In practice, any transformer model can be used.
|
||||
It doesn't require any additional training, and can be easily adapted to work as BM42 backend.
|
||||
|
||||
|
||||
### WordPiece retokenization
|
||||
|
||||
The final piece of the puzzle we need to solve is the tokenization issue. In order to get attention vectors, we need to use native transformer tokenization.
|
||||
But this tokenization is not suitable for the retrieval tasks. What can we do about it?
|
||||
|
||||
Actually, the solution we came up with is quite simple. We reverse the tokenization process after we get the attention vectors.
|
||||
|
||||
Transformers use [WordPiece](https://huggingface.co/learn/nlp-course/en/chapter6/6) tokenization.
|
||||
In case it sees the word, which is not in the vocabulary, it splits it into subwords.
|
||||
|
||||
Here is how that looks:
|
||||
|
||||
```text
|
||||
"unbelievable" -> ["un", "##believ", "##able"]
|
||||
```
|
||||
|
||||
What can merge the subwords back into the words. Luckily, the subwords are marked with the `##` prefix, so we can easily detect them.
|
||||
Since the attention weights are normalized, we can simply sum the attention weights of the subwords to get the attention weight of the word.
|
||||
|
||||
After that, we can apply the same traditional NLP techniques, as
|
||||
|
||||
- Removing of the stop-words
|
||||
- Removing of the punctuation
|
||||
- Lemmatization
|
||||
|
||||
In this way, we can significantly reduce the number of tokens, and therefore minimize the memory footprint of the sparse embeddings. We won't simultaneously compromise the ability to match (almost) exact tokens.
|
||||
|
||||
## Practical examples
|
||||
|
||||
|
||||
| Trait | BM25 | SPLADE | BM42 |
|
||||
|-------------------------|--------------|--------------|--------------|
|
||||
| Interpretability | High ✅ | Ok 🆗 | High ✅ |
|
||||
| Document Inference speed| Very high ✅ | Slow 🐌 | High ✅ |
|
||||
| Query Inference speed | Very high ✅ | Slow 🐌 | Very high ✅ |
|
||||
| Memory footprint | Low ✅ | High ❌ | Low ✅ |
|
||||
| In-domain accuracy | Ok 🆗 | High ✅ | High ✅ |
|
||||
| Out-of-domain accuracy | Ok 🆗 | Low ❌ | Ok 🆗 |
|
||||
| Small documents accuracy| Low ❌ | High ✅ | High ✅ |
|
||||
| Large documents accuracy| High ✅ | Low ❌ | Ok 🆗 |
|
||||
| Unknown tokens handling | Yes ✅ | Bad ❌ | Yes ✅ |
|
||||
| Multi-lingual support | Yes ✅ | No ❌ | Yes ✅ |
|
||||
| Best Match | Yes ✅ | No ❌ | Yes ✅ |
|
||||
|
||||
|
||||
Starting from Qdrant v1.10.0, BM42 can be used in Qdrant via FastEmbed inference.
|
||||
|
||||
Let's see how you can setup a collection for hybrid search with BM42 and [jina.ai](https://jina.ai/embeddings/) dense embeddings.
|
||||
|
||||
```http
|
||||
PUT collections/my-hybrid-collection
|
||||
{
|
||||
"vectors": {
|
||||
"jina": {
|
||||
"size": 768,
|
||||
"distance": "Cosine"
|
||||
}
|
||||
},
|
||||
"sparse_vectors": {
|
||||
"bm42": {
|
||||
"modifier": "idf" // <--- This parameter enables the IDF calculation
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient()
|
||||
|
||||
client.create_collection(
|
||||
collection_name="my-hybrid-collection",
|
||||
vectors_config={
|
||||
"jina": models.VectorParams(
|
||||
size=768,
|
||||
distance=models.Distance.COSINE,
|
||||
)
|
||||
},
|
||||
sparse_vectors_config={
|
||||
"bm42": models.SparseVectorParams(
|
||||
modifier=models.Modifier.IDF,
|
||||
)
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
The search query will retrieve the documents with both dense and sparse embeddings and combine the scores
|
||||
using the Reciprocal Rank Fusion (RRF) algorithm.
|
||||
|
||||
```python
|
||||
from fastembed import SparseTextEmbedding, TextEmbedding
|
||||
|
||||
query_text = "best programming language for beginners?"
|
||||
|
||||
model_bm42 = SparseTextEmbedding(model_name="Qdrant/bm42-all-minilm-l6-v2-attentions")
|
||||
model_jina = TextEmbedding(model_name="jinaai/jina-embeddings-v2-base-en")
|
||||
|
||||
sparse_embedding = list(embedding_model.embed_query(query_text))[0]
|
||||
dense_embedding = list(model_jina.embed_query(query_text))[0]
|
||||
|
||||
client.query_points(
|
||||
collection_name="my-hybrid-collection",
|
||||
prefetch=[
|
||||
models.Prefetch(query=sparse_embedding, using="bm42", limit=10),
|
||||
models.Prefetch(query=dense_embedding, using="jina", limit=10),
|
||||
],
|
||||
query=models.FusionQuery(fusion=models.Fusion.RRF), # <--- Combine the scores
|
||||
limit=10
|
||||
)
|
||||
|
||||
```
|
||||
|
||||
### Benchmarks
|
||||
|
||||
To prove the point further we have conducted some benchmarks to highlight the cases where BM42 outperforms BM25.
|
||||
Please note, that we didn't intend to make an exhaustive evaluation, as we are presenting a new approach, not a new model.
|
||||
|
||||
For out experiments we choose [quora](https://huggingface.co/datasets/BeIR/quora) dataset, as a good representative of the Question-Answering task.
|
||||
The typical example of the dataset is the following:
|
||||
|
||||
```text
|
||||
{"_id": "109", "text": "How GST affects the CAs and tax officers?"}
|
||||
{"_id": "110", "text": "Why can't I do my homework?"}
|
||||
{"_id": "111", "text": "How difficult is it get into RSI?"}
|
||||
```
|
||||
|
||||
As you can see, it has pretty short texts, there are not much of the statistics to rely on.
|
||||
|
||||
After encoding with BM42, the average vector size is only **5.6 elements per document**.
|
||||
|
||||
With `datatype: uint8` available in Qdrant, the total size of the sparse vector index is about **13Mb** for ~530k documents.
|
||||
|
||||
|
||||
| | BM25 | BM42 |
|
||||
|---------------|------|----------|
|
||||
|Precision @ 10 | 0.45 | **0.49** |
|
||||
|
||||
|
||||
Please note, that both BM25 and BM42 won't work well on their own in a production environment.
|
||||
Best results are achieved with a combination of sparse and dense embeddings in a hybrid approach.
|
||||
In this scenario, the two models are complementary to each other.
|
||||
The sparse model is responsible for exact token matching, while the dense model is responsible for semantic matching.
|
||||
|
||||
Some more advanced models might outperform default `sentence-transformers/all-MiniLM-L6-v2` model we were using.
|
||||
We encourage developers involved in training embedding models to include a way to extract attention weights and contribute to the BM42 backend.
|
||||
|
||||
## Fostering curiosity and experimentation
|
||||
|
||||
Despite all of its advantages, BM42 is not always a silver bullet.
|
||||
For large documents without chunks, BM25 might still be a better choice.
|
||||
|
||||
There might be a smarter way to extract the importance information from the transformer. There could be a better method to weigh IDF against attention scores.
|
||||
|
||||
Qdrant does not specialize in model training. Our core project is the search engine itself. However, we understand that we are not operating in a vacuum. By introducing BM42, we are stepping up to empower our community with novel tools for experimentation.
|
||||
|
||||
We truly believe that the sparse vectors method is at exact level of abstraction to yield both powerful and flexible results.
|
||||
|
||||
Many of you are sharing your recent Qdrant projects in our [Discord channel](https://discord.com/invite/qdrant). Feel free to try out BM42 and let us know what you come up with.
|
||||
|
||||
@@ -208,8 +208,6 @@ Now the out-of-memory happens when we allow using **600mb** RAM only
|
||||
|
||||
</details>
|
||||
|
||||
<br/>
|
||||
|
||||
At this point we have to switch from network-mounted storage to a faster disk, as the network-based storage is too slow to handle the amount of sequential reads that our system needs to serve the queries.
|
||||
|
||||
But let's first see how much RAM we need to serve 1 million vectors and then we will discuss the speed optimization as well.
|
||||
@@ -251,8 +249,6 @@ With this configuration we are able to serve 1 million vectors with **only 135mb
|
||||
|
||||
</details>
|
||||
|
||||
<br/>
|
||||
|
||||
At this point the importance of the disk speed becomes critical.
|
||||
We can serve the search requests with 135mb of RAM, but the speed of the requests makes it impossible to use the system in production.
|
||||
|
||||
|
||||
@@ -26,7 +26,5 @@ Here are the principles we followed while designing these benchmarks:
|
||||
|
||||
</details>
|
||||
|
||||
</br>
|
||||
|
||||
Some of our experiment design decisions are described in the [F.A.Q Section](/benchmarks/#benchmarks-faq).
|
||||
Reach out to us on our [Discord channel](https://qdrant.to/discord) if you want to discuss anything related Qdrant or these benchmarks.
|
||||
|
||||
@@ -0,0 +1,80 @@
|
||||
---
|
||||
title: "Community Highlights #1"
|
||||
draft: false
|
||||
slug: community-highlights-1 # Change this slug to your page slug if needed
|
||||
short_description: Celebrating top contributions and achievements in vector search, featuring standout projects, articles, and the Creator of the Month, Pavan Kumar. # Change this
|
||||
description: Celebrating top contributions and achievements in vector search, featuring standout projects, articles, and the Creator of the Month, Pavan Kumar!
|
||||
preview_image: /blog/community-highlights-1/preview-image.png
|
||||
social_preview_image: /blog/community-highlights-1/preview-image.png
|
||||
|
||||
date: 2024-06-20T11:57:37-03:00
|
||||
author: Sabrina Aquino
|
||||
featured: false
|
||||
tags:
|
||||
- news
|
||||
- vector search
|
||||
- qdrant
|
||||
- ambassador program
|
||||
- community
|
||||
- artificial intelligence
|
||||
---
|
||||
|
||||
Welcome to the very first edition of Community Highlights, where we celebrate the most impactful contributions and achievements of our vector search community! 🎉
|
||||
|
||||
## Content Highlights 🚀
|
||||
|
||||
Here are some standout projects and articles from our community this past month. If you're looking to learn more about vector search or build some great projects, we recommend you to check these guides:
|
||||
|
||||
* **[Implementing Advanced Agentic Vector Search](https://towardsdev.com/implementing-advanced-agentic-vector-search-a-comprehensive-guide-to-crewai-and-qdrant-ca214ca4d039): A Comprehensive Guide to CrewAI and Qdrant by [Pavan Kumar](https://www.linkedin.com/in/kameshwara-pavan-kumar-mantha-91678b21/)**
|
||||
* **Build Your Own RAG Using [Unstructured, Llama3 via Groq, Qdrant & LangChain](https://www.youtube.com/watch?v=m_3q3XnLlTI) by [Sudarshan Koirala](https://www.linkedin.com/in/sudarshan-koirala/)**
|
||||
* **Qdrant filtering and [self-querying retriever](https://www.youtube.com/watch?v=iaXFggqqGD0) retrieval with LangChain by [Daniel Romero](https://www.linkedin.com/in/infoslack/)**
|
||||
* **RAG Evaluation with [Arize Phoenix](https://superlinked.com/vectorhub/articles/retrieval-augmented-generation-eval-qdrant-arize) by [Atita Arora](https://www.linkedin.com/in/atitaarora/)**
|
||||
* **Building a Serverless Application with [AWS Lambda and Qdrant](https://medium.com/@benitomartin/building-a-serverless-application-with-aws-lambda-and-qdrant-for-semantic-search-ddb7646d4c2f) for Semantic Search by [Benito Martin](https://www.linkedin.com/in/benitomzh/)**
|
||||
* **Production ready Secure and [Powerful AI Implementations with Azure Services](https://towardsdev.com/production-ready-secure-and-powerful-ai-implementations-with-azure-services-671b68631212) by [Pavan Kumar](https://www.linkedin.com/in/kameshwara-pavan-kumar-mantha-91678b21/)**
|
||||
* **Building [Agentic RAG with Rust, OpenAI & Qdrant](https://medium.com/@joshmo_dev/building-agentic-rag-with-rust-openai-qdrant-d3a0bb85a267) by [Joshua Mo](https://www.linkedin.com/in/joshua-mo-4146aa220/)**
|
||||
* **Qdrant [Hybrid Search](https://medium.com/@nickprock/qdrant-hybrid-search-under-the-hood-using-haystack-355841225ac6) under the hood using Haystack by [Nicola Procopio](https://www.linkedin.com/in/nicolaprocopio/)**
|
||||
* **[Llama 3 Powered Voice Assistant](https://medium.com/@datadrifters/llama-3-powered-voice-assistant-integrating-local-rag-with-qdrant-whisper-and-langchain-b4d075b00ac5): Integrating Local RAG with Qdrant, Whisper, and LangChain by [Datadrifters](https://medium.com/@datadrifters)**
|
||||
* **[Distributed deployment](https://medium.com/@vardhanam.daga/distributed-deployment-of-qdrant-cluster-with-sharding-replicas-e7923d483ebc) of Qdrant cluster with sharding & replicas by [Vardhanam Daga](https://www.linkedin.com/in/vardhanam-daga/overlay/about-this-profile/)**
|
||||
* **Private [Healthcare AI Assistant](https://medium.com/aimpact-all-things-ai/building-private-healthcare-ai-assistant-for-clinics-using-qdrant-hybrid-cloud-jwt-rbac-dspy-and-089a772e08ae) using Qdrant Hybrid Cloud, DSPy, and Groq by [Sachin Khandewal](https://www.linkedin.com/in/sachink1729/)**
|
||||
|
||||
|
||||
## Creator of the Month 🌟
|
||||
|
||||
|
||||
<img src="/blog/community-highlights-1/creator-of-the-month-pavan.png" alt="Picture of Pavan Kumar with over 6 content contributions for the Creator of the Month" style="width: 70%;" />
|
||||
|
||||
|
||||
Congratulations to Pavan Kumar for being awarded **Creator of the Month!** Check out what were Pavan's most valuable contributions to the Qdrant vector search community this past month:
|
||||
|
||||
|
||||
* **[Implementing Advanced Agentic Vector Search](https://towardsdev.com/implementing-advanced-agentic-vector-search-a-comprehensive-guide-to-crewai-and-qdrant-ca214ca4d039): A Comprehensive Guide to CrewAI and Qdrant**
|
||||
* **Production ready Secure and [Powerful AI Implementations with Azure Services](https://towardsdev.com/production-ready-secure-and-powerful-ai-implementations-with-azure-services-671b68631212)**
|
||||
* **Building Neural Search Pipelines with Azure and Qdrant: A Step-by-Step Guide [Part-1](https://towardsdev.com/building-neural-search-pipelines-with-azure-and-qdrant-a-step-by-step-guide-part-1-40c191084258) and [Part-2](https://towardsdev.com/building-neural-search-pipelines-with-azure-and-qdrant-a-step-by-step-guide-part-2-fba287b49574)**
|
||||
* **Building a RAG System with [Ollama, Qdrant and Raspberry Pi](https://blog.gopenai.com/harnessing-ai-at-the-edge-building-a-rag-system-with-ollama-qdrant-and-raspberry-pi-45ac3212cf75)**
|
||||
* **Building a [Multi-Document ReAct Agent](https://blog.stackademic.com/building-a-multi-document-react-agent-for-financial-analysis-using-llamaindex-and-qdrant-72a535730ac3) for Financial Analysis using LlamaIndex and Qdrant**
|
||||
|
||||
Pavan is a seasoned technology expert with 14 years of extensive experience, passionate about sharing his knowledge through technical blogging, engaging in technical meetups, and staying active with cycling!
|
||||
|
||||
Thank you, Pavan, for your outstanding contributions and commitment to the community!
|
||||
|
||||
## Most Active Members 🏆
|
||||
|
||||
|
||||
<img src="/blog/community-highlights-1/most-active-members.png" alt="Picture of the 3 most active members of our vector search community" style="width: 70%;" />
|
||||
|
||||
|
||||
We're excited to recognize our most active community members, who have been a constant support to vector search builders, and sharing their knowledge and making our community more engaging:
|
||||
|
||||
* 🥇 **1st Place: Robert Caulk**
|
||||
* 🥈 **2nd Place: Nicola Procopio**
|
||||
* 🥉 **3rd Place: Joshua Mo**
|
||||
|
||||
Thank you all for your dedication and for making the Qdrant vector search community such a dynamic and valuable place!
|
||||
|
||||
Stay tuned for more highlights and updates in the next edition of Community Highlights! 🚀
|
||||
|
||||
**Join us for Office Hours! 🎙️**
|
||||
|
||||
Don't miss our next [Office Hours hangout on Discord](https://discord.gg/s9YxGeQK?event=1252726857753821236), happening next week on June 27th. This is a great opportunity to introduce yourself to the community, learn more about vector search, and engage with the people behind this awesome content!
|
||||
|
||||
See you there 👋
|
||||
@@ -0,0 +1,388 @@
|
||||
---
|
||||
title: "DSPy vs LangChain: A Comprehensive Framework Comparison" #required
|
||||
short_description: DSPy and LangChain are powerful frameworks for building AI applications leveraging LLMs and vector search technology.
|
||||
description: We dive deep into the capabilities of DSPy and LangChain and discuss scenarios where each of these frameworks shine. #required
|
||||
social_preview_image: /blog/dspy-vs-langchain/dspy-langchain.png # This image will be used in
|
||||
preview_image: /blog/dspy-vs-langchain/dspy-langchain.png
|
||||
author: Qdrant Team # Author of the article. Required.
|
||||
author_link: https://qdrant.tech/ # Link to the author's page. Required.
|
||||
date: 2024-02-23T08:00:00-03:00 # Date of the article. Required.
|
||||
draft: false # If true, the article will not be published
|
||||
keywords: # Keywords for SEO
|
||||
- DSPy
|
||||
- LangChain
|
||||
- AI frameworks
|
||||
- LLMs
|
||||
- vector search
|
||||
- RAG applications
|
||||
- chatbots
|
||||
---
|
||||
|
||||
As Large Language Models (LLMs) and vector stores have become steadily more powerful, a new generation of frameworks has appeared which can streamline the development of AI applications by leveraging LLMs and vector search technology. These frameworks simplify the process of building everything from Retrieval Augmented Generation (RAG) applications to complex chatbots with advanced conversational abilities, and even sophisticated reasoning-driven AI applications.
|
||||
|
||||
The most well-known of these frameworks is possibly [LangChain](https://github.com/langchain-ai/langchain). [Launched in October 2022](https://en.wikipedia.org/wiki/LangChain) as an open-source project by Harrison Chase, the project quickly gained popularity, attracting contributions from hundreds of developers on GitHub. LangChain excels in its broad support for documents, data sources, and APIs. This, along with seamless integration with vector stores like Qdrant and the ability to chain multiple LLMs, has allowed developers to build complex AI applications without reinventing the wheel.
|
||||
|
||||
However, despite the many capabilities unlocked by frameworks like LangChain, developers still needed expertise in [prompt engineering](https://en.wikipedia.org/wiki/Prompt_engineering) to craft optimal LLM prompts. Additionally, optimizing these prompts and adapting them to build multi-stage reasoning AI remained challenging with the existing frameworks.
|
||||
|
||||
In fact, as you start building production-grade AI applications, it becomes clear that a single LLM call isn’t enough to unlock the full capabilities of LLMs. Instead, you need to create a workflow where the model interacts with external tools like web browsers, fetches relevant snippets from documents, and compiles the results into a multi-stage reasoning pipeline.
|
||||
|
||||
This involves building an architecture that combines and reasons on intermediate outputs, with LLM prompts that adapt according to the task at hand, before producing a final output. A manual approach to prompt engineering quickly falls short in such scenarios.
|
||||
|
||||
In October 2023, researchers working in Stanford NLP released a library, [DSPy](https://github.com/stanfordnlp/dspy), which entirely automates the process of optimizing prompts and weights for large language models (LLMs), eliminating the need for manual prompting or prompt engineering.
|
||||
|
||||
One of DSPy's key features is its ability to automatically tune LLM prompts, an approach that is especially powerful when your application needs to call the LLM several times within a pipeline.
|
||||
|
||||
So, when building an LLM and vector store-backed AI application, which of these frameworks should you choose? In this article, we dive deep into the capabilities of each and discuss scenarios where each of these frameworks shine. Let’s get started!
|
||||
|
||||
## **LangChain: Features, Performance, and Use Cases**
|
||||
|
||||
LangChain, as discussed above, is an open-source orchestration framework available in both [Python](https://python.langchain.com/v0.2/docs/introduction/) and [JavaScript](https://js.langchain.com/v0.2/docs/introduction/), designed to simplify the development of AI applications leveraging LLMs. For developers working with one or multiple LLMs, it acts as a universal interface for these AI models. LangChain integrates with various external data sources, supports a wide range of data types and stores, streamlines the handling of vector embeddings and retrieval through similarity search, and simplifies the integration of AI applications with existing software workflows.
|
||||
|
||||
At a high level, LangChain abstracts the common steps required to work with language models into modular components, which serve as the building blocks of AI applications. These components can be "chained" together to create complex applications. Thanks to these abstractions, LangChain allows for rapid experimentation and prototyping of AI applications in a short timeframe.
|
||||
|
||||
LangChain breaks down the functionality required to build AI applications into three key sections:
|
||||
|
||||
- **Model I/O**: Building blocks to interface with the LLM.
|
||||
- **Retrieval**: Building blocks to streamline the retrieval of data used by the LLM for generation (such as the retrieval step in RAG applications).
|
||||
- **Composition**: Components to combine external APIs, services and other LangChain primitives.
|
||||
|
||||
These components are pulled together into ‘chains’ that are constructed using [LangChain Expression Language](https://python.langchain.com/v0.1/docs/expression_language/) (LCEL). We’ill first look at the various building blocks, and then see how they can be combined using LCEL.
|
||||
|
||||
### **LLM Model I/O**
|
||||
|
||||
LangChain offers broad compatibility with various LLMs, and its [LLM](https://python.langchain.com/v0.1/docs/modules/model_io/llms/) class provides a standard interface to these models. Leveraging proprietary models offered by platforms like OpenAI, Mistral, Cohere, or Gemini is straightforward and requires just an API key from the respective platform.
|
||||
|
||||
For instance, to use OpenAI models, you simply need to do the following:
|
||||
|
||||
```python
|
||||
from langchain_openai import OpenAI
|
||||
|
||||
llm = OpenAI(api_key="...")
|
||||
|
||||
llm.invoke("Where is Paris?")
|
||||
|
||||
```
|
||||
|
||||
|
||||
Open-source models like Meta AI’s Llama variants (such as Llama3-8B) or Mistral AI’s open models (like Mistral-7B) can be easily integrated using their Hugging Face endpoints or local LLM deployment tools like Ollama, vLLM, or LM Studio. You can also use the [CustomLLM](https://python.langchain.com/v0.1/docs/modules/model_io/llms/custom_llm/) class to build Custom LLM wrappers.
|
||||
|
||||
Here’s how simple it is to use LangChain with LlaMa3-8B, using [Ollama](https://ollama.com/).
|
||||
|
||||
```python
|
||||
from langchain_community.llms import Ollama
|
||||
|
||||
llm = Ollama(model="llama3")
|
||||
|
||||
llm.invoke("Where is Berlin?")
|
||||
|
||||
```
|
||||
|
||||
|
||||
LangChain also offers output parsers to structure the LLM output in a format that the application may need, such as structured data types like JSON, XML, CSV, and others. To understand LangChain’s interface with LLMs in detail, read the documentation [here](https://python.langchain.com/v0.1/docs/modules/model_io/).
|
||||
|
||||
### **Retrieval**
|
||||
|
||||
Most enterprise AI applications are built by augmenting the LLM context using data specific to the application’s use case. To accomplish this, the relevant data needs to be first retrieved, typically using vector similarity search, and then passed to the LLM context at the generation step. This architecture, known as [Retrieval Augmented Generation](/articles/what-is-rag-in-ai/) (RAG), can be used to build a wide range of AI applications.
|
||||
|
||||
While the retrieval process sounds simple, it involves a number of complex steps: loading data from a source, splitting it into chunks, converting it into vectors or vector embeddings, storing it in a vector store, and then retrieving results based on a query before the generation step.
|
||||
|
||||
LangChain offers a number of building blocks to make this retrieval process simpler.
|
||||
|
||||
- **Document Loaders**: LangChain offers over 100 different document loaders, including integrations with providers like Unstructured or Airbyte. It also supports loading various types of documents, such as PDFs, HTML, CSV, and code, from a range of locations like S3.
|
||||
- **Splitting**: During the retrieval step, you typically need to retrieve only the relevant section of a document. To do this, you need to split a large document into smaller chunks. LangChain offers various document transformers that make it easy to split, combine, filter, or manipulate documents.
|
||||
- **Text Embeddings**: A key aspect of the retrieval step is converting document chunks into vectors, which are high-dimensional numerical representations that capture the semantic meaning of the text. LangChain offers integrations with over 25 embedding providers and methods, such as [FastEmbed](https://github.com/qdrant/fastembed).
|
||||
- **Vector Store Integration**: LangChain integrates with over 50 vector stores, including specialized ones like [Qdrant](/documentation/frameworks/langchain/), and exposes a standard interface.
|
||||
- **Retrievers**: LangChain offers various retrieval algorithms and allows you to use third-party retrieval algorithms or create custom retrievers.
|
||||
- **Indexing**: LangChain also offers an indexing API that keeps data from any data source in sync with the vector store, helping to reduce complexities around managing unchanged content or avoiding duplicate content.
|
||||
|
||||
### **Composition**
|
||||
|
||||
Finally, LangChain also offers building blocks that help combine external APIs, services, and LangChain primitives. For instance, it provides tools to fetch data from Wikipedia or search using Google Lens. The list of tools it offers is [extremely varied](https://python.langchain.com/v0.1/docs/integrations/tools/).
|
||||
|
||||
LangChain also offers ways to build agents that use language models to decide on the sequence of actions to take.
|
||||
|
||||
### **LCEL**
|
||||
|
||||
The primary method of building an application in LangChain is through the use of [LCEL](https://python.langchain.com/v0.1/docs/expression_language/), the LangChain Expression Language. It is a declarative syntax designed to simplify the composition of chains within the LangChain framework. It provides a minimalist code layer that enables the rapid development of chains, leveraging advanced features such as streaming, asynchronous execution, and parallel processing.
|
||||
|
||||
LCEL is particularly useful for building chains that involve multiple language model calls, data transformations, and the integration of outputs from language models into downstream applications.
|
||||
|
||||
### **Some Use Cases of LangChain**
|
||||
|
||||
Given the flexibility that LangChain offers, a wide range of applications can be built using the framework. Here are some examples:
|
||||
|
||||
**RAG Applications**: LangChain provides all the essential building blocks needed to build Retrieval Augmented Generation (RAG) applications. It integrates with vector stores and LLMs, streamlining the entire process of loading, chunking, and retrieving relevant sections of a document in a few lines of code.
|
||||
|
||||
**Chatbots**: LangChain offers a suite of components that streamline the process of building conversational chatbots. These include chat models, which are specifically designed for message-based interactions and provide a conversational tone suitable for chatbots.
|
||||
|
||||
**Extracting Structured Outputs**: LangChain assists in extracting structured output from data using various tools and methods. It supports multiple extraction approaches, including tool/function calling mode, JSON mode, and prompting-based extraction.
|
||||
|
||||
**Agents**: LangChain simplifies the process of building agents by providing building blocks and integration with LLMs, enabling developers to construct complex, multi-step workflows. These agents can interact with external data sources and tools, and generate dynamic and context-aware responses for various applications.
|
||||
|
||||
If LangChain offers such a wide range of integrations and the primary building blocks needed to build AI applications, *why do we need another framework?*
|
||||
|
||||
As Omar Khattab, PhD, Stanford and researcher at Stanford NLP, said when introducing DSPy in his [talk](https://www.youtube.com/watch?v=Dt3H2ninoeY) at ‘Scale By the Bay’ in November 2023: “We can build good reliable systems with these new artifacts that are language models (LMs), but importantly, this is conditioned on us *adapting* them as well as *stacking* them well”.
|
||||
|
||||
## **DSPy: Features, Performance, and Use Cases**
|
||||
|
||||
When building AI systems, developers need to break down the task into multiple reasoning steps, adapt language model (LM) prompts for each step until they get the right results, and then ensure that the steps work together to achieve the desired outcome.
|
||||
|
||||
Complex multihop pipelines, where multiple LLM calls are stacked, are messy. They involve string-based prompting tricks or prompt hacks at each step, and getting the pipeline to work is even trickier.
|
||||
|
||||
Additionally, the manual prompting approach is highly unscalable, as any change in the underlying language model breaks the prompts and the pipeline. LMs are highly sensitive to prompts and slight changes in wording, context, or phrasing can significantly impact the model's output. Due to this, despite the functionality provided by frameworks like LangChain, developers often have to spend a lot of time engineering prompts to get the right results from LLMs.
|
||||
|
||||
How do you build a system that’s less brittle and more predictable? Enter DSPy!
|
||||
|
||||
[DSPy](https://github.com/stanfordnlp/dspy) is built on the paradigm that language models (LMs) should be programmed rather than prompted. The framework is designed for algorithmically optimizing and adapting LM prompts and weights, and focuses on replacing prompting techniques with a programming-centric approach.
|
||||
|
||||
DSPy treats the LM like a device and abstracts out the underlying complexities of prompting. To achieve this, DSPy introduces three simple building blocks:
|
||||
|
||||
### **Signatures**
|
||||
|
||||
[Signatures](https://dspy-docs.vercel.app/docs/building-blocks/signatures) replace handwritten prompts and are written in natural language. They are simply declarations or specs of the behavior that you expect from the language model. Some examples are:
|
||||
|
||||
- question -> answer
|
||||
- long_document -> summary
|
||||
- context, question -> rationale, response
|
||||
|
||||
Rather than manually crafting complex prompts or engaging in extensive fine-tuning of LLMs, signatures allow for the automatic generation of optimized prompts.
|
||||
|
||||
DSPy Signatures can be specified in two ways:
|
||||
|
||||
1. Inline Signatures: Simple tasks can be defined in a concise format, like "question -> answer" for question-answering or "document -> summary" for summarization.
|
||||
|
||||
2. Class-Based Signatures: More complex tasks might require class-based signatures, which can include additional instructions or descriptions about the inputs and outputs. For example, a class for emotion classification might clearly specify the range of emotions that can be classified.
|
||||
|
||||
### **Modules**
|
||||
|
||||
Modules take signatures as input, and automatically generate high-quality prompts. Inspired heavily from PyTorch, DSPy [modules](https://dspy-docs.vercel.app/docs/building-blocks/modules) eliminate the need for crafting prompts manually.
|
||||
|
||||
The framework supports advanced modules like [dspy.ChainOfThought](https://dspy-docs.vercel.app/api/modules/ChainOfThought), which adds step-by-step rationalization before producing an output. The output not only provides answers but also rationales. Other modules include [dspy.ProgramOfThought](https://dspy-docs.vercel.app/api/modules/ProgramOfThought), which outputs code whose execution results dictate the response, and [dspy.ReAct](https://dspy-docs.vercel.app/api/modules/ReAct), an agent that uses tools to implement signatures.
|
||||
|
||||
DSPy also offers modules like [dspy.MultiChainComparison](https://dspy-docs.vercel.app/api/modules/MultiChainComparison), which can compare multiple outputs from dspy.ChainOfThought in order to produce a final prediction. There are also utility modules like [dspy.majority](https://dspy-docs.vercel.app/docs/building-blocks/modules#what-other-dspy-modules-are-there-how-can-i-use-them) for aggregating responses through voting.
|
||||
|
||||
Modules can be composed into larger programs, and you can compose multiple modules into bigger modules. This allows you to create complex, behavior-rich applications using language models.
|
||||
|
||||
### **Optimizers**
|
||||
|
||||
[Optimizers](https://dspy-docs.vercel.app/docs/building-blocks/optimizers) take a set of modules that have been connected to create a pipeline, compile them into auto-optimized prompts, and maximize an outcome metric.
|
||||
|
||||
Essentially, optimizers are designed to generate, test, and refine prompts, and ensure that the final prompt is highly optimized for the specific dataset and task at hand. Using optimizers in the DSPy framework significantly simplifies the process of developing and refining LM applications by automating the prompt engineering process.
|
||||
|
||||
### **Building AI Applications with DSPy**
|
||||
|
||||
A typical DSPy program requires the developer to follow the following 8 steps:
|
||||
|
||||
1. **Defining the Task**: Identify the specific problem you want to solve, including the input and output formats.
|
||||
2. **Defining the Pipeline**: Plan the sequence of operations needed to solve the task. Then craft the signatures and the modules.
|
||||
3. **Testing with Examples**: Run the pipeline with a few examples to understand the initial performance. This helps in identifying immediate issues with the program and areas for improvement.
|
||||
4. **Defining Your Data**: Prepare and structure your training and validation datasets. This is needed by the optimizer for training the model and evaluating its performance accurately.
|
||||
5. **Defining Your Metric**: Choose metrics that will measure the success of your model. These metrics help the optimizer evaluate how well the model is performing.
|
||||
6. **Collecting Zero-Shot Evaluations**: Run initial evaluations without prior training to establish a baseline. This helps in understanding the model’s capabilities and limitations out of the box.
|
||||
7. **Compiling with a DSPy Optimizer**: Given the data and metric, you can now optimize the program. DSPy offers a variety of optimizers designed for different purposes. These optimizers can generate step-by-step examples, craft detailed instructions, and/or update language model prompts and weights as needed.
|
||||
8. **Iterating**: Continuously refine each aspect of your task, from the pipeline and data to the metrics and evaluations. Iteration helps in gradually improving the model’s performance and adapting to new requirements.
|
||||
9.
|
||||
|
||||
|
||||
{{< figure src=/blog/dspy-vs-langchain/process.jpg caption="Process" >}}
|
||||
|
||||
**Language Model Setup**
|
||||
|
||||
Setting up the LM in DSPy is easy.
|
||||
|
||||
```python
|
||||
# pip install dspy
|
||||
|
||||
import dspy
|
||||
|
||||
llm = dspy.OpenAI(model='gpt-3.5-turbo-1106', max_tokens=300)
|
||||
|
||||
dspy.configure(lm=llm)
|
||||
|
||||
# Let's test this. First define a module (ChainOfThought) and assign it a signature (return an answer, given a question).
|
||||
|
||||
qa = dspy.ChainOfThought('question -> answer')
|
||||
|
||||
# Then, run with the default LM configured.
|
||||
|
||||
response = qa(question="Where is Paris?")
|
||||
|
||||
print(response.answer)
|
||||
|
||||
```
|
||||
|
||||
You are not restricted to using one LLM in your program; you can use [multiple](https://dspy-docs.vercel.app/docs/building-blocks/language_models#using-multiple-lms-at-once). DSPy can be used with both managed models such as OpenAI, Cohere, Anyscale, Together, or PremAI as well as with local LLM deployments through vLLM, Ollama, or TGI server. All LLM calls are cached by default.
|
||||
|
||||
**Vector Store Integration (Retrieval Model)**
|
||||
|
||||
You can easily set up [Qdrant](/documentation/frameworks/dspy/) vector store to act as the retrieval model. To do so, follow these steps:
|
||||
|
||||
```python
|
||||
# pip install dspy-ai[qdrant]
|
||||
|
||||
import dspy
|
||||
|
||||
from dspy.retrieve.qdrant_rm import QdrantRM
|
||||
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
llm = dspy.OpenAI(model="gpt-3.5-turbo")
|
||||
|
||||
qdrant_client = QdrantClient()
|
||||
|
||||
qdrant_rm = QdrantRM("collection-name", qdrant_client, k=3)
|
||||
|
||||
dspy.settings.configure(lm=llm, rm=qdrant_rm)
|
||||
|
||||
```
|
||||
|
||||
The above code sets up DSPy to use Qdrant (localhost), with collection-name as the default retrieval client. You can now build a RAG module in the following way:
|
||||
|
||||
```python
|
||||
|
||||
class RAG(dspy.Module):
|
||||
def __init__(self, num_passages=5):
|
||||
super().__init__()
|
||||
|
||||
self.retrieve = dspy.Retrieve(k=num_passages)
|
||||
self.generate_answer = dspy.ChainOfThought('context, question -> answer') # using inline signature
|
||||
|
||||
def forward(self, question):
|
||||
context = self.retrieve(question).passages
|
||||
prediction = self.generate_answer(context=context, question=question)
|
||||
return dspy.Prediction(context=context, answer=prediction.answer)
|
||||
|
||||
```
|
||||
|
||||
Now you can use the RAG module like any Python module.
|
||||
|
||||
**Optimizing the Pipeline**
|
||||
|
||||
In this step, DSPy requires you to create a training dataset and a metric function, which can help validate the output of your program. Using this, DSPy tunes the parameters (i.e., the prompts and/or the LM weights) to maximize the accuracy of the RAG pipeline.
|
||||
|
||||
Using DSPy optimizers involves the following steps:
|
||||
|
||||
1. Set up your DSPy program with the desired signatures and modules.
|
||||
2. Create a training and validation dataset, with example input and output that you expect from your DSPy program.
|
||||
3. Choose an appropriate optimizer such as BootstrapFewShotWithRandomSearch, MIPRO, or BootstrapFinetune.
|
||||
4. Create a metric function that evaluates the performance of the DSPy program. You can evaluate based on accuracy or quality of responses, or on a metric that’s relevant to your program.
|
||||
5. Run the optimizer with the DSPy program, metric function, and training inputs. DSPy will compile the program and automatically adjust parameters and improve performance.
|
||||
6. Use the compiled program to perform the task. Iterate and adapt if required.
|
||||
|
||||
To learn more about optimizing DSPy programs, read [this](https://dspy-docs.vercel.app/docs/building-blocks/optimizers).
|
||||
|
||||
DSPy is heavily influenced by PyTorch, and replaces complex prompting with reusable modules for common tasks. Instead of crafting specific prompts, you write code that DSPy automatically translates for the LLM. This, along with built-in optimizers, makes working with LLMs more systematic and efficient.
|
||||
|
||||
### **Use Cases of DSPy**
|
||||
|
||||
As we saw above, DSPy can be used to create fairly complex applications which require stacking multiple LM calls without the need for prompt engineering. Even though the framework is comparatively new - it started gaining popularity since November 2023 when it was first introduced - it has created a promising new direction for LLM-based applications.
|
||||
|
||||
Here are some of the possible uses of DSPy:
|
||||
|
||||
**Automating Prompt Engineering**: DSPy automates the process of creating prompts for LLMs, and allows developers to focus on the core logic of their application. This is powerful as manual prompt engineering makes AI applications highly unscalable and brittle.
|
||||
|
||||
**Building Chatbots**: The modular design of DSPy makes it well-suited for creating chatbots with improved response quality and faster development cycles. DSPy's automatic prompting and optimizers can help ensure chatbots generate consistent and informative responses across different conversation contexts.
|
||||
|
||||
**Complex Information Retrieval Systems**: DSPy programs can be easily integrated with vector stores, and used to build multi-step information retrieval systems with stacked calls to the LLM. This can be used to build highly sophisticated retrieval systems. For example, DSPy can be used to develop custom search engines that understand complex user queries and retrieve the most relevant information from vector stores.
|
||||
|
||||
**Improving LLM Pipelines**: One of the best uses of DSPy is to optimize LLM pipelines. DSPy's modular design greatly simplifies the integration of LLMs into existing workflows. Additionally, DSPy's built-in optimizers can help fine-tune LLM pipelines based on desired metrics.
|
||||
|
||||
**Multi-Hop Question-Answering**: Multi-hop question-answering involves answering complex questions that require reasoning over multiple pieces of information, which are often scattered across different documents or sections of text. With DSPy, users can leverage its automated prompt engineering capabilities to develop prompts that effectively guide the model on how to piece together information from various sources.
|
||||
|
||||
## **Comparative Analysis: DSPy vs LangChain**
|
||||
|
||||
DSPy and LangChain are both powerful frameworks for building AI applications, leveraging large language models (LLMs) and vector search technology. Below is a comparative analysis of their key features, performance, and use cases:
|
||||
|
||||
| Feature | LangChain | DSPy |
|
||||
| --- | --- | --- |
|
||||
| Core Focus | Focus on providing a large number of building blocks to simplify the development of applications that use LLMs in conjunction with user-specified data sources. | Focus on automating and modularizing LLM interactions, eliminating manual prompt engineering and improving systematic reliability. |
|
||||
| Approach | Utilizes modular components and chains that can be linked together using the LangChain Expression Language (LCEL). | Streamlines LLM interaction by prioritizing programming instead of prompting, and automating prompt refinement and weight tuning. |
|
||||
| Complex Pipelines | Facilitates the creation of chains using LCEL, supporting asynchronous execution and integration with various data sources and APIs. | Simplifies multi-stage reasoning pipelines using modules and optimizers, and ensures scalability through less manual intervention. |
|
||||
| Optimization | Relies on user expertise for prompt engineering and chaining of multiple LLM calls. | Includes built-in optimizers that automatically tune prompts and weights, and helps bring efficiency and effectiveness in LLM pipelines. |
|
||||
| Community and Support | Large open-source community with extensive documentation and examples. | Emerging framework with growing community support, and bringing a paradigm-shift in LLM prompting. |
|
||||
|
||||
### **LangChain**
|
||||
|
||||
Strengths:
|
||||
|
||||
1. Data Sources and APIs: LangChain supports a wide variety of data sources and APIs, and allows seamless integration with different types of data. This makes it highly versatile for various AI applications.
|
||||
2. LangChain provides modular components that can be chained together and allows you to create complex AI workflows. LangChain Expression Language (LCEL) lets you use declarative syntax and makes it easier to build and manage workflows.
|
||||
3. Since LangChain is an older framework, it has extensive documentation and thousands of examples that developers can take inspiration from.
|
||||
|
||||
Weaknesses:
|
||||
|
||||
1. For projects involving complex, multi-stage reasoning tasks, LangChain requires significant manual prompt engineering. This can be time-consuming and prone to errors.
|
||||
2. Scalability Issues: Managing and scaling workflows that require multiple LLM calls can be pretty challenging.
|
||||
3. Developers need sound understanding of prompt engineering in order to build applications that require multiple calls to the LLM.
|
||||
|
||||
### **DSPy**
|
||||
|
||||
Strengths:
|
||||
|
||||
1. DSPy automates the process of prompt generation and optimization, and significantly reduces the need for manual prompt engineering. This makes working with LLMs easier and helps build scalable AI workflows.
|
||||
2. The framework includes built-in optimizers like BootstrapFewShot and MIPRO, which automatically refine prompts and adapt them to specific datasets.
|
||||
3. DSPy uses general-purpose modules and optimizers to simplify the complexities of prompt engineering. This can help you create complex multi-step reasoning applications easily, without worrying about the intricacies of dealing with LLMs.
|
||||
4. DSPy supports various LLMs, including the flexibility of using multiple LLMs in the same program.
|
||||
5. By focusing on programming rather than prompting, DSPy ensures higher reliability and performance for AI applications, particularly those that require complex multi-stage reasoning.
|
||||
|
||||
Weaknesses:
|
||||
|
||||
1. As a newer framework, DSPy has a smaller community compared to LangChain. This means you will have limited availability of resources, examples, and community support.
|
||||
2. Although DSPy offers tutorials and guides, its documentation is less extensive than LangChain’s, which can pose challenges when you start.
|
||||
3. When starting with DSPy, you may feel limited to the paradigms and modules it provides.
|
||||
|
||||
## **Selecting the Ideal Framework for Your AI Project**
|
||||
|
||||
When deciding between DSPy and LangChain for your AI project, you should consider the problem statement and choose the framework that best aligns with your project goals.
|
||||
|
||||
Here are some guidelines:
|
||||
|
||||
### **Project Type**
|
||||
|
||||
**LangChain**: LangChain is ideal for projects that require extensive integration with multiple data sources and APIs, especially projects that benefit from the wide range of document loaders, vector stores, and retrieval algorithms that it supports.
|
||||
|
||||
**DSPy**: DSPy is best suited for projects that involve complex multi-stage reasoning pipelines or those that may eventually need stacked LLM calls. DSPy’s systematic approach to prompt engineering and its ability to optimize LLM interactions can help create highly reliable AI applications.
|
||||
|
||||
### **Technical Expertise**
|
||||
|
||||
**LangChain**: As the complexity of the application grows, LangChain requires a good understanding of prompt engineering and expertise in chaining multiple LLM calls.
|
||||
|
||||
**DSPy**: Since DSPy is designed to abstract away the complexities of prompt engineering, it makes it easier for developers to focus on high-level logic rather than low-level prompt crafting.
|
||||
|
||||
### **Community and Support**
|
||||
|
||||
**LangChain**: LangChain boasts a large and active community with extensive documentation, examples, and active contributions, and you will find it easier to get going.
|
||||
|
||||
**DSPy**: Although newer and with a smaller community, DSPy is growing rapidly and offers tutorials and guides for some of the key use cases. DSPy may be more challenging to get started with, but its architecture makes it highly scalable.
|
||||
|
||||
### **Use Case Scenarios**
|
||||
|
||||
**Retrieval Augmented Generation (RAG) Applications**
|
||||
|
||||
**LangChain**: Excellent for building simple RAG applications due to its robust support for vector stores, document loaders, and retrieval algorithms.
|
||||
|
||||
**DSPy**: Suitable for RAG applications requiring high reliability and automated prompt optimization, ensuring consistent performance across complex retrieval tasks.
|
||||
|
||||
**Chatbots and Conversational AI**
|
||||
|
||||
**LangChain**: Provides a wide range of components for building conversational AI, making it easy to integrate LLMs with external APIs and services.
|
||||
|
||||
**DSPy**: Ideal for developing chatbots that need to handle complex, multi-stage conversations with high reliability and performance. DSPy’s automated optimizations ensure consistent and contextually accurate responses.
|
||||
|
||||
**Complex Information Retrieval Systems**
|
||||
|
||||
**LangChain**: Effective for projects that require seamless integration with various data sources and sophisticated retrieval capabilities.
|
||||
|
||||
**DSPy**: Best for systems that involve complex multi-step retrieval processes, where prompt optimization and modular design can significantly enhance performance and reliability.
|
||||
|
||||
You can also choose to combine and use the best features of both. In fact, LangChain has released an [integration with DSPy](https://python.langchain.com/v0.1/docs/integrations/providers/dspy/) to simplify this process. This allows you to use some of the utility functions that LangChain provides, such as text splitter, directory loaders, or integrations with other data sources while using DSPy for the LM interactions.
|
||||
|
||||
## **Level Up Your AI Projects with Advanced Frameworks**
|
||||
|
||||
LangChain and DSPy both offer unique capabilities and can help you build powerful AI applications. Qdrant integrates with both LangChain and DSPy, allowing you to leverage its performance, efficiency and security features in either scenario. LangChain is ideal for projects that require extensive integration with various data sources and APIs. On the other hand, DSPy offers a powerful paradigm for building complex multi-stage applications. For pulling together an AI application that doesn’t require much prompt engineering, use LangChain. However, pick DSPy when you need a systematic approach to prompt optimization and modular design, and need robustness and scalability for complex, multi-stage reasoning applications.
|
||||
|
||||
## **References**
|
||||
|
||||
[https://python.langchain.com/v0.1/docs/get_started/introduction](https://python.langchain.com/v0.1/docs/get_started/introduction)
|
||||
|
||||
[https://dspy-docs.vercel.app/docs/intro](https://dspy-docs.vercel.app/docs/intro)
|
||||
@@ -0,0 +1,574 @@
|
||||
---
|
||||
title: "Qdrant 1.10 - Universal Query, Built-in IDF & ColBERT Support"
|
||||
draft: false
|
||||
short_description: "Single search API. Server-side IDF. Native multivector support."
|
||||
description: "Consolidated search API, built-in IDF, and native multivector support. "
|
||||
preview_image: /blog/qdrant-1.10.x/social_preview.png
|
||||
social_preview_image: /blog/qdrant-1.10.x/social_preview.png
|
||||
date: 2024-07-01T00:00:00-08:00
|
||||
author: David Myriel
|
||||
featured: false
|
||||
tags:
|
||||
- vector search
|
||||
- ColBERT late interaction
|
||||
- BM25 algorithm
|
||||
- search API
|
||||
- new features
|
||||
---
|
||||
|
||||
[Qdrant 1.10.0 is out!](https://github.com/qdrant/qdrant/releases/tag/v1.10.0) This version introduces some major changes, so let's dive right in:
|
||||
|
||||
**Universal Query API:** All search APIs, including Hybrid Search, are now in one Query endpoint.</br>
|
||||
**Built-in IDF:** We added the IDF mechanism to Qdrant's core search and indexing processes.</br>
|
||||
**Multivector Support:** Native support for late interaction ColBERT is accessible via Query API.
|
||||
|
||||
## One Endpoint for All Queries
|
||||
**Query API** will consolidate all search APIs into a single request. Previously, you had to work outside of the API to combine different search requests. Now these approaches are reduced to parameters of a single request, so you can avoid merging individual results.
|
||||
|
||||
You can now configure the Query API request with the following parameters:
|
||||
|
||||
|Parameter|Description|
|
||||
|-|-|
|
||||
|no parameter|Returns points by `id`|
|
||||
|`nearest`|Queries nearest neighbors ([Search](/documentation/concepts/search/))|
|
||||
|`fusion`|Fuses sparse/dense prefetch queries ([Hybrid Search](/documentation/concepts/hybrid-queries/#hybrid-search))|
|
||||
|`discover`|Queries `target` with added `context` ([Discovery](/documentation/concepts/explore/#discovery-api))|
|
||||
|`context` |No target with `context` only ([Context](/documentation/concepts/explore/#context-search))|
|
||||
|`recommend`|Queries against `positive`/`negative` examples. ([Recommendation](/documentation/concepts/explore/#recommendation-api))|
|
||||
|`order_by`|Orders results by [payload field](/documentation/concepts/hybrid-queries/#re-ranking-with-payload-values)|
|
||||
|
||||
For example, you can configure Query API to run [Discovery search](/documentation/concepts/explore/#discovery-api). Let's see how that looks:
|
||||
|
||||
```http
|
||||
POST collections/{collection_name}/points/query
|
||||
{
|
||||
"query": {
|
||||
"discover": {
|
||||
"target": <vector_input>,
|
||||
"context": [
|
||||
{
|
||||
"positive": <vector_input>,
|
||||
"negative": <vector_input>
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
We will be publishing code samples in [docs](/documentation/concepts/hybrid-queries/) and our new [API specification](http://api.qdrant.tech).</br> *If you need additional support with this new method, our [Discord](https://qdrant.to/discord) on-call engineers can help you.*
|
||||
|
||||
### Native Hybrid Search Support
|
||||
Query API now also natively supports **sparse/dense fusion**. Up to this point, you had to combine the results of sparse and dense searches on your own. This is now sorted on the back-end, and you only have to configure them as basic parameters for Query API.
|
||||
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"prefetch": [
|
||||
{
|
||||
"query": {
|
||||
"indices": [1, 42], // <┐
|
||||
"values": [0.22, 0.8] // <┴─sparse vector
|
||||
},
|
||||
"using": "sparse",
|
||||
"limit": 20,
|
||||
},
|
||||
{
|
||||
"query": [0.01, 0.45, 0.67, ...], // <-- dense vector
|
||||
"using": "dense",
|
||||
"limit": 20,
|
||||
}
|
||||
],
|
||||
"query": { "fusion": "rrf" }, // <--- reciprocal rank fusion
|
||||
"limit": 10
|
||||
}
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{Fusion, PrefetchQueryBuilder, Query, QueryPointsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.add_prefetch(PrefetchQueryBuilder::default()
|
||||
.query(Query::new_nearest([(1, 0.22), (42, 0.8)].as_slice()))
|
||||
.using("sparse")
|
||||
.limit(20u64)
|
||||
)
|
||||
.add_prefetch(PrefetchQueryBuilder::default()
|
||||
.query(Query::new_nearest(vec![0.01, 0.45, 0.67]))
|
||||
.using("dense")
|
||||
.limit(20u64)
|
||||
)
|
||||
.query(Query::new_fusion(Fusion::Rrf))
|
||||
).await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
import static io.qdrant.client.QueryFactory.fusion;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.Fusion;
|
||||
import io.qdrant.client.grpc.Points.PrefetchQuery;
|
||||
import io.qdrant.client.grpc.Points.QueryPoints;
|
||||
|
||||
QdrantClient client = new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client.queryAsync(
|
||||
QueryPoints.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.addPrefetch(PrefetchQuery.newBuilder()
|
||||
.setQuery(nearest(List.of(0.22f, 0.8f), List.of(1, 42)))
|
||||
.setUsing("sparse")
|
||||
.setLimit(20)
|
||||
.build())
|
||||
.addPrefetch(PrefetchQuery.newBuilder()
|
||||
.setQuery(nearest(List.of(0.01f, 0.45f, 0.67f)))
|
||||
.setUsing("dense")
|
||||
.setLimit(20)
|
||||
.build())
|
||||
.setQuery(fusion(Fusion.RRF))
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
prefetch: new List < PrefetchQuery > {
|
||||
new() {
|
||||
Query = new(float, uint)[] {
|
||||
(0.22f, 1), (0.8f, 42),
|
||||
},
|
||||
Using = "sparse",
|
||||
Limit = 20
|
||||
},
|
||||
new() {
|
||||
Query = new float[] {
|
||||
0.01f, 0.45f, 0.67f
|
||||
},
|
||||
Using = "dense",
|
||||
Limit = 20
|
||||
}
|
||||
},
|
||||
query: Fusion.Rrf
|
||||
);
|
||||
```
|
||||
|
||||
Query API can now pre-fetch vectors for requests, which means you can run queries sequentially within the same API call. There are a lot of options here, so you will need to define a strategy to merge these requests using new parameters. For example, you can now include **rescoring within Hybrid Search**, which can open the door to strategies like iterative refinement via matryoshka embeddings.
|
||||
|
||||
*To learn more about this, read the [Query API documentation](/documentation/concepts/search/#query-api).*
|
||||
|
||||
## Inverse Document Frequency [IDF]
|
||||
|
||||
IDF is a critical component of the **TF-IDF (Term Frequency-Inverse Document Frequency)** weighting scheme used to evaluate the importance of a word in a document relative to a collection of documents (corpus).
|
||||
There are various ways in which IDF might be calculated, but the most commonly used formula is:
|
||||
|
||||
$$
|
||||
\text{IDF}(q_i) = \ln \left(\frac{N - n(q_i) + 0.5}{n(q_i) + 0.5}+1\right)
|
||||
$$
|
||||
|
||||
Where:</br>
|
||||
`N` is the total number of documents in the collection. </br>
|
||||
`n` is the number of documents containing non-zero values for the given vector.
|
||||
|
||||
This variant is also used in BM25, whose support was heavily requested by our users. We decided to move the IDF calculation into the Qdrant engine itself. This type of separation allows streaming updates of the sparse embeddings while keeping the IDF calculation up-to-date.
|
||||
|
||||
The values of IDF previously had to be calculated using all the documents on the client side. However, now that Qdrant does it out of the box, you won't need to implement it anywhere else and recompute the value if some documents are removed or newly added.
|
||||
|
||||
You can enable the IDF modifier in the collection configuration:
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"sparse_vectors": {
|
||||
"text": {
|
||||
"modifier": "idf"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
sparse_vectors={
|
||||
"text": models.SparseVectorParams(
|
||||
modifier=models.Modifier.IDF,
|
||||
),
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{CreateCollectionBuilder, sparse_vectors_config::SparseVectorsConfigBuilder, Modifier, SparseVectorParamsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
let mut config = SparseVectorsConfigBuilder::default();
|
||||
config.add_named_vector_params(
|
||||
"text",
|
||||
SparseVectorParamsBuilder::default().modifier(Modifier::Idf),
|
||||
);
|
||||
|
||||
client
|
||||
.create_collection(
|
||||
CreateCollectionBuilder::new("{collection_name}")
|
||||
.sparse_vectors_config(config),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||
import io.qdrant.client.grpc.Collections.Modifier;
|
||||
import io.qdrant.client.grpc.Collections.SparseVectorConfig;
|
||||
import io.qdrant.client.grpc.Collections.SparseVectorParams;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.createCollectionAsync(
|
||||
CreateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setSparseVectorsConfig(
|
||||
SparseVectorConfig.newBuilder()
|
||||
.putMap("text", SparseVectorParams.newBuilder().setModifier(Modifier.Idf).build()))
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
sparseVectorsConfig: ("text", new SparseVectorParams {
|
||||
Modifier = Modifier.Idf,
|
||||
})
|
||||
);
|
||||
```
|
||||
|
||||
### IDF as Part of BM42
|
||||
|
||||
This quarter, Qdrant also introduced BM42, a novel algorithm that combines the IDF element of BM25 with transformer-based attention matrices to improve text retrieval. It utilizes attention matrices from your embedding model to determine the importance of each token in the document based on the attention value it receives.
|
||||
|
||||
We've prepared the standard `all-MiniLM-L6-v2` Sentence Transformer so [it outputs the attention values](https://huggingface.co/Qdrant/all_miniLM_L6_v2_with_attentions). Still, you can use virtually any model of your choice, as long as you have access to its parameters. This is just another reason to stick with open source technologies over proprietary systems.
|
||||
|
||||
In practical terms, the BM42 method addresses the tokenization issues and computational costs associated with SPLADE. The model is both efficient and effective across different document types and lengths, offering enhanced search performance by leveraging the strengths of both BM25 and modern transformer techniques.
|
||||
|
||||
> To learn more about IDF and BM42, read our [dedicated technical article](/articles/bm42/).
|
||||
|
||||
**You can expect BM42 to excel in scalable RAG-based scenarios where short texts are more common.** Document inference speed is much higher with BM42, which is critical for large-scale applications such as search engines, recommendation systems, and real-time decision-making systems.
|
||||
|
||||
## Multivector Support
|
||||
We are adding native support for multivector search that is compatible, e.g., with the late-interaction [ColBERT](https://github.com/stanford-futuredata/ColBERT) model. If you are working with high-dimensional similarity searches, **ColBERT is highly recommended as a reranking step in the Universal Query search.** You will experience better quality vector retrieval since ColBERT’s approach allows for deeper semantic understanding.
|
||||
|
||||
This model retains contextual information during query-document interaction, leading to better relevance scoring. In terms of efficiency and scalability benefits, documents and queries will be encoded separately, which gives an opportunity for pre-computation and storage of document embeddings for faster retrieval.
|
||||
|
||||
**Note:** *This feature supports all the original quantization compression methods, just the same as the regular search method.*
|
||||
|
||||
**Run a query with ColBERT vectors:**
|
||||
|
||||
Query API can handle exceedingly complex requests. The following example prefetches 1000 entries most similar to the given query using the `mrl_byte` named vector, then reranks them to get the best 100 matches with `full` named vector and eventually reranks them again to extract the top 10 results with the named vector called `colbert`. A single API call can now implement complex reranking schemes.
|
||||
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"prefetch": {
|
||||
"prefetch": {
|
||||
"query": [1, 23, 45, 67], // <------ small byte vector
|
||||
"using": "mrl_byte"
|
||||
"limit": 1000,
|
||||
},
|
||||
"query": [0.01, 0.45, 0.67, ...], // <-- full dense vector
|
||||
"using": "full"
|
||||
"limit": 100,
|
||||
},
|
||||
"query": [ // <─┐
|
||||
[0.1, 0.2, ...], // < │
|
||||
[0.2, 0.1, ...], // < ├─ multi-vector
|
||||
[0.8, 0.9, ...] // < │
|
||||
], // <─┘
|
||||
"using": "colbert",
|
||||
"limit": 10
|
||||
}
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{PrefetchQueryBuilder, Query, QueryPointsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.add_prefetch(PrefetchQueryBuilder::default()
|
||||
.add_prefetch(PrefetchQueryBuilder::default()
|
||||
.query(Query::new_nearest(vec![1.0, 23.0, 45.0, 67.0]))
|
||||
.using("mlr_byte")
|
||||
.limit(1000u64)
|
||||
)
|
||||
.query(Query::new_nearest(vec![0.01, 0.45, 0.67]))
|
||||
.using("full")
|
||||
.limit(100u64)
|
||||
)
|
||||
.query(Query::new_nearest(vec![
|
||||
vec![0.1, 0.2],
|
||||
vec![0.2, 0.1],
|
||||
vec![0.8, 0.9],
|
||||
]))
|
||||
.using("colbert")
|
||||
.limit(10u64)
|
||||
).await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.PrefetchQuery;
|
||||
import io.qdrant.client.grpc.Points.QueryPoints;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.queryAsync(
|
||||
QueryPoints.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.addPrefetch(
|
||||
PrefetchQuery.newBuilder()
|
||||
.addPrefetch(
|
||||
PrefetchQuery.newBuilder()
|
||||
.setQuery(nearest(1, 23, 45, 67)) // <------------- small byte vector
|
||||
.setUsing("mrl_byte")
|
||||
.setLimit(1000)
|
||||
.build())
|
||||
.setQuery(nearest(0.01f, 0.45f, 0.67f)) // <-- dense vector
|
||||
.setUsing("full")
|
||||
.setLimit(100)
|
||||
.build())
|
||||
.setQuery(
|
||||
nearest(
|
||||
new float[][] {
|
||||
{0.1f, 0.2f}, // <─┐
|
||||
{0.2f, 0.1f}, // < ├─ multi-vector
|
||||
{0.8f, 0.9f} // < ┘
|
||||
}))
|
||||
.setUsing("colbert")
|
||||
.setLimit(10)
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
prefetch: new List <PrefetchQuery> {
|
||||
new() {
|
||||
Prefetch = {
|
||||
new List <PrefetchQuery> {
|
||||
new() {
|
||||
Query = new float[] { 1, 23, 45, 67 }, // <------------- small byte vector
|
||||
Using = "mrl_byte",
|
||||
Limit = 1000
|
||||
},
|
||||
}
|
||||
},
|
||||
Query = new float[] {0.01f, 0.45f, 0.67f}, // <-- dense vector
|
||||
Using = "full",
|
||||
Limit = 100
|
||||
}
|
||||
},
|
||||
query: new float[][] {
|
||||
[0.1f, 0.2f], // <─┐
|
||||
[0.2f, 0.1f], // < ├─ multi-vector
|
||||
[0.8f, 0.9f] // < ┘
|
||||
},
|
||||
usingVector: "colbert",
|
||||
limit: 10
|
||||
);
|
||||
```
|
||||
|
||||
**Note:** *The multivector feature is not only useful for ColBERT; it can also be used in other ways.*</br>
|
||||
For instance, in e-commerce, you can use multi-vector to store multiple images of the same item. This serves as an alternative to the [group-by](/documentation/concepts/search/#grouping-api) method.
|
||||
|
||||
## Sparse Vectors Compression
|
||||
|
||||
In version 1.9, we introduced the `uint8` [vector datatype](/documentation/concepts/vectors/#datatypes) for sparse vectors, in order to support pre-quantized embeddings from companies like JinaAI and Cohere.
|
||||
This time, we are introducing a new datatype **for both sparse and dense vectors**, as well as a different way of **storing** these vectors.
|
||||
|
||||
**Datatype:** Sparse and dense vectors were previously represented in larger `float32` values, but now they can be turned to the `float16`. `float16` vectors have a lower precision compared to `float32`, which means that there is less numerical accuracy in the vector values - but this is negligible for practical use cases.
|
||||
|
||||
These vectors will use half the memory of regular vectors, which can significantly reduce the footprint of large vector datasets. Operations can be faster due to reduced memory bandwidth requirements and better cache utilization. This can lead to faster vector search operations, especially in memory-bound scenarios.
|
||||
|
||||
When creating a collection, you need to specify the `datatype` upfront:
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"vectors": {
|
||||
"size": 1024,
|
||||
"distance": "Cosine",
|
||||
"datatype": "float16"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```java
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||
import io.qdrant.client.grpc.Collections.Datatype;
|
||||
import io.qdrant.client.grpc.Collections.Distance;
|
||||
import io.qdrant.client.grpc.Collections.VectorParams;
|
||||
import io.qdrant.client.grpc.Collections.VectorsConfig;
|
||||
|
||||
QdrantClient client = new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.createCollectionAsync(
|
||||
CreateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setVectorsConfig(VectorsConfig.newBuilder()
|
||||
.setParams(VectorParams.newBuilder()
|
||||
.setSize(1024)
|
||||
.setDistance(Distance.Cosine)
|
||||
.setDatatype(Datatype.Float16)
|
||||
.build())
|
||||
.build())
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{CreateCollectionBuilder, Datatype, Distance, VectorParamsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client
|
||||
.create_collection(
|
||||
CreateCollectionBuilder::new("{collection_name}").vectors_config(
|
||||
VectorParamsBuilder::new(1024, Distance::Cosine).datatype(Datatype::Float16),
|
||||
),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
vectorsConfig: new VectorParams {
|
||||
Size = 1024,
|
||||
Distance = Distance.Cosine,
|
||||
Datatype = Datatype.Float16
|
||||
}
|
||||
);
|
||||
```
|
||||
|
||||
**Storage:** On the backend, we implemented bit packing to minimize the bits needed to store data, crucial for handling sparse vectors in applications like machine learning and data compression. For sparse vectors with mostly zeros, this focuses on storing only the indices and values of non-zero elements.
|
||||
|
||||
You will benefit from a more compact storage and higher processing efficiency. This can also lead to reduced dataset sizes for faster processing and lower storage costs in data compression.
|
||||
|
||||
## New Rust Client
|
||||
|
||||
Qdrant’s Rust client has been fully reshaped. It is now more accessible and
|
||||
easier to use. We have focused on putting together a minimalistic API interface.
|
||||
All operations and their types now use the builder pattern, providing an easy
|
||||
and extensible interface, preventing breakage with future updates. See the Rust
|
||||
[ColBERT query](#multivector-support) as great example. Additionally,
|
||||
Rust supports safe concurrent execution, which is crucial for handling multiple
|
||||
simultaneous requests efficiently.
|
||||
|
||||
Documentation got a significant improvement as well. It is much better organized
|
||||
and provides usage examples across the board. Everything links back to our main
|
||||
documentation, making it easier to navigate and find the information you need.
|
||||
|
||||
<p align="center">
|
||||
Visit our
|
||||
<a href="https://docs.rs/qdrant-client/1.10/qdrant_client/">client</a> and
|
||||
<a href="https://docs.rs/qdrant-client/1.10/qdrant_client/struct.Qdrant.html">operations</a> documentation
|
||||
</p>
|
||||
|
||||
## S3 Snapshot Storage
|
||||
Qdrant **Collections**, **Shards** and **Storage** can be backed up with [Snapshots](/documentation/concepts/snapshots/) and saved in case of data loss or other data transfer purposes. These snapshots can be quite large and the resources required to maintain them can result in higher costs. AWS S3 and other S3-compatible implementations like [min.io](https://min.io/) is a great low-cost alternative that can hold snapshots without incurring high costs. It is globally reliable, scalable and resistant to data loss.
|
||||
|
||||
You can configure S3 storage settings in the [config.yaml](https://github.com/qdrant/qdrant/blob/master/config/config.yaml), specifically with `snapshots_storage`.
|
||||
|
||||
For example, to use AWS S3:
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
snapshots_config:
|
||||
# Use 's3' to store snapshots on S3
|
||||
snapshots_storage: s3
|
||||
|
||||
s3_config:
|
||||
# Bucket name
|
||||
bucket: your_bucket_here
|
||||
|
||||
# Bucket region (e.g. eu-central-1)
|
||||
region: your_bucket_region_here
|
||||
|
||||
# Storage access key
|
||||
# Can be specified either here or in the `AWS_ACCESS_KEY_ID` environment variable.
|
||||
access_key: your_access_key_here
|
||||
|
||||
# Storage secret key
|
||||
# Can be specified either here or in the `AWS_SECRET_ACCESS_KEY` environment variable.
|
||||
secret_key: your_secret_key_here
|
||||
```
|
||||
|
||||
*Read more about [S3 snapshot storage](/documentation/concepts/snapshots/#s3) and [configuration](/documentation/guides/configuration/).*
|
||||
|
||||
This integration allows for a more convenient distribution of snapshots. Users of **any S3-compatible object storage** can now benefit from other platform services, such as automated workflows and disaster recovery options. S3's encryption and access control ensure secure storage and regulatory compliance. Additionally, S3 supports performance optimization through various storage classes and efficient data transfer methods, enabling quick and effective snapshot retrieval and management.
|
||||
|
||||
## Issues API
|
||||
Issues API notifies you about potential performance issues and misconfigurations. This powerful new feature allows users (such as database admins) to efficiently manage and track issues directly within the system, ensuring smoother operations and quicker resolutions.
|
||||
|
||||
You can find the Issues button in the top right. When you click the bell icon, a sidebar will open to show ongoing issues.
|
||||
|
||||

|
||||
|
||||
## Minor Improvements
|
||||
|
||||
- Pre-configure collection parameters; quantization, vector storage & replication factor - [#4299](https://github.com/qdrant/qdrant/pull/4299)
|
||||
|
||||
- Overwrite global optimizer configuration for collections. Lets you separate roles for indexing and searching within the single qdrant cluster - [#4317](https://github.com/qdrant/qdrant/pull/4317)
|
||||
|
||||
- Delta encoding and bitpacking compression for sparse vectors reduces memory consumption for sparse vectors by up to 75% - [#4253](https://github.com/qdrant/qdrant/pull/4253), [#4350](https://github.com/qdrant/qdrant/pull/4350)
|
||||
|
||||
@@ -0,0 +1,205 @@
|
||||
---
|
||||
title: "What is Vector Similarity? Understanding its Role in AI Applications."
|
||||
draft: false
|
||||
short_description: "An in-depth exploration of vector similarity and its applications in AI."
|
||||
description: "Discover the significance of vector similarity in AI applications and how our vector database revolutionizes similarity search technology for enhanced performance and accuracy."
|
||||
preview_image: /blog/what-is-vector-similarity/social_preview.png
|
||||
social_preview_image: /blog/what-is-vector-similarity/social_preview.png
|
||||
date: 2024-02-24T00:00:00-08:00
|
||||
author: Qdrant Team
|
||||
featured: false
|
||||
tags:
|
||||
- vector search
|
||||
- vector similarity
|
||||
- similarity search
|
||||
- embeddings
|
||||
---
|
||||
|
||||
# Understanding Vector Similarity: Powering Next-Gen AI Applications
|
||||
|
||||
A core function of a wide range of AI applications is to first understand the *meaning* behind a user query, and then provide *relevant* answers to the questions that the user is asking. With increasingly advanced interfaces and applications, this query can be in the form of language, or an image, an audio, video, or other forms of *unstructured* data.
|
||||
|
||||
On an ecommerce platform, a user can, for instance, try to find ‘clothing for a trek’, when they actually want results around ‘waterproof jackets’, or ‘winter socks’. Keyword, or full-text, or even synonym search would fail to provide any response to such a query. Similarly, on a music app, a user might be looking for songs that sound similar to an audio clip they have heard. Or, they might want to look up furniture that has a similar look as the one they saw on a trip.
|
||||
|
||||
## How Does Vector Similarity Work?
|
||||
So, how does an algorithm capture the essence of a user’s query, and then unearth results that are relevant?
|
||||
|
||||
At a high level, here’s how:
|
||||
|
||||
- Unstructured data is first converted into a numerical representation, known as vectors, using a deep-learning model. The goal here is to capture the ‘semantics’ or the key features of this data.
|
||||
- The vectors are then stored in a vector database, along with references to their original data.
|
||||
- When a user performs a query, the query is first converted into its vector representation using the same model. Then search is performed using a metric, to find other vectors which are closest to the query vector.
|
||||
- The list of results returned corresponds to the vectors that were found to be the closest.
|
||||
|
||||
At the heart of all such searches lies the concept of *vector similarity*, which gives us the ability to measure how closely related two data points are, how similar or dissimilar they are, or find other related data points.
|
||||
|
||||
In this document, we will deep-dive into the essence of vector similarity, study how vector similarity search is used in the context of AI, look at some real-world use cases and show you how to leverage the power of vector similarity and vector similarity search for building AI applications.
|
||||
|
||||
## **Understanding Vectors, Vector Spaces and Vector Similarity**
|
||||
|
||||
ML and deep learning models require numerical data as inputs to accomplish their tasks. Therefore, when working with non-numerical data, we first need to convert them into a numerical representation that captures the key features of that data. This is where vectors come in.
|
||||
|
||||
A vector is a set of numbers that represents data, which can be text, image, or audio, or any multidimensional data. Vectors reside in a high-dimensional space, the vector space, where each dimension captures a specific aspect or feature of the data.
|
||||
|
||||
{{< figure width=80% src=/blog/what-is-vector-similarity/working.png caption="Working" >}}
|
||||
|
||||
The number of dimensions of a vector can range from tens or hundreds to thousands, and each dimension is stored as the element of an array. Vectors are, therefore, an array of numbers of fixed length, and in their totality, they encode the key features of the data they represent.
|
||||
|
||||
Vector embeddings are created by AI models, a process known as vectorization. They are then stored in vector stores like Qdrant, which have the capability to rapidly search through vector space, and find similar or dissimilar vectors, cluster them, find related ones, or even the ones which are complete outliers.
|
||||
|
||||
For example, in the case of text data, “coat” and “jacket” have similar meaning, even though the words are completely different. Vector representations of these two words should be such that they lie close to each other in the vector space. The process of measuring their proximity in vector space is vector similarity.
|
||||
|
||||
Vector similarity, therefore, is a measure of how closely related two data points are in a vector space. It quantifies how alike or different two data points are based on their respective vector representations.
|
||||
|
||||
Suppose we have the words "king", "queen" and “apple”. Given a model, words with similar meanings have vectors that are close to each other in the vector space. Vector representations of “king” and “queen” would be, therefore, closer together than "king" and "apple", or “queen” and “apple” due to their semantic relationship. Vector similarity is how you calculate this.
|
||||
|
||||
An extremely powerful aspect of vectors is that they are not limited to representing just text, image or audio. In fact, vector representations can be created out of any kind of data. You can create vector representations of 3D models, for instance. Or for video clips, or molecular structures, or even [protein sequences](https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-019-3220-8).
|
||||
|
||||
There are several methodologies through which vectorization is performed. In creating vector representations of text, for example, the process involves analyzing the text for its linguistic elements using a transformer model. These models essentially learn to capture the essence of the text by dissecting its language components.
|
||||
|
||||
## **How Is Vector Similarity Calculated?**
|
||||
|
||||
There are several ways to calculate the similarity (or distance) between two vectors, which we call metrics. The most popular ones are:
|
||||
|
||||
**Dot Product**: Obtained by multiplying corresponding elements of the vectors and then summing those products. A larger dot product indicates a greater degree of similarity.
|
||||
|
||||
**Cosine Similarity**: Calculated using the dot product of the two vectors divided by the product of their magnitudes (norms). Cosine similarity of 1 implies that the vectors are perfectly aligned, while a value of 0 indicates no similarity. A value of -1 means they are diametrically opposed (or dissimilar).
|
||||
|
||||
**Euclidean Distance**: Assuming two vectors act like arrows in vector space, Euclidean distance calculates the length of the straight line connecting the heads of these two arrows. The smaller the Euclidean distance, the greater the similarity.
|
||||
|
||||
**Manhattan Distance**: Also known as taxicab distance, it is calculated as the total distance between the two vectors in a vector space, if you follow a grid-like path. The smaller the Manhattan distance, the greater the similarity.
|
||||
|
||||
{{< figure width=80% src=/blog/what-is-vector-similarity/products.png caption="Metrics" >}}
|
||||
|
||||
As a rule of thumb, the choice of the best similarity metric depends on how the vectors were encoded.
|
||||
|
||||
Of the four metrics, Cosine Similarity is the most popular.
|
||||
|
||||
## **The Significance of Vector Similarity**
|
||||
|
||||
Vector Similarity is vital in powering machine learning applications. By comparing the vector representation of a query to the vectors of all data points, vector similarity search algorithms can retrieve the most relevant vectors. This helps in building powerful similarity search and recommendation systems, and has numerous applications in image and text analysis, in natural language processing, and in other domains that deal with high-dimensional data.
|
||||
|
||||
Let’s look at some of the key ways in which vector similarity can be leveraged.
|
||||
|
||||
**Image Analysis**
|
||||
|
||||
Once images are converted to their vector representations, vector similarity can help create systems to identify, categorize, and compare them. This can enable powerful reverse image search, facial recognition systems, or can be used for object detection and classification.
|
||||
|
||||
**Text Analysis**
|
||||
|
||||
Vector similarity in text analysis helps in understanding and processing language data. Vectorized text can be used to build semantic search systems, or in document clustering, or plagiarism detection applications.
|
||||
|
||||
**Retrieval Augmented Generation (RAG)**
|
||||
|
||||
Vector similarity can help in representing and comparing linguistic features, from single words to entire documents. This can help build retrieval augmented generation (RAG) applications, where the data is retrieved based on user intent. It also enables nuanced language tasks such as sentiment analysis, synonym detection, language translation, and more.
|
||||
|
||||
**Recommender Systems**
|
||||
|
||||
By converting user preference vectors into item vectors from a dataset, vector similarity can help build semantic search and recommendation systems. This can be utilized in a range of domains such e-commerce or OTT services, where it can help in suggesting relevant products, movies or songs.
|
||||
|
||||
Due to its varied applications, vector similarity has become a critical component in AI tooling. However, implementing it at scale, and in production settings, poses some hard problems. Below we will discuss some of them and explore how Qdrant helps solve these challenges.
|
||||
|
||||
## **Challenges with Vector Similarity Search**
|
||||
|
||||
The biggest challenge in this area comes from what researchers call the "[curse of dimensionality](https://en.wikipedia.org/wiki/Curse_of_dimensionality)." Algorithms like k-d trees may work well for finding exact matches in low dimensions (in 2D or 3D space). However, when you jump to high-dimensional spaces (hundreds or thousands of dimensions, which is common with vector embeddings), these algorithms become impractical. Traditional search methods and OLTP or OLAP databases struggle to handle this curse of dimensionality efficiently.
|
||||
|
||||
This means that building production applications that leverage vector similarity involves navigating several challenges. Here are some of the key challenges to watch out for.
|
||||
|
||||
### Scalability
|
||||
|
||||
Various vector search algorithms were originally developed to handle datasets small enough to be accommodated entirely within the memory of a single computer.
|
||||
|
||||
However, in real-world production settings, the datasets can encompass billions of high-dimensional vectors. As datasets grow, the storage and computational resources required to maintain and search through vector space increases dramatically.
|
||||
|
||||
For building scalable applications, leveraging vector databases that allow for a distributed architecture and have the capabilities of sharding, partitioning and load balancing is crucial.
|
||||
|
||||
### Efficiency
|
||||
|
||||
As the number of dimensions in vectors increases, algorithms that work in lower dimensions become less effective in measuring true similarity. This makes finding nearest neighbors computationally expensive and inaccurate in high-dimensional space.
|
||||
|
||||
For efficient query processing, it is important to choose vector search systems which use indexing techniques that help speed up search through high-dimensional vector space, and reduce latency.
|
||||
|
||||
### Security
|
||||
|
||||
For real-world applications, vector databases frequently house privacy-sensitive data. This can encompass Personally Identifiable Information (PII) in customer records, intellectual property (IP) like proprietary documents, or specialized datasets subject to stringent compliance regulations.
|
||||
|
||||
For data security, the vector search system should offer features that prevent unauthorized access to sensitive information. Also, it should empower organizations to retain data sovereignty, ensuring their data complies with their own regulations and legal requirements, independent of the platform or the cloud provider.
|
||||
|
||||
These are some of the many challenges that developers face when attempting to leverage vector similarity in production applications.
|
||||
|
||||
To address these challenges head-on, we have made several design choices at Qdrant which help power vector search use-cases that go beyond simple CRUD applications.
|
||||
|
||||
## How Qdrant Solves Vector Similarity Search Challenges
|
||||
|
||||
Qdrant is a highly performant and scalable vector search system, developed ground up in Rust. Qdrant leverages Rust’s famed memory efficiency and performance. It supports horizontal scaling, sharding, and replicas, and includes security features like role-based authentication. Additionally, Qdrant can be deployed in various environments, including [hybrid cloud setups](/hybrid-cloud/).
|
||||
|
||||
Here’s how we have taken on some of the key challenges that vector search applications face in production.
|
||||
|
||||
### Efficiency
|
||||
|
||||
Our [choice of Rust](/articles/why-rust/) significantly contributes to the efficiency of Qdrant’s vector similarity search capabilities. Rust’s emphasis on safety and performance, without the need for a garbage collector, helps with better handling of memory and resources. Rust is renowned for its performance and safety features, particularly in concurrent processing, and we leverage it heavily to handle high loads efficiently.
|
||||
|
||||
Also, a key feature of Qdrant is that we leverage both vector and traditional indexes (payload index). This means that vector index helps speed up vector search, while traditional indexes help filter the results.
|
||||
|
||||
The vector index in Qdrant employs the Hierarchical Navigable Small World (HNSW) algorithm for Approximate Nearest Neighbor (ANN) searches, which is one of the fastest algorithms according to [benchmarks](https://github.com/erikbern/ann-benchmarks).
|
||||
|
||||
### Scalability
|
||||
|
||||
For massive datasets and demanding workloads, Qdrant supports [distributed deployment](/documentation/guides/distributed_deployment/) from v0.8.0. In this mode, you can set up a Qdrant cluster and distribute data across multiple nodes, enabling you to maintain high performance and availability even under increased workloads. Clusters support sharding and replication, and harness the Raft consensus algorithm to manage node coordination.
|
||||
|
||||
Qdrant also supports vector [quantization](/documentation/guides/quantization/) to reduce memory footprint and speed up vector similarity searches, making it very effective for large-scale applications where efficient resource management is critical.
|
||||
|
||||
There are three quantization strategies you can choose from - scalar quantization, binary quantization and product quantization - which will help you control the trade-off between storage efficiency, search accuracy and speed.
|
||||
|
||||
### Security
|
||||
|
||||
Qdrant offers several [security features](/documentation/guides/security/) to help protect data and access to the vector store:
|
||||
|
||||
- API Key Authentication: This helps secure API access to Qdrant Cloud with static or read-only API keys.
|
||||
- JWT-Based Access Control: You can also enable more granular access control through JSON Web Tokens (JWT), and opt for restricted access to specific parts of the stored data while building Role-Based Access Control (RBAC).
|
||||
- TLS Encryption: Additionally, you can enable TLS Encryption on data transmission to ensure security of data in transit.
|
||||
|
||||
To help with data sovereignty, Qdrant can be run in a [Hybrid Cloud](/hybrid-cloud/) setup. Hybrid Cloud allows for seamless deployment and management of the vector database across various environments, and integrates Kubernetes clusters into a unified managed service. You can manage these clusters via Qdrant Cloud’s UI while maintaining control over your infrastructure and resources.
|
||||
|
||||
## Optimizing Similarity Search Performance
|
||||
|
||||
In order to achieve top performance in vector similarity searches, Qdrant employs a number of other tactics in addition to the features discussed above.**FastEmbed**: Qdrant supports [FastEmbed](/articles/fastembed/), a lightweight Python library for generating fast and efficient text embeddings. FastEmbed uses quantized transformer models integrated with ONNX Runtime, and is significantly faster than traditional methods of embedding generation.
|
||||
|
||||
**Support for Dense and Sparse Vectors**: Qdrant supports both dense and sparse vector representations. While dense vectors are most common, you may encounter situations where the dataset contains a range of specialized domain-specific keywords. [Sparse vectors](/articles/sparse-vectors/) shine in such scenarios. Sparse vectors are vector representations of data where most elements are zero.
|
||||
|
||||
**Multitenancy**: Qdrant supports [multitenancy](/documentation/guides/multiple-partitions/) by allowing vectors to be partitioned by payload within a single collection. Using this you can isolate each user's data, and avoid creating separate collections for each user. In order to ensure indexing performance, Qdrant also offers ways to bypass the construction of a global vector index, so that you can index vectors for each user independently.
|
||||
|
||||
**IO Optimizations**: If your data doesn’t fit into the memory, it may require storing on disk. To [optimize disk IO performance](/articles/io_uring/), Qdrant offers io_uring based *async uring* storage backend on Linux-based systems. Benchmarks show that it drastically helps reduce operating system overhead from disk IO.
|
||||
|
||||
**Data Integrity**: To ensure data integrity, Qdrant handles data changes in two stages. First, changes are recorded in the Write-Ahead Log (WAL). Then, changes are applied to segments, which store both the latest and individual point versions. In case of abnormal shutdowns, data is restored from WAL.
|
||||
|
||||
**Integrations**: Qdrant has integrations with most popular frameworks, such as LangChain, LlamaIndex, Haystack, Apache Spark, FiftyOne, and more. Qdrant also has several [trusted partners](/blog/hybrid-cloud-launch-partners/) for Hybrid Cloud deployments, such as Oracle Cloud Infrastructure, Red Hat OpenShift, Vultr, OVHcloud, Scaleway, and DigitalOcean.
|
||||
|
||||
We regularly run [benchmarks](/benchmarks/) comparing Qdrant against other vector databases like Elasticsearch, Milvus, and Weaviate. Our benchmarks show that Qdrant consistently achieves the highest requests-per-second (RPS) and lowest latencies across various scenarios, regardless of the precision threshold and metric used.
|
||||
|
||||
## Real-World Use Cases
|
||||
|
||||
Vector similarity is increasingly being used in a wide range of [real-world applications](/use-cases/). In e-commerce, it powers recommendation systems by comparing user behavior vectors to product vectors. In social media, it can enhance content recommendations and user connections by analyzing user interaction vectors. In image-oriented applications, vector similarity search enables reverse image search, similar image clustering, and efficient content-based image retrieval. In healthcare, vector similarity helps in genetic research by comparing DNA sequence vectors to identify similarities and variations. The possibilities are endless.
|
||||
|
||||
A unique example of real-world application of vector similarity is how VISUA uses Qdrant. A leading computer vision platform, VISUA faced two key challenges. First, a rapid and accurate method to identify images and objects within them for reinforcement learning. Second, dealing with the scalability issues of their quality control processes due to the rapid growth in data volume. Their previous quality control, which relied on meta-information and manual reviews, was no longer scalable, which prompted the VISUA team to explore vector databases as a solution.
|
||||
|
||||
After exploring a number of vector databases, VISUA picked Qdrant as the solution of choice. Vector similarity search helped identify similarities and deduplicate large volumes of images, videos, and frames. This allowed VISUA to uniquely represent data and prioritize frames with anomalies for closer examination, which helped scale their quality assurance and reinforcement learning processes. Read our [case study](/blog/case-study-visua/) to learn more.
|
||||
|
||||
## Future Directions and Innovations
|
||||
|
||||
As real-world deployments of vector similarity search technology grows, there are a number of promising directions where this technology is headed.
|
||||
|
||||
We are developing more efficient indexing and search algorithms to handle increasing data volumes and high-dimensional data more effectively. Simultaneously, in case of dynamic datasets, we are pushing to enhance our handling of real-time updates and low-latency search capabilities.
|
||||
|
||||
Qdrant is one of the most secure vector stores out there. However, we are working on bringing more privacy-preserving techniques in vector search implementations to protect sensitive data.
|
||||
|
||||
We have just about witnessed the tip of the iceberg in terms of what vector similarity can achieve. If you are working on an interesting use-case that uses vector similarity, we would like to hear from you.
|
||||
|
||||
## Getting Started with Qdrant
|
||||
|
||||
Ready to implement vector similarity in your AI applications? Explore Qdrant's vector database to enhance your data retrieval and AI capabilities. For additional resources and documentation, visit:
|
||||
|
||||
- [Quick Start Guide](/documentation/quick-start/)
|
||||
- [Documentation](/documentation/)
|
||||
|
||||
We are always available on our [Discord channel](https://qdrant.to/discord) to answer any questions you might have. You can also sign up for our [newsletter](/subscribe/) to stay ahead of the curve.
|
||||
@@ -30,6 +30,10 @@ A [Payload](/documentation/concepts/payload/) describes information that you can
|
||||
|
||||
[Explore](/documentation/concepts/explore/) includes several APIs for exploring data in your collections.
|
||||
|
||||
## Hybrid Queries
|
||||
|
||||
[Hybrid Queries](/documentation/concepts/hybrid-queries/) combines multiple queries or performs them in more than one stage.
|
||||
|
||||
## Filtering
|
||||
|
||||
[Filtering](/documentation/concepts/filtering/) defines various database-style clauses, conditions, and more.
|
||||
|
||||
@@ -1262,7 +1262,6 @@ await client.GetCollectionInfoAsync("{collection_name}");
|
||||
```
|
||||
|
||||
</details>
|
||||
<br/>
|
||||
|
||||
If you insert the vectors into the collection, the `status` field may become
|
||||
`yellow` whilst it is optimizing. It will become `green` once all the points are
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -604,6 +604,136 @@ Unlike a dense vector index, a sparse vector index does not require a pre-define
|
||||
|
||||
**Note:** A sparse vector index only supports dot-product similarity searches. It does not support other distance metrics.
|
||||
|
||||
### IDF Modifier
|
||||
|
||||
*Available as of v1.10.0*
|
||||
|
||||
For many search algorithms, it is important to consider how often an item occurs in a collection.
|
||||
Intuitively speaking, the less frequently an item appears in a collection, the more important it is in a search.
|
||||
|
||||
This is also known as the Inverse Document Frequency (IDF). It is used in text search engines to rank search results based on the rarity of a word in a collection.
|
||||
|
||||
IDF depends on the currently stored documents and therefore can't be pre-computed in the sparse vectors in streaming inference mode.
|
||||
In order to support IDF in the sparse vector index, Qdrant provides an option to modify the sparse vector query with the IDF statistics automatically.
|
||||
|
||||
The only requirement is to enable the IDF modifier in the collection configuration:
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"sparse_vectors": {
|
||||
"text": {
|
||||
"modifier": "idf"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
sparse_vectors={
|
||||
"text": models.SparseVectorParams(
|
||||
modifier=models.Modifier.IDF,
|
||||
),
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient, Schemas } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.createCollection("{collection_name}", {
|
||||
sparse_vectors: {
|
||||
"text": {
|
||||
modifier: "idf"
|
||||
}
|
||||
}
|
||||
});
|
||||
```
|
||||
|
||||
|
||||
```rust
|
||||
use qdrant_client::qdrant::{
|
||||
CreateCollectionBuilder, SparseVectorParamsBuilder,
|
||||
Modifier, sparse_vectors_config::SparseVectorsConfigBuilder
|
||||
};
|
||||
use qdrant_client::qdrant::;
|
||||
use qdrant_client::Qdrant;
|
||||
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
let mut sparse_vectors_config = SparseVectorsConfigBuilder::default();
|
||||
|
||||
sparse_vectors_config.add_named_vector_params(
|
||||
"text",
|
||||
SparseVectorParamsBuilder::default()
|
||||
.modifier(Modifier::Idf),
|
||||
);
|
||||
|
||||
client
|
||||
.create_collection(
|
||||
CreateCollectionBuilder::new("{collection_name}")
|
||||
.sparse_vectors_config(sparse_vectors_config)
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
|
||||
```java
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||
import io.qdrant.client.grpc.Collections.Modifier;
|
||||
import io.qdrant.client.grpc.Collections.SparseVectorConfig;
|
||||
import io.qdrant.client.grpc.Collections.SparseVectorParams;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.createCollectionAsync(
|
||||
CreateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setSparseVectorsConfig(
|
||||
SparseVectorConfig.newBuilder()
|
||||
.putMap("text", SparseVectorParams.newBuilder().setModifier(Modifier.Idf).build()))
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
sparseVectorsConfig: ("text", new SparseVectorParams {
|
||||
Modifier = Modifier.Idf,
|
||||
})
|
||||
);
|
||||
```
|
||||
|
||||
Qdrant uses the following formula to calculate the IDF modifier:
|
||||
|
||||
$$
|
||||
\text{IDF}(q_i) = \ln \left(\frac{N - n(q_i) + 0.5}{n(q_i) + 0.5}+1\right)
|
||||
$$
|
||||
|
||||
Where:
|
||||
|
||||
- `N` is the total number of documents in the collection.
|
||||
- `n` is the number of documents containing non-zero values for the given vector element.
|
||||
|
||||
## Filtrable Index
|
||||
|
||||
Separately, a payload index and a vector index cannot solve the problem of search using the filter completely.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: Payload
|
||||
weight: 40
|
||||
weight: 45
|
||||
aliases:
|
||||
- ../payload
|
||||
---
|
||||
|
||||
@@ -8,7 +8,18 @@ aliases:
|
||||
# Points
|
||||
|
||||
The points are the central entity that Qdrant operates with.
|
||||
A point is a record consisting of a vector and an optional [payload](../payload/).
|
||||
A point is a record consisting of a [vector](../vectors/) and an optional [payload](../payload/).
|
||||
|
||||
It looks like this:
|
||||
|
||||
```json
|
||||
// This is a simple point
|
||||
{
|
||||
"id": 129,
|
||||
"vector": [0.1, 0.2, 0.3, 0.4],
|
||||
"payload": {"color": "red"},
|
||||
}
|
||||
```
|
||||
|
||||
You can search among the points grouped in one [collection](../collections/) based on vector similarity.
|
||||
This procedure is described in more detail in the [search](../search/) and [filtering](../filtering/) sections.
|
||||
@@ -20,40 +31,6 @@ At the first stage, the operation is written to the Write-ahead-log.
|
||||
|
||||
After this moment, the service will not lose the data, even if the machine loses power supply.
|
||||
|
||||
## Awaiting result
|
||||
|
||||
If the API is called with the `&wait=false` parameter, or if it is not explicitly specified, the client will receive an acknowledgment of receiving data:
|
||||
|
||||
```json
|
||||
{
|
||||
"result": {
|
||||
"operation_id": 123,
|
||||
"status": "acknowledged"
|
||||
},
|
||||
"status": "ok",
|
||||
"time": 0.000206061
|
||||
}
|
||||
```
|
||||
|
||||
This response does not mean that the data is available for retrieval yet. This
|
||||
uses a form of eventual consistency. It may take a short amount of time before it
|
||||
is actually processed as updating the collection happens in the background. In
|
||||
fact, it is possible that such request eventually fails.
|
||||
If inserting a lot of vectors, we also recommend using asynchronous requests to take advantage of pipelining.
|
||||
|
||||
If the logic of your application requires a guarantee that the vector will be available for searching immediately after the API responds, then use the flag `?wait=true`.
|
||||
In this case, the API will return the result only after the operation is finished:
|
||||
|
||||
```json
|
||||
{
|
||||
"result": {
|
||||
"operation_id": 0,
|
||||
"status": "completed"
|
||||
},
|
||||
"status": "ok",
|
||||
"time": 0.000206061
|
||||
}
|
||||
```
|
||||
|
||||
## Point IDs
|
||||
|
||||
@@ -306,6 +283,26 @@ await client.UpsertAsync(
|
||||
|
||||
are both possible.
|
||||
|
||||
## Vectors
|
||||
|
||||
Each point in qdrant may have one or more vectors.
|
||||
Vectors are the central component of the Qdrant architecture,
|
||||
qdrant relies on different types of vectors to provide different types of data exploration and search.
|
||||
|
||||
Here is a list of supported vector types:
|
||||
|
||||
|||
|
||||
|-|-|
|
||||
| Dense Vectors | A regular vectors, generated by majority of the embedding models. |
|
||||
| Sparse Vectors | Vectors with no fixed length, but only a few non-zero elements. <br> Useful for exact token match and collaborative filtering recommendations. |
|
||||
| MultiVectors | Matrices of numbers with fixed length but variable height. <br> Usually obtained from late interraction models like ColBERT. |
|
||||
|
||||
It is possible to attach more than one type of vector to a single point.
|
||||
In Qdrant we call it Named Vectors.
|
||||
|
||||
Read more about vector types, how they are stored and optimized in the [vectors](../vectors/) section.
|
||||
|
||||
|
||||
## Upload points
|
||||
|
||||
To optimize performance, Qdrant supports batch loading of points. I.e., you can load several points into the service in one API call.
|
||||
@@ -2306,3 +2303,39 @@ client
|
||||
|
||||
To batch many points with a single operation type, please use batching
|
||||
functionality in that operation directly.
|
||||
|
||||
|
||||
## Awaiting result
|
||||
|
||||
If the API is called with the `&wait=false` parameter, or if it is not explicitly specified, the client will receive an acknowledgment of receiving data:
|
||||
|
||||
```json
|
||||
{
|
||||
"result": {
|
||||
"operation_id": 123,
|
||||
"status": "acknowledged"
|
||||
},
|
||||
"status": "ok",
|
||||
"time": 0.000206061
|
||||
}
|
||||
```
|
||||
|
||||
This response does not mean that the data is available for retrieval yet. This
|
||||
uses a form of eventual consistency. It may take a short amount of time before it
|
||||
is actually processed as updating the collection happens in the background. In
|
||||
fact, it is possible that such request eventually fails.
|
||||
If inserting a lot of vectors, we also recommend using asynchronous requests to take advantage of pipelining.
|
||||
|
||||
If the logic of your application requires a guarantee that the vector will be available for searching immediately after the API responds, then use the flag `?wait=true`.
|
||||
In this case, the API will return the result only after the operation is finished:
|
||||
|
||||
```json
|
||||
{
|
||||
"result": {
|
||||
"operation_id": 0,
|
||||
"status": "completed"
|
||||
},
|
||||
"status": "ok",
|
||||
"time": 0.000206061
|
||||
}
|
||||
```
|
||||
@@ -11,13 +11,171 @@ Searching for the nearest vectors is at the core of many representational learni
|
||||
Modern neural networks are trained to transform objects into vectors so that objects close in the real world appear close in vector space.
|
||||
It could be, for example, texts with similar meanings, visually similar pictures, or songs of the same genre.
|
||||
|
||||

|
||||
|
||||
{{< figure src="/docs/encoders.png" caption="This is how vector similarity works" width="70%" >}}
|
||||
|
||||
## Query API
|
||||
|
||||
*Available as of v1.10.0*
|
||||
|
||||
Qdrant provides a single interface for all kinds of search and exploration requests - the `Query API`.
|
||||
Here is a reference list of what kind of queries you can perform with the `Query API` in Qdrant:
|
||||
|
||||
Depending on the `query` parameter, Qdrant might prefer different strategies for the search.
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| Nearest Neighbors Search | Vector Similarity Search, also known as k-NN |
|
||||
| Search By Id | Search by an already stored vector - skip embedding model inference |
|
||||
| [Recommendations](../explore/#recommendation-api) | Provide positive and negative examples |
|
||||
| [Discovery Search](../explore/#discovery-api) | Guide the search using context as a one-shot training set |
|
||||
| [Scroll](../points/#scroll-points) | Get all points with optional filtering |
|
||||
| [Order By](../hybrid-queries/#re-ranking-with-stored-values) | Order points by payload key |
|
||||
| [Hybrid Search](../hybrid-queries/#hybrid-search) | Combine multiple queries to get better results |
|
||||
| [Multi-Stage Search](../hybrid-queries/#multi-stage-queries) | Optimize performance for large embeddings |
|
||||
|
||||
|
||||
**Nearest Neighbors Search**
|
||||
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"query": [0.2, 0.1, 0.9, 0.7] // <--- Dense vector
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query=[0.2, 0.1, 0.9, 0.7], # <--- Dense vector
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.query("{collection_name}", {
|
||||
query: [0.2, 0.1, 0.9, 0.7], // <--- Dense vector
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{Condition, Filter, Query, QueryPointsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client
|
||||
.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.query(Query::new_nearest(vec![0.2, 0.1, 0.9, 0.7]))
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import java.util.List;
|
||||
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.QueryPoints;
|
||||
|
||||
QdrantClient client = new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client.queryAsync(QueryPoints.newBuilder()
|
||||
.setCollectionName("{collectionName}")
|
||||
.setQuery(nearest(List.of(0.2f, 0.1f, 0.9f, 0.7f)))
|
||||
.build()).get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
query: new float[] { 0.2f, 0.1f, 0.9f, 0.7f }
|
||||
);
|
||||
```
|
||||
|
||||
**Search By Id**
|
||||
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"query": "43cf51e2-8777-4f52-bc74-c2cbde0c8b04" // <--- point id
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query="43cf51e2-8777-4f52-bc74-c2cbde0c8b04", # <--- point id
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.query("{collection_name}", {
|
||||
query: '43cf51e2-8777-4f52-bc74-c2cbde0c8b04', // <--- point id
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{Condition, Filter, PointId, Query, QueryPointsBuilder};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client
|
||||
.query(
|
||||
QueryPointsBuilder::new("{collection_name}")
|
||||
.query(Query::new_nearest(PointId::new("43cf51e2-8777-4f52-bc74-c2cbde0c8b04")))
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import java.util.UUID;
|
||||
|
||||
import static io.qdrant.client.QueryFactory.nearest;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.QueryPoints;
|
||||
|
||||
QdrantClient client = new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client.queryAsync(QueryPoints.newBuilder()
|
||||
.setCollectionName("{collectionName}")
|
||||
.setQuery(nearest(UUID.fromString("43cf51e2-8777-4f52-bc74-c2cbde0c8b04")))
|
||||
.build()).get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
query: Guid.Parse("43cf51e2-8777-4f52-bc74-c2cbde0c8b04")
|
||||
);
|
||||
```
|
||||
|
||||
## Metrics
|
||||
|
||||
There are many ways to estimate the similarity of vectors with each other.
|
||||
In Qdrant terms, these ways are called metrics.
|
||||
The choice of metric depends on vectors obtaining and, in particular, on the method of neural network encoder training.
|
||||
The choice of metric depends on the vectors obtained and, in particular, on the neural network encoder training method.
|
||||
|
||||
Qdrant supports these most popular types of metrics:
|
||||
|
||||
@@ -37,22 +195,8 @@ It happens only once for each vector.
|
||||
The second step is the comparison of vectors.
|
||||
In this case, it becomes equivalent to dot production - a very fast operation due to SIMD.
|
||||
|
||||
## Query planning
|
||||
|
||||
Depending on the filter used in the search - there are several possible scenarios for query execution.
|
||||
Qdrant chooses one of the query execution options depending on the available indexes, the complexity of the conditions and the cardinality of the filtering result.
|
||||
This process is called query planning.
|
||||
|
||||
The strategy selection process relies heavily on heuristics and can vary from release to release.
|
||||
However, the general principles are:
|
||||
|
||||
* planning is performed for each segment independently (see [storage](../storage/) for more information about segments)
|
||||
* prefer a full scan if the amount of points is below a threshold
|
||||
* estimate the cardinality of a filtered result before selecting a strategy
|
||||
* retrieve points using payload index (see [indexing](../indexing/)) if cardinality is below threshold
|
||||
* use filterable vector index if the cardinality is above a threshold
|
||||
|
||||
You can adjust the threshold using a [configuration file](https://github.com/qdrant/qdrant/blob/master/config/config.yaml), as well as independently for each collection.
|
||||
Depending on the query configuration, Qdrant might prefer different strategies for the search.
|
||||
Read more about it in the [query planning](#query-planning) section.
|
||||
|
||||
## Search API
|
||||
|
||||
@@ -1486,3 +1630,21 @@ The looked up result will show up under `lookup` in each group.
|
||||
```
|
||||
|
||||
Since the lookup is done by matching directly with the point id, any group id that is not an existing (and valid) point id in the lookup collection will be ignored, and the `lookup` field will be empty.
|
||||
|
||||
|
||||
## Query planning
|
||||
|
||||
Depending on the filter used in the search - there are several possible scenarios for query execution.
|
||||
Qdrant chooses one of the query execution options depending on the available indexes, the complexity of the conditions and the cardinality of the filtering result.
|
||||
This process is called query planning.
|
||||
|
||||
The strategy selection process relies heavily on heuristics and can vary from release to release.
|
||||
However, the general principles are:
|
||||
|
||||
* planning is performed for each segment independently (see [storage](../storage/) for more information about segments)
|
||||
* prefer a full scan if the amount of points is below a threshold
|
||||
* estimate the cardinality of a filtered result before selecting a strategy
|
||||
* retrieve points using payload index (see [indexing](../indexing/)) if cardinality is below threshold
|
||||
* use filterable vector index if the cardinality is above a threshold
|
||||
|
||||
You can adjust the threshold using a [configuration file](https://github.com/qdrant/qdrant/blob/master/config/config.yaml), as well as independently for each collection.
|
||||
|
||||
@@ -15,29 +15,6 @@ This feature can be used to archive data or easily replicate an existing deploym
|
||||
|
||||
For a step-by-step guide on how to use snapshots, see our [tutorial](/documentation/tutorials/create-snapshot/).
|
||||
|
||||
## Store snapshots
|
||||
|
||||
The target directory used to store generated snapshots is controlled through the [configuration](../../guides/configuration/) or using the ENV variable: `QDRANT__STORAGE__SNAPSHOTS_PATH=./snapshots`.
|
||||
|
||||
You can set the snapshots storage directory from the [config.yaml](https://github.com/qdrant/qdrant/blob/master/config/config.yaml) file. If no value is given, default is `./snapshots`.
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
# Specify where you want to store snapshots.
|
||||
snapshots_path: ./snapshots
|
||||
```
|
||||
|
||||
*Available as of v1.3.0*
|
||||
|
||||
While a snapshot is being created, temporary files are by default placed in the configured storage directory.
|
||||
This location may have limited capacity or be on a slow network-attached disk. You may specify a separate location for temporary files:
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
# Where to store temporary files
|
||||
temp_path: /tmp
|
||||
```
|
||||
|
||||
## Create snapshot
|
||||
|
||||
<aside role="status">If you work with a distributed deployment, you have to create snapshots for each node separately. A single snapshot will contain only the data stored on the node on which the snapshot was created.</aside>
|
||||
@@ -527,3 +504,68 @@ For example:
|
||||
```bash
|
||||
./qdrant --storage-snapshot /snapshots/full-snapshot-2022-07-18-11-20-51.snapshot
|
||||
```
|
||||
|
||||
## Storage
|
||||
|
||||
Created, uploaded and recovered snapshots are stored as `.snapshot` files. By
|
||||
default, they're stored on the [local file system](#local-file-system). You may
|
||||
also configure to use an [S3 storage](#s3) service for them.
|
||||
|
||||
### Local file system
|
||||
|
||||
By default, snapshots are stored at `./snapshots` or at `/qdrant/snapshots` when
|
||||
using our Docker image.
|
||||
|
||||
The target directory can be controlled through the [configuration](../../guides/configuration/):
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
# Specify where you want to store snapshots.
|
||||
snapshots_path: ./snapshots
|
||||
```
|
||||
|
||||
Alternatively you may use the environment variable `QDRANT__STORAGE__SNAPSHOTS_PATH=./snapshots`.
|
||||
|
||||
*Available as of v1.3.0*
|
||||
|
||||
While a snapshot is being created, temporary files are placed in the configured
|
||||
storage directory by default. In case of limited capacity or a slow
|
||||
network attached disk, you can specify a separate location for temporary files:
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
# Where to store temporary files
|
||||
temp_path: /tmp
|
||||
```
|
||||
|
||||
### S3
|
||||
|
||||
*Available as of v1.10.0*
|
||||
|
||||
Rather than storing snapshots on the local file system, you may also configure
|
||||
to store snapshots in an S3-compatible storage service. To enable this, you must
|
||||
configure it in the [configuration](../../guides/configuration/) file.
|
||||
|
||||
For example, to configure for AWS S3:
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
snapshots_config:
|
||||
# Use 's3' to store snapshots on S3
|
||||
snapshots_storage: s3
|
||||
|
||||
s3_config:
|
||||
# Bucket name
|
||||
bucket: your_bucket_here
|
||||
|
||||
# Bucket region (e.g. eu-central-1)
|
||||
region: your_bucket_region_here
|
||||
|
||||
# Storage access key
|
||||
# Can be specified either here or in the `AWS_ACCESS_KEY_ID` environment variable.
|
||||
access_key: your_access_key_here
|
||||
|
||||
# Storage secret key
|
||||
# Can be specified either here or in the `AWS_SECRET_ACCESS_KEY` environment variable.
|
||||
secret_key: your_secret_key_here
|
||||
```
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -7,6 +7,7 @@ weight: 33
|
||||
| ------------------------------------- | ---------------------------------------------------------------------------------------------------- |
|
||||
| [Airbyte](./airbyte/) | Data integration platform specialising in ELT pipelines. |
|
||||
| [Airflow](./airflow/) | Platform designed for developing, scheduling, and monitoring batch-oriented workflows. |
|
||||
| [Apify](./apify/) | Platform to build web scrapers and automate web browser tasks. |
|
||||
| [AutoGen](./autogen/) | Framework from Microsoft building LLM applications using multiple conversational agents. |
|
||||
| [Bubble](./bubble) | Development platform for application development with a no-code interface |
|
||||
| [Canopy](./canopy/) | Framework from Pinecone for building RAG applications using LLMs and knowledge bases. |
|
||||
|
||||
@@ -0,0 +1,82 @@
|
||||
---
|
||||
title: Apify
|
||||
weight: 3600
|
||||
---
|
||||
|
||||
# Apify
|
||||
|
||||
[Apify](https://apify.com/) is a web scraping and browser automation platform featuring an [app store](https://apify.com/store) with over 1,500 pre-built micro-apps known as Actors. These serverless cloud programs, which are essentially dockers under the hood, are designed for various web automation applications, including data collection.
|
||||
|
||||
One such Actor, built especially for AI and RAG applications, is [Website Content Crawler](https://apify.com/apify/website-content-crawler).
|
||||
|
||||
It's ideal for this purpose because it has built-in HTML processing and data-cleaning functions. That means you can easily remove fluff, duplicates, and other things on a web page that aren't relevant, and provide only the necessary data to the language model.
|
||||
|
||||
The Markdown can then be used to feed Qdrant to train AI models or supply them with fresh web content.
|
||||
|
||||
Qdrant is available as an [official integration](https://apify.com/apify/qdrant-integration) to load Apify datasets into a collection.
|
||||
|
||||
You can refer to the [Apify documentation](https://docs.apify.com/platform/integrations/qdrant) to set up the integration via the Apify UI.
|
||||
|
||||
## Programmatic Usage
|
||||
|
||||
Apify also supports programmatic access to integrations via the [Apify Python SDK](https://docs.apify.com/sdk/python/).
|
||||
|
||||
1. Install the Apify Python SDK by running the following command:
|
||||
|
||||
```sh
|
||||
pip install apify-client
|
||||
```
|
||||
|
||||
2. Create a Python script and import all the necessary modules:
|
||||
|
||||
```python
|
||||
from apify_client import ApifyClient
|
||||
|
||||
APIFY_API_TOKEN = "YOUR-APIFY-TOKEN"
|
||||
OPENAI_API_KEY = "YOUR-OPENAI-API-KEY"
|
||||
# COHERE_API_KEY = "YOUR-COHERE-API-KEY"
|
||||
|
||||
QDRANT_URL = "YOUR-QDRANT-URL"
|
||||
QDRANT_API_KEY = "YOUR-QDRANT-API-KEY"
|
||||
|
||||
client = ApifyClient(APIFY_API_TOKEN)
|
||||
```
|
||||
|
||||
3. Call the [Website Content Crawler](https://apify.com/apify/website-content-crawler) Actor to crawl the Qdrant documentation and extract text content from the web pages:
|
||||
|
||||
```python
|
||||
actor_call = client.actor("apify/website-content-crawler").call(
|
||||
run_input={"startUrls": [{"url": "https://qdrant.tech/documentation/"}]}
|
||||
)
|
||||
```
|
||||
|
||||
4. Call the Qdrant integration and store all data in the Qdrant Vector Database:
|
||||
|
||||
```python
|
||||
qdrant_integration_inputs = {
|
||||
"qdrantUrl": QDRANT_URL,
|
||||
"qdrantApiKey": QDRANT_API_KEY,
|
||||
"qdrantCollectionName": "apify",
|
||||
"qdrantAutoCreateCollection": True,
|
||||
"datasetId": actor_call["defaultDatasetId"],
|
||||
"datasetFields": ["text"],
|
||||
"enableDeltaUpdates": True,
|
||||
"deltaUpdatesPrimaryDatasetFields": ["url"],
|
||||
"expiredObjectDeletionPeriodDays": 30,
|
||||
"embeddingsProvider": "OpenAI", # "Cohere"
|
||||
"embeddingsApiKey": OPENAI_API_KEY,
|
||||
"performChunking": True,
|
||||
"chunkSize": 1000,
|
||||
"chunkOverlap": 0,
|
||||
}
|
||||
actor_call = client.actor("apify/qdrant-integration").call(run_input=qdrant_integration_inputs)
|
||||
|
||||
```
|
||||
|
||||
Upon running the script, the data from <https://qdrant.tech/documentation/> will be scraped, transformed into vector embeddings and stored in the Qdrant collection.
|
||||
|
||||
## Further Reading
|
||||
|
||||
- Apify [Documentation](https://docs.apify.com/)
|
||||
- Apify [Templates](https://apify.com/templates)
|
||||
- Integration [Source Code](https://github.com/apify/actor-vector-database-integrations)
|
||||
@@ -48,5 +48,5 @@ The data is persisted at the default MemGPT storage directory.
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [MemGPT Examples][https://github.com/cpacker/MemGPT/tree/main/examples]
|
||||
- [MemGPT Examples](https://github.com/cpacker/MemGPT/tree/main/examples)
|
||||
- [MemGPT Documentation](https://memgpt.readme.io/docs/index).
|
||||
|
||||
@@ -98,6 +98,15 @@ storage:
|
||||
# Where to store snapshots
|
||||
snapshots_path: ./snapshots
|
||||
|
||||
snapshots_config:
|
||||
# "local" or "s3" - where to store snapshots
|
||||
snapshots_storage: local
|
||||
# s3_config:
|
||||
# bucket: ""
|
||||
# region: ""
|
||||
# access_key: ""
|
||||
# secret_key: ""
|
||||
|
||||
# Where to store temporary files
|
||||
# If null, temporary snapshot are stored in: storage/snapshots_temp/
|
||||
temp_path: null
|
||||
@@ -130,10 +139,17 @@ storage:
|
||||
performance:
|
||||
# Number of parallel threads used for search operations. If 0 - auto selection.
|
||||
max_search_threads: 0
|
||||
# Max total number of threads, which can be used for running optimization processes across all collections.
|
||||
# Note: Each optimization thread will also use `max_indexing_threads` for index building.
|
||||
# So total number of threads used for optimization will be `max_optimization_threads * max_indexing_threads`
|
||||
max_optimization_threads: 1
|
||||
|
||||
# Max number of threads (jobs) for running optimizations across all collections, each thread runs one job.
|
||||
# If 0 - have no limit and choose dynamically to saturate CPU.
|
||||
# Note: each optimization job will also use `max_indexing_threads` threads by itself for index building.
|
||||
max_optimization_threads: 0
|
||||
|
||||
# CPU budget, how many CPUs (threads) to allocate for an optimization job.
|
||||
# If 0 - auto selection, keep 1 or more CPUs unallocated depending on CPU size
|
||||
# If negative - subtract this number of CPUs from the available CPUs.
|
||||
# If positive - use this exact number of CPUs.
|
||||
optimizer_cpu_budget: 0
|
||||
|
||||
# Prevent DDoS of too many concurrent updates in distributed mode.
|
||||
# One external update usually triggers multiple internal updates, which breaks internal
|
||||
@@ -141,6 +157,18 @@ storage:
|
||||
# If null - auto selection.
|
||||
update_rate_limit: null
|
||||
|
||||
# Limit for number of incoming automatic shard transfers per collection on this node, does not affect user-requested transfers.
|
||||
# The same value should be used on all nodes in a cluster.
|
||||
# Default is to allow 1 transfer.
|
||||
# If null - allow unlimited transfers.
|
||||
#incoming_shard_transfers_limit: 1
|
||||
|
||||
# Limit for number of outgoing automatic shard transfers per collection on this node, does not affect user-requested transfers.
|
||||
# The same value should be used on all nodes in a cluster.
|
||||
# Default is to allow 1 transfer.
|
||||
# If null - allow unlimited transfers.
|
||||
#outgoing_shard_transfers_limit: 1
|
||||
|
||||
optimizers:
|
||||
# The minimal fraction of deleted vectors in a segment, required to perform segment optimization
|
||||
deleted_threshold: 0.2
|
||||
@@ -186,33 +214,76 @@ storage:
|
||||
# Interval between forced flushes.
|
||||
flush_interval_sec: 5
|
||||
|
||||
# Max number of threads, which can be used for optimization per collection.
|
||||
# Note: Each optimization thread will also use `max_indexing_threads` for index building.
|
||||
# So total number of threads used for optimization will be `max_optimization_threads * max_indexing_threads`
|
||||
# If `max_optimization_threads = 0`, optimization will be disabled.
|
||||
max_optimization_threads: 1
|
||||
# Max number of threads (jobs) for running optimizations per shard.
|
||||
# Note: each optimization job will also use `max_indexing_threads` threads by itself for index building.
|
||||
# If null - have no limit and choose dynamically to saturate CPU.
|
||||
# If 0 - no optimization threads, optimizations will be disabled.
|
||||
max_optimization_threads: null
|
||||
|
||||
# This section has the same options as 'optimizers' above. All values specified here will overwrite the collections
|
||||
# optimizers configs regardless of the config above and the options specified at collection creation.
|
||||
#optimizers_overwrite:
|
||||
# deleted_threshold: 0.2
|
||||
# vacuum_min_vector_number: 1000
|
||||
# default_segment_number: 0
|
||||
# max_segment_size_kb: null
|
||||
# memmap_threshold_kb: null
|
||||
# indexing_threshold_kb: 20000
|
||||
# flush_interval_sec: 5
|
||||
# max_optimization_threads: null
|
||||
|
||||
# Default parameters of HNSW Index. Could be overridden for each collection or named vector individually
|
||||
hnsw_index:
|
||||
# Number of edges per node in the index graph. Larger the value - more accurate the search, more space required.
|
||||
m: 16
|
||||
|
||||
# Number of neighbours to consider during the index building. Larger the value - more accurate the search, more time required to build index.
|
||||
ef_construct: 100
|
||||
|
||||
# Minimal size (in KiloBytes) of vectors for additional payload-based indexing.
|
||||
# If payload chunk is smaller than `full_scan_threshold_kb` additional indexing won't be used -
|
||||
# in this case full-scan search should be preferred by query planner and additional indexing is not required.
|
||||
# Note: 1Kb = 1 vector of size 256
|
||||
full_scan_threshold_kb: 10000
|
||||
# Number of parallel threads used for background index building. If 0 - auto selection.
|
||||
|
||||
# Number of parallel threads used for background index building.
|
||||
# If 0 - automatically select.
|
||||
# Best to keep between 8 and 16 to prevent likelihood of building broken/inefficient HNSW graphs.
|
||||
# On small CPUs, less threads are used.
|
||||
max_indexing_threads: 0
|
||||
|
||||
# Store HNSW index on disk. If set to false, index will be stored in RAM. Default: false
|
||||
on_disk: false
|
||||
|
||||
# Custom M param for hnsw graph built for payload index. If not set, default M will be used.
|
||||
payload_m: null
|
||||
|
||||
# Default shard transfer method to use if none is defined.
|
||||
# If null - don't have a shard transfer preference, choose automatically.
|
||||
# If stream_records, snapshot or wal_delta - prefer this specific method.
|
||||
# More info: https://qdrant.tech/documentation/guides/distributed_deployment/#shard-transfer-method
|
||||
shard_transfer_method: null
|
||||
|
||||
# Default parameters for collections
|
||||
collection:
|
||||
# Number of replicas of each shard that network tries to maintain
|
||||
replication_factor: 1
|
||||
|
||||
# How many replicas should apply the operation for us to consider it successful
|
||||
write_consistency_factor: 1
|
||||
|
||||
# Default parameters for vectors.
|
||||
vectors:
|
||||
# Whether vectors should be stored in memory or on disk.
|
||||
on_disk: null
|
||||
|
||||
# shard_number_per_node: 1
|
||||
|
||||
# Default quantization configuration.
|
||||
# More info: https://qdrant.tech/documentation/guides/quantization
|
||||
quantization: null
|
||||
|
||||
service:
|
||||
|
||||
# Maximum size of POST data in a single request in megabytes
|
||||
max_request_size_mb: 32
|
||||
|
||||
@@ -253,7 +324,7 @@ service:
|
||||
#
|
||||
# Uncomment to enable.
|
||||
# api_key: your_secret_api_key_here
|
||||
|
||||
|
||||
# Set an api-key for read-only operations.
|
||||
# If set, all requests must include a header with the api-key.
|
||||
# example header: `api-key: <API-KEY>`
|
||||
@@ -265,6 +336,12 @@ service:
|
||||
# Uncomment to enable.
|
||||
# read_only_api_key: your_secret_read_only_api_key_here
|
||||
|
||||
# Uncomment to enable JWT Role Based Access Control (RBAC).
|
||||
# If enabled, you can generate JWT tokens with fine-grained rules for access control.
|
||||
# Use generated token instead of API key.
|
||||
#
|
||||
# jwt_rbac: true
|
||||
|
||||
cluster:
|
||||
# Use `enabled: true` to run Qdrant in distributed deployment mode
|
||||
enabled: false
|
||||
@@ -331,4 +408,4 @@ WARN - storage.hnsw_index.m: value 1 invalid, must be from 4 to 10000
|
||||
```
|
||||
|
||||
The server will continue to operate. Any validation errors should be fixed as
|
||||
soon as possible though to prevent problematic behavior.
|
||||
soon as possible though to prevent problematic behavior.
|
||||
|
||||
@@ -21,6 +21,6 @@ These tutorials demonstrate different ways you can build vector search into your
|
||||
| [Asynchronous API](../tutorials/async-api/) | Communicate with Qdrant server asynchronously with Python SDK. | Qdrant, Python |
|
||||
| [Create Dataset Snapshots](../tutorials/create-snapshot/) | Turn a dataset into a snapshot by exporting it from a collection. | Qdrant |
|
||||
| [Load HuggingFace Dataset](../tutorials/huggingface-datasets/) | Load a Hugging Face dataset to Qdrant | Qdrant, Python, datasets |
|
||||
| [Measure retrieval quality](../tutorials/retrieval-quality/) | Measure and fine-tune the retrieval quality | Qdrant, Python, datasets |
|
||||
| [Use semantic search to navigate your codebase](../tutorials/code-search/) | Implement semantic search application for code search task | Qdrant, Python, sentence-transformers, Jina |
|
||||
|
||||
| [Measure Retrieval Quality](../tutorials/retrieval-quality/) | Measure and fine-tune the retrieval quality | Qdrant, Python, datasets |
|
||||
| [Search Through Code](../tutorials/code-search/) | Implement semantic search application for code search tasks | Qdrant, Python, sentence-transformers, Jina |
|
||||
| [Setup Collaborative Filtering](../tutorials/collaborative-filtering/) | Implement a collaborative filtering system for recommendation engines | Qdrant|
|
||||
|
||||
@@ -0,0 +1,277 @@
|
||||
---
|
||||
title: Collaborative filtering
|
||||
short_description: "Build an effective movie recommendation system using collaborative filtering and Qdrant's similarity search."
|
||||
description: "Build an effective movie recommendation system using collaborative filtering and Qdrant's similarity search."
|
||||
preview_image: /blog/collaborative-filtering/social_preview.png
|
||||
social_preview_image: /blog/collaborative-filtering/social_preview.png
|
||||
weight: 23
|
||||
---
|
||||
|
||||
# Create a collaborative filtering system
|
||||
|
||||
| Time: 45 min | Level: Intermediate | [](https://githubtocolab.com/qdrant/examples/blob/master/collaborative-filtering/collaborative-filtering.ipynb) | |
|
||||
|--------------|---------------------|--|----|
|
||||
|
||||
Every time Spotify recommends the next song from a band you've never heard of, it uses a recommendation algorithm based on other users' interactions with that song. This type of algorithm is known as **collaborative filtering**.
|
||||
|
||||
Unlike content-based recommendations, collaborative filtering excels when the objects' semantics are loosely or unrelated to users' preferences. This adaptability is what makes it so fascinating. Movie, music, or book recommendations are good examples of such use cases. After all, we rarely choose which book to read purely based on the plot twists.
|
||||
|
||||
The traditional way to build a collaborative filtering engine involves training a model that converts the sparse matrix of user-to-item relations into a compressed, dense representation of user and item vectors. Some of the most commonly referenced algorithms for this purpose include [SVD (Singular Value Decomposition)](https://en.wikipedia.org/wiki/Singular_value_decomposition) and [Factorization Machines](https://en.wikipedia.org/wiki/Matrix_factorization_(recommender_systems)). However, the model training approach requires significant resource investments. Model training necessitates data, regular re-training, and a mature infrastructure.
|
||||
|
||||
## Methodology
|
||||
|
||||
Fortunately, there is a way to build collaborative filtering systems without any model training. You can obtain interpretable recommendations and have a scalable system using a technique based on similarity search. Let’s explore how this works with an example of building a movie recommendation system.
|
||||
|
||||
<p align="center"><iframe width="560" height="315" src="https://www.youtube.com/embed/9B7RrmQCQeQ?si=nHp-fM_szHynLcH8" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></p>
|
||||
|
||||
## Implementation
|
||||
|
||||
To implement this, you will use a simple yet powerful resource: [Qdrant with Sparse Vectors](https://qdrant.tech/articles/sparse-vectors/).
|
||||
|
||||
Notebook: [You can try this code here](https://githubtocolab.com/qdrant/examples/blob/master/collaborative-filtering/collaborative-filtering.ipynb)
|
||||
|
||||
|
||||
### Setup
|
||||
|
||||
You have to first import the necessary libraries and define the environment.
|
||||
|
||||
```python
|
||||
import os
|
||||
import pandas as pd
|
||||
import requests
|
||||
from qdrant_client import QdrantClient, models
|
||||
from qdrant_client.models import PointStruct, SparseVector, NamedSparseVector
|
||||
from collections import defaultdict
|
||||
|
||||
# OMDB API Key - for movie posters
|
||||
omdb_api_key = os.getenv("OMDB_API_KEY")
|
||||
|
||||
# Collection name
|
||||
collection_name = "movies"
|
||||
|
||||
# Set Qdrant Client
|
||||
qdrant_client = QdrantClient(
|
||||
os.getenv("QDRANT_HOST"),
|
||||
api_key=os.getenv("QDRANT_API_KEY")
|
||||
)
|
||||
```
|
||||
|
||||
### Define output
|
||||
|
||||
Here, you will configure the recommendation engine to retrieve movie posters as output.
|
||||
|
||||
```python
|
||||
# Function to get movie poster using OMDB API
|
||||
def get_movie_poster(imdb_id, api_key):
|
||||
url = f"https://www.omdbapi.com/?i={imdb_id}&apikey={api_key}"
|
||||
data = requests.get(url).json()
|
||||
return data.get('Poster'), data
|
||||
```
|
||||
|
||||
### Prepare the data
|
||||
|
||||
Load the movie datasets. These include three main CSV files: user ratings, movie titles, and OMDB IDs.
|
||||
|
||||
```python
|
||||
# Load CSV files
|
||||
ratings_df = pd.read_csv('data/ratings.csv', low_memory=False)
|
||||
movies_df = pd.read_csv('data/movies.csv', low_memory=False)
|
||||
|
||||
# Convert movieId in ratings_df and movies_df to string
|
||||
ratings_df['movieId'] = ratings_df['movieId'].astype(str)
|
||||
movies_df['movieId'] = movies_df['movieId'].astype(str)
|
||||
|
||||
rating = ratings_df['rating']
|
||||
|
||||
# Normalize ratings
|
||||
ratings_df['rating'] = (rating - rating.mean()) / rating.std()
|
||||
|
||||
# Merge ratings with movie metadata to get movie titles
|
||||
merged_df = ratings_df.merge(
|
||||
movies_df[['movieId', 'title']],
|
||||
left_on='movieId', right_on='movieId', how='inner'
|
||||
)
|
||||
|
||||
# Aggregate ratings to handle duplicate (userId, title) pairs
|
||||
ratings_agg_df = merged_df.groupby(['userId', 'movieId']).rating.mean().reset_index()
|
||||
|
||||
ratings_agg_df.head()
|
||||
```
|
||||
|
||||
| |userId |movieId |rating |
|
||||
|---|-----------|---------|---------|
|
||||
|0 |1 |1 |0.429960 |
|
||||
|1 |1 |1036 |1.369846 |
|
||||
|2 |1 |1049 |-0.509926|
|
||||
|3 |1 |1066 |0.429960 |
|
||||
|4 |1 |110 |0.429960 |
|
||||
|
||||
### Convert to sparse
|
||||
|
||||
If you want to search across numerous reviews from different users, you can represent these reviews in a sparse matrix.
|
||||
|
||||
```python
|
||||
# Convert ratings to sparse vectors
|
||||
user_sparse_vectors = defaultdict(lambda: {"values": [], "indices": []})
|
||||
for row in ratings_agg_df.itertuples():
|
||||
user_sparse_vectors[row.userId]["values"].append(row.rating)
|
||||
user_sparse_vectors[row.userId]["indices"].append(int(row.movieId))
|
||||
```
|
||||
|
||||

|
||||
|
||||
|
||||
### Upload the data
|
||||
|
||||
Here, you will initialize the Qdrant client and create a new collection to store the data.
|
||||
Convert the user ratings to sparse vectors and include the `movieId` in the payload.
|
||||
|
||||
```python
|
||||
# Define a data generator
|
||||
def data_generator():
|
||||
for user_id, sparse_vector in user_sparse_vectors.items():
|
||||
yield PointStruct(
|
||||
id=user_id,
|
||||
vector={"ratings": SparseVector(
|
||||
indices=sparse_vector["indices"],
|
||||
values=sparse_vector["values"]
|
||||
)},
|
||||
payload={"user_id": user_id, "movie_id": sparse_vector["indices"]}
|
||||
)
|
||||
|
||||
# Upload points using the data generator
|
||||
qdrant_client.upload_points(
|
||||
collection_name=collection_name,
|
||||
points=data_generator()
|
||||
)
|
||||
```
|
||||
|
||||
### Define query
|
||||
|
||||
In order to get recommendations, we need to find users with similar tastes to ours.
|
||||
Let's describe our preferences by providing ratings for some of our favorite movies.
|
||||
|
||||
`1` indicates that we like the movie, `-1` indicates that we dislike it.
|
||||
|
||||
```python
|
||||
my_ratings = {
|
||||
603: 1, # Matrix
|
||||
13475: 1, # Star Trek
|
||||
11: 1, # Star Wars
|
||||
1091: -1, # The Thing
|
||||
862: 1, # Toy Story
|
||||
597: -1, # Titanic
|
||||
680: -1, # Pulp Fiction
|
||||
13: 1, # Forrest Gump
|
||||
120: 1, # Lord of the Rings
|
||||
87: -1, # Indiana Jones
|
||||
562: -1 # Die Hard
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
<details>
|
||||
<summary>Click to see the code for <code>to_vector</code> </summary>
|
||||
|
||||
```python
|
||||
# Create sparse vector from my_ratings
|
||||
def to_vector(ratings):
|
||||
vector = SparseVector(
|
||||
values=[],
|
||||
indices=[]
|
||||
)
|
||||
for movie_id, rating in ratings.items():
|
||||
vector.values.append(rating)
|
||||
vector.indices.append(movie_id)
|
||||
return vector
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
|
||||
### Run the query
|
||||
|
||||
From the uploaded list of movies with ratings, we can perform a search in Qdrant to get the top most similar users to us.
|
||||
|
||||
```python
|
||||
# Perform the search
|
||||
results = qdrant_client.search(
|
||||
collection_name=collection_name,
|
||||
query_vector=NamedSparseVector(
|
||||
name="ratings",
|
||||
vector=to_vector(my_ratings)
|
||||
),
|
||||
limit=20
|
||||
)
|
||||
```
|
||||
|
||||
Now we can find the movies liked by the other similar users, but we haven't seen yet.
|
||||
Let's combine the results from found users, filter out seen movies, and sort by the score.
|
||||
|
||||
```python
|
||||
# Convert results to scores and sort by score
|
||||
def results_to_scores(results):
|
||||
movie_scores = defaultdict(lambda: 0)
|
||||
for result in results:
|
||||
for movie_id in result.payload["movie_id"]:
|
||||
movie_scores[movie_id] += result.score
|
||||
return movie_scores
|
||||
|
||||
# Convert results to scores and sort by score
|
||||
movie_scores = results_to_scores(results)
|
||||
top_movies = sorted(movie_scores.items(), key=lambda x: x[1], reverse=True)
|
||||
```
|
||||
|
||||
<details>
|
||||
|
||||
<summary> Visualize results in Jupyter Notebook </summary>
|
||||
|
||||
Finally, we display the top 5 recommended movies along with their posters and titles.
|
||||
|
||||
```python
|
||||
# Create HTML to display top 5 results
|
||||
html_content = "<div class='movies-container'>"
|
||||
|
||||
for movie_id, score in top_movies[:5]:
|
||||
imdb_id_row = links.loc[links['movieId'] == int(movie_id), 'imdbId']
|
||||
if not imdb_id_row.empty:
|
||||
imdb_id = imdb_id_row.values[0]
|
||||
poster_url, movie_info = get_movie_poster(imdb_id, omdb_api_key)
|
||||
movie_title = movie_info.get('Title', 'Unknown Title')
|
||||
|
||||
html_content += f"""
|
||||
<div class='movie-card'>
|
||||
<img src="{poster_url}" alt="Poster" class="movie-poster">
|
||||
<div class="movie-title">{movie_title}</div>
|
||||
<div class="movie-score">Score: {score}</div>
|
||||
</div>
|
||||
"""
|
||||
else:
|
||||
continue # Skip if imdb_id is not found
|
||||
|
||||
html_content += "</div>"
|
||||
|
||||
display(HTML(html_content))
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
## Recommendations
|
||||
|
||||
For a complete display of movie posters, check the [notebook output](https://github.com/qdrant/examples/blob/master/collaborative-filtering/collaborative-filtering.ipynb). Here are the results without html content.
|
||||
|
||||
```text
|
||||
Toy Story, Score: 131.2033799
|
||||
Monty Python and the Holy Grail, Score: 131.2033799
|
||||
Star Wars: Episode V - The Empire Strikes Back, Score: 131.2033799
|
||||
Star Wars: Episode VI - Return of the Jedi, Score: 131.2033799
|
||||
Men in Black, Score: 131.2033799
|
||||
```
|
||||
|
||||
On top of collaborative filtering, we can further enhance the recommendation system by incorporating other features like user demographics, movie genres, or movie tags.
|
||||
|
||||
Or, for example, only consider recent ratings via a time-based filter. This way, we can recommend movies that are currently popular among users.
|
||||
|
||||
## Conclusion
|
||||
|
||||
As demonstrated, it is possible to build an interesting movie recommendation system without intensive model training using Qdrant and Sparse Vectors. This approach not only simplifies the recommendation process but also makes it scalable and interpretable. In future tutorials, we can experiment more with this combination to further enhance our recommendation systems.
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
stats:
|
||||
githubStars: 18.6k
|
||||
discordMembers: 5.8k
|
||||
twitterFollowers: 6.2k
|
||||
githubStars: 18.7k
|
||||
discordMembers: 6.0k
|
||||
twitterFollowers: 7.5k
|
||||
---
|
||||
Reference in New Issue
Block a user