Merge branch 'master' into feat/bashofmann/hybrid-cloud

This commit is contained in:
Mike Jang
2024-03-15 07:43:46 -07:00
committed by GitHub
20 changed files with 359 additions and 47 deletions
@@ -21,12 +21,12 @@ We are thrilled to announce that Qdrant was 𝐍𝐎𝐓 𝐚𝐜𝐜𝐞𝐩
## Our project ideas.
We have prepared some excelent project ideas. Take a look and choose if you want to contribute in Rust or a Python-based project.
We have prepared some excellent project ideas. Take a look and choose if you want to contribute in Rust or a Python-based project.
➡ *WASM-based dimension reduction viz* 📊
Implement a dimension reduction algorithm in Rust and compile to WASM and integrate the WASM code with Qdrant Web UI.
Implement a dimension reduction algorithm in Rust, compile to WASM and integrate the WASM code with Qdrant Web UI.
➡ *Efficient BM25 and Okapi BM25, which uses the BERT Tokenizer* 🥇
@@ -38,7 +38,7 @@ Export a cross-encoder ranking models to operate on ONNX runtime and integrate t
➡ *Ranking Fusion Algorithms implementation in Rust* 🧪
Develop Rust implementations of various ranking fusion algorithms including but not limited to Reciprocal Rank Fusion (RRF). For complete list, see: https://github.com/AmenRa/ranx
Develop Rust implementations of various ranking fusion algorithms including but not limited to Reciprocal Rank Fusion (RRF). For a complete list, see: https://github.com/AmenRa/ranx
and create Python bindings for the implemented Rust modules.
➡ *Setup Jepsen to test Qdrant’s distributed guarantees* 💣
@@ -51,4 +51,4 @@ See all details on our Notion page: https://www.notion.so/qdrant/GSoC-2024-ideas
Contributor application period begins on March 18th. We will accept applications via email. Let's contribute and celebrate together!
In open-source, we trust! 🦀🤘🚀
In open-source, we trust! 🦀🤘🚀
@@ -0,0 +1,109 @@
---
draft: false
title: "Integrating Qdrant and LangChain for Advanced Vector Similarity Search"
short_description: Discover how Qdrant and LangChain can be integrated to enhance AI applications.
description: Discover how Qdrant and LangChain can be integrated to enhance AI applications with advanced vector similarity search technology.
preview_image: /blog/using-qdrant-and-langchain/qdrant-langchain.png
date: 2024-03-12T09:00:00Z
author: David Myriel
featured: false
weight: 0
tags:
- Qdrant
- LangChain
- LangChain integration
- Vector similarity search
- AI LLM (large language models)
- LangChain agents
- Large Language Models
---
> *"Building AI applications doesn't have to be complicated. You can leverage pre-trained models and support complex pipelines with a few lines of code. LangChain provides a unified interface, so that you can avoid writing boilerplate code and focus on the value you want to bring."* Kacper Lukawski, Developer Advocate, Qdrant
## Long-Term Memory for Your GenAI App
Qdrant's vector database quickly grew due to its ability to make Generative AI more effective. On its own, an LLM can be used to build a process-altering invention. With Qdrant, you can turn this invention into a production-level app that brings real business value.
The use of vector search in GenAI now has a name: **Retrieval Augmented Generation (RAG)**. [In our previous article](https://qdrant.tech/articles/rag-is-dead/), we argued why RAG is an essential component of AI setups, and why large-scale AI can't operate without it. Numerous case studies explain that AI applications are simply too costly and resource-intensive to run using only LLMs.
> Going forward, the solution is to leverage composite systems that use models and vector databases.
**What is RAG?** Essentially, a RAG setup turns Qdrant into long-term memory storage for LLMs. As a vector database, Qdrant manages the efficient storage and retrieval of user data.
Adding relevant context to LLMs can vastly improve user experience, leading to better retrieval accuracy, faster query speed and lower use of compute. Augmenting your AI application with vector search reduces hallucinations, a situation where AI models produce legitimate-sounding but made-up responses.
Qdrant streamlines this process of retrieval augmentation, making it faster, easier to scale and efficient. When you are accessing vast amounts of data (hundreds or thousands of documents), vector search helps your sort through relevant context. **This makes RAG a primary candidate for enterprise-scale use cases.**
## Why LangChain?
Retrieval Augmented Generation is not without its challenges and limitations. One of the main setbacks for app developers is managing the entire setup. The integration of a retriever and a generator into a single model can lead to a raised level of complexity, thus increasing the computational resources required.
[LangChain](https://www.langchain.com/) is a framework that makes developing RAG-based applications much easier. It unifies interfaces to different libraries, including major embedding providers like OpenAI or Cohere and vector stores like Qdrant. With LangChain, you can focus on creating tangible GenAI applications instead of writing your logic from the ground up.
> Qdrant is one of the **top supported vector stores** on LangChain, with [extensive documentation](https://python.langchain.com/docs/integrations/vectorstores/qdrant) and [examples](https://python.langchain.com/docs/integrations/retrievers/self_query/qdrant_self_query).
**How it Works:** LangChain receives a query and retrieves the query vector from an embedding model. Then, it dispatches the vector to a vector database, retrieving relevant documents. Finally, both the query and the retrieved documents are sent to the large language model to generate an answer.
![qdrant-langchain-rag](/blog/using-qdrant-and-langchain/flow-diagram.png)
When supported by LangChain, Qdrant can help you set up effective question-answer systems, detection systems and chatbots that leverage RAG to its full potential. When it comes to long-term memory storage, developers can use LangChain to easily add relevant documents, chat history memory & rich user data to LLM app prompts via Qdrant.
## Common Use Cases
Integrating Qdrant and LangChain can revolutionize your AI applications. Let's take a look at what this integration can do for you:
*Enhance Natural Language Processing (NLP):*
LangChain is great for developing question-answering **chatbots**, where Qdrant is used to contextualize and retrieve results for the LLM. We cover this in [our article](https://qdrant.tech/articles/langchain-integration/), and in OpenAI's [cookbook examples](https://cookbook.openai.com/examples/vector_databases/qdrant/qa_with_langchain_qdrant_and_openai) that use LangChain and GPT to process natural language.
*Improve Recommendation Systems:*
Food delivery services thrive on indecisive customers. Businesses need to accomodate a multi-aim search process, where customers seek recommendations though semantic search. With LangChain you can build systems for **e-commerce, content sharing, or even dating apps**.
*Advance Data Analysis and Insights:* Sometimes you just want to browse results that are not necessarily closest, but still relevant. Semantic search helps user discover products in **online stores**. Customers don't exactly know what they are looking for, but require constrained space in which a search is performed.
*Offer Content Similarity Analysis:* Ever been stuck seeing the same recommendations on your **local news portal**? You may be held in a similarity bubble! As inputs get more complex, diversity becomes scarce, and it becomes harder to force the system to show something different. LangChain developers can use semantic search to develop further context.
## Building a Chatbot with LangChain
_Now that you know how Qdrant and LangChain work together - it's time to build something!_
Follow Daniel Romero's video and create a RAG Chatbot completely from scratch. You will only use OpenAI, Qdrant and LangChain.
Here is what this basic tutorial will teach you:
**1. How to set up a chatbot using Qdrant and LangChain:** You will use LangChain to create a RAG pipeline that retrieves information from a dataset and generates output. This will demonstrate the difference between using an LLM by itself and leveraging a vector database like Qdrant for memory retrieval.
**2. Preprocess and format data for use by the chatbot:** First, you will download a sample dataset based on some academic journals. Then, you will process this data into embeddings and store it as vectors inside of Qdrant.
**3. Implement vector similarity search algorithms:** Second, you will create and test a chatbot that only uses the LLM. Then, you will enable the memory component offered by Qdrant. This will allow your chatbot to be modified and updated, giving it long-term memory.
**4. Optimize the chatbot's performance:** In the last step, you will query the chatbot in two ways. First query will retrieve parametric data from the LLM, while the second one will get contexual data via Qdrant.
The goal of this exercise is to show that RAG is simple to implement via LangChain and yields much better results than using LLMs by itself.
<iframe width="560" height="315" src="https://www.youtube.com/embed/O60-KuZZeQA?si=jkDsyJ52qA4ivXUy" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen></iframe>
## Scaling Qdrant and LangChain
If you are looking to scale up and keep the same level of performance, Qdrant and LangChain are a rock-solid combination. Getting started with both is a breeze and the [documentation](https://python.langchain.com/docs/integrations/vectorstores/qdrant) covers a broad number of cases. However, the main strength of Qdrant is that it can consistently support the user way past the prototyping and launch phases.
> *"We are all-in on performance and reliability. Every release we make Qdrant faster, more stable and cost-effective for the user. When others focus on prototyping, we are already ready for production. Very soon, our users will build successful products and go to market. At this point, I anticipate a great need for a reliable vector store. Qdrant will be there for LangChain and the entire community."*
Whether you are building a bank fraud-detection system, RAG for e-commerce, or services for the federal government - you will need to leverage a scalable architecture for your product. Qdrant offers different features to help you considerably increase your application’s performance and lower your hosting costs.
> Read more about out how we foster [best practices for large-scale deployments](https://qdrant.tech/articles/multitenancy/).
## Next Steps
Now that you know how Qdrant and LangChain can elevate your setup - it's time to try us out.
- Qdrant is open source and you can [quickstart locally](https://qdrant.tech/documentation/quick-start/), [install it via Docker](https://qdrant.tech/documentation/quick-start/), [or to Kubernetes](https://github.com/qdrant/qdrant-helm/).
- We also offer [a free-tier of Qdrant Cloud](https://cloud.qdrant.io/) for prototyping and testing.
- For best integration with LangChain, read the [official LangChain documentation](https://python.langchain.com/docs/integrations/vectorstores/qdrant/).
- For all other cases, [Qdrant documentation](https://qdrant.tech/documentation/integrations/langchain/) is the best place to get there.
> We offer additional support tailored to your business needs. [Contact us](https://qdrant.to/contact-us) to learn more about implementation strategies and integrations that suit your company.
@@ -55,6 +55,10 @@ If you're testing Qdrant, We recommend the Free Tier cluster. The capacity
should be enough to serve up to 1 M vectors of 768 dimensions. To calculate
your needs, refer to our documentation on [Capacity and sizing](/documentation/cloud/capacity-sizing/).
We recommend that you use the Free Tier cluster for testing purposes. The
capacity should be enough to serve up to 1 M vectors of 768 dimensions. To
calculate your needs, refer to our documentation on [Capacity and sizing](/documentation/cloud/capacity-sizing/).
## Support & Troubleshooting
All Qdrant Cloud users are welcome to join our [Discord community](https://qdrant.to/discord/).
@@ -9,25 +9,58 @@ This page shows you how to use the Qdrant Cloud Console to create a custom Qdran
> **Prerequisite:** Please make sure you have provided billing information before creating a custom cluster.
1. Start in the **Clusters** section of the [Cloud Dashboard](https://cloud.qdrant.io).
2. Select **Clusters** and then click **+ Create**.
3. A window will open. Enter a cluster **Name**.
4. Currently, you can deploy to AWS, GCP, or Azure.
5. Choose your data center region. If you have latency concerns or other topology-related requirements, [**let us know**](mailto:cloud@qdrant.io).
6. Configure RAM size for each node (1GB to 64GB).
> Please read [**Capacity and Sizing**](../../cloud/capacity-sizing/) to make the right choice. If you need more capacity per node, [**let us know**](mailto:cloud@qdrant.io).
7. Choose the number of CPUs per node (0.5 core to 16 cores). The max/min number of CPUs is coupled to the chosen RAM size.
8. Select the number of nodes you want the cluster to be deployed on.
> Each node is automatically attached with a disk space offering enough space for your data if you decide to put the metadata or even the index on the disk storage.
9. Click **Create** and wait for your cluster to be provisioned.
1. Start in the **Clusters** section of the [Cloud Dashboard](https://cloud.qdrant.io/).
1. Select **Clusters** and then click **+ Create**.
1. In the **Create a cluster** screen select **Free** or **Standard**
For more information on a free cluster, see the [Cloud quickstart](/documentation/cloud/quickstart-cloud/). The remaining steps assume you want a standard cluster.
1. Select a provider. Currently, you can deploy to:
Your cluster will be reachable on port 443 and 6333 (Rest) and 6334 (gRPC).
- Amazon Web Services (AWS)
- Google Cloud Platform (GCP)
- Microsoft Azure
![Embeddings](/docs/cloud/create-cluster.png)
1. Choose your data center region. If you have latency concerns or other topology-related requirements, [**let us know**](mailto:cloud@qdrant.io).
1. Configure RAM for each node (2 GB to 64 GB).
> For more information, see our [**Capacity and Sizing**](/documentation/cloud/capacity-sizing/) guidance. If you need more capacity per node, [**let us know**](mailto:cloud@qdrant.io).
1. Choose the number of vCPUs per node (0.5 core to 16 cores). If you add more
RAM, the menu provides different options for vCPUs.
1. Select the number of nodes you want the cluster to be deployed on.
> Each node is automatically attached with a disk space offering enough space for your data if you decide to put the metadata or even the index on the disk storage.
1. Select the disk space for your deployment. You can choose from 8 GB to 2 TB.
1. Review your cluster configuration and pricing.
1. When you're ready, select **Create**. It takes some time to provision your cluster.
Once provisioned, you can access your cluster on ports 443 and 6333 (REST)
and 6334 (gRPC).
![Cluster configured in the UI](/docs/cloud/create-cluster-test.png)
You should now see the new cluster in the **Clusters** menu.
A custom cluster includes the following resources. The values in the table are maximums.
| Resource | Value (max) |
|------------|-------------|
| RAM | 64 GB |
| vCPU | 16 vCPU |
| Disk space | 2 TB |
| Nodes | 10 |
### Included features (paid)
The features included with this cluster are:
- Dedicated resources
- Backup and disaster recovery
- Horizontal and vertical scaling
- Monitoring and log management
Learn more about these features in the [Qdrant Cloud dashboard](https://cloud.qdrant.io/).
## Next steps
You will need to connect to your new Qdrant Cloud cluster. Follow [**Authentication**](../../cloud/authentication/) to create one or more API keys.
You will need to connect to your new Qdrant Cloud cluster. Follow [**Authentication**](/documentation/cloud/authentication/) to create one or more API keys.
Your new cluster is highly available and responsive to your application requirements and resource load. Read more in [**Cluster Scaling**](../../cloud/cluster-scaling/).
Your new cluster is highly available and responsive to your application requirements and resource load. Read more in [**Cluster Scaling**](/documentation/cloud/cluster-scaling/).
@@ -1,26 +1,40 @@
---
title: Quickstart
title: Cloud Quickstart
weight: 10
aliases:
- ../cloud-quick-start
- cloud-quick-start
---
# Quickstart
# Cloud Quickstart
This page shows you how to use the Qdrant Cloud Console to create a free tier cluster and then connect to it with Qdrant Client.
## Step 1: Create a Free Tier cluster
## Create a Free Tier cluster
1. Start in the **Overview** section of the [Cloud Dashboard](https://cloud.qdrant.io).
2. Under **Set a Cluster Up** enter a **Cluster name**.
3. Click **Create Free Tier** and then **Continue**.
4. Under **Get an API Key**, select the cluster and click **Get API Key**.
5. Save the API key, as you won't be able to request it again. Click **Continue**.
6. Save the code snippet provided to access your cluster. Click **Complete** to finish setup.
1. Start in the **Overview** section of the [Cloud Dashboard](https://cloud.qdrant.io/).
1. Find the dashboard menu in the left-hand pane. If you do not see it, select
the icon with three horizonal lines in the upper-left of the screen
1. Select **Clusters**. On the Clusters page, select **Create**.
1. In the **Create a Cluster** page, select **Free**
1. Scroll down. Confirm your cluster configuration, and select **Create**.
![Embeddings](/docs/cloud/quickstart-cloud.png)
You should now see your new free tier cluster in the **Clusters** menu.
## Step 2: Test cluster access
A free tier cluster includes the following resources:
| Resource | Value |
|------------|-------|
| RAM | 1 GB |
| vCPU | 0.5 |
| Disk space | 4 GB |
| Nodes | 1 |
## Get an API key
To use your cluster, you need an API key. Read our documentation on [Cloud
Authentication](/documentation/cloud/authentication/) for the process.
## Test cluster access
After creation, you will receive a code snippet to access your cluster. Your generated request should look very similar to this one:
@@ -32,14 +46,18 @@ curl \
Open Terminal and run the request. You should get a response that looks like this:
```bash
{"title":"qdrant - vector search engine","version":"1.4.1"}
{"title":"qdrant - vector search engine","version":"1.8.1"}
```
> **Note:** The API key needs to be present in the request header every time you make a request via Rest or gRPC interface.
## Step 3: Authenticate via SDK
> **Note:** You need to include the API key in the request header for every
> request over REST or gRPC.
Now that you have created your first cluster and key, you might want to access Qdrant Cloud from within your application.
Our official Qdrant clients for Python, TypeScript, Go, Rust, and .NET all support the API key parameter.
## Authenticate via SDK
Now that you have created your first cluster and API key, you can access the
Qdrant Cloud from within your application.
Our official Qdrant clients for Python, TypeScript, Go, Rust, and .NET all
support the API key parameter.
```python
from qdrant_client import QdrantClient
@@ -17,7 +17,7 @@ Please read more about collections, isolation, and multiple users in our [Multit
### My search results contain vectors with null values. Why?
By default, Qdrant tries to minimize network traffic and doesn't return vectors in search results.
But you can force Qdrant to do so by setting the `with_vector` parameter of the Search/Scroll to `true`.
But you can force Qdrant to do so by setting the `with_vector` parameter of the Search/Scroll to `true`.
If you're still seeing `"vector": null` in your results, it might be that the vector you're passing is not in the correct format, or there's an issue with how you're calling the upsert method.
@@ -47,13 +47,13 @@ What Qdrant doesn't plan to support:
- Query analyzers and other NLP tools
Of course, you can always combine Qdrant with any specialized tool you need, including full-text search engines.
Read more about [our approach](../../../articles/hybrid-search/) to hybrid search.
Read more about [our approach](../../../articles/hybrid-search/) to hybrid search.
### How do I upload a large number of vectors into a Qdrant collection?
Read about our recommendations in the [bulk upload](../../tutorials/bulk-upload/) tutorial.
### Can I only store quantized vectors and discard full precision vectors?
### Can I only store quantized vectors and discard full precision vectors?
No, Qdrant requires full precision vectors for operations like reindexing, rescoring, etc.
@@ -66,12 +66,16 @@ But in some cases, we might be able to help you with that through manual interve
## Versioning
### Do you support downgrades?
We do not support downgrading a cluster on any of our products. If you deploy a newer version of Qdrant, your
data is automatically migrated to the newer storage format. This migration is not reversible.
### How do I avoid issues when updating to the latest version?
We only guarantee compatibility if you update between consequent versions. You would need to upgrade versions one at a time: `1.1 -> 1.2`, then `1.2 -> 1.3`, then `1.3 -> 1.4`.
We only guarantee compatibility if you update between consecutive versions. You would need to upgrade versions one at a time: `1.1 -> 1.2`, then `1.2 -> 1.3`, then `1.3 -> 1.4`.
### Do you guarantee compatibility across versions?
In case your version is older, we guarantee only compatibility between two consecutive minor versions.
In case your version is older, we only guarantee compatibility between two consecutive minor versions. This also applies to client versions. Ensure your client version is never more than one minor version away from your cluster version.
While we will assist with break/fix troubleshooting of issues and errors specific to our products, Qdrant is not accountable for reviewing, writing (or rewriting), or debugging custom code.
@@ -0,0 +1,69 @@
---
title: Spring AI
weight: 2200
---
# Spring AI
[Spring AI](https://docs.spring.io/spring-ai/reference/) is a Java framework that provides a [Spring-friendly](https://spring.io/) API and abstractions for developing AI applications.
Qdrant is available as supported vector database for use within your Spring AI projects.
## Installation
To acquire Spring AI artifacts, declare the Spring Snapshot repository in your `pom.xml`.
```xml
<repository>
<id>spring-snapshots</id>
<name>Spring Snapshots</name>
<url>https://repo.spring.io/snapshot</url>
<releases>
<enabled>false</enabled>
</releases>
</repository>
```
Add the `spring-ai-qdrant` package.
```xml
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-qdrant</artifactId>
<version>VERSION</version>
</dependency>
```
## Usage
You can set up the Qdrant vector store with the `QdrantVectorStoreConfig` options.
```java
@Bean
public QdrantVectorStoreConfig qdrantVectorStoreConfig() {
return QdrantVectorStoreConfig.builder()
.withHost("<QDRANT_HOSTNAME>")
.withPort(<QDRANT_GRPC_PORT>)
.withCollectionName("<QDRANT_COLLECTION_NAME>")
.withApiKey("<QDRANT_API_KEY>")
.build();
}
```
<aside role="status">You'll need to <a href="/documentation/concepts/collections/#create-a-collection">create a collection</a> with the appropriate vector dimensions and configurations in advance.</aside>
Build the vector store using the config and any of the support [Spring AI embedding providers](https://docs.spring.io/spring-ai/reference/api/embeddings.html#available-implementations).
```java
@Bean
public VectorStore vectorStore(QdrantVectorStoreConfig config, EmbeddingClient embeddingClient) {
return new QdrantVectorStore(config, embeddingClient);
}
```
You can now use the `VectorStore` instance backed by Qdrant as a vector store in the Spring AI APIs.
## Further Reading
- 📚 Spring AI [reference](https://docs.spring.io/spring-ai/reference/index.html)
@@ -11,6 +11,27 @@ aliases:
Since version v0.8.0 Qdrant supports a distributed deployment mode.
In this mode, multiple Qdrant services communicate with each other to distribute the data across the peers to extend the storage capabilities and increase stability.
## How many Qdrant nodes should I run?
The ideal number of Qdrant nodes depends on how much you value cost-saving, resilience, and performance/scalability in relation to each other.
- **Prioritizing cost-saving**: If cost is most important to you, run a single Qdrant node. This is not recommended for production environments. Drawbacks:
- Resilience: Users will experience downtime during node restarts, and recovery is not possible unless you have backups or snapshots.
- Performance: Limited to the resources of a single server.
- **Prioritizing resilience**: If resilience is most important to you, run a Qdrant cluster with three or more nodes and two or more shard replicas. Clusters with three or more nodes and replication can perform all operations even while one node is down. Additionally, they gain performance benefits from load-balancing and they can recover from the permanent loss of one node without the need for backups or snapshots (but backups are still strongly recommended). This is most recommended for production environments. Drawbacks:
- Cost: Larger clusters are more costly than smaller clusters, which is the only drawback of this configuration.
- **Balancing cost, resilience, and performance**: Running a two-node Qdrant cluster with replicated shards allows the cluster to respond to most read/write requests even when one node is down, such as during maintenance events. Having two nodes also means greater performance than a single-node cluster while still being cheaper than a three-node cluster. Drawbacks:
- Resilience (uptime): The cluster cannot perform operations on collections when one node is down. Those operations require >50% of nodes to be running, so this is only possible in a 3+ node cluster. Since creating, editing, and deleting collections are usually rare operations, many users find this drawback to be negligible.
- Resilience (data integrity): If the data on one of the two nodes is permanently lost or corrupted, it cannot be recovered aside from snapshots or backups. Only 3+ node clusters can recover from the permanent loss of a single node since recovery operations require >50% of the cluster to be healthy.
- Cost: Replicating your shards requires storing two copies of your data.
- Performance: The maximum performance of a Qdrant cluster increases as you add more nodes.
In summary, single-node clusters are best for non-production workloads, replicated 3+ node clusters are the gold standard, and replicated 2-node clusters strike a good balance.
## Enabling distributed mode in self-hosted Qdrant
To enable distributed deployment - enable the cluster mode in the [configuration](../configuration/) or using the ENV variable: `QDRANT__CLUSTER__ENABLED=true`.
```yaml
@@ -106,6 +127,24 @@ Example result:
}
```
Note that enabling distributed mode does not automatically replicate your data. See the section on [making use of a new distributed Qdrant cluster](#making-use-of-a-new-distributed-qdrant-cluster) for the next steps.
## Enabling distributed mode in Qdrant Cloud
For best results, first ensure your cluster is running Qdrant v1.7.4 or higher. Older versions of Qdrant do support distributed mode, but improvements in v1.7.4 make distributed clusters more resilient during outages.
In the [Qdrant Cloud console](https://cloud.qdrant.io/), click "Scale Up" to increase your cluster size to >1. Qdrant Cloud configures the distributed mode settings automatically.
After the scale-up process completes, you will have a new empty node running alongside your existing node(s). To replicate data into this new empty node, see the next section.
## Making use of a new distributed Qdrant cluster
When you enable distributed mode and scale up to two or more nodes, your data does not move to the new node automatically; it starts out empty. To make use of your new empty node, do one of the following:
* Replicate your existing data to the new node by [creating new shard replicas](#creating-new-shard-replicas)
* Create a new replicated collection by setting the [replication_factor](#replication-factor) to 2 or more (can only be set at creation time)
* Move data (without replicating it) onto the new node by [moving shards](#moving-shards)
## Raft
Qdrant uses the [Raft](https://raft.github.io/) consensus protocol to maintain consistency regarding the cluster topology and the collections structure.
@@ -780,11 +819,12 @@ Let's walk through them from best to worst.
**Recover with replicated collection**
If the number of failed nodes is less than the replication factor of the collection, then no data is lost.
Your cluster should still be able to perform read, search and update queries.
If the number of failed nodes is less than the replication factor of the collection, then your cluster should still be able to perform read, search and update queries.
Now, if the failed node restarts, consensus will trigger the replication process to update the recovering node with the newest updates it has missed.
If the failed node never restarts, you can recover the lost shards if you have a 3+ node cluster. You cannot recover lost shards in smaller clusters because recovery operations go through [raft](#raft) which requires >50% of the nodes to be healthy.
**Recreate node with replicated collections**
If a node fails and it is impossible to recover it, you should exclude the dead node from the consensus and create an empty node.
@@ -820,6 +860,17 @@ The service will download the specified snapshot of the collection and recover s
Once all shards of the collection are recovered, the collection will become operational again.
### Temporary node failure
If properly configured, running Qdrant in distributed mode can make your cluster resistant to outages when one node fails temporarily.
Here is how differently-configured Qdrant clusters respond:
* 1-node clusters: All operations time out or fail for up to a few minutes. It depends on how long it takes to restart and load data from disk.
* 2-node clusters where shards ARE NOT replicated: All operations will time out or fail for up to a few minutes. It depends on how long it takes to restart and load data from disk.
* 2-node clusters where all shards ARE replicated to both nodes: All requests except for operations on collections continue to work during the outage.
* 3+-node clusters where all shards are replicated to at least 2 nodes: All requests continue to work during the outage.
## Consistency guarantees
By default, Qdrant focuses on availability and maximum throughput of search operations.
@@ -20,8 +20,8 @@ source code itself, which is mostly written in Rust.
We want to search codebases using natural semantic queries, and searching for
code based on similar logic. You can set up these tasks with embeddings:
1. General usage neural encoder for natural-like queries, in our case `all-MiniLM-L6-v2`
from the
1. General usage neural encoder for Natural Language Processing (NLP), in our case
`all-MiniLM-L6-v2` from the
[sentence-transformers](https://www.sbert.net/docs/pretrained_models.html) library.
2. Specialized embeddings for code-to-code similarity search. We use the
`jina-embeddings-v2-base-code` model.
@@ -31,6 +31,11 @@ more closely resembles natural language. The Jina embeddings model supports a
variety of standard programming languages, so there is no need to preprocess the
snippets. We can use the code as is.
NLP-based search is based on function signatures, but code search may return
smaller pieces, such as loops. So, if we receive a particular function signature
from the NLP model and part of its implementation from the code model, we merge
the results and highlight the overlap.
## Data preparation
Chunking the application sources into smaller parts is a non-trivial task. In
@@ -411,14 +416,30 @@ This is one example of how you can use different models and combine the results.
In a real-world scenario, you might run some reranking and deduplication, as
well as additional processing of the results.
Our [Code search demo](https://github.com/qdrant/demo-code-search) uses
both models. In the screenshot, we search for `flush of wal`. The result
### Code search demo
Our [Code search demo](https://code-search.qdrant.tech/) uses the following process:
1. The user sends a query.
1. Both models vectorize that query simultaneously. We get two different
vectors.
1. Both vectors are used in parallel to find relevant snippets. We expect
5 examples from the NLP search and 20 examples from the code search.
1. Once we retrieve results for both vectors, we merge them in one of the
following scenarios:
1. If both methods return different results, we prefer the results from
the general usage model (NLP).
1. If there is an overlap between the search results, we merge overlapping
snippets.
In the screenshot, we search for `flush of wal`. The result
shows relevant code, merged from both models. Note the highlighted
code in lines 621-629. It's where both models agree.
![Results from both models, with overlap](/documentation/tutorials/code-search/code-search-demo-example.png)
Now you see semantic code intelligence, in action.
### Grouping the results
You can improve the search results, by grouping them by payload properties.
@@ -15,6 +15,9 @@ This tutorial will show you how to create a snapshot of a collection and restore
<aside role="status">Snapshots cannot be created in local mode of Python SDK. You need to spin up a Qdrant Docker container or use Qdrant Cloud.</aside>
You can use the techniques described in this page to migrate a cluster. Follow the instructions
in this tutorial to create and download snapshots. When you [Restore from snapshot](#restore-from-snapshot), restore your data to the new cluster.
## Prerequisites
Let's assume you already have a running Qdrant instance or a cluster. If not, you can follow the [installation guide](/documentation/guides/installation/) to set up a local Qdrant instance or use [Qdrant Cloud](https://cloud.qdrant.io/) to create a cluster in a few clicks.