mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-09 21:08:31 +02:00
Break up "Distributed Deployment" page into new "Scaling & Resilience" section (#2491)
* Add new Scaling landing page under Operations Introduces a Scaling section with vertical vs. horizontal scaling guidance and failover best practices, linking out to detail pages. * Add new Vertical Scaling page Dedicated how-to guidance for resizing existing nodes: when to scale vertically, RAM sizing formulas, and Cloud/self-hosted resize steps. * Add new Horizontal Scaling and Resilience page Covers Raft consensus, the replication model, consistency guarantees, Multi-AZ, and the resilience terminology used elsewhere in the docs. * Move Distributed Deployment under Scaling and update all incoming links Moves distributed_deployment.md into the new scaling/ section, trims its Raft/Replication/Consistency intros into cross-links to the new Horizontal Scaling and Resilience page, adds Multi-AZ and single-replica cross-link callouts in the Cloud docs, rewrites all internal references across ~30 files to the new canonical path instead of relying on aliases, and applies Title Case to Distributed Deployment's headers. * Split Resilience out of Horizontal Scaling and Resilience Adds a dedicated Resilience page covering fault tolerance, Multi-AZ, resilience terminology, and failover best practices (moved from the Scaling landing page). Horizontal Scaling is retitled and scoped to the underlying mechanics: Raft consensus, replication, and consistency. * Reorganize Horizontal Scaling's structure Moves "How Many Qdrant Nodes Should I Run?" from Distributed Deployment into Horizontal Scaling, adds a conceptual Sharding section, and reorders Sharding/Replication/Raft Consensus/Consistency. Moves the remaining conceptual content out of Distributed Deployment: Temporary Node Failure to Resilience, Error Handling folded into Replication, sharding heuristics folded into Sharding, and the Consensus Checkpointing explanation folded into Raft Consensus. * Rename Scaling section to Scaling & Resilience Renames the section and restructures the landing page: the vertical- vs-horizontal decision is now purely about scaling, with a dedicated Resilience section covering fault tolerance through sharding and multi-node deployments. * Polish Vertical Scaling and Resilience page content Reframes Vertical Scaling's "What Not to Do" as positive "Best Practices". Reworks Resilience's structure: moves the uptime/data- integrity terminology into the intro as three distinct aspects of resilience, and renames "How Resilience Works" to "Setting Up a Resilient Qdrant Cluster". * Add diagrams illustrating sharding and replication Adds cluster diagrams to the Sharding and Replication sections on Horizontal Scaling to make the shard/replica layout easier to follow. * Add new Node Failure Recovery page Extracts the node failure recovery scenarios out of Distributed Deployment into their own page, with each bolded sub-header converted to a proper heading, and links updated across Resilience and the Scaling landing page. * Add new Consistency Guarantees page Extracts write consistency factor, read consistency, and write ordering out of Distributed Deployment into their own page, positioned after Distributed Deployment. * Add new "Deploy Behind a Load Balancer" section Explains why a load balancer is needed in front of a multi-node Qdrant cluster: avoiding a single point of failure at the entry point and making sure replicas on every node actually serve reads. * Add new "Rebalancing" section Documents how Qdrant Cloud automatically rebalances shards across nodes, as its own subsection under Sharding. * Rewrite Multi-AZ vs. Replication Factor as Multi-AZ Deployments Defines an availability zone on first use, explains why multi-AZ deployments guard against a zone going down, clarifies that Qdrant Cloud is zone-aware once enabled, and that self-hosted deployments need to place and move replicas across zones manually. * Restructure node-count guidance into One/Two/Three-or-more Node subsections Splits "How Many Qdrant Nodes Should I Run?" into three subsections and drops the "balanced" framing for two nodes: it states plainly that two nodes give more capacity without true high availability. * Add new "Which Configuration Is Right for You?" section Summarizes the one/two/three-or-more node tradeoffs in one place right after the detailed breakdown. * Add explicit _redirects entry for legacy distributed_deployment URL Closes the redirect chain: the existing /guides/ and /operations/ legacy rules both terminate at /documentation/distributed_deployment/, which previously had no explicit _redirects entry and only resolved via the Hugo alias meta-refresh page. * Fix all incoming links to Distributed Deployment and pages under Scaling Repoints two same-page anchors in distributed_deployment.md that broke when Write Ordering moved to Consistency Guarantees, and one link in cloud/create-cluster.md that broke when a Resilience heading was reworded. * Update time-based sharding diagram and restructure section * Fix a couple of broken links * Move 'Consensus Checkpointing' to 'Node Failure Recovery' page
This commit is contained in:
@@ -48,7 +48,7 @@ features:
|
|||||||
description: Qdrant’s architecture is optimized for high-throughput embedding processing, minimizing CPU load and preventing performance bottlenecks. This enables AI agents in Agentic RAG workflows to execute complex, multi-step tasks efficiently, ensuring smooth operation even at scale.
|
description: Qdrant’s architecture is optimized for high-throughput embedding processing, minimizing CPU load and preventing performance bottlenecks. This enables AI agents in Agentic RAG workflows to execute complex, multi-step tasks efficiently, ensuring smooth operation even at scale.
|
||||||
link:
|
link:
|
||||||
text: Distributed Deployment
|
text: Distributed Deployment
|
||||||
url: /documentation/distributed_deployment/
|
url: /documentation/scaling/distributed_deployment/
|
||||||
- id: 4
|
- id: 4
|
||||||
icon:
|
icon:
|
||||||
src: /icons/outline/speedometer-blue.svg
|
src: /icons/outline/speedometer-blue.svg
|
||||||
|
|||||||
@@ -126,7 +126,7 @@ The same agent that speeds through a toy dataset with 10,000 points will become
|
|||||||
|
|
||||||
We’ll talk about three concepts you can take advantage of to improve your scale, but if you want even more information on how to scale, check out this [article](https://qdrant.tech/documentation/database-tutorials/large-scale-search/) on large scale search.
|
We’ll talk about three concepts you can take advantage of to improve your scale, but if you want even more information on how to scale, check out this [article](https://qdrant.tech/documentation/database-tutorials/large-scale-search/) on large scale search.
|
||||||
|
|
||||||
As your dataset and traffic grow, Qdrant Cloud offers a suite of features to ensure your system can scale effectively. Horizontal scaling is achieved through [sharding](https://qdrant.tech/articles/multitenancy/), which splits your collection across multiple nodes to distribute the load and improve performance. For high availability and fault tolerance, Qdrant supports [replication](https://qdrant.tech/documentation/distributed_deployment/), creating copies of your shards across the cluster.
|
As your dataset and traffic grow, Qdrant Cloud offers a suite of features to ensure your system can scale effectively. Horizontal scaling is achieved through [sharding](https://qdrant.tech/articles/multitenancy/), which splits your collection across multiple nodes to distribute the load and improve performance. For high availability and fault tolerance, Qdrant supports [replication](https://qdrant.tech/documentation/scaling/distributed_deployment/), creating copies of your shards across the cluster.
|
||||||
|
|
||||||
Qdrant provides robust tools for resource and cost optimization. Vector [quantization](https://qdrant.tech/documentation/manage-data/quantization/) compresses your data, significantly reducing its memory footprint and speeding up search.
|
Qdrant provides robust tools for resource and cost optimization. Vector [quantization](https://qdrant.tech/documentation/manage-data/quantization/) compresses your data, significantly reducing its memory footprint and speeding up search.
|
||||||
|
|
||||||
|
|||||||
@@ -19,14 +19,14 @@ category: production-ops
|
|||||||
|
|
||||||
# Scaling Your Machine Learning Setup: The Power of Multitenancy and Custom Sharding in Qdrant
|
# Scaling Your Machine Learning Setup: The Power of Multitenancy and Custom Sharding in Qdrant
|
||||||
|
|
||||||
We are seeing the topics of [multitenancy](/documentation/manage-data/multitenancy/) and [distributed deployment](/documentation/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup.
|
We are seeing the topics of [multitenancy](/documentation/manage-data/multitenancy/) and [distributed deployment](/documentation/scaling/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup.
|
||||||
|
|
||||||
Whether you are building a bank fraud-detection system, [RAG](https://qdrant.tech/articles/what-is-rag-in-ai/) for e-commerce, or services for the federal government - you will need to leverage a multitenant architecture to scale your product.
|
Whether you are building a bank fraud-detection system, [RAG](https://qdrant.tech/articles/what-is-rag-in-ai/) for e-commerce, or services for the federal government - you will need to leverage a multitenant architecture to scale your product.
|
||||||
In the world of SaaS and enterprise apps, this setup is the norm. It will considerably increase your application's performance and lower your hosting costs.
|
In the world of SaaS and enterprise apps, this setup is the norm. It will considerably increase your application's performance and lower your hosting costs.
|
||||||
|
|
||||||
## Multitenancy & custom sharding with Qdrant
|
## Multitenancy & custom sharding with Qdrant
|
||||||
|
|
||||||
We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](/documentation/manage-data/multitenancy/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](/documentation/distributed_deployment/#user-defined-sharding).
|
We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](/documentation/manage-data/multitenancy/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](/documentation/scaling/distributed_deployment/#user-defined-sharding).
|
||||||
|
|
||||||
Combining these two will result in an efficiently-partitioned architecture that further leverages the convenience of a single Qdrant cluster. This article will briefly explain the benefits and show how you can get started using both features.
|
Combining these two will result in an efficiently-partitioned architecture that further leverages the convenience of a single Qdrant cluster. This article will briefly explain the benefits and show how you can get started using both features.
|
||||||
|
|
||||||
@@ -41,7 +41,7 @@ Qdrant is built to excel in a single collection with a vast number of tenants. Y
|
|||||||
|
|
||||||
## Sharding your database
|
## Sharding your database
|
||||||
|
|
||||||
With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](/documentation/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node.
|
With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](/documentation/scaling/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node.
|
||||||
|
|
||||||
During vector search, your operations will be able to hit only the subset of shards they actually need. In massive-scale deployments, __this can significantly improve the performance of operations that do not require the whole collection to be scanned__.
|
During vector search, your operations will be able to hit only the subset of shards they actually need. In massive-scale deployments, __this can significantly improve the performance of operations that do not require the whole collection to be scanned__.
|
||||||
|
|
||||||
@@ -49,7 +49,7 @@ This works in the other direction as well. Whenever you search for something, yo
|
|||||||
|
|
||||||
### Common use cases
|
### Common use cases
|
||||||
|
|
||||||
A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](/documentation/distributed_deployment/#moving-shards).
|
A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](/documentation/scaling/distributed_deployment/#moving-shards).
|
||||||
|
|
||||||
**Figure 2:** Users can both upsert and query shards that are relevant to them, all within the same collection. Regional sharding can help avoid cross-continental traffic.
|
**Figure 2:** Users can both upsert and query shards that are relevant to them, all within the same collection. Regional sharding can help avoid cross-continental traffic.
|
||||||

|

|
||||||
@@ -79,7 +79,7 @@ client.create_shard_key("{tenant_data}", "germany")
|
|||||||
```
|
```
|
||||||
In this example, your cluster is divided between Germany and Canada. Canadian and German law differ when it comes to international data transfer. Let's say you are creating a RAG application that supports the healthcare industry. Your Canadian customer data will have to be clearly separated for compliance purposes from your German customer.
|
In this example, your cluster is divided between Germany and Canada. Canadian and German law differ when it comes to international data transfer. Let's say you are creating a RAG application that supports the healthcare industry. Your Canadian customer data will have to be clearly separated for compliance purposes from your German customer.
|
||||||
|
|
||||||
Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](/documentation/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech).
|
Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](/documentation/scaling/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech).
|
||||||
|
|
||||||
## Configure a multitenant setup for users
|
## Configure a multitenant setup for users
|
||||||
|
|
||||||
|
|||||||
@@ -92,7 +92,7 @@ POST /collections/my_collection/points/search
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
If you want to know more about the user-defined sharding, please refer to the [sharding documentation](/documentation/distributed_deployment/#sharding).
|
If you want to know more about the user-defined sharding, please refer to the [sharding documentation](/documentation/scaling/distributed_deployment/#sharding).
|
||||||
|
|
||||||
### Snapshot-based shard transfer
|
### Snapshot-based shard transfer
|
||||||
|
|
||||||
@@ -101,7 +101,7 @@ That's a really more in depth technical improvement for the distributed mode use
|
|||||||
Moving shards is required for dynamical scaling of the cluster. Your data can migrate between nodes, and the way you move it is crucial for the performance of the whole system. The good old `stream_records` method (still the default one) transmits all the records between the machines and indexes them on the target node.
|
Moving shards is required for dynamical scaling of the cluster. Your data can migrate between nodes, and the way you move it is crucial for the performance of the whole system. The good old `stream_records` method (still the default one) transmits all the records between the machines and indexes them on the target node.
|
||||||
In the case of moving the shard, it's necessary to recreate the HNSW index each time. However, with the introduction of the new `snapshot` approach, the snapshot itself, inclusive of all data and potentially quantized content, is transferred to the target node. This comprehensive snapshot includes the entire index, enabling the target node to seamlessly load it and promptly begin handling requests without the need for index recreation.
|
In the case of moving the shard, it's necessary to recreate the HNSW index each time. However, with the introduction of the new `snapshot` approach, the snapshot itself, inclusive of all data and potentially quantized content, is transferred to the target node. This comprehensive snapshot includes the entire index, enabling the target node to seamlessly load it and promptly begin handling requests without the need for index recreation.
|
||||||
|
|
||||||
There are multiple scenarios in which you may prefer one over the other. Please check out the docs of the [shard transfer method](/documentation/distributed_deployment/#shard-transfer-method) for more details and head-to-head comparison. As for now, the old `stream_records` method is still the default one, but we may decide to change it in the future.
|
There are multiple scenarios in which you may prefer one over the other. Please check out the docs of the [shard transfer method](/documentation/scaling/distributed_deployment/#shard-transfer-method) for more details and head-to-head comparison. As for now, the old `stream_records` method is still the default one, but we may decide to change it in the future.
|
||||||
|
|
||||||
## Minor improvements
|
## Minor improvements
|
||||||
|
|
||||||
|
|||||||
@@ -243,7 +243,7 @@ It depends. If you're just starting out - we have prepared a tool on our website
|
|||||||
|
|
||||||
A three-node setup provides a baseline for fault tolerance: if one node goes offline, the remaining two can continue serving queries and maintain a quorum for data consistency. This guards against hardware failures, rolling updates, and network disruptions. Fewer than three nodes leaves you vulnerable to single-point failures that can knock your entire cluster offline.
|
A three-node setup provides a baseline for fault tolerance: if one node goes offline, the remaining two can continue serving queries and maintain a quorum for data consistency. This guards against hardware failures, rolling updates, and network disruptions. Fewer than three nodes leaves you vulnerable to single-point failures that can knock your entire cluster offline.
|
||||||
|
|
||||||
> [**We follow the Raft Protocol**](https://qdrant.tech/documentation/distributed_deployment/#raft), so check out the docs and learn why this is important.
|
> [**We follow the Raft Protocol**](https://qdrant.tech/documentation/scaling/horizontal-scaling/#raft-consensus), so check out the docs and learn why this is important.
|
||||||
|
|
||||||
✅ **Set a replication factor of at least 2** to tolerate node failure without losing availability.
|
✅ **Set a replication factor of at least 2** to tolerate node failure without losing availability.
|
||||||
|
|
||||||
@@ -275,7 +275,7 @@ Development and staging environments often run experimental builds, tests, or si
|
|||||||
|
|
||||||
> It's quite possible that the user has multiple shards on one node, which end up handling most traffic while other nodes remain underutilized.
|
> It's quite possible that the user has multiple shards on one node, which end up handling most traffic while other nodes remain underutilized.
|
||||||
|
|
||||||
In this case, you should [**choose the right number of shards**](https://qdrant.tech/documentation/distributed_deployment/#sharding) based on your node count and expected RPS.
|
In this case, you should [**choose the right number of shards**](https://qdrant.tech/documentation/scaling/distributed_deployment/#sharding) based on your node count and expected RPS.
|
||||||
|
|
||||||
You need to implement a shard strategy that aligns with real usage patterns. First, distribute your shards across all available nodes. This will help balance the load more effectively. After redistributing the shards, run performance tests to see how it affects your system. Then add replicas and test again to see how that changes performance.
|
You need to implement a shard strategy that aligns with real usage patterns. First, distribute your shards across all available nodes. This will help balance the load more effectively. After redistributing the shards, run performance tests to see how it affects your system. Then add replicas and test again to see how that changes performance.
|
||||||
|
|
||||||
@@ -286,14 +286,14 @@ Proper sharding considers data distribution and query patterns. By default, shar
|
|||||||
|
|
||||||
||
|
||
|
||||||
|-|
|
|-|
|
||||||
|**Read More:** [**Sharding Documentation**](https://qdrant.tech/documentation/distributed_deployment/#sharding)|
|
|**Read More:** [**Sharding Documentation**](https://qdrant.tech/documentation/scaling/distributed_deployment/#sharding)|
|
||||||
|
|
||||||
### Manage Your Costs by Scaling Up or Down
|
### Manage Your Costs by Scaling Up or Down
|
||||||

|

|
||||||
|
|
||||||
Some teams scale up for daytime surges, then scale down overnight to save resources. If you do this, ensure data is sharded and replicated appropriately, so that scaling up and down won't result in service degradation.
|
Some teams scale up for daytime surges, then scale down overnight to save resources. If you do this, ensure data is sharded and replicated appropriately, so that scaling up and down won't result in service degradation.
|
||||||
|
|
||||||
If using Qdrant Cloud you could also do this using the [**Replication Factor**](https://qdrant.tech/documentation/distributed_deployment/#replication-factor), though it may be considered a bit of a hack.
|
If using Qdrant Cloud you could also do this using the [**Replication Factor**](https://qdrant.tech/documentation/scaling/distributed_deployment/#replication-factor), though it may be considered a bit of a hack.
|
||||||
|
|
||||||
> If you have 3 nodes with just 1 shard, and replication factor 6. It will create 3 replicas (one on each node) of that shard, because it can't host more. If you add 3 more nodes at peak times, it'll automatically replicate that shard 3 more times in an attempt to match the factor of 6.
|
> If you have 3 nodes with just 1 shard, and replication factor 6. It will create 3 replicas (one on each node) of that shard, because it can't host more. If you add 3 more nodes at peak times, it'll automatically replicate that shard 3 more times in an attempt to match the factor of 6.
|
||||||
|
|
||||||
@@ -309,7 +309,7 @@ If new nodes remain empty after joining, you waste resources. If departing nodes
|
|||||||
|
|
||||||
||
|
||
|
||||||
|-|
|
|-|
|
||||||
|**Read More:** [**Distributed Deployment Documentation**](https://qdrant.tech/documentation/distributed_deployment/)|
|
|**Read More:** [**Distributed Deployment Documentation**](https://qdrant.tech/documentation/scaling/distributed_deployment/)|
|
||||||
|**Read More:** [**Resharding**](https://qdrant.tech/documentation/cloud/cluster-scaling/#resharding)|
|
|**Read More:** [**Resharding**](https://qdrant.tech/documentation/cloud/cluster-scaling/#resharding)|
|
||||||
|
|
||||||
### How to Predict and Test Cluster Performance
|
### How to Predict and Test Cluster Performance
|
||||||
@@ -332,7 +332,7 @@ Remember, cold-starts and query behaviour are dataset dependent, which is why yo
|
|||||||
|
|
||||||
||
|
||
|
||||||
|-|
|
|-|
|
||||||
|**Read More:** [Distributed Deployment Documentation](https://qdrant.tech/documentation/distributed_deployment/)
|
|**Read More:** [Distributed Deployment Documentation](https://qdrant.tech/documentation/scaling/distributed_deployment/)
|
||||||
|
|
||||||
### How to Design Your Systems to Protect Against Failure
|
### How to Design Your Systems to Protect Against Failure
|
||||||
|
|
||||||
|
|||||||
@@ -325,7 +325,7 @@ Here’s how to choose the shard_number:
|
|||||||
| **Plan for Scalability** | Start with at least **2 shards per node** to allow room for future growth. |
|
| **Plan for Scalability** | Start with at least **2 shards per node** to allow room for future growth. |
|
||||||
| **Future-Proofing** | Starting with around **12 shards** is a good rule of thumb. This setup allows your system to scale seamlessly from 1 to 12 nodes without requiring re-sharding. |
|
| **Future-Proofing** | Starting with around **12 shards** is a good rule of thumb. This setup allows your system to scale seamlessly from 1 to 12 nodes without requiring re-sharding. |
|
||||||
|
|
||||||
Learn more about [**Sharding in Distributed Deployment**](/documentation/distributed_deployment/)
|
Learn more about [**Sharding in Distributed Deployment**](/documentation/scaling/distributed_deployment/)
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -413,7 +413,7 @@ client.create_collection(
|
|||||||
|
|
||||||
We recommend using sharding and replication together so that your data is both split across nodes and replicated for availability.
|
We recommend using sharding and replication together so that your data is both split across nodes and replicated for availability.
|
||||||
|
|
||||||
For more details on features like **user-defined sharding, node failure recovery**, and **consistency guarantees**, see our guide on [Distributed Deployment.](https://qdrant.tech/documentation/distributed_deployment/)
|
For more details on features like **user-defined sharding, node failure recovery**, and **consistency guarantees**, see our guide on [Distributed Deployment.](https://qdrant.tech/documentation/scaling/distributed_deployment/)
|
||||||
|
|
||||||
## Multitenancy: Data Isolation for Multi-Tenant Architectures
|
## Multitenancy: Data Isolation for Multi-Tenant Architectures
|
||||||
|
|
||||||
|
|||||||
@@ -58,7 +58,7 @@ As part of their selection process, Nyris evaluated several critical factors to
|
|||||||
Nyris has found several aspects of Qdrant particularly beneficial in their production environment:
|
Nyris has found several aspects of Qdrant particularly beneficial in their production environment:
|
||||||
|
|
||||||
- **Enhanced Security with JWT**: [JSON Web Tokens](https://qdrant.tech/documentation/security#granular-access-api-keys) provide enhanced security and performance, critical for safeguarding their data.
|
- **Enhanced Security with JWT**: [JSON Web Tokens](https://qdrant.tech/documentation/security#granular-access-api-keys) provide enhanced security and performance, critical for safeguarding their data.
|
||||||
- **Seamless Scalability**: Qdrant's ability to [scale effortlessly across nodes](https://qdrant.tech/documentation/distributed_deployment/) ensures consistent high performance, even as Nyris's data volume grows.
|
- **Seamless Scalability**: Qdrant's ability to [scale effortlessly across nodes](https://qdrant.tech/documentation/scaling/distributed_deployment/) ensures consistent high performance, even as Nyris's data volume grows.
|
||||||
- **Flexible Search Options**: The availability of both graph-based and brute-force search methods offers Nyris the flexibility to tailor the search approach to specific use case requirements.
|
- **Flexible Search Options**: The availability of both graph-based and brute-force search methods offers Nyris the flexibility to tailor the search approach to specific use case requirements.
|
||||||
- **Versatile Data Handling**: Qdrant imposes almost no restrictions on data types and vector sizes, allowing Nyris to manage diverse and complex datasets effectively.
|
- **Versatile Data Handling**: Qdrant imposes almost no restrictions on data types and vector sizes, allowing Nyris to manage diverse and complex datasets effectively.
|
||||||
- **Built with Rust**: The use of [Rust](https://qdrant.tech/articles/why-rust/) ensures superior performance and future-proofing, while its open-source nature allows Nyris to inspect and customize the code as necessary.
|
- **Built with Rust**: The use of [Rust](https://qdrant.tech/articles/why-rust/) ensures superior performance and future-proofing, while its open-source nature allows Nyris to inspect and customize the code as necessary.
|
||||||
|
|||||||
@@ -28,7 +28,7 @@ partition: case-studies
|
|||||||
|
|
||||||
As part of this development, the Voiceflow engineering team was looking for a [vector database](/qdrant-vector-database/) solution to power their RAG setup. They evaluated various vector databases based on several key factors:
|
As part of this development, the Voiceflow engineering team was looking for a [vector database](/qdrant-vector-database/) solution to power their RAG setup. They evaluated various vector databases based on several key factors:
|
||||||
|
|
||||||
- **Performance**: The ability to [handle the scale](/documentation/distributed_deployment/) required by Voiceflow, supporting hundreds of thousands of projects efficiently.
|
- **Performance**: The ability to [handle the scale](/documentation/scaling/distributed_deployment/) required by Voiceflow, supporting hundreds of thousands of projects efficiently.
|
||||||
- **Metadata**: The capability to tag data and chunks and retrieve based on those values, essential for organizing and accessing specific information swiftly.
|
- **Metadata**: The capability to tag data and chunks and retrieve based on those values, essential for organizing and accessing specific information swiftly.
|
||||||
- **Managed Solution**: The availability of a [managed service](/documentation/cloud/) with automated maintenance, scaling, and security, freeing the team from infrastructure concerns.
|
- **Managed Solution**: The availability of a [managed service](/documentation/cloud/) with automated maintenance, scaling, and security, freeing the team from infrastructure concerns.
|
||||||
|
|
||||||
|
|||||||
@@ -46,7 +46,7 @@ Qdrant is highly scalable and performant: it can handle billions of vectors effi
|
|||||||
|
|
||||||
- **Advanced Similarity Search:** Qdrant supports various similarity [search](https://qdrant.tech/documentation/search/search/) metrics like dot product, cosine similarity, Euclidean distance, and Manhattan distance. You can store additional information along with vectors, known as [payload](https://qdrant.tech/documentation/manage-data/payload/) in Qdrant terminology. A payload is any JSON formatted data.
|
- **Advanced Similarity Search:** Qdrant supports various similarity [search](https://qdrant.tech/documentation/search/search/) metrics like dot product, cosine similarity, Euclidean distance, and Manhattan distance. You can store additional information along with vectors, known as [payload](https://qdrant.tech/documentation/manage-data/payload/) in Qdrant terminology. A payload is any JSON formatted data.
|
||||||
- **Built Using Rust:** Qdrant is built with Rust, and leverages its performance and efficiency. Rust is famed for its [memory safety](https://arxiv.org/abs/2206.05503) without the overhead of a garbage collector, and rivals C and C++ in speed.
|
- **Built Using Rust:** Qdrant is built with Rust, and leverages its performance and efficiency. Rust is famed for its [memory safety](https://arxiv.org/abs/2206.05503) without the overhead of a garbage collector, and rivals C and C++ in speed.
|
||||||
- **Scaling and Multitenancy**: Qdrant supports both vertical and horizontal scaling and uses the Raft consensus protocol for [distributed deployments](https://qdrant.tech/documentation/distributed_deployment/). Developers can run Qdrant clusters with replicas and shards, and seamlessly scale to handle large datasets. Qdrant also supports [multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/) where developers can create single collections and partition them using payload.
|
- **Scaling and Multitenancy**: Qdrant supports both vertical and horizontal scaling and uses the Raft consensus protocol for [distributed deployments](https://qdrant.tech/documentation/scaling/distributed_deployment/). Developers can run Qdrant clusters with replicas and shards, and seamlessly scale to handle large datasets. Qdrant also supports [multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/) where developers can create single collections and partition them using payload.
|
||||||
- **Payload Indexing and Filtering:** Just as Qdrant allows attaching any JSON payload to vectors, it also supports payload indexing and [filtering](https://qdrant.tech/documentation/search/filtering/) with a wide range of data types and query conditions, including keyword matching, full-text filtering, numerical ranges, nested object filters, and [geo](https://qdrant.tech/documentation/search/filtering/#geo)filtering.
|
- **Payload Indexing and Filtering:** Just as Qdrant allows attaching any JSON payload to vectors, it also supports payload indexing and [filtering](https://qdrant.tech/documentation/search/filtering/) with a wide range of data types and query conditions, including keyword matching, full-text filtering, numerical ranges, nested object filters, and [geo](https://qdrant.tech/documentation/search/filtering/#geo)filtering.
|
||||||
- **Hybrid Search with Sparse Vectors:** Qdrant supports both dense and [sparse vectors](https://qdrant.tech/articles/sparse-vectors/), thereby enabling hybrid search capabilities. Sparse vectors are numerical representations of data where most of the elements are zero. Developers can combine search results from dense and sparse vectors, where sparse vectors ensure that results containing the specific keywords are returned and dense vectors identify semantically similar results.
|
- **Hybrid Search with Sparse Vectors:** Qdrant supports both dense and [sparse vectors](https://qdrant.tech/articles/sparse-vectors/), thereby enabling hybrid search capabilities. Sparse vectors are numerical representations of data where most of the elements are zero. Developers can combine search results from dense and sparse vectors, where sparse vectors ensure that results containing the specific keywords are returned and dense vectors identify semantically similar results.
|
||||||
- **Built-In Vector Quantization:** Qdrant offers three different [quantization](https://qdrant.tech/documentation/manage-data/quantization/) options to developers to optimize resource usage. Scalar quantization balances accuracy, speed, and compression by converting 32-bit floats to 8-bit integers. Binary quantization, the fastest method, significantly reduces memory usage. Product quantization offers the highest compression, and is perfect for memory-constrained scenarios.
|
- **Built-In Vector Quantization:** Qdrant offers three different [quantization](https://qdrant.tech/documentation/manage-data/quantization/) options to developers to optimize resource usage. Scalar quantization balances accuracy, speed, and compression by converting 32-bit floats to 8-bit integers. Binary quantization, the fastest method, significantly reduces memory usage. Product quantization offers the highest compression, and is perfect for memory-constrained scenarios.
|
||||||
|
|||||||
@@ -32,7 +32,7 @@ Additionally, version 1.16 introduces a new conditional update API, facilitating
|
|||||||
Multitenancy is a common requirement for SaaS applications, where multiple customers (tenants) share the same database instance. In Qdrant, when an instance is shared between multiple users, you may need to partition vectors by user. This is done so that each user can only access their own vectors and can’t see the vectors of other users. To implement multitenancy in Qdrant, there are two main approaches:
|
Multitenancy is a common requirement for SaaS applications, where multiple customers (tenants) share the same database instance. In Qdrant, when an instance is shared between multiple users, you may need to partition vectors by user. This is done so that each user can only access their own vectors and can’t see the vectors of other users. To implement multitenancy in Qdrant, there are two main approaches:
|
||||||
|
|
||||||
- [Payload-based multitenancy](/documentation/manage-data/multitenancy/), which works well when you have a large number of small tenants. This causes practically no overhead. Quite the opposite: a query with a tenant payload filter can be faster than a full search.
|
- [Payload-based multitenancy](/documentation/manage-data/multitenancy/), which works well when you have a large number of small tenants. This causes practically no overhead. Quite the opposite: a query with a tenant payload filter can be faster than a full search.
|
||||||
- [Shard-based multitenancy](/documentation/distributed_deployment/#user-defined-sharding), designed for when you have a smaller number of larger tenants. This works well when each tenant requires isolation and dedicated resources. Separating tenants by shard prevents a classic noisy neighbor problem where a single high-volume tenant can force the cluster to scale for everyone, increasing costs and reducing performance for smaller tenants. However, shard-based multitenancy is not a good solution when you have a large number of small tenants, as each shard incurs some overhead.
|
- [Shard-based multitenancy](/documentation/scaling/distributed_deployment/#user-defined-sharding), designed for when you have a smaller number of larger tenants. This works well when each tenant requires isolation and dedicated resources. Separating tenants by shard prevents a classic noisy neighbor problem where a single high-volume tenant can force the cluster to scale for everyone, increasing costs and reducing performance for smaller tenants. However, shard-based multitenancy is not a good solution when you have a large number of small tenants, as each shard incurs some overhead.
|
||||||
|
|
||||||
Real-world usage patterns often fall between these two use cases. It's common to have a small number of large tenants and a huge tail of smaller ones. You may even have tenants that grow over time, starting small and eventually becoming large enough to require dedicated resources.
|
Real-world usage patterns often fall between these two use cases. It's common to have a small number of large tenants and a huge tail of smaller ones. You may even have tenants that grow over time, starting small and eventually becoming large enough to require dedicated resources.
|
||||||
|
|
||||||
@@ -40,7 +40,7 @@ In version 1.16, Qdrant can now efficiently combine the two multitenancy approac
|
|||||||
|
|
||||||
The main principles behind Tiered Multitenancy are:
|
The main principles behind Tiered Multitenancy are:
|
||||||
|
|
||||||
- [User-defined Sharding](/documentation/distributed_deployment/#user-defined-sharding) allows you to create named shards within a collection. It enables you to isolate large tenants into their own shards. A multitenant collection can consist of a shared "fallback" shard for small tenants and multiple dedicated shards for large tenants.
|
- [User-defined Sharding](/documentation/scaling/distributed_deployment/#user-defined-sharding) allows you to create named shards within a collection. It enables you to isolate large tenants into their own shards. A multitenant collection can consist of a shared "fallback" shard for small tenants and multiple dedicated shards for large tenants.
|
||||||
- **Fallback shards** - a special routing mechanism that allows Qdrant to route a request to either a dedicated shard (if it exists) or to a shared fallback shard. This keeps requests unified, without the need to know whether a tenant is dedicated or shared.
|
- **Fallback shards** - a special routing mechanism that allows Qdrant to route a request to either a dedicated shard (if it exists) or to a shared fallback shard. This keeps requests unified, without the need to know whether a tenant is dedicated or shared.
|
||||||
- [Tenant promotion](/documentation/manage-data/multitenancy/#promote-tenant-to-dedicated-shard) - a mechanism that makes it possible to "promote" tenants from the shared Fallback Shard to their own dedicated shard when they grow large enough. This process is based on Qdrant’s internal shard transfer mechanism, which makes promotion completely transparent for the application. Both read and write requests are supported during the promotion process.
|
- [Tenant promotion](/documentation/manage-data/multitenancy/#promote-tenant-to-dedicated-shard) - a mechanism that makes it possible to "promote" tenants from the shared Fallback Shard to their own dedicated shard when they grow large enough. This process is based on Qdrant’s internal shard transfer mechanism, which makes promotion completely transparent for the application. Both read and write requests are supported during the promotion process.
|
||||||
|
|
||||||
|
|||||||
@@ -115,7 +115,7 @@ Many people have been asking about point filtering in web UI. And now it's back,
|
|||||||
As an open source project, we welcome contributions from the Qdrant community. This release features two contributions from community members:
|
As an open source project, we welcome contributions from the Qdrant community. This release features two contributions from community members:
|
||||||
|
|
||||||
- Not all payload field indexes are used in combination with dense vector queries. With this release, you can [specify whether individual payload field indexes should be reflected in the HNSW index](/documentation/manage-data/indexing/#disable-the-creation-of-extra-edges-for-payload-fields).
|
- Not all payload field indexes are used in combination with dense vector queries. With this release, you can [specify whether individual payload field indexes should be reflected in the HNSW index](/documentation/manage-data/indexing/#disable-the-creation-of-extra-edges-for-payload-fields).
|
||||||
- A new API endpoint is available to [list all user-defined shard keys](/documentation/distributed_deployment/#user-defined-sharding).
|
- A new API endpoint is available to [list all user-defined shard keys](/documentation/scaling/distributed_deployment/#user-defined-sharding).
|
||||||
|
|
||||||
Additionally, this release adds the following features:
|
Additionally, this release adds the following features:
|
||||||
|
|
||||||
|
|||||||
@@ -40,7 +40,7 @@ We highly recommend this feature to enterprises using [Qdrant Hybrid Cloud](/hyb
|
|||||||
|
|
||||||
## Faster shard transfers on node recovery
|
## Faster shard transfers on node recovery
|
||||||
|
|
||||||
We now offer a streamlined approach to [data synchronization between shards](/documentation/distributed_deployment/#shard-transfer-method) during node upgrades or recovery processes. Traditional methods used to transfer the entire dataset, but our new `wal_delta` method focuses solely on transmitting the difference between two existing shards. By leveraging the Write-Ahead Log (WAL) of both shards, this method selectively transmits missed operations to the target shard, ensuring data consistency.
|
We now offer a streamlined approach to [data synchronization between shards](/documentation/scaling/distributed_deployment/#shard-transfer-method) during node upgrades or recovery processes. Traditional methods used to transfer the entire dataset, but our new `wal_delta` method focuses solely on transmitting the difference between two existing shards. By leveraging the Write-Ahead Log (WAL) of both shards, this method selectively transmits missed operations to the target shard, ensuring data consistency.
|
||||||
|
|
||||||
In some cases, where transfers can take hours, this update **reduces transfers down to a few minutes.**
|
In some cases, where transfers can take hours, this update **reduces transfers down to a few minutes.**
|
||||||
|
|
||||||
@@ -48,7 +48,7 @@ The advantages of this approach are twofold:
|
|||||||
1. **It is faster** since only the differential data is transmitted, avoiding the transfer of redundant information.
|
1. **It is faster** since only the differential data is transmitted, avoiding the transfer of redundant information.
|
||||||
2. It upholds robust **ordering guarantees**, crucial for applications reliant on strict sequencing.
|
2. It upholds robust **ordering guarantees**, crucial for applications reliant on strict sequencing.
|
||||||
|
|
||||||
For more details on how this works, check out the [shard transfer documentation](/documentation/distributed_deployment/#shard-transfer-method).
|
For more details on how this works, check out the [shard transfer documentation](/documentation/scaling/distributed_deployment/#shard-transfer-method).
|
||||||
|
|
||||||
> **Note:** There are limitations to consider. First, this method only works with existing shards. Second, while the WALs typically retain recent operations, their capacity is finite, potentially impeding the transfer process if exceeded. Nevertheless, for scenarios like rapid node restarts or upgrades, where the WAL content remains manageable, WAL delta transfer is an efficient solution.
|
> **Note:** There are limitations to consider. First, this method only works with existing shards. Second, while the WALs typically retain recent operations, their capacity is finite, potentially impeding the transfer process if exceeded. Nevertheless, for scenarios like rapid node restarts or upgrades, where the WAL content remains manageable, WAL delta transfer is an efficient solution.
|
||||||
|
|
||||||
|
|||||||
@@ -145,7 +145,7 @@ The vector index in Qdrant employs the Hierarchical Navigable Small World (HNSW)
|
|||||||
|
|
||||||
### Scalability
|
### Scalability
|
||||||
|
|
||||||
For massive datasets and demanding workloads, Qdrant supports [distributed deployment](/documentation/distributed_deployment/) from v0.8.0. In this mode, you can set up a Qdrant cluster and distribute data across multiple nodes, enabling you to maintain high performance and availability even under increased workloads. Clusters support sharding and replication, and harness the Raft consensus algorithm to manage node coordination.
|
For massive datasets and demanding workloads, Qdrant supports [distributed deployment](/documentation/scaling/distributed_deployment/) from v0.8.0. In this mode, you can set up a Qdrant cluster and distribute data across multiple nodes, enabling you to maintain high performance and availability even under increased workloads. Clusters support sharding and replication, and harness the Raft consensus algorithm to manage node coordination.
|
||||||
|
|
||||||
Qdrant also supports vector [quantization](/documentation/manage-data/quantization/) to reduce memory footprint and speed up vector similarity searches, making it very effective for large-scale applications where efficient resource management is critical.
|
Qdrant also supports vector [quantization](/documentation/manage-data/quantization/) to reduce memory footprint and speed up vector similarity searches, making it very effective for large-scale applications where efficient resource management is critical.
|
||||||
|
|
||||||
|
|||||||
@@ -112,7 +112,7 @@ Qdrant is an AI-native vector search engine for storing, indexing, and searching
|
|||||||
- [Managed Cloud](/documentation/cloud/index.md) — Qdrant as a managed service on AWS, GCP, or Azure with automatic scaling, backups, and zero-downtime upgrades.
|
- [Managed Cloud](/documentation/cloud/index.md) — Qdrant as a managed service on AWS, GCP, or Azure with automatic scaling, backups, and zero-downtime upgrades.
|
||||||
- [Hybrid Cloud](/documentation/hybrid-cloud/index.md) — Deploy into your own Kubernetes cluster while managing through Qdrant Cloud.
|
- [Hybrid Cloud](/documentation/hybrid-cloud/index.md) — Deploy into your own Kubernetes cluster while managing through Qdrant Cloud.
|
||||||
- [Private Cloud](/documentation/private-cloud/index.md) — Fully air-gapped deployment in your own Kubernetes cluster with no Qdrant Cloud connectivity required.
|
- [Private Cloud](/documentation/private-cloud/index.md) — Fully air-gapped deployment in your own Kubernetes cluster with no Qdrant Cloud connectivity required.
|
||||||
- [Distributed Deployment](/documentation/distributed_deployment/index.md) — Multi-node clusters with horizontal sharding and replication for scale and fault tolerance.
|
- [Distributed Deployment](/documentation/scaling/distributed_deployment/index.md) — Multi-node clusters with horizontal sharding and replication for scale and fault tolerance.
|
||||||
- [Security](/documentation/security/index.md) — API keys, JWT-based collection-scoped access control, TLS encryption, and network binding.
|
- [Security](/documentation/security/index.md) — API keys, JWT-based collection-scoped access control, TLS encryption, and network binding.
|
||||||
- [Configuration](/documentation/ops-configuration/index.md) — Customize Qdrant via config files and environment variables; runtime administration tools; GPU-accelerated vector indexing.
|
- [Configuration](/documentation/ops-configuration/index.md) — Customize Qdrant via config files and environment variables; runtime administration tools; GPU-accelerated vector indexing.
|
||||||
- [Monitoring & Telemetry](/documentation/ops-monitoring/index.md) — Monitor Qdrant with Prometheus and Grafana via built-in OpenMetrics endpoints.
|
- [Monitoring & Telemetry](/documentation/ops-monitoring/index.md) — Monitor Qdrant with Prometheus and Grafana via built-in OpenMetrics endpoints.
|
||||||
|
|||||||
@@ -30,7 +30,7 @@ After setting up your account, you can create a Qdrant Cluster by following the
|
|||||||
|
|
||||||
## Preparing for Production
|
## Preparing for Production
|
||||||
|
|
||||||
For a production-ready environment, consider deploying a multi-node Qdrant cluster (at least three nodes) with replication enabled. More details are available in the [Distributed Deployment](/documentation/distributed_deployment/) guide. For more information on how to create a production-ready cluster, see our [Vector Search in Production](/articles/vector-search-production/) article.
|
For a production-ready environment, consider deploying a multi-node Qdrant cluster (at least three nodes) with replication enabled. More details are available in the [Distributed Deployment](/documentation/scaling/distributed_deployment/) guide. For more information on how to create a production-ready cluster, see our [Vector Search in Production](/articles/vector-search-production/) article.
|
||||||
|
|
||||||
If you are looking to optimize costs, you can reduce memory usage through [Quantization](/documentation/manage-data/quantization/) or by [offloading vectors to disk](/documentation/manage-data/storage/#configuring-memmap-storage).
|
If you are looking to optimize costs, you can reduce memory usage through [Quantization](/documentation/manage-data/quantization/) or by [offloading vectors to disk](/documentation/manage-data/storage/#configuring-memmap-storage).
|
||||||
|
|
||||||
|
|||||||
@@ -18,7 +18,7 @@ Qdrant Cloud offers an optional premium tier for customers who require additiona
|
|||||||
* **Single Sign-On (SSO)**: Premium customers can use their existing SSO provider to manage access to Qdrant Cloud.
|
* **Single Sign-On (SSO)**: Premium customers can use their existing SSO provider to manage access to Qdrant Cloud.
|
||||||
* **VPC Private Links**: Premium customers can connect their Qdrant Cloud clusters to their VPCs using private links.
|
* **VPC Private Links**: Premium customers can connect their Qdrant Cloud clusters to their VPCs using private links.
|
||||||
* **Storage encryption with shared keys**: Premium customers can encrypt their data at rest using their own keys.
|
* **Storage encryption with shared keys**: Premium customers can encrypt their data at rest using their own keys.
|
||||||
* **Topology Aware Multi-AZ Setup**: Premium customers can deploy their clusters across multiple availability zones for higher availability and resilience. This guarantees a **99.95% uptime SLA** for Multi-AZ clusters.
|
* **Topology Aware Multi-AZ Setup**: Premium customers can deploy their clusters across multiple availability zones for higher availability and resilience. This guarantees a **99.95% uptime SLA** for Multi-AZ clusters. Multi-AZ is independent of replication factor; see [Multi-AZ Deployments](/documentation/scaling/resilience/#multi-az-deployments) for the distinction.
|
||||||
|
|
||||||
Please refer to the [Qdrant Cloud SLA](https://qdrant.to/sla/) for a detailed definition on uptime and support SLAs.
|
Please refer to the [Qdrant Cloud SLA](https://qdrant.to/sla/) for a detailed definition on uptime and support SLAs.
|
||||||
|
|
||||||
|
|||||||
@@ -282,7 +282,7 @@ The account owner will receive automatic alerts via email if your cluster has an
|
|||||||
|
|
||||||
**Where can I learn more about this alert?**
|
**Where can I learn more about this alert?**
|
||||||
|
|
||||||
Learn more about distributed deployments and resharding [here](/documentation/distributed_deployment/#resharding).
|
Learn more about distributed deployments and resharding [here](/documentation/scaling/distributed_deployment/#resharding).
|
||||||
|
|
||||||
Learn more about cloud rebalancing [here](/documentation/cloud/configure-cluster/#shard-rebalancing).
|
Learn more about cloud rebalancing [here](/documentation/cloud/configure-cluster/#shard-rebalancing).
|
||||||
|
|
||||||
|
|||||||
@@ -29,7 +29,7 @@ Vertical scaling can be an effective way to improve the performance of a cluster
|
|||||||
|
|
||||||
In such cases, horizontal scaling may be a more effective solution.
|
In such cases, horizontal scaling may be a more effective solution.
|
||||||
|
|
||||||
Horizontal scaling is the process of increasing the capacity of a cluster by adding more nodes and distributing the load and data among them. The horizontal scaling at Qdrant starts on the collection level. You have to choose the number of shards you want to distribute your collection around while creating the collection. Please refer to the [sharding documentation](/documentation/distributed_deployment/#sharding) section for details.
|
Horizontal scaling is the process of increasing the capacity of a cluster by adding more nodes and distributing the load and data among them. The horizontal scaling at Qdrant starts on the collection level. You have to choose the number of shards you want to distribute your collection around while creating the collection. Please refer to the [sharding documentation](/documentation/scaling/distributed_deployment/#sharding) section for details.
|
||||||
|
|
||||||
When scaling up horizontally, the cloud platform will automatically rebalance all available shards across nodes to ensure that the data is evenly distributed. See [Configuring Clusters](/documentation/cloud/configure-cluster/#shard-rebalancing) for more details.
|
When scaling up horizontally, the cloud platform will automatically rebalance all available shards across nodes to ensure that the data is evenly distributed. See [Configuring Clusters](/documentation/cloud/configure-cluster/#shard-rebalancing) for more details.
|
||||||
|
|
||||||
|
|||||||
@@ -104,13 +104,14 @@ To create a production-ready cluster, you need to ensure the following:
|
|||||||
|
|
||||||
**High Availability**
|
**High Availability**
|
||||||
|
|
||||||
Your cluster should have at least 3 nodes, and each collection should have a replication factor of at least 2. This ensures that is one node fails, or is restarted due to maintenance, a version upgrade, or a scaling operation, that the cluster remains fully operational. You can ensure this by checking the **High Availability** checkbox when creating a cluster.
|
Your cluster should have at least 3 nodes, and each collection should have a replication factor of at least 2. This ensures that is one node fails, or is restarted due to maintenance, a version upgrade, or a scaling operation, that the cluster remains fully operational. You can ensure this by checking the **High Availability** checkbox when creating a cluster. A cluster without these settings runs with a single replica of each shard by default, which gets none of these guarantees; see [Setting Up a Resilient Qdrant Cluster](/documentation/scaling/resilience/#setting-up-a-resilient-qdrant-cluster) for details.
|
||||||
|
|
||||||
**Multi AZ Deployment (Premium only)**
|
**Multi AZ Deployment (Premium only)**
|
||||||
|
|
||||||
Premium tier customers can choose to deploy their cluster across multiple availability zones. This ensures that if one availability zone goes down, the cluster remains operational. You can ensure this by checking the **Multi AZ Deployment** checkbox when creating a cluster. This can not be changed later.
|
Premium tier customers can choose to deploy their cluster across multiple availability zones. This ensures that if one availability zone goes down, the cluster remains operational. You can ensure this by checking the **Multi AZ Deployment** checkbox when creating a cluster. This can not be changed later.
|
||||||
Multi AZ clusters need a minimum of 3 nodes, and can only scale to a multiple of 3 (e.g. 3, 6, 9, etc.) to ensure that nodes are evenly distributed across availability zones.
|
Multi AZ clusters need a minimum of 3 nodes, and can only scale to a multiple of 3 (e.g. 3, 6, 9, etc.) to ensure that nodes are evenly distributed across availability zones.
|
||||||
Your collections should have a replication factor of at least 2 (better 3) to ensure that all data is available across availability zones, so the outage of one zone does not compromise the availability of the cluster. Shards will be automatically distributed across availability zones, so that each shard has a replica in another availability zone. Traffic is routed between zones automatically, so that the cluster remains available even if one zone goes down.
|
Your collections should have a replication factor of at least 2 (better 3) to ensure that all data is available across availability zones, so the outage of one zone does not compromise the availability of the cluster. Shards will be automatically distributed across availability zones, so that each shard has a replica in another availability zone. Traffic is routed between zones automatically, so that the cluster remains available even if one zone goes down.
|
||||||
|
Replication factor and Multi-AZ are independent settings: a replicated cluster does not automatically span multiple zones unless Multi-AZ is enabled. See [Multi-AZ Deployments](/documentation/scaling/resilience/#multi-az-deployments) for the distinction.
|
||||||
|
|
||||||
**Disk Speed (AWS only)**
|
**Disk Speed (AWS only)**
|
||||||
|
|
||||||
@@ -126,7 +127,7 @@ You should create a backup schedule for your cluster. This ensures that you can
|
|||||||
|
|
||||||
**Collection Sharding**
|
**Collection Sharding**
|
||||||
|
|
||||||
To allow your cluster to easily scale horizontally, you should configure at least twice as many shards per collection than the number of nodes in your cluster. You can configure the number of shards when creating a collection. See [**Sharding**](/documentation/distributed_deployment/#sharding) for more information.
|
To allow your cluster to easily scale horizontally, you should configure at least twice as many shards per collection than the number of nodes in your cluster. You can configure the number of shards when creating a collection. See [**Sharding**](/documentation/scaling/distributed_deployment/#sharding) for more information.
|
||||||
|
|
||||||
If you did not configure enough shards in a collection, you can use the [**Resharding**](/documentation/cloud/cluster-scaling/#resharding) feature to change the number of shards in an existing collection.
|
If you did not configure enough shards in a collection, you can use the [**Resharding**](/documentation/cloud/cluster-scaling/#resharding) feature to change the number of shards in an existing collection.
|
||||||
|
|
||||||
|
|||||||
@@ -115,7 +115,7 @@ build:
|
|||||||
## Self-Hosted
|
## Self-Hosted
|
||||||
|
|
||||||
- [Installation](/documentation/installation/index.md) — Install Qdrant via Docker, Kubernetes, or binary on Linux, macOS, or Windows.
|
- [Installation](/documentation/installation/index.md) — Install Qdrant via Docker, Kubernetes, or binary on Linux, macOS, or Windows.
|
||||||
- [Distributed Deployment](/documentation/distributed_deployment/index.md) — Multi-node clusters with horizontal sharding and replication for scale and fault tolerance.
|
- [Distributed Deployment](/documentation/scaling/distributed_deployment/index.md) — Multi-node clusters with horizontal sharding and replication for scale and fault tolerance.
|
||||||
- [Capacity Planning](/documentation/capacity-planning/index.md) — Estimate RAM and disk requirements for vectors, payloads, indexes, and replication factors.
|
- [Capacity Planning](/documentation/capacity-planning/index.md) — Estimate RAM and disk requirements for vectors, payloads, indexes, and replication factors.
|
||||||
- [Snapshots](/documentation/snapshots/index.md) — Back up and restore collections for disaster recovery and cross-cluster replication.
|
- [Snapshots](/documentation/snapshots/index.md) — Back up and restore collections for disaster recovery and cross-cluster replication.
|
||||||
- [Production Checklist](/documentation/production-checklist/index.md) — Pre-launch review of sharding, replication, quantization, load balancing, and observability.
|
- [Production Checklist](/documentation/production-checklist/index.md) — Pre-launch review of sharding, replication, quantization, load balancing, and observability.
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
@@ -51,7 +51,7 @@ Each Qdrant instance requires three open ports:
|
|||||||
|
|
||||||
* `6333` - For the HTTP API, for the [Monitoring](/documentation/ops-monitoring/monitoring/) health and metrics endpoints
|
* `6333` - For the HTTP API, for the [Monitoring](/documentation/ops-monitoring/monitoring/) health and metrics endpoints
|
||||||
* `6334` - For the [gRPC](/documentation/interfaces/#grpc-interface) API
|
* `6334` - For the [gRPC](/documentation/interfaces/#grpc-interface) API
|
||||||
* `6335` - For [Distributed deployment](/documentation/distributed_deployment/)
|
* `6335` - For [Distributed deployment](/documentation/scaling/distributed_deployment/)
|
||||||
|
|
||||||
All Qdrant instances in a cluster must be able to:
|
All Qdrant instances in a cluster must be able to:
|
||||||
|
|
||||||
@@ -85,7 +85,7 @@ We provide a Qdrant Enterprise Operator for Kubernetes installations as part of
|
|||||||
|
|
||||||
### Kubernetes
|
### Kubernetes
|
||||||
|
|
||||||
You can use a ready-made [Helm Chart](https://helm.sh/docs/) to run Qdrant in your Kubernetes cluster. While it is possible to deploy Qdrant in a distributed setup with the Helm chart, it does not come with the same level of features for zero-downtime upgrades, up and down-scaling, monitoring, logging, and backup and disaster recovery as the Qdrant Cloud offering or the Qdrant Private Cloud Enterprise Operator. Instead you must manage and set this up [yourself](/documentation/distributed_deployment/). Support for the Helm chart is limited to community support.
|
You can use a ready-made [Helm Chart](https://helm.sh/docs/) to run Qdrant in your Kubernetes cluster. While it is possible to deploy Qdrant in a distributed setup with the Helm chart, it does not come with the same level of features for zero-downtime upgrades, up and down-scaling, monitoring, logging, and backup and disaster recovery as the Qdrant Cloud offering or the Qdrant Private Cloud Enterprise Operator. Instead you must manage and set this up [yourself](/documentation/scaling/distributed_deployment/). Support for the Helm chart is limited to community support.
|
||||||
|
|
||||||
The following table gives you an overview about the feature differences between the Qdrant Cloud and the Helm chart:
|
The following table gives you an overview about the feature differences between the Qdrant Cloud and the Helm chart:
|
||||||
|
|
||||||
@@ -128,7 +128,7 @@ In addition, you have to make sure:
|
|||||||
|
|
||||||
* To use a performant [persistent storage](#storage) for your data
|
* To use a performant [persistent storage](#storage) for your data
|
||||||
* To configure the [security settings](/documentation/security/) for your deployment
|
* To configure the [security settings](/documentation/security/) for your deployment
|
||||||
* To set up and configure Qdrant on multiple nodes for a highly available [distributed deployment](/documentation/distributed_deployment/)
|
* To set up and configure Qdrant on multiple nodes for a highly available [distributed deployment](/documentation/scaling/distributed_deployment/)
|
||||||
* To set up a load balancer for your Qdrant cluster
|
* To set up a load balancer for your Qdrant cluster
|
||||||
* To create a [backup and disaster recovery strategy](/documentation/snapshots/) for your data
|
* To create a [backup and disaster recovery strategy](/documentation/snapshots/) for your data
|
||||||
* To integrate Qdrant with your [monitoring](/documentation/ops-monitoring/monitoring/) and logging solutions
|
* To integrate Qdrant with your [monitoring](/documentation/ops-monitoring/monitoring/) and logging solutions
|
||||||
|
|||||||
@@ -44,7 +44,7 @@ In addition to the required options, you can also specify custom values for the
|
|||||||
* `hnsw_config` - see [indexing](/documentation/manage-data/indexing/#vector-index) for details.
|
* `hnsw_config` - see [indexing](/documentation/manage-data/indexing/#vector-index) for details.
|
||||||
* `wal_config` - Write-Ahead-Log related configuration. See more details about [WAL](/documentation/manage-data/storage/#versioning).
|
* `wal_config` - Write-Ahead-Log related configuration. See more details about [WAL](/documentation/manage-data/storage/#versioning).
|
||||||
* `optimizers_config` - see [optimizer](/documentation/ops-optimization/optimizer/) for details.
|
* `optimizers_config` - see [optimizer](/documentation/ops-optimization/optimizer/) for details.
|
||||||
* `shard_number` - which defines how many shards the collection should have. See [distributed deployment](/documentation/distributed_deployment/#sharding) section for details.
|
* `shard_number` - which defines how many shards the collection should have. See [distributed deployment](/documentation/scaling/distributed_deployment/#sharding) section for details.
|
||||||
* `on_disk_payload` - defines where to store payload data. If `true` - payload will be stored on disk only. Might be useful for limiting the RAM usage in case of large payload.
|
* `on_disk_payload` - defines where to store payload data. If `true` - payload will be stored on disk only. Might be useful for limiting the RAM usage in case of large payload.
|
||||||
* `quantization_config` - see [quantization](/documentation/manage-data/quantization/#setting-up-quantization-in-qdrant) for details.
|
* `quantization_config` - see [quantization](/documentation/manage-data/quantization/#setting-up-quantization-in-qdrant) for details.
|
||||||
* `strict_mode_config` - see [strict mode](/documentation/ops-configuration/administration/#strict-mode) for details.
|
* `strict_mode_config` - see [strict mode](/documentation/ops-configuration/administration/#strict-mode) for details.
|
||||||
|
|||||||
@@ -81,7 +81,7 @@ To address this problem, in v1.16.0 Qdrant provides a built-in mechanism for tie
|
|||||||
With tiered multitenancy, you can implement two levels of tenant isolation within a single collection, keeping small tenants together inside a shared Shard, while isolating large tenants into their own dedicated Shards.
|
With tiered multitenancy, you can implement two levels of tenant isolation within a single collection, keeping small tenants together inside a shared Shard, while isolating large tenants into their own dedicated Shards.
|
||||||
There are 3 components in Qdrant, that allows you to implement tiered multitenancy:
|
There are 3 components in Qdrant, that allows you to implement tiered multitenancy:
|
||||||
|
|
||||||
- [**User-defined Sharding**](/documentation/distributed_deployment/#user-defined-sharding) allows you to create named Shards within a collection. It allows to isolate large tenants into their own Shards.
|
- [**User-defined Sharding**](/documentation/scaling/distributed_deployment/#user-defined-sharding) allows you to create named Shards within a collection. It allows to isolate large tenants into their own Shards.
|
||||||
- **Fallback shards** - a special routing mechanism that allows to route request to either a dedicated Shard (if it exists) or to a shared Fallback Shard. It allows to keep requests unified, without the need to know whether a tenant is dedicated or shared.
|
- **Fallback shards** - a special routing mechanism that allows to route request to either a dedicated Shard (if it exists) or to a shared Fallback Shard. It allows to keep requests unified, without the need to know whether a tenant is dedicated or shared.
|
||||||
- **Tenant promotion** - a mechanism that allows to move tenants from the shared Fallback Shard to their own dedicated Shard when they grow large enough. This process is based on Qdrant's internal shard transfer mechanism, which makes promotion completely transparent for the application. Both read and write requests are supported during the promotion process.
|
- **Tenant promotion** - a mechanism that allows to move tenants from the shared Fallback Shard to their own dedicated Shard when they grow large enough. This process is based on Qdrant's internal shard transfer mechanism, which makes promotion completely transparent for the application. Both read and write requests are supported during the promotion process.
|
||||||
|
|
||||||
|
|||||||
@@ -315,7 +315,7 @@ storage:
|
|||||||
# Default shard transfer method to use if none is defined.
|
# Default shard transfer method to use if none is defined.
|
||||||
# If null - don't have a shard transfer preference, choose automatically.
|
# If null - don't have a shard transfer preference, choose automatically.
|
||||||
# If stream_records, snapshot or wal_delta - prefer this specific method.
|
# If stream_records, snapshot or wal_delta - prefer this specific method.
|
||||||
# More info: https://qdrant.tech/documentation/distributed_deployment/#shard-transfer-method
|
# More info: https://qdrant.tech/documentation/scaling/distributed_deployment/#shard-transfer-method
|
||||||
shard_transfer_method: null
|
shard_transfer_method: null
|
||||||
|
|
||||||
# Default parameters for collections
|
# Default parameters for collections
|
||||||
|
|||||||
@@ -158,7 +158,7 @@ As a reference point: 1 KB ≈ one vector of 256 dimensions. A value of `100000`
|
|||||||
|
|
||||||
The previous steps all operate within a fixed hardware budget: they reallocate CPU between the optimizer and queries. If you've exhausted those options and latency is still too high, the answer is more capacity.
|
The previous steps all operate within a fixed hardware budget: they reallocate CPU between the optimizer and queries. If you've exhausted those options and latency is still too high, the answer is more capacity.
|
||||||
|
|
||||||
[Adding nodes to your cluster](/documentation/distributed_deployment/) and increasing the replication factor distributes read traffic across more peers. Because every replica of a shard contains the same data, Qdrant can route read requests to any of them. More replicas mean more vCPUs available to serve queries, and the optimizer on each node only contends with the query load that node carries instead of the full cluster load.
|
[Adding nodes to your cluster](/documentation/scaling/distributed_deployment/) and increasing the replication factor distributes read traffic across more peers. Because every replica of a shard contains the same data, Qdrant can route read requests to any of them. More replicas mean more vCPUs available to serve queries, and the optimizer on each node only contends with the query load that node carries instead of the full cluster load.
|
||||||
|
|
||||||
For example, a collection with three shards and a replication factor of two has six replicas total. On a three-node cluster, each node handles two replicas. A read request hits one replica per shard, so query load is spread evenly across all three nodes.
|
For example, a collection with three shards and a replication factor of two has six replicas total. On a three-node cluster, each node handles two replicas. A read request hits one replica per shard, so query load is spread evenly across all three nodes.
|
||||||
|
|
||||||
@@ -179,5 +179,5 @@ Like Step 8, this step adds capacity rather than reallocating it. Where horizont
|
|||||||
- [Optimizer](/documentation/ops-optimization/optimizer/) covers all optimizer settings referenced in this guide, including how to monitor deferred points.
|
- [Optimizer](/documentation/ops-optimization/optimizer/) covers all optimizer settings referenced in this guide, including how to monitor deferred points.
|
||||||
- [Low-Latency Search](/documentation/search/low-latency-search/) covers delayed fan-outs and other techniques for reducing search latency.
|
- [Low-Latency Search](/documentation/search/low-latency-search/) covers delayed fan-outs and other techniques for reducing search latency.
|
||||||
- [Qdrant under the Hood: io_uring](/articles/io_uring/) explains how async I/O works in Qdrant.
|
- [Qdrant under the Hood: io_uring](/articles/io_uring/) explains how async I/O works in Qdrant.
|
||||||
- [Distributed Deployment](/documentation/distributed_deployment/) covers horizontal scaling with shards and replicas.
|
- [Distributed Deployment](/documentation/scaling/distributed_deployment/) covers horizontal scaling with shards and replicas.
|
||||||
- [Bulk Upload](/documentation/manage-data/bulk-upload/) covers best practices for high-throughput ingestion.
|
- [Bulk Upload](/documentation/manage-data/bulk-upload/) covers best practices for high-throughput ingestion.
|
||||||
|
|||||||
@@ -52,7 +52,7 @@ Qdrant collections are designed for horizontal and vertical scaling. You can lea
|
|||||||
* [Points](/documentation/manage-data/points/)
|
* [Points](/documentation/manage-data/points/)
|
||||||
* [Indexing](/documentation/manage-data/indexing/)
|
* [Indexing](/documentation/manage-data/indexing/)
|
||||||
* [Storage](/documentation/manage-data/storage/)
|
* [Storage](/documentation/manage-data/storage/)
|
||||||
* [Distributed Deployment](/documentation/distributed_deployment/)
|
* [Distributed Deployment](/documentation/scaling/distributed_deployment/)
|
||||||
* [Strict Mode](/documentation/ops-configuration/administration/#strict-mode)
|
* [Strict Mode](/documentation/ops-configuration/administration/#strict-mode)
|
||||||
|
|
||||||
## Deployments {#deployments}
|
## Deployments {#deployments}
|
||||||
@@ -114,7 +114,7 @@ Vertical scaling has natural limits \- eventually, you'll hit the maximum capaci
|
|||||||
|
|
||||||
Qdrant uses sharding to split collections across multiple nodes, where each shard is an independent store of points. A common recommendation is to start with 12 shards, which provides flexibility to scale from 1 node up to 2, 3, 6, or 12 nodes without resharding. However, this approach can limit throughput on small clusters since each node manages multiple shards.
|
Qdrant uses sharding to split collections across multiple nodes, where each shard is an independent store of points. A common recommendation is to start with 12 shards, which provides flexibility to scale from 1 node up to 2, 3, 6, or 12 nodes without resharding. However, this approach can limit throughput on small clusters since each node manages multiple shards.
|
||||||
|
|
||||||
For optimal throughput, set `shard_number` equal to your node count (read more [here](/documentation/distributed_deployment/#sharding)). If you want to have better control over sharding, Qdrant supports [custom shards](/documentation/distributed_deployment/#user-defined-sharding).
|
For optimal throughput, set `shard_number` equal to your node count (read more [here](/documentation/scaling/distributed_deployment/#sharding)). If you want to have better control over sharding, Qdrant supports [custom shards](/documentation/scaling/distributed_deployment/#user-defined-sharding).
|
||||||
|
|
||||||
#### Replication {#replication}
|
#### Replication {#replication}
|
||||||
|
|
||||||
|
|||||||
@@ -20,6 +20,8 @@ On top of the open source Qdrant database, it allows
|
|||||||
* Extended telemetry
|
* Extended telemetry
|
||||||
* Qdrant Enterprise Support Services
|
* Qdrant Enterprise Support Services
|
||||||
|
|
||||||
|
Multi-AZ deployment is independent of replication factor: replicating your data doesn't by itself spread it across zones. See [Multi-AZ Deployments](/documentation/scaling/resilience/#multi-az-deployments) for the distinction.
|
||||||
|
|
||||||
Since there is no communication or connection with Qdrant, you are fully responsible for the entire security of the Qdrant Private Cloud installation. This also means that you do not benefit from all the integrated management and observability features of Qdrant Managed Cloud and Hybrid Cloud, such as:
|
Since there is no communication or connection with Qdrant, you are fully responsible for the entire security of the Qdrant Private Cloud installation. This also means that you do not benefit from all the integrated management and observability features of Qdrant Managed Cloud and Hybrid Cloud, such as:
|
||||||
|
|
||||||
* A central management UI and API
|
* A central management UI and API
|
||||||
|
|||||||
@@ -17,7 +17,7 @@ A practical checklist to ensure Qdrant is optimized, stable, and ready to handle
|
|||||||
Architect for scale from day one. Retrofitting these patterns onto an existing deployment is costly.
|
Architect for scale from day one. Retrofitting these patterns onto an existing deployment is costly.
|
||||||
|
|
||||||
- **Ensure you have enough shards to scale.**
|
- **Ensure you have enough shards to scale.**
|
||||||
Qdrant [scales horizontally](/documentation/distributed_deployment/) through [sharding](/documentation/distributed_deployment/#sharding). Plan for enough shards to evenly distribute your data and load across the nodes in your cluster. At a minimum, you need one shard or replica per node.
|
Qdrant [scales horizontally](/documentation/scaling/distributed_deployment/) through [sharding](/documentation/scaling/distributed_deployment/#sharding). Plan for enough shards to evenly distribute your data and load across the nodes in your cluster. At a minimum, you need one shard or replica per node.
|
||||||
|
|
||||||
- **Ensure you don't have too many shards.** While sharding is essential for scale, having too many shards can lead to performance degradation. Each collection has its own shards, so if you have many collections, you may end up with an excessive number of shards.
|
- **Ensure you don't have too many shards.** While sharding is essential for scale, having too many shards can lead to performance degradation. Each collection has its own shards, so if you have many collections, you may end up with an excessive number of shards.
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,39 @@
|
|||||||
|
---
|
||||||
|
title: Scaling & Resilience
|
||||||
|
short_description: "Scale Qdrant vertically or horizontally as your data and traffic grow, and configure replication, node count, and Multi-AZ for production resilience."
|
||||||
|
description: "Learn when to scale Qdrant vertically versus horizontally, and how to configure replication, node count, Multi-AZ, and failover for production-ready resilience."
|
||||||
|
weight: 115
|
||||||
|
partition: deploy
|
||||||
|
---
|
||||||
|
|
||||||
|
# Scaling & Resilience
|
||||||
|
|
||||||
|
When you get started with Qdrant, you typically deploy a single node: one server that stores your vectors and handles queries. That's enough for most early-stage workloads, but as your dataset and traffic grow, you'll need to scale. Qdrant gives you two options: give your node more resources (vertical scaling), or add more nodes to the cluster (horizontal scaling).
|
||||||
|
|
||||||
|
Horizontal scaling also provides resilience: deployments with multiple nodes and replicas remain available for reads and writes, even when individual nodes fail.
|
||||||
|
|
||||||
|
<aside role="status">Before adding capacity, consider whether your existing cluster can be optimized. Quantization, moving vector storage to disk, and other techniques can significantly reduce resource usage. See <a href="/documentation/ops-optimization/optimize/">Optimization</a> for a full overview.</aside>
|
||||||
|
|
||||||
|
## Vertical vs. Horizontal Scaling
|
||||||
|
|
||||||
|
Vertical scaling means adding more CPU, RAM, or disk to an existing node. A single node can typically hold up to about 100 million vectors, depending on dimensionality and quantization. RAM usage approaching 80% is the main signal that it's time to resize. See [Vertical Scaling](/documentation/scaling/vertical-scaling/) for RAM sizing guidelines and resize steps for Qdrant Cloud and self-hosted deployments.
|
||||||
|
|
||||||
|
Qdrant can run in a distributed mode, where multiple nodes operate together as a single entity called a cluster. Horizontal scaling means adding more nodes to a cluster instead of resizing existing ones. It's necessary when your data no longer fits on a single node even with quantization, or when you're bottlenecked on disk I/O. See [Horizontal Scaling](/documentation/scaling/horizontal-scaling/) to understand how distributed mode works, and [Distributed Deployment](/documentation/scaling/distributed_deployment/) for the configuration steps.
|
||||||
|
|
||||||
|
Scale vertically first: it's simpler than distributing data across a cluster, avoids network overhead, and is easy to reverse. Move to horizontal scaling once vertical scaling isn't enough. If you're running in Qdrant Cloud, [Scale Clusters](/documentation/cloud/cluster-scaling/) covers the steps for both directions.
|
||||||
|
|
||||||
|
## Resilience
|
||||||
|
|
||||||
|
Resilience comes from replicating data across multiple nodes. A single node, or a single copy of your data, has no protection against failure: if you lose it, you lose both the data and the ability to serve it. A minimal fault-tolerant cluster needs three nodes and a replication factor of two or higher. Three nodes are the minimum to form a majority for Qdrant's Raft consensus, and two or more replicas ensure a single node failure won't take your data or availability down with it.
|
||||||
|
|
||||||
|
See [Resilience](/documentation/scaling/resilience/) for how replication factor and node count determine fault tolerance and failover best practices.
|
||||||
|
|
||||||
|
When a node fails, the recovery path depends on whether the lost shards have replicas on surviving nodes. A node that restarts rejoins consensus and catches up automatically. A permanently lost node can be replaced by provisioning a new one and rebalancing shards. See [Node Failure Recovery](/documentation/scaling/node-failure-recovery/) for step-by-step procedures.
|
||||||
|
|
||||||
|
## Where to Go Next
|
||||||
|
|
||||||
|
- [Vertical Scaling](/documentation/scaling/vertical-scaling/): resize existing nodes.
|
||||||
|
- [Horizontal Scaling](/documentation/scaling/horizontal-scaling/): how Qdrant's distributed model achieves scale.
|
||||||
|
- [Resilience](/documentation/scaling/resilience/): fault tolerance, multiple availability zones, and failover best practices.
|
||||||
|
- [Consistency Guarantees](/documentation/scaling/consistency-guarantees/): write consistency factor, read consistency, and write ordering.
|
||||||
|
- [Node Failure Recovery](/documentation/scaling/node-failure-recovery/): step-by-step procedures for recovering from a failed node.
|
||||||
@@ -0,0 +1,545 @@
|
|||||||
|
---
|
||||||
|
title: Consistency Guarantees
|
||||||
|
short_description: "Configure Qdrant's write consistency factor, read consistency, and write ordering to trade throughput for stronger consistency guarantees."
|
||||||
|
description: "Configure Qdrant's consistency guarantees: the write consistency factor, read consistency levels, and write ordering options, with API examples in every client library."
|
||||||
|
weight: 22
|
||||||
|
---
|
||||||
|
|
||||||
|
# Consistency Guarantees
|
||||||
|
|
||||||
|
By default, Qdrant focuses on availability and maximum throughput of search operations, which is a preferable trade-off for most use cases. During normal operation, you can search and modify data from any peer in the cluster: reads use a partial fan-out strategy to optimize latency and availability, and writes execute in parallel on all active sharded replicas.
|
||||||
|
|
||||||
|
This means concurrent updates on one point can result in an inconsistent state. For example, if two clients simultaneously update the same point in a collection with three replicas per shard. On some replicas, the point may reflect the update from one client, while on other replicas, the point may reflect the update from the other client.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
In some cases, it is necessary to ensure additional guarantees during possible hardware instabilities, mass concurrent updates of same documents, etc.
|
||||||
|
|
||||||
|
Qdrant provides a few options to control consistency guarantees:
|
||||||
|
|
||||||
|
- `write_consistency_factor` - defines the number of replicas that must acknowledge a write operation before responding to the client. Increasing this value will make write operations tolerant to network partitions in the cluster, but will require a higher number of replicas to be active to perform write operations.
|
||||||
|
- Read `consistency` param, can be used with search and retrieve operations to ensure that the results obtained from all replicas are the same. If this option is used, Qdrant will perform the read operation on multiple replicas and resolve the result according to the selected strategy. This option is useful to avoid data inconsistency in case of concurrent updates of the same documents. This options is preferred if the update operations are frequent and the number of replicas is low.
|
||||||
|
- Write `ordering` param, can be used with update and delete operations to ensure that the operations are executed in the same order on all replicas. If this option is used, Qdrant will route the operation to the leader replica of the shard and wait for the response before responding to the client. This option is useful to avoid data inconsistency in case of concurrent updates of the same documents. This options is preferred if read operations are more frequent than update and if search performance is critical.
|
||||||
|
|
||||||
|
|
||||||
|
## Write Consistency Factor
|
||||||
|
|
||||||
|
The `write_consistency_factor` represents the number of replicas that must acknowledge a write operation before responding to the client. It is set to 1 by default.
|
||||||
|
It can be configured at the collection's creation or when updating the
|
||||||
|
collection parameters.
|
||||||
|
|
||||||
|
This value can range from 1 to the number of replicas you have for each shard.
|
||||||
|
|
||||||
|
```http
|
||||||
|
PUT /collections/{collection_name}
|
||||||
|
{
|
||||||
|
"vectors": {
|
||||||
|
"size": 300,
|
||||||
|
"distance": "Cosine"
|
||||||
|
},
|
||||||
|
"shard_number": 6,
|
||||||
|
"replication_factor": 2,
|
||||||
|
"write_consistency_factor": 2
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
|
client = QdrantClient(url="http://localhost:6333")
|
||||||
|
|
||||||
|
client.create_collection(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
vectors_config=models.VectorParams(size=300, distance=models.Distance.COSINE),
|
||||||
|
shard_number=6,
|
||||||
|
replication_factor=2,
|
||||||
|
write_consistency_factor=2,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||||
|
|
||||||
|
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||||
|
|
||||||
|
client.createCollection("{collection_name}", {
|
||||||
|
vectors: {
|
||||||
|
size: 300,
|
||||||
|
distance: "Cosine",
|
||||||
|
},
|
||||||
|
shard_number: 6,
|
||||||
|
replication_factor: 2,
|
||||||
|
write_consistency_factor: 2,
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
```rust
|
||||||
|
use qdrant_client::qdrant::{CreateCollectionBuilder, Distance, VectorParamsBuilder};
|
||||||
|
use qdrant_client::Qdrant;
|
||||||
|
|
||||||
|
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||||
|
|
||||||
|
client
|
||||||
|
.create_collection(
|
||||||
|
CreateCollectionBuilder::new("{collection_name}")
|
||||||
|
.vectors_config(VectorParamsBuilder::new(300, Distance::Cosine))
|
||||||
|
.shard_number(6)
|
||||||
|
.replication_factor(2)
|
||||||
|
.write_consistency_factor(2),
|
||||||
|
)
|
||||||
|
.await?;
|
||||||
|
```
|
||||||
|
|
||||||
|
```java
|
||||||
|
import io.qdrant.client.QdrantClient;
|
||||||
|
import io.qdrant.client.QdrantGrpcClient;
|
||||||
|
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||||
|
import io.qdrant.client.grpc.Collections.Distance;
|
||||||
|
import io.qdrant.client.grpc.Collections.VectorParams;
|
||||||
|
import io.qdrant.client.grpc.Collections.VectorsConfig;
|
||||||
|
|
||||||
|
QdrantClient client =
|
||||||
|
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||||
|
|
||||||
|
client
|
||||||
|
.createCollectionAsync(
|
||||||
|
CreateCollection.newBuilder()
|
||||||
|
.setCollectionName("{collection_name}")
|
||||||
|
.setVectorsConfig(
|
||||||
|
VectorsConfig.newBuilder()
|
||||||
|
.setParams(
|
||||||
|
VectorParams.newBuilder()
|
||||||
|
.setSize(300)
|
||||||
|
.setDistance(Distance.Cosine)
|
||||||
|
.build())
|
||||||
|
.build())
|
||||||
|
.setShardNumber(6)
|
||||||
|
.setReplicationFactor(2)
|
||||||
|
.setWriteConsistencyFactor(2)
|
||||||
|
.build())
|
||||||
|
.get();
|
||||||
|
```
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
using Qdrant.Client;
|
||||||
|
using Qdrant.Client.Grpc;
|
||||||
|
|
||||||
|
var client = new QdrantClient("localhost", 6334);
|
||||||
|
|
||||||
|
await client.CreateCollectionAsync(
|
||||||
|
collectionName: "{collection_name}",
|
||||||
|
vectorsConfig: new VectorParams { Size = 300, Distance = Distance.Cosine },
|
||||||
|
shardNumber: 6,
|
||||||
|
replicationFactor: 2,
|
||||||
|
writeConsistencyFactor: 2
|
||||||
|
);
|
||||||
|
```
|
||||||
|
|
||||||
|
```go
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
|
||||||
|
"github.com/qdrant/go-client/qdrant"
|
||||||
|
)
|
||||||
|
|
||||||
|
client, err := qdrant.NewClient(&qdrant.Config{
|
||||||
|
Host: "localhost",
|
||||||
|
Port: 6334,
|
||||||
|
})
|
||||||
|
|
||||||
|
client.CreateCollection(context.Background(), &qdrant.CreateCollection{
|
||||||
|
CollectionName: "{collection_name}",
|
||||||
|
VectorsConfig: qdrant.NewVectorsConfig(&qdrant.VectorParams{
|
||||||
|
Size: 300,
|
||||||
|
Distance: qdrant.Distance_Cosine,
|
||||||
|
}),
|
||||||
|
ShardNumber: qdrant.PtrOf(uint32(6)),
|
||||||
|
ReplicationFactor: qdrant.PtrOf(uint32(2)),
|
||||||
|
WriteConsistencyFactor: qdrant.PtrOf(uint32(2)),
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
Write operations will fail if the number of active replicas is less than the
|
||||||
|
`write_consistency_factor`. In this case, the client is expected to send the
|
||||||
|
operation again to ensure a consistent state is reached.
|
||||||
|
|
||||||
|
Setting the `write_consistency_factor` to a lower value may allow accepting
|
||||||
|
writes even if there are unresponsive nodes. Unresponsive nodes are marked as
|
||||||
|
dead and will automatically be recovered once available to ensure data
|
||||||
|
consistency.
|
||||||
|
|
||||||
|
The configuration of the `write_consistency_factor` is important for adjusting the cluster's behavior when some nodes go offline due to restarts, upgrades, or failures.
|
||||||
|
|
||||||
|
By default, the cluster continues to accept updates as long as at least one replica of each shard is online. However, this behavior means that once an offline replica is restored, it will require additional synchronization with the rest of the cluster. In some cases, this synchronization can be resource-intensive and undesirable.
|
||||||
|
|
||||||
|
Setting the `write_consistency_factor` to match the replication factor modifies the cluster's behavior so that unreplicated updates are rejected, preventing the need for extra synchronization.
|
||||||
|
|
||||||
|
If the update is applied to enough replicas - according to the `write_consistency_factor` - the update will return a successful status. Any replicas that failed to apply the update will be temporarily disabled and are automatically recovered to keep data consistency. If the update could not be applied to enough replicas, it'll return an error and may be partially applied. The user must submit the operation again to ensure data consistency.
|
||||||
|
|
||||||
|
For asynchronous updates and injection pipelines capable of handling errors and retries, this strategy might be preferable.
|
||||||
|
|
||||||
|
|
||||||
|
## Read Consistency
|
||||||
|
|
||||||
|
Read `consistency` can be specified for most read requests and will ensure that the returned result
|
||||||
|
is consistent across cluster nodes.
|
||||||
|
|
||||||
|
- `all` will query all nodes and return points, which present on all of them
|
||||||
|
- `majority` will query all nodes and return points, which present on the majority of them
|
||||||
|
- `quorum` will query randomly selected majority of nodes and return points, which present on all of them
|
||||||
|
- `1`/`2`/`3`/etc - will query specified number of randomly selected nodes and return points which present on all of them
|
||||||
|
- default `consistency` is `1`
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /collections/{collection_name}/points/query?consistency=majority
|
||||||
|
{
|
||||||
|
"query": [0.2, 0.1, 0.9, 0.7],
|
||||||
|
"filter": {
|
||||||
|
"must": [
|
||||||
|
{
|
||||||
|
"key": "city",
|
||||||
|
"match": {
|
||||||
|
"value": "London"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"params": {
|
||||||
|
"hnsw_ef": 128,
|
||||||
|
"exact": false
|
||||||
|
},
|
||||||
|
"limit": 3
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.query_points(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
query=[0.2, 0.1, 0.9, 0.7],
|
||||||
|
query_filter=models.Filter(
|
||||||
|
must=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="city",
|
||||||
|
match=models.MatchValue(
|
||||||
|
value="London",
|
||||||
|
),
|
||||||
|
)
|
||||||
|
]
|
||||||
|
),
|
||||||
|
search_params=models.SearchParams(hnsw_ef=128, exact=False),
|
||||||
|
limit=3,
|
||||||
|
consistency="majority",
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
client.query("{collection_name}", {
|
||||||
|
query: [0.2, 0.1, 0.9, 0.7],
|
||||||
|
filter: {
|
||||||
|
must: [{ key: "city", match: { value: "London" } }],
|
||||||
|
},
|
||||||
|
params: {
|
||||||
|
hnsw_ef: 128,
|
||||||
|
exact: false,
|
||||||
|
},
|
||||||
|
limit: 3,
|
||||||
|
consistency: "majority",
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
```rust
|
||||||
|
use qdrant_client::qdrant::{
|
||||||
|
read_consistency::Value, Condition, Filter, QueryPointsBuilder, ReadConsistencyType,
|
||||||
|
SearchParamsBuilder,
|
||||||
|
};
|
||||||
|
use qdrant_client::{Qdrant, QdrantError};
|
||||||
|
|
||||||
|
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||||
|
|
||||||
|
client
|
||||||
|
.query(
|
||||||
|
QueryPointsBuilder::new("{collection_name}")
|
||||||
|
.query(vec![0.2, 0.1, 0.9, 0.7])
|
||||||
|
.limit(3)
|
||||||
|
.filter(Filter::must([Condition::matches(
|
||||||
|
"city",
|
||||||
|
"London".to_string(),
|
||||||
|
)]))
|
||||||
|
.params(SearchParamsBuilder::default().hnsw_ef(128).exact(false))
|
||||||
|
.read_consistency(Value::Type(ReadConsistencyType::Majority.into())),
|
||||||
|
)
|
||||||
|
.await?;
|
||||||
|
```
|
||||||
|
|
||||||
|
```java
|
||||||
|
import io.qdrant.client.QdrantClient;
|
||||||
|
import io.qdrant.client.QdrantGrpcClient;
|
||||||
|
import io.qdrant.client.grpc.Common.Filter;
|
||||||
|
import io.qdrant.client.grpc.Points.QueryPoints;
|
||||||
|
import io.qdrant.client.grpc.Points.ReadConsistency;
|
||||||
|
import io.qdrant.client.grpc.Points.ReadConsistencyType;
|
||||||
|
import io.qdrant.client.grpc.Points.SearchParams;
|
||||||
|
|
||||||
|
import static io.qdrant.client.QueryFactory.nearest;
|
||||||
|
import static io.qdrant.client.ConditionFactory.matchKeyword;
|
||||||
|
|
||||||
|
QdrantClient client =
|
||||||
|
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||||
|
|
||||||
|
client.queryAsync(
|
||||||
|
QueryPoints.newBuilder()
|
||||||
|
.setCollectionName("{collection_name}")
|
||||||
|
.setFilter(Filter.newBuilder().addMust(matchKeyword("city", "London")).build())
|
||||||
|
.setQuery(nearest(.2f, 0.1f, 0.9f, 0.7f))
|
||||||
|
.setParams(SearchParams.newBuilder().setHnswEf(128).setExact(false).build())
|
||||||
|
.setLimit(3)
|
||||||
|
.setReadConsistency(
|
||||||
|
ReadConsistency.newBuilder().setType(ReadConsistencyType.Majority).build())
|
||||||
|
.build())
|
||||||
|
.get();
|
||||||
|
```
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
using Qdrant.Client;
|
||||||
|
using Qdrant.Client.Grpc;
|
||||||
|
using static Qdrant.Client.Grpc.Conditions;
|
||||||
|
|
||||||
|
var client = new QdrantClient("localhost", 6334);
|
||||||
|
|
||||||
|
await client.QueryAsync(
|
||||||
|
collectionName: "{collection_name}",
|
||||||
|
query: new float[] { 0.2f, 0.1f, 0.9f, 0.7f },
|
||||||
|
filter: MatchKeyword("city", "London"),
|
||||||
|
searchParams: new SearchParams { HnswEf = 128, Exact = false },
|
||||||
|
limit: 3,
|
||||||
|
readConsistency: new ReadConsistency { Type = ReadConsistencyType.Majority }
|
||||||
|
);
|
||||||
|
```
|
||||||
|
|
||||||
|
```go
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
|
||||||
|
"github.com/qdrant/go-client/qdrant"
|
||||||
|
)
|
||||||
|
|
||||||
|
client, err := qdrant.NewClient(&qdrant.Config{
|
||||||
|
Host: "localhost",
|
||||||
|
Port: 6334,
|
||||||
|
})
|
||||||
|
|
||||||
|
client.Query(context.Background(), &qdrant.QueryPoints{
|
||||||
|
CollectionName: "{collection_name}",
|
||||||
|
Query: qdrant.NewQuery(0.2, 0.1, 0.9, 0.7),
|
||||||
|
Filter: &qdrant.Filter{
|
||||||
|
Must: []*qdrant.Condition{
|
||||||
|
qdrant.NewMatch("city", "London"),
|
||||||
|
},
|
||||||
|
},
|
||||||
|
Params: &qdrant.SearchParams{
|
||||||
|
HnswEf: qdrant.PtrOf(uint64(128)),
|
||||||
|
},
|
||||||
|
Limit: qdrant.PtrOf(uint64(3)),
|
||||||
|
ReadConsistency: qdrant.NewReadConsistencyType(qdrant.ReadConsistencyType_Majority),
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
## Write Ordering
|
||||||
|
|
||||||
|
Write `ordering` can be specified for any write request to serialize it through a single "leader" node,
|
||||||
|
which ensures that all write operations (issued with the same `ordering`) are performed and observed
|
||||||
|
sequentially.
|
||||||
|
|
||||||
|
- `weak` _(default)_ ordering does not provide any additional guarantees, so write operations can be freely reordered.
|
||||||
|
- `medium` ordering serializes all write operations through a dynamically elected leader, which might cause minor inconsistencies in case of leader change.
|
||||||
|
- `strong` ordering serializes all write operations through the permanent leader, which provides strong consistency, but write operations may be unavailable if the leader is down.
|
||||||
|
|
||||||
|
<aside role="status">Some <a href="/documentation/scaling/distributed_deployment/#shard-transfer-method">shard transfer methods</a> may affect ordering guarantees.</aside>
|
||||||
|
|
||||||
|
```http
|
||||||
|
PUT /collections/{collection_name}/points?ordering=strong
|
||||||
|
{
|
||||||
|
"batch": {
|
||||||
|
"ids": [1, 2, 3],
|
||||||
|
"payloads": [
|
||||||
|
{"color": "red"},
|
||||||
|
{"color": "green"},
|
||||||
|
{"color": "blue"}
|
||||||
|
],
|
||||||
|
"vectors": [
|
||||||
|
[0.9, 0.1, 0.1],
|
||||||
|
[0.1, 0.9, 0.1],
|
||||||
|
[0.1, 0.1, 0.9]
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.upsert(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
points=models.Batch(
|
||||||
|
ids=[1, 2, 3],
|
||||||
|
payloads=[
|
||||||
|
{"color": "red"},
|
||||||
|
{"color": "green"},
|
||||||
|
{"color": "blue"},
|
||||||
|
],
|
||||||
|
vectors=[
|
||||||
|
[0.9, 0.1, 0.1],
|
||||||
|
[0.1, 0.9, 0.1],
|
||||||
|
[0.1, 0.1, 0.9],
|
||||||
|
],
|
||||||
|
),
|
||||||
|
ordering=models.WriteOrdering.STRONG,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
client.upsert("{collection_name}", {
|
||||||
|
batch: {
|
||||||
|
ids: [1, 2, 3],
|
||||||
|
payloads: [{ color: "red" }, { color: "green" }, { color: "blue" }],
|
||||||
|
vectors: [
|
||||||
|
[0.9, 0.1, 0.1],
|
||||||
|
[0.1, 0.9, 0.1],
|
||||||
|
[0.1, 0.1, 0.9],
|
||||||
|
],
|
||||||
|
},
|
||||||
|
ordering: "strong",
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
```rust
|
||||||
|
use qdrant_client::qdrant::{
|
||||||
|
PointStruct, UpsertPointsBuilder, WriteOrdering, WriteOrderingType
|
||||||
|
};
|
||||||
|
use qdrant_client::Qdrant;
|
||||||
|
|
||||||
|
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||||
|
|
||||||
|
client
|
||||||
|
.upsert_points(
|
||||||
|
UpsertPointsBuilder::new(
|
||||||
|
"{collection_name}",
|
||||||
|
vec![
|
||||||
|
PointStruct::new(1, vec![0.9, 0.1, 0.1], [("color", "red".into())]),
|
||||||
|
PointStruct::new(2, vec![0.1, 0.9, 0.1], [("color", "green".into())]),
|
||||||
|
PointStruct::new(3, vec![0.1, 0.1, 0.9], [("color", "blue".into())]),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
.ordering(WriteOrdering {
|
||||||
|
r#type: WriteOrderingType::Strong.into(),
|
||||||
|
}),
|
||||||
|
)
|
||||||
|
.await?;
|
||||||
|
```
|
||||||
|
|
||||||
|
```java
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
|
||||||
|
import static io.qdrant.client.PointIdFactory.id;
|
||||||
|
import static io.qdrant.client.ValueFactory.value;
|
||||||
|
import static io.qdrant.client.VectorsFactory.vectors;
|
||||||
|
|
||||||
|
import io.qdrant.client.grpc.Points.PointStruct;
|
||||||
|
import io.qdrant.client.grpc.Points.UpsertPoints;
|
||||||
|
import io.qdrant.client.grpc.Points.WriteOrdering;
|
||||||
|
import io.qdrant.client.grpc.Points.WriteOrderingType;
|
||||||
|
|
||||||
|
client
|
||||||
|
.upsertAsync(
|
||||||
|
UpsertPoints.newBuilder()
|
||||||
|
.setCollectionName("{collection_name}")
|
||||||
|
.addAllPoints(
|
||||||
|
List.of(
|
||||||
|
PointStruct.newBuilder()
|
||||||
|
.setId(id(1))
|
||||||
|
.setVectors(vectors(0.9f, 0.1f, 0.1f))
|
||||||
|
.putAllPayload(Map.of("color", value("red")))
|
||||||
|
.build(),
|
||||||
|
PointStruct.newBuilder()
|
||||||
|
.setId(id(2))
|
||||||
|
.setVectors(vectors(0.1f, 0.9f, 0.1f))
|
||||||
|
.putAllPayload(Map.of("color", value("green")))
|
||||||
|
.build(),
|
||||||
|
PointStruct.newBuilder()
|
||||||
|
.setId(id(3))
|
||||||
|
.setVectors(vectors(0.1f, 0.1f, 0.94f))
|
||||||
|
.putAllPayload(Map.of("color", value("blue")))
|
||||||
|
.build()))
|
||||||
|
.setOrdering(WriteOrdering.newBuilder().setType(WriteOrderingType.Strong).build())
|
||||||
|
.build())
|
||||||
|
.get();
|
||||||
|
```
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
using Qdrant.Client;
|
||||||
|
using Qdrant.Client.Grpc;
|
||||||
|
|
||||||
|
var client = new QdrantClient("localhost", 6334);
|
||||||
|
|
||||||
|
await client.UpsertAsync(
|
||||||
|
collectionName: "{collection_name}",
|
||||||
|
points: new List<PointStruct>
|
||||||
|
{
|
||||||
|
new()
|
||||||
|
{
|
||||||
|
Id = 1,
|
||||||
|
Vectors = new[] { 0.9f, 0.1f, 0.1f },
|
||||||
|
Payload = { ["color"] = "red" }
|
||||||
|
},
|
||||||
|
new()
|
||||||
|
{
|
||||||
|
Id = 2,
|
||||||
|
Vectors = new[] { 0.1f, 0.9f, 0.1f },
|
||||||
|
Payload = { ["color"] = "green" }
|
||||||
|
},
|
||||||
|
new()
|
||||||
|
{
|
||||||
|
Id = 3,
|
||||||
|
Vectors = new[] { 0.1f, 0.1f, 0.9f },
|
||||||
|
Payload = { ["color"] = "blue" }
|
||||||
|
}
|
||||||
|
},
|
||||||
|
ordering: WriteOrderingType.Strong
|
||||||
|
);
|
||||||
|
```
|
||||||
|
|
||||||
|
```go
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
|
||||||
|
"github.com/qdrant/go-client/qdrant"
|
||||||
|
)
|
||||||
|
|
||||||
|
client, err := qdrant.NewClient(&qdrant.Config{
|
||||||
|
Host: "localhost",
|
||||||
|
Port: 6334,
|
||||||
|
})
|
||||||
|
|
||||||
|
client.Upsert(context.Background(), &qdrant.UpsertPoints{
|
||||||
|
CollectionName: "{collection_name}",
|
||||||
|
Points: []*qdrant.PointStruct{
|
||||||
|
{
|
||||||
|
Id: qdrant.NewIDNum(1),
|
||||||
|
Vectors: qdrant.NewVectors(0.9, 0.1, 0.1),
|
||||||
|
Payload: qdrant.NewValueMap(map[string]any{"color": "red"}),
|
||||||
|
},
|
||||||
|
{
|
||||||
|
Id: qdrant.NewIDNum(2),
|
||||||
|
Vectors: qdrant.NewVectors(0.1, 0.9, 0.1),
|
||||||
|
Payload: qdrant.NewValueMap(map[string]any{"color": "green"}),
|
||||||
|
},
|
||||||
|
{
|
||||||
|
Id: qdrant.NewIDNum(3),
|
||||||
|
Vectors: qdrant.NewVectors(0.1, 0.1, 0.9),
|
||||||
|
Payload: qdrant.NewValueMap(map[string]any{"color": "blue"}),
|
||||||
|
},
|
||||||
|
},
|
||||||
|
Ordering: &qdrant.WriteOrdering{
|
||||||
|
Type: qdrant.WriteOrderingType_Strong,
|
||||||
|
},
|
||||||
|
})
|
||||||
|
```
|
||||||
@@ -0,0 +1,689 @@
|
|||||||
|
---
|
||||||
|
title: Distributed Deployment
|
||||||
|
short_description: "Run Qdrant in distributed mode across multiple nodes for higher availability, scalable throughput, and fault-tolerant vector search."
|
||||||
|
description: "Configure distributed Qdrant deployments to scale storage, balance load, and tolerate node failures using sharding and replication across a cluster."
|
||||||
|
weight: 20
|
||||||
|
aliases:
|
||||||
|
- /documentation/distributed_deployment
|
||||||
|
- /guides/distributed_deployment
|
||||||
|
- /documentation/operations/distributed_deployment
|
||||||
|
---
|
||||||
|
|
||||||
|
# Distributed Deployment
|
||||||
|
|
||||||
|
*Available since Qdrant v0.8.0*
|
||||||
|
|
||||||
|
Qdrant supports a distributed deployment mode.
|
||||||
|
In this mode, multiple Qdrant services communicate with each other to distribute the data across the peers to extend the storage capabilities and increase stability.
|
||||||
|
|
||||||
|
## Enabling Distributed Mode in Self-Hosted Qdrant
|
||||||
|
|
||||||
|
To enable distributed deployment - enable the cluster mode in the [configuration](/documentation/ops-configuration/configuration/) or using the ENV variable: `QDRANT__CLUSTER__ENABLED=true`.
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
cluster:
|
||||||
|
# Use `enabled: true` to run Qdrant in distributed deployment mode
|
||||||
|
enabled: true
|
||||||
|
# Configuration of the inter-cluster communication
|
||||||
|
p2p:
|
||||||
|
# Port for internal communication between peers
|
||||||
|
port: 6335
|
||||||
|
|
||||||
|
# Configuration related to distributed consensus algorithm
|
||||||
|
consensus:
|
||||||
|
# How frequently peers should ping each other.
|
||||||
|
# Setting this parameter to lower value will allow consensus
|
||||||
|
# to detect disconnected node earlier, but too frequent
|
||||||
|
# tick period may create significant network and CPU overhead.
|
||||||
|
# We encourage you NOT to change this parameter unless you know what you are doing.
|
||||||
|
tick_period_ms: 100
|
||||||
|
```
|
||||||
|
|
||||||
|
By default, Qdrant will use port `6335` for its internal communication.
|
||||||
|
All peers should be accessible on this port from within the cluster, but make sure to isolate this port from outside access, as it might be used to perform write operations.
|
||||||
|
|
||||||
|
Additionally, you must provide the `--uri` flag to the first peer so it can tell other nodes how it should be reached:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
./qdrant --uri 'http://qdrant_node_1:6335'
|
||||||
|
```
|
||||||
|
|
||||||
|
Subsequent peers in a cluster must know at least one node of the existing cluster to synchronize through it with the rest of the cluster.
|
||||||
|
|
||||||
|
To do this, they need to be provided with a bootstrap URL:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
./qdrant --bootstrap 'http://qdrant_node_1:6335'
|
||||||
|
```
|
||||||
|
|
||||||
|
The URL of the new peers themselves will be calculated automatically from the IP address of their request.
|
||||||
|
But it is also possible to provide them individually using the `--uri` argument.
|
||||||
|
|
||||||
|
```text
|
||||||
|
USAGE:
|
||||||
|
qdrant [OPTIONS]
|
||||||
|
|
||||||
|
OPTIONS:
|
||||||
|
--bootstrap <URI>
|
||||||
|
Uri of the peer to bootstrap from in case of multi-peer deployment. If not specified -
|
||||||
|
this peer will be considered as a first in a new deployment
|
||||||
|
|
||||||
|
--uri <URI>
|
||||||
|
Uri of this peer. Other peers should be able to reach it by this uri.
|
||||||
|
|
||||||
|
This value has to be supplied if this is the first peer in a new deployment.
|
||||||
|
|
||||||
|
In case this is not the first peer and it bootstraps the value is optional. If not
|
||||||
|
supplied then qdrant will take internal grpc port from config and derive the IP address
|
||||||
|
of this peer on bootstrap peer (receiving side)
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
After a successful synchronization you can observe the state of the cluster through the [REST API](https://api.qdrant.tech/master/api-reference/distributed/cluster-status):
|
||||||
|
|
||||||
|
```http
|
||||||
|
GET /cluster
|
||||||
|
```
|
||||||
|
|
||||||
|
Example result:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"result": {
|
||||||
|
"status": "enabled",
|
||||||
|
"peer_id": 11532566549086892000,
|
||||||
|
"peers": {
|
||||||
|
"9834046559507417430": {
|
||||||
|
"uri": "http://172.18.0.3:6335/"
|
||||||
|
},
|
||||||
|
"11532566549086892528": {
|
||||||
|
"uri": "http://qdrant_node_1:6335/"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"raft_info": {
|
||||||
|
"term": 1,
|
||||||
|
"commit": 4,
|
||||||
|
"pending_operations": 1,
|
||||||
|
"leader": 11532566549086892000,
|
||||||
|
"role": "Leader"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"status": "ok",
|
||||||
|
"time": 5.731e-06
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Note that enabling distributed mode does not automatically replicate your data. See the section on [making use of a new distributed Qdrant cluster](#making-use-of-a-new-distributed-qdrant-cluster) for the next steps.
|
||||||
|
|
||||||
|
## Enabling Distributed Mode in Qdrant Cloud
|
||||||
|
|
||||||
|
For best results, first ensure your cluster is running Qdrant v1.7.4 or higher. Older versions of Qdrant do support distributed mode, but improvements in v1.7.4 make distributed clusters more resilient during outages.
|
||||||
|
|
||||||
|
To enable distributed mode, in the [Qdrant Cloud console](https://cloud.qdrant.io/), click "Scale Cluster" and select the desired node count.
|
||||||
|
|
||||||
|
Additionally, Qdrant Cloud also offers the ability to automatically rebalance and to reshard your collections, which is not available in self-hosted Qdrant. See the [Resharding](/documentation/cloud/cluster-scaling/#resharding) and [Shard Rebalancing](/documentation/cloud/configure-cluster/#shard-rebalancing) sections in for more details.
|
||||||
|
|
||||||
|
## Making Use of a New Distributed Qdrant Cluster
|
||||||
|
|
||||||
|
When you enable distributed mode and scale up to two or more nodes, your data does not move to the new node automatically; it starts out empty. To make use of your new empty node, do one of the following:
|
||||||
|
|
||||||
|
* Create a new replicated collection by setting the [replication_factor](#replication-factor) to 2 or more and setting the [number of shards](#choosing-the-right-number-of-shards) to a multiple of your number of nodes.
|
||||||
|
* If you have an existing collection which does not contain enough shards for each node, you must create a new collection as described in the previous bullet point.
|
||||||
|
* If you already have enough shards for each node, and you merely need to replicate your data, follow the directions for [creating new shard replicas](#creating-new-shard-replicas).
|
||||||
|
* If you already have enough shards for each node, and your data is already replicated, you can move data (without replicating it) onto the new node(s) by [moving shards](#moving-shards).
|
||||||
|
|
||||||
|
Qdrant uses the [Raft](https://raft.github.io/) consensus protocol to keep the cluster topology and collection structure consistent across nodes. For how consensus works and what it means for availability, see [Raft consensus](/documentation/scaling/horizontal-scaling/#raft-consensus) in Horizontal Scaling.
|
||||||
|
|
||||||
|
## Sharding
|
||||||
|
|
||||||
|
Qdrant distributes a collection's points across shards to scale horizontally. For how sharding works conceptually, see [Sharding](/documentation/scaling/horizontal-scaling/#sharding) in Horizontal Scaling.
|
||||||
|
|
||||||
|
### Choosing the Right Number of Shards
|
||||||
|
|
||||||
|
When you create a collection, Qdrant splits the collection into `shard_number` shards. If left unset, `shard_number` is set to the number of nodes in your cluster when the collection was created. The `shard_number` cannot be changed without recreating the collection.
|
||||||
|
|
||||||
|
```http
|
||||||
|
PUT /collections/{collection_name}
|
||||||
|
{
|
||||||
|
"vectors": {
|
||||||
|
"size": 300,
|
||||||
|
"distance": "Cosine"
|
||||||
|
},
|
||||||
|
"shard_number": 6
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
|
client = QdrantClient(url="http://localhost:6333")
|
||||||
|
|
||||||
|
client.create_collection(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
vectors_config=models.VectorParams(size=300, distance=models.Distance.COSINE),
|
||||||
|
shard_number=6,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||||
|
|
||||||
|
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||||
|
|
||||||
|
client.createCollection("{collection_name}", {
|
||||||
|
vectors: {
|
||||||
|
size: 300,
|
||||||
|
distance: "Cosine",
|
||||||
|
},
|
||||||
|
shard_number: 6,
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
```rust
|
||||||
|
use qdrant_client::qdrant::{CreateCollectionBuilder, Distance, VectorParamsBuilder};
|
||||||
|
use qdrant_client::Qdrant;
|
||||||
|
|
||||||
|
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||||
|
|
||||||
|
client
|
||||||
|
.create_collection(
|
||||||
|
CreateCollectionBuilder::new("{collection_name}")
|
||||||
|
.vectors_config(VectorParamsBuilder::new(300, Distance::Cosine))
|
||||||
|
.shard_number(6),
|
||||||
|
)
|
||||||
|
.await?;
|
||||||
|
```
|
||||||
|
|
||||||
|
```java
|
||||||
|
import io.qdrant.client.QdrantClient;
|
||||||
|
import io.qdrant.client.QdrantGrpcClient;
|
||||||
|
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||||
|
import io.qdrant.client.grpc.Collections.Distance;
|
||||||
|
import io.qdrant.client.grpc.Collections.VectorParams;
|
||||||
|
import io.qdrant.client.grpc.Collections.VectorsConfig;
|
||||||
|
|
||||||
|
QdrantClient client =
|
||||||
|
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||||
|
|
||||||
|
client
|
||||||
|
.createCollectionAsync(
|
||||||
|
CreateCollection.newBuilder()
|
||||||
|
.setCollectionName("{collection_name}")
|
||||||
|
.setVectorsConfig(
|
||||||
|
VectorsConfig.newBuilder()
|
||||||
|
.setParams(
|
||||||
|
VectorParams.newBuilder()
|
||||||
|
.setSize(300)
|
||||||
|
.setDistance(Distance.Cosine)
|
||||||
|
.build())
|
||||||
|
.build())
|
||||||
|
.setShardNumber(6)
|
||||||
|
.build())
|
||||||
|
.get();
|
||||||
|
```
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
using Qdrant.Client;
|
||||||
|
using Qdrant.Client.Grpc;
|
||||||
|
|
||||||
|
var client = new QdrantClient("localhost", 6334);
|
||||||
|
|
||||||
|
await client.CreateCollectionAsync(
|
||||||
|
collectionName: "{collection_name}",
|
||||||
|
vectorsConfig: new VectorParams { Size = 300, Distance = Distance.Cosine },
|
||||||
|
shardNumber: 6
|
||||||
|
);
|
||||||
|
```
|
||||||
|
|
||||||
|
```go
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
|
||||||
|
"github.com/qdrant/go-client/qdrant"
|
||||||
|
)
|
||||||
|
|
||||||
|
client, err := qdrant.NewClient(&qdrant.Config{
|
||||||
|
Host: "localhost",
|
||||||
|
Port: 6334,
|
||||||
|
})
|
||||||
|
|
||||||
|
client.CreateCollection(context.Background(), &qdrant.CreateCollection{
|
||||||
|
CollectionName: "{collection_name}",
|
||||||
|
VectorsConfig: qdrant.NewVectorsConfig(&qdrant.VectorParams{
|
||||||
|
Size: 300,
|
||||||
|
Distance: qdrant.Distance_Cosine,
|
||||||
|
}),
|
||||||
|
ShardNumber: qdrant.PtrOf(uint32(6)),
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
To ensure all nodes in your cluster are evenly utilized, the number of shards must be a multiple of the number of nodes you are currently running in your cluster.
|
||||||
|
|
||||||
|
> Aside: Advanced use cases such as multitenancy may require an uneven distribution of shards. See [Multitenancy](/articles/multitenancy/).
|
||||||
|
|
||||||
|
We recommend creating at least 2 shards per node to allow future expansion without having to re-shard. [Resharding](/documentation/cloud/cluster-scaling/#resharding) is possible on Qdrant Cloud, but should be avoided if hosting elsewhere as it would require creating a new collection.
|
||||||
|
|
||||||
|
If you anticipate a lot of growth, we recommend 12 shards since you can expand from 1 node up to 2, 3, 6, and 12 nodes without having to re-shard. Having more than 12 shards in a small cluster may not be worth the performance overhead.
|
||||||
|
|
||||||
|
### Rebalancing
|
||||||
|
|
||||||
|
Shards are evenly distributed across all existing nodes when a collection is first created.
|
||||||
|
|
||||||
|
When you add or remove nodes from the cluster, rebalancing of existing shards across the nodes depends on how you've deployed the cluster:
|
||||||
|
|
||||||
|
- In Qdrant Cloud, shards are [balanced across the nodes automatically](/documentation/cloud/configure-cluster/#shard-rebalancing).
|
||||||
|
- If your cluster is not running in Qdrant Cloud, you need to [manually balance shards](#moving-shards).
|
||||||
|
|
||||||
|
### Resharding
|
||||||
|
|
||||||
|
*Available as of v1.13.0 in Cloud*
|
||||||
|
|
||||||
|
<aside role="alert">Resharding a large collection can take a long time. It's better to set the desired shard count when <a href="#choosing-the-right-number-of-shards">creating a collection</a>.</aside>
|
||||||
|
|
||||||
|
Resharding allows you to change the number of shards in your existing collections if you're hosting with our [Cloud](/documentation/deploy-intro/) offering.
|
||||||
|
|
||||||
|
Resharding can change the number of shards both up and down, without having to recreate the collection from scratch.
|
||||||
|
|
||||||
|
Please refer to the [Resharding](/documentation/cloud/cluster-scaling/#resharding) section in our cloud documentation for more details.
|
||||||
|
|
||||||
|
### Moving Shards
|
||||||
|
|
||||||
|
*Available as of v0.9.0*
|
||||||
|
|
||||||
|
Qdrant allows moving shards between nodes in the cluster and removing nodes from the cluster. This functionality unlocks the ability to dynamically scale the cluster size without downtime. It also allows you to upgrade or migrate nodes without downtime.
|
||||||
|
|
||||||
|
If your cluster is running in Qdrant Cloud, shards are balanced across the cluster nodes automatically. For more information see the [Configuring Cloud Clusters](/documentation/cloud/configure-cluster/#shard-rebalancing) and [Cloud Cluster Scaling](/documentation/cloud/cluster-scaling/) documentation.
|
||||||
|
|
||||||
|
Qdrant provides the information regarding the current shard distribution in the cluster with the [Collection Cluster info API](https://api.qdrant.tech/master/api-reference/distributed/collection-cluster-info).
|
||||||
|
|
||||||
|
Use the [Update collection cluster setup API](https://api.qdrant.tech/master/api-reference/distributed/update-collection-cluster) to initiate the shard transfer:
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /collections/{collection_name}/cluster
|
||||||
|
{
|
||||||
|
"move_shard": {
|
||||||
|
"shard_id": 0,
|
||||||
|
"from_peer_id": 381894127,
|
||||||
|
"to_peer_id": 467122995
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
<aside role="status">You likely want to select a specific <a href="#shard-transfer-method">shard transfer method</a> to get desired performance and guarantees.</aside>
|
||||||
|
|
||||||
|
After the transfer is initiated, the service will process it based on the used
|
||||||
|
[transfer method](#shard-transfer-method) keeping both shards in sync. Once the
|
||||||
|
transfer is completed, the old shard is deleted from the source node.
|
||||||
|
|
||||||
|
In case you want to downscale the cluster, you can move all shards away from a peer and then remove the peer using the [remove peer API](https://api.qdrant.tech/master/api-reference/distributed/remove-peer).
|
||||||
|
|
||||||
|
```http
|
||||||
|
DELETE /cluster/peer/{peer_id}
|
||||||
|
```
|
||||||
|
|
||||||
|
After that, Qdrant will exclude the node from the consensus, and the instance will be ready for shutdown.
|
||||||
|
|
||||||
|
### User-Defined Sharding
|
||||||
|
|
||||||
|
*Available as of v1.7.0*
|
||||||
|
|
||||||
|
Qdrant allows you to specify the shard for each point individually. This feature is useful if you want to control the shard placement of your data, so that operations can hit only the subset of shards they actually need. In big clusters, this can significantly improve the performance of operations that do not require the whole collection to be scanned.
|
||||||
|
|
||||||
|
#### Multitenancy
|
||||||
|
|
||||||
|
A use-case for this feature is managing a [multi-tenant collection](/documentation/manage-data/multitenancy/), where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards.
|
||||||
|
|
||||||
|
{{< code-snippet path="/documentation/headless/snippets/create-collection/with-custom-sharding/" >}}
|
||||||
|
|
||||||
|
In this mode, the `shard_number` means the number of shards per shard key, where points will be distributed evenly. For example, if you have 10 shard keys and a collection config with these settings:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"shard_number": 1,
|
||||||
|
"sharding_method": "custom",
|
||||||
|
"replication_factor": 2
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Then you will have `1 * 10 * 2 = 20` total physical shards in the collection.
|
||||||
|
|
||||||
|
Physical shards require a large amount of resources, so make sure your custom sharding key has a low cardinality.
|
||||||
|
|
||||||
|
For large cardinality keys, it is recommended to use [partition by payload](/documentation/manage-data/multitenancy/#partition-by-payload) instead.
|
||||||
|
|
||||||
|
Now you need to create custom shards ([API reference](https://api.qdrant.tech/api-reference/distributed/create-shard-key#request)):
|
||||||
|
|
||||||
|
{{< code-snippet path="/documentation/headless/snippets/create-shard/create-named-shard/" >}}
|
||||||
|
|
||||||
|
You can list all custom shard keys in a collection:
|
||||||
|
|
||||||
|
{{< code-snippet path="/documentation/headless/snippets/list-shard-keys/" >}}
|
||||||
|
|
||||||
|
To specify the shard for each point, you need to provide the `shard_key` field in the upsert request:
|
||||||
|
|
||||||
|
{{< code-snippet path="/documentation/headless/snippets/insert-points/with-custom-shard/" >}}
|
||||||
|
|
||||||
|
<aside role="alert">
|
||||||
|
Using the same point ID across multiple shard keys is <strong>not supported<sup>*</sup></strong> and should be avoided.
|
||||||
|
</aside>
|
||||||
|
<sup>
|
||||||
|
<strong>*</strong> When using custom sharding, IDs are only enforced to be unique within a shard key. This means that you can have multiple points with the same ID, if they have different shard keys.
|
||||||
|
This is a limitation of the current implementation, and is an anti-pattern that should be avoided because it can create scenarios of points with the same ID to have different contents. In the future, we plan to add a global ID uniqueness check.
|
||||||
|
</sup>
|
||||||
|
|
||||||
|
Now you can target the operations to specific shard(s) by specifying the `shard_key` on any operation you do. Operations that do not specify the shard key will be executed on __all__ shards.
|
||||||
|
|
||||||
|
#### Time-Based Sharding
|
||||||
|
|
||||||
|
Another use case for user-defined sharding is time-based sharding, where you route points to a specific shard (or shards) based on timestamp. This enables efficient querying of recent data and efficient data lifecycle management by deleting old shards once they pass a certain age. See the [Time-Based Sharding](/documentation/tutorials-operations/time-based-sharding/) tutorial for more details.
|
||||||
|
|
||||||
|
<img src="/documentation/tutorials/time-based-sharding/time-based-sharding.png" alt="Sharding per day">
|
||||||
|
|
||||||
|
### Shard Transfer Method
|
||||||
|
|
||||||
|
*Available as of v1.7.0*
|
||||||
|
|
||||||
|
There are different methods for transferring a shard, such as moving or
|
||||||
|
replicating, to another node. Depending on what performance and guarantees you'd
|
||||||
|
like to have and how you'd like to manage your cluster, you likely want to
|
||||||
|
choose a specific method. Each method has its own pros and cons. Which is
|
||||||
|
fastest depends on the size and state of a shard.
|
||||||
|
|
||||||
|
Available shard transfer methods are:
|
||||||
|
|
||||||
|
- `stream_records`: _(default)_ transfer by streaming just its records to the target node in batches.
|
||||||
|
- `snapshot`: transfer including its index and quantized data by utilizing a [snapshot](/documentation/snapshots/) automatically.
|
||||||
|
- `wal_delta`: _(auto recovery default)_ transfer by resolving [WAL] difference; the operations that were missed.
|
||||||
|
|
||||||
|
Each has pros, cons and specific requirements, some of which are:
|
||||||
|
|
||||||
|
| Method: | Stream records | Snapshot | WAL delta |
|
||||||
|
|:---|:---|:---|:---|
|
||||||
|
| **Version** | v0.8.0+ | v1.7.0+ | v1.8.0+ |
|
||||||
|
| **Target** | New/existing shard | New/existing shard | Existing shard |
|
||||||
|
| **Connectivity** | Internal gRPC API <small>(<abbr title="port">6335</abbr>)</small> | REST API <small>(<abbr title="port">6333</abbr>)</small><br>Internal gRPC API <small>(<abbr title="port">6335</abbr>)</small> | Internal gRPC API <small>(<abbr title="port">6335</abbr>)</small> |
|
||||||
|
| **HNSW index** | Doesn't transfer, will reindex on target. | Does transfer, immediately ready on target. | Doesn't transfer, may index on target. |
|
||||||
|
| **Quantization** | Doesn't transfer, will requantize on target. | Does transfer, immediately ready on target. | Doesn't transfer, may quantize on target. |
|
||||||
|
| **Ordering** | Unordered updates on target[^unordered] | Ordered updates on target[^ordered] | Ordered updates on target[^ordered] |
|
||||||
|
| **Disk space** | No extra required | Extra required for snapshot on both nodes | No extra required |
|
||||||
|
|
||||||
|
[^unordered]: Weak ordering for updates: All records are streamed to the target node in order.
|
||||||
|
New updates are received on the target node in parallel, while the transfer
|
||||||
|
of records is still happening. We therefore have `weak` ordering, regardless
|
||||||
|
of what [ordering](/documentation/scaling/consistency-guarantees/#write-ordering) is used for updates.
|
||||||
|
[^ordered]: Strong ordering for updates: A snapshot of the shard
|
||||||
|
is created, it is transferred and recovered on the target node. That ensures
|
||||||
|
the state of the shard is kept consistent. New updates are queued on the
|
||||||
|
source node, and transferred in order to the target node. Updates therefore
|
||||||
|
have the same [ordering](/documentation/scaling/consistency-guarantees/#write-ordering) as the user selects, making
|
||||||
|
`strong` ordering possible.
|
||||||
|
|
||||||
|
To select a shard transfer method, specify the `method` like:
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /collections/{collection_name}/cluster
|
||||||
|
{
|
||||||
|
"move_shard": {
|
||||||
|
"shard_id": 0,
|
||||||
|
"from_peer_id": 381894127,
|
||||||
|
"to_peer_id": 467122995,
|
||||||
|
"method": "snapshot"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The `stream_records` transfer method is the simplest available. It simply
|
||||||
|
transfers all shard records in batches to the target node until it has
|
||||||
|
transferred all of them, keeping both shards in sync. It will also make sure the
|
||||||
|
transferred shard indexing process is keeping up before performing a final
|
||||||
|
switch. The method has two common disadvantages: 1. It does not transfer index
|
||||||
|
or quantization data, meaning that the shard has to be optimized again on the
|
||||||
|
new node, which can be very expensive. 2. The ordering guarantees are
|
||||||
|
`weak`[^unordered], which is not suitable for some applications. Because it is
|
||||||
|
so simple, it's also very robust, making it a reliable choice if the above cons
|
||||||
|
are acceptable in your use case. If your cluster is unstable and out of
|
||||||
|
resources, it's probably best to use the `stream_records` transfer method,
|
||||||
|
because it is unlikely to fail.
|
||||||
|
|
||||||
|
The `snapshot` transfer method utilizes [snapshots](/documentation/snapshots/)
|
||||||
|
to transfer a shard. A snapshot is created automatically. It is then transferred
|
||||||
|
and restored on the target node. After this is done, the snapshot is removed
|
||||||
|
from both nodes. While the snapshot/transfer/restore operation is happening, the
|
||||||
|
source node queues up all new operations. All queued updates are then sent in
|
||||||
|
order to the target shard to bring it into the same state as the source. There
|
||||||
|
are two important benefits: 1. It transfers index and quantization data, so that
|
||||||
|
the shard does not have to be optimized again on the target node, making them
|
||||||
|
immediately available. This way, Qdrant ensures that there will be no
|
||||||
|
degradation in performance at the end of the transfer. Especially on large
|
||||||
|
shards, this can give a huge performance improvement. 2. The ordering guarantees
|
||||||
|
can be `strong`[^ordered], required for some applications.
|
||||||
|
|
||||||
|
The `wal_delta` transfer method only transfers the difference between two
|
||||||
|
shards. More specifically, it transfers all operations that were missed to the
|
||||||
|
target shard. The [WAL] of both shards is used to resolve this. There are two
|
||||||
|
benefits: 1. It will be very fast because it only transfers the difference
|
||||||
|
rather than all data. 2. The ordering guarantees can be `strong`[^ordered],
|
||||||
|
required for some applications. Two disadvantages are: 1. It can only be used to
|
||||||
|
transfer to a shard that already exists on the other node. 2. Applicability is
|
||||||
|
limited because the WALs normally don't hold more than 64MB of recent
|
||||||
|
operations. But that should be enough for a node that quickly restarts, to
|
||||||
|
upgrade for example. If a delta cannot be resolved, this method automatically
|
||||||
|
falls back to `stream_records` which equals transferring the full shard.
|
||||||
|
|
||||||
|
The `stream_records` method is currently used as default. This may change in the
|
||||||
|
future. As of Qdrant 1.9.0 `wal_delta` is used for automatic shard replications
|
||||||
|
to recover dead shards.
|
||||||
|
|
||||||
|
[WAL]: /documentation/manage-data/storage/#versioning
|
||||||
|
|
||||||
|
## Replication
|
||||||
|
|
||||||
|
Shards can be [replicated](/documentation/scaling/horizontal-scaling/#replication) between nodes in the cluster, keeping several copies of a shard spread across the cluster. This enables you to scale your read throughput and tolerate node failures.
|
||||||
|
|
||||||
|
### Replication Factor
|
||||||
|
|
||||||
|
When you create a collection, you can control how many shard replicas you'd like to store by changing the `replication_factor`. By default, `replication_factor` is set to `1`, meaning no additional copy is maintained automatically. The default can be changed in the [Qdrant configuration](/documentation/ops-configuration/configuration/#configuration-options). You can change the default per-collection by setting the `replication_factor` when you create a collection.
|
||||||
|
|
||||||
|
The `replication_factor` can be updated for an existing collection, but the effect of this depends on how you're running Qdrant. If you're hosting the open source version of Qdrant yourself, changing the replication factor after collection creation doesn't do anything. You can manually [create](#creating-new-shard-replicas) or drop shard replicas to achieve your desired replication factor. In Qdrant Cloud (including Hybrid Cloud, Private Cloud) your shards will automatically be replicated or dropped to match your configured replication factor.
|
||||||
|
|
||||||
|
```http
|
||||||
|
PUT /collections/{collection_name}
|
||||||
|
{
|
||||||
|
"vectors": {
|
||||||
|
"size": 300,
|
||||||
|
"distance": "Cosine"
|
||||||
|
},
|
||||||
|
"shard_number": 6,
|
||||||
|
"replication_factor": 2
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
|
client = QdrantClient(url="http://localhost:6333")
|
||||||
|
|
||||||
|
client.create_collection(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
vectors_config=models.VectorParams(size=300, distance=models.Distance.COSINE),
|
||||||
|
shard_number=6,
|
||||||
|
replication_factor=2,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||||
|
|
||||||
|
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||||
|
|
||||||
|
client.createCollection("{collection_name}", {
|
||||||
|
vectors: {
|
||||||
|
size: 300,
|
||||||
|
distance: "Cosine",
|
||||||
|
},
|
||||||
|
shard_number: 6,
|
||||||
|
replication_factor: 2,
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
```rust
|
||||||
|
use qdrant_client::qdrant::{CreateCollectionBuilder, Distance, VectorParamsBuilder};
|
||||||
|
use qdrant_client::Qdrant;
|
||||||
|
|
||||||
|
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||||
|
|
||||||
|
client
|
||||||
|
.create_collection(
|
||||||
|
CreateCollectionBuilder::new("{collection_name}")
|
||||||
|
.vectors_config(VectorParamsBuilder::new(300, Distance::Cosine))
|
||||||
|
.shard_number(6)
|
||||||
|
.replication_factor(2),
|
||||||
|
)
|
||||||
|
.await?;
|
||||||
|
```
|
||||||
|
|
||||||
|
```java
|
||||||
|
import io.qdrant.client.QdrantClient;
|
||||||
|
import io.qdrant.client.QdrantGrpcClient;
|
||||||
|
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||||
|
import io.qdrant.client.grpc.Collections.Distance;
|
||||||
|
import io.qdrant.client.grpc.Collections.VectorParams;
|
||||||
|
import io.qdrant.client.grpc.Collections.VectorsConfig;
|
||||||
|
|
||||||
|
QdrantClient client =
|
||||||
|
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||||
|
|
||||||
|
client
|
||||||
|
.createCollectionAsync(
|
||||||
|
CreateCollection.newBuilder()
|
||||||
|
.setCollectionName("{collection_name}")
|
||||||
|
.setVectorsConfig(
|
||||||
|
VectorsConfig.newBuilder()
|
||||||
|
.setParams(
|
||||||
|
VectorParams.newBuilder()
|
||||||
|
.setSize(300)
|
||||||
|
.setDistance(Distance.Cosine)
|
||||||
|
.build())
|
||||||
|
.build())
|
||||||
|
.setShardNumber(6)
|
||||||
|
.setReplicationFactor(2)
|
||||||
|
.build())
|
||||||
|
.get();
|
||||||
|
```
|
||||||
|
|
||||||
|
```csharp
|
||||||
|
using Qdrant.Client;
|
||||||
|
using Qdrant.Client.Grpc;
|
||||||
|
|
||||||
|
var client = new QdrantClient("localhost", 6334);
|
||||||
|
|
||||||
|
await client.CreateCollectionAsync(
|
||||||
|
collectionName: "{collection_name}",
|
||||||
|
vectorsConfig: new VectorParams { Size = 300, Distance = Distance.Cosine },
|
||||||
|
shardNumber: 6,
|
||||||
|
replicationFactor: 2
|
||||||
|
);
|
||||||
|
```
|
||||||
|
|
||||||
|
```go
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
|
||||||
|
"github.com/qdrant/go-client/qdrant"
|
||||||
|
)
|
||||||
|
|
||||||
|
client, err := qdrant.NewClient(&qdrant.Config{
|
||||||
|
Host: "localhost",
|
||||||
|
Port: 6334,
|
||||||
|
})
|
||||||
|
|
||||||
|
client.CreateCollection(context.Background(), &qdrant.CreateCollection{
|
||||||
|
CollectionName: "{collection_name}",
|
||||||
|
VectorsConfig: qdrant.NewVectorsConfig(&qdrant.VectorParams{
|
||||||
|
Size: 300,
|
||||||
|
Distance: qdrant.Distance_Cosine,
|
||||||
|
}),
|
||||||
|
ShardNumber: qdrant.PtrOf(uint32(6)),
|
||||||
|
ReplicationFactor: qdrant.PtrOf(uint32(2)),
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
This code sample creates a collection with a total of 6 logical shards backed by a total of 12 physical shards.
|
||||||
|
|
||||||
|
Since a replication factor of "2" would require twice as much storage space, it is advised to make sure the hardware can host the additional shard replicas beforehand.
|
||||||
|
|
||||||
|
### Creating New Shard Replicas
|
||||||
|
|
||||||
|
It is possible to create or delete replicas manually on an existing collection using the [Update collection cluster setup API](https://api.qdrant.tech/master/api-reference/distributed/update-collection-cluster). This is usually only necessary if you run Qdrant open-source. In Qdrant Cloud shard replication is handled and updated automatically, matching the configured `replication_factor`.
|
||||||
|
|
||||||
|
A replica can be added on a specific peer by specifying the peer from which to replicate.
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /collections/{collection_name}/cluster
|
||||||
|
{
|
||||||
|
"replicate_shard": {
|
||||||
|
"shard_id": 0,
|
||||||
|
"from_peer_id": 381894127,
|
||||||
|
"to_peer_id": 467122995
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
<aside role="status">You likely want to select a specific <a href="#shard-transfer-method">shard transfer method</a> to get desired performance and guarantees.</aside>
|
||||||
|
|
||||||
|
And a replica can be removed on a specific peer.
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /collections/{collection_name}/cluster
|
||||||
|
{
|
||||||
|
"drop_replica": {
|
||||||
|
"shard_id": 0,
|
||||||
|
"peer_id": 381894127
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Keep in mind that a collection must contain at least one active replica of a shard.
|
||||||
|
|
||||||
|
## Deploy Behind a Load Balancer
|
||||||
|
|
||||||
|
In a multi-node Qdrant cluster, every node can accept requests and route them internally to the correct shards. To get the best performance and availability out of your cluster, distribute requests evenly across all nodes by routing requests through a load balancer.
|
||||||
|
|
||||||
|
Routing all traffic to a single node creates two problems:
|
||||||
|
|
||||||
|
- **Single point of failure:** If that node goes down, your application loses connectivity to the cluster even though the rest of it remains healthy. A load balancer automatically routes requests to any surviving node.
|
||||||
|
- **Idle replicas:** If the receiving node holds a local replica of every shard, it handles the entire query locally and the replicas on all other nodes sit idle. You pay the storage and write cost of replication without gaining any read throughput. A load balancer lets each node serve reads from its own local replicas.
|
||||||
|
|
||||||
|
Qdrant's clients don't distribute requests across nodes on their own. To take full advantage of replication, deploy Qdrant behind a load balancer. This:
|
||||||
|
|
||||||
|
- Eliminates the single point of failure at the entry point.
|
||||||
|
- Ensures replicas on all nodes serve reads.
|
||||||
|
- Distributes the coordinator role across nodes. Each node that receives a request fans out the query to the relevant shards, merges the partial results, and returns the final response so the CPU and memory cost of aggregation is shared rather than concentrated on one node.
|
||||||
|
|
||||||
|
On Qdrant Cloud, clusters are already behind a load balancer. Send all requests to the endpoint provided for each cluster, and Qdrant Cloud handles the load balancing. If you're self-hosting, you need to set up a load balancer yourself.
|
||||||
|
|
||||||
|
## Listener Mode
|
||||||
|
|
||||||
|
<aside role="alert">This is an experimental feature, its behavior may change in the future.</aside>
|
||||||
|
|
||||||
|
In some cases it might be useful to have a Qdrant node that only accumulates data and does not participate in search operations.
|
||||||
|
There are several scenarios where this can be useful:
|
||||||
|
|
||||||
|
- Listener option can be used to store data in a separate node, which can be used for backup purposes or to store data for a long time.
|
||||||
|
- Listener node can be used to synchronize data into another region, while still performing search operations in the local region.
|
||||||
|
|
||||||
|
|
||||||
|
To enable listener mode, set `node_type` to `Listener` in the config file:
|
||||||
|
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
storage:
|
||||||
|
node_type: "Listener"
|
||||||
|
```
|
||||||
|
|
||||||
|
Listener node will not participate in search operations, but will still accept write operations and will store the data in the local storage.
|
||||||
|
|
||||||
|
All shards, stored on the listener node, will be converted to the `Listener` state.
|
||||||
|
|
||||||
|
Additionally, all write requests sent to the listener node will be processed with `wait=false` option, which means that the write operations will be considered successful once they are written to WAL.
|
||||||
|
This mechanism should allow to minimize upsert latency in case of parallel snapshotting.
|
||||||
@@ -0,0 +1,96 @@
|
|||||||
|
---
|
||||||
|
title: Horizontal Scaling
|
||||||
|
short_description: "Understand how Qdrant scales out across nodes through sharding, replication, Raft consensus, and consistency guarantees."
|
||||||
|
description: "Learn how Qdrant's distributed model works: sharding, the default replication behavior, Raft consensus for cluster metadata, and the consistency trade-off."
|
||||||
|
weight: 10
|
||||||
|
---
|
||||||
|
|
||||||
|
# Horizontal Scaling
|
||||||
|
|
||||||
|
Horizontal scaling means adding more nodes to a Qdrant cluster instead of making existing nodes bigger. It's how Qdrant handles data that no longer fits on one node, and it's also the mechanism underlying Qdrant's fault tolerance. This page covers how it works under the hood: sharding, replication, Raft consensus, and consistency guarantees.
|
||||||
|
|
||||||
|
For what these mechanics buy you in terms of fault tolerance and failover, see [Resilience](/documentation/scaling/resilience/). For the practical configuration steps, see [Distributed Deployment](/documentation/scaling/distributed_deployment/).
|
||||||
|
|
||||||
|
## How Many Qdrant Nodes Should I Run?
|
||||||
|
|
||||||
|
The ideal number of Qdrant nodes depends on how much you value cost-saving, resilience, and performance/scalability in relation to each other.
|
||||||
|
|
||||||
|
### One Node
|
||||||
|
|
||||||
|
One node gives you the lowest cost and is easy to set up, but it offers no high availability. This is not recommended for production environments. Drawbacks:
|
||||||
|
|
||||||
|
- Resilience: Users will experience downtime during node restarts. If the data on the node is permanently lost or corrupted, it cannot be recovered aside from snapshots or backups.
|
||||||
|
- Performance: Limited to the resources of a single server.
|
||||||
|
|
||||||
|
### Two Nodes
|
||||||
|
|
||||||
|
Two nodes give you twice the capacity of a single node, and a two-node Qdrant cluster with replicated shards can respond to most read/write requests even when one node is down, such as during maintenance events. But it does not give you true high availability. Drawbacks:
|
||||||
|
|
||||||
|
- No high availability for cluster-wide operations, like creating, editing, or deleting collections. The cluster cannot perform collection operations when one node is down, since those require >50% of nodes to be running, which is only possible in a 3+ node cluster. This may be acceptable if you don't require true high availability.
|
||||||
|
- Data integrity: If the data on one of the two nodes is permanently lost or corrupted, it cannot be recovered aside from snapshots or backups. Only 3+ node clusters can recover from the permanent loss of a single node since recovery operations require >50% of the cluster to be healthy.
|
||||||
|
- Cost: Replicating your shards requires storing two copies of your data.
|
||||||
|
|
||||||
|
### Three or More Nodes
|
||||||
|
|
||||||
|
Three nodes or more provide a highly available cluster, as long as the replication factor is two or higher. Clusters with three or more nodes and replication can perform all operations even while one node is down. They also gain performance benefits from load-balancing, and can recover from the permanent loss of one node without needing backups or snapshots (though backups are still strongly recommended). This is the recommended configuration for production environments. Drawbacks:
|
||||||
|
|
||||||
|
- Cost: Replicating your shards requires storing two copies of your data.
|
||||||
|
- Cost: Larger clusters are more costly than smaller clusters.
|
||||||
|
|
||||||
|
### Which Configuration Is Right for You?
|
||||||
|
|
||||||
|
In summary:
|
||||||
|
|
||||||
|
- One node is suitable for non-production workloads.
|
||||||
|
- Two nodes give you more capacity than one, without true high availability.
|
||||||
|
- Three or more nodes and a replication factor of two or higher are the gold standard for production.
|
||||||
|
|
||||||
|
## Sharding
|
||||||
|
|
||||||
|
A Qdrant collection is partitioned into one or more shards. Each shard is an independent store of points capable of performing all the operations a collection supports. Each shard holds a distinct portion of the collection's points.
|
||||||
|
|
||||||
|
{{< figure src="/documentation/scaling/cluster-no-replication.png" alt="A three-node cluster with a collection with three shards. Each shard holds one-third of the collection's points." caption="A three-node cluster with a collection with three shards. Each shard holds one-third of the collection's points." >}}
|
||||||
|
|
||||||
|
Qdrant distributes points across shards in one of two ways:
|
||||||
|
|
||||||
|
- **Automatic sharding** (default): points are assigned to shards using a [consistent hashing](https://en.wikipedia.org/wiki/Consistent_hashing) algorithm, so each shard manages a non-overlapping subset of points without manual placement.
|
||||||
|
- **[User-defined sharding](/documentation/scaling/distributed_deployment/#user-defined-sharding)**: each point is uploaded to a shard you choose, so operations can target only the shard or shards they need. This is useful for isolating tenants or regions onto dedicated shards.
|
||||||
|
|
||||||
|
Every node knows where all shards are stored through [Raft consensus](#raft-consensus), so a search request sent to any single node automatically fans out to the rest of the cluster to gather the full result.
|
||||||
|
|
||||||
|
As a rule of thumb, create at least 2 shards per node so the cluster can grow without resharding, since [resharding](/documentation/scaling/distributed_deployment/#resharding) is only available in Qdrant Cloud and otherwise requires recreating the collection. If you anticipate significant growth, 12 shards is a common starting point, since it divides evenly as you scale from 1 node up to 2, 3, 6, and 12 nodes. Beyond that, more shards add overhead without much benefit on smaller clusters.
|
||||||
|
|
||||||
|
See [Sharding](/documentation/scaling/distributed_deployment/#sharding) in Distributed Deployment for how to set the shard count, enable user-defined sharding, and move shards between nodes.
|
||||||
|
|
||||||
|
## Replication
|
||||||
|
|
||||||
|
Qdrant allows you to replicate shards between nodes in the cluster, keeping several copies of a shard spread across the cluster. This ensures the availability of data in case of node failures, except if all replicas are lost.
|
||||||
|
|
||||||
|
{{< figure src="/documentation/scaling/cluster-with-replication.png" alt="A three-node cluster with a collection with three shards and a replication factor of two. Each of the three shards (0, 1, and 2) is replicated onto two nodes." caption="A three-node cluster with a collection with three shards and a replication factor of two. Each of the three shards (0, 1, and 2) is replicated onto two nodes." >}}
|
||||||
|
|
||||||
|
By default, Qdrant has no primary or secondary replicas. Writes execute in parallel on all active replicas of a shard, and any replica can serve reads or writes. A "leader" replica only exists when a collection is configured with `medium` or `strong` [write ordering](/documentation/scaling/consistency-guarantees/#write-ordering); even then, the leader is dynamically elected rather than fixed, and it exists to serialize writes for consistency, not to act as a permanent primary. See [Replication factor](/documentation/scaling/distributed_deployment/#replication-factor) and [Creating new shard replicas](/documentation/scaling/distributed_deployment/#creating-new-shard-replicas) for how to configure replication.
|
||||||
|
|
||||||
|
Each replica is in one of three states: active (healthy and serving traffic), dead (unresponsive to health checks or failing to serve traffic), or partial (resynchronizing before it can become active). A dead replica stops receiving traffic from other peers and may need manual intervention if it doesn't recover on its own. This state model is what keeps data consistent and available when only a subset of replicas fail during an update.
|
||||||
|
|
||||||
|
### Consistency Guarantees
|
||||||
|
|
||||||
|
Replication affects consistency. By default, Qdrant prioritizes availability and throughput, so concurrent updates to the same document can leave replicas temporarily inconsistent. The write consistency factor, read consistency, and write ordering options let you tighten those guarantees when your workload requires it. See [Consistency Guarantees](/documentation/scaling/consistency-guarantees/) for configuration details.
|
||||||
|
|
||||||
|
## Raft Consensus
|
||||||
|
|
||||||
|
Qdrant uses the [Raft](https://raft.github.io/) consensus protocol to maintain consistency regarding the cluster topology and the collections structure. Raft consensus only applies to cluster metadata, not to the data itself:
|
||||||
|
|
||||||
|
- Collection operations are part of the consensus. This guarantees that all operations are durable and eventually executed by all nodes. A majority of nodes must agree on what operations to apply before Qdrant performs them.
|
||||||
|
- Point operations don't go through the consensus infrastructure. Qdrant trades strong transaction guarantees for low overhead on point operations: it doesn't guarantee atomic distributed updates, but you can wait until an [operation is complete](/documentation/manage-data/points/#awaiting-result) to see the results of your writes.
|
||||||
|
|
||||||
|
Use the [Check Cluster Status API](https://api.qdrant.tech/master/api-reference/distributed/cluster-status) to check the consensus state.
|
||||||
|
|
||||||
|
For high availability, run at least three voting nodes. A two-node cluster can't form a majority if either node is unavailable, so Raft can't elect or confirm a leader until both nodes can communicate again. During this transition state — whether electing a new leader after a failure or starting up — Qdrant will deny collection update operations.
|
||||||
|
|
||||||
|
Qdrant keeps a Raft log of operations that have modified the cluster state. To keep the Raft log from growing indefinitely, Qdrant uses [consensus checkpointing](/documentation/scaling/node-failure-recovery/#consensus-checkpointing).
|
||||||
|
|
||||||
|
## Where to Go Next
|
||||||
|
|
||||||
|
- [Resilience](/documentation/scaling/resilience/) covers what these mechanics buy you in fault tolerance, Multi-AZ, and failover.
|
||||||
|
- [Distributed Deployment](/documentation/scaling/distributed_deployment/) covers the practical configuration: enabling distributed mode, sharding, replication, and node failure recovery.
|
||||||
|
- [Vertical Scaling](/documentation/scaling/vertical-scaling/) covers resizing existing nodes instead of adding more.
|
||||||
@@ -0,0 +1,55 @@
|
|||||||
|
---
|
||||||
|
title: Node Failure Recovery
|
||||||
|
short_description: "Recover a Qdrant cluster after a node fails, from restarting a replicated node to recreating one from a snapshot."
|
||||||
|
description: "Step-by-step recovery procedures for Qdrant node failures: restarting with replicated collections, recreating a failed node, and recovering from a snapshot when no replicas remain."
|
||||||
|
weight: 25
|
||||||
|
---
|
||||||
|
|
||||||
|
# Node Failure Recovery
|
||||||
|
|
||||||
|
Sometimes hardware malfunctions might render some nodes of the Qdrant cluster unrecoverable. Several recovery scenarios allow Qdrant to stay available for requests and even avoid performance degradation. Let's walk through them from best to worst.
|
||||||
|
|
||||||
|
## Recover with Replicated Collection
|
||||||
|
|
||||||
|
If the number of failed nodes is less than the replication factor of the collection, then your cluster should [still be able to perform read, search, and update queries](/documentation/scaling/resilience/).
|
||||||
|
|
||||||
|
If the failed node restarts, consensus will trigger the replication process to update the recovering node with any updates it missed.
|
||||||
|
|
||||||
|
If the failed node never restarts, you can recover the lost shards if you have a 3+ node cluster. You cannot recover lost shards in smaller clusters because recovery operations go through [Raft](/documentation/scaling/horizontal-scaling/#raft-consensus) which requires >50% of the nodes to be healthy.
|
||||||
|
|
||||||
|
## Recreate a Node with Replicated Collections
|
||||||
|
|
||||||
|
If a node fails and it's impossible to recover it, you should exclude the dead node from the consensus and create a new node:
|
||||||
|
|
||||||
|
1. To exclude failed nodes from the consensus, use [remove peer](https://api.qdrant.tech/master/api-reference/distributed/remove-peer) API. Apply the `force` flag if necessary.
|
||||||
|
1. Create a new node, make sure to attach it to the existing cluster by specifying `--bootstrap` CLI parameter with the URL of any of the running cluster nodes.
|
||||||
|
1. Once the new node is ready and synchronized with the cluster, verify that the collection shards are sufficiently replicated.
|
||||||
|
1. When self-hosting, Qdrant will not automatically balance shards since this is an expensive operation. Use the [Replicate Shard Operation](https://api.qdrant.tech/master/api-reference/distributed/update-collection-cluster) to create new replicas on the newly connected node. On Qdrant Cloud, the system will automatically replicate shards to the new node if the replication factor isn't met.
|
||||||
|
|
||||||
|
## Recover from a Snapshot
|
||||||
|
|
||||||
|
If shards are unrecoverable and there are no other copies in the cluster, you can still recover from a snapshot.
|
||||||
|
|
||||||
|
First, detach the failed node and create a new one:
|
||||||
|
|
||||||
|
1. To exclude failed nodes from the consensus, use [remove peer](https://api.qdrant.tech/master/api-reference/distributed/remove-peer) API. Apply the `force` flag if necessary.
|
||||||
|
1. Create a new node, making sure to attach it to the existing cluster by specifying the `--bootstrap` CLI parameter with the URL of any of the running cluster nodes.
|
||||||
|
1. Use the [Collection Snapshot Recovery API](/documentation/snapshots/#restore-snapshot) to restore the collection. The service will download the specified snapshot of the collection and recover shards with data from it. Snapshot recovery works differently in distributed mode than in single-node deployments. Consensus manages all collection metadata and doesn't require snapshots to restore it. But you can use snapshots to recover missing shards of the collections.
|
||||||
|
|
||||||
|
Once all shards of the collection are recovered, the collection will become operational again.
|
||||||
|
|
||||||
|
## Consensus Checkpointing
|
||||||
|
|
||||||
|
A cluster uses [Raft consensus](/documentation/scaling/horizontal-scaling/#raft-consensus) to manage cluster metadata and ensure that all nodes have a consistent view of the cluster state. Qdrant keeps a Raft log of operations that have modified the cluster state. When a node joins the cluster, it can replay the log to catch up with the current state.
|
||||||
|
|
||||||
|
To keep the Raft log from growing indefinitely, Qdrant uses consensus checkpointing: periodically creating a consistent snapshot of the cluster state that all nodes have agreed on, then truncating the log. Without this, a node that joins a long-running cluster would need to replay the entire log to catch up, which gets slower as the log grows.
|
||||||
|
|
||||||
|
To force a checkpoint, call the `/cluster/recover` API on the required node:
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /cluster/recover
|
||||||
|
```
|
||||||
|
|
||||||
|
This API can be triggered on any non-leader node, it will send a request to the current consensus leader to create a snapshot. The leader will in turn send the snapshot back to the requesting node for application.
|
||||||
|
|
||||||
|
In some cases, this API can be used to recover from an inconsistent cluster state by forcing a snapshot creation.
|
||||||
@@ -0,0 +1,69 @@
|
|||||||
|
---
|
||||||
|
title: Resilience
|
||||||
|
short_description: "Configure Qdrant's replication factor, node count, and Multi-AZ settings for fault tolerance, plus failover best practices for production."
|
||||||
|
description: "Learn how replication factor and node count together determine Qdrant's fault tolerance, how Multi-AZ differs from replication, and how to configure failover for production."
|
||||||
|
weight: 15
|
||||||
|
---
|
||||||
|
|
||||||
|
# Resilience
|
||||||
|
|
||||||
|
Qdrant's fault tolerance is driven by replication and node count. Together, they control whether your data survives a node loss and whether your cluster keeps serving requests when one goes down.
|
||||||
|
|
||||||
|
Resilience covers three distinct aspects of a Qdrant cluster, and it's worth keeping them separate since a cluster can have one without the others:
|
||||||
|
|
||||||
|
- **Availability for reading and writing data**: whether search and write requests keep succeeding while a node is down. See [Temporary Node Failure](#temporary-node-failure) for exactly how differently-configured clusters behave when a node goes down.
|
||||||
|
- **Availability for cluster-wide operations**: whether you can still create, edit, or delete collections while a node is down. This requires a majority of nodes to be healthy, regardless of replication factor. See [How many Qdrant nodes should I run?](/documentation/scaling/horizontal-scaling/#how-many-qdrant-nodes-should-i-run) for how this plays out at different cluster sizes.
|
||||||
|
- **Data integrity**: whether your data survives the permanent loss of a node. This depends on replication factor and node count together, not on either form of availability.
|
||||||
|
|
||||||
|
This page covers what determines Qdrant's fault tolerance and how to configure failover in production. For how the underlying replication and consensus mechanics work, see [Horizontal Scaling](/documentation/scaling/horizontal-scaling/).
|
||||||
|
|
||||||
|
## Setting Up a Resilient Qdrant Cluster
|
||||||
|
|
||||||
|
Two factors determine the resilience of a Qdrant cluster:
|
||||||
|
|
||||||
|
- **[Replication factor](/documentation/scaling/distributed_deployment/#replication)** controls how many copies of each shard exist. More copies mean your data survives the loss of more nodes.
|
||||||
|
- **[Node count](/documentation/scaling/horizontal-scaling/#how-many-qdrant-nodes-should-i-run)** controls how many independent machines those copies can be spread across, and how many nodes can vote in the [Raft consensus](/documentation/scaling/horizontal-scaling/#raft-consensus) that manages collection operations.
|
||||||
|
|
||||||
|
These two factors work together:
|
||||||
|
- A high replication factor on a single node cluster is meaningless: Qdrant won't assign multiple copies of the same shard to a single node. More importantly, it leaves you with a single point of failure. If the node fails, the cluster is unavailable. If the node is unrecoverable, you lose data.
|
||||||
|
- A large node count with a replication factor of one means a single node failure causes unavailability for any collection that had shards on that node. If the node can't be recovered, the data it held is lost, even though the rest of the cluster keeps running.
|
||||||
|
|
||||||
|
Resilience comes from provisioning both together: a highly available cluster consists of three or more nodes, and collections should have a replication factor of two or higher.
|
||||||
|
|
||||||
|
## Temporary Node Failure
|
||||||
|
|
||||||
|
Running Qdrant in distributed mode makes your cluster resistant to outages when one node fails temporarily. Here's how differently-configured Qdrant clusters respond:
|
||||||
|
|
||||||
|
- 1-node clusters: All operations time out or fail for up to a few minutes, depending on how long it takes to restart and load data from disk.
|
||||||
|
- 2-node clusters where shards **are not** replicated: All operations will time out or fail for up to a few minutes, depending on how long it takes to restart and load data from disk.
|
||||||
|
- 2-node clusters where all shards **are** replicated to both nodes: All requests except for operations on collections continue to work during the outage.
|
||||||
|
- 3+-node clusters where all shards are replicated to at least 2 nodes: All requests continue to work during the outage.
|
||||||
|
|
||||||
|
For the steps to recover from a permanent node loss, see [Node Failure Recovery](/documentation/scaling/node-failure-recovery/).
|
||||||
|
|
||||||
|
## Multi-AZ Deployments
|
||||||
|
|
||||||
|
An availability zone (AZ) is an isolated section of a cloud provider's infrastructure, typically a distinct data center or group of data centers within a region, with its own power and networking. A multi-AZ deployment spreads your cluster's nodes across more than one availability zone, so that the loss of a single zone, for example due to a power outage or a network failure, doesn't take down your whole cluster.
|
||||||
|
|
||||||
|
Creating a multi-node cluster with replication does not automatically distribute replicas across availability zones: all replicas may land in the same zone, so a single zone outage could take down every replica of a shard at once, even with a replication factor of two or higher.
|
||||||
|
|
||||||
|
On the [Qdrant Cloud premium tier](/documentation/cloud-premium/), checking the **Multi AZ** box at cluster creation makes Qdrant zone-aware: it automatically distributes shards across availability zones so that each shard has a replica in another zone, and it routes traffic between zones automatically so the cluster stays available if one zone goes down. See [Creating a Production-Ready Cluster](/documentation/cloud/create-cluster/#creating-a-production-ready-cluster).
|
||||||
|
|
||||||
|
Self-hosted Qdrant clusters have no zone awareness and won't automatically distribute replicas across them. You need to spread your nodes across zones yourself, then [move replicas](/documentation/scaling/distributed_deployment/#moving-shards) so that each shard has at least one replica in a different zone.
|
||||||
|
|
||||||
|
## Failover Best Practices
|
||||||
|
|
||||||
|
How Qdrant responds to a node failure depends entirely on how the cluster is provisioned beforehand:
|
||||||
|
|
||||||
|
- **Run at least three nodes with a replication factor of two or more in production.** A single-replica cluster, which is the default, gets none of Qdrant's high-availability guarantees: no automatic failover, no Multi-AZ protection, and no zero-downtime upgrades. All of these depend on having replicated data spread across multiple nodes.
|
||||||
|
- **Enable Multi-AZ if you need protection against a zone-level outage.** Replication alone doesn't guarantee that your nodes, and therefore your replicas, are spread across separate availability zones. Multi-AZ must be enabled explicitly at cluster creation.
|
||||||
|
- **Self-hosted Qdrant does not fail over automatically.** If a node fails permanently, you need to remove it from consensus and attach a replacement node yourself. See [Node Failure Recovery](/documentation/scaling/node-failure-recovery/) for the recovery steps.
|
||||||
|
- **Qdrant Cloud (Managed, Hybrid, and Private) adds automatic failover on top of replication**, so Qdrant Cloud detects and replaces a failed node without manual intervention, provided your cluster has enough replicas to tolerate the loss.
|
||||||
|
- **Monitor cluster health so you can react before a second failure compounds the first.** See [Monitoring & Telemetry](/documentation/ops-monitoring/) for setting up Prometheus and Grafana against your cluster.
|
||||||
|
|
||||||
|
## Where to Go Next
|
||||||
|
|
||||||
|
- [Horizontal Scaling](/documentation/scaling/horizontal-scaling/) covers Raft consensus, the replication model, and consistency guarantees.
|
||||||
|
- [Distributed Deployment](/documentation/scaling/distributed_deployment/) covers the practical configuration: enabling distributed mode, sharding, replication, and node failure recovery.
|
||||||
|
- [Node Failure Recovery](/documentation/scaling/node-failure-recovery/) covers the steps to remove a permanently failed node from consensus and attach a replacement.
|
||||||
|
|
||||||
@@ -0,0 +1,71 @@
|
|||||||
|
---
|
||||||
|
title: Vertical Scaling
|
||||||
|
short_description: "Scale Qdrant by increasing CPU, RAM, or disk on existing nodes, with RAM sizing guidelines and signs it's time to scale horizontally instead."
|
||||||
|
description: "Learn when and how to scale Qdrant vertically by resizing node resources, including RAM sizing formulas, Qdrant Cloud and self-hosted resize steps, and the signals that mean it's time to scale horizontally instead."
|
||||||
|
weight: 5
|
||||||
|
---
|
||||||
|
|
||||||
|
# Vertical Scaling
|
||||||
|
|
||||||
|
Vertical scaling means resizing CPU, RAM, or disk on an existing node. It's simpler than horizontal scaling, avoids distributed system complexity, and is reversible, which makes it the recommended first step whenever a single node's resources are the bottleneck.
|
||||||
|
|
||||||
|
## When to Scale Vertically
|
||||||
|
|
||||||
|
Scale vertically when your current node resources are insufficient but your workload doesn't yet require distribution:
|
||||||
|
|
||||||
|
- RAM usage is approaching 80% of available memory. Beyond this threshold, the operating system starts evicting pages from cache, which causes a sharp performance drop rather than a gradual one.
|
||||||
|
- CPU is saturated during query serving or indexing.
|
||||||
|
- Disk space is running low for on-disk vectors and payloads.
|
||||||
|
- Your workload is non-production or otherwise tolerant of a single point of failure.
|
||||||
|
|
||||||
|
A single node can typically hold up to about 100 million vectors, depending on vector dimensionality and whether quantization is enabled.
|
||||||
|
|
||||||
|
## How to Scale Vertically in Qdrant Cloud
|
||||||
|
|
||||||
|
Vertical scaling in Qdrant Cloud is managed through the [Cloud Console](https://cloud.qdrant.io/):
|
||||||
|
|
||||||
|
1. Select the cluster you want to resize.
|
||||||
|
2. Choose a larger node configuration, increasing CPU, RAM, or both.
|
||||||
|
3. Confirm the resize.
|
||||||
|
|
||||||
|
The resize runs as a rolling restart. If your collections have a replication factor of two or higher, the restart completes with no downtime, since other replicas keep serving traffic while each node restarts in turn. Set replication factor to two or higher before resizing if you need to avoid downtime.
|
||||||
|
|
||||||
|
Scaling up is generally safe. Scaling down needs more care: if your working set no longer fits in RAM after downsizing, performance degrades severely due to cache eviction. Load test before scaling down.
|
||||||
|
|
||||||
|
## How to Scale Vertically Self-Hosted
|
||||||
|
|
||||||
|
For self-hosted deployments, resize the underlying VM or container resources directly, then restart the affected nodes. If your collections have a replication factor of two or higher, you can resize nodes one at a time without downtime.
|
||||||
|
|
||||||
|
## RAM Sizing Guidelines
|
||||||
|
|
||||||
|
RAM is the resource that most directly affects Qdrant's search performance, since search is fastest when vectors and indexes fit in memory.
|
||||||
|
|
||||||
|
Exact RAM usage is difficult to predict precisely, but this formula gives a reasonable estimate for full-precision vectors kept in RAM:
|
||||||
|
|
||||||
|
```text
|
||||||
|
num_vectors * dimensions * 4 bytes * 1.5
|
||||||
|
```
|
||||||
|
|
||||||
|
Quantization can reduce this estimate by a factor of 4 to 32, depending on the quantization method.
|
||||||
|
|
||||||
|
On top of the vector data itself, budget for the HNSW index, which typically adds 20% to 30% overhead, along with payload indexes and the write-ahead log. Reserve about 20% headroom for optimizer operations and operating system cache.
|
||||||
|
|
||||||
|
See [Quantization](/documentation/manage-data/quantization/) for the tradeoffs between quantization methods, and monitor actual memory usage before and after resizing (see [Monitor Collection Memory Usage](/documentation/ops-monitoring/memory-usage/)).
|
||||||
|
|
||||||
|
## When Vertical Scaling Is No Longer Enough
|
||||||
|
|
||||||
|
These signals mean it's time to scale horizontally instead of resizing further:
|
||||||
|
|
||||||
|
- Your data volume exceeds what a single node can hold, even with quantization.
|
||||||
|
- CPU on your largest available node size is already maxed out with unacceptable query latency.
|
||||||
|
- Disk I/O is saturated. Adding nodes gives you more independent disk throughput.
|
||||||
|
- You need fault tolerance, which requires replicating data across nodes.
|
||||||
|
|
||||||
|
When you hit these limits, see [Horizontal Scaling](/documentation/scaling/horizontal-scaling/) and [Distributed Deployment](/documentation/scaling/distributed_deployment/) for how to scale out.
|
||||||
|
|
||||||
|
## Best Practices
|
||||||
|
|
||||||
|
- Load test before scaling down RAM. Cache eviction after downsizing can cause a latency regression.
|
||||||
|
- Keep RAM usage below 80%. Memory pressure in Qdrant causes a performance cliff, not a gradual slowdown.
|
||||||
|
- Set the replication factor to two or higher before resizing in Qdrant Cloud. A rolling restart without replicas causes downtime.
|
||||||
|
- Diagnose the bottleneck before adding CPU. Workloads bound on disk I/O won't improve from additional cores.
|
||||||
@@ -17,7 +17,7 @@ Queries that filter on unindexed fields are not only slower; they can also unnec
|
|||||||
|
|
||||||
## Scale Horizontally with Replicas
|
## Scale Horizontally with Replicas
|
||||||
|
|
||||||
Qdrant can be deployed in a [distributed configuration](/documentation/distributed_deployment/). In distributed mode, multiple instances of Qdrant, called peers, operate as a single entity, called a cluster. Data is stored in [collections](/documentation/manage-data/collections/), which are divided into [shards](/documentation/distributed_deployment/#sharding) that are distributed across the peers. Each shard can have multiple [replicas](/documentation/distributed_deployment/#replication) for redundancy and load balancing. Because every replica of the same shard contains the same data, read requests can be distributed across replicas, reducing latency and increasing throughput.
|
Qdrant can be deployed in a [distributed configuration](/documentation/scaling/distributed_deployment/). In distributed mode, multiple instances of Qdrant, called peers, operate as a single entity, called a cluster. Data is stored in [collections](/documentation/manage-data/collections/), which are divided into [shards](/documentation/scaling/distributed_deployment/#sharding) that are distributed across the peers. Each shard can have multiple [replicas](/documentation/scaling/distributed_deployment/#replication) for redundancy and load balancing. Because every replica of the same shard contains the same data, read requests can be distributed across replicas, reducing latency and increasing throughput.
|
||||||
|
|
||||||
For example, a collection with three shards and a replication factor of two would have six total replicas (two replicas for each of the three shards). On a cluster with three peers, these replicas can be evenly distributed across the peers, with each peer hosting two replicas.
|
For example, a collection with three shards and a replication factor of two would have six total replicas (two replicas for each of the three shards). On a cluster with three peers, these replicas can be evenly distributed across the peers, with each peer hosting two replicas.
|
||||||
|
|
||||||
|
|||||||
@@ -139,7 +139,7 @@ To recover from a URL, you specify an additional parameter in the request body:
|
|||||||
Sometimes it might be handy to create snapshot not just for a single collection, but for the whole storage, including collection aliases.
|
Sometimes it might be handy to create snapshot not just for a single collection, but for the whole storage, including collection aliases.
|
||||||
Qdrant provides a dedicated API for that as well. It is similar to collection-level snapshots, but does not require `collection_name`.
|
Qdrant provides a dedicated API for that as well. It is similar to collection-level snapshots, but does not require `collection_name`.
|
||||||
|
|
||||||
<aside role="alert">Full storage snapshots are only suitable for single-node deployments. <a href="/documentation/distributed_deployment/">Distributed</a> mode is not supported as it doesn't contain the necessary files for that.</aside>
|
<aside role="alert">Full storage snapshots are only suitable for single-node deployments. <a href="/documentation/scaling/distributed_deployment/">Distributed</a> mode is not supported as it doesn't contain the necessary files for that.</aside>
|
||||||
|
|
||||||
<aside role="status">Full storage snapshots can be created and downloaded from Qdrant Cloud, but you cannot restore a Qdrant Cloud cluster from a whole storage snapshot since that requires use of the Qdrant CLI. You can use <a href="/documentation/cloud/backups/">Backups</a> instead.</aside>
|
<aside role="status">Full storage snapshots can be created and downloaded from Qdrant Cloud, but you cannot restore a Qdrant Cloud cluster from a whole storage snapshot since that requires use of the Qdrant CLI. You can use <a href="/documentation/cloud/backups/">Backups</a> instead.</aside>
|
||||||
|
|
||||||
|
|||||||
+1
-1
@@ -184,7 +184,7 @@ qdrant_urls = [
|
|||||||
"/documentation/installation",
|
"/documentation/installation",
|
||||||
"/documentation/search/filtering",
|
"/documentation/search/filtering",
|
||||||
"/documentation/manage-data/indexing",
|
"/documentation/manage-data/indexing",
|
||||||
"/documentation/distributed_deployment",
|
"/documentation/scaling/distributed_deployment",
|
||||||
"/documentation/manage-data/quantization"
|
"/documentation/manage-data/quantization"
|
||||||
# Add more URLs as needed
|
# Add more URLs as needed
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -9,7 +9,7 @@ weight: 35
|
|||||||
|
|
||||||
When working with massive, fast-moving datasets, like social media or image/video streams, efficient storage and retrieval are critical. Often, only the most recent data is relevant, while older data can be archived or deleted. For instance, in sentiment analysis of social media posts, you might only need the last 7 days of data to capture current trends, with most queries focusing on the last 24 hours.
|
When working with massive, fast-moving datasets, like social media or image/video streams, efficient storage and retrieval are critical. Often, only the most recent data is relevant, while older data can be archived or deleted. For instance, in sentiment analysis of social media posts, you might only need the last 7 days of data to capture current trends, with most queries focusing on the last 24 hours.
|
||||||
|
|
||||||
Storing everything in Qdrant collection with default sharding can lead to expensive re-indexing across the entire dataset when deleting old points, impacting performance. A better solution is **time-based sharding**, where points are routed to a specific [shard (or shards)](/documentation/distributed_deployment/#sharding) based on timestamp. For use cases with a natural time-to-live (TTL) segmentation, sharding by a timestamp-based key enables efficient querying of recent data and allows users to seamlessly drop the old.
|
Storing everything in Qdrant collection with default sharding can lead to expensive re-indexing across the entire dataset when deleting old points, impacting performance. A better solution is **time-based sharding**, where points are routed to a specific [shard (or shards)](/documentation/scaling/distributed_deployment/#sharding) based on timestamp. For use cases with a natural time-to-live (TTL) segmentation, sharding by a timestamp-based key enables efficient querying of recent data and allows users to seamlessly drop the old.
|
||||||
|
|
||||||
For example, with daily shards, today's data is stored in today's shard, yesterday's data in yesterday's shard, and so on. Queries can target specific shards (today's shard, for example) or multiple shards to cover a date range.
|
For example, with daily shards, today's data is stored in today's shard, yesterday's data in yesterday's shard, and so on. Queries can target specific shards (today's shard, for example) or multiple shards to cover a date range.
|
||||||
|
|
||||||
@@ -46,7 +46,7 @@ This tutorial assumes you are using [Qdrant Cloud Inference](/documentation/infe
|
|||||||
|
|
||||||
## Create Collection
|
## Create Collection
|
||||||
|
|
||||||
Create a collection with [user-defined sharding](/documentation/distributed_deployment/#user-defined-sharding) by setting the sharding method to custom.
|
Create a collection with [user-defined sharding](/documentation/scaling/distributed_deployment/#user-defined-sharding) by setting the sharding method to custom.
|
||||||
|
|
||||||
{{< code-snippet path="/documentation/headless/snippets/time-based-sharding/" block="create-collection" >}}
|
{{< code-snippet path="/documentation/headless/snippets/time-based-sharding/" block="create-collection" >}}
|
||||||
|
|
||||||
|
|||||||
@@ -40,6 +40,9 @@
|
|||||||
/documentation/operations/security/* /documentation/security/:splat 301
|
/documentation/operations/security/* /documentation/security/:splat 301
|
||||||
/documentation/operations/common-errors/* /documentation/common-errors/:splat 301
|
/documentation/operations/common-errors/* /documentation/common-errors/:splat 301
|
||||||
|
|
||||||
|
# distributed_deployment moved under the Scaling & Resilience section
|
||||||
|
/documentation/distributed_deployment/* /documentation/scaling/distributed_deployment/:splat 301
|
||||||
|
|
||||||
# Operations sub-pages moved to new sub-sections (docs reorganization)
|
# Operations sub-pages moved to new sub-sections (docs reorganization)
|
||||||
/documentation/operations/configuration/* /documentation/ops-configuration/configuration/:splat 301
|
/documentation/operations/configuration/* /documentation/ops-configuration/configuration/:splat 301
|
||||||
/documentation/operations/administration/* /documentation/ops-configuration/administration/:splat 301
|
/documentation/operations/administration/* /documentation/ops-configuration/administration/:splat 301
|
||||||
|
|||||||
Binary file not shown.
|
Before Width: | Height: | Size: 18 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 18 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 24 KiB |
Reference in New Issue
Block a user