mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-09 21:08:31 +02:00
Break up "Distributed Deployment" page into new "Scaling & Resilience" section (#2491)
* Add new Scaling landing page under Operations Introduces a Scaling section with vertical vs. horizontal scaling guidance and failover best practices, linking out to detail pages. * Add new Vertical Scaling page Dedicated how-to guidance for resizing existing nodes: when to scale vertically, RAM sizing formulas, and Cloud/self-hosted resize steps. * Add new Horizontal Scaling and Resilience page Covers Raft consensus, the replication model, consistency guarantees, Multi-AZ, and the resilience terminology used elsewhere in the docs. * Move Distributed Deployment under Scaling and update all incoming links Moves distributed_deployment.md into the new scaling/ section, trims its Raft/Replication/Consistency intros into cross-links to the new Horizontal Scaling and Resilience page, adds Multi-AZ and single-replica cross-link callouts in the Cloud docs, rewrites all internal references across ~30 files to the new canonical path instead of relying on aliases, and applies Title Case to Distributed Deployment's headers. * Split Resilience out of Horizontal Scaling and Resilience Adds a dedicated Resilience page covering fault tolerance, Multi-AZ, resilience terminology, and failover best practices (moved from the Scaling landing page). Horizontal Scaling is retitled and scoped to the underlying mechanics: Raft consensus, replication, and consistency. * Reorganize Horizontal Scaling's structure Moves "How Many Qdrant Nodes Should I Run?" from Distributed Deployment into Horizontal Scaling, adds a conceptual Sharding section, and reorders Sharding/Replication/Raft Consensus/Consistency. Moves the remaining conceptual content out of Distributed Deployment: Temporary Node Failure to Resilience, Error Handling folded into Replication, sharding heuristics folded into Sharding, and the Consensus Checkpointing explanation folded into Raft Consensus. * Rename Scaling section to Scaling & Resilience Renames the section and restructures the landing page: the vertical- vs-horizontal decision is now purely about scaling, with a dedicated Resilience section covering fault tolerance through sharding and multi-node deployments. * Polish Vertical Scaling and Resilience page content Reframes Vertical Scaling's "What Not to Do" as positive "Best Practices". Reworks Resilience's structure: moves the uptime/data- integrity terminology into the intro as three distinct aspects of resilience, and renames "How Resilience Works" to "Setting Up a Resilient Qdrant Cluster". * Add diagrams illustrating sharding and replication Adds cluster diagrams to the Sharding and Replication sections on Horizontal Scaling to make the shard/replica layout easier to follow. * Add new Node Failure Recovery page Extracts the node failure recovery scenarios out of Distributed Deployment into their own page, with each bolded sub-header converted to a proper heading, and links updated across Resilience and the Scaling landing page. * Add new Consistency Guarantees page Extracts write consistency factor, read consistency, and write ordering out of Distributed Deployment into their own page, positioned after Distributed Deployment. * Add new "Deploy Behind a Load Balancer" section Explains why a load balancer is needed in front of a multi-node Qdrant cluster: avoiding a single point of failure at the entry point and making sure replicas on every node actually serve reads. * Add new "Rebalancing" section Documents how Qdrant Cloud automatically rebalances shards across nodes, as its own subsection under Sharding. * Rewrite Multi-AZ vs. Replication Factor as Multi-AZ Deployments Defines an availability zone on first use, explains why multi-AZ deployments guard against a zone going down, clarifies that Qdrant Cloud is zone-aware once enabled, and that self-hosted deployments need to place and move replicas across zones manually. * Restructure node-count guidance into One/Two/Three-or-more Node subsections Splits "How Many Qdrant Nodes Should I Run?" into three subsections and drops the "balanced" framing for two nodes: it states plainly that two nodes give more capacity without true high availability. * Add new "Which Configuration Is Right for You?" section Summarizes the one/two/three-or-more node tradeoffs in one place right after the detailed breakdown. * Add explicit _redirects entry for legacy distributed_deployment URL Closes the redirect chain: the existing /guides/ and /operations/ legacy rules both terminate at /documentation/distributed_deployment/, which previously had no explicit _redirects entry and only resolved via the Hugo alias meta-refresh page. * Fix all incoming links to Distributed Deployment and pages under Scaling Repoints two same-page anchors in distributed_deployment.md that broke when Write Ordering moved to Consistency Guarantees, and one link in cloud/create-cluster.md that broke when a Resilience heading was reworded. * Update time-based sharding diagram and restructure section * Fix a couple of broken links * Move 'Consensus Checkpointing' to 'Node Failure Recovery' page
This commit is contained in:
@@ -126,7 +126,7 @@ The same agent that speeds through a toy dataset with 10,000 points will become
|
||||
|
||||
We’ll talk about three concepts you can take advantage of to improve your scale, but if you want even more information on how to scale, check out this [article](https://qdrant.tech/documentation/database-tutorials/large-scale-search/) on large scale search.
|
||||
|
||||
As your dataset and traffic grow, Qdrant Cloud offers a suite of features to ensure your system can scale effectively. Horizontal scaling is achieved through [sharding](https://qdrant.tech/articles/multitenancy/), which splits your collection across multiple nodes to distribute the load and improve performance. For high availability and fault tolerance, Qdrant supports [replication](https://qdrant.tech/documentation/distributed_deployment/), creating copies of your shards across the cluster.
|
||||
As your dataset and traffic grow, Qdrant Cloud offers a suite of features to ensure your system can scale effectively. Horizontal scaling is achieved through [sharding](https://qdrant.tech/articles/multitenancy/), which splits your collection across multiple nodes to distribute the load and improve performance. For high availability and fault tolerance, Qdrant supports [replication](https://qdrant.tech/documentation/scaling/distributed_deployment/), creating copies of your shards across the cluster.
|
||||
|
||||
Qdrant provides robust tools for resource and cost optimization. Vector [quantization](https://qdrant.tech/documentation/manage-data/quantization/) compresses your data, significantly reducing its memory footprint and speeding up search.
|
||||
|
||||
|
||||
@@ -19,14 +19,14 @@ category: production-ops
|
||||
|
||||
# Scaling Your Machine Learning Setup: The Power of Multitenancy and Custom Sharding in Qdrant
|
||||
|
||||
We are seeing the topics of [multitenancy](/documentation/manage-data/multitenancy/) and [distributed deployment](/documentation/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup.
|
||||
We are seeing the topics of [multitenancy](/documentation/manage-data/multitenancy/) and [distributed deployment](/documentation/scaling/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup.
|
||||
|
||||
Whether you are building a bank fraud-detection system, [RAG](https://qdrant.tech/articles/what-is-rag-in-ai/) for e-commerce, or services for the federal government - you will need to leverage a multitenant architecture to scale your product.
|
||||
In the world of SaaS and enterprise apps, this setup is the norm. It will considerably increase your application's performance and lower your hosting costs.
|
||||
|
||||
## Multitenancy & custom sharding with Qdrant
|
||||
|
||||
We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](/documentation/manage-data/multitenancy/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](/documentation/distributed_deployment/#user-defined-sharding).
|
||||
We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](/documentation/manage-data/multitenancy/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](/documentation/scaling/distributed_deployment/#user-defined-sharding).
|
||||
|
||||
Combining these two will result in an efficiently-partitioned architecture that further leverages the convenience of a single Qdrant cluster. This article will briefly explain the benefits and show how you can get started using both features.
|
||||
|
||||
@@ -41,7 +41,7 @@ Qdrant is built to excel in a single collection with a vast number of tenants. Y
|
||||
|
||||
## Sharding your database
|
||||
|
||||
With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](/documentation/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node.
|
||||
With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](/documentation/scaling/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node.
|
||||
|
||||
During vector search, your operations will be able to hit only the subset of shards they actually need. In massive-scale deployments, __this can significantly improve the performance of operations that do not require the whole collection to be scanned__.
|
||||
|
||||
@@ -49,7 +49,7 @@ This works in the other direction as well. Whenever you search for something, yo
|
||||
|
||||
### Common use cases
|
||||
|
||||
A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](/documentation/distributed_deployment/#moving-shards).
|
||||
A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](/documentation/scaling/distributed_deployment/#moving-shards).
|
||||
|
||||
**Figure 2:** Users can both upsert and query shards that are relevant to them, all within the same collection. Regional sharding can help avoid cross-continental traffic.
|
||||

|
||||
@@ -79,7 +79,7 @@ client.create_shard_key("{tenant_data}", "germany")
|
||||
```
|
||||
In this example, your cluster is divided between Germany and Canada. Canadian and German law differ when it comes to international data transfer. Let's say you are creating a RAG application that supports the healthcare industry. Your Canadian customer data will have to be clearly separated for compliance purposes from your German customer.
|
||||
|
||||
Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](/documentation/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech).
|
||||
Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](/documentation/scaling/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech).
|
||||
|
||||
## Configure a multitenant setup for users
|
||||
|
||||
|
||||
@@ -92,7 +92,7 @@ POST /collections/my_collection/points/search
|
||||
}
|
||||
```
|
||||
|
||||
If you want to know more about the user-defined sharding, please refer to the [sharding documentation](/documentation/distributed_deployment/#sharding).
|
||||
If you want to know more about the user-defined sharding, please refer to the [sharding documentation](/documentation/scaling/distributed_deployment/#sharding).
|
||||
|
||||
### Snapshot-based shard transfer
|
||||
|
||||
@@ -101,7 +101,7 @@ That's a really more in depth technical improvement for the distributed mode use
|
||||
Moving shards is required for dynamical scaling of the cluster. Your data can migrate between nodes, and the way you move it is crucial for the performance of the whole system. The good old `stream_records` method (still the default one) transmits all the records between the machines and indexes them on the target node.
|
||||
In the case of moving the shard, it's necessary to recreate the HNSW index each time. However, with the introduction of the new `snapshot` approach, the snapshot itself, inclusive of all data and potentially quantized content, is transferred to the target node. This comprehensive snapshot includes the entire index, enabling the target node to seamlessly load it and promptly begin handling requests without the need for index recreation.
|
||||
|
||||
There are multiple scenarios in which you may prefer one over the other. Please check out the docs of the [shard transfer method](/documentation/distributed_deployment/#shard-transfer-method) for more details and head-to-head comparison. As for now, the old `stream_records` method is still the default one, but we may decide to change it in the future.
|
||||
There are multiple scenarios in which you may prefer one over the other. Please check out the docs of the [shard transfer method](/documentation/scaling/distributed_deployment/#shard-transfer-method) for more details and head-to-head comparison. As for now, the old `stream_records` method is still the default one, but we may decide to change it in the future.
|
||||
|
||||
## Minor improvements
|
||||
|
||||
|
||||
@@ -243,7 +243,7 @@ It depends. If you're just starting out - we have prepared a tool on our website
|
||||
|
||||
A three-node setup provides a baseline for fault tolerance: if one node goes offline, the remaining two can continue serving queries and maintain a quorum for data consistency. This guards against hardware failures, rolling updates, and network disruptions. Fewer than three nodes leaves you vulnerable to single-point failures that can knock your entire cluster offline.
|
||||
|
||||
> [**We follow the Raft Protocol**](https://qdrant.tech/documentation/distributed_deployment/#raft), so check out the docs and learn why this is important.
|
||||
> [**We follow the Raft Protocol**](https://qdrant.tech/documentation/scaling/horizontal-scaling/#raft-consensus), so check out the docs and learn why this is important.
|
||||
|
||||
✅ **Set a replication factor of at least 2** to tolerate node failure without losing availability.
|
||||
|
||||
@@ -275,7 +275,7 @@ Development and staging environments often run experimental builds, tests, or si
|
||||
|
||||
> It's quite possible that the user has multiple shards on one node, which end up handling most traffic while other nodes remain underutilized.
|
||||
|
||||
In this case, you should [**choose the right number of shards**](https://qdrant.tech/documentation/distributed_deployment/#sharding) based on your node count and expected RPS.
|
||||
In this case, you should [**choose the right number of shards**](https://qdrant.tech/documentation/scaling/distributed_deployment/#sharding) based on your node count and expected RPS.
|
||||
|
||||
You need to implement a shard strategy that aligns with real usage patterns. First, distribute your shards across all available nodes. This will help balance the load more effectively. After redistributing the shards, run performance tests to see how it affects your system. Then add replicas and test again to see how that changes performance.
|
||||
|
||||
@@ -286,14 +286,14 @@ Proper sharding considers data distribution and query patterns. By default, shar
|
||||
|
||||
||
|
||||
|-|
|
||||
|**Read More:** [**Sharding Documentation**](https://qdrant.tech/documentation/distributed_deployment/#sharding)|
|
||||
|**Read More:** [**Sharding Documentation**](https://qdrant.tech/documentation/scaling/distributed_deployment/#sharding)|
|
||||
|
||||
### Manage Your Costs by Scaling Up or Down
|
||||

|
||||
|
||||
Some teams scale up for daytime surges, then scale down overnight to save resources. If you do this, ensure data is sharded and replicated appropriately, so that scaling up and down won't result in service degradation.
|
||||
|
||||
If using Qdrant Cloud you could also do this using the [**Replication Factor**](https://qdrant.tech/documentation/distributed_deployment/#replication-factor), though it may be considered a bit of a hack.
|
||||
If using Qdrant Cloud you could also do this using the [**Replication Factor**](https://qdrant.tech/documentation/scaling/distributed_deployment/#replication-factor), though it may be considered a bit of a hack.
|
||||
|
||||
> If you have 3 nodes with just 1 shard, and replication factor 6. It will create 3 replicas (one on each node) of that shard, because it can't host more. If you add 3 more nodes at peak times, it'll automatically replicate that shard 3 more times in an attempt to match the factor of 6.
|
||||
|
||||
@@ -309,7 +309,7 @@ If new nodes remain empty after joining, you waste resources. If departing nodes
|
||||
|
||||
||
|
||||
|-|
|
||||
|**Read More:** [**Distributed Deployment Documentation**](https://qdrant.tech/documentation/distributed_deployment/)|
|
||||
|**Read More:** [**Distributed Deployment Documentation**](https://qdrant.tech/documentation/scaling/distributed_deployment/)|
|
||||
|**Read More:** [**Resharding**](https://qdrant.tech/documentation/cloud/cluster-scaling/#resharding)|
|
||||
|
||||
### How to Predict and Test Cluster Performance
|
||||
@@ -332,7 +332,7 @@ Remember, cold-starts and query behaviour are dataset dependent, which is why yo
|
||||
|
||||
||
|
||||
|-|
|
||||
|**Read More:** [Distributed Deployment Documentation](https://qdrant.tech/documentation/distributed_deployment/)
|
||||
|**Read More:** [Distributed Deployment Documentation](https://qdrant.tech/documentation/scaling/distributed_deployment/)
|
||||
|
||||
### How to Design Your Systems to Protect Against Failure
|
||||
|
||||
|
||||
@@ -325,7 +325,7 @@ Here’s how to choose the shard_number:
|
||||
| **Plan for Scalability** | Start with at least **2 shards per node** to allow room for future growth. |
|
||||
| **Future-Proofing** | Starting with around **12 shards** is a good rule of thumb. This setup allows your system to scale seamlessly from 1 to 12 nodes without requiring re-sharding. |
|
||||
|
||||
Learn more about [**Sharding in Distributed Deployment**](/documentation/distributed_deployment/)
|
||||
Learn more about [**Sharding in Distributed Deployment**](/documentation/scaling/distributed_deployment/)
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -413,7 +413,7 @@ client.create_collection(
|
||||
|
||||
We recommend using sharding and replication together so that your data is both split across nodes and replicated for availability.
|
||||
|
||||
For more details on features like **user-defined sharding, node failure recovery**, and **consistency guarantees**, see our guide on [Distributed Deployment.](https://qdrant.tech/documentation/distributed_deployment/)
|
||||
For more details on features like **user-defined sharding, node failure recovery**, and **consistency guarantees**, see our guide on [Distributed Deployment.](https://qdrant.tech/documentation/scaling/distributed_deployment/)
|
||||
|
||||
## Multitenancy: Data Isolation for Multi-Tenant Architectures
|
||||
|
||||
|
||||
Reference in New Issue
Block a user