mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-04 02:18:29 +02:00
Improvements for cloud docs
* Documentation page for the different cluster configuration options * Add RBAC and IaC to getting started * Docs for automatic shard rebalancing in cloud * Docs for restart policy in cloud
This commit is contained in:
@@ -12,7 +12,9 @@ Welcome to Qdrant Managed Cloud! This document contains all the information you
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Before creating a cluster, make sure you have a Qdrant Cloud account. Detailed instructions for signing up can be found in the [Qdrant Cloud Setup](/documentation/cloud/qdrant-cloud-setup/) guide. You also need to provide [payment details](/documentation/cloud/pricing-payments/). If you have a custom payment agreement, first create your account, then [contact our Support Team](https://support.qdrant.io/) to finalize the setup.
|
||||
Before creating a cluster, make sure you have a Qdrant Cloud account. Detailed instructions for signing up can be found in the [Qdrant Cloud Setup](/documentation/cloud/qdrant-cloud-setup/) guide. Qdrant Cloud supports granular [role-based access control](/documentation/cloud-rbac/).
|
||||
|
||||
You also need to provide [payment details](/documentation/cloud/pricing-payments/). If you have a custom payment agreement, first create your account, then [contact our Support Team](https://support.qdrant.io/) to finalize the setup.
|
||||
|
||||
Premium Plan subscribers can enable single sign-on (SSO) for their organizations. To activate SSO, please reach out to the Support Team at [https://support.qdrant.io/](https://support.qdrant.io/) for guidance.
|
||||
|
||||
@@ -26,6 +28,10 @@ After setting up your account, you can create a Qdrant Cluster by following the
|
||||
|
||||
## Preparing for Production
|
||||
|
||||
For a production-ready environment, consider deploying a multi-node Qdrant cluster (at least three nodes) with replication enabled. Instructions for configuring distributed clusters are available in the [Distributed Deployment](/documentation/guides/distributed_deployment/) guide.
|
||||
For a production-ready environment, consider deploying a multi-node Qdrant cluster (at least three nodes) with replication enabled. More details are available in the [Distributed Deployment](/documentation/guides/distributed_deployment/) guide. For more information on how to create a production-ready cluster, see our [Vector Search in Production](/articles/vector-search-production/) article.
|
||||
|
||||
If you are looking to optimize costs, you can reduce memory usage through [Quantization](/documentation/guides/quantization/) or by [offloading vectors to disk](/documentation/concepts/storage/#configuring-memmap-storage).
|
||||
|
||||
## Infrastructure as Code Automation
|
||||
|
||||
Qdrant Cloud can be fully automated using the [Qdrant Cloud API](/documentation/cloud-api/). This allows you to create, manage, and scale clusters programmatically. You can also use our [Terraform Provider](https://registry.terraform.io/providers/qdrant/qdrant-cloud) to automate your Qdrant Cloud infrastructure.
|
||||
|
||||
@@ -16,13 +16,14 @@ You can also attach your own infrastructure as a Hybrid Cloud Environment. For d
|
||||
|
||||
## Cluster Configuration
|
||||
|
||||
Each database cluster comes pre-configured with the following tools, features, and support services:
|
||||
|
||||
- Allows the creation of highly available clusters with automatic failover.
|
||||
- Supports upgrades to later versions of Qdrant as they are released.
|
||||
- Upgrades are zero-downtime on highly available clusters.
|
||||
- Includes monitoring and logging to observe the health of each cluster.
|
||||
- Horizontally and vertically scalable.
|
||||
- Available natively on AWS and GCP, and Azure.
|
||||
- Available on your own infrastructure and other providers if you use the Hybrid Cloud.
|
||||
Each database cluster comes pre-configured with the following features:
|
||||
|
||||
- Allows the creation of highly available clusters with automatic failover
|
||||
- Easy version upgrades, zero-downtime on highly available clusters
|
||||
- Monitoring, logging and alerting to observe the health of each cluster
|
||||
- Horizontal and vertical up and down scaling
|
||||
- Automatic shard rebalancing
|
||||
- Support for resharding
|
||||
- Backups and disaster recovery
|
||||
- Available natively on AWS and GCP, and Azure.
|
||||
- Available on your own infrastructure and other providers if you use the Hybrid Cloud
|
||||
|
||||
@@ -5,17 +5,17 @@ weight: 50
|
||||
|
||||
# Scaling Qdrant Cloud Clusters
|
||||
|
||||
The amount of data is always growing and at some point you might need to upgrade or downgrade the capacity of your cluster.
|
||||
The amount of data is always growing and at some point you might need to change the capacity of your cluster. You can easily scale your Qdrant cluster up or down from the Cluster detail page in the Qdrant Cloud console.
|
||||
|
||||

|
||||
|
||||
There are different options for how it can be done.
|
||||
|
||||
## Vertical Scaling
|
||||
|
||||
Vertical scaling is the process of increasing the capacity of a cluster by adding or removing CPU, storage and memory resources on each database node.
|
||||
|
||||
You can start with a minimal cluster configuration of 2GB of RAM and resize it up to 64GB of RAM (or even more if desired) over the time step by step with the growing amount of data in your application. If your cluster consists of several nodes each node will need to be scaled to the same size. Please note that vertical cluster scaling will require a short downtime period to restart your cluster. In order to avoid a downtime you can make use of data replication, which can be configured on the collection level. Vertical scaling can be initiated on the cluster detail page via the button "scale".
|
||||
You can start with a minimal cluster configuration and scale it up over the time to accomodate the growing amount of data in your application. If your cluster consists of several nodes each node will need to be scaled to the same size.
|
||||
|
||||
Note that vertical cluster scaling will require a short downtime, if the collections in your cluster are not replicated. This is because each node of the cluster needs to be restarted to apply the CPU, memory and disk size.
|
||||
|
||||
If you want to scale your cluster down, the new, smaller memory size must be still sufficient to store all the data in the cluster. Otherwise, the database cluster could run out of memory and crash. Therefore, the new memory size must be at least as large as the current memory usage of the database cluster including a bit of buffer. Qdrant Cloud will automatically prevent you from scaling down the Qdrant database cluster with a too small memory size.
|
||||
|
||||
@@ -27,14 +27,14 @@ Vertical scaling can be an effective way to improve the performance of a cluster
|
||||
|
||||
In such cases, horizontal scaling may be a more effective solution.
|
||||
|
||||
Horizontal scaling, also known as horizontal expansion, is the process of increasing the capacity of a cluster by adding more nodes and distributing the load and data among them. The horizontal scaling at Qdrant starts on the collection level. You have to choose the number of shards you want to distribute your collection around while creating the collection. Please refer to the [sharding documentation](/documentation/guides/distributed_deployment/#sharding) section for details.
|
||||
Horizontal scaling, is the process of increasing the capacity of a cluster by adding more nodes and distributing the load and data among them. The horizontal scaling at Qdrant starts on the collection level. You have to choose the number of shards you want to distribute your collection around while creating the collection. Please refer to the [sharding documentation](/documentation/guides/distributed_deployment/#sharding) section for details.
|
||||
|
||||
After that, you can configure, or change the amount of Qdrant database nodes within a cluster during cluster creation, or on the cluster detail page via "Scale" button.
|
||||
|
||||
Important: The number of shards means the maximum amount of nodes you can add to your cluster. In the beginning, all the shards can reside on one node. With the growing amount of data you can add nodes to your cluster and move shards to the dedicated nodes using the [cluster setup API](/documentation/guides/distributed_deployment/#cluster-scaling).
|
||||
When scaling up horizontally, the cloud paltform will automatically rebalance all available shards across nodes to ensure that the data is evenly distributed. See [Configuring Clusters](/documentation/cloud/configure-cluster/#shard-rebalancing) for more details.
|
||||
|
||||
When scaling down horizontally, the cloud platform will automatically ensure that any shards that are present on the nodes to be deleted, are moved to the remaining nodes.
|
||||
|
||||
Important: If you configure e.g. 2 shards for a collection, but then scale your cluster from 1 to 3 nodes, your cluster nodes can't be fully utilized. The cloud platform will automatically rebalance your shards, so that two nodes will have one shard each, but the third node will not have any shards at all. You can use the [resharding feature](/documentation/cloud/cluster-scaling/#resharding) to change the number of shards in an existing collection. Once the resharding is complete, the cloud platform will rebalance the shards across all nodes, ensuring that all nodes are utilized.
|
||||
|
||||
We will be glad to consult you on an optimal strategy for scaling.
|
||||
|
||||
[Let us know](/documentation/support/) your needs and decide together on a proper solution.
|
||||
|
||||
@@ -14,3 +14,5 @@ To update to a new version, go to the Cluster details page, choose the new versi
|
||||
If you have a multi-node cluster and if your collections have a replication factor of at least **2**, the update process will be zero-downtime and done in a rolling fashion. You will be able to use your database cluster normally.
|
||||
|
||||
If you have a single-node cluster or a collection with a replication factor of **1**, the update process will require a short downtime period to restart your cluster with the new version.
|
||||
|
||||
See also [Restart Mode](/documentation/cloud/configure-cluster/#restart-mode) for more details.
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
---
|
||||
title: Configure Clusters
|
||||
weight: 55
|
||||
---
|
||||
|
||||
# Configure Qdrant Cloud Clusters
|
||||
|
||||
Qdrant Cloud offers several advanced configuration options to optimizer your clusters to your specific needs. You can access these options from the Cluster details page in the Qdrant Cloud console.
|
||||
|
||||
## Collection Defaults
|
||||
|
||||
You can set default values for the configuration of new collections in your cluster. These defaults will be used when creating a new collection, unless you override them in the collection creation request.
|
||||
|
||||
You can configure the default *Replication Factor*, the default *Write Consistency Factor* and if vectors should be stored on disk only, instead of being cached in RAM.
|
||||
|
||||
Refer to [Qdrant Configuration](/documentation/guides/configuration/#configuration-options) for more details.
|
||||
|
||||
## Advanced Optimizations
|
||||
|
||||
You can change the *Optimzer CPU Budget* and the *Async Scorer* configurations for your cluster. These advanced settings will have an impact on performance and reliability. We recommend using the default values unless you are confident they are required for your use case.
|
||||
|
||||
See [Qdrant under the hood: io_uring](/articles/io_uring/#and-what-about-qdrant) and [Large Scale Search](/documentation/database-tutorials/large-scale-search/) for more details.
|
||||
|
||||
## Client IP Restrictions
|
||||
|
||||
If configured, only the chosen IP ranges will be allowed to access the cluster. This is useful for securing your cluster and ensuring that only trusted clients can connect to it.
|
||||
|
||||
## Restart Mode
|
||||
|
||||
The cloud platform will automatically choose the best restart mode during version upgrades or maintenance for your cluster. If you have a multi-node cluster and one or more collections with a replication factor of at least 2, the cloud platform will use a rolling restart mode. This means that the cluster will be restarted one node at a time, ensuring that the cluster remains available during the restart process.
|
||||
|
||||
If you have a multi-node cluster, but all collections have a replication factor of 1, the cloud platform will use a parallel restart mode. This means that the cluster will be restarted all at once, which will result in a short downtime period, but will be faster than a rolling restart.
|
||||
|
||||
However, you can override this setting if you want to use a specific restart mode.
|
||||
|
||||
## Shard Rebalancing
|
||||
|
||||
When you scale your cluster horizontally, the cloud platform will automatically rebalance the shards across the nodes to ensure that the data is evenly distributed. This is done to ensure that all nodes are utilized and that the performance of the cluster is optimal.
|
||||
|
||||
Qdrant Cloud offers three strategies for shard rebalancing:
|
||||
|
||||
* `by_count_and_size` (default): This strategy will rebalance the shards based on the number of shards and their size. It will ensure that all nodes have the same number of shards and that the size of the shards is evenly distributed across the nodes.
|
||||
* `by_count`: This strategy will rebalance the shards based on the number of shards only. It will ensure that all nodes have the same number of shards, but the size of the shards may not be evenly distributed across the nodes.
|
||||
* `by_size`: This strategy will rebalance the shards based on their size only. It will ensure that the size of the shards is evenly distributed across the nodes, but the number of shards may not be the same on all nodes.
|
||||
@@ -91,6 +91,26 @@ Once provisioned, you can access your cluster on ports 443 and 6333 (REST) and 6
|
||||
|
||||
You should now see the new cluster in the **Clusters** menu.
|
||||
|
||||
## Creating a Production-Ready Cluster
|
||||
|
||||
To create a production-ready cluster, you need to ensure the following:
|
||||
|
||||
**High Availability**
|
||||
|
||||
Your cluster should have at least 3 nodes, and each collection should have a replication factor of at least 2. This ensures that if one node fails, or is being restarted due to maintenance, version upgrades or scaling operations, the cluster remains fully operational. You can ensure this by checking the **High Availability** checkbox when creating a cluster.
|
||||
|
||||
**Backup and Disaster Recovery**
|
||||
|
||||
You should create a backup schedule for your cluster. This ensures that you can restore your data in case of a disaster. You can configure backups in the **Backups** section of the cluster detail page. See [**Backups**](/documentation/cloud/backups/) for more information.
|
||||
|
||||
**Collection Sharding**
|
||||
|
||||
To allow your cluster to scale horizontally easily, you should configure at least 2 times the shards per collection than the number of nodes in your cluster. You can configure the number of shards when creating a collection. See [**Sharding**](/documentation/guides/distributed_deployment/#sharding) for more information.
|
||||
|
||||
If you did not configure enough shards in a collection, you can use the [**Resharding**](/documentation/cloud/cluster-scaling/#resharding) feature to change the number of shards in an existing collection.
|
||||
|
||||
For more information on how to create a production-ready cluster, see our [**Vector Search in Production**](/articles/vector-search-production/) article.
|
||||
|
||||
## Deleting a Cluster
|
||||
|
||||
You can delete a Qdrant database cluster from the cluster's detail page.
|
||||
|
||||
@@ -135,6 +135,8 @@ For best results, first ensure your cluster is running Qdrant v1.7.4 or higher.
|
||||
|
||||
In the [Qdrant Cloud console](https://cloud.qdrant.io/), click "Scale Up" to increase your cluster size to >1. Qdrant Cloud configures the distributed mode settings automatically.
|
||||
|
||||
Additionally, Qdrant Cloud also offers the ability to automatically rebalance and to reshard your collections, which is not available in self-hosted Qdrant. See the [Resharding](/documentation/cloud/cluster-scaling/#resharding) and [Shard Rebalancing](/documentation/cloud/configure-cluster/#shard-rebalancing) sections in for more details.
|
||||
|
||||
After the scale-up process completes, you will have a new empty node running alongside your existing node(s). To replicate data into this new empty node, see the next section.
|
||||
|
||||
## Making use of a new distributed Qdrant cluster
|
||||
@@ -143,8 +145,8 @@ When you enable distributed mode and scale up to two or more nodes, your data do
|
||||
|
||||
* Create a new replicated collection by setting the [replication_factor](#replication-factor) to 2 or more and setting the [number of shards](#choosing-the-right-number-of-shards) to a multiple of your number of nodes.
|
||||
* If you have an existing collection which does not contain enough shards for each node, you must create a new collection as described in the previous bullet point.
|
||||
* If you already have enough shards for each node and you merely need to replicate your data, follow the directions for [creating new shard replicas](#creating-new-shard-replicas).
|
||||
* If you already have enough shards for each node and your data is already replicated, you can move data (without replicating it) onto the new node(s) by [moving shards](#moving-shards).
|
||||
* If you already have enough shards for each node, and you merely need to replicate your data, follow the directions for [creating new shard replicas](#creating-new-shard-replicas).
|
||||
* If you already have enough shards for each node, and your data is already replicated, you can move data (without replicating it) onto the new node(s) by [moving shards](#moving-shards).
|
||||
|
||||
## Raft
|
||||
|
||||
@@ -318,6 +320,8 @@ Please refer to the [Resharding](/documentation/cloud/cluster-scaling/#reshardin
|
||||
|
||||
Qdrant allows moving shards between nodes in the cluster and removing nodes from the cluster. This functionality unlocks the ability to dynamically scale the cluster size without downtime. It also allows you to upgrade or migrate nodes without downtime.
|
||||
|
||||
If your cluster is running in Qdrant Cloud, the shards are balanced across the cluster nodes automatically. For more information see the [Configuring Cloud Clusters](/documentation/cloud/configure-cluster/#shard-rebalancing) and [Cloud Cluster Scaling](/documentation/cloud/cluster-scaling/) documentation.
|
||||
|
||||
Qdrant provides the information regarding the current shard distribution in the cluster with the [Collection Cluster info API](https://api.qdrant.tech/master/api-reference/distributed/collection-cluster-info).
|
||||
|
||||
Use the [Update collection cluster setup API](https://api.qdrant.tech/master/api-reference/distributed/update-collection-cluster) to initiate the shard transfer:
|
||||
|
||||
Reference in New Issue
Block a user