mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-09 21:08:31 +02:00
Break up "Distributed Deployment" page into new "Scaling & Resilience" section (#2491)
* Add new Scaling landing page under Operations Introduces a Scaling section with vertical vs. horizontal scaling guidance and failover best practices, linking out to detail pages. * Add new Vertical Scaling page Dedicated how-to guidance for resizing existing nodes: when to scale vertically, RAM sizing formulas, and Cloud/self-hosted resize steps. * Add new Horizontal Scaling and Resilience page Covers Raft consensus, the replication model, consistency guarantees, Multi-AZ, and the resilience terminology used elsewhere in the docs. * Move Distributed Deployment under Scaling and update all incoming links Moves distributed_deployment.md into the new scaling/ section, trims its Raft/Replication/Consistency intros into cross-links to the new Horizontal Scaling and Resilience page, adds Multi-AZ and single-replica cross-link callouts in the Cloud docs, rewrites all internal references across ~30 files to the new canonical path instead of relying on aliases, and applies Title Case to Distributed Deployment's headers. * Split Resilience out of Horizontal Scaling and Resilience Adds a dedicated Resilience page covering fault tolerance, Multi-AZ, resilience terminology, and failover best practices (moved from the Scaling landing page). Horizontal Scaling is retitled and scoped to the underlying mechanics: Raft consensus, replication, and consistency. * Reorganize Horizontal Scaling's structure Moves "How Many Qdrant Nodes Should I Run?" from Distributed Deployment into Horizontal Scaling, adds a conceptual Sharding section, and reorders Sharding/Replication/Raft Consensus/Consistency. Moves the remaining conceptual content out of Distributed Deployment: Temporary Node Failure to Resilience, Error Handling folded into Replication, sharding heuristics folded into Sharding, and the Consensus Checkpointing explanation folded into Raft Consensus. * Rename Scaling section to Scaling & Resilience Renames the section and restructures the landing page: the vertical- vs-horizontal decision is now purely about scaling, with a dedicated Resilience section covering fault tolerance through sharding and multi-node deployments. * Polish Vertical Scaling and Resilience page content Reframes Vertical Scaling's "What Not to Do" as positive "Best Practices". Reworks Resilience's structure: moves the uptime/data- integrity terminology into the intro as three distinct aspects of resilience, and renames "How Resilience Works" to "Setting Up a Resilient Qdrant Cluster". * Add diagrams illustrating sharding and replication Adds cluster diagrams to the Sharding and Replication sections on Horizontal Scaling to make the shard/replica layout easier to follow. * Add new Node Failure Recovery page Extracts the node failure recovery scenarios out of Distributed Deployment into their own page, with each bolded sub-header converted to a proper heading, and links updated across Resilience and the Scaling landing page. * Add new Consistency Guarantees page Extracts write consistency factor, read consistency, and write ordering out of Distributed Deployment into their own page, positioned after Distributed Deployment. * Add new "Deploy Behind a Load Balancer" section Explains why a load balancer is needed in front of a multi-node Qdrant cluster: avoiding a single point of failure at the entry point and making sure replicas on every node actually serve reads. * Add new "Rebalancing" section Documents how Qdrant Cloud automatically rebalances shards across nodes, as its own subsection under Sharding. * Rewrite Multi-AZ vs. Replication Factor as Multi-AZ Deployments Defines an availability zone on first use, explains why multi-AZ deployments guard against a zone going down, clarifies that Qdrant Cloud is zone-aware once enabled, and that self-hosted deployments need to place and move replicas across zones manually. * Restructure node-count guidance into One/Two/Three-or-more Node subsections Splits "How Many Qdrant Nodes Should I Run?" into three subsections and drops the "balanced" framing for two nodes: it states plainly that two nodes give more capacity without true high availability. * Add new "Which Configuration Is Right for You?" section Summarizes the one/two/three-or-more node tradeoffs in one place right after the detailed breakdown. * Add explicit _redirects entry for legacy distributed_deployment URL Closes the redirect chain: the existing /guides/ and /operations/ legacy rules both terminate at /documentation/distributed_deployment/, which previously had no explicit _redirects entry and only resolved via the Hugo alias meta-refresh page. * Fix all incoming links to Distributed Deployment and pages under Scaling Repoints two same-page anchors in distributed_deployment.md that broke when Write Ordering moved to Consistency Guarantees, and one link in cloud/create-cluster.md that broke when a Resilience heading was reworded. * Update time-based sharding diagram and restructure section * Fix a couple of broken links * Move 'Consensus Checkpointing' to 'Node Failure Recovery' page
This commit is contained in:
@@ -0,0 +1,71 @@
|
||||
---
|
||||
title: Vertical Scaling
|
||||
short_description: "Scale Qdrant by increasing CPU, RAM, or disk on existing nodes, with RAM sizing guidelines and signs it's time to scale horizontally instead."
|
||||
description: "Learn when and how to scale Qdrant vertically by resizing node resources, including RAM sizing formulas, Qdrant Cloud and self-hosted resize steps, and the signals that mean it's time to scale horizontally instead."
|
||||
weight: 5
|
||||
---
|
||||
|
||||
# Vertical Scaling
|
||||
|
||||
Vertical scaling means resizing CPU, RAM, or disk on an existing node. It's simpler than horizontal scaling, avoids distributed system complexity, and is reversible, which makes it the recommended first step whenever a single node's resources are the bottleneck.
|
||||
|
||||
## When to Scale Vertically
|
||||
|
||||
Scale vertically when your current node resources are insufficient but your workload doesn't yet require distribution:
|
||||
|
||||
- RAM usage is approaching 80% of available memory. Beyond this threshold, the operating system starts evicting pages from cache, which causes a sharp performance drop rather than a gradual one.
|
||||
- CPU is saturated during query serving or indexing.
|
||||
- Disk space is running low for on-disk vectors and payloads.
|
||||
- Your workload is non-production or otherwise tolerant of a single point of failure.
|
||||
|
||||
A single node can typically hold up to about 100 million vectors, depending on vector dimensionality and whether quantization is enabled.
|
||||
|
||||
## How to Scale Vertically in Qdrant Cloud
|
||||
|
||||
Vertical scaling in Qdrant Cloud is managed through the [Cloud Console](https://cloud.qdrant.io/):
|
||||
|
||||
1. Select the cluster you want to resize.
|
||||
2. Choose a larger node configuration, increasing CPU, RAM, or both.
|
||||
3. Confirm the resize.
|
||||
|
||||
The resize runs as a rolling restart. If your collections have a replication factor of two or higher, the restart completes with no downtime, since other replicas keep serving traffic while each node restarts in turn. Set replication factor to two or higher before resizing if you need to avoid downtime.
|
||||
|
||||
Scaling up is generally safe. Scaling down needs more care: if your working set no longer fits in RAM after downsizing, performance degrades severely due to cache eviction. Load test before scaling down.
|
||||
|
||||
## How to Scale Vertically Self-Hosted
|
||||
|
||||
For self-hosted deployments, resize the underlying VM or container resources directly, then restart the affected nodes. If your collections have a replication factor of two or higher, you can resize nodes one at a time without downtime.
|
||||
|
||||
## RAM Sizing Guidelines
|
||||
|
||||
RAM is the resource that most directly affects Qdrant's search performance, since search is fastest when vectors and indexes fit in memory.
|
||||
|
||||
Exact RAM usage is difficult to predict precisely, but this formula gives a reasonable estimate for full-precision vectors kept in RAM:
|
||||
|
||||
```text
|
||||
num_vectors * dimensions * 4 bytes * 1.5
|
||||
```
|
||||
|
||||
Quantization can reduce this estimate by a factor of 4 to 32, depending on the quantization method.
|
||||
|
||||
On top of the vector data itself, budget for the HNSW index, which typically adds 20% to 30% overhead, along with payload indexes and the write-ahead log. Reserve about 20% headroom for optimizer operations and operating system cache.
|
||||
|
||||
See [Quantization](/documentation/manage-data/quantization/) for the tradeoffs between quantization methods, and monitor actual memory usage before and after resizing (see [Monitor Collection Memory Usage](/documentation/ops-monitoring/memory-usage/)).
|
||||
|
||||
## When Vertical Scaling Is No Longer Enough
|
||||
|
||||
These signals mean it's time to scale horizontally instead of resizing further:
|
||||
|
||||
- Your data volume exceeds what a single node can hold, even with quantization.
|
||||
- CPU on your largest available node size is already maxed out with unacceptable query latency.
|
||||
- Disk I/O is saturated. Adding nodes gives you more independent disk throughput.
|
||||
- You need fault tolerance, which requires replicating data across nodes.
|
||||
|
||||
When you hit these limits, see [Horizontal Scaling](/documentation/scaling/horizontal-scaling/) and [Distributed Deployment](/documentation/scaling/distributed_deployment/) for how to scale out.
|
||||
|
||||
## Best Practices
|
||||
|
||||
- Load test before scaling down RAM. Cache eviction after downsizing can cause a latency regression.
|
||||
- Keep RAM usage below 80%. Memory pressure in Qdrant causes a performance cliff, not a gradual slowdown.
|
||||
- Set the replication factor to two or higher before resizing in Qdrant Cloud. A rolling restart without replicas causes downtime.
|
||||
- Diagnose the bottleneck before adding CPU. Workloads bound on disk I/O won't improve from additional cores.
|
||||
Reference in New Issue
Block a user