Break up "Distributed Deployment" page into new "Scaling & Resilience" section (#2491)

* Add new Scaling landing page under Operations

Introduces a Scaling section with vertical vs. horizontal scaling
guidance and failover best practices, linking out to detail pages.

* Add new Vertical Scaling page

Dedicated how-to guidance for resizing existing nodes: when to scale
vertically, RAM sizing formulas, and Cloud/self-hosted resize steps.

* Add new Horizontal Scaling and Resilience page

Covers Raft consensus, the replication model, consistency guarantees,
Multi-AZ, and the resilience terminology used elsewhere in the docs.

* Move Distributed Deployment under Scaling and update all incoming links

Moves distributed_deployment.md into the new scaling/ section, trims
its Raft/Replication/Consistency intros into cross-links to the new
Horizontal Scaling and Resilience page, adds Multi-AZ and single-replica
cross-link callouts in the Cloud docs, rewrites all internal references
across ~30 files to the new canonical path instead of relying on
aliases, and applies Title Case to Distributed Deployment's headers.

* Split Resilience out of Horizontal Scaling and Resilience

Adds a dedicated Resilience page covering fault tolerance, Multi-AZ,
resilience terminology, and failover best practices (moved from the
Scaling landing page). Horizontal Scaling is retitled and scoped to
the underlying mechanics: Raft consensus, replication, and consistency.

* Reorganize Horizontal Scaling's structure

Moves "How Many Qdrant Nodes Should I Run?" from Distributed Deployment
into Horizontal Scaling, adds a conceptual Sharding section, and
reorders Sharding/Replication/Raft Consensus/Consistency. Moves the
remaining conceptual content out of Distributed Deployment: Temporary
Node Failure to Resilience, Error Handling folded into Replication,
sharding heuristics folded into Sharding, and the Consensus
Checkpointing explanation folded into Raft Consensus.

* Rename Scaling section to Scaling & Resilience

Renames the section and restructures the landing page: the vertical-
vs-horizontal decision is now purely about scaling, with a dedicated
Resilience section covering fault tolerance through sharding and
multi-node deployments.

* Polish Vertical Scaling and Resilience page content

Reframes Vertical Scaling's "What Not to Do" as positive "Best
Practices". Reworks Resilience's structure: moves the uptime/data-
integrity terminology into the intro as three distinct aspects of
resilience, and renames "How Resilience Works" to "Setting Up a
Resilient Qdrant Cluster".

* Add diagrams illustrating sharding and replication

Adds cluster diagrams to the Sharding and Replication sections on
Horizontal Scaling to make the shard/replica layout easier to follow.

* Add new Node Failure Recovery page

Extracts the node failure recovery scenarios out of Distributed
Deployment into their own page, with each bolded sub-header converted
to a proper heading, and links updated across Resilience and the
Scaling landing page.

* Add new Consistency Guarantees page

Extracts write consistency factor, read consistency, and write
ordering out of Distributed Deployment into their own page, positioned
after Distributed Deployment.

* Add new "Deploy Behind a Load Balancer" section

Explains why a load balancer is needed in front of a multi-node
Qdrant cluster: avoiding a single point of failure at the entry point
and making sure replicas on every node actually serve reads.

* Add new "Rebalancing" section

Documents how Qdrant Cloud automatically rebalances shards across
nodes, as its own subsection under Sharding.

* Rewrite Multi-AZ vs. Replication Factor as Multi-AZ Deployments

Defines an availability zone on first use, explains why multi-AZ
deployments guard against a zone going down, clarifies that Qdrant
Cloud is zone-aware once enabled, and that self-hosted deployments
need to place and move replicas across zones manually.

* Restructure node-count guidance into One/Two/Three-or-more Node subsections

Splits "How Many Qdrant Nodes Should I Run?" into three subsections
and drops the "balanced" framing for two nodes: it states plainly
that two nodes give more capacity without true high availability.

* Add new "Which Configuration Is Right for You?" section

Summarizes the one/two/three-or-more node tradeoffs in one place
right after the detailed breakdown.

* Add explicit _redirects entry for legacy distributed_deployment URL

Closes the redirect chain: the existing /guides/ and /operations/
legacy rules both terminate at /documentation/distributed_deployment/,
which previously had no explicit _redirects entry and only resolved
via the Hugo alias meta-refresh page.

* Fix all incoming links to Distributed Deployment and pages under Scaling

Repoints two same-page anchors in distributed_deployment.md that broke
when Write Ordering moved to Consistency Guarantees, and one link in
cloud/create-cluster.md that broke when a Resilience heading was
reworded.

* Update time-based sharding diagram and restructure section

* Fix a couple of broken links

* Move 'Consensus Checkpointing' to 'Node Failure Recovery' page
This commit is contained in:
Abdon Pijpelink
2026-07-16 09:11:58 +02:00
committed by GitHub
parent e413d5107c
commit 45f19f30ee
45 changed files with 1620 additions and 1405 deletions
@@ -0,0 +1,71 @@
---
title: Vertical Scaling
short_description: "Scale Qdrant by increasing CPU, RAM, or disk on existing nodes, with RAM sizing guidelines and signs it's time to scale horizontally instead."
description: "Learn when and how to scale Qdrant vertically by resizing node resources, including RAM sizing formulas, Qdrant Cloud and self-hosted resize steps, and the signals that mean it's time to scale horizontally instead."
weight: 5
---
# Vertical Scaling
Vertical scaling means resizing CPU, RAM, or disk on an existing node. It's simpler than horizontal scaling, avoids distributed system complexity, and is reversible, which makes it the recommended first step whenever a single node's resources are the bottleneck.
## When to Scale Vertically
Scale vertically when your current node resources are insufficient but your workload doesn't yet require distribution:
- RAM usage is approaching 80% of available memory. Beyond this threshold, the operating system starts evicting pages from cache, which causes a sharp performance drop rather than a gradual one.
- CPU is saturated during query serving or indexing.
- Disk space is running low for on-disk vectors and payloads.
- Your workload is non-production or otherwise tolerant of a single point of failure.
A single node can typically hold up to about 100 million vectors, depending on vector dimensionality and whether quantization is enabled.
## How to Scale Vertically in Qdrant Cloud
Vertical scaling in Qdrant Cloud is managed through the [Cloud Console](https://cloud.qdrant.io/):
1. Select the cluster you want to resize.
2. Choose a larger node configuration, increasing CPU, RAM, or both.
3. Confirm the resize.
The resize runs as a rolling restart. If your collections have a replication factor of two or higher, the restart completes with no downtime, since other replicas keep serving traffic while each node restarts in turn. Set replication factor to two or higher before resizing if you need to avoid downtime.
Scaling up is generally safe. Scaling down needs more care: if your working set no longer fits in RAM after downsizing, performance degrades severely due to cache eviction. Load test before scaling down.
## How to Scale Vertically Self-Hosted
For self-hosted deployments, resize the underlying VM or container resources directly, then restart the affected nodes. If your collections have a replication factor of two or higher, you can resize nodes one at a time without downtime.
## RAM Sizing Guidelines
RAM is the resource that most directly affects Qdrant's search performance, since search is fastest when vectors and indexes fit in memory.
Exact RAM usage is difficult to predict precisely, but this formula gives a reasonable estimate for full-precision vectors kept in RAM:
```text
num_vectors * dimensions * 4 bytes * 1.5
```
Quantization can reduce this estimate by a factor of 4 to 32, depending on the quantization method.
On top of the vector data itself, budget for the HNSW index, which typically adds 20% to 30% overhead, along with payload indexes and the write-ahead log. Reserve about 20% headroom for optimizer operations and operating system cache.
See [Quantization](/documentation/manage-data/quantization/) for the tradeoffs between quantization methods, and monitor actual memory usage before and after resizing (see [Monitor Collection Memory Usage](/documentation/ops-monitoring/memory-usage/)).
## When Vertical Scaling Is No Longer Enough
These signals mean it's time to scale horizontally instead of resizing further:
- Your data volume exceeds what a single node can hold, even with quantization.
- CPU on your largest available node size is already maxed out with unacceptable query latency.
- Disk I/O is saturated. Adding nodes gives you more independent disk throughput.
- You need fault tolerance, which requires replicating data across nodes.
When you hit these limits, see [Horizontal Scaling](/documentation/scaling/horizontal-scaling/) and [Distributed Deployment](/documentation/scaling/distributed_deployment/) for how to scale out.
## Best Practices
- Load test before scaling down RAM. Cache eviction after downsizing can cause a latency regression.
- Keep RAM usage below 80%. Memory pressure in Qdrant causes a performance cliff, not a gradual slowdown.
- Set the replication factor to two or higher before resizing in Qdrant Cloud. A rolling restart without replicas causes downtime.
- Diagnose the bottleneck before adding CPU. Workloads bound on disk I/O won't improve from additional cores.