* Add new Scaling landing page under Operations Introduces a Scaling section with vertical vs. horizontal scaling guidance and failover best practices, linking out to detail pages. * Add new Vertical Scaling page Dedicated how-to guidance for resizing existing nodes: when to scale vertically, RAM sizing formulas, and Cloud/self-hosted resize steps. * Add new Horizontal Scaling and Resilience page Covers Raft consensus, the replication model, consistency guarantees, Multi-AZ, and the resilience terminology used elsewhere in the docs. * Move Distributed Deployment under Scaling and update all incoming links Moves distributed_deployment.md into the new scaling/ section, trims its Raft/Replication/Consistency intros into cross-links to the new Horizontal Scaling and Resilience page, adds Multi-AZ and single-replica cross-link callouts in the Cloud docs, rewrites all internal references across ~30 files to the new canonical path instead of relying on aliases, and applies Title Case to Distributed Deployment's headers. * Split Resilience out of Horizontal Scaling and Resilience Adds a dedicated Resilience page covering fault tolerance, Multi-AZ, resilience terminology, and failover best practices (moved from the Scaling landing page). Horizontal Scaling is retitled and scoped to the underlying mechanics: Raft consensus, replication, and consistency. * Reorganize Horizontal Scaling's structure Moves "How Many Qdrant Nodes Should I Run?" from Distributed Deployment into Horizontal Scaling, adds a conceptual Sharding section, and reorders Sharding/Replication/Raft Consensus/Consistency. Moves the remaining conceptual content out of Distributed Deployment: Temporary Node Failure to Resilience, Error Handling folded into Replication, sharding heuristics folded into Sharding, and the Consensus Checkpointing explanation folded into Raft Consensus. * Rename Scaling section to Scaling & Resilience Renames the section and restructures the landing page: the vertical- vs-horizontal decision is now purely about scaling, with a dedicated Resilience section covering fault tolerance through sharding and multi-node deployments. * Polish Vertical Scaling and Resilience page content Reframes Vertical Scaling's "What Not to Do" as positive "Best Practices". Reworks Resilience's structure: moves the uptime/data- integrity terminology into the intro as three distinct aspects of resilience, and renames "How Resilience Works" to "Setting Up a Resilient Qdrant Cluster". * Add diagrams illustrating sharding and replication Adds cluster diagrams to the Sharding and Replication sections on Horizontal Scaling to make the shard/replica layout easier to follow. * Add new Node Failure Recovery page Extracts the node failure recovery scenarios out of Distributed Deployment into their own page, with each bolded sub-header converted to a proper heading, and links updated across Resilience and the Scaling landing page. * Add new Consistency Guarantees page Extracts write consistency factor, read consistency, and write ordering out of Distributed Deployment into their own page, positioned after Distributed Deployment. * Add new "Deploy Behind a Load Balancer" section Explains why a load balancer is needed in front of a multi-node Qdrant cluster: avoiding a single point of failure at the entry point and making sure replicas on every node actually serve reads. * Add new "Rebalancing" section Documents how Qdrant Cloud automatically rebalances shards across nodes, as its own subsection under Sharding. * Rewrite Multi-AZ vs. Replication Factor as Multi-AZ Deployments Defines an availability zone on first use, explains why multi-AZ deployments guard against a zone going down, clarifies that Qdrant Cloud is zone-aware once enabled, and that self-hosted deployments need to place and move replicas across zones manually. * Restructure node-count guidance into One/Two/Three-or-more Node subsections Splits "How Many Qdrant Nodes Should I Run?" into three subsections and drops the "balanced" framing for two nodes: it states plainly that two nodes give more capacity without true high availability. * Add new "Which Configuration Is Right for You?" section Summarizes the one/two/three-or-more node tradeoffs in one place right after the detailed breakdown. * Add explicit _redirects entry for legacy distributed_deployment URL Closes the redirect chain: the existing /guides/ and /operations/ legacy rules both terminate at /documentation/distributed_deployment/, which previously had no explicit _redirects entry and only resolved via the Hugo alias meta-refresh page. * Fix all incoming links to Distributed Deployment and pages under Scaling Repoints two same-page anchors in distributed_deployment.md that broke when Write Ordering moved to Consistency Guarantees, and one link in cloud/create-cluster.md that broke when a Resilience heading was reworded. * Update time-based sharding diagram and restructure section * Fix a couple of broken links * Move 'Consensus Checkpointing' to 'Node Failure Recovery' page
4.5 KiB
title, short_description, description, weight
| title | short_description | description | weight |
|---|---|---|---|
| Node Failure Recovery | Recover a Qdrant cluster after a node fails, from restarting a replicated node to recreating one from a snapshot. | Step-by-step recovery procedures for Qdrant node failures: restarting with replicated collections, recreating a failed node, and recovering from a snapshot when no replicas remain. | 25 |
Node Failure Recovery
Sometimes hardware malfunctions might render some nodes of the Qdrant cluster unrecoverable. Several recovery scenarios allow Qdrant to stay available for requests and even avoid performance degradation. Let's walk through them from best to worst.
Recover with Replicated Collection
If the number of failed nodes is less than the replication factor of the collection, then your cluster should still be able to perform read, search, and update queries.
If the failed node restarts, consensus will trigger the replication process to update the recovering node with any updates it missed.
If the failed node never restarts, you can recover the lost shards if you have a 3+ node cluster. You cannot recover lost shards in smaller clusters because recovery operations go through Raft which requires >50% of the nodes to be healthy.
Recreate a Node with Replicated Collections
If a node fails and it's impossible to recover it, you should exclude the dead node from the consensus and create a new node:
- To exclude failed nodes from the consensus, use remove peer API. Apply the
forceflag if necessary. - Create a new node, make sure to attach it to the existing cluster by specifying
--bootstrapCLI parameter with the URL of any of the running cluster nodes. - Once the new node is ready and synchronized with the cluster, verify that the collection shards are sufficiently replicated.
- When self-hosting, Qdrant will not automatically balance shards since this is an expensive operation. Use the Replicate Shard Operation to create new replicas on the newly connected node. On Qdrant Cloud, the system will automatically replicate shards to the new node if the replication factor isn't met.
Recover from a Snapshot
If shards are unrecoverable and there are no other copies in the cluster, you can still recover from a snapshot.
First, detach the failed node and create a new one:
- To exclude failed nodes from the consensus, use remove peer API. Apply the
forceflag if necessary. - Create a new node, making sure to attach it to the existing cluster by specifying the
--bootstrapCLI parameter with the URL of any of the running cluster nodes. - Use the Collection Snapshot Recovery API to restore the collection. The service will download the specified snapshot of the collection and recover shards with data from it. Snapshot recovery works differently in distributed mode than in single-node deployments. Consensus manages all collection metadata and doesn't require snapshots to restore it. But you can use snapshots to recover missing shards of the collections.
Once all shards of the collection are recovered, the collection will become operational again.
Consensus Checkpointing
A cluster uses Raft consensus to manage cluster metadata and ensure that all nodes have a consistent view of the cluster state. Qdrant keeps a Raft log of operations that have modified the cluster state. When a node joins the cluster, it can replay the log to catch up with the current state.
To keep the Raft log from growing indefinitely, Qdrant uses consensus checkpointing: periodically creating a consistent snapshot of the cluster state that all nodes have agreed on, then truncating the log. Without this, a node that joins a long-running cluster would need to replay the entire log to catch up, which gets slower as the log grows.
To force a checkpoint, call the /cluster/recover API on the required node:
POST /cluster/recover
This API can be triggered on any non-leader node, it will send a request to the current consensus leader to create a snapshot. The leader will in turn send the snapshot back to the requesting node for application.
In some cases, this API can be used to recover from an inconsistent cluster state by forcing a snapshot creation.