From 9e045ed9e2c8832cd30cfe7283c493b9b8095d06 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Tim=20Vis=C3=A9e?= Date: Mon, 22 Apr 2024 15:28:38 +0200 Subject: [PATCH] Document WAL delta transfer for Qdrant 1.9 (#736) * List Qdrant version that supports each shard transfer method * Update shard transfer method feature table * Describe WAL delta transfer in more detail * Improve first shard transfer method sentence * Improve WAL delta text * List possible targets in shard transfer method feature matrix * Update connectivity text in shard transfer method feature matrix * Remove unordered lists to give shard transfer feature matrix more space * Shorten text in shard transfer method feature matrix * Add notice, shard transfer methods may affect ordering guarantees * Fix minor capitalization issue * Update qdrant-landing/content/documentation/guides/distributed_deployment.md Co-authored-by: Anush * Apply suggestions from code review Co-authored-by: Anush --------- Co-authored-by: Anush --- .../guides/distributed_deployment.md | 52 +++++++++++++------ 1 file changed, 36 insertions(+), 16 deletions(-) diff --git a/qdrant-landing/content/documentation/guides/distributed_deployment.md b/qdrant-landing/content/documentation/guides/distributed_deployment.md index 104b01108..df4ca4dc0 100644 --- a/qdrant-landing/content/documentation/guides/distributed_deployment.md +++ b/qdrant-landing/content/documentation/guides/distributed_deployment.md @@ -560,26 +560,29 @@ Another use-case would be to have shards that track the data chronologically, so *Available as of v1.7.0* -There are different methods for transferring, such as moving or replicating, a -shard to another node. Depending on what performance and guarantees you'd like -to have and how you'd like to manage your cluster, you likely want to choose a -specific method. Each method has its own pros and cons. Which is fastest depends -on the size and state of a shard. +There are different methods for transferring a shard, such as moving or +replicating, to another node. Depending on what performance and guarantees you'd +like to have and how you'd like to manage your cluster, you likely want to +choose a specific method. Each method has its own pros and cons. Which is +fastest depends on the size and state of a shard. Available shard transfer methods are: -- `stream_records`: _(default)_ transfer shard by streaming just its records to the target node in batches. -- `snapshot`: transfer shard including its index and quantized data by utilizing a [snapshot](../../concepts/snapshots/) automatically. +- `stream_records`: _(default)_ transfer by streaming just its records to the target node in batches. +- `snapshot`: transfer including its index and quantized data by utilizing a [snapshot](../../concepts/snapshots/) automatically. +- `wal_delta`: _(auto recovery default)_ transfer by resolving [WAL] difference; the operations that were missed. -Each has pros, cons and specific requirements, which are: +Each has pros, cons and specific requirements, some of which are: -| Method: | Stream records | Snapshot | -|:---|:---|:---| -| **Connection** |
  • Requires internal gRPC API (port 6335)
|
  • Requires internal gRPC API (port 6335)
  • Requires REST API (port 6333)
| -| **HNSW index** |
  • Doesn't transfer index
  • Will reindex on target node
|
  • Index is transferred with a snapshot
  • Immediately ready on target node
| -| **Quantization** |
  • Doesn't transfer quantized data
  • Will re-quantize on target node
|
  • Quantized data is transferred with a snapshot
  • Immediately ready on target node
| -| **Ordering** |
  • Unordered updates on target node[^unordered]
|
  • Ordered updates on target node[^ordered]
| -| **Disk space** |
  • No extra disk space required
|
  • Extra disk space required for snapshot on both nodes
| +| Method: | Stream records | Snapshot | WAL delta | +|:---|:---|:---|:---| +| **Version** | v0.8.0+ | v1.7.0+ | v1.8.0+ | +| **Target** | New/existing shard | New/existing shard | Existing shard | +| **Connectivity** | Internal gRPC API (6335) | REST API (6333)
Internal gRPC API (6335) | Internal gRPC API (6335) | +| **HNSW index** | Doesn't transfer, will reindex on target. | Does transfer, immediately ready on target. | Doesn't transfer, may index on target. | +| **Quantization** | Doesn't transfer, will requantize on target. | Does transfer, immediately ready on target. | Doesn't transfer, may quantize on target. | +| **Ordering** | Unordered updates on target[^unordered] | Ordered updates on target[^ordered] | Ordered updates on target[^ordered] | +| **Disk space** | No extra required | Extra required for snapshot on both nodes | No extra required | [^unordered]: Weak ordering for updates: All records are streamed to the target node in order. New updates are received on the target node in parallel, while the transfer @@ -632,8 +635,23 @@ degradation in performance at the end of the transfer. Especially on large shards, this can give a huge performance improvement. 2. The ordering guarantees can be `strong`[^ordered], required for some applications. +The `wal_delta` transfer method only transfers the difference between two +shards. More specifically, it transfers all operations that were missed to the +target shard. The [WAL] of both shards is used to resolve this. There are two +benefits: 1. It will be very fast because it only transfers the difference +rather than all data. 2. The ordering guarantees can be `strong`[^ordered], +required for some applications. Two disadvantages are: 1. It can only be used to +transfer to a shard that already exists on the other node. 2. Applicability is +limited because the WALs normally don't hold more than 64MB of recent +operations. But that should be enough for a node that quickly restarts, to +upgrade for example. If a delta cannot be resolved, this method automatically +falls back to `stream_records` which equals transferring the full shard. + The `stream_records` method is currently used as default. This may change in the -future. +future. As of Qdrant 1.9.0 `wal_delta` is used for automatic shard replications +to recover dead shards. + +[WAL]: ../../concepts/storage/#versioning ## Replication @@ -1173,6 +1191,8 @@ sequentially. - `medium` ordering serializes all write operations through a dynamically elected leader, which might cause minor inconsistencies in case of leader change. - `strong` ordering serializes all write operations through the permanent leader, which provides strong consistency, but write operations may be unavailable if the leader is down. + + ```http PUT /collections/{collection_name}/points?ordering=strong {