Edge docs: add partial snapshots (#2087)

* Qdrant Edge partial snapshot docs

* Restructure and edits

* Break data syncing into two pages

* Add diagrams

* Broken links
This commit is contained in:
Abdon Pijpelink
2026-01-25 15:25:32 +05:30
committed by GitHub
parent 896a69b099
commit ae2c6cdd8b
6 changed files with 263 additions and 46 deletions
@@ -10,12 +10,14 @@ partition: qdrant
Qdrant Edge is a lightweight, embedded vector search engine for AI on devices like robots, kiosks, home assistants, and mobile phones. Designed for real-time vector search on edge devices with limited computational resources, Qdrant Edge allows applications to use Qdrant's functionality even with intermittent or no internet connectivity.
Qdrant Edge does not run as a separate process. Instead, it runs inside an application process. Data is stored and queried locally on the device, ensuring low-latency access and enhanced privacy since data does not need to be transmitted to an external server. That said, Qdrant Edge provides APIs to [synchronize data with a Qdrant server](/documentation/edge/edge-synchronizing-with-the-cloud/). This enables you to offload heavy computations such as indexing to more powerful server instances, back up and restore data, and centrally aggregate data from multiple edge devices.
Qdrant Edge does not run as a separate process. Instead, it runs inside an application process. Data is stored and queried locally on the device, ensuring low-latency access and enhanced privacy since data does not need to be transmitted to an external server. That said, Qdrant Edge provides APIs to [synchronize data with a Qdrant server](/documentation/edge/edge-data-synchronization-patterns/). This enables you to offload heavy computations such as indexing to more powerful server instances, back up and restore data, and centrally aggregate data from multiple edge devices.
## Qdrant Edge Shard
Qdrant Edge is built around the concept of an **Edge Shard**: a self-contained storage unit that can operate independently on edge devices. Each Edge Shard manages its own data, including vector and payload storage, and can perform local search and retrieval operations.
![Qdrant Edge Shards operate on edge devices](/documentation/edge/qdrant-edge.png)
To work with a Qdrant Edge Shard from a Python application, use the [Python Bindings for Qdrant Edge](https://pypi.org/project/qdrant-edge-py/) package. This package provides an `EdgeShard` class with methods to manage data, query it, and restore snapshots:
- `update`: Updates the data.
@@ -1,79 +1,84 @@
---
title: "Synchronizing with a Server"
title: "Data Synchronization Patterns"
weight: 20
---
# Synchronizing Qdrant Edge with a Server
# Data Synchronization Patterns
A Qdrant Edge Shard can be synchronized with a collection from an external Qdrant server to support use cases like:
This page describes patterns for synchronizing data between Qdrant Edge Shards and Qdrant server collections. For a practical end-to-end guide on implementing these patterns, refer to the [Qdrant Edge Synchronization Guide](/documentation/edge/edge-synchronization-guide/).
- **Offload indexing**: Indexing is a computationally expensive operation. By synchronizing an Edge Shard with a server collection, you can offload the indexing process to a more powerful server instance. The indexed data can then be synchronized back to the Edge Shard.
- **Back up and Restore**: Regularly back up your Edge Shard data to a central Qdrant instance to prevent data loss. In case of hardware failure or data corruption on the edge device, you can restore the data from the central instance.
- **Data Aggregation**: Collect data from multiple Edge Shards deployed in different locations and aggregate it into a central Qdrant instance for comprehensive analysis and reporting.
- **Synchronization between devices**: Keep data consistent across multiple edge devices by synchronizing their Edge Shards with a central Qdrant instance.
## Initialize Edge Shard from Existing Qdrant Collection
For an example implementation of the patterns described in this guide, refer to the [Qdrant Edge Demo GitHub repository](https://github.com/qdrant/qdrant-edge-demo).
Instead of starting with an empty Edge Shard, you may want to initialize it with pre-existing data from a collection on a Qdrant server. You can achieve this by restoring a snapshot of a shard in the server-side collection.
## Initialize Edge Shard from existing Qdrant Collection
![Qdrant Edge Shards can be initialized from snapshots of server-side shards](/documentation/edge/qdrant-edge-restore-snapshot.png)
First, create a snapshot on the server:
When creating a snapshot for synchronization, specify the applicable server-side shard ID in the snapshot URL. This allows for a single collection to serve multiple independent users or devices, each with its own Edge Shard. Read more about Qdrant's sharding strategy in the [Tiered Multitenancy Documentation](/documentation/guides/multitenancy/#tiered-multitenancy).
First, craft a snapshot URL:
```python
import requests
COLLECTION_NAME="edge-collection"
snapshot_url = f"{QDRANT_URL}/collections/{COLLECTION_NAME}/shards/0/snapshot"
```
Note, that Qdrant Edge operates on a single shard. Therefore, when creating a snapshot for synchronization, specify shard `0` in the snapshot URL (assuming the collection has a single shard).
Note that this example uses shard ID `0`.
This allows single collection to serve multiple independent users or devices, each with its own Edge Shard. Read more about qdrant sharding strategy in the [Tiered Multitenancy Documentation](/documentation/guides/multitenancy/#tiered-multitenancy).
Using the snapshot URL, you can download the snapshot, as shown in this helper function:
Using the snapshot URL, you can download the snapshot to the local disk and use its data to initialize a new Edge Shard.
```python
def download_snapshot(url: str, target_path: Path):
with requests.get(url, headers={"api-key": QDRANT_API_KEY}, stream=True) as r:
r.raise_for_status()
with open(target_path, "wb") as f:
for chunk in r.iter_content(chunk_size=8192):
f.write(chunk)
```
Finally, you can use this function to download the snapshot to the local disk and use the snapshot's data to initialize a new Edge Shard:
```python
import tempfile
from pathlib import Path
from qdrant_edge import EdgeShard
import requests
import shutil
import tempfile
STORAGE_DIRECTORY = "./qdrant-edge-directory"
data_dir = Path(STORAGE_DIRECTORY)
SHARD_DIRECTORY = "./qdrant-edge-directory"
data_dir = Path(SHARD_DIRECTORY)
with tempfile.TemporaryDirectory(dir=data_dir.parent) as restore_dir:
snapshot_path = Path(restore_dir) / "shard.snapshot"
download_snapshot(snapshot_url, snapshot_path)
with requests.get(snapshot_url, headers={"api-key": QDRANT_API_KEY}, stream=True) as r:
r.raise_for_status()
with open(snapshot_path, "wb") as f:
for chunk in r.iter_content(chunk_size=8192):
f.write(chunk)
edge_shard = None
if data_dir.exists():
shutil.rmtree(data_dir)
data_dir.mkdir(parents=True, exist_ok=True)
EdgeShard.unpack_snapshot(str(snapshot_path), str(data_dir))
edge_shard = EdgeShard(str(data_dir))
edge_shard = EdgeShard(SHARD_DIRECTORY)
```
This code first downloads the snapshot to a temporary directory. Next, the current instance of `EdgeShard` (if it existed) is destroyed by setting it to `None` and deleting its data directory. Finally, `EdgeShard.unpack_snapshot` unpacks the downloaded snapshot into the data directory, and a new instance of `EdgeShard` is created using the unpacked snapshot's data and configuration.
This code first downloads the snapshot to a temporary directory. Next, `EdgeShard.unpack_snapshot` unpacks the downloaded snapshot into the data directory, and a new instance of `EdgeShard` is created using the unpacked snapshot's data and configuration.
While restoring a snapshot, you may want to pause and buffer any ongoing data updates on the Edge Shard. Before taking the snapshot, ensure all queued data has been written to the server. After the restoration is complete, you can resume normal operations. Refer to the [Qdrant Edge Demo GitHub repository](https://github.com/qdrant/qdrant-edge-demo) for an example implementation.
The `edge_shard` will use the same configuration and the same file structure as the source collection from which the snapshot was created, including vector and payload indexes.
The `edge_shard` will use same configuration and same file structure as the source collection from which the snapshot was created, including vector and payload indexes.
## Update Qdrant Edge with Server-Side Changes
To keep an Edge Shard updated with new data from a server collection, you can periodically download and apply a snapshot. Restoring a full snapshot every time would create unnecessary overhead. Instead, you can use partial snapshots to restore changes since the last snapshot. A partial snapshot contains only those segments that have changed, based on the Edge Shard's manifest that describes all its segments and metadata. The `EdgeShard` class provides an `update_from_snapshot` method to update an Edge Shard from a partial snapshot.
<!-- ToDO -->
<!-- ## Synchronize Server-Side Changes with an Edge Shard -->
<!-- Talk about partial snapshots here -->
```Python
manifest = edge_shard.snapshot_manifest()
url = f"{QDRANT_URL}/collections/{COLLECTION_NAME}/shards/0/snapshot/partial/create"
with tempfile.TemporaryDirectory(dir=data_dir) as temp_dir:
partial_snapshot_path = Path(temp_dir) / "partial.snapshot"
response = requests.post(url, headers={"api-key": QDRANT_API_KEY}, json=manifest, stream=True)
response.raise_for_status()
with open(partial_snapshot_path, "wb") as f:
for chunk in response.iter_content(chunk_size=8192):
f.write(chunk)
edge_shard.update_from_snapshot(str(partial_snapshot_path))
```
## Update a Server Collection from an Edge Shard
@@ -83,7 +88,7 @@ Instead of writing to the server collection directly, you may want to set up a b
First, initialize:
- an Edge Shard from scratch or from server-side snapshot
- Qdrant server connection.
- a Qdrant server connection.
<details>
<summary>Details</summary>
@@ -185,8 +190,3 @@ if points_to_upload:
```
Make sure to properly handle errors and retries in case of network issues or server unavailability.
## Support
For explicit support in implementing Qdrant Edge in your project, please contact [Qdrant Sales](https://qdrant.tech/contact-us/).
@@ -0,0 +1,215 @@
---
title: "Synchronize with a Server"
weight: 30
---
# Synchronize Qdrant Edge with a Server
Qdrant Edge can be synchronized with a collection from an external Qdrant server to support use cases like:
- **Offload indexing**: Indexing is a computationally expensive operation. By synchronizing an Edge Shard with a server collection, you can offload the indexing process to a more powerful server instance. The indexed data can then be synchronized back to the Edge Shard.
- **Back up and Restore**: Regularly back up your Edge Shard data to a central Qdrant instance to prevent data loss. In case of hardware failure or data corruption on the edge device, you can restore the data from the central instance.
- **Data Aggregation**: Collect data from multiple Edge Shards deployed in different locations and aggregate it into a central Qdrant instance for comprehensive analysis and reporting.
- **Synchronization between devices**: Keep data consistent across multiple edge devices by synchronizing their Edge Shards with a central Qdrant instance.
## Synchronizing Qdrant Edge with a Server
To support having local updates from the device as well as updates from a server, you can implement a setup with two Edge Shards:
- A **mutable** Edge Shard that handles local data updates.
- An **immutable** Edge Shard that mirrors a shard from a collection on a server using partial snapshots.
When querying data, merge results from both Edge Shards to provide a unified view. This way, new points added on the device are available for search alongside the data synchronized from the server.
![Qdrant Edge Shards can be synchronized with a central server](/documentation/edge/qdrant-edge-sync-with-server.png)
Implementing a dual-write mechanism that writes data to both the mutable Edge Shard and the server collection ensures that data is indexed on the server and synchronized back to the immutable Edge Shard, benefitting search performance.
For an example implementation of the patterns described in this guide, refer to the [Qdrant Edge Demo GitHub repository](https://github.com/qdrant/qdrant-edge-demo).
### 1. Initialize a Mutable Edge Shard
The mutable Edge Shard will manage local data updates. It can be initialized from scratch, as detailed in the [Qdrant Edge Quickstart Guide](/documentation/edge/edge-quickstart/).
```python
from pathlib import Path
from qdrant_edge import (
Distance,
EdgeConfig,
EdgeShard,
VectorDataConfig,
)
MUTABLE_SHARD_DIR = "./qdrant-edge-directory/mutable"
Path(MUTABLE_SHARD_DIR).mkdir(parents=True, exist_ok=True)
VECTOR_NAME="my-vector"
VECTOR_DIMENSION=4
config = EdgeConfig(
vector_data={
VECTOR_NAME: VectorDataConfig(
size=VECTOR_DIMENSION,
distance=Distance.Cosine,
)
}
)
mutable_shard = EdgeShard(MUTABLE_SHARD_DIR, config)
```
### 2. Initialize an Immutable Edge Shard from a Server Snapshot
Next, create the immutable Edge Shard from a snapshot on the server, as outlined in [Initialize Edge Shard from existing Qdrant Collection](/documentation/edge/edge-data-synchronization-patterns/#initialize-edge-shard-from-existing-qdrant-collection):
```python
import requests
import tempfile
import shutil
COLLECTION_NAME="edge-collection"
snapshot_url = f"{QDRANT_URL}/collections/{COLLECTION_NAME}/shards/0/snapshot"
IMMUTABLE_SHARD_DIR = "./qdrant-edge-directory/mutable"
data_dir = Path(IMMUTABLE_SHARD_DIR)
with tempfile.TemporaryDirectory(dir=data_dir.parent) as restore_dir:
snapshot_path = Path(restore_dir) / "shard.snapshot"
with requests.get(snapshot_url, headers={"api-key": QDRANT_API_KEY}, stream=True) as r:
r.raise_for_status()
with open(snapshot_path, "wb") as f:
for chunk in r.iter_content(chunk_size=8192):
f.write(chunk)
immutable_shard = None
if data_dir.exists():
shutil.rmtree(data_dir)
data_dir.mkdir(parents=True, exist_ok=True)
EdgeShard.unpack_snapshot(str(snapshot_path), str(data_dir))
immutable_shard = EdgeShard(IMMUTABLE_SHARD_DIR)
```
### 3. Implement a Dual-Write Mechanism
With both Edge Shards initialized, you can implement a dual-write mechanism in your application as outlined in [Update a Server Collection from an Edge Shard](/documentation/edge/edge-data-synchronization-patterns/#update-a-server-collection-from-an-edge-shard). When adding or updating a point, write it to the mutable Edge Shard and enqueue it for writing to the server collection.
```python
from qdrant_edge import ( Point, UpdateOperation )
from qdrant_client import models
import time
SYNC_TIMESTAMP_KEY="timestamp"
id=2
vector=[0.4, 0.3, 0.2, 0.1]
payload={
"color": "green",
SYNC_TIMESTAMP_KEY: time.time()
}
point = Point(
id=id,
vector={VECTOR_NAME: vector},
payload=payload
)
mutable_shard.update(UpdateOperation.upsert_points([point]))
rest_point = models.PointStruct(id=id, vector={VECTOR_NAME: vector}, payload=payload)
upload_queue.put(rest_point)
```
Each point's payload should include a timestamp field (`SYNC_TIMESTAMP_KEY` in this example) that records when the point was upserted. This timestamp is used to deduplicate data when the immutable Edge Shard is synchronized with the server.
### 4. Periodically Update the Immutable Edge Shard
You can periodically update the immutable Edge Shard with changes from the server using partial snapshots, as described in [Update Qdrant Edge with Server-Side Changes](/documentation/edge/edge-data-synchronization-patterns/#update-qdrant-edge-with-server-side-changes).
While restoring a snapshot, you may want to pause and buffer any ongoing data updates on the mutable Edge Shard. Before taking the snapshot, ensure all queued data has been written to the server. After the restoration is complete, you can resume normal operations. Refer to the [Qdrant Edge Demo GitHub repository](https://github.com/qdrant/qdrant-edge-demo) for an example implementation.
```python
import time
manifest = immutable_shard.snapshot_manifest()
url = f"{QDRANT_URL}/collections/{COLLECTION_NAME}/shards/0/snapshot/partial/create"
sync_timestamp = time.time()
with tempfile.TemporaryDirectory(dir=data_dir) as temp_dir:
partial_snapshot_path = Path(temp_dir) / "partial.snapshot"
response = requests.post(url, headers={"api-key": QDRANT_API_KEY}, json=manifest, stream=True)
response.raise_for_status()
with open(partial_snapshot_path, "wb") as f:
for chunk in response.iter_content(chunk_size=8192):
f.write(chunk)
immutable_shard.update_from_snapshot(str(partial_snapshot_path))
```
This example records a `sync_timestamp` at the time of creating the partial snapshot. All points that were added to the mutable Edge Shard before this timestamp are now restored to the immutable Edge Shard. These duplicate points can now be deleted from the mutable Edge Shard:
```python
from qdrant_edge import (
Filter,
FieldCondition,
RangeFloat
)
mutable_shard.update(
UpdateOperation.delete_points_by_filter(Filter(
must=[
FieldCondition(
key=SYNC_TIMESTAMP_KEY, range=RangeFloat(lte=sync_timestamp)
)
])
)
```
### 5. Query Both Edge Shards
To provide a unified search experience across all data, query both the mutable and immutable Edge Shards and merge the two result sets. Since a point may exist in both Edge Shards, deduplicate the results based on point ID.
```python
from qdrant_edge import Query, QueryRequest
query_request = QueryRequest(
query=Query.Nearest([0.2, 0.1, 0.9, 0.7], using=VECTOR_NAME),
limit=10,
with_vector=False,
with_payload=True
)
mutable_results = mutable_shard.query(query_request)
immutable_results = immutable_shard.query(query_request)
all_results = list(mutable_results) + list(immutable_results)
all_results.sort(key=lambda x: x.score, reverse=True)
seen_ids = set()
unique_results = []
for result in all_results:
if result.id not in seen_ids:
seen_ids.add(result.id)
unique_results.append(result)
results= [
{
"id": result.id,
"score": result.score,
"payload": result.payload
}
for result in unique_results[:10]
]
```
## Support
For explicit support in implementing Qdrant Edge in your project, please contact [Qdrant Sales](https://qdrant.tech/contact-us/).
Binary file not shown.

After

Width:  |  Height:  |  Size: 25 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 39 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 16 KiB