mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-04 10:28:29 +02:00
Initial draft of tutorials restructure
This commit is contained in:
@@ -0,0 +1,22 @@
|
||||
---
|
||||
title: Operations & Scale
|
||||
weight: 21
|
||||
is_empty: false
|
||||
aliases:
|
||||
- how-to
|
||||
- tutorials
|
||||
partition: qdrant
|
||||
---
|
||||
|
||||
# Operations & Scale Tutorials
|
||||
*Production-grade management, monitoring, and high-volume optimization.*
|
||||
|
||||
| Tutorial | Objective | Stack | Time | Level |
|
||||
| :--- | :--- | :--- | :--- | :--- |
|
||||
| [Bulk Data Uploads](https://qdrant.tech/documentation/tutorials-operations/bulk-upload/) | High-scale ingestion tricks for power users. | Python | 20m | Intermediate |
|
||||
| [Snapshot & Backup](https://qdrant.tech/documentation/tutorials-operations/create-snapshot/) | Create and restore collection snapshots. | Python | 20m | Beginner |
|
||||
| [Billion-Scale Search](https://qdrant.tech/documentation/tutorials-operations/large-scale-search/) | Cost-efficient search for LAION-400M datasets. | None | 2 days | Advanced |
|
||||
| [Python Async API](https://qdrant.tech/documentation/tutorials-operations/async-api/) | Use Asynchronous programming for efficiency. | Python | 25m | Intermediate |
|
||||
| [Cloud Inference Search](https://qdrant.tech/documentation/tutorials-and-examples/cloud-inference-hybrid-search/) | Hybrid search using Qdrant's built-in inference. | Any | 20m | Beginner |
|
||||
| [Monitor Managed Cloud](https://qdrant.tech/documentation/tutorials-and-examples/managed-cloud-prometheus/) | Observability with Prometheus and Grafana. | Prometheus | 30m | Intermediate |
|
||||
| [Monitor Private Cloud](https://qdrant.tech/documentation/tutorials-and-examples/hybrid-cloud-prometheus/) | Observability for hybrid/private cloud setups. | Prometheus | 30m | Intermediate |
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
title: Build With Async API
|
||||
aliases:
|
||||
- /documentation/tutorials/async-api/
|
||||
weight: 4
|
||||
---
|
||||
|
||||
# Using Qdrant’s Async API for Efficient Python Applications
|
||||
|
||||
Asynchronous programming is being broadly adopted in the Python ecosystem. Tools such as FastAPI [have embraced this new
|
||||
paradigm](https://fastapi.tiangolo.com/async/), but it is also becoming a standard for ML models served as SaaS. For example, the Cohere SDK
|
||||
[provides an async client](https://github.com/cohere-ai/cohere-python/blob/856a4c3bd29e7a75fa66154b8ac9fcdf1e0745e0/src/cohere/client.py#L189) next to its synchronous counterpart.
|
||||
|
||||
Databases are often launched as separate services and are accessed via a network. All the interactions with them are IO-bound and can
|
||||
be performed asynchronously so as not to waste time actively waiting for a server response. In Python, this is achieved by
|
||||
using [`async/await`](https://docs.python.org/3/library/asyncio-task.html) syntax. That lets the interpreter switch to another task
|
||||
while waiting for a response from the server.
|
||||
|
||||
## When to use async API
|
||||
|
||||
There is no need to use async API if the application you are writing will never support multiple users at once (e.g it is a script that runs once per day). However, if you are writing a web service that multiple users will use simultaneously, you shouldn't be
|
||||
blocking the threads of the web server as it limits the number of concurrent requests it can handle. In this case, you should use
|
||||
the async API.
|
||||
|
||||
Modern web frameworks like [FastAPI](https://fastapi.tiangolo.com/) and [Quart](https://quart.palletsprojects.com/en/latest/) support
|
||||
async API out of the box. Mixing asynchronous code with an existing synchronous codebase might be a challenge. The `async/await` syntax
|
||||
cannot be used in synchronous functions. On the other hand, calling an IO-bound operation synchronously in async code is considered
|
||||
an antipattern. Therefore, if you build an async web service, exposed through an [ASGI](https://asgi.readthedocs.io/en/latest/) server,
|
||||
you should use the async API for all the interactions with Qdrant.
|
||||
|
||||
<aside role="status">
|
||||
All the async code has to be launched in an async context. Usually, it means you have to use <code>asyncio.run</code> or <code>asyncio.create_task</code> to run them.
|
||||
Please refer to the <a href="https://docs.python.org/3/library/asyncio.html">asyncio documentation</a> for more details.
|
||||
</aside>
|
||||
|
||||
### Using Qdrant asynchronously
|
||||
|
||||
The simplest way of running asynchronous code is to use define `async` function and use the `asyncio.run` in the following way to run it:
|
||||
|
||||
```python
|
||||
from qdrant_client import models
|
||||
|
||||
import qdrant_client
|
||||
import asyncio
|
||||
|
||||
|
||||
async def main():
|
||||
client = qdrant_client.AsyncQdrantClient("localhost")
|
||||
|
||||
# Create a collection
|
||||
await client.create_collection(
|
||||
collection_name="my_collection",
|
||||
vectors_config=models.VectorParams(size=4, distance=models.Distance.COSINE),
|
||||
)
|
||||
|
||||
# Insert a vector
|
||||
await client.upsert(
|
||||
collection_name="my_collection",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id="5c56c793-69f3-4fbf-87e6-c4bf54c28c26",
|
||||
payload={
|
||||
"color": "red",
|
||||
},
|
||||
vector=[0.9, 0.1, 0.1, 0.5],
|
||||
),
|
||||
],
|
||||
)
|
||||
|
||||
# Search for nearest neighbors
|
||||
points = await client.query_points(
|
||||
collection_name="my_collection",
|
||||
query=[0.9, 0.1, 0.1, 0.5],
|
||||
limit=2,
|
||||
).points
|
||||
|
||||
# Your async code using AsyncQdrantClient might be put here
|
||||
# ...
|
||||
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
The `AsyncQdrantClient` provides the same methods as the synchronous counterpart `QdrantClient`. If you already have a synchronous
|
||||
codebase, switching to async API is as simple as replacing `QdrantClient` with `AsyncQdrantClient` and adding `await` before each
|
||||
method call.
|
||||
|
||||
<aside role="status">
|
||||
Asynchronous client was introduced in <code>qdrant-client</code> version 1.6.1. If you are using an older version, you need to use autogenerated async clients directly.
|
||||
</aside>
|
||||
@@ -0,0 +1,647 @@
|
||||
---
|
||||
title: Bulk Upload Vectors
|
||||
aliases:
|
||||
- /documentation/tutorials/bulk-upload/
|
||||
weight: 1
|
||||
---
|
||||
|
||||
# Bulk Upload Vectors to a Qdrant Collection
|
||||
|
||||
Uploading a large-scale dataset fast might be a challenge, but Qdrant has a few tricks to help you with that.
|
||||
|
||||
The first important detail about data uploading is that the bottleneck is usually located on the client side, not on the server side.
|
||||
This means that if you are uploading a large dataset, you should prefer a high-performance client library.
|
||||
|
||||
We recommend using our [Rust client library](https://github.com/qdrant/rust-client) for this purpose, as it is the fastest client library available for Qdrant.
|
||||
|
||||
If you are not using Rust, you might want to consider parallelizing your upload process.
|
||||
|
||||
## Choose an Indexing Strategy
|
||||
|
||||
Qdrant incrementally builds an HNSW index for dense vectors as new data arrives. This ensures fast search, but indexing is memory- and CPU-intensive. During bulk ingestion, frequent index updates can reduce throughput and increase resource usage.
|
||||
|
||||
To control this behavior and optimize for your system’s limits, adjust the following parameters:
|
||||
|
||||
| Your Goal | What to Do | Configuration |
|
||||
|-------------------------------------------|-------------------------------------------------|----------------------------------------------------|
|
||||
| Fastest upload, tolerate high RAM usage | Disable indexing completely | `indexing_threshold: 0` |
|
||||
| Low memory usage during upload | Defer HNSW graph construction (recommended) | `m: 0` |
|
||||
| Faster index availability after upload | Keep indexing enabled (default behavior) | `m: 16`, `indexing_threshold: 20000` *(default)* |
|
||||
|
||||
Indexing must be re-enabled after upload to activate fast HNSW search if it was disabled during ingestion.
|
||||
|
||||
|
||||
### Defer HNSW graph construction (`m: 0`)
|
||||
|
||||
For dense vectors, setting the HNSW `m` parameter to `0` disables index building entirely. Vectors will still be stored, but not indexed until you enable indexing later.
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"vectors": {
|
||||
"size": 768,
|
||||
"distance": "Cosine"
|
||||
},
|
||||
"hnsw_config": {
|
||||
"m": 0
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config=models.VectorParams(size=768, distance=models.Distance.COSINE),
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=0,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.createCollection("{collection_name}", {
|
||||
vectors: {
|
||||
size: 768,
|
||||
distance: "Cosine",
|
||||
},
|
||||
hnsw_config: {
|
||||
m: 0,
|
||||
},
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::qdrant::{
|
||||
CreateCollectionBuilder, Distance, HnswConfigDiffBuilder, VectorParamsBuilder,
|
||||
};
|
||||
use qdrant_client::Qdrant;
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client
|
||||
.create_collection(
|
||||
CreateCollectionBuilder::new("{collection_name}")
|
||||
.vectors_config(VectorParamsBuilder::new(768, Distance::Cosine))
|
||||
.hnsw_config(HnswConfigDiffBuilder::default().m(0)),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||
import io.qdrant.client.grpc.Collections.Distance;
|
||||
import io.qdrant.client.grpc.Collections.HnswConfigDiff;
|
||||
import io.qdrant.client.grpc.Collections.VectorParams;
|
||||
import io.qdrant.client.grpc.Collections.VectorsConfig;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.createCollectionAsync(
|
||||
CreateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setVectorsConfig(
|
||||
VectorsConfig.newBuilder()
|
||||
.setParams(
|
||||
VectorParams.newBuilder()
|
||||
.setSize(768)
|
||||
.setDistance(Distance.Cosine)
|
||||
.build())
|
||||
.build())
|
||||
.setHnswConfig(HnswConfigDiff.newBuilder().setM(0).build())
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
vectorsConfig: new VectorParams { Size = 768, Distance = Distance.Cosine },
|
||||
hnswConfig: new HnswConfigDiff { M = 0 }
|
||||
);
|
||||
```
|
||||
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "localhost",
|
||||
Port: 6334,
|
||||
})
|
||||
|
||||
client.CreateCollection(context.Background(), &qdrant.CreateCollection{
|
||||
CollectionName: "{collection_name}",
|
||||
VectorsConfig: qdrant.NewVectorsConfig(&qdrant.VectorParams{
|
||||
Size: 768,
|
||||
Distance: qdrant.Distance_Cosine,
|
||||
}),
|
||||
HnswConfig: &qdrant.HnswConfigDiff{
|
||||
M: qdrant.PtrOf(uint64(0)),
|
||||
},
|
||||
})
|
||||
```
|
||||
|
||||
Once ingestion is complete, re-enable HNSW by setting `m` to your production value (usually 16 or 32).
|
||||
|
||||
```http
|
||||
PATCH /collections/{collection_name}
|
||||
{
|
||||
"vectors": {
|
||||
"size": 768,
|
||||
"distance": "Cosine"
|
||||
},
|
||||
"hnsw_config": {
|
||||
"m": 16
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.update_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config=models.VectorParams(size=768, distance=models.Distance.COSINE),
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=16,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.updateCollection("{collection_name}", {
|
||||
vectors: {
|
||||
size: 768,
|
||||
distance: "Cosine",
|
||||
},
|
||||
hnsw_config: {
|
||||
m: 16,
|
||||
},
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::qdrant::{
|
||||
UpdateCollectionBuilder, HnswConfigDiffBuilder,
|
||||
};
|
||||
use qdrant_client::Qdrant;
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client
|
||||
.update_collection(
|
||||
UpdateCollectionBuilder::new("{collection_name}")
|
||||
.hnsw_config(HnswConfigDiffBuilder::default().m(16)),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import io.qdrant.client.grpc.Collections.UpdateCollection;
|
||||
import io.qdrant.client.grpc.Collections.HnswConfigDiff;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client.updateCollectionAsync(
|
||||
UpdateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setHnswConfig(HnswConfigDiff.newBuilder().setM(16).build())
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.UpdateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
hnswConfig: new HnswConfigDiff { M = 16 }
|
||||
);
|
||||
```
|
||||
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
qdrant.NewClient(&qdrant.Config{
|
||||
Host: "localhost",
|
||||
Port: 6334,
|
||||
})
|
||||
|
||||
client, err := client.UpdateCollection(context.Background(), &qdrant.UpdateCollection{
|
||||
CollectionName: "{collection_name}",
|
||||
HnswConfig: &qdrant.HnswConfigDiff{
|
||||
M: qdrant.PtrOf(uint64(16)),
|
||||
},
|
||||
})
|
||||
```
|
||||
|
||||
### Disable indexing completely (`indexing_threshold: 0`)
|
||||
|
||||
In case you are doing an initial upload of a large dataset, you might want to disable indexing during upload. It will enable to avoid unnecessary indexing of vectors, which will be overwritten by the next batch.
|
||||
|
||||
Setting `indexing_threshold` to `0` disables indexing altogether:
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"vectors": {
|
||||
"size": 768,
|
||||
"distance": "Cosine"
|
||||
},
|
||||
"optimizers_config": {
|
||||
"indexing_threshold": 0
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config=models.VectorParams(size=768, distance=models.Distance.COSINE),
|
||||
optimizers_config=models.OptimizersConfigDiff(
|
||||
indexing_threshold=0,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.createCollection("{collection_name}", {
|
||||
vectors: {
|
||||
size: 768,
|
||||
distance: "Cosine",
|
||||
},
|
||||
optimizers_config: {
|
||||
indexing_threshold: 0,
|
||||
},
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::qdrant::{
|
||||
OptimizersConfigDiffBuilder, UpdateCollectionBuilder,
|
||||
};
|
||||
use qdrant_client::Qdrant;
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client
|
||||
.create_collection(
|
||||
CreateCollectionBuilder::new("{collection_name}")
|
||||
.optimizers_config(OptimizersConfigDiffBuilder::default().indexing_threshold(0)),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||
import io.qdrant.client.grpc.Collections.Distance;
|
||||
import io.qdrant.client.grpc.Collections.VectorParams;
|
||||
import io.qdrant.client.grpc.Collections.VectorsConfig;
|
||||
import io.qdrant.client.grpc.Collections.OptimizersConfigDiff;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client.createCollectionAsync(
|
||||
CreateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setVectorsConfig(
|
||||
VectorsConfig.newBuilder()
|
||||
.setParams(
|
||||
VectorParams.newBuilder()
|
||||
.setSize(768)
|
||||
.setDistance(Distance.Cosine)
|
||||
.build())
|
||||
.build())
|
||||
.setOptimizersConfig(
|
||||
OptimizersConfigDiff.newBuilder()
|
||||
.setIndexingThreshold(0)
|
||||
.build())
|
||||
.build()
|
||||
).get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
vectorsConfig: new VectorParams { Size = 768, Distance = Distance.Cosine },
|
||||
optimizersConfig: new OptimizersConfigDiff { IndexingThreshold = 0 }
|
||||
);
|
||||
```
|
||||
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "localhost",
|
||||
Port: 6334,
|
||||
})
|
||||
|
||||
client.CreateCollection(context.Background(), &qdrant.CreateCollection{
|
||||
CollectionName: "{collection_name}",
|
||||
VectorsConfig: qdrant.NewVectorsConfig(&qdrant.VectorParams{
|
||||
Size: 768,
|
||||
Distance: qdrant.Distance_Cosine,
|
||||
}),
|
||||
OptimizersConfig: &qdrant.OptimizersConfigDiff{
|
||||
IndexingThreshold: qdrant.PtrOf(uint64(0)),
|
||||
},
|
||||
})
|
||||
```
|
||||
|
||||
<aside role="status">
|
||||
With indexing_threshold set to 0, storage won't be optimized properly, which can lead to high RAM usage as segments accumulate in memory.
|
||||
</aside>
|
||||
|
||||
After upload is done, you can enable indexing by setting `indexing_threshold` to a desired value (default is 20000):
|
||||
|
||||
```http
|
||||
PATCH /collections/{collection_name}
|
||||
{
|
||||
"optimizers_config": {
|
||||
"indexing_threshold": 20000
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.update_collection(
|
||||
collection_name="{collection_name}",
|
||||
optimizers_config=models.OptimizersConfigDiff(indexing_threshold=20000),
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.updateCollection("{collection_name}", {
|
||||
optimizers_config: {
|
||||
indexing_threshold: 20000,
|
||||
},
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::qdrant::{
|
||||
OptimizersConfigDiffBuilder, UpdateCollectionBuilder,
|
||||
};
|
||||
use qdrant_client::Qdrant;
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client
|
||||
.update_collection(
|
||||
UpdateCollectionBuilder::new("{collection_name}")
|
||||
.optimizers_config(OptimizersConfigDiffBuilder::default().indexing_threshold(20000)),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import io.qdrant.client.grpc.Collections.UpdateCollection;
|
||||
import io.qdrant.client.grpc.Collections.OptimizersConfigDiff;
|
||||
|
||||
client.updateCollectionAsync(
|
||||
UpdateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setOptimizersConfig(
|
||||
OptimizersConfigDiff.newBuilder()
|
||||
.setIndexingThreshold(20000)
|
||||
.build()
|
||||
)
|
||||
.build()
|
||||
).get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.UpdateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
optimizersConfig: new OptimizersConfigDiff { IndexingThreshold = 20000 }
|
||||
);
|
||||
```
|
||||
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "localhost",
|
||||
Port: 6334,
|
||||
})
|
||||
|
||||
client.UpdateCollection(context.Background(), &qdrant.UpdateCollection{
|
||||
CollectionName: "{collection_name}",
|
||||
OptimizersConfig: &qdrant.OptimizersConfigDiff{
|
||||
IndexingThreshold: qdrant.PtrOf(uint64(20000)),
|
||||
},
|
||||
})
|
||||
```
|
||||
|
||||
|
||||
|
||||
At this point, Qdrant will begin indexing new and previously unindexed segments in the background.
|
||||
|
||||
## Upload directly to disk
|
||||
|
||||
When the vectors you upload do not all fit in RAM, you likely want to use
|
||||
[memmap](/documentation/concepts/storage/#configuring-memmap-storage)
|
||||
support.
|
||||
|
||||
During collection
|
||||
[creation](/documentation/concepts/collections/#create-collection),
|
||||
memmaps may be enabled on a per-vector basis using the `on_disk` parameter. This
|
||||
will store vector data directly on disk at all times. It is suitable for
|
||||
ingesting a large amount of data, essential for the billion scale benchmark.
|
||||
|
||||
Using `memmap_threshold` is not recommended in this case. It would require
|
||||
the [optimizer](/documentation/concepts/optimizer/) to constantly
|
||||
transform in-memory segments into memmap segments on disk. This process is
|
||||
slower, and the optimizer can be a bottleneck when ingesting a large amount of
|
||||
data.
|
||||
|
||||
Read more about this in
|
||||
[Configuring Memmap Storage](/documentation/concepts/storage/#configuring-memmap-storage).
|
||||
|
||||
## Parallel upload into multiple shards
|
||||
|
||||
In Qdrant, each collection is split into shards. Each shard has a separate Write-Ahead-Log (WAL), which is responsible for ordering operations.
|
||||
By creating multiple shards, you can parallelize upload of a large dataset. From 2 to 4 shards per one machine is a reasonable number.
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"vectors": {
|
||||
"size": 768,
|
||||
"distance": "Cosine"
|
||||
},
|
||||
"shard_number": 2
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config=models.VectorParams(size=768, distance=models.Distance.COSINE),
|
||||
shard_number=2,
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.createCollection("{collection_name}", {
|
||||
vectors: {
|
||||
size: 768,
|
||||
distance: "Cosine",
|
||||
},
|
||||
shard_number: 2,
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::qdrant::{CreateCollectionBuilder, Distance, VectorParamsBuilder};
|
||||
use qdrant_client::Qdrant;
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client
|
||||
.create_collection(
|
||||
CreateCollectionBuilder::new("{collection_name}")
|
||||
.vectors_config(VectorParamsBuilder::new(768, Distance::Cosine))
|
||||
.shard_number(2),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Collections.CreateCollection;
|
||||
import io.qdrant.client.grpc.Collections.Distance;
|
||||
import io.qdrant.client.grpc.Collections.VectorParams;
|
||||
import io.qdrant.client.grpc.Collections.VectorsConfig;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.createCollectionAsync(
|
||||
CreateCollection.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setVectorsConfig(
|
||||
VectorsConfig.newBuilder()
|
||||
.setParams(
|
||||
VectorParams.newBuilder()
|
||||
.setSize(768)
|
||||
.setDistance(Distance.Cosine)
|
||||
.build())
|
||||
.build())
|
||||
.setShardNumber(2)
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreateCollectionAsync(
|
||||
collectionName: "{collection_name}",
|
||||
vectorsConfig: new VectorParams { Size = 768, Distance = Distance.Cosine },
|
||||
shardNumber: 2
|
||||
);
|
||||
```
|
||||
|
||||
```go
|
||||
import (
|
||||
"context"
|
||||
|
||||
"github.com/qdrant/go-client/qdrant"
|
||||
)
|
||||
|
||||
client, err := qdrant.NewClient(&qdrant.Config{
|
||||
Host: "localhost",
|
||||
Port: 6334,
|
||||
})
|
||||
|
||||
client.CreateCollection(context.Background(), &qdrant.CreateCollection{
|
||||
CollectionName: "{collection_name}",
|
||||
VectorsConfig: qdrant.NewVectorsConfig(&qdrant.VectorParams{
|
||||
Size: 768,
|
||||
Distance: qdrant.Distance_Cosine,
|
||||
}),
|
||||
ShardNumber: qdrant.PtrOf(uint32(2)),
|
||||
})
|
||||
```
|
||||
@@ -0,0 +1,285 @@
|
||||
---
|
||||
title: Create & Restore Snapshots
|
||||
aliases:
|
||||
- /documentation/tutorials/create-snapshot/
|
||||
weight: 2
|
||||
---
|
||||
|
||||
# Backup and Restore Qdrant Collections Using Snapshots
|
||||
|
||||
| Time: 20 min | Level: Beginner | | |
|
||||
|--------------|-----------------|--|----|
|
||||
|
||||
A collection is a basic unit of data storage in Qdrant. It contains vectors, their IDs, and payloads. However, keeping the search efficient requires additional data structures to be built on top of the data. Building these data structures may take a while, especially for large collections.
|
||||
That's why using snapshots is the best way to export and import Qdrant collections, as they contain all the bits and pieces required to restore the entire collection efficiently.
|
||||
|
||||
This tutorial will show you how to create a snapshot of a collection and restore it. Since working with snapshots in a distributed environment might be thought to be a bit more complex, we will use a 3-node Qdrant cluster. However, the same approach applies to a single-node setup.
|
||||
|
||||
<aside role="status">Snapshots cannot be created in local mode of Python SDK. You need to spin up a Qdrant Docker container or use Qdrant Cloud.</aside>
|
||||
|
||||
You can use the techniques described in this page to migrate a cluster. Follow the instructions
|
||||
in this tutorial to create and download snapshots. When you [Restore from snapshot](#restore-from-snapshot), restore your data to the new cluster.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Let's assume you already have a running Qdrant instance or a cluster. If not, you can follow the [installation guide](/documentation/guides/installation/) to set up a local Qdrant instance or use [Qdrant Cloud](https://cloud.qdrant.io/) to create a cluster in a few clicks.
|
||||
|
||||
Once the cluster is running, let's install the required dependencies:
|
||||
|
||||
```shell
|
||||
pip install qdrant-client datasets
|
||||
```
|
||||
|
||||
### Establish a connection to Qdrant
|
||||
|
||||
We are going to use the Python SDK and raw HTTP calls to interact with Qdrant. Since we are going to use a 3-node cluster, we need to know the URLs of all the nodes. For the simplicity, let's keep them all in constants, along with the API key, so we can refer to them later:
|
||||
|
||||
```python
|
||||
QDRANT_MAIN_URL = "https://my-cluster.com:6333"
|
||||
QDRANT_NODES = (
|
||||
"https://node-0.my-cluster.com:6333",
|
||||
"https://node-1.my-cluster.com:6333",
|
||||
"https://node-2.my-cluster.com:6333",
|
||||
)
|
||||
QDRANT_API_KEY = "my-api-key"
|
||||
```
|
||||
|
||||
<aside role="status">If you are using Qdrant Cloud, you can find the URL and API key in the <a href="https://cloud.qdrant.io/">Qdrant Cloud dashboard</a>.</aside>
|
||||
|
||||
We can now create a client instance:
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient(QDRANT_MAIN_URL, api_key=QDRANT_API_KEY)
|
||||
```
|
||||
|
||||
First of all, we are going to create a collection from a precomputed dataset. If you already have a collection, you can skip this step and start by [creating a snapshot](#create-and-download-snapshots).
|
||||
|
||||
<details>
|
||||
<summary>(Optional) Create collection and import data</summary>
|
||||
|
||||
### Load the dataset
|
||||
|
||||
We are going to use a dataset with precomputed embeddings, available on Hugging Face Hub. The dataset is called [Qdrant/arxiv-titles-instructorxl-embeddings](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings) and was created using the [InstructorXL](https://huggingface.co/hkunlp/instructor-xl) model. It contains 2.25M embeddings for the titles of the papers from the [arXiv](https://arxiv.org/) dataset.
|
||||
|
||||
Loading the dataset is as simple as:
|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
|
||||
dataset = load_dataset(
|
||||
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
|
||||
)
|
||||
```
|
||||
|
||||
We used the streaming mode, so the dataset is not loaded into memory. Instead, we can iterate through it and extract the id and vector embedding:
|
||||
|
||||
```python
|
||||
for payload in dataset:
|
||||
id_ = payload.pop("id")
|
||||
vector = payload.pop("vector")
|
||||
print(id_, vector, payload)
|
||||
```
|
||||
|
||||
A single payload looks like this:
|
||||
|
||||
```json
|
||||
{
|
||||
'title': 'Dynamics of partially localized brane systems',
|
||||
'DOI': '1109.1415'
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
### Create a collection
|
||||
|
||||
First things first, we need to create our collection. We're not going to play with the configuration of it, but it makes sense to do it right now.
|
||||
The configuration is also a part of the collection snapshot.
|
||||
|
||||
```python
|
||||
from qdrant_client import models
|
||||
|
||||
if not client.collection_exists("test_collection"):
|
||||
client.create_collection(
|
||||
collection_name="test_collection",
|
||||
vectors_config=models.VectorParams(
|
||||
size=768, # Size of the embedding vector generated by the InstructorXL model
|
||||
distance=models.Distance.COSINE
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
### Upload the dataset
|
||||
|
||||
Calculating the embeddings is usually a bottleneck of the vector search pipelines, but we are happy to have them in place already. Since the goal of this tutorial is to show how to create a snapshot, **we are going to upload only a small part of the dataset**.
|
||||
|
||||
```python
|
||||
ids, vectors, payloads = [], [], []
|
||||
for payload in dataset:
|
||||
id_ = payload.pop("id")
|
||||
vector = payload.pop("vector")
|
||||
|
||||
ids.append(id_)
|
||||
vectors.append(vector)
|
||||
payloads.append(payload)
|
||||
|
||||
# We are going to upload only 1000 vectors
|
||||
if len(ids) == 1000:
|
||||
break
|
||||
|
||||
client.upsert(
|
||||
collection_name="test_collection",
|
||||
points=models.Batch(
|
||||
ids=ids,
|
||||
vectors=vectors,
|
||||
payloads=payloads,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
Our collection is now ready to be used for search. Let's create a snapshot of it.
|
||||
|
||||
</details>
|
||||
|
||||
If you already have a collection, you can skip the previous step and start by [creating a snapshot](#create-and-download-snapshots).
|
||||
|
||||
## Create and download snapshots
|
||||
|
||||
Qdrant exposes an HTTP endpoint to request creating a snapshot, but we can also call it with the Python SDK.
|
||||
Our setup consists of 3 nodes, so we need to call the endpoint **on each of them** and create a snapshot on each node. While using Python SDK, that means creating a separate client instance for each node.
|
||||
|
||||
|
||||
<aside role="status">You may get a timeout error, if the collection size is big. You can trigger the snapshot process in the background, without awaiting for the result, by using <code>wait=false</code> parameter. You can always <a href="/documentation/concepts/snapshots/#list-snapshot">list all the snapshots through the API</a> later on.</aside>
|
||||
|
||||
|
||||
```python
|
||||
snapshot_urls = []
|
||||
for node_url in QDRANT_NODES:
|
||||
node_client = QdrantClient(node_url, api_key=QDRANT_API_KEY)
|
||||
snapshot_info = node_client.create_snapshot(collection_name="test_collection")
|
||||
|
||||
snapshot_url = f"{node_url}/collections/test_collection/snapshots/{snapshot_info.name}"
|
||||
snapshot_urls.append(snapshot_url)
|
||||
```
|
||||
|
||||
```http
|
||||
// for `https://node-0.my-cluster.com:6333`
|
||||
POST /collections/test_collection/snapshots
|
||||
|
||||
// for `https://node-1.my-cluster.com:6333`
|
||||
POST /collections/test_collection/snapshots
|
||||
|
||||
// for `https://node-2.my-cluster.com:6333`
|
||||
POST /collections/test_collection/snapshots
|
||||
```
|
||||
|
||||
<details>
|
||||
<summary>Response</summary>
|
||||
|
||||
```json
|
||||
{
|
||||
"result": {
|
||||
"name": "test_collection-559032209313046-2024-01-03-13-20-11.snapshot",
|
||||
"creation_time": "2024-01-03T13:20:11",
|
||||
"size": 18956800
|
||||
},
|
||||
"status": "ok",
|
||||
"time": 0.307644965
|
||||
}
|
||||
```
|
||||
</details>
|
||||
|
||||
|
||||
|
||||
Once we have the snapshot URLs, we can download them. Please make sure to include the API key in the request headers.
|
||||
Downloading the snapshot **can be done only through the HTTP API**, so we are going to use the `requests` library.
|
||||
|
||||
```python
|
||||
import requests
|
||||
import os
|
||||
|
||||
# Create a directory to store snapshots
|
||||
os.makedirs("snapshots", exist_ok=True)
|
||||
|
||||
local_snapshot_paths = []
|
||||
for snapshot_url in snapshot_urls:
|
||||
snapshot_name = os.path.basename(snapshot_url)
|
||||
local_snapshot_path = os.path.join("snapshots", snapshot_name)
|
||||
|
||||
response = requests.get(
|
||||
snapshot_url, headers={"api-key": QDRANT_API_KEY}
|
||||
)
|
||||
with open(local_snapshot_path, "wb") as f:
|
||||
response.raise_for_status()
|
||||
f.write(response.content)
|
||||
|
||||
local_snapshot_paths.append(local_snapshot_path)
|
||||
```
|
||||
|
||||
Alternatively, you can use the `wget` command:
|
||||
|
||||
```bash
|
||||
wget https://node-0.my-cluster.com:6333/collections/test_collection/snapshots/test_collection-559032209313046-2024-01-03-13-20-11.snapshot \
|
||||
--header="api-key: ${QDRANT_API_KEY}" \
|
||||
-O node-0-shapshot.snapshot
|
||||
|
||||
wget https://node-1.my-cluster.com:6333/collections/test_collection/snapshots/test_collection-559032209313047-2024-01-03-13-20-12.snapshot \
|
||||
--header="api-key: ${QDRANT_API_KEY}" \
|
||||
-O node-1-shapshot.snapshot
|
||||
|
||||
wget https://node-2.my-cluster.com:6333/collections/test_collection/snapshots/test_collection-559032209313048-2024-01-03-13-20-13.snapshot \
|
||||
--header="api-key: ${QDRANT_API_KEY}" \
|
||||
-O node-2-shapshot.snapshot
|
||||
```
|
||||
|
||||
The snapshots are now stored locally. We can use them to restore the collection to a different Qdrant instance, or treat them as a backup. We will create another collection using the same data on the same cluster.
|
||||
|
||||
## Restore from snapshot
|
||||
|
||||
Our brand-new snapshot is ready to be restored. Typically, it is used to move a collection to a different Qdrant instance, but we are going to use it to create a new collection on the same cluster.
|
||||
It is just going to have a different name, `test_collection_import`. We do not need to create a collection first, as it is going to be created automatically.
|
||||
|
||||
Restoring collection is also done separately on each node, but our Python SDK does not support it yet. We are going to use the HTTP API instead,
|
||||
and send a request to each node using `requests` library.
|
||||
|
||||
```python
|
||||
for node_url, snapshot_path in zip(QDRANT_NODES, local_snapshot_paths):
|
||||
snapshot_name = os.path.basename(snapshot_path)
|
||||
requests.post(
|
||||
f"{node_url}/collections/test_collection_import/snapshots/upload?priority=snapshot",
|
||||
headers={
|
||||
"api-key": QDRANT_API_KEY,
|
||||
},
|
||||
files={"snapshot": (snapshot_name, open(snapshot_path, "rb"))},
|
||||
)
|
||||
```
|
||||
|
||||
Alternatively, you can use the `curl` command:
|
||||
|
||||
```bash
|
||||
curl -X POST 'https://node-0.my-cluster.com:6333/collections/test_collection_import/snapshots/upload?priority=snapshot' \
|
||||
-H 'api-key: ${QDRANT_API_KEY}' \
|
||||
-H 'Content-Type:multipart/form-data' \
|
||||
-F 'snapshot=@node-0-shapshot.snapshot'
|
||||
|
||||
curl -X POST 'https://node-1.my-cluster.com:6333/collections/test_collection_import/snapshots/upload?priority=snapshot' \
|
||||
-H 'api-key: ${QDRANT_API_KEY}' \
|
||||
-H 'Content-Type:multipart/form-data' \
|
||||
-F 'snapshot=@node-1-shapshot.snapshot'
|
||||
|
||||
curl -X POST 'https://node-2.my-cluster.com:6333/collections/test_collection_import/snapshots/upload?priority=snapshot' \
|
||||
-H 'api-key: ${QDRANT_API_KEY}' \
|
||||
-H 'Content-Type:multipart/form-data' \
|
||||
-F 'snapshot=@node-2-shapshot.snapshot'
|
||||
```
|
||||
|
||||
|
||||
**Important:** We selected `priority=snapshot` to make sure that the snapshot is preferred over the data stored on the node. You can read mode about the priority in the [documentation](/documentation/concepts/snapshots/#snapshot-priority).
|
||||
|
||||
Apart from Snapshots, Qdrant also provides the [Qdrant Migration Tool](https://github.com/qdrant/migration) that supports:
|
||||
- Migration between Qdrant Cloud instances.
|
||||
- Migrating vectors from other providers into Qdrant.
|
||||
- Migrating from Qdrant OSS to Qdrant Cloud.
|
||||
|
||||
Follow our [migration guide](/documentation/database-tutorials/migration/) to learn how to effectively use the Qdrant Migration tool.
|
||||
@@ -0,0 +1,355 @@
|
||||
---
|
||||
title: Large Scale Search
|
||||
weight: 2
|
||||
---
|
||||
|
||||
|
||||
# Upload and Search Large collections cost-efficiently
|
||||
|
||||
| Time: 2 days | Level: Advanced | | |
|
||||
|--------------|-----------------|--|----|
|
||||
|
||||
|
||||
In this tutorial, we will describe an approach to upload, index, and search a large volume of data cost-efficiently,
|
||||
on an example of the real-world dataset [LAION-400M](https://laion.ai/blog/laion-400-open-dataset/).
|
||||
|
||||
The goal of this tutorial is to demonstrate what minimal amount of resources is required to index and search a large dataset,
|
||||
while still maintaining a reasonable search latency and accuracy.
|
||||
|
||||
All relevant code snippets are available in the [GitHub repository](https://github.com/qdrant/laion-400m-benchmark).
|
||||
|
||||
The recommended Qdrant version for this tutorial is `v1.13.5` and higher.
|
||||
|
||||
|
||||
## Dataset
|
||||
|
||||
The dataset we will use is [LAION-400M](https://laion.ai/blog/laion-400-open-dataset/), a collection of approximately 400 million vectors obtained from
|
||||
images extracted from a Common Crawl dataset. Each vector is 512-dimensional and generated using a [CLIP](https://openai.com/blog/clip/) model.
|
||||
|
||||
Vectors are associated with a number of metadata fields, such as `url`, `caption`, `LICENSE`, etc.
|
||||
|
||||
The overall payload size is approximately 200 GB, and the vectors are 400 GB.
|
||||
|
||||
<aside role="status">
|
||||
Dataset doesn't store images themselves, and only contain URLs to the image origin. By the time of writing, some of the URLs are already unavailable.
|
||||
</aside>
|
||||
|
||||
The dataset is available in the form of 409 chunks, each containing approximately 1M vectors.
|
||||
We will use the following [python script](https://github.com/qdrant/laion-400m-benchmark/blob/master/upload.py) to upload dataset chunks one by one.
|
||||
|
||||
## Hardware
|
||||
|
||||
After some initial experiments, we figured out a minimal hardware configuration for the task:
|
||||
|
||||
- 8 CPU cores
|
||||
- 64Gb RAM
|
||||
- 650Gb Disk space
|
||||
|
||||
{{< figure src="/documentation/tutorials/large-scale-search/hardware.png" caption="Hardware configuration" >}}
|
||||
|
||||
|
||||
This configuration is enough to index and explore the dataset in a single-user mode; latency is reasonable enough to build interactive graphs and navigate in the dashboard.
|
||||
|
||||
Naturally, you might need more CPU cores and RAM for production-grade configurations.
|
||||
|
||||
It is important to ensure high network bandwidth for this experiment so you are running the client and server in the same region.
|
||||
|
||||
|
||||
## Uploading and Indexing
|
||||
|
||||
We will use the following [python script](https://github.com/qdrant/laion-400m-benchmark/blob/master/upload.py) to upload dataset chunks one by one.
|
||||
|
||||
```bash
|
||||
export QDRANT_URL="https://xxxx-xxxx.xxxx.cloud.qdrant.io"
|
||||
export QDRANT_API_KEY="xxxx-xxxx-xxxx-xxxx"
|
||||
|
||||
python upload.py
|
||||
```
|
||||
|
||||
This script will download chunks of the LAION dataset one by one and upload them to Qdrant. Intermediate data is not persisted on disk, so the script doesn't require much disk space on the client side.
|
||||
|
||||
Let's take a look at the collection configuration we used:
|
||||
|
||||
```python
|
||||
client.create_collection(
|
||||
QDRANT_COLLECTION_NAME,
|
||||
vectors_config=models.VectorParams(
|
||||
size=512, # CLIP model output size
|
||||
distance=models.Distance.COSINE, # CLIP model uses cosine distance
|
||||
datatype=models.Datatype.FLOAT16, # We only need 16 bits for float, otherwise disk usage would be 800Gb instead of 400Gb
|
||||
on_disk=True # We don't need original vectors in RAM
|
||||
),
|
||||
# Even though CLIP vectors don't work well with binary quantization, out of the box,
|
||||
# we can rely on query-time oversampling to get more accurate results
|
||||
quantization_config=models.BinaryQuantization(
|
||||
binary=models.BinaryQuantizationConfig(
|
||||
always_ram=True,
|
||||
)
|
||||
),
|
||||
optimizers_config=models.OptimizersConfigDiff(
|
||||
# Bigger size of segments are desired for faster search
|
||||
# However it might be slower for indexing
|
||||
max_segment_size=5_000_000,
|
||||
),
|
||||
# Having larger M value is desirable for higher accuracy,
|
||||
# but in our case we care more about memory usage
|
||||
# We could still achieve reasonable accuracy even with M=6 + oversampling
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=6, # decrease M for lower memory usage
|
||||
on_disk=False
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
There are a few important points to note:
|
||||
|
||||
- We use `FLOAT16` datatype for vectors, which allows us to store vectors in half the size compared to `FLOAT32`. There are no significant accuracy losses for this dataset.
|
||||
- We use `BinaryQuantization` with `always_ram=True` to enable query-time oversampling. This allows us to get an accurate and resource-efficient search, even though 512d CLIP vectors don't work well with binary quantization out of the box.
|
||||
- We use `HnswConfig` with `m=6` to reduce memory usage. We will look deeper into memory usage in the next section.
|
||||
|
||||
Goal of this configuration is to ensure that prefetch component of the search never needs to load data from disk, and at least a minimal version of vectors and vector index is always in RAM.
|
||||
The second stage of the search can explicitly determine how many times we can afford to load data from a disk.
|
||||
|
||||
|
||||
In our experiment, the upload process was going at 5000 points per second.
|
||||
The indexation process was going in parallel with the upload and was happening at the rate of approximately 4000 points per second.
|
||||
|
||||
|
||||
{{< figure src="/documentation/tutorials/large-scale-search/upload_process.png" caption="Upload and indexation process" >}}
|
||||
|
||||
## Memory Usage
|
||||
|
||||
After the upload and indexation process is finished, let's take a detailed look at the memory usage of the Qdrant server.
|
||||
|
||||
{{< figure src="/documentation/tutorials/large-scale-search/memory_usage.png" caption="Memory usage" >}}
|
||||
|
||||
On the high level, memory usage consists of 3 components:
|
||||
|
||||
- System memory - 8.34Gb - this is memory reserved for internal systems and OS, it doesn't depend on the dataset size.
|
||||
- Data memory - 39.27Gb - this is a resident memory of qdrant process, it can't be evicter and qdrant process will crash if it exceeds the limit.
|
||||
- Cache memory - 14.54Gb - this is a disk cache qdrant uses. It is necessary for fast search but can be evicted if needed.
|
||||
|
||||
|
||||
The most interest for us is Data and Cache memory. Let's look what exactly is stored in these components.
|
||||
|
||||
In our scenario, Qdrant uses memory to store the following components:
|
||||
|
||||
- Storing vectors
|
||||
- Storing vector index
|
||||
- Storing information about IDs and versions of points
|
||||
|
||||
<aside role="status">
|
||||
Please note, that payload indexes are out of scope for this tutorial. If you are using payload indexes in your collection, you might need to adjust the estimations accordingly.
|
||||
</aside>
|
||||
|
||||
### Size of vectors
|
||||
|
||||
In our scenario, we store only quantized vectors in RAM, so it is relatively easy to calculate the required size:
|
||||
|
||||
```text
|
||||
400_000_000 * 512d / 8 bits / 1024 (Kb) / 1024 (Mb) / 1024 (Gb) = 23.84Gb
|
||||
```
|
||||
|
||||
### Size of vector index
|
||||
|
||||
Vector index is a bit more complicated, as it is not a simple matrix.
|
||||
|
||||
Internally, it is stored as a list of connections in a graph, and each connection is a 4-byte integer.
|
||||
|
||||
The number of connections is defined by the `M` parameter of the HNSW index, and in our case, it is `6` on the high level and `2 x M` on level 0.
|
||||
|
||||
This gives us the following estimation:
|
||||
|
||||
```text
|
||||
400_000_000 * (6 * 2) * 4 bytes / 1024 (Kb) / 1024 (Mb) / 1024 (Gb) = 17.881Gb
|
||||
```
|
||||
|
||||
In practice the size of index is a bit smaller due to the [compression](https://qdrant.tech/blog/qdrant-1.13.x/#hnsw-graph-compression) we implemented in Qdrant v1.13.0, but it is still a good estimation.
|
||||
|
||||
The HNSW index in Qdrant is stored as a mmap, and it can be evicted from RAM if needed.
|
||||
So, the memory consumption of HNSW falls under the category of `Cache memory`.
|
||||
|
||||
|
||||
### Size of IDs and versions
|
||||
|
||||
Qdrant must store additional information about each point, such as ID and version.
|
||||
This information is needed on each request, so it is very important to keep it in RAM for fast access.
|
||||
|
||||
Let's take a look at Qdrant internals to understand how much memory is required for this information.
|
||||
|
||||
```rust
|
||||
|
||||
// This is s simplified version of the IdTracker struct
|
||||
// It omits all optimizations and small details,
|
||||
// but gives a good estimation of memory usage
|
||||
IdTracker {
|
||||
// Mapping of internal id to version (u64), compressed to 4 bytes
|
||||
// Required for versioning and conflict resolution between segments
|
||||
internal_to_version, // 400M x 4 = 1.5Gb
|
||||
|
||||
// Mapping of external id to internal id, 4 bytes per point.
|
||||
// Required to determine original point ID after search inside the segment
|
||||
internal_to_external: Vec<u128>, // 400M x 16 = 6.4Gb
|
||||
|
||||
// Mapping of external id to internal id. For numeric ids it uses 8 bytes,
|
||||
// UUIDs are stored as 16 bytes.
|
||||
// Required to determine sequential point ID inside the segment
|
||||
external_to_internal: Vec<u64, u32>, // 400M x (8 + 4) = 4.5Gb
|
||||
}
|
||||
```
|
||||
|
||||
In the v1.13.5 we introduced a [significant optimization](https://github.com/qdrant/qdrant/pull/6023) to reduce the memory usage of `IdTracker` by approximately 2 times.
|
||||
So the total memory usage of `IdTracker` in our case is approximately `12.4Gb`.
|
||||
|
||||
So total expected RAM usage of Qdrant server in our case is approximately `23.84Gb + 17.881Gb + 12.4Gb = 54.121Gb`, which is very close to the actual memory usage we observed: `39.27Gb + 14.54Gb = 53.81Gb`.
|
||||
|
||||
We had to apply some simplifications to the estimations, but they are good enough to understand the memory usage of the Qdrant server.
|
||||
|
||||
|
||||
## Search
|
||||
|
||||
After the dataset is uploaded and indexed, we can start searching for similar vectors.
|
||||
|
||||
We can start by exploring the dataset in Web-UI. So you can get an intuition into the search performance, not just table numbers.
|
||||
|
||||
{{< figure src="/documentation/tutorials/large-scale-search/web-ui-bear1.png" caption="Web-UI Bear image" width="80%" >}}
|
||||
|
||||
{{< figure src="/documentation/tutorials/large-scale-search/web-ui-bear2.png" caption="Web-UI similar Bear image" width="80%" >}}
|
||||
|
||||
Web-UI default requests do not use oversampling, but the observable results are still good enough to see the resemblance between images.
|
||||
|
||||
|
||||
### Ground truth data
|
||||
|
||||
However, to estimate the search performance more accurately, we need to compare search results with the ground truth.
|
||||
Unfortunately, the LAION dataset doesn't contain usable ground truth, so we had to generate it ourselves.
|
||||
|
||||
To do this, we need to perform a full-scan search for each vector in the dataset and store the results in a separate file.
|
||||
Unfortunately, this process is very time-consuming and requires a lot of resources, so we had to limit the number of queries to 100,
|
||||
we provide a ready-to-use [ground truth file](https://github.com/qdrant/laion-400m-benchmark/blob/master/expected.py) and the [script](https://github.com/qdrant/laion-400m-benchmark/blob/master/full_scan.py) to generate it (requires 512Gb RAM machine and about 20 hours of execution time).
|
||||
|
||||
|
||||
Our ground truth file contains 100 queries, each with 50 results. The first 100 vectors of the dataset itself were used to generate queries.
|
||||
|
||||
<aside role="status">
|
||||
Note, that this dataset contain a significant amount of exact duplicates, so ordering of the results might be different in different runs.
|
||||
</aside>
|
||||
|
||||
|
||||
### Search Query
|
||||
|
||||
To precisely control the amount of oversampling, we will use the following search query:
|
||||
|
||||
```python
|
||||
|
||||
limit = 50
|
||||
rescore_limit = 1000 # oversampling factor is 20
|
||||
|
||||
query = vectors[query_id] # One of existing vectors
|
||||
|
||||
response = client.query_points(
|
||||
collection_name=QDRANT_COLLECTION_NAME,
|
||||
query=query,
|
||||
limit=limit,
|
||||
# Go to disk
|
||||
search_params=models.SearchParams(
|
||||
quantization=models.QuantizationSearchParams(
|
||||
rescore=True,
|
||||
),
|
||||
),
|
||||
# Prefetch is performed using only in-RAM data,
|
||||
# so querying even large amount of data is fast
|
||||
prefetch=models.Prefetch(
|
||||
query=query,
|
||||
limit=rescore_limit,
|
||||
params=models.SearchParams(
|
||||
quantization=models.QuantizationSearchParams(
|
||||
# Avoid rescoring in prefetch
|
||||
# We should do it explicitly on the second stage
|
||||
rescore=False,
|
||||
),
|
||||
)
|
||||
)
|
||||
)
|
||||
```
|
||||
|
||||
As you can see, this query contains two stages:
|
||||
|
||||
- First stage is a prefetch, which is performed using only in-RAM data. It is very fast and allows us to get a large amount of candidates.
|
||||
- The second stage is a rescore, which is performed with full-size vectors stored on disks.
|
||||
|
||||
By using 2-stage search we can precisely control the amount of data loaded from disk and ensure the balance between search speed and accuracy.
|
||||
|
||||
You can find the complete code of the search process in the [eval.py](https://github.com/qdrant/laion-400m-benchmark/blob/master/eval.py)
|
||||
|
||||
|
||||
## Performance tweak
|
||||
|
||||
One important performance tweak we found useful for this dataset is to enable [Async IO](https://qdrant.tech/articles/io_uring) in Qdrant.
|
||||
|
||||
By default, Qdrant uses synchronous IO, which is good for in-memory datasets but can be a bottleneck when we want to read a lot of data from a disk.
|
||||
|
||||
Async IO (implemented with `io_uring`) allows to send parallel requests to the disk and saturate the disk bandwidth.
|
||||
|
||||
This is exactly what we are looking for when performing large-scale re-scoring with original vectors.
|
||||
|
||||
Instead of reading vectors one by one and waiting for the disk response 1000 times, we can send 1000 requests to the disk and wait for all of them to complete. This allows us to saturate the disk bandwidth and get faster results.
|
||||
|
||||
To enable Async IO in Qdrant, you need to set the following environment variable:
|
||||
|
||||
```bash
|
||||
QDRANT__STORAGE__PERFORMANCE__ASYNC_SCORER=true
|
||||
```
|
||||
|
||||
Or set parameter in config file:
|
||||
|
||||
```yaml
|
||||
storage:
|
||||
performance:
|
||||
async_scorer: true
|
||||
```
|
||||
|
||||
In Qdrant Managed cloud Async IO can be enabled via `Advanced optimizations` section in cluster `Configuration` tab.
|
||||
|
||||
{{< figure src="/documentation/tutorials/large-scale-search/async_io.png" caption="Async IO configuration in Cloud" width="80%" >}}
|
||||
|
||||
|
||||
## Running search requests
|
||||
|
||||
Once all the preparations are done, we can run the search requests and evaluate the results.
|
||||
|
||||
You can find the full code of the search process in the [eval.py](https://github.com/qdrant/laion-400m-benchmark/blob/master/eval.py)
|
||||
|
||||
This script will run 100 search requests with configured oversampling factor and compare the results with the ground truth.
|
||||
|
||||
```bash
|
||||
python eval.py --rescore_limit 1000
|
||||
```
|
||||
|
||||
In our request we achieved the following results:
|
||||
|
||||
| Rescore Limit | Precision@50 | Time per request |
|
||||
|---------------|--------------|------------------|
|
||||
| 1000 | 75.2% | 0.7s |
|
||||
| 5000 | 81.0% | 2.2s |
|
||||
|
||||
Additional experiments with `m=16` demonstrated that we can achieve `85%` precision with `rescore_limit=1000`, but they would require slightly more memory.
|
||||
|
||||
{{< figure src="/documentation/tutorials/large-scale-search/precision.png" caption="Log of search evaluation" width="50%">}}
|
||||
|
||||
|
||||
## Conclusion
|
||||
|
||||
In this tutorial we demonstrated how to upload, index and search a large dataset in Qdrant cost-efficiently.
|
||||
Binary quantization can be applied even on 512d vectors, if combined with query-time oversampling.
|
||||
|
||||
Qdrant allows to precisely control where each part of storage is located, which allows to achieve a good balance between search speed and memory usage.
|
||||
|
||||
### Potential improvements
|
||||
|
||||
In this experiment, we investigated in detail which parts of the storage are responsible for memory usage and how to control them.
|
||||
|
||||
One especially interesting part is the `VectorIndex` component, which is responsible for storing the graph of connections between vectors.
|
||||
|
||||
In our further research, we will investigate the possibility of making HNSW more disk-friendly so it can be offloaded to disk without significant performance losses.
|
||||
|
||||
Reference in New Issue
Block a user