mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-27 15:08:30 +02:00
separate tutorials by topic
This commit is contained in:
@@ -0,0 +1,104 @@
|
||||
---
|
||||
title: Bulk upload
|
||||
weight: 30
|
||||
---
|
||||
|
||||
# Bulk upload a large number of vectors
|
||||
|
||||
Uploading a large-scale dataset fast might be a challenge, but Qdrant has a few tricks to help you with that.
|
||||
|
||||
The first important detail about data uploading is that the bottleneck is usually located on the client side, not on the server side.
|
||||
This means that if you are uploading a large dataset, you should prefer a high-performance client library.
|
||||
|
||||
We recommend using our [Rust client library](https://github.com/qdrant/rust-client) for this purpose, as it is the fastest client library available for Qdrant.
|
||||
|
||||
If you are not using Rust, you might want to consider parallelizing your upload process.
|
||||
|
||||
## Disable indexing during upload
|
||||
|
||||
In case you are doing an initial upload of a large dataset, you might want to disable indexing during upload.
|
||||
It will enable to avoid unnecessary indexing of vectors, which will be overwritten by the next batch.
|
||||
|
||||
To disable indexing during upload, set `indexing_threshold` to `0`:
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
|
||||
{
|
||||
"vectors": {
|
||||
"size": 768,
|
||||
"distance": "Cosine"
|
||||
},
|
||||
"optimizers_config": {
|
||||
"indexing_threshold": 0
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
|
||||
client.recreate_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config=models.VectorParams(size=768, distance=models.Distance.COSINE),
|
||||
optimizers_config=models.OptimizersConfigDiff(
|
||||
indexing_threshold=0,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
After upload is done, you can enable indexing by setting `indexing_threshold` to a desired value (default is 20000):
|
||||
|
||||
```http
|
||||
PATCH /collections/{collection_name}
|
||||
|
||||
{
|
||||
"optimizers_config": {
|
||||
"indexing_threshold": 20000
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
|
||||
client.update_collection(
|
||||
collection_name="{collection_name}",
|
||||
optimizer_config=models.OptimizersConfigDiff(
|
||||
indexing_threshold=20000
|
||||
)
|
||||
)
|
||||
```
|
||||
|
||||
## Parallel upload into multiple shards
|
||||
|
||||
In Qdrant, each collection is split into shards. Each shard has a separate Write-Ahead-Log (WAL), which is responsible for ordering operations.
|
||||
By creating multiple shards, you can parallelize upload of a large dataset. From 2 to 4 shards per one machine is a reasonable number.
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
|
||||
{
|
||||
"vectors": {
|
||||
"size": 768,
|
||||
"distance": "Cosine"
|
||||
},
|
||||
"shard_number": 2
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("localhost", port=6333)
|
||||
|
||||
client.recreate_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config=models.VectorParams(size=768, distance=models.Distance.COSINE),
|
||||
shard_number=2,
|
||||
)
|
||||
```
|
||||
Reference in New Issue
Block a user