mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-28 23:48:31 +02:00
add content
This commit is contained in:
@@ -0,0 +1,22 @@
|
||||
---
|
||||
title: Using the Database
|
||||
weight: 18
|
||||
# If the index.md file is empty, the link to the section will be hidden from the sidebar
|
||||
is_empty: false
|
||||
aliases:
|
||||
- how-to
|
||||
- tutorials
|
||||
partition: qdrant
|
||||
---
|
||||
|
||||
# Database Tutorials
|
||||
|
||||
These tutorials demonstrate different ways you can build vector search into your applications.
|
||||
|
||||
| Essential How-Tos | Description | Stack |
|
||||
|---------------------------------------------------------------------------------|-------------------------------------------------------------------|---------------------------------------------|
|
||||
| [Bulk Upload Vectors](/documentation/tutorials/bulk-upload/) | Upload a large scale dataset. | Qdrant |
|
||||
| [Asynchronous API](/documentation/tutorials/async-api/) | Communicate with Qdrant server asynchronously with Python SDK. | Qdrant, Python |
|
||||
| [Create Dataset Snapshots](/documentation/tutorials/create-snapshot/) | Turn a dataset into a snapshot by exporting it from a collection. | Qdrant |
|
||||
| [Load HuggingFace Dataset](/documentation/tutorials/huggingface-datasets/) | Load a Hugging Face dataset to Qdrant | Qdrant, Python, datasets |
|
||||
|
||||
@@ -0,0 +1,94 @@
|
||||
---
|
||||
title: Using the Async API
|
||||
weight: 4
|
||||
---
|
||||
|
||||
# Using Qdrant asynchronously
|
||||
|
||||
Asynchronous programming is being broadly adopted in the Python ecosystem. Tools such as FastAPI [have embraced this new
|
||||
paradigm](https://fastapi.tiangolo.com/async/), but it is also becoming a standard for ML models served as SaaS. For example, the Cohere SDK
|
||||
[provides an async client](https://github.com/cohere-ai/cohere-python/blob/856a4c3bd29e7a75fa66154b8ac9fcdf1e0745e0/src/cohere/client.py#L189) next to its synchronous counterpart.
|
||||
|
||||
Databases are often launched as separate services and are accessed via a network. All the interactions with them are IO-bound and can
|
||||
be performed asynchronously so as not to waste time actively waiting for a server response. In Python, this is achieved by
|
||||
using [`async/await`](https://docs.python.org/3/library/asyncio-task.html) syntax. That lets the interpreter switch to another task
|
||||
while waiting for a response from the server.
|
||||
|
||||
## When to use async API
|
||||
|
||||
There is no need to use async API if the application you are writing will never support multiple users at once (e.g it is a script that runs once per day). However, if you are writing a web service that multiple users will use simultaneously, you shouldn't be
|
||||
blocking the threads of the web server as it limits the number of concurrent requests it can handle. In this case, you should use
|
||||
the async API.
|
||||
|
||||
Modern web frameworks like [FastAPI](https://fastapi.tiangolo.com/) and [Quart](https://quart.palletsprojects.com/en/latest/) support
|
||||
async API out of the box. Mixing asynchronous code with an existing synchronous codebase might be a challenge. The `async/await` syntax
|
||||
cannot be used in synchronous functions. On the other hand, calling an IO-bound operation synchronously in async code is considered
|
||||
an antipattern. Therefore, if you build an async web service, exposed through an [ASGI](https://asgi.readthedocs.io/en/latest/) server,
|
||||
you should use the async API for all the interactions with Qdrant.
|
||||
|
||||
<aside role="status">
|
||||
All the async code has to be launched in an async context. Usually, it means you have to use <code>asyncio.run</code> or <code>asyncio.create_task</code> to run them.
|
||||
Please refer to the <a href="https://docs.python.org/3/library/asyncio.html">asyncio documentation</a> for more details.
|
||||
</aside>
|
||||
|
||||
### Using Qdrant asynchronously
|
||||
|
||||
The simplest way of running asynchronous code is to use define `async` function and use the `asyncio.run` in the following way to run it:
|
||||
|
||||
```python
|
||||
from qdrant_client import models
|
||||
|
||||
import qdrant_client
|
||||
import asyncio
|
||||
|
||||
|
||||
async def main():
|
||||
client = qdrant_client.AsyncQdrantClient("localhost")
|
||||
|
||||
# Create a collection
|
||||
await client.create_collection(
|
||||
collection_name="my_collection",
|
||||
vectors_config=models.VectorParams(size=4, distance=models.Distance.COSINE),
|
||||
)
|
||||
|
||||
# Insert a vector
|
||||
await client.upsert(
|
||||
collection_name="my_collection",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id="5c56c793-69f3-4fbf-87e6-c4bf54c28c26",
|
||||
payload={
|
||||
"color": "red",
|
||||
},
|
||||
vector=[0.9, 0.1, 0.1, 0.5],
|
||||
),
|
||||
],
|
||||
)
|
||||
|
||||
# Search for nearest neighbors
|
||||
points = await client.query_points(
|
||||
collection_name="my_collection",
|
||||
query=[0.9, 0.1, 0.1, 0.5],
|
||||
limit=2,
|
||||
).points
|
||||
|
||||
# Your async code using AsyncQdrantClient might be put here
|
||||
# ...
|
||||
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
The `AsyncQdrantClient` provides the same methods as the synchronous counterpart `QdrantClient`. If you already have a synchronous
|
||||
codebase, switching to async API is as simple as replacing `QdrantClient` with `AsyncQdrantClient` and adding `await` before each
|
||||
method call.
|
||||
|
||||
<aside role="status">
|
||||
Asynchronous client was introduced in <code>qdrant-client</code> version 1.6.1. If you are using an older version, you need to use autogenerated async clients directly.
|
||||
</aside>
|
||||
|
||||
## Supported Python libraries
|
||||
|
||||
Qdrant integrates with numerous Python libraries. Until recently, only [Langchain](https://python.langchain.com) provided async Python API support.
|
||||
Qdrant is the only vector database with full coverage of async API in Langchain. Their documentation [describes how to use
|
||||
it](https://python.langchain.com/docs/modules/data_connection/vectorstores/#asynchronous-operations).
|
||||
@@ -0,0 +1,162 @@
|
||||
---
|
||||
title: Bulk Upload Vectors
|
||||
weight: 1
|
||||
---
|
||||
|
||||
# Bulk upload a large number of vectors
|
||||
|
||||
Uploading a large-scale dataset fast might be a challenge, but Qdrant has a few tricks to help you with that.
|
||||
|
||||
The first important detail about data uploading is that the bottleneck is usually located on the client side, not on the server side.
|
||||
This means that if you are uploading a large dataset, you should prefer a high-performance client library.
|
||||
|
||||
We recommend using our [Rust client library](https://github.com/qdrant/rust-client) for this purpose, as it is the fastest client library available for Qdrant.
|
||||
|
||||
If you are not using Rust, you might want to consider parallelizing your upload process.
|
||||
|
||||
## Disable indexing during upload
|
||||
|
||||
In case you are doing an initial upload of a large dataset, you might want to disable indexing during upload.
|
||||
It will enable to avoid unnecessary indexing of vectors, which will be overwritten by the next batch.
|
||||
|
||||
To disable indexing during upload, set `indexing_threshold` to `0`:
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"vectors": {
|
||||
"size": 768,
|
||||
"distance": "Cosine"
|
||||
},
|
||||
"optimizers_config": {
|
||||
"indexing_threshold": 0
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config=models.VectorParams(size=768, distance=models.Distance.COSINE),
|
||||
optimizers_config=models.OptimizersConfigDiff(
|
||||
indexing_threshold=0,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.createCollection("{collection_name}", {
|
||||
vectors: {
|
||||
size: 768,
|
||||
distance: "Cosine",
|
||||
},
|
||||
optimizers_config: {
|
||||
indexing_threshold: 0,
|
||||
},
|
||||
});
|
||||
```
|
||||
|
||||
After upload is done, you can enable indexing by setting `indexing_threshold` to a desired value (default is 20000):
|
||||
|
||||
```http
|
||||
PATCH /collections/{collection_name}
|
||||
{
|
||||
"optimizers_config": {
|
||||
"indexing_threshold": 20000
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.update_collection(
|
||||
collection_name="{collection_name}",
|
||||
optimizer_config=models.OptimizersConfigDiff(indexing_threshold=20000),
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.updateCollection("{collection_name}", {
|
||||
optimizers_config: {
|
||||
indexing_threshold: 20000,
|
||||
},
|
||||
});
|
||||
```
|
||||
|
||||
## Upload directly to disk
|
||||
|
||||
When the vectors you upload do not all fit in RAM, you likely want to use
|
||||
[memmap](/documentation/concepts/storage/#configuring-memmap-storage)
|
||||
support.
|
||||
|
||||
During collection
|
||||
[creation](/documentation/concepts/collections/#create-collection),
|
||||
memmaps may be enabled on a per-vector basis using the `on_disk` parameter. This
|
||||
will store vector data directly on disk at all times. It is suitable for
|
||||
ingesting a large amount of data, essential for the billion scale benchmark.
|
||||
|
||||
Using `memmap_threshold` is not recommended in this case. It would require
|
||||
the [optimizer](/documentation/concepts/optimizer/) to constantly
|
||||
transform in-memory segments into memmap segments on disk. This process is
|
||||
slower, and the optimizer can be a bottleneck when ingesting a large amount of
|
||||
data.
|
||||
|
||||
Read more about this in
|
||||
[Configuring Memmap Storage](/documentation/concepts/storage/#configuring-memmap-storage).
|
||||
|
||||
## Parallel upload into multiple shards
|
||||
|
||||
In Qdrant, each collection is split into shards. Each shard has a separate Write-Ahead-Log (WAL), which is responsible for ordering operations.
|
||||
By creating multiple shards, you can parallelize upload of a large dataset. From 2 to 4 shards per one machine is a reasonable number.
|
||||
|
||||
```http
|
||||
PUT /collections/{collection_name}
|
||||
{
|
||||
"vectors": {
|
||||
"size": 768,
|
||||
"distance": "Cosine"
|
||||
},
|
||||
"shard_number": 2
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config=models.VectorParams(size=768, distance=models.Distance.COSINE),
|
||||
shard_number=2,
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
client.createCollection("{collection_name}", {
|
||||
vectors: {
|
||||
size: 768,
|
||||
distance: "Cosine",
|
||||
},
|
||||
shard_number: 2,
|
||||
});
|
||||
```
|
||||
@@ -0,0 +1,276 @@
|
||||
---
|
||||
title: Create & Restore Snapshots
|
||||
weight: 2
|
||||
---
|
||||
|
||||
# Create and restore collections from snapshot
|
||||
|
||||
| Time: 20 min | Level: Beginner | | |
|
||||
|--------------|-----------------|--|----|
|
||||
|
||||
A collection is a basic unit of data storage in Qdrant. It contains vectors, their IDs, and payloads. However, keeping the search efficient requires additional data structures to be built on top of the data. Building these data structures may take a while, especially for large collections.
|
||||
That's why using snapshots is the best way to export and import Qdrant collections, as they contain all the bits and pieces required to restore the entire collection efficiently.
|
||||
|
||||
This tutorial will show you how to create a snapshot of a collection and restore it. Since working with snapshots in a distributed environment might be thought to be a bit more complex, we will use a 3-node Qdrant cluster. However, the same approach applies to a single-node setup.
|
||||
|
||||
<aside role="status">Snapshots cannot be created in local mode of Python SDK. You need to spin up a Qdrant Docker container or use Qdrant Cloud.</aside>
|
||||
|
||||
You can use the techniques described in this page to migrate a cluster. Follow the instructions
|
||||
in this tutorial to create and download snapshots. When you [Restore from snapshot](#restore-from-snapshot), restore your data to the new cluster.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Let's assume you already have a running Qdrant instance or a cluster. If not, you can follow the [installation guide](/documentation/guides/installation/) to set up a local Qdrant instance or use [Qdrant Cloud](https://cloud.qdrant.io/) to create a cluster in a few clicks.
|
||||
|
||||
Once the cluster is running, let's install the required dependencies:
|
||||
|
||||
```shell
|
||||
pip install qdrant-client datasets
|
||||
```
|
||||
|
||||
### Establish a connection to Qdrant
|
||||
|
||||
We are going to use the Python SDK and raw HTTP calls to interact with Qdrant. Since we are going to use a 3-node cluster, we need to know the URLs of all the nodes. For the simplicity, let's keep them all in constants, along with the API key, so we can refer to them later:
|
||||
|
||||
```python
|
||||
QDRANT_MAIN_URL = "https://my-cluster.com:6333"
|
||||
QDRANT_NODES = (
|
||||
"https://node-0.my-cluster.com:6333",
|
||||
"https://node-1.my-cluster.com:6333",
|
||||
"https://node-2.my-cluster.com:6333",
|
||||
)
|
||||
QDRANT_API_KEY = "my-api-key"
|
||||
```
|
||||
|
||||
<aside role="status">If you are using Qdrant Cloud, you can find the URL and API key in the <a href="https://cloud.qdrant.io/">Qdrant Cloud dashboard</a>.</aside>
|
||||
|
||||
We can now create a client instance:
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient(QDRANT_MAIN_URL, api_key=QDRANT_API_KEY)
|
||||
```
|
||||
|
||||
First of all, we are going to create a collection from a precomputed dataset. If you already have a collection, you can skip this step and start by [creating a snapshot](#create-and-download-snapshots).
|
||||
|
||||
<details>
|
||||
<summary>(Optional) Create collection and import data</summary>
|
||||
|
||||
### Load the dataset
|
||||
|
||||
We are going to use a dataset with precomputed embeddings, available on Hugging Face Hub. The dataset is called [Qdrant/arxiv-titles-instructorxl-embeddings](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings) and was created using the [InstructorXL](https://huggingface.co/hkunlp/instructor-xl) model. It contains 2.25M embeddings for the titles of the papers from the [arXiv](https://arxiv.org/) dataset.
|
||||
|
||||
Loading the dataset is as simple as:
|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
|
||||
dataset = load_dataset(
|
||||
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
|
||||
)
|
||||
```
|
||||
|
||||
We used the streaming mode, so the dataset is not loaded into memory. Instead, we can iterate through it and extract the id and vector embedding:
|
||||
|
||||
```python
|
||||
for payload in dataset:
|
||||
id_ = payload.pop("id")
|
||||
vector = payload.pop("vector")
|
||||
print(id_, vector, payload)
|
||||
```
|
||||
|
||||
A single payload looks like this:
|
||||
|
||||
```json
|
||||
{
|
||||
'title': 'Dynamics of partially localized brane systems',
|
||||
'DOI': '1109.1415'
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
### Create a collection
|
||||
|
||||
First things first, we need to create our collection. We're not going to play with the configuration of it, but it makes sense to do it right now.
|
||||
The configuration is also a part of the collection snapshot.
|
||||
|
||||
```python
|
||||
from qdrant_client import models
|
||||
|
||||
if not client.collection_exists("test_collection"):
|
||||
client.create_collection(
|
||||
collection_name="test_collection",
|
||||
vectors_config=models.VectorParams(
|
||||
size=768, # Size of the embedding vector generated by the InstructorXL model
|
||||
distance=models.Distance.COSINE
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
### Upload the dataset
|
||||
|
||||
Calculating the embeddings is usually a bottleneck of the vector search pipelines, but we are happy to have them in place already. Since the goal of this tutorial is to show how to create a snapshot, **we are going to upload only a small part of the dataset**.
|
||||
|
||||
```python
|
||||
ids, vectors, payloads = [], [], []
|
||||
for payload in dataset:
|
||||
id_ = payload.pop("id")
|
||||
vector = payload.pop("vector")
|
||||
|
||||
ids.append(id_)
|
||||
vectors.append(vector)
|
||||
payloads.append(payload)
|
||||
|
||||
# We are going to upload only 1000 vectors
|
||||
if len(ids) == 1000:
|
||||
break
|
||||
|
||||
client.upsert(
|
||||
collection_name="test_collection",
|
||||
points=models.Batch(
|
||||
ids=ids,
|
||||
vectors=vectors,
|
||||
payloads=payloads,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
Our collection is now ready to be used for search. Let's create a snapshot of it.
|
||||
|
||||
</details>
|
||||
|
||||
If you already have a collection, you can skip the previous step and start by [creating a snapshot](#create-and-download-snapshots).
|
||||
|
||||
## Create and download snapshots
|
||||
|
||||
Qdrant exposes an HTTP endpoint to request creating a snapshot, but we can also call it with the Python SDK.
|
||||
Our setup consists of 3 nodes, so we need to call the endpoint **on each of them** and create a snapshot on each node. While using Python SDK, that means creating a separate client instance for each node.
|
||||
|
||||
|
||||
<aside role="status">You may get a timeout error, if the collection size is big. You can trigger the snapshot process in the background, without awaiting for the result, by using <code>wait=false</code> parameter. You can always <a href="/documentation/concepts/snapshots/#list-snapshot">list all the snapshots through the API</a> later on.</aside>
|
||||
|
||||
|
||||
```python
|
||||
snapshot_urls = []
|
||||
for node_url in QDRANT_NODES:
|
||||
node_client = QdrantClient(node_url, api_key=QDRANT_API_KEY)
|
||||
snapshot_info = node_client.create_snapshot(collection_name="test_collection")
|
||||
|
||||
snapshot_url = f"{node_url}/collections/test_collection/snapshots/{snapshot_info.name}"
|
||||
snapshot_urls.append(snapshot_url)
|
||||
```
|
||||
|
||||
```http
|
||||
// for `https://node-0.my-cluster.com:6333`
|
||||
POST /collections/test_collection/snapshots
|
||||
|
||||
// for `https://node-1.my-cluster.com:6333`
|
||||
POST /collections/test_collection/snapshots
|
||||
|
||||
// for `https://node-2.my-cluster.com:6333`
|
||||
POST /collections/test_collection/snapshots
|
||||
```
|
||||
|
||||
<details>
|
||||
<summary>Response</summary>
|
||||
|
||||
```json
|
||||
{
|
||||
"result": {
|
||||
"name": "test_collection-559032209313046-2024-01-03-13-20-11.snapshot",
|
||||
"creation_time": "2024-01-03T13:20:11",
|
||||
"size": 18956800
|
||||
},
|
||||
"status": "ok",
|
||||
"time": 0.307644965
|
||||
}
|
||||
```
|
||||
</details>
|
||||
|
||||
|
||||
|
||||
Once we have the snapshot URLs, we can download them. Please make sure to include the API key in the request headers.
|
||||
Downloading the snapshot **can be done only through the HTTP API**, so we are going to use the `requests` library.
|
||||
|
||||
```python
|
||||
import requests
|
||||
import os
|
||||
|
||||
# Create a directory to store snapshots
|
||||
os.makedirs("snapshots", exist_ok=True)
|
||||
|
||||
local_snapshot_paths = []
|
||||
for snapshot_url in snapshot_urls:
|
||||
snapshot_name = os.path.basename(snapshot_url)
|
||||
local_snapshot_path = os.path.join("snapshots", snapshot_name)
|
||||
|
||||
response = requests.get(
|
||||
snapshot_url, headers={"api-key": QDRANT_API_KEY}
|
||||
)
|
||||
with open(local_snapshot_path, "wb") as f:
|
||||
response.raise_for_status()
|
||||
f.write(response.content)
|
||||
|
||||
local_snapshot_paths.append(local_snapshot_path)
|
||||
```
|
||||
|
||||
Alternatively, you can use the `wget` command:
|
||||
|
||||
```bash
|
||||
wget https://node-0.my-cluster.com:6333/collections/test_collection/snapshots/test_collection-559032209313046-2024-01-03-13-20-11.snapshot \
|
||||
--header="api-key: ${QDRANT_API_KEY}" \
|
||||
-O node-0-shapshot.snapshot
|
||||
|
||||
wget https://node-1.my-cluster.com:6333/collections/test_collection/snapshots/test_collection-559032209313047-2024-01-03-13-20-12.snapshot \
|
||||
--header="api-key: ${QDRANT_API_KEY}" \
|
||||
-O node-1-shapshot.snapshot
|
||||
|
||||
wget https://node-2.my-cluster.com:6333/collections/test_collection/snapshots/test_collection-559032209313048-2024-01-03-13-20-13.snapshot \
|
||||
--header="api-key: ${QDRANT_API_KEY}" \
|
||||
-O node-2-shapshot.snapshot
|
||||
```
|
||||
|
||||
The snapshots are now stored locally. We can use them to restore the collection to a different Qdrant instance, or treat them as a backup. We will create another collection using the same data on the same cluster.
|
||||
|
||||
## Restore from snapshot
|
||||
|
||||
Our brand-new snapshot is ready to be restored. Typically, it is used to move a collection to a different Qdrant instance, but we are going to use it to create a new collection on the same cluster.
|
||||
It is just going to have a different name, `test_collection_import`. We do not need to create a collection first, as it is going to be created automatically.
|
||||
|
||||
Restoring collection is also done separately on each node, but our Python SDK does not support it yet. We are going to use the HTTP API instead,
|
||||
and send a request to each node using `requests` library.
|
||||
|
||||
```python
|
||||
for node_url, snapshot_path in zip(QDRANT_NODES, local_snapshot_paths):
|
||||
snapshot_name = os.path.basename(snapshot_path)
|
||||
requests.post(
|
||||
f"{node_url}/collections/test_collection_import/snapshots/upload?priority=snapshot",
|
||||
headers={
|
||||
"api-key": QDRANT_API_KEY,
|
||||
},
|
||||
files={"snapshot": (snapshot_name, open(snapshot_path, "rb"))},
|
||||
)
|
||||
```
|
||||
|
||||
Alternatively, you can use the `curl` command:
|
||||
|
||||
```bash
|
||||
curl -X POST 'https://node-0.my-cluster.com:6333/collections/test_collection_import/snapshots/upload?priority=snapshot' \
|
||||
-H 'api-key: ${QDRANT_API_KEY}' \
|
||||
-H 'Content-Type:multipart/form-data' \
|
||||
-F 'snapshot=@node-0-shapshot.snapshot'
|
||||
|
||||
curl -X POST 'https://node-1.my-cluster.com:6333/collections/test_collection_import/snapshots/upload?priority=snapshot' \
|
||||
-H 'api-key: ${QDRANT_API_KEY}' \
|
||||
-H 'Content-Type:multipart/form-data' \
|
||||
-F 'snapshot=@node-1-shapshot.snapshot'
|
||||
|
||||
curl -X POST 'https://node-2.my-cluster.com:6333/collections/test_collection_import/snapshots/upload?priority=snapshot' \
|
||||
-H 'api-key: ${QDRANT_API_KEY}' \
|
||||
-H 'Content-Type:multipart/form-data' \
|
||||
-F 'snapshot=@node-2-shapshot.snapshot'
|
||||
```
|
||||
|
||||
|
||||
**Important:** We selected `priority=snapshot` to make sure that the snapshot is preferred over the data stored on the node. You can read mode about the priority in the [documentation](/documentation/concepts/snapshots/#snapshot-priority).
|
||||
@@ -0,0 +1,115 @@
|
||||
---
|
||||
title: Load a HuggingFace Dataset
|
||||
weight: 3
|
||||
---
|
||||
|
||||
# Loading a dataset from Hugging Face hub
|
||||
|
||||
[Hugging Face](https://huggingface.co/) provides a platform for sharing and using ML models and
|
||||
datasets. [Qdrant](https://huggingface.co/Qdrant) also publishes datasets along with the
|
||||
embeddings that you can use to practice with Qdrant and build your applications based on semantic
|
||||
search. **Please [let us know](https://qdrant.to/discord) if you'd like to see a specific dataset!**
|
||||
|
||||
## arxiv-titles-instructorxl-embeddings
|
||||
|
||||
[This dataset](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings) contains
|
||||
embeddings generated from the paper titles only. Each vector has a payload with the title used to
|
||||
create it, along with the DOI (Digital Object Identifier).
|
||||
|
||||
```json
|
||||
{
|
||||
"title": "Nash Social Welfare for Indivisible Items under Separable, Piecewise-Linear Concave Utilities",
|
||||
"DOI": "1612.05191"
|
||||
}
|
||||
```
|
||||
|
||||
You can find a detailed description of the dataset in the [Practice Datasets](/documentation/datasets/#journal-article-titles)
|
||||
section. If you prefer loading the dataset from a Qdrant snapshot, it also linked there.
|
||||
|
||||
Loading the dataset is as simple as using the `load_dataset` function from the `datasets` library:
|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
|
||||
dataset = load_dataset("Qdrant/arxiv-titles-instructorxl-embeddings")
|
||||
```
|
||||
|
||||
<aside role="status">The dataset has over 16 GB, so it might take a while to download.</aside>
|
||||
|
||||
The dataset contains 2,250,000 vectors. This is how you can check the list of the features in the dataset:
|
||||
|
||||
```python
|
||||
dataset.features
|
||||
```
|
||||
|
||||
### Streaming the dataset
|
||||
|
||||
Dataset streaming lets you work with a dataset without downloading it. The data is streamed as
|
||||
you iterate over the dataset. You can read more about it in the [Hugging Face
|
||||
documentation](https://huggingface.co/docs/datasets/stream).
|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
|
||||
dataset = load_dataset(
|
||||
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
|
||||
)
|
||||
```
|
||||
|
||||
### Loading the dataset into Qdrant
|
||||
|
||||
You can load the dataset into Qdrant using the [Python SDK](https://github.com/qdrant/qdrant-client).
|
||||
The embeddings are already precomputed, so you can store them in a collection, that we're going
|
||||
to create in a second:
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
vectors_config=models.VectorParams(
|
||||
size=768,
|
||||
distance=models.Distance.COSINE,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
It is always a good idea to use batching, while loading a large dataset, so let's do that.
|
||||
We are going to need a helper function to split the dataset into batches:
|
||||
|
||||
```python
|
||||
from itertools import islice
|
||||
|
||||
def batched(iterable, n):
|
||||
iterator = iter(iterable)
|
||||
while batch := list(islice(iterator, n)):
|
||||
yield batch
|
||||
```
|
||||
|
||||
If you are a happy user of Python 3.12+, you can use the [`batched` function from the `itertools`
|
||||
](https://docs.python.org/3/library/itertools.html#itertools.batched) package instead.
|
||||
|
||||
No matter what Python version you are using, you can use the `upsert` method to load the dataset,
|
||||
batch by batch, into Qdrant:
|
||||
|
||||
```python
|
||||
batch_size = 100
|
||||
|
||||
for batch in batched(dataset, batch_size):
|
||||
ids = [point.pop("id") for point in batch]
|
||||
vectors = [point.pop("vector") for point in batch]
|
||||
|
||||
client.upsert(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
points=models.Batch(
|
||||
ids=ids,
|
||||
vectors=vectors,
|
||||
payloads=batch,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
Your collection is ready to be used for search! Please [let us know using Discord](https://qdrant.to/discord)
|
||||
if you would like to see more datasets published on Hugging Face hub.
|
||||
Reference in New Issue
Block a user