mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-04 10:28:29 +02:00
Initial draft of tutorials restructure
This commit is contained in:
@@ -0,0 +1,22 @@
|
||||
---
|
||||
title: Ecosystem & Integrations
|
||||
weight: 20
|
||||
is_empty: false
|
||||
aliases:
|
||||
- how-to
|
||||
- tutorials
|
||||
partition: qdrant
|
||||
---
|
||||
|
||||
# Ecosystem & Integrations
|
||||
*Connect Qdrant to cloud providers, data streams, and ETL tools.*
|
||||
|
||||
| Tutorial | Objective | Stack | Time | Level |
|
||||
| :--- | :--- | :--- | :--- | :--- |
|
||||
| [Embedding Migration](https://qdrant.tech/documentation/tutorials-ecosystem/migration/) | Move dense and sparse embeddings to Qdrant. | CLI | 30m | Intermediate |
|
||||
| [S3 Ingestion with LangChain](https://qdrant.tech/documentation/data-ingestion-beginners/) | Stream data from AWS S3 to vector store. | LangChain | 30m | Beginner |
|
||||
| [Hugging Face Datasets](https://qdrant.tech/documentation/tutorials-ecosystem/huggingface-datasets/) | Load and search public ML datasets. | Python | 15m | Beginner |
|
||||
| [Databricks Integration](https://qdrant.tech/documentation/send-data/databricks/) | Vectorize datasets using FastEmbed on Databricks. | Databricks | 30m | Intermediate |
|
||||
| [Airflow & Astronomer](https://qdrant.tech/documentation/send-data/qdrant-airflow-astronomer/) | Orchestrate data engineering workflows. | Airflow | 45m | Intermediate |
|
||||
| [Kafka Data Streaming](https://qdrant.tech/documentation/send-data/data-streaming-kafka-qdrant/) | Setup Qdrant Sink Connector for real-time data. | Kafka | 60m | Advanced |
|
||||
| [No-Code Automation (n8n)](https://qdrant.tech/documentation/qdrant-n8n/) | Combine Qdrant with low-code n8n workflows. | n8n | 45m | Intermediate |
|
||||
@@ -0,0 +1,117 @@
|
||||
---
|
||||
title: Load a HuggingFace Dataset
|
||||
aliases:
|
||||
- /documentation/tutorials/huggingface-datasets/
|
||||
weight: 3
|
||||
---
|
||||
|
||||
# Load and Search Hugging Face Datasets with Qdrant
|
||||
|
||||
[Hugging Face](https://huggingface.co/) provides a platform for sharing and using ML models and
|
||||
datasets. [Qdrant](https://huggingface.co/Qdrant) also publishes datasets along with the
|
||||
embeddings that you can use to practice with Qdrant and build your applications based on semantic
|
||||
search. **Please [let us know](https://qdrant.to/discord) if you'd like to see a specific dataset!**
|
||||
|
||||
## arxiv-titles-instructorxl-embeddings
|
||||
|
||||
[This dataset](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings) contains
|
||||
embeddings generated from the paper titles only. Each vector has a payload with the title used to
|
||||
create it, along with the DOI (Digital Object Identifier).
|
||||
|
||||
```json
|
||||
{
|
||||
"title": "Nash Social Welfare for Indivisible Items under Separable, Piecewise-Linear Concave Utilities",
|
||||
"DOI": "1612.05191"
|
||||
}
|
||||
```
|
||||
|
||||
You can find a detailed description of the dataset in the [Practice Datasets](/documentation/datasets/#journal-article-titles)
|
||||
section. If you prefer loading the dataset from a Qdrant snapshot, it also linked there.
|
||||
|
||||
Loading the dataset is as simple as using the `load_dataset` function from the `datasets` library:
|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
|
||||
dataset = load_dataset("Qdrant/arxiv-titles-instructorxl-embeddings")
|
||||
```
|
||||
|
||||
<aside role="status">The dataset has over 16 GB, so it might take a while to download.</aside>
|
||||
|
||||
The dataset contains 2,250,000 vectors. This is how you can check the list of the features in the dataset:
|
||||
|
||||
```python
|
||||
dataset.features
|
||||
```
|
||||
|
||||
### Streaming the dataset
|
||||
|
||||
Dataset streaming lets you work with a dataset without downloading it. The data is streamed as
|
||||
you iterate over the dataset. You can read more about it in the [Hugging Face
|
||||
documentation](https://huggingface.co/docs/datasets/stream).
|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
|
||||
dataset = load_dataset(
|
||||
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
|
||||
)
|
||||
```
|
||||
|
||||
### Loading the dataset into Qdrant
|
||||
|
||||
You can load the dataset into Qdrant using the [Python SDK](https://github.com/qdrant/qdrant-client).
|
||||
The embeddings are already precomputed, so you can store them in a collection, that we're going
|
||||
to create in a second:
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("http://localhost:6333")
|
||||
|
||||
client.create_collection(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
vectors_config=models.VectorParams(
|
||||
size=768,
|
||||
distance=models.Distance.COSINE,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
It is always a good idea to use batching, while loading a large dataset, so let's do that.
|
||||
We are going to need a helper function to split the dataset into batches:
|
||||
|
||||
```python
|
||||
from itertools import islice
|
||||
|
||||
def batched(iterable, n):
|
||||
iterator = iter(iterable)
|
||||
while batch := list(islice(iterator, n)):
|
||||
yield batch
|
||||
```
|
||||
|
||||
If you are a happy user of Python 3.12+, you can use the [`batched` function from the `itertools`
|
||||
](https://docs.python.org/3/library/itertools.html#itertools.batched) package instead.
|
||||
|
||||
No matter what Python version you are using, you can use the `upsert` method to load the dataset,
|
||||
batch by batch, into Qdrant:
|
||||
|
||||
```python
|
||||
batch_size = 100
|
||||
|
||||
for batch in batched(dataset, batch_size):
|
||||
ids = [point.pop("id") for point in batch]
|
||||
vectors = [point.pop("vector") for point in batch]
|
||||
|
||||
client.upsert(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
points=models.Batch(
|
||||
ids=ids,
|
||||
vectors=vectors,
|
||||
payloads=batch,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
Your collection is ready to be used for search! Please [let us know using Discord](https://qdrant.to/discord)
|
||||
if you would like to see more datasets published on Hugging Face hub.
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
title: Migration to Qdrant
|
||||
weight: 180
|
||||
---
|
||||
|
||||
# Migration
|
||||
|
||||
Migrating data between vector databases, especially across regions, platforms, or deployment types, can be a hassle. That’s where the [Qdrant Migration Tool](https://github.com/qdrant/migration) comes in. It supports a wide range of migration needs, including transferring data between Qdrant instances and migrating from other vector database providers to Qdrant.
|
||||
|
||||
You can run the migration tool on any machine where you have connectivity to both the source and the target Qdrant databases. Direct connectivity between both databases is not required. For optimal performance, you should run the tool on a machine with a fast network connection and minimum latency to both databases.
|
||||
|
||||
In this tutorial, we will learn how to use the migration tool and walk through a practical example of migrating from other vector databases to Qdrant.
|
||||
|
||||
|
||||
## Why use this instead of Qdrant’s Native Snapshotting?
|
||||
|
||||
Qdrant supports [snapshot-based backups](https://qdrant.tech/documentation/concepts/snapshots/), low-level disk operations built for same cluster recovery or local backups. These snapshots:
|
||||
|
||||
* Require snapshot consistency across nodes.
|
||||
* Can be hard to port across machines or cloud zones.
|
||||
|
||||
On the other hand, the Qdrant Migration Tool:
|
||||
|
||||
* Streams data in live batches.
|
||||
* Can resume interrupted migrations.
|
||||
* Works even when data is being inserted.
|
||||
* Supports collection reconfiguration (e.g., change replication, and quantization)
|
||||
* Supports migrating from other vector DBs (Pinecone, Chroma, Weaviate, etc.)
|
||||
|
||||
## How to Use the Qdrant Migration Tool
|
||||
|
||||
You can run the tool via Docker.
|
||||
|
||||
Installation:
|
||||
|
||||
```shell
|
||||
docker pull registry.cloud.qdrant.io/library/qdrant-migration
|
||||
```
|
||||
|
||||
Here is an example of how to perform a Qdrant to Qdrant migration:
|
||||
|
||||
```bash
|
||||
docker run --rm -it \
|
||||
registry.cloud.qdrant.io/library/qdrant-migration qdrant \
|
||||
--source.url 'https://source-instance.cloud.qdrant.io:6334' \
|
||||
--source.api-key 'qdrant-source-key' \
|
||||
--source.collection 'benchmark' \
|
||||
--target.url 'https://target-instance.cloud.qdrant.io:6334' \
|
||||
--target.api-key 'qdrant-target-key' \
|
||||
--target.collection 'benchmark'
|
||||
```
|
||||
|
||||
<aside role="alert">
|
||||
Note: The migration CLI uses the Qdrant GRPC API, this means you must always configure the GRPC port for Qdrant URLs with the Migration CLI (default: 6334).
|
||||
</aside>
|
||||
|
||||
## Example: Migrate from Pinecone to Qdrant
|
||||
|
||||
Let’s now walk through an example of migrating from Pinecone to Qdrant. Assuming your Pinecone index looks like this:
|
||||
|
||||

|
||||
|
||||
The information you need from Pinecone is:
|
||||
|
||||
* Your Pinecone API key
|
||||
* The index name
|
||||
* The index host URL
|
||||
|
||||
With that information, you can migrate your vector database from Pinecone to Qdrant with the following command:
|
||||
|
||||
```bash
|
||||
docker run --net=host --rm -it registry.cloud.qdrant.io/library/qdrant-migration pinecone \
|
||||
--pinecone.index-host 'https://sample-movies-efgjrye.svc.aped-4627-b74a.pinecone.io' \
|
||||
--pinecone.index-name 'sample-movies' \
|
||||
--pinecone.api-key 'pcsk_7Dh5MW_…' \
|
||||
--qdrant.url 'https://5f1a5c6c-7d47-45c3-8d47-d7389b1fad66.eu-west-1-0.aws.cloud.qdrant.io:6334' \
|
||||
--qdrant.api-key 'eyJhbGciOiJIUzI1NiIsInR5c…' \
|
||||
--qdrant.collection 'sample-movies' \
|
||||
--migration.batch-size 64
|
||||
|
||||
|
||||
```
|
||||
When the migration is complete, you will see the new collection on Qdrant with all the vectors.
|
||||
|
||||
## Conclusion
|
||||
|
||||
The **Qdrant Migration Tool** makes data transfer across vector database instances effortless. Whether you're moving between cloud regions, upgrading from self-hosted to Qdrant Cloud, or switching from other databases such as Pinecone, this tool saves you hours of manual effort. [Try it today](https://github.com/qdrant/migration).
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user