mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-28 23:48:31 +02:00
Add "Practice datasets" section (#199)
* Add common datasets * Add instruction for titles * Add tutorial on how to use Hugging Face datasets * Add an information about Qdrant org at Hugging Face * Add links to Hugging Face datasets * Add preview images * Add CTA (Discord) * Fix typo * Add link to Discord * List available datasets on top of the page * Fix some typos and improve intro * Fix some typos and improve intro * Add fixed snapshot links * Add Wolt food dataset * Add API calls to restore the snapshots * Adjust the weight of the section * Reorder tutorials
This commit is contained in:
@@ -188,7 +188,6 @@ keywords = "search engine, vector database, neural network, matching, filter, Sa
|
|||||||
[menu.main.params]
|
[menu.main.params]
|
||||||
external = true
|
external = true
|
||||||
|
|
||||||
|
|
||||||
[[menu.main]]
|
[[menu.main]]
|
||||||
identifier = "community"
|
identifier = "community"
|
||||||
name = "Community"
|
name = "Community"
|
||||||
|
|||||||
@@ -1,12 +0,0 @@
|
|||||||
---
|
|
||||||
title: Common Datasets
|
|
||||||
draft: true
|
|
||||||
description: Ready snapshots of some standard datasets which you can easily import into Qdrant.
|
|
||||||
keywords:
|
|
||||||
- embeddings
|
|
||||||
- vector embeddings
|
|
||||||
- vectors
|
|
||||||
- precomputed embeddings
|
|
||||||
- datasets
|
|
||||||
- dataset snapshots
|
|
||||||
---
|
|
||||||
@@ -1,52 +0,0 @@
|
|||||||
---
|
|
||||||
draft: true
|
|
||||||
id: 2
|
|
||||||
title: Common datasets snapshots
|
|
||||||
weight: 1
|
|
||||||
---
|
|
||||||
|
|
||||||
Understanding that creating embeddings every time can be a resource-intensive task, we have
|
|
||||||
devised a more efficient solution for you. In this section, we will regularly publish
|
|
||||||
snapshots of common public datasets often used for general or educational purposes. These
|
|
||||||
snapshots contain pre-computed vectors, critical for semantic search, that you can easily
|
|
||||||
import into your Qdrant instance. Our objective is to streamline your process and accelerate
|
|
||||||
your progress. Say goodbye to redundant operations and harness the power of Qdrant with
|
|
||||||
a simple [snapshot import](/documentation/concepts/snapshots/).
|
|
||||||
|
|
||||||
<table>
|
|
||||||
<thead>
|
|
||||||
<tr>
|
|
||||||
<th>Dataset</th>
|
|
||||||
<th>Description</th>
|
|
||||||
<th>Model</th>
|
|
||||||
<th>Dimensionality</th>
|
|
||||||
<th>Size</th>
|
|
||||||
<th>Link</th>
|
|
||||||
</tr>
|
|
||||||
</thead>
|
|
||||||
<tbody>
|
|
||||||
<tr>
|
|
||||||
<th rowspan="2">Arxiv.org</th>
|
|
||||||
<th>Only titles</th>
|
|
||||||
<td><a href="https://huggingface.co/hkunlp/instructor-xl">InstructorXL</a></td>
|
|
||||||
<td>768</td>
|
|
||||||
<td>7.1 GB</td>
|
|
||||||
<td>
|
|
||||||
<a href="https://storage.googleapis.com/common-datasets-snapshots/arxiv_titles-3083016565637815127-2023-05-29-13-56-22.snapshot">
|
|
||||||
<img src="/images/icons/download.svg" alt="download" />
|
|
||||||
</a>
|
|
||||||
</td>
|
|
||||||
</tr>
|
|
||||||
<tr>
|
|
||||||
<th>Only abstracts</th>
|
|
||||||
<td><a href="https://huggingface.co/hkunlp/instructor-xl">InstructorXL</a></td>
|
|
||||||
<td>768</td>
|
|
||||||
<td>8.4 GB</td>
|
|
||||||
<td>
|
|
||||||
<a href="https://storage.googleapis.com/common-datasets-snapshots/arxiv_abstracts-3083016565637815127-2023-06-02-07-26-29.snapshot">
|
|
||||||
<img src="/images/icons/download.svg" alt="download" />
|
|
||||||
</a>
|
|
||||||
</td>
|
|
||||||
</tr>
|
|
||||||
</tbody>
|
|
||||||
</table>
|
|
||||||
@@ -20,6 +20,10 @@ There are three ways to use Qdrant:
|
|||||||
|
|
||||||
First, try Qdrant locally using the [Qdrant Client](https://github.com/qdrant/qdrant-client) and with the help of our [Tutorials](tutorials/) and Guides. Develop a sample app from our [Examples](examples/) list and try it using a [Qdrant Docker](guides/installation/) container. Then, when you are ready for production, deploy to a Free Tier [Qdrant Cloud](cloud/) cluster.
|
First, try Qdrant locally using the [Qdrant Client](https://github.com/qdrant/qdrant-client) and with the help of our [Tutorials](tutorials/) and Guides. Develop a sample app from our [Examples](examples/) list and try it using a [Qdrant Docker](guides/installation/) container. Then, when you are ready for production, deploy to a Free Tier [Qdrant Cloud](cloud/) cluster.
|
||||||
|
|
||||||
|
### Try Qdrant with Practice Data:
|
||||||
|
|
||||||
|
You may always use our [Practice Datasets](datasets/) to build with Qdrant. This page will be regularly updated with dataset snapshots you can use to bootstrap complete projects.
|
||||||
|
|
||||||
## Popular Topics:
|
## Popular Topics:
|
||||||
|
|
||||||
| Tutorial | Description | Tutorial| Description |
|
| Tutorial | Description | Tutorial| Description |
|
||||||
|
|||||||
@@ -0,0 +1,194 @@
|
|||||||
|
---
|
||||||
|
title: Practice Datasets
|
||||||
|
weight: 41
|
||||||
|
---
|
||||||
|
|
||||||
|
# Common Datasets in Snapshot Format
|
||||||
|
|
||||||
|
You may find that creating embeddings from datasets is a very resource-intensive task.
|
||||||
|
If you need a practice dataset, feel free to pick one of the ready-made snapshots on this page.
|
||||||
|
These snapshots contain pre-computed vectors that you can easily import into your Qdrant instance.
|
||||||
|
|
||||||
|
## Available datasets
|
||||||
|
|
||||||
|
Our snapshots are usually generated from publicly available datasets, which are often used for
|
||||||
|
non-commercial or academic purposes. The following datasets are currently available. Please click
|
||||||
|
on a dataset name to see its detailed description.
|
||||||
|
|
||||||
|
| Dataset | Model | Vector size | Documents | Size | Qdrant snapshot | HF Hub |
|
||||||
|
|--------------------------------------------|-----------------------------------------------------------------------------|-------------|-----------|--------|----------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------|
|
||||||
|
| [Arxiv.org titles](#arxivorg-titles) | [InstructorXL](https://huggingface.co/hkunlp/instructor-xl) | 768 | 2.3M | 7.1 GB | [Download](https://snapshots.qdrant.io/arxiv_titles-3083016565637815127-2023-05-29-13-56-22.snapshot) | [Open](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings) |
|
||||||
|
| [Arxiv.org abstracts](#arxivorg-abstracts) | [InstructorXL](https://huggingface.co/hkunlp/instructor-xl) | 768 | 2.3M | 8.4 GB | [Download](https://snapshots.qdrant.io/arxiv_abstracts-3083016565637815127-2023-06-02-07-26-29.snapshot) | [Open](https://huggingface.co/datasets/Qdrant/arxiv-abstracts-instructorxl-embeddings) |
|
||||||
|
| [Wolt food](#wolt-food) | [clip-ViT-B-32](https://huggingface.co/sentence-transformers/clip-ViT-B-32) | 512 | 1.7M | 6.1 GB | [Download](https://snapshots.qdrant.io/wolt-clip-ViT-B-32-2446808438011867-2023-12-14-15-55-26.snapshot) | Not available |
|
||||||
|
|
||||||
|
Once you download a snapshot, you need to [restore it](/documentation/concepts/snapshots/#restore-snapshot)
|
||||||
|
using the Qdrant CLI upon startup or through the API.
|
||||||
|
|
||||||
|
## Qdrant on Hugging Face
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<a href="https://huggingface.co/Qdrant">
|
||||||
|
<img style="max-width: 500px" src="/content/images/hf-logo-with-title.svg" alt="HuggingFace" title="HuggingFace">
|
||||||
|
</a>
|
||||||
|
</p>
|
||||||
|
|
||||||
|
[Hugging Face](https://huggingface.co/) provides a platform for sharing and using ML models and
|
||||||
|
datasets. [Qdrant](https://huggingface.co/Qdrant) is one of the organizations there! We aim to
|
||||||
|
provide you with datasets containing neural embeddings that you can use to practice with Qdrant
|
||||||
|
and build your applications based on semantic search. **Please let us know if you'd like to see
|
||||||
|
a specific dataset!**
|
||||||
|
|
||||||
|
If you are not familiar with [Hugging Face datasets](https://huggingface.co/docs/datasets/index),
|
||||||
|
or would like to know how to combine it with Qdrant, please refer to the [tutorial](/documentation/tutorials/huggingface-datasets/).
|
||||||
|
|
||||||
|
## Arxiv.org
|
||||||
|
|
||||||
|
[Arxiv.org](https://arxiv.org) is a highly-regarded open-access repository of electronic preprints in multiple
|
||||||
|
fields. Operated by Cornell University, arXiv allows researchers to share their findings with
|
||||||
|
the scientific community and receive feedback before they undergo peer review for formal
|
||||||
|
publication. Its archives host millions of scholarly articles, making it an invaluable resource
|
||||||
|
for those looking to explore the cutting edge of scientific research. With a high frequency of
|
||||||
|
daily submissions from scientists around the world, arXiv forms a comprehensive, evolving dataset
|
||||||
|
that is ripe for mining, analysis, and the development of future innovations.
|
||||||
|
|
||||||
|
<aside role="status">
|
||||||
|
Arxiv.org snapshots were created using precomputed embeddings exposed by <a href="https://alex.macrocosm.so/download">the Alexandria Index</a>.
|
||||||
|
</aside>
|
||||||
|
|
||||||
|
### Arxiv.org titles
|
||||||
|
|
||||||
|
This dataset contains embeddings generated from the paper titles only. Each vector has a
|
||||||
|
payload with the title used to create it, along with the DOI (Digital Object Identifier).
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"title": "Nash Social Welfare for Indivisible Items under Separable, Piecewise-Linear Concave Utilities",
|
||||||
|
"DOI": "1612.05191"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The embeddings generated with InstructorXL model have been generated using the following
|
||||||
|
instruction:
|
||||||
|
|
||||||
|
> Represent the Research Paper title for retrieval; Input:
|
||||||
|
|
||||||
|
The following code snippet shows how to generate embeddings using the InstructorXL model:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from InstructorEmbedding import INSTRUCTOR
|
||||||
|
|
||||||
|
model = INSTRUCTOR("hkunlp/instructor-xl")
|
||||||
|
sentence = "3D ActionSLAM: wearable person tracking in multi-floor environments"
|
||||||
|
instruction = "Represent the Research Paper title for retrieval; Input:"
|
||||||
|
embeddings = model.encode([[instruction, sentence]])
|
||||||
|
```
|
||||||
|
|
||||||
|
The snapshot of the dataset might be downloaded [here](https://snapshots.qdrant.io/arxiv_titles-3083016565637815127-2023-05-29-13-56-22.snapshot).
|
||||||
|
|
||||||
|
#### Importing the dataset
|
||||||
|
|
||||||
|
The easiest way to use the provided dataset is to recover it via the API by passing the
|
||||||
|
URL as a location. It works also in [Qdrant Cloud](https://cloud.qdrant.io/). The following
|
||||||
|
code snippet shows how to create a new collection and fill it with the snapshot data:
|
||||||
|
|
||||||
|
```http request
|
||||||
|
PUT /collections/{collection_name}/snapshots/recover
|
||||||
|
{
|
||||||
|
"location": "https://snapshots.qdrant.io/arxiv_titles-3083016565637815127-2023-05-29-13-56-22.snapshot"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Arxiv.org abstracts
|
||||||
|
|
||||||
|
This dataset contains embeddings generated from the paper abstracts. Each vector has a
|
||||||
|
payload with the abstract used to create it, along with the DOI (Digital Object Identifier).
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"abstract": "Recently Cole and Gkatzelis gave the first constant factor approximation\nalgorithm for the problem of allocating indivisible items to agents, under\nadditive valuations, so as to maximize the Nash Social Welfare. We give\nconstant factor algorithms for a substantial generalization of their problem --\nto the case of separable, piecewise-linear concave utility functions. We give\ntwo such algorithms, the first using market equilibria and the second using the\ntheory of stable polynomials.\n In AGT, there is a paucity of methods for the design of mechanisms for the\nallocation of indivisible goods and the result of Cole and Gkatzelis seemed to\nbe taking a major step towards filling this gap. Our result can be seen as\nanother step in this direction.\n",
|
||||||
|
"DOI": "1612.05191"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The embeddings generated with InstructorXL model have been generated using the following
|
||||||
|
instruction:
|
||||||
|
|
||||||
|
> Represent the Research Paper abstract for retrieval; Input:
|
||||||
|
|
||||||
|
The following code snippet shows how to generate embeddings using the InstructorXL model:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from InstructorEmbedding import INSTRUCTOR
|
||||||
|
|
||||||
|
model = INSTRUCTOR("hkunlp/instructor-xl")
|
||||||
|
sentence = "The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train."
|
||||||
|
instruction = "Represent the Research Paper abstract for retrieval; Input:"
|
||||||
|
embeddings = model.encode([[instruction, sentence]])
|
||||||
|
```
|
||||||
|
|
||||||
|
The snapshot of the dataset might be downloaded [here](https://snapshots.qdrant.io/arxiv_abstracts-3083016565637815127-2023-06-02-07-26-29.snapshot).
|
||||||
|
|
||||||
|
#### Importing the dataset
|
||||||
|
|
||||||
|
The easiest way to use the provided dataset is to recover it via the API by passing the
|
||||||
|
URL as a location. It works also in [Qdrant Cloud](https://cloud.qdrant.io/). The following
|
||||||
|
code snippet shows how to create a new collection and fill it with the snapshot data:
|
||||||
|
|
||||||
|
```http request
|
||||||
|
PUT /collections/{collection_name}/snapshots/recover
|
||||||
|
{
|
||||||
|
"location": "https://snapshots.qdrant.io/arxiv_abstracts-3083016565637815127-2023-06-02-07-26-29.snapshot"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Wolt food
|
||||||
|
|
||||||
|
Our [Food Discovery demo](https://food-discovery.qdrant.tech/) relies on the dataset of
|
||||||
|
food images from the Wolt app. Each point in the collection represents a dish with a single
|
||||||
|
image. The image is represented as a vector of 512 float numbers. There is also a JSON
|
||||||
|
payload attached to each point, which looks similar to this:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"cafe": {
|
||||||
|
"address": "VGX7+6R2 Vecchia Napoli, Valletta",
|
||||||
|
"categories": ["italian", "pasta", "pizza", "burgers", "mediterranean"],
|
||||||
|
"location": {"lat": 35.8980154, "lon": 14.5145106},
|
||||||
|
"menu_id": "610936a4ee8ea7a56f4a372a",
|
||||||
|
"name": "Vecchia Napoli Is-Suq Tal-Belt",
|
||||||
|
"rating": 9,
|
||||||
|
"slug": "vecchia-napoli-skyparks-suq-tal-belt"
|
||||||
|
},
|
||||||
|
"description": "Tomato sauce, mozzarella fior di latte, crispy guanciale, Pecorino Romano cheese and a hint of chilli",
|
||||||
|
"image": "https://wolt-menu-images-cdn.wolt.com/menu-images/610936a4ee8ea7a56f4a372a/005dfeb2-e734-11ec-b667-ced7a78a5abd_l_amatriciana_pizza_joel_gueller1.jpeg",
|
||||||
|
"name": "L'Amatriciana"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The embeddings generated with clip-ViT-B-32 model have been generated using the following
|
||||||
|
code snippet:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from PIL import Image
|
||||||
|
from sentence_transformers import SentenceTransformer
|
||||||
|
|
||||||
|
image_path = "5dbfd216-5cce-11eb-8122-de94874ad1c8_ns_takeaway_seelachs_ei_baguette.jpeg"
|
||||||
|
|
||||||
|
model = SentenceTransformer("clip-ViT-B-32")
|
||||||
|
embedding = model.encode(Image.open(image_path))
|
||||||
|
```
|
||||||
|
|
||||||
|
The snapshot of the dataset might be downloaded [here](https://snapshots.qdrant.io/wolt-clip-ViT-B-32-2446808438011867-2023-12-14-15-55-26.snapshot).
|
||||||
|
|
||||||
|
#### Importing the dataset
|
||||||
|
|
||||||
|
The easiest way to use the provided dataset is to recover it via the API by passing the
|
||||||
|
URL as a location. It works also in [Qdrant Cloud](https://cloud.qdrant.io/). The following
|
||||||
|
code snippet shows how to create a new collection and fill it with the snapshot data:
|
||||||
|
|
||||||
|
```http request
|
||||||
|
PUT /collections/{collection_name}/snapshots/recover
|
||||||
|
{
|
||||||
|
"location": "https://snapshots.qdrant.io/wolt-clip-ViT-B-32-2446808438011867-2023-12-14-15-55-26.snapshot"
|
||||||
|
}
|
||||||
|
```
|
||||||
@@ -24,6 +24,5 @@ These tutorials demonstrate different ways you can build vector search into your
|
|||||||
| [Mighty Semantic Search](../tutorials/mighty/) | Build a simple semantic search with an on-demand NLP service. | Qdrant, Mighty |
|
| [Mighty Semantic Search](../tutorials/mighty/) | Build a simple semantic search with an on-demand NLP service. | Qdrant, Mighty |
|
||||||
| [Asynchronous API](../tutorials/async-api/) | Communicate with Qdrant server asynchronously with Python SDK. | Qdrant, Python |
|
| [Asynchronous API](../tutorials/async-api/) | Communicate with Qdrant server asynchronously with Python SDK. | Qdrant, Python |
|
||||||
| [Multitenancy with LlamaIndex](../tutorials/llama-index-multitenancy/) | Handle data coming from multiple users in LlamaIndex. | Qdrant, Python, LlamaIndex |
|
| [Multitenancy with LlamaIndex](../tutorials/llama-index-multitenancy/) | Handle data coming from multiple users in LlamaIndex. | Qdrant, Python, LlamaIndex |
|
||||||
| [Troubleshooting](../tutorials/common-errors/) | Solutions to common errors and fixes | Qdrant |
|
| [HuggingFace datasets](../tutorials/huggingface-datasets/) | Load a Hugging Face dataset to Qdrant | Qdrant, Python, datasets |
|
||||||
|
| [Troubleshooting](../tutorials/common-errors/) | Solutions to common errors and fixes | Qdrant |
|
||||||
|
|
||||||
|
|||||||
@@ -104,7 +104,7 @@ qdrant_client.recreate_collection(
|
|||||||
vectors_config=VectorParams(
|
vectors_config=VectorParams(
|
||||||
size=len(vectors[0]),
|
size=len(vectors[0]),
|
||||||
distance=Distance.COSINE,
|
distance=Distance.COSINE,
|
||||||
)
|
),
|
||||||
)
|
)
|
||||||
qdrant_client.upsert(
|
qdrant_client.upsert(
|
||||||
collection_name="COCO",
|
collection_name="COCO",
|
||||||
@@ -112,7 +112,7 @@ qdrant_client.upsert(
|
|||||||
ids=ids,
|
ids=ids,
|
||||||
vectors=vectors,
|
vectors=vectors,
|
||||||
payloads=payloads,
|
payloads=payloads,
|
||||||
)
|
),
|
||||||
)
|
)
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,115 @@
|
|||||||
|
---
|
||||||
|
title: Load Hugging Face dataset
|
||||||
|
weight: 19
|
||||||
|
---
|
||||||
|
|
||||||
|
# Loading a dataset from Hugging Face hub
|
||||||
|
|
||||||
|
[Hugging Face](https://huggingface.co/) provides a platform for sharing and using ML models and
|
||||||
|
datasets. [Qdrant](https://huggingface.co/Qdrant) also publishes datasets along with the
|
||||||
|
embeddings that you can use to practice with Qdrant and build your applications based on semantic
|
||||||
|
search. **Please [let us know](https://qdrant.to/discord) if you'd like to see a specific dataset!**
|
||||||
|
|
||||||
|
## arxiv-titles-instructorxl-embeddings
|
||||||
|
|
||||||
|
[This dataset](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings) contains
|
||||||
|
embeddings generated from the paper titles only. Each vector has a payload with the title used to
|
||||||
|
create it, along with the DOI (Digital Object Identifier).
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"title": "Nash Social Welfare for Indivisible Items under Separable, Piecewise-Linear Concave Utilities",
|
||||||
|
"DOI": "1612.05191"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
You can find a detailed description of the dataset in the [Practice Datasets](/documentation/datasets/#journal-article-titles)
|
||||||
|
section. If you prefer loading the dataset from a Qdrant snapshot, it also linked there.
|
||||||
|
|
||||||
|
Loading the dataset is as simple as using the `load_dataset` function from the `datasets` library:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from datasets import load_dataset
|
||||||
|
|
||||||
|
dataset = load_dataset("Qdrant/arxiv-titles-instructorxl-embeddings")
|
||||||
|
```
|
||||||
|
|
||||||
|
<aside role="status">The dataset has over 16 GB, so it might take a while to download.</aside>
|
||||||
|
|
||||||
|
The dataset contains 2,250,000 vectors. This is how you can check the list of the features in the dataset:
|
||||||
|
|
||||||
|
```python
|
||||||
|
dataset.features
|
||||||
|
```
|
||||||
|
|
||||||
|
### Streaming the dataset
|
||||||
|
|
||||||
|
Dataset streaming lets you work with a dataset without downloading it. The data is streamed as
|
||||||
|
you iterate over the dataset. You can read more about it in the [Hugging Face
|
||||||
|
documentation](https://huggingface.co/docs/datasets/stream).
|
||||||
|
|
||||||
|
```python
|
||||||
|
from datasets import load_dataset
|
||||||
|
|
||||||
|
dataset = load_dataset(
|
||||||
|
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Loading the dataset into Qdrant
|
||||||
|
|
||||||
|
You can load the dataset into Qdrant using the [Python SDK](https://github.com/qdrant/qdrant-client).
|
||||||
|
The embeddings are already precomputed, so you can store them in a collection, that we're going
|
||||||
|
to create in a second:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
|
client = QdrantClient("http://localhost:6333")
|
||||||
|
|
||||||
|
client.create_collection(
|
||||||
|
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||||
|
vectors_config=models.VectorParams(
|
||||||
|
size=768,
|
||||||
|
distance=models.Distance.COSINE,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
It is always a good idea to use batching, while loading a large dataset, so let's do that.
|
||||||
|
We are going to need a helper function to split the dataset into batches:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from itertools import islice
|
||||||
|
|
||||||
|
def batched(iterable, n):
|
||||||
|
iterator = iter(iterable)
|
||||||
|
while batch := list(islice(iterator, n)):
|
||||||
|
yield batch
|
||||||
|
```
|
||||||
|
|
||||||
|
If you are a happy user of Python 3.12+, you can use the [`batched` function from the `itertools`
|
||||||
|
](https://docs.python.org/3/library/itertools.html#itertools.batched) package instead.
|
||||||
|
|
||||||
|
No matter what Python version you are using, you can use the `upsert` method to load the dataset,
|
||||||
|
batch by batch, into Qdrant:
|
||||||
|
|
||||||
|
```python
|
||||||
|
batch_size = 100
|
||||||
|
|
||||||
|
for batch in batched(dataset, batch_size):
|
||||||
|
ids = [point.pop("id") for point in batch]
|
||||||
|
vectors = [point.pop("vector") for point in batch]
|
||||||
|
|
||||||
|
client.upsert(
|
||||||
|
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||||
|
points=models.Batch(
|
||||||
|
ids=ids,
|
||||||
|
vectors=vectors,
|
||||||
|
payloads=batch,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
Your collection is ready to be used for search! Please [let us know using Discord](https://qdrant.to/discord)
|
||||||
|
if you would like to see more datasets published on Hugging Face hub.
|
||||||
@@ -214,16 +214,16 @@ class NeuralSearcher:
|
|||||||
|
|
||||||
```python
|
```python
|
||||||
def search(self, text: str):
|
def search(self, text: str):
|
||||||
search_result = self.qdrant_client.query(
|
search_result = self.qdrant_client.query(
|
||||||
collection_name=self.collection_name,
|
collection_name=self.collection_name,
|
||||||
query_text=text,
|
query_text=text,
|
||||||
query_filter=None, # If you don't want any filters for now
|
query_filter=None, # If you don't want any filters for now
|
||||||
limit=5 # 5 the most closest results is enough
|
limit=5, # 5 the closest results are enough
|
||||||
)
|
)
|
||||||
# `search_result` contains found vector ids with similarity scores along with the stored payload
|
# `search_result` contains found vector ids with similarity scores along with the stored payload
|
||||||
# In this function you are interested in payload only
|
# In this function you are interested in payload only
|
||||||
metadata = [hit.metadata for hit in search_result]
|
metadata = [hit.metadata for hit in search_result]
|
||||||
return metadata
|
return metadata
|
||||||
```
|
```
|
||||||
|
|
||||||
3. Add search filters.
|
3. Add search filters.
|
||||||
@@ -286,17 +286,17 @@ from neural_searcher import NeuralSearcher
|
|||||||
app = FastAPI()
|
app = FastAPI()
|
||||||
|
|
||||||
# Create a neural searcher instance
|
# Create a neural searcher instance
|
||||||
neural_searcher = NeuralSearcher(collection_name='startups')
|
neural_searcher = NeuralSearcher(collection_name="startups")
|
||||||
|
|
||||||
|
|
||||||
@app.get("/api/search")
|
@app.get("/api/search")
|
||||||
def search_startup(q: str):
|
def search_startup(q: str):
|
||||||
return {
|
return {"result": neural_searcher.search(text=q)}
|
||||||
"result": neural_searcher.search(text=q)
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
import uvicorn
|
import uvicorn
|
||||||
|
|
||||||
uvicorn.run(app, host="0.0.0.0", port=8000)
|
uvicorn.run(app, host="0.0.0.0", port=8000)
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -223,20 +223,20 @@ class NeuralSearcher:
|
|||||||
|
|
||||||
```python
|
```python
|
||||||
def search(self, text: str):
|
def search(self, text: str):
|
||||||
# Convert text query into vector
|
# Convert text query into vector
|
||||||
vector = self.model.encode(text).tolist()
|
vector = self.model.encode(text).tolist()
|
||||||
|
|
||||||
# Use `vector` for search for closest vectors in the collection
|
# Use `vector` for search for closest vectors in the collection
|
||||||
search_result = self.qdrant_client.search(
|
search_result = self.qdrant_client.search(
|
||||||
collection_name=self.collection_name,
|
collection_name=self.collection_name,
|
||||||
query_vector=vector,
|
query_vector=vector,
|
||||||
query_filter=None, # If you don't want any filters for now
|
query_filter=None, # If you don't want any filters for now
|
||||||
limit=5 # 5 the most closest results is enough
|
limit=5, # 5 the most closest results is enough
|
||||||
)
|
)
|
||||||
# `search_result` contains found vector ids with similarity scores along with the stored payload
|
# `search_result` contains found vector ids with similarity scores along with the stored payload
|
||||||
# In this function you are interested in payload only
|
# In this function you are interested in payload only
|
||||||
payloads = [hit.payload for hit in search_result]
|
payloads = [hit.payload for hit in search_result]
|
||||||
return payloads
|
return payloads
|
||||||
```
|
```
|
||||||
|
|
||||||
3. Add search filters.
|
3. Add search filters.
|
||||||
@@ -299,17 +299,17 @@ from neural_searcher import NeuralSearcher
|
|||||||
app = FastAPI()
|
app = FastAPI()
|
||||||
|
|
||||||
# Create a neural searcher instance
|
# Create a neural searcher instance
|
||||||
neural_searcher = NeuralSearcher(collection_name='startups')
|
neural_searcher = NeuralSearcher(collection_name="startups")
|
||||||
|
|
||||||
|
|
||||||
@app.get("/api/search")
|
@app.get("/api/search")
|
||||||
def search_startup(q: str):
|
def search_startup(q: str):
|
||||||
return {
|
return {"result": neural_searcher.search(text=q)}
|
||||||
"result": neural_searcher.search(text=q)
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
import uvicorn
|
import uvicorn
|
||||||
|
|
||||||
uvicorn.run(app, host="0.0.0.0", port=8000)
|
uvicorn.run(app, host="0.0.0.0", port=8000)
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
File diff suppressed because one or more lines are too long
|
After Width: | Height: | Size: 46 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 621 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 610 KiB |
Reference in New Issue
Block a user