Add "Practice datasets" section (#199)

* Add common datasets

* Add instruction for titles

* Add tutorial on how to use Hugging Face datasets

* Add an information about Qdrant org at Hugging Face

* Add links to Hugging Face datasets

* Add preview images

* Add CTA (Discord)

* Fix typo

* Add link to Discord

* List available datasets on top of the page

* Fix some typos and improve intro

* Fix some typos and improve intro

* Add fixed snapshot links

* Add Wolt food dataset

* Add API calls to restore the snapshots

* Adjust the weight of the section

* Reorder tutorials
This commit is contained in:
Kacper Łukawski
2024-01-02 16:24:09 +01:00
committed by GitHub
parent f9867f6840
commit bc20106b7d
13 changed files with 357 additions and 101 deletions
-1
View File
@@ -188,7 +188,6 @@ keywords = "search engine, vector database, neural network, matching, filter, Sa
[menu.main.params] [menu.main.params]
external = true external = true
[[menu.main]] [[menu.main]]
identifier = "community" identifier = "community"
name = "Community" name = "Community"
-12
View File
@@ -1,12 +0,0 @@
---
title: Common Datasets
draft: true
description: Ready snapshots of some standard datasets which you can easily import into Qdrant.
keywords:
- embeddings
- vector embeddings
- vectors
- precomputed embeddings
- datasets
- dataset snapshots
---
@@ -1,52 +0,0 @@
---
draft: true
id: 2
title: Common datasets snapshots
weight: 1
---
Understanding that creating embeddings every time can be a resource-intensive task, we have
devised a more efficient solution for you. In this section, we will regularly publish
snapshots of common public datasets often used for general or educational purposes. These
snapshots contain pre-computed vectors, critical for semantic search, that you can easily
import into your Qdrant instance. Our objective is to streamline your process and accelerate
your progress. Say goodbye to redundant operations and harness the power of Qdrant with
a simple [snapshot import](/documentation/concepts/snapshots/).
<table>
<thead>
<tr>
<th>Dataset</th>
<th>Description</th>
<th>Model</th>
<th>Dimensionality</th>
<th>Size</th>
<th>Link</th>
</tr>
</thead>
<tbody>
<tr>
<th rowspan="2">Arxiv.org</th>
<th>Only titles</th>
<td><a href="https://huggingface.co/hkunlp/instructor-xl">InstructorXL</a></td>
<td>768</td>
<td>7.1 GB</td>
<td>
<a href="https://storage.googleapis.com/common-datasets-snapshots/arxiv_titles-3083016565637815127-2023-05-29-13-56-22.snapshot">
<img src="/images/icons/download.svg" alt="download" />
</a>
</td>
</tr>
<tr>
<th>Only abstracts</th>
<td><a href="https://huggingface.co/hkunlp/instructor-xl">InstructorXL</a></td>
<td>768</td>
<td>8.4 GB</td>
<td>
<a href="https://storage.googleapis.com/common-datasets-snapshots/arxiv_abstracts-3083016565637815127-2023-06-02-07-26-29.snapshot">
<img src="/images/icons/download.svg" alt="download" />
</a>
</td>
</tr>
</tbody>
</table>
@@ -20,6 +20,10 @@ There are three ways to use Qdrant:
First, try Qdrant locally using the [Qdrant Client](https://github.com/qdrant/qdrant-client) and with the help of our [Tutorials](tutorials/) and Guides. Develop a sample app from our [Examples](examples/) list and try it using a [Qdrant Docker](guides/installation/) container. Then, when you are ready for production, deploy to a Free Tier [Qdrant Cloud](cloud/) cluster. First, try Qdrant locally using the [Qdrant Client](https://github.com/qdrant/qdrant-client) and with the help of our [Tutorials](tutorials/) and Guides. Develop a sample app from our [Examples](examples/) list and try it using a [Qdrant Docker](guides/installation/) container. Then, when you are ready for production, deploy to a Free Tier [Qdrant Cloud](cloud/) cluster.
### Try Qdrant with Practice Data:
You may always use our [Practice Datasets](datasets/) to build with Qdrant. This page will be regularly updated with dataset snapshots you can use to bootstrap complete projects.
## Popular Topics: ## Popular Topics:
| Tutorial | Description | Tutorial| Description | | Tutorial | Description | Tutorial| Description |
@@ -0,0 +1,194 @@
---
title: Practice Datasets
weight: 41
---
# Common Datasets in Snapshot Format
You may find that creating embeddings from datasets is a very resource-intensive task.
If you need a practice dataset, feel free to pick one of the ready-made snapshots on this page.
These snapshots contain pre-computed vectors that you can easily import into your Qdrant instance.
## Available datasets
Our snapshots are usually generated from publicly available datasets, which are often used for
non-commercial or academic purposes. The following datasets are currently available. Please click
on a dataset name to see its detailed description.
| Dataset | Model | Vector size | Documents | Size | Qdrant snapshot | HF Hub |
|--------------------------------------------|-----------------------------------------------------------------------------|-------------|-----------|--------|----------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------|
| [Arxiv.org titles](#arxivorg-titles) | [InstructorXL](https://huggingface.co/hkunlp/instructor-xl) | 768 | 2.3M | 7.1 GB | [Download](https://snapshots.qdrant.io/arxiv_titles-3083016565637815127-2023-05-29-13-56-22.snapshot) | [Open](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings) |
| [Arxiv.org abstracts](#arxivorg-abstracts) | [InstructorXL](https://huggingface.co/hkunlp/instructor-xl) | 768 | 2.3M | 8.4 GB | [Download](https://snapshots.qdrant.io/arxiv_abstracts-3083016565637815127-2023-06-02-07-26-29.snapshot) | [Open](https://huggingface.co/datasets/Qdrant/arxiv-abstracts-instructorxl-embeddings) |
| [Wolt food](#wolt-food) | [clip-ViT-B-32](https://huggingface.co/sentence-transformers/clip-ViT-B-32) | 512 | 1.7M | 6.1 GB | [Download](https://snapshots.qdrant.io/wolt-clip-ViT-B-32-2446808438011867-2023-12-14-15-55-26.snapshot) | Not available |
Once you download a snapshot, you need to [restore it](/documentation/concepts/snapshots/#restore-snapshot)
using the Qdrant CLI upon startup or through the API.
## Qdrant on Hugging Face
<p align="center">
<a href="https://huggingface.co/Qdrant">
<img style="max-width: 500px" src="/content/images/hf-logo-with-title.svg" alt="HuggingFace" title="HuggingFace">
</a>
</p>
[Hugging Face](https://huggingface.co/) provides a platform for sharing and using ML models and
datasets. [Qdrant](https://huggingface.co/Qdrant) is one of the organizations there! We aim to
provide you with datasets containing neural embeddings that you can use to practice with Qdrant
and build your applications based on semantic search. **Please let us know if you'd like to see
a specific dataset!**
If you are not familiar with [Hugging Face datasets](https://huggingface.co/docs/datasets/index),
or would like to know how to combine it with Qdrant, please refer to the [tutorial](/documentation/tutorials/huggingface-datasets/).
## Arxiv.org
[Arxiv.org](https://arxiv.org) is a highly-regarded open-access repository of electronic preprints in multiple
fields. Operated by Cornell University, arXiv allows researchers to share their findings with
the scientific community and receive feedback before they undergo peer review for formal
publication. Its archives host millions of scholarly articles, making it an invaluable resource
for those looking to explore the cutting edge of scientific research. With a high frequency of
daily submissions from scientists around the world, arXiv forms a comprehensive, evolving dataset
that is ripe for mining, analysis, and the development of future innovations.
<aside role="status">
Arxiv.org snapshots were created using precomputed embeddings exposed by <a href="https://alex.macrocosm.so/download">the Alexandria Index</a>.
</aside>
### Arxiv.org titles
This dataset contains embeddings generated from the paper titles only. Each vector has a
payload with the title used to create it, along with the DOI (Digital Object Identifier).
```json
{
"title": "Nash Social Welfare for Indivisible Items under Separable, Piecewise-Linear Concave Utilities",
"DOI": "1612.05191"
}
```
The embeddings generated with InstructorXL model have been generated using the following
instruction:
> Represent the Research Paper title for retrieval; Input:
The following code snippet shows how to generate embeddings using the InstructorXL model:
```python
from InstructorEmbedding import INSTRUCTOR
model = INSTRUCTOR("hkunlp/instructor-xl")
sentence = "3D ActionSLAM: wearable person tracking in multi-floor environments"
instruction = "Represent the Research Paper title for retrieval; Input:"
embeddings = model.encode([[instruction, sentence]])
```
The snapshot of the dataset might be downloaded [here](https://snapshots.qdrant.io/arxiv_titles-3083016565637815127-2023-05-29-13-56-22.snapshot).
#### Importing the dataset
The easiest way to use the provided dataset is to recover it via the API by passing the
URL as a location. It works also in [Qdrant Cloud](https://cloud.qdrant.io/). The following
code snippet shows how to create a new collection and fill it with the snapshot data:
```http request
PUT /collections/{collection_name}/snapshots/recover
{
"location": "https://snapshots.qdrant.io/arxiv_titles-3083016565637815127-2023-05-29-13-56-22.snapshot"
}
```
### Arxiv.org abstracts
This dataset contains embeddings generated from the paper abstracts. Each vector has a
payload with the abstract used to create it, along with the DOI (Digital Object Identifier).
```json
{
"abstract": "Recently Cole and Gkatzelis gave the first constant factor approximation\nalgorithm for the problem of allocating indivisible items to agents, under\nadditive valuations, so as to maximize the Nash Social Welfare. We give\nconstant factor algorithms for a substantial generalization of their problem --\nto the case of separable, piecewise-linear concave utility functions. We give\ntwo such algorithms, the first using market equilibria and the second using the\ntheory of stable polynomials.\n In AGT, there is a paucity of methods for the design of mechanisms for the\nallocation of indivisible goods and the result of Cole and Gkatzelis seemed to\nbe taking a major step towards filling this gap. Our result can be seen as\nanother step in this direction.\n",
"DOI": "1612.05191"
}
```
The embeddings generated with InstructorXL model have been generated using the following
instruction:
> Represent the Research Paper abstract for retrieval; Input:
The following code snippet shows how to generate embeddings using the InstructorXL model:
```python
from InstructorEmbedding import INSTRUCTOR
model = INSTRUCTOR("hkunlp/instructor-xl")
sentence = "The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train."
instruction = "Represent the Research Paper abstract for retrieval; Input:"
embeddings = model.encode([[instruction, sentence]])
```
The snapshot of the dataset might be downloaded [here](https://snapshots.qdrant.io/arxiv_abstracts-3083016565637815127-2023-06-02-07-26-29.snapshot).
#### Importing the dataset
The easiest way to use the provided dataset is to recover it via the API by passing the
URL as a location. It works also in [Qdrant Cloud](https://cloud.qdrant.io/). The following
code snippet shows how to create a new collection and fill it with the snapshot data:
```http request
PUT /collections/{collection_name}/snapshots/recover
{
"location": "https://snapshots.qdrant.io/arxiv_abstracts-3083016565637815127-2023-06-02-07-26-29.snapshot"
}
```
## Wolt food
Our [Food Discovery demo](https://food-discovery.qdrant.tech/) relies on the dataset of
food images from the Wolt app. Each point in the collection represents a dish with a single
image. The image is represented as a vector of 512 float numbers. There is also a JSON
payload attached to each point, which looks similar to this:
```json
{
"cafe": {
"address": "VGX7+6R2 Vecchia Napoli, Valletta",
"categories": ["italian", "pasta", "pizza", "burgers", "mediterranean"],
"location": {"lat": 35.8980154, "lon": 14.5145106},
"menu_id": "610936a4ee8ea7a56f4a372a",
"name": "Vecchia Napoli Is-Suq Tal-Belt",
"rating": 9,
"slug": "vecchia-napoli-skyparks-suq-tal-belt"
},
"description": "Tomato sauce, mozzarella fior di latte, crispy guanciale, Pecorino Romano cheese and a hint of chilli",
"image": "https://wolt-menu-images-cdn.wolt.com/menu-images/610936a4ee8ea7a56f4a372a/005dfeb2-e734-11ec-b667-ced7a78a5abd_l_amatriciana_pizza_joel_gueller1.jpeg",
"name": "L'Amatriciana"
}
```
The embeddings generated with clip-ViT-B-32 model have been generated using the following
code snippet:
```python
from PIL import Image
from sentence_transformers import SentenceTransformer
image_path = "5dbfd216-5cce-11eb-8122-de94874ad1c8_ns_takeaway_seelachs_ei_baguette.jpeg"
model = SentenceTransformer("clip-ViT-B-32")
embedding = model.encode(Image.open(image_path))
```
The snapshot of the dataset might be downloaded [here](https://snapshots.qdrant.io/wolt-clip-ViT-B-32-2446808438011867-2023-12-14-15-55-26.snapshot).
#### Importing the dataset
The easiest way to use the provided dataset is to recover it via the API by passing the
URL as a location. It works also in [Qdrant Cloud](https://cloud.qdrant.io/). The following
code snippet shows how to create a new collection and fill it with the snapshot data:
```http request
PUT /collections/{collection_name}/snapshots/recover
{
"location": "https://snapshots.qdrant.io/wolt-clip-ViT-B-32-2446808438011867-2023-12-14-15-55-26.snapshot"
}
```
@@ -24,6 +24,5 @@ These tutorials demonstrate different ways you can build vector search into your
| [Mighty Semantic Search](../tutorials/mighty/) | Build a simple semantic search with an on-demand NLP service. | Qdrant, Mighty | | [Mighty Semantic Search](../tutorials/mighty/) | Build a simple semantic search with an on-demand NLP service. | Qdrant, Mighty |
| [Asynchronous API](../tutorials/async-api/) | Communicate with Qdrant server asynchronously with Python SDK. | Qdrant, Python | | [Asynchronous API](../tutorials/async-api/) | Communicate with Qdrant server asynchronously with Python SDK. | Qdrant, Python |
| [Multitenancy with LlamaIndex](../tutorials/llama-index-multitenancy/) | Handle data coming from multiple users in LlamaIndex. | Qdrant, Python, LlamaIndex | | [Multitenancy with LlamaIndex](../tutorials/llama-index-multitenancy/) | Handle data coming from multiple users in LlamaIndex. | Qdrant, Python, LlamaIndex |
| [Troubleshooting](../tutorials/common-errors/) | Solutions to common errors and fixes | Qdrant | | [HuggingFace datasets](../tutorials/huggingface-datasets/) | Load a Hugging Face dataset to Qdrant | Qdrant, Python, datasets |
| [Troubleshooting](../tutorials/common-errors/) | Solutions to common errors and fixes | Qdrant |
@@ -104,7 +104,7 @@ qdrant_client.recreate_collection(
vectors_config=VectorParams( vectors_config=VectorParams(
size=len(vectors[0]), size=len(vectors[0]),
distance=Distance.COSINE, distance=Distance.COSINE,
) ),
) )
qdrant_client.upsert( qdrant_client.upsert(
collection_name="COCO", collection_name="COCO",
@@ -112,7 +112,7 @@ qdrant_client.upsert(
ids=ids, ids=ids,
vectors=vectors, vectors=vectors,
payloads=payloads, payloads=payloads,
) ),
) )
``` ```
@@ -0,0 +1,115 @@
---
title: Load Hugging Face dataset
weight: 19
---
# Loading a dataset from Hugging Face hub
[Hugging Face](https://huggingface.co/) provides a platform for sharing and using ML models and
datasets. [Qdrant](https://huggingface.co/Qdrant) also publishes datasets along with the
embeddings that you can use to practice with Qdrant and build your applications based on semantic
search. **Please [let us know](https://qdrant.to/discord) if you'd like to see a specific dataset!**
## arxiv-titles-instructorxl-embeddings
[This dataset](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings) contains
embeddings generated from the paper titles only. Each vector has a payload with the title used to
create it, along with the DOI (Digital Object Identifier).
```json
{
"title": "Nash Social Welfare for Indivisible Items under Separable, Piecewise-Linear Concave Utilities",
"DOI": "1612.05191"
}
```
You can find a detailed description of the dataset in the [Practice Datasets](/documentation/datasets/#journal-article-titles)
section. If you prefer loading the dataset from a Qdrant snapshot, it also linked there.
Loading the dataset is as simple as using the `load_dataset` function from the `datasets` library:
```python
from datasets import load_dataset
dataset = load_dataset("Qdrant/arxiv-titles-instructorxl-embeddings")
```
<aside role="status">The dataset has over 16 GB, so it might take a while to download.</aside>
The dataset contains 2,250,000 vectors. This is how you can check the list of the features in the dataset:
```python
dataset.features
```
### Streaming the dataset
Dataset streaming lets you work with a dataset without downloading it. The data is streamed as
you iterate over the dataset. You can read more about it in the [Hugging Face
documentation](https://huggingface.co/docs/datasets/stream).
```python
from datasets import load_dataset
dataset = load_dataset(
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
)
```
### Loading the dataset into Qdrant
You can load the dataset into Qdrant using the [Python SDK](https://github.com/qdrant/qdrant-client).
The embeddings are already precomputed, so you can store them in a collection, that we're going
to create in a second:
```python
from qdrant_client import QdrantClient, models
client = QdrantClient("http://localhost:6333")
client.create_collection(
collection_name="arxiv-titles-instructorxl-embeddings",
vectors_config=models.VectorParams(
size=768,
distance=models.Distance.COSINE,
),
)
```
It is always a good idea to use batching, while loading a large dataset, so let's do that.
We are going to need a helper function to split the dataset into batches:
```python
from itertools import islice
def batched(iterable, n):
iterator = iter(iterable)
while batch := list(islice(iterator, n)):
yield batch
```
If you are a happy user of Python 3.12+, you can use the [`batched` function from the `itertools`
](https://docs.python.org/3/library/itertools.html#itertools.batched) package instead.
No matter what Python version you are using, you can use the `upsert` method to load the dataset,
batch by batch, into Qdrant:
```python
batch_size = 100
for batch in batched(dataset, batch_size):
ids = [point.pop("id") for point in batch]
vectors = [point.pop("vector") for point in batch]
client.upsert(
collection_name="arxiv-titles-instructorxl-embeddings",
points=models.Batch(
ids=ids,
vectors=vectors,
payloads=batch,
),
)
```
Your collection is ready to be used for search! Please [let us know using Discord](https://qdrant.to/discord)
if you would like to see more datasets published on Hugging Face hub.
@@ -214,16 +214,16 @@ class NeuralSearcher:
```python ```python
def search(self, text: str): def search(self, text: str):
search_result = self.qdrant_client.query( search_result = self.qdrant_client.query(
collection_name=self.collection_name, collection_name=self.collection_name,
query_text=text, query_text=text,
query_filter=None, # If you don't want any filters for now query_filter=None, # If you don't want any filters for now
limit=5 # 5 the most closest results is enough limit=5, # 5 the closest results are enough
) )
# `search_result` contains found vector ids with similarity scores along with the stored payload # `search_result` contains found vector ids with similarity scores along with the stored payload
# In this function you are interested in payload only # In this function you are interested in payload only
metadata = [hit.metadata for hit in search_result] metadata = [hit.metadata for hit in search_result]
return metadata return metadata
``` ```
3. Add search filters. 3. Add search filters.
@@ -286,17 +286,17 @@ from neural_searcher import NeuralSearcher
app = FastAPI() app = FastAPI()
# Create a neural searcher instance # Create a neural searcher instance
neural_searcher = NeuralSearcher(collection_name='startups') neural_searcher = NeuralSearcher(collection_name="startups")
@app.get("/api/search") @app.get("/api/search")
def search_startup(q: str): def search_startup(q: str):
return { return {"result": neural_searcher.search(text=q)}
"result": neural_searcher.search(text=q)
}
if __name__ == "__main__": if __name__ == "__main__":
import uvicorn import uvicorn
uvicorn.run(app, host="0.0.0.0", port=8000) uvicorn.run(app, host="0.0.0.0", port=8000)
``` ```
@@ -223,20 +223,20 @@ class NeuralSearcher:
```python ```python
def search(self, text: str): def search(self, text: str):
# Convert text query into vector # Convert text query into vector
vector = self.model.encode(text).tolist() vector = self.model.encode(text).tolist()
# Use `vector` for search for closest vectors in the collection # Use `vector` for search for closest vectors in the collection
search_result = self.qdrant_client.search( search_result = self.qdrant_client.search(
collection_name=self.collection_name, collection_name=self.collection_name,
query_vector=vector, query_vector=vector,
query_filter=None, # If you don't want any filters for now query_filter=None, # If you don't want any filters for now
limit=5 # 5 the most closest results is enough limit=5, # 5 the most closest results is enough
) )
# `search_result` contains found vector ids with similarity scores along with the stored payload # `search_result` contains found vector ids with similarity scores along with the stored payload
# In this function you are interested in payload only # In this function you are interested in payload only
payloads = [hit.payload for hit in search_result] payloads = [hit.payload for hit in search_result]
return payloads return payloads
``` ```
3. Add search filters. 3. Add search filters.
@@ -299,17 +299,17 @@ from neural_searcher import NeuralSearcher
app = FastAPI() app = FastAPI()
# Create a neural searcher instance # Create a neural searcher instance
neural_searcher = NeuralSearcher(collection_name='startups') neural_searcher = NeuralSearcher(collection_name="startups")
@app.get("/api/search") @app.get("/api/search")
def search_startup(q: str): def search_startup(q: str):
return { return {"result": neural_searcher.search(text=q)}
"result": neural_searcher.search(text=q)
}
if __name__ == "__main__": if __name__ == "__main__":
import uvicorn import uvicorn
uvicorn.run(app, host="0.0.0.0", port=8000) uvicorn.run(app, host="0.0.0.0", port=8000)
``` ```
File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 46 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 621 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 610 KiB