mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-30 00:18:32 +02:00
add content
This commit is contained in:
@@ -0,0 +1,22 @@
|
||||
---
|
||||
title: Vector Search Basics
|
||||
weight: 16
|
||||
# If the index.md file is empty, the link to the section will be hidden from the sidebar
|
||||
is_empty: false
|
||||
aliases:
|
||||
- how-to
|
||||
- tutorials
|
||||
partition: qdrant
|
||||
---
|
||||
|
||||
# Beginner Tutorials
|
||||
|
||||
These tutorials demonstrate different ways you can build vector search into your applications.
|
||||
|
||||
| Essential How-Tos | Description | Stack |
|
||||
|---------------------------------------------------------------------------------|-------------------------------------------------------------------|---------------------------------------------|
|
||||
| [Semantic Search for Beginners](/documentation/tutorials/search-beginners/) | Create a simple search engine locally in minutes. | Qdrant |
|
||||
| [Simple Neural Search](/documentation/tutorials/neural-search/) | Build and deploy a neural search that browses startup data. | Qdrant, BERT, FastAPI |
|
||||
| [Neural Search with FastEmbed](/documentation/tutorials/neural-search-fastembed/) | Build and deploy a neural search with our FastEmbed library. | Qdrant |
|
||||
| [Measure Retrieval Quality](/documentation/tutorials/retrieval-quality/) | Measure and fine-tune the retrieval quality | Qdrant, Python, datasets |
|
||||
|
||||
@@ -0,0 +1,384 @@
|
||||
---
|
||||
title: Hybrid Search with FastEmbed
|
||||
weight: 3
|
||||
|
||||
aliases:
|
||||
- /documentation/tutorials/neural-search-fastembed/
|
||||
---
|
||||
|
||||
# Create a Hybrid Search Service with Fastembed
|
||||
|
||||
| Time: 20 min | Level: Beginner | Output: [GitHub](https://github.com/qdrant/qdrant_demo/) |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
This tutorial shows you how to build and deploy your own hybrid search service to look through descriptions of companies from [startups-list.com](https://www.startups-list.com/) and pick the most similar ones to your query.
|
||||
The website contains the company names, descriptions, locations, and a picture for each entry.
|
||||
|
||||
As we have already written on our [blog](/articles/hybrid-search/), there is no single definition of hybrid search.
|
||||
In this tutorial we are covering the case with a combination of dense and [sparse embeddings](/articles/sparse-vectors/).
|
||||
The former ones refer to the embeddings generated by such well-known neural networks as BERT, while the latter ones are more related to a traditional full-text search approach.
|
||||
|
||||
Our hybrid search service will use [Fastembed](https://github.com/qdrant/fastembed) package to generate embeddings of text descriptions and [FastAPI](https://fastapi.tiangolo.com/) to serve the search API.
|
||||
Fastembed natively integrates with Qdrant client, so you can easily upload the data into Qdrant and perform search queries.
|
||||
|
||||

|
||||
|
||||
|
||||
## Workflow
|
||||
|
||||
To create a hybrid search service, you will need to transform your raw data and then create a search function to manipulate it.
|
||||
First, you will 1) download and prepare a sample dataset using a modified version of the BERT ML model. Then, you will 2) load the data into Qdrant, 3) create a hybrid search API and 4) serve it using FastAPI.
|
||||
|
||||

|
||||
|
||||
## Prerequisites
|
||||
|
||||
To complete this tutorial, you will need:
|
||||
|
||||
- Docker - The easiest way to use Qdrant is to run a pre-built Docker image.
|
||||
- [Raw parsed data](https://storage.googleapis.com/generall-shared-data/startups_demo.json) from startups-list.com.
|
||||
- Python version >=3.8
|
||||
|
||||
## Prepare sample dataset
|
||||
|
||||
To conduct a hybrid search on startup descriptions, you must first encode the description data into vectors.
|
||||
Fastembed integration into qdrant client combines encoding and uploading into a single step.
|
||||
|
||||
It also takes care of batching and parallelization, so you don't have to worry about it.
|
||||
|
||||
Let's start by downloading the data and installing the necessary packages.
|
||||
|
||||
|
||||
1. First you need to download the dataset.
|
||||
|
||||
```bash
|
||||
wget https://storage.googleapis.com/generall-shared-data/startups_demo.json
|
||||
```
|
||||
|
||||
## Run Qdrant in Docker
|
||||
|
||||
Next, you need to manage all of your data using a vector engine. Qdrant lets you store, update or delete created vectors. Most importantly, it lets you search for the nearest vectors via a convenient API.
|
||||
|
||||
> **Note:** Before you begin, create a project directory and a virtual python environment in it.
|
||||
|
||||
1. Download the Qdrant image from DockerHub.
|
||||
|
||||
```bash
|
||||
docker pull qdrant/qdrant
|
||||
```
|
||||
2. Start Qdrant inside of Docker.
|
||||
|
||||
```bash
|
||||
docker run -p 6333:6333 \
|
||||
-v $(pwd)/qdrant_storage:/qdrant/storage \
|
||||
qdrant/qdrant
|
||||
```
|
||||
You should see output like this
|
||||
|
||||
```text
|
||||
...
|
||||
[2021-02-05T00:08:51Z INFO actix_server::builder] Starting 12 workers
|
||||
[2021-02-05T00:08:51Z INFO actix_server::builder] Starting "actix-web-service-0.0.0.0:6333" service on 0.0.0.0:6333
|
||||
```
|
||||
|
||||
Test the service by going to [http://localhost:6333/](http://localhost:6333/). You should see the Qdrant version info in your browser.
|
||||
|
||||
All data uploaded to Qdrant is saved inside the `./qdrant_storage` directory and will be persisted even if you recreate the container.
|
||||
|
||||
|
||||
## Upload data to Qdrant
|
||||
|
||||
1. Install the official Python client to best interact with Qdrant.
|
||||
|
||||
```bash
|
||||
pip install "qdrant-client[fastembed]>=1.8.2"
|
||||
```
|
||||
> **Note:** This tutorial requires fastembed of version >=0.2.6.
|
||||
|
||||
At this point, you should have startup records in the `startups_demo.json` file and Qdrant running on a local machine.
|
||||
|
||||
Now you need to write a script to upload all startup data and vectors into the search engine.
|
||||
|
||||
2. Create a client object for Qdrant.
|
||||
|
||||
```python
|
||||
# Import client library
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333")
|
||||
```
|
||||
|
||||
3. Select model to encode your data.
|
||||
|
||||
You will be using two pre-trained models to compute dense and sparse vectors correspondingly: `sentence-transformers/all-MiniLM-L6-v2` and `prithivida/Splade_PP_en_v1`.
|
||||
|
||||
<aside role="status">
|
||||
Hybrid search implementation can be easily switched to a dense vector search by omitting the lines related to sparse vectors.
|
||||
</aside>
|
||||
|
||||
```python
|
||||
client.set_model("sentence-transformers/all-MiniLM-L6-v2")
|
||||
# comment this line to use dense vectors only
|
||||
client.set_sparse_model("prithivida/Splade_PP_en_v1")
|
||||
```
|
||||
|
||||
4. Related vectors need to be added to a collection. Create a new collection for your startup vectors.
|
||||
|
||||
```python
|
||||
if not client.collection_exists("startups"):
|
||||
client.create_collection(
|
||||
collection_name="startups",
|
||||
vectors_config=client.get_fastembed_vector_params(),
|
||||
# comment this line to use dense vectors only
|
||||
sparse_vectors_config=client.get_fastembed_sparse_vector_params(),
|
||||
)
|
||||
```
|
||||
|
||||
Qdrant requires vectors to have their own names and configurations.
|
||||
|
||||
Methods `get_fastembed_vector_params` and `get_fastembed_sparse_vector_params` help you to get the corresponding parameters for the models you are using.
|
||||
These parameters include vector size, distance function, etc.
|
||||
|
||||
Without fastembed integration, you would need to specify the vector size and distance function manually. Read more about it [here](/documentation/tutorials/neural-search/).
|
||||
|
||||
Additionally, you can specify extended configuration for your vectors, like `quantization_config` or `hnsw_config`.
|
||||
|
||||
|
||||
5. Read data from the file.
|
||||
|
||||
```python
|
||||
import json
|
||||
|
||||
payload_path = "startups_demo.json"
|
||||
metadata = []
|
||||
documents = []
|
||||
|
||||
with open(payload_path) as fd:
|
||||
for line in fd:
|
||||
obj = json.loads(line)
|
||||
documents.append(obj.pop("description"))
|
||||
metadata.append(obj)
|
||||
```
|
||||
|
||||
In this block of code, we read data from `startups_demo.json` file and split it into 2 lists: `documents` and `metadata`.
|
||||
Documents are the raw text descriptions of startups. Metadata is the payload associated with each startup, such as the name, location, and picture.
|
||||
We will use `documents` to encode the data into vectors.
|
||||
|
||||
|
||||
6. Encode and upload data.
|
||||
|
||||
```python
|
||||
client.add(
|
||||
collection_name="startups",
|
||||
documents=documents,
|
||||
metadata=metadata,
|
||||
parallel=0, # Use all available CPU cores to encode data.
|
||||
# Requires wrapping code into if __name__ == '__main__' block
|
||||
)
|
||||
```
|
||||
|
||||
<aside role="status">
|
||||
Vector generation process might be time-consuming. In order to save time, you can skip this step by uploading already processed data (available under the spoiler).
|
||||
</aside>
|
||||
|
||||
<details>
|
||||
<summary>Upload processed data</summary>
|
||||
|
||||
Download and unpack the processed data from [here](https://storage.googleapis.com/dataset-startup-search/startup-list-com/startups_hybrid_search_processed_40k.tar.gz) or use the following script:
|
||||
|
||||
```bash
|
||||
wget https://storage.googleapis.com/dataset-startup-search/startup-list-com/startups_hybrid_search_processed_40k.tar.gz
|
||||
tar -xvf startups_hybrid_search_processed_40k.tar.gz
|
||||
```
|
||||
|
||||
Then you can upload the data to Qdrant.
|
||||
|
||||
```python
|
||||
from typing import List
|
||||
import json
|
||||
import numpy as np
|
||||
from qdrant_client import models
|
||||
|
||||
|
||||
def named_vectors(vectors: List[float], sparse_vectors: List[models.SparseVector]) -> dict:
|
||||
# make sure to use the same client object as previously
|
||||
# or `set_model_name` and `set_sparse_model_name` manually
|
||||
dense_vector_name = client.get_vector_field_name()
|
||||
sparse_vector_name = client.get_sparse_vector_field_name()
|
||||
for vector, sparse_vector in zip(vectors, sparse_vectors):
|
||||
yield {
|
||||
dense_vector_name: vector,
|
||||
sparse_vector_name: models.SparseVector(**sparse_vector),
|
||||
}
|
||||
|
||||
with open("dense_vectors.npy", "rb") as f:
|
||||
vectors = np.load(f)
|
||||
|
||||
with open("sparse_vectors.json", "r") as f:
|
||||
sparse_vectors = json.load(f)
|
||||
|
||||
with open("payload.json", "r",) as f:
|
||||
payload = json.load(f)
|
||||
|
||||
client.upload_collection(
|
||||
"startups", vectors=named_vectors(vectors, sparse_vectors), payload=payload
|
||||
)
|
||||
```
|
||||
</details>
|
||||
|
||||
The `add` method will encode all documents and upload them to Qdrant.
|
||||
This is one of the two fastembed-specific methods, that combines encoding and uploading into a single step.
|
||||
|
||||
The `parallel` parameter enables data-parallelism instead of built-in ONNX parallelism.
|
||||
|
||||
Additionally, you can specify ids for each document, if you want to use them later to update or delete documents.
|
||||
If you don't specify ids, they will be generated automatically and returned as a result of the `add` method.
|
||||
|
||||
You can monitor the progress of the encoding by passing tqdm progress bar to the `add` method.
|
||||
|
||||
```python
|
||||
from tqdm import tqdm
|
||||
|
||||
client.add(
|
||||
collection_name="startups",
|
||||
documents=documents,
|
||||
metadata=metadata,
|
||||
ids=tqdm(range(len(documents))),
|
||||
)
|
||||
```
|
||||
|
||||
## Build the search API
|
||||
|
||||
Now that all the preparations are complete, let's start building a neural search class.
|
||||
|
||||
In order to process incoming requests, the hybrid search class will need 3 things: 1) models to convert the query into a vector, 2) the Qdrant client to perform search queries, 3) fusion function to re-rank dense and sparse search results.
|
||||
|
||||
Fastembed integration encapsulates query encoding, search and fusion into a single method call.
|
||||
Fastembed leverages [reciprocal rank fusion](https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf) in order combine the results.
|
||||
|
||||
|
||||
1. Create a file named `hybrid_searcher.py` and specify the following.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
|
||||
class HybridSearcher:
|
||||
DENSE_MODEL = "sentence-transformers/all-MiniLM-L6-v2"
|
||||
SPARSE_MODEL = "prithivida/Splade_PP_en_v1"
|
||||
def __init__(self, collection_name):
|
||||
self.collection_name = collection_name
|
||||
# initialize Qdrant client
|
||||
self.qdrant_client = QdrantClient("http://localhost:6333")
|
||||
self.qdrant_client.set_model(self.DENSE_MODEL)
|
||||
# comment this line to use dense vectors only
|
||||
self.qdrant_client.set_sparse_model(self.SPARSE_MODEL)
|
||||
```
|
||||
|
||||
2. Write the search function.
|
||||
|
||||
```python
|
||||
def search(self, text: str):
|
||||
search_result = self.qdrant_client.query(
|
||||
collection_name=self.collection_name,
|
||||
query_text=text,
|
||||
query_filter=None, # If you don't want any filters for now
|
||||
limit=5, # 5 the closest results
|
||||
)
|
||||
# `search_result` contains found vector ids with similarity scores
|
||||
# along with the stored payload
|
||||
|
||||
# Select and return metadata
|
||||
metadata = [hit.metadata for hit in search_result]
|
||||
return metadata
|
||||
```
|
||||
|
||||
3. Add search filters.
|
||||
|
||||
With Qdrant it is also feasible to add some conditions to the search.
|
||||
For example, if you wanted to search for startups in a certain city, the search query could look like this:
|
||||
|
||||
```python
|
||||
from qdrant_client import models
|
||||
|
||||
...
|
||||
|
||||
city_of_interest = "Berlin"
|
||||
|
||||
# Define a filter for cities
|
||||
city_filter = models.Filter(
|
||||
must=[
|
||||
models.FieldCondition(
|
||||
key="city",
|
||||
match=models.MatchValue(value=city_of_interest)
|
||||
)
|
||||
]
|
||||
)
|
||||
|
||||
search_result = self.qdrant_client.query(
|
||||
collection_name=self.collection_name,
|
||||
query_text=text,
|
||||
query_filter=city_filter,
|
||||
limit=5
|
||||
)
|
||||
...
|
||||
```
|
||||
|
||||
You have now created a class for neural search queries. Now wrap it up into a service.
|
||||
|
||||
## Deploy the search with FastAPI
|
||||
|
||||
To build the service you will use the FastAPI framework.
|
||||
|
||||
1. Install FastAPI.
|
||||
|
||||
To install it, use the command
|
||||
|
||||
```bash
|
||||
pip install fastapi uvicorn
|
||||
```
|
||||
|
||||
2. Implement the service.
|
||||
|
||||
Create a file named `service.py` and specify the following.
|
||||
|
||||
The service will have only one API endpoint and will look like this:
|
||||
|
||||
```python
|
||||
from fastapi import FastAPI
|
||||
|
||||
# The file where HybridSearcher is stored
|
||||
from hybrid_searcher import HybridSearcher
|
||||
|
||||
app = FastAPI()
|
||||
|
||||
# Create a neural searcher instance
|
||||
hybrid_searcher = HybridSearcher(collection_name="startups")
|
||||
|
||||
|
||||
@app.get("/api/search")
|
||||
def search_startup(q: str):
|
||||
return {"result": hybrid_searcher.search(text=q)}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
import uvicorn
|
||||
|
||||
uvicorn.run(app, host="0.0.0.0", port=8000)
|
||||
```
|
||||
|
||||
3. Run the service.
|
||||
|
||||
```bash
|
||||
python service.py
|
||||
```
|
||||
|
||||
4. Open your browser at [http://localhost:8000/docs](http://localhost:8000/docs).
|
||||
|
||||
You should be able to see a debug interface for your service.
|
||||
|
||||

|
||||
|
||||
Feel free to play around with it, make queries regarding the companies in our corpus, and check out the results.
|
||||
|
||||
Join our [Discord community](https://qdrant.to/discord), where we talk about vector search and similarity learning, publish other examples of neural networks and neural search applications.
|
||||
@@ -0,0 +1,340 @@
|
||||
---
|
||||
title: Neural Search Service
|
||||
weight: 2
|
||||
---
|
||||
|
||||
# Create a Simple Neural Search Service
|
||||
|
||||
| Time: 30 min | Level: Beginner | Output: [GitHub](https://github.com/qdrant/qdrant_demo/tree/sentense-transformers) | [](https://colab.research.google.com/drive/1kPktoudAP8Tu8n8l-iVMOQhVmHkWV_L9?usp=sharing) |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
|
||||
|
||||
This tutorial shows you how to build and deploy your own neural search service to look through descriptions of companies from [startups-list.com](https://www.startups-list.com/) and pick the most similar ones to your query. The website contains the company names, descriptions, locations, and a picture for each entry.
|
||||
|
||||
A neural search service uses artificial neural networks to improve the accuracy and relevance of search results. Besides offering simple keyword results, this system can retrieve results by meaning. It can understand and interpret complex search queries and provide more contextually relevant output, effectively enhancing the user's search experience.
|
||||
|
||||
<aside role="status">
|
||||
There is a version of this tutorial that uses <a href="https://github.com/qdrant/fastembed">Fastembed</a> model inference engine instead of Sentence Transformers.
|
||||
Check it out <a href="/documentation/tutorials/hybrid-search-fastembed/">here</a>.
|
||||
</aside>
|
||||
|
||||
|
||||
## Workflow
|
||||
|
||||
To create a neural search service, you will need to transform your raw data and then create a search function to manipulate it. First, you will 1) download and prepare a sample dataset using a modified version of the BERT ML model. Then, you will 2) load the data into Qdrant, 3) create a neural search API and 4) serve it using FastAPI.
|
||||
|
||||

|
||||
|
||||
> **Note**: The code for this tutorial can be found here: | [Step 1: Data Preparation Process](https://colab.research.google.com/drive/1kPktoudAP8Tu8n8l-iVMOQhVmHkWV_L9?usp=sharing) | [Step 2: Full Code for Neural Search](https://github.com/qdrant/qdrant_demo/tree/sentense-transformers). |
|
||||
|
||||
## Prerequisites
|
||||
|
||||
To complete this tutorial, you will need:
|
||||
|
||||
- Docker - The easiest way to use Qdrant is to run a pre-built Docker image.
|
||||
- [Raw parsed data](https://storage.googleapis.com/generall-shared-data/startups_demo.json) from startups-list.com.
|
||||
- Python version >=3.8
|
||||
|
||||
## Prepare sample dataset
|
||||
|
||||
To conduct a neural search on startup descriptions, you must first encode the description data into vectors. To process text, you can use a pre-trained models like [BERT](https://en.wikipedia.org/wiki/BERT_(language_model)) or sentence transformers. The [sentence-transformers](https://github.com/UKPLab/sentence-transformers) library lets you conveniently download and use many pre-trained models, such as DistilBERT, MPNet, etc.
|
||||
|
||||
1. First you need to download the dataset.
|
||||
|
||||
```bash
|
||||
wget https://storage.googleapis.com/generall-shared-data/startups_demo.json
|
||||
```
|
||||
|
||||
2. Install the SentenceTransformer library as well as other relevant packages.
|
||||
|
||||
```bash
|
||||
pip install sentence-transformers numpy pandas tqdm
|
||||
```
|
||||
|
||||
3. Import the required modules.
|
||||
|
||||
```python
|
||||
from sentence_transformers import SentenceTransformer
|
||||
import numpy as np
|
||||
import json
|
||||
import pandas as pd
|
||||
from tqdm.notebook import tqdm
|
||||
```
|
||||
|
||||
You will be using a pre-trained model called `all-MiniLM-L6-v2`.
|
||||
This is a performance-optimized sentence embedding model and you can read more about it and other available models [here](https://www.sbert.net/docs/pretrained_models.html).
|
||||
|
||||
|
||||
4. Download and create a pre-trained sentence encoder.
|
||||
|
||||
```python
|
||||
model = SentenceTransformer(
|
||||
"all-MiniLM-L6-v2", device="cuda"
|
||||
) # or device="cpu" if you don't have a GPU
|
||||
```
|
||||
5. Read the raw data file.
|
||||
|
||||
```python
|
||||
df = pd.read_json("./startups_demo.json", lines=True)
|
||||
```
|
||||
6. Encode all startup descriptions to create an embedding vector for each. Internally, the `encode` function will split the input into batches, which will significantly speed up the process.
|
||||
|
||||
```python
|
||||
vectors = model.encode(
|
||||
[row.alt + ". " + row.description for row in df.itertuples()],
|
||||
show_progress_bar=True,
|
||||
)
|
||||
```
|
||||
All of the descriptions are now converted into vectors. There are 40474 vectors of 384 dimensions. The output layer of the model has this dimension
|
||||
|
||||
```python
|
||||
vectors.shape
|
||||
# > (40474, 384)
|
||||
```
|
||||
|
||||
7. Download the saved vectors into a new file named `startup_vectors.npy`
|
||||
|
||||
```python
|
||||
np.save("startup_vectors.npy", vectors, allow_pickle=False)
|
||||
```
|
||||
|
||||
## Run Qdrant in Docker
|
||||
|
||||
Next, you need to manage all of your data using a vector engine. Qdrant lets you store, update or delete created vectors. Most importantly, it lets you search for the nearest vectors via a convenient API.
|
||||
|
||||
> **Note:** Before you begin, create a project directory and a virtual python environment in it.
|
||||
|
||||
1. Download the Qdrant image from DockerHub.
|
||||
|
||||
```bash
|
||||
docker pull qdrant/qdrant
|
||||
```
|
||||
2. Start Qdrant inside of Docker.
|
||||
|
||||
```bash
|
||||
docker run -p 6333:6333 \
|
||||
-v $(pwd)/qdrant_storage:/qdrant/storage \
|
||||
qdrant/qdrant
|
||||
```
|
||||
You should see output like this
|
||||
|
||||
```text
|
||||
...
|
||||
[2021-02-05T00:08:51Z INFO actix_server::builder] Starting 12 workers
|
||||
[2021-02-05T00:08:51Z INFO actix_server::builder] Starting "actix-web-service-0.0.0.0:6333" service on 0.0.0.0:6333
|
||||
```
|
||||
|
||||
Test the service by going to [http://localhost:6333/](http://localhost:6333/). You should see the Qdrant version info in your browser.
|
||||
|
||||
All data uploaded to Qdrant is saved inside the `./qdrant_storage` directory and will be persisted even if you recreate the container.
|
||||
|
||||
## Upload data to Qdrant
|
||||
|
||||
1. Install the official Python client to best interact with Qdrant.
|
||||
|
||||
```bash
|
||||
pip install qdrant-client
|
||||
```
|
||||
|
||||
At this point, you should have startup records in the `startups_demo.json` file, encoded vectors in `startup_vectors.npy` and Qdrant running on a local machine.
|
||||
|
||||
Now you need to write a script to upload all startup data and vectors into the search engine.
|
||||
|
||||
2. Create a client object for Qdrant.
|
||||
|
||||
```python
|
||||
# Import client library
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.models import VectorParams, Distance
|
||||
|
||||
client = QdrantClient("http://localhost:6333")
|
||||
```
|
||||
|
||||
3. Related vectors need to be added to a collection. Create a new collection for your startup vectors.
|
||||
|
||||
```python
|
||||
if not client.collection_exists("startups"):
|
||||
client.create_collection(
|
||||
collection_name="startups",
|
||||
vectors_config=VectorParams(size=384, distance=Distance.COSINE),
|
||||
)
|
||||
```
|
||||
<aside role="status">
|
||||
|
||||
- The `vector_size` parameter defines the size of the vectors for a specific collection. If their size is different, it is impossible to calculate the distance between them. `384` is the encoder output dimensionality. You can also use `model.get_sentence_embedding_dimension()` to get the dimensionality of the model you are using.
|
||||
|
||||
- The `distance` parameter lets you specify the function used to measure the distance between two points.
|
||||
|
||||
</aside>
|
||||
|
||||
4. Create an iterator over the startup data and vectors.
|
||||
|
||||
The Qdrant client library defines a special function that allows you to load datasets into the service.
|
||||
However, since there may be too much data to fit a single computer memory, the function takes an iterator over the data as input.
|
||||
|
||||
```python
|
||||
fd = open("./startups_demo.json")
|
||||
|
||||
# payload is now an iterator over startup data
|
||||
payload = map(json.loads, fd)
|
||||
|
||||
# Load all vectors into memory, numpy array works as iterable for itself.
|
||||
# Other option would be to use Mmap, if you don't want to load all data into RAM
|
||||
vectors = np.load("./startup_vectors.npy")
|
||||
```
|
||||
|
||||
5. Upload the data
|
||||
|
||||
```python
|
||||
client.upload_collection(
|
||||
collection_name="startups",
|
||||
vectors=vectors,
|
||||
payload=payload,
|
||||
ids=None, # Vector ids will be assigned automatically
|
||||
batch_size=256, # How many vectors will be uploaded in a single request?
|
||||
)
|
||||
```
|
||||
|
||||
Vectors are now uploaded to Qdrant.
|
||||
|
||||
## Build the search API
|
||||
|
||||
Now that all the preparations are complete, let's start building a neural search class.
|
||||
|
||||
In order to process incoming requests, neural search will need 2 things: 1) a model to convert the query into a vector and 2) the Qdrant client to perform search queries.
|
||||
|
||||
1. Create a file named `neural_searcher.py` and specify the following.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from sentence_transformers import SentenceTransformer
|
||||
|
||||
|
||||
class NeuralSearcher:
|
||||
def __init__(self, collection_name):
|
||||
self.collection_name = collection_name
|
||||
# Initialize encoder model
|
||||
self.model = SentenceTransformer("all-MiniLM-L6-v2", device="cpu")
|
||||
# initialize Qdrant client
|
||||
self.qdrant_client = QdrantClient("http://localhost:6333")
|
||||
```
|
||||
|
||||
2. Write the search function.
|
||||
|
||||
```python
|
||||
def search(self, text: str):
|
||||
# Convert text query into vector
|
||||
vector = self.model.encode(text).tolist()
|
||||
|
||||
# Use `vector` for search for closest vectors in the collection
|
||||
search_result = self.qdrant_client.query_points(
|
||||
collection_name=self.collection_name,
|
||||
query=vector,
|
||||
query_filter=None, # If you don't want any filters for now
|
||||
limit=5, # 5 the most closest results is enough
|
||||
).points
|
||||
# `search_result` contains found vector ids with similarity scores along with the stored payload
|
||||
# In this function you are interested in payload only
|
||||
payloads = [hit.payload for hit in search_result]
|
||||
return payloads
|
||||
```
|
||||
|
||||
3. Add search filters.
|
||||
|
||||
With Qdrant it is also feasible to add some conditions to the search.
|
||||
For example, if you wanted to search for startups in a certain city, the search query could look like this:
|
||||
|
||||
```python
|
||||
from qdrant_client.models import Filter
|
||||
|
||||
...
|
||||
|
||||
city_of_interest = "Berlin"
|
||||
|
||||
# Define a filter for cities
|
||||
city_filter = Filter(**{
|
||||
"must": [{
|
||||
"key": "city", # Store city information in a field of the same name
|
||||
"match": { # This condition checks if payload field has the requested value
|
||||
"value": city_of_interest
|
||||
}
|
||||
}]
|
||||
})
|
||||
|
||||
search_result = self.qdrant_client.query_points(
|
||||
collection_name=self.collection_name,
|
||||
query=vector,
|
||||
query_filter=city_filter,
|
||||
limit=5
|
||||
).points
|
||||
...
|
||||
```
|
||||
|
||||
You have now created a class for neural search queries. Now wrap it up into a service.
|
||||
|
||||
## Deploy the search with FastAPI
|
||||
|
||||
To build the service you will use the FastAPI framework.
|
||||
|
||||
1. Install FastAPI.
|
||||
|
||||
To install it, use the command
|
||||
|
||||
```bash
|
||||
pip install fastapi uvicorn
|
||||
```
|
||||
|
||||
2. Implement the service.
|
||||
|
||||
Create a file named `service.py` and specify the following.
|
||||
|
||||
The service will have only one API endpoint and will look like this:
|
||||
|
||||
```python
|
||||
from fastapi import FastAPI
|
||||
|
||||
# The file where NeuralSearcher is stored
|
||||
from neural_searcher import NeuralSearcher
|
||||
|
||||
app = FastAPI()
|
||||
|
||||
# Create a neural searcher instance
|
||||
neural_searcher = NeuralSearcher(collection_name="startups")
|
||||
|
||||
|
||||
@app.get("/api/search")
|
||||
def search_startup(q: str):
|
||||
return {"result": neural_searcher.search(text=q)}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
import uvicorn
|
||||
|
||||
uvicorn.run(app, host="0.0.0.0", port=8000)
|
||||
```
|
||||
|
||||
3. Run the service.
|
||||
|
||||
```bash
|
||||
python service.py
|
||||
```
|
||||
|
||||
4. Open your browser at [http://localhost:8000/docs](http://localhost:8000/docs).
|
||||
|
||||
You should be able to see a debug interface for your service.
|
||||
|
||||

|
||||
|
||||
Feel free to play around with it, make queries regarding the companies in our corpus, and check out the results.
|
||||
|
||||
## Next steps
|
||||
|
||||
The code from this tutorial has been used to develop a [live online demo](https://qdrant.to/semantic-search-demo).
|
||||
You can try it to get an intuition for cases when the neural search is useful.
|
||||
The demo contains a switch that selects between neural and full-text searches.
|
||||
You can turn the neural search on and off to compare your result with a regular full-text search.
|
||||
|
||||
> **Note**: The code for this tutorial can be found here: | [Step 1: Data Preparation Process](https://colab.research.google.com/drive/1kPktoudAP8Tu8n8l-iVMOQhVmHkWV_L9?usp=sharing) | [Step 2: Full Code for Neural Search](https://github.com/qdrant/qdrant_demo/tree/sentense-transformers). |
|
||||
|
||||
Join our [Discord community](https://qdrant.to/discord), where we talk about vector search and similarity learning, publish other examples of neural networks and neural search applications.
|
||||
@@ -0,0 +1,227 @@
|
||||
---
|
||||
title: Measure Search Quality
|
||||
weight: 4
|
||||
---
|
||||
|
||||
# Measure retrieval quality
|
||||
|
||||
| Time: 30 min | Level: Intermediate | | |
|
||||
|--------------|---------------------|--|----|
|
||||
|
||||
Semantic search pipelines are as good as the embeddings they use. If your model cannot properly represent input data, similar objects might
|
||||
be far away from each other in the vector space. No surprise, that the search results will be poor in this case. There is, however, another
|
||||
component of the process which can also degrade the quality of the search results. It is the ANN algorithm itself.
|
||||
|
||||
In this tutorial, we will show how to measure the quality of the semantic retrieval and how to tune the parameters of the HNSW, the ANN
|
||||
algorithm used in Qdrant, to obtain the best results.
|
||||
|
||||
## Embeddings quality
|
||||
|
||||
The quality of the embeddings is a topic for a separate tutorial. In a nutshell, it is usually measured and compared by benchmarks, such as
|
||||
[Massive Text Embedding Benchmark (MTEB)](https://huggingface.co/spaces/mteb/leaderboard). The evaluation process itself is pretty
|
||||
straightforward and is based on a ground truth dataset built by humans. We have a set of queries and a set of the documents we would expect
|
||||
to receive for each of them. In the [evaluation process](https://qdrant.tech/rag/rag-evaluation-guide/), we take a query, find the most similar documents in the vector space and compare
|
||||
them with the ground truth. In that setup, **finding the most similar documents is implemented as full kNN search, without any approximation**.
|
||||
As a result, we can measure the quality of the embeddings themselves, without the influence of the ANN algorithm.
|
||||
|
||||
## Retrieval quality
|
||||
|
||||
Embeddings quality is indeed the most important factor in the semantic search quality. However, vector search engines, such as Qdrant, do not
|
||||
perform pure kNN search. Instead, they use **Approximate Nearest Neighbors** (ANN) algorithms, which are much faster than the exact search,
|
||||
but can return suboptimal results. We can also **measure the retrieval quality of that approximation** which also contributes to the overall
|
||||
search quality.
|
||||
|
||||
### Quality metrics
|
||||
|
||||
There are various ways of how quantify the quality of semantic search. Some of them, such as [Precision@k](https://en.wikipedia.org/wiki/Evaluation_measures_(information_retrieval)#Precision_at_k),
|
||||
are based on the number of relevant documents in the top-k search results. Others, such as [Mean Reciprocal Rank (MRR)](https://en.wikipedia.org/wiki/Mean_reciprocal_rank),
|
||||
take into account the position of the first relevant document in the search results. [DCG and NDCG](https://en.wikipedia.org/wiki/Discounted_cumulative_gain)
|
||||
metrics are, in turn, based on the relevance score of the documents.
|
||||
|
||||
If we treat the search pipeline as a whole, we could use them all. The same is true for the embeddings quality evaluation. However, for the
|
||||
ANN algorithm itself, anything based on the relevance score or ranking is not applicable. Ranking in vector search relies on the distance
|
||||
between the query and the document in the vector space, however distance is not going to change due to approximation, as the function is
|
||||
still the same.
|
||||
|
||||
Therefore, it only makes sense to measure the quality of the ANN algorithm by the number of relevant documents in the top-k search results,
|
||||
such as `precision@k`. It is calculated as the number of relevant documents in the top-k search results divided by `k`. In case of testing
|
||||
just the ANN algorithm, we can use the exact kNN search as a ground truth, with `k` being fixed. It will be a measure on **how well the ANN
|
||||
algorithm approximates the exact search**.
|
||||
|
||||
## Measure the quality of the search results
|
||||
|
||||
Let's build a quality [evaluation](https://qdrant.tech/rag/rag-evaluation-guide/) of the ANN algorithm in Qdrant. We will, first, call the search endpoint in a standard way to obtain
|
||||
the approximate search results. Then, we will call the exact search endpoint to obtain the exact matches, and finally compare both results
|
||||
in terms of precision.
|
||||
|
||||
Before we start, let's create a collection, fill it with some data and then start our evaluation. We will use the same dataset as in the
|
||||
[Loading a dataset from Hugging Face hub](/documentation/tutorials/huggingface-datasets/) tutorial, `Qdrant/arxiv-titles-instructorxl-embeddings`
|
||||
from the [Hugging Face hub](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings). Let's download it in a streaming
|
||||
mode, as we are only going to use part of it.
|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
|
||||
dataset = load_dataset(
|
||||
"Qdrant/arxiv-titles-instructorxl-embeddings", split="train", streaming=True
|
||||
)
|
||||
```
|
||||
|
||||
We need some data to be indexed and another set for the testing purposes. Let's get the first 50000 items for the training and the next 1000
|
||||
for the testing.
|
||||
|
||||
```python
|
||||
dataset_iterator = iter(dataset)
|
||||
train_dataset = [next(dataset_iterator) for _ in range(60000)]
|
||||
test_dataset = [next(dataset_iterator) for _ in range(1000)]
|
||||
```
|
||||
|
||||
Now, let's create a collection and index the training data. This collection will be created with the default configuration. Please be aware that
|
||||
it might be different from your collection settings, and it's always important to test exactly the same configuration you are going to use later
|
||||
in production.
|
||||
|
||||
<aside role="status">
|
||||
Distance function is another parameter that may impact the retrieval quality. If the embedding model was not trained to minimize cosine
|
||||
distance, you can get suboptimal search results by using it. Please test different distance functions to find the best one for your embeddings,
|
||||
if you don't know the specifics of the model training.
|
||||
</aside>
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient("http://localhost:6333")
|
||||
client.create_collection(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
vectors_config=models.VectorParams(
|
||||
size=768, # Size of the embeddings generated by InstructorXL model
|
||||
distance=models.Distance.COSINE,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
We are now ready to index the training data. Uploading the records is going to trigger the indexing process, which will build the HNSW graph.
|
||||
The indexing process may take some time, depending on the size of the dataset, but your data is going to be available for search immediately
|
||||
after receiving the response from the `upsert` endpoint. **As long as the indexing is not finished, and HNSW not built, Qdrant will perform
|
||||
the exact search**. We have to wait until the indexing is finished to be sure that the approximate search is performed.
|
||||
|
||||
```python
|
||||
client.upload_points( # upload_points is available as of qdrant-client v1.7.1
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=item["id"],
|
||||
vector=item["vector"],
|
||||
payload=item,
|
||||
)
|
||||
for item in train_dataset
|
||||
]
|
||||
)
|
||||
|
||||
while True:
|
||||
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
|
||||
if collection_info.status == models.CollectionStatus.GREEN:
|
||||
# Collection status is green, which means the indexing is finished
|
||||
break
|
||||
```
|
||||
|
||||
## Standard mode vs exact search
|
||||
|
||||
Qdrant has a built-in exact search mode, which can be used to measure the quality of the search results. In this mode, Qdrant performs a
|
||||
full kNN search for each query, without any approximation. It is not suitable for production use with high load, but it is perfect for the
|
||||
evaluation of the ANN algorithm and its parameters. It might be triggered by setting the `exact` parameter to `True` in the search request.
|
||||
We are simply going to use all the examples from the test dataset as queries and compare the results of the approximate search with the
|
||||
results of the exact search. Let's create a helper function with `k` being a parameter, so we can calculate the `precision@k` for different
|
||||
values of `k`.
|
||||
|
||||
```python
|
||||
def avg_precision_at_k(k: int):
|
||||
precisions = []
|
||||
for item in test_dataset:
|
||||
ann_result = client.query_points(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
query=item["vector"],
|
||||
limit=k,
|
||||
).points
|
||||
|
||||
knn_result = client.query_points(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
query=item["vector"],
|
||||
limit=k,
|
||||
search_params=models.SearchParams(
|
||||
exact=True, # Turns on the exact search mode
|
||||
),
|
||||
).points
|
||||
|
||||
# We can calculate the precision@k by comparing the ids of the search results
|
||||
ann_ids = set(item.id for item in ann_result)
|
||||
knn_ids = set(item.id for item in knn_result)
|
||||
precision = len(ann_ids.intersection(knn_ids)) / k
|
||||
precisions.append(precision)
|
||||
|
||||
return sum(precisions) / len(precisions)
|
||||
```
|
||||
|
||||
Calculating the `precision@5` is as simple as calling the function with the corresponding parameter:
|
||||
|
||||
```python
|
||||
print(f"avg(precision@5) = {avg_precision_at_k(k=5)}")
|
||||
```
|
||||
|
||||
Response:
|
||||
|
||||
```text
|
||||
avg(precision@5) = 0.9935999999999995
|
||||
```
|
||||
|
||||
As we can see, the precision of the approximate search vs exact search is pretty high. There are, however, some scenarios when we
|
||||
need higher precision and can accept higher latency. HNSW is pretty tunable, and we can increase the precision by changing its parameters.
|
||||
|
||||
## Tweaking the HNSW parameters
|
||||
|
||||
HNSW is a hierarchical graph, where each node has a set of links to other nodes. The number of edges per node is called the `m` parameter.
|
||||
The larger the value of it, the higher the precision of the search, but more space required. The `ef_construct` parameter is the number of
|
||||
neighbours to consider during the index building. Again, the larger the value, the higher the precision, but the longer the indexing time.
|
||||
The default values of these parameters are `m=16` and `ef_construct=100`. Let's try to increase them to `m=32` and `ef_construct=200` and
|
||||
see how it affects the precision. Of course, we need to wait until the indexing is finished before we can perform the search.
|
||||
|
||||
```python
|
||||
client.update_collection(
|
||||
collection_name="arxiv-titles-instructorxl-embeddings",
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=32, # Increase the number of edges per node from the default 16 to 32
|
||||
ef_construct=200, # Increase the number of neighbours from the default 100 to 200
|
||||
)
|
||||
)
|
||||
|
||||
while True:
|
||||
collection_info = client.get_collection(collection_name="arxiv-titles-instructorxl-embeddings")
|
||||
if collection_info.status == models.CollectionStatus.GREEN:
|
||||
# Collection status is green, which means the indexing is finished
|
||||
break
|
||||
```
|
||||
|
||||
The same function can be used to calculate the average `precision@5`:
|
||||
|
||||
```python
|
||||
print(f"avg(precision@5) = {avg_precision_at_k(k=5)}")
|
||||
```
|
||||
|
||||
Response:
|
||||
|
||||
```text
|
||||
avg(precision@5) = 0.9969999999999998
|
||||
```
|
||||
|
||||
The precision has obviously increased, and we know how to control it. However, there is a trade-off between the precision and the search
|
||||
latency and memory requirements. In some specific cases, we may want to increase the precision as much as possible, so now we know how
|
||||
to do it.
|
||||
|
||||
## Wrapping up
|
||||
|
||||
Assessing the quality of retrieval is a critical aspect of [evaluating](https://qdrant.tech/rag/rag-evaluation-guide/) semantic search performance. It is imperative to measure retrieval quality when aiming for optimal quality of.
|
||||
your search results. Qdrant provides a built-in exact search mode, which can be used to measure the quality of the ANN algorithm itself,
|
||||
even in an automated way, as part of your CI/CD pipeline.
|
||||
|
||||
Again, **the quality of the embeddings is the most important factor**. HNSW does a pretty good job in terms of precision, and it is
|
||||
parameterizable and tunable, when required. There are some other ANN algorithms available out there, such as [IVF*](https://github.com/facebookresearch/faiss/wiki/Faiss-indexes#cell-probe-methods-indexivf-indexes),
|
||||
but they usually [perform worse than HNSW in terms of quality and performance](https://nirantk.com/writing/pgvector-vs-qdrant/#correctness).
|
||||
@@ -0,0 +1,250 @@
|
||||
---
|
||||
title: Semantic Search 101
|
||||
weight: 1
|
||||
aliases:
|
||||
- /documentation/tutorials/mighty.md/
|
||||
---
|
||||
|
||||
# Semantic Search for Beginners
|
||||
|
||||
| Time: 5 - 15 min | Level: Beginner | | |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
<p align="center"><iframe width="560" height="315" src="https://www.youtube.com/embed/AASiqmtKo54" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen></iframe></p>
|
||||
|
||||
## Overview
|
||||
|
||||
If you are new to vector databases, this tutorial is for you. In 5 minutes you will build a semantic search engine for science fiction books. After you set it up, you will ask the engine about an impending alien threat. Your creation will recommend books as preparation for a potential space attack.
|
||||
|
||||
Before you begin, you need to have a [recent version of Python](https://www.python.org/downloads/) installed. If you don't know how to run this code in a virtual environment, follow Python documentation for [Creating Virtual Environments](https://docs.python.org/3/tutorial/venv.html#creating-virtual-environments) first.
|
||||
|
||||
This tutorial assumes you're in the bash shell. Use the Python documentation to activate a virtual environment, with commands such as:
|
||||
|
||||
```bash
|
||||
source tutorial-env/bin/activate
|
||||
```
|
||||
|
||||
## 1. Installation
|
||||
|
||||
You need to process your data so that the search engine can work with it. The [Sentence Transformers](https://www.sbert.net/) framework gives you access to common Large Language Models that turn raw data into embeddings.
|
||||
|
||||
```bash
|
||||
pip install -U sentence-transformers
|
||||
```
|
||||
|
||||
Once encoded, this data needs to be kept somewhere. Qdrant lets you store data as embeddings. You can also use Qdrant to run search queries against this data. This means that you can ask the engine to give you relevant answers that go way beyond keyword matching.
|
||||
|
||||
```bash
|
||||
pip install -U qdrant-client
|
||||
```
|
||||
|
||||
<aside role="status">
|
||||
This tutorial requires qdrant-client version 1.7.1 or higher.
|
||||
</aside>
|
||||
|
||||
### Import the models
|
||||
|
||||
Once the two main frameworks are defined, you need to specify the exact models this engine will use. Before you do, activate the Python prompt (`>>>`) with the `python` command.
|
||||
|
||||
```python
|
||||
from qdrant_client import models, QdrantClient
|
||||
from sentence_transformers import SentenceTransformer
|
||||
```
|
||||
|
||||
The [Sentence Transformers](https://www.sbert.net/index.html) framework contains many embedding models. However, [all-MiniLM-L6-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) is the fastest encoder for this tutorial.
|
||||
|
||||
```python
|
||||
encoder = SentenceTransformer("all-MiniLM-L6-v2")
|
||||
```
|
||||
|
||||
## 2. Add the dataset
|
||||
|
||||
[all-MiniLM-L6-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) will encode the data you provide. Here you will list all the science fiction books in your library. Each book has metadata, a name, author, publication year and a short description.
|
||||
|
||||
```python
|
||||
documents = [
|
||||
{
|
||||
"name": "The Time Machine",
|
||||
"description": "A man travels through time and witnesses the evolution of humanity.",
|
||||
"author": "H.G. Wells",
|
||||
"year": 1895,
|
||||
},
|
||||
{
|
||||
"name": "Ender's Game",
|
||||
"description": "A young boy is trained to become a military leader in a war against an alien race.",
|
||||
"author": "Orson Scott Card",
|
||||
"year": 1985,
|
||||
},
|
||||
{
|
||||
"name": "Brave New World",
|
||||
"description": "A dystopian society where people are genetically engineered and conditioned to conform to a strict social hierarchy.",
|
||||
"author": "Aldous Huxley",
|
||||
"year": 1932,
|
||||
},
|
||||
{
|
||||
"name": "The Hitchhiker's Guide to the Galaxy",
|
||||
"description": "A comedic science fiction series following the misadventures of an unwitting human and his alien friend.",
|
||||
"author": "Douglas Adams",
|
||||
"year": 1979,
|
||||
},
|
||||
{
|
||||
"name": "Dune",
|
||||
"description": "A desert planet is the site of political intrigue and power struggles.",
|
||||
"author": "Frank Herbert",
|
||||
"year": 1965,
|
||||
},
|
||||
{
|
||||
"name": "Foundation",
|
||||
"description": "A mathematician develops a science to predict the future of humanity and works to save civilization from collapse.",
|
||||
"author": "Isaac Asimov",
|
||||
"year": 1951,
|
||||
},
|
||||
{
|
||||
"name": "Snow Crash",
|
||||
"description": "A futuristic world where the internet has evolved into a virtual reality metaverse.",
|
||||
"author": "Neal Stephenson",
|
||||
"year": 1992,
|
||||
},
|
||||
{
|
||||
"name": "Neuromancer",
|
||||
"description": "A hacker is hired to pull off a near-impossible hack and gets pulled into a web of intrigue.",
|
||||
"author": "William Gibson",
|
||||
"year": 1984,
|
||||
},
|
||||
{
|
||||
"name": "The War of the Worlds",
|
||||
"description": "A Martian invasion of Earth throws humanity into chaos.",
|
||||
"author": "H.G. Wells",
|
||||
"year": 1898,
|
||||
},
|
||||
{
|
||||
"name": "The Hunger Games",
|
||||
"description": "A dystopian society where teenagers are forced to fight to the death in a televised spectacle.",
|
||||
"author": "Suzanne Collins",
|
||||
"year": 2008,
|
||||
},
|
||||
{
|
||||
"name": "The Andromeda Strain",
|
||||
"description": "A deadly virus from outer space threatens to wipe out humanity.",
|
||||
"author": "Michael Crichton",
|
||||
"year": 1969,
|
||||
},
|
||||
{
|
||||
"name": "The Left Hand of Darkness",
|
||||
"description": "A human ambassador is sent to a planet where the inhabitants are genderless and can change gender at will.",
|
||||
"author": "Ursula K. Le Guin",
|
||||
"year": 1969,
|
||||
},
|
||||
{
|
||||
"name": "The Three-Body Problem",
|
||||
"description": "Humans encounter an alien civilization that lives in a dying system.",
|
||||
"author": "Liu Cixin",
|
||||
"year": 2008,
|
||||
},
|
||||
]
|
||||
```
|
||||
|
||||
## 3. Define storage location
|
||||
|
||||
You need to tell Qdrant where to store embeddings. This is a basic demo, so your local computer will use its memory as temporary storage.
|
||||
|
||||
```python
|
||||
client = QdrantClient(":memory:")
|
||||
```
|
||||
|
||||
## 4. Create a collection
|
||||
|
||||
All data in Qdrant is organized by collections. In this case, you are storing books, so we are calling it `my_books`.
|
||||
|
||||
```python
|
||||
client.create_collection(
|
||||
collection_name="my_books",
|
||||
vectors_config=models.VectorParams(
|
||||
size=encoder.get_sentence_embedding_dimension(), # Vector size is defined by used model
|
||||
distance=models.Distance.COSINE,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
- The `vector_size` parameter defines the size of the vectors for a specific collection. If their size is different, it is impossible to calculate the distance between them. 384 is the encoder output dimensionality. You can also use model.get_sentence_embedding_dimension() to get the dimensionality of the model you are using.
|
||||
|
||||
- The `distance` parameter lets you specify the function used to measure the distance between two points.
|
||||
|
||||
|
||||
## 5. Upload data to collection
|
||||
|
||||
Tell the database to upload `documents` to the `my_books` collection. This will give each record an id and a payload. The payload is just the metadata from the dataset.
|
||||
|
||||
```python
|
||||
client.upload_points(
|
||||
collection_name="my_books",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=idx, vector=encoder.encode(doc["description"]).tolist(), payload=doc
|
||||
)
|
||||
for idx, doc in enumerate(documents)
|
||||
],
|
||||
)
|
||||
```
|
||||
|
||||
## 6. Ask the engine a question
|
||||
|
||||
Now that the data is stored in Qdrant, you can ask it questions and receive semantically relevant results.
|
||||
|
||||
```python
|
||||
hits = client.query_points(
|
||||
collection_name="my_books",
|
||||
query=encoder.encode("alien invasion").tolist(),
|
||||
limit=3,
|
||||
).points
|
||||
|
||||
for hit in hits:
|
||||
print(hit.payload, "score:", hit.score)
|
||||
```
|
||||
|
||||
**Response:**
|
||||
|
||||
The search engine shows three of the most likely responses that have to do with the alien invasion. Each of the responses is assigned a score to show how close the response is to the original inquiry.
|
||||
|
||||
```text
|
||||
{'name': 'The War of the Worlds', 'description': 'A Martian invasion of Earth throws humanity into chaos.', 'author': 'H.G. Wells', 'year': 1898} score: 0.570093257022374
|
||||
{'name': "The Hitchhiker's Guide to the Galaxy", 'description': 'A comedic science fiction series following the misadventures of an unwitting human and his alien friend.', 'author': 'Douglas Adams', 'year': 1979} score: 0.5040468703143637
|
||||
{'name': 'The Three-Body Problem', 'description': 'Humans encounter an alien civilization that lives in a dying system.', 'author': 'Liu Cixin', 'year': 2008} score: 0.45902943411768216
|
||||
```
|
||||
|
||||
### Narrow down the query
|
||||
|
||||
How about the most recent book from the early 2000s?
|
||||
|
||||
```python
|
||||
hits = client.query_points(
|
||||
collection_name="my_books",
|
||||
query=encoder.encode("alien invasion").tolist(),
|
||||
query_filter=models.Filter(
|
||||
must=[models.FieldCondition(key="year", range=models.Range(gte=2000))]
|
||||
),
|
||||
limit=1,
|
||||
).points
|
||||
|
||||
for hit in hits:
|
||||
print(hit.payload, "score:", hit.score)
|
||||
```
|
||||
|
||||
**Response:**
|
||||
|
||||
The query has been narrowed down to one result from 2008.
|
||||
|
||||
```text
|
||||
{'name': 'The Three-Body Problem', 'description': 'Humans encounter an alien civilization that lives in a dying system.', 'author': 'Liu Cixin', 'year': 2008} score: 0.45902943411768216
|
||||
```
|
||||
|
||||
## Next Steps
|
||||
|
||||
Congratulations, you have just created your very first search engine! Trust us, the rest of Qdrant is not that complicated, either. For your next tutorial you should try building an actual [Neural Search Service with a complete API and a dataset](/documentation/tutorials/neural-search/).
|
||||
|
||||
## Return to the bash shell
|
||||
|
||||
To return to the bash prompt:
|
||||
|
||||
1. Press Ctrl+D to exit the Python prompt (`>>>`).
|
||||
1. Enter the `deactivate` command to deactivate the virtual environment.
|
||||
Reference in New Issue
Block a user