new: add hybrid search with fastembed tutorial (#790)

* new: add hybrid search with fastembed demo

* fix: fix link, replace neural search

* fix: update article weight, fix upload code

* new: add param to enable or disable hybrid search

* fix: rephrasing

* fix: increase limit

* refactoring: rephrasing, updating code examples

* refactoring: add break, rephrase repeated sentence

* new: add package versions

* fix: merge neural search and hybrid search articles

* fix: fix rrf link

* fix: re-add spoiler to upload already processed data

* new: add link to the processed sample

* fix: fix recreate collection

* fix: fix filename, warn about multiprocessing

* fix: remove limit mention

* fix: fix link to the processed data

* fix: fix indentation in the spoiler

* fix: replace type hints for python3.8, add comments, file reading"

* review fixes

* upd image style

* rephrase comment in code

---------

Co-authored-by: generall <andrey@vasnetsov.com>
This commit is contained in:
George
2024-04-13 22:24:32 +02:00
committed by GitHub
co-authored by generall
parent cf127ef223
commit 6f1d649a52
5 changed files with 121 additions and 64 deletions
@@ -140,7 +140,7 @@ We plan to go deeper into selecting the best model based on performance, cost, i
## Create a Neural Search Service with Fastembed ## Create a Neural Search Service with Fastembed
Now that you’re familiar with the core concepts around vector embeddings, how about start building your own [Neural Search Service](/documentation/tutorials/neural-search-fastembed/)? Now that you’re familiar with the core concepts around vector embeddings, how about start building your own [Neural Search Service](/documentation/tutorials/neural-search/)?
Tutorial guides you through a practical application of how to use Qdrant for document management based on descriptions of companies from [startups-list.com](https://www.startups-list.com/). From embedding data, integrating it with Qdrant's vector database, constructing a search API, and finally deploying your solution with FastAPI. Tutorial guides you through a practical application of how to use Qdrant for document management based on descriptions of companies from [startups-list.com](https://www.startups-list.com/). From embedding data, integrating it with Qdrant's vector database, constructing a search API, and finally deploying your solution with FastAPI.
@@ -1,36 +1,35 @@
--- ---
title: Neural Search with Fastembed title: Hybrid Search with Fastembed
weight: 2 weight: 2
aliases:
- /documentation/tutorials/neural-search-fastembed/
--- ---
# Create a Neural Search Service with Fastembed # Create a Hybrid Search Service with Fastembed
| Time: 20 min | Level: Beginner | Output: [GitHub](https://github.com/qdrant/qdrant_demo/) | | Time: 20 min | Level: Beginner | Output: [GitHub](https://github.com/qdrant/qdrant_demo/) |
| --- | ----------- | ----------- |----------- | | --- | ----------- | ----------- |----------- |
This tutorial shows you how to build and deploy your own neural search service to look through descriptions of companies from [startups-list.com](https://www.startups-list.com/) and pick the most similar ones to your query. This tutorial shows you how to build and deploy your own hybrid search service to look through descriptions of companies from [startups-list.com](https://www.startups-list.com/) and pick the most similar ones to your query.
The website contains the company names, descriptions, locations, and a picture for each entry. The website contains the company names, descriptions, locations, and a picture for each entry.
Alternatively, you can use datasources such as [Crunchbase](https://www.crunchbase.com/), but that would require obtaining an API key from them. As we have already written on our [blog](/articles/hybrid-search/), there is no single definition of hybrid search.
In this tutorial we are covering the case with a combination of dense and [sparse embeddings](/articles/sparse-vectors/).
The former ones refer to the embeddings generated by such well-known neural networks as BERT, while the latter ones are more related to a traditional full-text search approach.
Our neural search service will use [Fastembed](https://github.com/qdrant/fastembed) package to generate embeddings of text descriptions and [FastAPI](https://fastapi.tiangolo.com/) to serve the search API. Our hybrid search service will use [Fastembed](https://github.com/qdrant/fastembed) package to generate embeddings of text descriptions and [FastAPI](https://fastapi.tiangolo.com/) to serve the search API.
Fastembed natively integrates with Qdrant client, so you can easily upload the data into Qdrant and perform search queries. Fastembed natively integrates with Qdrant client, so you can easily upload the data into Qdrant and perform search queries.
![Hybrid Search Schema](/documentation/tutorials/hybrid-search-with-fastembed/hybrid-search-schema.png)
<aside role="status">
There is a version of this tutorial that uses <a href="https://www.sbert.net/">SentenceTransformers</a> model inference engine instead of Fastembed.
Check it out <a href="/documentation/tutorials/neural-search/">here</a>.
</aside>
## Workflow ## Workflow
To create a neural search service, you will need to transform your raw data and then create a search function to manipulate it. To create a hybrid search service, you will need to transform your raw data and then create a search function to manipulate it.
First, you will 1) download and prepare a sample dataset using a modified version of the BERT ML model. Then, you will 2) load the data into Qdrant, 3) create a neural search API and 4) serve it using FastAPI. First, you will 1) download and prepare a sample dataset using a modified version of the BERT ML model. Then, you will 2) load the data into Qdrant, 3) create a hybrid search API and 4) serve it using FastAPI.
![Neural Search Workflow](/docs/workflow-neural-search.png) ![Hybrid Search Workflow](/docs/workflow-neural-search.png)
> **Note**: The code for this tutorial can be found here: [Step 2: Full Code for Neural Search](https://github.com/qdrant/qdrant_demo/).
## Prerequisites ## Prerequisites
@@ -42,7 +41,7 @@ To complete this tutorial, you will need:
## Prepare sample dataset ## Prepare sample dataset
To conduct a neural search on startup descriptions, you must first encode the description data into vectors. To conduct a hybrid search on startup descriptions, you must first encode the description data into vectors.
Fastembed integration into qdrant client combines encoding and uploading into a single step. Fastembed integration into qdrant client combines encoding and uploading into a single step.
It also takes care of batching and parallelization, so you don't have to worry about it. It also takes care of batching and parallelization, so you don't have to worry about it.
@@ -92,10 +91,10 @@ All data uploaded to Qdrant is saved inside the `./qdrant_storage` directory and
1. Install the official Python client to best interact with Qdrant. 1. Install the official Python client to best interact with Qdrant.
```bash ```bash
pip install qdrant-client[fastembed] pip install "qdrant-client[fastembed]>=1.8.2"
``` ```
> **Note:** This tutorial requires fastembed of version >=0.2.6.
Note, that you need to install the `fastembed` extra to enable Fastembed integration.
At this point, you should have startup records in the `startups_demo.json` file and Qdrant running on a local machine. At this point, you should have startup records in the `startups_demo.json` file and Qdrant running on a local machine.
Now you need to write a script to upload all startup data and vectors into the search engine. Now you need to write a script to upload all startup data and vectors into the search engine.
@@ -106,38 +105,50 @@ Now you need to write a script to upload all startup data and vectors into the s
# Import client library # Import client library
from qdrant_client import QdrantClient from qdrant_client import QdrantClient
client = QdrantClient("http://localhost:6333") client = QdrantClient(url="http://localhost:6333")
``` ```
3. Select model to encode your data. 3. Select model to encode your data.
You will be using a pre-trained model called `sentence-transformers/all-MiniLM-L6-v2`. You will be using two pre-trained models to compute dense and sparse vectors correspondingly: `sentence-transformers/all-MiniLM-L6-v2` and `prithivida/Splade_PP_en_v1`.
<aside role="status">
Hybrid search implementation can be easily switched to a dense vector search by omitting the lines related to sparse vectors.
</aside>
```python ```python
client.set_model("sentence-transformers/all-MiniLM-L6-v2") client.set_model("sentence-transformers/all-MiniLM-L6-v2")
# comment this line to use dense vectors only
client.set_sparse_model("prithivida/Splade_PP_en_v1")
``` ```
4. Related vectors need to be added to a collection. Create a new collection for your startup vectors. 4. Related vectors need to be added to a collection. Create a new collection for your startup vectors.
```python ```python
client.recreate_collection( client.recreate_collection(
collection_name="startups", collection_name="startups",
vectors_config=client.get_fastembed_vector_params(), vectors_config=client.get_fastembed_vector_params(),
# comment this line to use dense vectors only
sparse_vectors_config=client.get_fastembed_sparse_vector_params(),
) )
``` ```
Note, that we use `get_fastembed_vector_params` to get the vector size and distance function from the model. Qdrant requires vectors to have their own names and configurations.
This method automatically generates configuration, compatible with the model you are using.
Methods `get_fastembed_vector_params` and `get_fastembed_sparse_vector_params` help you to get the corresponding parameters for the models you are using.
These parameters include vector size, distance function, etc.
Without fastembed integration, you would need to specify the vector size and distance function manually. Read more about it [here](/documentation/tutorials/neural-search/). Without fastembed integration, you would need to specify the vector size and distance function manually. Read more about it [here](/documentation/tutorials/neural-search/).
Additionally, you can specify extended configuration for our vectors, like `quantization_config` or `hnsw_config`. Additionally, you can specify extended configuration for your vectors, like `quantization_config` or `hnsw_config`.
5. Read data from the file. 5. Read data from the file.
```python ```python
payload_path = os.path.join(DATA_DIR, "startups_demo.json") import json
payload_path = "startups_demo.json"
metadata = [] metadata = []
documents = [] documents = []
@@ -148,7 +159,7 @@ with open(payload_path) as fd:
metadata.append(obj) metadata.append(obj)
``` ```
In this block of code, we read data we read data from `startups_demo.json` file and split it into 2 lists: `documents` and `metadata`. In this block of code, we read data from `startups_demo.json` file and split it into 2 lists: `documents` and `metadata`.
Documents are the raw text descriptions of startups. Metadata is the payload associated with each startup, such as the name, location, and picture. Documents are the raw text descriptions of startups. Metadata is the payload associated with each startup, such as the name, location, and picture.
We will use `documents` to encode the data into vectors. We will use `documents` to encode the data into vectors.
@@ -160,14 +171,64 @@ client.add(
collection_name="startups", collection_name="startups",
documents=documents, documents=documents,
metadata=metadata, metadata=metadata,
parallel=0, # Use all available CPU cores to encode data parallel=0, # Use all available CPU cores to encode data.
# Requires wrapping code into if __name__ == '__main__' block
) )
``` ```
The `add` method will encode all documents and upload them to Qdrant. <aside role="status">
This is one of two fastembed-specific methods, that combines encoding and uploading into a single step. Vector generation process might be time-consuming. In order to save time, you can skip this step by uploading already processed data (available under the spoiler).
</aside>
The `parallel` parameter controls the number of CPU cores used to encode data. <details>
<summary>Upload processed data</summary>
Download and unpack the processed data from [here](https://storage.googleapis.com/dataset-startup-search/startup-list-com/startups_hybrid_search_processed_40k.tar.gz) or use the following script:
```bash
wget https://storage.googleapis.com/dataset-startup-search/startup-list-com/startups_hybrid_search_processed_40k.tar.gz
tar -xvf startups_hybrid_search_processed_40k.tar.gz
```
Then you can upload the data to Qdrant.
```python
from typing import List
import json
import numpy as np
from qdrant_client import models
def named_vectors(vectors: List[float], sparse_vectors: List[models.SparseVector]) -> dict:
# make sure to use the same client object as previously
# or `set_model_name` and `set_sparse_model_name` manually
dense_vector_name = client.get_vector_field_name()
sparse_vector_name = client.get_sparse_vector_field_name()
for vector, sparse_vector in zip(vectors, sparse_vectors):
yield {
dense_vector_name: vector,
sparse_vector_name: models.SparseVector(**sparse_vector),
}
with open("dense_vectors.npy", "rb") as f:
vectors = np.load(f)
with open("sparse_vectors.json", "r") as f:
sparse_vectors = json.load(f)
with open("payload.json", "r",) as f:
payload = json.load(f)
client.upload_collection(
"startups", vectors=named_vectors(vectors, sparse_vectors), payload=payload
)
```
</details>
The `add` method will encode all documents and upload them to Qdrant.
This is one of the two fastembed-specific methods, that combines encoding and uploading into a single step.
The `parallel` parameter enables data-parallelism instead of built-in ONNX parallelism.
Additionally, you can specify ids for each document, if you want to use them later to update or delete documents. Additionally, you can specify ids for each document, if you want to use them later to update or delete documents.
If you don't specify ids, they will be generated automatically and returned as a result of the `add` method. If you don't specify ids, they will be generated automatically and returned as a result of the `add` method.
@@ -185,29 +246,32 @@ client.add(
) )
``` ```
> **Note**: See the full code for this step [here](https://github.com/qdrant/qdrant_demo/blob/master/qdrant_demo/init_collection_startups.py).
## Build the search API ## Build the search API
Now that all the preparations are complete, let's start building a neural search class. Now that all the preparations are complete, let's start building a neural search class.
In order to process incoming requests, neural search will need 2 things: 1) a model to convert the query into a vector and 2) the Qdrant client to perform search queries. In order to process incoming requests, the hybrid search class will need 3 things: 1) models to convert the query into a vector, 2) the Qdrant client to perform search queries, 3) fusion function to re-rank dense and sparse search results.
Fastembed integration into qdrant client combines encoding and uploading into a single method call.
Fastembed integration encapsulates query encoding, search and fusion into a single method call.
Fastembed leverages [reciprocal rank fusion](https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf) in order combine the results.
1. Create a file named `neural_searcher.py` and specify the following. 1. Create a file named `hybrid_searcher.py` and specify the following.
```python ```python
from qdrant_client import QdrantClient from qdrant_client import QdrantClient
class NeuralSearcher: class HybridSearcher:
DENSE_MODEL = "sentence-transformers/all-MiniLM-L6-v2"
SPARSE_MODEL = "prithivida/Splade_PP_en_v1"
def __init__(self, collection_name): def __init__(self, collection_name):
self.collection_name = collection_name self.collection_name = collection_name
# initialize Qdrant client # initialize Qdrant client
self.qdrant_client = QdrantClient("http://localhost:6333") self.qdrant_client = QdrantClient("http://localhost:6333")
self.qdrant_client.set_model("sentence-transformers/all-MiniLM-L6-v2") self.qdrant_client.set_model(self.DENSE_MODEL)
# comment this line to use dense vectors only
self.qdrant_client.set_sparse_model(self.SPARSE_MODEL)
``` ```
2. Write the search function. 2. Write the search function.
@@ -218,10 +282,12 @@ def search(self, text: str):
collection_name=self.collection_name, collection_name=self.collection_name,
query_text=text, query_text=text,
query_filter=None, # If you don't want any filters for now query_filter=None, # If you don't want any filters for now
limit=5, # 5 the closest results are enough limit=5, # 5 the closest results
) )
# `search_result` contains found vector ids with similarity scores along with the stored payload # `search_result` contains found vector ids with similarity scores
# In this function you are interested in payload only # along with the stored payload
# Select and return metadata
metadata = [hit.metadata for hit in search_result] metadata = [hit.metadata for hit in search_result]
return metadata return metadata
``` ```
@@ -239,14 +305,14 @@ from qdrant_client.models import Filter
city_of_interest = "Berlin" city_of_interest = "Berlin"
# Define a filter for cities # Define a filter for cities
city_filter = Filter(**{ city_filter = models.Filter(
"must": [{ must=[
"key": "city", # Store city information in a field of the same name models.FieldCondition(
"match": { # This condition checks if payload field has the requested value key="city",
"value": "city_of_interest" match=models.MatchValue(value=city_of_interest)
} )
}] ]
}) )
search_result = self.qdrant_client.query( search_result = self.qdrant_client.query(
collection_name=self.collection_name, collection_name=self.collection_name,
@@ -280,18 +346,18 @@ The service will have only one API endpoint and will look like this:
```python ```python
from fastapi import FastAPI from fastapi import FastAPI
# The file where NeuralSearcher is stored # The file where HybridSearcher is stored
from neural_searcher import NeuralSearcher from hybrid_searcher import HybridSearcher
app = FastAPI() app = FastAPI()
# Create a neural searcher instance # Create a neural searcher instance
neural_searcher = NeuralSearcher(collection_name="startups") hybrid_searcher = HybridSearcher(collection_name="startups")
@app.get("/api/search") @app.get("/api/search")
def search_startup(q: str): def search_startup(q: str):
return {"result": neural_searcher.search(text=q)} return {"result": hybrid_searcher.search(text=q)}
if __name__ == "__main__": if __name__ == "__main__":
@@ -314,13 +380,4 @@ You should be able to see a debug interface for your service.
Feel free to play around with it, make queries regarding the companies in our corpus, and check out the results. Feel free to play around with it, make queries regarding the companies in our corpus, and check out the results.
## Next steps
The code from this tutorial has been used to develop a [live online demo](https://qdrant.to/semantic-search-demo).
You can try it to get an intuition for cases when the neural search is useful.
The demo contains a switch that selects between neural and full-text searches.
You can turn the neural search on and off to compare your result with a regular full-text search.
> **Note**: The code for this tutorial can be found here: [Full Code for Neural Search](https://github.com/qdrant/qdrant_demo/).
Join our [Discord community](https://qdrant.to/discord), where we talk about vector search and similarity learning, publish other examples of neural networks and neural search applications. Join our [Discord community](https://qdrant.to/discord), where we talk about vector search and similarity learning, publish other examples of neural networks and neural search applications.
@@ -14,7 +14,7 @@ A neural search service uses artificial neural networks to improve the accuracy
<aside role="status"> <aside role="status">
There is a version of this tutorial that uses <a href="https://github.com/qdrant/fastembed">Fastembed</a> model inference engine instead of Sentence Transformers. There is a version of this tutorial that uses <a href="https://github.com/qdrant/fastembed">Fastembed</a> model inference engine instead of Sentence Transformers.
Check it out <a href="/documentation/tutorials/neural-search-fastembed/">here</a>. Check it out <a href="/documentation/tutorials/hybrid-search-fastembed/">here</a>.
</aside> </aside>
Binary file not shown.

Before

Width:  |  Height:  |  Size: 122 KiB

After

Width:  |  Height:  |  Size: 40 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 71 KiB