show how to use local inference in reranking hybrid search (#1591)

* show how to use local inference in reranking hybrid search

* Update reranking-hybrid-search.md

* fix: remove accidentally added code

---------

Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
This commit is contained in:
George
2025-04-26 13:10:19 +03:00
committed by GitHub
co-authored by Andrey Vasnetsov
parent 488f34382e
commit ad82b164ea
@@ -158,6 +158,37 @@ operation_info = client.upsert(
)
```
<aside role="status">
Check how points can be uploaded with builtin Fastembed integration.
</aside>
<details>
<summary>Upload with implicit embeddings computation</summary>
```python
from qdrant_client.models import PointStruct
points = []
for idx, doc in enumerate(documents):
point = PointStruct(
id=idx,
vector={
"all-MiniLM-L6-v2": models.Document(text=doc, model="sentence-transformers/all-MiniLM-L6-v2"),
"bm25": models.Document(text=doc, model="Qdrant/bm25"),
"colbertv2.0": models.Document(text=doc, model="colbert-ir/colbertv2.0"),
},
payload={"document": doc}
)
points.append(point)
operation_info = client.upsert(
collection_name="hybrid-search",
points=points
)
```
</details>
---
This code pulls everything together by creating a list of **PointStruct** objects, each containing the embeddings and corresponding documents.
@@ -223,6 +254,39 @@ results = client.query_points(
)
```
<aside role="status">
Check how queries can be made with builtin Fastembed integration.
</aside>
<details>
<summary>Query points with implicit embeddings computation</summary>
```python
prefetch = [
models.Prefetch(
query=models.Document(text=query, model="sentence-transformers/all-MiniLM-L6-v2"),
using="all-MiniLM-L6-v2",
limit=20,
),
models.Prefetch(
query=models.Document(text=query, model="Qdrant/bm25"),
using="bm25",
limit=20,
),
]
results = client.query_points(
"hybrid-search",
prefetch=prefetch,
query=models.Document(text=query, model="colbert-ir/colbertv2.0"),
using="colbertv2.0",
with_payload=True,
limit=10,
)
```
</details>
---
Let’s look at how the positions change after applying reranking. Notice how some documents shift in rank based on their relevance according to the late interaction embeddings.
@@ -248,4 +312,4 @@ Reranking can dramatically improve the relevance of search results, especially w
Reranking is a powerful tool that boosts the relevance of search results, especially when combined with hybrid search methods. While it can add some latency due to its complexity, applying it to a smaller, pre-filtered subset of results ensures both speed and relevance.
Qdrant offers an easy-to-use API to get started with your own search engine, so if you’re ready to dive in, sign up for free at [Qdrant Cloud](https://qdrant.tech/) and start building
Qdrant offers an easy-to-use API to get started with your own search engine, so if you’re ready to dive in, sign up for free at [Qdrant Cloud](https://qdrant.tech/) and start building