Cloud inference example (#1768)

* Cloud inference example

* Cloud inference example

* update URL

* Update qdrant-landing/content/documentation/examples/cloud-inference-hybrid-search.md

Co-authored-by: Bastian Hofmann <mail@bastianhofmann.de>

* Update qdrant-landing/content/documentation/examples/cloud-inference-hybrid-search.md

Co-authored-by: Bastian Hofmann <mail@bastianhofmann.de>

* Update qdrant-landing/content/documentation/examples/cloud-inference-hybrid-search.md

Co-authored-by: Bastian Hofmann <mail@bastianhofmann.de>

* Update qdrant-landing/content/documentation/examples/cloud-inference-hybrid-search.md

Co-authored-by: Bastian Hofmann <mail@bastianhofmann.de>

* Update qdrant-landing/content/documentation/examples/cloud-inference-hybrid-search.md

Co-authored-by: Bastian Hofmann <mail@bastianhofmann.de>

* Update qdrant-landing/content/documentation/examples/cloud-inference-hybrid-search.md

Co-authored-by: Bastian Hofmann <mail@bastianhofmann.de>

* Update qdrant-landing/content/documentation/examples/cloud-inference-hybrid-search.md

Co-authored-by: Bastian Hofmann <mail@bastianhofmann.de>

* Update qdrant-landing/content/documentation/examples/cloud-inference-hybrid-search.md

Co-authored-by: Bastian Hofmann <mail@bastianhofmann.de>

* Update qdrant-landing/content/documentation/examples/cloud-inference-hybrid-search.md

Co-authored-by: Bastian Hofmann <mail@bastianhofmann.de>

* Move cloud inference tutorial

* move below support

* rename folder

* update cloud example

* update cloud example

* update cloud example

* update cloud example

* update cloud example

* update cloud example

* update cloud example

* update cloud example

* Link inference docs

* Link inference docs

* Create collection snippet

* Create snippets

* Create snippets

* Create snippets

* Create snippets

* Create snippets

* Create snippets

* Create snippets

* Create snippets

* Update qdrant-landing/content/documentation/headless/snippets/cloud-inference/vector-search/upload-data/python.md

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* Update qdrant-landing/content/documentation/tutorials-and-examples/cloud-inference-hybrid-search.md

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* Update qdrant-landing/content/documentation/tutorials-and-examples/_index.md

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* Update qdrant-landing/content/documentation/headless/snippets/cloud-inference/vector-search/upload-data/python.md

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* Update qdrant-landing/content/documentation/tutorials-and-examples/cloud-inference-hybrid-search.md

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* Update qdrant-landing/content/documentation/headless/snippets/cloud-inference/vector-search/run-vector-search/python.md

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* Update qdrant-landing/content/documentation/tutorials-and-examples/cloud-inference-hybrid-search.md

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* Update qdrant-landing/content/documentation/headless/snippets/cloud-inference/vector-search/initialize-client/_description.md

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* Update qdrant-landing/content/documentation/headless/snippets/cloud-inference/vector-search/run-vector-search/_description.md

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* Update qdrant-landing/content/documentation/headless/snippets/cloud-inference/vector-search/create-sample-query/python.md

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* Update qdrant-landing/content/documentation/headless/snippets/cloud-inference/vector-search/create-collection/python.md

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* Update qdrant-landing/content/documentation/tutorials-and-examples/cloud-inference-hybrid-search.md

---------

Co-authored-by: Bastian Hofmann <mail@bastianhofmann.de>
Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>
This commit is contained in:
Derrick Mwiti
2025-07-15 14:08:37 +03:00
committed by GitHub
co-authored by Bastian Hofmann Kacper Łukawski
parent b40514678f
commit 5eff7dfc24
13 changed files with 169 additions and 1 deletions
@@ -36,3 +36,4 @@ Our Notebooks offer complex instructions that are supported with a throrough exp
| [Extractive QA System](https://githubtocolab.com/qdrant/examples/blob/master/extractive_qa/extractive-question-answering.ipynb) | Extract answers directly from context to generate highly relevant answers. | Qdrant |
| [Ecommerce Reverse Image Search](https://githubtocolab.com/qdrant/examples/blob/master/ecommerce_reverse_image_search/ecommerce-reverse-image-search.ipynb) | Accept images as search queries to receive semantically appropriate answers. | Qdrant |
| [Basic RAG](https://githubtocolab.com/qdrant/examples/blob/master/rag-openai-qdrant/rag-openai-qdrant.ipynb) | Basic RAG pipeline with Qdrant and OpenAI SDKs. | OpenAI, Qdrant, FastEmbed |
@@ -0,0 +1 @@
This code snippet creates a collection configured for hybrid search using Qdrant's Cloud Inference. It defines a sparse BM25 vector and a dense vector for MiniLM. This setup allows Qdrant to perform hybrid search using the dense and sparse vectors.
@@ -0,0 +1,21 @@
```python
from qdrant_client import models
collection_name = "my_collection_name"
if not client.collection_exists(collection_name=collection_name):
client.create_collection(
collection_name=collection_name,
vectors_config={
"dense_vector": models.VectorParams(
size=384,
distance=models.Distance.COSINE
)
},
sparse_vectors_config={
"bm25_sparse_vector": models.SparseVectorParams(
modifier=models.Modifier.IDF # Enable Inverse Document Frequency
)
}
)
```
@@ -0,0 +1 @@
This code snippet creates sample query for hybrid search using Qdrant's cloud inference.
@@ -0,0 +1,3 @@
```python
query_text = "What is relapsing polychondritis?"
```
@@ -0,0 +1 @@
This code snippet connects to Qdrant Cloud with Cloud Inference enabled. This is done by setting `cloud_inference` to `True` in the initializer of the `QdrantClient` class.
@@ -0,0 +1,8 @@
```python
client = QdrantClient(
url="https://YOUR_URL.eastus-0.azure.cloud.qdrant.io:6333/",
api_key="YOUR_API_KEY",
cloud_inference=True,
timeout=30.0
)
```
@@ -0,0 +1 @@
This code snippet demonstrates how to use the Universal Query API to prefetch results with dense and sparse vector search, and then rerank them with Reciprocal Rank Fusion. It uses Cloud Inference to create embeddings by passing document text along with the model names, instead of vectors.
@@ -0,0 +1,28 @@
```python
results = client.query_points(
collection_name=collection_name,
prefetch=[
models.Prefetch(
query=Document(
text=query_text,
model=dense_model
),
using="dense_vector",
limit=5
),
models.Prefetch(
query=Document(
text=query_text,
model=bm25_model
),
using="bm25_sparse_vector",
limit=5
)
],
query=models.FusionQuery(fusion=models.Fusion.RRF),
limit=5,
with_payload=True
)
print(results.points)
```
@@ -0,0 +1 @@
This code snipet shows how to load a dataset and upload dense and sparse vectors to Qdrant. While uploading the vectors we also include a payload known as text.
@@ -0,0 +1,38 @@
```python
from qdrant_client.http.models import PointStruct, Document
from datasets import load_dataset
import uuid
dense_model = "sentence-transformers/all-minilm-l6-v2"
bm25_model = "qdrant/bm25"
ds = load_dataset("miriad/miriad-4.4M", split="train[0:100]")
points = []
for idx, item in enumerate(ds):
passage = item["passage_text"]
point = PointStruct(
id=uuid.uuid4().hex, # use unique string ID
payload=item,
vector={
"dense_vector": Document(
text=passage,
model=dense_model
),
"bm25_sparse_vector": Document(
text=passage,
model=bm25_model
)
}
)
points.append(point)
client.upload_points(
collection_name=collection_name,
points=points,
batch_size=8
)
```
@@ -0,0 +1,12 @@
---
title: Tutorials & Examples
weight: 40
partition: cloud
---
## Cloud Tutorials & Examples
| Example | Description |
| ----------------------------------- | ------------------------------------------------------------------------------------------- |
| [Using Cloud Inference to Build Hybrid Search](/documentation/tutorials-and-examples/cloud-inference-hybrid-search/) | Cloud inference hybrid example |
@@ -0,0 +1,52 @@
---
title: Using Cloud Inference to Build Hybrid Search
weight: 35
---
# Using Cloud Inference with Qdrant for Vector Search
In this tutorial, we'll walkthrough building a **hybrid semantic search engine** using Qdrant Cloud's built-in [inference](/documentation/cloud/inference/) capabilities. You'll learn how to:
- Automatically embed your data using [cloud Inference](/documentation/cloud/inference/) without needing to run local models,
- Combine dense semantic embeddings with [sparse BM25 keywords](https://qdrant.tech/documentation/advanced-tutorials/reranking-hybrid-search/), and
- Perform hybrid search using [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/concepts/hybrid-queries/) to retrieve the most relevant results.
## Install Qdrant Client
```bash
pip install qdrant-client datasets
```
## Initialize the Client
Initialize the Qdrant client after creating a [Qdrant Cloud account](/documentation/cloud/) and a [dedicated paid cluster](/documentation/cloud/create-cluster/). Set `cloud_inference` to `True` to enable [cloud inference](/documentation/cloud/inference/).
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/vector-search/initialize-client/" >}}
## Create a Collection
Qdrant stores vectors and associated metadata in collections. A collection requires vector parameters to be set during creation. In this case, let's set up a collection using `BM25` for sparse vectors and `all-minilm-l6-v2` for dense vectors. BM25 uses the Inverse Document Frequency to reduce the weight of common terms that appear in many documents while boosting the importance of rare terms that are more discriminative for retrieval. Qdrant will handle the calculations of the IDF term if we enable that in the configuration of the `bm25_sparse_vector` named sparse vector.
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/vector-search/create-collection/" >}}
## Add Data
Now you can add sample documents, their associated metadata, and a point id for each. Here's a sample of the [miriad/miriad-4.4M](https://huggingface.co/datasets/miriad/miriad-4.4M) dataset:
| qa_id | paper_id | question | year | venue | specialty | passage_text |
|--------------------|----------|-------------------------------------------------------|------|--------------------------------------|--------------|--------------------------------------------------------|
| 38_77498699_0_1 | 77498699 | What are the clinical features of relapsing polychondritis? | 2006 | Internet Journal of Otorhinolaryngology | Rheumatology | A 45-year-old man presented with painful swelling... |
| 38_77498699_0_2 | 77498699 | What treatments are available for relapsing polychondritis? | 2006 | Internet Journal of Otorhinolaryngology | Rheumatology | Patient showed improvement after treatment with... |
| 38_88124321_0_3 | 88124321 | How is Takayasu arteritis diagnosed? | 2015 | Journal of Autoimmune Diseases | Rheumatology | A 32-year-old woman with fatigue and limb pain... |
We won't ingest all the entries from the dataset, but for demo purposes, just take the first hundred ones:
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/vector-search/upload-data/" >}}
## Set Up Input Query
Create a sample query:
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/vector-search/create-sample-query/" >}}
## Run Vector Search
Here, you will ask a question that will allow you to retrieve semantically relevant results. The final results are obtained by reranking using [Reciprocal Rank Fusion](https://qdrant.tech/documentation/concepts/hybrid-queries/#hybrid-search).
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/vector-search/run-vector-search/" >}}
The semantic search engine will retrieve the most similar result in order of relevance.
```markdown
[ScoredPoint(id='9968a760-fbb5-4d91-8549-ffbaeb3ebdba',
version=0, score=14.545895,
payload={'text': "Relapsing Polychondritis is a rare..."},
vector=None, shard_key=None, order_value=None)]
```