mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-09 04:48:30 +02:00
* Break Inference page into several pages * Make all inference code snippets testable and clean up * Make more snippets testable * Edits * Document automatic query and passage prefix injection in Cloud Inference Qdrant Cloud Inference silently applies model-specific prefixes (e.g. "query: "/"passage: " for E5, BGE-style instruction prefix for BGE/mxbai/ Snowflake arctic-embed) so users don't need to manage them manually. Add a section explaining this behavior, the idempotency guarantee, and the scope (Qdrant-hosted models only; external providers handle their own). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * Document short query optimization in Cloud Inference Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * Update links * Expand on external provider API key usage * Add section about external provider API keys * Default to header for external API keys --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
67 lines
4.5 KiB
Markdown
67 lines
4.5 KiB
Markdown
---
|
|
title: Cloud Inference
|
|
weight: 30
|
|
---
|
|
|
|
# Qdrant Cloud Inference
|
|
|
|

|
|
|
|
Clusters on Qdrant Managed Cloud can use [Qdrant Cloud Inference](/documentation/cloud/inference/) to generate embeddings. For a list of available models and their dimensions, visit the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. You can also enable Cloud Inference for a cluster from this tab.
|
|
|
|
Several embedding models are free to use with Qdrant Cloud Inference, including on free-tier clusters. The “Cost: Free” label in the Inference tab identifies these models.
|
|
|
|
Qdrant Cloud Inference automatically handles short search queries locally, reducing latency for search inputs. This applies to supported Qdrant-hosted models and is fully transparent. Longer search inputs and upserts are handled by a dedicated remote inference service.
|
|
|
|
Before using a Cloud-hosted embedding model, ensure that your collection has been configured for vectors with the correct dimensionality. The Inference tab of the Cluster Detail page in the Qdrant Cloud Console lists the dimensionality for each supported embedding model.
|
|
|
|
## Text Inference
|
|
|
|
Let's consider an example of using Cloud Inference with a text model that produces dense vectors. This example creates one point and uses a simple search query with a `Document` Inference Object.
|
|
|
|
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/simple/" >}}
|
|
|
|
Usage examples, specific to each cluster and model, can also be found in the Inference tab of the Cluster Detail page in the Qdrant Cloud Console.
|
|
|
|
Note that each model has a context window, which is the maximum number of tokens that can be processed by the model in a single request. If the input text exceeds the context window, it is truncated to fit within the limit. The context window size is displayed in the Inference tab of the Cluster Detail page.
|
|
|
|
For dense vector models, you also have to ensure that the vector size configured in the collection matches the output size of the model. If the vector size does not match, the upsert will fail with an error.
|
|
|
|
## Image Inference
|
|
|
|
Here is another example of using Cloud Inference with an image model. This example uses the `CLIP` model to encode an image and then uses a text query to search for it.
|
|
|
|
Since the `CLIP` model is multimodal, we can use both image and text inputs on the same vector field.
|
|
|
|
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/image/" >}}
|
|
|
|
The Qdrant Cloud Inference server will download the images using the provided URL. Alternatively, you can provide the image as a base64-encoded string. Each model has limitations on the file size and extensions it can work with. Refer to the model card for details.
|
|
|
|
## Automatic Query and Passage Prefixes
|
|
|
|
Some embedding models are trained to expect different text prefixes for query inputs versus passage or document inputs. Using the wrong prefix or omitting one can degrade retrieval quality.
|
|
|
|
Qdrant Cloud Inference automatically applies the correct prefix for the model you're using, based on whether the request is a search query or an upsert. For example, when you upsert a document using a model from the E5 family, the proxy prepends `passage: ` to each input text. When you query with the same model, it prepends `query: ` instead. You don't need to add these prefixes yourself.
|
|
|
|
If your input text already starts with the correct prefix, it's left unchanged, so you don't end up with a double prefix.
|
|
|
|
This behavior applies to supported Qdrant-hosted models. Models accessed through [external providers](/documentation/inference/external-inference-providers/), such as `openai/`, `cohere/`, or `jinaai/`, are handled by those providers.
|
|
|
|
## Local Inference Compatibility
|
|
|
|
The Python SDK offers a unique capability: it supports both [local](/documentation/fastembed/fastembed-semantic-search/) and cloud inference through an identical interface.
|
|
|
|
You can easily switch between local and cloud inference by setting the `cloud_inference` flag when initializing the QdrantClient. For example:
|
|
|
|
```python
|
|
client = QdrantClient(
|
|
url="https://your-cluster.qdrant.io",
|
|
api_key="<your-api-key>",
|
|
cloud_inference=True, # Set to False to use local inference
|
|
)
|
|
```
|
|
|
|
This flexibility allows you to develop and test your applications locally or in continuous integration (CI) environments without requiring access to cloud inference resources.
|
|
|
|
* When `cloud_inference` is set to `False`, inference is performed locally using `fastembed`.
|
|
* When set to `True`, inference requests are handled by Qdrant Cloud. |