Files
landing_page/qdrant-landing/content/documentation/cloud/inference.md
T
1a40961d62 fix: linkchecker include filter port mismatch (#2242)
* initial commit; fixed anchor links on internal docs pages

* add back in absolute paths for links in code comments

* fix: update linkchecker include filter to match server port 1314

PR #1629 changed the Hugo server to port 1314 but forgot to update
the --include filter, which still matched port 1313. This caused all
links to be excluded, making the checker a no-op (0 checked, 82277 excluded).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Fix url rewrite regex so images are not impacted

* Fix links from non-documentation pages

* Fix broken links

* more broken links

* more broken links

* broken link

* Add srcset width descriptor to .lycheeignore

* Ignore URLs that contain a % character

* Anchor regex so it matches the entire URL

---------

Co-authored-by: kanungle <neil.kanungo@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com>
2026-03-30 17:21:32 +02:00

2.7 KiB

title, weight
title weight
Inference 81

Inference in Qdrant Managed Cloud

Inference is the process of creating vector embeddings from text, images, or other data types using a machine learning model.

Qdrant Managed Cloud allows you to use inference directly in the cloud, without the need to set up and maintain your own inference infrastructure. You can use embedding models hosted on Qdrant Cloud, or use externally hosted models.

Cluster Cluster UI

Enabling/Disabling Inference

Inference is enabled by default for all new clusters, created after July, 7th 2025. You can enable it for existing clusters directly from the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. Activating inference will trigger a restart of your cluster to apply the new configuration.

Using Inference

Inference can be easily used through the Qdrant SDKs and the REST or GRPC APIs when upserting points and when querying the database. Refer to the Inference documentation for details.

Cloud Inference

Clusters on Qdrant Managed Cloud can access embedding models that are hosted on Qdrant Cloud.

You can see the list of supported models in the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. The list includes models for text, both to produce dense and sparse vectors, as well as multi-modal models for images.

Free Embedding Models

Several embedding models can be used for free with Qdrant Cloud Inference, also in combination with clusters on the Qdrant Cloud free tier. Free models are identified by the "Cost: Free" label in the Inference tab of the Cluster Detail page.

Billing

Usage of non-free embedding models is billed based on the number of tokens processed by the model. The cost is calculated per 1,000,000 tokens. The price depends on the model and is displayed on the Inference tab of the Cluster Detail page. You also can see the current usage of each model there.

Use External Models

Qdrant Cloud can act as a proxy for the following external embedding providers:

  • OpenAI
  • Cohere
  • Jina AI
  • OpenRouter

This enables you to access any of the embedding models provided by these providers through the Qdrant API.

Billing

To use an external provider's embedding model, you need an API key from that provider. Billing is managed directly through the external provider, based on API key usage. Refer to each external embedding model provider's website for pricing details.