Files
landing_page/qdrant-landing/content/documentation/inference/_index.md
T
Abdon PijpelinkandClaude Sonnet 4.6 478b96554f Restructure inference docs (#2225)
* Break Inference page into several pages

* Make all inference code snippets testable and clean up

* Make more snippets testable

* Edits

* Document automatic query and passage prefix injection in Cloud Inference

Qdrant Cloud Inference silently applies model-specific prefixes (e.g.
"query: "/"passage: " for E5, BGE-style instruction prefix for BGE/mxbai/
Snowflake arctic-embed) so users don't need to manage them manually.
Add a section explaining this behavior, the idempotency guarantee, and
the scope (Qdrant-hosted models only; external providers handle their own).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* Document short query optimization in Cloud Inference

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* Update links

* Expand on external provider API key usage

* Add section about external provider API keys

* Default to header for external API keys

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-24 08:11:03 +02:00

3.1 KiB

title, short_description, description, weight, partition, aliases
title short_description description weight partition aliases
Inference Generate vector embeddings inside Qdrant — server-side inference removes the need to run separate embedding infrastructure. Use Qdrant's built-in inference to generate text, image, and multimodal embeddings server-side — no separate embedding stack required. 220 develop
../inference
/documentation/concepts/inference/

Inference

Inference generates vectors embeddings from documents or images

Inference is the process of using a machine learning model to create vector embeddings from text, images, or other data types. While you can create embeddings on the client side, you can also use Qdrant's Inference API to generate them server-side using a single, unified API across sparse, managed, and externally hosted models.

Inference Options

Where and how you generate embeddings depends on the model you want to use and whether you prefer to manage your own embedding infrastructure:

  • Client-side inference: Manage your own inference pipeline locally, for example, using Qdrant's Python FastEmbed library. This provides full control over the model and its configuration without external network calls and is ideal when you prefer to manage the embedding infrastructure yourself.
  • Qdrant Cluster (BM25): Generate sparse embeddings using the BM25 model directly within the Qdrant cluster. This keeps the embedding logic close to the data, eliminating the need for a separate inference service for keyword-based retrieval.
  • Qdrant Cloud Inference: Managed deployments on Qdrant Cloud have access to Cloud Inference. Qdrant Cloud hosts a range of embedding models, some for free, allowing you to generate embeddings without managing the infrastructure.
  • Externally Hosted Models (Qdrant Cloud): Access embedding models hosted by third-party embedding model providers (OpenAI, Cohere, Jina AI, and OpenRouter). Use a wide range of state-of-the-art models through a single, unified Qdrant API, without the need to manage a separate embedding pipeline.

Choose Your Approach

The right option depends on your deployment and what you need to embed. Use this table to find the best fit for your use case.

If you… Use…
Need sparse BM25 embeddings Qdrant Cluster
Already manage your own inference service Client-side inference
Self-host Qdrant Client-side inference, for example using FastEmbed
Use Qdrant Cloud and want to use one of the supported embedding models Qdrant Cloud Inference
Use Qdrant Cloud and want to use a model from OpenAI, Cohere, Jina AI, or OpenRouter Qdrant Cloud Inference with an external provider (requires API key)