From 1d8a4169b80fd1dee3588400868937421d894931 Mon Sep 17 00:00:00 2001 From: daniel-azoulai Date: Mon, 14 Jul 2025 08:46:10 -0700 Subject: [PATCH] qdrant-cloud-inference-blog --- .../Qdrant Cloud Inference - Launch Blog.md | 69 ++++++++++++++++++ .../qdrant-cloud-inference/how-it-works.jpg | Bin 0 -> 252410 bytes .../inference-architecture.jpg | Bin 0 -> 102821 bytes .../inference-how-it-works.png | Bin 0 -> 133506 bytes .../qdrant-cloud-inference/inference-ui.png | Bin 0 -> 53994 bytes 5 files changed, 69 insertions(+) create mode 100644 qdrant-landing/content/blog/Qdrant Cloud Inference - Launch Blog.md create mode 100644 qdrant-landing/static/blog/qdrant-cloud-inference/how-it-works.jpg create mode 100644 qdrant-landing/static/blog/qdrant-cloud-inference/inference-architecture.jpg create mode 100644 qdrant-landing/static/blog/qdrant-cloud-inference/inference-how-it-works.png create mode 100644 qdrant-landing/static/blog/qdrant-cloud-inference/inference-ui.png diff --git a/qdrant-landing/content/blog/Qdrant Cloud Inference - Launch Blog.md b/qdrant-landing/content/blog/Qdrant Cloud Inference - Launch Blog.md new file mode 100644 index 000000000..ad341324a --- /dev/null +++ b/qdrant-landing/content/blog/Qdrant Cloud Inference - Launch Blog.md @@ -0,0 +1,69 @@ +--- +draft: false +title: "Introducing Qdrant Cloud Inference" +short_description: "Run inferencing natively in Qdrant Cloud" +description: "Learn how to generate embeddings natively in Qdrant Cloud" +preview_image: /blog/qdrant-cloud-inference/inference-how-it-works.png +social_preview_image: /blog/qdrant-cloud-inference/inference-how-it-works.png +date: 2025-07-14T00:00:00Z +author: Daniel Azoulai +featured: false +tags: +- vector search +- enterprise +- monitoring +- RBAC +- observability +--- + +# Introducing Qdrant Cloud Inference + +Today, we’re announcing the launch of Qdrant Cloud Inference. With Qdrant Cloud Inference, users can generate, store and index embeddings in a single API call, turning unstructured text and images into search-ready vectors in a single environment. Directly integrating model inference into Qdrant Cloud removes the need for separate inference infrastructure, manual pipelines, and redundant data transfers. + +This simplifies workflows, accelerates development cycles, and eliminates unnecessary network hops for developers. With a single API call, you can now embed, store, and index your data more quickly and more simply. This speeds up application development for RAG, Multimodal, Hybrid search, and more. + +## Unify embedding and search + +Traditionally, building application data pipelines means juggling separate embedding services and a vector database, introducing unnecessary complexity, latency, and network costs. Qdrant Cloud Inference brings everything into one system. Embeddings are generated inside the network of your cluster, which removes external API overhead, resulting in lower latency and faster response times. Additionally, you can now track vector database and inference costs in one place. + +![architecture](/blog/qdrant-cloud-inference/inference-architecture.jpg) + +## Supported Models for Multimodal and Hybrid Search Applications + +At launch, Qdrant Cloud Inference includes six curated models to start with. Choose from dense models like `all-MiniLM-L6-v2` for fast semantic matching, `mxbai/embed-large-v1` for richer understanding, or sparse models like `splade-pp-en-v1` and `bm25`. For multimodal workloads, Qdrant uniquely supports `OpenAI CLIP`\-style models for both text and images. + +*Want to request a different model to integrate? You can do this at [https://support.qdrant.io/](https://support.qdrant.io/).* + +![architecture](/blog/qdrant-cloud-inference/inference-ui.png) + +## Get up to 5M free tokens per model per month + +To make onboarding even easier, we’re offering 5 million free tokens per text model, 1 million for our image model, and unlimited for `bm25` to all paid Qdrant Cloud users. These token allowances renew monthly so long as you have a paid Qdrant Cloud cluster. These free monthly tokens are perfect for development, staging, or even running initial production workloads without added cost. + +## Inference is automatically enabled for paid accounts + +Getting started is easy. Inference is automatically enabled for any new paid clusters with version 1.14.0 or higher. It can be activated for existing clusters with a click on the inference tab on the Cluster Detail page in the Qdrant Cloud console. You will see examples of how to use inference with our different Qdrant SDKs. + +*\