From 1d8a4169b80fd1dee3588400868937421d894931 Mon Sep 17 00:00:00 2001 From: daniel-azoulai Date: Mon, 14 Jul 2025 08:46:10 -0700 Subject: [PATCH 01/11] qdrant-cloud-inference-blog --- .../Qdrant Cloud Inference - Launch Blog.md | 69 ++++++++++++++++++ .../qdrant-cloud-inference/how-it-works.jpg | Bin 0 -> 252410 bytes .../inference-architecture.jpg | Bin 0 -> 102821 bytes .../inference-how-it-works.png | Bin 0 -> 133506 bytes .../qdrant-cloud-inference/inference-ui.png | Bin 0 -> 53994 bytes 5 files changed, 69 insertions(+) create mode 100644 qdrant-landing/content/blog/Qdrant Cloud Inference - Launch Blog.md create mode 100644 qdrant-landing/static/blog/qdrant-cloud-inference/how-it-works.jpg create mode 100644 qdrant-landing/static/blog/qdrant-cloud-inference/inference-architecture.jpg create mode 100644 qdrant-landing/static/blog/qdrant-cloud-inference/inference-how-it-works.png create mode 100644 qdrant-landing/static/blog/qdrant-cloud-inference/inference-ui.png diff --git a/qdrant-landing/content/blog/Qdrant Cloud Inference - Launch Blog.md b/qdrant-landing/content/blog/Qdrant Cloud Inference - Launch Blog.md new file mode 100644 index 000000000..ad341324a --- /dev/null +++ b/qdrant-landing/content/blog/Qdrant Cloud Inference - Launch Blog.md @@ -0,0 +1,69 @@ +--- +draft: false +title: "Introducing Qdrant Cloud Inference" +short_description: "Run inferencing natively in Qdrant Cloud" +description: "Learn how to generate embeddings natively in Qdrant Cloud" +preview_image: /blog/qdrant-cloud-inference/inference-how-it-works.png +social_preview_image: /blog/qdrant-cloud-inference/inference-how-it-works.png +date: 2025-07-14T00:00:00Z +author: Daniel Azoulai +featured: false +tags: +- vector search +- enterprise +- monitoring +- RBAC +- observability +--- + +# Introducing Qdrant Cloud Inference + +Today, we’re announcing the launch of Qdrant Cloud Inference. With Qdrant Cloud Inference, users can generate, store and index embeddings in a single API call, turning unstructured text and images into search-ready vectors in a single environment. Directly integrating model inference into Qdrant Cloud removes the need for separate inference infrastructure, manual pipelines, and redundant data transfers. + +This simplifies workflows, accelerates development cycles, and eliminates unnecessary network hops for developers. With a single API call, you can now embed, store, and index your data more quickly and more simply. This speeds up application development for RAG, Multimodal, Hybrid search, and more. + +## Unify embedding and search + +Traditionally, building application data pipelines means juggling separate embedding services and a vector database, introducing unnecessary complexity, latency, and network costs. Qdrant Cloud Inference brings everything into one system. Embeddings are generated inside the network of your cluster, which removes external API overhead, resulting in lower latency and faster response times. Additionally, you can now track vector database and inference costs in one place. + +![architecture](/blog/qdrant-cloud-inference/inference-architecture.jpg) + +## Supported Models for Multimodal and Hybrid Search Applications + +At launch, Qdrant Cloud Inference includes six curated models to start with. Choose from dense models like `all-MiniLM-L6-v2` for fast semantic matching, `mxbai/embed-large-v1` for richer understanding, or sparse models like `splade-pp-en-v1` and `bm25`. For multimodal workloads, Qdrant uniquely supports `OpenAI CLIP`\-style models for both text and images. + +*Want to request a different model to integrate? You can do this at [https://support.qdrant.io/](https://support.qdrant.io/).* + +![architecture](/blog/qdrant-cloud-inference/inference-ui.png) + +## Get up to 5M free tokens per model per month + +To make onboarding even easier, we’re offering 5 million free tokens per text model, 1 million for our image model, and unlimited for `bm25` to all paid Qdrant Cloud users. These token allowances renew monthly so long as you have a paid Qdrant Cloud cluster. These free monthly tokens are perfect for development, staging, or even running initial production workloads without added cost. + +## Inference is automatically enabled for paid accounts + +Getting started is easy. Inference is automatically enabled for any new paid clusters with version 1.14.0 or higher. It can be activated for existing clusters with a click on the inference tab on the Cluster Detail page in the Qdrant Cloud console. You will see examples of how to use inference with our different Qdrant SDKs. + +*\ ## Start Embedding Today From d91fc67a53487751ec9009f83f29b416393f6951 Mon Sep 17 00:00:00 2001 From: daniel-azoulai Date: Mon, 14 Jul 2025 13:23:52 -0700 Subject: [PATCH 07/11] Update qdrant-cloud-inference-launch.md --- .../content/blog/qdrant-cloud-inference-launch.md | 11 ++++++----- 1 file changed, 6 insertions(+), 5 deletions(-) diff --git a/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md b/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md index 39e9a4560..77673ce10 100644 --- a/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md +++ b/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md @@ -60,9 +60,10 @@ Kacper Łukawski, Senior Developer Advocate, is hosting a live session on Thursd We'll show you how to: -* Generate embeddings for text or images using pre-integrated models -* Store and search embeddings in the same Qdrant Cloud environment -* Power multimodal (an industry first) and hybrid search with just one API -* Reduce network egress (for non on-prem deployments) fees and simplify your AI stack - + [**Save your spot**](https://try.qdrant.tech/cloud-inference). \ No newline at end of file From 1145f220e63ee4299d6b838b8f2ae35f353ff45b Mon Sep 17 00:00:00 2001 From: daniel-azoulai Date: Mon, 14 Jul 2025 13:29:00 -0700 Subject: [PATCH 08/11] Update qdrant-cloud-inference-launch.md --- qdrant-landing/content/blog/qdrant-cloud-inference-launch.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md b/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md index 77673ce10..a15e318b0 100644 --- a/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md +++ b/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md @@ -29,7 +29,7 @@ Traditionally, building application data pipelines means juggling separate embed ## Supported Models for Multimodal and Hybrid Search Applications -At launch, Qdrant Cloud Inference includes six curated models to start with. Choose from dense models like `all-MiniLM-L6-v2` for fast semantic matching, `mxbai/embed-large-v1` for richer understanding, or sparse models like `splade-pp-en-v1` and `bm25`. For multimodal workloads, Qdrant uniquely supports `OpenAI CLIP`-style models for both text and images. +At launch, Qdrant Cloud Inference includes six curated models to start with. Choose from dense models like `all-MiniLM-L6-v2` for fast semantic matching, `mxbai/embed-large-v1` for richer understanding, or sparse models like `splade-pp-en-v1` and `bm25` ([Check out this hybrid search tutorial to see it in action](https://qdrant.tech/documentation/tutorials-and-examples/cloud-inference-hybrid-search/)). For multimodal workloads, Qdrant uniquely supports `OpenAI CLIP`-style models for both text and images. *Want to request a different model to integrate? You can do this at [https://support.qdrant.io/](https://support.qdrant.io/).* @@ -66,4 +66,5 @@ We'll show you how to:
  • Power multimodal (an industry first) and hybrid search with just one API
  • Reduce network egress fees and simplify your AI stack
  • + [**Save your spot**](https://try.qdrant.tech/cloud-inference). \ No newline at end of file From 9ef6241e5c47343a3f2b62285afd7d3dc76f21cb Mon Sep 17 00:00:00 2001 From: daniel-azoulai Date: Mon, 14 Jul 2025 13:37:18 -0700 Subject: [PATCH 09/11] Update qdrant-cloud-inference-launch.md --- qdrant-landing/content/blog/qdrant-cloud-inference-launch.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md b/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md index a15e318b0..367f24039 100644 --- a/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md +++ b/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md @@ -67,4 +67,5 @@ We'll show you how to:
  • Reduce network egress fees and simplify your AI stack
  • -[**Save your spot**](https://try.qdrant.tech/cloud-inference). \ No newline at end of file + +### [**Save your spot**](https://try.qdrant.tech/cloud-inference). \ No newline at end of file From 0a35b77177e5caab1c21d8bbf4fb699fbe110980 Mon Sep 17 00:00:00 2001 From: daniel-azoulai Date: Mon, 14 Jul 2025 13:39:34 -0700 Subject: [PATCH 10/11] Update qdrant-cloud-inference-launch.md --- qdrant-landing/content/blog/qdrant-cloud-inference-launch.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md b/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md index 367f24039..abe757070 100644 --- a/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md +++ b/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md @@ -68,4 +68,4 @@ We'll show you how to: -### [**Save your spot**](https://try.qdrant.tech/cloud-inference). \ No newline at end of file +### [Save your spot](https://try.qdrant.tech/cloud-inference). \ No newline at end of file From 80298d03c7966a2c48c3ec2b2d3855d8e9446bd8 Mon Sep 17 00:00:00 2001 From: daniel-azoulai Date: Mon, 14 Jul 2025 14:29:50 -0700 Subject: [PATCH 11/11] Update qdrant-cloud-inference-launch.md --- qdrant-landing/content/blog/qdrant-cloud-inference-launch.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md b/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md index abe757070..e2e8cf64d 100644 --- a/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md +++ b/qdrant-landing/content/blog/qdrant-cloud-inference-launch.md @@ -5,7 +5,7 @@ short_description: "Run inferencing natively in Qdrant Cloud" description: "Learn how to generate embeddings natively in Qdrant Cloud" preview_image: /blog/qdrant-cloud-inference/inference-social-preview.jpg social_preview_image: /blog/qdrant-cloud-inference/inference-social-preview.jpg -date: 2025-07-14T00:00:00Z +date: 2025-07-15T00:00:00Z author: Daniel Azoulai featured: true tags: