Files
meinsta 0b2252b8c6 Rebuild Cloud Inference landing page
Rebuilds the existing /cloud-inference/ page against the signed-off Figma
frame (Inference LP): hero, Why It Matters, What You Get, How It Works,
Cloud Inference Approaches with the comparison table, FAQs, and the shared
CTA banner.

The slug and its nav entry are unchanged, so nothing is orphaned. The
previous hero, features, cluster, read-doc, get-started and FAQ sections
are replaced. Copy comes from the doc linked on MKT-425; the model-request
callout reuses the existing pricing-banner partial.

MKT-425
2026-09-11 10:35:17 -07:00

2.2 KiB

label, title, subheading, description, link, cards, sitemapExclude
label title subheading description link cards sitemapExclude
Why It Matters Pick the Right Embedding Model, Or Bring Your Own Use the embedding model that fits your use case. Start at no cost: BM25 and a dense model like all-MiniLM-L6-v2 are free to use, including on free-tier clusters. On paid clusters, models like mxbai and SPLADE run on metered tokens with a monthly free allowance of up to 5 million tokens per model. Choose from supported hosted models or bring your own API key for an external provider. Send text or images; Qdrant embeds and searches in one call.
text url
Follow the embedding model migration guide /documentation/tutorials-operations/embedding-model-migration/
id icon title description
0
src alt
/icons/outline/shield-check-blue.svg Shield
Embed without leaving your cluster Inference runs inside your Qdrant Cloud cluster's network, so every upsert and query stays on one path, with no external hops, no extra egress, and fewer moving parts to maintain.
id icon title description
1
src alt
/icons/outline/layers-blue.svg Layers
Text, image, and sparse vector models included Managed Cloud gives you access to multimodal embeddings, plus sparse vector support for BM25-style retrieval, all callable through the same API as your database.
id icon title description
2
src alt
/icons/outline/puzzle-blue.svg Puzzle
Bring your own model or provider Point the client at an <a href="/documentation/inference/external-inference-providers/">externally hosted model</a>, or run your own <a href="/documentation/fastembed/fastembed-postprocessing/">client-side inference</a> locally using Qdrant's FastEmbed library. You're never locked to a fixed model catalog.
id icon title description
3
src alt
/icons/outline/square-pen-blue.svg Edit
Swap and test without rebuilding Retrieval performance and domain specificity both depend on the embedding model you choose. Swapping models on Managed Cloud makes migration significantly easier.
true