Files
meinsta 0b2252b8c6 Rebuild Cloud Inference landing page
Rebuilds the existing /cloud-inference/ page against the signed-off Figma
frame (Inference LP): hero, Why It Matters, What You Get, How It Works,
Cloud Inference Approaches with the comparison table, FAQs, and the shared
CTA banner.

The slug and its nav entry are unchanged, so nothing is orphaned. The
previous hero, features, cluster, read-doc, get-started and FAQ sections
are replaced. Copy comes from the doc linked on MKT-425; the model-request
callout reuses the existing pricing-banner partial.

MKT-425
2026-09-11 10:35:17 -07:00

1.2 KiB

label, title, description, image, items, button, sitemapExclude
label title description image items button sitemapExclude
What You Get Inference Runs Inside Your Cluster Qdrant Cloud Inference ships a set of hosted models you call through the same API as your database.
src alt
/img/cloud-inference/inference-flow.png Your application sending text and images to an embedding model inside a Qdrant Cloud cluster, writing to collections A, B, and C
id title description
0 One call from query to result. Send raw text, image or multivectors, get ranked results back. Your application code handles one request type, covering both vectorization and retrieval in a single operation.
id title description
1 Run hybrid search at no inference cost. Pair a free dense model like all-MiniLM-L6-v2 with BM25; free models carry no token charges and are available even on free-tier clusters. SPLADE and other larger models are metered. Sparse and dense embeddings run together, so keyword-precision and semantic recall are available in the same query through the same managed endpoint. Cluster resources bill as usual.
text url
Read About Inference /documentation/inference/
true