Files
landing_page/qdrant-landing/content/qdrant-cloud/qdrant-cloud-faq.md
T
nastyapashandtrean f145f0a500 Update Cloud page (#2542)
* update Cloud page

* fix image paths

* fix image paths

* update images and text on the button

* add links and target=_blank

* add missing links, updated the cards so the entire card is clickable

* add graphana dashboard gh link

---------

Co-authored-by: trean <trean.mi@gmail.com>
2026-07-31 20:08:21 +02:00

9.2 KiB

title, list, sitemapExclude
title list sitemapExclude
FAQs
id title questions
0 Getting Started
id question answer
0 Is there a free tier? Yes. The free tier is a single-node cluster with 0.5 vCPU, 1GB RAM, and 4GB disk. That fits roughly one million 768-dimension vectors. No credit card. Free clusters suspend after 1 week of inactivity and delete after 4 weeks if you don't reactivate them.
id question answer
1 What clouds and regions are supported? AWS, GCP, and Azure across multiple regions. The current questions shows up in the cluster creation flow and on the pricing page. New regions get added when customers ask for them.
id question answer
2 Which clients and SDKs do you support? Official SDKs for Python, TypeScript, Rust, Go, Java, and .NET. The REST and gRPC APIs are documented and identical to open-source Qdrant, so any community client built against the OSS engine works against Cloud.
id question answer
3 Can I bring my own embedding model? Yes. Qdrant Cloud is model-agnostic, so you can bring vectors from OpenAI, Cohere, Voyage, or your own fine-tunes. If you want one less vendor, Qdrant Cloud Inference generates text and image embeddings inside the cluster using models like MiniLM, SPLADE, BM25, Mixedbread Embed-Large, and CLIP. Paid clusters get up to 5 million free tokens per model per month, with no token cap on BM25.
id title questions
1 Pricing and Billing
id question answer
0 How is pricing calculated? Resource-based: vCPU, RAM, and storage, billed hourly. Cloud Inference adds usage charges only when you call paid embedding models. The pricing page has a sizing calculator.
id question answer
1 Are there overage charges or surprise bills? No. Cluster sizes are explicit and any change needs your authorization. Capacity alerts fire at 80% so you can upgrade on your own terms before anything breaks.
id question answer
2 Do you offer startup or research discounts? Yes. The Qdrant for Startups program offers a 20% Qdrant Cloud discount for 12 months. Eligibility: pre-seed, seed, or Series A, under 5 years old, under $5M in funding, building an AI product (agencies and dev shops don't qualify). Apply at qdrant.tech/qdrant-for-startups.
id title questions
2 Performance and Scale
id question answer
0 What query latency should I expect? For a typical 10M-vector collection at 768 dimensions with moderate filtering, P50 lands in the single-digit milliseconds and P99 stays under 50ms. Numbers shift with dimension count, recall target, filter selectivity, and cluster size. Reproducible benchmarks at qdrant.tech/benchmarks.
id question answer
1 How fast can I ingest data? CPU indexing handles tens of thousands of vectors per second per node, depending on dimension and HNSW parameters. GPU-accelerated HNSW indexing delivers up to 4x faster index construction for bulk loads. GPU clusters are available today on AWS, with other clouds on the roadmap.
id question answer
2 How does quantization affect recall? Scalar quantization typically loses 1% to 2% recall and cuts RAM by 4x. TurboQuant (Google's algorithm, integrated into Qdrant) sits in the middle: 4-bit gives roughly 8x compression with recall close to scalar, up to 32x at higher ratios. Binary with rescoring keeps recall close to full precision while cutting memory by up to 32x for compatible embedding models. The cluster UI shows recall before and after so you can pick the trade-off you want.
id title questions
3 Clusters and Scaling
id question answer
0 How do I scale up or down? Vertical scaling adds vCPU, RAM, or disk to existing nodes. Horizontal scaling adds nodes and rebalances shards automatically. Both run from the dashboard or the API. If your collections aren't replicated, a vertical scale takes a short downtime window. Replicated collections scale without interruption.
id question answer
1 What happens when my cluster gets full? If RAM or disk usage stays above 80% for 5 minutes, the account owner gets an email alert. From there you can scale vertically (more capacity per node) or horizontally (more nodes), or delete data. No surprise lockouts, no surprise bills.
id question answer
2 Is multi-region supported? Multi-AZ within a single region is available on the Premium Multi-AZ tier. It needs a minimum of three nodes and scales in multiples of three. Multi-region active-active replication isn't a managed feature yet. Teams that need it run separate clusters per region today.
id title questions
4 Reliability
id question answer
0 What's your SLA? 99.5% uptime on Standard. 99.9% on Premium. Up to 99.95% on Premium Multi-AZ. The full SLA, including uptime definitions and service-credit terms, lives at qdrant.to/sla.
id question answer
1 How do backups work? Scheduled incremental snapshots on AWS and GCP, with configurable retention. Azure backups bill based on total disk usage. You can also take on-demand snapshots before risky changes and restore to the same cluster or a new one. For long-term retention or compliance, you can export snapshots to your own object storage.
id question answer
2 What's your disaster recovery posture? You set the snapshot cadence to match your RPO. Premium Multi-AZ deployments replicate across three availability zones (cross-AZ replication, not failover) with no failover delay. If a zone goes down, reads and writes continue from the surviving zones with no customer action required. For cross-region retention, export snapshots to your own object storage in any region.
id title questions
5 Security and Compliance
id question answer
0 Are you SOC 2, GDPR, and HIPAA compliant? Yes. SOC 2 Type II report and HIPAA certification on file. GDPR-compliant Data Processing Agreement available. Email Solutions Engineering for current compliance documentation and BAA scope.
id question answer
1 How is data encrypted? TLS in transit. Storage volumes encrypted at rest. Premium customers can use customer-managed keys for disk encryption, plus SSO and VPC private links. Snapshots and backups inherit the same encryption.
id question answer
2 What does audit logging capture? Every API operation: queries, upserts, deletes, collection management, and snapshot operations. Each entry is structured JSON with caller identity, timestamp, target collection, and the decision (allowed or denied). You can retrieve logs through an API endpoint, configure retention to match your policy, and download them for long-term storage in your own systems. Available on all paid clusters.
id question answer
3 Are you EU-based? Yes. Qdrant is headquartered in Berlin. EU customers can keep data in EU regions exclusively, which addresses US Cloud Act concerns and similar extraterritorial regimes.
id title questions
6 Migration and Lock-In
id question answer
0 Is migrating from Qdrant Open Source to Cloud difficult? No. Our open-source migration tool turns it into a configuration change, not a rewrite. Application code points at a new endpoint; business logic stays intact.
id question answer
1 How do I migrate from another vector database? The migration tool covers Pinecone, Weaviate, Milvus, Chroma, Redis, MongoDB, OpenSearch, Elasticsearch, pgvector, S3 Vectors, FAISS, Apache Solr, and Qdrant-to-Qdrant (for example, OSS to Cloud). The bigger lift is usually re-running the ingestion pipeline; the data move itself is incremental and resumes if interrupted. Solutions Engineering will pair on a migration plan if you ask.
id question answer
2 Can I move workloads back from Cloud to OSS later? Yes. Export a snapshot, restore it on your own infrastructure running open-source Qdrant. The engine and data format are identical. Your data is yours.
id question answer
3 What happens if Qdrant Cloud goes away? The engine is open source under Apache 2.0, with 30k+ GitHub stars and a 60k-member Discord community. You can run it yourself indefinitely. Engine parity keeps your options open whatever happens to the managed service.
id title questions
7 Support
id question answer
0 What support tiers do you offer? Community (Discord, free), Standard (10x5 business hours, Mon-Fri 08:00 to 18:00 CET), and Premium (24x7 critical incident response with priority response times).
id question answer
1 How fast do you respond? We use four severity levels. Standard customers get a Sev 1 response in 4 business hours, Sev 2 in 6, Sev 3 in 24. Premium customers get Sev 1 in 1 hour, Sev 2 in 2, Sev 3 in 4 business hours. Real engineers, not a tier-1 chatbot.
id question answer
2 Do you have a community channel? Yes. The Qdrant Discord is open to all developers, including the engineering team. discord.gg/qdrant
true