Merge branch 'master' into retrival-quality-guide-improvement

This commit is contained in:
Dylan Couzon
2026-04-23 12:24:29 -04:00
325 changed files with 9038 additions and 727 deletions
+3
View File
@@ -2,3 +2,6 @@
__pycache__
qdrant-landing/.DS_Store
.DS_Store
.tools/
package-lock.json
.claude/
@@ -5,6 +5,8 @@ edition = "2024"
[dependencies]
anyhow = "1.0.100"
chrono = "0.4"
csv = "1.3"
qdrant-edge = "0.6.0"
fs-err = "3"
ordered-float = "5"
@@ -33,8 +33,8 @@ features:
icon:
src: /icons/outline/filter-blue.svg
alt: Filter
title: Single Stage filtering That Works
description: Qdrant enhances search speeds and control and context understanding through filtering on any nested entry in our payload. Unique architecture allows Qdrant to avoid expensive pre-filtering and post-filtering stages, making search faster and accurate.
title: Pre-Filtering that Works
description: Qdrant enhances search speeds and control and context understanding through pre-filtering on any nested entry in our payload. This filter works first – making your search faster and combines with our hybrid search and multi-modal search seamlessly.
link:
text: Learn More
url: /articles/filterable-hnsw/
@@ -5,10 +5,11 @@ startFree:
text: Get Started
url: https://cloud.qdrant.io/signup
learnMore:
text: Contact Us
text: Learn More
url: /contact-us/
image:
src: /img/vectors/vector-0.svg
src: /img/home/horizontal-slider/advanced-search.png
webp: /img/home/horizontal-slider/advanced-search.webp
alt: Advanced search
sitemapExclude: true
---
@@ -48,7 +48,7 @@ features:
description: Qdrant’s architecture is optimized for high-throughput embedding processing, minimizing CPU load and preventing performance bottlenecks. This enables AI agents in Agentic RAG workflows to execute complex, multi-step tasks efficiently, ensuring smooth operation even at scale.
link:
text: Distributed Deployment
url: /documentation/operations/distributed_deployment/
url: /documentation/distributed_deployment/
- id: 4
icon:
src: /icons/outline/speedometer-blue.svg
@@ -109,13 +109,13 @@ Once you’ve built a fast, accurate, secure, and scalable agent, how do you kno
Your agent’s ability to complete complex tasks is only as good as the context it can retrieve. It is crucial to closely and continuously monitor the agent’s performance with grounding checks to spot and prevent hallucinations, recall@k to ensure relevance in search, and MMR parameters for diversity.
#### [**Performance-Cost Tradeoff**](https://qdrant.tech/documentation/operations/optimize/)
#### [**Performance-Cost Tradeoff**](https://qdrant.tech/documentation/ops-optimization/optimize/)
In production agents, efficiency is one of, if not the, most important metric to track. First, in enterprise environments you must meet strict latency budgets. Evaluations also let you track the cost per task by tracking token usage and end-to-end compute time. Finally, you can track the effectiveness of your memory layer by monitoring cache hit rates for your memory banks.
![Tradeoff Triangle](/articles_data/agentic-builders-guide/tradeoff-triangle.png)
#### [**Guardrails & Fallbacks**](https://qdrant.tech/documentation/operations/security/)
#### [**Guardrails & Fallbacks**](https://qdrant.tech/documentation/security/)
Just like humans, agents aren’t perfect. A production agentic system should expect and anticipate failures and have guardrails to handle them gracefully. You can also use a human in the loop when confidence scores are below a chosen threshold or the query touches on a high-stakes or sensitive topic. To handle a wide range of queries, your agent should use hybrid search. Hybrid search combines the power of semantic search for understanding meaning with the precision of keyword search for exact matches, giving you the best of both worlds.
@@ -125,7 +125,7 @@ The same agent that speeds through a toy dataset with 10,000 points will become
We’ll talk about three concepts you can take advantage of to improve your scale, but if you want even more information on how to scale, check out this [article](https://qdrant.tech/documentation/database-tutorials/large-scale-search/) on large scale search.
As your dataset and traffic grow, Qdrant Cloud offers a suite of features to ensure your system can scale effectively. Horizontal scaling is achieved through [sharding](https://qdrant.tech/articles/multitenancy/), which splits your collection across multiple nodes to distribute the load and improve performance. For high availability and fault tolerance, Qdrant supports [replication](https://qdrant.tech/documentation/operations/distributed_deployment/), creating copies of your shards across the cluster.
As your dataset and traffic grow, Qdrant Cloud offers a suite of features to ensure your system can scale effectively. Horizontal scaling is achieved through [sharding](https://qdrant.tech/articles/multitenancy/), which splits your collection across multiple nodes to distribute the load and improve performance. For high availability and fault tolerance, Qdrant supports [replication](https://qdrant.tech/documentation/distributed_deployment/), creating copies of your shards across the cluster.
Qdrant provides robust tools for resource and cost optimization. Vector [quantization](https://qdrant.tech/documentation/manage-data/quantization/) compresses your data, significantly reducing its memory footprint and speeding up search.
@@ -143,7 +143,7 @@ Workflows that are effective for a single user or a handful of users in a develo
Qdrant provides production-grade [authorization and authentication](https://qdrant.tech/documentation/cloud/authentication/) via API keys, multitenancy, and Role-Based Access Control (RBAC) to make sure your agent doesn’t go rogue.
Authentication for agents is handled by [API keys](https://qdrant.tech/documentation/operations/security/#api-keys). With Qdrant, API keys are more than just a password. They act as smart credentials that also carry the details of the authorization rules that are enforced once the agent’s identity is confirmed. These keys can be dynamically created as temporary credentials for each user session, making access limited to the session and secure.
Authentication for agents is handled by [API keys](https://qdrant.tech/documentation/security/#api-keys). With Qdrant, API keys are more than just a password. They act as smart credentials that also carry the details of the authorization rules that are enforced once the agent’s identity is confirmed. These keys can be dynamically created as temporary credentials for each user session, making access limited to the session and secure.
Note: Qdrant also supports concurrent queries, so your search won’t slow down as more users are writing queries simultaneously.
@@ -165,7 +165,7 @@ Your AI agent needs to be able to run without issues wherever your users are run
Qdrant lets you deploy in a variety of ways so you can build the system that best fits your use case. You can deploy via the cloud or self host on Docker or your own machine. Read more below:
* [Qdrant Cloud](https://qdrant.tech/documentation/cloud-intro/)
* [Qdrant Cloud](https://qdrant.tech/documentation/deploy-intro/)
* [Hybrid Cloud](https://qdrant.tech/documentation/hybrid-cloud/)
* [Self-Hosted](https://qdrant.tech/documentation/quickstart/)
@@ -55,7 +55,7 @@ On Qdrant Cloud, you can create API keys using the [Cloud Dashboard](https://qdr
For on-premise or local deployments, you'll need to configure API key authentication. This involves specifying a key in either the Qdrant configuration file or as an environment variable. This ensures that all requests to the server must include a valid API key sent in the header.
When using the simple API key-based authentication, you should also turn on TLS encryption. Otherwise, you are exposing the connection to sniffing and MitM attacks. To secure your connection using TLS, you would need to create a certificate and private key, and then [enable TLS](/documentation/operations/security/#tls) in the configuration.
When using the simple API key-based authentication, you should also turn on TLS encryption. Otherwise, you are exposing the connection to sniffing and MitM attacks. To secure your connection using TLS, you would need to create a certificate and private key, and then [enable TLS](/documentation/security/#tls) in the configuration.
API authentication, coupled with TLS encryption, offers a first layer of security for your Qdrant instance. However, to enable more granular access control, the recommended approach is to leverage JSON Web Tokens (JWTs).
@@ -128,7 +128,7 @@ With index maintenance divided between segments, Qdrant can ensure high performa
| | |
|---------------------|-------------|
| **Mutable Segments** | These are used for quickly ingesting new data and handling changes (updates) to existing data. |
| **Immutable Segments** | Once a mutable segment reaches a certain size, an optimization process converts it into an immutable segment, constructing an HNSW index – you could [**read about these optimizers here**](/documentation/operations/optimizer/#optimizer) in detail. This immutability trick allowed us, for example, to ensure effective [**tenant isolation**](/documentation/manage-data/indexing/#tenant-index). |
| **Immutable Segments** | Once a mutable segment reaches a certain size, an optimization process converts it into an immutable segment, constructing an HNSW index – you could [**read about these optimizers here**](/documentation/ops-optimization/optimizer/#optimizer) in detail. This immutability trick allowed us, for example, to ensure effective [**tenant isolation**](/documentation/manage-data/indexing/#tenant-index). |
Immutable segments are an implementation detail transparent for users — they can delete vectors at any time, while additions and updates are applied to a mutable segment instead. This combination of mutability and immutability allows search and indexing to smoothly run simultaneously, even under heavy loads. This approach minimizes the performance impact of indexing time and allows on-the-fly configuration changes on a collection level (such as enabling or disabling data quantization) without downtimes.
@@ -224,7 +224,7 @@ docker-compose up -d
```
The demo will be available at `http://localhost:8001`, but you won't be able to search anything until you [import the snapshot into your Qdrant
instance](/documentation/operations/snapshots/#recover-via-api). If you don't want to bother with hosting a local one, you can use the [Qdrant
instance](/documentation/snapshots/#recover-via-api). If you don't want to bother with hosting a local one, you can use the [Qdrant
Cloud](https://cloud.qdrant.io/) cluster. 4 GB RAM is enough to load all the 2 million entries.
## Fork and reuse
@@ -111,10 +111,13 @@ configuration:
```yaml
# within the storage config
storage:
performance:
# enable the async scorer which uses io_uring
async_scorer: true
```
or `QDRANT__STORAGE__PERFORMANCE__ASYNC_SCORER=true`.
You can return to the mmap based backend by either deleting the `async_scorer`
entry or setting the value to `false`.
@@ -19,14 +19,14 @@ category: vector-search-manuals
# Scaling Your Machine Learning Setup: The Power of Multitenancy and Custom Sharding in Qdrant
We are seeing the topics of [multitenancy](/documentation/manage-data/multitenancy/) and [distributed deployment](/documentation/operations/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup.
We are seeing the topics of [multitenancy](/documentation/manage-data/multitenancy/) and [distributed deployment](/documentation/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup.
Whether you are building a bank fraud-detection system, [RAG](https://qdrant.tech/articles/what-is-rag-in-ai/) for e-commerce, or services for the federal government - you will need to leverage a multitenant architecture to scale your product.
In the world of SaaS and enterprise apps, this setup is the norm. It will considerably increase your application's performance and lower your hosting costs.
## Multitenancy & custom sharding with Qdrant
We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](/documentation/manage-data/multitenancy/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](/documentation/operations/distributed_deployment/#user-defined-sharding).
We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](/documentation/manage-data/multitenancy/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](/documentation/distributed_deployment/#user-defined-sharding).
Combining these two will result in an efficiently-partitioned architecture that further leverages the convenience of a single Qdrant cluster. This article will briefly explain the benefits and show how you can get started using both features.
@@ -41,7 +41,7 @@ Qdrant is built to excel in a single collection with a vast number of tenants. Y
## Sharding your database
With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](/documentation/operations/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node.
With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](/documentation/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node.
During vector search, your operations will be able to hit only the subset of shards they actually need. In massive-scale deployments, __this can significantly improve the performance of operations that do not require the whole collection to be scanned__.
@@ -49,7 +49,7 @@ This works in the other direction as well. Whenever you search for something, yo
### Common use cases
A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](/documentation/operations/distributed_deployment/#moving-shards).
A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](/documentation/distributed_deployment/#moving-shards).
**Figure 2:** Users can both upsert and query shards that are relevant to them, all within the same collection. Regional sharding can help avoid cross-continental traffic.
![Qdrant Multitenancy](/articles_data/multitenancy/shards.png)
@@ -79,7 +79,7 @@ client.create_shard_key("{tenant_data}", "germany")
```
In this example, your cluster is divided between Germany and Canada. Canadian and German law differ when it comes to international data transfer. Let's say you are creating a RAG application that supports the healthcare industry. Your Canadian customer data will have to be clearly separated for compliance purposes from your German customer.
Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](/documentation/operations/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech).
Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](/documentation/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech).
## Configure a multitenant setup for users
@@ -195,7 +195,7 @@ Out-of-Memory errors.
Qdrant 1.2 enters recovery mode, if enabled, when it detects a failure on startup.
That makes the service halt the loading of collection data and commence operations in a partial state.
This state allows for removing collections but doesn't support search or update functions.
**Recovery mode [has to be enabled by user](/documentation/operations/administration/#recovery-mode).**
**Recovery mode [has to be enabled by user](/documentation/ops-configuration/administration/#recovery-mode).**
### Appendable mmap
@@ -206,7 +206,7 @@ to read more about segments, check out our docs on [vector storage](/documentati
## Security
There are two major changes in terms of [security](/documentation/operations/security/):
There are two major changes in terms of [security](/documentation/security/):
1. **API-key support** - basic authentication with a static API key to prevent unwanted access. Previously
API keys were only supported in [Qdrant Cloud](https://cloud.qdrant.io/).
@@ -92,7 +92,7 @@ POST /collections/my_collection/points/search
}
```
If you want to know more about the user-defined sharding, please refer to the [sharding documentation](/documentation/operations/distributed_deployment/#sharding).
If you want to know more about the user-defined sharding, please refer to the [sharding documentation](/documentation/distributed_deployment/#sharding).
### Snapshot-based shard transfer
@@ -101,7 +101,7 @@ That's a really more in depth technical improvement for the distributed mode use
Moving shards is required for dynamical scaling of the cluster. Your data can migrate between nodes, and the way you move it is crucial for the performance of the whole system. The good old `stream_records` method (still the default one) transmits all the records between the machines and indexes them on the target node.
In the case of moving the shard, it's necessary to recreate the HNSW index each time. However, with the introduction of the new `snapshot` approach, the snapshot itself, inclusive of all data and potentially quantized content, is transferred to the target node. This comprehensive snapshot includes the entire index, enabling the target node to seamlessly load it and promptly begin handling requests without the need for index recreation.
There are multiple scenarios in which you may prefer one over the other. Please check out the docs of the [shard transfer method](/documentation/operations/distributed_deployment/#shard-transfer-method) for more details and head-to-head comparison. As for now, the old `stream_records` method is still the default one, but we may decide to change it in the future.
There are multiple scenarios in which you may prefer one over the other. Please check out the docs of the [shard transfer method](/documentation/distributed_deployment/#shard-transfer-method) for more details and head-to-head comparison. As for now, the old `stream_records` method is still the default one, but we may decide to change it in the future.
## Minor improvements
@@ -65,7 +65,7 @@ This isn't mandatory, as Qdrant is by default tuned to strike the right balance
This version introduces a `optimizer_cpu_budget` parameter to control the maximum number of CPUs used for indexing.
> Read more about `config.yaml` in the [configuration file](/documentation/operations/configuration/).
> Read more about `config.yaml` in the [configuration file](/documentation/ops-configuration/configuration/).
```yaml
# CPU budget, how many CPUs (threads) to allocate for an optimization job.
@@ -88,7 +88,7 @@ Sometimes users don't properly balance HNSW search parameters. Setting the HNSW
❓ **Use Case:** A customer ran advanced similarity searches across their vast dataset of nearly 800 million vectors. Initially, they found that queries took anywhere from 10 to 20 seconds, especially when combining multiple filters and metadata fields.
> ✅ How can they retain accuracy and keep things fast? [**The answer is optimization.**](https://qdrant.tech/documentation/operations/optimize/)
> ✅ How can they retain accuracy and keep things fast? [**The answer is optimization.**](https://qdrant.tech/documentation/ops-optimization/optimize/)
**Figure 1:** Qdrant is highly configurable. You can configure it for speed, precision or resource use.
![qdrant resource tradeoffs](/docs/tradeoff.png)
@@ -99,7 +99,7 @@ This strategy balanced memory usage with performance: only the compact vectors n
||
|-|
|**Read More:** [**Optimization Guide**](https://qdrant.tech/documentation/operations/optimize/)|#optimizing-qdrant-performance-three-scenarios
|**Read More:** [**Optimization Guide**](https://qdrant.tech/documentation/ops-optimization/optimize/)|#optimizing-qdrant-performance-three-scenarios
|**Read More:** [**HNSW Documentation**](https://qdrant.tech/documentation/manage-data/indexing/#vector-index)|
### Compress Your Data with Quantization Strategies
@@ -155,7 +155,7 @@ Once all records are inserted, you can rebuild the index in a single pass. Consi
||
|-|
|**Read More:** [**Configuration Documentation**](https://qdrant.tech/documentation/operations/configuration/)|
|**Read More:** [**Configuration Documentation**](https://qdrant.tech/documentation/ops-configuration/configuration/)|
### When Indexing Falls Behind Ingestion
![vector-search-production](/articles_data/vector-search-production/vector-search-production-3.jpg)
@@ -230,7 +230,7 @@ Figure: For many-tenant setups, spinning up a new collection per tenant can ball
|:-:|
|**"How many nodes, CPUs, RAM and storage do I need for my Qdrant Cluster?"**|
It depends. If you're just starting out - we have prepared a tool on our website to help you figure this out. For more information, [**check out the Capacity Planning document as well.**](https://qdrant.tech/documentation/operations/capacity-planning/)
It depends. If you're just starting out - we have prepared a tool on our website to help you figure this out. For more information, [**check out the Capacity Planning document as well.**](https://qdrant.tech/documentation/capacity-planning/)
✅ [**Use the sizing calculator**](https://cloud.qdrant.io/calculator) or performance testing to ensure node specs (RAM/CPU) match your workload.
@@ -242,7 +242,7 @@ It depends. If you're just starting out - we have prepared a tool on our website
A three-node setup provides a baseline for fault tolerance: if one node goes offline, the remaining two can continue serving queries and maintain a quorum for data consistency. This guards against hardware failures, rolling updates, and network disruptions. Fewer than three nodes leaves you vulnerable to single-point failures that can knock your entire cluster offline.
> [**We follow the Raft Protocol**](https://qdrant.tech/documentation/operations/distributed_deployment/#raft), so check out the docs and learn why this is important.
> [**We follow the Raft Protocol**](https://qdrant.tech/documentation/distributed_deployment/#raft), so check out the docs and learn why this is important.
✅ **Set a replication factor of at least 2** to tolerate node failure without losing availability.
@@ -274,7 +274,7 @@ Development and staging environments often run experimental builds, tests, or si
> It's quite possible that the user has multiple shards on one node, which end up handling most traffic while other nodes remain underutilized.
In this case, you should [**choose the right number of shards**](https://qdrant.tech/documentation/operations/distributed_deployment/#sharding) based on your node count and expected RPS.
In this case, you should [**choose the right number of shards**](https://qdrant.tech/documentation/distributed_deployment/#sharding) based on your node count and expected RPS.
You need to implement a shard strategy that aligns with real usage patterns. First, distribute your shards across all available nodes. This will help balance the load more effectively. After redistributing the shards, run performance tests to see how it affects your system. Then add replicas and test again to see how that changes performance.
@@ -285,14 +285,14 @@ Proper sharding considers data distribution and query patterns. By default, shar
||
|-|
|**Read More:** [**Sharding Documentation**](https://qdrant.tech/documentation/operations/distributed_deployment/#sharding)|
|**Read More:** [**Sharding Documentation**](https://qdrant.tech/documentation/distributed_deployment/#sharding)|
### Manage Your Costs by Scaling Up or Down
![vector-search-production](/articles_data/vector-search-production/vector-search-production-5.jpg)
Some teams scale up for daytime surges, then scale down overnight to save resources. If you do this, ensure data is sharded and replicated appropriately, so that scaling up and down won't result in service degradation.
If using Qdrant Cloud you could also do this using the [**Replication Factor**](https://qdrant.tech/documentation/operations/distributed_deployment/#replication-factor), though it may be considered a bit of a hack.
If using Qdrant Cloud you could also do this using the [**Replication Factor**](https://qdrant.tech/documentation/distributed_deployment/#replication-factor), though it may be considered a bit of a hack.
> If you have 3 nodes with just 1 shard, and replication factor 6. It will create 3 replicas (one on each node) of that shard, because it can't host more. If you add 3 more nodes at peak times, it'll automatically replicate that shard 3 more times in an attempt to match the factor of 6.
@@ -308,7 +308,7 @@ If new nodes remain empty after joining, you waste resources. If departing nodes
||
|-|
|**Read More:** [**Distributed Deployment Documentation**](https://qdrant.tech/documentation/operations/distributed_deployment/)|
|**Read More:** [**Distributed Deployment Documentation**](https://qdrant.tech/documentation/distributed_deployment/)|
|**Read More:** [**Resharding**](https://qdrant.tech/documentation/cloud/cluster-scaling/#resharding)|
### How to Predict and Test Cluster Performance
@@ -331,7 +331,7 @@ Remember, cold-starts and query behaviour are dataset dependent, which is why yo
||
|-|
|**Read More:** [Distributed Deployment Documentation](https://qdrant.tech/documentation/operations/distributed_deployment/)
|**Read More:** [Distributed Deployment Documentation](https://qdrant.tech/documentation/distributed_deployment/)
### How to Design Your Systems to Protect Against Failure
@@ -375,7 +375,7 @@ By following these comprehensive load testing practices, you'll be able to ident
||
|-|
|**Read More:** [**Telemetry and Monitoring Documentation**](https://qdrant.tech/documentation/operations/monitoring/)|
|**Read More:** [**Telemetry and Monitoring Documentation**](https://qdrant.tech/documentation/ops-monitoring/monitoring/)|
|**Read More:** [**Cloud Monitoring Documentation**](https://qdrant.tech/documentation/hybrid-cloud/networking-logging-monitoring/)
## 4. Ensuring Disaster Recovery With Database Backups and Snapshots
@@ -413,7 +413,7 @@ If you host tens of billions of vectors, store backups off-node in a different d
||
|-|
|**Read More:** [**Snapshot Documentation**](https://qdrant.tech/documentation/operations/snapshots/)|
|**Read More:** [**Snapshot Documentation**](https://qdrant.tech/documentation/snapshots/)|
|**Read More:** [**Managed Cloud Backup Documentation**](https://qdrant.tech/documentation/cloud/backups/)|
|**Read More:** [**Private Cloud Backup Documentation**](https://qdrant.tech/documentation/private-cloud/backups/)|
@@ -438,7 +438,7 @@ Investigations showed they hadn't adjusted the default configuration or reserved
||
|-|
|**Read More:** [**Qdrant Configuration Documentation**](https://qdrant.tech/documentation/operations/configuration/)|
|**Read More:** [**Qdrant Configuration Documentation**](https://qdrant.tech/documentation/ops-configuration/configuration/)|
### Security & Governance
@@ -450,7 +450,7 @@ Enabling TLS/HTTPS is essential for meeting compliance requirements in regulated
> You need to protect data in transit. To enable TLS/HTTPS for encrypted traffic in production, you need to configure secure communication between clients and your Qdrant database, as well as individual cluster nodes. This involves implementing Transport Layer Security (TLS) certificates to encrypt all traffic, preventing unauthorized access and data interception.
If self-hosting, you can set up encryption yourself by [**incorporating TLS directly from the configuration**](https://qdrant.tech/documentation/operations/security/#tls)
If self-hosting, you can set up encryption yourself by [**incorporating TLS directly from the configuration**](https://qdrant.tech/documentation/security/#tls)
```text
service:
@@ -469,7 +469,7 @@ tls:
||
|-|
|**Read More:** [**Security Documentation**](https://qdrant.tech/documentation/operations/security/)|
|**Read More:** [**Security Documentation**](https://qdrant.tech/documentation/security/)|
### Setting up Access Controls in Production
@@ -32,12 +32,12 @@ Let's take a look at some common goals and optimization strategies:
| Intended Result | Optimization Strategy |
|--------------------------------|------------------------------|
| [**High Search Precision + Low Memory Expenditure**](/documentation/operations/optimize/#1-high-speed-search-with-low-memory-usage) | [**On-Disk Indexing**](/documentation/operations/optimize/#1-high-speed-search-with-low-memory-usage) |
| [**High Search Precision + Low Memory Expenditure**](/documentation/ops-optimization/optimize/#1-high-speed-search-with-low-memory-usage) | [**On-Disk Indexing**](/documentation/ops-optimization/optimize/#1-high-speed-search-with-low-memory-usage) |
| [**Low Memory Expenditure + Fast Search Speed**](/documentation/manage-data/quantization/) | [**Quantization**](/documentation/manage-data/quantization/) |
| [**High Search Precision + Fast Search Speed**](/documentation/operations/optimize/#3-high-precision-with-high-speed-search) | [**RAM Storage + Quantization**](/documentation/operations/optimize/#3-high-precision-with-high-speed-search) |
| [**Balance Latency vs Throughput**](/documentation/operations/optimize/#balancing-latency-and-throughput) | [**Segment Configuration**](/documentation/operations/optimize/#balancing-latency-and-throughput) |
| [**High Search Precision + Fast Search Speed**](/documentation/ops-optimization/optimize/#3-high-precision-with-high-speed-search) | [**RAM Storage + Quantization**](/documentation/ops-optimization/optimize/#3-high-precision-with-high-speed-search) |
| [**Balance Latency vs Throughput**](/documentation/ops-optimization/optimize/#balancing-latency-and-throughput) | [**Segment Configuration**](/documentation/ops-optimization/optimize/#balancing-latency-and-throughput) |
After this article, check out the code samples in our docs on [**Qdrant’s Optimization Methods**](/documentation/operations/optimize/).
After this article, check out the code samples in our docs on [**Qdrant’s Optimization Methods**](/documentation/ops-optimization/optimize/).
---
@@ -57,7 +57,7 @@ Qdrant uses the [**HNSW (Hierarchical Navigable Small World Graph) algorithm**](
Working with massive datasets that contain billions of vectors demands significant resources—and those resources come with a price. While Qdrant provides reasonable defaults, tailoring them to your specific use case can unlock optimal performance. Here’s what you need to know.
The following parameters give you the flexibility to fine-tune Qdrant’s performance for your specific workload. You can modify them directly in Qdrant's [**configuration**](https://qdrant.tech/documentation/operations/configuration/) files or at the collection and named vector levels for more granular control.
The following parameters give you the flexibility to fine-tune Qdrant’s performance for your specific workload. You can modify them directly in Qdrant's [**configuration**](https://qdrant.tech/documentation/ops-configuration/configuration/) files or at the collection and named vector levels for more granular control.
**Figure 3:** A description of three key HNSW parameters.
@@ -325,7 +325,7 @@ Here’s how to choose the shard_number:
| **Plan for Scalability** | Start with at least **2 shards per node** to allow room for future growth. |
| **Future-Proofing** | Starting with around **12 shards** is a good rule of thumb. This setup allows your system to scale seamlessly from 1 to 12 nodes without requiring re-sharding. |
Learn more about [**Sharding in Distributed Deployment**](/documentation/operations/distributed_deployment/)
Learn more about [**Sharding in Distributed Deployment**](/documentation/distributed_deployment/)
---
@@ -584,7 +584,7 @@ Here are some important metrics to monitor:
| grpc_responses_avg_duration_seconds | | Average response duration in gRPC API |
| rest_responses_fail_total | | Total number of failed responses (REST) |
Read more about [**Qdrant Open Source Monitoring**](/documentation/operations/monitoring/) and [**Qdrant Cloud Monitoring**](/documentation/cloud/cluster-monitoring/) for managed clusters.
Read more about [**Qdrant Open Source Monitoring**](/documentation/ops-monitoring/monitoring/) and [**Qdrant Cloud Monitoring**](/documentation/cloud/cluster-monitoring/) for managed clusters.
_________________________________________________________________________
## Recap: When Should You Optimize?
@@ -414,7 +414,7 @@ client.create_collection(
We recommend using sharding and replication together so that your data is both split across nodes and replicated for availability.
For more details on features like **user-defined sharding, node failure recovery**, and **consistency guarantees**, see our guide on [Distributed Deployment.](https://qdrant.tech/documentation/operations/distributed_deployment/)
For more details on features like **user-defined sharding, node failure recovery**, and **consistency guarantees**, see our guide on [Distributed Deployment.](https://qdrant.tech/documentation/distributed_deployment/)
## Multitenancy: Data Isolation for Multi-Tenant Architectures
@@ -477,7 +477,7 @@ You can easily setup your access tokens and secure access to sensitive data thro
<img src="/articles_data/what-is-a-vector-database/jwt-web-ui.png" alt="Qdrant Web UI for generating a new access token." width="1000">
By default, Qdrant instances are **unsecured**, so it's important to configure security measures before moving to production. To learn more about how to configure security for your Qdrant instance and other advanced options, please check out the [official Qdrant documentation on security.](https://qdrant.tech/documentation/operations/security/)
By default, Qdrant instances are **unsecured**, so it's important to configure security measures before moving to production. To learn more about how to configure security for your Qdrant instance and other advanced options, please check out the [official Qdrant documentation on security.](https://qdrant.tech/documentation/security/)
## Time to Experiment
+2 -2
View File
@@ -57,8 +57,8 @@ In 2025, we focused on giving teams explicit control over retrieval quality as a
To support large, cost-sensitive workloads, we targeted the biggest performance bottlenecks in production systems. New improvements help teams scale indexing and querying without over-provisioning memory or compute.
**Related enhancements:**
• [GPU-Accelerated HNSW Indexing](https://qdrant.tech/documentation/operations/running-with-gpu/) unlocks up to an order-of-magnitude faster ingestion
• [Inline Storage](https://qdrant.tech/documentation/operations/optimize/#inline-storage-in-hnsw-index) embedded quantized vectors directly into the graph to dramatically improve disk-based search performance
• [GPU-Accelerated HNSW Indexing](https://qdrant.tech/documentation/ops-configuration/running-with-gpu/) unlocks up to an order-of-magnitude faster ingestion
• [Inline Storage](https://qdrant.tech/documentation/ops-optimization/optimize/#inline-storage-in-hnsw-index) embedded quantized vectors directly into the graph to dramatically improve disk-based search performance
• [Custom storage engine](https://qdrant.tech/articles/gridstore-key-value-storage/) optimized for predictable low-latency access
• [Incremental HNSW indexing](https://qdrant.tech/documentation/database-tutorials/bulk-upload/?q=incremental+hnsw#choose-an-indexing-strategy) for upsert-heavy workloads
• HNSW graph compression to reduce memory footprint
@@ -19,7 +19,7 @@ We’ve launched the **beta** of our Qdrant **Vector Data Migration Tool**, desi
This powerful tool streams all vectors from a source collection to a target Qdrant instance in live batches. It supports migrations from one Qdrant deployment to another, including from open source to Qdrant Cloud or between cloud regions. But that's not all. You can also migrate your data from other vector databases directly into Qdrant. All with a single command.
Unlike Qdrant’s included [snapshot migration method](https://qdrant.tech/documentation/operations/snapshots/), which requires consistent node-specific snapshots, our migration tool enables you to easily migrate data between different Qdrant database clusters in streaming batches. The only requirement is that the vector size and distance function must match.
Unlike Qdrant’s included [snapshot migration method](https://qdrant.tech/documentation/snapshots/), which requires consistent node-specific snapshots, our migration tool enables you to easily migrate data between different Qdrant database clusters in streaming batches. The only requirement is that the vector size and distance function must match.
This is especially useful if you want to change the collection configuration on the target, for example by choosing a different replication factor or quantization method.
@@ -57,7 +57,7 @@ Beyond coding, Anima uses Qdrant to understand documents at scale. By working wi
Several factors made Qdrant a strong fit for healthcare workloads.
[Deployment flexibility](https://qdrant.tech/documentation/operations/installation/) was non-negotiable. Anima requires data to remain at rest in the UK, allowing them to make strong guarantees to customers about compliance and residency. Qdrant’s self-hosted and region-controlled deployment options enabled this without compromising performance.
[Deployment flexibility](https://qdrant.tech/documentation/installation/) was non-negotiable. Anima requires data to remain at rest in the UK, allowing them to make strong guarantees to customers about compliance and residency. Qdrant’s self-hosted and region-controlled deployment options enabled this without compromising performance.
Cost predictability also played a critical role. With a fixed infrastructure cost for vector search, Anima could use retrieval across multiple passes in their pipelines. This unlocked higher-quality results without eroding margins.
@@ -48,7 +48,7 @@ For the engineering team, these failures had two serious implications. First, mi
After evaluating alternatives, Fieldy selected [Qdrant](http://qdrant.tech) for its stability, straightforward configuration, and suitability for self-hosted deployment. They opted to run Qdrant in the same environment as their backend services, ensuring low-latency access and avoiding the cross-region connectivity issues that had contributed to failures in the previous architecture.
The migration process took less than a week. The team reused their existing vector schema to avoid redesigning their indexing logic during the transition. Embeddings, generated using Cohere’s multilingual v3 model, were batched locally before ingestion, replacing the integrated embedding calls previously handled within Weaviate. Following [Qdrant best practices](https://qdrant.tech/documentation/operations/optimize/), they disabled indexing during bulk import to maximize throughput, re-enabling it only after the migration was complete. The backend API calls were updated to use Qdrant’s gRPC interface, and [hybrid search](https://qdrant.tech/articles/hybrid-search/) with BM25 was combined with dense vector retrieval via [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/search/hybrid-queries/#hybrid-search) for relevance scoring.
The migration process took less than a week. The team reused their existing vector schema to avoid redesigning their indexing logic during the transition. Embeddings, generated using Cohere’s multilingual v3 model, were batched locally before ingestion, replacing the integrated embedding calls previously handled within Weaviate. Following [Qdrant best practices](https://qdrant.tech/documentation/ops-optimization/optimize/), they disabled indexing during bulk import to maximize throughput, re-enabling it only after the migration was complete. The backend API calls were updated to use Qdrant’s gRPC interface, and [hybrid search](https://qdrant.tech/articles/hybrid-search/) with BM25 was combined with dense vector retrieval via [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/search/hybrid-queries/#hybrid-search) for relevance scoring.
### Architecture after migration
@@ -57,8 +57,8 @@ As part of their selection process, Nyris evaluated several critical factors to
Nyris has found several aspects of Qdrant particularly beneficial in their production environment:
- **Enhanced Security with JWT**: [JSON Web Tokens](https://qdrant.tech/documentation/operations/security/#granular-access-control-with-jwt) provide enhanced security and performance, critical for safeguarding their data.
- **Seamless Scalability**: Qdrant's ability to [scale effortlessly across nodes](https://qdrant.tech/documentation/operations/distributed_deployment/) ensures consistent high performance, even as Nyris's data volume grows.
- **Enhanced Security with JWT**: [JSON Web Tokens](https://qdrant.tech/documentation/security/#granular-access-control-with-jwt) provide enhanced security and performance, critical for safeguarding their data.
- **Seamless Scalability**: Qdrant's ability to [scale effortlessly across nodes](https://qdrant.tech/documentation/distributed_deployment/) ensures consistent high performance, even as Nyris's data volume grows.
- **Flexible Search Options**: The availability of both graph-based and brute-force search methods offers Nyris the flexibility to tailor the search approach to specific use case requirements.
- **Versatile Data Handling**: Qdrant imposes almost no restrictions on data types and vector sizes, allowing Nyris to manage diverse and complex datasets effectively.
- **Built with Rust**: The use of [Rust](https://qdrant.tech/articles/why-rust/) ensures superior performance and future-proofing, while its open-source nature allows Nyris to inspect and customize the code as necessary.
@@ -28,7 +28,7 @@ partition: case-studies
As part of this development, the Voiceflow engineering team was looking for a [vector database](/qdrant-vector-database/) solution to power their RAG setup. They evaluated various vector databases based on several key factors:
- **Performance**: The ability to [handle the scale](/documentation/operations/distributed_deployment/) required by Voiceflow, supporting hundreds of thousands of projects efficiently.
- **Performance**: The ability to [handle the scale](/documentation/distributed_deployment/) required by Voiceflow, supporting hundreds of thousands of projects efficiently.
- **Metadata**: The capability to tag data and chunks and retrieve based on those values, essential for organizing and accessing specific information swiftly.
- **Managed Solution**: The availability of a [managed service](/documentation/cloud/) with automated maintenance, scaling, and security, freeing the team from infrastructure concerns.
@@ -46,12 +46,12 @@ Qdrant is highly scalable and performant: it can handle billions of vectors effi
- **Advanced Similarity Search:** Qdrant supports various similarity [search](https://qdrant.tech/documentation/search/search/) metrics like dot product, cosine similarity, Euclidean distance, and Manhattan distance. You can store additional information along with vectors, known as [payload](https://qdrant.tech/documentation/manage-data/payload/) in Qdrant terminology. A payload is any JSON formatted data.
- **Built Using Rust:** Qdrant is built with Rust, and leverages its performance and efficiency. Rust is famed for its [memory safety](https://arxiv.org/abs/2206.05503) without the overhead of a garbage collector, and rivals C and C++ in speed.
- **Scaling and Multitenancy**: Qdrant supports both vertical and horizontal scaling and uses the Raft consensus protocol for [distributed deployments](https://qdrant.tech/documentation/operations/distributed_deployment/). Developers can run Qdrant clusters with replicas and shards, and seamlessly scale to handle large datasets. Qdrant also supports [multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/) where developers can create single collections and partition them using payload.
- **Scaling and Multitenancy**: Qdrant supports both vertical and horizontal scaling and uses the Raft consensus protocol for [distributed deployments](https://qdrant.tech/documentation/distributed_deployment/). Developers can run Qdrant clusters with replicas and shards, and seamlessly scale to handle large datasets. Qdrant also supports [multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/) where developers can create single collections and partition them using payload.
- **Payload Indexing and Filtering:** Just as Qdrant allows attaching any JSON payload to vectors, it also supports payload indexing and [filtering](https://qdrant.tech/documentation/search/filtering/) with a wide range of data types and query conditions, including keyword matching, full-text filtering, numerical ranges, nested object filters, and [geo](https://qdrant.tech/documentation/search/filtering/#geo)filtering.
- **Hybrid Search with Sparse Vectors:** Qdrant supports both dense and [sparse vectors](https://qdrant.tech/articles/sparse-vectors/), thereby enabling hybrid search capabilities. Sparse vectors are numerical representations of data where most of the elements are zero. Developers can combine search results from dense and sparse vectors, where sparse vectors ensure that results containing the specific keywords are returned and dense vectors identify semantically similar results.
- **Built-In Vector Quantization:** Qdrant offers three different [quantization](https://qdrant.tech/documentation/manage-data/quantization/) options to developers to optimize resource usage. Scalar quantization balances accuracy, speed, and compression by converting 32-bit floats to 8-bit integers. Binary quantization, the fastest method, significantly reduces memory usage. Product quantization offers the highest compression, and is perfect for memory-constrained scenarios.
- **Flexible Deployment Options:** Qdrant offers a range of deployment options. Developers can easily set up Qdrant (or Qdrant cluster) [locally](https://qdrant.tech/documentation/quickstart/#download-and-run) using Docker for free. [Qdrant Cloud](https://qdrant.tech/cloud/), on the other hand, is a scalable, managed solution that provides easy access with flexible pricing. Additionally, Qdrant offers [Hybrid Cloud](https://qdrant.tech/hybrid-cloud/) which integrates Kubernetes clusters from cloud, on-premises, or edge, into an enterprise-grade managed service.
- **Security through API Keys, JWT and RBAC:** Qdrant offers developers various ways to [secure](https://qdrant.tech/documentation/operations/security/) their instances. For simple authentication, developers can use API keys (including Read Only API keys). For more granular access control, it offers JSON Web Tokens (JWT) and the ability to build Role-Based Access Control (RBAC). TLS can be enabled to secure connections. Qdrant is also [SOC 2 Type II](https://qdrant.tech/blog/qdrant-soc2-type2-audit/) certified.
- **Security through API Keys, JWT and RBAC:** Qdrant offers developers various ways to [secure](https://qdrant.tech/documentation/security/) their instances. For simple authentication, developers can use API keys (including Read Only API keys). For more granular access control, it offers JSON Web Tokens (JWT) and the ability to build Role-Based Access Control (RBAC). TLS can be enabled to secure connections. Qdrant is also [SOC 2 Type II](https://qdrant.tech/blog/qdrant-soc2-type2-audit/) certified.
Additionally, Qdrant integrates seamlessly with popular machine learning frameworks such as [LangChain](https://qdrant.tech/blog/using-qdrant-and-langchain/), LlamaIndex, and Haystack; and Qdrant Hybrid Cloud integrates seamlessly with AWS, DigitalOcean, Google Cloud, Linode, Oracle Cloud, OpenShift, and Azure, among others.
@@ -45,13 +45,13 @@ guide](/documentation/cloud/authentication/#test-cluster-access).
If your Qdrant deployment is local, you do not need an API key.
Your next step depends on how you installed Qdrant. For details, read the
[Qdrant Installation](/documentation/operations/installation/)
[Qdrant Installation](/documentation/installation/)
guide.
#### If you use the Qdrant container or binary
Upgrade your deployment. Run the commands in the applicable section of the
[Qdrant Installation](/documentation/operations/installation/)
[Qdrant Installation](/documentation/installation/)
guide. The default commands automatically pull the latest version of Qdrant.
#### If you use the Qdrant helm chart
@@ -45,13 +45,13 @@ guide](https://qdrant.tech/documentation/cloud/quickstart-cloud/#step-2-test-clu
If your Qdrant deployment is local, you do not need an API key.
Your next step depends on how you installed Qdrant. For details, read the
[Qdrant Installation](https://qdrant.tech/documentation/operations/installation/)
[Qdrant Installation](https://qdrant.tech/documentation/installation/)
guide.
#### If you use the Qdrant container or binary
Upgrade your deployment. Run the commands in the applicable section of the
[Qdrant Installation](https://qdrant.tech/documentation/operations/installation/)
[Qdrant Installation](https://qdrant.tech/documentation/installation/)
guide. The default commands automatically pull the latest version of Qdrant.
#### If you use the Qdrant helm chart
@@ -85,7 +85,7 @@ Leveraging Late-Interaction Models for Rich Documents
Traditional OCR pipelines can add complexity and create accuracy challenges. But late-interaction models simplify the ingestion pipeline by running at the reranking stage.
Models like ([ColPali](https://qdrant.tech/blog/qdrant-colpali/) and ColQwen) bypass traditional OCR pipelines, directly processing images of complex documents. They enhance accuracy by maintaining original layouts and contextual integrity, simplifying your retrieval pipelines. The tradeoff is a heavier application, but these challenges can be addressed with further [optimization](https://qdrant.tech/documentation/operations/optimize/)*.*
Models like ([ColPali](https://qdrant.tech/blog/qdrant-colpali/) and ColQwen) bypass traditional OCR pipelines, directly processing images of complex documents. They enhance accuracy by maintaining original layouts and contextual integrity, simplifying your retrieval pipelines. The tradeoff is a heavier application, but these challenges can be addressed with further [optimization](https://qdrant.tech/documentation/ops-optimization/optimize/)*.*
#### Enabling highly granular accuracy for complex legal searches
+2 -2
View File
@@ -669,7 +669,7 @@ documentation, making it easier to navigate and find the information you need.
## S3 Snapshot Storage
Qdrant **Collections**, **Shards** and **Storage** can be backed up with [Snapshots](/documentation/operations/snapshots/) and saved in case of data loss or other data transfer purposes. These snapshots can be quite large and the resources required to maintain them can result in higher costs. AWS S3 and other S3-compatible implementations like [min.io](https://min.io/) is a great low-cost alternative that can hold snapshots without incurring high costs. It is globally reliable, scalable and resistant to data loss.
Qdrant **Collections**, **Shards** and **Storage** can be backed up with [Snapshots](/documentation/snapshots/) and saved in case of data loss or other data transfer purposes. These snapshots can be quite large and the resources required to maintain them can result in higher costs. AWS S3 and other S3-compatible implementations like [min.io](https://min.io/) is a great low-cost alternative that can hold snapshots without incurring high costs. It is globally reliable, scalable and resistant to data loss.
You can configure S3 storage settings in the [config.yaml](https://github.com/qdrant/qdrant/blob/master/config/config.yaml), specifically with `snapshots_storage`.
@@ -697,7 +697,7 @@ storage:
secret_key: your_secret_key_here
```
*Read more about [S3 snapshot storage](/documentation/operations/snapshots/#s3) and [configuration](/documentation/operations/configuration/).*
*Read more about [S3 snapshot storage](/documentation/snapshots/#s3) and [configuration](/documentation/ops-configuration/configuration/).*
This integration allows for a more convenient distribution of snapshots. Users of **any S3-compatible object storage** can now benefit from other platform services, such as automated workflows and disaster recovery options. S3's encryption and access control ensure secure storage and regulatory compliance. Additionally, S3 supports performance optimization through various storage classes and efficient data transfer methods, enabling quick and effective snapshot retrieval and management.
+1 -1
View File
@@ -255,7 +255,7 @@ PUT /collections/{collection_name}/index
![geo-index-disk](/blog/qdrant-1.12.x/geo-index-disk.png)
> To learn how to get the best performance from Qdrant, read the [**Optimization Guide**](/documentation/operations/optimize/).
> To learn how to get the best performance from Qdrant, read the [**Optimization Guide**](/documentation/ops-optimization/optimize/).
## Just the Beginning
+3 -3
View File
@@ -62,13 +62,13 @@ This experiment didn't require any changes to the codebase, and everything worke
- **Full Feature Support:** GPU indexing supports **all quantization options and datatypes** implemented in Qdrant.
- **Large-Scale Benefits:** Fast indexing unlocks larger size of segments, which leads to **higher RPS on the same hardware**.
### [Instructions & Documentation](/documentation/operations/running-with-gpu/)
### [Instructions & Documentation](/documentation/ops-configuration/running-with-gpu/)
The setup is simple, with pre-configured Docker images [**(check Docker Registry)**](https://hub.docker.com/r/qdrant/qdrant/tags) for GPU environments like NVIDIA and AMD.
We've made it so you can enable GPU indexing with minimal configuration changes.
> Note: Logs will clearly indicate GPU detection and usage for transparency.
*Read more about this feature in the [**GPU Indexing Documentation**](/documentation/operations/running-with-gpu/)*
*Read more about this feature in the [**GPU Indexing Documentation**](/documentation/ops-configuration/running-with-gpu/)*
#### Interview With the Creator of GPU Indexing
@@ -212,7 +212,7 @@ client.CreateCollection(context.Background(), &qdrant.CreateCollection{
```
> You may also use the `PATCH` request to enable Strict Mode on an existing collection.
*Read more about Strict Mode in the [**Database Administration Guide**](/documentation/operations/administration/#strict-mode)*
*Read more about Strict Mode in the [**Database Administration Guide**](/documentation/ops-configuration/administration/#strict-mode)*
## HNSW Graph Compression
+7 -7
View File
@@ -32,7 +32,7 @@ Additionally, version 1.16 introduces a new conditional update API, facilitating
Multitenancy is a common requirement for SaaS applications, where multiple customers (tenants) share the same database instance. In Qdrant, when an instance is shared between multiple users, you may need to partition vectors by user. This is done so that each user can only access their own vectors and can’t see the vectors of other users. To implement multitenancy in Qdrant, there are two main approaches:
- [Payload-based multitenancy](/documentation/manage-data/multitenancy/), which works well when you have a large number of small tenants. This causes practically no overhead. Quite the opposite: a query with a tenant payload filter can be faster than a full search.
- [Shard-based multitenancy](/documentation/operations/distributed_deployment/#user-defined-sharding), designed for when you have a smaller number of larger tenants. This works well when each tenant requires isolation and dedicated resources. Separating tenants by shard prevents a classic noisy neighbor problem where a single high-volume tenant can force the cluster to scale for everyone, increasing costs and reducing performance for smaller tenants. However, shard-based multitenancy is not a good solution when you have a large number of small tenants, as each shard incurs some overhead.
- [Shard-based multitenancy](/documentation/distributed_deployment/#user-defined-sharding), designed for when you have a smaller number of larger tenants. This works well when each tenant requires isolation and dedicated resources. Separating tenants by shard prevents a classic noisy neighbor problem where a single high-volume tenant can force the cluster to scale for everyone, increasing costs and reducing performance for smaller tenants. However, shard-based multitenancy is not a good solution when you have a large number of small tenants, as each shard incurs some overhead.
Real-world usage patterns often fall between these two use cases. It's common to have a small number of large tenants and a huge tail of smaller ones. You may even have tenants that grow over time, starting small and eventually becoming large enough to require dedicated resources.
@@ -40,7 +40,7 @@ In version 1.16, Qdrant can now efficiently combine the two multitenancy approac
The main principles behind Tiered Multitenancy are:
- [User-defined Sharding](/documentation/operations/distributed_deployment/#user-defined-sharding) allows you to create named shards within a collection. It enables you to isolate large tenants into their own shards. A multitenant collection can consist of a shared "fallback" shard for small tenants and multiple dedicated shards for large tenants.
- [User-defined Sharding](/documentation/distributed_deployment/#user-defined-sharding) allows you to create named shards within a collection. It enables you to isolate large tenants into their own shards. A multitenant collection can consist of a shared "fallback" shard for small tenants and multiple dedicated shards for large tenants.
- **Fallback shards** - a special routing mechanism that allows Qdrant to route a request to either a dedicated shard (if it exists) or to a shared fallback shard. This keeps requests unified, without the need to know whether a tenant is dedicated or shared.
- [Tenant promotion](/documentation/manage-data/multitenancy/#promote-tenant-to-dedicated-shard) - a mechanism that makes it possible to "promote" tenants from the shared Fallback Shard to their own dedicated shard when they grow large enough. This process is based on Qdrant’s internal shard transfer mechanism, which makes promotion completely transparent for the application. Both read and write requests are supported during the promotion process.
@@ -54,7 +54,7 @@ To use Tiered Multitenancy, after [setting up a collection with a shared fallbac
To enhance the scalability and speed of vector search, Qdrant employs a graph-based index structure known as [HNSW (Hierarchical Navigable Small World)](/documentation/manage-data/indexing/#vector-index). While traditional HNSW is primarily designed for unfiltered searches, Qdrant has addressed this limitation by implementing a [filterable HSNW index](/articles/filterable-hnsw/). This innovative approach extends the HNSW graph with additional edges that correspond to indexed payload values. This enables Qdrant to maintain search quality even with high filtering selectivity, without introducing any runtime overhead during the search process.
Even with filterable HNSW graphs, there are instances where the quality of search results can deteriorate significantly. This can happen when you use a combination of high cardinality filters, leading to the HNSW graph becoming [disconnected](/documentation/manage-data/indexing/#filterable-index). It is impractical to build additional links for every possible combination of filters in advance due to the potentially vast number of combinations. Another case where filterable HNSW may break down is in situations where filtering criteria are not known in advance.
Even with filterable HNSW graphs, there are instances where the quality of search results can deteriorate significantly. This can happen when you use a combination of high cardinality filters, leading to the HNSW graph becoming [disconnected](/documentation/manage-data/indexing/#filterable-hnsw-index). It is impractical to build additional links for every possible combination of filters in advance due to the potentially vast number of combinations. Another case where filterable HNSW may break down is in situations where filtering criteria are not known in advance.
To address these limitations, in version 1.16 we are introducing support for [ACORN](/documentation/search/search/#acorn-search-algorithm), based on the ACORN-1 algorithm described in the paper [ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data](https://arxiv.org/abs/2403.04871). With ACORN enabled, Qdrant not only traverses direct neighbors (the first hop) in the HNSW graph but also examines neighbors of neighbors (the second hop) if the direct neighbors have been filtered out. This enhancement improves search accuracy at the expense of performance, especially when multiple low-selectivity filters are applied.
@@ -126,9 +126,9 @@ For instance, querying 1 million vectors with the HNSW parameters `m` set to 16
However, disk-based storage has a property we can exploit to reduce the number of random access reads: paged reading. Disk devices typically read a full page (4KB or more) of data at once. Traditional tree-based data structures, such as B-trees, have used this property effectively. However, in graph-based structures like the HNSW index, grouping connected nodes into pages is not straightforward due to each node potentially having an arbitrary number of connections to other nodes.
With Qdrant version 1.16, you can make use of paged reading through a new feature called [inline storage](/documentation/operations/optimize/#inline-storage-in-hnsw-index). Inline storage allows for storing quantized vector data directly inside the HNSW nodes. This offers faster read access, at the cost of additional storage space.
With Qdrant version 1.16, you can make use of paged reading through a new feature called [inline storage](/documentation/ops-optimization/optimize/#inline-storage-in-hnsw-index). Inline storage allows for storing quantized vector data directly inside the HNSW nodes. This offers faster read access, at the cost of additional storage space.
Inline storage can be enabled by [setting a collection's HNSW configuration `inline_storage` option to `true`](/documentation/operations/optimize/#inline-storage-in-hnsw-index). It requires quantization to be enabled.
Inline storage can be enabled by [setting a collection's HNSW configuration `inline_storage` option to `true`](/documentation/ops-optimization/optimize/#inline-storage-in-hnsw-index). It requires quantization to be enabled.
<figure>
<img src="/blog/qdrant-1.16.x/no-inline-storage.png">
@@ -316,8 +316,8 @@ In version 1.16, we have revamped the Web UI with a fresh new look and improved
![Section 7](/blog/qdrant-1.16.x/section-7.png)
- The constant `k` that determines how Reciprocal Rank Fusion (RRF) fuses result sets [is now configurable](/documentation/search/hybrid-queries/#parametrized-rrf).
- The Metrics API now exposes [additional metrics](/documentation/operations/monitoring/#metrics) that help monitor your deployment's health.
- In strict mode, it is now possible to [configure the maximum number of payload indices](/documentation/operations/administration/#maximum-number-of-payload-index-count).
- The Metrics API now exposes [additional metrics](/documentation/ops-monitoring/monitoring/#metrics) that help monitor your deployment's health.
- In strict mode, it is now possible to [configure the maximum number of payload indices](/documentation/ops-configuration/administration/#maximum-number-of-payload-index-count).
- It's now possible to [attach custom metadata to collections](/documentation/manage-data/collections/#collection-metadata).
For a full list of all changes in version 1.16, please refer to the [change log](https://github.com/qdrant/qdrant/releases/tag/v1.16.0).
+4 -4
View File
@@ -78,11 +78,11 @@ We are continuously working to enhance the operational observability of Qdrant c
Qdrant’s API exposes a `/telemetry` endpoint which provides information about the current state of a peer in a cluster, including the number of vectors, shards, and other useful information. However, obtaining a complete view of the entire cluster using this endpoint is not straightforward, requiring querying each peer and piecing together a complete view yourself.
In version 1.17, we’re introducing a new [`/cluster/telemetry` endpoint](/documentation/operations/monitoring/#cluster-wide-telemetry). This API provides information about all peers in a cluster, offering insights into cluster-wide operations such as leader elections, resharding, and shard transfers.
In version 1.17, we’re introducing a new [`/cluster/telemetry` endpoint](/documentation/ops-monitoring/monitoring/#cluster-wide-telemetry). This API provides information about all peers in a cluster, offering insights into cluster-wide operations such as leader elections, resharding, and shard transfers.
### Segment Optimization Monitoring
Optimization is a background process where Qdrant removes data marked for deletion, merges segments, and creates indexes. To improve visibility into this process, this release introduces [segment optimization monitoring capabilities](/documentation/operations/optimizer/#optimization-monitoring).
Optimization is a background process where Qdrant removes data marked for deletion, merges segments, and creates indexes. To improve visibility into this process, this release introduces [segment optimization monitoring capabilities](/documentation/ops-optimization/optimizer/#optimization-monitoring).
A new `/collections/{collection_name}/optimizations` API endpoint provides cluster-wide information about the current optimization status, as well as detailed information for current and past optimization operations. Because the output of the API can be verbose, we’ve added a new Optimizations tab to the Collections interface in the Web UI that makes it easier to analyze the data. Here, you can find an overview of the current optimization status, a timeline of current and past optimization operations, and a breakdown of the tasks in a specific cycle and their durations.
@@ -115,7 +115,7 @@ Many people have been asking about point filtering in web UI. And now it's back,
As an open source project, we welcome contributions from the Qdrant community. This release features two contributions from community members:
- Not all payload field indexes are used in combination with dense vector queries. With this release, you can [specify whether individual payload field indexes should be reflected in the HNSW index](/documentation/manage-data/indexing/#disable-the-creation-of-extra-edges-for-payload-fields).
- A new API endpoint is available to [list all user-defined shard keys](/documentation/operations/distributed_deployment/#user-defined-sharding).
- A new API endpoint is available to [list all user-defined shard keys](/documentation/distributed_deployment/#user-defined-sharding).
Additionally, this release adds the following features:
@@ -123,7 +123,7 @@ Additionally, this release adds the following features:
- To speed up the recovery of the replicas after they’ve been down, shards will [increase the size of their write-ahead log](https://github.com/qdrant/qdrant/pull/7834) when they detect that one of their remote replicas is unavailable.
- Reciprocal Rank Fusion (RRF) combines multiple query results into one list, but its default equal weighting can let weaker rankers dilute stronger ones. [Weighted RRF](/documentation/search/hybrid-queries/#reciprocal-rank-fusion-rrf) in Qdrant 1.17 addresses this by letting you assign weights to individual queries.
- A new [user interface in the Web UI enables resharding collections](https://github.com/qdrant/qdrant-web-ui/pull/341) on Qdrant Cloud.
- Qdrant now supports [audit logging](/documentation/operations/security/#audit-logging) to track all API operations that require authentication or authorization.
- Qdrant now supports [audit logging](/documentation/security/#audit-logging) to track all API operations that require authentication or authorization.
- [External provider API keys for inference requests](/documentation/inference/#external-embedding-model-providers) can now be provided in the request header.
For a full list of all changes in version 1.17, please refer to the [change log](https://github.com/qdrant/qdrant/releases/tag/v1.17.0).
+4 -4
View File
@@ -28,7 +28,7 @@ tags:
Historically, our API key supported basic read and write operations. However, recognizing the evolving needs of our user base, especially large organizations, we've implemented additional options for finer control over data access within internal environments.
Qdrant now supports [granular access control using JSON Web Tokens (JWT)](/documentation/operations/security/#granular-access-control-with-jwt). JWT will let you easily limit a user's access to the specific data they are permitted to view. Specifically, JWT-based authentication leverages tokens with restricted access to designated data segments, laying the foundation for implementing role-based access control (RBAC) on top of it. **You will be able to define permissions for users and restrict access to sensitive endpoints.**
Qdrant now supports [granular access control using JSON Web Tokens (JWT)](/documentation/security/#granular-access-control-with-jwt). JWT will let you easily limit a user's access to the specific data they are permitted to view. Specifically, JWT-based authentication leverages tokens with restricted access to designated data segments, laying the foundation for implementing role-based access control (RBAC) on top of it. **You will be able to define permissions for users and restrict access to sensitive endpoints.**
**Dashboard users:** For your convenience, we have added a JWT generation tool the Qdrant Web UI under the 🔑 tab. If you're using the default url, you will find it at `http://localhost:6333/dashboard#/jwt`.
@@ -36,11 +36,11 @@ Qdrant now supports [granular access control using JSON Web Tokens (JWT)](/docum
We highly recommend this feature to enterprises using [Qdrant Hybrid Cloud](/hybrid-cloud/), as it is tailored to those who need additional control over company data and user access. RBAC empowers administrators to define roles and assign specific privileges to users based on their roles within the organization. In combination with [Hybrid Cloud's data sovereign architecture](/documentation/hybrid-cloud/), this feature reinforces internal security and efficient collaboration by granting access only to relevant resources.
> **Documentation:** [Read the access level breakdown](/documentation/operations/security/#table-of-access) to see which actions are allowed or denied.
> **Documentation:** [Read the access level breakdown](/documentation/security/#table-of-access) to see which actions are allowed or denied.
## Faster shard transfers on node recovery
We now offer a streamlined approach to [data synchronization between shards](/documentation/operations/distributed_deployment/#shard-transfer-method) during node upgrades or recovery processes. Traditional methods used to transfer the entire dataset, but our new `wal_delta` method focuses solely on transmitting the difference between two existing shards. By leveraging the Write-Ahead Log (WAL) of both shards, this method selectively transmits missed operations to the target shard, ensuring data consistency.
We now offer a streamlined approach to [data synchronization between shards](/documentation/distributed_deployment/#shard-transfer-method) during node upgrades or recovery processes. Traditional methods used to transfer the entire dataset, but our new `wal_delta` method focuses solely on transmitting the difference between two existing shards. By leveraging the Write-Ahead Log (WAL) of both shards, this method selectively transmits missed operations to the target shard, ensuring data consistency.
In some cases, where transfers can take hours, this update **reduces transfers down to a few minutes.**
@@ -48,7 +48,7 @@ The advantages of this approach are twofold:
1. **It is faster** since only the differential data is transmitted, avoiding the transfer of redundant information.
2. It upholds robust **ordering guarantees**, crucial for applications reliant on strict sequencing.
For more details on how this works, check out the [shard transfer documentation](/documentation/operations/distributed_deployment/#shard-transfer-method).
For more details on how this works, check out the [shard transfer documentation](/documentation/distributed_deployment/#shard-transfer-method).
> **Note:** There are limitations to consider. First, this method only works with existing shards. Second, while the WALs typically retain recent operations, their capacity is finite, potentially impeding the transfer process if exceeded. Nevertheless, for scenarios like rapid node restarts or upgrades, where the WAL content remains manageable, WAL delta transfer is an efficient solution.
@@ -69,4 +69,4 @@ As large companies continue to integrate sophisticated AI and machine learning t
Qdrant is open source and offers a complete SaaS solution, hosted on AWS, GCP, and Azure.
Getting started is easy, either spin up a [container image](https://hub.docker.com/r/qdrant/qdrant) or start a [free Cloud instance](https://cloud.qdrant.io/login). The documentation covers [adding the data](/documentation/tutorials-develop/bulk-upload/) to your Qdrant instance as well as [creating your indices](/documentation/operations/optimize/). We would love to hear about what you are building and please connect with our engineering team on [Github](https://github.com/qdrant/qdrant), [Discord](https://discord.com/invite/tdtYvXjC4h), or [LinkedIn](https://www.linkedin.com/company/qdrant).
Getting started is easy, either spin up a [container image](https://hub.docker.com/r/qdrant/qdrant) or start a [free Cloud instance](https://cloud.qdrant.io/login). The documentation covers [adding the data](/documentation/tutorials-develop/bulk-upload/) to your Qdrant instance as well as [creating your indices](/documentation/ops-optimization/optimize/). We would love to hear about what you are building and please connect with our engineering team on [Github](https://github.com/qdrant/qdrant), [Discord](https://discord.com/invite/tdtYvXjC4h), or [LinkedIn](https://www.linkedin.com/company/qdrant).
@@ -0,0 +1,75 @@
---
title: "Announcing Vector Space Day 2026 in San Francisco"
draft: false
slug: vector-space-day-sf-2026
short_description: "Join 300+ AI builders in San Francisco to talk about search, AI retrieval, agents, memory, edge & robotics AI, and more."
description: "From building scalable RAG pipelines to enabling real-time AI memory and next-gen context engineering, we’re covering the full spectrum of modern vector-native search."
preview_image: /blog/vector-space-day-2026-sf/sf-hero-20-april.jpg
social_preview_image: /blog/vector-space-day-2026-sf/sf-hero-20-april.jpg
date: 2026-04-21
author: Qdrant
featured: false
tags:
- news
- blog
---
## Vector Space Day 2026: Powered by Qdrant
### About
We’re hosting our second-ever full-day in-person Vector Space Day (https://luma.com/vsd-sf) on June 11th at The Midway in San Francisco, and you’re invited.
Last year in Berlin we brought together 400+ engineers, researchers, and AI builders to explore the cutting edge of retrieval, vector search infrastructure, and agentic AI. And we’re excited to now bring that to San Francisco.
![Berlin](/blog/vector-space-day-2026-sf/berlin1.jpeg)
### Why You Should Attend
* **Deep-dives, lightning talks, sessions you won’t want to miss**: Learn from the teams solving hard problems in AI infrastructure, search relevance, and semantic retrieval in production.
* **Meet the community**: This isn’t just a conference. It’s a gathering of the developers rethinking how AI systems search, retrieve, and reason at scale. Network or not, productive conversations will be had.
### Topics We’ll Explore
If you work on any of the following, the Vector Space Day 2026 is your event:
* Search & AI Retrieval
* Agents & Memory
* Edge & Robotics AI
* Happy Hour
### Happy Hour
The day won’t end when the last session wraps. Your ticket includes access to the happy hour where we will unwind and keep the conversations going over drinks, music, and light bites.
### Call for Speakers
We’ve locked in a strong lineup including Llamaindex, mem0, Neo4j, and more, but we’re saving a few select slots for standout talks from the community. If you’re building something novel in vector search, AI memory, context engineering, or retrieval infra, we want to hear from you.
[Submit your proposal.](https://docs.google.com/forms/d/e/1FAIpQLSfefAtmGP59-0IhNxCCMbNRCDlU-JRkJnSja3GfOrBnTZw-CA/viewform?usp=dialog)
Due May 6th.
### Get Your Ticket
General admission: $199
Early bird pricing: $99 through May 11
[**Reserve your spot now.**](https://luma.com/vsd-sf)
Space is limited.
### Global Hackathon
In the lead-up to Vector Space Day, we're hosting **Think Outside the Bot**, a global, virtual hackathon challenging devs to reimagine what's possible with vector search. Forget the classical RAG chatbot! Explore multi-modal applications, intelligent recommendations, and advanced vector search that go far beyond conversational interfaces.
[Submit your project.](https://try.qdrant.tech/hackathon-vsd)
### Need your manager’s approval to attend the event on June 11?
We’ve got you covered. Download this ready-to-send request letter to help explain why attending Vector Search Day is a valuable use of your time (and budget). [Download now](https://docs.google.com/document/d/1EivCVK47XEFXAhyoo8QaCBX0Op6uicUODAxTGXhZxrs/edit?usp=sharing).
@@ -145,7 +145,7 @@ The vector index in Qdrant employs the Hierarchical Navigable Small World (HNSW)
### Scalability
For massive datasets and demanding workloads, Qdrant supports [distributed deployment](/documentation/operations/distributed_deployment/) from v0.8.0. In this mode, you can set up a Qdrant cluster and distribute data across multiple nodes, enabling you to maintain high performance and availability even under increased workloads. Clusters support sharding and replication, and harness the Raft consensus algorithm to manage node coordination.
For massive datasets and demanding workloads, Qdrant supports [distributed deployment](/documentation/distributed_deployment/) from v0.8.0. In this mode, you can set up a Qdrant cluster and distribute data across multiple nodes, enabling you to maintain high performance and availability even under increased workloads. Clusters support sharding and replication, and harness the Raft consensus algorithm to manage node coordination.
Qdrant also supports vector [quantization](/documentation/manage-data/quantization/) to reduce memory footprint and speed up vector similarity searches, making it very effective for large-scale applications where efficient resource management is critical.
@@ -153,7 +153,7 @@ There are three quantization strategies you can choose from - scalar quantizatio
### Security
Qdrant offers several [security features](/documentation/operations/security/) to help protect data and access to the vector store:
Qdrant offers several [security features](/documentation/security/) to help protect data and access to the vector store:
- API Key Authentication: This helps secure API access to Qdrant Cloud with static or read-only API keys.
- JWT-Based Access Control: You can also enable more granular access control through JSON Web Tokens (JWT), and opt for restricted access to specific parts of the stored data while building Role-Based Access Control (RBAC).
@@ -85,7 +85,7 @@ you’ll get a detailed view with these tabs:
* **Search Quality Tab**: Evaluate and benchmark retrieval precision against ground truth. Tune parameters and measure the impact on accuracy.
* **Snapshots Tab**: Manage backups for this collection. Create a [snapshot](/documentation/operations/snapshots/), restore it later, or migrate it to another cluster.
* **Snapshots Tab**: Manage backups for this collection. Create a [snapshot](/documentation/snapshots/), restore it later, or migrate it to another cluster.
* **Visualize Tab**: Explore your vector space with an interactive 2D projection. See clusters, spot outliers, and build intuition about your embeddings.
@@ -526,6 +526,6 @@ print("=" * 60)
- [Qdrant Documentation](/documentation/) - Complete technical reference
- [HNSW Paper](https://arxiv.org/abs/1603.09320) - Original algorithm research
- [Qdrant Cloud](https://cloud.qdrant.io/) - Managed vector search service
- [Performance Tuning Guide](/documentation/operations/optimize/) - Advanced optimization techniques
- [Performance Tuning Guide](/documentation/ops-optimization/optimize/) - Advanced optimization techniques
**Ready for the pitstop project?** Now it's your turn to optimize performance with your own dataset and use case. You'll apply these same techniques to your domain-specific data and measure the real-world impact of different HNSW parameters and indexing strategies.
@@ -9,7 +9,7 @@ isLesson: true
# Combining Vector Search and Filtering
We've talked about how Qdrant uses the [HNSW](/documentation/manage-data/indexing/#filterable-index) graph to efficiently search dense vectors. But in real-world applications, you'll often want to constrain your search using filters. This creates unique challenges for graph traversal that Qdrant solves elegantly.
We've talked about how Qdrant uses the [HNSW](/documentation/manage-data/indexing/#filterable-hnsw-index) graph to efficiently search dense vectors. But in real-world applications, you'll often want to constrain your search using filters. This creates unique challenges for graph traversal that Qdrant solves elegantly.
<div class="video">
<iframe
@@ -119,7 +119,7 @@ accurate_search = SearchParams(hnsw_ef=256) # Higher recall, slower
### Memory & Indexing Behavior
Some vectors can remain unindexed depending on [optimizer](/documentation/operations/optimizer/) settings e.g. when the unindexed part stays below the `indexing_threshold` (kB).
Some vectors can remain unindexed depending on [optimizer](/documentation/ops-optimization/optimizer/) settings e.g. when the unindexed part stays below the `indexing_threshold` (kB).
Small collections or low-dimensional vectors may not trigger HNSW indexing at all. In such cases, full-scan search (brute force) is used instead until indexing becomes beneficial
@@ -80,7 +80,7 @@ Once running, you can access the Web UI at `http://localhost:6333/dashboard` to
### Alternative Installation Methods
For production deployments or other installation methods, see the [Qdrant Installation Guide](/documentation/operations/installation/).
For production deployments or other installation methods, see the [Qdrant Installation Guide](/documentation/installation/).
## Verifying Your Setup
+48 -31
View File
@@ -1,5 +1,5 @@
---
title: Home
title: Documentation
weight: 2
hideTOC: true
breadcrumb: false
@@ -16,56 +16,73 @@ content:
url: /documentation/quickstart/
contained: true
- partial: documentation/banners/banner-d
developingTitle: Ready to start developing?
developingDescription: Qdrant is open-source and can be self-hosted. However, the quickest way to get started is with our <a href="https://qdrant.to/cloud" target="_blank">free tier</a> on Qdrant Cloud. It scales easily and provides a UI where you can interact with data.
developingTitle: Introducing Qdrant Edge
developingDescription: Qdrant Edge is a lightweight, embedded vector search engine for in-process retrieval — no background services, minimal memory footprint, and no network required. Built for robots, kiosks, mobile devices, and any environment requiring offline-capable AI search.
developingBlock:
title: Create your first Qdrant Cloud cluster today
title: Run vector search anywhere, even offline
button:
text: Get Started
url: https://qdrant.to/cloud
url: /documentation/edge/edge-quickstart/
image:
src: /img/rocket.svg
alt: Rocket
- partial: documentation/sections/cards-section
title: Optimize Qdrant's performance
description: Boost search speed, reduce latency, and improve the accuracy and memory usage of your Qdrant deployment.
button:
text: Learn More
url: /documentation/operations/optimize/
title: Qdrant User Manual
description: Learn how to manage your data, run powerful searches, and leverage inference to build AI-native applications.
cardsPartial: documentation/cards/docs-cards
cards:
- id: 1
tag: Documents
icon:
src: /icons/outline/documentation-blue.svg
alt: Documents
title: Distributed Deployment
description: Scale Qdrant beyond a single node and optimize for high availability, fault tolerance, and billion-scale performance.
src: /icons/outline/vectors-blue.svg
alt: Vectors
title: Manage Data
description: Create collections, manage vectors, payloads, and storage. Learn about indexing, quantization, and multitenancy.
link:
url: /documentation/operations/distributed_deployment/
url: /documentation/manage-data/
text: Read More
- id: 2
tag: Documents
icon:
src: /icons/outline/documentation-blue.svg
alt: Documents
title: Multitenancy
description: Build vector search apps that serve millions of users. Learn about data isolation, security, and performance tuning.
src: /icons/outline/search-blue.svg
alt: Search
title: Search
description: Learn about similarity search, filtering, hybrid queries, and advanced retrieval techniques.
link:
url: /documentation/manage-data/multitenancy/
url: /documentation/search/
text: Read More
- id: 3
tag: Blog
tagColor: violet
icon:
src: /icons/outline/blog-purple.svg
alt: Blog
title: Vector Quantization
description: Learn about cutting-edge techniques for vector quantization and how they can be used to improve search performance.
src: /icons/outline/integration-blue.svg
alt: Inference
title: Inference
description: Configure dense, sparse, and multi-vector embeddings. Use cloud-hosted embedding models directly with Qdrant.
link:
url: /articles/what-is-vector-quantization/
url: /documentation/inference/
text: Read More
partition: qdrant
- partial: documentation/sections/cards-section
title: Support
description: Get help from the Qdrant community or contact our support team.
cardsPartial: documentation/cards/docs-cards
cardsPerRow: 2
cards:
- id: 1
icon:
src: /icons/outline/discord-purple.svg
alt: Discord icon
title: Community Support
description: Join 6,000+ active members to learn, collaborate, and participate in Qdrant's latest activities.
link:
text: Join our Discord
url: https://qdrant.to/discord
- id: 2
icon:
src: /icons/outline/support-blue.svg
alt: Support icon
title: Qdrant Cloud Support
description: Paying customers have access to our Support team. Links to the support portal are available in the Qdrant Cloud Console.
link:
text: Join Qdrant
url: https://qdrant.to/cloud
partition: develop
---
THIS CONTENT IS GOING TO BE IGNORED FOR NOW
@@ -89,7 +106,7 @@ Qdrant is an AI-native vector search and a semantic search engine. You can use i
||||
|:-|:-|:-|
|[Filterable HNSW](/documentation/search/filtering/) </br> Single-stage payload filtering | [Recommendations & Context Search](/documentation/search/explore/#explore-the-data) </br> Exploratory advanced search| [Pure-Vector Hybrid Search](/documentation/search/hybrid-queries/)</br>Full text and semantic search in one|
|[Multitenancy](/documentation/manage-data/multitenancy/) </br> Payload-based partitioning|[Custom Sharding](/documentation/operations/distributed_deployment/#sharding) </br> For data isolation and distribution|[Role Based Access Control](/documentation/operations/security/?q=jwt#granular-access-control-with-jwt)</br>Secure JWT-based access |
|[Multitenancy](/documentation/manage-data/multitenancy/) </br> Payload-based partitioning|[Custom Sharding](/documentation/distributed_deployment/#sharding) </br> For data isolation and distribution|[Role Based Access Control](/documentation/security/?q=jwt#granular-access-control-with-jwt)</br>Secure JWT-based access |
|[Quantization](/documentation/manage-data/quantization/) </br> Compress data for drastic speedups|[Multivector Support](/documentation/manage-data/vectors/?q=multivect#multivectors) </br> For ColBERT late interaction |[Built-in IDF](/documentation/manage-data/indexing/?q=inverse+docu#idf-modifier) </br> Advanced similarity calculation|
## Developer guidebooks:
@@ -1,9 +1,12 @@
---
title: Capacity Planning
weight: 5
partition: deploy
weight: 238
aliases:
- capacity
- /documentation/cloud/capacity-sizing
- /documentation/capacity-planning
- /documentation/operations/capacity-planning
---
# Capacity Planning
@@ -1,7 +1,7 @@
---
title: Account Setup
weight: 13
partition: cloud
weight: 210
partition: deploy
aliases:
- /documentation/cloud/qdrant-cloud-setup/
---
@@ -1,7 +1,7 @@
---
title: Qdrant Cloud API
weight: 27
partition: cloud
weight: 245
partition: deploy
aliases:
- /documentation/qdrant-cloud-api/
---
@@ -1,7 +1,7 @@
---
title: Qdrant Cloud CLI
weight: 28
partition: cloud
weight: 250
partition: deploy
---
# Qdrant Cloud CLI
@@ -1,7 +1,7 @@
---
title: Getting Started
weight: 12
partition: cloud
weight: 205
partition: deploy
aliases:
- /documentation/cloud/getting-started/
---
@@ -20,7 +20,7 @@ Premium Plan subscribers can enable single sign-on (SSO) for their organizations
## Cluster Sizing
Before deploying any cluster, consider the resources needed for your specific workload. Our [Capacity Planning guide](/documentation/operations/capacity-planning/) describes how to assess the required CPU, memory, and storage. Additionally, the [Pricing Calculator](https://cloud.qdrant.io/calculator) helps you estimate associated costs based on your projected usage.
Before deploying any cluster, consider the resources needed for your specific workload. Our [Capacity Planning guide](/documentation/capacity-planning/) describes how to assess the required CPU, memory, and storage. Additionally, the [Pricing Calculator](https://cloud.qdrant.io/calculator) helps you estimate associated costs based on your projected usage.
## Creating and Managing Clusters
@@ -28,7 +28,7 @@ After setting up your account, you can create a Qdrant Cluster by following the
## Preparing for Production
For a production-ready environment, consider deploying a multi-node Qdrant cluster (at least three nodes) with replication enabled. More details are available in the [Distributed Deployment](/documentation/operations/distributed_deployment/) guide. For more information on how to create a production-ready cluster, see our [Vector Search in Production](/articles/vector-search-production/) article.
For a production-ready environment, consider deploying a multi-node Qdrant cluster (at least three nodes) with replication enabled. More details are available in the [Distributed Deployment](/documentation/distributed_deployment/) guide. For more information on how to create a production-ready cluster, see our [Vector Search in Production](/articles/vector-search-production/) article.
If you are looking to optimize costs, you can reduce memory usage through [Quantization](/documentation/manage-data/quantization/) or by [offloading vectors to disk](/documentation/manage-data/storage/#configuring-memmap-storage).
@@ -1,7 +1,7 @@
---
title: Premium Tier
weight: 19
partition: cloud
weight: 260
partition: deploy
aliases:
- /documentation/cloud/premium/
---
@@ -1,7 +1,7 @@
---
title: Billing & Payments
weight: 18
partition: cloud
weight: 255
partition: deploy
aliases:
- aws-marketplace
- gcp-marketplace
@@ -1,7 +1,7 @@
---
title: Cloud Quickstart
weight: 4
partition: cloud
weight: 115
partition: develop
aliases:
- ../cloud-quick-start
- cloud-quick-start
@@ -22,7 +22,7 @@ Learn how to set up Qdrant Cloud and perform your first semantic search in just
2. Under **Create a Free Cluster**, enter a cluster name and select your preferred cloud provider and region. Click **Create Free Cluster**.
3. Copy your **API key** when prompted - you'll need it to connect. Store it somewhere safe as it won't be displayed again.
For detailed cluster setup instructions, see the [Cloud documentation](/documentation/cloud-intro/).
For detailed cluster setup instructions, see the [Cloud documentation](/documentation/deploy-intro/).
## 2. Install the Qdrant Client
@@ -1,7 +1,7 @@
---
title: Cloud RBAC
weight: 14
partition: cloud
weight: 215
partition: deploy
---
# Cloud RBAC
@@ -1,6 +1,6 @@
---
title: Permission Reference
weight: 3
weight: 15
---
# **Permission Reference**
@@ -57,8 +57,8 @@ Permissions for API Keys, backups, clusters, and backup schedules.
### **Cluster Data**
| Permission | Description |
|------------|------------|
| `read:cluster_data` | View cluster data, used for the Cluster UI button on Cluster Details. [Maps to global `read-only` JWT access for the cluster.](/documentation/operations/security/) |
| `write:cluster_data` | View and modify cluster data, used for the Cluster UI button on Cluster Details. [Maps to global `read-write` JWT access for the cluster.](/documentation/operations/security/) |
| `read:cluster_data` | View cluster data, used for the Cluster UI button on Cluster Details. [Maps to global `read-only` JWT access for the cluster.](/documentation/security/) |
| `write:cluster_data` | View and modify cluster data, used for the Cluster UI button on Cluster Details. [Maps to global `read-write` JWT access for the cluster.](/documentation/security/) |
### **Backup Schedules**
| Permission | Description |
@@ -1,6 +1,6 @@
---
title: Role Management
weight: 1
weight: 5
---
# Role Management
@@ -1,6 +1,6 @@
---
title: User Management
weight: 2
weight: 10
---
# User Management
@@ -1,7 +1,7 @@
---
title: Security
weight: 36
partition: cloud
title: Cloud Security
weight: 240
partition: deploy
aliases:
- /documentation/cloud/security/
---
@@ -4,8 +4,8 @@ slug: cloud-intro
breadcrumb: false
content:
- partial: documentation/banners/banner-b
title: Welcome to Qdrant Cloud
description: Dev-portal Cloud
title: Deploy & Operate Qdrant
description: Deploy & Operate Qdrant
image:
src: /img/dev-portal-cloud/dev-portal-cloud-hero.png
alt: Qdrant cloud dashboard
@@ -13,7 +13,39 @@ content:
text: Get Started
url: https://qdrant.to/cloud
- partial: documentation/sections/cards-section
title: Managed Services
title: Operations
description: Install, configure, optimize, and monitor your Qdrant deployment across any environment.
cardsPartial: documentation/cards/docs-cards
cards:
- id: 1
icon:
src: /icons/outline/server-rack-blue.svg
alt: Installation
title: Installation
description: Deploy Qdrant on any infrastructure. Get requirements, configuration options, and GPU setup guides.
link:
url: /documentation/installation/
text: Read More
- id: 2
icon:
src: /icons/outline/switches-blue.svg
alt: Configuration
title: Configuration
description: Tune storage, network, performance, and runtime settings for your Qdrant instance.
link:
url: /documentation/ops-configuration/configuration/
text: Read More
- id: 3
icon:
src: /icons/outline/chart-bar-blue.svg
alt: Monitoring
title: Monitoring & Telemetry
description: Monitor cluster health, collect metrics with Prometheus and Grafana, and configure telemetry.
link:
url: /documentation/ops-monitoring/monitoring/
text: Read More
- partial: documentation/sections/cards-section
title: Cloud
description: Deploy and manage high-performance vector search clusters across cloud environments. Easily scale with fully managed cloud solutions, integrate seamlessly across hybrid setups, or maintain complete control with private cloud deployments in Kubernetes.
cardsPartial: documentation/cards/docs-cards
cards:
@@ -45,8 +77,8 @@ content:
url: /documentation/private-cloud/
text: Read More
- partial: documentation/sections/cards-section
title: Customer Support
description: Stream, index, and migrate data to Qdrant with these essential tools and strategies.
title: Support
description: Get help from the Qdrant community or contact our support team.
cardsPartial: documentation/cards/docs-cards
cardsPerRow: 2
cards:
@@ -68,6 +100,6 @@ content:
link:
text: Join Qdrant
url: https://qdrant.to/cloud
partition: cloud
partition: deploy
hideInSidebar: true
---
@@ -1,7 +1,7 @@
---
title: Infrastructure Tools
weight: 28
partition: cloud
weight: 235
partition: deploy
---
## Cloud Tools
@@ -1,5 +1,6 @@
---
title: Pulumi
weight: 5
aliases:
- /documentation/infrastructure/pulumi/
---
@@ -1,5 +1,6 @@
---
title: Terraform
weight: 10
aliases:
- /documentation/infrastructure/terraform/
---
@@ -1,9 +1,9 @@
---
title: Managed Cloud
weight: 15
weight: 220
aliases:
- /documentation/overview/qdrant-alternatives/documentation/cloud/
partition: cloud
partition: deploy
---
# About Qdrant Managed Cloud
@@ -1,13 +1,13 @@
---
title: Authentication
weight: 30
weight: 10
---
# Database Authentication in Qdrant Managed Cloud
This page describes what Database API keys are and shows you how to use the Qdrant Cloud Console to create a Database API key for a cluster. You will learn how to connect to your cluster using the new API key.
Database API keys can be configured with granular access control. Database API keys with granular access control can be recognized by starting with `eyJhb`. Please refer to the [Table of access](/documentation/operations/security/#table-of-access) to understand what permissions you can configure.
Database API keys can be configured with granular access control. Database API keys with granular access control can be recognized by starting with `eyJhb`. Please refer to the [Table of access](/documentation/security/#table-of-access) to understand what permissions you can configure.
Database API keys with granular access control are available for clusters using version **v1.11.0** and above.
@@ -1,6 +1,6 @@
---
title: Backup Clusters
weight: 61
weight: 40
---
# Backing up Qdrant Cloud Clusters
@@ -71,7 +71,7 @@ Or you can restore the backup into a new cluster.
Qdrant also offers a snapshot API which allows you to create a snapshot
of a specific collection or your entire cluster. For more information, see our
[snapshot documentation](/documentation/operations/snapshots/).
[snapshot documentation](/documentation/snapshots/).
Here is how you can take a snapshot and recover a collection:
@@ -79,11 +79,11 @@ Here is how you can take a snapshot and recover a collection:
- For a single node cluster, call the snapshot endpoint on the exposed URL.
- For a multi node cluster call a snapshot on each node of the collection.
Specifically, prepend `node-{num}-` to your cluster URL.
Then call the [snapshot endpoint](/documentation/operations/snapshots/#create-snapshot) on the individual hosts. Start with node 0.
Then call the [snapshot endpoint](/documentation/snapshots/#create-snapshot) on the individual hosts. Start with node 0.
- In the response, you'll see the name of the snapshot.
2. Delete and recreate the collection.
3. Recover the snapshot:
- Call the [recover endpoint](/documentation/operations/snapshots/#recover-in-cluster-deployment). Set a location which points to the snapshot file (`file:///qdrant/snapshots/{collection_name}/{snapshot_file_name}`) for each host.
- Call the [recover endpoint](/documentation/snapshots/#recover-in-cluster-deployment). Set a location which points to the snapshot file (`file:///qdrant/snapshots/{collection_name}/{snapshot_file_name}`) for each host.
## Backup Considerations
@@ -1,6 +1,6 @@
---
title: Cluster Access
weight: 35
weight: 15
---
# Accessing Qdrant Cloud Clusters
@@ -1,6 +1,6 @@
---
title: Monitor Clusters
weight: 55
weight: 30
---
# Monitoring Qdrant Cloud Clusters
@@ -120,7 +120,7 @@ The account owner will receive automatic alerts via email if your cluster has an
**Where can I learn more about this alert?**
You can learn more about disk capacity [here](/documentation/operations/capacity-planning/#scaling-disk-space-in-qdrant-cloud).
You can learn more about disk capacity [here](/documentation/capacity-planning/#scaling-disk-space-in-qdrant-cloud).
You can learn about vertical scaling [here](/documentation/cloud/cluster-scaling/#vertical-scaling).
@@ -202,7 +202,7 @@ The account owner will receive automatic alerts via email if your cluster has an
Learn about the SDKs [here](/documentation/interfaces/).
Learn more about JWT Keys and permissions [here](/documentation/operations/security/?q=jwt#granular-access-control-with-jwt).
Learn more about JWT Keys and permissions [here](/documentation/security/?q=jwt#granular-access-control-with-jwt).
- title: A Node is CPU Throttled
content: |
@@ -230,9 +230,9 @@ The account owner will receive automatic alerts via email if your cluster has an
**Where can I learn more about this alert?**
Learn how to optimize Qdrant for performance and configure indexing [here](/documentation/manage-data/indexing/) and [here](/documentation/operations/optimize/).
Learn how to optimize Qdrant for performance and configure indexing [here](/documentation/manage-data/indexing/) and [here](/documentation/ops-optimization/optimize/).
Learn about optimizers [here](/documentation/operations/optimizer/).
Learn about optimizers [here](/documentation/ops-optimization/optimizer/).
Learn more about hybrid search [here](/documentation/search/hybrid-queries/).
@@ -280,7 +280,7 @@ The account owner will receive automatic alerts via email if your cluster has an
**Where can I learn more about this alert?**
Learn more about distributed deployments and resharding [here](/documentation/operations/distributed_deployment/#resharding).
Learn more about distributed deployments and resharding [here](/documentation/distributed_deployment/#resharding).
Learn more about cloud rebalancing [here](/documentation/cloud/configure-cluster/#shard-rebalancing).
@@ -294,11 +294,11 @@ To scrape metrics from a Qdrant cluster running in Qdrant Cloud, an [API key](/d
### Qdrant Node Metrics
Metrics in a Prometheus-compatible format are available at the `/metrics` endpoint of each Qdrant database node. When scraping, you should use the [node specific URLs](/documentation/cloud/cluster-access/#node-specific-endpoints) to ensure that you are scraping metrics from all nodes in each cluster. For more information, see [Qdrant monitoring](/documentation/operations/monitoring/).
Metrics in a Prometheus-compatible format are available at the `/metrics` endpoint of each Qdrant database node. When scraping, you should use the [node specific URLs](/documentation/cloud/cluster-access/#node-specific-endpoints) to ensure that you are scraping metrics from all nodes in each cluster. For more information, see [Qdrant monitoring](/documentation/ops-monitoring/monitoring/).
You can also access the `/telemetry` [endpoint](https://api.qdrant.tech/api-reference/service/telemetry) of your database. This endpoint is available on the cluster endpoint and provides information about the current state of the database, including the number of vectors, shards, and other useful information.
For more information, see [Qdrant monitoring](/documentation/operations/monitoring/).
For more information, see [Qdrant monitoring](/documentation/ops-monitoring/monitoring/).
### Cluster System Metrics
@@ -1,6 +1,6 @@
---
title: Scale Clusters
weight: 50
weight: 20
---
# Scaling Qdrant Cloud Clusters
@@ -27,7 +27,7 @@ Vertical scaling can be an effective way to improve the performance of a cluster
In such cases, horizontal scaling may be a more effective solution.
Horizontal scaling is the process of increasing the capacity of a cluster by adding more nodes and distributing the load and data among them. The horizontal scaling at Qdrant starts on the collection level. You have to choose the number of shards you want to distribute your collection around while creating the collection. Please refer to the [sharding documentation](/documentation/operations/distributed_deployment/#sharding) section for details.
Horizontal scaling is the process of increasing the capacity of a cluster by adding more nodes and distributing the load and data among them. The horizontal scaling at Qdrant starts on the collection level. You have to choose the number of shards you want to distribute your collection around while creating the collection. Please refer to the [sharding documentation](/documentation/distributed_deployment/#sharding) section for details.
When scaling up horizontally, the cloud platform will automatically rebalance all available shards across nodes to ensure that the data is evenly distributed. See [Configuring Clusters](/documentation/cloud/configure-cluster/#shard-rebalancing) for more details.
@@ -43,7 +43,7 @@ We will be glad to consult you on an optimal strategy for scaling.
*Available as of Qdrant v1.13.0*
<aside role="status">Resharding is exclusively available on multi-node clusters across our <a href="/documentation/cloud-intro/">cloud</a> offering, including <a href="/documentation/hybrid-cloud/">Hybrid</a> and <a href="/documentation/private-cloud/">Private</a> Cloud.</aside>
<aside role="status">Resharding is exclusively available on multi-node clusters across our <a href="/documentation/deploy-intro/">cloud</a> offering, including <a href="/documentation/hybrid-cloud/">Hybrid</a> and <a href="/documentation/private-cloud/">Private</a> Cloud.</aside>
When creating a collection, it has a specific number of shards. The ideal number of shards might change as your cluster evolves.
@@ -1,6 +1,6 @@
---
title: Update Clusters
weight: 55
weight: 35
---
# Updating Qdrant Cloud Clusters
@@ -1,18 +1,18 @@
---
title: Configure Clusters
weight: 55
weight: 25
---
# Configure Qdrant Cloud Clusters
Qdrant Cloud offers several advanced configuration options to optimize clusters for your specific needs. You can access these options from the Cluster Details page in the Qdrant Cloud console.
The cloud platform does not expose all [configuration options](/documentation/operations/configuration/) available in Qdrant. We have selected the relevant options that are explained in detail below.
The cloud platform does not expose all [configuration options](/documentation/ops-configuration/configuration/) available in Qdrant. We have selected the relevant options that are explained in detail below.
In addition, the cloud platform automatically configures the following settings for your cluster to ensure optimal performance and reliability:
* The maximum number of collections in a cluster is set to 1000. Larger numbers of collections lead to performance degradation. For more information see [Multitenancy](/documentation/manage-data/multitenancy/).
* Strict mode is activated by default for new collections enforcing that all filters being used in retrieve and update queries are indexed. This improves performance and reliability. You can disable this individually for each collection. For more information see [Strict Mode](/documentation/operations/administration/#strict-mode).
* Strict mode is activated by default for new collections enforcing that all filters being used in retrieve and update queries are indexed. This improves performance and reliability. You can disable this individually for each collection. For more information see [Strict Mode](/documentation/ops-configuration/administration/#strict-mode).
* The cluster mode is automatically enabled to allow distributed deployments and horizontal scaling.
* The maximum amount of payload indexes per collection is set to 100. Larger numbers of payload indexes lead to performance degradation (starting with Qdrant v1.16.0).
@@ -24,7 +24,7 @@ You can set default values for the configuration of new collections in your clus
You can configure the default *Replication Factor*, the default *Write Consistency Factor*, and if vectors should be stored on disk only, instead of being cached in RAM.
Refer to [Qdrant Configuration](/documentation/operations/configuration/#configuration-options) for more details.
Refer to [Qdrant Configuration](/documentation/ops-configuration/configuration/#configuration-options) for more details.
## Advanced Optimizations
@@ -1,6 +1,6 @@
---
title: Create a Cluster
weight: 20
weight: 5
---
# Creating a Qdrant Cloud Cluster
@@ -20,7 +20,7 @@ A free tier cluster only includes 1 single node with the following resources:
| Disk space | 4 GB |
| Nodes | 1 |
This configuration supports serving about 1 M vectors of 768 dimensions. To calculate your needs, refer to our documentation on [Capacity Planning](/documentation/operations/capacity-planning/).
This configuration supports serving about 1 M vectors of 768 dimensions. To calculate your needs, refer to our documentation on [Capacity Planning](/documentation/capacity-planning/).
The choice of cloud providers and regions is limited.
@@ -77,7 +77,7 @@ This page shows you how to use the Qdrant Cloud Console to create a custom Qdran
1. Choose your data center region or Hybrid Cloud environment.
1. Configure RAM for each node.
> For more information, see our [Capacity Planning](/documentation/operations/capacity-planning/) guidance.
> For more information, see our [Capacity Planning](/documentation/capacity-planning/) guidance.
1. Choose the number of vCPUs and GPUs per node. If you add more
RAM, the menu provides different options for vCPUs. For higher RAM configurations, you can also choose to add a GPU to optimize indexing performance (AWS only).
1. Select the number of nodes you want the cluster to be deployed on.
@@ -116,7 +116,7 @@ We recommend the **Balanced** tier for disks >= 32 GiB, and the **Performance**
**GPUs (AWS only)**
If you have a write-heavy workload, you can add a GPU to each node to optimize indexing performance. See [**GPUs for Indexing**](/documentation/operations/running-with-gpu/) for more information. All GPU settings will be configured automatically by the cloud platform.
If you have a write-heavy workload, you can add a GPU to each node to optimize indexing performance. See [**GPUs for Indexing**](/documentation/ops-configuration/running-with-gpu/) for more information. All GPU settings will be configured automatically by the cloud platform.
**Backup and Disaster Recovery**
@@ -124,7 +124,7 @@ You should create a backup schedule for your cluster. This ensures that you can
**Collection Sharding**
To allow your cluster to easily scale horizontally, you should configure at least twice as many shards per collection than the number of nodes in your cluster. You can configure the number of shards when creating a collection. See [**Sharding**](/documentation/operations/distributed_deployment/#sharding) for more information.
To allow your cluster to easily scale horizontally, you should configure at least twice as many shards per collection than the number of nodes in your cluster. You can configure the number of shards when creating a collection. See [**Sharding**](/documentation/distributed_deployment/#sharding) for more information.
If you did not configure enough shards in a collection, you can use the [**Resharding**](/documentation/cloud/cluster-scaling/#resharding) feature to change the number of shards in an existing collection.
@@ -1,6 +1,6 @@
---
title: Inference
weight: 81
weight: 45
---
# Inference in Qdrant Managed Cloud
@@ -1,10 +1,13 @@
---
title: Troubleshooting
weight: 45
partition: deploy
weight: 150
aliases:
- ../tutorials/common-errors
- /documentation/troubleshooting/
- /documentation/guides/common-errors/
- /documentation/common-errors
- /documentation/operations/common-errors
---
# Solving common errors
@@ -35,7 +38,7 @@ Please note, the command should be executed before you run Qdrant server.
## Incompatible file system
Qdrant have a [set of requirements](/documentation/operations/installation/#storage) for persistent file storage.
Qdrant have a [set of requirements](/documentation/installation/#storage) for persistent file storage.
The most important requirement is that file system **must** be [POSIX-compatible](https://www.quobyte.com/storage-explained/posix-filesystem/).
@@ -26,7 +26,7 @@ Before you start, make sure you have the following:
1. Airbyte instance, either [Open Source](https://airbyte.com/solutions/airbyte-open-source),
[Self-Managed](https://airbyte.com/solutions/airbyte-enterprise), or [Cloud](https://airbyte.com/solutions/airbyte-cloud).
2. Running instance of Qdrant. It has to be accessible by URL from the machine where Airbyte is running.
You can follow the [installation guide](/documentation/operations/installation/) to set up Qdrant.
You can follow the [installation guide](/documentation/installation/) to set up Qdrant.
## Setting up Qdrant as a destination
@@ -13,7 +13,7 @@ Qdrant is available as a [provider](https://airflow.apache.org/docs/apache-airfl
Before configuring Airflow, you need:
1. A Qdrant instance to connect to. You can set one up in our [installation guide](/documentation/operations/installation/).
1. A Qdrant instance to connect to. You can set one up in our [installation guide](/documentation/installation/).
2. A running Airflow instance. You can use their [Quick Start Guide](https://airflow.apache.org/docs/apache-airflow/stable/start.html).
@@ -22,7 +22,7 @@ on a dataset name to see its detailed description.
| [Arxiv.org abstracts](#arxivorg-abstracts) | [InstructorXL](https://huggingface.co/hkunlp/instructor-xl) | 768 | 2.3M | 8.4 GB | [Download](https://snapshots.qdrant.io/arxiv_abstracts-3083016565637815127-2023-06-02-07-26-29.snapshot) | [Open](https://huggingface.co/datasets/Qdrant/arxiv-abstracts-instructorxl-embeddings) |
| [Wolt food](#wolt-food) | [clip-ViT-B-32](https://huggingface.co/sentence-transformers/clip-ViT-B-32) | 512 | 1.7M | 7.9 GB | [Download](https://snapshots.qdrant.io/wolt-clip-ViT-B-32-2446808438011867-2023-12-14-15-55-26.snapshot) | [Open](https://huggingface.co/datasets/Qdrant/wolt-food-clip-ViT-B-32-embeddings) |
Once you download a snapshot, you need to [restore it](/documentation/operations/snapshots/#restore-snapshot)
Once you download a snapshot, you need to [restore it](/documentation/snapshots/#restore-snapshot)
using the Qdrant CLI upon startup or through the API.
## Qdrant on Hugging Face
@@ -0,0 +1,108 @@
---
title: Deploy Qdrant
slug: deploy-intro
breadcrumb: false
aliases:
- /documentation/deploy-intro
- /documentation/cloud-intro
content:
- partial: documentation/banners/banner-b
title: Deploy Qdrant
description: Deploy Qdrant
image:
src: /img/dev-portal-cloud/dev-portal-cloud-hero.png
alt: Qdrant cloud dashboard
startedButton:
text: Get Started
url: https://qdrant.to/cloud
- partial: documentation/sections/cards-section
title: Operations
description: Install, configure, optimize, and monitor your Qdrant deployment across any environment.
cardsPartial: documentation/cards/docs-cards
cards:
- id: 1
icon:
src: /icons/outline/server-rack-blue.svg
alt: Installation
title: Installation
description: Deploy Qdrant on any infrastructure. Get requirements, configuration options, and GPU setup guides.
link:
url: /documentation/installation/
text: Read More
- id: 2
icon:
src: /icons/outline/switches-blue.svg
alt: Configuration
title: Configuration
description: Tune storage, network, performance, and runtime settings for your Qdrant instance.
link:
url: /documentation/ops-configuration/configuration/
text: Read More
- id: 3
icon:
src: /icons/outline/chart-bar-blue.svg
alt: Monitoring
title: Monitoring & Telemetry
description: Monitor cluster health, collect metrics with Prometheus and Grafana, and configure telemetry.
link:
url: /documentation/ops-monitoring/monitoring/
text: Read More
- partial: documentation/sections/cards-section
title: Cloud
description: Deploy and manage high-performance vector search clusters across cloud environments. Easily scale with fully managed cloud solutions, integrate seamlessly across hybrid setups, or maintain complete control with private cloud deployments in Kubernetes.
cardsPartial: documentation/cards/docs-cards
cards:
- id: 1
image:
src: /img/dev-portal-cloud/managed-cloud.png
alt: Managed Cloud
title: Managed Cloud
description: Qdrant Managed Cloud is our SaaS solution, providing managed Qdrant database clusters on the cloud.
link:
url: /documentation/cloud/
text: Read More
- id: 2
image:
src: /img/dev-portal-cloud/hybrid-cloud.png
alt: Hybrid Cloud
title: Hybrid Cloud
description: Deploy and manage your vector database across diverse environments, ensuring performance, security, and cost efficiency.
link:
url: /documentation/hybrid-cloud/
text: Read More
- id: 3
image:
src: /img/dev-portal-cloud/private-cloud.png
alt: Private Cloud
title: Private Cloud
description: Qdrant Private Cloud allows you to manage Qdrant database clusters in any Kubernetes cluster on any infrastructure.
link:
url: /documentation/private-cloud/
text: Read More
- partial: documentation/sections/cards-section
title: Support
description: Get help from the Qdrant community or contact our support team.
cardsPartial: documentation/cards/docs-cards
cardsPerRow: 2
cards:
- id: 1
icon:
src: /icons/outline/discord-purple.svg
alt: Discord icon
title: Community Support
description: Join 6,000+ active members to learn, collaborate, and participate in Qdrant’s latest activities.
link:
text: Join our Discord
url: https://qdrant.to/discord
- id: 2
icon:
src: /icons/outline/support-blue.svg
alt: Support icon
title: Qdrant Cloud Support
description: Paying customers have access to our Support team. Links to the support portal are available in the Qdrant Cloud Console.
link:
text: Join Qdrant
url: https://qdrant.to/cloud
partition: deploy
hideInSidebar: true
---
@@ -1,9 +1,11 @@
---
title: Distributed Deployment
weight: 60
partition: deploy
weight: 115
aliases:
- ../distributed_deployment
- /documentation/distributed_deployment
- /guides/distributed_deployment
- /documentation/operations/distributed_deployment
---
# Distributed deployment
@@ -32,7 +34,7 @@ In summary, single-node clusters are best for non-production workloads, replicat
## Enabling distributed mode in self-hosted Qdrant
To enable distributed deployment - enable the cluster mode in the [configuration](/documentation/operations/configuration/) or using the ENV variable: `QDRANT__CLUSTER__ENABLED=true`.
To enable distributed deployment - enable the cluster mode in the [configuration](/documentation/ops-configuration/configuration/) or using the ENV variable: `QDRANT__CLUSTER__ENABLED=true`.
```yaml
cluster:
@@ -313,7 +315,7 @@ When you add or remove nodes from the cluster, rebalancing of existing shards ac
*Available as of v1.13.0 in Cloud*
Resharding allows you to change the number of shards in your existing collections if you're hosting with our [Cloud](/documentation/cloud-intro/) offering.
Resharding allows you to change the number of shards in your existing collections if you're hosting with our [Cloud](/documentation/deploy-intro/) offering.
Resharding can change the number of shards both up and down, without having to recreate the collection from scratch.
@@ -423,7 +425,7 @@ fastest depends on the size and state of a shard.
Available shard transfer methods are:
- `stream_records`: _(default)_ transfer by streaming just its records to the target node in batches.
- `snapshot`: transfer including its index and quantized data by utilizing a [snapshot](/documentation/operations/snapshots/) automatically.
- `snapshot`: transfer including its index and quantized data by utilizing a [snapshot](/documentation/snapshots/) automatically.
- `wal_delta`: _(auto recovery default)_ transfer by resolving [WAL] difference; the operations that were missed.
Each has pros, cons and specific requirements, some of which are:
@@ -476,7 +478,7 @@ are acceptable in your use case. If your cluster is unstable and out of
resources, it's probably best to use the `stream_records` transfer method,
because it is unlikely to fail.
The `snapshot` transfer method utilizes [snapshots](/documentation/operations/snapshots/)
The `snapshot` transfer method utilizes [snapshots](/documentation/snapshots/)
to transfer a shard. A snapshot is created automatically. It is then transferred
and restored on the target node. After this is done, the snapshot is removed
from both nodes. While the snapshot/transfer/restore operation is happening, the
@@ -516,7 +518,7 @@ This ensures the availability of the data in case of node failures, except if al
### Replication factor
When you create a collection, you can control how many shard replicas you'd like to store by changing the `replication_factor`. By default, `replication_factor` is set to "1", meaning no additional copy is maintained automatically. The default can be changed in the [Qdrant configuration](/documentation/operations/configuration/#configuration-options). You can change that by setting the `replication_factor` when you create a collection.
When you create a collection, you can control how many shard replicas you'd like to store by changing the `replication_factor`. By default, `replication_factor` is set to "1", meaning no additional copy is maintained automatically. The default can be changed in the [Qdrant configuration](/documentation/ops-configuration/configuration/#configuration-options). You can change that by setting the `replication_factor` when you create a collection.
The `replication_factor` can be updated for an existing collection, but the effect of this depends on how you're running Qdrant. If you're hosting the open source version of Qdrant yourself, changing the replication factor after collection creation doesn't do anything. You can manually [create](#creating-new-shard-replicas) or drop shard replicas to achieve your desired replication factor. In Qdrant Cloud (including Hybrid Cloud, Private Cloud) your shards will automatically be replicated or dropped to match your configured replication factor.
@@ -739,7 +741,7 @@ Snapshot recovery, used in single-node deployment, is different from cluster one
Consensus manages all metadata about all collections and does not require snapshots to recover it.
But you can use snapshots to recover missing shards of the collections.
Use the [Collection Snapshot Recovery API](/documentation/operations/snapshots/#recover-in-cluster-deployment) to do it.
Use the [Collection Snapshot Recovery API](/documentation/snapshots/#recover-in-cluster-deployment) to do it.
The service will download the specified snapshot of the collection and recover shards with data from it.
Once all shards of the collection are recovered, the collection will become operational again.
@@ -3,9 +3,10 @@
title: "Getting Started"
type: delimiter
weight: 1 # Change this weight to change order of sections
hideInSidebar: true
sitemapExclude: True
_build:
publishResources: false
render: never
partition: cloud
partition: deploy
---
@@ -3,9 +3,10 @@
title: "Interfaces & Tools"
type: delimiter
weight: 25 # Change this weight to change order of sections
hideInSidebar: true
sitemapExclude: True
_build:
publishResources: false
render: never
partition: cloud
partition: deploy
---
@@ -2,10 +2,10 @@
#Delimiter files are used to separate the list of documentation pages into sections.
title: "Support"
type: delimiter
weight: 30 # Change this weight to change order of sections
weight: 300
sitemapExclude: True
_build:
publishResources: false
render: never
partition: cloud
partition: deploy
---
@@ -7,5 +7,5 @@ sitemapExclude: True
_build:
publishResources: false
render: never
partition: qdrant
partition: develop
---
@@ -1,11 +1,11 @@
---
#Delimiter files are used to separate the list of documentation pages into sections.
title: "Managed Services"
title: "Cloud"
type: delimiter
weight: 11 # Change this weight to change order of sections
weight: 200 # Change this weight to change order of sections
sitemapExclude: True
_build:
publishResources: false
render: never
partition: cloud
partition: deploy
---
@@ -7,5 +7,5 @@ sitemapExclude: True
_build:
publishResources: false
render: never
partition: qdrant
partition: develop
---
@@ -7,5 +7,5 @@ sitemapExclude: True
_build:
publishResources: false
render: never
partition: qdrant
partition: develop
---
@@ -7,5 +7,5 @@ sitemapExclude: True
_build:
publishResources: false
render: never
partition: qdrant
partition: develop
---
@@ -7,5 +7,5 @@ sitemapExclude: True
_build:
publishResources: false
render: never
partition: qdrant
partition: develop
---
@@ -1,7 +1,7 @@
---
title: "Qdrant Edge"
weight: 225
partition: qdrant
weight: 220
partition: develop
---
<aside role="status">Qdrant Edge is in beta. The API and functionality may change in future releases.</aside>
@@ -1,7 +1,7 @@
---
title: "Data Synchronization Patterns"
weight: 20
partition: qdrant
partition: develop
---
# Data Synchronization Patterns
@@ -1,7 +1,7 @@
---
title: "On-Device Embeddings"
weight: 15
partition: qdrant
partition: develop
---
# On-Device Embeddings with Qdrant Edge and FastEmbed
@@ -1,7 +1,7 @@
---
title: "Quickstart"
weight: 10
partition: qdrant
partition: develop
---
# Qdrant Edge Quickstart
@@ -1,7 +1,7 @@
---
title: "Synchronize with a Server"
weight: 30
partition: qdrant
partition: develop
---
# Synchronize Qdrant Edge with a Server
@@ -22,7 +22,7 @@ The following example shows how to integrate Gemini embeddings with Qdrant:
Let's see how to use the Embedding Model API to embed documents for retrieval.
The following example shows how to embed multiple documents with the `gemini-embedding-2-preview` model using the `RETRIEVAL_DOCUMENT` [task type](#supported-task-types):
The following example shows how to embed multiple documents with the `gemini-embedding-2` model using the `RETRIEVAL_DOCUMENT` [task type](#supported-task-types):
## Embedding a document
@@ -40,7 +40,7 @@ texts = [
]
result = gemini_client.models.embed_content(
model="gemini-embedding-2-preview",
model="gemini-embedding-2",
contents=texts,
config=types.EmbedContentConfig(task_type="RETRIEVAL_DOCUMENT"),
)
@@ -59,7 +59,7 @@ const texts = [
];
const result = await geminiClient.models.embedContent({
model: "gemini-embedding-2-preview",
model: "gemini-embedding-2",
contents: texts,
config: { taskType: "RETRIEVAL_DOCUMENT" },
});
@@ -90,7 +90,7 @@ const points = texts.map((text, idx) => ({
### Create Collection
By default, `gemini-embedding-2-preview` outputs a 3072-dimensional embedding vector. You can reduce it to a smaller size (e.g., 768 or 1536) using the `output_dimensionality` configuration to save storage space. In this example, we keep the default 3072 dimensions.
By default, `gemini-embedding-2` outputs a 3072-dimensional embedding vector. You can reduce it to a smaller size (e.g., 768 or 1536) using the `output_dimensionality` configuration to save storage space. In this example, we keep the default 3072 dimensions.
```python
client.create_collection(
@@ -124,7 +124,7 @@ Once the documents are indexed, you can search for the most relevant documents u
```python
query_result = gemini_client.models.embed_content(
model="gemini-embedding-2-preview",
model="gemini-embedding-2",
contents="Is Qdrant compatible with Gemini?",
config=types.EmbedContentConfig(task_type="RETRIEVAL_QUERY"),
)
@@ -137,7 +137,7 @@ client.query_points(
```typescript
const queryResult = await geminiClient.models.embedContent({
model: "gemini-embedding-2-preview",
model: "gemini-embedding-2",
contents: "Is Qdrant compatible with Gemini?",
config: { taskType: "RETRIEVAL_QUERY" },
});
@@ -161,7 +161,7 @@ pdf_part = types.Part.from_bytes(
)
gemini_client.models.embed_content(
model="gemini-embedding-2-preview",
model="gemini-embedding-2",
contents=[pdf_part],
)
```
@@ -173,7 +173,7 @@ const pdfBytes = readFileSync("filename.pdf");
const base64 = pdfBytes.toString("base64");
await geminiClient.models.embedContent({
model: "gemini-embedding-2-preview",
model: "gemini-embedding-2",
contents: [{
parts: [{ inlineData: { mimeType: "application/pdf", data: base64 } }],
}],
@@ -1,7 +1,7 @@
---
title: FAQ
weight: 505
partition: qdrant
weight: 510
partition: develop
# If the index.md file `is_empty`, the sidebar will display the first child link as the main entry
is_empty: true
build:
@@ -13,7 +13,7 @@ The primary source of memory usage is vector data. There are several ways to add
- Configure on-disk vector storage
The choice of the approach depends on your requirements.
Read more about [configuring the optimal](/documentation/operations/optimize/) use of Qdrant.
Read more about [configuring the optimal](/documentation/ops-optimization/optimize/) use of Qdrant.
### How do you choose the machine configuration?
@@ -18,7 +18,7 @@ In dense vectors, Qdrant supports up to 65,535 dimensions.
### What is the maximum size of vector metadata that can be stored?
There is no inherent limitation on metadata size, but it should be [optimized for performance and resource usage](/documentation/operations/optimize/). Users can set upper limits in the configuration.
There is no inherent limitation on metadata size, but it should be [optimized for performance and resource usage](/documentation/ops-optimization/optimize/). Users can set upper limits in the configuration.
### Can the same similarity search query yield different results on different machines?
@@ -1,7 +1,7 @@
---
title: "FastEmbed"
weight: 305
partition: qdrant
partition: develop
---
# What is FastEmbed?
@@ -27,7 +27,7 @@ Cheshire Cat takes great advantage of the following features of Qdrant:
* [Collection Aliases](/documentation/manage-data/collections/#collection-aliases) to manage the change from one embedder to another.
* [Quantization](/documentation/manage-data/quantization/) to obtain a good balance between speed, memory usage and quality of the results.
* [Snapshots](/documentation/operations/snapshots/) to not miss any information.
* [Snapshots](/documentation/snapshots/) to not miss any information.
* [Community](https://discord.com/invite/tdtYvXjC4h)
![RAG Pipeline](/documentation/frameworks/cheshire-cat/stregatto.jpg)
@@ -81,7 +81,7 @@ qdrant = Qdrant.from_documents(
### On-premise server deployment
No matter if you choose to launch QdrantVectorStore locally with [a Docker container](/documentation/operations/installation/), or
No matter if you choose to launch QdrantVectorStore locally with [a Docker container](/documentation/installation/), or
select a Kubernetes deployment with [the official Helm chart](https://github.com/qdrant/qdrant-helm), the way you're
going to connect to such an instance will be identical. You'll need to provide a URL pointing to the service.
@@ -3,6 +3,7 @@
| [Snapshots](/documentation/tutorials-operations/create-snapshot/) | Create and restore collection snapshots. | <span class="pill">Python</span> | 20m | <span class="text-green">Beginner</span> |
| [Data Migration](/documentation/tutorials-operations/migration/) | Move embeddings to Qdrant. | <span class="pill">CLI</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Embedding Model Migration](/documentation/tutorials-operations/embedding-model-migration/) | Use your new model with zero downtime. | <span class="pill">None</span> | 40m | <span class="text-yellow">Intermediate</span> |
| [Time-Based Sharding](/documentation/tutorials-operations/time-based-sharding/) | Efficiently manage time-series data with user-defined sharding. | <span class="pill">None</span> | 1h | <span class="text-yellow">Intermediate</span> |
| [Large-Scale Search](/documentation/tutorials-operations/large-scale-search/) | Cost-efficient search for LAION-400M datasets. | <span class="pill">None</span> | 48h | <span class="text-red">Advanced</span> |
| [Qdrant Cloud Prometheus Monitoring](/documentation/tutorials-and-examples/managed-cloud-prometheus/) | Observability with Prometheus and Grafana. | <span class="pill">Prometheus</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Self-Hosted Prometheus Monitoring](/documentation/tutorials-and-examples/hybrid-cloud-prometheus/) | Observability for hybrid/private cloud setups. | <span class="pill">Prometheus</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Qdrant Cloud Prometheus Monitoring](/documentation/ops-monitoring/managed-cloud-prometheus/) | Observability with Prometheus and Grafana. | <span class="pill">Prometheus</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Self-Hosted Prometheus Monitoring](/documentation/ops-monitoring/hybrid-cloud-prometheus/) | Observability for hybrid/private cloud setups. | <span class="pill">Prometheus</span> | 30m | <span class="text-yellow">Intermediate</span> |
@@ -0,0 +1,238 @@
using System.Net.Http;
using Microsoft.VisualBasic.FileIO;
using Qdrant.Client;
using Qdrant.Client.Grpc;
public class Snippet
{
public static async Task Run()
{
// @hide-start
string QDRANT_URL = "";
string QDRANT_API_KEY = "";
// @hide-end
// @block-start initialize-client
var client = new QdrantClient(
host: QDRANT_URL,
https: true,
apiKey: QDRANT_API_KEY
);
// @block-end initialize-client
// @block-start create-collection
string collectionName = "my_collection";
if (await client.CollectionExistsAsync(collectionName))
await client.DeleteCollectionAsync(collectionName);
await client.CreateCollectionAsync(
collectionName: collectionName,
vectorsConfig: new VectorParamsMap
{
Map = {
["dense_vector"] = new VectorParams { Size = 384, Distance = Distance.Cosine }
}
},
shardingMethod: ShardingMethod.Custom
);
// @block-end create-collection
// @block-start parse-csv
async IAsyncEnumerable<(string text, string datetime)> ParseCsv(string url)
{
using var httpClient = new HttpClient();
using var stream = await httpClient.GetStreamAsync(url);
using var parser = new TextFieldParser(new StreamReader(stream));
parser.TextFieldType = Microsoft.VisualBasic.FileIO.FieldType.Delimited;
parser.SetDelimiters(",");
string[]? headers = parser.ReadFields();
int textIdx = Array.IndexOf(headers!, "text");
int datetimeIdx = Array.IndexOf(headers!, "datetime");
while (!parser.EndOfData)
{
var fields = parser.ReadFields()!;
yield return (fields[textIdx], fields[datetimeIdx]);
}
}
// @block-end parse-csv
// @block-start upload-vectors
string csvUrl = "https://raw.githubusercontent.com/qdrant/examples/refs/heads/master/time-based-sharding/social-media-posts.csv";
// Retrieve a list of existing shard keys in the collection
var existingShardKeys = (await client.ListShardKeysAsync(collectionName))
.Select(sk => sk.Key.Keyword)
.ToHashSet();
string denseModel = "sentence-transformers/all-MiniLM-L6-v2";
int batchSize = 100;
string? currentDate = null;
var buffer = new List<PointStruct>();
await foreach (var (text, datetime) in ParseCsv(csvUrl))
{
string shardDate = datetime[..10]; // Extract YYYY-MM-DD
if (shardDate != currentDate)
{
// Flush buffer for the previous date before switching
if (buffer.Count > 0)
{
await client.UpsertAsync(
collectionName: collectionName,
points: buffer,
shardKeySelector: new ShardKeySelector
{
ShardKeys = { new List<ShardKey> { currentDate! } }
}
);
buffer.Clear();
}
// Create shard for the new date if it doesn't exist yet
if (!existingShardKeys.Contains(shardDate))
{
await client.CreateShardKeyAsync(
collectionName,
new CreateShardKey { ShardKey = new ShardKey { Keyword = shardDate } }
);
existingShardKeys.Add(shardDate);
}
currentDate = shardDate;
}
// Add point to buffer
buffer.Add(new PointStruct
{
Id = Guid.NewGuid(),
Vectors = new Dictionary<string, Vector>
{
["dense_vector"] = new Document { Text = text, Model = denseModel }
},
Payload = { ["text"] = text, ["datetime"] = datetime }
});
// Flush batch if buffer size exceeds batch size
if (buffer.Count >= batchSize)
{
await client.UpsertAsync(
collectionName: collectionName,
points: buffer,
shardKeySelector: new ShardKeySelector
{
ShardKeys = { new List<ShardKey> { currentDate! } }
}
);
buffer.Clear();
}
}
// Flush remaining partial batch
if (buffer.Count > 0)
{
await client.UpsertAsync(
collectionName: collectionName,
points: buffer,
shardKeySelector: new ShardKeySelector
{
ShardKeys = { new List<ShardKey> { currentDate! } }
}
);
}
// @block-end upload-vectors
// @block-start search-single-shard
string queryText = "coffee";
var result = await client.QueryAsync(
collectionName: collectionName,
query: new Document { Text = queryText, Model = denseModel },
usingVector: "dense_vector",
limit: 5,
shardKeySelector: new ShardKeySelector
{
ShardKeys = { new List<ShardKey> { "2026-04-07" } }
}
);
foreach (var hit in result)
Console.WriteLine(hit);
// @block-end search-single-shard
// @block-start search-multiple-shards
result = await client.QueryAsync(
collectionName: collectionName,
query: new Document { Text = queryText, Model = denseModel },
usingVector: "dense_vector",
limit: 5,
shardKeySelector: new ShardKeySelector
{
ShardKeys = { new List<ShardKey> { "2026-04-06", "2026-04-07" } }
}
);
foreach (var hit in result)
Console.WriteLine(hit);
// @block-end search-multiple-shards
// @block-start search-all-shards
result = await client.QueryAsync(
collectionName: collectionName,
query: new Document { Text = queryText, Model = denseModel },
usingVector: "dense_vector",
limit: 5
);
foreach (var hit in result)
Console.WriteLine(hit);
// @block-end search-all-shards
// @block-start pruning-shards
string today = "2026-04-08";
string oldestShardKey = DateOnly.ParseExact(today, "yyyy-MM-dd")
.AddDays(-7)
.ToString("yyyy-MM-dd");
await client.CreateShardKeyAsync(
collectionName,
new CreateShardKey { ShardKey = new ShardKey { Keyword = today } }
);
await client.DeleteShardKeyAsync(
collectionName,
new DeleteShardKey { ShardKey = new ShardKey { Keyword = oldestShardKey } }
);
// @block-end pruning-shards
// @block-start ingest-new-data
await client.UpsertAsync(
collectionName: collectionName,
points: new List<PointStruct>
{
new()
{
Id = Guid.NewGuid(),
Vectors = new Dictionary<string, Vector>
{
["dense_vector"] = new Document
{
Text = "The best way to start a Wednesday is with a cup of coffee",
Model = denseModel
}
},
Payload =
{
["text"] = "The best way to start a Wednesday is with a cup of coffee",
["datetime"] = "2026-04-08T07:57:47"
}
}
},
shardKeySelector: new ShardKeySelector
{
ShardKeys = { new List<ShardKey> { today } }
}
);
// @block-end ingest-new-data
}
}
@@ -0,0 +1,17 @@
```csharp
string collectionName = "my_collection";
if (await client.CollectionExistsAsync(collectionName))
await client.DeleteCollectionAsync(collectionName);
await client.CreateCollectionAsync(
collectionName: collectionName,
vectorsConfig: new VectorParamsMap
{
Map = {
["dense_vector"] = new VectorParams { Size = 384, Distance = Distance.Cosine }
}
},
shardingMethod: ShardingMethod.Custom
);
```
@@ -0,0 +1,21 @@
```go
collectionName := "my_collection"
exists, err := client.CollectionExists(context.Background(), collectionName)
if exists {
client.DeleteCollection(context.Background(), collectionName)
}
client.CreateCollection(context.Background(), &qdrant.CreateCollection{
CollectionName: collectionName,
VectorsConfig: qdrant.NewVectorsConfigMap(
map[string]*qdrant.VectorParams{
"dense_vector": {
Size: 384,
Distance: qdrant.Distance_Cosine,
},
},
),
ShardingMethod: qdrant.ShardingMethod_Custom.Enum(),
})
```

Some files were not shown because too many files have changed in this diff Show More