mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-01 08:58:31 +02:00
Multitenancy
This commit is contained in:
@@ -1,9 +1,9 @@
|
||||
---
|
||||
title: "Qdrant 1.16 - Scalable Multitenancy, Disk-Efficient Vector Search and Edge Beta"
|
||||
title: "Qdrant 1.16 - Scalable Multitenancy & Disk-Efficient Vector Search"
|
||||
draft: false
|
||||
slug: qdrant-1.16.x
|
||||
short_description: "v1.16 of Qdrant focuses on scalability and operational efficiency"
|
||||
description: "ToDO: add description"
|
||||
short_description: "v1.16 of Qdrant focuses on scalable multitenancy with tenant promotion and disk-efficient vector search."
|
||||
description: "v1.16 of Qdrant focuses on scalable multitenancy with tenant promotion, disk-efficient vector search with inline storage, and improved filtered vector search with ACORN."
|
||||
date: 2025-11-19T00:00:00-08:00
|
||||
author: Abdon Pijpelink
|
||||
featured: true
|
||||
@@ -15,58 +15,62 @@ tags:
|
||||
|
||||
[**Qdrant 1.16.0 is out!**](https://github.com/qdrant/qdrant/releases/tag/v1.16.0) Let’s look at the main features for this version:
|
||||
|
||||
**Scalable Multitenancy:** TODO
|
||||
**Scalable Multitenancy:** An improved approach to multitenancy that enables you to combine small and large tenants in a single collection, with the ability to promote growing tenants to dedicated shards.
|
||||
|
||||
**ACORN**: TODO
|
||||
**ACORN**: A new search algorithm that improves the quality of filtered vector search in cases of high filtering selectivity.
|
||||
|
||||
**Inline Storage**: TODO
|
||||
**Inline Storage**: A new HNSW index storage mode that stores vector data directly inside HNSW nodes, enabling efficient disk-based vector search.
|
||||
|
||||
<!-- On top of that Qdrant release includes improvements to conditional update API, which allows simpler collection migrations into newer versions of embedding models,
|
||||
fills the gap in text search API with `text_any` condition in addition to existing "match all" and "match phrase" conditions. -->
|
||||
Additionally, version 1.16 introduces a new conditional update API, facilitating easier migration of embedding models to a newer version. And, this version improved Qdrant's full-text search capabilities with a new `text_any` condition and ASCII folding support.
|
||||
|
||||
|
||||
## Tenant Promotion - Scalable Multitenancy
|
||||
## Scalable Multitenancy Using Tenant Promotion
|
||||
|
||||

|
||||
|
||||
Currently, Qdrant supports 2 aproaches to multitenancy:
|
||||
Multitenancy is a common requirement for SaaS applications, where multiple customers (tenants) share the same database instance. In Qdrant, you may be tempted to create a separate collection for each tenant, but that is not recommended. Each collection incurs a bit of resource overhead. When you have a large number of collections, this leads to increased costs, and at some point, you may see performance degradation and cluster instability. Instead, Qdrant offers two approaches to multitenancy:
|
||||
|
||||
- Payload-based multitenancy (link) - perfect for huge amount of small tenants, have practically zero overhead for tenant (and actually even more optimized than full search).
|
||||
- Shard-based multitenancy (link to custom sharding docs) - designed for smaller amount of larger tenants, which require isolation and dedicated resources.
|
||||
- [Payload-based multitenancy](/documentation/guides/multiple-partitions/), which works well when you have a large number of small tenants. This causes practically no overhead. Quite the opposite: a query with a tenant payload filter can be faster than a full search.
|
||||
- [Shard-based multitenancy](/documentation/guides/distributed_deployment/#user-defined-sharding), designed for when you have a smaller number of larger tenants. This works well when each tenant requires isolation and dedicated resources.
|
||||
|
||||
Real-world usage patterns often fall between these two use cases. It's common to have a small number of large tenants and a huge tail of smaller ones. You may even have tenants that grow over time, starting small and eventually becoming large enough to require dedicated resources.
|
||||
|
||||
But the real-world usage patterns are often more complex. Most likely the use-case will contain small amount of large tenants and huge tail of smaller ones.
|
||||
Now Qdrant can efficintly combine those two approaches with Tenant Promotion feature.
|
||||
In version 1.16, Qdrant can now efficiently combine the two multitenancy approaches with a new feature called Tenant Promotion.
|
||||
|
||||
Main principles:
|
||||
The main principles behind Tenant Promotion are:
|
||||
|
||||
- Collection should have a "shared" shard, which is used for small tenants, and large tenants occupy dedicated shards.
|
||||
- Each query to the database contain both: routing to the dedicated shard and filter. So it doesn't matter where the data is located, the query will return correct results. Aka placement of tenants is transparent to the application.
|
||||
- When a tenant grows beyond certain threshold, it is now possible to "promote" it to a dedicated shard, moving all data from shared shard to the new dedicated one. This process is implemented as a background operation which maintains consistency of data and doesn't block other operations on the collection.
|
||||
- A multitenant collection can consist of a shared "fallback" shard, which is used for small tenants, and multiple dedicated shards for large tenants.
|
||||
- Each query specifies routing to a dedicated shard, as well as a tenant filter for the fallback shard. This ensures that it doesn't matter where the data resides: the query returns the correct results. In other words, the location of tenants is transparent to the application.
|
||||
- When a tenant grows beyond a certain threshold, it is now possible to "promote" it to a dedicated shard, moving all of the tenant's data from the shared shard to the new dedicated shard. This process is implemented as a background operation that maintains data consistency. It doesn't block other operations on the collection.
|
||||
|
||||
Being able to promote a tenant to its own dedicated shard prevents a classic noisy neighbor problem where a single high-volume tenant can force the cluster to scale for everyone, increasing costs and reducing performance for smaller tenants.
|
||||
|
||||
Snippets:
|
||||
To query a collection that contains a shared fallback shard and dedicated shards, use both a tenant filter and fallback routing:
|
||||
|
||||
- One snippet which demostrates a request to collection with "fallback" routing.
|
||||
- One snippet which demostrates tenant promotion request.
|
||||
```json
|
||||
TODO: One snippet that demonstrates a request to collection with "fallback" routing.
|
||||
```
|
||||
|
||||
Once a tenant grows beyond a certain threshold, you can promote that tenant to its own dedicated shard:
|
||||
|
||||
```json
|
||||
- One snippet which demonstrates a tenant promotion request.
|
||||
Examples can be found in the integration test: https://github.com/qdrant/qdrant/blob/dev/tests/consensus_tests/test_tenant_promotion.py
|
||||
|
||||
```
|
||||
|
||||
Known limitations:
|
||||
|
||||
- Default shard-key can have only one shard-id. We will improve it in future releases to support multiple shard-ids per shard-key.
|
||||
- tenant promotion must be triggered manually. We plan to add auto-promotion in cloud offering in future.
|
||||
- The default shard key can have only one shard ID. Future releases will support multiple shard IDs per shard key.
|
||||
- Tenant promotion must be triggered manually. We plan to add auto-promotion to Qdrant Cloud in the future.
|
||||
|
||||
## ACORN - Filtered Vector Search Improvements
|
||||
|
||||

|
||||
|
||||
To enhance the scalability and speed of vector search, Qdrant employs a graph-based indexing structure known as [HNSW (Hierarchical Navigable Small World) graph-based index](/documentation/concepts/indexing/#vector-index) structure. While traditional HNSW is primarily designed for unfiltered searches, Qdrant has addressed this limitation by implementing a [filterable HSNW index](/articles/filtrable-hnsw/). This innovative approach extends the HNSW graph with additional edges that correspond to indexed payload values. The filterable HSNW index enables Qdrant to maintain search quality even with high filtering selectivity, without introducing any runtime overhead during the search process.
|
||||
To enhance the scalability and speed of vector search, Qdrant employs a graph-based index structure known as [HNSW (Hierarchical Navigable Small World)](/documentation/concepts/indexing/#vector-index). While traditional HNSW is primarily designed for unfiltered searches, Qdrant has addressed this limitation by implementing a [filterable HSNW index](/articles/filtrable-hnsw/). This innovative approach extends the HNSW graph with additional edges that correspond to indexed payload values. This enables Qdrant to maintain search quality even with high filtering selectivity, without introducing any runtime overhead during the search process.
|
||||
|
||||
Even with filterable HNSW graphs, there are instances where the quality of search results can deteriorate significantly. This can happen if a filter discards too many vectors, leading to the HNSW graph becoming [disconnected](/documentation/concepts/indexing/#filtrable-index), especially when you use a combination of high cardinality filters. It is impractical to build additional links for every possible combination of filters in advance due to the potentially vast number of combinations. Another case where filterable HNSW may break down is in situations where the filtering criteria are not known in advance.
|
||||
Even with filterable HNSW graphs, there are instances where the quality of search results can deteriorate significantly. This can happen if a filter discards too many vectors, leading to the HNSW graph becoming [disconnected](/documentation/concepts/indexing/#filtrable-index), especially when you use a combination of high cardinality filters. It is impractical to build additional links for every possible combination of filters in advance due to the potentially vast number of combinations. Another case where filterable HNSW may break down is in situations where filtering criteria are not known in advance.
|
||||
|
||||
To address these situations, in version 1.16 we are introducing support for [ACORN](documentation/concepts/search/#acorn-search-algorithm), based on the ACORN-1 algorithm described in the paper [ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data](https://arxiv.org/abs/2403.04871). When enabled, Qdrant not only traverses direct neighbors (the first hop) in the HNSW graph but also examines neighbors of neighbors (the second hop) if the direct neighbors have been filtered out. This enhancement improves search accuracy at the expense of performance, especially when multiple low-selectivity filters are applied.
|
||||
To address these limitations, in version 1.16 we are introducing support for [ACORN](documentation/concepts/search/#acorn-search-algorithm), based on the ACORN-1 algorithm described in the paper [ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data](https://arxiv.org/abs/2403.04871). With ACORN enabled, Qdrant not only traverses direct neighbors (the first hop) in the HNSW graph but also examines neighbors of neighbors (the second hop) if the direct neighbors have been filtered out. This enhancement improves search accuracy at the expense of performance, especially when multiple low-selectivity filters are applied.
|
||||
|
||||
You can enable ACORN on a per-query basis, via the optional query-time `acorn` parameter. This doesn't require any changes at index time.
|
||||
|
||||
@@ -78,7 +82,7 @@ TODO: Benchmarks of ACORN ...
|
||||
|
||||
### When Should You Use ACORN?
|
||||
|
||||
Enabling ACORN allows Qdrant to explore more nodes within the HNSW graph, which results in the evaluation of a larger number of vectors. However, this does come with some runtime overhead. It should not be enabled on every query.
|
||||
Enabling ACORN allows Qdrant to explore more nodes within the HNSW graph, which results in the evaluation of a larger number of vectors. This does come with some runtime overhead and should not be enabled on every query.
|
||||
|
||||
To help you choose when to use ACORN, refer to the following decision matrix:
|
||||
|
||||
@@ -93,13 +97,13 @@ To help you choose when to use ACORN, refer to the following decision matrix:
|
||||
|
||||

|
||||
|
||||
Deploying a vector search engine into Production often requires striking a balance between performance and cost. A good example of this trade-off is the decision between using RAM-based storage and disk-based storage for the HNSW index. HSNW was designed to be an in-memory index structure. Traversing the HNSW graph involves a lot of random access reads, which is fast in RAM, but slow on disk.
|
||||
Deploying a vector search engine into Production often requires striking a balance between performance and cost. A good example of this trade-off is the decision between RAM-based storage and disk-based storage for the HNSW index. HNSW was designed to be an in-memory index structure. Traversing the HNSW graph involves a lot of random access reads, which is fast in RAM, but slow on disk.
|
||||
|
||||
For instance, querying 1 million vectors with HNSW parameters `m` set to 16 and `ef` to 100 requires approximately 1200 vector comparisons. This is fine in RAM, but it is not acceptable on disk, where each random access read can take up to 1ms, or even longer when using HDDs instead of SSDs.
|
||||
For instance, querying 1 million vectors with the HNSW parameters `m` set to 16 and `ef` to 100 requires approximately 1200 vector comparisons. This is fine in RAM, but it is slow on disk, where each random access read can take up to 1ms, or even longer when using HDDs instead of SSDs.
|
||||
|
||||
However, disk-based storage has a property we can exploit to reduce the number of random access reads: paged reading. Disk devices typically read a full page (4KB or more) of data at once. Traditional tree-based data structures, such as B-trees, have used this property effectively. However, in graph-based structures like the HNSW index, grouping connected nodes into pages is not straightforward due to each node potentially having an arbitrary number of connections to other nodes.
|
||||
|
||||
That is, unless we can duplicate data associated with each node. That is what Qdrant version 1.16 supports with [inline storage](documentation/guides/optimize/#inline-storage-in-hnsw-index): storing vector data directly inside the HNSW nodes. This offers faster read access, at the cost of additional storage space.
|
||||
That is, unless we can duplicate the data associated with each node. This is now possible with a new feature in Qdrant version 1.16: [inline storage](documentation/guides/optimize/#inline-storage-in-hnsw-index), storing vector data directly inside the HNSW nodes. This offers faster read access, at the cost of additional storage space.
|
||||
|
||||
Let's do some napkin math:
|
||||
|
||||
@@ -138,11 +142,11 @@ TODO, after it's merged, add code snippet < code-snippet path="/documentation/he
|
||||
|
||||

|
||||
|
||||
While Qdrant is primarily a vector search engine, many applications require a combination of vector search and traditional full-text search. For that reason, we are continuously enhancing our full-text search features. In version 1.16, we have introduced two new features to improve the full-text search experience in Qdrant.
|
||||
While Qdrant is primarily a vector search engine, many applications require a combination of vector search and traditional full-text search. For that reason, we are continuously enhancing our full-text search capabilities. In version 1.16, we have introduced two new features to improve the full-text search experience in Qdrant.
|
||||
|
||||
### Match Any Condition for Text Search
|
||||
|
||||
Prior to version 1.16, Qdrant supported two way of searching for multiple search terms in text fields: the `text` condition that searches for all search terms, and the `phrase` condition that searches for an exact phrase match.
|
||||
Prior to version 1.16, Qdrant supported two ways of searching for multiple search terms in text fields: the `text` condition that searches for all search terms, and the `phrase` condition that searches for an exact phrase match.
|
||||
|
||||
However, there was no convenient way to search to match *at least one* of the provided query terms. You would have to tokenize a multi-term query yourself on the client side and build a complex boolean condition with multiple `match` conditions:
|
||||
|
||||
@@ -168,7 +172,7 @@ The `text_any` condition matches text fields that contain any of the query terms
|
||||
}
|
||||
```
|
||||
|
||||
A good example of using the `text_any` condition is in e-commerce applications, where users often search for products using multiple keywords. By combining a vector query with series of increasingly lenient full-text filters, you can ensure that users receive relevant results even if their initial search terms are too restrictive.
|
||||
A good example of using the `text_any` condition is in e-commerce applications, where users often search for products using multiple keywords. By combining a vector query with a series of increasingly lenient full-text filters, you can ensure that users receive relevant results even if their initial search terms are too restrictive.
|
||||
|
||||
```json
|
||||
batch [
|
||||
@@ -211,7 +215,7 @@ batch [
|
||||
|
||||
### ASCII Folding - Improved Search for Multilingual Texts
|
||||
|
||||
Many Latin languages use diacritical marks (accents) to indicate different pronunciations or meanings of letters. For example, the letter "é" in French is pronounced differently than "e" and can change the meaning of a word. Users, when searching for terms with diacritics, may not always include these marks in their queries, which can lead to missed matches.
|
||||
Many Latin languages use diacritical marks (accents) to indicate different pronunciations or meanings of letters. For example, the letter "é" in French is pronounced differently from "e" and can change the meaning of a word. Users, when searching for terms with diacritics, may not always include these marks in their queries, which can lead to missed matches.
|
||||
|
||||
A solution to this problem is to normalize characters with diacritics to their base ASCII equivalents, a process known as ASCII folding. For example, "café" becomes "cafe" and "naïve" becomes "naive." This normalization allows for more flexible and inclusive search results, improving search recall for multilingual texts.
|
||||
|
||||
@@ -257,7 +261,7 @@ TODO, after it's merged, add image < figure src="/docs/embedding-model-migration
|
||||
In version 1.16, we have revamped the Web UI with a fresh new look and improved user experience. The new design features the following enhancements:
|
||||
|
||||
- A new welcome page that offers quick access to tutorials and reference documentation.
|
||||
- Redesigned Point, Visualize, and Graph views in the Collections manager, making it easier to work with your data by presenting them in a more compact format.
|
||||
- Redesigned Point, Visualize, and Graph views in the Collections manager, making it easier to work with your data by presenting it in a more compact format.
|
||||
- In the tutorials, code snippets are now executed inline, which frees up screen space for better usability.
|
||||
|
||||

|
||||
@@ -266,11 +270,11 @@ In version 1.16, we have revamped the Web UI with a fresh new look and improved
|
||||
|
||||

|
||||
|
||||
All other notable improvements in a list with links to further reading.
|
||||
TODO: All other notable improvements in a list with links to further reading.
|
||||
|
||||
- RRF with configurable parameters.
|
||||
- Metrics upgrade
|
||||
- More perfrormance improvements & bug fixes, but we will know about them after changelog is out.
|
||||
- More performance improvements & bug fixes, but we will know about them after the changelog is out.
|
||||
|
||||
|
||||
## Engage
|
||||
|
||||
Reference in New Issue
Block a user