Restructure Docs - Stage 4a (#2280)

* create Develop and Deploy tabs; move Operations; re-weight pages

* move capacity planning page; create section dropdown content

* added aliases to frontmatter

* update link references to new canonical links; maintain anchoring

* address remaining link issues and errors

* fix outlier tutorial reference issue

* Treat 'develop' and 'deploy' as a unified search space

* fix some frontmatter aliases

* add section header redirects

* fix 'Operations' redirect to go to 'Deploy' tab

* update redirects file for * pattern

* add :splat to redirect references

* Add wildcard to each entry in _redirects file

---------

Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com>
This commit is contained in:
kanungle
2026-04-20 18:03:45 +02:00
committed by GitHub
co-authored by Abdon Pijpelink
parent 7b463c5ae5
commit 544708f293
155 changed files with 654 additions and 414 deletions
+6 -6
View File
@@ -32,7 +32,7 @@ Additionally, version 1.16 introduces a new conditional update API, facilitating
Multitenancy is a common requirement for SaaS applications, where multiple customers (tenants) share the same database instance. In Qdrant, when an instance is shared between multiple users, you may need to partition vectors by user. This is done so that each user can only access their own vectors and can’t see the vectors of other users. To implement multitenancy in Qdrant, there are two main approaches:
- [Payload-based multitenancy](/documentation/manage-data/multitenancy/), which works well when you have a large number of small tenants. This causes practically no overhead. Quite the opposite: a query with a tenant payload filter can be faster than a full search.
- [Shard-based multitenancy](/documentation/operations/distributed_deployment/#user-defined-sharding), designed for when you have a smaller number of larger tenants. This works well when each tenant requires isolation and dedicated resources. Separating tenants by shard prevents a classic noisy neighbor problem where a single high-volume tenant can force the cluster to scale for everyone, increasing costs and reducing performance for smaller tenants. However, shard-based multitenancy is not a good solution when you have a large number of small tenants, as each shard incurs some overhead.
- [Shard-based multitenancy](/documentation/distributed_deployment/#user-defined-sharding), designed for when you have a smaller number of larger tenants. This works well when each tenant requires isolation and dedicated resources. Separating tenants by shard prevents a classic noisy neighbor problem where a single high-volume tenant can force the cluster to scale for everyone, increasing costs and reducing performance for smaller tenants. However, shard-based multitenancy is not a good solution when you have a large number of small tenants, as each shard incurs some overhead.
Real-world usage patterns often fall between these two use cases. It's common to have a small number of large tenants and a huge tail of smaller ones. You may even have tenants that grow over time, starting small and eventually becoming large enough to require dedicated resources.
@@ -40,7 +40,7 @@ In version 1.16, Qdrant can now efficiently combine the two multitenancy approac
The main principles behind Tiered Multitenancy are:
- [User-defined Sharding](/documentation/operations/distributed_deployment/#user-defined-sharding) allows you to create named shards within a collection. It enables you to isolate large tenants into their own shards. A multitenant collection can consist of a shared "fallback" shard for small tenants and multiple dedicated shards for large tenants.
- [User-defined Sharding](/documentation/distributed_deployment/#user-defined-sharding) allows you to create named shards within a collection. It enables you to isolate large tenants into their own shards. A multitenant collection can consist of a shared "fallback" shard for small tenants and multiple dedicated shards for large tenants.
- **Fallback shards** - a special routing mechanism that allows Qdrant to route a request to either a dedicated shard (if it exists) or to a shared fallback shard. This keeps requests unified, without the need to know whether a tenant is dedicated or shared.
- [Tenant promotion](/documentation/manage-data/multitenancy/#promote-tenant-to-dedicated-shard) - a mechanism that makes it possible to "promote" tenants from the shared Fallback Shard to their own dedicated shard when they grow large enough. This process is based on Qdrant’s internal shard transfer mechanism, which makes promotion completely transparent for the application. Both read and write requests are supported during the promotion process.
@@ -126,9 +126,9 @@ For instance, querying 1 million vectors with the HNSW parameters `m` set to 16
However, disk-based storage has a property we can exploit to reduce the number of random access reads: paged reading. Disk devices typically read a full page (4KB or more) of data at once. Traditional tree-based data structures, such as B-trees, have used this property effectively. However, in graph-based structures like the HNSW index, grouping connected nodes into pages is not straightforward due to each node potentially having an arbitrary number of connections to other nodes.
With Qdrant version 1.16, you can make use of paged reading through a new feature called [inline storage](/documentation/operations/optimize/#inline-storage-in-hnsw-index). Inline storage allows for storing quantized vector data directly inside the HNSW nodes. This offers faster read access, at the cost of additional storage space.
With Qdrant version 1.16, you can make use of paged reading through a new feature called [inline storage](/documentation/ops-optimization/optimize/#inline-storage-in-hnsw-index). Inline storage allows for storing quantized vector data directly inside the HNSW nodes. This offers faster read access, at the cost of additional storage space.
Inline storage can be enabled by [setting a collection's HNSW configuration `inline_storage` option to `true`](/documentation/operations/optimize/#inline-storage-in-hnsw-index). It requires quantization to be enabled.
Inline storage can be enabled by [setting a collection's HNSW configuration `inline_storage` option to `true`](/documentation/ops-optimization/optimize/#inline-storage-in-hnsw-index). It requires quantization to be enabled.
<figure>
<img src="/blog/qdrant-1.16.x/no-inline-storage.png">
@@ -316,8 +316,8 @@ In version 1.16, we have revamped the Web UI with a fresh new look and improved
![Section 7](/blog/qdrant-1.16.x/section-7.png)
- The constant `k` that determines how Reciprocal Rank Fusion (RRF) fuses result sets [is now configurable](/documentation/search/hybrid-queries/#parametrized-rrf).
- The Metrics API now exposes [additional metrics](/documentation/operations/monitoring/#metrics) that help monitor your deployment's health.
- In strict mode, it is now possible to [configure the maximum number of payload indices](/documentation/operations/administration/#maximum-number-of-payload-index-count).
- The Metrics API now exposes [additional metrics](/documentation/ops-monitoring/monitoring/#metrics) that help monitor your deployment's health.
- In strict mode, it is now possible to [configure the maximum number of payload indices](/documentation/ops-configuration/administration/#maximum-number-of-payload-index-count).
- It's now possible to [attach custom metadata to collections](/documentation/manage-data/collections/#collection-metadata).
For a full list of all changes in version 1.16, please refer to the [change log](https://github.com/qdrant/qdrant/releases/tag/v1.16.0).