Restructure Docs - Stage 4a (#2280)

* create Develop and Deploy tabs; move Operations; re-weight pages

* move capacity planning page; create section dropdown content

* added aliases to frontmatter

* update link references to new canonical links; maintain anchoring

* address remaining link issues and errors

* fix outlier tutorial reference issue

* Treat 'develop' and 'deploy' as a unified search space

* fix some frontmatter aliases

* add section header redirects

* fix 'Operations' redirect to go to 'Deploy' tab

* update redirects file for * pattern

* add :splat to redirect references

* Add wildcard to each entry in _redirects file

---------

Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com>
This commit is contained in:
kanungle
2026-04-20 18:03:45 +02:00
committed by GitHub
co-authored by Abdon Pijpelink
parent 7b463c5ae5
commit 544708f293
155 changed files with 654 additions and 414 deletions
@@ -1,65 +1,10 @@
---
title: Operations
weight: 215
partition: qdrant
type: delimiter
weight: 100
partition: deploy
sitemapExclude: True
_build:
publishResources: false
render: never
---
# Operations
Everything you need to deploy, configure, and run Qdrant in production. These pages cover installation, capacity planning, distributed setups, security, monitoring, performance optimization, and troubleshooting.
## Capacity Planning
[Capacity Planning](/documentation/operations/capacity-planning/) helps you estimate memory and storage requirements based on your vector dimensions, quantization settings, and dataset size.
## Installation
[Installation](/documentation/operations/installation/) covers how to run Qdrant using Docker, from packages, or from source, including basic configuration options.
## Upgrading
[Upgrading](/documentation/operations/upgrades/) explains how to safely upgrade your Qdrant cluster to a new version with zero downtime, including version compatibility and client SDK updates.
## Snapshots
[Snapshots](/documentation/operations/snapshots/) describe how to back up and restore collections at a point in time, for individual nodes or the full cluster.
## Usage Statistics
[Usage Statistics](/documentation/operations/usage-statistics/) explains the anonymized telemetry Qdrant collects and how to opt out.
## Monitoring & Telemetry
[Monitoring](/documentation/operations/monitoring/) covers Qdrant's metrics endpoint, Prometheus integration, and health-check APIs for observability.
## Security
[Security](/documentation/operations/security/) explains how to enable API key authentication and TLS to secure your Qdrant instance.
## Troubleshooting
[Troubleshooting](/documentation/operations/common-errors/) lists common errors and their solutions to help diagnose issues quickly.
## Configuration
[Configuration](/documentation/operations/configuration/) documents all available Qdrant configuration file settings and environment variable overrides.
## Administration
[Administration](/documentation/operations/administration/) covers runtime administration tools, including recovery mode and collection locking.
## Distributed Deployment
[Distributed Deployment](/documentation/operations/distributed_deployment/) explains how to run Qdrant as a multi-node cluster, including sharding, replication, and consensus.
## Running with GPU
[Running with GPU](/documentation/operations/running-with-gpu/) describes how to enable GPU-accelerated indexing using dedicated Qdrant Docker images for NVIDIA and AMD GPUs.
## Optimize Performance
[Optimize Performance](/documentation/operations/optimize/) walks through three main strategies for tuning Qdrant for high speed, low memory, or high precision workloads.
## Optimizer
[Optimizer](/documentation/operations/optimizer/) describes the background optimizer process that rebuilds segments to improve search performance over time.
@@ -1,142 +0,0 @@
---
title: Administration
weight: 55
aliases:
- ../administration
---
# Administration
Qdrant exposes administration tools which enable to modify at runtime the behavior of a qdrant instance without changing its configuration manually.
## Recovery mode
*Available as of v1.2.0*
Recovery mode can help in situations where Qdrant fails to start repeatedly.
When starting in recovery mode, Qdrant only loads collection metadata to prevent
going out of memory. This allows you to resolve out of memory situations, for
example, by deleting a collection. After resolving Qdrant can be restarted
normally to continue operation.
In recovery mode, collection operations are limited to
[deleting](/documentation/manage-data/collections/#delete-collection) a
collection. That is because only collection metadata is loaded during recovery.
To enable recovery mode with the Qdrant Docker image you must set the
environment variable `QDRANT_ALLOW_RECOVERY_MODE=true`. The container will try
to start normally first, and restarts in recovery mode if initialisation fails
due to an out of memory error. This behavior is disabled by default.
If using a Qdrant binary, recovery mode can be enabled by setting a recovery
message in an environment variable, such as
`QDRANT__STORAGE__RECOVERY_MODE="My recovery message"`.
## Strict mode
*Available as of v1.13.0*
Strict mode is a feature to restrict certain type of operations on a collection in order to protect the Qdrant cluster.
The goal is to prevent inefficient usage patterns that could overload the system.
Strict mode ensures a more predictable and responsive service when you do not have control over the queries that are being executed.
Upon crossing a limit, the server will return a client side error with the information about the limit that was crossed.
The `strict_mode_config` can be enabled when [creating](#create-a-collection) a new collection, see [schema definitions](https://api.qdrant.tech/api-reference/collections/create-collection#request.body.strict_mode_config) for all the available `strict_mode_config` parameters.
As part of the config, the `enabled` field act as a toggle to enable or disable the strict mode dynamically.
It is possible to raise the default limits and/or disable strict mode entirely. Though, in order to ensure a stable cluster we strongly recommend to keep strict mode enabled using its default configuration. For disabling strict mode on an existing collection use:
{{< code-snippet path="/documentation/headless/snippets/strict-mode/disable/" >}}
### Disable retrieving via non indexed payload
Setting `unindexed_filtering_retrieve` to false prevents retrieving points by filtering on a non indexed payload key which can be very slow.
{{< code-snippet path="/documentation/headless/snippets/strict-mode/unindexed-filtering-retrieve/" >}}
Or turn it off later on an existing collection through the [update collection parameters](/documentation/manage-data/collections/#update-collection-parameters) API.
{{< code-snippet path="/documentation/headless/snippets/strict-mode/unindexed-filtering-retrieve-off/" >}}
### Disable updating via non indexed payload
Setting `unindexed_filtering_update` to false prevents updating points by filtering on a non indexed payload key which can be very slow.
{{< code-snippet path="/documentation/headless/snippets/strict-mode/unindexed-filtering-update/" >}}
### Maximum number of payload index count
Setting `max_payload_index_count` caps the maximum number of payload index that can exist on a collection.
{{< code-snippet path="/documentation/headless/snippets/strict-mode/max-payload-index-count/" >}}
### Maximum query `limit` parameter
Retrieving large result set is expensive.
Setting `max_query_limit` caps the maximum number of points that can be retrieved in a single query.
{{< code-snippet path="/documentation/headless/snippets/strict-mode/max-query-limit/" >}}
### Maximum `timeout` parameter
Long running operations are often symptomatic of a deeper issue.
Setting `max_timeout` caps the maximum value in seconds for the `timeout` parameter in all API operations.
{{< code-snippet path="/documentation/headless/snippets/strict-mode/max-timeout/" >}}
### Maximum size of a filtering condition
Large filtering conditions are expensive to evaluate.
Setting `condition_max_size` caps the maximum number of element a filtering condition can have.
e.g. the number of elements in `MatchAny`
{{< code-snippet path="/documentation/headless/snippets/strict-mode/condition-max-size/" >}}
### Maximum number of conditions in a filter
A large number of filtering conditions are expensive to evaluate.
Setting `filter_max_conditions` caps the maximum number of conditions filters can have.
{{< code-snippet path="/documentation/headless/snippets/strict-mode/filter-max-conditions/" >}}
### Maximum batch size when inserting vectors
Sending very large batch upserts can create internal congestion.
Setting `upsert_max_batchsize` caps the maximum size in bytes of a batch during vector upserts.
{{< code-snippet path="/documentation/headless/snippets/strict-mode/upsert-max-batchsize/" >}}
### Maximum collection storage size
It is possible to set the maximum size of a collection in terms of vectors and/or payload storage size.
Setting `max_collection_vector_size_bytes` and/or `max_collection_payload_size_bytes` caps the maximum byte size of a collection.
{{< code-snippet path="/documentation/headless/snippets/strict-mode/max-collection-storage-size-bytes/" >}}
### Maximum points count
Setting `max_points_count` caps the maximum number of points for a collection.
{{< code-snippet path="/documentation/headless/snippets/strict-mode/max-points-count/" >}}
### Rate limiting
An extremely high rate of incoming requests can have a negative impact on the latency.
Setting `read_rate_limit` and/or `write_rate_limit` to cap the maximum number of operations per minute per replica.
When exceeding the maximum number of operations, the client will receive an HTTP 429 error code with a suggested delay before retrying.
{{< code-snippet path="/documentation/headless/snippets/strict-mode/rate-limiting/" >}}
@@ -1,109 +0,0 @@
---
title: Capacity Planning
weight: 5
aliases:
- capacity
- /documentation/cloud/capacity-sizing
---
# Capacity Planning
When setting up your cluster, you'll need to figure out the right balance of **RAM** and **disk storage**. The best setup depends on a few things:
- How many vectors you have and their dimensions.
- The amount of payload data you're using and their indexes.
- What data you want to store in memory versus on disk.
- Your cluster's replication settings.
- Whether you're using quantization and how you’ve set it up.
## Calculating RAM size
You should store frequently accessed data in RAM for faster retrieval. If you want to keep all vectors in memory for optimal performance, you can use this rough formula for estimation:
```text
memory_size = number_of_vectors * vector_dimension * 4 bytes * 1.5
```
At the end, we multiply everything by 1.5. This extra 50% accounts for metadata (such as indexes and point versions) and temporary segments created during optimization.
Let's say you want to store 1 million vectors with 1024 dimensions:
```text
memory_size = 1,000,000 * 1024 * 4 bytes * 1.5
```
The memory_size is approximately 6,144,000,000 bytes, or about 5.72 GB.
Depending on the use case, large datasets can benefit from reduced memory requirements via [quantization](/documentation/manage-data/quantization/).
## Calculating payload size
This is always different. The size of the payload depends on the [structure and content of your data](/documentation/manage-data/payload/#payload-types). For instance:
- **Text fields** consume space based on length and encoding (e.g. a large chunk of text vs a few words).
- **Floats** have fixed sizes of 8 bytes for `int64` or `float64`.
- **Boolean fields** typically consume 1 byte.
<aside role="alert">
The easiest way to calculate your payload size is to use a JSON size calculator.
</aside>
Calculating total payload size is similar to vectors. We have to multiply it by 1.5 for back-end indexing processes.
```text
total_payload_size = number_of_points * payload_size * 1.5
```
Let's say you want to store 1 million points with JSON payloads of 5KB:
```text
total_payload_size = 1,000,000 * 5KB * 1.5
```
The total_payload_size is approximately 5,000,000 bytes, or about 4.77 GB.
## Choosing disk over RAM
For optimal performance, you should store only frequently accessed data in RAM. The rest should be offloaded to the disk. For example, extra payload fields that you don't use for filtering can be stored on disk.
Only [indexed fields](/documentation/manage-data/indexing/#payload-index) should be stored in RAM. You can read more about payload storage in the [Storage](/documentation/manage-data/storage/#payload-storage) section.
### Storage-focused configuration
If your priority is to handle large volumes of vectors with average search latency, it's recommended to configure [memory-mapped (mmap) storage](/documentation/manage-data/storage/#configuring-memmap-storage). In this setup, vectors are stored on disk in memory-mapped files, while only the most frequently accessed vectors are cached in RAM.
The amount of available RAM greatly impacts search performance. As a general rule, if you store half as many vectors in RAM, search latency will roughly double.
Disk speed is also crucial. [Contact us](/documentation/support/) if you have specific requirements for high-volume searches in our Cloud.
### Subgroup-oriented configuration
If your use case involves splitting vectors into multiple collections or subgroups based on payload values (e.g., serving searches for multiple users, each with their own subset of vectors), memory-mapped storage is recommended.
In this scenario, only the active subset of vectors will be cached in RAM, allowing for fast searches for the most recent and active users. You can estimate the required memory size as:
```text
memory_size = number_of_active_vectors * vector_dimension * 4 bytes * 1.5
```
Please refer to our [multitenancy](/documentation/manage-data/multitenancy/) documentation for more details on partitioning data in a Qdrant.
## Scaling disk space in Qdrant Cloud
Clusters supporting vector search require substantial disk space compared to other search systems. If you're running low on disk space, you can use the UI at [cloud.qdrant.io](https://cloud.qdrant.io/) to **Scale Up** your cluster.
<aside role="status">Note: If you increase disk space via the Qdrant UI, you cannot reduce it later.</aside>
When running low on disk space, consider the following benefits of scaling up:
- **Larger Datasets**: Supports larger datasets, which can improve the relevance and quality of search results.
- **Improved Indexing**: Enables the use of advanced indexing strategies like HNSW.
- **Caching**: Enhances speed by having more RAM, allowing more frequently accessed data to be cached.
- **Backups and Redundancy**: Facilitates more frequent backups, which is a key advantage for data safety.
Always remember to add 50% of the vector size. This would account for things like indexes and auxiliary data used during operations such as vector insertion, deletion, and search. Thus, the estimated memory size including metadata is:
```text
total_vector_size = number_of_dimensions * 4 bytes * 1.5
```
**Disclaimer**
The above calculations are estimates at best. If you're looking for more accurate numbers, you should always test your data set in practice.
@@ -1,143 +0,0 @@
---
title: Troubleshooting
weight: 45
aliases:
- ../tutorials/common-errors
- /documentation/troubleshooting/
- /documentation/guides/common-errors/
---
# Solving common errors
## Too many files open (OS error 24)
Each collection segment needs some files to be open. At some point you may encounter the following errors in your server log:
```text
Error: Too many files open (OS error 24)
```
In such a case you may need to increase the limit of the open files. It might be done, for example, while you launch the Docker container:
```bash
docker run --ulimit nofile=10000:10000 qdrant/qdrant:latest
```
The command above will set both soft and hard limits to `10000`.
If you are not using Docker, the following command will change the limit for the current user session:
```bash
ulimit -n 10000
```
Please note, the command should be executed before you run Qdrant server.
## Incompatible file system
Qdrant have a [set of requirements](/documentation/operations/installation/#storage) for persistent file storage.
The most important requirement is that file system **must** be [POSIX-compatible](https://www.quobyte.com/storage-explained/posix-filesystem/).
Starting from v1.15.0 Qdrant performs runtime check of file system compatibility on start.
If it detects an unknown file system, you can see a warning like this:
```text
WARN qdrant: There is a potential issue with
the filesystem for storage path ./storage. Details:
HFS/HFS+ filesystem support is untested
```
If runtime check fails, you might see an error message:
```text
ERROR qdrant: Filesystem check failed for storage path ./storage.
Details: FUSE filesystems may cause data corruption due to caching issues
```
If an error like this is reported, it is NOT safe to continue working with current configuration and you're at risk of losing your data.
Most common errors you might see, if you continue using Qdrant with incompatible file system:
```text
ERROR
Panic occurred in file /qdrant/lib/gridstore/src/gridstore.rs at line 53:
called `Result::unwrap()` on an `Err` value: OutputTooSmall { expected: 4, actual: 0 }
```
or
```text
ERROR
Service internal error: task XXX panicked with message
"called `Result::unwrap()` on an `Err` value: OutputTooSmall { expected: 4, actual: 0 }"
```
It might be also possible that vector data will be lost (set to all zeros) after service restart.
### How to avoid Incompatible file system?
Most common used configuration of incompatible file system is usage of WSL-baced Docker containers in Windows.
When you mount Windows folder into Qdrant docker container, the Windows hyper visor creates a shared mount, which is not fully POSIX-compatible.
Prefer to use docker volumes instead of bind mount:
```bash
# Create named volume
docker volume create qdrant-storage
# Use named volume with qdrant container
docker run --rm -it \
-p 6333:6333 -p 6334:6334 \
-v qdrant-storage:/qdrant/storage qdrant/qdrant:v1.15.3
```
The above keeps the volume inside the Linux container, preventing issues with a mount shared with Windows.
## Can't open Collections meta Wal
When starting a Qdrant instance as part of a distributed deployment, you may
come across an error message similar to this:
```bash
Can't open Collections meta Wal: Os { code: 11, kind: WouldBlock, message: "Resource temporarily unavailable" }
```
It means that Qdrant cannot start because a collection cannot be loaded. Its
associated [WAL](/documentation/manage-data/storage/#versioning) files are currently
unavailable, likely because the same files are already being used by another
Qdrant instance.
Each node must have their own separate storage directory, volume or mount.
The formed cluster will take care of sharing all data with each node, putting it
all in the correct places for you. If using Kubernetes, each node must have
their own volume. If using Docker, each node must have their own storage mount
or volume. If using Qdrant directly, each node must have their own storage
directory.
## Using python gRPC client with `multiprocessing`
When using the Python gRPC client with `multiprocessing`, you may encounter an error like this:
```text
<_InactiveRpcError of RPC that terminated with:
status = StatusCode.UNAVAILABLE
details = "sendmsg: Socket operation on non-socket (88)"
debug_error_string = "UNKNOWN:Error received from peer {grpc_message:"sendmsg: Socket operation on non-socket (88)", grpc_status:14, created_time:"....."}"
```
This error happens, because `multiprocessing` creates copies of gRPC channels, which share the same socket. When the parent process closes the channel, it closes the socket, and the child processes try to use a closed socket.
To prevent this error, you can use the `forkserver` or `spawn` start methods for `multiprocessing`.
```python
import multiprocessing
multiprocessing.set_start_method("forkserver") # or "spawn"
```
Alternatively, you can switch to `REST` API, async client, or use built-in parallelization in the Python client - functions like `qdrant.upload_points(...)`
@@ -1,507 +0,0 @@
---
title: Configuration
weight: 50
aliases:
- ../configuration
- /guides/configuration/
---
# Configuration
Qdrant ships with sensible defaults for collection and network settings that are suitable for most use cases. You can view these defaults in the [Qdrant source](https://github.com/qdrant/qdrant/blob/master/config/config.yaml). If you need to customize the settings, you can do so using configuration files and environment variables.
<aside role="status">
Qdrant Cloud does not allow modifying the Qdrant configuration.
</aside>
## Configuration Files
To customize Qdrant, you can mount your configuration file in any of the following locations. This guide uses `.yaml` files, but Qdrant also supports other formats such as `.toml`, `.json`, and `.ini`.
1. **Main Configuration: `qdrant/config/config.yaml`**
Mount your custom `config.yaml` file to override default settings:
```bash
docker run -p 6333:6333 \
-v $(pwd)/config.yaml:/qdrant/config/config.yaml \
qdrant/qdrant
```
2. **Environment-Specific Configuration: `config/{RUN_MODE}.yaml`**
Qdrant looks for an environment-specific configuration file based on the `RUN_MODE` variable. By default, the [official Docker image](https://hub.docker.com/r/qdrant/qdrant) uses `RUN_MODE=production`, meaning it will look for `config/production.yaml`.
You can override this by setting `RUN_MODE` to another value (e.g., `dev`), and providing the corresponding file:
```bash
docker run -p 6333:6333 \
-v $(pwd)/dev.yaml:/qdrant/config/dev.yaml \
-e RUN_MODE=dev \
qdrant/qdrant
```
3. **Local Configuration: `config/local.yaml`**
The `local.yaml` file is typically used for machine-specific settings that are not tracked in version control:
```bash
docker run -p 6333:6333 \
-v $(pwd)/local.yaml:/qdrant/config/local.yaml \
qdrant/qdrant
```
4. **Custom Configuration via `--config-path`**
You can specify a custom configuration file path using the `--config-path` argument. This will override other configuration files:
```bash
docker run -p 6333:6333 \
-v $(pwd)/config.yaml:/path/to/config.yaml \
qdrant/qdrant \
./qdrant --config-path /path/to/config.yaml
```
For details on how these configurations are loaded and merged, see the [loading order and priority](#loading-order-and-priority). The full list of available configuration options can be found [below](#configuration-options).
## Environment Variables
You can also configure Qdrant using environment variables, which always take the highest priority and override any file-based settings.
Environment variables follow this format: they should be prefixed with `QDRANT__`, and nested properties should be separated by double underscores (`__`). For example:
```bash
docker run -p 6333:6333 \
-e QDRANT__LOG_LEVEL=INFO \
-e QDRANT__SERVICE__API_KEY=<MY_SECRET_KEY> \
-e QDRANT__SERVICE__ENABLE_TLS=1 \
-e QDRANT__TLS__CERT=./tls/cert.pem \
qdrant/qdrant
```
This results in the following configuration:
```yaml
log_level: INFO
service:
enable_tls: true
api_key: <MY_SECRET_KEY>
tls:
cert: ./tls/cert.pem
```
## Loading Order and Priority
During startup, Qdrant merges multiple configuration sources into a single effective configuration. The loading order is as follows (from least to most significant):
1. Embedded default configuration
2. `config/config.yaml`
3. `config/{RUN_MODE}.yaml`
4. `config/local.yaml`
5. Custom configuration file
6. Environment variables
### Overriding Behavior
Settings from later sources in the list override those from earlier sources:
- Settings in `config/{RUN_MODE}.yaml` (3) will override those in `config/config.yaml` (2).
- A custom configuration file provided via `--config-path` (5) will override all other file-based settings.
- Environment variables (6) have the highest priority and will override any settings from files.
## Configuration Validation
Qdrant validates the configuration during startup. If any issues are found, the server will terminate immediately, providing information about the error. For example:
```console
Error: invalid type: 64-bit integer `-1`, expected an unsigned 64-bit or smaller integer for key `storage.hnsw_index.max_indexing_threads` in config/production.yaml
```
This ensures that misconfigurations are caught early, preventing Qdrant from running with invalid settings.
## Configuration Options
The following YAML example describes the available configuration options.
```yaml
log_level: INFO
# Logging configuration
# Qdrant logs to stdout. You may configure to also write logs to a file on disk.
# Be aware that this file may grow indefinitely.
# logger:
# # Logging format, supports `text` and `json`
# format: text
# on_disk:
# enabled: true
# log_file: path/to/log/file.log
# log_level: INFO
# # Logging format, supports `text` and `json`
# format: text
# buffer_size_bytes: 1024
storage:
# Where to store all the data
storage_path: ./storage
# Where to store snapshots
snapshots_path: ./snapshots
snapshots_config:
# "local" or "s3" - where to store snapshots
snapshots_storage: local
# s3_config:
# bucket: ""
# region: ""
# access_key: ""
# secret_key: ""
# Where to store temporary files
# If null, temporary snapshots are stored in: storage/snapshots_temp/
temp_path: null
# If true - point payloads will not be stored in memory.
# It will be read from the disk every time it is requested.
# This setting saves RAM by (slightly) increasing the response time.
# Note: those payload values that are involved in filtering and are indexed - remain in RAM.
#
# Default: true
on_disk_payload: true
# Maximum number of concurrent updates to shard replicas
# If `null` - maximum concurrency is used.
update_concurrency: null
# Write-ahead-log related configuration
wal:
# Size of a single WAL segment
wal_capacity_mb: 32
# Number of WAL segments to create ahead of actual data requirement
wal_segments_ahead: 0
# Normal node - receives all updates and answers all queries
node_type: "Normal"
# Listener node - receives all updates, but does not answer search/read queries
# Useful for setting up a dedicated backup node
# node_type: "Listener"
performance:
# Number of parallel threads used for search operations. If 0 - auto selection.
max_search_threads: 0
# CPU budget, how many CPUs (threads) to allocate for an optimization job.
# If 0 - auto selection, keep 1 or more CPUs unallocated depending on CPU size
# If negative - subtract this number of CPUs from the available CPUs.
# If positive - use this exact number of CPUs.
optimizer_cpu_budget: 0
# Prevent DDoS of too many concurrent updates in distributed mode.
# One external update usually triggers multiple internal updates, which breaks internal
# timings. For example, the health check timing and consensus timing.
# If null - auto selection.
update_rate_limit: null
# Limit for number of incoming automatic shard transfers per collection on this node, does not affect user-requested transfers.
# The same value should be used on all nodes in a cluster.
# Default is to allow 1 transfer.
# If null - allow unlimited transfers.
#incoming_shard_transfers_limit: 1
# Limit for number of outgoing automatic shard transfers per collection on this node, does not affect user-requested transfers.
# The same value should be used on all nodes in a cluster.
# Default is to allow 1 transfer.
# If null - allow unlimited transfers.
#outgoing_shard_transfers_limit: 1
# Enable async scorer which uses io_uring when rescoring.
# Only supported on Linux, must be enabled in your kernel.
# See: <https://qdrant.tech/articles/io_uring/#and-what-about-qdrant>
#async_scorer: false
# Maximum number of collections to load concurrently.
#max_concurrent_collection_loads: 1
# Maximum number of local shards to load concurrently when loading a collection.
#max_concurrent_shard_loads: 1
# Maximum number of segments to load concurrently when loading a local shard.
#max_concurrent_segment_loads: 8
optimizers:
# The minimal fraction of deleted vectors in a segment, required to perform segment optimization
deleted_threshold: 0.2
# The minimal number of vectors in a segment, required to perform segment optimization
vacuum_min_vector_number: 1000
# Target amount of segments optimizer will try to keep.
# Real amount of segments may vary depending on multiple parameters:
# - Amount of stored points
# - Current write RPS
#
# It is recommended to select default number of segments as a factor of the number of search threads,
# so that each segment would be handled evenly by one of the threads.
# If `default_segment_number = 0`, will be automatically selected by the number of available CPUs
default_segment_number: 0
# Do not create segments larger this size (in KiloBytes).
# Large segments might require disproportionately long indexation times,
# therefore it makes sense to limit the size of segments.
#
# If indexation speed have more priority for your - make this parameter lower.
# If search speed is more important - make this parameter higher.
# Note: 1Kb = 1 vector of size 256
# If not set, will be automatically selected considering the number of available CPUs.
max_segment_size_kb: null
# Maximum size (in KiloBytes) of vectors allowed for plain index.
# Default value based on experiments and observations.
# Note: 1Kb = 1 vector of size 256
# To explicitly disable vector indexing, set to `0`.
# If not set, the default value will be used.
indexing_threshold_kb: 10000
# Interval between forced flushes.
flush_interval_sec: 5
# Max number of threads (jobs) for running optimizations per shard.
# Note: each optimization job will also use `max_indexing_threads` threads by itself for index building.
# If null - have no limit and choose dynamically to saturate CPU.
# If 0 - no optimization threads, optimizations will be disabled.
max_optimization_threads: null
# This section has the same options as 'optimizers' above. All values specified here will overwrite the collections
# optimizers configs regardless of the config above and the options specified at collection creation.
#optimizers_overwrite:
# deleted_threshold: 0.2
# vacuum_min_vector_number: 1000
# default_segment_number: 0
# max_segment_size_kb: null
# indexing_threshold_kb: 10000
# flush_interval_sec: 5
# max_optimization_threads: null
# Default parameters of HNSW Index. Could be overridden for each collection or named vector individually
hnsw_index:
# Number of edges per node in the index graph. Larger the value - more accurate the search, more space required.
m: 16
# Number of neighbours to consider during the index building. Larger the value - more accurate the search, more time required to build index.
ef_construct: 100
# Minimal size threshold (in KiloBytes) below which full-scan is preferred over HNSW search.
# This measures the total size of vectors being queried against.
# When the maximum estimated amount of points that a condition satisfies is smaller than
# `full_scan_threshold_kb`, the query planner will use full-scan search instead of HNSW index
# traversal for better performance.
# Note: 1Kb = 1 vector of size 256
full_scan_threshold_kb: 10000
# Number of parallel threads used for background index building.
# If 0 - automatically select.
# Best to keep between 8 and 16 to prevent likelihood of building broken/inefficient HNSW graphs.
# On small CPUs, less threads are used.
max_indexing_threads: 0
# Store HNSW index on disk. If set to false, index will be stored in RAM. Default: false
on_disk: false
# Custom M param for hnsw graph built for payload index. If not set, default M will be used.
payload_m: null
# Default shard transfer method to use if none is defined.
# If null - don't have a shard transfer preference, choose automatically.
# If stream_records, snapshot or wal_delta - prefer this specific method.
# More info: https://qdrant.tech/documentation/operations/distributed_deployment/#shard-transfer-method
shard_transfer_method: null
# Default parameters for collections
collection:
# Number of replicas of each shard that network tries to maintain
replication_factor: 1
# How many replicas should apply the operation for us to consider it successful
write_consistency_factor: 1
# Default parameters for vectors.
vectors:
# Whether vectors should be stored in memory or on disk.
on_disk: null
# shard_number_per_node: 1
# Default quantization configuration.
# More info: https://qdrant.tech/documentation/manage-data/quantization
quantization: null
# Default strict mode parameters for newly created collections.
#strict_mode:
# Whether strict mode is enabled for a collection or not.
#enabled: false
# Max allowed `limit` parameter for all APIs that don't have their own max limit.
#max_query_limit: null
# Max allowed `timeout` parameter.
#max_timeout: null
# Allow usage of unindexed fields in retrieval based (eg. search) filters.
#unindexed_filtering_retrieve: null
# Allow usage of unindexed fields in filtered updates (eg. delete by payload).
#unindexed_filtering_update: null
# Max HNSW value allowed in search parameters.
#search_max_hnsw_ef: null
# Whether exact search is allowed or not.
#search_allow_exact: null
# Max oversampling value allowed in search.
#search_max_oversampling: null
# Maximum number of collections allowed to be created
# If null - no limit.
max_collections: null
service:
# Maximum size of POST data in a single request in megabytes
max_request_size_mb: 32
# Number of parallel workers used for serving the api. If 0 - equal to the number of available cores.
# If missing - Same as storage.max_search_threads
max_workers: 0
# Host to bind the service on
host: 0.0.0.0
# HTTP(S) port to bind the service on
http_port: 6333
# gRPC port to bind the service on.
# If `null` - gRPC is disabled. Default: null
# Comment to disable gRPC:
grpc_port: 6334
# Enable CORS headers in REST API.
# If enabled, browsers would be allowed to query REST endpoints regardless of query origin.
# More info: https://developer.mozilla.org/en-US/docs/Web/HTTP/CORS
# Default: true
enable_cors: true
# Enable HTTPS for the REST and gRPC API
enable_tls: false
# Check user HTTPS client certificate against CA file specified in tls config
verify_https_client_certificate: false
# Set an api-key.
# If set, all requests must include a header with the api-key.
# example header: `api-key: <API-KEY>`
#
# If you enable this you should also enable TLS.
# (Either above or via an external service like nginx.)
# Sending an api-key over an unencrypted channel is insecure.
#
# Uncomment to enable.
# api_key: your_secret_api_key_here
# Set an api-key for read-only operations.
# If set, all requests must include a header with the api-key.
# example header: `api-key: <API-KEY>`
#
# If you enable this you should also enable TLS.
# (Either above or via an external service like nginx.)
# Sending an api-key over an unencrypted channel is insecure.
#
# Uncomment to enable.
# read_only_api_key: your_secret_read_only_api_key_here
# Uncomment to enable JWT Role Based Access Control (RBAC).
# If enabled, you can generate JWT tokens with fine-grained rules for access control.
# Use generated token instead of API key.
#
# jwt_rbac: true
# Hardware reporting adds information to the API responses with a
# hint on how many resources were used to execute the request.
#
# Warning: experimental, this feature is still under development and is not supported yet.
#
# Uncomment to enable.
# hardware_reporting: true
#
# Uncomment to enable.
# Prefix for the names of metrics in the /metrics API.
# metrics_prefix: qdrant_
cluster:
# Use `enabled: true` to run Qdrant in distributed deployment mode
enabled: false
# Configuration of the inter-cluster communication
p2p:
# Port for internal communication between peers
port: 6335
# Use TLS for communication between peers
enable_tls: false
# Configuration related to distributed consensus algorithm
consensus:
# How frequently peers should ping each other.
# Setting this parameter to lower value will allow consensus
# to detect disconnected nodes earlier, but too frequent
# tick period may create significant network and CPU overhead.
# We encourage you NOT to change this parameter unless you know what you are doing.
tick_period_ms: 100
# Compact consensus operations once we have this amount of applied
# operations. Allows peers to join quickly with a consensus snapshot without
# replaying a huge amount of operations.
# If 0 - disable compaction
compact_wal_entries: 128
# Set to true to prevent service from sending usage statistics to the developers.
# Read more: https://qdrant.tech/documentation/operations/usage-statistics
telemetry_disabled: false
# TLS configuration.
# Required if either service.enable_tls or cluster.p2p.enable_tls is true.
tls:
# Server certificate chain file
cert: ./tls/cert.pem
# Server private key file
key: ./tls/key.pem
# Certificate authority certificate file.
# This certificate will be used to validate the certificates
# presented by other nodes during inter-cluster communication.
#
# If verify_https_client_certificate is true, it will verify
# HTTPS client certificate
#
# Required if cluster.p2p.enable_tls is true.
ca_cert: ./tls/cacert.pem
# TTL in seconds to reload certificate from disk, useful for certificate rotations.
# Only works for HTTPS endpoints. Does not support gRPC (and intra-cluster communication).
# If `null` - TTL is disabled.
cert_ttl: 3600
# Audit logging configuration.
# When enabled, Qdrant writes structured JSON audit log entries for every
# access-checked API request.
#
# audit:
# enabled: false
# dir: ./storage/audit
# rotation: daily
# max_log_files: 7
# # If true, use X-Forwarded-For header to determine client IP in audit logs.
# # Only enable this when running behind a trusted reverse proxy or load balancer.
# # WARNING: Enabling this without a trusted proxy allows clients to spoof their IP.
# # Default: false
# trust_forwarded_headers: false
```
File diff suppressed because it is too large Load Diff
@@ -1,234 +0,0 @@
---
title: Installation
weight: 10
aliases:
- ../install
- ../installation
---
# Installation requirements
The following sections describe the requirements for deploying Qdrant.
## CPU and memory
The preferred size of your CPU and RAM depends on:
- Number of vectors
- Vector dimensions
- [Payloads](/documentation/manage-data/payload/) and their indexes
- Storage
- Replication
- How you configure quantization
Our [Cloud Pricing Calculator](https://cloud.qdrant.io/calculator) can help you estimate required resources without payload or index data.
### Supported CPU architectures:
**64-bit system:**
- x86_64/amd64
- AArch64/arm64
**32-bit system:**
- Not supported
### Storage
For persistent storage, Qdrant requires block-level access to storage devices with a [POSIX-compatible file system](https://www.quobyte.com/storage-explained/posix-filesystem/). Network systems such as [iSCSI](https://en.wikipedia.org/wiki/ISCSI) that provide block-level access are also acceptable.
Qdrant won't work with [Network file systems](https://en.wikipedia.org/wiki/File_system#Network_file_systems) such as NFS, or [Object storage](https://en.wikipedia.org/wiki/Object_storage) systems such as S3.
If you offload vectors to a local disk, we recommend you use a solid-state (SSD or NVMe) drive.
<aside role="status">Using Docker/WSL on Windows with mounts is known to have file system problems causing data loss. See <a href="/documentation/operations/common-errors/#incompatible-file-system">troubleshooting</a>.</aside>
### Networking
Each Qdrant instance requires three open ports:
* `6333` - For the HTTP API, for the [Monitoring](/documentation/operations/monitoring/) health and metrics endpoints
* `6334` - For the [gRPC](/documentation/interfaces/#grpc-interface) API
* `6335` - For [Distributed deployment](/documentation/operations/distributed_deployment/)
All Qdrant instances in a cluster must be able to:
- Communicate with each other over these ports
- Allow incoming connections to ports `6333` and `6334` from clients that use Qdrant.
### Security
The default configuration of Qdrant might not be secure enough for every situation. Please see [our security documentation](/documentation/operations/security/) for more information.
## Installation options
Qdrant can be installed in different ways depending on your needs:
For production, you can use our Qdrant Cloud to run Qdrant either fully managed in our infrastructure or with Hybrid Cloud in yours.
If you want to run Qdrant in your own infrastructure, without any cloud connection, we recommend to install Qdrant in a Kubernetes cluster with our Qdrant Private Cloud Enterprise Operator.
For testing or development setups, you can run the Qdrant container or as a binary executable. We also provide a Helm chart for an easy installation in Kubernetes.
## Production
### Qdrant Cloud
You can set up production with the [Qdrant Cloud](https://qdrant.to/cloud), which provides fully managed Qdrant databases.
It provides horizontal and vertical scaling, one click installation and upgrades, monitoring, logging, as well as backup and disaster recovery. For more information, see the [Qdrant Cloud documentation](/documentation/cloud/).
### Qdrant Kubernetes Operator
We provide a Qdrant Enterprise Operator for Kubernetes installations as part of our [Qdrant Private Cloud](/documentation/private-cloud/) offering. For more information, [use this form](https://qdrant.to/contact-us) to contact us.
### Kubernetes
You can use a ready-made [Helm Chart](https://helm.sh/docs/) to run Qdrant in your Kubernetes cluster. While it is possible to deploy Qdrant in a distributed setup with the Helm chart, it does not come with the same level of features for zero-downtime upgrades, up and down-scaling, monitoring, logging, and backup and disaster recovery as the Qdrant Cloud offering or the Qdrant Private Cloud Enterprise Operator. Instead you must manage and set this up [yourself](/documentation/operations/distributed_deployment/). Support for the Helm chart is limited to community support.
The following table gives you an overview about the feature differences between the Qdrant Cloud and the Helm chart:
| Feature | Qdrant Helm Chart | Qdrant Cloud |
|--------------------------------------------------------|:-----------------:|:-------------:|
| Open-source | ✅ | |
| Community support only | ✅ | |
| Quick to get started | ✅ | ✅ |
| Vertical and horizontal scaling | ✅ | ✅ |
| API keys with granular access control | ✅ | ✅ |
| Qdrant version upgrades | ✅ | ✅ |
| Support for transit and storage encryption | ✅ | ✅ |
| Zero-downtime upgrades with optimized restart strategy | | ✅ |
| Production ready out-of the box | | ✅ |
| Dataloss prevention on downscaling | | ✅ |
| Full cluster backup and disaster recovery | | ✅ |
| Automatic shard rebalancing | | ✅ |
| Re-sharding support | | ✅ |
| Automatic persistent volume scaling | | ✅ |
| Advanced telemetry | | ✅ |
| One-click API key revoking | | ✅ |
| Recreating nodes with new volumes in existing cluster | | ✅ |
| Enterprise support | | ✅ |
To install the helm chart:
```bash
helm repo add qdrant https://qdrant.to/helm
helm install qdrant qdrant/qdrant
```
For more information, see the [qdrant-helm](https://github.com/qdrant/qdrant-helm/tree/main/charts/qdrant) README.
### Docker and Docker Compose
Usually, we recommend to run Qdrant in Kubernetes, or use the Qdrant Cloud for production setups. This makes setting up highly available and scalable Qdrant clusters with backups and disaster recovery a lot easier.
However, you can also use Docker and Docker Compose to run Qdrant in production, by following the setup instructions in the [Docker](#docker) and [Docker Compose](#docker-compose) Development sections.
In addition, you have to make sure:
* To use a performant [persistent storage](#storage) for your data
* To configure the [security settings](/documentation/operations/security/) for your deployment
* To set up and configure Qdrant on multiple nodes for a highly available [distributed deployment](/documentation/operations/distributed_deployment/)
* To set up a load balancer for your Qdrant cluster
* To create a [backup and disaster recovery strategy](/documentation/operations/snapshots/) for your data
* To integrate Qdrant with your [monitoring](/documentation/operations/monitoring/) and logging solutions
## Development
For development and testing, we recommend that you set up Qdrant in Docker. We also have different client libraries.
### Docker
The easiest way to start using Qdrant for testing or development is to run the Qdrant container image.
The latest versions are always available on [DockerHub](https://hub.docker.com/r/qdrant/qdrant/tags?page=1&ordering=last_updated).
Make sure that [Docker](https://docs.docker.com/engine/install/), [Podman](https://podman.io/docs/installation) or the container runtime of your choice is installed and running. The following instructions use Docker.
Pull the image:
```bash
docker pull qdrant/qdrant
```
In the following command, revise `$(pwd)/path/to/data` for your Docker configuration. Then use the updated command to run the container:
```bash
docker run -p 6333:6333 \
-v $(pwd)/path/to/data:/qdrant/storage \
qdrant/qdrant
```
With this command, you start a Qdrant instance with the default configuration.
It stores all data in the `./path/to/data` directory.
By default, Qdrant uses port 6333, so at [localhost:6333](http://localhost:6333) you should see the welcome message.
To change the Qdrant configuration, you can overwrite the production configuration:
```bash
docker run -p 6333:6333 \
-v $(pwd)/path/to/data:/qdrant/storage \
-v $(pwd)/path/to/custom_config.yaml:/qdrant/config/production.yaml \
qdrant/qdrant
```
Alternatively, you can use your own `custom_config.yaml` configuration file:
```bash
docker run -p 6333:6333 \
-v $(pwd)/path/to/data:/qdrant/storage \
-v $(pwd)/path/to/custom_config.yaml:/qdrant/config/custom_config.yaml \
qdrant/qdrant \
./qdrant --config-path config/custom_config.yaml
```
For more information, see the [Configuration](/documentation/operations/configuration/) documentation.
### Docker Compose
You can also use [Docker Compose](https://docs.docker.com/compose/) to run Qdrant.
Here is an example customized compose file for a single node Qdrant cluster:
```yaml
services:
qdrant:
image: qdrant/qdrant:latest
restart: always
container_name: qdrant
ports:
- 6333:6333
- 6334:6334
expose:
- 6333
- 6334
- 6335
configs:
- source: qdrant_config
target: /qdrant/config/production.yaml
volumes:
- ./qdrant_data:/qdrant/storage
configs:
qdrant_config:
content: |
log_level: INFO
```
<aside role="status">Providing the inline <code>content</code> in the <a href="https://docs.docker.com/compose/compose-file/08-configs/">configs top-level element</a> requires <a href="https://docs.docker.com/compose/release-notes/#2231">Docker Compose v2.23.1</a> or above. This functionality is supported starting <a href="https://docs.docker.com/engine/release-notes/25.0/#2500">Docker Engine v25.0.0</a> and <a href="https://docs.docker.com/desktop/release-notes/#4260">Docker Desktop v4.26.0</a> onwards.</aside>
### From source
Qdrant is written in Rust and can be compiled into a binary executable.
This installation method can be helpful if you want to compile Qdrant for a specific processor architecture or if you do not want to use Docker.
Before compiling, make sure that the necessary libraries and the [rust toolchain](https://www.rust-lang.org/tools/install) are installed.
The current list of required libraries can be found in the [Dockerfile](https://github.com/qdrant/qdrant/blob/master/Dockerfile).
Build Qdrant with Cargo:
```bash
cargo build --release --bin qdrant
```
After a successful build, you can find the binary in the following subdirectory `./target/release/qdrant`.
## Client libraries
In addition to the service, Qdrant provides a variety of client libraries for different programming languages. For a full list, see our [Client libraries](/documentation/interfaces/#client-libraries) documentation.
@@ -1,170 +0,0 @@
---
title: Monitoring & Telemetry
weight: 35
aliases:
- ../monitoring
---
# Monitoring & Telemetry
Qdrant exposes its metrics in [Prometheus](https://prometheus.io/docs/instrumenting/exposition_formats/#text-based-format)/[OpenMetrics](https://github.com/OpenObservability/OpenMetrics) format, so you can integrate them easily
with the compatible tools and monitor Qdrant with your own monitoring system. You can
use the `/metrics` endpoint and configure it as a scrape target.
Metrics endpoint: <http://localhost:6333/metrics>
The integration with Qdrant is easy to
[configure](https://prometheus.io/docs/prometheus/latest/getting_started/#configure-prometheus-to-monitor-the-sample-targets)
with Prometheus and Grafana.
## Metrics
Qdrant exposes various metrics in Prometheus/OpenMetrics format, commonly used together with Grafana for monitoring.
Two endpoints are available:
- `/metrics` for metrics of a Qdrant node/peer, see [all metrics](#node-metrics-metrics).
- `/sys_metrics` (Qdrant Cloud only) for metrics about your cluster, like CPU, memory, disk utilisation, collection metrics and load balancer telemetry. For more information, see [Qdrant Cloud Monitoring](/documentation/cloud/cluster-monitoring/).
Note that `/metrics` only reports metrics for the peer connected to. It is therefore important to scrape from each peer individually, even if a load balancer is involved.
### Node metrics `/metrics`
Each Qdrant node will expose the following metrics.
Counters - such as the number of created snapshots - are reset when the node is restarted.
**Application metrics**
| Name | Type | Meaning |
| ----------------------------------- | ------- | ------------------------------ |
| app_info | gauge | Qdrant server name and version |
| app_status_recovery_mode | gauge | If started in recovery mode |
**Collection metrics**
| Name | Type | Meaning |
| ------------------------------------------------- | ------- | ----------------------------------------------------------------------------------------------------- |
| collections_total | gauge | Number of collections |
| collection_points | gauge | Number of points, per collection <sup>(v1.16+)</sup> |
| collection_vectors | gauge | Number of vectors, per collection and vector name <sup>(v1.16+)</sup> |
| collections_vector_total | gauge | Number of vectors in all collections |
| collection_indexed_only_excluded_points | gauge | Number of points excluded in [`indexed_only`](/documentation/search/search/#search-api) search, per collection and vector name <sup>(v1.16+)</sup> |
| collection_active_replicas_min | gauge | Minimum number of active replicas across all collections and shards <sup>(v1.16+)</sup> |
| collection_active_replicas_max | gauge | Maximum number of active replicas across all collections and shards <sup>(v1.16+)</sup> |
| collection_dead_replicas | gauge | Number of non-active replicas across all collections and shards <sup>(v1.16+)</sup> |
| collection_running_optimizations | gauge | Number of running optimization tasks, per collection <sup>(v1.16+)</sup> |
| collection_hardware_metric_cpu | counter | CPU measurements of a collection, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
| collection_hardware_metric_payload_io_read | counter | Payload IO read operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
| collection_hardware_metric_payload_io_write | counter | Payload IO write operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
| collection_hardware_metric_payload_index_io_read | counter | Payload index read operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
| collection_hardware_metric_payload_index_io_write | counter | Payload index write operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
| collection_hardware_metric_vector_io_read | counter | Vector IO read operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
| collection_hardware_metric_vector_io_write | counter | Vector IO write operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
[^metrics-hwreporting]: Only reported if hardware metrics are enabled in the configuration. See `service.hardware_reporting` in the [configuration](/documentation/operations/configuration/).
**Snapshot metrics**
| Name | Type | Meaning |
| --------------------------------------- | ------- | --------------------------------------------------------------------------- |
| snapshot_creation_running | gauge | Number of snapshots being created, per collection <sup>(v1.16+)</sup> |
| snapshot_recovery_running | gauge | Number of snapshots being recovered, per collection <sup>(v1.16+)</sup> |
| snapshot_created_total | counter | Number of created snapshots since start, per collection <sup>(v1.16+)</sup> |
**API response metrics**
| Name | Type | Meaning |
| ----------------------------------- | --------- | ------------------------------------------------------------------ |
| rest_responses_total | counter | Number of responses through REST API |
| rest_responses_fail_total | counter | Number of failed responses through REST API |
| rest_responses_avg_duration_seconds | gauge | Average response duration in REST API |
| rest_responses_min_duration_seconds | gauge | Minimum response duration in REST API |
| rest_responses_max_duration_seconds | gauge | Maximum response duration in REST API |
| rest_responses_duration_seconds | histogram | Histogram of response durations in the REST API <sup>(v1.8+)</sup> |
| grpc_responses_total | counter | Number of responses through gRPC API |
| grpc_responses_fail_total | counter | Number of failed responses through REST API |
| grpc_responses_avg_duration_seconds | gauge | Average response duration in gRPC API |
| grpc_responses_min_duration_seconds | gauge | Minimum response duration in gRPC API |
| grpc_responses_max_duration_seconds | gauge | Maximum response duration in gRPC API |
| grpc_responses_duration_seconds | histogram | Histogram of response durations in the gRPC API <sup>(v1.8+)</sup> |
**Process metrics**
| Name | Type | Meaning |
| ----------------------------------- | ------- | ----------------------------------------------------------------------------------------------------------------------------- |
| memory_active_bytes | gauge | Total number of bytes in active pages allocated by the application ([ref](https://jemalloc.net/jemalloc.3.html#stats.active)) |
| memory_allocated_bytes | gauge | Total number of bytes allocated by the application ([ref](https://jemalloc.net/jemalloc.3.html#stats.allocated)) |
| memory_metadata_bytes | gauge | Total number of bytes dedicated to allocator metadata ([ref](https://jemalloc.net/jemalloc.3.html#stats.metadata)) |
| memory_resident_bytes | gauge | Maximum number of bytes in physically resident data pages mapped ([ref](https://jemalloc.net/jemalloc.3.html#stats.resident)) |
| memory_retained_bytes | gauge | Total number of bytes in virtual memory mappings ([ref](https://jemalloc.net/jemalloc.3.html#stats.retained)) |
| process_threads | gauge | Number of used system threads <sup>(v1.16+)</sup> |
| process_open_mmaps | gauge | Number of open memory maps <sup>(v1.16+)</sup> |
| system_max_mmaps | gauge | System wide maximum number of open memory maps <sup>(v1.16+)</sup> |
| process_open_fds | gauge | Number of open file descriptors <sup>(v1.16+)</sup> |
| process_max_fds | gauge | Maximum number of open file descriptors <sup>(v1.16+)</sup> |
| process_minor_page_faults_total | counter | Number of minor page faults encountered by the process <sup>(v1.16+)</sup> |
| process_major_page_faults_total | counter | Number of major page faults encountered by the process <sup>(v1.16+)</sup> |
**Cluster metrics (consensus)**
Metrics reporting the current cluster consensus state of the node. Exposed only
when distributed mode is enabled.
| Name | Type | Meaning |
| -------------------------------- | ------- | ----------------------------------------------------------------------- |
| cluster_enabled | gauge | If distributed mode is enabled [^metrics-distributed] |
| cluster_peers_total | gauge | Number of cluster peers [^metrics-distributed] |
| cluster_term | counter | Raft consensus term [^metrics-distributed] |
| cluster_commit | counter | Raft consensus commit - last committed operation [^metrics-distributed] |
| cluster_pending_operations_total | gauge | Number of pending consensus operations [^metrics-distributed] |
| cluster_voter | gauge | If a consensus voter (`1`) or learner (`0`) [^metrics-distributed] |
[^metrics-distributed]: Only reported if distributed mode (cluster mode) is enabled. Enabled by default in all Qdrant Cloud environments. See `cluster.enabled` in the [configuration](/documentation/operations/configuration/).
### Metrics configuration
*Available as of v1.16.0*
In self-hosted environments you have further configuration options for metrics.
By default, all Qdrant metrics have no application namespace prefix. You may set
a prefix with `service.metrics_prefix` in the
[configuration](/documentation/operations/configuration/).
To achieve this you may use the following environment variable for example:
```bash
QDRANT__SERVICE__METRICS_PREFIX="qdrant_"
```
## Telemetry endpoint
Qdrant also provides a `/telemetry` endpoint, which provides information about the current state of the database, including the number of vectors, shards, and other useful information. You can find the full documentation for this endpoint in the [API reference](https://api.qdrant.tech/api-reference/service/telemetry).
## Cluster-wide telemetry
The `/telemetry` endpoint reports from the point of view of the peer being queried. Qdrant also provides a `/cluster/telemetry` endpoint, which aggregates telemetry from all peers.
This includes less information than `/telemetry`, but provides information like shard transfer progress more reliably.
You can find the full documentation for this endpoint in the [API reference](https://api.qdrant.tech/api-reference/service/cluster-telemetry).
## Kubernetes health endpoints
*Available as of v1.5.0*
Qdrant exposes three endpoints, namely
[`/healthz`](http://localhost:6333/healthz),
[`/livez`](http://localhost:6333/livez) and
[`/readyz`](http://localhost:6333/readyz), to indicate the current status of the
Qdrant server.
These currently provide the most basic status response, returning HTTP 200 if
Qdrant is started and ready to be used.
Regardless of whether an [API key](/documentation/operations/security/#authentication) is configured,
the endpoints are always accessible.
You can read more about Kubernetes health endpoints
[here](https://kubernetes.io/docs/reference/using-api/health-checks/).
@@ -1,126 +0,0 @@
---
title: Optimize Performance
weight: 70
aliases:
- ../tutorials/optimize
---
# Optimizing Qdrant Performance: Three Scenarios
Different use cases require different balances between memory usage, search speed, and precision. Qdrant is designed to be flexible and customizable so you can tune it to your specific needs.
This guide will walk you three main optimization strategies:
- High Speed Search & Low Memory Usage
- High Precision & Low Memory Usage
- High Precision & High Speed Search
![qdrant resource tradeoffs](/docs/tradeoff.png)
## 1. High-Speed Search with Low Memory Usage
To achieve high search speed with minimal memory usage, you can store vectors on disk while minimizing the number of disk reads. Vector quantization is a technique that compresses vectors, allowing more of them to be stored in memory, thus reducing the need to read from disk.
To configure in-memory quantization, with on-disk original vectors, you need to create a collection with the following parameters:
- `on_disk`: Stores original vectors on disk.
- `quantization_config`: Compresses quantized vectors to `int8` using the `scalar` method.
- `always_ram`: Keeps quantized vectors in RAM.
{{< code-snippet path="/documentation/headless/snippets/create-collection/scalar-quantization-in-ram/" >}}
### Disable Rescoring for Faster Search (optional)
This is completely optional. Disabling rescoring with search `params` can further reduce the number of disk reads. Note that this might slightly decrease precision.
{{< code-snippet path="/documentation/headless/snippets/query-points/disable-quantization-rescoring/" >}}
## 2. High Precision with Low Memory Usage
If you require high precision but have limited RAM, you can store both vectors and the HNSW index on disk. This setup reduces memory usage while maintaining search precision.
To store the vectors `on_disk`, you need to configure both the vectors and the HNSW index:
{{< code-snippet path="/documentation/headless/snippets/create-collection/with-vectors-and-hnsw-on-disk/" >}}
### Improving Precision
Increase the `ef` and `m` parameters of the HNSW index to improve precision, even with limited RAM:
```json
...
"hnsw_config": {
"m": 64,
"ef_construct": 512,
"on_disk": true
}
...
```
**Note:** The speed of this setup depends on the disk’s IOPS (Input/Output Operations Per Second).</br>
You can use [fio](https://gist.github.com/superboum/aaa45d305700a7873a8ebbab1abddf2b) to measure disk IOPS.
### Inline Storage in HNSW Index
*Available as of v1.16.0*
When storing vectors and the HNSW index on disk, you can improve search performance by enabling the `inline_storage` option in the `hnsw_config`.
With inline storage, Qdrant stores copies of vectors directly within the HNSW index file.
It makes searches faster by reducing the number of IO operations, at the cost of 3-4x increased storage usage.
It requires quantization to be enabled.
{{< code-snippet path="/documentation/headless/snippets/create-collection/with-inline-storage/" >}}
## 3. High Precision with High-Speed Search
For scenarios requiring both high speed and high precision, keep as much data in RAM as possible. Apply quantization with re-scoring for tunable accuracy.
Here is how you can configure scalar quantization for a collection:
{{< code-snippet path="/documentation/headless/snippets/create-collection/scalar-quantization-and-vectors-in-ram/" >}}
### Fine-Tuning Search Parameters
You can adjust search parameters like `hnsw_ef` and `exact` to balance between speed and precision:
**Key Parameters:**
- `hnsw_ef`: Number of neighbors to visit during search (higher value = better accuracy, slower speed).
- `exact`: Set to `true` for exact search, which is slower but more accurate. You can use it to compare results of the search with different `hnsw_ef` values versus the ground truth.
{{< code-snippet path="/documentation/headless/snippets/query-points/with-params/" >}}
## Balancing Latency and Throughput
When optimizing search performance, latency and throughput are two main metrics to consider:
- **Latency:** Time taken for a single request.
- **Throughput:** Number of requests handled per second.
The following optimization approaches are not mutually exclusive, but in some cases it might be preferable to optimize for one or another.
### Minimizing Latency
To minimize latency, you can set up Qdrant to use as many cores as possible for a single request.
You can do this by setting the number of segments in the collection to be equal to the number of cores in the system.
In this case, each segment will be processed in parallel, and the final result will be obtained faster.
{{< code-snippet path="/documentation/headless/snippets/create-collection/with-high-number-of-segments/" >}}
### Maximizing Throughput
To maximize throughput, configure Qdrant to use as many cores as possible to process multiple requests in parallel.
To do that, use fewer segments (usually 2) of larger size (default 200Mb per segment) to handle more requests in parallel.
Large segments benefit from the size of the index and overall smaller number of vector comparisons required to find the nearest neighbors. However, they will require more time to build the HNSW index.
{{< code-snippet path="/documentation/headless/snippets/create-collection/with-large-segments/" >}}
## Summary
By adjusting configurations like vector storage, quantization, and search parameters, you can optimize Qdrant for different use cases:
- **Low Memory + High Speed:** Use vector quantization.
- **High Precision + Low Memory:** Store vectors and HNSW index on disk.
- **High Precision + High Speed:** Keep data in RAM, use quantization with re-scoring.
- **Latency vs. Throughput:** Adjust segment numbers based on the priority.
Choose the strategy that best fits your use case to get the most out of Qdrant’s performance capabilities.
@@ -1,196 +0,0 @@
---
title: Optimizer
weight: 75
aliases:
- ../optimizer
---
# Optimizer
It is much more efficient to apply changes in batches than perform each change individually, as many other databases do. Qdrant here is no exception. Since Qdrant operates with data structures that are not always easy to change, it is sometimes necessary to rebuild those structures completely.
Storage optimization in Qdrant occurs at the segment level (see [storage](/documentation/manage-data/storage/)).
In this case, the segment to be optimized remains readable for the time of the rebuild.
![Segment optimization](/articles_data/immutable-data-structures/optimization.png)
The availability is achieved by wrapping the segment into a proxy that transparently handles data changes.
Changed data is placed in the copy-on-write segment, which has priority for retrieval and subsequent updates.
## Vacuum Optimizer
The simplest example of a case where you need to rebuild a segment repository is to remove points.
Like many other databases, Qdrant does not delete entries immediately after a query.
Instead, it marks records as deleted and ignores them for future queries.
This strategy allows us to minimize disk access - one of the slowest operations.
However, a side effect of this strategy is that, over time, deleted records accumulate, occupy memory and slow down the system.
To avoid these adverse effects, Vacuum Optimizer is used.
It is used if the segment has accumulated too many deleted records.
The criteria for starting the optimizer are defined in the configuration file.
Here is an example of parameter values:
```yaml
storage:
optimizers:
# The minimal fraction of deleted vectors in a segment, required to perform segment optimization
deleted_threshold: 0.2
# The minimal number of vectors in a segment, required to perform segment optimization
vacuum_min_vector_number: 1000
```
## Merge Optimizer
The service may require the creation of temporary segments.
Such segments, for example, are created as copy-on-write segments during optimization itself.
It is also essential to have at least one small segment that Qdrant will use to store frequently updated data.
On the other hand, too many small segments lead to suboptimal search performance.
The merge optimizer constantly tries to reduce the number of segments if there
currently are too many. The desired number of segments is specified
with `default_segment_number` and defaults to the number of CPUs. The optimizer
may takes at least the three smallest segments and merges them into one.
Segments will not be merged if they'll exceed the maximum configured segment
size with `max_segment_size_kb`. It prevents creating segments that are too
large to efficiently index. Increasing this number may help to reduce the number
of segments if you have a lot of data, and can potentially improve search performance.
The criteria for starting the optimizer are defined in the configuration file.
Here is an example of parameter values:
```yaml
storage:
optimizers:
# Target amount of segments optimizer will try to keep.
# Real amount of segments may vary depending on multiple parameters:
# - Amount of stored points
# - Current write RPS
#
# It is recommended to select default number of segments as a factor of the number of search threads,
# so that each segment would be handled evenly by one of the threads.
# If `default_segment_number = 0`, will be automatically selected by the number of available CPUs
default_segment_number: 0
# Do not create segments larger this size (in KiloBytes).
# Large segments might require disproportionately long indexation times,
# therefore it makes sense to limit the size of segments.
#
# If indexation speed have more priority for your - make this parameter lower.
# If search speed is more important - make this parameter higher.
# Note: 1Kb = 1 vector of size 256
# If not set, will be automatically selected considering the number of available CPUs.
max_segment_size_kb: null
```
## Indexing Optimizer
Qdrant allows you to choose the type of indexes and data storage methods used depending on the number of records.
So, for example, if the number of points is less than 10000, using any index would be less efficient than a brute force scan.
The Indexing Optimizer is used to implement the enabling of indexes and memmap storage when the minimal amount of records is reached.
The criteria for starting the optimizer are defined in the configuration file.
Here is an example of parameter values:
```yaml
storage:
optimizers:
# Maximum size (in kilobytes) of vectors to store in-memory per segment.
# Segments larger than this threshold will be stored as read-only memmaped file.
# Memmap storage is disabled by default, to enable it, set this threshold to a reasonable value.
# To disable memmap storage, set this to `0`.
# Note: 1Kb = 1 vector of size 256
memmap_threshold: 200000
# Maximum size (in KiloBytes) of vectors allowed for plain index.
# Default value based on experiments and observations.
# Note: 1Kb = 1 vector of size 256
# To explicitly disable vector indexing, set to `0`.
# If not set, the default value will be used.
indexing_threshold_kb: 10000
```
In addition to the configuration file, you can also set optimizer parameters separately for each [collection](/documentation/manage-data/collections/).
Dynamic parameter updates may be useful, for example, for more efficient initial loading of points. You can disable indexing during the upload process with these settings and enable it immediately after it is finished. As a result, you will not waste extra computation resources on rebuilding the index.
## Prevent Reads from Unindexed Segments
*Available as of v1.17.1*
<aside role="alert"><code>prevent_unoptimized</code> is an experimental feature; its behavior may change slightly in future releases and it must be used with care.</aside>
When a collection receives a high volume of updates, for example, during nightly batch updates or when processing a large backlog of updates after a period of downtime, the optimizer might not be able to index new points fast enough to keep up. When this happens, searches may slow down as Qdrant has to scan through large amounts of unindexed data for every query.
To address this, Qdrant supports [querying indexed data only](/documentation/search/low-latency-search/#query-indexed-data-only), by setting `indexed_only` to `true`. Because updates in Qdrant are implemented as a delete followed by an insert, a side effect of searching indexed data only is that it can cause recently updated data to temporarily disappear from search results until it is indexed again. This is because the delete operation immediately removes the old point from the index, while the insert operation adds the new point to an unindexed segment that is not yet visible to searches.
To mitigate this, the optimizer supports a `prevent_unoptimized` mode. When enabled, points written to an unindexed segment that is larger than `indexing_threshold` are accepted and durably stored but are not visible in search results until the optimizer has indexed the segment. These are called deferred points. Not until the optimizer finishes indexing a segment containing deferred points, do those points become visible.
Set `prevent_unoptimized` to `true` when creating or updating a collection:
{{< code-snippet path="/documentation/headless/snippets/update-collection/prevent-unoptimized/" >}}
<aside role="status">
Enabling <code>prevent_unoptimized</code> only affects newly created segments. Existing segments are not changed retroactively. Similarly, changing <code>indexing_threshold</code> does not affect existing segments. Only new segments will use the updated threshold.
</aside>
With `prevent_unoptimized` enabled, setting `indexed_only` to `true` is not necessary to avoid slow searches, as unindexed segments do not return deferred points.
| `prevent_unoptimized` | `indexed_only` | Effect |
|----------------------|----------------|--------|
| `false` (default) | `false` | All points are searchable, but searches may be slow if there are many unindexed points. |
| `false` (default) | `true` | Only indexed points are searchable, but recently updated points may temporarily disappear from search results until they are indexed. |
| `true` | `false` (default) | Reads return indexed points and unindexed points that are not deferred. Deferred points are not visible to reads until indexed. |
| `true` | `true` | Only indexed points are visible to reads. |
### Effect on `wait=true`
Qdrant processes updates in strict order: each update is written to the write-ahead log and then applied sequentially by the update worker, preserving this order.
Under normal conditions, setting `wait=true` on a write request returns after the update has been applied to a segment. After enabling `prevent_unoptimized`, the response is held until every deferred point, including the current update, has been indexed and is visible for search. Depending on the volume of updates and the speed of the optimizer, this can take a significant amount of time and may lead to timeouts on the client side. If the client times out, the update can be expected to be durably stored and eventually indexed, but the client will not receive a confirmation for that specific request.
Because the update worker must finish indexing before continuing to consume the queue, a blocked `wait=true` request also delays all subsequent updates that use `wait=true`. Updates with `wait=false` are written to the write-ahead log immediately, but they are not applied until the blocked request unblocks. This head-of-line blocking means that `wait=true` can stall the entire update pipeline for as long as indexing takes. Use it with caution when `prevent_unoptimized` is enabled and the cluster is under heavy write load.
### Monitoring Deferred Points
You can check the number of deferred points in a collection via the `update_queue` section in the response of the [collection info API](/documentation/manage-data/collections/#collection-info). The same information is also available in [telemetry and metrics](/documentation/operations/monitoring/), enabling dashboards and alerting.
A non-zero deferred point count means the optimizer is processing a backlog. This is expected under heavy write load; monitor the count to confirm that it is decreasing over time.
## Optimization Monitoring
*Available as of v1.17.0*
The `/collections/{collection_name}/optimizations` API endpoint returns information about the optimization of a specific collection, including:
- A summary of optimization activity, with the number of queued optimizations, queued segments, queued points, and idle segments (segments that need no optimization).
- Details about any currently running optimization, including:
- the specific optimizer
- its status
- the segments involved
- its progress
Optionally, you can use the `with` query parameter with one or more of the following comma-separated values to retrieve additional information:
- `queued`, to return a list of queued optimizations
- `completed`, to return a list of completed optimizations
- `idle_segments`, to return a list of idle segments
For example:
{{< code-snippet path="/documentation/headless/snippets/optimizations/" >}}
### Web UI
The same information is also accessible via the **Optimizations** tab within the **Collections** interface in [the Web UI](/documentation/web-ui/). For a specific collection, this tab provides an overview of the current optimization status and a timeline of current and past optimization cycles:
![The Optimizations tab in Web UI shows progress and a timeline of optimization cycles](/docs/web-ui-optimizations-progress-timeline.png)
Selecting a specific optimization cycle from the timeline provides detailed information about the tasks performed during that cycle, including their durations:
![The Optimizations tab in Web UI provides access to detailed information about optimization tasks and their durations](/docs/web-ui-optimizations-tree.png)
@@ -1,211 +0,0 @@
---
title: Running with GPU
weight: 65
aliases:
- /documentation/guides/running-with-GPU/
---
# Running Qdrant with GPU Support
Starting from version v1.13.0, Qdrant offers support for GPU acceleration.
However, GPU support is not included in the default Qdrant binary due to additional dependencies and libraries. Instead, you will need to use dedicated Docker images with GPU support ([NVIDIA](#nvidia-gpus), [AMD](#amd-gpus)).
## Configuration
Qdrant includes a number of configuration options to control GPU usage. The following options are available:
```yaml
gpu:
# Enable GPU indexing.
indexing: false
# Force half precision for `f32` values while indexing.
# `f16` conversion will take place
# only inside GPU memory and won't affect storage type.
force_half_precision: false
# Used vulkan "groups" of GPU.
# In other words, how many parallel points can be indexed by GPU.
# Optimal value might depend on the GPU model.
# Proportional, but doesn't necessary equal
# to the physical number of warps.
# Do not change this value unless you know what you are doing.
# Default: 512
groups_count: 512
# Filter for GPU devices by hardware name. Case insensitive.
# Comma-separated list of substrings to match
# against the gpu device name.
# Example: "nvidia"
# Default: "" - all devices are accepted.
device_filter: ""
# List of explicit GPU devices to use.
# If host has multiple GPUs, this option allows to select specific devices
# by their index in the list of found devices.
# If `device_filter` is set, indexes are applied after filtering.
# By default, all devices are accepted.
devices: null
# How many parallel indexing processes are allowed to run.
# Default: 1
parallel_indexes: 1
# Allow to use integrated GPUs.
# Default: false
allow_integrated: false
# Allow to use emulated GPUs like LLVMpipe. Useful for CI.
# Default: false
allow_emulated: false
```
It is not recommended to change these options unless you are familiar with the Qdrant internals and the Vulkan API.
## Standalone GPU Support
For standalone usage, you can build Qdrant with GPU support by running the following command:
```bash
cargo build --release --features gpu
```
Ensure your device supports Vulkan API v1.3. This includes compatibility with Apple Silicon, Intel GPUs, and CPU emulators. Note that `gpu.indexing: true` must be set in your configuration to use GPUs at runtime.
## NVIDIA GPUs
### Prerequisites
To use Docker with NVIDIA GPU support, ensure the following are installed on your host:
- Latest NVIDIA drivers
- [nvidia-container-toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html)
Most AI or CUDA images on Amazon/GCP come pre-configured with the NVIDIA container toolkit.
### Docker images with NVIDIA GPU support
Docker images with NVIDIA GPU support use the tag suffix `gpu-nvidia`, e.g., `qdrant/qdrant:v1.13.0-gpu-nvidia`. These images include all necessary dependencies.
To enable GPU support, use the `--gpus=all` flag with Docker settings. Example:
```bash
# `--gpus=all` flag says to Docker that we want to use GPUs.
# `-e QDRANT__GPU__INDEXING=1` flag says to Qdrant that we want to use GPUs for indexing.
docker run \
--rm \
--gpus=all \
-p 6333:6333 \
-p 6334:6334 \
-e QDRANT__GPU__INDEXING=1 \
qdrant/qdrant:gpu-nvidia-latest
```
To ensure that the GPU was initialized correctly, you may check it in logs. First Qdrant prints all found GPU devices without filtering and then prints list of all created devices:
```text
2025-01-13T11:58:29.124087Z INFO gpu::instance: Found GPU device: NVIDIA GeForce RTX 3090
2025-01-13T11:58:29.124118Z INFO gpu::instance: Found GPU device: llvmpipe (LLVM 15.0.7, 256 bits)
2025-01-13T11:58:29.124138Z INFO gpu::device: Create GPU device NVIDIA GeForce RTX 3090
```
Here you can see that two devices were found: RTX 3090 and llvmpipe (a CPU-emulated GPU which is included in the Docker image). Later, you will see that only RTX was initialized.
This concludes the setup. Now, you can start using this Qdrant instance.
### Troubleshooting NVIDIA GPUs
If your GPU is not detected in Docker, make sure your driver and `nvidia-container-toolkit` are up-to-date.
If needed, you can install latest version of `nvidia-container-toolkit` from it's GitHub Releases [page](https://github.com/NVIDIA/nvidia-container-toolkit/releases)
Verify Vulkan API visibility in the Docker container using:
```bash
docker run --rm --gpus=all qdrant/qdrant:gpu-nvidia-latest vulkaninfo --summary
```
The system may show you an error message explaining why the NVIDIA device is not visible.
Note that if your NVIDIA GPU is not visible in Docker, the Docker image cannot use libGLX_nvidia.so.0 on your host. Here is what an error message could look like:
```text
ERROR: [Loader Message] Code 0 : loader_scanned_icd_add: Could not get `vkCreateInstance` via `vk_icdGetInstanceProcAddr` for ICD libGLX_nvidia.so.0
WARNING: [Loader Message] Code 0 : terminator_CreateInstance: Failed to CreateInstance in ICD 0. Skipping ICD.
```
To resolve errors, update your NVIDIA container runtime configuration:
```bash
sudo nano /etc/nvidia-container-runtime/config.toml
```
Set `no-cgroups=false`, save the configuration, and restart Docker:
```bash
sudo systemctl restart docker
```
## AMD GPUs
### Prerequisites
Running Qdrant with AMD GPUs requires [ROCm](https://rocm.docs.amd.com/projects/install-on-linux/en/latest/install/detailed-install.html) to be installed on your host.
### Docker images with AMD GPU support
Docker images for AMD GPUs use the tag suffix `gpu-amd`, e.g., `qdrant/qdrant:v1.13.0-gpu-amd`. These images include all required dependencies.
To enable GPU for Docker, you need additional `--device /dev/kfd --device /dev/dri` flags. To enable GPU for Qdrant you need to set the enable flag. Here is an example:
```bash
# `--device /dev/kfd --device /dev/dri` flags say to Docker that we want to use GPUs.
# `-e QDRANT__GPU__INDEXING=1` flag says to Qdrant that we want to use GPUs for indexing.
docker run \
--rm \
--device /dev/kfd --device /dev/dri \
-p 6333:6333 \
-p 6334:6334 \
-e QDRANT__LOG_LEVEL=debug \
-e QDRANT__GPU__INDEXING=1 \
qdrant/qdrant:gpu-amd-latest
```
Check logs to confirm GPU initialization. Example log output:
```text
2025-01-10T11:56:55.926466Z INFO gpu::instance: Found GPU device: AMD Radeon Graphics (RADV GFX1103_R1)
2025-01-10T11:56:55.926485Z INFO gpu::instance: Found GPU device: llvmpipe (LLVM 17.0.6, 256 bits)
2025-01-10T11:56:55.926504Z INFO gpu::device: Create GPU device AMD Radeon Graphics (RADV GFX1103_R1)
```
This concludes the setup. In a basic scenario, you won't need to configure anything else.
## Known limitations
* **Platform Support:** Docker images are only available for Linux x86_64. Windows, macOS, ARM, and other platforms are not supported.
* **Memory Limits:** Each GPU can process up to 16GB of vector data per indexing iteration.
Due to this limitation, you should not create segments where either original vectors OR quantized vectors are larger than 16GB.
For example, a collection with 1536d vectors and scalar quantization can have at most:
```text
16Gb / 1536 ~= 11 million vectors per segment
```
And without quantization:
```text
16Gb / 1536 * 4 ~= 2.7 million vectors per segment
```
The maximum size of each segment can be configured in the collection settings.
Use the following operation to [change](/documentation/manage-data/collections/#update-collection-parameters) on your existing collection:
```http
PATCH collections/{collection_name}
{
"optimizers_config": {
"max_segment_size": 1000000
}
}
```
Note that `max_segment_size` is specified in KiloBytes.
@@ -1,643 +0,0 @@
---
title: Security
weight: 40
aliases:
- ../security
---
# Security
Qdrant supports various security features to help you secure your instance. Most
of these must to be explicitly configured to make your instance production
ready. Please read the following section carefully.
## Secure Your Instance
<aside role="alert">Custom deployments are <b>not</b> secure by default and are <b>not</b> production ready. Qdrant Cloud deployments are always secure and production ready.</aside>
By default, all self-deployed Qdrant instances are not secure. They are open to
all network interfaces and do not have any kind of authentication configured. They
may be open to everybody on the internet without any restrictions. You must
therefore take security measures to make your instance production-ready.
Please read through this section carefully for instructions on how to secure
your instance.
Instances deployed via Qdrant Cloud are always secure by default. Refer to
[Authentication](/documentation/cloud/authentication/) and [Client IP
Restrictions](/documentation/cloud/configure-cluster/#client-ip-restrictions).
To properly secure your own instance, we strongly recommend taking the following steps:
1. [Authentication](#authentication): set up an API key to prevent unauthorized access.
The most important step to prevent unauthenticated actors from accessing your data.
2. [Network Bind](#network-bind): bind to a specific network interface or IP address.
When developing locally, bind to `127.0.0.1` to prevent all external access.
When deploying to production, bind to a private network interface or IP.
3. [TLS](#tls): enable encrypted traffic everywhere using TLS.
## Authentication
*Available as of v1.2.0*
Qdrant supports a simple form of client authentication using a static API key.
This can be used to secure your instance.
To enable API key based authentication in your own Qdrant instance you must
specify a key in the configuration:
```yaml
service:
# Set an api-key.
# If set, all requests must include a header with the api-key.
# example header: `api-key: <API-KEY>`
#
# If you enable this you should also enable TLS.
# (Either above or via an external service like nginx.)
# Sending an api-key over an unencrypted channel is insecure.
api_key: your_secret_api_key_here
```
Or alternatively, you can use the environment variable:
```bash
docker run -p 6333:6333 \
-e QDRANT__SERVICE__API_KEY=your_secret_api_key_here \
qdrant/qdrant
```
<aside role="alert"><a href="#tls">TLS</a> must be used to prevent leaking the API key over an unencrypted connection.</aside>
For using API key based authentication in Qdrant Cloud see the cloud
[Authentication](/documentation/cloud/authentication/)
section.
The API key then needs to be present in all REST or gRPC requests to your instance.
All official Qdrant clients for Python, Go, Rust, .NET and Java support the API key parameter.
<!---
Examples with clients
-->
```bash
curl \
-X GET https://localhost:6333 \
--header 'api-key: your_secret_api_key_here'
```
```python
from qdrant_client import QdrantClient
client = QdrantClient(
url="https://localhost:6333",
api_key="your_secret_api_key_here",
)
```
```typescript
import { QdrantClient } from "@qdrant/js-client-rest";
const client = new QdrantClient({
url: "http://localhost",
port: 6333,
apiKey: "your_secret_api_key_here",
});
```
```rust
use qdrant_client::Qdrant;
let client = Qdrant::from_url("https://xyz-example.eu-central.aws.cloud.qdrant.io:6334")
.api_key("<paste-your-api-key-here>")
.build()?;
```
```java
import io.qdrant.client.QdrantClient;
import io.qdrant.client.QdrantGrpcClient;
QdrantClient client =
new QdrantClient(
QdrantGrpcClient.newBuilder(
"xyz-example.eu-central.aws.cloud.qdrant.io",
6334,
true)
.withApiKey("<paste-your-api-key-here>")
.build());
```
```csharp
using Qdrant.Client;
var client = new QdrantClient(
host: "xyz-example.eu-central.aws.cloud.qdrant.io",
https: true,
apiKey: "<paste-your-api-key-here>"
);
```
```go
import "github.com/qdrant/go-client/qdrant"
client, err := qdrant.NewClient(&qdrant.Config{
Host: "xyz-example.eu-central.aws.cloud.qdrant.io",
Port: 6334,
APIKey: "<paste-your-api-key-here>",
UseTLS: true,
})
```
<aside role="alert">Internal communication channels are <strong>never</strong> protected by an API key nor bearer tokens. Internal gRPC uses port 6335 by default if running in distributed mode. You must ensure that this port is not publicly reachable and can only be used for node communication. By default, this setting is disabled for Qdrant Cloud and the Qdrant Helm chart.</aside>
### Read-Only API Key
*Available as of v1.7.0*
In addition to the regular API key, Qdrant also supports a read-only API key.
This key can be used to access read-only operations on the instance.
```yaml
service:
read_only_api_key: your_secret_read_only_api_key_here
```
Or with the environment variable:
```bash
export QDRANT__SERVICE__READ_ONLY_API_KEY=your_secret_read_only_api_key_here
```
Both API keys can be used simultaneously.
### Rotate an API Key
*Available as of v1.17.0*
In a distributed deployment, you can rotate an API key without downtime. Use the `alt_api_key` setting to temporarily configure a second API key that acts identically to the primary `api_key`, allowing both the old and new API keys to be active at the same time.
```yaml
service:
api_key: your_current_api_key_here
alt_api_key: your_new_api_key_here
```
To rotate an API key without downtime:
1. Configure each peer with the new key set as `alt_api_key`. Restart only one peer at a time to avoid downtime (rolling restart). During the rotation window, requests authenticated with either key are accepted.
2. Switch clients to the new key.
3. Perform another rolling restart of the peers, promoting the new key to `api_key` and removing `alt_api_key`.
<aside role="alert">JWT tokens are tied to the key they were signed with and are <strong>not</strong> automatically migrated. They must be re-created after switching to the new key.</aside>
### Granular Access Control with JWT
*Available as of v1.9.0*
For more complex cases, Qdrant supports granular access control with [JSON Web Tokens (JWT)](https://jwt.io/).
This allows you to create tokens which restrict access to data stored in your cluster, and build [Role-based access control (RBAC)](https://en.wikipedia.org/wiki/Role-based_access_control) on top of that.
In this way, you can define permissions for users and restrict access to sensitive endpoints.
To enable JWT-based authentication in your own Qdrant instance you need to specify the `api-key` and enable the `jwt_rbac` feature in the configuration:
```yaml
service:
api_key: you_secret_api_key_here
jwt_rbac: true
```
Or with the environment variables:
```bash
export QDRANT__SERVICE__API_KEY=your_secret_api_key_here
export QDRANT__SERVICE__JWT_RBAC=true
```
The `api_key` you set in the configuration will be used to encode and decode the JWTs, so –needless to say– keep it secure. If your `api_key` changes, all existing tokens will be invalid.
To use JWT-based authentication, you need to provide it as a bearer token in the `Authorization` header, or as an key in the `Api-Key` header of your requests.
```http
Authorization: Bearer <JWT>
// or
Api-Key: <JWT>
```
```python
from qdrant_client import QdrantClient
qdrant_client = QdrantClient(
"xyz-example.eu-central.aws.cloud.qdrant.io",
api_key="<JWT>",
)
```
```typescript
import { QdrantClient } from "@qdrant/js-client-rest";
const client = new QdrantClient({
host: "xyz-example.eu-central.aws.cloud.qdrant.io",
apiKey: "<JWT>",
});
```
```rust
use qdrant_client::Qdrant;
let client = Qdrant::from_url("https://xyz-example.eu-central.aws.cloud.qdrant.io:6334")
.api_key("<JWT>")
.build()?;
```
```java
import io.qdrant.client.QdrantClient;
import io.qdrant.client.QdrantGrpcClient;
QdrantClient client =
new QdrantClient(
QdrantGrpcClient.newBuilder(
"xyz-example.eu-central.aws.cloud.qdrant.io",
6334,
true)
.withApiKey("<JWT>")
.build());
```
```csharp
using Qdrant.Client;
var client = new QdrantClient(
host: "xyz-example.eu-central.aws.cloud.qdrant.io",
https: true,
apiKey: "<JWT>"
);
```
```go
import "github.com/qdrant/go-client/qdrant"
client, err := qdrant.NewClient(&qdrant.Config{
Host: "xyz-example.eu-central.aws.cloud.qdrant.io",
Port: 6334,
APIKey: "<JWT>",
UseTLS: true,
})
```
#### Generating JSON Web Tokens
Due to the nature of JWT, anyone who knows the `api_key` can generate tokens by using any of the existing libraries and tools, it is not necessary for them to have access to the Qdrant instance to generate them.
For convenience, we have added a JWT generation tool the Qdrant Web UI under the 🔑 tab, if you're using the default url, it will be at `http://localhost:6333/dashboard#/jwt`.
- **JWT Header** - Qdrant uses the `HS256` algorithm to decode the tokens.
```json
{
"alg": "HS256",
"typ": "JWT"
}
```
- **JWT Payload** - You can include any combination of the [parameters available](#jwt-configuration) in the payload. Keep reading for more info on each one.
```json
{
"exp": 1640995200, // Expiration time
"value_exists": ..., // Validate this token by looking for a point with a payload value
"access": "r", // Define the access level.
}
```
**Signing the token** - To confirm that the generated token is valid, it needs to be signed with the `api_key` you have set in the configuration.
That would mean, that someone who knows the `api_key` gives the authorization for the new token to be used in the Qdrant instance.
Qdrant can validate the signature, because it knows the `api_key` and can decode the token.
The process of token generation can be done on the client side offline, and doesn't require any communication with the Qdrant instance.
Here is an example of libraries that can be used to generate JWT tokens:
- Python: [PyJWT](https://pyjwt.readthedocs.io/en/stable/)
- JavaScript: [jsonwebtoken](https://www.npmjs.com/package/jsonwebtoken)
- Rust: [jsonwebtoken](https://crates.io/crates/jsonwebtoken)
- CLI: [jwt-cli](https://github.com/mike-engel/jwt-cli)
Here is an example using `jwt-cli`:
```bash
jwt encode --payload '{
"access": "r",
"exp": 1766055305
}' --secret 'your-api-key'
```
#### JWT Configuration
These are the available options, or **claims** in the JWT lingo. You can use them in the JWT payload to define its functionality.
- **`exp`** - The expiration time of the token. This is a Unix timestamp in seconds. The token will be invalid after this time. The check for this claim includes a 30-second leeway to account for clock skew.
```json
{
"exp": 1640995200, // Expiration time
}
```
- **`value_exists`** - This is a claim that can be used to validate the token against the data stored in a collection. Structure of this claim is as follows:
```json
{
"value_exists": {
"collection": "my_validation_collection",
"matches": [
{ "key": "my_key", "value": "value_that_must_exist" }
],
},
}
```
If this claim is present, Qdrant will check if there is a point in the collection with the specified key-values. If it does, the token is valid.
This claim is especially useful if you want to have an ability to revoke tokens without changing the `api_key`.
Consider a case where you have a collection of users, and you want to revoke access to a specific user.
```json
{
"value_exists": {
"collection": "users",
"matches": [
{ "key": "user_id", "value": "andrey" },
{ "key": "role", "value": "manager" }
],
},
}
```
You can create a token with this claim, and when you want to revoke access, you can change the `role` of the user to something else, and the token will be invalid.
- **`access`** - This claim defines the [access level](#table-of-access) of the token. If this claim is present, Qdrant will check if the token has the required access level to perform the operation. If this claim is **not** present, **manage** access is assumed.
It can provide global access with `r` for read-only, or `m` for manage. For example:
```json
{
"access": "r"
}
```
It can also be specific to one or more collections. The `access` level for each collection is `r` for read-only, or `rw` for read-write, like this:
```json
{
"access": [
{
"collection": "my_collection",
"access": "rw"
}
]
}
```
### Table of Access
Check out this table to see which actions are allowed or denied based on the access level.
This is also applicable to using api keys instead of tokens. In that case, `api_key` maps to **manage**, while `read_only_api_key` maps to **read-only**.
<div style="text-align: right"> <strong>Symbols:</strong> ✅ Allowed | ❌ Denied | 🟡 Allowed, but filtered </div>
| Action | manage | read-only | collection read-write | collection read-only |
|--------|--------|-----------|----------------------|-----------------------|
| list collections | ✅ | ✅ | 🟡 | 🟡 |
| get collection info | ✅ | ✅ | ✅ | ✅ |
| create collection | ✅ | ❌ | ❌ | ❌ |
| delete collection | ✅ | ❌ | ❌ | ❌ |
| update collection params | ✅ | ❌ | ❌ | ❌ |
| get collection cluster info | ✅ | ✅ | ✅ | ✅ |
| collection exists | ✅ | ✅ | ✅ | ✅ |
| update collection cluster setup | ✅ | ❌ | ❌ | ❌ |
| update aliases | ✅ | ❌ | ❌ | ❌ |
| list collection aliases | ✅ | ✅ | 🟡 | 🟡 |
| list aliases | ✅ | ✅ | 🟡 | 🟡 |
| create shard key | ✅ | ❌ | ❌ | ❌ |
| delete shard key | ✅ | ❌ | ❌ | ❌ |
| create payload index | ✅ | ❌ | ✅ | ❌ |
| delete payload index | ✅ | ❌ | ✅ | ❌ |
| list collection snapshots | ✅ | ✅ | ✅ | ✅ |
| create collection snapshot | ✅ | ❌ | ✅ | ❌ |
| delete collection snapshot | ✅ | ❌ | ✅ | ❌ |
| download collection snapshot | ✅ | ✅ | ✅ | ✅ |
| upload collection snapshot | ✅ | ❌ | ❌ | ❌ |
| recover collection snapshot | ✅ | ❌ | ❌ | ❌ |
| list shard snapshots | ✅ | ✅ | ✅ | ✅ |
| create shard snapshot | ✅ | ❌ | ✅ | ❌ |
| delete shard snapshot | ✅ | ❌ | ✅ | ❌ |
| download shard snapshot | ✅ | ✅ | ✅ | ✅ |
| upload shard snapshot | ✅ | ❌ | ❌ | ❌ |
| recover shard snapshot | ✅ | ❌ | ❌ | ❌ |
| list full snapshots | ✅ | ✅ | ❌ | ❌ |
| create full snapshot | ✅ | ❌ | ❌ | ❌ |
| delete full snapshot | ✅ | ❌ | ❌ | ❌ |
| download full snapshot | ✅ | ✅ | ❌ | ❌ |
| get cluster info | ✅ | ✅ | ❌ | ❌ |
| recover raft state | ✅ | ❌ | ❌ | ❌ |
| delete peer | ✅ | ❌ | ❌ | ❌ |
| get point | ✅ | ✅ | ✅ | ✅ |
| get points | ✅ | ✅ | ✅ | ✅ |
| upsert points | ✅ | ❌ | ✅ | ❌ |
| update points batch | ✅ | ❌ | ✅ | ❌ |
| delete points | ✅ | ❌ | ✅ | ❌ | ❌ /
| update vectors | ✅ | ❌ | ✅ | ❌ |
| delete vectors | ✅ | ❌ | ✅ | ❌ | ❌ /
| set payload | ✅ | ❌ | ✅ | ❌ |
| overwrite payload | ✅ | ❌ | ✅ | ❌ |
| delete payload | ✅ | ❌ | ✅ | ❌ |
| clear payload | ✅ | ❌ | ✅ | ❌ |
| scroll points | ✅ | ✅ | ✅ | ✅ |
| query points | ✅ | ✅ | ✅ | ✅ |
| search points | ✅ | ✅ | ✅ | ✅ |
| search groups | ✅ | ✅ | ✅ | ✅ |
| recommend points | ✅ | ✅ | ✅ | ✅ |
| recommend groups | ✅ | ✅ | ✅ | ✅ |
| discover points | ✅ | ✅ | ✅ | ✅ |
| count points | ✅ | ✅ | ✅ | ✅ |
| version | ✅ | ✅ | ✅ | ✅ |
| readyz, healthz, livez | ✅ | ✅ | ✅ | ✅ |
| telemetry | ✅ | ✅ | ❌ | ❌ |
| metrics | ✅ | ✅ | ❌ | ❌ |
## Audit Logging
*Available as of v1.17.0*
Audit logging records all API operations that require authentication or authorization, and writes them to a log file in JSON format.
Audit logging is not enabled by default. To enable it, use the following configuration options:
```yaml
audit:
enabled: false
dir: ./storage/audit
rotation: daily
max_log_files: 7
```
By default, audit logs are rotated daily, and the seven most recent log files are kept. To configure hourly rotation, set `rotation` to `hourly`. When the number of log files exceeds `max_log_files`, the oldest log file is deleted.
<aside role="alert">Audit logging is verbose and audit logs can grow in size rapidly. Ensure that you have sufficient disk space.</aside>
## Network Bind
By default, a custom Qdrant deployment binds to all network interfaces. Your
instance may be open to everybody on the internet. On a local development
machine you likely have a firewall in place to prevent public access, but that
may not be the case on a public VPS or dedicated server.
It is highly recommended to bind to a specific interface or IP address to
prevent unwanted access:
- when developing locally, bind to `127.0.0.1` so no external access is possible
- or, when deploying to production, bind to a private network interface or IP
When using Docker, you may use the publish flag to bind to a specific interface.
For example:
```bash
docker run -p 127.0.0.1:6333:6333 qdrant/qdrant
```
If using another type of deployment you may configure the bind address in Qdrant
itself. Either set `service.host: 127.0.0.1` in the configuration, or use an
environment variable like this:
```bash
QDRANT__SERVICE__HOST=127.0.0.1 ./qdrant
```
Managed Qdrant Cloud deployments are always secure by default. They are publicly
accessible and bound to the endpoint that is assigned to the cluster. You may
configure authentication with [API keys](/documentation/cloud/authentication/),
and restrict access to specific IP addresses through [Client IP
Restrictions](/documentation/cloud/configure-cluster/#client-ip-restrictions).
[Hybrid Cloud](/documentation/hybrid-cloud/networking-logging-monitoring/) and
[Private
Cloud](/documentation/private-cloud/qdrant-cluster-management/#exposing-a-cluster)
deployments have their own kind of configuration.
## TLS
*Available as of v1.2.0*
TLS for encrypted connections can be enabled on your Qdrant instance to secure
connections.
<aside role="alert">Connections are unencrypted by default. This allows sniffing and <a href="https://en.wikipedia.org/wiki/Man-in-the-middle_attack">MitM</a> attacks.</aside>
First make sure you have a certificate and private key for TLS, usually in
`.pem` format. On your local machine you may use
[mkcert](https://github.com/FiloSottile/mkcert#readme) to generate a self signed
certificate.
To enable TLS, set the following properties in the Qdrant configuration with the
correct paths and restart:
```yaml
service:
# Enable HTTPS for the REST and gRPC API
enable_tls: true
# TLS configuration.
# Required if either service.enable_tls or cluster.p2p.enable_tls is true.
tls:
# Server certificate chain file
cert: ./tls/cert.pem
# Server private key file
key: ./tls/key.pem
```
For internal communication when running cluster mode, TLS can be enabled with:
```yaml
cluster:
# Configuration of the inter-cluster communication
p2p:
# Use TLS for communication between peers
enable_tls: true
```
With TLS enabled, you must start using HTTPS connections. For example:
```bash
curl -X GET https://localhost:6333
```
```python
from qdrant_client import QdrantClient
client = QdrantClient(
url="https://localhost:6333",
)
```
```typescript
import { QdrantClient } from "@qdrant/js-client-rest";
const client = new QdrantClient({ url: "https://localhost", port: 6333 });
```
```rust
use qdrant_client::Qdrant;
let client = Qdrant::from_url("http://localhost:6334").build()?;
```
Certificate rotation is enabled with a default refresh time of one hour. This
reloads certificate files every hour while Qdrant is running. This way changed
certificates are picked up when they get updated externally. The refresh time
can be tuned by changing the `tls.cert_ttl` setting. You can leave this on, even
if you don't plan to update your certificates. Currently this is only supported
for the REST API.
Optionally, you can enable client certificate validation on the server against a
local certificate authority. Set the following properties and restart:
```yaml
service:
# Check user HTTPS client certificate against CA file specified in tls config
verify_https_client_certificate: false
# TLS configuration.
# Required if either service.enable_tls or cluster.p2p.enable_tls is true.
tls:
# Certificate authority certificate file.
# This certificate will be used to validate the certificates
# presented by other nodes during inter-cluster communication.
#
# If verify_https_client_certificate is true, it will verify
# HTTPS client certificate
#
# Required if cluster.p2p.enable_tls is true.
ca_cert: ./tls/cacert.pem
```
## Hardening
We recommend reducing the amount of permissions granted to Qdrant containers so that you can reduce the risk of exploitation. Here are some ways to reduce the permissions of a Qdrant container:
* Run Qdrant as a non-root user. This can help mitigate the risk of future container breakout vulnerabilities. Qdrant does not need the privileges of the root user for any purpose.
- You can use the image `qdrant/qdrant:<version>-unprivileged` instead of the default Qdrant image.
- You can use the flag `--user=1000:2000` when running [`docker run`](https://docs.docker.com/reference/cli/docker/container/run/).
- You can set [`user: 1000`](https://docs.docker.com/compose/compose-file/05-services/#user) when using Docker Compose.
- You can set [`runAsUser: 1000`](https://kubernetes.io/docs/tasks/configure-pod-container/security-context) when running in Kubernetes (our [Helm chart](https://github.com/qdrant/qdrant-helm) does this by default).
* Run Qdrant with a read-only root filesystem. This can help mitigate vulnerabilities that require the ability to modify system files, which is a permission Qdrant does not need. As long as the container uses mounted volumes for storage (`/qdrant/storage` and `/qdrant/snapshots` by default), Qdrant can continue to operate while being prevented from writing data outside of those volumes.
- You can use the flag `--read-only` when running [`docker run`](https://docs.docker.com/reference/cli/docker/container/run/).
- You can set [`read_only: true`](https://docs.docker.com/compose/compose-file/05-services/#read_only) when using Docker Compose.
- You can set [`readOnlyRootFilesystem: true`](https://kubernetes.io/docs/tasks/configure-pod-container/security-context) when running in Kubernetes (our [Helm chart](https://github.com/qdrant/qdrant-helm) does this by default).
* Block Qdrant's external network access. This can help mitigate [server side request forgery attacks](https://owasp.org/www-community/attacks/Server_Side_Request_Forgery), like via the [snapshot recovery API](https://api.qdrant.tech/api-reference/snapshots/recover-from-snapshot). Single-node Qdrant clusters do not require any outbound network access. Multi-node Qdrant clusters only need the ability to connect to other Qdrant nodes via TCP ports 6333, 6334, and 6335.
- You can use [`docker network create --internal <name>`](https://docs.docker.com/reference/cli/docker/network/create/#internal) and use that network when running [`docker run --network <name>`](https://docs.docker.com/reference/cli/docker/container/run/#network).
- You can create an [internal network](https://docs.docker.com/compose/compose-file/06-networks/#internal) when using Docker Compose.
- You can create a [NetworkPolicy](https://kubernetes.io/docs/concepts/services-networking/network-policies/) when using Kubernetes. Note that multi-node Qdrant clusters [will also need access to cluster DNS in Kubernetes](https://github.com/ahmetb/kubernetes-network-policy-recipes/blob/master/11-deny-egress-traffic-from-an-application.md#allowing-dns-traffic).
There are other techniques for reducing the permissions such as dropping [Linux capabilities](https://www.man7.org/linux/man-pages/man7/capabilities.7.html) depending on your deployment method, but the methods mentioned above are the most important.
@@ -1,246 +0,0 @@
---
title: Snapshots
weight: 25
aliases:
- ../snapshots
---
# Snapshots
*Available as of v0.8.4*
Snapshots are `tar` archive files that contain data and configuration of a specific collection on a specific node at a specific time. In a distributed setup, when you have multiple nodes in your cluster, you must create snapshots for each node separately when dealing with a single collection.
This feature can be used to archive data or easily replicate an existing deployment. For disaster recovery, Qdrant Cloud users may prefer to use [Backups](/documentation/cloud/backups/) instead, which are physical disk-level copies of your data.
A collection level snapshot only contains data within that collection, including the collection configuration, all points and payloads. Collection aliases are not included and can be migrated or recovered [separately](/documentation/manage-data/collections/#collection-aliases).
For a step-by-step guide on how to use snapshots, see our [tutorial](/documentation/tutorials-operations/create-snapshot/).
## Create snapshot
<aside role="status">If you work with a distributed deployment, you have to create snapshots for each node separately. A single snapshot will contain only the data stored on the node on which the snapshot was created.</aside>
To create a new snapshot for an existing collection:
{{< code-snippet path="/documentation/headless/snippets/snapshots/create-collection-snapshot/" >}}
This is a synchronous operation for which a `tar` archive file will be generated into the `snapshot_path`.
### Delete snapshot
*Available as of v1.0.0*
{{< code-snippet path="/documentation/headless/snippets/snapshots/delete-collection-snapshot/" >}}
## List snapshot
List of snapshots for a collection:
{{< code-snippet path="/documentation/headless/snippets/snapshots/list-collection-snapshots/" >}}
## Retrieve snapshot
<aside role="status">Only available through the REST API for the time being.</aside>
To download a specified snapshot from a collection as a file:
{{< code-snippet path="/documentation/headless/snippets/snapshots/download-collection-snapshot/" >}}
## Restore snapshot
<aside role="status">Snapshots generated in one Qdrant cluster can only be restored to other Qdrant clusters that share the same minor version. For instance, a snapshot captured from a v1.4.1 cluster can only be restored to clusters running version v1.4.x, where x is equal to or greater than 1.</aside>
Snapshots can be restored in three possible ways:
1. [Recovering from a URL or local file](#recover-from-a-url-or-local-file) (useful for restoring a snapshot file that is on a remote server or already stored on the node)
3. [Recovering from an uploaded file](#recover-from-an-uploaded-file) (useful for migrating data to a new cluster)
3. [Recovering during start-up](#recover-during-start-up) (useful when running a self-hosted single-node Qdrant instance)
Regardless of the method used, Qdrant will extract the shard data from the snapshot and properly register shards in the cluster.
If there are other active replicas of the recovered shards in the cluster, Qdrant will replicate them to the newly recovered node by default to maintain data consistency.
### Recover from a URL or local file
*Available as of v0.11.3*
This method of recovery requires the snapshot file to be downloadable from a URL or exist as a local file on the node (like if you [created the snapshot](#create-snapshot) on this node previously). If instead you need to upload a snapshot file, see the next section.
To recover from a URL or local file use the [snapshot recovery endpoint](https://api.qdrant.tech/master/api-reference/snapshots/recover-from-snapshot). This endpoint accepts either a URL like `https://example.com` or a [file URI](https://en.wikipedia.org/wiki/File_URI_scheme) like `file:///tmp/snapshot-2022-10-10.snapshot`. If the target collection does not exist, it will be created.
{{< code-snippet path="/documentation/headless/snippets/snapshots/recover-collection-snapshot-from-url/" >}}
<aside role="status">When recovering from a URL, the URL must be reachable by the Qdrant node that you are restoring. In Qdrant Cloud, restoring via URL is not supported since all outbound traffic is blocked for security purposes. You may still restore via file URI or via an uploaded file.</aside>
### Recover from an uploaded file
The snapshot file can also be uploaded as a file and restored using the [recover from uploaded snapshot](https://api.qdrant.tech/master/api-reference/snapshots/recover-from-uploaded-snapshot). This endpoint accepts the raw snapshot data in the request body. If the target collection does not exist, it will be created.
```bash
curl -X POST 'http://{qdrant-url}:6333/collections/{collection_name}/snapshots/upload?priority=snapshot' \
-H 'api-key: ********' \
-H 'Content-Type:multipart/form-data' \
-F 'snapshot=@/path/to/snapshot-2022-10-10.snapshot'
```
This method is typically used to migrate data from one cluster to another, so we recommend setting the [priority](#snapshot-priority) to "snapshot" for that use-case.
### Recover during start-up
<aside role="alert">This method cannot be used in a multi-node deployment and cannot be used in Qdrant Cloud.</aside>
If you have a single-node deployment, you can recover any collection at start-up and it will be immediately available.
Restoring snapshots is done through the Qdrant CLI at start-up time via the `--snapshot` argument which accepts a list of pairs such as `<snapshot_file_path>:<target_collection_name>`
For example:
```bash
./qdrant --snapshot /snapshots/test-collection-archive.snapshot:test-collection --snapshot /snapshots/test-collection-archive.snapshot:test-copy-collection
```
The target collection **must** be absent otherwise the program will exit with an error.
If you wish instead to overwrite an existing collection, use the `--force_snapshot` flag with caution.
### Snapshot priority
When recovering a snapshot to a non-empty node, there may be conflicts between the snapshot data and the existing data. The "priority" setting controls how Qdrant handles these conflicts. The priority setting is important because different priorities can give very
different end results. The default priority may not be best for all situations.
The available snapshot recovery priorities are:
- `replica`: _(default)_ prefer existing data over the snapshot.
- `snapshot`: prefer snapshot data over existing data.
- `no_sync`: restore snapshot without any additional synchronization.
To recover a new collection from a snapshot, you need to set
the priority to `snapshot`. With `snapshot` priority, all data from the snapshot
will be recovered onto the cluster. With `replica` priority _(default)_, you'd
end up with an empty collection because the collection on the cluster did not
contain any points and that source was preferred.
`no_sync` is for specialized use cases and is not commonly used. It allows
managing shards and transferring shards between clusters manually without any
additional synchronization. Using it incorrectly will leave your cluster in a
broken state.
To recover from a URL, you specify an additional parameter in the request body:
{{< code-snippet path="/documentation/headless/snippets/snapshots/recover-snapshot-with-priority/" >}}
## Snapshots for the whole storage
*Available as of v0.8.5*
Sometimes it might be handy to create snapshot not just for a single collection, but for the whole storage, including collection aliases.
Qdrant provides a dedicated API for that as well. It is similar to collection-level snapshots, but does not require `collection_name`.
<aside role="alert">Full storage snapshots are only suitable for single-node deployments. <a href="/documentation/operations/distributed_deployment/">Distributed</a> mode is not supported as it doesn't contain the necessary files for that.</aside>
<aside role="status">Full storage snapshots can be created and downloaded from Qdrant Cloud, but you cannot restore a Qdrant Cloud cluster from a whole storage snapshot since that requires use of the Qdrant CLI. You can use <a href="/documentation/cloud/backups/">Backups</a> instead.</aside>
### Create full storage snapshot
{{< code-snippet path="/documentation/headless/snippets/snapshots/create-full-snapshot/" >}}
### Delete full storage snapshot
*Available as of v1.0.0*
{{< code-snippet path="/documentation/headless/snippets/snapshots/delete-full-snapshot/" >}}
### List full storage snapshots
{{< code-snippet path="/documentation/headless/snippets/snapshots/list-full-snapshots/" >}}
### Download full storage snapshot
<aside role="status">Only available through the REST API for the time being.</aside>
{{< code-snippet path="/documentation/headless/snippets/snapshots/download-full-snapshot/" >}}
## Restore full storage snapshot
Restoring snapshots can only be done through the Qdrant CLI at startup time.
For example:
```bash
./qdrant --storage-snapshot /snapshots/full-snapshot-2022-07-18-11-20-51.snapshot
```
## Storage
Created, uploaded and recovered snapshots are stored as `.snapshot` files. By
default, they're stored on the [local file system](#local-file-system). You may
also configure to use an [S3 storage](#s3) service for them.
### Local file system
By default, snapshots are stored at `./snapshots` or at `/qdrant/snapshots` when
using our Docker image.
The target directory can be controlled through the [configuration](/documentation/operations/configuration/):
```yaml
storage:
# Specify where you want to store snapshots.
snapshots_path: ./snapshots
```
Alternatively you may use the environment variable `QDRANT__STORAGE__SNAPSHOTS_PATH=./snapshots`.
*Available as of v1.3.0*
While a snapshot is being created, temporary files are placed in the configured
storage directory by default. In case of limited capacity or a slow
network attached disk, you can specify a separate location for temporary files:
```yaml
storage:
# Where to store temporary files
temp_path: /tmp
```
### S3
*Available as of v1.10.0*
Rather than storing snapshots on the local file system, you may also configure
to store snapshots in an S3-compatible storage service. To enable this, you must
configure it in the [configuration](/documentation/operations/configuration/) file.
For example, to configure for AWS S3:
```yaml
storage:
snapshots_config:
# Use 's3' to store snapshots on S3
snapshots_storage: s3
s3_config:
# Bucket name
bucket: your_bucket_here
# Bucket region (e.g. eu-central-1)
region: your_bucket_region_here
# Storage access key
# Can be specified either here or in the `QDRANT__STORAGE__SNAPSHOTS_CONFIG__S3_CONFIG__ACCESS_KEY` environment variable.
access_key: your_access_key_here
# Storage secret key
# Can be specified either here or in the `QDRANT__STORAGE__SNAPSHOTS_CONFIG__S3_CONFIG__SECRET_KEY` environment variable.
secret_key: your_secret_key_here
# S3-Compatible Storage URL
# Can be specified either here or in the `QDRANT__STORAGE__SNAPSHOTS_CONFIG__S3_CONFIG__ENDPOINT_URL` environment variable.
endpoint_url: your_url_here
```
Apart from Snapshots, Qdrant also provides the [Qdrant Migration Tool](https://github.com/qdrant/migration) that supports:
- Migration between Qdrant Cloud instances.
- Migrating vectors from other providers into Qdrant.
- Migrating from Qdrant OSS to Qdrant Cloud.
Follow our [migration guide](/documentation/tutorials-operations/migration/) to learn how to effectively use the Qdrant Migration tool.
@@ -1,44 +0,0 @@
---
title: Upgrades
weight: 11
---
# Upgrading Qdrant
If you are several versions behind, multiple updates might be required to reach the latest version. When upgrading Qdrant, upgrade to the latest patch version of each intermediate minor version first. For example, if you are running version 1.15 and want to upgrade to 1.17, you must first upgrade all cluster nodes to 1.16.3 before upgrading to 1.17. A Qdrant node with version 1.17 will be compatible with a node with version 1.16, but not with a node with version 1.15. If you run a single node cluster, you also can not skip versions to ensure that all data migrations are properly applied. Qdrant Cloud does this automatically for you.
If all collections of a cluster have a replication factor of at least **2**, the update process will be zero-downtime as long as you restart the Qdrant nodes in a rolling fashion. This means that you should restart the nodes one after another, allowing the cluster to maintain availability during the update process. If you have a single-node cluster or a collection with a replication factor of **1**, the update process will require a short downtime period to restart your cluster with the new version.
You need to ensure that your client applications and used SDKs are compatible with the target version.
We recommend first updating the client SDKs, and after that update the cluster to ensure a smooth update process. All client SDKs are tested to be backwards compatible with the latest 3 minor versions of Qdrant.
## Qdrant Cloud
For Qdrant Cloud, see [Cluster Upgrades](/documentation/cloud/cluster-upgrades/).
## Kubernetes
If you are using Helm, you can upgrade your Qdrant cluster by upgrading the Helm release. For example:
```bash
helm upgrade qdrant qdrant/qdrant --version <target-version> -n <namespace>
```
Kubernetes will automatically perform a rolling update of the Qdrant StatefulSet.
## Docker
If you run Qdrant using Docker, you can upgrade to a new version by pulling the new image and restarting the Qdrant containers one after another. For example:
```bash
docker pull qdrant/qdrant
docker stop qdrant
docker run --name qdrant -d -p 6333:6333 \
-v $(pwd)/path/to/data:/qdrant/storage \
qdrant/qdrant
```
## Docker Compose
If you run Qdrant using Docker Compose, you can upgrade to a new version by updating the image version in your `docker-compose.yml` file and restarting the Qdrant service.
@@ -1,82 +0,0 @@
---
title: Usage Statistics
weight: 30
aliases:
- ../telemetry
- /documentation/guides/telemetry
- /documentation/guides/usage-statistics/
---
# Usage statistics
The Qdrant open-source container image collects anonymized usage statistics from users in order to improve the engine by default. You can [deactivate](#deactivate-telemetry) at any time, and any data that has already been collected can be [deleted on request](#request-information-deletion).
Deactivating this will not affect your ability to monitor the Qdrant database yourself by accessing the `/metrics` or `/telemetry` endpoints of your database. It will just stop sending independent, anonymized usage statistics to the Qdrant team.
<aside role="status">When using Qdrant Cloud, this setting does not apply and anonymized usage statistics are disabled by default.</aside>
## Why do we collect usage statistics?
We want to make Qdrant fast and reliable. To do this, we need to understand how it performs in real-world scenarios.
We do a lot of benchmarking internally, but it is impossible to cover all possible use cases, hardware, and configurations.
In order to identify bottlenecks and improve Qdrant, we need to collect information about how it is used.
Additionally, Qdrant uses a bunch of internal heuristics to optimize the performance.
To better set up parameters for these heuristics, we need to collect timings and counters of various pieces of code.
With this information, we can make Qdrant faster for everyone.
## What information is collected?
There are 3 types of information that we collect:
* System information - general information about the system, such as CPU, RAM, and disk type. As well as the configuration of the Qdrant instance.
* Performance - information about timings and counters of various pieces of code.
* Critical error reports - information about critical errors, such as backtraces, that occurred in Qdrant. This information would allow to identify problems nobody yet reported to us.
### We **never** collect the following information:
- User's IP address
- Any data that can be used to identify the user or the user's organization
- Any data, stored in the collections
- Any names of the collections
- Any URLs
## How do we anonymize data?
We understand that some users may be concerned about the privacy of their data.
That is why we make an extra effort to ensure your privacy.
There are several different techniques that we use to anonymize the data:
- We use a random UUID to identify instances. This UUID is generated on each startup and is not stored anywhere. There are no other ways to distinguish between different instances.
- We round all big numbers, so that the last digits are always 0. For example, if the number is 123456789, we will store 123456000.
- We replace all names with irreversibly hashed values. So no collection or field names will leak into the telemetry.
- All urls are hashed as well.
You can see exact version of anomymized collected data by accessing the [telemetry API](https://api.qdrant.tech/master/api-reference/service/telemetry) with `anonymize=true` parameter.
For example, <http://localhost:6333/telemetry?details_level=6&anonymize=true>
## Deactivate usage statistics
You can deactivate usage statistics by:
- setting the `QDRANT__TELEMETRY_DISABLED` environment variable to `true`
- setting the config option `telemetry_disabled` to `true` in the `config/production.yaml` or `config/config.yaml` files
- using cli option `--disable-telemetry`
Any of these options will prevent Qdrant from sending any usage statistics data.
If you decide to deactivate usage statistics, we kindly ask you to share your feedback with us in the [Discord community](https://qdrant.to/discord) or GitHub [discussions](https://github.com/qdrant/qdrant/discussions)
## Request information deletion
We provide an email address so that users can request the complete removal of their data from all of our tools.
To do so, send an email to privacy@qdrant.com containing the unique identifier generated for your Qdrant installation.
You can find this identifier in the telemetry API response (`"id"` field), or in the logs of your Qdrant instance.
Any questions regarding the management of the data we collect can also be sent to this email address.