mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-05 19:08:32 +02:00
Merge pull request #1995 from qdrant/qdrant-1.16-metrics
[1.16] Metrics documentation
This commit is contained in:
@@ -17,65 +17,127 @@ The integration with Qdrant is easy to
|
||||
[configure](https://prometheus.io/docs/prometheus/latest/getting_started/#configure-prometheus-to-monitor-the-sample-targets)
|
||||
with Prometheus and Grafana.
|
||||
|
||||
## Monitoring multi-node clusters
|
||||
## Metrics
|
||||
|
||||
When scraping metrics from multi-node Qdrant clusters, it is important to scrape from
|
||||
each node individually instead of using a load-balanced URL. Otherwise, your metrics will appear inconsistent after each scrape.
|
||||
Qdrant exposes various metrics in Prometheus/OpenMetrics format, commonly used together with Grafana for monitoring.
|
||||
|
||||
## Monitoring in Qdrant Cloud
|
||||
Two endpoints are available:
|
||||
|
||||
Qdrant Cloud offers additional metrics and telemetry that are not available in the open-source version. For more information, see [Qdrant Cloud Monitoring](/documentation/cloud/cluster-monitoring/).
|
||||
- `/metrics` for metrics of a Qdrant node/peer, see [all metrics](#node-metrics-metrics).
|
||||
|
||||
## Exposed metrics
|
||||
|
||||
There are two endpoints avaliable:
|
||||
|
||||
- `/metrics` is the direct endpoint of the underlying Qdrant database node.
|
||||
|
||||
- `/sys_metrics` is a Qdrant cloud-only endpoint that provides additional operational and infrastructure metrics about your cluster, like CPU, memory and disk utilisation, collection metrics and load balancer telemetry. For more information, see [Qdrant Cloud Monitoring](/documentation/cloud/cluster-monitoring/).
|
||||
- `/sys_metrics` (Qdrant Cloud only) for metrics about your cluster, like CPU, memory, disk utilisation, collection metrics and load balancer telemetry. For more information, see [Qdrant Cloud Monitoring](/documentation/cloud/cluster-monitoring/).
|
||||
|
||||
Note that `/metrics` only reports metrics for the peer connected to. It is therefore important to scrape from each peer individually, even if a load balancer is involved.
|
||||
|
||||
### Node metrics `/metrics`
|
||||
|
||||
Each Qdrant server will expose the following metrics.
|
||||
Each Qdrant node will expose the following metrics.
|
||||
|
||||
| Name | Type | Meaning |
|
||||
| ----------------------------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| app_info | gauge | Information about Qdrant server |
|
||||
| app_status_recovery_mode | gauge | If Qdrant is currently started in recovery mode |
|
||||
| collections_total | gauge | Number of collections |
|
||||
| collections_vector_total | gauge | Total number of vectors in all collections |
|
||||
| collections_full_total | gauge | Number of full collections |
|
||||
| collections_aggregated_total | gauge | Number of aggregated collections |
|
||||
| rest_responses_total | counter | Total number of responses through REST API |
|
||||
| rest_responses_fail_total | counter | Total number of failed responses through REST API |
|
||||
| rest_responses_avg_duration_seconds | gauge | Average response duration in REST API |
|
||||
| rest_responses_min_duration_seconds | gauge | Minimum response duration in REST API |
|
||||
| rest_responses_max_duration_seconds | gauge | Maximum response duration in REST API |
|
||||
| grpc_responses_total | counter | Total number of responses through gRPC API |
|
||||
| grpc_responses_fail_total | counter | Total number of failed responses through REST API |
|
||||
| grpc_responses_avg_duration_seconds | gauge | Average response duration in gRPC API |
|
||||
| grpc_responses_min_duration_seconds | gauge | Minimum response duration in gRPC API |
|
||||
| grpc_responses_max_duration_seconds | gauge | Maximum response duration in gRPC API |
|
||||
| cluster_enabled | gauge | Whether the cluster support is enabled. 1 - YES |
|
||||
| memory_active_bytes | gauge | Total number of bytes in active pages allocated by the application. [Reference](https://jemalloc.net/jemalloc.3.html#stats.active) |
|
||||
| memory_allocated_bytes | gauge | Total number of bytes allocated by the application. [Reference](https://jemalloc.net/jemalloc.3.html#stats.allocated) |
|
||||
| memory_metadata_bytes | gauge | Total number of bytes dedicated to allocator metadata. [Reference](https://jemalloc.net/jemalloc.3.html#stats.metadata) |
|
||||
| memory_resident_bytes | gauge | Maximum number of bytes in physically resident data pages mapped. [Reference](https://jemalloc.net/jemalloc.3.html#stats.resident) |
|
||||
| memory_retained_bytes | gauge | Total number of bytes in virtual memory mappings. [Reference](https://jemalloc.net/jemalloc.3.html#stats.retained) |
|
||||
| collection_hardware_metric_cpu | gauge | CPU measurements of a collection (Experimental) |
|
||||
Counters - such as the number of created snapshots - are reset when the node is restarted.
|
||||
|
||||
**Cluster-related metrics**
|
||||
**Application metrics**
|
||||
|
||||
There are also some metrics which are exposed in distributed mode only.
|
||||
| Name | Type | Meaning |
|
||||
| ----------------------------------- | ------- | ------------------------------ |
|
||||
| app_info | gauge | Qdrant server name and version |
|
||||
| app_status_recovery_mode | gauge | If started in recovery mode |
|
||||
|
||||
| Name | Type | Meaning |
|
||||
| -------------------------------- | ------- | ---------------------------------------------------------------------- |
|
||||
| cluster_peers_total | gauge | Total number of cluster peers |
|
||||
| cluster_term | counter | Current cluster term |
|
||||
| cluster_commit | counter | Index of last committed (finalized) operation cluster peer is aware of |
|
||||
| cluster_pending_operations_total | gauge | Total number of pending operations for cluster peer |
|
||||
| cluster_voter | gauge | Whether the cluster peer is a voter or learner. 1 - VOTER |
|
||||
**Collection metrics**
|
||||
|
||||
| Name | Type | Meaning |
|
||||
| ------------------------------------------------- | ------- | ----------------------------------------------------------------------------------------------------- |
|
||||
| collections_total | gauge | Number of collections |
|
||||
| collection_points | gauge | Number of points, per collection <sup>(v1.16+)</sup> |
|
||||
| collection_vectors | gauge | Number of vectors, per collection and vector name <sup>(v1.16+)</sup> |
|
||||
| collections_vector_total | gauge | Number of vectors in all collections |
|
||||
| collection_indexed_only_excluded_points | gauge | Number of points excluded in [`indexed_only`](/documentation/concepts/search/#search-api) search, per collection and vector name <sup>(v1.16+)</sup> |
|
||||
| collection_active_replicas_min | gauge | Minimum number of active replicas across all collections and shards <sup>(v1.16+)</sup> |
|
||||
| collection_active_replicas_max | gauge | Maximum number of active replicas across all collections and shards <sup>(v1.16+)</sup> |
|
||||
| collection_dead_replicas | gauge | Number of non-active replicas across all collections and shards <sup>(v1.16+)</sup> |
|
||||
| collection_running_optimizations | gauge | Number of running optimization tasks, per collection <sup>(v1.16+)</sup> |
|
||||
| collection_hardware_metric_cpu | counter | CPU measurements of a collection, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
|
||||
| collection_hardware_metric_payload_io_read | counter | Payload IO read operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
|
||||
| collection_hardware_metric_payload_io_write | counter | Payload IO write operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
|
||||
| collection_hardware_metric_payload_index_io_read | counter | Payload index read operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
|
||||
| collection_hardware_metric_payload_index_io_write | counter | Payload index write operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
|
||||
| collection_hardware_metric_vector_io_read | counter | Vector IO read operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
|
||||
| collection_hardware_metric_vector_io_write | counter | Vector IO write operations measurement, per collection <sup>(v1.13+)</sup> [^metrics-hwreporting] |
|
||||
|
||||
[^metrics-hwreporting]: Only reported if hardware metrics are enabled in the configuration. See `service.hardware_reporting` in the [configuration](/documentation/guides/configuration/).
|
||||
|
||||
**Snapshot metrics**
|
||||
|
||||
| Name | Type | Meaning |
|
||||
| --------------------------------------- | ------- | --------------------------------------------------------------------------- |
|
||||
| snapshot_creation_running | gauge | Number of snapshots being created, per collection <sup>(v1.16+)</sup> |
|
||||
| snapshot_recovery_running | gauge | Number of snapshots being recovered, per collection <sup>(v1.16+)</sup> |
|
||||
| snapshot_created_total | counter | Number of created snapshots since start, per collection <sup>(v1.16+)</sup> |
|
||||
|
||||
**API response metrics**
|
||||
|
||||
| Name | Type | Meaning |
|
||||
| ----------------------------------- | --------- | ------------------------------------------------------------------ |
|
||||
| rest_responses_total | counter | Number of responses through REST API |
|
||||
| rest_responses_fail_total | counter | Number of failed responses through REST API |
|
||||
| rest_responses_avg_duration_seconds | gauge | Average response duration in REST API |
|
||||
| rest_responses_min_duration_seconds | gauge | Minimum response duration in REST API |
|
||||
| rest_responses_max_duration_seconds | gauge | Maximum response duration in REST API |
|
||||
| rest_responses_duration_seconds | histogram | Histogram of response durations in the REST API <sup>(v1.8+)</sup> |
|
||||
| grpc_responses_total | counter | Number of responses through gRPC API |
|
||||
| grpc_responses_fail_total | counter | Number of failed responses through REST API |
|
||||
| grpc_responses_avg_duration_seconds | gauge | Average response duration in gRPC API |
|
||||
| grpc_responses_min_duration_seconds | gauge | Minimum response duration in gRPC API |
|
||||
| grpc_responses_max_duration_seconds | gauge | Maximum response duration in gRPC API |
|
||||
| grpc_responses_duration_seconds | histogram | Histogram of response durations in the gRPC API <sup>(v1.8+)</sup> |
|
||||
|
||||
**Process metrics**
|
||||
|
||||
| Name | Type | Meaning |
|
||||
| ----------------------------------- | ------- | ----------------------------------------------------------------------------------------------------------------------------- |
|
||||
| memory_active_bytes | gauge | Total number of bytes in active pages allocated by the application ([ref](https://jemalloc.net/jemalloc.3.html#stats.active)) |
|
||||
| memory_allocated_bytes | gauge | Total number of bytes allocated by the application ([ref](https://jemalloc.net/jemalloc.3.html#stats.allocated)) |
|
||||
| memory_metadata_bytes | gauge | Total number of bytes dedicated to allocator metadata ([ref](https://jemalloc.net/jemalloc.3.html#stats.metadata)) |
|
||||
| memory_resident_bytes | gauge | Maximum number of bytes in physically resident data pages mapped ([ref](https://jemalloc.net/jemalloc.3.html#stats.resident)) |
|
||||
| memory_retained_bytes | gauge | Total number of bytes in virtual memory mappings ([ref](https://jemalloc.net/jemalloc.3.html#stats.retained)) |
|
||||
| process_threads | gauge | Number of used system threads <sup>(v1.16+)</sup> |
|
||||
| process_open_mmaps | gauge | Number of open memory maps <sup>(v1.16+)</sup> |
|
||||
| system_max_mmaps | gauge | System wide maximum number of open memory maps <sup>(v1.16+)</sup> |
|
||||
| process_open_fds | gauge | Number of open file descriptors <sup>(v1.16+)</sup> |
|
||||
| process_max_fds | gauge | Maximum number of open file descriptors <sup>(v1.16+)</sup> |
|
||||
| process_minor_page_faults_total | counter | Number of minor page faults encountered by the process <sup>(v1.16+)</sup> |
|
||||
| process_major_page_faults_total | counter | Number of major page faults encountered by the process <sup>(v1.16+)</sup> |
|
||||
|
||||
**Cluster metrics (consensus)**
|
||||
|
||||
Metrics reporting the current cluster consensus state of the node. Exposed only
|
||||
when distributed mode is enabled.
|
||||
|
||||
| Name | Type | Meaning |
|
||||
| -------------------------------- | ------- | ----------------------------------------------------------------------- |
|
||||
| cluster_enabled | gauge | If distributed mode is enabled [^metrics-distributed] |
|
||||
| cluster_peers_total | gauge | Number of cluster peers [^metrics-distributed] |
|
||||
| cluster_term | counter | Raft consensus term [^metrics-distributed] |
|
||||
| cluster_commit | counter | Raft consensus commit - last committed operation [^metrics-distributed] |
|
||||
| cluster_pending_operations_total | gauge | Number of pending consensus operations [^metrics-distributed] |
|
||||
| cluster_voter | gauge | If a consensus voter (`1`) or learner (`0`) [^metrics-distributed] |
|
||||
|
||||
[^metrics-distributed]: Only reported if distributed mode (cluster mode) is enabled. Enabled by default in all Qdrant Cloud environments. See `cluster.enabled` in the [configuration](/documentation/guides/configuration/).
|
||||
|
||||
### Metrics configuration
|
||||
|
||||
*Available as of v1.16.0*
|
||||
|
||||
In self-hosted environments you have further configuration options for metrics.
|
||||
|
||||
By default, all Qdrant metrics have no application namespace prefix. You may set
|
||||
a prefix with `service.metrics_prefix` in the
|
||||
[configuration](/documentation/guides/configuration/).
|
||||
|
||||
To achieve this you may use the following environment variable for example:
|
||||
|
||||
```bash
|
||||
QDRANT__SERVICE__METRICS_PREFIX="qdrant_"
|
||||
```
|
||||
|
||||
## Telemetry endpoint
|
||||
|
||||
|
||||
Reference in New Issue
Block a user