Files
landing_page/qdrant-landing/content/documentation/guides/monitoring.md
T
AnushandTim Visée a8c9ec2f1c docs: List hardware/memory metrics in /guides/monitoring.md (#1353)
* docs: Updated hardware/memory metrics

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* Update qdrant-landing/content/documentation/guides/monitoring.md

* Update qdrant-landing/content/documentation/guides/monitoring.md

Co-authored-by: Tim Visée <tim@visee.me>

* chore: formatting

Signed-off-by: Anush008 <anushshetty90@gmail.com>

---------

Signed-off-by: Anush008 <anushshetty90@gmail.com>
Co-authored-by: Tim Visée <tim@visee.me>
2024-12-20 15:00:13 +05:30

7.7 KiB

title, weight, aliases
title weight aliases
Monitoring & Telemetry 155
../monitoring

Monitoring & Telemetry

Qdrant exposes its metrics in Prometheus/OpenMetrics format, so you can integrate them easily with the compatible tools and monitor Qdrant with your own monitoring system. You can use the /metrics endpoint and configure it as a scrape target.

Metrics endpoint: http://localhost:6333/metrics

The integration with Qdrant is easy to configure with Prometheus and Grafana.

Monitoring multi-node clusters

When scraping metrics from multi-node Qdrant clusters, it is important to scrape from each node individually instead of using a load-balanced URL. Otherwise, your metrics will appear inconsistent after each scrape.

Monitoring in Qdrant Cloud

To scrape metrics from a Qdrant cluster running in Qdrant Cloud, note that an API key is required to access /metrics. Qdrant Cloud also supports supplying the API key as a Bearer token, which may be required by some providers.

Exposed metrics

Each Qdrant server will expose the following metrics.

Name Type Meaning
app_info gauge Information about Qdrant server
app_status_recovery_mode gauge If Qdrant is currently started in recovery mode
collections_total gauge Number of collections
collections_vector_total gauge Total number of vectors in all collections
collections_full_total gauge Number of full collections
collections_aggregated_total gauge Number of aggregated collections
rest_responses_total counter Total number of responses through REST API
rest_responses_fail_total counter Total number of failed responses through REST API
rest_responses_avg_duration_seconds gauge Average response duration in REST API
rest_responses_min_duration_seconds gauge Minimum response duration in REST API
rest_responses_max_duration_seconds gauge Maximum response duration in REST API
grpc_responses_total counter Total number of responses through gRPC API
grpc_responses_fail_total counter Total number of failed responses through REST API
grpc_responses_avg_duration_seconds gauge Average response duration in gRPC API
grpc_responses_min_duration_seconds gauge Minimum response duration in gRPC API
grpc_responses_max_duration_seconds gauge Maximum response duration in gRPC API
cluster_enabled gauge Whether the cluster support is enabled. 1 - YES
memory_active_bytes gauge Total number of bytes in active pages allocated by the application (jemalloc)
memory_allocated_bytes gauge Total number of bytes allocated by the application (jemalloc)
memory_metadata_bytes gauge Total number of bytes dedicated to allocator metadata (jemalloc)
memory_resident_bytes gauge Maximum number of bytes in physically resident data pages mapped (resident)
memory_retained_bytes gauge Total number of bytes in virtual memory mappings (retained)
collection_hardware_metric_cpu gauge CPU measurements of a collection

There are also some metrics which are exposed in distributed mode only.

Name Type Meaning
cluster_peers_total gauge Total number of cluster peers
cluster_term counter Current cluster term
cluster_commit counter Index of last committed (finalized) operation cluster peer is aware of
cluster_pending_operations_total gauge Total number of pending operations for cluster peer
cluster_voter gauge Whether the cluster peer is a voter or learner. 1 - VOTER

Telemetry endpoint

Qdrant also provides a /telemetry endpoint, which provides information about the current state of the database, including the number of vectors, shards, and other useful information. You can find a full documentation of this endpoint in the API reference.

Kubernetes health endpoints

Available as of v1.5.0

Qdrant exposes three endpoints, namely /healthz, /livez and /readyz, to indicate the current status of the Qdrant server.

These currently provide the most basic status response, returning HTTP 200 if Qdrant is started and ready to be used.

Regardless of whether an API key is configured, the endpoints are always accessible.

You can read more about Kubernetes health endpoints here.