diff --git a/qdrant-landing/content/documentation/private-cloud/api-reference.md b/qdrant-landing/content/documentation/private-cloud/api-reference.md index f3df44d84..15a2808ff 100644 --- a/qdrant-landing/content/documentation/private-cloud/api-reference.md +++ b/qdrant-landing/content/documentation/private-cloud/api-reference.md @@ -57,6 +57,26 @@ _Appears in:_ +#### ComponentReference + + + + + + + +_Appears in:_ +- [QdrantCloudRegionSpec](#qdrantcloudregionspec) + +| Field | Description | Default | Validation | +| --- | --- | --- | --- | +| `apiVersion` _string_ | APIVersion is the group and version of the component being referenced. | | | +| `kind` _string_ | Kind is the type of component being referenced | | | +| `name` _string_ | Name is the name of component being referenced | | | +| `namespace` _string_ | Namespace is the namespace of component being referenced. | | | +| `markedForDeletion` _boolean_ | MarkedForDeletion specifies whether the component is marked for deletion | | | + + #### ComponentStatus @@ -77,6 +97,35 @@ _Appears in:_ | `message` _string_ | Message specifies the info explaining the current phase of the component | | | +#### GPU + + + + + + + +_Appears in:_ +- [QdrantClusterSpec](#qdrantclusterspec) + +| Field | Description | Default | Validation | +| --- | --- | --- | --- | +| `gpuType` _[GPUType](#gputype)_ | GPUType specifies the type of the GPU to use. | | Enum: [nvidia amd]
| + + +#### GPUType + +_Underlying type:_ _string_ + + + + + +_Appears in:_ +- [GPU](#gpu) + + + #### HelmRelease @@ -111,6 +160,22 @@ _Appears in:_ | `object` _[HelmRepository](#helmrepository)_ | Object specifies the helm repository object | | EmbeddedResource: {}
| +#### InferenceConfig + + + + + + + +_Appears in:_ +- [QdrantConfiguration](#qdrantconfiguration) + +| Field | Description | Default | Validation | +| --- | --- | --- | --- | +| `enabled` _boolean_ | Enabled specifies whether to enable inference for the cluster or not. | false | | + + #### Ingress @@ -246,6 +311,47 @@ _Appears in:_ | `grpcHost` _string_ | GRPCHost specifies the host name for the GRPC ingress. | | | +#### NodeInfo + + + + + + + +_Appears in:_ +- [QdrantCloudRegionStatus](#qdrantcloudregionstatus) + +| Field | Description | Default | Validation | +| --- | --- | --- | --- | +| `name` _string_ | Name specifies the name of the node | | | +| `region` _string_ | Region specifies the region of the node | | | +| `zone` _string_ | Zone specifies the zone of the node | | | +| `instanceType` _string_ | InstanceType specifies the instance type of the node | | | +| `arch` _string_ | Arch specifies the CPU architecture of the node | | | +| `capacity` _[NodeResourceInfo](#noderesourceinfo)_ | Capacity specifies the capacity of the node | | | +| `allocatable` _[NodeResourceInfo](#noderesourceinfo)_ | Allocatable specifies the allocatable resources of the node | | | + + +#### NodeResourceInfo + + + + + + + +_Appears in:_ +- [NodeInfo](#nodeinfo) + +| Field | Description | Default | Validation | +| --- | --- | --- | --- | +| `cpu` _string_ | CPU specifies the CPU resources of the node | | | +| `memory` _string_ | Memory specifies the memory resources of the node | | | +| `pods` _string_ | Pods specifies the pods resources of the node | | | +| `ephemeralStorage` _string_ | EphemeralStorage specifies the ephemeral storage resources of the node | | | + + #### NodeStatus @@ -265,74 +371,6 @@ _Appears in:_ | `version` _string_ | Version specifies the version of Qdrant running on the node | | | -#### Operation - - - - - - - -_Appears in:_ -- [QdrantClusterStatus](#qdrantclusterstatus) - -| Field | Description | Default | Validation | -| --- | --- | --- | --- | -| `type` _[OperationType](#operationtype)_ | Type specifies the type of the operation | | | -| `phase` _[OperationPhase](#operationphase)_ | Phase specifies the phase of the operation | | | -| `id` _integer_ | Id specifies the id of the operation | | | -| `startTime` _string_ | StartTime specifies the time when the operation started | | | -| `completionTime` _string_ | CompletionTime specifies the time when the operation completed | | | -| `message` _string_ | Message specifies the message of the operation | | | -| `subOperation` _boolean_ | SubOperation specifies whether the operation is a sub-operation of another operation | | | -| `steps` _[OperationStep](#operationstep) array_ | Steps specifies the steps the operation has performed | | | - - -#### OperationPhase - -_Underlying type:_ _string_ - - - - - -_Appears in:_ -- [Operation](#operation) - - - -#### OperationStep - - - - - - - -_Appears in:_ -- [Operation](#operation) - -| Field | Description | Default | Validation | -| --- | --- | --- | --- | -| `name` _string_ | Name specifies the name of the step | | | -| `id` _integer_ | Id specifies the id of the step | | | -| `phase` _[StepPhase](#stepphase)_ | Phase specifies the phase of the step | | | -| `message` _string_ | Message specifies the reason in case of failure | | | - - -#### OperationType - -_Underlying type:_ _string_ - - - - - -_Appears in:_ -- [Operation](#operation) - - - #### Pause @@ -402,8 +440,9 @@ _Appears in:_ | Field | Description | Default | Validation | | --- | --- | --- | --- | | `id` _string_ | Id specifies the unique identifier of the region | | | -| `helmRepositories` _[HelmRepository](#helmrepository) array_ | HelmRepositories specifies the list of helm repositories to be created to the region | | | -| `helmReleases` _[HelmRelease](#helmrelease) array_ | HelmReleases specifies the list of helm releases to be created to the region | | | +| `components` _[ComponentReference](#componentreference) array_ | Components specifies the list of components to be installed in the region | | | +| `helmRepositories` _[HelmRepository](#helmrepository) array_ | HelmRepositories specifies the list of helm repositories to be created to the region
Deprecated: Use "Components" instead | | | +| `helmReleases` _[HelmRelease](#helmrelease) array_ | HelmReleases specifies the list of helm releases to be created to the region
Deprecated: Use "Components" instead | | | @@ -650,7 +689,6 @@ _Appears in:_ | `clusterManager` _boolean_ | ClusterManager specifies whether to use the cluster manager for this cluster.
The Python-operator will deploy a dedicated cluster manager instance.
The Go-operator will use a shared instance.
If not set, the default will be taken from the operator config. | | | | `suspend` _boolean_ | Suspend specifies whether to suspend the cluster.
If enabled, the cluster will be suspended and all related resources will be removed except the PVCs. | false | | | `pauses` _[Pause](#pause) array_ | Pauses specifies a list of pause request by developer for manual maintenance.
Operator will skip handling any changes in the CR if any pause request is present. | | | -| `distributed` _boolean_ | Deprecated | | | | `image` _[QdrantImage](#qdrantimage)_ | Image specifies the image to use for each Qdrant node. | | | | `resources` _[Resources](#resources)_ | Resources specifies the resources to allocate for each Qdrant node. | | | | `security` _[QdrantSecurityContext](#qdrantsecuritycontext)_ | Security specifies the security context for each Qdrant node. | | | @@ -659,11 +697,13 @@ _Appears in:_ | `config` _[QdrantConfiguration](#qdrantconfiguration)_ | Config specifies the Qdrant configuration setttings for the clusters. | | | | `ingress` _[Ingress](#ingress)_ | Ingress specifies the ingress for the cluster. | | | | `service` _[KubernetesService](#kubernetesservice)_ | Service specifies the configuration of the Qdrant Kubernetes Service. | | | +| `gpu` _[GPU](#gpu)_ | GPU specifies GPU configuration for the cluster. If this field is not set, no GPU will be used. | | | | `statefulSet` _[KubernetesStatefulSet](#kubernetesstatefulset)_ | StatefulSet specifies the configuration of the Qdrant Kubernetes StatefulSet. | | | | `storageClassNames` _[StorageClassNames](#storageclassnames)_ | StorageClassNames specifies the storage class names for db and snapshots. | | | | `topologySpreadConstraints` _[TopologySpreadConstraint](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#topologyspreadconstraint-v1-core)_ | TopologySpreadConstraints specifies the topology spread constraints for the cluster. | | | | `podDisruptionBudget` _[PodDisruptionBudgetSpec](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#poddisruptionbudgetspec-v1-policy)_ | PodDisruptionBudget specifies the pod disruption budget for the cluster. | | | | `restartAllPodsConcurrently` _boolean_ | RestartAllPodsConcurrently specifies whether to restart all pods concurrently (also called one-shot-restart).
If enabled, all the pods in the cluster will be restarted concurrently in situations where multiple pods
need to be restarted like when RestartedAtAnnotationKey is added/updated or the Qdrant version need to be upgraded.
This helps sharded but not replicated clusters to reduce downtime to possible minimum during restart. | | | +| `startupDelaySeconds` _integer_ | If StartupDelaySeconds is set (> 0), an additional 'sleep ' will be emitted to the pod startup.
The sleep will be added when a pod is restarted, it will not force any pod to restart.
This feature can be used for debugging the core, e.g. if a pod is in crash loop, it provided a way
to inspect the attached storage. | | | @@ -686,6 +726,7 @@ _Appears in:_ | `service` _[QdrantConfigurationService](#qdrantconfigurationservice)_ | Service specifies the service level configuration for Qdrant. | | | | `tls` _[QdrantConfigurationTLS](#qdrantconfigurationtls)_ | TLS specifies the TLS configuration for Qdrant. | | | | `storage` _[StorageConfig](#storageconfig)_ | Storage specifies the storage configuration for Qdrant. | | | +| `inference` _[InferenceConfig](#inferenceconfig)_ | Inference configuration. This is used in Qdrant Managed Cloud only. If not set Inference is not available to this cluster. | | | #### QdrantConfigurationCollection @@ -757,6 +798,7 @@ _Appears in:_ | --- | --- | --- | --- | | `cert` _[QdrantSecretKeyRef](#qdrantsecretkeyref)_ | Reference to the secret containing the server certificate chain file | | | | `key` _[QdrantSecretKeyRef](#qdrantsecretkeyref)_ | Reference to the secret containing the server private key file | | | +| `caCert` _[QdrantSecretKeyRef](#qdrantsecretkeyref)_ | Reference to the secret containing the CA certificate file | | | #### QdrantImage @@ -999,17 +1041,25 @@ _Appears in:_ -#### StepPhase +#### StorageClass + -_Underlying type:_ _string_ _Appears in:_ -- [OperationStep](#operationstep) +- [QdrantCloudRegionStatus](#qdrantcloudregionstatus) +| Field | Description | Default | Validation | +| --- | --- | --- | --- | +| `name` _string_ | Name specifies the name of the storage class | | | +| `default` _boolean_ | Default specifies whether the storage class is the default storage class | | | +| `provisioner` _string_ | Provisioner specifies the provisioner of the storage class | | | +| `allowVolumeExpansion` _boolean_ | AllowVolumeExpansion specifies whether the storage class allows volume expansion | | | +| `reclaimPolicy` _string_ | ReclaimPolicy specifies the reclaim policy of the storage class | | | +| `parameters` _object (keys:string, values:string)_ | Parameters specifies the parameters of the storage class | | | #### StorageClassNames @@ -1078,6 +1128,23 @@ _Appears in:_ | `allowedSourceRanges` _string array_ | AllowedSourceRanges specifies the allowed CIDR source ranges for the ingress. | | | +#### VolumeSnapshotClass + + + + + + + +_Appears in:_ +- [QdrantCloudRegionStatus](#qdrantcloudregionstatus) + +| Field | Description | Default | Validation | +| --- | --- | --- | --- | +| `name` _string_ | Name specifies the name of the volume snapshot class | | | +| `driver` _string_ | Driver specifies the driver of the volume snapshot class | | | + + #### VolumeSnapshotInfo diff --git a/qdrant-landing/content/documentation/private-cloud/changelog.md b/qdrant-landing/content/documentation/private-cloud/changelog.md index 9639994d3..9039a42ef 100644 --- a/qdrant-landing/content/documentation/private-cloud/changelog.md +++ b/qdrant-landing/content/documentation/private-cloud/changelog.md @@ -5,6 +5,11 @@ weight: 5 # Changelog +## 1.6.0 (2025-03-07) + +* Add support for GPU instances +* Experimental support for automatic shard balancing + ## 1.5.1 (2025-03-04) * Fix scaling down clusters that have TLS with self-signed certificates configured diff --git a/qdrant-landing/content/documentation/private-cloud/configuration.md b/qdrant-landing/content/documentation/private-cloud/configuration.md index 4688978aa..9380d0732 100644 --- a/qdrant-landing/content/documentation/private-cloud/configuration.md +++ b/qdrant-landing/content/documentation/private-cloud/configuration.md @@ -91,7 +91,7 @@ operator: appEnvironment: kubernetes # The log level for the operator # Available options: DEBUG | INFO | WARN | ERROR - logLevel: INFO + logLevel: INFO # Metrics contains the operator config related the metrics metrics: # The port used for metrics @@ -103,7 +103,7 @@ operator: # Controller related settings controller: # The period a forced recync is done by the controller (if watches are missed / nothing happened) - forceResyncPeriod: 10h + forceResyncPeriod: 2m # QPS indicates the maximum QPS to the master from this client. # Default is 200 qps: 200 @@ -130,7 +130,7 @@ operator: # Qdrant config contains settings specific for the database qdrant: # The config where to find the image for qdrant - image: + image: # The repository where to find the image for qdrant # Default is "qdrant/qdrant" repository: registry.cloud.qdrant.io/qdrant/qdrant @@ -154,7 +154,7 @@ operator: # See: asyncScorer: false # Qdrant DB log level - # Available options: DEBUG | INFO | WARN | ERROR + # Available options: DEBUG | INFO | WARN | ERROR # Default is "INFO" logLevel: INFO # Default Qdrant security context configuration @@ -199,8 +199,7 @@ operator: # Default is false. enable: true # The endpoint address the cluster manager could be reached - # If set, this should be a full URL like: http://cluster-manager.qdrant-cloud-ns.svc.cluster.local:7333 - endpointAddress: http://cluster-manager + endpointAddress: "http://qdrant-cluster-manager" # InvocationInterval is the interval between calls (started after the previous call is retured) # Default is 10 seconds invocationInterval: 10s @@ -209,7 +208,7 @@ operator: timeout: 30s # Specifies overrides for the manage rules manageRulesOverrides: - #dry_run: + #dry_run: #max_transfers: #max_transfers_per_collection: #rebalance: @@ -248,7 +247,7 @@ operator: # Default is 3 seconds telemetryTimeout: 3s # MaxConcurrentReconciles is the maximum number of concurrent Reconciles which can be run. Defaults to 20. - maxConcurrentReconciles: 20 + maxConcurrentReconciles: 20 # VolumeExpansionMode specifies the expansion mode, which can be online or offline (e.g. in case of Azure). # Available options: Online, Offline # Default is Online @@ -277,13 +276,17 @@ operator: # Whether or not the ScheduledSnapshot feature is enabled. # Default is true. enable: true + # RemoveCronJobs can be enabled when the previous [Python] operator (qdrant-operator) has been run and this + # operator should remove the cron jobs it created (not used by this operator anymore). + # Default is true. + removeCronJobs: true # MaxConcurrentReconciles is the maximum number of concurrent Reconciles which can be run. Defaults to 1. maxConcurrentReconciles: 1 # Restores contains the settings for restoring (a snapshot) as part of backup management. restores: # Whether or not the Restore feature is enabled. # Default is true. - enable: true + enable: true # MaxConcurrentReconciles is the maximum number of concurrent Reconciles which can be run. Defaults to 1. maxConcurrentReconciles: 1 diff --git a/qdrant-landing/content/documentation/private-cloud/private-cloud-setup.md b/qdrant-landing/content/documentation/private-cloud/private-cloud-setup.md index 74d7b9068..8371f4381 100644 --- a/qdrant-landing/content/documentation/private-cloud/private-cloud-setup.md +++ b/qdrant-landing/content/documentation/private-cloud/private-cloud-setup.md @@ -80,8 +80,8 @@ Once you are onboarded to Qdrant Private Cloud, you will receive credentials to kubectl create namespace qdrant-private-cloud kubectl create secret docker-registry qdrant-registry-creds --docker-server=registry.cloud.qdrant.io --docker-username='your-username' --docker-password='your-password' --namespace qdrant-private-cloud helm registry login 'registry.cloud.qdrant.io' --username 'your-username' --password 'your-password' -helm upgrade --install qdrant-private-cloud-crds oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-kubernetes-api --namespace qdrant-private-cloud --version v1.12.0 --wait -helm upgrade --install qdrant-private-cloud oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-private-cloud --namespace qdrant-private-cloud --version 1.5.1 +helm upgrade --install qdrant-private-cloud-crds oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-kubernetes-api --namespace qdrant-private-cloud --version v1.14.2 --wait +helm upgrade --install qdrant-private-cloud oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-private-cloud --namespace qdrant-private-cloud --version 1.6.0 ``` For a list of available versions consult the [Private Cloud Changelog](/documentation/private-cloud/changelog/). diff --git a/qdrant-landing/content/documentation/private-cloud/qdrant-cluster-management.md b/qdrant-landing/content/documentation/private-cloud/qdrant-cluster-management.md index 7d7b0c948..cc6b0ee9d 100644 --- a/qdrant-landing/content/documentation/private-cloud/qdrant-cluster-management.md +++ b/qdrant-landing/content/documentation/private-cloud/qdrant-cluster-management.md @@ -287,4 +287,39 @@ step certificate create mydomain.com qdrant-nodes.crt qdrant-nodes.key \ --san qdrant-my-cluster-0.qdrant-headless-my-cluster \ --san qdrant-my-cluster-1.qdrant-headless-my-cluster ``` - \ No newline at end of file + + +## Creating a cluster with GPU support + +Starting with Qdrant 1.13 you can create a cluster that uses GPUs to accelarate indexing. Starting with private-cloud version 1.6.0 you can make use of this in private cloud. + +As a prerequisite, you need to have a Kubernetes cluster with GPU support. You can check the [Kubernetes documentation](https://kubernetes.io/docs/tasks/manage-gpus/scheduling-gpus/) for generic information on GPUs and Kubernetes, or the documentation of your specific Kubernetes distribution. + +Examples: + +* [AWS EKS GPU support](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/amazon-eks.html) +* [Azure AKS GPU support](https://docs.microsoft.com/en-us/azure/aks/gpu-cluster) +* [GCP GKE GPU support](https://cloud.google.com/kubernetes-engine/docs/how-to/gpus) +* [Vultr Kubernetes GPU support](https://blogs.vultr.com/whats-new-vultr-q2-2023) + +Once you have a Kubernetes cluster with GPU support, you can create a QdrantCluster with GPU support: + +```yaml +apiVersion: qdrant.io/v1 +kind: QdrantCluster +metadata: + name: qdrant-a7d8d973-0cc5-42de-8d7b-c29d14d24840 + labels: + cluster-id: "a7d8d973-0cc5-42de-8d7b-c29d14d24840" + customer-id: "acme-industries" +spec: + id: "a7d8d973-0cc5-42de-8d7b-c29d14d24840" + version: "v1.13.4" + size: 1 + resources: + cpu: 2 + memory: "8Gi" + storage: "40Gi" + gpu: + gpuType: "nvidia" +```