Merge remote-tracking branch 'upstream' into v1.16-release-blog

This commit is contained in:
Abdon Pijpelink
2025-11-17 09:14:30 +01:00
12 changed files with 289 additions and 10 deletions
@@ -0,0 +1,83 @@
---
draft: false
title: "How Dragonfruit AI scaled real-time computer vision with Qdrant"
short_description: "Dragonfruit AI scales AI video analytics across thousands of cameras."
description: "Discover how Dragonfruit AI leveraged Qdrant’s per-collection configurability and float16 optimizations to achieve real-time, multi-camera computer vision analytics at enterprise scale."
preview_image: /blog/case-study-dragonfruit/social_preview_partnership-dragonfruit.jpg
social_preview_image: /blog/case-study-dragonfruit/social_preview_partnership-dragonfruit.jpg
date: 2025-11-13
author: "Daniel Azoulai"
featured: true
tags:
- Dragonfruit AI
- vector search
- computer vision
- real-time analytics
- Split AI
- float16 optimization
- case study
---
![Dragonfruit Overview](/blog/case-study-dragonfruit/dragonfruit-bento-box-dark.png)
## Dragonfruit AI scales real-time computer vision with Qdrant
### Building enterprise-ready computer vision
<a href="https://www.dragonfruit.ai/" target="_blank">Dragonfruit AI</a> builds enterprise-ready computer vision solutions, turning ordinary IP camera feeds into actionable insights for security, safety, operations, and compliance. Their platform ships a suite of AI “agents,” including retail loss prevention and warehouse safety, that run with a patented “Split AI” approach: real-time inference on-prem for speed and bandwidth efficiency, paired with cloud services for aggregation and search. Dragonfruit needed to keep total cost of ownership low, meet strict latency targets, and operate reliably across hundreds of sites with thousands of cameras; all without asking customers to rip and replace existing infrastructure.
*“We use customers’ existing camera networks and deliver the lowest total cost of ownership we can. On-prem inference plus smart use of the cloud is what makes it practical at retail scale.”*
— Karissa Price, Chief Customer Officer, Dragonfruit AI
### The challenge: Real-time at messy, planetary scale
Retail and warehouse environments are bandwidth-constrained and heterogeneous. A single store may run 25–130 IP cameras, and many Dragonfruit agents are real-time, such as burglar alarms or self-checkout monitoring. The engineering requirements were:
* Process \~30 FPS per camera on edge devices (Mac Minis) and move only compact inference data to the cloud.
* Track people and objects across multiple cameras and locations, which requires robust embeddings, re-identification, and fast similarity search.
* Sustain high ingestion and high query throughput simultaneously, with strict tail latencies.
* Operate a vector store at enterprise scale: thousands of locations → thousands of cameras, accumulating into tens to hundreds of billions of vectors and multi-terabyte storage.
### Why Qdrant: Performance headroom and operational control
Dragonfruit chose the open-source version of Qdrant as its vector search engine to meet the twin pressures of real-time reads and high-velocity writes. In head-to-head experiments, Qdrant delivered the QPS targets they needed while giving the team granular, [per-collection](https://qdrant.tech/documentation/concepts/collections/) tuning to match workload diversity.
Key reasons the team highlighted:
* **Per-collection configurability.** Collections with heavy reads and low writes use different settings than write-heavy pipelines. Tuning shard counts and HNSW parameters by collection helped hit latency Service Level Objectives (SLOs) without overprovisioning.
* **Efficient numeric formats.** For most vision workloads, [float16](https://qdrant.tech/documentation/concepts/vectors/) vectors were sufficient, improving memory efficiency and cache behavior with no material loss in retrieval accuracy for their use cases.
* **Open source and ecosystem fit.** Qdrant’s OSS model aligned with Dragonfruit’s platform strategy and let them co-evolve the deployment with their edge and cloud stack.
*“With Qdrant’s collection-level controls, we matched very different workloads, some ingestion-heavy, some read-intensive, and still hit real-time query performance.”*
— **Shivang Agarwal, VP Engineering, Dragonfruit AI**
### Solution architecture: Split AI \+ vector search
At the edge, Mac Minis ingest RTSP streams from IP cameras, run on-prem inference, and emit compact embeddings and event metadata. In the cloud, Qdrant serves multiple, distinct retrieval patterns:
* **Multi-camera person re-ID.** Embeddings link tracks across cameras to build “spaghetti charts” of in-store movement, enabling spatial analytics and alerting.
* **Self-checkout product verification.** Frames around scan events are embedded and matched against clustered product libraries to detect missed or incorrect scans in real time.
* **Full-frame semantic search.** Video frames are embedded for [text-to-image retrieval](https://qdrant.tech/advanced-search/) (for example “emergency door open”), enabling rapid incident triage and safety reviews.
### Results: New agents, faster delivery, lower total cost of ownership
Qdrant became an enabling layer for Dragonfruit’s agent roadmap. With real-time retrieval performance and cost-efficient storage, the team launched and iterated on new domain-specific agents more quickly, spanning loss prevention, occupational safety, and warehouse operations.
From a go-to-market perspective, Dragonfruit focuses on enterprise buyers and also works through major IT and security integrators, giving them reach into retail (their largest segment), manufacturing, government, entertainment, and healthcare. As Price summarized:
*“Our work with Qdrant gives us the agility to say ‘yes’ faster, adding new agents and answering new use cases, while staying cost-effective.”*
Quantitative highlights (self-reported by the team):
* Real-time processing of many cameras per site at high FPS on edge, with only inference data transmitted.
* Format optimization: float16 embeddings standard for most pipelines to reduce memory and improve throughput while maintaining retrieval quality.
### Lessons learned
Operating retrieval for vision at enterprise scale is as much an operations problem as an algorithms problem. Dragonfruit learned to tune per workload, not per system. Treat each collection as a workload with its own performance profile; shard counts, flush intervals, and vector precision matter.
### Conclusion
Dragonfruit shows how a Split-AI architecture plus a purpose-built vector search engine can turn ubiquitous cameras into reliable, low-latency analytics at enterprise scale. Qdrant’s performance headroom, per-collection controls, and operational simplicity helped the team meet real-time constraints and expand their agent portfolio without ballooning costs or bandwidth. As the platform grows, tighter, native data-lifecycle tooling will further reduce operational toil, but the core result is already clear: fast, affordable, and scalable computer vision for the real world.
@@ -0,0 +1,108 @@
---
draft: false
title: "How Xaver scaled personalized financial advice with Qdrant"
short_description: "Xaver built a compliant AI advisory engine with Qdrant."
description: "Discover how Xaver used Qdrant to power its two-tier knowledge engine, enabling sub-second, compliant AI-assisted financial consultations across chat, video, and phone."
preview_image: /blog/case-study-xaver/social_preview_partnership-xaver.jpg
social_preview_image: /blog/case-study-xaver/social_preview_partnership-xaver.jpg
date: 2025-11-13
author: "Daniel Azoulai"
featured: true
tags:
- Xaver
- vector search
- financial services
- AI agents
- knowledge engine
- latency optimization
- case study
---
![Xaver Overview](/blog/case-study-xaver/xaver-bento-box-dark.jpg)
## How Xaver Built its AI Knowledge Engine with Qdrant
<a href="https://www.xaver.com/" target="_blank">Xaver</a> is tackling a core challenge in the financial industry: scaling personalized financial and retirement advice. As demographic shifts increase demand for private pensions, traditional, manual consultation models are proving too slow and costly to support everyone who needs help.
To solve this, Xaver provides banks, insurers and distributors with a vertically specialized and compliant agentic sales platform. This technology acts as both an AI sales assistant for human advisors and as an autonomous agent to deliver compliant, personalized financial guidance to consumers via phone, video avatars, messengers and web journeys.
One core component of the platform is a fast, flexible knowledge engine designed to provide instant, contextually accurate answers for consulting AI agents.
*“Every second of latency matters in a phone or video consultation. Qdrant gave us the performance foundation to serve knowledge in real time without sacrificing quality.”*
Ole Breulmann, Founder / CPTO**, Xaver**
### The challenge: Bringing scale and speed to pension consultation
Private pension products are critical to addressing old-age poverty, yet the process of advising consumers remains highly manual and fragmented. Sales and consultation are typically handled by human brokers or advisers, each with their own methods, tools, and systems. This creates barriers to access, especially for consumers who expect digital-first financial services available on demand.
Xaver’s goal was to empower financial institutions to meet this demand with an always-available, AI-assisted advisory experience that maintains the trust and compliance standards of traditional consultation. That required:
* Real-time performance for interactive use cases like voice or video conversations.
* Reliable knowledge retrieval across channels such as WhatsApp, web chat, and advisor co-pilots.
* High-quality responses with guardrails for confidence and accuracy.
To make this work, Xaver needed a system that could manage knowledge retrieval and reasoning under tight latency constraints while remaining transparent, explainable, and easy to scale.
### The solution: Semantic caching, or a two-layer knowledge engine
[Qdrant](https://qdrant.tech/documentation/overview/) was selected after extensive evaluation for several reasons:
Xaver’s AI platform includes a “knowledge engine,” an [indexing](https://qdrant.tech/documentation/concepts/indexing/) and [retrieval](https://qdrant.tech/documentation/beginner-tutorials/retrieval-quality/) layer that feeds contextually relevant insights to both automated and human-assisted consultations. It powers two key functions:
1. Automated consultation through AI-led sessions via phone, video avatar, messengers, or web chat.
2. Advisor co-pilot that provides real-time context and recommendations for human consultants during live sessions.
To meet latency goals and maintain precision, Xaver implemented a two-tier retrieval architecture:
* **Tier 1: Condensed knowledge base (CKB).** A curated index of pre-summarized answers for the most common use case specific questions. This enables near-instant recall without triggering new large language model (LLM) summarization steps.
* **Tier 2: Full knowledge base.** A deeper corpus containing regulatory, financial, and policy documents, used when Tier 1 confidence is low or when a query requires more detailed reasoning.
This architecture enables the Xaver platform to deliver knowledge for the most common conversational situations instantly, while supporting rare or complex cases through a second-tier retrieval layer. The approach minimizes computational overhead, reduces response time by two to three seconds in typical cases, and preserves conversational flow, which is crucial for voice and video experiences.
![Xaver Retrieval Process](/blog/case-study-xaver/xaver-retrieval-process.jpg)
*Figure: Xaver’s retrieval process*
### Why Qdrant
When designing the knowledge engine, Xaver deliberately chose to focus on its application layer instead of building core infrastructure from scratch. The team wanted to dedicate engineering effort to customer-facing innovation, not database maintenance.
[Qdrant](https://qdrant.tech/documentation/overview/) was selected after extensive evaluation for several reasons:
* **High performance at low latency.** Its [Rust-based core](https://qdrant.tech/articles/why-rust/) allowed Xaver to meet strict real-time response thresholds for conversational use cases.
* **Developer simplicity.** Qdrant’s [intuitive API](https://api.qdrant.tech/api-reference) and clean operational model helped the team integrate quickly without diverting resources to maintain infrastructure.
* **Flexible deployment.** The platform could run both the condensed and full knowledge bases side by side with custom confidence thresholds and filters.
* **Future readiness.** Qdrant’s stability and scalability aligned with Xaver’s roadmap to expand across channels and markets without adding infrastructure complexity.
By relying on Qdrant for the vector storage and retrieval layer, Xaver was able to stay focused on what truly differentiated its product: the AI logic, compliance workflows, and personalized advisory experience that sit above the database layer.
*“Qdrant turned vector search from a bottleneck into a runway; faster results and more room to innovate with indexing strategies in production.”*
Arijit Das, PhD, Senior AI Engineer, Xaver
### Results: Real-time guidance and trusted digital advisory
With Qdrant underpinning the knowledge engine, Xaver achieved a responsive and reliable consultation experience across multiple modalities:
* Sub-second retrieval for common intents in real-time phone and video sessions.
* Improved advisor efficiency through contextual recommendations delivered in co-pilot mode.
* Unified knowledge management across all channels from chat to audio.
These gains allowed financial institutions to modernize customer interactions while preserving the human quality and regulatory rigor that define financial advice.
![multi-channel approach for financial advisory](/blog/case-study-xaver/xaver_qdrant.png)
*Xaver’s multi-channel approach for financial advisory*
### Providing the speed, reliability, and simplicity for efficient AI-driven consultations
For Xaver, success is about prioritization: investing deeply in the vertical AI expertise and the user experience while outsourcing the heavy lifting of vector search infrastructure to a trusted partner. Qdrant provided the speed, reliability, and simplicity needed to make efficient AI-driven consultation a reality for financial institutions.
The result is a scalable, human-centered system that brings trusted financial guidance into the digital age, fast, compliant, and ready for the future.
@@ -785,7 +785,7 @@ _Appears in:_
| --- | --- | --- | --- |
| `id` _string_ | Id specifies the unique identifier of the cluster | | |
| `version` _string_ | Version specifies the version of Qdrant to deploy | | |
| `size` _integer_ | Size specifies the desired number of Qdrant nodes in the cluster | | Maximum: 30 <br />Minimum: 1 <br /> |
| `size` _integer_ | Size specifies the desired number of Qdrant nodes in the cluster | | Maximum: 100 <br />Minimum: 1 <br /> |
| `servicePerNode` _boolean_ | ServicePerNode specifies whether the cluster should start a dedicated service for each node. | true | |
| `clusterManager` _boolean_ | ClusterManager specifies whether to use the cluster manager for this cluster.<br />The Python-operator will deploy a dedicated cluster manager instance.<br />The Go-operator will use a shared instance.<br />If not set, the default will be taken from the operator config. | | |
| `suspend` _boolean_ | Suspend specifies whether to suspend the cluster.<br />If enabled, the cluster will be suspended and all related resources will be removed except the PVCs. | false | |
@@ -801,11 +801,14 @@ _Appears in:_
| `gpu` _[GPU](#gpu)_ | GPU specifies GPU configuration for the cluster. If this field is not set, no GPU will be used. | | |
| `statefulSet` _[KubernetesStatefulSet](#kubernetesstatefulset)_ | StatefulSet specifies the configuration of the Qdrant Kubernetes StatefulSet. | | |
| `storageClassNames` _[StorageClassNames](#storageclassnames)_ | StorageClassNames specifies the storage class names for db and snapshots. | | |
| `storageTier` _[StorageTier](#storagetier)_ | StorageTier specifies the performance tier to use for the disk | | Enum: [budget balanced performance] <br /> |
| `topologySpreadConstraints` _[TopologySpreadConstraint](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#topologyspreadconstraint-v1-core)_ | TopologySpreadConstraints specifies the topology spread constraints for the cluster. | | |
| `podDisruptionBudget` _[PodDisruptionBudgetSpec](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#poddisruptionbudgetspec-v1-policy)_ | PodDisruptionBudget specifies the pod disruption budget for the cluster. | | |
| `restartAllPodsConcurrently` _boolean_ | RestartAllPodsConcurrently specifies whether to restart all pods concurrently (also called one-shot-restart).<br />If enabled, all the pods in the cluster will be restarted concurrently in situations where multiple pods<br />need to be restarted, like when RestartedAtAnnotationKey is added/updated or the Qdrant version needs to be upgraded.<br />This helps sharded but not replicated clusters to reduce downtime to a possible minimum during restart.<br />If unset, the operator is going to restart nodes concurrently if none of the collections if replicated. | | |
| `startupDelaySeconds` _integer_ | If StartupDelaySeconds is set (> 0), an additional 'sleep <value>' will be emitted to the pod startup.<br />The sleep will be added when a pod is restarted, it will not force any pod to restart.<br />This feature can be used for debugging the core, e.g. if a pod is in crash loop, it provided a way<br />to inspect the attached storage. | | |
| `rebalanceStrategy` _[RebalanceStrategy](#rebalancestrategy)_ | RebalanceStrategy specifies the strategy to use for automaticially rebalancing shards the cluster.<br />Cluster-manager needs to be enabled for this feature to work. | | Enum: [by_count by_size by_count_and_size] <br /> |
| `readClusters` _[ReadCluster](#readcluster) array_ | ReadClusters specifies the read clusters for this cluster to synchronize.<br />Cluster-manager needs to be enabled for this feature to work. | | |
| `writeCluster` _[WriteCluster](#writecluster)_ | WriteCluster specifies the write cluster for this cluster. This configures the NetworkPolicy to allow egress to the write cluster. | | |
@@ -847,6 +850,23 @@ _Appears in:_
| `replication_factor` _integer_ | ReplicationFactor specifies the default number of replicas of each shard | | |
| `write_consistency_factor` _integer_ | WriteConsistencyFactor specifies how many replicas should apply the operation to consider it successful | | |
| `vectors` _[QdrantConfigurationCollectionVectors](#qdrantconfigurationcollectionvectors)_ | Vectors specifies the default parameters for vectors | | |
| `strict_mode` _[QdrantConfigurationCollectionStrictMode](#qdrantconfigurationcollectionstrictmode)_ | StrictMode specifies the strict mode configuration for the collection | | |
#### QdrantConfigurationCollectionStrictMode
_Appears in:_
- [QdrantConfigurationCollection](#qdrantconfigurationcollection)
| Field | Description | Default | Validation |
| --- | --- | --- | --- |
| `max_payload_index_count` _integer_ | MaxPayloadIndexCount represents the maximal number of payload indexes allowed to be created.<br />It can be set for Qdrant version >= 1.16.0<br />Default to 100 if omitted and Qdrant version >= 1.16.0 | | Minimum: 1 <br /> |
#### QdrantConfigurationCollectionVectors
@@ -1097,6 +1117,22 @@ _Appears in:_
| `fsGroup` _integer_ | FsGroup specifies file system group to run the Qdrant process as. | | |
#### ReadCluster
_Appears in:_
- [QdrantClusterSpec](#qdrantclusterspec)
| Field | Description | Default | Validation |
| --- | --- | --- | --- |
| `id` _string_ | Id specifies the unique identifier of the read cluster | | |
#### RebalanceStrategy
_Underlying type:_ _string_
@@ -1202,6 +1238,7 @@ _Appears in:_
| --- | --- | --- | --- |
| `name` _string_ | Name of the destination cluster | | |
| `namespace` _string_ | Namespace of the destination cluster | | |
| `create` _boolean_ | Create when set to true indicates that<br />a new cluster with the specified name should be created.<br />Otherwise, if set to false, the existing cluster is going to be restored<br />to the specified state. | | |
#### RestorePhase
@@ -1221,6 +1258,7 @@ _Appears in:_
| `Skipped` | |
| `Failed` | |
| `Succeeded` | |
| `Pending` | |
#### RestoreSource
@@ -1329,6 +1367,25 @@ _Appears in:_
| `async_scorer` _boolean_ | AsyncScorer enables io_uring when rescoring | | |
#### StorageTier
_Underlying type:_ _string_
StorageTier specifies the performance profile for the disk to use.
_Validation:_
- Enum: [budget balanced performance]
_Appears in:_
- [QdrantClusterSpec](#qdrantclusterspec)
| Field | Description |
| --- | --- |
| `budget` | |
| `balanced` | |
| `performance` | |
#### TraefikConfig
@@ -1381,3 +1438,19 @@ _Appears in:_
| `readyToUse` _boolean_ | ReadyToUse indicates if the volume snapshot is ready to use | | |
| `snapshotHandle` _string_ | SnapshotHandle is the identifier of the volume snapshot in the respective cloud provider | | |
#### WriteCluster
_Appears in:_
- [QdrantClusterSpec](#qdrantclusterspec)
| Field | Description | Default | Validation |
| --- | --- | --- | --- |
| `id` _string_ | Id specifies the unique identifier of the write cluster | | |
@@ -5,6 +5,21 @@ weight: 5
# Changelog
## 1.9.1 (2025-11-14)
| Component | Version |
|-------------------------|---------|
| qdrant-kubernetes-api | v1.20.0 |
| operator | 2.8.1 |
| qdrant-cluster-manager | v0.3.9 |
| qdrant-cluster-exporter | 1.7.2 |
* Automated storage class migration
* Support for max_payload_index_count in strict mode
* Experimental support for read replica clusters
* Enable restoring snapshots into another cluster
* Performance and stability improvements
## 1.8.0 (2025-08-08)
| Component | Version |
@@ -66,13 +66,13 @@ helm registry login 'registry.cloud.qdrant.io' --username 'your-username' --pass
4. Install the Qdrant Kubernetes Operator Custom Resource Definitions (CRDs):
```bash
helm upgrade --install qdrant-private-cloud-crds oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-kubernetes-api --namespace qdrant-private-cloud --version v1.17.2 --wait
helm upgrade --install qdrant-private-cloud-crds oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-kubernetes-api --namespace qdrant-private-cloud --version v1.20.0 --wait
```
5. Install Qdrant Private Cloud:
```bash
helm upgrade --install qdrant-private-cloud oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-private-cloud --namespace qdrant-private-cloud --version 1.8.0
helm upgrade --install qdrant-private-cloud oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-private-cloud --namespace qdrant-private-cloud --version 1.9.1
```
Ensure that the `qdrant-kubernetes-api` version is compatible with the `qdrant-private-cloud` version you are installing.
@@ -81,8 +81,8 @@ For a list of available versions consult the [Private Cloud Changelog](/document
Current default versions are:
* qdrant-kubernetes-api v1.17.2
* qdrant-private-cloud 1.8.0
* qdrant-kubernetes-api v1.20.0
* qdrant-private-cloud 1.9.1
For more information also see the [Helm Install Documentation](https://helm.sh/docs/helm/helm_install/).
@@ -110,7 +110,7 @@ operator:
You can configure Qdrant Private Cloud like this:
```bash
helm upgrade --install qdrant-private-cloud oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-private-cloud --namespace qdrant-private-cloud --version 1.8.0 -f values.yaml
helm upgrade --install qdrant-private-cloud oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-private-cloud --namespace qdrant-private-cloud --version 1.9.1 -f values.yaml
```
## Upgrades
@@ -118,13 +118,13 @@ helm upgrade --install qdrant-private-cloud oci://registry.cloud.qdrant.io/qdran
To upgrade Qdrant Private Cloud to a new version, first upgrade the Qdrant Kubernetes Operator Custom Resource Definitions (CRDs):
```bash
helm upgrade --install qdrant-private-cloud-crds oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-kubernetes-api --namespace qdrant-private-cloud --version v1.17.2 --wait
helm upgrade --install qdrant-private-cloud-crds oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-kubernetes-api --namespace qdrant-private-cloud --version v1.20.0 --wait
```
Then upgrade the Qdrant Private Cloud Helm chart using the same configuration values, e.g.:
```bash
helm upgrade --install qdrant-private-cloud oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-private-cloud --namespace qdrant-private-cloud --version 1.8.0 -f values.yaml
helm upgrade --install qdrant-private-cloud oci://registry.cloud.qdrant.io/qdrant-charts/qdrant-private-cloud --namespace qdrant-private-cloud --version 1.9.1 -f values.yaml
```
Note, that the image tag values are automatically derived from the chart's appVersions and should not be overridden in the `values.yaml`.
+2 -2
View File
@@ -1,6 +1,6 @@
---
stats:
githubStars: 27.0k
discordMembers: 8.8k
githubStars: 27.1k
discordMembers: 8.9k
twitterFollowers: 7.5k
---
Binary file not shown.

After

Width:  |  Height:  |  Size: 232 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 275 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 251 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 757 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 51 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 687 KiB