Files
landing_page/qdrant-landing/content/blog/case-study-voiceflow.md
T
Abdon Pijpelink 45f19f30ee Break up "Distributed Deployment" page into new "Scaling & Resilience" section (#2491)
* Add new Scaling landing page under Operations

Introduces a Scaling section with vertical vs. horizontal scaling
guidance and failover best practices, linking out to detail pages.

* Add new Vertical Scaling page

Dedicated how-to guidance for resizing existing nodes: when to scale
vertically, RAM sizing formulas, and Cloud/self-hosted resize steps.

* Add new Horizontal Scaling and Resilience page

Covers Raft consensus, the replication model, consistency guarantees,
Multi-AZ, and the resilience terminology used elsewhere in the docs.

* Move Distributed Deployment under Scaling and update all incoming links

Moves distributed_deployment.md into the new scaling/ section, trims
its Raft/Replication/Consistency intros into cross-links to the new
Horizontal Scaling and Resilience page, adds Multi-AZ and single-replica
cross-link callouts in the Cloud docs, rewrites all internal references
across ~30 files to the new canonical path instead of relying on
aliases, and applies Title Case to Distributed Deployment's headers.

* Split Resilience out of Horizontal Scaling and Resilience

Adds a dedicated Resilience page covering fault tolerance, Multi-AZ,
resilience terminology, and failover best practices (moved from the
Scaling landing page). Horizontal Scaling is retitled and scoped to
the underlying mechanics: Raft consensus, replication, and consistency.

* Reorganize Horizontal Scaling's structure

Moves "How Many Qdrant Nodes Should I Run?" from Distributed Deployment
into Horizontal Scaling, adds a conceptual Sharding section, and
reorders Sharding/Replication/Raft Consensus/Consistency. Moves the
remaining conceptual content out of Distributed Deployment: Temporary
Node Failure to Resilience, Error Handling folded into Replication,
sharding heuristics folded into Sharding, and the Consensus
Checkpointing explanation folded into Raft Consensus.

* Rename Scaling section to Scaling & Resilience

Renames the section and restructures the landing page: the vertical-
vs-horizontal decision is now purely about scaling, with a dedicated
Resilience section covering fault tolerance through sharding and
multi-node deployments.

* Polish Vertical Scaling and Resilience page content

Reframes Vertical Scaling's "What Not to Do" as positive "Best
Practices". Reworks Resilience's structure: moves the uptime/data-
integrity terminology into the intro as three distinct aspects of
resilience, and renames "How Resilience Works" to "Setting Up a
Resilient Qdrant Cluster".

* Add diagrams illustrating sharding and replication

Adds cluster diagrams to the Sharding and Replication sections on
Horizontal Scaling to make the shard/replica layout easier to follow.

* Add new Node Failure Recovery page

Extracts the node failure recovery scenarios out of Distributed
Deployment into their own page, with each bolded sub-header converted
to a proper heading, and links updated across Resilience and the
Scaling landing page.

* Add new Consistency Guarantees page

Extracts write consistency factor, read consistency, and write
ordering out of Distributed Deployment into their own page, positioned
after Distributed Deployment.

* Add new "Deploy Behind a Load Balancer" section

Explains why a load balancer is needed in front of a multi-node
Qdrant cluster: avoiding a single point of failure at the entry point
and making sure replicas on every node actually serve reads.

* Add new "Rebalancing" section

Documents how Qdrant Cloud automatically rebalances shards across
nodes, as its own subsection under Sharding.

* Rewrite Multi-AZ vs. Replication Factor as Multi-AZ Deployments

Defines an availability zone on first use, explains why multi-AZ
deployments guard against a zone going down, clarifies that Qdrant
Cloud is zone-aware once enabled, and that self-hosted deployments
need to place and move replicas across zones manually.

* Restructure node-count guidance into One/Two/Three-or-more Node subsections

Splits "How Many Qdrant Nodes Should I Run?" into three subsections
and drops the "balanced" framing for two nodes: it states plainly
that two nodes give more capacity without true high availability.

* Add new "Which Configuration Is Right for You?" section

Summarizes the one/two/three-or-more node tradeoffs in one place
right after the detailed breakdown.

* Add explicit _redirects entry for legacy distributed_deployment URL

Closes the redirect chain: the existing /guides/ and /operations/
legacy rules both terminate at /documentation/distributed_deployment/,
which previously had no explicit _redirects entry and only resolved
via the Hugo alias meta-refresh page.

* Fix all incoming links to Distributed Deployment and pages under Scaling

Repoints two same-page anchors in distributed_deployment.md that broke
when Write Ordering moved to Consistency Guarantees, and one link in
cloud/create-cluster.md that broke when a Resilience heading was
reworded.

* Update time-based sharding diagram and restructure section

* Fix a couple of broken links

* Move 'Consensus Checkpointing' to 'Node Failure Recovery' page
2026-07-16 09:11:58 +02:00

8.1 KiB
Raw Blame History

draft, title, short_description, description, preview_image, social_preview_image, date, author, featured, tags, partition
draft title short_description description preview_image social_preview_image date author featured tags partition
false Voiceflow & Qdrant: Powering No-Code AI Agent Creation with Scalable Vector Search Enabling scalable, no-code AI agent creation. Learn how Voiceflow builds scalable, customizable, no-code AI agent solutions for enterprises. /blog/case-study-voiceflow/image0.png /blog/case-study-voiceflow/image0.png 2024-12-10T00:02:00Z Qdrant false
voiceflow
qdrant
no-code
AI agents
vector search
RAG
case study
case-studies

voiceflow/image2.png

Voiceflow enables enterprises to create AI agents in a no-code environment by designing workflows through a drag-and-drop interface. The platform allows developers to host and customize chatbot interfaces without needing to build their own RAG pipeline, working out of the box and being easily adaptable to specific use cases. “Powered by technologies like Natural Language Understanding (NLU), Large Language Models (LLM), and Qdrant as a vector search engine, Voiceflow serves a diverse range of customers, including enterprises that develop chatbots for internal and external AI use cases,” says Xavier Portillo Edo, Head of Cloud Infrastructure at Voiceflow.

Evaluation Criteria

Denys Linkov, Machine Learning Team Lead at Voiceflow, explained the journey of building a managed RAG solution. "Initially, our product focused on users manually defining steps on the canvas. After the release of ChatGPT, we added AI-based responses, leading to the launch of our managed RAG solution in the spring of 2023," Linkov said.

As part of this development, the Voiceflow engineering team was looking for a vector database solution to power their RAG setup. They evaluated various vector databases based on several key factors:

  • Performance: The ability to handle the scale required by Voiceflow, supporting hundreds of thousands of projects efficiently.
  • Metadata: The capability to tag data and chunks and retrieve based on those values, essential for organizing and accessing specific information swiftly.
  • Managed Solution: The availability of a managed service with automated maintenance, scaling, and security, freeing the team from infrastructure concerns.

"We started with Pinecone but eventually switched to Qdrant," Linkov noted. The reasons for the switch included:

  • Scaling Capabilities: Qdrant offers a robust multi-node setup with horizontal scaling, allowing clusters to grow by adding more nodes and distributing data and load among them. This ensures high performance and resilience, which is crucial for handling large-scale projects.
  • Infrastructure: “Qdrant provides robust infrastructure support, allowing integration with virtual private clouds on AWS using AWS Private Links and ensuring encryption with AWS KMS. This setup ensures high security and reliability,” says Portillo Edo.
  • Responsive Qdrant Team: "The Qdrant team is very responsive, ships features quickly and is a great partner to build with," Linkov added.

Migration and Onboarding

Voiceflow began its migration to Qdrant by creating backups and ensuring data consistency through random checks and key customer verifications. "Once we were confident in the stability, we transitioned the primary database to Qdrant, completing the migration smoothly," Linkov explained.

During onboarding, Voiceflow transitioned from namespaces to Qdrant's collections, which offer enhanced flexibility and advanced vector search capabilities. They also implemented Quantization to enhance data processing efficiency. This comprehensive process ensured a seamless transition to Qdrant's robust infrastructure.

RAG Pipeline Setup

Voiceflow's RAG pipeline setup provides a streamlined process for uploading and managing data from various sources, designed to offer flexibility and customization at each step.

  • Data Upload: Customers can upload data via API from sources such as URLs, PDFs, Word documents, and plain text formats. Integration with platforms like Zendesk is supported, and users can choose between single uploads or refresh-based uploads.
  • Data Ingestion: Once data is ingested, Voiceflow offers preset strategies for data checking. Users can utilize these strategies or opt for more customization through the API to tailor the ingestion process as needed.
  • Metadata Tagging: Metadata tags can be applied during the ingestion process, which helps organize and facilitate efficient data retrieval later on.
  • Data Retrieval: At retrieval time, Voiceflow provides prompts that can modify user questions by adding context, variables, or other modifications. This customization includes adding personas or structuring responses as markdown. Depending on the type of interaction (e.g., button, carousel with an image for image retrieval), these prompts are displayed to users in a structured format.

This comprehensive setup ensures that Voiceflow users can efficiently manage and customize their data workflows, providing a robust solution for building AI-driven applications.

How Voiceflow Uses Qdrant

Voiceflow leverages Qdrant's robust features and infrastructure to optimize their AI assistant platform. Here’s a breakdown of how they utilize these capabilities:

Database Features:

  • Quantization: This feature helps Voiceflow to perform efficient data processing by reducing the size of vectors, making searches faster. The team uses Product Quantization in particular.
  • Chunking Search: Voiceflow uses chunking search to improve search efficiency by breaking down large datasets into manageable chunks, which allows for faster and more efficient data retrieval.
  • Sparse Vector Search: Although not yet implemented, this feature is being explored for more precise keyword searches. "This is an encouraging direction the Qdrant team is taking here as many users seek more exact keyword search," said Linkov.

Architecture:

  • Node Pool: A large node pool is used for public cloud users, ensuring scalability, while several smaller, isolated instances cater to private cloud users, providing enhanced security.

Infrastructure:

  • Private Link: The ability to use Private Link connections across different instances is a significant advantage, requiring robust infrastructure support from Qdrant. "This setup was crucial for SOC2 compliance, and Qdrant's support team made the process seamless by ensuring feasibility and aiding in the implementation," Linkov explained.

By utilizing these features, Voiceflow ensures that its platform is scalable, secure, and efficient, meeting the diverse needs of its users.

The Outcome

Voiceflow achieved significant improvements and efficiencies by leveraging Qdrant's capabilities:

  • Enhanced Metadata Tagging: Implemented robust metadata tagging, allowing for custom fields and tags that facilitate efficient search filtering.
  • Optimized Performance: Resolved concerns about retrieval times with a high number of tags by optimizing indexing strategies, achieving efficient performance.
  • Minimal Operational Overhead: Experienced minimal overhead, streamlining their operational processes.
  • Future-Ready: Anticipates further innovation in hybrid search with multi-token attention.
  • Multitenancy Support: Utilized Qdrant's efficient and isolated data management to support diverse user needs.

Overall, Qdrant's features and infrastructure provided Voiceflow with a stable, scalable, and efficient solution for their data processing and retrieval needs.

What’s Next

Voiceflow plans to enhance its platform with more filtering and customization options, allowing developers to host and customize chatbot interfaces without building their own RAG pipeline.