* Add new Scaling landing page under Operations Introduces a Scaling section with vertical vs. horizontal scaling guidance and failover best practices, linking out to detail pages. * Add new Vertical Scaling page Dedicated how-to guidance for resizing existing nodes: when to scale vertically, RAM sizing formulas, and Cloud/self-hosted resize steps. * Add new Horizontal Scaling and Resilience page Covers Raft consensus, the replication model, consistency guarantees, Multi-AZ, and the resilience terminology used elsewhere in the docs. * Move Distributed Deployment under Scaling and update all incoming links Moves distributed_deployment.md into the new scaling/ section, trims its Raft/Replication/Consistency intros into cross-links to the new Horizontal Scaling and Resilience page, adds Multi-AZ and single-replica cross-link callouts in the Cloud docs, rewrites all internal references across ~30 files to the new canonical path instead of relying on aliases, and applies Title Case to Distributed Deployment's headers. * Split Resilience out of Horizontal Scaling and Resilience Adds a dedicated Resilience page covering fault tolerance, Multi-AZ, resilience terminology, and failover best practices (moved from the Scaling landing page). Horizontal Scaling is retitled and scoped to the underlying mechanics: Raft consensus, replication, and consistency. * Reorganize Horizontal Scaling's structure Moves "How Many Qdrant Nodes Should I Run?" from Distributed Deployment into Horizontal Scaling, adds a conceptual Sharding section, and reorders Sharding/Replication/Raft Consensus/Consistency. Moves the remaining conceptual content out of Distributed Deployment: Temporary Node Failure to Resilience, Error Handling folded into Replication, sharding heuristics folded into Sharding, and the Consensus Checkpointing explanation folded into Raft Consensus. * Rename Scaling section to Scaling & Resilience Renames the section and restructures the landing page: the vertical- vs-horizontal decision is now purely about scaling, with a dedicated Resilience section covering fault tolerance through sharding and multi-node deployments. * Polish Vertical Scaling and Resilience page content Reframes Vertical Scaling's "What Not to Do" as positive "Best Practices". Reworks Resilience's structure: moves the uptime/data- integrity terminology into the intro as three distinct aspects of resilience, and renames "How Resilience Works" to "Setting Up a Resilient Qdrant Cluster". * Add diagrams illustrating sharding and replication Adds cluster diagrams to the Sharding and Replication sections on Horizontal Scaling to make the shard/replica layout easier to follow. * Add new Node Failure Recovery page Extracts the node failure recovery scenarios out of Distributed Deployment into their own page, with each bolded sub-header converted to a proper heading, and links updated across Resilience and the Scaling landing page. * Add new Consistency Guarantees page Extracts write consistency factor, read consistency, and write ordering out of Distributed Deployment into their own page, positioned after Distributed Deployment. * Add new "Deploy Behind a Load Balancer" section Explains why a load balancer is needed in front of a multi-node Qdrant cluster: avoiding a single point of failure at the entry point and making sure replicas on every node actually serve reads. * Add new "Rebalancing" section Documents how Qdrant Cloud automatically rebalances shards across nodes, as its own subsection under Sharding. * Rewrite Multi-AZ vs. Replication Factor as Multi-AZ Deployments Defines an availability zone on first use, explains why multi-AZ deployments guard against a zone going down, clarifies that Qdrant Cloud is zone-aware once enabled, and that self-hosted deployments need to place and move replicas across zones manually. * Restructure node-count guidance into One/Two/Three-or-more Node subsections Splits "How Many Qdrant Nodes Should I Run?" into three subsections and drops the "balanced" framing for two nodes: it states plainly that two nodes give more capacity without true high availability. * Add new "Which Configuration Is Right for You?" section Summarizes the one/two/three-or-more node tradeoffs in one place right after the detailed breakdown. * Add explicit _redirects entry for legacy distributed_deployment URL Closes the redirect chain: the existing /guides/ and /operations/ legacy rules both terminate at /documentation/distributed_deployment/, which previously had no explicit _redirects entry and only resolved via the Hugo alias meta-refresh page. * Fix all incoming links to Distributed Deployment and pages under Scaling Repoints two same-page anchors in distributed_deployment.md that broke when Write Ordering moved to Consistency Guarantees, and one link in cloud/create-cluster.md that broke when a Resilience heading was reworded. * Update time-based sharding diagram and restructure section * Fix a couple of broken links * Move 'Consensus Checkpointing' to 'Node Failure Recovery' page
8.1 KiB
draft, title, short_description, description, preview_image, social_preview_image, date, author, featured, tags, partition
| draft | title | short_description | description | preview_image | social_preview_image | date | author | featured | tags | partition | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| false | Voiceflow & Qdrant: Powering No-Code AI Agent Creation with Scalable Vector Search | Enabling scalable, no-code AI agent creation. | Learn how Voiceflow builds scalable, customizable, no-code AI agent solutions for enterprises. | /blog/case-study-voiceflow/image0.png | /blog/case-study-voiceflow/image0.png | 2024-12-10T00:02:00Z | Qdrant | false |
|
case-studies |
Voiceflow enables enterprises to create AI agents in a no-code environment by designing workflows through a drag-and-drop interface. The platform allows developers to host and customize chatbot interfaces without needing to build their own RAG pipeline, working out of the box and being easily adaptable to specific use cases. “Powered by technologies like Natural Language Understanding (NLU), Large Language Models (LLM), and Qdrant as a vector search engine, Voiceflow serves a diverse range of customers, including enterprises that develop chatbots for internal and external AI use cases,” says Xavier Portillo Edo, Head of Cloud Infrastructure at Voiceflow.
Evaluation Criteria
Denys Linkov, Machine Learning Team Lead at Voiceflow, explained the journey of building a managed RAG solution. "Initially, our product focused on users manually defining steps on the canvas. After the release of ChatGPT, we added AI-based responses, leading to the launch of our managed RAG solution in the spring of 2023," Linkov said.
As part of this development, the Voiceflow engineering team was looking for a vector database solution to power their RAG setup. They evaluated various vector databases based on several key factors:
- Performance: The ability to handle the scale required by Voiceflow, supporting hundreds of thousands of projects efficiently.
- Metadata: The capability to tag data and chunks and retrieve based on those values, essential for organizing and accessing specific information swiftly.
- Managed Solution: The availability of a managed service with automated maintenance, scaling, and security, freeing the team from infrastructure concerns.
"We started with Pinecone but eventually switched to Qdrant," Linkov noted. The reasons for the switch included:
- Scaling Capabilities: Qdrant offers a robust multi-node setup with horizontal scaling, allowing clusters to grow by adding more nodes and distributing data and load among them. This ensures high performance and resilience, which is crucial for handling large-scale projects.
- Infrastructure: “Qdrant provides robust infrastructure support, allowing integration with virtual private clouds on AWS using AWS Private Links and ensuring encryption with AWS KMS. This setup ensures high security and reliability,” says Portillo Edo.
- Responsive Qdrant Team: "The Qdrant team is very responsive, ships features quickly and is a great partner to build with," Linkov added.
Migration and Onboarding
Voiceflow began its migration to Qdrant by creating backups and ensuring data consistency through random checks and key customer verifications. "Once we were confident in the stability, we transitioned the primary database to Qdrant, completing the migration smoothly," Linkov explained.
During onboarding, Voiceflow transitioned from namespaces to Qdrant's collections, which offer enhanced flexibility and advanced vector search capabilities. They also implemented Quantization to enhance data processing efficiency. This comprehensive process ensured a seamless transition to Qdrant's robust infrastructure.
RAG Pipeline Setup
Voiceflow's RAG pipeline setup provides a streamlined process for uploading and managing data from various sources, designed to offer flexibility and customization at each step.
- Data Upload: Customers can upload data via API from sources such as URLs, PDFs, Word documents, and plain text formats. Integration with platforms like Zendesk is supported, and users can choose between single uploads or refresh-based uploads.
- Data Ingestion: Once data is ingested, Voiceflow offers preset strategies for data checking. Users can utilize these strategies or opt for more customization through the API to tailor the ingestion process as needed.
- Metadata Tagging: Metadata tags can be applied during the ingestion process, which helps organize and facilitate efficient data retrieval later on.
- Data Retrieval: At retrieval time, Voiceflow provides prompts that can modify user questions by adding context, variables, or other modifications. This customization includes adding personas or structuring responses as markdown. Depending on the type of interaction (e.g., button, carousel with an image for image retrieval), these prompts are displayed to users in a structured format.
This comprehensive setup ensures that Voiceflow users can efficiently manage and customize their data workflows, providing a robust solution for building AI-driven applications.
How Voiceflow Uses Qdrant
Voiceflow leverages Qdrant's robust features and infrastructure to optimize their AI assistant platform. Here’s a breakdown of how they utilize these capabilities:
Database Features:
- Quantization: This feature helps Voiceflow to perform efficient data processing by reducing the size of vectors, making searches faster. The team uses Product Quantization in particular.
- Chunking Search: Voiceflow uses chunking search to improve search efficiency by breaking down large datasets into manageable chunks, which allows for faster and more efficient data retrieval.
- Sparse Vector Search: Although not yet implemented, this feature is being explored for more precise keyword searches. "This is an encouraging direction the Qdrant team is taking here as many users seek more exact keyword search," said Linkov.
Architecture:
- Node Pool: A large node pool is used for public cloud users, ensuring scalability, while several smaller, isolated instances cater to private cloud users, providing enhanced security.
Infrastructure:
- Private Link: The ability to use Private Link connections across different instances is a significant advantage, requiring robust infrastructure support from Qdrant. "This setup was crucial for SOC2 compliance, and Qdrant's support team made the process seamless by ensuring feasibility and aiding in the implementation," Linkov explained.
By utilizing these features, Voiceflow ensures that its platform is scalable, secure, and efficient, meeting the diverse needs of its users.
The Outcome
Voiceflow achieved significant improvements and efficiencies by leveraging Qdrant's capabilities:
- Enhanced Metadata Tagging: Implemented robust metadata tagging, allowing for custom fields and tags that facilitate efficient search filtering.
- Optimized Performance: Resolved concerns about retrieval times with a high number of tags by optimizing indexing strategies, achieving efficient performance.
- Minimal Operational Overhead: Experienced minimal overhead, streamlining their operational processes.
- Future-Ready: Anticipates further innovation in hybrid search with multi-token attention.
- Multitenancy Support: Utilized Qdrant's efficient and isolated data management to support diverse user needs.
Overall, Qdrant's features and infrastructure provided Voiceflow with a stable, scalable, and efficient solution for their data processing and retrieval needs.
What’s Next
Voiceflow plans to enhance its platform with more filtering and customization options, allowing developers to host and customize chatbot interfaces without building their own RAG pipeline.
