Merge branch 'master' into fix/monitoring-logging-doc-typos

This commit is contained in:
Mohamed Arbi
2026-03-31 02:11:16 +02:00
committed by GitHub
441 changed files with 5707 additions and 1189 deletions
@@ -2,6 +2,6 @@
title: Brand Resources
seeOpenRoles:
text: View Brand Resources
url: /brand-resources
url: /brand-resources/
image: /img/about-us/style-guide.svg
---
@@ -17,13 +17,13 @@ values:
icon:
src: /img/about-us/rust-logo.svg
alt: rust-logo
link: /articles/why-rust
link: /articles/why-rust/
- id: 3
title: 9k Members
icon:
src: /img/about-us/discord-logo.svg
alt: discord-logo
link: /community
link: /community/
- id: 4
title: 100+ employees
icon:
@@ -10,7 +10,7 @@ features:
description: Qdrant optimizes similarity search, identifying the closest database items to any query vector for applications like recommendation systems, RAG and image retrieval, enhancing accuracy and user experience.
link:
text: Learn More
url: /documentation/concepts/search/
url: /documentation/search/search/
- id: 1
icon:
src: /icons/outline/search-text-blue.svg
@@ -48,7 +48,7 @@ features:
description: Qdrant’s architecture is optimized for high-throughput embedding processing, minimizing CPU load and preventing performance bottlenecks. This enables AI agents in Agentic RAG workflows to execute complex, multi-step tasks efficiently, ensuring smooth operation even at scale.
link:
text: Distributed Deployment
url: /documentation/guides/distributed_deployment/
url: /documentation/operations/distributed_deployment/
- id: 4
icon:
src: /icons/outline/speedometer-blue.svg
@@ -14,7 +14,7 @@ integrations:
alt: Open AI logo
title: Swarm
description: Decentralized platform enabling collaboration among AI agents for task completion.
url: /documentation/frameworks/swarm/
url: /documentation/frameworks/
- id: 2
icon:
src: /img/integrations/integration-crew-ai.svg
@@ -20,7 +20,7 @@ features:
description: Learn how to build OpenAI Swarm agents using Qdrant for fast, scalable vector search and real-time actions.
link:
text: Read the Docs
url: /documentation/frameworks/swarm/
url: /documentation/frameworks/
- id: 2
image:
src: /img/ai-agents-use-cases/ai-scheduler.svg
@@ -38,7 +38,7 @@ Reliable agentic workflows require a clear plan executed with precise tools. A v
* [**Real-time memory layer:**](https://qdrant.tech/blog/case-study-fieldy/) fast access to prior steps, actions, and knowledge
* [**Multimodal support**](https://qdrant.tech/blog/case-study-mixpeek/): text, image, videos, audio, and code
* [**Hybrid search**](https://qdrant.tech/articles/hybrid-search/)**:** combining dense \+ sparse vectors
* [**Advanced filtering**](https://qdrant.tech/documentation/concepts/filtering/)**:** semantic \+ metadata \+ keyword constraints
* [**Advanced filtering**](https://qdrant.tech/documentation/search/filtering/)**:** semantic \+ metadata \+ keyword constraints
* [**Millisecond vector retrieval**](https://qdrant.tech/articles/vector-search-production/)**:** Fast retrieval at \>billion vector scale
Throughout this article, we’ll use TripAdvisor’s [TripBuilder](https://www.tripadvisor.com/TripBuilder) to illustrate how each of the four concepts above is critical for building and deploying agentic search at scale.
@@ -55,7 +55,7 @@ When your agent has to design the itinerary for the trip to Berlin, TripBuilder
## Why Speed Matters in Agentic Retrieval
Agents may run thousands of searches to answer a single complex question. Sometimes those searches are run in parallel, but sometimes they’re run sequentially. For those sequential searches, each millisecond saved compounds. That’s when Qdrant's [Rust-based](https://qdrant.tech/articles/why-rust/) [HNSW indexing](https://qdrant.tech/documentation/concepts/indexing/#hnsw-graph-index), which enables millisecond-level retrieval at scale, becomes a differentiator.
Agents may run thousands of searches to answer a single complex question. Sometimes those searches are run in parallel, but sometimes they’re run sequentially. For those sequential searches, each millisecond saved compounds. That’s when Qdrant's [Rust-based](https://qdrant.tech/articles/why-rust/) [HNSW indexing](https://qdrant.tech/documentation/manage-data/indexing/#hnsw-graph-index), which enables millisecond-level retrieval at scale, becomes a differentiator.
Let’s imagine that you see a 75ms retrieval performance gain with Qdrant. On its own, that might not make a demonstrable difference in the search experience. But if an agent performs four sequential searches, that 300ms starts to have a user experience impact. With that delay, a user can quickly get bored or distracted, leading to frustration. With agentic AI, the retrieval speed optimizations are compounded.
@@ -81,7 +81,7 @@ Combining semantic search, metadata, and keyword [filters](https://qdrant.tech/a
## Real-Time Memory Layer for Agents
Users expect agents to remember the details of their conversation. Your agent's memory must be updated every time it gets new information, not just at the start of each session. Qdrant supports [real-time upserts](https://qdrant.tech/documentation/concepts/points/#upsert-points), giving your agent access to the freshest data for short-term memory and long-term memory.
Users expect agents to remember the details of their conversation. Your agent's memory must be updated every time it gets new information, not just at the start of each session. Qdrant supports [real-time upserts](https://qdrant.tech/documentation/manage-data/points/#upsert-points), giving your agent access to the freshest data for short-term memory and long-term memory.
Not all information is timeless. For an agent, knowing what's recent is as important as knowing what's relevant. This is where [decay functions](https://qdrant.tech/blog/decay-functions/) come in, acting as a "recency boost" during a search.
@@ -109,13 +109,13 @@ Once you’ve built a fast, accurate, secure, and scalable agent, how do you kno
Your agent’s ability to complete complex tasks is only as good as the context it can retrieve. It is crucial to closely and continuously monitor the agent’s performance with grounding checks to spot and prevent hallucinations, recall@k to ensure relevance in search, and MMR parameters for diversity.
#### [**Performance-Cost Tradeoff**](https://qdrant.tech/documentation/guides/optimize/)
#### [**Performance-Cost Tradeoff**](https://qdrant.tech/documentation/operations/optimize/)
In production agents, efficiency is one of, if not the, most important metric to track. First, in enterprise environments you must meet strict latency budgets. Evaluations also let you track the cost per task by tracking token usage and end-to-end compute time. Finally, you can track the effectiveness of your memory layer by monitoring cache hit rates for your memory banks.
![Tradeoff Triangle](/articles_data/agentic-builders-guide/tradeoff-triangle.png)
#### [**Guardrails & Fallbacks**](https://qdrant.tech/documentation/guides/security/)
#### [**Guardrails & Fallbacks**](https://qdrant.tech/documentation/operations/security/)
Just like humans, agents aren’t perfect. A production agentic system should expect and anticipate failures and have guardrails to handle them gracefully. You can also use a human in the loop when confidence scores are below a chosen threshold or the query touches on a high-stakes or sensitive topic. To handle a wide range of queries, your agent should use hybrid search. Hybrid search combines the power of semantic search for understanding meaning with the precision of keyword search for exact matches, giving you the best of both worlds.
@@ -125,13 +125,13 @@ The same agent that speeds through a toy dataset with 10,000 points will become
We’ll talk about three concepts you can take advantage of to improve your scale, but if you want even more information on how to scale, check out this [article](https://qdrant.tech/documentation/database-tutorials/large-scale-search/) on large scale search.
As your dataset and traffic grow, Qdrant Cloud offers a suite of features to ensure your system can scale effectively. Horizontal scaling is achieved through [sharding](https://qdrant.tech/articles/multitenancy/), which splits your collection across multiple nodes to distribute the load and improve performance. For high availability and fault tolerance, Qdrant supports [replication](https://qdrant.tech/documentation/guides/distributed_deployment/), creating copies of your shards across the cluster.
As your dataset and traffic grow, Qdrant Cloud offers a suite of features to ensure your system can scale effectively. Horizontal scaling is achieved through [sharding](https://qdrant.tech/articles/multitenancy/), which splits your collection across multiple nodes to distribute the load and improve performance. For high availability and fault tolerance, Qdrant supports [replication](https://qdrant.tech/documentation/operations/distributed_deployment/), creating copies of your shards across the cluster.
Qdrant provides robust tools for resource and cost optimization. Vector [quantization](https://qdrant.tech/documentation/guides/quantization/) compresses your data, significantly reducing its memory footprint and speeding up search.
Qdrant provides robust tools for resource and cost optimization. Vector [quantization](https://qdrant.tech/documentation/manage-data/quantization/) compresses your data, significantly reducing its memory footprint and speeding up search.
![Quantization](/articles_data/agentic-builders-guide/quantization.png)
To further manage costs as your dataset expands, [on-disk storage](https://qdrant.tech/documentation/concepts/storage/) allows you to keep the full vectors on more affordable SSDs while the necessary index data remains in RAM. This hybrid approach enables searches over billions of vectors without the high cost of keeping all data in memory.
To further manage costs as your dataset expands, [on-disk storage](https://qdrant.tech/documentation/manage-data/storage/) allows you to keep the full vectors on more affordable SSDs while the necessary index data remains in RAM. This hybrid approach enables searches over billions of vectors without the high cost of keeping all data in memory.
To enhance the relevance and precision of search queries, Qdrant natively supports hybrid search, which combines traditional keyword-based search with semantic vector search. By doing so, your application can find documents that match exact terms, such as product codes or names, while also discovering documents that are semantically similar in meaning. This ensures that you can find the most relevant information, even in massive and complex datasets.
@@ -143,11 +143,11 @@ Workflows that are effective for a single user or a handful of users in a develo
Qdrant provides production-grade [authorization and authentication](https://qdrant.tech/documentation/cloud/authentication/) via API keys, multitenancy, and Role-Based Access Control (RBAC) to make sure your agent doesn’t go rogue.
Authentication for agents is handled by [API keys](https://qdrant.tech/documentation/guides/security/#api-keys). With Qdrant, API keys are more than just a password. They act as smart credentials that also carry the details of the authorization rules that are enforced once the agent’s identity is confirmed. These keys can be dynamically created as temporary credentials for each user session, making access limited to the session and secure.
Authentication for agents is handled by [API keys](https://qdrant.tech/documentation/operations/security/#api-keys). With Qdrant, API keys are more than just a password. They act as smart credentials that also carry the details of the authorization rules that are enforced once the agent’s identity is confirmed. These keys can be dynamically created as temporary credentials for each user session, making access limited to the session and secure.
Note: Qdrant also supports concurrent queries, so your search won’t slow down as more users are writing queries simultaneously.
Authorization is handled by [RBAC](https://qdrant.tech/articles/data-privacy/) and [multitenancy](https://qdrant.tech/documentation/guides/multiple-partitions/), which work hand-in-hand to define and enforce permissions specific to the agent. RBAC is a set of rules that defines the allowed permissions inlcuding read-only, read-write, and admin controls. It answers the question, “What is this agent allowed to do?” For instance, can it only search for hotels (read-only), or can it also add, update, and delete them (read-write)?
Authorization is handled by [RBAC](https://qdrant.tech/articles/data-privacy/) and [multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/), which work hand-in-hand to define and enforce permissions specific to the agent. RBAC is a set of rules that defines the allowed permissions including read-only, read-write, and admin controls. It answers the question, “What is this agent allowed to do?” For instance, can it only search for hotels (read-only), or can it also add, update, and delete them (read-write)?
![Multi-tenancy](/articles_data/agentic-builders-guide/multi-tenancy.png)
@@ -212,7 +212,7 @@ Some of the key concepts of CrewAI include:
Qdrant comes into play, as it might be used as a long-term memory layer.**
CrewAI provides a rich set of tools integrated into the framework. That may be a huge advantage for those who want to
combine RAG with e.g. code execution, or image generation. The ecosystem is rich, however brining your own tools is
combine RAG with e.g. code execution, or image generation. The ecosystem is rich, however bringing your own tools is
not a big deal, as CrewAI is designed to be extensible.
A simple agentic RAG application implemented in CrewAI could look like this:
@@ -216,6 +216,6 @@ We recommend the following best practices for leveraging Binary Quantization to
Binary quantization is exceptional if you need to work with large volumes of data under high recall expectations. You can try this feature either by spinning up a [Qdrant container image](https://hub.docker.com/r/qdrant/qdrant) locally or, having us create one for you through a [free account](https://cloud.qdrant.io/login) in our cloud hosted service.
The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials/bulk-upload/) to your Qdrant instance as well as [more quantization methods](/documentation/guides/quantization/).
The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials-develop/bulk-upload/) to your Qdrant instance as well as [more quantization methods](/documentation/manage-data/quantization/).
Want to discuss these findings and learn more about Binary Quantization? [Join our Discord community.](https://discord.gg/qdrant)
@@ -231,6 +231,6 @@ If you determine that binary quantization is appropriate for your datasets and q
Binary quantization is exceptional if you need to work with large volumes of data under high recall expectations. You can try this feature either by spinning up a [Qdrant container image](https://hub.docker.com/r/qdrant/qdrant) locally or, having us create one for you through a [free account](https://cloud.qdrant.io/signup) in our cloud hosted service.
The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials/bulk-upload/) to your Qdrant instance as well as [more quantization methods](/documentation/guides/quantization/).
The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials-develop/bulk-upload/) to your Qdrant instance as well as [more quantization methods](/documentation/manage-data/quantization/).
If you have any feedback, drop us a note on Twitter or LinkedIn to tell us about your results. [Join our lively Discord Server](https://discord.gg/Qy6HCJK9Dc) if you want to discuss BQ with like-minded people!
@@ -34,7 +34,7 @@ However, similarity learning comes with its own difficulties such as:
Quaterion is a fine tuning framework built to tackle such problems in similarity learning.
It uses [PyTorch Lightning](https://www.pytorchlightning.ai/)
as a backend, which is advertized with the motto, "spend more time on research, less on engineering."
as a backend, which is advertised with the motto, "spend more time on research, less on engineering."
This is also true for Quaterion, and it includes:
1. Trainable and servable model classes,
@@ -55,7 +55,7 @@ On Qdrant Cloud, you can create API keys using the [Cloud Dashboard](https://qdr
For on-premise or local deployments, you'll need to configure API key authentication. This involves specifying a key in either the Qdrant configuration file or as an environment variable. This ensures that all requests to the server must include a valid API key sent in the header.
When using the simple API key-based authentication, you should also turn on TLS encryption. Otherwise, you are exposing the connection to sniffing and MitM attacks. To secure your connection using TLS, you would need to create a certificate and private key, and then [enable TLS](/documentation/guides/security/#tls) in the configuration.
When using the simple API key-based authentication, you should also turn on TLS encryption. Otherwise, you are exposing the connection to sniffing and MitM attacks. To secure your connection using TLS, you would need to create a certificate and private key, and then [enable TLS](/documentation/operations/security/#tls) in the configuration.
API authentication, coupled with TLS encryption, offers a first layer of security for your Qdrant instance. However, to enable more granular access control, the recommended approach is to leverage JSON Web Tokens (JWTs).
@@ -107,7 +107,7 @@ In practice, this means that your main database becomes burdened with high memor
Fortunately, the data synchronization problem is not new and definitely not unique to vector search.
There are many well-known solutions, starting with message queues and ending with specialized ETL tools.
For example, we recently released our [integration with Airbyte](/documentation/integrations/airbyte/), allowing you to synchronize data from various sources into Qdrant incrementally.
For example, we recently released our [integration with Airbyte](/documentation/data-management/airbyte/), allowing you to synchronize data from various sources into Qdrant incrementally.
###### You have to pay for a vector service uptime and data transfer of both solutions.
@@ -115,7 +115,7 @@ In the open-source world, you pay for the resources you use, not the number of d
Resources depend more on the optimal solution for each use case.
As a result, running a dedicated vector search engine can be even cheaper, as it allows optimization specifically for vector search use cases.
For instance, Qdrant implements a number of [quantization techniques](/documentation/guides/quantization/) that can significantly reduce the memory footprint of embeddings.
For instance, Qdrant implements a number of [quantization techniques](/documentation/manage-data/quantization/) that can significantly reduce the memory footprint of embeddings.
In terms of data transfer costs, on most cloud providers, network use within a region is usually free. As long as you put the original source data and the vector store in the same region, there are no added data transfer costs.
@@ -22,9 +22,9 @@ In this article, we will describe the unique challenges vector search poses and
## Vectors
![vectors](/articles_data/dedicated-vector-search/image1.jpg)
Let's look at the central concept of vector databases — [**vectors**](/documentation/concepts/vectors/).
Let's look at the central concept of vector databases — [**vectors**](/documentation/manage-data/vectors/).
Vectors (also known as embeddings) are high-dimensional representations of various data points — texts, images, videos, etc. Many state-of-the-art (SOTA) embedding models generate representations of over 1,500 dimensions. When it comes to state-of-the-art PDF retrieval, the representations can reach [**over 100,000 dimensions per page**](/documentation/advanced-tutorials/pdf-retrieval-at-scale/).
Vectors (also known as embeddings) are high-dimensional representations of various data points — texts, images, videos, etc. Many state-of-the-art (SOTA) embedding models generate representations of over 1,500 dimensions. When it comes to state-of-the-art PDF retrieval, the representations can reach [**over 100,000 dimensions per page**](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/).
This brings us to the first challenge of vector search — vectors are heavy.
@@ -52,7 +52,7 @@ However, vectors have positive properties as well. One of the most important is
Embedding models are designed to produce vectors of a fixed size. We have to use it to our advantage.
For fast search, vectors need to be instantly accessible. Whether in [**RAM or disk**](/documentation/concepts/storage/), vectors should be stored in a format that allows quick access and comparison. This is essential, as vector comparison is a very hot operation in vector search workloads. It is often performed thousands of times per search query, so even a small overhead can lead to a significant slowdown.
For fast search, vectors need to be instantly accessible. Whether in [**RAM or disk**](/documentation/manage-data/storage/), vectors should be stored in a format that allows quick access and comparison. This is essential, as vector comparison is a very hot operation in vector search workloads. It is often performed thousands of times per search query, so even a small overhead can lead to a significant slowdown.
For dedicated storage, vectors' fixed size comes as a blessing. Knowing how much space one data point needs, we don't have to deal with the usual overhead of locating data — the location of elements in storage is straightforward to calculate.
@@ -103,7 +103,7 @@ A strictly consistent transactional approach also loses its attractiveness when
## Vector Index
![vector-index](/articles_data/dedicated-vector-search/image3.jpg)
[**Vector search**](/documentation/concepts/search/) relies on high-dimensional vector mathematics, making it computationally heavy at scale. A brute-force similarity search would require comparing a query against every vector in the database. In a database with 100 million 1536-dimensional vectors, performing 100 million comparisons per one query is unfeasible for production scenarios. Instead of a brute-force approach, vector databases have specialized approximate nearest neighbour (ANN) indexes that balance search precision and speed. These indexes require carefully designed architectures to make their maintenance in production feasible.
[**Vector search**](/documentation/search/search/) relies on high-dimensional vector mathematics, making it computationally heavy at scale. A brute-force similarity search would require comparing a query against every vector in the database. In a database with 100 million 1536-dimensional vectors, performing 100 million comparisons per one query is unfeasible for production scenarios. Instead of a brute-force approach, vector databases have specialized approximate nearest neighbour (ANN) indexes that balance search precision and speed. These indexes require carefully designed architectures to make their maintenance in production feasible.
{{< figure src=/articles_data/dedicated-vector-search/hnsw.png caption="HNSW Index" width=80% >}}
@@ -111,7 +111,7 @@ One of the most popular vector indexes is **HNSW (Hierarchical Navigable Small W
### Index Complexity
[**HNSW**](/documentation/concepts/indexing/) is structured as a multi-layered graph. With a new data point inserted, the algorithm must compare it to existing nodes across several layers to index it. As the number of vectors grows, these comparisons will noticeably slow down the construction process, making updates increasingly time-consuming. The indexing operation can quickly become the bottleneck in the system, slowing down search requests.
[**HNSW**](/documentation/manage-data/indexing/) is structured as a multi-layered graph. With a new data point inserted, the algorithm must compare it to existing nodes across several layers to index it. As the number of vectors grows, these comparisons will noticeably slow down the construction process, making updates increasingly time-consuming. The indexing operation can quickly become the bottleneck in the system, slowing down search requests.
Building an HNSW monolith means limiting the scalability of your solution — its size has to be capped, as its construction time scales **non-linearly** with the number of elements. To keep the construction process feasible and ensure it doesn't affect the search time, we came up with a layered architecture that breaks down all data management into small units called **segments**.
@@ -128,7 +128,7 @@ With index maintenance divided between segments, Qdrant can ensure high performa
| | |
|---------------------|-------------|
| **Mutable Segments** | These are used for quickly ingesting new data and handling changes (updates) to existing data. |
| **Immutable Segments** | Once a mutable segment reaches a certain size, an optimization process converts it into an immutable segment, constructing an HNSW index – you could [**read about these optimizers here**](/documentation/concepts/optimizer/#optimizer) in detail. This immutability trick allowed us, for example, to ensure effective [**tenant isolation**](/documentation/concepts/indexing/#tenant-index). |
| **Immutable Segments** | Once a mutable segment reaches a certain size, an optimization process converts it into an immutable segment, constructing an HNSW index – you could [**read about these optimizers here**](/documentation/operations/optimizer/#optimizer) in detail. This immutability trick allowed us, for example, to ensure effective [**tenant isolation**](/documentation/manage-data/indexing/#tenant-index). |
Immutable segments are an implementation detail transparent for users — they can delete vectors at any time, while additions and updates are applied to a mutable segment instead. This combination of mutability and immutability allows search and indexing to smoothly run simultaneously, even under heavy loads. This approach minimizes the performance impact of indexing time and allows on-the-fly configuration changes on a collection level (such as enabling or disabling data quantization) without downtimes.
@@ -145,7 +145,7 @@ In many vector search solutions, filtering is approached in two ways: **pre-filt
| ❌ | **Pre-filtering** | Has the linear complexity of computing the vector mask and becomes a bottleneck for large datasets. |
| ❌ | **Post-filtering** | The problem with **post-filtering** is tied to vector search "*everything fits and doesn't at the same time*" nature: imagine a low-cardinality filter that leaves only a few matching elements in the database. If none of them are similar enough to the query to appear in the top-X retrieved results, they'll all be filtered out. |
Qdrant [**took filtering in vector search further**](/articles/vector-search-filtering/), recognizing the limitations of pre-filtering & post-filtering strategies. We developed an adaptation of HNSW — [**filterable HNSW**](/articles/filterable-hnsw/) — that also enables **in-place filtering** during graph traversal. To make this possible, we condition HNSW index construction on possible filtering conditions reflected by [**payload indexes**](/documentation/concepts/indexing/#payload-index) (inverted indexes built on vectors' [**metadata**](/documentation/concepts/payload/)).
Qdrant [**took filtering in vector search further**](/articles/vector-search-filtering/), recognizing the limitations of pre-filtering & post-filtering strategies. We developed an adaptation of HNSW — [**filterable HNSW**](/articles/filterable-hnsw/) — that also enables **in-place filtering** during graph traversal. To make this possible, we condition HNSW index construction on possible filtering conditions reflected by [**payload indexes**](/documentation/manage-data/indexing/#payload-index) (inverted indexes built on vectors' [**metadata**](/documentation/manage-data/payload/)).
**Qdrant was designed with a vector index being a central component of the system.** That made it possible to organize optimizers, payload indexes and other components around the vector index, unlocking the possibility of building a filterable HNSW.
@@ -165,7 +165,7 @@ The strength of vector search lies in its ability to facilitate [**discovery**](
### Recommendations
Vector search is perfect for [**recommendations**](/documentation/concepts/explore/#recommendation-api). Imagine browsing for a new book or movie. Instead of searching for an exact match, you might look for stories that capture a certain mood or theme but differ in key aspects from what you already know. For example, you may [**want a film featuring wizards without the familiar feel of the "Harry Potter" series**](https://www.youtube.com/watch?v=O5mT8M7rqQQ). This flexibility is possible because vector search is not tied to the binary "match/not match" concept but operates on distances in a vector space.
Vector search is perfect for [**recommendations**](/documentation/search/explore/#recommendation-api). Imagine browsing for a new book or movie. Instead of searching for an exact match, you might look for stories that capture a certain mood or theme but differ in key aspects from what you already know. For example, you may [**want a film featuring wizards without the familiar feel of the "Harry Potter" series**](https://www.youtube.com/watch?v=O5mT8M7rqQQ). This flexibility is possible because vector search is not tied to the binary "match/not match" concept but operates on distances in a vector space.
### Big Unstructured Data Analysis
@@ -193,17 +193,17 @@ Consider some of the advanced features implemented in Qdrant:
GPU acceleration in Qdrant is a custom solution developed by an enthusiast from our core team. It's vendor-free and natively supports all Qdrant's unique architectural features, from FIlterable HNSW to multivectors.
- [**Multivectors**](/documentation/concepts/vectors/?q=multivectors#multivectors)
- [**Multivectors**](/documentation/manage-data/vectors/?q=multivectors#multivectors)
Some modern embedding models produce an entire matrix (a list of vectors) as output rather than a single vector. Qdrant supports multivectors natively.
This feature is critical when using state-of-the-art retrieval models such as [**ColBERT**](/documentation/fastembed/fastembed-colbert/), ColPali, or ColQwen. For instance, ColPali and ColQwen produce multivector outputs, and supporting them natively is crucial for [**state-of-the-art (SOTA) PDF-retrieval**](/documentation/advanced-tutorials/pdf-retrieval-at-scale/).
This feature is critical when using state-of-the-art retrieval models such as [**ColBERT**](/documentation/fastembed/fastembed-colbert/), ColPali, or ColQwen. For instance, ColPali and ColQwen produce multivector outputs, and supporting them natively is crucial for [**state-of-the-art (SOTA) PDF-retrieval**](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/).
In addition to that, we continuously look for improvements in:
| | |
|----------------------------------|-------------|
| **Memory Efficiency & Compression** | Techniques such as [**quantization**](/documentation/guides/quantization/) and [**HNSW compression**](/blog/qdrant-1.13.x/#hnsw-graph-compression) to reduce storage requirements |
| **Retrieval Algorithms** | Support for the latest retrieval algorithms, including [**sparse neural retrieval**](/articles/modern-sparse-neural-retrieval/), [**hybrid search**](/documentation/concepts/hybrid-queries/) methods, and [**re-rankers**](/documentation/fastembed/fastembed-rerankers/). |
| **Memory Efficiency & Compression** | Techniques such as [**quantization**](/documentation/manage-data/quantization/) and [**HNSW compression**](/blog/qdrant-1.13.x/#hnsw-graph-compression) to reduce storage requirements |
| **Retrieval Algorithms** | Support for the latest retrieval algorithms, including [**sparse neural retrieval**](/articles/modern-sparse-neural-retrieval/), [**hybrid search**](/documentation/search/hybrid-queries/) methods, and [**re-rankers**](/documentation/fastembed/fastembed-rerankers/). |
| **Vector Data Analysis & Visualization** | Tools like the [**distance matrix API**](/blog/qdrant-1.12.x/#distance-matrix-api-for-data-insights) provide insights into vectorized data, and a [**Web UI**](/blog/qdrant-1.11.x/#web-ui-search-quality-tool) allows for intuitive exploration of data. |
| **Search Speed & Scalability** | Includes optimizations for [**multi-tenant environments**](/articles/multitenancy/) to ensure efficient and scalable search. |
@@ -38,7 +38,7 @@ This is where a __vector _context___ can help. We define _context_ as a list of
![Discovery search visualization](/articles_data/discovery-search/discovery-search.png)
While positive and negative vectors might suggest the use of the <a href="/documentation/concepts/explore/#recommendation-api" target="_blank">recommendation interface</a>, in the case of _context_ they require to be paired up in a positive-negative fashion. This is inspired from the machine-learning concept of <a href="https://en.wikipedia.org/wiki/Triplet_loss" target="_blank">_triplet loss_</a>, where you have three vectors: an anchor, a positive, and a negative. Triplet loss is an evaluation of how much the anchor is closer to the positive than to the negative vector, so that learning happens by "moving" the positive and negative points to try to get a better evaluation. However, during discovery, we consider the positive and negative vectors as static points, and we search through the whole dataset for the "anchors", or result candidates, which fit this characteristic better.
While positive and negative vectors might suggest the use of the <a href="/documentation/search/explore/#recommendation-api" target="_blank">recommendation interface</a>, in the case of _context_ they require to be paired up in a positive-negative fashion. This is inspired from the machine-learning concept of <a href="https://en.wikipedia.org/wiki/Triplet_loss" target="_blank">_triplet loss_</a>, where you have three vectors: an anchor, a positive, and a negative. Triplet loss is an evaluation of how much the anchor is closer to the positive than to the negative vector, so that learning happens by "moving" the positive and negative points to try to get a better evaluation. However, during discovery, we consider the positive and negative vectors as static points, and we search through the whole dataset for the "anchors", or result candidates, which fit this characteristic better.
![Triplet loss](/articles_data/discovery-search/triplet-loss.png)
@@ -100,4 +100,4 @@ This way you can give refreshing recommendations, while still being in control b
- Discovery search is a powerful tool for controlled exploration in vector spaces.
Context, consisting of positive and negative vectors constrain the search space, while a target guides the search.
- Real-world applications include multimodal search, diverse recommendations, and context-driven exploration.
- Ready to learn more about the math behind it and how to use it? Check out the [documentation](/documentation/concepts/explore/#discovery-api)
- Ready to learn more about the math behind it and how to use it? Check out the [documentation](/documentation/search/explore/#discovery-api)
@@ -11,7 +11,7 @@ draft: false
keywords:
- clusterization
- dimensionality reduction
- vizualization
- visualization
category: data-exploration
---
@@ -25,8 +25,8 @@ Examining data points individually is not always the best way to grasp the struc
As numbers in a table obtain meaning when plotted on a graph, visualising distances (similar/dissimilar) between unstructured data items can reveal hidden structures and patterns.
{{< figure src="/articles_data/distance-based-exploration/data-on-chart.png" alt="Data visualization" caption="Vizualized chart, very intuitive" >}}
There are many tools to investigate data similarity, and Qdrant's [1.12 release](https://qdrant.tech/blog/qdrant-1.12.x/) made it much easier to start this investigation. With the new [Distance Matrix API](/documentation/concepts/explore/#distance-matrix), Qdrant handles the most computationally expensive part of the process—calculating the distances between data points.
{{< figure src="/articles_data/distance-based-exploration/data-on-chart.png" alt="Data visualization" caption="Visualized chart, very intuitive" >}}
There are many tools to investigate data similarity, and Qdrant's [1.12 release](https://qdrant.tech/blog/qdrant-1.12.x/) made it much easier to start this investigation. With the new [Distance Matrix API](/documentation/search/explore/#distance-matrix), Qdrant handles the most computationally expensive part of the process—calculating the distances between data points.
In many implementations, the distance matrix calculation was part of the clustering or visualization processes, requiring either brute-force computation or building a temporary index. With Qdrant, however, the data is already indexed, and the distance matrix can be computed relatively cheaply.
@@ -62,7 +62,7 @@ PUT /collections/midlib/snapshots/recover
```
<details>
<summary>We also need to prepare our python enviroment:</summary>
<summary>We also need to prepare our python environment:</summary>
```bash
pip install umap-learn seaborn matplotlib qdrant-client
@@ -77,7 +77,7 @@ from qdrant_client import QdrantClient
from umap import UMAP
# Python implementation for sparse matrices
from scipy.sparse import csr_matrix
# For vizualization
# For visualization
import seaborn as sns
```
@@ -55,7 +55,7 @@ As embeddings are vectors, one can apply a simple function to calculate the simi
So with similarity learning, all we need to do is provide pairs of correct questions and answers.
And then, the model will learn to distinguish proper answers by the similarity of embeddings.
>If you want to learn more about similarity learning and applications, check out this [article](/documentation/tutorials/neural-search/) which might be an asset.
>If you want to learn more about similarity learning and applications, check out this [article](/documentation/tutorials-search-engineering/neural-search/) which might be an asset.
## Let's build
+1 -1
View File
@@ -238,7 +238,7 @@ If you're curious about how FastEmbed and Qdrant can make your search tasks a br
1. **Cloud**: Get started with a free plan on the [Qdrant Cloud](https://qdrant.to/cloud?utm_source=qdrant&utm_medium=website&utm_campaign=fastembed&utm_content=article).
2. **Docker Container**: If you're the DIY type, you can set everything up on your own machine. Here's a quick guide to help you out: [Quick Start with Docker](/documentation/quick-start/?utm_source=qdrant&utm_medium=website&utm_campaign=fastembed&utm_content=article).
2. **Docker Container**: If you're the DIY type, you can set everything up on your own machine. Here's a quick guide to help you out: [Quick Start with Docker](/documentation/quickstart/?utm_source=qdrant&utm_medium=website&utm_campaign=fastembed&utm_content=article).
So, go ahead, take it for a test drive. We're excited to hear what you think!
@@ -1,7 +1,7 @@
---
title: Filterable HNSW
short_description: How to make ANN search with custom filtering?
description: How to make ANN search with custom filtering? Search in selected subsets without loosing the results.
description: How to make ANN search with custom filtering? Search in selected subsets without losing the results.
# external_link: https://blog.vasnetsov.com/posts/categorical-hnsw/
social_preview_image: /articles_data/filterable-hnsw/social_preview.jpg
preview_dir: /articles_data/filterable-hnsw/preview
@@ -27,7 +27,7 @@ Otherwise, read on to learn more about the demo and how it works!
In general, our application consists of three parts: a [FastAPI](https://fastapi.tiangolo.com/) backend, a [React](https://react.dev/) frontend, and
a [Qdrant](/) instance. The architecture diagram below shows how these components interact with each other:
![Archtecture diagram](/articles_data/food-discovery-demo/architecture-diagram.png)
![Architecture diagram](/articles_data/food-discovery-demo/architecture-diagram.png)
## Why did we use a CLIP model?
@@ -92,7 +92,7 @@ in the vector space.
![Random points selection](/articles_data/food-discovery-demo/textual-search.png)
This is implemented as [a group search query to Qdrant](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L44).
We didn't use a simple search, but performed grouping by the restaurant to get more diverse results. [Search groups](/documentation/concepts/search/#search-groups)
We didn't use a simple search, but performed grouping by the restaurant to get more diverse results. [Search groups](/documentation/search/search/#search-groups)
is a mechanism similar to `GROUP BY` clause in SQL, and it's useful when you want to get a specific number of result per group (in our case just one).
```python
@@ -121,7 +121,7 @@ and the demo will update the search results accordingly.
#### Negative feedback only
Qdrant [Recommendation API](/documentation/concepts/search/#recommendation-api) needs at least one positive example to work. However, in our demo
Qdrant [Recommendation API](/documentation/search/search/#recommendation-api) needs at least one positive example to work. However, in our demo
we want to be able to provide only negative examples. This is because we want to be able to say “I don’t like this dish” without having to like anything first.
To achieve this, we use a trick. We negate the vectors of the disliked dishes and use their mean as a query. This way, the disliked dishes will be pushed away
from the search results. **This works because the cosine distance is based on the angle between two vectors, and the angle between a vector and its negation is 180 degrees.**
@@ -129,8 +129,8 @@ from the search results. **This works because the cosine distance is based on th
![CLIP model](/articles_data/food-discovery-demo/negated-vector.png)
Food Discovery Demo [implements that trick](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L122)
by calling Qdrant twice. Initially, we use the [Scroll API](/documentation/concepts/points/#scroll-points) to find disliked items,
and then calculate a negated mean of all their vectors. That allows using the [Search Groups API](/documentation/concepts/search/#search-groups)
by calling Qdrant twice. Initially, we use the [Scroll API](/documentation/manage-data/points/#scroll-points) to find disliked items,
and then calculate a negated mean of all their vectors. That allows using the [Search Groups API](/documentation/search/search/#search-groups)
to find the nearest neighbors of the negated mean vector.
```python
@@ -163,7 +163,7 @@ response = client.search_groups(
#### Positive and negative feedback
Since the [Recommendation API](/documentation/concepts/search/#recommendation-api) requires at least one positive example, we can use it only when
Since the [Recommendation API](/documentation/search/search/#recommendation-api) requires at least one positive example, we can use it only when
the user has liked at least one dish. We could theoretically use the same trick as above and negate the disliked dishes, but it would be a bit weird, as Qdrant has
that feature already built-in, and we can call it just once to do the job. It's always better to perform the search server-side. Thus, in this case [we just call
the Qdrant server with a list of positive and negative examples](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L166),
@@ -185,7 +185,7 @@ From the user perspective nothing changes comparing to the previous case.
Last but not least, location plays an important role in the food discovery process. You are definitely looking for something you can find nearby, not on the other
side of the globe. Therefore, your current location can be toggled as a filtering condition. You can enable it by clicking on “Find near me” icon
in the top right. This way you can find the best pizza in your neighborhood, not in the whole world. Qdrant [geo radius filter](/documentation/concepts/filtering/#geo-radius) is a perfect choice for this. It lets you
in the top right. This way you can find the best pizza in your neighborhood, not in the whole world. Qdrant [geo radius filter](/documentation/search/filtering/#geo-radius) is a perfect choice for this. It lets you
filter the results by distance from a given point.
```python
@@ -208,7 +208,7 @@ query_filter = models.Filter(
)
```
Such a filter needs [a payload index](/documentation/concepts/indexing/#payload-index) to work efficiently, and it was created on a collection
Such a filter needs [a payload index](/documentation/manage-data/indexing/#payload-index) to work efficiently, and it was created on a collection
we used to create the snapshot. When you import it into your instance, the index will be already there.
## Using the demo
@@ -224,7 +224,7 @@ docker-compose up -d
```
The demo will be available at `http://localhost:8001`, but you won't be able to search anything until you [import the snapshot into your Qdrant
instance](/documentation/concepts/snapshots/#recover-via-api). If you don't want to bother with hosting a local one, you can use the [Qdrant
instance](/documentation/operations/snapshots/#recover-via-api). If you don't want to bother with hosting a local one, you can use the [Qdrant
Cloud](https://cloud.qdrant.io/) cluster. 4 GB RAM is enough to load all the 2 million entries.
## Fork and reuse
@@ -75,4 +75,4 @@ Being selected for Google Summer of Code 2023 and collaborating with Arnaud and
Without a doubt, I'm eager to continue growing alongside this community and contribute to new features and enhancements that elevate the product. I've also become an advocate for Qdrant, introducing this project to numerous coworkers and friends in the tech industry. I'm excited to witness new users and contributors emerge from within my own network!
If you want to try out my work, read the [documentation](/documentation/concepts/filtering/#geo-polygon) and then, either sign up for a free [cloud account](https://cloud.qdrant.io) or download the [Docker image](https://hub.docker.com/r/qdrant/qdrant). I look forward to seeing how people are using my work in their own applications!
If you want to try out my work, read the [documentation](/documentation/search/filtering/#geo-polygon) and then, either sign up for a free [cloud account](https://cloud.qdrant.io) or download the [Docker image](https://hub.docker.com/r/qdrant/qdrant). I look forward to seeing how people are using my work in their own applications!
@@ -97,7 +97,7 @@ Real-world systems don’t operate in a vacuum. Failures happen: software bugs c
If one component is updated but another isn’t, the entire system could become inconsistent. Worse, if an operation is only partially written to disk, it could lead to orphaned data, unusable space, or even data corruption.
### Stability Through Idempotency: Recovering With WAL
To guard against these risks, Qdrant relies on a [**Write-Ahead Log (WAL)**](/documentation/concepts/storage/). Before committing an operation, Qdrant ensures that it is at least recorded in the WAL. If a crash happens before all updates are flushed, the system can safely replay operations from the log.
To guard against these risks, Qdrant relies on a [**Write-Ahead Log (WAL)**](/documentation/manage-data/storage/). Before committing an operation, Qdrant ensures that it is at least recorded in the WAL. If a crash happens before all updates are flushed, the system can safely replay operations from the log.
This recovery mechanism introduces another essential property: [**idempotence**](https://en.wikipedia.org/wiki/Idempotence).
@@ -167,7 +167,7 @@ For even sharper debugging, Property-Based Testing adds automated test generatio
Designing for crash resilience is one thing, and proving it works under stress is another. To push Qdrant’s data integrity to the limit, we built [**Crasher**](https://github.com/qdrant/crasher), a test bench that brutally kills and restarts Qdrant while it handles a heavy update workload.
Crasher runs a loop that continuously writes data, then randomly crashes Qdrant. On each restart, Qdrant replays its [**Write-Ahead Log (WAL)**](/documentation/concepts/storage/), and we verify if data integrity holds. Possible anomalies include:
Crasher runs a loop that continuously writes data, then randomly crashes Qdrant. On each restart, Qdrant replays its [**Write-Ahead Log (WAL)**](/documentation/manage-data/storage/), and we verify if data integrity holds. Possible anomalies include:
- Missing data (points, vectors, or payloads)
- Corrupt payload values
@@ -194,7 +194,7 @@ This shows a clear boost in performance. As we can see, the investment in Gridst
### End-to-End Benchmarking
Now, let’s test the impact on a real Qdrant instance. So far, we’ve only integrated Gridstore for [**payloads**](/documentation/concepts/payload/) and [**sparse vectors**](/documentation/concepts/vectors/#sparse-vectors), but even this partial switch should show noticeable improvements.
Now, let’s test the impact on a real Qdrant instance. So far, we’ve only integrated Gridstore for [**payloads**](/documentation/manage-data/payload/) and [**sparse vectors**](/documentation/manage-data/vectors/#sparse-vectors), but even this partial switch should show noticeable improvements.
For benchmarking, we used our in-house [**bfb tool**](https://github.com/qdrant/bfb) to generate a workload. Our configuration:
@@ -21,7 +21,7 @@ required to combine the results from different methods on your end.
**Qdrant 1.10 introduces a new Query API that lets you build a search system by combining different search methods
to improve retrieval quality**. Everything is now done on the server side, and you can focus on building the best search
experience for your users. In this article, we will show you how to utilize the new [Query
API](/documentation/concepts/search/#query-api) to build a hybrid search system.
API](/documentation/search/search/#query-api) to build a hybrid search system.
## Introducing the new Query API
@@ -169,12 +169,12 @@ If we knew some additional information about the data, we could combine all rele
{{< figure src="/articles_data/immutable-data-structures/defragmentation.png" alt="Defragmentation" caption="Defragmentation" width="70%" >}}
This additional information is available to Qdrant via the [payload index](/documentation/concepts/indexing/#payload-index).
This additional information is available to Qdrant via the [payload index](/documentation/manage-data/indexing/#payload-index).
By specifying the payload index, which is going to be used for filtering most of the time, we can put all vectors with the same payload together.
This way, reading a single page will also read nearby vectors, which will be used in the search.
This approach is especially efficient for [multi-tenant systems](/documentation/guides/multiple-partitions/), where only a small subset of vectors is actively used for search.
This approach is especially efficient for [multi-tenant systems](/documentation/manage-data/multitenancy/), where only a small subset of vectors is actively used for search.
The capacity of such a deployment is typically defined by the size of the hot subset, which is much smaller than the total number of vectors.
> Grouping relevant vectors together allows us to optimize the size of the hot subset by avoiding caching of irrelevant data.
@@ -200,7 +200,7 @@ All benchmarks are made with minimal RAM allocation to demonstrate disk cache ef
As you can see, the biggest impact is on the small tenant size, where defragmentation allows us to achieve **100x more RPS**.
Of course, the real-world impact of defragmentation depends on the specific workload and the size of the hot subset, but enabling this feature can significantly improve the performance of Qdrant.
Please find more details on how to enable defragmentation in the [indexing documentation](/documentation/concepts/indexing/#tenant-index).
Please find more details on how to enable defragmentation in the [indexing documentation](/documentation/manage-data/indexing/#tenant-index).
## Updating Immutable Data Structures
+2 -2
View File
@@ -227,7 +227,7 @@ Here are the specific characteristics of the miniCOIL model we trained based on
| **Input Dense Encoder** | [`jina-embeddings-v2-small-en`](https://huggingface.co/jinaai/jina-embeddings-v2-small-en) (512 dimensions) |
| **miniCOIL Vectors Size** | 4 dimensions |
| **miniCOIL Vocabulary** | List of 30,000 of the most common English words, cleaned of stop words and words shorter than 3 letters, [taken from here](https://github.com/arstgit/high-frequency-vocabulary/tree/master). Words are stemmed to align miniCOIL with our BM25 implementation. |
| **Training Data** | 40 million sentences — a random subset of the [OpenWebText dataset](https://paperswithcode.com/dataset/openwebtext). To make triplet sampling convenient, we uploaded sentences and their [`mxbai-embed-large-v1`](https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1) embeddings to Qdrant and built a [full-text payload index](https://qdrant.tech/documentation/concepts/indexing/#full-text-index) on sentences with a tokenizer of type `word`. |
| **Training Data** | 40 million sentences — a random subset of the [OpenWebText dataset](https://paperswithcode.com/dataset/openwebtext). To make triplet sampling convenient, we uploaded sentences and their [`mxbai-embed-large-v1`](https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1) embeddings to Qdrant and built a [full-text payload index](https://qdrant.tech/documentation/manage-data/indexing/#full-text-index) on sentences with a tokenizer of type `word`. |
| **Training Data per Word** | We sample 8000 sentences per word and form triplets with a margin of at least **0.1**.<br>Additionally, we apply **augmentation** — take a sentence and cut out the target word plus its 1–3 neighbours. We reuse the same similarity score between original and augmented sentences for simplicity. |
| **Training Parameters** | **Epochs**: 60<br>**Optimizer**: Adam with a learning rate of 1e-4<br>**Validation set**: 20% |
@@ -237,7 +237,7 @@ We included this `minicoil-v1` version in the [v0.7.0 release of our FastEmbed l
You can check an example of `minicoil-v1` usage with FastEmbed in the [HuggingFace card](https://huggingface.co/Qdrant/minicoil-v1).
<aside role="status">
To use <span style="font-weight: bold;">minicoil-v1</span> correctly, make sure to configure sparse vectors with <a href="https://qdrant.tech/documentation/concepts/indexing/?q=modifier#idf-modifier">Modifier.IDF</a>
To use <span style="font-weight: bold;">minicoil-v1</span> correctly, make sure to configure sparse vectors with <a href="https://qdrant.tech/documentation/manage-data/indexing/?q=modifier#idf-modifier">Modifier.IDF</a>
</aside>
## Results
@@ -234,7 +234,7 @@ internal document expansion idea, which made the retrieval quality noticeably be
- The SPARTA model is not sparse enough by construction, so authors of the SPLADE family of models introduced explicit **sparsity regularisation**,
preventing the model from producing too many non-zero values.
- The SPARTA model mostly uses the BERT model as-is, without any additional neural network to capture the specifity of Information Retrieval problem,
- The SPARTA model mostly uses the BERT model as-is, without any additional neural network to capture the specificity of Information Retrieval problem,
so SPLADE models introduce a trainable neural network on top of BERT with a specific architecture choice to make it perfectly fit the task.
- SPLADE family of models, finally, uses **knowledge distillation**, which is learning from a bigger
(and therefore much slower, not-so-fit for production tasks) model how to predict good representations.
@@ -397,8 +397,8 @@ from qdrant_client import QdrantClient, models
qdrant_client = QdrantClient(":memory:") # Qdrant is running from RAM.
```
Now, let's create a [collection](https://qdrant.tech/documentation/concepts/collections/) in which could upload our sparse SPLADE++ embeddings. \
For that, we will use the [sparse vectors](https://qdrant.tech/documentation/concepts/vectors/#sparse-vectors) representation supported in Qdrant.
Now, let's create a [collection](https://qdrant.tech/documentation/manage-data/collections/) in which could upload our sparse SPLADE++ embeddings. \
For that, we will use the [sparse vectors](https://qdrant.tech/documentation/manage-data/vectors/#sparse-vectors) representation supported in Qdrant.
```python
qdrant_client.create_collection(
@@ -19,14 +19,14 @@ category: vector-search-manuals
# Scaling Your Machine Learning Setup: The Power of Multitenancy and Custom Sharding in Qdrant
We are seeing the topics of [multitenancy](/documentation/guides/multiple-partitions/) and [distributed deployment](/documentation/guides/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup.
We are seeing the topics of [multitenancy](/documentation/manage-data/multitenancy/) and [distributed deployment](/documentation/operations/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup.
Whether you are building a bank fraud-detection system, [RAG](https://qdrant.tech/articles/what-is-rag-in-ai/) for e-commerce, or services for the federal government - you will need to leverage a multitenant architecture to scale your product.
In the world of SaaS and enterprise apps, this setup is the norm. It will considerably increase your application's performance and lower your hosting costs.
## Multitenancy & custom sharding with Qdrant
We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](/documentation/guides/multiple-partitions/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](/documentation/guides/distributed_deployment/#user-defined-sharding).
We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](/documentation/manage-data/multitenancy/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](/documentation/operations/distributed_deployment/#user-defined-sharding).
Combining these two will result in an efficiently-partitioned architecture that further leverages the convenience of a single Qdrant cluster. This article will briefly explain the benefits and show how you can get started using both features.
@@ -41,7 +41,7 @@ Qdrant is built to excel in a single collection with a vast number of tenants. Y
## Sharding your database
With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](/documentation/guides/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node.
With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](/documentation/operations/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node.
During vector search, your operations will be able to hit only the subset of shards they actually need. In massive-scale deployments, __this can significantly improve the performance of operations that do not require the whole collection to be scanned__.
@@ -49,7 +49,7 @@ This works in the other direction as well. Whenever you search for something, yo
### Common use cases
A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](/documentation/guides/distributed_deployment/#moving-shards).
A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](/documentation/operations/distributed_deployment/#moving-shards).
**Figure 2:** Users can both upsert and query shards that are relevant to them, all within the same collection. Regional sharding can help avoid cross-continental traffic.
![Qdrant Multitenancy](/articles_data/multitenancy/shards.png)
@@ -79,7 +79,7 @@ client.create_shard_key("{tenant_data}", "germany")
```
In this example, your cluster is divided between Germany and Canada. Canadian and German law differ when it comes to international data transfer. Let's say you are creating a RAG application that supports the healthcare industry. Your Canadian customer data will have to be clearly separated for compliance purposes from your German customer.
Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](/documentation/guides/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech).
Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](/documentation/operations/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech).
## Configure a multitenant setup for users
@@ -182,7 +182,7 @@ client.create_payload_index(
## Explore multitenancy and custom sharding in Qdrant for scalable solutions
Qdrant is ready to support a massive-scale architecture for your machine learning project. If you want to see whether our [vector database](https://qdrant.tech/) is right for you, try the [quickstart tutorial](/documentation/quick-start/) or read our [docs and tutorials](/documentation/).
Qdrant is ready to support a massive-scale architecture for your machine learning project. If you want to see whether our [vector database](https://qdrant.tech/) is right for you, try the [quickstart tutorial](/documentation/quickstart/) or read our [docs and tutorials](/documentation/).
To spin up a free instance of Qdrant, sign up for [Qdrant Cloud](https://qdrant.to/cloud) - no strings attached.
@@ -14,7 +14,7 @@ category: vector-search-manuals
Multi-vector representations are superior to single-vector embeddings in many benchmarks. It might be tempting to use
them right away, but there is a catch: they are slower to search. Traditional vector search structures like
[HNSW](/documentation/concepts/indexing/#vector-index) are optimized for retrieving the nearest neighbors of a single
[HNSW](/documentation/manage-data/indexing/#vector-index) are optimized for retrieving the nearest neighbors of a single
query vector using simple metrics such as cosine similarity. These indexes are not suitable for multi-vector retrieval
strategies, such as MaxSim, where a query and document are each represented by multiple vectors and the final score is
computed as the maximum similarity over all cross-pairings. MaxSim is inherently asymmetric and non-metric, so HNSW
@@ -66,7 +66,7 @@ done by computing the dot product of the input vector with each hyperplane norma
result. Since each of our regions can be represented as a binary string of length `k_sim` (where each bit indicates
which side of a hyperplane the vector is on), we can interpret this binary string as an integer to get a cluster ID.
![SimHash cluster assignement](/articles_data/muvera-embeddings/simhash-cluster-assignment.png)
![SimHash cluster assignment](/articles_data/muvera-embeddings/simhash-cluster-assignment.png)
### Fixed Dimensional Encoding (FDE) creation
@@ -17,15 +17,15 @@ does exist, and recommendation systems are a great example. Recommendations migh
to find items close to positive and far from negative examples. This use of vector databases has many applications, including
recommendation systems for e-commerce, content, or even dating apps.
Qdrant has provided the [Recommendation API](/documentation/concepts/search/#recommendation-api) for a while, and with the latest release, [Qdrant 1.6](https://github.com/qdrant/qdrant/releases/tag/v1.6.0),
Qdrant has provided the [Recommendation API](/documentation/search/search/#recommendation-api) for a while, and with the latest release, [Qdrant 1.6](https://github.com/qdrant/qdrant/releases/tag/v1.6.0),
we're glad to give you more flexibility and control over the Recommendation API.
Here, we'll discuss some internals and show how they may be used in practice.
### Recap of the old recommendations API
The previous [Recommendation API](/documentation/concepts/search/#recommendation-api) in Qdrant came with some limitations. First of all, it was required to pass vector IDs for
The previous [Recommendation API](/documentation/search/search/#recommendation-api) in Qdrant came with some limitations. First of all, it was required to pass vector IDs for
both positive and negative example points. If you wanted to use vector embeddings directly, you had to either create a new point
in a collection or mimic the behaviour of the Recommendation API by using the [Search API](/documentation/concepts/search/#search-api).
in a collection or mimic the behaviour of the Recommendation API by using the [Search API](/documentation/search/search/#search-api).
Moreover, in the previous releases of Qdrant, you were always asked to provide at least one positive example. This requirement
was based on the algorithm used to combine multiple samples into a single query vector. It was a simple, yet effective approach.
However, if the only information you had was that your user dislikes some items, you couldn't use it directly.
@@ -149,7 +149,7 @@ If you want to know more about the internals of HNSW, you can check out the arti
## Food Discovery demo
Our [Food Discovery demo](/articles/food-discovery-demo/) is an application built on top of the new [Recommendation API](/documentation/concepts/search/#recommendation-api).
Our [Food Discovery demo](/articles/food-discovery-demo/) is an application built on top of the new [Recommendation API](/documentation/search/search/#recommendation-api).
It allows you to find a meal based on liked and disliked photos. There are some updates, enabled by the new Qdrant release:
* **Ability to include multiple textual queries in the recommendation request.** Previously, we only allowed passing a single
@@ -224,7 +224,7 @@ In circumstances that do not align with the above, Scalar Quantization should be
## Using Qdrant for Product Quantization
If you’re already a Qdrant user, we have, documentation on [Product Quantization](/documentation/guides/quantization/#setting-up-product-quantization) that will help you to set and configure the new quantization for your data and achieve even
If you’re already a Qdrant user, we have, documentation on [Product Quantization](/documentation/manage-data/quantization/#setting-up-product-quantization) that will help you to set and configure the new quantization for your data and achieve even
up to 64x memory reduction.
Ready to experience the power of Product Quantization? [Sign up now](https://cloud.qdrant.io/signup) for a free Qdrant demo and optimize your data management today!
@@ -27,7 +27,7 @@ set up your collections.
Previously, you had to send multiple requests to the Qdrant API to perform multiple non-related tasks. However, this
can cause significant network overhead and slow down the process, especially if you have a poor connection speed.
Fortunately, the [new batch search feature](/documentation/concepts/search/#batch-search-api) allows
Fortunately, the [new batch search feature](/documentation/search/search/#batch-search-api) allows
you to avoid this issue. With just one API call, Qdrant will handle multiple search requests in the most efficient way
possible. This means that you can perform multiple tasks simultaneously without having to worry about network overhead
or slow performance.
@@ -44,6 +44,6 @@ both ARM and non-ARM architectures using similar setups to understand the potent
Qdrant is a vector database that allows you to quickly search for the nearest neighbors. However, you may need to apply
additional filters on top of the semantic search. Up until version 0.10, Qdrant only supported keyword filters. With the
release of Qdrant 0.10, [you can now use full-text filters](/documentation/concepts/filtering/#full-text-match)
release of Qdrant 0.10, [you can now use full-text filters](/documentation/search/filtering/#full-text-match)
as well. This new filter type can be used on its own or in combination with other filter types to provide even more
flexibility in your searches.
@@ -39,7 +39,7 @@ why we have introduced the [Scalar Quantization](/articles/scalar-quantization/)
which makes it possible to reduce the memory requirements by up to four times.
Today, we are bringing a new quantization mechanism to life. A separate article on [Product
Quantization](/documentation/quantization/#product-quantization) will describe that feature in more
Quantization](/documentation/manage-data/quantization/#product-quantization) will describe that feature in more
detail. In a nutshell, you can **reduce the memory requirements by up to 64 times**!
### Optional named vectors
@@ -83,7 +83,7 @@ returned.
Unlike some other vector databases, Qdrant accepts any arbitrary JSON payload, including
arrays, objects, and arrays of objects. You can also [filter the search results using nested
keys](/documentation/filtering/#nested-key), even though arrays (using the `[]` syntax).
keys](/documentation/search/filtering/#nested-key), even though arrays (using the `[]` syntax).
Before Qdrant 1.2 it was impossible to express some more complex conditions for the
nested structures. For example, let's assume we have the following payload:
@@ -195,18 +195,18 @@ Out-of-Memory errors.
Qdrant 1.2 enters recovery mode, if enabled, when it detects a failure on startup.
That makes the service halt the loading of collection data and commence operations in a partial state.
This state allows for removing collections but doesn't support search or update functions.
**Recovery mode [has to be enabled by user](/documentation/administration/#recovery-mode).**
**Recovery mode [has to be enabled by user](/documentation/operations/administration/#recovery-mode).**
### Appendable mmap
For a long time, segments using mmap storage were `non-appendable` and could only be constructed by
the optimizer. Dynamically adding vectors to the mmap file is fairly complicated and thus not
implemented in Qdrant, but we did our best to implement it in the recent release. If you want
to read more about segments, check out our docs on [vector storage](/documentation/storage/#vector-storage).
to read more about segments, check out our docs on [vector storage](/documentation/manage-data/storage/#vector-storage).
## Security
There are two major changes in terms of [security](/documentation/security/):
There are two major changes in terms of [security](/documentation/operations/security/):
1. **API-key support** - basic authentication with a static API key to prevent unwanted access. Previously
API keys were only supported in [Qdrant Cloud](https://cloud.qdrant.io/).
@@ -33,7 +33,7 @@ Your feedback is valuable to us, and are always tying to include some of your fe
## New features
### Asychronous I/O interface
### Asynchronous I/O interface
Going forward, we will support the `io_uring` asychnronous interface for storage devices on Linux-based systems. Since its introduction, `io_uring` has been proven to speed up slow-disk deployments as it decouples kernel work from the IO process.
@@ -59,7 +59,7 @@ Please keep in mind that this feature is experimental and that the interface may
### Oversampling for quantization
We are introducing [oversampling](/documentation/guides/quantization/#oversampling) as a new way to help you improve the accuracy and performance of similarity search algorithms. With this method, you are able to significantly compress high-dimensional vectors in memory and then compensate the accuracy loss by re-scoring additional points with the original vectors.
We are introducing [oversampling](/documentation/manage-data/quantization/#oversampling) as a new way to help you improve the accuracy and performance of similarity search algorithms. With this method, you are able to significantly compress high-dimensional vectors in memory and then compensate the accuracy loss by re-scoring additional points with the original vectors.
You will experience much faster performance with quantization due to parallel disk usage when reading vectors. Much better IO means that you can keep quantized vectors in RAM, so the pre-selection will be even faster. Finally, once pre-selection is done, you can use parallel IO to retrieve original vectors, which is significantly faster than traversing HNSW on slow disks.
@@ -180,7 +180,7 @@ client.search_groups(
We are excited to announce a more user-friendly way to organize and work with your collections inside of Qdrant. Our dashboard's design is simple, but very intuitive and easy to access.
Try it out now! If you have Docker running, you can [quickstart Qdrant](/documentation/quick-start/) and access the Dashboard locally from [http://localhost:6333/dashboard](http://localhost:6333/dashboard). You should see this simple access point to Qdrant:
Try it out now! If you have Docker running, you can [quickstart Qdrant](/documentation/quickstart/) and access the Dashboard locally from [http://localhost:6333/dashboard](http://localhost:6333/dashboard). You should see this simple access point to Qdrant:
![Qdrant Web UI](/articles_data/qdrant-1.3.x/web-ui.png)
@@ -201,7 +201,7 @@ Internally, `is_empty` was not using the index when it was called, so it had to
### Faster read access with mmap
If you used mmap, you most likely found that segments were always created with cold caches. The first request to the database needed to request the disk, which made startup slower despite plenty of RAM being available. We have implemeneted a way to ask the kernel to "heat up" the disk cache and make initialization much faster.
If you used mmap, you most likely found that segments were always created with cold caches. The first request to the database needed to request the disk, which made startup slower despite plenty of RAM being available. We have implemented a way to ask the kernel to "heat up" the disk cache and make initialization much faster.
The function is expected to be used on startup and after segment optimization and reloading of newly indexed segment. So far this is only implemented for "immutable" memmaps.
@@ -51,11 +51,11 @@ Things have changed since then, as so many of you wanted a single tool for spars
If you're coming across the topic of sparse vectors for the first time, our [Brief History of Search](/documentation/overview/vector-search/) explains the difference between sparse and dense vectors.
Check out the [sparse vectors article](/articles/sparse-vectors/) and [sparse vectors index docs](/documentation/concepts/indexing/#sparse-vector-index) for more details on what this new index means for Qdrant users.
Check out the [sparse vectors article](/articles/sparse-vectors/) and [sparse vectors index docs](/documentation/manage-data/indexing/#sparse-vector-index) for more details on what this new index means for Qdrant users.
### Discovery API
The recently launched [Discovery API](/documentation/concepts/explore/#discovery-api) extends the range of scenarios for leveraging vectors. While its interface mirrors the [Recommendation API](/documentation/concepts/explore/#recommendation-api), it focuses on refining the search parameters for greater precision.
The recently launched [Discovery API](/documentation/search/explore/#discovery-api) extends the range of scenarios for leveraging vectors. While its interface mirrors the [Recommendation API](/documentation/search/explore/#recommendation-api), it focuses on refining the search parameters for greater precision.
The concept of 'context' refers to a collection of positive-negative pairs that define zones within a space. Each pair effectively divides the space into positive or negative segments. This concept guides the search operation to prioritize points based on their inclusion within positive zones or their avoidance of negative zones. Essentially, the search algorithm favors points that fall within multiple positive zones or steer clear of negative ones.
The Discovery API can be used in two ways - either with or without the target point. The first case is called a **discovery search**, while the second is called a **context search**.
@@ -66,7 +66,7 @@ The Discovery API can be used in two ways - either with or without the target po
![Discovery search visualization](/articles_data/qdrant-1.7.x/discovery-search.png)
Please refer to the [Discovery API documentation on discovery search](/documentation/concepts/explore/#discovery-search) for more details and the internal mechanics of the operation.
Please refer to the [Discovery API documentation on discovery search](/documentation/search/explore/#discovery-search) for more details and the internal mechanics of the operation.
#### Context search
@@ -92,7 +92,7 @@ POST /collections/my_collection/points/search
}
```
If you want to know more about the user-defined sharding, please refer to the [sharding documentation](/documentation/guides/distributed_deployment/#sharding).
If you want to know more about the user-defined sharding, please refer to the [sharding documentation](/documentation/operations/distributed_deployment/#sharding).
### Snapshot-based shard transfer
@@ -101,7 +101,7 @@ That's a really more in depth technical improvement for the distributed mode use
Moving shards is required for dynamical scaling of the cluster. Your data can migrate between nodes, and the way you move it is crucial for the performance of the whole system. The good old `stream_records` method (still the default one) transmits all the records between the machines and indexes them on the target node.
In the case of moving the shard, it's necessary to recreate the HNSW index each time. However, with the introduction of the new `snapshot` approach, the snapshot itself, inclusive of all data and potentially quantized content, is transferred to the target node. This comprehensive snapshot includes the entire index, enabling the target node to seamlessly load it and promptly begin handling requests without the need for index recreation.
There are multiple scenarios in which you may prefer one over the other. Please check out the docs of the [shard transfer method](/documentation/guides/distributed_deployment/#shard-transfer-method) for more details and head-to-head comparison. As for now, the old `stream_records` method is still the default one, but we may decide to change it in the future.
There are multiple scenarios in which you may prefer one over the other. Please check out the docs of the [shard transfer method](/documentation/operations/distributed_deployment/#shard-transfer-method) for more details and head-to-head comparison. As for now, the old `stream_records` method is still the default one, but we may decide to change it in the future.
## Minor improvements
@@ -27,7 +27,7 @@ This time around, we have focused on Qdrant's internals. Our goal was to optimiz
- **Faster [sparse vectors](https://qdrant.tech/articles/sparse-vectors/):** [Hybrid search](https://qdrant.tech/articles/hybrid-search/) is up to 16x faster now!
- **CPU resource management:** You can allocate CPU threads for faster indexing.
- **Better indexing performance:** We optimized text [indexing](https://qdrant.tech/documentation/concepts/indexing/) on the backend.
- **Better indexing performance:** We optimized text [indexing](https://qdrant.tech/documentation/manage-data/indexing/) on the backend.
## Faster search with sparse vectors
@@ -55,7 +55,7 @@ The colors within both scatter plots show the frequency of results. The red dots
This performance increase can have a dramatic effect on hybrid search implementations. [Read more about how to set this up.](/articles/sparse-vectors/)
FYI, sparse vectors were released in [Qdrant v.1.7.0](/articles/qdrant-1.7.x/#sparse-vectors). They are stored using a different index, so first [check out the documentation](/documentation/concepts/indexing/#sparse-vector-index) if you want to try an implementation.
FYI, sparse vectors were released in [Qdrant v.1.7.0](/articles/qdrant-1.7.x/#sparse-vectors). They are stored using a different index, so first [check out the documentation](/documentation/manage-data/indexing/#sparse-vector-index) if you want to try an implementation.
## CPU resource management
@@ -65,7 +65,7 @@ This isn't mandatory, as Qdrant is by default tuned to strike the right balance
This version introduces a `optimizer_cpu_budget` parameter to control the maximum number of CPUs used for indexing.
> Read more about `config.yaml` in the [configuration file](/documentation/guides/configuration/).
> Read more about `config.yaml` in the [configuration file](/documentation/operations/configuration/).
```yaml
# CPU budget, how many CPUs (threads) to allocate for an optimization job.
@@ -99,8 +99,8 @@ This approach ensures stability in the [vector search](https://qdrant.tech/docum
Beyond these enhancements, [Qdrant v1.8.0](https://github.com/qdrant/qdrant/releases/tag/v1.8.0) adds and improves on several smaller features:
1. **Order points by payload:** In addition to searching for semantic results, you might want to retrieve results by specific metadata (such as price). You can now use Scroll API to [order points by payload key](/documentation/concepts/points/#order-points-by-payload-key).
2. **Datetime support:** We have implemented [datetime support for the payload index](/documentation/concepts/filtering/#datetime-range). Prior to this, if you wanted to search for a specific datetime range, you would have had to convert dates to UNIX timestamps. ([PR#3320](https://github.com/qdrant/qdrant/issues/3320))
1. **Order points by payload:** In addition to searching for semantic results, you might want to retrieve results by specific metadata (such as price). You can now use Scroll API to [order points by payload key](/documentation/manage-data/points/#order-points-by-payload-key).
2. **Datetime support:** We have implemented [datetime support for the payload index](/documentation/search/filtering/#datetime-range). Prior to this, if you wanted to search for a specific datetime range, you would have had to convert dates to UNIX timestamps. ([PR#3320](https://github.com/qdrant/qdrant/issues/3320))
3. **Check collection existence:** You can check whether a collection exists via the `/exists` endpoint to the `/collections/{collection_name}`. You will get a true/false response. ([PR#3472](https://github.com/qdrant/qdrant/pull/3472)).
4. **Find points** whose payloads match more than the minimal amount of conditions. We included the `min_should` match feature for a condition to be `true` ([PR#3331](https://github.com/qdrant/qdrant/pull/3466/)).
5. **Modify nested fields:** We have improved the `set_payload` API, adding the ability to update nested fields ([PR#3548](https://github.com/qdrant/qdrant/pull/3548)).
@@ -467,7 +467,7 @@ We will reprocess the data with the updated parameters above:
```python
## for iteration 2 - lets modify chunk configuration
## We will start with creating seperate collection to store vectors
## We will start with creating separate collection to store vectors
chunk_size = 1024
chunk_overlap = 128
@@ -408,11 +408,11 @@ Out of all retriever–feedback model pairs, three leaders emerged with the foll
## Relevance Feedback Query
The results were convincing enough to justify implementing Relevance Feedback Query. And [here it is](https://qdrant.tech/documentation/concepts/search-relevance/#relevance-feedback), ready for your retrieval pipelines!
The results were convincing enough to justify implementing Relevance Feedback Query. And [here it is](https://qdrant.tech/documentation/search/search-relevance/#relevance-feedback), ready for your retrieval pipelines!
### When to Use It
Use Relevance Feedback Query once a basic retrieval pipeline is in place and you're looking for additional techniques to boost result relevance, such as [Maximal Marginal Relevance (MMR)](https://qdrant.tech/documentation/concepts/search-relevance/#maximal-marginal-relevance-mmr), [Reranking](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries), or [Score Boosting](https://qdrant.tech/documentation/concepts/search-relevance/#score-boosting).
Use Relevance Feedback Query once a basic retrieval pipeline is in place and you're looking for additional techniques to boost result relevance, such as [Maximal Marginal Relevance (MMR)](https://qdrant.tech/documentation/search/search-relevance/#maximal-marginal-relevance-mmr), [Reranking](https://qdrant.tech/documentation/search/hybrid-queries/#multi-stage-queries), or [Score Boosting](https://qdrant.tech/documentation/search/search-relevance/#score-boosting).
**It's here not to replace but to complement other search relevance tools.**
For example, Relevance Feedback Query can be a great aid for search agents, letting you propagate the agent's understanding of your use case directly to the vector search index.
@@ -434,7 +434,7 @@ If no training queries are supplied, the package will train the formula directly
> **Warning:** If your use case doesn't involve document-to-document semantic similarity search, training on sampled documents alone may completely cancel the effect of relevance feedback scoring on real data.
> It's far more effective to use real queries.
Once you've obtained the weights, simply plug them into your [Qdrant Client of choice](https://qdrant.tech/documentation/concepts/search-relevance/#relevance-feedback).
Once you've obtained the weights, simply plug them into your [Qdrant Client of choice](https://qdrant.tech/documentation/search/search-relevance/#relevance-feedback).
### Evaluating Your Gains
@@ -290,6 +290,6 @@ expensive setup if you can agree to a small decrease in the search precision.
### Accessing best practices
Qdrant documentation on [Scalar Quantization](/documentation/quantization/#setting-up-quantization-in-qdrant)
Qdrant documentation on [Scalar Quantization](/documentation/manage-data/quantization/#setting-up-quantization-in-qdrant)
is a great resource describing different scenarios and strategies to achieve up to 4x
lower memory footprint and even up to 2x performance increase.
@@ -118,7 +118,7 @@ sparse_vectors_config={
}
```
**Hybrid in one request.** Combine sparse precision with dense semantics via native [RRF/prefetch](https://qdrant.tech/documentation/concepts/hybrid-queries/) - no external reranker:
**Hybrid in one request.** Combine sparse precision with dense semantics via native [RRF/prefetch](https://qdrant.tech/documentation/search/hybrid-queries/) - no external reranker:
```python
client.query_points(
@@ -100,7 +100,7 @@ This launches a web dashboard with tabs for each stage of the pipeline:
**Evaluate.** Point at a trained model and test queries. Get metric cards for nDCG@10, MRR@10, Recall, and Precision.
**[Collections](https://qdrant.tech/documentation/concepts/collections/).** Browse your Qdrant collections, check point counts, and run test searches against indexed products. Useful for sanity-checking that indexing worked before evaluation.
**[Collections](https://qdrant.tech/documentation/manage-data/collections/).** Browse your Qdrant collections, check point counts, and run test searches against indexed products. Useful for sanity-checking that indexing worked before evaluation.
**Publish.** Enter a model path and HuggingFace repo name. Click publish.
@@ -83,7 +83,7 @@ Most people use default settings and build vector search apps that aren't proper
#### Remember to run all tutorial code in Qdrant's Dashboard
The easiest way to reach that "Hello World" moment is to [**try filtering in a live cluster**](/documentation/quickstart-cloud/). Our interactive tutorial will show you how to create a cluster, add data and try some filtering clauses.
The easiest way to reach that "Hello World" moment is to [**try filtering in a live cluster**](/documentation/cloud-quickstart/). Our interactive tutorial will show you how to create a cluster, add data and try some filtering clauses.
![qdrant-filtering-tutorial](/articles_data/vector-search-filtering/qdrant-filtering-tutorial.png)
@@ -93,13 +93,13 @@ Qdrant follows a specific method of searching and filtering through dense vector
Let's take a look at this **3-stage diagram**. In this case, we are trying to find the nearest neighbour to the query vector **(green)**. Your search journey starts at the bottom **(orange)**.
By default, Qdrant connects all your data points within the [**vector index**](/documentation/concepts/indexing/). After you [**introduce filters**](/documentation/concepts/filtering/), some data points become disconnected. Vector search can't cross the grayed out area and it won't reach the nearest neighbor.
By default, Qdrant connects all your data points within the [**vector index**](/documentation/manage-data/indexing/). After you [**introduce filters**](/documentation/search/filtering/), some data points become disconnected. Vector search can't cross the grayed out area and it won't reach the nearest neighbor.
How can we bridge this gap?
**Figure 1:** How Qdrant maintains a filterable vector index.
![filterable-vector-index](/articles_data/vector-search-filtering/filterable-vector-index.png)
[**Filterable vector index**](/documentation/concepts/indexing/): This technique builds additional links **(orange)** between leftover data points. The filtered points which stay behind are now traversible once again. Qdrant uses special category-based methods to connect these data points.
[**Filterable vector index**](/documentation/manage-data/indexing/): This technique builds additional links **(orange)** between leftover data points. The filtered points which stay behind are now traversible once again. Qdrant uses special category-based methods to connect these data points.
### Qdrant's approach vs traditional filtering methods
@@ -201,7 +201,7 @@ As you can see, Qdrant's filtering method has a greater chance of capturing all
This specific example uses the `range` condition for filtering. Qdrant, however, offers many other possible ways to structure a filter
**For detailed usage examples, [filtering](/documentation/concepts/filtering/) docs are the best resource.**
**For detailed usage examples, [filtering](/documentation/search/filtering/) docs are the best resource.**
### Scrolling instead of searching
@@ -240,7 +240,7 @@ POST /collections/online_store/points/scroll
```
The response contains a batch of points that match the criteria and a reference (offset or next page token) to retrieve the next set of points.
> [**Scrolling**](/documentation/concepts/points/#scroll-points) is designed to be efficient. It minimizes the load on the server and reduces memory consumption on the client side by returning only manageable chunks of data at a time.
> [**Scrolling**](/documentation/manage-data/points/#scroll-points) is designed to be efficient. It minimizes the load on the server and reduces memory consumption on the client side by returning only manageable chunks of data at a time.
#### Available filtering conditions
@@ -254,7 +254,7 @@ The response contains a batch of points that match the criteria and a reference
| **Full Text Match** | Search in text fields. | **Is Empty** | Filter empty fields. |
| **Has ID** | Filter by unique ID. | **Is Null** | Filter null values. |
> All clauses and conditions are outlined in Qdrant's [filtering](/documentation/concepts/filtering/) documentation.
> All clauses and conditions are outlined in Qdrant's [filtering](/documentation/search/filtering/) documentation.
#### Filtering clauses to remember
@@ -559,10 +559,10 @@ You can use filters to retrieve data points without knowing their `id`. You can
| Action | Description | Action | Description |
|--------|-------------|--------|-------------|
| [Delete Points](/documentation/concepts/points/#delete-points) | Deletes all points matching the filter. | [Set Payload](/documentation/concepts/payload/#set-payload) | Adds payload fields to all points matching the filter. |
| [Scroll Points](/documentation/concepts/points/#scroll-points) | Lists all points matching the filter. | [Update Payload](/documentation/concepts/payload/#overwrite-payload) | Updates payload fields for points matching the filter. |
| [Order Points](/documentation/concepts/points/#order-points-by-payload-key) | Lists all points, sorted by the filter. | [Delete Payload](/documentation/concepts/payload/#delete-payload-keys) | Deletes fields for points matching the filter. |
| [Count Points](/documentation/concepts/points/#counting-points) | Totals the points matching the filter. | | |
| [Delete Points](/documentation/manage-data/points/#delete-points) | Deletes all points matching the filter. | [Set Payload](/documentation/manage-data/payload/#set-payload) | Adds payload fields to all points matching the filter. |
| [Scroll Points](/documentation/manage-data/points/#scroll-points) | Lists all points matching the filter. | [Update Payload](/documentation/manage-data/payload/#overwrite-payload) | Updates payload fields for points matching the filter. |
| [Order Points](/documentation/manage-data/points/#order-points-by-payload-key) | Lists all points, sorted by the filter. | [Delete Payload](/documentation/manage-data/payload/#delete-payload-keys) | Deletes fields for points matching the filter. |
| [Count Points](/documentation/manage-data/points/#counting-points) | Totals the points matching the filter. | | |
## Filtering with the payload index
@@ -577,7 +577,7 @@ Just how the vector index organizes vectors, the payload index will structure yo
![payload-index-vector-search](/articles_data/vector-search-filtering/payload-index-vector-search.png)
On its own, semantic searching over terabytes of data can take up lots of RAM. [**Filtering**](/documentation/concepts/filtering/) and [**Indexing**](/documentation/concepts/indexing/) are two easy strategies to reduce your compute usage and still get the best results. Remember, this is only a guide. For an exhaustive list of filtering options, you should read the [filtering documentation](/documentation/concepts/filtering/).
On its own, semantic searching over terabytes of data can take up lots of RAM. [**Filtering**](/documentation/search/filtering/) and [**Indexing**](/documentation/manage-data/indexing/) are two easy strategies to reduce your compute usage and still get the best results. Remember, this is only a guide. For an exhaustive list of filtering options, you should read the [filtering documentation](/documentation/search/filtering/).
Here is how you can create a single index for a metadata field "category":
@@ -660,11 +660,11 @@ If your users are often filtering by **laptop** when looking up a product **cate
| Index Type | Description |
|---------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------|
| [Full-text Index](/documentation/concepts/indexing/#full-text-index) | Enables efficient text search in large datasets. |
| [Tenant Index](/documentation/concepts/indexing/#tenant-index) | For data isolation and retrieval efficiency in multi-tenant architectures. |
| [Principal Index](/documentation/concepts/indexing/#principal-index) | Manages data based on primary entities like users or accounts. |
|[On-Disk Index](/documentation/concepts/indexing/#on-disk-payload-index) | Stores indexes on disk to manage large datasets without memory usage. |
| [Parameterized Index](/documentation/concepts/indexing/#parameterized-index) | Allows for dynamic querying, where the index can adapt based on different parameters or conditions provided by the user. Useful for numeric data like prices or timestamps. |
| [Full-text Index](/documentation/manage-data/indexing/#full-text-index) | Enables efficient text search in large datasets. |
| [Tenant Index](/documentation/manage-data/indexing/#tenant-index) | For data isolation and retrieval efficiency in multi-tenant architectures. |
| [Principal Index](/documentation/manage-data/indexing/#principal-index) | Manages data based on primary entities like users or accounts. |
|[On-Disk Index](/documentation/manage-data/indexing/#on-disk-payload-index) | Stores indexes on disk to manage large datasets without memory usage. |
| [Parameterized Index](/documentation/manage-data/indexing/#parameterized-index) | Allows for dynamic querying, where the index can adapt based on different parameters or conditions provided by the user. Useful for numeric data like prices or timestamps. |
### Indexing payloads in multitenant setups
@@ -688,7 +688,7 @@ PUT /collections/{collection_name}/index
```
Additionally, we offer a way of organizing data efficiently by means of the tenant index. This is another variant of the payload index that makes tenant data more accessible. This time, the request will specify the field as a tenant. This means that you can mark various customer types and user id’s as `is_tenant: true`.
Read more about setting up [tenant defragmentation](/documentation/concepts/indexing/?q=tenant#tenant-index) in multitenant environments,
Read more about setting up [tenant defragmentation](/documentation/manage-data/indexing/?q=tenant#tenant-index) in multitenant environments,
## Key takeaways in filtering and indexing
![best-practices](/articles_data/vector-search-filtering/best-practices.png)
@@ -744,7 +744,7 @@ As a conclusion to this guide, let's look at some real-life use cases where filt
#### Before you go - all the code is in Qdrant's Dashboard
The easiest way to reach that "Hello World" moment is to [**try filtering in a live cluster**](/documentation/quickstart-cloud/). Our interactive tutorial will show you how to create a cluster, add data and try some filtering clauses.
The easiest way to reach that "Hello World" moment is to [**try filtering in a live cluster**](/documentation/cloud-quickstart/). Our interactive tutorial will show you how to create a cluster, add data and try some filtering clauses.
**It's all in your free cluster!**
@@ -52,17 +52,17 @@ This article will help you successfully deploy and maintain vector search system
### Ensure your hot dataset fits in RAM for low-latency queries.
If not, then you'll have to [**offload data 'on_disk'**](/documentation/concepts/storage/#configuring-memmap-storage). If this parameter is enabled, Qdrant caches your most frequently accessed vectors loaded into RAM, and the rest is memory-mapped onto the disk.
If not, then you'll have to [**offload data 'on_disk'**](/documentation/manage-data/storage/#configuring-memmap-storage). If this parameter is enabled, Qdrant caches your most frequently accessed vectors loaded into RAM, and the rest is memory-mapped onto the disk.
This ensures minimal disk access during queries, significantly reducing latency and boosting overall performance. By monitoring query patterns and usage metrics, you can identify which subsets of your data deserve dedicated in-memory storage, reserving disk access only for colder, less frequently queried vectors.
||
|-|
|**Read More:** [**Storage Documentation**](https://qdrant.tech/documentation/concepts/storage/)|
|**Read More:** [**Storage Documentation**](https://qdrant.tech/documentation/manage-data/storage/)|
### Index Your Important Metadata to Avoid Costly Queries
✅ You should always [**create payload indexes**](https://qdrant.tech/documentation/concepts/indexing/#payload-index) for all fields used in filters or sorting.
✅ You should always [**create payload indexes**](https://qdrant.tech/documentation/manage-data/indexing/#payload-index) for all fields used in filters or sorting.
Many users configure complex filters but may not be aware of the need to create corresponding payload indexes.
@@ -72,11 +72,11 @@ Filtering after retrieving thousands of vectors can get expensive. If you don't
Unlike some other engines, Qdrant lets you make the optimal choice of which fields to index for your use case rather than creating indexes for every field by default.
> **Note:** Don't forget to use the correct [**payload index type**](https://qdrant.tech/documentation/concepts/indexing/#payload-index). If there are numeric values, the you must use a numeric index. If you represent numbers in strings ("123"), a numeric index will not work.
> **Note:** Don't forget to use the correct [**payload index type**](https://qdrant.tech/documentation/manage-data/indexing/#payload-index). If there are numeric values, the you must use a numeric index. If you represent numbers in strings ("123"), a numeric index will not work.
||
|-|
|**Read More:** [**Filtering Documentation**](https://qdrant.tech/documentation/concepts/filtering/)|
|**Read More:** [**Filtering Documentation**](https://qdrant.tech/documentation/search/filtering/)|
### Don't Forget to Tune HNSW Search Parameters
@@ -88,7 +88,7 @@ Sometimes users don't properly balance HNSW search parameters. Setting the HNSW
❓ **Use Case:** A customer ran advanced similarity searches across their vast dataset of nearly 800 million vectors. Initially, they found that queries took anywhere from 10 to 20 seconds, especially when combining multiple filters and metadata fields.
> ✅ How can they retain accuracy and keep things fast? [**The answer is optimization.**](https://qdrant.tech/documentation/guides/optimize/)
> ✅ How can they retain accuracy and keep things fast? [**The answer is optimization.**](https://qdrant.tech/documentation/operations/optimize/)
**Figure 1:** Qdrant is highly configurable. You can configure it for speed, precision or resource use.
![qdrant resource tradeoffs](/docs/tradeoff.png)
@@ -99,8 +99,8 @@ This strategy balanced memory usage with performance: only the compact vectors n
||
|-|
|**Read More:** [**Optimization Guide**](https://qdrant.tech/documentation/guides/optimize/)|#optimizing-qdrant-performance-three-scenarios
|**Read More:** [**HNSW Documentation**](https://qdrant.tech/documentation/concepts/indexing/#vector-index)|
|**Read More:** [**Optimization Guide**](https://qdrant.tech/documentation/operations/optimize/)|#optimizing-qdrant-performance-three-scenarios
|**Read More:** [**HNSW Documentation**](https://qdrant.tech/documentation/manage-data/indexing/#vector-index)|
### Compress Your Data with Quantization Strategies
@@ -108,21 +108,21 @@ This strategy balanced memory usage with performance: only the compact vectors n
|:-:|
|**"We're using too much memory for our massive dataset."**|
Many users skip [**quantization**](https://qdrant.tech/documentation/guides/quantization/), causing their index to consume excessive RAM and produce uneven performance. Some users hesitate to compromise precision, but this is not always the case.
Many users skip [**quantization**](https://qdrant.tech/documentation/manage-data/quantization/), causing their index to consume excessive RAM and produce uneven performance. Some users hesitate to compromise precision, but this is not always the case.
If your workload can tolerate a moderate drop in embedding precision, data compression offers a powerful way to shrink vector size and slash memory usage. By converting high-dimensional floating-point values into lower-bit formats (such as 8-bit scalar or even a single bit-sized representations), you can keep far more vectors in RAM while reducing disk footprint.
> ✅ [**You should evaluate and apply quantization**](https://qdrant.tech/documentation/guides/quantization/#how-to-choose-the-right-quantization-method) if your use case allows. Quantization seriously improves performance and reduces storage costs.
> ✅ [**You should evaluate and apply quantization**](https://qdrant.tech/documentation/manage-data/quantization/#how-to-choose-the-right-quantization-method) if your use case allows. Quantization seriously improves performance and reduces storage costs.
This not only speeds up query throughput for large-scale datasets, but also cuts hardware costs and storage overhead. While Scalar Quantization is a midrange compression alternative, Binary quantization is more drastic, so be sure to test your accuracy requirements for each thoroughly.
When using [**quantization**](https://qdrant.tech/documentation/guides/quantization/), you can store only the compressed vectors in memory while leaving the original floating-point versions on disk for reference. This approach dramatically lowers RAM consumption—since quantized vectors take far less space—yet still allows you to retrieve full-precision vectors if needed for downstream tasks like re-ranking.
When using [**quantization**](https://qdrant.tech/documentation/manage-data/quantization/), you can store only the compressed vectors in memory while leaving the original floating-point versions on disk for reference. This approach dramatically lowers RAM consumption—since quantized vectors take far less space—yet still allows you to retrieve full-precision vectors if needed for downstream tasks like re-ranking.
>**Sidenote:** You can always enable `async_io` scorer when the linux kernel supports it and if you have `on_disk` vectors.
||
|-|
|**Read More:** [**Quantization Documentation**](https://qdrant.tech/documentation/guides/quantization/)|
|**Read More:** [**Quantization Documentation**](https://qdrant.tech/documentation/manage-data/quantization/)|
## 2. How do I Ingest and Index Large Amounts of Data?
![vector-search-production](/articles_data/vector-search-production/vector-search-production-2.jpg)
@@ -141,7 +141,7 @@ Once all records are inserted, you can rebuild the index in a single pass. Consi
||
|-|
|**Read More:** [**Configuring the Vector Index**](https://qdrant.tech/documentation/concepts/indexing/#vector-index)|
|**Read More:** [**Configuring the Vector Index**](https://qdrant.tech/documentation/manage-data/indexing/#vector-index)|
### Other Solutions to Alleviate Indexing Bottleneck
@@ -155,7 +155,7 @@ Once all records are inserted, you can rebuild the index in a single pass. Consi
||
|-|
|**Read More:** [**Configuration Documentation**](https://qdrant.tech/documentation/guides/configuration/)|
|**Read More:** [**Configuration Documentation**](https://qdrant.tech/documentation/operations/configuration/)|
### When Indexing Falls Behind Ingestion
![vector-search-production](/articles_data/vector-search-production/vector-search-production-3.jpg)
@@ -166,7 +166,7 @@ By default, searches include unindexed data. However, a large number of unindexe
If the maximum number of indexed points remains consistently low, this is likely not an issue. If you anticipate periods with many unindexed points, you should take measures to prevent search disruptions in production.
One option is to [**set `indexed_only=true` in search requests**](https://qdrant.tech/documentation/concepts/search/#search-api). This will ensure fast searches by only considering indexed data, at the expense of eventual consistency (new data becomes searchable only after indexing).
One option is to [**set `indexed_only=true` in search requests**](https://qdrant.tech/documentation/search/search/#search-api). This will ensure fast searches by only considering indexed data, at the expense of eventual consistency (new data becomes searchable only after indexing).
Alternatively, you can perform [**bulk vector uploads**](https://qdrant.tech/documentation/database-tutorials/bulk-upload/) during low-traffic periods to allow indexing to complete before increased traffic.
@@ -174,7 +174,7 @@ Alternatively, you can perform [**bulk vector uploads**](https://qdrant.tech/doc
||
|-|
|**Read More:** [**Indexing Documentation**](https://qdrant.tech/documentation/concepts/indexing/)|
|**Read More:** [**Indexing Documentation**](https://qdrant.tech/documentation/manage-data/indexing/)|
### How to Arrange Metadata and Schema for Consistency
@@ -184,7 +184,7 @@ Alternatively, you can perform [**bulk vector uploads**](https://qdrant.tech/doc
In some cases, the payload schema is inconsistent across data pipelines, so some fields have mismatched types or are missing altogether.
❓ **Use Case:** A healthcare firm discovered that some pipelines inserted strings where others inserted integers. Filters broke silently or returned inconsistent results, signalling that [**a unified payload schema**](https://qdrant.tech/documentation/concepts/indexing/#payload-index) was not in place.
❓ **Use Case:** A healthcare firm discovered that some pipelines inserted strings where others inserted integers. Filters broke silently or returned inconsistent results, signalling that [**a unified payload schema**](https://qdrant.tech/documentation/manage-data/indexing/#payload-index) was not in place.
> When payload fields are typed inconsistently across your ingestion pipelines, filters can break in unpredictable ways.
@@ -194,13 +194,13 @@ For example, **some services might write a "status" field as a string ("active")
||
|-|
|**Read More:** [**Payload Documentation**](https://qdrant.tech/documentation/concepts/payload/)|
|**Read More:** [**Payload Documentation**](https://qdrant.tech/documentation/manage-data/payload/)|
### Decide How to Set Up a Multitenant Collection
❓ **Use Case:** When implementing vector databases, healthcare organizations need to ensure isolation between users' data. Our customer needed to make sure that when they filtered queries to only show a particular patient's documents, and no other patient's documents appeared in the query results.
✅ [**You should almost always consolidate tenants to a single collection**](https://qdrant.tech/documentation/guides/multiple-partitions/) if possible, tagging by tenant.
✅ [**You should almost always consolidate tenants to a single collection**](https://qdrant.tech/documentation/manage-data/multitenancy/) if possible, tagging by tenant.
```text
PUT /collections/{collection_name}/index
@@ -221,7 +221,7 @@ Figure: For many-tenant setups, spinning up a new collection per tenant can ball
||
|-|
|**Read More:** [**Multitenancy Documentation**](https://qdrant.tech/documentation/concepts/multitenancy/)|
|**Read More:** [**Multitenancy Documentation**](https://qdrant.tech/documentation/manage-data/multitenancy/)|
## 3. What's the Best Way to Scale the Database and Optimize Resources?
![vector-search-production](/articles_data/vector-search-production/vector-search-production-4.jpg)
@@ -230,7 +230,7 @@ Figure: For many-tenant setups, spinning up a new collection per tenant can ball
|:-:|
|**"How many nodes, CPUs, RAM and storage do I need for my Qdrant Cluster?"**|
It depends. If you're just starting out - we have prepared a tool on our website to help you figure this out. For more information, [**check out the Capacity Planning document as well.**](https://qdrant.tech/documentation/guides/capacity-planning/)
It depends. If you're just starting out - we have prepared a tool on our website to help you figure this out. For more information, [**check out the Capacity Planning document as well.**](https://qdrant.tech/documentation/operations/capacity-planning/)
✅ [**Use the sizing calculator**](https://cloud.qdrant.io/calculator) or performance testing to ensure node specs (RAM/CPU) match your workload.
@@ -242,7 +242,7 @@ It depends. If you're just starting out - we have prepared a tool on our website
A three-node setup provides a baseline for fault tolerance: if one node goes offline, the remaining two can continue serving queries and maintain a quorum for data consistency. This guards against hardware failures, rolling updates, and network disruptions. Fewer than three nodes leaves you vulnerable to single-point failures that can knock your entire cluster offline.
> [**We follow the Raft Protocol**](https://qdrant.tech/documentation/guides/distributed_deployment/#raft), so check out the docs and learn why this is important.
> [**We follow the Raft Protocol**](https://qdrant.tech/documentation/operations/distributed_deployment/#raft), so check out the docs and learn why this is important.
✅ **Set a replication factor of at least 2** to tolerate node failure without losing availability.
@@ -274,7 +274,7 @@ Development and staging environments often run experimental builds, tests, or si
> It's quite possible that the user has multiple shards on one node, which end up handling most traffic while other nodes remain underutilized.
In this case, you should [**choose the right number of shards**](https://qdrant.tech/documentation/guides/distributed_deployment/#sharding) based on your node count and expected RPS.
In this case, you should [**choose the right number of shards**](https://qdrant.tech/documentation/operations/distributed_deployment/#sharding) based on your node count and expected RPS.
You need to implement a shard strategy that aligns with real usage patterns. First, distribute your shards across all available nodes. This will help balance the load more effectively. After redistributing the shards, run performance tests to see how it affects your system. Then add replicas and test again to see how that changes performance.
@@ -285,14 +285,14 @@ Proper sharding considers data distribution and query patterns. By default, shar
||
|-|
|**Read More:** [**Sharding Documentation**](https://qdrant.tech/documentation/concepts/sharding/)|
|**Read More:** [**Sharding Documentation**](https://qdrant.tech/documentation/operations/distributed_deployment/#sharding)|
### Manage Your Costs by Scaling Up or Down
![vector-search-production](/articles_data/vector-search-production/vector-search-production-5.jpg)
Some teams scale up for daytime surges, then scale down overnight to save resources. If you do this, ensure data is sharded and replicated appropriately, so that scaling up and down won't result in service degradation.
If using Qdrant Cloud you could also do this using the [**Replication Factor**](https://qdrant.tech/documentation/guides/distributed_deployment/#replication-factor), though it may be considered a bit of a hack.
If using Qdrant Cloud you could also do this using the [**Replication Factor**](https://qdrant.tech/documentation/operations/distributed_deployment/#replication-factor), though it may be considered a bit of a hack.
> If you have 3 nodes with just 1 shard, and replication factor 6. It will create 3 replicas (one on each node) of that shard, because it can't host more. If you add 3 more nodes at peak times, it'll automatically replicate that shard 3 more times in an attempt to match the factor of 6.
@@ -308,7 +308,7 @@ If new nodes remain empty after joining, you waste resources. If departing nodes
||
|-|
|**Read More:** [**Distributed Deployment Documentation**](https://qdrant.tech/documentation/guides/distributed_deployment/)|
|**Read More:** [**Distributed Deployment Documentation**](https://qdrant.tech/documentation/operations/distributed_deployment/)|
|**Read More:** [**Resharding**](https://qdrant.tech/documentation/cloud/cluster-scaling/#resharding)|
### How to Predict and Test Cluster Performance
@@ -331,7 +331,7 @@ Remember, cold-starts and query behaviour are dataset dependent, which is why yo
||
|-|
|**Read More:** [Distributed Deployment Documentation](https://qdrant.tech/documentation/guides/distributed_deployment/)
|**Read More:** [Distributed Deployment Documentation](https://qdrant.tech/documentation/operations/distributed_deployment/)
### How to Design Your Systems to Protect Against Failure
@@ -375,7 +375,7 @@ By following these comprehensive load testing practices, you'll be able to ident
||
|-|
|**Read More:** [**Telemetry and Monitoring Documentation**](https://qdrant.tech/documentation/guides/monitoring/)|
|**Read More:** [**Telemetry and Monitoring Documentation**](https://qdrant.tech/documentation/operations/monitoring/)|
|**Read More:** [**Cloud Monitoring Documentation**](https://qdrant.tech/documentation/hybrid-cloud/networking-logging-monitoring/)
## 4. Ensuring Disaster Recovery With Database Backups and Snapshots
@@ -413,7 +413,7 @@ If you host tens of billions of vectors, store backups off-node in a different d
||
|-|
|**Read More:** [**Snapshot Documentation**](https://qdrant.tech/documentation/concepts/snapshots/)|
|**Read More:** [**Snapshot Documentation**](https://qdrant.tech/documentation/operations/snapshots/)|
|**Read More:** [**Managed Cloud Backup Documentation**](https://qdrant.tech/documentation/cloud/backups/)|
|**Read More:** [**Private Cloud Backup Documentation**](https://qdrant.tech/documentation/private-cloud/backups/)|
@@ -438,7 +438,7 @@ Investigations showed they hadn't adjusted the default configuration or reserved
||
|-|
|**Read More:** [**Qdrant Configuration Documentation**](https://qdrant.tech/documentation/guides/configuration/)|
|**Read More:** [**Qdrant Configuration Documentation**](https://qdrant.tech/documentation/operations/configuration/)|
### Security & Governance
@@ -450,7 +450,7 @@ Enabling TLS/HTTPS is essential for meeting compliance requirements in regulated
> You need to protect data in transit. To enable TLS/HTTPS for encrypted traffic in production, you need to configure secure communication between clients and your Qdrant database, as well as individual cluster nodes. This involves implementing Transport Layer Security (TLS) certificates to encrypt all traffic, preventing unauthorized access and data interception.
If self-hosting, you can set up encryption yourself by [**incorporating TLS directly from the configuration**](https://qdrant.tech/documentation/guides/security/#tls)
If self-hosting, you can set up encryption yourself by [**incorporating TLS directly from the configuration**](https://qdrant.tech/documentation/operations/security/#tls)
```text
service:
@@ -469,7 +469,7 @@ tls:
||
|-|
|**Read More:** [**Security Documentation**](https://qdrant.tech/documentation/guides/security/)|
|**Read More:** [**Security Documentation**](https://qdrant.tech/documentation/operations/security/)|
### Setting up Access Controls in Production
@@ -32,12 +32,12 @@ Let's take a look at some common goals and optimization strategies:
| Intended Result | Optimization Strategy |
|--------------------------------|------------------------------|
| [**High Search Precision + Low Memory Expenditure**](/documentation/guides/optimize/#1-high-speed-search-with-low-memory-usage) | [**On-Disk Indexing**](/documentation/guides/optimize/#1-high-speed-search-with-low-memory-usage) |
| [**Low Memory Expenditure + Fast Search Speed**](/documentation/guides/quantization/) | [**Quantization**](/documentation/guides/quantization/) |
| [**High Search Precision + Fast Search Speed**](/documentation/guides/optimize/#3-high-precision-with-high-speed-search) | [**RAM Storage + Quantization**](/documentation/guides/optimize/#3-high-precision-with-high-speed-search) |
| [**Balance Latency vs Throughput**](/documentation/guides/optimize/#balancing-latency-and-throughput) | [**Segment Configuration**](/documentation/guides/optimize/#balancing-latency-and-throughput) |
| [**High Search Precision + Low Memory Expenditure**](/documentation/operations/optimize/#1-high-speed-search-with-low-memory-usage) | [**On-Disk Indexing**](/documentation/operations/optimize/#1-high-speed-search-with-low-memory-usage) |
| [**Low Memory Expenditure + Fast Search Speed**](/documentation/manage-data/quantization/) | [**Quantization**](/documentation/manage-data/quantization/) |
| [**High Search Precision + Fast Search Speed**](/documentation/operations/optimize/#3-high-precision-with-high-speed-search) | [**RAM Storage + Quantization**](/documentation/operations/optimize/#3-high-precision-with-high-speed-search) |
| [**Balance Latency vs Throughput**](/documentation/operations/optimize/#balancing-latency-and-throughput) | [**Segment Configuration**](/documentation/operations/optimize/#balancing-latency-and-throughput) |
After this article, check out the code samples in our docs on [**Qdrant’s Optimization Methods**](/documentation/guides/optimize/).
After this article, check out the code samples in our docs on [**Qdrant’s Optimization Methods**](/documentation/operations/optimize/).
---
@@ -47,7 +47,7 @@ After this article, check out the code samples in our docs on [**Qdrant’s Opti
A vector index is the central location where Qdrant calculates vector similarity. It is the backbone of your search process, retrieving relevant results from vast amounts of data.
Qdrant uses the [**HNSW (Hierarchical Navigable Small World Graph) algorithm**](/documentation/concepts/indexing/#vector-index) as its dense vector index, which is both powerful and scalable.
Qdrant uses the [**HNSW (Hierarchical Navigable Small World Graph) algorithm**](/documentation/manage-data/indexing/#vector-index) as its dense vector index, which is both powerful and scalable.
**Figure 2:** A sample HNSW vector index with three layers. Follow the blue arrow on the top layer to see how a query travels throughout the database index. The closest result is on the bottom level, nearest to the gray query point.
@@ -57,7 +57,7 @@ Qdrant uses the [**HNSW (Hierarchical Navigable Small World Graph) algorithm**](
Working with massive datasets that contain billions of vectors demands significant resources—and those resources come with a price. While Qdrant provides reasonable defaults, tailoring them to your specific use case can unlock optimal performance. Here’s what you need to know.
The following parameters give you the flexibility to fine-tune Qdrant’s performance for your specific workload. You can modify them directly in Qdrant's [**configuration**](https://qdrant.tech/documentation/guides/configuration/) files or at the collection and named vector levels for more granular control.
The following parameters give you the flexibility to fine-tune Qdrant’s performance for your specific workload. You can modify them directly in Qdrant's [**configuration**](https://qdrant.tech/documentation/operations/configuration/) files or at the collection and named vector levels for more granular control.
**Figure 3:** A description of three key HNSW parameters.
@@ -101,7 +101,7 @@ client.query_points(
)
```
---
These are just the basics of HNSW. Learn More about [**Indexing**](/documentation/concepts/indexing/).
These are just the basics of HNSW. Learn More about [**Indexing**](/documentation/manage-data/indexing/).
---
@@ -110,7 +110,7 @@ These are just the basics of HNSW. Learn More about [**Indexing**](/documentatio
Efficient data compression is a cornerstone of resource optimization in vector databases. By reducing memory usage, you can achieve faster query performance without sacrificing too much accuracy.
One powerful technique is [**quantization**](/documentation/guides/quantization/), which transforms high-dimensional vectors into compact representations while preserving relative similarity. Let’s explore the quantization options available in Qdrant.
One powerful technique is [**quantization**](/documentation/manage-data/quantization/), which transforms high-dimensional vectors into compact representations while preserving relative similarity. Let’s explore the quantization options available in Qdrant.
#### Scalar Quantization
@@ -156,7 +156,7 @@ When working with Qdrant, you can fine-tune the quantization configuration to op
Adjust these settings to strike the right balance between precision and efficiency for your specific workload.
---
Learn More about [**Scalar Quantization**](/documentation/guides/quantization/)
Learn More about [**Scalar Quantization**](/documentation/manage-data/quantization/)
---
@@ -194,7 +194,7 @@ client.create_collection(
> By default, quantized vectors load like original vectors unless you set `always_ram` to `True` for instant access and faster queries.
---
Learn more about [**Binary Quantization**](/documentation/guides/quantization/)
Learn more about [**Binary Quantization**](/documentation/manage-data/quantization/)
---
@@ -259,7 +259,7 @@ client.upsert(
To ensure proper data isolation in a multitenant environment, you can assign a unique identifier, such as a **group_id**, to each vector. This approach ensures that each user's data remains segregated, allowing users to access only their own data. You can further enhance this setup by applying filters during queries to restrict access to the relevant data.
---
Learn More about [**Multitenancy**](/documentation/guides/multiple-partitions/)
Learn More about [**Multitenancy**](/documentation/manage-data/multitenancy/)
---
@@ -325,7 +325,7 @@ Here’s how to choose the shard_number:
| **Plan for Scalability** | Start with at least **2 shards per node** to allow room for future growth. |
| **Future-Proofing** | Starting with around **12 shards** is a good rule of thumb. This setup allows your system to scale seamlessly from 1 to 12 nodes without requiring re-sharding. |
Learn more about [**Sharding in Distributed Deployment**](/documentation/guides/distributed_deployment/)
Learn more about [**Sharding in Distributed Deployment**](/documentation/operations/distributed_deployment/)
---
@@ -358,10 +358,10 @@ results = client.search(
![filterable-vector-index](/articles_data/vector-search-resource-optimization/filterable-vector-index.png)
[**Filterable vector index**](/documentation/concepts/indexing/): This technique builds additional links **(orange)** between leftover data points. The filtered points which stay behind are now traversible once again. Qdrant uses special category-based methods to connect these data points.
[**Filterable vector index**](/documentation/manage-data/indexing/): This technique builds additional links **(orange)** between leftover data points. The filtered points which stay behind are now traversible once again. Qdrant uses special category-based methods to connect these data points.
---
Read more about [**Filtering Docs**](/documentation/concepts/filtering/) and check out the [**Complete Filtering Guide**](/articles/vector-search-filtering/).
Read more about [**Filtering Docs**](/documentation/search/filtering/) and check out the [**Complete Filtering Guide**](/articles/vector-search-filtering/).
---
#### Batch Processing
@@ -417,7 +417,7 @@ ___
#### Hybrid Search
Hybrid search combines **keyword filtering** with **vector similarity search**, enabling faster and more precise results. Keywords help narrow down the dataset quickly, while vector similarity ensures semantic accuracy. This search method combines [**dense and sparse vectors**](/documentation/concepts/vectors/).
Hybrid search combines **keyword filtering** with **vector similarity search**, enabling faster and more precise results. Keywords help narrow down the dataset quickly, while vector similarity ensures semantic accuracy. This search method combines [**dense and sparse vectors**](/documentation/manage-data/vectors/).
Hybrid search in Qdrant uses both fusion and reranking. The former is about combining the results from different search methods, based solely on the scores returned by each method. That usually involves some normalization, as the scores returned by different methods might be in different ranges.
@@ -428,7 +428,7 @@ Hybrid search in Qdrant uses both fusion and reranking. The former is about comb
After that, there is a formula that takes the relevancy measures and calculates the final score that we use later on to reorder the documents. Qdrant has built-in support for the Reciprocal Rank Fusion method, which is the de facto standard in the field.
---
Learn more about [**Hybrid Search**](/articles/hybrid-search/) and read out [**Hybrid Queries docs**](/documentation/concepts/hybrid-queries/).
Learn more about [**Hybrid Search**](/articles/hybrid-search/) and read out [**Hybrid Queries docs**](/documentation/search/hybrid-queries/).
---
@@ -494,7 +494,7 @@ client.query_points(
)
```
___
Learn more about [**Reranking**](/documentation/search-precision/reranking-hybrid-search/#rerank).
Learn more about [**Reranking**](/documentation/tutorials-search-engineering/reranking-hybrid-search/#rerank).
---
@@ -584,7 +584,7 @@ Here are some important metrics to monitor:
| grpc_responses_avg_duration_seconds | | Average response duration in gRPC API |
| rest_responses_fail_total | | Total number of failed responses (REST) |
Read more about [**Qdrant Open Source Monitoring**](/documentation/guides/monitoring/) and [**Qdrant Cloud Monitoring**](/documentation/cloud/cluster-monitoring/) for managed clusters.
Read more about [**Qdrant Open Source Monitoring**](/documentation/operations/monitoring/) and [**Qdrant Cloud Monitoring**](/documentation/cloud/cluster-monitoring/) for managed clusters.
_________________________________________________________________________
## Recap: When Should You Optimize?
@@ -169,7 +169,7 @@ In this loss, the model is trained by fitting the information of relative simila
Using the same mechanics, we can look at the training process from the other side.
Given a trained model, the user can provide positive and negative examples, and the goal of the discovery process is then to find suitable anchors across the stored collection of vectors.
<!-- ToDo: image where we know positive and nagative -->
<!-- ToDo: image where we know positive and negative -->
{{< figure width=60% src=/articles_data/vector-similarity-beyond-search/discovery.png caption="Reversed triplet loss" >}}
Multiple positive-negative pairs can be provided to make the discovery process more accurate.
@@ -140,7 +140,7 @@ We plan to go deeper into selecting the best model based on performance, cost, i
## Create a neural search service with Fastmbed
Now that you’re familiar with the core concepts around vector embeddings, how about start building your own [Neural Search Service](/documentation/tutorials/neural-search/)?
Now that you’re familiar with the core concepts around vector embeddings, how about start building your own [Neural Search Service](/documentation/tutorials-search-engineering/neural-search/)?
Tutorial guides you through a practical application of how to use Qdrant for document management based on descriptions of companies from [startups-list.com](https://www.startups-list.com/). From embedding data, integrating it with Qdrant's vector database, constructing a search API, and finally deploying your solution with FastAPI.
@@ -121,7 +121,7 @@ A vector database is made of multiple different entities and relations. Let's un
### Collections
A [collection](https://qdrant.tech/documentation/concepts/collections/) is essentially a group of **vectors** (or “[points](https://qdrant.tech/documentation/concepts/points/)”) that are logically grouped together **based on similarity or a specific task**. Every vector within a collection shares the same dimensionality and can be compared using a single metric. Avoid creating multiple collections unless necessary; instead, consider techniques like **sharding** for scaling across nodes or **multitenancy** for handling different use cases within the same infrastructure.
A [collection](https://qdrant.tech/documentation/manage-data/collections/) is essentially a group of **vectors** (or “[points](https://qdrant.tech/documentation/manage-data/points/)”) that are logically grouped together **based on similarity or a specific task**. Every vector within a collection shares the same dimensionality and can be compared using a single metric. Avoid creating multiple collections unless necessary; instead, consider techniques like **sharding** for scaling across nodes or **multitenancy** for handling different use cases within the same infrastructure.
### Distance Metrics
@@ -154,7 +154,7 @@ client.create_collection(
)
```
For other configurations like `hnsw_config.on_disk` or `memmap_threshold`, see the Qdrant documentation for [Storage.](https://qdrant.tech/documentation/concepts/storage/)
For other configurations like `hnsw_config.on_disk` or `memmap_threshold`, see the Qdrant documentation for [Storage.](https://qdrant.tech/documentation/manage-data/storage/)
### SDKs
@@ -186,7 +186,7 @@ In Qdrant, indexing is modular. You can configure indexes for **both vectors and
You need to build the payload index for **each field** you'd like to search. The magic here is in the combination: HNSW finds similar vectors, and the payload index makes sure only the ones that fit your criteria come through. Learn more about Qdrant's [Filterable HNSW](https://qdrant.tech/articles/filterable-hnsw/) and why it was built like this.
> Combining [full-text search](https://qdrant.tech/documentation/concepts/indexing/#full-text-index) with vector-based search gives you even more versatility. You can simultaneously search for conceptually similar documents while ensuring specific keywords are present, all within the same query.
> Combining [full-text search](https://qdrant.tech/documentation/manage-data/indexing/#full-text-index) with vector-based search gives you even more versatility. You can simultaneously search for conceptually similar documents while ensuring specific keywords are present, all within the same query.
### 2. Searching: Approximate Nearest Neighbors (ANN) Search
@@ -334,7 +334,7 @@ It works by converting high-dimensional vectors, which typically use `4 bytes` p
Quantization reduces data precision, and yes, this does lead to some loss of accuracy. However, for binary quantization, **OpenAI embeddings** achieves this performance improvement at a cost of only 5% of accuracy. If you apply techniques like **oversampling** and **rescoring**, this loss can be brought down even further.
However, binary quantization isn’t the only available option. Techniques like [**Scalar Quantization**](https://qdrant.tech/documentation/guides/quantization/#scalar-quantization) and [**Product Quantization**](https://qdrant.tech/documentation/guides/quantization/#product-quantization) are also popular alternatives when optimizing vector compression.
However, binary quantization isn’t the only available option. Techniques like [**Scalar Quantization**](https://qdrant.tech/documentation/manage-data/quantization/#scalar-quantization) and [**Product Quantization**](https://qdrant.tech/documentation/manage-data/quantization/#product-quantization) are also popular alternatives when optimizing vector compression.
You can set up your chosen quantization method using the `quantization_config` parameter when creating a new collection:
@@ -414,7 +414,7 @@ client.create_collection(
We recommend using sharding and replication together so that your data is both split across nodes and replicated for availability.
For more details on features like **user-defined sharding, node failure recovery**, and **consistency guarantees**, see our guide on [Distributed Deployment.](https://qdrant.tech/documentation/guides/distributed_deployment/)
For more details on features like **user-defined sharding, node failure recovery**, and **consistency guarantees**, see our guide on [Distributed Deployment.](https://qdrant.tech/documentation/operations/distributed_deployment/)
## Multitenancy: Data Isolation for Multi-Tenant Architectures
@@ -477,7 +477,7 @@ You can easily setup your access tokens and secure access to sensitive data thro
<img src="/articles_data/what-is-a-vector-database/jwt-web-ui.png" alt="Qdrant Web UI for generating a new access token." width="1000">
By default, Qdrant instances are **unsecured**, so it's important to configure security measures before moving to production. To learn more about how to configure security for your Qdrant instance and other advanced options, please check out the [official Qdrant documentation on security.](https://qdrant.tech/documentation/guides/security/)
By default, Qdrant instances are **unsecured**, so it's important to configure security measures before moving to production. To learn more about how to configure security for your Qdrant instance and other advanced options, please check out the [official Qdrant documentation on security.](https://qdrant.tech/documentation/operations/security/)
## Time to Experiment
@@ -28,7 +28,7 @@ When working with high-dimensional vectors, such as embeddings from providers li
With 1 million vectors needing around 6 GB of memory, as your dataset grows to multiple **millions of vectors**, the memory and processing demands increase significantly.
To understand why this process is so computationally demanding, let's take a look at the nature of the [HNSW index](https://qdrant.tech/documentation/concepts/indexing/#vector-index).
To understand why this process is so computationally demanding, let's take a look at the nature of the [HNSW index](https://qdrant.tech/documentation/manage-data/indexing/#vector-index).
The **HNSW (Hierarchical Navigable Small World) index** organizes vectors in a layered graph, connecting each vector to its nearest neighbors. At each layer, the algorithm narrows down the search area until it reaches the lower layers, where it efficiently finds the closest matches to the query.
@@ -54,7 +54,7 @@ There are several methods to achieve this, and here we will focus on three main
![](/articles_data/what-is-vector-quantization/astronaut-mars.jpg)
In Qdrant, each dimension is represented by a `float32` value, which uses **4 bytes** of memory. When using [Scalar Quantization](https://qdrant.tech/documentation/guides/quantization/#scalar-quantization), we map our vectors to a range that the smaller `int8` type can represent. An `int8` is only **1 byte** and can represent 256 values (from -128 to 127, or 0 to 255). This results in a **75% reduction** in memory size.
In Qdrant, each dimension is represented by a `float32` value, which uses **4 bytes** of memory. When using [Scalar Quantization](https://qdrant.tech/documentation/manage-data/quantization/#scalar-quantization), we map our vectors to a range that the smaller `int8` type can represent. An `int8` is only **1 byte** and can represent 256 values (from -128 to 127, or 0 to 255). This results in a **75% reduction** in memory size.
For example, if our data lies in the range of -1.0 to 1.0, Scalar Quantization will transform these values to a range that `int8` can represent, i.e., within -128 to 127. The system **maps** the `float32` values into this range.
@@ -107,7 +107,7 @@ While the performance gains of Scalar Quantization may not match those achieved
![Astronaut in surreal white environment](/articles_data/what-is-vector-quantization/astronaut-white-surreal.jpg)
[Binary Quantization](https://qdrant.tech/documentation/guides/quantization/#binary-quantization) is an excellent option if you're looking to **reduce memory** usage while also achieving a significant **boost in speed**. It works by converting high-dimensional vectors into simple binary (0 or 1) representations.
[Binary Quantization](https://qdrant.tech/documentation/manage-data/quantization/#binary-quantization) is an excellent option if you're looking to **reduce memory** usage while also achieving a significant **boost in speed**. It works by converting high-dimensional vectors into simple binary (0 or 1) representations.
- Values greater than zero are converted to 1.
- Values less than or equal to zero are converted to 0.
@@ -176,7 +176,7 @@ If you're interested in exploring Binary Quantization in more detail—including
![](/articles_data/what-is-vector-quantization/astronaut-centroids.jpg)
[Product Quantization](https://qdrant.tech/documentation/guides/quantization/#product-quantization) is a method used to compress high-dimensional vectors by representing them with a smaller set of representative points.
[Product Quantization](https://qdrant.tech/documentation/manage-data/quantization/#product-quantization) is a method used to compress high-dimensional vectors by representing them with a smaller set of representative points.
The process begins by splitting the original high-dimensional vectors into smaller **sub-vectors.** Each sub-vector represents a segment of the original vector, capturing different characteristics of the data.
@@ -268,7 +268,7 @@ Product Quantization can significantly reduce memory usage, potentially offering
If your application requires high precision or real-time performance, Product Quantization may not be the best choice. However, if **memory savings** are critical and some accuracy loss is acceptable, it could still be an ideal solution.
Here’s a comparison of speed, accuracy, and compression for all three methods, adapted from [Qdrant's documentation](https://qdrant.tech/documentation/guides/quantization/#how-to-choose-the-right-quantization-method):
Here’s a comparison of speed, accuracy, and compression for all three methods, adapted from [Qdrant's documentation](https://qdrant.tech/documentation/manage-data/quantization/#how-to-choose-the-right-quantization-method):
| Quantization method | Accuracy | Speed | Compression |
|---------------------|----------|------------|-------------|
@@ -519,7 +519,7 @@ Here are some final thoughts to help you choose the right quantization method fo
### Learn More
If you want to learn more about improving accuracy, memory efficiency, and speed when using quantization in Qdrant, we have a dedicated [Quantization tips](https://qdrant.tech/documentation/guides/quantization/#quantization-tips) section in our docs that explains all the quantization tips you can use to enhance your results.
If you want to learn more about improving accuracy, memory efficiency, and speed when using quantization in Qdrant, we have a dedicated [Quantization tips](https://qdrant.tech/documentation/manage-data/quantization/#quantization-tips) section in our docs that explains all the quantization tips you can use to enhance your results.
Learn more about optimizing real-time precision with oversampling in Binary Quantization by watching this interview with Qdrant’s CTO, Andrey Vasnetsov:
@@ -133,7 +133,7 @@ The LLM is typically a model like GPT, BART or T5, trained on massive datasets t
![How a Generator works](/articles_data/what-is-rag-in-ai/how-generation-works.png)
The retriever and generator don't operate in isolation. The image bellow shows how the output of the retrieval feeds the generator to produce the final generated response.
The retriever and generator don't operate in isolation. The image below shows how the output of the retrieval feeds the generator to produce the final generated response.
![The entire architecture of a RAG system](/articles_data/what-is-rag-in-ai/rag-system.jpg)
@@ -18,7 +18,7 @@ As you can see from the charts, there are three main patterns:
- **Speed downturn** - some engines struggle to keep high RPS, it might be related to the requirement of building a filtering mask for the dataset, as described above.
- **Accuracy collapse** - some engines are loosing accuracy dramatically under some filters. It is related to the fact that the HNSW graph becomes disconnected, and the search becomes unreliable.
- **Accuracy collapse** - some engines are losing accuracy dramatically under some filters. It is related to the fact that the HNSW graph becomes disconnected, and the search becomes unreliable.
Qdrant avoids all these problems and also benefits from the speed boost, as it implements an advanced [query planning strategy](/documentation/search/#query-planning).
@@ -17,7 +17,7 @@ Unlisted: false
Most of the engines have improved since [our last run](/benchmarks/single-node-speed-benchmark-2022/). Both life and software have trade-offs but some clearly do better:
* **`Qdrant` achives highest RPS and lowest latencies in almost all the scenarios, no matter the precision threshold and the metric we choose.** It has also shown 4x RPS gains on one of the datasets.
* **`Qdrant` achieves highest RPS and lowest latencies in almost all the scenarios, no matter the precision threshold and the metric we choose.** It has also shown 4x RPS gains on one of the datasets.
* `Elasticsearch` has become considerably fast for many cases but it's very slow in terms of indexing time. It can be 10x slower when storing 10M+ vectors of 96 dimensions! (32mins vs 5.5 hrs)
* `Milvus` is the fastest when it comes to indexing time and maintains good precision. However, it's not on-par with others when it comes to RPS or latency when you have higher dimension embeddings or more number of vectors.
* `Redis` is able to achieve good RPS but mostly for lower precision. It also achieved low latency with single thread, however its latency goes up quickly with more parallel requests. Part of this speed gain comes from their custom protocol.
+7 -7
View File
@@ -46,9 +46,9 @@ In response, our 2025 roadmap centered on four tightly connected capability area
In 2025, we focused on giving teams explicit control over retrieval quality as applications moved beyond basic semantic search. Our new capabilities make relevance more explainable, tunable, and aligned with real user intent, especially in agentic and hybrid search workflows.
**Related enhancements:**
• [Score-Boosting Reranking](https://qdrant.tech/documentation/concepts/search-relevance/#score-boosting) allowing the blending of vector similarity with business signals
• [Full-Text Filtering](https://qdrant.tech/documentation/concepts/filtering/) which brought native multilingual tokenization, stemming, and phrase matching
• [ACORN algorithm](https://qdrant.tech/documentation/concepts/search/#acorn-search-algorithm) for higher-quality filtered HNSW queries
• [Score-Boosting Reranking](https://qdrant.tech/documentation/search/search-relevance/#score-boosting) allowing the blending of vector similarity with business signals
• [Full-Text Filtering](https://qdrant.tech/documentation/search/filtering/) which brought native multilingual tokenization, stemming, and phrase matching
• [ACORN algorithm](https://qdrant.tech/documentation/search/search/#acorn-search-algorithm) for higher-quality filtered HNSW queries
• [Maximal Marginal Relevance (MMR)](https://qdrant.tech/blog/mmr-diversity-aware-reranking/) to balance relevance and diversity
• ASCII folding for improved multilingual recall
@@ -57,19 +57,19 @@ In 2025, we focused on giving teams explicit control over retrieval quality as a
To support large, cost-sensitive workloads, we targeted the biggest performance bottlenecks in production systems. New improvements help teams scale indexing and querying without over-provisioning memory or compute.
**Related enhancements:**
• [GPU-Accelerated HNSW Indexing](https://qdrant.tech/documentation/guides/running-with-gpu/) unlocks up to an order-of-magnitude faster ingestion
• [Inline Storage](https://qdrant.tech/documentation/guides/optimize/#inline-storage-in-hnsw-index) embedded quantized vectors directly into the graph to dramatically improve disk-based search performance
• [GPU-Accelerated HNSW Indexing](https://qdrant.tech/documentation/operations/running-with-gpu/) unlocks up to an order-of-magnitude faster ingestion
• [Inline Storage](https://qdrant.tech/documentation/operations/optimize/#inline-storage-in-hnsw-index) embedded quantized vectors directly into the graph to dramatically improve disk-based search performance
• [Custom storage engine](https://qdrant.tech/articles/gridstore-key-value-storage/) optimized for predictable low-latency access
• [Incremental HNSW indexing](https://qdrant.tech/documentation/database-tutorials/bulk-upload/?q=incremental+hnsw#choose-an-indexing-strategy) for upsert-heavy workloads
• HNSW graph compression to reduce memory footprint
• Expanded [Quantization](https://qdrant.tech/documentation/guides/quantization/#15-bit-and-2-bit-quantization]) options, including 1.5-bit, 2-bit, and asymmetric quantization
• Expanded [Quantization](https://qdrant.tech/documentation/manage-data/quantization/#15-bit-and-2-bit-quantization]) options, including 1.5-bit, 2-bit, and asymmetric quantization
### Enterprise Scaling & Isolation
As Qdrant became shared infrastructure inside larger organizations, we focused on multitenancy, governance, and enterprise needs.
**Related enhancements:**
• [Tiered Multitenancy](https://qdrant.tech/documentation/guides/multitenancy/#tiered-multitenancy) enables efficient support for both small and large tenants within a single system
• [Tiered Multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/#tiered-multitenancy) enables efficient support for both small and large tenants within a single system
• [Single Sign-On (SSO) and role-based access control (RBAC)](https://qdrant.tech/enterprise-solutions/)
• Granular database API keys
• [Terraform-enabled Cloud API](https://qdrant.tech/enterprise-solutions/) for automation and governance
@@ -41,8 +41,8 @@ Ready to experience the benefits of Qdrant on Azure Marketplace? Getting started
1. **Visit the Azure Marketplace**: Navigate to [Qdrant's Marketplace listing](https://azuremarketplace.microsoft.com/en-en/marketplace/apps/qdrantsolutionsgmbh1698769709989.qdrant-db).
2. **Deploy Qdrant**: Follow the simple deployment instructions to set up your instance.
3. **Start Using Qdrant**: Once deployed, start exploring the [features and capabilities of Qdrant](/documentation/concepts/) on Azure.
4. **Read Documentation**: Read Qdrant's [Documentation](/documentation/) and build demo apps using [Tutorials](/documentation/tutorials/).
3. **Start Using Qdrant**: Once deployed, start exploring the [features and capabilities of Qdrant](/documentation/overview/) on Azure.
4. **Read Documentation**: Read Qdrant's [Documentation](/documentation/) and build demo apps using [Tutorials](/documentation/tutorials-lp-overview/).
## Join Us on this Exciting Journey:
@@ -19,7 +19,7 @@ We’ve launched the **beta** of our Qdrant **Vector Data Migration Tool**, desi
This powerful tool streams all vectors from a source collection to a target Qdrant instance in live batches. It supports migrations from one Qdrant deployment to another, including from open source to Qdrant Cloud or between cloud regions. But that's not all. You can also migrate your data from other vector databases directly into Qdrant. All with a single command.
Unlike Qdrant’s included [snapshot migration method](https://qdrant.tech/documentation/concepts/snapshots/), which requires consistent node-specific snapshots, our migration tool enables you to easily migrate data between different Qdrant database clusters in streaming batches. The only requirement is that the vector size and distance function must match.
Unlike Qdrant’s included [snapshot migration method](https://qdrant.tech/documentation/operations/snapshots/), which requires consistent node-specific snapshots, our migration tool enables you to easily migrate data between different Qdrant database clusters in streaming batches. The only requirement is that the vector size and distance function must match.
This is especially useful if you want to change the collection configuration on the target, for example by choosing a different replication factor or quantization method.
@@ -71,7 +71,7 @@ By using [Qdrant Cloud](https://qdrant.tech/cloud/), \&AI avoided the need to ma
"Patent litigation has huge stakes, one result could influence a billion-dollar case," said Turner. "Accuracy is the top priority, and Qdrant let us optimize for that without compromising on cost or performance."
Qdrant’s support for [payload filters](https://qdrant.tech/documentation/concepts/filtering/), [multitenancy](https://qdrant.tech/documentation/guides/multiple-partitions/), and quantization let \&AI optimize deeply. Their AI patent agent, Andy, uses natural language to guide attorneys through patent analysis tasks, drastically cutting time-to-result.
Qdrant’s support for [payload filters](https://qdrant.tech/documentation/search/filtering/), [multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/), and quantization let \&AI optimize deeply. Their AI patent agent, Andy, uses natural language to guide attorneys through patent analysis tasks, drastically cutting time-to-result.
*"With Qdrant, we scaled to a billion vectors and still respond in sub-second latency. That lets us power workflows that used to take hours in just a few minutes."*
@@ -56,14 +56,14 @@ Beyond coding, Anima uses Qdrant to understand documents at scale. By working wi
Several factors made Qdrant a strong fit for healthcare workloads.
[Deployment flexibility](https://qdrant.tech/documentation/guides/installation/) was non-negotiable. Anima requires data to remain at rest in the UK, allowing them to make strong guarantees to customers about compliance and residency. Qdrant’s self-hosted and region-controlled deployment options enabled this without compromising performance.
[Deployment flexibility](https://qdrant.tech/documentation/operations/installation/) was non-negotiable. Anima requires data to remain at rest in the UK, allowing them to make strong guarantees to customers about compliance and residency. Qdrant’s self-hosted and region-controlled deployment options enabled this without compromising performance.
Cost predictability also played a critical role. With a fixed infrastructure cost for vector search, Anima could use retrieval across multiple passes in their pipelines. This unlocked higher-quality results without eroding margins.
>“Knowing that retrieval is reliable and low cost changed how we build. We do not think twice about using vector search as part of our pipelines.”
-Colin Cooke, Lead AI Engineer, Anima Health
Finally, Qdrant’s vector-native capabilities mattered. [Payload-based filtering](https://qdrant.tech/documentation/concepts/payload/) allows Anima to scope searches precisely across different electronic health record systems. [Multivector support](https://qdrant.tech/documentation/concepts/payload/) enables experimentation with multiple embedding strategies and providers, reducing long-term lock-in and easing future transitions.
Finally, Qdrant’s vector-native capabilities mattered. [Payload-based filtering](https://qdrant.tech/documentation/manage-data/payload/) allows Anima to scope searches precisely across different electronic health record systems. [Multivector support](https://qdrant.tech/documentation/manage-data/payload/) enables experimentation with multiple embedding strategies and providers, reducing long-term lock-in and easing future transitions.
### Results: Scalable, privacy-first AI in production
@@ -55,8 +55,8 @@ With over a billion vectors, post-filtering meant searching far more data than n
Qdrant stood out for a few key reasons:
* [Multitenancy](https://qdrant.tech/documentation/guides/multitenancy/) with payload-based partitioning, allowing searches to be scoped by client, product, or category at query time
* [Quantization](https://qdrant.tech/documentation/guides/quantization/), enabling dramatic reductions in storage and RAM requirements which translates directly to cost
* [Multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/) with payload-based partitioning, allowing searches to be scoped by client, product, or category at query time
* [Quantization](https://qdrant.tech/documentation/manage-data/quantization/), enabling dramatic reductions in storage and RAM requirements which translates directly to cost
* [Hybrid cloud deployment](https://qdrant.tech/hybrid-cloud/), running inside Bazaarvoice’s VPC on Kubernetes
* Operational simplicity, eliminating manual partition management entirely
@@ -167,7 +167,7 @@ They aim to work on Perspective feed next and say
![perspective-feed-with-qdrant](/case-studies/dailymotion/perspective-feed-qdrant.jpg)
The team is also interested in leveraging advanced features like [Qdrant’s Discovery API](/documentation/concepts/explore/#recommendation-api) to promote exploration of content to enable finding not only similar but dissimilar content too by using positive and negative vectors in the queries and making it work with the existing collaborative recommendation model.
The team is also interested in leveraging advanced features like [Qdrant’s Discovery API](/documentation/search/explore/#recommendation-api) to promote exploration of content to enable finding not only similar but dissimilar content too by using positive and negative vectors in the queries and making it work with the existing collaborative recommendation model.
### References
@@ -76,7 +76,7 @@ When Deutsche Telekom began searching for a scalable, high-performance vector da
The team structured its evaluation around two key metrics:
1. **Qualitative metrics**: developer experience, ease of use, memory efficiency features.
2. **Operational simplicity**: how well it fit into their PaaS-first approach and [multitenancy requirements](https://qdrant.tech/documentation/guides/multiple-partitions/).
2. **Operational simplicity**: how well it fit into their PaaS-first approach and [multitenancy requirements](https://qdrant.tech/documentation/manage-data/multitenancy/).
Deutsche Telekom's engineers also cited several standout features that made Qdrant the right fit:
@@ -41,13 +41,13 @@ Retail and warehouse environments are bandwidth-constrained and heterogeneous. A
* Operate a vector store at enterprise scale: thousands of locations → thousands of cameras, accumulating into tens to hundreds of billions of vectors and multi-terabyte storage.
### Why Qdrant: Performance headroom and operational control
Dragonfruit chose the open-source version of Qdrant as its vector search engine to meet the twin pressures of real-time reads and high-velocity writes. In head-to-head experiments, Qdrant delivered the QPS targets they needed while giving the team granular, [per-collection](https://qdrant.tech/documentation/concepts/collections/) tuning to match workload diversity.
Dragonfruit chose the open-source version of Qdrant as its vector search engine to meet the twin pressures of real-time reads and high-velocity writes. In head-to-head experiments, Qdrant delivered the QPS targets they needed while giving the team granular, [per-collection](https://qdrant.tech/documentation/manage-data/collections/) tuning to match workload diversity.
Key reasons the team highlighted:
* **Per-collection configurability.** Collections with heavy reads and low writes use different settings than write-heavy pipelines. Tuning shard counts and HNSW parameters by collection helped hit latency Service Level Objectives (SLOs) without overprovisioning.
* **Efficient numeric formats.** For most vision workloads, [float16](https://qdrant.tech/documentation/concepts/vectors/) vectors were sufficient, improving memory efficiency and cache behavior with no material loss in retrieval accuracy for their use cases.
* **Efficient numeric formats.** For most vision workloads, [float16](https://qdrant.tech/documentation/manage-data/vectors/) vectors were sufficient, improving memory efficiency and cache behavior with no material loss in retrieval accuracy for their use cases.
* **Open source and ecosystem fit.** Qdrant’s OSS model aligned with Dragonfruit’s platform strategy and let them co-evolve the deployment with their edge and cloud stack.
@@ -83,8 +83,8 @@ billing and increase security by having the instance live within the same VPC.
2. **Scale and optimize:** As the load grew, Dust started to take advantage of Qdrant’s
features to tune the setup for optimization and scale. They started to look into
how they map and cache data, as well as applying some of Qdrant’s [built-in
compression features](/documentation/guides/quantization/). In particular, Dust leveraged the control of the [MMAP
payload threshold](/documentation/concepts/storage/#configuring-memmap-storage) as well as [Scalar Quantization](/articles/scalar-quantization/), which enabled Dust to manage
compression features](/documentation/manage-data/quantization/). In particular, Dust leveraged the control of the [MMAP
payload threshold](/documentation/manage-data/storage/#configuring-memmap-storage) as well as [Scalar Quantization](/articles/scalar-quantization/), which enabled Dust to manage
the balance between storing vectors on disk and keeping quantized vectors in RAM,
more effectively. “This allowed us to scale smoothly from there,” Polu says.
@@ -47,7 +47,7 @@ For the engineering team, these failures had two serious implications. First, mi
After evaluating alternatives, Fieldy selected [Qdrant](http://qdrant.tech) for its stability, straightforward configuration, and suitability for self-hosted deployment. They opted to run Qdrant in the same environment as their backend services, ensuring low-latency access and avoiding the cross-region connectivity issues that had contributed to failures in the previous architecture.
The migration process took less than a week. The team reused their existing vector schema to avoid redesigning their indexing logic during the transition. Embeddings, generated using Cohere’s multilingual v3 model, were batched locally before ingestion, replacing the integrated embedding calls previously handled within Weaviate. Following [Qdrant best practices](https://qdrant.tech/documentation/guides/optimize/), they disabled indexing during bulk import to maximize throughput, re-enabling it only after the migration was complete. The backend API calls were updated to use Qdrant’s gRPC interface, and [hybrid search](https://qdrant.tech/articles/hybrid-search/) with BM25 was combined with dense vector retrieval via [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/concepts/hybrid-queries/#hybrid-search) for relevance scoring.
The migration process took less than a week. The team reused their existing vector schema to avoid redesigning their indexing logic during the transition. Embeddings, generated using Cohere’s multilingual v3 model, were batched locally before ingestion, replacing the integrated embedding calls previously handled within Weaviate. Following [Qdrant best practices](https://qdrant.tech/documentation/operations/optimize/), they disabled indexing during bulk import to maximize throughput, re-enabling it only after the migration was complete. The backend API calls were updated to use Qdrant’s gRPC interface, and [hybrid search](https://qdrant.tech/articles/hybrid-search/) with BM25 was combined with dense vector retrieval via [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/search/hybrid-queries/#hybrid-search) for relevance scoring.
### Architecture after migration
@@ -63,4 +63,4 @@ Fieldy also achieved a two-thirds reduction in infrastructure costs after moving
### Next steps for retrieval quality
With reliability and cost efficiency achieved, Fieldy’s engineering focus is shifting toward retrieval quality. Planned improvements include adding [location filtering](https://qdrant.tech/documentation/concepts/filtering/#geo) and [datetime filtering](https://qdrant.tech/documentation/concepts/filtering/#datetime-range) within Qdrant to refine result sets, experimenting with late chunking strategies, and testing parallel hybrid searches to increase recall on complex multi-faceted queries. The team is also exploring embedding summaries alongside raw transcript segments to improve retrieval performance on high-level or thematic searches.
With reliability and cost efficiency achieved, Fieldy’s engineering focus is shifting toward retrieval quality. Planned improvements include adding [location filtering](https://qdrant.tech/documentation/search/filtering/#geo) and [datetime filtering](https://qdrant.tech/documentation/search/filtering/#datetime-range) within Qdrant to refine result sets, experimenting with late chunking strategies, and testing parallel hybrid searches to increase recall on complex multi-faceted queries. The team is also exploring embedding summaries alongside raw transcript segments to improve retrieval performance on high-level or thematic searches.
@@ -4,7 +4,6 @@ title: "Building real-time multimodal similarity search in Flipkart Trust & Safe
short_description: "Tackling fraud and abuse with scalable similarity search."
description: "Tackling fraud and abuse with scalable similarity search."
preview_image: /blog/case-study-flipkart/social_preview_partnership-flipkart.png
social_preview_image: /blog/case-study/social_preview_partnership-flipkart.png
date: 2026-01-09
author: "Daniel Azoulai"
featured: true
@@ -48,7 +48,7 @@ Instead of optimizing around a single response time target, GlassDollar focused
## Accuracy Meant End-to-end Results, not Just Retriever Scores
For GlassDollar, accuracy was measured at the workflow level: did the system surface the right companies for a given enterprise need, and did users act on the results. That meant optimizing the full architecture, including [query expansion](https://qdrant.tech/documentation/concepts/hybrid-queries/), retrieval, ranking, and contextual ranking, rather than chasing marginal gains in any single component. Faster retrieval mattered most because it enabled more queries to improve query expansion, which raised recall and improved the final shortlist quality.
For GlassDollar, accuracy was measured at the workflow level: did the system surface the right companies for a given enterprise need, and did users act on the results. That meant optimizing the full architecture, including [query expansion](https://qdrant.tech/documentation/search/hybrid-queries/), retrieval, ranking, and contextual ranking, rather than chasing marginal gains in any single component. Faster retrieval mattered most because it enabled more queries to improve query expansion, which raised recall and improved the final shortlist quality.
## Contextual Embeddings Improved Matching Between User Intent and Company Descriptions
@@ -47,9 +47,9 @@ After evaluating multiple vector databases, the Connectivity Platform team selec
The decision came down to a combination of search quality, performance, and operational fit.
Qdrant’s hybrid search capabilities were a key factor. By supporting both dense vectors for semantic search and sparse vectors for keyword-based retrieval—combined using [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/concepts/hybrid-queries/#reciprocal-rank-fusion-rrf), the team could address both conceptual questions and exact-match queries in a single system. Named Vectors made it possible to manage multiple vector types within the same collection.
Qdrant’s hybrid search capabilities were a key factor. By supporting both dense vectors for semantic search and sparse vectors for keyword-based retrieval—combined using [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/search/hybrid-queries/#reciprocal-rank-fusion-rrf), the team could address both conceptual questions and exact-match queries in a single system. Named Vectors made it possible to manage multiple vector types within the same collection.
Performance was another major consideration. Qdrant’s [Rust-based architecture](https://qdrant.tech/articles/why-rust/), efficient [HNSW implementation](https://qdrant.tech/course/essentials/day-2/what-is-hnsw/), and support for [scalar quantization (INT8)](https://qdrant.tech/documentation/guides/quantization/#scalar-quantization) provided low-latency search while optimizing memory usage. This was crucial for an internal service expected to scale over time.
Performance was another major consideration. Qdrant’s [Rust-based architecture](https://qdrant.tech/articles/why-rust/), efficient [HNSW implementation](https://qdrant.tech/course/essentials/day-2/what-is-hnsw/), and support for [scalar quantization (INT8)](https://qdrant.tech/documentation/manage-data/quantization/#scalar-quantization) provided low-latency search while optimizing memory usage. This was crucial for an internal service expected to scale over time.
From an operational standpoint, Qdrant fit naturally into Kakao’s environment. Its single-binary design simplified deployment, it ran reliably on Kubernetes, and it allowed Kakao to retain full control over data by self-hosting within internal infrastructure.
@@ -63,7 +63,7 @@ The team integrated Qdrant using the asynchronous Python client (`AsyncQdrantCli
Collections were designed around data sources, with separate collections for internal technical documentation, historical inquiry data, and a semantic cache used to speed up repeated queries. Metadata filtering allows the system to narrow search scope by service or time period, while maintaining fast response times.
Each collection stores both dense and sparse vectors using [Named Vectors](https://qdrant.tech/documentation/concepts/vectors/#named-vectors). Hybrid search results are merged using RRF to produce more accurate answers across different query types.
Each collection stores both dense and sparse vectors using [Named Vectors](https://qdrant.tech/documentation/manage-data/vectors/#named-vectors). Hybrid search results are merged using RRF to produce more accurate answers across different query types.
An automated indexing pipeline handles document ingestion end-to-end. This ranges from cleansing and chunking, to embedding generation, to batch upserts into Qdrant.
@@ -38,7 +38,7 @@ Enterprises in regulated sectors deal with extensive, complex documentation feat
One component of the build was the vector database. Lettria evaluated Weaviate, Milvus, and Qdrant based on their hybrid search capability, deployment simplicity (Docker, Kubernetes), and search performance (latency, RAM usage).
Ultimately, Lettria chose Qdrant. First, it had a simple Kubernetes deployment, superior latency and lower memory footprint in competitive benchmarks. Additionally, there were unique features, such as the grouping API and detailed [payload indexing](https://qdrant.tech/documentation/concepts/payload/), that made Qdrant stand out.
Ultimately, Lettria chose Qdrant. First, it had a simple Kubernetes deployment, superior latency and lower memory footprint in competitive benchmarks. Additionally, there were unique features, such as the grouping API and detailed [payload indexing](https://qdrant.tech/documentation/manage-data/payload/), that made Qdrant stand out.
## Building the document understanding and extraction pipeline
@@ -288,7 +288,7 @@ Note that they duplicate the english tags for string properties as they are cons
#### Filtering
based on a filter definition. Lettria flattens properties on Neo4J so that they can use similar filters. The nested structure ([more on that here](https://qdrant.tech/documentation/concepts/filtering/#nested-key)) {"foo": { "bar": "qux" }} is kept in Qdrant and dot separated in NeoJ: foo.bar=qux so that they can perform match queries with similar keys from Qdrant. This introduces some complexity as they need to be careful in the handling of url and other 'dot rich' values in properties.
based on a filter definition. Lettria flattens properties on Neo4J so that they can use similar filters. The nested structure ([more on that here](https://qdrant.tech/documentation/search/filtering/#nested-key)) {"foo": { "bar": "qux" }} is kept in Qdrant and dot separated in NeoJ: foo.bar=qux so that they can perform match queries with similar keys from Qdrant. This introduces some complexity as they need to be careful in the handling of url and other 'dot rich' values in properties.
If they want to filter based on the onto:surface value, the same keys are used in Qdrant and Neo4J:
@@ -48,7 +48,7 @@ Everything changed when OpenAI released its embedding model. Instead of hoping u
>"That was a transformational moment for us. Now we could have hundreds of help articles ingested in the system, and a user can ask a question, and we can answer that really specifically and cheaply and quickly."
Over time, the team also learned that semantic search was strong but not universally sufficient, especially when tickets contained product names, error codes, or specific identifiers that benefit from lexical matching. That realization led My AskAI toward experimentation with [hybrid search](https://qdrant.tech/documentation/concepts/hybrid-queries/) as a way to blend semantic similarity with keyword signals, while keeping the operational footprint small.
Over time, the team also learned that semantic search was strong but not universally sufficient, especially when tickets contained product names, error codes, or specific identifiers that benefit from lexical matching. That realization led My AskAI toward experimentation with [hybrid search](https://qdrant.tech/documentation/search/hybrid-queries/) as a way to blend semantic similarity with keyword signals, while keeping the operational footprint small.
## Why My AskAI Chose Qdrant: Scalability, Integrations, and Developer Experience
@@ -74,7 +74,7 @@ Once on Qdrant Cloud, My AskAI leaned into a workflow where scaling and day-to-d
The ideal infrastructure, as Alex puts it, is the kind you don't have to think about. "I didn't want to have to think about it."
My AskAI also began running customer-specific proofs of concept for [hybrid search](https://qdrant.tech/documentation/concepts/hybrid-queries/), aiming to find the right blend that improved retrieval in the edge cases where semantic-only results were not enough. Before Qdrant, managing hybrid search had required spinning up separate infrastructure on AWS and handling reranking externally. With Qdrant, the team could enable hybrid search per collection and iterate without managing additional systems.
My AskAI also began running customer-specific proofs of concept for [hybrid search](https://qdrant.tech/documentation/search/hybrid-queries/), aiming to find the right blend that improved retrieval in the edge cases where semantic-only results were not enough. Before Qdrant, managing hybrid search had required spinning up separate infrastructure on AWS and handling reranking externally. With Qdrant, the team could enable hybrid search per collection and iterate without managing additional systems.
>"Just being able to turn on hybrid search is super useful. It removes that headache and pushes management of hybrid search down to the vendor."
@@ -50,14 +50,14 @@ As part of their selection process, Nyris evaluated several critical factors to
- **Insert Speed**: Nyris assessed how quickly data could be inserted into the database, including the performance during simultaneous data ingests and query requests. Qdrant excelled in this area, providing the necessary efficiency for their operations.
- **Total Cost of Ownership**: Nyris analyzed the infrastructure costs and licensing fees associated with each solution. Qdrant offered a competitive total cost of ownership, making it an economically viable option.
- **Data Sovereignty**: The ability to deploy Qdrant in their own clusters was a key aspect for Nyris, ensuring they maintained control over their data and complied with relevant data sovereignty requirements.
- **Dedicated Vector Search Engine:** One of the key advantages of Qdrant, as Lukasson highlights, is its specialization as a dedicated, native vector search engine. "Qdrant, being purpose-built for vector search, can introduce relevant features much faster, like [quantization](https://qdrant.tech/documentation/guides/quantization/), integer8 support, and float32 rescoring. These advancements make searches more precise and cost-effective without sacrificing accuracy—exactly what Nyris needs," said Lukasson. "When optimizing for search accuracy and speed, compromises aren't an option. Just as you wouldn't use a truck to race in Formula 1, we needed a solution designed specifically for vector search, not just a general database with vector search tacked on. With every Qdrant release, we gain new, tailored features that directly enhance our use case.”
- **Dedicated Vector Search Engine:** One of the key advantages of Qdrant, as Lukasson highlights, is its specialization as a dedicated, native vector search engine. "Qdrant, being purpose-built for vector search, can introduce relevant features much faster, like [quantization](https://qdrant.tech/documentation/manage-data/quantization/), integer8 support, and float32 rescoring. These advancements make searches more precise and cost-effective without sacrificing accuracy—exactly what Nyris needs," said Lukasson. "When optimizing for search accuracy and speed, compromises aren't an option. Just as you wouldn't use a truck to race in Formula 1, we needed a solution designed specifically for vector search, not just a general database with vector search tacked on. With every Qdrant release, we gain new, tailored features that directly enhance our use case.”
## Key Benefits of Qdrant in Production
Nyris has found several aspects of Qdrant particularly beneficial in their production environment:
- **Enhanced Security with JWT**: [JSON Web Tokens](https://qdrant.tech/documentation/guides/security/#granular-access-control-with-jwt) provide enhanced security and performance, critical for safeguarding their data.
- **Seamless Scalability**: Qdrant's ability to [scale effortlessly across nodes](https://qdrant.tech/documentation/guides/distributed_deployment/) ensures consistent high performance, even as Nyris's data volume grows.
- **Enhanced Security with JWT**: [JSON Web Tokens](https://qdrant.tech/documentation/operations/security/#granular-access-control-with-jwt) provide enhanced security and performance, critical for safeguarding their data.
- **Seamless Scalability**: Qdrant's ability to [scale effortlessly across nodes](https://qdrant.tech/documentation/operations/distributed_deployment/) ensures consistent high performance, even as Nyris's data volume grows.
- **Flexible Search Options**: The availability of both graph-based and brute-force search methods offers Nyris the flexibility to tailor the search approach to specific use case requirements.
- **Versatile Data Handling**: Qdrant imposes almost no restrictions on data types and vector sizes, allowing Nyris to manage diverse and complex datasets effectively.
- **Built with Rust**: The use of [Rust](https://qdrant.tech/articles/why-rust/) ensures superior performance and future-proofing, while its open-source nature allows Nyris to inspect and customize the code as necessary.
@@ -43,7 +43,7 @@ Qdrant’s fast, scalable [vector search](/advanced-search/) enables QA.tech to
## Why QA.tech chose Qdrant for its AI Agent platform
QA.tech’s AI Agents handle high-velocity web actions, requiring efficient real-time operations and scalable infrastructure. The team faced challenges with managing network overhead, CPU load, and the need to store [multiple embeddings](/documentation/concepts/vectors/#multivectors) for different use cases. Qdrant provided the solution to address these issues.
QA.tech’s AI Agents handle high-velocity web actions, requiring efficient real-time operations and scalable infrastructure. The team faced challenges with managing network overhead, CPU load, and the need to store [multiple embeddings](/documentation/manage-data/vectors/#multivectors) for different use cases. Qdrant provided the solution to address these issues.
**Reducing Network Overhead with Batch Operations**
@@ -35,7 +35,7 @@ Qovery’s ambitious vision for the DevOps Copilot ([read more here](https://www
### Seamless Integration of Scalable and Efficient Vector Search
Qovery chose [Qdrant Cloud](https://qdrant.tech/cloud/) after carefully evaluating several options. Romaric Philogène, CEO and co-founder of Qovery, highlighted the importance of open-source credibility, performance, ease of use, and scalability. Qdrant’s native support for [real-time indexing](https://qdrant.tech/documentation/concepts/indexing/) and low-latency queries made it ideal for handling Qovery’s significant data volume and frequency of updates.
Qovery chose [Qdrant Cloud](https://qdrant.tech/cloud/) after carefully evaluating several options. Romaric Philogène, CEO and co-founder of Qovery, highlighted the importance of open-source credibility, performance, ease of use, and scalability. Qdrant’s native support for [real-time indexing](https://qdrant.tech/documentation/manage-data/indexing/) and low-latency queries made it ideal for handling Qovery’s significant data volume and frequency of updates.
The integration process was straightforward, with minimal operational overhead, enabling the Qovery team to focus their resources on enhancing the Copilot's capabilities rather than maintaining complex database infrastructure. With its Rust-based architecture, Qdrant delivered the speed, accuracy, and low resource utilization Qovery required.
@@ -46,8 +46,8 @@ After evaluating several options of vector DBs, including Pinecone, Weaviate, an
- **High Customizability:** Qdrant provided Sprinklr with essential flexibility through high-level abstractions that allowed for extensive customizations. The diverse teams at Sprinklr, working on various GenAI applications, needed a solution that could adapt to different workloads. “The ability to fine-tune configurations at the collection level was crucial for our varied AI applications,” says Sonavane. Qdrant met this need by offering:
- **Configuration for high-speed search** that fine-tunes settings for optimal performance.
- [**Quantized vectors**](https://qdrant.tech/documentation/guides/quantization/) for high-dimensional data workloads
- [**Memory map**](https://qdrant.tech/documentation/concepts/storage/#configuring-memmap-storage) for efficient search optimizing memory usage.
- [**Quantized vectors**](https://qdrant.tech/documentation/manage-data/quantization/) for high-dimensional data workloads
- [**Memory map**](https://qdrant.tech/documentation/manage-data/storage/#configuring-memmap-storage) for efficient search optimizing memory usage.
- **Speed and Cost Efficiency:** Qdrant provided the best combination of speed and cost, making it the most viable solution for Sprinklr’s needs. “We needed a solution that wouldn’t just meet our performance requirements but also keep costs in check, and Qdrant delivered on both fronts,” says Sonavane.
- **Enhanced Monitoring:** Qdrant’s monitoring tools further boosted system efficiency, allowing Sprinklr to maintain high performance across their platforms.
@@ -55,7 +55,7 @@ After evaluating several options of vector DBs, including Pinecone, Weaviate, an
Sprinklr’s transition to Qdrant was carefully managed, starting with 10% of their workloads before gradually scaling up. The transition was seamless, thanks in part to Qdrant’s configurable [Web UI](https://qdrant.tech/documentation/interfaces/web-ui/), which allowed Sprinklr to fully utilize its capabilities within the existing infrastructure.
“Qdrant’s ability to index [multiple vectors](https://qdrant.tech/documentation/concepts/vectors/#multivectors) simultaneously and retrieve and re-rank with precision brought significant improvements to our workflow,” Sonavane remarks. This feature reduced the need for repeated retrieval processes, significantly improving efficiency. Additionally, Qdrant’s [quantization](https://qdrant.tech/documentation/guides/quantization/) and [memory mapping](https://qdrant.tech/documentation/concepts/storage/#configuring-memmap-storage) features enabled Sprinklr to reduce RAM usage, leading to substantial cost savings.
“Qdrant’s ability to index [multiple vectors](https://qdrant.tech/documentation/manage-data/vectors/#multivectors) simultaneously and retrieve and re-rank with precision brought significant improvements to our workflow,” Sonavane remarks. This feature reduced the need for repeated retrieval processes, significantly improving efficiency. Additionally, Qdrant’s [quantization](https://qdrant.tech/documentation/manage-data/quantization/) and [memory mapping](https://qdrant.tech/documentation/manage-data/storage/#configuring-memmap-storage) features enabled Sprinklr to reduce RAM usage, leading to substantial cost savings.
Qdrant now plays a key supportive role in enhancing Sprinklr’s vector search capabilities within its AI-driven applications, which is designed to be cloud- and LLM-agnostic. The platform supports various AI-driven tasks, from retrieval and re-ranking to serving advanced customer experiences. “Retrieval is the foundation of all our AI tasks, and Qdrant’s resilience and speed have made it an integral part of our system,” Sonavane emphasizes. Sprinklr operates [Qdrant as a managed service on AWS](https://qdrant.tech/cloud/), ensuring scalability, reliability, and ease of use.
@@ -81,6 +81,6 @@ Integrating Qdrant into VISUA's quality control operations has delivered measura
#### Expanding Qdrant’s Use Beyond Anomaly Detection
While the primary application of Qdrant is focused on quality control, VISUA's team is actively exploring additional use cases with Qdrant. VISUA's use of Qdrant has inspired new opportunities, notably in content moderation. "The moment we started to experiment with Qdrant, opened up a lot of ideas within the team for new applications,” said Prest on the potential unlocked by Qdrant. For example, this has led them to actively explore the Qdrant [Discovery API](/documentation/concepts/explore/?q=discovery#discovery-api), with an eye on enhancing content moderation processes.
While the primary application of Qdrant is focused on quality control, VISUA's team is actively exploring additional use cases with Qdrant. VISUA's use of Qdrant has inspired new opportunities, notably in content moderation. "The moment we started to experiment with Qdrant, opened up a lot of ideas within the team for new applications,” said Prest on the potential unlocked by Qdrant. For example, this has led them to actively explore the Qdrant [Discovery API](/documentation/search/explore/?q=discovery#discovery-api), with an eye on enhancing content moderation processes.
Beyond content moderation, VISUA is set for significant growth by broadening its copyright infringement detection services. As the demand for detecting a wider range of infringements, like unauthorized use of popular characters on merchandise, increases, VISUA plans to expand its technology capabilities. Qdrant will be pivotal in this expansion, enabling VISUA to meet the complex and growing challenges of moderating copyrighted content effectively and ensuring comprehensive protection for brands and creators.
@@ -27,7 +27,7 @@ tags:
As part of this development, the Voiceflow engineering team was looking for a [vector database](/qdrant-vector-database/) solution to power their RAG setup. They evaluated various vector databases based on several key factors:
- **Performance**: The ability to [handle the scale](/documentation/guides/distributed_deployment/) required by Voiceflow, supporting hundreds of thousands of projects efficiently.
- **Performance**: The ability to [handle the scale](/documentation/operations/distributed_deployment/) required by Voiceflow, supporting hundreds of thousands of projects efficiently.
- **Metadata**: The capability to tag data and chunks and retrieve based on those values, essential for organizing and accessing specific information swiftly.
- **Managed Solution**: The availability of a [managed service](/documentation/cloud/) with automated maintenance, scaling, and security, freeing the team from infrastructure concerns.
@@ -82,7 +82,7 @@ Voiceflow achieved significant improvements and efficiencies by leveraging Qdran
- **Optimized Performance**: Resolved concerns about retrieval times with a high number of tags by optimizing indexing strategies, achieving efficient performance.
- **Minimal Operational Overhead**: Experienced minimal overhead, streamlining their operational processes.
- **Future-Ready**: Anticipates further innovation in hybrid search with multi-token attention.
- **Multitenancy Support**: Utilized Qdrant's efficient and [isolated data management](/documentation/guides/multiple-partitions/) to support diverse user needs.
- **Multitenancy Support**: Utilized Qdrant's efficient and [isolated data management](/documentation/manage-data/multitenancy/) to support diverse user needs.
Overall, Qdrant's features and infrastructure provided Voiceflow with a stable, scalable, and efficient solution for their data processing and retrieval needs.
@@ -49,7 +49,7 @@ To make this work, Xaver needed a system that could manage knowledge retrieval a
### The solution: Semantic caching, or a two-layer knowledge engine
[Qdrant](https://qdrant.tech/documentation/overview/) was selected after extensive evaluation for several reasons:
Xaver’s AI platform includes a “knowledge engine,” an [indexing](https://qdrant.tech/documentation/concepts/indexing/) and [retrieval](https://qdrant.tech/documentation/beginner-tutorials/retrieval-quality/) layer that feeds contextually relevant insights to both automated and human-assisted consultations. It powers two key functions:
Xaver’s AI platform includes a “knowledge engine,” an [indexing](https://qdrant.tech/documentation/manage-data/indexing/) and [retrieval](https://qdrant.tech/documentation/beginner-tutorials/retrieval-quality/) layer that feeds contextually relevant insights to both automated and human-assisted consultations. It powers two key functions:
1. Automated consultation through AI-led sessions via phone, video avatar, messengers, or web chat.
@@ -27,7 +27,7 @@ Traditional databases, while effective at handling structured data, fall short w
- **Indexing Limitations**: Database indexing methods like B-Trees or hash indexes, typically used in relational databases, are inefficient for high-dimensional data and show poor query performance.
- **Curse of Dimensionality**: As dimensions increase, data points become sparse, and distance metrics like Euclidean distance lose their effectiveness, leading to poor search query performance.
- **Lack of Specialized Algorithms**: Traditional databases do not incorporate advanced algorithms designed to handle high-dimensional data, resulting in slow query processing times.
- **Scalability Challenges**: Managing and querying high-dimensional [vectors](https://qdrant.tech/documentation/concepts/vectors/) require optimized data structures, which traditional databases are not built to handle.
- **Scalability Challenges**: Managing and querying high-dimensional [vectors](https://qdrant.tech/documentation/manage-data/vectors/) require optimized data structures, which traditional databases are not built to handle.
- **Storage Inefficiency**: Traditional databases are not optimized for efficiently storing large volumes of high-dimensional data, facing significant challenges in managing space complexity and [retrieval efficiency](https://qdrant.tech/documentation/tutorials/retrieval-quality/).
Vector databases address these challenges by efficiently storing and querying high-dimensional vectors. They offer features such as high-dimensional vector storage and retrieval, efficient similarity search, sophisticated indexing algorithms, advanced compression techniques, and integration with various machine learning frameworks.
@@ -44,14 +44,14 @@ Qdrant is highly scalable and performant: it can handle billions of vectors effi
### Key Features of Qdrant Vector Database
- **Advanced Similarity Search:** Qdrant supports various similarity [search](https://qdrant.tech/documentation/concepts/search/) metrics like dot product, cosine similarity, Euclidean distance, and Manhattan distance. You can store additional information along with vectors, known as [payload](https://qdrant.tech/documentation/concepts/payload/) in Qdrant terminology. A payload is any JSON formatted data.
- **Advanced Similarity Search:** Qdrant supports various similarity [search](https://qdrant.tech/documentation/search/search/) metrics like dot product, cosine similarity, Euclidean distance, and Manhattan distance. You can store additional information along with vectors, known as [payload](https://qdrant.tech/documentation/manage-data/payload/) in Qdrant terminology. A payload is any JSON formatted data.
- **Built Using Rust:** Qdrant is built with Rust, and leverages its performance and efficiency. Rust is famed for its [memory safety](https://arxiv.org/abs/2206.05503) without the overhead of a garbage collector, and rivals C and C++ in speed.
- **Scaling and Multitenancy**: Qdrant supports both vertical and horizontal scaling and uses the Raft consensus protocol for [distributed deployments](https://qdrant.tech/documentation/guides/distributed_deployment/). Developers can run Qdrant clusters with replicas and shards, and seamlessly scale to handle large datasets. Qdrant also supports [multitenancy](https://qdrant.tech/documentation/guides/multiple-partitions/) where developers can create single collections and partition them using payload.
- **Payload Indexing and Filtering:** Just as Qdrant allows attaching any JSON payload to vectors, it also supports payload indexing and [filtering](https://qdrant.tech/documentation/concepts/filtering/) with a wide range of data types and query conditions, including keyword matching, full-text filtering, numerical ranges, nested object filters, and [geo](https://qdrant.tech/documentation/concepts/filtering/#geo)filtering.
- **Scaling and Multitenancy**: Qdrant supports both vertical and horizontal scaling and uses the Raft consensus protocol for [distributed deployments](https://qdrant.tech/documentation/operations/distributed_deployment/). Developers can run Qdrant clusters with replicas and shards, and seamlessly scale to handle large datasets. Qdrant also supports [multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/) where developers can create single collections and partition them using payload.
- **Payload Indexing and Filtering:** Just as Qdrant allows attaching any JSON payload to vectors, it also supports payload indexing and [filtering](https://qdrant.tech/documentation/search/filtering/) with a wide range of data types and query conditions, including keyword matching, full-text filtering, numerical ranges, nested object filters, and [geo](https://qdrant.tech/documentation/search/filtering/#geo)filtering.
- **Hybrid Search with Sparse Vectors:** Qdrant supports both dense and [sparse vectors](https://qdrant.tech/articles/sparse-vectors/), thereby enabling hybrid search capabilities. Sparse vectors are numerical representations of data where most of the elements are zero. Developers can combine search results from dense and sparse vectors, where sparse vectors ensure that results containing the specific keywords are returned and dense vectors identify semantically similar results.
- **Built-In Vector Quantization:** Qdrant offers three different [quantization](https://qdrant.tech/documentation/guides/quantization/) options to developers to optimize resource usage. Scalar quantization balances accuracy, speed, and compression by converting 32-bit floats to 8-bit integers. Binary quantization, the fastest method, significantly reduces memory usage. Product quantization offers the highest compression, and is perfect for memory-constrained scenarios.
- **Flexible Deployment Options:** Qdrant offers a range of deployment options. Developers can easily set up Qdrant (or Qdrant cluster) [locally](https://qdrant.tech/documentation/quick-start/#download-and-run) using Docker for free. [Qdrant Cloud](https://qdrant.tech/cloud/), on the other hand, is a scalable, managed solution that provides easy access with flexible pricing. Additionally, Qdrant offers [Hybrid Cloud](https://qdrant.tech/hybrid-cloud/) which integrates Kubernetes clusters from cloud, on-premises, or edge, into an enterprise-grade managed service.
- **Security through API Keys, JWT and RBAC:** Qdrant offers developers various ways to [secure](https://qdrant.tech/documentation/guides/security/) their instances. For simple authentication, developers can use API keys (including Read Only API keys). For more granular access control, it offers JSON Web Tokens (JWT) and the ability to build Role-Based Access Control (RBAC). TLS can be enabled to secure connections. Qdrant is also [SOC 2 Type II](https://qdrant.tech/blog/qdrant-soc2-type2-audit/) certified.
- **Built-In Vector Quantization:** Qdrant offers three different [quantization](https://qdrant.tech/documentation/manage-data/quantization/) options to developers to optimize resource usage. Scalar quantization balances accuracy, speed, and compression by converting 32-bit floats to 8-bit integers. Binary quantization, the fastest method, significantly reduces memory usage. Product quantization offers the highest compression, and is perfect for memory-constrained scenarios.
- **Flexible Deployment Options:** Qdrant offers a range of deployment options. Developers can easily set up Qdrant (or Qdrant cluster) [locally](https://qdrant.tech/documentation/quickstart/#download-and-run) using Docker for free. [Qdrant Cloud](https://qdrant.tech/cloud/), on the other hand, is a scalable, managed solution that provides easy access with flexible pricing. Additionally, Qdrant offers [Hybrid Cloud](https://qdrant.tech/hybrid-cloud/) which integrates Kubernetes clusters from cloud, on-premises, or edge, into an enterprise-grade managed service.
- **Security through API Keys, JWT and RBAC:** Qdrant offers developers various ways to [secure](https://qdrant.tech/documentation/operations/security/) their instances. For simple authentication, developers can use API keys (including Read Only API keys). For more granular access control, it offers JSON Web Tokens (JWT) and the ability to build Role-Based Access Control (RBAC). TLS can be enabled to secure connections. Qdrant is also [SOC 2 Type II](https://qdrant.tech/blog/qdrant-soc2-type2-audit/) certified.
Additionally, Qdrant integrates seamlessly with popular machine learning frameworks such as [LangChain](https://qdrant.tech/blog/using-qdrant-and-langchain/), LlamaIndex, and Haystack; and Qdrant Hybrid Cloud integrates seamlessly with AWS, DigitalOcean, Google Cloud, Linode, Oracle Cloud, OpenShift, and Azure, among others.
@@ -167,7 +167,7 @@ References:
- [Pinecone Documentation](https://docs.pinecone.io/)
- [Qdrant Documentation](https://qdrant.tech/documentation/)
- If you aren't ready yet, [try out Qdrant locally](/documentation/quick-start/) or sign up for [Qdrant Cloud](https://cloud.qdrant.io/signup).
- If you aren't ready yet, [try out Qdrant locally](/documentation/quickstart/) or sign up for [Qdrant Cloud](https://cloud.qdrant.io/signup).
- For more basic information on Qdrant read our [Overview](/documentation/overview/) section or learn more about Qdrant Cloud's [Free Tier](/documentation/cloud/).
@@ -45,13 +45,13 @@ guide](/documentation/cloud/authentication/#test-cluster-access).
If your Qdrant deployment is local, you do not need an API key.
Your next step depends on how you installed Qdrant. For details, read the
[Qdrant Installation](/documentation/guides/installation/)
[Qdrant Installation](/documentation/operations/installation/)
guide.
#### If you use the Qdrant container or binary
Upgrade your deployment. Run the commands in the applicable section of the
[Qdrant Installation](/documentation/guides/installation/)
[Qdrant Installation](/documentation/operations/installation/)
guide. The default commands automatically pull the latest version of Qdrant.
#### If you use the Qdrant helm chart
@@ -45,13 +45,13 @@ guide](https://qdrant.tech/documentation/cloud/quickstart-cloud/#step-2-test-clu
If your Qdrant deployment is local, you do not need an API key.
Your next step depends on how you installed Qdrant. For details, read the
[Qdrant Installation](https://qdrant.tech/documentation/guides/installation/)
[Qdrant Installation](https://qdrant.tech/documentation/operations/installation/)
guide.
#### If you use the Qdrant container or binary
Upgrade your deployment. Run the commands in the applicable section of the
[Qdrant Installation](https://qdrant.tech/documentation/guides/installation/)
[Qdrant Installation](https://qdrant.tech/documentation/operations/installation/)
guide. The default commands automatically pull the latest version of Qdrant.
#### If you use the Qdrant helm chart
@@ -53,7 +53,7 @@ In the podcast, we addressed the following:
- **Model evaluation(LLM)** - Understanding the model at the domain-level for the given use case, supporting required context length and terminology/concept understanding.
- **Ingestion pipeline evaluation** - Evaluating factors related to data ingestion and processing such as chunk strategies, chunk size, chunk overlap, and more.
- **Retrieval evaluation** - Understanding factors such as average precision, [Distributed cumulative gain](https://en.wikipedia.org/wiki/Discounted_cumulative_gain) (DCG), as well as normalized DCG.
- **Generation evaluation(E2E)** - Establishing guardrails. Evaulating prompts. Evaluating the number of chunks needed to set up the context for generation.
- **Generation evaluation(E2E)** - Establishing guardrails. Evaluating prompts. Evaluating the number of chunks needed to set up the context for generation.
### The recording
@@ -5,7 +5,6 @@ slug: decay-functions # Change this slug to your page slug if needed
short_description: Why's and how's of decay functions in Qdrant's relevance score boosting. # Change this
description: Understanding decay functions for relevance score boosting. # Change this
preview_image: /blog/decay-functions/preview/preview.jpg # Change this
social_preview_image: /blog/decay-functions/preview/social_preview.jpg # Optional image used for link previews
title_preview_image: /blog/decay-functions/preview/title.jpg # Optional image used for blog post title
date: 2025-09-01T14:55:45+02:00
author: Evgeniya Sukhodolskaya
@@ -17,7 +16,7 @@ tags:
---
A problem we've noticed while monitoring the [Qdrant Discord Community](https://discord.gg/d4MPnX3s) is that due to the extensive list of expressions that the [score boosting](https://qdrant.tech/documentation/concepts/search-relevance/#score-boosting) functionality provides, there's room for confusion on how it's supposed to be applied. And that might block you from moving the business logic behind relevance scoring into the Qdrant search engine. We don't want that!
A problem we've noticed while monitoring the [Qdrant Discord Community](https://discord.gg/d4MPnX3s) is that due to the extensive list of expressions that the [score boosting](https://qdrant.tech/documentation/search/search-relevance/#score-boosting) functionality provides, there's room for confusion on how it's supposed to be applied. And that might block you from moving the business logic behind relevance scoring into the Qdrant search engine. We don't want that!
In this blog, we'd like to de-spooky-fy the **decay functions** part of the score boosting, or, more precisely: `LinDecayExpression`, `ExpDecayExpression`, and `GaussDecayExpression` -- frequent guests on the Discord *#ask-for-help* channel.
@@ -144,13 +143,13 @@ Anything longer than 9 minutes or shorter than 1 minute quickly becomes less rel
**Explanation:**
Out of all promo codes for different products/events, users will strongly prefer ones uploaded *just now*, as they’re most likely to work. But that relevance drops quickly over time: within a week, it reaches a midpoint of 0.1. After that, if a promo code is still active, it’s a gamble anyway: might work, might not. So old-but-not-expired codes are roughly equally irrelevant.
**Note #5.** For Qdrant [datetime](https://qdrant.tech/documentation/concepts/payload/#datetime) payloads, `scale` should always be provided in seconds!
**Note #5.** For Qdrant [datetime](https://qdrant.tech/documentation/manage-data/payload/#datetime) payloads, `scale` should always be provided in seconds!
### I Don't Know All the Parameters in Advance
As you can see, using decay functions in Qdrant's score boosting means you'll have to know the parameters in advance.
What we've seen in our Discord Community quite a few times is that people try to apply decay functions to normalize similarity scores from [prefetches](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries), usually as a way to fuse results from different types of similarity searches.
What we've seen in our Discord Community quite a few times is that people try to apply decay functions to normalize similarity scores from [prefetches](https://qdrant.tech/documentation/search/hybrid-queries/#multi-stage-queries), usually as a way to fuse results from different types of similarity searches.
The common question is:
@@ -170,10 +169,10 @@ But here's the problem: That 36 might not be a "high" score at all. Maybe your d
Now let's see how using decay functions looks in Qdrant.
We'll provide HTTP request examples, but you can use decay functions [analogously in the Python, TypeScript, Rust, Java, C#, and Go clients](/documentation/concepts/search-relevance/#time-based-score-boosting).
We'll provide HTTP request examples, but you can use decay functions [analogously in the Python, TypeScript, Rust, Java, C#, and Go clients](/documentation/search/search-relevance/#time-based-score-boosting).
**Note #6.**
Payload variables used within the formula benefit from having [payload indexes](https://qdrant.tech/documentation/concepts/indexing/#payload-index). So, we require you to set up a payload index for any variable used in a formula.
Payload variables used within the formula benefit from having [payload indexes](https://qdrant.tech/documentation/manage-data/indexing/#payload-index). So, we require you to set up a payload index for any variable used in a formula.
Let's take our "educational videos in the German language" example and see how it takes shape in Qdrant:
@@ -248,7 +247,7 @@ We truly hope this write-up helped untangle things a bit. Now the only thing lef
Use the snippets in the article as a starting point and experiment with the relevance score boosting in [Qdrant Cloud](https://qdrant.tech/). We offer a free-forever 1GB cluster: enough to test, tweak, and see how the decay functions behave on your data.
And if you feel like diving deeper into decay functions or score boosting in general, check out our [documentation](/documentation/concepts/search-relevance/#score-boosting), which includes a decay-on-distance example and plenty more to learn from.
And if you feel like diving deeper into decay functions or score boosting in general, check out our [documentation](/documentation/search/search-relevance/#score-boosting), which includes a decay-on-distance example and plenty more to learn from.
### Tell Us What You're Building
@@ -44,7 +44,7 @@ ___
## Architecture
**Search Engine & DB:** [**Qdrant**](https://qdrant.tech) stands out as a high-performance [**vector database**](/qdrant-vector-database/) built in Rust, known for its reliability and speed. Its advanced features, such as [**vector visualization**](/documentation/web-ui/) and efficient [**querying**](/documentation/concepts/search/), make it a go-to choice for developers working on embedding-based projects.
**Search Engine & DB:** [**Qdrant**](https://qdrant.tech) stands out as a high-performance [**vector database**](/qdrant-vector-database/) built in Rust, known for its reliability and speed. Its advanced features, such as [**vector visualization**](/documentation/web-ui/) and efficient [**querying**](/documentation/search/search/), make it a go-to choice for developers working on embedding-based projects.
![architecture](/blog/facial-recognition/architecture.png)
@@ -121,7 +121,7 @@ If your data is properly embedded, then the visualization tool will appropriatel
## Lessons and Takeaways
Scalability poses challenges when working with large datasets, such as 20,000+ images. Consider optimizations like [**quantization**](/documentation/guides/quantization/) to reduce memory usage or precomputing average embeddings for clusters can significantly minimize storage and computational costs. These strategies ensure the system remains performant as the dataset grows.
Scalability poses challenges when working with large datasets, such as 20,000+ images. Consider optimizations like [**quantization**](/documentation/manage-data/quantization/) to reduce memory usage or precomputing average embeddings for clusters can significantly minimize storage and computational costs. These strategies ensure the system remains performant as the dataset grows.
The potential real-world applications of this technology extend far beyond entertainment. Similar systems can be used in security applications for embedding-based facial recognition to secure access to buildings or devices.
@@ -43,7 +43,7 @@ We put together an end-to-end tutorial to show you how to build a GenAI applicat
Learn how to set up a private AI service that addresses customer support issues with high accuracy and effectiveness. By leveraging Airbyte’s data pipelines with Qdrant Hybrid Cloud, you will create a customer support system that is always synchronized with up-to-date knowledge.
[Try the Tutorial](/documentation/tutorials/rag-customer-support-cohere-airbyte-aws/)
[Try the Tutorial](/documentation/examples/rag-customer-support-cohere-airbyte-aws/)
#### Documentation: Deploy Qdrant in a Few Clicks
@@ -39,7 +39,7 @@ We put together an end-to-end tutorial to show you how to build a GenAI applicat
Learn how to set up a private AI service that addresses customer support issues with high accuracy and effectiveness. By leveraging Cohere’s models with Qdrant Hybrid Cloud, you will create a fully private customer support system.
[Try the Tutorial](/documentation/tutorials/rag-customer-support-cohere-airbyte-aws/)
[Try the Tutorial](/documentation/examples/rag-customer-support-cohere-airbyte-aws/)
#### Documentation: Deploy Qdrant in a Few Clicks
@@ -47,7 +47,7 @@ To get Qdrant Hybrid Cloud setup on DigitalOcean, just follow these steps:
We created a tutorial that guides you through setting up and leveraging Qdrant Hybrid Cloud on DigitalOcean for a RAG application. It highlights practical steps to integrate vector search with Jina AI's LLMs, optimizing the generation of high-quality, relevant AI content, while ensuring data sovereignty is maintained throughout. This specific system is tied together via the LlamaIndex framework.
[Try the Tutorial](/documentation/tutorials/hybrid-search-llamaindex-jinaai/)
[Try the Tutorial](/documentation/examples/hybrid-search-llamaindex-jinaai/)
For a comprehensive guide, our documentation provides detailed instructions on setting up Qdrant on DigitalOcean.
@@ -41,7 +41,7 @@ To get you started, we created a comprehensive tutorial that shows how to build
Learn how to develop a tutor chatbot from online course materials. You will create a Retrieval Augmented Generation (RAG) pipeline with Haystack for enhanced generative AI capabilities and Qdrant Hybrid Cloud for vector search. By deploying every tool on RedHat OpenShift, you will ensure complete privacy and data sovereignty, whereby no course content leaves your cloud.
[Try the Tutorial](/documentation/tutorials/rag-chatbot-red-hat-openshift-haystack/)
[Try the Tutorial](/documentation/examples/rag-chatbot-red-hat-openshift-haystack/)
#### Documentation: Deploy Qdrant in a Few Clicks
@@ -41,7 +41,7 @@ To get you started, we created a comprehensive tutorial that shows how to build
Learn how to build an app that retrieves information from PDF user manuals to enhance user experience for companies that sell household appliances. The system will leverage Jina AI embeddings and Qdrant Hybrid Cloud for enhanced generative AI capabilities, while the RAG pipeline will be tied together using the LlamaIndex framework. This example demonstrates how complex tables in PDF documentation can be processed as high quality embeddings with no extra configuration. By introducing Hybrid Search from Qdrant, the RAG functionality is highly accurate.
[Try the Tutorial](/documentation/tutorials/hybrid-search-llamaindex-jinaai/)
[Try the Tutorial](/documentation/examples/hybrid-search-llamaindex-jinaai/)
#### Documentation: Deploy Qdrant in a Few Clicks
@@ -41,7 +41,7 @@ To get you started, we’ve put together a tutorial that shows how to create nex
We created a comprehensive tutorial to show how you can build a RAG-based system with Qdrant Hybrid Cloud, LangChain and Cohere’s embeddings. This use case is focused on building a question-answering system for internal corporate employee onboarding.
[Try the Tutorial](/documentation/tutorials/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/)
[Try the Tutorial](/documentation/examples/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/)
#### Documentation: Deploy Qdrant in a Few Clicks
@@ -34,49 +34,49 @@ Together with our launch partners, we created in-depth tutorials and use cases f
> This tutorial shows how to build a private AI customer support system using Cohere's AI models on AWS, Airbyte, and Qdrant Hybrid Cloud for efficient and secure query automation.
[View Tutorial](/documentation/tutorials/rag-customer-support-cohere-airbyte-aws/)
[View Tutorial](/documentation/examples/rag-customer-support-cohere-airbyte-aws/)
**RAG System for Employee Onboarding** with Qdrant Hybrid Cloud, Oracle Cloud Infrastructure (OCI), Cohere, and LangChain
> This tutorial demonstrates how to use Oracle Cloud Infrastructure (OCI) for a secure setup that integrates Cohere's language models with Qdrant Hybrid Cloud, using LangChain to orchestrate natural language search for corporate documents, enhancing resource discovery and onboarding.
[View Tutorial](/documentation/tutorials/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/)
[View Tutorial](/documentation/examples/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/)
**Hybrid Search for Product PDF Manuals** with Qdrant Hybrid Cloud, LlamaIndex, and JinaAI
> Create a RAG-based chatbot that enhances customer support by parsing product PDF manuals using Qdrant Hybrid Cloud, LlamaIndex, and JinaAI, with DigitalOcean as the cloud host. This tutorial will guide you through the setup and integration process, enabling your system to deliver precise, context-aware responses for household appliance inquiries.
[View Tutorial](/documentation/tutorials/hybrid-search-llamaindex-jinaai/)
[View Tutorial](/documentation/examples/hybrid-search-llamaindex-jinaai/)
**Region-Specific RAG System for Contract Management** with Qdrant Hybrid Cloud, Aleph Alpha, and STACKIT
> Learn how to streamline contract management with a RAG-based system in this tutorial, which utilizes Aleph Alpha’s embeddings and a region-specific cloud setup. Hosted on STACKIT with Qdrant Hybrid Cloud, this solution ensures secure, GDPR-compliant storage and processing of data, ideal for businesses with intensive contractual needs.
[View Tutorial](/documentation/tutorials/rag-contract-management-stackit-aleph-alpha/)
[View Tutorial](/documentation/examples/rag-contract-management-stackit-aleph-alpha/)
**Movie Recommendation System** with Qdrant Hybrid Cloud and OVHcloud
> Discover how to build a recommendation system with our guide on collaborative filtering, using sparse vectors and the Movielens dataset.
[View Tutorial](/documentation/tutorials/recommendation-system-ovhcloud/)
[View Tutorial](/documentation/examples/recommendation-system-ovhcloud/)
**Private RAG Information Extraction Engine** with Qdrant Hybrid Cloud and Vultr using DSPy and Ollama
> This tutorial teaches you how to handle and structure private documents with large unstructured data. Learn to use DSPy for information extraction, run your LLM with Ollama on Vultr, and manage data with Qdrant Hybrid Cloud on Vultr, perfect for regulated environments needing data privacy.
[View Tutorial](/documentation/tutorials/rag-chatbot-vultr-dspy-ollama/)
[View Tutorial](/documentation/examples/rag-chatbot-vultr-dspy-ollama/)
**RAG System That Chats with Blog Contents** with Qdrant Hybrid Cloud and Scaleway using LangChain.
> Build a RAG system that combines blog scanning with the capabilities of semantic search. RAG enhances the generation of answers by retrieving relevant documents to aid the question-answering process. This setup showcases the integration of advanced search and AI language processing to improve information retrieval and generation tasks.
[View Tutorial](/documentation/tutorials/rag-chatbot-scaleway/)
[View Tutorial](/documentation/examples/rag-chatbot-scaleway/)
**Private Chatbot for Interactive Learning** with Qdrant Hybrid Cloud and Red Hat OpenShift using Haystack.
> In this tutorial, you will build a chatbot without public internet access. The goal is to keep sensitive data secure and isolated. Your RAG system will be built with Qdrant Hybrid Cloud on Red Hat OpenShift, leveraging Haystack for enhanced generative AI capabilities. This tutorial especially explores how this setup ensures that not a single data point leaves the environment.
[View Tutorial](/documentation/tutorials/rag-chatbot-red-hat-openshift-haystack/)
[View Tutorial](/documentation/examples/rag-chatbot-red-hat-openshift-haystack/)
#### Supporting Documentation
@@ -41,7 +41,7 @@ To get you started, we created a comprehensive tutorial that shows how to build
Use this end-to-end tutorial to create a system that retrieves information from complex user manuals in PDF format to enhance user experience for companies that sell household appliances. You will build a RAG pipeline with LlamaIndex leveraging Qdrant Hybrid Cloud for enhanced generative AI capabilities. The LlamaIndex integration shows how complex tables inside of items’ PDF documents can be processed via hybrid vector search with no additional configuration.
[Try the Tutorial](/documentation/tutorials/hybrid-search-llamaindex-jinaai/)
[Try the Tutorial](/documentation/examples/hybrid-search-llamaindex-jinaai/)
#### Documentation: Deploy Qdrant in a Few Clicks
@@ -37,7 +37,7 @@ Deploying Qdrant Hybrid Cloud on OCI facilitates vector search in production env
We created a comprehensive tutorial to show how to leverage the benefits of Qdrant Hybrid Cloud on OCI and build AI applications with a focus on data sovereignty. This use case is focused on building a RAG system for FAQ, leveraging the strengths of Qdrant Hybrid Cloud for vector search, Oracle Cloud Infrastructure (OCI) as a managed Kubernetes provider, Cohere models for embedding, and LangChain as a framework.
[Try the Tutorial](/documentation/tutorials/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/)
[Try the Tutorial](/documentation/examples/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/)
Deploying Qdrant Hybrid Cloud on Oracle Cloud Infrastructure only takes a few minutes due to the seamless Kubernetes-native integration. You can get started by following these three steps:
@@ -37,7 +37,7 @@ Through the seamless integration between Qdrant Hybrid Cloud and OVHcloud, devel
To show how Qdrant Hybrid Cloud deployed on OVHcloud allows developers to leverage the benefits of an AI use case that is completely run within the existing infrastructure, we put together a comprehensive use case tutorial. This tutorial guides you through creating a recommendation system using collaborative filtering and sparse vectors with Qdrant Hybrid Cloud on OVHcloud. It employs the Movielens dataset for practical application, providing insights into building efficient, scalable recommendation engines suitable for developers and data scientists looking to leverage advanced vector search technologies within a secure, GDPR-compliant European cloud infrastructure.
[Try the Tutorial](/documentation/tutorials/recommendation-system-ovhcloud/)
[Try the Tutorial](/documentation/examples/recommendation-system-ovhcloud/)
#### Get Started Today and Leverage the Benefits of Qdrant Hybrid Cloud
@@ -52,7 +52,7 @@ To get started, we created a comprehensive tutorial that shows how to build next
In this tutorial, you will build a chatbot without public internet access. The goal is to keep sensitive data secure and isolated. Your RAG system will be built with Qdrant Hybrid Cloud on Red Hat OpenShift, leveraging Haystack for enhanced generative AI capabilities. This tutorial especially explores how this setup ensures that not a single data point leaves the environment.
[Try the Tutorial](/documentation/tutorials/rag-chatbot-red-hat-openshift-haystack/)
[Try the Tutorial](/documentation/examples/rag-chatbot-red-hat-openshift-haystack/)
#### Documentation: Deploy Qdrant in a Few Clicks
@@ -29,7 +29,7 @@ RAG applications often rely on sensitive or proprietary internal data, emphasizi
We created a tutorial that guides you through setting up and leveraging Qdrant Hybrid Cloud on Scaleway for a RAG application, providing insights into efficiently managing data within a secure, sovereign framework. It highlights practical steps to integrate vector search with LLMs, optimizing the generation of high-quality, relevant AI content, while ensuring data sovereignty is maintained throughout.
[Try the Tutorial](/documentation/tutorials/rag-chatbot-scaleway/)
[Try the Tutorial](/documentation/examples/rag-chatbot-scaleway/)
#### The Benefits of Running Qdrant Hybrid Cloud on Scaleway
@@ -33,7 +33,7 @@ Qdrant Hybrid Cloud is the first managed vector database that can be deployed in
To demonstrate the power of Qdrant Hybrid Cloud on STACKIT, we’ve developed a comprehensive tutorial showcasing how to build secure, AI-driven applications focusing on data sovereignty. This tutorial specifically shows how to build a contract management platform that enables users to upload documents (PDF or DOCx), which are then segmented for searchable access. Designed with multitenancy, users can only access their team or organization's documents. It also features custom sharding for location-specific document storage. Beyond search, the application offers rephrasing of document excerpts for clarity to those without context.
[Try the Tutorial](/documentation/tutorials/rag-contract-management-stackit-aleph-alpha/)
[Try the Tutorial](/documentation/examples/rag-contract-management-stackit-aleph-alpha/)
#### Start Using Qdrant with STACKIT
@@ -45,7 +45,7 @@ We've compiled an in-depth guide for leveraging Qdrant Hybrid Cloud on Vultr to
This tutorial outlines creating a personalized AI assistant using Qdrant Hybrid Cloud on Vultr, incorporating advanced vector search to power dynamic, interactive experiences. We will develop a RAG pipeline powered by DSPy and detail how to maintain data privacy within your Vultr environment.
[Try the Tutorial](/documentation/tutorials/rag-chatbot-vultr-dspy-ollama/)
[Try the Tutorial](/documentation/examples/rag-chatbot-vultr-dspy-ollama/)
#### Documentation: Effortless Deployment with Qdrant
@@ -85,7 +85,7 @@ Leveraging Late-Interaction Models for Rich Documents
Traditional OCR pipelines can add complexity and create accuracy challenges. But late-interaction models simplify the ingestion pipeline by running at the reranking stage.
Models like ([ColPali](https://qdrant.tech/blog/qdrant-colpali/) and ColQwen) bypass traditional OCR pipelines, directly processing images of complex documents. They enhance accuracy by maintaining original layouts and contextual integrity, simplifying your retrieval pipelines. The tradeoff is a heavier application, but these challenges can be addressed with further [optimization](https://qdrant.tech/documentation/guides/optimize/)*.*
Models like ([ColPali](https://qdrant.tech/blog/qdrant-colpali/) and ColQwen) bypass traditional OCR pipelines, directly processing images of complex documents. They enhance accuracy by maintaining original layouts and contextual integrity, simplifying your retrieval pipelines. The tradeoff is a heavier application, but these challenges can be addressed with further [optimization](https://qdrant.tech/documentation/operations/optimize/)*.*
#### Enabling highly granular accuracy for complex legal searches
@@ -125,7 +125,7 @@ final_results = reranked[:5]
Not every clause is created equal. Legal professionals often care more about specific provisions, jurisdictions, or case types, for example.
Qdrant's [Score Boosting Reranker](/documentation/concepts/search-relevance/#score-boosting) lets you integrate domain-specific logic (e.g., jurisdiction or recent cases) directly into search rankings, ensuring results align precisely with legal business rules.
Qdrant's [Score Boosting Reranker](/documentation/search/search-relevance/#score-boosting) lets you integrate domain-specific logic (e.g., jurisdiction or recent cases) directly into search rankings, ensuring results align precisely with legal business rules.
```json
POST /collections/legal-docs/points/query
@@ -163,13 +163,13 @@ Legal datasets are growing, and so are the compute bills. From GPU acceleration
* [GPU indexing](https://qdrant.tech/blog/qdrant-1.13.x/) accelerates indexing by up to 10x compared to CPU methods, offering vendor-agnostic compatibility with modern GPUs via Vulkan API.
* [Vector quantization](https://qdrant.tech/documentation/guides/quantization/) compresses embeddings, significantly reducing memory and operational costs. It results in lower accuracy, so carefully consider this option. For example, [LawMe](http://qdrant.tech/blog/case-study-lawme), a Qdrant user, uses Binary Quantization to cost-effectively add more data for its AI Legal Assistants.
* [Vector quantization](https://qdrant.tech/documentation/manage-data/quantization/) compresses embeddings, significantly reducing memory and operational costs. It results in lower accuracy, so carefully consider this option. For example, [LawMe](http://qdrant.tech/blog/case-study-lawme), a Qdrant user, uses Binary Quantization to cost-effectively add more data for its AI Legal Assistants.
### Getting Started: Choosing Your Search Infrastructure
#### Deploy in private, cloud, or hybrid environments without sacrificing control
No matter your stage—prototype or production—your stack will have to meet both engineering and compliance needs. Qdrant supports flexible deployment strategies, including [managed cloud](https://qdrant.tech/cloud/) and [hybrid cloud](https://qdrant.tech/hybrid-cloud/), along with open-source solutions via [Docker](https://qdrant.tech/documentation/quick-start/), enabling easy scaling and secure management of legal data.
No matter your stage—prototype or production—your stack will have to meet both engineering and compliance needs. Qdrant supports flexible deployment strategies, including [managed cloud](https://qdrant.tech/cloud/) and [hybrid cloud](https://qdrant.tech/hybrid-cloud/), along with open-source solutions via [Docker](https://qdrant.tech/documentation/quickstart/), enabling easy scaling and secure management of legal data.
#### Build and iterate quickly with responsive support and built-in tooling
@@ -0,0 +1,66 @@
---
title: "Master Multi-Vector Search With Qdrant"
draft: false
slug: multi-vector-search-course
short_description: "Go beyond single-vector embeddings. Our new advanced course covers ColBERT, ColPali, MaxSim, and production-grade multi-vector pipelines in Qdrant."
description: "Go beyond single-vector embeddings. Our new advanced course covers ColBERT, ColPali, MaxSim, and production-grade multi-vector pipelines in Qdrant."
preview_image: /blog/multi-vector-course-release/hero.png
social_preview_image: /blog/multi-vector-course-release/hero.png
date: 2026-03-24
author: Neil Kanungo
featured: true
tags:
- qdrant-course
- multi-vector-search
- colbert
- colpali
- certification
---
Most vector search tutorials stop at single-vector embeddings: one document, one vector, one similarity score. That works for demos. It falls apart when your retrieval pipeline needs to capture fine-grained token-level interactions across text, images, and PDFs at production scale.
Until now, engineers who wanted to go deeper had to piece together scattered papers, blog posts, and half-documented repos. There was no structured, hands-on resource that connected the theory of late interaction models to real implementation in a production search engine.
We built one.
## Introducing the Multi-Vector Search Course
[Qdrant's Multi-Vector Search Course](https://qdrant.tech/course/multi-vector-search/) is a free, advanced course created by [Kacper Łukawski](https://www.linkedin.com/in/kacperlukawski/). Kacper designed this course to fill a real gap in the developer community: practical, production-focused education on multi-vector retrieval that goes well beyond "here's how embeddings work."
This is an advanced course! It's built for ML engineers, backend engineers, and search engineers who already understand vector search fundamentals and want to master what comes next.
## What You'll Learn
The course is organized into four modules, each taking roughly one to two hours:
**Module 0: Setup.** Configure a Qdrant Cloud or local instance and install Python dependencies.
**Module 1: Text Multi-Vectors.** Understand the late interaction paradigm, learn the MaxSim distance metric, explore real use cases and challenges, and implement ColBERT-based search with Qdrant.
**Module 2: Multi-Modal Search.** Apply multi-vector representations to images and PDFs using ColPali. Explore model variants in the ColPali family and use visual interpretability for debugging retrieval results.
**Module 3: Optimization and Evaluation.** Master vector quantization, pooling techniques, and MUVERA indexing for memory-efficient search at billion scale. Build multi-stage retrieval pipelines with Qdrant's Universal Query API and evaluate with industry-standard metrics (Recall@k, NDCG, MRR).
The course wraps up with a final project: build your own production-ready multi-modal search system from scratch.
## How It Works
Every module combines video lessons from the Qdrant team with hands-on Google Colab notebooks. You watch, you build, you evaluate. The progressive structure means each module builds directly on the previous one, so you finish with a complete, working pipeline rather than a collection of disconnected concepts.
## Earn a Qdrant Certification
Complete the course and pass the certification exam to earn a shareable Qdrant Multi-Vector Search certificate. It's a concrete way to demonstrate that you can design, implement, and optimize multi-vector retrieval pipelines in production.
Certifications are available through [Qdrant Academy](https://qdrant.tech/course/multi-vector-search/certification/).
## Free Swag for the First 20 Certified
Dive in now: the **first 20 people** who complete their Multi-Vector Search certification and post on LinkedIn with the hashtag **#QdrantCertified** will receive free Qdrant swag. Share your certificate, tag us, and we'll reach out.
## Start Now
The course is free, self-paced, and available today. If you've been looking for a structured path from "I understand embeddings" to "I can build and evaluate multi-vector retrieval at scale," this is it.
[Start the Multi-Vector Search Course](https://qdrant.tech/course/multi-vector-search/)
Questions? Join our [Discord community](https://discord.gg/qdrant) where Qdrant experts collaborate.
@@ -40,19 +40,19 @@ Here are the six conditions:
**1. Your vector dataset is under ~1M vectors.** The community's empirical ceiling is around 10M, but the comfortable range is much lower. Above 1M you'll start hitting index-build times, memory pressure, and recall degradation under load.
This threshold is also easier to hit than you'd expect — especially if you're working with **multivectors**. Techniques like ColBERT-style late interaction generate one embedding *per token* rather than one per document, so a corpus of 100K documents can easily balloon into tens of millions of vectors overnight. Qdrant, by contrast, has [native multivector support](/documentation/concepts/vectors/#multivectors) with dedicated documentation and query APIs built around it.
This threshold is also easier to hit than you'd expect — especially if you're working with **multivectors**. Techniques like ColBERT-style late interaction generate one embedding *per token* rather than one per document, so a corpus of 100K documents can easily balloon into tens of millions of vectors overnight. Qdrant, by contrast, has [native multivector support](/documentation/manage-data/vectors/#multivectors) with dedicated documentation and query APIs built around it.
**2. You don't need accurate metadata filtering.** If every search is against the full collection, post-filtering won't limit you. But the moment you need to scope searches to a user, tenant, category, or any selective predicate, pgvector generates unnecessary search overhead.
> *"I think the most relevant weakness for pgvector is the lack of 'proper' prefiltering on metadata while leveraging the vector index."*
Qdrant takes an entirely different approach to filtering. Specifically, Qdrant utilizes a [filterable HNSW](/documentation/concepts/indexing/#filtrable-index) which lets you traverse the nearest-neighbor graph while maintaining metadata filters.
Qdrant takes an entirely different approach to filtering. Specifically, Qdrant utilizes a [filterable HNSW](/documentation/manage-data/indexing/#filtrable-index) which lets you traverse the nearest-neighbor graph while maintaining metadata filters.
**3. Your embeddings are tightly coupled to relational data.** If vectors are just an attribute of a row (e.g., a product description embedding alongside the product), colocation helps. If vectors are first-class entities, the argument for co-location weakens.
**4. You don't need hybrid search.** While pgvector supports dense vector similarity search via HNSW, the Postgres extension ecosystem still lacks a high-quality BM25 implementation — a critical component for hybrid search.
Postgres *does* have full-text search via `tsvector`/`tsquery`, and it's excellent for what it does. But that's lexical search — exact term matches, stemming, and stop words. BM25 is a probabilistic model that considers term frequency, inverse document frequency, and document length. Qdrant supports [native BM25 via sparse vectors](/documentation/concepts/vectors/#sparse-vectors). They're not the same thing.
Postgres *does* have full-text search via `tsvector`/`tsquery`, and it's excellent for what it does. But that's lexical search — exact term matches, stemming, and stop words. BM25 is a probabilistic model that considers term frequency, inverse document frequency, and document length. Qdrant supports [native BM25 via sparse vectors](/documentation/manage-data/vectors/#sparse-vectors). They're not the same thing.
**5. Postgres is already doing the heavy lifting for your business logic.** You have existing transactions, schemas, and ACID guarantees that matter. Adding a second data store splits that concern. If Postgres is truly central, the operational argument for colocation is real.
@@ -75,7 +75,7 @@ That's three conditions gone before you've even thought about scale. People are
When the conditions above don't all hold, dedicated vector stores offer concrete advantages:
- **Efficient metadata filtering** — pre-filter on metadata fields before computing similarity, avoiding wasted work on irrelevant vectors
- **Native hybrid search** — combine dense similarity and BM25 keyword matching in a single query with [reciprocal rank fusion](/documentation/concepts/hybrid-queries/)
- **Native hybrid search** — combine dense similarity and BM25 keyword matching in a single query with [reciprocal rank fusion](/documentation/search/hybrid-queries/)
- **Scale beyond 10M vectors** — purpose-built sharding, distributed indexing, and memory management
- **Decoupled architecture** — scale, optimize, and evolve your search layer independently of your relational database
@@ -59,7 +59,7 @@ When looking at the overview of your cluster, we’ve added new tabs with an imp
* **Logs**: get a real-time window into what’s happening inside cluste for transparency, diagnostics, and control (especially important during debugging, performance tuning, or infrastructure troubleshooting\!)
* **Backups:** View snapshots of your vector data and metadata that can be used to restore your collections in case of data loss, migration, or rollbacks (not available on free clusters)
* **Configuration**: Check your collection defaults and add advanced optimizations (after reading Docs of course)
* For example, we advise against setting up a ton of different collections. Instead segment with [payloads](https://qdrant.tech/documentation/concepts/payload/).
* For example, we advise against setting up a ton of different collections. Instead segment with [payloads](https://qdrant.tech/documentation/manage-data/payload/).
When viewing the details of your clusters, you can now view the Cluster UI Dashboard regardless of where you are, and also have easier access to tutorials and resources.
@@ -74,7 +74,7 @@ Next we’ve done a major overhaul to the “Get Started” page. Our goal is to
![Image of Get Started webpage](/blog/product-ui-changes/get-started-overview.jpg)
**Explore Your Data or Start with Samples**
You’ll see immediately pertinent information to help you get the most out of Qdrant quickly, including the [Cloud Quickstart guide](https://qdrant.tech/documentation/quickstart-cloud/), and resources to help you get your data into Qdrant, or use sample data.
You’ll see immediately pertinent information to help you get the most out of Qdrant quickly, including the [Cloud Quickstart guide](https://qdrant.tech/documentation/cloud-quickstart/), and resources to help you get your data into Qdrant, or use sample data.
Learn about the different ways to connect to your cluster, use the Qdrant API, try out sample data, and our personal favorite, use the Qdrant Cluster UI to view your collection data and access tutorials.
+13 -13
View File
@@ -31,14 +31,14 @@ You can now configure the Query API request with the following parameters:
|Parameter|Description|
|-|-|
|no parameter|Returns points by `id`|
|`nearest`|Queries nearest neighbors ([Search](/documentation/concepts/search/))|
|`fusion`|Fuses sparse/dense prefetch queries ([Hybrid Search](/documentation/concepts/hybrid-queries/#hybrid-search))|
|`discover`|Queries `target` with added `context` ([Discovery](/documentation/concepts/explore/#discovery-api))|
|`context` |No target with `context` only ([Context](/documentation/concepts/explore/#context-search))|
|`recommend`|Queries against `positive`/`negative` examples. ([Recommendation](/documentation/concepts/explore/#recommendation-api))|
|`order_by`|Orders results by [payload field](/documentation/concepts/hybrid-queries/#re-ranking-with-payload-values)|
|`nearest`|Queries nearest neighbors ([Search](/documentation/search/search/))|
|`fusion`|Fuses sparse/dense prefetch queries ([Hybrid Search](/documentation/search/hybrid-queries/#hybrid-search))|
|`discover`|Queries `target` with added `context` ([Discovery](/documentation/search/explore/#discovery-api))|
|`context` |No target with `context` only ([Context](/documentation/search/explore/#context-search))|
|`recommend`|Queries against `positive`/`negative` examples. ([Recommendation](/documentation/search/explore/#recommendation-api))|
|`order_by`|Orders results by [payload field](/documentation/search/hybrid-queries/#re-ranking-with-payload-values)|
For example, you can configure Query API to run [Discovery search](/documentation/concepts/explore/#discovery-api). Let's see how that looks:
For example, you can configure Query API to run [Discovery search](/documentation/search/explore/#discovery-api). Let's see how that looks:
```http
POST collections/{collection_name}/points/query
@@ -57,7 +57,7 @@ POST collections/{collection_name}/points/query
}
```
We will be publishing code samples in [docs](/documentation/concepts/hybrid-queries/) and our new [API specification](http://api.qdrant.tech).</br> *If you need additional support with this new method, our [Discord](https://qdrant.to/discord) on-call engineers can help you.*
We will be publishing code samples in [docs](/documentation/search/hybrid-queries/) and our new [API specification](http://api.qdrant.tech).</br> *If you need additional support with this new method, our [Discord](https://qdrant.to/discord) on-call engineers can help you.*
### Native Hybrid Search Support
@@ -221,7 +221,7 @@ await client.QueryAsync(
Query API can now pre-fetch vectors for requests, which means you can run queries sequentially within the same API call. There are a lot of options here, so you will need to define a strategy to merge these requests using new parameters. For example, you can now include **rescoring within Hybrid Search**, which can open the door to strategies like iterative refinement via matryoshka embeddings.
*To learn more about this, read the [Query API documentation](/documentation/concepts/search/#query-api).*
*To learn more about this, read the [Query API documentation](/documentation/search/search/#query-api).*
## Inverse Document Frequency [IDF]
@@ -535,11 +535,11 @@ await client.QueryAsync(
```
**Note:** *The multivector feature is not only useful for ColBERT; it can also be used in other ways.*</br>
For instance, in e-commerce, you can use multi-vector to store multiple images of the same item. This serves as an alternative to the [group-by](/documentation/concepts/search/#grouping-api) method.
For instance, in e-commerce, you can use multi-vector to store multiple images of the same item. This serves as an alternative to the [group-by](/documentation/search/search/#grouping-api) method.
## Sparse Vectors Compression
In version 1.9, we introduced the `uint8` [vector datatype](/documentation/concepts/vectors/#datatypes) for sparse vectors, in order to support pre-quantized embeddings from companies like JinaAI and Cohere.
In version 1.9, we introduced the `uint8` [vector datatype](/documentation/manage-data/vectors/#datatypes) for sparse vectors, in order to support pre-quantized embeddings from companies like JinaAI and Cohere.
This time, we are introducing a new datatype **for both sparse and dense vectors**, as well as a different way of **storing** these vectors.
**Datatype:** Sparse and dense vectors were previously represented in larger `float32` values, but now they can be turned to the `float16`. `float16` vectors have a lower precision compared to `float32`, which means that there is less numerical accuracy in the vector values - but this is negligible for practical use cases.
@@ -669,7 +669,7 @@ documentation, making it easier to navigate and find the information you need.
## S3 Snapshot Storage
Qdrant **Collections**, **Shards** and **Storage** can be backed up with [Snapshots](/documentation/concepts/snapshots/) and saved in case of data loss or other data transfer purposes. These snapshots can be quite large and the resources required to maintain them can result in higher costs. AWS S3 and other S3-compatible implementations like [min.io](https://min.io/) is a great low-cost alternative that can hold snapshots without incurring high costs. It is globally reliable, scalable and resistant to data loss.
Qdrant **Collections**, **Shards** and **Storage** can be backed up with [Snapshots](/documentation/operations/snapshots/) and saved in case of data loss or other data transfer purposes. These snapshots can be quite large and the resources required to maintain them can result in higher costs. AWS S3 and other S3-compatible implementations like [min.io](https://min.io/) is a great low-cost alternative that can hold snapshots without incurring high costs. It is globally reliable, scalable and resistant to data loss.
You can configure S3 storage settings in the [config.yaml](https://github.com/qdrant/qdrant/blob/master/config/config.yaml), specifically with `snapshots_storage`.
@@ -697,7 +697,7 @@ storage:
secret_key: your_secret_key_here
```
*Read more about [S3 snapshot storage](/documentation/concepts/snapshots/#s3) and [configuration](/documentation/guides/configuration/).*
*Read more about [S3 snapshot storage](/documentation/operations/snapshots/#s3) and [configuration](/documentation/operations/configuration/).*
This integration allows for a more convenient distribution of snapshots. Users of **any S3-compatible object storage** can now benefit from other platform services, such as automated workflows and disaster recovery options. S3's encryption and access control ensure secure storage and regulatory compliance. Additionally, S3 supports performance optimization through various storage classes and efficient data transfer methods, enabling quick and effective snapshot retrieval and management.
+7 -7
View File
@@ -36,7 +36,7 @@ New Web UI Tools:</br>
Before we dive into the specifics of our optimizations, let's first go over Multitenancy. This is one of our most significant features, [best used for scaling and data isolation](https://qdrant.tech/articles/multitenancy/).
If you’re using Qdrant to manage data for multiple users, regions, or workspaces (tenants), we suggest setting up a [multitenant environment](/documentation/guides/multiple-partitions/). This approach keeps all tenant data in a single global collection, with points separated and isolated by their payload.
If you’re using Qdrant to manage data for multiple users, regions, or workspaces (tenants), we suggest setting up a [multitenant environment](/documentation/manage-data/multitenancy/). This approach keeps all tenant data in a single global collection, with points separated and isolated by their payload.
To avoid slow and unnecessary indexing, it’s better to create an index for each relevant payload rather than indexing the entire collection globally. Since some data is indexed more frequently, you can focus on building indexes for specific regions, workspaces, or users.
@@ -156,7 +156,7 @@ await client.CreatePayloadIndexAsync(
As a result, the storage structure will be organized in a way to co-locate vectors of the same tenant together at the next optimization.
*To learn more about defragmentation, read the [Multitenancy documentation](/documentation/guides/multiple-partitions/).*
*To learn more about defragmentation, read the [Multitenancy documentation](/documentation/manage-data/multitenancy/).*
### On-Disk Support for the Payload Index
@@ -281,7 +281,7 @@ await client.CreatePayloadIndexAsync(
By moving the index to disk, Qdrant can handle larger datasets that exceed the capacity of RAM, making the system more scalable and capable of storing more data without being constrained by memory limitations.
*To learn more about this, read the [Indexing documentation](/documentation/concepts/indexing/).*
*To learn more about this, read the [Indexing documentation](/documentation/manage-data/indexing/).*
### UUID Datatype for the Payload Index
@@ -311,7 +311,7 @@ PUT /collections/{collection_name}/points
> For organizations that have numerous users and UUIDs, this simple fix can significantly reduce the cluster size and improve efficiency.
*To learn more about this, read the [Payload documentation](/documentation/concepts/payload/).*
*To learn more about this, read the [Payload documentation](/documentation/manage-data/payload/).*
### Query API: Groups Endpoint
@@ -411,7 +411,7 @@ await client.QueryGroupsAsync(
This endpoint will retrieve the best N points for each document, assuming that the payload of the points contains the document ID. Sometimes, the best N points cannot be fulfilled due to lack of points or a big distance with respect to the query. In every case, the `group_size` is a best-effort parameter, similar to the limit parameter.
*For more information on grouping capabilities refer to our [Hybrid Queries documentation](/documentation/concepts/hybrid-queries/).*
*For more information on grouping capabilities refer to our [Hybrid Queries documentation](/documentation/search/hybrid-queries/).*
### Query API: Random Sampling
@@ -490,7 +490,7 @@ await client.QueryAsync(
);
```
*To learn more, check out the [Query API documentation](/documentation/concepts/hybrid-queries/).*
*To learn more, check out the [Query API documentation](/documentation/search/hybrid-queries/).*
### Query API: Distribution-Based Score Fusion
@@ -658,7 +658,7 @@ await client.QueryAsync(
Note that `dbsf` is stateless and calculates the normalization limits only based on the results of each query, not on all the scores that it has seen.
*To learn more, check out the [Hybrid Queries documentation](/documentation/concepts/hybrid-queries/).*
*To learn more, check out the [Hybrid Queries documentation](/documentation/search/hybrid-queries/).*
## Web UI: Search Quality Tool
+9 -9
View File
@@ -113,7 +113,7 @@ Two arrays, `offsets_row` and `offsets_col`, represent the positions of non-zero
}
}
```
*To learn more about the distance matrix, read [**The Distance Matrix documentation**](/documentation/concepts/explore/#distance-matrix).*
*To learn more about the distance matrix, read [**The Distance Matrix documentation**](/documentation/search/explore/#distance-matrix).*
## Distance Matrix API in the Graph UI
@@ -134,7 +134,7 @@ The new graphing method is cleaner and reveals **relationships and outliers:**
![distance-matrix](/blog/qdrant-1.12.x/distance-matrix.png)
*To learn more about the Web UI Dashboard, read the [**Interfaces documentation**](/documentation/interfaces/web-ui/).*
*To learn more about the Web UI Dashboard, read the [**Interfaces documentation**](/documentation/web-ui/).*
## Facet API for Metadata Cardinality
@@ -142,7 +142,7 @@ The new graphing method is cleaner and reveals **relationships and outliers:**
In modern applications like e-commerce, users often rely on [**filters**](/articles/vector-search-filtering/), such as **brand** or **color**, to refine search results. The **Facet API** is designed to help users understand the distribution of values in a dataset.
The `facet` endpoint can efficiently count and aggregate values for a specific [**payload field**](/documentation/concepts/payload/) in your dataset.
The `facet` endpoint can efficiently count and aggregate values for a specific [**payload field**](/documentation/manage-data/payload/) in your dataset.
You can use it to retrieve unique values for a field, along with the number of points that contain each value. This functionality is similar to `GROUP BY` with `COUNT(*)` in SQL databases.
@@ -195,12 +195,12 @@ POST /collections/{collection_name}/facet
```
This feature provides flexibility between performance and precision, depending on the needs of your application.
*To learn more about faceting, read the [**Facet API documentation**](/documentation/concepts/payload/#facet-counts).*
*To learn more about faceting, read the [**Facet API documentation**](/documentation/manage-data/payload/#facet-counts).*
## Text Index on Disk Support
![text-index-disk](/blog/qdrant-1.12.x/text-index-disk.png)
[**Qdrant text indexing**](/documentation/concepts/indexing/#full-text-index) tokenizes text into smaller units (tokens) based on chosen settings (e.g., tokenizer type, token length). These tokens are stored in an inverted index for fast text searches.
[**Qdrant text indexing**](/documentation/manage-data/indexing/#full-text-index) tokenizes text into smaller units (tokens) based on chosen settings (e.g., tokenizer type, token length). These tokens are stored in an inverted index for fast text searches.
> With `on_disk` text indexing, the inverted index is stored on disk, reducing memory usage.
@@ -222,11 +222,11 @@ PUT /collections/{collection_name}/index
}
```
*To learn more about indexes, read the [**Indexing documentation**](/documentation/concepts/indexing/).*
*To learn more about indexes, read the [**Indexing documentation**](/documentation/manage-data/indexing/).*
## Geo Index on Disk Support
For [**large-scale geographic datasets**](/documentation/concepts/payload/#geo) where storing all indexes in memory is impractical, **geo indexing** allows efficient filtering of points based on geographic coordinates.
For [**large-scale geographic datasets**](/documentation/manage-data/payload/#geo) where storing all indexes in memory is impractical, **geo indexing** allows efficient filtering of points based on geographic coordinates.
With `on_disk` geo indexing, the index is written to disk instead of residing in memory, making it possible to handle large datasets without exhausting system memory.
@@ -255,11 +255,11 @@ PUT /collections/{collection_name}/index
![geo-index-disk](/blog/qdrant-1.12.x/geo-index-disk.png)
> To learn how to get the best performance from Qdrant, read the [**Optimization Guide**](/documentation/guides/optimize/).
> To learn how to get the best performance from Qdrant, read the [**Optimization Guide**](/documentation/operations/optimize/).
## Just the Beginning
The easiest way to reach that **Hello World** moment is to [**try vector search in a live cluster**](/documentation/quickstart-cloud/). Our **interactive tutorial** will show you how to create a cluster, add data and try some filtering clauses.
The easiest way to reach that **Hello World** moment is to [**try vector search in a live cluster**](/documentation/cloud-quickstart/). Our **interactive tutorial** will show you how to create a cluster, add data and try some filtering clauses.
**All of the new features from version 1.12 can be tested in the Web UI:**

Some files were not shown because too many files have changed in this diff Show More