From 6b13919e9d4a085b02ad692156fdbbb97574f4ee Mon Sep 17 00:00:00 2001
From: kartik-gupta-ij
Date: Wed, 8 May 2024 20:19:08 +0530
Subject: [PATCH 01/32] Optimizers configuration question added
---
.../documentation/faq/database-optimization.md | 15 +++++++++++++++
1 file changed, 15 insertions(+)
diff --git a/qdrant-landing/content/documentation/faq/database-optimization.md b/qdrant-landing/content/documentation/faq/database-optimization.md
index 857c092f9..9dcc34f6b 100644
--- a/qdrant-landing/content/documentation/faq/database-optimization.md
+++ b/qdrant-landing/content/documentation/faq/database-optimization.md
@@ -42,3 +42,18 @@ There are several possible reasons for that:
- **Using filters without payload index** -- If you're performing a search with a filter but you don't have a payload index, Qdrant will have to load whole payload data from disk to check the filtering condition. Ensure you have adequately configured [payload indexes](../../concepts/indexing/#payload-index).
- **Usage of on-disk vector storage with slow disks** -- If you're using on-disk vector storage, ensure you have fast enough disks. We recommend using local SSDs with at least 50k IOPS. Read more about the influence of the disk speed on the search latency in the article about [Memory Consumption](../../../articles/memory-consumption/).
- **Large limit or non-optimal query parameters** -- A large limit or offset might lead to significant performance degradation. Please pay close attention to the query/collection parameters that significantly diverge from the defaults. They might be the reason for the performance issues.
+
+
+
+### How can I optimize optimizers configuration settings for better accuracy and speed, especially for large-scale collections?
+
+To optimize optimizers config in Qdrant for better accuracy and speed, consider the following:
+- For low memory footprint with high-speed search, utilize vector quantization with disk storage for vectors and in-memory quantized vectors. Configure memmap_threshold and always_ram accordingly.
+- To prioritize high precision with a low memory footprint, enable on-disk vectors and HNSW index. Adjust HNSW parameters for precision while considering disk IOPS.
+- For high precision with high-speed search, focus on keeping data in RAM. Utilize quantization with re-scoring and adjust search-time parameters like hnsw_ef for accuracy and speed balance.
+- Balance latency vs. throughput based on your needs. Configure the number of segments to utilize CPU cores effectively for either minimizing latency or maximizing throughput.
+- Explore different quantization methods like scalar, binary, and product quantization based on accuracy, speed, and compression requirements.
+- Fine-tune quantization parameters such as quantile for scalar quantization to optimize search precision and memory usage.
+- Adjust storage modes to balance between memory footprint and search speed, considering factors like disk reads and storage type (RAM, SSD, or HDD).
+
+Read more about [optimizing](../../guides/optimize/), and [quantization](../../guides/quantization/) in Qdrant.
\ No newline at end of file
From 4bb2b9df66b02ba2f7470beccb7f605e5779c415 Mon Sep 17 00:00:00 2001
From: kartik-gupta-ij
Date: Wed, 8 May 2024 20:58:31 +0530
Subject: [PATCH 02/32] docs: Add Qdrant search performance enhancement with
filters
---
.../faq/database-optimization.md | 26 ++++++++++++++-----
1 file changed, 19 insertions(+), 7 deletions(-)
diff --git a/qdrant-landing/content/documentation/faq/database-optimization.md b/qdrant-landing/content/documentation/faq/database-optimization.md
index 9dcc34f6b..dc419b42b 100644
--- a/qdrant-landing/content/documentation/faq/database-optimization.md
+++ b/qdrant-landing/content/documentation/faq/database-optimization.md
@@ -12,8 +12,8 @@ The primary source of memory usage vector data. There are several ways to addres
- Configure [Quantization](../../guides/quantization/) to reduce the memory usage of vectors.
- Configure on-disk vector storage
-The choice of the approach depends on your requirements.
-Read more about [configuring the optimal](../../tutorials/optimize/) use of Qdrant.
+The choice of the approach depends on your requirements.
+Read more about [configuring the optimal](../../tutorials/optimize/) use of Qdrant.
### How do you choose machine configuration?
@@ -22,7 +22,7 @@ There are two main scenarios of Qdrant usage in terms of resource consumption:
- **Performance-optimized** -- when you need to serve vector search as fast (many) as possible. In this case, you need to have as much vector data in RAM as possible. Use our [calculator](https://cloud.qdrant.io/calculator) to estimate the required RAM.
- **Storage-optimized** -- when you need to store many vectors and minimize costs by compromising some search speed. In this case, pay attention to the disk speed instead. More about it in the article about [Memory Consumption](../../../articles/memory-consumption/).
-### I configured on-disk vector storage, but memory usage is still high. Why?
+### I configured on-disk vector storage, but memory usage is still high. Why?
Firstly, memory usage metrics as reported by `top` or `htop` may be misleading. They are not showing the minimal amount of memory required to run the service.
If the RSS memory usage is 10 GB, it doesn't mean that it won't work on a machine with 8 GB of RAM.
@@ -34,7 +34,6 @@ As a result, the Qdrant process might use more memory than the minimum required
If you want to limit the memory usage of the service, we recommend using [limits in Docker](https://docs.docker.com/config/containers/resource_constraints/#memory) or Kubernetes.
-
### My requests are very slow or time out. What should I do?
There are several possible reasons for that:
@@ -43,11 +42,10 @@ There are several possible reasons for that:
- **Usage of on-disk vector storage with slow disks** -- If you're using on-disk vector storage, ensure you have fast enough disks. We recommend using local SSDs with at least 50k IOPS. Read more about the influence of the disk speed on the search latency in the article about [Memory Consumption](../../../articles/memory-consumption/).
- **Large limit or non-optimal query parameters** -- A large limit or offset might lead to significant performance degradation. Please pay close attention to the query/collection parameters that significantly diverge from the defaults. They might be the reason for the performance issues.
-
-
### How can I optimize optimizers configuration settings for better accuracy and speed, especially for large-scale collections?
To optimize optimizers config in Qdrant for better accuracy and speed, consider the following:
+
- For low memory footprint with high-speed search, utilize vector quantization with disk storage for vectors and in-memory quantized vectors. Configure memmap_threshold and always_ram accordingly.
- To prioritize high precision with a low memory footprint, enable on-disk vectors and HNSW index. Adjust HNSW parameters for precision while considering disk IOPS.
- For high precision with high-speed search, focus on keeping data in RAM. Utilize quantization with re-scoring and adjust search-time parameters like hnsw_ef for accuracy and speed balance.
@@ -56,4 +54,18 @@ To optimize optimizers config in Qdrant for better accuracy and speed, consider
- Fine-tune quantization parameters such as quantile for scalar quantization to optimize search precision and memory usage.
- Adjust storage modes to balance between memory footprint and search speed, considering factors like disk reads and storage type (RAM, SSD, or HDD).
-Read more about [optimizing](../../guides/optimize/), and [quantization](../../guides/quantization/) in Qdrant.
\ No newline at end of file
+Read more about [optimizing](../../guides/optimize/), and [quantization](../../guides/quantization/) in Qdrant.
+
+### How can Qdrant I enhance search performance when applying filters to their queries?
+
+To improve search performance with filters in Qdrant:
+
+- **Utilize Payload Index:** Implement payload indexes for fields used in filtering conditions. This enhances point retrieval speed based on filter criteria.
+- **Consider Full-text Index:** For string payloads, utilize full-text indexes for word or phrase presence filtering, adjusting tokenization parameters as needed.
+- **Parameterized Integer Index:** Opt for parameterized integer indexes for specific range filter requirements, balancing between lookup and range operations.
+- **Leverage Vector Index:** Employ HNSW for dense vector indexing, adjusting parameters like `m` and `ef_construct` for improved search accuracy and efficiency.
+- **Sparse Vector Index:** For sparse vectors, ensure efficient indexing and search mechanisms. Configure memory or disk storage based on segment characteristics and query patterns.
+- **Filtrable Index Extension:** Enhance HNSW graph with additional edges based on payload values to optimize search efficiency with varying filter stringency.
+
+By strategically combining these indexing techniques, you can address performance issues and achieve faster search results with applied filters in Qdrant.
+Read more about [indexing](../../concepts/indexing/) in Qdrant.
From 2f2643c02be65922003840f8233baa384d5da04e Mon Sep 17 00:00:00 2001
From: kartik-gupta-ij
Date: Thu, 9 May 2024 00:30:47 +0530
Subject: [PATCH 03/32] fix
---
.../content/documentation/faq/database-optimization.md | 6 +++---
1 file changed, 3 insertions(+), 3 deletions(-)
diff --git a/qdrant-landing/content/documentation/faq/database-optimization.md b/qdrant-landing/content/documentation/faq/database-optimization.md
index dc419b42b..15139a85a 100644
--- a/qdrant-landing/content/documentation/faq/database-optimization.md
+++ b/qdrant-landing/content/documentation/faq/database-optimization.md
@@ -42,9 +42,9 @@ There are several possible reasons for that:
- **Usage of on-disk vector storage with slow disks** -- If you're using on-disk vector storage, ensure you have fast enough disks. We recommend using local SSDs with at least 50k IOPS. Read more about the influence of the disk speed on the search latency in the article about [Memory Consumption](../../../articles/memory-consumption/).
- **Large limit or non-optimal query parameters** -- A large limit or offset might lead to significant performance degradation. Please pay close attention to the query/collection parameters that significantly diverge from the defaults. They might be the reason for the performance issues.
-### How can I optimize optimizers configuration settings for better accuracy and speed, especially for large-scale collections?
+### How can I optimize optimizer's configuration settings for better accuracy and speed, especially for large-scale collections?
-To optimize optimizers config in Qdrant for better accuracy and speed, consider the following:
+To optimize optimizer's config in Qdrant for better accuracy and speed, consider the following:
- For low memory footprint with high-speed search, utilize vector quantization with disk storage for vectors and in-memory quantized vectors. Configure memmap_threshold and always_ram accordingly.
- To prioritize high precision with a low memory footprint, enable on-disk vectors and HNSW index. Adjust HNSW parameters for precision while considering disk IOPS.
@@ -56,7 +56,7 @@ To optimize optimizers config in Qdrant for better accuracy and speed, consider
Read more about [optimizing](../../guides/optimize/), and [quantization](../../guides/quantization/) in Qdrant.
-### How can Qdrant I enhance search performance when applying filters to their queries?
+### How can I enhance search performance when applying filters to their queries?
To improve search performance with filters in Qdrant:
From 0b22e2008d6f3ef706c53ec26f8fac9038f9897e Mon Sep 17 00:00:00 2001
From: kartik-gupta-ij
Date: Thu, 9 May 2024 00:53:05 +0530
Subject: [PATCH 04/32] Fix answer 2
---
.../documentation/faq/database-optimization.md | 11 +----------
1 file changed, 1 insertion(+), 10 deletions(-)
diff --git a/qdrant-landing/content/documentation/faq/database-optimization.md b/qdrant-landing/content/documentation/faq/database-optimization.md
index 15139a85a..3d70326c4 100644
--- a/qdrant-landing/content/documentation/faq/database-optimization.md
+++ b/qdrant-landing/content/documentation/faq/database-optimization.md
@@ -58,14 +58,5 @@ Read more about [optimizing](../../guides/optimize/), and [quantization](../../g
### How can I enhance search performance when applying filters to their queries?
-To improve search performance with filters in Qdrant:
-
-- **Utilize Payload Index:** Implement payload indexes for fields used in filtering conditions. This enhances point retrieval speed based on filter criteria.
-- **Consider Full-text Index:** For string payloads, utilize full-text indexes for word or phrase presence filtering, adjusting tokenization parameters as needed.
-- **Parameterized Integer Index:** Opt for parameterized integer indexes for specific range filter requirements, balancing between lookup and range operations.
-- **Leverage Vector Index:** Employ HNSW for dense vector indexing, adjusting parameters like `m` and `ef_construct` for improved search accuracy and efficiency.
-- **Sparse Vector Index:** For sparse vectors, ensure efficient indexing and search mechanisms. Configure memory or disk storage based on segment characteristics and query patterns.
-- **Filtrable Index Extension:** Enhance HNSW graph with additional edges based on payload values to optimize search efficiency with varying filter stringency.
-
-By strategically combining these indexing techniques, you can address performance issues and achieve faster search results with applied filters in Qdrant.
+To enhance the search performance with the filters in qdrant, it's important to optimize the indexing strategies. Users can get a combination of vector and traditional indexes from qdrant, where the vector indexes reduce the time taken for the vector search and the payload indexes quicken the pace of filtering. Users need to strategically mark the fields as indexable as well as prioritize the fields that appear frequently in the filtering conditions in order to efficiently utilize the memory resources. Through thorough cosideration of the memory constraints as well as careful index configuration, users can effectively enhance search performance with filters in qdrant.
Read more about [indexing](../../concepts/indexing/) in Qdrant.
From ee1f9554167c143ebc334d57ac96bab2c3554535 Mon Sep 17 00:00:00 2001
From: kartik-gupta-ij
Date: Mon, 13 May 2024 00:51:44 +0530
Subject: [PATCH 05/32] docs: Which embedding models does FastEmbed support,
and how does it utilize memory?
---
.../documentation/faq/database-optimization.md | 1 +
.../content/documentation/faq/qdrant-fundamentals.md | 11 +++++++++++
2 files changed, 12 insertions(+)
diff --git a/qdrant-landing/content/documentation/faq/database-optimization.md b/qdrant-landing/content/documentation/faq/database-optimization.md
index 3d70326c4..7d32d6923 100644
--- a/qdrant-landing/content/documentation/faq/database-optimization.md
+++ b/qdrant-landing/content/documentation/faq/database-optimization.md
@@ -59,4 +59,5 @@ Read more about [optimizing](../../guides/optimize/), and [quantization](../../g
### How can I enhance search performance when applying filters to their queries?
To enhance the search performance with the filters in qdrant, it's important to optimize the indexing strategies. Users can get a combination of vector and traditional indexes from qdrant, where the vector indexes reduce the time taken for the vector search and the payload indexes quicken the pace of filtering. Users need to strategically mark the fields as indexable as well as prioritize the fields that appear frequently in the filtering conditions in order to efficiently utilize the memory resources. Through thorough cosideration of the memory constraints as well as careful index configuration, users can effectively enhance search performance with filters in qdrant.
+
Read more about [indexing](../../concepts/indexing/) in Qdrant.
diff --git a/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md b/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
index fb623701b..12185abb7 100644
--- a/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
+++ b/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
@@ -81,3 +81,14 @@ We only guarantee compatibility if you update between consecutive versions. You
In case your version is older, we only guarantee compatibility between two consecutive minor versions. This also applies to client versions. Ensure your client version is never more than one minor version away from your cluster version.
While we will assist with break/fix troubleshooting of issues and errors specific to our products, Qdrant is not accountable for reviewing, writing (or rewriting), or debugging custom code.
+
+### Which embedding models does FastEmbed support, and how does it utilize memory?
+
+You can find the list of supported models [here](https://qdrant.github.io/fastembed/examples/Supported_Models/).
+
+Memory usage in fastEmbed depends on several factors:
+1. **Size of the Text**: Everything, including the text and its vectors, is loaded into memory. As vectors are computed, RAM consumption increases accordingly.
+2. **Data Parallelism**: fastEmbed utilizes Python's multiprocessing to split large lists of strings into smaller ones and process them in parallel. While this enhances processing speed, it can also lead to higher RAM consumption.
+3. **Model Used**: Memory usage varies depending on the model used to embed your data.
+
+For optimal performance and memory management, consider these factors when using fastEmbed.
\ No newline at end of file
From 476de391c34fb6ac15d4f93c34ce3911e8e4e836 Mon Sep 17 00:00:00 2001
From: kartik-gupta-ij
Date: Mon, 13 May 2024 11:17:08 +0530
Subject: [PATCH 06/32] Docs : Update Qdrant documentation with FAQs and
answers
---
.../documentation/faq/qdrant-fundamentals.md | 23 ++++++++++++++++++-
1 file changed, 22 insertions(+), 1 deletion(-)
diff --git a/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md b/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
index 12185abb7..04385a9a2 100644
--- a/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
+++ b/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
@@ -91,4 +91,25 @@ Memory usage in fastEmbed depends on several factors:
2. **Data Parallelism**: fastEmbed utilizes Python's multiprocessing to split large lists of strings into smaller ones and process them in parallel. While this enhances processing speed, it can also lead to higher RAM consumption.
3. **Model Used**: Memory usage varies depending on the model used to embed your data.
-For optimal performance and memory management, consider these factors when using fastEmbed.
\ No newline at end of file
+For optimal performance and memory management, consider these factors when using fastEmbed.
+
+### To ensure all shards are replicated to the listener node, can I use the "Creating new shard replicas" API right after bootstrapping a new cluster and loading data, but before clients start sending new requests?
+
+Yes, that's correct. Before clients start sending requests, you can use the "Creating new shard replicas" API to replicate all shards to the listener node. This ensures that all new data will be replicated to the data node, facilitating whole cluster backup or scaling out.
+
+### Does a listener node in a cluster receive data from all shards?
+
+Yes, a listener node can indeed receive data from all shards. However, it's important to note that a listener node doesn't participate in searches.
+
+### When creating snapshots from the listener node, will they include data from all shards by default?
+
+No, snapshots from the listener node won't automatically include data from all shards. You'll need to manually replicate all shards to the listener node. You can use the `replicate_shard` API for this purpose.
+
+
+### If I want to use a listener node solely for data storage and scaling out my cluster, can I do so without affecting other nodes with query traffic?
+
+Absolutely. You can configure a listener node to act as a pure data node, thereby scaling out your cluster without impacting other nodes with query traffic.
+
+### How does your cloud handle shard rebalancing when increasing the number of nodes?
+
+Our [cloud platform](https://cloud.qdrant.io) handles shard rebalancing when scaling out. However, it's essential to create enough shards beforehand to facilitate the scaling process effectively. For instance, if you have 3 nodes, it's advisable to choose 6 or 9 shards to allow for rebalancing upon extending your cluster, which wouldn't be possible with just 3 shards.
\ No newline at end of file
From b834bd1a5a13e84d87105e74f06720c6ff872dff Mon Sep 17 00:00:00 2001
From: kartik-gupta-ij
Date: Mon, 13 May 2024 13:49:29 +0530
Subject: [PATCH 07/32] docs: removing listener node related question from the
FAQ
---
.../documentation/faq/qdrant-fundamentals.md | 17 -----------------
1 file changed, 17 deletions(-)
diff --git a/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md b/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
index 04385a9a2..ba983b62a 100644
--- a/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
+++ b/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
@@ -93,23 +93,6 @@ Memory usage in fastEmbed depends on several factors:
For optimal performance and memory management, consider these factors when using fastEmbed.
-### To ensure all shards are replicated to the listener node, can I use the "Creating new shard replicas" API right after bootstrapping a new cluster and loading data, but before clients start sending new requests?
-
-Yes, that's correct. Before clients start sending requests, you can use the "Creating new shard replicas" API to replicate all shards to the listener node. This ensures that all new data will be replicated to the data node, facilitating whole cluster backup or scaling out.
-
-### Does a listener node in a cluster receive data from all shards?
-
-Yes, a listener node can indeed receive data from all shards. However, it's important to note that a listener node doesn't participate in searches.
-
-### When creating snapshots from the listener node, will they include data from all shards by default?
-
-No, snapshots from the listener node won't automatically include data from all shards. You'll need to manually replicate all shards to the listener node. You can use the `replicate_shard` API for this purpose.
-
-
-### If I want to use a listener node solely for data storage and scaling out my cluster, can I do so without affecting other nodes with query traffic?
-
-Absolutely. You can configure a listener node to act as a pure data node, thereby scaling out your cluster without impacting other nodes with query traffic.
-
### How does your cloud handle shard rebalancing when increasing the number of nodes?
Our [cloud platform](https://cloud.qdrant.io) handles shard rebalancing when scaling out. However, it's essential to create enough shards beforehand to facilitate the scaling process effectively. For instance, if you have 3 nodes, it's advisable to choose 6 or 9 shards to allow for rebalancing upon extending your cluster, which wouldn't be possible with just 3 shards.
\ No newline at end of file
From 610312cc5189ea409d936bdb39cc731ce223615d Mon Sep 17 00:00:00 2001
From: kartik-gupta-ij
Date: Mon, 13 May 2024 14:02:56 +0530
Subject: [PATCH 08/32] docs: removing optimizer's configuration from FAq
---
.../documentation/faq/database-optimization.md | 14 --------------
1 file changed, 14 deletions(-)
diff --git a/qdrant-landing/content/documentation/faq/database-optimization.md b/qdrant-landing/content/documentation/faq/database-optimization.md
index 7d32d6923..ceea6b727 100644
--- a/qdrant-landing/content/documentation/faq/database-optimization.md
+++ b/qdrant-landing/content/documentation/faq/database-optimization.md
@@ -42,20 +42,6 @@ There are several possible reasons for that:
- **Usage of on-disk vector storage with slow disks** -- If you're using on-disk vector storage, ensure you have fast enough disks. We recommend using local SSDs with at least 50k IOPS. Read more about the influence of the disk speed on the search latency in the article about [Memory Consumption](../../../articles/memory-consumption/).
- **Large limit or non-optimal query parameters** -- A large limit or offset might lead to significant performance degradation. Please pay close attention to the query/collection parameters that significantly diverge from the defaults. They might be the reason for the performance issues.
-### How can I optimize optimizer's configuration settings for better accuracy and speed, especially for large-scale collections?
-
-To optimize optimizer's config in Qdrant for better accuracy and speed, consider the following:
-
-- For low memory footprint with high-speed search, utilize vector quantization with disk storage for vectors and in-memory quantized vectors. Configure memmap_threshold and always_ram accordingly.
-- To prioritize high precision with a low memory footprint, enable on-disk vectors and HNSW index. Adjust HNSW parameters for precision while considering disk IOPS.
-- For high precision with high-speed search, focus on keeping data in RAM. Utilize quantization with re-scoring and adjust search-time parameters like hnsw_ef for accuracy and speed balance.
-- Balance latency vs. throughput based on your needs. Configure the number of segments to utilize CPU cores effectively for either minimizing latency or maximizing throughput.
-- Explore different quantization methods like scalar, binary, and product quantization based on accuracy, speed, and compression requirements.
-- Fine-tune quantization parameters such as quantile for scalar quantization to optimize search precision and memory usage.
-- Adjust storage modes to balance between memory footprint and search speed, considering factors like disk reads and storage type (RAM, SSD, or HDD).
-
-Read more about [optimizing](../../guides/optimize/), and [quantization](../../guides/quantization/) in Qdrant.
-
### How can I enhance search performance when applying filters to their queries?
To enhance the search performance with the filters in qdrant, it's important to optimize the indexing strategies. Users can get a combination of vector and traditional indexes from qdrant, where the vector indexes reduce the time taken for the vector search and the payload indexes quicken the pace of filtering. Users need to strategically mark the fields as indexable as well as prioritize the fields that appear frequently in the filtering conditions in order to efficiently utilize the memory resources. Through thorough cosideration of the memory constraints as well as careful index configuration, users can effectively enhance search performance with filters in qdrant.
From 0fa6ff954f5a3ab6f203719b3d2509b90c5f86a8 Mon Sep 17 00:00:00 2001
From: kartik-gupta-ij
Date: Mon, 13 May 2024 14:19:36 +0530
Subject: [PATCH 09/32] docs: remove cloud handle shard rebalancing from FAQ
---
.../content/documentation/faq/qdrant-fundamentals.md | 6 +-----
1 file changed, 1 insertion(+), 5 deletions(-)
diff --git a/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md b/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
index ba983b62a..12185abb7 100644
--- a/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
+++ b/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md
@@ -91,8 +91,4 @@ Memory usage in fastEmbed depends on several factors:
2. **Data Parallelism**: fastEmbed utilizes Python's multiprocessing to split large lists of strings into smaller ones and process them in parallel. While this enhances processing speed, it can also lead to higher RAM consumption.
3. **Model Used**: Memory usage varies depending on the model used to embed your data.
-For optimal performance and memory management, consider these factors when using fastEmbed.
-
-### How does your cloud handle shard rebalancing when increasing the number of nodes?
-
-Our [cloud platform](https://cloud.qdrant.io) handles shard rebalancing when scaling out. However, it's essential to create enough shards beforehand to facilitate the scaling process effectively. For instance, if you have 3 nodes, it's advisable to choose 6 or 9 shards to allow for rebalancing upon extending your cluster, which wouldn't be possible with just 3 shards.
\ No newline at end of file
+For optimal performance and memory management, consider these factors when using fastEmbed.
\ No newline at end of file
From 3901c4a698db57c7842a098997d42738be4612fa Mon Sep 17 00:00:00 2001
From: timvisee
Date: Fri, 5 Jul 2024 12:03:16 +0200
Subject: [PATCH 10/32] Remove note about getting duplicate points if matching
multiple payloads
---
qdrant-landing/content/documentation/concepts/points.md | 2 --
1 file changed, 2 deletions(-)
diff --git a/qdrant-landing/content/documentation/concepts/points.md b/qdrant-landing/content/documentation/concepts/points.md
index f86fcee3d..97f308d0f 100755
--- a/qdrant-landing/content/documentation/concepts/points.md
+++ b/qdrant-landing/content/documentation/concepts/points.md
@@ -1751,8 +1751,6 @@ new OrderBy
};
```
-**Note:** for payloads with more than one value (such as arrays), the same point may show up more than once. Each point can appear as many times as the number of elements in the array. For example, if you have a point payload with a `timestamp` key, and the value for the key is an array of 3 elements, the same point will appear 3 times in the results, one for each timestamp.
-
When sorting is based on a non-unique value, it is not possible to rely on an ID offset. Thus, next_page_offset is not returned within the response. However, you can still do pagination by combining `"order_by": { "start_from": ... }` with a `{ "must_not": [{ "has_id": [...] }] }` filter.
From d62b30f995a12f2174c618d3dccaaf0a6078c8ab Mon Sep 17 00:00:00 2001
From: Bastian Hofmann
Date: Fri, 12 Jul 2024 15:52:05 +0200
Subject: [PATCH 11/32] Small improvements for the Hybrid Cloud docs
This clarifies prerequisites and adds information about the new config UI
---
.../hybrid-cloud/hybrid-cloud-setup.md | 38 +++++++++++++++++--
1 file changed, 35 insertions(+), 3 deletions(-)
diff --git a/qdrant-landing/content/documentation/hybrid-cloud/hybrid-cloud-setup.md b/qdrant-landing/content/documentation/hybrid-cloud/hybrid-cloud-setup.md
index edb9002b9..670476a78 100644
--- a/qdrant-landing/content/documentation/hybrid-cloud/hybrid-cloud-setup.md
+++ b/qdrant-landing/content/documentation/hybrid-cloud/hybrid-cloud-setup.md
@@ -5,7 +5,7 @@ weight: 1
# Creating a Hybrid Cloud Environment
-The following instruction set will show you how to properly setup a **Qdrant cluster** in your **Hybrid Cloud Environment**.
+The following instruction set will show you how to properly set up a **Qdrant cluster** in your **Hybrid Cloud Environment**.
To learn how Hybrid Cloud works, [read the overview document](/documentation/hybrid-cloud/).
@@ -22,6 +22,15 @@ To learn how Hybrid Cloud works, [read the overview document](/documentation/hyb
> **Note:** You can also mirror these images and charts into your own registry and pull them from there.
+### CLI tools
+
+During the onboarding, you will need to deploy the Qdrant Kubernetes Operator and Agent using Helm. Make sure you have the following tools installed:
+
+* [kubectl](https://kubernetes.io/docs/tasks/tools/install-kubectl/)
+* [helm](https://helm.sh/docs/intro/install/)
+
+You will need to have access to the Kubernetes cluster with `kubectl` and `helm` configured to connect to it. Please refer the documentation of your Kubernetes distribution for more information.
+
### Required artifacts
Container images:
@@ -29,7 +38,7 @@ Container images:
- `docker.io/qdrant/qdrant`
- `registry.cloud.qdrant.io/qdrant/qdrant-cloud-agent`
- `registry.cloud.qdrant.io/qdrant/qdrant-operator`
-- `registry.cloud.qdrant.io/qdrant/qdrant-cloud-cluster-manager`
+- `registry.cloud.qdrant.io/qdrant/cluster-manager`
- `registry.cloud.qdrant.io/qdrant/prometheus`
- `registry.cloud.qdrant.io/qdrant/prometheus-config-reloader`
- `registry.cloud.qdrant.io/qdrant/kube-state-metrics`
@@ -38,6 +47,7 @@ Open Containers Initiative (OCI) Helm charts:
- `registry.cloud.qdrant.io/qdrant-charts/qdrant-cloud-agent`
- `registry.cloud.qdrant.io/qdrant-charts/qdrant-operator`
+- `registry.cloud.qdrant.io/qdrant-charts/qdrant-cluster-manager`
- `registry.cloud.qdrant.io/qdrant-charts/prometheus`
## Installation
@@ -53,6 +63,8 @@ Open Containers Initiative (OCI) Helm charts:
- **Name:** A name for the Hybrid Cloud Environment
- **Kubernetes Namespace:** The Kubernetes namespace for the operator and agent. Once you select a namespace, you can't change it.
+You can also configure the StorageClass and VolumeSnapshotClass to use for the Qdrant databases, if you want to deviate from the default settings of your cluster.
+
4. You can then enter the YAML configuration for your Kubernetes operator. Qdrant supports a specific list of configuration options, as described in the [Qdrant Operator configuration](/documentation/hybrid-cloud/operator-configuration/) section.
5. (Optional) If you have special requirements for any of the following, activate the **Show advanced configuration** option:
@@ -84,6 +96,15 @@ You need this command only for the initial installation. After that, you can upd
Once you have created a Hybrid Cloud Environment, you can create a Qdrant cluster in that enviroment. Use the same process to [Create a cluster](/documentation/cloud/create-cluster/). Make sure to select your Hybrid Cloud Environment as the target.
+Note that in the "Kubernetes Configuration" section you can configure:
+
+* Node selectors for the Qdrant database pods
+* Toleration for the Qdrant database pods
+* Additional labels for the Qdrant database pods
+* A service type and annotations for the Qdrant database service
+
+These settings can also be changed after the cluster is created on the cluster detail page.
+
### Authentication at your Qdrant clusters
In Hybrid Cloud the authentication information is provided with Kubernetes secrets.
@@ -124,7 +145,18 @@ kubectl -n qdrant-namespace port-forward service/qdrant-9a9f48c7-bb90-4fb2-816f-
You can also expose the database outside the Kubernetes cluster with a `LoadBalancer` (if supported in your Kubernetes environment) or `NodePort` service or an ingress.
-A simple Loadbalancer service could look like this:
+The service type and necessary annotations can be configured in the "Kubernetes Configuration" section during cluster creation, or on the cluster detail page.
+
+Especially if you create a LoadBalancer Service, you may need to provider annotations for the loadbalancer configration. Please refer to the documention of your cloud provider for more details.
+
+Examples:
+
+* [AWS EKS LoadBalancer annotations](https://kubernetes-sigs.github.io/aws-load-balancer-controller/latest/guide/ingress/annotations/)
+* [Azure AKS Public LoadBalancer annotations](https://learn.microsoft.com/en-us/azure/aks/load-balancer-standard)
+* [Azure AKS Internal LoadBalancer annotations](https://learn.microsoft.com/en-us/azure/aks/internal-lb)
+* [GCP GKE LoadBalancer annotations](https://cloud.google.com/kubernetes-engine/docs/concepts/service-load-balancer-parameters)
+
+You could also create a Loadbalancer service manually like this:
```yaml
apiVersion: v1
From c04ee922d0234d516cb24758645c72b140941a8f Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?Luis=20Coss=C3=ADo?=
Date: Fri, 12 Jul 2024 17:46:17 -0400
Subject: [PATCH 12/32] proofread article changes
---
.../content/articles/discovery-search.md | 24 ++++++++-----------
1 file changed, 10 insertions(+), 14 deletions(-)
diff --git a/qdrant-landing/content/articles/discovery-search.md b/qdrant-landing/content/articles/discovery-search.md
index 64dbbe1bb..ef05ba5dc 100644
--- a/qdrant-landing/content/articles/discovery-search.md
+++ b/qdrant-landing/content/articles/discovery-search.md
@@ -1,7 +1,7 @@
---
-title: "Discovery Search: A New Approach to Vector Space"
-short_description: Discovery Search, an innovative API for precise, tailored search results.
-description: Explore the next frontier in search technology with Discovery Search. Learn how this innovative API provides precise and tailored results.
+title: "Discovery needs context"
+short_description: Discover points by constraining the vector space.
+description: Discovery Search, an innovative way to constrain the vector space in which a search is performed, relying only on vectors.
social_preview_image: /articles_data/discovery-search/social_preview.jpg
small_preview_image: /articles_data/discovery-search/icon.svg
preview_dir: /articles_data/discovery-search/preview
@@ -14,28 +14,24 @@ keywords:
- why use a vector database
- specialty
- search
- - discovery
+ - multimodal
- state-of-the-art
- vector-search
---
-# How to Master Vector Space Exploration with Discovery Search
+# Discovery needs context
-When Christopher Columbus and his crew sailed to cross the Atlantic Ocean, they were not looking for America. They were looking for a new route to India, and they were convinced that the Earth was round. They didn't know anything about America, but since they were going west, they stumbled upon it.
+When Christopher Columbus and his crew sailed to cross the Atlantic Ocean, they were not looking for the Americas. They were looking for a new route to India because they were convinced that the Earth was round. They didn't know anything about a new continent, but since they were going west, they stumbled upon it.
They couldn't reach their _target_, because the geography didn't let them, but once they realized it wasn't India, they claimed it a new "discovery" for their crown. If we consider that sailors need water to sail, then we can establish a _context_ which is positive in the water, and negative on land. Once the sailor's search was stopped by the land, they could not go any further, and a new route was found. Let's keep these concepts of _target_ and _context_ in mind as we explore the new functionality of Qdrant: __Discovery search__.
## What is discovery search?
-Discovery search is a powerful tool that lets you explore the vector space in a more controlled way. It can be used to find points that are not necessarily close to the target but are still relevant to the search. It can also be used to represent complex tastes and break out of the similarity bubble. Check out the documentation to learn more about the math behind it and how to use it.
-
-## Qdrant's discovery search: version 1.7 release
-
In version 1.7, Qdrant [released](/articles/qdrant-1.7.x/) this novel API that lets you constrain the space in which a search is performed, relying only on pure vectors. This is a powerful tool that lets you explore the vector space in a more controlled way. It can be used to find points that are not necessarily closest to the target, but are still relevant to the search.
You can already select which points are available to the search by using payload filters. This by itself is very versatile because it allows us to craft complex filters that show only the points that satisfy their criteria deterministically. However, the payload associated with each point is arbitrary and cannot tell us anything about their position in the vector space. In other words, filtering out irrelevant points can be seen as creating a _mask_ rather than a hyperplane –cutting in between the positive and negative vectors– in the space.
-## Understanding context in discovery search
+## Understanding context
This is where a __vector _context___ can help. We define _context_ as a list of pairs. Each pair is made up of a positive and a negative vector. With a context, we can define hyperplanes within the vector space, which always prefer the positive over the negative vectors. This effectively partitions the space where the search is performed. After the space is partitioned, we then need a _target_ to return the points that are more similar to it.
@@ -99,8 +95,8 @@ Creating complex tastes in a high-dimensional space becomes easier since you can
This way you can give refreshing recommendations, while still being in control by providing positive and negative feedback, or even by trying out different permutations of pairs.
-## Key rakeaways:
+## Key takeaways:
- Discovery search is a powerful tool for controlled exploration in vector spaces.
-Context, positive, and negative vectors guide search parameters and refine results.
+Context, consisting of positive and negative vectors constrain the search space, while a target guides the search.
- Real-world applications include multimodal search, diverse recommendations, and context-driven exploration.
-- Ready to experience the power of Qdrant's Discovery search for yourself? [Try a free demo](https://qdrant.tech/contact-us/) now and unlock the full potential of controlled exploration in vector spaces!
\ No newline at end of file
+- Ready to learn more about the math behind it and how to use it? Check out the [documentation](/documentation/concepts/explore/#discovery-api)
\ No newline at end of file
From f72a036b5be45c1f2b69a520e1c8bf3677f8ec1d Mon Sep 17 00:00:00 2001
From: generall
Date: Sun, 14 Jul 2024 17:02:17 +0000
Subject: [PATCH 13/32] Update GitHub Stars
---
qdrant-landing/content/headless/stats.md | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/qdrant-landing/content/headless/stats.md b/qdrant-landing/content/headless/stats.md
index 3e4c70d35..aa13b4f2e 100644
--- a/qdrant-landing/content/headless/stats.md
+++ b/qdrant-landing/content/headless/stats.md
@@ -1,6 +1,6 @@
---
stats:
- githubStars: 18.8k
- discordMembers: 6.0k
+ githubStars: 18.9k
+ discordMembers: 6.1k
twitterFollowers: 7.5k
---
\ No newline at end of file
From 84417ae5b2292a85cfe9d023ede90607d8bdf90c Mon Sep 17 00:00:00 2001
From: Anush
Date: Tue, 16 Jul 2024 17:38:00 +0530
Subject: [PATCH 14/32] docs: fix links vectors.md
---
.../content/documentation/concepts/vectors.md | 10 +++++-----
1 file changed, 5 insertions(+), 5 deletions(-)
diff --git a/qdrant-landing/content/documentation/concepts/vectors.md b/qdrant-landing/content/documentation/concepts/vectors.md
index 22f8eb1e2..3ea117cfe 100644
--- a/qdrant-landing/content/documentation/concepts/vectors.md
+++ b/qdrant-landing/content/documentation/concepts/vectors.md
@@ -1313,16 +1313,16 @@ await client.CreateCollectionAsync(
## Quantization
Apart from changing the datatype of the original vectors, Qdrant can create quantized representations of vectors alongside the original ones.
-This quantized representation can be used to quickly select candidates for rescoring with the original vectors, or even used directly for search.
+This quantized representation can be used to quickly select candidates for rescoring with the original vectors or even used directly for search.
Quantization is applied in the background, during the optimization process.
-More information about the quantization process can be found in the [Quantization](../guides/quantization/) section.
+More information about the quantization process can be found in the [Quantization](../../guides/quantization/) section.
## Vector Storage
-Depending on the requirements of the application, Qdrant can use one of the data storage options.
-Keep in mind that youu will have to tradeoff between search speed and the size of RAM used.
+Depending on the application's requirements, Qdrant can use one of the data storage options.
+Keep in mind that you will have to tradeoff between search speed and the size of RAM used.
-More information about the storage options can be found in the [Storage](../concepts/storage/#vector-storage) section.
+More information about the storage options can be found in the [Storage](../../concepts/storage/#vector-storage) section.
From a221efb71ca6274a18cbcfa90950715338578bcf Mon Sep 17 00:00:00 2001
From: Anush
Date: Tue, 16 Jul 2024 17:43:02 +0530
Subject: [PATCH 15/32] docs: Updated vectors.md
---
qdrant-landing/content/documentation/concepts/vectors.md | 5 +----
1 file changed, 1 insertion(+), 4 deletions(-)
diff --git a/qdrant-landing/content/documentation/concepts/vectors.md b/qdrant-landing/content/documentation/concepts/vectors.md
index 3ea117cfe..66dc82463 100644
--- a/qdrant-landing/content/documentation/concepts/vectors.md
+++ b/qdrant-landing/content/documentation/concepts/vectors.md
@@ -1322,7 +1322,4 @@ More information about the quantization process can be found in the [Quantizatio
## Vector Storage
-Depending on the application's requirements, Qdrant can use one of the data storage options.
-Keep in mind that you will have to tradeoff between search speed and the size of RAM used.
-
-More information about the storage options can be found in the [Storage](../../concepts/storage/#vector-storage) section.
+Information about the vector storage options with their benefits and trade-offs can be found in the [Storage](../../concepts/storage/#vector-storage) section.
From 0c2b70b9c88d0dad3f454fbbeca6a3196bd821d9 Mon Sep 17 00:00:00 2001
From: davidmyriel
Date: Tue, 16 Jul 2024 05:23:55 -0700
Subject: [PATCH 16/32] fix vectors page links
---
.../content/documentation/concepts/vectors.md | 13 ++++++++-----
1 file changed, 8 insertions(+), 5 deletions(-)
diff --git a/qdrant-landing/content/documentation/concepts/vectors.md b/qdrant-landing/content/documentation/concepts/vectors.md
index 66dc82463..4d540d666 100644
--- a/qdrant-landing/content/documentation/concepts/vectors.md
+++ b/qdrant-landing/content/documentation/concepts/vectors.md
@@ -2,7 +2,7 @@
title: Vectors
weight: 41
aliases:
- - ../vectors
+ - /vectors
---
@@ -19,7 +19,7 @@ If two images are similar, their vectors will be close to each other in the vect
In order to obtain a vector representation of an object, you need to apply a vectorization algorithm to the object.
Usually, this algorithm is a neural network that converts the object into a fixed-size vector.
-The neural network is usually [trained](../articles/metric-learning-tips/) on a pairs or [triplets](../articles/triplet-loss/) of similar and dissimilar objects, so it learns to recognize a specific type of similarity.
+The neural network is usually [trained](/articles/metric-learning-tips/) on a pairs or [triplets](/articles/triplet-loss/) of similar and dissimilar objects, so it learns to recognize a specific type of similarity.
By using this property of vectors, you can explore your data in a number of ways; e.g. by searching for similar objects, clustering objects, and more.
@@ -58,7 +58,7 @@ It looks like this:
```
The majority of neural networks create dense vectors, so you can use them with Qdrant without any additional processing.
-Although compatible with most embedding models out there, Qdrant has been tested with the following [verified embedding providers](../embeddings/).
+Although compatible with most embedding models out there, Qdrant has been tested with the following [verified embedding providers](/documentation/embeddings/).
### Sparse Vectors
@@ -1317,9 +1317,12 @@ This quantized representation can be used to quickly select candidates for resco
Quantization is applied in the background, during the optimization process.
-More information about the quantization process can be found in the [Quantization](../../guides/quantization/) section.
+More information about the quantization process can be found in the [Quantization](/documentation/guides/quantization/) section.
## Vector Storage
-Information about the vector storage options with their benefits and trade-offs can be found in the [Storage](../../concepts/storage/#vector-storage) section.
+Depending on the requirements of the application, Qdrant can use one of the data storage options.
+Keep in mind that you will have to tradeoff between search speed and the size of RAM used.
+
+More information about the storage options can be found in the [Storage](/documentation/concepts/storage/#vector-storage) section.
From df617d2a24f132500751f5fe642bebfe98838f5b Mon Sep 17 00:00:00 2001
From: Anush008
Date: Tue, 16 Jul 2024 22:00:30 +0530
Subject: [PATCH 17/32] docs: Added TS, Py snippets to 1.10 blog
---
qdrant-landing/content/blog/qdrant-1.10.x.md | 184 +++++++++++++++++--
1 file changed, 164 insertions(+), 20 deletions(-)
diff --git a/qdrant-landing/content/blog/qdrant-1.10.x.md b/qdrant-landing/content/blog/qdrant-1.10.x.md
index 2b1c190de..76d45a73c 100644
--- a/qdrant-landing/content/blog/qdrant-1.10.x.md
+++ b/qdrant-landing/content/blog/qdrant-1.10.x.md
@@ -23,7 +23,8 @@ tags:
**Multivector Support:** Native support for late interaction ColBERT is accessible via Query API.
## One Endpoint for All Queries
-**Query API** will consolidate all search APIs into a single request. Previously, you had to work outside of the API to combine different search requests. Now these approaches are reduced to parameters of a single request, so you can avoid merging individual results.
+
+**Query API** will consolidate all search APIs into a single request. Previously, you had to work outside of the API to combine different search requests. Now these approaches are reduced to parameters of a single request, so you can avoid merging individual results.
You can now configure the Query API request with the following parameters:
@@ -59,7 +60,8 @@ POST collections/{collection_name}/points/query
We will be publishing code samples in [docs](/documentation/concepts/hybrid-queries/) and our new [API specification](http://api.qdrant.tech). *If you need additional support with this new method, our [Discord](https://qdrant.to/discord) on-call engineers can help you.*
### Native Hybrid Search Support
-Query API now also natively supports **sparse/dense fusion**. Up to this point, you had to combine the results of sparse and dense searches on your own. This is now sorted on the back-end, and you only have to configure them as basic parameters for Query API.
+
+Query API now also natively supports **sparse/dense fusion**. Up to this point, you had to combine the results of sparse and dense searches on your own. This is now sorted on the back-end, and you only have to configure them as basic parameters for Query API.
```http
POST /collections/{collection_name}/points/query
@@ -84,6 +86,56 @@ POST /collections/{collection_name}/points/query
}
```
+```python
+from qdrant_client import QdrantClient, models
+
+client = QdrantClient(url="http://localhost:6333")
+
+client.query_points(
+ collection_name="{collection_name}",
+ prefetch=[
+ models.Prefetch(
+ query=models.SparseVector(indices=[1, 42], values=[0.22, 0.8]),
+ using="sparse",
+ limit=20,
+ ),
+ models.Prefetch(
+ query=[0.01, 0.45, 0.67],
+ using="dense",
+ limit=20,
+ ),
+ ],
+ query=models.FusionQuery(fusion=models.Fusion.RRF),
+)
+```
+
+```typescript
+import { QdrantClient } from "@qdrant/js-client-rest";
+
+const client = new QdrantClient({ host: "localhost", port: 6333 });
+
+client.query("{collection_name}", {
+ prefetch: [
+ {
+ query: {
+ values: [0.22, 0.8],
+ indices: [1, 42],
+ },
+ using: 'sparse',
+ limit: 20,
+ },
+ {
+ query: [0.01, 0.45, 0.67],
+ using: 'dense',
+ limit: 20,
+ },
+ ],
+ query: {
+ fusion: 'rrf',
+ },
+});
+```
+
```rust
use qdrant_client::Qdrant;
use qdrant_client::qdrant::{Fusion, PrefetchQueryBuilder, Query, QueryPointsBuilder};
@@ -167,7 +219,7 @@ await client.QueryAsync(
);
```
-Query API can now pre-fetch vectors for requests, which means you can run queries sequentially within the same API call. There are a lot of options here, so you will need to define a strategy to merge these requests using new parameters. For example, you can now include **rescoring within Hybrid Search**, which can open the door to strategies like iterative refinement via matryoshka embeddings.
+Query API can now pre-fetch vectors for requests, which means you can run queries sequentially within the same API call. There are a lot of options here, so you will need to define a strategy to merge these requests using new parameters. For example, you can now include **rescoring within Hybrid Search**, which can open the door to strategies like iterative refinement via matryoshka embeddings.
*To learn more about this, read the [Query API documentation](/documentation/concepts/search/#query-api).*
@@ -214,6 +266,20 @@ client.create_collection(
)
```
+```typescript
+import { QdrantClient } from "@qdrant/js-client-rest";
+
+const client = new QdrantClient({ host: "localhost", port: 6333 });
+
+client.createCollection("{collection_name}", {
+ sparse_vectors: {
+ "text": {
+ modifier: "idf"
+ }
+ }
+});
+```
+
```rust
use qdrant_client::Qdrant;
use qdrant_client::qdrant::{CreateCollectionBuilder, sparse_vectors_config::SparseVectorsConfigBuilder, Modifier, SparseVectorParamsBuilder};
@@ -282,12 +348,13 @@ In practical terms, the BM42 method addresses the tokenization issues and comput
**You can expect BM42 to excel in scalable RAG-based scenarios where short texts are more common.** Document inference speed is much higher with BM42, which is critical for large-scale applications such as search engines, recommendation systems, and real-time decision-making systems.
-## Multivector Support
-We are adding native support for multivector search that is compatible, e.g., with the late-interaction [ColBERT](https://github.com/stanford-futuredata/ColBERT) model. If you are working with high-dimensional similarity searches, **ColBERT is highly recommended as a reranking step in the Universal Query search.** You will experience better quality vector retrieval since ColBERT’s approach allows for deeper semantic understanding.
+## Multivector Support
-This model retains contextual information during query-document interaction, leading to better relevance scoring. In terms of efficiency and scalability benefits, documents and queries will be encoded separately, which gives an opportunity for pre-computation and storage of document embeddings for faster retrieval.
+We are adding native support for multivector search that is compatible, e.g., with the late-interaction [ColBERT](https://github.com/stanford-futuredata/ColBERT) model. If you are working with high-dimensional similarity searches, **ColBERT is highly recommended as a reranking step in the Universal Query search.** You will experience better quality vector retrieval since ColBERT’s approach allows for deeper semantic understanding.
-**Note:** *This feature supports all the original quantization compression methods, just the same as the regular search method.*
+This model retains contextual information during query-document interaction, leading to better relevance scoring. In terms of efficiency and scalability benefits, documents and queries will be encoded separately, which gives an opportunity for pre-computation and storage of document embeddings for faster retrieval.
+
+**Note:** *This feature supports all the original quantization compression methods, just the same as the regular search method.*
**Run a query with ColBERT vectors:**
@@ -316,6 +383,55 @@ POST /collections/{collection_name}/points/query
}
```
+```python
+from qdrant_client import QdrantClient, models
+
+client = QdrantClient(url="http://localhost:6333")
+
+client.query_points(
+ collection_name="{collection_name}",
+ prefetch=models.Prefetch(
+ prefetch=models.Prefetch(query=[1, 23, 45, 67], using="mrl_byte", limit=1000),
+ query=[0.01, 0.45, 0.67],
+ using="full",
+ limit=100,
+ ),
+ query=[
+ [0.1, 0.2],
+ [0.2, 0.1],
+ [0.8, 0.9],
+ ],
+ using="colbert",
+ limit=10,
+)
+```
+
+```typescript
+import { QdrantClient } from "@qdrant/js-client-rest";
+
+const client = new QdrantClient({ host: "localhost", port: 6333 });
+
+client.query("{collection_name}", {
+ prefetch: {
+ prefetch: {
+ query: [1, 23, 45, 67],
+ using: 'mrl_byte',
+ limit: 1000
+ },
+ query: [0.01, 0.45, 0.67],
+ using: 'full',
+ limit: 100,
+ },
+ query: [
+ [0.1, 0.2],
+ [0.2, 0.1],
+ [0.8, 0.9],
+ ],
+ using: 'colbert',
+ limit: 10,
+});
+```
+
```rust
use qdrant_client::Qdrant;
use qdrant_client::qdrant::{PrefetchQueryBuilder, Query, QueryPointsBuilder};
@@ -327,7 +443,7 @@ client.query(
.add_prefetch(PrefetchQueryBuilder::default()
.add_prefetch(PrefetchQueryBuilder::default()
.query(Query::new_nearest(vec![1.0, 23.0, 45.0, 67.0]))
- .using("mlr_byte")
+ .using("mrl_byte")
.limit(1000u64)
)
.query(Query::new_nearest(vec![0.01, 0.45, 0.67]))
@@ -363,7 +479,7 @@ client
PrefetchQuery.newBuilder()
.addPrefetch(
PrefetchQuery.newBuilder()
- .setQuery(nearest(1, 23, 45, 67)) // <------------- small byte vector
+ .setQuery(nearest(1, 23, 45, 67)) // <------------- small byte vector
.setUsing("mrl_byte")
.setLimit(1000)
.build())
@@ -374,9 +490,9 @@ client
.setQuery(
nearest(
new float[][] {
- {0.1f, 0.2f}, // <─┐
- {0.2f, 0.1f}, // < ├─ multi-vector
- {0.8f, 0.9f} // < ┘
+ {0.1f, 0.2f}, // <─┐
+ {0.2f, 0.1f}, // < ├─ multi-vector
+ {0.8f, 0.9f} // < ┘
}))
.setUsing("colbert")
.setLimit(10)
@@ -418,15 +534,15 @@ await client.QueryAsync(
);
```
-**Note:** *The multivector feature is not only useful for ColBERT; it can also be used in other ways.*
-For instance, in e-commerce, you can use multi-vector to store multiple images of the same item. This serves as an alternative to the [group-by](/documentation/concepts/search/#grouping-api) method.
+**Note:** *The multivector feature is not only useful for ColBERT; it can also be used in other ways.*
+For instance, in e-commerce, you can use multi-vector to store multiple images of the same item. This serves as an alternative to the [group-by](/documentation/concepts/search/#grouping-api) method.
## Sparse Vectors Compression
-In version 1.9, we introduced the `uint8` [vector datatype](/documentation/concepts/vectors/#datatypes) for sparse vectors, in order to support pre-quantized embeddings from companies like JinaAI and Cohere.
-This time, we are introducing a new datatype **for both sparse and dense vectors**, as well as a different way of **storing** these vectors.
+In version 1.9, we introduced the `uint8` [vector datatype](/documentation/concepts/vectors/#datatypes) for sparse vectors, in order to support pre-quantized embeddings from companies like JinaAI and Cohere.
+This time, we are introducing a new datatype **for both sparse and dense vectors**, as well as a different way of **storing** these vectors.
-**Datatype:** Sparse and dense vectors were previously represented in larger `float32` values, but now they can be turned to the `float16`. `float16` vectors have a lower precision compared to `float32`, which means that there is less numerical accuracy in the vector values - but this is negligible for practical use cases.
+**Datatype:** Sparse and dense vectors were previously represented in larger `float32` values, but now they can be turned to the `float16`. `float16` vectors have a lower precision compared to `float32`, which means that there is less numerical accuracy in the vector values - but this is negligible for practical use cases.
These vectors will use half the memory of regular vectors, which can significantly reduce the footprint of large vector datasets. Operations can be faster due to reduced memory bandwidth requirements and better cache utilization. This can lead to faster vector search operations, especially in memory-bound scenarios.
@@ -443,6 +559,33 @@ PUT /collections/{collection_name}
}
```
+```python
+from qdrant_client import QdrantClient, models
+
+client = QdrantClient(url="http://localhost:6333")
+
+client.create_collection(
+ "{collection_name}",
+ vectors_config=models.VectorParams(
+ size=1024, distance=models.Distance.COSINE, datatype=models.Datatype.FLOAT16
+ ),
+)
+```
+
+```typescript
+import { QdrantClient } from "@qdrant/js-client-rest";
+
+const client = new QdrantClient({ host: "localhost", port: 6333 });
+
+client.createCollection("{collection_name}", {
+ vectors: {
+ size: 1024,
+ distance: "Cosine",
+ datatype: "float16"
+ }
+});
+```
+
```java
import io.qdrant.client.QdrantClient;
import io.qdrant.client.QdrantGrpcClient;
@@ -500,7 +643,7 @@ await client.CreateCollectionAsync(
);
```
-**Storage:** On the backend, we implemented bit packing to minimize the bits needed to store data, crucial for handling sparse vectors in applications like machine learning and data compression. For sparse vectors with mostly zeros, this focuses on storing only the indices and values of non-zero elements.
+**Storage:** On the backend, we implemented bit packing to minimize the bits needed to store data, crucial for handling sparse vectors in applications like machine learning and data compression. For sparse vectors with mostly zeros, this focuses on storing only the indices and values of non-zero elements.
You will benefit from a more compact storage and higher processing efficiency. This can also lead to reduced dataset sizes for faster processing and lower storage costs in data compression.
@@ -525,6 +668,7 @@ documentation, making it easier to navigate and find the information you need.