Fixing links again (#697)

* use relative links instead of absolute

* add trailing slashes to avoid 301 redirect

* add trailing slashes to avoid 301 redirect

* make link checker unhappy with local redirects

* test if ci fails (should fail)

* rollback: test if ci fails (should fail)
This commit is contained in:
Andrey Vasnetsov
2024-03-07 20:31:05 +01:00
committed by GitHub
parent 2285373bc8
commit cae8456ce1
72 changed files with 142 additions and 140 deletions
+4 -2
View File
@@ -17,12 +17,14 @@ jobs:
with: with:
hugo-version: "latest" hugo-version: "latest"
- name: Run hugo - name: Run hugo
run: cd qdrant-landing && hugo -b '' run: |
cd qdrant-landing && hugo --gc -b 'http://localhost:1313' && hugo serve &
sleep 5 # wait for server to start
- name: Link Checker - name: Link Checker
id: lychee id: lychee
uses: lycheeverse/lychee-action@v1.8.0 uses: lycheeverse/lychee-action@v1.8.0
with: with:
args: --offline --base qdrant-landing/public qdrant-landing/public args: --max-redirects 0 --exclude '.*' --include 'http://localhost:1313/.*' qdrant-landing/public/
fail: true fail: true
env: env:
GITHUB_TOKEN: ${{secrets.GITHUB_TOKEN}} GITHUB_TOKEN: ${{secrets.GITHUB_TOKEN}}
+1 -1
View File
@@ -52,7 +52,7 @@ disableKinds = ["taxonomy", "term"]
quick_start = "/documentation/quick-start/" quick_start = "/documentation/quick-start/"
tutorial = "/articles/neural-search-tutorial/" tutorial = "/articles/neural-search-tutorial/"
benchmarks = "/benchmarks/" benchmarks = "/benchmarks/"
demos = "/demo" demos = "/demo/"
cloud = "https://qdrant.to/cloud" cloud = "https://qdrant.to/cloud"
linkedin = "https://qdrant.to/linkedin" linkedin = "https://qdrant.to/linkedin"
@@ -34,7 +34,7 @@ In this post, we discuss:
- Implications of these findings for real-world applications - Implications of these findings for real-world applications
- Best practices for leveraging Binary Quantization to enhance OpenAI embeddings - Best practices for leveraging Binary Quantization to enhance OpenAI embeddings
If you're new to Binary Quantization, consider reading our article which walks you through the concept and [how to use it with Qdrant](https://qdrant.tech/articles/binary-quantization/) If you're new to Binary Quantization, consider reading our article which walks you through the concept and [how to use it with Qdrant](/articles/binary-quantization/)
You can also try out these techniques as described in [Binary Quantization OpenAI](https://github.com/qdrant/examples/blob/openai-3/binary-quantization-openai/README.md), which includes Jupyter notebooks. You can also try out these techniques as described in [Binary Quantization OpenAI](https://github.com/qdrant/examples/blob/openai-3/binary-quantization-openai/README.md), which includes Jupyter notebooks.
@@ -215,6 +215,6 @@ We recommend the following best practices for leveraging Binary Quantization to
Binary quantization is exceptional if you need to work with large volumes of data under high recall expectations. You can try this feature either by spinning up a [Qdrant container image](https://hub.docker.com/r/qdrant/qdrant) locally or, having us create one for you through a [free account](https://cloud.qdrant.io/login) in our cloud hosted service. Binary quantization is exceptional if you need to work with large volumes of data under high recall expectations. You can try this feature either by spinning up a [Qdrant container image](https://hub.docker.com/r/qdrant/qdrant) locally or, having us create one for you through a [free account](https://cloud.qdrant.io/login) in our cloud hosted service.
The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials/bulk-upload/) to your Qdrant instance as well as [more quantization methods](https://qdrant.tech/documentation/guides/quantization/). The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials/bulk-upload/) to your Qdrant instance as well as [more quantization methods](/documentation/guides/quantization/).
Want to discuss these findings and learn more about Binary Quantization? [Join our Discord community.](https://discord.gg/qdrant) Want to discuss these findings and learn more about Binary Quantization? [Join our Discord community.](https://discord.gg/qdrant)
@@ -228,6 +228,6 @@ If you determine that binary quantization is appropriate for your datasets and q
Binary quantization is exceptional if you need to work with large volumes of data under high recall expectations. You can try this feature either by spinning up a [Qdrant container image](https://hub.docker.com/r/qdrant/qdrant) locally or, having us create one for you through a [free account](https://cloud.qdrant.io/login) in our cloud hosted service. Binary quantization is exceptional if you need to work with large volumes of data under high recall expectations. You can try this feature either by spinning up a [Qdrant container image](https://hub.docker.com/r/qdrant/qdrant) locally or, having us create one for you through a [free account](https://cloud.qdrant.io/login) in our cloud hosted service.
The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials/bulk-upload/) to your Qdrant instance as well as [more quantization methods](https://qdrant.tech/documentation/guides/quantization/). The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials/bulk-upload/) to your Qdrant instance as well as [more quantization methods](/documentation/guides/quantization/).
If you have any feedback, drop us a note on Twitter or LinkedIn to tell us about your results. [Join our lively Discord Server](https://discord.gg/Qy6HCJK9Dc) if you want to discuss BQ with like-minded people! If you have any feedback, drop us a note on Twitter or LinkedIn to tell us about your results. [Join our lively Discord Server](https://discord.gg/Qy6HCJK9Dc) if you want to discuss BQ with like-minded people!
@@ -107,7 +107,7 @@ Diversity:
{{< figure src=https://storage.googleapis.com/demo-dataset-quality-public/article/diversity_transparent.png caption="Diversity search" >}} {{< figure src=https://storage.googleapis.com/demo-dataset-quality-public/article/diversity_transparent.png caption="Diversity search" >}}
Diversity search utilizes the very same embeddings, and you can reuse them. Diversity search utilizes the very same embeddings, and you can reuse them.
If your data is huge and does not fit into memory, vector search engines like [Qdrant](https://qdrant.tech/) might be helpful. If your data is huge and does not fit into memory, vector search engines like [Qdrant](https://github.com/qdrant/qdrant) might be helpful.
Although the described methods can be used independently. But they are simple to combine and improve detection capabilities. Although the described methods can be used independently. But they are simple to combine and improve detection capabilities.
If the quality remains insufficient, you can fine-tune the models using a similarity learning approach (e.g. with [Quaterion](https://quaterion.qdrant.tech) both to provide a better representation of your data and pull apart dissimilar objects in space. If the quality remains insufficient, you can fine-tune the models using a similarity learning approach (e.g. with [Quaterion](https://quaterion.qdrant.tech) both to provide a better representation of your data and pull apart dissimilar objects in space.
@@ -638,7 +638,7 @@ if __name__ == "__main__":
``` ```
We stored our collection of answer embeddings in memory and perform search directly in Python. We stored our collection of answer embeddings in memory and perform search directly in Python.
For production purposes, it's better to use some sort of vector search engine like [Qdrant](https://qdrant.tech/). For production purposes, it's better to use some sort of vector search engine like [Qdrant](https://github.com/qdrant/qdrant).
It provides durability, speed boost, and a bunch of other features. It provides durability, speed boost, and a bunch of other features.
So far, we've implemented a whole training process, prepared model for serving and even applied a So far, we've implemented a whole training process, prepared model for serving and even applied a
+2 -2
View File
@@ -139,7 +139,7 @@ If anything changes, you'll see a new version number pop up, like going from 0.0
## Usage with Qdrant ## Usage with Qdrant
Qdrant is a Vector Store, offering a comprehensive, efficient, and scalable solution for modern machine learning and AI applications. Whether you are dealing with billions of data points, require a low latency performant vector solution, or specialized quantization methods – [Qdrant is engineered](https://qdrant.tech/documentation/overview/) to meet those demands head-on. Qdrant is a Vector Store, offering a comprehensive, efficient, and scalable solution for modern machine learning and AI applications. Whether you are dealing with billions of data points, require a low latency performant vector solution, or specialized quantization methods – [Qdrant is engineered](/documentation/overview/) to meet those demands head-on.
The fusion of FastEmbed with Qdrant’s vector store capabilities enables a transparent workflow for seamless embedding generation, storage, and retrieval. This simplifies the API design — while still giving you the flexibility to make significant changes e.g. you can use FastEmbed to make your own embedding other than the DefaultEmbedding and use that with Qdrant. The fusion of FastEmbed with Qdrant’s vector store capabilities enables a transparent workflow for seamless embedding generation, storage, and retrieval. This simplifies the API design — while still giving you the flexibility to make significant changes e.g. you can use FastEmbed to make your own embedding other than the DefaultEmbedding and use that with Qdrant.
@@ -237,7 +237,7 @@ If you're curious about how FastEmbed and Qdrant can make your search tasks a br
1. **Cloud**: Get started with a free plan on the [Qdrant Cloud](https://qdrant.to/cloud?utm_source=qdrant&utm_medium=website&utm_campaign=fastembed&utm_content=article). 1. **Cloud**: Get started with a free plan on the [Qdrant Cloud](https://qdrant.to/cloud?utm_source=qdrant&utm_medium=website&utm_campaign=fastembed&utm_content=article).
2. **Docker Container**: If you're the DIY type, you can set everything up on your own machine. Here's a quick guide to help you out: [Quick Start with Docker](https://qdrant.tech/documentation/quick-start/?utm_source=qdrant&utm_medium=website&utm_campaign=fastembed&utm_content=article). 2. **Docker Container**: If you're the DIY type, you can set everything up on your own machine. Here's a quick guide to help you out: [Quick Start with Docker](/documentation/quick-start/?utm_source=qdrant&utm_medium=website&utm_campaign=fastembed&utm_content=article).
So, go ahead, take it for a test drive. We're excited to hear what you think! So, go ahead, take it for a test drive. We're excited to hear what you think!
@@ -24,7 +24,7 @@ so you can easily deploy it on your own and play with it. If you prefer to dive
Otherwise, read on to learn more about the demo and how it works! Otherwise, read on to learn more about the demo and how it works!
In general, our application consists of three parts: a [FastAPI](https://fastapi.tiangolo.com/) backend, a [React](https://react.dev/) frontend, and In general, our application consists of three parts: a [FastAPI](https://fastapi.tiangolo.com/) backend, a [React](https://react.dev/) frontend, and
a [Qdrant](https://qdrant.tech/) instance. The architecture diagram below shows how these components interact with each other: a [Qdrant](/) instance. The architecture diagram below shows how these components interact with each other:
![Archtecture diagram](/articles_data/food-discovery-demo/architecture-diagram.png) ![Archtecture diagram](/articles_data/food-discovery-demo/architecture-diagram.png)
@@ -91,7 +91,7 @@ in the vector space.
![Random points selection](/articles_data/food-discovery-demo/textual-search.png) ![Random points selection](/articles_data/food-discovery-demo/textual-search.png)
This is implemented as [a group search query to Qdrant](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L44). This is implemented as [a group search query to Qdrant](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L44).
We didn't use a simple search, but performed grouping by the restaurant to get more diverse results. [Search groups](https://qdrant.tech/documentation/concepts/search/#search-groups) We didn't use a simple search, but performed grouping by the restaurant to get more diverse results. [Search groups](/documentation/concepts/search/#search-groups)
is a mechanism similar to `GROUP BY` clause in SQL, and it's useful when you want to get a specific number of result per group (in our case just one). is a mechanism similar to `GROUP BY` clause in SQL, and it's useful when you want to get a specific number of result per group (in our case just one).
```python ```python
@@ -120,7 +120,7 @@ and the demo will update the search results accordingly.
#### Negative feedback only #### Negative feedback only
Qdrant [Recommendation API](https://qdrant.tech/documentation/concepts/search/#recommendation-api) needs at least one positive example to work. However, in our demo Qdrant [Recommendation API](/documentation/concepts/search/#recommendation-api) needs at least one positive example to work. However, in our demo
we want to be able to provide only negative examples. This is because we want to be able to say “I don’t like this dish” without having to like anything first. we want to be able to provide only negative examples. This is because we want to be able to say “I don’t like this dish” without having to like anything first.
To achieve this, we use a trick. We negate the vectors of the disliked dishes and use their mean as a query. This way, the disliked dishes will be pushed away To achieve this, we use a trick. We negate the vectors of the disliked dishes and use their mean as a query. This way, the disliked dishes will be pushed away
from the search results. **This works because the cosine distance is based on the angle between two vectors, and the angle between a vector and its negation is 180 degrees.** from the search results. **This works because the cosine distance is based on the angle between two vectors, and the angle between a vector and its negation is 180 degrees.**
@@ -128,8 +128,8 @@ from the search results. **This works because the cosine distance is based on th
![CLIP model](/articles_data/food-discovery-demo/negated-vector.png) ![CLIP model](/articles_data/food-discovery-demo/negated-vector.png)
Food Discovery Demo [implements that trick](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L122) Food Discovery Demo [implements that trick](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L122)
by calling Qdrant twice. Initially, we use the [Scroll API](https://qdrant.tech/documentation/concepts/points/#scroll-points) to find disliked items, by calling Qdrant twice. Initially, we use the [Scroll API](/documentation/concepts/points/#scroll-points) to find disliked items,
and then calculate a negated mean of all their vectors. That allows using the [Search Groups API](https://qdrant.tech/documentation/concepts/search/#search-groups) and then calculate a negated mean of all their vectors. That allows using the [Search Groups API](/documentation/concepts/search/#search-groups)
to find the nearest neighbors of the negated mean vector. to find the nearest neighbors of the negated mean vector.
```python ```python
@@ -162,7 +162,7 @@ response = client.search_groups(
#### Positive and negative feedback #### Positive and negative feedback
Since the [Recommendation API](https://qdrant.tech/documentation/concepts/search/#recommendation-api) requires at least one positive example, we can use it only when Since the [Recommendation API](/documentation/concepts/search/#recommendation-api) requires at least one positive example, we can use it only when
the user has liked at least one dish. We could theoretically use the same trick as above and negate the disliked dishes, but it would be a bit weird, as Qdrant has the user has liked at least one dish. We could theoretically use the same trick as above and negate the disliked dishes, but it would be a bit weird, as Qdrant has
that feature already built-in, and we can call it just once to do the job. It's always better to perform the search server-side. Thus, in this case [we just call that feature already built-in, and we can call it just once to do the job. It's always better to perform the search server-side. Thus, in this case [we just call
the Qdrant server with a list of positive and negative examples](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L166), the Qdrant server with a list of positive and negative examples](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L166),
@@ -184,7 +184,7 @@ From the user perspective nothing changes comparing to the previous case.
Last but not least, location plays an important role in the food discovery process. You are definitely looking for something you can find nearby, not on the other Last but not least, location plays an important role in the food discovery process. You are definitely looking for something you can find nearby, not on the other
side of the globe. Therefore, your current location can be toggled as a filtering condition. You can enable it by clicking on “Find near me” icon side of the globe. Therefore, your current location can be toggled as a filtering condition. You can enable it by clicking on “Find near me” icon
in the top right. This way you can find the best pizza in your neighborhood, not in the whole world. Qdrant [geo radius filter](https://qdrant.tech/documentation/concepts/filtering/#geo-radius) is a perfect choice for this. It lets you in the top right. This way you can find the best pizza in your neighborhood, not in the whole world. Qdrant [geo radius filter](/documentation/concepts/filtering/#geo-radius) is a perfect choice for this. It lets you
filter the results by distance from a given point. filter the results by distance from a given point.
```python ```python
@@ -207,7 +207,7 @@ query_filter = models.Filter(
) )
``` ```
Such a filter needs [a payload index](https://qdrant.tech/documentation/concepts/indexing/#payload-index) to work efficiently, and it was created on a collection Such a filter needs [a payload index](/documentation/concepts/indexing/#payload-index) to work efficiently, and it was created on a collection
we used to create the snapshot. When you import it into your instance, the index will be already there. we used to create the snapshot. When you import it into your instance, the index will be already there.
## Using the demo ## Using the demo
@@ -75,4 +75,4 @@ Being selected for Google Summer of Code 2023 and collaborating with Arnaud and
Without a doubt, I'm eager to continue growing alongside this community and contribute to new features and enhancements that elevate the product. I've also become an advocate for Qdrant, introducing this project to numerous coworkers and friends in the tech industry. I'm excited to witness new users and contributors emerge from within my own network! Without a doubt, I'm eager to continue growing alongside this community and contribute to new features and enhancements that elevate the product. I've also become an advocate for Qdrant, introducing this project to numerous coworkers and friends in the tech industry. I'm excited to witness new users and contributors emerge from within my own network!
If you want to try out my work, read the [documentation](https://qdrant.tech/documentation/concepts/filtering/#geo-polygon) and then, either sign up for a free [cloud account](https://cloud.qdrant.io) or download the [Docker image](https://hub.docker.com/r/qdrant/qdrant). I look forward to seeing how people are using my work in their own applications! If you want to try out my work, read the [documentation](/documentation/concepts/filtering/#geo-polygon) and then, either sign up for a free [cloud account](https://cloud.qdrant.io) or download the [Docker image](https://hub.docker.com/r/qdrant/qdrant). I look forward to seeing how people are using my work in their own applications!
@@ -14,7 +14,7 @@ date: 2023-02-15T10:48:00.000Z
There is not a single definition of hybrid search. Actually, if we use more than one search algorithm, it There is not a single definition of hybrid search. Actually, if we use more than one search algorithm, it
might be described as some sort of hybrid. Some of the most popular definitions are: might be described as some sort of hybrid. Some of the most popular definitions are:
1. A combination of vector search with [attribute filtering](https://qdrant.tech/documentation/filtering/). 1. A combination of vector search with [attribute filtering](/documentation/filtering/).
We won't dive much into details, as we like to call it just filtered vector search. We won't dive much into details, as we like to call it just filtered vector search.
2. Vector search with keyword-based search. This one is covered in this article. 2. Vector search with keyword-based search. This one is covered in this article.
3. A mix of dense and sparse vectors. That strategy will be covered in the upcoming article. 3. A mix of dense and sparse vectors. That strategy will be covered in the upcoming article.
+2 -2
View File
@@ -98,8 +98,8 @@ switching overhead plus the wait time until the disk IO is finished. Ultimately,
this works very well with the asynchronous nature of Qdrant's core. this works very well with the asynchronous nature of Qdrant's core.
One of the great optimizations Qdrant offers is quantization (either One of the great optimizations Qdrant offers is quantization (either
[scalar](https://qdrant.tech/articles/scalar-quantization/) or [scalar](/articles/scalar-quantization/) or
[product](https://qdrant.tech/articles/product-quantization/)-based). [product](/articles/product-quantization/)-based).
However unless the collection resides fully in memory, this optimization However unless the collection resides fully in memory, this optimization
method generates significant disk IO, so it is a prime candidate for possible method generates significant disk IO, so it is a prime candidate for possible
improvements. improvements.
@@ -16,12 +16,12 @@ keywords:
- vector database - vector database
--- ---
We are seeing the topics of [multitenancy](https://qdrant.tech/documentation/guides/multiple-partitions/) and [distributed deployment](https://qdrant.tech/documentation/guides/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup. We are seeing the topics of [multitenancy](/documentation/guides/multiple-partitions/) and [distributed deployment](/documentation/guides/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup.
Whether you are building a bank fraud-detection system, RAG for e-commerce, or services for the federal government - you will need to leverage a multitenant architecture to scale your product. Whether you are building a bank fraud-detection system, RAG for e-commerce, or services for the federal government - you will need to leverage a multitenant architecture to scale your product.
In the world of SaaS and enterprise apps, this setup is the norm. It will considerably increase your application's performance and lower your hosting costs. In the world of SaaS and enterprise apps, this setup is the norm. It will considerably increase your application's performance and lower your hosting costs.
We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](https://qdrant.tech/documentation/guides/multiple-partitions/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](https://qdrant.tech/documentation/guides/distributed_deployment/#user-defined-sharding). We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](/documentation/guides/multiple-partitions/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](/documentation/guides/distributed_deployment/#user-defined-sharding).
Combining these two will result in an efficiently-partitioned architecture that further leverages the convenience of a single Qdrant cluster. This article will briefly explain the benefits and show how you can get started using both features. Combining these two will result in an efficiently-partitioned architecture that further leverages the convenience of a single Qdrant cluster. This article will briefly explain the benefits and show how you can get started using both features.
@@ -36,7 +36,7 @@ Qdrant is built to excel in a single collection with a vast number of tenants. Y
## Sharding your database ## Sharding your database
With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](https://qdrant.tech/documentation/guides/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node. With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](/documentation/guides/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node.
During vector search, your operations will be able to hit only the subset of shards they actually need. In massive-scale deployments, __this can significantly improve the performance of operations that do not require the whole collection to be scanned__. During vector search, your operations will be able to hit only the subset of shards they actually need. In massive-scale deployments, __this can significantly improve the performance of operations that do not require the whole collection to be scanned__.
@@ -44,7 +44,7 @@ This works in the other direction as well. Whenever you search for something, yo
### Common use cases ### Common use cases
A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](https://qdrant.tech/documentation/guides/distributed_deployment/#moving-shards). A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](/documentation/guides/distributed_deployment/#moving-shards).
**Figure 2:** Users can both upsert and query shards that are relevant to them, all within the same collection. Regional sharding can help avoid cross-continental traffic. **Figure 2:** Users can both upsert and query shards that are relevant to them, all within the same collection. Regional sharding can help avoid cross-continental traffic.
![Qdrant Multitenancy](/articles_data/multitenancy/shards.png) ![Qdrant Multitenancy](/articles_data/multitenancy/shards.png)
@@ -74,7 +74,7 @@ client.create_shard_key("{tenant_data}", "germany")
``` ```
In this example, your cluster is divided between Germany and Canada. Canadian and German law differ when it comes to international data transfer. Let's say you are creating a RAG application that supports the healthcare industry. Your Canadian customer data will have to be clearly separated for compliance purposes from your German customer. In this example, your cluster is divided between Germany and Canada. Canadian and German law differ when it comes to international data transfer. Let's say you are creating a RAG application that supports the healthcare industry. Your Canadian customer data will have to be clearly separated for compliance purposes from your German customer.
Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](https://qdrant.tech/documentation/guides/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech). Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](/documentation/guides/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech).
## Configure a multitenant setup for users ## Configure a multitenant setup for users
@@ -177,7 +177,7 @@ client.create_payload_index(
## Next steps ## Next steps
Qdrant is ready to support a massive-scale architecture for your machine learning project. If you want to see whether our vector database is right for you, try the [quickstart tutorial](https://qdrant.tech/documentation/quick-start/) or read our [docs and tutorials](https://qdrant.tech/documentation/). Qdrant is ready to support a massive-scale architecture for your machine learning project. If you want to see whether our vector database is right for you, try the [quickstart tutorial](/documentation/quick-start/) or read our [docs and tutorials](/documentation/).
To spin up a free instance of Qdrant, sign up for [Qdrant Cloud](https://qdrant.to/cloud) - no strings attached. To spin up a free instance of Qdrant, sign up for [Qdrant Cloud](https://qdrant.to/cloud) - no strings attached.
@@ -17,15 +17,15 @@ does exist, and recommendation systems are a great example. Recommendations migh
to find items close to positive and far from negative examples. This use of vector databases has many applications, including to find items close to positive and far from negative examples. This use of vector databases has many applications, including
recommendation systems for e-commerce, content, or even dating apps. recommendation systems for e-commerce, content, or even dating apps.
Qdrant has provided the [Recommendation API](https://qdrant.tech/documentation/concepts/search/#recommendation-api) for a while, and with the latest release, [Qdrant 1.6](https://github.com/qdrant/qdrant/releases/tag/v1.6.0), Qdrant has provided the [Recommendation API](/documentation/concepts/search/#recommendation-api) for a while, and with the latest release, [Qdrant 1.6](https://github.com/qdrant/qdrant/releases/tag/v1.6.0),
we're glad to give you more flexibility and control over the Recommendation API. we're glad to give you more flexibility and control over the Recommendation API.
Here, we'll discuss some internals and show how they may be used in practice. Here, we'll discuss some internals and show how they may be used in practice.
### Recap of the old recommendations API ### Recap of the old recommendations API
The previous [Recommendation API](https://qdrant.tech/documentation/concepts/search/#recommendation-api) in Qdrant came with some limitations. First of all, it was required to pass vector IDs for The previous [Recommendation API](/documentation/concepts/search/#recommendation-api) in Qdrant came with some limitations. First of all, it was required to pass vector IDs for
both positive and negative example points. If you wanted to use vector embeddings directly, you had to either create a new point both positive and negative example points. If you wanted to use vector embeddings directly, you had to either create a new point
in a collection or mimic the behaviour of the Recommendation API by using the [Search API](https://qdrant.tech/documentation/concepts/search/#search-api). in a collection or mimic the behaviour of the Recommendation API by using the [Search API](/documentation/concepts/search/#search-api).
Moreover, in the previous releases of Qdrant, you were always asked to provide at least one positive example. This requirement Moreover, in the previous releases of Qdrant, you were always asked to provide at least one positive example. This requirement
was based on the algorithm used to combine multiple samples into a single query vector. It was a simple, yet effective approach. was based on the algorithm used to combine multiple samples into a single query vector. It was a simple, yet effective approach.
However, if the only information you had was that your user dislikes some items, you couldn't use it directly. However, if the only information you had was that your user dislikes some items, you couldn't use it directly.
@@ -61,7 +61,7 @@ negative examples.
## HNSW ANN example and strategy ## HNSW ANN example and strategy
Let’s start with an example to help you understand the [HNSW graph](https://qdrant.tech/articles/filtrable-hnsw/). Assume you want Let’s start with an example to help you understand the [HNSW graph](/articles/filtrable-hnsw/). Assume you want
to travel to a small city on another continent: to travel to a small city on another continent:
1. You start from your hometown and take a bus to the local airport. 1. You start from your hometown and take a bus to the local airport.
@@ -142,11 +142,11 @@ significantly lower. However, if the best negative score is higher than the best
further away from the negatives. That procedure effectively **pulls the traversal procedure away from the negative examples**. further away from the negatives. That procedure effectively **pulls the traversal procedure away from the negative examples**.
If you want to know more about the internals of HNSW, you can check out the article about the If you want to know more about the internals of HNSW, you can check out the article about the
[Filtrable HNSW](https://qdrant.tech/articles/filtrable-hnsw/) that covers the topic thoroughly. [Filtrable HNSW](/articles/filtrable-hnsw/) that covers the topic thoroughly.
## Food Discovery demo ## Food Discovery demo
Our [Food Discovery demo](https://qdrant.tech/articles/food-discovery-demo/) is an application built on top of the new [Recommendation API](https://qdrant.tech/documentation/concepts/search/#recommendation-api). Our [Food Discovery demo](/articles/food-discovery-demo/) is an application built on top of the new [Recommendation API](/documentation/concepts/search/#recommendation-api).
It allows you to find a meal based on liked and disliked photos. There are some updates, enabled by the new Qdrant release: It allows you to find a meal based on liked and disliked photos. There are some updates, enabled by the new Qdrant release:
* **Ability to include multiple textual queries in the recommendation request.** Previously, we only allowed passing a single * **Ability to include multiple textual queries in the recommendation request.** Previously, we only allowed passing a single
@@ -100,7 +100,7 @@ Product Quantization comes with a cost - there are some additional operations to
that the performance might be reduced. However, memory usage might be reduced drastically as that the performance might be reduced. However, memory usage might be reduced drastically as
well. As usual, we did some benchmarks to give you a brief understanding of what you may expect. well. As usual, we did some benchmarks to give you a brief understanding of what you may expect.
Again, we reused the same pipeline as in [the other benchmarks we published](/benchmarks). We Again, we reused the same pipeline as in [the other benchmarks we published](/benchmarks/). We
selected [Arxiv-titles-384-angular-no-filters](https://github.com/qdrant/ann-filtering-benchmark-datasets) selected [Arxiv-titles-384-angular-no-filters](https://github.com/qdrant/ann-filtering-benchmark-datasets)
and [Glove-100](https://github.com/erikbern/ann-benchmarks/) datasets to measure the impact and [Glove-100](https://github.com/erikbern/ann-benchmarks/) datasets to measure the impact
of Product Quantization on precision and time. Both experiments were launched with $ EF = 128 $. of Product Quantization on precision and time. Both experiments were launched with $ EF = 128 $.
@@ -35,7 +35,7 @@ feedback, and we tried to include the features most requested by our community.
The primary focus of Qdrant was always performance. That's why we built it in Rust, but we were The primary focus of Qdrant was always performance. That's why we built it in Rust, but we were
always concerned about making vector search affordable. From the very beginning, Qdrant offered always concerned about making vector search affordable. From the very beginning, Qdrant offered
support for disk-stored collections, as storage space is way cheaper than memory. That's also support for disk-stored collections, as storage space is way cheaper than memory. That's also
why we have introduced the [Scalar Quantization](/articles/scalar-quantization) mechanism recently, why we have introduced the [Scalar Quantization](/articles/scalar-quantization/) mechanism recently,
which makes it possible to reduce the memory requirements by up to four times. which makes it possible to reduce the memory requirements by up to four times.
Today, we are bringing a new quantization mechanism to life. A separate article on [Product Today, we are bringing a new quantization mechanism to life. A separate article on [Product
@@ -180,7 +180,7 @@ client.search_groups(
We are excited to announce a more user-friendly way to organize and work with your collections inside of Qdrant. Our dashboard's design is simple, but very intuitive and easy to access. We are excited to announce a more user-friendly way to organize and work with your collections inside of Qdrant. Our dashboard's design is simple, but very intuitive and easy to access.
Try it out now! If you have Docker running, you can [quickstart Qdrant](https://qdrant.tech/documentation/quick-start/) and access the Dashboard locally from [http://localhost:6333/dashboard](http://localhost:6333/dashboard). You should see this simple access point to Qdrant: Try it out now! If you have Docker running, you can [quickstart Qdrant](/documentation/quick-start/) and access the Dashboard locally from [http://localhost:6333/dashboard](http://localhost:6333/dashboard). You should see this simple access point to Qdrant:
![Qdrant Web UI](/articles_data/qdrant-1.3.x/web-ui.png) ![Qdrant Web UI](/articles_data/qdrant-1.3.x/web-ui.png)
@@ -49,7 +49,7 @@ Until now, Qdrant has not been able to handle sparse vectors natively. Some were
Things have changed since then, as so many of you wanted a single tool for sparse and dense vectors. And responding to this [popular](https://github.com/qdrant/qdrant/issues/1678) [demand](https://github.com/qdrant/qdrant/issues/1135), we've now introduced sparse vectors! Things have changed since then, as so many of you wanted a single tool for sparse and dense vectors. And responding to this [popular](https://github.com/qdrant/qdrant/issues/1678) [demand](https://github.com/qdrant/qdrant/issues/1135), we've now introduced sparse vectors!
If you're coming across the topic of sparse vectors for the first time, our [Brief History of Search](https://qdrant.tech/documentation/overview/vector-search/) explains the difference between sparse and dense vectors. If you're coming across the topic of sparse vectors for the first time, our [Brief History of Search](/documentation/overview/vector-search/) explains the difference between sparse and dense vectors.
Check out the [sparse vectors article](../sparse-vectors/) and [sparse vectors index docs](/documentation/concepts/indexing/#sparse-vector-index) for more details on what this new index means for Qdrant users. Check out the [sparse vectors article](../sparse-vectors/) and [sparse vectors index docs](/documentation/concepts/indexing/#sparse-vector-index) for more details on what this new index means for Qdrant users.
@@ -29,7 +29,7 @@ This time around, we have focused on Qdrant's internals. Our goal was to optimiz
## Faster search with sparse vectors ## Faster search with sparse vectors
Search throughput is now up to 16 times faster for sparse vectors. If you are [using Qdrant for hybrid search](https://qdrant.tech/articles/sparse-vectors/), this means that you can now handle up to sixteen times as many queries. This improvement comes from extensive backend optimizations aimed at increasing efficiency and capacity. Search throughput is now up to 16 times faster for sparse vectors. If you are [using Qdrant for hybrid search](/articles/sparse-vectors/), this means that you can now handle up to sixteen times as many queries. This improvement comes from extensive backend optimizations aimed at increasing efficiency and capacity.
What this means for your setup: What this means for your setup:
@@ -51,9 +51,9 @@ Latency (y-axis) has dropped significantly for queries. You can see the before/a
The colors within both scatter plots show the frequency of results. The red dots show that the highest concentration is around 2200ms (before) and 135ms (after). This tells us that latency for sparse vectors queries dropped by about a factor of 16. Therefore, the time it takes to retrieve an answer with Qdrant is that much shorter. The colors within both scatter plots show the frequency of results. The red dots show that the highest concentration is around 2200ms (before) and 135ms (after). This tells us that latency for sparse vectors queries dropped by about a factor of 16. Therefore, the time it takes to retrieve an answer with Qdrant is that much shorter.
This performance increase can have a dramatic effect on hybrid search implementations. [Read more about how to set this up.](https://qdrant.tech/articles/sparse-vectors/) This performance increase can have a dramatic effect on hybrid search implementations. [Read more about how to set this up.](/articles/sparse-vectors/)
FYI, sparse vectors were released in [Qdrant v.1.7.0](https://qdrant.tech/articles/qdrant-1.7.x/#sparse-vectors). They are stored using a different index, so first [check out the documentation](https://qdrant.tech/documentation/concepts/indexing/#sparse-vector-index) if you want to try an implementation. FYI, sparse vectors were released in [Qdrant v.1.7.0](/articles/qdrant-1.7.x/#sparse-vectors). They are stored using a different index, so first [check out the documentation](/documentation/concepts/indexing/#sparse-vector-index) if you want to try an implementation.
## CPU resource management ## CPU resource management
@@ -63,7 +63,7 @@ This isn't mandatory, as Qdrant is by default tuned to strike the right balance
This version introduces a `optimizer_cpu_budget` parameter to control the maximum number of CPUs used for indexing. This version introduces a `optimizer_cpu_budget` parameter to control the maximum number of CPUs used for indexing.
> Read more about `config.yaml` in the [configuration file](https://qdrant.tech/documentation/guides/configuration/). > Read more about `config.yaml` in the [configuration file](/documentation/guides/configuration/).
```yaml ```yaml
# CPU budget, how many CPUs (threads) to allocate for an optimization job. # CPU budget, how many CPUs (threads) to allocate for an optimization job.
@@ -98,7 +98,7 @@ This approach ensures stability in the vector search index, with faster and more
Beyond these enhancements, [Qdrant v1.8.0](https://github.com/qdrant/qdrant/releases/tag/v1.8.0) adds and improves on several smaller features: Beyond these enhancements, [Qdrant v1.8.0](https://github.com/qdrant/qdrant/releases/tag/v1.8.0) adds and improves on several smaller features:
1. **Order points by payload:** In addition to searching for semantic results, you might want to retrieve results by specific metadata (such as price). You can now use Scroll API to [order points by payload key](/documentation/concepts/points/#order-points-by-payload-key). 1. **Order points by payload:** In addition to searching for semantic results, you might want to retrieve results by specific metadata (such as price). You can now use Scroll API to [order points by payload key](/documentation/concepts/points/#order-points-by-payload-key).
2. **Datetime support:** We have implemented [datetime support for the payload index](https://qdrant.tech/documentation/concepts/filtering/#datetime-range). Prior to this, if you wanted to search for a specific datetime range, you would have had to convert dates to UNIX timestamps. ([PR#3320](https://github.com/qdrant/qdrant/issues/3320)) 2. **Datetime support:** We have implemented [datetime support for the payload index](/documentation/concepts/filtering/#datetime-range). Prior to this, if you wanted to search for a specific datetime range, you would have had to convert dates to UNIX timestamps. ([PR#3320](https://github.com/qdrant/qdrant/issues/3320))
3. **Check collection existence:** You can check whether a collection exists via the `/exists` endpoint to the `/collections/{collection_name}`. You will get a true/false response. ([PR#3472](https://github.com/qdrant/qdrant/pull/3472)). 3. **Check collection existence:** You can check whether a collection exists via the `/exists` endpoint to the `/collections/{collection_name}`. You will get a true/false response. ([PR#3472](https://github.com/qdrant/qdrant/pull/3472)).
4. **Find points** whose payloads match more than the minimal amount of conditions. We included the `min_should` match feature for a condition to be `true` ([PR#3331](https://github.com/qdrant/qdrant/pull/3466/)). 4. **Find points** whose payloads match more than the minimal amount of conditions. We included the `min_should` match feature for a condition to be `true` ([PR#3331](https://github.com/qdrant/qdrant/pull/3466/)).
5. **Modify nested fields:** We have improved the `set_payload` API, adding the ability to update nested fields ([PR#3548](https://github.com/qdrant/qdrant/pull/3548)). 5. **Modify nested fields:** We have improved the `set_payload` API, adding the ability to update nested fields ([PR#3548](https://github.com/qdrant/qdrant/pull/3548)).
@@ -49,7 +49,7 @@ We don’t think 60-80% is good enough. The LLM might retrieve enough relevant f
> The whole point of vector search is to circumvent this process by efficiently picking the information your app needs to generate the best response. A vector database keeps the compute load low and the query response fast. You don’t need to wait for the LLM at all. > The whole point of vector search is to circumvent this process by efficiently picking the information your app needs to generate the best response. A vector database keeps the compute load low and the query response fast. You don’t need to wait for the LLM at all.
Qdrant’s benchmark results are strongly in favor of accuracy and efficiency. We recommend that you consider them before deciding that an LLM is enough. Take a look at our [open-source benchmark reports](https://qdrant.tech/benchmarks/) and [try out the tests](https://github.com/qdrant/vector-db-benchmark) yourself. Qdrant’s benchmark results are strongly in favor of accuracy and efficiency. We recommend that you consider them before deciding that an LLM is enough. Take a look at our [open-source benchmark reports](/benchmarks/) and [try out the tests](https://github.com/qdrant/vector-db-benchmark) yourself.
## Vector search in compound systems ## Vector search in compound systems
@@ -74,7 +74,7 @@ That’s a buck a question.
> According to our estimations, vector search queries are **at least** 100 million times cheaper than queries made by LLMs. > According to our estimations, vector search queries are **at least** 100 million times cheaper than queries made by LLMs.
Conversely, the only up-front investment with vector databases is the indexing (which requires more compute). After this step, everything else is a breeze. Once setup, Qdrant easily scales via [features like Multitenancy and Sharding](https://qdrant.tech/articles/multitenancy/). This lets you scale up your reliance on the vector retrieval process and minimize your use of the compute-heavy LLMs. As an optimization measure, Qdrant is irreplaceable. Conversely, the only up-front investment with vector databases is the indexing (which requires more compute). After this step, everything else is a breeze. Once setup, Qdrant easily scales via [features like Multitenancy and Sharding](/articles/multitenancy/). This lets you scale up your reliance on the vector retrieval process and minimize your use of the compute-heavy LLMs. As an optimization measure, Qdrant is irreplaceable.
Julien Simon from HuggingFace says it best: Julien Simon from HuggingFace says it best:
@@ -89,4 +89,4 @@ Our customers remind us of this fact every day. As a product, our vector databas
We want to keep Qdrant compact, efficient and with a focused purpose. This purpose is to empower our customers to use it however they see fit. We want to keep Qdrant compact, efficient and with a focused purpose. This purpose is to empower our customers to use it however they see fit.
When large enterprises release their generative AI into production, they need to keep costs under control, while retaining the best possible quality of responses. Qdrant has the tools to do just that. Whether through [RAG, Semantic Search, Dissimilarity Search, Recommendations or Multimodality](https://qdrant.tech/articles/vector-similarity-beyond-search/) - Qdrant will continue to journey on. When large enterprises release their generative AI into production, they need to keep costs under control, while retaining the best possible quality of responses. Qdrant has the tools to do just that. Whether through [RAG, Semantic Search, Dissimilarity Search, Recommendations or Multimodality](/articles/vector-similarity-beyond-search/) - Qdrant will continue to journey on.
@@ -130,7 +130,7 @@ performance. As usual, we performed some benchmarks to support this statement!
## Benchmarks ## Benchmarks
We simply used the same approach as we use in all [the other benchmarks we publish](/benchmarks). We simply used the same approach as we use in all [the other benchmarks we publish](/benchmarks/).
Both [Arxiv-titles-384-angular-no-filters](https://github.com/qdrant/ann-filtering-benchmark-datasets) Both [Arxiv-titles-384-angular-no-filters](https://github.com/qdrant/ann-filtering-benchmark-datasets)
and [Gist-960](https://github.com/erikbern/ann-benchmarks/) datasets were chosen to make and [Gist-960](https://github.com/erikbern/ann-benchmarks/) datasets were chosen to make
the comparison between non-quantized and quantized vectors. The results are summarized the comparison between non-quantized and quantized vectors. The results are summarized
@@ -280,7 +280,7 @@ And another group with more strict memory limits:
In those experiments, throughput was mainly defined by the number of disk reads, and quantization efficiently reduces it by allowing more vectors in RAM. In those experiments, throughput was mainly defined by the number of disk reads, and quantization efficiently reduces it by allowing more vectors in RAM.
Read more about on-disk storage in Qdrant and how we measure its performance in our article: [Minimal RAM you need to serve a million vectors Read more about on-disk storage in Qdrant and how we measure its performance in our article: [Minimal RAM you need to serve a million vectors
](https://qdrant.tech/articles/memory-consumption/). ](/articles/memory-consumption/).
The mechanism of Scalar Quantization with rescoring disabled pushes the limits of low-end The mechanism of Scalar Quantization with rescoring disabled pushes the limits of low-end
machines even further. It seems like handling lots of requests does not require an machines even further. It seems like handling lots of requests does not require an
@@ -288,6 +288,6 @@ expensive setup if you can agree to a small decrease in the search precision.
### Good practices ### Good practices
Qdrant documentation on [Scalar Quantization](https://qdrant.tech/documentation/quantization/#setting-up-quantization-in-qdrant) Qdrant documentation on [Scalar Quantization](/documentation/quantization/#setting-up-quantization-in-qdrant)
is a great resource describing different scenarios and strategies to achieve up to 4x is a great resource describing different scenarios and strategies to achieve up to 4x
lower memory footprint and even up to 2x performance increase. lower memory footprint and even up to 2x performance increase.
@@ -54,7 +54,7 @@ POST collections/site/points/recommend
Now I have, in the best Rust tradition, a blazingly fast semantic search. Now I have, in the best Rust tradition, a blazingly fast semantic search.
To demo it, I used our [Qdrant documentation website](https://qdrant.tech/documentation)'s page search, replacing our previous Python implementation. So in order to not just spew empty words, here is a benchmark, showing different queries that exercise different code paths. To demo it, I used our [Qdrant documentation website](/documentation/)'s page search, replacing our previous Python implementation. So in order to not just spew empty words, here is a benchmark, showing different queries that exercise different code paths.
Since the operations themselves are far faster than the network whose fickle nature would have swamped most measurable differences, I benchmarked both the Python and Rust services locally. I'm measuring both versions on the same AMD Ryzen 9 5900HX with 16GB RAM running Linux. The table shows the average time and error bound in milliseconds. I only measured up to a thousand concurrent requests. None of the services showed any slowdown with more requests in that range. I do not expect our service to become DDOS'd, so I didn't benchmark with more load. Since the operations themselves are far faster than the network whose fickle nature would have swamped most measurable differences, I benchmarked both the Python and Rust services locally. I'm measuring both versions on the same AMD Ryzen 9 5900HX with 16GB RAM running Linux. The table shows the average time and error bound in milliseconds. I only measured up to a thousand concurrent requests. None of the services showed any slowdown with more requests in that range. I do not expect our service to become DDOS'd, so I didn't benchmark with more load.
@@ -46,11 +46,11 @@ The current Qdrant ecosystem consists of excellent products to work with vector
{{< figure src=/articles_data/seed-round/ecosystem.png caption="Qdrant Ecosystem" alt="Qdrant Vector Database Ecosystem" >}} {{< figure src=/articles_data/seed-round/ecosystem.png caption="Qdrant Ecosystem" alt="Qdrant Vector Database Ecosystem" >}}
Our plan for the current [open-source roadmap](https://github.com/qdrant/qdrant/blob/master/docs/roadmap/README.md) is to make billion-scale vector search affordable. Our recent release of the [Scalar Quantization](https://qdrant.tech/articles/scalar-quantization/) improves both memory usage (x4) as well as speed (x2). Upcoming [Product Quantization](https://www.irisa.fr/texmex/people/jegou/papers/jegou_searching_with_quantization.pdf) will introduce even another option with more memory saving. Stay tuned. Our plan for the current [open-source roadmap](https://github.com/qdrant/qdrant/blob/master/docs/roadmap/README.md) is to make billion-scale vector search affordable. Our recent release of the [Scalar Quantization](/articles/scalar-quantization/) improves both memory usage (x4) as well as speed (x2). Upcoming [Product Quantization](https://www.irisa.fr/texmex/people/jegou/papers/jegou_searching_with_quantization.pdf) will introduce even another option with more memory saving. Stay tuned.
Qdrant started more than two years ago with the mission of building a vector database powered by a well-thought-out tech stack. Using Rust as the system programming language and technical architecture decision during the development of the engine made Qdrant the leading and one of the most popular vector database solutions. Qdrant started more than two years ago with the mission of building a vector database powered by a well-thought-out tech stack. Using Rust as the system programming language and technical architecture decision during the development of the engine made Qdrant the leading and one of the most popular vector database solutions.
Our unique custom modification of the [HNSW algorithm](https://qdrant.tech/articles/filtrable-hnsw/) for Approximate Nearest Neighbor Search (ANN) allows querying the result with a state-of-the-art speed and applying filters without compromising on results. Cloud-native support for distributed deployment and replications makes the engine suitable for high-throughput applications with real-time latency requirements. Rust brings stability, efficiency, and the possibility to make optimization on a very low level. In general, we always aim for the best possible results in [performance](https://qdrant.tech/benchmarks/), code quality, and feature set. Our unique custom modification of the [HNSW algorithm](/articles/filtrable-hnsw/) for Approximate Nearest Neighbor Search (ANN) allows querying the result with a state-of-the-art speed and applying filters without compromising on results. Cloud-native support for distributed deployment and replications makes the engine suitable for high-throughput applications with real-time latency requirements. Rust brings stability, efficiency, and the possibility to make optimization on a very low level. In general, we always aim for the best possible results in [performance](/benchmarks/), code quality, and feature set.
Most importantly, we want to say a big thank you to our [open-source community](https://qdrant.to/discord), our adopters, our contributors, and our customers. Your active participation in the development of our products has helped make Qdrant the best vector database on the market. I cannot imagine how we could do what we’re doing without the community or without being open-source and having the TRUST of the engineers. Thanks to all of you! Most importantly, we want to say a big thank you to our [open-source community](https://qdrant.to/discord), our adopters, our contributors, and our customers. Your active participation in the development of our products has helped make Qdrant the best vector database on the market. I cannot imagine how we could do what we’re doing without the community or without being open-source and having the TRUST of the engineers. Thanks to all of you!
@@ -23,7 +23,7 @@ You may find all of the assets for this tutorial on [GitHub](https://github.com/
* [cargo lambda](https://cargo-lambda.info) (install via package manager, [download](https://github.com/cargo-lambda/cargo-lambda/releases) binary or `cargo install cargo-lambda`) * [cargo lambda](https://cargo-lambda.info) (install via package manager, [download](https://github.com/cargo-lambda/cargo-lambda/releases) binary or `cargo install cargo-lambda`)
* The [AWS CLI](https://aws.amazon.com/cli) * The [AWS CLI](https://aws.amazon.com/cli)
* Qdrant instance ([free tier](https://cloud.qdrant.io) available) * Qdrant instance ([free tier](https://cloud.qdrant.io) available)
* An embedding provider service of your choice (see our [Embeddings docs](https://qdrant.tech/documentation/embeddings). You may be able to get credits from [AI Grant](https://aigrant.org), also Cohere has a [rate-limited non-commercial free tier](https://cohere.com/pricing)) * An embedding provider service of your choice (see our [Embeddings docs](/documentation/embeddings/). You may be able to get credits from [AI Grant](https://aigrant.org), also Cohere has a [rate-limited non-commercial free tier](https://cohere.com/pricing))
* AWS Lambda account (12-month free tier available) * AWS Lambda account (12-month free tier available)
## What you're going to build ## What you're going to build
@@ -176,7 +176,7 @@ pub async fn embed(client: &Client, text: &str, api_key: &str) -> Result<Vec<Vec
Note that this may return multiple vectors if the text overflows the input dimensions. Note that this may return multiple vectors if the text overflows the input dimensions.
Cohere's `small` model has 1024 output dimensions. Cohere's `small` model has 1024 output dimensions.
Other providers have similar interfaces. Consult our [Embeddings docs](https://qdrant.tech/documentation/embeddings) for further information. See how little code it took to get the embedding? Other providers have similar interfaces. Consult our [Embeddings docs](/documentation/embeddings/) for further information. See how little code it took to get the embedding?
While you're at it, it's a good idea to write a small test to check if embedding works and the vectors are of the expected size: While you're at it, it's a good idea to write a small test to check if embedding works and the vectors are of the expected size:
@@ -514,6 +514,6 @@ Alright, folks, let's wrap it up. Better search isn't a 'nice-to-have,' it's a g
Got questions? Our [Discord community](https://qdrant.to/discord?utm_source=qdrant&utm_medium=website&utm_campaign=sparse-vectors&utm_content=article&utm_term=sparse-vectors) is teeming with answers. Got questions? Our [Discord community](https://qdrant.to/discord?utm_source=qdrant&utm_medium=website&utm_campaign=sparse-vectors&utm_content=article&utm_term=sparse-vectors) is teeming with answers.
If you enjoyed reading this, why not sign up for our [newsletter](https://qdrant.tech/subscribe/?utm_source=qdrant&utm_medium=website&utm_campaign=sparse-vectors&utm_content=article&utm_term=sparse-vectors) to stay ahead of the curve. If you enjoyed reading this, why not sign up for our [newsletter](/subscribe/?utm_source=qdrant&utm_medium=website&utm_campaign=sparse-vectors&utm_content=article&utm_term=sparse-vectors) to stay ahead of the curve.
And, of course, a big thanks to you, our readers, for pushing us to make ranking better for everyone. And, of course, a big thanks to you, our readers, for pushing us to make ranking better for everyone.
@@ -73,7 +73,7 @@ To ensure our catalog is accurate, we can use a dissimilarity search to highligh
To do this, we only need to search for the most dissimilar items using the To do this, we only need to search for the most dissimilar items using the
embedding of the category title itself as a query. embedding of the category title itself as a query.
This can be too broad, so, combining it with filters —a [Qdrant superpower](/articles/filtrable-hnsw)—, we can narrow down the search to a specific category. This can be too broad, so, combining it with filters —a [Qdrant superpower](/articles/filtrable-hnsw/)—, we can narrow down the search to a specific category.
{{< figure src=/articles_data/vector-similarity-beyond-search/mislabelling.png caption="Mislabeling Detection" >}} {{< figure src=/articles_data/vector-similarity-beyond-search/mislabelling.png caption="Mislabeling Detection" >}}
@@ -58,7 +58,7 @@ This capability is crucial for creating search systems, recommendation engines,
## How do embeddings work? ## How do embeddings work?
Embeddings are created through neural networks. They capture complex relationships and semantics into [dense vectors](https://www1.se.cuhk.edu.hk/~seem5680/lecture/semantics-with-dense-vectors-2018.pdf) which are more suitable for machine learning and data processing applications. They can then project these vectors into a proper **high-dimensional** space, specifically, a [Vector Database](https://qdrant.tech/articles/what-is-a-vector-database/). Embeddings are created through neural networks. They capture complex relationships and semantics into [dense vectors](https://www1.se.cuhk.edu.hk/~seem5680/lecture/semantics-with-dense-vectors-2018.pdf) which are more suitable for machine learning and data processing applications. They can then project these vectors into a proper **high-dimensional** space, specifically, a [Vector Database](/articles/what-is-a-vector-database/).
@@ -130,7 +130,7 @@ The model creates a vector embedding for "biophilic design" that encapsulates th
### Integration with Embedding APIs ### Integration with Embedding APIs
Selecting the right embedding model for your use case is crucial to your application performance. Qdrant makes it easier by offering seamless integration with the best selection of embedding APIs, including [Cohere](https://qdrant.tech/documentation/embeddings/cohere/), [Gemini](https://qdrant.tech/documentation/embeddings/gemini/), [Jina Embeddings](https://qdrant.tech/documentation/embeddings/jina-embeddings/), [OpenAI](https://qdrant.tech/documentation/embeddings/openai/), [Aleph Alpha](https://qdrant.tech/documentation/embeddings/aleph-alpha/), [Fastembed](https://github.com/qdrant/fastembed), and [AWS Bedrock](https://qdrant.tech/documentation/embeddings/bedrock/). Selecting the right embedding model for your use case is crucial to your application performance. Qdrant makes it easier by offering seamless integration with the best selection of embedding APIs, including [Cohere](/documentation/embeddings/cohere/), [Gemini](/documentation/embeddings/gemini/), [Jina Embeddings](/documentation/embeddings/jina-embeddings/), [OpenAI](/documentation/embeddings/openai/), [Aleph Alpha](/documentation/embeddings/aleph-alpha/), [Fastembed](https://github.com/qdrant/fastembed), and [AWS Bedrock](/documentation/embeddings/bedrock/).
If you’re looking for NLP and rapid prototyping, including language translation, question-answering, and text generation, OpenAI is a great choice. Gemini is ideal for image search, duplicate detection, and clustering tasks. If you’re looking for NLP and rapid prototyping, including language translation, question-answering, and text generation, OpenAI is a great choice. Gemini is ideal for image search, duplicate detection, and clustering tasks.
@@ -140,7 +140,7 @@ We plan to go deeper into selecting the best model based on performance, cost, i
## Create a Neural Search Service with Fastembed ## Create a Neural Search Service with Fastembed
Now that you’re familiar with the core concepts around vector embeddings, how about start building your own [Neural Search Service](https://qdrant.tech/documentation/tutorials/neural-search-fastembed/)? Now that you’re familiar with the core concepts around vector embeddings, how about start building your own [Neural Search Service](/documentation/tutorials/neural-search-fastembed/)?
Tutorial guides you through a practical application of how to use Qdrant for document management based on descriptions of companies from [startups-list.com](https://www.startups-list.com/). From embedding data, integrating it with Qdrant's vector database, constructing a search API, and finally deploying your solution with FastAPI. Tutorial guides you through a practical application of how to use Qdrant for document management based on descriptions of companies from [startups-list.com](https://www.startups-list.com/). From embedding data, integrating it with Qdrant's vector database, constructing a search API, and finally deploying your solution with FastAPI.
@@ -94,14 +94,14 @@ This way, finding similar images becomes a quick hop across related groups, inst
![](/articles_data/what-is-a-vector-database/Indexing.jpg) ![](/articles_data/what-is-a-vector-database/Indexing.jpg)
Different indexing methods exist, each with its strengths. [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) balances speed and accuracy like a well-connected network of shortcuts in the crowd. Others, like IVF or Product Quantization, focus on specific tasks or memory efficiency. Different indexing methods exist, each with its strengths. [HNSW](/articles/filtrable-hnsw/) balances speed and accuracy like a well-connected network of shortcuts in the crowd. Others, like IVF or Product Quantization, focus on specific tasks or memory efficiency.
#### What is Binary Quantization? #### What is Binary Quantization?
Quantization is a technique used for reducing the total size of the database. It works by compressing vectors into a more compact representation at the cost of accuracy. Quantization is a technique used for reducing the total size of the database. It works by compressing vectors into a more compact representation at the cost of accuracy.
[Binary Quantization](https://qdrant.tech/articles/binary-quantization/) is a fast indexing and data compression method used by Qdrant. It supports vector comparisons, which can dramatically speed up query processing times (up to 40x faster!). [Binary Quantization](/articles/binary-quantization/) is a fast indexing and data compression method used by Qdrant. It supports vector comparisons, which can dramatically speed up query processing times (up to 40x faster!).
Think of each data point as a ruler. Binary quantization splits this ruler in half at a certain point, marking everything above as "1" and everything below as "0". This [binarization](https://deepai.org/machine-learning-glossary-and-terms/binarization) process results in a string of bits, representing the original vector. Think of each data point as a ruler. Binary quantization splits this ruler in half at a certain point, marking everything above as "1" and everything below as "0". This [binarization](https://deepai.org/machine-learning-glossary-and-terms/binarization) process results in a string of bits, representing the original vector.
@@ -115,9 +115,9 @@ This "quantized" code is much smaller and easier to compare. Especially for Open
### What is Similarity Search? ### What is Similarity Search?
[Similarity search](https://qdrant.tech/documentation/concepts/search/) allows you to search not by keywords but by meaning. This way you can do searches such as similar songs that evoke the same mood, finding images that match your artistic vision, or even exploring emotional patterns in text. [Similarity search](/documentation/concepts/search/) allows you to search not by keywords but by meaning. This way you can do searches such as similar songs that evoke the same mood, finding images that match your artistic vision, or even exploring emotional patterns in text.
The way it works is, when the user queries the database, this query is also converted into a vector (the query vector). The [vector search](https://qdrant.tech/documentation/overview/vector-search/) starts at the top layer of the HNSW index, where the algorithm quickly identifies the area of the graph likely to contain vectors closest to the query vector. The algorithm compares your query vector to all the others, using metrics like "distance" or "similarity" to gauge how close they are. The way it works is, when the user queries the database, this query is also converted into a vector (the query vector). The [vector search](/documentation/overview/vector-search/) starts at the top layer of the HNSW index, where the algorithm quickly identifies the area of the graph likely to contain vectors closest to the query vector. The algorithm compares your query vector to all the others, using metrics like "distance" or "similarity" to gauge how close they are.
The search then moves down progressively narrowing down to more closely related vectors. The goal is to narrow down the dataset to the most relevant items. The image below illustrates this. The search then moves down progressively narrowing down to more closely related vectors. The goal is to narrow down the dataset to the most relevant items. The image below illustrates this.
@@ -178,11 +178,11 @@ A vector database is made of multiple different entities and relations. Here's a
![](/articles_data/what-is-a-vector-database/Architecture-of-a-Vector-Database.jpg) ![](/articles_data/what-is-a-vector-database/Architecture-of-a-Vector-Database.jpg)
**Collections**: [Collections](https://qdrant.tech/documentation/concepts/collections/) are a named set of data points, where each point is a vector with an associated payload. All vectors within a collection must have the same dimensionality and be comparable using a single metric. **Collections**: [Collections](/documentation/concepts/collections/) are a named set of data points, where each point is a vector with an associated payload. All vectors within a collection must have the same dimensionality and be comparable using a single metric.
**Distance Metrics**: These metrics are used to measure the similarity between vectors. The choice of distance metric is made when creating a collection. It depends on the nature of the vectors and how they were generated, considering the neural network used for the encoding. **Distance Metrics**: These metrics are used to measure the similarity between vectors. The choice of distance metric is made when creating a collection. It depends on the nature of the vectors and how they were generated, considering the neural network used for the encoding.
**Points**: Each [point](https://qdrant.tech/documentation/concepts/points/) consists of a **vector** and can also include an optional **identifier** (ID) and **[payload](https://qdrant.tech/documentation/concepts/payload/)**. The vector represents the high-dimensional data and the payload carries metadata information in a JSON format, giving the data point more context or attributes. **Points**: Each [point](/documentation/concepts/points/) consists of a **vector** and can also include an optional **identifier** (ID) and **[payload](/documentation/concepts/payload/)**. The vector represents the high-dimensional data and the payload carries metadata information in a JSON format, giving the data point more context or attributes.
**Storage Options**: There are two primary storage options. The in-memory storage option keeps all vectors in RAM, which allows for the highest speed in data access since disk access is only required for persistence. **Storage Options**: There are two primary storage options. The in-memory storage option keeps all vectors in RAM, which allows for the highest speed in data access since disk access is only required for persistence.
@@ -206,13 +206,13 @@ Here’s some examples on how to take advantage of using vector databases:
There are many other use cases like for **fraud detection and anomaly analysis** used in sectors like finance and cybersecurity, to detect anomalies and potential fraud. And **Content-Based Image Retrieval (CBIR)** for images by comparing vector representations rather than metadata or tags. There are many other use cases like for **fraud detection and anomaly analysis** used in sectors like finance and cybersecurity, to detect anomalies and potential fraud. And **Content-Based Image Retrieval (CBIR)** for images by comparing vector representations rather than metadata or tags.
Those are just a few examples. The ability of vector databases to “match” data with queries makes them essential for multiple types of applications. Here are some more [use cases examples](https://qdrant.tech/use-cases/) you can take a look at. Those are just a few examples. The ability of vector databases to “match” data with queries makes them essential for multiple types of applications. Here are some more [use cases examples](/use-cases/) you can take a look at.
### Starting Your First Vector Database Project ### Starting Your First Vector Database Project
Now that you're familiar with the core concepts around vector databases, it’s time to get our hands dirty. [Start by building your own semantic search engine](https://qdrant.tech/documentation/tutorials/search-beginners/) for science fiction books in just about 5 minutes with the help of Qdrant. You can also watch our [video tutorial](https://www.youtube.com/watch?v=AASiqmtKo54). Now that you're familiar with the core concepts around vector databases, it’s time to get our hands dirty. [Start by building your own semantic search engine](/documentation/tutorials/search-beginners/) for science fiction books in just about 5 minutes with the help of Qdrant. You can also watch our [video tutorial](https://www.youtube.com/watch?v=AASiqmtKo54).
Feeling ready to dive into a more complex project? Take the next step and get started building an actual [Neural Search Service with a complete API and a dataset](https://qdrant.tech/documentation/tutorials/neural-search/). Feeling ready to dive into a more complex project? Take the next step and get started building an actual [Neural Search Service with a complete API and a dataset](/documentation/tutorials/neural-search/).
Let’s get into action! Let’s get into action!
+1 -1
View File
@@ -27,7 +27,7 @@ So Andrey looked around at what younger languages would fit the challenge. After
This early decision has been validated time and again. When first learning Rust, the compiler’s error messages are very helpful (and have only improved in the meantime). It’s easy to keep memory profile low when one doesn’t have to wrestle a garbage collector and has complete control over stack and heap. Apart from the much advertised memory safety, many footguns one can run into when writing C++ have been meticulously designed out. And it’s much easier to parallelize a task if one doesn’t have to fear data races. This early decision has been validated time and again. When first learning Rust, the compiler’s error messages are very helpful (and have only improved in the meantime). It’s easy to keep memory profile low when one doesn’t have to wrestle a garbage collector and has complete control over stack and heap. Apart from the much advertised memory safety, many footguns one can run into when writing C++ have been meticulously designed out. And it’s much easier to parallelize a task if one doesn’t have to fear data races.
With Qdrant written in Rust, we can offer cloud services that don’t keep us awake at night, thanks to Rust’s famed robustness. A current qdrant docker container comes in at just a bit over 50MB — try that for size. As for performance… have some [benchmarks](https://qdrant.tech/benchmarks). With Qdrant written in Rust, we can offer cloud services that don’t keep us awake at night, thanks to Rust’s famed robustness. A current qdrant docker container comes in at just a bit over 50MB — try that for size. As for performance… have some [benchmarks](/benchmarks/).
And we don’t have to compromise on ergonomics either, not for us nor for our users. Of course, there are downsides: Rust compile times are usually similar to C++’s, and though the learning curve has been considerably softened in the last years, it’s still no match for easy-entry languages like Python or Go. But learning it is a one-time cost. Contrast this with Go, where you may find [the apparent simplicity is only skin-deep](https://fasterthanli.me/articles/i-want-off-mr-golangs-wild-ride). And we don’t have to compromise on ergonomics either, not for us nor for our users. Of course, there are downsides: Rust compile times are usually similar to C++’s, and though the learning curve has been considerably softened in the last years, it’s still no match for easy-entry languages like Python or Go. But learning it is a one-time cost. Contrast this with Go, where you may find [the apparent simplicity is only skin-deep](https://fasterthanli.me/articles/i-want-off-mr-golangs-wild-ride).
@@ -7,7 +7,7 @@ weight: 1
# Benchmarking Vector Databases # Benchmarking Vector Databases
At Qdrant, performance is the top-most priority. We always make sure that we use system resources efficiently so you get the **fastest and most accurate results at the cheapest cloud costs**. So all of our decisions from [choosing Rust](/articles/why-rust), [io optimisations](/articles/io_uring), [serverless support](/articles/serverless), [binary quantization](/articles/binary-quantization), to our [fastembed library](/articles/fastembed) are all based on our principle. In this article, we will compare how Qdrant performs against the other vector search engines. At Qdrant, performance is the top-most priority. We always make sure that we use system resources efficiently so you get the **fastest and most accurate results at the cheapest cloud costs**. So all of our decisions from [choosing Rust](/articles/why-rust/), [io optimisations](/articles/io_uring/), [serverless support](/articles/serverless/), [binary quantization](/articles/binary-quantization/), to our [fastembed library](/articles/fastembed/) are all based on our principle. In this article, we will compare how Qdrant performs against the other vector search engines.
Here are the principles we followed while designing these benchmarks: Here are the principles we followed while designing these benchmarks:
@@ -15,7 +15,7 @@ Unlisted: false
## Observations ## Observations
Most of the engines have improved since [our last run](/benchmarks/single-node-speed-benchmark-2022). Both life and software have trade-offs but some clearly do better: Most of the engines have improved since [our last run](/benchmarks/single-node-speed-benchmark-2022/). Both life and software have trade-offs but some clearly do better:
* **`Qdrant` achives highest RPS and lowest latencies in almost all the scenarios, no matter the precision threshold and the metric we choose.** It has also shown 4x RPS gains on one of the datasets. * **`Qdrant` achives highest RPS and lowest latencies in almost all the scenarios, no matter the precision threshold and the metric we choose.** It has also shown 4x RPS gains on one of the datasets.
* `Elasticsearch` has become considerably fast for many cases but it's very slow in terms of indexing time. It can be 10x slower when storing 10M+ vectors of 96 dimensions! (32mins vs 5.5 hrs) * `Elasticsearch` has become considerably fast for many cases but it's very slow in terms of indexing time. It can be 10x slower when storing 10M+ vectors of 96 dimensions! (32mins vs 5.5 hrs)
@@ -106,7 +106,7 @@ Looking at the complexity, scale and adaptability of the desired solution, the t
**6. Impressive Benchmarks:** **6. Impressive Benchmarks:**
[Qdrant’s benchmarks](https://qdrant.tech/benchmarks/) has definitely been one of the key motivations for the Dailymotion’s team to try the solution and the team comments that the performance has been only better than the benchmarks. [Qdrant’s benchmarks](/benchmarks/) has definitely been one of the key motivations for the Dailymotion’s team to try the solution and the team comments that the performance has been only better than the benchmarks.
**7. Ease of usage:** **7. Ease of usage:**
@@ -167,7 +167,7 @@ They aim to work on Perspective feed next and say
![perspective-feed-with-qdrant](/case-studies/dailymotion/perspective-feed-qdrant.jpg) ![perspective-feed-with-qdrant](/case-studies/dailymotion/perspective-feed-qdrant.jpg)
The team is also interested in leveraging advanced features like [Qdrant’s Discovery API](https://qdrant.tech/documentation/concepts/explore/#recommendation-api) to promote exploration of content to enable finding not only similar but dissimilar content too by using positive and negative vectors in the queries and making it work with the existing collaborative recommendation model. The team is also interested in leveraging advanced features like [Qdrant’s Discovery API](/documentation/concepts/explore/#recommendation-api) to promote exploration of content to enable finding not only similar but dissimilar content too by using positive and negative vectors in the queries and making it work with the existing collaborative recommendation model.
### References ### References
@@ -82,8 +82,8 @@ billing and increase security by having the instance live within the same VPC.
2. **Scale and optimize:** As the load grew, Dust started to take advantage of Qdrant’s 2. **Scale and optimize:** As the load grew, Dust started to take advantage of Qdrant’s
features to tune the setup for optimization and scale. They started to look into features to tune the setup for optimization and scale. They started to look into
how they map and cache data, as well as applying some of Qdrant’s [built-in how they map and cache data, as well as applying some of Qdrant’s [built-in
compression features](https://qdrant.tech/documentation/guides/quantization/). In particular, Dust leveraged the control of the [MMAP compression features](/documentation/guides/quantization/). In particular, Dust leveraged the control of the [MMAP
payload threshold](https://qdrant.tech/documentation/concepts/storage/#configuring-memmap-storage) as well as [Scalar Quantization](https://qdrant.tech/articles/scalar-quantization/), which enabled Dust to manage payload threshold](/documentation/concepts/storage/#configuring-memmap-storage) as well as [Scalar Quantization](/articles/scalar-quantization/), which enabled Dust to manage
the balance between storing vectors on disk and keeping quantized vectors in RAM, the balance between storing vectors on disk and keeping quantized vectors in RAM,
more effectively. “This allowed us to scale smoothly from there,” Polu says. more effectively. “This allowed us to scale smoothly from there,” Polu says.
@@ -32,7 +32,7 @@ To develop Quaterion, we utilized PyTorch Lightning, leveraging a high-performin
![quaterion](/blog/from_cms/new-cmp-demo.gif) ![quaterion](/blog/from_cms/new-cmp-demo.gif)
This framework empowers vector search [solutions](https://qdrant.tech/solutions/), such as semantic search, anomaly detection, and others, by advanced coaching mechanism, specially designed head layers for pre-trained models, and high flexibility in terms of customization according to large-scale training pipelines and other features. This framework empowers vector search [solutions](/solutions/), such as semantic search, anomaly detection, and others, by advanced coaching mechanism, specially designed head layers for pre-trained models, and high flexibility in terms of customization according to large-scale training pipelines and other features.
Here you can read why similarity learning is preferable to the traditional machine learning approach and how Quaterion can help benefit <https://quaterion.qdrant.tech/getting_started/why_quaterion.html#why-quaterion>    Here you can read why similarity learning is preferable to the traditional machine learning approach and how Quaterion can help benefit <https://quaterion.qdrant.tech/getting_started/why_quaterion.html#why-quaterion>   
@@ -75,7 +75,7 @@ Now as we have a vector representation for all our records, we need to store the
The vector search engine can take care of all these tasks. It provides a convenient API for searching and managing vectors. The vector search engine can take care of all these tasks. It provides a convenient API for searching and managing vectors.
In our tutorial we will use [Qdrant](https://qdrant.tech/) vector search engine. It not only supports all necessary operations with vectors but also allows to store additional payload along with vectors and use it to perform filtering of the search result. Qdrant has a client for python and also defines the API schema if you need to use it from other languages. In our tutorial we will use [Qdrant](/) vector search engine. It not only supports all necessary operations with vectors but also allows to store additional payload along with vectors and use it to perform filtering of the search result. Qdrant has a client for python and also defines the API schema if you need to use it from other languages.
The easiest way to use Qdrant is to run a pre-built image. So make sure you have Docker installed on your system. The easiest way to use Qdrant is to run a pre-built image. So make sure you have Docker installed on your system.
@@ -21,31 +21,31 @@ Together, Pienso and Qdrant will empower enterprises to harness the full potenti
Qdrant enhances the accuracy of large language models (LLMs) by offering an alternative to relying solely on patterns identified during the training phase. By integrating with Qdrant, Pienso will empower customer LLMs with dynamic long-term storage, which will ultimately enable them to generate concrete and factual responses. Qdrant effectively preserves the extensive context windows managed by advanced LLMs, allowing for a broader analysis of the conversation or document at hand. By leveraging this extended context, LLMs can achieve a more comprehensive understanding and produce contextually relevant outputs. Qdrant enhances the accuracy of large language models (LLMs) by offering an alternative to relying solely on patterns identified during the training phase. By integrating with Qdrant, Pienso will empower customer LLMs with dynamic long-term storage, which will ultimately enable them to generate concrete and factual responses. Qdrant effectively preserves the extensive context windows managed by advanced LLMs, allowing for a broader analysis of the conversation or document at hand. By leveraging this extended context, LLMs can achieve a more comprehensive understanding and produce contextually relevant outputs.
## [](https://qdrant.tech/case-studies/pienso/#joint-dedication-to-scalability-efficiency-and-reliability)Joint Dedication to Scalability, Efficiency and Reliability ## [](/case-studies/pienso/#joint-dedication-to-scalability-efficiency-and-reliability)Joint Dedication to Scalability, Efficiency and Reliability
> “Every commercial generative AI use case we encounter benefits from faster training and inference, whether mining customer interactions for next best actions or sifting clinical data to speed a therapeutic through trial and patent processes.” - Birago Jones, CEO, Pienso > “Every commercial generative AI use case we encounter benefits from faster training and inference, whether mining customer interactions for next best actions or sifting clinical data to speed a therapeutic through trial and patent processes.” - Birago Jones, CEO, Pienso
Pienso chose Qdrant for its exceptional LLM interoperability, recognizing the potential it offers in maximizing the power of large language models and interactive deep learning for large enterprises. Qdrant excels in efficient nearest neighbor search, which is an expensive and computationally demanding task. Our ability to store and search high-dimensional vectors with remarkable performance and precision will offer a significant peace of mind to Pienso’s customers. Through intelligent indexing and partitioning techniques, Qdrant will significantly boost the speed of these searches, accelerating both training and inference processes for users. Pienso chose Qdrant for its exceptional LLM interoperability, recognizing the potential it offers in maximizing the power of large language models and interactive deep learning for large enterprises. Qdrant excels in efficient nearest neighbor search, which is an expensive and computationally demanding task. Our ability to store and search high-dimensional vectors with remarkable performance and precision will offer a significant peace of mind to Pienso’s customers. Through intelligent indexing and partitioning techniques, Qdrant will significantly boost the speed of these searches, accelerating both training and inference processes for users.
### [](https://qdrant.tech/case-studies/pienso/#scalability-preparing-for-sustained-growth-in-data-volumes)Scalability: Preparing for Sustained Growth in Data Volumes ### [](/case-studies/pienso/#scalability-preparing-for-sustained-growth-in-data-volumes)Scalability: Preparing for Sustained Growth in Data Volumes
Qdrant’s distributed deployment mode plays a vital role in empowering large enterprises dealing with massive data volumes. It ensures that increasing data volumes do not hinder performance but rather enrich the model’s capabilities, making scalability a seamless process. Moreover, Qdrant is well-suited for Pienso’s enterprise customers as it operates best on bare metal infrastructure, enabling them to maintain complete control over their data sovereignty and autonomous LLM regimes. This ensures that enterprises can maintain their full span of control while leveraging the scalability and performance benefits of Qdrant’s solution. Qdrant’s distributed deployment mode plays a vital role in empowering large enterprises dealing with massive data volumes. It ensures that increasing data volumes do not hinder performance but rather enrich the model’s capabilities, making scalability a seamless process. Moreover, Qdrant is well-suited for Pienso’s enterprise customers as it operates best on bare metal infrastructure, enabling them to maintain complete control over their data sovereignty and autonomous LLM regimes. This ensures that enterprises can maintain their full span of control while leveraging the scalability and performance benefits of Qdrant’s solution.
### [](https://qdrant.tech/case-studies/pienso/#efficiency-maximizing-the-customer-value-proposition)Efficiency: Maximizing the Customer Value Proposition ### [](/case-studies/pienso/#efficiency-maximizing-the-customer-value-proposition)Efficiency: Maximizing the Customer Value Proposition
Qdrant’s storage efficiency delivers cost savings on hardware while ensuring a responsive system even with extensive data sets. In an independent benchmark stress test, Pienso discovered that Qdrant could efficiently store 128 million documents, consuming a mere 20.4GB of storage and only 1.25GB of memory. This storage efficiency not only minimizes hardware expenses for Pienso’s customers, but also ensures optimal performance, making Qdrant an ideal solution for managing large-scale data with ease and efficiency. Qdrant’s storage efficiency delivers cost savings on hardware while ensuring a responsive system even with extensive data sets. In an independent benchmark stress test, Pienso discovered that Qdrant could efficiently store 128 million documents, consuming a mere 20.4GB of storage and only 1.25GB of memory. This storage efficiency not only minimizes hardware expenses for Pienso’s customers, but also ensures optimal performance, making Qdrant an ideal solution for managing large-scale data with ease and efficiency.
### [](https://qdrant.tech/case-studies/pienso/#reliability-fast-performance-in-a-secure-environment)Reliability: Fast Performance in a Secure Environment ### [](/case-studies/pienso/#reliability-fast-performance-in-a-secure-environment)Reliability: Fast Performance in a Secure Environment
Qdrant’s utilization of Rust, coupled with its memmap storage and write-ahead logging, offers users a powerful combination of high-performance operations, robust data protection, and enhanced data safety measures. Our memmap storage feature offers Pienso fast performance comparable to in-memory storage. In the context of machine learning, where rapid data access and retrieval are crucial for training and inference tasks, this capability proves invaluable. Furthermore, our write-ahead logging (WAL), is critical to ensuring changes are logged before being applied to the database. This approach adds additional layers of data safety, further safeguarding the integrity of the stored information. Qdrant’s utilization of Rust, coupled with its memmap storage and write-ahead logging, offers users a powerful combination of high-performance operations, robust data protection, and enhanced data safety measures. Our memmap storage feature offers Pienso fast performance comparable to in-memory storage. In the context of machine learning, where rapid data access and retrieval are crucial for training and inference tasks, this capability proves invaluable. Furthermore, our write-ahead logging (WAL), is critical to ensuring changes are logged before being applied to the database. This approach adds additional layers of data safety, further safeguarding the integrity of the stored information.
> “We chose Qdrant because it’s fast to query, has a small memory footprint and allows for instantaneous setup of a new vector collection that is going to be queried. Other solutions we evaluated had long bootstrap times and also long collection initialization times {..} This partnership comes at a great time, because it allows Pienso to use Qdrant to its maximum potential, giving our customers a seamless experience while they explore and get meaningful insights about their data.” - Felipe Balduino Cassar, Senior Software Engineer, Pienso > “We chose Qdrant because it’s fast to query, has a small memory footprint and allows for instantaneous setup of a new vector collection that is going to be queried. Other solutions we evaluated had long bootstrap times and also long collection initialization times {..} This partnership comes at a great time, because it allows Pienso to use Qdrant to its maximum potential, giving our customers a seamless experience while they explore and get meaningful insights about their data.” - Felipe Balduino Cassar, Senior Software Engineer, Pienso
## [](https://qdrant.tech/case-studies/pienso/#whats-next)What’s Next? ## [](/case-studies/pienso/#whats-next)What’s Next?
Pienso and Qdrant are dedicated to jointly develop the most reliable customer offering for the long term. Our partnership will deliver a combination of no-code/low-code interactive deep learning with efficient vector computation engineered for open source models and libraries. Pienso and Qdrant are dedicated to jointly develop the most reliable customer offering for the long term. Our partnership will deliver a combination of no-code/low-code interactive deep learning with efficient vector computation engineered for open source models and libraries.
### [](https://qdrant.tech/case-studies/pienso/#to-learn-more-about-how-we-plan-on-achieving-this-join-the-founders-for-a-technical-fireside-chat-at-0930-pst-thursday-20th-july-on-discordhttpsdiscordggvnvg3fheevent1128331722270969909)To learn more about how we plan on achieving this, join the founders for a [technical fireside chat at 09:30 PST Thursday, 20th July on Discord](https://discord.gg/Vnvg3fHE?event=1128331722270969909). ### [](/case-studies/pienso/#to-learn-more-about-how-we-plan-on-achieving-this-join-the-founders-for-a-technical-fireside-chat-at-0930-pst-thursday-20th-july-on-discordhttpsdiscordggvnvg3fheevent1128331722270969909)To learn more about how we plan on achieving this, join the founders for a [technical fireside chat at 09:30 PST Thursday, 20th July on Discord](https://discord.gg/Vnvg3fHE?event=1128331722270969909).
<!--EndFragment--> <!--EndFragment-->
@@ -28,7 +28,7 @@ What this means for you:
**"With Qdrant, we found the missing piece to develop our own provider independent multimodal generative AI platform at enterprise scale."** -- Jeremy Teichmann (AI Squad Technical Lead & Generative AI Expert), Daly Singh (AI Squad Lead & Product Owner) - Bosch Digital. **"With Qdrant, we found the missing piece to develop our own provider independent multimodal generative AI platform at enterprise scale."** -- Jeremy Teichmann (AI Squad Technical Lead & Generative AI Expert), Daly Singh (AI Squad Lead & Product Owner) - Bosch Digital.
Get started by [signing up for a Qdrant Cloud account](https://cloud.qdrant.io). And learn more about Qdrant Cloud in our [docs](https://qdrant.tech/documentation/cloud/). Get started by [signing up for a Qdrant Cloud account](https://cloud.qdrant.io). And learn more about Qdrant Cloud in our [docs](/documentation/cloud/).
<video autoplay="true" loop="true" width="100%" controls><source src="/blog/qdrant-cloud-on-azure/azure-cluster-deployment-short.mp4" type="video/mp4"></video> <video autoplay="true" loop="true" width="100%" controls><source src="/blog/qdrant-cloud-on-azure/azure-cluster-deployment-short.mp4" type="video/mp4"></video>
+1 -1
View File
@@ -20,7 +20,7 @@ Let's go through the process of building a workflow. We'll build a chat with a c
## Prerequisites ## Prerequisites
- A running Qdrant instance. If you need one, use our [Quick start guide](https://qdrant.tech/documentation/quick-start/) to set it up. - A running Qdrant instance. If you need one, use our [Quick start guide](/documentation/quick-start/) to set it up.
- An OpenAI API Key. Retrieve your key from the [OpenAI API page](https://platform.openai.com/account/api-keys) for your account. - An OpenAI API Key. Retrieve your key from the [OpenAI API page](https://platform.openai.com/account/api-keys) for your account.
- A GitHub access token. If you need to generate one, start at the [GitHub Personal access tokens page](https://github.com/settings/tokens/). - A GitHub access token. If you need to generate one, start at the [GitHub Personal access tokens page](https://github.com/settings/tokens/).
@@ -18,7 +18,7 @@ In this blog post, we'll demonstrate how to load data into Qdrant from the chann
### Prerequisites ### Prerequisites
- A running Qdrant instance. Refer to our [Quickstart guide](https://qdrant.tech/documentation/quick-start/) to set up an instance. - A running Qdrant instance. Refer to our [Quickstart guide](/documentation/quick-start/) to set up an instance.
- A Discord bot token. Generate one [here](https://discord.com/developers/applications) after adding the bot to your server. - A Discord bot token. Generate one [here](https://discord.com/developers/applications) after adding the bot to your server.
- Unstructured CLI with the required extras. For more information, see the Discord [Getting Started guide](https://discord.com/developers/docs/getting-started). Install it with the following command: - Unstructured CLI with the required extras. For more information, see the Discord [Getting Started guide](https://discord.com/developers/docs/getting-started). Install it with the following command:
@@ -30,7 +30,7 @@ We've compared how Qdrant performs against the other vector search engines to gi
Since the last time we ran our benchmarks, we received a bunch of suggestions on how to run other engines more efficiently, and we applied them. Since the last time we ran our benchmarks, we received a bunch of suggestions on how to run other engines more efficiently, and we applied them.
This has resulted in significant improvements across all engines. As a result, we have achieved an impressive improvement of nearly four times in certain cases. You can view the previous benchmark results [here](https://qdrant.tech/benchmarks/single-node-speed-benchmark-2022/). This has resulted in significant improvements across all engines. As a result, we have achieved an impressive improvement of nearly four times in certain cases. You can view the previous benchmark results [here](/benchmarks/single-node-speed-benchmark-2022/).
#### Introducing a New Dataset #### Introducing a New Dataset
@@ -65,7 +65,7 @@ We use the same benchmark datasets as the [ann-benchmarks](https://github.com/er
### Detailed Report and Access ### Detailed Report and Access
For an in-depth look at our latest benchmark results, we invite you to read the [detailed report](https://qdrant.tech/benchmarks). For an in-depth look at our latest benchmark results, we invite you to read the [detailed report](/benchmarks/).
If you're interested in testing the benchmark yourself or want to contribute to its development, head over to our [benchmark repository](https://github.com/qdrant/vector-db-benchmark). We appreciate your support and involvement in improving the performance of vector databases. If you're interested in testing the benchmark yourself or want to contribute to its development, head over to our [benchmark repository](https://github.com/qdrant/vector-db-benchmark). We appreciate your support and involvement in improving the performance of vector databases.
@@ -27,9 +27,9 @@ The rise of generative AI in the last few years has shone a spotlight on vector
## What sets Qdrant apart? ## What sets Qdrant apart?
To meet the needs of the next generation of AI applications, Qdrant has always been built with four keys in mind: efficiency, scalability, performance, and flexibility. Our goal is to give our users unmatched speed and reliability, even when they are building massive-scale AI applications requiring the handling of billions of vectors. We did so by building Qdrant on Rust for performance, memory safety, and scale. Additionally, [our custom HNSW search algorithm](https://qdrant.tech/articles/filtrable-hnsw/) and unique [filtering](https://qdrant.tech/documentation/concepts/filtering/) capabilities consistently lead to [highest RPS](https://qdrant.tech/benchmarks/), minimal latency, and high control with accuracy when running large-scale, high-dimensional operations. To meet the needs of the next generation of AI applications, Qdrant has always been built with four keys in mind: efficiency, scalability, performance, and flexibility. Our goal is to give our users unmatched speed and reliability, even when they are building massive-scale AI applications requiring the handling of billions of vectors. We did so by building Qdrant on Rust for performance, memory safety, and scale. Additionally, [our custom HNSW search algorithm](/articles/filtrable-hnsw/) and unique [filtering](/documentation/concepts/filtering/) capabilities consistently lead to [highest RPS](/benchmarks/), minimal latency, and high control with accuracy when running large-scale, high-dimensional operations.
Beyond performance, we provide our users with the most flexibility in cost savings and deployment options. A combination of cutting-edge efficiency features, like [built-in compression options](https://qdrant.tech/documentation/guides/quantization/), [multitenancy](https://qdrant.tech/documentation/guides/multiple-partitions/) and the ability to [offload data to disk](https://qdrant.tech/documentation/concepts/storage/), dramatically reduce memory consumption. Committed to privacy and security, crucial for modern AI applications, Qdrant now also offers on-premise and hybrid SaaS solutions, meeting diverse enterprise needs in a data-sensitive world. This approach, coupled with our open-source foundation, builds trust and reliability with engineers and developers, making Qdrant a game-changer in the vector database domain. Beyond performance, we provide our users with the most flexibility in cost savings and deployment options. A combination of cutting-edge efficiency features, like [built-in compression options](/documentation/guides/quantization/), [multitenancy](/documentation/guides/multiple-partitions/) and the ability to [offload data to disk](/documentation/concepts/storage/), dramatically reduce memory consumption. Committed to privacy and security, crucial for modern AI applications, Qdrant now also offers on-premise and hybrid SaaS solutions, meeting diverse enterprise needs in a data-sensitive world. This approach, coupled with our open-source foundation, builds trust and reliability with engineers and developers, making Qdrant a game-changer in the vector database domain.
## What's next? ## What's next?
@@ -137,8 +137,8 @@ In addition to the required options, you can also specify custom values for the
* `hnsw_config` - see [indexing](../indexing/#vector-index) for details. * `hnsw_config` - see [indexing](../indexing/#vector-index) for details.
* `wal_config` - Write-Ahead-Log related configuration. See more details about [WAL](../storage/#versioning) * `wal_config` - Write-Ahead-Log related configuration. See more details about [WAL](../storage/#versioning)
* `optimizers_config` - see [optimizer](../optimizer) for details. * `optimizers_config` - see [optimizer](../optimizer/) for details.
* `shard_number` - which defines how many shards the collection should have. See [distributed deployment](../../guides/distributed_deployment#sharding) section for details. * `shard_number` - which defines how many shards the collection should have. See [distributed deployment](../../guides/distributed_deployment/#sharding) section for details.
* `on_disk_payload` - defines where to store payload data. If `true` - payload will be stored on disk only. Might be useful for limiting the RAM usage in case of large payload. * `on_disk_payload` - defines where to store payload data. If `true` - payload will be stored on disk only. Might be useful for limiting the RAM usage in case of large payload.
* `quantization_config` - see [quantization](../../guides/quantization/#setting-up-quantization-in-qdrant) for details. * `quantization_config` - see [quantization](../../guides/quantization/#setting-up-quantization-in-qdrant) for details.
@@ -1167,7 +1167,7 @@ _Note: these numbers may be removed in a future version of Qdrant._
### Indexing vectors in HNSW ### Indexing vectors in HNSW
In some cases, you might be surprised the value of `indexed_vectors_count` is lower than `vectors_count`. This is an intended behaviour and In some cases, you might be surprised the value of `indexed_vectors_count` is lower than `vectors_count`. This is an intended behaviour and
depends on the [optimizer configuration](../optimizer). A new index segment is built if the size of non-indexed vectors is higher than the depends on the [optimizer configuration](../optimizer/). A new index segment is built if the size of non-indexed vectors is higher than the
value of `indexing_threshold`(in kB). If your collection is very small or the dimensionality of the vectors is low, there might be no HNSW segment value of `indexing_threshold`(in kB). If your collection is very small or the dimensionality of the vectors is low, there might be no HNSW segment
created and `indexed_vectors_count` might be equal to `0`. created and `indexed_vectors_count` might be equal to `0`.
@@ -7,7 +7,7 @@ aliases:
# Explore the data # Explore the data
After mastering the concepts in [search](../search), you can start exploring your data in other ways. Qdrant provides a stack of APIs that allow you to find similar vectors in a different fashion, as well as to find the most dissimilar ones. These are useful tools for recommendation systems, data exploration, and data cleaning. After mastering the concepts in [search](../search/), you can start exploring your data in other ways. Qdrant provides a stack of APIs that allow you to find similar vectors in a different fashion, as well as to find the most dissimilar ones. These are useful tools for recommendation systems, data exploration, and data cleaning.
## Recommendation API ## Recommendation API
@@ -8,7 +8,7 @@ aliases:
# Filtering # Filtering
With Qdrant, you can set conditions when searching or retrieving points. With Qdrant, you can set conditions when searching or retrieving points.
For example, you can impose conditions on both the [payload](../payload) and the `id` of the point. For example, you can impose conditions on both the [payload](../payload/) and the `id` of the point.
Setting additional conditions is important when it is impossible to express all the features of the object in the embedding. Setting additional conditions is important when it is impossible to express all the features of the object in the embedding.
Examples include a variety of business requirements: stock availability, user location, or desired price range. Examples include a variety of business requirements: stock availability, user location, or desired price range.
@@ -12,14 +12,14 @@ A key feature of Qdrant is the effective combination of vector and traditional i
The indexes in the segments exist independently, but the parameters of the indexes themselves are configured for the whole collection. The indexes in the segments exist independently, but the parameters of the indexes themselves are configured for the whole collection.
Not all segments automatically have indexes. Not all segments automatically have indexes.
Their necessity is determined by the [optimizer](../optimizer) settings and depends, as a rule, on the number of stored points. Their necessity is determined by the [optimizer](../optimizer/) settings and depends, as a rule, on the number of stored points.
## Payload Index ## Payload Index
Payload index in Qdrant is similar to the index in conventional document-oriented databases. Payload index in Qdrant is similar to the index in conventional document-oriented databases.
This index is built for a specific field and type, and is used for quick point requests by the corresponding filtering condition. This index is built for a specific field and type, and is used for quick point requests by the corresponding filtering condition.
The index is also used to accurately estimate the filter cardinality, which helps the [query planning](../search#query-planning) choose a search strategy. The index is also used to accurately estimate the filter cardinality, which helps the [query planning](../search/#query-planning) choose a search strategy.
Creating an index requires additional computational resources and memory, so choosing fields to be indexed is essential. Qdrant does not make this choice but grants it to the user. Creating an index requires additional computational resources and memory, so choosing fields to be indexed is essential. Qdrant does not make this choice but grants it to the user.
@@ -299,7 +299,7 @@ storage:
``` ```
And so in the process of creating a [collection](../collections). The `ef` parameter is configured during [the search](../search) and by default is equal to `ef_construct`. And so in the process of creating a [collection](../collections/). The `ef` parameter is configured during [the search](../search/) and by default is equal to `ef_construct`.
HNSW is chosen for several reasons. HNSW is chosen for several reasons.
First, HNSW is well-compatible with the modification that allows Qdrant to use filters during a search. First, HNSW is well-compatible with the modification that allows Qdrant to use filters during a search.
@@ -9,7 +9,7 @@ aliases:
It is much more efficient to apply changes in batches than perform each change individually, as many other databases do. Qdrant here is no exception. Since Qdrant operates with data structures that are not always easy to change, it is sometimes necessary to rebuild those structures completely. It is much more efficient to apply changes in batches than perform each change individually, as many other databases do. Qdrant here is no exception. Since Qdrant operates with data structures that are not always easy to change, it is sometimes necessary to rebuild those structures completely.
Storage optimization in Qdrant occurs at the segment level (see [storage](../storage)). Storage optimization in Qdrant occurs at the segment level (see [storage](../storage/)).
In this case, the segment to be optimized remains readable for the time of the rebuild. In this case, the segment to be optimized remains readable for the time of the rebuild.
![Segment optimization](/docs/optimization.svg) ![Segment optimization](/docs/optimization.svg)
@@ -91,6 +91,6 @@ storage:
indexing_threshold_kb: 20000 indexing_threshold_kb: 20000
``` ```
In addition to the configuration file, you can also set optimizer parameters separately for each [collection](../collections). In addition to the configuration file, you can also set optimizer parameters separately for each [collection](../collections/).
Dynamic parameter updates may be useful, for example, for more efficient initial loading of points. You can disable indexing during the upload process with these settings and enable it immediately after it is finished. As a result, you will not waste extra computation resources on rebuilding the index. Dynamic parameter updates may be useful, for example, for more efficient initial loading of points. You can disable indexing during the upload process with these settings and enable it immediately after it is finished. As a result, you will not waste extra computation resources on rebuilding the index.
@@ -50,7 +50,7 @@ For example, you will get an empty output if you apply the [range condition](../
However, arrays (multiple values of the same type) are treated a little bit different. When we apply a filter to an array, it will succeed if at least one of the values inside the array meets the condition. However, arrays (multiple values of the same type) are treated a little bit different. When we apply a filter to an array, it will succeed if at least one of the values inside the array meets the condition.
The filtering process is discussed in detail in the section [Filtering](../filtering). The filtering process is discussed in detail in the section [Filtering](../filtering/).
Let's look at the data types that Qdrant supports for searching: Let's look at the data types that Qdrant supports for searching:
@@ -959,7 +959,7 @@ await client.DeletePayloadAsync(
To search more efficiently with filters, Qdrant allows you to create indexes for payload fields by specifying the name and type of field it is intended to be. To search more efficiently with filters, Qdrant allows you to create indexes for payload fields by specifying the name and type of field it is intended to be.
The indexed fields also affect the vector index. See [Indexing](../indexing) for details. The indexed fields also affect the vector index. See [Indexing](../indexing/) for details.
In practice, we recommend creating an index on those fields that could potentially constrain the results the most. In practice, we recommend creating an index on those fields that could potentially constrain the results the most.
For example, using an index for the object ID will be much more efficient, being unique for each record, than an index by its color, which has only a few possible values. For example, using an index for the object ID will be much more efficient, being unique for each record, than an index by its color, which has only a few possible values.
@@ -8,10 +8,10 @@ aliases:
# Points # Points
The points are the central entity that Qdrant operates with. The points are the central entity that Qdrant operates with.
A point is a record consisting of a vector and an optional [payload](../payload). A point is a record consisting of a vector and an optional [payload](../payload/).
You can search among the points grouped in one [collection](../collections) based on vector similarity. You can search among the points grouped in one [collection](../collections/) based on vector similarity.
This procedure is described in more detail in the [search](../search) and [filtering](../filtering) sections. This procedure is described in more detail in the [search](../search/) and [filtering](../filtering/) sections.
This section explains how to create and manage vectors. This section explains how to create and manage vectors.
@@ -46,10 +46,10 @@ This process is called query planning.
The strategy selection process relies heavily on heuristics and can vary from release to release. The strategy selection process relies heavily on heuristics and can vary from release to release.
However, the general principles are: However, the general principles are:
* planning is performed for each segment independently (see [storage](../storage) for more information about segments) * planning is performed for each segment independently (see [storage](../storage/) for more information about segments)
* prefer a full scan if the amount of points is below a threshold * prefer a full scan if the amount of points is below a threshold
* estimate the cardinality of a filtered result before selecting a strategy * estimate the cardinality of a filtered result before selecting a strategy
* retrieve points using payload index (see [indexing](../indexing)) if cardinality is below threshold * retrieve points using payload index (see [indexing](../indexing/)) if cardinality is below threshold
* use filterable vector index if the cardinality is above a threshold * use filterable vector index if the cardinality is above a threshold
You can adjust the threshold using a [configuration file](https://github.com/qdrant/qdrant/blob/master/config/config.yaml), as well as independently for each collection. You can adjust the threshold using a [configuration file](https://github.com/qdrant/qdrant/blob/master/config/config.yaml), as well as independently for each collection.
@@ -211,7 +211,7 @@ Currently, it could be:
* `indexed_only` - With this option you can disable the search in those segments where vector index is not built yet. This may be useful if you want to minimize the impact to the search performance whilst the collection is also being updated. Using this option may lead to a partial result if the collection is not fully indexed yet, consider using it only if eventual consistency is acceptable for your use case. * `indexed_only` - With this option you can disable the search in those segments where vector index is not built yet. This may be useful if you want to minimize the impact to the search performance whilst the collection is also being updated. Using this option may lead to a partial result if the collection is not fully indexed yet, consider using it only if eventual consistency is acceptable for your use case.
Since the `filter` parameter is specified, the search is performed only among those points that satisfy the filter condition. Since the `filter` parameter is specified, the search is performed only among those points that satisfy the filter condition.
See details of possible filters and their work in the [filtering](../filtering) section. See details of possible filters and their work in the [filtering](../filtering/) section.
Example result of this API would be Example result of this API would be
@@ -17,7 +17,7 @@ For a step-by-step guide on how to use snapshots, see our [tutorial](/documentat
## Store snapshots ## Store snapshots
The target directory used to store generated snapshots is controlled through the [configuration](../../guides/configuration) or using the ENV variable: `QDRANT__STORAGE__SNAPSHOTS_PATH=./snapshots`. The target directory used to store generated snapshots is controlled through the [configuration](../../guides/configuration/) or using the ENV variable: `QDRANT__STORAGE__SNAPSHOTS_PATH=./snapshots`.
You can set the snapshots storage directory from the [config.yaml](https://github.com/qdrant/qdrant/blob/master/config/config.yaml) file. If no value is given, default is `./snapshots`. You can set the snapshots storage directory from the [config.yaml](https://github.com/qdrant/qdrant/blob/master/config/config.yaml) file. If no value is given, default is `./snapshots`.
@@ -13,7 +13,7 @@ Each segment has its independent vector and payload storage as well as indexes.
Data stored in segments usually do not overlap. Data stored in segments usually do not overlap.
However, storing the same point in different segments will not cause problems since the search contains a deduplication mechanism. However, storing the same point in different segments will not cause problems since the search contains a deduplication mechanism.
The segments consist of vector and payload storages, vector and payload [indexes](../indexing), and id mapper, which stores the relationship between internal and external ids. The segments consist of vector and payload storages, vector and payload [indexes](../indexing/), and id mapper, which stores the relationship between internal and external ids.
A segment can be `appendable` or `non-appendable` depending on the type of storage and index used. A segment can be `appendable` or `non-appendable` depending on the type of storage and index used.
You can freely add, delete and query data in the `appendable` segment. You can freely add, delete and query data in the `appendable` segment.
@@ -36,7 +36,7 @@ qdrant_client.upsert(
``` ```
If you are interested in seeing an end-to-end project created with co.embed API and Qdrant, please check out the If you are interested in seeing an end-to-end project created with co.embed API and Qdrant, please check out the
"[Question Answering as a Service with Cohere and Qdrant](https://qdrant.tech/articles/qa-with-cohere-and-qdrant/)" article. "[Question Answering as a Service with Cohere and Qdrant](/articles/qa-with-cohere-and-qdrant/)" article.
## Embed v3 ## Embed v3
@@ -13,7 +13,7 @@ Qdrant is available as a [provider](https://airflow.apache.org/docs/apache-airfl
Before configuring Airflow, you need: Before configuring Airflow, you need:
1. A Qdrant instance to connect to. You can set one up in our [installation guide](https://qdrant.tech/documentation/guides/installation). 1. A Qdrant instance to connect to. You can set one up in our [installation guide](/documentation/guides/installation/).
2. A running Airflow instance. You can use their [Quick Start Guide](https://airflow.apache.org/docs/apache-airflow/stable/start.html). 2. A running Airflow instance. You can use their [Quick Start Guide](https://airflow.apache.org/docs/apache-airflow/stable/start.html).
@@ -23,7 +23,7 @@ Before you use the following code sample, customize the following values for you
- `YOUR_QDRANT_REST_URL`: If you've set up Qdrant using the [Quick Start](/documentation/quick-start/) guide, - `YOUR_QDRANT_REST_URL`: If you've set up Qdrant using the [Quick Start](/documentation/quick-start/) guide,
set this value to `http://localhost:6333`. set this value to `http://localhost:6333`.
- `YOUR_COLLECTION_NAME`: Use our [Collections](/documentation/concepts/collections) guide to create or - `YOUR_COLLECTION_NAME`: Use our [Collections](/documentation/concepts/collections/) guide to create or
list collections. list collections.
```go ```go
@@ -25,7 +25,7 @@ Add the `langchain4j-qdrant` to your project dependencies.
Before you use the following code sample, customize the following values for your configuration: Before you use the following code sample, customize the following values for your configuration:
- `YOUR_COLLECTION_NAME`: Use our [Collections](/documentation/concepts/collections) guide to create or - `YOUR_COLLECTION_NAME`: Use our [Collections](/documentation/concepts/collections/) guide to create or
list collections. list collections.
- `YOUR_HOST_URL`: Use the GRPC URL for your system. If you used the [Quick Start](/documentation/quick-start/) guide, - `YOUR_HOST_URL`: Use the GRPC URL for your system. If you used the [Quick Start](/documentation/quick-start/) guide,
it may be http://localhost:6334. If you've deployed in the [Qdrant Cloud](/documentation/cloud/), you may have a it may be http://localhost:6334. If you've deployed in the [Qdrant Cloud](/documentation/cloud/), you may have a
@@ -25,7 +25,7 @@ Before you start, make sure you have the following:
Navigate to your scenario on the Make dashboard and select a Qdrant app module to start a connection. Navigate to your scenario on the Make dashboard and select a Qdrant app module to start a connection.
![Qdrant Make connection](/documentation/frameworks/make/connection.png) ![Qdrant Make connection](/documentation/frameworks/make/connection.png)
You can now establish a connection to Qdrant using your [instance credentials](https://qdrant.tech/documentation/cloud/authentication/). You can now establish a connection to Qdrant using your [instance credentials](/documentation/cloud/authentication/).
![Qdrant Make form](/documentation/frameworks/make/connection-form.png) ![Qdrant Make form](/documentation/frameworks/make/connection-form.png)
@@ -24,7 +24,7 @@ You can now configure the vectorstore node according to your workflow requiremen
![Qdrant Config](/documentation/frameworks/n8n/config.png) ![Qdrant Config](/documentation/frameworks/n8n/config.png)
Create a connection to Qdrant using your [instance credentials](https://qdrant.tech/documentation/cloud/authentication/). Create a connection to Qdrant using your [instance credentials](/documentation/cloud/authentication/).
![Qdrant Credentials](/documentation/frameworks/n8n/credentials.png) ![Qdrant Credentials](/documentation/frameworks/n8n/credentials.png)
@@ -18,7 +18,7 @@ production mode, you could also choose to overwrite `config/production.yaml`.
See [ordering](#order-and-priority) for details on how configurations are See [ordering](#order-and-priority) for details on how configurations are
loaded. loaded.
The [Installation](../installation) guide contains examples of how to set up Qdrant with a custom configuration for the different deployment methods. The [Installation](../installation/) guide contains examples of how to set up Qdrant with a custom configuration for the different deployment methods.
## Order and priority ## Order and priority
@@ -11,7 +11,7 @@ aliases:
Since version v0.8.0 Qdrant supports a distributed deployment mode. Since version v0.8.0 Qdrant supports a distributed deployment mode.
In this mode, multiple Qdrant services communicate with each other to distribute the data across the peers to extend the storage capabilities and increase stability. In this mode, multiple Qdrant services communicate with each other to distribute the data across the peers to extend the storage capabilities and increase stability.
To enable distributed deployment - enable the cluster mode in the [configuration](../configuration) or using the ENV variable: `QDRANT__CLUSTER__ENABLED=true`. To enable distributed deployment - enable the cluster mode in the [configuration](../configuration/) or using the ENV variable: `QDRANT__CLUSTER__ENABLED=true`.
```yaml ```yaml
cluster: cluster:
@@ -529,7 +529,7 @@ on the size and state of a shard.
Available shard transfer methods are: Available shard transfer methods are:
- `stream_records`: _(default)_ transfer shard by streaming just its records to the target node in batches. - `stream_records`: _(default)_ transfer shard by streaming just its records to the target node in batches.
- `snapshot`: transfer shard including its index and quantized data by utilizing a [snapshot](../../concepts/snapshots) automatically. - `snapshot`: transfer shard including its index and quantized data by utilizing a [snapshot](../../concepts/snapshots/) automatically.
Each has pros, cons and specific requirements, which are: Each has pros, cons and specific requirements, which are:
@@ -579,7 +579,7 @@ the above cons are acceptable in your use case. If your cluster is unstable and
out of resources, it's probably best to use the `stream_records` transfer out of resources, it's probably best to use the `stream_records` transfer
method, because it is unlikely to fail. method, because it is unlikely to fail.
The `snapshot` transfer method utilizes [snapshots](../../concepts/snapshots) to The `snapshot` transfer method utilizes [snapshots](../../concepts/snapshots/) to
transfer a shard. A snapshot is created automatically. It is then transferred transfer a shard. A snapshot is created automatically. It is then transferred
and restored on the target node. After this is done, the snapshot is removed and restored on the target node. After this is done, the snapshot is removed
from both nodes. While the snapshot/transfer/restore operation is happening, the from both nodes. While the snapshot/transfer/restore operation is happening, the
@@ -60,7 +60,7 @@ For production, we recommend that you configure Qdrant in the cloud, with Kubern
### Qdrant Cloud ### Qdrant Cloud
You can set up production with the [Qdrant Cloud](https://qdrant.to/cloud), which provides fully managed Qdrant databases. You can set up production with the [Qdrant Cloud](https://qdrant.to/cloud), which provides fully managed Qdrant databases.
It provides horizontal and vertical scaling, one click installation and upgrades, monitoring, logging, as well as backup and disaster recovery. For more information, see the [Qdrant Cloud documentation](/documentation/cloud). It provides horizontal and vertical scaling, one click installation and upgrades, monitoring, logging, as well as backup and disaster recovery. For more information, see the [Qdrant Cloud documentation](/documentation/cloud/).
### Kubernetes ### Kubernetes
@@ -66,7 +66,7 @@ Qdrant server.
These currently provide the most basic status response, returning HTTP 200 if These currently provide the most basic status response, returning HTTP 200 if
Qdrant is started and ready to be used. Qdrant is started and ready to be used.
Regardless of whether an [API key](../security#authentication) is configured, Regardless of whether an [API key](../security/#authentication) is configured,
the endpoints are always accessible. the endpoints are always accessible.
You can read more about Kubernetes health endpoints You can read more about Kubernetes health endpoints
@@ -43,7 +43,7 @@ export QDRANT__SERVICE__API_KEY=your_secret_api_key_here
<aside role="alert"><a href="#tls">TLS</a> must be used to prevent leaking the API key over an unencrypted connection.</aside> <aside role="alert"><a href="#tls">TLS</a> must be used to prevent leaking the API key over an unencrypted connection.</aside>
For using API key based authentication in Qdrant cloud see the cloud For using API key based authentication in Qdrant cloud see the cloud
[Authentication](https://qdrant.tech/documentation/cloud/authentication) [Authentication](/documentation/cloud/authentication/)
section. section.
The API key then needs to be present in all REST or gRPC requests to your instance. The API key then needs to be present in all REST or gRPC requests to your instance.
@@ -7,7 +7,7 @@ aliases:
# Introduction # Introduction
![qdrant](https://qdrant.tech/images/logo_with_text.png) ![qdrant](/images/logo_with_text.png)
Vector databases are a relatively new way for interacting with abstract data representations Vector databases are a relatively new way for interacting with abstract data representations
derived from opaque machine learning models such as deep learning architectures. These derived from opaque machine learning models such as deep learning architectures. These
@@ -8,7 +8,7 @@ social_preview_image: /docs/gettingstarted/vector-social.png
If you are still trying to figure out how vector search works, please read ahead. This document describes how vector search is used, covers Qdrant's place in the larger ecosystem, and outlines how you can use Qdrant to augment your existing projects. If you are still trying to figure out how vector search works, please read ahead. This document describes how vector search is used, covers Qdrant's place in the larger ecosystem, and outlines how you can use Qdrant to augment your existing projects.
For those who want to start writing code right away, visit our [Complete Beginners tutorial](/documentation/tutorials/search-beginners) to build a search engine in 5-15 minutes. For those who want to start writing code right away, visit our [Complete Beginners tutorial](/documentation/tutorials/search-beginners/) to build a search engine in 5-15 minutes.
## A Brief History of Search ## A Brief History of Search
@@ -60,11 +60,11 @@ While doing a semantic search at scale, because this is what we sometimes call t
Vector search is an exciting alternative to sparse methods. It solves the issues we had with the keyword-based search without needing to maintain lots of heuristics manually. It requires an additional component, a neural encoder, to convert text into vectors. Vector search is an exciting alternative to sparse methods. It solves the issues we had with the keyword-based search without needing to maintain lots of heuristics manually. It requires an additional component, a neural encoder, to convert text into vectors.
[**Tutorial 1 - Qdrant for Complete Beginners**](../../tutorials/search-beginners) [**Tutorial 1 - Qdrant for Complete Beginners**](/documentation/tutorials/search-beginners/)
Despite its complicated background, vectors search is extraordinarily simple to set up. With Qdrant, you can have a search engine up-and-running in five minutes. Our [Complete Beginners tutorial](../../tutorials/search-beginners) will show you how. Despite its complicated background, vectors search is extraordinarily simple to set up. With Qdrant, you can have a search engine up-and-running in five minutes. Our [Complete Beginners tutorial](../../tutorials/search-beginners/) will show you how.
[**Tutorial 2 - Question and Answer System**](../../../articles/qa-with-cohere-and-qdrant) [**Tutorial 2 - Question and Answer System**](/articles/qa-with-cohere-and-qdrant/)
However, you can also choose SaaS tools to generate them and avoid building your model. Setting up a vector search project with Qdrant Cloud and Cohere co.embed API is fairly easy if you follow the [Question and Answer system tutorial](../../../articles/qa-with-cohere-and-qdrant). However, you can also choose SaaS tools to generate them and avoid building your model. Setting up a vector search project with Qdrant Cloud and Cohere co.embed API is fairly easy if you follow the [Question and Answer system tutorial](/articles/qa-with-cohere-and-qdrant/).
There is another exciting thing about vector search. You can search for any kind of data as long as there is a neural network that would vectorize your data type. Do you think about a reverse image search? That’s also possible with vector embeddings. There is another exciting thing about vector search. You can search for any kind of data as long as there is a neural network that would vectorize your data type. Do you think about a reverse image search? That’s also possible with vector embeddings.
@@ -517,7 +517,7 @@ version: 1
``` ```
The results are returned in decreasing similarity order. Note that payload and vector data is missing in these results by default. The results are returned in decreasing similarity order. Note that payload and vector data is missing in these results by default.
See [payload and vector in the result](../concepts/search#payload-and-vector-in-the-result) on how to enable it. See [payload and vector in the result](../concepts/search/#payload-and-vector-in-the-result) on how to enable it.
## Add a filter ## Add a filter
@@ -43,7 +43,7 @@ similar files for given query.
There are two things you need to set up before you start: There are two things you need to set up before you start:
1. You need to have a Qdrant instance running. If you want to launch it locally, 1. You need to have a Qdrant instance running. If you want to launch it locally,
[Docker is the fastest way to do that](https://qdrant.tech/documentation/quick_start/#installation). [Docker is the fastest way to do that](/documentation/quick_start/#installation).
2. You need to have a registered [Aleph Alpha account](https://app.aleph-alpha.com/). 2. You need to have a registered [Aleph Alpha account](https://app.aleph-alpha.com/).
3. Upon registration, create an API key (see: [API Tokens](https://app.aleph-alpha.com/profile)). 3. Upon registration, create an API key (see: [API Tokens](https://app.aleph-alpha.com/profile)).
@@ -254,7 +254,7 @@ pip install qdrant-client
``` ```
Of course, we need a running Qdrant server for vector search. If you need one, Of course, we need a running Qdrant server for vector search. If you need one,
you can [use a local Docker container](https://qdrant.tech/documentation/quick-start/) you can [use a local Docker container](/documentation/quick-start/)
or deploy it using the [Qdrant Cloud](https://cloud.qdrant.io/). or deploy it using the [Qdrant Cloud](https://cloud.qdrant.io/).
You can use either to follow this tutorial. Configure the connection parameters: You can use either to follow this tutorial. Configure the connection parameters:
@@ -17,7 +17,7 @@ This tutorial will show you how to create a snapshot of a collection and restore
## Prerequisites ## Prerequisites
Let's assume you already have a running Qdrant instance or a cluster. If not, you can follow the [installation guide](/documentation/guides/installation) to set up a local Qdrant instance or use [Qdrant Cloud](https://cloud.qdrant.io/) to create a cluster in a few clicks. Let's assume you already have a running Qdrant instance or a cluster. If not, you can follow the [installation guide](/documentation/guides/installation/) to set up a local Qdrant instance or use [Qdrant Cloud](https://cloud.qdrant.io/) to create a cluster in a few clicks.
Once the cluster is running, let's install the required dependencies: Once the cluster is running, let's install the required dependencies:
@@ -19,7 +19,7 @@ Fastembed natively integrates with Qdrant client, so you can easily upload the d
<aside role="status"> <aside role="status">
There is a version of this tutorial that uses <a href="https://www.sbert.net/">SentenceTransformers</a> model inference engine instead of Fastembed. There is a version of this tutorial that uses <a href="https://www.sbert.net/">SentenceTransformers</a> model inference engine instead of Fastembed.
Check it out <a href="/documentation/tutorials/neural-search">here</a>. Check it out <a href="/documentation/tutorials/neural-search/">here</a>.
</aside> </aside>
@@ -129,7 +129,7 @@ qdrant_client.recreate_collection(
Note, that we use `get_fastembed_vector_params` to get the vector size and distance function from the model. Note, that we use `get_fastembed_vector_params` to get the vector size and distance function from the model.
This method automatically generates configuration, compatible with the model you are using. This method automatically generates configuration, compatible with the model you are using.
Without fastembed integration, you would need to specify the vector size and distance function manually. Read more about it [here](/documentation/tutorials/neural-search). Without fastembed integration, you would need to specify the vector size and distance function manually. Read more about it [here](/documentation/tutorials/neural-search/).
Additionally, you can specify extended configuration for our vectors, like `quantization_config` or `hnsw_config`. Additionally, you can specify extended configuration for our vectors, like `quantization_config` or `hnsw_config`.
@@ -14,7 +14,7 @@ A neural search service uses artificial neural networks to improve the accuracy
<aside role="status"> <aside role="status">
There is a version of this tutorial that uses <a href="https://github.com/qdrant/fastembed">Fastembed</a> model inference engine instead of Sentence Transformers. There is a version of this tutorial that uses <a href="https://github.com/qdrant/fastembed">Fastembed</a> model inference engine instead of Sentence Transformers.
Check it out <a href="/documentation/tutorials/neural-search-fastembed">here</a>. Check it out <a href="/documentation/tutorials/neural-search-fastembed/">here</a>.
</aside> </aside>
+2 -2
View File
@@ -78,7 +78,7 @@ You accept these conditions and agree not to use the Solution or its content for
### 5. Financial terms ### 5. Financial terms
The prices applicable at the date of subscription to the Solution are accessible through the following [link](https://qdrant.com/pricing). The prices applicable at the date of subscription to the Solution are accessible through the following [link](/pricing/).
Unless otherwise stated, prices are in dollars and exclusive of any applicable taxes. Unless otherwise stated, prices are in dollars and exclusive of any applicable taxes.
@@ -179,7 +179,7 @@ In the event of a violation of any provision of these T&Cs or, more generally, i
In the context of the use of the Solution and the Website, Qdrant may collect and process certain personal data, including your name, surname, email address, banking information, address, telephone number, IP address, connection, and navigation data and data recorded in cookies (the “Data”). In the context of the use of the Solution and the Website, Qdrant may collect and process certain personal data, including your name, surname, email address, banking information, address, telephone number, IP address, connection, and navigation data and data recorded in cookies (the “Data”).
Qdrant ensures that the Data is collected and processed in compliance with the provisions of German law and in accordance with its Privacy Policy, available at the following [link](https://qdrant.tech/legal/privacy-policy). Qdrant ensures that the Data is collected and processed in compliance with the provisions of German law and in accordance with its Privacy Policy, available at the following [link](/legal/privacy-policy/).
The Privacy Policy is an integral part of the T&Cs. The Privacy Policy is an integral part of the T&Cs.
@@ -89,7 +89,7 @@
<section class="qb-related row clearfix"> <section class="qb-related row clearfix">
<header class="col-12 qb-related__title-outer"> <header class="col-12 qb-related__title-outer">
<h2 class="qb-related__title">Related Posts</h2> <h2 class="qb-related__title">Related Posts</h2>
<a class="qb-related__link" href="/blog">View all blog posts <a class="qb-related__link" href="/blog/">View all blog posts
<i> <i>
<svg width="14" height="10" viewBox="0 0 14 10" fill="none" xmlns="http://www.w3.org/2000/svg"> <svg width="14" height="10" viewBox="0 0 14 10" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M1 5H13" stroke="#102252" stroke-width="2" stroke-linecap="round" <path d="M1 5H13" stroke="#102252" stroke-width="2" stroke-linecap="round"
@@ -5,7 +5,7 @@
{{ $len := sub (len $paths) 2 }} {{ $len := sub (len $paths) 2 }}
{{ range first $len $paths }} {{ range first $len $paths }}
{{ if gt (len . ) 0 }} {{ if gt (len . ) 0 }}
{{ $rellink = printf "%s/%s" $rellink . }} {{ $rellink = printf "%s/%s/" $rellink . }}
<li class="breadcrumbs__crumb">/</li> <li class="breadcrumbs__crumb">/</li>
<li class="breadcrumbs__crumb"><a href="{{$rellink}}">{{ humanize . }}</a></li> <li class="breadcrumbs__crumb"><a href="{{$rellink}}">{{ humanize . }}</a></li>
{{ end }} {{ end }}