Merge branch 'master' of https://github.com/qdrant/landing_page into vst-deasylabs

This commit is contained in:
sabrinaaquino
2025-02-21 19:17:13 -03:00
54 changed files with 594 additions and 58 deletions
+11
View File
@@ -128,6 +128,12 @@ You can install `cwebp` with the following command:
curl -s https://raw.githubusercontent.com/Intervox/node-webp/latest/bin/install_webp | sudo bash
```
For **macOS**, you'll have to install `coreutils` too.
```
brew install coreutils
```
#### Prepare preview image
For the preview use an image with an aspect ratio of 3 to 1 in JPG or PNG format. With a resolution not smaller than 1200x630px. The image should illustrate in some way the article's core idea. Fill free got creative. Check out that the most important part of the image is in the center.
@@ -148,6 +154,11 @@ bash -x automation/process-article-img.sh ~/Pictures/my_preview.jpg filtrable-hn
This command will create a directory `preview` in `static/article_data/filtrable-hnsw` and generate preview images in it. If the directory `static/article_data/filtrable-hnsw` doesn't exist, it will be created. If it exists, only files in the children `preview` directory will be affected. In this case, preview images will be overwritten. Your original image will not be affected.
For **macOS** you'll have to make 2 adjustements to `process-img.sh` script which is run by `process-article-img.sh` script:
1. Exchange `stat -c %Y` with `stat -f %m`;
2. Exchange `realpath` with `grealpath`.
#### Preview images set
Preview images set consists of the following images:
+2 -2
View File
@@ -1,6 +1,6 @@
---
title: "AI Agents"
description: "AI Agents"
title: "AI Agents with Qdrant"
description: "AI agents powered by Qdrant leverage advanced vector search to access and retrieve high-dimensional data in real-time, enabling intelligent, Agentic-RAG driven, multi-step decision-making across dynamic environments."
build:
render: always
cascade:
@@ -215,7 +215,6 @@ If you're working with OpenAI or Cohere embeddings, we recommend the following o
|OpenAI text-embedding-3-large|3072|[DBpedia 1M](https://huggingface.co/datasets/Qdrant/dbpedia-entities-openai3-text-embedding-3-large-3072-1M) | 0.9966|3x|
|OpenAI text-embedding-3-small|1536|[DBpedia 100K](https://huggingface.co/datasets/Qdrant/dbpedia-entities-openai3-text-embedding-3-small-1536-100K)| 0.9847|3x|
|OpenAI text-embedding-3-large|1536|[DBpedia 1M](https://huggingface.co/datasets/Qdrant/dbpedia-entities-openai3-text-embedding-3-large-1536-1M)| 0.9826|3x|
|Cohere AI embed-english-v2.0|4096|[Wikipedia](https://huggingface.co/datasets/nreimers/wikipedia-22-12-large/tree/main) 1M|0.98|2x|
|OpenAI text-embedding-ada-002|1536|[DbPedia 1M](https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M) |0.98|4x|
|Gemini|768|No Open Data| 0.9563|3x|
|Mistral Embed|768|No Open Data| 0.9445 |3x|
@@ -0,0 +1,231 @@
---
title: "Built for Vector Search"
short_description: "Why add-on vector search looks good — until you actually use it."
description: "Why add-on vector search looks good — until you actually use it."
social_preview_image: /articles_data/dedicated-vector-search/preview/social_preview.jpg
preview_dir: /articles_data/dedicated-vector-search/preview
weight: -170
author: Evgeniya Sukhodolskaya & Andrey Vasnetsov
date: 2025-02-17T10:00:00+03:00
draft: false
keywords:
- system architecture
- vector search
- vector database
category: qdrant-internals
---
Any problem with even a bit of complexity requires a specialized solution. You can use a Swiss Army knife to open a bottle or poke a hole in a cardboard box, but you will need an axe to chop wood — the same goes for software.
In this article, we will describe the unique challenges vector search poses and why a dedicated solution is the best way to tackle them.
## Vectors
![vectors](/articles_data/dedicated-vector-search/image1.jpg)
Let's look at the central concept of vector databases — [**vectors**](/documentation/concepts/vectors/).
Vectors (also known as embeddings) are high-dimensional representations of various data points — texts, images, videos, etc. Many state-of-the-art (SOTA) embedding models generate representations of over 1,500 dimensions. When it comes to state-of-the-art PDF retrieval, the representations can reach [**over 100,000 dimensions per page**](/documentation/advanced-tutorials/pdf-retrieval-at-scale/).
This brings us to the first challenge of vector search — vectors are heavy.
### Vectors are Heavy
To put this in perspective, consider one million records stored in a relational database. It's a relatively small amount of data for modern databases, which a free tier of many cloud providers could easily handle.
Now, generate a 1536-dimensional embedding with OpenAI's `text-embedding-ada-002` model from each record, and you are looking at around **6GB of storage**. As a result, vector search workloads, especially if not optimized, will quickly dominate the main use cases of a non-vector database.
Having vectors as a part of a main database is a potential issue for another reason — vectors are always a transformation of other data.
### Vectors are a Transformation
Vectors are obtained from some other source-of-truth data. They can be restored if lost with the same embedding model previously used. At the same time, even small changes in that model can shift the geometry of the vector space, so if you update or change the embedding model, you need to update and reindex all the data to maintain accurate vector comparisons.
If coupled with the main database, this update process can lead to significant complications and even unavailability of the whole system.
<aside role="status">
Decouple vector workloads even if you plan to use a general-purpose database for vectors.
</aside>
However, vectors have positive properties as well. One of the most important is that vectors are fixed-size.
### Vectors are Fixed-Size
Embedding models are designed to produce vectors of a fixed size. We have to use it to our advantage.
For fast search, vectors need to be instantly accessible. Whether in [**RAM or disk**](/documentation/concepts/storage/), vectors should be stored in a format that allows quick access and comparison. This is essential, as vector comparison is a very hot operation in vector search workloads. It is often performed thousands of times per search query, so even a small overhead can lead to a significant slowdown.
For dedicated storage, vectors' fixed size comes as a blessing. Knowing how much space one data point needs, we don't have to deal with the usual overhead of locating data — the location of elements in storage is straightforward to calculate.
Everything becomes far less intuitive if vectors are stored together with other data types, for example, texts or JSONs. The size of a single data point is not fixed anymore, so accessing it becomes non-trivial, especially if data is added, updated, and deleted over time.
{{<figure src=/articles_data/dedicated-vector-search/dedicated_storage.png caption="Fixed size columns VS Variable length table" width=80% >}}
**Storing vectors together with other types of data, we lose all the benefits of their characteristics**; however, we fully "enjoy" their drawbacks, polluting the storage with an extremely heavy transformation of data already existing in that storage.
## Vector Search
![vector-search](/articles_data/dedicated-vector-search/image2.jpg)
Unlike traditional databases that serve as data stores, **vector databases are more like search engines**. They are designed to be **scalable**, always **available**, and capable of delivering high-speed search results even under heavy loads. Just as Google or Bing can handle billions of queries at once, vector databases are designed for scenarios where rapid, high-throughput, low-latency retrieval is a must.
{{<figure src=/articles_data/dedicated-vector-search/compass.png caption="Database Compass" width=80% >}}
### Pick Any Two
Distributed systems are perfect for scalability — horizontal scaling in these systems allows you to add more machines as needed. In the world of distributed systems, one well-known principle — the **CAP theorem** — illustrates that you cannot have it all. The theorem states that a distributed system can guarantee only two out of three properties: **Consistency**, **Availability**, and **Partition Tolerance**.
As network partitions are inevitable in any real-world distributed system, all modern distributed databases are designed with partition tolerance in mind, forcing a trade-off between **consistency** (providing the most up-to-date data) and **availability** (remaining responsive).
<aside role="status">
<strong>CP systems</strong> are still available to clients under normal operation — they prioritize data correctness over availability during failures. <br/>
<strong>AP systems</strong> deliver quick responses by relaxing immediate consistency guarantees but eventually converge to a correct state.
</aside>
There are two main design philosophies for databases in this context:
### ACID: Prioritizing Consistency
The ACID model ensures that every transaction (a group of operations treated as a single unit, such as transferring money between accounts) is executed fully or not at all (reverted), leaving the database in a valid state. When a system is distributed, achieving ACID properties requires complex coordination between nodes. Each node must communicate and agree on the state of a transaction, which can **limit system availability** — if a node is uncertain about the state of another, it may refuse to process a transaction until consistency is assured. This coordination also makes **scaling more challenging**.
Financial institutions use ACID-compliant databases when dealing with money transfers, where even a momentary discrepancy in an account balance is unacceptable.
### BASE: Prioritizing Availability
On the other hand, the BASE model favors high availability and partition tolerance. BASE systems distribute data and workload across multiple nodes, enabling them to respond to read and write requests immediately. They operate under the principle of **eventual consistency** — although data may be temporarily out-of-date, the system will converge on a consistent state given time.
Social media platforms, streaming services, and search engines all benefit from the BASE approach. For these applications, having immediate responsiveness is more critical than strict consistency.
### BASEd Vector Search
Considering the specifics of vector search — its nature demanding availability & scalability — it should be served on BASE-oriented architecture. This choice is made due to the need for horizontal scaling, high availability, low latency, and high throughput. For example, having BASE-focused architecture allows us to [**easily manage resharding**](/documentation/cloud/cluster-scaling/#resharding).
A strictly consistent transactional approach also loses its attractiveness when we remember that vectors are heavy transformations of data at our disposal — what's the point in limiting data protection mechanisms if we can always restore vectorized data through a transformation?
## Vector Index
![vector-index](/articles_data/dedicated-vector-search/image3.jpg)
[**Vector search**](/documentation/concepts/search/) relies on high-dimensional vector mathematics, making it computationally heavy at scale. A brute-force similarity search would require comparing a query against every vector in the database. In a database with 100 million 1536-dimensional vectors, performing 100 million comparisons per one query is unfeasible for production scenarios. Instead of a brute-force approach, vector databases have specialized approximate nearest neighbour (ANN) indexes that balance search precision and speed. These indexes require carefully designed architectures to make their maintenance in production feasible.
{{< figure src=/articles_data/dedicated-vector-search/hnsw.png caption="HNSW Index" width=80% >}}
One of the most popular vector indexes is **HNSW (Hierarchical Navigable Small World)**, which we picked for its capability to provide simultaneously high search speed and accuracy. High performance came with a cost — implementing it in production is untrivial due to several challenges, so to make it shine all the system's architecture has to be structured around it, serving the capricious index.
### Index Complexity
[**HNSW**](/documentation/concepts/indexing/) is structured as a multi-layered graph. With a new data point inserted, the algorithm must compare it to existing nodes across several layers to index it. As the number of vectors grows, these comparisons will noticeably slow down the construction process, making updates increasingly time-consuming. The indexing operation can quickly become the bottleneck in the system, slowing down search requests.
Building an HNSW monolith means limiting the scalability of your solution — its size has to be capped, as its construction time scales **non-linearly** with the number of elements. To keep the construction process feasible and ensure it doesn't affect the search time, we came up with a layered architecture that breaks down all data management into small units called **segments**.
{{<figure src=/articles_data/dedicated-vector-search/segments.png caption="Storage structure" width=80% >}}
Each segment isolates a subset of vectorized corpora and supports all collection-level operations on it, from searching to indexing, for example segments build their own index on the subset of data available to them. For users working on a collection level, the specifics of segmentation are unnoticeable. The search results they get span the whole collection, as sub-results are gathered from segments and then merged & deduplicated.
By balancing between size and number of segments, we can ensure the right balance between search speed and indexing time, making the system flexible for different workloads.
### Immutability
With index maintenance divided between segments, Qdrant can ensure high performance even during heavy load, and additional optimizations secure that further. These optimizations come from an idea that working with immutable structures introduces plenty of benefits: the possibility of using internally fixed sized lists (so no dynamic updates), ordering stored data accordingly to access patterns (so no unpredictable random accesses). With this in mind, to optimize search speed and memory management further, we use a strategy that combines and manages [**mutable and immutable segments**](/articles/immutable-data-structures/).
| | |
|---------------------|-------------|
| **Mutable Segments** | These are used for quickly ingesting new data and handling changes (updates) to existing data. |
| **Immutable Segments** | Once a mutable segment reaches a certain size, an optimization process converts it into an immutable segment, constructing an HNSW index – you could [**read about these optimizers here**](/documentation/concepts/optimizer/#optimizer) in detail. This immutability trick allowed us, for example, to ensure effective [**tenant isolation**](/documentation/concepts/indexing/#tenant-index). |
Immutable segments are an implementation detail transparent for users — they can delete vectors at any time, while additions and updates are applied to a mutable segment instead. This combination of mutability and immutability allows search and indexing to smoothly run simultaneously, even under heavy loads. This approach minimizes the performance impact of indexing time and allows on-the-fly configuration changes on a collection level (such as enabling or disabling data quantization) without downtimes.
### Filterable Index
Vector search wasn't historically designed for filtering — imposing strict constraints on results. It's inherently fuzzy; every document is, to some extent, both similar and dissimilar to any query — there's no binary "*fits/doesn't fit*" segregation. As a result, vector search algorithms weren't originally built with filtering in mind.
At the same time, filtering is unavoidable in many vector search applications, such as [**e-commerce search/recommendations**](/recommendations/). Searching for a Christmas present, you might want to filter out everything over 100 euros while still benefiting from the vector search's semantic nature.
In many vector search solutions, filtering is approached in two ways: **pre-filtering** (computes a binary mask for all vectors fitting the condition before running HNSW search) or **post-filtering** (running HNSW as usual and then filtering the results).
| | | |
|----|------------------|---------|
| ❌ | **Pre-filtering** | Has the linear complexity of computing the vector mask and becomes a bottleneck for large datasets. |
| ❌ | **Post-filtering** | The problem with **post-filtering** is tied to vector search "*everything fits and doesn't at the same time*" nature: imagine a low-cardinality filter that leaves only a few matching elements in the database. If none of them are similar enough to the query to appear in the top-X retrieved results, they'll all be filtered out. |
Qdrant [**took filtering in vector search further**](/articles/vector-search-filtering/), recognizing the limitations of pre-filtering & post-filtering strategies. We developed an adaptation of HNSW — [**filterable HNSW**](/articles/filtrable-hnsw/) — that also enables **in-place filtering** during graph traversal. To make this possible, we condition HNSW index construction on possible filtering conditions reflected by [**payload indexes**](/documentation/concepts/indexing/#payload-index) (inverted indexes built on vectors' [**metadata**](/documentation/concepts/payload/)).
**Qdrant was designed with a vector index being a central component of the system.** That made it possible to organize optimizers, payload indexes and other components around the vector index, unlocking the possibility of building a filterable HNSW.
{{<figure src=/articles_data/dedicated-vector-search/filterable-vector-index.png caption="Filterable Vector Index" width=80% >}}
In general, optimizing vector search requires a custom, finely tuned approach to data and index management that secures high performance even as data grows and changes dynamically. This specialized architecture is the key reason why **dedicated vector databases will always outperform general-purpose databases in production settings**.
## Vector Search Beyond RAG
{{<figure src=/articles_data/dedicated-vector-search/venn-diagram.png caption="Vector Search is not Text Search Extension" width=80% >}}
Many discussions about the purpose of vector databases focus on Retrieval-Augmented Generation (RAG) — or its more advanced variant, agentic RAG — where vector databases are used as a knowledge source to retrieve context for large language models (LLMs). This is a legitimate use case, however, the hype wave of RAG solutions has overshadowed the broader potential of vector search, which goes [**beyond augmenting generative AI**](/articles/vector-similarity-beyond-search/).
### Discovery
The strength of vector search lies in its ability to facilitate [**discovery**](/articles/discovery-search/). Vector search allows you to refine your choices as you search rather than starting with a fixed query. Say, [**you're ordering food not knowing exactly what you want**](/articles/food-discovery-demo/) — just that it should contain meat & not a burger, or that it should be meat with cheese & not tacos. Instead of searching for a specific dish, vector search helps you navigate options based on similarity and dissimilarity, guiding you toward something that matches your taste without requiring you to define it upfront.
### Recommendations
Vector search is perfect for [**recommendations**](/documentation/concepts/explore/#recommendation-api). Imagine browsing for a new book or movie. Instead of searching for an exact match, you might look for stories that capture a certain mood or theme but differ in key aspects from what you already know. For example, you may [**want a film featuring wizards without the familiar feel of the "Harry Potter" series**](https://www.youtube.com/watch?v=O5mT8M7rqQQ). This flexibility is possible because vector search is not tied to the binary "match/not match" concept but operates on distances in a vector space.
### Big Unstructured Data Analysis
Vector search nature makes it also ideal for [**big unstructured data analysis**](https://www.youtube.com/watch?v=_BQTnXpuH-E), for instance, anomaly detection. In large, unstructured, and often unlabelled datasets, vector search can help identify clusters and outliers by analyzing distance relationships between data points.
### Fundamentally Different
**Vector search beyond RAG isn't just another feature — it's a fundamental shift in how we interact with data**. Dedicated solutions integrate these capabilities natively and are designed from the ground up to handle high-dimensional math and (dis-)similarity-based retrieval. In contrast, databases with vector extensions are built around a different data paradigm, making it impossible to efficiently support advanced vector search capabilities.
Even if you want to retrofit these capabilities, it's not just a matter of adding a new feature — it's a structural problem. Supporting advanced vector search requires **dedicated interfaces** that enable flexible usage of vector search from multi-stage filtering to dynamic exploration of high-dimensional spaces.
When the underlying architecture wasn't initially designed for this kind of interaction, integrating interfaces is a **software engineering team nightmare**. You end up breaking existing assumptions, forcing inefficient workarounds, and often introducing backwards-compatibility problems. It's why attempts to patch vector search onto traditional databases won't match the efficiency of purpose-built systems.
## Making Vector Search State-of-the-Art
![vector-search-state-of-the-art](/articles_data/dedicated-vector-search/image4.jpg)
Now, let's shift focus to another key advantage of dedicated solutions — their ability to keep up with state-of-the-art solutions in the field.
[**Vector databases**](/qdrant-vector-database/) are purpose-built for vector retrieval, and as a result, they offer cutting-edge features that are often critical for AI businesses relying on vector search. Vector database engineers invest significant time and effort into researching and implementing the most optimal ways to perform vector search. Many of these innovations come naturally to vector-native architectures, while general-purpose databases with added vector capabilities may struggle to adapt and replicate these benefits efficiently.
Consider some of the advanced features implemented in Qdrant:
- [**GPU-Accelerated Indexing**](/blog/qdrant-1.13.x/#gpu-accelerated-indexing)
By offloading index construction tasks to the GPU, Qdrant can significantly speed up the process of data indexing while keeping costs low. This becomes especially valuable when working with large datasets in hot data scenarios.
GPU acceleration in Qdrant is a custom solution developed by an enthusiast from our core team. It's vendor-free and natively supports all Qdrant's unique architectural features, from FIlterable HNSW to multivectors.
- [**Multivectors**](/documentation/concepts/vectors/?q=multivectors#multivectors)
Some modern embedding models produce an entire matrix (a list of vectors) as output rather than a single vector. Qdrant supports multivectors natively.
This feature is critical when using state-of-the-art retrieval models such as [**ColBERT**](/documentation/fastembed/fastembed-colbert/), ColPali, or ColQwen. For instance, ColPali and ColQwen produce multivector outputs, and supporting them natively is crucial for [**state-of-the-art (SOTA) PDF-retrieval**](/documentation/advanced-tutorials/pdf-retrieval-at-scale/).
In addition to that, we continuously look for improvements in:
| | |
|----------------------------------|-------------|
| **Memory Efficiency & Compression** | Techniques such as [**quantization**](documentation/guides/quantization/) and [**HNSW compression**](/blog/qdrant-1.13.x/#hnsw-graph-compression) to reduce storage requirements |
| **Retrieval Algorithms** | Support for the latest retrieval algorithms, including [**sparse neural retrieval**](/articles/modern-sparse-neural-retrieval/), [**hybrid search**](/documentation/concepts/hybrid-queries/) methods, and [**re-rankers**](/documentation/fastembed/fastembed-rerankers/). |
| **Vector Data Analysis & Visualization** | Tools like the [**distance matrix API**](/blog/qdrant-1.12.x/#distance-matrix-api-for-data-insights) provide insights into vectorized data, and a [**Web UI**](/blog/qdrant-1.11.x/#web-ui-search-quality-tool) allows for intuitive exploration of data. |
| **Search Speed & Scalability** | Includes optimizations for [**multi-tenant environments**](/articles/multitenancy/) to ensure efficient and scalable search. |
**These advancements are not just incremental improvements — they define the difference between a system optimized for vector search and one that accommodates it.**
Staying at the cutting edge of vector search is not just about performance — it's also about keeping pace with an evolving AI landscape.
## Summing up
![conclusion-vector-search](/articles_data/dedicated-vector-search/image5.jpg)
When it comes to vector search, there's a clear distinction between using a dedicated vector search solution and extending a database to support vector operations.
**For small-scale applications or prototypes handling up to a million data points, a non-optimized architecture might suffice.** However, as the volume of vectors grows, an unoptimized solution will quickly become a bottleneck — slowing down search operations and limiting scalability. Dedicated vector search solutions are engineered from the ground up to handle massive amounts of high-dimensional data efficiently.
State-of-the-art (SOTA) vector search evolves rapidly. If you plan to build on the latest advances, using a vector extension will eventually hold you back. Dedicated vector search solutions integrate these features natively, ensuring that you benefit from continuous innovations without compromising performance.
The power of vector search extends into areas such as big data analysis, recommendation systems, and discovery-based applications, and to support these vector search capabilities, a dedicated solution is needed.
### When to Choose a Dedicated Database over an Extension:
- **High-Volume, Real-Time Search**: Ideal for applications with many simultaneous users who require fast, continuous access to search results—think search engines, e-commerce recommendations, social media, or media streaming services.
- **Dynamic, Unstructured Data**: Perfect for scenarios where data is continuously evolving and where the goal is to discover insights from data patterns.
- **Innovative Applications**: If you're looking to implement advanced use cases such as recommendation engines, hybrid search solutions, or exploratory data analysis where traditional exact or token-based searches hold short.
Investing in a dedicated vector search engine will deliver the performance and flexibility necessary for success if your application relies on vector search at scale, keeps up with trends, or requires more than just a simple small-scale similarity search.
@@ -16,8 +16,6 @@ category: practicle-examples
Do you want to insert a semantic search function into your website or online app? Now you can do so - without spending any money! In this example, you will learn how to create a free prototype search engine for your own non-commercial purposes.
You may find all of the assets for this tutorial on [GitHub](https://github.com/qdrant/examples/tree/master/lambda-search).
## Ingredients
* A [Rust](https://rust-lang.org) toolchain
@@ -244,7 +242,7 @@ fn setup<'i>(
}
```
Depending on whether you want to efficiently filter the data, you can also add some indexes. I'm leaving this out for brevity, but you can look at the [example code](https://github.com/qdrant/examples/tree/master/lambda-search) containing this operation. Also this does not implement chunking (splitting the data to upsert in multiple requests, which avoids timeout errors).
Depending on whether you want to efficiently filter the data, you can also add some indexes. I'm leaving this out for brevity. Also this does not implement chunking (splitting the data to upsert in multiple requests, which avoids timeout errors).
Add a suitable `main` method and you can run this code to insert the points (or just use the binary from the example). Be sure to include the port in the `qdrant_url`.
@@ -274,7 +272,7 @@ You can also filter by adding a `filter: ...` field to the `SearchPoints`, and y
## Putting it all together
Now that you have all the parts, it's time to join them up. Now copying and wiring up the snippets above is left as an exercise to the reader. Impatient minds can peruse the [example repo](https://github.com/qdrant/examples/tree/master/lambda-search) instead.
Now that you have all the parts, it's time to join them up. Now copying and wiring up the snippets above is left as an exercise to the reader.
You'll want to extend the `main` method a bit to connect with the Client once at the start, also get API keys from the environment so you don't need to compile them into the code. To do that, you can get them with `std::env::var(_)` from the rust code and set the environment from the AWS console.
@@ -140,7 +140,7 @@ DSPy treats the LM like a device and abstracts out the underlying complexities o
### **Signatures**
[Signatures](https://dspy-docs.vercel.app/docs/building-blocks/signatures) replace handwritten prompts and are written in natural language. They are simply declarations or specs of the behavior that you expect from the language model. Some examples are:
[Signatures](https://dspy.ai/learn/programming/signatures/) replace handwritten prompts and are written in natural language. They are simply declarations or specs of the behavior that you expect from the language model. Some examples are:
- question -> answer
- long_document -> summary
@@ -156,17 +156,17 @@ DSPy Signatures can be specified in two ways:
### **Modules**
Modules take signatures as input, and automatically generate high-quality prompts. Inspired heavily from PyTorch, DSPy [modules](https://dspy-docs.vercel.app/docs/building-blocks/modules) eliminate the need for crafting prompts manually.
Modules take signatures as input, and automatically generate high-quality prompts. Inspired heavily from PyTorch, DSPy [modules](https://dspy.ai/learn/programming/modules/) eliminate the need for crafting prompts manually.
The framework supports advanced modules like [dspy.ChainOfThought](https://dspy-docs.vercel.app/api/modules/ChainOfThought), which adds step-by-step rationalization before producing an output. The output not only provides answers but also rationales. Other modules include [dspy.ProgramOfThought](https://dspy-docs.vercel.app/api/modules/ProgramOfThought), which outputs code whose execution results dictate the response, and [dspy.ReAct](https://dspy-docs.vercel.app/api/modules/ReAct), an agent that uses tools to implement signatures.
DSPy also offers modules like [dspy.MultiChainComparison](https://dspy-docs.vercel.app/api/modules/MultiChainComparison), which can compare multiple outputs from dspy.ChainOfThought in order to produce a final prediction. There are also utility modules like [dspy.majority](https://dspy-docs.vercel.app/docs/building-blocks/modules#what-other-dspy-modules-are-there-how-can-i-use-them) for aggregating responses through voting.
DSPy also offers modules like [dspy.MultiChainComparison](https://dspy-docs.vercel.app/api/modules/MultiChainComparison), which can compare multiple outputs from dspy.ChainOfThought in order to produce a final prediction. There are also utility modules like [dspy.majority](https://dspy.ai/learn/programming/modules/?h=modul#what-other-dspy-modules-are-there-how-can-i-use-them) for aggregating responses through voting.
Modules can be composed into larger programs, and you can compose multiple modules into bigger modules. This allows you to create complex, behavior-rich applications using language models.
### **Optimizers**
[Optimizers](https://dspy-docs.vercel.app/docs/building-blocks/optimizers) take a set of modules that have been connected to create a pipeline, compile them into auto-optimized prompts, and maximize an outcome metric.
[Optimizers](https://dspy.ai/learn/optimization/optimizers/) take a set of modules that have been connected to create a pipeline, compile them into auto-optimized prompts, and maximize an outcome metric.
Essentially, optimizers are designed to generate, test, and refine prompts, and ensure that the final prompt is highly optimized for the specific dataset and task at hand. Using optimizers in the DSPy framework significantly simplifies the process of developing and refining LM applications by automating the prompt engineering process.
@@ -212,7 +212,7 @@ print(response.answer)
```
You are not restricted to using one LLM in your program; you can use [multiple](https://dspy-docs.vercel.app/docs/building-blocks/language_models#using-multiple-lms-at-once). DSPy can be used with both managed models such as OpenAI, Cohere, Anyscale, Together, or PremAI as well as with local LLM deployments through vLLM, Ollama, or TGI server. All LLM calls are cached by default.
You are not restricted to using one LLM in your program; you can use [multiple](https://dspy.ai/learn/programming/language_models/?h=language#using-multiple-lms). DSPy can be used with both managed models such as OpenAI, Cohere, Anyscale, Together, or PremAI as well as with local LLM deployments through vLLM, Ollama, or TGI server. All LLM calls are cached by default.
**Vector Store Integration (Retrieval Model)**
@@ -270,7 +270,7 @@ Using DSPy optimizers involves the following steps:
5. Run the optimizer with the DSPy program, metric function, and training inputs. DSPy will compile the program and automatically adjust parameters and improve performance.
6. Use the compiled program to perform the task. Iterate and adapt if required.
To learn more about optimizing DSPy programs, read [this](https://dspy-docs.vercel.app/docs/building-blocks/optimizers).
To learn more about optimizing DSPy programs, read [this](https://dspy.ai/learn/optimization/optimizers/).
DSPy is heavily influenced by PyTorch, and replaces complex prompting with reusable modules for common tasks. Instead of crafting specific prompts, you write code that DSPy automatically translates for the LLM. This, along with built-in optimizers, makes working with LLMs more systematic and efficient.
@@ -400,4 +400,4 @@ LangChain and DSPy both offer unique capabilities and can help you build powerfu
[https://python.langchain.com/v0.1/docs/get_started/introduction](https://python.langchain.com/v0.1/docs/get_started/introduction)
[https://dspy-docs.vercel.app/docs/intro](https://dspy-docs.vercel.app/docs/intro)
[DSPy Introduction](https://dspy.ai/)
@@ -57,7 +57,7 @@ To simplify the evaluation process, several powerful frameworks are available. B
### Ragas: Testing RAG with questions and answers
[Ragas](https://docs.ragas.io/en/v0.0.17/index.html) (or RAG Assessment) uses a dataset of questions, ideal answers, and relevant context to compare a RAG system's generated answers with the ground truth. It provides metrics like faithfulness, relevance, and semantic similarity to assess retrieval and answer quality.
[Ragas](https://docs.ragas.io/en/stable/) (or RAG Assessment) uses a dataset of questions, ideal answers, and relevant context to compare a RAG system's generated answers with the ground truth. It provides metrics like faithfulness, relevance, and semantic similarity to assess retrieval and answer quality.
**Figure 1:** *Output of the Ragas framework, showcasing metrics like faithfulness, answer relevancy, context recall, precision, relevancy, entity recall, and answer similarity. These are used to evaluate the quality of RAG system responses.*
@@ -175,7 +175,7 @@ First, create question and ground-truth answer pairs from source documents for t
- **Hand-crafting your dataset:** Manually create questions and answers.
- **Use LLM to create synthetic data:** Leverage LLMs like [T5](https://huggingface.co/docs/transformers/en/model_doc/t5) or OpenAI APIs.
- **Use the Ragas framework**: [This method](https://docs.ragas.io/en/latest/getstarted/testset_generation.html) uses an LLM to generate various question types for evaluating RAG systems.
- **Use the Ragas framework**: [This method](https://docs.ragas.io/en/stable/concepts/test_data_generation/) uses an LLM to generate various question types for evaluating RAG systems.
- **Use FiddleCube**: [FiddleCube](https://www.fiddlecube.ai/) is a system that can help generate a range of question types aimed at different aspects of the testing process.
Once you have created a dataset, collect the retrieved context and the final answer generated by your RAG pipeline for each question.
@@ -1295,7 +1295,7 @@ To get the facet counts for a field, you can use the following:
<aside role="status">By default, the number of <code>hits</code> returned is limited to 10. To change this, use the <code>limit</code> parameter. Keep this in mind when checking the number of unique values a payload field contains.</aside>
REST API ([Facet](https://api.qdrant.tech/api-reference/search/facet))
REST API ([Facet](https://api.qdrant.tech/v-1-13-x/api-reference/points/facet))
```http
POST /collections/{collection_name}/facet
@@ -87,4 +87,4 @@ QdrantIngestOperator(
## Reference
- 📦 [Provider package PyPI](https://pypi.org/project/apache-airflow-providers-qdrant/)
- 📚 [Provider docs](https://airflow.apache.org/docs/apache-airflow-providers-qdrant/stable/index.html)
- 📄 [Source Code](https://github.com/apache/airflow/tree/main/airflow/providers/qdrant)
- 📄 [Source Code](https://github.com/apache/airflow/tree/main/providers/qdrant)
@@ -97,4 +97,4 @@ if __name__ == "__main__":
- Unstructured API [reference](https://unstructured-io.github.io/unstructured/api.html).
- Qdrant ingestion destination [reference](https://unstructured-io.github.io/unstructured/ingest/destination_connectors/qdrant.html).
- [Source Code](https://github.com/Unstructured-IO/unstructured/blob/main/unstructured/ingest/connector/qdrant.py)
- [Source Code](https://github.com/Unstructured-IO/unstructured-ingest/blob/main/unstructured_ingest/connector/qdrant.py)
@@ -88,9 +88,3 @@ method call.
<aside role="status">
Asynchronous client was introduced in <code>qdrant-client</code> version 1.6.1. If you are using an older version, you need to use autogenerated async clients directly.
</aside>
## Supported Python libraries
Qdrant integrates with numerous Python libraries. Until recently, only [Langchain](https://python.langchain.com) provided async Python API support.
Qdrant is the only vector database with full coverage of async API in Langchain. Their documentation [describes how to use
it](https://python.langchain.com/docs/modules/data_connection/vectorstores/#asynchronous-operations).
@@ -73,9 +73,9 @@ client.upsert(
Once the documents are indexed, you can search for the most relevant documents using the Embed v3 model:
```python
client.search(
client.query_points(
collection_name="MyCollection",
query_vector=cohere_client.embed(
query=cohere_client.embed(
model="embed-english-v3.0", # New Embed v3 model
input_type="search_query", # Input type for search queries
texts=["The best vector database"],
@@ -24,7 +24,7 @@ the documents used to generate the response.
*Source: https://docs.cohere.com/docs/retrieval-augmented-generation-rag*
The connectors have to implement a specific interface and expose the data source as HTTP REST API. Cohere documentation
[describes a general process of creating a connector](https://docs.cohere.com/docs/creating-and-deploying-a-connector).
[describes a general process of creating a connector](https://docs.cohere.com/v1/docs/creating-and-deploying-a-connector).
This tutorial guides you step by step on building such a service around Qdrant.
## Qdrant connector
@@ -153,7 +153,7 @@ class SearchQuery(BaseModel):
query: str
```
RAG connector does not have to return the documents in any specific format. There are [some good practices to follow](https://docs.cohere.com/docs/creating-and-deploying-a-connector#configure-the-connection-between-the-connector-and-the-chat-api),
RAG connector does not have to return the documents in any specific format. There are [some good practices to follow](https://docs.cohere.com/v1/docs/creating-and-deploying-a-connector#configure-the-connection-between-the-connector-and-the-chat-api),
but Cohere models are quite flexible here. Results just have to be returned as JSON, with a list of objects in a
`results` property of the output. We will use the same document structure as we did for the Qdrant payloads, so there
is no conversion required. That requires two additional models to be created.
@@ -8,14 +8,14 @@ aliases:
# Blog-Reading Chatbot with GPT-4o
| Time: 90 min | Level: Advanced |[GitHub](https://github.com/qdrant/examples/blob/master/langchain-lcel-rag/Langchain-LCEL-RAG-Demo.ipynb)| |
| Time: 90 min | Level: Advanced |[GitHub](https://github.com/qdrant/examples/blob/langchain-lcel-rag/langchain-lcel-rag/Langchain-LCEL-RAG-Demo.ipynb)| |
|--------------|-----------------|--|----|
In this tutorial, you will build a RAG system that combines blog content ingestion with the capabilities of semantic search. **OpenAI's GPT-4o LLM** is powerful, but scaling its use requires us to supply context systematically.
RAG enhances the LLM's generation of answers by retrieving relevant documents to aid the question-answering process. This setup showcases the integration of advanced search and AI language processing to improve information retrieval and generation tasks.
A notebook for this tutorial is available on [GitHub](https://github.com/qdrant/examples/blob/master/langchain-lcel-rag/Langchain-LCEL-RAG-Demo.ipynb).
A notebook for this tutorial is available on [GitHub](https://github.com/qdrant/examples/blob/langchain-lcel-rag/langchain-lcel-rag/Langchain-LCEL-RAG-Demo.ipynb).
**Data Privacy and Sovereignty:** RAG applications often rely on sensitive or proprietary internal data. Running the entire stack within your own environment becomes crucial for maintaining control over this data. Qdrant Hybrid Cloud deployed on [Scaleway](https://www.scaleway.com/) addresses this need perfectly, offering a secure, scalable platform that still leverages the full potential of RAG. Scaleway offers serverless [Functions](https://www.scaleway.com/en/serverless-functions/) and serverless [Jobs](https://www.scaleway.com/en/serverless-jobs/), both of which are ideal for embedding creation in large-scale RAG cases.
@@ -278,7 +278,7 @@ Output:
The task was solved successfully, even without any optimization. However, each of the events has the "Event Name: "
prefix that we might want to remove. DSPy allows optimizing the module, so we can improve the results. Optimization
might be done in different ways, and it's [well covered in the DSPy
documentation](https://dspy-docs.vercel.app/docs/building-blocks/optimizers).
documentation](https://dspy.ai/learn/optimization/optimizers/).
We are not going to go through the optimization process in this tutorial. However, we encourage you to experiment with
it, as it might significantly improve the performance of your pipeline.
@@ -27,10 +27,10 @@ Directory.
> **Note:** In this tutorial, we are going to build a solid foundation for such a system. However, it is up to your organization's setup to implement the entire solution.
- **Dataset** - a collection of documents, using different formats, such as PDF or DOCx, scraped from internet
- **Asymmetric semantic embeddings** - [Aleph Alpha embedding](https://docs.aleph-alpha.com/api/semantic-embed/) to
- **Asymmetric semantic embeddings** - [Aleph Alpha embedding](https://docs.aleph-alpha.com/api/pharia-inference/semantic-embed/) to
convert the queries and the documents into vectors
- **Large Language Model** - the [Luminous-extended-control
model](https://docs.aleph-alpha.com/docs/introduction/model-card/), but you can play with a different one from the
model](https://docs.aleph-alpha.com/api/pharia-inference/available-models/), but you can play with a different one from the
Luminous family
- **Qdrant Hybrid Cloud** - a knowledge base to store the vectors and search over the documents
- **STACKIT** - a [German business cloud](https://www.stackit.de) to run the Qdrant Hybrid Cloud and the application
@@ -44,7 +44,7 @@ interacts with the system with some set of permissions, and can only access the
### Aleph Alpha account
Since you will be using Aleph Alpha's models, [sign up](https://app.aleph-alpha.com/signup) with their managed service and generate an API token in the [User Profile](https://app.aleph-alpha.com/profile). Once you have it ready, store it as an environment variable:
Since you will be using Aleph Alpha's models, [sign up](https://aleph-alpha.com) with their managed service and obtain an API token. Once you have it ready, store it as an environment variable:
```shell
export ALEPH_ALPHA_API_KEY="<your-token>"
@@ -97,8 +97,7 @@ os.environ["QDRANT_API_KEY"] = "your-api-key"
### Airbyte Open Source
Airbyte is an open-source data integration platform that helps you replicate your data in your warehouses, lakes, and
databases. You can install it on your infrastructure and use it to load the data into Qdrant. The installation process
for AWS EC2 is described in the [official documentation](https://docs.airbyte.com/deploying-airbyte/on-aws-ec2).
databases. You can install it on your infrastructure and use it to load the data into Qdrant. The installation process is described in the [official documentation](https://docs.airbyte.com/deploying-airbyte/).
Please follow the instructions to set up your own instance.
#### Setting up the connection
@@ -27,6 +27,7 @@ partition: build
| [LangGraph](/documentation/frameworks/langgraph/) | Python, Javascript libraries for building stateful, multi-actor applications. |
| [LlamaIndex](/documentation/frameworks/llama-index/) | A data framework for building LLM applications with modular integrations. |
| [Mastra](/documentation/frameworks/mastra/) | Typescript framework to build AI applications and features quickly. |
| [Mirror Security](/documentation/frameworks/mirror-security/) | Python framework for vector encryption and access control. |
| [Mem0](/documentation/frameworks/mem0/) | Self-improving memory layer for LLM applications, enabling personalized AI experiences. |
| [MemGPT](/documentation/frameworks/memgpt/) | System to build LLM agents with long term memory & custom tools |
| [Neo4j GraphRAG](/documentation/frameworks/neo4j-graphrag/) | Package to build graph retrieval augmented generation (GraphRAG) applications using Neo4j and Python. |
@@ -35,7 +35,7 @@ online_store:
write_batch_size: 100
```
You can refer to the Feast [reference](https://rtd.feast.dev/en/master/index.html#) for the full list of configuration options.
You can refer to the Feast [documentation](https://docs.feast.dev/reference/alpha-vector-database#configuration-and-installation) for the full list of configuration options.
## Retrieving Documents
@@ -58,6 +58,5 @@ feature_values = feature_store.retrieve_online_documents(
## 📚 Further Reading
- [Feast Docs](http://docs.feast.dev/)
- [Feast Reference](https://rtd.feast.dev/en/master/index.html/)
- [Feast Documentation](http://docs.feast.dev/)
- [Source](https://github.com/feast-dev/feast/tree/master/sdk/python/feast/infra/online_stores/)
@@ -0,0 +1,185 @@
---
title: VectaX - Mirror Security
---
![VectaX Logo](/documentation/frameworks/mirror-security/vectax-logo.png)
[VectaX](https://mirrorsecurity.io/vectax) by Mirror Security is an AI-centric access control and encryption system designed for managing and protecting vector embeddings. It combines similarity-preserving encryption with fine-grained RBAC to enable secure storage, retrieval, and operations on vector data.
It can be integrated with Qdrant to secure vector searches.
We'll see how to do so using basic VectaX vector encryption and the sophisticated RBAC mechanism. You can obtain an API key and the Mirror SDK from the [Mirror Security Platform](https://platform.mirrorsecurity.io/en/login).
Let's set up both the VectaX and Qdrant clients.
```python
from mirror_sdk.core.mirror_core import MirrorSDK, MirrorConfig
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams
# Get your API key from
# https://platform.mirrorsecurity.io
config = MirrorConfig(
api_key="<your_api_key>",
server_url="https://mirrorapi.azure-api.net/v1",
secret="<your_encrypt_secret>",
)
mirror_sdk = MirrorSDK(config)
# Connects to http://localhost:6333/ by default
qdrant = QdrantClient()
```
## Vector Encryption
Now, let's secure vector embeddings using VectaX encryption.
```python
from qdrant_client.models import PointStruct
from mirror_sdk.core.models import VectorData
# Generate or retrieve vector embeddings
# embedding = generate_document_embedding()
vector_data = VectorData(vector=embedding, id="doc1")
encrypted = mirror_sdk.vectax.encrypt(vector_data)
point = PointStruct(
id=0,
vector=encrypted.ciphertext,
payload={
"content": "Document content",
"iv": encrypted.iv,
"auth_hash": encrypted.auth_hash
}
)
qdrant.upsert(collection_name="vectax", points=[point])
# Encrypt a query vector for secure search
# query_embedding = generate_query_embedding(...)
encrypted_query = mirror_sdk.vectax.encrypt(
VectorData(vector=query_embedding, id="query")
)
results = qdrant.query_points(
collection_name="vectax",
query=encrypted_query.ciphertext,
limit=5
).points
```
## Vector Search with RBAC
RBAC allows fine-grained access control over encrypted vector data based on roles, groups, and departments.
### Defining Access Policies
```python
app_policy = {
"roles": ["admin", "analyst", "user"],
"groups": ["team_a", "team_b"],
"departments": ["research", "engineering"],
}
mirror_sdk.set_policy(app_policy)
```
### Generating Access Keys
```python
# Generate a secret key for use by the 'admin' role holders.
admin_key = mirror_sdk.rbac.generate_user_secret_key(
{"roles": ["admin"], "groups": ["team_a"], "departments": ["research"]}
)
```
### Storing Encrypted Data with RBAC Policies
We can now store data that is only accessible to users with the "admin" role.
```python
from mirror_sdk.core.models import RBACVectorData
from mirror_sdk.utils import encode_binary_data
policy = {
"roles": ["admin"],
"groups": ["team_a"],
"departments": ["research"],
}
# vector_embedding = generate_vector_embedding(...)
vector_data = RBACVectorData(
# Generate or retrieve vector embeddings
vector=vector_embedding,
id=1,
access_policy=policy,
)
encrypted = mirror_sdk.rbac.encrypt(vector_data)
qdrant.upsert(
collection_name="vectax",
points=[
models.PointStruct(
id=1,
vector=encrypted.crypto.ciphertext,
payload={
"encrypted_header": encrypted.encrypted_header,
"encrypted_vector_metadata": encode_binary_data(
encrypted.crypto.serialize()
),
"content": "My content",
},
)
],
)
```
### Querying with Role-Based Decryption
Using the admin key, only accessible data will be decrypted.
```python
from mirror_sdk.core import MirrorError
from mirror_sdk.core.models import MirrorCrypto
from mirror_sdk.utils import decode_binary_data
# Encrypt a query vector for secure search
# query_embedding = generate_query_embedding(...)
query_data = RBACVectorData(vector=query_embedding, id="query", access_policy=policy)
encrypted_query = mirror_sdk.rbac.encrypt(query_data)
results = qdrant.query_points(
collection_name="vectax", query=encrypted_query.crypto.ciphertext, limit=10
)
accessible_results = []
for point in results.points:
try:
encrypted_vector_metadata = decode_binary_data(
point.payload["encrypted_vector_metadata"]
)
mirror_data = MirrorCrypto.deserialize(encrypted_vector_metadata)
admin_decrypted = mirror_sdk.rbac.decrypt(
mirror_data,
point.payload["encrypted_header"],
admin_key,
)
accessible_results.append(
{
"id": point.id,
"content": point.payload["content"],
"score": point.score,
"accessible": True,
}
)
except MirrorError as e:
print(f"Access denied for point {point.id}: {e}")
# Proceed to only use results within `accessible_results`.
```
## Further Reading
- [Mirror Security Docs](https://docs.mirrorsecurity.io/introduction)
- [Mirror Security Blog](https://mirrorsecurity.io/blog)
@@ -92,4 +92,4 @@ agent.train(queries=[query], codes=[response])
- [Getting Started with Pandas-AI](https://pandasai-docs.readthedocs.io/en/latest/getting-started/)
- [Pandas-AI Reference](https://pandasai-docs.readthedocs.io/en/latest/)
- [Source Code](https://github.com/Sinaptik-AI/pandas-ai/blob/main/pandasai/ee/vectorstores/qdrant.py)
- [Source Code](https://github.com/sinaptik-ai/pandas-ai/tree/main/extensions/ee/vectorstores/qdrant)
@@ -59,7 +59,7 @@ However, binary quantization is only efficient for high-dimensional vectors and
At the moment, binary quantization shows good accuracy results with the following models:
- OpenAI `text-embedding-ada-002` - 1536d tested with [dbpedia dataset](https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M) achieving 0.98 recall@100 with 4x oversampling
- Cohere AI `embed-english-v2.0` - 4096d tested on [wikipedia embeddings](https://huggingface.co/datasets/nreimers/wikipedia-22-12-large/tree/main) - 0.98 recall@50 with 2x oversampling
- Cohere AI `embed-english-v2.0` - 4096d tested on Wikipedia embeddings - 0.98 recall@50 with 2x oversampling
Models with a lower dimensionality or a different distribution of vector components may require additional experiments to find the optimal quantization parameters.
@@ -37,10 +37,6 @@ First, install the required libraries `qdrant-client` and `llama-index-embedding
pip install qdrant-client llama-index-embeddings-huggingface
```
<aside role="status">
The code for this tutorial can be found <a href="https://github.com/qdrant/examples/multimodal-search">here</a>.
</aside>
## Dataset
To make the demonstration simple, we created a tiny dataset of images and their captions for you.
@@ -11,7 +11,7 @@ Qdrant is supported as a vectorstore in DocsGPT to ingest and semantically retri
## Configuration
Learn how to setup DocsGPT in their [Quickstart guide](https://docs.docsgpt.co.uk/Deploying/Quickstart).
Learn how to setup DocsGPT in their [Quickstart guide](https://docs.docsgpt.cloud/quickstart).
You can configure DocsGPT with environment variables in a `.env` file.
@@ -139,7 +139,7 @@ spec:
If you set the `jwt_rbac` flag, you will also be able to create granular [JWT tokens for role based access control](/documentation/guides/security/#granular-access-control-with-jwt).
### Configuring TLS
### Configuring TLS for Database Access
If you want to configure TLS for accessing your Qdrant database, there are two options:
@@ -195,3 +195,96 @@ spec:
name: qdrant-tls
key: tls.key
```
### Configuring TLS for Inter-cluster Communication
*Available as of Operator v2.2.0*
<aside role="alert">
The feature can be enabled only at the cluster creation. Later changes are not possible.
</aside>
If you want to encrypt communication between Qdrant nodes, you need to enable TLS by providing
certificate, key, and root CA certificate used for generating the former.
Similar to the instruction stated in the previous section, you need to create a secret:
```shell
kubectl create secret generic qdrant-p2p-tls \
--from-file=tls.crt=qdrant-nodes.crt \
--from-file=tls.key=qdrant-nodes.key \
--from-file=ca.crt=root-ca.crt
--namespace the-qdrant-namespace
```
The resulting secret will look like this:
```yaml
apiVersion: v1
data:
tls.crt: ...
tls.key: ...
ca.crt: ...
kind: Secret
metadata:
name: qdrant-p2p-tls
namespace: the-qdrant-namespace
type: Opaque
```
You can reference the secret in the QdrantCluster spec:
```yaml
apiVersion: qdrant.io/v1
kind: QdrantCluster
metadata:
name: test-cluster
labels:
cluster-id: "my-cluster"
customer-id: "acme-industries"
spec:
id: "my-cluster"
version: "v1.13.3"
size: 2
resources:
cpu: 100m
memory: "1Gi"
storage: "2Gi"
config:
service:
enable_tls: true
tls:
caCert:
secretKeyRef:
name: qdrant-p2p-tls
key: ca.crt
cert:
secretKeyRef:
name: qdrant-p2p-tls
key: tls.crt
key:
secretKeyRef:
name: qdrant-p2p-tls
key: tls.key
```
<aside role="status">
The operator assigns the names to nodes in cluster according to the following convention:
```
qdrant-{spec.id}-{node-index}.qdrant-headless-{spec.id}
```
Therefore, in addition to the domain used for accessing the database,
the provided certificate must contain Subjective Alternative Names (SAN) for all foreseen nodes.
It can be created with a tool of your choice, e.g.,
using [step CLI](https://smallstep.com/docs/step-cli/installation/).
Following the example `QdrantCluster`, the proper certificate can be obtained with:
```shell
step certificate create mydomain.com qdrant-nodes.crt qdrant-nodes.key \
--profile leaf --not-after 43800h \
--ca root-ca.crt --ca-key root-ca.key \
--san qdrant-my-cluster-0.qdrant-headless-my-cluster \
--san qdrant-my-cluster-1.qdrant-headless-my-cluster
```
</aside>
+2 -2
View File
@@ -1,6 +1,6 @@
---
stats:
githubStars: 21.7k
discordMembers: 7.5k
githubStars: 21.9k
discordMembers: 7.6k
twitterFollowers: 7.5k
---
Binary file not shown.

After

Width:  |  Height:  |  Size: 24 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 14 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 787 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 208 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 219 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 252 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 252 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 324 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 194 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 31 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 26 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 207 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 84 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 68 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 370 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 76 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 163 KiB

@@ -43,6 +43,13 @@
p {
margin-bottom: 0;
}
& > p {
width: 49.5%;
&:last-child {
text-align: right;
}
}
}
&-link:hover {
@@ -75,6 +75,12 @@
right: 0;
}
.hs_recaptcha.hs-recaptcha {
display: flex;
align-items: end;
margin-bottom: -1px;
}
&__top-overlay {
position: absolute;
right: 0;
@@ -214,3 +214,10 @@ form.hs-form {
line-height: pxToRem(24);
text-align: right;
}
.hs_recaptcha.hs-recaptcha {
.grecaptcha-badge {
box-shadow: none !important;
}
}
@@ -101,10 +101,13 @@
}
blockquote {
padding: pxToRem(32) pxToRem(40) 1px;
padding: pxToRem(32) pxToRem(40);
border-radius: $spacer * 0.5;
background: linear-gradient(180deg, #161e33 0%, #0e1424 100%);
color: $neutral-98;
p:last-child {
margin-bottom: 0;
}
}
details {
@@ -67,4 +67,8 @@
max-width: pxToRem(537);
}
}
.hs_recaptcha.hs-recaptcha {
order: 3;
}
}
@@ -111,6 +111,16 @@
}
}
blockquote {
padding: pxToRem(32) pxToRem(24);
border-radius: $spacer * 0.5;
background: linear-gradient(180deg, #161e33 0%, #0e1424 100%);
color: $neutral-98;
p:last-child {
margin-bottom: 0;
}
}
@include media-breakpoint-up(xl) {
width: 100%;
@@ -75,7 +75,6 @@
<p class="post-description">{{ .Params.description }}</p>
<div class="post-about">
<p>{{ .Params.author }}</p>
<span>&middot;</span>
<p>{{ time.Format "January 02, 2006" .Date }}</p>
</div>
</a>
@@ -118,7 +117,6 @@
<p class="post-description">{{ .Params.description }}</p>
<div class="post-about">
<p>{{ .Params.author }}</p>
<span>&middot;</span>
<p>{{ time.Format "January 02, 2006" .Date }}</p>
</div>
</a>
@@ -18,7 +18,6 @@
<p class="post-description">{{ .Params.description }}</p>
<div class="post-about">
<p>{{ .Params.author }}</p>
<span>&middot;</span>
<p>{{ time.Format "January 02, 2006" .Date }}</p>
</div>
</a>
@@ -19,7 +19,6 @@
<p class="post-description">{{ .Params.description }}</p>
<div class="post-about">
<p>{{ .Params.author }}</p>
<span>&middot;</span>
<p>{{ time.Format "January 02, 2006" .Date }}</p>
</div>
</a>
@@ -29,7 +29,6 @@
<h5 class="post-title">{{ .Params.title }}</h5>
<div class="post-about">
<p>{{ .Params.author }}</p>
<span>&middot;</span>
<p>{{ time.Format "January 02, 2006" .Date }}</p>
</div>
</a>
@@ -44,7 +43,6 @@
<h6 class="post-title">{{ .Params.title }}</h6>
<div class="post-about">
<p>{{ .Params.author }}</p>
<span>&middot;</span>
<p>{{ time.Format "January 02, 2006" .Date }}</p>
</div>
</a>
@@ -24,7 +24,6 @@
<p class="post-description">{{ .Params.description }}</p>
<div class="post-about">
<p>{{ .Params.author }}</p>
<span>&middot;</span>
<p>{{ time.Format "January 02, 2006" .Date }}</p>
</div>
</a>