mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-03 01:48:32 +02:00
Fix Docs : minor grammar fixes (#2218)
* fix(docs): fix typos and some links in documentation * fix(docs): correct typo in filtering.md * Update qdrant-landing/content/documentation/headless/snippets/inference/jinaai-upsert/generated/typescript.md Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com> * Update qdrant-landing/content/documentation/headless/snippets/inference/multiple/generated/typescript.md Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com> * Update qdrant-landing/content/documentation/hybrid-cloud/configure-scale-upgrade.md Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com> * Update qdrant-landing/content/documentation/cloud-api.md Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com> --------- Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com>
This commit is contained in:
co-authored by
Abdon Pijpelink
parent
3f3ef1ad20
commit
c0db45ed8f
@@ -147,7 +147,7 @@ Authentication for agents is handled by [API keys](https://qdrant.tech/documenta
|
||||
|
||||
Note: Qdrant also supports concurrent queries, so your search won’t slow down as more users are writing queries simultaneously.
|
||||
|
||||
Authorization is handled by [RBAC](https://qdrant.tech/articles/data-privacy/) and [multitenancy](https://qdrant.tech/documentation/guides/multiple-partitions/), which work hand-in-hand to define and enforce permissions specific to the agent. RBAC is a set of rules that defines the allowed permissions inlcuding read-only, read-write, and admin controls. It answers the question, “What is this agent allowed to do?” For instance, can it only search for hotels (read-only), or can it also add, update, and delete them (read-write)?
|
||||
Authorization is handled by [RBAC](https://qdrant.tech/articles/data-privacy/) and [multitenancy](https://qdrant.tech/documentation/guides/multiple-partitions/), which work hand-in-hand to define and enforce permissions specific to the agent. RBAC is a set of rules that defines the allowed permissions including read-only, read-write, and admin controls. It answers the question, “What is this agent allowed to do?” For instance, can it only search for hotels (read-only), or can it also add, update, and delete them (read-write)?
|
||||
|
||||

|
||||
|
||||
|
||||
@@ -212,7 +212,7 @@ Some of the key concepts of CrewAI include:
|
||||
Qdrant comes into play, as it might be used as a long-term memory layer.**
|
||||
|
||||
CrewAI provides a rich set of tools integrated into the framework. That may be a huge advantage for those who want to
|
||||
combine RAG with e.g. code execution, or image generation. The ecosystem is rich, however brining your own tools is
|
||||
combine RAG with e.g. code execution, or image generation. The ecosystem is rich, however bringing your own tools is
|
||||
not a big deal, as CrewAI is designed to be extensible.
|
||||
|
||||
A simple agentic RAG application implemented in CrewAI could look like this:
|
||||
|
||||
@@ -34,7 +34,7 @@ However, similarity learning comes with its own difficulties such as:
|
||||
|
||||
Quaterion is a fine tuning framework built to tackle such problems in similarity learning.
|
||||
It uses [PyTorch Lightning](https://www.pytorchlightning.ai/)
|
||||
as a backend, which is advertized with the motto, "spend more time on research, less on engineering."
|
||||
as a backend, which is advertised with the motto, "spend more time on research, less on engineering."
|
||||
This is also true for Quaterion, and it includes:
|
||||
|
||||
1. Trainable and servable model classes,
|
||||
|
||||
@@ -11,7 +11,7 @@ draft: false
|
||||
keywords:
|
||||
- clusterization
|
||||
- dimensionality reduction
|
||||
- vizualization
|
||||
- visualization
|
||||
category: data-exploration
|
||||
---
|
||||
|
||||
@@ -25,7 +25,7 @@ Examining data points individually is not always the best way to grasp the struc
|
||||
|
||||
As numbers in a table obtain meaning when plotted on a graph, visualising distances (similar/dissimilar) between unstructured data items can reveal hidden structures and patterns.
|
||||
|
||||
{{< figure src="/articles_data/distance-based-exploration/data-on-chart.png" alt="Data visualization" caption="Vizualized chart, very intuitive" >}}
|
||||
{{< figure src="/articles_data/distance-based-exploration/data-on-chart.png" alt="Data visualization" caption="Visualized chart, very intuitive" >}}
|
||||
There are many tools to investigate data similarity, and Qdrant's [1.12 release](https://qdrant.tech/blog/qdrant-1.12.x/) made it much easier to start this investigation. With the new [Distance Matrix API](/documentation/concepts/explore/#distance-matrix), Qdrant handles the most computationally expensive part of the process—calculating the distances between data points.
|
||||
|
||||
In many implementations, the distance matrix calculation was part of the clustering or visualization processes, requiring either brute-force computation or building a temporary index. With Qdrant, however, the data is already indexed, and the distance matrix can be computed relatively cheaply.
|
||||
@@ -62,7 +62,7 @@ PUT /collections/midlib/snapshots/recover
|
||||
```
|
||||
|
||||
<details>
|
||||
<summary>We also need to prepare our python enviroment:</summary>
|
||||
<summary>We also need to prepare our python environment:</summary>
|
||||
|
||||
```bash
|
||||
pip install umap-learn seaborn matplotlib qdrant-client
|
||||
@@ -77,7 +77,7 @@ from qdrant_client import QdrantClient
|
||||
from umap import UMAP
|
||||
# Python implementation for sparse matrices
|
||||
from scipy.sparse import csr_matrix
|
||||
# For vizualization
|
||||
# For visualization
|
||||
import seaborn as sns
|
||||
```
|
||||
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
title: Filterable HNSW
|
||||
short_description: How to make ANN search with custom filtering?
|
||||
description: How to make ANN search with custom filtering? Search in selected subsets without loosing the results.
|
||||
description: How to make ANN search with custom filtering? Search in selected subsets without losing the results.
|
||||
# external_link: https://blog.vasnetsov.com/posts/categorical-hnsw/
|
||||
social_preview_image: /articles_data/filterable-hnsw/social_preview.jpg
|
||||
preview_dir: /articles_data/filterable-hnsw/preview
|
||||
|
||||
@@ -27,7 +27,7 @@ Otherwise, read on to learn more about the demo and how it works!
|
||||
In general, our application consists of three parts: a [FastAPI](https://fastapi.tiangolo.com/) backend, a [React](https://react.dev/) frontend, and
|
||||
a [Qdrant](/) instance. The architecture diagram below shows how these components interact with each other:
|
||||
|
||||

|
||||

|
||||
|
||||
## Why did we use a CLIP model?
|
||||
|
||||
|
||||
@@ -234,7 +234,7 @@ internal document expansion idea, which made the retrieval quality noticeably be
|
||||
|
||||
- The SPARTA model is not sparse enough by construction, so authors of the SPLADE family of models introduced explicit **sparsity regularisation**,
|
||||
preventing the model from producing too many non-zero values.
|
||||
- The SPARTA model mostly uses the BERT model as-is, without any additional neural network to capture the specifity of Information Retrieval problem,
|
||||
- The SPARTA model mostly uses the BERT model as-is, without any additional neural network to capture the specificity of Information Retrieval problem,
|
||||
so SPLADE models introduce a trainable neural network on top of BERT with a specific architecture choice to make it perfectly fit the task.
|
||||
- SPLADE family of models, finally, uses **knowledge distillation**, which is learning from a bigger
|
||||
(and therefore much slower, not-so-fit for production tasks) model how to predict good representations.
|
||||
|
||||
@@ -66,7 +66,7 @@ done by computing the dot product of the input vector with each hyperplane norma
|
||||
result. Since each of our regions can be represented as a binary string of length `k_sim` (where each bit indicates
|
||||
which side of a hyperplane the vector is on), we can interpret this binary string as an integer to get a cluster ID.
|
||||
|
||||

|
||||

|
||||
|
||||
### Fixed Dimensional Encoding (FDE) creation
|
||||
|
||||
|
||||
@@ -33,7 +33,7 @@ Your feedback is valuable to us, and are always tying to include some of your fe
|
||||
|
||||
## New features
|
||||
|
||||
### Asychronous I/O interface
|
||||
### Asynchronous I/O interface
|
||||
|
||||
Going forward, we will support the `io_uring` asychnronous interface for storage devices on Linux-based systems. Since its introduction, `io_uring` has been proven to speed up slow-disk deployments as it decouples kernel work from the IO process.
|
||||
|
||||
@@ -201,7 +201,7 @@ Internally, `is_empty` was not using the index when it was called, so it had to
|
||||
|
||||
### Faster read access with mmap
|
||||
|
||||
If you used mmap, you most likely found that segments were always created with cold caches. The first request to the database needed to request the disk, which made startup slower despite plenty of RAM being available. We have implemeneted a way to ask the kernel to "heat up" the disk cache and make initialization much faster.
|
||||
If you used mmap, you most likely found that segments were always created with cold caches. The first request to the database needed to request the disk, which made startup slower despite plenty of RAM being available. We have implemented a way to ask the kernel to "heat up" the disk cache and make initialization much faster.
|
||||
|
||||
The function is expected to be used on startup and after segment optimization and reloading of newly indexed segment. So far this is only implemented for "immutable" memmaps.
|
||||
|
||||
|
||||
@@ -467,7 +467,7 @@ We will reprocess the data with the updated parameters above:
|
||||
|
||||
```python
|
||||
## for iteration 2 - lets modify chunk configuration
|
||||
## We will start with creating seperate collection to store vectors
|
||||
## We will start with creating separate collection to store vectors
|
||||
|
||||
chunk_size = 1024
|
||||
chunk_overlap = 128
|
||||
|
||||
@@ -169,7 +169,7 @@ In this loss, the model is trained by fitting the information of relative simila
|
||||
Using the same mechanics, we can look at the training process from the other side.
|
||||
Given a trained model, the user can provide positive and negative examples, and the goal of the discovery process is then to find suitable anchors across the stored collection of vectors.
|
||||
|
||||
<!-- ToDo: image where we know positive and nagative -->
|
||||
<!-- ToDo: image where we know positive and negative -->
|
||||
{{< figure width=60% src=/articles_data/vector-similarity-beyond-search/discovery.png caption="Reversed triplet loss" >}}
|
||||
|
||||
Multiple positive-negative pairs can be provided to make the discovery process more accurate.
|
||||
|
||||
@@ -133,7 +133,7 @@ The LLM is typically a model like GPT, BART or T5, trained on massive datasets t
|
||||

|
||||
|
||||
|
||||
The retriever and generator don't operate in isolation. The image bellow shows how the output of the retrieval feeds the generator to produce the final generated response.
|
||||
The retriever and generator don't operate in isolation. The image below shows how the output of the retrieval feeds the generator to produce the final generated response.
|
||||
|
||||
|
||||

|
||||
|
||||
@@ -18,7 +18,7 @@ As you can see from the charts, there are three main patterns:
|
||||
|
||||
- **Speed downturn** - some engines struggle to keep high RPS, it might be related to the requirement of building a filtering mask for the dataset, as described above.
|
||||
|
||||
- **Accuracy collapse** - some engines are loosing accuracy dramatically under some filters. It is related to the fact that the HNSW graph becomes disconnected, and the search becomes unreliable.
|
||||
- **Accuracy collapse** - some engines are losing accuracy dramatically under some filters. It is related to the fact that the HNSW graph becomes disconnected, and the search becomes unreliable.
|
||||
|
||||
Qdrant avoids all these problems and also benefits from the speed boost, as it implements an advanced [query planning strategy](/documentation/search/#query-planning).
|
||||
|
||||
|
||||
@@ -17,7 +17,7 @@ Unlisted: false
|
||||
|
||||
Most of the engines have improved since [our last run](/benchmarks/single-node-speed-benchmark-2022/). Both life and software have trade-offs but some clearly do better:
|
||||
|
||||
* **`Qdrant` achives highest RPS and lowest latencies in almost all the scenarios, no matter the precision threshold and the metric we choose.** It has also shown 4x RPS gains on one of the datasets.
|
||||
* **`Qdrant` achieves highest RPS and lowest latencies in almost all the scenarios, no matter the precision threshold and the metric we choose.** It has also shown 4x RPS gains on one of the datasets.
|
||||
* `Elasticsearch` has become considerably fast for many cases but it's very slow in terms of indexing time. It can be 10x slower when storing 10M+ vectors of 96 dimensions! (32mins vs 5.5 hrs)
|
||||
* `Milvus` is the fastest when it comes to indexing time and maintains good precision. However, it's not on-par with others when it comes to RPS or latency when you have higher dimension embeddings or more number of vectors.
|
||||
* `Redis` is able to achieve good RPS but mostly for lower precision. It also achieved low latency with single thread, however its latency goes up quickly with more parallel requests. Part of this speed gain comes from their custom protocol.
|
||||
|
||||
@@ -53,7 +53,7 @@ In the podcast, we addressed the following:
|
||||
- **Model evaluation(LLM)** - Understanding the model at the domain-level for the given use case, supporting required context length and terminology/concept understanding.
|
||||
- **Ingestion pipeline evaluation** - Evaluating factors related to data ingestion and processing such as chunk strategies, chunk size, chunk overlap, and more.
|
||||
- **Retrieval evaluation** - Understanding factors such as average precision, [Distributed cumulative gain](https://en.wikipedia.org/wiki/Discounted_cumulative_gain) (DCG), as well as normalized DCG.
|
||||
- **Generation evaluation(E2E)** - Establishing guardrails. Evaulating prompts. Evaluating the number of chunks needed to set up the context for generation.
|
||||
- **Generation evaluation(E2E)** - Establishing guardrails. Evaluating prompts. Evaluating the number of chunks needed to set up the context for generation.
|
||||
|
||||
### The recording
|
||||
|
||||
|
||||
@@ -89,7 +89,7 @@ Dataset: Laion 1 million 512d vectors
|
||||

|
||||
|
||||
Full-text filtering in Qdrant in an efficient way to combine Vector-based scoring with exact keyword match.
|
||||
And in v1.15 full-text index recieved a number of upgrades which make vector similarity evem more useful.
|
||||
And in v1.15 full-text index received a number of upgrades which make vector similarity even more useful.
|
||||
|
||||
### Multilingual Tokenization
|
||||
|
||||
@@ -172,7 +172,7 @@ PUT /collections/{collection_name}/index
|
||||
With [phrase matching](/documentation/concepts/filtering/#phrase-match), you can now perform exact phrase search.
|
||||
It allows you to search for a specific phrase, words in exact order, within a text field.
|
||||
|
||||
For efficient phrase seach Qdrant requires to build an additional data structure,
|
||||
For efficient phrase search Qdrant requires to build an additional data structure,
|
||||
so it needs to be configured during creation of the full-text index:
|
||||
|
||||
```http
|
||||
@@ -283,7 +283,7 @@ This modification, in combinations with [incremental HNSW indexing](/blog/qdrant
|
||||
|
||||
### HNSW Graph connectivity estimation
|
||||
|
||||
Qdrant builds [addtitional HNSW links](/articles/filterable-hnsw/) to ensure that filtered searches are performed fast and accurate.
|
||||
Qdrant builds [additional HNSW links](/articles/filterable-hnsw/) to ensure that filtered searches are performed fast and accurate.
|
||||
|
||||
It does, however, introduce an overhead for indexing complexity, especially when the number of payload indexes is large.
|
||||
With v1.15, Qdrant introduces an optimization, which quickly estimates graph connectivity before creating additional links.
|
||||
|
||||
@@ -75,7 +75,7 @@ Our inaugural Qdrant Stars are a diverse and talented lineup who have shown exce
|
||||
<div style="display: flex; align-items: center; margin-bottom: 20px;">
|
||||
<img src="/blog/qdrant-stars-announcement/Prof-Owen-Colegrove.jpeg" alt="Owen Colegrove" style="width: 200px; height: 200px; object-fit: cover; object-position: center; margin-right: 20px; margin-top: 20px;">
|
||||
<div>
|
||||
<p>Owen Colegrove is the Co-Founder of <a href="https://www.sciphi.ai/">SciPhi</a>, making it easy build, deploy, and scale RAG systems using Qdrant vector search tecnology. He has Ph.D. in Physics and was previously a Quantitative Strategist at Citadel and a Researcher at CERN.</p>
|
||||
<p>Owen Colegrove is the Co-Founder of <a href="https://www.sciphi.ai/">SciPhi</a>, making it easy build, deploy, and scale RAG systems using Qdrant vector search technology. He has Ph.D. in Physics and was previously a Quantitative Strategist at Citadel and a Researcher at CERN.</p>
|
||||
</div>
|
||||
</div>
|
||||
<blockquote>
|
||||
@@ -172,7 +172,7 @@ Share your journey with vector search technologies and how you plan to contribut
|
||||
|
||||
#### Nominate a Qdrant Star
|
||||
|
||||
Do you know someone who could be our next Qdrant Star? Please submit your nomination through our [nomination form](hhttps://forms.gle/jsEJ9zjdaxqk7F5b9), explaining why they're a great fit. Your recommendation could help us find the next standout ambassador.
|
||||
Do you know someone who could be our next Qdrant Star? Please submit your nomination through our [nomination form](https://forms.gle/jsEJ9zjdaxqk7F5b9), explaining why they're a great fit. Your recommendation could help us find the next standout ambassador.
|
||||
|
||||
#### Learn More
|
||||
|
||||
|
||||
@@ -74,7 +74,7 @@ Here is what this basic tutorial will teach you:
|
||||
|
||||
**3. Implement vector similarity search algorithms:** Second, you will create and test a chatbot that only uses the LLM. Then, you will enable the memory component offered by Qdrant. This will allow your chatbot to be modified and updated, giving it long-term memory.
|
||||
|
||||
**4. Optimize the chatbot's performance:** In the last step, you will query the chatbot in two ways. First query will retrieve parametric data from the LLM, while the second one will get contexual data via Qdrant.
|
||||
**4. Optimize the chatbot's performance:** In the last step, you will query the chatbot in two ways. First query will retrieve parametric data from the LLM, while the second one will get contextual data via Qdrant.
|
||||
|
||||
The goal of this exercise is to show that RAG is simple to implement via LangChain and yields much better results than using LLMs by itself.
|
||||
|
||||
|
||||
@@ -126,7 +126,7 @@ Noe Acache:
|
||||
So the training task was quite simple. And at the end, the embeddings was not learning any very complex features, so it was not really improving it. So jumping onto the areas of improvement, knowing all of that, the first thing I would do if I had to do it again will be to use the managed milboss for a better fine tuning, it would be to labyd hard examples, hard pairs. So, for instance, you know that when you have a matching pair where the similarity score is not too high or not too low, you know, it's where the model kind of struggles and you will find some good matching and also some mistakes. So it's where it kind of is interesting to level to then be able to fine tune your model and make it learn more complex things according to your tasks. Another possibility for fine tuning will be some sort of multilabel classification. So for instance, if you consider tab close, you could say, all right, those disclose contain buttons. It have a color, it have stripes.
|
||||
|
||||
Noe Acache:
|
||||
And for all of these categories, you'll get a score between zero and one. And concatenating all these scores together, you can get an embedding which you can put in a vector database for your vector search. It's kind of hard to scale because you need to do a specific model and labeling for each type of object. And I really wonder how Google lens does because their algorithm work very well. So are they working more like with this kind of functioning or this kind of functioning? So if anyone had any thought on that or any idea, again, I'd be happy to talk about it afterwards. And finally, I feel like we made a lot of advancements in multimodal training, trying to combine text inputs with image. We've made input to build some kind of complex embeddings. And how great would it be to have an image embeding you could guide with text.
|
||||
And for all of these categories, you'll get a score between zero and one. And concatenating all these scores together, you can get an embedding which you can put in a vector database for your vector search. It's kind of hard to scale because you need to do a specific model and labeling for each type of object. And I really wonder how Google lens does because their algorithm work very well. So are they working more like with this kind of functioning or this kind of functioning? So if anyone had any thought on that or any idea, again, I'd be happy to talk about it afterwards. And finally, I feel like we made a lot of advancements in multimodal training, trying to combine text inputs with image. We've made input to build some kind of complex embeddings. And how great would it be to have an image embedding you could guide with text.
|
||||
|
||||
Noe Acache:
|
||||
So you could just like when creating an embedding of your image, just say, all right, here, I don't care about the movements, I only care about the features on the object, for instance. And then it will learn an embedding according to your task without any fine tuning. I really feel like with the current state of the arts we are able to do this. I mean, we need to do it, but the technology is ready.
|
||||
|
||||
+1
-1
@@ -158,7 +158,7 @@ Demetrios:
|
||||
And so you kind of touched on this earlier, but can you say it again? Because I don't know if I fully grasped it. Where are all the places in the system that you are evaluating? Because it's not just the output. Right. And how do you look at evaluation as a system rather than just evaluating the output every once in a while?
|
||||
|
||||
Sourabh Agrawal:
|
||||
Yeah, so I mean, what we do is we plug with every part. So even if you start with retrieval, so we have a high level check where we look at the quality of retrieved context. And then we also have evaluations for every part of this retrieval pipeline. So if you're doing query rewrite, if you're doing re ranking, if you're doing sub question, we have evaluations for all of them. In fact, we have worked closely with the llama index team to kind of integrate with all of their modular pipelines. Secondly, once we cross the retrieval step, we have around five to six matrices on this retrieval part. Then we look at the response generation. We have their evaluations for different criterias.
|
||||
Yeah, so I mean, what we do is we plug with every part. So even if you start with retrieval, so we have a high level check where we look at the quality of retrieved context. And then we also have evaluations for every part of this retrieval pipeline. So if you're doing query rewrite, if you're doing re ranking, if you're doing sub question, we have evaluations for all of them. In fact, we have worked closely with the llama index team to kind of integrate with all of their modular pipelines. Secondly, once we cross the retrieval step, we have around five to six matrices on this retrieval part. Then we look at the response generation. We have their evaluations for different criteria.
|
||||
|
||||
Sourabh Agrawal:
|
||||
So conciseness, completeness, safety, jailbreaks, prompt injections, as well as you can define your custom guidelines. So you can say that, okay, if the user is asking anything and related to code, the output should also give an example code snippet so you can just in plain English, define this guideline. And we check for that. And then finally, like zooming out, we also have checks. We look at conversations as a whole, how the user is satisfied, how many turns it requires for them to, for the chatbot or the LLM to answer the user. Yeah, that's how we look at the whole evaluations as a whole.
|
||||
|
||||
@@ -54,7 +54,7 @@ André highlighted the underlying forces driving this shift:
|
||||
|
||||
We are convinced that if AI is going to evolve beyond static assistants, it needs a **retrieval layer built for unstructured data and agent workflows**.
|
||||
|
||||
Next on stage, our Co-Founder and CTO [**Andrey Vasnetsov**](https://www.linkedin.com/in/andrey-vasnetsov-75268897/) emphazised our belief that ‘vector database’ is actually the wrong term to describe what we are building at Qdrant. **Qdrant is not “a vector database”** because vectors themselves are not data, but representations.
|
||||
Next on stage, our Co-Founder and CTO [**Andrey Vasnetsov**](https://www.linkedin.com/in/andrey-vasnetsov-75268897/) emphasized our belief that ‘vector database’ is actually the wrong term to describe what we are building at Qdrant. **Qdrant is not “a vector database”** because vectors themselves are not data, but representations.
|
||||
|
||||

|
||||
|
||||
|
||||
@@ -27,7 +27,7 @@ Every individual upsert call initiates a transaction that consumes memory and di
|
||||
|
||||
## Choosing Your Ingestion Strategy
|
||||
|
||||
Qdrant provides several methods for data ingestion, each tailored to different scales and use cases. It should be noted that only the Python client supports the upload_points and upload_collection methods. If you're using Qdrant on a different client then we reccomend using upsert with batch upload for large scale ingestion. [Learn more about bulk operations](/documentation/guides/bulk-operations/).
|
||||
Qdrant provides several methods for data ingestion, each tailored to different scales and use cases. It should be noted that only the Python client supports the upload_points and upload_collection methods. If you're using Qdrant on a different client then we recommend using upsert with batch upload for large scale ingestion. [Learn more about bulk operations](/documentation/guides/bulk-operations/).
|
||||
|
||||
- **upsert (Individual or Batched)**: This is the fundamental operation for adding or updating points. Individual upserts are best suited for real-time updates while batching works best for larger workloads.
|
||||
|
||||
|
||||
@@ -69,7 +69,7 @@ If you use multiple accounts for different purposes, it is a good idea to give t
|
||||
|
||||
### Changing the Account Owner
|
||||
|
||||
Every account has one owner. The owner is granted full admin permissions for the account as well as futher unique permissions allowing them to either delete the account or transfer account ownership.
|
||||
Every account has one owner. The owner is granted full admin permissions for the account as well as further unique permissions allowing them to either delete the account or transfer account ownership.
|
||||
|
||||
To transfer ownership of an account, as the owner, visit the *Access Management* page. In the actions menu of the user you wish to transfer to, you will find the option 'Make Account Owner' which begins the transfer.
|
||||
|
||||
|
||||
@@ -19,7 +19,7 @@ To cater to diverse integration needs, the Qdrant Cloud API offers two primary i
|
||||
* **REST/JSON API**: A conventional HTTP/1.1 (and HTTP/2) interface with JSON payloads. This API is provided via a gRPC Gateway, translating RESTful calls into gRPC messages, offering ease of use for web clients, scripts, and broader tool compatibility.
|
||||
|
||||
You can find the API definitions and generated client libraries in our Qdrant Cloud Public API [GitHub repository](https://github.com/qdrant/qdrant-cloud-public-api).
|
||||
**Note:** The API is splitted into multiple services to make it easier to use.
|
||||
**Note:** The API is split into multiple services to make it easier to use.
|
||||
|
||||
### Qdrant Cloud API Endpoints
|
||||
|
||||
|
||||
@@ -58,7 +58,7 @@ To subscribe:
|
||||
2. Select **GCP Marketplace** as the payment method. You will be redirected to the GCP Marketplace listing for Qdrant.
|
||||
3. Select **Subscribe**. (If you have already subscribed, select **Manage on Provider**.)
|
||||
4. On the next screen, choose options as required, and select **Subscribe**.
|
||||
5. On the pop-up window that appers, select **Sign up with Qdrant**.
|
||||
5. On the pop-up window that appears, select **Sign up with Qdrant**.
|
||||
|
||||
You will be redirected to the Billing Details screen in the [Qdrant Cloud Console](https://cloud.qdrant.io/). From there you can start to create Qdrant database clusters.
|
||||
|
||||
|
||||
@@ -42,7 +42,7 @@ Qdrant clusters in Hybrid Cloud also run in hardened, unprivileged containers wi
|
||||
|
||||
### Private Cloud
|
||||
|
||||
In Qdrant Private Cloud, Qdrant clusters run completely isolated and air-gapped within your infrastucture without any connection to the Qdrant Cloud Console.
|
||||
In Qdrant Private Cloud, Qdrant clusters run completely isolated and air-gapped within your infrastructure without any connection to the Qdrant Cloud Console.
|
||||
|
||||
Since there is no connection or communication with Qdrant, you are fully responsible for the security of the entire Qdrant Private Cloud installation. This also means that you do not benefit from the integrated management and observability features of Qdrant Managed Cloud and Hybrid Cloud.
|
||||
|
||||
|
||||
@@ -27,9 +27,9 @@ Have a look at the [API reference](/documentation/interfaces/#api-reference) and
|
||||
|
||||
## Node Specific Endpoints
|
||||
|
||||
Next to the cluster endpoint which loadbalances requests across all healthy Qdrant nodes, each node in the cluster has its own endpoint as well. This is mainly usefull for monitoring or manual shard management purpuses.
|
||||
Next to the cluster endpoint which loadbalances requests across all healthy Qdrant nodes, each node in the cluster has its own endpoint as well. This is mainly useful for monitoring or manual shard management purposes.
|
||||
|
||||
You can finde the node specific endpoints on the cluster detail page in the Qdrant Cloud Console.
|
||||
You can find the node specific endpoints on the cluster detail page in the Qdrant Cloud Console.
|
||||
|
||||

|
||||
|
||||
|
||||
@@ -9,10 +9,10 @@ Qdrant Cloud offers several advanced configuration options to optimize clusters
|
||||
|
||||
The cloud platform does not expose all [configuration options](/documentation/guides/configuration/) available in Qdrant. We have selected the relevant options that are explained in detail below.
|
||||
|
||||
In adition the cloud platform automatically configures the following settings for your cluster to ensure optimal performance and reliability:
|
||||
In addition the cloud platform automatically configures the following settings for your cluster to ensure optimal performance and reliability:
|
||||
|
||||
* The maximum number of collections in a cluster is set to 1000. Larger numbers of collections lead to performance degradation. For more information see [Multitenancy](/documentation/guides/multiple-partitions/).
|
||||
* Strict mode is activated by default for new collections enforcing that all filters being used in retrieve and udpate queries are indexed. This improves performance and reliability. You can disable this individually for each collection. For more information see [Strict Mode](/documentation/guides/administration/#strict-mode).
|
||||
* Strict mode is activated by default for new collections enforcing that all filters being used in retrieve and update queries are indexed. This improves performance and reliability. You can disable this individually for each collection. For more information see [Strict Mode](/documentation/guides/administration/#strict-mode).
|
||||
* The cluster mode is automatically enabled to allow distributed deployments and horizontal scaling.
|
||||
* The maximum amount of payload indexes per collection is set to 100. Larger numbers of payload indexes lead to performance degradation (starting with Qdrant v1.16.0).
|
||||
|
||||
|
||||
@@ -28,7 +28,7 @@ The following guided samples help you get started with real-world projects using
|
||||
|
||||
## Example Notebooks
|
||||
|
||||
Our Notebooks offer complex instructions that are supported with a throrough explanation. Follow along by trying out the code and get the most out of each example.
|
||||
Our Notebooks offer complex instructions that are supported with a thorough explanation. Follow along by trying out the code and get the most out of each example.
|
||||
|
||||
| Example | Description | Stack |
|
||||
|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------|----------------------------|
|
||||
|
||||
@@ -258,7 +258,7 @@ The output should look like following:
|
||||
Our web service is implemented, yet running only on our local machine. It has to be exposed to the public before
|
||||
Command-R can interact with it. For a quick experiment, it might be enough to set up tunneling using services such as
|
||||
[ngrok](https://ngrok.com/). We won't cover all the details in the tutorial, but their
|
||||
[Quickstart](https://ngrok.com/docs/guides/getting-started/) is a great resource describing the process step-by-step.
|
||||
[Quickstart](https://ngrok.com/docs/getting-started/) is a great resource describing the process step-by-step.
|
||||
Alternatively, you can also deploy the service with a public URL.
|
||||
|
||||
Once it's done, we can create the connector first, and then tell the model to use it, while interacting through the chat
|
||||
|
||||
+1
-1
@@ -237,7 +237,7 @@ search_pipeline = Pipeline()
|
||||
```
|
||||
|
||||
Our second process takes user input, converts it into embeddings and then searches for the most relevant documents
|
||||
using the query embedding. This might look familiar, but we arent working with `Document` instances
|
||||
using the query embedding. This might look familiar, but we aren't working with `Document` instances
|
||||
anymore, since the query only accepts raw text. Thus, some of the components will be different, especially the embedder,
|
||||
as it has to accept a single string as an input and produce a single embedding as an output:
|
||||
|
||||
|
||||
@@ -13,7 +13,7 @@ $$
|
||||
$$
|
||||
|
||||
A detailed breakdown of the idea behind miniCOIL can be found in the
|
||||
["miniCOIL: on the road to Usable Sparse Neural Retreival" article](https://qdrant.tech/articles/minicoil/) or, in a [recorded talk "miniCOIL: Sparse Neural Retrieval Done Right"](https://youtu.be/f1sBJMSgBXA?si=G3C5--UVRKAW5WJ0).
|
||||
["miniCOIL: on the road to Usable Sparse Neural Retrieval" article](https://qdrant.tech/articles/minicoil/) or, in a [recorded talk "miniCOIL: Sparse Neural Retrieval Done Right"](https://youtu.be/f1sBJMSgBXA?si=G3C5--UVRKAW5WJ0).
|
||||
|
||||
This tutorial will demonstrate how miniCOIL-based sparse neural retrieval performs compared to BM25-based lexical retrieval.
|
||||
|
||||
|
||||
@@ -41,7 +41,7 @@ TextCrossEncoder.list_supported_models()
|
||||
This command displays the available models, including details such as output embedding dimensions, model description, model size, model sources, and model file.
|
||||
|
||||
<details>
|
||||
<summary> <span style="background-color: gray; color: black;"> Avaliable models </span> </summary>
|
||||
<summary> <span style="background-color: gray; color: black;"> Available models </span> </summary>
|
||||
|
||||
|
||||
```python
|
||||
|
||||
@@ -92,7 +92,7 @@ retrieved_info = ar.run_vector_retriever(
|
||||
print(retrieved_info)
|
||||
```
|
||||
|
||||
You can refer to the Camel [documentation](https://docs.camel-ai.org/index.html) for more information about the retrieval mechansims.
|
||||
You can refer to the Camel [documentation](https://docs.camel-ai.org/index.html) for more information about the retrieval mechanisms.
|
||||
|
||||
## End-To-End Examples
|
||||
|
||||
|
||||
@@ -82,5 +82,5 @@ You can scale this process with a dataset (e.g. from Hugging Face) and evaluate
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [End-to-end Evalutation Example](https://github.com/qdrant/qdrant-rag-eval/blob/master/workshop-rag-eval-qdrant-deepeval/notebook/rag_eval_qdrant_deepeval.ipynb)
|
||||
- [End-to-end Evaluation Example](https://github.com/qdrant/qdrant-rag-eval/blob/master/workshop-rag-eval-qdrant-deepeval/notebook/rag_eval_qdrant_deepeval.ipynb)
|
||||
- [DeepEval documentation](https://deepeval.com)
|
||||
|
||||
@@ -11,7 +11,7 @@ By integrating Qdrant with HoneyHive, you can:
|
||||
- Trace vector database operations
|
||||
- Monitor latency, embedding quality, and context relevance
|
||||
- Evaluate retrieval performance in your RAG pipelines
|
||||
- Optimize paramaters such as `chunk_size` or `chunk_overlap`
|
||||
- Optimize parameters such as `chunk_size` or `chunk_overlap`
|
||||
|
||||
## Prerequisites
|
||||
|
||||
|
||||
@@ -18,7 +18,7 @@ It might be installed with pip:
|
||||
pip install langchain-qdrant
|
||||
```
|
||||
|
||||
The integration supports searching for relevant documents usin dense/sparse and hybrid retrieval.
|
||||
The integration supports searching for relevant documents using dense/sparse and hybrid retrieval.
|
||||
|
||||
Qdrant acts as a vector index that may store the embeddings with the documents used to generate them. There are various ways to use it, but calling `QdrantVectorStore.from_texts` or `QdrantVectorStore.from_documents` is probably the most straightforward way to get started:
|
||||
|
||||
|
||||
@@ -79,4 +79,4 @@ vn.ask(question="<YOUR_QUESTION>")
|
||||
|
||||
- [Getting started with Vanna.AI](https://vanna.ai/docs/app/)
|
||||
- [Vanna.AI documentation](https://vanna.ai/docs/)
|
||||
- [Source Code](https://github.com/vanna-ai/vanna/tree/main/src/vanna/qdrant)
|
||||
- [Source Code](https://github.com/vanna-ai/vanna/tree/main/src/vanna/integrations/qdrant)
|
||||
|
||||
+1
-1
@@ -1 +1 @@
|
||||
This code snipet shows how to load a dataset and upload dense and sparse vectors to Qdrant. While uploading the vectors we also include a payload known as text.
|
||||
This code snippet shows how to load a dataset and upload dense and sparse vectors to Qdrant. While uploading the vectors we also include a payload known as text.
|
||||
+1
-1
@@ -1 +1 @@
|
||||
This code snippet demonstrates how to create a full-text index for a specified field in a collection with stopwords configuration. There are 2 examples of Stopwords configuration, simple and explicit. Simple configuration specifies only a signle language, only pre-defined stopwords from this language will be used. Explicit configurations configures combination of stopwords from 2 languages and one custom stopword.
|
||||
This code snippet demonstrates how to create a full-text index for a specified field in a collection with stopwords configuration. There are 2 examples of Stopwords configuration, simple and explicit. Simple configuration specifies only a single language, only pre-defined stopwords from this language will be used. Explicit configurations configures combination of stopwords from 2 languages and one custom stopword.
|
||||
+1
-1
@@ -1,2 +1,2 @@
|
||||
This code snippet is for a PUT request to insert points into a collection, where each point has an ID, a payload containing a group ID, and a vector. The code illustrates partitioning vectors by user to ensure that each user can only access their own vectors. It emphasizes adding a `group_id` field to each vector in the collection, facilitating user-specific data access control. Additionally, it suggests using an appropriate naming convention for the key in the payload for flexibility in data structures.
|
||||
In addition, the snippet includes a shard key selector, allowing a dymamic routing between shared and dedicated shards based on the existance of `target` shard in the collection.
|
||||
In addition, the snippet includes a shard key selector, allowing a dynamic routing between shared and dedicated shards based on the existence of `target` shard in the collection.
|
||||
|
||||
@@ -11,7 +11,7 @@ Alongside Hybrid Cloud specific scheduling options, you can also adjust various
|
||||
|
||||
## Scale Clusters
|
||||
|
||||
Hybrid cloud clusters can be scaled up and down, horizontall and vertically, at any time. For more details see [Scale Clusters](/documentation/cloud/cluster-scaling/).
|
||||
Hybrid cloud clusters can be scaled up and down, horizontally and vertically, at any time. For more details see [Scale Clusters](/documentation/cloud/cluster-scaling/).
|
||||
|
||||
### Automatic Shard Rebalancing
|
||||
|
||||
|
||||
@@ -196,7 +196,7 @@ By default, Qdrant Cloud will reserve 20% of available CPU and memory on each Po
|
||||
|
||||
You can modify this reservation in the “Configuration” section of the Qdrant Cluster detail page.
|
||||
|
||||
If you want to check how much resources are availabe on an empty Kubernetes node, you can use the following command:
|
||||
If you want to check how much resources are available on an empty Kubernetes node, you can use the following command:
|
||||
|
||||
```shell
|
||||
kubectl describe node <node-name>
|
||||
|
||||
@@ -60,7 +60,7 @@ By default, Qdrant Cloud will provision two volumes per Qdrant Pod: One for the
|
||||
|
||||
5. (Optional) If you have special requirements for any of the following, activate the **Show advanced configuration** option:
|
||||
|
||||
- If you use a proxy to connect from your infrastructure to the Qdrant Cloud API, you can specify the proxy URL, credentials and cetificates.
|
||||
- If you use a proxy to connect from your infrastructure to the Qdrant Cloud API, you can specify the proxy URL, credentials and certificates.
|
||||
- Container registry URL for Qdrant services (like Agent, Operator, Cluster-manager and monitoring stack) images. The default is <https://registry.cloud.qdrant.io/qdrant/>.
|
||||
- Helm chart repository URL for the Qdrant services. The default is <oci://registry.cloud.qdrant.io/qdrant-charts>.
|
||||
- An optional secret with credentials to access your own container registry.
|
||||
|
||||
@@ -144,6 +144,6 @@ At this point it is safe to delete the tenant's data from the shared Fallback Sh
|
||||
|
||||
### Limitations
|
||||
|
||||
- Currently, `fallback` Shard may only contain a single shard ID on its own. That means all small tenants must fit a single peer of the cluser. This restriction will be improved in future releases.
|
||||
- Currently, `fallback` Shard may only contain a single shard ID on its own. That means all small tenants must fit a single peer of the cluster. This restriction will be improved in future releases.
|
||||
- Similar to collections, dedicated Shards introduce some resource overhead. It is not recommended to create more than a thousand dedicated Shards per cluster. Recommended threshold of promoting a tenant is the same as the indexing threshold for a single collection, which is around 20K points.
|
||||
|
||||
|
||||
@@ -183,7 +183,7 @@ Similarly, you can use inference at query time by providing the text or image to
|
||||
|
||||
## Datatypes
|
||||
|
||||
Newest versions of embeddings models generate vectors with very large dimentionalities.
|
||||
Newest versions of embeddings models generate vectors with very large dimensionalities.
|
||||
With OpenAI's `text-embedding-3-large` embedding model, the dimensionality can go up to 3072.
|
||||
|
||||
The amount of memory required to store such vectors grows linearly with the dimensionality,
|
||||
|
||||
@@ -40,5 +40,5 @@ With the LLM Observability data now being collected by OpenLIT, the next step is
|
||||
|
||||
To begin exploring your LLM Application's performance data within the OpenLIT UI, please see the [Quickstart Guide](https://docs.openlit.io/latest/quickstart).
|
||||
|
||||
If you want to integrate and send the generated metrics and traces to your existing observability tools like Promethues+Jaeger, Grafana or more, refer to the [Official Documentation for OpenLIT Connections](https://docs.openlit.io/latest/connections/intro) for detailed instructions.
|
||||
If you want to integrate and send the generated metrics and traces to your existing observability tools like Prometheus+Jaeger, Grafana or more, refer to the [Official Documentation for OpenLIT Connections](https://docs.openlit.io/) for detailed instructions.
|
||||
|
||||
|
||||
@@ -305,10 +305,10 @@ If you anticipate a lot of growth, we recommend 12 shards since you can expand f
|
||||
|
||||
Shards are evenly distributed across all existing nodes when a collection is first created.
|
||||
|
||||
When you add or remove nodes from the cluster, rebalancing of existing shards accross the nodes depends on how you've deployed the cluster:
|
||||
When you add or remove nodes from the cluster, rebalancing of existing shards across the nodes depends on how you've deployed the cluster:
|
||||
|
||||
- In Qdrant Cloud, shards are [balanced across the nodes automatically](/documentation/cloud/configure-cluster/#shard-rebalancing).
|
||||
- If your cluster is not runnning in Qdrant Cloud, you need to [manually balance shards](#moving-shards).
|
||||
- If your cluster is not running in Qdrant Cloud, you need to [manually balance shards](#moving-shards).
|
||||
|
||||
### Resharding
|
||||
|
||||
@@ -1326,7 +1326,7 @@ Listener node will not participate in search operations, but will still accept w
|
||||
|
||||
All shards, stored on the listener node, will be converted to the `Listener` state.
|
||||
|
||||
Additionally, all write requests sent to the listener node will be processed with `wait=false` option, which means that the write oprations will be considered successful once they are written to WAL.
|
||||
Additionally, all write requests sent to the listener node will be processed with `wait=false` option, which means that the write operations will be considered successful once they are written to WAL.
|
||||
This mechanism should allow to minimize upsert latency in case of parallel snapshotting.
|
||||
|
||||
## Consensus Checkpointing
|
||||
|
||||
@@ -11,7 +11,7 @@ aliases:
|
||||
|
||||
The Qdrant open-source container image collects anonymized usage statistics from users in order to improve the engine by default. You can [deactivate](#deactivate-telemetry) at any time, and any data that has already been collected can be [deleted on request](#request-information-deletion).
|
||||
|
||||
Deactivating this will not affect your ability to monitor the Qdrant database yourself by accessing the `/metrics` or `/telemetry` endpoints of your database. It will just stop sending independend, anonymized usage statistics to the Qdrant team.
|
||||
Deactivating this will not affect your ability to monitor the Qdrant database yourself by accessing the `/metrics` or `/telemetry` endpoints of your database. It will just stop sending independent, anonymized usage statistics to the Qdrant team.
|
||||
|
||||
<aside role="status">When using Qdrant Cloud, this setting does not apply and anonymized usage statistics are disabled by default.</aside>
|
||||
|
||||
|
||||
@@ -24,7 +24,7 @@ If you are looking for a specific topic in a particular book, you can try to fin
|
||||
|
||||
Time passed, and we haven’t had much change in that area for quite a long time. But our textual data collection started to grow at a greater pace. So we also started building up many processes around those inverted indexes. For example, we allowed our users to provide many words and started splitting them into pieces. That allowed finding some documents which do not necessarily contain all the query words, but possibly part of them. We also started converting words into their root forms to cover more cases, removing stopwords, etc. Effectively we were becoming more and more user-friendly. Still, the idea behind the whole process is derived from the most straightforward keyword-based search known since the Middle Ages, with some tweaks.
|
||||
|
||||
{{< figure src=/docs/gettingstarted/tokenization.png caption="The process of tokenization with an additional stopwords removal and converstion to root form of a word." >}}
|
||||
{{< figure src=/docs/gettingstarted/tokenization.png caption="The process of tokenization with an additional stopwords removal and conversion to root form of a word." >}}
|
||||
|
||||
Technically speaking, we encode the documents and queries into so-called sparse vectors where each position has a corresponding word from the whole dictionary. If the input text contains a specific word, it gets a non-zero value at that position. But in reality, none of the texts will contain more than hundreds of different words. So the majority of vectors will have thousands of zeros and a few non-zero values. That’s why we call them sparse. And they might be already used to calculate some word-based similarity by finding the documents which have the biggest overlap.
|
||||
|
||||
|
||||
@@ -5,9 +5,9 @@ aliases: [ ../frameworks/buildship/ ]
|
||||
|
||||
# BuildShip
|
||||
|
||||
[BuildShip](https://buildship.com/) is a low-code visual builder to create APIs, scheduled jobs, and backend workflows with AI assitance.
|
||||
[BuildShip](https://buildship.com/) is a low-code visual builder to create APIs, scheduled jobs, and backend workflows with AI assistance.
|
||||
|
||||
You can use the [Qdrant integration](https://buildship.com/integrations/qdrant) to development workflows with semantic-search capabilites.
|
||||
You can use the [Qdrant integration](https://buildship.com/integrations/qdrant) to development workflows with semantic-search capabilities.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
|
||||
@@ -915,7 +915,7 @@ _Appears in:_
|
||||
| `security` _[QdrantSecurityContext](#qdrantsecuritycontext)_ | Security specifies the security context for each Qdrant node. | | |
|
||||
| `tolerations` _[Toleration](https://kubernetes.io/docs/reference/generated/kubernetes-api/v1.28/#toleration-v1-core) array_ | Tolerations specifies the tolerations for each Qdrant node. | | |
|
||||
| `nodeSelector` _object (keys:string, values:string)_ | NodeSelector specifies the node selector for each Qdrant node. | | |
|
||||
| `config` _[QdrantConfiguration](#qdrantconfiguration)_ | Config specifies the Qdrant configuration setttings for the clusters. | | |
|
||||
| `config` _[QdrantConfiguration](#qdrantconfiguration)_ | Config specifies the Qdrant configuration settings for the clusters. | | |
|
||||
| `ingress` _[Ingress](#ingress)_ | Ingress specifies the ingress for the cluster. | | |
|
||||
| `service` _[KubernetesService](#kubernetesservice)_ | Service specifies the configuration of the Qdrant Kubernetes Service. | | |
|
||||
| `gpu` _[GPU](#gpu)_ | GPU specifies GPU configuration for the cluster. If this field is not set, no GPU will be used. | | |
|
||||
|
||||
@@ -291,7 +291,7 @@ step certificate create mydomain.com qdrant-nodes.crt qdrant-nodes.key \
|
||||
|
||||
## GPU support
|
||||
|
||||
Starting with Qdrant 1.13 and private-cloud version 1.6.1 you can create a cluster that uses GPUs to accelarate indexing.
|
||||
Starting with Qdrant 1.13 and private-cloud version 1.6.1 you can create a cluster that uses GPUs to accelerate indexing.
|
||||
|
||||
As a prerequisite, you need to have a Kubernetes cluster with GPU support. You can check the [Kubernetes documentation](https://kubernetes.io/docs/tasks/manage-gpus/scheduling-gpus/) for generic information on GPUs and Kubernetes, or the documentation of your specific Kubernetes distribution.
|
||||
|
||||
|
||||
@@ -366,7 +366,7 @@ Can be applied to [float](/documentation/concepts/payload/#float) and [integer](
|
||||
### Datetime Range
|
||||
|
||||
The datetime range is a unique range condition, used for [datetime](/documentation/concepts/payload/#datetime) payloads, which supports RFC 3339 formats.
|
||||
You do not need to convert dates to UNIX timestaps. During comparison, timestamps are parsed and converted to UTC.
|
||||
You do not need to convert dates to UNIX timestamps. During comparison, timestamps are parsed and converted to UTC.
|
||||
|
||||
_Available as of v1.8.0_
|
||||
|
||||
|
||||
+1
-1
@@ -78,7 +78,7 @@ spec:
|
||||
app.kubernetes.io/name: operator
|
||||
```
|
||||
|
||||
The example aboves assumes that your Qdrant database and the cloud platform exporter are deployed in the `qdrant` namespace. Adjust the `namespaceSelector` and `namespace` fields according to your deployment.
|
||||
The example above assumes that your Qdrant database and the cloud platform exporter are deployed in the `qdrant` namespace. Adjust the `namespaceSelector` and `namespace` fields according to your deployment.
|
||||
|
||||
## Step 3: Access Grafana
|
||||
|
||||
|
||||
@@ -70,7 +70,7 @@ Qdrant will use vector embeddings of our facts to enrich the original prompt wit
|
||||
|
||||
We'll be using the [bge-base-en-v1.5](https://huggingface.co/BAAI/bge-small-en-v1.5) model via [FastEmbed](https://github.com/qdrant/fastembed/) - A lightweight, fast, Python library for embeddings generation.
|
||||
|
||||
The Qdrant client provides a handy integration with FastEmbed that makes building a knowledge base very straighforward.
|
||||
The Qdrant client provides a handy integration with FastEmbed that makes building a knowledge base very straightforward.
|
||||
|
||||
First, we need to create a collection, so Qdrant would know what vectors it will be dealing with, and then, we just pass our raw documents
|
||||
wrapped into `models.Document` to compute and upload the embeddings.
|
||||
|
||||
@@ -223,15 +223,15 @@ Alternatively, you can use the `wget` command:
|
||||
```bash
|
||||
wget https://node-0.my-cluster.com:6333/collections/test_collection/snapshots/test_collection-559032209313046-2024-01-03-13-20-11.snapshot \
|
||||
--header="api-key: ${QDRANT_API_KEY}" \
|
||||
-O node-0-shapshot.snapshot
|
||||
-O node-0-snapshot.snapshot
|
||||
|
||||
wget https://node-1.my-cluster.com:6333/collections/test_collection/snapshots/test_collection-559032209313047-2024-01-03-13-20-12.snapshot \
|
||||
--header="api-key: ${QDRANT_API_KEY}" \
|
||||
-O node-1-shapshot.snapshot
|
||||
-O node-1-snapshot.snapshot
|
||||
|
||||
wget https://node-2.my-cluster.com:6333/collections/test_collection/snapshots/test_collection-559032209313048-2024-01-03-13-20-13.snapshot \
|
||||
--header="api-key: ${QDRANT_API_KEY}" \
|
||||
-O node-2-shapshot.snapshot
|
||||
-O node-2-snapshot.snapshot
|
||||
```
|
||||
|
||||
The snapshots are now stored locally. We can use them to restore the collection to a different Qdrant instance, or treat them as a backup. We will create another collection using the same data on the same cluster.
|
||||
@@ -262,17 +262,17 @@ Alternatively, you can use the `curl` command:
|
||||
curl -X POST 'https://node-0.my-cluster.com:6333/collections/test_collection_import/snapshots/upload?priority=snapshot' \
|
||||
-H 'api-key: ${QDRANT_API_KEY}' \
|
||||
-H 'Content-Type:multipart/form-data' \
|
||||
-F 'snapshot=@node-0-shapshot.snapshot'
|
||||
-F 'snapshot=@node-0-snapshot.snapshot'
|
||||
|
||||
curl -X POST 'https://node-1.my-cluster.com:6333/collections/test_collection_import/snapshots/upload?priority=snapshot' \
|
||||
-H 'api-key: ${QDRANT_API_KEY}' \
|
||||
-H 'Content-Type:multipart/form-data' \
|
||||
-F 'snapshot=@node-1-shapshot.snapshot'
|
||||
-F 'snapshot=@node-1-snapshot.snapshot'
|
||||
|
||||
curl -X POST 'https://node-2.my-cluster.com:6333/collections/test_collection_import/snapshots/upload?priority=snapshot' \
|
||||
-H 'api-key: ${QDRANT_API_KEY}' \
|
||||
-H 'Content-Type:multipart/form-data' \
|
||||
-F 'snapshot=@node-2-shapshot.snapshot'
|
||||
-F 'snapshot=@node-2-snapshot.snapshot'
|
||||
```
|
||||
|
||||
|
||||
|
||||
+1
-1
@@ -134,7 +134,7 @@ if not client.collection_exists("startups"):
|
||||
```
|
||||
|
||||
Qdrant requires vectors to have their own names and configurations.
|
||||
Parameters `size` and `distance` are mandatory, however, you can additionaly specify extended configuration for your vectors, like `quantization_config` or `hnsw_config`.
|
||||
Parameters `size` and `distance` are mandatory, however, you can additionally specify extended configuration for your vectors, like `quantization_config` or `hnsw_config`.
|
||||
|
||||
|
||||
4. Read data from the file.
|
||||
|
||||
Reference in New Issue
Block a user