mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-28 23:48:31 +02:00
Merge branch 'master' into gui-quickstart
This commit is contained in:
@@ -5,14 +5,14 @@ hideTOC: true
|
||||
---
|
||||
# Documentation
|
||||
|
||||
Qdrant is an AI-native vector dabatase and a semantic search engine. You can use it to extract meaningful information from unstructured data. **[Learn more about vector search](/documentation/overview/)** and how it works with AI.
|
||||
Qdrant is an AI-native vector dabatase and a semantic search engine. You can use it to extract meaningful information from unstructured data. Want to see how it works? [Clone this repo now](https://github.com/qdrant/qdrant_demo/) and build a search engine in five minutes.
|
||||
|
||||
|||
|
||||
|-:|:-|
|
||||
|[Docker Quickstart](/documentation/quick-start/)|[Visual Quickstart](/documentation/quickstart-gui/)|
|
||||
|Use Qdrant Client SDKs|Try the GUI Dashboard|
|
||||
|[Local Quickstart](/documentation/quick-start/)|[Cloud Quickstart](/documentation/cloud/quickstart-cloud/)|
|
||||
|
||||
## Ready to start developing?
|
||||
|
||||
## Ready to start developing?
|
||||
|
||||
***<p style="text-align: center;">Qdrant is open-source and can be self-hosted. However, the quickest way to get started is with our [free tier](https://qdrant.to/cloud) on Qdrant Cloud. It scales easily and provides an UI where you can interact with data.</p>***
|
||||
|
||||
|
||||
@@ -1,9 +1,14 @@
|
||||
---
|
||||
title: Aleph Alpha
|
||||
weight: 900
|
||||
aliases: [ ../integrations/aleph-alpha/ ]
|
||||
aliases:
|
||||
- /documentation/examples/aleph-alpha-search/
|
||||
- /documentation/tutorials/aleph-alpha-search/
|
||||
- /documentation/integrations/aleph-alpha/
|
||||
---
|
||||
|
||||
# Using Aleph Alpha Embeddings with Qdrant
|
||||
|
||||
Aleph Alpha is a multimodal and multilingual embeddings' provider. Their API allows creating the embeddings for text and images, both
|
||||
in the same latent space. They maintain an [official Python client](https://github.com/Aleph-Alpha/aleph-alpha-client) that might be
|
||||
installed with pip:
|
||||
|
||||
@@ -6,8 +6,6 @@ weight: 35
|
||||
|
||||
| End-to-End Code Samples | Description | Stack |
|
||||
|---------------------------------------------------------------------------------|-------------------------------------------------------------------|---------------------------------------------|
|
||||
| [Aleph Alpha Search](../examples/aleph-alpha-search/) | Build a multimodal search that combines text and image data. | Qdrant, Aleph Alpha |
|
||||
| [Mighty Semantic Search](../examples/mighty/) | Build a simple semantic search with an on-demand NLP service. | Qdrant, Mighty |
|
||||
| [Multitenancy with LlamaIndex](../examples/llama-index-multitenancy/) | Handle data coming from multiple users in LlamaIndex. | Qdrant, Python, LlamaIndex |
|
||||
| [Implement custom connector for Cohere RAG](../examples/cohere-rag-connector/) | Bring data stored in Qdrant to Cohere RAG | Qdrant, Cohere, FastAPI |
|
||||
| [Chatbot for Interactive Learning](../examples/rag-chatbot-red-hat-openshift-haystack/) | Build a Private RAG Chatbot for Interactive Learning | Qdrant, Haystack, OpenShift |
|
||||
@@ -42,3 +40,4 @@ Our Notebooks offer complex instructions that are supported with a throrough exp
|
||||
| Example | Description | Stack |
|
||||
|---------------------------------------------------------------------------------|-------------------------------------------------------------------|---------------------------------------------|
|
||||
| [Pinecone to Qdrant Data Transfer](https://githubtocolab.com/qdrant/examples/blob/master/data-migration/from-pinecone-to-qdrant.ipynb) | Migrate your vector data from Pinecone to Qdrant. | Qdrant, Vector-io |
|
||||
| [Stream Data to Qdrant with Kafka](../examples/data-streaming-kafka-qdrant/) | Use Confluent to Stream Data to Qdrant via Managed Kafka. | Qdrant, Kafka |
|
||||
@@ -1,8 +1,7 @@
|
||||
---
|
||||
title: Aleph Alpha Search
|
||||
weight: 16
|
||||
aliases:
|
||||
- /documentation/tutorials/aleph-alpha-search/
|
||||
draft: true
|
||||
---
|
||||
|
||||
# Multimodal Semantic Search with Aleph Alpha
|
||||
|
||||
@@ -0,0 +1,270 @@
|
||||
---
|
||||
title: How to Setup Seamless Data Streaming with Kafka and Qdrant
|
||||
weight: 49
|
||||
---
|
||||
|
||||
# Setup Data Streaming with Kafka via Confluent
|
||||
|
||||
**Author:** [M K Pavan Kumar](https://www.linkedin.com/in/kameshwara-pavan-kumar-mantha-91678b21/) , research scholar at [IIITDM, Kurnool](https://iiitk.ac.in). Specialist in hallucination mitigation techniques and RAG methodologies.
|
||||
• [GitHub](https://github.com/pavanjava) • [Medium](https://medium.com/@manthapavankumar11)
|
||||
|
||||
## Introduction
|
||||
|
||||
This guide will walk you through the detailed steps of installing and setting up the [Qdrant Sink Connector](https://github.com/qdrant/qdrant-kafka), building the necessary infrastructure, and creating a practical playground application. By the end of this article, you will have a deep understanding of how to leverage this powerful integration to streamline your data workflows, ultimately enhancing the performance and capabilities of your data-driven real-time semantic search and RAG applications.
|
||||
|
||||
In this example, original data will be sourced from Azure Blob Storage and MongoDB.
|
||||
|
||||

|
||||
|
||||
Figure 1: [Real time Change Data Capture (CDC)](https://www.confluent.io/learn/change-data-capture/) with Kafka and Qdrant.
|
||||
|
||||
## The Architecture:
|
||||
|
||||
## Source Systems
|
||||
|
||||
The architecture begins with the **source systems**, represented by MongoDB and Azure Blob Storage. These systems are vital for storing and managing raw data. MongoDB, a popular NoSQL database, is known for its flexibility in handling various data formats and its capability to scale horizontally. It is widely used for applications that require high performance and scalability. Azure Blob Storage, on the other hand, is Microsoft’s object storage solution for the cloud. It is designed for storing massive amounts of unstructured data, such as text or binary data. The data from these sources is extracted using **source connectors**, which are responsible for capturing changes in real-time and streaming them into Kafka.
|
||||
|
||||
## Kafka
|
||||
|
||||
At the heart of this architecture lies **Kafka**, a distributed event streaming platform capable of handling trillions of events a day. Kafka acts as a central hub where data from various sources can be ingested, processed, and distributed to various downstream systems. Its fault-tolerant and scalable design ensures that data can be reliably transmitted and processed in real-time. Kafka’s capability to handle high-throughput, low-latency data streams makes it an ideal choice for real-time data processing and analytics. The use of **Confluent** enhances Kafka’s functionalities, providing additional tools and services for managing Kafka clusters and stream processing.
|
||||
|
||||
## Qdrant
|
||||
|
||||
The processed data is then routed to **Qdrant**, a highly scalable vector search engine designed for similarity searches. Qdrant excels at managing and searching through high-dimensional vector data, which is essential for applications involving machine learning and AI, such as recommendation systems, image recognition, and natural language processing. The **Qdrant Sink Connector** for Kafka plays a pivotal role here, enabling seamless integration between Kafka and Qdrant. This connector allows for the real-time ingestion of vector data into Qdrant, ensuring that the data is always up-to-date and ready for high-performance similarity searches.
|
||||
|
||||
## Integration and Pipeline Importance
|
||||
|
||||
The integration of these components forms a powerful and efficient data streaming pipeline. The **Qdrant Sink Connector** ensures that the data flowing through Kafka is continuously ingested into Qdrant without any manual intervention. This real-time integration is crucial for applications that rely on the most current data for decision-making and analysis. By combining the strengths of MongoDB and Azure Blob Storage for data storage, Kafka for data streaming, and Qdrant for vector search, this pipeline provides a robust solution for managing and processing large volumes of data in real-time. The architecture’s scalability, fault-tolerance, and real-time processing capabilities are key to its effectiveness, making it a versatile solution for modern data-driven applications.
|
||||
|
||||
## Installation of Confluent Kafka Platform
|
||||
|
||||
To install the Confluent Kafka Platform (self-managed locally), follow these 3 simple steps:
|
||||
|
||||
**Download and Extract the Distribution Files:**
|
||||
|
||||
- Visit [Confluent Installation Page](https://www.confluent.io/installation/).
|
||||
- Download the distribution files (tar, zip, etc.).
|
||||
- Extract the downloaded file using:
|
||||
|
||||
```bash
|
||||
tar -xvf confluent-<version>.tar.gz
|
||||
```
|
||||
or
|
||||
```bash
|
||||
unzip confluent-<version>.zip
|
||||
```
|
||||
|
||||
**Configure Environment Variables:**
|
||||
|
||||
```bash
|
||||
# Set CONFLUENT_HOME to the installation directory:
|
||||
export CONFLUENT_HOME=/path/to/confluent-<version>
|
||||
|
||||
# Add Confluent binaries to your PATH
|
||||
export PATH=$CONFLUENT_HOME/bin:$PATH
|
||||
```
|
||||
|
||||
**Run Confluent Platform Locally:**
|
||||
|
||||
```bash
|
||||
# Start the Confluent Platform services:
|
||||
confluent local start
|
||||
# Stop the Confluent Platform services:
|
||||
confluent local stop
|
||||
```
|
||||
|
||||
## Installation of Qdrant:
|
||||
|
||||
To install and run Qdrant (self-managed locally), you can use Docker, which simplifies the process. First, ensure you have Docker installed on your system. Then, you can pull the Qdrant image from Docker Hub and run it with the following commands:
|
||||
|
||||
```bash
|
||||
docker pull qdrant/qdrant
|
||||
docker run -p 6334:6334 -p 6333:6333 qdrant/qdrant
|
||||
```
|
||||
|
||||
This will download the Qdrant image and start a Qdrant instance accessible at `http://localhost:6333`. For more detailed instructions and alternative installation methods, refer to the [Qdrant installation documentation](https://qdrant.tech/documentation/quick-start/).
|
||||
|
||||
## Installation of Qdrant-Kafka Sink Connector:
|
||||
|
||||
To install the Qdrant Kafka connector using [Confluent Hub](https://www.confluent.io/hub/), you can utilize the straightforward `confluent-hub install` command. This command simplifies the process by eliminating the need for manual configuration file manipulations. To install the Qdrant Kafka connector version 1.1.0, execute the following command in your terminal:
|
||||
|
||||
```bash
|
||||
confluent-hub install qdrant/qdrant-kafka:1.1.0
|
||||
```
|
||||
|
||||
This command downloads and installs the specified connector directly from Confluent Hub into your Confluent Platform or Kafka Connect environment. The installation process ensures that all necessary dependencies are handled automatically, allowing for a seamless integration of the Qdrant Kafka connector with your existing setup. Once installed, the connector can be configured and managed using the Confluent Control Center or the Kafka Connect REST API, enabling efficient data streaming between Kafka and Qdrant without the need for intricate manual setup.
|
||||
|
||||

|
||||
|
||||
*Figure 2: Local Confluent platform showing the Source and Sink connectors after installation.*
|
||||
|
||||
Ensure the configuration of the connector once it's installed as below. keep in mind that your `key.converter` and `value.converter` are very important for kafka to safely deliver the messages from topic to qdrant.
|
||||
|
||||
```bash
|
||||
{
|
||||
"name": "QdrantSinkConnectorConnector_0",
|
||||
"config": {
|
||||
"value.converter.schemas.enable": "false",
|
||||
"name": "QdrantSinkConnectorConnector_0",
|
||||
"connector.class": "io.qdrant.kafka.QdrantSinkConnector",
|
||||
"key.converter": "org.apache.kafka.connect.storage.StringConverter",
|
||||
"value.converter": "org.apache.kafka.connect.json.JsonConverter",
|
||||
"topics": "topic_62,qdrant_kafka.docs",
|
||||
"errors.deadletterqueue.topic.name": "dead_queue",
|
||||
"errors.deadletterqueue.topic.replication.factor": "1",
|
||||
"qdrant.grpc.url": "http://localhost:6334",
|
||||
"qdrant.api.key": "************"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Installation of MongoDB
|
||||
|
||||
For the Kafka to connect MongoDB as source, your MongoDB instance should be running in a `replicaSet` mode. below is the `docker compose` file which will spin a single node `replicaSet` instance of MongoDB.
|
||||
|
||||
```bash
|
||||
version: "3.8"
|
||||
|
||||
services:
|
||||
mongo1:
|
||||
image: mongo:7.0
|
||||
command: ["--replSet", "rs0", "--bind_ip_all", "--port", "27017"]
|
||||
ports:
|
||||
- 27017:27017
|
||||
healthcheck:
|
||||
test: echo "try { rs.status() } catch (err) { rs.initiate({_id:'rs0',members:[{_id:0,host:'host.docker.internal:27017'}]}) }" | mongosh --port 27017 --quiet
|
||||
interval: 5s
|
||||
timeout: 30s
|
||||
start_period: 0s
|
||||
start_interval: 1s
|
||||
retries: 30
|
||||
volumes:
|
||||
- "mongo1_data:/data/db"
|
||||
- "mongo1_config:/data/configdb"
|
||||
|
||||
volumes:
|
||||
mongo1_data:
|
||||
mongo1_config:
|
||||
```
|
||||
|
||||
Similarly, install and configure source connector as below.
|
||||
|
||||
```bash
|
||||
confluent-hub install mongodb/kafka-connect-mongodb:latest
|
||||
```
|
||||
|
||||
After installing the `MongoDB` connector, connector configuration should look like this:
|
||||
|
||||
```bash
|
||||
{
|
||||
"name": "MongoSourceConnectorConnector_0",
|
||||
"config": {
|
||||
"connector.class": "com.mongodb.kafka.connect.MongoSourceConnector",
|
||||
"key.converter": "org.apache.kafka.connect.storage.StringConverter",
|
||||
"value.converter": "org.apache.kafka.connect.storage.StringConverter",
|
||||
"connection.uri": "mongodb://127.0.0.1:27017/?replicaSet=rs0&directConnection=true",
|
||||
"database": "qdrant_kafka",
|
||||
"collection": "docs",
|
||||
"publish.full.document.only": "true",
|
||||
"topic.namespace.map": "{\"*\":\"qdrant_kafka.docs\"}",
|
||||
"copy.existing": "true"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Playground Application
|
||||
|
||||
As the infrastructure set is completely done, now it's time for us to create a simple application and check our setup. the objective of our application is the data is inserted to Mongodb and eventually it will get ingested into Qdrant also using [Change Data Capture (CDC)](https://www.confluent.io/learn/change-data-capture/).
|
||||
|
||||
`requirements.txt`
|
||||
|
||||
```bash
|
||||
fastembed==0.3.1
|
||||
pymongo==4.8.0
|
||||
qdrant_client==1.10.1
|
||||
```
|
||||
|
||||
`project_root_folder/main.py`
|
||||
|
||||
This is just sample code. Nevertheless it can be extended to millions of operations based on your use case.
|
||||
|
||||
```python
|
||||
from pymongo import MongoClient
|
||||
from utils.app_utils import create_qdrant_collection
|
||||
from fastembed import TextEmbedding
|
||||
|
||||
collection_name: str = 'test'
|
||||
embed_model_name: str = 'snowflake/snowflake-arctic-embed-s'
|
||||
```
|
||||
```python
|
||||
# Step 0: create qdrant_collection
|
||||
create_qdrant_collection(collection_name=collection_name, embed_model=embed_model_name)
|
||||
|
||||
# Step 1: Connect to MongoDB
|
||||
client = MongoClient('mongodb://127.0.0.1:27017/?replicaSet=rs0&directConnection=true')
|
||||
|
||||
# Step 2: Select Database
|
||||
db = client['qdrant_kafka']
|
||||
|
||||
# Step 3: Select Collection
|
||||
collection = db['docs']
|
||||
|
||||
# Step 4: Create a Document to Insert
|
||||
|
||||
description = "qdrant is a high available vector search engine"
|
||||
embedding_model = TextEmbedding(model_name=embed_model_name)
|
||||
vector = next(embedding_model.embed(documents=description)).tolist()
|
||||
document = {
|
||||
"collection_name": collection_name,
|
||||
"id": 1,
|
||||
"vector": vector,
|
||||
"payload": {
|
||||
"name": "qdrant",
|
||||
"description": description,
|
||||
"url": "https://qdrant.tech/documentation"
|
||||
}
|
||||
}
|
||||
|
||||
# Step 5: Insert the Document into the Collection
|
||||
result = collection.insert_one(document)
|
||||
|
||||
# Step 6: Print the Inserted Document's ID
|
||||
print("Inserted document ID:", result.inserted_id)
|
||||
```
|
||||
|
||||
`project_root_folder/utils/app_utils.py`
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(url="http://localhost:6333", api_key="<YOUR_KEY>")
|
||||
dimension_dict = {"snowflake/snowflake-arctic-embed-s": 384}
|
||||
|
||||
def create_qdrant_collection(collection_name: str, embed_model: str):
|
||||
|
||||
if not client.collection_exists(collection_name=collection_name):
|
||||
client.create_collection(
|
||||
collection_name=collection_name,
|
||||
vectors_config=models.VectorParams(size=dimension_dict.get(embed_model), distance=models.Distance.COSINE)
|
||||
)
|
||||
```
|
||||
|
||||
Before we run the application, below is the state of MongoDB and Qdrant databases.
|
||||
|
||||

|
||||
|
||||
Figure 3: Initial state: no collection named `test` & `no data` in the `docs` collection of MongodDB.
|
||||
|
||||
Once you run the code the data goes into Mongodb and the CDC gets triggered and eventually Qdrant will receive this data.
|
||||
|
||||

|
||||
|
||||
Figure 4: The test Qdrant collection is created automatically.
|
||||
|
||||

|
||||
|
||||
Figure 5: Data is inserted into both MongoDB and Qdrant.
|
||||
|
||||
## Conclusion:
|
||||
|
||||
In conclusion, the integration of **Kafka** with **Qdrant** using the **Qdrant Sink Connector** provides a seamless and efficient solution for real-time data streaming and processing. This setup not only enhances the capabilities of your data pipeline but also ensures that high-dimensional vector data is continuously indexed and readily available for similarity searches. By following the installation and setup guide, you can easily establish a robust data flow from your **source systems** like **MongoDB** and **Azure Blob Storage**, through **Kafka**, and into **Qdrant**. This architecture empowers modern applications to leverage real-time data insights and advanced search capabilities, paving the way for innovative data-driven solutions.
|
||||
@@ -6,8 +6,7 @@ weight: 17
|
||||
author: Andre Bogus
|
||||
author_link: https://llogiq.github.io
|
||||
date: 2023-06-01T11:24:20+01:00
|
||||
aliases:
|
||||
- /documentation/tutorials/mighty.md/
|
||||
draft: true
|
||||
keywords:
|
||||
- vector search
|
||||
- embeddings
|
||||
|
||||
@@ -12,6 +12,7 @@ weight: 33
|
||||
| [Bubble](./bubble) | Development platform for application development with a no-code interface |
|
||||
| [Canopy](./canopy/) | Framework from Pinecone for building RAG applications using LLMs and knowledge bases. |
|
||||
| [Cheshire Cat](./cheshire-cat/) | Framework to create personalized AI assistants using custom data. |
|
||||
| [Confluent](./confluent/) | Fully-managed data streaming platform with a cloud-native Apache Kafka engine. |
|
||||
| [DLT](./dlt/) | Python library to simplify data loading processes between several sources and destinations. |
|
||||
| [DocArray](./docarray/) | Python library for managing data in multi-modal AI applications. |
|
||||
| [DocsGPT](./docsgpt/) | Tool for ingesting documentation sources and enabling conversations and queries. |
|
||||
|
||||
@@ -0,0 +1,283 @@
|
||||
---
|
||||
title: Confluent
|
||||
weight: 3700
|
||||
---
|
||||
|
||||

|
||||
|
||||
[Confluent Cloud](https://www.confluent.io/confluent-cloud/?utm_campaign=tm.pmm_cd.cwc_partner_Qdrant_generic&utm_source=Qdrant&utm_medium=partnerref) is a fully-managed data streaming platform, available on AWS, GCP, and Azure, with a cloud-native Apache Kafka engine for elastic scaling, enterprise-grade security, stream processing, and governance.
|
||||
|
||||
With our [Qdrant-Kafka Sink Connector](https://github.com/qdrant/qdrant-kafka), Qdrant is part of the [Connect with Confluent](https://www.confluent.io/partners/connect/) technology partner program. It brings fully managed data streams directly to organizations through the Confluent Cloud platform. Making it easier for organizations to stream any data to Qdrant with a fully managed Apache Kafka service.
|
||||
|
||||
## Usage
|
||||
|
||||
### Pre-requisites
|
||||
|
||||
- A Confluent Cloud account. You can begin with a [free trial](https://www.confluent.io/confluent-cloud/tryfree/?utm_campaign=tm.pmm_cd.cwc_partner_qdrant_tryfree&utm_source=qdrant&utm_medium=partnerref) with credits for the first 30 days.
|
||||
- Qdrant instance to connect to. You can get a free cloud instance at [cloud.qdrant.io](https://cloud.qdrant.io/).
|
||||
|
||||
### Installation
|
||||
|
||||
1) Download the latest connector zip file from [Confluent Hub](https://www.confluent.io/hub/qdrant/qdrant-kafka).
|
||||
|
||||
2) Configure an environment and cluster on Confluent and create a topic to produce messages for.
|
||||
|
||||
3) Navigate to the `Connectors` section of the Confluent cluster and click `Add Plugin`. Upload the zip file with the following info.
|
||||
|
||||

|
||||
|
||||
4) Once installed, navigate to the connector and set the following configuration values.
|
||||
|
||||

|
||||
|
||||
Replace the placeholder values with your credentials.
|
||||
|
||||
5) Add the Qdrant instance host to the allowed networking endpoints.
|
||||
|
||||

|
||||
|
||||
7) Start the connector.
|
||||
|
||||
## Producing Messages
|
||||
|
||||
You can now produce messages for the configured topic, and they'll be written into the configured Qdrant instance.
|
||||
|
||||

|
||||
|
||||
## Message Formats
|
||||
|
||||
The connector supports messages in the following formats.
|
||||
|
||||
_Click each to expand._
|
||||
|
||||
<details>
|
||||
<summary><b>Unnamed/Default vector</b></summary>
|
||||
|
||||
Reference: [Creating a collection with a default vector](https://qdrant.tech/documentation/concepts/collections/#create-a-collection).
|
||||
|
||||
```json
|
||||
{
|
||||
"collection_name": "{collection_name}",
|
||||
"id": 1,
|
||||
"vector": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8
|
||||
],
|
||||
"payload": {
|
||||
"name": "kafka",
|
||||
"description": "Kafka is a distributed streaming platform",
|
||||
"url": "https://kafka.apache.org/"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Named multiple vectors</b></summary>
|
||||
|
||||
Reference: [Creating a collection with multiple vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-multiple-vectors).
|
||||
|
||||
```json
|
||||
{
|
||||
"collection_name": "{collection_name}",
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"some-dense": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8
|
||||
],
|
||||
"some-other-dense": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8
|
||||
]
|
||||
},
|
||||
"payload": {
|
||||
"name": "kafka",
|
||||
"description": "Kafka is a distributed streaming platform",
|
||||
"url": "https://kafka.apache.org/"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Sparse vectors</b></summary>
|
||||
|
||||
Reference: [Creating a collection with sparse vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-sparse-vectors).
|
||||
|
||||
```json
|
||||
{
|
||||
"collection_name": "{collection_name}",
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"some-sparse": {
|
||||
"indices": [
|
||||
0,
|
||||
1,
|
||||
2,
|
||||
3,
|
||||
4,
|
||||
5,
|
||||
6,
|
||||
7,
|
||||
8,
|
||||
9
|
||||
],
|
||||
"values": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8,
|
||||
0.9,
|
||||
1.0
|
||||
]
|
||||
}
|
||||
},
|
||||
"payload": {
|
||||
"name": "kafka",
|
||||
"description": "Kafka is a distributed streaming platform",
|
||||
"url": "https://kafka.apache.org/"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Multi-vectors</b></summary>
|
||||
|
||||
Reference:
|
||||
|
||||
- [Multi-vectors](https://qdrant.tech/documentation/concepts/vectors/#multivectors)
|
||||
|
||||
```json
|
||||
{
|
||||
"collection_name": "{collection_name}",
|
||||
"id": 1,
|
||||
"vector": {
|
||||
"some-multi": [
|
||||
[
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8,
|
||||
0.9,
|
||||
1.0
|
||||
],
|
||||
[
|
||||
1.0,
|
||||
0.9,
|
||||
0.8,
|
||||
0.5,
|
||||
0.4,
|
||||
0.8,
|
||||
0.6,
|
||||
0.4,
|
||||
0.2,
|
||||
0.1
|
||||
]
|
||||
]
|
||||
},
|
||||
"payload": {
|
||||
"name": "kafka",
|
||||
"description": "Kafka is a distributed streaming platform",
|
||||
"url": "https://kafka.apache.org/"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><b>Combination of named dense and sparse vectors</b></summary>
|
||||
|
||||
Reference:
|
||||
|
||||
- [Creating a collection with multiple vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-multiple-vectors).
|
||||
|
||||
- [Creating a collection with sparse vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-sparse-vectors).
|
||||
|
||||
```json
|
||||
{
|
||||
"collection_name": "{collection_name}",
|
||||
"id": "a10435b5-2a58-427a-a3a0-a5d845b147b7",
|
||||
"vector": {
|
||||
"some-other-dense": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8
|
||||
],
|
||||
"some-sparse": {
|
||||
"indices": [
|
||||
0,
|
||||
1,
|
||||
2,
|
||||
3,
|
||||
4,
|
||||
5,
|
||||
6,
|
||||
7,
|
||||
8,
|
||||
9
|
||||
],
|
||||
"values": [
|
||||
0.1,
|
||||
0.2,
|
||||
0.3,
|
||||
0.4,
|
||||
0.5,
|
||||
0.6,
|
||||
0.7,
|
||||
0.8,
|
||||
0.9,
|
||||
1.0
|
||||
]
|
||||
}
|
||||
},
|
||||
"payload": {
|
||||
"name": "kafka",
|
||||
"description": "Kafka is a distributed streaming platform",
|
||||
"url": "https://kafka.apache.org/"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Kafka Connect Docs](https://docs.confluent.io/platform/current/connect/index.html)
|
||||
- [Confluent Connectors Docs](https://docs.confluent.io/cloud/current/connectors/bring-your-connector/custom-connector-qs.html)
|
||||
@@ -168,7 +168,12 @@ storage:
|
||||
# Default is to allow 1 transfer.
|
||||
# If null - allow unlimited transfers.
|
||||
#outgoing_shard_transfers_limit: 1
|
||||
|
||||
|
||||
# Enable async scorer which uses io_uring when rescoring.
|
||||
# Only supported on Linux, must be enabled in your kernel.
|
||||
# See: <https://qdrant.tech/articles/io_uring/#and-what-about-qdrant>
|
||||
#async_scorer: false
|
||||
|
||||
optimizers:
|
||||
# The minimal fraction of deleted vectors in a segment, required to perform segment optimization
|
||||
deleted_threshold: 0.2
|
||||
|
||||
@@ -40,15 +40,15 @@ databases (as seen in the image above), data is organized in rows and columns (a
|
||||
called **Tables**), and queries are performed based on the values in those columns. However,
|
||||
in certain applications including image recognition, natural language processing, and recommendation
|
||||
systems, data is often represented as vectors in a high-dimensional space, and these vectors, plus
|
||||
an id and a payload, are the elements we store in something called a **Collection** a vector
|
||||
an id and a payload, are the elements we store in something called a **Collection** within a vector
|
||||
database like Qdrant.
|
||||
|
||||
A vector in this context is a mathematical representation of an object or data point, where each
|
||||
element of the vector corresponds to a specific feature or attribute of the object. For example,
|
||||
A vector in this context is a mathematical representation of an object or data point, where elements of
|
||||
the vector implicitly or explicitly correspond to specific features or attributes of the object. For example,
|
||||
in an image recognition system, a vector could represent an image, with each element of the vector
|
||||
representing a pixel value or a descriptor/characteristic of that pixel. In a music recommendation
|
||||
system, each vector would represent a song, and each element of the vector would represent a
|
||||
characteristic song such as tempo, genre, lyrics, and so on.
|
||||
system, each vector could represent a song, and elements of the vector would capture song characteristics
|
||||
such as tempo, genre, lyrics, and so on.
|
||||
|
||||
Vector databases are optimized for **storing** and **querying** these high-dimensional vectors
|
||||
efficiently, and they often using specialized data structures and indexing techniques such as
|
||||
@@ -60,24 +60,22 @@ Distance, Cosine Similarity, and Dot Product, and these three are fully supporte
|
||||
|
||||
Here's a quick overview of the three:
|
||||
- [**Cosine Similarity**](https://en.wikipedia.org/wiki/Cosine_similarity) - Cosine similarity
|
||||
is a way to measure how similar two things are. Think of it like a ruler that tells you how far
|
||||
apart two points are, but instead of measuring distance, it measures how similar two things
|
||||
are. It's often used with text to compare how similar two documents or sentences are to each
|
||||
other. The output of the cosine similarity ranges from -1 to 1, where -1 means the two things
|
||||
are completely dissimilar, and 1 means the two things are exactly the same. It's a straightforward
|
||||
and effective way to compare two things!
|
||||
- [**Dot Product**](https://en.wikipedia.org/wiki/Dot_product) - The dot product similarity
|
||||
metric is another way of measuring how similar two things are, like cosine similarity. It's
|
||||
often used in machine learning and data science when working with numbers. The dot product
|
||||
similarity is calculated by multiplying the values in two sets of numbers, and then adding
|
||||
up those products. The higher the sum, the more similar the two sets of numbers are. So, it's
|
||||
like a scale that tells you how closely two sets of numbers match each other.
|
||||
is a way to measure how similar two vectors are. To simplify, it reflects whether the vectors
|
||||
have the same direction (similar) or are poles apart. Cosine similarity is often used with text representations
|
||||
to compare how similar two documents or sentences are to each other. The output of cosine similarity ranges
|
||||
from -1 to 1, where -1 means the two vectors are completely dissimilar, and 1 indicates maximum similarity.
|
||||
- [**Dot Product**](https://en.wikipedia.org/wiki/Dot_product) - The dot product similarity metric is another way
|
||||
of measuring how similar two vectors are. Unlike cosine similarity, it also considers the length of the vectors.
|
||||
This might be important when, for example, vector representations of your documents are built
|
||||
based on the term (word) frequencies. The dot product similarity is calculated by multiplying the respective values
|
||||
in the two vectors and then summing those products. The higher the sum, the more similar the two vectors are.
|
||||
If you normalize the vectors (so the numbers in them sum up to 1), the dot product similarity will become
|
||||
the cosine similarity.
|
||||
- [**Euclidean Distance**](https://en.wikipedia.org/wiki/Euclidean_distance) - Euclidean
|
||||
distance is a way to measure the distance between two points in space, similar to how we
|
||||
measure the distance between two places on a map. It's calculated by finding the square root
|
||||
of the sum of the squared differences between the two points' coordinates. This distance metric
|
||||
is commonly used in machine learning to measure how similar or dissimilar two data points are
|
||||
or, in other words, to understand how far apart they are.
|
||||
is also commonly used in machine learning to measure how similar or dissimilar two vectors are.
|
||||
|
||||
Now that we know what vector databases are and how they are structurally different than other
|
||||
databases, let's go over why they are important.
|
||||
|
||||
@@ -1,10 +1,14 @@
|
||||
---
|
||||
title: Vector Search Basics
|
||||
title: Understanding Vector Search in Qdrant
|
||||
weight: 1
|
||||
social_preview_image: /docs/gettingstarted/vector-social.png
|
||||
---
|
||||
|
||||
# Vector Search Basics
|
||||
# How Does Vector Search Work in Qdrant?
|
||||
|
||||
<p align="center"><iframe width="560" height="315" src="https://www.youtube.com/embed/mXNrhyw4q84?si=wruP9wWSa8JW4t78" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></p>
|
||||
|
||||
|
||||
|
||||
If you are still trying to figure out how vector search works, please read ahead. This document describes how vector search is used, covers Qdrant's place in the larger ecosystem, and outlines how you can use Qdrant to augment your existing projects.
|
||||
|
||||
|
||||
@@ -1,6 +1,8 @@
|
||||
---
|
||||
title: Semantic Search 101
|
||||
weight: -100
|
||||
aliases:
|
||||
- /documentation/tutorials/mighty.md/
|
||||
---
|
||||
|
||||
# Semantic Search for Beginners
|
||||
|
||||
Reference in New Issue
Block a user