mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-09 21:08:31 +02:00
Address review feedback
This commit is contained in:
@@ -1,14 +1,14 @@
|
||||
---
|
||||
title: "What is a Vector Database?"
|
||||
draft: false
|
||||
short_description: What is a Vector Database? Use Cases & Examples | Qdrant
|
||||
description: Discover what a vector database is, its core functionalities, and real-world applications.
|
||||
short_description: What Is a Vector Database? Concepts, Architecture & Use Cases | Qdrant
|
||||
description: Discover how vector databases power semantic search by understanding the meaning behind your data. This guide breaks down how they work and why they've become the backbone of modern AI search.
|
||||
preview_dir: /articles_data/what-is-a-vector-database/preview
|
||||
weight: 30
|
||||
social_preview_image: /articles_data/what-is-a-vector-database/preview/social_preview.png
|
||||
date: 2024-10-09T09:29:33-03:00
|
||||
aliases: [ /blog/what-is-a-vector-database/ ]
|
||||
author: Sabrina Aquino
|
||||
author: Sabrina Aquino & Chadha Sridi
|
||||
featured: true
|
||||
tags:
|
||||
- vector-search
|
||||
@@ -19,8 +19,6 @@ category: core-concepts
|
||||
|
||||
## An Introduction to Vector Databases
|
||||
|
||||

|
||||
|
||||
Most of the millions of terabytes of data we generate each day is **unstructured**. Think of the meal photos you snap, the PDFs shared at work, or the podcasts you save but may never listen to. None of it fits neatly into rows and columns.
|
||||
|
||||
Unstructured data lacks a strict format or schema, making it challenging for conventional databases to manage. Yet, this unstructured data holds immense potential for **AI**, **machine learning**, and **modern search engines**.
|
||||
@@ -29,7 +27,7 @@ Unstructured data lacks a strict format or schema, making it challenging for con
|
||||
|
||||
Traditional [OLTP](https://www.ibm.com/topics/oltp) and [OLAP](https://www.ibm.com/topics/olap) databases have been the backbone of data storage for decades. They are great at managing structured data with well-defined schemas, like `name`, `address`, `phone number`, and `purchase history`.
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/oltp-and-olap.png" alt="Structure of OLTP and OLAP databases" width="500">
|
||||
<img src="/articles_data/what-is-a-vector-database/oltp-vs-olap.png" alt="Structure of OLTP and OLAP databases" width="600">
|
||||
|
||||
But when data can't be easily categorized, like the content inside a PDF file, things start to get complicated.
|
||||
|
||||
@@ -37,7 +35,7 @@ You can always store the PDF file as raw data, perhaps with some metadata attach
|
||||
|
||||
Also, this applies to more than just PDF documents. Think about the vast amounts of text, audio, and image data you generate every day. If a database can’t grasp the **meaning** of this data, how can you search for or find relationships within the data?
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/vector-db-structure.png" alt="Structure of a Vector Database" width="400">
|
||||
<img src="/articles_data/what-is-a-vector-database/vector-database-structure.png" alt="Structure of a Vector Database" width="600">
|
||||
|
||||
This is where vector databases come in. They index and query unstructured data as **vectors** that capture patterns and relationships, enabling applications to search and retrieve information based on meaning rather than exact matches.
|
||||
|
||||
@@ -58,8 +56,6 @@ Traditional and vector databases aren't rivals; they solve different problems. T
|
||||
|
||||
## What Is a Vector?
|
||||
|
||||

|
||||
|
||||
When a machine needs to process unstructured data - an image, a piece of text, or an audio file, it first has to translate that data into a format it can work with: **vectors**.
|
||||
|
||||
> A **vector** is a numerical representation of data in a **multi-dimensional space** that captures the **context** and **semantics** of data.
|
||||
@@ -68,11 +64,11 @@ These numbers are generated by **embedding models**, such as deep learning algor
|
||||
|
||||
To represent textual data, for example, an embedding will encapsulate the nuances of language, such as semantics and context within its dimensions.
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/embedding-model.png" alt="Creation of a vector based on a sentence with an embedding model" width="500">
|
||||
<img src="/articles_data/what-is-a-vector-database/embedding-model-diagram.png" alt="Creation of a vector based on a sentence with an embedding model" width="500">
|
||||
|
||||
For that reason, when comparing two similar sentences, their embeddings will turn out to be very similar, because they have similar **linguistic elements**.
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/two-similar-vectors.png" alt="Comparison of the embeddings of 2 similar sentences" width="500">
|
||||
<img src="/articles_data/what-is-a-vector-database/embedding-similarity.png" alt="Comparison of the embeddings of 2 similar sentences" width="500">
|
||||
|
||||
That’s the beauty of embeddings. The complexity of the data is distilled into something that can be compared across a multi-dimensional space.
|
||||
|
||||
@@ -94,7 +90,7 @@ Before diving into the individual concepts, it helps to see the full picture. Re
|
||||
|
||||
5. **Power your application:** The retrieved results can then be used in applications such as semantic search, recommendation systems, anomaly detection, and Retrieval-Augmented Generation (RAG).
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/search-workflow.png" alt=" Vector search workflow diagram" width="1000">
|
||||
<img src="/articles_data/what-is-a-vector-database/workflow.png" alt=" Vector search workflow diagram" width="800">
|
||||
|
||||
## Common Applications of Vector Search
|
||||
|
||||
@@ -115,13 +111,13 @@ Vector search enables a wide range of applications by allowing systems to retrie
|
||||
> A quick note on naming. You'll almost always hear these systems called vector databases, but the term is a little misleading. Traditional databases like Postgres or MySQL are built on ACID principles: transactions, strong consistency, and atomicity. Most "vector databases" aren't databases in that sense. They're really **vector search engines**, designed for horizontal scalability, low-latency queries, and high availability. Those priorities lead to different architectural decisions that are not reproducible in general-purpose databases. The label vector database stuck for marketing reasons, so we use it throughout this article, but vector search engine is the more accurate description.
|
||||
|
||||
A vector database is made of multiple different entities and relations. Let's understand a bit of what's happening here:
|
||||
<img src="/articles_data/what-is-a-vector-database/vector-database-architecture.png" alt="Architecture Diagram of a Vector Database" width="900">
|
||||
<img src="/articles_data/what-is-a-vector-database/architecture.png" alt="Architecture Diagram of a Vector Database" width="900">
|
||||
|
||||
### Points
|
||||
[Points](https://qdrant.tech/documentation/manage-data/points/) are the core units of data stored and retrieved. They are the central entity that Qdrant operates with.
|
||||
|
||||
There are three key elements that form a point: the **ID**, one or more **vector(s)**, and the **payload**.
|
||||
<img src="/articles_data/what-is-a-vector-database/point.png" alt="Representation of a Point in Qdrant" width="700">
|
||||
<img src="/articles_data/what-is-a-vector-database/point-structure.png" alt="Representation of a Point in Qdrant" width="700">
|
||||
|
||||
Each one of these parts plays an important role in how a point is stored, retrieved, and interpreted. Let's see how.
|
||||
|
||||
@@ -171,7 +167,7 @@ These metrics define how similarity between vectors is calculated. The choice of
|
||||
|
||||
- **Cosine Similarity:** This one is about the angle, not the length. It measures how two vectors point in the same direction, so it works well for text or documents when you care more about meaning than magnitude. For example, if two things are *similar*, *opposite*, or *unrelated*:
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/cosine-similarity.png" alt="Cosine Similarity Example" width="700">
|
||||
<img src="/articles_data/what-is-a-vector-database/cosine-sim.png" alt="Cosine Similarity Example" width="700">
|
||||
|
||||
- **Dot Product:** This looks at how much two vectors align. It’s popular in recommendation systems where you're interested in how much two things “agree” with each other.
|
||||
|
||||
@@ -202,8 +198,6 @@ Qdrant offers a range of SDKs. You can use the programming language you're most
|
||||
|
||||
## The Core Functionalities of Vector Databases
|
||||
|
||||

|
||||
|
||||
When you think of a traditional database, the operations are familiar: you **create**, **read**, **update**, and **delete** records. These are the fundamentals. And guess what? In many ways, vector databases work the same way, but the operations are translated for the complexity of vectors.
|
||||
|
||||
### 1. Indexing
|
||||
@@ -215,7 +209,7 @@ Indexing your vectors is like creating an entry in a traditional database. But f
|
||||
|
||||
It builds a multi-layered graph, where each vector is a node and connections represent similarity. The higher layers connect broadly similar vectors, while lower layers link vectors that are closely related, making searches progressively more refined as they go deeper.
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/hnsw.png" alt="Indexing Data with the HNSW algorithm" width="500">
|
||||
<img src="/articles_data/what-is-a-vector-database/hnsw-indexing.png" alt="Indexing Data with the HNSW algorithm" width="500">
|
||||
|
||||
When you run a search, HNSW starts at the top, quickly narrowing down the search by hopping between layers. It focuses only on relevant vectors as it goes deeper, refining the search with each step.
|
||||
|
||||
@@ -233,7 +227,7 @@ You need to build the payload index for **each field** you'd like to search. The
|
||||
|
||||
Similarity search allows you to search by **meaning**. This way you can do searches such as similar songs that evoke the same mood, finding images that match your artistic vision, or even exploring emotional patterns in text.
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/similarity.png" alt="Similar words grouped together" width="800">
|
||||
<img src="/articles_data/what-is-a-vector-database/word-clusters.png" alt="Similar words grouped together" width="800">
|
||||
|
||||
The way it works is, when the user queries the database, this query is also converted into a vector. The algorithm quickly identifies the area of the graph likely to contain vectors closest to the **query vector**.
|
||||
|
||||
@@ -289,7 +283,6 @@ You can use deletion to remove outdated data, clean up duplicates, and manage th
|
||||
|
||||
|
||||
## Hybrid Search
|
||||

|
||||
|
||||
Sometimes context alone isn’t enough. Sometimes you need precision, too. Dense vectors (embeddings we've talked about so far) are fantastic when you need to retrieve results based on the context or meaning behind the data. Sparse vectors are useful when you also need **keyword or specific attribute matching**.
|
||||
|
||||
@@ -303,7 +296,7 @@ Dense and sparse vectors represent data in fundamentally different ways, and eac
|
||||
|
||||
Dense vectors are, quite literally, dense with information. Every element in the vector contributes to the **semantic meaning**, **relationships** and **nuances** of the data. A dense vector representation of this sentence might look like this:
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/dense-1.png" alt="Representation of a Dense Vector" width="500">
|
||||
<img src="/articles_data/what-is-a-vector-database/dense-embedding.png" alt="Representation of a Dense Vector" width="600">
|
||||
|
||||
Each number holds weight. Together, they convey the overall meaning of the sentence, and are better for identifying contextually similar items, even if the words don’t match exactly.
|
||||
|
||||
@@ -313,7 +306,7 @@ Sparse vectors operate differently. They focus only on the essentials. In most s
|
||||
|
||||
In the image, you can see a sentence, *“I love Vector Similarity,”* broken down into tokens like *“i,” “love,” “vector”* through tokenization. Each token is assigned a unique `ID` from a large vocabulary. For example, *“i”* becomes `193`, and *“vector”* becomes `15012`.
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/sparse.png" alt="How Sparse Vectors are Created" width="700">
|
||||
<img src="/articles_data/what-is-a-vector-database/sparse-vector.png" alt="How Sparse Vectors are Created" width="800">
|
||||
|
||||
Sparse vectors are used for **exact matching** and specific token-based identification. The values on the right, such as `193: 0.04` and `9182: 0.12`, are the scores or weights for each token, showing how relevant or important each token is in the context. The final result is a sparse vector:
|
||||
|
||||
@@ -337,12 +330,12 @@ Sparse vectors are ideal for tasks like **keyword search** where you need to che
|
||||
|
||||
Qdrant uses **normalization** and **fusion** techniques to blend results from multiple search methods, such as dense and sparse. One common approach is **Reciprocal Rank Fusion (RRF)**, where results from different methods are merged, giving higher importance to items ranked highly by both methods. This ensures that the best candidates, whether identified through dense or sparse vectors, appear at the top of the results.
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/hybrid-search-2.png" alt="Hybrid Search API - How it works" width="500">
|
||||
<img src="/articles_data/what-is-a-vector-database/fusion.png" alt="Hybrid Search API - How it works" width="500">
|
||||
|
||||
### How to Use Hybrid Search in Qdrant
|
||||
Qdrant makes it easy to implement hybrid search through its Query API. Here’s how you can make it happen in your own project:
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/hybrid-query2.png" alt="Hybrid Query Example" width="700">
|
||||
<img src="/articles_data/what-is-a-vector-database/hybrid-pipeline.png" alt="Hybrid Query Example" width="800">
|
||||
|
||||
**Example Hybrid Query:**
|
||||
|
||||
@@ -367,43 +360,39 @@ client.query_points(
|
||||
```
|
||||
This is just a simple example and there's so much more you can do with it. See our complete [article on Hybrid Search](https://qdrant.tech/articles/hybrid-search/) guide to see what's happening behind the scenes and all the possibilities when building a hybrid search system.
|
||||
|
||||
## Quantization: Get 40x Faster Results
|
||||
## Quantization
|
||||
|
||||

|
||||
As your vector dataset grows larger, so do the memory and compute demands of searching through it. Quantization compresses vectors into a more compact form that is cheaper to store and faster to compare, trading a small amount of precision for large gains in efficiency.
|
||||
|
||||
As your vector dataset grows larger, so do the computational demands of searching through it.
|
||||
Qdrant 1.18 ships [**TurboQuant**](https://qdrant.tech/articles/turboquant-quantization/), a rotation-based vector quantization method from Google Research, with extensions that make it work on real production embeddings.
|
||||
|
||||
Quantized vectors are much smaller and easier to compare. With methods like [**Binary Quantization**](https://qdrant.tech/articles/binary-quantization/), you can see **search speeds improve by up to 40x while memory usage decreases by 32x**, improvements that can be decisive when dealing with large datasets or needing low-latency results.
|
||||
Instead of storing each dimension of your vectors at full 32-bit precision, quantization compresses it down to just a few bits. TurboQuant's 4-bit variant gives you 8x compression while matching Scalar Quantization's recall at half the memory.
|
||||
When you need to squeeze memory further, the 2-bit and 1-bit variants push the compression rate to 16x and 32x while still beating Binary Quantization at the same storage budget.
|
||||
|
||||
It works by converting high-dimensional vectors, which typically use `4 bytes` per dimension, into binary representations, using just `1 bit` per dimension. Values above zero become "1", and everything else becomes "0".
|
||||
<img src="/articles_data/what-is-a-vector-database/TQ.png" alt="TurboQuant vs Scalar Quantization vs full precision baseline" width="700">
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/binary-quantization-2.png" alt=" Binary Quantization example" width="600">
|
||||
|
||||
Quantization reduces data precision, and yes, this does lead to some loss of accuracy. However, for binary quantization, **OpenAI embeddings** achieve this performance improvement at a cost of only 5% drop of accuracy. If you apply techniques like **oversampling** and **rescoring**, this loss can be brought down even further.
|
||||
|
||||
However, binary quantization isn’t the only available option. Techniques like [**TurboQuant Quantization**](https://qdrant.tech/documentation/manage-data/quantization/#turboquant-quantization), [**Scalar Quantization**](https://qdrant.tech/documentation/manage-data/quantization/#scalar-quantization) and [**Product Quantization**](https://qdrant.tech/documentation/manage-data/quantization/#product-quantization) are also popular alternatives when optimizing vector compression.
|
||||
|
||||
You can set up your chosen quantization method using the `quantization_config` parameter when creating a new collection:
|
||||
You can set up quantization using the `quantization_config` parameter when creating a new collection:
|
||||
|
||||
```python
|
||||
client.create_collection(
|
||||
collection_name="{collection_name}",
|
||||
vectors_config=models.VectorParams(
|
||||
size=1536,
|
||||
size=1536,
|
||||
distance=models.Distance.COSINE
|
||||
),
|
||||
|
||||
# Choose your preferred quantization method
|
||||
quantization_config=models.BinaryQuantization(
|
||||
binary=models.BinaryQuantizationConfig(
|
||||
always_ram=True, # Store the quantized vectors in RAM for faster access
|
||||
# Enable TurboQuant (4-bit by default)
|
||||
quantization_config=models.TurboQuantization(
|
||||
turbo=models.TurboQuantQuantizationConfig(
|
||||
always_ram=True, # Keep quantized vectors in RAM for faster access
|
||||
),
|
||||
),
|
||||
)
|
||||
```
|
||||
You can store original vectors on disk within the `vectors_config` by setting `on_disk=True` to save RAM space, while keeping quantized vectors in RAM for faster access.
|
||||
|
||||
We recommend checking out our [Vector Quantization guide](https://qdrant.tech/articles/what-is-vector-quantization/) for a full breakdown of methods and tips on **optimizing performance** for your specific use case.
|
||||
You can store the original vectors on disk within `vectors_config` by setting `on_disk=True` to save RAM, while keeping the quantized vectors in RAM for faster access.
|
||||
|
||||
TurboQuant isn't the only option. [**Binary Quantization**](https://qdrant.tech/articles/binary-quantization/) is the most aggressive choice for speed, converting each dimension to a single bit so that search speeds can improve by up to 40x with memory reduced by 32x, at an accuracy cost that oversampling and rescoring can recover. [**Scalar Quantization**](https://qdrant.tech/documentation/manage-data/quantization/#scalar-quantization) and [**Product Quantization**](https://qdrant.tech/documentation/manage-data/quantization/#product-quantization) round out the alternatives. Check out our [Vector Quantization guide](https://qdrant.tech/articles/what-is-vector-quantization/) for a full breakdown and tips on **optimizing performance** for your use case.
|
||||
|
||||
## Distributed Deployment
|
||||
|
||||
@@ -415,7 +404,7 @@ In a distributed Qdrant cluster, data is split into smaller units called **shard
|
||||
|
||||
Each collection can be split into non-overlapping subsets, which are then managed by different nodes.
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/sharding-raft.png" alt=" Distributed vector database with sharding and Raft consensus" width="1000">
|
||||
<img src="/articles_data/what-is-a-vector-database/raft.png" alt=" Distributed vector database with sharding and Raft consensus" width="1000">
|
||||
|
||||
**Raft Consensus** ensures that all the nodes stay in sync and have a consistent view of the data. Each node knows where every shard is, and Raft ensures that all nodes are in sync. If one node fails, the others know where the missing data is located and can take over.
|
||||
|
||||
@@ -436,7 +425,7 @@ There are two main types of sharding:
|
||||
|
||||
Each shard is divided into **segments**. They are a smaller storage unit within a shard, storing a subset of vectors and their associated payloads (metadata). When a query is executed, it targets only the relevant segments, processing them in parallel.
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/segments.png" alt="Segments act as smaller storage units within a shard" width="700">
|
||||
<img src="/articles_data/what-is-a-vector-database/shard-segments.png" alt="Segments act as smaller storage units within a shard" width="700">
|
||||
|
||||
### Replication: High Availability and Data Integrity
|
||||
|
||||
@@ -444,7 +433,7 @@ You don’t want a single failure to take down your system, right? Replication k
|
||||
|
||||
In Qdrant, **Replica Sets** manage these copies of shards across different nodes. If one replica becomes unavailable, others are there to take over and keep the system running. Whether the data is local or remote is mainly influenced by how you've configured the cluster.
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/replication.png" alt=" Replica Set and Replication diagram" width="1000">
|
||||
<img src="/articles_data/what-is-a-vector-database/replication-diagram.png" alt=" Replica Set and Replication diagram" width="1000">
|
||||
|
||||
When a query is made, if the relevant data is stored locally, the local shard handles the operation. If the data is on a remote shard, it’s retrieved via gRPC.
|
||||
|
||||
@@ -465,13 +454,11 @@ For more details on features like **user-defined sharding, node failure recovery
|
||||
|
||||
## Multitenancy: Data Isolation for Multi-Tenant Architectures
|
||||
|
||||

|
||||
|
||||
Sharding efficiently distributes data across nodes, while replication guarantees redundancy and fault tolerance. But what happens when you’ve got multiple clients or user groups, and you need to keep their data isolated within the same infrastructure?
|
||||
|
||||
**Multitenancy** allows you to keep data for different tenants (users, clients, or organizations) isolated within a single cluster. Instead of creating separate collections for `Tenant 1` and `Tenant 2`, you store their data in the same collection but tag each vector with a `group_id` to identify which tenant it belongs to.
|
||||
|
||||
<img src="/articles_data/what-is-a-vector-database/multitenancy-1.png" alt="Multitenancy dividing data between 2 tenants" width="1000">
|
||||
<img src="/articles_data/what-is-a-vector-database/multitenancy.png" alt="Multitenancy dividing data between 2 tenants" width="1000">
|
||||
|
||||
In the backend, Qdrant can store `Tenant 1`’s data in Shard 1 located in Canada (perhaps for compliance reasons like GDPR), while `Tenant 2`’s data is stored in Shard 2 located in Germany. The data will be physically separated but still within the same infrastructure.
|
||||
|
||||
@@ -534,10 +521,6 @@ As we've seen in this article, a vector database is definitely not **just** a da
|
||||
|
||||
But there’s no better way to learn than by doing. Try building a [semantic search engine](https://qdrant.tech/documentation/tutorials/search-beginners/) or experiment deploying a [hybrid search service](https://qdrant.tech/documentation/tutorials/hybrid-search-fastembed/) from zero. You'll realize there are endless ways you can take advantage of vectors.
|
||||
|
||||
You can also watch our video tutorial and get started with Qdrant to generate semantic search results and recommendations from a sample dataset.
|
||||
|
||||
<iframe width="560" height="315" src="https://www.youtube.com/embed/LRcZ9pbGnno?si=sO5oX9mc-QDTBNrV" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
|
||||
|
||||
Phew! I hope you found some of the concepts here useful. If you have any questions feel free to send them in our [Discord Community](https://discord.com/invite/qdrant) where our team will be more than happy to help you out!
|
||||
|
||||
> Remember, don't get lost in vector space! 🚀
|
||||
|
||||
Reference in New Issue
Block a user