update doc

This commit is contained in:
Derrick Mwiti
2025-06-05 08:54:24 +03:00
parent c919950b2b
commit 583f9f52e8
@@ -10,11 +10,15 @@ Multivector Representations are one of the most powerful features of Qdrant. How
In this tutorial, you'll discover how to effectively use multivector representations in Qrant.
## What are Multivector Representations?
In most vector engines, each document results in one vector. However, this is not always effective when you have very long documents. A single document can have multiple vectors in multivector representations, leading to more precise matching between queries and parts of the document. This is particularly useful for Late Interaction models such as [ColBERT](https://qdrant.tech/documentation/fastembed/fastembed-colbert/), where each document is represented using multiple token-level vectors.
In most vector engines, each document is represented by a single vector — an approach that works well for short texts but often struggles with longer documents. To mitigate this, it's common practice to split documents into smaller chunks (e.g., paragraphs or sentences) and embed each chunk individually. While this chunking strategy improves retrieval granularity, it can miss broader context across chunks.
Multivector representations offer a more fine-grained alternative: instead of chunking, a single document is represented using multiple vectors, often at the token or phrase level. This enables more precise matching between specific query terms and relevant parts of the document. This is especially effective in Late Interaction models like [ColBERT](https://qdrant.tech/documentation/fastembed/fastembed-colbert/), which retain token-level embeddings and perform interaction during query time to boost retrieval precision.
As you will see later in the tutorial, Qdrant supports multivectors and late interaction models natively.
## Why Token-level Vectors are Useful
With token-level vectors, your system can find the exact part of a document matching your query and support Late Interaction Models with high accuracy.
With token-level vectors, models like ColBERT can match specific query tokens to the most relevant parts of a document, enabling high-accuracy retrieval through Late Interaction.
Each document is converted into multiple token-level vectors instead of a single vector in Late Interaction. The query is also tokenized and embedded into various vectors. Then, the query and document vectors are matched using a similarity function. In traditional retrieval, the query and document are converted into single embeddings, after which similarity is computed. This is an early interaction because the information is compressed before retrieval.
@@ -24,11 +28,16 @@ Rescoring is two-fold:
- Rerank them using a more accurate but slower model such as ColBERT.
## Why Indexing Every Vector by Default is a Problem
Multivector documents can have hundreds of vectors per document. Indexing each vector in Qdrant using HNSW leads to:
- Usage of a lot of RAM.
- Slow inserts.
In multivector representations (such as those used by Late Interaction models like ColBERT), a single logical document — say, a PDF or full article — can result in hundreds of token-level vectors. Indexing each of these vectors individually with HNSW in Qdrant can lead to:
Since multivector search is used for reranking only, you don't need to index the vectors.
- High RAM usage
- Slow insert times due to the complexity of maintaining the HNSW graph
However, because multivector search is typically used in the reranking stage (after a first-pass retrieval using dense vectors), there's often no need to index these token-level vectors with HNSW.
Instead, they can be stored as multi-vector fields (without HNSW indexing) and used at query-time for reranking, which reduces resource overhead and improves performance.
For more on this, check out Qdrant's detailed breakdown in our [Scaling PDF Retrieval with Qdrant tutorial](https://qdrant.tech/documentation/advanced-tutorials/pdf-retrieval-at-scale/#math-behind-the-scaling).
With Qdrant, you have full control of how indexing works. You can disable indexing by setting the HNSW `m` parameter to `0`:
```python
@@ -88,6 +97,7 @@ colbert_query_vector = list(colbert_model.embed([query_text]))[0]
### 3. Create a Qdrant collection
Then create a Qdrant collection with both vector types. Note that we leave indexing on for the `dense` vector but turn it off for the `colbert` vector that will be used for reranking.
```python
collection_name = "dense_multivector_demo"
client.create_collection(
collection_name=collection_name,
vectors_config={
@@ -122,16 +132,16 @@ points = [
payload={"text": documents[i]}
) for i in range(len(documents))
]
client.upsert(collection_name="hybrid_dense_multivector_demo", points=points)
client.upsert(collection_name="dense_multivector_demo", points=points)
```
### Query with Retrieval + Reranking in One Call
Now let’s run a hybrid search:
Now let’s run a search:
```python
results = client.query_points(
collection_name="hybrid_dense_multivector_demo",
collection_name="dense_multivector_demo",
prefetch=models.Prefetch(
query=dense_query_vector,
using="dense",
@@ -146,14 +156,14 @@ results = client.query_points(
```
- The dense vector retrieves the top 100 candidates quickly.
- The Colbert multivector reranks them using token-level `MaxSim`.
- The Colbert multivector reranks them using token-level `MaxSim` with fine-grained precision.
- Returns the top 10 results.
## Conclusion
Multivector search is one of the most powerful features of a vector database when used correctly. With this functionality in Qdrant, you can:
- Store token-level embeddings natively.
- Disable indexing to reduce overhead.
- Run fast and accurate hybrid search in one API call.
- Run fast and accurate search in one API call.
- Efficiently scale late interaction.
Combining FastEmbed and Qdrant leads to a production-ready pipeline for ColBERT-style reranking without wasting resources. You can do this locally or use Qdrant Cloud. Qdrant offers an easy-to-use API to get started with your search engine, so if you’re ready to dive in, sign up for free at [Qdrant Cloud](https://qdrant.tech/cloud/) and start building.