mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-03 09:58:30 +02:00
Remaining drafts
This commit is contained in:
@@ -26,4 +26,18 @@ This lesson provides a framework for making data-driven decisions about your mul
|
||||
|
||||
---
|
||||
|
||||
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-3/evaluating-pipelines.ipynb">
|
||||
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||
</a>
|
||||
|
||||
---
|
||||
|
||||
TODO: introduce the concepts like qrels
|
||||
|
||||
TODO: evaluate the search quality with ranx
|
||||
|
||||
## Summary
|
||||
|
||||
TODO: summarize the entire course
|
||||
|
||||
Congratulations! You now have the knowledge to build, optimize, and deploy production-ready multi-vector search systems with Qdrant.
|
||||
|
||||
@@ -10,8 +10,6 @@ weight: 3
|
||||
|
||||
While quantization reduces the size of each vector, pooling reduces the number of vectors per document. By intelligently combining token embeddings, you can achieve significant memory savings while preserving retrieval quality.
|
||||
|
||||
Pooling is particularly effective when combined with quantization for maximum memory efficiency.
|
||||
|
||||
---
|
||||
|
||||
<div class="video">
|
||||
@@ -26,4 +24,41 @@ Pooling is particularly effective when combined with quantization for maximum me
|
||||
|
||||
---
|
||||
|
||||
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-3/pooling-techniques.ipynb">
|
||||
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||
</a>
|
||||
|
||||
---
|
||||
|
||||
## Pooling in Embedding Models
|
||||
|
||||
Pooling isn't new to vector search - it's fundamental to how most embedding models work. When you encode text with models like Sentence Transformers, the model first generates embeddings for each token in your input. But to create a single vector representing the entire text, the model must **pool** these token embeddings together.
|
||||
|
||||
Common pooling strategies in dense embedding models include:
|
||||
|
||||
- **Mean pooling**: Average all token embeddings into a single vector
|
||||
- **CLS token pooling**: Use the special `[CLS]` token's embedding as the document representation
|
||||
- **Max pooling**: Take the maximum value for each dimension across all tokens
|
||||
- **Weighted pooling**: Assign different importance to different tokens (e.g., using attention weights)
|
||||
|
||||
These techniques compress variable-length sequences of token embeddings into fixed-size vectors, making them compatible with traditional vector search systems.
|
||||
|
||||
With multi-vector representations, we face a similar but more nuanced challenge. Instead of reducing tokens to a single vector upfront, we maintain multiple vectors per document to preserve richer semantic information. However, as you learned in the previous lessons, this creates memory and performance challenges. **Pooling techniques for multi-vector search** let you strategically reduce the number of vectors while retaining the benefits of late interaction.
|
||||
|
||||
## Pooling for Multi-Vector Representations
|
||||
|
||||
### Image-Specific Methods
|
||||
|
||||
TODO: describe row/column pooling we can implement because of the spatial relationships in data
|
||||
|
||||

|
||||
|
||||
### Generic Methods
|
||||
|
||||
TODO: hierarchical token pooling as a universal method, based on clustering
|
||||
|
||||
## What's Next
|
||||
|
||||
TODO: summarize the lesson
|
||||
|
||||
Now let's tackle the indexing challenge with MUVERA, enabling fast approximate search for multi-vector representations.
|
||||
|
||||
@@ -61,13 +61,6 @@ $$
|
||||
|
||||
**Multi-vector representations use ~170x more memory** than single-vector models for the same number of documents. This is where quantization becomes essential.
|
||||
|
||||
<!-- TODO: Add comparison diagram showing memory footprint
|
||||
- Side-by-side bar chart: Single-vector (3 KB) vs Multi-vector (512 KB) per document
|
||||
- Scale visualization showing 1M documents: 3 GB vs 512 GB
|
||||
- Color code: Green for single-vector, Orange/Red for multi-vector to emphasize the difference
|
||||
- Add annotation showing "170x memory increase"
|
||||
-->
|
||||
|
||||
## Quantization: Compressing Without Losing Quality
|
||||
|
||||
**Vector quantization** reduces memory by representing vectors with fewer bits while preserving the relative distances between them. Qdrant supports several quantization methods optimized for different scenarios.
|
||||
@@ -151,16 +144,6 @@ $$
|
||||
|
||||
For ColModernVBERT's **128 dimensions**, binary quantization presents unique challenges. With such low dimensionality, each bit of precision has a larger impact on the representation. The choice between scalar and binary quantization - and which binary variant to use - depends on your specific use case and quality requirements. We'll explore how to evaluate these trade-offs systematically in the final lesson of this module.
|
||||
|
||||
<!-- TODO: Add binary quantization comparison diagram
|
||||
- Three-column comparison showing 1-bit, 1.5-bit, and 2-bit
|
||||
- For each: show example vector component quantization
|
||||
- Display memory savings: 32x, 24x, 16x respectively
|
||||
- Include dimension recommendations for each
|
||||
- Use color coding: 1-bit (red/blue binary), 1.5-bit (3 colors), 2-bit (4 colors)
|
||||
- Add note that ColModernVBERT's 128 dims is below recommended range for binary quantization
|
||||
- Highlight that scalar quantization is preferred for low-dimensional vectors
|
||||
-->
|
||||
|
||||
### Real-World Impact for ColPali Collections
|
||||
|
||||
Let's compare all options for a **1 million document** ColModernVBERT collection:
|
||||
|
||||
Reference in New Issue
Block a user