mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-04 10:28:29 +02:00
Finish ColPali family lesson
This commit is contained in:
@@ -26,12 +26,6 @@ Let's explore what the options are and which model to choose depending on the da
|
||||
|
||||
---
|
||||
|
||||
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-2/colpali-family.ipynb">
|
||||
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||
</a>
|
||||
|
||||
---
|
||||
|
||||
The ColPali family includes several model variants. When selecting a model for your application, you'll need to consider factors like model size, supported languages, computational requirements, and licensing constraints - each variant offers different trade-offs along these dimensions.
|
||||
|
||||
## Model Size
|
||||
@@ -40,7 +34,7 @@ While the original ColPali model delivers good performance, its multi-billion pa
|
||||
|
||||
### ColSmol: Efficient Small-Scale Models
|
||||
|
||||
The **[ColSmol](https://huggingface.co/vidore/colSmol-256M)** family offers compact variants built on SmolVLM, available in [256M](https://huggingface.co/vidore/colSmol-256M) and [500M](https://huggingface.co/vidore/colSmol-500M) parameter sizes that achieve 80.1 and 82.3 nDCG@5 respectively on the ViDoRe benchmark. These Apache 2.0 licensed models generate ColBERT-style multi-vector representations while being small enough for browser-based applications, edge computing, and resource-constrained environments.
|
||||
The **[ColSmol](https://huggingface.co/vidore/colSmol-256M)** family offers compact variants built on SmolVLM, available in [256M](https://huggingface.co/vidore/colSmol-256M) and [500M](https://huggingface.co/vidore/colSmol-500M) parameter sizes. These Apache 2.0 licensed models generate ColBERT-style multi-vector representations while being small enough for resource-constrained environments, like browser-based applications, or edge computing.
|
||||
|
||||
### ColFlor: Ultra-Compact Retrieval
|
||||
|
||||
@@ -60,6 +54,16 @@ However, **potential users should be aware of licensing restrictions**. The mode
|
||||
|
||||
For commercial applications, **[Nomic AI's ColNomic-Embed-Multimodal-7B](https://huggingface.co/nomic-ai/colnomic-embed-multimodal-7b)** offers a compelling fully open-source alternative. Released in early 2025, this model demonstrated competitive performance on multilingual retrieval benchmarks. The model is available under an open-source license that permits commercial use, making it suitable for production deployments without licensing concerns. Nomic AI released a complete suite including both multi-vector (ColNomic) and single-vector variants in 3B and 7B parameter sizes, giving developers flexibility in choosing the right trade-off between performance and resource requirements.
|
||||
|
||||
## Benchmarking Visual Document Retrieval
|
||||
|
||||
The **ViDoRe (Visual Document Retrieval) Benchmark** has emerged as the leading evaluation framework for visual retrieval models on document understanding tasks. As of January 2026, it stands as the largest and most comprehensive benchmark in the field, evaluating models across multiple domains, 6 languages (including English and French), and realistic retrieval scenarios including cross-document and long-form queries. You can explore the latest model performance and compare different approaches on the [ViDoRe leaderboard](https://huggingface.co/spaces/vidore/vidore-leaderboard).
|
||||
|
||||

|
||||
|
||||
***Source:** https://huggingface.co/vidore*
|
||||
|
||||
The current version, **ViDoRe V3**, represents the benchmark's scale and ambition with 26,000+ pages across 3,099 queries in 6 languages, spanning 10 datasets (8 public, 2 private). The benchmark uses **nDCG** (Normalized Discounted Cumulative Gain) as its primary metric and includes challenging datasets spanning diverse domains - including HR, finance, industrial, pharmaceuticals, physics, computer science, and energy sectors. What sets ViDoRe V3 apart is its focus on real-world complexity: models are tested on truly challenging retrieval tasks with human-verified annotations that mirror actual user behavior in enterprise document retrieval scenarios.
|
||||
|
||||
## The Impact of Bidirectional Attention
|
||||
|
||||
Bidirectional attention has emerged as a promising approach for multi-vector representations of multi-modal data.
|
||||
@@ -72,70 +76,54 @@ Most large language models (LLMs) and vision-language models (VLMs) like GPT, Ll
|
||||
|
||||
This design makes perfect sense for generative tasks: when predicting the next token, the model should only use past context, not future information it hasn't generated yet. However, this constraint creates limitations for embedding tasks. Token representations encode only information from previous context, missing crucial contextual information from subsequent tokens in the sequence.
|
||||
|
||||
<!-- TODO: Add diagram for unidirectional attention
|
||||

|
||||
|
||||
Show:
|
||||
- Sequence of tokens (T1, T2, T3, T4, T5)
|
||||
- Focus on T3 as the current token being processed
|
||||
- Show arrows from T3 pointing backward to T1, T2, T3
|
||||
- Gray out or fade T4, T5 to show they're not accessible
|
||||
- Label: "Unidirectional Attention: Only Past Context"
|
||||
- Add annotation: "When processing T3, the model can only see tokens T1, T2, T3"
|
||||
The default ColPali model uses the unidirectional attention, derived from the underlying VLM. As a reminder, here is how we load the ColPali v.1.3 with FastEmbed.
|
||||
|
||||
Use color coding:
|
||||
- Blue arrows for backward attention flow
|
||||
- Darker highlight on current token (T3)
|
||||
- Lighter/grayed future tokens to emphasize they're hidden
|
||||
-->
|
||||
```python
|
||||
from fastembed import LateInteractionMultimodalEmbedding
|
||||
|
||||
# Load the Qdrant/colpali-v1.3-fp16 model from HF hub
|
||||
colpali_model = LateInteractionMultimodalEmbedding(
|
||||
model_name="Qdrant/colpali-v1.3-fp16"
|
||||
)
|
||||
```
|
||||
|
||||
### Bidirectional Attention: Optimized for Embeddings
|
||||
|
||||
**Bidirectional attention**, as used in encoder models like BERT, allows each token to attend to the entire input sequence - both past and future tokens. This creates richer, more contextually informed representations since each token's embedding incorporates information from the complete surrounding context.
|
||||
|
||||
<!-- TODO: Add diagram for bidirectional attention
|
||||

|
||||
|
||||
Show:
|
||||
- Same sequence of tokens (T1, T2, T3, T4, T5)
|
||||
- Focus on T3 as the current token being processed
|
||||
- Show arrows from T3 pointing in BOTH directions to all tokens (T1, T2, T3, T4, T5)
|
||||
- All tokens highlighted/visible to show full context available
|
||||
- Label: "Bidirectional Attention: Full Context"
|
||||
- Add annotation: "When processing T3, the model can see ALL tokens in both directions"
|
||||
|
||||
Use color coding:
|
||||
- Blue arrows for backward attention (T3 → T1, T2)
|
||||
- Green arrows for forward attention (T3 → T4, T5)
|
||||
- Darker highlight on current token (T3)
|
||||
- All tokens equally visible/highlighted to show full accessibility
|
||||
- Add contrast note: "Captures richer semantic representations by seeing the entire sequence"
|
||||
-->
|
||||
|
||||
For multi-vector retrieval tasks, this architectural choice proves crucial. Research has shown that [bidirectional encoder models are often the best option when training visual retrievers](https://arxiv.org/html/2510.01149). The [**ColModernVBERT**](https://huggingface.co/ModernVBERT/colmodernvbert) model demonstrates this advantage: despite having over 10 times fewer parameters than ColPali, it achieves performance only slightly lower on the ViDoRe benchmark.
|
||||
|
||||
The ModernVBERT architecture leverages bidirectional attention to create compact yet powerful visual document retrievers. By allowing full context flow in both directions, these models generate embeddings that capture nuanced semantic relationships - critical for accurate multi-modal retrieval where visual and textual information must be jointly understood.
|
||||
|
||||
## Benchmark Results
|
||||
|
||||
The **ViDoRe (Visual Document Retrieval) Benchmark** is a comprehensive evaluation framework designed to measure the performance of visual retrieval models on document understanding tasks. It evaluates models across multiple domains, languages (English, French, Spanish, and German), and realistic retrieval scenarios including cross-document and long-form queries. You can explore the latest model performance and compare different approaches on the [ViDoRe leaderboard](https://huggingface.co/spaces/vidore/vidore-leaderboard).
|
||||
|
||||
The benchmark uses **nDCG@5** (Normalized Discounted Cumulative Gain at 5) as its primary metric and includes challenging datasets spanning various domains - from biomedical research papers to insurance documents and ESG reports. Unlike earlier benchmarks, ViDoRe V2 emphasizes real-world complexity through blind contextual querying and human-in-the-loop evaluation, ensuring that models are tested on truly challenging retrieval tasks that mirror actual user behavior.
|
||||
Research has shown that [bidirectional encoder models are often the best option when training visual retrievers](https://arxiv.org/html/2510.01149). The [**ColModernVBERT**](https://huggingface.co/ModernVBERT/colmodernvbert) model exemplifies this approach: with only ~250 million parameters - over 10 times fewer than ColPali - it achieves performance only slightly lower on the ViDoRe benchmark.
|
||||
|
||||

|
||||
|
||||
***Source:** Teiletche, P., Macé, Q., Conti, M., Loison, A., Viaud, G., Colombo, P., & Faysse, M. (2025). *ModernVBERT: Towards Smaller Visual Document Retrievers*. arXiv preprint arXiv:2510.01149. https://arxiv.org/abs/2510.01149*
|
||||
|
||||
The chart above shows the performance of various models on the ViDoRe benchmark. Models using bidirectional attention, particularly those in the ColPali and ColQwen families, consistently demonstrate superior performance on visual document retrieval tasks.
|
||||
The chart above compares performance on the ViDoRe V2 benchmark. While ColPali and ColQwen achieve strong results with unidirectional attention, **ColModernVBERT stands out as the only bidirectional model** - proving that full-context attention enables competitive performance with dramatically fewer parameters. By allowing context to flow in both directions, these compact models generate embeddings that capture the nuanced semantic relationships critical for accurate multi-modal retrieval where visual and textual information must be jointly understood.
|
||||
|
||||
It's important to note that **ColModernVBERT achieves competitive results with only ~250 million parameters** - significantly smaller than ColPali and ColQwen models which contain billions of parameters. This makes ModernVBERT particularly attractive for resource-constrained environments or CPU-only deployments, where the slight performance trade-off is worthwhile for the substantial efficiency gains.
|
||||
This efficiency makes ModernVBERT particularly attractive for resource-constrained environments or CPU-only deployments, where the slight performance trade-off is worthwhile for the substantial gains in speed and reduced computational requirements.
|
||||
|
||||
ColModernVBERT is available in FastEmbed and might be used like any other late interaction model for multi-modal data.
|
||||
|
||||
```python
|
||||
from fastembed import LateInteractionMultimodalEmbedding
|
||||
|
||||
# Load the Qdrant/colmodernvbert model from HF hub
|
||||
colpali_model = LateInteractionMultimodalEmbedding(
|
||||
model_name="Qdrant/colmodernvbert"
|
||||
)
|
||||
```
|
||||
|
||||
## Future of Multi-Vector Representations
|
||||
|
||||
The multi-vector approach extends naturally to new modalities beyond images. Models like **[TomoroAI/tomoro-colqwen3-embed-8b](https://huggingface.co/TomoroAI/tomoro-colqwen3-embed-8b)**, based on ColQwen3, demonstrate this extensibility by adding support for short video retrieval. Based on preliminary findings, ColQwen3 generalizes to videos while learning from image-text retrieval tasks - it samples video clips, encodes frames, then pools frame embeddings with per-dimension max before MaxSim scoring.
|
||||
|
||||
While not yet fine-tuned on large-scale video retrieval datasets, this lightweight approach highlights a key strength of the multi-vector paradigm: **new modalities can be added by adapting the preprocessing pipeline** while preserving the core late interaction architecture.
|
||||
While not yet fine-tuned on large-scale video retrieval datasets, this lightweight approach highlights a key strength of the multi-vector paradigm, where **new modalities can be added** while preserving the core late interaction architecture.
|
||||
|
||||
## What's next
|
||||
|
||||
TODO: summarize the lesson quickly and introduce the next part
|
||||
You've explored the ColPali family ecosystem - from ultra-compact models like ColFlor (174M parameters) to multilingual alternatives from NVIDIA and Nomic AI. You've learned how bidirectional attention architectures can achieve competitive results with far fewer parameters, and seen the benchmark results that guide model selection based on your performance and resource requirements.
|
||||
|
||||
Now that you know which model to use, let's learn how to interpret what ColPali "sees" in your documents.
|
||||
|
||||
@@ -26,4 +26,10 @@ This transparency is invaluable for building trust in multi-modal search systems
|
||||
|
||||
---
|
||||
|
||||
TODO: Colab link
|
||||
|
||||
---
|
||||
|
||||
TODO: write the lesson content
|
||||
|
||||
With a solid understanding of ColPali, let's start building a real application in Module 3.
|
||||
|
||||
Reference in New Issue
Block a user