Add TODOs for diagrams + video embeddings

This commit is contained in:
Kacper Łukawski
2026-01-05 10:27:16 +01:00
parent fa2412d966
commit 07964ca88f
@@ -60,54 +60,63 @@ However, **potential users should be aware of licensing restrictions**. The mode
For commercial applications, **[Nomic AI's ColNomic-Embed-Multimodal-7B](https://huggingface.co/nomic-ai/colnomic-embed-multimodal-7b)** offers a compelling fully open-source alternative. Released in early 2025, this model demonstrated competitive performance on multilingual retrieval benchmarks. The model is available under an open-source license that permits commercial use, making it suitable for production deployments without licensing concerns. Nomic AI released a complete suite including both multi-vector (ColNomic) and single-vector variants in 3B and 7B parameter sizes, giving developers flexibility in choosing the right trade-off between performance and resource requirements.
## An Impact of the Bidirectional Attention
## The Impact of Bidirectional Attention
As of the end of 2025, the use of bidirectional attention seems to be a promising approach to multi-vector representations for multi-modal data.
Bidirectional attention has emerged as a promising approach for multi-vector representations of multi-modal data.
The choice between **unidirectional** and **bidirectional** attention mechanisms significantly impacts model performance for embedding tasks. Understanding this distinction helps explain why certain architectures excel at retrieval while others are optimized for generation.
The choice between **unidirectional** and **bidirectional** attention mechanisms significantly impacts model performance for embedding tasks. Understanding this distinction helps explain why certain architectures excel at retrieval while others are optimized for generation. Let's first examine unidirectional attention and its limitations, then see how bidirectional attention addresses these constraints.
### Unidirectional Attention: Designed for Generation
Most large language models (LLMs) and vision-language models (VLMs) like GPT, Llama, and the base models of [ColPali](https://huggingface.co/vidore/colpali) use **unidirectional (causal) attention**. In this approach, each token can only attend to tokens that came before it in the sequence - the model looks backward but never forward.
This design makes perfect sense for generative tasks: when predicting the next token, the model should only use past context, not future information it hasn't generated yet. However, this constraint creates limitations for embedding tasks. Word representations encode only information from previous context, missing the valuable signal from tokens that appear later in the sequence.
This design makes perfect sense for generative tasks: when predicting the next token, the model should only use past context, not future information it hasn't generated yet. However, this constraint creates limitations for embedding tasks. Token representations encode only information from previous context, missing crucial contextual information from subsequent tokens in the sequence.
<!-- TODO: Add diagram comparing unidirectional vs bidirectional attention
<!-- TODO: Add diagram for unidirectional attention
Show side-by-side comparison:
- Left panel: Unidirectional Attention
- Sequence of tokens (T1, T2, T3, T4, T5)
- Show that T3 can only attend to T1, T2, T3 (backwards arrows)
- Label: "For generation (predicting next token)"
- Gray out future tokens to show they're not accessible
Show:
- Sequence of tokens (T1, T2, T3, T4, T5)
- Focus on T3 as the current token being processed
- Show arrows from T3 pointing backward to T1, T2, T3
- Gray out or fade T4, T5 to show they're not accessible
- Label: "Unidirectional Attention: Only Past Context"
- Add annotation: "When processing T3, the model can only see tokens T1, T2, T3"
- Right panel: Bidirectional Attention
- Same sequence of tokens (T1, T2, T3, T4, T5)
- Show that T3 can attend to ALL tokens (arrows in both directions)
- Label: "For embeddings (understanding full context)"
- All tokens highlighted to show full context
- Use color coding:
- Blue arrows for backward attention
- Green arrows for forward attention
- Darker highlight on current token being processed
- Add annotation: "Bidirectional models capture richer semantic representations by seeing the entire sequence"
Use color coding:
- Blue arrows for backward attention flow
- Darker highlight on current token (T3)
- Lighter/grayed future tokens to emphasize they're hidden
-->
### Bidirectional Attention: Optimized for Embeddings
**Bidirectional attention**, as used in encoder models like BERT, allows each token to attend to the entire input sequence - both past and future tokens. This creates richer, more contextually informed representations since each token's embedding incorporates information from the complete surrounding context.
For multi-vector retrieval tasks, this architectural choice proves crucial. Research has shown that [bidirectional encoder models are often the best option when training visual retrievers](https://arxiv.org/html/2510.01149). The [**ColModernVBERT**](https://huggingface.co/ModernVBERT/colmodernvbert) model demonstrates this advantage: despite having over 10 times fewer parameters than ColPali, it achieves performance only 0.6 nDCG@5 points lower on the ViDoRe benchmark.
<!-- TODO: Add diagram for bidirectional attention
Show:
- Same sequence of tokens (T1, T2, T3, T4, T5)
- Focus on T3 as the current token being processed
- Show arrows from T3 pointing in BOTH directions to all tokens (T1, T2, T3, T4, T5)
- All tokens highlighted/visible to show full context available
- Label: "Bidirectional Attention: Full Context"
- Add annotation: "When processing T3, the model can see ALL tokens in both directions"
Use color coding:
- Blue arrows for backward attention (T3 → T1, T2)
- Green arrows for forward attention (T3 → T4, T5)
- Darker highlight on current token (T3)
- All tokens equally visible/highlighted to show full accessibility
- Add contrast note: "Captures richer semantic representations by seeing the entire sequence"
-->
For multi-vector retrieval tasks, this architectural choice proves crucial. Research has shown that [bidirectional encoder models are often the best option when training visual retrievers](https://arxiv.org/html/2510.01149). The [**ColModernVBERT**](https://huggingface.co/ModernVBERT/colmodernvbert) model demonstrates this advantage: despite having over 10 times fewer parameters than ColPali, it achieves performance only slightly lower on the ViDoRe benchmark.
The ModernVBERT architecture leverages bidirectional attention to create compact yet powerful visual document retrievers. By allowing full context flow in both directions, these models generate embeddings that capture nuanced semantic relationships - critical for accurate multi-modal retrieval where visual and textual information must be jointly understood.
## Benchmark Results
The **ViDoRe (Visual Document Retrieval) Benchmark** is a comprehensive evaluation framework designed to measure the performance of visual retrieval models on document understanding tasks. It evaluates models across multiple domains, languages (English, French, Spanish, and German), and realistic retrieval scenarios including cross-document and long-form queries.
TODO: mention the leaderboard https://huggingface.co/spaces/vidore/vidore-leaderboard
The **ViDoRe (Visual Document Retrieval) Benchmark** is a comprehensive evaluation framework designed to measure the performance of visual retrieval models on document understanding tasks. It evaluates models across multiple domains, languages (English, French, Spanish, and German), and realistic retrieval scenarios including cross-document and long-form queries. You can explore the latest model performance and compare different approaches on the [ViDoRe leaderboard](https://huggingface.co/spaces/vidore/vidore-leaderboard).
The benchmark uses **nDCG@5** (Normalized Discounted Cumulative Gain at 5) as its primary metric and includes challenging datasets spanning various domains - from biomedical research papers to insurance documents and ESG reports. Unlike earlier benchmarks, ViDoRe V2 emphasizes real-world complexity through blind contextual querying and human-in-the-loop evaluation, ensuring that models are tested on truly challenging retrieval tasks that mirror actual user behavior.
@@ -121,7 +130,9 @@ It's important to note that **ColModernVBERT achieves competitive results with o
## Future of Multi-Vector Representations
TODO: focus on video retrieval of https://huggingface.co/TomoroAI/tomoro-colqwen3-embed-8b
The multi-vector approach extends naturally to new modalities beyond images. Models like **[TomoroAI/tomoro-colqwen3-embed-8b](https://huggingface.co/TomoroAI/tomoro-colqwen3-embed-8b)**, based on ColQwen3, demonstrate this extensibility by adding support for short video retrieval. Based on preliminary findings, ColQwen3 generalizes to videos while learning from image-text retrieval tasks - it samples video clips, encodes frames, then pools frame embeddings with per-dimension max before MaxSim scoring.
While not yet fine-tuned on large-scale video retrieval datasets, this lightweight approach highlights a key strength of the multi-vector paradigm: **new modalities can be added by adapting the preprocessing pipeline** while preserving the core late interaction architecture.
## What's next