mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-04 10:28:29 +02:00
Add diagrams to late interaction basics
This commit is contained in:
+13
-6
@@ -44,6 +44,8 @@ We can categorize approaches based on when this interaction occurs:
|
||||
- **Early interaction:** Query and document are encoded together, with each word attending to the other during the encoding process. Maximum interaction, but no pre-computation.
|
||||
- **Late interaction:** Query and document are encoded independently (like no interaction), but we preserve fine-grained representations that interact during scoring (late in the process).
|
||||
|
||||

|
||||
|
||||
Let's examine each paradigm to understand the trade-offs. A simple example will illustrate the differences between all the methods.
|
||||
|
||||
```python
|
||||
@@ -55,7 +57,6 @@ documents = [
|
||||
query = "What is Qdrant?"
|
||||
```
|
||||
|
||||
|
||||
### Single-Vector Embeddings (No Interaction)
|
||||
|
||||
The most common approach encodes each document and query into a single dense vector, then compares them using similarity (most often cosine similarity).
|
||||
@@ -106,7 +107,9 @@ Let's calculate the similarity with the second document as well.
|
||||
np.dot(dense_query_vector, next(dense_generator))
|
||||
```
|
||||
|
||||
Notice how each document and query produces exactly **one vector** of 384 dimensions.
|
||||
Notice how each document and query produces exactly **one vector** of 384 dimensions. To achieve this compression, the model internally generates embeddings for each token, then uses **pooling** (typically mean pooling or a special [CLS] token) to aggregate them into a single fixed-size representation.
|
||||
|
||||

|
||||
|
||||
### Cross-Encoders (Early Interaction)
|
||||
|
||||
@@ -144,14 +147,16 @@ Late interaction solves this challenge through a simple but powerful idea: **enc
|
||||
|
||||
### How It Works
|
||||
|
||||
Instead of compressing a document into a single vector, late interaction represents it as a collection of contextualized token embeddings:
|
||||
Instead of compressing a document into a single vector, late interaction represents it as a collection of contextualized token embeddings
|
||||
|
||||
1. **Encode independently:** Pass the document through an encoder (like BERT) to generate one embedding vector per token
|
||||
2. **Store multi-vector representations:** Keep all token vectors rather than aggregating them into a single vector
|
||||
1. **Encode:** Pass the document through an encoder (like ColBERT) to generate one embedding vector per token
|
||||
2. **Store multi-vector representations:** Keep all token vectors instead of aggregating them into a single vector
|
||||
3. **Defer comparison:** At search time, compare query token vectors against document token vectors
|
||||
4. **Late interaction:** The actual "interaction" between query and document happens late - only when computing relevance scores
|
||||
|
||||
This is fundamentally different from both single-vector and cross-encoder approaches:
|
||||

|
||||
|
||||
This isn't that much fundamentally different from both single-vector and cross-encoder approaches, yet there are some differences:
|
||||
- Unlike single-vector: We preserve fine-grained, token-level information
|
||||
- Unlike cross-encoders: We encode documents independently, enabling pre-computation
|
||||
|
||||
@@ -163,6 +168,8 @@ This is fundamentally different from both single-vector and cross-encoder approa
|
||||
|
||||
**Contextual understanding:** Token embeddings are contextualized by the surrounding text. The word "bank" has different embeddings in "river bank" versus "financial bank."
|
||||
|
||||

|
||||
|
||||
**Scalability:** While requiring more storage than single vectors, the deferred comparison enables searching large collections efficiently.
|
||||
|
||||
### The ColBERT Approach
|
||||
|
||||
Reference in New Issue
Block a user