Merge pull request #2053 from qdrant/course/multi-vector-search
Course: multi vector search
@@ -1,5 +1,5 @@
|
|||||||
---
|
---
|
||||||
title: "Welcome to Qdrant Academy"
|
title: "Qdrant Academy"
|
||||||
description: Master vector search and AI-powered applications with Qdrant Academy. Free, self-paced courses guide you from beginner to expert with hands-on projects, code notebooks, and certification.
|
description: Master vector search and AI-powered applications with Qdrant Academy. Free, self-paced courses guide you from beginner to expert with hands-on projects, code notebooks, and certification.
|
||||||
weight: 50
|
weight: 50
|
||||||
---
|
---
|
||||||
@@ -12,13 +12,11 @@ Qdrant Academy is your step-by-step learning hub for mastering vector search, hy
|
|||||||
|
|
||||||
Whether you’re new to Qdrant or building production-grade systems, our guided courses help you go from beginner to expert, one module at a time.
|
Whether you’re new to Qdrant or building production-grade systems, our guided courses help you go from beginner to expert, one module at a time.
|
||||||
|
|
||||||
Qdrant Academy currently offers one comprehensive course, but more are on the way! Register your interest for upcoming courses below, or take the available course and [get certified](https://train.qdrant.dev)!
|
|
||||||
|
|
||||||
## Available Now
|
## Available Now
|
||||||
|
|
||||||
{{< course-card
|
{{< course-card
|
||||||
title="Qdrant Essentials Course"
|
title="Qdrant Essentials Course"
|
||||||
image="/icons/outline/rocket-white-light.svg"
|
image="/icons/outline/rocket-white-light.svg"
|
||||||
link="/course/essentials/"
|
link="/course/essentials/"
|
||||||
>}}
|
>}}
|
||||||
**What you’ll gain:**
|
**What you’ll gain:**
|
||||||
@@ -30,7 +28,24 @@ Qdrant Academy currently offers one comprehensive course, but more are on the wa
|
|||||||
- Ecosystem Integrations (Bonus)
|
- Ecosystem Integrations (Bonus)
|
||||||
<br><br>
|
<br><br>
|
||||||
Time to Complete: 9-12 hours<br>
|
Time to Complete: 9-12 hours<br>
|
||||||
Includes: videos, code notebooks, projects, walkthroughs
|
Includes: videos, code notebooks, projects, certification
|
||||||
|
{{< /course-card >}}
|
||||||
|
|
||||||
|
{{< course-card
|
||||||
|
title="Multi-Vector Search Course"
|
||||||
|
image="/icons/outline/similarity-blue.svg"
|
||||||
|
link="/course/multi-vector-search/"
|
||||||
|
>}}
|
||||||
|
**What you’ll gain:**
|
||||||
|
- Late Interaction Models and MaxSim Scoring
|
||||||
|
- ColBERT for Text Search
|
||||||
|
- ColPali for Visual Document Search
|
||||||
|
- Multi-Stage Retrieval Pipelines
|
||||||
|
- Quantization and Pooling Techniques
|
||||||
|
- MUVERA Indexing for Large-Scale Search
|
||||||
|
<br><br>
|
||||||
|
Time to Complete: 4-6 hours<br>
|
||||||
|
Includes: videos, code notebooks, projects, certification
|
||||||
{{< /course-card >}}
|
{{< /course-card >}}
|
||||||
|
|
||||||
## Upcoming Courses
|
## Upcoming Courses
|
||||||
|
|||||||
@@ -1,6 +1,7 @@
|
|||||||
---
|
---
|
||||||
title: "Qdrant Essentials Certification"
|
title: "Qdrant Essentials Certification"
|
||||||
description: Get officially certified by Qdrant today!
|
description: Get officially certified by Qdrant today!
|
||||||
|
isLesson: true
|
||||||
weight: 100
|
weight: 100
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Implementing a Basic Vector Search"
|
title: "Implementing a Basic Vector Search"
|
||||||
description: Learn how to build a basic vector search in Qdrant. Create collections, insert vectors, and run your first similarity search step-by-step with Python.
|
description: Learn how to build a basic vector search in Qdrant. Create collections, insert vectors, and run your first similarity search step-by-step with Python.
|
||||||
weight: 3
|
weight: 3
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 0 {{< /date >}}
|
{{< date >}} Day 0 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Project: Building Your First Vector Search System"
|
title: "Project: Building Your First Vector Search System"
|
||||||
description: Apply your Qdrant skills to build a complete vector search system. Create collections, insert data, run similarity and filtered searches, and share your results.
|
description: Apply your Qdrant skills to build a complete vector search system. Create collections, insert data, run similarity and filtered searches, and share your results.
|
||||||
weight: 4
|
weight: 4
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 0 {{< /date >}}
|
{{< date >}} Day 0 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Qdrant Setup"
|
title: "Qdrant Setup"
|
||||||
description: Set up your Qdrant Cloud cluster in minutes. Learn to create collections, manage data, access the Web UI, and connect securely from Python.
|
description: Set up your Qdrant Cloud cluster in minutes. Learn to create collections, manage data, access the Web UI, and connect securely from Python.
|
||||||
weight: 2
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 0 {{< /date >}}
|
{{< date >}} Day 0 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Text Chunking Strategies"
|
title: "Text Chunking Strategies"
|
||||||
description: Learn how to split text into meaningful chunks for vector search. Compare six chunking strategies and discover how metadata improves retrieval precision in Qdrant.
|
description: Learn how to split text into meaningful chunks for vector search. Compare six chunking strategies and discover how metadata improves retrieval precision in Qdrant.
|
||||||
weight: 4
|
weight: 4
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 1 {{< /date >}}
|
{{< date >}} Day 1 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Distance Metrics"
|
title: "Distance Metrics"
|
||||||
description: Learn how distance metrics like cosine, Euclidean, Manhattan, and dot product shape vector similarity in Qdrant. Discover which metric fits your data and use case.
|
description: Learn how distance metrics like cosine, Euclidean, Manhattan, and dot product shape vector similarity in Qdrant. Discover which metric fits your data and use case.
|
||||||
weight: 3
|
weight: 3
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 1 {{< /date >}}
|
{{< date >}} Day 1 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Points, Vectors and Payloads"
|
title: "Points, Vectors and Payloads"
|
||||||
description: Learn Qdrant’s core data model with points, vectors, payloads, and named vectors. Compare dense, sparse, and multivectors, understand dimensionality trade-offs, and master filtering with payload indexes for precise retrieval.
|
description: Learn Qdrant’s core data model with points, vectors, payloads, and named vectors. Compare dense, sparse, and multivectors, understand dimensionality trade-offs, and master filtering with payload indexes for precise retrieval.
|
||||||
weight: 2
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 1 {{< /date >}}
|
{{< date >}} Day 1 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Demo: Semantic Movie Search"
|
title: "Demo: Semantic Movie Search"
|
||||||
description: Build a semantic movie search with Qdrant. Compare chunking strategies, embed descriptions, and combine cosine similarity with metadata filters and grouping for accurate, theme-aware recommendations.
|
description: Build a semantic movie search with Qdrant. Compare chunking strategies, embed descriptions, and combine cosine similarity with metadata filters and grouping for accurate, theme-aware recommendations.
|
||||||
weight: 5
|
weight: 5
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 1 {{< /date >}}
|
{{< date >}} Day 1 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Project: Building a Semantic Search Engine"
|
title: "Project: Building a Semantic Search Engine"
|
||||||
description: Build a semantic search engine with Qdrant. Compare chunking strategies, index embeddings, and query by meaning to discover what works best for your domain.
|
description: Build a semantic search engine with Qdrant. Compare chunking strategies, index embeddings, and query by meaning to discover what works best for your domain.
|
||||||
weight: 6
|
weight: 6
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 1 {{< /date >}}
|
{{< date >}} Day 1 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Demo: HNSW Performance Tuning"
|
title: "Demo: HNSW Performance Tuning"
|
||||||
description: Tune Qdrant’s HNSW index for speed and precision. Optimize bulk uploads, test filters, and benchmark performance on a real 100K OpenAI embedding dataset.
|
description: Tune Qdrant’s HNSW index for speed and precision. Optimize bulk uploads, test filters, and benchmark performance on a real 100K OpenAI embedding dataset.
|
||||||
weight: 4
|
weight: 4
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 2 {{< /date >}}
|
{{< date >}} Day 2 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Combining Vector Search and Filtering"
|
title: "Combining Vector Search and Filtering"
|
||||||
description: Learn how Qdrant combines HNSW vector search with payload filtering. Understand Filterable HNSW, query planning, and payload indexing for accurate, high-performance retrieval.
|
description: Learn how Qdrant combines HNSW vector search with payload filtering. Understand Filterable HNSW, query planning, and payload indexing for accurate, high-performance retrieval.
|
||||||
weight: 3
|
weight: 3
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 2 {{< /date >}}
|
{{< date >}} Day 2 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Project: HNSW Performance Benchmarking"
|
title: "Project: HNSW Performance Benchmarking"
|
||||||
description: Optimize vector search with Qdrant. Test multiple HNSW configurations, time uploads and queries, and evaluate filtering with and without payload indexes to find the best settings for your domain.
|
description: Optimize vector search with Qdrant. Test multiple HNSW configurations, time uploads and queries, and evaluate filtering with and without payload indexes to find the best settings for your domain.
|
||||||
weight: 5
|
weight: 5
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 2 {{< /date >}}
|
{{< date >}} Day 2 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "HNSW Indexing Fundamentals"
|
title: "HNSW Indexing Fundamentals"
|
||||||
description: Learn how HNSW indexing powers fast, scalable vector search in Qdrant. Understand parameters like m, ef_construct, and hnsw_ef to balance recall, speed, and memory efficiency.
|
description: Learn how HNSW indexing powers fast, scalable vector search in Qdrant. Understand parameters like m, ef_construct, and hnsw_ef to balance recall, speed, and memory efficiency.
|
||||||
weight: 2
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 2 {{< /date >}}
|
{{< date >}} Day 2 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Demo: Implementing a Hybrid Search System"
|
title: "Demo: Implementing a Hybrid Search System"
|
||||||
description: Step-by-step demo on implementing hybrid search using Qdrant’s Universal Query API. Explore dense vs. sparse search, score fusion algorithms, and real-world evaluation techniques.
|
description: Step-by-step demo on implementing hybrid search using Qdrant’s Universal Query API. Explore dense vs. sparse search, score fusion algorithms, and real-world evaluation techniques.
|
||||||
weight: 5
|
weight: 5
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 3 {{< /date >}}
|
{{< date >}} Day 3 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Hybrid Search and the Universal Query API"
|
title: "Hybrid Search and the Universal Query API"
|
||||||
description: Master hybrid search in Qdrant using dense and sparse vectors. Explore retrieval, reranking, and Reciprocal Rank Fusion (RRF) to build efficient, adaptive search experiences.
|
description: Master hybrid search in Qdrant using dense and sparse vectors. Explore retrieval, reranking, and Reciprocal Rank Fusion (RRF) to build efficient, adaptive search experiences.
|
||||||
weight: 4
|
weight: 4
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 3 {{< /date >}}
|
{{< date >}} Day 3 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Project: Building a Hybrid Search Engine"
|
title: "Project: Building a Hybrid Search Engine"
|
||||||
description: Build a hybrid search engine in Qdrant combining dense and sparse vectors with Reciprocal Rank Fusion. Compare performance, optimize retrieval, and understand when hybrid search outperforms single-vector methods.
|
description: Build a hybrid search engine in Qdrant combining dense and sparse vectors with Reciprocal Rank Fusion. Compare performance, optimize retrieval, and understand when hybrid search outperforms single-vector methods.
|
||||||
weight: 6
|
weight: 6
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 3 {{< /date >}}
|
{{< date >}} Day 3 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Demo: Keyword Search with Sparse Vectors"
|
title: "Demo: Keyword Search with Sparse Vectors"
|
||||||
description: Hands-on sparse retrieval in Qdrant—create BM25 collections, enable IDF, index with FastEmbed, try SPLADE++ expansion, and execute keyword queries via the Universal Query API.
|
description: Hands-on sparse retrieval in Qdrant—create BM25 collections, enable IDF, index with FastEmbed, try SPLADE++ expansion, and execute keyword queries via the Universal Query API.
|
||||||
weight: 3
|
weight: 3
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 3 {{< /date >}}
|
{{< date >}} Day 3 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Sparse Vectors and Inverted Indexes"
|
title: "Sparse Vectors and Inverted Indexes"
|
||||||
description: Learn sparse vectors and inverted indexes in Qdrant, create named sparse vectors, store index–value pairs, run exact dot-product search, and prepare for hybrid search with dense vectors.
|
description: Learn sparse vectors and inverted indexes in Qdrant, create named sparse vectors, store index–value pairs, run exact dot-product search, and prepare for hybrid search with dense vectors.
|
||||||
weight: 2
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 3 {{< /date >}}
|
{{< date >}} Day 3 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Large-Scale Data Ingestion"
|
title: "Large-Scale Data Ingestion"
|
||||||
description: Master large-scale vector ingestion in Qdrant. Explore batching, upload_points, and upload_collection methods, on-disk storage, and parallel streaming for billion-scale AI data pipelines.
|
description: Master large-scale vector ingestion in Qdrant. Explore batching, upload_points, and upload_collection methods, on-disk storage, and parallel streaming for billion-scale AI data pipelines.
|
||||||
weight: 4
|
weight: 4
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 4 {{< /date >}}
|
{{< date >}} Day 4 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Project: Quantization Performance Optimization"
|
title: "Project: Quantization Performance Optimization"
|
||||||
description: Apply vector quantization in Qdrant to boost search speed, reduce memory, and balance accuracy. Test scalar, binary, and 2-bit quantization with oversampling and rescoring optimization.
|
description: Apply vector quantization in Qdrant to boost search speed, reduce memory, and balance accuracy. Test scalar, binary, and 2-bit quantization with oversampling and rescoring optimization.
|
||||||
weight: 5
|
weight: 5
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 4 {{< /date >}}
|
{{< date >}} Day 4 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Accuracy Recovery with Rescoring"
|
title: "Accuracy Recovery with Rescoring"
|
||||||
description: Learn how oversampling and rescoring restore accuracy in quantized vector search. Improve Qdrant search precision while maintaining high performance using efficient reranking on original vectors.
|
description: Learn how oversampling and rescoring restore accuracy in quantized vector search. Improve Qdrant search precision while maintaining high performance using efficient reranking on original vectors.
|
||||||
weight: 3
|
weight: 3
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 4 {{< /date >}}
|
{{< date >}} Day 4 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Vector Quantization Methods"
|
title: "Vector Quantization Methods"
|
||||||
description: Explore scalar, binary, and product quantization in Qdrant. Learn how compression boosts vector search speed, cuts memory costs, and balances accuracy for large-scale AI retrieval.
|
description: Explore scalar, binary, and product quantization in Qdrant. Learn how compression boosts vector search speed, cuts memory costs, and balances accuracy for large-scale AI retrieval.
|
||||||
weight: 2
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 4 {{< /date >}}
|
{{< date >}} Day 4 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Multivectors for Late Interaction Models"
|
title: "Multivectors for Late Interaction Models"
|
||||||
description: Learn how Qdrant supports late interaction models like ColBERT and ColPali using multivectors for token-level precision, enabling fine-grained, context-aware text and visual document retrieval.
|
description: Learn how Qdrant supports late interaction models like ColBERT and ColPali using multivectors for token-level precision, enabling fine-grained, context-aware text and visual document retrieval.
|
||||||
weight: 2
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 5 {{< /date >}}
|
{{< date >}} Day 5 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Project: Building a Recommendation System"
|
title: "Project: Building a Recommendation System"
|
||||||
description: Build a hybrid AI recommendation system with Qdrant’s Universal Query API—combining dense, sparse, and multivector retrieval, ColBERT reranking, and RRF fusion in one atomic query.
|
description: Build a hybrid AI recommendation system with Qdrant’s Universal Query API—combining dense, sparse, and multivector retrieval, ColBERT reranking, and RRF fusion in one atomic query.
|
||||||
weight: 5
|
weight: 5
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 5 {{< /date >}}
|
{{< date >}} Day 5 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "The Universal Query API"
|
title: "The Universal Query API"
|
||||||
description: Learn how to run dense, sparse, and ColBERT multivector retrieval with Qdrant’s Universal Query API—fusing, filtering, and reranking results in a single atomic request.
|
description: Learn how to run dense, sparse, and ColBERT multivector retrieval with Qdrant’s Universal Query API—fusing, filtering, and reranking results in a single atomic request.
|
||||||
weight: 3
|
weight: 3
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 5 {{< /date >}}
|
{{< date >}} Day 5 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Demo: Universal Query for Hybrid Retrieval"
|
title: "Demo: Universal Query for Hybrid Retrieval"
|
||||||
description: Build a hybrid research discovery system using Qdrant’s Universal Query API—combine dense, sparse, and ColBERT vectors for semantic, keyword, and reranked retrieval in one query.
|
description: Build a hybrid research discovery system using Qdrant’s Universal Query API—combine dense, sparse, and ColBERT vectors for semantic, keyword, and reranked retrieval in one query.
|
||||||
weight: 4
|
weight: 4
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 5 {{< /date >}}
|
{{< date >}} Day 5 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Course Completion and Next Steps"
|
title: "Course Completion and Next Steps"
|
||||||
description: Complete your Qdrant course by earning certification and mastering hybrid, multivector, and production-ready vector search techniques—skills to design, evaluate, and deploy real-world AI search systems.
|
description: Complete your Qdrant course by earning certification and mastering hybrid, multivector, and production-ready vector search techniques—skills to design, evaluate, and deploy real-world AI search systems.
|
||||||
weight: 3
|
weight: 3
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 6 {{< /date >}}
|
{{< date >}} Day 6 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Final Project: Production-Ready Documentation Search Engine"
|
title: "Final Project: Production-Ready Documentation Search Engine"
|
||||||
description: Create a complete documentation search system with Qdrant, featuring hybrid retrieval, multivector reranking, and performance evaluation for a portfolio-ready, production-quality vector search application.
|
description: Create a complete documentation search system with Qdrant, featuring hybrid retrieval, multivector reranking, and performance evaluation for a portfolio-ready, production-quality vector search application.
|
||||||
weight: 2
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 6 {{< /date >}}
|
{{< date >}} Day 6 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Integrating with Camel AI"
|
title: "Integrating with Camel AI"
|
||||||
description: Learn how Camel AI and Qdrant enable automated RAG pipelines with multi-agent communication, vector-based memory, and seamless integration into live environments like Discord bots.
|
description: Learn how Camel AI and Qdrant enable automated RAG pipelines with multi-agent communication, vector-based memory, and seamless integration into live environments like Discord bots.
|
||||||
weight: 8
|
weight: 8
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 7 {{< /date >}}
|
{{< date >}} Day 7 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Integrating with Haystack"
|
title: "Integrating with Haystack"
|
||||||
description: Learn how Qdrant and Haystack combine to deliver end-to-end search and recommendation systems with hybrid retrieval, semantic filtering, and agentic AI orchestration.
|
description: Learn how Qdrant and Haystack combine to deliver end-to-end search and recommendation systems with hybrid retrieval, semantic filtering, and agentic AI orchestration.
|
||||||
weight: 2
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 7 {{< /date >}}
|
{{< date >}} Day 7 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Integrating with Jina AI"
|
title: "Integrating with Jina AI"
|
||||||
description: Learn how Jina AI’s Embeddings v4 and Qdrant enable advanced multimodal retrieval, supporting text-to-image, image-to-text, and hybrid search with high-performance vector storage.
|
description: Learn how Jina AI’s Embeddings v4 and Qdrant enable advanced multimodal retrieval, supporting text-to-image, image-to-text, and hybrid search with high-performance vector storage.
|
||||||
weight: 9
|
weight: 9
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 7 {{< /date >}}
|
{{< date >}} Day 7 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Integrating with LlamaIndex"
|
title: "Integrating with LlamaIndex"
|
||||||
description: Learn how LlamaIndex and Qdrant power intelligent RAG pipelines, function-calling agents, and cloud-synced vector search systems with structured workflows and dynamic query handling.
|
description: Learn how LlamaIndex and Qdrant power intelligent RAG pipelines, function-calling agents, and cloud-synced vector search systems with structured workflows and dynamic query handling.
|
||||||
weight: 6
|
weight: 6
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 7 {{< /date >}}
|
{{< date >}} Day 7 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Integrating with Quotient"
|
title: "Integrating with Quotient"
|
||||||
description: Learn how Quotient and Qdrant combine to deliver end-to-end AI monitoring, hallucination detection, and performance analytics for reliable, high-quality retrieval-augmented generation systems.
|
description: Learn how Quotient and Qdrant combine to deliver end-to-end AI monitoring, hallucination detection, and performance analytics for reliable, high-quality retrieval-augmented generation systems.
|
||||||
weight: 7
|
weight: 7
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 7 {{< /date >}}
|
{{< date >}} Day 7 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Integrating with Superlinked"
|
title: "Integrating with Superlinked"
|
||||||
description: Learn how Superlinked’s Mixture of Encoders and Qdrant enable rich, multi-modal embeddings that fuse semantic, numerical, and temporal data for optimized vector retrieval.
|
description: Learn how Superlinked’s Mixture of Encoders and Qdrant enable rich, multi-modal embeddings that fuse semantic, numerical, and temporal data for optimized vector retrieval.
|
||||||
weight: 5
|
weight: 5
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 7 {{< /date >}}
|
{{< date >}} Day 7 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Integrating with Tensorlake"
|
title: "Integrating with Tensorlake"
|
||||||
description: Learn how TensorLake and Qdrant combine document parsing, knowledge graphs, and vector search to build scalable, structured data lakes for advanced RAG and research discovery applications.
|
description: Learn how TensorLake and Qdrant combine document parsing, knowledge graphs, and vector search to build scalable, structured data lakes for advanced RAG and research discovery applications.
|
||||||
weight: 4
|
weight: 4
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 7 {{< /date >}}
|
{{< date >}} Day 7 {{< /date >}}
|
||||||
|
|||||||
@@ -2,6 +2,7 @@
|
|||||||
title: "Integrating with Unstructured.io"
|
title: "Integrating with Unstructured.io"
|
||||||
description: Learn how Unstructured.io and Qdrant transform unstructured enterprise data into structured embeddings through VLM document understanding, smart chunking, and secure, production-ready ETL pipelines.
|
description: Learn how Unstructured.io and Qdrant transform unstructured enterprise data into structured embeddings through VLM document understanding, smart chunking, and secure, production-ready ETL pipelines.
|
||||||
weight: 3
|
weight: 3
|
||||||
|
isLesson: true
|
||||||
---
|
---
|
||||||
|
|
||||||
{{< date >}} Day 7 {{< /date >}}
|
{{< date >}} Day 7 {{< /date >}}
|
||||||
|
|||||||
@@ -0,0 +1,166 @@
|
|||||||
|
---
|
||||||
|
title: "Multi-Vector Search Course"
|
||||||
|
page_title: "Qdrant Multi-Vector Search Course"
|
||||||
|
description: Master late interaction models, ColPali, and production optimization. Build scalable multi-vector search pipelines.
|
||||||
|
content:
|
||||||
|
sidebarTitle: "Multi-Vector Search Course"
|
||||||
|
menuTitle:
|
||||||
|
text: Course Overview
|
||||||
|
url: /course/multi-vector-search/
|
||||||
|
nextButton: Continue to Next Step
|
||||||
|
nextDay: Complete
|
||||||
|
title: "Multi-Vector Search"
|
||||||
|
description: Master late interaction models, ColPali, and production optimization. Build scalable multi-vector search pipelines.
|
||||||
|
partition: course
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
# Multi-Vector Search
|
||||||
|
|
||||||
|
**Build production-ready multi-vector search pipelines**
|
||||||
|
|
||||||
|
Go beyond single-vector embeddings with late interaction models like ColBERT and ColPali. Learn the MaxSim distance metric, optimize for billion-scale search, and evaluate your retrieval pipelines with industry-standard metrics.
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/WXEf6K6DBq8?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<br/>
|
||||||
|
|
||||||
|
{{< cards-list >}}
|
||||||
|
- icon: /icons/outline/play-white.svg
|
||||||
|
title: 4 modules
|
||||||
|
content: Focused lessons building from fundamentals to production
|
||||||
|
- icon: /icons/outline/cloud-check-blue.svg
|
||||||
|
title: Shareable certificate
|
||||||
|
content: Earn a digital certificate upon completion
|
||||||
|
- icon: /icons/outline/time-blue.svg
|
||||||
|
title: Flexible schedule
|
||||||
|
content: Learn at your own pace (1–2 hours/module)
|
||||||
|
- icon: /icons/outline/plan.svg
|
||||||
|
title: Advanced level
|
||||||
|
content: Assumes familiarity with vector search basics
|
||||||
|
|
||||||
|
{{< /cards-list >}}
|
||||||
|
|
||||||
|
<br/>
|
||||||
|
|
||||||
|
## What you'll learn
|
||||||
|
{{< course-card
|
||||||
|
title="Skills you'll gain:"
|
||||||
|
image="/icons/outline/training-white.svg"
|
||||||
|
type="wide-list">}}
|
||||||
|
|
||||||
|
- Late interaction paradigm and MaxSim distance metric
|
||||||
|
- ColBERT for text and ColPali for visual documents
|
||||||
|
- Multi-stage retrieval with prefetch and reranking
|
||||||
|
- Quantization and pooling techniques for memory optimization
|
||||||
|
- MUVERA indexing for billion-scale search
|
||||||
|
- Evaluation metrics: Recall@k, NDCG, MRR
|
||||||
|
|
||||||
|
{{< /course-card >}}
|
||||||
|
|
||||||
|
### The Path
|
||||||
|
|
||||||
|
**Module 0**: Setup. Configure Qdrant Cloud or local instance and install Python dependencies.
|
||||||
|
|
||||||
|
**Module 1**: Text multi-vectors. Understand the late interaction paradigm, learn the MaxSim distance metric, explore use cases and challenges, and implement ColBERT with Qdrant.
|
||||||
|
|
||||||
|
**Module 2**: Multi-modal search. Apply multi-vector representations to images and PDFs with ColPali. Explore model variants and leverage visual interpretability for debugging.
|
||||||
|
|
||||||
|
**Module 3**: Optimization and evaluation. Master quantization, pooling, and MUVERA for memory-efficient search. Build multi-stage retrieval pipelines and evaluate with standard metrics.
|
||||||
|
|
||||||
|
## How the course works
|
||||||
|
|
||||||
|
{{< cards-list >}}
|
||||||
|
|
||||||
|
- icon: /icons/outline/training-purple.svg
|
||||||
|
title: Video-first lessons
|
||||||
|
content: Clear, concise modules by the Qdrant team
|
||||||
|
- icon: /icons/outline/hacker-purple.svg
|
||||||
|
title: Final project
|
||||||
|
content: Build a production-ready multi-modal search system
|
||||||
|
- icon: /icons/outline/similarity-blue.svg
|
||||||
|
title: Hands-on notebooks
|
||||||
|
content: Practice each concept with Colab notebooks
|
||||||
|
- icon: /icons/outline/copy.svg
|
||||||
|
title: Progressive learning
|
||||||
|
content: Build from fundamentals to advanced optimization
|
||||||
|
{{< /cards-list >}}
|
||||||
|
|
||||||
|
<br/>
|
||||||
|
|
||||||
|
## Syllabus
|
||||||
|
|
||||||
|
{{< accordion >}}
|
||||||
|
- title: "Module 0: Setting Up Dependencies"
|
||||||
|
content: |
|
||||||
|
- Qdrant Setup
|
||||||
|
- Installing Dependencies
|
||||||
|
<br>
|
||||||
|
<br>
|
||||||
|
<p style="margin-left: 0px;"><a href="/course/multi-vector-search/module-0/">→ Start Module 0</a></p>
|
||||||
|
|
||||||
|
- title: "Module 1: Multi-Vector Representations for Textual Data"
|
||||||
|
content: |
|
||||||
|
- Late Interaction Basics
|
||||||
|
- MaxSim Distance Metric
|
||||||
|
- Use Cases for Multi-Vector Search
|
||||||
|
- Problems of Multi-Vector Search
|
||||||
|
- Multi-Vector Embeddings in Qdrant
|
||||||
|
<br>
|
||||||
|
<br>
|
||||||
|
<p style="margin-left: 0px;"><a href="/course/multi-vector-search/module-1/">→ Start Module 1</a></p>
|
||||||
|
|
||||||
|
- title: "Module 2: Multi-Vector Representations for Multi-Modal Data"
|
||||||
|
content: |
|
||||||
|
- How ColPali Models Work
|
||||||
|
- ColPali Family Overview
|
||||||
|
- Visual Interpretability of ColPali
|
||||||
|
<br>
|
||||||
|
<br>
|
||||||
|
<p style="margin-left: 0px;"><a href="/course/multi-vector-search/module-2/">→ Start Module 2</a></p>
|
||||||
|
|
||||||
|
- title: "Module 3: Scalability and Optimization"
|
||||||
|
content: |
|
||||||
|
- Multi-Stage Retrieval with Universal Query API
|
||||||
|
- Vector Quantization Techniques
|
||||||
|
- Pooling Techniques
|
||||||
|
- MUVERA Indexing
|
||||||
|
- Evaluating Search Pipelines
|
||||||
|
- Final Project
|
||||||
|
<br>
|
||||||
|
<br>
|
||||||
|
<p style="margin-left: 0px;"><a href="/course/multi-vector-search/module-3/">→ Start Module 3</a></p>
|
||||||
|
{{< /accordion >}}
|
||||||
|
|
||||||
|
|
||||||
|
## Who it's for
|
||||||
|
|
||||||
|
ML, backend, and search engineers who want to go beyond single-vector embeddings. Requires intermediate Python, basic familiarity with vector search concepts (embeddings, similarity metrics), and comfort with APIs.
|
||||||
|
|
||||||
|
## Time commitment
|
||||||
|
|
||||||
|
- Duration: 4 modules at 1 hour/module
|
||||||
|
- Video learning: 1.5 hours
|
||||||
|
- Hands-on notebooks: 1.5 hours
|
||||||
|
- Final project: 1-3 hours
|
||||||
|
- Total: 4-6 hours
|
||||||
|
|
||||||
|
|
||||||
|
{{< course-card
|
||||||
|
title="Ready to master multi-vector search?"
|
||||||
|
image="/icons/outline/rocket-white-light.svg"
|
||||||
|
link="/course/multi-vector-search/module-0/">}}
|
||||||
|
**What you'll get**
|
||||||
|
- Build production-ready multi-vector pipelines
|
||||||
|
- Practice with real Colab notebooks
|
||||||
|
- Learn optimization techniques for scale
|
||||||
|
- Portfolio project and community support
|
||||||
|
{{< /course-card >}}
|
||||||
@@ -0,0 +1,26 @@
|
|||||||
|
---
|
||||||
|
title: "Qdrant Multi-Vector Certification"
|
||||||
|
description: "Get officially certified in multi-vector search by Qdrant."
|
||||||
|
url: /course/multi-vector-search/certification/
|
||||||
|
isLesson: true
|
||||||
|
weight: 50
|
||||||
|
---
|
||||||
|
|
||||||
|
# Qdrant Multi-Vector Search Certification
|
||||||
|
|
||||||
|
Congratulations! You’ve completed the **Multi-Vector Search course**. You didn’t just learn how to store vectors; you learned how to build high-performance retrieval systems using late interaction models and multi-vector representations.
|
||||||
|
|
||||||
|
You’ve moved past single-vector embeddings and dove deep into ColBERT, ColPali, MaxSim scoring, MUVERA, and production-grade multi-vector pipelines. That effort deserves more than just a “finished” status. It deserves professional recognition!
|
||||||
|
|
||||||
|
## Get #QdrantCertified
|
||||||
|
|
||||||
|
Your expertise is now production-ready. It’s time to validate those skills with our official certification.
|
||||||
|
|
||||||
|
Head over to [train.qdrant.dev](https://train.qdrant.dev) to take the exam.
|
||||||
|
|
||||||
|
Passing this exam proves you aren’t just a user; you are a Search Engineer capable of:
|
||||||
|
|
||||||
|
- Designing multi-vector retrieval pipelines with late interaction models.
|
||||||
|
- Applying ColPali and its variants for visual document search.
|
||||||
|
- Optimizing multi-vector search for both memory and latency.
|
||||||
|
- Mastering MaxSim scoring and the nuances of multi-vector architectures in Qdrant.
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
---
|
||||||
|
title: "Module 0: Setting Up Dependencies"
|
||||||
|
description: "Set up your development environment for multi-vector search. Install required dependencies and prepare your workspace for the course."
|
||||||
|
isLesson: true
|
||||||
|
weight: 10
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 0 {{< /date >}}
|
||||||
|
|
||||||
|
# Setting Up Dependencies
|
||||||
|
|
||||||
|
Get your environment ready for exploring multi-vector search with Qdrant.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Today's path
|
||||||
|
|
||||||
|
1. Qdrant Setup
|
||||||
|
2. Installing Dependencies
|
||||||
|
|
||||||
|
By the end, you'll have a working development environment ready for multi-vector search experiments.
|
||||||
@@ -0,0 +1,132 @@
|
|||||||
|
---
|
||||||
|
title: "Installing Dependencies"
|
||||||
|
description: Install Python dependencies including FastEmbed and Qdrant client.
|
||||||
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 0 {{< /date >}}
|
||||||
|
|
||||||
|
# Installing Dependencies
|
||||||
|
|
||||||
|
To work with multi-vector search in Qdrant, you'll need several Python libraries: Qdrant client for search and FastEmbed for multi-vector embeddings.
|
||||||
|
|
||||||
|
We'll set up a clean Python environment and install everything you need to start experimenting with multi-vector representations.
|
||||||
|
|
||||||
|
## Python Environment Setup
|
||||||
|
|
||||||
|
### Using uv (Recommended)
|
||||||
|
|
||||||
|
For this course, we recommend using [uv](https://docs.astral.sh/uv/), a modern Python package manager that's significantly faster and more reliable than traditional pip. It handles virtual environments and dependencies with better performance and dependency resolution.
|
||||||
|
|
||||||
|
**Install uv:**
|
||||||
|
|
||||||
|
On macOS and Linux:
|
||||||
|
```bash
|
||||||
|
curl -LsSf https://astral.sh/uv/install.sh | sh
|
||||||
|
```
|
||||||
|
|
||||||
|
On Windows:
|
||||||
|
```bash
|
||||||
|
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Create a new virtual environment:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
uv venv
|
||||||
|
source .venv/bin/activate # On macOS/Linux
|
||||||
|
# or
|
||||||
|
.venv\Scripts\activate # On Windows
|
||||||
|
```
|
||||||
|
|
||||||
|
### Alternative: Using Poetry
|
||||||
|
|
||||||
|
If you prefer Poetry for dependency management, it offers robust project management with automatic virtual environment handling and dependency lock files.
|
||||||
|
|
||||||
|
**Install Poetry:**
|
||||||
|
|
||||||
|
On macOS and Linux:
|
||||||
|
```bash
|
||||||
|
curl -sSL https://install.python-poetry.org | python3 -
|
||||||
|
```
|
||||||
|
|
||||||
|
On Windows (PowerShell):
|
||||||
|
```bash
|
||||||
|
(Invoke-WebRequest -Uri https://install.python-poetry.org -UseBasicParsing).Content | py -
|
||||||
|
```
|
||||||
|
|
||||||
|
**Create a new project or add dependencies:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Initialize a new Poetry project
|
||||||
|
poetry init
|
||||||
|
|
||||||
|
# Activate the virtual environment
|
||||||
|
poetry shell
|
||||||
|
```
|
||||||
|
|
||||||
|
**Python Version Requirements:**
|
||||||
|
You'll need Python 3.10 or higher, as required by the qdrant-client library.
|
||||||
|
|
||||||
|
## Installing Dependencies
|
||||||
|
|
||||||
|
With your virtual environment activated, install the required libraries:
|
||||||
|
|
||||||
|
**Using uv (recommended):**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
uv pip install "qdrant-client>=1.16.2" "fastembed>=0.8.0"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Using Poetry:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
poetry add "qdrant-client>=1.16.2" "fastembed>=0.8.0"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Using pip:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pip install "qdrant-client>=1.16.2" "fastembed>=0.8.0"
|
||||||
|
```
|
||||||
|
|
||||||
|
### What These Libraries Do
|
||||||
|
|
||||||
|
- **qdrant-client**: The official Python client for Qdrant, providing both synchronous and asynchronous APIs for vector search operations. This library contains full type definitions and supports all Qdrant features.
|
||||||
|
|
||||||
|
- **fastembed**: A fast, lightweight library for generating embeddings, maintained by the Qdrant team. It includes support for multi-vector embeddings which we'll use extensively in this course. **(Note: fastembed=0.7.5 or above required for this course)**
|
||||||
|
|
||||||
|
## Verification Steps
|
||||||
|
|
||||||
|
Let's verify that everything is installed correctly.
|
||||||
|
|
||||||
|
**Test your imports:**
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
from fastembed import TextEmbedding
|
||||||
|
|
||||||
|
print("All dependencies installed successfully!")
|
||||||
|
```
|
||||||
|
|
||||||
|
**Quick connection test:**
|
||||||
|
|
||||||
|
If you set up Qdrant in the previous lesson, verify you can connect:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# For Qdrant Cloud
|
||||||
|
client = QdrantClient(
|
||||||
|
url="https://your-cluster-url.cloud.qdrant.io",
|
||||||
|
api_key="your-api-key"
|
||||||
|
)
|
||||||
|
|
||||||
|
# For local Qdrant
|
||||||
|
# client = QdrantClient(url="http://localhost:6333")
|
||||||
|
|
||||||
|
print(f"Connected to Qdrant: {client.get_collections()}")
|
||||||
|
```
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
With your Python environment configured and dependencies installed, you're ready to dive into Module 1, where we'll explore the fundamentals of multi-vector search and understand how it differs from traditional single-vector approaches.
|
||||||
@@ -0,0 +1,95 @@
|
|||||||
|
---
|
||||||
|
title: "Qdrant Setup"
|
||||||
|
description: Set up Qdrant for multi-vector search. Learn how to create a collection and configure it for multi-vector embeddings.
|
||||||
|
weight: 1
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 0 {{< /date >}}
|
||||||
|
|
||||||
|
# Qdrant Setup
|
||||||
|
|
||||||
|
Before diving into multi-vector search, you need a running Qdrant instance. Whether you choose Qdrant Cloud for a managed solution or a local deployment, this lesson will get you up and running.
|
||||||
|
|
||||||
|
Multi-vector search requires specific collection configurations that differ from traditional single-vector setups. We'll cover the essentials to prepare your environment.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Qdrant Cloud Setup (Recommended)
|
||||||
|
|
||||||
|
Qdrant Cloud is the fastest way to get started with multi-vector search. It provides a fully managed, production-ready vector database with automatic backups, high availability, and secure TLS connections. Both Qdrant Cloud and the open-source version provide the same feature set - Cloud simply handles the infrastructure for you.
|
||||||
|
|
||||||
|
### Create Your Cluster
|
||||||
|
|
||||||
|
1. Sign up at [cloud.qdrant.io](https://cloud.qdrant.io/signup) using your email, Google, or GitHub account.
|
||||||
|
|
||||||
|
2. Navigate to **Clusters** -> **Create a Free Cluster**. The Free Tier provides sufficient resources for this course.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
3. Select a region closest to your location or application.
|
||||||
|
|
||||||
|
4. Once your cluster is ready, copy the API key from the cluster dashboard and store it securely. You can generate additional keys later from the **API Keys** section.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
### Access the Web UI
|
||||||
|
|
||||||
|
Click **Cluster UI** in the top-right corner of your cluster page to open the dashboard.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
The Web UI provides several useful tools:
|
||||||
|
- **Console**: Test REST API calls directly in your browser
|
||||||
|
- **Collections**: Manage all your collections and their configurations
|
||||||
|
- **Tutorial**: Interactive walkthrough with sample data
|
||||||
|
|
||||||
|
### Save Your Credentials
|
||||||
|
|
||||||
|
Store your cluster URL and API key for use in upcoming lessons. Create an `.env` file in your working directory:
|
||||||
|
|
||||||
|
```env
|
||||||
|
QDRANT_URL=https://YOUR-CLUSTER.cloud.qdrant.io:6333
|
||||||
|
QDRANT_API_KEY=YOUR_API_KEY
|
||||||
|
```
|
||||||
|
|
||||||
|
Replace `YOUR-CLUSTER` with your actual cluster URL from the dashboard, and `YOUR_API_KEY` with the API key you copied earlier.
|
||||||
|
|
||||||
|
You'll use these credentials in the next lesson when we install and configure the Python client.
|
||||||
|
|
||||||
|
## Local Qdrant Installation
|
||||||
|
|
||||||
|
Qdrant's open-source version provides the same features as Qdrant Cloud but requires you to manage the infrastructure yourself. This option works well for development, testing, or when you need full control over your deployment.
|
||||||
|
|
||||||
|
### Docker Installation (Recommended)
|
||||||
|
|
||||||
|
The fastest way to run Qdrant locally is with Docker:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run -p 6333:6333 -p 6334:6334 \
|
||||||
|
-v $(pwd)/qdrant_storage:/qdrant/storage:z \
|
||||||
|
qdrant/qdrant
|
||||||
|
```
|
||||||
|
|
||||||
|
This command:
|
||||||
|
- Exposes port `6333` for the REST API
|
||||||
|
- Exposes port `6334` for the gRPC API
|
||||||
|
- Mounts a local directory for persistent storage
|
||||||
|
|
||||||
|
Once running, you can access the Web UI at `http://localhost:6333/dashboard` to verify the installation.
|
||||||
|
|
||||||
|
### Alternative Installation Methods
|
||||||
|
|
||||||
|
For production deployments or other installation methods, see the [Qdrant Installation Guide](/documentation/guides/installation/).
|
||||||
|
|
||||||
|
## Verifying Your Setup
|
||||||
|
|
||||||
|
Open the Qdrant Web UI:
|
||||||
|
- **Cloud users**: Click **Cluster UI** in the top-right corner of your cluster dashboard
|
||||||
|
- **Local users**: Navigate to `http://localhost:6333/dashboard`
|
||||||
|
|
||||||
|
If the Web UI loads and you can see the **Collections** tab, your setup is complete. In the next lesson, you'll install the Python dependencies to connect programmatically.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
Next, you'll install the Python dependencies needed to work with multi-vector embeddings.
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
---
|
||||||
|
title: "Module 1: Multi-Vector Representations for Textual Data"
|
||||||
|
description: "Learn about multi-vector representations for text with ColBERT. Understand how they differ from single vector embeddings and when to use them."
|
||||||
|
isLesson: true
|
||||||
|
weight: 20
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 1 {{< /date >}}
|
||||||
|
|
||||||
|
# Multi-Vector Representations for Textual Data
|
||||||
|
|
||||||
|
Dive into multi-vector text representations and discover how ColBERT changes the vector search landscape.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Today's path
|
||||||
|
|
||||||
|
1. Late Interaction Basics
|
||||||
|
2. MaxSim Distance Metric
|
||||||
|
3. Use Cases for Multi-Vector Search
|
||||||
|
4. Problems of Multi-Vector Search
|
||||||
|
5. Multi-Vector Embeddings in Qdrant
|
||||||
|
|
||||||
|
You'll understand when multi-vector representations outperform traditional single-vector embeddings, and what kind of
|
||||||
|
problems to expect when you start working with multi-vector search at scale.
|
||||||
@@ -0,0 +1,222 @@
|
|||||||
|
---
|
||||||
|
title: "Late Interaction Basics"
|
||||||
|
description: Understand the late interaction paradigm and how it differs from traditional dense embeddings for text search.
|
||||||
|
weight: 1
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 1 {{< /date >}}
|
||||||
|
|
||||||
|
# Late Interaction Basics
|
||||||
|
|
||||||
|
When building a search system, one fundamental question emerges: **when should a query and document interact?** The answer to this question may affect both the quality of search results and the system's scalability.
|
||||||
|
|
||||||
|
This lesson introduces the late interaction paradigm - the foundation of multi-vector search - and explores how it compares to other approaches.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/8HrvD5o2w8M?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-1/late-interaction-basics.ipynb">
|
||||||
|
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||||
|
</a>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Understanding the Alternatives
|
||||||
|
|
||||||
|
Before diving into late interaction, let's establish what we mean by "interaction."
|
||||||
|
|
||||||
|
In search systems, **interaction** refers to when and how the query and document representations influence each other. Do they interact during encoding, or only during comparison? This timing fundamentally shapes the system's architecture.
|
||||||
|
|
||||||
|
We can categorize approaches based on when this interaction occurs:
|
||||||
|
|
||||||
|
- **No interaction:** Query and document are encoded independently into fixed representations, then compared. They never "see" each other during encoding.
|
||||||
|
- **Early interaction:** Query and document are encoded together, with each word attending to the other during the encoding process. Maximum interaction, but no pre-computation.
|
||||||
|
- **Late interaction:** Query and document are encoded independently (like no interaction), but we preserve fine-grained representations that interact during scoring (late in the process).
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Let's examine each paradigm to understand the trade-offs. A simple example will illustrate the differences between all the methods.
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Example documents and query we'll use throughout this lesson
|
||||||
|
documents = [
|
||||||
|
"Qdrant is an AI-native vector database and a semantic search engine",
|
||||||
|
"Relational databases are not well-suited for search",
|
||||||
|
]
|
||||||
|
query = "What is Qdrant?"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Single-Vector Embeddings (No Interaction)
|
||||||
|
|
||||||
|
The most common approach encodes each document and query into a single dense vector, then compares them using similarity (most often cosine similarity).
|
||||||
|
|
||||||
|
**The strength:** This method is simple, fast, and scales well. Document vectors can be pre-computed and stored, making search efficient even across billions of documents.
|
||||||
|
|
||||||
|
**The limitation:** Compressing an entire document into a single vector means losing fine-grained details. Think of it like summarizing a book in one sentence - you capture the general theme but miss the nuances that might be relevant to a specific query.
|
||||||
|
|
||||||
|
Let's load a dense embedding model and generate vector representations for our documents.
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import TextEmbedding
|
||||||
|
|
||||||
|
# Load the BAAI/bge-small-en-v1.5 model
|
||||||
|
dense_model = TextEmbedding("BAAI/bge-small-en-v1.5")
|
||||||
|
# Pass the documents through the model. The .passage_embed
|
||||||
|
# method returns a generator we can iterate over and is
|
||||||
|
# supposed to be used for the documents only.
|
||||||
|
dense_generator = dense_model.passage_embed(documents)
|
||||||
|
# Running next on the generator yields one vector at
|
||||||
|
# the time, representing a single document.
|
||||||
|
dense_vector = next(dense_generator)
|
||||||
|
```
|
||||||
|
|
||||||
|
We also need to generate a vector for the query using the same model.
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Generate a dense vector for the query as well, using
|
||||||
|
# the .query_embed method this time.
|
||||||
|
dense_query_vector = next(dense_model.query_embed(query))
|
||||||
|
```
|
||||||
|
|
||||||
|
Now we can calculate the similarity between the query and each document using the dot product.
|
||||||
|
|
||||||
|
```python
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
# Calculate the dot product between the query
|
||||||
|
# and the first document vector
|
||||||
|
np.dot(dense_query_vector, dense_vector)
|
||||||
|
```
|
||||||
|
|
||||||
|
Let's calculate the similarity with the second document as well.
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Calculate the dot product between the same query
|
||||||
|
# and the second document vectors
|
||||||
|
np.dot(dense_query_vector, next(dense_generator))
|
||||||
|
```
|
||||||
|
|
||||||
|
Notice how each document and query produces exactly **one vector** of 384 dimensions. To achieve this compression, the model internally generates embeddings for each token, then uses **pooling** (typically mean pooling or a special [CLS] token) to aggregate them into a single fixed-size representation.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
### Cross-Encoders (Early Interaction)
|
||||||
|
|
||||||
|
At the other extreme, cross-encoders process the query and document together through a neural network, producing a relevance score. This is "early interaction" because the query and document interact during the encoding phase itself.
|
||||||
|
|
||||||
|
**The strength:** This approach achieves deep contextual understanding. Every query word can "attend to" every document word during encoding, enabling precise relevance judgments.
|
||||||
|
|
||||||
|
**The limitation:** You must process every query-document pair from scratch. For a collection of a million documents, that means a million forward passes through a neural network for each query - prohibitively expensive for initial retrieval.
|
||||||
|
|
||||||
|
Cross-encoders excel at re-ranking a small candidate set but don't scale for searching large collections.
|
||||||
|
|
||||||
|
Let's see how cross-encoders work differently by processing the query and documents together.
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed.rerank.cross_encoder import TextCrossEncoder
|
||||||
|
|
||||||
|
# Load the Xenova/ms-marco-MiniLM-L-6-v2 cross encoder model
|
||||||
|
cross_encoder = TextCrossEncoder("Xenova/ms-marco-MiniLM-L-6-v2")
|
||||||
|
# Run .rerank method on the query and all the documents.
|
||||||
|
# It does not create any vector representations, but gives
|
||||||
|
# the score indicating the relevance of the document for
|
||||||
|
# the provided query.
|
||||||
|
cross_encoder.rerank(query, documents)
|
||||||
|
```
|
||||||
|
|
||||||
|
The key difference: cross-encoders **cannot pre-compute** document representations. Each query requires fresh computation for every candidate, making them impractical for initial retrieval over large collections.
|
||||||
|
|
||||||
|
### The Gap
|
||||||
|
|
||||||
|
We need an approach that combines the best of both worlds: the scalability of pre-computed single-vector representations and the fine-grained matching capability of cross-encoders.
|
||||||
|
|
||||||
|
## Late Interaction: The Core Paradigm
|
||||||
|
|
||||||
|
Late interaction solves this challenge through a simple but powerful idea: **encode documents and queries into multiple token-level vectors, then defer the comparison until search time.**
|
||||||
|
|
||||||
|
### How It Works
|
||||||
|
|
||||||
|
Instead of compressing a document into a single vector, late interaction represents it as a collection of contextualized token embeddings
|
||||||
|
|
||||||
|
1. **Encode:** Pass the document through an encoder (like ColBERT) to generate one embedding vector per token
|
||||||
|
2. **Store multi-vector representations:** Keep all token vectors instead of aggregating them into a single vector
|
||||||
|
3. **Defer comparison:** At search time, compare query token vectors against document token vectors
|
||||||
|
4. **Late interaction:** The actual "interaction" between query and document happens late - only when computing relevance scores
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
This isn't that much fundamentally different from both single-vector and cross-encoder approaches, yet there are some differences:
|
||||||
|
- Unlike single-vector: We preserve fine-grained, token-level information
|
||||||
|
- Unlike cross-encoders: We encode documents independently, enabling pre-computation
|
||||||
|
|
||||||
|
### Key Benefits
|
||||||
|
|
||||||
|
**Pre-computation:** Document embeddings can be computed once and stored, just like single-vector approaches. You don't need to re-encode documents for every query.
|
||||||
|
|
||||||
|
**Fine-grained matching:** Different query terms can match different parts of the document. A query about "apple computer" can distinguish contextual meaning from "apple fruit" based on which document tokens match strongly.
|
||||||
|
|
||||||
|
**Contextual understanding:** Token embeddings are contextualized by the surrounding text. The word "bank" has different embeddings in "river bank" versus "financial bank."
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
**Scalability:** While requiring more storage than single vectors, the deferred comparison enables searching large collections efficiently.
|
||||||
|
|
||||||
|
### The ColBERT Approach
|
||||||
|
|
||||||
|
The canonical implementation of late interaction is **ColBERT** (Contextualized Late Interaction over BERT), developed at Stanford. ColBERT popularized this paradigm and demonstrated that you can achieve cross-encoder-level effectiveness with single-vector-level efficiency.
|
||||||
|
|
||||||
|
The core innovation: maintaining bags of contextualized embeddings and delaying the interaction computation until the final stage.
|
||||||
|
|
||||||
|
Let's load a ColBERT model and generate multi-vector representations for our documents.
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import LateInteractionTextEmbedding
|
||||||
|
|
||||||
|
# Load the colbert-ir/colbertv2.0 model
|
||||||
|
colbert_model = LateInteractionTextEmbedding("colbert-ir/colbertv2.0")
|
||||||
|
# Run .passage_embed on all the documents and create
|
||||||
|
# a generator of the multi-vector representations
|
||||||
|
colbert_generator = colbert_model.passage_embed(documents)
|
||||||
|
colbert_vector = next(colbert_generator)
|
||||||
|
```
|
||||||
|
|
||||||
|
Similarly, we create a multi-vector representation for the query.
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Create multi-vector representation for the query
|
||||||
|
colbert_query_vector = next(colbert_model.query_embed(query))
|
||||||
|
```
|
||||||
|
|
||||||
|
**Key observation:** Unlike single-vector search, each document is represented by **multiple vectors**. At search time, we compare each query token against all document tokens to compute a relevance score.
|
||||||
|
|
||||||
|
## Why This Matters for Multi-Vector Search
|
||||||
|
|
||||||
|
Late interaction isn't just a technical optimization - it represents a fundamental shift in how we think about semantic search.
|
||||||
|
|
||||||
|
**Captures semantic nuance:** Because we maintain multiple vectors per document, the system can capture complex, multi-faceted content. A document about "Python programming for data science" can match queries about programming languages, data analysis, and scientific computing - each matching different token sets.
|
||||||
|
|
||||||
|
**Enables scale:** Pre-computed multi-vector representations mean you can build practical search systems over large document collections. The computational cost grows with collection size, not quadratically with query-document pairs.
|
||||||
|
|
||||||
|
**Foundation for this course:** Everything we'll explore in subsequent lessons builds on this paradigm - from the distance metrics that enable multi-vector comparison to multi-modal extensions like ColPali to optimization techniques for production deployment.
|
||||||
|
|
||||||
|
**Beyond text:** The late interaction paradigm extends naturally to other modalities. Module 2 explores how ColPali applies these same principles to visual documents, enabling semantic search over images and PDFs.
|
||||||
|
|
||||||
|
## What's Next
|
||||||
|
|
||||||
|
Understanding the conceptual foundation of late interaction is the first step. But how exactly do we compare sets of query vectors against sets of document vectors?
|
||||||
|
|
||||||
|
In the next lesson, you'll learn about **MaxSim** - the distance metric that powers late interaction search. MaxSim defines the specific mathematical operation for computing similarity between multi-vector representations.
|
||||||
|
|
||||||
|
From there, we'll explore use cases where multi-vector search excels, challenges you'll face in production, and how to implement these techniques in Qdrant.
|
||||||
@@ -0,0 +1,194 @@
|
|||||||
|
---
|
||||||
|
title: "MaxSim Distance Metric"
|
||||||
|
description: Learn about the MaxSim distance metric used in multi-vector search and how it computes similarity between multi-vector representations.
|
||||||
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 1 {{< /date >}}
|
||||||
|
|
||||||
|
# MaxSim Distance Metric
|
||||||
|
|
||||||
|
MaxSim (Maximum Similarity) is the core distance metric for late interaction models. Unlike traditional vector similarity metrics that operate on pairs of single vectors, MaxSim computes similarity between sequences of vectors.
|
||||||
|
|
||||||
|
Understanding MaxSim is important for working with multi-vector search effectively and understanding its performance characteristics.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/JvSvuK19m8A?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-1/maxsim-distance.ipynb">
|
||||||
|
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||||
|
</a>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The MaxSim Formula
|
||||||
|
|
||||||
|
In late interaction, we represent documents and queries as sequences of token vectors. But how do we measure similarity between two sets of vectors?
|
||||||
|
|
||||||
|
The answer is **MaxSim** (Maximum Similarity), defined mathematically as:
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{MaxSim}(Q, D) = \sum_{i=1}^{|Q|} \max_{j=1}^{|D|} \text{sim}(q_i, d_j)
|
||||||
|
$$
|
||||||
|
|
||||||
|
Where:
|
||||||
|
- $Q$ represents the query token vectors
|
||||||
|
- $D$ represents the document token vectors
|
||||||
|
- $\text{sim}(q_i, d_j)$ is a base similarity function that measures the distance between two vectors
|
||||||
|
|
||||||
|
The $\text{sim}(q_i, d_j)$ function can be any distance metric - dot product, cosine similarity, Euclidean distance, or others. Let's break down what this formula actually computes.
|
||||||
|
|
||||||
|
## Understanding MaxSim Step-by-Step
|
||||||
|
|
||||||
|
### The Computation Process
|
||||||
|
|
||||||
|
Let's make this concrete with an example. Consider:
|
||||||
|
|
||||||
|
- **Query**: "apple computer"
|
||||||
|
- **Document**: "Apple makes the MacBook laptop"
|
||||||
|
|
||||||
|
MaxSim formula will choose the strongest connection from the query tokens to document tokens and sum the strengths of each connection for all the query tokens.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
The key insight here: **each query token finds its best match in the document**. This enables **fine-grained semantic matching** - instead of comparing holistic document representations, we're allowing each part of the query to independently find its most relevant counterpart.
|
||||||
|
|
||||||
|
Let's walk through this with a concrete example:
|
||||||
|
|
||||||
|
```python
|
||||||
|
query = "apple computer"
|
||||||
|
document = "Apple makes the MacBook laptop"
|
||||||
|
```
|
||||||
|
|
||||||
|
Next, we load the ColBERT model to generate multi-vector representations for both the query and document:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import LateInteractionTextEmbedding
|
||||||
|
|
||||||
|
# Load the colbert-ir/colbertv2.0 model
|
||||||
|
colbert_model = LateInteractionTextEmbedding("colbert-ir/colbertv2.0")
|
||||||
|
|
||||||
|
# Create multi-vector representations of the query and document
|
||||||
|
query_vector = next(colbert_model.query_embed(query))
|
||||||
|
document_vector = next(colbert_model.passage_embed(document))
|
||||||
|
```
|
||||||
|
|
||||||
|
### Understanding Tokenization
|
||||||
|
|
||||||
|
Before we compute MaxSim, let's understand how ColBERT tokenizes our query and document. This tokenization is crucial because each token gets its own embedding vector.
|
||||||
|
|
||||||
|
```python
|
||||||
|
query_tokenization = colbert_model.model.tokenize([query])[0]
|
||||||
|
query_tokenization.tokens
|
||||||
|
```
|
||||||
|
|
||||||
|
This shows how the query tokenizes into individual tokens including special tokens like `[CLS]` and `[SEP]`.
|
||||||
|
|
||||||
|
```python
|
||||||
|
document_tokenization = colbert_model.model.tokenize([document])[0]
|
||||||
|
document_tokenization.tokens
|
||||||
|
```
|
||||||
|
|
||||||
|
Similarly, the document is tokenized. Notice that words like 'MacBook' may be split into subword tokens using WordPiece tokenization (e.g., 'mac' and '##book'). Each token will get its own embedding vector.
|
||||||
|
|
||||||
|
Now let's compute MaxSim step-by-step. For each query token, we'll find its maximum similarity score across all document tokens, then sum these maxima:
|
||||||
|
|
||||||
|
```python
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
similarity = 0.0
|
||||||
|
for qt, qt_vector in zip(query_tokenization.tokens,
|
||||||
|
query_vector):
|
||||||
|
max_idx, max_sim = 0, np.dot(qt_vector, document_vector[0])
|
||||||
|
for i, dt_vector in enumerate(document_vector[1:], start=1):
|
||||||
|
distance = np.dot(qt_vector, dt_vector)
|
||||||
|
if distance > max_sim:
|
||||||
|
max_idx, max_sim = i, distance
|
||||||
|
|
||||||
|
print(qt, max_idx, max_sim)
|
||||||
|
similarity += max_sim
|
||||||
|
```
|
||||||
|
|
||||||
|
The code iterates through each query token, computes its dot product similarity with every document token, and keeps track of the maximum. The `print` statement shows which document position each query token matched best with and the similarity score. Each query token independently seeks its strongest counterpart in the document.
|
||||||
|
|
||||||
|
```python
|
||||||
|
print("MaxSim(Q, D) =", similarity)
|
||||||
|
```
|
||||||
|
|
||||||
|
The final MaxSim(Q, D) score is the sum of all these maximum similarities.
|
||||||
|
|
||||||
|
### The Intuition Behind MaxSim
|
||||||
|
|
||||||
|
MaxSim implements **token-level relevance matching**. Each query term actively seeks its most relevant counterpart in the document. Unlike single-vector search where you compare one query vector to one document vector (one-to-one), MaxSim performs many-to-many matching. This enables **contextual precision**. When you search for "apple computer", the word "apple" in your query will match strongly with "Apple" (the company name) in documents about technology, not "apple" (the fruit) in documents about nutrition, as contextualized token embeddings should capture these semantic distinctions.
|
||||||
|
|
||||||
|
But why maximum and not average? Consider a query about "Python programming". A document might discuss Python extensively in one section but also mention cooking recipes in another. MaxSim focuses on the **strong matches** (Python-related tokens) rather than diluting the score with irrelevant tokens. This design choice reflects a fundamental insight: relevance is often concentrated, not uniform.
|
||||||
|
|
||||||
|
## The HNSW Challenge
|
||||||
|
|
||||||
|
There's a significant challenge related to MaxSim when it comes to indexing these representations efficiently. HNSW (Hierarchical Navigable Small World) graphs enable fast approximate nearest neighbor search by building static proximity graphs. They rely on the assumption that distance functions must be symmetric and query-independent. This allows HNSW to construct fixed neighbor relationships that makes the effective graph traversal possible.
|
||||||
|
|
||||||
|
**MaxSim breaks this assumption by design.** Looking back at the formula, notice that $Q$ and $D$ play fundamentally different roles:
|
||||||
|
|
||||||
|
- We **iterate over query tokens** - each query token contributes to the final sum
|
||||||
|
- Documents are **what we search within** - we take the maximum similarity from document tokens for each query token
|
||||||
|
|
||||||
|
This non-symmetrical structure means `MaxSim(Q, D) ≠ MaxSim(D, Q)`. When you swap the parameters, you change which tokens contribute to the sum.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
To illustrate this asymmetry concretely, let's compute MaxSim in reverse - iterating over document tokens instead of query tokens:
|
||||||
|
|
||||||
|
```python
|
||||||
|
similarity = 0.0
|
||||||
|
for dt, dt_vector in zip(document_tokenization.tokens,
|
||||||
|
document_vector):
|
||||||
|
max_idx, max_sim = 0, np.dot(dt_vector, query_vector[0])
|
||||||
|
for i, qt_vector in enumerate(query_vector[1:], start=1):
|
||||||
|
distance = np.dot(dt_vector, qt_vector)
|
||||||
|
if distance > max_sim:
|
||||||
|
max_idx, max_sim = i, distance
|
||||||
|
|
||||||
|
print(dt, max_idx, max_sim)
|
||||||
|
similarity += max_sim
|
||||||
|
```
|
||||||
|
|
||||||
|
Now the iteration is different: we loop over document tokens instead of query tokens. Since the document has more tokens than the query, we're summing over more terms, which fundamentally changes the computation.
|
||||||
|
|
||||||
|
```python
|
||||||
|
print("MaxSim(D, Q) =", similarity)
|
||||||
|
```
|
||||||
|
|
||||||
|
The `MaxSim(D, Q)` score will be different from `MaxSim(Q, D)` because we iterated over a different number of tokens. This asymmetry occurs because:
|
||||||
|
- We summed over document tokens instead of query tokens (different number of terms)
|
||||||
|
- The iteration direction changed the fundamental computation
|
||||||
|
- Query and document play fundamentally different roles in the formula
|
||||||
|
|
||||||
|
This non-symmetry is why HNSW indexing becomes problematic for MaxSim - nearest neighbor relationships change depending on which direction you query.
|
||||||
|
|
||||||
|
A document's nearest neighbors are query-dependent and change based on which query you're processing, making it impossible to build the static proximity graph that HNSW requires. The practical reality is you must compute MaxSim against every document at query time with brute force comparison. For large collections, this becomes slow, or even impossible.
|
||||||
|
|
||||||
|
The common solution is **two-stage retrieval**, and we'll explore that pattern in later lessons and see how Qdrant optimizes multi-vector search.
|
||||||
|
|
||||||
|
## Connecting the Concepts
|
||||||
|
|
||||||
|
In the previous lesson, you learned about the late interaction paradigm - encoding queries and documents into multiple token vectors and deferring comparison until search time. MaxSim is the mathematical operation that makes this deferred comparison work. It's the bridge between the conceptual model (multi-vector representations) and practical implementation. Without MaxSim or a similar aggregation function, we'd have no way to score documents against queries when both are represented as sequences of vectors.
|
||||||
|
|
||||||
|
## What's Next
|
||||||
|
|
||||||
|
MaxSim gives us a way to compare multi-vector representations, capturing fine-grained semantic matching that single-vector search cannot achieve. But as we've seen, it also introduces significant computational challenges - particularly the asymmetry that breaks traditional indexing approaches like HNSW.
|
||||||
|
|
||||||
|
In the next lesson, we'll explore **use cases where multi-vector search excels** despite these challenges. You'll see scenarios where the improved relevance and semantic precision justify the additional computational cost.
|
||||||
|
|
||||||
|
From there, we'll learn how to implement efficient multi-vector search in Qdrant, including the hybrid retrieval patterns that make late interaction practical at scale.
|
||||||
@@ -0,0 +1,188 @@
|
|||||||
|
---
|
||||||
|
title: "Multi-Vector Embeddings in Qdrant"
|
||||||
|
description: Configure Qdrant collections for multi-vector embeddings and learn how to index and query multi-vector data.
|
||||||
|
weight: 5
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 1 {{< /date >}}
|
||||||
|
|
||||||
|
# Multi-Vector Embeddings in Qdrant
|
||||||
|
|
||||||
|
You've learned how MaxSim enables fine-grained token-level matching and explored both the benefits and challenges of multi-vector search. Now it's time to put that knowledge into practice.
|
||||||
|
|
||||||
|
Qdrant provides first-class support for multi-vector embeddings, making it straightforward to build search systems that leverage late interaction. In this lesson, you'll learn how to configure Qdrant collections for multi-vector search, index documents with token-level embeddings, and execute queries using MaxSim distance.
|
||||||
|
|
||||||
|
By the end, you'll understand the key configuration parameters, know when to use storage optimization strategies, and be ready to build your own multi-vector search applications.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/THZE2O4kMDg?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-1/multi-vector-in-qdrant.ipynb">
|
||||||
|
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||||
|
</a>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Creating a Multi-Vector Collection
|
||||||
|
|
||||||
|
Setting up a collection for multi-vector search requires specific configuration to enable late interaction and MaxSim distance calculation. Here's how to create a collection configured for ColBERT embeddings:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
|
client = QdrantClient("http://localhost:6333")
|
||||||
|
|
||||||
|
client.create_collection(
|
||||||
|
collection_name="colbert-search",
|
||||||
|
vectors_config={
|
||||||
|
"colbert": models.VectorParams(
|
||||||
|
# Size of an individual token vector
|
||||||
|
size=128,
|
||||||
|
# Distance function for token similarity
|
||||||
|
distance=models.Distance.DOT,
|
||||||
|
# Enable multi-vector mode with MaxSim
|
||||||
|
multivector_config=models.MultiVectorConfig(
|
||||||
|
comparator=models.MultiVectorComparator.MAX_SIM,
|
||||||
|
),
|
||||||
|
# Disable HNSW indexing
|
||||||
|
hnsw_config=models.HnswConfigDiff(m=0),
|
||||||
|
),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
Let's break down each parameter:
|
||||||
|
|
||||||
|
- **`size`**: The dimensionality of each individual token embedding.
|
||||||
|
- **`distance`**: The similarity metric used to compare individual token vectors. This is the per-token comparison, and MaxSim will aggregate these scores.
|
||||||
|
- **`multivector_config`**: Enables multi-vector mode for this collection. The `comparator=MultiVectorComparator.MAX_SIM` parameter tells Qdrant to use MaxSim distance when comparing multi-vector documents to queries.
|
||||||
|
- **`hnsw_config`**: Setting `m=0` disables HNSW indexing. HNSW graphs don't work with MaxSim either way, so it should be disabled to not create an index we won't use.
|
||||||
|
|
||||||
|
You now have a collection ready for multi-vector search. The configuration handles MaxSim automatically - you just need to provide the embeddings.
|
||||||
|
|
||||||
|
## Indexing Multi-Vector Documents
|
||||||
|
|
||||||
|
With your collection configured, you can now index documents. Each document needs multiple token embeddings - one for each token in the text:
|
||||||
|
|
||||||
|
```python
|
||||||
|
import uuid
|
||||||
|
|
||||||
|
documents = [
|
||||||
|
# Document A: Highly relevant - addresses all query aspects
|
||||||
|
"When async tasks fail to return database connections, the pool "
|
||||||
|
"becomes exhausted and requests start failing. Ensuring "
|
||||||
|
"connections are closed after awaits prevents this.",
|
||||||
|
|
||||||
|
# Document B: Partially relevant - mentions some concepts
|
||||||
|
"Database resource exhaustion can occur due to limited pool sizes.",
|
||||||
|
|
||||||
|
# Document C: Keyword-stuffed - contains related terms without substance
|
||||||
|
"Understanding concurrency, async IO, and database performance in "
|
||||||
|
"Python web applications.",
|
||||||
|
|
||||||
|
# Document D: Completely irrelevant
|
||||||
|
"Handling training for pythons should be done gradually, starting "
|
||||||
|
"with short sessions and increasing duration as the snake becomes "
|
||||||
|
"more comfortable.",
|
||||||
|
]
|
||||||
|
|
||||||
|
client.upsert(
|
||||||
|
collection_name="colbert-search",
|
||||||
|
points=[
|
||||||
|
models.PointStruct(
|
||||||
|
id=uuid.uuid4().hex,
|
||||||
|
vector={
|
||||||
|
"colbert": models.Document(
|
||||||
|
text=doc,
|
||||||
|
model="colbert-ir/colbertv2.0",
|
||||||
|
)
|
||||||
|
},
|
||||||
|
payload={
|
||||||
|
"text": doc,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
for doc in documents
|
||||||
|
]
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
This example uses `models.Document` for convenience - Qdrant's **local inference** feature powered by FastEmbed integration. You provide the text and model name, and Qdrant handles tokenization and encoding automatically. You can also generate embeddings externally and provide them directly as lists of vectors.
|
||||||
|
|
||||||
|
The `payload` field stores the original text and metadata, returned with search results.
|
||||||
|
|
||||||
|
## Querying with MaxSim
|
||||||
|
|
||||||
|
Querying works the same way - provide your query text and let Qdrant handle MaxSim computation:
|
||||||
|
|
||||||
|
```python
|
||||||
|
query = "How can I prevent Python database connection pool exhaustion?"
|
||||||
|
|
||||||
|
results = client.query_points(
|
||||||
|
collection_name="colbert-search",
|
||||||
|
query=models.Document(
|
||||||
|
text=query,
|
||||||
|
model="colbert-ir/colbertv2.0",
|
||||||
|
),
|
||||||
|
using="colbert",
|
||||||
|
limit=2,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
The results show which documents have the highest token-level semantic overlap with your query. Documents about async connection pool exhaustion rank higher than generic database documents because more query tokens find strong matches.
|
||||||
|
|
||||||
|
## Storage Optimization: Offloading to Disk
|
||||||
|
|
||||||
|
As discussed in the previous lesson, multi-vector search requires significantly more memory than traditional single-vector search. A document with 500 tokens stores 500 separate embeddings, creating substantial memory pressure for large collections.
|
||||||
|
|
||||||
|
When you prioritize maximum precision over low latency, Qdrant offers a solution: **offload vectors to disk**. This strategy keeps embeddings on disk rather than loading them into RAM, dramatically reducing memory requirements at the cost of slower query performance:
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.create_collection(
|
||||||
|
collection_name="colbert-search-on-disk",
|
||||||
|
vectors_config={
|
||||||
|
"colbert": models.VectorParams(
|
||||||
|
size=128,
|
||||||
|
distance=models.Distance.DOT,
|
||||||
|
multivector_config=models.MultiVectorConfig(
|
||||||
|
comparator=models.MultiVectorComparator.MAX_SIM,
|
||||||
|
),
|
||||||
|
hnsw_config=models.HnswConfigDiff(m=0),
|
||||||
|
# Offload vectors to disk
|
||||||
|
on_disk=True,
|
||||||
|
),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
The single `on_disk=True` parameter changes the storage strategy. Here's the trade-off:
|
||||||
|
|
||||||
|
**Benefits**: Dramatically reduced memory footprint. Instead of storing thousands of token embeddings in RAM, Qdrant reads them from disk during search. This enables multi-vector search on collections that would otherwise exceed available memory.
|
||||||
|
|
||||||
|
**Costs**: Higher query latency due to disk I/O. Every MaxSim calculation requires reading document embeddings from disk, which is orders of magnitude slower than RAM access. Query times increase from milliseconds to potentially seconds, depending on collection size and hardware.
|
||||||
|
|
||||||
|
**When to use it**: Disk offloading works best when precision is critical but latency constraints are relaxed. Research applications, offline batch processing, or scenarios where you're willing to wait a few seconds for the most accurate results all benefit from this approach. For production systems requiring sub-second response times, you'll need other optimization strategies.
|
||||||
|
|
||||||
|
This is just one optimization technique. Module 3 covers additional approaches including quantization, pooling, and multi-stage retrieval that balance precision, memory, and latency more effectively for production deployments.
|
||||||
|
|
||||||
|
## What's Next
|
||||||
|
|
||||||
|
You've learned how to configure Qdrant for multi-vector search, from creating collections with MaxSim comparators to indexing documents and executing queries. The key takeaways:
|
||||||
|
|
||||||
|
- **Multi-vector collections** require specific configuration: `multivector_config` with `MAX_SIM` comparator
|
||||||
|
- **HNSW indexing is disabled** because MaxSim doesn't work with static proximity graphs
|
||||||
|
- **Disk offloading** (`on_disk=True`) reduces memory usage when precision matters more than latency
|
||||||
|
- **Local inference** with FastEmbed integration lets you provide text directly instead of pre-computed embeddings
|
||||||
|
|
||||||
|
The examples in this lesson focused on text search using ColBERT. But late interaction and MaxSim aren't limited to text. In Module 2, you'll discover how multi-vector embeddings extend to multi-modal data - searching PDFs and images using visual token embeddings with ColPali.
|
||||||
@@ -0,0 +1,64 @@
|
|||||||
|
---
|
||||||
|
title: "Problems of Multi-Vector Search"
|
||||||
|
description: Understand the challenges and limitations of multi-vector search at scale, including memory and performance considerations.
|
||||||
|
weight: 4
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 1 {{< /date >}}
|
||||||
|
|
||||||
|
# Problems of Multi-Vector Search
|
||||||
|
|
||||||
|
Multi-vector search delivers impressive retrieval quality, but it comes with significant challenges. Before deploying multi-vector search in production, you need to understand these limitations and plan accordingly.
|
||||||
|
|
||||||
|
The good news: Module 3 covers optimization techniques that address many of these challenges.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/vQKF1hO7hzU?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The Indexing Challenge: Why HNSW Doesn't Work
|
||||||
|
|
||||||
|
One of the fundamental challenges with multi-vector search stems from **HNSW indexing incompatibility**. As you learned in the MaxSim lesson, traditional vector search relies on HNSW (Hierarchical Navigable Small World) graphs to enable fast approximate nearest neighbor search. HNSW works by building static proximity graphs that connect similar documents, allowing efficient traversal during queries.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
However, HNSW requires distance functions to be **symmetric and query-independent**. MaxSim breaks both assumptions by design. Remember that `MaxSim(Q, D) ≠ MaxSim(D, Q)` - when you swap the parameters, you iterate over different token sets, fundamentally changing the computation. A query with 3 tokens produces a different score than iterating over a document's 200 tokens.
|
||||||
|
|
||||||
|
This asymmetry creates **query-dependent nearest neighbor relationships**. A document's closest neighbors change based on which query you're processing, making it impossible to build the static proximity graph that HNSW requires. The practical consequence: **you must compute MaxSim against every document at query time using brute force comparison**. For collections with millions of documents, this becomes prohibitively slow without optimization strategies.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
## The Resource Challenge: Memory and Computational Overhead
|
||||||
|
|
||||||
|
Beyond indexing challenges, multi-vector search demands significantly more resources than traditional single-vector approaches. You've seen the precision benefits in the previous lesson on use cases - now let's understand the costs.
|
||||||
|
|
||||||
|
The overhead manifests in three ways: **storage, memory, and computation**. While single-vector search represents each document with one embedding (typically at least 384 dimensions), late interaction models like ColBERT use smaller per-token dimensions - around **128 dimensions per token** - but maintain separate embeddings for every token.
|
||||||
|
|
||||||
|
Even with these smaller dimensions, the total storage explodes. Consider a technical article with 500 tokens: you're storing 500 × 128 = 64,000 floats versus just 384 floats for a single dense embedding - roughly **167x more storage** for that document. The lower dimensionality per token doesn't compensate for the sheer number of vectors.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Each query-document pair requires computing multiple dot products (query tokens × document tokens), creating substantial computational overhead. A query with 10 tokens against a document with 200 tokens means 2,000 dot product operations instead of a single comparison.
|
||||||
|
|
||||||
|
This **memory overhead** combined with brute-force scanning creates fundamental scalability limits. While multi-vector search may work acceptably for collections with thousands or even millions of documents, **brute-force MaxSim doesn't scale indefinitely**. At billion-scale collections, computing MaxSim against every document becomes prohibitively expensive - both in terms of memory required to keep all vectors and latency from scanning the entire dataset. Whether you can deploy multi-vector search in production depends on your collection size, infrastructure budget, and latency requirements. For some applications with manageable collection sizes, the precision gains justify the resource investment. For billion-scale systems, the optimization techniques covered in Module 3 become essential rather than optional.
|
||||||
|
|
||||||
|
## These Challenges Are Solvable
|
||||||
|
|
||||||
|
While these limitations are real, **solutions exist**. The precision benefits of multi-vector search sometimes justify the costs for applications that demand fine-grained semantic matching. Understanding these trade-offs helps you make informed decisions about when and how to deploy multi-vector search in production.
|
||||||
|
|
||||||
|
## What's Next
|
||||||
|
|
||||||
|
Understanding these fundamental limitations prepares you for practical implementation. In the next lesson, you'll learn how to use **Qdrant multi-vector search**, including collection configuration and storage optimization strategies.
|
||||||
|
|
||||||
|
From theory to practice - let's see how Qdrant makes multi-vector search work.
|
||||||
@@ -0,0 +1,200 @@
|
|||||||
|
---
|
||||||
|
title: "Use Cases for Multi-Vector Search"
|
||||||
|
description: Discover scenarios where multi-vector search outperforms single-vector embeddings and provides better retrieval quality.
|
||||||
|
weight: 3
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 1 {{< /date >}}
|
||||||
|
|
||||||
|
# Use Cases for Multi-Vector Search
|
||||||
|
|
||||||
|
**When is the added complexity of multi-vector search actually worth it?** Multi-vector representations require more storage, more computation, and more careful implementation than simple single-vector embeddings. So why bother?
|
||||||
|
|
||||||
|
The answer comes down to one core capability: **fine-grained matching**. In the previous lessons, you learned how late interaction preserves token-level representations and how MaxSim computes similarity through independent token matching. Now you'll see when this precision actually matters - and when it doesn't.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/35WO2_Q_0C0?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-1/use-cases-multi-vector.ipynb">
|
||||||
|
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||||
|
</a>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The Power of Fine-Grained Matching
|
||||||
|
|
||||||
|
Single-vector embeddings compress entire documents and queries into single points in embedding space. This compression is a form of **lossy averaging** - all the nuanced details, specific requirements, and fine-grained semantics get blended into one representative vector. For many tasks, this works fine. But it fundamentally loses information.
|
||||||
|
|
||||||
|
**Multi-vector search preserves token-level information.** Instead of averaging "Python async database connection pooling" into one vector that represents the general topic, it maintains separate representations for each concept. As you learned in the MaxSim lesson, each query token independently finds its best match in the document, and these matches are summed, not averaged, to produce the final score.
|
||||||
|
|
||||||
|
This distinction is crucial. With single-vector search, a document that mentions Python and databases will likely get a moderate similarity score, even if it never discusses the specific combination you're looking for. With multi-vector search, **every query token must find a strong match** for the overall score to be high. This is token-level verification, not just topical matching.
|
||||||
|
|
||||||
|
## A Concrete Demonstration
|
||||||
|
|
||||||
|
Let's move from theory to practice with a real-world example that shows exactly when multi-vector search makes a difference.
|
||||||
|
|
||||||
|
### The Scenario: A Technical Support Query
|
||||||
|
|
||||||
|
Consider a developer searching for a specific technical solution:
|
||||||
|
|
||||||
|
```python
|
||||||
|
query = "How can I prevent Python database connection " \
|
||||||
|
"pool exhaustion in async web applications?"
|
||||||
|
```
|
||||||
|
|
||||||
|
This query has multiple specific requirements: Python, database connections, connection pooling, exhaustion problems, and async web applications. A truly relevant document should address all of these aspects together.
|
||||||
|
|
||||||
|
Now consider four documents with varying relevance:
|
||||||
|
|
||||||
|
```python
|
||||||
|
documents = [
|
||||||
|
# Document A: Highly relevant - addresses all query aspects
|
||||||
|
"When async tasks fail to return database connections, the pool "
|
||||||
|
"becomes exhausted and requests start failing. Ensuring "
|
||||||
|
"connections are closed after awaits prevents this.",
|
||||||
|
|
||||||
|
# Document B: Partially relevant - mentions some concepts
|
||||||
|
"Database resource exhaustion can occur due to limited pool sizes.",
|
||||||
|
|
||||||
|
# Document C: Keyword-stuffed - contains related terms without substance
|
||||||
|
"Understanding concurrency, async IO, and database performance in "
|
||||||
|
"Python web applications.",
|
||||||
|
|
||||||
|
# Document D: Completely irrelevant
|
||||||
|
"Handling training for pythons should be done gradually, starting "
|
||||||
|
"with short sessions and increasing duration as the snake becomes "
|
||||||
|
"more comfortable.",
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
**What should happen?** Document A should rank highest - it's the only one that directly addresses the problem. Document B is partially relevant. Document C sounds relevant (it mentions Python, async, database, web applications) but provides no substantive answer. Document D is obviously irrelevant (despite containing "python").
|
||||||
|
|
||||||
|
Let's see how single-vector and multi-vector approaches handle this.
|
||||||
|
|
||||||
|
### Single-Vector Embeddings: Missing the Details
|
||||||
|
|
||||||
|
First, let's try the traditional approach using a single dense vector per document:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import TextEmbedding
|
||||||
|
|
||||||
|
# Load the BAAI/bge-small-en-v1.5 model
|
||||||
|
dense_model = TextEmbedding("BAAI/bge-small-en-v1.5")
|
||||||
|
```
|
||||||
|
|
||||||
|
Encode the query into a single 384-dimensional vector:
|
||||||
|
|
||||||
|
```python
|
||||||
|
dense_query_vector = next(dense_model.query_embed(query))
|
||||||
|
dense_query_vector.shape
|
||||||
|
```
|
||||||
|
|
||||||
|
Encode all documents into single vectors:
|
||||||
|
|
||||||
|
```python
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
dense_vectors = np.array(list(dense_model.passage_embed(documents)))
|
||||||
|
dense_vectors.shape
|
||||||
|
```
|
||||||
|
|
||||||
|
Compute similarity scores using dot product:
|
||||||
|
|
||||||
|
```python
|
||||||
|
np.dot(dense_query_vector, dense_vectors.T)
|
||||||
|
```
|
||||||
|
|
||||||
|
**What happened?** The single-vector approach assigns very similar scores to Document A (highly relevant) and Document C (keyword-stuffed). Document C ranks nearly as high as Document A, even though it provides no actual solution!
|
||||||
|
|
||||||
|
The problem: by compressing all information into a single vector, the model captures **topical similarity** but misses whether the document actually addresses the specific requirements. Document C mentions the right topics (Python, async, database, web applications) but doesn't connect them meaningfully.
|
||||||
|
|
||||||
|
### Multi-Vector with ColBERT: Token-Level Verification
|
||||||
|
|
||||||
|
Now let's use ColBERT's multi-vector approach:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import LateInteractionTextEmbedding
|
||||||
|
|
||||||
|
# Load the colbert-ir/colbertv2.0 model
|
||||||
|
colbert_model = LateInteractionTextEmbedding("colbert-ir/colbertv2.0")
|
||||||
|
```
|
||||||
|
|
||||||
|
Encode the query into multiple token-level vectors:
|
||||||
|
|
||||||
|
```python
|
||||||
|
colbert_query_vector = next(colbert_model.query_embed(query))
|
||||||
|
colbert_query_vector.shape
|
||||||
|
```
|
||||||
|
|
||||||
|
Encode documents - note that each has a different number of token vectors:
|
||||||
|
|
||||||
|
```python
|
||||||
|
colbert_vectors = list(colbert_model.passage_embed(documents))
|
||||||
|
[cv.shape for cv in colbert_vectors]
|
||||||
|
```
|
||||||
|
|
||||||
|
Compute MaxSim scores for each document:
|
||||||
|
|
||||||
|
```python
|
||||||
|
for colbert_doc_vector in colbert_vectors:
|
||||||
|
# For each document, compute similarity between all query-doc token pairs
|
||||||
|
dot_product = np.dot(colbert_query_vector, colbert_doc_vector.T)
|
||||||
|
# For each query token, take the maximum similarity with any doc token
|
||||||
|
max_scores = dot_product.max(axis=1)
|
||||||
|
# Sum these maximum similarities to get the final MaxSim score
|
||||||
|
print(max_scores.sum())
|
||||||
|
```
|
||||||
|
|
||||||
|
**What happened?** ColBERT produces a clear ranking:
|
||||||
|
1. **Document A** (highly relevant) - highest score
|
||||||
|
2. **Document B** (partially relevant) - moderate score
|
||||||
|
3. **Document C** (keyword-stuffed) - lower score
|
||||||
|
4. **Document D** (irrelevant) - lowest score
|
||||||
|
|
||||||
|
Document C now correctly ranks **lower** than Documents A and B. The token-level verification catches that while Document C mentions related keywords, it lacks the specific technical details the query requires.
|
||||||
|
|
||||||
|
### Why the Difference?
|
||||||
|
|
||||||
|
**Single-vector behavior**: Document C gets inflated scores because it contains many topically-related terms. The averaging process captures "this document is about Python web development and databases" but can't verify whether it actually addresses connection pool exhaustion.
|
||||||
|
|
||||||
|
**ColBERT behavior**: Each query token must find strong matches in the document:
|
||||||
|
- "prevent" needs a match -> Document C has no solution-oriented content
|
||||||
|
- "connection pool exhaustion" needs specific matches -> Document C mentions these words separately but not in context
|
||||||
|
- "async" + "database" + "Python" must all connect -> Document C has them but not in the right relationship
|
||||||
|
|
||||||
|
**The key insight**: MaxSim's requirement that **every query token finds a strong match** prevents keyword-stuffing from inflating scores. It's not enough to mention the right topics - the document must contain those concepts with the right semantic relationships.
|
||||||
|
|
||||||
|
## When NOT to Use Multi-Vector Search
|
||||||
|
|
||||||
|
Multi-vector search isn't always necessary. Two scenarios where single-vector embeddings are preferable:
|
||||||
|
|
||||||
|
**Simple, broad queries**: When you're searching for general topical relevance rather than verifying specific requirements, single-vector embeddings work well. Queries like "Python tutorials" or "machine learning basics" seek documents about a topic, not documents containing multiple specific concepts. For these queries, the added precision of token-level matching doesn't provide significant value - topical similarity is exactly what you need.
|
||||||
|
|
||||||
|
**Resource-constrained environments**: Multi-vector search comes with substantial overhead:
|
||||||
|
- **Storage**: Multiple vectors per document instead of one (even thousands of them)
|
||||||
|
- **Memory**: All token vectors must be accessible during search
|
||||||
|
- **Computation**: MaxSim is incompatible with HNSW and requires computing similarities across all query-document token pairs
|
||||||
|
|
||||||
|
When deploying to systems with strict latency requirements, these costs may be prohibitive. In these cases, you're making a trade-off: accepting lower precision on complex queries to meet resource constraints. The precision benefits of multi-vector search apply regardless of collection size. Collection size affects whether you can afford the overhead, not whether you need the precision.
|
||||||
|
|
||||||
|
These resource considerations are real challenges you'll face in production. The next lesson examines them in detail.
|
||||||
|
|
||||||
|
## Conclusion
|
||||||
|
|
||||||
|
Multi-vector search excels at **fine-grained matching** - scenarios where you need to verify that ALL aspects of a query are present in a document, not just that they're topically related. By preserving token-level information rather than averaging it away, ColBERT and other late interaction models enable precision that single-vector embeddings cannot achieve.
|
||||||
|
|
||||||
|
When you have multi-requirement queries, need contextual precision, or must distinguish between partial and complete information, MaxSim's token-level verification provides clear advantages. Each query token independently finds its best match, and the aggregation ensures all requirements are strongly present.
|
||||||
|
|
||||||
|
But this power comes at a cost. In the next lesson, we'll examine the challenges: storing hundreds of vectors per document, computing MaxSim efficiently, and the indexing limitations you learned about in the MaxSim lesson.
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
---
|
||||||
|
title: "Module 2: Multi-Vector Representations for Multi-Modal Data"
|
||||||
|
description: "Explore multi-modal multi-vector search with ColPali. Learn how to search across images and text, and configure Qdrant for multi-vector embeddings."
|
||||||
|
isLesson: true
|
||||||
|
weight: 30
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 2 {{< /date >}}
|
||||||
|
|
||||||
|
# Multi-Vector Representations for Multi-Modal Data
|
||||||
|
|
||||||
|
Extend multi-vector representations beyond text to unlock powerful multi-modal search capabilities.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Today's path
|
||||||
|
|
||||||
|
1. How ColPali Models Work
|
||||||
|
2. ColPali Family Overview
|
||||||
|
3. Visual Interpretability of ColPali
|
||||||
|
|
||||||
|
You'll learn to build multi-modal search systems that understand both images and text.
|
||||||
@@ -0,0 +1,130 @@
|
|||||||
|
---
|
||||||
|
title: "ColPali Family Overview"
|
||||||
|
description: Explore the ColPali model family and their capabilities for multi-modal document understanding and retrieval.
|
||||||
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 2 {{< /date >}}
|
||||||
|
|
||||||
|
# ColPali Family Overview
|
||||||
|
|
||||||
|
The ColPali is not only the name of a model. Still, it is also often used to refer to an entire family of models that convert images and text into multi-vector representations, based on Vision Language Models.
|
||||||
|
|
||||||
|
Let's explore what the options are and which model to choose depending on the data you work with.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/5ypH0t_X-4k?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
The ColPali family includes several model variants. When selecting a model for your application, you'll need to consider factors like model size, supported languages, computational requirements, and licensing constraints - each variant offers different trade-offs along these dimensions.
|
||||||
|
|
||||||
|
## Model Size
|
||||||
|
|
||||||
|
While the original ColPali model delivers good performance, its multi-billion parameter size can be challenging for resource-constrained environments, demos, or CPU-only deployments. Fortunately, smaller alternatives maintain competitive performance while dramatically reducing computational requirements.
|
||||||
|
|
||||||
|
### ColSmol: Efficient Small-Scale Models
|
||||||
|
|
||||||
|
The **[ColSmol](https://huggingface.co/vidore/colSmol-256M)** family offers compact variants built on SmolVLM, available in [256M](https://huggingface.co/vidore/colSmol-256M) and [500M](https://huggingface.co/vidore/colSmol-500M) parameter sizes. These Apache 2.0 licensed models generate ColBERT-style multi-vector representations while being small enough for resource-constrained environments, like browser-based applications, or edge computing.
|
||||||
|
|
||||||
|
### ColFlor: Ultra-Compact Retrieval
|
||||||
|
|
||||||
|
**[ColFlor](https://huggingface.co/ahmed-masry/ColFlor)** pushes efficiency even further with only 174 million parameters, achieving performance 17× smaller and up to 9.8× faster than ColPali with only a 1.8% drop in accuracy on text-rich English documents. Built on Florence-2's architecture, ColFlor is particularly attractive for demo environments, educational purposes, and scenarios where computational efficiency outweighs marginal performance differences.
|
||||||
|
|
||||||
|
## Supported Languages
|
||||||
|
|
||||||
|
Original ColPali model primarily focuses on English documents, but several multilingual alternatives have emerged that extend multi-vector capabilities across different languages.
|
||||||
|
|
||||||
|
### NVIDIA Multilingual Models
|
||||||
|
|
||||||
|
The **[NVIDIA Llama-NeMoRetriever-ColEmbed-3B-v1](https://huggingface.co/nvidia/llama-nemoretriever-colembed-3b-v1)** was among the leading multilingual visual document retrieval models when it was released. Built on top of Google's SigLIP-2 vision encoder and Meta's Llama 3.2-3B language model, this late interaction embedding model demonstrated strong performance on multilingual retrieval benchmarks including ViDoRe and MIRACL-VISION.
|
||||||
|
|
||||||
|
However, **potential users should be aware of licensing restrictions**. The model is available **for non-commercial and research use only** due to multiple overlapping licenses: NVIDIA's Non-Commercial License, Apache 2.0 for the SigLIP-2 component, and Meta's Llama 3.2 Community License Agreement. Organizations requiring commercial deployment should carefully review these license terms or consider alternatives.
|
||||||
|
|
||||||
|
### Open-Source Multilingual Alternative
|
||||||
|
|
||||||
|
For commercial applications, **[Nomic AI's ColNomic-Embed-Multimodal-7B](https://huggingface.co/nomic-ai/colnomic-embed-multimodal-7b)** offers a compelling fully open-source alternative. Released in early 2025, this model demonstrated competitive performance on multilingual retrieval benchmarks. The model is available under an open-source license that permits commercial use, making it suitable for production deployments without licensing concerns. Nomic AI released a complete suite including both multi-vector (ColNomic) and single-vector variants in 3B and 7B parameter sizes, giving developers flexibility in choosing the right trade-off between performance and resource requirements.
|
||||||
|
|
||||||
|
## Benchmarking Visual Document Retrieval
|
||||||
|
|
||||||
|
The **ViDoRe (Visual Document Retrieval) Benchmark** has emerged as the leading evaluation framework for visual retrieval models on document understanding tasks. As of January 2026, it stands as the largest and most comprehensive benchmark in the field, evaluating models across multiple domains, 6 languages (including English and French), and realistic retrieval scenarios including cross-document and long-form queries. You can explore the latest model performance and compare different approaches on the [ViDoRe leaderboard](https://huggingface.co/spaces/vidore/vidore-leaderboard).
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
***Source:** https://huggingface.co/vidore*
|
||||||
|
|
||||||
|
The current version, **ViDoRe V3**, represents the benchmark's scale and ambition with 26,000+ pages across 3,099 queries in 6 languages, spanning 10 datasets (8 public, 2 private). The benchmark uses **nDCG** (Normalized Discounted Cumulative Gain) as its primary metric and includes challenging datasets spanning diverse domains - including HR, finance, industrial, pharmaceuticals, physics, computer science, and energy sectors. What sets ViDoRe V3 apart is its focus on real-world complexity: models are tested on truly challenging retrieval tasks with human-verified annotations that mirror actual user behavior in enterprise document retrieval scenarios.
|
||||||
|
|
||||||
|
## The Impact of Bidirectional Attention
|
||||||
|
|
||||||
|
Bidirectional attention has emerged as a promising approach for multi-vector representations of multi-modal data.
|
||||||
|
|
||||||
|
The choice between **unidirectional** and **bidirectional** attention mechanisms significantly impacts model performance for embedding tasks. Understanding this distinction helps explain why certain architectures excel at retrieval while others are optimized for generation. Let's first examine unidirectional attention and its limitations, then see how bidirectional attention addresses these constraints.
|
||||||
|
|
||||||
|
### Unidirectional Attention: Designed for Generation
|
||||||
|
|
||||||
|
Most large language models (LLMs) and vision-language models (VLMs) like GPT, Llama, and the base models of [ColPali](https://huggingface.co/vidore/colpali) use **unidirectional (causal) attention**. In this approach, each token can only attend to tokens that came before it in the sequence - the model looks backward but never forward.
|
||||||
|
|
||||||
|
This design makes perfect sense for generative tasks: when predicting the next token, the model should only use past context, not future information it hasn't generated yet. However, this constraint creates limitations for embedding tasks. Token representations encode only information from previous context, missing crucial contextual information from subsequent tokens in the sequence.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
The default ColPali model uses the unidirectional attention, derived from the underlying VLM. As a reminder, here is how we load the ColPali v.1.3 with FastEmbed.
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import LateInteractionMultimodalEmbedding
|
||||||
|
|
||||||
|
# Load the Qdrant/colpali-v1.3-fp16 model from HF hub
|
||||||
|
colpali_model = LateInteractionMultimodalEmbedding(
|
||||||
|
model_name="Qdrant/colpali-v1.3-fp16"
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Bidirectional Attention: Optimized for Embeddings
|
||||||
|
|
||||||
|
**Bidirectional attention**, as used in encoder models like BERT, allows each token to attend to the entire input sequence - both past and future tokens. This creates richer, more contextually informed representations since each token's embedding incorporates information from the complete surrounding context.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
Research has shown that [bidirectional encoder models are often the best option when training visual retrievers](https://arxiv.org/html/2510.01149). The [**ColModernVBERT**](https://huggingface.co/ModernVBERT/colmodernvbert) model exemplifies this approach: with only ~250 million parameters - over 10 times fewer than ColPali - it achieves performance only slightly lower on the ViDoRe benchmark.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
***Source:** Teiletche, P., Macé, Q., Conti, M., Loison, A., Viaud, G., Colombo, P., & Faysse, M. (2025). *ModernVBERT: Towards Smaller Visual Document Retrievers*. arXiv preprint arXiv:2510.01149. https://arxiv.org/abs/2510.01149*
|
||||||
|
|
||||||
|
The chart above compares performance on the ViDoRe V2 benchmark. While ColPali and ColQwen achieve strong results with unidirectional attention, **ColModernVBERT stands out as the only bidirectional model** - proving that full-context attention enables competitive performance with dramatically fewer parameters. By allowing context to flow in both directions, these compact models generate embeddings that capture the nuanced semantic relationships critical for accurate multi-modal retrieval where visual and textual information must be jointly understood.
|
||||||
|
|
||||||
|
This efficiency makes ModernVBERT particularly attractive for resource-constrained environments or CPU-only deployments, where the slight performance trade-off is worthwhile for the substantial gains in speed and reduced computational requirements.
|
||||||
|
|
||||||
|
ColModernVBERT is available in FastEmbed and might be used like any other late interaction model for multi-modal data.
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import LateInteractionMultimodalEmbedding
|
||||||
|
|
||||||
|
# Load the Qdrant/colmodernvbert model from HF hub
|
||||||
|
colpali_model = LateInteractionMultimodalEmbedding(
|
||||||
|
model_name="Qdrant/colmodernvbert"
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Future of Multi-Vector Representations
|
||||||
|
|
||||||
|
The multi-vector approach extends naturally to new modalities beyond images. Models like **[TomoroAI/tomoro-colqwen3-embed-8b](https://huggingface.co/TomoroAI/tomoro-colqwen3-embed-8b)**, based on ColQwen3, demonstrate this extensibility by adding support for short video retrieval. Based on preliminary findings, ColQwen3 generalizes to videos while learning from image-text retrieval tasks - it samples video clips, encodes frames, then pools frame embeddings with per-dimension max before MaxSim scoring.
|
||||||
|
|
||||||
|
While not yet fine-tuned on large-scale video retrieval datasets, this lightweight approach highlights a key strength of the multi-vector paradigm, where **new modalities can be added** while preserving the core late interaction architecture.
|
||||||
|
|
||||||
|
## What's next
|
||||||
|
|
||||||
|
You've explored the ColPali family ecosystem - from ultra-compact models like ColFlor (174M parameters) to multilingual alternatives from NVIDIA and Nomic AI. You've learned how bidirectional attention architectures can achieve competitive results with far fewer parameters, and seen the benchmark results that guide model selection based on your performance and resource requirements.
|
||||||
|
|
||||||
|
Now that you know which model to use, let's learn how to interpret what ColPali "sees" in your documents.
|
||||||
@@ -0,0 +1,305 @@
|
|||||||
|
---
|
||||||
|
title: "How ColPali Models Work"
|
||||||
|
description: Understand the inner workings of ColPali models and how they generate multi-vector representations for images and documents.
|
||||||
|
weight: 1
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 2 {{< /date >}}
|
||||||
|
|
||||||
|
# How ColPali Models Work
|
||||||
|
|
||||||
|
ColPali extends the late interaction paradigm from text to visual documents. It can process PDFs, images, and scanned documents, generating multi-vector representations that capture both textual and visual information.
|
||||||
|
|
||||||
|
Understanding ColPali's architecture helps you leverage its full potential for multi-modal document retrieval.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/Fai9aY1PMCA?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-2/how-colpali-works.ipynb">
|
||||||
|
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||||
|
</a>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## From Text to Visual Documents
|
||||||
|
|
||||||
|
**What about documents that aren't just text?** PDFs often contain diagrams, tables, charts, equations, and complex layouts where the visual presentation carries as much meaning as the text itself.
|
||||||
|
|
||||||
|
Traditional approaches to searching visual documents typically involve two steps: first, extract text using OCR (Optical Character Recognition), then search the extracted text. This pipeline has significant limitations:
|
||||||
|
|
||||||
|
- **Lost layout information**: OCR converts documents to plain text, discarding spatial relationships and visual structure
|
||||||
|
- **Diagram blindness**: Charts, graphs, and diagrams are either ignored or poorly represented
|
||||||
|
- **Format fragility**: Tables, mathematical notation, and multi-column layouts often get mangled during text extraction
|
||||||
|
|
||||||
|
**ColPali takes a fundamentally different approach**: it treats the document image itself as the primary representation. ColPali "tokenizes" images into spatial patches. Each patch becomes a visual token with its own embedding, enabling token-level matching without ever extracting text.
|
||||||
|
|
||||||
|
This means ColPali can match queries directly to visual regions of document pages - no OCR required.
|
||||||
|
|
||||||
|
## The Vision Language Model Foundation
|
||||||
|
|
||||||
|
Before diving into how ColPali processes images, let's understand its architectural foundation. **ColPali isn't built from scratch** - it's based on sophisticated Vision Language Models (VLMs) that already understand the relationship between visual and textual information.
|
||||||
|
|
||||||
|
<aside role="status">
|
||||||
|
This lesson focuses specifically on ColPali v1.3. The preprocessing pipeline may differ between different models, yet the principles should be the same.
|
||||||
|
</aside>
|
||||||
|
|
||||||
|
### Building on PaliGemma
|
||||||
|
|
||||||
|
ColPali v1.3 is built on **PaliGemma-3B**, a Vision Language Model that combines two powerful components:
|
||||||
|
|
||||||
|
- **SigLIP-So400m** (Vision Encoder): Processes images into visual features
|
||||||
|
- **Gemma-2B** (Language Model): Contextualizes those features using transformer layers
|
||||||
|
|
||||||
|
Why use a VLM instead of training a vision model from scratch? VLMs are pre-trained on massive datasets of image-text pairs, so they already "understand" how visual content relates to language. This makes them ideal for document retrieval: they can naturally connect text queries to corresponding visual content.
|
||||||
|
|
||||||
|
PaliGemma was specifically designed for tasks requiring both visual understanding and language processing - perfect for searching documents that blend text, diagrams, tables, and equations.
|
||||||
|
|
||||||
|
**Fine-tuning for Document Retrieval**: The base PaliGemma model is adapted for document retrieval using LoRA (Low-Rank Adaptation), a parameter-efficient technique that specializes the model without retraining everything. The vision encoder stays frozen while attention layers are fine-tuned to optimize for document search.
|
||||||
|
|
||||||
|
### The Two-Component Architecture
|
||||||
|
|
||||||
|
Think of ColPali's architecture as having "eyes" and a "brain":
|
||||||
|
|
||||||
|
**Vision Encoder (SigLIP) - The Eyes:**
|
||||||
|
- Takes the document image and divides it into patches (the 32×32 grid we'll explore next)
|
||||||
|
- Processes each patch through a vision transformer
|
||||||
|
- Outputs visual features that capture what's in each patch: text characters, diagram elements, table cells, etc.
|
||||||
|
|
||||||
|
**Language Model (Gemma-2B) - The Brain:**
|
||||||
|
- Receives the visual features from SigLIP
|
||||||
|
- Runs them through transformer layers to add context and semantic understanding
|
||||||
|
- Each patch's features get enriched by understanding neighboring patches
|
||||||
|
- Outputs contextualized representations that understand relationships across the document
|
||||||
|
|
||||||
|
**The Flow:** Image patches are first processed through the SigLIP encoder to generate visual features. These visual features then pass through the Gemma-2B transformer, which produces contextualized representations. Finally, the contextualized representations flow through a projection layer to produce 128-dimensional embeddings. This pipeline ensures that each 128-dim embedding doesn't just capture what's in one patch - it understands that patch in the context of the entire document.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
For text inputs, there isn't any additional preprocessing step, but it is tokenized and passed through the language model. This architecture can use the same "brain" to process different modalities.
|
||||||
|
|
||||||
|
## Patch-Based Image Processing
|
||||||
|
|
||||||
|
Vision transformers, the foundation of ColPali, don't process entire images at once. Instead, they divide images into a grid of fixed-size patches - think of it like a checkerboard overlaid on your document.
|
||||||
|
|
||||||
|
The ColPali pipeline starts with a document. It is usually a screenshot of a single PDF page, or anything you want to encode. While you might convert a PDF page to a high-resolution screenshot to preserve visual details during conversion, **the ColPali preprocessing itself resizes images to a fixed resolution** of **448×448 pixels** before processing.
|
||||||
|
|
||||||
|
This fixed-size input is then divided into patches:
|
||||||
|
- **Patch size**: 14×14 pixels per patch
|
||||||
|
- **Patch grid**: 32×32 patches (448 / 14 = 32 patches per side)
|
||||||
|
- **Total patches**: 1024 patches per image (32 × 32 = 1024)
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
**Each patch becomes a visual "token"** representing a spatial region of the document. A patch might contain part of a diagram, a few words of text, a piece of an equation, or even just whitespace - each gets its own embedding. This consistent structure ensures predictable embedding sizes and simplifies the search pipeline.
|
||||||
|
|
||||||
|
## Token Structure and Embeddings
|
||||||
|
|
||||||
|
Now that you understand ColPali's VLM foundation - SigLIP processing patches, Gemma-2B contextualizing features, and the projection layer generating embeddings - let's see how this architecture processes images into the final multi-vector representation.
|
||||||
|
|
||||||
|
When ColPali processes an image, the PaliGemma model doesn't just generate patch embeddings directly. It creates a sequence of tokens that flows through the entire architecture we just described. This sequence includes both the image patches and additional instruction tokens that help guide the model's behavior.
|
||||||
|
|
||||||
|
For a single document image, the token sequence contains:
|
||||||
|
- **1024 special `<image>` tokens** - one for each patch in the 32×32 grid
|
||||||
|
- **Instruction tokens** - additional tokens that help the model understand its task
|
||||||
|
|
||||||
|
`<image><image>...<image><bos>Describe the image.\n`
|
||||||
|
|
||||||
|
The final embedding output contains **1030 vectors** of 128 dimensions each: 1024 for the image patches plus 6 additional instruction tokens. For search purposes, all vectors participate in the MaxSim calculation - each query token finds its best match among all 1030 vectors.
|
||||||
|
|
||||||
|
With this multi-vector representation, each query word finds its best match among all 1030 visual vectors using **late interaction**. Query terms match to the most semantically similar patches - whether those patches contain diagrams, text, tables, or equations - exactly the same MaxSim scoring we learned about in Module 1, but extended to visual documents.
|
||||||
|
|
||||||
|
## Implementing Multi-Modal Search with FastEmbed and Qdrant
|
||||||
|
|
||||||
|
Now that you understand ColPali's architecture - from its VLM foundation to patch-based processing to multi-vector embeddings - let's build a practical multi-modal search system. We'll again use **FastEmbed**, as it also provides a unified interface for multi-modal late interaction models like ColPali.
|
||||||
|
|
||||||
|
### Loading ColPali with FastEmbed
|
||||||
|
|
||||||
|
Let's start by loading a ColPali model. FastEmbed supports several ColPali variants, but in this lesson we'll use ColPali v1.3:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import LateInteractionMultimodalEmbedding
|
||||||
|
|
||||||
|
# Load the Qdrant/colpali-v1.3-fp16 model from HF hub
|
||||||
|
colpali_model = LateInteractionMultimodalEmbedding(
|
||||||
|
model_name="Qdrant/colpali-v1.3-fp16"
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
This single initialization loads the entire ColPali architecture we discussed: SigLIP vision encoder, Gemma-2B language model, and the projection layer. The `LateInteractionMultimodalEmbedding` class handles both image and text encoding through a unified interface, encapsulating the image and text processing logic so you can just pass image paths or queries.
|
||||||
|
|
||||||
|
### Processing Images and Understanding Embeddings
|
||||||
|
|
||||||
|
Now let's process a document image through ColPali to see the multi-vector embeddings in action:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from PIL import Image
|
||||||
|
|
||||||
|
image = Image.open("images/einstein-newspaper.jpg")
|
||||||
|
image
|
||||||
|
```
|
||||||
|
|
||||||
|
Now let's generate the embeddings:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Create the representation of the image
|
||||||
|
colpali_generator = colpali_model.embed_image(["images/einstein-newspaper.jpg"])
|
||||||
|
document_vectors = next(colpali_generator)
|
||||||
|
document_vectors.shape
|
||||||
|
```
|
||||||
|
|
||||||
|
The output shape shows **1030 vectors of 128 dimensions each**. These 1030 vectors include the 1024 image patch embeddings (from the 32×32 grid) plus 6 additional instruction tokens used by the model.
|
||||||
|
|
||||||
|
**Processing query text:**
|
||||||
|
|
||||||
|
ColPali also encodes text queries into multi-vector representations:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Create the representation of the query
|
||||||
|
query = "When did dr. Einstein die?"
|
||||||
|
query_vectors = next(colpali_model.embed_text(query))
|
||||||
|
query_vectors.shape
|
||||||
|
```
|
||||||
|
|
||||||
|
Let's define a helper function to compute the MaxSim score:
|
||||||
|
|
||||||
|
```python
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
def maxsim(Q, D):
|
||||||
|
sims = np.dot(Q, D.T)
|
||||||
|
max_sims = sims.max(axis=1)
|
||||||
|
return max_sims.sum()
|
||||||
|
```
|
||||||
|
|
||||||
|
Now finally compute the score:
|
||||||
|
|
||||||
|
```python
|
||||||
|
maxsim(query_vectors, document_vectors)
|
||||||
|
```
|
||||||
|
|
||||||
|
During search, each query token finds its best match among all 1030 document vectors using **MaxSim** - the same late interaction scoring we learned in Module 1. A higher score indicates better semantic overlap between the query and the visual document.
|
||||||
|
|
||||||
|
### Setting Up Qdrant for Multi-Vector Search
|
||||||
|
|
||||||
|
To store and search these multi-vector embeddings, we need a Qdrant collection configured for late interaction:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
|
client = QdrantClient("http://localhost:6333")
|
||||||
|
|
||||||
|
client.create_collection(
|
||||||
|
collection_name="colpali",
|
||||||
|
vectors_config={
|
||||||
|
"colpali-v1.3": models.VectorParams(
|
||||||
|
size=128,
|
||||||
|
distance=models.Distance.DOT,
|
||||||
|
multivector_config=models.MultiVectorConfig(
|
||||||
|
comparator=models.MultiVectorComparator.MAX_SIM,
|
||||||
|
),
|
||||||
|
hnsw_config=models.HnswConfigDiff(m=0),
|
||||||
|
),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
This configuration mirrors what you learned in [Module 1](/course/multi-vector-search/module-1/multi-vector-in-qdrant/) - multi-vector config with MAX_SIM comparator, dot product distance, and HNSW disabled.
|
||||||
|
|
||||||
|
### Indexing Visual Documents
|
||||||
|
|
||||||
|
Now let's index visual documents using Qdrant's **local inference** feature. This approach uses `models.Image()` to let Qdrant handle embedding generation automatically via FastEmbed:
|
||||||
|
|
||||||
|
```python
|
||||||
|
import uuid
|
||||||
|
|
||||||
|
documents = [
|
||||||
|
"images/einstein-newspaper.jpg",
|
||||||
|
"images/titanic-newspaper.jpg",
|
||||||
|
"images/men-walk-on-moon-newspaper.jpg",
|
||||||
|
]
|
||||||
|
client.upsert(
|
||||||
|
collection_name="colpali",
|
||||||
|
points=[
|
||||||
|
models.PointStruct(
|
||||||
|
id=uuid.uuid4().hex,
|
||||||
|
vector={
|
||||||
|
"colpali-v1.3": models.Image(
|
||||||
|
image=doc,
|
||||||
|
model="Qdrant/colpali-v1.3-fp16",
|
||||||
|
)
|
||||||
|
},
|
||||||
|
payload={
|
||||||
|
"image": doc,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
for doc in documents
|
||||||
|
]
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
This uses **local inference** - similar to `models.Document()` for text that you saw in Module 1. Qdrant's FastEmbed integration processes images locally (not on a remote server), generating the multi-vector embeddings automatically. This is much simpler than manually generating embeddings with FastEmbed and then upserting them.
|
||||||
|
|
||||||
|
Once indexed, each document image is represented by its 1030 vectors (1024 patches + 6 instruction tokens), ready for late interaction search. When a query comes in, Qdrant will compute MaxSim scores between the query tokens and these vectors.
|
||||||
|
|
||||||
|
### Querying with Late Interaction
|
||||||
|
|
||||||
|
The power of ColPali becomes clear when searching. Using local inference, you can query with text and match against visual documents:
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.query_points(
|
||||||
|
collection_name="colpali",
|
||||||
|
query=models.Document(
|
||||||
|
text="Who was the first man on the moon?",
|
||||||
|
model="Qdrant/colpali-v1.3-fp16",
|
||||||
|
),
|
||||||
|
using="colpali-v1.3",
|
||||||
|
limit=2,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
This uses the same local inference pattern you saw in Module 1 with `models.Document()`. Qdrant handles query tokenization, embedding generation, and MaxSim computation automatically.
|
||||||
|
|
||||||
|
**Let's try different queries:**
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.query_points(
|
||||||
|
collection_name="colpali",
|
||||||
|
query=models.Document(
|
||||||
|
text="Why did Titanic sink?",
|
||||||
|
model="Qdrant/colpali-v1.3-fp16",
|
||||||
|
),
|
||||||
|
using="colpali-v1.3",
|
||||||
|
limit=2,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
**How late interaction works for visual search:**
|
||||||
|
|
||||||
|
1. Your query is tokenized into query embeddings (e.g., 6-8 tokens for "Who was the first man on the moon?")
|
||||||
|
2. For each query token, Qdrant finds the maximum similarity across all 1030 visual vectors of each document image
|
||||||
|
3. The maximum similarities are summed (MaxSim) to score each image
|
||||||
|
4. Images with content matching the query score highest
|
||||||
|
|
||||||
|
This is **text-to-image search without OCR**. You're matching the semantic meaning of text queries directly to visual content - headlines, photos, captions, and text in its original layout. The late interaction paradigm enables fine-grained matching at the token level, just like ColBERT for text, but extended to visual documents.
|
||||||
|
|
||||||
|
## What's Next
|
||||||
|
|
||||||
|
Now that you understand **how ColPali works** and can build basic multi-modal search systems, several questions naturally arise:
|
||||||
|
|
||||||
|
- **Which ColPali model should you use?** The ColPali family includes variants optimized for different hardware constraints and accuracy requirements.
|
||||||
|
- **How can you interpret what the model is finding?** Unlike black-box embeddings, ColPali offers powerful visualization capabilities to see exactly which image regions match your query terms.
|
||||||
|
- **How do you optimize for production?** Multi-vector search can be resource-intensive at scale - you'll need techniques like quantization, pooling, and multi-stage retrieval.
|
||||||
|
|
||||||
|
In the next lesson, we'll explore the **ColPali family** and learn when to use each model variant. Then, we'll dive into visual interpretability to understand exactly what ColPali sees in your documents.
|
||||||
@@ -0,0 +1,399 @@
|
|||||||
|
---
|
||||||
|
title: "Visual Interpretability of ColPali"
|
||||||
|
description: Learn how to visualize and interpret ColPali embeddings to understand what the model focuses on in images.
|
||||||
|
weight: 3
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 2 {{< /date >}}
|
||||||
|
|
||||||
|
# Visual Interpretability of ColPali
|
||||||
|
|
||||||
|
**Why did this document match my query?** Unlike traditional black-box embedding models that produce a single opaque vector, ColPali's multi-vector architecture offers something remarkable: you can see exactly where the model "looks" when matching a query to a document.
|
||||||
|
|
||||||
|
This visual interpretability is invaluable for building trust in multi-modal search systems, debugging unexpected results, and understanding model behavior and limitations.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/sQcuYWMS4bo?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-2/visual-interpretability.ipynb">
|
||||||
|
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||||
|
</a>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The Key Insight: Spatial Correspondence
|
||||||
|
|
||||||
|
In the [previous lesson on ColPali's architecture](/course/multi-vector-search/module-2/how-colpali-works/), you learned that ColPali divides images into a 32×32 grid of patches:
|
||||||
|
|
||||||
|
- **Input image**: 448×448 pixels
|
||||||
|
- **Patch size**: 14×14 pixels
|
||||||
|
- **Patch grid**: 32×32 patches
|
||||||
|
- **Total patch embeddings**: 1024 vectors
|
||||||
|
|
||||||
|
The crucial insight for interpretability is that **each embedding maintains a known spatial location**. Patch index `i` (where `i` ranges from 0 to 1023) maps directly to a position in the grid:
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{row} = \left\lfloor \frac{i}{32} \right\rfloor
|
||||||
|
$$
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{col} = i \mod 32
|
||||||
|
$$
|
||||||
|
|
||||||
|
This position corresponds to a specific pixel region in the original image:
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{pixel\_region} = \text{image}\left[\text{row} \cdot 14 : (\text{row}+1) \cdot 14, \; \text{col} \cdot 14 : (\text{col}+1) \cdot 14\right]
|
||||||
|
$$
|
||||||
|
|
||||||
|
<!--
|
||||||
|

|
||||||
|
|
||||||
|
TODO: Add diagram showing 32×32 grid overlaid on a document image
|
||||||
|
- Show a document page with the 32×32 grid overlaid
|
||||||
|
- Highlight a few specific patches (e.g., one containing text, one containing a diagram element)
|
||||||
|
- Label the row/col coordinates for the highlighted patches
|
||||||
|
- Show the formula: patch_index → (row, col) → pixel_region
|
||||||
|
-->
|
||||||
|
|
||||||
|
This spatial correspondence is what makes visual interpretability possible. When a query token has high similarity with a particular patch embedding, we know exactly where in the document that match occurred.
|
||||||
|
|
||||||
|
## Computing Token-Patch Similarities
|
||||||
|
|
||||||
|
To visualize what ColPali focuses on, we compute the similarity between each query token and all document patches. For a given query token embedding, we can calculate its similarity with each of the 1024 patch embeddings:
|
||||||
|
|
||||||
|
```python
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
def compute_similarity_map(query_token_vec, doc_vectors):
|
||||||
|
"""Compute similarity map for a single query token."""
|
||||||
|
# Take only the 1024 patch embeddings (excluding instruction tokens)
|
||||||
|
patch_vectors = doc_vectors[:1024]
|
||||||
|
|
||||||
|
# Compute dot product similarity with all patches
|
||||||
|
similarities = np.dot(patch_vectors, query_token_vec)
|
||||||
|
|
||||||
|
# Reshape to 32×32 spatial grid
|
||||||
|
return similarities.reshape(32, 32)
|
||||||
|
```
|
||||||
|
|
||||||
|
The result is a 32×32 similarity map - essentially a heatmap showing where in the document this particular query token has the strongest matches. High values indicate regions where the model finds semantic relevance to that token.
|
||||||
|
|
||||||
|
## Manual Implementation with FastEmbed
|
||||||
|
|
||||||
|
Let's build a complete interpretability system from scratch using only FastEmbed and standard Python libraries. This approach works with any late interaction model and gives you full control over the visualization process.
|
||||||
|
|
||||||
|
### Step 1: Generate Embeddings
|
||||||
|
|
||||||
|
First, let's load a model and generate embeddings for both a document image and a query:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import LateInteractionMultimodalEmbedding
|
||||||
|
|
||||||
|
# Load ColPali model
|
||||||
|
model = LateInteractionMultimodalEmbedding(
|
||||||
|
model_name="Qdrant/colpali-v1.3-fp16"
|
||||||
|
)
|
||||||
|
|
||||||
|
# Load and embed a document image
|
||||||
|
image_path = "images/einstein-newspaper.jpg"
|
||||||
|
doc_vectors = next(model.embed_image([image_path]))
|
||||||
|
|
||||||
|
# Embed a query
|
||||||
|
query = "When did Einstein die?"
|
||||||
|
query_vectors = next(model.embed_text(query))
|
||||||
|
|
||||||
|
print(f"Document embeddings shape: {doc_vectors.shape}") # (1030, 128)
|
||||||
|
print(f"Query embeddings shape: {query_vectors.shape}") # (N, 128) where N = number of tokens
|
||||||
|
```
|
||||||
|
|
||||||
|
### Step 2: Compute Similarity Maps for All Query Tokens
|
||||||
|
|
||||||
|
Now we compute a similarity map for each query token:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def compute_all_similarity_maps(query_vectors, doc_vectors):
|
||||||
|
"""Compute similarity maps for all query tokens."""
|
||||||
|
return np.array([
|
||||||
|
compute_similarity_map(query_token_vec, doc_vectors)
|
||||||
|
for query_token_vec in query_vectors
|
||||||
|
])
|
||||||
|
|
||||||
|
# Compute similarity maps
|
||||||
|
similarity_maps = compute_all_similarity_maps(query_vectors, doc_vectors)
|
||||||
|
print(f"Similarity maps shape: {similarity_maps.shape}") # (N, 32, 32)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Creating Heatmap Visualizations
|
||||||
|
|
||||||
|
To visualize the similarity maps, we need to:
|
||||||
|
1. Upsample the 32×32 map to match the image dimensions
|
||||||
|
2. Overlay it as a semi-transparent heatmap on the original image
|
||||||
|
|
||||||
|
### Step 3: Upsample and Overlay
|
||||||
|
|
||||||
|
```python
|
||||||
|
from scipy.ndimage import zoom
|
||||||
|
import matplotlib.pyplot as plt
|
||||||
|
import matplotlib.cm as cm
|
||||||
|
|
||||||
|
def create_heatmap_overlay(image, similarity_map, alpha=0.5):
|
||||||
|
"""Create a heatmap overlay on the original image."""
|
||||||
|
# Ensure image is in RGB and resized to 448×448 (ColPali's input size)
|
||||||
|
if isinstance(image, str):
|
||||||
|
image = Image.open(image)
|
||||||
|
image = image.convert("RGB").resize((448, 448))
|
||||||
|
image_array = np.array(image)
|
||||||
|
|
||||||
|
# Upsample similarity map from 32×32 to 448×448
|
||||||
|
# zoom factor = 448/32 = 14
|
||||||
|
# Convert to float64 as scipy.ndimage.zoom doesn't support all dtypes (e.g., float16)
|
||||||
|
upsampled_map = zoom(similarity_map.astype(np.float64), 14, order=1)
|
||||||
|
|
||||||
|
# Normalize to [0, 1] range
|
||||||
|
min_val = upsampled_map.min()
|
||||||
|
max_val = upsampled_map.max()
|
||||||
|
if max_val > min_val:
|
||||||
|
normalized_map = (upsampled_map - min_val) / (max_val - min_val)
|
||||||
|
else:
|
||||||
|
normalized_map = np.zeros_like(upsampled_map)
|
||||||
|
|
||||||
|
# Apply colormap (using 'jet' for red=high, blue=low)
|
||||||
|
heatmap = cm.jet(normalized_map)[:, :, :3] # Remove alpha channel
|
||||||
|
heatmap = (heatmap * 255).astype(np.uint8)
|
||||||
|
|
||||||
|
# Blend with original image
|
||||||
|
blended = (alpha * heatmap + (1 - alpha) * image_array).astype(np.uint8)
|
||||||
|
|
||||||
|
return Image.fromarray(blended), normalized_map
|
||||||
|
```
|
||||||
|
|
||||||
|
### Step 4: Visualize Multiple Query Tokens
|
||||||
|
|
||||||
|
Let's create a side-by-side visualization of what each query token focuses on:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def visualize_query_tokens(image_path, query, model, num_tokens_to_show=5):
|
||||||
|
"""Visualize similarity maps for each query token."""
|
||||||
|
# Generate embeddings
|
||||||
|
doc_vectors = next(model.embed_image([image_path]))
|
||||||
|
query_vectors = next(model.embed_text(query))
|
||||||
|
|
||||||
|
# Compute similarity maps
|
||||||
|
similarity_maps = compute_all_similarity_maps(query_vectors, doc_vectors)
|
||||||
|
|
||||||
|
# Load original image
|
||||||
|
original_image = Image.open(image_path).convert("RGB").resize((448, 448))
|
||||||
|
|
||||||
|
# Limit number of tokens to display
|
||||||
|
n_tokens = min(len(similarity_maps), num_tokens_to_show)
|
||||||
|
|
||||||
|
# Create figure
|
||||||
|
fig, axes = plt.subplots(1, n_tokens + 1, figsize=(4 * (n_tokens + 1), 4))
|
||||||
|
|
||||||
|
# Show original image
|
||||||
|
axes[0].imshow(original_image)
|
||||||
|
axes[0].set_title("Original")
|
||||||
|
axes[0].axis("off")
|
||||||
|
|
||||||
|
# Show heatmap for each token
|
||||||
|
for i in range(n_tokens):
|
||||||
|
overlay, _ = create_heatmap_overlay(original_image, similarity_maps[i])
|
||||||
|
axes[i + 1].imshow(overlay)
|
||||||
|
axes[i + 1].set_title(f"Token {i}")
|
||||||
|
axes[i + 1].axis("off")
|
||||||
|
|
||||||
|
plt.suptitle(f'Query: "{query}"', fontsize=14)
|
||||||
|
plt.tight_layout()
|
||||||
|
plt.show()
|
||||||
|
|
||||||
|
# Visualize what each token focuses on
|
||||||
|
visualize_query_tokens(
|
||||||
|
"images/einstein-newspaper.jpg",
|
||||||
|
"When did Einstein die?",
|
||||||
|
model,
|
||||||
|
num_tokens_to_show=6
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
## Practical Example: Debugging a Search
|
||||||
|
|
||||||
|
Visual interpretability becomes powerful when debugging search results. Let's walk through a complete example to understand why certain documents match (or don't match) specific queries.
|
||||||
|
|
||||||
|
### Scenario: Investigating an Unexpected Match
|
||||||
|
|
||||||
|
Imagine you're searching for "bar chart showing revenue" and get an unexpected result. Let's visualize what's happening:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from transformers import AutoTokenizer
|
||||||
|
|
||||||
|
# Load tokenizer for ColPali (based on PaliGemma)
|
||||||
|
tokenizer = AutoTokenizer.from_pretrained("google/paligemma-3b-pt-224")
|
||||||
|
|
||||||
|
def debug_search_result(image_path, query, model, tokenizer):
|
||||||
|
"""Debug why a document matched a query."""
|
||||||
|
# Generate embeddings
|
||||||
|
doc_vectors = next(model.embed_image([image_path]))
|
||||||
|
query_vectors = next(model.embed_text(query))
|
||||||
|
|
||||||
|
# Tokenize query to get actual token strings
|
||||||
|
# ColPali uses "Query: " prefix internally
|
||||||
|
query_with_prefix = f"Query: {query}"
|
||||||
|
tokens = tokenizer.tokenize(query_with_prefix)
|
||||||
|
|
||||||
|
# Compute MaxSim score
|
||||||
|
similarities = np.dot(query_vectors, doc_vectors.T)
|
||||||
|
max_sims = similarities.max(axis=1)
|
||||||
|
total_score = max_sims.sum()
|
||||||
|
|
||||||
|
print(f"Query: {query}")
|
||||||
|
print(f"Total MaxSim Score: {total_score:.2f}")
|
||||||
|
print(f"\nPer-token contributions:")
|
||||||
|
|
||||||
|
# Show contribution of each token
|
||||||
|
for i, (max_sim, token_sims) in enumerate(zip(max_sims, similarities)):
|
||||||
|
# Find which patch this token matched best with
|
||||||
|
best_patch_idx = token_sims[:1024].argmax()
|
||||||
|
row, col = best_patch_idx // 32, best_patch_idx % 32
|
||||||
|
# Display actual token text (fall back to index if out of range)
|
||||||
|
token_str = tokens[i] if i < len(tokens) else f"[pad_{i}]"
|
||||||
|
print(f" '{token_str}': score={max_sim:.3f}, best match at patch ({row}, {col})")
|
||||||
|
|
||||||
|
return total_score, max_sims
|
||||||
|
|
||||||
|
# Debug the search result
|
||||||
|
score, token_scores = debug_search_result(
|
||||||
|
"images/financial-report.png",
|
||||||
|
"bar chart showing revenue",
|
||||||
|
model,
|
||||||
|
tokenizer
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
This analysis shows you:
|
||||||
|
- The total relevance score
|
||||||
|
- How much each query token contributes
|
||||||
|
- Where in the document each token found its best match
|
||||||
|
|
||||||
|
If a token like "revenue" is matching in an unexpected location, the visualization reveals whether the model is correctly identifying revenue-related content or making an error.
|
||||||
|
|
||||||
|
### Interpreting the Results
|
||||||
|
|
||||||
|
When analyzing heatmaps:
|
||||||
|
|
||||||
|
- **Concentrated heat**: The token is focusing on a specific region - good for precise matches
|
||||||
|
- **Diffuse heat**: The token finds multiple relevant regions or isn't strongly matched anywhere
|
||||||
|
- **Unexpected locations**: May indicate the model is matching based on visual similarity rather than semantic meaning
|
||||||
|
|
||||||
|
## Aggregated MaxSim Visualization
|
||||||
|
|
||||||
|
Sometimes you want to see which patches contribute most to the overall score, regardless of which query token they matched. This **aggregated MaxSim view** shows the document-level relevance:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def compute_maxsim_contribution(query_vectors, doc_vectors):
|
||||||
|
"""Compute how much each patch contributes to the MaxSim score."""
|
||||||
|
# Use only patch embeddings
|
||||||
|
patch_vectors = doc_vectors[:1024]
|
||||||
|
|
||||||
|
# Compute all pairwise similarities
|
||||||
|
similarities = np.dot(query_vectors, patch_vectors.T) # (n_query, 1024)
|
||||||
|
|
||||||
|
# For each patch, take the maximum contribution across all query tokens
|
||||||
|
# This shows which patches are most "useful" for any query token
|
||||||
|
max_contribution = similarities.max(axis=0) # (1024,)
|
||||||
|
|
||||||
|
# Reshape to spatial grid
|
||||||
|
contribution_map = max_contribution.reshape(32, 32)
|
||||||
|
|
||||||
|
return contribution_map
|
||||||
|
|
||||||
|
def visualize_maxsim_contribution(image_path, query, model):
|
||||||
|
"""Visualize which patches contribute most to the MaxSim score."""
|
||||||
|
# Generate embeddings
|
||||||
|
doc_vectors = next(model.embed_image([image_path]))
|
||||||
|
query_vectors = next(model.embed_text(query))
|
||||||
|
|
||||||
|
# Compute contribution map
|
||||||
|
contribution_map = compute_maxsim_contribution(query_vectors, doc_vectors)
|
||||||
|
|
||||||
|
# Create visualization
|
||||||
|
original_image = Image.open(image_path).convert("RGB").resize((448, 448))
|
||||||
|
overlay, _ = create_heatmap_overlay(original_image, contribution_map)
|
||||||
|
|
||||||
|
fig, axes = plt.subplots(1, 2, figsize=(10, 5))
|
||||||
|
|
||||||
|
axes[0].imshow(original_image)
|
||||||
|
axes[0].set_title("Original Document")
|
||||||
|
axes[0].axis("off")
|
||||||
|
|
||||||
|
axes[1].imshow(overlay)
|
||||||
|
axes[1].set_title(f"MaxSim Contribution\nQuery: \"{query}\"")
|
||||||
|
axes[1].axis("off")
|
||||||
|
|
||||||
|
plt.tight_layout()
|
||||||
|
plt.show()
|
||||||
|
|
||||||
|
# Visualize overall contribution
|
||||||
|
visualize_maxsim_contribution(
|
||||||
|
"images/einstein-newspaper.jpg",
|
||||||
|
"When did Einstein die?",
|
||||||
|
model
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
This aggregated view is particularly useful for:
|
||||||
|
- **Understanding document-level relevance**: See which regions make this document match the query
|
||||||
|
- **Identifying key content**: Highlights the most semantically important patches
|
||||||
|
- **Quality assessment**: Check if the model focuses on relevant content (text, diagrams) rather than noise
|
||||||
|
|
||||||
|
## A Note on Newer Architectures
|
||||||
|
|
||||||
|
The interpretability techniques we've covered work directly with ColPali because of its simple spatial mapping: 448×448 pixels -> 32×32 patches -> 1024 embeddings. Each patch index maps directly to a spatial location.
|
||||||
|
|
||||||
|
However, **newer architectures use more complex image processing** that makes precise visualization more challenging.
|
||||||
|
|
||||||
|
### Split-Image Processing
|
||||||
|
|
||||||
|
Models like ColModernVBERT and ColIdefics3 use a **split-image approach**:
|
||||||
|
|
||||||
|
1. **Resize**: The image is resized to fit a maximum edge constraint (e.g., 1344 pixels)
|
||||||
|
2. **Split into sub-patches**: The resized image is divided into fixed-size sub-patches (typically 512×512 pixels)
|
||||||
|
3. **Token grids per sub-patch**: Each sub-patch becomes a grid of tokens (e.g., 8×8 = 64 tokens)
|
||||||
|
4. **Global patch**: A downscaled view of the entire image is appended as a final set of tokens
|
||||||
|
|
||||||
|
This means tokens arrive in **sub-patch-sequential order** rather than row-major spatial order. To reconstruct spatial correspondence for visualization, you need to:
|
||||||
|
|
||||||
|
1. Exclude the global patch tokens (they lack spatial correspondence to specific regions)
|
||||||
|
2. Rearrange tokens from sub-patch order back to a 2D spatial grid
|
||||||
|
3. Account for varying image dimensions (different images produce different numbers of sub-patches)
|
||||||
|
|
||||||
|
The core insight remains: **multi-vector representations enable interpretability** because each embedding has semantic meaning. The mapping from embedding to image location just becomes more involved with advanced architectures.
|
||||||
|
|
||||||
|
## What's Next
|
||||||
|
|
||||||
|
You've now learned one of ColPali's most powerful features: the ability to see exactly where the model focuses when matching queries to documents. This transparency helps you:
|
||||||
|
|
||||||
|
- Debug unexpected search results
|
||||||
|
- Build trust in your retrieval system
|
||||||
|
- Understand model behavior and limitations
|
||||||
|
- Validate that the model focuses on relevant content
|
||||||
|
|
||||||
|
With a solid understanding of how ColPali works, the model variants available, and how to interpret what the model sees, you're ready to tackle the next challenge: **making these systems production-ready**.
|
||||||
|
|
||||||
|
In Module 3, we'll explore the scalability and optimization techniques you need for real-world deployments - from memory optimization and quantization to multi-stage retrieval pipelines that can handle millions of documents efficiently.
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
---
|
||||||
|
title: "Module 3: Scalability and Optimization"
|
||||||
|
description: "Address scalability challenges in multi-vector search. Learn optimization techniques including quantization, pooling, MUVERA, and multi-stage retrieval."
|
||||||
|
isLesson: true
|
||||||
|
weight: 40
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 3 {{< /date >}}
|
||||||
|
|
||||||
|
# Scalability and Optimization
|
||||||
|
|
||||||
|
Tackle the memory and performance challenges of production-scale multi-vector search.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Today's path
|
||||||
|
|
||||||
|
1. Multi-Stage Retrieval with Universal Query API
|
||||||
|
2. Vector Quantization Techniques
|
||||||
|
3. Pooling Techniques
|
||||||
|
4. MUVERA
|
||||||
|
5. Evaluating Search Pipelines
|
||||||
|
6. Final Project
|
||||||
|
|
||||||
|
You'll master the optimization strategies needed to deploy multi-vector search at scale. The module concludes with a final project where you'll apply all the learned skills to build a real-world multi-vector search system for multi-modal data.
|
||||||
@@ -0,0 +1,589 @@
|
|||||||
|
---
|
||||||
|
title: "Evaluating Search Pipelines"
|
||||||
|
description: Learn how to evaluate different search configurations in terms of cost, latency, and retrieval quality using ground truth datasets and standardized metrics.
|
||||||
|
weight: 5
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 3 {{< /date >}}
|
||||||
|
|
||||||
|
# Evaluating Search Pipelines
|
||||||
|
|
||||||
|
Throughout this module, you've learned many optimization techniques: quantization to reduce memory, pooling to compress representations, MUVERA for efficient indexing, and multi-stage retrieval to balance speed with accuracy. But how do you know which combination is right for *your* data?
|
||||||
|
|
||||||
|
The answer lies in systematic evaluation across three dimensions: **cost** (memory and compute resources), **latency** (query response time), and **quality** (retrieval accuracy). Cost and latency are straightforward to measure - you can observe memory usage and time queries directly. Quality, however, requires a more principled approach: you need to measure whether your system returns the *right* documents.
|
||||||
|
|
||||||
|
> **Note:** This lesson demonstrates evaluation methodology on a **small, comprehensible dataset** that you can manually inspect and understand. We use 4 document images with 8 queries where you can verify relevance judgments yourself. Real production benchmarks would use larger datasets like ViDoRe, but the methodology remains the same.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/1gHbp9c01iE?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-3/evaluating-pipelines.ipynb">
|
||||||
|
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||||
|
</a>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Quick Reference: Evaluation Metrics
|
||||||
|
|
||||||
|
Before diving in, here's the essential guide to choosing metrics. For multi-stage pipelines, focus on **Recall@k for the prefetch stage** (did we capture candidates?) and **NDCG@k for the final ranking** (did we order them correctly?).
|
||||||
|
|
||||||
|
| Metric | When to Use | Key Insight |
|
||||||
|
|-----------------|----------------------------|----------------------------------------------|
|
||||||
|
| **Recall@k** | Prefetch stage (k=50, 100) | Did we capture relevant candidates? |
|
||||||
|
| **NDCG@k** | Final ranking (k=5, 10) | Are most relevant results ranked highest? |
|
||||||
|
| **MRR** | Single-answer retrieval | How quickly do we find THE right answer? |
|
||||||
|
| **Precision@k** | Result page quality | What fraction of shown results are relevant? |
|
||||||
|
|
||||||
|
## Ground Truth: Defining "Correct" Retrieval
|
||||||
|
|
||||||
|
To measure retrieval quality, you need **relevance judgments** (called **qrels**) - data that defines which documents are relevant for which queries.
|
||||||
|
|
||||||
|
### What Are Qrels?
|
||||||
|
|
||||||
|
Each qrel is a triplet: a query, a document, and a relevance score indicating how relevant that document is to the query.
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Ground truth: which documents are relevant to which queries?
|
||||||
|
qrels_dict = {
|
||||||
|
"company quarterly financial results and revenue": {
|
||||||
|
"images/financial-report.png": 3, # Highly relevant
|
||||||
|
},
|
||||||
|
"historic ship disaster at sea": {
|
||||||
|
"images/titanic-newspaper.jpg": 3, # Highly relevant
|
||||||
|
},
|
||||||
|
"news headline from early 1900s": {
|
||||||
|
"images/titanic-newspaper.jpg": 3, # Highly relevant
|
||||||
|
"images/einstein-newspaper.jpg": 2, # Somewhat relevant
|
||||||
|
},
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Relevance can be **binary** (0 = not relevant, 1 = relevant) or **graded** (0-3 scale where higher means more relevant). Graded relevance provides more nuance for ranking evaluation.
|
||||||
|
|
||||||
|
### Building Ground Truth
|
||||||
|
|
||||||
|
Three practical approaches:
|
||||||
|
|
||||||
|
1. **Manual Annotation** - Have domain experts review query-document pairs. Highest quality but time-intensive. Even 50-100 queries provides valuable signal.
|
||||||
|
|
||||||
|
2. **Synthetic Generation with LLMs** - Use language models to generate relevant queries for documents. Scales well but may not capture real user query patterns.
|
||||||
|
|
||||||
|
3. **Existing Benchmarks** - For visual document retrieval, **ViDoRe** provides standardized evaluation. For text retrieval, **BEIR** offers diverse domains.
|
||||||
|
|
||||||
|
For this lesson, we'll use a small manually-annotated dataset where you can inspect the relevance judgments yourself.
|
||||||
|
|
||||||
|
## The Degrees of Freedom Problem
|
||||||
|
|
||||||
|
With so many optimization techniques, the configuration space explodes quickly:
|
||||||
|
|
||||||
|
| Dimension | Options |
|
||||||
|
|------------------|---------------------------------------------------------------------|
|
||||||
|
| **Quantization** | None, Scalar (int8), Binary (1-bit, 1.5-bit, 2-bit), Product |
|
||||||
|
| **Pooling** | None, Hierarchical (k=16, 32, 64, ...) |
|
||||||
|
| **MUVERA** | Disabled, or Enabled (differenr k_sim, dim_proj, r_reps variations) |
|
||||||
|
| **Multi-stage** | Single-stage, Two-stage (prefetch: 50, 100, 500, ...) |
|
||||||
|
|
||||||
|
Just considering these options yields hundreds of possible configurations. The key insight: **you don't need to test them all**. Instead, you need a systematic approach to test *representative* configurations efficiently.
|
||||||
|
|
||||||
|
## Unified Collection Architecture
|
||||||
|
|
||||||
|
The traditional approach creates a separate collection for each configuration - tedious and storage-intensive. A better approach: **store all vector representations in a single collection** using named vectors.
|
||||||
|
|
||||||
|
### Four Named Vectors
|
||||||
|
|
||||||
|
We'll store four different representations of each document:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
from qdrant_client.models import (
|
||||||
|
VectorParams, Distance, MultiVectorConfig, MultiVectorComparator,
|
||||||
|
ScalarQuantization, ScalarQuantizationConfig, ScalarType,
|
||||||
|
)
|
||||||
|
|
||||||
|
client = QdrantClient(url="http://localhost:6333")
|
||||||
|
|
||||||
|
COLLECTION_NAME = "eval-multi-vector"
|
||||||
|
|
||||||
|
client.create_collection(
|
||||||
|
collection_name=COLLECTION_NAME,
|
||||||
|
vectors_config={
|
||||||
|
# Full ColModernVBERT multi-vector (no quantization)
|
||||||
|
"colmodernvbert": VectorParams(
|
||||||
|
size=128,
|
||||||
|
distance=Distance.DOT,
|
||||||
|
multivector_config=MultiVectorConfig(
|
||||||
|
comparator=MultiVectorComparator.MAX_SIM
|
||||||
|
),
|
||||||
|
hnsw_config=models.HnswConfigDiff(m=0), # Disable HNSW for multi-vector
|
||||||
|
),
|
||||||
|
# ColModernVBERT with scalar quantization enabled
|
||||||
|
"colmodernvbert_sq": VectorParams(
|
||||||
|
size=128,
|
||||||
|
distance=Distance.DOT,
|
||||||
|
multivector_config=MultiVectorConfig(
|
||||||
|
comparator=MultiVectorComparator.MAX_SIM
|
||||||
|
),
|
||||||
|
hnsw_config=models.HnswConfigDiff(m=0),
|
||||||
|
quantization_config=ScalarQuantization(
|
||||||
|
scalar=ScalarQuantizationConfig(
|
||||||
|
type=ScalarType.INT8,
|
||||||
|
quantile=0.99,
|
||||||
|
always_ram=True,
|
||||||
|
)
|
||||||
|
),
|
||||||
|
),
|
||||||
|
# MUVERA single-vector approximation for fast HNSW search
|
||||||
|
"muvera": VectorParams(
|
||||||
|
size=40960, # muvera.embedding_size from k_sim=6, dim_proj=32, r_reps=20
|
||||||
|
distance=Distance.COSINE,
|
||||||
|
),
|
||||||
|
# Hierarchical pooled multi-vector (k=32 clusters)
|
||||||
|
"hierarchical": VectorParams(
|
||||||
|
size=128,
|
||||||
|
distance=Distance.DOT,
|
||||||
|
multivector_config=MultiVectorConfig(
|
||||||
|
comparator=MultiVectorComparator.MAX_SIM
|
||||||
|
),
|
||||||
|
hnsw_config=models.HnswConfigDiff(m=0),
|
||||||
|
),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Key insight:** The `colmodernvbert_sq` vector stores the same embeddings as `colmodernvbert` but with scalar quantization enabled. This allows direct comparison of quantized vs. non-quantized search without re-indexing.
|
||||||
|
|
||||||
|
### Embedding and Uploading Documents
|
||||||
|
|
||||||
|
A single function generates all four representations for each document:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import LateInteractionMultimodalEmbedding
|
||||||
|
from scipy.cluster.vq import kmeans2
|
||||||
|
import numpy as np
|
||||||
|
from fastembed.postprocess import Muvera
|
||||||
|
|
||||||
|
# Load the embedding model
|
||||||
|
model = LateInteractionMultimodalEmbedding(
|
||||||
|
model_name="Qdrant/colmodernvbert"
|
||||||
|
)
|
||||||
|
|
||||||
|
# Initialize MUVERA with the same configuration as the collection
|
||||||
|
muvera = Muvera.from_multivector_model(model=model, k_sim=6, dim_proj=32, r_reps=20)
|
||||||
|
|
||||||
|
|
||||||
|
def hierarchical_pool(embeddings: np.ndarray, k: int = 32) -> np.ndarray:
|
||||||
|
"""Pool multi-vector to k centroids using k-means clustering."""
|
||||||
|
if len(embeddings) <= k:
|
||||||
|
return embeddings # No pooling needed
|
||||||
|
centroids, labels = kmeans2(embeddings.astype(np.float64), k, minit="++")
|
||||||
|
# Return mean of embeddings in each cluster
|
||||||
|
pooled = np.array([
|
||||||
|
embeddings[labels == i].mean(axis=0)
|
||||||
|
for i in range(k)
|
||||||
|
if (labels == i).any()
|
||||||
|
])
|
||||||
|
return pooled.astype(np.float32)
|
||||||
|
|
||||||
|
|
||||||
|
def embed_and_upload_document(doc_path: str, doc_id: int) -> None:
|
||||||
|
"""Embed a document and upload all four vector representations."""
|
||||||
|
# Generate full multi-vector embeddings
|
||||||
|
full_multivec = np.array(list(model.embed_image([doc_path]))[0])
|
||||||
|
|
||||||
|
# Generate MUVERA approximation
|
||||||
|
muvera_vec = muvera.process_document(full_multivec)
|
||||||
|
|
||||||
|
# Generate hierarchical pooled version (k=32)
|
||||||
|
hierarchical_vec = hierarchical_pool(full_multivec, k=32)
|
||||||
|
|
||||||
|
# Upload all representations in one point
|
||||||
|
client.upsert(
|
||||||
|
collection_name=COLLECTION_NAME,
|
||||||
|
points=[
|
||||||
|
models.PointStruct(
|
||||||
|
id=doc_id,
|
||||||
|
payload={"filename": doc_path},
|
||||||
|
vector={
|
||||||
|
"colmodernvbert": full_multivec.tolist(),
|
||||||
|
"colmodernvbert_sq": full_multivec.tolist(), # Same data, quantized config
|
||||||
|
"muvera": muvera_vec.tolist(),
|
||||||
|
"hierarchical": hierarchical_vec.tolist(),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
],
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Sample Dataset
|
||||||
|
|
||||||
|
For this lesson, we use a small dataset you can manually inspect:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Small dataset of document images you can manually inspect
|
||||||
|
DOC_PATHS = [
|
||||||
|
"images/financial-report.png",
|
||||||
|
"images/titanic-newspaper.jpg",
|
||||||
|
"images/men-walk-on-moon-newspaper.jpg",
|
||||||
|
"images/einstein-newspaper.jpg",
|
||||||
|
]
|
||||||
|
|
||||||
|
# Upload all documents with all 4 vector representations
|
||||||
|
for doc_id, doc_path in enumerate(DOC_PATHS):
|
||||||
|
embed_and_upload_document(doc_path, doc_id)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Building Evaluation Pipelines
|
||||||
|
|
||||||
|
With all vectors stored in a single collection, we can build different pipelines **without re-indexing** - we simply query different named vectors.
|
||||||
|
|
||||||
|
### The Search Function
|
||||||
|
|
||||||
|
```python
|
||||||
|
def search_pipeline(
|
||||||
|
query_embedding: np.ndarray,
|
||||||
|
using: str,
|
||||||
|
prefetch_using: str | None = None,
|
||||||
|
prefetch_limit: int = 50,
|
||||||
|
limit: int = 10,
|
||||||
|
) -> list[tuple[str, float]]:
|
||||||
|
"""
|
||||||
|
Execute a search pipeline with optional prefetch stage.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
query_embedding: The query's multi-vector embedding
|
||||||
|
using: Named vector for final ranking
|
||||||
|
prefetch_using: Named vector for prefetch (None = single-stage)
|
||||||
|
prefetch_limit: How many candidates to retrieve in prefetch
|
||||||
|
limit: Final number of results
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
List of (filename, score) tuples
|
||||||
|
"""
|
||||||
|
if prefetch_using is None:
|
||||||
|
# Single-stage search
|
||||||
|
response = client.query_points(
|
||||||
|
collection_name=COLLECTION_NAME,
|
||||||
|
query=query_embedding.tolist(),
|
||||||
|
using=using,
|
||||||
|
limit=limit,
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
# Two-stage search: prefetch with one vector, rerank with another
|
||||||
|
# For MUVERA prefetch, we need the MUVERA query embedding
|
||||||
|
if prefetch_using == "muvera":
|
||||||
|
prefetch_query = muvera.process_query(query_embedding).tolist()
|
||||||
|
else:
|
||||||
|
prefetch_query = query_embedding.tolist()
|
||||||
|
|
||||||
|
response = client.query_points(
|
||||||
|
collection_name=COLLECTION_NAME,
|
||||||
|
prefetch=[
|
||||||
|
models.Prefetch(
|
||||||
|
query=prefetch_query,
|
||||||
|
using=prefetch_using,
|
||||||
|
limit=prefetch_limit,
|
||||||
|
)
|
||||||
|
],
|
||||||
|
query=query_embedding.tolist(),
|
||||||
|
using=using,
|
||||||
|
limit=limit,
|
||||||
|
)
|
||||||
|
|
||||||
|
return [
|
||||||
|
(point.payload["filename"], point.score)
|
||||||
|
for point in response.points
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
### Representative Pipeline Configurations
|
||||||
|
|
||||||
|
Rather than testing all combinations, we select **6 representative pipelines** that cover the key trade-offs:
|
||||||
|
|
||||||
|
```python
|
||||||
|
PIPELINES = {
|
||||||
|
# Baseline: full quality, no optimization
|
||||||
|
"baseline": {
|
||||||
|
"using": "colmodernvbert",
|
||||||
|
"prefetch_using": None,
|
||||||
|
},
|
||||||
|
|
||||||
|
# Scalar quantized: reduced memory, minimal quality loss
|
||||||
|
"scalar_quantized": {
|
||||||
|
"using": "colmodernvbert_sq",
|
||||||
|
"prefetch_using": None,
|
||||||
|
},
|
||||||
|
|
||||||
|
# Hierarchical pooling: fewer vectors per document
|
||||||
|
"hierarchical": {
|
||||||
|
"using": "hierarchical",
|
||||||
|
"prefetch_using": None,
|
||||||
|
},
|
||||||
|
|
||||||
|
# Two-stage: fast MUVERA prefetch + full quality rerank
|
||||||
|
"muvera_rerank": {
|
||||||
|
"using": "colmodernvbert",
|
||||||
|
"prefetch_using": "muvera",
|
||||||
|
"prefetch_limit": 50,
|
||||||
|
},
|
||||||
|
|
||||||
|
# Two-stage with quantized rerank
|
||||||
|
"muvera_quantized": {
|
||||||
|
"using": "colmodernvbert_sq",
|
||||||
|
"prefetch_using": "muvera",
|
||||||
|
"prefetch_limit": 50,
|
||||||
|
},
|
||||||
|
|
||||||
|
# Maximum compression: MUVERA prefetch + pooled rerank
|
||||||
|
"muvera_hierarchical": {
|
||||||
|
"using": "hierarchical",
|
||||||
|
"prefetch_using": "muvera",
|
||||||
|
"prefetch_limit": 50,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Key insight:** All six pipelines query the **same indexed data** - we just configure which named vectors to use for prefetch and final ranking.
|
||||||
|
|
||||||
|
### Running All Pipelines
|
||||||
|
|
||||||
|
```python
|
||||||
|
def evaluate_all_pipelines(
|
||||||
|
queries: dict[str, np.ndarray],
|
||||||
|
qrels: dict,
|
||||||
|
) -> dict[str, dict]:
|
||||||
|
"""Run all pipeline configurations and collect results."""
|
||||||
|
from ranx import Qrels, Run, evaluate
|
||||||
|
import time
|
||||||
|
|
||||||
|
ranx_qrels = Qrels(qrels)
|
||||||
|
results = {}
|
||||||
|
|
||||||
|
for pipeline_name, config in PIPELINES.items():
|
||||||
|
# Collect search results for all queries
|
||||||
|
pipeline_results = {}
|
||||||
|
latencies = []
|
||||||
|
|
||||||
|
for query_text, query_embedding in queries.items():
|
||||||
|
start = time.perf_counter()
|
||||||
|
search_results = search_pipeline(query_embedding, **config)
|
||||||
|
latencies.append((time.perf_counter() - start) * 1000)
|
||||||
|
|
||||||
|
# Convert to ranx format: {doc_id: score}
|
||||||
|
pipeline_results[query_text] = {
|
||||||
|
filename: score for filename, score in search_results
|
||||||
|
}
|
||||||
|
|
||||||
|
# Create ranx Run and evaluate
|
||||||
|
run = Run(pipeline_results, name=pipeline_name)
|
||||||
|
metrics = evaluate(ranx_qrels, run, ["ndcg@10", "recall@10", "mrr"])
|
||||||
|
|
||||||
|
results[pipeline_name] = {
|
||||||
|
"metrics": metrics,
|
||||||
|
"avg_latency_ms": np.mean(latencies),
|
||||||
|
}
|
||||||
|
|
||||||
|
return results
|
||||||
|
```
|
||||||
|
|
||||||
|
## Evaluation with ranx
|
||||||
|
|
||||||
|
The **ranx** library provides battle-tested implementations of IR metrics. Here's how to compare all our pipelines:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from ranx import Qrels, Run, compare
|
||||||
|
import time
|
||||||
|
|
||||||
|
# Define ground truth - which documents are relevant to which queries
|
||||||
|
qrels_dict = {
|
||||||
|
"company quarterly financial results and revenue": {
|
||||||
|
"images/financial-report.png": 3, # Highly relevant
|
||||||
|
},
|
||||||
|
"historic ship disaster at sea": {
|
||||||
|
"images/titanic-newspaper.jpg": 3,
|
||||||
|
},
|
||||||
|
"space exploration and astronauts": {
|
||||||
|
"images/men-walk-on-moon-newspaper.jpg": 3,
|
||||||
|
},
|
||||||
|
"physics theory and scientist": {
|
||||||
|
"images/einstein-newspaper.jpg": 3,
|
||||||
|
},
|
||||||
|
"news headline from early 1900s": {
|
||||||
|
"images/titanic-newspaper.jpg": 3,
|
||||||
|
"images/einstein-newspaper.jpg": 2, # Somewhat relevant
|
||||||
|
},
|
||||||
|
"business earnings report": {
|
||||||
|
"images/financial-report.png": 3,
|
||||||
|
},
|
||||||
|
"NASA moon landing mission": {
|
||||||
|
"images/men-walk-on-moon-newspaper.jpg": 3,
|
||||||
|
},
|
||||||
|
"ocean liner sinking": {
|
||||||
|
"images/titanic-newspaper.jpg": 3,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
qrels = Qrels(qrels_dict)
|
||||||
|
|
||||||
|
# Generate query embeddings
|
||||||
|
query_embeddings = {
|
||||||
|
query: np.array(list(model.embed_text([query]))[0])
|
||||||
|
for query in qrels_dict.keys()
|
||||||
|
}
|
||||||
|
|
||||||
|
# Collect runs from each pipeline
|
||||||
|
runs = []
|
||||||
|
latency_results = {}
|
||||||
|
|
||||||
|
for pipeline_name, config in PIPELINES.items():
|
||||||
|
pipeline_results = {}
|
||||||
|
latencies = []
|
||||||
|
|
||||||
|
for query_text, query_embedding in query_embeddings.items():
|
||||||
|
start = time.perf_counter()
|
||||||
|
search_results = search_pipeline(query_embedding, **config, limit=10)
|
||||||
|
latencies.append((time.perf_counter() - start) * 1000)
|
||||||
|
|
||||||
|
# Convert to ranx format: {doc_id: score}
|
||||||
|
pipeline_results[query_text] = {
|
||||||
|
filename: score for filename, score in search_results
|
||||||
|
}
|
||||||
|
|
||||||
|
runs.append(Run(pipeline_results, name=pipeline_name))
|
||||||
|
latency_results[pipeline_name] = np.mean(latencies)
|
||||||
|
|
||||||
|
# Compare all pipelines
|
||||||
|
report = compare(
|
||||||
|
qrels=qrels,
|
||||||
|
runs=runs,
|
||||||
|
metrics=["ndcg@10", "recall@10", "mrr"],
|
||||||
|
max_p=0.05, # Statistical significance threshold
|
||||||
|
)
|
||||||
|
print(report)
|
||||||
|
```
|
||||||
|
|
||||||
|
This produces a formatted comparison table showing how each pipeline performs across all metrics, with statistical significance indicators.
|
||||||
|
|
||||||
|
## Choosing a Pipeline: Trade-off Analysis
|
||||||
|
|
||||||
|
With multiple pipelines evaluated across quality and latency, how do you choose? The concept of **Pareto optimality** helps frame the decision. A pipeline is Pareto-optimal if no other pipeline beats it in *all* dimensions simultaneously. For example, `muvera_rerank` might offer 94% of baseline quality at 4x the speed - you can't get both faster *and* higher quality without trade-offs.
|
||||||
|
|
||||||
|
When analyzing your results, plot quality (NDCG@10) against latency or memory usage. Pipelines that fall below the "frontier" of best options are **dominated** - there's always a better choice available. Focus your attention on configurations along this frontier, then choose based on your constraints: quality-first applications should stay closer to baseline, while latency-critical systems can move toward faster approximations like `muvera_hierarchical`.
|
||||||
|
|
||||||
|
In practice, multi-stage pipelines with MUVERA prefetch often provide the best balance - they leverage fast HNSW search for candidate retrieval while preserving full multi-vector quality for final ranking. Start with `muvera_rerank` as a strong default, then adjust the prefetch limit based on your recall requirements.
|
||||||
|
|
||||||
|
## Putting It All Together
|
||||||
|
|
||||||
|
Here's the complete evaluation workflow:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import LateInteractionMultimodalEmbedding
|
||||||
|
from fastembed.postprocess import Muvera
|
||||||
|
from ranx import Qrels, Run, compare
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
# 1. Load model and initialize MUVERA
|
||||||
|
model = LateInteractionMultimodalEmbedding(
|
||||||
|
model_name="Qdrant/colmodernvbert"
|
||||||
|
)
|
||||||
|
muvera = Muvera.from_multivector_model(model=model, k_sim=6, dim_proj=32, r_reps=20)
|
||||||
|
|
||||||
|
# 2. Create unified collection with all named vectors
|
||||||
|
# (See collection creation code above)
|
||||||
|
|
||||||
|
# 3. Embed and upload documents (generates all 4 representations)
|
||||||
|
DOC_PATHS = [
|
||||||
|
"images/financial-report.png",
|
||||||
|
"images/titanic-newspaper.jpg",
|
||||||
|
"images/men-walk-on-moon-newspaper.jpg",
|
||||||
|
"images/einstein-newspaper.jpg",
|
||||||
|
]
|
||||||
|
|
||||||
|
for doc_id, doc_path in enumerate(DOC_PATHS):
|
||||||
|
embed_and_upload_document(doc_path, doc_id)
|
||||||
|
|
||||||
|
# 4. Define ground truth (qrels)
|
||||||
|
qrels_dict = {
|
||||||
|
"company quarterly financial results and revenue": {
|
||||||
|
"images/financial-report.png": 3,
|
||||||
|
},
|
||||||
|
"historic ship disaster at sea": {
|
||||||
|
"images/titanic-newspaper.jpg": 3,
|
||||||
|
},
|
||||||
|
# ... more queries with relevance judgments
|
||||||
|
}
|
||||||
|
qrels = Qrels(qrels_dict)
|
||||||
|
|
||||||
|
# 5. Embed queries
|
||||||
|
query_embeddings = {
|
||||||
|
query: np.array(list(model.embed_text([query]))[0])
|
||||||
|
for query in qrels_dict.keys()
|
||||||
|
}
|
||||||
|
|
||||||
|
# 6. Evaluate all pipeline configurations and collect runs
|
||||||
|
runs = []
|
||||||
|
for pipeline_name, config in PIPELINES.items():
|
||||||
|
pipeline_results = {}
|
||||||
|
for query_text, query_embedding in query_embeddings.items():
|
||||||
|
search_results = search_pipeline(query_embedding, **config, limit=10)
|
||||||
|
pipeline_results[query_text] = {
|
||||||
|
filename: score for filename, score in search_results
|
||||||
|
}
|
||||||
|
runs.append(Run(pipeline_results, name=pipeline_name))
|
||||||
|
|
||||||
|
# 7. Compare all pipelines
|
||||||
|
report = compare(
|
||||||
|
qrels=qrels,
|
||||||
|
runs=runs,
|
||||||
|
metrics=["ndcg@10", "recall@10", "mrr"],
|
||||||
|
)
|
||||||
|
print(report)
|
||||||
|
|
||||||
|
# 8. Select pipeline based on requirements:
|
||||||
|
# - Quality-first? Use baseline or muvera_rerank
|
||||||
|
# - Latency-critical? Use muvera_hierarchical
|
||||||
|
# - Balanced? Use muvera_quantized
|
||||||
|
```
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
Evaluating multi-vector search pipelines requires:
|
||||||
|
|
||||||
|
1. **Ground truth data** (qrels) that defines what "correct" retrieval means
|
||||||
|
2. **A unified collection architecture** that stores all vector representations together
|
||||||
|
3. **Representative pipeline configurations** that cover the key trade-offs
|
||||||
|
4. **Appropriate metrics** chosen for your use case (NDCG for ranking, Recall for prefetch)
|
||||||
|
5. **Trade-off analysis** using Pareto frontiers to identify optimal configurations
|
||||||
|
|
||||||
|
The key insight from this lesson: **you don't need separate collections for each configuration**. By storing multiple named vectors (full multi-vector, quantized, MUVERA, hierarchical pooled), you can evaluate many pipeline configurations efficiently without re-indexing.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Course Conclusion
|
||||||
|
|
||||||
|
Congratulations on completing the Multi-Vector Search course!
|
||||||
|
|
||||||
|
You've journeyed from the foundations of late interaction and MaxSim distance, through multi-modal applications with ColPali, to production optimization techniques. You now have the knowledge to:
|
||||||
|
|
||||||
|
- **Understand** why multi-vector representations capture richer semantic relationships
|
||||||
|
- **Implement** multi-vector search with Qdrant's native support
|
||||||
|
- **Apply** these techniques to visual document retrieval using ColPali models
|
||||||
|
- **Optimize** for production with quantization, pooling, MUVERA, and multi-stage retrieval
|
||||||
|
- **Evaluate** different configurations to find the right trade-offs for your use case
|
||||||
|
|
||||||
|
Multi-vector search is a powerful paradigm that's particularly well-suited for complex retrieval tasks where single-vector representations fall short. As you apply these techniques to your own projects, remember that the best configuration depends on your specific data, queries, and constraints. Use the evaluation framework from this lesson to make data-driven decisions.
|
||||||
|
|
||||||
|
We encourage you to experiment with your own datasets, try different model variants, and share what you learn with the community. The field of multi-vector search is evolving rapidly, and practical insights from real-world applications are invaluable.
|
||||||
|
|
||||||
|
Happy searching!
|
||||||
@@ -0,0 +1,146 @@
|
|||||||
|
---
|
||||||
|
title: "Final Project: Build Your Own Multi-Vector Search System"
|
||||||
|
description: Apply everything you've learned to build a multi-vector search system that solves a real problem of your choosing.
|
||||||
|
weight: 7
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 3 {{< /date >}}
|
||||||
|
|
||||||
|
# Final Project: Build Your Own Multi-Vector Search System
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Your Mission
|
||||||
|
|
||||||
|
It's time to bring together everything you've learned about multi-vector search, late interaction models, and production optimization. You'll build a sophisticated document retrieval system that leverages late interaction's token-level matching for superior search quality.
|
||||||
|
|
||||||
|
Your search engine will understand the nuanced relationships between query terms and document content. When someone searches for "machine learning applications in healthcare," your system will find documents that discuss relevant concepts even when they use different terminology, thanks to late interaction's fine-grained matching.
|
||||||
|
|
||||||
|
This mirrors real-world challenges in enterprise search, research libraries, and knowledge management. You'll implement the complete pipeline: multi-vector embedding using **ColModernVBERT** or **ColPali**, memory-efficient storage with quantization or pooling, and optimized retrieval with multi-stage search.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## What You'll Build
|
||||||
|
|
||||||
|
A working multi-vector search system that:
|
||||||
|
|
||||||
|
- Indexes real documents using late interaction embeddings
|
||||||
|
- Applies optimization techniques to manage memory and latency
|
||||||
|
- Retrieves relevant results with measurable quality
|
||||||
|
- Documents your design decisions and trade-offs
|
||||||
|
|
||||||
|
Both ColModernVBERT and ColPali work well for visual document understanding - choose based on your preference or experiment with both.
|
||||||
|
|
||||||
|
The specific dataset, optimization strategy, and retrieval configuration are up to you.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Choose Your Challenge
|
||||||
|
|
||||||
|
This is your project. Pick a problem that matters to you.
|
||||||
|
|
||||||
|
### Use Your Own Data
|
||||||
|
|
||||||
|
The most valuable learning comes from working with documents you actually care about:
|
||||||
|
|
||||||
|
- **Technical documentation** you reference frequently
|
||||||
|
- **Research papers** in your field of interest
|
||||||
|
- **Internal documents** (reports, manuals, wikis) from your work
|
||||||
|
- **Personal collection** of PDFs, articles, or notes
|
||||||
|
|
||||||
|
Aim for **at least 50 documents** to have enough variety for meaningful evaluation. More is better for seeing how your system scales.
|
||||||
|
|
||||||
|
### Public Datasets (If You Prefer)
|
||||||
|
|
||||||
|
If you don't have a suitable personal dataset:
|
||||||
|
|
||||||
|
- **ArXiv papers** on a topic you're curious about
|
||||||
|
- **Wikipedia articles** from a category you'd like to explore
|
||||||
|
- **Open-source documentation** from projects you use
|
||||||
|
- **News articles** from a domain you follow
|
||||||
|
|
||||||
|
The key is picking something where you can judge search quality intuitively. You'll need to create evaluation queries, and that's easier when you understand the content.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Project Requirements
|
||||||
|
|
||||||
|
Your project should demonstrate:
|
||||||
|
|
||||||
|
### 1. Working Multi-Vector Search
|
||||||
|
|
||||||
|
**Use ColModernVBERT** to index your documents and retrieve results using MaxSim scoring. The system should return relevant documents for natural language queries. Alternatively, consider **ColPali**.
|
||||||
|
|
||||||
|
### 2. At Least One Optimization Technique
|
||||||
|
|
||||||
|
Apply something you learned in this module:
|
||||||
|
|
||||||
|
- Binary or scalar quantization
|
||||||
|
- Token pooling (clustering, attention-based, or hierarchical)
|
||||||
|
- MUVERA indexing
|
||||||
|
- Multi-stage retrieval pipeline
|
||||||
|
|
||||||
|
Measure the impact of your chosen technique on memory, latency, or search quality.
|
||||||
|
|
||||||
|
### 3. Evaluation with Ground Truth
|
||||||
|
|
||||||
|
Create a test set of queries with known relevant documents. This doesn't need to be exhaustive - 10-20 queries with 3-5 relevant documents each is enough to see meaningful patterns. Measure at least one retrieval metric (precision@k, recall@k, or MRR).
|
||||||
|
|
||||||
|
### 4. Brief Write-Up
|
||||||
|
|
||||||
|
Document your decisions:
|
||||||
|
|
||||||
|
- Why you chose your dataset
|
||||||
|
- Whether you used ColModernVBERT or ColPali, and why
|
||||||
|
- What optimization technique(s) you applied and why
|
||||||
|
- What worked well and what surprised you
|
||||||
|
- Key metrics from your evaluation
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Suggested Approach
|
||||||
|
|
||||||
|
These hints are optional. Feel free to chart your own path.
|
||||||
|
|
||||||
|
**Start simple.** Get a basic multi-vector search working before adding optimizations. It's easier to measure the impact of changes when you have a baseline.
|
||||||
|
|
||||||
|
**Create ground truth early.** Before optimizing, write your evaluation queries and identify relevant documents. This lets you measure whether changes improve or hurt quality.
|
||||||
|
|
||||||
|
**Compare configurations.** Try at least two different setups (e.g., with and without quantization, or different pooling strategies). The comparison will teach you more than a single configuration.
|
||||||
|
|
||||||
|
**Keep notes as you go.** Document what you try and what happens. Your future self will thank you, and it makes the write-up easier.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Share Your Results
|
||||||
|
|
||||||
|
We'd love to see what you build. Share your project on the [Qdrant Discord](https://discord.gg/qdrant) in the [#course-submissions](https://discord.com/channels/907569970500743200/1429673887590776832) channel.
|
||||||
|
|
||||||
|
Tell us about:
|
||||||
|
|
||||||
|
- **Your dataset** and why you chose it
|
||||||
|
- **Your model choice** (ColModernVBERT or ColPali, and why)
|
||||||
|
- **Your collection configuration** (quantization, pooling, indexing)
|
||||||
|
- **Your retrieval metrics** (precision@10, recall, latency)
|
||||||
|
- **What you learned** and what surprised you
|
||||||
|
|
||||||
|
Seeing how others approached the same challenge is one of the best ways to deepen your understanding.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## What You've Accomplished
|
||||||
|
|
||||||
|
By completing this project, you've built a production-ready multi-vector search system that:
|
||||||
|
|
||||||
|
* **Leverages late interaction** for superior search quality through token-level matching
|
||||||
|
* **Scales efficiently** using quantization, pooling, or optimized indexing
|
||||||
|
* **Delivers fast results** through multi-stage retrieval and HNSW optimization
|
||||||
|
* **Measures quality** with comprehensive evaluation metrics
|
||||||
|
* **Documents trade-offs** between accuracy, speed, and memory
|
||||||
|
|
||||||
|
You've mastered the full pipeline from multi-vector embeddings to production optimization - skills directly applicable to enterprise search, document management, research platforms, and AI-powered knowledge systems.
|
||||||
|
|
||||||
|
**Congratulations on completing the Multi-Vector Search course.**
|
||||||
|
|
||||||
|
Continue your learning journey by exploring advanced topics like fine-tuning late interaction models for domain-specific documents, or building Retrieval Augmented Generation pipelines to derive insights from your retrieved documents.
|
||||||
@@ -0,0 +1,334 @@
|
|||||||
|
---
|
||||||
|
title: "Multi-Stage Retrieval with Universal Query API"
|
||||||
|
description: Combine multiple optimization techniques in multi-stage retrieval pipelines using Qdrant's Universal Query API.
|
||||||
|
weight: 1
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 3 {{< /date >}}
|
||||||
|
|
||||||
|
# Multi-Stage Retrieval with Universal Query API
|
||||||
|
|
||||||
|
The most effective production deployments combine multiple optimization techniques in multi-stage pipelines. Fast approximate methods retrieve candidates, which are then reranked with higher-quality methods.
|
||||||
|
|
||||||
|
Qdrant's Universal Query API makes it easy to build sophisticated multi-stage retrieval systems.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/qIjPepsY35E?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-3/multi-stage-retrieval.ipynb">
|
||||||
|
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||||
|
</a>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Why Multi-Stage Retrieval?
|
||||||
|
|
||||||
|
You've learned that multi-vector representations like ColBERT provide superior search quality compared to single-vector embeddings. But there's a challenge: **computing MaxSim for every document in a large collection is expensive**.
|
||||||
|
|
||||||
|
Here's the dilemma: single-vector models are fast but less accurate, while multi-vector models are accurate but computationally intensive. What if you could combine the strengths of both?
|
||||||
|
|
||||||
|
Multi-stage retrieval offers an elegant solution: use a fast method to narrow down candidates, then apply a high-quality method to rerank only those candidates.
|
||||||
|
|
||||||
|
## The Multi-Stage Retrieval Pattern
|
||||||
|
|
||||||
|
The key insight is that you don't need to use your most expensive model on every document in your collection. Instead, you can split the search into stages:
|
||||||
|
|
||||||
|
1. **Stage 1 (Prefetch)**: Use a fast single-vector embedding model to retrieve a large set of candidates
|
||||||
|
2. **Stage 2 (Rerank)**: Use ColBERT's multi-vector representations to rerank only those candidates
|
||||||
|
|
||||||
|
This pattern dramatically reduces computational cost while maintaining high search quality.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
## The Critical Role of Oversampling
|
||||||
|
|
||||||
|
Here's the most important concept in multi-stage retrieval: **you must retrieve more candidates in the prefetch stage than you want in your final results**.
|
||||||
|
|
||||||
|
This is called **oversampling**, and it's essential for maintaining search quality.
|
||||||
|
|
||||||
|
### Why Oversample?
|
||||||
|
|
||||||
|
Consider what happens if you retrieve exactly 10 results in both stages:
|
||||||
|
|
||||||
|
- **Stage 1**: Single-vector model retrieves its "top 10" documents
|
||||||
|
- **Stage 2**: ColBERT reranks those same 10 documents
|
||||||
|
|
||||||
|
The problem? You're limited to ColBERT reranking only the 10 documents the single-vector model selected. If the truly best document is ranked 11th by the single-vector model, ColBERT will never see it.
|
||||||
|
|
||||||
|
By oversampling - retrieving 100, 500, or even 1000 candidates in the prefetch stage - you give ColBERT a much larger pool to work with. This dramatically increases the chance that the truly best documents are in the candidate set.
|
||||||
|
|
||||||
|
### Choosing the Oversampling Factor
|
||||||
|
|
||||||
|
The oversampling factor (how many candidates to retrieve vs. how many final results you want) is a trade-off:
|
||||||
|
|
||||||
|
- **Higher oversampling** (e.g., retrieve 1000 to return 10):
|
||||||
|
- Better final search quality
|
||||||
|
- Higher computational cost in Stage 2
|
||||||
|
|
||||||
|
- **Lower oversampling** (e.g., retrieve 50 to return 10):
|
||||||
|
- Faster overall query time
|
||||||
|
- Lower computational cost
|
||||||
|
- Risk of missing relevant documents
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
## Implementing Multi-Stage Retrieval in Qdrant
|
||||||
|
|
||||||
|
Qdrant's Universal Query API makes multi-stage retrieval straightforward using the `prefetch` parameter. The pattern is simple: whenever a query has at least one prefetch, Qdrant:
|
||||||
|
|
||||||
|
1. Performs the prefetch query (or queries)
|
||||||
|
2. Applies the main query over the results of the prefetch
|
||||||
|
|
||||||
|
### Basic Example: Single-Vector to ColBERT
|
||||||
|
|
||||||
|
Let's say you want to search for documents about "quantum computing applications" and return the top 10 results.
|
||||||
|
|
||||||
|
First, create a collection with both single-vector and multi-vector representations. We'll also add sparse vectors for BM25 to enable hybrid search later:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
|
client = QdrantClient("http://localhost:6333")
|
||||||
|
|
||||||
|
# Create collection with both single-vector and multi-vector representations
|
||||||
|
client.create_collection(
|
||||||
|
collection_name="hybrid-search",
|
||||||
|
vectors_config={
|
||||||
|
# Fast single-vector for prefetch stage
|
||||||
|
"bge-small-en-v1.5": models.VectorParams(
|
||||||
|
size=384,
|
||||||
|
distance=models.Distance.COSINE,
|
||||||
|
),
|
||||||
|
# High-quality multi-vector for reranking stage
|
||||||
|
"colbert": models.VectorParams(
|
||||||
|
size=128,
|
||||||
|
distance=models.Distance.DOT,
|
||||||
|
multivector_config=models.MultiVectorConfig(
|
||||||
|
comparator=models.MultiVectorComparator.MAX_SIM,
|
||||||
|
),
|
||||||
|
hnsw_config=models.HnswConfigDiff(m=0),
|
||||||
|
),
|
||||||
|
},
|
||||||
|
sparse_vectors_config={
|
||||||
|
"bm25": models.SparseVectorParams(
|
||||||
|
modifier=models.Modifier.IDF,
|
||||||
|
),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
Next, ingest your documents. Using `models.Document`, Qdrant handles the embedding automatically on the server side:
|
||||||
|
|
||||||
|
```python
|
||||||
|
documents = [
|
||||||
|
("research", "Quantum computing applications are emerging in cryptography..."),
|
||||||
|
("research", "Researchers are exploring quantum computing applications in drug discovery..."),
|
||||||
|
("finance", "Quantum computing applications in finance include portfolio optimization..."),
|
||||||
|
# ... more documents
|
||||||
|
]
|
||||||
|
|
||||||
|
# Ingest data to the collection - Qdrant embeds automatically
|
||||||
|
client.upsert(
|
||||||
|
collection_name="hybrid-search",
|
||||||
|
points=[
|
||||||
|
models.PointStruct(
|
||||||
|
id=i,
|
||||||
|
vector={
|
||||||
|
"bge-small-en-v1.5": models.Document(
|
||||||
|
text=doc,
|
||||||
|
model="BAAI/bge-small-en-v1.5",
|
||||||
|
),
|
||||||
|
"colbert": models.Document(
|
||||||
|
text=doc,
|
||||||
|
model="colbert-ir/colbertv2.0",
|
||||||
|
),
|
||||||
|
"bm25": models.Document(
|
||||||
|
text=doc,
|
||||||
|
model="Qdrant/bm25",
|
||||||
|
),
|
||||||
|
},
|
||||||
|
payload={
|
||||||
|
"text": doc,
|
||||||
|
"category": category,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
for i, (category, doc) in enumerate(documents)
|
||||||
|
]
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
Now, here's how you structure the multi-stage query. Notice that we use `models.Document` for the query as well - Qdrant embeds it server-side using the appropriate model for each stage:
|
||||||
|
|
||||||
|
```python
|
||||||
|
query = "quantum computing applications"
|
||||||
|
|
||||||
|
# Multi-stage query: prefetch with single-vector, rerank with ColBERT
|
||||||
|
results = client.query_points(
|
||||||
|
collection_name="hybrid-search",
|
||||||
|
prefetch=[
|
||||||
|
models.Prefetch(
|
||||||
|
query=models.Document(
|
||||||
|
text=query,
|
||||||
|
model="BAAI/bge-small-en-v1.5",
|
||||||
|
),
|
||||||
|
using="bge-small-en-v1.5",
|
||||||
|
limit=500, # Retrieve 500 candidates for reranking
|
||||||
|
),
|
||||||
|
],
|
||||||
|
query=models.Document(
|
||||||
|
text=query,
|
||||||
|
model="colbert-ir/colbertv2.0",
|
||||||
|
),
|
||||||
|
using="colbert",
|
||||||
|
limit=10, # Return top 10 after reranking
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Understanding the Query Structure
|
||||||
|
|
||||||
|
The key parts of a multi-stage query:
|
||||||
|
|
||||||
|
- **`prefetch`**: Defines the first stage search
|
||||||
|
- Uses `models.Document` to specify the text and embedding model
|
||||||
|
- The `using` parameter specifies which named vector to search
|
||||||
|
- Has its own `limit` parameter (this is your oversampling size)
|
||||||
|
- Quickly narrows down the candidate set
|
||||||
|
|
||||||
|
- **Main query**: Defines the reranking stage
|
||||||
|
- Uses `models.Document` with the ColBERT model for multi-vector embedding
|
||||||
|
- The `using` parameter specifies the multi-vector named vector (e.g., `"colbert"`)
|
||||||
|
- Its `limit` parameter determines final result count
|
||||||
|
- Only runs on the candidates from prefetch
|
||||||
|
|
||||||
|
Using `models.Document` lets Qdrant handle all embedding with FastEmbed, simplifying your client code and ensuring consistency between indexing and querying.
|
||||||
|
|
||||||
|
## Advanced Multi-Stage Patterns
|
||||||
|
|
||||||
|
Multi-stage retrieval isn't limited to just two stages. You can chain multiple prefetch operations for three or more stages by nesting prefetch operations.
|
||||||
|
|
||||||
|
### Combining Multiple Weak Retrievers
|
||||||
|
|
||||||
|
You can also use multiple retrieval methods in the prefetch stage. For example, you might combine both dense and sparse vectors using query fusion to create a stronger initial candidate set, then rerank with ColBERT.
|
||||||
|
|
||||||
|
This hybrid approach in the prefetch stage can improve recall - ensuring that the candidate pool contains relevant documents that might be missed by either dense or sparse search alone. Since we configured BM25 sparse vectors when creating the collection, we can use them directly in the prefetch:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Multi-stage with hybrid prefetch: combine dense and sparse retrieval
|
||||||
|
results = client.query_points(
|
||||||
|
collection_name="hybrid-search",
|
||||||
|
prefetch=[
|
||||||
|
# Dense retrieval using single-vector embeddings
|
||||||
|
models.Prefetch(
|
||||||
|
query=models.Document(
|
||||||
|
text=query,
|
||||||
|
model="BAAI/bge-small-en-v1.5",
|
||||||
|
),
|
||||||
|
using="bge-small-en-v1.5",
|
||||||
|
limit=500,
|
||||||
|
),
|
||||||
|
# Sparse retrieval using BM25
|
||||||
|
models.Prefetch(
|
||||||
|
query=models.Document(
|
||||||
|
text=query,
|
||||||
|
model="Qdrant/bm25",
|
||||||
|
),
|
||||||
|
using="bm25",
|
||||||
|
limit=500,
|
||||||
|
),
|
||||||
|
],
|
||||||
|
# Results from both prefetch queries are combined, then reranked
|
||||||
|
query=models.Document(
|
||||||
|
text=query,
|
||||||
|
model="colbert-ir/colbertv2.0",
|
||||||
|
),
|
||||||
|
using="colbert",
|
||||||
|
limit=10,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
Multi-stage retrieval works seamlessly with filters. An important behavior to understand: **filters in the main query are automatically propagated to all prefetch stages**. This means when you add a filter to your main query, it applies to the entire multi-stage pipeline. This is efficient because it narrows the candidate pool early in the prefetch stage, reducing computational cost throughout.
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Filters in the main query automatically propagate to prefetch stages
|
||||||
|
results = client.query_points(
|
||||||
|
collection_name="hybrid-search",
|
||||||
|
prefetch=[
|
||||||
|
models.Prefetch(
|
||||||
|
query=models.Document(
|
||||||
|
text=query,
|
||||||
|
model="BAAI/bge-small-en-v1.5",
|
||||||
|
),
|
||||||
|
using="bge-small-en-v1.5",
|
||||||
|
limit=500,
|
||||||
|
),
|
||||||
|
],
|
||||||
|
query=models.Document(
|
||||||
|
text=query,
|
||||||
|
model="colbert-ir/colbertv2.0",
|
||||||
|
),
|
||||||
|
using="colbert",
|
||||||
|
limit=10,
|
||||||
|
# This filter applies to BOTH prefetch and reranking stages
|
||||||
|
query_filter=models.Filter(
|
||||||
|
must=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="category",
|
||||||
|
match=models.MatchValue(value="research"),
|
||||||
|
),
|
||||||
|
],
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
## When to Use Multi-Stage Retrieval
|
||||||
|
|
||||||
|
Multi-stage retrieval is valuable for **large collections** (100K+ documents) where you need **fast queries** while maintaining **multi-vector quality**. It's particularly effective when combining different embedding models or optimizing compute costs.
|
||||||
|
|
||||||
|
Skip multi-stage retrieval for small collections (< 10K documents), scenarios where single-vector embeddings suffice, or real-time indexing where maintaining dual embeddings adds excessive overhead.
|
||||||
|
|
||||||
|
## Performance Characteristics
|
||||||
|
|
||||||
|
Multi-stage retrieval's key advantage is **reducing the number of documents the multi-vector model scans**. Since late interaction performs full scans, limiting candidates dramatically improves performance.
|
||||||
|
|
||||||
|
**Direct late interaction**: 1M documents = 1M MaxSim calculations
|
||||||
|
**Multi-stage**: 1M documents -> prefetch 500 candidates = 500 MaxSim calculations
|
||||||
|
|
||||||
|
Speedup is roughly **size of the collection / number of candidates**:
|
||||||
|
- 1M documents, prefetch 1000 -> ~1000x fewer calculations
|
||||||
|
- 100K documents, prefetch 500 -> ~200x fewer calculations
|
||||||
|
|
||||||
|
Higher oversampling improves quality but increases Stage 2 cost
|
||||||
|
|
||||||
|
## Bringing It All Together
|
||||||
|
|
||||||
|
Multi-stage retrieval is a powerful optimization technique, but it's just one tool in your arsenal. Real-world production systems often combine multiple techniques:
|
||||||
|
|
||||||
|
- **Multi-stage retrieval** for computational efficiency
|
||||||
|
- **Score boosting** to adjust relevance based on metadata
|
||||||
|
- **Diversification** to reduce redundancy in results
|
||||||
|
- **Filtering** to narrow results by business rules
|
||||||
|
- **Query fusion** to combine multiple search strategies
|
||||||
|
|
||||||
|
The beauty of Qdrant's Universal Query API is that all these techniques can be composed together in a single query. You might use multi-stage retrieval with oversampling, apply filters at each stage, boost scores based on recency, and diversify final results - all in one request.
|
||||||
|
|
||||||
|
However, **there's no one-size-fits-all solution**. The optimal search pipeline depends on your specific data characteristics, quality requirements, latency constraints, and computational budget. This is why **evaluation is critical**. You need to systematically measure and compare different pipeline configurations to find what works best for your use case.
|
||||||
|
|
||||||
|
In the final lesson of this module, you'll learn exactly how to evaluate and compare different search pipelines - giving you the tools to make informed, data-driven decisions about your search architecture.
|
||||||
|
|
||||||
|
## What's Next
|
||||||
|
|
||||||
|
You've learned how to build sophisticated multi-stage retrieval pipelines that combine different techniques for optimal results. But multi-vector representations can be memory-intensive, especially at scale.
|
||||||
|
|
||||||
|
In the next lesson, you'll discover **vector quantization techniques** - powerful compression methods that can reduce memory usage by 4-32x while maintaining search quality.
|
||||||
@@ -0,0 +1,199 @@
|
|||||||
|
---
|
||||||
|
title: "MUVERA"
|
||||||
|
description: Understand MUVERA and how it enables HNSW indexing for multi-vector search despite MaxSim asymmetry.
|
||||||
|
weight: 4
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 3 {{< /date >}}
|
||||||
|
|
||||||
|
# MUVERA
|
||||||
|
|
||||||
|
MUVERA (Multi-Vector Retrieval with Approximation) solves a fundamental problem: MaxSim's asymmetry makes traditional indexing methods like HNSW ineffective. MUVERA enables fast approximate search for multi-vector representations.
|
||||||
|
|
||||||
|
Understanding MUVERA is key to scaling multi-vector search to millions of documents.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/-r0Apuy0c8k?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-3/muvera.ipynb">
|
||||||
|
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||||
|
</a>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The HNSW Incompatibility Problem
|
||||||
|
|
||||||
|
Traditional vector indexes like HNSW are designed for single-vector search with symmetric distance metrics. Multi-vector representations break this assumption: **MaxSim is inherently asymmetric and non-metric**.
|
||||||
|
|
||||||
|
When comparing query tokens to document tokens, the direction matters - searching documents with a query gives different results than searching queries with a document. This asymmetry makes HNSW's graph-based navigation ineffective, forcing us back to full scans across millions of documents.
|
||||||
|
|
||||||
|
## MUVERA: Making Multi-Vector Search Fast
|
||||||
|
|
||||||
|
MUVERA (Multi-Vector Retrieval Algorithm) solves this incompatibility by creating a **single approximation vector** for each document that HNSW can efficiently index. The algorithm works in three stages:
|
||||||
|
|
||||||
|
1. **SimHash Clustering**: Groups token vectors into spatial regions using random hyperplanes
|
||||||
|
2. **Fixed Dimensional Encoding (FDE)**: Aggregates clustered vectors into a single representative vector per document
|
||||||
|
3. **Dimensionality Reduction**: Applies random projection to create compact, robust representations
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
The resulting single vector approximates the multi-vector representation well enough for fast retrieval, achieving massive speedup. A reranking step with the original multi-vector data then aims to recover full accuracy.
|
||||||
|
|
||||||
|
For a deep dive into MUVERA's architecture and mathematical foundations, see our detailed article: [MUVERA: Making Multivectors More Performant](/articles/muvera-embeddings/).
|
||||||
|
|
||||||
|
## Applying MUVERA Multi-Stage Retrieval
|
||||||
|
|
||||||
|
FastEmbed provides built-in support for MUVERA postprocessing, making it straightforward to implement this optimization. Let's see how to apply MUVERA to multi-vector embeddings for fast retrieval with multi-stage reranking.
|
||||||
|
|
||||||
|
### Setting Up MUVERA with ColModernVBERT
|
||||||
|
|
||||||
|
FastEmbed 0.7.2+ includes MUVERA as a postprocessing technique that transforms variable-length multi-vector sequences into fixed-dimensional single vectors. You can apply MUVERA to any late interaction model, including **ColModernVBERT**.
|
||||||
|
|
||||||
|
The workflow involves loading a multi-vector embedding model and wrapping it with a MUVERA processor:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import LateInteractionMultimodalEmbedding
|
||||||
|
from fastembed.postprocess import Muvera
|
||||||
|
|
||||||
|
# Load ColModernVBERT model
|
||||||
|
model = LateInteractionMultimodalEmbedding(model_name="Qdrant/colmodernvbert")
|
||||||
|
|
||||||
|
# Wrap with MUVERA processor
|
||||||
|
muvera = Muvera.from_multivector_model(
|
||||||
|
model=model,
|
||||||
|
k_sim=6, # 2^6 = 64 similarity buckets
|
||||||
|
dim_proj=32, # Projection dimensionality
|
||||||
|
r_reps=20 # Random projection repetitions
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
The MUVERA processor accepts several key parameters that control the speed-accuracy tradeoff:
|
||||||
|
|
||||||
|
- **`k_sim`**: Number of similarity buckets for SimHash clustering (more buckets = higher precision)
|
||||||
|
- **`dim_proj`**: Target dimensionality for the projection (lower = faster search, higher = better accuracy)
|
||||||
|
- **`r_reps`**: Number of random projections for robust encoding (higher = more stable representations)
|
||||||
|
|
||||||
|
### Creating the Collection
|
||||||
|
|
||||||
|
For multi-stage retrieval to work, you need to index **both** the MUVERA approximation vectors and the original multi-vector representations in Qdrant. The MUVERA vectors enable fast HNSW retrieval, while the multi-vector representations provide accurate reranking.
|
||||||
|
|
||||||
|
First, create a collection with both vector types:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
|
# Create collection with both vector types
|
||||||
|
client = QdrantClient("http://localhost:6333")
|
||||||
|
client.create_collection(
|
||||||
|
collection_name="documents-muvera",
|
||||||
|
vectors_config={
|
||||||
|
"muvera": models.VectorParams(
|
||||||
|
size=muvera.embedding_size,
|
||||||
|
distance=models.Distance.COSINE
|
||||||
|
),
|
||||||
|
"colmodernvbert": models.VectorParams(
|
||||||
|
size=model.embedding_size,
|
||||||
|
distance=models.Distance.COSINE,
|
||||||
|
multivector_config=models.MultiVectorConfig(
|
||||||
|
comparator=models.MultiVectorComparator.MAX_SIM
|
||||||
|
)
|
||||||
|
)
|
||||||
|
}
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
The dual-vector approach stores:
|
||||||
|
1. **MUVERA embeddings**: Single vector per document, fast HNSW retrieval
|
||||||
|
2. **Multi-vector representation**: Full token sequences per document, precise MaxSim scoring
|
||||||
|
|
||||||
|
### Embedding and Indexing Documents
|
||||||
|
|
||||||
|
Next, embed your documents and upload both representations. Generate multi-vector embeddings with ColModernVBERT, then process them through MUVERA:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Embed documents
|
||||||
|
image_path = "images/financial-report.png"
|
||||||
|
doc_multivec = list(model.embed_image([image_path]))[0]
|
||||||
|
doc_muvera = muvera.process_document(doc_multivec)
|
||||||
|
```
|
||||||
|
|
||||||
|
Upload both the MUVERA approximation and the original multi-vector representation:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Upload both representations
|
||||||
|
client.upsert(
|
||||||
|
collection_name="documents-muvera",
|
||||||
|
points=[
|
||||||
|
models.PointStruct(
|
||||||
|
id=0,
|
||||||
|
payload={"source": image_path},
|
||||||
|
vector={"muvera": doc_muvera, "colmodernvbert": doc_multivec}
|
||||||
|
)
|
||||||
|
]
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
### Querying with Multi-Stage Retrieval
|
||||||
|
|
||||||
|
At query time, you first retrieve candidates using the fast MUVERA vectors, then rescore those candidates using the original multi-vector representations. This hybrid approach maintains search quality while dramatically reducing compute requirements.
|
||||||
|
|
||||||
|
First, encode the query in both formats:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Encode query in both formats
|
||||||
|
query = "quarterly revenue growth"
|
||||||
|
query_multivec = list(model.embed_text(query))[0]
|
||||||
|
query_muvera = muvera.process_query(query_multivec)
|
||||||
|
```
|
||||||
|
|
||||||
|
Then perform two-stage retrieval using Qdrant's prefetch mechanism:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Two-stage retrieval with Qdrant's prefetch
|
||||||
|
results = client.query_points(
|
||||||
|
collection_name="documents-muvera",
|
||||||
|
prefetch=models.Prefetch(
|
||||||
|
query=query_muvera,
|
||||||
|
using="muvera",
|
||||||
|
limit=100, # Stage 1: Fast MUVERA retrieval
|
||||||
|
),
|
||||||
|
query=query_multivec,
|
||||||
|
using="colmodernvbert", # Stage 2: Precise MaxSim reranking
|
||||||
|
limit=10,
|
||||||
|
with_payload=True
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
This two-stage process achieves near-identical accuracy to full multi-vector search across all documents, but only computes expensive MaxSim operations for a small candidate set.
|
||||||
|
|
||||||
|
For complete implementation details and parameter tuning guidance, see the [FastEmbed Postprocessing documentation](/documentation/fastembed/fastembed-postprocessing/).
|
||||||
|
|
||||||
|
### Trade-offs and Considerations
|
||||||
|
|
||||||
|
**Storage Requirements**: MUVERA requires storing both representations - the single approximation vector and the full multi-vector sequence. This doubles storage compared to single-vector search, but remains practical for production systems. Offloading original vectors to disk might be a solution, if you can afford using that much memory.
|
||||||
|
|
||||||
|
**Speed Gains**: The performance improvement is substantial. Instead of computing MaxSim against millions of documents, you only compute it for your candidate set (typically 100-1000 documents). The MUVERA-powered HNSW retrieval scales logarithmically, making multi-vector search viable at scale.
|
||||||
|
|
||||||
|
**Accuracy Preservation**: Research shows MUVERA maintains nearly the same accuracy as full multi-vector search when properly configured. The key is choosing appropriate parameter values based on your dataset and quality requirements.
|
||||||
|
|
||||||
|
## What's Next
|
||||||
|
|
||||||
|
You've learned why MaxSim's asymmetry prevents HNSW from efficiently indexing multi-vector representations, and how MUVERA solves this with single-vector approximations. By combining MUVERA's fast retrieval with multi-vector reranking, you can achieve significant speedups while maintaining search quality.
|
||||||
|
|
||||||
|
The optimization techniques covered in this module - quantization, pooling, and MUVERA - are **complementary techniques** that can be combined for maximum efficiency. For example, you can test quantization on both your MUVERA vectors and your stored multi-vector representations, reducing memory footprint while maintaining the speed benefits of HNSW indexing. Or keep MUVERA vectors unchanged and play with quantization and pooling for the original vectors. These optimizations stack together, allowing you to build highly efficient production pipelines.
|
||||||
|
|
||||||
|
Now that you have multiple optimization tools in your toolkit, how do you evaluate which combinations work best for your use case? Let's learn how to measure and compare different search pipeline configurations.
|
||||||
@@ -0,0 +1,209 @@
|
|||||||
|
---
|
||||||
|
title: "Pooling Techniques"
|
||||||
|
description: Reduce the number of vectors per document using row/column pooling and hierarchical token pooling strategies.
|
||||||
|
weight: 3
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 3 {{< /date >}}
|
||||||
|
|
||||||
|
# Pooling Techniques
|
||||||
|
|
||||||
|
While quantization reduces the size of each vector, pooling reduces the number of vectors per document. By intelligently combining token embeddings, you can achieve significant memory savings while preserving retrieval quality.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/idDXBOrIuik?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-3/pooling-techniques.ipynb">
|
||||||
|
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||||
|
</a>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Pooling in Embedding Models
|
||||||
|
|
||||||
|
Pooling isn't new to vector search - it's fundamental to how most embedding models work. When you encode text with models like Sentence Transformers, the model first generates embeddings for each token in your input. But to create a single vector representing the entire text, the model must **pool** these token embeddings together.
|
||||||
|
|
||||||
|
Common pooling strategies in dense embedding models include:
|
||||||
|
|
||||||
|
- **Mean pooling**: Average all token embeddings into a single vector
|
||||||
|
- **CLS token pooling**: Use the special `[CLS]` token's embedding as the document representation
|
||||||
|
- **Max pooling**: Take the maximum value for each dimension across all tokens
|
||||||
|
- **Weighted pooling**: Assign different importance to different tokens (e.g., using attention weights)
|
||||||
|
|
||||||
|
These techniques compress variable-length sequences of token embeddings into fixed-size vectors, making them compatible with traditional vector search systems.
|
||||||
|
|
||||||
|
With multi-vector representations, we face a similar but more nuanced challenge. Instead of reducing tokens to a single vector upfront, we maintain multiple vectors per document to preserve richer semantic information. However, as you learned in the previous lessons, this creates memory and performance challenges. **Pooling techniques for multi-vector search** let you strategically reduce the number of vectors while retaining the benefits of late interaction.
|
||||||
|
|
||||||
|
**Important:** Pooling is typically applied only to **document embeddings**, not queries. Why? Queries are usually short (a few tokens), so there's little memory to save. More importantly, we want to preserve full query resolution - every query token should have the opportunity to find its best match among document tokens. The memory savings come from compressing the large document collection, not the ephemeral query vectors.
|
||||||
|
|
||||||
|
## Pooling for Multi-Vector Representations
|
||||||
|
|
||||||
|
### Image-Specific Methods
|
||||||
|
|
||||||
|
For visual document representations like ColPali, spatial relationships in the patch grid enable effective pooling strategies. As you learned in Module 2's visual interpretability lesson, patches in the same row or column often capture semantically related content - a row might contain a line of text, while a column might capture a vertical element like a table border or sidebar.
|
||||||
|
|
||||||
|
**Row pooling** groups patches by their horizontal position:
|
||||||
|
|
||||||
|
1. Organize the 1024 patch embeddings into a 32×32 grid
|
||||||
|
2. Apply mean pooling across each row (combining 32 patches)
|
||||||
|
3. Result: 32 vectors instead of 1024
|
||||||
|
|
||||||
|
Mathematically:
|
||||||
|
|
||||||
|
<p>$$\text{RowPool}_i = \text{Mean}(\lbrace p_{i,j} : j \in [0, 31] \rbrace)$$</p>
|
||||||
|
|
||||||
|
Where $p_{i,j}$ is the patch embedding at row $i$, column $j$.
|
||||||
|
|
||||||
|
**Column pooling** works similarly but along the vertical axis:
|
||||||
|
|
||||||
|
<p>$$\text{ColPool}_j = \text{Mean}(\lbrace p_{i,j} : i \in [0, 31] \rbrace)$$</p>
|
||||||
|
|
||||||
|
This also produces 32 vectors, but captures vertical content relationships instead.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
**Memory savings** are substantial. FastEmbed returns embeddings in float16 format by default, which already halves the memory compared to float32:
|
||||||
|
|
||||||
|
| Representation | Vectors | Memory (float16) | Memory (float32) |
|
||||||
|
|----------------|---------|------------------|------------------|
|
||||||
|
| Full patches | 1024 | 256 KB | 512 KB |
|
||||||
|
| Row pooling | 32 | 8 KB | 16 KB |
|
||||||
|
| Column pooling | 32 | 8 KB | 16 KB |
|
||||||
|
|
||||||
|
That's a **32× reduction** in vector count and memory footprint.
|
||||||
|
|
||||||
|
**Trade-offs to consider:**
|
||||||
|
|
||||||
|
- **Loss of fine-grained resolution**: Small details that span partial rows may blend together
|
||||||
|
- **Row pooling** may work better for horizontally-oriented content, like text
|
||||||
|
- **Column pooling** may better capture vertical structures like tables, sidebars, or vertically-oriented text
|
||||||
|
- You can combine both (64 vectors) for a balanced approach
|
||||||
|
|
||||||
|
```python
|
||||||
|
import numpy as np
|
||||||
|
from fastembed import LateInteractionMultimodalEmbedding
|
||||||
|
|
||||||
|
# Load ColPali model
|
||||||
|
model = LateInteractionMultimodalEmbedding(model_name="Qdrant/colpali-v1.3-fp16")
|
||||||
|
|
||||||
|
# Embed a document image (returns ~1030 vectors × 128 dimensions)
|
||||||
|
image_path = "images/financial-report.png" # Your document image
|
||||||
|
embeddings = list(model.embed_image([image_path]))[0]
|
||||||
|
print(f"Original shape: {embeddings.shape}") # (1030, 128)
|
||||||
|
|
||||||
|
# Reshape to spatial grid: (rows, columns, embedding_dim)
|
||||||
|
# Get only the first 1024 embeddings, as instruction tokens do
|
||||||
|
# not represent images
|
||||||
|
grid = embeddings[:1024].reshape(32, 32, 128)
|
||||||
|
|
||||||
|
# Row pooling: average across columns (axis=1)
|
||||||
|
row_pooled = grid.mean(axis=1) # Shape: (32, 128)
|
||||||
|
|
||||||
|
# Column pooling: average across rows (axis=0)
|
||||||
|
col_pooled = grid.mean(axis=0) # Shape: (32, 128)
|
||||||
|
|
||||||
|
# Combined approach (optional): concatenate row and column pooled
|
||||||
|
combined = np.vstack([row_pooled, col_pooled]) # Shape: (64, 128)
|
||||||
|
|
||||||
|
# Memory comparison (FastEmbed uses float16 by default)
|
||||||
|
original_memory = embeddings.nbytes # 1030 × 128 × 2 = 263,680 bytes
|
||||||
|
pooled_memory = row_pooled.nbytes # 32 × 128 × 2 = 8,192 bytes
|
||||||
|
|
||||||
|
print(f"Original: {original_memory:,} bytes ({original_memory // 1024} KB)")
|
||||||
|
print(f"Row pooled: {pooled_memory:,} bytes ({pooled_memory // 1024} KB)")
|
||||||
|
print(f"Reduction: {original_memory // pooled_memory}×")
|
||||||
|
```
|
||||||
|
|
||||||
|
### Generic Methods
|
||||||
|
|
||||||
|
While row/column pooling exploits the spatial structure of image embeddings, **hierarchical token pooling** works for any multi-vector representation - text, images, or hybrid documents. The core idea: instead of grouping by fixed spatial positions, cluster tokens by **semantic similarity**.
|
||||||
|
|
||||||
|
**How hierarchical pooling works:**
|
||||||
|
|
||||||
|
1. Apply k-means clustering to group similar token embeddings
|
||||||
|
2. Pool within each cluster using mean pooling
|
||||||
|
3. Output: $k$ vectors instead of $n$ original tokens
|
||||||
|
|
||||||
|
This approach adapts to the content itself. For a document with dense text and sparse images, clustering naturally allocates more representative vectors to the text regions where semantic variation is higher.
|
||||||
|
|
||||||
|
**Key parameters:**
|
||||||
|
|
||||||
|
- **Number of clusters ($k$)**: Controls the compression ratio. $k=32$ gives similar compression to row pooling, while $k=64$ preserves more detail
|
||||||
|
- **Clustering algorithm**: k-means is fast and effective. Although hierarchical clustering can capture nested semantic structures but adds overhead
|
||||||
|
|
||||||
|
**Comparison: Row/Column vs. Hierarchical Pooling**
|
||||||
|
|
||||||
|
| Aspect | Row/Column Pooling | Hierarchical Pooling |
|
||||||
|
|-----------------------|-------------------------------------|---------------------------------|
|
||||||
|
| **Works with** | Images only (requires spatial grid) | Any multi-vector representation |
|
||||||
|
| **Grouping strategy** | Fixed spatial positions | Semantic similarity |
|
||||||
|
| **Compression ratio** | Fixed (32×) | Configurable via $k$ |
|
||||||
|
| **Indexing overhead** | None | Clustering computation |
|
||||||
|
| **Preserves** | Spatial structure | Semantic diversity |
|
||||||
|
|
||||||
|
**Trade-offs:**
|
||||||
|
|
||||||
|
- **Higher indexing cost**: Clustering adds computational overhead during document encoding
|
||||||
|
- **Content-adaptive**: Allocates representation capacity where semantic variation is highest
|
||||||
|
- **Loses spatial interpretability**: Unlike row pooling, you can't easily map pooled vectors back to document regions
|
||||||
|
- **Hyperparameter sensitivity**: The choice of $k$ affects retrieval quality and must be tuned
|
||||||
|
|
||||||
|
```python
|
||||||
|
from scipy.cluster.vq import kmeans2
|
||||||
|
|
||||||
|
# Embed a document image
|
||||||
|
image_path = "images/financial-report.png"
|
||||||
|
embeddings = list(model.embed_image([image_path]))[0]
|
||||||
|
|
||||||
|
def hierarchical_pool(embeddings: np.ndarray, k: int) -> np.ndarray:
|
||||||
|
"""Pool embeddings using k-means clustering."""
|
||||||
|
# Cluster embeddings into k groups
|
||||||
|
centroids, labels = kmeans2(embeddings, k, minit='++')
|
||||||
|
|
||||||
|
# Pool within each cluster using mean
|
||||||
|
pooled = np.array([
|
||||||
|
embeddings[labels == i].mean(axis=0)
|
||||||
|
for i in range(k)
|
||||||
|
])
|
||||||
|
return pooled
|
||||||
|
|
||||||
|
# Compare different compression levels
|
||||||
|
for k in [16, 32, 64, 128]:
|
||||||
|
pooled = hierarchical_pool(embeddings, k)
|
||||||
|
reduction = len(embeddings) / k
|
||||||
|
print(f"k={k:3d}: {len(embeddings)} → {k} vectors ({reduction:.0f}× reduction)")
|
||||||
|
```
|
||||||
|
|
||||||
|
## What's Next
|
||||||
|
|
||||||
|
This lesson covered two complementary strategies for reducing the number of vectors per document:
|
||||||
|
|
||||||
|
- **Row/column pooling**: Exploits spatial structure in image embeddings for a fixed reduction (32x for ColPali)
|
||||||
|
- **Hierarchical pooling**: Content-adaptive clustering that works for any multi-vector representation
|
||||||
|
|
||||||
|
Combined with quantization from the previous lesson, you can achieve dramatic memory savings:
|
||||||
|
|
||||||
|
| Technique | Memory per Document |
|
||||||
|
|-----------------------------------|---------------------|
|
||||||
|
| Baseline (1024 vectors × float32) | 512 KB |
|
||||||
|
| Row pooling only | 16 KB |
|
||||||
|
| Row pooling + scalar quantization | 4 KB |
|
||||||
|
| Row pooling + binary quantization | 512 bytes |
|
||||||
|
|
||||||
|
That's a **1000× reduction** from baseline to the most aggressive combination - making multi-vector search practical even for large document collections.
|
||||||
|
|
||||||
|
However, there's still one challenge we haven't addressed: **indexing**. Even with pooled representations, we're still performing brute-force MaxSim comparisons. For millions of documents, this becomes a bottleneck.
|
||||||
|
|
||||||
|
In the next lesson, you'll learn about **MUVERA** - a technique that enables HNSW indexing for multi-vector representations, unlocking fast approximate search at scale.
|
||||||
@@ -0,0 +1,336 @@
|
|||||||
|
---
|
||||||
|
title: "Vector Quantization Techniques"
|
||||||
|
description: Learn how to reduce memory usage with scalar quantization, binary quantization, and other compression methods.
|
||||||
|
weight: 2
|
||||||
|
isLesson: true
|
||||||
|
---
|
||||||
|
|
||||||
|
{{< date >}} Module 3 {{< /date >}}
|
||||||
|
|
||||||
|
# Vector Quantization Techniques
|
||||||
|
|
||||||
|
Vector quantization compresses vectors by reducing the precision of each component. Qdrant supports several quantization methods that can reduce memory usage by 4-64x, sometimes with minimal quality loss.
|
||||||
|
|
||||||
|
Choosing the right quantization method depends on your quality requirements and memory constraints.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<div class="video">
|
||||||
|
<iframe
|
||||||
|
src="https://www.youtube-nocookie.com/embed/we-AEfiXaow?rel=0"
|
||||||
|
frameborder="0"
|
||||||
|
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||||
|
referrerpolicy="strict-origin-when-cross-origin"
|
||||||
|
allowfullscreen>
|
||||||
|
</iframe>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course-multi-vector-search/module-3/quantization-techniques.ipynb">
|
||||||
|
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||||
|
</a>
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The Memory Challenge with Multi-Vector Models
|
||||||
|
|
||||||
|
By default, embedding models produce vectors with **float32 precision** - each component uses 32 bits (4 bytes) of memory. For single-vector embeddings, this is manageable. But multi-vector models like **ColModernVBERT** change the equation dramatically.
|
||||||
|
|
||||||
|
Consider a typical ColPali scenario using **ColModernVBERT**:
|
||||||
|
- **~1024 vectors per document** (one per visual patch)
|
||||||
|
- **128 dimensions per vector** (model embedding size)
|
||||||
|
- **float32 precision** (4 bytes per component)
|
||||||
|
|
||||||
|
Let's calculate the memory for a single document:
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{Memory per document} = 1024 \text{ vectors} \times 128 \text{ dims} \times 4 \text{ bytes} = 524{,}288 \text{ bytes} = 512 \text{ KB}
|
||||||
|
$$
|
||||||
|
|
||||||
|
For a collection of **1 million documents**:
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{Total memory} = 1{,}000{,}000 \times 512 \text{ KB} = 512 \text{ GB}
|
||||||
|
$$
|
||||||
|
|
||||||
|
Compare this to a traditional single-vector model (e.g., 768-dimensional):
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{Single-vector memory} = 768 \text{ dims} \times 4 \text{ bytes} = 3{,}072 \text{ bytes} = 3 \text{ KB per document}
|
||||||
|
$$
|
||||||
|
|
||||||
|
**Multi-vector representations use ~170x more memory** than single-vector models for the same number of documents. This is where quantization becomes essential.
|
||||||
|
|
||||||
|
## Quantization: Compressing Without Losing Quality
|
||||||
|
|
||||||
|
**Vector quantization** reduces memory by representing vectors with fewer bits while preserving the relative distances between them. Qdrant supports several quantization methods optimized for different scenarios.
|
||||||
|
|
||||||
|
**Important: What Quantization Does (and Doesn't) Do**
|
||||||
|
|
||||||
|
Quantization is a **memory optimization technique**, not an indexing solution:
|
||||||
|
- ✅ **Reduces memory footprint** by 4-32x for storing multi-vector representations
|
||||||
|
- ✅ **Reduces infrastructure costs** by requiring less RAM
|
||||||
|
- ✅ **Provides speed improvements** through SIMD operations and smaller data transfers
|
||||||
|
- ❌ **Does NOT enable HNSW indexing** for multi-vector search - brute force scan is still required
|
||||||
|
|
||||||
|
Multi-vector search with MaxSim fundamentally requires comparing query tokens against all document tokens. HNSW and other graph-based indexes cannot efficiently navigate this token-level comparison space. Quantization makes the brute force search faster and cheaper, but the search strategy remains exhaustive.
|
||||||
|
|
||||||
|
The key insight: **you don't need perfect precision to find the right matches**. If document A is closer to a query than document B in full precision, it usually remains closer after quantization.
|
||||||
|
|
||||||
|
### Scalar Quantization: The Reliable Default
|
||||||
|
|
||||||
|
**Scalar quantization** converts float32 values to 8-bit integers (uint8), reducing memory by **4x**.
|
||||||
|
|
||||||
|
**How it works:**
|
||||||
|
1. Find the min and max values across all vector components
|
||||||
|
2. Map the float range to [0, 255]
|
||||||
|
3. Store the scaling parameters for reconstruction
|
||||||
|
|
||||||
|
For our ColModernVBERT example:
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{Quantized memory} = \frac{512 \text{ GB}}{4} = 128 \text{ GB}
|
||||||
|
$$
|
||||||
|
|
||||||
|
**Benefits:**
|
||||||
|
- **4x memory reduction** with <1% accuracy loss
|
||||||
|
- **Up to 2x faster brute force search** via SIMD optimization (still exhaustive, but more efficient)
|
||||||
|
- **Lower infrastructure costs** by reducing RAM requirements
|
||||||
|
- Works universally across all vector types and dimensions
|
||||||
|
|
||||||
|
**Configuration parameter:**
|
||||||
|
- `quantile`: Excludes outliers (e.g., 0.99 excludes 1% of extreme values for better scaling)
|
||||||
|
|
||||||
|
### Binary Quantization: Maximum Compression
|
||||||
|
|
||||||
|
**Binary quantization** represents each component as a single bit (positive/negative), achieving **32x compression**. Qdrant also supports **1.5-bit** and **2-bit** variants for better accuracy with moderate compression.
|
||||||
|
|
||||||
|
For ColModernVBERT vectors:
|
||||||
|
|
||||||
|
**1-bit binary quantization:**
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{Memory per document} = 1024 \text{ vectors} \times 128 \text{ dims} \times \frac{1}{8} \text{ bytes} = 16{,}384 \text{ bytes} = 16 \text{ KB}
|
||||||
|
$$
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{Total for 1M docs} = \frac{512 \text{ GB}}{32} = 16 \text{ GB}
|
||||||
|
$$
|
||||||
|
|
||||||
|
**1.5-bit binary quantization:**
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{Total for 1M docs} = \frac{512 \text{ GB}}{24} \approx 21.3 \text{ GB}
|
||||||
|
$$
|
||||||
|
|
||||||
|
**2-bit binary quantization:**
|
||||||
|
|
||||||
|
$$
|
||||||
|
\text{Total for 1M docs} = \frac{512 \text{ GB}}{16} = 32 \text{ GB}
|
||||||
|
$$
|
||||||
|
|
||||||
|
**Typical dimension ranges for binary quantization:**
|
||||||
|
- **1-bit**: Often used with high-dimensional vectors (1536+ dimensions)
|
||||||
|
- **1.5-bit**: Commonly applied to 1024-1536 dimensions
|
||||||
|
- **2-bit**: Frequently used with 768-1024 dimensions
|
||||||
|
|
||||||
|
For ColModernVBERT's **128 dimensions**, binary quantization presents unique challenges. With such low dimensionality, each bit of precision has a larger impact on the representation. The choice between scalar and binary quantization - and which binary variant to use - depends on your specific use case and quality requirements. We'll explore how to evaluate these trade-offs systematically in the final lesson of this module.
|
||||||
|
|
||||||
|
### Real-World Impact for ColPali Collections
|
||||||
|
|
||||||
|
Let's compare all options for a **1 million document** ColModernVBERT collection:
|
||||||
|
|
||||||
|
| Method | Memory | Compression | Speed Boost |
|
||||||
|
|----------------------|---------|---------------|-------------|
|
||||||
|
| **No quantization** | 512 GB | 1x (baseline) | 1x |
|
||||||
|
| **Scalar (int8)** | 128 GB | 4x | ~2x |
|
||||||
|
| **Binary (2-bit)** | 32 GB | 16x | ~20x |
|
||||||
|
| **Binary (1.5-bit)** | 21.3 GB | 24x | ~30x |
|
||||||
|
| **Binary (1-bit)** | 16 GB | 32x | ~40x |
|
||||||
|
|
||||||
|
<aside role="status">
|
||||||
|
<b>Note:</b> The "Speed Boost" column refers to improvements in brute force search performance. Quantization does not enable HNSW or other graph-based indexing for multi-vector search - all documents are still scanned exhaustively. The speed improvements come from faster distance computations and reduced memory bandwidth requirements during the brute force scan.
|
||||||
|
</aside>
|
||||||
|
|
||||||
|
The table above shows the theoretical memory savings and search performance characteristics for each quantization method. The actual impact on retrieval quality is **not included** because it varies significantly based on your specific documents, queries, and quality requirements.
|
||||||
|
|
||||||
|
**The choice between these methods requires systematic evaluation** of your complete search pipeline. We'll cover evaluation methodologies in detail in the final lesson of this module, where you'll learn how to measure the impact of quantization on your specific use case.
|
||||||
|
|
||||||
|
## Enabling Quantization in Qdrant
|
||||||
|
|
||||||
|
One of Qdrant's powerful features: **you can enable quantization on an existing collection** without changing your inference or ingestion pipelines. The quantization happens transparently during indexing.
|
||||||
|
|
||||||
|
|
||||||
|
First, let's load the ColPali model and prepare some sample documents:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from fastembed import LateInteractionMultimodalEmbedding
|
||||||
|
|
||||||
|
# Load ColPali model for generating multi-vector embeddings
|
||||||
|
model = LateInteractionMultimodalEmbedding(
|
||||||
|
model_name="Qdrant/colpali-v1.3-fp16"
|
||||||
|
)
|
||||||
|
|
||||||
|
# Sample document images and metadata
|
||||||
|
image_paths = [
|
||||||
|
"images/financial-report.png",
|
||||||
|
"images/titanic-newspaper.jpg",
|
||||||
|
"images/moon-landing.jpg",
|
||||||
|
"images/einstein-newspaper.jpg",
|
||||||
|
]
|
||||||
|
|
||||||
|
documents = [
|
||||||
|
{"title": "Financial Report", "type": "report", "topic": "finance"},
|
||||||
|
{"title": "Titanic Sinking", "type": "newspaper", "topic": "history"},
|
||||||
|
{"title": "Moon Landing", "type": "newspaper", "topic": "space"},
|
||||||
|
{"title": "Einstein Theory", "type": "newspaper", "topic": "science"},
|
||||||
|
]
|
||||||
|
|
||||||
|
# Generate embeddings for all images
|
||||||
|
image_embeddings = list(model.embed_image(image_paths))
|
||||||
|
```
|
||||||
|
|
||||||
|
Now create a collection with scalar quantization:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
|
client = QdrantClient("http://localhost:6333")
|
||||||
|
|
||||||
|
client.create_collection(
|
||||||
|
collection_name="colpali-scalar",
|
||||||
|
vectors_config={
|
||||||
|
"colpali": models.VectorParams(
|
||||||
|
size=128,
|
||||||
|
distance=models.Distance.DOT,
|
||||||
|
multivector_config=models.MultiVectorConfig(
|
||||||
|
comparator=models.MultiVectorComparator.MAX_SIM,
|
||||||
|
),
|
||||||
|
hnsw_config=models.HnswConfigDiff(m=0), # Disable HNSW for multi-vector
|
||||||
|
),
|
||||||
|
},
|
||||||
|
quantization_config=models.ScalarQuantization(
|
||||||
|
scalar=models.ScalarQuantizationConfig(
|
||||||
|
type=models.ScalarType.INT8,
|
||||||
|
quantile=0.99, # Exclude 1% outliers for better scaling
|
||||||
|
always_ram=True,
|
||||||
|
),
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
Ingest the documents into the scalar-quantized collection:
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.upsert(
|
||||||
|
collection_name="colpali-scalar",
|
||||||
|
points=[
|
||||||
|
models.PointStruct(
|
||||||
|
id=i,
|
||||||
|
vector={"colpali": embedding.tolist()},
|
||||||
|
payload=documents[i],
|
||||||
|
)
|
||||||
|
for i, embedding in enumerate(image_embeddings)
|
||||||
|
],
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
Enabling a different type of quantization requires setting a different quantization configuration.
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.create_collection(
|
||||||
|
collection_name="colpali-binary",
|
||||||
|
vectors_config={
|
||||||
|
"colpali": models.VectorParams(
|
||||||
|
size=128,
|
||||||
|
distance=models.Distance.DOT,
|
||||||
|
multivector_config=models.MultiVectorConfig(
|
||||||
|
comparator=models.MultiVectorComparator.MAX_SIM,
|
||||||
|
),
|
||||||
|
hnsw_config=models.HnswConfigDiff(m=0),
|
||||||
|
),
|
||||||
|
},
|
||||||
|
quantization_config=models.BinaryQuantization(
|
||||||
|
binary=models.BinaryQuantizationConfig(
|
||||||
|
always_ram=True,
|
||||||
|
),
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
# Ingest the same data into the binary-quantized collection
|
||||||
|
client.upsert(
|
||||||
|
collection_name="colpali-binary",
|
||||||
|
points=[
|
||||||
|
models.PointStruct(
|
||||||
|
id=i,
|
||||||
|
vector={"colpali": embedding.tolist()},
|
||||||
|
payload=documents[i],
|
||||||
|
)
|
||||||
|
for i, embedding in enumerate(image_embeddings)
|
||||||
|
],
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Search-Time Control with Rescoring
|
||||||
|
|
||||||
|
Qdrant provides **automatic rescoring**: the quantized index quickly finds candidates, then re-ranks them using the original float32 vectors for accuracy.
|
||||||
|
|
||||||
|
**Key search parameters:**
|
||||||
|
- `rescore`: Re-evaluate top candidates with original vectors (default: true)
|
||||||
|
- `oversampling`: Fetch more candidates before rescoring (e.g., 2.0 = fetch 2x results)
|
||||||
|
|
||||||
|
For ColPali searches:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Generate query embeddings from a text query
|
||||||
|
query = "financial quarterly results revenue"
|
||||||
|
query_embeddings = list(model.embed_text([query]))[0]
|
||||||
|
|
||||||
|
# Search with rescoring enabled
|
||||||
|
results = client.query_points(
|
||||||
|
collection_name="colpali-scalar",
|
||||||
|
query=query_embeddings.tolist(),
|
||||||
|
using="colpali",
|
||||||
|
limit=10,
|
||||||
|
search_params=models.SearchParams(
|
||||||
|
quantization=models.QuantizationSearchParams(
|
||||||
|
ignore=False, # Use quantized vectors for initial search
|
||||||
|
rescore=True, # Re-rank with original float32 vectors
|
||||||
|
oversampling=2.0, # Fetch 2x candidates before rescoring
|
||||||
|
),
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
The rescoring step is **critical for multi-vector search** because MaxSim aggregates many token-level similarities - small quantization errors can compound. Rescoring with original vectors ensures your final results maintain high quality.
|
||||||
|
|
||||||
|
## Quantization Impact on Search Quality
|
||||||
|
|
||||||
|
The compression ratios and speed improvements shown above are only part of the story. **The real question is how quantization affects your retrieval quality** - and that answer depends entirely on your specific use case.
|
||||||
|
|
||||||
|
Several factors influence quantization's impact:
|
||||||
|
|
||||||
|
1. **Model dimensionality** - ColModernVBERT's 128-dimensional vectors behave differently under quantization than higher-dimensional models (768+ dims)
|
||||||
|
2. **Document characteristics** - Text-heavy documents, image-heavy pages, and structured forms each respond differently to compression
|
||||||
|
3. **Query patterns** - Keyword-like queries vs. semantic questions may show different sensitivity to quantization errors
|
||||||
|
4. **Quality thresholds** - Your application's tolerance for retrieval quality changes
|
||||||
|
|
||||||
|
Published benchmarks provide general guidance, but **your specific documents and queries will behave differently**. A quantization method that works well for one dataset might perform poorly on another.
|
||||||
|
|
||||||
|
This is why the final lesson in this module focuses entirely on **evaluating multi-vector search pipelines**. You'll learn systematic approaches to measure quantization's impact on your specific use case, helping you make informed trade-offs between memory savings and retrieval quality.
|
||||||
|
|
||||||
|
## What's Next
|
||||||
|
|
||||||
|
In this lesson, you learned how quantization dramatically reduces memory usage and improves brute force search performance for multi-vector representations. Key takeaways:
|
||||||
|
- ColModernVBERT's memory footprint: **512 KB per document** without quantization
|
||||||
|
- Scalar quantization reduces this to **128 KB** (4x compression) with ~2x faster brute force search
|
||||||
|
- Binary quantization can achieve **16-32 GB** total memory for 1 million documents (vs 512 GB uncompressed) with up to 40x faster brute force search
|
||||||
|
- **Quantization is a memory and speed optimization**, not an indexing solution - HNSW remains incompatible with multi-vector search
|
||||||
|
- The search strategy remains exhaustive (brute force), but becomes significantly cheaper and faster
|
||||||
|
|
||||||
|
Crucially, **quantization can be enabled on existing collections** without modifying your ingestion or inference code, making it straightforward to experiment with different approaches.
|
||||||
|
|
||||||
|
However, choosing the right quantization method requires measuring its impact on your specific retrieval quality. We'll cover systematic evaluation approaches in the final lesson of this module.
|
||||||
|
|
||||||
|
Next, we'll explore **pooling techniques** that reduce the number of vectors per document - a complementary approach to reducing memory that works alongside quantization.
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
---
|
||||||
|
#Delimiter files are used to separate the list of documentation pages into sections.
|
||||||
|
type: reference
|
||||||
|
reference: /course/multi-vector-search
|
||||||
|
weight: 220
|
||||||
|
sitemapExclude: True
|
||||||
|
_build:
|
||||||
|
publishResources: false
|
||||||
|
render: never
|
||||||
|
partition: learn
|
||||||
|
---
|
||||||
@@ -1,7 +0,0 @@
|
|||||||
---
|
|
||||||
title: "Multi-Vector Retrieval"
|
|
||||||
weight: 230
|
|
||||||
type: external-link
|
|
||||||
external_url: https://www.deeplearning.ai/short-courses/multi-vector-image-retrieval/
|
|
||||||
partition: learn
|
|
||||||
---
|
|
||||||
|
After Width: | Height: | Size: 10 KiB |
|
After Width: | Height: | Size: 77 KiB |
|
After Width: | Height: | Size: 75 KiB |
|
After Width: | Height: | Size: 152 KiB |
|
After Width: | Height: | Size: 29 KiB |
|
After Width: | Height: | Size: 42 KiB |
|
After Width: | Height: | Size: 28 KiB |
|
After Width: | Height: | Size: 48 KiB |
|
After Width: | Height: | Size: 74 KiB |
|
After Width: | Height: | Size: 340 KiB |
|
After Width: | Height: | Size: 35 KiB |
|
After Width: | Height: | Size: 90 KiB |
|
After Width: | Height: | Size: 574 KiB |
|
After Width: | Height: | Size: 585 KiB |
|
After Width: | Height: | Size: 104 KiB |
|
After Width: | Height: | Size: 732 KiB |
|
After Width: | Height: | Size: 1.4 MiB |
|
After Width: | Height: | Size: 36 KiB |
|
After Width: | Height: | Size: 283 KiB |
|
After Width: | Height: | Size: 96 KiB |
|
After Width: | Height: | Size: 37 KiB |
|
After Width: | Height: | Size: 63 KiB |
|
After Width: | Height: | Size: 80 KiB |
|
After Width: | Height: | Size: 429 KiB |
@@ -21,7 +21,7 @@
|
|||||||
<div class="d-flex justify-content-lg-end pb-3 pt-4">
|
<div class="d-flex justify-content-lg-end pb-3 pt-4">
|
||||||
{{/* Determine the course section */}}
|
{{/* Determine the course section */}}
|
||||||
{{ $subject := "" }}
|
{{ $subject := "" }}
|
||||||
{{ if eq .Params.isLesson true }}
|
{{ if .IsSection }}
|
||||||
{{/* Current page is a lesson (_index.md), parent is the course section */}}
|
{{/* Current page is a lesson (_index.md), parent is the course section */}}
|
||||||
{{ $subject = .Parent }}
|
{{ $subject = .Parent }}
|
||||||
{{ else }}
|
{{ else }}
|
||||||
|
|||||||
@@ -40,7 +40,19 @@
|
|||||||
{{ end }}
|
{{ end }}
|
||||||
|
|
||||||
{{ if or (eq $currentNode.Params.isLesson true) (eq $currentNode.Parent.Params.isLesson true) }}
|
{{ if or (eq $currentNode.Params.isLesson true) (eq $currentNode.Parent.Params.isLesson true) }}
|
||||||
{{ $allLessons := where $subject.Pages ".Params.isLesson" true }}
|
{{ $allLessons := slice }}
|
||||||
|
{{ range $subject.Pages }}
|
||||||
|
{{ if .Params.isLesson }}
|
||||||
|
{{ $allLessons = $allLessons | append . }}
|
||||||
|
{{ end }}
|
||||||
|
{{ if .IsSection }}
|
||||||
|
{{ range .Pages }}
|
||||||
|
{{ if .Params.isLesson }}
|
||||||
|
{{ $allLessons = $allLessons | append . }}
|
||||||
|
{{ end }}
|
||||||
|
{{ end }}
|
||||||
|
{{ end }}
|
||||||
|
{{ end }}
|
||||||
{{ $total := len $allLessons }}
|
{{ $total := len $allLessons }}
|
||||||
|
|
||||||
{{ if $total }}
|
{{ if $total }}
|
||||||
|
|||||||