diff --git a/qdrant-landing/content/course/essentials/_index.md b/qdrant-landing/content/course/essentials/_index.md
index 71604f51d..0190cbc98 100644
--- a/qdrant-landing/content/course/essentials/_index.md
+++ b/qdrant-landing/content/course/essentials/_index.md
@@ -175,7 +175,7 @@ Build the vector search skills that matter: hybrid retrieval, multivector rerank
- ML Platforms & Analytics (Tensorlake, Vectorize.io, Superlinked, Quotient)
-
→ Start day 7
+ → Start Day 7
{{< /accordion >}}
diff --git a/qdrant-landing/content/course/essentials/day-5/colbert-multivectors.md b/qdrant-landing/content/course/essentials/day-5/colbert-multivectors.md
index 735b7b975..8bfe64626 100644
--- a/qdrant-landing/content/course/essentials/day-5/colbert-multivectors.md
+++ b/qdrant-landing/content/course/essentials/day-5/colbert-multivectors.md
@@ -28,7 +28,7 @@ Late-interaction models such as ColBERT retain per-token document vectors. At se
Qdrant implements this powerful technique through [multivector representations](/documentation/concepts/vectors/#multivectors). A multivector field holds an ordered list of subvectors, each of which captures a different token of the document.
-At query time, Qdrant performs the late interaction scoring. It compares every query token embedding $ q_i $ with each document token embedding $ d_j $. Only the highest score per query token is retained and these top scores are then summed. This mechanism, called MaxSim, delivers fine-grained relevance that respects the structure of your content.
+At query time, Qdrant performs late interaction scoring. It compares every query token embedding $ q_i $ with each document token embedding $ d_j $. Only the highest score per query token is retained, and those top scores are then summed. This mechanism, called MaxSim, delivers fine-grained relevance that respects the structure of your content.
$$
MaxSim_{\text{norm}}(Q, D) = \frac{1}{|Q|} \sum_{i=1}^{|Q|} \max_{j=1}^{|D|} \text{sim}(q_i, d_j)
@@ -48,7 +48,7 @@ doc_multivectors = list(encoder.embed(["A long document about AI in medicine."])
# Returns [[token_vec1, token_vec2, ...]]
```
-The model `colbert-ir/colbertv2.0` outputs 128-dimensional vectors and is available through FastEmbed's optimized ONNX runtime. Use `.embed` for documents and `.query_embed` for queries.
+The model `colbert-ir/colbertv2.0` outputs 128-dimensional vectors and is available through FastEmbed's optimized ONNX Runtime. Use `.embed` for documents and `.query_embed` for queries.
## Collection Configuration: Multivector Setup
@@ -80,7 +80,7 @@ client.create_collection(
)
```
-By specifying `MAX_SIM`, you tell Qdrant to apply the late interaction scoring at query time. We explicitly disable HNSW indexing with `m=0` because the graph typically won't be used for multivectors (except in rare edge cases), so disabling it saves RAM. Without HNSW, queries use brute-force MaxSim scoring across all points, which provides maximum precision but may be slower on large collections. For better performance on larger datasets, you'll learn about retrieval-reranking patterns in the next lesson.
+By specifying `MAX_SIM`, you tell Qdrant to apply late interaction scoring at query time. We explicitly disable HNSW indexing with `m=0` because the graph typically won't be used for multivectors except in rare edge cases, so disabling it saves RAM. Without HNSW, queries use brute-force MaxSim scoring across all points, which provides maximum precision but may be slower on large collections. For better performance on larger datasets, you'll learn about retrieval and reranking patterns in the next lesson.
## Querying with ColBERT
@@ -102,13 +102,13 @@ hits = client.query_points(
)
```
-Qdrant performs brute-force MaxSim scoring between your query tokens and the document tokens for all points in the collection. This delivers highly precise results based on fine-grained token-level matching. Keep in mind that without HNSW indexing, this approach may be slower on large collections - in the next lesson you'll see how to combine fast approximate retrieval with ColBERT reranking for better performance.
+Qdrant performs brute-force MaxSim scoring between your query tokens and the document tokens for all points in the collection. This delivers highly precise results based on fine-grained token-level matching. Keep in mind that without HNSW indexing, this approach may be slower on large collections. In the next lesson, you'll see how to combine fast approximate retrieval with ColBERT reranking for better performance.
## ColPali for Visual Documents
-For documents with rich layouts, PDFs, invoices, slide decks, ColPali (Contextualized Late Interaction over PaliGemma) extends the same idea to vision. ColPali divides each page into a 32×32 grid (1,024 patches), encodes each patch with a vision-language model into 128-dimensional vectors, and treats those patch embeddings as subvectors. You use the identical multivector configuration, and Qdrant applies MaxSim on token embeddings, no matter how they were created.
+For documents with rich layouts such as PDFs, invoices, and slide decks, ColPali (Contextualized Late Interaction over PaliGemma) extends the same idea to vision. ColPali divides each page into a 32x32 grid (1,024 patches), encodes each patch with a vision-language model into 128-dimensional vectors, and treats those patch embeddings as subvectors. You use the same multivector configuration, and Qdrant applies MaxSim to those embeddings regardless of how they were created.
The visual approach eliminates traditional OCR and layout detection steps, processing document images directly to capture both textual content and visual structure in a single pass. This makes ColPali particularly effective for complex documents where layout and visual elements are crucial for understanding.
## Next
-With multivectors in your toolkit, you can achieve high-precision retrieval for both text and visual documents. In the next lesson, we'll explore the Universal Query API, where you'll learn how to combine multiple retrieval strategies and use ColBERT for reranking - a more common production pattern that balances speed and precision when working with large collections.
\ No newline at end of file
+With multivectors in your toolkit, you can achieve high-precision retrieval for both text and visual documents. In the next lesson, we'll explore the Universal Query API, where you'll learn how to combine multiple retrieval strategies and use ColBERT for reranking, a more common production pattern that balances speed and precision when working with large collections.
diff --git a/qdrant-landing/content/course/essentials/day-5/universal-query-demo.md b/qdrant-landing/content/course/essentials/day-5/universal-query-demo.md
index 23619a520..1ff241a29 100644
--- a/qdrant-landing/content/course/essentials/day-5/universal-query-demo.md
+++ b/qdrant-landing/content/course/essentials/day-5/universal-query-demo.md
@@ -108,7 +108,7 @@ client.create_payload_index(
## Prepare and Ingest Research Paper Data
-Now that our collection is configured with vectors and payload indexes, let's take some sample research papers:
+Now that our collection is configured with vectors and payload indexes, let's define a few sample research papers:
```python
sample_data = [
@@ -133,13 +133,13 @@ sample_data = [
"open_access": True,
},
{
- "title": "Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace",
- "authors": ["Andre Rusli", "Shoma Ishimoto", "Sho Akiyama", "Aman Kumar Singh"],
- "abstract": "Visual search offers an intuitive way for customers to explore diverse product catalogs, particularly in consumer-to-consumer (C2C) marketplaces where listings are often unstructured and visually driven. This paper presents a scalable visual search system deployed in Mercari's C2C marketplace...",
- "research_area": "computer_vision",
- "published_date": "2025-07-31",
- "impact_score": 0.78,
- "citation_count": 12,
+ "title": "MUVERA: Multi-Vector Retrieval via Fixed Dimensional Encodings",
+ "authors": ["Jason Lee", "Vahab Mirrokni", "Rajesh Jayaram"],
+ "abstract": "We present MUVERA, a retrieval approach that compresses multi-vector representations into fixed-dimensional encodings for efficient first-stage retrieval while preserving the quality benefits of late interaction models. The method reduces serving costs and latency without giving up strong reranking performance...",
+ "research_area": "machine_learning",
+ "published_date": "2024-05-29",
+ "impact_score": 0.84,
+ "citation_count": 27,
"open_access": True,
},
]
@@ -188,7 +188,7 @@ client.upload_points(
)
```
-## Step 3: The Universal Query in Action
+## Step 2: The Universal Query in Action
Let's build a sophisticated research discovery query step by step. We'll orchestrate dense search, sparse search, RRF fusion, and ColBERT reranking - all in a single API call.
@@ -327,11 +327,11 @@ for i, hit in enumerate(response.points or [], 1):
print(f" Score: {hit.score:.4f}\n")
```
-And there you have it - a sophisticated multi-stage research discovery system in a single declarative query!
+And there you have it: a sophisticated multi-stage research discovery system in a single declarative query.
## Real ArXiv Dataset Integration
-Here's how you could populate the collection with real data (if the endpoint wasn't broken):
+Here's how you could populate the collection with real arXiv data:
```python
# ! pip install arxiv
@@ -391,10 +391,10 @@ print(f"Uploaded {len(points)} research papers to collection")
- **Single Request**: Complex multi-stage research discovery in one API call
- **Parallel Execution**: Dense and sparse searches run concurrently
- **Smart Filtering**: Apply research quality filters at optimal stages
-- **Real Data**: Works with actual arXiv dataset and research metadata
+- **Real Data**: Works with actual arXiv data and research metadata
- **Production Ready**: Scales to millions of papers with sub-second latency
-The Universal Query API eliminates the complexity of building multi-turn retrieval systems. What used to require coordination between semantic search engines, keyword systems, and reranking models now happens in a single, optimized request - perfect for academic search, literature reviews, and research recommendation systems.
+The Universal Query API eliminates the complexity of building multi-turn retrieval systems. What used to require coordination between semantic search engines, keyword systems, and reranking models now happens in a single, optimized request, which makes it a good fit for academic search, literature reviews, and research recommendation systems.
## Next