docs: Add course content for Day 5 (#1932)

* docs: Add course content for Day 5

* Apply suggestions from code review

colbert-mulivectors changes from kacper

Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>

* multi vector intro

* integrated kacpers ideas

* universal query api doc fixes and recall to retrieval in all files

* fixed demo

* added discord link

* updated structure

* fix: update colbert-multivectors.md for clarity and formatting improvements

* fix: update universal-query-api.md filter explanation

* fix: split longer cells into pieces in universal-query-demo.md

* fix: detailed query construction steps in universal-query-demo.md

* fix: divide pitstop-project.md into smaller steps

* fix: a helper function for building filters

* fix: refine wording in universal-query-demo.md

* fix: move multivector reranking to the day 5 lesson 2

* fix: add Colab link and badge to universal-query-demo.md

* fix: remove header

* standart init for client

* pitstop Reflect on Your Findings

* pitstop Reflect on Your Findings

* pitstop fix

---------

Co-authored-by: Kirstin <kirstin.taufertshoefer@qdrant.com>
Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com>
Co-authored-by: Kacper Łukawski <lukawski.kacper@gmail.com>
This commit is contained in:
Kirstin
2025-10-22 13:27:16 +02:00
committed by GitHub
co-authored by Kacper Łukawski Kirstin Kacper Łukawski
parent bcaa46c8ef
commit a8f39b3f4b
5 changed files with 1371 additions and 0 deletions
@@ -0,0 +1,23 @@
---
title: "Day 5: Advanced APIs"
isLesson: true
weight: 6
---
{{< date >}} Day 5 {{< /date >}}
# Advanced APIs
Multi‑vector patterns and the Universal Query API.
---
## Today’s path
1. Multivectors for Late Interaction Models
2. The Universal Query API
3. Demo: Universal Query for Hybrid Retrieval
4. Project: Building a Recommendation System
You’ll synthesize these APIs into a practical recommendation system.
@@ -0,0 +1,115 @@
---
title: Multivectors for Late Interaction Models
weight: 1
---
{{< date >}} Day 5 {{< /date >}}
# Multivectors for Late Interaction Models
<div class="video">
<iframe
src="https://www.youtube.com/embed/8ptlXSsSEPk?si=TzsWlastazBQPWWb"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
allowfullscreen>
</iframe>
</div>
<br/>
Many embedding models represent data as a single vector. Transformer-based encoders achieve this by pooling the per-token vector matrix from the final layer into a single vector. That works great for most cases. But when your documents get more complex, cover multiple topics, or require context sensitivity, that one-size-fits-all compression starts to break down. You lose granularity and semantic alignment (though chunking and learned pooling mitigate this to an extent).
Late-interaction models such as ColBERT retain per-token document vectors. At search time, they identify the best matches by comparing all the query tokens with all the document tokens. This token-level precision preserves local semantic matches and delivers superior relevance, especially for complex queries and documents.
## Late Interaction: Token-Level Precision
Qdrant implements this powerful technique through [multivector representations](/documentation/concepts/vectors/#multivectors). A multivector field holds an ordered list of subvectors, each of which captures a different token of the document.
At query time, Qdrant performs the late interaction scoring. It compares every query token embedding $ q_i $ with each document token embedding $ d_j $. Only the highest score per query token is retained and these top scores are then summed. This mechanism, called MaxSim, delivers fine-grained relevance that respects the structure of your content.
$$
MaxSim_{\text{norm}}(Q, D) = \frac{1}{|Q|} \sum_{i=1}^{|Q|} \max_{j=1}^{|D|} \text{sim}(q_i, d_j)
$$
To enable this in Qdrant, you create a collection with a dense vector that has a multivector comparator provided.
## Generating Token-Level Embeddings
You can generate the required token-level embeddings with FastEmbed:
```python
from fastembed import LateInteractionTextEmbedding
encoder = LateInteractionTextEmbedding("colbert-ir/colbertv2.0")
doc_multivectors = list(encoder.embed(["A long document about AI in medicine."]))
# Returns [[token_vec1, token_vec2, ...]]
```
The model `colbert-ir/colbertv2.0` outputs 128-dimensional vectors and is available through FastEmbed's optimized ONNX runtime. Use `.embed` for documents and `.query_embed` for queries.
## Collection Configuration: Multivector Setup
To use ColBERT for retrieval, create a collection with a multivector field configured for late interaction scoring:
```python
from qdrant_client import QdrantClient, models
import os
from dotenv import load_dotenv
load_dotenv()
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
# For Colab:
# from google.colab import userdata
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
client.create_collection(
collection_name="my_colbert_collection",
vectors_config={
"colbert": models.VectorParams(
size=128,
distance=models.Distance.COSINE,
multivector_config=models.MultiVectorConfig(
comparator=models.MultiVectorComparator.MAX_SIM
),
# Disable HNSW to save RAM - it won't typically be used with multivectors
hnsw_config=models.HnswConfigDiff(m=0),
)
}
)
```
By specifying `MAX_SIM`, you tell Qdrant to apply the late interaction scoring at query time. We explicitly disable HNSW indexing with `m=0` because the graph typically won't be used for multivectors (except in rare edge cases), so disabling it saves RAM. Without HNSW, queries use brute-force MaxSim scoring across all points, which provides maximum precision but may be slower on large collections. For better performance on larger datasets, you'll learn about retrieval-reranking patterns in the next lesson.
## Querying with ColBERT
To search using your ColBERT multivector field, embed your query and pass it directly to `query_points`:
```python
from fastembed import LateInteractionTextEmbedding
# Encode the query
colbert = LateInteractionTextEmbedding("colbert-ir/colbertv2.0")
colbert_query = next(colbert.query_embed(["what is the policy?"])).tolist()
# Search using ColBERT multivector
hits = client.query_points(
collection_name="my_colbert_collection",
query=colbert_query,
using="colbert",
limit=20,
)
```
Qdrant performs brute-force MaxSim scoring between your query tokens and the document tokens for all points in the collection. This delivers highly precise results based on fine-grained token-level matching. Keep in mind that without HNSW indexing, this approach may be slower on large collections - in the next lesson you'll see how to combine fast approximate retrieval with ColBERT reranking for better performance.
## ColPali for Visual Documents
For documents with rich layouts, PDFs, invoices, slide decks, ColPali (Contextualized Late Interaction over PaliGemma) extends the same idea to vision. ColPali divides each page into a 32×32 grid (1,024 patches), encodes each patch with a vision-language model into 128-dimensional vectors, and treats those patch embeddings as subvectors. You use the identical multivector configuration, and Qdrant applies MaxSim on token embeddings, no matter how they were created.
The visual approach eliminates traditional OCR and layout detection steps, processing document images directly to capture both textual content and visual structure in a single pass. This makes ColPali particularly effective for complex documents where layout and visual elements are crucial for understanding.
## Next
With multivectors in your toolkit, you can achieve high-precision retrieval for both text and visual documents. In the next lesson, we'll explore the Universal Query API, where you'll learn how to combine multiple retrieval strategies and use ColBERT for reranking - a more common production pattern that balances speed and precision when working with large collections.
@@ -0,0 +1,659 @@
---
title: "Project: Building a Recommendation System"
weight: 4
---
{{< date >}} Day 5 {{< /date >}}
# Project: Building a Recommendation System
Bring together dense, sparse, and multivectors in one atomic Universal Query. You'll retrieve candidates, fuse signals, rerank with ColBERT, and apply business filters - in a single request.
## Your Mission
Build a complete recommendation system using Qdrant’s Universal Query API with dense, sparse, and ColBERT multivectors in one request.
**Estimated Time:** 90 minutes
## What You'll Build
A hybrid recommendation system using:
- **Multi-vector architecture** with dense, sparse, and ColBERT vectors
- **Universal Query API** for atomic multi-stage search
- **RRF fusion** for combining candidates
- **ColBERT reranking** for fine-grained relevance scoring
- **Business rule filtering** at multiple pipeline stages
- **Production-ready patterns** for recommendation systems
## Setup
### Prerequisites
* Qdrant Cloud cluster (URL + API key)
* Python 3.9+ (or Google Colab)
* Packages: `qdrant-client`, `fastembed`, `python-dotenv`
### Models
- **Dense**: `sentence-transformers/all-MiniLM-L6-v2` (384-dim)
- **Sparse**: `prithivida/Splade_PP_en_v1` (SPLADE)
- **Multivector**: `colbert-ir/colbertv2.0` (128-dim tokens)
### Dataset
- **Scope**: A small set of sample items (e.g., 10-20 movies).
- **Payload Fields**: `title`, `description`, `category`, `genre`, `year`, `rating`, `user_segment`, `popularity_score`, `release_date`.
- **Filters Used**: `category`, `user_segment`, `release_date`, `popularity_score`.
## Build Steps
### Step 1: Set Up the Hybrid Collection
#### Initialize Client and Collection
First, connect to Qdrant and create a clean collection for our recommendation system:
```python
from datetime import datetime
from qdrant_client import QdrantClient, models
import os
from dotenv import load_dotenv
load_dotenv()
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
# For Colab:
# from google.colab import userdata
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
collection_name = "day5_recommendations_hybrid"
# Clean state
if client.collection_exists(collection_name=collection_name):
client.delete_collection(collection_name=collection_name)
```
Now configure a collection with three vectors - each serving a different purpose in our recommendation pipeline:
```python
client.create_collection(
collection_name=collection_name,
vectors_config={
# Dense vectors for semantic understanding
"dense": models.VectorParams(size=384, distance=models.Distance.COSINE),
# ColBERT multivectors for fine-grained reranking
"colbert": models.VectorParams(
size=128,
distance=models.Distance.COSINE,
multivector_config=models.MultiVectorConfig(
comparator=models.MultiVectorComparator.MAX_SIM
),
hnsw_config=models.HnswConfigDiff(
m=0 # Disable HNSW - used only for reranking
),
),
},
sparse_vectors_config={
# Sparse vectors for exact keyword matching
"sparse": models.SparseVectorParams(
index=models.SparseIndexParams(on_disk=False)
)
},
)
```
**Why this setup**: ColBERT uses `MAX_SIM` for token-level comparison and `m=0` since it's only used for reranking, not initial retrieval, and setting `m=0` effectively disables HNSW indexing to save memory and compute.
#### Create Payload Indexes
Before ingesting data, create indexes for the fields we'll filter by. This enables efficient filtering during vector search:
```python
# Business metadata indexes
client.create_payload_index(
collection_name=collection_name,
field_name="category",
field_schema="keyword",
)
client.create_payload_index(
collection_name=collection_name,
field_name="user_segment",
field_schema="keyword",
)
# Quality and recency indexes
client.create_payload_index(
collection_name=collection_name,
field_name="release_date",
field_schema="datetime",
)
client.create_payload_index(
collection_name=collection_name,
field_name="popularity_score",
field_schema="float",
)
client.create_payload_index(
collection_name=collection_name,
field_name="rating",
field_schema="float",
)
```
### Step 2: Prepare and Upload Recommendation Data
#### Create Sample Recommendation Data
Let's create sample movie data with business metadata for filtering:
```python
# Example: Create sample movie/content data
sample_data = [
{
"title": "The Matrix",
"description": "A computer hacker learns about the true nature of reality and his role in the war against its controllers.",
"category": "movie",
"genre": ["sci-fi", "action"],
"year": 1999,
"rating": 8.7,
"user_segment": "premium",
"popularity_score": 0.9,
"release_date": "1999-03-31T00:00:00Z",
},
# Add more sample data...
]
texts = [it["description"] for it in sample_data]
```
#### Initialize Embedding Models
We'll use FastEmbed to generate all three vector types:
```python
from fastembed import TextEmbedding, SparseTextEmbedding, LateInteractionTextEmbedding
# Model configurations
DENSE_MODEL_ID = "sentence-transformers/all-MiniLM-L6-v2" # 384-dim
SPARSE_MODEL_ID = "prithivida/Splade_PP_en_v1" # SPLADE sparse
COLBERT_MODEL_ID = "colbert-ir/colbertv2.0" # 128-dim multivector
dense_model = TextEmbedding(DENSE_MODEL_ID)
sparse_model = SparseTextEmbedding(SPARSE_MODEL_ID)
colbert_model = LateInteractionTextEmbedding(COLBERT_MODEL_ID)
```
#### Generate Embeddings
Encode all descriptions with all three embedding models:
```python
# Generate embeddings for all items
dense_embeds = list(
dense_model.embed(texts, parallel=0)
) # list[np.ndarray] shape (384,)
sparse_embeds = list(
sparse_model.embed(texts, parallel=0)
) # list[SparseEmbedding] with .indices/.values
colbert_multivectors = list(
colbert_model.embed(texts, parallel=0)
) # list[np.ndarray] shape (tokens, 128)
```
#### Create Points and Upload
Now package everything into points and upload to Qdrant:
```python
# Generate vectors for each item
points = []
for i, item in enumerate(sample_data):
# Create sparse vector (keyword matching)
sparse_vector = sparse_embeds[i].as_object()
# Create dense vector (semantic understanding)
dense_vector = dense_embeds[i]
# Create ColBERT multivector (token-level understanding)
colbert_vector = colbert_multivectors[i]
points.append(
models.PointStruct(
id=i,
vector={
"dense": dense_vector,
"sparse": sparse_vector,
"colbert": colbert_vector,
},
payload=item,
)
)
client.upload_points(collection_name=collection_name, points=points)
print(f"Uploaded {len(points)} recommendation items")
```
### Step 3: One Universal Query — Retrieve → Fuse → Rerank → Filter
Now let's build a complete multi-stage search that combines everything in a single request.
#### Prepare Query Embeddings
First, encode the user's search intent using all three embedding models:
```python
from datetime import timedelta
# Example user intent
user_query = "premium user likes sci-fi action movies with strong hacker themes"
user_dense_vector = next(dense_model.query_embed(user_query))
user_sparse_vector = next(sparse_model.query_embed(user_query)).as_object()
user_multivector = next(colbert_model.query_embed(user_query))
```
#### Define Global Filter with Automatic Propagation
Create a single filter that will automatically propagate to all stages of the pipeline:
```python
# Global filter - this will be propagated to ALL prefetch stages
global_filter = models.Filter(
must=[
# Content type and user segment
models.FieldCondition(
key="category", match=models.MatchValue(value="movie")
),
models.FieldCondition(
key="user_segment", match=models.MatchValue(value="premium")
),
# Quality and recency constraints
models.FieldCondition(
key="release_date",
range=models.DatetimeRange(
gte=(datetime.now() - timedelta(days=365 * 30)).isoformat()
),
),
models.FieldCondition(key="popularity_score", range=models.Range(gte=0.7)),
]
)
```
**Key insight**: This filter automatically propagates to ALL prefetch stages. Qdrant doesn't do "late filtering" or "post-filtering" - filters are applied at the HNSW search level for maximum efficiency.
#### Set Up Prefetch and Fusion
Configure parallel retrieval and fusion strategy:
```python
# Prefetch queries - global filter will be automatically applied to both
hybrid_query = [
models.Prefetch(query=user_dense_vector, using="dense", limit=100),
models.Prefetch(query=user_sparse_vector, using="sparse", limit=100),
]
# Fusion stage - combine candidates with RRF
fusion_query = models.Prefetch(
prefetch=hybrid_query,
query=models.FusionQuery(fusion=models.Fusion.RRF),
limit=100,
)
```
These prefetch queries run in parallel, and the global filter from the main query will automatically propagate to both dense and sparse searches.
#### Execute Universal Query with Reranking
Finally, send the complete query with ColBERT reranking:
```python
# The Universal Query: Global filter propagates through all stages
response = client.query_points(
collection_name=collection_name,
prefetch=fusion_query,
query=user_multivector,
using="colbert",
query_filter=global_filter, # Propagates to all prefetch stages
limit=10,
with_payload=True,
)
for hit in response.points or []:
print(hit.payload)
```
**Why this works**: Dense and sparse retrieval happen in parallel (with filters applied), RRF fuses the results, ColBERT reranks with token-level precision, and the global filter is applied at every stage via automatic propagation - all in one atomic API call.
### Step 4: Build a Recommendation Service
Now let's wrap everything into a production-ready function that handles dynamic user preferences.
#### Helper Function for Building Filters
First, create a reusable helper function to build filters from user profiles:
```python
def build_recommendation_filter(user_profile, user_preference=None):
"""
Build a global filter from user profile and preferences.
This filter will automatically propagate to all prefetch stages.
Args:
user_profile: {
"liked_titles": list[str], # optional
"preferred_genres": list[str], # e.g. ["sci-fi","action"]
"segment": str, # e.g. "premium"
"query": str # free-text intent, e.g. "smart sci-fi with hacker vibe"
}
user_preference: {
"category": str | None, # e.g. "movie"
"min_rating": float | None, # e.g. 8.0
"released_within_days": int | None # e.g. 365
}
Returns:
models.Filter object or None if no conditions
"""
from datetime import datetime, timedelta
filter_conditions = []
# User segment filtering
if user_profile.get("segment"):
filter_conditions.append(
models.FieldCondition(
key="user_segment",
match=models.MatchValue(value=user_profile["segment"]),
)
)
if user_preference:
# Category filtering
if user_preference.get("category"):
filter_conditions.append(
models.FieldCondition(
key="category",
match=models.MatchValue(value=user_preference["category"]),
)
)
# Rating filtering
if user_preference.get("min_rating") is not None:
filter_conditions.append(
models.FieldCondition(
key="rating",
range=models.Range(gte=user_preference["min_rating"])
)
)
# Recency filtering
if user_preference.get("released_within_days"):
days = int(user_preference["released_within_days"])
filter_conditions.append(
models.FieldCondition(
key="release_date",
range=models.DatetimeRange(
gte=(datetime.utcnow() - timedelta(days=days)).isoformat()
),
)
)
return models.Filter(must=filter_conditions) if filter_conditions else None
```
#### Recommendation Function
Now create the main recommendation function using the helper:
```python
def get_recommendations(user_profile, user_preference=None, limit=10):
"""
Get personalized recommendations using Universal Query API
Args:
user_profile: {
"liked_titles": list[str], # optional
"preferred_genres": list[str], # e.g. ["sci-fi","action"]
"segment": str, # e.g. "premium"
"query": str # free-text intent, e.g. "smart sci-fi with hacker vibe"
}
user_preference: {
"category": str | None, # e.g. "movie"
"min_rating": float | None, # e.g. 8.0
"released_within_days": int | None # e.g. 365
}
limit: top-k to return
"""
# Generate query embeddings
user_dense_vector = next(dense_model.query_embed(user_profile["query"]))
user_sparse_vector = next(
sparse_model.query_embed(user_profile["query"])
).as_object()
user_multivector = next(colbert_model.query_embed(user_profile["query"]))
# Build global filter using helper function
global_filter = build_recommendation_filter(user_profile, user_preference)
# Prefetch queries - global filter will propagate automatically
hybrid_query = [
models.Prefetch(query=user_dense_vector, using="dense", limit=100),
models.Prefetch(query=user_sparse_vector, using="sparse", limit=100),
]
# Combine candidates with RRF
fusion_query = models.Prefetch(
prefetch=hybrid_query,
query=models.FusionQuery(fusion=models.Fusion.RRF),
limit=100,
)
# Universal query - global filter propagates to all stages
response = client.query_points(
collection_name=collection_name,
prefetch=fusion_query,
query=user_multivector,
using="colbert",
query_filter=global_filter, # Propagates to all prefetch stages
limit=limit,
with_payload=True,
)
return [
{
"title": hit.payload["title"],
"description": hit.payload["description"],
"score": hit.score,
"metadata": {
k: v
for k, v in hit.payload.items()
if k not in ["title", "description"]
},
}
for hit in (response.points or [])
]
```
#### Test the Service
Let's test the recommendation function:
```python
# Test the recommendation service
user_profile = {
"liked_titles": ["The Matrix", "Blade Runner"],
"preferred_genres": ["sci-fi", "action"],
"segment": "premium",
"query": "highly rated cyberpunk movies",
}
recommendations = get_recommendations(
user_profile,
user_preference={
"category": "movie",
"min_rating": 8.0,
"released_within_days": 365 * 30,
},
limit=10,
)
for i, rec in enumerate(recommendations, 1):
print(f"{i}. {rec['title']} (Score: {rec['score']:.3f})")
```
**What happened under the hood**: Qdrant retrieved 100 candidates from dense and 100 from sparse in parallel, fused them with RRF, reranked with ColBERT's MaxSim over token‑level subvectors, applied final business filters, and returned the top 10 - all in one call.
## Success Criteria
You'll know you've succeeded when:
<input type="checkbox"> Your collection contains dense, sparse, and ColBERT vectors
<input type="checkbox"> You can execute complex multi-stage searches in a single API call
<input type="checkbox"> RRF fusion effectively combines different vector types
<input type="checkbox"> ColBERT reranking improves result relevance
<input type="checkbox"> Business filters propagate automatically to all prefetch stages
<input type="checkbox"> Your recommendation service provides personalized, high-quality results
## Share Your Discovery
### Step 1: Reflect on Your Findings
1. How did the Universal Query API simplify your recommendation pipeline (fewer calls, fewer joins, less code)?
2. Which fusion strategy (RRF vs. DBSF) ranked better for your queries?
3. What was the effect of ColBERT reranking on recommendation quality (top-k changes, click-like signals, nDCG/MRR)?
4. What was the latency impact of multi-stage filtering (prefetch, rerank, total)?
### Step 2: Post Your Results
**Post your results in** <a href="https://discord.com/invite/qdrant" target="_blank" rel="noopener noreferrer" aria-label="Qdrant Discord"> <img src="https://img.shields.io/badge/Qdrant%20Discord-5865F2?style=flat&logo=discord&logoColor=white&labelColor=5865F2&color=5865F2"
alt="Post your results in Discord"
style="display:inline; margin:0; vertical-align:middle; border-radius:9999px;" /> </a> **using this:**
```markdown
**[Day N] Recommendations with the Universal Query API**
**High-Level Summary**
- **Domain:** "I built recommendations for [domain]"
- **Key Result:** "Using [RRF/DBSF] + ColBERT rerank improved [metric] to [value] with total latency [Z] ms."
**Reproducibility**
- **Collection:** [name]
- **Models:** dense=[id, dim], sparse=[method], colbert=[id, dim]
- **Dataset:** [N items] (snapshot: YYYY-MM-DD)
**Settings (today)**
- **Fusion:** [RRF/DBSF], k_dense=[..], k_sparse=[..]
- **Reranker:** ColBERT (MaxSim), top-k=[..]
- **Filters:** category=[...], segment=[...], rating≥..., release_date≥...
- **Filter propagation:** applied to all prefetch stages
**Results (demo query: "[user intent]")**
- **Top picks (rank → title → score):**
1) ...
2) ...
3) ...
- **Why these won:** [token hit like “hacker”, genre match, rating signal]
- **Latency:** prefetch ~[X] ms | rerank ~[Y] ms | total ~[Z] ms
- **Dropped by rules:** [ids/titles → which rule]
**Surprise**
- "[one thing you didn’t expect]"
**Next step**
- "[what you’ll try next]"
```
## Optional: Go Further
### Experiment with Fusion Strategies
Compare RRF with Distribution-Based Score Fusion:
```python
# Test DBSF vs RRF
# Filters and other prefetch params same as above
fusion_query = models.Prefetch(
prefetch=hybrid_query,
query=models.FusionQuery(fusion=models.Fusion.DBSF),
limit=100,
)
response = client.query_points(
collection_name=collection_name,
prefetch=fusion_query,
query=user_multivector,
using="colbert",
query_filter=global_filter, # Same global filter propagates to all stages
limit=10,
with_payload=True,
)
for hit in response.points or []:
print(hit.payload)
```
### A/B Testing Framework
Build a framework to systematically compare fusion strategies across multiple user profiles.
#### Initialize Testing Function
Set up the function structure and results tracking:
```python
def ab_test_fusion_strategies(user_profiles, user_preferences):
"""Compare RRF vs DBSF performance"""
results = {"RRF": [], "DBSF": []}
for user_profile, user_preference in zip(user_profiles, user_preferences):
for fusion_type in [models.Fusion.RRF, models.Fusion.DBSF]:
# Run recommendation query
user_dense_vector = next(dense_model.query_embed(user_profile["query"]))
user_sparse_vector = next(
sparse_model.query_embed(user_profile["query"])
).as_object()
user_multivector = next(colbert_model.query_embed(user_profile["query"]))
# Build global filter using helper function
global_filter = build_recommendation_filter(user_profile, user_preference)
# Prefetch queries - global filter will propagate automatically
hybrid_query = [
models.Prefetch(query=user_dense_vector, using="dense", limit=100),
models.Prefetch(query=user_sparse_vector, using="sparse", limit=100),
]
# Combine candidates with chosen fusion strategy
fusion_query = models.Prefetch(
prefetch=hybrid_query,
query=models.FusionQuery(fusion=fusion_type),
limit=100,
)
response = client.query_points(
collection_name=collection_name,
prefetch=fusion_query,
query=user_multivector,
using="colbert",
query_filter=global_filter, # Propagates to all prefetch stages
limit=10,
with_payload=True,
)
strategy_name = "RRF" if fusion_type == models.Fusion.RRF else "DBSF"
results[strategy_name].append(
{
"user_id": user_profile["user_id"],
"recommendations": [hit.id for hit in response.points],
"scores": [hit.score for hit in response.points],
}
)
return results
```
## Next Steps
Turn this into a mini‑service: ingest your own items and user signals, write a tiny function that takes profile vectors and returns top‑k via the Universal Query API, then experiment with filters and fusion strategies (try DBSF vs RRF) to see how ranking shifts.
@@ -0,0 +1,172 @@
---
title: The Universal Query API
weight: 2
---
{{< date >}} Day 5 {{< /date >}}
# The Universal Query API
Picture this: a customer types "leather jackets" into your store's search bar. You want to show items that match the style semantically - so a bomber jacket surfaces even if it doesn't mention "leather jackets" verbatim - but you also need to enforce your business rules. Only products under $200, only items in stock, only jackets released within the past year. Traditionally, you'd fire off a search, gather results, then apply filters and glue code. With Qdrant's [Universal Query API](/documentation/concepts/hybrid-queries/), all of that happens in one declarative request.
## Run dense + sparse retrieval in parallel with RRF
First, you retrieve candidates from multiple sources in parallel and fuse their ranks. Below, we blend dense semantics from a BGE model with sparse keyword matching from SPLADE by using [Reciprocal Rank Fusion](/documentation/concepts/hybrid-queries/#hybrid-search) to merge the two lists:
```python
from qdrant_client import QdrantClient, models
import os
from dotenv import load_dotenv
load_dotenv()
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
# For Colab:
# from google.colab import userdata
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
response = client.query_points(
collection_name="products",
prefetch=[
models.Prefetch(
query=dense_vector,
using="dense_bge",
limit=20
),
models.Prefetch(
query=sparse_vector,
using="sparse_splade",
limit=20
)
],
query=models.FusionQuery(fusion=models.Fusion.RRF),
limit=10
)
```
Qdrant sends both prefetches concurrently, fuses the two ranked lists by reciprocal rank, and returns your top ten products that satisfy both semantic and keyword relevance.
## Dense Retrieval + ColBERT Rerank
While the previous lesson showed how to use ColBERT directly for retrieval, **in production systems reranking is the more common pattern**. The reason is practical: ColBERT's brute-force MaxSim scoring on an entire large collection can be slow. Instead, you combine two strengths - fast approximate search to narrow candidates, then precise token-level scoring on that smaller set.
You create a collection with two vector fields: a dense vector with HNSW indexing for speed, and a ColBERT multivector with HNSW disabled for precision:
```python
client.create_collection(
collection_name="articles",
vectors_config={
# Fast HNSW-indexed dense retrieval
"bge-dense": models.VectorParams(
size=384,
distance=models.Distance.COSINE,
),
# Precise multivector reranking (HNSW disabled to save RAM)
"colbert": models.VectorParams(
size=128,
distance=models.Distance.COSINE,
multivector_config=models.MultiVectorConfig(
comparator=models.MultiVectorComparator.MAX_SIM
),
hnsw_config=models.HnswConfigDiff(m=0),
)
}
)
```
Now you retrieve 100 candidates quickly with your HNSW-indexed dense field, then apply ColBERT's MaxSim scoring to rerank those hundred and select the very best ten:
```python
from fastembed import LateInteractionTextEmbedding, TextEmbedding
# Encode with both models
dense = TextEmbedding("BAAI/bge-small-en-v1.5")
dense_query_vector = next(dense.query_embed(["what is the policy?"])).tolist()
colbert = LateInteractionTextEmbedding("colbert-ir/colbertv2.0")
colbert_query_multivector = next(colbert.query_embed(["what is the policy?"])).tolist()
# Fast retrieval + precise reranking in one call
response = client.query_points(
collection_name="articles",
prefetch=[
models.Prefetch(
query=dense_query_vector,
using="bge-dense",
limit=100
)
],
query=colbert_query_multivector,
using="colbert",
limit=10
)
```
Behind the scenes, Qdrant fetches the 100 nearest points from the HNSW-indexed dense field, then applies the MaxSim late-interaction score from your ColBERT multivector to only those 100 candidates, returning the ten highest-scoring results. This two-stage approach delivers ColBERT's precision while keeping query latency practical for large-scale deployments.
## Global and Prefetch-Specific Filters
Finally, you layer in filtering wherever it makes sense. In the snippet below, you specify global filters at the query level - price under $200 and release date after January 1, 2023 - which automatically propagate to all prefetches. Then, you add an additional prefetch-specific filter on the dense search to only retrieve products that are in stock and in the "jackets" category:
```python
response = client.query_points(
collection_name="products",
prefetch=[
models.Prefetch(
query=dense_query_vector,
using="bge-dense",
limit=100,
filter=models.Filter(
must=[
models.FieldCondition(
key="in_stock",
match=models.MatchValue(value=True)
),
models.FieldCondition(
key="category",
match=models.MatchValue(value="jackets")
)
]
)
),
models.Prefetch(
query=sparse_query_vector,
using="sparse-splade",
limit=100
)
],
query=models.FusionQuery(fusion=models.Fusion.RRF),
filter=models.Filter(
must=[
models.FieldCondition(
key="price",
range=models.Range(lt=200.0)
),
models.FieldCondition(
key="release_date",
range=models.DatetimeRange(
gte="2023-01-01T00:00:00Z"
)
)
]
),
limit=10
)
```
The global filters (price and release_date) apply to both prefetches automatically. The first prefetch adds extra constraints (in_stock and category), while the second prefetch only uses the global filters. This eliminates the need to repeat common filters across every prefetch. All filtering happens efficiently during the retrieval phase - there's no separate post-processing step. The entire pipeline - hybrid retrieval, filtering, fusion, and reranking - executes in one API call.
## The Universal Query API Structure
The Universal Query API enables complex search patterns through a simple declarative interface:
- **Prefetch Stage**: Execute multiple searches in parallel against different vector fields. Each prefetch can have its own filters, limits, and vector types used.
- **Fusion Stage**: Combine results from multiple prefetches using algorithms like Reciprocal Rank Fusion (RRF) or Distribution-Based Score Fusion (DBSF).
- **Reranking Stage**: Or rerank candidates from a single prefetch with a stronger scorer such as ColBERT. Fusion and reranking are alternative final steps in most pipelines.
- **Filtering**: Apply filters globally at the query level (propagated to all prefetches) or add prefetch-specific filters for additional constraints on individual searches.
This architecture eliminates the need for multiple API calls, client-side result merging, and complex orchestration code. Everything happens server-side in a single, optimized request.
## Next
In our next lesson, you'll bring these building blocks together in a personalized recommendation pipeline. You'll learn how to retrieve candidate from three different vector representations, fuse their signals, rerank with ColBERT, apply user-segment filters, and deliver highly relevant suggestions without writing a single line of glue code.
@@ -0,0 +1,402 @@
---
title: "Demo: Universal Query for Hybrid Retrieval"
weight: 3
---
{{< date >}} Day 5 {{< /date >}}
# Demo: Universal Query for Hybrid Retrieval
In this hands-on demo, we'll build a research paper discovery system using the arXiv dataset that showcases the full power of Qdrant's Universal Query API. You'll see how to combine dense semantics, sparse keywords, and ColBERT reranking to help researchers find exactly the papers they need - all in a single query.
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course/day_5/universal-query-demo.ipynb">
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
</a>
## The Challenge: Intelligent Research Discovery
Imagine you're a machine learning researcher looking for "transformer architectures for multimodal learning with attention mechanisms." You need to:
1. **Retrieve broadly** using semantic understanding of research concepts (dense vectors)
2. **Match precisely** on technical terms like "transformer" and "attention" (sparse vectors)
3. **Rerank intelligently** using fine-grained text understanding (ColBERT)
4. **Apply research filters** like publication date, citation count, and research domain
Traditionally, this would require multiple searches across different systems, manual result merging, and complex ranking logic. With the Universal Query API, it's one declarative request.
## Step 1: Create the Research Paper Collection
### Initialize the Collection with Vector Configurations
First, let's set up a collection with three vector types - each serving a different purpose in our research discovery pipeline:
```python
from datetime import datetime, timedelta
from qdrant_client import QdrantClient, models
import os
from dotenv import load_dotenv
load_dotenv()
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
# For Colab:
# from google.colab import userdata
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
collection_name = "research-papers"
# Clean state
if client.collection_exists(collection_name=collection_name):
client.delete_collection(collection_name=collection_name)
# Create collection with three vector types
client.create_collection(
collection_name=collection_name,
vectors_config={
# Dense vectors for semantic understanding of research concepts
"dense": models.VectorParams(size=384, distance=models.Distance.COSINE),
# ColBERT multivectors for fine-grained text understanding
"colbert": models.VectorParams(
size=128,
distance=models.Distance.COSINE,
multivector_config=models.MultiVectorConfig(
comparator=models.MultiVectorComparator.MAX_SIM
),
),
},
sparse_vectors_config={
# Sparse vectors for exact technical term matching
"sparse": models.SparseVectorParams(
index=models.SparseIndexParams(on_disk=False)
)
},
)
```
### Create Payload Indexes
Before ingesting any data, we create payload indexes for the fields we'll filter by. Qdrant's version of HNSW incorporates payload filtering directly into the search process for efficiency.
```python
# Index fields that will be used for filtering
client.create_payload_index(
collection_name=collection_name,
field_name="research_area",
field_schema="keyword", # For filtering by domain (ML, CV, NLP)
)
client.create_payload_index(
collection_name=collection_name,
field_name="open_access",
field_schema="bool", # For filtering open access papers
)
client.create_payload_index(
collection_name=collection_name,
field_name="published_date",
field_schema="datetime",
)
client.create_payload_index(
collection_name=collection_name,
field_name="impact_score",
field_schema="float",
)
client.create_payload_index(
collection_name=collection_name,
field_name="citation_count",
field_schema="integer",
)
```
## Prepare and Ingest Research Paper Data
Now that our collection is configured with vectors and payload indexes, let's take some sample research papers:
```python
sample_data = [
{
"title": "Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace",
"authors": ["Andre Rusli", "Shoma Ishimoto", "Sho Akiyama", "Aman Kumar Singh"],
"abstract": "Visual search offers an intuitive way for customers to explore diverse product catalogs, particularly in consumer-to-consumer (C2C) marketplaces where listings are often unstructured and visually driven. This paper presents a scalable visual search system deployed in Mercari's C2C marketplace...",
"research_area": "computer_vision",
"published_date": "2025-07-31",
"impact_score": 0.78,
"citation_count": 12,
"open_access": True,
},
{
"title": "TALI: Towards A Lightweight Information Retrieval Framework for Neural Search",
"authors": ["Chaoqun Liu", "Yuanming Zhang", "Jianmin Zhang", "Jiawei Han"],
"abstract": "Neural search systems have emerged as a promising approach to enhance user engagement in information retrieval. However, their high computational costs and memory usage have limited their widespread adoption. In this paper, we present TALI, a lightweight information retrieval framework for neural search that efficiently addresses these challenges...",
"research_area": "machine_learning",
"published_date": "2025-07-31",
"impact_score": 0.78,
"citation_count": 12,
"open_access": True,
},
{
"title": "Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace",
"authors": ["Andre Rusli", "Shoma Ishimoto", "Sho Akiyama", "Aman Kumar Singh"],
"abstract": "Visual search offers an intuitive way for customers to explore diverse product catalogs, particularly in consumer-to-consumer (C2C) marketplaces where listings are often unstructured and visually driven. This paper presents a scalable visual search system deployed in Mercari's C2C marketplace...",
"research_area": "computer_vision",
"published_date": "2025-07-31",
"impact_score": 0.78,
"citation_count": 12,
"open_access": True,
},
]
texts = [it["abstract"] for it in sample_data]
```
We'll use FastEmbed to generate dense, sparse, and ColBERT embeddings for the abstracts. Then we upload everything to Qdrant:
```python
from fastembed import TextEmbedding, SparseTextEmbedding, LateInteractionTextEmbedding
DENSE_MODEL_ID = "sentence-transformers/all-MiniLM-L6-v2" # 384-dim
SPARSE_MODEL_ID = "prithivida/Splade_PP_en_v1" # SPLADE sparse
COLBERT_MODEL_ID = "colbert-ir/colbertv2.0" # 128-dim multivector
dense_model = TextEmbedding(DENSE_MODEL_ID)
sparse_model = SparseTextEmbedding(SPARSE_MODEL_ID)
colbert_model = LateInteractionTextEmbedding(COLBERT_MODEL_ID)
dense_embeds = list(dense_model.embed(texts, parallel=0))
sparse_embeds = list(sparse_model.embed(texts, parallel=0))
colbert_multivectors = list(colbert_model.embed(texts, parallel=0))
points = []
for i, text in enumerate(texts):
sparse_embed = sparse_embeds[i].as_object()
dense_embed = dense_embeds[i]
colbert_embed = colbert_multivectors[i]
points.append(
models.PointStruct(
id=i,
vector={
"dense": dense_embed,
"sparse": sparse_embed,
"colbert": colbert_embed,
},
payload=sample_data[i],
)
)
client.upload_points(
collection_name=collection_name,
points=points,
)
```
## Step 3: The Universal Query in Action
Let's build a sophisticated research discovery query step by step. We'll orchestrate dense search, sparse search, RRF fusion, and ColBERT reranking - all in a single API call.
### Prepare Query Embeddings
First, we encode our research query using all three embedding models:
```python
research_query = "transformer architectures for multimodal learning"
research_query_dense = next(dense_model.query_embed(research_query))
research_query_sparse = next(sparse_model.query_embed(research_query)).as_object()
research_query_colbert = next(colbert_model.query_embed(research_query))
```
We generate three different representations of the same query - each optimized for a different stage of our retrieval pipeline.
### Define Global Filter with Automatic Propagation
Now we define quality constraints that will apply throughout our entire search pipeline:
```python
# Define global filter - this will be propagated to ALL prefetch stages
global_filter = models.Filter(
must=[
# Research domain filtering
models.FieldCondition(
key="research_area",
match=models.MatchAny(any=[
"machine_learning",
"computer_vision",
"nlp",
]),
),
# Open access only
models.FieldCondition(
key="open_access",
match=models.MatchValue(value=True)
),
# Recent research only (last 6 years)
models.FieldCondition(
key="published_date",
range=models.DatetimeRange(
gte=(datetime.now() - timedelta(days=365 * 6)).isoformat()
),
),
# High-impact papers
models.FieldCondition(key="impact_score", range=models.Range(gte=0.6)),
# Well-cited work
models.FieldCondition(key="citation_count", range=models.Range(gte=5)),
]
)
```
**Key insight**: This filter will be automatically propagated to ALL prefetch stages. Qdrant doesn't do "late filtering" or "post-filtering" - the filters are applied at the HNSW search level for maximum efficiency. This is enabled by the payload indexes we created in Step 1.
### Set Up Parallel Prefetch Queries
Next, we configure our hybrid retrieval with two concurrent searches:
```python
# Prefetch queries - global filter will be automatically applied to both
hybrid_query = [
# Dense retrieval: semantic understanding
models.Prefetch(query=research_query_dense, using="dense", limit=100),
# Sparse retrieval: exact technical term matching
models.Prefetch(query=research_query_sparse, using="sparse", limit=100),
]
```
These two prefetch queries run in parallel:
- **Dense search**: Finds semantically similar papers based on research concepts - but only considers papers matching our global filter (ML/CV/NLP domains, open access, recent, high-impact, well-cited)
- **Sparse search**: Matches exact technical terms like "transformer," "multimodal," "attention" - but only within the same filtered subset
- Both searches execute concurrently for maximum speed
- Filter propagation happens automatically - no manual coordination needed
We could retrieve up to 200 candidates total (100 from each search), but many will overlap.
### Add Fusion Stage
Now we combine the two filtered candidate lists using Reciprocal Rank Fusion:
```python
# Fusion stage - combines dense and sparse results
fusion_query = models.Prefetch(
prefetch=hybrid_query,
query=models.FusionQuery(fusion=models.Fusion.RRF),
limit=100,
)
```
The RRF algorithm:
- Combines the two filtered candidate lists intelligently
- Papers appearing high in both lists get boosted scores
- Creates a unified ranking that balances semantic similarity with technical term relevance
- All results still satisfy the global filter constraints
### Execute the Universal Query with ColBERT Reranking
Finally, we send our complete query that ties everything together:
```python
# The Universal Query: Global filter propagates through all stages
response = client.query_points(
collection_name=collection_name,
prefetch=fusion_query,
query=research_query_colbert,
using="colbert",
query_filter=global_filter, # Propagates to all prefetch stages
limit=10,
with_payload=True,
)
```
This final stage applies ColBERT reranking to the fused results:
- Token-level late interaction scoring examines fine-grained text alignment
- Each query token is compared against each abstract token
- MaxSim aggregation finds the best conceptual alignments
- Returns the top 10 papers with precise relevance scores
**Why this matters**: By applying filters at every stage (not after retrieval), Qdrant maintains high accuracy while avoiding wasted computation on papers that would be filtered out anyway.
### Display Results
```python
# Display results
print("Top Research Papers:")
for i, hit in enumerate(response.points or [], 1):
paper = hit.payload
print(f"{i}. {paper['title']}")
print(f" Authors: {', '.join(paper['authors'][:3])}{'...' if len(paper['authors']) > 3 else ''}")
print(f" Published: {paper['published_date']} | Citations: {paper['citation_count']}")
print(f" Research Area: {paper['research_area']}")
print(f" Open Access: {paper['open_access']}")
print(f" Score: {hit.score:.4f}\n")
```
And there you have it - a sophisticated multi-stage research discovery system in a single declarative query!
## Real ArXiv Dataset Integration
Here's how you could populate the collection with real data (if the endpoint wasn't broken):
```python
# ! pip install arxiv
import arxiv
arxiv_client = arxiv.Client()
search = arxiv.Search(
query="transformer AND multimodal",
max_results=2,
sort_by=arxiv.SortCriterion.SubmittedDate,
)
points = []
for i, paper in enumerate(arxiv_client.results(search)):
print(paper)
# Create dense embedding from abstract
dense_vector = next(dense_model.embed(paper["abstract"]))
# Create sparse vector from technical terms (simplified)
# In practice, you'd use a proper sparse encoder like SPLADE or BM25
sparse_vector = next(sparse_model.embed(paper["abstract"])).as_object()
# Create ColBERT multivector (simplified)
colbert_vector = next(colbert_model.embed(paper["abstract"]))
point = models.PointStruct(
id=i, # Extract arXiv ID
payload={
"title": paper["title"],
"authors": [author for author in paper["authors"]],
"abstract": paper["abstract"],
"published_date": datetime.strptime(
paper["published_date"], "%Y-%m-%d"
).isoformat(),
"citation_count": 0, # Would need external API
"venue": "arXiv",
"research_area": paper["research_area"],
"impact_score": paper["impact_score"],
"open_access": True,
},
vector={
"dense": dense_vector,
"sparse": sparse_vector,
"colbert": colbert_vector,
},
)
points.append(point)
# Upload to Qdrant
client.upsert(collection_name=collection_name, points=points)
print(f"Uploaded {len(points)} research papers to collection")
```
## Key Takeaways
- **Single Request**: Complex multi-stage research discovery in one API call
- **Parallel Execution**: Dense and sparse searches run concurrently
- **Smart Filtering**: Apply research quality filters at optimal stages
- **Real Data**: Works with actual arXiv dataset and research metadata
- **Production Ready**: Scales to millions of papers with sub-second latency
The Universal Query API eliminates the complexity of building multi-turn retrieval systems. What used to require coordination between semantic search engines, keyword systems, and reranking models now happens in a single, optimized request - perfect for academic search, literature reviews, and research recommendation systems.
## Next
In the next lesson, you'll take this foundation and build a complete recommendation service with real data ingestion and user profiling.