mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-02 01:18:30 +02:00
docs: Add course content for Day 5 (#1932)
* docs: Add course content for Day 5 * Apply suggestions from code review colbert-mulivectors changes from kacper Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> * multi vector intro * integrated kacpers ideas * universal query api doc fixes and recall to retrieval in all files * fixed demo * added discord link * updated structure * fix: update colbert-multivectors.md for clarity and formatting improvements * fix: update universal-query-api.md filter explanation * fix: split longer cells into pieces in universal-query-demo.md * fix: detailed query construction steps in universal-query-demo.md * fix: divide pitstop-project.md into smaller steps * fix: a helper function for building filters * fix: refine wording in universal-query-demo.md * fix: move multivector reranking to the day 5 lesson 2 * fix: add Colab link and badge to universal-query-demo.md * fix: remove header * standart init for client * pitstop Reflect on Your Findings * pitstop Reflect on Your Findings * pitstop fix --------- Co-authored-by: Kirstin <kirstin.taufertshoefer@qdrant.com> Co-authored-by: Kacper Łukawski <kacperlukawski@users.noreply.github.com> Co-authored-by: Kacper Łukawski <lukawski.kacper@gmail.com>
This commit is contained in:
co-authored by
Kacper Łukawski
Kirstin
Kacper Łukawski
parent
bcaa46c8ef
commit
a8f39b3f4b
@@ -0,0 +1,23 @@
|
||||
---
|
||||
title: "Day 5: Advanced APIs"
|
||||
isLesson: true
|
||||
weight: 6
|
||||
---
|
||||
|
||||
{{< date >}} Day 5 {{< /date >}}
|
||||
|
||||
# Advanced APIs
|
||||
|
||||
Multi‑vector patterns and the Universal Query API.
|
||||
|
||||
---
|
||||
|
||||
## Today’s path
|
||||
|
||||
1. Multivectors for Late Interaction Models
|
||||
2. The Universal Query API
|
||||
3. Demo: Universal Query for Hybrid Retrieval
|
||||
4. Project: Building a Recommendation System
|
||||
|
||||
You’ll synthesize these APIs into a practical recommendation system.
|
||||
|
||||
@@ -0,0 +1,115 @@
|
||||
---
|
||||
title: Multivectors for Late Interaction Models
|
||||
weight: 1
|
||||
---
|
||||
|
||||
{{< date >}} Day 5 {{< /date >}}
|
||||
|
||||
# Multivectors for Late Interaction Models
|
||||
|
||||
<div class="video">
|
||||
<iframe
|
||||
src="https://www.youtube.com/embed/8ptlXSsSEPk?si=TzsWlastazBQPWWb"
|
||||
frameborder="0"
|
||||
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||
referrerpolicy="strict-origin-when-cross-origin"
|
||||
allowfullscreen>
|
||||
</iframe>
|
||||
</div>
|
||||
|
||||
<br/>
|
||||
|
||||
Many embedding models represent data as a single vector. Transformer-based encoders achieve this by pooling the per-token vector matrix from the final layer into a single vector. That works great for most cases. But when your documents get more complex, cover multiple topics, or require context sensitivity, that one-size-fits-all compression starts to break down. You lose granularity and semantic alignment (though chunking and learned pooling mitigate this to an extent).
|
||||
|
||||
Late-interaction models such as ColBERT retain per-token document vectors. At search time, they identify the best matches by comparing all the query tokens with all the document tokens. This token-level precision preserves local semantic matches and delivers superior relevance, especially for complex queries and documents.
|
||||
|
||||
## Late Interaction: Token-Level Precision
|
||||
|
||||
Qdrant implements this powerful technique through [multivector representations](/documentation/concepts/vectors/#multivectors). A multivector field holds an ordered list of subvectors, each of which captures a different token of the document.
|
||||
|
||||
At query time, Qdrant performs the late interaction scoring. It compares every query token embedding $ q_i $ with each document token embedding $ d_j $. Only the highest score per query token is retained and these top scores are then summed. This mechanism, called MaxSim, delivers fine-grained relevance that respects the structure of your content.
|
||||
|
||||
$$
|
||||
MaxSim_{\text{norm}}(Q, D) = \frac{1}{|Q|} \sum_{i=1}^{|Q|} \max_{j=1}^{|D|} \text{sim}(q_i, d_j)
|
||||
$$
|
||||
|
||||
To enable this in Qdrant, you create a collection with a dense vector that has a multivector comparator provided.
|
||||
|
||||
## Generating Token-Level Embeddings
|
||||
|
||||
You can generate the required token-level embeddings with FastEmbed:
|
||||
|
||||
```python
|
||||
from fastembed import LateInteractionTextEmbedding
|
||||
|
||||
encoder = LateInteractionTextEmbedding("colbert-ir/colbertv2.0")
|
||||
doc_multivectors = list(encoder.embed(["A long document about AI in medicine."]))
|
||||
# Returns [[token_vec1, token_vec2, ...]]
|
||||
```
|
||||
|
||||
The model `colbert-ir/colbertv2.0` outputs 128-dimensional vectors and is available through FastEmbed's optimized ONNX runtime. Use `.embed` for documents and `.query_embed` for queries.
|
||||
|
||||
## Collection Configuration: Multivector Setup
|
||||
|
||||
To use ColBERT for retrieval, create a collection with a multivector field configured for late interaction scoring:
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
|
||||
load_dotenv()
|
||||
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
|
||||
|
||||
# For Colab:
|
||||
# from google.colab import userdata
|
||||
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
|
||||
|
||||
client.create_collection(
|
||||
collection_name="my_colbert_collection",
|
||||
vectors_config={
|
||||
"colbert": models.VectorParams(
|
||||
size=128,
|
||||
distance=models.Distance.COSINE,
|
||||
multivector_config=models.MultiVectorConfig(
|
||||
comparator=models.MultiVectorComparator.MAX_SIM
|
||||
),
|
||||
# Disable HNSW to save RAM - it won't typically be used with multivectors
|
||||
hnsw_config=models.HnswConfigDiff(m=0),
|
||||
)
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
By specifying `MAX_SIM`, you tell Qdrant to apply the late interaction scoring at query time. We explicitly disable HNSW indexing with `m=0` because the graph typically won't be used for multivectors (except in rare edge cases), so disabling it saves RAM. Without HNSW, queries use brute-force MaxSim scoring across all points, which provides maximum precision but may be slower on large collections. For better performance on larger datasets, you'll learn about retrieval-reranking patterns in the next lesson.
|
||||
|
||||
## Querying with ColBERT
|
||||
|
||||
To search using your ColBERT multivector field, embed your query and pass it directly to `query_points`:
|
||||
|
||||
```python
|
||||
from fastembed import LateInteractionTextEmbedding
|
||||
|
||||
# Encode the query
|
||||
colbert = LateInteractionTextEmbedding("colbert-ir/colbertv2.0")
|
||||
colbert_query = next(colbert.query_embed(["what is the policy?"])).tolist()
|
||||
|
||||
# Search using ColBERT multivector
|
||||
hits = client.query_points(
|
||||
collection_name="my_colbert_collection",
|
||||
query=colbert_query,
|
||||
using="colbert",
|
||||
limit=20,
|
||||
)
|
||||
```
|
||||
|
||||
Qdrant performs brute-force MaxSim scoring between your query tokens and the document tokens for all points in the collection. This delivers highly precise results based on fine-grained token-level matching. Keep in mind that without HNSW indexing, this approach may be slower on large collections - in the next lesson you'll see how to combine fast approximate retrieval with ColBERT reranking for better performance.
|
||||
|
||||
## ColPali for Visual Documents
|
||||
|
||||
For documents with rich layouts, PDFs, invoices, slide decks, ColPali (Contextualized Late Interaction over PaliGemma) extends the same idea to vision. ColPali divides each page into a 32×32 grid (1,024 patches), encodes each patch with a vision-language model into 128-dimensional vectors, and treats those patch embeddings as subvectors. You use the identical multivector configuration, and Qdrant applies MaxSim on token embeddings, no matter how they were created.
|
||||
|
||||
The visual approach eliminates traditional OCR and layout detection steps, processing document images directly to capture both textual content and visual structure in a single pass. This makes ColPali particularly effective for complex documents where layout and visual elements are crucial for understanding.
|
||||
|
||||
## Next
|
||||
With multivectors in your toolkit, you can achieve high-precision retrieval for both text and visual documents. In the next lesson, we'll explore the Universal Query API, where you'll learn how to combine multiple retrieval strategies and use ColBERT for reranking - a more common production pattern that balances speed and precision when working with large collections.
|
||||
@@ -0,0 +1,659 @@
|
||||
---
|
||||
title: "Project: Building a Recommendation System"
|
||||
weight: 4
|
||||
---
|
||||
|
||||
{{< date >}} Day 5 {{< /date >}}
|
||||
|
||||
# Project: Building a Recommendation System
|
||||
|
||||
Bring together dense, sparse, and multivectors in one atomic Universal Query. You'll retrieve candidates, fuse signals, rerank with ColBERT, and apply business filters - in a single request.
|
||||
|
||||
## Your Mission
|
||||
|
||||
Build a complete recommendation system using Qdrant’s Universal Query API with dense, sparse, and ColBERT multivectors in one request.
|
||||
|
||||
**Estimated Time:** 90 minutes
|
||||
|
||||
## What You'll Build
|
||||
|
||||
A hybrid recommendation system using:
|
||||
|
||||
- **Multi-vector architecture** with dense, sparse, and ColBERT vectors
|
||||
- **Universal Query API** for atomic multi-stage search
|
||||
- **RRF fusion** for combining candidates
|
||||
- **ColBERT reranking** for fine-grained relevance scoring
|
||||
- **Business rule filtering** at multiple pipeline stages
|
||||
- **Production-ready patterns** for recommendation systems
|
||||
|
||||
## Setup
|
||||
### Prerequisites
|
||||
|
||||
* Qdrant Cloud cluster (URL + API key)
|
||||
* Python 3.9+ (or Google Colab)
|
||||
* Packages: `qdrant-client`, `fastembed`, `python-dotenv`
|
||||
|
||||
### Models
|
||||
- **Dense**: `sentence-transformers/all-MiniLM-L6-v2` (384-dim)
|
||||
- **Sparse**: `prithivida/Splade_PP_en_v1` (SPLADE)
|
||||
- **Multivector**: `colbert-ir/colbertv2.0` (128-dim tokens)
|
||||
|
||||
### Dataset
|
||||
- **Scope**: A small set of sample items (e.g., 10-20 movies).
|
||||
- **Payload Fields**: `title`, `description`, `category`, `genre`, `year`, `rating`, `user_segment`, `popularity_score`, `release_date`.
|
||||
- **Filters Used**: `category`, `user_segment`, `release_date`, `popularity_score`.
|
||||
|
||||
## Build Steps
|
||||
|
||||
### Step 1: Set Up the Hybrid Collection
|
||||
|
||||
#### Initialize Client and Collection
|
||||
|
||||
First, connect to Qdrant and create a clean collection for our recommendation system:
|
||||
|
||||
```python
|
||||
from datetime import datetime
|
||||
from qdrant_client import QdrantClient, models
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
|
||||
load_dotenv()
|
||||
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
|
||||
|
||||
# For Colab:
|
||||
# from google.colab import userdata
|
||||
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
|
||||
|
||||
collection_name = "day5_recommendations_hybrid"
|
||||
|
||||
# Clean state
|
||||
if client.collection_exists(collection_name=collection_name):
|
||||
client.delete_collection(collection_name=collection_name)
|
||||
```
|
||||
|
||||
Now configure a collection with three vectors - each serving a different purpose in our recommendation pipeline:
|
||||
|
||||
```python
|
||||
client.create_collection(
|
||||
collection_name=collection_name,
|
||||
vectors_config={
|
||||
# Dense vectors for semantic understanding
|
||||
"dense": models.VectorParams(size=384, distance=models.Distance.COSINE),
|
||||
# ColBERT multivectors for fine-grained reranking
|
||||
"colbert": models.VectorParams(
|
||||
size=128,
|
||||
distance=models.Distance.COSINE,
|
||||
multivector_config=models.MultiVectorConfig(
|
||||
comparator=models.MultiVectorComparator.MAX_SIM
|
||||
),
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=0 # Disable HNSW - used only for reranking
|
||||
),
|
||||
),
|
||||
},
|
||||
sparse_vectors_config={
|
||||
# Sparse vectors for exact keyword matching
|
||||
"sparse": models.SparseVectorParams(
|
||||
index=models.SparseIndexParams(on_disk=False)
|
||||
)
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
**Why this setup**: ColBERT uses `MAX_SIM` for token-level comparison and `m=0` since it's only used for reranking, not initial retrieval, and setting `m=0` effectively disables HNSW indexing to save memory and compute.
|
||||
|
||||
#### Create Payload Indexes
|
||||
|
||||
Before ingesting data, create indexes for the fields we'll filter by. This enables efficient filtering during vector search:
|
||||
|
||||
```python
|
||||
# Business metadata indexes
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="category",
|
||||
field_schema="keyword",
|
||||
)
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="user_segment",
|
||||
field_schema="keyword",
|
||||
)
|
||||
|
||||
# Quality and recency indexes
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="release_date",
|
||||
field_schema="datetime",
|
||||
)
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="popularity_score",
|
||||
field_schema="float",
|
||||
)
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="rating",
|
||||
field_schema="float",
|
||||
)
|
||||
```
|
||||
|
||||
### Step 2: Prepare and Upload Recommendation Data
|
||||
|
||||
#### Create Sample Recommendation Data
|
||||
|
||||
Let's create sample movie data with business metadata for filtering:
|
||||
|
||||
```python
|
||||
# Example: Create sample movie/content data
|
||||
sample_data = [
|
||||
{
|
||||
"title": "The Matrix",
|
||||
"description": "A computer hacker learns about the true nature of reality and his role in the war against its controllers.",
|
||||
"category": "movie",
|
||||
"genre": ["sci-fi", "action"],
|
||||
"year": 1999,
|
||||
"rating": 8.7,
|
||||
"user_segment": "premium",
|
||||
"popularity_score": 0.9,
|
||||
"release_date": "1999-03-31T00:00:00Z",
|
||||
},
|
||||
# Add more sample data...
|
||||
]
|
||||
|
||||
texts = [it["description"] for it in sample_data]
|
||||
```
|
||||
|
||||
#### Initialize Embedding Models
|
||||
|
||||
We'll use FastEmbed to generate all three vector types:
|
||||
|
||||
```python
|
||||
from fastembed import TextEmbedding, SparseTextEmbedding, LateInteractionTextEmbedding
|
||||
|
||||
# Model configurations
|
||||
DENSE_MODEL_ID = "sentence-transformers/all-MiniLM-L6-v2" # 384-dim
|
||||
SPARSE_MODEL_ID = "prithivida/Splade_PP_en_v1" # SPLADE sparse
|
||||
COLBERT_MODEL_ID = "colbert-ir/colbertv2.0" # 128-dim multivector
|
||||
|
||||
dense_model = TextEmbedding(DENSE_MODEL_ID)
|
||||
sparse_model = SparseTextEmbedding(SPARSE_MODEL_ID)
|
||||
colbert_model = LateInteractionTextEmbedding(COLBERT_MODEL_ID)
|
||||
```
|
||||
|
||||
#### Generate Embeddings
|
||||
|
||||
Encode all descriptions with all three embedding models:
|
||||
|
||||
```python
|
||||
# Generate embeddings for all items
|
||||
dense_embeds = list(
|
||||
dense_model.embed(texts, parallel=0)
|
||||
) # list[np.ndarray] shape (384,)
|
||||
|
||||
sparse_embeds = list(
|
||||
sparse_model.embed(texts, parallel=0)
|
||||
) # list[SparseEmbedding] with .indices/.values
|
||||
|
||||
colbert_multivectors = list(
|
||||
colbert_model.embed(texts, parallel=0)
|
||||
) # list[np.ndarray] shape (tokens, 128)
|
||||
```
|
||||
|
||||
#### Create Points and Upload
|
||||
|
||||
Now package everything into points and upload to Qdrant:
|
||||
|
||||
```python
|
||||
# Generate vectors for each item
|
||||
points = []
|
||||
for i, item in enumerate(sample_data):
|
||||
# Create sparse vector (keyword matching)
|
||||
sparse_vector = sparse_embeds[i].as_object()
|
||||
|
||||
# Create dense vector (semantic understanding)
|
||||
dense_vector = dense_embeds[i]
|
||||
|
||||
# Create ColBERT multivector (token-level understanding)
|
||||
colbert_vector = colbert_multivectors[i]
|
||||
|
||||
points.append(
|
||||
models.PointStruct(
|
||||
id=i,
|
||||
vector={
|
||||
"dense": dense_vector,
|
||||
"sparse": sparse_vector,
|
||||
"colbert": colbert_vector,
|
||||
},
|
||||
payload=item,
|
||||
)
|
||||
)
|
||||
|
||||
client.upload_points(collection_name=collection_name, points=points)
|
||||
print(f"Uploaded {len(points)} recommendation items")
|
||||
```
|
||||
|
||||
### Step 3: One Universal Query — Retrieve → Fuse → Rerank → Filter
|
||||
|
||||
Now let's build a complete multi-stage search that combines everything in a single request.
|
||||
|
||||
#### Prepare Query Embeddings
|
||||
|
||||
First, encode the user's search intent using all three embedding models:
|
||||
|
||||
```python
|
||||
from datetime import timedelta
|
||||
|
||||
# Example user intent
|
||||
user_query = "premium user likes sci-fi action movies with strong hacker themes"
|
||||
|
||||
user_dense_vector = next(dense_model.query_embed(user_query))
|
||||
user_sparse_vector = next(sparse_model.query_embed(user_query)).as_object()
|
||||
user_multivector = next(colbert_model.query_embed(user_query))
|
||||
```
|
||||
|
||||
#### Define Global Filter with Automatic Propagation
|
||||
|
||||
Create a single filter that will automatically propagate to all stages of the pipeline:
|
||||
|
||||
```python
|
||||
# Global filter - this will be propagated to ALL prefetch stages
|
||||
global_filter = models.Filter(
|
||||
must=[
|
||||
# Content type and user segment
|
||||
models.FieldCondition(
|
||||
key="category", match=models.MatchValue(value="movie")
|
||||
),
|
||||
models.FieldCondition(
|
||||
key="user_segment", match=models.MatchValue(value="premium")
|
||||
),
|
||||
# Quality and recency constraints
|
||||
models.FieldCondition(
|
||||
key="release_date",
|
||||
range=models.DatetimeRange(
|
||||
gte=(datetime.now() - timedelta(days=365 * 30)).isoformat()
|
||||
),
|
||||
),
|
||||
models.FieldCondition(key="popularity_score", range=models.Range(gte=0.7)),
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
**Key insight**: This filter automatically propagates to ALL prefetch stages. Qdrant doesn't do "late filtering" or "post-filtering" - filters are applied at the HNSW search level for maximum efficiency.
|
||||
|
||||
#### Set Up Prefetch and Fusion
|
||||
|
||||
Configure parallel retrieval and fusion strategy:
|
||||
|
||||
```python
|
||||
# Prefetch queries - global filter will be automatically applied to both
|
||||
hybrid_query = [
|
||||
models.Prefetch(query=user_dense_vector, using="dense", limit=100),
|
||||
models.Prefetch(query=user_sparse_vector, using="sparse", limit=100),
|
||||
]
|
||||
|
||||
# Fusion stage - combine candidates with RRF
|
||||
fusion_query = models.Prefetch(
|
||||
prefetch=hybrid_query,
|
||||
query=models.FusionQuery(fusion=models.Fusion.RRF),
|
||||
limit=100,
|
||||
)
|
||||
```
|
||||
|
||||
These prefetch queries run in parallel, and the global filter from the main query will automatically propagate to both dense and sparse searches.
|
||||
|
||||
#### Execute Universal Query with Reranking
|
||||
|
||||
Finally, send the complete query with ColBERT reranking:
|
||||
|
||||
```python
|
||||
# The Universal Query: Global filter propagates through all stages
|
||||
response = client.query_points(
|
||||
collection_name=collection_name,
|
||||
prefetch=fusion_query,
|
||||
query=user_multivector,
|
||||
using="colbert",
|
||||
query_filter=global_filter, # Propagates to all prefetch stages
|
||||
limit=10,
|
||||
with_payload=True,
|
||||
)
|
||||
|
||||
for hit in response.points or []:
|
||||
print(hit.payload)
|
||||
```
|
||||
|
||||
**Why this works**: Dense and sparse retrieval happen in parallel (with filters applied), RRF fuses the results, ColBERT reranks with token-level precision, and the global filter is applied at every stage via automatic propagation - all in one atomic API call.
|
||||
|
||||
|
||||
### Step 4: Build a Recommendation Service
|
||||
|
||||
Now let's wrap everything into a production-ready function that handles dynamic user preferences.
|
||||
|
||||
#### Helper Function for Building Filters
|
||||
|
||||
First, create a reusable helper function to build filters from user profiles:
|
||||
|
||||
```python
|
||||
def build_recommendation_filter(user_profile, user_preference=None):
|
||||
"""
|
||||
Build a global filter from user profile and preferences.
|
||||
This filter will automatically propagate to all prefetch stages.
|
||||
|
||||
Args:
|
||||
user_profile: {
|
||||
"liked_titles": list[str], # optional
|
||||
"preferred_genres": list[str], # e.g. ["sci-fi","action"]
|
||||
"segment": str, # e.g. "premium"
|
||||
"query": str # free-text intent, e.g. "smart sci-fi with hacker vibe"
|
||||
}
|
||||
user_preference: {
|
||||
"category": str | None, # e.g. "movie"
|
||||
"min_rating": float | None, # e.g. 8.0
|
||||
"released_within_days": int | None # e.g. 365
|
||||
}
|
||||
|
||||
Returns:
|
||||
models.Filter object or None if no conditions
|
||||
"""
|
||||
from datetime import datetime, timedelta
|
||||
|
||||
filter_conditions = []
|
||||
|
||||
# User segment filtering
|
||||
if user_profile.get("segment"):
|
||||
filter_conditions.append(
|
||||
models.FieldCondition(
|
||||
key="user_segment",
|
||||
match=models.MatchValue(value=user_profile["segment"]),
|
||||
)
|
||||
)
|
||||
|
||||
if user_preference:
|
||||
# Category filtering
|
||||
if user_preference.get("category"):
|
||||
filter_conditions.append(
|
||||
models.FieldCondition(
|
||||
key="category",
|
||||
match=models.MatchValue(value=user_preference["category"]),
|
||||
)
|
||||
)
|
||||
|
||||
# Rating filtering
|
||||
if user_preference.get("min_rating") is not None:
|
||||
filter_conditions.append(
|
||||
models.FieldCondition(
|
||||
key="rating",
|
||||
range=models.Range(gte=user_preference["min_rating"])
|
||||
)
|
||||
)
|
||||
|
||||
# Recency filtering
|
||||
if user_preference.get("released_within_days"):
|
||||
days = int(user_preference["released_within_days"])
|
||||
filter_conditions.append(
|
||||
models.FieldCondition(
|
||||
key="release_date",
|
||||
range=models.DatetimeRange(
|
||||
gte=(datetime.utcnow() - timedelta(days=days)).isoformat()
|
||||
),
|
||||
)
|
||||
)
|
||||
|
||||
return models.Filter(must=filter_conditions) if filter_conditions else None
|
||||
```
|
||||
|
||||
#### Recommendation Function
|
||||
|
||||
Now create the main recommendation function using the helper:
|
||||
|
||||
```python
|
||||
def get_recommendations(user_profile, user_preference=None, limit=10):
|
||||
"""
|
||||
Get personalized recommendations using Universal Query API
|
||||
|
||||
Args:
|
||||
user_profile: {
|
||||
"liked_titles": list[str], # optional
|
||||
"preferred_genres": list[str], # e.g. ["sci-fi","action"]
|
||||
"segment": str, # e.g. "premium"
|
||||
"query": str # free-text intent, e.g. "smart sci-fi with hacker vibe"
|
||||
}
|
||||
user_preference: {
|
||||
"category": str | None, # e.g. "movie"
|
||||
"min_rating": float | None, # e.g. 8.0
|
||||
"released_within_days": int | None # e.g. 365
|
||||
}
|
||||
limit: top-k to return
|
||||
"""
|
||||
|
||||
# Generate query embeddings
|
||||
user_dense_vector = next(dense_model.query_embed(user_profile["query"]))
|
||||
user_sparse_vector = next(
|
||||
sparse_model.query_embed(user_profile["query"])
|
||||
).as_object()
|
||||
user_multivector = next(colbert_model.query_embed(user_profile["query"]))
|
||||
|
||||
# Build global filter using helper function
|
||||
global_filter = build_recommendation_filter(user_profile, user_preference)
|
||||
|
||||
# Prefetch queries - global filter will propagate automatically
|
||||
hybrid_query = [
|
||||
models.Prefetch(query=user_dense_vector, using="dense", limit=100),
|
||||
models.Prefetch(query=user_sparse_vector, using="sparse", limit=100),
|
||||
]
|
||||
|
||||
# Combine candidates with RRF
|
||||
fusion_query = models.Prefetch(
|
||||
prefetch=hybrid_query,
|
||||
query=models.FusionQuery(fusion=models.Fusion.RRF),
|
||||
limit=100,
|
||||
)
|
||||
|
||||
# Universal query - global filter propagates to all stages
|
||||
response = client.query_points(
|
||||
collection_name=collection_name,
|
||||
prefetch=fusion_query,
|
||||
query=user_multivector,
|
||||
using="colbert",
|
||||
query_filter=global_filter, # Propagates to all prefetch stages
|
||||
limit=limit,
|
||||
with_payload=True,
|
||||
)
|
||||
|
||||
return [
|
||||
{
|
||||
"title": hit.payload["title"],
|
||||
"description": hit.payload["description"],
|
||||
"score": hit.score,
|
||||
"metadata": {
|
||||
k: v
|
||||
for k, v in hit.payload.items()
|
||||
if k not in ["title", "description"]
|
||||
},
|
||||
}
|
||||
for hit in (response.points or [])
|
||||
]
|
||||
```
|
||||
|
||||
#### Test the Service
|
||||
|
||||
Let's test the recommendation function:
|
||||
|
||||
```python
|
||||
# Test the recommendation service
|
||||
user_profile = {
|
||||
"liked_titles": ["The Matrix", "Blade Runner"],
|
||||
"preferred_genres": ["sci-fi", "action"],
|
||||
"segment": "premium",
|
||||
"query": "highly rated cyberpunk movies",
|
||||
}
|
||||
|
||||
recommendations = get_recommendations(
|
||||
user_profile,
|
||||
user_preference={
|
||||
"category": "movie",
|
||||
"min_rating": 8.0,
|
||||
"released_within_days": 365 * 30,
|
||||
},
|
||||
limit=10,
|
||||
)
|
||||
|
||||
for i, rec in enumerate(recommendations, 1):
|
||||
print(f"{i}. {rec['title']} (Score: {rec['score']:.3f})")
|
||||
```
|
||||
|
||||
**What happened under the hood**: Qdrant retrieved 100 candidates from dense and 100 from sparse in parallel, fused them with RRF, reranked with ColBERT's MaxSim over token‑level subvectors, applied final business filters, and returned the top 10 - all in one call.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
You'll know you've succeeded when:
|
||||
|
||||
<input type="checkbox"> Your collection contains dense, sparse, and ColBERT vectors
|
||||
<input type="checkbox"> You can execute complex multi-stage searches in a single API call
|
||||
<input type="checkbox"> RRF fusion effectively combines different vector types
|
||||
<input type="checkbox"> ColBERT reranking improves result relevance
|
||||
<input type="checkbox"> Business filters propagate automatically to all prefetch stages
|
||||
<input type="checkbox"> Your recommendation service provides personalized, high-quality results
|
||||
|
||||
## Share Your Discovery
|
||||
|
||||
### Step 1: Reflect on Your Findings
|
||||
|
||||
1. How did the Universal Query API simplify your recommendation pipeline (fewer calls, fewer joins, less code)?
|
||||
2. Which fusion strategy (RRF vs. DBSF) ranked better for your queries?
|
||||
3. What was the effect of ColBERT reranking on recommendation quality (top-k changes, click-like signals, nDCG/MRR)?
|
||||
4. What was the latency impact of multi-stage filtering (prefetch, rerank, total)?
|
||||
|
||||
### Step 2: Post Your Results
|
||||
|
||||
**Post your results in** <a href="https://discord.com/invite/qdrant" target="_blank" rel="noopener noreferrer" aria-label="Qdrant Discord"> <img src="https://img.shields.io/badge/Qdrant%20Discord-5865F2?style=flat&logo=discord&logoColor=white&labelColor=5865F2&color=5865F2"
|
||||
alt="Post your results in Discord"
|
||||
style="display:inline; margin:0; vertical-align:middle; border-radius:9999px;" /> </a> **using this:**
|
||||
|
||||
```markdown
|
||||
**[Day N] Recommendations with the Universal Query API**
|
||||
|
||||
**High-Level Summary**
|
||||
- **Domain:** "I built recommendations for [domain]"
|
||||
- **Key Result:** "Using [RRF/DBSF] + ColBERT rerank improved [metric] to [value] with total latency [Z] ms."
|
||||
|
||||
**Reproducibility**
|
||||
- **Collection:** [name]
|
||||
- **Models:** dense=[id, dim], sparse=[method], colbert=[id, dim]
|
||||
- **Dataset:** [N items] (snapshot: YYYY-MM-DD)
|
||||
|
||||
**Settings (today)**
|
||||
- **Fusion:** [RRF/DBSF], k_dense=[..], k_sparse=[..]
|
||||
- **Reranker:** ColBERT (MaxSim), top-k=[..]
|
||||
- **Filters:** category=[...], segment=[...], rating≥..., release_date≥...
|
||||
- **Filter propagation:** applied to all prefetch stages
|
||||
|
||||
**Results (demo query: "[user intent]")**
|
||||
- **Top picks (rank → title → score):**
|
||||
1) ...
|
||||
2) ...
|
||||
3) ...
|
||||
- **Why these won:** [token hit like “hacker”, genre match, rating signal]
|
||||
- **Latency:** prefetch ~[X] ms | rerank ~[Y] ms | total ~[Z] ms
|
||||
- **Dropped by rules:** [ids/titles → which rule]
|
||||
|
||||
**Surprise**
|
||||
- "[one thing you didn’t expect]"
|
||||
|
||||
**Next step**
|
||||
- "[what you’ll try next]"
|
||||
```
|
||||
|
||||
## Optional: Go Further
|
||||
|
||||
### Experiment with Fusion Strategies
|
||||
|
||||
Compare RRF with Distribution-Based Score Fusion:
|
||||
|
||||
```python
|
||||
# Test DBSF vs RRF
|
||||
# Filters and other prefetch params same as above
|
||||
|
||||
fusion_query = models.Prefetch(
|
||||
prefetch=hybrid_query,
|
||||
query=models.FusionQuery(fusion=models.Fusion.DBSF),
|
||||
limit=100,
|
||||
)
|
||||
|
||||
response = client.query_points(
|
||||
collection_name=collection_name,
|
||||
prefetch=fusion_query,
|
||||
query=user_multivector,
|
||||
using="colbert",
|
||||
query_filter=global_filter, # Same global filter propagates to all stages
|
||||
limit=10,
|
||||
with_payload=True,
|
||||
)
|
||||
|
||||
for hit in response.points or []:
|
||||
print(hit.payload)
|
||||
```
|
||||
|
||||
### A/B Testing Framework
|
||||
|
||||
Build a framework to systematically compare fusion strategies across multiple user profiles.
|
||||
|
||||
#### Initialize Testing Function
|
||||
|
||||
Set up the function structure and results tracking:
|
||||
|
||||
```python
|
||||
def ab_test_fusion_strategies(user_profiles, user_preferences):
|
||||
"""Compare RRF vs DBSF performance"""
|
||||
|
||||
results = {"RRF": [], "DBSF": []}
|
||||
|
||||
for user_profile, user_preference in zip(user_profiles, user_preferences):
|
||||
for fusion_type in [models.Fusion.RRF, models.Fusion.DBSF]:
|
||||
# Run recommendation query
|
||||
user_dense_vector = next(dense_model.query_embed(user_profile["query"]))
|
||||
user_sparse_vector = next(
|
||||
sparse_model.query_embed(user_profile["query"])
|
||||
).as_object()
|
||||
user_multivector = next(colbert_model.query_embed(user_profile["query"]))
|
||||
|
||||
# Build global filter using helper function
|
||||
global_filter = build_recommendation_filter(user_profile, user_preference)
|
||||
|
||||
# Prefetch queries - global filter will propagate automatically
|
||||
hybrid_query = [
|
||||
models.Prefetch(query=user_dense_vector, using="dense", limit=100),
|
||||
models.Prefetch(query=user_sparse_vector, using="sparse", limit=100),
|
||||
]
|
||||
|
||||
# Combine candidates with chosen fusion strategy
|
||||
fusion_query = models.Prefetch(
|
||||
prefetch=hybrid_query,
|
||||
query=models.FusionQuery(fusion=fusion_type),
|
||||
limit=100,
|
||||
)
|
||||
|
||||
response = client.query_points(
|
||||
collection_name=collection_name,
|
||||
prefetch=fusion_query,
|
||||
query=user_multivector,
|
||||
using="colbert",
|
||||
query_filter=global_filter, # Propagates to all prefetch stages
|
||||
limit=10,
|
||||
with_payload=True,
|
||||
)
|
||||
|
||||
strategy_name = "RRF" if fusion_type == models.Fusion.RRF else "DBSF"
|
||||
results[strategy_name].append(
|
||||
{
|
||||
"user_id": user_profile["user_id"],
|
||||
"recommendations": [hit.id for hit in response.points],
|
||||
"scores": [hit.score for hit in response.points],
|
||||
}
|
||||
)
|
||||
|
||||
return results
|
||||
```
|
||||
|
||||
## Next Steps
|
||||
|
||||
Turn this into a mini‑service: ingest your own items and user signals, write a tiny function that takes profile vectors and returns top‑k via the Universal Query API, then experiment with filters and fusion strategies (try DBSF vs RRF) to see how ranking shifts.
|
||||
@@ -0,0 +1,172 @@
|
||||
---
|
||||
title: The Universal Query API
|
||||
weight: 2
|
||||
---
|
||||
|
||||
{{< date >}} Day 5 {{< /date >}}
|
||||
|
||||
# The Universal Query API
|
||||
|
||||
Picture this: a customer types "leather jackets" into your store's search bar. You want to show items that match the style semantically - so a bomber jacket surfaces even if it doesn't mention "leather jackets" verbatim - but you also need to enforce your business rules. Only products under $200, only items in stock, only jackets released within the past year. Traditionally, you'd fire off a search, gather results, then apply filters and glue code. With Qdrant's [Universal Query API](/documentation/concepts/hybrid-queries/), all of that happens in one declarative request.
|
||||
|
||||
## Run dense + sparse retrieval in parallel with RRF
|
||||
|
||||
First, you retrieve candidates from multiple sources in parallel and fuse their ranks. Below, we blend dense semantics from a BGE model with sparse keyword matching from SPLADE by using [Reciprocal Rank Fusion](/documentation/concepts/hybrid-queries/#hybrid-search) to merge the two lists:
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
|
||||
load_dotenv()
|
||||
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
|
||||
|
||||
# For Colab:
|
||||
# from google.colab import userdata
|
||||
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
|
||||
|
||||
response = client.query_points(
|
||||
collection_name="products",
|
||||
prefetch=[
|
||||
models.Prefetch(
|
||||
query=dense_vector,
|
||||
using="dense_bge",
|
||||
limit=20
|
||||
),
|
||||
models.Prefetch(
|
||||
query=sparse_vector,
|
||||
using="sparse_splade",
|
||||
limit=20
|
||||
)
|
||||
],
|
||||
query=models.FusionQuery(fusion=models.Fusion.RRF),
|
||||
limit=10
|
||||
)
|
||||
```
|
||||
|
||||
Qdrant sends both prefetches concurrently, fuses the two ranked lists by reciprocal rank, and returns your top ten products that satisfy both semantic and keyword relevance.
|
||||
|
||||
## Dense Retrieval + ColBERT Rerank
|
||||
|
||||
While the previous lesson showed how to use ColBERT directly for retrieval, **in production systems reranking is the more common pattern**. The reason is practical: ColBERT's brute-force MaxSim scoring on an entire large collection can be slow. Instead, you combine two strengths - fast approximate search to narrow candidates, then precise token-level scoring on that smaller set.
|
||||
|
||||
You create a collection with two vector fields: a dense vector with HNSW indexing for speed, and a ColBERT multivector with HNSW disabled for precision:
|
||||
|
||||
```python
|
||||
client.create_collection(
|
||||
collection_name="articles",
|
||||
vectors_config={
|
||||
# Fast HNSW-indexed dense retrieval
|
||||
"bge-dense": models.VectorParams(
|
||||
size=384,
|
||||
distance=models.Distance.COSINE,
|
||||
),
|
||||
# Precise multivector reranking (HNSW disabled to save RAM)
|
||||
"colbert": models.VectorParams(
|
||||
size=128,
|
||||
distance=models.Distance.COSINE,
|
||||
multivector_config=models.MultiVectorConfig(
|
||||
comparator=models.MultiVectorComparator.MAX_SIM
|
||||
),
|
||||
hnsw_config=models.HnswConfigDiff(m=0),
|
||||
)
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
Now you retrieve 100 candidates quickly with your HNSW-indexed dense field, then apply ColBERT's MaxSim scoring to rerank those hundred and select the very best ten:
|
||||
|
||||
```python
|
||||
from fastembed import LateInteractionTextEmbedding, TextEmbedding
|
||||
|
||||
# Encode with both models
|
||||
dense = TextEmbedding("BAAI/bge-small-en-v1.5")
|
||||
dense_query_vector = next(dense.query_embed(["what is the policy?"])).tolist()
|
||||
|
||||
colbert = LateInteractionTextEmbedding("colbert-ir/colbertv2.0")
|
||||
colbert_query_multivector = next(colbert.query_embed(["what is the policy?"])).tolist()
|
||||
|
||||
# Fast retrieval + precise reranking in one call
|
||||
response = client.query_points(
|
||||
collection_name="articles",
|
||||
prefetch=[
|
||||
models.Prefetch(
|
||||
query=dense_query_vector,
|
||||
using="bge-dense",
|
||||
limit=100
|
||||
)
|
||||
],
|
||||
query=colbert_query_multivector,
|
||||
using="colbert",
|
||||
limit=10
|
||||
)
|
||||
```
|
||||
|
||||
Behind the scenes, Qdrant fetches the 100 nearest points from the HNSW-indexed dense field, then applies the MaxSim late-interaction score from your ColBERT multivector to only those 100 candidates, returning the ten highest-scoring results. This two-stage approach delivers ColBERT's precision while keeping query latency practical for large-scale deployments.
|
||||
|
||||
## Global and Prefetch-Specific Filters
|
||||
|
||||
Finally, you layer in filtering wherever it makes sense. In the snippet below, you specify global filters at the query level - price under $200 and release date after January 1, 2023 - which automatically propagate to all prefetches. Then, you add an additional prefetch-specific filter on the dense search to only retrieve products that are in stock and in the "jackets" category:
|
||||
|
||||
```python
|
||||
response = client.query_points(
|
||||
collection_name="products",
|
||||
prefetch=[
|
||||
models.Prefetch(
|
||||
query=dense_query_vector,
|
||||
using="bge-dense",
|
||||
limit=100,
|
||||
filter=models.Filter(
|
||||
must=[
|
||||
models.FieldCondition(
|
||||
key="in_stock",
|
||||
match=models.MatchValue(value=True)
|
||||
),
|
||||
models.FieldCondition(
|
||||
key="category",
|
||||
match=models.MatchValue(value="jackets")
|
||||
)
|
||||
]
|
||||
)
|
||||
),
|
||||
models.Prefetch(
|
||||
query=sparse_query_vector,
|
||||
using="sparse-splade",
|
||||
limit=100
|
||||
)
|
||||
],
|
||||
query=models.FusionQuery(fusion=models.Fusion.RRF),
|
||||
filter=models.Filter(
|
||||
must=[
|
||||
models.FieldCondition(
|
||||
key="price",
|
||||
range=models.Range(lt=200.0)
|
||||
),
|
||||
models.FieldCondition(
|
||||
key="release_date",
|
||||
range=models.DatetimeRange(
|
||||
gte="2023-01-01T00:00:00Z"
|
||||
)
|
||||
)
|
||||
]
|
||||
),
|
||||
limit=10
|
||||
)
|
||||
```
|
||||
|
||||
The global filters (price and release_date) apply to both prefetches automatically. The first prefetch adds extra constraints (in_stock and category), while the second prefetch only uses the global filters. This eliminates the need to repeat common filters across every prefetch. All filtering happens efficiently during the retrieval phase - there's no separate post-processing step. The entire pipeline - hybrid retrieval, filtering, fusion, and reranking - executes in one API call.
|
||||
|
||||
## The Universal Query API Structure
|
||||
|
||||
The Universal Query API enables complex search patterns through a simple declarative interface:
|
||||
|
||||
- **Prefetch Stage**: Execute multiple searches in parallel against different vector fields. Each prefetch can have its own filters, limits, and vector types used.
|
||||
- **Fusion Stage**: Combine results from multiple prefetches using algorithms like Reciprocal Rank Fusion (RRF) or Distribution-Based Score Fusion (DBSF).
|
||||
- **Reranking Stage**: Or rerank candidates from a single prefetch with a stronger scorer such as ColBERT. Fusion and reranking are alternative final steps in most pipelines.
|
||||
- **Filtering**: Apply filters globally at the query level (propagated to all prefetches) or add prefetch-specific filters for additional constraints on individual searches.
|
||||
|
||||
This architecture eliminates the need for multiple API calls, client-side result merging, and complex orchestration code. Everything happens server-side in a single, optimized request.
|
||||
|
||||
## Next
|
||||
|
||||
In our next lesson, you'll bring these building blocks together in a personalized recommendation pipeline. You'll learn how to retrieve candidate from three different vector representations, fuse their signals, rerank with ColBERT, apply user-segment filters, and deliver highly relevant suggestions without writing a single line of glue code.
|
||||
@@ -0,0 +1,402 @@
|
||||
---
|
||||
title: "Demo: Universal Query for Hybrid Retrieval"
|
||||
weight: 3
|
||||
---
|
||||
|
||||
{{< date >}} Day 5 {{< /date >}}
|
||||
|
||||
# Demo: Universal Query for Hybrid Retrieval
|
||||
|
||||
In this hands-on demo, we'll build a research paper discovery system using the arXiv dataset that showcases the full power of Qdrant's Universal Query API. You'll see how to combine dense semantics, sparse keywords, and ColBERT reranking to help researchers find exactly the papers they need - all in a single query.
|
||||
|
||||
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course/day_5/universal-query-demo.ipynb">
|
||||
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||
</a>
|
||||
|
||||
## The Challenge: Intelligent Research Discovery
|
||||
|
||||
Imagine you're a machine learning researcher looking for "transformer architectures for multimodal learning with attention mechanisms." You need to:
|
||||
|
||||
1. **Retrieve broadly** using semantic understanding of research concepts (dense vectors)
|
||||
2. **Match precisely** on technical terms like "transformer" and "attention" (sparse vectors)
|
||||
3. **Rerank intelligently** using fine-grained text understanding (ColBERT)
|
||||
4. **Apply research filters** like publication date, citation count, and research domain
|
||||
|
||||
Traditionally, this would require multiple searches across different systems, manual result merging, and complex ranking logic. With the Universal Query API, it's one declarative request.
|
||||
|
||||
## Step 1: Create the Research Paper Collection
|
||||
|
||||
### Initialize the Collection with Vector Configurations
|
||||
|
||||
First, let's set up a collection with three vector types - each serving a different purpose in our research discovery pipeline:
|
||||
|
||||
```python
|
||||
from datetime import datetime, timedelta
|
||||
|
||||
from qdrant_client import QdrantClient, models
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
|
||||
load_dotenv()
|
||||
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
|
||||
|
||||
# For Colab:
|
||||
# from google.colab import userdata
|
||||
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
|
||||
|
||||
collection_name = "research-papers"
|
||||
|
||||
# Clean state
|
||||
if client.collection_exists(collection_name=collection_name):
|
||||
client.delete_collection(collection_name=collection_name)
|
||||
|
||||
# Create collection with three vector types
|
||||
client.create_collection(
|
||||
collection_name=collection_name,
|
||||
vectors_config={
|
||||
# Dense vectors for semantic understanding of research concepts
|
||||
"dense": models.VectorParams(size=384, distance=models.Distance.COSINE),
|
||||
# ColBERT multivectors for fine-grained text understanding
|
||||
"colbert": models.VectorParams(
|
||||
size=128,
|
||||
distance=models.Distance.COSINE,
|
||||
multivector_config=models.MultiVectorConfig(
|
||||
comparator=models.MultiVectorComparator.MAX_SIM
|
||||
),
|
||||
),
|
||||
},
|
||||
sparse_vectors_config={
|
||||
# Sparse vectors for exact technical term matching
|
||||
"sparse": models.SparseVectorParams(
|
||||
index=models.SparseIndexParams(on_disk=False)
|
||||
)
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
### Create Payload Indexes
|
||||
|
||||
Before ingesting any data, we create payload indexes for the fields we'll filter by. Qdrant's version of HNSW incorporates payload filtering directly into the search process for efficiency.
|
||||
|
||||
```python
|
||||
# Index fields that will be used for filtering
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="research_area",
|
||||
field_schema="keyword", # For filtering by domain (ML, CV, NLP)
|
||||
)
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="open_access",
|
||||
field_schema="bool", # For filtering open access papers
|
||||
)
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="published_date",
|
||||
field_schema="datetime",
|
||||
)
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="impact_score",
|
||||
field_schema="float",
|
||||
)
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="citation_count",
|
||||
field_schema="integer",
|
||||
)
|
||||
```
|
||||
|
||||
## Prepare and Ingest Research Paper Data
|
||||
|
||||
Now that our collection is configured with vectors and payload indexes, let's take some sample research papers:
|
||||
|
||||
```python
|
||||
sample_data = [
|
||||
{
|
||||
"title": "Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace",
|
||||
"authors": ["Andre Rusli", "Shoma Ishimoto", "Sho Akiyama", "Aman Kumar Singh"],
|
||||
"abstract": "Visual search offers an intuitive way for customers to explore diverse product catalogs, particularly in consumer-to-consumer (C2C) marketplaces where listings are often unstructured and visually driven. This paper presents a scalable visual search system deployed in Mercari's C2C marketplace...",
|
||||
"research_area": "computer_vision",
|
||||
"published_date": "2025-07-31",
|
||||
"impact_score": 0.78,
|
||||
"citation_count": 12,
|
||||
"open_access": True,
|
||||
},
|
||||
{
|
||||
"title": "TALI: Towards A Lightweight Information Retrieval Framework for Neural Search",
|
||||
"authors": ["Chaoqun Liu", "Yuanming Zhang", "Jianmin Zhang", "Jiawei Han"],
|
||||
"abstract": "Neural search systems have emerged as a promising approach to enhance user engagement in information retrieval. However, their high computational costs and memory usage have limited their widespread adoption. In this paper, we present TALI, a lightweight information retrieval framework for neural search that efficiently addresses these challenges...",
|
||||
"research_area": "machine_learning",
|
||||
"published_date": "2025-07-31",
|
||||
"impact_score": 0.78,
|
||||
"citation_count": 12,
|
||||
"open_access": True,
|
||||
},
|
||||
{
|
||||
"title": "Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace",
|
||||
"authors": ["Andre Rusli", "Shoma Ishimoto", "Sho Akiyama", "Aman Kumar Singh"],
|
||||
"abstract": "Visual search offers an intuitive way for customers to explore diverse product catalogs, particularly in consumer-to-consumer (C2C) marketplaces where listings are often unstructured and visually driven. This paper presents a scalable visual search system deployed in Mercari's C2C marketplace...",
|
||||
"research_area": "computer_vision",
|
||||
"published_date": "2025-07-31",
|
||||
"impact_score": 0.78,
|
||||
"citation_count": 12,
|
||||
"open_access": True,
|
||||
},
|
||||
]
|
||||
|
||||
texts = [it["abstract"] for it in sample_data]
|
||||
```
|
||||
|
||||
We'll use FastEmbed to generate dense, sparse, and ColBERT embeddings for the abstracts. Then we upload everything to Qdrant:
|
||||
|
||||
```python
|
||||
from fastembed import TextEmbedding, SparseTextEmbedding, LateInteractionTextEmbedding
|
||||
|
||||
DENSE_MODEL_ID = "sentence-transformers/all-MiniLM-L6-v2" # 384-dim
|
||||
SPARSE_MODEL_ID = "prithivida/Splade_PP_en_v1" # SPLADE sparse
|
||||
COLBERT_MODEL_ID = "colbert-ir/colbertv2.0" # 128-dim multivector
|
||||
|
||||
dense_model = TextEmbedding(DENSE_MODEL_ID)
|
||||
sparse_model = SparseTextEmbedding(SPARSE_MODEL_ID)
|
||||
colbert_model = LateInteractionTextEmbedding(COLBERT_MODEL_ID)
|
||||
|
||||
dense_embeds = list(dense_model.embed(texts, parallel=0))
|
||||
sparse_embeds = list(sparse_model.embed(texts, parallel=0))
|
||||
colbert_multivectors = list(colbert_model.embed(texts, parallel=0))
|
||||
|
||||
points = []
|
||||
for i, text in enumerate(texts):
|
||||
sparse_embed = sparse_embeds[i].as_object()
|
||||
dense_embed = dense_embeds[i]
|
||||
colbert_embed = colbert_multivectors[i]
|
||||
|
||||
points.append(
|
||||
models.PointStruct(
|
||||
id=i,
|
||||
vector={
|
||||
"dense": dense_embed,
|
||||
"sparse": sparse_embed,
|
||||
"colbert": colbert_embed,
|
||||
},
|
||||
payload=sample_data[i],
|
||||
)
|
||||
)
|
||||
|
||||
client.upload_points(
|
||||
collection_name=collection_name,
|
||||
points=points,
|
||||
)
|
||||
```
|
||||
|
||||
## Step 3: The Universal Query in Action
|
||||
|
||||
Let's build a sophisticated research discovery query step by step. We'll orchestrate dense search, sparse search, RRF fusion, and ColBERT reranking - all in a single API call.
|
||||
|
||||
### Prepare Query Embeddings
|
||||
|
||||
First, we encode our research query using all three embedding models:
|
||||
|
||||
```python
|
||||
research_query = "transformer architectures for multimodal learning"
|
||||
|
||||
research_query_dense = next(dense_model.query_embed(research_query))
|
||||
research_query_sparse = next(sparse_model.query_embed(research_query)).as_object()
|
||||
research_query_colbert = next(colbert_model.query_embed(research_query))
|
||||
```
|
||||
|
||||
We generate three different representations of the same query - each optimized for a different stage of our retrieval pipeline.
|
||||
|
||||
### Define Global Filter with Automatic Propagation
|
||||
|
||||
Now we define quality constraints that will apply throughout our entire search pipeline:
|
||||
|
||||
```python
|
||||
# Define global filter - this will be propagated to ALL prefetch stages
|
||||
global_filter = models.Filter(
|
||||
must=[
|
||||
# Research domain filtering
|
||||
models.FieldCondition(
|
||||
key="research_area",
|
||||
match=models.MatchAny(any=[
|
||||
"machine_learning",
|
||||
"computer_vision",
|
||||
"nlp",
|
||||
]),
|
||||
),
|
||||
# Open access only
|
||||
models.FieldCondition(
|
||||
key="open_access",
|
||||
match=models.MatchValue(value=True)
|
||||
),
|
||||
# Recent research only (last 6 years)
|
||||
models.FieldCondition(
|
||||
key="published_date",
|
||||
range=models.DatetimeRange(
|
||||
gte=(datetime.now() - timedelta(days=365 * 6)).isoformat()
|
||||
),
|
||||
),
|
||||
# High-impact papers
|
||||
models.FieldCondition(key="impact_score", range=models.Range(gte=0.6)),
|
||||
# Well-cited work
|
||||
models.FieldCondition(key="citation_count", range=models.Range(gte=5)),
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
**Key insight**: This filter will be automatically propagated to ALL prefetch stages. Qdrant doesn't do "late filtering" or "post-filtering" - the filters are applied at the HNSW search level for maximum efficiency. This is enabled by the payload indexes we created in Step 1.
|
||||
|
||||
### Set Up Parallel Prefetch Queries
|
||||
|
||||
Next, we configure our hybrid retrieval with two concurrent searches:
|
||||
|
||||
```python
|
||||
# Prefetch queries - global filter will be automatically applied to both
|
||||
hybrid_query = [
|
||||
# Dense retrieval: semantic understanding
|
||||
models.Prefetch(query=research_query_dense, using="dense", limit=100),
|
||||
# Sparse retrieval: exact technical term matching
|
||||
models.Prefetch(query=research_query_sparse, using="sparse", limit=100),
|
||||
]
|
||||
```
|
||||
|
||||
These two prefetch queries run in parallel:
|
||||
|
||||
- **Dense search**: Finds semantically similar papers based on research concepts - but only considers papers matching our global filter (ML/CV/NLP domains, open access, recent, high-impact, well-cited)
|
||||
- **Sparse search**: Matches exact technical terms like "transformer," "multimodal," "attention" - but only within the same filtered subset
|
||||
- Both searches execute concurrently for maximum speed
|
||||
- Filter propagation happens automatically - no manual coordination needed
|
||||
|
||||
We could retrieve up to 200 candidates total (100 from each search), but many will overlap.
|
||||
|
||||
### Add Fusion Stage
|
||||
|
||||
Now we combine the two filtered candidate lists using Reciprocal Rank Fusion:
|
||||
|
||||
```python
|
||||
# Fusion stage - combines dense and sparse results
|
||||
fusion_query = models.Prefetch(
|
||||
prefetch=hybrid_query,
|
||||
query=models.FusionQuery(fusion=models.Fusion.RRF),
|
||||
limit=100,
|
||||
)
|
||||
```
|
||||
|
||||
The RRF algorithm:
|
||||
- Combines the two filtered candidate lists intelligently
|
||||
- Papers appearing high in both lists get boosted scores
|
||||
- Creates a unified ranking that balances semantic similarity with technical term relevance
|
||||
- All results still satisfy the global filter constraints
|
||||
|
||||
### Execute the Universal Query with ColBERT Reranking
|
||||
|
||||
Finally, we send our complete query that ties everything together:
|
||||
|
||||
```python
|
||||
# The Universal Query: Global filter propagates through all stages
|
||||
response = client.query_points(
|
||||
collection_name=collection_name,
|
||||
prefetch=fusion_query,
|
||||
query=research_query_colbert,
|
||||
using="colbert",
|
||||
query_filter=global_filter, # Propagates to all prefetch stages
|
||||
limit=10,
|
||||
with_payload=True,
|
||||
)
|
||||
```
|
||||
|
||||
This final stage applies ColBERT reranking to the fused results:
|
||||
- Token-level late interaction scoring examines fine-grained text alignment
|
||||
- Each query token is compared against each abstract token
|
||||
- MaxSim aggregation finds the best conceptual alignments
|
||||
- Returns the top 10 papers with precise relevance scores
|
||||
|
||||
**Why this matters**: By applying filters at every stage (not after retrieval), Qdrant maintains high accuracy while avoiding wasted computation on papers that would be filtered out anyway.
|
||||
|
||||
### Display Results
|
||||
|
||||
```python
|
||||
# Display results
|
||||
print("Top Research Papers:")
|
||||
for i, hit in enumerate(response.points or [], 1):
|
||||
paper = hit.payload
|
||||
print(f"{i}. {paper['title']}")
|
||||
print(f" Authors: {', '.join(paper['authors'][:3])}{'...' if len(paper['authors']) > 3 else ''}")
|
||||
print(f" Published: {paper['published_date']} | Citations: {paper['citation_count']}")
|
||||
print(f" Research Area: {paper['research_area']}")
|
||||
print(f" Open Access: {paper['open_access']}")
|
||||
print(f" Score: {hit.score:.4f}\n")
|
||||
```
|
||||
|
||||
And there you have it - a sophisticated multi-stage research discovery system in a single declarative query!
|
||||
|
||||
## Real ArXiv Dataset Integration
|
||||
|
||||
Here's how you could populate the collection with real data (if the endpoint wasn't broken):
|
||||
|
||||
```python
|
||||
# ! pip install arxiv
|
||||
import arxiv
|
||||
|
||||
arxiv_client = arxiv.Client()
|
||||
|
||||
search = arxiv.Search(
|
||||
query="transformer AND multimodal",
|
||||
max_results=2,
|
||||
sort_by=arxiv.SortCriterion.SubmittedDate,
|
||||
)
|
||||
|
||||
points = []
|
||||
for i, paper in enumerate(arxiv_client.results(search)):
|
||||
print(paper)
|
||||
# Create dense embedding from abstract
|
||||
dense_vector = next(dense_model.embed(paper["abstract"]))
|
||||
|
||||
# Create sparse vector from technical terms (simplified)
|
||||
# In practice, you'd use a proper sparse encoder like SPLADE or BM25
|
||||
sparse_vector = next(sparse_model.embed(paper["abstract"])).as_object()
|
||||
|
||||
# Create ColBERT multivector (simplified)
|
||||
colbert_vector = next(colbert_model.embed(paper["abstract"]))
|
||||
|
||||
point = models.PointStruct(
|
||||
id=i, # Extract arXiv ID
|
||||
payload={
|
||||
"title": paper["title"],
|
||||
"authors": [author for author in paper["authors"]],
|
||||
"abstract": paper["abstract"],
|
||||
"published_date": datetime.strptime(
|
||||
paper["published_date"], "%Y-%m-%d"
|
||||
).isoformat(),
|
||||
"citation_count": 0, # Would need external API
|
||||
"venue": "arXiv",
|
||||
"research_area": paper["research_area"],
|
||||
"impact_score": paper["impact_score"],
|
||||
"open_access": True,
|
||||
},
|
||||
vector={
|
||||
"dense": dense_vector,
|
||||
"sparse": sparse_vector,
|
||||
"colbert": colbert_vector,
|
||||
},
|
||||
)
|
||||
points.append(point)
|
||||
|
||||
# Upload to Qdrant
|
||||
client.upsert(collection_name=collection_name, points=points)
|
||||
print(f"Uploaded {len(points)} research papers to collection")
|
||||
```
|
||||
|
||||
## Key Takeaways
|
||||
|
||||
- **Single Request**: Complex multi-stage research discovery in one API call
|
||||
- **Parallel Execution**: Dense and sparse searches run concurrently
|
||||
- **Smart Filtering**: Apply research quality filters at optimal stages
|
||||
- **Real Data**: Works with actual arXiv dataset and research metadata
|
||||
- **Production Ready**: Scales to millions of papers with sub-second latency
|
||||
|
||||
The Universal Query API eliminates the complexity of building multi-turn retrieval systems. What used to require coordination between semantic search engines, keyword systems, and reranking models now happens in a single, optimized request - perfect for academic search, literature reviews, and research recommendation systems.
|
||||
|
||||
## Next
|
||||
|
||||
In the next lesson, you'll take this foundation and build a complete recommendation service with real data ingestion and user profiling.
|
||||
Reference in New Issue
Block a user