mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-04 18:38:30 +02:00
docs: Add course content for Day 2 (#1928)
* docs: Add course content for Day 2 * included feedback * cleaned up spelling and removed repetitive info * added warning about reindexing * pitstop project fixed * finished what-is-hnsw * filterable hnsw file done * updated image embedding * fixed what is hnsw code * header level * updated filterable hnsw * fixed pitstop and added waiting for index (which breaks the whole thing, because index doesnt build). * added discord link * updated discord * updated project structur * updated structure * last fixes aligning notebook and markdown * standart init for client * pitstop structure * pitstop fix * pitstop fix --------- Co-authored-by: Kirstin <kirstin.taufertshoefer@qdrant.com> Co-authored-by: Evgeniya Sukhodolskaya <suxodolskaya97@gmail.com>
This commit is contained in:
co-authored by
Kirstin
Evgeniya Sukhodolskaya
parent
c19e60bf77
commit
b72da43166
@@ -1,9 +1,23 @@
|
||||
---
|
||||
title: Day 2
|
||||
title: "Day 2: Indexing and Performance"
|
||||
isLesson: true
|
||||
weight: 6
|
||||
weight: 3
|
||||
---
|
||||
|
||||
{{< date >}} Day 2 {{< /date >}}
|
||||
|
||||
# Day 2
|
||||
# Indexing and Performance
|
||||
|
||||
Master [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) indexing and practical tuning for fast retrieval.
|
||||
|
||||
---
|
||||
|
||||
## Today’s path
|
||||
|
||||
1. HNSW Indexing Fundamentals
|
||||
2. Combining Vector Search and Filtering
|
||||
3. Demo: HNSW Performance Tuning
|
||||
4. Project: HNSW Performance Benchmarking
|
||||
|
||||
You’ll measure recall and latency impacts to make data‑driven tuning choices.
|
||||
|
||||
|
||||
@@ -0,0 +1,508 @@
|
||||
---
|
||||
title: "Demo: HNSW Performance Tuning"
|
||||
weight: 3
|
||||
---
|
||||
|
||||
{{< date >}} Day 2 {{< /date >}}
|
||||
|
||||
# Demo: HNSW Performance Tuning
|
||||
|
||||
Learn how to improve vector search speed with [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) tuning and payload indexing on a real 100K dataset.
|
||||
|
||||
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course/day_2/hnsw_performance_tuning.ipynb">
|
||||
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
|
||||
</a>
|
||||
|
||||
## What You’ll Do
|
||||
|
||||
Yesterday you learned the theory behind HNSW indexing. Today you'll see it in action on a 100,000-vector dataset, measuring performance differences and applying optimization strategies that work in production.
|
||||
|
||||
**You'll learn to:**
|
||||
- Optimize bulk upload speed with strategic HNSW configuration
|
||||
- Measure the performance impact of payload indexes
|
||||
- Tune HNSW params
|
||||
- Compare full-scan vs. HNSW search performance
|
||||
|
||||
## The Performance Challenge
|
||||
|
||||
Working with 100K high-dimensional vectors (1536 dimensions from OpenAI's text-embedding-3-large) presents real performance challenges:
|
||||
- **Upload speed**: How fast can we ingest vectors?
|
||||
- **Search speed**: How quickly can we find similar vectors?
|
||||
- **Filtering speed**: How much overhead do payload filters add?
|
||||
- **Memory efficiency**: How do different configurations affect RAM needs?
|
||||
|
||||
## Step 1: Environment Setup
|
||||
|
||||
### Install Required Libraries
|
||||
|
||||
|
||||
**Library purposes:**
|
||||
|
||||
`datasets`: Access to Hugging Face datasets, specifically our [DBpedia 100K dataset](https://huggingface.co/datasets/Qdrant/dbpedia-entities-openai3-text-embedding-3-large-1536-100K)
|
||||
`qdrant-client`: Official Qdrant Python client for vector search operations
|
||||
`tqdm`: Progress bars for bulk operations (essential for 100K upload tracking)
|
||||
`openai`: Generate query embeddings compatible with the dataset
|
||||
`python-dotenv`: Secure environment variable management
|
||||
|
||||
### Set Up API Keys
|
||||
|
||||
You'll need an OpenAI API key for query embeddings:
|
||||
|
||||
- Visit [OpenAI's API platform](https://platform.openai.com)
|
||||
- Create an account or sign in
|
||||
- Navigate to [API Keys](https://platform.openai.com/api-keys) and create a new key
|
||||
- **Important**: You'll need credits in your OpenAI account (~$1 should be sufficient for this demo)
|
||||
|
||||
### Environment Configuration
|
||||
|
||||
Create a `.env` file your project directory or use Google Colab secrets.
|
||||
|
||||
```bash
|
||||
# .env file
|
||||
QDRANT_URL=https://your-cluster-url.cloud.qdrant.io
|
||||
QDRANT_API_KEY=your-qdrant-api-key-here
|
||||
OPENAI_API_KEY=sk-your-openai-api-key-here
|
||||
```
|
||||
|
||||
**Security Note**: Never commit the .env file.
|
||||
|
||||
## Step 2: Connect to Qdrant Cloud
|
||||
|
||||
We’ll use Qdrant Cloud for stable resources at 100K scale.
|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
from qdrant_client import QdrantClient, models
|
||||
from tqdm import tqdm
|
||||
import openai
|
||||
import time
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
|
||||
load_dotenv()
|
||||
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
|
||||
|
||||
# For Colab:
|
||||
# from google.colab import userdata
|
||||
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
|
||||
|
||||
# Verify connection
|
||||
try:
|
||||
collections = client.get_collections()
|
||||
print(f"Connected to Qdrant Cloud successfully!")
|
||||
print(f"Current collections: {len(collections.collections)}")
|
||||
except Exception as e:
|
||||
print(f"Connection failed: {e}")
|
||||
print("Check your QDRANT_URL and QDRANT_API_KEY in .env file")
|
||||
```
|
||||
|
||||
**Why Cloud:**
|
||||
- **Convenience**: No local setup hassles
|
||||
- **Free Tier**: We are well within the free tier with a 100k dataset
|
||||
- **Realistic testing**: Production-like environment for accurate benchmarks
|
||||
- **Scalability**: Easy to scale up later
|
||||
|
||||
## Step 3: Load the DBpedia Dataset
|
||||
|
||||
We're using a curated dataset of 100K Wikipedia articles with pre-computed 1536-dimensional embeddings from OpenAI's `text-embedding-3-large` model:
|
||||
|
||||
```python
|
||||
# Load the dataset (this may take a few minutes for first download)
|
||||
print("Loading DBpedia 100K dataset...")
|
||||
ds = load_dataset("Qdrant/dbpedia-entities-openai3-text-embedding-3-large-1536-100K")
|
||||
collection_name = "dbpedia_100K"
|
||||
|
||||
print("Dataset loaded successfully!")
|
||||
print(f"Dataset size: {len(ds['train'])} articles")
|
||||
|
||||
# Explore the dataset structure
|
||||
print("\nDataset structure:")
|
||||
print("Available columns:", ds['train'].column_names)
|
||||
|
||||
# Look at a sample entry
|
||||
sample = ds['train'][0]
|
||||
print(f"\nSample article:")
|
||||
print(f"Title: {sample['title']}")
|
||||
print(f"Text preview: {sample['text'][:200]}...")
|
||||
print(f"Embedding dimensions: {len(sample['text-embedding-3-large-1536-embedding'])}")
|
||||
```
|
||||
|
||||
**About this dataset:**
|
||||
- **Source**: [Hugging Face DBpedia dataset](https://huggingface.co/datasets/Qdrant/dbpedia-entities-openai3-text-embedding-3-large-1536-100K)
|
||||
- **Content**: Pre-computed Wikipedia article embeddings
|
||||
- **Size**: 100,000 articles
|
||||
- **Embeddings**: with OpenAI's `text-embedding-3-large` truncated to 1536 dims
|
||||
- **Metadata**: `_id`, `titles` and `text`
|
||||
|
||||
|
||||
## Step 4: Strategic Collection Creation
|
||||
|
||||
Set `m=0` to skip HNSW graph links during bulk upload. Switch to a normal `m` after ingest to build the graph. This speeds up inserts 5-10x because link creation is deferred.
|
||||
|
||||
**Warning:** Do not toggle back to `m=0` on a collection that already has an HNSW index if you care about keeping that index. Rebuilding from scratch is slow and uses more resources.
|
||||
|
||||
|
||||
```python
|
||||
# Delete collection if it exists (for clean restart)
|
||||
try:
|
||||
client.delete_collection(collection_name)
|
||||
print(f"Deleted existing collection: {collection_name}")
|
||||
except Exception:
|
||||
pass # Collection doesn't exist, which is fine
|
||||
|
||||
# Create collection with optimized settings
|
||||
print(f"Creating collection: {collection_name}")
|
||||
|
||||
client.create_collection(
|
||||
collection_name=collection_name,
|
||||
vectors_config=models.VectorParams(
|
||||
size=1536, # Matches dataset dims
|
||||
distance=models.Distance.COSINE # Good for normalized embeddings
|
||||
),
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=0, # Skip links during upload for speed
|
||||
ef_construct=100, # Used after we set m>0
|
||||
full_scan_threshold=10000
|
||||
),
|
||||
strict_mode_config=models.StrictModeConfig(
|
||||
enabled=False, # More flexible while testing
|
||||
unindexed_filtering_retrieve=True # Allow filters without payload indexes
|
||||
)
|
||||
)
|
||||
|
||||
print(f"Collection '{collection_name}' created successfully!")
|
||||
|
||||
# Verify collection settings
|
||||
collection_info = client.get_collection(collection_name)
|
||||
print(f"Vector size: {collection_info.config.params.vectors.size}")
|
||||
print(f"Distance metric: {collection_info.config.params.vectors.distance}")
|
||||
print(f"HNSW m: {collection_info.config.hnsw_config.m}")
|
||||
```
|
||||
|
||||
|
||||
**Configuration details:**
|
||||
- **`size=1536`**: To match the dimensions parameter we set for the OpenAI `text-embedding-3-large`
|
||||
- **`distance=COSINE`**: Standart for normalized embeddings and semantic similarity
|
||||
- **`full_scan_threshold=10000`**: Uses exact search for smaller result sets
|
||||
- **`strict_mode_config`**: Managed Cloud runs in strict mode by default. We set `enabled=False` to let you experiment with unindexed payload keys during the demo.
|
||||
|
||||
**Side note:** `text-embedding-3-large` outputs 3072 dims. Trunkating that to only 1536 dimensions cuts compute and memory, with some accuracy loss.
|
||||
|
||||
## Step 5: Bulk Upload with Rich Payloads
|
||||
|
||||
We'll upload 100K vectors in 10K batches. The payload includes `title`, `length`, and `has_numbers` for filter tests.
|
||||
|
||||
```python
|
||||
def upload_batch(start_idx, end_idx):
|
||||
points = []
|
||||
for i in range(start_idx, min(end_idx, total_points)):
|
||||
example = ds["train"][i]
|
||||
|
||||
# Get the pre-computed embedding
|
||||
embedding = example["text-embedding-3-large-1536-embedding"]
|
||||
|
||||
# Create payload with fields for filtering tests
|
||||
payload = {
|
||||
"text": example["text"],
|
||||
"title": example["title"],
|
||||
"_id": example["_id"],
|
||||
"length": len(example["text"]),
|
||||
"has_numbers": any(char.isdigit() for char in example["text"]),
|
||||
}
|
||||
|
||||
points.append(models.PointStruct(id=i, vector=embedding, payload=payload))
|
||||
|
||||
if points:
|
||||
client.upload_points(collection_name=collection_name, points=points)
|
||||
return len(points)
|
||||
return 0
|
||||
|
||||
|
||||
batch_size = 10000
|
||||
total_points = len(ds["train"])
|
||||
print(f"Uploading {total_points} points in batches of {batch_size}")
|
||||
|
||||
# Upload all batches with progress tracking
|
||||
total_uploaded = 0
|
||||
for i in tqdm(range(0, total_points, batch_size), desc="Uploading points"):
|
||||
uploaded = upload_batch(i, i + batch_size)
|
||||
total_uploaded += uploaded
|
||||
|
||||
print(f"Upload completed! Total points uploaded: {total_uploaded}")
|
||||
```
|
||||
|
||||
## Step 6: Enable HNSW Indexing
|
||||
|
||||
Now switch from `m=0` to `m=16` to build HNSW connections and improve search time.
|
||||
|
||||
```python
|
||||
client.update_collection(
|
||||
collection_name=collection_name,
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=16 # Each node connects to 16 neighbors
|
||||
)
|
||||
)
|
||||
|
||||
print("HNSW indexing enabled with m=16")
|
||||
```
|
||||
|
||||
**What happens now?** Qdrant builds a navigable graph so search becomes near‑logarithmic instead of linear scanning.
|
||||
|
||||
## Step 7: Create Query Embeddings
|
||||
|
||||
We need to use the same model and dimensions as our dataset to ensure compatibility.
|
||||
|
||||
If you do not have an OpenAI key, use the commented fallback below.
|
||||
|
||||
```python
|
||||
# Optional fallback without an API key:
|
||||
# import requests
|
||||
# test_query = "artificial intelligence"
|
||||
# url = "https://storage.googleapis.com/qdrant-examples/query_embedding_day_2.json"
|
||||
# resp = requests.get(url)
|
||||
# query_embedding = resp.json()["query_vector"]
|
||||
# print(f"Generated embedding for: '{test_query}'")
|
||||
# print(f"Embedding dimensions: {len(query_embedding)}")
|
||||
# print(f"First 5 values: {query_embedding[:5]}")
|
||||
|
||||
|
||||
# Initialize OpenAI client
|
||||
openai_client = openai.OpenAI(api_key=os.getenv('OPENAI_API_KEY'))
|
||||
|
||||
# for colab:
|
||||
# openai_client = openai.OpenAI(api_key=userdata.get('OPENAI_API_KEY'))
|
||||
|
||||
def get_query_embedding(text):
|
||||
"""Generate embedding using the same model as the dataset"""
|
||||
try:
|
||||
response = openai_client.embeddings.create(
|
||||
model="text-embedding-3-large", # Must match dataset model
|
||||
input=text,
|
||||
dimensions=1536 # Must match dataset dimensions
|
||||
)
|
||||
return response.data[0].embedding
|
||||
except Exception as e:
|
||||
print(f"Error getting OpenAI embedding: {e}")
|
||||
print("Common issues:")
|
||||
print(" - Check your OPENAI_API_KEY in .env file")
|
||||
print(" - Ensure you have credits in your OpenAI account")
|
||||
print(" - Verify your API key has embedding permissions")
|
||||
print("Using random vector as fallback for demo purposes...")
|
||||
import numpy as np
|
||||
return np.random.normal(0, 1, 1536).tolist()
|
||||
|
||||
# Test embedding generation
|
||||
print("Generating query embedding...")
|
||||
test_query = "artificial intelligence"
|
||||
query_embedding = get_query_embedding(test_query)
|
||||
print(f"Generated embedding for: '{test_query}'")
|
||||
print(f"Embedding dimensions: {len(query_embedding)}")
|
||||
print(f"First 5 values: {query_embedding[:5]}")
|
||||
```
|
||||
|
||||
**Query Embedding Compatibility:**
|
||||
- **Model**: Must use `text-embedding-3-large` (same as dataset)
|
||||
- **Dimensions**: Must be 1536 (same as dataset)
|
||||
|
||||
## Step 8: Baseline Performance Testing
|
||||
|
||||
Let's measure search performance on the HNSW‑enabled collection.
|
||||
|
||||
```python
|
||||
print("Running baseline performance test...")
|
||||
|
||||
# Warm up the RAM index/vectors cache with a test query
|
||||
print("Warming up caches...")
|
||||
client.query_points(collection_name=collection_name, query=query_embedding, limit=1)
|
||||
|
||||
# Measure vector search performance
|
||||
search_times = []
|
||||
for _ in range(3): # Multiple runs for a stable average
|
||||
start_time = time.time()
|
||||
response = client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query_embedding,
|
||||
limit=10
|
||||
)
|
||||
search_time = (time.time() - start_time) * 1000
|
||||
search_times.append(search_time)
|
||||
|
||||
baseline_time = sum(search_times) / len(search_times)
|
||||
|
||||
print(f"Average search time: {baseline_time:.2f}ms")
|
||||
print(f"Search times: {[f'{t:.2f}ms' for t in search_times]}")
|
||||
print(f"Found {len(response.points)} results")
|
||||
print(f"Top result: '{response.points[0].payload['title']}' (score: {response.points[0].score:.4f})")
|
||||
|
||||
# Show a few more results for context
|
||||
print(f"\nTop 3 results:")
|
||||
for i, point in enumerate(response.points[:3], 1):
|
||||
title = point.payload['title']
|
||||
score = point.score
|
||||
text_preview = point.payload['text'][:100] + "..."
|
||||
print(f" {i}. {title} (score: {score:.4f})")
|
||||
print(f" {text_preview}")
|
||||
```
|
||||
|
||||
**Performance factors:**
|
||||
- **Cache warming**: First query loads relevant index parts/vectors into memory, subsequent queries are faster
|
||||
- **HNSW with m=16**: Graph-based search is much faster than full scan
|
||||
- **MRepeated runs**: Average of several queries gives more reliable timing results
|
||||
|
||||
## Step 9: Filtering Without Payload Indexes
|
||||
|
||||
Now, let's test filtering performance without indexes. This forces Qdrant to scan through vectors and check each one against the filter:
|
||||
|
||||
```python
|
||||
print("Testing filtering without payload indexes")
|
||||
|
||||
# Create a text-based filter
|
||||
text_filter = models.Filter(
|
||||
must=[
|
||||
models.FieldCondition(
|
||||
key="text",
|
||||
match=models.MatchText(text="data")
|
||||
)
|
||||
]
|
||||
)
|
||||
|
||||
# Run multiple times for more reliable measurement
|
||||
unindexed_times = []
|
||||
for i in range(3):
|
||||
start_time = time.time()
|
||||
response = client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query_embedding,
|
||||
limit=10,
|
||||
search_params=models.SearchParams(hnsw_ef=100),
|
||||
query_filter=text_filter
|
||||
)
|
||||
unindexed_times.append((time.time() - start_time) * 1000)
|
||||
|
||||
unindexed_filter_time = sum(unindexed_times) / len(unindexed_times)
|
||||
|
||||
print(f"Filtered search (WITHOUT index): {unindexed_filter_time:.2f}ms")
|
||||
print(f"Individual times: {[f'{t:.2f}ms' for t in unindexed_times]}")
|
||||
print(f"Overhead vs baseline: {unindexed_filter_time - baseline_time:.2f}ms")
|
||||
print(f"Found {len(response.points)} matching results")
|
||||
if response.points:
|
||||
print(f"Top result: '{response.points[0].payload['text']}'\nScore: {response.points[0].score:.4f}")
|
||||
else:
|
||||
print("No results found - try a different filter term")
|
||||
```
|
||||
|
||||
## Step 10: Create Payload Indexes
|
||||
|
||||
Create a [full‑text index](/documentation/concepts/indexing/#full-text-index) for faster filtering.
|
||||
|
||||
```python
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="text",
|
||||
wait=True,
|
||||
field_schema=models.TextIndexParams(
|
||||
type="text",
|
||||
tokenizer="word",
|
||||
phrase_matching=False
|
||||
)
|
||||
)
|
||||
|
||||
print("Payload index created for 'text' field")
|
||||
|
||||
# If you want filter‑aware HNSW and you built the graph before creating payload indexes,
|
||||
# rebuild the graph to attach filter data structures.
|
||||
# Note: Reindexing takes up a lot of resources, and it is advised to set payload
|
||||
# indexes only once, before building HNSW.
|
||||
# client.update_collection(collection_name=collection_name, hnsw_config=models.HnswConfigDiff(m=0))
|
||||
# client.update_collection(collection_name=collection_name, hnsw_config=models.HnswConfigDiff(m=16))
|
||||
```
|
||||
|
||||
## Step 11: Filtering With Payload Indexes
|
||||
|
||||
Run the same query with the index in place.
|
||||
|
||||
```python
|
||||
print("Testing filtering WITH payload indexes...")
|
||||
|
||||
# Run multiple times for more reliable measurement
|
||||
indexed_times = []
|
||||
for i in range(3):
|
||||
start_time = time.time()
|
||||
response = client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query_embedding,
|
||||
limit=10,
|
||||
search_params=models.SearchParams(hnsw_ef=100),
|
||||
query_filter=text_filter
|
||||
)
|
||||
indexed_times.append((time.time() - start_time) * 1000)
|
||||
|
||||
indexed_filter_time = sum(indexed_times) / len(indexed_times)
|
||||
|
||||
print(f"Filtered search (WITH index): {indexed_filter_time:.2f}ms")
|
||||
print(f"Individual times: {[f'{t:.2f}ms' for t in indexed_times]}")
|
||||
print(f"Overhead vs baseline: {indexed_filter_time - baseline_time:.2f}ms")
|
||||
print(f"Found {len(response.points)} matching results")
|
||||
if response.points:
|
||||
print(f"Top result: '{response.points[0].payload['text']}'\nScore: {response.points[0].score:.4f}")
|
||||
else:
|
||||
print("No results found - try a different filter term")
|
||||
```
|
||||
|
||||
## Performance Analysis
|
||||
|
||||
Compare your results and see the effect of each optimization:
|
||||
|
||||
```python
|
||||
print("\n" + "="*60)
|
||||
print("FINAL PERFORMANCE SUMMARY")
|
||||
print("="*60)
|
||||
|
||||
# Key metrics
|
||||
if unindexed_filter_time > 0 and indexed_filter_time > 0:
|
||||
index_speedup = unindexed_filter_time / indexed_filter_time
|
||||
filter_overhead_without = unindexed_filter_time - baseline_time
|
||||
filter_overhead_with = indexed_filter_time - baseline_time
|
||||
else:
|
||||
index_speedup = 0
|
||||
filter_overhead_without = 0
|
||||
filter_overhead_with = 0
|
||||
|
||||
print(f"Baseline search (HNSW m=16): {baseline_time:.2f}ms")
|
||||
print(f"Filtering WITHOUT index: {unindexed_filter_time:.2f}ms")
|
||||
print(f"Filtering WITH index: {indexed_filter_time:.2f}ms")
|
||||
print("")
|
||||
print(f"Performance improvements:")
|
||||
print(f" • Index speedup: {index_speedup:.1f}x faster")
|
||||
print(f" • Filter overhead (no index): +{filter_overhead_without:.2f}ms")
|
||||
print(f" • Filter overhead (with index): +{filter_overhead_with:.2f}ms")
|
||||
print("")
|
||||
print(f"Key insights:")
|
||||
print(f" • HNSW (m=16) enables fast vector search")
|
||||
print(f" • Payload indexes dramatically improve filtering")
|
||||
print(f" • Upload strategy (m=0→m=16) optimizes ingestion")
|
||||
print("="*60)
|
||||
```
|
||||
|
||||
|
||||
## Next Steps & Resources
|
||||
|
||||
**What you've learned:**
|
||||
- Strategic optimization of initial bulk upload with `m=0` → `m=16` switching
|
||||
- Real-world performance measurement techniques
|
||||
- The dramatic impact of payload indexes on filtering
|
||||
- Production-ready configuration patterns
|
||||
|
||||
**Recommended next steps:**
|
||||
1. **Experiment with parameters**: Try different `m` values (8, 32, 64) and `ef_construct` settings
|
||||
2. **Test with your data**: Apply these techniques to your own domain datasets
|
||||
3. **Production deployment**: Use these patterns in real applications
|
||||
4. **Advanced features**: Explore quantization, sharding, and replication
|
||||
|
||||
**Additional resources:**
|
||||
- [Qdrant Documentation](https://qdrant.tech/documentation/) - Complete technical reference
|
||||
- [HNSW Paper](https://arxiv.org/abs/1603.09320) - Original algorithm research
|
||||
- [Qdrant Cloud](https://cloud.qdrant.io/) - Managed vector search service
|
||||
- [Performance Tuning Guide](https://qdrant.tech/documentation/guides/optimization/) - Advanced optimization techniques
|
||||
|
||||
**Ready for the pitstop project?** Now it's your turn to optimize performance with your own dataset and use case. You'll apply these same techniques to your domain-specific data and measure the real-world impact of different HNSW parameters and indexing strategies.
|
||||
@@ -0,0 +1,203 @@
|
||||
---
|
||||
title: Combining Vector Search and Filtering
|
||||
weight: 2
|
||||
---
|
||||
|
||||
{{< date >}} Day 2 {{< /date >}}
|
||||
|
||||
# Combining Vector Search and Filtering
|
||||
|
||||
We've talked about how Qdrant uses the [HNSW](documentation/concepts/indexing/#filtrable-index) graph to efficiently search dense vectors. But in real-world applications, you'll often want to constrain your search using filters. This creates unique challenges for graph traversal that Qdrant solves elegantly.
|
||||
|
||||
<div class="video">
|
||||
<iframe
|
||||
src="https://www.youtube.com/embed/VJVHU47IAik?si=nkaHShqrlaO4fj7O"
|
||||
frameborder="0"
|
||||
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||
referrerpolicy="strict-origin-when-cross-origin"
|
||||
allowfullscreen>
|
||||
</iframe>
|
||||
</div>
|
||||
|
||||
## The Challenge: Filters Break Graph Connectivity
|
||||
|
||||
Consider retrieving items from an online store collection where you only want to show laptops priced under $1,000. That price information, along with the category 'laptop', isn't part of the vector - it lives in the [payload](/documentation/concepts/payload/).
|
||||
|
||||

|
||||
|
||||
When you apply a filter like `price < 1000`, you're essentially restricting which points are eligible during search. This introduces challenges for graph traversal because HNSW depends on both short- and long-range edges to efficiently explore the vector space. It requires any point in the graph to be reachable. But if filtering removes a large portion of those points, the search path can break. You risk missing relevant results - not because they weren't similar, but because they were unreachable under the filter.
|
||||
|
||||
## Naive Approaches and Their Problems
|
||||
|
||||
### Post-Filtering
|
||||
|
||||
One approach is to ignore the filter initially: run the search across the whole dataset, get the top K most similar vectors, then apply the filter afterward.
|
||||
|
||||
**The problem:** You might discard most of the top K results. If the best match that satisfies the filter wasn't in that top K, you won't retrieve it. You waste compute and lose recall because relevant points were never retrieved.
|
||||
|
||||
### Pre-filtering
|
||||
|
||||
Another approach is to filter first, Then search inside the filtered set.
|
||||
|
||||
**The problem:** When filters are too restrictive, they fragment the HNSW graph, breaking connectivity and making traversal inefficient or impossible.
|
||||
|
||||
## Qdrant's Solution: Filterable HNSW
|
||||
|
||||
Qdrant solves this with a smarter approach. We guarantee that the HNSW graph remains connected by creating additional edges to maintain connectivity under filtering. Qdrant builds subgraphs per payload value, then merges them back into the full graph.
|
||||
|
||||
So if you filter `brand = Apple`, Qdrant has already built a connected subgraph of just Apple points, and traversal works fine within that subset.
|
||||
|
||||

|
||||
|
||||
## The Query Planner: Adaptive Strategy
|
||||
|
||||
At query time, Qdrant uses a query planner to determine the appropriate strategy. This planning happens per segment and is based on filter cardinality, index availability, and thresholds like `full_scan_threshold`.
|
||||
|
||||
- **Filter matches many points:** Qdrant performs regular HNSW search but skips over nodes that don't match during traversal. This avoids the cost of pre-filtering a large result set while keeping the search fast and approximate.
|
||||
- **Filter matches few points:** If the filter matches just a tiny portion of the collection, Qdrant might skip HNSW entirely and fall back to a simple full scan if that's faster.
|
||||
|
||||
## Payload Indexing: The Base Layer
|
||||
|
||||
**Critical point:** Qdrant does not index payload fields by default. You must explicitly define which fields to index. Best practice is to create payload indexes before uploading any data so HNSW can build filter-aware links. If you add payload indexes after HNSW was built, searches may run less efficiently until you rebuild. To rebuild, change `m` to recreate HNSW per segment but bear in mind that rebuilding is compute heavy on large collections.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
|
||||
load_dotenv()
|
||||
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
|
||||
|
||||
# For Colab:
|
||||
# from google.colab import userdata
|
||||
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
|
||||
|
||||
collection_name = "store"
|
||||
vector_size = 768
|
||||
|
||||
if client.collection_exists(collection_name=collection_name):
|
||||
client.delete_collection(collection_name=collection_name)
|
||||
|
||||
client.create_collection(
|
||||
collection_name=collection_name,
|
||||
vectors_config=models.VectorParams(
|
||||
size=vector_size,
|
||||
distance=models.Distance.COSINE,
|
||||
),
|
||||
optimizers_config=models.OptimizersConfigDiff(
|
||||
indexing_threshold=100,
|
||||
),
|
||||
)
|
||||
|
||||
# Index frequently filtered fields
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="category",
|
||||
field_schema=models.PayloadSchemaType.KEYWORD,
|
||||
)
|
||||
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="price",
|
||||
field_schema=models.PayloadSchemaType.FLOAT,
|
||||
)
|
||||
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name,
|
||||
field_name="brand",
|
||||
field_schema=models.PayloadSchemaType.KEYWORD,
|
||||
)
|
||||
```
|
||||
|
||||
Add sample data:
|
||||
|
||||
```python
|
||||
# Upload data
|
||||
import random
|
||||
|
||||
points = []
|
||||
for i in range(1000):
|
||||
points.append(
|
||||
models.PointStruct(
|
||||
id=i,
|
||||
vector=[random.random() for _ in range(vector_size)],
|
||||
payload={
|
||||
"category": random.choice(["laptop", "phone", "tablet"]),
|
||||
"price": random.randint(0, 1000),
|
||||
"brand": random.choice(
|
||||
["Apple", "Dell", "HP", "Lenovo", "Asus", "Acer", "Samsung"]
|
||||
),
|
||||
},
|
||||
)
|
||||
)
|
||||
client.upload_points(
|
||||
collection_name=collection_name,
|
||||
points=points,
|
||||
)
|
||||
```
|
||||
|
||||
### Memory Considerations
|
||||
|
||||
Payload indexes consume additional memory, so it's recommended to only index fields used in filtering conditions.
|
||||
|
||||
If memory is limited, prioritise the field that produces the most specific search results. The more varied and granular the payload field values are, the more effective it's index will be.
|
||||
|
||||
## Practical Implementation
|
||||
|
||||
### Setting Up Filterable Search
|
||||
|
||||
The following example illustrates how filtering is carried out in practice:
|
||||
|
||||
```python
|
||||
# Create filter combining multiple conditions
|
||||
filter_conditions = models.Filter(
|
||||
must=[
|
||||
models.FieldCondition(key="category", match=models.MatchValue(value="laptop")),
|
||||
models.FieldCondition(key="price", range=models.Range(lte=1000)),
|
||||
models.FieldCondition(key="brand", match=models.MatchAny(any=["Apple", "Dell", "HP"])),
|
||||
]
|
||||
)
|
||||
|
||||
query_vector = [random.random() for _ in range(vector_size)]
|
||||
|
||||
# Execute filtered search
|
||||
results = client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query_vector,
|
||||
query_filter=filter_conditions,
|
||||
limit=10,
|
||||
search_params=models.SearchParams(hnsw_ef=128),
|
||||
)
|
||||
```
|
||||
|
||||
See more in [the docs](/documentation/concepts/filtering/).
|
||||
|
||||
### Query Planner Decision Matrix
|
||||
|
||||
| Filter Cardinality | Strategy | When Used |
|
||||
| -------------------------- | ------------------------- | -------------------------- |
|
||||
| **High (many matches)** | HNSW with node skipping | Filter matches many points |
|
||||
| **Very low (few matches)** | Full scan over candidates | Tiny result set |
|
||||
|
||||
## Performance Optimization Tips
|
||||
|
||||
|
||||
1. **Index early:** Create payload indexes before HNSW build.
|
||||
1. **Index the right fields:** Create payload indexes for all filtered fields. Prefer high-selectivity fields if memory is limited.
|
||||
3. **Test filter combinations:** Complex multi-field filters gain the most from proper indexing.
|
||||
4. **Tune thresholds**: Adjust `full_scan_threshold` based on your data distribution and query patterns
|
||||
5. **Measure real performance**: Benchmark with your actual data and query patterns to validate planner decisions
|
||||
|
||||
## Key Takeaways
|
||||
|
||||
**Filterable HNSW is not a separate indexing mechanism**, it extends the HNSW graph by adding extra edges based on stored payload values to maintain traversal performance under filtering constraints.
|
||||
|
||||
**The query planner** automatically selects the optimal strategy based on filter selectivity, available indexes, and segment characteristics.
|
||||
|
||||
**Payload indexing** is essential for filtered search performance, especially with large datasets and complex filter conditions.
|
||||
|
||||
**Always benchmark** with your specific data to understand how the planner behaves and optimize accordingly.
|
||||
|
||||
In the next section, we'll define a collection with structured payloads, configure payload indexing, and evaluate how different HNSW parameters impact filtered search performance.
|
||||
|
||||
Learn more: [Filterable HNSW Article](https://qdrant.tech/articles/filtrable-hnsw/)
|
||||
@@ -0,0 +1,384 @@
|
||||
---
|
||||
title: "Project: HNSW Performance Benchmarking"
|
||||
weight: 4
|
||||
---
|
||||
|
||||
{{< date >}} Day 2 {{< /date >}}
|
||||
|
||||
# Project: HNSW Performance Benchmarking
|
||||
|
||||
Now that you've seen how [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) parameters and payload indexes affect performance with the DBpedia dataset, it's time to optimize for your own domain and use case.
|
||||
|
||||
## Your Mission
|
||||
|
||||
Build on your Day 1 search engine by adding performance optimization. You'll discover which HNSW settings work best for your specific data and queries, and measure the real impact of payload indexing.
|
||||
|
||||
**Estimated Time:** 90 minutes
|
||||
|
||||
## What You'll Build
|
||||
|
||||
A performance-optimized version of your Day 1 search engine that demonstrates:
|
||||
|
||||
- **Fast bulk load**: Load with `m=0`, then switch to HNSW
|
||||
- **HNSW parameter tuning**: Try different `m` and `ef_construct`
|
||||
- **Payload indexing impact**: Time filtering with and without indexes
|
||||
- **Domain findings**: What works best for your content
|
||||
|
||||
## Setup
|
||||
|
||||
### Prerequisites
|
||||
|
||||
* Qdrant Cloud cluster (URL + API key)
|
||||
* Python 3.9+ (or Google Colab)
|
||||
* Packages: `qdrant-client`, `sentence-transformers`, `python-dotenv`, `numpy`
|
||||
|
||||
### Models
|
||||
|
||||
* Embeddings: `sentence-transformers/all-MiniLM-L6-v2` (384-dim)
|
||||
|
||||
### Dataset
|
||||
|
||||
* Reuse your Day 1 domain data or prepare a dataset with **1,000+ items** and a rich text field (e.g., `description`).
|
||||
* Include a few numeric fields for filtering (e.g., `length`, `word_count`) so payload indexing impact can be measured.
|
||||
|
||||
## Build Steps
|
||||
|
||||
### Step 1: Extend Your Day 1 Project
|
||||
|
||||
Start with your domain search engine from Day 1, or create a new one with 1000+ items:
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
from sentence_transformers import SentenceTransformer
|
||||
import time
|
||||
import numpy as np
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
|
||||
load_dotenv()
|
||||
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
|
||||
|
||||
# For Colab:
|
||||
# from google.colab import userdata
|
||||
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
|
||||
|
||||
encoder = SentenceTransformer("all-MiniLM-L6-v2")
|
||||
```
|
||||
|
||||
### Step 2: Create Multiple Test Collections
|
||||
|
||||
Test different HNSW configurations to find what works best:
|
||||
|
||||
```python
|
||||
# Test configurations
|
||||
configs = [
|
||||
{"name": "fast_initial_upload", "m": 0, "ef_construct": 100},
|
||||
{"name": "memory_optimized", "m": 8, "ef_construct": 100},
|
||||
{"name": "balanced", "m": 16, "ef_construct": 200},
|
||||
{"name": "high_quality", "m": 32, "ef_construct": 400},
|
||||
]
|
||||
|
||||
for config in configs:
|
||||
collection_name = f"my_domain_{config['name']}"
|
||||
if client.collection_exists(collection_name=collection_name):
|
||||
client.delete_collection(collection_name=collection_name)
|
||||
|
||||
client.create_collection(
|
||||
collection_name=collection_name,
|
||||
vectors_config=models.VectorParams(size=384, distance=models.Distance.COSINE),
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=config["m"], ef_construct=config["ef_construct"], full_scan_threshold=10
|
||||
),
|
||||
optimizers_config=models.OptimizersConfigDiff(indexing_threshold=0),
|
||||
strict_mode_config=models.StrictModeConfig(
|
||||
unindexed_filtering_retrieve=True, unindexed_filtering_update=True
|
||||
),
|
||||
)
|
||||
print(f"Created collection: {collection_name}")
|
||||
```
|
||||
|
||||
### Step 3: Upload and Time
|
||||
|
||||
Measure upload performance for each configuration:
|
||||
|
||||
```python
|
||||
def upload_with_timing(collection_name, data, config_name):
|
||||
embeddings = [encoder.encode(dat["description"]).tolist() for dat in data]
|
||||
points = []
|
||||
for i, item in enumerate(data):
|
||||
embedding = embeddings[i]
|
||||
|
||||
points.append(
|
||||
models.PointStruct(
|
||||
id=i,
|
||||
vector=embedding,
|
||||
payload={
|
||||
**item,
|
||||
"length": len(item["description"]),
|
||||
"word_count": len(item["description"].split()),
|
||||
"has_keywords": any(
|
||||
keyword in item["description"].lower()
|
||||
for keyword in ["important", "key", "main"]
|
||||
),
|
||||
},
|
||||
)
|
||||
)
|
||||
|
||||
start_time = time.time()
|
||||
client.upload_points(collection_name=collection_name, points=points)
|
||||
upload_time = time.time() - start_time
|
||||
|
||||
print(f"{config_name}: Uploaded {len(points)} points in {upload_time:.2f}s")
|
||||
return upload_time
|
||||
|
||||
|
||||
# Load your dataset here
|
||||
# your_dataset = [{"description": "This is a description of a product"}, ...]
|
||||
|
||||
# Upload to each collection
|
||||
upload_times = {}
|
||||
for config in configs:
|
||||
collection_name = f"my_domain_{config['name']}"
|
||||
upload_times[config["name"]] = upload_with_timing(
|
||||
collection_name, your_dataset, config["name"]
|
||||
)
|
||||
|
||||
# Wait for index to be built
|
||||
def wait_for_index_built(collection_name, vectors_per_point=1):
|
||||
info = client.get_collection(collection_name=collection_name)
|
||||
count = 0
|
||||
while info.points_count * vectors_per_point - info.indexed_vectors_count != 0 and count < 10:
|
||||
time.sleep(1)
|
||||
info = client.get_collection(collection_name=collection_name)
|
||||
count += 1
|
||||
if count == 10:
|
||||
raise Exception(
|
||||
f"Indexed vectors count ({info.indexed_vectors_count}) is not equal to points count ({info.points_count}). Upload enough points to trigger index rebuild."
|
||||
)
|
||||
|
||||
|
||||
for config in configs:
|
||||
collection_name = f"my_domain_{config['name']}"
|
||||
wait_for_index_built(collection_name)
|
||||
|
||||
```
|
||||
|
||||
### Step 4: Benchmark Search Performance
|
||||
|
||||
Test search speed with different `hnsw_ef` values:
|
||||
|
||||
```python
|
||||
def benchmark_search(collection_name, query_embedding, ef_values=[64, 128, 256]):
|
||||
# Warmup
|
||||
_ = client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query_embedding,
|
||||
limit=10,
|
||||
search_params=models.SearchParams(hnsw_ef=ef_values[0]),
|
||||
)
|
||||
|
||||
results = {}
|
||||
for hnsw_ef in ef_values:
|
||||
times = []
|
||||
|
||||
# Run multiple queries for more reliable timing
|
||||
for _ in range(5):
|
||||
start_time = time.time()
|
||||
|
||||
_ = client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query_embedding,
|
||||
limit=10,
|
||||
search_params=models.SearchParams(hnsw_ef=hnsw_ef),
|
||||
)
|
||||
|
||||
times.append((time.time() - start_time) * 1000)
|
||||
|
||||
results[hnsw_ef] = {
|
||||
"avg_time": np.mean(times),
|
||||
"min_time": np.min(times),
|
||||
"max_time": np.max(times),
|
||||
}
|
||||
|
||||
return results
|
||||
|
||||
|
||||
test_query = "your test query"
|
||||
query_embedding = encoder.encode(test_query).tolist()
|
||||
|
||||
performance_results = {}
|
||||
for config in configs:
|
||||
if config["m"] > 0: # Skip m=0 collections for search
|
||||
collection_name = f"my_domain_{config['name']}"
|
||||
performance_results[config["name"]] = benchmark_search(
|
||||
collection_name, query_embedding
|
||||
)
|
||||
```
|
||||
|
||||
### Step 5: Measure Payload Indexing Impact
|
||||
|
||||
Measure filtering performance with and without indexes:
|
||||
|
||||
```python
|
||||
def test_filtering_performance(collection_name):
|
||||
query_embedding = encoder.encode("your filter test query").tolist()
|
||||
|
||||
# Test filter without index
|
||||
filter_condition = models.Filter(
|
||||
must=[models.FieldCondition(key="length", range=models.Range(gte=100, lte=500))]
|
||||
)
|
||||
|
||||
# Timing without payload index
|
||||
start_time = time.time()
|
||||
_ = client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query_embedding,
|
||||
query_filter=filter_condition,
|
||||
limit=10,
|
||||
)
|
||||
time_without_index = (time.time() - start_time) * 1000
|
||||
|
||||
# Create payload index
|
||||
client.create_payload_index(
|
||||
collection_name=collection_name, field_name="length", field_schema="integer"
|
||||
)
|
||||
|
||||
# Rebuild HNSW to attach filter data structures.
|
||||
# Note: This is not advised for production. Better create payload index before uploading any data to avoid rebuild.
|
||||
suffix = collection_name.replace("my_domain_", "")
|
||||
config = next((c for c in configs if c["name"] == suffix), None)
|
||||
|
||||
client.update_collection(
|
||||
collection_name=collection_name, hnsw_config=models.HnswConfigDiff(m=0)
|
||||
)
|
||||
|
||||
client.update_collection(
|
||||
collection_name=collection_name,
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=16,
|
||||
ef_construct=config["ef_construct"],
|
||||
full_scan_threshold=10,
|
||||
payload_m=None,
|
||||
max_indexing_threads=1,
|
||||
),
|
||||
optimizers_config=models.OptimizersConfigDiff(vacuum_min_vector_number=0),
|
||||
)
|
||||
|
||||
# Wait for index to be built
|
||||
wait_for_index_built(collection_name)
|
||||
|
||||
# Timing with index
|
||||
start_time = time.time()
|
||||
_ = client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query_embedding,
|
||||
query_filter=filter_condition,
|
||||
limit=10,
|
||||
)
|
||||
time_with_index = (time.time() - start_time) * 1000
|
||||
|
||||
return {
|
||||
"without_index": time_without_index,
|
||||
"with_index": time_with_index,
|
||||
"speedup": time_without_index / time_with_index,
|
||||
}
|
||||
|
||||
|
||||
# Test on your best performing collection
|
||||
best_collection = "my_domain_balanced" # Choose based on your results
|
||||
filtering_results = test_filtering_performance(best_collection)
|
||||
```
|
||||
|
||||
### Step 6: Analyze Your Results
|
||||
|
||||
Create a summary of your findings:
|
||||
|
||||
```python
|
||||
print("=" * 60)
|
||||
print("PERFORMANCE OPTIMIZATION RESULTS")
|
||||
print("=" * 60)
|
||||
|
||||
print("\n1) Upload Performance:")
|
||||
for config_name, time_taken in upload_times.items():
|
||||
print(f" {config_name}: {time_taken:.2f}s")
|
||||
|
||||
print("\n2) Search Performance (hnsw_ef=128):")
|
||||
for config_name, results in performance_results.items():
|
||||
if 128 in results:
|
||||
print(f" {config_name}: {results[128]['avg_time']:.2f}ms")
|
||||
|
||||
print("\n3) Filtering Impact:")
|
||||
print(f" Without index: {filtering_results['without_index']:.2f}ms")
|
||||
print(f" With index: {filtering_results['with_index']:.2f}ms")
|
||||
print(f" Speedup: {filtering_results['speedup']:.1f}x")
|
||||
```
|
||||
|
||||
## Success Criteria
|
||||
|
||||
You'll know you've succeeded when:
|
||||
|
||||
<input type="checkbox"> You've tested multiple HNSW configurations with real timing data
|
||||
<input type="checkbox"> You can explain which settings work best for your domain and why
|
||||
<input type="checkbox"> You've measured the concrete impact of payload indexing
|
||||
<input type="checkbox"> You have clear recommendations for production deployment
|
||||
|
||||
|
||||
## Share Your Discovery
|
||||
|
||||
### Step 1: Reflect on Your Findings
|
||||
|
||||
1. Which HNSW configuration (`m`, `ef_construct`) worked best for your domain?
|
||||
2. How did the balance between upload time and search speed look?
|
||||
3. What was the impact of adding a payload index?
|
||||
4. How do your results compare to the DBpedia demo?
|
||||
|
||||
### Step 2: Post Your Results
|
||||
|
||||
**Post your results in** <a href="https://discord.com/invite/qdrant" target="_blank" rel="noopener noreferrer" aria-label="Qdrant Discord"> <img src="https://img.shields.io/badge/Qdrant%20Discord-5865F2?style=flat&logo=discord&logoColor=white&labelColor=5865F2&color=5865F2"
|
||||
alt="Post your results in Discord"
|
||||
style="display:inline; margin:0; vertical-align:middle; border-radius:9999px;" /> </a> **using this:**
|
||||
|
||||
```markdown
|
||||
**[Day 2] HNSW Performance Benchmarking**
|
||||
|
||||
**High-Level Summary**
|
||||
- **Domain:** "[your domain]"
|
||||
- **Key Result:** "m=[..], ef_construct=[..], hnsw_ef=[..] gave [X] ms search and [Y] s upload (best balance)."
|
||||
|
||||
**Reproducibility**
|
||||
- **Collections:** ...
|
||||
- **Model:** sentence-transformers/all-MiniLM-L6-v2 (384-dim)
|
||||
- **Dataset:** [N items] (snapshot: YYYY-MM-DD)
|
||||
|
||||
**Configuration Results**
|
||||
| m | ef_construct | Upload_s | Search_ms@ef=128 |
|
||||
|----|--------------|----------|------------------|
|
||||
| 0 | 100 | X.X | — |
|
||||
| 8 | 100 | Y.Y | A.A |
|
||||
| 16 | 200 | Z.Z | B.B |
|
||||
| 32 | 400 | W.W | C.C |
|
||||
|
||||
**Filtering Impact**
|
||||
- Payload index on `length`: **[speedup]×**
|
||||
Without index: [T1] ms → With index: [T2] ms
|
||||
|
||||
**Recommendations**
|
||||
- Best config for this domain: [m, ef_construct, hnsw_ef]
|
||||
- When to pick another setting: [short guidance]
|
||||
- Notes for production: [one line on indexing order / filters]
|
||||
|
||||
**Surprise**
|
||||
- "[one unexpected finding]"
|
||||
|
||||
**Next Step**
|
||||
- "[one concrete action you’ll try next]"
|
||||
```
|
||||
|
||||
## Optional: Go Further
|
||||
|
||||
* Test more granular parameters:
|
||||
|
||||
* **ef_construct** impact on recall & build time
|
||||
* **hnsw_ef** per-query tuning by complexity
|
||||
* Track memory usage differences (RAM/on-disk, payload indexes)
|
||||
* Add accuracy metrics vs. a small labeled query set to see if higher `m` truly improves quality for your domain
|
||||
@@ -0,0 +1,337 @@
|
||||
---
|
||||
title: HNSW Indexing Fundamentals
|
||||
weight: 1
|
||||
---
|
||||
|
||||
{{< date >}} Day 2 {{< /date >}}
|
||||
|
||||
# HNSW Indexing Fundamentals
|
||||
|
||||
At this point, you've learned how vector search retrieves the nearest vectors to a query using cosine similarity, dot product, or Euclidean distance. How does this work at scale?
|
||||
|
||||
<div class="video">
|
||||
<iframe
|
||||
src="https://www.youtube.com/embed/-q-pLgGDYr4?si=Ln0WYymciqPxQKJl"
|
||||
frameborder="0"
|
||||
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
|
||||
referrerpolicy="strict-origin-when-cross-origin"
|
||||
allowfullscreen>
|
||||
</iframe>
|
||||
</div>
|
||||
|
||||
## Why Vector Search Needs Indexing
|
||||
### The Vector Search Challenge
|
||||
|
||||
You might wonder if Qdrant calculates the distance to every single vector in your collection for each query. This method, known as brute force search, technically works but with millions or billions of vectors this is too slow per query.
|
||||
|
||||
Fortunately, Qdrant speeds things up with **[HNSW — Hierarchical Navigable Small World](https://qdrant.tech/articles/filtrable-hnsw/)**.
|
||||
|
||||
### The Library Analogy
|
||||
|
||||
To understand HNSW, imagine walking into a library with millions of books, looking for a specific book based on a brief description of the content. If the books were all piled up randomly, you'd have to check every book individually to find the one that matches your description best - that's brute force. While it would eventually work, it would take forever.
|
||||
|
||||
But libraries aren't organized like that. They're structured to naturally guide you to the right section. You walk into a library and at the top level, you choose whether your book is in the fiction or nonfiction section. Then, you narrow it down by genre - history, science, or biography. After that, you move to more specific subcategories or alphabetical order until you reach the exact shelf where your book is located.
|
||||
|
||||
## How HNSW Works
|
||||
|
||||
### Graph Structure
|
||||
|
||||
HNSW works similarly by building a multi-layered graph where each vector is a node. The idea is that the graph has a hierarchical structure, where the top layer contains a smaller number of nodes that are broadly connected, and each lower layer has more nodes with increasingly specific connections.
|
||||
|
||||

|
||||
|
||||
### The Search Process
|
||||
|
||||
When a query is performed, HNSW starts from an entry point at the top layer and navigates down through the graph, progressively moving from broader to more precise connections. At each layer, the algorithm explores the nearest neighbors of the current node to determine the best path forward. It continues this process, refining the search as it descends through the layers, until it reaches the lowest level, where it selects the final nearest neighbors.
|
||||
|
||||
This way, Qdrant avoids brute force and quickly narrows down the search space on large datasets.
|
||||
|
||||
## Configuring HNSW
|
||||
|
||||
We can fine-tune how our HNSW graph is built to balance search speed, accuracy, memory usage, and indexing time, depending on our use case needs. Qdrant allows you to control how the HNSW index behaves with three key parameters: `m`, `hnsw_ef`, and `ef_construct`.
|
||||
|
||||
### Graph Connectivity: `m`
|
||||
|
||||
The `m` parameter controls the maximum number of connections per node in the graph.
|
||||
|
||||
- Higher `m`: results in a denser graph where each vector is connected to more neighbors, which improves search accuracy because the graph has more pathways to traverse, making it less likely to miss relevant vectors. However, this also increases memory usage and indexing time since more connections must be maintained.
|
||||
- Lower `m`: makes the graph sparser, reducing memory and speeding up insertion. However, search may become less accurate since fewer paths are available for traversal.
|
||||
- Typical values: between 8 and 64
|
||||
|
||||
```python
|
||||
from qdrant_client.models import HnswConfig
|
||||
|
||||
# Example m values
|
||||
fast_config = HnswConfig(m=8, ef_construct=100, full_scan_threshold=10000) # Lower recall, less memory, faster build
|
||||
balanced_config = HnswConfig(m=16, ef_construct=100, full_scan_threshold=10000) # Default - good balance
|
||||
accurate_config = HnswConfig(m=32, ef_construct=100, full_scan_threshold=10000) # Better recall, more memory, slower build
|
||||
```
|
||||
|
||||
### Build Thoroughness: `ef_construct`
|
||||
|
||||
The `ef_construct` parameter controls how many candidates are checked while inserting a new vector.
|
||||
|
||||
- Higher `ef_construct`: means more neighbors are evaluated, resulting in a more comprehensive and accurate graph. However, this also makes the indexing process slower and more computationally demanding.
|
||||
- Lower `ef_construct`: speeds up the insertion process, but the graph may end up with less optimal connections, which can impact search accuracy.
|
||||
- Common range: between 100 and 500. Complex data can require higher values to maintain reliable connections.
|
||||
|
||||
|
||||
```python
|
||||
# Example ef_construct values
|
||||
fast_build = HnswConfig(m=16, ef_construct=100, full_scan_threshold=10000) # Default - Faster indexing, lower quality
|
||||
balanced_build = HnswConfig(m=16, ef_construct=200, full_scan_threshold=10000) # Good balance
|
||||
quality_build = HnswConfig(m=16, ef_construct=400, full_scan_threshold=10000) # Slower indexing, higher quality
|
||||
```
|
||||
|
||||
### Search Thoroughness: `hnsw_ef`
|
||||
|
||||
The `hnsw_ef` parameter determines the number of candidates evaluated during a search query.
|
||||
|
||||
- Higher `hnsw_ef`: lead to more accurate search results since the algorithm explores a larger neighborhood. However, they also increase query time because more nodes are processed.
|
||||
- Lower `hnsw_ef`: speed up the search but may reduce accuracy since fewer candidate vectors are considered.
|
||||
- Typical range: 50–200+ depending on latency targets.
|
||||
|
||||
```python
|
||||
from qdrant_client.models import SearchParams
|
||||
|
||||
# hnsw_ef is set at search time, not build time
|
||||
fast_search = SearchParams(hnsw_ef=32) # Very fast, lower recall
|
||||
balanced_search = SearchParams(hnsw_ef=128) # Default - good balance
|
||||
accurate_search = SearchParams(hnsw_ef=256) # Higher recall, slower
|
||||
```
|
||||
|
||||
### Parameter Summary
|
||||
|
||||
| Parameter | Purpose | Effect |
|
||||
| ---------------- | ---------------------------- | ------------------------------------------------ |
|
||||
| **m** | Links per node | Up improves recall; uses more RAM and build time |
|
||||
| **ef_construct** | Candidates checked on insert | Up improves graph quality; slows indexing |
|
||||
| **hnsw_ef** | Candidates checked on search | Up improves recall; slows queries |
|
||||
|
||||
## Choosing Settings
|
||||
### Optimizing for Different Workloads
|
||||
|
||||
- **High-speed retrieval:** lower `m` and `hnsw_ef`; set `ef_construct` just high enough for acceptable recall.
|
||||
- **Maximum recall:** raise `m`, `hnsw_ef`, and `ef_construct` and accept slower queries and builds.
|
||||
- **Tight RAM:** reduce `m`; keep `ef_construct` high enough to avoid poor links.
|
||||
|
||||
### Memory & Indexing Behavior
|
||||
|
||||
Some vectors can remain unindexed depending on optimizer settings e.g. when the unindexed part stays below the `indexing_threshold` (kB).
|
||||
|
||||
Small collections or low-dimensional vectors may not trigger HNSW indexing at all. In such cases, full-scan search (brute force) is used instead until indexing becomes beneficial
|
||||
|
||||
|
||||
## HNSW in Action
|
||||
### Why It Is Fast and Scalable
|
||||
|
||||
- **Sublinear Search Scaling**: Unlike O(N) brute force search, HNSW search grows roughly logarithmically with the number of vectors. This makes million-scale datasets searchable in milliseconds rather than minutes.
|
||||
|
||||
- **Filter‑Aware**: Qdrant extends HNSW with filter-aware indexing (Filterable HNSW), allowing fast searches under structured conditions. This avoids costly full scans when filtering by metadata.
|
||||
|
||||
- **Large-Scale Use**:
|
||||
- Supports real-time updates while maintaining high recall
|
||||
- Fits semantic search and recommendation systems
|
||||
- Scales from thousands to billions of vectors
|
||||
|
||||
### Practical Configuration Examples
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
|
||||
load_dotenv()
|
||||
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
|
||||
|
||||
# For Colab:
|
||||
# from google.colab import userdata
|
||||
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
|
||||
|
||||
# Production configuration
|
||||
client.create_collection(
|
||||
collection_name="production_vectors",
|
||||
vectors_config=models.VectorParams(
|
||||
size=768,
|
||||
distance=models.Distance.COSINE,
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=16, # Balanced connections (default)
|
||||
ef_construct=200, # Good build quality (default)
|
||||
full_scan_threshold=10000, # Use brute force below this size (default)
|
||||
),
|
||||
),
|
||||
)
|
||||
|
||||
# Development / testing: faster builds
|
||||
client.create_collection(
|
||||
collection_name="dev_vectors",
|
||||
vectors_config=models.VectorParams(
|
||||
size=384,
|
||||
distance=models.Distance.COSINE,
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=8, # Fewer connections
|
||||
ef_construct=100, # Faster builds
|
||||
full_scan_threshold=10000, # Use brute force below this size (default)
|
||||
),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
## Performance Benchmarking
|
||||
|
||||
Let's test the performance. First we upload some toy data to a new collection:
|
||||
|
||||
```python
|
||||
import time
|
||||
from qdrant_client import QdrantClient, models
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
|
||||
load_dotenv()
|
||||
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
|
||||
|
||||
# For Colab:
|
||||
# from google.colab import userdata
|
||||
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
|
||||
|
||||
collection_name = "my_collection"
|
||||
|
||||
if client.collection_exists(collection_name=collection_name):
|
||||
client.delete_collection(collection_name=collection_name)
|
||||
|
||||
|
||||
# Development / testing: faster builds
|
||||
client.create_collection(
|
||||
collection_name=collection_name,
|
||||
vectors_config=models.VectorParams(
|
||||
size=4,
|
||||
distance=models.Distance.COSINE,
|
||||
hnsw_config=models.HnswConfigDiff(
|
||||
m=8, # Fewer connections
|
||||
ef_construct=100, # Faster builds
|
||||
full_scan_threshold=100, # Use brute force below this size (default)
|
||||
),
|
||||
),
|
||||
optimizers_config=models.OptimizersConfigDiff(
|
||||
indexing_threshold=100, # Use brute force below this size (default)
|
||||
),
|
||||
)
|
||||
|
||||
# upload data
|
||||
import random
|
||||
|
||||
points = []
|
||||
for i in range(20000):
|
||||
points.append(
|
||||
models.PointStruct(id=i, vector=[random.random() for _ in range(4)], payload={})
|
||||
)
|
||||
client.upload_points(
|
||||
collection_name=collection_name,
|
||||
points=points,
|
||||
)
|
||||
```
|
||||
|
||||
### Search Performance
|
||||
|
||||
This simple local benchmark, intended as a toy example, since this approach does not take into account other factors, such as network speed.
|
||||
|
||||
```python
|
||||
def benchmark_search_performance(collection_name, test_queries, ef_values):
|
||||
"""Compare latency across hnsw_ef values"""
|
||||
|
||||
results = {}
|
||||
for hnsw_ef in ef_values:
|
||||
start_time = time.time()
|
||||
for query in test_queries:
|
||||
client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query,
|
||||
limit=10,
|
||||
search_params=models.SearchParams(hnsw_ef=hnsw_ef),
|
||||
)
|
||||
|
||||
avg_time = (time.time() - start_time) / len(test_queries)
|
||||
results[hnsw_ef] = avg_time
|
||||
print(f"hnsw_ef={hnsw_ef}: {avg_time:.3f}s per query")
|
||||
|
||||
return results
|
||||
|
||||
|
||||
# Test different hnsw_ef values
|
||||
test_queries = [
|
||||
[30, 60, 90, 120],
|
||||
[150, 180, 210, 240],
|
||||
[270, 300, 330, 360],
|
||||
[390, 420, 450, 480],
|
||||
[510, 540, 570, 600],
|
||||
]
|
||||
|
||||
ef_values = [32, 64, 128, 256]
|
||||
performance = benchmark_search_performance(collection_name, test_queries, ef_values)
|
||||
```
|
||||
|
||||
### Inspecting Performance and Index Use
|
||||
|
||||
Use [`get_collection`](/api-reference/collections/get-collection) to inspect your collection. It returns Current statistics and configuration of the collection like `points_count`, `indexed_vectors_count` or `hnsw_config`. It also lists `payload_schema` for payload indexes you created.
|
||||
|
||||
To see whether your data is actually indexed check vector and point counts: if `indexed_vectors_count` is far below `points_count * vectors_per_point`, a large part of your data is not in HNSW yet.
|
||||
|
||||
If queries feel slow check:
|
||||
- whether filter fields have [payload indexes](/documentation/concepts/indexing/#payload-index).
|
||||
- if payload indexes have been set before building HNSW graph with setting `m>0`
|
||||
- if the payload indexes have been set before building the HNSW graph (HNSW graph building begins when you switch from `m = 0` to `m > 0`)
|
||||
- if `hnsw_config.full_scan_threshold` is too high.
|
||||
|
||||
|
||||
```python
|
||||
# Inspect collection status
|
||||
info = client.get_collection(collection_name)
|
||||
|
||||
vectors_per_point = 1 # set per your vectors_config
|
||||
vectors_count = info.points_count * vectors_per_point
|
||||
|
||||
print(f"Total vectors: {vectors_count}")
|
||||
print(f"Indexed vectors: {info.indexed_vectors_count}")
|
||||
print(f"HNSW config: {info.config.hnsw_config}")
|
||||
|
||||
if vectors_count:
|
||||
proportion_unindexed = 1 - (info.indexed_vectors_count / vectors_count)
|
||||
else:
|
||||
proportion_unindexed = 0
|
||||
|
||||
print(f"Proportion unindexed: {proportion_unindexed:.2%}")
|
||||
```
|
||||
|
||||
## When Not to Use HNSW
|
||||
|
||||
### Small Collections
|
||||
For collections with fewer than 10,000 vectors, brute force is often faster and uses less RAM than building HNSW.
|
||||
|
||||
### Exact Search Requirements
|
||||
HNSW is approximate. If you need exact results, use brute force.
|
||||
|
||||
### Extreme Memory Constraints
|
||||
|
||||
For very tight RAM budgets consider these solutions:
|
||||
- **Lower `m`**: HNSW uses additional memory proportional to `m × vectors_count`.
|
||||
- **Vector quantization (VQ):**
|
||||
- Scalar quantization (SQ) often cuts RAM ~4×
|
||||
- Binary quantization (BQ) compresses to 1‑bit per dimension and can cut RAM by large factors.
|
||||
- **On Disk Storage**: Set `on_disk=true` for vectors and the HNSW index to use mmap files where only the most frequently accessed vectors are cached in RAM.
|
||||
- **Disable HNSW for Reranking Embeddings:** Useful since reranking vectors such as multi‑vectors are large.
|
||||
|
||||
## Key Takeaways
|
||||
|
||||
1. **HNSW**: reduces O(N) scans to roughly O(log N) graph search
|
||||
2. **`m`**: controls graph density - more connections improve accuracy but use more memory
|
||||
3. **`ef_construct`**: affects graph quality - higher values create more granular graphs but take longer to build
|
||||
4. **`hnsw_ef`**: controls search thoroughness - tune at query time for the speed/accuracy trade-off you need
|
||||
5. **`indexed_vectors_count`**: track to confirm HNSW indexing
|
||||
|
||||
## What's Next
|
||||
|
||||
Now you understand how HNSW makes vector search fast and scalable. Next we'll combine fast search with complex filters using Qdrant’s filter‑aware HNSW.
|
||||
|
||||
Learn more: [HNSW in Qdrant Documentation](https://qdrant.tech/documentation/concepts/indexing/#vector-index)
|
||||
|
||||
Ready to see how HNSW handles real-world filtering scenarios? Let's continue!
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 166 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 142 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 42 KiB |
Reference in New Issue
Block a user