diff --git a/qdrant-landing/content/course/essentials/_index.md b/qdrant-landing/content/course/essentials/_index.md
index 9274064f4..6a38ab68a 100644
--- a/qdrant-landing/content/course/essentials/_index.md
+++ b/qdrant-landing/content/course/essentials/_index.md
@@ -58,7 +58,7 @@ Build the vector search skills that matter: hybrid retrieval, multivector rerank
- Qdrant data modeling: points, payloads, and schemas
- Embeddings, chunking, and similarity metrics
-- Indexing and retrieval tuning ([HNSW](https://qdrant.tech/articles/filtrable-hnsw/), filters, recall/latency)
+- Indexing and retrieval tuning (HNSW, filters, recall/latency)
- Hybrid search with sparse + dense vectors and re-ranking
- Performance optimization, compression, and quantization
- Scaling, sharding/replication, and security
@@ -67,9 +67,9 @@ Build the vector search skills that matter: hybrid retrieval, multivector rerank
### The Path
-**Days 0–2**: Foundations. Connect to Qdrant Cloud, work with points and payloads, compute semantic similarity, chunk text, and tune HNSW for speed and recall.
+**Days 0-2**: Foundations. Connect to Qdrant Cloud, work with points and payloads, compute semantic similarity, chunk text, and tune HNSW for speed and recall.
-**Days 3–5**: Advanced retrieval. Combine dense and sparse signals, do hybrid search with server-side fusion, use multivectors (ColBERT) with the Universal Query API, and build recommendations.
+**Days 3-5**: Advanced retrieval. Combine dense and sparse signals, do hybrid search with server-side fusion, use multivectors (ColBERT) with the Universal Query API, and build recommendations.
**Day 6**: Ship. Wire ingestion, hybrid retrieval, multivector re-ranking, and evaluation (Recall@10, MRR, latency P50/P95).
@@ -185,11 +185,11 @@ ML, backend, data, and search engineers building RAG, semantic search, or recomm
## Time commitment
-- Duration: 7 days at 1–2 hours/day + optional bonus day
+- Duration: 6 days at 1-2 hours/day + 1 optional bonus day
- Video learning: ~3 hours
- Hands-on learning: 4-5 hours
-- Final project: 2–4 hours
-- Total: 9–12 hours
+- Final project: 2-4 hours
+- Total: 9-12 hours
{{< course-card
diff --git a/qdrant-landing/content/course/essentials/certification/_index.md b/qdrant-landing/content/course/essentials/certification/_index.md
index bb3b92a61..e34f2cf9a 100644
--- a/qdrant-landing/content/course/essentials/certification/_index.md
+++ b/qdrant-landing/content/course/essentials/certification/_index.md
@@ -5,4 +5,4 @@ weight: 100
# Qdrant Essentials Certification
-Coming soon!
\ No newline at end of file
+Coming soon! [Click here](https://forms.gle/QPSfdMjs3QpUCtGT9) to be notified when certifications become available.
\ No newline at end of file
diff --git a/qdrant-landing/content/course/essentials/day-0/pitstop-project.md b/qdrant-landing/content/course/essentials/day-0/pitstop-project.md
index 334a6afee..d53a5d43a 100644
--- a/qdrant-landing/content/course/essentials/day-0/pitstop-project.md
+++ b/qdrant-landing/content/course/essentials/day-0/pitstop-project.md
@@ -40,7 +40,7 @@ Before creating data, decide what each of the four dimensions in your vectors wi
**Example Ideas:**
- **Product categories**: Create vectors where each dimension represents a feature (affordability, quality, popularity, innovation). Electronics might be `[0.8, 0.7, 0.9, 0.6]`, while books could be `[0.3, 0.9, 0.4, 0.8]`.
-- **Color palettes**: Each dimension represents color intensity (red, green, blue, brightness). Bright red: `[0.9, 0.1, 0.1, 0.8]`, forest green: `[0.1, 0.8, 0.2, 0.5]`.
+- **Color palettes**: Each dimension represents color (red, green, blue). Bright red: `[0.9, 0.1, 0.1]`, forest green: `[0.1, 0.8, 0.2]`.
- **Data types**: Dimensions for structure, size, complexity, frequency. Spreadsheets: `[0.9, 0.6, 0.3, 0.7]`, images: `[0.2, 0.8, 0.5, 0.4]`.
- **Movie genres**: Action, drama, comedy, sci-fi intensities. Action thriller: `[0.9, 0.3, 0.1, 0.7]`, romantic comedy: `[0.1, 0.6, 0.9, 0.2]`.
@@ -148,7 +148,7 @@ For a new concept (not the Product Categories concept) run the code above and do
### Step 2: Post Your Results
-Show what you built and compare notes with others. **Post your results in**
+Show what you built and compare notes with others. **Post your results in**
diff --git a/qdrant-landing/content/course/essentials/day-1/pitstop-project.md b/qdrant-landing/content/course/essentials/day-1/pitstop-project.md
index 0e097fd6b..52832ca23 100644
--- a/qdrant-landing/content/course/essentials/day-1/pitstop-project.md
+++ b/qdrant-landing/content/course/essentials/day-1/pitstop-project.md
@@ -291,7 +291,7 @@ Now it's time to analyze your results and share what you've learned. Follow thes
### Step 2: Post Your Results
-**Post your results in**
+**Post your results in**
diff --git a/qdrant-landing/content/course/essentials/day-2/collection-tuning-demo.md b/qdrant-landing/content/course/essentials/day-2/collection-tuning-demo.md
index 96ea44aac..d988b496f 100644
--- a/qdrant-landing/content/course/essentials/day-2/collection-tuning-demo.md
+++ b/qdrant-landing/content/course/essentials/day-2/collection-tuning-demo.md
@@ -117,10 +117,10 @@ print(f"Dataset size: {len(ds['train'])} articles")
# Explore the dataset structure
print("\nDataset structure:")
-print("Available columns:", ds['train'].column_names)
+print("Available columns:", ds["train"].column_names)
# Look at a sample entry
-sample = ds['train'][0]
+sample = ds["train"][0]
print(f"\nSample article:")
print(f"Title: {sample['title']}")
print(f"Text preview: {sample['text'][:200]}...")
@@ -132,7 +132,7 @@ print(f"Embedding dimensions: {len(sample['text-embedding-3-large-1536-embedding
- **Content**: Pre-computed Wikipedia article embeddings
- **Size**: 100,000 articles
- **Embeddings**: with OpenAI's `text-embedding-3-large` truncated to 1536 dims
-- **Metadata**: `_id`, `titles` and `text`
+- **Metadata**: `_id`, `title` and `text`
## Step 4: Strategic Collection Creation
@@ -156,18 +156,20 @@ print(f"Creating collection: {collection_name}")
client.create_collection(
collection_name=collection_name,
vectors_config=models.VectorParams(
- size=1536, # Matches dataset dims
- distance=models.Distance.COSINE # Good for normalized embeddings
+ size=1536, # Matches dataset dims
+ distance=models.Distance.COSINE, # Good for normalized embeddings
),
hnsw_config=models.HnswConfigDiff(
- m=0, # Skip links during upload for speed
- ef_construct=100, # Used after we set m>0
- full_scan_threshold=10000
+ m=0, # Bulk load fast: m=0 (build links after ingest).
+ ef_construct=100, # Build quality: used after we set m>0
+ full_scan_threshold=10, # force HNSW instead of full scan
),
+ optimizers_config=models.OptimizersConfigDiff(
+ indexing_threshold=10
+ ), # Force indexing even on small sets for demo
strict_mode_config=models.StrictModeConfig(
- enabled=False, # More flexible while testing
- unindexed_filtering_retrieve=True # Allow filters without payload indexes
- )
+ enabled=False,
+ ), # More flexible while testing
)
print(f"Collection '{collection_name}' created successfully!")
@@ -182,11 +184,11 @@ print(f"HNSW m: {collection_info.config.hnsw_config.m}")
**Configuration details:**
- **`size=1536`**: To match the dimensions parameter we set for the OpenAI `text-embedding-3-large`
-- **`distance=COSINE`**: Standart for normalized embeddings and semantic similarity
+- **`distance=COSINE`**: Standard for normalized embeddings and semantic similarity
- **`full_scan_threshold=10000`**: Uses exact search for smaller result sets
- **`strict_mode_config`**: Managed Cloud runs in strict mode by default. We set `enabled=False` to let you experiment with unindexed payload keys during the demo.
-**Side note:** `text-embedding-3-large` outputs 3072 dims. Trunkating that to only 1536 dimensions cuts compute and memory, with some accuracy loss.
+**Side note:** `text-embedding-3-large` outputs 3072 dims. Truncating that to only 1536 dimensions cuts compute and memory, with some accuracy loss.
## Step 5: Bulk Upload with Rich Payloads
@@ -218,7 +220,7 @@ def upload_batch(start_idx, end_idx):
return 0
-batch_size = 10000
+batch_size = 64 * 10
total_points = len(ds["train"])
print(f"Uploading {total_points} points in batches of {batch_size}")
@@ -239,8 +241,8 @@ Now switch from `m=0` to `m=16` to build HNSW connections and improve search tim
client.update_collection(
collection_name=collection_name,
hnsw_config=models.HnswConfigDiff(
- m=16 # Each node connects to 16 neighbors
- )
+ m=16 # Build HNSW now: m=16 after the bulk load.
+ ),
)
print("HNSW indexing enabled with m=16")
@@ -312,17 +314,14 @@ Let's measure search performance on the HNSW‑enabled collection.
print("Running baseline performance test...")
# Warm up the RAM index/vectors cache with a test query
-print("Warming up caches...")
client.query_points(collection_name=collection_name, query=query_embedding, limit=1)
# Measure vector search performance
search_times = []
-for _ in range(3): # Multiple runs for a stable average
+for _ in range(25): # Multiple runs for a stable average
start_time = time.time()
response = client.query_points(
- collection_name=collection_name,
- query=query_embedding,
- limit=10
+ collection_name=collection_name, query=query_embedding, limit=10
)
search_time = (time.time() - start_time) * 1000
search_times.append(search_time)
@@ -332,14 +331,16 @@ baseline_time = sum(search_times) / len(search_times)
print(f"Average search time: {baseline_time:.2f}ms")
print(f"Search times: {[f'{t:.2f}ms' for t in search_times]}")
print(f"Found {len(response.points)} results")
-print(f"Top result: '{response.points[0].payload['title']}' (score: {response.points[0].score:.4f})")
+print(
+ f"Top result: '{response.points[0].payload['title']}' (score: {response.points[0].score:.4f})"
+)
# Show a few more results for context
print(f"\nTop 3 results:")
for i, point in enumerate(response.points[:3], 1):
- title = point.payload['title']
+ title = point.payload["title"]
score = point.score
- text_preview = point.payload['text'][:100] + "..."
+ text_preview = point.payload["text"][:100] + "..."
print(f" {i}. {title} (score: {score:.4f})")
print(f" {text_preview}")
```
@@ -347,7 +348,7 @@ for i, point in enumerate(response.points[:3], 1):
**Performance factors:**
- **Cache warming**: First query loads relevant index parts/vectors into memory, subsequent queries are faster
- **HNSW with m=16**: Graph-based search is much faster than full scan
-- **MRepeated runs**: Average of several queries gives more reliable timing results
+- **Repeated runs**: Average of several queries gives more reliable timing results
## Step 9: Filtering Without Payload Indexes
@@ -356,26 +357,31 @@ Now, let's test filtering performance without indexes. This forces Qdrant to sca
```python
print("Testing filtering without payload indexes")
+# Warning: We enable unindexed_filtering_retrieve only for demonstration purposes. In production, don’t use it.
+# Demo only: allow filtering without an index by scanning. Turn this off later.
+client.update_collection(
+ collection_name=collection_name,
+ strict_mode_config=models.StrictModeConfig(unindexed_filtering_retrieve=True),
+)
+
# Create a text-based filter
text_filter = models.Filter(
- must=[
- models.FieldCondition(
- key="text",
- match=models.MatchText(text="data")
- )
- ]
+ must=[models.FieldCondition(key="text", match=models.MatchText(text="data"))]
)
+# Warmup
+client.query_points(collection_name=collection_name, query=query_embedding, limit=1)
+
# Run multiple times for more reliable measurement
unindexed_times = []
-for i in range(3):
+for i in range(25):
start_time = time.time()
response = client.query_points(
collection_name=collection_name,
query=query_embedding,
limit=10,
search_params=models.SearchParams(hnsw_ef=100),
- query_filter=text_filter
+ query_filter=text_filter,
)
unindexed_times.append((time.time() - start_time) * 1000)
@@ -386,7 +392,9 @@ print(f"Individual times: {[f'{t:.2f}ms' for t in unindexed_times]}")
print(f"Overhead vs baseline: {unindexed_filter_time - baseline_time:.2f}ms")
print(f"Found {len(response.points)} matching results")
if response.points:
- print(f"Top result: '{response.points[0].payload['text']}'\nScore: {response.points[0].score:.4f}")
+ print(
+ f"Top result: '{response.points[0].payload['text']}'\nScore: {response.points[0].score:.4f}"
+ )
else:
print("No results found - try a different filter term")
```
@@ -396,25 +404,25 @@ else:
Create a [full‑text index](/documentation/concepts/indexing/#full-text-index) for faster filtering.
```python
+# Create a payload index for 'text' so filters use an index, not a scan.
client.create_payload_index(
collection_name=collection_name,
field_name="text",
wait=True,
field_schema=models.TextIndexParams(
- type="text",
- tokenizer="word",
- phrase_matching=False
- )
- )
+ type="text", tokenizer="word", phrase_matching=False
+ ),
+)
+
+client.update_collection(
+ collection_name=collection_name,
+ hnsw_config=models.HnswConfigDiff(
+ ef_construct=101
+ ), # Added payload index after HNSW; bump ef_construct (+1) to rebuild with filter data.
+ strict_mode_config=models.StrictModeConfig(unindexed_filtering_retrieve=False),
+)
print("Payload index created for 'text' field")
-
-# If you want filter‑aware HNSW and you built the graph before creating payload indexes,
-# rebuild the graph to attach filter data structures.
-# Note: Reindexing takes up a lot of resources, and it is advised to set payload
-# indexes only once, before building HNSW.
-# client.update_collection(collection_name=collection_name, hnsw_config=models.HnswConfigDiff(m=0))
-# client.update_collection(collection_name=collection_name, hnsw_config=models.HnswConfigDiff(m=16))
```
## Step 11: Filtering With Payload Indexes
@@ -424,16 +432,20 @@ Run the same query with the index in place.
```python
print("Testing filtering WITH payload indexes...")
+
+# Warmup
+client.query_points(collection_name=collection_name, query=query_embedding, limit=1)
+
# Run multiple times for more reliable measurement
indexed_times = []
-for i in range(3):
+for i in range(25):
start_time = time.time()
response = client.query_points(
collection_name=collection_name,
query=query_embedding,
limit=10,
search_params=models.SearchParams(hnsw_ef=100),
- query_filter=text_filter
+ query_filter=text_filter,
)
indexed_times.append((time.time() - start_time) * 1000)
@@ -444,7 +456,9 @@ print(f"Individual times: {[f'{t:.2f}ms' for t in indexed_times]}")
print(f"Overhead vs baseline: {indexed_filter_time - baseline_time:.2f}ms")
print(f"Found {len(response.points)} matching results")
if response.points:
- print(f"Top result: '{response.points[0].payload['text']}'\nScore: {response.points[0].score:.4f}")
+ print(
+ f"Top result: '{response.points[0].payload['text']}'\nScore: {response.points[0].score:.4f}"
+ )
else:
print("No results found - try a different filter term")
```
@@ -454,9 +468,9 @@ else:
Compare your results and see the effect of each optimization:
```python
-print("\n" + "="*60)
+print("\n" + "=" * 60)
print("FINAL PERFORMANCE SUMMARY")
-print("="*60)
+print("=" * 60)
# Key metrics
if unindexed_filter_time > 0 and indexed_filter_time > 0:
@@ -481,7 +495,7 @@ print(f"Key insights:")
print(f" • HNSW (m=16) enables fast vector search")
print(f" • Payload indexes dramatically improve filtering")
print(f" • Upload strategy (m=0→m=16) optimizes ingestion")
-print("="*60)
+print("=" * 60)
```
diff --git a/qdrant-landing/content/course/essentials/day-2/pitstop-project.md b/qdrant-landing/content/course/essentials/day-2/pitstop-project.md
index 3698c385e..7b03011de 100644
--- a/qdrant-landing/content/course/essentials/day-2/pitstop-project.md
+++ b/qdrant-landing/content/course/essentials/day-2/pitstop-project.md
@@ -72,10 +72,10 @@ Test different HNSW configurations to find what works best:
```python
# Test configurations
configs = [
- {"name": "fast_initial_upload", "m": 0, "ef_construct": 100},
- {"name": "memory_optimized", "m": 8, "ef_construct": 100},
- {"name": "balanced", "m": 16, "ef_construct": 200},
- {"name": "high_quality", "m": 32, "ef_construct": 400},
+ {"name": "fast_initial_upload", "m": 0, "ef_construct": 100}, # m=0 = ingest-only
+ {"name": "memory_optimized", "m": 8, "ef_construct": 100}, # m=8 = lower RAM
+ {"name": "balanced", "m": 16, "ef_construct": 200}, # m=16 = balanced
+ {"name": "high_quality", "m": 32, "ef_construct": 400}, # m=32 = higher recall, slower build
]
for config in configs:
@@ -87,12 +87,13 @@ for config in configs:
collection_name=collection_name,
vectors_config=models.VectorParams(size=384, distance=models.Distance.COSINE),
hnsw_config=models.HnswConfigDiff(
- m=config["m"], ef_construct=config["ef_construct"], full_scan_threshold=10
- ),
- optimizers_config=models.OptimizersConfigDiff(indexing_threshold=0),
- strict_mode_config=models.StrictModeConfig(
- unindexed_filtering_retrieve=True, unindexed_filtering_update=True
+ m=config["m"],
+ ef_construct=config["ef_construct"],
+ full_scan_threshold=10, # force HNSW instead of full scan
),
+ optimizers_config=models.OptimizersConfigDiff(
+ indexing_threshold=10
+ ), # Force indexing even on small sets for demo
)
print(f"Created collection: {collection_name}")
```
@@ -103,7 +104,8 @@ Measure upload performance for each configuration:
```python
def upload_with_timing(collection_name, data, config_name):
- embeddings = [encoder.encode(dat["description"]).tolist() for dat in data]
+ embeddings = encoder.encode([d["description"] for d in data], show_progress_bar=True).tolist()
+
points = []
for i, item in enumerate(data):
embedding = embeddings[i]
@@ -117,13 +119,15 @@ def upload_with_timing(collection_name, data, config_name):
"length": len(item["description"]),
"word_count": len(item["description"].split()),
"has_keywords": any(
- keyword in item["description"].lower()
- for keyword in ["important", "key", "main"]
+ keyword in item["description"].lower() for keyword in ["important", "key", "main"]
),
},
)
)
+ # Warmup
+ client.query_points(collection_name=collection_name, query=points[0].vector, limit=1)
+
start_time = time.time()
client.upload_points(collection_name=collection_name, points=points)
upload_time = time.time() - start_time
@@ -132,35 +136,43 @@ def upload_with_timing(collection_name, data, config_name):
return upload_time
-# Load your dataset here
+# Load your dataset here. The larger the dataset, the more accurate the benchmark will be.
# your_dataset = [{"description": "This is a description of a product"}, ...]
# Upload to each collection
upload_times = {}
for config in configs:
collection_name = f"my_domain_{config['name']}"
- upload_times[config["name"]] = upload_with_timing(
- collection_name, your_dataset, config["name"]
- )
+ upload_times[config["name"]] = upload_with_timing(collection_name, your_dataset, config["name"])
-# Wait for index to be built
-def wait_for_index_built(collection_name, vectors_per_point=1):
- info = client.get_collection(collection_name=collection_name)
- count = 0
- while info.points_count * vectors_per_point - info.indexed_vectors_count != 0 and count < 10:
- time.sleep(1)
+
+def wait_for_indexing(collection_name, timeout=60, poll_interval=1):
+ print(f"Waiting for collection '{collection_name}' to be indexed...")
+ start_time = time.time()
+
+ while time.time() - start_time < timeout:
info = client.get_collection(collection_name=collection_name)
- count += 1
- if count == 10:
- raise Exception(
- f"Indexed vectors count ({info.indexed_vectors_count}) is not equal to points count ({info.points_count}). Upload enough points to trigger index rebuild."
- )
+
+ if info.indexed_vectors_count > 0 and info.status == models.CollectionStatus.GREEN:
+ print(f"Success! Collection '{collection_name}' is indexed and ready.")
+ print(f" - Status: {info.status.value}")
+ print(f" - Indexed vectors: {info.indexed_vectors_count}")
+ return
+
+ print(f" - Status: {info.status.value}, Indexed vectors: {info.indexed_vectors_count}. Waiting...")
+ time.sleep(poll_interval)
+
+ info = client.get_collection(collection_name=collection_name)
+ raise Exception(
+ f"Timeout reached after {timeout} seconds. Collection '{collection_name}' is not ready. "
+ f"Final status: {info.status.value}, Indexed vectors: {info.indexed_vectors_count}"
+ )
for config in configs:
- collection_name = f"my_domain_{config['name']}"
- wait_for_index_built(collection_name)
-
+ if config["m"] > 0: # m=0 has no HNSW to wait for
+ collection_name = f"my_domain_{config['name']}"
+ wait_for_indexing(collection_name)
```
### Step 4: Benchmark Search Performance
@@ -170,19 +182,15 @@ Test search speed with different `hnsw_ef` values:
```python
def benchmark_search(collection_name, query_embedding, ef_values=[64, 128, 256]):
# Warmup
- _ = client.query_points(
- collection_name=collection_name,
- query=query_embedding,
- limit=10,
- search_params=models.SearchParams(hnsw_ef=ef_values[0]),
- )
+ client.query_points(collection_name=collection_name, query=query_embedding, limit=1)
+ # hnsw_ef: higher = better recall, but slower. Tune per your latency goal.
results = {}
for hnsw_ef in ef_values:
times = []
# Run multiple queries for more reliable timing
- for _ in range(5):
+ for _ in range(25):
start_time = time.time()
_ = client.query_points(
@@ -190,6 +198,7 @@ def benchmark_search(collection_name, query_embedding, ef_values=[64, 128, 256])
query=query_embedding,
limit=10,
search_params=models.SearchParams(hnsw_ef=hnsw_ef),
+ with_payload=False,
)
times.append((time.time() - start_time) * 1000)
@@ -225,57 +234,73 @@ def test_filtering_performance(collection_name):
# Test filter without index
filter_condition = models.Filter(
- must=[models.FieldCondition(key="length", range=models.Range(gte=100, lte=500))]
+ must=[models.FieldCondition(key="length", range=models.Range(gte=10, lte=200))]
)
- # Timing without payload index
- start_time = time.time()
- _ = client.query_points(
+ # Demo only: unindexed_filtering_retrieve=True forces a scan; turn it off right after measuring.
+ client.update_collection(
collection_name=collection_name,
- query=query_embedding,
- query_filter=filter_condition,
- limit=10,
+ strict_mode_config=models.StrictModeConfig(unindexed_filtering_retrieve=True),
)
- time_without_index = (time.time() - start_time) * 1000
+
+ # Warmup
+ client.query_points(collection_name=collection_name, query=query_embedding, limit=1)
+
+ # Timing without payload index
+ times = []
+ for _ in range(25):
+ start_time = time.time()
+ _ = client.query_points(
+ collection_name=collection_name,
+ query=query_embedding,
+ query_filter=filter_condition,
+ limit=10,
+ with_payload=False,
+ )
+ times.append((time.time() - start_time) * 1000)
+ time_without_index = np.mean(times)
# Create payload index
client.create_payload_index(
- collection_name=collection_name, field_name="length", field_schema="integer"
+ collection_name=collection_name,
+ field_name="length",
+ field_schema=models.PayloadSchemaType.INTEGER,
+ wait=True,
)
- # Rebuild HNSW to attach filter data structures.
- # Note: This is not advised for production. Better create payload index before uploading any data to avoid rebuild.
- suffix = collection_name.replace("my_domain_", "")
- config = next((c for c in configs if c["name"] == suffix), None)
-
- client.update_collection(
- collection_name=collection_name, hnsw_config=models.HnswConfigDiff(m=0)
- )
+ # HNSW was already built; adding the payload index doesn’t rebuild it.
+ # Bump ef_construct (+1) once to trigger a safe rebuild.
+ base_ef = client.get_collection(
+ collection_name=collection_name
+ ).config.hnsw_config.ef_construct
+ new_ef_construct = base_ef + 1
client.update_collection(
collection_name=collection_name,
- hnsw_config=models.HnswConfigDiff(
- m=16,
- ef_construct=config["ef_construct"],
- full_scan_threshold=10,
- payload_m=None,
- max_indexing_threads=1,
- ),
- optimizers_config=models.OptimizersConfigDiff(vacuum_min_vector_number=0),
+ hnsw_config=models.HnswConfigDiff(ef_construct=new_ef_construct),
+ strict_mode_config=models.StrictModeConfig(
+ unindexed_filtering_retrieve=False
+ ), # Turn off scanning and use payload index instead.
)
- # Wait for index to be built
- wait_for_index_built(collection_name)
+ wait_for_indexing(collection_name)
+
+ # Warmup
+ client.query_points(collection_name=collection_name, query=query_embedding, limit=1)
# Timing with index
- start_time = time.time()
- _ = client.query_points(
- collection_name=collection_name,
- query=query_embedding,
- query_filter=filter_condition,
- limit=10,
- )
- time_with_index = (time.time() - start_time) * 1000
+ times = []
+ for _ in range(25):
+ start_time = time.time()
+ _ = client.query_points(
+ collection_name=collection_name,
+ query=query_embedding,
+ query_filter=filter_condition,
+ limit=10,
+ with_payload=False,
+ )
+ times.append((time.time() - start_time) * 1000)
+ time_with_index = np.mean(times)
return {
"without_index": time_without_index,
@@ -334,7 +359,7 @@ You'll know you've succeeded when:
### Step 2: Post Your Results
-**Post your results in**
**using this:**
diff --git a/qdrant-landing/content/course/essentials/day-2/what-is-hnsw.md b/qdrant-landing/content/course/essentials/day-2/what-is-hnsw.md
index 452065e6b..d04991bad 100644
--- a/qdrant-landing/content/course/essentials/day-2/what-is-hnsw.md
+++ b/qdrant-landing/content/course/essentials/day-2/what-is-hnsw.md
@@ -274,7 +274,7 @@ performance = benchmark_search_performance(collection_name, test_queries, ef_val
Use [`get_collection`](/api-reference/collections/get-collection) to inspect your collection. It returns Current statistics and configuration of the collection like `points_count`, `indexed_vectors_count` or `hnsw_config`. It also lists `payload_schema` for payload indexes you created.
-To see whether your data is actually indexed check vector and point counts: if `indexed_vectors_count` is far below `points_count * vectors_per_point`, a large part of your data is not in HNSW yet.
+To see whether your data is actually indexed, you need to check two things: the number of indexed vectors and the collection's status. If `indexed_vectors_count` is low, indexing may not have completed. More importantly, you should check the collection `status`. A `YELLOW` status means optimization (indexing) is still in progress, while a `GREEN` status confirms it is complete and ready for optimal performance.
If queries feel slow check:
- whether filter fields have [payload indexes](/documentation/concepts/indexing/#payload-index).
@@ -290,9 +290,9 @@ info = client.get_collection(collection_name)
vectors_per_point = 1 # set per your vectors_config
vectors_count = info.points_count * vectors_per_point
-print(f"Total vectors: {vectors_count}")
+print(f"Collection status: {info.status}")
+print(f"Total points: {info.points_count}")
print(f"Indexed vectors: {info.indexed_vectors_count}")
-print(f"HNSW config: {info.config.hnsw_config}")
if vectors_count:
proportion_unindexed = 1 - (info.indexed_vectors_count / vectors_count)
@@ -300,6 +300,13 @@ else:
proportion_unindexed = 0
print(f"Proportion unindexed: {proportion_unindexed:.2%}")
+
+if info.status == models.CollectionStatus.GREEN:
+ print("\n✅ Collection is indexed and ready!")
+elif info.status == models.CollectionStatus.YELLOW:
+ print("\n⚠️ Collection is still being indexed (optimizing).")
+else:
+ print(f"\n❌ Collection status is {info.status}.")
```
## When Not to Use HNSW
diff --git a/qdrant-landing/content/course/essentials/day-4/large-scale-ingestion.md b/qdrant-landing/content/course/essentials/day-4/large-scale-ingestion.md
index 872c20f92..1de4e01c3 100644
--- a/qdrant-landing/content/course/essentials/day-4/large-scale-ingestion.md
+++ b/qdrant-landing/content/course/essentials/day-4/large-scale-ingestion.md
@@ -77,7 +77,7 @@ client.recreate_collection(
max_segment_size=5_000_000, # Create larger segments for faster search
),
hnsw_config=models.HnswConfigDiff(
- m=6, # Lower M to reduce memory usage
+ m=6, # Lower m to reduce memory usage
on_disk=False # Keep the HNSW index graph in RAM
),
)
diff --git a/qdrant-landing/content/course/essentials/day-5/pitstop-project.md b/qdrant-landing/content/course/essentials/day-5/pitstop-project.md
index 661e02212..007893f1e 100644
--- a/qdrant-landing/content/course/essentials/day-5/pitstop-project.md
+++ b/qdrant-landing/content/course/essentials/day-5/pitstop-project.md
@@ -525,7 +525,7 @@ You'll know you've succeeded when:
### Step 2: Post Your Results
-**Post your results in**
**using this:**
diff --git a/qdrant-landing/content/course/essentials/day-6/congratulations.md b/qdrant-landing/content/course/essentials/day-6/congratulations.md
index e23e7d6d8..6aed36bf1 100644
--- a/qdrant-landing/content/course/essentials/day-6/congratulations.md
+++ b/qdrant-landing/content/course/essentials/day-6/congratulations.md
@@ -15,7 +15,7 @@ You've built and shipped a complete vector search application and gained the exp
You've progressed from vector search fundamentals to production-ready expertise:
-**Foundation Building** (Days 0-2): You mastered the core concepts of vector search, learned how similarity metrics work, and understood how [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) indexing enables fast retrieval at scale.
+**Foundation Building** (Days 0-2): You mastered the core concepts of vector search, learned how similarity metrics work, and understood how HNSW indexing enables fast retrieval at scale.
**Advanced Retrieval** (Days 3-5): You implemented hybrid search combining semantic and keyword signals, explored quantization for performance optimization, and mastered the Universal Query API with multivector reranking.
@@ -40,13 +40,13 @@ Your final project demonstrates several production-critical capabilities:
{{< course-card
title="Earn your Qdrant Essentials Certificate"
image="/icons/outline/training-white.svg"
- link="/course/certification/" >}}
+ link="/course/essentials/certification/" >}}
Get recognized for completing Day 0–6 and the final project. Add it to your LinkedIn and portfolio.
{{< /course-card >}}
## What's Next?
-**Explore Advanced Integrations**: Check out [Day 9 Partner Integrations](../../day-9/) to see how Qdrant works with leading AI frameworks and data platforms.
+**Explore Advanced Integrations**: Check out [Day 7 Partner Integrations](../../day-7/) to see how Qdrant works with leading AI frameworks and data platforms.
**Join the Community**: Share your final project results and connect with other practitioners building vector search systems. The Qdrant community is always excited to see what people build.
diff --git a/qdrant-landing/content/course/essentials/day-6/final-project.md b/qdrant-landing/content/course/essentials/day-6/final-project.md
index a58256ebe..3705ca0b7 100644
--- a/qdrant-landing/content/course/essentials/day-6/final-project.md
+++ b/qdrant-landing/content/course/essentials/day-6/final-project.md
@@ -144,7 +144,7 @@ Transform raw results into user-friendly output: page title, section title, URLs
### Step 7: Analyze Your Results
-Build a small eval set and measure quality and latency. Use results to guide tuning (fusion strategy, candidate sizes, search-time `ef`, etc.).
+Build a small eval set and measure quality and latency. Use results to guide tuning (fusion strategy, candidate sizes, search-time `hnsw_ef`, etc.).
**Ground Truth**:
- Create 20–30 realistic queries with expected section URLs/anchors.
@@ -199,7 +199,7 @@ As you test your search engine, consider:
* **Rerank or not:** Are multivectors worth it, or is fusion alone enough?
* **Performance tuning:** Which search and HNSW settings hit your accuracy/latency goals?
- * Search time: raise `ef` from 64 → 128 → 256 until gains flatten.
+ * Search time: raise `hnsw_ef` from 64 → 128 → 256 until gains flatten.
* Index time (if rebuilding): try higher `m` (16, 32) and `ef_construct` (200, 400).
### Step 2: Post Your Results
@@ -228,7 +228,7 @@ Show your run and learn from others. **Post your results in** "
diff --git a/qdrant-landing/content/course/essentials/day-7/_index.md b/qdrant-landing/content/course/essentials/day-7/_index.md
index 958ee81c7..abad9e51d 100644
--- a/qdrant-landing/content/course/essentials/day-7/_index.md
+++ b/qdrant-landing/content/course/essentials/day-7/_index.md
@@ -23,9 +23,9 @@ Learn about the Qdrant ecosystem and integration strategies.
## Choose Your Integration
{{< cards-list >}}
-- icon: /courses/course-integrations/haystack.svg
+- icon: /courses/course-integrations/haystack.png
title: Haystack
- content: Build end-to-end NLP pipelines with Qdrant
+ content: Build end-to-end agentic pipelines with Qdrant
link: haystack/
- icon: /courses/course-integrations/tensorlake.svg
diff --git a/qdrant-landing/content/course/essentials/day-7/haystack.md b/qdrant-landing/content/course/essentials/day-7/haystack.md
index d643039c2..eecd844cc 100644
--- a/qdrant-landing/content/course/essentials/day-7/haystack.md
+++ b/qdrant-landing/content/course/essentials/day-7/haystack.md
@@ -7,7 +7,7 @@ weight: 2
# Integrating with Haystack
-Build end-to-end NLP pipelines with Haystack and Qdrant.
+Build end-to-end agentic pipelines with Qdrant.
{{< youtube "lMinhPZufTc" >}}
diff --git a/qdrant-landing/static/courses/course-integrations/haystack.png b/qdrant-landing/static/courses/course-integrations/haystack.png
new file mode 100644
index 000000000..4bc4d339d
Binary files /dev/null and b/qdrant-landing/static/courses/course-integrations/haystack.png differ