docs: Add course content for Day 4 (#1931)

* docs: Add course content for Day 4

* improved strucutre

* fixed what-is-quantization

* fixed rescoring-oversampling-indexing

* fixed pitstop project

* standart init for client

* pitstop Reflect on Your Findings

* pitstop fix

* pitstop fix

* add new video and md for day 3

* Update Discord link for posting results

---------

Co-authored-by: Kirstin <kirstin.taufertshoefer@qdrant.com>
Co-authored-by: Evgeniya Sukhodolskaya <suxodolskaya97@gmail.com>
Co-authored-by: Thierry Damiba <thierrydamiba@gmail.com>
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
This commit is contained in:
Kirstin
2025-10-22 13:26:10 +02:00
committed by GitHub
co-authored by Kirstin Evgeniya Sukhodolskaya Thierry Damiba Andrey Vasnetsov
parent ab884c16d5
commit bcaa46c8ef
5 changed files with 858 additions and 3 deletions
@@ -1,9 +1,23 @@
---
title: Day 4
title: "Day 4: Optimization and Scale"
isLesson: true
weight: 8
weight: 5
---
{{< date >}} Day 4 {{< /date >}}
# Day 4
# Optimization and Scale
Compression, advanced tuning, and high‑throughput ingestion.
---
## Today’s path
1. Vector Quantization Methods
2. Accuracy Recovery with Rescoring
3. High-Throughput Data Ingestion
4. Project: Quantization Performance Optimization
You’ll compare memory footprint, recall, and throughput across configurations.
@@ -0,0 +1,136 @@
---
title: Large-Scale Data Ingestion
weight: 3
---
{{< date >}} Day 4 {{< /date >}}
# Large-Scale Data Ingestion
<div class="video">
<iframe
src="https://www.youtube.com/embed/Rawvm7TP1XI"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
allowfullscreen>
</iframe>
</div>
<br/>
In vector search applications inserting a few thousand data points is straightforward but the dynamics change completely when dealing with millions or billions of records. Tiny inefficiencies in the ingestion process compound into significant time losses, increased memory pressure, and degraded search performance.
Every individual upsert call initiates a transaction that consumes memory and disk I/O to build parts of the index. At scale, this naive approach can overwhelm your system, causing upload times to spike and search quality to decrease. Efficiently preparing and loading your data into Qdrant is paramount for building a robust and scalable AI application.
## Choosing Your Ingestion Strategy
Qdrant provides several methods for data ingestion, each tailored to different scales and use cases. It should be noted that only the Python client supports the upload_points and upload_collection methods. If you're using Qdrant on a different client then we reccomend using upsert with batch upload for large scale ingestion. [Learn more about bulk operations](/documentation/guides/bulk-operations/).
- **upsert (Individual or Batched)**: This is the fundamental operation for adding or updating points. Individual upserts are best suited for real-time updates while batching works best for larger workloads.
- **upload_points**: This method is optimized for uploading an entire batch of points that can comfortably fit into your client's memory. It leverages features like lazy batching, retries, and parallelism, making it a strong choice for medium-sized datasets.
- **upload_collection**: For truly large-scale datasets, upload_collection is the most powerful tool. It streams data directly from an iterator, meaning the entire dataset doesn't need to be loaded into memory at once. This memory-efficient approach is ideal for ingesting millions or billions of points.
> **<font color='red'>Note:</font>** For clients in other languages like **TypeScript**, **Rust**, and **Go**, batched upsert calls are the recommended method for efficient data loading.
## Heuristics for Scale
Deciding which method to use can be guided by a few simple rules of thumb. While every use case is different, these heuristics provide a solid starting point:
- **Less than 100,000 points**: A single-threaded, batched upsert operation will generally perform well.
- **100,000 to 1 million points**: `upload_points` is recommended, using batch sizes between 1,000 and 10,000 to balance network overhead and memory usage.
- **More than 1 million points**: `upload_collection` is the ideal choice for streaming data from disk. To maximize throughput, you should enable parallelism by setting the `parallel` parameter to the number of available CPU cores (e.g., 4 or 8).
> **<font color='red'>Best Practice:</font>** Start small and test. Before attempting to upload your entire dataset, ingest a smaller chunk to validate your configuration and process.
## A Real-World Example: Ingesting LAION-400M
To illustrate these principles, let's examine the process of ingesting the LAION-400M dataset, which contains approximately 400 million image-text pairs with 512-dimensional CLIP embeddings. This massive dataset, with 400 GB of vectors and 200 GB of payload, requires a carefully optimized strategy.
### The Optimal Collection Configuration
The foundation of scalable ingestion is a well-designed collection configuration. For a dataset of this magnitude, the goal is to intelligently balance memory usage, disk I/O, and search performance.
```python
from qdrant_client import QdrantClient, models
import os
from dotenv import load_dotenv
load_dotenv()
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
client.recreate_collection(
collection_name="laion400m_collection",
vectors_config=models.VectorParams(
size=512, # CLIP embedding dimensions
distance=models.Distance.COSINE,
on_disk=True, # Store original vectors on disk
),
quantization_config=models.BinaryQuantization(
binary=models.BinaryQuantizationConfig(
always_ram=True, # Keep quantized vectors in RAM
)
),
optimizers_config=models.OptimizersConfigDiff(
max_segment_size=5_000_000, # Create larger segments for faster search
),
hnsw_config=models.HnswConfigDiff(
m=6, # Lower M to reduce memory usage
on_disk=False # Keep the HNSW index graph in RAM
),
)
```
This configuration employs several key optimizations:
- **`on_disk=True`**: This is the most critical setting for large datasets. It instructs Qdrant to store the full-precision original vectors on disk ([memmap storage](/documentation/guides/storage/#on-disk-storage)) instead of in RAM, dramatically reducing memory requirements.
- **Binary Quantization with `always_ram=True`**: While the original vectors are on disk, we enable [binary quantization](/documentation/guides/quantization/#binary-quantization) and force the compressed vectors to remain in RAM. This provides a lightweight in-memory representation for fast initial candidate searches.
- **Large Segment Size**: The `max_segment_size` is increased to create fewer, larger segments. This can improve search performance at the cost of slightly slower indexing.
- **In-Memory HNSW Index**: By setting `on_disk=False` for the [HNSW config](/documentation/guides/quantization/#hnsw-config), we keep the graph index in RAM. This ensures that navigating vector relationships during a search is extremely fast, avoiding disk latency. The `m` value is lowered to 6 to further conserve memory.
### The Upload Process
With the collection configured, the upload can proceed using a memory-efficient streaming approach. The LAION dataset is split into 409 parts, each containing about 1 million records. The script processes one part at a time, downloading the data, preparing the points, and streaming them to Qdrant.
```python
def upload_data_to_qdrant(client, embeddings, metadata, parallel=4):
"""
Uploads data to Qdrant using the upload_collection method.
"""
client.upload_collection(
collection_name="laion400m_collection",
points=zip(range(len(metadata)), embeddings, metadata),
batch_size=256,
parallel=parallel,
show_progress=True,
)
# --- Simplified logic for processing chunks ---
# for part in dataset_parts:
# embeddings, metadata = download_and_process_part(part)
# upload_data_to_qdrant(client, embeddings, metadata)
# cleanup_local_files(part)
```
This method processes the dataset in manageable chunks without ever loading the entire 400 million points into memory. Using `parallel=4` allows the client to upload multiple batches concurrently, saturating the network connection and maximizing ingestion speed.
## The Payoff: An Efficient Architecture at Scale
This combined strategy of a hybrid storage configuration and streaming ingestion creates a highly efficient system. By keeping only the most essential components in RAM: the quantized vectors and the HNSW index, Qdrant can index and serve a 400 million vector dataset on a machine with just 64GB of RAM. The original vectors, which would consume hundreds of gigabytes, are efficiently accessed from disk only when needed for rescoring top candidates.
This architecture strikes a balance by keeping infrastructure costs low by minimizing RAM usage while maintaining fast and accurate search performance. By understanding and applying these ingestion strategies, you can confidently scale your Qdrant-powered applications to handle real-world data volumes.
> Learn more in a complete hands-on guide in our **[Large-Scale Search tutorial](https://qdrant.tech/documentation/database-tutorials/large-scale-search/)**.
> **Check out the reference implementation:**
> [qdrant/laion-400m-benchmark on GitHub](https://github.com/qdrant/laion-400m-benchmark)
> This open-source repository includes full scripts for downloading, processing, and uploading the LAION-400M dataset to Qdrant using efficient, production-ready patterns.
> **Want to try this workflow hands-on?**
> Run the [Google Colab notebook](https://colab.research.google.com/drive/1X4EW-nymqcsyhwFYS2ZrmE8MPSnKKj62?usp=sharing) to see large-scale vector ingestion, quantized search, and efficient RAM/disk optimization in action!
@@ -0,0 +1,477 @@
---
title: "Project: Quantization Performance Optimization"
weight: 4
---
{{< date >}} Day 4 {{< /date >}}
# Project: Quantization Performance Optimization
Apply quantization techniques to your domain search engine and measure the real-world impact on speed, memory, and accuracy. You'll discover how different quantization methods affect your specific use case and learn to optimize the accuracy recovery pipeline.
## Your Mission
Transform your search engine from previous days into a production-ready system by implementing quantization optimization. You'll test different quantization methods, measure performance impacts, and tune the oversampling + rescoring pipeline for optimal results.
**Estimated Time:** 120 minutes
## What You'll Build
A quantization-optimized search system that demonstrates:
- **Performance comparison**: Before and after quantization metrics
- **Method evaluation**: Testing scalar and binary quantization on your data
- **Accuracy recovery**: Implementing oversampling and rescoring pipeline
- **Production deployment**: Memory-optimized storage configuration
### Prerequisites
* Qdrant Cloud cluster (URL + API key)
* Python 3.9+ (or Google Colab)
* Packages: `qdrant-client`, `numpy`, `python-dotenv`
### Models
* Use the same embedding model and dimension as your existing collection.
* If your vectors are **1536-dim**, keep `size=1536` below.
* Otherwise, change the `VectorParams(size=...)` to your model’s dim.
### Dataset
* Reuse your Day 1/2 domain dataset (ideally **1,000+** items) with a primary text field for embeddings.
* Include at least one numeric field (e.g., `length`, `word_count`) to measure payload index impact.
## Build Steps
### Step 1: Baseline Measurement
Start by measuring your current system's performance without quantization:
```python
import time
import numpy as np
from qdrant_client import QdrantClient, models
import os
from dotenv import load_dotenv
load_dotenv()
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
# For Colab:
# from google.colab import userdata
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
def measure_search_performance(collection_name, test_queries, label="Baseline"):
"""Measure search performance across multiple queries"""
latencies = []
# Don't forget to warm up caches!
#response = client.query_points(
# collection_name=collection_name,
# query=query,
# limit=10
# )
for query in test_queries:
start_time = time.time()
response = client.query_points(
collection_name=collection_name,
query=query,
limit=10
)
latency = (time.time() - start_time) * 1000
latencies.append(latency)
avg_latency = np.mean(latencies)
p95_latency = np.percentile(latencies, 95)
print(f"{label}:")
print(f" Average latency: {avg_latency:.2f}ms")
print(f" P95 latency: {p95_latency:.2f}ms")
print(f" Memory usage: Check Qdrant Cloud dashboard")
return {"avg": avg_latency, "p95": p95_latency}
# Measure baseline performance
baseline_metrics = measure_search_performance(
"your_domain_collection",
your_test_queries,
"Baseline (No Quantization)"
)
```
### Step 2: Test Quantization Methods
Create collections with different quantization methods to compare their impact:
> Note: When creating several collections for educational purposes with different quantization configurations (e.g., original, binary quantized, scalar quantized, 2-bit binary quantized), make sure to monitor available resources. The original vectors are stored for each collection (on disk in this case), in addition to their quantized versions.
```python
# Test configurations
quantization_configs = {
"scalar": {
"config": models.ScalarQuantization(
scalar=models.ScalarQuantizationConfig(
type=models.ScalarType.INT8,
quantile=0.99,
always_ram=True,
)
),
"expected_speedup": "2x",
"expected_compression": "4x"
},
"binary": {
"config": models.BinaryQuantization(
binary=models.BinaryQuantizationConfig(
encoding=models.BinaryQuantizationEncoding.ONE_BIT,
always_ram=True,
)
),
"expected_speedup": "40x",
"expected_compression": "32x"
},
"binary_2bit": {
"config": models.BinaryQuantization(
binary=models.BinaryQuantizationConfig(
encoding=models.BinaryQuantizationEncoding.TWO_BITS,
always_ram=True,
)
),
"expected_speedup": "20x",
"expected_compression": "16x"
}
}
# Create quantized collections
for method_name, config_info in quantization_configs.items():
collection_name = f"quantized_{method_name}"
client.create_collection(
collection_name=collection_name,
vectors_config=models.VectorParams(
size=1536, # Adjust to your embedding size
distance=models.Distance.COSINE,
on_disk=True, # Store originals on disk
),
quantization_config=config_info["config"]
)
print(f"Created {method_name} quantized collection: {collection_name}")
```
### Step 3: Upload Data and Measure Impact
Upload your dataset to each quantized collection and measure the performance differences:
```python
def benchmark(collection_name, your_test_queries, method_name):
"""Measure quantized search performance"""
# Test without oversampling/rescoring first
no_rescoring_metrics = measure_search_performance(
collection_name,
your_test_queries,
f"{method_name} (No Rescoring)"
)
# Test with oversampling and rescoring
def search_with_rescoring(collection_name, query, oversampling_factor=3.0):
start_time = time.time()
response = client.query_points(
collection_name=collection_name,
query=query,
limit=10,
search_params=models.SearchParams(
quantization=models.QuantizationSearchParams(
rescore=True,
oversampling=oversampling_factor,
)
),
)
return (time.time() - start_time) * 1000, response
# Measure with rescoring
rescoring_latencies = []
for query in your_test_queries:
latency, response = search_with_rescoring(collection_name, query)
rescoring_latencies.append(latency)
avg_rescoring = np.mean(rescoring_latencies)
p95_rescoring = np.percentile(rescoring_latencies, 95)
print(f"{method_name} (With Rescoring):")
print(f" Average latency: {avg_rescoring:.2f}ms")
print(f" P95 latency: {p95_rescoring:.2f}ms")
return {
"no_rescoring": no_rescoring_metrics,
"with_rescoring": {"avg": avg_rescoring, "p95": p95_rescoring}
}
# Upload your data (same as it was done in the previous days for a basic unquantized collection) in each collection
# Test each quantization method
quantization_results = {}
for method_name in quantization_configs.keys():
collection_name = f"quantized_{method_name}"
quantization_results[method_name] = benchmark(
collection_name, your_test_queries, method_name
)
```
### Step 4: Optimize Oversampling Factors
Find the optimal oversampling factor for your best-performing quantization method, based on the balance between latency and retained accuracy:
```python
def measure_accuracy_retention(original_collection, quantized_collection, test_queries, factors=[2, 3, 5, 8, 10]):
"""Compare search results between original and quantized collections"""
results = {}
for factor in factors:
accuracy_scores = []
for query in test_queries:
# Get baseline results
baseline_results = client.query_points(
collection_name=original_collection,
query=query,
limit=10
)
baseline_ids = [point.id for point in baseline_results.points]
# Get quantized results with rescoring
quantized_results = client.query_points(
collection_name=quantized_collection,
query=query,
limit=10,
search_params=models.SearchParams(
quantization=models.QuantizationSearchParams(
rescore=True,
oversampling=factor,
)
),
)
quantized_ids = [point.id for point in quantized_results.points]
# Calculate overlap (simple accuracy measure)
overlap = len(set(baseline_ids) & set(quantized_ids))
accuracy = overlap / len(baseline_ids)
accuracy_scores.append(accuracy)
results[factor] = {
"avg_accuracy": np.mean(accuracy_scores)
}
return results
def tune_oversampling(collection_name, test_queries, factors=[2, 3, 5, 8, 10]):
"""Find optimal oversampling factor"""
results = {}
for factor in factors:
latencies = []
for query in test_queries:
start_time = time.time()
response = client.query_points(
collection_name=collection_name,
query=query,
limit=10,
search_params=models.SearchParams(
quantization=models.QuantizationSearchParams(
rescore=True,
oversampling=factor,
)
),
)
latencies.append((time.time() - start_time) * 1000)
results[factor] = {
"avg_latency": np.mean(latencies),
"p95_latency": np.percentile(latencies, 95)
}
return results
# Tune oversampling for your method of choice
best_method = "binary" # Choose based on your results
oversampling_factors = [2, 3, 5, 8, 10]
oversampling_results_latency = tune_oversampling(
f"quantized_{best_method}",
your_test_queries,
oversampling_factors
)
oversampling_results_accuracy = measure_accuracy_retention(
"your_domain_collection",
f"quantized_{best_method}",
your_test_queries,
oversampling_factors
)
print("Oversampling Factor Optimization:")
for factor in oversampling_factors:
print(f" {factor}x:")
print(f" {oversampling_results_latency[factor]['avg_latency']:.2f}ms avg latency, {oversampling_results_latency[factor]['p95_latency']:.2f}ms P95 latency")
print(f" {oversampling_results_accuracy[factor]['avg_accuracy']:.2f} avg accuracy retention")
```
### Step 5: Analyze Your Results
Create a comprehensive analysis of your quantization experiments:
```python
print("=" * 60)
print("QUANTIZATION PERFORMANCE ANALYSIS")
print("=" * 60)
print(f"\nBaseline Performance:")
print(f" Average latency: {baseline_metrics['avg']:.2f}ms")
print(f" P95 latency: {baseline_metrics['p95']:.2f}ms")
print(f"\nQuantization Results:")
for method, results in quantization_results.items():
no_rescoring = results['no_rescoring']
with_rescoring = results['with_rescoring']
speedup_no_rescoring = baseline_metrics['avg'] / no_rescoring['avg']
speedup_with_rescoring = baseline_metrics['avg'] / with_rescoring['avg']
print(f"\n{method.upper()}:")
print(f" Without rescoring: {no_rescoring['avg']:.2f}ms ({speedup_no_rescoring:.1f}x speedup)")
print(f" With rescoring: {with_rescoring['avg']:.2f}ms ({speedup_with_rescoring:.1f}x speedup)")
```
## Success Criteria
You'll know you've succeeded when:
<input type="checkbox"> You've achieved measurable search speed improvements
<input type="checkbox"> You've maintained acceptable accuracy through oversampling optimization
<input type="checkbox"> You've demonstrated significant hot memory savings with `on_disk` configuration
<input type="checkbox"> You can make informed recommendations about quantization for your domain
## Share Your Discovery
### Step 1: Reflect on Your Findings
1. Which quantization method gave the best balance between speed and accuracy?
2. How did the oversampling factor change latency and accuracy?
3. What was the real memory and cost impact?
4. How do your results compare to the reference maximums (≈40× speed, ≈32× compression)?
### Step 2: Post Your Results
**Post your results in** <a href="https://discord.com/channels/907569970500743200/1429673887590776832" target="_blank" rel="noopener noreferrer" aria-label="Qdrant Discord"> <img src="https://img.shields.io/badge/Qdrant%20Discord-5865F2?style=flat&logo=discord&logoColor=white&labelColor=5865F2&color=5865F2"
alt="Post your results in Discord"
style="display:inline; margin:0; vertical-align:middle; border-radius:9999px;" /> </a> **using this:**
```markdown
**[Day 4] Quantization Performance Optimization**
**High-Level Summary**
- **Domain:** "I optimized [your domain] search with quantization"
- **Key Result:** "Best was [Scalar/Binary/(2-bit Binary)] with oversampling [x]× → [Z]× faster, [A]% accuracy retained."
**Reproducibility**
- **Collections:** day4_baseline_collection, day4_quantized_scalar, day4_quantized_binary (and/or day4_quantized_2bit)
- **Model:** [name, dim]
- **Dataset:** [N items] (snapshot: YYYY-MM-DD)
- **Search settings:** hnsw_ef=[..] (if used)
**Results**
- **Baseline latency:** [X] ms
- **Quantized latency (rescoring on):** [Y] ms
- **Oversampling:** [factor]×
- **Accuracy retention:** [..]%
- **Memory:** [before GB] → [after GB] (**[compression]×**)
- **(Optional) Cost:** ~$[before]/mo → ~$[after]/mo, save ~$[delta]/mo
**Method Notes**
- **Scalar (INT8):** [one line]
- **Binary (1-bit / 2-bit):** [one line]
**Surprise**
- "[most unexpected finding]"
**Next step**
- "[one concrete action for tomorrow]"
```
## Optional: Go Further
### Dynamic Oversampling
Implement adaptive oversampling based on query characteristics:
```python
def adaptive_oversampling(query, base_factor=3.0):
"""Adjust oversampling based on query complexity"""
# Simple heuristic: longer queries may need more oversampling (adapt to your domain/use case)
query_length = len(query) if isinstance(query, str) else len([x for x in query if x != 0])
if query_length > 1000: # Complex query
return base_factor * 1.5
elif query_length < 100: # Simple query
return base_factor * 0.8
else:
return base_factor
# Test adaptive oversampling vs fixed oversampling
```
### Cost-Performance Analysis
Calculate the true cost impact of quantization:
```python
def calculate_cost_savings(baseline_memory_gb, compression_ratio, ram_cost_per_gb_monthly=10):
"""Calculate monthly cost savings from quantization"""
quantized_memory_gb = baseline_memory_gb / compression_ratio
monthly_savings = (baseline_memory_gb - quantized_memory_gb) * ram_cost_per_gb_monthly
return {
"baseline_cost": baseline_memory_gb * ram_cost_per_gb_monthly,
"quantized_cost": quantized_memory_gb * ram_cost_per_gb_monthly,
"monthly_savings": monthly_savings,
"annual_savings": monthly_savings * 12
}
# Calculate cost impact for your deployment
cost_analysis = calculate_cost_savings(
baseline_memory_gb=10, # Your baseline memory usage
compression_ratio=32, # Your best quantization compression
)
print(f"Annual cost savings: ${cost_analysis['annual_savings']:.2f}")
```
### Memory Usage Monitoring
Track actual memory usage changes:
```python
# Monitor collection memory usage
collection_info = client.get_collection("quantized_binary")
print(f"Vectors count: {collection_info.points_count}")
print(f"Memory usage: Check Qdrant Cloud metrics")
# Compare RAM usage with and without on_disk configuration
```
@@ -0,0 +1,77 @@
---
title: Accuracy Recovery with Rescoring
weight: 2
---
{{< date >}} Day 4 {{< /date >}}
# Accuracy Recovery with Rescoring
<div class="video">
<iframe
src="https://www.youtube.com/embed/ksw3Ok-XXqo"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
allowfullscreen>
</iframe>
</div>
<br/>
When we use quantization methods like Scalar, Binary, or Product Quantization, we're compressing our vectors to save memory and improve performance. However, this compression can slightly reduce the accuracy of our similarity searches because the quantized vectors are approximations of the original data. To mitigate this loss of accuracy, you can use oversampling and rescoring, which help improve the accuracy of the final search results.
So let's say we are performing a search in a collection with Binary Quantization. Qdrant retrieves the top candidates using the quantized vectors based on their similarity to the query vector, as determined by the quantized data. This step is fast because we're using the quantized vectors.
## Oversampling
Since some relevant matches could be missed in the initial search, to compensate for that we will apply oversampling. Which means that you will retrieve more candidates, increasing the chances that the most relevant vectors make it into the final results.
For example, if your desired number of results (`limit`) is 4 and you set an oversampling factor of 2, Qdrant will retrieve 8 candidates (4 × 2). More candidates mean a better chance of obtaining high-quality top-K results.
## Rescoring
After oversampling to gather more potential matches, each candidate is re-evaluated based on additional criteria to ensure higher accuracy and relevance to the query. The rescoring process maps the quantized vectors to their corresponding original vectors, allowing you to consider factors like context, metadata, or additional relevance that wasn't included in the initial search, leading to more accurate results.
During rescoring, one of the lower-ranked candidates from oversampling might turn out to be a better match than some of the original top-K candidates. Even though rescoring uses the original, larger vectors, the process remains much faster because only a very small number of vectors are read.
**Reranking as a Result of Rescoring**
With the new similarity scores from rescoring, reranking is where the final top-K candidates are determined based on the updated similarity scores.
For example, in our case with a limit of 4, a candidate that ranked 6th in the initial quantized search might improve its score after rescoring because the original vectors capture more context or metadata. As a result, this candidate could move into the final top 4 after reranking, replacing a less relevant option from the initial search.
## Implementation
```python
from qdrant_client import QdrantClient, models
import os
from dotenv import load_dotenv
load_dotenv()
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
# For Colab:
# from google.colab import userdata
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
response = client.query_points(
collection_name="quantized_collection",
query=[0.12] * 1536,
limit=10,
search_params=models.SearchParams(
hnsw_ef=128,
quantization=models.QuantizationSearchParams(
ignore=False, # Use quantization for initial search
rescore=True, # Enable original vectors-based rescoring
oversampling=3.0, # Retrieve 3x candidates for rescoring
),
),
with_payload=True,
)
```
> Check out **[how to set up oversampling and rescoring](/documentation/guides/quantization/#searching-with-quantization)** in **TypeScript**, **Rust**, **Java**, **C#**, and **Go** clients.
If quantization is impacting performance in an application that requires high accuracy, combining oversampling with rescoring is a great choice. However, if you need faster searches and can tolerate some loss in accuracy, you might choose to use oversampling without rescoring, or adjust the oversampling factor to a lower value.
> Check out our **[quantization tips](/documentation/guides/quantization/#quantization-tips)**
@@ -0,0 +1,151 @@
---
title: Vector Quantization Methods
weight: 1
---
{{< date >}} Day 4 {{< /date >}}
# Vector Quantization Methods
<div class="video">
<iframe
src="https://www.youtube.com/embed/oExGyAEOpP4"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
allowfullscreen>
</iframe>
</div>
Production vector search engines face an inevitable scaling challenge: memory requirements grow with dataset size, while search latency demands vectors remain in fast storage. [Quantization](/documentation/guides/quantization/) provides the solution by compressing vector representations while maintaining retrieval quality - but the method you choose fundamentally determines your system's performance characteristics.
## The Memory Economics
Consider the mathematics of scale. OpenAI's `text-embedding-3-small` produces 1536-dimensional vectors requiring 6 KB each (1536 × 4 bytes per float32). This scales predictably: 1 million vectors consume 6 GB, 10 million require 60 GB, and 100 million demand 600 GB of memory.
For production RAG systems serving real-time queries, these vectors must reside in RAM or high-speed SSD storage. At cloud pricing, this translates to substantial infrastructure costs - hundreds to thousands of dollars monthly for enterprise-scale deployments.
Quantization breaks this cost scaling by compressing vectors while maintaining search effectiveness. The three primary quantization families - scalar, binary, and product - offer different points on the speed-memory-accuracy tradeoff curve, each optimized for specific deployment scenarios.
## Scalar Quantization
[Scalar quantization](/articles/scalar-quantization/) maps each float32 dimension (4 bytes) to an int8 representation (1 byte), achieving 4x memory compression through learned range mapping. The algorithm analyzes your vector distribution and determines optimal bounds - typically using quantiles to exclude outliers - then linearly maps the float32 range to the int8 range (-128 to 127).
The technique leverages SIMD optimizations available for int8 operations, enabling up to 2x speed improvements beyond the memory benefits. Distance calculations using int8 values are computationally simpler than float32 operations, particularly for dot product and cosine similarity computations that dominate vector search workloads.
Scalar quantization excels as the production default because it [maintains 99%+ accuracy](/articles/scalar-quantization/) across diverse embedding models while providing predictable compression ratios. Unlike binary quantization, which requires specific model characteristics, scalar quantization works reliably with embeddings from commercial providers (OpenAI, Cohere, Anthropic) and open-source models alike.
```python
from qdrant_client import QdrantClient, models
import os
from dotenv import load_dotenv
load_dotenv()
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
# For Colab:
# from google.colab import userdata
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
# Scalar quantization setup
client.create_collection(
collection_name="scalar_collection",
vectors_config=models.VectorParams(
size=1536,
distance=models.Distance.COSINE,
on_disk=True, # Move originals to disk
),
quantization_config=models.ScalarQuantization(
scalar=models.ScalarQuantizationConfig(
type=models.ScalarType.INT8,
quantile=0.99, # Exclude extreme 1% of values
always_ram=True, # Keep quantized vectors in RAM
)
),
)
```
> [Check out](/documentation/guides/quantization/#setting-up-scalar-quantization) how to set up scalar quantization in **TypeScript**, **Rust**, **Java**, **C#**, and **Go** clients.
## Binary Quantization
[Binary quantization](https://youtu.be/wRnSDDzLQmk) represents the extreme compression approach, reducing each dimension to a single bit through sign-based thresholding: values greater than zero become 1, values less than or equal to zero become 0. This transforms a 1536-dimensional vector from 6 KB (1536 × 4 bytes) to 192 bytes (1536 bits ÷ 8), achieving 32x memory compression.
The computational advantages are substantial. Bitwise operations enable distance calculations using native CPU instructions, delivering up to 40x speed improvements over float32 computations. Modern processors excel at parallel bitwise operations, making binary quantization particularly effective for high-throughput search scenarios.
However, binary quantization demands specific model characteristics for optimal performance. The technique works best with high-dimensional vectors (≥1024 dimensions) that exhibit centered value distributions around zero. Models like OpenAI's text-embedding-ada-002 and Cohere's embed-english-v2.0 have been validated for binary compatibility, but other models may experience significant accuracy degradation.
> **<font color='red'>Update:</font>** Starting from Qdrant **v1.15.0**, [two additional quantization types](/documentation/guides/quantization/#15-bit-and-2-bit-quantization) were introduced: **1.5-bit** and **2-bit binary quantization**.
> These methods provide a useful middle ground: they are more aggressive than scalar quantization but offer better precision than standard binary quantization. They also address one of binary quantization’s main weaknesses: **handling values close to zero**.
>
> **<font color='red'>Additionally:</font>** [Asymmetric quantization](#asymmetric-quantization) was added. This method allows combining different quantization strategies for queries and documents, helping balance **accuracy** and **compression efficiency**.
```python
# Binary quantization setup
client.create_collection(
collection_name="binary_collection",
vectors_config=models.VectorParams(
size=1536,
distance=models.Distance.COSINE,
on_disk=True,
),
quantization_config=models.BinaryQuantization(
binary=models.BinaryQuantizationConfig(
encoding=models.BinaryQuantizationEncoding.ONE_BIT,
always_ram=True,
)
),
)
```
> [Check out](/documentation/guides/quantization/#setting-up-binary-quantization) how to set up binary quantization in **TypeScript**, **Rust**, **Java**, **C#**, and **Go** clients.
## Product Quantization
[Product quantization](/articles/product-quantization/) employs a divide-and-conquer approach, segmenting vectors into sub-vectors and encoding each segment using learned codebooks. The algorithm splits a high-dimensional vector into equal-sized sub-vectors, then applies k-means clustering to each segment independently, creating separate codebooks of 256 centroids per segment.
The compression mechanism stores centroid indices rather than original values. A 1024-dimensional vector divided into 128 sub-vectors (8 dimensions each) requires only 128 bytes for storage (128 indices × 1 byte), achieving approximately 32x compression. In extreme configurations, product quantization can reach 64x compression ratios.
The tradeoff is computational complexity and accuracy degradation. Distance calculations become non-SIMD-friendly, often resulting in slower query performance than unquantized vectors. The segmented encoding introduces approximation errors that compound across sub-vectors, leading to more significant accuracy penalties compared to scalar or binary methods. Product quantization serves specialized use cases where extreme compression outweighs accuracy and speed considerations.
```python
# Product quantization setup
client.create_collection(
collection_name="pq_collection",
vectors_config=models.VectorParams(
size=1024,
distance=models.Distance.COSINE,
on_disk=True,
),
quantization_config=models.ProductQuantization(
product=models.ProductQuantizationConfig(
compression=models.CompressionRatio.X32, #or X4, X8, X16, X32 and X64
always_ram=True,
)
),
)
```
> [Check out](/documentation/guides/quantization/#setting-up-product-quantization) how to set up product quantization in **TypeScript**, **Rust**, **Java**, **C#**, and **Go** clients.
## Quantization Comparison
| Quantization Method | Accuracy | Speed | Compression |
|---------------------|----------|----------|-------------|
| Scalar | 0.99 | up to 2x | 4x |
| Binary | 0.95* | up to 40x| 32x |
| Product | 0.7 | 0.5x | up to 64x |å
*For compatible models
> [Check out](/documentation/guides/quantization/#how-to-choose-the-right-quantization-method) how the new **1.5-bit** and **2-bit binary quantization** methods compare to classical binary quantization. They offer a balanced middle ground between **binary** and **scalar** approaches.
## Dual Storage Architecture
Qdrant's quantization implementation maintains both compressed and original vectors, enabling flexible deployment strategies and safe experimentation. This dual storage approach allows you to switch quantization methods, adjust parameters, or disable quantization entirely without data re-ingestion - a critical advantage for production systems where data pipeline complexity must be minimized.
> Check out our **[quantization tips](/documentation/guides/quantization/#quantization-tips)**
The default configuration stores both representations in RAM, providing fast search with quantized vectors and exact scoring with originals when needed. However, this negates memory savings. The optimal production pattern places original vectors on disk (`on_disk=True`) while keeping quantized vectors in RAM (`always_ram=True`). This configuration delivers the best of both worlds: rapid quantized search with the ability to perform exact rescoring by reading only the small candidate set from disk storage.
This storage strategy is particularly effective because quantized search reliably identifies the neighborhood of relevant documents, making the selective disk reads for rescoring both efficient and accurate. The result is dramatic RAM reduction with minimal latency impact, exactly what production RAG systems require for cost-effective scaling.