Merge remote-tracking branch 'origin/master' into sync-beginners-course-module-3

# Conflicts:
#	qdrant-landing/content/course/beginners/_index.md
#	qdrant-landing/content/course/beginners/module-0/_index.md
#	qdrant-landing/content/course/beginners/module-0/qdrant-cloud.md
This commit is contained in:
kanungle
2026-09-08 18:38:58 -07:00
1680 changed files with 42601 additions and 79759 deletions
+18 -22
View File
@@ -15,6 +15,23 @@ Whether you’re new to Qdrant or building production-grade systems, our guided
## Available Now
{{< course-card
title="Qdrant Beginner Course"
image="/icons/outline/training-white.svg"
link="/course/beginners/"
>}}
**What you'll gain:**
- Why Traditional Search Falls Short
- Embeddings and Distance Metrics
- Vector Search First Principles
- Sparse, Dense, and Hybrid Search
- Designing a Vector Search System
- Capstone: Multimodal Supplier Risk Intelligence
<br><br>
Time to Complete: under 5 hours<br>
Includes: videos, code notebooks, projects, certification
{{< /course-card >}}
{{< course-card
title="Qdrant Essentials Course"
image="/icons/outline/rocket-white-light.svg"
@@ -51,27 +68,6 @@ Includes: videos, code notebooks, projects, certification
## Upcoming Courses
### Beginner Level
Beginner courses require no previous experience with Qdrant and are useful for building a strong foundation in vector search.
{{< accordion >}}
- title: "Qdrant Fundamentals"
content: |
- Vector Search Concepts
- Setting Up Qdrant
- Creating and Managing Collections
- Ingesting Vector Embeddings
- Running Your First Query
<br>
<br>
Time to Complete: 2 hours (TBD)<br>
Includes: videos, code notebooks
<br>
<br>
<p style="margin-left: 0px;"><a href="https://forms.gle/jiBmDcXr8j9tAb5UA" target="_blank">→ Register Interest</a></p>
{{< /accordion >}}
### Intermediate Level
Intermediate courses are recommended for those that have completed the Beginner Level Modules first, and extend knowledge into more practical usage of Qdrant in the real-world.
@@ -148,4 +144,4 @@ Advanced courses are recommended for those that have completed the Beginner and
</p>
{{< /accordion >}}
**Want something not mentioned above? Email [devrel@qdrant.com](emailto:devrel@qdrant.com) and let us know!**
**Want something not mentioned above? Email [devrel@qdrant.com](mailto:devrel@qdrant.com) and let us know!**
@@ -0,0 +1,206 @@
---
title: "Beginner Course"
page_title: "Qdrant Beginner Course"
short_description: "Learn the fundamentals of vector search: why keyword search struggles, how semantic search improves it, embeddings, distance metrics, and hybrid systems."
description: "Understand why traditional search struggles and how modern semantic search improves it, and build your first search system."
content:
sidebarTitle: "Beginner Course"
menuTitle:
text: Course Overview
url: /course/beginners/
nextButton: Continue to Next Step
nextDay: Complete
title: "Beginner Course"
description: "Understand why traditional search struggles and how modern semantic search improves it, and build your first search system."
partition: course
---
# Beginner Course
**Learn the fundamentals of vector search**
Understand why traditional search struggles and how modern semantic search improves it. Learn about embeddings, distance metrics, and hybrid search systems.
<br/>
<div class="video">
<iframe
src="https://www.youtube.com/embed/9XKQ3NSch_s?rel=0"
title="YouTube video player"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
allowfullscreen>
</iframe>
</div>
{{< cards-list >}}
- icon: /icons/outline/play-white.svg
title: 6 modules
content: From setting up dependencies to a hands-on capstone project
- icon: /icons/outline/cloud-check-blue.svg
title: Shareable certificate
content: Earn a digital certificate upon completion
- icon: /icons/outline/time-blue.svg
title: Flexible schedule
content: Learn at your own pace
- icon: /icons/outline/plan.svg
title: Beginner level
content: No prior experience required
{{< /cards-list >}}
<br/>
## What You'll Learn
{{< course-card
title="Skills you'll gain:"
image="/icons/outline/training-white.svg"
type="wide-list">}}
- Why traditional search struggles and how modern semantic search improves it
- How embeddings convert text to vectors that capture meaning
- Distance metrics: cosine similarity, dot product, Euclidean and Manhattan
- Hybrid search: combining dense and sparse retrieval
- Building your first Qdrant collection and queries
{{< /course-card >}}
### The Path
**Module 0**: Setting Up Dependencies. Configure your environment and get started with the basics.
**Module 1**: Let's Understand Search. Understand why traditional search struggles and how modern semantic search improves it.
**Module 2**: First Principles of Vector Search. Anatomy of a vector - how data is stored, indexed, and retrieved in Qdrant.
**Module 3**: Sparse vs Dense vs Hybrid Search. Understand dense vs sparse search, when each fails, and how hybrid systems combine them.
**Module 4**: Designing a Vector Search System. How to design a vector search system - layers, filtering, RAG, and deployment.
**Module 5**: Capstone - Multimodal Supplier Risk Intelligence. Ingest, cluster, and query multimodal supplier signals across languages.
**Bonus Module**: Further Reading. A roundup of advanced techniques for further reading: score boosting, relevance feedback, MMR, and re-ranking.
## How the Course Works
{{< cards-list >}}
- icon: /icons/outline/training-purple.svg
title: Bite-sized lessons
content: Short, friendly modules you can finish in one sitting
- icon: /icons/outline/hacker-purple.svg
title: Learn by doing
content: Follow along with real examples and hands-on exercises
- icon: /icons/outline/similarity-blue.svg
title: One step at a time
content: Each module builds on the last, so nothing feels out of reach
- icon: /icons/outline/copy.svg
title: Go at your own pace
content: Pause anytime and pick up right where you left off
{{< /cards-list >}}
<br/>
## Syllabus
{{< accordion >}}
- title: "Module 0: Setting Up Dependencies"
content: |
- Qdrant Cloud Setup
- Implementing a Basic Vector Search
<br>
<br>
<p style="margin-left: 0px;"><a href="/course/beginners/module-0/">→ Start Module 0</a></p>
- title: "Module 1: Let's Understand Search"
content: |
- The Problem: Why Traditional Search Struggles
- How Traditional Search Improved
- Enter Semantic Search
- How It Works: Embeddings
- Comparing Meaning: Distance Metrics
- Why Similarity Alone Is Not Enough
- Modern Search = Hybrid Systems
- References & Further Reading
<br>
<br>
<p style="margin-left: 0px;"><a href="/course/beginners/module-1/">→ Start Module 1</a></p>
- title: "Module 2: First Principles of Vector Search"
content: |
- What is a Vector?
- How Dimensions Represent Meaning
- Similarity Under the Hood
- Your First Qdrant Collection
- Points, Payloads, and Queries
<br>
<br>
<p style="margin-left: 0px;"><a href="/course/beginners/module-2/">→ Start Module 2</a></p>
- title: "Module 3: Sparse vs Dense vs Hybrid Search"
content: |
- The Two Families of Search
- Hybrid Search: Dense + Sparse
- Setting Up Hybrid Search in Qdrant
- Fusion Strategies
- Beyond Text: Multimodal Search
<br>
<br>
<p style="margin-left: 0px;"><a href="/course/beginners/module-3/">→ Start Module 3</a></p>
- title: "Module 4: Designing a Vector Search System"
content: |
- Architecture Layers of a Search System
- Filtering and Metadata Strategies
- Retrieval-Augmented Generation (RAG) Patterns
- Deployment Considerations
<br>
<br>
<p style="margin-left: 0px;"><a href="/course/beginners/module-4/">→ Start Module 4</a></p>
- title: "Module 5: Capstone - Multimodal Supplier Risk Intelligence"
content: |
- Ingesting Multimodal Supplier Signals
- Clustering Signals Across Languages
- Querying the Capstone System
- Putting It All Together
<br>
<br>
<p style="margin-left: 0px;"><a href="/course/beginners/module-5/">→ Start Module 5</a></p>
- title: "Bonus Module: Further Reading"
content: |
- Score Boosting
- Relevance Feedback
- Maximal Marginal Relevance (MMR)
- Re-ranking
- Other Advanced Techniques
<br>
<br>
<p style="margin-left: 0px;"><a href="/course/beginners/module-6/">→ Start Module 6</a></p>
{{< /accordion >}}
## Who It's For
Anyone new to vector search who wants to understand the fundamentals. No prior experience with Qdrant or vector search engines required.
## Time Commitment
- Core course (Modules 0-4): under 2 hours
- Capstone project (Module 5): ~3 hours
- **Total: under 5 hours**
- Bonus module: optional, not included in the total above
- Self-paced, flexible schedule
{{< course-card
title="Ready to start your vector search journey?"
image="/icons/outline/rocket-white-light.svg"
link="/course/beginners/module-0/">}}
**What you'll get**
- Understand the fundamentals of vector search
- Learn why semantic search outperforms keyword search
- Build your first Qdrant collection
- Foundation for advanced courses
{{< /course-card >}}
@@ -0,0 +1,26 @@
---
title: "Qdrant Beginner Certification"
short_description: "Validate your vector search fundamentals with an official certification exam covering semantic search, embeddings, and hybrid retrieval."
description: "Earn the official Qdrant Beginner certification: prove you can build collections, choose distance metrics, filter payloads, and run hybrid search."
isLesson: true
weight: 100
---
# Qdrant Beginner Certification
Congratulations! You've completed the **Qdrant Beginner** course.
Along the way you learned why keyword search falls short, how embeddings capture meaning, how distance metrics compare that meaning, and how hybrid search brings dense and sparse retrieval together. That effort deserves professional recognition.
## 🏆 Get #QdrantCertified
You've got the fundamentals down, and now it's time to prove it. Validate your skills with our official certification.
**Head over to [train.qdrant.dev](https://train.qdrant.dev) to take the exam.**
Passing this exam shows you can:
* **Explain** when and why semantic search beats keyword search.
* **Turn** text into embeddings and choose the right distance metric for the job.
* **Build** a Qdrant collection and run vector, filtered, and hybrid queries.
* **Design** a complete search system end to end, all the way to a multimodal capstone project.
@@ -0,0 +1,20 @@
---
title: "Module 0: Setting Up Dependencies"
short_description: "Module 0 of the Beginner Course: set up Qdrant Cloud, build a first vector search, and get started with the basics."
description: "Set up Qdrant and build your first vector search app. Learn how to configure Qdrant Cloud, run a basic search, and get started with the fundamentals."
isLesson: true
weight: 10
---
{{< date >}} Module 0 {{< /date >}}
# Setting Up Dependencies
Get started with Qdrant by setting up your environment and building your first vector search application.
## Today's Path
1. Qdrant Cloud Setup
2. Implementing a Basic Vector Search
By the end, you'll have a working Qdrant setup and a complete first search running.
@@ -0,0 +1,85 @@
---
title: "Implementing a Basic Vector Search"
short_description: "Walk through your first vector search: connect to Qdrant, create a collection, insert points, and run similarity queries with the Python client."
description: Learn how to build a basic vector search in Qdrant. Create collections, insert vectors, and run your first similarity search step-by-step with Python.
weight: 3
isLesson: true
---
{{< date >}} Module 0 {{< /date >}}
# Implementing a Basic Vector Search
<!--
TODO (video): the previous embed here (_83L9ZIoOjM) is the Essentials-course
recording — the narration names "Day 0 of the Essentials course," which
contradicts this Beginner Course Module 0 page. Drop in a beginner-specific cut here,
or leave this commented out until one exists.
-->
In this lesson you'll build your very first search, one small step at a time. You'll connect to Qdrant, create a place to store data, add a few example vectors, and then ask Qdrant to find the closest match. Every step has runnable code, so follow along in a notebook or script.
A quick vocabulary note before you start: a **vector** is just a list of numbers that represents something (a piece of text, an image, a product). Searching by vectors means finding the entries whose numbers are closest to your query's numbers. That's the whole idea, and the code below makes it concrete.
## Before You Start
This course requires Python 3.11 or above installed
## Step 1: Install the Qdrant Client
The **client** is the Python library that lets your code talk to Qdrant. Install it first:
```python
!pip install qdrant-client
```
## Step 2: Import the Libraries You'll Need
Import two things from the package: `QdrantClient`, which opens the connection, and `models`, which holds the building blocks you'll use to describe collections and points.
```python
from qdrant_client import QdrantClient, models
```
## Step 3: Connect to Qdrant Cloud
Use the cluster URL and API key from the previous lesson. If you saved them in a `.env` file, this reads them automatically:
```python
import os
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
# For Colab:
# from google.colab import userdata
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
```
**Tip:** For quick experiments with no cloud account at all, you can use `client = QdrantClient(":memory:")`. It runs entirely in memory, but your data disappears when the program stops.
## Step 4: Create a Collection
A [collection](/documentation/manage-data/collections/) is where your vectors live. It's a lot like a table in a regular database: a named container for related data. When you create one, you tell Qdrant two things:
- **Size:** how many numbers each vector has.
- **Distance metric:** how Qdrant measures whether two vectors are "close."
```python
# Name your collection
collection_name = "my_first_collection"
# Create it, describing the vectors it will hold
client.create_collection(
collection_name=collection_name,
vectors_config=models.VectorParams(
size=4, # each vector has 4 numbers
distance=models.Distance.COSINE # how we measure closeness
)
)
```
This returns `True` when it works.
If completed correctly, you will now have an established Qdrant environment for the rest of the course. Later modules will explain collections, points, distance metrics, and more. Keep going to find out more!
**Congratulations! You've completed Module 0.** 🎉
@@ -0,0 +1,174 @@
---
title: "Qdrant Setup"
short_description: "Spin up a managed Qdrant Cloud cluster, generate API keys, and explore the Web UI for collections, points, and cluster monitoring."
description: Set up your Qdrant Cloud cluster in minutes. Learn to create collections, manage data, access the Web UI, and connect securely from Python.
weight: 2
isLesson: true
---
{{< date >}} Module 0 {{< /date >}}
# Qdrant Setup
<div class="video">
<iframe
src="https://www.youtube.com/embed/-8TpooWX8kE?rel=0"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
allowfullscreen>
</iframe>
</div>
<br/>
Welcome to your first hands-on step. Before you can search anything, you need a place to store your vectors. That's what Qdrant Cloud gives you: a managed Qdrant environment that runs in the cloud, so there's nothing to install and nothing to keep running on your own local machine. It comes with a secure connection, backups, easier updates, and a clean interface you'll use throughout this course.
Don't worry if some terms here are new. You'll set up a cluster, get a key that lets your code talk to it, and run one quick check to confirm it's working. That's the whole goal for this lesson.
## Create Your Cluster
A **cluster** is your personal Qdrant instance in the cloud. Here's how to create one:
1. Sign up at [cloud.qdrant.io](https://cloud.qdrant.io/signup) with email, Google, or GitHub.
2. Open **Clusters** and select **Create a Free Cluster**. The Free Tier is enough for this whole course, and you won't be asked for a card.
![Screenshot of the Qdrant Cloud page for creating a new cluster](/docs/gettingstarted/gui-quickstart/create-cluster.png)
3. Pick a region close to you or your users. This keeps things fast.
4. When the cluster is ready, copy the **API key** and store it somewhere safe. An API key is like a password your code uses to prove it's allowed to reach your cluster, so treat it like one. You can always create new keys later from the **API Keys** section on the cluster page.
![Screenshot of the API key panel in Qdrant Cloud](/docs/gettingstarted/gui-quickstart/api-key.png)
## Access the Web UI
The **Web UI** is a dashboard for looking at your data and running searches without writing code. It's the fastest way to see what's happening inside your cluster while you learn.
1. Select **Cluster UI** in the top corner of the cluster page to open the dashboard.
![Screenshot of the Qdrant Cloud dashboard](/docs/gettingstarted/gui-quickstart/access-dashboard.png)
### What You Can Do in the Web UI
Use the Web UI to manage collections, inspect data, and check how your searches perform.
#### Main Navigation
- **Console:** Run commands against Qdrant right in the browser. Great for testing and seeing responses without writing a program.
- **Collections:** See and manage all your collections in one place, and track their status, size, and settings at a glance.
- **Tutorial:** Follow a guided walkthrough with sample data. You create a collection, add vectors, and run a search with live results.
![Screenshot of the interactive tutorial in the Qdrant Web UI](/docs/gettingstarted/gui-quickstart/interactive-tutorial.png)
- **Datasets:** Load ready-made public datasets into your cluster with one click.
#### Inside a Collection
When you open a collection by selecting its name,
![Screenshot showing how to select a collection in the Web UI](/courses/day0/select-collection.png)
you'll see a detailed view with several tabs. You don't need all of these yet, so here's a plain-language tour you can come back to later:
![Screenshot of the points view inside a collection](/courses/day0/collection-points.png)
- **Points Tab:** Look at, search, and manage your individual data entries. You can view each entry's data, run a quick "find similar" search, or open a graph view of how it connects to its neighbors.
- **Info Tab:** A health check for the collection. The one field to know for now is `status` — `green` means everything is healthy.
- **Cluster Tab:** Shows how your data is spread across machines. You'll care about this only once you scale up.
- **Search Quality Tab:** Measures how accurate your searches are. Useful later, when you start tuning.
- **Snapshots Tab:** Manage backups of the collection. You can create a [collection snapshot](/documentation/snapshots/), restore it, or move it to another cluster.
- **Visualize Tab:** See your vectors as a 2D map. A nice way to build intuition once you have real data loaded.
- **Graph Tab:** Explore how points connect to their nearest neighbors.
## Connect from Python
Now let's connect from code. First, store your credentials in a file named `.env` at the root of your project (or set them in Colab). Keeping them in a separate file means you won't accidentally paste your key into shared code:
```env
QDRANT_URL=https://YOUR-CLUSTER.cloud.qdrant.io:6333
QDRANT_API_KEY=YOUR_API_KEY
```
Then load those values and create a client. The **client** is the object your Python code uses to send requests to Qdrant:
```python
from qdrant_client import QdrantClient, models
import os
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
# For Colab:
# from google.colab import userdata
# client = QdrantClient(url=userdata.get("QDRANT_URL"), api_key=userdata.get("QDRANT_API_KEY"))
# Quick health check
collections = client.get_collections()
print(f"Connected to Qdrant Cloud: {len(collections.collections)} collections")
```
If that prints a line about being connected, you're done. That's the whole setup.
## Other Ways to Connect
You can also reach your cluster directly over the web, without Python. This is handy for a quick test:
```bash
# Using the api-key header
curl -X GET https://xyz-example.eu-central.aws.cloud.qdrant.io:6333/collections \
--header 'api-key: <your-api-key>'
# Using the Authorization header
curl -X GET https://xyz-example.eu-central.aws.cloud.qdrant.io:6333/collections \
--header 'Authorization: Bearer <your-api-key>'
```
## Quick Validation
If you want to double-check the connection, these two commands confirm your cluster is up and reachable:
```bash
# Service health
curl -s "$QDRANT_URL/healthz" -H "api-key: $QDRANT_API_KEY"
# List collections
curl -s "$QDRANT_URL/collections" -H "api-key: $QDRANT_API_KEY"
```
## Good Practices
A few habits worth starting now:
- Keep your key out of your code. Use an environment variable or a secrets manager.
- Rotate your API keys now and then from the cluster **Access** tab.
- Use HTTPS only, and tighten access before you expose a cluster to the public internet.
## Common Issues
- **Authentication error:** Recheck the API key and the `api-key` header. A stray space or a missing character is the usual cause.
- **Connection error:** Confirm the cluster is running and the region URL is correct. Some workplace networks block outbound connections, so try from a personal network if a request hangs.
## Qdrant Cloud Inference
This part is optional, but good to know it exists. Normally you turn text or images into vectors yourself before storing them. **[Cloud Inference](/cloud-inference/)** does that step for you inside Qdrant Cloud: you send raw text or images, and Qdrant creates the vectors and stores them in one call. You'll create vectors by hand in the next lessons so you understand what's happening, but this is a shortcut you can reach for later.
<div class="video">
<iframe
src="https://www.youtube.com/embed/nJIX0zhrBL4?rel=0"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
allowfullscreen>
</iframe>
</div>
Learn more in the [Qdrant Cloud Inference documentation](/documentation/cloud/inference/).
## Qdrant Agent Skills
If you're using an AI coding assistant (Claude Code, Cursor, and others) alongside this course, install the [Qdrant Advisor skill](https://qdrant.tech/documentation/skills/#the-qdrant-advisor) early with this simple command:
```bash
npx skills add qdrant/skills/meta/qdrant-advisor
```
It's a single assistant that can troubleshoot and advise on any Qdrant deployment: when you describe a problem like slow search, memory climbing toward an out-of-memory crash, a stuck optimizer, a scaling decision, it searches live documentation, pulls only the branch of guidance that matches your symptom, and grounds its diagnosis in that current, official guidance instead of stale training data.
@@ -3,6 +3,7 @@ title: "Qdrant Essentials Course"
page_title: Qdrant Essentials Course
short_description: "Build production vector search skills in seven days: hybrid retrieval, multivector reranking, quantization, sharding, and multitenancy."
description: Learn hybrid search, multivectors, and production deployment in 7 days. Build and ship a docs search engine.
weight: 10
content:
sidebarTitle: Qdrant Essentials
menuTitle:
@@ -173,7 +174,7 @@ Build the vector search skills that matter: hybrid retrieval, multivector rerank
content: |
- AI & LLM Frameworks (Haystack, Jina AI, TwelveLabs)
- Data Processing (Unstructured.io)
- ML Platforms & Analytics (Tensorlake, Vectorize.io, Superlinked, Quotient)
- ML Platforms & Analytics (Tensorlake, Vectorize.io, Superlinked)
<br>
<br>
<p style="margin-left: 0px;"><a href="/course/essentials/day-7/">→ Start Day 7</a></p>
@@ -12,7 +12,7 @@ isLesson: true
<div class="video">
<iframe
src="https://www.youtube.com/embed/9JBlgNBQoOY?si=7t3LAvMsUUtlUMN7&rel=0"
src="https://www.youtube.com/embed/-8TpooWX8kE?rel=0"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
@@ -1,7 +1,7 @@
---
title: "Large-Scale Data Ingestion"
short_description: "Choose the right ingestion strategy for Qdrant: batched upserts, upload_points, and streaming uploads for million- and billion-scale workloads."
description: Master large-scale vector ingestion in Qdrant. Explore batching, upload_points, and upload_collection methods, on-disk storage, and parallel streaming for billion-scale AI data pipelines.
description: Master large-scale vector ingestion in Qdrant. Compare upsert, upload_points, and upload_collection, and learn how to stream a large dataset into a collection without loading it into memory.
weight: 4
isLesson: true
---
@@ -12,7 +12,7 @@ isLesson: true
<div class="video">
<iframe
src="https://www.youtube.com/embed/Rawvm7TP1XI"
src="https://www.youtube.com/embed/EhSZkGwTWsA"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
@@ -24,37 +24,37 @@ isLesson: true
In vector search applications inserting a few thousand data points is straightforward but the dynamics change completely when dealing with millions or billions of records. Tiny inefficiencies in the ingestion process compound into significant time losses, increased memory pressure, and degraded search performance.
Every individual upsert call initiates a transaction that consumes memory and disk I/O to build parts of the index. At scale, this naive approach can overwhelm your system, causing upload times to spike and search quality to decrease. Efficiently preparing and loading your data into Qdrant is paramount for building a robust and scalable AI application.
Every individual upsert call initiates a transaction that consumes memory and disk I/O to build parts of the index. At scale, this naive approach can overwhelm your system, causing upload times to spike and search quality to decrease. Efficiently preparing and loading your data into Qdrant is paramount for building a reliable and scalable AI application.
## Choosing Your Ingestion Strategy
Qdrant provides several methods for data ingestion, each tailored to different scales and use cases. It should be noted that only the Python client supports the upload_points and upload_collection methods. If you're using Qdrant on a different client then we recommend using upsert with batch upload for large scale ingestion. [Learn more about bulk operations](/documentation/manage-data/points/#batch-update).
The Qdrant client gives you three ways to get points in. The first has you managing the batching; the other two hand that to the client.
- **upsert (Individual or Batched)**: This is the fundamental operation for adding or updating points. Individual upserts are best suited for real-time updates while batching works best for larger workloads.
- **upsert** is the basic write operation, and the one every client library has. Send points one at a time for real-time updates, or [in batches](/documentation/manage-data/points/#upload-points) for a bulk load, which minimizes the overhead of opening a connection per point. You decide the batch size and you send the requests.
- **upload_points**: This method is optimized for uploading an entire batch of points that can comfortably fit into your client's memory. It leverages features like lazy batching, retries, and parallelism, making it a strong choice for medium-sized datasets.
- **upload_points** takes an iterable of `models.PointStruct`, the record-oriented shape: one object per point, carrying its own id, vector, and payload.
- **upload_collection**: For truly large-scale datasets, upload_collection is the most powerful tool. It streams data directly from an iterator, meaning the entire dataset doesn't need to be loaded into memory at once. This memory-efficient approach is ideal for ingesting millions or billions of points.
- **upload_collection** takes `vectors`, `payload`, and `ids` as separate arguments, the column-oriented shape.
> **<font color='red'>Note:</font>** For clients in other languages like **TypeScript**, **Rust**, and **Go**, batched upsert calls are the recommended method for efficient data loading.
![upsert takes models.Batch or a list and you batch it yourself. upload_points is record-oriented, an iterable of PointStruct. upload_collection is column-oriented, taking vectors, payload and ids as parallel columns. All three write into the collection.](/courses/day4/choosing-an-upload-method.svg)
## Heuristics for Scale
Those last two are the same tool in two shapes. Both add [parallelization, retries, and lazy batching](/documentation/manage-data/points/#python-client-optimizations) on top of upsert, and because both accept iterators, neither needs the whole dataset in memory. Pick whichever matches how your data already sits: the docs note the two formats are equivalent internally and offered for convenience.
Deciding which method to use can be guided by a few simple rules of thumb. While every use case is different, these heuristics provide a solid starting point:
You can also skip generating embeddings yourself. With [inference](/documentation/inference/), you send the text or image and the model name, and Qdrant produces the vector on upsert.
- **Less than 100,000 points**: A single-threaded, batched upsert operation will generally perform well.
- **100,000 to 1 million points**: `upload_points` is recommended, using batch sizes between 1,000 and 10,000 to balance network overhead and memory usage.
- **More than 1 million points**: `upload_collection` is the ideal choice for streaming data from disk. To maximize throughput, you should enable parallelism by setting the `parallel` parameter to the number of available CPU cores (e.g., 4 or 8).
`upload_points` and `upload_collection` are helpers in the client library rather than server endpoints, so what's available depends on your language:
> **<font color='red'>Best Practice:</font>** Start small and test. Before attempting to upload your entire dataset, ingest a smaller chunk to validate your configuration and process.
| Client | Bulk helper |
|---|---|
| Python | `upload_points`, `upload_collection` |
| Rust | `upsert_points_chunked(request, chunk_size)` |
| TypeScript, Go, Java, C# | Batched `upsert` calls |
## A Real-World Example: Ingesting LAION-400M
The bottleneck during upload is usually the client library, not the Qdrant server. If ingestion speed is your priority, the [Rust client](https://github.com/qdrant/rust-client) is the fastest option.
To illustrate these principles, let's examine the process of ingesting the LAION-400M dataset, which contains approximately 400 million image-text pairs with 512-dimensional CLIP embeddings. This massive dataset, with 400 GB of vectors and 200 GB of payload, requires a carefully optimized strategy.
## The Collection Configuration
### The Optimal Collection Configuration
The foundation of scalable ingestion is a well-designed collection configuration. For a dataset of this magnitude, the goal is to intelligently balance memory usage, disk I/O, and search performance.
When a collection is too large to hold in memory, each structure takes a `memory` parameter that says where it lives. `pinned` stays on the heap, `cached` is memory-mapped and pre-warmed, and `cold` is memory-mapped and read on demand.
```python
from qdrant_client import QdrantClient, models
@@ -62,76 +62,55 @@ import os
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
client.recreate_collection(
collection_name="laion400m_collection",
client.create_collection(
collection_name="my_collection",
vectors_config=models.VectorParams(
size=512, # CLIP embedding dimensions
size=512,
distance=models.Distance.COSINE,
on_disk=True, # Store original vectors on disk
datatype=models.Datatype.FLOAT16,
memory=models.Memory.COLD,
),
payload=models.PayloadStorageParams(memory=models.Memory.COLD),
quantization_config=models.BinaryQuantization(
binary=models.BinaryQuantizationConfig(
always_ram=True, # Keep quantized vectors in RAM
)
),
optimizers_config=models.OptimizersConfigDiff(
max_segment_size=5_000_000, # Create larger segments for faster search
),
hnsw_config=models.HnswConfigDiff(
m=6, # Lower m to reduce memory usage
on_disk=False # Keep the HNSW index graph in RAM
binary=models.BinaryQuantizationConfig(memory=models.Memory.PINNED),
),
hnsw_config=models.HnswConfigDiff(memory=models.Memory.PINNED),
)
```
This configuration employs several key optimizations:
Two things catch people out. `pinned` is rejected for dense vectors, which support only `cached` or `cold`. And `max_segment_size` and `indexing_threshold` are both measured in kilobytes rather than points. See [Memory Tiers](/documentation/ops-configuration/memory-tiers/) for which tier suits which structure.
- **`on_disk=True`**: This is the most critical setting for large datasets. It instructs Qdrant to store the full-precision original vectors on disk ([memmap storage](/documentation/manage-data/storage/#on-disk-storage)) instead of in RAM, dramatically reducing memory requirements.
`memory` arrived in Qdrant v1.19. If you are following older material, it replaces `on_disk` on the vectors, `on_disk_payload` on the collection, `always_ram` in the quantization config, and `on_disk` in the HNSW config. Those still work but are deprecated, and `pinned` has no equivalent among them.
- **Binary Quantization with `always_ram=True`**: While the original vectors are on disk, we enable [binary quantization](/documentation/manage-data/quantization/#binary-quantization) and force the compressed vectors to remain in RAM. This provides a lightweight in-memory representation for fast initial candidate searches.
## The Upload Process
- **Large Segment Size**: The `max_segment_size` is increased to create fewer, larger segments. This can improve search performance at the cost of slightly slower indexing.
- **In-Memory HNSW Index**: By setting `on_disk=False` for the [HNSW config](/documentation/manage-data/quantization/#hnsw-config), we keep the graph index in RAM. This ensures that navigating vector relationships during a search is extremely fast, avoiding disk latency. The `m` value is lowered to 6 to further conserve memory.
### The Upload Process
With the collection configured, the upload can proceed using a memory-efficient streaming approach. The LAION dataset is split into 409 parts, each containing about 1 million records. The script processes one part at a time, downloading the data, preparing the points, and streaming them to Qdrant.
Hand the method an iterable and it takes care of the requests. Because it accepts an iterator, you can feed it a generator that reads from disk as it goes, rather than materializing the whole set first.
```python
def upload_data_to_qdrant(client, embeddings, metadata, parallel=4):
"""
Uploads data to Qdrant using the upload_collection method.
"""
client.upload_collection(
collection_name="laion400m_collection",
points=zip(range(len(metadata)), embeddings, metadata),
batch_size=256,
parallel=parallel,
show_progress=True,
)
import tqdm
# --- Simplified logic for processing chunks ---
# for part in dataset_parts:
# embeddings, metadata = download_and_process_part(part)
# upload_data_to_qdrant(client, embeddings, metadata)
# cleanup_local_files(part)
client.upload_collection(
collection_name="my_collection",
vectors=embeddings,
payload=payloads,
ids=tqdm.tqdm(ids),
batch_size=256,
parallel=4,
)
```
This method processes the dataset in manageable chunks without ever loading the entire 400 million points into memory. Using `parallel=4` allows the client to upload multiple batches concurrently, saturating the network connection and maximizing ingestion speed.
A few things worth knowing about these parameters:
## The Payoff: An Efficient Architecture at Scale
- **`ids`** must be unique across the whole upload. Writes are [idempotent](/documentation/manage-data/points/#idempotence), so a point sent under an id that already exists overwrites it instead of erroring. That is what you want on a retry, and what bites you if two batches reuse the same numbers.
- **`batch_size`** controls how many points go in each request. The [Bulk Upload guide](/documentation/manage-data/bulk-upload/) covers how to pick it, along with sharding and payload indexes.
- **`parallel`** starts worker processes. Each one opens its own connection, so if batches begin failing after you raise it, drop back to `1`.
- **`tqdm`** around any of the iterables gives you progress. There is no `show_progress` parameter.
- **`prefer_grpc=True`** on the client skips JSON serialization on every batch.
- **`update_mode=models.UpdateMode.INSERT_ONLY`** (v1.17) makes a resumed upload skip points already in the collection rather than rewriting them. See [Update Mode](/documentation/manage-data/points/#update-mode).
- **On Windows and macOS**, `parallel` greater than 1 needs the upload call behind a `if __name__ == "__main__":` guard in a script, because Python starts workers by re-importing your module. Without it the upload hangs instead of failing. Notebooks are unaffected.
This combined strategy of a hybrid storage configuration and streaming ingestion creates a highly efficient system. By keeping only the most essential components in RAM: the quantized vectors and the HNSW index, Qdrant can index and serve a 400 million vector dataset on a machine with just 64GB of RAM. The original vectors, which would consume hundreds of gigabytes, are efficiently accessed from disk only when needed for rescoring top candidates.
Start small and test. Before attempting to upload your entire dataset, ingest a smaller chunk to validate your configuration and process.
This architecture strikes a balance by keeping infrastructure costs low by minimizing RAM usage while maintaining fast and accurate search performance. By understanding and applying these ingestion strategies, you can confidently scale your Qdrant-powered applications to handle real-world data volumes.
> Learn more in a complete hands-on guide in our **[Large-Scale Search tutorial](/documentation/tutorials-operations/large-scale-search/)**.
> **Check out the reference implementation:**
> [qdrant/laion-400m-benchmark on GitHub](https://github.com/qdrant/laion-400m-benchmark)
> This open-source repository includes full scripts for downloading, processing, and uploading the LAION-400M dataset to Qdrant using efficient, production-ready patterns.
> **Want to try this workflow hands-on?**
> Run the [Google Colab notebook](https://colab.research.google.com/drive/1X4EW-nymqcsyhwFYS2ZrmE8MPSnKKj62?usp=sharing) to see large-scale vector ingestion, quantized search, and efficient RAM/disk optimization in action!
Try it hands-on in the [Google Colab notebook](https://colab.research.google.com/github/qdrant/examples/blob/master/course/day_4/large_scale_ingestion.ipynb).
See it at 400 million points in the [LAION-400M benchmark](https://github.com/qdrant/laion-400m-benchmark).
@@ -45,11 +45,6 @@ Learn about the Qdrant ecosystem and integration strategies.
content: Process and vectorize documents from any format
link: /course/essentials/day-7/unstructured/
- icon: /courses/course-integrations/quotient.svg
title: Quotient
content: Advanced analytics with vector data
link: /course/essentials/day-7/quotient/
- icon: /courses/course-integrations/superlinked.svg
title: Superlinked
content: Advanced feature engineering for vectors
@@ -1,127 +0,0 @@
---
title: "Integrating with Quotient"
short_description: "Monitor RAG quality with Quotient AI and Qdrant: detect hallucinations, score document relevance, and debug retrieval-augmented agents."
description: Learn how Quotient and Qdrant combine to deliver end-to-end AI monitoring, hallucination detection, and performance analytics for reliable, high-quality retrieval-augmented generation systems.
weight: 7
isLesson: true
---
{{< date >}} Day 7 {{< /date >}}
# Integrating with Quotient
Advanced analytics with vector data using Quotient platform.
{{< youtube "QeQuCsh1SHs" >}}
## What You'll Learn
- Analytics platform integration
- Vector data analysis techniques
- Business intelligence applications
- Reporting and visualization
- RAG monitoring and quality assurance
- AI application monitoring and debugging
- Hallucination detection and document relevance scoring
## Quotient AI Monitoring Platform
Quotient AI provides critical monitoring capabilities for AI applications and agents, automatically detecting quality issues and providing comprehensive insights into system performance. The platform serves as an essential monitoring layer for RAG (Retrieval Augmented Generation) applications, helping maintain reliability and enabling effective debugging of AI systems.
### Core Architecture
The Quotient AI monitoring platform offers several key components:
- **Automatic Quality Monitoring**: Continuously monitors agent performance and response quality
- **Hallucination Detection**: Identifies when AI responses contain fabricated or incorrect information
- **Document Relevance Scoring**: Evaluates how well retrieved documents match user queries
- **Monitoring Dashboards**: Provides visual insights into system performance and quality metrics
- **Root Cause Analysis**: Offers tools to identify and debug issues in AI applications
### Documentation Q&A Agent with Monitoring
The demonstration showcases building a comprehensive documentation Q&A agent with built-in monitoring using a modular stack:
**Core Components**:
- **Tavilli**: Crawling and extracting documentation from company websites
- **Qdrant**: Fast semantic search over document chunks
- **LangChain**: Orchestrating document processing, retrieval, and LLM calls
- **OpenAI**: Generating answers and creating high-quality embeddings
- **Quotient AI**: Critical monitoring layer for quality assurance
### Implementation Workflow
The complete system follows these steps:
1. **Environment Setup**:
- Configure API keys for OpenAI, Tavilli, and Quotient AI
- Install necessary libraries and dependencies
- Set up monitoring configurations
2. **Core Component Configuration**:
- Qdrant as vector store for document embeddings
- OpenAI for embeddings (`text-embedding-3-small`) and answer generation (`GPT-4o`)
- Quotient for monitoring with hallucination detection and document relevance scoring enabled by default
3. **Content Extraction and Processing**:
- Tavilli extracts real text content from live documentation websites
- Crawls and collects relevant URLs automatically
- Processes structured content for embedding
4. **Document Processing Pipeline**:
- LangChain's recursive character text splitter creates overlapping 700-token chunks
- 50-token overlap preserves context across chunks
- Optimized chunking strategy for embedding models
5. **Embedding and Storage**:
- Chunked documents are embedded using OpenAI's embedding model
- Embeddings and metadata stored in Qdrant collections
- Optimized for fast retrieval and similarity search
6. **RAG Chain Implementation**:
- Orchestrates question-answering process
- Retrieves relevant chunks based on user queries
- Formats retrieved content as context for answer generation
- Constrains answers to use only relevant documentation
7. **Monitoring Integration**:
- Every interaction logged to Quotient AI
- Asynchronous detection pipelines identify potential hallucinations
- Document relevance scoring provides retrieval quality insights
- Real-time monitoring dashboards track system performance
### Key Monitoring Features
**Hallucination Detection**:
- Automatically identifies fabricated or incorrect information in AI responses
- Provides confidence scores for response accuracy
- Enables proactive quality control
**Document Relevance Scoring**:
- Evaluates how well retrieved documents match user queries
- Identifies retrieval quality issues
- Helps optimize document chunking and embedding strategies
**Performance Dashboards**:
- Visual insights into system performance metrics
- Quality trends and anomaly detection
- Root cause analysis tools for debugging
### Real-World Applications
This monitoring architecture enables various enterprise use cases:
- **Enterprise Documentation**: Monitoring Q&A systems for technical documentation
- **Customer Support**: Quality assurance for AI-powered customer service agents
- **Content Management**: Monitoring AI systems that process and respond to content queries
- **Research Applications**: Ensuring accuracy in AI-powered research and analysis tools
## Resources
- [Optimizing RAG Through an Evaluation-Based Methodology](/articles/rapid-rag-optimization-with-qdrant-and-quotient/):
Learn how to optimize RAG systems using Qdrant and Quotient through systematic evaluation. Covers experimentation with chunking, retrieval strategies, and model selection.
- [Building High-Quality RAG Applications with Qdrant and Quotient](https://blog.quotientai.co/building-high-quality-rag-applications-with-qdrant-and-quotient/):
Comprehensive guide on building production-ready RAG applications with comprehensive monitoring using Quotient AI and Qdrant for quality assurance and performance tracking.
**Note**: Visit [quotientai.co](https://www.quotientai.co/) for more information.
@@ -3,6 +3,7 @@ title: "Multi-Vector Search Course"
page_title: "Qdrant Multi-Vector Search Course"
short_description: "Master multi-vector search with ColBERT and ColPali: late interaction, MaxSim scoring, multi-stage retrieval, MUVERA indexing, and evaluation."
description: Master late interaction models, ColPali, and production optimization. Build scalable multi-vector search pipelines.
weight: 20
content:
sidebarTitle: "Multi-Vector Search Course"
menuTitle:
@@ -10,6 +10,18 @@ isLesson: true
# Qdrant Setup
<div class="video">
<iframe
src="https://www.youtube.com/embed/-8TpooWX8kE?rel=0"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
allowfullscreen>
</iframe>
</div>
<br/>
Before diving into multi-vector search, you need a running Qdrant instance. Whether you choose Qdrant Cloud for a managed solution or a local deployment, this lesson will get you up and running.
Multi-vector search requires specific collection configurations that differ from traditional single-vector setups. We'll cover the essentials to prepare your environment.