converted section dropdowns to pages, updated Build page nav and context

This commit is contained in:
kanungle
2026-01-05 20:37:02 -08:00
parent 402f930956
commit 0af8d62f9f
34 changed files with 177 additions and 50 deletions
@@ -1,7 +1,8 @@
---
title: Ecosystem & Integrations
weight: 22
is_empty: false
weight: 36
is_empty: true
hideInSidebar: true
aliases:
- how-to
- tutorials
@@ -0,0 +1,315 @@
---
title: Data Ingestion for Beginners
weight: 2
#partition: build
social_preview_image: /documentation/examples/data-ingestion-beginners/social_preview.png
aliases:
- /documentation/data-ingestion-beginners/
---
![data-ingestion-beginners-7](/documentation/examples/data-ingestion-beginners/data-ingestion-7.png)
# Send S3 Data to Qdrant Vector Store with LangChain
| Time: 30 min | Level: Beginner | | |
| --- | ----------- | ----------- |----------- |
**Data ingestion into a vector store** is essential for building effective search and retrieval algorithms, especially since nearly 80% of data is unstructured, lacking any predefined format.
In this tutorial, we’ll create a streamlined data ingestion pipeline, pulling data directly from **AWS S3** and feeding it into Qdrant. We’ll dive into vector embeddings, transforming unstructured data into a format that allows you to search documents semantically. Prepare to discover new ways to uncover insights hidden within unstructured data!
## Ingestion Workflow Architecture
We’ll set up a powerful document ingestion and analysis pipeline in this workflow using cloud storage, natural language processing (NLP) tools, and embedding technologies. Starting with raw data in an S3 bucket, we'll preprocess it with LangChain, apply embedding APIs for both text and images and store the results in Qdrant – a vector database optimized for similarity search.
**Figure 1: Data Ingestion Workflow Architecture**
![data-ingestion-beginners-5](/documentation/examples/data-ingestion-beginners/data-ingestion-5.png)
Let's break down each component of this workflow:
- **S3 Bucket:** This is our starting point—a centralized, scalable storage solution for various file types like PDFs, images, and text.
- **LangChain:** Acting as the pipeline’s orchestrator, LangChain handles extraction, preprocessing, and manages data flow for embedding generation. It simplifies processing PDFs, so you won’t need to worry about applying OCR (Optical Character Recognition) here.
- **Qdrant:** As your vector database, Qdrant stores embeddings and their [payloads](https://qdrant.tech/documentation/concepts/payload/), enabling efficient similarity search and retrieval across all content types.
## Prerequisites
![data-ingestion-beginners-11](/documentation/examples/data-ingestion-beginners/data-ingestion-11.png)
In this section, you’ll get a step-by-step guide on ingesting data from an S3 bucket. But before we dive in, let’s make sure you’re set up with all the prerequisites:
| | |
| -------------- | ------------------------------------------------------------------------------------------------------------------------ |
| Sample Data | We’ll use a sample dataset, where each folder includes product reviews in text format along with corresponding images. |
| AWS Account | An active [AWS account](https://aws.amazon.com/free/) with access to S3 services. |
| Qdrant Cloud | A [Qdrant Cloud account](https://cloud.qdrant.io) with access to the WebUI for managing collections and running queries. |
| LangChain | You will use this [popular framework](https://www.langchain.com) to tie everything together. |
#### Supported Document Types
The documents used for ingestion can be of various types, such as PDFs, text files, or images. We will organize a structured S3 bucket with folders with the supported document types for testing and experimentation.
#### Python Environment
Ensure you have a Python environment (Python 3.9 or higher) with these libraries installed:
```python
boto3
langchain-community
langchain
python-dotenv
unstructured
unstructured[pdf]
qdrant_client
fastembed
```
---
**Access Keys:** Store your AWS access key, S3 secret key, and Qdrant API key in a .env file for easy access. Here’s a sample `.env` file.
```text
ACCESS_KEY = ""
SECRET_ACCESS_KEY = ""
QDRANT_KEY = ""
```
---
<aside role="alert"> Although the code includes support for processing PDFs, the sample data currently has no PDF files included. </aside>
## Step 1: Ingesting Data from S3
![data-ingestion-beginners-9.png](/documentation/examples/data-ingestion-beginners/data-ingestion-9.png)
The LangChain framework makes it easy to ingest data from storage services like AWS S3, with built-in support for loading documents in formats such as PDFs, images, and text files.
To connect LangChain with S3, you’ll use the `S3DirectoryLoader`, which lets you load files directly from an S3 bucket into LangChain’s pipeline.
### Example: Configuring LangChain to Load Files from S3
Here’s how to set up LangChain to ingest data from an S3 bucket:
```python
from langchain_community.document_loaders import S3DirectoryLoader
# Initialize the S3 document loader
loader = S3DirectoryLoader(
"product-dataset", # S3 bucket name
"p_1", #S3 Folder name containing the data for the first product
aws_access_key_id=aws_access_key_id, # AWS Access Key
aws_secret_access_key=aws_secret_access_key # AWS Secret Access Key
)
# Load documents from the specified S3 bucket
docs = loader.load()
```
---
## Step 2. Turning Documents into Embeddings
[Embeddings](/articles/what-are-embeddings/) are the secret sauce here—they’re numerical representations of data (like text, images, or audio) that capture the “meaning” in a form that’s easy to compare. By converting text and images into embeddings, you’ll be able to perform similarity searches quickly and efficiently. Think of embeddings as the bridge to storing and retrieving meaningful insights from your data in Qdrant.
### Models We’ll Use for Generating Embeddings
To get things rolling, we’ll use two powerful models:
1. **`sentence-transformers/all-MiniLM-L6-v2` Embeddings** for transforming text data.
2. **`CLIP` (Contrastive Language-Image Pretraining)** for image data.
---
### Document Processing Function
![data-ingestion-beginners-8.png](/documentation/examples/data-ingestion-beginners/data-ingestion-8.png)
Next, we’ll define two functions — `process_text` and `process_image` to handle different file types in our document pipeline. The `process_text` function extracts and returns the raw content from a text-based document, while `process_image` retrieves an image from an S3 source and loads it into memory.
```python
from PIL import Image
def process_text(doc):
source = doc.metadata['source'] # Extract document source (e.g., S3 URL)
text = doc.page_content # Extract the content from the text file
print(f"Processing text from {source}")
return source, text
def process_image(doc):
source = doc.metadata['source'] # Extract document source (e.g., S3 URL)
print(f"Processing image from {source}")
bucket_name, object_key = parse_s3_url(source) # Parse the S3 URL
response = s3.get_object(Bucket=bucket_name, Key=object_key) # Fetch image from S3
img_bytes = response['Body'].read()
img = Image.open(io.BytesIO(img_bytes))
return source, img
```
### Helper Functions for Document Processing
To retrieve images from S3, a helper function `parse_s3_url` breaks down the S3 URL into its bucket and critical components. This is essential for fetching the image from S3 storage.
```python
def parse_s3_url(s3_url):
parts = s3_url.replace("s3://", "").split("/", 1)
bucket_name = parts[0]
object_key = parts[1]
return bucket_name, object_key
```
---
## Step 3: Loading Embeddings into Qdrant
![data-ingestion-beginners-10](/documentation/examples/data-ingestion-beginners/data-ingestion-10.png)
Now that your documents have been processed and converted into embeddings, the next step is to load these embeddings into Qdrant.
### Creating a Collection in Qdrant
In Qdrant, data is organized in collections, each representing a set of embeddings (or points) and their associated metadata (payload). To store the embeddings generated earlier, you’ll first need to create a collection.
Here’s how to create a collection in Qdrant to store both text and image embeddings:
```python
def create_collection(collection_name):
qdrant_client.create_collection(
collection_name,
vectors_config={
"text_embedding": models.VectorParams(
size=384, # Dimension of text embeddings
distance=models.Distance.COSINE, # Cosine similarity is used for comparison
),
"image_embedding": models.VectorParams(
size=512, # Dimension of image embeddings
distance=models.Distance.COSINE, # Cosine similarity is used for comparison
),
},
)
create_collection("products-data")
```
---
This function creates a collection for storing text (384 dimensions) and image (512 dimensions) embeddings, using cosine similarity to compare embeddings within the collection.
Once the collection is set up, you can load the embeddings into Qdrant. This involves inserting (or updating) the embeddings and their associated metadata (payload) into the specified collection.
Here’s the code for loading embeddings into Qdrant:
```python
def ingest_data(points):
operation_info = qdrant_client.upsert(
collection_name="products-data", # Collection where data is being inserted
points=points
)
return operation_info
```
---
**Explanation of Ingestion**
1. **Upserting the Data Point:** The upsert method on the `qdrant_client` inserts each PointStruct into the specified collection. If a point with the same ID already exists, it will be updated with the new values.
2. **Operation Info:** The function returns `operation_info`, which contains details about the upsert operation, such as success status or any potential errors.
**Running the Ingestion Code**
Here’s how to call the function and ingest data:
```python
from qdrant_client import models
if __name__ == "__main__":
collection_name = "products-data"
create_collection(collection_name)
for i in range(1,6): # Five documents
folder = f"p_{i}"
loader = S3DirectoryLoader(
"product-dataset",
folder,
aws_access_key_id=aws_access_key_id,
aws_secret_access_key=aws_secret_access_key
)
docs = loader.load()
points, text_review, product_image = [], "", ""
for idx, doc in enumerate(docs):
source = doc.metadata['source']
if source.endswith(".txt") or source.endswith(".pdf"):
_text_review_source, text_review = process_text(doc)
elif source.endswith(".png"):
product_image_source, product_image = process_image(doc)
if text_review:
point = models.PointStruct(
id=idx, # Unique identifier for each point
vector={
"text_embedding": models.Document(
text=text_review, model="sentence-transformers/all-MiniLM-L6-v2"
),
"image_embedding": models.Image(
image=product_image, model="Qdrant/clip-ViT-B-32-vision"
),
},
payload={"review": text_review, "product_image": product_image_source},
)
points.append(point)
operation_info = ingest_data(points)
print(operation_info)
```
The `PointStruct` is instantiated with these key parameters:
- **id:** A unique identifier for each embedding, typically an incremental index.
- **vector:** A dictionary holding the text and image inputs to be embedded. `qdrant-client` uses [FastEmbed](https://github.com/qdrant/fastembed) under the hood to automatically generate vector representations from these inputs locally.
- **payload:** A dictionary storing additional metadata, like product reviews and image references, which is invaluable for retrieval and context during searches.
The code dynamically loads folders from an S3 bucket, processes text and image files separately, and stores their embeddings and associated data in dedicated lists. It then creates a `PointStruct` for each data entry and calls the ingestion function to load it into Qdrant.
### Exploring the Qdrant WebUI Dashboard
Once the embeddings are loaded into Qdrant, you can use the WebUI dashboard to visualize and manage your collections. The dashboard provides a clear, structured interface for viewing collections and their data. Let’s take a closer look in the next section.
## Step 4: Visualizing Data in Qdrant WebUI
To start visualizing your data in the Qdrant WebUI, head to the **Overview** section and select **Access the database**.
**Figure 2: Accessing the Database from the Qdrant UI**
![data-ingestion-beginners-2.png](/documentation/examples/data-ingestion-beginners/data-ingestion-2.png)
When prompted, enter your API key. Once inside, you’ll be able to view your collections and the corresponding data points. You should see your collection displayed like this:
**Figure 3: The product-data Collection in Qdrant**
![data-ingestion-beginners-4.png](/documentation/examples/data-ingestion-beginners/data-ingestion-4.png)
Here’s a look at the most recent point ingested into Qdrant:
**Figure 4: The Latest Point Added to the product-data Collection**
![data-ingestion-beginners-6.png](/documentation/examples/data-ingestion-beginners/data-ingestion-6.png)
The Qdrant WebUI’s search functionality allows you to perform vector searches across your collections. With options to apply filters and parameters, retrieving relevant embeddings and exploring relationships within your data becomes easy. To start, head over to the **Console** in the left panel, where you can create queries:
**Figure 5: Overview of Console in Qdrant**
![data-ingestion-beginners-1.png](/documentation/examples/data-ingestion-beginners/data-ingestion-1.png)
The first query retrieves all collections, the second fetches points from the product-data collection, and the third performs a sample query. This demonstrates how straightforward it is to interact with your data in the Qdrant UI.
Now, let’s retrieve some documents from the database using a query!.
**Figure 6: Querying the Qdrant Client to Retrieve Relevant Documents**
![data-ingestion-beginners-3.png](/documentation/examples/data-ingestion-beginners/data-ingestion-3.png)
In this example, we queried **Phones with improved design**. Then, we converted the text to vectors using OpenAI and retrieved a relevant phone review highlighting design improvements.
## Conclusion
In this guide, we set up an S3 bucket, ingested various data types, and stored embeddings in Qdrant. Using LangChain, we dynamically processed text and image files, making it easy to work with each file type.
Now, it’s your turn. Try experimenting with different data types, such as videos, and explore Qdrant’s advanced features to enhance your applications. To get started, [sign up](https://cloud.qdrant.io/signup) for Qdrant today.
![data-ingestion-beginners-12](/documentation/examples/data-ingestion-beginners/data-ingestion-12.png)
@@ -0,0 +1,272 @@
---
title: Automating Processes with Qdrant and n8n
weight: 7
#partition: build
social_preview_image: /documentation/examples/qdrant-n8n-2/preview/social_preview.png
aliases:
- /blog/qdrant-n8n-beyond-simple-similarity-search/
- /documentation/qdrant-n8n/
---
![n8n-qdrant](/documentation/examples/qdrant-n8n-2/cover.png)
# Automating Processes with Qdrant and n8n beyond simple RAG
| Time: 45 min | Level: Intermediate |
| --- | ----------- |
This tutorial shows how to combine Qdrant with [n8n](https://n8n.io/) low-code automation platform to cover **use cases beyond basic Retrieval-Augmented Generation (RAG)**. You'll learn how to use vector search for **recommendations** and **unstructured big data analysis**.
<aside role="status">
Since this tutorial was created, <a href="https://qdrant.tech/documentation/platforms/n8n/">an official Qdrant node for n8n</a> has been released. It simplifies workflows and replaces the HTTP request nodes used in the examples below. Watch <a href="https://youtu.be/sYP_kHWptHY"> a quick video introduction</a> to it.
</aside>
## Setting Up Qdrant in n8n
To start using Qdrant with n8n, you need to provide your Qdrant instance credentials in the [credentials](https://docs.n8n.io/integrations/builtin/credentials/qdrant/#using-api-key) tab. Select `QdrantApi` from the list.
### Qdrant Cloud
To connect [Qdrant Cloud](https://qdrant.tech/documentation/cloud/) to n8n:
1. Open the [Cloud Dashboard](https://qdrant.to/cloud) and select a cluster.
2. From the **Cluster Details**, copy the `Endpoint` address—this will be used as the `Qdrant URL` in n8n.
3. Navigate to the **API Keys** tab and copy your API key—this will be the `API Key` in n8n.
For a walkthrough, see this [step-by-step video guide](https://youtu.be/fYMGpXyAsfQ?feature=shared&t=177).
### Local Mode
For a fully local experimnets-driven setup, a valuable option is n8n's [Self-hosted AI Starter Kit](https://github.com/n8n-io/self-hosted-ai-starter-kit). This is an open-source Docker Compose template for local AI & low-code development environment.
This kit includes a [local instance of Qdrant](https://qdrant.tech/documentation/quickstart/). To get started:
1. Follow the instructions in the repository to install the AI Starter Kit.
2. Use the values from the `docker-compose.yml` file to fill in the connection details.
<aside role="status">
Remember to update to the latest Qdrant Docker image using <code>docker-compose pull</code>.
</aside>
The default Qdrant configuration in AI Starter Kit's `docker-compose.yml` looks like this:
```yaml
qdrant:
image: qdrant/qdrant
hostname: qdrant
container_name: qdrant
networks: ['demo']
restart: unless-stopped
ports:
- 6333:6333
volumes:
- qdrant_storage:/qdrant/storage
```
From this configuration, the `Qdrant URL` in n8n Qdrant credentials is `http://qdrant:6333/`.
To set up a local Qdrant API key, add the following lines to the YAML file:
```yaml
qdrant:
...
volumes:
- qdrant_storage:/qdrant/storage
environment:
- QDRANT_API_KEY=test
```
After saving the configuration and running the Starter Kit, use `QDRANT_API_KEY` value (e.g., `test`) as the `API Key` and `http://qdrant:6333/` as the `Qdrant URL`.
## Qdrant + n8n Beyond Simple Similarity Search
Vector search's ability to determine semantic similarity between objects is often used to address models' hallucinations, powering the memory of Retrieval-Augmented Generation-based applications. Yet there's more to vector search than just a "knowledge base" role.
The combination of similarity and dissimilarity metrics in vector space expands vector search to recommendations, discovery search, and large-scale unstructured data analysis.
![overview](/documentation/examples/qdrant-n8n-2/overview.png)
### Recommendations
When searching for new music, films, books, or food, it can be difficult to articulate exactly what we want. Instead, we often rely on discovering new content through comparison to examples of what we like or dislike.
The [Qdrant Recommendation API](https://qdrant.tech/articles/new-recommendation-api/) is built to make these discovery searches possible by using positive and negative examples as anchors. It helps find new relevant results based on your preferences.
![recommendations](/documentation/examples/qdrant-n8n-2/recommendations.png)
#### Movie Recommendations
Imagine a home cinema night—you've already watched Harry Potter 666 times and crave a new series featuring young wizards. Your favorite streaming service repetitively recommends all seven parts of the millennial saga. Frustrated, you turn to n8n to create an **Agentic Movie Recommendation tool**.
**Setup:**
1. **Dataset**: We use movie descriptions from the [IMDB Top 1000 Kaggle dataset](https://www.kaggle.com/datasets/omarhanyy/imdb-top-1000).
2. **Embedding Model**: We'll use OpenAI `text-embedding-3-small`, but you can opt for any other suitable embedding model.
**Workflow:**
A [Template Agentic Movie Recommendation Workflow](https://n8n.io/workflows/2440-building-rag-chatbot-for-movie-recommendations-with-qdrant-and-open-ai/) consists of three parts:
1. **Movie Data Uploader**: Embeds movie descriptions and uploads them to Qdrant using the [Qdrant Vector Store Node](https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.vectorstoreqdrant) (now this can also be done using the [official Qdrant Node for n8n](https://github.com/qdrant/n8n-nodes-qdrant)). In the template workflow, the dataset is fetched from GitHub, but you can use any supported storage, for example [Google Cloud Storage node](https://docs.n8n.io/integrations/builtin/app-nodes/n8n-nodes-base.googlecloudstorage).
2. **AI Agent**: Uses the [AI Agent Node](https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.agent) to formulate Recommendation API calls based on your natural language requests. Choose an LLM as a "brain" and define a [JSON schema](https://docs.n8n.io/integrations/builtin/cluster-nodes/sub-nodes/n8n-nodes-langchain.toolworkflow/#specify-input-schema) for the recommendations tool powered by Qdrant. This schema lets the LLM map your requests to the tool input format.
3. **Recommendations Tool**: A [subworkflow](https://docs.n8n.io/flow-logic/subworkflows/) that calls the Qdrant Recommendation API using the [HTTP Request Node](https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.httprequest) (now this can also be done using the [official Qdrant Node for n8n](https://github.com/qdrant/n8n-nodes-qdrant)). The agent extracts relevant and irrelevant movie descriptions from your chat message and passes them to the tool. The tool embeds them with `text-embedding-3-small` and uses the Qdrant Recommendation API to get movie recommendations, which are passed back to the agent.
Set it up, run a chat and ask for "*something about wizards but not Harry Potter*."
What results do you get?
---
If you'd like a detailed walkthrough of building this workflow step-by-step, watch the video below:
<iframe width="560" height="315" src="https://www.youtube.com/embed/O5mT8M7rqQQ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
This recommendation scenario is easily adaptable to any language or data type (images, audio, video).
### Big Data Analysis
The ability to map data to a vector space that reflects items' similarity and dissimilarity relationships provides a range of mathematical tools for data analysis.
Vector search dedicated solutions are built to handle billions of data points and quickly compute distances between them, simplifying **clustering, classification, dissimilarity sampling, deduplication, interpolation**, and **anomaly detection at scale**.
The combination of this vector search feature with automation tools like n8n creates production-level solutions capable of monitoring data temporal shifts, managing data drift, and discovering patterns in seemingly unstructured data.
A practical example is worth a thousand words. Let's look at **Qdrant-based anomaly detection and classification tools**, which are designed to be used by the [n8n AI Agent node](https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.agent) for data analysis automation.
To make it more interesting, this time we'll focus on image data.
#### Anomaly Detection Tool
One definition of "anomaly" comes intuitively after projecting vector representations of data points into a 2D space—[Qdrant webUI](https://qdrant.tech/documentation/web-ui/) provides this functionality.
Points that don't belong to any clusters are more likely to be anomalous.
![anomalies-on-2D](/documentation/examples/qdrant-n8n-2/anomalies-2D.png)
With that intuition comes the recipe for building an anomaly detection tool. We will demonstrate it on anomaly detection in agricultural crops. Qdrant will be used to:
1. Store vectorized images.
2. Identify a "center" (representative) for each crop cluster.
3. Define the borders of each cluster.
4. Check if new images fall within these boundaries. If an image does not fit within any cluster, it is flagged as anomalous. Alternatively, you can check if an image is anomalous to a specific cluster.
![anomaly-detection](/documentation/examples/qdrant-n8n-2/anomaly-detection.png)
**Setup:**
1. **Dataset**: We use the [Agricultural Crops Image Classification dataset](https://www.kaggle.com/datasets/mdwaquarazam/agricultural-crops-image-classification).
2. **Embedding Model**: The [Voyage AI multimodal embedding model](https://docs.voyageai.com/docs/multimodal-embeddings). It can project images and text data into a shared vector space.
**1. Uploading Images to Qdrant**
Since the [Qdrant Vector Store node](https://docs.n8n.io/integrations/builtin/cluster-nodes/root-nodes/n8n-nodes-langchain.vectorstoreqdrant/) does not support embedding models outside the predefined list (which doesn't include Voyage AI), we embed and upload data to Qdrant via direct API calls in [HTTP Request nodes](https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.httprequest/).
With the release of the [official Qdrant node](https://github.com/qdrant/n8n-nodes-qdrant), which supports arbitrary vectorized input, the HTTP Request node can now be replaced with this native integration.
**Workflow:**
*There are three workflows: (1) Uploading images to Qdrant (2) Setting up cluster centers and thresholds (3) Anomaly detection tool itself.*
An [1/3 Uploading Images to Qdrant Template Workflow](https://n8n.io/workflows/2654-vector-database-as-a-big-data-analysis-tool-for-ai-agents-13-anomaly12-knn/) consists of the following blocks:
1. **Check Collection**: Verifies if a collection with the specified name exists in Qdrant. If not, it creates one.
2. **Payload Index**: Adds a [payload index](https://qdrant.tech/documentation/concepts/indexing/#payload-index) on the `crop_name` payload (metadata) field. This field stores crop class labels, and indexing it improves the speed of filterable searches in Qdrant. It changes the way a vector index is constructed, adapting it for fast vector search under filtering constraints. For more details, refer to this [guide on filtering in Qdrant](https://qdrant.tech/articles/vector-search-filtering/).
3. **Fetch Images**: Fetches images from Google Cloud Storage using the [Google Cloud Storage node](https://docs.n8n.io/integrations/builtin/app-nodes/n8n-nodes-base.googlecloudstorage).
4. **Generate IDs**: Assigns UUIDs to each data point.
5. **Embed Images**: Embeds the images using the Voyage API.
6. **Batch Upload**: Uploads the embeddings to Qdrant in batches.
**2. Defining a Cluster Representative**
We used two approaches (it's not an exhaustive list) to defining a cluster representative, depending on the availability of labeled data:
| Method | Description |
|----------------------|-----------------------------------------------------------------------------|
| **Medoids** | A point within the cluster that has the smallest total distance to all other cluster points. This approach needs labeled data for each cluster. |
| **Perfect Representative** | A representative defined by a textual description of the ideal cluster member—the multimodality of Voyage AI embeddings allows for this trick. For example, for cherries: *"Small, glossy red fruits on a medium-sized tree with slender branches and serrated leaves."* The closest image to this description in the vector space is selected as the representative. This method requires experimentation to align descriptions with real data. |
**Workflow:**
Both methods are demonstrated in the [2/3 Template Workflow for Anomaly Detection](https://n8n.io/workflows/2655-vector-database-as-a-big-data-analysis-tool-for-ai-agents-23-anomaly/).
| **Method** | **Steps** |
|------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **Medoids** | 1. Sample labeled cluster points from Qdrant. <br> 2. Compute a **pairwise distance matrix** for the cluster using Qdrant's [Distance Matrix API](https://qdrant.tech/documentation/concepts/explore/?q=distance+#distance-matrix). This API helps with scalable cluster analysis and data points relationship exploration. Learn more in [this article](https://qdrant.tech/articles/distance-based-exploration/). <br> 3. For each point, calculate the sum of its distances to all other points. The point with the smallest total distance (or highest similarity for COSINE distance metric) is the medoid. <br> 4. Mark this point as the cluster representative. |
| **Perfect Representative** | 1. Define textual descriptions for each cluster (e.g., AI-generated). <br> 2. Embed these descriptions using Voyage. <br> 3. Find the image embedding closest to the description one. <br> 4. Mark this image as the cluster representative. |
**3. Defining the Cluster Border**
**Workflow:**
The approach demonstrated in [2/3 Template Workflow for Anomaly Detection](https://n8n.io/workflows/2655-vector-database-as-a-big-data-analysis-tool-for-ai-agents-23-anomaly/) works similarly for both types of cluster representatives.
1. Within a cluster, identify the furthest data point from the cluster representative (it can also be the 2nd or Xth furthest point; the best way to define it is through experimentation—for us, the 5th furthest point worked well). Since we use COSINE similarity, this is equivalent to the most similar point to the [opposite](https://mathinsight.org/image/vector_opposite) of the cluster representative (its vector multiplied by -1).
2. Save the distance between the representative and respective furthest point as the cluster border (threshold).
**4. Anomaly Detection Tool**
**Workflow:**
With the preparatory steps complete, you can set up the anomaly detection tool, demonstrated in the [3/3 Template Workflow for Anomaly Detection](https://n8n.io/workflows/2656-vector-database-as-a-big-data-analysis-tool-for-ai-agents-33-anomaly/).
Steps:
1. Choose the method of the cluster representative definition.
2. Fetch all the clusters to compare the candidate image against.
3. Using Voyage AI, embed the candidate image in the same vector space.
4. Calculate the candidate's similarity to each cluster representative. The image is flagged as anomalous if the similarity is below the threshold for all clusters (outside the cluster borders). Alternatively, you can check if it's anomalous to a particular cluster, for example, the cherries one.
---
Anomaly detection in image data has diverse applications, including:
- Moderation of advertisements.
- Anomaly detection in vertical farming.
- Quality control in the food industry, such as [detecting anomalies in coffee beans](https://qdrant.tech/articles/detecting-coffee-anomalies/).
- Identifying anomalies in map tiles for tasks like automated map updates or ecological monitoring.
This tool is easily adaptable to these use cases.
#### Classification Tool
The anomaly detection tool can also be used for classification, but there's a simpler approach: K-Nearest Neighbors (KNN) classification.
> "Show me your friends, and I will tell you who you are."
![KNN-2D](/documentation/examples/qdrant-n8n-2/classification.png)
The KNN method labels a data point by analyzing its classified neighbors and assigning this point the majority class in the neighborhood. This approach doesn't require all data points to be labeled—a subset of labeled examples can serve as anchors to propagate labels across the dataset.
Let's build a KNN-based image classification tool.
**Setup**
1. **Dataset**: We'll use the [Land-Use Scene Classification dataset](https://www.kaggle.com/datasets/apollo2506/landuse-scene-classification). Satellite imagery analysis has applications in ecology, rescue operations, and map updates.
2. **Embedding Model**: As for anomaly detection, we'll use the [Voyage AI multimodal embedding model](https://docs.voyageai.com/docs/multimodal-embeddings).
Additionally, it's good to have test and validation data to determine the optimal value of K for your dataset.
**Workflow:**
Uploading images to Qdrant can be done using the same workflow—[1/3 Uploading Images to Qdrant Template Workflow](https://n8n.io/workflows/2654-vector-database-as-a-big-data-analysis-tool-for-ai-agents-13-anomaly12-knn/), just by swapping the dataset.
The [KNN-Classification Tool Template](https://n8n.io/workflows/2657-vector-database-as-a-big-data-analysis-tool-for-ai-agents-22-knn/) has the following steps:
1. **Embed Image**: Embeds the candidate for classification using Voyage.
2. **Fetch neighbors**: Retrieves the K closest labeled neighbors from Qdrant.
3. **Majority Voting**: Determines the prevailing class in the neighborhood by simple majority voting.
4. **Optional: Ties Resolving**: In case of ties, expands the neighborhood radius.
Of course, this is a simple solution, and there exist more advanced approaches with higher precision & no need for labeled data—for example, you could try [metric learning with Qdrant](https://qdrant.tech/articles/metric-learning-tips/).
Though classification seems like a task that was solved in machine learning decades ago, it's not so trivial to deal with in production. Issues like data drift, shifting class definitions, mislabeled data, and fuzzy differences between classes create unexpected problems, which require continuous adjustments of classifiers, and vector search can be an unusual but effective solution, due to its scalability.
#### Live Walkthrough
To see how n8n agents use these tools in practice, and to revisit the main ideas of the "*Big Data Analysis*" section, watch our integration webinar:
<iframe width="560" height="315" src="https://www.youtube.com/embed/_BQTnXpuH-E" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
## Conclusion
Vector search is not limited to similarity search or basic RAG. When combined with automation platforms like n8n, it becomes a powerful tool for building smarter systems. Think dynamic routing in customer support, content moderation based on user behavior, or AI-driven alerts in data monitoring dashboards.
This tutorial showed how to use Qdrant and n8n for AI-backed recommendations, classification, and anomaly detection. But that's just the start—try vector search for:
- **Deduplication**
- **Dissimilarity search**
- **Diverse sampling**
With Qdrant and n8n, there's plenty of room to create something unique!