mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-29 07:58:31 +02:00
add scaleway and ovhcloud tutorials
This commit is contained in:
@@ -1,51 +1,55 @@
|
||||
---
|
||||
title: Movie Recommendation System
|
||||
title: Movie Recommendation System on OVHcloud
|
||||
weight: 34
|
||||
aliases:
|
||||
- /documentation/tutorials/recommendation-system-ovhcloud/
|
||||
---
|
||||
|
||||
# Build a Movie Recommendation System
|
||||
# Movie Recommendation System on OVHcloud
|
||||
|
||||
| Time: 120 min | Level: Advanced | Output: [GitHub](https://github.com/infoslack/qdrant-example/blob/main/HC-demo/HC-OVH.ipynb) |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
This notebook aims to create a recommendation system using the MovieLens dataset and Qdrant. Vector databases like Qdrant are crucial for storing high-dimensional data, such as user and item embeddings, enabling personalized recommendations by quickly retrieving similar users or items based on advanced indexing techniques. We'll leverage collaborative filtering with a MovieLens dataset, identifying similar users based on ratings represented as vectors in Qdrant, and suggesting movies they liked but we haven't seen yet. The suggested items or content should closely align with the user's interests, leading to more personalized and relevant recommendations.
|
||||
In this tutorial, you will build a mechanism that recommends movies based on defined preferences. Vector databases like Qdrant are good for storing high-dimensional data, such as user and item embeddings. They can enable personalized recommendations by quickly retrieving similar entries based on advanced indexing techniques. In this specific case, we will use [sparse vectors](/articles/sparse-vectors/) to create an efficient and accurate recommendation system.
|
||||
|
||||
Collaborative filtering works on the principle that users with similar tastes will enjoy similar movies. To implement this, we'll represent each user's ratings as vectors in a high-dimensional space using Qdrant. By indexing these vectors, we can find users with similar tastes to ours and recommend movies they liked but we haven't seen yet.
|
||||
**Privacy and Sovereignty:** Since preference data is proprietary, it should be stored in a secure and controlled environment. Our vector database can easily be hosted on [OVHcloud](https://ovhcloud.com/), our trusted [Qdrant Hybrid Cloud](/documentation/hybrid-cloud/) partner. This means that Qdrant can be run from your OVHcloud region, but the database itself can still be managed from within Qdrant Cloud's interface. Both products have been tested for compatibility and scalability, and we recommend their [managed Kubernetes](https://www.ovhcloud.com/en/public-cloud/kubernetes/) service.
|
||||
|
||||
> To see the entire output, use our [notebook with complete instructions](https://github.com/infoslack/qdrant-example/blob/main/HC-demo/HC-OVH.ipynb).
|
||||
|
||||
## Components
|
||||
|
||||
- **Dataset:** [Red Hat Interactive Learning Portal](https://developers.redhat.com/learn)
|
||||
- **Vector DB:** [Qdrant Hybrid Cloud](https://qdrant.tech) running on OpenShift.
|
||||
- **Web Host:** [OVHcloud](https://haystack.deepset.ai/)
|
||||
- **Dataset:** The [MovieLens dataset](https://grouplens.org/datasets/movielens/) contains a list of movies and ratings given by users.
|
||||
- **Cloud:** [OVHcloud](https://ovhcloud.com/), with managed Kubernetes.
|
||||
- **Vector DB:** [Qdrant Hybrid Cloud](https://qdrant.tech) running on [OVHcloud](https://ovhcloud.com/).
|
||||
|
||||
**Methodology:** We're adopting a collaborative filtering approach to construct a recommendation system from the dataset provided. Collaborative filtering works on the premise that if two users share similar tastes, they're likely to enjoy similar movies. Leveraging this concept, we'll identify users whose ratings align closely with ours, and explore the movies they liked but we haven't seen yet. To do this, we'll represent each user's ratings as a vector in a high-dimensional, sparse space. Using Qdrant, we'll index these vectors and search for users whose ratings vectors closely match ours. Ultimately, we will see which movies were enjoyed by users similar to us.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
First, download and unzip the MovieLens dataset into a local directory.
|
||||
Download and unzip the MovieLens dataset:
|
||||
|
||||
```bash
|
||||
```shell
|
||||
mkdir -p data
|
||||
wget https://files.grouplens.org/datasets/movielens/ml-1m.zip
|
||||
unzip ml-1m.zip -d data
|
||||
```
|
||||
|
||||
The necessary Python libraries are installed using `pip`, including `pandas` for data manipulation, `qdrant-client` for interfacing with Qdrant, and `python-dotenv` for managing environment variables.
|
||||
The necessary * libraries are installed using `pip`, including `pandas` for data manipulation, `qdrant-client` for interfacing with Qdrant, and `*-dotenv` for managing environment variables.
|
||||
|
||||
```python
|
||||
!pip install -U \
|
||||
pandas \
|
||||
qdrant-client \
|
||||
python-dotenv
|
||||
*-dotenv
|
||||
```
|
||||
|
||||
The `.env` file is used to store sensitive information like the Qdrant host URL and API key securely.
|
||||
|
||||
```bash
|
||||
```shell
|
||||
QDRANT_HOST
|
||||
QDRANT_API_KEY
|
||||
```
|
||||
Load all environment variables into the setup.
|
||||
Load all environment variables into the setup:
|
||||
|
||||
```python
|
||||
import os
|
||||
@@ -55,57 +59,48 @@ load_dotenv('./.env')
|
||||
|
||||
## Implementation
|
||||
|
||||
Load the user, movie, and rating data from the MovieLens dataset into pandas DataFrames to facilitate data manipulation and analysis.
|
||||
Load the data from the MovieLens dataset into pandas DataFrames to facilitate data manipulation and analysis.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
import pandas as pd
|
||||
```
|
||||
|
||||
Load user data:
|
||||
```python
|
||||
# load users
|
||||
users = pd.read_csv('data/ml-1m/users.dat', sep='::', names=['user_id', 'gender', 'age', 'occupation', 'zip'], engine='python')
|
||||
users = pd.read_csv('data/ml-1m/users.dat', sep='::', names=['user_id', 'gender', 'age', 'occupation', 'zip'], engine='*')
|
||||
users.head()
|
||||
```
|
||||
|
||||
Add movies:
|
||||
```python
|
||||
# load movies
|
||||
movies = pd.read_csv('data/ml-1m/movies.dat', sep='::', names=['movie_id', 'title', 'genres'], engine='python', encoding='latin-1')
|
||||
movies = pd.read_csv('data/ml-1m/movies.dat', sep='::', names=['movie_id', 'title', 'genres'], engine='*', encoding='latin-1')
|
||||
movies.head()
|
||||
```
|
||||
|
||||
Finally, add the ratings:
|
||||
```python
|
||||
#load ratings
|
||||
ratings = pd.read_csv( 'data/ml-1m/ratings.dat', sep='::', names=['user_id', 'movie_id', 'rating', 'timestamp'], engine='python')
|
||||
ratings = pd.read_csv( 'data/ml-1m/ratings.dat', sep='::', names=['user_id', 'movie_id', 'rating', 'timestamp'], engine='*')
|
||||
ratings.head()
|
||||
```
|
||||
|
||||
**Normalize ratings**
|
||||
### Normalize the ratings
|
||||
|
||||
Sparse vectors can use advantage of negative values, so we can normalize ratings to have a mean of 0 and a standard deviation of 1
|
||||
This normalization ensures that ratings are consistent and centered around zero, enabling accurate similarity calculations.
|
||||
In this scenario we can take into account movies that we don't like.
|
||||
Sparse vectors can use advantage of negative values, so we can normalize ratings to have a mean of 0 and a standard deviation of 1. This normalization ensures that ratings are consistent and centered around zero, enabling accurate similarity calculations. In this scenario we can take into account movies that we don't like.
|
||||
|
||||
```python
|
||||
ratings.rating = (ratings.rating - ratings.rating.mean()) / ratings.rating.std()
|
||||
```
|
||||
To get the results:
|
||||
|
||||
```python
|
||||
ratings.head()
|
||||
```
|
||||
|
||||
## Preparing the data and creating a collection
|
||||
### Data preparation
|
||||
|
||||
Transform user ratings into sparse vectors, where each vector represents ratings for different movies. This step prepares the data for indexing in Qdrant.
|
||||
Now you will transform user ratings into sparse vectors, where each vector represents ratings for different movies. This step prepares the data for indexing in Qdrant.
|
||||
|
||||
First, create a collection with configured sparse vectors
|
||||
- Sparse vectors don't require to specify dimension, because it's extracted from the data automatically
|
||||
|
||||
> An explanation of using hybrid cloud with OVH can be inserted here!
|
||||
First, create a collection with configured sparse vectors. For sparse vectors, you don't need to specify the dimension, because it's extracted from the data automatically.
|
||||
|
||||
```python
|
||||
# Convert ratings to sparse vectors
|
||||
|
||||
from collections import defaultdict
|
||||
|
||||
user_sparse_vectors = defaultdict(lambda: {"values": [], "indices": []})
|
||||
@@ -114,6 +109,8 @@ for row in ratings.itertuples():
|
||||
user_sparse_vectors[row.user_id]["values"].append(row.rating)
|
||||
user_sparse_vectors[row.user_id]["indices"].append(row.movie_id)
|
||||
```
|
||||
Connect to Qdrant and create a collection called **movielens**:
|
||||
|
||||
```python
|
||||
client = QdrantClient(
|
||||
url = os.getenv("QDRANT_HOST"),
|
||||
@@ -129,7 +126,7 @@ client.create_collection(
|
||||
)
|
||||
```
|
||||
|
||||
Upload user ratings to the "movielens" collection in Qdrant as sparse vectors, along with user metadata. This step populates the database with the necessary data for recommendation generation.
|
||||
Upload user ratings to the **movielens** collection in Qdrant as sparse vectors, along with user metadata. This step populates the database with the necessary data for recommendation generation.
|
||||
|
||||
```python
|
||||
def data_generator():
|
||||
@@ -148,7 +145,7 @@ client.upload_points(
|
||||
)
|
||||
```
|
||||
|
||||
## Running the Recommendation System
|
||||
## Recommendations
|
||||
|
||||
Personal movie ratings are specified, where positive ratings indicate likes and negative ratings indicate dislikes. These ratings serve as the basis for finding similar users with comparable tastes.
|
||||
|
||||
@@ -159,9 +156,10 @@ Let's try to recommend something for ourselves:
|
||||
1 = Like
|
||||
-1 = dislike
|
||||
|
||||
Search with movies[movies.title.str.contains("Matrix", case=False)]
|
||||
|
||||
```python
|
||||
# Search with movies[movies.title.str.contains("Matrix", case=False)].
|
||||
|
||||
my_ratings = {
|
||||
2571: 1, # Matrix
|
||||
329: 1, # Star Trek
|
||||
@@ -231,7 +229,7 @@ for movie_id, score in top_movies[:5]:
|
||||
|
||||
## Result
|
||||
|
||||
```bash
|
||||
```shell
|
||||
Star Wars: Episode V - The Empire Strikes Back (1980) 20.02387858
|
||||
Star Wars: Episode VI - Return of the Jedi (1983) 16.443184379999998
|
||||
Princess Bride, The (1987) 15.840068229999996
|
||||
|
||||
Reference in New Issue
Block a user