add scaleway and ovhcloud tutorials

This commit is contained in:
davidmyriel
2024-04-11 17:45:01 -07:00
parent 880eea1fc3
commit 9032b377ef
7 changed files with 71 additions and 57 deletions
@@ -1,51 +1,55 @@
---
title: Movie Recommendation System
title: Movie Recommendation System on OVHcloud
weight: 34
aliases:
- /documentation/tutorials/recommendation-system-ovhcloud/
---
# Build a Movie Recommendation System
# Movie Recommendation System on OVHcloud
| Time: 120 min | Level: Advanced | Output: [GitHub](https://github.com/infoslack/qdrant-example/blob/main/HC-demo/HC-OVH.ipynb) |
| --- | ----------- | ----------- |----------- |
This notebook aims to create a recommendation system using the MovieLens dataset and Qdrant. Vector databases like Qdrant are crucial for storing high-dimensional data, such as user and item embeddings, enabling personalized recommendations by quickly retrieving similar users or items based on advanced indexing techniques. We'll leverage collaborative filtering with a MovieLens dataset, identifying similar users based on ratings represented as vectors in Qdrant, and suggesting movies they liked but we haven't seen yet. The suggested items or content should closely align with the user's interests, leading to more personalized and relevant recommendations.
In this tutorial, you will build a mechanism that recommends movies based on defined preferences. Vector databases like Qdrant are good for storing high-dimensional data, such as user and item embeddings. They can enable personalized recommendations by quickly retrieving similar entries based on advanced indexing techniques. In this specific case, we will use [sparse vectors](/articles/sparse-vectors/) to create an efficient and accurate recommendation system.
Collaborative filtering works on the principle that users with similar tastes will enjoy similar movies. To implement this, we'll represent each user's ratings as vectors in a high-dimensional space using Qdrant. By indexing these vectors, we can find users with similar tastes to ours and recommend movies they liked but we haven't seen yet.
**Privacy and Sovereignty:** Since preference data is proprietary, it should be stored in a secure and controlled environment. Our vector database can easily be hosted on [OVHcloud](https://ovhcloud.com/), our trusted [Qdrant Hybrid Cloud](/documentation/hybrid-cloud/) partner. This means that Qdrant can be run from your OVHcloud region, but the database itself can still be managed from within Qdrant Cloud's interface. Both products have been tested for compatibility and scalability, and we recommend their [managed Kubernetes](https://www.ovhcloud.com/en/public-cloud/kubernetes/) service.
> To see the entire output, use our [notebook with complete instructions](https://github.com/infoslack/qdrant-example/blob/main/HC-demo/HC-OVH.ipynb).
## Components
- **Dataset:** [Red Hat Interactive Learning Portal](https://developers.redhat.com/learn)
- **Vector DB:** [Qdrant Hybrid Cloud](https://qdrant.tech) running on OpenShift.
- **Web Host:** [OVHcloud](https://haystack.deepset.ai/)
- **Dataset:** The [MovieLens dataset](https://grouplens.org/datasets/movielens/) contains a list of movies and ratings given by users.
- **Cloud:** [OVHcloud](https://ovhcloud.com/), with managed Kubernetes.
- **Vector DB:** [Qdrant Hybrid Cloud](https://qdrant.tech) running on [OVHcloud](https://ovhcloud.com/).
**Methodology:** We're adopting a collaborative filtering approach to construct a recommendation system from the dataset provided. Collaborative filtering works on the premise that if two users share similar tastes, they're likely to enjoy similar movies. Leveraging this concept, we'll identify users whose ratings align closely with ours, and explore the movies they liked but we haven't seen yet. To do this, we'll represent each user's ratings as a vector in a high-dimensional, sparse space. Using Qdrant, we'll index these vectors and search for users whose ratings vectors closely match ours. Ultimately, we will see which movies were enjoyed by users similar to us.
## Prerequisites
First, download and unzip the MovieLens dataset into a local directory.
Download and unzip the MovieLens dataset:
```bash
```shell
mkdir -p data
wget https://files.grouplens.org/datasets/movielens/ml-1m.zip
unzip ml-1m.zip -d data
```
The necessary Python libraries are installed using `pip`, including `pandas` for data manipulation, `qdrant-client` for interfacing with Qdrant, and `python-dotenv` for managing environment variables.
The necessary * libraries are installed using `pip`, including `pandas` for data manipulation, `qdrant-client` for interfacing with Qdrant, and `*-dotenv` for managing environment variables.
```python
!pip install -U \
pandas \
qdrant-client \
python-dotenv
*-dotenv
```
The `.env` file is used to store sensitive information like the Qdrant host URL and API key securely.
```bash
```shell
QDRANT_HOST
QDRANT_API_KEY
```
Load all environment variables into the setup.
Load all environment variables into the setup:
```python
import os
@@ -55,57 +59,48 @@ load_dotenv('./.env')
## Implementation
Load the user, movie, and rating data from the MovieLens dataset into pandas DataFrames to facilitate data manipulation and analysis.
Load the data from the MovieLens dataset into pandas DataFrames to facilitate data manipulation and analysis.
```python
from qdrant_client import QdrantClient, models
import pandas as pd
```
Load user data:
```python
# load users
users = pd.read_csv('data/ml-1m/users.dat', sep='::', names=['user_id', 'gender', 'age', 'occupation', 'zip'], engine='python')
users = pd.read_csv('data/ml-1m/users.dat', sep='::', names=['user_id', 'gender', 'age', 'occupation', 'zip'], engine='*')
users.head()
```
Add movies:
```python
# load movies
movies = pd.read_csv('data/ml-1m/movies.dat', sep='::', names=['movie_id', 'title', 'genres'], engine='python', encoding='latin-1')
movies = pd.read_csv('data/ml-1m/movies.dat', sep='::', names=['movie_id', 'title', 'genres'], engine='*', encoding='latin-1')
movies.head()
```
Finally, add the ratings:
```python
#load ratings
ratings = pd.read_csv( 'data/ml-1m/ratings.dat', sep='::', names=['user_id', 'movie_id', 'rating', 'timestamp'], engine='python')
ratings = pd.read_csv( 'data/ml-1m/ratings.dat', sep='::', names=['user_id', 'movie_id', 'rating', 'timestamp'], engine='*')
ratings.head()
```
**Normalize ratings**
### Normalize the ratings
Sparse vectors can use advantage of negative values, so we can normalize ratings to have a mean of 0 and a standard deviation of 1
This normalization ensures that ratings are consistent and centered around zero, enabling accurate similarity calculations.
In this scenario we can take into account movies that we don't like.
Sparse vectors can use advantage of negative values, so we can normalize ratings to have a mean of 0 and a standard deviation of 1. This normalization ensures that ratings are consistent and centered around zero, enabling accurate similarity calculations. In this scenario we can take into account movies that we don't like.
```python
ratings.rating = (ratings.rating - ratings.rating.mean()) / ratings.rating.std()
```
To get the results:
```python
ratings.head()
```
## Preparing the data and creating a collection
### Data preparation
Transform user ratings into sparse vectors, where each vector represents ratings for different movies. This step prepares the data for indexing in Qdrant.
Now you will transform user ratings into sparse vectors, where each vector represents ratings for different movies. This step prepares the data for indexing in Qdrant.
First, create a collection with configured sparse vectors
- Sparse vectors don't require to specify dimension, because it's extracted from the data automatically
> An explanation of using hybrid cloud with OVH can be inserted here!
First, create a collection with configured sparse vectors. For sparse vectors, you don't need to specify the dimension, because it's extracted from the data automatically.
```python
# Convert ratings to sparse vectors
from collections import defaultdict
user_sparse_vectors = defaultdict(lambda: {"values": [], "indices": []})
@@ -114,6 +109,8 @@ for row in ratings.itertuples():
user_sparse_vectors[row.user_id]["values"].append(row.rating)
user_sparse_vectors[row.user_id]["indices"].append(row.movie_id)
```
Connect to Qdrant and create a collection called **movielens**:
```python
client = QdrantClient(
url = os.getenv("QDRANT_HOST"),
@@ -129,7 +126,7 @@ client.create_collection(
)
```
Upload user ratings to the "movielens" collection in Qdrant as sparse vectors, along with user metadata. This step populates the database with the necessary data for recommendation generation.
Upload user ratings to the **movielens** collection in Qdrant as sparse vectors, along with user metadata. This step populates the database with the necessary data for recommendation generation.
```python
def data_generator():
@@ -148,7 +145,7 @@ client.upload_points(
)
```
## Running the Recommendation System
## Recommendations
Personal movie ratings are specified, where positive ratings indicate likes and negative ratings indicate dislikes. These ratings serve as the basis for finding similar users with comparable tastes.
@@ -159,9 +156,10 @@ Let's try to recommend something for ourselves:
1 = Like
-1 = dislike
Search with movies[movies.title.str.contains("Matrix", case=False)]
```python
# Search with movies[movies.title.str.contains("Matrix", case=False)].
my_ratings = {
2571: 1, # Matrix
329: 1, # Star Trek
@@ -231,7 +229,7 @@ for movie_id, score in top_movies[:5]:
## Result
```bash
```shell
Star Wars: Episode V - The Empire Strikes Back (1980) 20.02387858
Star Wars: Episode VI - Return of the Jedi (1983) 16.443184379999998
Princess Bride, The (1987) 15.840068229999996