mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-29 07:58:31 +02:00
reorganize sidebar examples + tutorials
This commit is contained in:
@@ -0,0 +1,38 @@
|
||||
---
|
||||
title: Examples
|
||||
weight: 34
|
||||
# If the index.md file is empty, the link to the section will be hidden from the sidebar
|
||||
is_empty: false
|
||||
---
|
||||
# Examples
|
||||
|
||||
| End-to-End Code Samples | Description | Stack |
|
||||
|---------------------------------------------------------------------------------|-------------------------------------------------------------------|---------------------------------------------|
|
||||
| [Aleph Alpha Search](../examples/aleph-alpha-search/) | Build a multimodal search that combines text and image data. | Qdrant, Aleph Alpha |
|
||||
| [Mighty Semantic Search](../examples/mighty/) | Build a simple semantic search with an on-demand NLP service. | Qdrant, Mighty |
|
||||
| [Multitenancy with LlamaIndex](../examples/llama-index-multitenancy/) | Handle data coming from multiple users in LlamaIndex. | Qdrant, Python, LlamaIndex |
|
||||
| [Implement custom connector for Cohere RAG](../examples/cohere-rag-connector/) | Bring data stored in Qdrant to Cohere RAG | Qdrant, Cohere, FastAPI |
|
||||
| [Chatbot for Interactive Learning](../examples/rag-chatbot-red-hat-openshift-haystack/) | Build a Private RAG Chatbot for Interactive Learning | Qdrant, Haystack, OpenShift |
|
||||
| [Information Extraction Engine](../examples/rag-chatbot-vultr-dspy-ollama/) | Build a Private RAG Information Extraction Engine | Qdrant, Vultr, DSPy, Ollama |
|
||||
| [System for Employee Onboarding](../examples/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/) | Build a RAG System for Employee Onboarding | Qdrant, Cohere, LangChain |
|
||||
| [System for Contract Management](../examples/rag-contract-management-stackit-aleph-alpha/) | Build a Region-Specific RAG System for Contract Management | Qdrant, Aleph Alpha, STACKIT |
|
||||
| [Question-Answering System for Customer Support](../examples/rag-customer-support-cohere-airbyte-aws/) | Build a RAG System for AI Customer Support | Qdrant, Cohere, Airbyte, AWS |
|
||||
| [Hybrid Search on PDF Documents](../examples/hybrid-search-llamaindex-jinaai/) | Develop a Hybrid Search System for Product PDF Manuals | Qdrant, LlamaIndex, Jina AI
|
||||
| [Build a RAG-based Chatbot](../examples/rag-chatbot-scaleway) | Develop Build a RAG-based Chatbot on Scaleway and with LangChain | Qdrant, LlamaIndex, Jina AI
|
||||
| [Movie Recommendation System](../examples/recommendation-system-ovhcloud/) | Build a Movie Recommendation System with LlamaIndex and With JinaAI | Qdrant, LlamaIndex, Jina AI
|
||||
|
||||
|
||||
## Notebooks
|
||||
|
||||
Our Notebooks offer complex instructions that are supported with a throrough explanation. Follow along by trying out the code and get the most out of each example.
|
||||
|
||||
| Example | Description | Stack |
|
||||
|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------|----------------------------|
|
||||
| [Intro to Semantic Search and Recommendations Systems](https://githubtocolab.com/qdrant/examples/blob/master/qdrant_101_getting_started/getting_started.ipynb) | Learn how to get started building semantic search and recommendation systems. | Qdrant |
|
||||
| [Search and Recommend Newspaper Articles](https://githubtocolab.com/qdrant/examples/blob/master/qdrant_101_text_data/qdrant_and_text_data.ipynb) | Work with text data to develop a semantic search and a recommendation engine for news articles. | Qdrant |
|
||||
| [Recommendation System for Songs](https://githubtocolab.com/qdrant/examples/blob/master/qdrant_101_audio_data/03_qdrant_101_audio.ipynb) | Use Qdrant to develop a music recommendation engine based on audio embeddings. | Qdrant |
|
||||
| [Image Comparison System for Skin Conditions](https://colab.research.google.com/github/qdrant/examples/blob/master/qdrant_101_image_data/04_qdrant_101_cv.ipynb) | Use Qdrant to compare challenging images with labels representing different skin diseases. | Qdrant |
|
||||
| [Question and Answer System with LlamaIndex](https://githubtocolab.com/qdrant/examples/blob/master/llama_index_recency/Qdrant%20and%20LlamaIndex%20%E2%80%94%20A%20new%20way%20to%20keep%20your%20Q%26A%20systems%20up-to-date.ipynb) | Combine Qdrant and LlamaIndex to create a self-updating Q&A system. | Qdrant, LlamaIndex, Cohere |
|
||||
| [Extractive QA System](https://githubtocolab.com/qdrant/examples/blob/master/extractive_qa/extractive-question-answering.ipynb) | Extract answers directly from context to generate highly relevant answers. | Qdrant |
|
||||
| [Ecommerce Reverse Image Search](https://githubtocolab.com/qdrant/examples/blob/master/ecommerce_reverse_image_search/ecommerce-reverse-image-search.ipynb) | Accept images as search queries to receive semantically appropriate answers. | Qdrant |
|
||||
| [Basic RAG](https://githubtocolab.com/qdrant/examples/blob/master/rag-openai-qdrant/rag-openai-qdrant.ipynb) | Basic RAG pipeline with Qdrant and OpenAI SDKs | OpenAI, Qdrant, FastEmbed |
|
||||
@@ -0,0 +1,180 @@
|
||||
---
|
||||
title: Aleph Alpha Search
|
||||
weight: 16
|
||||
aliases:
|
||||
- /documentation/tutorials/aleph-alpha-search/
|
||||
---
|
||||
|
||||
# Multimodal Semantic Search with Aleph Alpha
|
||||
|
||||
| Time: 30 min | Level: Beginner | | |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
This tutorial shows you how to run a proper multimodal semantic search system with a few lines of code, without the need to annotate the data or train your networks.
|
||||
|
||||
In most cases, semantic search is limited to homogenous data types for both documents and queries (text-text, image-image, audio-audio, etc.). With the recent growth of multimodal architectures, it is now possible to encode different data types into the same latent space. That opens up some great possibilities, as you can finally explore non-textual data, for example visual, with text queries.
|
||||
|
||||
In the past, this would require labelling every image with a description of what it presents. Right now, you can rely on vector embeddings, which can represent all
|
||||
the inputs in the same space.
|
||||
|
||||
*Figure 1: Two examples of text-image pairs presenting a similar object, encoded by a multimodal network into the same
|
||||
2D latent space. Both texts are examples of English [pangrams](https://en.wikipedia.org/wiki/Pangram).
|
||||
https://deepai.org generated the images with pangrams used as input prompts.*
|
||||
|
||||

|
||||
|
||||
|
||||
## Sample dataset
|
||||
|
||||
You will be using [COCO](https://cocodataset.org/), a large-scale object detection, segmentation, and captioning dataset. It provides
|
||||
various splits, 330,000 images in total. For demonstration purposes, this tutorials uses the
|
||||
[2017 validation split](http://images.cocodataset.org/zips/train2017.zip) that contains 5000 images from different
|
||||
categories with total size about 19GB.
|
||||
```terminal
|
||||
wget http://images.cocodataset.org/zips/train2017.zip
|
||||
```
|
||||
|
||||
## Prerequisites
|
||||
|
||||
There is no need to curate your datasets and train the models. [Aleph Alpha](https://www.aleph-alpha.com/), already has multimodality and multilinguality already built-in. There is an [official Python client](https://github.com/Aleph-Alpha/aleph-alpha-client) that simplifies the integration.
|
||||
|
||||
In order to enable the search capabilities, you need to build the search index to query on. For this example,
|
||||
you are going to vectorize the images and store their embeddings along with the filenames. You can then return the most
|
||||
similar files for given query.
|
||||
|
||||
There are two things you need to set up before you start:
|
||||
|
||||
1. You need to have a Qdrant instance running. If you want to launch it locally,
|
||||
[Docker is the fastest way to do that](/documentation/quick_start/#installation).
|
||||
2. You need to have a registered [Aleph Alpha account](https://app.aleph-alpha.com/).
|
||||
3. Upon registration, create an API key (see: [API Tokens](https://app.aleph-alpha.com/profile)).
|
||||
|
||||
Now you can store the Aleph Alpha API key in a variable and choose the model your are going to use.
|
||||
|
||||
```python
|
||||
aa_token = "<< your_token >>"
|
||||
model = "luminous-base"
|
||||
```
|
||||
|
||||
## Vectorize the dataset
|
||||
|
||||
In this example, images have been extracted and are stored in the `val2017` directory:
|
||||
|
||||
```python
|
||||
from aleph_alpha_client import (
|
||||
Prompt,
|
||||
AsyncClient,
|
||||
SemanticEmbeddingRequest,
|
||||
SemanticRepresentation,
|
||||
Image,
|
||||
)
|
||||
|
||||
from glob import glob
|
||||
|
||||
ids, vectors, payloads = [], [], []
|
||||
async with AsyncClient(token=aa_token) as aa_client:
|
||||
for i, image_path in enumerate(glob("./val2017/*.jpg")):
|
||||
# Convert the JPEG file into the embedding by calling
|
||||
# Aleph Alpha API
|
||||
prompt = Image.from_file(image_path)
|
||||
prompt = Prompt.from_image(prompt)
|
||||
query_params = {
|
||||
"prompt": prompt,
|
||||
"representation": SemanticRepresentation.Symmetric,
|
||||
"compress_to_size": 128,
|
||||
}
|
||||
query_request = SemanticEmbeddingRequest(**query_params)
|
||||
query_response = await aa_client.semantic_embed(request=query_request, model=model)
|
||||
|
||||
# Finally store the id, vector and the payload
|
||||
ids.append(i)
|
||||
vectors.append(query_response.embedding)
|
||||
payloads.append({"filename": image_path})
|
||||
```
|
||||
|
||||
## Load embeddings into Qdrant
|
||||
|
||||
Add all created embeddings, along with their ids and payloads into the `COCO` collection.
|
||||
|
||||
```python
|
||||
import qdrant_client
|
||||
from qdrant_client.models import Batch, VectorParams, Distance
|
||||
|
||||
client = qdrant_client.QdrantClient()
|
||||
client.recreate_collection(
|
||||
collection_name="COCO",
|
||||
vectors_config=VectorParams(
|
||||
size=len(vectors[0]),
|
||||
distance=Distance.COSINE,
|
||||
),
|
||||
)
|
||||
client.upsert(
|
||||
collection_name="COCO",
|
||||
points=Batch(
|
||||
ids=ids,
|
||||
vectors=vectors,
|
||||
payloads=payloads,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
## Query the database
|
||||
|
||||
The `luminous-base`, model can provide you the vectors for both texts and images, which means you can run both
|
||||
text queries and reverse image search. Assume you want to find images similar to the one below:
|
||||
|
||||

|
||||
|
||||
With the following code snippet create its vector embedding and then perform the lookup in Qdrant:
|
||||
|
||||
```python
|
||||
async with AsyncCliet(token=aa_token) as aa_client:
|
||||
prompt = ImagePrompt.from_file("query.jpg")
|
||||
prompt = Prompt.from_image(prompt)
|
||||
|
||||
query_params = {
|
||||
"prompt": prompt,
|
||||
"representation": SemanticRepresentation.Symmetric,
|
||||
"compress_to_size": 128,
|
||||
}
|
||||
query_request = SemanticEmbeddingRequest(**query_params)
|
||||
query_response = await aa_client.semantic_embed(request=query_request, model=model)
|
||||
|
||||
results = client.search(
|
||||
collection_name="COCO",
|
||||
query_vector=query_response.embedding,
|
||||
limit=3,
|
||||
)
|
||||
print(results)
|
||||
```
|
||||
|
||||
Here are the results:
|
||||
|
||||

|
||||
|
||||
**Note:** AlephAlpha models can provide embeddings for English, French, German, Italian
|
||||
and Spanish. Your search is not only multimodal, but also multilingual, without any need for translations.
|
||||
|
||||
```python
|
||||
text = "Surfing"
|
||||
|
||||
async with AsyncClient(token=aa_token) as aa_client:
|
||||
query_params = {
|
||||
"prompt": Prompt.from_text(text),
|
||||
"representation": SemanticRepresentation.Symmetric,
|
||||
"compres_to_size": 128,
|
||||
}
|
||||
query_request = SemanticEmbeddingRequest(**query_params)
|
||||
query_response = await aa_client.semantic_embed(request=query_request, model=model)
|
||||
|
||||
results = client.search(
|
||||
collection_name="COCO",
|
||||
query_vector=query_response.embedding,
|
||||
limit=3,
|
||||
)
|
||||
print(results)
|
||||
```
|
||||
|
||||
Here are the top 3 results for “Surfing”:
|
||||
|
||||

|
||||
@@ -0,0 +1,306 @@
|
||||
---
|
||||
title: Implement Cohere RAG connector
|
||||
weight: 24
|
||||
aliases:
|
||||
- /documentation/tutorials/cohere-rag-connector/
|
||||
---
|
||||
|
||||
# Implement custom connector for Cohere RAG
|
||||
|
||||
| Time: 45 min | Level: Intermediate | | |
|
||||
|--------------|---------------------|-|----|
|
||||
|
||||
The usual approach to implementing Retrieval Augmented Generation requires users to build their prompts with the
|
||||
relevant context the LLM may rely on, and manually sending them to the model. Cohere is quite unique here, as their
|
||||
models can now speak to the external tools and extract meaningful data on their own. You can virtually connect any data
|
||||
source and let the Cohere LLM know how to access it. Obviously, vector search goes well with LLMs, and enabling semantic
|
||||
search over your data is a typical case.
|
||||
|
||||
Cohere RAG has lots of interesting features, such as inline citations, which help you to refer to the specific parts of
|
||||
the documents used to generate the response.
|
||||
|
||||

|
||||
|
||||
*Source: https://docs.cohere.com/docs/retrieval-augmented-generation-rag*
|
||||
|
||||
The connectors have to implement a specific interface and expose the data source as HTTP REST API. Cohere documentation
|
||||
[describes a general process of creating a connector](https://docs.cohere.com/docs/creating-and-deploying-a-connector).
|
||||
This tutorial guides you step by step on building such a service around Qdrant.
|
||||
|
||||
## Qdrant connector
|
||||
|
||||
You probably already have some collections you would like to bring to the LLM. Maybe your pipeline was set up using some
|
||||
of the popular libraries such as Langchain, Llama Index, or Haystack. Cohere connectors may implement even more complex
|
||||
logic, e.g. hybrid search. In our case, we are going to start with a fresh Qdrant collection, index data using Cohere
|
||||
Embed v3, build the connector, and finally connect it with the [Command-R model](https://txt.cohere.com/command-r/).
|
||||
|
||||
### Building the collection
|
||||
|
||||
First things first, let's build a collection and configure it for the Cohere `embed-multilingual-v3.0` model. It
|
||||
produces 1024-dimensional embeddings, and we can choose any of the distance metrics available in Qdrant. Our connector
|
||||
will act as a personal assistant of a software engineer, and it will expose our notes to suggest the priorities or
|
||||
actions to perform.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
"https://my-cluster.cloud.qdrant.io:6333",
|
||||
api_key="my-api-key",
|
||||
)
|
||||
client.create_collection(
|
||||
collection_name="personal-notes",
|
||||
vectors_config=models.VectorParams(
|
||||
size=1024,
|
||||
distance=models.Distance.DOT,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
Our notes will be represented as simple JSON objects with a `title` and `text` of the specific note. The embeddings will
|
||||
be created from the `text` field only.
|
||||
|
||||
```python
|
||||
notes = [
|
||||
{
|
||||
"title": "Project Alpha Review",
|
||||
"text": "Review the current progress of Project Alpha, focusing on the integration of the new API. Check for any compatibility issues with the existing system and document the steps needed to resolve them. Schedule a meeting with the development team to discuss the timeline and any potential roadblocks."
|
||||
},
|
||||
{
|
||||
"title": "Learning Path Update",
|
||||
"text": "Update the learning path document with the latest courses on React and Node.js from Pluralsight. Schedule at least 2 hours weekly to dedicate to these courses. Aim to complete the React course by the end of the month and the Node.js course by mid-next month."
|
||||
},
|
||||
{
|
||||
"title": "Weekly Team Meeting Agenda",
|
||||
"text": "Prepare the agenda for the weekly team meeting. Include the following topics: project updates, review of the sprint backlog, discussion on the new feature requests, and a brainstorming session for improving remote work practices. Send out the agenda and the Zoom link by Thursday afternoon."
|
||||
},
|
||||
{
|
||||
"title": "Code Review Process Improvement",
|
||||
"text": "Analyze the current code review process to identify inefficiencies. Consider adopting a new tool that integrates with our version control system. Explore options such as GitHub Actions for automating parts of the process. Draft a proposal with recommendations and share it with the team for feedback."
|
||||
},
|
||||
{
|
||||
"title": "Cloud Migration Strategy",
|
||||
"text": "Draft a plan for migrating our current on-premise infrastructure to the cloud. The plan should cover the selection of a cloud provider, cost analysis, and a phased migration approach. Identify critical applications for the first phase and any potential risks or challenges. Schedule a meeting with the IT department to discuss the plan."
|
||||
},
|
||||
{
|
||||
"title": "Quarterly Goals Review",
|
||||
"text": "Review the progress towards the quarterly goals. Update the documentation to reflect any completed objectives and outline steps for any remaining goals. Schedule individual meetings with team members to discuss their contributions and any support they might need to achieve their targets."
|
||||
},
|
||||
{
|
||||
"title": "Personal Development Plan",
|
||||
"text": "Reflect on the past quarter's achievements and areas for improvement. Update the personal development plan to include new technical skills to learn, certifications to pursue, and networking events to attend. Set realistic timelines and check-in points to monitor progress."
|
||||
},
|
||||
{
|
||||
"title": "End-of-Year Performance Reviews",
|
||||
"text": "Start preparing for the end-of-year performance reviews. Collect feedback from peers and managers, review project contributions, and document achievements. Consider areas for improvement and set goals for the next year. Schedule preliminary discussions with each team member to gather their self-assessments."
|
||||
},
|
||||
{
|
||||
"title": "Technology Stack Evaluation",
|
||||
"text": "Conduct an evaluation of our current technology stack to identify any outdated technologies or tools that could be replaced for better performance and productivity. Research emerging technologies that might benefit our projects. Prepare a report with findings and recommendations to present to the management team."
|
||||
},
|
||||
{
|
||||
"title": "Team Building Event Planning",
|
||||
"text": "Plan a team-building event for the next quarter. Consider activities that can be done remotely, such as virtual escape rooms or online game nights. Survey the team for their preferences and availability. Draft a budget proposal for the event and submit it for approval."
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
Storing the embeddings along with the metadata is fairly simple.
|
||||
|
||||
```python
|
||||
import cohere
|
||||
import uuid
|
||||
|
||||
cohere_client = cohere.Client(api_key="my-cohere-api-key")
|
||||
|
||||
response = cohere_client.embed(
|
||||
texts=[
|
||||
note.get("text")
|
||||
for note in notes
|
||||
],
|
||||
model="embed-multilingual-v3.0",
|
||||
input_type="search_document",
|
||||
)
|
||||
|
||||
client.upload_points(
|
||||
collection_name="personal-notes",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=uuid.uuid4().hex,
|
||||
vector=embedding,
|
||||
payload=note,
|
||||
)
|
||||
for note, embedding in zip(notes, response.embeddings)
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
Our collection is now ready to be searched over. In the real world, the set of notes would be changing over time, so the
|
||||
ingestion process won't be as straightforward. This data is not yet exposed to the LLM, but we will build the connector
|
||||
in the next step.
|
||||
|
||||
### Connector web service
|
||||
|
||||
[FastAPI](https://fastapi.tiangolo.com/) is a modern web framework and perfect a choice for a simple HTTP API. We are
|
||||
going to use it for the purposes of our connector. There will be just one endpoint, as required by the model. It will
|
||||
accept POST requests at the `/search` path. There is a single `query` parameter required. Let's define a corresponding
|
||||
model.
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
|
||||
class SearchQuery(BaseModel):
|
||||
query: str
|
||||
```
|
||||
|
||||
RAG connector does not have to return the documents in any specific format. There are [some good practices to follow](https://docs.cohere.com/docs/creating-and-deploying-a-connector#configure-the-connection-between-the-connector-and-the-chat-api),
|
||||
but Cohere models are quite flexible here. Results just have to be returned as JSON, with a list of objects in a
|
||||
`results` property of the output. We will use the same document structure as we did for the Qdrant payloads, so there
|
||||
is no conversion required. That requires two additional models to be created.
|
||||
|
||||
```python
|
||||
from typing import List
|
||||
|
||||
class Document(BaseModel):
|
||||
title: str
|
||||
text: str
|
||||
|
||||
class SearchResults(BaseModel):
|
||||
results: List[Document]
|
||||
```
|
||||
|
||||
Once our model classes are ready, we can implement the logic that will get the query and provide the notes that are
|
||||
relevant to it. Please note the LLM is not going to define the number of documents to be returned. That's completely
|
||||
up to you how many of them you want to bring to the context.
|
||||
|
||||
There are two services we need to interact with - Qdrant server and Cohere API. FastAPI has a concept of a [dependency
|
||||
injection](https://fastapi.tiangolo.com/tutorial/dependencies/#dependencies), and we will use it to provide both
|
||||
clients into the implementation.
|
||||
|
||||
In case of queries, we need to set the `input_type` to `search_query` in the calls to Cohere API.
|
||||
|
||||
```python
|
||||
from fastapi import FastAPI, Depends
|
||||
from typing import Annotated
|
||||
|
||||
app = FastAPI()
|
||||
|
||||
def client() -> QdrantClient:
|
||||
return QdrantClient(config.QDRANT_URL, api_key=config.QDRANT_API_KEY)
|
||||
|
||||
def cohere_client() -> cohere.Client:
|
||||
return cohere.Client(api_key=config.COHERE_API_KEY)
|
||||
|
||||
@app.post("/search")
|
||||
def search(
|
||||
query: SearchQuery,
|
||||
client: Annotated[QdrantClient, Depends(client)],
|
||||
cohere_client: Annotated[cohere.Client, Depends(cohere_client)],
|
||||
) -> SearchResults:
|
||||
response = cohere_client.embed(
|
||||
texts=[query.query],
|
||||
model="embed-multilingual-v3.0",
|
||||
input_type="search_query",
|
||||
)
|
||||
results = client.search(
|
||||
collection_name="personal-notes",
|
||||
query_vector=response.embeddings[0],
|
||||
limit=2,
|
||||
)
|
||||
return SearchResults(
|
||||
results=[
|
||||
Document(**point.payload)
|
||||
for point in results
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
Our app might be launched locally for the development purposes, given we have the `uvicorn` server installed:
|
||||
|
||||
```shell
|
||||
uvicorn main:app
|
||||
```
|
||||
|
||||
FastAPI exposes an interactive documentation at `http://localhost:8000/docs`, where we can test our endpoint. The
|
||||
`/search` endpoint is available there.
|
||||
|
||||

|
||||
|
||||
We can interact with it and check the documents that will be returned for a specific query. For example, we want to know
|
||||
recall what we are supposed to do regarding the infrastructure for your projects.
|
||||
|
||||
```shell
|
||||
curl -X "POST" \
|
||||
-H "Content-type: application/json" \
|
||||
-d '{"query": "Is there anything I have to do regarding the project infrastructure?"}' \
|
||||
"http://localhost:8000/search"
|
||||
```
|
||||
|
||||
The output should look like following:
|
||||
|
||||
```json
|
||||
{
|
||||
"results": [
|
||||
{
|
||||
"title": "Cloud Migration Strategy",
|
||||
"text": "Draft a plan for migrating our current on-premise infrastructure to the cloud. The plan should cover the selection of a cloud provider, cost analysis, and a phased migration approach. Identify critical applications for the first phase and any potential risks or challenges. Schedule a meeting with the IT department to discuss the plan."
|
||||
},
|
||||
{
|
||||
"title": "Project Alpha Review",
|
||||
"text": "Review the current progress of Project Alpha, focusing on the integration of the new API. Check for any compatibility issues with the existing system and document the steps needed to resolve them. Schedule a meeting with the development team to discuss the timeline and any potential roadblocks."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### Connecting to Command-R
|
||||
|
||||
Our web service is implemented, yet running only on our local machine. It has to be exposed to the public before
|
||||
Command-R can interact with it. For a quick experiment, it might be enough to set up tunneling using services such as
|
||||
[ngrok](https://ngrok.com/). We won't cover all the details in the tutorial, but their
|
||||
[Quickstart](https://ngrok.com/docs/guides/getting-started/) is a great resource describing the process step-by-step.
|
||||
Alternatively, you can also deploy the service with a public URL.
|
||||
|
||||
Once it's done, we can create the connector first, and then tell the model to use it, while interacting through the chat
|
||||
API. Creating a connector is a single call to Cohere client:
|
||||
|
||||
```python
|
||||
connector_response = cohere_client.connectors.create(
|
||||
name="personal-notes",
|
||||
url="https:/this-is-my-domain.app/search",
|
||||
)
|
||||
```
|
||||
|
||||
The `connector_response.connector` will be a descriptor, with `id` being one of the attributes. We'll use this
|
||||
identifier for our interactions like this:
|
||||
|
||||
```python
|
||||
response = cohere_client.chat(
|
||||
message=(
|
||||
"Is there anything I have to do regarding the project infrastructure? "
|
||||
"Please mention the tasks briefly."
|
||||
),
|
||||
connectors=[
|
||||
cohere.ChatConnector(id=connector_response.connector.id)
|
||||
],
|
||||
model="command-r",
|
||||
)
|
||||
```
|
||||
|
||||
We changed the `model` to `command-r`, as this is currently the best Cohere model available to public. The
|
||||
`response.text` is the output of the model:
|
||||
|
||||
```text
|
||||
Here are some of the tasks related to project infrastructure that you might have to perform:
|
||||
- You need to draft a plan for migrating your on-premise infrastructure to the cloud and come up with a plan for the selection of a cloud provider, cost analysis, and a gradual migration approach.
|
||||
- It's important to evaluate your current technology stack to identify any outdated technologies. You should also research emerging technologies and the benefits they could bring to your projects.
|
||||
```
|
||||
|
||||
You only need to create a specific connector once! Please do not call `cohere_client.connectors.create` for every single
|
||||
message you send to the `chat` method.
|
||||
|
||||
## Wrapping up
|
||||
|
||||
We have built a Cohere RAG connector that integrates with your existing knowledge base stored in Qdrant. We covered just
|
||||
the basic flow, but in real world scenarios, you should also consider e.g. [building the authentication
|
||||
system](https://docs.cohere.com/docs/connector-authentication) to prevent unauthorized access.
|
||||
@@ -0,0 +1,220 @@
|
||||
---
|
||||
title: Chat With Product PDF Manuals Using Hybrid Search
|
||||
weight: 27
|
||||
aliases:
|
||||
- /documentation/tutorials/hybrid-search-llamaindex-jinaai/
|
||||
---
|
||||
|
||||
# Chat With Product PDF Manuals Using Hybrid Search
|
||||
|
||||
| Time: 120 min | Level: Advanced | Output: [GitHub](https://github.com/infoslack/qdrant-example/blob/main/HC-demo/HC-DO-LlamaIndex-Jina-v2.ipynb) | [](https://githubtocolab.com/infoslack/qdrant-example/blob/main/HC-demo/HC-DO-LlamaIndex-Jina-v2.ipynb) |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
With the proliferation of digital manuals and the increasing demand for quick and accurate customer support, having a chatbot capable of efficiently parsing through complex PDF documents and delivering precise information can be a game-changer for any business.
|
||||
|
||||
In this tutorial, we'll walk you through the process of building a RAG-based chatbot, designed specifically to assist users with understanding the operation of various household appliances.
|
||||
We'll cover the essential steps required to build your system, including data ingestion, natural language understanding, and response generation for customer support use cases.
|
||||
|
||||
## Components
|
||||
|
||||
- **Embeddings:** Jina Embeddings, served via the [Jina Embeddings API](https://jina.ai/embeddings/#apiform)
|
||||
- **Database:** [Qdrant Hybrid Cloud](/documentation/hybrid-cloud/), deployed in an environment of your own choice
|
||||
- **LLM:** [Mixtral-8x7B-Instruct-v0.1](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1) language model on HuggingFace
|
||||
- **Framework:** [LlamaIndex](https://www.llamaindex.ai/) for extended RAG functionality and [Hybrid Search support](https://docs.llamaindex.ai/en/stable/examples/vector_stores/qdrant_hybrid/).
|
||||
- **Parser:** [LlamaParse](https://github.com/run-llama/llama_parse) as a way to parse complex documents with embedded objects such as tables and figures.
|
||||
|
||||
### Procedure
|
||||
|
||||
Retrieval Augmented Generation (RAG) combines search with language generation. An external information retrieval system is used to identify documents likely to provide information relevant to the user's query. These documents, along with the user's request, are then passed on to a text-generating language model, producing a natural response.
|
||||
|
||||
This method enables a language model to respond to questions and access information from a much larger set of documents than it could see otherwise. The language model only looks at a few relevant sections of the documents when generating responses, which also helps to reduce inexplicable errors.
|
||||
|
||||
## Prerequisites
|
||||
First, install all dependencies:
|
||||
|
||||
```python
|
||||
!pip install -U \
|
||||
llama-index \
|
||||
llama-parse \
|
||||
python-dotenv \
|
||||
llama-index-embeddings-jinaai \
|
||||
llama-index-llms-huggingface \
|
||||
llama-index-vector-stores-qdrant \
|
||||
"huggingface_hub[inference]" \
|
||||
datasets
|
||||
```
|
||||
|
||||
Set up secret key values on `.env` file:
|
||||
|
||||
```bash
|
||||
JINAAI_API_KEY
|
||||
HF_INFERENCE_API_KEY
|
||||
LLAMA_CLOUD_API_KEY
|
||||
QDRANT_HOST
|
||||
QDRANT_API_KEY
|
||||
```
|
||||
|
||||
Load all environment variables:
|
||||
|
||||
```python
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
load_dotenv('./.env')
|
||||
```
|
||||
## Implementation
|
||||
|
||||
### Connect Jina Embeddings and Mixtral LLM
|
||||
|
||||
LlamaIndex provides built-in support for the [Jina Embeddings API](https://jina.ai/embeddings/#apiform). To use it, you need to initialize the `JinaEmbedding` object with your API Key and model name.
|
||||
|
||||
For the LLM, you need wrap it in a subclass of `llama_index.llms.CustomLLM` to make it compatible with LlamaIndex.
|
||||
|
||||
```python
|
||||
# connect embeddings
|
||||
from llama_index.embeddings.jinaai import JinaEmbedding
|
||||
|
||||
jina_embedding_model = JinaEmbedding(
|
||||
model="jina-embeddings-v2-base-en",
|
||||
api_key=os.getenv("JINAAI_API_KEY"),
|
||||
)
|
||||
|
||||
# connect LLM
|
||||
from llama_index.llms.huggingface import HuggingFaceInferenceAPI
|
||||
|
||||
mixtral_llm = HuggingFaceInferenceAPI(
|
||||
model_name = "mistralai/Mixtral-8x7B-Instruct-v0.1",
|
||||
token=os.getenv("HF_INFERENCE_API_KEY"),
|
||||
)
|
||||
```
|
||||
|
||||
### Prepare data for RAG
|
||||
|
||||
This example will use household appliance manuals, which are generally available as PDF documents.
|
||||
LlamaPar
|
||||
In the `data` folder, we have three documents, and we will use it to extract the textual content from the PDF and use it as a knowledge base in a simple RAG.
|
||||
|
||||
The free LlamaIndex Cloud plan is sufficient for our example:
|
||||
|
||||
```python
|
||||
import nest_asyncio
|
||||
nest_asyncio.apply()
|
||||
from llama_parse import LlamaParse
|
||||
|
||||
llamaparse_api_key = os.getenv("LLAMA_CLOUD_API_KEY")
|
||||
|
||||
llama_parse_documents = LlamaParse(api_key=llamaparse_api_key, result_type="markdown").load_data([
|
||||
"data/DJ68-00682F_0.0.pdf",
|
||||
"data/F500E_WF80F5E_03445F_EN.pdf",
|
||||
"data/O_ME4000R_ME19R7041FS_AA_EN.pdf"
|
||||
])
|
||||
```
|
||||
|
||||
### Store data into Qdrant
|
||||
The code below does the following:
|
||||
|
||||
- create a vector store with Qdrant client;
|
||||
- get an embedding for each chunk using Jina Embeddings API;
|
||||
- combines `sparse` and `dense` vectors for hybrid search;
|
||||
- stores all data into Qdrant;
|
||||
|
||||
Hybrid search with Qdrant must be enabled from the beginning - we can simply set `enable_hybrid=True`.
|
||||
|
||||
```python
|
||||
# By default llamaindex uses OpenAI models
|
||||
# setting embed_model to Jina and llm model to Mixtral
|
||||
from llama_index.core import Settings
|
||||
Settings.embed_model = jina_embedding_model
|
||||
Settings.llm = mixtral_llm
|
||||
|
||||
from llama_index.core import VectorStoreIndex, StorageContext
|
||||
from llama_index.vector_stores.qdrant import QdrantVectorStore
|
||||
import qdrant_client
|
||||
|
||||
client = qdrant_client.QdrantClient(
|
||||
url = os.getenv("QDRANT_HOST"),
|
||||
api_key = os.getenv("QDRANT_API_KEY")
|
||||
)
|
||||
|
||||
vector_store = QdrantVectorStore(
|
||||
client=client, collection_name="demo", enable_hybrid=True, batch_size=20
|
||||
)
|
||||
Settings.chunk_size = 512
|
||||
|
||||
storage_context = StorageContext.from_defaults(vector_store=vector_store)
|
||||
index = VectorStoreIndex.from_documents(
|
||||
documents=llama_parse_documents,
|
||||
storage_context=storage_context
|
||||
)
|
||||
```
|
||||
|
||||
### Prepare a prompt
|
||||
Here we will create a custom prompt template. This prompt asks the LLM to use only the context information retrieved from Qdrant. When querying with hybrid mode, we can set `similarity_top_k` and `sparse_top_k` separately:
|
||||
|
||||
- `sparse_top_k` represents how many nodes will be retrieved from each dense and sparse query.
|
||||
- `similarity_top_k` controls the final number of returned nodes. In the above setting, we end up with 10 nodes.
|
||||
|
||||
Then, we assemble the query engine using the prompt.
|
||||
|
||||
```python
|
||||
from llama_index.core import PromptTemplate
|
||||
|
||||
qa_prompt_tmpl = (
|
||||
"Context information is below.\n"
|
||||
"-------------------------------"
|
||||
"{context_str}\n"
|
||||
"-------------------------------"
|
||||
"Given the context information and not prior knowledge,"
|
||||
"answer the query. Please be concise, and complete.\n"
|
||||
"If the context does not contain an answer to the query,"
|
||||
"respond with \"I don't know!\"."
|
||||
"Query: {query_str}\n"
|
||||
"Answer: "
|
||||
)
|
||||
qa_prompt = PromptTemplate(qa_prompt_tmpl)
|
||||
|
||||
from llama_index.core.retrievers import VectorIndexRetriever
|
||||
from llama_index.core.query_engine import RetrieverQueryEngine
|
||||
from llama_index.core import get_response_synthesizer
|
||||
from llama_index.core import Settings
|
||||
Settings.embed_model = jina_embedding_model
|
||||
Settings.llm = mixtral_llm
|
||||
|
||||
# retriever
|
||||
retriever = VectorIndexRetriever(
|
||||
index=index,
|
||||
similarity_top_k=2,
|
||||
sparse_top_k=12,
|
||||
vector_store_query_mode="hybrid"
|
||||
)
|
||||
|
||||
# response synthesizer
|
||||
response_synthesizer = get_response_synthesizer(
|
||||
llm=mixtral_llm,
|
||||
text_qa_template=qa_prompt,
|
||||
response_mode="compact",
|
||||
)
|
||||
|
||||
# query engine
|
||||
query_engine = RetrieverQueryEngine(
|
||||
retriever=retriever,
|
||||
response_synthesizer=response_synthesizer,
|
||||
)
|
||||
```
|
||||
|
||||
## Run a test query
|
||||
Now you can ask questions and receive answers based on the data:
|
||||
|
||||
**Question**
|
||||
|
||||
```python
|
||||
result = query_engine.query("What temperature should I use for my laundry?")
|
||||
print(result.response)
|
||||
```
|
||||
|
||||
**Answer**
|
||||
|
||||
```python
|
||||
The water temperature is set to 70 ˚C during the Eco Drum Clean cycle. You cannot change the water temperature. However, the temperature for other cycles is not specified in the context.
|
||||
```
|
||||
|
||||
And that's it! Feel free to scale this up to as many documents and complex PDFs as you like.
|
||||
@@ -0,0 +1,232 @@
|
||||
---
|
||||
title: Multitenancy with LlamaIndex
|
||||
weight: 18
|
||||
aliases:
|
||||
- /documentation/tutorials/llama-index-multitenancy/
|
||||
---
|
||||
|
||||
# Multitenancy with LlamaIndex
|
||||
|
||||
If you are building a service that serves vectors for many independent users, and you want to isolate their
|
||||
data, the best practice is to use a single collection with payload-based partitioning. This approach is
|
||||
called **multitenancy**. Our guide on the [Separate Partitions](/documentation/guides/multiple-partitions/) describes
|
||||
how to set it up in general, but if you use [LlamaIndex](/documentation/integrations/llama-index/) as a
|
||||
backend, you may prefer reading a more specific instruction. So here it is!
|
||||
|
||||
## Prerequisites
|
||||
|
||||
This tutorial assumes that you have already installed Qdrant and LlamaIndex. If you haven't, please run the
|
||||
following commands:
|
||||
|
||||
```bash
|
||||
pip install llama-index llama-index-vector-stores-qdrant
|
||||
```
|
||||
|
||||
We are going to use a local Docker-based instance of Qdrant. If you want to use a remote instance, please
|
||||
adjust the code accordingly. Here is how we can start a local instance:
|
||||
|
||||
```bash
|
||||
docker run -d --name qdrant -p 6333:6333 -p 6334:6334 qdrant/qdrant:latest
|
||||
```
|
||||
|
||||
## Setting up LlamaIndex pipeline
|
||||
|
||||
We are going to implement an end-to-end example of multitenant application using LlamaIndex. We'll be
|
||||
indexing the documentation of different Python libraries, and we definitely don't want any users to see the
|
||||
results coming from a library they are not interested in. In real case scenarios, this is even more dangerous,
|
||||
as the documents may contain sensitive information.
|
||||
|
||||
### Creating vector store
|
||||
|
||||
[QdrantVectorStore](https://docs.llamaindex.ai/en/stable/examples/vector_stores/QdrantIndexDemo.html) is a
|
||||
wrapper around Qdrant that provides all the necessary methods to work with your vector database in LlamaIndex.
|
||||
Let's create a vector store for our collection. It requires setting a collection name and passing an instance
|
||||
of `QdrantClient`.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient
|
||||
from llama_index.vector_stores.qdrant import QdrantVectorStore
|
||||
|
||||
|
||||
client = QdrantClient("http://localhost:6333")
|
||||
|
||||
vector_store = QdrantVectorStore(
|
||||
collection_name="my_collection",
|
||||
client=client,
|
||||
)
|
||||
```
|
||||
|
||||
### Defining chunking strategy and embedding model
|
||||
|
||||
Any semantic search application requires a way to convert text queries into vectors - an embedding model.
|
||||
`ServiceContext` is a bundle of commonly used resources used during the indexing and querying stage in any
|
||||
LlamaIndex application. We can also use it to set up an embedding model - in our case, a local
|
||||
[BAAI/bge-small-en-v1.5](https://huggingface.co/BAAI/bge-small-en-v1.5).
|
||||
set up
|
||||
|
||||
```python
|
||||
from llama_index.core import ServiceContext
|
||||
|
||||
service_context = ServiceContext.from_defaults(
|
||||
embed_model="local:BAAI/bge-small-en-v1.5",
|
||||
)
|
||||
```
|
||||
*Note*, in case you are using Large Language Model different from OpenAI's ChatGPT, you should specify
|
||||
`llm` parameter for `ServiceContext`.
|
||||
|
||||
We can also control how our documents are split into chunks, or nodes using LLamaIndex's terminology.
|
||||
The `SimpleNodeParser` splits documents into fixed length chunks with an overlap. The defaults are
|
||||
reasonable, but we can also adjust them if we want to. Both values are defined in tokens.
|
||||
|
||||
```python
|
||||
from llama_index.core.node_parser import SimpleNodeParser
|
||||
|
||||
node_parser = SimpleNodeParser.from_defaults(chunk_size=512, chunk_overlap=32)
|
||||
```
|
||||
|
||||
Now we also need to inform the `ServiceContext` about our choices:
|
||||
|
||||
```python
|
||||
service_context = ServiceContext.from_defaults(
|
||||
embed_model="local:BAAI/bge-large-en-v1.5",
|
||||
node_parser=node_parser,
|
||||
)
|
||||
```
|
||||
|
||||
Both embedding model and selected node parser will be implicitly used during the indexing and querying.
|
||||
|
||||
### Combining everything together
|
||||
|
||||
The last missing piece, before we can start indexing, is the `VectorStoreIndex`. It is a wrapper around
|
||||
`VectorStore` that provides a convenient interface for indexing and querying. It also requires a
|
||||
`ServiceContext` to be initialized.
|
||||
|
||||
```python
|
||||
from llama_index.core import VectorStoreIndex
|
||||
|
||||
index = VectorStoreIndex.from_vector_store(
|
||||
vector_store=vector_store, service_context=service_context
|
||||
)
|
||||
```
|
||||
|
||||
## Indexing documents
|
||||
|
||||
No matter how our documents are generated, LlamaIndex will automatically split them into nodes, if
|
||||
required, encode using selected embedding model, and then store in the vector store. Let's define
|
||||
some documents manually and insert them into Qdrant collection. Our documents are going to have
|
||||
a single metadata attribute - a library name they belong to.
|
||||
|
||||
```python
|
||||
from llama_index.core.schema import Document
|
||||
|
||||
documents = [
|
||||
Document(
|
||||
text="LlamaIndex is a simple, flexible data framework for connecting custom data sources to large language models.",
|
||||
metadata={
|
||||
"library": "llama-index",
|
||||
},
|
||||
),
|
||||
Document(
|
||||
text="Qdrant is a vector database & vector similarity search engine.",
|
||||
metadata={
|
||||
"library": "qdrant",
|
||||
},
|
||||
),
|
||||
]
|
||||
```
|
||||
|
||||
Now we can index them using our `VectorStoreIndex`:
|
||||
|
||||
```python
|
||||
for document in documents:
|
||||
index.insert(document)
|
||||
```
|
||||
|
||||
### Performance considerations
|
||||
|
||||
Our documents have been split into nodes, encoded using the embedding model, and stored in the vector
|
||||
store. However, we don't want to allow our users to search for all the documents in the collection,
|
||||
but only for the documents that belong to a library they are interested in. For that reason, we need
|
||||
to set up the Qdrant [payload index](/documentation/concepts/indexing/#payload-index), so the search
|
||||
is more efficient.
|
||||
|
||||
```python
|
||||
from qdrant_client import models
|
||||
|
||||
client.create_payload_index(
|
||||
collection_name="my_collection",
|
||||
field_name="metadata.library",
|
||||
field_type=models.PayloadSchemaType.KEYWORD,
|
||||
)
|
||||
```
|
||||
|
||||
The payload index is not the only thing we want to change. Since none of the search
|
||||
queries will be executed on the whole collection, we can also change its configuration, so the HNSW
|
||||
graph is not built globally. This is also done due to [performance reasons](/documentation/guides/multiple-partitions/#calibrate-performance).
|
||||
**You should not be changing these parameters, if you know there will be some global search operations
|
||||
done on the collection.**
|
||||
|
||||
```python
|
||||
client.update_collection(
|
||||
collection_name="my_collection",
|
||||
hnsw_config=models.HnswConfigDiff(payload_m=16, m=0),
|
||||
)
|
||||
```
|
||||
|
||||
Once both operations are completed, we can start searching for our documents.
|
||||
|
||||
<aside role="status">These steps are done just once, when you index your first documents!</aside>
|
||||
|
||||
## Querying documents with constraints
|
||||
|
||||
Let's assume we are searching for some information about large language models, but are only allowed to
|
||||
use Qdrant documentation. LlamaIndex has a concept of retrievers, responsible for finding the most
|
||||
relevant nodes for a given query. Our `VectorStoreIndex` can be used as a retriever, with some additional
|
||||
constraints - in our case value of the `library` metadata attribute.
|
||||
|
||||
```python
|
||||
from llama_index.core.vector_stores.types import MetadataFilters, ExactMatchFilter
|
||||
|
||||
qdrant_retriever = index.as_retriever(
|
||||
filters=MetadataFilters(
|
||||
filters=[
|
||||
ExactMatchFilter(
|
||||
key="library",
|
||||
value="qdrant",
|
||||
)
|
||||
]
|
||||
)
|
||||
)
|
||||
|
||||
nodes_with_scores = qdrant_retriever.retrieve("large language models")
|
||||
for node in nodes_with_scores:
|
||||
print(node.text, node.score)
|
||||
# Output: Qdrant is a vector database & vector similarity search engine. 0.60551536
|
||||
```
|
||||
|
||||
The description of Qdrant was the best match, even though it didn't mention large language models
|
||||
at all. However, it was the only document that belonged to the `qdrant` library, so there was no
|
||||
other choice. Let's try to search for something that is not present in the collection.
|
||||
|
||||
Let's define another retrieve, this time for the `llama-index` library:
|
||||
|
||||
```python
|
||||
llama_index_retriever = index.as_retriever(
|
||||
filters=MetadataFilters(
|
||||
filters=[
|
||||
ExactMatchFilter(
|
||||
key="library",
|
||||
value="llama-index",
|
||||
)
|
||||
]
|
||||
)
|
||||
)
|
||||
|
||||
nodes_with_scores = llama_index_retriever.retrieve("large language models")
|
||||
for node in nodes_with_scores:
|
||||
print(node.text, node.score)
|
||||
# Output: LlamaIndex is a simple, flexible data framework for connecting custom data sources to large language models. 0.63576734
|
||||
```
|
||||
|
||||
The results returned by both retrievers are different, due to the different constraints, so we implemented
|
||||
a real multitenant search application!
|
||||
@@ -0,0 +1,141 @@
|
||||
---
|
||||
title: "Inference with Mighty"
|
||||
short_description: "Mighty offers a speedy scalable embedding, a perfect fit for the speedy scalable Qdrant search. Let's combine them!"
|
||||
description: "We combine Mighty and Qdrant to create a semantic search service in Rust with just a few lines of code."
|
||||
weight: 17
|
||||
author: Andre Bogus
|
||||
author_link: https://llogiq.github.io
|
||||
date: 2023-06-01T11:24:20+01:00
|
||||
aliases:
|
||||
- /documentation/tutorials/mighty.md/
|
||||
keywords:
|
||||
- vector search
|
||||
- embeddings
|
||||
- mighty
|
||||
- rust
|
||||
- semantic search
|
||||
---
|
||||
|
||||
# Semantic Search with Mighty and Qdrant
|
||||
|
||||
Much like Qdrant, the [Mighty](https://max.io/) inference server is written in Rust and promises to offer low latency and high scalability. This brief demo combines Mighty and Qdrant into a simple semantic search service that is efficient, affordable and easy to setup. We will use [Rust](https://rust-lang.org) and our [qdrant\_client crate](https://docs.rs/qdrant_client) for this integration.
|
||||
|
||||
## Initial setup
|
||||
|
||||
For Mighty, start up a [docker container](https://hub.docker.com/layers/maxdotio/mighty-sentence-transformers/0.9.9/images/sha256-0d92a89fbdc2c211d927f193c2d0d34470ecd963e8179798d8d391a4053f6caf?context=explore) with an open port 5050. Just loading the port in a window shows the following:
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "sentence-transformers/all-MiniLM-L6-v2",
|
||||
"architectures": [
|
||||
"BertModel"
|
||||
],
|
||||
"model_type": "bert",
|
||||
"max_position_embeddings": 512,
|
||||
"labels": null,
|
||||
"named_entities": null,
|
||||
"image_size": null,
|
||||
"source": "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2"
|
||||
}
|
||||
```
|
||||
|
||||
Note that this uses the `MiniLM-L6-v2` model from Hugging Face. As per their website, the model "maps sentences & paragraphs to a 384 dimensional dense vector space and can be used for tasks like clustering or semantic search". The distance measure to use is cosine similarity.
|
||||
|
||||
Verify that mighty works by calling `curl https://<address>:5050/sentence-transformer?q=hello+mighty`. This will give you a result like (formatted via `jq`):
|
||||
|
||||
```json
|
||||
{
|
||||
"outputs": [
|
||||
[
|
||||
-0.05019686743617058,
|
||||
0.051746174693107605,
|
||||
0.048117730766534805,
|
||||
... (381 values skipped)
|
||||
]
|
||||
],
|
||||
"shape": [
|
||||
1,
|
||||
384
|
||||
],
|
||||
"texts": [
|
||||
"Hello mighty"
|
||||
],
|
||||
"took": 77
|
||||
}
|
||||
```
|
||||
|
||||
For Qdrant, follow our [cloud documentation](../../cloud/cloud-quick-start/) to spin up a [free tier](https://cloud.qdrant.io/). Make sure to retrieve an API key.
|
||||
|
||||
## Implement model API
|
||||
|
||||
For mighty, you will need a way to emit HTTP(S) requests. This version uses the [reqwest](https://docs.rs/reqwest) crate, so add the following to your `Cargo.toml`'s dependencies section:
|
||||
|
||||
```toml
|
||||
[dependencies]
|
||||
reqwest = { version = "0.11.18", default-features = false, features = ["json", "rustls-tls"] }
|
||||
```
|
||||
|
||||
Mighty offers a variety of model APIs which will download and cache the model on first use. For semantic search, use the `sentence-transformer` API (as in the above `curl` command). The Rust code to make the call is:
|
||||
|
||||
```rust
|
||||
use anyhow::anyhow;
|
||||
use reqwest::Client;
|
||||
use serde::Deserialize;
|
||||
use serde_json::Value as JsonValue;
|
||||
|
||||
#[derive(Deserialize)]
|
||||
struct EmbeddingsResponse {
|
||||
pub outputs: Vec<Vec<f32>>,
|
||||
}
|
||||
|
||||
pub async fn get_mighty_embedding(
|
||||
client: &Client,
|
||||
url: &str,
|
||||
text: &str
|
||||
) -> anyhow::Result<Vec<f32>> {
|
||||
let response = client.get(url).query(&[("text", text)]).send().await?;
|
||||
|
||||
if !response.status().is_success() {
|
||||
return Err(anyhow!(
|
||||
"Mighty API returned status code {}",
|
||||
response.status()
|
||||
));
|
||||
}
|
||||
|
||||
let embeddings: EmbeddingsResponse = response.json().await?;
|
||||
// ignore multiple embeddings at the moment
|
||||
embeddings.get(0).ok_or_else(|| anyhow!("mighty returned empty embedding"))
|
||||
}
|
||||
```
|
||||
|
||||
Note that mighty can return multiple embeddings (if the input is too long to fit the model, it is automatically split).
|
||||
|
||||
## Create embeddings and run a query
|
||||
|
||||
Use this code to create embeddings both for insertion and search. On the Qdrant side, take the embedding and run a query:
|
||||
|
||||
```rust
|
||||
use anyhow::anyhow;
|
||||
use qdrant_client::prelude::*;
|
||||
|
||||
pub const SEARCH_LIMIT: u64 = 5;
|
||||
const COLLECTION_NAME: &str = "mighty";
|
||||
|
||||
pub async fn qdrant_search_embeddings(
|
||||
qdrant_client: &QdrantClient,
|
||||
vector: Vec<f32>,
|
||||
) -> anyhow::Result<Vec<ScoredPoint>> {
|
||||
qdrant_client
|
||||
.search_points(&SearchPoints {
|
||||
collection_name: COLLECTION_NAME.to_string(),
|
||||
vector,
|
||||
limit: SEARCH_LIMIT,
|
||||
with_payload: Some(true.into()),
|
||||
..Default::default()
|
||||
})
|
||||
.await
|
||||
.map_err(|err| anyhow!("Failed to search Qdrant: {}", err))
|
||||
}
|
||||
```
|
||||
|
||||
You can convert the [`ScoredPoint`](https://docs.rs/qdrant-client/latest/qdrant_client/qdrant/struct.ScoredPoint.html)s to fit your desired output format.
|
||||
+310
@@ -0,0 +1,310 @@
|
||||
---
|
||||
title: RAG System for Employee Onboarding
|
||||
weight: 30
|
||||
aliases:
|
||||
- /documentation/tutorials/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/
|
||||
---
|
||||
|
||||
# RAG System for Employee Onboarding
|
||||
|
||||
Public websites are a great way to share information with a wide audience. However, finding the right information can be
|
||||
challenging, if you are not familiar with the website's structure or the terminology used. That's what the search bar is
|
||||
for, but it is not always easy to formulate a query that will return the desired results, if you are not yet familiar
|
||||
with the content. This is even more important in a corporate environment, and for the new employees, who are just
|
||||
starting to learn the ropes, and don't even know how to ask the right questions yet. You may have even the best intranet
|
||||
pages, but onboarding is more than just reading the documentation, it is about understanding the processes. Semantic
|
||||
search can help with finding right resources easier, but wouldn't it be easier to just chat with the website, like you
|
||||
would with a colleague?
|
||||
|
||||
Technological advancements have made it possible to interact with websites using natural language. This tutorial will
|
||||
guide you through the process of integrating [Cohere](https://cohere.com/)'s language models with Qdrant to enable
|
||||
natural language search on your documentation. We are going to use [Langchain](https://langchain.com/) as an
|
||||
orchestrator. Everything will be hosted on [Oracle Cloud Infrastructure (OCI)](https://www.oracle.com/cloud/), so you
|
||||
can scale your application as needed, and do not send your data to third parties. That is especially important when you
|
||||
are working with confidential or sensitive data.
|
||||
|
||||
## Building up the application
|
||||
|
||||
Our application will consist of two main processes: indexing and searching. Langchain will glue everything together,
|
||||
as we will use a few components, including Cohere and Qdrant, as well as some OCI services. Here is a high-level
|
||||
overview of the architecture:
|
||||
|
||||

|
||||
|
||||
### Prerequisites
|
||||
|
||||
Before we dive into the implementation, make sure to set up all the necessary accounts and tools.
|
||||
|
||||
#### Libraries
|
||||
|
||||
We are going to use a few Python libraries. Of course, Langchain will be our main framework, but the Cohere models on
|
||||
OCI are accessible via the [OCI SDK](https://docs.oracle.com/en-us/iaas/tools/python/2.125.1/). Let's install all the
|
||||
necessary libraries:
|
||||
|
||||
```shell
|
||||
pip install langchain oci qdrant-client
|
||||
```
|
||||
|
||||
#### Oracle Cloud
|
||||
|
||||
Our application will be fully running on Oracle Cloud Infrastructure (OCI). It's up to you to choose how you want to
|
||||
deploy your application. Qdrant Hybrid Cloud will be running in your [Kubernetes cluster running on Oracle Cloud
|
||||
(OKE)](https://www.oracle.com/cloud/cloud-native/container-engine-kubernetes/), so all the processes might be also
|
||||
deployed there. You can get started with signing up for an account on [Oracle Cloud](https://signup.cloud.oracle.com/).
|
||||
|
||||
Cohere models are available on OCI as a part of the [Generative AI
|
||||
Service](https://www.oracle.com/artificial-intelligence/generative-ai/generative-ai-service/). We need both the
|
||||
[Generation models](https://docs.oracle.com/en-us/iaas/Content/generative-ai/use-playground-generate.htm) and the
|
||||
[Embedding models](https://docs.oracle.com/en-us/iaas/Content/generative-ai/use-playground-embed.htm). Please follow the
|
||||
linked tutorials to grasp the basics of using Cohere models there.
|
||||
|
||||
Accessing the models programmatically requires knowing the compartment OCID. Please refer to the [documentation that
|
||||
describes how to find it](https://docs.oracle.com/en-us/iaas/Content/GSG/Tasks/contactingsupport_topic-Locating_Oracle_Cloud_Infrastructure_IDs.htm#Finding_the_OCID_of_a_Compartment).
|
||||
For the further reference, we will assume that the compartment OCID is stored in the environment variable:
|
||||
|
||||
```shell
|
||||
export COMPARTMENT_OCID="<your-compartment-ocid>"
|
||||
```
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
os.environ["COMPARTMENT_OCID"] = "<your-compartment-ocid>"
|
||||
```
|
||||
|
||||
#### Qdrant Hybrid Cloud
|
||||
|
||||
Qdrant Hybrid Cloud running on Oracle Cloud helps you build a solution without sending your data to external services.
|
||||
Our documentation provides a step-by-step guide on how to [deploy Qdrant Hybrid Cloud on Oracle
|
||||
Cloud](...).
|
||||
|
||||
[//]: # (TODO: add a correct link to the documentation deployment guide)
|
||||
|
||||
Qdrant will be running on a specific URL and access will be restricted by the API key. Make sure to store them both as
|
||||
environment variables as well:
|
||||
|
||||
```shell
|
||||
export QDRANT_URL="https://qdrant.example.com"
|
||||
export QDRANT_API_KEY="your-api-key"
|
||||
```
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
os.environ["QDRANT_URL"] = "https://qdrant.example.com"
|
||||
os.environ["QDRANT_API_KEY"] = "your-api-key"
|
||||
```
|
||||
|
||||
Let's create the collection that will store the indexed documents. We will use the `qdrant-client` library, and our
|
||||
collection will be named `oracle-cloud-website`. Our embedding model, `cohere.embed-english-v3.0`, produces embeddings
|
||||
of size 1024, and we have to specify that when creating the collection.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
location=os.environ.get("QDRANT_URL"),
|
||||
api_key=os.environ.get("QDRANT_API_KEY"),
|
||||
)
|
||||
client.create_collection(
|
||||
collection_name="oracle-cloud-website",
|
||||
vectors_config=models.VectorParams(
|
||||
size=1024,
|
||||
distance=models.Distance.COSINE,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
### Indexing process
|
||||
|
||||
We have all the necessary tools set up, so let's start with the indexing process. We will use the Cohere Embedding
|
||||
models to convert the text into vectors, and then store them in Qdrant. Langchain is integrated with OCI Generative AI
|
||||
Service, so we can easily access the models.
|
||||
|
||||
Our dataset will be fairly simple, as it will consist of the questions and answers from the [Oracle Cloud Free Tier
|
||||
FAQ page](https://www.oracle.com/cloud/free/faq/).
|
||||
|
||||

|
||||
|
||||
Questions and answers are presented in an HTML format, but we don't want to manually extract the text and adapt it for
|
||||
each subpage. Instead, we will use the `WebBaseLoader` that just loads the HTML content from given URL and converts it
|
||||
to text.
|
||||
|
||||
```python
|
||||
from langchain_community.document_loaders.web_base import WebBaseLoader
|
||||
|
||||
loader = WebBaseLoader("https://www.oracle.com/cloud/free/faq/")
|
||||
documents = loader.load()
|
||||
```
|
||||
|
||||
Our `documents` is a list with just a single element, which is the text of the whole page. We need to split it into
|
||||
meaningful parts, so we will use the `RecursiveCharacterTextSplitter` component. It will try to keep all paragraphs (and
|
||||
then sentences, and then words) together as long as possible, as those would generically seem to be the strongest
|
||||
semantically related pieces of text. The chunk size and overlap are both parameters that can be adjusted to fit the
|
||||
specific use case.
|
||||
|
||||
```python
|
||||
from langchain_text_splitters import RecursiveCharacterTextSplitter
|
||||
|
||||
splitter = RecursiveCharacterTextSplitter(chunk_size=300, chunk_overlap=100)
|
||||
split_documents = splitter.split_documents(documents)
|
||||
```
|
||||
|
||||
Our documents might be now indexed, but we need to convert them into vectors. Let's configure the embeddings so the
|
||||
`cohere.embed-english-v3.0` is used. Not all the regions support the Generative AI Service, so we need to specify the
|
||||
region where the models are stored. We will use the `us-chicago-1`, but please check the
|
||||
[documentation](https://docs.oracle.com/en-us/iaas/Content/generative-ai/overview.htm#regions) for the most up-to-date
|
||||
list of supported regions.
|
||||
|
||||
```python
|
||||
from langchain_community.embeddings.oci_generative_ai import OCIGenAIEmbeddings
|
||||
|
||||
embeddings = OCIGenAIEmbeddings(
|
||||
model_id="cohere.embed-english-v3.0",
|
||||
service_endpoint="https://inference.generativeai.us-chicago-1.oci.oraclecloud.com",
|
||||
compartment_id=os.environ.get("COMPARTMENT_OCID"),
|
||||
)
|
||||
```
|
||||
|
||||
Now we can embed the documents and store them in Qdrant. We will create an instance of `Qdrant` and add the split
|
||||
documents to the collection.
|
||||
|
||||
```python
|
||||
from langchain.vectorstores.qdrant import Qdrant
|
||||
|
||||
qdrant = Qdrant(
|
||||
client=client,
|
||||
collection_name="oracle-cloud-website",
|
||||
embeddings=embeddings,
|
||||
)
|
||||
|
||||
qdrant.add_documents(split_documents, batch_size=20)
|
||||
```
|
||||
|
||||
Our documents should be now indexed and ready for searching. Let's move to the next step.
|
||||
|
||||
### Speaking to the website
|
||||
|
||||
The intended method of interaction with the website is through the chatbot. Large Language Model, in our case [Cohere
|
||||
Command](https://cohere.com/command), will be answering user's questions based on the relevant documents that Qdrant
|
||||
will return using the question as a query. Our LLM is also hosted on OCI, so we can access it similarly to the embedding
|
||||
model:
|
||||
|
||||
```python
|
||||
from langchain_community.llms.oci_generative_ai import OCIGenAI
|
||||
|
||||
llm = OCIGenAI(
|
||||
model_id="cohere.command",
|
||||
service_endpoint="https://inference.generativeai.us-chicago-1.oci.oraclecloud.com",
|
||||
compartment_id=os.environ.get("COMPARTMENT_OCID"),
|
||||
)
|
||||
```
|
||||
|
||||
Connection to Qdrant might be established in the same way as we did during the indexing process. We can use it to create
|
||||
an instance of `RetrievalQA`, which implements the question-answering process.
|
||||
|
||||
```python
|
||||
from langchain.chains.retrieval_qa.base import RetrievalQA
|
||||
|
||||
retriever = qdrant.as_retriever()
|
||||
|
||||
retrieval_qa = RetrievalQA.from_chain_type(
|
||||
llm=llm,
|
||||
retriever=retriever,
|
||||
return_source_documents=True,
|
||||
)
|
||||
response = retrieval_qa.invoke({"query": "What is the Oracle Cloud Free Tier?"})
|
||||
```
|
||||
|
||||
The output of the `.invoke` method is a dictionary-like structure with the query and response, but we can also access
|
||||
the source documents used to generate the response. This might be useful for debugging or for further processing.
|
||||
|
||||
```python
|
||||
{
|
||||
"query": "What is the Oracle Cloud Free Tier?",
|
||||
"result": " The Oracle Cloud Free Tier is a subscription that gives you access to Oracle Cloud's various services, including Always Free services and a Free Trial with $300 of free credit that can be used on all eligible Oracle Cloud Infrastructure services for up to 30 days. It is designed to allow users to learn, explore, build, and test in the Oracle Cloud environment for free. \n\nThis is particularly aimed at those who want to experiment with cloud capabilities, such as:\n- Developers who want to try building and deploying cloud-based applications before committing to a paid plan.\n- Students and academics who want to learn cloud computing and practice with hands-on exercises. \n\nThe Free Tier is initially available in most regions where commercial Oracle Cloud Infrastructure services are available; specific regions may vary during the sign-up process. You can use the $300 free credits for a limited time, and Always Free services are unlimited but without SLAs or Oracle Support. \n\nPlease note that I am an AI chatbot, and I do not have access to real-time information. My knowledge only covers details up to January 2023. If you want the most up-to-date information on the Oracle Cloud Free Tier, you can visit Oracle's official website for the latest details. ",
|
||||
"source_documents": [
|
||||
Document(
|
||||
page_content="* Free Tier is generally available in regions where commercial Oracle Cloud Infrastructure service is available. See the data regions page for detailed service availability (the exact regions available for Free Tier may differ during the sign-up process). The US$300 cloud credit is available in",
|
||||
metadata={
|
||||
"language": "en-US",
|
||||
"source": "https://www.oracle.com/cloud/free/faq/",
|
||||
"title": "FAQ on Oracle's Cloud Free Tier",
|
||||
"_id": "a20bada5-def8-4e6e-af87-b7b5cbd08dc7",
|
||||
"_collection_name": "oracle-cloud-website"
|
||||
}
|
||||
),
|
||||
Document(
|
||||
page_content="Oracle Cloud Free Tier allows you to sign up for an Oracle Cloud account which provides a number of Always Free services and a Free Trial with US$300 of free credit to use on all eligible Oracle Cloud Infrastructure services for up to 30 days. The Always Free services are available for an unlimited",
|
||||
metadata={
|
||||
"language": "en-US",
|
||||
"source": "https://www.oracle.com/cloud/free/faq/",
|
||||
"title": "FAQ on Oracle's Cloud Free Tier",
|
||||
"_id": "bba5f27a-e41e-4b69-9c79-76140523f600",
|
||||
"_collection_name": "oracle-cloud-website"
|
||||
}
|
||||
),
|
||||
Document(
|
||||
page_content="Oracle Cloud Free Tier does not include SLAs. Community support through our forums is available to all customers. Customers using only Always Free resources are not eligible for Oracle Support. Limited support is available for Oracle Cloud Free Tier with Free Trial credits. After you use all of",
|
||||
metadata={
|
||||
"language": "en-US",
|
||||
"source": "https://www.oracle.com/cloud/free/faq/",
|
||||
"title": "FAQ on Oracle's Cloud Free Tier",
|
||||
"_id": "e1873826-e6df-41b9-8dea-ec1de43bf633",
|
||||
"_collection_name": "oracle-cloud-website"
|
||||
}),
|
||||
Document(
|
||||
page_content="looking to test things before moving to cloud, a student wanting to learn, or an academic developing curriculum in the cloud, Oracle Cloud Free Tier enables you to learn, explore, build and test for free.",
|
||||
metadata={
|
||||
"language": "en-US",
|
||||
"source": "https://www.oracle.com/cloud/free/faq/",
|
||||
"title": "FAQ on Oracle's Cloud Free Tier",
|
||||
"_id": "73f17f07-c594-463b-9d55-663c7b7d54fc",
|
||||
"_collection_name": "oracle-cloud-website"
|
||||
}
|
||||
)
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
#### Other experiments
|
||||
|
||||
Asking the basic questions is just the beginning. What you want to avoid is a hallucination, where the model generates
|
||||
an answer that is not based on the actual content. The default prompt of Langchain should already prevent this, but you
|
||||
might still want to check it. Let's ask a question that is not directly answered on the FAQ page:
|
||||
|
||||
```python
|
||||
response = retrieval_qa.invoke({
|
||||
"query": "Is Oracle Generative AI Service included in the free tier?"
|
||||
})
|
||||
```
|
||||
|
||||
Output:
|
||||
|
||||
> Unfortunately, I don't know the answer to this, but it could be found on the company's website or in the provided text.
|
||||
>
|
||||
> I cannot search the internet since I lack an internet connection. If you would like, you are welcome to look for this
|
||||
> answer and share it with me.
|
||||
>
|
||||
> Otherwise, we can interpret the context to try and guess the answer.
|
||||
>
|
||||
> In general, it seems like the Oracle Cloud Free Tier includes a variety of free services that are available to use
|
||||
> indefinitely, and then additionally a trial with credits that last for up to 30 days. It seems like in order to get
|
||||
> continued support past the 30 days, you'd need to upgrade to a paid account.
|
||||
>
|
||||
> It is quite possible that Oracle Generative AI Service is included in the free tier for the 30 day trial period, but
|
||||
> not indefinitely.
|
||||
>
|
||||
> Unfortunately, I don't have the context or knowledge required to give you a certain answer, and this is just my best
|
||||
> guess based on what I have seen.
|
||||
|
||||
It seems that Cohere Command model could not find the exact answer in the provided documents, but it tried to interpret
|
||||
the context and provide a reasonable answer, without making up the information. This is a good sign that the model is
|
||||
not hallucinating in that case.
|
||||
|
||||
## Wrapping up
|
||||
|
||||
This tutorial has shown how to integrate Cohere's language models with Qdrant to enable natural language search on your
|
||||
website. We have used Langchain as an orchestrator, and everything was hosted on Oracle Cloud Infrastructure (OCI).
|
||||
Real world would require integrating this mechanism into your organization's systems, but we built a solid foundation
|
||||
that can be further developed.
|
||||
+459
@@ -0,0 +1,459 @@
|
||||
---
|
||||
title: Private Chatbot for Interactive Learning
|
||||
weight: 23
|
||||
aliases:
|
||||
- /documentation/tutorials/rag-chatbot-red-hat-openshift-haystack/
|
||||
---
|
||||
|
||||
# Private Chatbot for Interactive Learning
|
||||
|
||||
| Time: 120 min | Level: Advanced | Output: [GitHub](https://github.com/qdrant/) |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
With chatbots, companies can scale their training programs to accommodate a large workforce, delivering consistent and standardized learning experiences across departments, locations, and time zones. Furthermore, having already completed their online training, corporate employees might want to refer back old course materials. Most of this information is proprietary to the company, and manually searching through an entire library of materials takes time. However, a chatbot built on this knowledge can respond in the blink of an eye.
|
||||
|
||||
With a simple RAG pipeline, you can build a private chatbot. In this tutorial, you will combine open source tools inside of a closed infrastructure and tie them together with a reliable framework. This custom solution lets you run a chatbot without public internet access. You will be able to keep sensitive data secure without compromising privacy.
|
||||
|
||||

|
||||
**Figure 1:** The LLM and Qdrant Hybrid Cloud are containerized as separate services. Haystack combines them into a RAG pipeline and exposes the API via Hayhooks.
|
||||
|
||||
## Components
|
||||
To maintain complete data isolation, we need to limit ourselves to open-source tools and use them in a private environment, such as [Red Hat OpenShift](https://www.redhat.com/en/technologies/cloud-computing/openshift). The pipeline will run internally and will be inaccessible from the internet.
|
||||
|
||||
- **Dataset:** [Red Hat Interactive Learning Portal](https://developers.redhat.com/learn), an online library of RedHat course materials.
|
||||
- **LLM:** `mistralai/Mistral-7B-Instruct-v0.1`, deployed as a standalone service on OpenShift.
|
||||
- **Embedding Model:** `BAAI/bge-m3`, lightweight embedding model deployed from within the Haystack pipeline.
|
||||
- **Vector DB:** [Qdrant Hybrid Cloud](https://qdrant.tech) running on OpenShift.
|
||||
- **Framework:** [Haystack 2.x](https://haystack.deepset.ai/) to connect all and [Hayhooks](https://docs.haystack.deepset.ai/docs/hayhooks) to serve the app through HTTP endpoints.
|
||||
|
||||
### Procedure
|
||||
The [Haystack](https://haystack.deepset.ai/) framework leverages two pipelines, which combine our components sequentially to process data.
|
||||
|
||||
1. The **Indexing Pipeline** will run offline in batches, when new data is added or updated.
|
||||
2. The **Search Pipeline** will retrieve information from Qdrant and use an LLM to produce an answer.
|
||||
|
||||
> **Note:** We will define the pipelines in Python and then export them to YAML format, so that [Hayhooks](https://docs.haystack.deepset.ai/docs/hayhooks) can run them as a web service.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
### Deploy the LLM to OpenShift
|
||||
|
||||
Follow the steps in [Chapter 6. Serving large language models](https://access.redhat.com/documentation/en-us/red_hat_openshift_ai_self-managed/2.5/html/working_on_data_science_projects/serving-large-language-models_serving-large-language-models#doc-wrapper). This will download the LLM from the [HuggingFace](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.1), and deploy it to OpenShift using a *single model serving platform*.
|
||||
|
||||
Your LLM service will have a URL, which you need to store as an environment variable.
|
||||
|
||||
```shell
|
||||
export INFERENCE_ENDPOINT_URL="http://mistral-service.default.svc.cluster.local"
|
||||
```
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
os.environ["INFERENCE_ENDPOINT_URL"] = "http://mistral-service.default.svc.cluster.local"
|
||||
```
|
||||
|
||||
### Launch Qdrant Hybrid Cloud
|
||||
|
||||
Complete **How to Set Up Qdrant on RedHat OpenShift**. When in Hybrid Cloud, your Qdrant instance is private and and its nodes run on the same OpenShift infrastructure as your other components.
|
||||
|
||||
Retrieve your Qdrant URL and API key and store them as environment variables:
|
||||
|
||||
```shell
|
||||
export QDRANT_URL="https://qdrant.example.com"
|
||||
export QDRANT_API_KEY="your-api-key"
|
||||
```
|
||||
|
||||
```python
|
||||
os.environ["QDRANT_URL"] = "https://qdrant.example.com"
|
||||
os.environ["QDRANT_API_KEY"] = "your-api-key"
|
||||
```
|
||||
## Implementation
|
||||
|
||||
We will first create an indexing pipeline to add documents to the system.
|
||||
Then, the search pipeline will retrieve relevant data from our documents.
|
||||
After the pipelines are tested, we will export them to YAML files.
|
||||
|
||||
### Indexing pipeline
|
||||
|
||||
[Haystack 2.x](https://haystack.deepset.ai/) comes packed with a lot of useful components, from data fetching, through
|
||||
HTML parsing, up to the vector storage. Before we start, there are a few Python packages that we need to install:
|
||||
|
||||
```shell
|
||||
pip install haystack-ai \
|
||||
qdrant-haystack \
|
||||
transformers \
|
||||
torch --index-url https://download.pytorch.org/whl/cpu
|
||||
```
|
||||
|
||||
<aside role="status">
|
||||
We set the index URL for PyTorch to CPU, as we are going to run the indexing process, including the embedding model, on
|
||||
the processor. If you have a compatible GPU, you can install the corresponding version of PyTorch (CUDA or ROCm).
|
||||
</aside>
|
||||
|
||||
Our environment is now ready, so we can jump right into the code. Let's define an empty pipeline and gradually add
|
||||
components to it:
|
||||
|
||||
```python
|
||||
from haystack import Pipeline
|
||||
|
||||
indexing_pipeline = Pipeline()
|
||||
```
|
||||
|
||||
#### Data fetching and conversion
|
||||
|
||||
In this step, we will use Haystack's `LinkContentFetcher` to download course content from a list of URLs and store it in Qdrant for retrieval.
|
||||
As we don't want to store raw HTML, this tool will extract text content from each webpage. Then, the fetcher will divide them into digestible chunks, since the documents might be pretty long.
|
||||
|
||||
Let's start with data fetching and text conversion:
|
||||
|
||||
```python
|
||||
from haystack.components.fetchers import LinkContentFetcher
|
||||
from haystack.components.converters import HTMLToDocument
|
||||
|
||||
fetcher = LinkContentFetcher()
|
||||
converter = HTMLToDocument()
|
||||
|
||||
indexing_pipeline.add_component("fetcher", fetcher)
|
||||
indexing_pipeline.add_component("converter", converter)
|
||||
```
|
||||
|
||||
Our pipeline knows there are two components, but they are not connected yet. We need to define the flow between them:
|
||||
|
||||
```python
|
||||
indexing_pipeline.connect("fetcher.streams", "converter.sources")
|
||||
```
|
||||
|
||||
Each component has a set of inputs and outputs which might be combined in a directed graph. The definitions of the
|
||||
inputs and outputs are usually provided in the documentation of the component. The `LinkContentFetcher` has the
|
||||
following parameters:
|
||||
|
||||

|
||||
|
||||
*Source: https://docs.haystack.deepset.ai/docs/linkcontentfetcher*
|
||||
|
||||
#### Chunking and creating the embeddings
|
||||
|
||||
We used `HTMLToDocument` to convert the HTML sources into `Document` instances of Haystack, which is a
|
||||
base class containing some data to be queried. However, a single document might be too long to be processed by the
|
||||
embedding model, and it also carries way too much information to make the search relevant.
|
||||
|
||||
Therefore, we need to split the document into smaller parts and convert them into embeddings.
|
||||
For this, we will use the `DocumentSplitter` and `HuggingFaceTEIDocumentEmbedder` pointed to our `BAAI/bge-m3` model:
|
||||
|
||||
```python
|
||||
from haystack.components.preprocessors import DocumentSplitter
|
||||
from haystack.components.embedders import HuggingFaceTEIDocumentEmbedder
|
||||
|
||||
splitter = DocumentSplitter(split_by="sentence", split_length=5, split_overlap=2)
|
||||
embedder = HuggingFaceTEIDocumentEmbedder(model="BAAI/bge-m3")
|
||||
|
||||
indexing_pipeline.add_component("splitter", splitter)
|
||||
indexing_pipeline.add_component("embedder", embedder)
|
||||
|
||||
indexing_pipeline.connect("converter.documents", "splitter.documents")
|
||||
indexing_pipeline.connect("splitter.documents", "embedder.documents")
|
||||
```
|
||||
|
||||
#### Writing data to Qdrant
|
||||
|
||||
The splitter will be producing chunks with a maximum length of 5 sentences, with an overlap of 2 sentences. Then, these
|
||||
smaller portions will be converted into embeddings.
|
||||
|
||||
Finally, we need to store our embeddings in Qdrant.
|
||||
|
||||
```python
|
||||
from haystack.utils import Secret
|
||||
from haystack_integrations.document_stores.qdrant import QdrantDocumentStore
|
||||
from haystack.components.writers import DocumentWriter
|
||||
|
||||
document_store = QdrantDocumentStore(
|
||||
os.environ["QDRANT_URL"],
|
||||
api_key=Secret.from_env_var("QDRANT_API_KEY"),
|
||||
index="red-hat-learning",
|
||||
return_embedding=True,
|
||||
embedding_dim=1024,
|
||||
)
|
||||
writer = DocumentWriter(document_store=document_store)
|
||||
|
||||
indexing_pipeline.add_component("writer", writer)
|
||||
|
||||
indexing_pipeline.connect("embedder.documents", "writer.documents")
|
||||
```
|
||||
|
||||
Our pipeline is now complete. Haystack comes with a handy visualization of the pipeline, so you can see and verify the
|
||||
connections between the components. It is displayed in the Jupyter notebook, but you can also export it to a file:
|
||||
|
||||
```python
|
||||
indexing_pipeline.draw("indexing_pipeline.png")
|
||||
```
|
||||
|
||||

|
||||
|
||||
#### Test the entire pipeline
|
||||
|
||||
We can finally run it on a list of URLs to index the content in Qdrant. We have a bunch of URLs to all the Red Hat
|
||||
OpenShift Foundations course lessons, so let's use them:
|
||||
|
||||
```python
|
||||
course_urls = [
|
||||
"https://developers.redhat.com/learn/openshift/foundations-openshift",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:openshift-and-developer-sandbox",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:overview-web-console",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:use-terminal-window-within-red-hat-openshift-web-console",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:install-application-source-code-github-repository-using-openshift-web-console",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:install-application-linux-container-image-repository-using-openshift-web-console",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:install-application-linux-container-image-using-oc-cli-tool",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:install-application-source-code-using-oc-cli-tool",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:scale-applications-using-openshift-web-console",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:scale-applications-using-oc-cli-tool",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:work-databases-openshift-using-oc-cli-tool",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:work-databases-openshift-web-console",
|
||||
"https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:view-performance-information-using-openshift-web-console",
|
||||
]
|
||||
|
||||
indexing_pipeline.run(data={
|
||||
"fetcher": {
|
||||
"urls": course_urls,
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
The execution might take a while, as the model needs to process all the documents. After the process is finished, we
|
||||
should have all the documents stored in Qdrant, ready for search. You should see a short summary of processed documents:
|
||||
|
||||
```shell
|
||||
{'writer': {'documents_written': 381}}
|
||||
```
|
||||
|
||||
### Search pipeline
|
||||
|
||||
Our documents are now indexed and ready for search. The next pipeline is a bit simpler, but we still need to define a
|
||||
few components. Let's start again with an empty pipeline:
|
||||
|
||||
```python
|
||||
search_pipeline = Pipeline()
|
||||
```
|
||||
|
||||
Our second process takes user input, converts it into embeddings and then searches for the most relevant documents
|
||||
using the query embedding. This might look familiar, but we arent working with `Document` instances
|
||||
anymore, since the query only accepts raw text. Thus, some of the components will be different, especially the embedder,
|
||||
as it has to accept a single string as an input and produce a single embedding as an output:
|
||||
|
||||
```python
|
||||
from haystack.components.embedders import HuggingFaceTEITextEmbedder
|
||||
from haystack_integrations.components.retrievers.qdrant import QdrantEmbeddingRetriever
|
||||
|
||||
query_embedder = HuggingFaceTEITextEmbedder(model="BAAI/bge-m3")
|
||||
retriever = QdrantEmbeddingRetriever(
|
||||
document_store=document_store, # The same document store as the one used for indexing
|
||||
top_k=3, # Number of documents to return
|
||||
)
|
||||
|
||||
search_pipeline.add_component("query_embedder", query_embedder)
|
||||
search_pipeline.add_component("retriever", retriever)
|
||||
|
||||
search_pipeline.connect("query_embedder.embedding", "retriever.query_embedding")
|
||||
```
|
||||
|
||||
#### Run a test query
|
||||
|
||||
If our goal was to just retrieve the relevant documents, we could stop here. Let's try the current pipeline on a simple
|
||||
query:
|
||||
|
||||
```python
|
||||
query = "How to install an application using the OpenShift web console?"
|
||||
|
||||
search_pipeline.run(data={
|
||||
"query_embedder": {
|
||||
"text": query
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
We set the `top_k` parameter to 3, so the retriever should return the three most relevant documents. Your output should look like this:
|
||||
|
||||
```text
|
||||
{
|
||||
'retriever': {
|
||||
'documents': [
|
||||
Document(id=d127499a751f01969e76874049de7bea2dae077184eed9129d98fc97844c4bf2, content: ' Enter the application’s name (Figure 8). Figure 8: Deleting an application using the OpenShift web...', meta: {'content_type': 'text/html', 'source_id': '2a0759f3ce4a37d9f5c2af9c0ffcc80879077c102fb8e41e576e04833c9d24ce', 'url': 'https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:install-application-linux-container-image-repository-using-openshift-web-console'}, score: 0.87855008),
|
||||
Document(id=59095a8e45e656ec7ead299907407c9e63819813ec4a044fb33d953f79448437, content: 'For example, OpenShift lets you install a web application directly from source code or from a conta...', meta: {'content_type': 'text/html', 'source_id': '97f3aed6ff6712d980d6a501def31752317134eabc9897f900f9e2190fbbe186', 'url': 'https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:openshift-and-developer-sandbox'}, score: 0.8763400299999999),
|
||||
Document(id=5204e1b6adf3cbd2a3984db7bc9768ac4f27e2786f1a2682745b33ab8f5f686e, content: ' You declared the URL of the application’s source code in GitHub, then instigated the build process ...', meta: {'content_type': 'text/html', 'source_id': 'a4c4cd62d07c0d9d240e3289d2a1cc0a3d1127ae70704529967f715601559089', 'url': 'https://developers.redhat.com/learning/learn:openshift:foundations-openshift/resource/resources:install-application-source-code-github-repository-using-openshift-web-console'}, score: 0.8730886)
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Generating the answer
|
||||
|
||||
Retrieval should serve more than just documents. Therefore, we will need to use an LLM to generate exact answers to our question.
|
||||
This is the final component of our second pipeline.
|
||||
|
||||
Haystack will create a prompt which adds your documents to the model's context.
|
||||
|
||||
```python
|
||||
from haystack.components.builders.prompt_builder import PromptBuilder
|
||||
from haystack.components.generators import HuggingFaceTGIGenerator
|
||||
|
||||
prompt_builder = PromptBuilder("""
|
||||
Given the following information, answer the question.
|
||||
|
||||
Context:
|
||||
{% for document in documents %}
|
||||
{{ document.content }}
|
||||
{% endfor %}
|
||||
|
||||
Question: {{ query }}
|
||||
""")
|
||||
llm = HuggingFaceTGIGenerator(
|
||||
model="mistralai/Mistral-7B-Instruct-v0.1",
|
||||
url=os.environ["INFERENCE_ENDPOINT_URL"],
|
||||
generation_kwargs={
|
||||
"max_new_tokens": 1000, # Allow longer responses
|
||||
},
|
||||
)
|
||||
|
||||
search_pipeline.add_component("prompt_builder", prompt_builder)
|
||||
search_pipeline.add_component("llm", llm)
|
||||
|
||||
search_pipeline.connect("retriever.documents", "prompt_builder.documents")
|
||||
search_pipeline.connect("prompt_builder.prompt", "llm.prompt")
|
||||
```
|
||||
|
||||
The `PromptBuilder` is a Jinja2 template that will be filled with the documents and the query. The
|
||||
`HuggingFaceTGIGenerator` connects to the LLM service and generates the answer. Let's run the pipeline again:
|
||||
|
||||
```python
|
||||
query = "How to install an application using the OpenShift web console?"
|
||||
|
||||
response = search_pipeline.run(data={
|
||||
"query_embedder": {
|
||||
"text": query
|
||||
},
|
||||
"prompt_builder": {
|
||||
"query": query
|
||||
},
|
||||
})
|
||||
```
|
||||
|
||||
The LLM may provide multiple replies, if asked to do so, so let's iterate over and print them out:
|
||||
|
||||
```python
|
||||
for reply in response["llm"]["replies"]:
|
||||
print(reply.strip())
|
||||
```
|
||||
|
||||
In our case there is a single response, which should be the answer to the question:
|
||||
|
||||
```text
|
||||
Answer: To install an application using the OpenShift web console, you need to follow these steps:
|
||||
|
||||
1. Enter the application’s name.
|
||||
2. Access the Deploy Image web page.
|
||||
3. Declare the URL for a container image hosted on a public container image repository.
|
||||
4. Click the Create button.
|
||||
5. OpenShift downloads the container image and creates a Linux container using that container image.
|
||||
6. View the application by using a URL that OpenShift creates.
|
||||
```
|
||||
|
||||
Our final search pipeline might also be visualized, so we can see how the components are glued together:
|
||||
|
||||
```python
|
||||
search_pipeline.draw("search_pipeline.png")
|
||||
```
|
||||
|
||||

|
||||
|
||||
## Deployment
|
||||
|
||||
The pipelines are now ready, and we can export them to YAML. Hayhooks will use these files to run the
|
||||
pipelines as HTTP endpoints. To do this, specify both file paths and your environment variables.
|
||||
|
||||
> Note: The indexing pipeline might be run inside your ETL tool, but search should be definitely exposed as an HTTP endpoint.
|
||||
|
||||
Let's run it on the local machine:
|
||||
|
||||
```shell
|
||||
pip install hayhooks
|
||||
```
|
||||
|
||||
First of all, we need to save the pipelines to the YAML file:
|
||||
|
||||
```python
|
||||
with open("search-pipeline.yaml", "w") as fp:
|
||||
search_pipeline.dump(fp)
|
||||
```
|
||||
|
||||
And now we are able to run the Hayhooks service:
|
||||
|
||||
```shell
|
||||
hayhooks run
|
||||
```
|
||||
|
||||
The command should start the service on the default port, so you can access it at `http://localhost:1416`. The pipeline
|
||||
is not deployed yet, but we can do it with just another command:
|
||||
|
||||
```shell
|
||||
hayhooks deploy search-pipeline.yaml
|
||||
```
|
||||
|
||||
Once it's finished, you should be able to see the OpenAPI documentation at
|
||||
[http://localhost:1416/docs](http://localhost:1416/docs), and test the newly created endpoint.
|
||||
|
||||

|
||||
|
||||
Our search is now accessible through the HTTP endpoint, so we can integrate it with any other service. We can even
|
||||
control the other parameters, like the number of documents to return:
|
||||
|
||||
```shell
|
||||
curl -X 'POST' \
|
||||
'http://localhost:1416/search-pipeline' \
|
||||
-H 'Accept: application/json' \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"llm": {
|
||||
},
|
||||
"prompt_builder": {
|
||||
"query": "How can I remove an application?"
|
||||
},
|
||||
"query_embedder": {
|
||||
"text": "How can I remove an application?"
|
||||
},
|
||||
"retriever": {
|
||||
"top_k": 5
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
The response should be similar to the one we got in the Python before:
|
||||
|
||||
```json
|
||||
{
|
||||
"llm": {
|
||||
"replies": [
|
||||
"\n\nAnswer: You can remove an application running in OpenShift by right-clicking on the circular graphic representing the application in Topology view and selecting the Delete Application text from the dialog that appears when you click the graphic’s outer ring. Alternatively, you can use the oc CLI tool to delete an installed application using the oc delete all command."
|
||||
],
|
||||
"meta": [
|
||||
{
|
||||
"model": "mistralai/Mistral-7B-Instruct-v0.1",
|
||||
"index": 0,
|
||||
"finish_reason": "eos_token",
|
||||
"usage": {
|
||||
"completion_tokens": 75,
|
||||
"prompt_tokens": 642,
|
||||
"total_tokens": 717
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Next steps
|
||||
|
||||
- In this example, [RedHat OpenShift](https://www.redhat.com/en/technologies/cloud-computing/openshift) is the infrastructure of choice for proprietary chatbots. [Read more](https://access.redhat.com/documentation/en-us/red_hat_openshift_ai_self-managed/2.8) about how to host AI projects in their [extensive documentation](https://access.redhat.com/documentation/en-us/red_hat_openshift_ai_self-managed/2.8).
|
||||
|
||||
- [Haystack's documentation](https://docs.haystack.deepset.ai/docs/kubernetes) describes [how to deploy the Hayhooks service in a Kubernetes
|
||||
environment](https://docs.haystack.deepset.ai/docs/kubernetes), so you can easily move it to your own OpenShift infrastructure.
|
||||
|
||||
- If you are just getting started and need more guidance on Qdrant, read the [quickstart](https://qdrant.tech/documentation/quick-start/) or try out our [beginner tutorial](https://qdrant.tech/documentation/tutorials/neural-search/).
|
||||
@@ -0,0 +1,121 @@
|
||||
---
|
||||
title: Build a RAG-Based Chatbot on Scaleway
|
||||
weight: 35
|
||||
aliases:
|
||||
- /documentation/tutorials/rag-chatbot-scaleway/
|
||||
---
|
||||
|
||||
# Build a RAG-Based Chatbot on Scaleway
|
||||
|
||||
| Time: 90 min | Level: Advanced | | |
|
||||
|--------------|-----------------|--|----|
|
||||
|
||||
## Langchain x Qdrant: RAG Demo with Web Scraping
|
||||
|
||||
This section introduces the demonstration of building a Retrieval-Augmented Generation (RAG) model that combines web scraping with the capabilities of Langchain and Qdrant. The RAG model enhances the generation of answers by first retrieving relevant documents. Qdrant serves as the vector search engine for retrieval, while GPT-3.5, developed by OpenAI, is utilized as the generator for producing answers. This setup showcases the integration of advanced search and AI language processing to improve information retrieval and generation tasks.
|
||||
|
||||
|
||||
|
||||
## Prerequisites
|
||||
|
||||
To prepare the environment for working with Qdrant and related libraries, it's necessary to install all required Python packages. This can be done using Poetry, a tool for dependency management and packaging in Python. The code snippet imports various libraries essential for the tasks ahead, including `bs4` for parsing HTML and XML documents, `langchain` and its community extensions for working with language models and document loaders, and `Qdrant` for vector storage and retrieval. These imports lay the groundwork for utilizing Qdrant alongside other tools for natural language processing and machine learning tasks.
|
||||
|
||||
```python
|
||||
import getpass
|
||||
import os
|
||||
|
||||
import bs4
|
||||
from langchain import hub
|
||||
from langchain_community.document_loaders import WebBaseLoader
|
||||
from langchain_community.vectorstores import Qdrant
|
||||
from langchain_core.output_parsers import StrOutputParser
|
||||
from langchain_core.runnables import RunnablePassthrough
|
||||
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
|
||||
from langchain_text_splitters import RecursiveCharacterTextSplitter
|
||||
```
|
||||
|
||||
### Setting Up the OpenAI API Key
|
||||
|
||||
```python
|
||||
os.environ["OPENAI_API_KEY"] = getpass.getpass()
|
||||
```
|
||||
|
||||
### Initializing the Language Model
|
||||
|
||||
```python
|
||||
llm = ChatOpenAI(model="gpt-3.5-turbo-0125")
|
||||
```
|
||||
|
||||
It is here that we configure both the Embeddings and LLM. You can replace this with your own models using Ollama or other services. Scaleway has some great [GPU Instances](https://www.scaleway.com/en/gpu-instances/) too - including H100 on the higher end, and soon L4 for everything small.
|
||||
|
||||
## Download and Index
|
||||
|
||||
To begin working with blog post contents, the process involves loading and parsing the HTML content. This is achieved using `urllib` and `BeautifulSoup`, which are tools designed for such tasks. After the content is loaded and parsed, it is indexed using Qdrant, a powerful tool for managing and querying vector data. The code snippet demonstrates how to load, chunk, and index the contents of a blog post by specifying the URL of the blog and the specific HTML elements to parse. This step is crucial for preparing the data for further processing and analysis with Qdrant.
|
||||
|
||||
```python
|
||||
# Load, chunk and index the contents of the blog.
|
||||
loader = WebBaseLoader(
|
||||
web_paths=("https://lilianweng.github.io/posts/2023-06-23-agent/",),
|
||||
bs_kwargs=dict(
|
||||
parse_only=bs4.SoupStrainer(
|
||||
class_=("post-content", "post-title", "post-header")
|
||||
)
|
||||
),
|
||||
)
|
||||
docs = loader.load()
|
||||
|
||||
```
|
||||
|
||||
### Chunking before Indexing
|
||||
|
||||
When dealing with large documents, such as a blog post exceeding 42,000 characters, it's crucial to manage the data efficiently for processing. Many models have a limited context window and struggle with long inputs, making it difficult to extract or find relevant information. To overcome this, the document is divided into smaller chunks. This approach enhances the model's ability to process and retrieve the most pertinent sections of the document effectively.
|
||||
|
||||
In this scenario, the document is split into chunks using the `RecursiveCharacterTextSplitter` with a specified chunk size and overlap. This method ensures that no critical information is lost between chunks. Following the splitting, these chunks are then indexed into Qdrant—a vector database for efficient similarity search and storage of embeddings. The `Qdrant.from_documents` function is utilized for indexing, with documents being the split chunks and embeddings generated through `OpenAIEmbeddings`. The entire process is facilitated within an in-memory database, signifying that the operations are performed without the need for persistent storage, and the collection is named "lilianweng" for reference.
|
||||
|
||||
This chunking and indexing strategy significantly improves the management and retrieval of information from large documents, making it a practical solution for handling extensive texts in data processing workflows.
|
||||
|
||||
```python
|
||||
text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
|
||||
splits = text_splitter.split_documents(docs)
|
||||
|
||||
vectorstore = Qdrant.from_documents(
|
||||
documents=splits, embedding=OpenAIEmbeddings(), location=":memory:", collection_name="lilianweng"
|
||||
)
|
||||
```
|
||||
|
||||
## Retrieve and Generate
|
||||
|
||||
In this section, the process of retrieving information and generating content using a vector store and a language model is outlined. The `vectorstore` is utilized as a retriever to fetch relevant documents based on vector similarity. The `hub.pull("rlm/rag-prompt")` function is used to pull a specific prompt from a repository, which is designed to work with retrieved documents and a question to generate a response.
|
||||
|
||||
The `format_docs` function formats the retrieved documents into a single string, preparing them for further processing. This formatted string, along with a question, is passed through a chain of operations. Firstly, the context (formatted documents) and the question are processed by the retriever and the prompt. Then, the result is fed into a large language model (`llm`) for content generation. Finally, the output is parsed into a string format using `StrOutputParser()`.
|
||||
|
||||
This chain of operations demonstrates a sophisticated approach to information retrieval and content generation, leveraging both the semantic understanding capabilities of vector search and the generative prowess of large language models.
|
||||
|
||||
```python
|
||||
# Retrieve and generate using the relevant snippets of the blog.
|
||||
retriever = vectorstore.as_retriever()
|
||||
prompt = hub.pull("rlm/rag-prompt")
|
||||
|
||||
|
||||
def format_docs(docs):
|
||||
return "\n\n".join(doc.page_content for doc in docs)
|
||||
|
||||
|
||||
rag_chain = (
|
||||
{"context": retriever | format_docs, "question": RunnablePassthrough()}
|
||||
| prompt
|
||||
| llm
|
||||
| StrOutputParser()
|
||||
)
|
||||
```
|
||||
|
||||
### Invoking the RAG Chain
|
||||
|
||||
```python
|
||||
rag_chain.invoke("What is Task Decomposition?")
|
||||
```
|
||||
|
||||
## Deploying Langchain Applications on Scaleway
|
||||
Scaleway has serverless [Functions](https://www.scaleway.com/en/serverless-functions/) and serverless [Jobs](https://www.scaleway.com/en/serverless-jobs/) -- ideal for embedding creation when doing a bulk operation.
|
||||
|
||||
Their French deployment regions e.g. France are excellent for network latency and data sovereignty. Need a GPU? [Render with P100](https://www.scaleway.com/en/gpu-render-instances/) is there for you.
|
||||
@@ -0,0 +1,313 @@
|
||||
---
|
||||
title: Private RAG Information Extraction Engine
|
||||
weight: 32
|
||||
aliases:
|
||||
- /documentation/tutorials/rag-chatbot-vultr-dspy-ollama/
|
||||
---
|
||||
|
||||
# Private RAG Information Extraction Engine
|
||||
|
||||
| Time: 90 min | Level: Advanced | | |
|
||||
|--------------|-----------------|--|----|
|
||||
|
||||
Handling private documents is a common task in many industries. Various businesses possess a large amount of
|
||||
unstructured data stored as huge files that must be processed and analyzed. Industry reports, financial analysis, legal
|
||||
documents, and many other documents are stored in PDF, Word, and other formats. Conversational chatbots built on top of
|
||||
RAG pipelines are one of the viable solutions for finding the relevant answers in such documents. However, if we want to
|
||||
extract structured information from these documents, and pass them to downstream systems, we need to use a different
|
||||
approach.
|
||||
|
||||
Information extraction is a process of structuring unstructured data into a format that can be easily processed by
|
||||
machines. In this tutorial, we will show you how to use [DSPy](https://dspy-docs.vercel.app/) to perform that process on
|
||||
a set of documents. Assuming we cannot send our data to an external service, we will use [Ollama](https://ollama.com/)
|
||||
to run our own LLM model on our premises, using [Vultr](https://www.vultr.com/) as a cloud provider. Qdrant, acting in
|
||||
this setup as a knowledge base providing the relevant pieces of documents for a given query, will also be hosted in the
|
||||
Hybrid Cloud mode on Vultr. The last missing piece, the DSPy application will be also running in the same environment.
|
||||
If you work in a regulated industry, or just need to keep your data private, this tutorial is for you.
|
||||
|
||||

|
||||
|
||||
## Configuring the environment
|
||||
|
||||
All the services we are going to use in this tutorial will be running on [Vultr Kubernetes
|
||||
Engine](https://www.vultr.com/kubernetes/). That gives us a lot of flexibility in terms of scaling and managing the
|
||||
resources. Before we go further, make sure you have a Vultr account and a Kubernetes cluster running. Please follow the
|
||||
[official documentation](https://docs.vultr.com/vultr-kubernetes-engine) to get everything up and running.
|
||||
|
||||
### Installing the necessary packages
|
||||
|
||||
We are going to need a couple of Python packages to run our application. They might be installed together with the
|
||||
`dspy-ai` package and `qdrant` extra:
|
||||
|
||||
```shell
|
||||
pip install dspy-ai[qdrant]
|
||||
```
|
||||
|
||||
### Qdrant Hybrid Cloud
|
||||
|
||||
Our documentation contains a comprehensive guide on how to set up Qdrant in the Hybrid Cloud mode on Vultr. Please
|
||||
follow it carefully to get your Qdrant instance up and running. Once it's done, we need to store the Qdrant URL and the
|
||||
API key in the environment variables. You can do it by running the following commands:
|
||||
|
||||
[//]: # (TODO: add a link to the Qdrant Hybrid Cloud documentation above)
|
||||
|
||||
```shell
|
||||
export QDRANT_URL="https://qdrant.example.com"
|
||||
export QDRANT_API_KEY="your-api-key"
|
||||
```
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
os.environ["QDRANT_URL"] = "https://qdrant.example.com"
|
||||
os.environ["QDRANT_API_KEY"] = "your-api-key"
|
||||
```
|
||||
|
||||
DSPy is framework we are going to use. It's integrated with Qdrant already, but it assumes you use
|
||||
[FastEmbed](https://qdrant.github.io/fastembed/) to create the embeddings. DSPy does not provide a way to index the
|
||||
data, but leaves this task to the user. We are going to create a collection on our own, and fill it with the embeddings
|
||||
of our document chunks.
|
||||
|
||||
#### Data indexing
|
||||
|
||||
FastEmbed uses the `BAAI/bge-small-en` as the default embedding model. We are going to use it as well. Our collection
|
||||
will be created automatically if we call the `.add` method on an existing `QdrantClient` instance. In this tutorial we
|
||||
are not going to focus much on the document parsing, as there are plenty of tools that can help with that. The
|
||||
[`unstructured`](https://github.com/Unstructured-IO/unstructured) library is one of the options you can launch on your
|
||||
infrastructure. In our simplified example, we are going to use a list of strings as our documents. These are the
|
||||
descriptions of the made up technical events. Each of them should contain the name of the event along with the location
|
||||
and start and end dates.
|
||||
|
||||
```python
|
||||
documents = [
|
||||
"Taking place in San Francisco, USA, from the 10th to the 12th of June, 2024, the Global Developers Conference is the annual gathering spot for developers worldwide, offering insights into software engineering, web development, and mobile applications.",
|
||||
"The AI Innovations Summit, scheduled for 15-17 September 2024 in London, UK, aims at professionals and researchers advancing artificial intelligence and machine learning.",
|
||||
"Berlin, Germany will host the CyberSecurity World Conference between November 5th and 7th, 2024, serving as a key forum for cybersecurity professionals to exchange strategies and research on threat detection and mitigation.",
|
||||
"Data Science Connect in New York City, USA, occurring from August 22nd to 24th, 2024, connects data scientists, analysts, and engineers to discuss data science's innovative methodologies, tools, and applications.",
|
||||
"Set for July 14-16, 2024, in Tokyo, Japan, the Frontend Developers Fest invites developers to delve into the future of UI/UX design, web performance, and modern JavaScript frameworks.",
|
||||
"The Blockchain Expo Global, happening May 20-22, 2024, in Dubai, UAE, focuses on blockchain technology's applications, opportunities, and challenges for entrepreneurs, developers, and investors.",
|
||||
"Singapore's Cloud Computing Summit, scheduled for October 3-5, 2024, is where IT professionals and cloud experts will convene to discuss strategies, architectures, and cloud solutions.",
|
||||
"The IoT World Forum, taking place in Barcelona, Spain from December 1st to 3rd, 2024, is the premier conference for those focused on the Internet of Things, from smart cities to IoT security.",
|
||||
"Los Angeles, USA, will become the hub for game developers, designers, and enthusiasts at the Game Developers Arcade, running from April 18th to 20th, 2024, to showcase new games and discuss development tools.",
|
||||
"The TechWomen Summit in Sydney, Australia, from March 8-10, 2024, aims to empower women in tech with workshops, keynotes, and networking opportunities.",
|
||||
"Seoul, South Korea's Mobile Tech Conference, happening from September 29th to October 1st, 2024, will explore the future of mobile technology, including 5G networks and app development trends.",
|
||||
"The Open Source Summit, to be held in Helsinki, Finland from August 11th to 13th, 2024, celebrates open source technologies and communities, offering insights into the latest software and collaboration techniques.",
|
||||
"Vancouver, Canada will play host to the VR/AR Innovation Conference from June 20th to 22nd, 2024, focusing on the latest in virtual and augmented reality technologies.",
|
||||
"Scheduled for May 5-7, 2024, in London, UK, the Fintech Leaders Forum brings together experts to discuss the future of finance, including innovations in blockchain, digital currencies, and payment technologies.",
|
||||
"The Digital Marketing Summit, set for April 25-27, 2024, in New York City, USA, is designed for marketing professionals and strategists to discuss digital marketing and social media trends.",
|
||||
"EcoTech Symposium in Paris, France, unfolds over 2024-10-09 to 2024-10-11, spotlighting sustainable technologies and green innovations for environmental scientists, tech entrepreneurs, and policy makers.",
|
||||
"Set in Tokyo, Japan, from 16th to 18th May '24, the Robotic Innovations Conference showcases automation, robotics, and AI-driven solutions, appealing to enthusiasts and engineers.",
|
||||
"The Software Architecture World Forum in Dublin, Ireland, occurring 22-24 Sept 2024, gathers software architects and IT managers to discuss modern architecture patterns.",
|
||||
"Quantum Computing Summit, convening in Silicon Valley, USA from 2024/11/12 to 2024/11/14, is a rendezvous for exploring quantum computing advancements with physicists and technologists.",
|
||||
"From March 3 to 5, 2024, the Global EdTech Conference in London, UK, discusses the intersection of education and technology, featuring e-learning and digital classrooms.",
|
||||
"Bangalore, India's NextGen DevOps Days, from 28 to 30 August 2024, is a hotspot for IT professionals keen on the latest DevOps tools and innovations.",
|
||||
"The UX/UI Design Conference, slated for April 21-23, 2024, in New York City, USA, invites discussions on the latest in user experience and interface design among designers and developers.",
|
||||
"Big Data Analytics Summit, taking place 2024 July 10-12 in Amsterdam, Netherlands, brings together data professionals to delve into big data analysis and insights.",
|
||||
"Toronto, Canada, will see the HealthTech Innovation Forum from June 8 to 10, '24, focusing on technology's impact on healthcare with professionals and innovators.",
|
||||
"Blockchain for Business Summit, happening in Singapore from 2024-05-02 to 2024-05-04, focuses on blockchain's business applications, from finance to supply chain.",
|
||||
"Las Vegas, USA hosts the Global Gaming Expo from October 18th to 20th, 2024, a premiere event for game developers, publishers, and enthusiasts.",
|
||||
"The Renewable Energy Tech Conference in Copenhagen, Denmark, from 2024/09/05 to 2024/09/07, discusses renewable energy innovations and policies.",
|
||||
"Set for 2024 Apr 9-11 in Boston, USA, the Artificial Intelligence in Healthcare Summit gathers healthcare professionals to discuss AI's healthcare applications.",
|
||||
"Nordic Software Engineers Conference, happening in Stockholm, Sweden from June 15 to 17, 2024, focuses on software development in the Nordic region.",
|
||||
"The International Space Exploration Symposium, scheduled in Houston, USA from 2024-08-05 to 2024-08-07, invites discussions on space exploration technologies and missions."
|
||||
]
|
||||
```
|
||||
|
||||
We'll be able to ask general questions, for example, about topics we are interested in or events happening in a specific
|
||||
location, but expect the results to be returned in a structured format.
|
||||
|
||||

|
||||
|
||||
Indexing in Qdrant is a single call if we have the documents defined:
|
||||
|
||||
```python
|
||||
client.add(
|
||||
collection_name="document-parts",
|
||||
documents=documents,
|
||||
metadata=[{"document": document} for document in documents],
|
||||
)
|
||||
```
|
||||
|
||||
Our collection is ready to be queried. We can now move to the next step, which is setting up the Ollama model.
|
||||
|
||||
### Ollama on Vultr
|
||||
|
||||
Ollama is a great tool for running the LLM models on your own infrastructure. It's designed to be lightweight and easy
|
||||
to use, and [an official Docker image](https://hub.docker.com/r/ollama/ollama) is available. We can use it to run Ollama
|
||||
on our Vultr Kubernetes cluster. In case of LLMs we may have some special requirements, like a GPU, and Vultr provides
|
||||
the [Vultr Kubernetes Engine for Cloud GPU](https://www.vultr.com/products/cloud-gpu/) so the model can be run on a
|
||||
specialized machine. Please refer to the official documentation to get Ollama up and running within your environment.
|
||||
Once it's done, we need to store the Ollama URL in the environment variable:
|
||||
|
||||
```shell
|
||||
export OLLAMA_URL="https://ollama.example.com"
|
||||
```
|
||||
|
||||
```python
|
||||
os.environ["OLLAMA_URL"] = "https://ollama.example.com"
|
||||
```
|
||||
|
||||
We will refer to this URL later on when configuring the Ollama model in our application.
|
||||
|
||||
#### Setting up the Large Language Model
|
||||
|
||||
We are going to use one of the lightweight LLMs available in Ollama, a `gemma:2b` model. It was developed by Google
|
||||
DeepMind team and has 3B parameters. The [Ollama version](https://ollama.com/library/gemma:2b) uses 4-bit quantization.
|
||||
Installing the model is as simple as running the following command on the machine where Ollama is running:
|
||||
|
||||
```shell
|
||||
ollama run gemma:2b
|
||||
```
|
||||
|
||||
Ollama models are also integrated with DSPy, so we can use them directly in our application.
|
||||
|
||||
## Implementing the information extraction pipeline
|
||||
|
||||
DSPy is a bit different from the other LLM frameworks. It's designed to optimize the prompts and weights of LMs in a
|
||||
pipeline. It's a bit like a compiler for LMs: you write a pipeline in a high-level language, and DSPy generates the
|
||||
prompts and weights for you. This means you can build complex systems without having to worry about the details of how
|
||||
to prompt your LMs, as DSPy will do that for you. It is somehow similar to PyTorch but for LLMs.
|
||||
|
||||
First of all, we will define the Language Model we are going to use:
|
||||
|
||||
```python
|
||||
import dspy
|
||||
|
||||
gemma_model = dspy.OllamaLocal(
|
||||
model="gemma:2b",
|
||||
base_url=os.environ.get("OLLAMA_URL"),
|
||||
max_tokens=500,
|
||||
)
|
||||
```
|
||||
|
||||
Similarly, we have to define connection to our Qdrant Hybrid Cloud cluster:
|
||||
|
||||
```python
|
||||
from dspy.retrieve.qdrant_rm import QdrantRM
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
os.environ.get("QDRANT_URL"),
|
||||
api_key=os.environ.get("QDRANT_API_KEY"),
|
||||
)
|
||||
qdrant_retriever = QdrantRM(
|
||||
qdrant_collection_name="document-parts",
|
||||
qdrant_client=client,
|
||||
)
|
||||
```
|
||||
|
||||
Finally, both components have to be configured in DSPy with a simple call to one of the functions:
|
||||
|
||||
```python
|
||||
dspy.configure(lm=gemma_model, rm=qdrant_retriever)
|
||||
```
|
||||
|
||||
### Application logic
|
||||
|
||||
There is a concept of signatures which defines input and output formats of the pipeline. We are going to define a simple
|
||||
signature for the event:
|
||||
|
||||
```python
|
||||
class Event(dspy.Signature):
|
||||
description = dspy.InputField(
|
||||
desc="Textual description of the event, including name, location and dates"
|
||||
)
|
||||
event_name = dspy.OutputField(desc="Name of the event")
|
||||
location = dspy.OutputField(desc="Location of the event")
|
||||
start_date = dspy.OutputField(desc="Start date of the event, YYYY-MM-DD")
|
||||
end_date = dspy.OutputField(desc="End date of the event, YYYY-MM-DD")
|
||||
```
|
||||
|
||||
It is designed to derive the structured information from the textual description of the event. Now, we can build our
|
||||
module that will use it, along with Qdrant and Ollama model. Let's call it `EventExtractor`:
|
||||
|
||||
```python
|
||||
class EventExtractor(dspy.Module):
|
||||
|
||||
def __init__(self):
|
||||
super().__init__()
|
||||
# Retrieve module to get relevant documents
|
||||
self.retriever = dspy.Retrieve(k=3)
|
||||
# Predict module for the created signature
|
||||
self.predict = dspy.Predict(Event)
|
||||
|
||||
def forward(self, query: str):
|
||||
# Retrieve the most relevant documents
|
||||
results = self.retriever.forward(query)
|
||||
|
||||
# Try to extract events from the retrieved documents
|
||||
events = []
|
||||
for document in results.passages:
|
||||
event = self.predict(description=document)
|
||||
events.append(event)
|
||||
|
||||
return events
|
||||
```
|
||||
|
||||
The logic is simple: we retrieve the most relevant documents from Qdrant, and then try to extract the structured
|
||||
information from them using the `Event` signature. We can simply call it and see the results:
|
||||
|
||||
```python
|
||||
extractor = EventExtractor()
|
||||
extractor.forward("Blockchain events close to Europe")
|
||||
```
|
||||
|
||||
Output:
|
||||
|
||||
```python
|
||||
[
|
||||
Prediction(
|
||||
event_name='Event Name: Blockchain Expo Global',
|
||||
location='Dubai, UAE',
|
||||
start_date='2024-05-20',
|
||||
end_date='2024-05-22'
|
||||
),
|
||||
Prediction(
|
||||
event_name='Event Name: Blockchain for Business Summit',
|
||||
location='Singapore',
|
||||
start_date='2024-05-02',
|
||||
end_date='2024-05-04'
|
||||
),
|
||||
Prediction(
|
||||
event_name='Event Name: Open Source Summit',
|
||||
location='Helsinki, Finland',
|
||||
start_date='2024-08-11',
|
||||
end_date='2024-08-13'
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
The task was solved successfully, even without any optimization. However, each of the events has the "Event Name: "
|
||||
prefix that we might want to remove. DSPy allows optimizing the module, so we can improve the results. Optimization
|
||||
might be done in different ways, and it's [well covered in the DSPy
|
||||
documentation](https://dspy-docs.vercel.app/docs/building-blocks/optimizers).
|
||||
|
||||
We are not going to go through the optimization process in this tutorial. However, we encourage you to experiment with
|
||||
it, as it might significantly improve the performance of your pipeline.
|
||||
|
||||
Created module might be easily stored on a specific path, and loaded later on:
|
||||
|
||||
```python
|
||||
extractor.save("event_extractor")
|
||||
```
|
||||
|
||||
To load, just create an instance of the module and call the `load` method:
|
||||
|
||||
```python
|
||||
second_extractor = EventExtractor()
|
||||
second_extractor.load("event_extractor")
|
||||
```
|
||||
|
||||
This is especially useful when you optimize the module, as the optimized version might be stored and loaded later on
|
||||
without redoing the optimization process each time you run the application.
|
||||
|
||||
### Deploying the extraction pipeline
|
||||
|
||||
Vultr gives us a lot of flexibility in terms of deploying the applications. Perfectly, we would use the Kubernetes
|
||||
cluster we set up earlier to run it. The deployment is as simple as running any other Python application. This time we
|
||||
don't need a GPU, as Ollama is already running on a separate machine, and DSPy just interacts with it.
|
||||
|
||||
## Wrapping up
|
||||
|
||||
In this tutorial, we showed you how to set up a private environment for information extraction using DSPy, Ollama, and
|
||||
Qdrant. All the components might be securely hosted on the Vultr cloud, giving you full control over your data.
|
||||
+302
@@ -0,0 +1,302 @@
|
||||
---
|
||||
title: Region-Specific Contract Management System
|
||||
weight: 28
|
||||
aliases:
|
||||
- /documentation/tutorials/rag-contract-management-stackit-aleph-alpha/
|
||||
---
|
||||
|
||||
# Region-Specific Contract Management System
|
||||
|
||||
| Time: 90 min | Level: Advanced | |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
Contract management benefits greatly from Retrieval Augmented Generation (RAG), streamlining the handling of lengthy business contract texts. With AI assistance, complex questions can be asked and well-informed answers generated, facilitating efficient document management. This proves invaluable for businesses with extensive relationships, like shipping companies, construction firms, and consulting practices. Access to such contracts is often restricted to authorized team members due to security and regulatory requirements, such as GDPR in Europe, necessitating secure storage practices.
|
||||
|
||||
Companies want their data to be kept and processed within specific geographical boundaries. For that reason, this RAG-centric tutorial focuses on dealing with a region-specific cloud provider. You will set up a contract management system using [Aleph Alpha's](https://aleph-alpha.com/) embeddings and LLM. You will host everything on [STACKIT](https://www.stackit.de/), a German business cloud provider. On this platform, you will run Qdrant Hybrid Cloud as well as the rest of your RAG application. This setup will ensure that your data is stored and processed in Germany.
|
||||
|
||||
[//]: # (TODO: add link to Qdrant Hybrid Cloud above)
|
||||
|
||||

|
||||
|
||||
## Components
|
||||
|
||||
A contract management platform is not a simple CLI tool, but an application that should be available to all team
|
||||
members. It needs an interface to upload, search, and manage the documents. Ideally, the system should be
|
||||
integrated with org's existing stack, and the permissions/access controls inherited from LDAP or Active
|
||||
Directory.
|
||||
|
||||
> **Note:** In this tutorial, we are going to build a solid foundation for such a system. However, it is up to your organization's setup to implement the entire solution.
|
||||
|
||||
- **Dataset** - a collection of documents, using different formats, such as PDF or DOCx, scraped from internet
|
||||
- **Asymmetric semantic embeddings** - [Aleph Alpha embedding](https://docs.aleph-alpha.com/api/semantic-embed/) to
|
||||
convert the queries and the documents into vectors
|
||||
- **Large Language Model** - the [Luminous-extended-control
|
||||
model](https://docs.aleph-alpha.com/docs/introduction/model-card/), but you can play with a different one from the
|
||||
Luminous family
|
||||
- **Qdrant Hybrid Cloud** - a knowledge base to store the vectors and search over the documents
|
||||
- **STACKIT** - a [German business cloud](https://www.stackit.de) to run the Qdrant Hybrid Cloud and the application
|
||||
processes
|
||||
|
||||
We will implement the process of uploading the documents, converting them into vectors, and storing them in Qdrant.
|
||||
Then, we will build a search interface to query the documents and get the answers. All that, assuming the user
|
||||
interacts with the system with some set of permissions, and can only access the documents they are allowed to.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
### Aleph Alpha account
|
||||
|
||||
Since you will be using Aleph Alpha's models, [sign up](https://app.aleph-alpha.com/signup) with their managed service and generate an API token in the [User Profile](https://app.aleph-alpha.com/profile). Once you have it ready, store it as an environment variable:
|
||||
|
||||
```shell
|
||||
export ALEPH_ALPHA_API_KEY="<your-token>"
|
||||
```
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
os.environ["ALEPH_ALPHA_API_KEY"] = "<your-token>"
|
||||
```
|
||||
|
||||
### Qdrant Hybrid Cloud on STACKIT
|
||||
|
||||
Please refer to our documentation to see how to deploy Qdrant Hybrid Cloud on STACKIT. Once you finish the deployment,
|
||||
you will have the API endpoint to interact with the Qdrant server. Let's store it in the environment variable as well:
|
||||
|
||||
[//]: # (TODO: refer to the documentation on how to deploy Qdrant on Stackit)
|
||||
|
||||
```shell
|
||||
export QDRANT_URL="https://qdrant.example.com"
|
||||
export QDRANT_API_KEY="your-api-key"
|
||||
```
|
||||
|
||||
```python
|
||||
os.environ["QDRANT_URL"] = "https://qdrant.example.com"
|
||||
os.environ["QDRANT_API_KEY"] = "your-api-key"
|
||||
```
|
||||
|
||||
## Implementation
|
||||
|
||||
To build the application, we can use the official SDKs of Aleph Alpha and Qdrant. However, to streamline the process but let's use [Langchain](https://python.langchain.com/docs/get_started/introduction). This framework is already integrated with both services, so we can focus our efforts on developing business logic.
|
||||
|
||||
### Qdrant collection
|
||||
|
||||
Aleph Alpha embeddings are high dimensional vectors by default, with a dimensionality of 5120. Qdrant can store such
|
||||
vector easily, but that also sounds like a good idea to enable [Binary
|
||||
Quantization](../../../documentation/guides/quantization/#binary-quantization) to save space and make the retrieval
|
||||
faster. Let's create a collection with such settings:
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(
|
||||
location=os.environ["QDRANT_URL"],
|
||||
api_key=os.environ["QDRANT_API_KEY"],
|
||||
)
|
||||
client.create_collection(
|
||||
collection_name="contracts",
|
||||
vectors_config=models.VectorParams(
|
||||
size=5120,
|
||||
distance=models.Distance.COSINE,
|
||||
quantization_config=models.BinaryQuantization(
|
||||
binary=models.BinaryQuantizationConfig(
|
||||
always_ram=True,
|
||||
)
|
||||
)
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
We are going to use the `contracts` collection to store the vectors of the documents. The `always_ram` flag is set to
|
||||
`True` to keep the quantized vectors in RAM, which will speed up the search process. We also wanted to restrict access
|
||||
to the individual documents, so only users with the proper permissions can see them. In Qdrant that should be solved by
|
||||
adding a payload field that defines who can access the document. We'll call this field `roles` and set it to an array
|
||||
of strings with the roles that can access the document.
|
||||
|
||||
```python
|
||||
client.create_payload_index(
|
||||
collection_name="contracts",
|
||||
field_name="metadata.roles",
|
||||
field_schema=models.PayloadSchemaType.KEYWORD,
|
||||
)
|
||||
```
|
||||
|
||||
Since we use Langchain, the `roles` field is a nested field of the `metadata`, so we have to define it as
|
||||
`metadata.roles`. The schema says that the field is a keyword, which means it is a string or an array of strings. We are
|
||||
going to use the name of the customers as the roles, so the access control will be based on the customer name.
|
||||
|
||||
### Ingestion pipeline
|
||||
|
||||
Semantic search systems rely on high-quality data as their foundation. With the [unstructured integration of Langchain](https://python.langchain.com/docs/integrations/providers/unstructured), ingestion of various document formats like PDFs, Microsoft Word files, and PowerPoint presentations becomes effortless. However, it's crucial to split the text intelligently to avoid converting entire documents into vectors; instead, they should be divided into meaningful chunks. Subsequently, the extracted documents are converted into vectors using Aleph Alpha embeddings and stored in the Qdrant collection.
|
||||
|
||||
Let's start by defining the components and connecting them together:
|
||||
|
||||
```python
|
||||
embeddings = AlephAlphaAsymmetricSemanticEmbedding(
|
||||
model="luminous-base",
|
||||
aleph_alpha_api_key=os.environ["ALEPH_ALPHA_API_KEY"],
|
||||
normalize=True,
|
||||
)
|
||||
|
||||
qdrant = Qdrant(
|
||||
client=client,
|
||||
collection_name="contracts",
|
||||
embeddings=embeddings,
|
||||
)
|
||||
```
|
||||
|
||||
Now it's high time to index our documents. Each of the documents is a separate file, and we also have to know the
|
||||
customer name to set the access control properly. There might be several roles for a single document, so let's keep them
|
||||
in a list.
|
||||
|
||||
```python
|
||||
documents = {
|
||||
"data/Data-Processing-Agreement_STACKIT_Cloud_version-1.2.pdf": ["stackit"],
|
||||
"data/langchain-terms-of-service.pdf": ["langchain"],
|
||||
}
|
||||
```
|
||||
|
||||
This is how the documents might look like:
|
||||
|
||||

|
||||
|
||||
Each has to be split into chunks first; there is no silver bullet. Our chunking algorithm will be simple and based on
|
||||
recursive splitting, with the maximum chunk size of 500 characters and the overlap of 100 characters.
|
||||
|
||||
```python
|
||||
from langchain_text_splitters import RecursiveCharacterTextSplitter
|
||||
|
||||
text_splitter = RecursiveCharacterTextSplitter(
|
||||
chunk_size=500,
|
||||
chunk_overlap=100,
|
||||
)
|
||||
```
|
||||
|
||||
Now we can iterate over the documents, split them into chunks, convert them into vectors with Aleph Alpha embedding
|
||||
model, and store them in the Qdrant.
|
||||
|
||||
```python
|
||||
from langchain_community.document_loaders.unstructured import UnstructuredFileLoader
|
||||
|
||||
for document_path, roles in documents.items():
|
||||
document_loader = UnstructuredFileLoader(file_path=document_path)
|
||||
|
||||
# Unstructured loads each file into a single Document object
|
||||
loaded_documents = document_loader.load()
|
||||
for doc in loaded_documents:
|
||||
doc.metadata["roles"] = roles
|
||||
|
||||
# Chunks will have the same metadata as the original document
|
||||
document_chunks = text_splitter.split_documents(loaded_documents)
|
||||
|
||||
# Add the documents to the Qdrant collection
|
||||
qdrant.add_documents(document_chunks, batch_size=20)
|
||||
```
|
||||
|
||||
Our collection is filled with data, and we can start searching over it. In a real-world scenario, the ingestion process
|
||||
should be automated and triggered by the new documents uploaded to the system. Since we already use Qdrant Hybrid Cloud
|
||||
running on Kubernetes, we can easily deploy the ingestion pipeline as a job to the same environment. On STACKIT, you
|
||||
probably use the [STACKIT Kubernetes Engine (SKE)](https://www.stackit.de/en/product/kubernetes/) and launch it in a
|
||||
container. The [Compute Engine](https://www.stackit.de/en/product/stackit-compute-engine/) is also an option, but
|
||||
everything depends on the specifics of your organization.
|
||||
|
||||
### Search application
|
||||
|
||||
Specialized Document Management Systems have a lot of features, but semantic search is not yet a standard. We are going
|
||||
to build a simple search mechanism which could be possibly integrated with the existing system. The search process is
|
||||
quite simple: we convert the query into a vector using the same Aleph Alpha model, and then search for the most similar
|
||||
documents in the Qdrant collection. The access control is also applied, so the user can only see the documents they are
|
||||
allowed to.
|
||||
|
||||
We start with creating an instance of the LLM of our choice, and set the maximum number of tokens to 200, as the default
|
||||
value is 64, which might be too low for our purposes.
|
||||
|
||||
```python
|
||||
from langchain.llms.aleph_alpha import AlephAlpha
|
||||
|
||||
llm = AlephAlpha(
|
||||
model="luminous-extended-control",
|
||||
aleph_alpha_api_key=os.environ["ALEPH_ALPHA_API_KEY"],
|
||||
maximum_tokens=200,
|
||||
)
|
||||
```
|
||||
|
||||
Then, we can glue the components together and build the search process. `RetrievalQA` is a class that takes implements
|
||||
the Question Retrieval process, with a specified retriever and Large Language Model. The instance of `Qdrant` might be
|
||||
converted into a retriever, with additional filter that will be passed to the `similarity_search` method. The filter
|
||||
is created as [in a regular Qdrant query](../../../documentation/concepts/filtering/), with the `roles` field set to the
|
||||
user's roles.
|
||||
|
||||
```python
|
||||
user_roles = ["stackit", "aleph-alpha"]
|
||||
|
||||
qdrant_retriever = qdrant.as_retriever(
|
||||
search_kwargs={
|
||||
"filter": models.Filter(
|
||||
must=[
|
||||
models.FieldCondition(
|
||||
key="metadata.roles",
|
||||
match=models.MatchAny(any=user_roles)
|
||||
)
|
||||
]
|
||||
)
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
We set the user roles to `stackit` and `aleph-alpha`, so the user can see the documents that are accessible to these
|
||||
customers, but not to the others. The final step is to create the `RetrievalQA` instance and use it to search over the
|
||||
documents, with the custom prompt.
|
||||
|
||||
```python
|
||||
from langchain.prompts import PromptTemplate
|
||||
from langchain.chains.retrieval_qa.base import RetrievalQA
|
||||
|
||||
prompt_template = """
|
||||
### Instruction:
|
||||
{question} If there's no answer, say "Provided context does not clarify it".
|
||||
### Input:
|
||||
Text:{context}
|
||||
Question:{question}
|
||||
### Response:
|
||||
"""
|
||||
prompt = PromptTemplate(
|
||||
template=prompt_template, input_variables=["context", "question"]
|
||||
)
|
||||
|
||||
retrieval_qa = RetrievalQA.from_chain_type(
|
||||
llm=llm,
|
||||
chain_type="stuff",
|
||||
retriever=qdrant_retriever,
|
||||
return_source_documents=True,
|
||||
chain_type_kwargs={"prompt": prompt},
|
||||
)
|
||||
|
||||
response = retrieval_qa.invoke({"query": "What are the rules of performing the audit?"})
|
||||
print(response["result"])
|
||||
```
|
||||
|
||||
Output:
|
||||
|
||||
```text
|
||||
The rules for performing the audit are as follows:
|
||||
|
||||
1. The Customer must inform the Contractor in good time (usually at least two weeks in advance) about any and all circumstances related to the performance of the audit.
|
||||
2. The Customer is entitled to perform one audit per calendar year. Any additional audits may be performed if agreed with the Contractor and are subject to reimbursement of expenses.
|
||||
3. If the Customer engages a third party to perform the audit, the Customer must obtain the Contractor's consent and ensure that the confidentiality agreements with the third party are observed.
|
||||
4. The Contractor may object to any third party deemed unsuitable.
|
||||
```
|
||||
|
||||
There are some other parameters that might be tuned to optimize the search process. The `k` parameter defines how many
|
||||
documents should be returned, but Langchain allows us also to control the retrieval process by choosing the type of the
|
||||
search operation. The default is `similarity`, which is just vector search, but we can also use `mmr` which stands for
|
||||
Maximal Marginal Relevance. It is a technique to diversify the search results, so the user gets the most relevant
|
||||
documents, but also the most diverse ones. The `mmr` search is slower, but might be more user-friendly.
|
||||
|
||||
Our search application is ready, and we can deploy it to the same environment as the ingestion pipeline on STACKIT. The
|
||||
same rules apply here, so you can use the SKE or the Compute Engine, depending on the specifics of your organization.
|
||||
|
||||
## Next steps
|
||||
|
||||
We built a solid foundation for the contract management system, but there is still a lot to do. If you want to make the
|
||||
system production-ready, you should consider implementing the mechanism into your existing stack. If you have any
|
||||
questions, feel free to ask on our [Discord community](https://qdrant.to/discord).
|
||||
+234
@@ -0,0 +1,234 @@
|
||||
---
|
||||
title: Question-Answering System for AI Customer Support
|
||||
weight: 26
|
||||
aliases:
|
||||
- /documentation/tutorials/rag-customer-support-cohere-airbyte-aws/
|
||||
---
|
||||
|
||||
# Question-Answering System for AI Customer Support
|
||||
|
||||
| Time: 120 min | Level: Advanced | |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
Maintaining top-notch customer service is vital to business success. As your operation expands, so does the influx of customer queries. Many of these queries are repetitive, making automation a time-saving solution.
|
||||
Your support team's expertise is typically kept private, but you can still use AI to automate responses securely.
|
||||
|
||||
In this tutorial we will setup a private AI service that answers customer support queries with high accuracy and effectiveness. By leveraging Cohere's powerful models (deployed to [AWS](https://cohere.com/deployment-options/aws)) with Qdrant Hybrid Cloud, you can create a fully private customer support system. Data synchronization, facilitated by [Airbyte](https://airbyte.com/), will complete the setup.
|
||||
|
||||
[//]: # (TODO: add a link to the corresponding Qdrant Hybrid Cloud documentation: deployment on AWS)
|
||||
|
||||

|
||||
|
||||
## System design
|
||||
|
||||
The history of past interactions with your customers is not a static dataset. It is constantly evolving, as new
|
||||
questions are coming in. You probably have a ticketing system that stores all the interactions, or use a different way
|
||||
to communicate with your customers. No matter what is the communication channel, you need to bring the correct answers
|
||||
to the selected Large Language Model, and have an established way to do it in a continuous manner. Thus, we will build
|
||||
an ingestion pipeline and then a Retrieval Augmented Generation application that will use the data.
|
||||
|
||||
- **Dataset:** a [set of Frequently Asked Questions from Qdrant
|
||||
users](https://qdrant.tech/documentation/faq/qdrant-fundamentals/) as an incrementally updated Excel sheet
|
||||
- **Embedding model:** Cohere `embed-multilingual-v3.0`, to support different languages with the same pipeline
|
||||
- **Knowledge base:** Qdrant, running in Hybrid Cloud mode
|
||||
- **Ingestion pipeline:** [Airbyte](https://airbyte.com/), loading the data into Qdrant
|
||||
- **Large Language Model:** Cohere [Command-R](https://docs.cohere.com/docs/command-r)
|
||||
- **RAG:** Cohere [RAG](https://docs.cohere.com/docs/retrieval-augmented-generation-rag) using our knowledge base
|
||||
through a custom connector
|
||||
|
||||
All the selected components are compatible with the [AWS](https://aws.amazon.com/) infrastructure. Thanks to Cohere
|
||||
models' availability, you can build a fully private customer support system completely isolates data within your
|
||||
infrastructure. Also, if you have AWS credits, you can now use them without spending additional money on the models or
|
||||
semantic search layer.
|
||||
|
||||
### Data ingestion
|
||||
|
||||
Building a RAG starts with a well-curated dataset. In your specific case you may prefer loading the data directly from
|
||||
a ticketing system, such as [Zendesk Support](https://airbyte.com/connectors/zendesk-support),
|
||||
[Freshdesk](https://airbyte.com/connectors/freshdesk), or maybe integrate it with a shared inbox. However, in case of
|
||||
customer questions quality over quantity is the key. There should be a conscious decision on what data to include in the
|
||||
knowledge base, so we do not confuse the model with possibly irrelevant information. We'll assume there is an [Excel
|
||||
sheet](https://docs.airbyte.com/integrations/sources/file) available over HTTP/FTP that Airbyte can access and load into
|
||||
Qdrant in an incremental manner.
|
||||
|
||||
### Cohere <> Qdrant Connector for RAG
|
||||
|
||||
Cohere RAG relies on [connectors](https://docs.cohere.com/docs/connectors) which brings additional context to the model.
|
||||
The connector is a web service that implements a specific interface, and exposes its data through HTTP API. With that
|
||||
setup, the Large Language Model becomes responsible for communicating with the connectors, so building a prompt with the
|
||||
context is not needed anymore.
|
||||
|
||||
### Answering bot
|
||||
|
||||
Finally, we want to automate the responses and send them automatically when we are sure that the model is confident
|
||||
enough. Again, the way such an application should be created strongly depends on the system you are using within the
|
||||
customer support team. If it exposes a way to set up a webhook whenever a new question is coming in, you can create a
|
||||
web service and use it to automate the responses. In general, our bot should be created specifically for the platform
|
||||
you use, so we'll just cover the general idea here and build a simple CLI tool.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
### Cohere models on AWS
|
||||
|
||||
One of the possible ways to deploy Cohere models on AWS is to use AWS SageMaker. Cohere's website has [a detailed
|
||||
guide on how to deploy the models in that way](https://docs.cohere.com/docs/amazon-sagemaker-setup-guide), so you can
|
||||
follow the steps described there to set up your own instance.
|
||||
|
||||
### Qdrant Hybrid Cloud on AWS
|
||||
|
||||
Our documentation covers the deployment of Qdrant on AWS in your private region, so you can follow the steps described
|
||||
there to set up your own instance. The deployment process is quite straightforward, and you can have your Qdrant cluster
|
||||
up and running in a few minutes.
|
||||
|
||||
[//]: # (TODO: refer to the documentation on how to deploy Qdrant on AWS)
|
||||
|
||||
Once you perform all the steps, your Qdrant cluster should be running on a specific URL. You will need this URL and the
|
||||
API key to interact with Qdrant, so let's store them both in the environment variables:
|
||||
|
||||
```shell
|
||||
export QDRANT_URL="https://qdrant.example.com"
|
||||
export QDRANT_API_KEY="your-api-key"
|
||||
```
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
os.environ["QDRANT_URL"] = "https://qdrant.example.com"
|
||||
os.environ["QDRANT_API_KEY"] = "your-api-key"
|
||||
```
|
||||
|
||||
### Airbyte Open Source
|
||||
|
||||
Airbyte is an open-source data integration platform that helps you replicate your data in your warehouses, lakes, and
|
||||
databases. You can install it on your infrastructure and use it to load the data into Qdrant. The installation process
|
||||
for AWS EC2 is described in the [official documentation](https://docs.airbyte.com/deploying-airbyte/on-aws-ec2).
|
||||
Please follow the instructions to set up your own instance.
|
||||
|
||||
#### Setting up the connection
|
||||
|
||||
Once you have an Airbyte up and running, you can configure the connection to load the data from the respective source
|
||||
into Qdrant. The configuration will require setting up the source and destination connectors. In this tutorial we will
|
||||
use the following connectors:
|
||||
|
||||
- **Source:** [File](https://docs.airbyte.com/integrations/sources/file) to load the data from an Excel sheet
|
||||
- **Destination:** [Qdrant](https://docs.airbyte.com/integrations/destinations/qdrant) to load the data into Qdrant
|
||||
|
||||
Airbyte UI will guide you through the process of setting up the source and destination and connecting them. Here is how
|
||||
the configuration of the source might look like:
|
||||
|
||||

|
||||
|
||||
Qdrant is our target destination, so we need to set up the connection to it. We need to specify which fields should be
|
||||
included to generate the embeddings. In our case it makes complete sense to embed just the questions, as we are going
|
||||
to look for similar questions asked in the past and provide the answers.
|
||||
|
||||

|
||||
|
||||
Once we have the destination set up, we can finally configure a connection. The connection will define the schedule
|
||||
of the data synchronization.
|
||||
|
||||

|
||||
|
||||
Airbyte should now be ready to accept any data updates from the source and load them into Qdrant. You can monitor the
|
||||
progress of the synchronization in the UI.
|
||||
|
||||
## RAG connector
|
||||
|
||||
One of our previous tutorials, guides you step-by-step on [implementing custom connector for Cohere
|
||||
RAG](../cohere-rag-connector/) with Cohere Embed v3 and Qdrant. You can just point it to use your Hybrid Cloud
|
||||
Qdrant instance running on AWS. Created connector might be deployed to Amazon Web Services in various ways, even in a
|
||||
[Serverless](https://aws.amazon.com/serverless/) manner using [AWS
|
||||
Lambda](https://aws.amazon.com/lambda/?c=ser&sec=srv).
|
||||
|
||||
In general, RAG connector has to expose a single endpoint that will accept POST requests with `query` parameter and
|
||||
return the matching documents as JSON document with a specific structure. Our FastAPI implementation created [in the
|
||||
related tutorial](../cohere-rag-connector/) is a perfect fit for this task. The only difference is that you
|
||||
should point it to the Cohere models and Qdrant running on AWS infrastructure.
|
||||
|
||||
> Our connector is a lightweight web service that exposes a single endpoint and glues the Cohere embedding model with
|
||||
> our Qdrant Hybrid Cloud instance. Thus, it perfectly fits the serverless architecture, requiring no additional
|
||||
> infrastructure to run.
|
||||
|
||||
You can also run the connector as another service within your [Kubernetes cluster running on AWS
|
||||
(EKS)](https://aws.amazon.com/eks/), or by launching an [EC2](https://aws.amazon.com/ec2/) compute instance. This step
|
||||
is dependent on the way you deploy your other services, so we'll leave it to you to decide how to run the connector.
|
||||
|
||||
Eventually, the web service should be available under a specific URL, and it's a good practice to store it in the
|
||||
environment variable, so the other services can easily access it.
|
||||
|
||||
```shell
|
||||
export RAG_CONNECTOR_URL="https://rag-connector.example.com/search"
|
||||
```
|
||||
|
||||
```python
|
||||
os.environ["RAG_CONNECTOR_URL"] = "https://rag-connector.example.com/search"
|
||||
```
|
||||
|
||||
## Customer interface
|
||||
|
||||
At this part we have all the data loaded into Qdrant, and the RAG connector is ready to serve the relevant context. The
|
||||
last missing piece is the customer interface, that will call the Command model to create the answer. Such a system
|
||||
should be built specifically for the platform you use and integrated into its workflow, but we will build the strong
|
||||
foundation for it and show how to use it in a simple CLI tool.
|
||||
|
||||
> Our application does not have to connect to Qdrant anymore, as the model will connect to the RAG connector directly.
|
||||
|
||||
First of all, we have to create a connection to Cohere services through the Cohere SDK.
|
||||
|
||||
```python
|
||||
import cohere
|
||||
|
||||
# Create a Cohere client pointing to the AWS instance
|
||||
cohere_client = cohere.Client(...)
|
||||
```
|
||||
|
||||
Next, our connector should be registered. **Please make sure to do it once, and store the id of the connector in the
|
||||
environment variable or in any other way that will be accessible to the application.**
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
connector_response = cohere_client.connectors.create(
|
||||
name="customer-support",
|
||||
url=os.environ["RAG_CONNECTOR_URL"],
|
||||
)
|
||||
|
||||
# The id returned by the API should be stored for future use
|
||||
connector_id = connector_response.connector.id
|
||||
```
|
||||
|
||||
Finally, we can create a prompt and get the answer from the model. Additionally, we define which of the connectors
|
||||
should be used to provide the context, as we may have multiple connectors and want to use specific ones, depending on
|
||||
some conditions. Let's start with asking a question.
|
||||
|
||||
```python
|
||||
query = "Why Qdrant does not return my vectors?"
|
||||
```
|
||||
|
||||
Now we can send the query to the model, get the response, and possibly send it back to the customer.
|
||||
|
||||
```python
|
||||
response = cohere_client.chat(
|
||||
message=query,
|
||||
connectors=[
|
||||
cohere.ChatConnector(id=connector_id),
|
||||
],
|
||||
model="command-r",
|
||||
)
|
||||
|
||||
print(response.text)
|
||||
```
|
||||
|
||||
The output should be the answer to the question, generated by the model, for example:
|
||||
|
||||
> Qdrant is set up by default to minimize network traffic and therefore doesn't return vectors in search results. However, you can make Qdrant return your vectors by setting the 'with_vector' parameter of the Search/Scroll function to true.
|
||||
|
||||
Customer support should not be fully automated, as some completely new issues might require human intervention. We
|
||||
should play with prompt engineering and expect the model to provide the answer with a certain confidence level. If the
|
||||
confidence is too low, we should not send the answer automatically but present it to the support team for review.
|
||||
|
||||
## Wrapping up
|
||||
|
||||
This tutorial shows how to build a fully private customer support system using Cohere models, Qdrant Hybrid Cloud, and
|
||||
Airbyte, which runs on AWS infrastructure. You can ensure your data does not leave your premises and focus on providing
|
||||
the best customer support experience without bothering your team with repetitive tasks.
|
||||
@@ -0,0 +1,240 @@
|
||||
---
|
||||
title: Movie Recommendation System
|
||||
weight: 34
|
||||
aliases:
|
||||
- /documentation/tutorials/recommendation-system-ovhcloud/
|
||||
---
|
||||
|
||||
# Build a Movie Recommendation System
|
||||
|
||||
| Time: 120 min | Level: Advanced | Output: [GitHub](https://github.com/infoslack/qdrant-example/blob/main/HC-demo/HC-OVH.ipynb) |
|
||||
| --- | ----------- | ----------- |----------- |
|
||||
|
||||
This notebook aims to create a recommendation system using the MovieLens dataset and Qdrant. Vector databases like Qdrant are crucial for storing high-dimensional data, such as user and item embeddings, enabling personalized recommendations by quickly retrieving similar users or items based on advanced indexing techniques. We'll leverage collaborative filtering with a MovieLens dataset, identifying similar users based on ratings represented as vectors in Qdrant, and suggesting movies they liked but we haven't seen yet. The suggested items or content should closely align with the user's interests, leading to more personalized and relevant recommendations.
|
||||
|
||||
Collaborative filtering works on the principle that users with similar tastes will enjoy similar movies. To implement this, we'll represent each user's ratings as vectors in a high-dimensional space using Qdrant. By indexing these vectors, we can find users with similar tastes to ours and recommend movies they liked but we haven't seen yet.
|
||||
|
||||
## Components
|
||||
|
||||
- **Dataset:** [Red Hat Interactive Learning Portal](https://developers.redhat.com/learn)
|
||||
- **Vector DB:** [Qdrant Hybrid Cloud](https://qdrant.tech) running on OpenShift.
|
||||
- **Web Host:** [OVHcloud](https://haystack.deepset.ai/)
|
||||
|
||||
## Prerequisites
|
||||
|
||||
First, download and unzip the MovieLens dataset into a local directory.
|
||||
|
||||
```bash
|
||||
mkdir -p data
|
||||
wget https://files.grouplens.org/datasets/movielens/ml-1m.zip
|
||||
unzip ml-1m.zip -d data
|
||||
```
|
||||
|
||||
The necessary Python libraries are installed using `pip`, including `pandas` for data manipulation, `qdrant-client` for interfacing with Qdrant, and `python-dotenv` for managing environment variables.
|
||||
|
||||
```python
|
||||
!pip install -U \
|
||||
pandas \
|
||||
qdrant-client \
|
||||
python-dotenv
|
||||
```
|
||||
|
||||
The `.env` file is used to store sensitive information like the Qdrant host URL and API key securely.
|
||||
|
||||
```bash
|
||||
QDRANT_HOST
|
||||
QDRANT_API_KEY
|
||||
```
|
||||
Load all environment variables into the setup.
|
||||
|
||||
```python
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
load_dotenv('./.env')
|
||||
```
|
||||
|
||||
## Implementation
|
||||
|
||||
Load the user, movie, and rating data from the MovieLens dataset into pandas DataFrames to facilitate data manipulation and analysis.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
import pandas as pd
|
||||
```
|
||||
|
||||
```python
|
||||
# load users
|
||||
users = pd.read_csv('data/ml-1m/users.dat', sep='::', names=['user_id', 'gender', 'age', 'occupation', 'zip'], engine='python')
|
||||
users.head()
|
||||
```
|
||||
|
||||
```python
|
||||
# load movies
|
||||
movies = pd.read_csv('data/ml-1m/movies.dat', sep='::', names=['movie_id', 'title', 'genres'], engine='python', encoding='latin-1')
|
||||
movies.head()
|
||||
```
|
||||
|
||||
```python
|
||||
#load ratings
|
||||
ratings = pd.read_csv( 'data/ml-1m/ratings.dat', sep='::', names=['user_id', 'movie_id', 'rating', 'timestamp'], engine='python')
|
||||
ratings.head()
|
||||
```
|
||||
|
||||
**Normalize ratings**
|
||||
|
||||
Sparse vectors can use advantage of negative values, so we can normalize ratings to have a mean of 0 and a standard deviation of 1
|
||||
This normalization ensures that ratings are consistent and centered around zero, enabling accurate similarity calculations.
|
||||
In this scenario we can take into account movies that we don't like.
|
||||
|
||||
```python
|
||||
ratings.rating = (ratings.rating - ratings.rating.mean()) / ratings.rating.std()
|
||||
```
|
||||
|
||||
```python
|
||||
ratings.head()
|
||||
```
|
||||
|
||||
## Preparing the data and creating a collection
|
||||
|
||||
Transform user ratings into sparse vectors, where each vector represents ratings for different movies. This step prepares the data for indexing in Qdrant.
|
||||
|
||||
First, create a collection with configured sparse vectors
|
||||
- Sparse vectors don't require to specify dimension, because it's extracted from the data automatically
|
||||
|
||||
> An explanation of using hybrid cloud with OVH can be inserted here!
|
||||
|
||||
```python
|
||||
# Convert ratings to sparse vectors
|
||||
|
||||
from collections import defaultdict
|
||||
|
||||
user_sparse_vectors = defaultdict(lambda: {"values": [], "indices": []})
|
||||
|
||||
for row in ratings.itertuples():
|
||||
user_sparse_vectors[row.user_id]["values"].append(row.rating)
|
||||
user_sparse_vectors[row.user_id]["indices"].append(row.movie_id)
|
||||
```
|
||||
```python
|
||||
client = QdrantClient(
|
||||
url = os.getenv("QDRANT_HOST"),
|
||||
api_key = os.getenv("QDRANT_API_KEY")
|
||||
)
|
||||
|
||||
client.create_collection(
|
||||
"movielens",
|
||||
vectors_config={},
|
||||
sparse_vectors_config={
|
||||
"ratings": models.SparseVectorParams()
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
Upload user ratings to the "movielens" collection in Qdrant as sparse vectors, along with user metadata. This step populates the database with the necessary data for recommendation generation.
|
||||
|
||||
```python
|
||||
def data_generator():
|
||||
for user in users.itertuples():
|
||||
yield models.PointStruct(
|
||||
id=user.user_id,
|
||||
vector={
|
||||
"ratings": user_sparse_vectors[user.user_id]
|
||||
},
|
||||
payload=user._asdict()
|
||||
)
|
||||
|
||||
client.upload_points(
|
||||
"movielens",
|
||||
data_generator()
|
||||
)
|
||||
```
|
||||
|
||||
## Running the Recommendation System
|
||||
|
||||
Personal movie ratings are specified, where positive ratings indicate likes and negative ratings indicate dislikes. These ratings serve as the basis for finding similar users with comparable tastes.
|
||||
|
||||
Personal ratings are converted into a sparse vector representation suitable for querying Qdrant. This vector represents the user's preferences across different movies.
|
||||
|
||||
Let's try to recommend something for ourselves:
|
||||
|
||||
1 = Like
|
||||
-1 = dislike
|
||||
|
||||
Search with movies[movies.title.str.contains("Matrix", case=False)]
|
||||
|
||||
```python
|
||||
my_ratings = {
|
||||
2571: 1, # Matrix
|
||||
329: 1, # Star Trek
|
||||
260: 1, # Star Wars
|
||||
2288: -1, # The Thing
|
||||
1: 1, # Toy Story
|
||||
1721: -1, # Titanic
|
||||
296: -1, # Pulp Fiction
|
||||
356: 1, # Forrest Gump
|
||||
2116: 1, # Lord of the Rings
|
||||
1291: -1, # Indiana Jones
|
||||
1036: -1 # Die Hard
|
||||
}
|
||||
|
||||
inverse_ratings = {k: -v for k, v in my_ratings.items()}
|
||||
|
||||
def to_vector(ratings):
|
||||
vector = models.SparseVector(
|
||||
values=[],
|
||||
indices=[]
|
||||
)
|
||||
for movie_id, rating in ratings.items():
|
||||
vector.values.append(rating)
|
||||
vector.indices.append(movie_id)
|
||||
return vector
|
||||
```
|
||||
|
||||
Query Qdrant to find users with similar tastes based on the provided personal ratings. The search returns a list of similar users along with their ratings, facilitating collaborative filtering.
|
||||
|
||||
```python
|
||||
results = client.search(
|
||||
"movielens",
|
||||
query_vector=models.NamedSparseVector(
|
||||
name="ratings",
|
||||
vector=to_vector(my_ratings)
|
||||
),
|
||||
with_vectors=True, # We will use those to find new movies
|
||||
limit=20
|
||||
)
|
||||
```
|
||||
|
||||
Movie scores are computed based on how frequently each movie appears in the ratings of similar users, weighted by their ratings. This step identifies popular movies among users with similar tastes. Calculate how frequently each movie is found in similar users' ratings
|
||||
|
||||
```python
|
||||
def results_to_scores(results):
|
||||
movie_scores = defaultdict(lambda: 0)
|
||||
|
||||
for user in results:
|
||||
user_scores = user.vector['ratings']
|
||||
for idx, rating in zip(user_scores.indices, user_scores.values):
|
||||
if idx in my_ratings:
|
||||
continue
|
||||
movie_scores[idx] += rating
|
||||
|
||||
return movie_scores
|
||||
```
|
||||
|
||||
The top-rated movies are sorted based on their scores and printed as recommendations for the user. These recommendations are tailored to the user's preferences and aligned with their tastes. Sort movies by score and print top five:
|
||||
|
||||
```python
|
||||
movie_scores = results_to_scores(results)
|
||||
top_movies = sorted(movie_scores.items(), key=lambda x: x[1], reverse=True)
|
||||
|
||||
for movie_id, score in top_movies[:5]:
|
||||
print(movies[movies.movie_id == movie_id].title.values[0], score)
|
||||
```
|
||||
|
||||
## Result
|
||||
|
||||
```bash
|
||||
Star Wars: Episode V - The Empire Strikes Back (1980) 20.02387858
|
||||
Star Wars: Episode VI - Return of the Jedi (1983) 16.443184379999998
|
||||
Princess Bride, The (1987) 15.840068229999996
|
||||
Raiders of the Lost Ark (1981) 14.94489462
|
||||
Sixth Sense, The (1999) 14.570322149999999
|
||||
```
|
||||
Reference in New Issue
Block a user