Add first draft of the Stackit & Aleph Alpha tutorial

This commit is contained in:
Kacper Łukawski
2024-04-03 22:02:35 +02:00
parent c137af4dc6
commit bf3c391a6b
3 changed files with 325 additions and 0 deletions
@@ -28,4 +28,5 @@ These tutorials demonstrate different ways you can build vector search into your
| [Measure retrieval quality](../tutorials/retrieval-quality/) | Measure and fine-tune the retrieval quality | Qdrant, Python, datasets |
| [Use semantic search to navigate your codebase](../tutorials/code-search/) | Implement semantic search application for code search task | Qdrant, Python, sentence-transformers, Jina |
| [Implement custom connector for Cohere RAG](../tutorials/cohere-rag-connector/) | Bring data stored in Qdrant to Cohere RAG | Qdrant, Cohere, FastAPI |
| [](../tutorials/contract-management-stackit-aleph-alpha/) | | Qdrant, Aleph Alpha, Stackit |
| [Troubleshooting](../tutorials/common-errors/) | Solutions to common errors and fixes | Qdrant |
@@ -0,0 +1,324 @@
---
title: Ease contract management in your organization
weight: 28
---
# Ease contract management in your organization
| Time: 90 min | Level: Advanced | |
| --- | ----------- | ----------- |----------- |
The flow of the documents in the organization might be chaotic if you don't use a proper tool to manage them. Knowing
all the contracts and agreements by heart is impossible, so having an established system to store and search over them
would be best. However, simple retrieval is not enough in times of Large Language Models. An ability to ask questions
about the documents and get the answers is a game-changer for any business. On the other hand, you want granular access
control to the documents so that only the authorized team members can access them.
Storing sensitive information requires fulfilling the security requirements restricted by your organization's policies
and local law regulations. If you operate in Europe, you might need to comply with GDPR, which requires storing personal
data securely. The physical location of the data is essential, so choosing the right cloud provider is crucial. The same
applies to the other components of the system, including the Large Language Model. You don't want to send your data
anywhere. You want it to be kept and processed within specific geographical boundaries.
In this tutorial, we will guide you through the process of setting up a contract management system using [Aleph
Alpha](https://aleph-alpha.com/) embeddings and LLM. We will use [Stackit](https://www.stackit.de/), the German business
cloud, to run Qdrant Hybrid Cloud and the all the application processes. This setup will ensure that your data is stored
and processed in Germany, and nowhere else.
[//]: # (TODO: add link to Qdrant Hybrid Cloud above)
TODO: add a system diagram with all the components
## Target system scope
A contract management platform is not a simple CLI tool, but an application that should be available to all the team
members. It should have an interface to upload, search, and manage the documents. Perfectly, the system should be
integrated with an existing organization's stack, and the permissions and access control inherited from LDAP or Active
Directory. We are going to build a solid foundation for such a system in this tutorial, but leave the integration
details to the reader, as that really depends on the organization's specifics.
There is a couple of moving parts in the system:
[//]: # (TODO: upload the dataset to GCP and link it below)
- **Dataset** - a collection of documents, using different formats, such as PDF or DOCx
- **Asymmetric semantic embeddings** - [Aleph Alpha embedding](https://docs.aleph-alpha.com/api/semantic-embed/) to
convert the queries and the documents into vectors
- **Large Language Model** - the [Luminous-extended-control
model](https://docs.aleph-alpha.com/docs/introduction/model-card/), but you can play with a different one from the
Luminous family
- **Qdrant Hybrid Cloud** - a knowledge base to store the vectors and search over the documents
- **Stackit** - a [German business cloud](https://www.stackit.de) to run the Qdrant Hybrid Cloud and the application
processes
We will implement the process of uploading the documents, converting them into vectors, and storing them in Qdrant.
Then, we will build a search interface to query the documents and get the answers. All that, assuming the user
interacts with the system with some set of permissions, and can only access the documents they are allowed to.
## Prerequisites
There are a couple of things you need to have in place before starting the tutorial:
### Aleph Alpha account
Our tutorial relies on the Aleph Alpha models, so it is essential to have an account there. You can [sign
up](https://app.aleph-alpha.com/signup) at their playground and start with generating the API token in the [User
Profile](https://app.aleph-alpha.com/profile). Once you have it generated, let's store it in the environment variable:
```shell
export ALEPH_ALPHA_API_KEY="<your-token>"
```
```python
import os
os.environ["ALEPH_ALPHA_API_KEY"] = "<your-token>"
```
### Qdrant Hybrid Cloud on Stackit
Please refer to our documentation to see how to deploy Qdrant Hybrid Cloud on Stackit. Once you finish the deployment,
you will have the API endpoint to interact with the Qdrant server. Let's store it in the environment variable as well:
[//]: # (TODO: refer to the documentation on how to deploy Qdrant on Stackit)
```shell
export QDRANT_URL="https://qdrant.example.com"
export QDRANT_API_KEY="your-api-key"
```
```python
os.environ["QDRANT_URL"] = "https://qdrant.example.com"
os.environ["QDRANT_API_KEY"] = "your-api-key"
```
## Building the application
We could build our application with just the official SDKs of Aleph Alpha and Qdrant, but let's do not reinvent the
wheel and use [Langchain](https://python.langchain.com/docs/get_started/introduction) to streamline the process.
Langchain is already integrated with both services, so we can focus on the business logic. Moreover, it also
### Setting up Qdrant collection
Aleph Alpha embeddings are high dimensional vectors by default, with a dimensionality of 5120. Qdrant can store such
vector easily, but that also sounds like a good idea to enable [Binary
Quantization](../../../documentation/guides/quantization/#binary-quantization) to save space and make the retrieval
faster. Let's create a collection with such settings:
```python
from qdrant_client import QdrantClient, models
client = QdrantClient(
location=os.environ["QDRANT_URL"],
api_key=os.environ["QDRANT_API_KEY"],
)
client.create_collection(
collection_name="contracts",
vectors_config=models.VectorParams(
size=5120,
distance=models.Distance.COSINE,
quantization_config=models.BinaryQuantization(
binary=models.BinaryQuantizationConfig(
always_ram=True,
)
)
),
)
```
We are going to use the `contracts` collection to store the vectors of the documents. The `always_ram` flag is set to
`True` to keep the quantized vectors in RAM, which will speed up the search process. We also wanted to restrict access
to the individual documents, so only users with the proper permissions can see them. In Qdrant that should be solved by
adding a payload field that defines who can access the document. We'll call this field `roles` and set it to an array
of strings with the roles that can access the document.
```python
client.create_payload_index(
collection_name="contracts",
field_name="metadata.roles",
field_schema=models.PayloadSchemaType.KEYWORD,
)
```
Since we use Langchain, the `roles` field is a nested field of the `metadata`, so we had to define it as
`metadata.roles`. The schema says that the field is a keyword, which means it is a string or an array of strings. We are
going to use the name of the customers as the roles, so the access control will be based on the customer name.
### Ingestion pipeline
Every single semantic search system starts with good quality data. In our case, the data is a set of documents in
different formats. Thanks to the [Unstructured integration of
Langchain](https://python.langchain.com/docs/integrations/providers/unstructured), we can easily ingest PDFs, Microsoft
Word files or even PowerPoint presentations, so things are fairly simple. We have to take care of the text splitting, so
we do not try to convert a single document into a vector, but rather divide it into meaningful chunks. Finally, the
extracted documents have to be converted into vectors using the Aleph Alpha embeddings and stored in the Qdrant
collection.
Let's start with defining the components and connecting them together:
```python
embeddings = AlephAlphaAsymmetricSemanticEmbedding(
model="luminous-base",
aleph_alpha_api_key=os.environ["ALEPH_ALPHA_API_KEY"],
normalize=True,
)
qdrant = Qdrant(
client=client,
collection_name="contracts",
embeddings=embeddings,
)
```
Now it's high time to index our documents. Each of the documents is a separate file, and we also have to know the
customer name to set the access control properly. There might be several roles for a single document, so let's keep them
in a list.
```python
documents = {
"data/Data-Processing-Agreement_STACKIT_Cloud_version-1.2.pdf": ["stackit"],
"data/langchain-terms-of-service.pdf": ["langchain"],
}
```
Each has to be split into chunks first; there is no silver bullet. Our chunking algorithm will be simple and based on
recursive splitting, with the maximum chunk size of 500 characters and the overlap of 100 characters.
```python
from langchain_text_splitters import RecursiveCharacterTextSplitter
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=500,
chunk_overlap=100,
)
```
Now we can iterate over the documents, split them into chunks, convert them into vectors with Aleph Alpha embedding
model, and store them in the Qdrant.
```python
from langchain_community.document_loaders.unstructured import UnstructuredFileLoader
for document_path, roles in documents.items():
document_loader = UnstructuredFileLoader(file_path=document_path)
# Unstructured loads each file into a single Document object
loaded_documents = document_loader.load()
for doc in loaded_documents:
doc.metadata["roles"] = roles
# Chunks will have the same metadata as the original document
document_chunks = text_splitter.split_documents(loaded_documents)
# Add the documents to the Qdrant collection
qdrant.add_documents(document_chunks, batch_size=20)
```
Our collection is filled with data, and we can start searching over it. In a real-world scenario, the ingestion process
should be automated and triggered by the new documents uploaded to the system. Since we already use Qdrant Hybrid Cloud
running on Kubernetes, we can easily deploy the ingestion pipeline as a job to the same environment. On Stackit, you
probably use the [STACKIT Kubernetes Engine (SKE)](https://www.stackit.de/en/product/kubernetes/) and launch it in a
container. The [Compute Engine](https://www.stackit.de/en/product/stackit-compute-engine/) is also an option, but
everything depends on the specifics of your organization.
### Search application
Specialized Document Management Systems have a lot of features, but semantic search is not yet a standard. We are going
to build a simple search mechanism which could be possibly integrated with the existing system. The search process is
quite simple: we convert the query into a vector using the same Aleph Alpha model, and then search for the most similar
documents in the Qdrant collection. The access control is also applied, so the user can only see the documents they are
allowed to.
We start with creating an instance of the LLM of our choice, and set the maximum number of tokens to 200, as the default
value is 64, which might be too low for our purposes.
```python
from langchain.llms.aleph_alpha import AlephAlpha
llm = AlephAlpha(
model="luminous-extended-control",
aleph_alpha_api_key=os.environ["ALEPH_ALPHA_API_KEY"],
maximum_tokens=200,
)
```
Then, we can glue the components together and build the search process. `RetrievalQA` is a class that takes implements
the Question Retrieval process, with a specified retriever and Large Language Model. The instance of `Qdrant` might be
converted into a retriever, with additional filter that will be passed to the `similarity_search` method. The filter
is created as [in a regular Qdrant query](../../../documentation/concepts/filtering/), with the `roles` field set to the
user's roles.
```python
user_roles = ["stackit", "aleph-alpha"]
qdrant_retriever = qdrant.as_retriever(
search_kwargs={
"filter": models.Filter(
must=[
models.FieldCondition(
key="metadata.roles",
match=models.MatchAny(any=user_roles)
)
]
)
}
)
```
We set the user roles to `stackit` and `aleph-alpha`, so the user can see the documents that are accessible to these
customers, but not to the others. The final step is to create the `RetrievalQA` instance and use it to search over the
documents, with the custom prompt.
```python
from langchain.prompts import PromptTemplate
from langchain.chains.retrieval_qa.base import RetrievalQA
prompt_template = """
### Instruction:
{question} If there's no answer, say "Provided context does not clarify it".
### Input:
Text:{context}
Question:{question}
### Response:
"""
prompt = PromptTemplate(
template=prompt_template, input_variables=["context", "question"]
)
retrieval_qa = RetrievalQA.from_chain_type(
llm=llm,
chain_type="stuff",
retriever=qdrant_retriever,
return_source_documents=True,
chain_type_kwargs={"prompt": prompt},
)
response = retrieval_qa.invoke({"query": "What is the purpose of the contract?"})
print(response["result"])
```
Output:
```text
The rules for performing the audit are as follows:
1. The Customer must inform the Contractor in good time (usually at least two weeks in advance) about any and all circumstances related to the performance of the audit.
2. The Customer is entitled to perform one audit per calendar year. Any additional audits may be performed if agreed with the Contractor and are subject to reimbursement of expenses.
3. If the Customer engages a third party to perform the audit, the Customer must obtain the Contractor's consent and ensure that the confidentiality agreements with the third party are observed.
4. The Contractor may object to any third party deemed unsuitable.
```
There are some other parameters that might be tuned to optimize the search process. The `k` parameter defines how many
documents should be returned, but Langchain allows us also to control the retrieval process by choosing the type of the
search operation. The default is `similarity`, which is just vector search, but we can also use `mmr` which stands for
Maximal Marginal Relevance. It is a technique to diversify the search results, so the user gets the most relevant
documents, but also the most diverse ones. The `mmr` search is slower, but might be more user-friendly.
Our search application is ready, and we can deploy it to the same environment as the ingestion pipeline on Stackit. The
same rules apply here, so you can use the SKE or the Compute Engine, depending on the specifics of your organization.
## Next steps
We built a solid foundation for the contract management system, but there is still a lot to do. If you want to make the
system production-ready, you should consider implementing the mechanism into your existing stack. If you have any
questions, feel free to ask on our [Discord community](https://qdrant.to/discord).