diff --git a/qdrant-landing/content/documentation/tutorials/_index.md b/qdrant-landing/content/documentation/tutorials/_index.md index 27b5ed2ef..085954cc5 100644 --- a/qdrant-landing/content/documentation/tutorials/_index.md +++ b/qdrant-landing/content/documentation/tutorials/_index.md @@ -28,4 +28,5 @@ These tutorials demonstrate different ways you can build vector search into your | [Measure retrieval quality](../tutorials/retrieval-quality/) | Measure and fine-tune the retrieval quality | Qdrant, Python, datasets | | [Use semantic search to navigate your codebase](../tutorials/code-search/) | Implement semantic search application for code search task | Qdrant, Python, sentence-transformers, Jina | | [Implement custom connector for Cohere RAG](../tutorials/cohere-rag-connector/) | Bring data stored in Qdrant to Cohere RAG | Qdrant, Cohere, FastAPI | +| [](../tutorials/contract-management-stackit-aleph-alpha/) | | Qdrant, Aleph Alpha, Stackit | | [Troubleshooting](../tutorials/common-errors/) | Solutions to common errors and fixes | Qdrant | diff --git a/qdrant-landing/content/documentation/tutorials/contract-management-stackit-aleph-alpha.md b/qdrant-landing/content/documentation/tutorials/contract-management-stackit-aleph-alpha.md new file mode 100644 index 000000000..9f45bdce9 --- /dev/null +++ b/qdrant-landing/content/documentation/tutorials/contract-management-stackit-aleph-alpha.md @@ -0,0 +1,324 @@ +--- +title: Ease contract management in your organization +weight: 28 +--- + +# Ease contract management in your organization + +| Time: 90 min | Level: Advanced | | +| --- | ----------- | ----------- |----------- | + +The flow of the documents in the organization might be chaotic if you don't use a proper tool to manage them. Knowing +all the contracts and agreements by heart is impossible, so having an established system to store and search over them +would be best. However, simple retrieval is not enough in times of Large Language Models. An ability to ask questions +about the documents and get the answers is a game-changer for any business. On the other hand, you want granular access +control to the documents so that only the authorized team members can access them. + +Storing sensitive information requires fulfilling the security requirements restricted by your organization's policies +and local law regulations. If you operate in Europe, you might need to comply with GDPR, which requires storing personal +data securely. The physical location of the data is essential, so choosing the right cloud provider is crucial. The same +applies to the other components of the system, including the Large Language Model. You don't want to send your data +anywhere. You want it to be kept and processed within specific geographical boundaries. + +In this tutorial, we will guide you through the process of setting up a contract management system using [Aleph +Alpha](https://aleph-alpha.com/) embeddings and LLM. We will use [Stackit](https://www.stackit.de/), the German business +cloud, to run Qdrant Hybrid Cloud and the all the application processes. This setup will ensure that your data is stored +and processed in Germany, and nowhere else. + +[//]: # (TODO: add link to Qdrant Hybrid Cloud above) + +TODO: add a system diagram with all the components + +## Target system scope + +A contract management platform is not a simple CLI tool, but an application that should be available to all the team +members. It should have an interface to upload, search, and manage the documents. Perfectly, the system should be +integrated with an existing organization's stack, and the permissions and access control inherited from LDAP or Active +Directory. We are going to build a solid foundation for such a system in this tutorial, but leave the integration +details to the reader, as that really depends on the organization's specifics. + +There is a couple of moving parts in the system: + +[//]: # (TODO: upload the dataset to GCP and link it below) + +- **Dataset** - a collection of documents, using different formats, such as PDF or DOCx +- **Asymmetric semantic embeddings** - [Aleph Alpha embedding](https://docs.aleph-alpha.com/api/semantic-embed/) to + convert the queries and the documents into vectors +- **Large Language Model** - the [Luminous-extended-control + model](https://docs.aleph-alpha.com/docs/introduction/model-card/), but you can play with a different one from the + Luminous family +- **Qdrant Hybrid Cloud** - a knowledge base to store the vectors and search over the documents +- **Stackit** - a [German business cloud](https://www.stackit.de) to run the Qdrant Hybrid Cloud and the application + processes + +We will implement the process of uploading the documents, converting them into vectors, and storing them in Qdrant. +Then, we will build a search interface to query the documents and get the answers. All that, assuming the user +interacts with the system with some set of permissions, and can only access the documents they are allowed to. + +## Prerequisites + +There are a couple of things you need to have in place before starting the tutorial: + +### Aleph Alpha account + +Our tutorial relies on the Aleph Alpha models, so it is essential to have an account there. You can [sign +up](https://app.aleph-alpha.com/signup) at their playground and start with generating the API token in the [User +Profile](https://app.aleph-alpha.com/profile). Once you have it generated, let's store it in the environment variable: + +```shell +export ALEPH_ALPHA_API_KEY="" +``` + +```python +import os + +os.environ["ALEPH_ALPHA_API_KEY"] = "" +``` + +### Qdrant Hybrid Cloud on Stackit + +Please refer to our documentation to see how to deploy Qdrant Hybrid Cloud on Stackit. Once you finish the deployment, +you will have the API endpoint to interact with the Qdrant server. Let's store it in the environment variable as well: + +[//]: # (TODO: refer to the documentation on how to deploy Qdrant on Stackit) + +```shell +export QDRANT_URL="https://qdrant.example.com" +export QDRANT_API_KEY="your-api-key" +``` + +```python +os.environ["QDRANT_URL"] = "https://qdrant.example.com" +os.environ["QDRANT_API_KEY"] = "your-api-key" +``` + +## Building the application + +We could build our application with just the official SDKs of Aleph Alpha and Qdrant, but let's do not reinvent the +wheel and use [Langchain](https://python.langchain.com/docs/get_started/introduction) to streamline the process. +Langchain is already integrated with both services, so we can focus on the business logic. Moreover, it also + +### Setting up Qdrant collection + +Aleph Alpha embeddings are high dimensional vectors by default, with a dimensionality of 5120. Qdrant can store such +vector easily, but that also sounds like a good idea to enable [Binary +Quantization](../../../documentation/guides/quantization/#binary-quantization) to save space and make the retrieval +faster. Let's create a collection with such settings: + +```python +from qdrant_client import QdrantClient, models + +client = QdrantClient( + location=os.environ["QDRANT_URL"], + api_key=os.environ["QDRANT_API_KEY"], +) +client.create_collection( + collection_name="contracts", + vectors_config=models.VectorParams( + size=5120, + distance=models.Distance.COSINE, + quantization_config=models.BinaryQuantization( + binary=models.BinaryQuantizationConfig( + always_ram=True, + ) + ) + ), +) +``` + +We are going to use the `contracts` collection to store the vectors of the documents. The `always_ram` flag is set to +`True` to keep the quantized vectors in RAM, which will speed up the search process. We also wanted to restrict access +to the individual documents, so only users with the proper permissions can see them. In Qdrant that should be solved by +adding a payload field that defines who can access the document. We'll call this field `roles` and set it to an array +of strings with the roles that can access the document. + +```python +client.create_payload_index( + collection_name="contracts", + field_name="metadata.roles", + field_schema=models.PayloadSchemaType.KEYWORD, +) +``` + +Since we use Langchain, the `roles` field is a nested field of the `metadata`, so we had to define it as +`metadata.roles`. The schema says that the field is a keyword, which means it is a string or an array of strings. We are +going to use the name of the customers as the roles, so the access control will be based on the customer name. + +### Ingestion pipeline + +Every single semantic search system starts with good quality data. In our case, the data is a set of documents in +different formats. Thanks to the [Unstructured integration of +Langchain](https://python.langchain.com/docs/integrations/providers/unstructured), we can easily ingest PDFs, Microsoft +Word files or even PowerPoint presentations, so things are fairly simple. We have to take care of the text splitting, so +we do not try to convert a single document into a vector, but rather divide it into meaningful chunks. Finally, the +extracted documents have to be converted into vectors using the Aleph Alpha embeddings and stored in the Qdrant +collection. + +Let's start with defining the components and connecting them together: + +```python +embeddings = AlephAlphaAsymmetricSemanticEmbedding( + model="luminous-base", + aleph_alpha_api_key=os.environ["ALEPH_ALPHA_API_KEY"], + normalize=True, +) + +qdrant = Qdrant( + client=client, + collection_name="contracts", + embeddings=embeddings, +) +``` + +Now it's high time to index our documents. Each of the documents is a separate file, and we also have to know the +customer name to set the access control properly. There might be several roles for a single document, so let's keep them +in a list. + +```python +documents = { + "data/Data-Processing-Agreement_STACKIT_Cloud_version-1.2.pdf": ["stackit"], + "data/langchain-terms-of-service.pdf": ["langchain"], +} +``` + +Each has to be split into chunks first; there is no silver bullet. Our chunking algorithm will be simple and based on +recursive splitting, with the maximum chunk size of 500 characters and the overlap of 100 characters. + +```python +from langchain_text_splitters import RecursiveCharacterTextSplitter + +text_splitter = RecursiveCharacterTextSplitter( + chunk_size=500, + chunk_overlap=100, +) +``` + +Now we can iterate over the documents, split them into chunks, convert them into vectors with Aleph Alpha embedding +model, and store them in the Qdrant. + +```python +from langchain_community.document_loaders.unstructured import UnstructuredFileLoader + +for document_path, roles in documents.items(): + document_loader = UnstructuredFileLoader(file_path=document_path) + + # Unstructured loads each file into a single Document object + loaded_documents = document_loader.load() + for doc in loaded_documents: + doc.metadata["roles"] = roles + + # Chunks will have the same metadata as the original document + document_chunks = text_splitter.split_documents(loaded_documents) + + # Add the documents to the Qdrant collection + qdrant.add_documents(document_chunks, batch_size=20) +``` + +Our collection is filled with data, and we can start searching over it. In a real-world scenario, the ingestion process +should be automated and triggered by the new documents uploaded to the system. Since we already use Qdrant Hybrid Cloud +running on Kubernetes, we can easily deploy the ingestion pipeline as a job to the same environment. On Stackit, you +probably use the [STACKIT Kubernetes Engine (SKE)](https://www.stackit.de/en/product/kubernetes/) and launch it in a +container. The [Compute Engine](https://www.stackit.de/en/product/stackit-compute-engine/) is also an option, but +everything depends on the specifics of your organization. + +### Search application + +Specialized Document Management Systems have a lot of features, but semantic search is not yet a standard. We are going +to build a simple search mechanism which could be possibly integrated with the existing system. The search process is +quite simple: we convert the query into a vector using the same Aleph Alpha model, and then search for the most similar +documents in the Qdrant collection. The access control is also applied, so the user can only see the documents they are +allowed to. + +We start with creating an instance of the LLM of our choice, and set the maximum number of tokens to 200, as the default +value is 64, which might be too low for our purposes. + +```python +from langchain.llms.aleph_alpha import AlephAlpha + +llm = AlephAlpha( + model="luminous-extended-control", + aleph_alpha_api_key=os.environ["ALEPH_ALPHA_API_KEY"], + maximum_tokens=200, +) +``` + +Then, we can glue the components together and build the search process. `RetrievalQA` is a class that takes implements +the Question Retrieval process, with a specified retriever and Large Language Model. The instance of `Qdrant` might be +converted into a retriever, with additional filter that will be passed to the `similarity_search` method. The filter +is created as [in a regular Qdrant query](../../../documentation/concepts/filtering/), with the `roles` field set to the +user's roles. + +```python +user_roles = ["stackit", "aleph-alpha"] + +qdrant_retriever = qdrant.as_retriever( + search_kwargs={ + "filter": models.Filter( + must=[ + models.FieldCondition( + key="metadata.roles", + match=models.MatchAny(any=user_roles) + ) + ] + ) + } +) +``` + +We set the user roles to `stackit` and `aleph-alpha`, so the user can see the documents that are accessible to these +customers, but not to the others. The final step is to create the `RetrievalQA` instance and use it to search over the +documents, with the custom prompt. + +```python +from langchain.prompts import PromptTemplate +from langchain.chains.retrieval_qa.base import RetrievalQA + +prompt_template = """ +### Instruction: +{question} If there's no answer, say "Provided context does not clarify it". +### Input: +Text:{context} +Question:{question} +### Response: +""" +prompt = PromptTemplate( + template=prompt_template, input_variables=["context", "question"] +) + +retrieval_qa = RetrievalQA.from_chain_type( + llm=llm, + chain_type="stuff", + retriever=qdrant_retriever, + return_source_documents=True, + chain_type_kwargs={"prompt": prompt}, +) + +response = retrieval_qa.invoke({"query": "What is the purpose of the contract?"}) +print(response["result"]) +``` + +Output: + +```text +The rules for performing the audit are as follows: + +1. The Customer must inform the Contractor in good time (usually at least two weeks in advance) about any and all circumstances related to the performance of the audit. +2. The Customer is entitled to perform one audit per calendar year. Any additional audits may be performed if agreed with the Contractor and are subject to reimbursement of expenses. +3. If the Customer engages a third party to perform the audit, the Customer must obtain the Contractor's consent and ensure that the confidentiality agreements with the third party are observed. +4. The Contractor may object to any third party deemed unsuitable. +``` + +There are some other parameters that might be tuned to optimize the search process. The `k` parameter defines how many +documents should be returned, but Langchain allows us also to control the retrieval process by choosing the type of the +search operation. The default is `similarity`, which is just vector search, but we can also use `mmr` which stands for +Maximal Marginal Relevance. It is a technique to diversify the search results, so the user gets the most relevant +documents, but also the most diverse ones. The `mmr` search is slower, but might be more user-friendly. + +Our search application is ready, and we can deploy it to the same environment as the ingestion pipeline on Stackit. The +same rules apply here, so you can use the SKE or the Compute Engine, depending on the specifics of your organization. + +## Next steps + +We built a solid foundation for the contract management system, but there is still a lot to do. If you want to make the +system production-ready, you should consider implementing the mechanism into your existing stack. If you have any +questions, feel free to ask on our [Discord community](https://qdrant.to/discord). \ No newline at end of file diff --git a/qdrant-landing/static/documentation/tutorials/contract-management-stackit-aleph-alpha-social-preview.png b/qdrant-landing/static/documentation/tutorials/contract-management-stackit-aleph-alpha-social-preview.png new file mode 100644 index 000000000..c898ca87a Binary files /dev/null and b/qdrant-landing/static/documentation/tutorials/contract-management-stackit-aleph-alpha-social-preview.png differ