Merge branch 'hybrid-cloud-dev' into hybrid-cloud/tutorial/customer-support-oci-cohere-airbyte
# Conflicts: # qdrant-landing/content/documentation/tutorials/_index.md
@@ -0,0 +1,57 @@
|
||||
---
|
||||
draft: false
|
||||
title: "Qdrant is Now Available on Azure Marketplace!"
|
||||
short_description: Discover the power of Qdrant on Azure Marketplace!
|
||||
description: Discover the power of Qdrant on Azure Marketplace! Get started today and streamline your operations with ease.
|
||||
preview_image: /blog/azure-marketplace/azure-marketplace.png
|
||||
date: 2024-03-26T10:30:00Z
|
||||
author: David Myriel
|
||||
featured: false
|
||||
weight: 0
|
||||
tags:
|
||||
- Qdrant
|
||||
- Azure Marketplace
|
||||
- Enterprise
|
||||
- Vector Database
|
||||
---
|
||||
|
||||
We're thrilled to announce that Qdrant is now [officially available on Azure Marketplace](https://azuremarketplace.microsoft.com/en-en/marketplace/apps/qdrantsolutionsgmbh1698769709989.qdrant-db), bringing enterprise-level vector search directly to Azure's vast community of users. This integration marks a significant milestone in our journey to make Qdrant more accessible and convenient for businesses worldwide.
|
||||
|
||||
> *With the landscape of AI being complex for most customers, Qdrant's ease of use provides an easy approach for customers' implementation of RAG patterns for Generative AI solutions and additional choices in selecting AI components on Azure,* - Tara Walker, Principal Software Engineer at Microsoft.
|
||||
|
||||
## Why Azure Marketplace?
|
||||
|
||||
[Azure Marketplace](https://azuremarketplace.microsoft.com/en-us/) is renowned for its robust ecosystem, trusted by millions of users globally. By listing Qdrant on Azure Marketplace, we're not only expanding our reach but also ensuring seamless integration with Azure's suite of tools and services. This collaboration opens up new possibilities for our users, enabling them to leverage the power of Azure alongside the capabilities of Qdrant.
|
||||
|
||||
> *Enterprises like Bosch can now use the power of Microsoft Azure to host Qdrant, unleashing unparalleled performance and massive-scale vector search. "With Qdrant, we found the missing piece to develop our own provider independent multimodal generative AI platform at enterprise scale,* - Jeremy Teichmann (AI Squad Technical Lead & Generative AI Expert), Daly Singh (AI Squad Lead & Product Owner) - Bosch Digital.
|
||||
|
||||
## Key Benefits for Users:
|
||||
|
||||
- **Rapid Application Development:** Deploying a cluster on Microsoft Azure via the Qdrant Cloud console only takes a few seconds and can scale up as needed, giving developers maximal flexibility for their production deployments.
|
||||
|
||||
- **Billion Vector Scale:** Seamlessly grow and handle large-scale datasets with billions of vectors by leveraging Qdrant's features like vertical and horizontal scaling or binary quantization with Microsoft Azure's scalable infrastructure.
|
||||
|
||||
- **Unparalleled Performance:** Qdrant is built to handle scaling challenges, high throughput, low latency, and efficient indexing. Written in Rust makes Qdrant fast and reliable even under high load. See benchmarks.
|
||||
|
||||
- **Versatile Applications:** From recommendation systems to similarity search, Qdrant's integration with Microsoft Azure provides a versatile tool for a diverse set of AI applications.
|
||||
|
||||
## Getting Started:
|
||||
|
||||
Ready to experience the benefits of Qdrant on Azure Marketplace? Getting started is easy:
|
||||
|
||||
1. **Visit the Azure Marketplace**: Navigate to [Qdrant's Marketplace listing](https://azuremarketplace.microsoft.com/en-en/marketplace/apps/qdrantsolutionsgmbh1698769709989.qdrant-db).
|
||||
2. **Deploy Qdrant**: Follow the simple deployment instructions to set up your instance.
|
||||
3. **Start Using Qdrant**: Once deployed, start exploring the [features and capabilities of Qdrant](https://qdrant.tech/documentation/concepts/) on Azure.
|
||||
4. **Read Documentation**: Read Qdrant's [Documentation](https://qdrant.tech/documentation/) and build demo apps using [Tutorials](https://qdrant.tech/documentation/tutorials/).
|
||||
|
||||
## Join Us on this Exciting Journey:
|
||||
|
||||
We're incredibly excited about this collaboration with Azure Marketplace and the opportunities it brings for our users. As we continue to innovate and enhance Qdrant, we invite you to join us on this journey towards greater efficiency, scalability, and success.
|
||||
|
||||
Ready to elevate your business with Qdrant? **Click the banner and get started today!**
|
||||
|
||||
[](https://azuremarketplace.microsoft.com/en-en/marketplace/apps/qdrantsolutionsgmbh1698769709989.qdrant-db)
|
||||
|
||||
### About Qdrant:
|
||||
|
||||
Qdrant is the leading, high-performance, scalable, open-source vector database and search engine, essential for building the next generation of AI/ML applications. Qdrant is able to handle billions of vectors, supports the matching of semantically complex objects, and is implemented in Rust for performance, memory safety, and scale.
|
||||
@@ -1,16 +1,16 @@
|
||||
---
|
||||
draft: true
|
||||
draft: false
|
||||
title: Insight Generation Platform for LifeScience Corporation - Hooman
|
||||
Sedghamiz | Vector Space Talks
|
||||
slug: insight-generation-platform
|
||||
short_description: Hooman Sedghamiz explores the untapped potential of large
|
||||
language models in creating cutting-edge applications.
|
||||
description: Hooman Sedghamiz unfolds the potential of AI in life sciences, from
|
||||
custom knowledge applications to improving crop yield predictions, while
|
||||
teasing apart the nuances of in-house AI deployment for multi-faceted
|
||||
short_description: Hooman Sedghamiz explores the potential of large language
|
||||
models in creating cutting-edge AI applications.
|
||||
description: Hooman Sedghamiz discloses the potential of AI in life sciences,
|
||||
from custom knowledge applications to improving crop yield predictions, while
|
||||
tearing apart the nuances of in-house AI deployment for multi-faceted
|
||||
enterprise efficiency.
|
||||
preview_image: /blog/from_cms/hooman-sedghamiz-bp-cropped.png
|
||||
date: 2024-03-08T09:45:59.753Z
|
||||
date: 2024-03-25T08:46:28.227Z
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
@@ -18,11 +18,13 @@ tags:
|
||||
- Retrieval Augmented Generation
|
||||
- Insight Generation Platform
|
||||
---
|
||||
> *"So there is this really great vector db comparison that came out recently. I saw there are like maybe more than 40 vector stores in 2024. When we started back in 2023 was only a few. And what I see, which is really lacking in this pipeline of retrieval augmented generation is major innovation around data pipeline.”*\
|
||||
> *"There is this really great vector db comparison that came out recently. I saw there are like maybe more than 40 vector stores in 2024. When we started back in 2023, there were only a few. What I see, which is really lacking in this pipeline of retrieval augmented generation is major innovation around data pipeline.”*\
|
||||
-- Hooman Sedghamiz
|
||||
>
|
||||
|
||||
Hooman Sedghamiz**,** Sr. Director AI/ML - Insights at Bayer AG is a distinguished figure in AI and ML in the life sciences field. With a wealth of experience, he has led teams and projects that have greatly advanced medical products, including implantable and wearable devices. Notably, he served as the Generative AI product owner and senior director at Bayer Pharmaceuticals, where he played a pivotal role in developing a GPT-based central platform for precision medicine. In 2023, he assumed the role of Co-Chair for the EMNLP 2023 GEM industrial track, furthering his contributions to the field. Hooman has also been an AI/ML advisor and scientist at the University of California, San Diego, leveraging his expertise in deep learning to drive biomedical research and innovation. His strengths lie in guiding data science initiatives from inception to commercialization and bridging the gap between medical and healthcare applications through MLOps, LlmOps, and deep learning product management. Engaging with research institutions and collaborating closely with Dr. Nemati at Harvard University and UCSD, Hooman continues to be a dynamic and influential figure in the data science community.
|
||||
Hooman Sedghamiz, Sr. Director AI/ML - Insights at Bayer AG is a distinguished figure in AI and ML in the life sciences field. With years of experience, he has led teams and projects that have greatly advanced medical products, including implantable and wearable devices. Notably, he served as the Generative AI product owner and Senior Director at Bayer Pharmaceuticals, where he played a pivotal role in developing a GPT-based central platform for precision medicine.
|
||||
|
||||
In 2023, he assumed the role of Co-Chair for the EMNLP 2023 GEM industrial track, furthering his contributions to the field. Hooman has also been an AI/ML advisor and scientist at the University of California, San Diego, leveraging his expertise in deep learning to drive biomedical research and innovation. His strengths lie in guiding data science initiatives from inception to commercialization and bridging the gap between medical and healthcare applications through MLOps, LLMOps, and deep learning product management. Engaging with research institutions and collaborating closely with Dr. Nemati at Harvard University and UCSD, Hooman continues to be a dynamic and influential figure in the data science community.
|
||||
|
||||
***Listen to the episode on [Spotify](https://open.spotify.com/episode/2oj2ne5l9qrURQSV0T1Hft?si=DMJRTAt7QXibWiQ9CEKTJw), Apple Podcast, Podcast addicts, Castbox. You can also watch this episode on [YouTube](https://youtu.be/yfzLaH5SFX0).***
|
||||
|
||||
@@ -32,9 +34,9 @@ Hooman Sedghamiz**,** Sr. Director AI/ML - Insights at Bayer AG is a distinguish
|
||||
|
||||
## **Top takeaways:**
|
||||
|
||||
Why is real-time evaluation critical in maintaining the integrity of chatbot interactions and preventing issues like promoting competitors or making false promises? What strategies do developers employ to minimize cost while maximizing the effectiveness of model evaluations, specifically when dealing with LLMs? These might be just some of the many questions people in the industry are asking themselves. Worry not! Because Demetrios and Sourabh will break it down for you.
|
||||
Why is real-time evaluation critical in maintaining the integrity of chatbot interactions and preventing issues like promoting competitors or making false promises? What strategies do developers employ to minimize cost while maximizing the effectiveness of model evaluations, specifically when dealing with LLMs? These might be just some of the many questions people in the industry are asking themselves. We aim to cover most of it in this talk.
|
||||
|
||||
Check out their conversation as they dive into the intricate world of AI chatbot evaluations. Discover the nuances of ensuring your chatbot's quality and continuous improvement across various metrics.
|
||||
Check out their conversation as they peek into world of AI chatbot evaluations. Discover the nuances of ensuring your chatbot's quality and continuous improvement across various metrics.
|
||||
|
||||
Here are the key topics of this episode:
|
||||
|
||||
@@ -44,7 +46,8 @@ Here are the key topics of this episode:
|
||||
4. **Cost-Effective Evaluation Models**: Discussion on employing smaller models for evaluation to reduce costs without compromising the depth of analysis, focusing on failure cases and root-cause assessments.
|
||||
5. **Tailored Evaluation Metrics**: Emphasis on the necessity of customizing evaluation criteria to suit specific use case requirements, including an exploration of the different metrics applicable to diverse scenarios.
|
||||
|
||||
> Fun Fact: Sourabh discussed the use of Uptrend, an innovative API that provides scores and explanations for various data checks, facilitating logical and informed decision-making when evaluating AI models.
|
||||
>Fun Fact: Large language models like Mistral, Llama, and Nexus Raven have improved in their ability to perform function calling with low hallucination and high-quality output.
|
||||
|
||||
>
|
||||
|
||||
## Show notes:
|
||||
@@ -73,13 +76,16 @@ Here are the key topics of this episode:
|
||||
|
||||
## Transcript:
|
||||
Demetrios:
|
||||
We are here and I couldn't think of a better way to spend my Valentine's Day than with you humen this is absolutely incredible. I'm so excited for this talk that you're going to bring and I want to let everyone that is out there listening know what caliber of of speaker we have with us today because you have done a lot of stuff. Folks out there do not let this man's young look fool you. You look like you are not in your. When it comes to your bio, it looks like you should be in your very excited. You've got a lot of experience running data science projects, ML projects, LLM projects, all that fun stuff. You're working at Bayern Munich, sorry, not Bayern Munich, Bayer AG. And you're the senior director of AI and ML.
|
||||
We are here and I couldn't think of a better way to spend my Valentine's Day than with you Hooman this is absolutely incredible. I'm so excited for this talk that you're going to bring and I want to let everyone that is out there listening know what caliber of a speaker we have with us today because you have done a lot of stuff. Folks out there do not let this man's young look fool you. You look like you are not in your fifty's or sixty's. But when it comes to your bio, it looks like you should be in your seventy's. I am very excited. You've got a lot of experience running data science projects, ML projects, LLM projects, all that fun stuff. You're working at Bayern Munich, sorry, not Bayern Munich, Bayer AG. And you're the senior director of AI and ML.
|
||||
|
||||
|
||||
Demetrios:
|
||||
And I think that there is a ton of other stuff that you've done when it comes to machine learning, artificial intelligence. You've got both like the traditional ML background, I think, and then you've also got this new generative AI background and so you can leverage both. But you also think about things in data engineering way. You understand the whole lifecycle. And so today we get to talk all about some of this fun. I know you've got some slides prepared for us. I'll let you throw those on and I'll let anyone else in the chat. Feel free to ask questions while humin is going through the presentation and I'll jump in and stop them when needed.
|
||||
And I think that there is a ton of other stuff that you've done when it comes to machine learning, artificial intelligence. You've got both like the traditional ML background, I think, and then you've also got this new generative AI background and so you can leverage both. But you also think about things in data engineering way. You understand the whole lifecycle. And so today we get to talk all about some of this fun. I know you've got some slides prepared for us. I'll let you throw those on and I'll let anyone else in the chat. Feel free to ask questions while Hooman is going through the presentation and I'll jump in and stop them when needed.
|
||||
|
||||
|
||||
Demetrios:
|
||||
But also we can have a little discussion after a few minutes of slides. So for everyone looking, we're going to be watching this and then we're going to be checking out like really talking about what 2024 AI in the enterprise looks like and what is needed to really take advantage of that. So human, I'm dropping off to you, man, and I'll jump in when needed.
|
||||
But also we can have a little discussion after a few minutes of slides. So for everyone looking, we're going to be watching this and then we're going to be checking out like really talking about what 2024 AI in the enterprise looks like and what is needed to really take advantage of that. So Hooman, I'm dropping off to you, man, and I'll jump in when needed.
|
||||
|
||||
|
||||
Hooman Sedghamiz:
|
||||
Thanks a lot for the introduction. Let me get started. Do you have my screen already?
|
||||
@@ -94,10 +100,10 @@ Hooman Sedghamiz:
|
||||
So now you can imagine via is really important to us because it has the potential of unlocking a future where good health is a reality and hunger is a memory. So I maybe start about maybe giving you a hint of what are really the numerous use cases that AI or challenges that AI could help out with. In life science industry. You can think of adverse event detection when patients are taking a medication, too much of it. The patients might report adverse events, stomach bleeding and go to social media post about it. A few years back, it was really difficult to process automatically all this sort of natural text in a kind of scalable manner. But nowadays, thanks to large language models, it's possible to automate this and identify if there is a medication or anything that might have negatively an adverse event on a patient population. Similarly, you can now create a lot of marketing content using these large language models for products.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
At the same time, drug discovery is making really big strides when it comes to identifying new compounds. You can essentially describe these compounds using formats like smiles, which could be represented as really text. And these large language models can be trained on them and they can predict the sequences. At the same time, you have this clinical trial outcome prediction, which is huge for pharmaceutical companies. If you could predict what will be the outcome of a trial, it would be a huge time and resource saving for a lot of companies. And of course, a lot of us already see in the market a lot of medical virtual assistants using large language models that can answer medical inquiries and give consultations around them. And there is really, I believe the biggest potential here is around real world data, like most of us nowadays, have some sort of sensor or watch that's measuring our health maybe at a minute by minute level, or it's measuring our heart rate. You go to the hospital, you have all your medical records recorded there, and these large language models have their capacity to process this complex data, and you will be able to drive better insights for individualized insights for patients.
|
||||
At the same time, drug discovery is making really big strides when it comes to identifying new compounds. You can essentially describe these compounds using formats like smiles, which could be represented as real text. And these large language models can be trained on them and they can predict the sequences. At the same time, you have this clinical trial outcome prediction, which is huge for pharmaceutical companies. If you could predict what will be the outcome of a trial, it would be a huge time and resource saving for a lot of companies. And of course, a lot of us already see in the market a lot of medical virtual assistants using large language models that can answer medical inquiries and give consultations around them. And there is really, I believe the biggest potential here is around real world data, like most of us nowadays, have some sort of sensor or watch that's measuring our health maybe at a minute by minute level, or it's measuring our heart rate. You go to the hospital, you have all your medical records recorded there, and these large language models have their capacity to process this complex data, and you will be able to drive better insights for individualized insights for patients.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
And our company is also in crop science, as I mentioned, and crop yield prediction. If you could help farmers improve their crop yield, it means that they can produce better products faster with higher quality. So maybe I could start with maybe a history in 2023, what happened? How companies like ours were looking at large language models and opportunities. They bring, I think in 2023, everyone was excited to bring these efficiency games, right? Everyone wanted to use them for creating content, drafting emails, all these really low hanging fruit use cases. That was around. And one of the earlier really nice architectures that came up that I really like was from a 16 z enterprise that was, I think, back in really, really early 2023. LangChain was new, we had land chain and we had all this. Of course, Kudran been there for a long time, but it was the first time that you could see vector store products could be integrated into applications.
|
||||
And our company is also in crop science, as I mentioned, and crop yield prediction. If you could help farmers improve their crop yield, it means that they can produce better products faster with higher quality. So maybe I could start with maybe a history in 2023, what happened? How companies like ours were looking at large language models and opportunities. They bring, I think in 2023, everyone was excited to bring these efficiency games, right? Everyone wanted to use them for creating content, drafting emails, all these really low hanging fruit use cases. That was around. And one of the earlier really nice architectures that came up that I really like was from a 16 z enterprise that was, I think, back in really, really early 2023. LangChain was new, we had land chain and we had all this. Of course, Qdrant been there for a long time, but it was the first time that you could see vector store products could be integrated into applications.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
Really at large scale. There are different components. It's quite complex architecture. So on the right side you see how you can host large language models. On the top you see how you can augment them using external data. Of course, we had these plugins, right? So you can connect these large language models with Google search APIs, all those sort of things, and some validation that are in the middle that you could use to validate the responses fast forward. Maybe I can kind of spend, let me check out the time. Maybe I can spend a few minutes about the components of LLM APIs and hosting because that I think has a lot of potential in terms of applications that need to be really scalable.
|
||||
@@ -109,10 +115,10 @@ Hooman Sedghamiz:
|
||||
And we know that most of the users won't use that. We know that it's a usage based application. You just probably go there. Depending on your daily work, you probably use it. Some people don't use it heavily. I kind of did some calculation. If you build it in house using APIs that you can access yourself, and large language models that corporations can deploy internally and locally, that cost saving could be huge, really magnitudes cheaper, maybe 30 to 20 to 30 times cheaper. So looking, comparing 2024 to 2023, a lot of things have changed.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
Like if you look at the open source large language models that came out really great models from Mistral, now we have models like Llama, two based model, all of these models came out. You can now kind of take a look and see that the performance of them is really, really getting close, if not better than GPT 3.5 already at same level and really approaching step by step to GPT four. And looking at the price on the right side and speed or throughput, you can see that like for example, Mixtrawl seven eight B could be a really cheap option to deploy. And also the performance of it gets really close to GPT 3.5 for many use cases in the enterprise companies. I think two of the big things this year, end of last year that came out that make this kind of really a reality are really a few large language models. I don't know if I can call them large language models. They are like 7 billion to 13 billion compared to GPT four, GT 3.5. I don't think they are really large.
|
||||
Like if you look at the open source large language models that came out really great models from Mistral, now we have models like Llama, two based model, all of these models came out. You can now kind of take a look and see that the performance of them is really, really getting close, if not better than GPT 3.5 already at same level and really approaching step by step to GPT 4. And looking at the price on the right side and speed or throughput, you can see that like for example, Mistral seven eight B could be a really cheap option to deploy. And also the performance of it gets really close to GPT 3.5 for many use cases in the enterprise companies. I think two of the big things this year, end of last year that came out that make this kind of really a reality are really a few large language models. I don't know if I can call them large language models. They are like 7 billion to 13 billion compared to GPT four, GT 3.5. I don't think they are really large.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
But one was Nexus, Raven. We know that applications, if they want to be robust, they really need function calling. We are seeing this paradigm of function calling, which essentially you ask a language model to generate structured output, you give it a function signature, right? You ask it to generate an output, structured output argument for that function. Next was Raven came out last year, that, as you can see here, really is getting really close to GPT four, right? And GPT four being magnitude bigger than this model. This model only being 13 billion parameters really provides really less hallucination, but at the same time really high quality of function calling. So this makes me really excited for the open source and also the companies that want to build their own applications that requires function calling. That was really lacking maybe just five months ago. At the same time, we have really dedicated large language models to programming languages or scripting like SQL, that we are also seeing like SQL coder that's already beating GPT four.
|
||||
But one was Nexus Raven. We know that applications, if they want to be robust, they really need function calling. We are seeing this paradigm of function calling, which essentially you ask a language model to generate structured output, you give it a function signature, right? You ask it to generate an output, structured output argument for that function. Next was Raven came out last year, that, as you can see here, really is getting really close to GPT four, right? And GPT four being magnitude bigger than this model. This model only being 13 billion parameters really provides really less hallucination, but at the same time really high quality of function calling. So this makes me really excited for the open source and also the companies that want to build their own applications that requires function calling. That was really lacking maybe just five months ago. At the same time, we have really dedicated large language models to programming languages or scripting like SQL, that we are also seeing like SQL coder that's already beating GPT four.
|
||||
|
||||
Hooman Sedghamiz:
|
||||
So maybe we can now quickly take a look at how model solving will look like for a large company like ours, like companies that have a lot of people across the globe again, in this aspect also, the community has made really big progress, right? So we have text generation inference from hugging face is open source for most purposes, can be used and it's the choice of mine and probably my group prefers this option. But we have Olama, which is great, a lot of people are using it. We have llama CPP which really optimizes the large language models for local deployment as well, and edge devices. I was really amazed seeing Raspberry PI running a large language model, right? Using Llama CPP. And you have this text generation inference that offers quantization support, continuous patching, all those sort of things that make these large LLMs more quantized or more compressed and also more suitable for deployment to large group of people. Maybe I can kind of give you kind of a quick summary of how, if you decide to deploy these large language models, what techniques you could use to make them more efficient, cost friendly and more scalable. So we have a lot of great open source projects like we have Lite LLM which essentially creates an open AI kind of signature on top of your large language models that you have deployed. Let's say you want to use Azure to host or to access GPT four gypty 3.5 or OpenAI to access OpenAI API.
|
||||
|
||||
@@ -1,17 +1,17 @@
|
||||
---
|
||||
draft: true
|
||||
draft: false
|
||||
title: Production-scale RAG for Real-Time News Distillation - Robert Caulk |
|
||||
Vector Space Talks
|
||||
slug: real-time-news-distillation-rag
|
||||
short_description: Robert Caulk dives into the challenges and innovations in
|
||||
open source AI and news article modeling
|
||||
short_description: Robert Caulk tackles the challenges and innovations in open
|
||||
source AI and news article modeling.
|
||||
description: Robert Caulk, founder of Emergent Methods, discusses the
|
||||
intricacies of context engineering, the power of Newscatcher API for broader
|
||||
complexities of context engineering, the power of Newscatcher API for broader
|
||||
news access, and the sophisticated use of tools like Qdrant for improved
|
||||
recommendation systems, all while emphasizing the importance of efficiency and
|
||||
modularity in technology stacks for real-time data management.
|
||||
preview_image: /blog/from_cms/robert-caulk-bp-cropped.png
|
||||
date: 2024-03-08T10:07:45.278Z
|
||||
date: 2024-03-25T08:49:22.422Z
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
@@ -36,14 +36,14 @@ Robert, Founder of Emergent Methods is a scientist by trade, dedicating his care
|
||||
|
||||
How do Robert Caulk and Emergent Methods contribute to the open-source community, particularly in AI systems and news article modeling?
|
||||
|
||||
In this episode, we're getting under the hood of open-source projects that are reshaping how we interact with AI systems and news article modeling. Robert takes us on a deep dive into the evolving landscape of news distribution and the tech making it more efficient and balanced.
|
||||
In this episode, we'll be learning stuff about open-source projects that are reshaping how we interact with AI systems and news article modeling. Robert takes us on an exploration into the evolving landscape of news distribution and the tech making it more efficient and balanced.
|
||||
|
||||
Here are some takeaways from this episode:
|
||||
|
||||
1. **Context Matters**: Discover the importance of context engineering in news and how it ensures a diversified and consumable information flow.
|
||||
2. **Introducing Newscatcher API**: Get the lowdown on how this tool taps into 50,000 news sources for more thorough and up-to-date reporting.
|
||||
3. **The Magic of Embedding**: Learn about article summarization and semantic search, and how they're crucial for discovering content that truly resonates.
|
||||
4. **Quadrant & Cloud**: Explore how Qdrant's cloud offering and its single responsibility principle support a robust, modular approach to managing news data.
|
||||
4. **Qdrant & Cloud**: Explore how Qdrant's cloud offering and its single responsibility principle support a robust, modular approach to managing news data.
|
||||
5. **Startup Superpowers**: Find out why startups have an edge in implementing new tech solutions and how incumbents are tied down by legacy products.
|
||||
|
||||
> Fun Fact: Did you know that startups' lack of established practices is actually a superpower in the face of new tech paradigms? Legacy products can't keep up!
|
||||
@@ -75,7 +75,7 @@ Here are some takeaways from this episode:
|
||||
|
||||
## Transcript:
|
||||
Demetrios:
|
||||
Robert, it's great to have you here for the vector space talks. I don't know if you're familiar with some of this fun stuff that we do here, but we get to talk with all kinds of experts like yourself on what they're doing when it comes to the vector space and how you've overcome challenges, how you're working through things, because this is a very new field and it is not the most intuitive, as you will tell us more in this upcoming talk. I really am excited because you've been a scientist by trade. Now, you're currently founder at emergent Methods and you've dedicated your career to a variety of open source projects that range from the large scale AI systems to the discrete element modeling. Now at emergent methods, you are adaptively modeling over 1 million news articles per day. That sounds like a whole lot of news articles. And you've been talking and working through production grade rag, which is basically everyone's favorite topic these days. So I know you got to talk for us, man.
|
||||
Robert, it's great to have you here for the vector space talks. I don't know if you're familiar with some of this fun stuff that we do here, but we get to talk with all kinds of experts like yourself on what they're doing when it comes to the vector space and how you've overcome challenges, how you're working through things, because this is a very new field and it is not the most intuitive, as you will tell us more in this upcoming talk. I really am excited because you've been a scientist by trade. Now, you're currently founder at Emergent Methods and you've dedicated your career to a variety of open source projects that range from the large scale AI systems to the discrete element modeling. Now at emergent methods, you are adaptively modeling over 1 million news articles per day. That sounds like a whole lot of news articles. And you've been talking and working through production grade RAG, which is basically everyone's favorite topic these days. So I know you got to talk for us, man.
|
||||
|
||||
Demetrios:
|
||||
I'm going to hand it over to you. I'll bring up your screen right now, and when someone wants to answer or ask a question, feel free to throw it in the chat and I'll jump out at Robert and stop him if needed.
|
||||
@@ -87,22 +87,22 @@ Demetrios:
|
||||
Great to have you here, man. I'm excited for this one.
|
||||
|
||||
Robert Caulk:
|
||||
Thanks for having me, Demetrius. Yeah, it's a great opportunity. I love talking about vector spaces, parameter spaces. So to talk on the show is great. We've got a lot of fun challenges ahead of us in the industry, I think, and the industry is establishing best practices. Like you said, everybody's just trying to figure out what's going on. And some of these base layer tools like Qdrant really enable products and enable companies and they enable us. So let me start.
|
||||
Thanks for having me, Demetrios. Yeah, it's a great opportunity. I love talking about vector spaces, parameter spaces. So to talk on the show is great. We've got a lot of fun challenges ahead of us in the industry, I think, and the industry is establishing best practices. Like you said, everybody's just trying to figure out what's going on. And some of these base layer tools like Qdrant really enable products and enable companies and they enable us. So let me start.
|
||||
|
||||
Robert Caulk:
|
||||
Yeah, like you said, I'm Robert and I'm a founder of emergent methods. Our background, like you said, we are really committed to free and open source software. We started with a lot of narrow AI. Freak AI was one of our original projects, which is AI ML for algo trading very narrow AI, but we came together and built flowdapt. It's a really nice cluster orchestration software, and I'll talk a little bit about that during this presentation. But some of our background goes into, like you said, large scale deep learning for supercomputers. Really cool, interesting stuff. We have some cloud experience.
|
||||
|
||||
Robert Caulk:
|
||||
We really like configuration, so let's dive into it. Why do we actually need to engineer context in the news? There's a lot of reasons why news is important and why it needs to be distributed in a way that's balanced and diversified, but also consumable. Right, let's look at Chat GPT on the left. This is Chat GPT plus it's kind of hanging out searching for Gaza news on Bing, trying to find the top three articles live. Web search is powerful, but it's slow and ultimately inaccurate. What we're building is real time indexing and we couldn't do that without Qdrant, and there's a lot of reasons which I'll be perfectly happy to dive into, but eventually Chappa Chi PT will pull something together here. There it is. And the first thing it reports is 25 day old article with 25 day old nudes.
|
||||
We really like configuration, so let's dive into it. Why do we actually need to engineer context in the news? There's a lot of reasons why news is important and why it needs to be distributed in a way that's balanced and diversified, but also consumable. Right, let's look at Chat GPT on the left. This is Chat GPT plus it's kind of hanging out searching for Gaza news on Bing, trying to find the top three articles live. Web search is powerful, but it's slow and ultimately inaccurate. What we're building is real time indexing and we couldn't do that without Qdrant, and there's a lot of reasons which I'll be perfectly happy to dive into, but eventually Chat GPT will pull something together here. There it is. And the first thing it reports is 25 day old article with 25 day old nudes.
|
||||
|
||||
Robert Caulk:
|
||||
Old news. So it's just inaccurate. So it's borderline dangerous, what's happening here. Right, so this is a very delicate topic. Engineering context in news properly, which takes a lot of energy, a lot of time and dedication and focus, and not every company really has this sort of resource. So we're talking about enforcing journalistic standards, right? OpenAI and Chachipt, they just don't have the time and energy to build a dedicated prompt for this sort of thing. It's fine, they're doing great stuff, they're helping you code. But someone needs to step in and really do enforce some journalistic standards here.
|
||||
Old news. So it's just inaccurate. So it's borderline dangerous, what's happening here. Right, so this is a very delicate topic. Engineering context in news properly, which takes a lot of energy, a lot of time and dedication and focus, and not every company really has this sort of resource. So we're talking about enforcing journalistic standards, right? OpenAI and Chat GPt, they just don't have the time and energy to build a dedicated prompt for this sort of thing. It's fine, they're doing great stuff, they're helping you code. But someone needs to step in and really do enforce some journalistic standards here.
|
||||
|
||||
Robert Caulk:
|
||||
And that includes enforcing diversity, languages, regions and sources. If I'm going to read about Gaza, what's happening over there, you can bet I want to know what Egypt is saying and what France is saying and what Algeria is saying. So let's do this right. That's kind of what we're suggesting, and the only way to do that is to parse a lot of articles. That's how you avoid outdated, stale reporting. And that's a real danger, which is kind of what we saw on that first slide. Everyone here knows hallucination is a problem and it's something you got to minimize, especially when you're talking about the news. It's just a really high cost if you get it wrong.
|
||||
|
||||
Robert Caulk:
|
||||
And so you need people dedicated to this. And if you're going to dedicate a ton of resources and ton of people, you might as well scale that properly. So that's kind of where this comes into. We call this context engineering news context engineering, to be precise, before llama two, which also is enabling products left and right. As we all know, the traditional pipeline was chunk it up, take 512 tokens, put it through a translator, put it through distillbart, do some sentence extraction, and maybe text classification, if you're lucky, get some sentiment out of it and it works. It gets you something. But after we're talking about reading full articles, getting real rich, context, flexible output, translating, summarizing, really deciding that custom extraction on the fly as your product evolves, that's something that the traditional pipeline really just doesn't support. Right.
|
||||
And so you need people dedicated to this. And if you're going to dedicate a ton of resources and ton of people, you might as well scale that properly. So that's kind of where this comes into. We call this context engineering news context engineering, to be precise, before llama two, which also is enabling products left and right. As we all know, the traditional pipeline was chunk it up, take 512 tokens, put it through a translator, put it through distill art, do some sentence extraction, and maybe text classification, if you're lucky, get some sentiment out of it and it works. It gets you something. But after we're talking about reading full articles, getting real rich, context, flexible output, translating, summarizing, really deciding that custom extraction on the fly as your product evolves, that's something that the traditional pipeline really just doesn't support. Right.
|
||||
|
||||
Robert Caulk:
|
||||
We're talking being able to on the fly say, you know what, actually we want to ask this very particular question of all articles and get this very particular field out. And it's really just a prompt modification. This all is based on having some very high quality, base level, diversified news. And so we'll talk a little bit more. But newscatchers is one of the sources that we're using, which opens up 50,000 different sources. So check them out. That's newscatcherapi.com. They even give free access to researchers if you're doing research in this.
|
||||
@@ -114,10 +114,10 @@ Robert Caulk:
|
||||
And then how do we connect the dots here? Of course, there are many ways to go about it. One way which is interesting and fun to talk about is ide. So that's basically a hypothetical document embedding. And what you do is you use the LLM directly to generate a fake article. And that's what we're showing here on the right. So let's say if the user says, what's going on in New York City government, well, you could say, hey, write me just a hypothetical summary based, it could completely fake and use that to create a fake embedding page and use that for the search. Right. So then you're getting a lot closer to where you want to go.
|
||||
|
||||
Robert Caulk:
|
||||
There's some limitations to this, to it's, there's a computational cost also, it's not updated. It's based on whatever. It's basically diving into what it knows about the New York City government and just creating keywords for you. So there's definitely optimizations here as well. When you talk about ambiguity, well, what if the user follows up and says, well, why did they change the rules? Of course, that's where you can start prompt engineering a little bit more and saying, okay, given this historic conversation and the current question, give me some explicit question without ambiguity, and then do the high de, if that's something you want to do. The real goal here is to stay in a single parameter space, a single vector space. Stay as close as possible when you're doing your search as when you do your embedding. So we're talking here about production scale of stuff.
|
||||
There's some limitations to this, to it's, there's a computational cost also, it's not updated. It's based on whatever. It's basically diving into what it knows about the New York City government and just creating keywords for you. So there's definitely optimizations here as well. When you talk about ambiguity, well, what if the user follows up and says, well, why did they change the rules? Of course, that's where you can start prompt engineering a little bit more and saying, okay, given this historic conversation and the current question, give me some explicit question without ambiguity, and then do the high, if that's something you want to do. The real goal here is to stay in a single parameter space, a single vector space. Stay as close as possible when you're doing your search as when you do your embedding. So we're talking here about production scale of stuff.
|
||||
|
||||
Robert Caulk:
|
||||
So I really am happy to geek out about the stack, the open source stack that we're relying on, which includes Qdrant here. But let's start with Vllm. I don't know if you guys have heard of it. This is a really great new project, and their focus on continuous batching and page detention. And if I'm being completely honest with you, it's really above my pay grade in the technicals and how they're actually implementing all of that inside the GPU memory. But what we do is we outsource that to that project and we really like what they're doing, and we've seen really good results. It's increasing throughput. So when you're talking about trying to parse through a million articles, you're going to need a lot of throughput.
|
||||
So I really am happy to geek out about the stack, the open source stack that we're relying on, which includes Qdrant here. But let's start with VLLM. I don't know if you guys have heard of it. This is a really great new project, and their focus on continuous batching and page detention. And if I'm being completely honest with you, it's really above my pay grade in the technicals and how they're actually implementing all of that inside the GPU memory. But what we do is we outsource that to that project and we really like what they're doing, and we've seen really good results. It's increasing throughput. So when you're talking about trying to parse through a million articles, you're going to need a lot of throughput.
|
||||
|
||||
Robert Caulk:
|
||||
The other is text embedding inference. This is a great server. A lot of vector databases will say, okay, we'll do all the embedding for you and we'll do all everything. But when you move to production scale, I'll talk a bit about this later. You need to be using micro service architecture, so it's not super smart to have your database bogged down with doing sorting out the embeddings and sorting out other things. So honestly, I'm a real big fan of single responsibility principle, and that's what Tei does for you. And it also does dynamic batching, which is great in this world where everything is heterogeneous lengths of what's coming in and what's going out. So it's great.
|
||||
@@ -129,7 +129,7 @@ Robert Caulk:
|
||||
The filters are huge. We're talking about real time filtering. We can't be searching on news articles from a month ago, two months ago, if the user is asking for a question that's related to the last 24 hours. So having that timestamp filtering and having it be efficient, which is what it is in Qdrant, is huge. Keyword filtering really opens up a massive realm of product opportunities for us. And then the sparse vectors, we hopped on this train immediately and are just seeing benefits. I don't want to say replacement of elasticsearch, but elasticsearch is using sparse vectors as well. So you can add splade into elasticsearch, and splade is great.
|
||||
|
||||
Robert Caulk:
|
||||
It's a really great alternative to BM 25. It's based on that Burt architecture, and that really opens up a lot of opportunities for filtering out keywords that are kind of useless to the search when the user uses the and a, and then there, these words that are less important splays a bit of a hybrid into semantics, but sparse retrieval. So it's really interesting. And then the idea of hybrid search with semantic and a sparse vector also opens up the ability to do ranking, and you got a higher quality product at the end, which is really the goal, right, especially in production. Point number four here, I would say, is probably one of the most important to us, because we're dealing in a world where latency is king, and being able to deploy Qdrant inside of the same cluster as all the other services. So we're just talking through the switch. That's huge. We're never getting bogged down by network.
|
||||
It's a really great alternative to BM 25. It's based on that architecture, and that really opens up a lot of opportunities for filtering out keywords that are kind of useless to the search when the user uses the and a, and then there, these words that are less important splays a bit of a hybrid into semantics, but sparse retrieval. So it's really interesting. And then the idea of hybrid search with semantic and a sparse vector also opens up the ability to do ranking, and you got a higher quality product at the end, which is really the goal, right, especially in production. Point number four here, I would say, is probably one of the most important to us, because we're dealing in a world where latency is king, and being able to deploy Qdrant inside of the same cluster as all the other services. So we're just talking through the switch. That's huge. We're never getting bogged down by network.
|
||||
|
||||
Robert Caulk:
|
||||
We're never worried about a cloud provider potentially getting overloaded or noisy neighbor problems, stuff like that, completely removed. And then you got high privacy, right. All the data is completely isolated from the external world. So this point number four, I'd say, is one of the biggest value adds for us. But then distributing deployment is huge because high availability is important, and deep storage, which when you're in the business of news archival, and that's one of our main missions here, is archiving the news forever. That's an ever growing database, and so you need a database that's going to be able to grow with you as your data grows. So what's the TLDR to this context? Engineering? Well, service orchestration is really just based on service orchestration in a very heterogeneous and parallel event driven environment. On the right side, we've got the user requests coming in.
|
||||
@@ -141,7 +141,7 @@ Robert Caulk:
|
||||
Open source projects like Qdrant, like Tei, like VLLM and Kubernetes, it's huge. Kubernetes is opening up doors for security and for latency. And of course, if you're going to be getting involved in this game, you got to find the strong DevOps. There's no escaping that. So let's step through kind of piece by piece and talk about flow Dapp. So that's our project. That's our open source project. We've spent about two years building this for our needs, and we're really excited because we did a public open sourcing maybe last week or the week before.
|
||||
|
||||
Robert Caulk:
|
||||
So finally, after all of our testing and rewrites and refactors, we're open. We're open for business. And it's running asknews app right now, and we're really excited for where it's going to go and how it's going to help other people orchestrate their clusters. Our goal and our priorities were highly paralyzed compute and we were running tests using all sorts of different executors, comparing them. So when you use Flowdapt, you can choose ray or dask. And that's key. Especially with vanilla Python, zero code changes, you don't need to know how ray or dask works. In the back end, floatapt is vanilla Python.
|
||||
So finally, after all of our testing and rewrites and refactors, we're open. We're open for business. And it's running asknews app right now, and we're really excited for where it's going to go and how it's going to help other people orchestrate their clusters. Our goal and our priorities were highly paralyzed compute and we were running tests using all sorts of different executors, comparing them. So when you use Flowdapt, you can choose ray or dask. And that's key. Especially with vanilla Python, zero code changes, you don't need to know how ray or dask works. In the back end, flowdapt is vanilla Python.
|
||||
|
||||
Robert Caulk:
|
||||
That was a key goal for us to ensure that we're optimizing how data is moving around the cluster. Automatic resource management this goes back to Ray and dask. They're helping manage the resources of the cluster, allocating a GPU to a task, or allocating multiple tasks to one GPU. These can come in very, very handy when you're dealing with very heterogeneous workloads like the ones that we discussed in those previous slides. For us, the biggest priority was ensuring rapid prototyping and debugging locally. When you're dealing with clusters of 1015 servers, 40 or 5100 with ray, honestly, ray just scales as far as you want. So when you're dealing with that big of a cluster, it's really imperative that what you see on your laptop is also what you are going to see once you deploy. And being able to debug anything you see in the cluster is big for us, we really found the need for easy cluster wide data sharing methods between tasks.
|
||||
@@ -150,19 +150,19 @@ Robert Caulk:
|
||||
So essentially what we've done is made it very easy to get and put values. And so this makes it extremely easy to move data and share data between tasks and make it highly available and stay in cluster memory or persist it to disk, so that when you do the inevitable version update or debug, you're reloading from a persisted state in the real time. News business scheduling is huge. Scheduling, making sure that various workflows are scheduled at different points and different periods or frequencies rather, and that they're being scheduled correctly, and that their triggers are triggering exactly what you need when you need it. Huge for real time. And then one of our biggest selling points, if you will, for this project is Kubernetes style. Everything. Our goal is everything's Kubernetes style, so that if you're coming from Kubernetes, everything's familiar, everything's resource oriented.
|
||||
|
||||
Robert Caulk:
|
||||
We even have our own flow ectyl, which would be the Kubectl style command schemas. A lot of what we've done is ensuring deployment cycle efficiency here. So the goal is that flowdapt can schedule everything and manage all these services for you, create workflows. But why these services? For this particular use case, I'll kind of skip through quickly. I know I'm kind of running out of time here, but of course you're going to need some proprietary remote models. That's just how it works. You're going to of course share that load with on premise llms to reduce cost and to have some reasoning engine on premise. But there's obviously advantages and disadvantages to these.
|
||||
We even have our own flowectl, which would be the Kubectl style command schemas. A lot of what we've done is ensuring deployment cycle efficiency here. So the goal is that flowdapt can schedule everything and manage all these services for you, create workflows. But why these services? For this particular use case, I'll kind of skip through quickly. I know I'm kind of running out of time here, but of course you're going to need some proprietary remote models. That's just how it works. You're going to of course share that load with on premise llms to reduce cost and to have some reasoning engine on premise. But there's obviously advantages and disadvantages to these.
|
||||
|
||||
Robert Caulk:
|
||||
I'm not going to go through them. I'm happy to make these slides available, and you're welcome to kind of parse through the details. Yeah, for sure. You need to start thinking about persistence and search and making sure those services are robust. That's where Qdrant comes into play. And we found that the all in one solutions kind of sacrifice performance for convenience, or sacrifice accuracy for convenience, but it really wasn't for us. We'd rather just orchestrate it ourselves and let Qdrant do what Qdrant does, instead of kind of just hope that an all in one solution is handling it for us and that allows for modularity performance. And we'll dump Qdrant if we want to.
|
||||
|
||||
Robert Caulk:
|
||||
Probably we won't. Or we'll dump minio if we need to, or we'll swap out for whatever replaces bllm. Trying to keep things modular so that future engineers are able to adapt with the tech that's just blowing up and exploding right now. Right. The last thing to talk about here in a production scale environment is really minimizing the latency. I touched on this with Kubernetes ensuring that these services are sitting on the same network, and that is huge. But that talks about decommunication latency. But when you start talking about getting hit with a ton of traffic, production scale, tons of people asking a question all simultaneously, and you needing to go hit a variety of services, well, this is where you really need to isolate that to an asynchronous environment.
|
||||
Probably we won't. Or we'll dump if we need to, or we'll swap out for whatever replaces vllm. Trying to keep things modular so that future engineers are able to adapt with the tech that's just blowing up and exploding right now. Right. The last thing to talk about here in a production scale environment is really minimizing the latency. I touched on this with Kubernetes ensuring that these services are sitting on the same network, and that is huge. But that talks about decommunication latency. But when you start talking about getting hit with a ton of traffic, production scale, tons of people asking a question all simultaneously, and you needing to go hit a variety of services, well, this is where you really need to isolate that to an asynchronous environment.
|
||||
|
||||
Robert Caulk:
|
||||
And of course, if you could write this all in Golang, that's probably going to be your best bet for us. We have some services written in Golang, but predominantly, especially the endpoints that the ML engineers need to work with. We're using fast API on pydantic and honestly, it's powerful. Pydantic V 2.0 now runs on Rust, and as anyone in the Qdrant community knows, Rust is really valuable when you're dealing with highly parallelized environments that require high security and protections for immutability and atomicity. Forgive me for the pronunciation, that kind of sums up the production scale talk, and I'm happy to answer questions. I love diving into this sort of stuff. I do have some just general thoughts on why startups are so much more well positioned right now than some of these incumbents, and I'll just do kind of a quick run through, less than a minute just to kind of get it out there. We can talk about it, see if we agree or disagree.
|
||||
|
||||
Robert Caulk:
|
||||
But you touched on it, Demetrius, in the introduction, which was the best practices have not been established. That's it. That is why startups have such a big advantage. And the reason they're not established is because, well, the new paradigm of technology is just underexplored. We don't really know what the limits are and how to properly handle these things. And that's huge. Meanwhile, some of these incumbents, they're dealing with all sorts of limitations and resistance to change and stuff, and then just market expectations for incumbents maintaining these kind of legacy products and trying to keep them hobbling along on this old tech. In my opinion, startups, you got your reasoning engine building everything around a reasoning engine, using that reasoning engine for every aspect of your system to really open up the adaptivity of your product.
|
||||
But you touched on it, Demetrios, in the introduction, which was the best practices have not been established. That's it. That is why startups have such a big advantage. And the reason they're not established is because, well, the new paradigm of technology is just underexplored. We don't really know what the limits are and how to properly handle these things. And that's huge. Meanwhile, some of these incumbents, they're dealing with all sorts of limitations and resistance to change and stuff, and then just market expectations for incumbents maintaining these kind of legacy products and trying to keep them hobbling along on this old tech. In my opinion, startups, you got your reasoning engine building everything around a reasoning engine, using that reasoning engine for every aspect of your system to really open up the adaptivity of your product.
|
||||
|
||||
Robert Caulk:
|
||||
And okay, I won't put elasticsearch in the incumbent world. I'll keep elasticsearch in the middle. I understand it still has a lot of value, but some of these vendor lock ins, not a huge fan of. But anyway, that's it. That's kind of all I have to say. But I'm happy to take questions or chat a bit.
|
||||
@@ -189,7 +189,7 @@ Robert Caulk:
|
||||
This is another logistical point that we think needs to get sorted properly and there's a few layers to it. So for us, as we're parsing that data coming in from Newscatcher, so newscatcher is doing a good job of always feeding the latest buckets to us. Sometimes one will be kind of arrive, but generally speaking, it's always the latest news. So we're taking five minute buckets, and then with those buckets, we're going through and doing all of our enrichment on that, adding it to Qdrant. And that is the point where we use that timestamp filtering, which is such an important point. So in the metadata of Qdrant, we're using the range filter, which is where we call that the timestamp filter, but it's really range filter, and that helps. So when we're going back to update things, we're sorting and ensuring that we're filtering out only what we haven't seen.
|
||||
|
||||
Demetrios:
|
||||
Okay, that makes complete sense. And basically you could generalize this to something like what I was talking to with people yesterday about, which was, hey, I've got an HR policy that gets updated every other month or every quarter, and I want to make sure that if my HR chat bot is telling people what their vacation policy is, it's pulling from the most recent HR policy. So how do I make sure and do that? And how do I make sure that my vector database isn't like a landmine where it's pulling any information, but we don't necessarily have that control to be able to pull the correct information? And this comes down to that retrieval evaluation, which is such a hot topic, too.
|
||||
Okay, that makes complete sense. And basically you could generalize this to something like what I was talking to with people yesterday about, which was, hey, I've got an HR policy that gets updated every other month or every quarter, and I want to make sure that if my HR chatbot is telling people what their vacation policy is, it's pulling from the most recent HR policy. So how do I make sure and do that? And how do I make sure that my vector database isn't like a landmine where it's pulling any information, but we don't necessarily have that control to be able to pull the correct information? And this comes down to that retrieval evaluation, which is such a hot topic, too.
|
||||
|
||||
Robert Caulk:
|
||||
That's true. No, I think that's a key piece of the puzzle. Now, in that particular example, maybe you actually want to go in and start cleansing a bit, your database, just to make sure if it's really something you're never going to need again. You got to get rid of it. This is a piece I didn't add to the presentation, but it's tangential. You got to keep multiple databases and you got to making sure to isolate resources and cleaning out a database, especially in real time. So ensuring that your database is representative of what you want to be searching on. And you can do this with collections too, if you want.
|
||||
@@ -198,10 +198,10 @@ Robert Caulk:
|
||||
But we find there's sometimes a good opportunity to isolate resources in that sense, 100%.
|
||||
|
||||
Demetrios:
|
||||
So, another question that I had for you was, I noticed Mongo was in the stack. Why did you not just use the Mongo vector option? Is it because of what you were mentioning, where it's like, yeah, you have these all in one options, but you sacrifice that performance for the convenience?
|
||||
So, another question that I had for you was, I noticed Mongo was in the stack. Why did you not just use the Mongo vector option? Is it because of what you were mentioning, where it's like, yeah, you have these all-in-one options, but you sacrifice that performance for the convenience?
|
||||
|
||||
Robert Caulk:
|
||||
We didn't test that, to be honest, I can't say. All I know is we tested weavyt, we tested one other, and I just really like. Although I was going to say I like that it's written in rust, although I believe Mongo is also written in rust, if I'm not mistaken. But for us, the document DB is more of a representation of state and what's happening, especially for our configurations and workflows. Meanwhile, we really like keeping and relying on Qdrant and all the features. Qdrant is updating, so, yeah, I'd say single responsibility principle is key to that. But I saw some chat in Qdrant discord about this, which I think the only way to use vector is actually to use their cloud offering, if I'm not mistaken. Do you know about this?
|
||||
We didn't test that, to be honest, I can't say. All I know is we tested weaviate, we tested one other, and I just really like. Although I was going to say I like that it's written in rust, although I believe Mongo is also written in rust, if I'm not mistaken. But for us, the document DB is more of a representation of state and what's happening, especially for our configurations and workflows. Meanwhile, we really like keeping and relying on Qdrant and all the features. Qdrant is updating, so, yeah, I'd say single responsibility principle is key to that. But I saw some chat in Qdrant discord about this, which I think the only way to use vector is actually to use their cloud offering, if I'm not mistaken. Do you know about this?
|
||||
|
||||
Demetrios:
|
||||
Yeah, I think so, too.
|
||||
|
||||
@@ -1,16 +1,16 @@
|
||||
---
|
||||
draft: true
|
||||
draft: false
|
||||
title: Talk with YouTube without paying a cent - Francesco Saverio Zuppichini |
|
||||
Vector Space Talks
|
||||
slug: youtube-without-paying-cent
|
||||
short_description: Dive deep into the tech world as Francesco shares his
|
||||
insights and processes on coding innovative solutions.
|
||||
description: Francesco Zuppichini intricately outlines the process of converting
|
||||
YouTube video subtitles into searchable vector databases, leveraging tools
|
||||
like YouTube DL and Hugging Face, and addressing the challenges of coding
|
||||
without conventional frameworks in machine learning engineering.
|
||||
short_description: A sneak peek into the tech world as Francesco shares his
|
||||
ideas and processes on coding innovative solutions.
|
||||
description: Francesco Zuppichini outlines the process of converting YouTube
|
||||
video subtitles into searchable vector databases, leveraging tools like
|
||||
YouTube DL and Hugging Face, and addressing the challenges of coding without
|
||||
conventional frameworks in machine learning engineering.
|
||||
preview_image: /blog/from_cms/francesco-saverio-zuppichini-bp-cropped.png
|
||||
date: 2024-03-08T10:12:59.752Z
|
||||
date: 2024-03-27T12:37:55.643Z
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
@@ -35,11 +35,11 @@ Francesco Saverio Zuppichini is a Senior Full Stack Machine Learning Engineer at
|
||||
|
||||
Curious about transforming YouTube content into searchable elements? Francesco Zuppichini unpacks the journey of coding a RAG by using subtitles as input, harnessing technologies like YouTube DL, Hugging Face, and Qdrant, while debating framework reliance and the fine art of selecting the right software tools.
|
||||
|
||||
Here are some insights from this great episode:
|
||||
Here are some insights from this episode:
|
||||
|
||||
1. **Behind the Code**: Francesco unravels how to create a RAG using YouTube videos. Get ready to geek out on the nuts and bolts that make this magic happen.
|
||||
2. **Vector Voodoo**: Ever wonder how embedding vectors carry out their similarity searches? Francesco's got you covered with his brilliant explanation of vector databases and the mind-bending distance method that seeks out those matches.
|
||||
3. **Function over Class**: The debate is as old as stardust. Francesco shares why he prefers using functions over classes for better code organization and demonstrates how this approach crystallizes when running language models with Ollama.
|
||||
3. **Function over Class**: The debate is as old as stardust. Francesco shares why he prefers using functions over classes for better code organization and demonstrates how this approach solidifies when running language models with Ollama.
|
||||
4. **Metadata Magic**: Find out how metadata isn't just a sidekick but plays a pivotal role in the realm of Qdrant and RAGs. Learn why Francesco values metadata as payload and the challenges it presents in developing domain-specific applications.
|
||||
5. **Tool Selection Tips**: Deciding on the right software tool can feel like navigating an asteroid belt. Francesco shares his criteria—ease of installation, robust documentation, and a little help from friends—to ensure a safe landing.
|
||||
|
||||
@@ -132,10 +132,10 @@ Demetrios:
|
||||
Yes, for sure.
|
||||
|
||||
Francesco Zuppichini:
|
||||
That's perfect. Okay, so today we're going to talk about talk with YouTube without paying a cent, no framework bs. So the goal of today is to showcase how to code a RAG given as an input a YouTube video without using any framework like language, et cetera, et cetera. And I want to show you that it's straightforward, using a bunch of technologies and Qdrants as well. And you can do all of this without actually pay to any service. Right. So we are going to run our PetrodB locally and also the language model. We are going to run our machines.
|
||||
That's perfect. Okay, so today we're going to talk about talk with YouTube without paying a cent, no framework bs. So the goal of today is to showcase how to code a RAG given as an input a YouTube video without using any framework like language, et cetera, et cetera. And I want to show you that it's straightforward, using a bunch of technologies and Qdrants as well. And you can do all of this without actually pay to any service. Right. So we are going to run our PEDro DB locally and also the language model. We are going to run our machines.
|
||||
|
||||
Francesco Zuppichini:
|
||||
And yeah, it's going to be a technical talk, so I will kind of guide you through the code. Feel free to interrupt me at any time if you have questions, if you want to ask why I did that, et cetera, et cetera. So very quickly, before we get started, I just want you not to introduce myself. So yeah, senior full stack machine engineer. That's just a bunch of funny work to basically say that I do a little bit of everything. Start. So when I was working, I start as computer vision engineer, I work at PwC, then a bunch of startups, and now I sold my soul to insurance companies working at insurance. And before I was doing computer vision, now I'm doing due to chat CPT, hyper language model, I'm doing more of that.
|
||||
And yeah, it's going to be a technical talk, so I will kind of guide you through the code. Feel free to interrupt me at any time if you have questions, if you want to ask why I did that, et cetera, et cetera. So very quickly, before we get started, I just want you not to introduce myself. So yeah, senior full stack machine engineer. That's just a bunch of funny work to basically say that I do a little bit of everything. Start. So when I was working, I start as computer vision engineer, I work at PwC, then a bunch of startups, and now I sold my soul to insurance companies working at insurance. And before I was doing computer vision, now I'm doing due to Chat GPT, hyper language model, I'm doing more of that.
|
||||
|
||||
Francesco Zuppichini:
|
||||
But I'm always involved in bringing the full product together. So from zero to something that is deployed and running. So I always be interested in web dev. I can also do website servers, a little bit of infrastructure as well. So now I'm just doing a little bit of everything. So this is why there is full stack there. Yeah. Okay, let's get started to something a little bit more interesting than myself.
|
||||
@@ -165,13 +165,13 @@ Francesco Zuppichini:
|
||||
Wonderful. Okay, so in order to get the embedding. So to translate from text to vectors, right, so we're going to use hugging face just an embedding model so we can actually get some vectors. Then as soon as we got our vectors, we need to store and search them. So we're going to use our beloved Qdrant to do so. We also need to keep a little bit of stage right because we need to know which video we have processed so we don't redo the old embeddings and the storing every time we see the same video. So for this part, I'm just going to use SQLite, which is just basically an SQL database in just a file. So very easy to use, very kind of lightweight, and it's only your computer, so it's safe to run the language model.
|
||||
|
||||
Francesco Zuppichini:
|
||||
We're going to use Olama. That is a very simple way and very well done way to just get a language model that is running on your computer. And you can also call it using the OpenAI Python library because they have implemented the same endpoint as. It's like, it's super convenient, super easy to use. If you already have some code that is calling OpenAI, you can just run a different language model using Olama. And you just need to basically change two lines of code. So what we're going to do, basically, I'm going to take a video. So here it's a video from Fireship IO.
|
||||
We're going to use Ollama. That is a very simple way and very well done way to just get a language model that is running on your computer. And you can also call it using the OpenAI Python library because they have implemented the same endpoint as. It's like, it's super convenient, super easy to use. If you already have some code that is calling OpenAI, you can just run a different language model using Ollama. And you just need to basically change two lines of code. So what we're going to do, basically, I'm going to take a video. So here it's a video from Fireship IO.
|
||||
|
||||
Francesco Zuppichini:
|
||||
We're going to run our command line and we're going to ask some questions. Now, if you can still, in theory, you should be able to see my full screen. Yeah. So very quickly to showcase that to you, I already processed this video from the good sound YouTube channel and I have already here my command line. So I can already kind of see, you know, I can ask a question like what is the contact size of Germany? And we're going to get the reply. Yeah. And here we're going to get a reply. And now I want to walk you through how you can do something similar.
|
||||
|
||||
Francesco Zuppichini:
|
||||
Now, the goal is not to create the best rack in the world. It's just to showcase like show zero to something that is actually working. How you can do that in a fully local way without using any framework so you can really understand what's going on under the hood. Because I think a lot of people, they try to copy, to just copy and paste stuff on langchain and then they end up in a situation when they need to change something, but they don't really know where the stuff is. So this is why I just want to just show like Windfield zero to hero. So the first step will be I get a YouTube video and now I need to get the subtitle. So you could actually use a model to take the audio from the video and get the text. Like a whisper model from OpenAI, for example.
|
||||
Now, the goal is not to create the best rack in the world. It's just to showcase like show zero to something that is actually working. How you can do that in a fully local way without using any framework so you can really understand what's going on under the hood. Because I think a lot of people, they try to copy, to just copy and paste stuff on Langchain and then they end up in a situation when they need to change something, but they don't really know where the stuff is. So this is why I just want to just show like Windfield zero to hero. So the first step will be I get a YouTube video and now I need to get the subtitle. So you could actually use a model to take the audio from the video and get the text. Like a whisper model from OpenAI, for example.
|
||||
|
||||
Francesco Zuppichini:
|
||||
In this case, we are taking advantage that YouTube allow people to upload subtitles and YouTube will automatically generate the subtitles. So here using YouTube dial, I'm just going to get my video URL. I'm going to set up a bunch of options like the format they want, et cetera, et cetera. And then basically I'm going to download and get the subtitles. And they look something like this. Let me show you an example. Something similar to this one, right? We have the timestamps and we do have all text inside. Now the next step.
|
||||
@@ -195,7 +195,7 @@ Francesco Zuppichini:
|
||||
You just need to do this once. I was very lazy so I just assumed that if this is going to fail, it means that it's because I've already created a collection. So I'm just going to pass it and call it a day. Okay, so this is basically all the preprocess this setup you need to do to have your Qdrant ready to store and search vectors. To store vectors. Straightforward, very straightforward as well. Just need again the client. So the connection to the database here I'm passing my embedding so sentence transformer model and I'm passing my chunks as a list of documents.
|
||||
|
||||
Francesco Zuppichini:
|
||||
So documents in my code is just a type dict that will contain just this metadata here. Very simple. It's similar to Lang chain here. I just have attacked it because it's lightweight. To store them we call the upload records function. We encode them here. There is a little bit of bad variable names from my side which I replacing that. So you shouldn't do that.
|
||||
So documents in my code is just a type that will contain just this metadata here. Very simple. It's similar to Lang chain here. I just have attacked it because it's lightweight. To store them we call the upload records function. We encode them here. There is a little bit of bad variable names from my side which I replacing that. So you shouldn't do that.
|
||||
|
||||
Francesco Zuppichini:
|
||||
Apologize about that and you just send the records. Another very cool thing about Qdrant. So the second things that I really like is that they have types for what you send through the library. So this models record is a Qdrant type. So you use it and you know immediately. So what you need to put inside. So let me give you an example. Right? So assuming that I'm programming, right, I'm going to say model record bank.
|
||||
@@ -213,13 +213,13 @@ Francesco Zuppichini:
|
||||
We need to recreate to embed in a vector and then we need to compare with the vectors in the vector Db using a distance method, in this case considered similarity in order to get the right matches right, the closest one in our vector DB, in our vector search base. So passing a query string, I'm passing a video id and I pass in a label. So how many hits I want to get from the metadb. Now to create a filter again you're going to use the model package from the Qdrant framework. So here I'm just creating a filter class for the model and I'm saying okay, this filter must match this key, right? So metadata video id with this video id. So when we search, before we do the similarity search, we are going to filter away all the vectors that are not from that video. Wonderful. Now super easy as well.
|
||||
|
||||
Francesco Zuppichini:
|
||||
We just call the DB search, right pass. Our collection name here is star coded. Apologies about that, I think I forgot to put the right global variable our coded, we create a query, we set the limit, we pass the query filter, we get the it back as a dictionary in the payload field of each it and we recreate our document a dictionary. I have types, right? So I know what this function is going to return. Now if you were to use a framework, right this part, it will be basically the same thing. If I were to use lamb chain and I want to specify a filter, I would have to write the same amount of code. So most of the times you don't really need to use a framework. One thing that is nice about not using a framework here is that I add control on the indexes.
|
||||
We just call the DB search, right pass. Our collection name here is star coded. Apologies about that, I think I forgot to put the right global variable our coded, we create a query, we set the limit, we pass the query filter, we get the it back as a dictionary in the payload field of each it and we recreate our document a dictionary. I have types, right? So I know what this function is going to return. Now if you were to use a framework, right this part, it will be basically the same thing. If I were to use langchain and I want to specify a filter, I would have to write the same amount of code. So most of the times you don't really need to use a framework. One thing that is nice about not using a framework here is that I add control on the indexes.
|
||||
|
||||
Francesco Zuppichini:
|
||||
Long chain, for instance, will create the indexes only while you call a classmate like from document. And that is kind of cumbersome because sometimes I wasn't quoting bugs in which I was not understanding why one index was created before, after, et cetera, et cetera. So yes, just try to keep things simple and not always write on frameworks. Wonderful. Now I have a way to ask a query to get back the relative parts from that video. Now we need to translate this list of chunks to something that we can read as human. Before we do that, I was almost going to forget we need to keep state. Now, one of the last missing part is something in which I can store data.
|
||||
Lang chain, for instance, will create the indexes only while you call a classmate like from document. And that is kind of cumbersome because sometimes I wasn't quoting bugs in which I was not understanding why one index was created before, after, et cetera, et cetera. So yes, just try to keep things simple and not always write on frameworks. Wonderful. Now I have a way to ask a query to get back the relative parts from that video. Now we need to translate this list of chunks to something that we can read as human. Before we do that, I was almost going to forget we need to keep state. Now, one of the last missing part is something in which I can store data.
|
||||
|
||||
Francesco Zuppichini:
|
||||
Here I just have a setup function in which I'm going to create an SQL lite database, create a table called videos in which I have an id and a title. So later I can check, hey, is this video already in my database? Yes. I don't need to process that. I can just start immediately to q a on that video. If not, I'm going to do the chunking and embeddings. Got a couple of functions here to get video from Db to save video from and to save video to Db. So notice now I only use functions. I'm not using classes here.
|
||||
Here I just have a setup function in which I'm going to create an SQL lite database, create a table called videos in which I have an id and a title. So later I can check, hey, is this video already in my database? Yes. I don't need to process that. I can just start immediately to QA on that video. If not, I'm going to do the chunking and embeddings. Got a couple of functions here to get video from Db to save video from and to save video to Db. So notice now I only use functions. I'm not using classes here.
|
||||
|
||||
Francesco Zuppichini:
|
||||
I'm not a fan of object writing programming because it's very easy to kind of reach inheritance health in which we have like ten levels of inheritance. And here if a function needs to have state, here we do need to have state because we need a connection. So I will just have a function that initialize that state. I return tat to me, and me as a caller, I'm just going to call it and pass my state. Very simple tips allow you really to divide your code properly. You don't need to think about is my class to couple with another class, et cetera, et cetera. Very simple, very effective. So what I suggest when you're coding, just start with function and share states across just pass down state.
|
||||
@@ -228,13 +228,13 @@ Francesco Zuppichini:
|
||||
And when you realize that you can cluster a lot of function together with a common behavior, you can go ahead and put state in a class and have key function as methods. So try to not start first by trying to understand which class I need to use around how I connect them, because in my opinion it's just a waste of time. So just start with function and then try to cluster them together if you need to. Okay, last part, the juicy part as well. Language models. So we need the language model. Why do we need the language model? Because I'm going to ask a question, right. I'm going to get a bunch of relevant chunks from a video and the language model.
|
||||
|
||||
Francesco Zuppichini:
|
||||
It needs to answer that to me. So it needs to get information from the chunks and reply that to me using that information as a context. To run language model, the easiest way in my opinion is using Olama. There are a lot of models that are available. I put a link here and you can also bring your own model. There are a lot of videos and tutorial how to do that. You run this command as soon as you install it on Linux. It's a one line to install o llama.
|
||||
It needs to answer that to me. So it needs to get information from the chunks and reply that to me using that information as a context. To run language model, the easiest way in my opinion is using Ollama. There are a lot of models that are available. I put a link here and you can also bring your own model. There are a lot of videos and tutorial how to do that. You run this command as soon as you install it on Linux. It's a one line to install Ollama.
|
||||
|
||||
Francesco Zuppichini:
|
||||
You run this command here, it's going to download Mistral seven B very good model and run it on your gpu if you have one, or your cpu if you don't have a gpu, run it on GPU. Here you can see it yet. It's around 6gb. So even with a low tier gpu, you should be able to run a seven minute model on your gpu. Okay, so this is the prompt just for also to show you how easy is this, this prompt was just very lazy. Copy and paste from langchain source code here prompt use the following piece of context to answer the question at the end. Blah blah blah variable to inject the context inside question variable to get question and then we're going to get an answer. How do we call it? Is it peasy? I have a function here called getanswer passing a bunch of stuff, passing also the OpenAI from the OpenAI Python package model client passing a question, passing a vdb, my Db client, my embeddings, reading my prompt, getting my matching documents, calling the search function we have just seen before, creating my context.
|
||||
You run this command here, it's going to download Mistral 7B very good model and run it on your gpu if you have one, or your cpu if you don't have a gpu, run it on GPU. Here you can see it yet. It's around 6gb. So even with a low tier gpu, you should be able to run a seven minute model on your gpu. Okay, so this is the prompt just for also to show you how easy is this, this prompt was just very lazy. Copy and paste from langchain source code here prompt use the following piece of context to answer the question at the end. Blah blah blah variable to inject the context inside question variable to get question and then we're going to get an answer. How do we call it? Is it easy? I have a function here called getanswer passing a bunch of stuff, passing also the OpenAI from the OpenAI Python package model client passing a question, passing a vdb, my DB client, my embeddings, reading my prompt, getting my matching documents, calling the search function we have just seen before, creating my context.
|
||||
|
||||
Francesco Zuppichini:
|
||||
So just joining the text in the chunks on a new line, calling the format function in Python. As simple as that. Just calling the format function in Python because the format function will look at a string and kitty will inject variables that match inside these parentheses. Passing context passing question using the Openi model client APIs and getting a reply back. Super easy. And here I'm returning the reply from the language model and also the list of documents. So this should be documents. I think I did a mistake.
|
||||
So just joining the text in the chunks on a new line, calling the format function in Python. As simple as that. Just calling the format function in Python because the format function will look at a string and kitty will inject variables that match inside these parentheses. Passing context passing question using the OpenAI model client APIs and getting a reply back. Super easy. And here I'm returning the reply from the language model and also the list of documents. So this should be documents. I think I did a mistake.
|
||||
|
||||
Francesco Zuppichini:
|
||||
When I copy and paste this to get this image and we are done right. We have a way to get some answers from a video by putting everything together. This can seem scary because there is no comment here, but I can show you tson code. I think it's easier so I can highlight stuff. I'm creating my embeddings, I'm getting my database, I'm getting my vector DB login, some stuff I'm getting my model client, I'm getting my vid. So here I'm defining the state that I need. You don't need comments because I get it straightforward. Like here I'm getting the vector db, good function name.
|
||||
@@ -288,7 +288,7 @@ Francesco Zuppichini:
|
||||
Nice to everyone by the way.
|
||||
|
||||
Demetrios:
|
||||
So from my side, I'm wondering, do you have any specific design decisions criteria that you use when you are building out your stack? Like you chose Mistral, you chose Olama, you chose Qdrant. It sounds like with Qdrant you did some testing and you appreciated the capabilities. With Qdrant, was it similar with Olama and Mistral?
|
||||
So from my side, I'm wondering, do you have any specific design decisions criteria that you use when you are building out your stack? Like you chose Mistral, you chose Ollama, you chose Qdrant. It sounds like with Qdrant you did some testing and you appreciated the capabilities. With Qdrant, was it similar with Ollama and Mistral?
|
||||
|
||||
Francesco Zuppichini:
|
||||
So my test is how long it's going to take to install that tool. If it's taking too much time and it's hard to install because documentation is bad, so that it's a red flag, right? Because if it's hard to install and documentation is bad for the installation, that's the first thing people are going to read. So probably it's not going to be great for something down the road to use Olama. It took me two minutes, took me two minutes, it was incredible. But just install it, run it and it was done. Same thing with Qualent as well and same thing with the hacking phase library. So to me, usually as soon as if I see that something is easy to install, that's usually means that is good. And if the documentation to install it, it's good.
|
||||
@@ -318,7 +318,7 @@ Sabrina Aquino:
|
||||
Yeah, that's a great explanation of collections. And I do love your approach of having everything locally and having everything in a structured way that you can really understand what you're doing. And I know you mentioned sometimes frameworks are not necessary. And I wonder also from your side, when do you think a framework would be necessary and does it have to do with scaling? What do you think?
|
||||
|
||||
Francesco Zuppichini:
|
||||
So that's a great question. So what frameworks in theory should give you is good interfaces, right? So a good interface means that if I'm following that interface, I know that I can always call something that implements that interface in the same way. Like for instance in Lanchain, if I call a betterdb, I can just swap the betterdb and I can call it in the same way. If the interfaces are good, the framework is useful. If you know that you are going to change stuff. In my case, I know from the beginning that I'm going to use Qdrant, I'm going to use Holama, and I'm going to use SQL lite. So why should I go to the hello reading framework documentation? I install libraries, and then you need to install a bunch of packages from the framework that you don't even know why you need them. Maybe you have a conflict package, et cetera, et cetera.
|
||||
So that's a great question. So what frameworks in theory should give you is good interfaces, right? So a good interface means that if I'm following that interface, I know that I can always call something that implements that interface in the same way. Like for instance in Langchain, if I call a betterdb, I can just swap the betterdb and I can call it in the same way. If the interfaces are good, the framework is useful. If you know that you are going to change stuff. In my case, I know from the beginning that I'm going to use Qdrant, I'm going to use Ollama, and I'm going to use SQL lite. So why should I go to the hello reading framework documentation? I install libraries, and then you need to install a bunch of packages from the framework that you don't even know why you need them. Maybe you have a conflict package, et cetera, et cetera.
|
||||
|
||||
Francesco Zuppichini:
|
||||
If you know ready. So what you want to do then just code it and call it a day? Like in this case, I know I'm not going to change the vector DB. If you think that you're going to change something, even if it's a simple approach, it's fair enough, simple to change stuff. Like I will say that if you know that you want to change your vector DB providers, either you define your own interface or you use a framework with an already defined interface. But be careful because right too much on framework will. First of all, basically you don't know what's going on inside the hood for launching because it's so kudos to them. They were the first one. They are very smart people, et cetera, et cetera.
|
||||
|
||||
@@ -1,16 +1,16 @@
|
||||
---
|
||||
draft: true
|
||||
draft: false
|
||||
title: "VirtualBrain: Best RAG to unleash the real power of AI - Guillaume
|
||||
Marquis | Vector Space Talks"
|
||||
slug: virtualbrain-best-rag
|
||||
short_description: Delve into the complex world of information retrieval with
|
||||
Guillaume Marquis, CTO & Co-founder at VirtualBrain.
|
||||
description: Guillaume Marquis, CTO & Co-founder at VirtualBrain, reveals the
|
||||
short_description: Let's explore information retrieval with Guillaume Marquis,
|
||||
CTO & Co-Founder at VirtualBrain.
|
||||
description: Guillaume Marquis, CTO & Co-Founder at VirtualBrain, reveals the
|
||||
mechanics of advanced document retrieval with RAG technology, discussing the
|
||||
challenges of scalability, up-to-date information, and navigating user
|
||||
feedback to enhance the productivity of knowledge workers.
|
||||
preview_image: /blog/from_cms/guillaume-marquis-2-cropped.png
|
||||
date: 2024-03-11T08:35:54.200Z
|
||||
date: 2024-03-27T12:41:51.859Z
|
||||
author: Demetrios Brinkmann
|
||||
featured: false
|
||||
tags:
|
||||
@@ -23,19 +23,19 @@ tags:
|
||||
— Guillaume Marquis
|
||||
>
|
||||
|
||||
Guillaume Marquis, a dedicated Engineer and AI enthusiast, serves as the Chief Technology Officer and Co-founder of VirtualBrain, an innovative AI company. He is committed to exploring novel approaches to integrating artificial intelligence into everyday life, driven by a passion for advancing the field and its applications.
|
||||
Guillaume Marquis, a dedicated Engineer and AI enthusiast, serves as the Chief Technology Officer and Co-Founder of VirtualBrain, an innovative AI company. He is committed to exploring novel approaches to integrating artificial intelligence into everyday life, driven by a passion for advancing the field and its applications.
|
||||
|
||||
***Listen to the episode on Spotify, Apple Podcast, Podcast addicts, Castbox. You can also watch this episode on YouTube.***
|
||||
***Listen to the episode on [Spotify](https://open.spotify.com/episode/20iFzv2sliYRSHRy1QHq6W?si=xZqW2dF5QxWsAN4nhjYGmA), Apple Podcast, Podcast addicts, Castbox. You can also watch this episode on [YouTube](https://youtu.be/v85HqNqLQcI?feature=shared).***
|
||||
|
||||
[embed YouTube video here]
|
||||
<iframe width="560" height="315" src="https://www.youtube.com/embed/v85HqNqLQcI?si=hjUiIhWxsDVO06-H" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
|
||||
|
||||
[embed anchor.fm podcast here]
|
||||
<iframe src="https://podcasters.spotify.com/pod/show/qdrant-vector-space-talk/embed/episodes/VirtualBrain-Best-RAG-to-unleash-the-real-power-of-AI---Guillaume-Marquis--Vector-Space-Talks-017-e2grbfg/a-ab22dgt" height="102px" width="400px" frameborder="0" scrolling="no"></iframe>
|
||||
|
||||
## **Top takeaways:**
|
||||
|
||||
Who knew that document retrieval could be creative? Guillaume and VirtualBrain help draft sales proposals using past reports. Fascinating how tech aids deep work beyond basic search tasks.
|
||||
Who knew that document retrieval could be creative? Guillaume and VirtualBrain help draft sales proposals using past reports. It's fascinating how tech aids deep work beyond basic search tasks.
|
||||
|
||||
Delving into document retrieval and AI assistance, Guillaume furthermore unpacks the intricacies of sifting through vast data using a scoring system, the virtue of RAG for deep work, and combating the 'illusion of work', enhancing insights for knowledge workers while confronting the challenges of scalability and user feedback on hallucinations.
|
||||
Tackling document retrieval and AI assistance, Guillaume furthermore unpacks the ins and outs of searching through vast data using a scoring system, the virtue of RAG for deep work, and going through the 'illusion of work', enhancing insights for knowledge workers while confronting the challenges of scalability and user feedback on hallucinations.
|
||||
|
||||
Here are some key insight from this episode you need to look out for:
|
||||
|
||||
@@ -51,7 +51,7 @@ Here are some key insight from this episode you need to look out for:
|
||||
|
||||
## Show notes:
|
||||
|
||||
00:00 Hosts’ and guest recommendations.\
|
||||
00:00 Hosts and guest recommendations.\
|
||||
09:01 Leveraging past knowledge to create new proposals.\
|
||||
12:33 Ingesting and parsing documents for context retrieval.\
|
||||
14:26 Creating and storing data, performing advanced searches.\
|
||||
@@ -69,7 +69,7 @@ Here are some key insight from this episode you need to look out for:
|
||||
*"We only exclusively use open source tools because of security aspects and stuff like that. That's why also we are using Qdrant one of the important point on that. So we have a system, we are using this serverless stuff to ingest document over time.”*\
|
||||
— Guillaume Marquis
|
||||
|
||||
*"One of the challenging part was the scalability of the system. We have clients that come with terra octave of data and want to be parsed really fast and so you have the ingestion, but even after the semantic search, even on a large data set can be slow. And today Chat GPT answer really fast. So your users, even if the question is way more complicated to answer than a basic Chat GPT question, they want to have their answer in seconds. So you have also this challenge that is really you have to take care.”*\
|
||||
*"One of the challenging part was the scalability of the system. We have clients that come with terra octave of data and want to be parsed really fast and so you have the ingestion, but even after the semantic search, even on a large data set can be slow. And today ChatGPT answers really fast. So your users, even if the question is way more complicated to answer than a basic ChatGPT question, they want to have their answer in seconds. So you have also this challenge that you really have to take care.”*\
|
||||
— Guillaume Marquis
|
||||
|
||||
*"Our AI is not trained to write you a speech based on Shakespeare and with the style of Martin Luther King. It's not the purpose of the tool. So if you ask something that is out of the box, he will just say like, okay, I don't know how to answer that. And that's an important point. That's a feature by itself to be able to not go outside of the box.”*\
|
||||
@@ -80,7 +80,7 @@ Demetrios:
|
||||
So, dude, I'm excited for this talk. Before we get into it, I want to make sure that we have some pre conversation housekeeping items that go out, one of which being, as always, we're doing these vector space talks and everyone is encouraged and invited to join in. Ask your questions, let us know where you're calling in from, let us know what you're up to, what your use case is, and feel free to drop any questions that you may have in the chat. We will be monitoring it like a hawk. Today I am joined by none other than Sabrina. How are you doing, Sabrina?
|
||||
|
||||
Sabrina Aquino:
|
||||
What's up, dementias? I'm doing great. Excited to be here. I just love seeing what amazing stuff people are building with Qdrant and. Yeah, let's get into it.
|
||||
What's up, Demetrios? I'm doing great. Excited to be here. I just love seeing what amazing stuff people are building with Qdrant and. Yeah, let's get into it.
|
||||
|
||||
Demetrios:
|
||||
Yeah. So I think I see Sabrina's wearing a special shirt which is don't get lost in vector space shirt. If anybody wants a shirt like that. There we go. Well, we got you covered, dude. You will get one at your front door soon enough. If anybody else wants one, come on here. Present at the next vector space talks.
|
||||
@@ -89,7 +89,7 @@ Demetrios:
|
||||
We're excited to have you. And we've got one last thing that I think is fun that we can talk about before we jump into the tech piece of the conversation. And that is I told Sabrina to get ready with some recommendations. Know vector databases, they can be used occasionally for recommendation systems, but nothing's better than getting that hidden gem from your friend. And right now what we're going to try and do is give you a few hidden gems so that the next time the recommendation engine is working for you, it's working in your favor. And Sabrina, I asked you to give me one music that you can recommend, one show and one rando. So basically one random thing that you can recommend to us.
|
||||
|
||||
Sabrina Aquino:
|
||||
So I've picked. I thought about this. Okay, I give it some thought. The movie would be catch me if you can by Leo DiCaprio and tone Hanks. Have you guys watched it? Really good movie. The song would be oh, children by knee cave and the bad scenes. Also very good song. And the random recommendation is my favorite scented candle, which is citrus notes, sea salt and cedar.
|
||||
So I've picked. I thought about this. Okay, I give it some thought. The movie would be Catch Me If You Can by Leo DiCaprio and Tom Hanks. Have you guys watched it? Really good movie. The song would be oh, children by knee cave and the bad scenes. Also very good song. And the random recommendation is my favorite scented candle, which is citrus notes, sea salt and cedar.
|
||||
|
||||
Sabrina Aquino:
|
||||
So there you go.
|
||||
@@ -98,13 +98,13 @@ Demetrios:
|
||||
A scented candle as a recommendation. I like it. I think that's cool. I didn't exactly tell you to get ready with that. So I'll go next, then you can have some more time to think. So for anybody that's joining in, we're just giving a few recommendations to help your own recommendation engines at home. And we're going to get into this conversation about rags in just a moment. But my song is with.
|
||||
|
||||
Demetrios:
|
||||
Oh, my God. I've been listening to it because I didn't think that they had it on Spotify, but I found it this morning and I was so happy that they did. And it is Bill Evans and Chet Baker. Basically, their whole album, the legendary sessions, is just like, incredible. But the first song on that album is called alone together. And when Chet Baker starts playing his little trombone, my God, it is like you can feel emotion. You can touch it. That is what I would recommend.
|
||||
Oh, my God. I've been listening to it because I didn't think that they had it on Spotify, but I found it this morning and I was so happy that they did. And it is Bill Evans and Chet Baker. Basically, their whole album, the legendary sessions, is just like, incredible. But the first song on that album is called Alone Together. And when Chet Baker starts playing his little trombone, my God, it is like you can feel emotion. You can touch it. That is what I would recommend.
|
||||
|
||||
Demetrios:
|
||||
Anyone out there? I'll drop a link in the chat if you like it. The film or series. This fool, if you speak Spanish, it's even better. It is amazing series. Get that, do it. And as the rando thing, I've been having Rishi mushroom powder in my coffee in the mornings. I highly recommend it. All right, last one, let's get into your recommendations and then we'll get into this rag chat.
|
||||
|
||||
Guillaume Marquis:
|
||||
So, yeah, I sucked a little bit. So for the song, I think I will give something like, because I'm french, I think you can hear it. So I will choose get lucky of daft Punk and because I am a little bit sad of the end of their collaboration. So, yeah, just like, I cannot forget it. And it's a really good music. Like, miss them as a movie, maybe something like I really enjoy. So we have a lot of french movies that are really nice, but something more international maybe, and more mainstream. Jungle of Tarantino, that is really a good movie and really enjoy it.
|
||||
So, yeah, I sucked a little bit. So for the song, I think I will give something like, because I'm french, I think you can hear it. So I will choose Get Lucky of Daft Punk and because I am a little bit sad of the end of their collaboration. So, yeah, just like, I cannot forget it. And it's a really good music. Like, miss them as a movie, maybe something like I really enjoy. So we have a lot of french movies that are really nice, but something more international maybe, and more mainstream. Jungle of Tarantino, that is really a good movie and really enjoy it.
|
||||
|
||||
Guillaume Marquis:
|
||||
I watched it several times and still a good movie to watch. And random thing, maybe a city. A city to go to visit. I really enjoyed. It's hard to choose. Really hard to choose a place in general. Okay, Florence, like in Italy.
|
||||
@@ -140,10 +140,10 @@ Demetrios:
|
||||
I have the million dollar question that I think is probably coming through everyone's head is like, you're retrieving so many documents, how are you evaluating your retrieval?
|
||||
|
||||
Guillaume Marquis:
|
||||
That's definitely the $1 million question. It's a toss task to do, to be honest. To be fair. Currently what we are doing is that we monitor every tasks of the process, so we have the output of every tasks. On each tasks we use a scoring system to evaluate if it's relevant to the initial question or the initial task of the user. And we have a global scoring system on all the system. So it's quite od, it's a little bit empiric, but it works for now. And it really help us to also improve over time all the tasks and all the processes that are done by the tool.
|
||||
That's definitely the $1 million question. It's a toss task to do, to be honest. To be fair. Currently what we are doing is that we monitor every tasks of the process, so we have the output of every tasks. On each tasks we use a scoring system to evaluate if it's relevant to the initial question or the initial task of the user. And we have a global scoring system on all the system. So it's quite odd, it's a little bit empiric, but it works for now. And it really help us to also improve over time all the tasks and all the processes that are done by the tool.
|
||||
|
||||
Guillaume Marquis:
|
||||
So it's really important. And for instance, you have this kind of framework that is called ragtriad. That is a way to evaluate rag on the accuracy of the context you retrieve on the link with the initial question and so on, several parameters. And you can really have a first way to evaluate the quality of answers and the quality of everything on each steps.
|
||||
So it's really important. And for instance, you have this kind of framework that is called RAGtriad. That is a way to evaluate rag on the accuracy of the context you retrieve on the link with the initial question and so on, several parameters. And you can really have a first way to evaluate the quality of answers and the quality of everything on each steps.
|
||||
|
||||
Sabrina Aquino:
|
||||
I love it. Can you go more into the tech that you use for each one of these steps in architecture?
|
||||
@@ -158,7 +158,7 @@ Guillaume Marquis:
|
||||
So basically we are creating unbelieving, we are storing it into Qdrant. We are performing similarity search to retrieve documents based on title summary filtering, on tags, on the semantic context. And we have also some keyword search, but it's more for specific tasks, like when we know that we need a specific document, at some point we are searching it with a keyword search. So it's like a kind of ebrid system that is using deterministic approach with filtering with tags, and a probabilistic approach with selecting document with this ebot search, and doing a scoring system after that to get what is the most relevant document and to select how much content we will take from each document. It's a little bit techy, but it's really cool to create and we have a way to evolve it and to improve it.
|
||||
|
||||
Demetrios:
|
||||
That's what we like around here, man. We want the techie stuff. That's what I think everybody signed up for. So that's very cool. One question that definitely comes up a lot when it comes to rags and when you're ingesting documents, and then when you're retrieving documents and updating documents, how do you make sure that the documents that you are, let's say, I know there's probably a hypothetical HR scenario where the company has a certain policy and they say you can have european style holidays, you get like three months of holidays a year, or even french style holidays. Basically, you just don't work. And whenever you want, you can work, you don't work. And then all of a sudden a US company comes and takes it over and they say, no, you guys don't get holidays.
|
||||
That's what we like around here, man. We want the techie stuff. That's what I think everybody signed up for. So that's very cool. One question that definitely comes up a lot when it comes to rags and when you're ingesting documents, and then when you're retrieving documents and updating documents, how do you make sure that the documents that you are, let's say, I know there's probably a hypothetical HR scenario where the company has a certain policy and they say you can have European style holidays, you get like three months of holidays a year, or even French style holidays. Basically, you just don't work. And whenever you want, you can work, you don't work. And then all of a sudden a US company comes and takes it over and they say, no, you guys don't get holidays.
|
||||
|
||||
Demetrios:
|
||||
Even when you do get holidays, you're not working or you are working and so you have to update all the HR documents, right? So now when you have this knowledge worker that is creating something, or when you have anyone that is getting help, like this copilot help, how do you make sure that the information that person is getting is the most up to date information possible?
|
||||
@@ -170,7 +170,7 @@ Demetrios:
|
||||
I'm coming with the hits today. I don't know what you were looking for.
|
||||
|
||||
Guillaume Marquis:
|
||||
That's a really good question. So basically you have several possibilities on that. First one you have like this PowerPoint presentation, v one, v two, vf vf one, vf two, et cetera. That's a mess in the knowledge bases and sometimes you just want to use the most updated up to date documents. So basically we can filter on the created ad and the date of the documents. Sometimes you want to also compare the evolution of the process over time. So that's another use case. Basically we base.
|
||||
That's a really good question. So basically you have several possibilities on that. First one you have like this PowerPoint presentation. That's a mess in the knowledge bases and sometimes you just want to use the most updated up to date documents. So basically we can filter on the created ad and the date of the documents. Sometimes you want to also compare the evolution of the process over time. So that's another use case. Basically we base.
|
||||
|
||||
Guillaume Marquis:
|
||||
So during the ingestion we are analyzing if date is inside the document, because sometimes in documentation you have like the date at the end of the document or at the beginning of the document. That's a first way to do it. We have the date of the creation of the document, but it's not a source of truth because sometimes you created it after or you duplicated it and the date is not the same, depending if you are working on Windows, Microsoft, stuff like that. It's definitely a mess. And also we compare documents. So when we retry the documents and documents are really similar one to each other, we keep it in mind and we try to give more information as possible. Sometimes it's not possible, so it's not 100%, it's not bulletproof, but it's a real question of that. So it's a partial answer of your question, but it's like some way we are today filtering and answering on this special topic.
|
||||
@@ -185,13 +185,13 @@ Sabrina Aquino:
|
||||
Challenging.
|
||||
|
||||
Guillaume Marquis:
|
||||
One of the challenging part was the scalability of the system. We have clients that come with terra octave of data and want to be parsed really fast and so you have the ingestion, but even after the semantic search, even on a large data set can be slow. And today chgptpt answer really fast. So your users, even if the question is way more complicated to answer than a basic Chat GPT question, they want to have their answer in seconds. So you have also this challenge that is really you have to take care. So it's quite challenging and it's like this industrial supply chain. So when you upgrade something, you have to be sure that everything is working well on the other side. And that's a real challenge to handle.
|
||||
One of the challenging part was the scalability of the system. We have clients that come with terra octave of data and want to be parsed really fast and so you have the ingestion, but even after the semantic search, even on a large data set can be slow. And today Chat GPT answer really fast. So your users, even if the question is way more complicated to answer than a basic Chat GPT question, they want to have their answer in seconds. So you have also this challenge that is really you have to take care. So it's quite challenging and it's like this industrial supply chain. So when you upgrade something, you have to be sure that everything is working well on the other side. And that's a real challenge to handle.
|
||||
|
||||
Guillaume Marquis:
|
||||
And we are still on it because we are still evolving and getting more data. And at the end of the day, you have to be sure that everything is working well in terms of LLM, but in terms of research and in terms also a few weeks to give some insight to the user of what is working under the hood, to give them the possibility to wait a few seconds more, but starting to give them pieces of answer.
|
||||
|
||||
Demetrios:
|
||||
Yeah, it's funny you say that because I remember talking to somebody that was working@u.com and they were saying how there's like the actual time. So they were calling it something like perceived time and real, like actual time. So you as an end user, if you get asked a question or maybe there's like a trivia quiz while the question is coming up, then it seems like it's not actually taking as long as it is. Even if it takes 5 seconds, it's a little bit cooler. Or as you were mentioning, I remember reading some paper, I think, on how people are a lot less anxious if they see the words starting to pop up like that and they see like, okay, it's not just I'm waiting and then the whole answer gets spit back out at me. It's like I see the answer forming as it is in real time. And so that can calm people's nerves too.
|
||||
Yeah, it's funny you say that because I remember talking to somebody that was working at you.com and they were saying how there's like the actual time. So they were calling it something like perceived time and real, like actual time. So you as an end user, if you get asked a question or maybe there's like a trivia quiz while the question is coming up, then it seems like it's not actually taking as long as it is. Even if it takes 5 seconds, it's a little bit cooler. Or as you were mentioning, I remember reading some paper, I think, on how people are a lot less anxious if they see the words starting to pop up like that and they see like, okay, it's not just I'm waiting and then the whole answer gets spit back out at me. It's like I see the answer forming as it is in real time. And so that can calm people's nerves too.
|
||||
|
||||
Guillaume Marquis:
|
||||
Yeah, definitely. Human's brain is like marvelous on that. And you have a lot of stuff. Like, one of my favorites is the illusion of work. Do you know it? It's the total opposite. If you have something that seems difficult to do, adding more time of processing. So the user will imagine that it's really an OD task to do. And so that's really funny.
|
||||
@@ -200,7 +200,7 @@ Demetrios:
|
||||
So funny like that.
|
||||
|
||||
Guillaume Marquis:
|
||||
Yeah. Yes. It's the opposite of what you will think if you create a product, but that's real stuff. And sometimes just to output them that you are performing toss tasks in the background, it helps them to. Oh, yes. My question was really like a complex question, like you have a lot of work to do. It's Axx word like. If you answer too fast, they will not trust the answer.
|
||||
Yeah. Yes. It's the opposite of what you will think if you create a product, but that's real stuff. And sometimes just to output them that you are performing toss tasks in the background, it helps them to. Oh, yes. My question was really like a complex question, like you have a lot of work to do. It's Axe word like. If you answer too fast, they will not trust the answer.
|
||||
|
||||
Guillaume Marquis:
|
||||
And it's the opposite if you answer too slow. You can have this. Okay. But it should be dumb because it's really slow. So it's a dumb AI or stuff like that. So that's really funny. My co founder actually was a product guy, so really focused on product, and he really loves this kind of stuff.
|
||||
@@ -218,7 +218,7 @@ Demetrios:
|
||||
Some tell me more. Yeah.
|
||||
|
||||
Guillaume Marquis:
|
||||
So we tried the classic postgre page vectors, that is, I think we tried it like 30 minutes, and we realized really fast that it was really not good for our use case. We tried wavyt, we tried Milvus, we tried Qdrant, we tried a lot. We prefer use open source because of security issues. We tried Pinecone initially, we were on Pinecone at the beginning of the company. And so the most important point, so we have the speed of the tool, we have the scalability we have also, maybe it's a little bit dumb to say that, but we have also the API. I remember using Pinecone and trying just to get all vectors and it was not possible somehow, and you have this dumb stuff that are sometimes really strange. And if you have a tool that is 100% made for your use case with people that are working on it, really dedicated on that, and that are aligned with your vision of what is the evolution of this. I think it's like the best tool you have to choose.
|
||||
So we tried the classic postgres page vectors, that is, I think we tried it like 30 minutes, and we realized really fast that it was really not good for our use case. We tried Weaviate, we tried Milvus, we tried Qdrant, we tried a lot. We prefer use open source because of security issues. We tried Pinecone initially, we were on Pinecone at the beginning of the company. And so the most important point, so we have the speed of the tool, we have the scalability we have also, maybe it's a little bit dumb to say that, but we have also the API. I remember using Pinecone and trying just to get all vectors and it was not possible somehow, and you have this dumb stuff that are sometimes really strange. And if you have a tool that is 100% made for your use case with people that are working on it, really dedicated on that, and that are aligned with your vision of what is the evolution of this. I think it's like the best tool you have to choose.
|
||||
|
||||
Demetrios:
|
||||
So one thing that I would love to hear about too, is when you're looking at your system and you're looking at just the product in general, what are some of the key metrics that you are constantly monitoring, and how do you know that you're hitting them or you're not? And then if you're not hitting them, what are some ways that you debug the situation?
|
||||
@@ -290,7 +290,7 @@ Sabrina Aquino:
|
||||
Yeah, I think I'm just very interesting to know from a user perspective, from a virtual brain, how are traditional models worse or what kind of errors virtual brain fixes in their structure, that users find it better that way.
|
||||
|
||||
Guillaume Marquis:
|
||||
I think in this particular, so we talked about hallucinations, I think it's like one of the main issues people have on classic elements. We really think that when you create a one size fit all tool, you have some chole because you have to manage different approaches, like when you are creating copilot as Microsoft, you have to under the use cases of, and I really think so. Our AI is not trained to write you a speech based on Shakespeare and with the style of Martin Luther Luther King. It's not the purpose of the tool. So if you ask something that is out of the box, he will just say like, okay, I don't know how to answer that. And that's an important point. That's a feature by itself to be able to not go outside of the box. And so we did this choice of putting the AI inside the box, the box that is containing basically all the knowledge of your company, all the retrieved knowledge.
|
||||
I think in this particular, so we talked about hallucinations, I think it's like one of the main issues people have on classic elements. We really think that when you create a one size fit all tool, you have some chole because you have to manage different approaches, like when you are creating copilot as Microsoft, you have to under the use cases of, and I really think so. Our AI is not trained to write you a speech based on Shakespeare and with the style of Martin Luther King. It's not the purpose of the tool. So if you ask something that is out of the box, he will just say like, okay, I don't know how to answer that. And that's an important point. That's a feature by itself to be able to not go outside of the box. And so we did this choice of putting the AI inside the box, the box that is containing basically all the knowledge of your company, all the retrieved knowledge.
|
||||
|
||||
Guillaume Marquis:
|
||||
Actually we do not have a lot of hallucination, I will not say like 0%, but it's close to zero. Because we analyze a question, we put the AI in a box, we enforce the AI to think about the answer before answering, and we analyze also the answer to know if the answer is relevant. And that's an important point that we are fixing and we fix for our user and we prefer yes, to give like non answers and a bad answer.
|
||||
|
||||
@@ -0,0 +1,188 @@
|
||||
---
|
||||
title: Nvidia
|
||||
weight: 1200
|
||||
---
|
||||
|
||||
# Nvidia
|
||||
|
||||
Qdrant supports working with [Nvidia embeddings](https://build.nvidia.com/explore/retrieval).
|
||||
|
||||
You can generate an API key to authenticate the requests from the [Nvidia Playground](<https://build.nvidia.com/nvidia/embed-qa-4>).
|
||||
|
||||
### Setting up the Qdrant client and Nvidia session
|
||||
|
||||
```python
|
||||
import requests
|
||||
from qdrant_client import QdrantClient
|
||||
|
||||
NVIDIA_BASE_URL = "https://ai.api.nvidia.com/v1/retrieval/nvidia/embeddings"
|
||||
|
||||
NVIDIA_API_KEY = "<YOUR_API_KEY>"
|
||||
|
||||
nvidia_session = requests.Session()
|
||||
|
||||
qdrant_client = QdrantClient(":memory:")
|
||||
|
||||
headers = {
|
||||
"Authorization": f"Bearer {NVIDIA_API_KEY}",
|
||||
"Accept": "application/json",
|
||||
}
|
||||
|
||||
texts = [
|
||||
"Qdrant is the best vector search engine!",
|
||||
"Loved by Enterprises and everyone building for low latency, high performance, and scale.",
|
||||
]
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from '@qdrant/js-client-rest';
|
||||
|
||||
const NVIDIA_BASE_URL = "https://ai.api.nvidia.com/v1/retrieval/nvidia/embeddings"
|
||||
const NVIDIA_API_KEY = "<YOUR_API_KEY>"
|
||||
|
||||
const client = new QdrantClient({ url: 'http://localhost:6333' });
|
||||
|
||||
const headers = {
|
||||
"Authorization": "Bearer " + NVIDIA_API_KEY,
|
||||
"Accept": "application/json",
|
||||
"Content-Type": "application/json"
|
||||
}
|
||||
|
||||
const texts = [
|
||||
"Qdrant is the best vector search engine!",
|
||||
"Loved by Enterprises and everyone building for low latency, high performance, and scale.",
|
||||
]
|
||||
```
|
||||
|
||||
The following example shows how to embed documents with the `embed-qa-4` model that generates sentence embeddings of size 1024.
|
||||
|
||||
### Embedding documents
|
||||
|
||||
```python
|
||||
payload = {
|
||||
"input": texts,
|
||||
"input_type": "passage",
|
||||
"model": "NV-Embed-QA",
|
||||
}
|
||||
|
||||
response_body = nvidia_session.post(
|
||||
NVIDIA_BASE_URL, headers=headers, json=payload
|
||||
).json()
|
||||
```
|
||||
|
||||
```typescript
|
||||
let body = {
|
||||
"input": texts,
|
||||
"input_type": "passage",
|
||||
"model": "NV-Embed-QA"
|
||||
}
|
||||
|
||||
let response = await fetch(NVIDIA_BASE_URL, {
|
||||
method: "POST",
|
||||
body: JSON.stringify(body),
|
||||
headers
|
||||
});
|
||||
|
||||
let response_body = await response.json()
|
||||
```
|
||||
|
||||
### Converting the model outputs to Qdrant points
|
||||
|
||||
```python
|
||||
from qdrant_client.http.models import PointStruct
|
||||
|
||||
points = [
|
||||
PointStruct(
|
||||
id=idx,
|
||||
vector=data["embedding"],
|
||||
payload={"text": text},
|
||||
)
|
||||
for idx, (data, text) in enumerate(zip(response_body["data"], texts))
|
||||
]
|
||||
```
|
||||
|
||||
```typescript
|
||||
let points = response_body.data.map((data, i) => {
|
||||
return {
|
||||
id: i,
|
||||
vector: data.embedding,
|
||||
payload: {
|
||||
text: texts[i]
|
||||
}
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
### Creating a collection to insert the documents
|
||||
|
||||
```python
|
||||
from qdrant_client.models import VectorParams, Distance
|
||||
|
||||
collection_name = "example_collection"
|
||||
|
||||
qdrant_client.create_collection(
|
||||
collection_name,
|
||||
vectors_config=VectorParams(
|
||||
size=1024,
|
||||
distance=Distance.COSINE,
|
||||
),
|
||||
)
|
||||
qdrant_client.upsert(collection_name, points)
|
||||
```
|
||||
|
||||
```typescript
|
||||
const COLLECTION_NAME = "example_collection"
|
||||
|
||||
await client.createCollection(COLLECTION_NAME, {
|
||||
vectors: {
|
||||
size: 1024,
|
||||
distance: 'Cosine',
|
||||
}
|
||||
});
|
||||
|
||||
await client.upsert(COLLECTION_NAME, {
|
||||
wait: true,
|
||||
points
|
||||
})
|
||||
```
|
||||
|
||||
## Searching for documents with Qdrant
|
||||
|
||||
Once the documents are added, you can search for the most relevant documents.
|
||||
|
||||
```python
|
||||
payload = {
|
||||
"input": "What is the best to use for vector search scaling?",
|
||||
"input_type": "query",
|
||||
"model": "NV-Embed-QA",
|
||||
}
|
||||
|
||||
response_body = nvidia_session.post(
|
||||
NVIDIA_BASE_URL, headers=headers, json=payload
|
||||
).json()
|
||||
|
||||
qdrant_client.search(
|
||||
collection_name=collection_name,
|
||||
query_vector=response_body["data"][0]["embedding"],
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
body = {
|
||||
"input": "What is the best to use for vector search scaling?",
|
||||
"input_type": "query",
|
||||
"model": "NV-Embed-QA",
|
||||
}
|
||||
|
||||
response = await fetch(NVIDIA_BASE_URL, {
|
||||
method: "POST",
|
||||
body: JSON.stringify(body),
|
||||
headers
|
||||
});
|
||||
|
||||
response_body = await response.json()
|
||||
|
||||
await client.search(COLLECTION_NAME, {
|
||||
vector: response_body.data[0].embedding,
|
||||
});
|
||||
```
|
||||
@@ -0,0 +1,94 @@
|
||||
---
|
||||
title: Pandas-AI
|
||||
weight: 2900
|
||||
---
|
||||
|
||||
# Pandas-AI
|
||||
|
||||
Pandas-AI is a Python library that uses a generative AI model to interpret natural language queries and translate them into Python code to interact with pandas data frames and return the final results to the user.
|
||||
|
||||
## Installation
|
||||
|
||||
```console
|
||||
pip install pandasai[qdrant]
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
You can begin a conversation by instantiating an `Agent` instance based on your Pandas data frame. The default Pandas-AI LLM requires an [API key](https://pandabi.ai.).
|
||||
|
||||
You can find the list of all supported LLMs [here](https://docs.pandas-ai.com/en/latest/LLMs/llms/)
|
||||
|
||||
```python
|
||||
import os
|
||||
import pandas as pd
|
||||
from pandasai import Agent
|
||||
|
||||
# Sample DataFrame
|
||||
sales_by_country = pd.DataFrame(
|
||||
{
|
||||
"country": [
|
||||
"United States",
|
||||
"United Kingdom",
|
||||
"France",
|
||||
"Germany",
|
||||
"Italy",
|
||||
"Spain",
|
||||
"Canada",
|
||||
"Australia",
|
||||
"Japan",
|
||||
"China",
|
||||
],
|
||||
"sales": [5000, 3200, 2900, 4100, 2300, 2100, 2500, 2600, 4500, 7000],
|
||||
}
|
||||
)
|
||||
|
||||
os.environ["PANDASAI_API_KEY"] = "YOUR_API_KEY"
|
||||
|
||||
agent = Agent(sales_by_country)
|
||||
agent.chat("Which are the top 5 countries by sales?")
|
||||
# OUTPUT: China, United States, Japan, Germany, Australia
|
||||
```
|
||||
|
||||
## Qdrant support
|
||||
|
||||
You can train Pandas-AI to understand your data better and improve the quality of the results.
|
||||
|
||||
Qdrant can be configured as a vector store to ingest training data and retrieve semantically relevant content.
|
||||
|
||||
```python
|
||||
from pandasai.ee.vectorstores.qdrant import Qdrant
|
||||
|
||||
qdrant = Qdrant(
|
||||
collection_name="<SOME_COLLECTION>",
|
||||
embedding_model="sentence-transformers/all-MiniLM-L6-v2",
|
||||
location="http://localhost:6334",
|
||||
prefer_grpc=True
|
||||
)
|
||||
|
||||
agent = Agent(df, vector_store=qdrant)
|
||||
|
||||
# Train with custom information
|
||||
agent.train(docs="The fiscal year starts in April")
|
||||
|
||||
# Train the q/a pairs of code snippets
|
||||
query = "What are the total sales for the current fiscal year?"
|
||||
response = """
|
||||
import pandas as pd
|
||||
|
||||
df = dfs[0]
|
||||
|
||||
# Calculate the total sales for the current fiscal year
|
||||
total_sales = df[df['date'] >= pd.to_datetime('today').replace(month=4, day=1)]['sales'].sum()
|
||||
result = { "type": "number", "value": total_sales }
|
||||
"""
|
||||
agent.train(queries=[query], codes=[response])
|
||||
|
||||
# # The model will use the information provided in the training to generate a response
|
||||
|
||||
```
|
||||
|
||||
## Further reading
|
||||
|
||||
- [Getting Started with Pandas-AI](https://pandasai-docs.readthedocs.io/en/latest/getting-started/)
|
||||
- [Pandas-AI Reference](https://pandasai-docs.readthedocs.io/en/latest/)
|
||||
@@ -0,0 +1,105 @@
|
||||
---
|
||||
title: Semantic-Router
|
||||
weight: 2700
|
||||
---
|
||||
|
||||
# Semantic-Router
|
||||
|
||||
[Semantic-Router](https://www.aurelio.ai/semantic-router/) is a library to build decision-making layers for your LLMs and agents. It uses vector embeddings to make tool-use decisions rather than LLM generations, routing our requests using semantic meaning.
|
||||
|
||||
Qdrant is available as a supported index in Semantic-Router for you to ingest route data and perform retrievals.
|
||||
|
||||
## Installation
|
||||
|
||||
To use Semantic-Router with Qdrant, install the `qdrant` extra:
|
||||
|
||||
```console
|
||||
pip install semantic-router[qdrant]
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Set up `QdrantIndex` with the appropriate configurations:
|
||||
|
||||
```python
|
||||
from semantic_router.index import QdrantIndex
|
||||
|
||||
qdrant_index = QdrantIndex(
|
||||
url="https://xyz-example.eu-central.aws.cloud.qdrant.io", api_key="<your-api-key>"
|
||||
)
|
||||
```
|
||||
|
||||
Once the Qdrant index is set up with the appropriate configurations, we can pass it to the `RouteLayer`.
|
||||
|
||||
```python
|
||||
from semantic_router.layer import RouteLayer
|
||||
|
||||
RouteLayer(encoder=some_encoder, routes=some_routes, index=qdrant_index)
|
||||
```
|
||||
|
||||
## Complete Example
|
||||
|
||||
<details>
|
||||
|
||||
<summary><b>Click to expand</b></summary>
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
from semantic_router import Route
|
||||
from semantic_router.encoders import OpenAIEncoder
|
||||
from semantic_router.index import QdrantIndex
|
||||
from semantic_router.layer import RouteLayer
|
||||
|
||||
# we could use this as a guide for our chatbot to avoid political conversations
|
||||
politics = Route(
|
||||
name="politics value",
|
||||
utterances=[
|
||||
"isn't politics the best thing ever",
|
||||
"why don't you tell me about your political opinions",
|
||||
"don't you just love the president",
|
||||
"they're going to destroy this country!",
|
||||
"they will save the country!",
|
||||
],
|
||||
)
|
||||
|
||||
# this could be used as an indicator to our chatbot to switch to a more
|
||||
# conversational prompt
|
||||
chitchat = Route(
|
||||
name="chitchat",
|
||||
utterances=[
|
||||
"how's the weather today?",
|
||||
"how are things going?",
|
||||
"lovely weather today",
|
||||
"the weather is horrendous",
|
||||
"let's go to the chippy",
|
||||
],
|
||||
)
|
||||
|
||||
# we place both of our decisions together into single list
|
||||
routes = [politics, chitchat]
|
||||
|
||||
os.environ["OPENAI_API_KEY"] = "<YOUR_API_KEY>"
|
||||
encoder = OpenAIEncoder()
|
||||
|
||||
rl = RouteLayer(
|
||||
encoder=encoder,
|
||||
routes=routes,
|
||||
index=QdrantIndex(location=":memory:"),
|
||||
)
|
||||
|
||||
print(rl("What have you been upto?").name)
|
||||
```
|
||||
|
||||
This returns:
|
||||
|
||||
```console
|
||||
[Out]: 'chitchat'
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
## 📚 Further Reading
|
||||
|
||||
- Semantic-Router [Documentation](https://github.com/aurelio-labs/semantic-router/tree/main/docs)
|
||||
- Semantic-Router [Video Course](https://www.aurelio.ai/course/semantic-router)
|
||||
@@ -244,5 +244,6 @@ Qdrant supports all the Spark data types, and the appropriate data types are map
|
||||
| `sparse_vector_index_fields` | Comma-separated names of columns holding the sparse vector indices. | `ArrayType(IntegerType)` | ❌ |
|
||||
| `sparse_vector_value_fields` | Comma-separated names of columns holding the sparse vector values. | `ArrayType(FloatType)` | ❌ |
|
||||
| `sparse_vector_names` | Comma-separated names of the sparse vectors in the collection. | - | ❌ |
|
||||
| `shard_key_selector` | Comma-separated names of custom shard keys to use during upsert. | - | ❌ |
|
||||
|
||||
For more information, be sure to check out the [Qdrant-Spark GitHub repository](https://github.com/qdrant/qdrant-spark). The Apache Spark guide is available [here](https://spark.apache.org/docs/latest/quick-start.html). Happy data processing!
|
||||
|
||||
@@ -12,20 +12,21 @@ aliases:
|
||||
|
||||
These tutorials demonstrate different ways you can build vector search into your applications.
|
||||
|
||||
| Tutorial | Description | Stack |
|
||||
|--------------------------------------------------------------------------------|------------------------------------------------------------------------------------------|--------------------------------------------------|
|
||||
| [Configure Optimal Use](../tutorials/optimize/) | Configure Qdrant collections for best resource use. | Qdrant |
|
||||
| [Separate Partitions](../tutorials/multiple-partitions/) | Serve vectors for many independent users. | Qdrant |
|
||||
| [Bulk Upload Vectors](../tutorials/bulk-upload/) | Upload a large scale dataset. | Qdrant |
|
||||
| [Create Dataset Snapshots](../tutorials/create-snapshot/) | Turn a dataset into a snapshot by exporting it from a collection. | Qdrant |
|
||||
| [Semantic Search for Beginners](../tutorials/search-beginners/) | Create a simple search engine locally in minutes. | Qdrant |
|
||||
| [Simple Neural Search](../tutorials/neural-search/) | Build and deploy a neural search that browses startup data. | Qdrant, BERT, FastAPI |
|
||||
| [Aleph Alpha Search](../tutorials/aleph-alpha-search/) | Build a multimodal search that combines text and image data. | Qdrant, Aleph Alpha |
|
||||
| [Mighty Semantic Search](../tutorials/mighty/) | Build a simple semantic search with an on-demand NLP service. | Qdrant, Mighty |
|
||||
| [Asynchronous API](../tutorials/async-api/) | Communicate with Qdrant server asynchronously with Python SDK. | Qdrant, Python |
|
||||
| [Multitenancy with LlamaIndex](../tutorials/llama-index-multitenancy/) | Handle data coming from multiple users in LlamaIndex. | Qdrant, Python, LlamaIndex |
|
||||
| [HuggingFace datasets](../tutorials/huggingface-datasets/) | Load a Hugging Face dataset to Qdrant | Qdrant, Python, datasets |
|
||||
| [Measure retrieval quality](../tutorials/retrieval-quality/) | Measure and fine-tune the retrieval quality | Qdrant, Python, datasets |
|
||||
| [Use semantic search to navigate your codebase](../tutorials/code-search/) | Implement semantic search application for code search task | Qdrant, Python, sentence-transformers, Jina |
|
||||
| [Automate customer support](../tutorials/customer-support-oci-cohere-airbyte/) | Unleash your customer support team from answering the same questions over and over again | Qdrant, Cohere, Command-R, Oracle Cloud, Airbyte |
|
||||
| [Troubleshooting](../tutorials/common-errors/) | Solutions to common errors and fixes | Qdrant |
|
||||
| Tutorial | Description | Stack |
|
||||
|---------------------------------------------------------------------------------|------------------------------------------------------------------------------------------|--------------------------------------------------|
|
||||
| [Configure Optimal Use](../tutorials/optimize/) | Configure Qdrant collections for best resource use. | Qdrant |
|
||||
| [Separate Partitions](../tutorials/multiple-partitions/) | Serve vectors for many independent users. | Qdrant |
|
||||
| [Bulk Upload Vectors](../tutorials/bulk-upload/) | Upload a large scale dataset. | Qdrant |
|
||||
| [Create Dataset Snapshots](../tutorials/create-snapshot/) | Turn a dataset into a snapshot by exporting it from a collection. | Qdrant |
|
||||
| [Semantic Search for Beginners](../tutorials/search-beginners/) | Create a simple search engine locally in minutes. | Qdrant |
|
||||
| [Simple Neural Search](../tutorials/neural-search/) | Build and deploy a neural search that browses startup data. | Qdrant, BERT, FastAPI |
|
||||
| [Aleph Alpha Search](../tutorials/aleph-alpha-search/) | Build a multimodal search that combines text and image data. | Qdrant, Aleph Alpha |
|
||||
| [Mighty Semantic Search](../tutorials/mighty/) | Build a simple semantic search with an on-demand NLP service. | Qdrant, Mighty |
|
||||
| [Asynchronous API](../tutorials/async-api/) | Communicate with Qdrant server asynchronously with Python SDK. | Qdrant, Python |
|
||||
| [Multitenancy with LlamaIndex](../tutorials/llama-index-multitenancy/) | Handle data coming from multiple users in LlamaIndex. | Qdrant, Python, LlamaIndex |
|
||||
| [HuggingFace datasets](../tutorials/huggingface-datasets/) | Load a Hugging Face dataset to Qdrant | Qdrant, Python, datasets |
|
||||
| [Measure retrieval quality](../tutorials/retrieval-quality/) | Measure and fine-tune the retrieval quality | Qdrant, Python, datasets |
|
||||
| [Use semantic search to navigate your codebase](../tutorials/code-search/) | Implement semantic search application for code search task | Qdrant, Python, sentence-transformers, Jina |
|
||||
| [Implement custom connector for Cohere RAG](../tutorials/cohere-rag-connector/) | Bring data stored in Qdrant to Cohere RAG | Qdrant, Cohere, FastAPI |
|
||||
| [Automate customer support](../tutorials/customer-support-oci-cohere-airbyte/) | Unleash your customer support team from answering the same questions over and over again | Qdrant, Cohere, Command-R, Oracle Cloud, Airbyte |
|
||||
| [Troubleshooting](../tutorials/common-errors/) | Solutions to common errors and fixes | Qdrant |
|
||||
|
||||
@@ -0,0 +1,292 @@
|
||||
---
|
||||
title: Implement Cohere RAG connector
|
||||
weight: 24
|
||||
---
|
||||
|
||||
# Implement custom connector for Cohere RAG
|
||||
|
||||
| Time: 45 min | Level: Intermediate | | |
|
||||
|--------------|---------------------|-|----|
|
||||
|
||||
The usual approach to implementing Retrieval Augmented Generation requires users to build their prompts with the
|
||||
relevant context the LLM may rely on, and manually sending them to the model. Cohere is quite unique here, as their
|
||||
models can now speak to the external tools and extract meaningful data on their own. You can virtually connect any data
|
||||
source and let the Cohere LLM know how to access it. Obviously, vector search goes well with LLMs, and enabling semantic
|
||||
search over your data is a typical case.
|
||||
|
||||
The connectors have to implement a specific interface and expose the data source as HTTP REST API. Cohere documentation
|
||||
[describes a general process of creating a connector](https://docs.cohere.com/docs/creating-and-deploying-a-connector).
|
||||
This tutorial guides you step by step on building such a service around Qdrant.
|
||||
|
||||
## Qdrant connector
|
||||
|
||||
You probably already have some collections you would like to bring to the LLM. Maybe your pipeline was set up using some
|
||||
of the popular libraries such as Langchain, Llama Index, or Haystack. Cohere connectors may implement even more complex
|
||||
logic, e.g. hybrid search. In our case, we are going to start with a fresh Qdrant collection, index data using Cohere
|
||||
Embed v3, build the connector, and finally connect it with the [Command-R model](https://txt.cohere.com/command-r/).
|
||||
|
||||
### Building the collection
|
||||
|
||||
First things first, let's build a collection and configure it for the Cohere `embed-multilingual-v3.0` model. It
|
||||
produces 1024-dimensional embeddings, and we can choose any of the distance metrics available in Qdrant. Our connector
|
||||
will act as a personal assistant of a software engineer, and it will expose our notes to suggest the priorities or
|
||||
actions to perform.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
qdrant_client = QdrantClient(
|
||||
"https://my-cluster.cloud.qdrant.io:6333",
|
||||
api_key="my-api-key",
|
||||
)
|
||||
qdrant_client.create_collection(
|
||||
collection_name="personal-notes",
|
||||
vectors_config=models.VectorParams(
|
||||
size=1024,
|
||||
distance=models.Distance.DOT,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
Our notes will be represented as simple JSON objects with a `title` and `text` of the specific note. The embeddings will
|
||||
be created from the `text` field only.
|
||||
|
||||
```python
|
||||
notes = [
|
||||
{
|
||||
"title": "Project Alpha Review",
|
||||
"text": "Review the current progress of Project Alpha, focusing on the integration of the new API. Check for any compatibility issues with the existing system and document the steps needed to resolve them. Schedule a meeting with the development team to discuss the timeline and any potential roadblocks."
|
||||
},
|
||||
{
|
||||
"title": "Learning Path Update",
|
||||
"text": "Update the learning path document with the latest courses on React and Node.js from Pluralsight. Schedule at least 2 hours weekly to dedicate to these courses. Aim to complete the React course by the end of the month and the Node.js course by mid-next month."
|
||||
},
|
||||
{
|
||||
"title": "Weekly Team Meeting Agenda",
|
||||
"text": "Prepare the agenda for the weekly team meeting. Include the following topics: project updates, review of the sprint backlog, discussion on the new feature requests, and a brainstorming session for improving remote work practices. Send out the agenda and the Zoom link by Thursday afternoon."
|
||||
},
|
||||
{
|
||||
"title": "Code Review Process Improvement",
|
||||
"text": "Analyze the current code review process to identify inefficiencies. Consider adopting a new tool that integrates with our version control system. Explore options such as GitHub Actions for automating parts of the process. Draft a proposal with recommendations and share it with the team for feedback."
|
||||
},
|
||||
{
|
||||
"title": "Cloud Migration Strategy",
|
||||
"text": "Draft a plan for migrating our current on-premise infrastructure to the cloud. The plan should cover the selection of a cloud provider, cost analysis, and a phased migration approach. Identify critical applications for the first phase and any potential risks or challenges. Schedule a meeting with the IT department to discuss the plan."
|
||||
},
|
||||
{
|
||||
"title": "Quarterly Goals Review",
|
||||
"text": "Review the progress towards the quarterly goals. Update the documentation to reflect any completed objectives and outline steps for any remaining goals. Schedule individual meetings with team members to discuss their contributions and any support they might need to achieve their targets."
|
||||
},
|
||||
{
|
||||
"title": "Personal Development Plan",
|
||||
"text": "Reflect on the past quarter's achievements and areas for improvement. Update the personal development plan to include new technical skills to learn, certifications to pursue, and networking events to attend. Set realistic timelines and check-in points to monitor progress."
|
||||
},
|
||||
{
|
||||
"title": "End-of-Year Performance Reviews",
|
||||
"text": "Start preparing for the end-of-year performance reviews. Collect feedback from peers and managers, review project contributions, and document achievements. Consider areas for improvement and set goals for the next year. Schedule preliminary discussions with each team member to gather their self-assessments."
|
||||
},
|
||||
{
|
||||
"title": "Technology Stack Evaluation",
|
||||
"text": "Conduct an evaluation of our current technology stack to identify any outdated technologies or tools that could be replaced for better performance and productivity. Research emerging technologies that might benefit our projects. Prepare a report with findings and recommendations to present to the management team."
|
||||
},
|
||||
{
|
||||
"title": "Team Building Event Planning",
|
||||
"text": "Plan a team-building event for the next quarter. Consider activities that can be done remotely, such as virtual escape rooms or online game nights. Survey the team for their preferences and availability. Draft a budget proposal for the event and submit it for approval."
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
Storing the embeddings along with the metadata is fairly simple.
|
||||
|
||||
```python
|
||||
import cohere
|
||||
import uuid
|
||||
|
||||
cohere_client = cohere.Client(api_key="my-cohere-api-key")
|
||||
|
||||
response = cohere_client.embed(
|
||||
texts=[
|
||||
note.get("text")
|
||||
for note in notes
|
||||
],
|
||||
model="embed-multilingual-v3.0",
|
||||
input_type="search_document",
|
||||
)
|
||||
|
||||
qdrant_client.upload_points(
|
||||
collection_name="personal-notes",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=uuid.uuid4().hex,
|
||||
vector=embedding,
|
||||
payload=note,
|
||||
)
|
||||
for note, embedding in zip(notes, response.embeddings)
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
Our collection is now ready to be searched over. In the real world, the set of notes would be changing over time, so the
|
||||
ingestion process won't be as straightforward. This data is not yet exposed to the LLM, but we will build the connector
|
||||
in the next step.
|
||||
|
||||
### Connector web service
|
||||
|
||||
[FastAPI](https://fastapi.tiangolo.com/) is a modern web framework and perfect a choice for a simple HTTP API. We are
|
||||
going to use it for the purposes of our connector. There will be just one endpoint, as required by the model. It will
|
||||
accept POST requests at the `/search` path. There is a single `query` parameter required. Let's define a corresponding
|
||||
model.
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
|
||||
class SearchQuery(BaseModel):
|
||||
query: str
|
||||
```
|
||||
|
||||
RAG connector does not have to return the documents in any specific format. There are [some good practices to follow](https://docs.cohere.com/docs/creating-and-deploying-a-connector#configure-the-connection-between-the-connector-and-the-chat-api),
|
||||
but Cohere models are quite flexible here. Results just have to be returned as JSON, with a list of objects in a
|
||||
`results` property of the output. We will use the same document structure as we did for the Qdrant payloads, so there
|
||||
is no conversion required. That requires two additional models to be created.
|
||||
|
||||
```python
|
||||
from typing import List
|
||||
|
||||
class Document(BaseModel):
|
||||
title: str
|
||||
text: str
|
||||
|
||||
class SearchResults(BaseModel):
|
||||
results: List[Document]
|
||||
```
|
||||
|
||||
Once our model classes are ready, we can implement the logic that will get the query and provide the notes that are
|
||||
relevant to it. Please note the LLM is not going to define the number of documents to be returned. That's completely
|
||||
up to you how many of them you want to bring to the context.
|
||||
|
||||
There are two services we need to interact with - Qdrant server and Cohere API. FastAPI has a concept of a [dependency
|
||||
injection](https://fastapi.tiangolo.com/tutorial/dependencies/#dependencies), and we will use it to provide both
|
||||
clients into the implementation.
|
||||
|
||||
In case of queries, we need to set the `input_type` to `search_query` in the calls to Cohere API.
|
||||
|
||||
```python
|
||||
from fastapi import FastAPI, Depends
|
||||
from typing import Annotated
|
||||
|
||||
app = FastAPI()
|
||||
|
||||
def qdrant_client() -> QdrantClient:
|
||||
return QdrantClient(config.QDRANT_URL, api_key=config.QDRANT_API_KEY)
|
||||
|
||||
def cohere_client() -> cohere.Client:
|
||||
return cohere.Client(api_key=config.COHERE_API_KEY)
|
||||
|
||||
@app.post("/search")
|
||||
def search(
|
||||
query: SearchQuery,
|
||||
qdrant_client: Annotated[QdrantClient, Depends(qdrant_client)],
|
||||
cohere_client: Annotated[cohere.Client, Depends(cohere_client)],
|
||||
) -> SearchResults:
|
||||
response = cohere_client.embed(
|
||||
texts=[query.query],
|
||||
model="embed-multilingual-v3.0",
|
||||
input_type="search_query",
|
||||
)
|
||||
results = qdrant_client.search(
|
||||
collection_name="personal-notes",
|
||||
query_vector=response.embeddings[0],
|
||||
limit=2,
|
||||
)
|
||||
return SearchResults(
|
||||
results=[
|
||||
Document(**point.payload)
|
||||
for point in results
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
Our app might be launched locally for the development purposes, given we have the `uvicorn` server installed:
|
||||
|
||||
```shell
|
||||
uvicorn main:app
|
||||
```
|
||||
|
||||
We can interact with it and check the documents that will be returned for a specific query. For example, we want to know
|
||||
recall what we are supposed to do regarding the infrastructure for your projects.
|
||||
|
||||
```shell
|
||||
curl -X "POST" \
|
||||
-H "Content-type: application/json" \
|
||||
-d '{"query": "Is there anything I have to do regarding the project infrastructure?"}' \
|
||||
"http://localhost:8000/search"
|
||||
```
|
||||
|
||||
The output should look like following:
|
||||
|
||||
```json
|
||||
{
|
||||
"results": [
|
||||
{
|
||||
"title": "Cloud Migration Strategy",
|
||||
"text": "Draft a plan for migrating our current on-premise infrastructure to the cloud. The plan should cover the selection of a cloud provider, cost analysis, and a phased migration approach. Identify critical applications for the first phase and any potential risks or challenges. Schedule a meeting with the IT department to discuss the plan."
|
||||
},
|
||||
{
|
||||
"title": "Project Alpha Review",
|
||||
"text": "Review the current progress of Project Alpha, focusing on the integration of the new API. Check for any compatibility issues with the existing system and document the steps needed to resolve them. Schedule a meeting with the development team to discuss the timeline and any potential roadblocks."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### Connecting to Command-R
|
||||
|
||||
Our web service is implemented, yet running only on our local machine. It has to be exposed to the public before
|
||||
Command-R can interact with it. For a quick experiment, it might be enough to set up tunneling using services such as
|
||||
[ngrok](https://ngrok.com/). We won't cover all the details in the tutorial, but their
|
||||
[Quickstart](https://ngrok.com/docs/guides/getting-started/) is a great resource describing the process step-by-step.
|
||||
Alternatively, you can also deploy the service with a public URL.
|
||||
|
||||
Once it's done, we can create the connector first, and then tell the model to use it, while interacting through the chat
|
||||
API. Creating a connector is a single call to Cohere client:
|
||||
|
||||
```python
|
||||
connector_response = cohere_client.connectors.create(
|
||||
name="personal-notes",
|
||||
url="https:/this-is-my-domain.app/search",
|
||||
)
|
||||
```
|
||||
|
||||
The `connector_response.connector` will be a descriptor, with `id` being one of the attributes. We'll use this
|
||||
identifier for our interactions like this:
|
||||
|
||||
```python
|
||||
response = cohere_client.chat(
|
||||
message=(
|
||||
"Is there anything I have to do regarding the project infrastructure? "
|
||||
"Please mention the tasks briefly."
|
||||
),
|
||||
connectors=[
|
||||
cohere.ChatConnector(id=connector_response.connector.id)
|
||||
],
|
||||
model="command-r",
|
||||
)
|
||||
```
|
||||
|
||||
We changed the `model` to `command-r`, as this is currently the best Cohere model available to public. The
|
||||
`response.text` is the output of the model:
|
||||
|
||||
```text
|
||||
Here are some of the tasks related to project infrastructure that you might have to perform:
|
||||
- You need to draft a plan for migrating your on-premise infrastructure to the cloud and come up with a plan for the selection of a cloud provider, cost analysis, and a gradual migration approach.
|
||||
- It's important to evaluate your current technology stack to identify any outdated technologies. You should also research emerging technologies and the benefits they could bring to your projects.
|
||||
```
|
||||
|
||||
You only need to create a specific connector once! Please do not call `cohere_client.connectors.create` for every single
|
||||
message you send to the `chat` method.
|
||||
|
||||
## Wrapping up
|
||||
|
||||
We have built a Cohere RAG connector that integrates with your existing knowledge base stored in Qdrant. We covered just
|
||||
the basic flow, but in real world scenarios, you should also consider e.g. [building the authentication
|
||||
system](https://docs.cohere.com/docs/connector-authentication) to prevent unauthorized access.
|
||||
|
After Width: | Height: | Size: 89 KiB |
|
After Width: | Height: | Size: 58 KiB |
|
After Width: | Height: | Size: 89 KiB |
|
After Width: | Height: | Size: 58 KiB |
|
After Width: | Height: | Size: 7.0 KiB |
|
After Width: | Height: | Size: 4.2 KiB |
|
After Width: | Height: | Size: 44 KiB |
|
After Width: | Height: | Size: 89 KiB |
|
After Width: | Height: | Size: 30 KiB |
|
After Width: | Height: | Size: 14 KiB |
|
After Width: | Height: | Size: 26 KiB |
|
After Width: | Height: | Size: 21 KiB |
|
After Width: | Height: | Size: 158 KiB |
|
After Width: | Height: | Size: 104 KiB |
|
After Width: | Height: | Size: 74 KiB |
|
After Width: | Height: | Size: 23 KiB |
|
After Width: | Height: | Size: 17 KiB |
|
After Width: | Height: | Size: 147 KiB |
|
After Width: | Height: | Size: 93 KiB |
|
After Width: | Height: | Size: 65 KiB |
|
After Width: | Height: | Size: 24 KiB |
|
After Width: | Height: | Size: 19 KiB |
|
After Width: | Height: | Size: 148 KiB |
|
After Width: | Height: | Size: 96 KiB |
|
After Width: | Height: | Size: 68 KiB |
|
After Width: | Height: | Size: 20 KiB |
|
After Width: | Height: | Size: 16 KiB |
|
After Width: | Height: | Size: 128 KiB |
|
After Width: | Height: | Size: 85 KiB |
|
After Width: | Height: | Size: 57 KiB |
|
After Width: | Height: | Size: 649 KiB |
|
After Width: | Height: | Size: 611 KiB |
|
After Width: | Height: | Size: 604 KiB |
|
After Width: | Height: | Size: 612 KiB |